Inference.net: Schematron V2 Turbo
About this model
Schematron V2 Turbo is a 3-billion-parameter specialized language model from Inference.net, designed to convert noisy HTML documents into strictly schema-conforming JSON. Unlike general-purpose chat models, it is trained for long-context web extraction: callers supply raw or cleaned HTML together with a JSON Schema, Pydantic model, or equivalent structured output definition, and the model returns typed JSON via schema-constrained decoding so outputs are valid by construction rather than through post-hoc repair or retries.
The Turbo variant sits on the throughput and cost tier of the Schematron V2 family. Inference.net reports that on a single NVIDIA H100 with roughly 10k input tokens and 500 output tokens per request, V2 Turbo sustains about 4.14 requests per second, roughly 2.5x the first-generation 8B model and about 1.7x the original 3B, while still beating V1 3B on LLM-as-judge extraction quality (4.039 vs 3.909 on a 1–5 scale). It supports up to 128K tokens of context and up to 8K tokens of output, with list pricing around $0.03 per million input tokens and $0.15 per million output tokens on the hosted API.
V2 Turbo weights are closed and served only through Inference.net’s serverless API (model id inference-net/schematron-v2-turbo); legacy inference-net/schematron-3b calls route to this model. In Inference.net’s SimpleQA pipeline benchmark—measuring factual QA when Schematron acts as the extraction layer over Exa search with GPT-5 Nano—V2 Turbo scores 79.42%, reinforcing the use case of making web-augmented LLM pipelines cheaper by extracting compact structured facts instead of feeding full HTML into large general models.
Benchmark Scores
Technical Specs
- Parameters: 3.0B
- Architecture: Transformer
- Context Window: 128,000 tokens
- Input Modalities: text
Hardware Requirements
- Compute: API only
Pricing
| Input | Output | Currency |
|---|---|---|
| 0.03 / 1M tokens | 0.15 / 1M tokens | USD |