Inference.net: Schematron V2 Small
About this model
Inference.net Schematron V2 Small is a 3-billion-parameter specialist model fine-tuned from Meta Llama 3.2 3B Instruct for HTML-to-JSON extraction. It accepts raw or cleaned HTML together with a JSON Schema (via structured response_format) and returns schema-conformant JSON through constrained decoding, so outputs are valid JSON by construction rather than by prompt engineering or retries. The model targets web scraping, product catalog ingestion, financial document parsing, and other pipelines that must turn noisy, long web pages into typed records at high volume.
Within the Schematron V2 family, Small is the quality-oriented tier: Inference.net reports it nearly matches first-generation Schematron 8B on an internal extraction judge benchmark while keeping roughly 3B-class latency (~2.47 requests per second on their setup). It supports up to about 128K input tokens and up to 4K completion tokens, with API pricing designed for cost-efficient batch extraction compared with frontier general-purpose chat models.
Open weights are published on Hugging Face (inference-net/schematron-v2-llama-3.2-3b) for self-hosting, while the managed identifier inference-net/schematron-v2-small is available through Inference.net and OpenAI-compatible providers such as OpenRouter. On Inference.net's SimpleQA pipeline benchmark (web search plus structured extraction layered on a small base LLM), V2 Small scores 83.10, reflecting strong gains in factual retrieval workflows when used as a dedicated extraction layer.
Benchmark Scores
Technical Specs
- Parameters: 3.0B
- Architecture: Transformer
- Context Window: 128,000 tokens
- Input Modalities: text
Hardware Requirements
- VRAM: 8.0 GB
- Compute: Single NVIDIA GPU with 8GB+ VRAM (16GB+ recommended for long-context batches)
Pricing
| Input | Output | Currency |
|---|---|---|
| 0.05 / 1M tokens | 0.23 / 1M tokens | USD |