LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
MiniCPM Tech Report MiniCPM Wiki(Chinese) GitHub Repo UltraData MiniCPM Desk Pet Online Demo
The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.
--- --- Developers Granite Team, IBM Model Type Decoder-only Dense Transformer (Reasoning) Architecture GraniteForCausalLM Base Model Granite-4.1-3B-Base Parameters 3B Context Length Natively Supports 128K (Long-context extension to 512K) Precision bfloat16 Tested Languages Engli…
MiniCPM Tech Report MiniCPM Wiki(Chinese) GitHub Repo UltraData MiniCPM Desk Pet Online Demo
MiniCPM Tech Report MiniCPM Wiki(Chinese) GitHub Repo UltraData MiniCPM Desk Pet Online Demo
- 🚀 Online Demo : Explore Step3-VL-10B on Hugging Face Spaces ! - 📢 [Notice] FP8 Quantization Support : FP8 quantized weights are now available. (Download link) - 📢 [Notice] vLLM Support: vLLM integration is now officially supported! (PR 32329) - ✅ [Fixed] HF Inference: Resolved…
- 🚀 Online Demo : Explore Step3-VL-10B on Hugging Face Spaces ! - 📢 [Notice] FP8 Quantization Support : FP8 quantized weights are now available. (Download link) - 📢 [Notice] vLLM Support: vLLM integration is now officially supported! (PR 32329) - ✅ [Fixed] HF Inference: Resolved…
STEP3-VL-10B is a lightweight open-source foundation model designed to redefine the trade-off between compact efficiency and frontier-level multimodal intelligence. Despite its compact 10B parameter footprint , STEP3-VL-10B excels in visual perception , complex reasoning , and hu…
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
This model is presented in the paper Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model.