LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
Today, we're announcing Qwen3-Coder-Next-FP8 , an open-weight language model designed specifically for coding agents and local development. It features the following key enhancements:
The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.
Today, we're announcing Qwen3-Coder-Next , an open-weight language model designed specifically for coding agents and local development. It features the following key enhancements:
The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.
--- --- Developers Granite Team, IBM Model Type Decoder-only Dense Transformer (Reasoning) Architecture GraniteForCausalLM Base Model Granite-4.1-3B-Base Parameters 3B Context Length Natively Supports 128K (Long-context extension to 512K) Precision bfloat16 Tested Languages Engli…
This is an EAGLE3 draft model trained from scratch (random initialization) using the Aurora inference-time training framework for speculative decoding. Unlike traditional approaches that fine-tune pre-trained models, this model is built entirely through Aurora's online training p…
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
LFM2.5-2.6B is part of LFM2.5, a family of hybrid models designed for on-device deployment . It builds on the LFM2 architecture with a 128K context window and agentic post-training.
🎉 License Updated! We are pleased to announce our more flexible licensing terms 🤗 ✈️ Try on FriendliAI (licensed under commercial purposes) 📢 EXAONE 4.0 is officially supported by HuggingFace transformers! Please check out the guide below
- 🚀 Online Demo : Explore Step3-VL-10B on Hugging Face Spaces ! - 📢 [Notice] FP8 Quantization Support : FP8 quantized weights are now available. (Download link) - 📢 [Notice] vLLM Support: vLLM integration is now officially supported! (PR 32329) - ✅ [Fixed] HF Inference: Resolved…
LFM2.5-DSpark is a family of speculative-decoding draft models that adapt DSpark for the LFM2.5 architecture. They allow LFM2.5 models to run faster without degrading quality.
LFM2.5 is a new family of hybrid models designed for on-device deployment . It builds on the LFM2 architecture with extended pre-training and reinforcement learning.