LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
--- --- Developers Granite Team, IBM Model Type Decoder-only Dense Transformer (Reasoning) Architecture GraniteForCausalLM Base Model Granite-4.1-3B-Base Parameters 3B Context Length Natively Supports 128K (Long-context extension to 512K) Precision bfloat16 Tested Languages Engli…
Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces. It targets complex tasks across reasoning, code generation, knowledge QA, and multilingual...
Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic...
Hugging Face GitHub Launch Blog Documentation License : Apache 2.0 Authors : Google DeepMind
--- pipeline tag: text-generation base model: google/diffusiongemma-26B-A4B-it license: apache-2.0 license name: apache-license-2.0 license link: https://ai.google.dev/gemma/apache 2 tags: - nvidia - ModelOpt - DiffusionGemma-26B-A4B-IT - quantized - NVFP4 - nvfp4 ---
MiniCPM Tech Report MiniCPM Wiki(Chinese) GitHub Repo UltraData MiniCPM Desk Pet Online Demo
Llama-3.3-Nemotron-Super-49B-v1.5-FP8 is a significantly upgraded version of Llama-3.3-Nemotron-Super-49B-v1 and is a large language model (LLM) which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model). It is a reasoning model that is post trained for reason…
MiniCPM Tech Report MiniCPM Wiki(Chinese) GitHub Repo UltraData MiniCPM Desk Pet Online Demo
MiniCPM Tech Report MiniCPM Wiki(Chinese) GitHub Repo UltraData MiniCPM Desk Pet Online Demo
Sarvam-1 is a 2-billion parameter language model specifically optimized for Indian languages. It provides best in-class performance in 10 Indic languages (bn, gu, hi, kn, ml, mr, or, pa, ta, te) when compared with popular models like Gemma-2-2B and Llama-3.2-3B. It is also compet…
Want a smaller model? Download Sarvam-30B!
LFM2.5-2.6B is part of LFM2.5, a family of hybrid models designed for on-device deployment . It builds on the LFM2 architecture with a 128K context window and agentic post-training.