LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
🤗 Hugging Face 🤖 ModelScope 🐙 OpenRouter
🤗 Hugging Face 🤖 ModelScope 🐙 OpenRouter
Qwen2.5 Bakeneko 32B Instruct (rinna/qwen2.5-bakeneko-32b-instruct)
DeepSeek R1 Distill Qwen2.5 Bakeneko 32B (rinna/deepseek-r1-distill-qwen2.5-bakeneko-32b)
QwQ Bakeneko 32B (rinna/qwq-bakeneko-32b)
Ling 3.0 Flash Sante is a health and medicine-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for...
Description: The NVIDIA Qwen3.6-35B-A3B-NVFP4 model is the quantized version of Alibaba's Qwen3.6-35B-A3B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.6-35B-A3B-NVFP4 m…
Description: The NVIDIA Qwen3.5-122B-A10B-NVFP4 model is the quantized version of Alibaba's Qwen3.5-122B-A10B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.5-122B-A10B N…
Description: Gemma 4 26B IT is an open multimodal model built by Google DeepMind that handles text and image inputs, can process video as sequences of frames, and generates text output. It is designed to deliver frontier-level performance for reasoning, agentic workflows, coding,…
Description: Gemma 4 31B IT is an open multimodal model built by Google DeepMind that handles text and image inputs, can process video as sequences of frames, and generates text output. It is designed to deliver frontier-level performance for reasoning, agentic workflows, coding,…
--- --- Developers Granite Team, IBM Model Type Decoder-only Dense Transformer (Reasoning) Architecture GraniteForCausalLM Base Model Granite-4.1-30B-Base Parameters 30B Context Length Natively Supports 128K (Long-context extension to 512K) Precision bfloat16 Tested Languages Eng…
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4