LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
py from mistral common.tokens.tokenizers.mistral import MistralTokenizer from mistral common.protocol.instruct.messages import UserMessage from mistral common.protocol.instruct.request import ChatCompletionRequest
Join Our 💬 WeChat 🧩 Discord community. MiniMax Agent ⚡️ API CLI MiniMax Website 🤗 Hugging Face 🐙 GitHub 🤖️ ModelScope 📄 LICENSE
We introduce the updated version of the Qwen3-4B-FP8 non-thinking mode , named Qwen3-4B-Instruct-2507-FP8 , featuring the following key enhancements:
Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct-FP8 . This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements:
DeepSeek-V3-0324 demonstrates notable improvements over its predecessor, DeepSeek-V3, in several key aspects.
We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, w…
🌟 Github 📥 Model Download 📄 Paper Link 📄 Arxiv Paper Link DeepSeek-OCR 2: Visual Causal Flow Explore more human-like visual encoding.
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
We introduce OLMo 2 1B, the smallest model in the OLMo 2 family. OLMo 2 was pre-trained on OLMo-mix-1124 and uses Dolmino-mix-1124 for mid-training.
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Building upon extensive advancements in training data, model architecture, and optimization techniques, Qwen3 delivers the followin…