LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities,…
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities,…
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities,…
[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.
1. Model Summary 2. Limitations 3. Training 4. License 5. Citation
1. Model Summary 2. Limitations 3. Training 4. License 5. Citation
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities,…
We introduce the updated version of the Qwen3-4B-FP8 non-thinking mode , named Qwen3-4B-Instruct-2507-FP8 , featuring the following key enhancements:
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities,…
We introduce the updated version of the Qwen3-30B-A3B non-thinking mode , named Qwen3-30B-A3B-Instruct-2507 , featuring the following key enhancements:
Join Our 💬 WeChat 🧩 Discord community. MiniMax Agent ⚡️ API CLI MiniMax Website 🤗 Hugging Face 🐙 GitHub 🤖️ ModelScope 📄 LICENSE
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better