LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct-FP8 . This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements:
Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct . This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements:
Today, we're announcing Qwen3-Coder , our most agentic code model to date. Qwen3-Coder is available in multiple sizes, but we're excited to introduce its most powerful variant first: Qwen3-Coder-480B-A35B-Instruct . featuring the following key enhancements:
Over the past three months, we have continued to scale the thinking capability of Qwen3-4B, improving both the quality and depth of reasoning. We are pleased to introduce Qwen3-4B-Thinking-2507 , featuring the following key enhancements:
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities,…
--- license: apache-2.0 language: - en pipeline tag: image-text-to-text tags: - multimodal - gui library name: transformers ---
1. Model Summary 2. Limitations 3. Training 4. License 5. Citation
------------------------- ------------------------------------------------------------------------------- Developers Microsoft Research Description phi-4 is a state-of-the-art open model built upon a blend of synthetic datasets, data from filtered public domain websites, and acqu…
🎉 Phi-3.5 : [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)
🎉 Phi-4 : [multimodal-instruct onnx]; [mini-instruct onnx]
📰 Tech Blog 📄 Paper
We introduce OLMo 2 1B, the smallest model in the OLMo 2 family. OLMo 2 was pre-trained on OLMo-mix-1124 and uses Dolmino-mix-1124 for mid-training.