LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.
Description: The NVIDIA MiniMax-M2.5-NVFP4 model is the quantized version of MiniMax's MiniMax-M2.5 model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA MiniMax-M2.5 NVFP4 model is q…
Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct . This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements:
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
Phi-3.5-vision is a lightweight, state-of-the-art open multimodal model built upon datasets which include - synthetic data and filtered publicly available websites - with a focus on very high-quality, reasoning dense data both on text and vision. The model belongs to the Phi-3 mo…
🤗 huggingchat 📰 Tech Blog
🎉 Phi-3.5 : [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…
📰 Tech Blog 📄 Paper
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…
Join Our 💬 WeChat 🧩 Discord community. MiniMax Agent ⚡️ API MCP MiniMax Website 🤗 Hugging Face 🚀 Hugging Face API 🐙 GitHub 🤖️ ModelScope 📄 License: Modified-MIT
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…