LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
[!Warning] 🚨 Qwen2.5-Math mainly supports solving English and Chinese math problems through CoT and TIR. We do not recommend using this series of models for other tasks.
Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.
👋 Join our WeChat or Discord community. 📖 Check out the GLM-5 Technical report . 📍 Use GLM-5 API services on Z.ai API Platform. 👉 One click to GLM-5 .
[🏠Homepage] [🤖 Chat with DeepSeek Coder] [Discord] [Wechat(微信)]
Description: The NVIDIA MiniMax-M2.5-NVFP4 model is the quantized version of MiniMax's MiniMax-M2.5 model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA MiniMax-M2.5 NVFP4 model is q…
Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct . This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements:
Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5-Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, 32 billion parameters, to meet the needs of different developers. Qwen2.5-Coder brings…
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model