LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
SmolVLM2-500M-Video is a lightweight multimodal model designed to analyze video content. The model processes videos, images, and text inputs to generate text outputs - whether answering questions about media files, comparing visual content, or transcribing text from images. Despi…
DeepSeek-V3.2: Efficient Reasoning & Agentic AI
Latest Updates: In addition to the original formula, we have further enhanced Qwen2.5-VL-32B's mathematical and problem-solving abilities through reinforcement learning. This has also significantly improved the model's subjective user experience, with response styles adjusted to…
Phi-2 is a Transformer with 2.7 billion parameters. It was trained using the same data sources as Phi-1.5, augmented with a new data source that consists of various NLP synthetic texts and filtered websites (for safety and educational value). When assessed against benchmarks test…
1. Model Summary 2. Limitations 3. Training 4. License 5. Citation
Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao, Maxwell Horton, Yanzi Jin, Chenfan Sun, Iman Mirzadeh, Mahyar Najibi, Dmitry Belenko, Peter Zatloukal, Mohammad Rastegari
Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5-Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, 32 billion parameters, to meet the needs of different developers. Qwen2.5-Coder brings…
We introduce the updated version of the Qwen3-30B-A3B non-thinking mode , named Qwen3-30B-A3B-Instruct-2507 , featuring the following key enhancements:
We're excited to unveil Qwen2-VL , the latest iteration of our Qwen-VL model, representing nearly a year of innovation.