LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
🎉 Phi-4 : [multimodal-instruct onnx]; [mini-instruct onnx]
Mage-VL An Efficient Codec-Native Streaming Multimodal Foundation Model
The Meta Llama 3.1 collection of multilingual large language models (LLMs) is a collection of pretrained and instruction tuned generative models in 8B, 70B and 405B sizes (text in/text out). The Llama 3.1 instruction tuned text only models (8B, 70B, 405B) are optimized for multil…
🎉 Phi-4 : [multimodal-instruct onnx]; [mini-instruct onnx]
Phi-tiny-MoE is a lightweight Mixture of Experts (MoE) model with 3.8B total parameters and 1.1B activated parameters. It is compressed and distilled from the base model shared by Phi-3.5-MoE and GRIN-MoE using the SlimMoE approach, then post-trained via supervised fine-tuning an…
Introducing DeepSeek-VL2, an advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL. DeepSeek-VL2 demonstrates superior capabilities across various tasks, including but not limited to visual question…
OLMoE-1B-7B-0125-Instruct January 2025 is post-trained variant of the OLMoE-1B-7B January 2025 model, which has undergone supervised finetuning on an OLMo-specific variant of the Tülu 3 dataset and further DPO training on this dataset, and finally RLVR training using this data. T…
Quantized to FP8 Version of olmOCR-2-7B-1025, using llmcompressor.
Join Our 💬 WeChat 🧩 Discord community. MiniMax Agent ⚡️ API (Now Free for a limited time!) MCP MiniMax Website 🤗 Hugging Face 🐙 GitHub 🤖️ ModelScope 📄 License: MIT