LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
OLMoE-1B-7B-0125-Instruct January 2025 is post-trained variant of the OLMoE-1B-7B January 2025 model, which has undergone supervised finetuning on an OLMo-specific variant of the Tülu 3 dataset and further DPO training on this dataset, and finally RLVR training using this data. T…
Quantized to FP8 Version of olmOCR-2-7B-1025, using llmcompressor.
Join Our 💬 WeChat 🧩 Discord community. MiniMax Agent ⚡️ API (Now Free for a limited time!) MCP MiniMax Website 🤗 Hugging Face 🐙 GitHub 🤖️ ModelScope 📄 License: MIT
[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…
Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks
This is a 164M parameters model with the same architecture as StarCoder (8k context length, MQA & FIM). It was trained on the Python data from StarCoderData for ~6 epochs which amounts to 100B tokens.
DeepSeek-V3.1 is a hybrid model that supports both thinking mode and non-thinking mode. Compared to the previous version, this upgrade brings improvements in multiple aspects:
LLaDA2.0-mini is a diffusion language model featuring a 16BA1B Mixture-of-Experts (MoE) architecture. As an enhanced, instruction-tuned iteration of the LLaDA series, it is optimized for practical applications.
📖 View the GLM-4.1V-9B-Thinking paper . 📍 Using GLM-4.1V-9B-Thinking API at Zhipu Foundation Model Open Platform
Llama-3.1-Nemotron-Nano-8B-v1 is a large language model (LLM) which is a derivative of Meta Llama-3.1-8B-Instruct (AKA the reference model). It is a reasoning model that is post trained for reasoning, human chat preferences, and tasks, such as RAG and tool calling.