LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
OLMoE-1B-7B is a Mixture-of-Experts LLM with 1B active and 7B total parameters released in January 2025 (0125) that is 100% open-source. It is an improved version of OLMoE-09-24, see the paper appendix for details.
📰 Tech Blog 📄 Paper Link (comming soon)
[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…
This is the base version of the Jamba model. We’ve since released a better, instruct-tuned version, Jamba-1.5-Mini. For even greater performance, check out the scaled-up Jamba-1.5-Large.
🤗 Hugging Face      🤖 ModelScope      🐙 Experience Now
👋 Join our Discord community. 📖 Check out the GLM-4.5 technical blog , technical report , and Zhipu AI technical documentation . 📍 Use GLM-4.5 API services on Zhipu AI Open Platform . 👉 One click to GLM-4.5 .
Solar Pro Preview: The most intelligent LLM on a single GPU
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
Ling is a MoE LLM provided and open-sourced by InclusionAI. We introduce two different sizes, which are Ling-lite and Ling-plus. Ling-lite has 16.8 billion parameters with 2.75 billion activated parameters, while Ling-plus has 290 billion parameters with 28.8 billion activated pa…
Model Summary: Granite-3.0-8B-Base is a decoder-only language model to support a variety of text-to-text generation tasks. It is trained from scratch following a two-stage training strategy. In the first stage, it is trained on 10 trillion tokens sourced from diverse domains. Dur…
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
FlexOlmo is a new kind of LM that unlocks a new paradigm of data collaboration. With FlexOlmo, data owners can contribute to the development of open language models without giving up control of their data. There is no need to share raw data directly, and data contributors can dec…