LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
LFM2.5‑VL-1.6B is Liquid AI's refreshed version of the first vision-language model, LFM2-VL-1.6B, built on an updated backbone LFM2.5-1.2B-Base and tuned for stronger real-world performance. Find more about LFM2.5 family of models in our blog post.
We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.
LFM2.5 is a new family of hybrid models designed for on-device deployment . It builds on the LFM2 architecture with extended pre-training and reinforcement learning.
Model Summary: Granite-4.0-Micro-Base is a decoder-only, long-context language model designed for a wide range of text-to-text generation tasks. It also supports Fill-in-the-Middle (FIM) code completion through the use of specialized prefix and suffix tokens. The model is trained…
Model Summary: Granite‑4.1‑3B‑Base is a decoder‑only language model with long‑context capabilities, designed to support a broad range of general text‑to‑text generation tasks, as well as fill‑in‑the‑Middle (FIM) code completion. This model shares the same underlying architecture…
Model Summary: Granite-4.0-1B-Base is a lightweight decoder-only language model designed for scenarios where efficiency and speed are critical. They can run on resource-constrained devices such as smartphones or IoT hardware, enabling offline and privacy-preserving applications.…
Longevity-LLM - LFM2-2.6B Longevity-LLM (L-LLM) is a family of compact, domain-adapted language models for interpreting heterogeneous aging biology data. Longevity LFMs are available in two sizes:
py from mistral common.tokens.tokenizers.mistral import MistralTokenizer from mistral common.protocol.instruct.messages import UserMessage from mistral common.protocol.instruct.request import ChatCompletionRequest
We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, w…
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…