AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

351 models for "DPO" Compare
ERNIE-4.5-300B-A47B-2Bits-TP2-Paddle logo
ERNIE-4.5-300B-A47B-2Bits-TP2-Paddle
baidu

The advanced capabilities of the ERNIE 4.5 models, particularly the MoE-based A47B and A3B series, are underpinned by several key technical innovations:

Open Source 300.0B ★ 8.0 ↓ 128
ERNIE-4.5-300B-A47B-2Bits-TP4-Paddle logo
ERNIE-4.5-300B-A47B-2Bits-TP4-Paddle
baidu

The advanced capabilities of the ERNIE 4.5 models, particularly the MoE-based A47B and A3B series, are underpinned by several key technical innovations:

Open Source 300.0B ★ 8.0 ↓ 85
Qwen3.5-2B logo
Qwen3.5-2B
Qwen

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. In light of its parameter scale, the intende…

Open Source 2.0B ★ 7.0 ↓ 4M
NVIDIA-Nemotron-3-Nano-4B-BF16 logo
NVIDIA-Nemotron-3-Nano-4B-BF16
nvidia

The pretraining data has a cutoff date of September 2024\.

Open Source 4.0B ★ 7.0 ↓ 1.9M
Kimi-Linear-48B-A3B-Instruct logo
Kimi-Linear-48B-A3B-Instruct
moonshotai

Kimi Linear: An Expressive, Efficient Attention Architecture

Open Source 48.0B ★ 7.0 ↓ 223K
LFM2.5-8B-A1B logo
LFM2.5-8B-A1B
LiquidAI

LFM2.5 is a new family of hybrid models designed for on-device deployment. It builds on the LFM2 architecture with extended pre-training and reinforcement learning.

Open Source 8.0B ★ 7.0 ↓ 25.8K
Olmo-3.1-32B-Think logo
Olmo-3.1-32B-Think
allenai

We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.

Open Source 32.0B ★ 7.0 ↓ 19.2K
LFM2.5-8B-A1B-Base logo
LFM2.5-8B-A1B-Base
LiquidAI

LFM2.5 is a new family of hybrid models designed for on-device deployment. It builds on the LFM2 architecture with extended pre-training and reinforcement learning.

Open Source 8.3B ★ 7.0 ↓ 2.9K
LFM2.5-8B-A1B-DSpark logo
LFM2.5-8B-A1B-DSpark
LiquidAI

LFM2.5-DSpark is a family of speculative-decoding draft models that adapt DSpark for the LFM2.5 architecture. They allow LFM2.5 models to run faster without degrading quality.

Open Source 8.0B ★ 7.0 ↓ 1.9K
Hermes-3-Llama-3.1-405B logo
Hermes-3-Llama-3.1-405B
NousResearch

Hermes 3 405B is the latest flagship model in the Hermes series of LLMs by Nous Research, and the first full parameter finetune since the release of Llama-3.1 405B.

Open Source 405.0B ★ 7.0 ↓ 1K
Ring-flash-2.0 logo
Ring-flash-2.0
inclusionAI

This model is presented in the paper Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model.

Open Source ★ 7.0 ↓ 653
Qwen3.5-0.8B logo
Qwen3.5-0.8B
Qwen

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. In light of its parameter scale, the intende…

Open Source 0.8B ★ 6.0 ↓ 2.5M