AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

351 models for "DPO" Compare
MiniCPM-2B-dpo-bf16 logo
MiniCPM-2B-dpo-bf16
openbmb

MiniCPM 技术报告 Technical Report OmniLMM 多模态模型 Multi-modal Model CPM-C 千亿模型试用 ~100B Model Trial

Open Source 2.0B ↓ 625
Falcon-H1-Tiny-90M-Instruct-pre-DPO logo
Falcon-H1-Tiny-90M-Instruct-pre-DPO
tiiuae

0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation

Open Source 0.09B ↓ 272
Falcon-H1-Tiny-90M-Instruct-Curriculum-pre-DPO logo
Falcon-H1-Tiny-90M-Instruct-Curriculum-pre-DPO
tiiuae

0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation

Open Source 0.09B ↓ 243
calm2-7b-chat-dpo-experimental logo
calm2-7b-chat-dpo-experimental
cyberagent

Model Card for "calm2-7b-chat-dpo-experimental"

Open Source 7.0B ↓ 130
gemma-2-9b-it-MoAA-DPO logo
gemma-2-9b-it-MoAA-DPO
togethercomputer

This is the DPO model in our Mixture of Agents Alignment (MoAA) pipeline. This model is tuned on the Gemma-2-9b-it. MoAA is an approach that leverages collective intelligence from open‑source LLMs to advance alignment.

Open Source 9.0B ↓ 83
Llama-3.1-8B-Instruct-MoAA-DPO logo
Llama-3.1-8B-Instruct-MoAA-DPO
togethercomputer

This is the DPO model in our Mixture of Agents Alignment (MoAA) pipeline. This model is tuned on the Llama-3.1-8b-Instruct. MoAA is an approach that leverages collective intelligence from open‑source LLMs to advance alignment.

Open Source 8.0B ↓ 81
SmolVLM-Instruct-DPO logo
SmolVLM-Instruct-DPO
HuggingFaceTB

SmolVLM is a compact open multimodal model that accepts arbitrary sequences of image and text inputs to produce text outputs. Designed for efficiency, SmolVLM can answer questions about images, describe visual content, create stories grounded on multiple images, or function as a…

Multimodal ↓ 39
Anthropic: Claude Opus 5.5 (batch) logo
Anthropic: Claude Opus 5.5 (batch)
anthropic

Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particularly strong at multi-step changes in large codebases, code...

Closed Source ★ 58.0
Anthropic: Claude Sonnet 5.5 (batch) logo
Anthropic: Claude Sonnet 5.5 (batch)
anthropic

Claude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade. It is especially strong at building features, fixing bugs, and producing...

Closed Source ★ 56.0
OpenAI: GPT-6 Astra Pro (batch) logo
OpenAI: GPT-6 Astra Pro (batch)
openai

GPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Closed Source ★ 53.0
OpenAI: GPT-6 Astra Pro logo
OpenAI: GPT-6 Astra Pro
openai

GPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Closed Source ★ 53.0
OpenAI: GPT-6.1 Sol Pro (batch) logo
OpenAI: GPT-6.1 Sol Pro (batch)
openai

GPT-6.1 Sol Pro is the same underlying model as [GPT-6.1 Sol](https://openrouter.ai/openai/gpt-6.1-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note:** pro mode spends far more...

Closed Source ★ 52.0