AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

351 models for "DPO" Compare
Olmo-3-32B-Think-DPO logo
Olmo-3-32B-Think-DPO
allenai

We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.

Open Source 32.0B ★ 31.0 ↓ 4K
Olmo-3-7B-Think-DPO logo
Olmo-3-7B-Think-DPO
allenai

We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.

Open Source 7.0B ★ 6.0 ↓ 8.7K
Olmo-3.1-32B-Instruct-DPO logo
Olmo-3.1-32B-Instruct-DPO
allenai

Model Card for Olmo-3.1-32B-Instruct-DPO

Open Source 32.0B ★ 6.0 ↓ 2.5K
Olmo-3-7B-Instruct-DPO logo
Olmo-3-7B-Instruct-DPO
allenai

We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.

Open Source 7.0B ★ 5.0 ↓ 14.7K
Nous-Hermes-2-Mixtral-8x7B-DPO logo
Nous-Hermes-2-Mixtral-8x7B-DPO
NousResearch

Nous Hermes 2 Mixtral 8x7B DPO is the new flagship Nous Research model trained over the Mixtral 8x7B MoE LLM.

Open Source 7.0B ↓ 11.9K
OLMo-2-0425-1B-DPO logo
OLMo-2-0425-1B-DPO
allenai

OLMo 2 1B DPO April 2025 is post-trained variant of the allenai/OLMo-2-0425-1B-SFT model, which has undergone supervised finetuning on an OLMo-specific variant of the Tülu 3 dataset and further DPO training on this dataset. Tülu 3 is designed for state-of-the-art performance on a…

Open Source 1.0B ↓ 7.8K
Llama-3.1-Tulu-3-8B-DPO logo
Llama-3.1-Tulu-3-8B-DPO
allenai

Tülu3 is a leading instruction following model family, offering fully open-source data, code, and recipes designed to serve as a comprehensive guide for modern post-training techniques. Tülu3 is designed for state-of-the-art performance on a diversity of tasks in addition to chat…

Open Source 8.0B ↓ 5K
OLMo-2-1124-7B-DPO logo
OLMo-2-1124-7B-DPO
allenai

Upon the initial release of OLMo-2 models, we realized the post-trained models did not share the pre-tokenization logic that the base models use. As a result, we have trained new post-trained models. The new models are available under the same names as the original models, but we…

Open Source 7.0B ↓ 3.7K
UI-TARS-7B-DPO logo
UI-TARS-7B-DPO
ByteDance-Seed

UI-TARS-7B-DPO UI-TARS-2B-SFT     UI-TARS-7B-SFT     UI-TARS-7B-DPO (Recommended)     UI-TARS-72B-SFT     UI-TARS-72B-DPO (Recommended) Introduction

Multimodal 7.0B ↓ 1.4K
Nous-Hermes-2-Mistral-7B-DPO logo
Nous-Hermes-2-Mistral-7B-DPO
NousResearch

Nous Hermes 2 on Mistral 7B DPO is the new flagship 7B Hermes! This model was DPO'd from Teknium/OpenHermes-2.5-Mistral-7B and has improved across the board on all benchmarks tested - AGIEval, BigBench Reasoning, GPT4All, and TruthfulQA.

Open Source 7.0B ↓ 848
MiniCPM-2B-dpo-bf16-llama-format logo
MiniCPM-2B-dpo-bf16-llama-format
openbmb

MiniCPM 技术报告 Technical Report OmniLMM 多模态模型 Multi-modal Model CPM-C 千亿模型试用 ~100B Model Trial

Open Source 2.4B ↓ 766
UI-TARS-72B-DPO logo
UI-TARS-72B-DPO
ByteDance-Seed

UI-TARS-72B-DPO UI-TARS-2B-SFT     UI-TARS-7B-SFT     UI-TARS-7B-DPO (Recommended)     UI-TARS-72B-SFT     UI-TARS-72B-DPO (Recommended) Introduction

Open Source 72.0B ↓ 643