AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

351 models for "DPO" Compare
Baichuan-M2-32B-GPTQ-Int4 logo
Baichuan-M2-32B-GPTQ-Int4
baichuan-inc

Baichuan-M2-32B is Baichuan AI's medical-enhanced reasoning model, the second medical model released by Baichuan. Designed for real-world medical reasoning tasks, this model builds upon Qwen2.5-32B with an innovative Large Verifier System. Through domain-specific fine-tuning on r…

Open Source 32.0B ↓ 243
Intern-S2-Preview-FP8 logo
Intern-S2-Preview-FP8
internlm

💻Github Repo • 🤗Model Collections • 💬Online Chat

Open Source ↓ 226
ERNIE-4.5-VL-28B-A3B-Base-PT logo
ERNIE-4.5-VL-28B-A3B-Base-PT
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 28.0B ↓ 212
UI-TARS-72B-SFT logo
UI-TARS-72B-SFT
ByteDance-Seed

UI-TARS-72B-SFT UI-TARS-2B-SFT     UI-TARS-7B-SFT     UI-TARS-7B-DPO (Recommended)     UI-TARS-72B-SFT     UI-TARS-72B-DPO (Recommended) Introduction

Open Source 72.0B ↓ 204
GTA1-7B logo
GTA1-7B
Salesforce

Reinforcement learning (RL) (e.g., GRPO) helps with grounding because of its inherent objective alignment—rewarding successful clicks—rather than encouraging long textual Chain-of-Thought (CoT) reasoning. Unlike approaches that rely heavily on verbose CoT reasoning, GRPO directly…

Open Source 7.0B ↓ 204
ERNIE-4.5-21B-A3B-Paddle logo
ERNIE-4.5-21B-A3B-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 21.0B ↓ 194
Swallow-7b-instruct-v0.1 logo
Swallow-7b-instruct-v0.1
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 7.0B ↓ 174
ERNIE-4.5-VL-28B-A3B-Paddle logo
ERNIE-4.5-VL-28B-A3B-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 28.0B ↓ 172
Swallow-13b-instruct-v0.1 logo
Swallow-13b-instruct-v0.1
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 13.0B ↓ 172
smollm-360M-instruct-add-basics logo
smollm-360M-instruct-add-basics
HuggingFaceTB

Model Summary Chat with the model at: https://huggingface.co/spaces/HuggingFaceTB/instant-smol

Open Source ↓ 169
ERNIE-4.5-0.3B-Base-Paddle logo
ERNIE-4.5-0.3B-Base-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 0.3B ↓ 167
ERNIE-4.5-VL-28B-A3B-Base-Paddle logo
ERNIE-4.5-VL-28B-A3B-Base-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 28.0B ↓ 160