AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

351 models for "DPO" Compare
ERNIE-4.5-0.3B-Paddle logo
ERNIE-4.5-0.3B-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 0.3B ↓ 520
InternVL3_5-2B-Flash logo
InternVL3_5-2B-Flash
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 2.3B ↓ 488
xgen-mm-phi3-mini-instruct-r-v1 logo
xgen-mm-phi3-mini-instruct-r-v1
Salesforce

📣 News 📌 [08/19/2024] xGen-MM-v1.5 released: - 🤗 xgen-mm-phi3-mini-instruct-interleave-r-v1.5 - 🤗 xgen-mm-phi3-mini-base-r-v1.5 - 🤗 xgen-mm-phi3-mini-instruct-singleimg-r-v1.5 - 🤗 xgen-mm-phi3-mini-instruct-dpo-r-v1.5

Multimodal 5.0B ↓ 472
Intern-S2-Preview-397B logo
Intern-S2-Preview-397B
internlm

💻Github Repo • 🤗HF Model Collections • 🤖ModelScope Collections • 💬Online Chat

Open Source 397.0B ↓ 445
llama-3-youko-8b-instruct logo
llama-3-youko-8b-instruct
rinna

Llama 3 Youko 8B Instruct (rinna/llama-3-youko-8b-instruct)

Open Source 8.0B ↓ 402
xLAM-7b-fc-r logo
xLAM-7b-fc-r
Salesforce

[Homepage] [Paper] [Discord] [Dataset] [Github]

Open Source 7.0B ↓ 384
Nous-Hermes-2-Mixtral-8x7B-SFT logo
Nous-Hermes-2-Mixtral-8x7B-SFT
NousResearch

Nous Hermes 2 Mixtral 8x7B SFT is the supervised finetune only version of our new flagship Nous Research model trained over the Mixtral 8x7B MoE LLM.

Open Source 7.0B ↓ 375
GTA1-7B-2507 logo
GTA1-7B-2507
Salesforce

Reinforcement learning (RL) (e.g., GRPO) helps with grounding because of its inherent objective alignment—rewarding successful clicks—rather than encouraging long textual Chain-of-Thought (CoT) reasoning. Unlike approaches that rely heavily on verbose CoT reasoning, GRPO directly…

Multimodal 7.0B ↓ 372
MiniCPM-2B-sft-fp32 logo
MiniCPM-2B-sft-fp32
openbmb

MiniCPM 技术报告 Technical Report OmniLMM 多模态模型 Multi-modal Model CPM-C 千亿模型试用 ~100B Model Trial

Open Source 2.0B ↓ 371
SuperApriel-15B-Instruct logo
SuperApriel-15B-Instruct
ServiceNow-AI

A 15B-parameter token-mixer supernet with 8 optimized deployment presets spanning 1.0× to 10.7× decode throughput at 32K sequence length, all from a single checkpoint. Derived from Apriel-1.6 through stochastic distillation and targeted supervised fine-tuning.

Open Source 15.0B ↓ 365
Yi-6B-Chat-4bits logo
Yi-6B-Chat-4bits
01-ai

Building the Next Generation of Open-Source and Bilingual LLMs

Open Source 6.0B ↓ 356
Baichuan-M3-235B-GPTQ-INT4 logo
Baichuan-M3-235B-GPTQ-INT4
baichuan-inc

From Inquiry to Decision: Building Trustworthy Medical AI

Open Source 235.0B ↓ 349