AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

351 models for "DPO" Compare
ERNIE-4.5-VL-424B-A47B-Paddle logo
ERNIE-4.5-VL-424B-A47B-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 424.0B ↓ 159
Swallow-70b-instruct-v0.1 logo
Swallow-70b-instruct-v0.1
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 70.0B ↓ 151
ERNIE-4.5-VL-424B-A47B-Base-Paddle logo
ERNIE-4.5-VL-424B-A47B-Base-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 424.0B ↓ 147
ERNIE-4.5-21B-A3B-Base-Paddle logo
ERNIE-4.5-21B-A3B-Base-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 21.0B ↓ 140
Klear-46B-A2.5B-Instruct logo
Klear-46B-A2.5B-Instruct
Kwai-Klear

🤗 Hugging Face 💻 Github Repository 📑 Technique Report 💬 Issues & Discussions

Open Source 46.0B ↓ 133
AI21-Jamba2-Mini-FP8 logo
AI21-Jamba2-Mini-FP8
ai21labs

Jamba2 Mini is an open source small language model built for enterprise reliability. With 12B active parameters (52B total), it delivers precise question answering without the computational overhead of reasoning models. The model's SSM-Transformer architecture provides a memory-e…

Open Source ↓ 133
llama-3-youko-70b-instruct logo
llama-3-youko-70b-instruct
rinna

Llama 3 Youko 70B Instruct (rinna/llama-3-youko-70b-instruct)

Open Source 70.0B ↓ 130
Intern-S2-Preview-397B-FP8 logo
Intern-S2-Preview-397B-FP8
internlm

💻Github Repo • 🤗HF Model Collections • 🤖ModelScope Collections • 💬Online Chat

Open Source 397.0B ↓ 130
Klear-46B-A2.5B-Base logo
Klear-46B-A2.5B-Base
Kwai-Klear

🤗 Hugging Face 💻 Github Repository 📑 Technique Report 💬 Issues & Discussions

Open Source 46.0B ↓ 123
Baichuan-M3-235B-FP8 logo
Baichuan-M3-235B-FP8
baichuan-inc

From Inquiry to Decision: Building Trustworthy Medical AI

Open Source 235.0B ↓ 109
Cola-DLM logo
Cola-DLM
ByteDance-Seed

Cola DLM ( Co ntinuous La tent D iffusion L anguage M odel) is a hierarchical continuous latent-space diffusion language model. It combines a Text VAE with a block-causal Diffusion Transformer (DiT) prior: the VAE maps text into continuous latent sequences and decodes latents bac…

Open Source ↓ 108
Nemotron-SEA-LION-v4.8-30B-A3B-NVFP4 logo
Nemotron-SEA-LION-v4.8-30B-A3B-NVFP4
aisingapore

This repository contains the NVFP4 (4-bit floating point, E2M1) quantized weights for aisingapore/Nemotron-SEA-LION-v4.8-30B-A3B .

Open Source 30.0B ↓ 97