AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

351 models for "DPO" Compare
Hy-Embodied-VLM-1.0 logo
Hy-Embodied-VLM-1.0
tencent

Hy-Embodied-VLM-1.0 Efficient Physical-World Agents Tencent Robotics X × Hy Vision Team × Futian Laboratory

Multimodal ↓ 659
InternVL3_5-1B-Instruct logo
InternVL3_5-1B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 1.0B ↓ 641
Hunyuan-0.5B-Pretrain logo
Hunyuan-0.5B-Pretrain
tencent

🤗  HuggingFace     🤖  ModelScope     🪡  AngelSlim

Open Source 0.5B ↓ 629
ERNIE-4.5-21B-A3B-Base-PT logo
ERNIE-4.5-21B-A3B-Base-PT
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 21.0B ↓ 611
xgen-mm-phi3-mini-instruct-interleave-r-v1.5 logo
xgen-mm-phi3-mini-instruct-interleave-r-v1.5
Salesforce

Model description xGen-MM is a series of the latest foundational Large Multimodal Models (LMMs) developed by Salesforce AI Research. This series advances upon the successful designs of the BLIP series, incorporating fundamental enhancements that ensure a more robust and superior…

Open Source ↓ 573
Hunyuan-4B-Pretrain logo
Hunyuan-4B-Pretrain
tencent

🤗  HuggingFace     🤖  ModelScope     🪡  AngelSlim

Open Source 4.0B ↓ 560
InternVL3_5-14B-Instruct logo
InternVL3_5-14B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 15.1B ↓ 555
Baichuan-M3-235B logo
Baichuan-M3-235B
baichuan-inc

From Inquiry to Decision: Building Trustworthy Medical AI

Open Source 235.0B ↓ 547
Swallow-MS-7b-instruct-v0.1 logo
Swallow-MS-7b-instruct-v0.1
tokyotech-llm

Our Swallow-MS-7b-v0.1 model has undergone continual pre-training from the Mistral-7B-v0.1, primarily with the addition of Japanese language data.

Open Source 7.0B ↓ 546
MiniCPM-2B-128k logo
MiniCPM-2B-128k
openbmb

MiniCPM is an End-Size LLM developed by ModelBest Inc. and TsinghuaNLP, with only 2.4B parameters excluding embeddings. MiniCPM-2B-128k is a long context extension trial of MiniCPM-2B. To our best knowledge, MiniCPM-2B-128k is the first long context( =128k) SLM smaller than 3B。 I…

Open Source 2.0B ↓ 546
InternVL3_5-38B-Flash logo
InternVL3_5-38B-Flash
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 38.0B ↓ 537
SmolLM2-1.7B-Instruct-16k logo
SmolLM2-1.7B-Instruct-16k
HuggingFaceTB

This is a 16k context version of SmolLM2-1.7B-Instruct, which originnaly only supported 8k context. We finetune the model on 15k samples consisting of a subset of SmolTalk, LongAlign and SeaLong datasets and increase RoPE from 100k to 500k. This improves the evaluation on HELMET…

Open Source 1.7B ↓ 533