AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

351 models for "DPO" Compare
🤖
Apodex: Apodex 1.1 Mini (free)
apodex

Apodex 1.1 Mini is a reasoning-first model from Apodex, built for complex, long-horizon research and forecasting tasks. It works directly with files, data, code, and tools to produce verifiable results,...

Open Source 35.0B ★ 26.0
Solar-Open2-250B logo
Solar-Open2-250B
upstage

Solar Open 2 is Upstage’s 250B-A15B open-weight large language model, built for agentic use cases such as office productivity, document-intensive work, and coding. Its Hybrid-Attention Mixture-of-Experts (MoE) architecture with linear attention delivers highly efficient inference…

Open Source 250.0B ★ 25.0 ↓ 5.3K
Qwen3.5-35B-A3B logo
Qwen3.5-35B-A3B
Qwen

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

Open Source 35.0B ★ 24.0 ↓ 1.5M
Qwen3.5-35B-A3B-FP8 logo
Qwen3.5-35B-A3B-FP8
Qwen

[!Note] This repository contains FP8-quantized model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. The quantization method is fin…

Multimodal 35.0B ★ 24.0 ↓ 1.3M
NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 logo
NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
nvidia

:--- :--- Total Parameters 550B (55B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 8x GB200/B200/GB300/B300, 16x H100, 8x H200 Supported Languages English, French, Spanish…

Open Source 550.0B ★ 23.0 ↓ 683.5K
NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 logo
NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
nvidia

:--- :--- Total Parameters 550B (55B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 4xGB200, 4xB200, 4x GB300, 4x B300, 8xH100 Supported Languages English, French, Spanish,…

Open Source 550.0B ★ 23.0 ↓ 159.4K
Olmo-3-1025-7B logo
Olmo-3-1025-7B
allenai

We introduce Olmo 3, a new family of 7B and 32B models. This suite includes Base, Instruct, and Think variants. The Base models were trained using a staged training approach.

Open Source 7.0B ★ 20.0 ↓ 127.5K
Olmo-3-32B-Think-SFT logo
Olmo-3-32B-Think-SFT
allenai

We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.

Open Source 32.0B ★ 20.0 ↓ 43.5K
Olmo-3-1125-32B logo
Olmo-3-1125-32B
allenai

We introduce Olmo 3, a new family of 7B and 32B models. This suite includes Base, Instruct, and Think variants. The Base models were trained using a staged training approach.

Open Source 32.0B ★ 20.0 ↓ 21.9K
Olmo-3-32B-Think logo
Olmo-3-32B-Think
allenai

We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.

Open Source 32.0B ★ 20.0 ↓ 14.6K
Qwen3.6-35B-A3B-FP8 logo
Qwen3.6-35B-A3B-FP8
Qwen

[!Note] This repository contains FP8-quantized model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. The quantization method is fin…

Open Source 35.0B ★ 18.0 ↓ 4.5M
Qwen3.6-35B-A3B logo
Qwen3.6-35B-A3B
Qwen

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

Open Source 35.0B ★ 18.0 ↓ 3.5M