AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

351 models for "DPO" Compare
Qwen3-Coder-Next-FP8 logo
Qwen3-Coder-Next-FP8
Qwen

Today, we're announcing Qwen3-Coder-Next-FP8 , an open-weight language model designed specifically for coding agents and local development. It features the following key enhancements:

Code 80.0B ★ 9.0 ↓ 1.1M
NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 logo
NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
nvidia

The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.

Open Source 30.0B ★ 9.0 ↓ 905.2K
Qwen3-Coder-Next logo
Qwen3-Coder-Next
Qwen

Today, we're announcing Qwen3-Coder-Next , an open-weight language model designed specifically for coding agents and local development. It features the following key enhancements:

Code 80.0B ★ 9.0 ↓ 599.2K
NVIDIA-Nemotron-3-Nano-30B-A3B-FP8 logo
NVIDIA-Nemotron-3-Nano-30B-A3B-FP8
nvidia

The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.

Open Source 30.0B ★ 9.0 ↓ 246K
granite-4.2-3b logo
granite-4.2-3b
ibm-granite

--- --- Developers Granite Team, IBM Model Type Decoder-only Dense Transformer (Reasoning) Architecture GraniteForCausalLM Base Model Granite-4.1-3B-Base Parameters 3B Context Length Natively Supports 128K (Long-context extension to 512K) Precision bfloat16 Tested Languages Engli…

Open Source 3.0B ★ 9.0 ↓ 49.3K
Aurora-Spec-Qwen3-Coder-Next-FP8 logo
Aurora-Spec-Qwen3-Coder-Next-FP8
togethercomputer

This is an EAGLE3 draft model trained from scratch (random initialization) using the Aurora inference-time training framework for speculative decoding. Unlike traditional approaches that fine-tune pre-trained models, this model is built entirely through Aurora's online training p…

Code ★ 9.0 ↓ 354
Mistral: Mistral Large 3 2512 (batch) logo
Mistral: Mistral Large 3 2512 (batch)
mistralai

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.

Open Source 675.0B ★ 9.0
LFM2.5-2.6B logo
LFM2.5-2.6B
LiquidAI

LFM2.5-2.6B is part of LFM2.5, a family of hybrid models designed for on-device deployment . It builds on the LFM2 architecture with a 128K context window and agentic post-training.

Open Source 2.6B ★ 8.0 ↓ 121.6K
EXAONE-4.0-32B logo
EXAONE-4.0-32B
LGAI-EXAONE

🎉 License Updated! We are pleased to announce our more flexible licensing terms 🤗 ✈️ Try on FriendliAI (licensed under commercial purposes) 📢 EXAONE 4.0 is officially supported by HuggingFace transformers! Please check out the guide below

Reasoning 32.0B ★ 8.0 ↓ 29.8K
Step3-VL-10B logo
Step3-VL-10B
stepfun-ai

- 🚀 Online Demo : Explore Step3-VL-10B on Hugging Face Spaces ! - 📢 [Notice] FP8 Quantization Support : FP8 quantized weights are now available. (Download link) - 📢 [Notice] vLLM Support: vLLM integration is now officially supported! (PR 32329) - ✅ [Fixed] HF Inference: Resolved…

Multimodal 10.0B ★ 8.0 ↓ 26.9K
LFM2.5-2.6B-DSpark logo
LFM2.5-2.6B-DSpark
LiquidAI

LFM2.5-DSpark is a family of speculative-decoding draft models that adapt DSpark for the LFM2.5 architecture. They allow LFM2.5 models to run faster without degrading quality.

Open Source 2.6B ★ 8.0 ↓ 22.5K
LFM2.5-2.6B-Base logo
LFM2.5-2.6B-Base
LiquidAI

LFM2.5 is a new family of hybrid models designed for on-device deployment . It builds on the LFM2 architecture with extended pre-training and reinforcement learning.

Open Source 2.6B ★ 8.0 ↓ 10.3K