AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

63 models for "Open Weights" Compare
gemma-2-2b-it logo
gemma-2-2b-it
google

Open Source 2.0B ↓ 530.9K
gemma-3-27b-it logo
gemma-3-27b-it
google

Multimodal 27.0B ↓ 417.6K
OLMo-2-0425-1B logo
OLMo-2-0425-1B
allenai

We introduce OLMo 2 1B, the smallest model in the OLMo 2 family. OLMo 2 was pre-trained on OLMo-mix-1124 and uses Dolmino-mix-1124 for mid-training.

Open Source 1.0B ↓ 314.9K
SmolLM3-3B-Base logo
SmolLM3-3B-Base
HuggingFaceTB

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. License

Open Source 3.0B ↓ 103.5K
NVIDIA-Nemotron-Parse-2.0 logo
NVIDIA-Nemotron-Parse-2.0
nvidia

Description: NVIDIA Nemotron Parse 2.0 transforms document images into structured, machine-readable representations with text, layout classes, bounding boxes, and reading-order information. Given a Red, Green, Blue (RGB) document image and a task prompt, the model produces format…

Multimodal 0.905B ↓ 101.7K
Nemotron-Labs-Diffusion-3B logo
Nemotron-Labs-Diffusion-3B
nvidia

Nemotron-Labs-Diffusion is a tri-mode language model that supports both AR decoding and diffusion-based parallel decoding by simply switching the attention pattern of the same model during inference. The synergy between these two modes enables a third mode, called self-speculatio…

Open Source 3.0B ↓ 70.7K
olmOCR-7B-0225-preview logo
olmOCR-7B-0225-preview
allenai

This is a preview release of the olmOCR model that's fine tuned from Qwen2-VL-7B-Instruct using the olmOCR-mix-0225 dataset.

Multimodal 7.0B ↓ 62.6K
granite-4.0-3b-vision logo
granite-4.0-3b-vision
ibm-granite

Model Summary: Granite-4.0-3B-Vision is a vision-language model (VLM) designed for enterprise-grade document data extraction. It focuses on specialized, complex extraction tasks that ultracompact models often struggle with:

Multimodal 4.0B ↓ 56.3K
OLMo-2-1124-7B-Instruct logo
OLMo-2-1124-7B-Instruct
allenai

Upon the initial release of OLMo-2 models, we realized the post-trained models did not share the pre-tokenization logic that the base models use. As a result, we have trained new post-trained models. The new models are available under the same names as the original models, but we…

Open Source 7.0B ↓ 36.1K
OLMo-2-1124-13B-Instruct logo
OLMo-2-1124-13B-Instruct
allenai

Upon the initial release of OLMo-2 models, we realized the post-trained models did not share the pre-tokenization logic that the base models use. As a result, we have trained new post-trained models. The new models are available under the same names as the original models, but we…

Open Source 13.0B ↓ 12.5K
OLMo-2-0325-32B-Instruct logo
OLMo-2-0325-32B-Instruct
allenai

OLMo 2 32B Instruct March 2025 is post-trained variant of the OLMo-2 32B March 2025 model, which has undergone supervised finetuning on an OLMo-specific variant of the Tülu 3 dataset, further DPO training on this dataset, and final RLVR training on this dataset. Tülu 3 is designe…

Open Source 32.0B ↓ 9.9K
d1-omni-600M logo
d1-omni-600M
LiquidAI

d1-omni-600M is a 600M parameter decision model built on LFM2.5-Encoder-350M. You give it a state (text or JSON, with images or a voice clip) and a set of named questions. It returns typed answers with zero output tokens : every answer is read directly from the model's distributi…

Multimodal 0.587B ↓ 8.5K