AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

440 models for "Base Model" Compare
deepseek-vl-1.3b-chat logo
deepseek-vl-1.3b-chat
deepseek-ai

Introducing DeepSeek-VL, an open-source Vision-Language (VL) Model designed for real-world vision and language understanding applications. DeepSeek-VL possesses general multimodal understanding capabilities, capable of processing logical diagrams, web pages, formula recognition,…

Multimodal 1.3B ↓ 5.9K
granitelib-core-r1.0 logo
granitelib-core-r1.0
ibm-granite

The Granite Core Library includes three families of LoRA adapters, each developed for a specific task that enhances the base model's capabilities for explainability, calibration, and constraint verification. We provide adapters for: ibm-granite/granite-4.0-micro ibm-granite/grani…

Open Source ↓ 5.9K
GLM-Z1-32B-0414 logo
GLM-Z1-32B-0414
zai-org

The GLM family welcomes a new generation of open-source models, the GLM-4-32B-0414 series, featuring 32 billion parameters. Its performance is comparable to OpenAI's GPT series and DeepSeek's V3/R1 series, and it supports very user-friendly local deployment features. GLM-4-32B-Ba…

Reasoning 32.0B ↓ 5.9K
deepseek-vl-7b-chat logo
deepseek-vl-7b-chat
deepseek-ai

Introducing DeepSeek-VL, an open-source Vision-Language (VL) Model designed for real-world vision and language understanding applications. DeepSeek-VL possesses general multimodal understanding capabilities, capable of processing logical diagrams, web pages, formula recognition,…

Multimodal 7.0B ↓ 5.8K
Baichuan2-7B-Base logo
Baichuan2-7B-Base
baichuan-inc

🦉GitHub 💬WeChat 百川API支持搜索增强和192K长窗口,新增百川搜索增强知识库、限时免费! 🚀 百川大模型在线对话平台 已正式向公众开放 🎉

Open Source 7.0B ↓ 5.7K
OLMo-2-0325-32B logo
OLMo-2-0325-32B
allenai

We introduce OLMo 2 32B, the largest model in the OLMo 2 family. OLMo 2 was pre-trained on OLMo-mix-1124 and uses Dolmino-mix-1124 for mid-training.

Open Source 32.0B ↓ 5.7K
GLM-4-32B-0414 logo
GLM-4-32B-0414
zai-org

The GLM family welcomes new members, the GLM-4-32B-0414 series models, featuring 32 billion parameters. Its performance is comparable to OpenAI’s GPT series and DeepSeek’s V3/R1 series. It also supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-tra…

Open Source 32.0B ↓ 5.4K
d1-3B logo
d1-3B
LiquidAI

d1-3B is a 3B parameter decision model built on LFM2.5-VL-3B. You give it a state (text, JSON, images, or a mix) and a set of questions. It returns calibrated, typed answers in one forward pass with zero output tokens .

Multimodal 3.12B ↓ 5.4K
Llama-3.1-Swallow-8B-Instruct-v0.3 logo
Llama-3.1-Swallow-8B-Instruct-v0.3
tokyotech-llm

Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…

Open Source 8.0B ↓ 5.3K
Llama-3.1-Tulu-3-8B-DPO logo
Llama-3.1-Tulu-3-8B-DPO
allenai

Tülu3 is a leading instruction following model family, offering fully open-source data, code, and recipes designed to serve as a comprehensive guide for modern post-training techniques. Tülu3 is designed for state-of-the-art performance on a diversity of tasks in addition to chat…

Open Source 8.0B ↓ 5K
Yarn-Llama-2-7b-128k logo
Yarn-Llama-2-7b-128k
NousResearch

Nous-Yarn-Llama-2-13b-128k is a state-of-the-art language model for long context, further pretrained on long context data for 600 steps. This model is the Flash Attention 2 patched version of the original model: https://huggingface.co/conceptofmind/Yarn-Llama-2-13b-128k

Open Source 7.0B ↓ 4.8K
InternVL3-14B-hf logo
InternVL3-14B-hf
OpenGVLab

InternVL3-14B Transformers 🤗 Implementation

Multimodal 15.1B ↓ 4.8K