AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,877 models Compare
🤖
Llama-3.2-3B-Instruct-QLORA_INT4_EO8
meta-llama

Open Source 3.0B ↓ 149
🤖
bloomz-p3
bigscience

1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation

Open Source ↓ 148
🤖
smollm-360M-instruct-add-basics
HuggingFaceTB

Model Summary Chat with the model at: https://huggingface.co/spaces/HuggingFaceTB/instant-smol

Open Source ↓ 148
🤖
PaCoRe-8B
stepfun-ai

PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning

Open Source 8.0B ↓ 148
🤖
AHN-Mamba2-for-Qwen-2.5-Instruct-7B
ByteDance-Seed

AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling

Open Source 7.0B ↓ 148
🤖
internlm2-chat-20b-sft
internlm

💻Github Repo • 🤔Reporting Issues • 📜Technical Report

Open Source 20.0B ↓ 147
🤖
deepseekcoder-33b-codeqwen-align-subset
bigcode

This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.

Code 33.0B ↓ 147
🤖
LongCat-Flash-Prover
meituan-longcat

We introduce LongCat-Flash-Prover , a flagship $560$-billion-parameter open-source Mixture-of-Experts (MoE) model that advances Native Formal Reasoning in Lean4 through agentic tool-integrated reasoning (TIR). We decompose the native formal reasoning task into three independent f…

Open Source ↓ 146
🤖
SmolVLM-Synthetic
HuggingFaceTB

SmolVLM is a compact open multimodal model that accepts arbitrary sequences of image and text inputs to produce text outputs. Designed for efficiency, SmolVLM can answer questions about images, describe visual content, create stories grounded on multiple images, or function as a…

Multimodal ↓ 146
🤖
BFS-Prover-V1-7B
ByteDance-Seed

🚀 BFS-Prover: Scalable Best-First Tree Search for LLM-based Automatic Theorem Proving State-of-the-art tactic generation model in Lean4

Open Source 7.0B ↓ 145
🤖
nekomata-7b
rinna

Overview We conduct continual pre-training of qwen-7b on 30B tokens from a mixture of Japanese and English datasets. The continual pre-training significantly improves the model's performance on Japanese tasks. It also enjoys the following great features provided by the original Q…

Open Source 7.0B ↓ 144
🤖
Llama-3.2-3B-Instruct-SpinQuant_INT4_EO8
meta-llama

Open Source 3.0B ↓ 143