AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,108 models for "Chat" Compare
🤖
LongCat-Flash-Lite
meituan-longcat

Model Introduction We introduce LongCat-Flash-Lite, a non-thinking 68.5B parameter Mixture-of-Experts (MoE) model with approximately 3B activated parameters, supporting a 256k context length through the YaRN method. Building upon the LongCat-Flash architecture, LongCat-Flash-Lite…

Open Source ★ 17.0 ↓ 15.4K
🤖
LongCat-Flash-Lite-FP8
meituan-longcat

Model Introduction We introduce LongCat-Flash-Lite, a non-thinking 68.5B parameter Mixture-of-Experts (MoE) model with approximately 3B activated parameters, supporting a 256k context length through the YaRN method. Building upon the LongCat-Flash architecture, LongCat-Flash-Lite…

Open Source ★ 17.0 ↓ 14.3K
🤖
LongCat-Flash-Lite-Sparse
meituan-longcat

LongCat-Flash-Lite-Sparse is a non-thinking Mixture-of-Experts (MoE) model with 69B total parameters and approximately 3B activated parameters per token. Built on LongCat-Flash-Lite, it replaces dense MLA with LongCat Sparse Attention (LSA) and natively supports context lengths o…

Open Source ★ 17.0 ↓ 1.3K
🤖
Mistral Large
mistralai

This is Mistral AI's flagship model, Mistral Large 2 (version `mistral-large-2407`). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more. Read the launch announcement [here](https://mistral.ai/news/mistral-large-2407/)....

Closed Source ★ 16.0
🤖
gpt-oss-20b
openai

Try gpt-oss · Guides · Model card · OpenAI blog

Open Source 20.0B ★ 15.0 ↓ 6.5M
🤖
NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
nvidia

The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.

Open Source 30.0B ★ 15.0 ↓ 871K
🤖
NVIDIA-Nemotron-3-Nano-30B-A3B-FP8
nvidia

The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.

Open Source 30.0B ★ 15.0 ↓ 685.6K
🤖
NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4
nvidia

The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.

Open Source 30.0B ★ 15.0 ↓ 580.3K
🤖
NVIDIA-Nemotron-3-Nano-30B-A3B-Base-BF16
nvidia

NVIDIA-Nemotron-3-Nano-30B-A3B-Base-BF16

Open Source 30.0B ★ 15.0 ↓ 121.6K
🤖
InternVL3_5-GPT-OSS-20B-A4B-Preview
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 20.0B ★ 15.0 ↓ 30.8K
🤖
Solar-Open-100B
upstage

Solar Open is Upstage's flagship 102B-parameter large language model, trained entirely from scratch and released under the Upstage Solar License (see LICENSE for details). As a Mixture-of-Experts (MoE) architecture, it delivers enterprise-grade performance in reasoning, instructi…

Open Source 100.0B ★ 15.0 ↓ 14.8K
🤖
Llama-4-Maverick-17B-128E-Instruct-FP8
meta-llama

Multimodal 400.0B ★ 14.0 ↓ 75.2K