AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,109 models for "Chat" Compare
🤖
GTA1-7B
Salesforce

Reinforcement learning (RL) (e.g., GRPO) helps with grounding because of its inherent objective alignment—rewarding successful clicks—rather than encouraging long textual Chain-of-Thought (CoT) reasoning. Unlike approaches that rely heavily on verbose CoT reasoning, GRPO directly…

Open Source 7.0B ↓ 204
🤖
bloom-560m-intermediate
bigscience

WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).

Open Source ↓ 201
🤖
AquilaMoE-SFT
BAAI

AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies Language Foundation Model & Software Team Beijing Academy of Artificial Intelligence (BAAI) [Paper(released soon)] [Code] [github]

Open Source ↓ 199
🤖
internlm2-math-20b
internlm

State-of-the-art bilingual open-sourced Math reasoning LLMs. A solver , prover , verifier , augmentor .

Open Source 20.0B ↓ 195
🤖
AI21-Jamba2-Mini-FP8
ai21labs

Jamba2 Mini is an open source small language model built for enterprise reliability. With 12B active parameters (52B total), it delivers precise question answering without the computational overhead of reasoning models. The model's SSM-Transformer architecture provides a memory-e…

Open Source ↓ 193
🤖
Gemma-SEA-Guard-12B-2602
aisingapore

SEA-Guard is a collection of safety-focused Large Language Models (LLMs) designed specifically for the Southeast Asia (SEA) region. While the collection comprises four distinct models, we currently offer a single API endpoint that exclusively serves the Gemma-based model. You can…

Open Source 12.0B ↓ 192
🤖
Spatial-SSRL-3B
internlm

📖 Paper 🏠 Github 🤗 Spatial-SSRL-7B Model 🤗 Spatial-SSRL-3B Model 🤗 Spatial-SSRL-Qwen3VL-4B Model 🤗 Spatial-SSRL-81k Dataset 📰 Daily Paper

Open Source 3.0B ↓ 189
🤖
OREAL-7B
internlm

- Arxiv - Github - Model Collection - Data

Open Source 7.0B ↓ 188
🤖
Medical-Qwen3-Swallow-32B
tokyotech-llm

Medical-Qwen3-Swallow-32B is a medical-domain language model based on tokyotech-llm/Qwen3-Swallow-32B-RL-v0.2. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.

Open Source 32.0B ↓ 184
🤖
GPT-NeoXT-Chat-Base-20B
togethercomputer

Feel free to try out our OpenChatKit feedback app!

Open Source 20.0B ↓ 183
🤖
JanusCoder-8B
internlm

💻Github Repo • 🤗Model Collections • 📜Technical Report

Code 8.0B ↓ 183
🤖
Redmond-Hermes-Coder
NousResearch

Redmond-Hermes-Coder 15B is a state-of-the-art language model fine-tuned on over 300,000 instructions. This model was fine-tuned by Nous Research, with Teknium and Karan4D leading the fine tuning process and dataset curation, Redmond AI sponsoring the compute, and several other c…

Code ↓ 183