AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,108 models for "Chat" Compare
🤖
Baichuan2-7B-Chat-4bits
baichuan-inc

🦉GitHub 💬WeChat 百川API支持搜索增强和192K长窗口,新增百川搜索增强知识库、限时免费! 🚀 百川大模型在线对话平台 已正式向公众开放 🎉

Open Source 7.0B ↓ 95
🤖
mistral-coreml
apple

[!IMPORTANT] ❗ This repo requires the use of the macOS Sequoia (15) Developer Beta to utilize the latest and greatest CoreML has to offer! Sign up for the Apple Beta Software Program here to get access. Check out the companion blog post to learn more about what's new in iOS 18 &…

Open Source ↓ 93
🤖
vicuna-13b-delta-finetuned-langchain-MRKL
rinna

NOTE: This "delta model" cannot be used directly. Users have to apply it on top of the original LLaMA weights to get actual vicuna-13b-finetuned-langchain-MRKL weights. See https://github.com/rinnakk/vicuna-13b-delta-finetuned-langchain-MRKL model-weights for instructions.

Open Source 13.0B ↓ 92
🤖
LongCat-Flash-Thinking-ZigZag
meituan-longcat

[2026.1.28] We have provided the TileLang kernels supporting prefill (chunked-prefill as well) and decode (multi-token prediction as well). The full attention version is placed at flash mla interface.py while the streaming sparse attention version is placed at streaming sparse at…

Open Source ↓ 91
🤖
gemma-2-baku-2b-it
rinna

Gemma 2 Baku 2B Instruct (rinna/gemma-2-baku-2b-it)

Open Source 2.0B ↓ 88
🤖
LongCat-Flash-Chat-FP8
meituan-longcat

Model Introduction We introduce LongCat-Flash, a powerful and efficient language model with 560 billion total parameters, featuring an innovative Mixture-of-Experts (MoE) architecture. The model incorporates a dynamic computation mechanism that activates 18.6B∼31.3B parameters (a…

Open Source ↓ 88
🤖
calm2-7b-chat-dpo-experimental
cyberagent

Model Card for "calm2-7b-chat-dpo-experimental"

Open Source 7.0B ↓ 84
🤖
Llama-3.1-Swallow-70B-Instruct-v0.1
tokyotech-llm

Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…

Open Source 70.0B ↓ 81
🤖
youri-7b-chat
rinna

Overview The model is the instruction-tuned version of rinna/youri-7b . It adopts a chat-style input format.

Open Source 7.0B ↓ 79
🤖
DeepSeek-R1-Distill-Qwen-14B-Japanese
cyberagent

This is a Japanese finetuned model based on deepseek-ai/DeepSeek-R1-Distill-Qwen-14B.

Reasoning 14.0B ↓ 77
🤖
LongCat-HeavyMode-Summary
meituan-longcat

We introduce an updated version of LongCat-Flash-Thinking, named LongCat-Flash-Thinking-2601, a powerful and efficient Large Reasoning Model (LRM) with 560 billion total parameters, built upon an innovative Mixture-of-Experts (MoE) architecture.

Open Source ↓ 75
🤖
sage-ft-mixtral-8x7b
apple

Authors : Yizhe Zhang, Navdeep Jaitly (Apple)

Open Source 7.0B ↓ 75