AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,111 models for "Chat" Compare
🤖
MiMo-VL-7B-SFT-2508
XiaomiMiMo

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ MiMo-VL Technical Report ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Multimodal 7.0B ↓ 769
🤖
Gemma-2-Llama-Swallow-9b-it-v0.1
tokyotech-llm

Gemma-2-Llama-Swallow series was built by continual pre-training on the gemma-2 models. Gemma 2 Swallow enhanced the Japanese language capabilities of the original Gemma 2 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled…

Open Source 9.0B ↓ 751
🤖
Yi-Coder-1.5B-Chat
01-ai

🐙 GitHub • 👾 Discord • 🐤 Twitter • 💬 WeChat 📝 Paper • 💪 Tech Blog • 🙌 FAQ • 📗 Learning Hub

Code 1.5B ↓ 746
🤖
sarashina2.2-1b-instruct-v0.1
sbintuitions

sbintuitions/sarashina2.2-1b-instruct-v0.1

Open Source 1.0B ↓ 745
🤖
GPT-OSS-Swallow-120B-RL-v0.1-MXFP4
tokyotech-llm

GPT-OSS-Swallow v0.1 is a family of large language models available in 20B and 120B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifiable Rewa…

Open Source 120.0B ↓ 744
🤖
Gemma-2-Llama-Swallow-2b-it-v0.1
tokyotech-llm

Gemma-2-Llama-Swallow series was built by continual pre-training on the gemma-2 models. Gemma 2 Swallow enhanced the Japanese language capabilities of the original Gemma 2 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled…

Open Source 2.0B ↓ 743
🤖
instructblip-vicuna-13b
Salesforce

InstructBLIP model using Vicuna-13b as language model. InstructBLIP was introduced in the paper InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning by Dai et al.

Open Source 13.0B ↓ 741
🤖
UI-Mate-9B
tencent

UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations

Open Source 9.0B ↓ 740
🤖
MiMo-7B-SFT
XiaomiMiMo

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Unlocking the Reasoning Potential of Language Model From Pretraining to Posttraining ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Open Source 7.0B ↓ 737
🤖
InternVL3_5-8B-Pretrained
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 8.0B ↓ 735
🤖
FastVLM-7B
apple

FastVLM: Efficient Vision Encoding for Vision Language Models

Multimodal 7.0B ↓ 732
🤖
GLM-4.5-Base
zai-org

👋 Join our Discord community. 📖 Check out the GLM-4.5 technical blog , technical report , and Zhipu AI technical documentation . 📍 Use GLM-4.5 API services on Z.ai API Platform (Global) or Zhipu AI Open Platform (Mainland China) . 👉 One click to GLM-4.5 .

Open Source ↓ 729