AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,307 models for "Transformer" Compare
🤖
Yarn-Mistral-7b-128k
NousResearch

Nous-Yarn-Mistral-7b-128k is a state-of-the-art language model for long context, further pretrained on long context data for 1500 steps using the YaRN extension method. It is an extension of Mistral-7B-v0.1 and supports a 128k token context window.

Open Source 7.0B ↓ 757
🤖
Gemma-2-Llama-Swallow-9b-it-v0.1
tokyotech-llm

Gemma-2-Llama-Swallow series was built by continual pre-training on the gemma-2 models. Gemma 2 Swallow enhanced the Japanese language capabilities of the original Gemma 2 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled…

Open Source 9.0B ↓ 751
🤖
Yi-Coder-1.5B-Chat
01-ai

🐙 GitHub • 👾 Discord • 🐤 Twitter • 💬 WeChat 📝 Paper • 💪 Tech Blog • 🙌 FAQ • 📗 Learning Hub

Code 1.5B ↓ 746
🤖
codet5-large
Salesforce

CodeT5 is a family of encoder-decoder language models for code from the paper: CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation by Yue Wang, Weishi Wang, Shafiq Joty, and Steven C.H. Hoi.

Code ↓ 746
🤖
sarashina2.2-1b-instruct-v0.1
sbintuitions

sbintuitions/sarashina2.2-1b-instruct-v0.1

Open Source 1.0B ↓ 745
🤖
Gemma-2-Llama-Swallow-2b-it-v0.1
tokyotech-llm

Gemma-2-Llama-Swallow series was built by continual pre-training on the gemma-2 models. Gemma 2 Swallow enhanced the Japanese language capabilities of the original Gemma 2 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled…

Open Source 2.0B ↓ 743
🤖
instructblip-vicuna-13b
Salesforce

InstructBLIP model using Vicuna-13b as language model. InstructBLIP was introduced in the paper InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning by Dai et al.

Open Source 13.0B ↓ 741
🤖
MiMo-7B-SFT
XiaomiMiMo

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Unlocking the Reasoning Potential of Language Model From Pretraining to Posttraining ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Open Source 7.0B ↓ 737
🤖
InternVL3_5-8B-Pretrained
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 8.0B ↓ 735
🤖
FastVLM-7B
apple

FastVLM: Efficient Vision Encoding for Vision Language Models

Multimodal 7.0B ↓ 732
🤖
sarashina2-7b
sbintuitions

This repository provides large language models trained by SB Intuitions.

Open Source 7.0B ↓ 731
🤖
GLM-4.5-Base
zai-org

👋 Join our Discord community. 📖 Check out the GLM-4.5 technical blog , technical report , and Zhipu AI technical documentation . 📍 Use GLM-4.5 API services on Z.ai API Platform (Global) or Zhipu AI Open Platform (Mainland China) . 👉 One click to GLM-4.5 .

Open Source ↓ 729