AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,877 models Compare
🤖
llama-30b-instruct-2048
upstage

Developed by : Upstage Backbone Model : LLaMA Variations : It has different model parameter sizes and sequence lengths: 30B/1024, 30B/2048, 65B/1024 Language(s) : English Library : HuggingFace Transformers License : This model is under a Non-commercial Bespoke License and governe…

Open Source 30.0B ↓ 781
🤖
Emu3-Chat
BAAI

Emu3: Next-Token Prediction is All You Need

Open Source ↓ 770
🤖
stablelm-tuned-alpha-7b
stabilityai

StableLM-Tuned-Alpha is a suite of 3B and 7B parameter decoder-only language models built on top of the StableLM-Base-Alpha models and further fine-tuned on various chat and instruction-following datasets.

Open Source 7.0B ↓ 770
🤖
MiMo-VL-7B-SFT-2508
XiaomiMiMo

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ MiMo-VL Technical Report ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Multimodal 7.0B ↓ 769
🤖
Yarn-Mistral-7b-128k
NousResearch

Nous-Yarn-Mistral-7b-128k is a state-of-the-art language model for long context, further pretrained on long context data for 1500 steps using the YaRN extension method. It is an extension of Mistral-7B-v0.1 and supports a 128k token context window.

Open Source 7.0B ↓ 757
🤖
Gemma-2-Llama-Swallow-9b-it-v0.1
tokyotech-llm

Gemma-2-Llama-Swallow series was built by continual pre-training on the gemma-2 models. Gemma 2 Swallow enhanced the Japanese language capabilities of the original Gemma 2 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled…

Open Source 9.0B ↓ 751
🤖
Yi-Coder-1.5B-Chat
01-ai

🐙 GitHub • 👾 Discord • 🐤 Twitter • 💬 WeChat 📝 Paper • 💪 Tech Blog • 🙌 FAQ • 📗 Learning Hub

Code 1.5B ↓ 746
🤖
codet5-large
Salesforce

CodeT5 is a family of encoder-decoder language models for code from the paper: CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation by Yue Wang, Weishi Wang, Shafiq Joty, and Steven C.H. Hoi.

Code ↓ 746
🤖
sarashina2.2-1b-instruct-v0.1
sbintuitions

sbintuitions/sarashina2.2-1b-instruct-v0.1

Open Source 1.0B ↓ 745
🤖
GPT-OSS-Swallow-120B-RL-v0.1-MXFP4
tokyotech-llm

GPT-OSS-Swallow v0.1 is a family of large language models available in 20B and 120B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifiable Rewa…

Open Source 120.0B ↓ 744
🤖
Gemma-2-Llama-Swallow-2b-it-v0.1
tokyotech-llm

Gemma-2-Llama-Swallow series was built by continual pre-training on the gemma-2 models. Gemma 2 Swallow enhanced the Japanese language capabilities of the original Gemma 2 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled…

Open Source 2.0B ↓ 743
🤖
instructblip-vicuna-13b
Salesforce

InstructBLIP model using Vicuna-13b as language model. InstructBLIP was introduced in the paper InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning by Dai et al.

Open Source 13.0B ↓ 741