AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,295 models for "Transformer" Compare
🤖
Apriel-5B-Base
ServiceNow-AI

1. Model Summary 2. Evaluation 3. Intended Use 4. Limitations 5. Security and Responsible Use 6. License 7. Citation

Open Source 5.0B ↓ 101
🤖
nekomata-14b-instruction
rinna

Overview The model is the instruction-tuned version of rinna/nekomata-14b . It adopts the Alpaca input format.

Open Source 14.0B ↓ 100
🤖
Swallow-MS-7b-v0.1
tokyotech-llm

Our Swallow-MS-7b-v0.1 model has undergone continual pre-training from the Mistral-7B-v0.1, primarily with the addition of Japanese language data.

Open Source 7.0B ↓ 97
🤖
Seed-Coder-8B-Reasoning-bf16
ByteDance-Seed

Introduction We are thrilled to introduce Seed-Coder, a powerful, transparent, and parameter-efficient family of open-source code models at the 8B scale, featuring base, instruct, and reasoning variants. Seed-Coder contributes to promote the evolution of open code models through…

Code 8.0B ↓ 96
🤖
Baichuan2-7B-Chat-4bits
baichuan-inc

🦉GitHub 💬WeChat 百川API支持搜索增强和192K长窗口,新增百川搜索增强知识库、限时免费! 🚀 百川大模型在线对话平台 已正式向公众开放 🎉

Open Source 7.0B ↓ 95
🤖
mt0-xxl-p3
bigscience

1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation

Open Source ↓ 95
🤖
gemma-2-baku-2b
rinna

We conduct continual pre-training of google/gemma-2-2b on 80B tokens from a mixture of Japanese and English datasets. The continual pre-training improves the model's performance on Japanese tasks.

Open Source 2.0B ↓ 93
🤖
mistral-coreml
apple

[!IMPORTANT] ❗ This repo requires the use of the macOS Sequoia (15) Developer Beta to utilize the latest and greatest CoreML has to offer! Sign up for the Apple Beta Software Program here to get access. Check out the companion blog post to learn more about what's new in iOS 18 &…

Open Source ↓ 93
🤖
vicuna-13b-delta-finetuned-langchain-MRKL
rinna

NOTE: This "delta model" cannot be used directly. Users have to apply it on top of the original LLaMA weights to get actual vicuna-13b-finetuned-langchain-MRKL weights. See https://github.com/rinnakk/vicuna-13b-delta-finetuned-langchain-MRKL model-weights for instructions.

Open Source 13.0B ↓ 92
🤖
LongCat-Flash-Thinking-ZigZag
meituan-longcat

[2026.1.28] We have provided the TileLang kernels supporting prefill (chunked-prefill as well) and decode (multi-token prediction as well). The full attention version is placed at flash mla interface.py while the streaming sparse attention version is placed at streaming sparse at…

Open Source ↓ 91
🤖
bilingual-gpt-neox-4b-instruction-ppo
rinna

Overview This repository provides an English-Japanese bilingual GPT-NeoX model of 3.8 billion parameters.

Open Source 4.0B ↓ 89
🤖
gemma-2-baku-2b-it
rinna

Gemma 2 Baku 2B Instruct (rinna/gemma-2-baku-2b-it)

Open Source 2.0B ↓ 88