AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,877 models Compare
🤖
BaichuanMed-OCR-72B
baichuan-inc

BaichuanMed-OCR-72B is a model fine-tuned from the Qwen2.5-VL-72B-Instruct with our constructed and curated medical report datasets consists of medical report images and related questions and answers (QAs). It has been specifically adapted to perform Optical Character Recognition…

Open Source 72.0B ↓ 111
🤖
StripedHyena-Hessian-7B
togethercomputer

One of the focus areas at Together Research is new architectures for long context, improved training, and inference performance over the Transformer architecture. Spinning out of a research program from our team and academic collaborators, with roots in signal processing-inspired…

Open Source 7.0B ↓ 110
🤖
Llama-3.1-Swallow-8B-v0.2
tokyotech-llm

Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…

Open Source 8.0B ↓ 108
🤖
CAT-Paws-8B
cyberagent

CAT-Paws is an agentic LLM that thinks in Japanese (e.g., reasoning trace is in Japanese). The model is based on Qwen3-Swallow-v0.2 which is a continual pretraining model based on Qwen3 to read and write fluently in Japanese.

Open Source 8.0B ↓ 108
🤖
Swallow-MS-7b-instruct-v0.1
tokyotech-llm

Our Swallow-MS-7b-v0.1 model has undergone continual pre-training from the Mistral-7B-v0.1, primarily with the addition of Japanese language data.

Open Source 7.0B ↓ 106
🤖
Apriel-5B-Base
ServiceNow-AI

1. Model Summary 2. Evaluation 3. Intended Use 4. Limitations 5. Security and Responsible Use 6. License 7. Citation

Open Source 5.0B ↓ 101
🤖
nekomata-14b-instruction
rinna

Overview The model is the instruction-tuned version of rinna/nekomata-14b . It adopts the Alpaca input format.

Open Source 14.0B ↓ 100
🤖
Qwen-SEA-LION-v4-32B-IT-OV-4BIT
aisingapore

- Model creator: AI Singapore - Original model: Qwen-SEA-LION-v4-32B-IT

Open Source 32.0B ↓ 99
🤖
Swallow-MS-7b-v0.1
tokyotech-llm

Our Swallow-MS-7b-v0.1 model has undergone continual pre-training from the Mistral-7B-v0.1, primarily with the addition of Japanese language data.

Open Source 7.0B ↓ 97
🤖
Seed-Coder-8B-Reasoning-bf16
ByteDance-Seed

Introduction We are thrilled to introduce Seed-Coder, a powerful, transparent, and parameter-efficient family of open-source code models at the 8B scale, featuring base, instruct, and reasoning variants. Seed-Coder contributes to promote the evolution of open code models through…

Code 8.0B ↓ 96
🤖
Baichuan2-7B-Chat-4bits
baichuan-inc

🦉GitHub 💬WeChat 百川API支持搜索增强和192K长窗口,新增百川搜索增强知识库、限时免费! 🚀 百川大模型在线对话平台 已正式向公众开放 🎉

Open Source 7.0B ↓ 95
🤖
mt0-xxl-p3
bigscience

1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation

Open Source ↓ 95