AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,294 models for "Transformer" Compare
🤖
SuperApriel-15b-Base
ServiceNow-AI

A 15B-parameter token-mixer supernet derived from Apriel-1.6 via stochastic distillation. Every decoder layer exposes four trained mixer options —Full Attention, Sliding Window Attention, Gated DeltaNet, and Kimi Delta Attention—enabling flexible architecture selection from a sin…

Open Source 15.0B ↓ 83
🤖
Swallow-70b-NVE-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 70.0B ↓ 83
🤖
Llama-3.1-Swallow-70B-Instruct-v0.1
tokyotech-llm

Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…

Open Source 70.0B ↓ 81
🤖
BFS-Prover-V2-32B
ByteDance-Seed

BFS-Prover-V2: Scaling up Multi-Turn Off-Policy RL and Multi-Agent Tree Search for LLM Step-Provers

Open Source 32.0B ↓ 81
🤖
youri-7b-chat
rinna

Overview The model is the instruction-tuned version of rinna/youri-7b . It adopts a chat-style input format.

Open Source 7.0B ↓ 79
🤖
DeepSeek-R1-Distill-Qwen-14B-Japanese
cyberagent

This is a Japanese finetuned model based on deepseek-ai/DeepSeek-R1-Distill-Qwen-14B.

Reasoning 14.0B ↓ 77
🤖
sage-ft-mixtral-8x7b
apple

Authors : Yizhe Zhang, Navdeep Jaitly (Apple)

Open Source 7.0B ↓ 75
🤖
Swallow-70b-instruct-v0.1
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 70.0B ↓ 74
🤖
diafill-sarashina2.2-3b-instruct
sbintuitions

Model Summary DiaFill is a Japanese dialogue script generation model designed to produce natural, spoken-style dialogue scripts rich in fillers and brief utternaces. Unlike typical assistant models that respond to users, this model is fine-tuned to generate a multi-turn dialogue…

Open Source 3.0B ↓ 74
🤖
nekomata-7b-instruction
rinna

Overview The model is the instruction-tuned version of rinna/nekomata-7b . It adopts the Alpaca input format.

Open Source 7.0B ↓ 73
🤖
Apriel-H1-15b-Thinker-SFT
ServiceNow-AI

A 15B-parameter hybrid reasoning model combining Transformer attention and Mamba State Space layers for high efficiency and scalability. Derived from Apriel-Nemotron-15B-Thinker through progressive distillation, Apriel-H1 replaces less critical attention layers with linear Mamba…

Open Source 15.0B ↓ 73
🤖
AHN-GDN-for-Qwen-2.5-Instruct-3B
ByteDance-Seed

AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling

Open Source 3.0B ↓ 73