AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

440 models for "Base Model" Compare
bloom-560m-intermediate logo
bloom-560m-intermediate
bigscience

WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).

Open Source ↓ 225
sage-ft-mixtral-8x7b logo
sage-ft-mixtral-8x7b
apple

Authors : Yizhe Zhang, Navdeep Jaitly (Apple)

Open Source 7.0B ↓ 224
StripedHyena-Hessian-7B logo
StripedHyena-Hessian-7B
togethercomputer

One of the focus areas at Together Research is new architectures for long context, improved training, and inference performance over the Transformer architecture. Spinning out of a research program from our team and academic collaborators, with roots in signal processing-inspired…

Open Source 7.0B ↓ 218
Swallow-70b-instruct-hf logo
Swallow-70b-instruct-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 70.0B ↓ 216
bloom-1b1-intermediate logo
bloom-1b1-intermediate
bigscience

WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).

Open Source ↓ 215
ERNIE-4.5-VL-28B-A3B-Base-PT logo
ERNIE-4.5-VL-28B-A3B-Base-PT
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 28.0B ↓ 212
gemma-2-baku-2b-it logo
gemma-2-baku-2b-it
rinna

Gemma 2 Baku 2B Instruct (rinna/gemma-2-baku-2b-it)

Open Source 2.0B ↓ 206
Gemma-2-Llama-Swallow-27b-pt-v0.1 logo
Gemma-2-Llama-Swallow-27b-pt-v0.1
tokyotech-llm

Gemma-2-Llama-Swallow series was built by continual pre-training on the gemma-2 models. Gemma 2 Swallow enhanced the Japanese language capabilities of the original Gemma 2 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled…

Open Source 27.0B ↓ 200
Medical-Qwen3-Swallow-32B logo
Medical-Qwen3-Swallow-32B
tokyotech-llm

Medical-Qwen3-Swallow-32B is a medical-domain language model based on tokyotech-llm/Qwen3-Swallow-32B-RL-v0.2. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.

Open Source 32.0B ↓ 190
LongCat-Flash-Chat-FP8 logo
LongCat-Flash-Chat-FP8
meituan-longcat

Model Introduction We introduce LongCat-Flash, a powerful and efficient language model with 560 billion total parameters, featuring an innovative Mixture-of-Experts (MoE) architecture. The model incorporates a dynamic computation mechanism that activates 18.6B∼31.3B parameters (a…

Open Source ↓ 185
Swallow-MS-7b-v0.1 logo
Swallow-MS-7b-v0.1
tokyotech-llm

Our Swallow-MS-7b-v0.1 model has undergone continual pre-training from the Mistral-7B-v0.1, primarily with the addition of Japanese language data.

Open Source 7.0B ↓ 184
Redmond-Hermes-Coder logo
Redmond-Hermes-Coder
NousResearch

Redmond-Hermes-Coder 15B is a state-of-the-art language model fine-tuned on over 300,000 instructions. This model was fine-tuned by Nous Research, with Teknium and Karan4D leading the fine tuning process and dataset curation, Redmond AI sponsoring the compute, and several other c…

Code ↓ 178