AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

362 models for "Fine-tuned" Compare
calm3-22b-chat logo
calm3-22b-chat
cyberagent

CyberAgentLM3 is a decoder-only language model pre-trained on 2.0 trillion tokens from scratch. CyberAgentLM3-Chat is a fine-tuned model specialized for dialogue use cases.

Open Source 22.0B ↓ 664
LLaDA2.0-flash logo
LLaDA2.0-flash
inclusionAI

LLaDA2.0-flash is a diffusion language model featuring a 100BA6B Mixture-of-Experts (MoE) architecture. As an enhanced, instruction-tuned iteration of the LLaDA2.0 series, it is optimized for practical applications.

Open Source ↓ 658
Apriel-1.5-15b-Thinker logo
Apriel-1.5-15b-Thinker
ServiceNow-AI

Apriel-1.5-15b-Thinker - Mid training is all you need!

Open Source 15.0B ↓ 635
Falcon-E-1B-Instruct logo
Falcon-E-1B-Instruct
tiiuae

0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation

Open Source 1.0B ↓ 615
Nous-Hermes-13b logo
Nous-Hermes-13b
NousResearch

Nous-Hermes-13b is a state-of-the-art language model fine-tuned on over 300,000 instructions. This model was fine-tuned by Nous Research, with Teknium and Karan4D leading the fine tuning process and dataset curation, Redmond AI sponsoring the compute, and several other contributo…

Open Source 13.0B ↓ 614
nanowhale-100m logo
nanowhale-100m
HuggingFaceTB

A small ~110M parameter language model implementing the DeepSeek-V4 architecture , fine-tuned for chat/instruction following. Trained from scratch — no weights from DeepSeek-V4 were used.

Open Source ↓ 613
ERNIE-4.5-21B-A3B-Base-PT logo
ERNIE-4.5-21B-A3B-Base-PT
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 21.0B ↓ 611
japanese-stablelm-base-ja_vocab-beta-7b logo
japanese-stablelm-base-ja_vocab-beta-7b
stabilityai

A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL

Open Source 7.0B ↓ 568
japanese-stablelm-base-beta-70b logo
japanese-stablelm-base-beta-70b
stabilityai

A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL

Open Source 70.0B ↓ 547
japanese-stablelm-base-beta-7b logo
japanese-stablelm-base-beta-7b
stabilityai

A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL

Open Source 7.0B ↓ 538
ERNIE-4.5-0.3B-Paddle logo
ERNIE-4.5-0.3B-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 0.3B ↓ 520
GPT-JT-6B-v1 logo
GPT-JT-6B-v1
togethercomputer

With a new decentralized training algorithm, we fine-tuned GPT-J (6B) on 3.53 billion tokens, resulting in GPT-JT (6B), a model that outperforms many 100B+ parameter models on classification benchmarks.

Open Source 6.0B ↓ 506