AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

362 models for "Fine-tuned" Compare
GPT-JT-Moderation-6B logo
GPT-JT-Moderation-6B
togethercomputer

This model card introduces a moderation model, a GPT-JT model fine-tuned on Ontocord.ai's OIG-moderation dataset v0.1.

Open Source 6.0B ↓ 216
bloom-1b1-intermediate logo
bloom-1b1-intermediate
bigscience

WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).

Open Source ↓ 215
ERNIE-4.5-VL-28B-A3B-Base-PT logo
ERNIE-4.5-VL-28B-A3B-Base-PT
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 28.0B ↓ 212
Qwen-SEA-LION-v4-32B-IT-4BIT logo
Qwen-SEA-LION-v4-32B-IT-4BIT
aisingapore

Qwen-SEA-LION-v4-32B-IT-4BIT (GPTQ model)

Open Source 32.0B ↓ 204
ERNIE-4.5-21B-A3B-Paddle logo
ERNIE-4.5-21B-A3B-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 21.0B ↓ 194
LongCat-Flash-Chat-FP8 logo
LongCat-Flash-Chat-FP8
meituan-longcat

Model Introduction We introduce LongCat-Flash, a powerful and efficient language model with 560 billion total parameters, featuring an innovative Mixture-of-Experts (MoE) architecture. The model incorporates a dynamic computation mechanism that activates 18.6B∼31.3B parameters (a…

Open Source ↓ 185
Redmond-Hermes-Coder logo
Redmond-Hermes-Coder
NousResearch

Redmond-Hermes-Coder 15B is a state-of-the-art language model fine-tuned on over 300,000 instructions. This model was fine-tuned by Nous Research, with Teknium and Karan4D leading the fine tuning process and dataset curation, Redmond AI sponsoring the compute, and several other c…

Code ↓ 178
SOLAR-0-70b-8bit logo
SOLAR-0-70b-8bit
upstage

This is a 8bit quantized version of upstage/SOLAR-0-70b-16bit

Open Source 70.0B ↓ 176
AquilaDense-16B logo
AquilaDense-16B
BAAI

AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies Language Foundation Model & Software Team Beijing Academy of Artificial Intelligence (BAAI) [Paper(released soon)] [Code] [github]

Open Source 16.0B ↓ 176
BaichuanMed-OCR-7B logo
BaichuanMed-OCR-7B
baichuan-inc

BaichuanMed-OCR-7B is a model fine-tuned from the Qwen2.5-VL-7B-Instruct with our constructed and curated medical report datasets consists of medical report images and related questions and answers (QAs). It has been specifically adapted to perform Optical Character Recognition (…

Open Source 7.0B ↓ 176
ERNIE-4.5-VL-28B-A3B-Paddle logo
ERNIE-4.5-VL-28B-A3B-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 28.0B ↓ 172
bloom-7b1-intermediate logo
bloom-7b1-intermediate
bigscience

WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).

Open Source ↓ 172