AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

362 models for "Fine-tuned" Compare
deepseek-coder-7b-instruct-v1.5 logo
deepseek-coder-7b-instruct-v1.5
deepseek-ai

[🏠Homepage] [🤖 Chat with DeepSeek Coder] [Discord] [Wechat(微信)]

Code 7.0B ↓ 642.4K
phi-2 logo
phi-2
microsoft

Phi-2 is a Transformer with 2.7 billion parameters. It was trained using the same data sources as Phi-1.5, augmented with a new data source that consists of various NLP synthetic texts and filtered websites (for safety and educational value). When assessed against benchmarks test…

Open Source ↓ 621.8K
blip2-opt-2.7b logo
blip2-opt-2.7b
Salesforce

BLIP-2 model, leveraging OPT-2.7b (a large language model with 2.7 billion parameters). It was introduced in the paper BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models by Li et al. and first released in this repository.

Open Source 2.7B ↓ 606.6K
bloom-560m logo
bloom-560m
bigscience

BLOOM LM BigScience Large Open-science Open-access Multilingual Language Model Model Card

Open Source 0.56B ↓ 459.2K
DeepSeek-R1-Distill-Qwen-32B logo
DeepSeek-R1-Distill-Qwen-32B
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning 32.0B ↓ 423.3K
Llama-3.1-70B-Instruct logo
Llama-3.1-70B-Instruct
meta-llama

Open Source 70.0B ↓ 391.2K
Florence-2-large logo
Florence-2-large
microsoft

Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks

Multimodal 0.77B ↓ 387.6K
Phi-3-mini-4k-instruct logo
Phi-3-mini-4k-instruct
microsoft

🎉 Phi-3.5 : [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)

Open Source ↓ 366.6K
Llama-2-7b-chat-hf logo
Llama-2-7b-chat-hf
meta-llama

Open Source 7.0B ↓ 347.5K
DeepSeek-R1-Distill-Qwen-14B logo
DeepSeek-R1-Distill-Qwen-14B
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning 14.0B ↓ 332K
granite-3.3-2b-instruct logo
granite-3.3-2b-instruct
ibm-granite

Model Summary: Granite-3.3-2B-Instruct is a 2-billion parameter 128K context length language model fine-tuned for improved reasoning and instruction-following capabilities. Built on top of Granite-3.3-2B-Base, the model delivers significant gains on benchmarks for measuring gener…

Open Source 2.0B ↓ 328.7K
OLMo-2-0425-1B logo
OLMo-2-0425-1B
allenai

We introduce OLMo 2 1B, the smallest model in the OLMo 2 family. OLMo 2 was pre-trained on OLMo-mix-1124 and uses Dolmino-mix-1124 for mid-training.

Open Source 1.0B ↓ 314.9K