AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,309 models for "Transformer" Compare
🤖
Step-3.5-Flash-FP8
stepfun-ai

Step 3.5 Flash (visit website) is our most capable open-source foundation model, engineered to deliver frontier reasoning and agentic capabilities with exceptional efficiency. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B p…

Open Source ↓ 1.5K
🤖
Intern-S2-Mobius
internlm

💻Github Repo • 🤗Model Collections • 🌳Arch Space

Open Source ↓ 1.5K
🤖
glm-4-9b-chat-1m-hf
zai-org

If you are using the weights from this repository, please update to

Open Source 9.0B ↓ 1.5K
🤖
InternVL3_5-14B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 14.0B ↓ 1.5K
🤖
open-calm-small
cyberagent

OpenCALM is a suite of decoder-only language models pre-trained on Japanese datasets, developed by CyberAgent, Inc.

Open Source ↓ 1.5K
🤖
Llama-SEA-LION-v3.5-8B-R
aisingapore

SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

Open Source 8.0B ↓ 1.5K
🤖
ERNIE-4.5-0.3B-Base-PT
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 0.3B ↓ 1.5K
🤖
InternVL2_5-8B-MPO
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 8.0B ↓ 1.4K
🤖
glm-edge-v-2b
zai-org

Install the transformers library from the source code:

Open Source 2.0B ↓ 1.4K
🤖
internlm-xcomposer2-vl-7b-4bit
internlm

InternLM-XComposer2 is a vision-language large model (VLLM) based on InternLM2 for advanced text-image comprehension and composition.

Multimodal 7.0B ↓ 1.4K
🤖
SmolVLM-500M-Base
HuggingFaceTB

This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.

Multimodal ↓ 1.3K
🤖
InternVL3-1B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 1.0B ↓ 1.3K