AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,313 models for "Transformer" Compare
🤖
falcon-11B
tiiuae

Falcon2-11B is an 11B parameters causal decoder-only model built by TII and trained on over 5,000B tokens of RefinedWeb enhanced with curated corpora. The model is made available under the TII Falcon License 2.0, the permissive Apache 2.0-based software license which includes an…

Open Source 11.0B ↓ 3.5K
🤖
InternVL3_5-38B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 38.0B ↓ 3.2K
🤖
InternVL3_5-241B-A28B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 241.0B ↓ 3.2K
🤖
InternVL2_5-8B-MPO-hf
OpenGVLab

This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.

Multimodal 8.0B ↓ 3.2K
🤖
Llama-3.1-Swallow-8B-Instruct-v0.2
tokyotech-llm

Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…

Open Source 8.0B ↓ 3.1K
🤖
LFM2-VL-3B
LiquidAI

LFM2-VL-3B is the newest and most capable model in Liquid AI's multimodal LFM2-VL series, designed to process text and images with variable resolutions. Built on the LFM2 backbone, it extends the architecture for higher-capacity reasoning and stronger visual understanding while r…

Multimodal 3.0B ↓ 3.1K
🤖
japanese-gpt2-small
rinna

This repository provides a small-sized Japanese GPT-2 model. The model was trained using code from Github repository rinnakk/japanese-pretrained-models by rinna Co., Ltd.

Open Source ↓ 3.1K
🤖
Nous-Hermes-Llama2-13b
NousResearch

Compute provided by our project sponsor Redmond AI, thank you! Follow RedmondAI on Twitter @RedmondAI.

Open Source 13.0B ↓ 3.1K
🤖
Youtu-VL-4B-Instruct
tencent

🏠 Project Page • 📃 License • 💻 Code • 📑 Technical Report • 📊 Benchmarks • 🚀 Getting Started

Multimodal 4.0B ↓ 3.1K
🤖
InternVL3-38B-hf
OpenGVLab

InternVL3-38B Transformers 🤗 Implementation

Multimodal 38.0B ↓ 3.1K
🤖
MiniCPM-V-4_5-int4
openbmb

A GPT-4o Level MLLM for Single Image, Multi Image and High-FPS Video Understanding on Your Phone

Open Source ↓ 3.1K
🤖
internlm2_5-7b
internlm

💻Github Repo • 🤔Reporting Issues • 📜Technical Report

Open Source 7.0B ↓ 3K