AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,313 models for "Transformer" Compare
🤖
InternVL3-14B-AWQ
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 14.0B ↓ 3K
🤖
LongCat-Flash-Thinking-2601
meituan-longcat

We introduce an updated version of LongCat-Flash-Thinking, a powerful and efficient Large Reasoning Model (LRM) with 560 billion total parameters, built upon an innovative Mixture-of-Experts (MoE) architecture. Beyond inheriting the domain-parallel training recipe in our previous…

Open Source ↓ 3K
🤖
DeepSeek-Coder-V2-Base
deepseek-ai

DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

Code ↓ 2.9K
🤖
GLM-4.5V-FP8
zai-org

👋 Join our Discord communities. 📖 Check out the paper . 📍 Access the GLM-V series models via API on the ZhipuAI Open Platform .

Open Source ↓ 2.9K
🤖
stablelm-2-12b
stabilityai

Stable LM 2 12B is a 12.1 billion parameter decoder-only language model pre-trained on 2 trillion tokens of diverse multilingual and code datasets for two epochs.

Open Source 12.0B ↓ 2.8K
🤖
MiniMax-Text-01
MiniMaxAI

MiniMax-Text-01 is a powerful language model with 456 billion total parameters, of which 45.9 billion are activated per token. To better unlock the long context capabilities of the model, MiniMax-Text-01 adopts a hybrid architecture that combines Lightning Attention, Softmax Atte…

Open Source ↓ 2.8K
🤖
Nous-Hermes-2-Mixtral-8x7B-SFT
NousResearch

Nous Hermes 2 Mixtral 8x7B SFT is the supervised finetune only version of our new flagship Nous Research model trained over the Mixtral 8x7B MoE LLM.

Open Source 7.0B ↓ 2.8K
🤖
Ling-1T
inclusionAI

--- license: mit pipeline tag: text-generation library name: transformers ---

Open Source ↓ 2.8K
🤖
Hermes-3-Llama-3.1-70B
NousResearch

Hermes 3 is the latest version of our flagship Hermes series of LLMs by Nous Research.

Open Source 70.0B ↓ 2.8K
🤖
MolmoPoint-8B
allenai

MolmoPoint-8B MolmoPoint-8B is a fully-open VLM developed by the Allen Institute for AI (Ai2) that support image, video and multi-image understanding and grounding. It has new pointing mechansim that improves image pointing, video pointing, and video tracking, see our technical r…

Open Source 8.0B ↓ 2.7K
🤖
tiny-lm-chat
sbintuitions

This repository provides a tiny 16M parameters language model for debugging and testing purposes. This is created by tuning sbintuitions/tiny-lm with oasset1 datasets in Japanese and English.

Open Source ↓ 2.7K
🤖
internlm2-chat-7b-sft
internlm

💻Github Repo • 🤔Reporting Issues • 📜Technical Report

Open Source 7.0B ↓ 2.7K