AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,308 models for "Transformer" Compare
🤖
EXAONE-Deep-7.8B
LGAI-EXAONE

We introduce EXAONE Deep, which exhibits superior capabilities in various reasoning tasks including math and coding benchmarks, ranging from 2.4B to 32B parameters developed and released by LG AI Research. Evaluation results show that 1) EXAONE Deep 2.4B outperforms other models…

Open Source 7.8B ↓ 984
🤖
StepFun-Formalizer-7B
stepfun-ai

StepFun-Formalizer: Unlocking the Autoformalization Potential of LLMs through Knowledge-Reasoning Fusion

Open Source 7.0B ↓ 984
🤖
blip2-opt-6.7b-coco
Salesforce

BLIP-2 model, leveraging OPT-6.7b (a large language model with 6.7 billion parameters). It was introduced in the paper BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models by Li et al. and first released in this repository.

Open Source 6.7B ↓ 968
🤖
nanowhale-100m
HuggingFaceTB

A small ~110M parameter language model implementing the DeepSeek-V4 architecture , fine-tuned for chat/instruction following. Trained from scratch — no weights from DeepSeek-V4 were used.

Open Source ↓ 939
🤖
codegen-6B-mono
Salesforce

CodeGen is a family of autoregressive language models for program synthesis from the paper: A Conversational Paradigm for Program Synthesis by Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, Caiming Xiong. The models are originally releas…

Code 6.0B ↓ 929
🤖
Meta-Llama-3.1-70B
NousResearch

The Meta Llama 3.1 collection of multilingual large language models (LLMs) is a collection of pretrained and instruction tuned generative models in 8B, 70B and 405B sizes (text in/text out). The Llama 3.1 instruction tuned text only models (8B, 70B, 405B) are optimized for multil…

Open Source 70.0B ↓ 927
🤖
UI-Venus-Ground-7B
inclusionAI

UI-Venus This repository contains the UI-Venus model from the report UI-Venus Technical Report: Building High-performance UI Agents with RFT.

Open Source 7.0B ↓ 909
🤖
Gemma2-9b-WangchanLIONv2-instruct
aisingapore

WangchanLION is a joint effort between VISTEC and AI Singapore to develop a Thai-specific collection of Large Language Models (LLMs), pre-trained for Southeast Asian (SEA) languages, and instruct-tuned specifically for the Thai language.

Open Source 9.0B ↓ 905
🤖
InternVL3_5-1B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 1.0B ↓ 882
🤖
bloomz
bigscience

1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation

Open Source ↓ 867
🤖
OpenELM-1_1B
apple

Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao, Maxwell Horton, Yanzi Jin, Chenfan Sun, Iman Mirzadeh, Mahyar Najibi, Dmitry Belenko, Peter Zatloukal, Mohammad Rastegari

Open Source 1.0B ↓ 864
🤖
Gemma-SEA-LION-v4-27B-IT
aisingapore

SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

Open Source 27.0B ↓ 859