AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,296 models for "Transformer" Compare
🤖
llama-3-youko-8b-instruct
rinna

Llama 3 Youko 8B Instruct (rinna/llama-3-youko-8b-instruct)

Open Source 8.0B ↓ 163
🤖
StepFun-Formalizer-32B
stepfun-ai

StepFun-Formalizer: Unlocking the Autoformalization Potential of LLMs through Knowledge-Reasoning Fusion

Open Source 32.0B ↓ 163
🤖
nekomata-14b
rinna

Overview We conduct continual pre-training of qwen-14b on 66B tokens from a mixture of Japanese and English datasets. The continual pre-training significantly improves the model's performance on Japanese tasks. It also enjoys the following great features provided by the original…

Open Source 14.0B ↓ 159
🤖
internlm2-math-7b
internlm

State-of-the-art bilingual open-sourced Math reasoning LLMs. A solver , prover , verifier , augmentor .

Open Source 7.0B ↓ 156
🤖
japanese-stablelm-base-beta-7b
stabilityai

A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL

Open Source 7.0B ↓ 154
🤖
bloomz-p3
bigscience

1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation

Open Source ↓ 148
🤖
smollm-360M-instruct-add-basics
HuggingFaceTB

Model Summary Chat with the model at: https://huggingface.co/spaces/HuggingFaceTB/instant-smol

Open Source ↓ 148
🤖
AHN-Mamba2-for-Qwen-2.5-Instruct-7B
ByteDance-Seed

AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling

Open Source 7.0B ↓ 148
🤖
internlm2-chat-20b-sft
internlm

💻Github Repo • 🤔Reporting Issues • 📜Technical Report

Open Source 20.0B ↓ 147
🤖
deepseekcoder-33b-codeqwen-align-subset
bigcode

This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.

Code 33.0B ↓ 147
🤖
LongCat-Flash-Prover
meituan-longcat

We introduce LongCat-Flash-Prover , a flagship $560$-billion-parameter open-source Mixture-of-Experts (MoE) model that advances Native Formal Reasoning in Lean4 through agentic tool-integrated reasoning (TIR). We decompose the native formal reasoning task into three independent f…

Open Source ↓ 146
🤖
SmolVLM-Synthetic
HuggingFaceTB

SmolVLM is a compact open multimodal model that accepts arbitrary sequences of image and text inputs to produce text outputs. Designed for efficiency, SmolVLM can answer questions about images, describe visual content, create stories grounded on multiple images, or function as a…

Multimodal ↓ 146