AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,877 models Compare
🤖
TinySolar-248m-4k-py-instruct
upstage

Used dataset for fine-tuning - sahil2801/CodeAlpaca-20k - m-a-p/CodeFeedback-Filtered-Instruction

Open Source ↓ 94
🤖
gemma-2-baku-2b
rinna

We conduct continual pre-training of google/gemma-2-2b on 80B tokens from a mixture of Japanese and English datasets. The continual pre-training improves the model's performance on Japanese tasks.

Open Source 2.0B ↓ 93
🤖
mistral-coreml
apple

[!IMPORTANT] ❗ This repo requires the use of the macOS Sequoia (15) Developer Beta to utilize the latest and greatest CoreML has to offer! Sign up for the Apple Beta Software Program here to get access. Check out the companion blog post to learn more about what's new in iOS 18 &…

Open Source ↓ 93
🤖
vicuna-13b-delta-finetuned-langchain-MRKL
rinna

NOTE: This "delta model" cannot be used directly. Users have to apply it on top of the original LLaMA weights to get actual vicuna-13b-finetuned-langchain-MRKL weights. See https://github.com/rinnakk/vicuna-13b-delta-finetuned-langchain-MRKL model-weights for instructions.

Open Source 13.0B ↓ 92
🤖
LongCat-Flash-Thinking-ZigZag
meituan-longcat

[2026.1.28] We have provided the TileLang kernels supporting prefill (chunked-prefill as well) and decode (multi-token prediction as well). The full attention version is placed at flash mla interface.py while the streaming sparse attention version is placed at streaming sparse at…

Open Source ↓ 91
🤖
bilingual-gpt-neox-4b-instruction-ppo
rinna

Overview This repository provides an English-Japanese bilingual GPT-NeoX model of 3.8 billion parameters.

Open Source 4.0B ↓ 89
🤖
gemma-2-baku-2b-it
rinna

Gemma 2 Baku 2B Instruct (rinna/gemma-2-baku-2b-it)

Open Source 2.0B ↓ 88
🤖
LongCat-Flash-Chat-FP8
meituan-longcat

Model Introduction We introduce LongCat-Flash, a powerful and efficient language model with 560 billion total parameters, featuring an innovative Mixture-of-Experts (MoE) architecture. The model incorporates a dynamic computation mechanism that activates 18.6B∼31.3B parameters (a…

Open Source ↓ 88
🤖
cudaLLM-8B
ByteDance-Seed

CudaLLM: A Language Model for High-Performance CUDA Kernel Generation

Open Source 8.0B ↓ 86
🤖
Llama-2-7b
meta-llama

Open Source 7.0B ↓ 85
🤖
calm2-7b-chat-dpo-experimental
cyberagent

Model Card for "calm2-7b-chat-dpo-experimental"

Open Source 7.0B ↓ 84
🤖
SuperApriel-15b-Base
ServiceNow-AI

A 15B-parameter token-mixer supernet derived from Apriel-1.6 via stochastic distillation. Every decoder layer exposes four trained mixer options —Full Attention, Sliding Window Attention, Gated DeltaNet, and Kimi Delta Attention—enabling flexible architecture selection from a sin…

Open Source 15.0B ↓ 83