AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,296 models for "Transformer" Compare
🤖
Medical-GPT-OSS-Swallow-120B
tokyotech-llm

Medical-GPT-OSS-Swallow-120B is a medical-domain language model based on tokyotech-llm/GPT-OSS-Swallow-120B-RL-v0.1. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.

Open Source 120.0B ↓ 175
🤖
llama-65b-instruct
upstage

Developed by : Upstage Backbone Model : LLaMA Variations : It has different model parameter sizes and sequence lengths: 30B/1024, 30B/2048, 65B/1024 Language(s) : English Library : HuggingFace Transformers License : This model is under a Non-commercial Bespoke License and governe…

Open Source 65.0B ↓ 175
🤖
internlm2-chat-1_8b-sft
internlm

💻Github Repo • 🤔Reporting Issues • 📜Technical Report

Open Source 8.0B ↓ 173
🤖
octogeex
bigcode

1. Model Summary 2. Use 3. Training 4. License 5. Citation

Open Source ↓ 173
🤖
Ring-lite
inclusionAI

Ring-lite is a lightweight, fully open-sourced MoE (Mixture of Experts) LLM designed for complex reasoning tasks. It is built upon the publicly available Ling-lite-1.5 model, which has 16.8B parameters with 2.75B activated parameters.. We use a joint training pipeline combining k…

Open Source ↓ 172
🤖
SimpleSD-30B-instruct
apple

This model is an example of the Simple Self-Distillation (SimpleSD) method that improves code generation by fine-tuning a language model on its own sampled outputs—without rewards, verifiers, teacher models, or reinforcement learning. Please see the paper below for more informati…

Open Source 30.0B ↓ 172
🤖
AHN-Mamba2-for-Qwen-2.5-Instruct-14B
ByteDance-Seed

AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling

Open Source 14.0B ↓ 172
🤖
Step-3.5-Flash-Base-Midtrain
stepfun-ai

Step 3.5 Flash (visit website) is our most capable open-source foundation model, engineered to deliver frontier reasoning and agentic capabilities with exceptional efficiency. We also open-sourced the training codebase (SteptronOss), with support for continue pretrain, SFT, RL (W…

Open Source ↓ 168
🤖
SmolVLM2-2.2B-Base
HuggingFaceTB

This is the base model for SmolVLM2-2.2B, a lightweight multimodal model designed to analyze video content. The model processes videos, images, and text inputs to generate text outputs - whether answering questions about media files, comparing visual content, or transcribing text…

Multimodal 2.2B ↓ 167
🤖
OREAL-7B-SFT
internlm

--- license: apache-2.0 library name: transformers base model: - Qwen/Qwen2.5-7B pipeline tag: text-generation ---

Open Source 7.0B ↓ 165
🤖
SmolLM2-1.7B-Instruct-16k
HuggingFaceTB

This is a 16k context version of SmolLM2-1.7B-Instruct, which originnaly only supported 8k context. We finetune the model on 15k samples consisting of a subset of SmolTalk, LongAlign and SeaLong datasets and increase RoPE from 100k to 500k. This improves the evaluation on HELMET…

Open Source 1.7B ↓ 164
🤖
GPT-JT-6B-v1
togethercomputer

With a new decentralized training algorithm, we fine-tuned GPT-J (6B) on 3.53 billion tokens, resulting in GPT-JT (6B), a model that outperforms many 100B+ parameter models on classification benchmarks.

Open Source 6.0B ↓ 164