AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
instructblip-vicuna-13b logo
instructblip-vicuna-13b
Salesforce

InstructBLIP model using Vicuna-13b as language model. InstructBLIP was introduced in the paper InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning by Dai et al.

Open Source 13.0B ↓ 377
JudgeLM-7B-v1.0 logo
JudgeLM-7B-v1.0
BAAI

--- inference: false language: - en tags: - instruction-finetuning pretty name: JudgeLM-100K task categories: - text-generation ---

Open Source 7.0B ↓ 374
GTA1-7B-2507 logo
GTA1-7B-2507
Salesforce

Reinforcement learning (RL) (e.g., GRPO) helps with grounding because of its inherent objective alignment—rewarding successful clicks—rather than encouraging long textual Chain-of-Thought (CoT) reasoning. Unlike approaches that rely heavily on verbose CoT reasoning, GRPO directly…

Multimodal 7.0B ↓ 372
CapRL-3B logo
CapRL-3B
internlm

CapRL 📖 Paper 🏠 Github 🤗 CapRL Collection 🤗 Daily Paper

Open Source 3.0B ↓ 370
SuperApriel-15B-Instruct logo
SuperApriel-15B-Instruct
ServiceNow-AI

A 15B-parameter token-mixer supernet with 8 optimized deployment presets spanning 1.0× to 10.7× decode throughput at 32K sequence length, all from a single checkpoint. Derived from Apriel-1.6 through stochastic distillation and targeted supervised fine-tuning.

Open Source 15.0B ↓ 365
JudgeLM-33B-v1.0 logo
JudgeLM-33B-v1.0
BAAI

--- inference: false language: - en tags: - instruction-finetuning pretty name: JudgeLM-100K task categories: - text-generation ---

Open Source 33.0B ↓ 364
AREX-Base logo
AREX-Base
BAAI

Towards a Recursively Self-Improving Agent for Deep Research

Open Source ↓ 362
CapRL-InternVL3.5-8B logo
CapRL-InternVL3.5-8B
internlm

CapRL 📖 Paper 🏠 Github 🤗 CapRL Collection 🤗 Daily Paper

Multimodal 8.0B ↓ 351
Spatial-SSRL-3B logo
Spatial-SSRL-3B
internlm

📖 Paper 🏠 Github 🤗 Spatial-SSRL-7B Model 🤗 Spatial-SSRL-3B Model 🤗 Spatial-SSRL-Qwen3VL-4B Model 🤗 Spatial-SSRL-81k Dataset 📰 Daily Paper

Open Source 3.0B ↓ 349
Falcon-E-3B-Base-prequantized logo
Falcon-E-3B-Base-prequantized
tiiuae

0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation

Open Source 3.0B ↓ 346
sarashina1-13b logo
sarashina1-13b
sbintuitions

This repository provides Japanese language models trained by SB Intuitions.

Open Source 13.0B ↓ 342
Bunny-v1_0-2B-zh logo
Bunny-v1_0-2B-zh
BAAI

Bunny is a family of lightweight but powerful multimodal models. It offers multiple plug-and-play vision encoders, like EVA-CLIP, SigLIP and language backbones, including Llama-3-8B, Phi-1.5, StableLM-2, Qwen1.5, MiniCPM and Phi-2. To compensate for the decrease in model size, we…

Multimodal 2.24B ↓ 339