AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
BaichuanMed-OCR-7B logo
BaichuanMed-OCR-7B
baichuan-inc

BaichuanMed-OCR-7B is a model fine-tuned from the Qwen2.5-VL-7B-Instruct with our constructed and curated medical report datasets consists of medical report images and related questions and answers (QAs). It has been specifically adapted to perform Optical Character Recognition (…

Open Source 7.0B ↓ 176
ERNIE-4.5-VL-28B-A3B-Paddle logo
ERNIE-4.5-VL-28B-A3B-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 28.0B ↓ 172
bloom-7b1-intermediate logo
bloom-7b1-intermediate
bigscience

WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).

Open Source ↓ 172
Gemma-SEA-LION-v4-27B-VL logo
Gemma-SEA-LION-v4-27B-VL
aisingapore

SEA-LION-VL is an instruct-tuned vision-text model for the Southeast Asia (SEA) region.

Multimodal 27.0B ↓ 163
ERNIE-4.5-VL-28B-A3B-Base-Paddle logo
ERNIE-4.5-VL-28B-A3B-Base-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 28.0B ↓ 160
ERNIE-4.5-VL-424B-A47B-Paddle logo
ERNIE-4.5-VL-424B-A47B-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 424.0B ↓ 159
Baichuan2-13B-Chat-4bits logo
Baichuan2-13B-Chat-4bits
baichuan-inc

🦉GitHub 💬WeChat 百川API支持搜索增强和192K长窗口,新增百川搜索增强知识库、限时免费! 🚀 百川大模型在线对话平台 已正式向公众开放 🎉

Open Source 13.0B ↓ 152
SuperApriel-15b-Base logo
SuperApriel-15b-Base
ServiceNow-AI

A 15B-parameter token-mixer supernet derived from Apriel-1.6 via stochastic distillation. Every decoder layer exposes four trained mixer options —Full Attention, Sliding Window Attention, Gated DeltaNet, and Kimi Delta Attention—enabling flexible architecture selection from a sin…

Open Source 15.0B ↓ 150
ERNIE-4.5-VL-424B-A47B-Base-Paddle logo
ERNIE-4.5-VL-424B-A47B-Base-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 424.0B ↓ 147
LongCat-Flash-Omni-FP8 logo
LongCat-Flash-Omni-FP8
meituan-longcat

Model Introduction We introduce LongCat-Flash-Omni , a state-of-the-art open-source omni-modal model with 560 billion parameters (with 27B activated), excelling at real-time audio-visual interaction, which is attained by leveraging LongCat-Flash's high-performance Shortcut-connec…

Open Source ↓ 141
ERNIE-4.5-21B-A3B-Base-Paddle logo
ERNIE-4.5-21B-A3B-Base-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 21.0B ↓ 140
Intern-S2-Preview-397B-FP8 logo
Intern-S2-Preview-397B-FP8
internlm

💻Github Repo • 🤗HF Model Collections • 🤖ModelScope Collections • 💬Online Chat

Open Source 397.0B ↓ 130