AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

88 models for "Audio" Compare
LFM2.5-1.2B-Thinking logo
LFM2.5-1.2B-Thinking
LiquidAI

LFM2.5 is a new family of hybrid models designed for on-device deployment . It builds on the LFM2 architecture with extended pre-training and reinforcement learning.

Reasoning 1.2B ★ 5.0 ↓ 8.5K
granite-4.0-micro-base logo
granite-4.0-micro-base
ibm-granite

Model Summary: Granite-4.0-Micro-Base is a decoder-only, long-context language model designed for a wide range of text-to-text generation tasks. It also supports Fill-in-the-Middle (FIM) code completion through the use of specialized prefix and suffix tokens. The model is trained…

Open Source 3.0B ★ 5.0 ↓ 4.3K
granite-4.0-h-350m logo
granite-4.0-h-350m
ibm-granite

Model Summary: Granite-4.0-H-350M is a lightweight instruct model finetuned from Granite-4.0-H-350M-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of tec…

Open Source 0.35B ★ 5.0 ↓ 4.2K
granite-4.1-3b logo
granite-4.1-3b
ibm-granite

Model Summary: Granite-4.1-3B is a 3B parameter long-context instruct model finetuned from Granite-4.1-3B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. Granite 4.1 models have gone through an impr…

Open Source 3.0B ★ 4.0 ↓ 560.5K
granite-4.1-3b-base logo
granite-4.1-3b-base
ibm-granite

Model Summary: Granite‑4.1‑3B‑Base is a decoder‑only language model with long‑context capabilities, designed to support a broad range of general text‑to‑text generation tasks, as well as fill‑in‑the‑Middle (FIM) code completion. This model shares the same underlying architecture…

Open Source 3.0B ★ 4.0 ↓ 14.6K
granite-4.0-1b-base logo
granite-4.0-1b-base
ibm-granite

Model Summary: Granite-4.0-1B-Base is a lightweight decoder-only language model designed for scenarios where efficiency and speed are critical. They can run on resource-constrained devices such as smartphones or IoT hardware, enabling offline and privacy-preserving applications.…

Open Source 1.6B ★ 2.0 ↓ 44.7K
granite-4.0-1b logo
granite-4.0-1b
ibm-granite

Model Summary: Granite-4.0-1B is a lightweight instruct model finetuned from Granite-4.0-1B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques…

Open Source 1.0B ★ 2.0 ↓ 11.7K
granite-4.0-h-micro logo
granite-4.0-h-micro
ibm-granite

📣 Update [10-07-2025]: Added a default system prompt to the chat template to guide the model towards more professional, accurate, and safe responses.

Open Source 3.0B ★ 2.0 ↓ 7.9K
Qwen2.5-VL-7B-Instruct logo
Qwen2.5-VL-7B-Instruct
Qwen

--- license: apache-2.0 language: - en pipeline tag: image-text-to-text tags: - multimodal library name: transformers ---

Multimodal 7.0B ↓ 5.4M
Qwen2.5-VL-3B-Instruct logo
Qwen2.5-VL-3B-Instruct
Qwen

--- license name: qwen-research license link: https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct/blob/main/LICENSE language: - en pipeline tag: image-text-to-text tags: - multimodal library name: transformers ---

Multimodal 3.0B ↓ 2.3M
Qwen2.5-VL-32B-Instruct-AWQ logo
Qwen2.5-VL-32B-Instruct-AWQ
Qwen

Latest Updates: In addition to the original formula, we have further enhanced Qwen2.5-VL-32B's mathematical and problem-solving abilities through reinforcement learning. This has also significantly improved the model's subjective user experience, with response styles adjusted to…

Multimodal 32.0B ↓ 1.8M
Qwen3-VL-8B-Instruct-FP8 logo
Qwen3-VL-8B-Instruct-FP8
Qwen

This repository contains an FP8 quantized version of the Qwen3-VL-8B-Instruct model. The quantization method is fine-grained fp8 quantization with block size of 128, and its performance metrics are nearly identical to those of the original BF16 model. Enjoy!

Multimodal 8.0B ↓ 1.8M