AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

88 models for "Audio" Compare
granite-vision-3.2-2b logo
granite-vision-3.2-2b
ibm-granite

Model Summary: granite-vision-3.2-2b is a compact and efficient vision-language model, specifically designed for visual document understanding, enabling automated content extraction from tables, charts, infographics, plots, diagrams, and more. The model was trained on a meticulou…

Multimodal 2.0B ↓ 4.3K
granite-3.1-1b-a400m-base logo
granite-3.1-1b-a400m-base
ibm-granite

Model Summary: Granite-3.1-1B-A400M-Base extends the context length of Granite-3.0-1B-A400M-Base from 4K to 128K using a progressive training strategy by increasing the supported context length in increments while adjusting RoPE theta until the model has successfully adapted to d…

Open Source 1.0B ↓ 4.2K
SingGuard-NSFA-4B logo
SingGuard-NSFA-4B
inclusionAI

SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification

Open Source 4.0B ↓ 726
xgen-mm-phi3-mini-instruct-interleave-r-v1.5 logo
xgen-mm-phi3-mini-instruct-interleave-r-v1.5
Salesforce

Model description xGen-MM is a series of the latest foundational Large Multimodal Models (LMMs) developed by Salesforce AI Research. This series advances upon the successful designs of the BLIP series, incorporating fundamental enhancements that ensure a more robust and superior…

Open Source ↓ 573
SingGuard-NSFA-0.8B logo
SingGuard-NSFA-0.8B
inclusionAI

SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification

Open Source 0.8B ↓ 558
xgen-mm-phi3-mini-instruct-r-v1 logo
xgen-mm-phi3-mini-instruct-r-v1
Salesforce

📣 News 📌 [08/19/2024] xGen-MM-v1.5 released: - 🤗 xgen-mm-phi3-mini-instruct-interleave-r-v1.5 - 🤗 xgen-mm-phi3-mini-base-r-v1.5 - 🤗 xgen-mm-phi3-mini-instruct-singleimg-r-v1.5 - 🤗 xgen-mm-phi3-mini-instruct-dpo-r-v1.5

Multimodal 5.0B ↓ 472
LongCat-Flash-Omni-FP8 logo
LongCat-Flash-Omni-FP8
meituan-longcat

Model Introduction We introduce LongCat-Flash-Omni , a state-of-the-art open-source omni-modal model with 560 billion parameters (with 27B activated), excelling at real-time audio-visual interaction, which is attained by leveraging LongCat-Flash's high-performance Shortcut-connec…

Open Source ↓ 141
Google: Gemini Flash Latest logo
Google: Gemini Flash Latest
~google

This model always redirects to the latest model in the Gemini Flash family.

Closed Source
Google: Gemini Pro Latest logo
Google: Gemini Pro Latest
~google

This model always redirects to the latest model in the Gemini Pro family.

Closed Source
Qwen: Qwen3.8 Omni Flash logo
Qwen: Qwen3.8 Omni Flash
qwen

Qwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio-video understanding. It is suited for audio-video analysis and summarization,...

Closed Source 125.0B
Perceptron: Perceptron Mk1.5 logo
Perceptron: Perceptron Mk1.5
perceptron

Perceptron Mk1.5 is Perceptron's embodied reasoning model for physical agents. It accepts text, image, video, and audio input, and answers with text plus optional structured annotations: points, boxes, polygons, tracks,...

Multimodal
🤖
TypeSafe: Jev Router
typesafe

Jev Router picks the best model and reasoning effort for each request, balancing quality, speed, and cost. It runs on [Jev](https://openrouter.ai/~typesafe/jev-latest), TypeSafe's first System One model, and adapts as your...

Closed Source