AI Agent Hub
Back to models
🤖

Qwen: Qwen2.5 VL 72B Instruct

Multimodal qwen
7.0 / 10 128K context Proprietary

About this model

Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing texts, charts, icons, graphics, and layouts within images.

Benchmark Scores

MMMU
70.2

Technical Specs

  • Architecture: text+image->text
  • Context Window: 128,000 tokens
  • Input Modalities: text

Hardware Requirements

  • API-only (no local hardware needed)

Pricing

Input Output Currency
0.25 / 1M tokens 0.75 / 1M tokens USD