This page covers Cursor’s legacy request-based pricing model that was used before the transition to usage-based pricing.

Overview

The request-based pricing model charged users based on the number of AI requests made rather than token usage. Each license had a monthly allotment of requests. If you exceeded your included usage, you could purchase additional usage on-demand.

Request

A request represents a single message sent to most models, which includes your message, any relevant context from your codebase, and the model’s response. View the model table to see request counts for each model.

  • On-demand usage is available at the model’s API rate plus 20%.
  • Max Mode is available at the model’s API rate plus 20%. Max Mode enables larger context windows, subagents, image generation, and access to the latest frontier models on request-based plans.

Models

Model Provider Default context Max context Capabilities Requests Notes
Claude 4 Sonnet Anthropic 200k - Agent, Thinking, Images 1 Hidden by default; Thinking variant counts as 2 requests in legacy pricing
Claude 4 Sonnet 1M Anthropic - 1M Agent, Thinking, Images 1 Hidden by default; Thinking variant counts as 2 requests in legacy pricing; This model can be very expensive due to the large context window; The cost is 2x when the input exceeds 200k tokens
Claude 4.5 Haiku Anthropic 200k - Thinking, Images 1 Hidden by default; Bedrock/Vertex: regional endpoints +10% surcharge; Cache: writes 1.25x, reads 0.1x
Claude 4.5 Opus Anthropic 200k 200k Agent, Thinking, Images - Hidden by default; Requires Max Mode on legacy request-based plans
Claude 4.5 Sonnet Anthropic 200k 1M Agent, Thinking, Images - Hidden by default; Requires Max Mode on legacy request-based plans; Up to 1M tokens with extended context at the same per-token rates (no long-context surcharge)
Claude 4.6 Opus Anthropic 200k 1M Agent, Thinking, Images - Hidden by default; Requires Max Mode on legacy request-based plans; Up to 1M tokens with extended context at the same per-token rates (no long-context surcharge)
Claude 4.6 Sonnet Anthropic 200k 1M Agent, Thinking, Images - Hidden by default; Requires Max Mode on legacy request-based plans; Up to 1M tokens with extended context at the same per-token rates (no long-context surcharge)
Claude 4.7 Opus Anthropic 300k 1M Agent, Thinking, Images - Hidden by default; Requires Max Mode on legacy request-based plans; Up to 1M tokens with extended context at the same per-token rates (no long-context surcharge)
Claude Fable 5 Anthropic 300k 1M Agent, Thinking, Images - Requires data retention approval for Enterprise customers, Teams and individual customers with Privacy Mode enabled; Anthropic stores agent input and output data for harm-prevention processes; this data is not used to train or improve Anthropic models or products; Requests that trip a security guardrail are automatically routed to Claude Opus; About 2x the cost of Claude Opus 5; Requires Max Mode on legacy request-based plans
Claude Opus 4.7 (fast mode) Anthropic 200k 1M Agent, Thinking, Images - Hidden by default; Requires Max Mode on legacy request-based plans; Limited research preview; Up to 1M tokens with extended context at the same per-token rates as shorter context
Claude Opus 4.8 Anthropic 300k 1M Agent, Thinking, Images - Hidden by default; Requires Max Mode on legacy request-based plans; Fast mode (`claude-opus-4-8-fast`) requires Max Mode on legacy request-based plans; Fast mode is 3x lower per-token pricing than Opus 4.7 fast mode; Up to 1M tokens with extended context at the same per-token rates (no long-context surcharge)
Claude Opus 5 Anthropic 300k 1M Agent, Thinking, Images - Requires Max Mode on legacy request-based plans; Fast mode (`claude-opus-5-fast`) requires Max Mode on legacy request-based plans; Up to 1M tokens with extended context at the same per-token rates (no long-context surcharge)
Claude Sonnet 5 Anthropic 200k 1M Agent, Thinking, Images - Requires Max Mode on legacy request-based plans; Up to 1M tokens with extended context at the same per-token rates (no long-context surcharge); Uses an updated tokenizer, so the same input can map to more tokens
Composer 1 Cursor 200k - Agent, Images 1 Hidden by default
Composer 2.5 Cursor 200k - Agent, Thinking, Images 2 -
Gemini 2.5 Flash Google 200k 1M Agent, Thinking, Images 1 Hidden by default
Gemini 3 Flash Google 200k 1M Agent, Thinking, Images 1 Hidden by default
Gemini 3 Pro Google 200k 1M Agent, Thinking, Images 1 Hidden by default
Gemini 3 Pro Image Preview Google 200k 1M Images 1 Hidden by default; Native image generation model optimized for speed, flexibility, and contextual understanding; Text input and output priced the same as Gemini 3 Pro; Image output: \(120/1M tokens (\~\)0.134 per 1K/2K image, \~$0.24 per 4K image); Preview models may change before becoming stable and have more restrictive rate limits
Gemini 3.1 Pro Google 200k 1M Agent, Thinking, Images 1 -
Gemini 3.5 Flash Google 200k 1M Agent, Thinking, Images 1 Hidden by default
Gemini 3.6 Flash Google 200k 1M Agent, Thinking, Images 1 Hidden by default
Gemini 3.7 Flash Google 200k 1M Agent, Thinking, Images 1 -
GLM 5.2 Z.ai 200k - Agent, Thinking 1 Hidden by default
GPT-5 OpenAI 272k - Agent, Thinking, Images 1 Hidden by default; Agentic and reasoning capabilities; Available reasoning effort variant is gpt-5-high
GPT-5 Fast OpenAI 272k - Agent, Thinking, Images 2 Hidden by default; Faster speed but 2x price; Available reasoning effort variants are gpt-5-high-fast, gpt-5-low-fast
GPT-5 Mini OpenAI 272k - Agent, Thinking, Images 1 Hidden by default
GPT-5-Codex OpenAI 272k - Agent, Thinking, Images 1 Hidden by default; Agentic and reasoning capabilities
GPT-5.1 Codex OpenAI 272k - Agent, Thinking, Images 1 Hidden by default; Agentic and reasoning capabilities
GPT-5.1 Codex Max OpenAI 272k - Agent, Thinking, Images 1 Hidden by default
GPT-5.1 Codex Mini OpenAI 272k - Agent, Thinking, Images 1 Hidden by default; Agentic and reasoning capabilities; 4x rate limits compared to GPT-5.1 Codex
GPT-5.2 OpenAI 272k - Agent, Thinking, Images 1 Hidden by default; Agentic and reasoning capabilities; Available reasoning effort variant is gpt-5.2-high
GPT-5.2 Codex OpenAI 272k - Agent, Thinking, Images 1 Hidden by default; Agentic and reasoning capabilities
GPT-5.3 Codex OpenAI 272k - Agent, Thinking, Images - Hidden by default; Requires Max Mode on legacy request-based plans; Agentic and reasoning capabilities; Available reasoning effort variant is gpt-5.3-codex-high
GPT-5.4 OpenAI 272k 1M Agent, Thinking, Images - Hidden by default; Requires Max Mode on legacy request-based plans; Agentic and reasoning capabilities; 90% discount on cached input tokens; Fast mode is 15% faster with 2x pricing; Long context supports up to 1M tokens with 2x input pricing
GPT-5.4 Mini OpenAI 272k - Agent, Thinking, Images 1 Hidden by default; Smaller, faster variant of GPT-5.4; 90% discount on cached input tokens
GPT-5.4 Nano OpenAI 272k - Agent, Thinking, Images 1 Hidden by default; Smallest GPT-5.4 variant, optimized for cost; 90% discount on cached input tokens
GPT-5.5 OpenAI 272k 1M Agent, Thinking, Images - Hidden by default; Requires Max Mode on legacy request-based plans; Agentic and reasoning capabilities; More token-efficient than GPT-5.4 on comparable tasks; Improved persistence on long-running tasks; Fast mode is available at higher rates; Long context supports up to 1M tokens with 2x input pricing
GPT-5.6 Luna OpenAI 272k - Agent, Thinking, Images - Smallest GPT-5.6 variant, optimized for cost and speed; Agentic and reasoning capabilities; Fast mode is available at 2x pricing; Cache writes are billed at 1.25x the uncached input rate
GPT-5.6 Sol OpenAI 272k 1M Agent, Thinking, Images - Requires Max Mode on legacy request-based plans; Agentic and reasoning capabilities; Fast mode is available at 2x pricing; Long context supports up to 1M tokens with 2x input pricing; Cache writes are billed at 1.25x the uncached input rate; Promotional pricing through November 21, 2026
GPT-5.6 Terra OpenAI 272k - Agent, Thinking, Images - Mid-tier GPT-5.6 variant between Sol and Luna; Agentic and reasoning capabilities; Fast mode is available at 2x pricing; Cache writes are billed at 1.25x the uncached input rate
Grok 4.5 Cursor 256k - Agent, Thinking - Jointly trained by Cursor and SpaceXAI
Grok 4.6 Cursor 256k - Agent, Thinking - Jointly trained by Cursor and SpaceXAI
Kimi K2.7 Code Moonshot 262k - Agent, Thinking, Images 1 Hidden by default
Kimi K3 Moonshot 200k 1M Agent, Thinking, Images 1 Hidden by default; Requires Max Mode on legacy request-based plans; Up to 1M tokens with extended context at the same per-token rates (no long-context surcharge); No separate cache-write fee
Legacy Enterprise Auto Cursor - - Agent - Hidden by default

Legacy customers

If you’re on a legacy request-based plan, you can continue using it until your next renewal. At renewal, you’ll be migrated to the usage-based pricing model.

Migration support

Our team is available to help with the transition from legacy pricing to the current model. We can provide:

  • Detailed cost analysis comparing old vs new pricing
  • Migration timeline and planning
  • Custom solutions for enterprise customers

Contact enterprise@cursor.com for migration assistance.