AI Agent Hub
Back to models
OpenAI: GPT Astra Latest logo

OpenAI: GPT Astra Latest

Closed Source ~openai Released 2026-09-03
53.0 / 100 1.1M context Proprietary

About this model

GPT-6 Astra (API id gpt-6-astra) is OpenAI's flagship model announced on September 3, 2026, positioned as the company's most capable and aligned general model for agentic work. It targets long-horizon computer use, browsing, software engineering, cybersecurity defense, science, and professional knowledge work, with rollout across ChatGPT paid tiers, Codex, and the OpenAI API (plus Azure and AWS Bedrock). Standard API pricing is $10 per million input tokens and $50 per million output tokens, with an optional Fast mode at 2x speed and 2x price.

On public evaluations, Astra sets strong marks on graduate science reasoning (GPQA Diamond), agentic browsing (BrowseComp), multimodal understanding (MMMU-Pro), terminal and repository-style coding agents (DeepSWE), and realistic professional agent workflows (Agents' Last Exam and AutomationBench). It also reports leading scores on FrontierMath Tier 4 and ARC-AGI-3 under OpenAI's evaluation harnesses, and state-of-the-art computer-use metrics such as OSWorld 2.0 and ScreenSpot-Pro in the launch materials. OpenAI did not publish a SWE-Bench Verified score at launch; software engineering evidence centers on DeepSWE and Terminal-Bench 4.0 (not mapped here because that benchmark name is outside the catalog's canonical Terminal-Bench versions).

Astra supports very long contexts (documented at about 1.05M tokens) with strong needle-in-haystack retrieval on OpenAI MRCR v2 out to roughly 1M tokens. It is multimodal for vision-grounded computer use and document workflows. OpenAI classifies it as reaching Critical cybersecurity capability under its Preparedness Framework, with tightened safeguards for offensive cyber tasks while enabling defensive security workflows. The model emphasizes alignment improvements, scope respect, and lower rates of capability hallucination relative to prior frontier models, alongside ongoing work on reasoning monitorability.

Benchmark Scores

HLE
57.2
DeepSWE
74.1
MMMU-Pro
86.9
GDPval-AA
1542.0
BrowseComp
94.2
GPQA-Diamond
96.1
AutomationBench
41.4
Agents-Last-Exam
59.3

Technical Specs

  • Architecture: Transformer
  • Context Window: 1,050,000 tokens
  • Input Modalities: text, image

Hardware Requirements

  • Compute: API only

Pricing

Input Output Currency
10.00 / 1M tokens 50.00 / 1M tokens USD