On September 3, 2026, OpenAI began quietly previewing GPT-6 Astra to a handful of trusted organizations. Less than 24 hours later, on September 4, the model went semi-public: available immediately to ChatGPT Pro, Business Premium, and Enterprise subscribers, with Plus and Business users told to expect access 'in the coming days.' It also landed simultaneously through the OpenAI API, Microsoft Azure, and AWS Bedrock. After months of speculation about a model once rumored under the codename 'Spud,' GPT-6 finally has a name, a price tag, and a real benchmark sheet.
The key numbers: Astra shipped Sept 3-4, 2026 with a 1,050,000-token context window (128,000 max output) and an April 30, 2026 knowledge cutoff. API pricing is $10 per million input tokens and $50 per million output tokens (cached input $1/M, cache writes $12.50/M), with a 'Fast' mode running about 2x faster at roughly 2x the price. OpenAI says it scored 74.1% on its DeepSWE v1.1 coding benchmark (vs. 70.8% for GPT-5.6 'Sol'), 72.6% on OSWorld V2-Offline computer-use tasks (vs. 65.7% for Sol, while cutting average task time from about 75 to about 40 minutes), and surpassed human action-efficiency on 96% of ARC-AGI-3 levels. Independent tracker Artificial Analysis puts Astra's Intelligence Index at 55 versus 51 for Sol.
A real jump, not the leap the marketing suggests
OpenAI is not being shy about how it wants Astra perceived. The benchmark sheet leans hard on tasks where the jump looks dramatic: a 96% pass rate on ARC-AGI-3 (a benchmark specifically designed to resist memorization and reward genuine reasoning), a coding score that beats the previous flagship by more than three points, and a computer-use score that both improved accuracy and nearly doubled speed. Those are genuine, measurable gains, especially the OSWorld result — cutting task completion time from roughly 75 minutes to 40 minutes matters a lot if you're running an agent that has to actually finish work rather than just attempt it. But the number that matters most for cutting through the hype is the one OpenAI doesn't control: Artificial Analysis's Intelligence Index, an independent aggregate benchmark score. Astra lands at 55, up from 51 for GPT-5.6 Sol. That's a real, meaningful improvement — but it's incremental, not the kind of step-change the 'most intelligent and aligned model in the world' framing implies. If you've been burned before by a launch that promised a new era and delivered a modest bump, that pattern is repeating here too, just with better manners about it.
OpenAI president says we're 'in the AGI era' — not everyone agrees
The most quoted line from launch week didn't come from a benchmark chart. OpenAI president Greg Brockman said it's 'not unreasonable to feel that we are now in the AGI era' - a remark widely reported across tech press that quickly drew pushback from AI safety researchers and commentators. The friction is straightforward: a 55 on an intelligence index that was 51 a few months ago is progress, but it isn't the kind of discontinuity that 'AGI era' rhetoric usually implies. Calling an incremental release the dawn of a new era either means the definition of that era has quietly gotten looser, or the marketing has outrun the measurement. This isn't just a semantic squabble. How a leading lab talks about its own capability trajectory shapes public expectations, investor behavior, and — more consequentially — how much scrutiny regulators and safety researchers think the moment deserves. Astra's own internals give that scrutiny some teeth: the model uses a new 'recurrent depth' reasoning technique that, according to safety researchers, makes its chain-of-thought harder to monitor than previous architectures. Harder-to-audit reasoning arriving in the same release where leadership is talking up AGI is exactly the combination that makes safety-focused observers uneasy.
The caution pattern continues from August
None of this is happening in a vacuum. LumiChats covered the run-up to this release back on August 15, when OpenAI slowed down an earlier version of Astra after internal testing suggested it was nearing 'Critical' cyber capability — the kind of threshold that triggers extra review before a model reaches the public. Alongside that delay, OpenAI released a separate, more restricted model (internally 'Daybreak,' publicly GPT-5.6-Cyber) available only to vetted security firms. That caution shows up again in this week's release. The public version of Astra ships in a restricted form that rejects certain cybersecurity-related prompts outright, rather than answering them with a generic safety disclaimer. It's a direct continuation of the same pattern: OpenAI's internal safety evaluations are running ahead of what the public model is allowed to do, and the company is choosing to gate capability rather than ship it uniformly. Whether that's a sufficient guardrail for a model with reasoning safety researchers already say is harder to audit is very much an open question — and one the industry will keep circling back to as recurrent-depth-style architectures spread to other labs.
Where Astra actually sits in the field
Context matters here too. Astra launched two to three days after Anthropic shipped Claude Fable 5.1 on September 1, 2026 — a reminder that the frontier is currently being contested week by week, not year by year. For anyone choosing a model based on capability alone, that's good news: competition between labs keeps pushing real benchmark numbers up, even when the marketing language around any single release should be taken with a grain of salt. For developers, the API economics are worth noting on their own. At $10/M input and $50/M output tokens, Astra sits at a premium price point, with the 2x-speed 'Fast' mode roughly doubling cost again for latency-sensitive use cases. The million-plus-token context window is a genuine practical upgrade for anyone working with large codebases or long documents, even setting aside the reasoning benchmark debate entirely.
| Metric | GPT-5.6 Sol | GPT-6 Astra |
|---|---|---|
| Artificial Analysis Intelligence Index | 51 | 55 |
| DeepSWE v1.1 (agentic coding) | 70.8% | 74.1% |
| OSWorld V2-Offline (computer use) | 65.7% | 72.6% |
| Avg. time per computer-use task | ~75 min | ~40 min |
| Context window | — | 1,050,000 tokens |
- GPT-6 Astra previewed to trusted orgs Sept 3, 2026, then went semi-public Sept 4 for ChatGPT Pro, Business Premium, and Enterprise users, with Plus/Business access 'coming days' later.
- Also live now via the OpenAI API, Microsoft Azure, and AWS Bedrock at $10/M input and $50/M output tokens (Fast mode: ~2x speed, ~2x price).
- Context window: 1,050,000 tokens, 128,000 max output, knowledge cutoff April 30, 2026.
- OpenAI's benchmarks show real gains — 96% on ARC-AGI-3, 74.1% on DeepSWE v1.1, 72.6% on OSWorld V2-Offline — but independent tracker Artificial Analysis scores the overall jump as incremental (55 vs. 51 Intelligence Index).
- OpenAI says this was its largest training run ever, the first pretrained on more than 100,000 GPUs, run at its Stargate site in Texas.
- President Greg Brockman's 'AGI era' comment has drawn pushback from safety researchers, partly because Astra's new 'recurrent depth' reasoning is reportedly harder to monitor.
- The public release restricts certain cybersecurity prompts outright, continuing the caution pattern from August's Astra delay and the restricted 'Daybreak' cyber model.
- The launch lands just two to three days after Anthropic's Claude Fable 5.1 (Sept 1, 2026), keeping the frontier race extremely close heading into fall 2026.
01Who can use GPT-6 Astra right now?
As of September 4, 2026, it's available to ChatGPT Pro, Business Premium, and Enterprise subscribers, plus anyone using the OpenAI API, Microsoft Azure, or AWS Bedrock. Plus and Business users are expected to get access within days.
02How much does GPT-6 Astra cost via the API?
$10 per million input tokens and $50 per million output tokens, with cached input at $1/M and cache writes at $12.50/M. A 'Fast' mode runs about twice as fast for roughly double the price.
03Is GPT-6 Astra actually a big leap over GPT-5.6?
It's a real improvement, but not the dramatic leap the launch messaging suggests. OpenAI's own benchmarks (DeepSWE, OSWorld, ARC-AGI-3) show solid gains, but independent tracker Artificial Analysis puts the overall Intelligence Index jump at 51 to 55 — meaningful, but incremental.
04Why is Astra's public version restricted for cybersecurity prompts?
OpenAI flagged an earlier version of Astra as nearing 'Critical' cyber capability back in early August 2026, delayed its release, and issued a separate vetted-only cyber model called Daybreak. The publicly released Astra continues that caution by rejecting certain cybersecurity-related prompts outright.
05Does this mean we're in the 'AGI era' as OpenAI's president suggested?
That claim is contested. Independent benchmarks show Astra as a solid but incremental step up from GPT-5.6, and safety researchers have pushed back on framing an iterative release as an AGI milestone — especially given that Astra's new reasoning technique is reportedly harder to audit, not easier.
If this launch cycle proves anything, it's that no single lab holds the frontier for long — Astra shipped just days after Claude Fable 5.1, and the next leapfrog is probably already in training somewhere. That's exactly the situation LumiChats exists for: instead of betting your workflow on one model's release calendar, you get Claude, GPT, Gemini, DeepSeek, and more under a single login, so when GPT-6 Astra's coding gains matter for one task and Claude handles another better, you're not stuck picking sides — or paying for multiple subscriptions to keep your options open, all for pay-per-day pricing under $1/day.
