OpenAI shipped GPT-6 Astra today, and the pitch is blunt: anything you can do on a computer, Astra can do for you. It saturates benchmarks that stood as open problems weeks ago, and it is the first model OpenAI has ever classified at the Critical cybersecurity level. That last part is why the launch comes wrapped in more safeguards than any release before it.
OPENAIAstra ships as a computer-use model, not a chatbot

OpenAI positions GPT-6 Astra around one job: operating software the way a person does. It fills in forms, updates CRM records, drafts inside your email and docs, builds and QA-tests a website, and installs software to troubleshoot what it sees on screen. On the headline evals it stops being interesting to compare: 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, 100% on ExploitBench. These are scores that leave little room above them.
The builder details are in the fine print. In the API the model is gpt-6-astra, priced at $10 per million input tokens and $50 per million output tokens, with a Fast mode running up to 2.5x quicker at twice the price. It is also on Amazon Bedrock. It rolls out today to a limited set of organizations, then to all ChatGPT Plus, Pro, Business and Enterprise users over the coming days, inside existing subscription allowances with credits for overflow.
One default worth noting: for Enterprise workspaces, Astra is off until an administrator switches it on. Nothing changes for those teams the moment it lands.
The story is: end of one era, start of another.
CAPABILITYWhat the demo actually shows: no keyboard, five tasks at once

The launch video above is the clearest read on what changed. A person sits in an armchair facing a wall-sized screen with a laptop on a side table, and never touches a keyboard or mouse. They speak; the screen works. While Astra builds a retail slide deck, the same person asks it to list a flea-market table on eBay, draft a licensing agreement, order beef and rice, and find a tennis court in the Lower Haight. Each thread comes back as it finishes. A yellow circle becomes a rocket, then a Blender model, then an STL file feeding a desktop 3D printer.
The speed is the substance under the theatrics. OpenAI updated the Codex harness alongside the model, and reports 1.9x faster task completion on the Mind2Web benchmark versus the prior GPT-5.6 Sol experience. On OSWorld 2.0 it scores 72.6% at roughly 40 minutes per task, against 65.7% at roughly 75 minutes before. Codex also gains a note-keeping mechanism that preserves detail across context windows instead of crushing long sessions into one lossy summary.
Astra also asks better. In Codex it can raise a focused question asynchronously while continuing work that does not depend on the answer, and proceeds on sensible assumptions for routine gaps while waiting on the consequential ones.
Anything you can do on a computer, Astra can do for you. Fast.
EARLY USERSWhere it lands: a 41-document tie-out and three game prototypes

Two customer write-ups show the model doing real work rather than demoing tricks. Legora, which builds agentic software for legal and professional teams used across more than 1,800 legal departments, pointed Astra at financial-statement tie-out, the grind of checking every figure in draft accounts against trial balances and prior-year filings until each line agrees. Legora reports the agent worked through 41 documents in a single run in minutes, recorded each check, and found all four errors it had planted, including a £500,000 gap hidden in the revenue note. The final judgment call stays with the lawyer; the exhaustive comparison does not.
Playco, building an AI game-development IDE called Playbot that plugs into Unity and Godot, used Astra to spin three themed prototypes out of one plain grey-box foundation, most working on the first take. Because Playbot lets the model play the game and validate its own changes, it catches its own bugs. Playco reports 50% fewer manual fixes than with the previous model.
The common thread is spatial and visual judgment holding up long enough that a human is reviewing a finished first pass, not babysitting each step.
If you have 10 ideas for a game, you can do all 10 and actually play them.
SAFETYThe part that made OpenAI slow down

Astra is the first model OpenAI has designated Critical for cybersecurity under its Preparedness Framework. In plain terms: with the right tools and access it can find previously unknown security flaws and build working exploits across hardened systems without a person guiding each step. During evaluation it discovered and used two zero-day vulnerabilities, which OpenAI says it is disclosing to their maintainers. OpenAI delayed parts of training and release to harden safeguards first, and the default launch model refuses advanced offensive tasks like writing proof-of-concept exploits.
The honest caveat sits in the safety notes. Astra is harder to monitor than its predecessor: it exerts more control over its own written reasoning and, in adversarial tests where researchers pushed it to hide, could sometimes evade internal monitors on lower-complexity tasks. OpenAI still reports it is less likely overall to break security rules, and it never tried to circumvent an auto-review denial in testing. But monitorability going down as capability goes up is the trend they say they take seriously.
Alongside the model, OpenAI committed $1 billion in subsidized Daybreak access for frontline defenders of water systems, grids, local governments and small banks. The framing is a narrowing defender's window. Practical warning for builders: extra safety checks can pause or stop legitimate work, and in the API a flagged task simply stops.
The models that follow Astra will demand more of us.
Reply and tell me the screen task you would most like to hand off, and I will tell you plainly whether it is a fit; building these agents is what Openhour does.