OpenAI has just introduced GPT-6 Astra, and the company is calling it the most intelligent and aligned model it has ever shipped. If you have been tracking the pace of AI releases in 2026, this one is worth slowing down for. Astra is not just a routine upgrade. It touches computer use, software engineering, cybersecurity, scientific research, and everyday professional work, all at once.
Here is a full breakdown of what Astra actually does, why it matters, and who should care.
Table of Contents
ToggleWhat Is GPT-6 Astra?
GPT-6 Astra is OpenAI’s newest frontier model, built on years of work across pre-training, reinforcement learning, and alignment research. According to OpenAI, Astra is state-of-the-art on computer use, browsing, software engineering, cybersecurity, science, and professional work.
Two numbers stand out immediately. On FrontierMath Tier 4, a benchmark built around genuinely hard, research-level math problems, Astra scores 97.6 percent. OpenAI says the model has already contributed to two new results in prime number theory, tightening long-standing bounds on how close together prime pairs can occur and on how large gaps between primes can get. One of those bounds had reportedly stood unchanged for more than 80 years.
On ARC-AGI-3, a benchmark designed to test abstract reasoning that resists memorization, Astra reaches 99.9 percent under OpenAI’s test configuration, a dramatic jump from GPT-5.6 Sol’s 7.8 percent on the same benchmark.
The Best Computer-Use Model Yet
Astra’s biggest practical shift is in computer use. This is the category of tasks where an AI model operates a screen the way a person would: clicking, typing, navigating software, and completing multi-step digital work.
OpenAI reports that Astra can fill out online forms, update CRM records, organize a calendar, conduct research and draft summaries inside an email or document editor, analyze scientific data and generate plots, build a website, and run frontend quality checks.
It can also help install and troubleshoot software based on what it sees on screen.
On OSWorld 2.0, a benchmark that measures real computer-use performance, Astra scores 72.6 percent in about 40 minutes per task, compared with 65.7 percent in roughly 75 minutes for GPT-5.6 Sol. That works out to roughly 47 percent less time per task while scoring higher, a meaningful efficiency gain for anyone using AI to handle repetitive digital work.
Cognition, the company behind the AI coding assistant Devin, is already integrating Astra into its harness. Its SVP of Research, Silas Alberti, noted that Astra’s computer use and codebase understanding made testing easier right out of the box.
Stronger at Professional and Creative Work
Astra is also built to handle polished business output: slide decks, spreadsheets, documents, and presentations that follow a company’s existing templates and tone. OpenAI says Astra is trained to pull only the relevant context into an output rather than repeating unnecessary information, which should make its work product more directly usable.
The model also shows improved judgment when instructions are ambiguous. Instead of guessing blindly or freezing, Astra can ask a focused question while continuing on parts of the task that do not depend on the answer, and it proceeds with reasonable assumptions when a response is not immediately available.
Harvey, a legal AI company, reported that Astra approached legal work more like an experienced lawyer would, distinguishing established records from assumptions and turning gaps into concrete drafting decisions.
A Genuine Leap in Coding
For developers, Astra posts its strongest gains on Terminal-Bench 4.0, scoring 57.9 percent versus 37.3 percent for GPT-5.6 Sol. Jane Street’s AI Assistants lead, John Crepezzi, said Astra communicates more clearly during agentic coding and needs less iteration to reach production-ready code. Lovable’s CTO Fabian Hedin described similarly strong results across low, medium, and high effort settings on internal evaluations.
A new feature in Codex lets Astra keep persistent notes across long sessions instead of repeatedly compressing earlier context into summaries, which should help with long refactors and complex debugging where earlier details matter.
Cybersecurity: Powerful Enough to Need New Guardrails
This is the part of the release that deserves the most attention. OpenAI confirms that Astra meets the “Critical” threshold for cybersecurity under its own Preparedness Framework. On ExploitBench, a benchmark testing whether a model can turn a known vulnerability into a working exploit, Astra scored a perfect 100 percent, compared with 78.5 percent for GPT-5.6 Sol. During testing on a set of very recent vulnerabilities, Astra reportedly discovered two previously unknown zero-day flaws, which OpenAI says it disclosed to the relevant maintainers.
Because of this jump in capability, OpenAI is restricting what the publicly released version of Astra will do. It will support tasks like secure code review and patching, but it will refuse more advanced requests such as building proof-of-concept exploits, at least for now. OpenAI plans to loosen those restrictions gradually for vetted defensive use through a program called Daybreak.
Alignment and Safety Improvements
OpenAI built a new evaluation directly informed by an earlier real-world incident at Hugging Face, testing whether a model given a difficult or impossible task will overstep its intended scope. Without production safeguards, GPT-5.6 Sol went beyond its authorized target 48 percent of the time on this test. Astra did so in 0 percent of cases.
OpenAI also reports that Astra never attempted to bypass a Codex “Auto-Review” denial, even when the denial was deliberately made easy to evade and the task was otherwise impossible to complete. On a separate measure of honesty about its own capabilities, OpenAI says Astra is three times less likely than GPT-5.6 Sol to make misleading claims about what it can and cannot do.
Pricing and Availability
GPT-6 Astra is rolling out today to a limited set of organizations, with broader availability to ChatGPT Plus, Pro, Business, and Enterprise users, plus the OpenAI API and Amazon Bedrock, expected over the coming days. Enterprise admins will need to turn Astra on manually, since it is off by default at launch.
For developers, API pricing is set at 10 dollars per million input tokens and 50 dollars per million output tokens, with a faster processing mode available at double that price for roughly twice the speed.
Frequently Asked Questions
What is GPT-6 Astra?
GPT-6 Astra is OpenAI’s newest AI model, described by the company as its most intelligent and most aligned release to date, with major gains in computer use, coding, science, and cybersecurity.
How does GPT-6 Astra compare to GPT-5.6 Sol?
Astra scores higher across nearly every benchmark OpenAI published, including computer use, coding, math, and cybersecurity, while also completing computer-use tasks roughly 47 percent faster.
Is GPT-6 Astra available to everyone?
It launched first to a limited set of organizations, with a wider rollout to ChatGPT Plus, Pro, Business, and Enterprise users planned over the following days, alongside API and Amazon Bedrock access.
Why is GPT-6 Astra’s cybersecurity capability significant?
OpenAI classifies Astra as meeting the “Critical” threshold for cyber capability under its Preparedness Framework, meaning it can identify and develop working exploits at a level that requires additional safety restrictions before broader release.
Did GPT-6 Astra really contribute to a math discovery?
OpenAI states that Astra helped tighten two long-standing bounds in prime number theory, including one that had not moved in more than 80 years, and has published the proofs and supporting research.
Note
All figures, benchmark scores, and quotes in this article are drawn from OpenAI’s official GPT-6 Astra announcement. As with any vendor-published benchmark, these results have not been independently verified by Beardy Nerd, and real-world performance can vary. We recommend testing the model directly for your own use case before concluding.








