GPT-6 Astra: What Frontier AI That Uses Your Software Means in 2026
GPT-6 Astra can drive software. Not describe it, and not write a script that calls its API, but read what’s on the screen and click through it the way one of your people would. Fill in the form. Update the record in the system nobody ever built an integration for. Run the tests, pull the results together, hand back something finished.
OpenAI released it on 3 September and it’s already live across Microsoft Foundry, Copilot Cowork, Copilot Studio and GitHub Copilot.
The benchmark scores are worth a closer look. On independent testing Astra isn’t the highest scorer for raw reasoning. Artificial Analysis puts it at 61 on the Intelligence Index against Claude Fable 5.1 at 66, and behind on the Coding Agent Index too. What it wins is everything that involves finishing a job: computer use, terminal work, automation, maths, cyber. And it gets there spending a fraction of the tokens.
For two years, choosing a model meant asking which one was cleverest. That question is getting less useful by the month. The one that pays rent now is how much completed work you get back per dollar, and by that measure Astra is a serious step forward.
Note: all dollar figures are in US dollars unless otherwise stated.
Key Takeaways
- Astra’s standout feature is computer use, meaning it can operate software through the interface, including systems with no API
- On independent testing it sits behind Claude Fable 5.1 for general reasoning and leads on computer use, automation, terminal tasks, maths and cyber
- It finishes an Intelligence Index task for around $1.67 against Fable 5.1’s $3.76, because it spends far fewer tokens getting there
- Available now in Microsoft Foundry, Copilot Cowork, Copilot Studio and GitHub Copilot, with different availability and admin controls on each
- It’s the first model OpenAI has rated Critical for cyber capability, and the shipping version is built for defensive work like secure code review and patching
- Automation becomes possible in places that never justified an integration project, which is where the early wins are
- APRA’s April letter reads well as a practical checklist, and most of what it asks for is a few weeks of work
- Accountability stays with the regulated entity, as it always has
What the Benchmarks Are Telling Buyers
OpenAI’s launch table has Astra ahead of Claude Fable 5.1 almost everywhere. Those are OpenAI’s own results, run in OpenAI’s research environment, and a few of the comparison rows carry footnotes about modified evaluations. Useful, and worth reading as a vendor’s account of its own product.
The independent numbers split more interestingly. Artificial Analysis has Fable 5.1 ahead on both of its flagship indices, while Astra takes Terminal-Bench 4.0 at 57.9 against 55.8, AutomationBench at 41.4 against 31.4, FrontierMath Tier 4 at 97.6 against 87.8, and GPQA Diamond at 96.0 against 93.7. It also leads every computer-use measure published at launch.
Efficiency is where the commercial case sits. Artificial Analysis measured Astra using a third of the tokens of GPT-5.6 Sol at max effort in Codex, and a fifth of the tokens of Claude Opus 5. List prices for Astra and Fable 5.1 are identical at $10 per million input tokens and $50 per million output, so the difference shows up in cost per completed task rather than on the rate card.
For anyone building on Foundry this is good news rather than a horse race. You get both models in the same catalogue, under the same identity and audit controls, and you can point each one at the work it’s best at. Reasoning-heavy research to one, high-volume execution to the other.
Software That Can Use Your Software
Computer use is the capability to pay attention to. Astra interprets what’s on screen and interacts with applications through their interfaces, which Microsoft describes in its Foundry announcement as covering things like updating records, working through development tools, testing software and assembling results into reports.
What makes that commercially interesting is where it applies. Every organisation has a handful of processes that were never worth automating, not because the process was complicated, but because two of the systems involved don’t talk to each other and the integration was never going to clear a business case. Those are exactly the processes this opens up, and there are usually more of them than anyone expects.
What it looks like on the floor
Take an arrears team at a non-bank lender. A case officer works a queue: open the servicing platform, check the payment history, look up whether a hardship note is sitting in a separate system, draft a letter from a template, log the outcome. Twenty minutes a file, repeated all morning, and it has stayed manual because the systems were never wired together.
An agent with computer use can work that queue through the same screens, without an integration project in front of it. The officer picks up completed files and spends their morning on the handful of cases that need an actual conversation with a customer, which is the part of the job they were hired for and the part that changes outcomes.
There’s a design decision worth making deliberately here, and it’s about where the officer’s judgement now lands. Five sequential steps gave you five natural moments to notice something odd. One reviewed file gives you one, so it’s worth making that review a good one. Reviewers do better work when they can see what the agent opened and where it made a call on an ambiguous instruction, rather than only the tidy artefact at the end. That’s a configuration choice, and it’s much easier to make at the start than to retrofit.
Permissions Finally Get Their Business Case
Microsoft grounds Astra in your files, meetings, chats and business data through Work IQ, and does it within existing permissions. That’s the right design, and it puts your permissions model to work in a way it has never really been tested before.
Most organisations are carrying a few years of SharePoint sprawl and Teams memberships nobody has looked at, and it has never caused much trouble, because no human was ever going to read all of it. An agent reading across everything a user can access changes the maths. The silver lining is that the tidy-up you’ve been deferring now has a clear return attached to it, and it’s foundational work that pays off across every AI project you run afterwards rather than just this one.
Most of the Governance Work Is Already Written Down
APRA wrote to every regulated entity on 30 April with its expectations on AI, and the letter is more useful than most regulatory correspondence because it’s specific. Frameworks and reporting lines. Ownership across the lifecycle, from design through to decommissioning. An inventory of AI tooling and use cases. Human involvement in high-risk decisions. Staff training. Visibility over the supply chain, including fourth parties. Monitoring that runs continuously rather than sampling once a quarter.
Read as a to-do list rather than a warning, that’s a manageable piece of work, and the inventory on its own answers most of what a board will ask you this year.
APRA found entities leaning on policy direction and after-the-fact detection instead of enforceable technical restrictions. A written policy is a reasonable control when the thing you’re guarding against is a person doing something once. Agents move at a different pace, so the controls that hold up are the ones configured into the platform itself.
Those platform controls are in better shape than they were a year ago. Agent 365 gives agents identities in Entra, so they can be governed with Conditional Access and access reviews like any other principal. Copilot Studio has an agent inventory schema in preview for discovering and auditing agents across a tenant. Purview DLP can stop an agent processing a prompt or returning a response. It’s useful to know where the line falls: Communication Compliance, Insider Risk Management, eDiscovery and audit surface risky activity rather than preventing it in real time, which makes them your evidence layer. You want both, and they do different jobs.
The Cyber Story Runs Both Ways
Astra is the first OpenAI model to meet the Critical cyber capability threshold under the company’s Preparedness Framework. OpenAI’s description is that with the right tools and access, it can find previously unknown security flaws and develop new ways to exploit them across many well-protected systems without a person guiding each step. It scored 100% on ExploitBench, and turned up two previously unknown vulnerabilities during evaluation, which OpenAI says it has disclosed to the maintainers. The version customers get refuses to build proof-of-concept exploits.
ASIC called this shift back in May. Commissioner Simone Constant’s open letter to licensees warned that frontier AI could expose vulnerabilities at unprecedented speed and scale, with the memorable line that “the clock is at a minute to midnight”. Entities were required to table it at their ultimate board and risk governance committees.
The defensive side of this gets less airtime than it deserves. The ASD guidance that APRA points entities to is refreshingly practical, and one of its recommendations is to use frontier models to strengthen code before it reaches production, alongside a patch every day mentality and applying patches regardless of severity rating, since low-severity flaws can be chained together. Astra supports secure code review and patching today. The same capability that makes your risk committee sit up will find weaknesses in your own code faster than your current process manages, and hardly anyone has pointed it there yet.
Where This Leaves Things
Two years of AI conversation have been about how clever these models are getting. Astra is a good moment to change the subject, because the measure that matters commercially is how much finished work comes back and what it cost to get it. On that measure the ground has moved this month.
The organisations that get the most out of it will be the ones that already know what’s running in their tenant, who owns it, and what it’s allowed to touch. That groundwork holds its value no matter what ships next quarter, and it’s available to anyone willing to spend a fortnight on it.
If you’d like to talk through what Astra opens up in your environment, or where to start with agent governance on the Microsoft stack, get in touch with the 365 Mechanix team. We work with organisations across Australia and New Zealand on Dynamics 365, Power Platform, Copilot and Azure, with a focus on regulated industries.
FAQs
What is GPT-6 Astra?
OpenAI’s newest frontier model, released on 3 September 2026. It’s built for multi-step work rather than conversation, and its defining feature is computer use, meaning it can operate software through the interface rather than only through APIs. You can reach it via Microsoft Foundry, Copilot Cowork, Copilot Studio, GitHub Copilot, the OpenAI API and AWS Bedrock.
Is GPT-6 Astra the most intelligent AI model available?
Depends what you mean by intelligent. On independent testing from Artificial Analysis, Claude Fable 5.1 leads the Intelligence Index at 66 to Astra’s 61, and the Coding Agent Index at 70 to 67. Astra leads on computer use, automation, terminal tasks, maths and cyber, and uses far fewer tokens to get its scores. The better question is which one suits the job in front of you, and on Foundry you don’t have to choose just one.
How do we get access on the Microsoft stack?
It varies by surface. Astra is generally available in Microsoft Foundry, though some Azure subscription tiers need to request quota first. Rollout to Copilot Cowork and Copilot Studio is gradual and varies by region and tenant, managed through the Microsoft 365 admin center. In Copilot Studio, administrators control preview, experimental and external models separately, and it’s worth knowing that data processed by preview or experimental models may sit outside your geographic boundary.
What does a Critical cyber rating mean?
It’s a designation in OpenAI’s own Preparedness Framework rather than a marketing term. It means the model can identify and develop working zero-day exploits in many hardened real-world systems without a person guiding each step, and Astra is the first OpenAI model to reach it. The commercial release carries safeguards that block the most advanced offensive work while still supporting defensive tasks.
What is computer use, and where does it help most?
The model reads what’s on screen and drives applications the way a person would, including systems with no API. The sweet spot is any process where the work itself is simple but the systems were never connected, which covers a surprising amount of back-office and operations work. Scoped credentials, approved resources and human checkpoints on consequential actions are the settings to get right before you scale it.
What does APRA expect from us on AI right now?
Board literacy sufficient to challenge AI risk, governance frameworks with clear reporting lines, an inventory of AI tooling and use cases, named ownership across the lifecycle, human involvement in high-risk decisions, staff training, visibility over the full AI supply chain including fourth parties, and continuous monitoring rather than point-in-time sampling. APRA has said it will apply supervisory focus to AI adoption and act where entities aren’t managing the risk proportionately.
What should we do first?
Find out what’s switched on. Most organisations can’t currently say which models their agents are calling, which systems those agents can reach, or who approved either. That inventory is quick to produce, it underpins everything else APRA expects, and it doesn’t go stale when the next model ships.
How does 365 Mechanix help?
We work with organisations across ANZ to design, build and govern AI on the Microsoft stack, across Copilot, Copilot Studio, Power Platform, Dynamics 365 and Foundry, with a focus on financial services and regulated environments. If any of this is live for you, get in touch.
This blog is intended as general guidance only and does not constitute legal or compliance advice. We recommend consulting your compliance team or legal advisors for advice specific to your organisation.