Agentic Security and CPS 230: What Microsoft’s Project Perception Means for Regulated Firms

Someone finds a weakness in your systems. In the old world that took an attacker weeks of patient probing, and you had a fair chance of spotting them before they got anywhere.

That world is closing. AI has collapsed the cost of finding a flaw and the time it takes to build something that exploits it. Attackers can now generate exploits faster, run more campaigns at once, and do it for a fraction of what it used to cost. Scanning your estate every so often and patching when you get around to it was always a bit optimistic. Against machine-speed attacks it’s simply finished.

Microsoft’s answer, unveiled at the end of July, is to put AI on the other side of the fight. Autonomous agents that hunt for weaknesses, judge what’s a real threat, and fix things, continuously, without waiting for a human to notice. It’s a genuine step change in what defence can do.

It also drops a live question on the desk of every compliance officer in financial services. If software is now taking defensive action on its own, who authorised it, can you explain it, and where exactly are you in the loop? Here’s what Microsoft has built, and why the answer to that second set of questions matters more than the first.

Key Takeaways

  • Microsoft has launched Project Perception, an agentic security system that detects, judges and remediates threats autonomously. Public preview opens 3 August
  • It runs three kinds of agent together: red teams that probe for weaknesses, blue teams that assess real risk, green teams that take corrective action
  • Behind it is MAI-Cyber-1-Flash, Microsoft’s first cyber model, which scores 96% on the CyberGym benchmark at around half the cost of Microsoft’s current setup
  • The defensive upside is real: threat response at machine speed, less manual triage, fewer weaknesses left sitting open
  • For APRA-regulated firms, autonomous action collides with CPS 230, the Financial Accountability Regime and ASIC’s focus on explainable decisions
  • Microsoft says humans stay in control and existing governance controls are inherited. That’s a claim to verify against your own obligations
  • The firms that win with this won’t be the fastest to deploy. They’ll be the ones who can act autonomously and still evidence who was accountable

 

Why the Maths of an Attack Changed

Security has always been a race between the people breaking in and the people keeping them out. What AI changed is the speed and the price of running that race, and it changed them on the attacker’s side first.

Finding a vulnerability used to take skill, patience and time. Now a capable model can read through an enormous codebase looking for the one weak spot, and do it cheaply enough to try against thousands of targets. The economics that used to protect you, the sheer effort involved in mounting a serious attack, have largely evaporated.

Defenders have been stuck with tools built for a slower, more human world. You scan on a schedule. Alerts pile up. Analysts work through them as fast as they can, which is never quite fast enough. By the time a person gets to a genuine threat buried in the noise, it’s had hours or days to sit there. None of that is a criticism of your team. It’s a mismatch of speed, and you can’t hire your way out of it.

The only thing that keeps pace with an AI-speed attack is an AI-speed defence. Which is the case Microsoft is making.

 

What Project Perception Actually Does

Project Perception enters public preview on 3 August. Strip away the framing and it’s a system that runs teams of specialised agents against your environment, around the clock, in a loop that keeps tightening.

There are three kinds of agent, and the names borrow from how security teams already think.

  • Red team agents go looking for trouble. They probe for the paths an attacker could take, trying to find the way in before someone hostile does.
  • Blue team agents do the judging. They investigate what’s happening, reason over the context, and work out what actually represents meaningful risk rather than noise.
  • Green team agents do the fixing. They take corrective action and harden defences across the environment.

Working together, they form a closed loop. The system continually discovers weaknesses, decides which ones matter, acts on them, and learns from how that went, then goes round again. It’s the difference between a smoke detector that beeps and a system that finds the frayed wire, decides it’s a real hazard, and shuts off the circuit before anything catches.

The reason this is possible now, and affordable, comes down to the model underneath.

The model doing the heavy lifting

Perception’s first job is software vulnerability management, and for that it leans on MAI-Cyber-1-Flash, Microsoft’s first purpose-built cyber model. Running inside a multi-agent harness Microsoft calls MDASH, it scores 96% on CyberGym, an industry benchmark for finding real vulnerabilities in large, messy codebases. Microsoft puts that 12 points above its Mythos configuration.

The number that matters for whoever signs the budget is the cost. Microsoft says the same setup runs at roughly half the price of its current configuration. It manages that by having the compact, specialised model handle up to nine tasks in ten, and only calling on a larger, dearer model for the small slice of genuinely hard problems. You get top-tier detection without paying top-tier prices on every scan, which is the only way continuous defence makes economic sense when it’s running every hour of every day.

 

The Part That Should Give Regulated Firms Pause

Everything above is a good story. For most businesses it’s an easy yes. If you run a bank, an insurer, a super fund or a non-bank lender, it’s more complicated, and the complication is the whole point of this article.

An agent that takes corrective action on its own is, by definition, making and acting on decisions without a person in the moment. That is precisely the territory your regulators have spent the last two years walking into.

CPS 230 and operational resilience

APRA’s CPS 230 is in force, with targeted amendments taking effect on 1 July 2026, and its whole thrust is that you remain accountable for the systems material to how you operate, including the ones a third party provides. An AI system taking autonomous defensive action across your estate is about as material as it gets. “The security platform did it” is not an account of a decision. If a green team agent takes an action that has knock-on effects, you need to be able to show how and why, and demonstrate the controls around it were real rather than described in a document nobody tested.

Accountability sits with a named person

Under the Financial Accountability Regime, responsibility for operational risk lands on identified senior executives, personally. Someone’s name is against the risk this system introduces and the controls that hold it in check, and that person needs to understand what the agents can do at more than a conceptual level. Buying capability you can’t explain to your own accountable executive is a governance gap, however good the capability is.

ASIC wants the decision explained

ASIC’s interest runs alongside APRA’s and it’s fixed on outcomes and explainability. Where an automated action touches how customers are treated, “the model decided” won’t survive scrutiny. You need a defensible account of what was considered and why. The two regulators’ interests increasingly overlap here, so a system that satisfies one and trips the other isn’t a comfortable place to sit.

Microsoft has clearly built with this in mind. It states that Perception keeps humans firmly in control, is built to its Responsible AI principles, and inherits the security, compliance and governance controls customers already rely on, with enterprise controls like role-based access, tenant isolation, auditability and sandboxed execution. That’s the right list of things to have addressed. It is also a set of claims, and the work of matching each one to your specific obligations under CPS 230, FAR and the rest is yours, not the vendor’s. Powerful defence and demonstrable control aren’t in tension here. You just have to be able to evidence both, and the evidence is the bit people leave until an auditor asks.

 

So How Should You Approach It?

None of this is a reason to sit it out. Machine-speed attacks are coming for regulated firms whether or not those firms modernise their defences, and refusing capable tooling on governance grounds just leaves you slower than the threat. The task is to adopt it in a way that holds up. Where you start depends on where you sit.

If your security operations are mature

You’re well placed to pilot this, and the preview from 3 August is the moment to get hands on. Run it against a defined slice of your environment, and pay as much attention to the audit trail the agents produce as to the threats they catch. The question to answer in the pilot isn’t just “does it work”, it’s “can we reconstruct and defend every autonomous action it took”. If you can, you have something. If you can’t, you’ve found the gap before it found you.

If you’re earlier in the journey

Don’t lead with the autonomous remediation. Start where the agents surface and prioritise risk while a human still makes the call to act. You get a large share of the benefit, faster detection and far less manual triage, while you build the governance scaffolding, the sign-off paths, the logging, the accountability mapping, that autonomous action will eventually need. Grow into the loop rather than jumping into the deep end of it.

If you’re the accountable executive

Ask three questions before anything goes live, and keep asking them. Can we explain, after the fact, why the system took a given action? Is there a clear point where a human can intervene, and is it real rather than theoretical? And if a regulator asked us tomorrow to walk through how this is governed, could we, today, without a scramble? If the honest answer to any of those is “not yet”, that’s your roadmap, and it’s worth more than the deployment date.

 

The Real Test

Autonomous defence is going to become normal faster than most people expect, because the alternative is losing a race you can’t win by hand. That part isn’t really in question.

What separates the firms that do this well from the ones that end up explaining themselves to APRA isn’t the sophistication of the tool. Everyone will have access to similar capability soon enough. It’s whether you can let software act at machine speed and still answer, calmly and on demand, who was accountable and how the decision was made. Get that right and agentic security is one of the better things to happen to your risk posture in years. Get it wrong and it’s a very fast way to create a problem you’ll be a long time cleaning up.

If you want to work through what agentic security means for your environment and your obligations, get in touch with the 365 Mechanix team. We help financial services organisations across Australia and New Zealand adopt Microsoft’s capabilities in a way that stands up to the scrutiny they operate under.

 

FAQs

What is Project Perception?

Microsoft’s new agentic security system, entering public preview on 3 August. It runs red, blue and green team agents together to continuously find weaknesses, judge which represent real risk, and take corrective action, rather than waiting for a human analyst to work through alerts.

What is MAI-Cyber-1-Flash?

Microsoft’s first purpose-built cyber security model. Inside its MDASH harness it scores 96% on the CyberGym benchmark for spotting real vulnerabilities in large codebases, at roughly half the cost of Microsoft’s current configuration, by handling most of the work itself and reserving a larger model for the hardest cases.

Does agentic security break CPS 230 compliance?

Not on its own, but it raises the bar for how you govern it. CPS 230 keeps you accountable for systems material to your operations, including third-party ones. An autonomous security system qualifies, so you need documented, tested controls around it and the ability to explain the actions it takes. The platform supports that. It doesn’t deliver it for you.

Can autonomous agents take security actions without a human involved?

They’re designed to, which is the source of both the benefit and the governance challenge. Microsoft states humans remain in control and can intervene. For a regulated firm, the practical work is defining where the human checkpoints sit, making sure they’re genuine rather than nominal, and logging every autonomous action so it can be reviewed later.

Why does this matter more for financial services than other industries?

Because regulated firms carry obligations most businesses don’t. APRA’s CPS 230, the Financial Accountability Regime and ASIC’s focus on explainable, fair outcomes all bear on systems that make and act on decisions autonomously. The security upside is the same for everyone. The accountability requirements are heavier in FSI, so the adoption approach has to be more deliberate.

Should we wait until the technology is more mature?

Waiting has its own risk, because the attacks this defends against aren’t waiting. A more sensible middle path is to adopt the parts that surface and prioritise risk under human control now, build the governance around them, and move to autonomous action as your controls and confidence catch up. That gets you moving without getting ahead of your own accountability.

How does 365 Mechanix help?

We work with financial services organisations across ANZ to adopt Microsoft’s security and AI capabilities in a way that fits the regulatory environment they operate in, spanning Azure, Copilot, Dynamics 365 and the wider stack. If agentic security is on your radar, get in touch and we’ll help you separate what’s worth piloting now from what needs governance groundwork first.

This blog is intended as general guidance only and does not constitute legal or compliance advice. We recommend consulting your compliance team or legal advisors for advice specific to your organisation.