AI Models in 2026: Why Financial Services Shouldn’t Bet on Just One
Ask most people which AI model is best and they’ll name one. Ask a bank the same question and the honest answer is: best at what, for whose data, and for how long is that going to stay true?
The market has pulled apart this year. At the top, the frontier models keep getting cleverer and, unavoidably, dearer. Underneath them, a wave of Chinese models now does most of the same work for a tenth, a twentieth, sometimes less than one percent of the price. Neither trend is a footnote. The distance between the two is where the real decisions now get made.
For a lot of businesses that’s a procurement curiosity. For financial services in Australia and New Zealand it’s a genuine bind, because the cheapest models on the market happen to be the ones you can least afford to hand your data to, and the cleverest ones are priced so you’d never dream of running them across everything. Then there’s the newer wrinkle nobody had to think about a year ago: whether you can even get hold of the best models, and for how long, is starting to depend on decisions made in Washington and Canberra rather than in your architecture review.
Which is why “which model” turns out to be the wrong thing to be arguing about. The question worth your time is how you use several of them well, keep your data where it’s supposed to be, and avoid tying your operations to a single bet that a regulator or an export rule can unpick without asking you first. That’s a platform question, not a model one. Here’s the lie of the land in mid-2026, and why it keeps pointing back to Microsoft.
Key Takeaways
- The model market has split. American labs hold the capability lead and charge for it; Chinese labs match them on plenty of everyday work for a fraction of the money.
- Claude Fable 5 tops the independent intelligence rankings right now, but at US$10 in and US$50 out per million tokens it’s about 180 times the price of the cheapest capable models. You use it where it earns that, not everywhere.
- The bargain-basement models are Chinese, and the data-jurisdiction worries that got DeepSeek pulled from every Australian government device take direct use off the table for most regulated work.
- Access to the top models is now political. GPT-5.6’s public launch was held back at the US government’s request; Anthropic’s Mythos tier sits behind export controls.
- The sensible pattern is routing: a cheap model for the volume, a frontier model only for the handful of steps where its judgement actually changes the answer.
- Microsoft Foundry is the only cloud running both Claude and GPT frontier models side by side, alongside Meta, Mistral, DeepSeek and Microsoft’s own models.
- For a bank, the platform matters more than the model. Foundry keeps identity, audit and data residency the same no matter which model does the work.
Two Races, Not One
For a couple of years the model business behaved like a single leaderboard. Everyone chased the same score, and the only live question each quarter was whose name sat at the top. That framing is finished. There are two races now, and they’re being run for different prizes.
The one everyone watches is the frontier. On the independent Artificial Analysis index, Anthropic’s Claude Fable 5 and OpenAI’s GPT-5.6 Sol trade the top spot, with Claude Opus 4.8 a whisker behind. These are extraordinary pieces of engineering. They reason across long, tangled problems and drive agentic work that would have read as wishful thinking in 2023.
What’s changed is the sticker. Fable 5 runs to roughly US$10 per million tokens going in and US$50 coming out. Opus 4.8 sits at US$5 and US$25. GPT-5.6 Sol is in the same bracket. Point one of these at every task in the business and the invoice does something unpleasant very quickly. They were built for the hard problems, and the economics only make sense if that’s what you save them for.
The other race is about price, and Chinese labs are winning it comfortably. A July 2026 analysis put frontier-grade models from DeepSeek, MiniMax, Kimi, Qwen and Xiaomi at 10 to 60 times cheaper than the Western equivalents, with some tiers past 200 times. DeepSeek V4 Flash lists around US$0.14 in and US$0.28 out. Set that against Fable 5’s US$50 output and you’re looking at a model that costs less than one percent as much yet lands within a few points of the flagships on a lot of benchmarks.
And this isn’t a launch promotion that’ll expire. The savings are baked into the design: architectures that fire only a sliver of the model per token, caching that borders on free, cheaper compute, and an open-weight strategy that treats price as a way to win developers rather than a source of margin. For most of the world it’s an easy trade. Push the high-volume, low-stakes work to the cheap model and keep the change. In regulated financial services, it’s nothing like that simple.
The Catch Nobody in FSI Can Ignore
The trouble with the cheapest models has nothing to do with how good they are. It’s about where your data comes to rest, and who can reach it once it’s there.
Send a prompt to a Chinese provider’s hosted API and it’s processed on infrastructure that answers to Chinese law, which can compel access to data held on domestic servers. That’s not a hypothetical in this country. Early in 2025 the Department of Home Affairs judged DeepSeek to carry an unacceptable level of security risk and ordered it off every non-corporate Commonwealth device, with the states and territories falling in behind within days. The reasons given were data security, personal information sitting offshore, and privacy-law compliance.
Now picture the follow-up. If that’s the government’s stance on its own laptops, try explaining to APRA or ASIC why customer financial data was routing through a model hosted in a jurisdiction your own government has flagged. The saving doesn’t survive first contact with a third-party-risk review. It was never really a saving; it was a deferred problem with a low headline number.
There is a legitimate route in. Most of these models are open-weight, so you can self-host them on infrastructure you control and keep the data onshore. Worth knowing. But that’s a build with real governance and engineering behind it, not a swipe of the corporate card, and the gap between the advertised price and the compliant price is the whole story.
You Can’t Assume You’ll Keep the Keys
Stay with the Western frontier and you sidestep the jurisdiction problem, but there’s a fresher risk worth saying out loud. The best AI has become something governments want a hand on.
OpenAI held GPT-5.6 back from public release at the US government’s request, opening it first to a vetted handful of partners before the Department of Commerce cleared a wider launch in July. Anthropic’s most capable Mythos-tier models have been sitting behind export controls that govern who can touch them, and where.
The moral for anyone who’s bet the shop on one model is not comfortable. Frontier access now rides on decisions taken a long way above your architecture board, and it can move fast. Build a critical process on one specific model from one specific lab and a policy shift you had no part in can throttle it or turn it off. In any other part of the business you’d call that concentration risk and be told to fix it. AI shouldn’t get a pass.
The Answer Isn’t one Model.
Stack those three pressures together, the frontier too dear to run flat out, the cheap stuff too risky to feed, and access that can be revoked, and the one-model strategy collapses under its own weight. What takes its place is routing, and it’s less a product than a way of thinking about the work.
Almost no business process is a single task. It’s a chain. Read the document, pull out the fields, check them against policy, draft the reply, make the call. Those links don’t all need the same muscle. Routing just means handing each one to the cheapest model that can do it properly, and holding the expensive frontier reasoning back for the step or two where it genuinely tips the outcome.
Get that right and a model that would bankrupt you at a million cases becomes affordable, because it only ever touches the five percent of the job where its judgement matters. The rest runs on something fast and cheap. The catch is that the maths only works if you can move between models cleanly, under one set of controls, without bolting a fresh vendor contract, login system and audit trail onto every one of them.
What it looks like on the ground
Take a document-verification job we rebuilt with a leading Australian bank. Their settlement team was eyeballing uploaded documents by hand, cross-checking customer details, working through a list of criteria one file at a time. Long handling times, outcomes that drifted depending on who was on shift, and a mountain of admin nobody enjoyed.
What replaced it doesn’t hang off one clever model doing everything. A capable document-understanding model reads each file and lifts the details that matter, names, addresses, dates, policy specifics. That structured output drops into rule-based logic that makes the pass-or-fail call, with the lot written to structured records and a plain-language summary so it can be audited later. The light extraction work and the higher-order judgement each go to the tool suited to them. Not one model stretched thin across both.
The payoff was the sort of thing that reads well in a board pack: checks that used to eat hours now finishing on upload, steadier outcomes, and an audit trail that actually stands up. The models involved aren’t the point. The point is that a workflow built with care uses more than one of them, on purpose, and that doing so in a regulated shop needs a platform where the routing is safe and governable rather than a science experiment.
Why All of This Keeps Landing on Microsoft
If the game is routing between many models under serious governance, then “which model” stops being the interesting question and “where do I run all of them without losing control” takes over. On that question the Microsoft stack has moved ahead for regulated business.
Since June 2026, Anthropic’s Claude models have been generally available in Microsoft Foundry, which makes Azure the only cloud offering both Claude and GPT frontier models in the same place. Meta’s Llama, Mistral, DeepSeek and Microsoft’s own MAI models sit alongside them, well over 1,900 in the catalogue. You’re not wedded to one lab. You get most of the field, and you can swap or blend as the market reshuffles, which it will.
The piece that matters most for routing is already in the box. Foundry’s model router reads an incoming request and sends it to the right model in a pool you’ve configured, cheap and quick for the routine, frontier for the hard yards, with fallback if something’s unavailable. That fallback is the answer to the access problem: gate one model and the workflow reroutes rather than falls over.
So Where Does That Leave You?
Already on the Microsoft stack
You’ve got the best hand at the table. The multi-model capability is already sitting inside the platform you run. The move is to stop hunting for the one right model and start finding the workflows where routing between a cheap model and a frontier one changes the numbers. Pick one high-volume process, prove the pattern, then spread it.
Still weighing it up
For regulated business the case for consolidating on Microsoft has rarely been this clean. Every major frontier model, routing at the task level, data kept in-region, and governance that lines up with what APRA and ASIC will ask for, all under one roof. In financial services especially, you will struggle to assemble that from parts and end up ahead.
Chasing the lowest possible token price
Be straight with yourself about the full cost. For non-sensitive, high-volume work the value models are genuinely tempting, and Foundry will give you governed access to several of them. But the moment customer or regulated data is in play, the real price of the cheapest option includes all the residency, jurisdiction and audit work it takes to use it without getting burned. Cheap tokens attached to an expensive compliance problem is not a bargain, whatever the spreadsheet says on the first pass.
The Short Version
We’ve been making the same argument to clients for a while, and this year hasn’t dented it. The win was never about crowning a model. It was about earning the ability to use many of them, matched to the job, governed properly, and shielded from a market that rearranges itself every few months.
What’s shifted in 2026 is the cost of getting it wrong. The price gap between models is vast, the security lines around them are hardening, and access itself has turned into something that can be handed over or taken back. For financial services in Australia and New Zealand, standing on a platform that gives you every model and the controls to run them safely isn’t a nicety. It’s the risk-managed choice, and it happens to be the cheaper one once you count everything.
If you want to talk through what a multi-model approach looks like in your environment, or how to get governed routing working on a real process rather than a slide, get in touch with the 365 Mechanix team. We work with organisations across Australia and New Zealand to turn all of this into something that runs in production.
FAQs
Which AI model is best in 2026?
Wrong question, and that’s not a dodge. On the independent rankings Claude Fable 5 and GPT-5.6 Sol lead on raw intelligence, but they’re priced for the hardest problems, not the daily grind. Plenty of far cheaper models handle everyday work fine. The useful move is matching the model to the task instead of looking for one winner.
Why are Chinese AI models so much cheaper?
Structural reasons, mostly. Architectures that activate only a fraction of the model per request, caching that’s close to free, cheaper compute and power, and an open-weight strategy that prices for adoption rather than margin. It’s real, not a temporary discount, which is exactly why it’s worth understanding rather than dismissing.
Can Australian financial services actually use models like DeepSeek?
Not casually. Using the hosted API means your data is processed under Chinese jurisdiction, which is why DeepSeek got pulled from Australian government devices on security grounds. Because many of these models are open-weight, you can self-host them onshore, but that’s a governance and engineering project, not a sign-up form.
What is model routing?
Sending each step of a workflow to the model that suits it, rather than forcing everything through one. Cheap, fast models take the routine volume; frontier models handle only the complex reasoning where they change the result. It’s what makes frontier capability affordable, by using it sparingly.
What is Microsoft Foundry?
Microsoft’s platform for building, running and governing AI models and agents on Azure. Large model catalogue, including Claude and GPT frontier models in one place, plus a router, enterprise identity through Microsoft Entra, unified audit logging, and deployment options that decide where your data is processed.
Why does the platform matter more than the model for FSI?
Because the models change every few months and your obligations don’t. A platform like Foundry keeps identity, access, audit and data residency consistent whichever model runs a task. That consistency is what makes a multi-model approach governable, and what supports the operational-resilience outcomes regulators expect to see.
How does 365 Mechanix help?
We work with organisations across ANZ to design, build and govern AI on the Microsoft stack, from Copilot and Power Platform through to custom agents and multi-model routing in Foundry, with a focus on financial services and regulated environments. If any of this is live for you, get in touch.
This blog is intended as general guidance only and does not constitute legal or compliance advice. We recommend consulting your compliance team or legal advisors for advice specific to your organisation.