Field Notes
The most capable AI models now sit behind vetting programs. Here's what that tells a bank or an insurer about where the real risk in an AI program sits.
A few weeks ago a CIO at a regional bank asked us a simple question. She'd just read that the most capable model from one of the frontier labs wouldn't be sold to her. Not at any price. If the people building these things are nervous, why shouldn't she be?
It's a fair question. The answer is more useful than the headline.
Every major lab now pairs its most capable model with a vetting program, and they all did it for the same reason.
Anthropic split its June release into two versions: one safe for general use, and a second restricted to vetted partners because it could find and exploit software vulnerabilities on its own. In September, the general version was allowed to find those vulnerabilities too, just not to write the exploit code.
OpenAI told the White House in August it would delay its next model because it "could not rule out critical cyber capabilities." By September it confirmed the model had crossed its own "Critical" threshold for cybersecurity risk: it can find and exploit unknown flaws without a person walking it through each step. Access to that capability now runs through a separate approval program.
Google's most cyber-capable model is gated the same way, limited to governments, healthcare and telecom customers who go through its own vetting program.
Three labs, three names for the same decision. The models didn't get worse. They got good enough at finding and exploiting flaws that nobody wants them loose without a way to know who's using them and why.
Here's the part worth sitting with: none of that is about the model you'll actually deploy. It's about a narrow slice of frontier capability that has nothing to do with the agent your team is building to read contracts or draft claims summaries.
The real risk in most enterprise AI programs isn't the model. It's what the agent can touch.
An agent that reads your documents will read whatever an attacker puts in front of it too. Two Microsoft 365 Copilot flaws surfaced this year that did exactly that: one let an attacker exfiltrate mail and files through a single click, the other did it with no click at all, just by opening a spreadsheet. Neither needed a "critical" cyber model. Both needed an agent with more reach than anyone had actually reviewed.
Put yourself in a Tuesday morning at a mid-size insurer. An agent triages inbound claims correspondence, pulls policy details from three systems, and drafts a response. Nobody set out to give it broad access. It accumulated broad access, one integration at a time, because each one made sense on its own. Now ask who can say, today, exactly what that agent can read and what it can act on. Most security teams can't. A CISO survey this year found fewer than half could even identify every AI agent running in their environment, let alone control what each one could reach.
That's the actual gap. Six national cyber agencies, led by CISA and the NSA, published joint guidance this spring built around exactly this: least privilege for every agent, a human in the loop for anything consequential, and a preference for reversibility over speed. Not because the models are dangerous in the way a headline suggests, but because the wiring around them usually isn't reviewed the way the rest of the stack is.
Three moves, in order.
Inventory what every agent can touch. Not what it was designed to touch. What it can actually reach today, including anything a later integration quietly added.
Log every action back to a source. If an agent pulled a number, made a decision, or sent a message, you should be able to trace it to the document or the person that authorized it. That's the same discipline a credible audit trail has always required. It just now applies to a system that moves faster than a person can watch in real time.
Put someone accountable at the end of the decision, not just at the start of the build. The agent proposes. A named person owns what happens next, especially anywhere the outcome is hard to reverse.
None of this makes the underlying model safer. That part is the labs' job, and it's genuinely out of your hands. What it does is make your use of the model defensible, which is the part that's entirely in your hands.
If you want to walk through where your own program stands against that list, we're happy to spend an hour on it. If the honest answer is that you're in good shape, we'll tell you that too.
Subscribe
One short email when a new piece goes up.