// Insight
The agent you buy is a model you cannot inventory
The 2025 AI Agent Index annotates 30 deployed agents across 45 fields each, from public documentation only. Nine authors across institutions including Cambridge, MIT, Harvard Law and Hebrew University, accepted to FAccT 2026, cutoff 31 December 2025. That is 1,350 fields, 227 of which came back empty. The subset that decides whether a desk can deploy anything is narrower and worse: 135 of the 240 safety, evaluation and impact fields hold no public information at all.
Read the shape before any single bar. The tall one is the certification set a vendor questionnaire already runs at. 25 of 30 clear it. The short ones are evidence that this particular agent was ever tested. A procurement process written for software-as-a-service passes the top bar and never reaches the bottom four, which is how an undocumented model walks through a control that is working as designed. Capability gets a number and safety gets a paragraph, the same asymmetry this blog keeps meeting from the other side, where production teams cannot cheaply verify their own agents either.
The gap is worst in the tier a bank or a fund would actually buy. Enterprise workflow agents are 13 of the 30 and leave 69 of their own 104 safety fields empty, against 42 of 96 for chat agents. They are not chat toys. They act through CRM connectors and record updates (8 of 13). Six of the 30 run at high autonomy once deployed, fired by an event like a new email or a database change, with no human involved during execution. Four carry no documented stop control despite running autonomously. Read that as an orchestration component wired into your systems of record, with an undocumented failure mode and, in four cases, no off switch anyone has written down.
Hold that against what a model has to survive on the way in. Every model touching a regulated decision lands in an inventory with a named owner, a validation date, and documentation a reviewer who did not build it can follow. That is the bar every model I have taken through a validation committee had to clear. The Index has a field for each of those questions. Across the whole corpus its non-empty fields average 14 words.
The concentration underneath is the part a quant will recognize fastest. Nearly every indexed agent runs on GPT, Claude or Gemini. Only 9 of 30 let the buyer select the provider at all. An agent estate assembled from five different vendors can still resolve to one common factor. That is a familiar way to be surprised. It does not show up on any procurement checklist I have seen.
None of this is hypothetical. 8 of 30 agents carry documented incidents or reported security concerns. Prompt injection is documented for 2 of the 5 browser agents, which is the failure class that reports zero when the defence is watching the wrong channel.
One caveat belongs here. It strengthens the paper. The Index measures disclosure, not safety. An agent with an empty row may be beautifully engineered by people who never wrote it down. The authors contacted every company and allowed four weeks for corrections. 23% offered some form of response. Only 4 of 30 said anything substantive. The silence is not an artifact of nobody asking. For a validation function the distinction collapses regardless. Undocumented and unsafe are different things. Neither of them clears a review.
What I would take into the next vendor call is short. Ask for the agent-specific evaluation, since only four vendors have one and a base model system card is a different document. Ask which provider sits underneath and whether you can change it. Ask for the stop control and make them demonstrate it, because 4 of 30 have shipped autonomous execution without documenting one. Then price the gap. Whatever the vendor will not tell you, your own monitoring and guardrails have to cover. That cost lands on your side of the contract.
A vendor agent arrives as a model with no documentation, no independent test, and often no disclosed off switch. The certifications your procurement process already checks say nothing about any of that.
Working on AI that needs to ship?
I help funds, fintechs, and data teams take AI from prototype to production.