Skip to content
Tim Frenzel

// Insight

The grader never opens the sources

7 min read
evaluationbenchmarkagentsmodel-risk

Benchmarks for knowledge work have a measurement problem that benchmarks for code do not. A compiler settles whether a patch builds. Nothing settles whether a 12-page memo is right. Somebody has to judge it. The design of that judgment decides what the score can certify. AA-Briefcase, announced by Artificial Analysis on 18 June 2026 and now carrying 15% of its Intelligence Index, is the most serious attempt at this I have read. Its headline finding is that the leading model satisfied every rubric check on 3% of tasks, with no model clearing 50% on 31 of the 91. Two deliberate design choices sit underneath those numbers. Both push them up.

What the benchmark runs

Four private scenarios covering data science, product management, banking operations and heavy industry strategy, built over months by practitioners from firms including Google, McKinsey and BCG. Each is a multi-week project holding 2 to 5 tasks per week, 91 tasks in all. The agent gets nearly 2,000 source files, among them more than 3,500 emails and 25,000 Slack messages, which Artificial Analysis describes as fragmented, messy and often carrying realistic contradiction. It then works in an offline sandbox for up to 500 turns with a single code-execution tool and submits real artefacts: spreadsheets, slide decks, PDFs, Word documents.

That is a much harder setup than a question-answering suite. The coverage I have read of it stops at the leaderboard. The harness is open source and the grading prompts are published, which is more disclosure than most benchmarks offer. Read those prompts and the ceiling becomes visible.

What each side of the benchmark can read
Agent, per task~2,000 source files3,500 emails, 25,000 Slack messages500 turns, offline sandboxone deliverableRubric judge, per checkthe deliverable alonetask text and one rubric itembinary pass or fail
The agent reconciles a contradictory corpus. The judge scoring that work is shown the output and never the corpus. Amber marks the side that issues the verdict.

The grader never opens the sources

Correctness on AA-Briefcase is scored by binary rubric checks. Those checks are decided by a language model. Artificial Analysis publishes the grading prompt. One line in it does most of the work. The judge is told that beyond the task instructions and the rubric item, “you only ever receive the submitted artifact itself, never the external source files it cites.”

The prompt then tells the judge what to do about that. It must not fail an item merely because it cannot open a cited source. It should assess citations on whether they are present, specific and well formed in the submission. In a benchmark built on contradictory sources, a confident and well-formed citation of a misread number passes the check that was supposed to catch it. The grader can test whether a deliverable is internally coherent and properly evidenced in appearance. It cannot test whether the evidence is true.

Two things keep this from being sloppiness. The rubric panel is three judges, Claude Opus 4.8 at max effort, GPT-5.5 at high reasoning and Gemini 3.1 Pro Preview at high reasoning, with one sampled per check and a given check always drawn by the same judge, which Artificial Analysis says reduces bias toward submissions from the same model family. Pairwise grading for analytical quality and presentation runs its own three-model panel. The design is careful. The constraint looks economic to me. Cost per task already reaches $31 at the top of the table. That covers generation alone. Putting the relevant corpus in front of three judges for every check on every submission would move grading from a rounding error to the dominant line item.

So the ceiling is structural. Grading the shape of an argument scales, while auditing that argument against its evidence does not. That is the same constraint this blog keeps running into from other directions, where the cheapest reliable verifier decides how far a production agent is allowed to go.

How one AA-Briefcase Elo is assembled
AA-Briefcase EloRubric pass rateBinary, deliverable onlyAnalytical quality EloPairwise, two outputsPresentation EloPairwise, two outputs
Rubric results become Elo through synthetic head-to-head matches. One branch tests correctness. Two rank preference. Every branch is judged by a three-model panel. None of them opens a source file.

The multi-week benchmark runs one week at a time

The second choice is stated plainly in the methodology. Models “currently complete each task in an independent run, without carrying over their own prior submissions.” Later weeks can receive standardized base-case files, the same reference work products handed to every model. Each task stays independently runnable while the content keeps its continuity.

The reasoning is sound. If week 3 started from whatever each model produced in week 2, you would be scoring accumulated drift and no two models would face the same task. Artificial Analysis bought comparability, which is the right trade for a leaderboard.

It also switches off the failure channel that decides real deployments. A 396-task reliability study covering 23,392 episodes across 10 models found aggregate pass@1 falling from 76.3% on short tasks to 52.1% on very long ones, a 24.3 percentage-point decline that runs faster than an independent-errors model predicts. Its stated mechanism is that errors correlate positively across steps, since a confused agent tends to stay confused. AA-Briefcase resets that correlation at every task boundary by construction.

Put the two choices together and the direction is consistent. The grader cannot fail a fabricated citation. The agent never inherits its own mistakes. A model that satisfies every rubric check on 3% of tasks under those conditions would do worse on the same work carried forward and checked against source. The 3% is the optimistic reading.

What the number is good for

Plenty, once it is read as a relative instrument. 32 of 196 tracked models have been run. The cost spread is the most immediately usable finding for anyone scoping a build: cost per task varies by more than 800x across the field, from over $31 for the June leader to roughly $0.04 for the cheapest entrant, which is the same economics that decides whether a routing tier pays for itself.

Privacy is the other trade. All four scored scenarios are private, including the task instructions, the input files and the rubrics, which is what blunts contamination. It also means nobody outside Artificial Analysis can reproduce a published score. What is reproducible is the machinery around it: an open-source harness called Stirrup, a Lite dataset on Hugging Face, plus a fifth public scenario that does not count toward results. That is a better disclosure posture than most of the agent market manages. It still leaves the result itself unauditable from outside.

The index integration deserves one note. Intelligence Index v4.3.2 is a weighted average over 10 evaluations in four categories, Agents at 30%, General at 30%, Coding at 20% and Scientific Reasoning at 20%. AA-Briefcase is 15% of the whole thing, half the Agents category on its own. An AA-Briefcase Elo is frozen at the moment a model joins and then normalised through a fixed range, clamp((Elo - 500) / 2000). Freezing keeps a model’s index contribution stable as the field moves, which is the point. It also means two models sitting side by side in one index were rated against different opponent pools at different dates. That is a reasonable engineering compromise. It is not a like-for-like comparison. The index does not claim to be one.

What I would add before trusting it as a deployment signal

Two things, and both are affordable at a smaller scale than the full benchmark. Grade a sampled subset against source, by giving the judge the specific files a submission cites and asking whether the claim survives. That converts citation form into citation truth on a slice. The gap between the two rates is itself the number worth publishing. Then run one scenario end to end with carry-forward, where week 3 starts from the model’s own week 2 output. Comparability dies on that slice, which is fine, because the question it answers is different. It measures whether the work compounds or decays.

Until someone runs those, treat a knowledge-work leaderboard the way a desk treats a paper Sharpe. The number is real and the protocol built it. That same protocol excluded the two things most likely to hurt you in production.

AA-Briefcase grades whether a deliverable looks properly evidenced, because grading whether it is properly evidenced would cost more than the models being tested. Read the score as a ceiling on knowledge work. Assume production sits below it.

Working on AI that needs to ship?

I help funds, fintechs, and data teams take AI from prototype to production.