Most AI benchmarks now bunch the top models within a rounding error of each other. A new evaluation aimed at the work analysts and consultants actually do tells a much less flattering story.
What Does AA-Briefcase Actually Measure?
AA-Briefcase is an agentic benchmark, launched June 18, 2026 by Artificial Analysis, that measures frontier models on realistic multi-week knowledge-work projects rather than short, self-contained problems . Instead of one-shot questions, models must turn fragmented corporate context into finished deliverables — spreadsheets, board presentations, financial models, design mock-ups, memos, and videos . The point is to test the kind of output an enterprise would otherwise pay a person to produce.
The benchmark exists because the old ones stopped discriminating. Suites like GPQA, math sets, and coding evals have largely saturated at the frontier, so top models cluster near the ceiling and no longer separate cleanly on economically relevant tasks . AA-Briefcase was deliberately designed to be unsaturated at launch — and it is. Across its task set, no model gets close to solving the problem reliably, which means the leaderboard is measuring failure modes rather than diminishing returns .
It also targets a different population than coding or science evals. The focus is on analyst, consultant, product manager, banker, and corporate strategist workflows — precisely the desk jobs where enterprise automation demand runs highest. The structure reflects that ambition:
- 91 tasks spread across four held-out, multi-week scenarios: Data Science, Product Management, Banking Operations, and Heavy Industry Strategy .
- Tasks developed over months by domain specialists from companies including Google, McKinsey & Company, and Boston Consulting Group .
- The four scored scenarios are kept private to prevent training-data contamination, with a separate public example scenario that does not contribute to official scores .
For Artificial Analysis, the release fills a gap in its own lineup. The firm now runs AA-Briefcase as a dedicated knowledge-work benchmark alongside its broader Intelligence Index (v4.1) and the GDPval-AA suite, part of a 2026 pivot toward agentic, long-horizon workloads instead of static question banks . The sections that follow unpack how the test is built, how it is graded, and why even the best model finishes far short of the bar.
Private Scenarios: Slack Threads, Emails, and Messy Corporate Archives

AA-Briefcase keeps its four scored scenarios — Data Science, Product Management, Banking Operations, and Heavy Industry Strategy — entirely private, so the prompts, source files, and rubrics never enter a training set . That secrecy is the point: a benchmark testing multi-week knowledge work is only meaningful if no model has memorized the answer. A separate public scenario, AA-Briefcase-Lite, exists for inspection but contributes nothing to official Elo, so anyone can study the format without compromising the held-out set .
What makes each scenario hard is the deliberate mess. A model is dropped into a fragmented corporate archive of nearly 2,000 source files — more than 3,500 emails, roughly 25,000 Slack messages, plus meeting transcripts, company documents, spreadsheets, standards documents, and market research . Nothing is pre-summarized. The required facts for a board deck or a financial model are scattered across threads and attachments, and part of the test is whether an agent can locate them at all before it starts producing a deliverable . This mirrors actual analyst work far more closely than a clean, self-contained prompt — and it is where most failures begin.
Execution happens in a controlled box. Each agent runs inside a week-scoped E2B sandbox driven by Artificial Analysis's Stirrup framework, with no internet access and a ceiling of 500 turns per task . The toolset is intentionally narrow:
- Code execution for parsing files, building spreadsheets, and rendering outputs.
- A finish/abandon submission tool, so a model can either commit a deliverable or give up on a task it cannot complete.
- A view-image tool for vision-capable models that need to read charts, mock-ups, or scanned documents .
No internet means a model cannot paper over gaps with a web search; the answer has to come from the supplied archive. That isolation is what lets the score reflect reasoning over messy context rather than retrieval luck.
There is one structural caveat Artificial Analysis flags openly. The scenarios are framed as weekly workflows — a sequence meant to resemble a multi-week project — but current runs are independent per task. A model does not carry its own earlier-week submissions into later tasks, so it never builds on a financial model or memo it produced the week before . For now that strips out genuine long-horizon continuity — arguably the hardest part of real analyst work — and means the benchmark tests deep single-task competence rather than accumulated project state . It is a known limitation, and a likely target for future versions.
Composite Grading via Rubric, Pairwise Elo, and a Judge Panel
AA-Briefcase grades each task on three separate dimensions rather than a single accuracy number. The first is a binary pass/fail rubric that checks instruction adherence, discovery of hidden requirements, correct evidence usage, and conclusion accuracy . The second is analytical quality, scored by pairwise model-vs-model comparison; the third is presentation quality, scored the same way . Splitting the score this way lets the benchmark separate "did the model do what was asked" from "is the deliverable any good" — two failure modes that a single rubric tends to blur.
The three dimensions then collapse into one headline number, the AA-Briefcase Elo. Analytical-quality Elo and presentation Elo come straight out of the pairwise comparisons, but rubric pass rates are not natively an Elo quantity. Artificial Analysis converts them by running synthetic head-to-head matches derived from each model's pass rate, then applying maximum-likelihood Elo aggregation to fold rubric performance in alongside the two comparison-based scores . The result is a single rank, but one assembled from three measurement systems with different statistical assumptions.
Grading is done by models, not humans. A balanced panel of three frontier judges does the scoring: Claude Opus 4.8 (max), GPT-5.5 (xhigh), and Gemini 3.1 Pro Preview, sampled across judgements to reduce single-judge bias and same-family bias . The three-vendor mix is a deliberate hedge against a judge favouring outputs that resemble its own family's style — a real risk when, for instance, a Claude model is among the candidates being ranked and a Claude model is also on the panel.
For developers reading the leaderboard, the methodology is where the caveats live. The rubric-to-Elo conversion is a novel step: collapsing pass/fail rates into synthetic matches and feeding them through maximum-likelihood aggregation is a design choice Artificial Analysis made, not an established standard, and it has not yet been through external peer review or third-party replication . The same applies to the pairwise-comparison weighting and the judge sampling. Artificial Analysis is also explicit that this remains model-judge-based evaluation, not blinded human expert review, so judge-model bias and grading drift stay open concerns despite the three-vendor panel .
None of that invalidates the ranking, but it does change how to read it. The AA-Briefcase Elo is best treated as Artificial Analysis's internal, consistently-applied composite — useful for relative comparison across models run through the same pipeline — rather than an externally audited ground truth. The weakest link is not the rubric or the comparisons individually, but the unscrutinised glue that combines them.
The Uncomfortable Stats Behind the Elo Ranking

Read in isolation, the AA-Briefcase leaderboard looks like a familiar frontier-model pecking order — but the absolute scores tell a harsher story than the ordinal one. As of June 18, 2026, Anthropic's Claude Fable 5 (Adaptive Reasoning) tops the table at roughly 1586 Elo (±15), well ahead of Claude Opus 4.8 (max) at ~1353 (±11), Claude Opus 4.7 at ~1291 (±11), GLM-5.2 from Z AI (max) at 1261 (±11), and OpenAI's GPT-5.5 (xhigh) at ~1159 (±10) . The gap between first and second place is large — over 230 Elo — and the leaderboard is dynamic, so these figures are Artificial Analysis's as-at-launch measurements rather than a fixed ranking.
| Model | AA-Briefcase Elo (±) |
|---|---|
| Claude Fable 5 (Adaptive Reasoning) | ~1586 (±15) |
| Claude Opus 4.8 (max) | ~1353 (±11) |
| Claude Opus 4.7 | ~1291 (±11) |
| GLM-5.2 (Z AI, max) | 1261 (±11) |
| GPT-5.5 (xhigh) | ~1159 (±10) |
Source: Artificial Analysis, AA-Briefcase, as at 2026-06-18.
The Elo ordering matters less than what sits underneath it. The headline number from the release is not who won, but how little winning means: Claude Fable 5 fully satisfied every evaluation criterion on only about 3% of the 91 tasks . That is the highest full-compliance rate of any model tested — and still near zero on an absolute scale. Treat 3% as the high-water mark, not a median. The leader of a knowledge-work benchmark completes the entire job, end to end, on roughly one task in thirty.
The distribution is just as telling. On 31 of the 91 tasks, no evaluated model exceeded a 50% rubric pass rate . Roughly a third of the benchmark is effectively unsolved at current capability levels — which is why AA-Briefcase launched unsaturated and is measuring frontier failure modes rather than separating near-perfect contenders.
"Even the best models fully satisfy every requirement on only a small fraction of tasks, and on roughly a third of the benchmark no model clears a 50% rubric pass rate," — Artificial Analysis (source: AA-Briefcase).
How models fail splits cleanly by capability tier, and the split is useful for anyone scoping an agentic deployment:
- Weaker models fail at the floor. They cannot reliably retrieve the required inputs from the messy source pool or produce a usable deliverable at all — the failure is mechanical, not analytical .
- Stronger models fail at the ceiling. They more often miss embedded or subtle requirements, produce incomplete analysis, or ship deliverables with formatting and file-integrity problems — the errors that surface only when a model must synthesize evidence across fragmented emails, Slack threads, and documents .
The practical reading: a higher Elo buys you fewer basic execution failures, not a model that can be trusted to finish multi-week analyst work unattended. The error mode shifts from "didn't open the right file" to "missed the one buried constraint that changes the conclusion" — and the second failure is far harder to catch in review.
The Economic Argument for Open-Weight AI: GLM-5.2 vs. Fable 5
That reliability gap costs real money, and the spread is enormous. On AA-Briefcase, per-task spend varies by more than 800x — from over $31 on average for Claude Fable 5, the leaderboard leader, down to roughly $0.04 for DeepSeek V4 Flash (Max), a gap of about 775x . The cheapest models also score far lower, so the headline number is not free intelligence. The more useful question for a deployment decision is where the price/performance curve bends — and Artificial Analysis points to the open-weight tier.
The standout case is GLM-5.2 (max) from Z AI, the leading open-weight model on the benchmark. It scored roughly 90 Elo below Claude Opus 4.8 (max) while costing under 25% as much per task . For cost-sensitive pipelines that tolerate a modest quality step-down, that is a defensible trade. DeepSeek V4 Pro (max) lands in a similar value band, and MiniMax-M3 registered 1113 Elo, just behind GLM-5.2 . None of these match the top closed models on rubric pass rate, but they reframe the question from "what is best" to "what is good enough per dollar."
Wall-clock time is the second axis, and it tracks reasoning depth rather than tool work. Frontier models take roughly 11–23 minutes per task: Claude Opus 4.8 about 23 minutes, GLM-5.2 about 16.3 minutes, and GPT-5.5 (xhigh) about 11 minutes . Tool execution — code runs, file I/O — accounts for only about 12% of that time; the rest comes from output verbosity, turn count, and inference speed .
| Model | AA-Briefcase Elo | Avg cost / task | Wall-clock / task |
|---|---|---|---|
| Claude Fable 5 (Adaptive Reasoning) | ~1586 | >$31 | ~28.5 min (est.) |
| Claude Opus 4.8 (max) | ~1353 | baseline (high) | ~23 min |
| GLM-5.2 (max) — open-weight | 1261 | <25% of Opus 4.8 | ~16.3 min |
| GPT-5.5 (xhigh) | ~1159 | — | ~11 min |
| MiniMax-M3 | 1113 | — | — |
| DeepSeek V4 Flash (Max) | low tier | ~$0.04 | — |
Extended reasoning carries a throughput cost that the leader makes concrete. Claude Fable 5 is estimated at about 28.5 minutes per task, built from roughly 139,000 output tokens at about 91 tokens per second, plus only about 3.1 minutes of actual tool execution . In other words, the top score is paid for in tokens and minutes spent thinking across fragmented, multi-document scenarios, not in heavier tooling. For teams budgeting an agentic analyst pipeline, the practical takeaway is that the highest Elo is also the slowest and most expensive line item — and an open-weight model at a quarter of the cost may be the rational default until the workload genuinely demands the top tier.
Blind Spots in the Judge-Panel Grading Methodology

The biggest caveat to every AA-Briefcase number is that you cannot check the work. The four scored scenarios — Data Science, Product Management, Banking Operations, and Heavy Industry Strategy — and their rubrics are held private to prevent training-data contamination, which means no independent researcher can reproduce or audit the headline Elo figures (source: Artificial Analysis, 2026-06). That secrecy is defensible — a public benchmark leaks into the next training run — but it is particularly consequential here because the scoring is composite rather than a single accuracy number.
Recall how the headline metric is built: a binary rubric pass rate, an analytical-quality Elo from pairwise model-vs-model comparison, and a presentation Elo, fused into one AA-Briefcase Elo through synthetic head-to-head matches and maximum-likelihood aggregation . Each of those steps — the pairwise aggregation, the synthetic rubric-to-Elo conversion, and the multi-judge weighting — is a methodology choice that has not been through peer review, third-party replication, or an independently published grading protocol. Different reasonable choices could reorder the leaderboard.
Then there is the grader itself. Despite the balanced design, this remains AI-as-judge evaluation. Grading runs on a three-model panel — Claude Opus 4.8 (max), GPT-5.5 (xhigh), and Gemini 3.1 Pro Preview — sampled to dampen single-judge and same-family bias (source: Artificial Analysis, 2026-06). The panel reduces correlated error; it does not eliminate judge-model bias or grading drift, and no blinded human expert review was conducted to anchor the model judges against ground truth. When the top-ranked model, Claude Fable 5, is itself produced by the same vendor whose flagship sits on the judging panel, the absence of a human control is worth keeping in view.
Artificial Analysis is candid about this boundary. As the team notes, AA-Briefcase "remains model-judge-based evaluation rather than blinded human expert review, so judge-model bias and grading drift remain open concerns despite the panel design" (source: Artificial Analysis, 2026-06). The practical reading: treat the leaderboard as a dynamic, vendor-published snapshot, not an audited fact.
That framing matters for how you cite these results. The roughly 1586 Elo for Claude Fable 5, the ~3% full-criteria pass rate, and the 800x cost spread are all Artificial Analysis's own measurements as at June 18, 2026 — directionally informative, internally consistent, but not externally verified. Rankings may have shifted since launch, and independent coverage beyond Artificial Analysis, Hugging Face, GitHub, and early secondary reporting from The Decoder remains thin. Use the figures as a useful frontier signal, and quote them as one firm's snapshot rather than settled ground truth.
The Demonstration Scenario: AA-Briefcase-Lite and the Stirrup Library
If you want to inspect AA-Briefcase before trusting its numbers, start with AA-Briefcase-Lite — a non-scored due-diligence scenario Artificial Analysis published on Hugging Face alongside the benchmark . It casts a fictional buyer, Halberd Capital Partners, evaluating a fictional target, Aurora Eggs Ltd, over a single week, grounded in real New Zealand egg-market data and deliberately mixing authentic and synthetic documents — including planted contradictions that an agent must reconcile . Because the four scored scenarios are kept private to prevent contamination, Lite is the only part of the suite you can actually open and read.
The scale is concrete: 4 tasks and 63 checks drawn from a pool of 67 sources and 147 files, split across shared and week-specific folders . The required deliverables show what "knowledge work" means in practice — not a chat answer but a set of files:
market_overview.tex/.pdf— a typeset market overviewmarket_model.xlsx— a working financial/market modeltarget_assessment.pptx— a board-style assessment deckbriefing.mp4andbriefing.srt— a recorded briefing with captions
The dataset is Apache-2.0 licensed, totals 104 MB, and showed roughly 465 downloads in the launch-window snapshot . It contributes nothing to the official Elo standings, but it is a faithful reference for the task structure, rubric format, and file-integrity expectations that drive the scored scenarios.
The runner is also open. Stirrup — the MIT-licensed Python agent framework that powers AA-Briefcase runs — is published on GitHub and supports code execution, web tools, MCP, document I/O, context management, and multimodal processing . For a team building its own agentic evaluation, that combination is the useful part: a permissively licensed harness plus a worked example of how messy inputs become graded deliverables, both adaptable to an internal pipeline without reverse-engineering anything.
The practical takeaway: treat AA-Briefcase's headline Elo as one firm's June-18-2026 signal, but treat Lite and Stirrup as reusable infrastructure. Clone the dataset, run it under Stirrup, and you can reproduce the deliverable-grading loop on your own tasks — which is a more durable benefit than any single leaderboard position. If you are weighing analyst-style automation, build the harness first and let your own rubric pass rates, not a vendor ranking, decide what ships.
Frequently asked questions
What makes AA-Briefcase different from benchmarks like GPQA or SWE-bench?
AA-Briefcase measures long-horizon, cross-document knowledge work — multi-week analyst and consultant scenarios built from messy, realistic inputs like emails, Slack threads, spreadsheets, and meeting transcripts. Each scored scenario draws on nearly 2,000 source files, including more than 3,500 emails and roughly 25,000 Slack messages . Benchmarks like GPQA, math sets, and coding suites test short, self-contained problems and have largely saturated at the frontier, no longer separating top models. AA-Briefcase was deliberately unsaturated at launch on June 18, 2026 , which is what makes it a harder, more economically relevant signal.
Which AI model scores highest on AA-Briefcase as of June 2026?
Anthropic's Claude Fable 5 (Adaptive Reasoning) leads at roughly 1586 Elo (±15), ahead of Claude Opus 4.8 (max) at about 1353, Claude Opus 4.7 at about 1291, GLM-5.2 from Z AI (max) at about 1261, and OpenAI's GPT-5.5 (xhigh) at about 1159 . Fable 5 led both rubric pass rate and analytical quality, while Opus 4.8 tied for the presentation lead . Treat these as Artificial Analysis's as-at-June-18-2026 measurements: the leaderboard is dynamic and will shift as more models are evaluated.
What does it mean that Fable 5 only fully passed 3% of tasks?
The rubric grades each task on binary pass/fail checks covering instruction adherence, discovery of hidden requirements, correct evidence use, and conclusion accuracy — and full satisfaction means clearing every check at once. Claude Fable 5 fully satisfied all criteria on only about 3% of the 91 tasks . Missing a single embedded requirement counts the whole task as incomplete. That 3% is the ceiling, not a median: the top-ranked model almost always drops at least one criterion in a realistic analyst scenario, and on 31 of the 91 tasks no tested model exceeded a 50% rubric pass rate .
Can I reproduce or run AA-Briefcase tasks myself?
Not the scored scenarios. The four official scenarios — Data Science, Product Management, Banking Operations, and Heavy Industry Strategy — are kept private to prevent training-data contamination, so they cannot be independently audited or rerun . What you can use is AA-Briefcase-Lite, an Apache-2.0 demonstration scenario on Hugging Face with 4 tasks and 63 checks, available for inspection and non-scored experimentation . The grading harness, the MIT-licensed Stirrup agent framework, is on GitHub and supports code execution, web tools, MCP, and document I/O for your own pipelines .
Is open-weight AI competitive with frontier proprietary models on AA-Briefcase?
On capability, not yet; on cost, clearly. GLM-5.2 (max), the leading open-weight model, scored roughly 90 Elo below Claude Opus 4.8 (max) but at under 25% of the per-task cost . That gap matters given how expensive the top end is: Claude Fable 5 averaged over $31 per task, versus about $0.04 for DeepSeek V4 Flash (Max) — more than an 800x spread . DeepSeek V4 Pro (max) offers another notable value tradeoff. For teams where per-task spend is the binding constraint, open-weight models are the clearest alternative.