The Benchmark You Cannot Audit
Return to Part 0: Table of ContentsPrevious Article: Part 28, The Sandbox Was Never a Wall: A Government Institute Measured the Failure Four Months Before It Happened
Anthropic's launch page for Claude Opus 5 says the model "surpasses all other models" on a benchmark called Frontier-Bench v0.1, and "more than doubles Opus 4.8's performance at a lower cost per task".
Go and look at that benchmark's public leaderboard. It lists one model. Verified results: zero. Self-reported results: one. Status: unverified. The single entry is Opus 5, scoring 0.433, reported by Anthropic, on a harness Anthropic selected and infrastructure Anthropic runs.
Comparative figures for the competitors do exist. They come from Anthropic's own system card. So "surpasses all other models" rests on numbers one party produced for every model in the comparison, and no independent party has scored any of them.
Nothing about that is a scandal, and nothing in what follows suggests dishonesty. The numbers may well be accurate. What they are not, yet, is audited. And that distinction is the one your procurement process almost certainly does not make.
Who ran the numbers
On 24 July 2026, Anthropic launched Opus 5 at five US dollars per million input tokens and twenty-five per million output, the same price as Opus 4.8 and half of Fable 5's rate. It was the company's fourth model in under two months.
The launch materials led with benchmark results: Frontier-Bench v0.1, CursorBench 3.2, Zapier AutomationBench, OSWorld 2.0, ARC-AGI-3, and a set of life-sciences evaluations. They are presented as one continuous wall of evidence. They are not one thing.
Frontier-Bench v0.1 is the self-reported result described above. CursorBench 3.2 is built, run and scored by Cursor, the company whose coding-agent product the benchmark measures; independent benchmark trackers list it as display-only for exactly that reason. AutomationBench is administered by Zapier. In each case the party keeping score has a commercial interest in the score.
Then there is the exception, and it is worth dwelling on because it shows what the alternative looks like. ARC Prize, the nonprofit that designs the ARC-AGI series, did not accept a figure from Anthropic. It ran the evaluation itself, on its own harness, and published the result with its own verification badge attached. Opus 5 scored 30.2% on ARC-AGI-3, against a previous record of 7.8%. That is close to a fourfold improvement on a benchmark built specifically to resist the kind of training that inflates the others, and ARC Prize's own notes record the model solving five environments no model had beaten before, translating tasks into algebraic notation and deriving reflection equations unprompted.
That result is real, it is independently administered, and it is the most impressive number of the launch. It was also not the centrepiece of the marketing. The strongest evidence Anthropic had was the evidence Anthropic did not commission.
So the first question for a buyer is not "how did the model score". It is "who held the stopwatch", and the answer differs for every row on the slide.
Independent is not one thing
Ask that question properly and it turns out to have more than two answers.
Within days of launch, an independent researcher, Guanghan Ning, ran Opus 5 on Witness, a separate interactive puzzle platform, and found a much narrower gain: a score of 43.4, statistically close to Kimi K3 and to Fable 5 rather than clearly ahead of them. Ning's reading is that the ARC-AGI-3 leap may reflect training on data adjacent to that task format rather than a broad improvement in reasoning. The AI researcher Greg Kamradt publicly disagreed, arguing that one puzzle type using familiar mechanics cannot settle the question in either direction.
Both of those researchers are independent of Anthropic. They ran different tests, measured different things and reached different conclusions about what the same model can do. Neither result cancels the other, and ARC Prize's 30.2% stands. What the disagreement establishes is that "independently verified" tells you who administered a test, not what the result means.
The gradations get finer. Artificial Analysis, the benchmark aggregator whose Intelligence Index is widely quoted in enterprise evaluations, scored Opus 5 at 61, narrowly ahead of Fable 5 at 60, GPT-5.6 Sol at 59 and Kimi K3 at 57, with substantial leads on two agentic benchmarks at materially lower cost per task. Its own published article on the model opens the relevant passage like this: "We supported Anthropic to evaluate Claude Opus 5 ahead of release: it sets the highest GDPval-AA v2 and AA-Briefcase scores so far."
Read the first half of that sentence carefully, because it is doing as much work as the second. It is a disclosure, and Artificial Analysis deserves credit for making it in plain language at the top of its own analysis. It also means the evaluation was not conducted at arm's length by a party with no advance knowledge of the model. Whether money changed hands is not stated either way. This is independent aggregation with a disclosed pre-release relationship, which is a real evidentiary category, and it is not the same category as a blind third-party audit.
So go looking for the evaluator with no relationships at all. You will not find one. Epoch AI, the research nonprofit that runs its own composite capability index, publishes a transparency page listing its commercial work: benchmark development, verification and a data audit for OpenAI, model evaluations and a joint report with Google DeepMind, evaluations for xAI, and a small benchmark pilot with Anthropic. The methodology behind its capability index was funded by Google DeepMind.
That page is not the disclosure of a problem. It is the answer to one. At this level of the industry nobody is unentangled, and an evaluation market where the serious evaluators all have vendor relationships is not a scandal either; it is what a specialised field looks like. Which means "is this evaluator independent" was always the wrong question, because the honest answer is never a clean yes. The question that survives is "what has this evaluator published about its entanglements, and where can I read it". Artificial Analysis answered that in one sentence at the top of its article. Epoch AI answered it with a standing page. A vendor's own leaderboard does not answer it, because the question does not arise: everyone already knows who ran it, and that is exactly the problem with using it as evidence.
So "Opus 5 tops the leaderboard" is a sentence whose meaning depends on which leaderboard, measured how, by whom, and under what relationship. On a launch page, none of that is presented to the reader as a choice, because on a launch page it never is.
One more thing did not appear on any launch slide. On Artificial Analysis's closed-book factuality benchmark, Opus 5's hallucination rate rose to roughly 50%, some fourteen percentage points higher than Opus 4.8, while its raw accuracy also improved. The mechanism is behavioural rather than a loss of knowledge: the model declines to answer less often and attempts more, which produces more correct answers and more fabrications at the same time. Whether that trade is good or bad depends entirely on what you are using it for. A research assistant that guesses more is a different product from one that says it does not know, and nothing in the aggregate index tells you which one you are buying.
The window closed in six days
Then the ground moved.
On 30 July, six days after the Opus 5 launch, OpenAI cut the price of GPT-5.6 Luna by 80%, from one dollar and six dollars per million input and output tokens to twenty cents and one dollar twenty, and cut Terra by 20%. OpenAI attributed the cuts to serving-cost efficiency gains from kernel-level work. Press coverage tied them directly to competitive pressure from Anthropic's pricing.
Put the week together from a buyer's seat. On 24 July there was a leaderboard position and a price. By 26 July there was a second independent evaluation reaching the opposite ranking and a replication attempt narrowing the flagship result. By 30 July a competitor's cheapest tier had fallen by four fifths. Any organisation that anchored a shortlist to the launch-day picture was working from a snapshot that had materially changed on both the capability and the price axis before the paperwork cleared legal.
This is where the pattern has a name. The Vendor-Decides Pattern describes enterprise governance emerging from accumulated vendor defaults rather than deliberate architectural choice: deploying before governing. It is usually observed after the purchase, in the settings nobody changed and the permissions nobody scoped. What this launch shows is that it starts earlier than that. When the only evaluation evidence available at the decision point is the evidence the vendor chose to commission, publish and frame, the vendor's disclosure decisions have already become your evaluation criteria. You did not choose Frontier-Bench. It was chosen for you, and it was the only thing on the table.
That is a governance gap at the procurement stage, one step upstream of where this series has usually found it. The familiar version is organisations deploying AI faster than they govern it. This is organisations buying AI faster than they can evidence it.
What this looks like from New Zealand
Two specifics for readers here, one about guidance and one about geography.
The Department of Internal Affairs publishes procurement guidance for generative AI as part of the Public Service AI Framework, last updated in February 2025. It tells agencies to know the market, the suppliers in it and their offerings, and to build evaluation criteria that are explicit about pre-conditions and include an exit strategy. Read against the week described above, what is worth noticing is a matter of vintage rather than quality: the guidance predates the practice of launching a frontier model against benchmarks the vendor built, and it contains no specific treatment of vendor-supplied performance claims. That is a description of what the document covers, not an assessment of it, and the advice it does give, evaluation criteria set in advance with an exit strategy attached, is exactly the discipline this article is arguing for.
The second specific is simpler and harder. Opus 5 reached Amazon Bedrock at launch in four regions: Northern Virginia, Melbourne, Ireland and Stockholm. There is no AWS Bedrock region in New Zealand, because there is no AWS Bedrock region in New Zealand to have. For a New Zealand organisation with data-residency constraints, the nearest regional endpoint for this model is in Australia. That is not specific to Opus 5 and it is not new, but it is the practical shape of the decision here, and it belongs in the evaluation alongside the benchmark table rather than being discovered afterwards.
For completeness, the most developed regulatory answer to any of this sits in Europe, where general-purpose AI providers must maintain technical documentation covering evaluation methodology and results. That documentation runs to regulators and to downstream providers building products on the model. A prospective enterprise customer deciding which model to buy has no equivalent statutory right to inspect the methodology behind a public benchmark claim. New Zealand, with a voluntary framework, has no equivalent either. Where those levers should sit is a question for others; this is a description of where they currently sit.
Four things to change before the next evaluation
None of this argues against buying frontier models, and none of it argues that Opus 5 is a poor one. The evidence suggests it is very good. The argument is about what you are entitled to conclude from a launch page, and it resolves into four changes you can make to a procurement process this month.
Ask who administered each number. Not who published it. Who ran it, on whose harness, on whose infrastructure. Build a column for it in your evaluation matrix. On the launch described here, that single column separates one genuinely independent result from five vendor-run ones, and it takes an afternoon to fill in.
Ask every evaluator what it has published about its vendor relationships. Not "are you independent", which invites a yes that means nothing. Artificial Analysis disclosed in one sentence at the top of its own analysis; Epoch AI maintains a standing page. Treat a published answer as the standard, and treat the absence of one as information.
Weight disaggregated results over composite scores. A composite index is a weighted opinion about what matters. The hallucination trade-off above is invisible in the aggregate and decisive for some use cases. Find the two or three benchmarks closest to your actual workload and read those instead.
Put a re-evaluation checkpoint in the contract. Not a review clause, a dated checkpoint with a named owner, timed for roughly ninety days after signature, when the independent replications exist and the launch-week pricing has settled. The market moved twice in the six days after the event described here. A three-year commitment that cannot be revisited for three years is a bet that it will stop.
There is an open-source dimension here, and it points at the fix. Part of why no independent party has reproduced the Frontier-Bench number is structural: reproducing a closed model's score means buying access on the vendor's terms, and the vendor alone decides what to publish. Open evaluation infrastructure removes half of that gate. The harness Frontier-Bench itself runs on, mini-SWE-agent, is open source; so is EleutherAI's language-model evaluation harness; so is the ARC-AGI task suite that let ARC Prize run its own test rather than accept a submitted figure. Open-weight models remove the other half, because anyone can score a model whose parameters are public without asking the maker's permission. None of that guarantees an honest number. It does mean the number can be checked by someone other than the party that produced it, which is the entire argument in one sentence.
The implication on the sovereignty and defence-procurement side is worth stating directly. The problem this article describes, a buyer with no way to check the supplier's own performance claim, was answered in defence acquisition decades ago by refusing to let the supplier keep the score. The United States Congress created the office of the Director of Operational Test and Evaluation in the early 1980s to assess major systems independently of both the program office that buys them and the contractor that builds them. Independent verification and validation is a standing discipline across defence, aerospace and safety-critical software for the same reason: the evaluator must not report to the builder. Frontier models are entering the same procurement channels with none of that machinery attached. For a small country the question is whether it can evaluate the systems it buys, or inherits the disclosure choices of offshore vendors and offshore evaluators.
- The benchmark Anthropic's launch page says Opus 5 "surpasses all other models" on, Frontier-Bench v0.1, has a public leaderboard listing one model: zero verified results, one self-reported, status unverified. The comparative figures for competitors come from Anthropic's own system card. Every number in that comparison was produced by one party.
- Most headline benchmarks at this launch were run by parties with a commercial interest in the result: Anthropic's own, Cursor's, Zapier's. The exception, ARC-AGI-3, was administered by ARC Prize on its own harness, scored 30.2% against a prior record of 7.8%, and was not the marketing centrepiece.
- "Independent" has gradations, and at this level of the industry no serious evaluator is unentangled. Artificial Analysis disclosed that it "supported Anthropic to evaluate Claude Opus 5 ahead of release". Epoch AI publishes a standing page listing its work for OpenAI, Google DeepMind, xAI and Anthropic. Both disclosures are the point. The right procurement question is not whether an evaluator is independent but what it has published about its relationships.
- The leaderboard position came with a reliability cost that no headline claim surfaced: hallucination rate up roughly fourteen points on a closed-book factuality benchmark, because the model answers more often when uncertain.
- Six days after launch, OpenAI cut GPT-5.6 Luna by 80% and Terra by 20%. Any shortlist anchored to the 24 July picture was stale before it reached a signature.
- For New Zealand buyers: DIA's generative AI procurement guidance dates from February 2025 and does not address vendor-supplied benchmark claims specifically. Opus 5's nearest Bedrock region is Melbourne; there is no New Zealand region.
So look at whatever benchmark table is currently sitting in your AI vendor evaluation. For each row, can you name who ran the test and what their relationship to the vendor is, and if you cannot, what exactly is that row doing in your decision?
The views expressed in this article are entirely my own, informed by more than 30 years of professional experience in architecture, security, and technology leadership in New Zealand. They do not represent the views of my employer, any government agency, or the New Zealand government. My commentary on legislation and policy is analytical, drawing on publicly available sources and my professional expertise in architecture, security, and AI governance. I follow the Public Service Commissioner's Code of Conduct for the Public Sector and social media guidance.
Andreas Hamberger is a New Zealand leader in Architecture & Security and Associate Member of the Institute of Directors. The Hamberger Report: Generative AI 2026 provides enterprise leaders with evidence-based analysis of the AI landscape.
I use AI tools, including Sudowrite, Claude, Perplexity AI, DeepSeek AI, ChatGPT, Grok, Copilot, Openart and Gemini, as deliberate production tools, not ghostwriters. This is consistent with my position: AI amplifies human judgement; it does not replace it. The frameworks, arguments, and editorial decisions in this series are original work. AI accelerated the process. The thinking is mine.
References
[1] llm-stats.com. "Frontier-Bench v0.1 leaderboard." Last updated 25 July 2026. Re-fetched at write time 30 July 2026. https://llm-stats.com/benchmarks/frontier-bench-v0.1
[2] benchmarklist.com. "Frontier-Bench comparative figures (attributed to an Anthropic system card, dated 23 to 24 July 2026)." 2026. https://benchmarklist.com/benchmarks/frontier_bench/
[3] Anthropic. "Introducing Claude Opus 5." 24 July 2026. https://anthropic.com/news/claude-opus-5
[4] Anthropic. "Pricing (platform documentation)." Retrieved 30 July 2026. https://platform.claude.com/docs/en/about-claude/pricing
[5] Digital Applied. "CursorBench: a first-party vendor benchmark analysis." 2026. https://digitalapplied.com/blog/cursorbench-v3-1-vendor-benchmark-analysis
[6] ARC Prize. "Claude Opus 5 results (ARC-AGI-3, independently administered)." 24 July 2026. https://arcprize.org/results/anthropic-claude-opus-5
[7] The Decoder. "Anthropic's Opus 5 and the benchmark designed to measure real intelligence (Guanghan Ning and Greg Kamradt)." July 2026. https://the-decoder.com/anthropics-opus-5-blows-past-fable-5-and-gpt-5-6-sol-on-the-benchmark-designed-to-measure-real-intelligence/
[8] Artificial Analysis. "Claude Opus 5." July 2026. https://artificialanalysis.ai/articles/opus-5
[9] Epoch AI. "Transparency (commercial relationships)." Retrieved 30 July 2026. https://epoch.ai/about/transparency
[10] The Decoder. "Claude Opus 5 costs well below Fable 5 while matching or beating it across most benchmarks (AA-Omniscience hallucination figures)." July 2026. https://the-decoder.com/anthropics-claude-opus-5-costs-well-below-fable-5-while-matching-or-beating-it-across-most-benchmarks/
[11] OpenAI. "Advancing the price-performance frontier with GPT-5.6." 30 July 2026. https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/
[12] Axios. "OpenAI cuts prices on GPT-5.6 Terra and Luna." 30 July 2026. https://axios.com/2026/07/30/openai-cuts-prices-gpt-terra-luna5
[13] Amazon Web Services. "Introducing Claude Opus 5 on AWS." 24 July 2026. https://aws.amazon.com/blogs/machine-learning/introducing-claude-opus-5-on-aws-anthropics-most-capable-opus-model/
[14] Department of Internal Affairs / Digital Government New Zealand. "Procurement and GenAI (Responsible AI guidance for the public service)." Last updated February 2025. https://www.digital.govt.nz/standards-and-guidance/technology-and-architecture/artificial-intelligence/responsible-ai-guidance-for-the-public-service-genai/genai-foundations/procurement
[15] EU AI Act, Article 53 (obligations for providers of general-purpose AI models: technical documentation including the testing process and evaluation results, provided to the AI Office and national competent authorities on request, and information made available to downstream providers). Reproduced primary text independently retrieved and verified at write time, 4 August 2026 production. https://artificialintelligenceact.eu/article/53/
[16] Director, Operational Test and Evaluation (DOT&E), United States Department of Defense. Statutory office established by the United States Congress in the FY1984 National Defense Authorization Act to assess major defence systems independently of the acquiring program office and the contractor. (Statutory office; no document URL captured in the research package.)
[17] Independent verification and validation (IV&V) as an assurance discipline in defence, aerospace, and safety-critical software: verification and validation performed by a party organisationally independent of the system's developer. (Established engineering practice; general reference, no single document URL captured in the research package.)

