AI Verification: Two Harnesses, One Model, Scores 37 Points Apart
Same weights. Same benchmark, same organisation, same day of publication. On 3 September 2026, ARC Prize scored OpenAI's GPT-6 Astra on its ARC-AGI-3 Semi-Private evaluation twice, and reported two numbers: 62.7% under what it calls its Standard harness, and 99.9% under a Provider Adapter harness. Thirty-seven points, and two tenths of a point, apart, from one model tested by one organisation on one day.
Read the coverage that followed and you would think somebody buried the smaller number in a footnote. That is not what happened. ARC Prize published both figures itself, in its own post, with the costs attached to the dollar: $26,098 for the harder-scoring run, $18,817 for the easier one. It also said, in the same breath, that even a saturated benchmark would not prove artificial general intelligence.
What changed between 62.7% and 99.9% was not the model. It was the software wrapped around it during the test. That software let the model keep working memory its own maker built for it, or forced it to answer cold, the way every competitor's model had to. A verified result is a property of the system and the apparatus that measured it together, not the system alone.
What ARC Prize Actually Did, and Did Not, Hide
ARC Prize's own language for the two harnesses is precise. The Standard harness, in ARC Prize's own word, is "provider-neutral," offering "a consistent, apples-to-apples comparison across providers": every model answers through the same minimal interface, with none of the memory or context-management machinery any individual vendor has built for its own product. The Provider Adapter harness drops that constraint. It "preserves opaque reasoning state between requests," and lets the model reuse work from earlier turns by compacting longer conversations. Astra scored 99.9% when it was allowed to keep its own notes between questions, and 62.7% when it was not.
Both numbers are real. Both were disclosed by the party that ran the test, in the same post, on release day. ARC Prize also stated its own limit on what either score means: saturating a benchmark would not amount to "proof of achieving AGI," and it was explicit that it was not claiming Astra had reached that bar.
The framing that spread fastest after launch says independent analysts caught a hidden footnote OpenAI had buried. It does not survive a direct read of ARC Prize's own post. Two independent outlets checked on this specific point corroborate the same account. One reported that ARC Prize "published both numbers on the day GPT-6 Astra launched," printing a full table of every reasoning level it had run, alongside its own refusal to call the result proof of AGI. A second found that ARC Prize published three figures for the same model on its results page, the different evaluation conditions stated plainly rather than concealed, "precisely so the distinction is visible." Neither source describes a cover-up. The disclosure worked as intended. The public conversation still collapsed two honestly labelled numbers into one convenient headline.
ARC Prize went further than a one-off disclosure. It committed, in the same post, to keep reporting both harness conditions on its leaderboard going forward, each one clearly labelled.
None of this makes 99.9% fake. It makes it a different kind of number from a cross-vendor comparison, one measured under conditions a competitor's model would not have been allowed to use against it. A benchmark can be run with full transparency. The headline can still miss it. A second failure this week is more serious.
The Story That Actually Needed Catching
A separate story from the same launch week involves a different actor entirely. It is the one that actually fits the "nobody would have known without an outside check" description. Fortune's Emily Forlini compared an embargoed draft of OpenAI's own launch blog post, the version provided to press ahead of publication, against the live page and against later snapshots of it. The figures did not hold still. Astra's own ARC-AGI-3 score, as OpenAI first wrote it, read 98.6%. By the time the post went live, the same figure read 99.99%. Forlini's reporting documents at least five metrics moving on OpenAI's own page across those two days. A hallucination rate shifted from 4.2% to 2% and back to 4.2%. An ExploitBench cybersecurity figure moved from 5.5% to 11.5%. A competing model's separately reported FrontierMath score shifted twice, and a coding-capability figure was nudged from 57.7% to 57.9%.
OpenAI's own response, reported in the same article, belongs next to the finding. The company said it always verifies evaluations before publication, so "adjustments between draft and final version are normal," and that most evaluations carry noise within a few percentage points, depending on the exact checkpoint, scaffold and evaluation run used. That is a plausible account of ordinary pre-publication correction. It is also a description of a process nobody outside the company could see happening. The only reason the public knows it happened at all is that a journalist kept the embargoed draft and checked it against what shipped.
This is a different failure from ARC Prize's harness distinction. ARC Prize's numbers were disclosed, labelled and stable from the moment they were published. The problem was which one people repeated. OpenAI's own numbers were not disclosed as moving at all. The problem was that they moved, on the company's own page, with no public account of the change until someone compared archives against the live version. One is a story about audience attention. The other is a story about self-reporting.
A Different Institutional Shape: NIST's Own Disclosure
A third example from the same fortnight comes from neither a benchmark operator nor a vendor. On 17 September 2026, the National Institute of Standards and Technology's Center for AI Standards and Innovation, CAISI, published its own assessment of Z.ai's GLM-5.3. GLM-5.3 is an open-weight model from the Chinese laboratory previously known as Zhipu AI. CAISI's finding was that GLM-5.3's cyber capabilities are "significantly lower than those of current U.S. frontier models," lagging the American frontier by roughly four months on an aggregate measure.
What matters here is the method CAISI published alongside the verdict. CAISI stated exactly how it tested. Every model ran as an agent inside a harness with bash and python tools available and, in its own words, "a nudge to continue if stopped," every model set to maximum reasoning settings, scored against four named benchmarks, with results reported using 95% Wilson confidence intervals rather than a bare percentage. CAISI named its own testing setup and its own statistical method before anyone asked.
One qualification belongs here. That CAISI ran this test, and ran it this way, is confirmed by a direct reading of CAISI's own publication. Whether GLM-5.3's true relative capability sits exactly where CAISI's own index places it is a different question. On the evidence available, that is a conclusion resting on one organisation's own assessment, with no independent replication located anywhere this week. The two claims differ in size. Collapsing them into one flattens the distinction CAISI's own confidence intervals exist to preserve.
Three institutions had three different reasons to disclose, inside one short month. A non-profit benchmark operator disclosed its own methodology for a benchmark it runs. A government standards body disclosed its own methodology for a model it did not build and has no commercial stake in. A commercial vendor's own numbers moved without a disclosed reason, caught only because someone had kept a copy of what came before.
What This Does, and Does Not, Mean for V.E.R.A.
This series named the shape of this problem at Episode 7, long before GPT-6 Astra existed. The Verification Gap is the claim that a vendor's assertion about what a system can do is a different kind of statement from an independently checkable demonstration of what it actually does. The distance between them is structural, not a matter of any one vendor's honesty. Episode 17 gave that claim its first empirical anchor: a case where an outside party checked a vendor's claim and found it short. This week's evidence opens the gap somewhere Episode 17 never reached, inside a single, fully disclosed measurement. ARC Prize did nothing wrong. Both its numbers were honest, labelled and stable. The gap still opened, downstream of publication, in which of two truthful figures the wider conversation chose to repeat.
A second piece of V.E.R.A.'s own architecture gets a life outside the series entirely this week. The Existence-Predication Firewall is the reason V.E.R.A. does not verify its own reasoning: a system asserting a fact about itself is not a source independent of that fact. ARC Prize's own word for its Standard harness, "provider-neutral," is a defence built on the same intuition, aimed at a different target. No model, under that harness, gets to keep the proprietary memory or context machinery its own maker built for it, because a claim checked using the claimant's own apparatus is not an independent check. Neither ARC Prize nor the National Institute of Standards and Technology built that idea from V.E.R.A. V.E.R.A. was not consulted in designing either harness. The parallel is a shared design instinct, arrived at independently, not a lineage.
This is not what happens when nobody has a V.E.R.A. Neither ARC Prize's dual-harness reporting nor CAISI's confidence-interval reporting is the kind of check V.E.R.A.'s own predicate logic performs. V.E.R.A.'s NTP and E! architecture checks a claim about an entity's existence against a corpus outside the claim itself. ARC Prize and CAISI are each running a live system under stated conditions and reporting the result: an experimental activity with disclosed methodology. The parallel sits at the level of design principle only. Both activities refuse to treat a number produced under undisclosed or self-selected conditions as meaning anything yet. Both build a mechanism to force the conditions into the open before the number is allowed to travel. The underlying problem is recognised well beyond this series. It is not evidence that V.E.R.A. solves it for benchmark evaluation, which V.E.R.A. does not attempt and does not claim to do.
What This Asks of a Buyer
Put the same shape in front of someone who has to sign a purchase order, and the abstraction gets concrete fast. A New Zealand technology procurement lead is comparing two AI vendor proposals with a colleague from the data team. The choice gets summarised in one sentence: one vendor claims 99.9% on a reasoning benchmark, the other something in the low sixties on the same test. The higher number looks like the easy call.
It is not. This month supplies the example, not a hypothetical one. The exact same model scored 62.7% under one evaluation setup and 99.9% under another, run by the same benchmark organisation, published the same day. Nobody lied. The second setup let the model keep working memory the first one withheld, which the benchmark's own operator says makes it the wrong number for a fair comparison across vendors. The 99.9% figure is not fake. It is just not the same kind of number a cross-vendor comparison needs.
There is a second question to ask before either figure goes into a business case, and this month supplies that example too. Has the vendor's own number been independently reproduced by anyone with no stake in the answer? Or does it exist only on the vendor's own blog, subject to revision nobody outside the company would notice?
Three questions do the practical work, and none needs new tooling to ask. Which harness, scaffold or configuration produced this number. Is that the configuration the benchmark's own operator calls the fair cross-vendor comparison. Has anyone other than the vendor reproduced it. A proposal that cannot answer the first question has not lied. It has told you something true about a condition it never named. The number in front of you is not wrong, exactly. It is just not telling you what it looks like it is telling you.
Two failures came out of the same launch week, and neither excuses the other. A benchmark can be run and reported with complete transparency, on the day, with the costs attached and the AGI question explicitly declined. The public conversation can still flatten it to whichever number travelled further. A vendor's own self-reported figures can move quietly after publication, regardless of how transparent a third party's parallel assessment was running alongside it. The only reason anyone knows is that a journalist happened to keep the draft.
NIST's Center for AI Standards and Innovation shows a third way to hold the same line: name the setup and the statistical method before anyone asks, and let the confidence interval say what a bare point estimate cannot. V.E.R.A.'s founding claim sits underneath all three examples: a verified result is a property of the apparatus that produced it, as much as of the thing being measured. Publication alone is not verification. This series has also now returned to the same model twice in three episodes, on two genuinely distinct findings. This particular launch week produced that pattern, not a rule anyone should expect to hold generally.
V.E.R.A.'s own codebase is public on GitHub under GNU General Public License version 3, inspectable rather than merely asserted, and that remains true. This week's open-source evidence is different: ARC Prize's own forward commitment, made in the same post that reported both of Astra's scores, to keep publishing every evaluation condition on its leaderboard, each one clearly labelled, rather than folding them into a single headline figure. It is a standing, checkable promise: anyone can return to the leaderboard next quarter and confirm whether the labelling held. A public codebase and a public leaderboard are different kinds of artefact. Both convert a claim about capability into something a party with no stake in the answer can go and check: the open-source instinct, applied one level above the source code itself.
Move the same requirement into a domain where getting it wrong is a legal failure. The disclosure discipline applies with little room for error. Lethal autonomous weapons accountability frameworks have asked this week's question for longer than any benchmark dispute. Can a system's own account of its performance stand in for an external check someone else can verify? The Verification Gap, named at Episode 7, is that question in general form. The North Atlantic Treaty Organisation's 2021 Principles of Responsible Use of Artificial Intelligence in Defence answer it for one domain. They name reliability and traceability as conditions a system must demonstrably meet, terms a vendor's assertion cannot substitute for. Neither the Alliance nor any national defence ministry was consulted on this argument. The same discipline, disclose the conditions before the number means anything, is load-bearing wherever a claimed capability meets a decision that matters.
Neither number this week was fake. The last time your own team compared two vendors' benchmark claims side by side, did anyone in the room ask which harness produced either number, or did the higher one just win?
This is one of seven weekly series in The Hamberger Report; if your board or legal team needs a technical verification of an AI system's runtime logic or retrieval pipeline, message me and I will send the scope and the fixed fee.
• • •
The views expressed in this article are entirely my own, informed by more than 30 years of professional experience in architecture, security, and technology leadership in New Zealand. I write as director of Te Pono Limited; the views are personal and do not represent the position of any client, any government agency, or the New Zealand government. My commentary on legislation and policy is analytical, drawing on publicly available sources and my professional expertise in architecture, security, and AI governance, and it is politically neutral.
• • •
Andreas Hamberger is a New Zealand leader in Architecture & Security and Associate Member of the Institute of Directors. V.E.R.A. (Verified Existence & Reason Architecture) is an open-source logic engine available on GitHub. Through Te Pono he provides technical verification of AI runtime logic and retrieval pipelines for boards and legal teams; contact andreas@thehambergerreport.com for the scope and fixed fee.
This article was produced with AI assistance under my direction. Research, drafting and images pass through a pipeline I built and govern: automated gates for source verification, forbidden language and political neutrality, and my own review before anything is published. The tools include Claude, Gemini and Openart. The frameworks, arguments and editorial judgements are mine and are the same discipline I apply to the AI systems I audit for clients. AI accelerated the work; the thinking, and the responsibility for it, are mine.
[1] ARC Prize. "GPT-6 Astra." 3 September 2026. https://arcprize.org/blog/astra
[2] ARC Prize. "OpenAI GPT-6 Astra results." 3 September 2026. https://arcprize.org/results/openai-gpt-6-astra
[3] TheNextWeb. Report on ARC Prize's dual-harness scoring of GPT-6 Astra. 2026. https://thenextweb.com/news/openai-astra-arc-agi-3-harness-62-7-vs-99-9-benchmark-revisions
[4] ibl.ai. Report on ARC Prize's evaluation-condition disclosure for GPT-6 Astra. 2026. https://ibl.ai/blog/gpt-6-astra-arc-agi-3-model-agnostic-architecture
[5] Forlini, E. "OpenAI quietly boosts some of Astra's evaluation metrics, and continues to change others post-launch." Fortune. 4 September 2026. https://fortune.com/2026/09/04/openai-quietly-boosts-some-of-astras-evaluation-metrics-amid-rare-delay-in-publication-of-the-modeblog-post-announcement/
[6] National Institute of Standards and Technology, Center for AI Standards and Innovation. "CAISI's assessment of Z.ai's GLM-5.3 cyber capabilities." 17 September 2026. https://www.nist.gov/news-events/news/2026/09/caisis-assessment-zais-glm-53-cyber-capabilities

