AI Verification: Sixty-Three Per Cent or Four Point Two: The Same Model, Two Instruments
[No navigation links: V.E.R.A. Saturday runs direct-from-research with no Part/Chapter sequence for this episode.]
Two numbers circulated together this week under one heading, "hallucination rate," and they invited a comparison neither of them can support. Claude Fable 5 hallucinates 63.6 per cent of the time it does not answer correctly, according to Artificial Analysis's AA-Omniscience benchmark, re-scored in August 2026. A few lines further down the same sweep sat a "4.2 per cent factual-recall floor," reported by Digital Applied, an AI-consulting firm, in a study published 23 April 2026. Read together, under a shared heading, the natural inference is that one model is roughly fifteen times worse than the best available. That inference is wrong, and not because either figure is false.
Both numbers are accurate. Both are sourced to real, named, methodologically disclosed studies. Neither source, read on its own page, claims to describe the other's subject. The 4.2 per cent figure is not Claude Fable 5's result on anything. It belongs to GPT-5.5 Pro, tested by a different organisation, on a different question set, five weeks before Claude Fable 5 existed to be tested by anyone. Nobody in the chain that produced this week's pairing said anything false. The false impression sits only in the space between two true sentences, and that space is this week's subject.
What sixty-three point six per cent actually means
Artificial Analysis built AA-Omniscience as a six-thousand-question cross-domain knowledge test, and its central design choice is stated on its own evaluation page: the benchmark "rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer." Hallucination rate, in the benchmark's own words, is "the proportion of incorrect answers out of all non-correct responses, i.e. incorrect divided by incorrect plus partial answers plus not attempted." A model that says "I don't know" more often improves this figure by construction, because its abstentions leave the denominator. A model that attempts more questions, as Claude Fable 5 does, is scored against a denominator that has already excluded the safer path its more cautious rivals took.
Claude Fable 5 launched 9 June 2026 at an Omniscience Index of 40, seven points ahead of the previous leader, on higher accuracy rather than lower hallucination; Artificial Analysis's own launch coverage notes the model falls back to an earlier Claude Opus model in 9 per cent of questions. The August re-score lifted it to Index 43, 65 per cent accuracy, 63.6 per cent hallucination. A second, independent report puts that accuracy figure at 65.4 per cent, against Claude Fable 5.1's newer 67.2 per cent, and notes Fable 5.1 attempts 93.4 per cent of questions against Claude Opus 5's 87.8 per cent, trading toward confident wrong answers rather than cautious refusals. The pattern holds inside one model family before it ever reaches a comparison across two: attempt more, and the hallucination figure rises, because the discipline scoring it never claimed to reward caution over correctness. It claims only to measure what happens when a model chooses to answer.
What four point two per cent actually means, and whose number it is
Digital Applied's study, published five weeks before Claude Fable 5 launched, tested five 2026 frontier models across 5,000 prompts: 1,800 factual-recall questions, 1,600 citation-accuracy questions, 1,600 code-reference questions. The five models were GPT-5.5 in two configurations, Claude Opus 4.7, Gemini 3 Pro Deep Think, Grok 4.5 and DeepSeek V4. No Claude Fable model of any kind appears anywhere in the study, tested or named.
Factual recall, in the study's own words, means "SimpleQA-style factual questions where the answer is a single verifiable fact": a date, a birthplace, a capital city, a year of publication. Hallucination is defined as "a factual claim emitted by the model that is contradicted by ground-truth data," and the study states plainly that it does not count a refusal or an "I don't know" as a hallucination; those are treated as correct uncertainty signals. Grading combined automated string matching against curated ground-truth answers with a sampled human review of eight per cent of runs, which caught and corrected 1.7 per cent of the automated grades. Against that methodology, GPT-5.5 Pro, with extended thinking enabled, recorded the lowest factual-recall hallucination rate in the study: 4.2 per cent. One genuine gap in the published methodology is worth stating rather than resolving by guesswork: the page does not say whether a refused prompt is excluded from the denominator entirely, as AA-Omniscience explicitly excludes "not attempted" responses, or retained and simply not counted toward the hallucination figure. Nothing in this article should be read as asserting one construction over the other; the study does not say, and neither will this sentence.
Set the two studies beside each other and the second one never had Claude Fable 5 in it. The 4.2 per cent figure is GPT-5.5 Pro's own result, on its own test, months before the model this week's pairing implies it against had a launch date.
The number that cannot exist
One further fact settles the question without needing either study's methodology re-argued. Digital Applied published its hallucination-benchmark study on 23 April 2026. Claude Fable 5 launched forty-seven days later, on 9 June 2026. A study fixed in time in April cannot contain a measurement of a model that did not exist until June, under any reading of either page's methodology. This is not a construct-validity problem, where two real instruments answer different questions about the same subject. It is a temporal impossibility, and it holds regardless of anything else in this article.
That is the moment this week's working title earns its own correction. The natural reading of two numbers filed under one "hallucination rate" heading is that they describe the same model, seen through two instruments. They do not. They describe two different models, tested by two different organisations, on two different question sets, at two points in time that never overlapped.
The aggregator that already told you so
Suprmind, a commercial platform that carries Claude Fable 5's AA-Omniscience figures into wider circulation, states this directly on the same page that publishes them: "Every number below comes from a different benchmark measuring a different aspect of hallucination... No single column tells the full story. Cross-reference at least two." The same page explains that Vectara's benchmark tests summarisation faithfulness while AA-Omniscience tests refusal against fabrication, two different constructs entirely, and notes that Claude Fable 5's 63.6 per cent "reflects only non-correct responses, not overall error rates." A search this week for any single page presenting the 63.6 per cent and 4.2 per cent figures together found none. The pairing that produced this week's working title appears to be an artefact of placing two individually sourced figures under one domain heading, not a claim any one publisher made.
None of the three organisations behind these figures did anything wrong. Artificial Analysis discloses its scoring formula in more technical detail than most vendor-adjacent research this series has examined. Digital Applied states its scope, its grading method and, where it falls short, its own unresolved gap. Suprmind's own hub warns against exactly the comparison this week's material invited. The error belongs to the proximity two honest figures were placed in, not to any single publisher's own page.
What a real same-model comparison looks like
One model in this week's material was tested on both instruments, and the contrast is worth deploying precisely because it is genuine. Claude Opus 4.7 scores 5.1 per cent hallucination on Digital Applied's factual-recall task with extended thinking enabled, 9.4 per cent in its default configuration; both figures come from a direct reading of the study's own table. On AA-Omniscience, secondary reporting places Opus 4.7 at Omniscience Index 26, a 36 per cent hallucination rate and a 70 per cent attempt rate, down from Claude Opus 4.6's reported 82 per cent attempt rate and 61 per cent hallucination rate. Artificial Analysis's own model page did not return its underlying table to two separate attempts to fetch it this week, so those AA-Omniscience decimals rest on the company's own social reporting and independent commentary rather than a directly confirmed primary table; they are reported here as reported, not as confirmed to the decimal.
Even allowing for that caveat, the direction is unmistakable: one real model, one real gap, somewhere between twenty-seven and thirty-one percentage points apart, driven by the same denominator mechanism the Fable 5 figure illustrates. A model tested on a narrow, single-fact question set that it answers with confidence looks very different from the same model tested on a six-thousand-question sweep that rewards saying "I don't know." That is what a legitimate same-model, two-instrument comparison looks like: one entity, a real and sizeable gap, and a nameable mechanism behind it. It sharpens the distinction with this week's actual finding rather than softening it. Even a legitimate comparison needs the denominator question answered before the gap means anything. An illegitimate one needs a prior question answered first: is this even the same model.
The same error, twice more, and once where it did not resolve
This week's pattern was not confined to AI benchmarks. McKinsey's State of AI 2026 survey found approximately 88 per cent of organisations now use AI in regular practice in at least one business function. Gartner's 1 September 2026 survey of 1,303 organisations with at least fifty million US dollars in annual revenue found 22 per cent have successfully scaled AI across multiple business units. One commentary this week proposed subtracting the second figure from the first to produce a "sixty-six-point gap." The two surveys measure different populations, answer different questions, and were never on a shared scale; the subtraction is arithmetic performed on numbers that share nothing but the word "AI."
The Linux kernel offers the same lesson with no vendor and nothing to sell. Linux 7.3-rc1, tagged 30 August 2026, measures at approximately 40.98 million lines by the standard cloc counting tool: roughly 31 million lines of code, 4.9 million comment lines, 5.1 million blank lines. A single directory, the AMD graphics driver's register-header files, contributes about 6.52 million lines on its own, roughly 16 per cent of the entire tree. Whether that figure belongs in the same headline as hand-written logic is a counting-methodology question, not a commercial one, and it fails the same way a hallucination-rate pairing does: a number that depends entirely on what the instrument counts, quoted as if the counting method were self-evident.
Not every check this week resolved cleanly, and the honest exception deserves as much space as the three clean corrections above it. California's SB 813 carries two conflicting vote counts: the Legislature's own record gives 67 to 6 in the Assembly and 39 to 0 on Senate concurrence, while figures of 53 to 4 and 37 to 0 circulate only through the bill's sponsor and a sponsor-backed press release. That conflict is not resolved here, and it is included precisely because it is not. A verification discipline that only ever shows its successes teaches the wrong lesson. Sometimes the honest answer is that the conflict stands.
What the classifier already knew, and the question that comes before it
V.E.R.A.'s NTP claim classification asks, first, whether a claim's source exists. Episode 31 of this series ran that check against five circulating AI-agent-security statistics and found four with no traceable source at all. Episode 32 asked a second, finer question: given that a source exists, does it answer the claim it is cited for, or a different one. Gallup and CARMA, that episode found, both existed and both measured something real; they were simply asked different questions about the same broad subject, trust in AI.
This week's failure sits one layer under both checks. Artificial Analysis and Digital Applied are not measuring different constructs about the same subject, the way Gallup and CARMA were. In the specific pairing this week's material produced, they are measuring two different subjects entirely: Claude Fable 5, and GPT-5.5 Pro. No amount of construct-validity checking would have caught this, because a construct-validity check already assumes the subject has been confirmed shared. This article proposes naming the failure mode Subject Substitution, a candidate for the same review process already open for Episode 32's "Instrument Substitution" frame; it is offered here as a recommendation for that review, not as an already-settled term.
One architectural point is worth stating precisely, because it cuts a different way from Episode 32's own boundary note. Episode 32's construct-validity check sat outside V.E.R.A.'s E! Verification Service, which resolves whether a named entity exists, not whether two studies ask the same question. This week's failure sits closer to home: "does this figure describe Claude Fable 5, or a different named entity" is, at root, an entity-identity question, the kind E! already resolves for named things in its corpus. That proximity is a genuine observation about where this failure mode could eventually be automated. It is not a claim that it was. Nothing in this week's checking ran through V.E.R.A.'s NTP or E! architecture. A human researcher traced two model pages by hand, the same method any careful reader has available, using no tool more specialised than a browser and the discipline of reading past the shared heading.
The question a New Zealand buyer will meet first
This week's evidence carries no New Zealand regulatory angle, and none is manufactured to supply one; every organisation in it, Artificial Analysis, Digital Applied, Suprmind, McKinsey, Gartner, is American, and the Linux kernel and the California Legislature carry no New Zealand dimension either. The practical angle is a buyer's one, not a government one, and it is genuinely useful on its own terms.
Picture a New Zealand insurer's head of technology strategy comparing two frontier models for a customer-facing claims-triage assistant. A vendor pitch slide states that one model scores 63.6 per cent hallucination on a named leaderboard while "industry low" sits at 4.2 per cent, and frames the gap as a reason to press the vendor for answers. Before that framing survives a first question, it is worth asking which model the 4.2 per cent describes. If the slide cannot answer, and this week's tracing shows the honest answer is that it describes an unrelated model tested before the model under evaluation existed, the 63.6 per cent figure has not been shown to be nine times worse than anything. It has been compared to a number that was never about the model being bought. The 63.6 per cent figure still deserves scrutiny on its own terms, because it measures something genuinely relevant to a claims-triage tool: how often a model guesses rather than declines. That is a real question. It is a different question from the one an unrelated model's figure can answer.
The method behind every figure in this article depended on something ordinary and open: both Artificial Analysis and Digital Applied publish their scoring formulas, their denominators and their question designs in plain prose, on pages anyone can read without a login or a fee. That disclosure is what made this week's tracing possible in an afternoon rather than an investigation. It is the same principle open source has always rested on for code: a method you cannot inspect is a method you are asked to trust rather than check, and the same test applies to a benchmark as to a codebase. A scoring formula published openly can be argued with. One kept behind a vendor relationship cannot.
Move the same requirement into a domain where getting the subject wrong is not an editorial embarrassment but a legal failure, and it already has a name. Article 36 of the 1977 Additional Protocol I requires a legal review of a specific weapon, not a class of weapon in general; a review of one system cannot transfer to a different one sharing its name or type. The United Nations' Group of Governmental Experts on lethal autonomous weapons systems, meeting under the Convention on Certain Conventional Weapons, applies the same requirement at the targeting level: a system's determination that an object is a lawful target is an existence claim about that object, checkable against an external reference, not a class carried over from a different case. Confirming the subject before trusting the verdict is not a caveat to either framework. It is the first thing both ask.
Arc 4 of this series has now asked seven questions about what a verifier should output, how it reaches its grounding, whether verification effort scales, and where the check must live. This week adds an eighth, sitting logically beneath even Episode 31's existence check: before asking whether a source exists, or whether it answers the right question, confirm it is describing the entity the sentence around it claims it describes. Two honestly built, honestly published benchmarks did everything right on their own pages this week. The number that reached a reader's feed still described the wrong model, because nobody checked the subject before comparing the instruments.
Where has your own organisation compared two numbers, from two sources, without first confirming they were about the same thing, and how would you have caught it before the comparison reached a decision?
If your board or legal team needs a technical verification of an AI system's runtime logic or retrieval pipeline, message me and I will send the scope and the fixed fee.
The views expressed in this article are entirely my own, informed by more than 30 years of professional experience in architecture, security, and technology leadership in New Zealand. I write as director of Te Pono Limited; the views are personal and do not represent the position of any client, any government agency, or the New Zealand government. My commentary on legislation and policy is analytical, drawing on publicly available sources and my professional expertise in architecture, security, and AI governance, and it is politically neutral.
Andreas Hamberger is a New Zealand leader in Architecture & Security and Associate Member of the Institute of Directors. V.E.R.A. (Verified Existence & Reason Architecture) is an open-source logic engine available on GitHub. Through Te Pono he provides technical verification of AI runtime logic and retrieval pipelines for boards and legal teams; contact andreas@thehambergerreport.com for the scope and fixed fee.
This article was produced with AI assistance under my direction. Research, drafting and images pass through a pipeline I built and govern: automated gates for source verification, forbidden language and political neutrality, and my own review before anything is published. The tools include Claude, Gemini and Openart. The frameworks, arguments and editorial judgements are mine and are the same discipline I apply to the AI systems I audit for clients. AI accelerated the work; the thinking, and the responsibility for it, are mine.
[1] Artificial Analysis. "AA-Omniscience Evaluation." Accessed 3 September 2026. https://artificialanalysis.ai/evaluations/omniscience
[2] Artificial Analysis. "Claude Fable 5: Setting a New Intelligence Index Record." 9 June 2026. https://artificialanalysis.ai/articles/claude-fable-5-mythos-intelligence-index
[3] 247wallst.com. Report on Claude Fable 5.1's AA-Omniscience re-score. 1 September 2026. https://247wallst.com/cards/xpost-01m1faqkznas0rwascqg4f9p4y
[4] Digital Applied. "AI Model Hallucination Rate Benchmarks: 2026 Study." 23 April 2026. https://www.digitalapplied.com/blog/ai-model-hallucination-rate-benchmarks-2026-study
[5] Suprmind. "AI Hallucination Rates and Benchmarks." Accessed 3 September 2026. https://suprmind.ai/hub/ai-hallucination-rates-and-benchmarks/
[6] Artificial Analysis (own social reporting) and BridgeMind (independent commentary). Claude Opus 4.7 AA-Omniscience benchmark results, reported August 2026. No primary-page URL was independently confirmed this session; the underlying leaderboard table did not return to two direct fetch attempts.
[7] McKinsey & Company. State of AI in 2026 survey findings. Primary figure independently verified via Research_Index_Unified.md entries 445 to 447. No URL captured this session.
[8] Gartner. AI-scaling adoption survey, 1 September 2026 press release, 1,303 respondents. Primary figure independently verified via Research_Index_Unified.md entries 445 to 447. No URL captured this session (a direct fetch attempt returned HTTP 403).
[9] Phoronix. Linux 7.3-rc1 code statistics (cloc breakdown). 30 August 2026. Corroborated by Techzine, XDA Developers and LinuxCompatible. No verbatim URL was captured this session; see Research_Index_Unified.md entry 439 for the primary sourcing chain.
[10] California State Legislature. Assembly and Senate floor vote records, California SB 813. Re-checked 3 September 2026; see Research_Index_Unified.md entry 457. No URL captured this session.
[11] V.E.R.A. documentation. Andreas Hamberger, Te Pono Limited. January 2026. GitHub. No page numbers exist; referenced by component and architecture name.

