AI Verification: The Study Is Real. It Is Measuring Something Else.

[No navigation links: V.E.R.A. Saturday runs direct-from-research with no Part/Chapter sequence for this episode.]


Two American research organisations published AI-trust figures in the same week of August 2026, and for a few days commentators treated the pair as a contradiction. Gallup, working with Bentley University, found that only 27 per cent of Americans trust businesses to use AI responsibly, down from 31 per cent a year earlier. CARMA, an international media-intelligence firm, found that public trust in AI itself had risen from 26 to 30 per cent over roughly the same period. One number falling, one number rising, both nominally about "AI" and "trust." The easy read is that one of the two surveys must be wrong.

Neither is. Gallup asked whether Americans trust businesses to deploy AI responsibly. CARMA asked a global panel whether they trust AI itself. Those are different questions, answered by different instruments, and a board that quotes rising trust in the technology as licence for its own deployment plans is citing a real study to support a claim that study never made.

That is this week's argument, and it runs deeper than one pair of surveys. Non-Traditional Predication Theory (NTP), the logic underneath V.E.R.A., asks first whether a claim's source exists at all. That check matters, and it catches a great deal: Episode 31 of this series used it to trace five circulating AI-agent-security statistics to their sources and found four with none. It does not catch what follows below. A source can exist, be honestly built, and still be cited to answer a question it never asked.

The study that says so itself

Start with the cleanest case, because it comes from the vendor's own words. Suprmind, a multi-model AI platform, publishes a Multi-Model AI Divergence Index built from 1,324 real conversation turns across 700 sessions and 299 external users of its own product, gathered over a stated 45-day window from March to April 2026. Fifty-four per cent of turns showed disagreement between the five models on the platform; 72 per cent contained one model correcting another. Financial questions produced the most disagreement, at 72.1 per cent; education questions produced the least, at 28.6 per cent.

None of that is in dispute, and none of it is hidden. Suprmind's own page states, in its own words, that the study does not measure correctness against an external ground truth. That is not a lawyer's hedge added after the fact. It is the study's own description of what it can and cannot tell a reader. It measures how often five models disagree with each other inside one company's chat product. By its own account, it does not measure how often any of them is actually wrong.

That distinction has not survived contact with the news cycle. The index's headline percentages have circulated, in several slightly different forms, framed as evidence about AI accuracy in general. The disagreement rate is real. The accuracy claim built on top of it is not the study's own claim; it was added somewhere between the vendor's page and the version that reached a reader's feed. NTP's existence check would pass this source without hesitation: Suprmind exists, the study exists, and the methodology is disclosed in more detail than most vendor research this series has examined. The failure sits one layer past existence, in the sentence that quietly swaps "how often models disagree with each other" for "how often AI is wrong."

Two surveys, two questions

Put the two surveys from the opening side by side and the shape becomes clear. Gallup's figure comes from a probability-based panel of 3,270 American adults, surveyed 4 to 11 May 2026, with a stated margin of error of plus or minus 2.4 percentage points. It asks about businesses: whether Americans trust businesses to use AI responsibly. The answer is falling, and two further Gallup findings from the same survey move the same direction: 39 per cent now say AI does more harm than good, up from 31 per cent, and 79 per cent expect AI to reduce jobs in the United States, up from 73 per cent.

CARMA's survey, the second edition of its Perceptions of AI tracking study, covers a public sample of 6,300 respondents across 19 markets alongside a media analysis of 6,006 articles across 500 outlets. It asks about the technology directly: whether people trust AI itself. That answer is rising, from 26 to 30 per cent, and for the first time trust has overtaken excitement as the public's dominant stated feeling about AI, 33 per cent against 29 per cent.

Both figures are accurately reported. Five separate outlets carried the CARMA numbers this research checked, and all five agree with each other, which sounds like independent confirmation until it becomes clear that all five are relaying the same press release rather than five separate measurements of the underlying survey. That is a real gap, separate from whether the reported figures are accurate, and it deserves to be said plainly rather than smoothed over: it confirms the release exists and says what it says, not that five different research teams weighed in. CARMA's own methodology, unlike Gallup's, does not publish fieldwork dates, a sampling method, or a stated margin of error anywhere this research reached.

None of that changes the point this article is built on. Gallup measured confidence in an industry's conduct. CARMA measured confidence in a technology. A commentator who reads "trust in AI is rising" against "trust in businesses using AI is falling" and calls it a contradiction has combined two real, honestly reported numbers into a sentence that neither survey supports, because neither one was asked the other's question.

The one that doesn't belong

Not every plausible-looking case this week turns out to belong in this cluster, and the one that does not belong is worth walking through in detail, because the mistake it would be easy to make is instructive.

A claim about agentic AI deployment failure rates has circulated widely enough this year that several commentary pieces treat it as settled background fact. Traced to its origin, it comes from a single Medium post by an individual consultant, describing that person's own client work supplemented by unspecified secondary reading. There is no institution behind it, no disclosed sampling method beyond a description of the author's own business, and no independent review of any kind. A cluster of other figures from other individual authors and vendor blog posts circulates alongside it, disagreeing with each other by wide margins, which is itself a sign that none of them rests on a shared, checkable measurement. None of those figures is repeated here; a number with no institution behind it does not become safer to cite because it is set beside three others in the same condition.

The only figure in this cluster carrying a named institution is Gartner's own stated expectation that more than 40 per cent of agentic AI projects will be cancelled by 2027, for reasons Gartner attributes to unclear business value, rising cost and thin risk controls. That is explicitly a forward-looking expectation about future cancellations. It is a different kind of claim from a measurement of past deployment failure, and treating the two as interchangeable would be its own construct-validity error, of exactly the kind this article is about.

This case looks, on first encounter, like this week's pattern: a number about AI that turns out to be less solid than its confident phrasing suggests. It is not. There is no real, disclosed study here being asked the wrong question. There is no study at all, only a personal blog post dressed in the kind of numerical precision that makes research look sturdier than it is. That is a different failure, and NTP already has a name for it: an existence check, not a construct-validity check, and Episode 31 of this series walked through exactly that discipline against four other circulating figures.

New Zealand's own version, and who is allowed to measure it

New Zealand's own version of this week's pattern needs no importing from overseas, and it sits inside a market many NZ businesses are dealing with directly this year: cyber insurance.

Marketing material from several competing SMB1001 and CyberCert certification providers states that certification may reduce cyber insurance premiums, with a commonly quoted range attached to the claim in wider circulation. Reading the providers' own pages directly tells a more careful story than the number that has detached from them. IT Live's own page does not state a percentage at all; it says the company introduces certified clients to a broker who has arranged a discount, and argues that the larger benefit is that many insurers now decline SMB cyber cover altogether without at least a baseline certification tier. Acronym IT's own page uses the hedged phrase "may reduce" premiums, with no number on the page itself. A third provider's marketing copy appears, by search result, to be where the specific percentage range originates, though that page returned an access error to a direct check this session and could not be confirmed first-hand.

The structural problem is this: every party quoting a number for the value of SMB1001 certification is also a seller of SMB1001 certification. None of them is an insurer. None has published how the figure was calculated. That does not make the figure false. It means the figure has never been measured by anyone independent of its own sale, which is a weaker evidential position than "reportedly" or "early data suggests" implies, and a different failure again from either of the two above: not a wrong question, and not a missing institution, but no independent measurer at all.

The direction underneath the marketing appears to be real and is worth keeping, even after the specific number is set aside: multi-factor authentication, endpoint detection and response, tested backups and a documented incident-response plan appear to be moving, across this market, from discount levers to conditions insurers will accept at all. A New Zealand business weighing certification has a genuine financial question to answer. It is "can we buy the cover we need without this," not the percentage a reseller's pitch deck currently carries, because nobody outside that reseller's own sale has measured the percentage.

When the instrument changes its mind

The fourth case does not involve a study answering the wrong question, a source that does not exist, or a circular sale. It involves an instrument that changed its own answer.

CVE-2026-69836 is a maximum-severity flaw in Microsoft's Entra ID cloud identity service, caused by deserialisation of untrusted data, disclosed 20 August 2026. Because Entra ID runs on Microsoft's own infrastructure, the fix was applied on Microsoft's backend, and no customer-side patch exists or was needed. Microsoft's security bulletin originally tagged the flaw's exploitation status as "yes, exploited." One day later, following a journalist's inquiry, Microsoft revised that status to "no," stating the flaw was not exploited in the wild, and offered a statement about releasing the advisory "for greater transparency" without explaining why the first published status differed from the second.

Nothing here suggests concealment. The correction is traceable to an outside question rather than hidden, and Microsoft's own statement is a transparency framing, not a denial that anything changed. The finding is narrower than that, and more useful: a severity assessment published once is not a fixed fact the moment it appears. Any board briefed on the strength of a vendor's first published statement about a flaw of this severity carried more risk, for one day, than the vendor's own second statement would have justified. A number that changes without explanation is not evidence that the number is wrong. It is evidence that the number was not yet stable enough to brief a board on, which a first announcement rarely admits to being.

What the classifier already knew, and what it doesn't check yet

V.E.R.A.'s NTP claim classification splits a claim by the kind of verification it needs. Existence claims, e-type in the framework's own vocabulary, assert that something is real, occurred, or was measured at a stated value, and they require a traceable source. Episode 31 ran that check against five circulating AI-agent-security statistics and found four of them failed it outright: no traceable source existed at all.

This week's five cases pass that first check cleanly. Suprmind exists. Gallup exists. CARMA exists. Microsoft's advisory exists. The certification resellers exist, and their pages say what this article quotes them as saying. An existence check, run correctly, would wave every one of them through. The failure this week lives one layer deeper, in whether the claim built on a real source actually matches what that source measured, which is a different verification act from checking that the source is there. V.E.R.A.'s E! Verification Service answers "does a named entity exist," against a corpus built from Wikipedia and Wikidata. It does not, and does not claim to, answer "does this survey measure what this sentence says it measures." That is a genuine gap in the framework as it stands today, not one hidden from view: nothing in this week's checking ran through V.E.R.A.'s NTP or E! architecture at all. Five ordinary sessions of fetching a page and reading what it actually said would have caught every finding in this article, using no tool more specialised than a browser and a willingness to read past the headline number.

The five cases nominated for this week's cluster do not, on inspection, share one mechanism. Two of them, Suprmind and the Gallup and CARMA pair, fit a single description: a real, disclosed measurement cited to answer a question it never asked. The other three each fail for a distinct and separately nameable reason. The Medium post fails because no real study exists behind it at all, which is last episode's problem wearing this episode's clothing. The insurance figure fails because nobody pricing it is independent of selling it. Microsoft's advisory fails because the instrument changed its own output without saying why. Collapsing four different failures into one label would make this series' own claim classification less precise, not more, and precision has been the argument this series has made since Episode 1.

For a reader deciding what to do with any of this, the working discipline is the same one this article has applied five times over: before repeating a number, find where it was actually measured, read what the source itself says it measured, and ask who benefits if it is believed uncritically. A study passing existence and disclosure does not make the sentence built on top of it true. Gallup and CARMA both cleared every check a careful reader could ask of a real, honest, well-documented instrument, and the sentence combining them was still wrong, because it asked both instruments a question neither one was built to answer.

The method behind every finding in this article is itself an open one. V.E.R.A.'s logic core ships under the GPL-3.0 licence on GitHub, and the discipline applied this week, read the source, compare its stated scope against the claim built on it, is reproducible by anyone with a browser and the patience to do it. None of it depends on a proprietary scoring model or a closed dataset. That reproducibility is the same argument the open-source movement has always made about code: a claim you cannot inspect is a claim you are asked to trust rather than verify. The same test applies to a cited statistic. A number attached to a locked methodology is asking for the trust a locked codebase asks for, and the honest answer, in both cases, is to open it before relying on it.

The same discipline, applied to a battlefield sensor rather than a survey, is not a hypothetical extension. NATO's own Test, Evaluation, Verification and Validation principles for military AI systems exist precisely because a sensor or targeting assessment can be real, technically sound and still be asked to answer an operational question outside the conditions it was tested under, which is this week's failure in a domain where the cost of missing it is measured in lives rather than headlines. Governance asymmetry becomes concrete here rather than abstract: a state whose framework requires disclosed test conditions before an AI system's output is trusted operationally is carrying a materially different risk profile from one that treats a vendor's first published assessment as final. The gap this article has spent several thousand words on, between a source existing and a source answering the question asked of it, does not shrink because the domain becomes lethal. It becomes the entire question.

Where has your own organisation quoted a real, credible source to answer a question that source never actually asked, and how would you have caught it before the number reached a board paper?

If your board or legal team needs a technical verification of an AI system's runtime logic or retrieval pipeline, message me and I will send the scope and the fixed fee.


The views expressed in this article are entirely my own, informed by more than 30 years of professional experience in architecture, security, and technology leadership in New Zealand. I write as director of Te Pono Limited; the views are personal and do not represent the position of any client, any government agency, or the New Zealand government. My commentary on legislation and policy is analytical, drawing on publicly available sources and my professional expertise in architecture, security, and AI governance, and it is politically neutral.


Andreas Hamberger is a New Zealand leader in Architecture & Security and Associate Member of the Institute of Directors. V.E.R.A. (Verified Existence & Reason Architecture) is an open-source logic engine available on GitHub. Through Te Pono he provides technical verification of AI runtime logic and retrieval pipelines for boards and legal teams; contact andreas@thehambergerreport.com for the scope and fixed fee.


This article was produced with AI assistance under my direction. Research, drafting and images pass through a pipeline I built and govern: automated gates for source verification, forbidden language and political neutrality, and my own review before anything is published. The tools include Claude, Gemini and Openart. The frameworks, arguments and editorial judgements are mine and are the same discipline I apply to the AI systems I audit for clients. AI accelerated the work; the thinking, and the responsibility for it, are mine.


[1] Suprmind. "The Confidence Trap: Multi-Model AI Divergence Index, April 2026 Edition." April 2026. https://suprmind.ai/hub/multi-model-ai-divergence-index/

[2] Gallup, in partnership with Bentley University. "Americans Cool Toward AI." 2026. https://news.gallup.com/poll/712751/americans-cool-toward.aspx

[3] CARMA. "Perceptions of AI" (second edition). 2026. No stable, directly hosted URL was located at time of writing; reported by outlets including reference 4.

[4] Branding in Asia. "AI Trust Rises as People Move Beyond the Hype." 2026. https://www.brandinginasia.com/ai-trust-rises-as-people-move-beyond-the-hype/

[5] Individual consultant, publishing on Medium under the handle "@snehal_singh." Personal analysis of agentic AI deployment outcomes, circulated as an unattributed failure-rate statistic; the specific figures are not used in this article, as the analysis rests on no institutional authority (see body text). 20 February 2026. https://medium.com/@snehal_singh/i-analyzed-847-ai-agent-deployments-in-2026-76-failed-heres-why-0b69d962ec8b

[6] dig.watch. "Gartner warns that more than 40 percent of agentic AI projects could be cancelled by 2027." 2026. https://dig.watch/updates/gartner-warns-that-more-than-40-percent-of-agentic-ai-projects-could-be-cancelled-by-2027

[7] IT Live. "SMB1001 Gold NZ." 2026. https://itlive.co.nz/smb1001-gold-nz/

[8] Acronym IT. "Cyber Security Certification, SMB1001." 2026. https://acronym.co.nz/security/cyber-security-certification-smb-1001/

[9] Help Net Security. "Microsoft Entra ID Vulnerability (CVE-2026-69836)." 21 August 2026. https://www.helpnetsecurity.com/2026/08/21/microsoft-entra-id-vulnerability-cve-2026-69836/

[10] The Hacker News. "Microsoft Entra ID Flaw, CVSS 10.0." 21 August 2026. https://thehackernews.com/2026/08/microsoft-entra-id-flaw-cvss-100.html

Next
Next

Smarter, and Wronger: The Deeper-Reasoning Paradox and the Case for Proof