The Hamberger Report Weekly #11: Same System, Different Number
VerifiedIntelligence
Same System, Different Number
Three numbers changed this week depending only on who measured them.
My new book, The Hallucination Economy, is out today. It is about how organisations act on AI claims no one can check, and what boards should ask instead: five parts, nineteen chapters, closing on the questions a board should put before it relies on any AI claim. The Kindle edition is available now; the paperback and hardcover follow once Amazon's review clears.
Seven pieces this week share one mechanic: the same fact produces two honest, differently labelled answers depending on who checked it, and the real claim lives in that gap. A model scores 62.7 per cent or 99.9 per cent depending on which test harness ran it. A reported 40 per cent price cut is really two cuts, 20 and 60 per cent, depending on which rate is isolated. A government portfolio has every framework it should, and an independent reviewer finds those frameworks inconsistently applied. A satellite constellation is filed with a regulator; none of it has flown. An interface sat usable for years before a court was asked whether it could be owned. On this week's evidence, across benchmarking, pricing, governance and ownership, the apparatus measuring a claim carries as much weight as the claim.
This week's Project V.E.R.A. analysis, on the same AI model scoring 37 points apart under two different test conditions, is worth ten minutes before you next compare two vendors' benchmark claims.Read the full analysis
Vendor Assurance: Contracted, Not Confirmed
Health New Zealand's new compliance notice requires an independent, demonstrated assessment of a vendor's security before patient data moves, not a vendor-completed questionnaire.
Orbital Compute: Filed, Not Flown
Two companies have filed with the United States government for more satellites than currently exist in Earth orbit combined, and neither has published a measured watt of on-orbit compute: a filing proves a claim, not a system.
The Cache Got Cheaper
Anthropic's "40 per cent cheaper" claim isn't on its own rate card: the real cuts are 20 per cent on tokens and 60 per cent on cached memory, a bet on which AI workloads it wants more of.
Cost Engineering, Not Vendor Loyalty
AT&T now routes 40 per cent of employee AI requests to open-weight models, up from 20 per cent in May, because it built a live cost measurement instead of trusting a year-long contract.
The Control Existed. The Evidence Didn't.
MBIE's own stocktake and the Privacy Commissioner's notices to Health NZ, published the same week, both found governance frameworks in place with no independent evidence anyone checked whether they were actually followed under pressure.
Nobody Decided Who Owns It
An eleven-year United States copyright fight over 11,500 lines of code reached the Supreme Court and still left open whether an interface can be owned: the same question now sits under every AI model's training data.
Two Harnesses, One Model
The same AI model scored 62.7 per cent and 99.9 per cent on the same benchmark, the same day, depending only on which test harness ran it.
Which of your organisation's own numbers, a cost saving, a vendor's benchmark claim, would survive being measured a second time, by someone with no reason to prefer the answer you have?

