AI Verification: The Study Is Real. It Is Measuring Something Else.
When a Real Study Still Answers the Wrong Question
Smarter, and Wronger: The Deeper-Reasoning Paradox and the Case for Proof
One model proved a 27-year-old theorem. Reasoning models got worse at facts, same week.
Smarter, and Wronger: The Deeper-Reasoning Paradox and the Case for Proof
One model proved a 27-year-old theorem. Reasoning models got worse at facts, same week.
The Model Agrees With You When You're Wrong
Sycophancy, and why verification must be external to the model, and to the user
The Paradox of Instruction: Why Safe AI Needs Architectural Gates, Not Better Prompts
What a 2026 ablation reveals: prompting and structural gates fail in different places
Look It Up: Why the Model That Searched Beat the Model That Remembered
A hallucination benchmark just made the difference measurable: a system that reads a document and one that searches it are not two grades of the same thing.
Verification Becomes a Scaling Axis: The Week the Judge Got a Probability Distribution
Why one grader scored 70.8% and another 87.4% on the same task, and what it cannot fix
Scoring the Reasoning: When the Verifier Becomes Probabilistic
A score, a proof, or a verdict: what should an external AI verifier actually output?
Governing the Path, Not Just the Answer: Runtime Verification for Agents
Prompting shifts what an agent is likely to do; only the execution path shows what it did
Thinking to Recall: When One Made-Up Fact Poisons the Whole Answer
One false step, one wrong answer: why verification must sit outside the model
Broken by Default: Why the Scanner Can't See What the Proof Can
Te Pono V.E.R.A. Episode 23: Six industry scanners read the same AI-generated code a formal solver did. The solver proved most of it broken. The scanners saw almost nothing.
The Code That Finds the Lie
Te Pono V.E.R.A. Episode 21: The claim extraction module is running. Mean Ffact score is 0.35. That is not a boast. It is a measurement.
The Government Wants to Read the Model First
What a US Executive Order, an Anthropic S-1, and an orbital jurisdiction gap tell us about the architecture of AI trust
V.E.R.A. Saturday, Episode 19: The Reasoning Trap
How Reinforcement Learning Creates the Failure Mode V.E.R.A. Is Built to Catch
The External Referent Arrived: What the Curl Test Tells Us About Vendor Verification
V.E.R.A. Saturday | Episode 17 | Coding Arc 6 / Close
Language Is Not Capability: Why AI Governance Documents Cannot Audit Themselves
Episode 14 of the V.E.R.A. Coding Arc: The Vocabulary-Substance Gap
V.E.R.A. Episode 13: The Mechanism Behind the Lie
Why AI Systems Cannot Distinguish Confidence from Accuracy
We Shipped: V.E.R.A. v0.1 Is in the World
Episode 11: Ten episodes of architecture. Here is what we built.

