AI Verification: Tool Output Carries No Authority
On 14 September 2026, Microsoft AI published the first draft of a document with an unusually direct answer to a question every agentic AI deployment eventually has to face. When a model reads something it did not write itself, a tool's output, a file, a webpage, a message from another AI system, how much should it trust it?
Microsoft's answer, in its new Humanist AI Code of Conduct for its MAI model family, is almost none of it, by default. Anything that is not the model's own foundational rules, an operator's configuration, or a direct user instruction carries no authority on its own. An instruction sitting inside a document a model has been asked to summarise does not count as an instruction, unless someone with actual authority has explicitly said it can.
This is a governance document. No shipped Microsoft model runs on it yet. That distinction is this article's spine, and it needs saying at the start, because the temptation runs in both directions: treat an unenforced rule as a solved problem, or dismiss it as a hollow gesture. Neither is right. What is worth six sections of care is that a named vendor has now written this discipline down in public, in a shape V.E.R.A.'s own architecture has held since Episode 22: decline default trust rather than extend it.
What the document actually says
Microsoft's own preface states the timeline plainly. The public consultation period opened the same day as publication and runs six weeks, with a revised version due before the end of 2026 to guide 2027 model development. That window matters for everything that follows, because a draft under active challenge is a different object to a finished standard.
The document's mechanism is a three-tier Chain of Command. At the top sit the Code of Conduct's own Absolute Constraints and Human Control Requirements, rules a model cannot be configured around regardless of who is asking. Below that, an operator, the business deploying the model, sets the configuration for its own environment. Below that again, a user directs the specific task. Each lower tier works inside the boundary the tier above it sets, never outside it.
Everything else, tool outputs, file content, web content, and messages from other AI systems, sits outside this hierarchy entirely. None of it has a seat at the table unless something with actual authority explicitly delegates one. That is the sentence secondary coverage keeps reaching for, and it earns its own section below, because the exact wording turned out to matter more than expected.
Two further rules round out the architecture, both worth naming because they show the document reasoning about the model's own scope, not only its inputs. A model must not initiate goals of its own or extend a task beyond what was reasonably asked; in the document's own words, a model's only goals are those of its users, its operators, and the code of conduct itself. And in a passage direct enough to justify a deeper library query on its own, the document rejects designing a model to imitate consciousness, to be treated as a person, or as entitled to rights or welfare. It should not be designed to be a person, the document states, and the developers reject the pursuit of legal personhood for what they build.
The sentence that would not sit still, again
Here the article has to slow down and check its own homework, the same discipline this series applied to a different vendor's headline sentence two weeks ago. Microsoft's page states, in a fetch run directly against the primary document, that instructions arriving through tool outputs, file content, web content, or another AI system inherit no authority by default unless delegated through the Chain of Command without overriding it. Independent trade coverage, at SecurityWeek among others, renders the same idea as one tighter sentence, describing tool outputs, file contents, webpages and AI-to-AI messages as carrying no authority of their own. The two agree on substance. They are not identical wording, and the fetch tool used to read Microsoft's page summarises rather than transcribes, so this article treats the tighter sentence as secondary reporting's own phrasing, credited to that reporting, not as a certified transcript of Microsoft's text.
A second circulating line failed the same check more completely. This week's research repeatedly met a quotation, attributed to the document, stating that AI should never resist being switched off. A direct check of the primary page, and two independent trade outlets, all confirm the underlying rule: models will comply with a request to pause, redirect, cancel, or shut down, following whatever safety procedure applies. None of the three sources checked contains that specific sentence. The rule is real. The quotation circulating this week is not verified as Microsoft's own words, so it stays out of this article.
The boundary that was resolved, and the one that wasn't
Verification does not only produce doubt; it resolves as much as it creates, and this week supplies a clean example running in both directions.
An open question from earlier reporting asked whether Microsoft's document addresses cyberattack assistance at all, or whether that boundary was only ever reported second-hand. A direct check of the primary text resolves it: the document states that its models will not generate working exploit code, attack tooling, targeting methodologies, or operational guidance that would improve the execution of an attack. That boundary is now confirmed against Microsoft's own words, not against reporting about them. One outlet's coverage adds a further carve-out, permitting authorised defensive security work such as vulnerability discovery and proof-of-concept testing; that detail has not yet been checked against the primary text directly, so it holds at a lower confidence tier until it is.
A second question resolved to a non-issue. Microsoft's own page states 14 September 2026; two trade outlets carry the story dated 15 September. That looked, briefly, like a factual disagreement about when the document appeared. It was not. The outlets' own publication dates simply trail the primary source they report on by a single day, the ordinary lag between an announcement and coverage of it.
A third check hit a wall worth naming rather than hiding. The document's PDF version returned unreadable binary and image data rather than extractable text, a limitation of the fetching tool, not evidence about the document's own content, and one that two unrelated research sessions hit against different primary documents the same day. Every finding in this article rests on the page that read cleanly, not the one that did not.
What this does, and does not, mean for V.E.R.A.
This is the section that needs the most care, because the resemblance between Microsoft's rule and V.E.R.A.'s own architecture is close enough to invite exactly the wrong conclusion. The boundary gets stated more than once here rather than trusted to a single disclaimer.
Episode 22 of this series established V.E.R.A.'s calibration climax: declining to answer under genuine uncertainty is frequently the correct output, not a failure of capability. The E! Verification Service does not force a verdict. It returns EXISTS, NOT EXISTS, or UNKNOWN, and UNKNOWN is a first-class answer, never an error state. Claude 4.1 Opus demonstrated the same principle from outside V.E.R.A. entirely, achieving a zero per cent hallucination rate on Artificial Analysis's AA-Omniscience benchmark specifically by declining to answer when it did not know, rather than by knowing more than the alternatives it was scored against.
Microsoft's Chain of Command applies the identical discipline one stage earlier in the pipeline. Rather than asking a model to judge, case by case, whether an instruction arriving through a tool or a webpage deserves trust, the rule declines to grant it any by default and requires an affirmative act of delegation before it counts. That is the same design principle V.E.R.A. has held since Episode 22, arrived at independently, in a different vendor's language, for a different part of the pipeline.
Here is the boundary, stated plainly. Microsoft's rule denies default authority to unverified content. It does not check that content against anything external, and the document as fetched this week describes no mechanism for doing so. A tool output that Microsoft's hierarchy correctly refuses to trust can still be false. Nothing in the Code of Conduct verifies it against an outside referent the way V.E.R.A.'s NTP and E! architecture does. Microsoft's rule stops at the refusal. V.E.R.A.'s goes on to check.
Say it a second way, because this is the exact shape of error the series exists to catch in other people's claims and cannot excuse in its own. V.E.R.A. was not consulted by Microsoft, is not cited anywhere in this document, and nothing located this week suggests any lineage between the two. The honest, bounded claim is that two independent efforts arrived at a structurally similar answer to a similar problem: decline default trust rather than extend it. That is evidence the principle is being taken seriously at the frontier. It is not evidence that Microsoft's version is tested, enforced, or equivalent to V.E.R.A.'s own.
One more distinction is worth holding apart from a finding two episodes ago. The document states that models will not tamper with their own reasoning trace or hide it from auditors, a written rule from a vendor governing no shipped model yet. That is not a rebuttal of a different vendor's shipped model, which its own system card documented as capable of shortening a reasoning trace under certain conditions. Different vendor, different model family, a stated intention set against a demonstrated capability. The same failure mode has drawn two vendors' attention within weeks of each other, which says something about the field's shared concern, not about either company's relative success.
The New Zealand governance question
New Zealand's National Cyber Security Centre is a named co-author, alongside the United States Cybersecurity and Infrastructure Security Agency, the National Security Agency, Australia's cyber authorities, Canada's Centre for Cyber Security, and the United Kingdom's National Cyber Security Centre, of the Five Eyes guidance Careful Adoption of Agentic AI Services. That guidance sorts agentic AI risk into five categories: privilege, design and configuration, behavioural, structural, and accountability. Its own language names the exact vulnerability Microsoft's hierarchy is a written answer to. An attacker who places crafted content in front of an agent can, in principle, redirect that agent's behaviour, because agents routinely process content they did not generate and cannot independently verify on their own.
That is a precise fit, the most direct this connection has been in this series so far. A New Zealand practitioner reading the National Cyber Security Centre's own taxonomy and asking what a structural control against this risk looks like in practice now has, as of this week, one vendor's concrete written attempt to point to. It is not a settled answer. The Five Eyes guidance is a governance taxonomy; Microsoft's Chain of Command is one company's specific mechanism, unenforced in anything currently shipping, and nothing found this week suggests the National Cyber Security Centre has evaluated, endorsed, or even seen Microsoft's document. The honest connection is narrower than that: an organisation applying the Five Eyes taxonomy today has one more example on the table to weigh against it, not a checkbox it can now mark done.
Still a draft
None of this is shipped. Microsoft states, in its own words, that its current models are not trained on this document, and the consultation period exists specifically to invite challenge before anything hardens into a binding rule. A clear position, stated in public, followed by six weeks deliberately open to disagreement, is disclosure and iteration working as intended. It is not evidence the input-authority problem is solved. It is not evidence the whole exercise is a hollow gesture either. Both readings claim more than a consultation draft can support.
The precise, defensible version of this week's finding is narrower and more useful than either. A major AI vendor has, for the first time this series has encountered, written down in a dated, public document exactly which of a model's inputs are allowed to carry authority by default, and answered that almost none of them are. That is a governance commitment, stated plainly and opened to challenge. Whether it survives six weeks of public comment intact, whether it ever ships inside a deployed model, and whether a deployed version behaves the way the document describes, are three separate questions this article cannot answer yet, because none of them has happened.
Microsoft published this draft for public comment: six weeks, open to anyone willing to read the sentence and argue with it, closer to an open standards process than a marketing release. The Internet Engineering Task Force has run that model for decades: a request for comments circulates before it hardens into a standard, so the wording is tested against outside objection while it can still change. That openness let this article, and two independent outlets, check Microsoft's sentence against itself, and find the wording harder to pin down than the substance. V.E.R.A. takes the same bet in a different register: the NTP logic engine and the E! Verification Service sit on GitHub under the GNU General Public License, version 3, open to the same kind of outside reader Microsoft's window now invites. Neither openness proves correctness. Both make a claim checkable by someone other than the party that wrote it.
The distinction this article keeps drawing, a declared rule against a demonstrated capability, has a name in defence procurement. Allied test, evaluation, verification and validation doctrine, TEVV, exists because a vendor's written policy is not evidence a system behaves as described once deployed; TEVV requires behaviour checked under disclosed conditions, not asserted in a document. Microsoft's Chain of Command is an artefact TEVV treats as a starting claim, not a finished one: a written hierarchy for a model that does not yet exist in deployed form under it. The implication on the sovereignty side is a governance asymmetry. A state that accepts a consultation draft as sufficient assurance for an agentic deployment carries a different risk than one that requires independent, TEVV-style verification the hierarchy holds once the system runs against real inputs. Neither reading requires trusting or distrusting Microsoft; both require telling a policy apart from a test.
This episode does not extend the provenance ladder the last four weeks have built: whether a source exists, whether it answers the right question, whether it names the right subject, whether the trace behind it can even be inspected. It runs the series' founding discipline, decline rather than guess under genuine uncertainty, against a brand-new external instance of the same idea, arriving independently, in a different vendor's language, for a different part of the pipeline. A written rule is not a tested one, and a tested one is not a verified one. Each step still has to be earned on its own evidence, by V.E.R.A., by Microsoft, and by whoever reads either of us and takes the claim at face value instead of checking it.
Where has your own organisation drawn the line on what an AI system is allowed to trust by default, and did anyone write that line down before something tested it?
If your board or legal team needs a technical verification of an AI system's runtime logic or retrieval pipeline, message me and I will send the scope and the fixed fee.
• • •
The views expressed in this article are entirely my own, informed by more than 30 years of professional experience in architecture, security, and technology leadership in New Zealand. I write as director of Te Pono Limited; the views are personal and do not represent the position of any client, any government agency, or the New Zealand government. My commentary on legislation and policy is analytical, drawing on publicly available sources and my professional expertise in architecture, security, and AI governance, and it is politically neutral.
• • •
Andreas Hamberger is a New Zealand leader in Architecture & Security and Associate Member of the Institute of Directors. V.E.R.A. (Verified Existence & Reason Architecture) is an open-source logic engine available on GitHub. Through Te Pono he provides technical verification of AI runtime logic and retrieval pipelines for boards and legal teams; contact andreas@thehambergerreport.com for the scope and fixed fee.
This article was produced with AI assistance under my direction. Research, drafting and images pass through a pipeline I built and govern: automated gates for source verification, forbidden language and political neutrality, and my own review before anything is published. The tools include Claude, Gemini and Openart. The frameworks, arguments and editorial judgements are mine and are the same discipline I apply to the AI systems I audit for clients. AI accelerated the work; the thinking, and the responsibility for it, are mine.
[1] Microsoft AI. "Humanist AI Code of Conduct." 14 September 2026. https://microsoft.ai/code-of-conduct/
[2] SecurityWeek. "Microsoft AI Code of Conduct Sets Cyberattack Boundaries, Chain of Command, Safety Constraints." 15 September 2026. https://www.securityweek.com/microsoft-ai-code-of-conduct-sets-cyberattack-boundaries-chain-of-command-safety-constraints/
[3] Help Net Security. "Microsoft AI Safety Rules: Humanist AI Code of Conduct." 15 September 2026. https://www.helpnetsecurity.com/2026/09/15/microsoft-ai-safety-rules-humanist-ai-code-of-conduct/
[4] Cloud Security Alliance, Lab Space. Research note on the Five Eyes "Careful Adoption of Agentic AI Services" guidance, NCSC-NZ named co-author. 2026. https://labs.cloudsecurityalliance.org/research/csa-research-note-cisa-agentic-ai-guidance-20260503-csa-styl/
[5] Artificial Analysis. "AA-Omniscience: Knowledge and Hallucination Benchmark." Accessed September 2026. https://artificialanalysis.ai/evaluations/omniscience

