Forty-Seven Seconds of National Darkness

Pre-Flight Metrics Card


Episode 6

I was tightening access lists on a core border router at Telecom XTRA when I typed the line that turned off New Zealand's internet.

access-list 101 deny ip any any

It was 2001. I was Security and General Systems Team Leader for the country's largest ISP, responsible for XTRA's entire server base and network infrastructure. The task was routine: hardening the access control list on a border router that carried traffic for 300,000 customers. The kind of work I had done dozens of times before. Except this time, I put the deny rule in the wrong position.

On a Cisco router, access lists are processed sequentially. The router reads each rule from top to bottom, and the first match wins. A deny ip any any rule is supposed to sit at the bottom of the list, as the final catch-all after all the permit statements have done their work. Put it at the top, or even in the wrong position partway through the list, and it matches everything. Every packet. Every protocol. Every source and destination. The router does exactly what you told it to do. It denies all traffic.

I looked at my watch. The seconds stretched. One minute. Then forty-seven seconds more. One minute and forty-seven seconds of total darkness for an entire country's internet, though in the retelling over the years I have compressed it to forty-seven seconds, because that is how memory works with moments of pure adrenaline.

Then the internet came back.

Andrew and the Power Button

The reason New Zealand's internet outage lasted less than two minutes instead of becoming a multi-hour national incident was not sophisticated failover technology. It was not automated recovery. It was not a redundant architecture designed for graceful degradation.

It was Andrew.

Andrew was a core router network specialist for Telecom. Before I started the access-list change, I had walked over to his desk and given him a simple instruction: if connectivity drops, power-cycle the router. Do not wait for me to diagnose the problem. Do not call anyone. Just hit the power button.

This was not standard procedure. There was no runbook for it. It was the kind of pre-emptive human judgement call that comes from working with infrastructure long enough to know that the most likely failure mode in a configuration change is the change itself. If the router lost connectivity, the only person who could fix it via the console was me, and I would be locked out by my own access list. The fastest recovery path was a cold reboot, which would reload the previous saved configuration.

Andrew saw the monitoring screens go dark. He counted to ten, confirming it was not a transient blip. Then he hit the power button. The router rebooted. The saved configuration loaded. Traffic resumed.

One minute and forty-seven seconds. That was the cost of one misplaced line in an access list on the core router for New Zealand's dominant ISP.

The Silence That Followed

Here is the part of the story that still surprises me twenty-five years later.

Nothing happened.

I walked to the Customer Service Centre after the router came back online. I expected a queue of calls, a stack of complaints, an inquest. Instead, the CSC team said they had received no calls about any outage. Not one. In 2001, if your internet dropped for less than two minutes, most people assumed their dial-up connection had simply done what dial-up connections did: disconnected. The few broadband customers on Telecom's new JetStream ADSL service would have seen a brief interruption and assumed it was normal. There was no real-time Twitter firestorm. No Downdetector graph going vertical. No push notifications alerting a nation that their connectivity had just been erased and restored by one line of Cisco IOS configuration and one person's finger on a power button.

The incident vanished into the daily noise of early 2000s internet reliability. But it did not vanish from my thinking.

No Test Environment, No Changes

The reason the outage happened was not incompetence. It was architecture. Or rather, the absence of it.

Telecom XTRA in 2001 had no test router environment. There was no lab where I could apply an access-list change to an identical router, verify that traffic still flowed, and then replicate the change in production. The only way to test a configuration change on the core border router was to apply it to the core border router. In production. Carrying traffic for 300,000 customers and, by extension, a significant proportion of the country's internet connectivity.

This was normal for the era. Test environments cost money. Identical hardware for a lab meant purchasing routers that would never carry a single customer packet. In the early 2000s, with ISPs running on thin margins and infrastructure budgets under constant pressure, the business case for a test lab was difficult to make.

After the outage, I made the case anyway. I requested a dedicated test router environment, and a few months later, I got one. Between the outage and the arrival of that test lab, I made no configuration changes to the production routers. None. The risk calculus had changed permanently. If there was no safe way to test, there was no safe way to deploy.

The Linux Pockets

XTRA was, on the surface, a Solaris and Windows shop. The core infrastructure, the mail servers, the web platform, the business applications: all ran on Sun Microsystems hardware or Microsoft systems. Linux was not part of the official architecture.

But Linux was there. Small pockets, running quietly on commodity hardware, doing the work that nobody thought to ask the enterprise platforms to do. Monitoring servers. The systems that watched the systems. When the access-list took down the border router and Andrew's power-cycle brought it back, the first confirmation I had that connectivity had been restored came from a Linux monitoring box. It had noticed the outage. It noticed the recovery. It logged both events with timestamps that let me reconstruct exactly what had happened and how long it had lasted.

This pattern, Linux running the oversight layer while commercial platforms ran the visible services, was not unique to XTRA. Across New Zealand's ISP landscape in the early 2000s, Linux had found its niche not as the primary platform but as the reliable, low-cost infrastructure that made everything else observable. It was the monitoring. The logging. The network management tools that ran on a repurposed desktop under someone's desk. The commercial platforms got the procurement budgets and the vendor support contracts. Linux got the jobs that actually required reliability.

The irony was not lost on me. The most expensive systems in the building went dark because of a single misconfigured access list. The cheapest system in the building, a Linux box running open-source monitoring software, told me exactly when they went dark and exactly when they came back.

Telecom XTRA in 2001: The Scale of What Could Break

To understand why a single router misconfiguration could take out a country's internet, you need to understand Telecom XTRA's position in 2001.

Telecom New Zealand was the country's dominant telecommunications company, worth roughly a third of the entire New Zealand Stock Exchange. XTRA, its ISP subsidiary, had signed up its 300,000th customer in 2000. Microsoft had invested more than NZ$200 million in Telecom. The Southern Cross Cable, connecting New Zealand directly to the United States via fibre optic, had just been switched on. JetStream ADSL was rolling out. The XtraMSN portal partnership was about to launch.

XTRA was not a small ISP. It was the gateway through which a significant proportion of New Zealand accessed the internet. And in 2001, that gateway's resilience against configuration errors on its core routing infrastructure consisted of one network specialist named Andrew, pre-briefed to hit a power button if the screens went dark.

There was no automated rollback. No canary deployment. No staged rollout of configuration changes across redundant paths. The same pattern that made the outage possible, a single point of configuration failure with no safe testing mechanism, existed at ISPs and enterprises around the world. We just happened to be the ones who proved it that day.

The Modern Bridge: When AI Agents Have No Test Lab

Twenty-five years after I typed access-list 101 deny ip any any in the wrong place, the same architectural failure is playing out at vastly larger scale.

In 2026, autonomous AI agents are being deployed into enterprise production environments with the same structural weakness that made my XTRA outage inevitable: no safe way to test before deployment, no reliable rollback when things go wrong, and insufficient human oversight at the moments that matter most.

The numbers tell the story clearly. According to Gartner, over 40% of agentic AI projects are expected to fail or be cancelled by 2027. A Deloitte study found that while 38% of organisations are piloting agentic AI solutions, only 11% have systems actually running in production. The failure pattern is consistent: agents deployed without defined failure modes, without graceful degradation paths, and without human-in-the-loop safeguards at critical decision points. When something goes wrong, the system keeps going, compounding the error with each autonomous step.

This is the access-list problem at machine speed. In 2001, my misconfigured rule affected traffic for less than two minutes because Andrew was standing by with a pre-agreed recovery plan. In 2026, an autonomous AI agent executing a multi-step workflow can cascade a minor error into a systemic failure before any human reviewer has time to intervene.

The parallel to New Zealand's own infrastructure is immediate. In February 2026, the MediMap health platform breach demonstrated what happens when a critical system has no integrity controls. An attacker gained unauthorised access and modified patient records across facilities serving roughly 60% of New Zealand's aged care sector, marking living patients as deceased and altering prescriber information. The system had no mechanism to detect that bulk demographic modifications were anomalous. Nurses reverted to paper-based medication rounds. The platform has been offline for over a week at the time of writing. The pattern is the same one I lived through in 2001: a system designed without adequate safeguards against the most predictable failure mode, one set of credentials doing something unexpected.

The lesson from that Cisco router has not changed. If you cannot test it safely, you should not deploy it. If you have no rollback plan, you have no plan. And if there is no human standing by with a finger on the power button, with pre-agreed authority to act when the screens go dark, then you are trusting your infrastructure to luck.

Linux, the operating system that quietly monitored XTRA's infrastructure from those small pockets of commodity hardware in 2001, now runs the infrastructure layer beneath virtually every AI system in production. The GPU clusters training frontier models run Linux. The Kubernetes clusters orchestrating agentic workflows run Linux. The monitoring and observability platforms that watch those agents for anomalous behaviour run on the same open-source foundations that told me my border router had crashed and recovered twenty-five years ago. Linux kernel 7.0, which entered its release candidate phase in February 2026 with Rust support now officially stable, carries forward the same role it played at XTRA: the reliable layer that makes everything else observable, recoverable, and governable.

Andrew's finger on the power button was not automation. It was governance. The simplest, most effective form of human-in-the-loop oversight: a competent person, briefed on the risk, empowered to act, watching the system at the moment of highest vulnerability. Twenty-five years later, that is still the governance model most AI deployments lack.

Your Outage Story

Have you ever caused, or fixed, a major outage?

The stories that teach us the most about infrastructure resilience are rarely the ones in the incident reports. They are the ones told at conference after-parties, in server rooms at 3am, in the knowing silence between two engineers who both understand what it feels like to watch a monitoring dashboard go completely dark and know that it was your command that did it.

What was your "access-list 101" moment? What did it teach you about testing, about human judgement, about the gap between what you intended and what the system actually did? Share your story. Someone reading this is about to make their first production change without a test environment, and they need to hear it.


Next in this series: Episode 7, "The Auckland Power Crisis: When Nature Humbles Technology"



The views expressed in this article are entirely my own, informed by more than 30 years of professional experience in architecture, security, and technology leadership in New Zealand. They do not represent the views of my employer, any government agency, or the New Zealand government. My commentary on legislation and policy is analytical, drawing on publicly available sources and my professional expertise in architecture, security, and AI governance. I follow the Public Service Commissioner's Code of Conduct for the Public Sector and social media guidance.


Andreas Hamberger is a New Zealand leader in Architecture & Security and Associate Member of the Institute of Directors. A Concise History of Linux chronicles the operating system that changed the world, and the lessons it holds for the AI era. He founded Yoper Linux, served as a technology specialist for Novell during the Linux Wars, and is the author of "Generative AI: Skynet or Heaven" and "Space Mafia." He can be reached at linux@linux.co.nz.


I use AI tools, including Sudowrite, Claude, Perplexity AI, DeepSeek AI, ChatGPT, Grok, Copilot, Openart and Gemini, as deliberate production tools, not ghostwriters. This is consistent with my position: AI amplifies human judgement; it does not replace it. The frameworks, arguments, and editorial decisions in this series are original work. AI accelerated the process. The thinking is mine.


References

  1. Direct testimony, Andreas Hamberger, February 2026. Access-list configuration error, Telecom XTRA core border router, 2001.
  2. Direct testimony, Andreas Hamberger, February 2026. Andrew (core router network specialist, Telecom) pre-briefed for power-cycle recovery.
  3. Direct testimony, Andreas Hamberger, February 2026. Customer Service Centre received no calls following the outage.
  4. Direct testimony, Andreas Hamberger, February 2026. No test router environment at XTRA; dedicated test lab obtained months after the incident.
  5. Direct testimony, Andreas Hamberger, February 2026. Linux monitoring servers running on commodity hardware at XTRA.
  6. Wikipedia, "Spark New Zealand." XTRA 300,000 customers by 2000; Southern Cross Cable; JetStream ADSL. https://en.wikipedia.org/wiki/Spark_New_Zealand
  7. FundingUniverse, "Telecom Corporation of New Zealand Limited History." Microsoft NZ$200M+ investment; ~1/3 of NZ Stock Exchange. https://www.fundinguniverse.com/company-histories/telecom-corporation-of-new-zealand-limited-history/
  8. Microsoft press release, 10 May 2001. XtraMSN partnership announcement. https://news.microsoft.com/2001/05/10/
  9. Cisco, "Configuring IP Access Lists." Sequential processing, first-match-wins ACL behaviour. https://www.cisco.com/c/en/us/support/docs/security/ios-firewall/23602-confaccesslists.html
  10. Gartner, mid-2025 press release. Over 40% of agentic AI projects expected to fail or be cancelled by 2027.
  11. Deloitte, "Emerging Technology Trends 2025." 38% of organisations piloting agentic AI; only 11% in production. https://www.deloitte.com/us/en/insights/topics/technology-management/tech-trends/2026/agentic-ai-strategy.html
  12. RNZ, Stuff, NZ Herald, multiple sources, February 2026. MediMap breach: unauthorised access, patient records modified across ~60% of NZ aged care facilities. https://www.rnz.co.nz/news/alert-nat/587773/patient-data-changed-as-major-nz-health-app-medimap-hacked
  13. Phoronix, 22 February 2026. Linux 7.0-rc1 released; Rust support now officially stable. https://www.phoronix.com/news/Linux-7.0-rc1-Released
Previous
Previous

The Auckland Power Crisis: When Nature Humbles Technology

Next
Next

Linux.co.nz and a Number Plate: Betting on the Future