AI Cost Engineering: The Forty Per Cent AT&T Now Routes to Open Source

Return to Part 0: Table of ContentsPrevious Article: Episode 32, Linux 7.3 and the EU Cyber Resilience Act's Manufacturer/Steward Split


In 2009 I joined 2 Degrees, New Zealand's third mobile network, as a Senior Design Engineer. The title never named the actual job. Find where a new telco's infrastructure spend, the operating systems, the middleware, the F5 load balancers, could be cut without breaking anything a customer would notice. Every one of those lines was a contract somebody had negotiated a year in advance. The whole company lived with those assumptions for twelve months, whether or not they still held. Cost was decided a year before it mattered, the day the contract was signed.

Seventeen years later, a much bigger telco has built that job into its infrastructure instead of paying someone to do it once a year. AT&T's internal AI gateway now decides which model answers every request an employee sends. It chooses between frontier systems from Anthropic or OpenAI and open-weight models the community builds and gives away. Forty per cent of the time, right now, it chooses open weight, up from twenty per cent in May.[1] AT&T has said publicly it wants that share at sixty to seventy per cent within a few years.[1] Cost decided the number, not conviction.

Open weight means the underlying model parameters are published for anyone to download, inspect and run themselves, the closest AI equivalent to source code. A frontier model's weights stay locked inside the vendor that trained them.

How the gateway decides

The mechanism is called an AI gateway. It runs on LiteLLM, an open-source routing layer any organisation can install for itself. Instead of employees calling one vendor's API and living with whatever that vendor charges, every request passes through a layer AT&T controls. That layer can send a coding query to Nvidia's Nemotron, a general query to Meta's Llama, or a lighter task to Google's Gemma, weighing cost, speed and expected quality at each turn. The routing is cache-aware, so a question close enough to one already answered can skip the expensive model altogether. The gateway can also switch models mid-conversation, so a task that starts simple and gets harder is not stuck with whichever model answered the first message.

AT&T's own executives have been specific about why the number is what it is: cost, not conviction. Mark Austin, the vice president who oversees AI for the company's roughly one hundred thousand employees, has put a figure on the trade-off directly. Routing coding and other advanced tasks to open-weight models cut AT&T's costs on that work by as much as fifty-six per cent, for a stated quality decline of about two per cent.[1] Fifty-six per cent cheaper. Two per cent worse. AT&T said both numbers, not just the first one. A separate and much larger figure, savings of as much as eighty per cent, has also been attributed to AT&T, this time to Andy Markus, the company's Chief Data and AI Officer, in a different interview weeks later.[2][3] Austin's fifty-six per cent names its scope and its trade-off. Markus's eighty per cent names neither.

The order of magnitude, and what is actually confirmed

What makes this a Linux Wednesday story rather than a straightforward cost-cutting one is the shape of the decision, not its size. AT&T did not run a tender for a cheaper AI vendor. It built the choice into its own infrastructure as a runtime property. That is how the kernel treats which scheduler runs, or which filesystem serves a request: a decision made at the moment it is needed, not fixed a year ahead. The community did not persuade AT&T that open weights were better in principle. An engineer measured what community-maintained code was worth against the alternative, and the answer routed itself.

This series has spent the last three editions asking who governs the shared registry and who can simply buy it outright. It has also asked who answers for a line of text once a human name is no longer reliably attached to it. This edition asks a blunter question none of the previous three needed to: what is any of it actually worth, in the currency a chief financial officer reads. AT&T just answered, in numbers, for itself.

The same month, the kernel had its own reckoning

A few thousand kilometres of fibre away from AT&T's gateway, the code that gateway increasingly routes to was having its own reckoning with what generative tools now cost it. Linus Torvalds released Linux 7.3-rc4 on 20 September 2026. Roughly four hundred and forty-nine commits, from around two hundred and thirty-nine contributors. Filesystem work led the cycle, SMB and NTFS hardening carrying most of it, Johannes Berg alone contributing thirty-four Wi-Fi patches.[5][6][7] Torvalds, measured rather than alarmed, credited large language model tooling with catching a genuine share of this cycle's error-path cleanup work. That is the unglamorous class of bug that rarely shows up in ordinary use but matters when it does.[5]

Stable-tree maintainer Greg Kroah-Hartman had warned it was coming eighteen days earlier. Filtering the USB subsystem's incoming review queue by hand, he narrowed it from one thousand seven hundred and thirty-two messages down to one thousand and ninety-four. He still called the result "still crazy though". He expects the workload to run a year and a half.[8] Nobody pays Kroah-Hartman by the message reviewed. Stable Linux 7.3 is targeted for 18 October 2026, or 25 October if a further release candidate is needed.[6][7]

The same open, inspectable, endlessly forkable codebase just saved a Tier-1 telco fifty-six per cent on its AI coding costs. This same month, it is absorbing a volume of AI-generated patches and bug reports its own maintainers describe as barely manageable. The bazaar is the cheapest supplier one enterprise has found this year. It is also the unpaid reviewer of the tooling now writing what gets submitted to it. Both are true of the same code, in the same release cycle.

What this means for the rest of us

For a New Zealand organisation watching this from outside a telco's scale, the practical lesson sits one level up from the model choice itself. The decision about which model handles a given request has stopped being a procurement question answered once a year. It has become an architecture question answered continuously, the same shift Linux forced on operating system choice two decades ago. None of the three steps below need a telco's budget. Three things follow. First, know what a request actually costs before deciding what should handle it. AT&T's defensible figure states the saving and the quality cost together, never one without the other. Second, build the ability to route, even modestly, rather than commit a whole workload to one vendor's roadmap and pricing. The routing layer is now worth more than any single model behind it. Third, treat the maintainers of whatever open-source layer that routing depends on as infrastructure, not as a free resource that will always be there. The same bazaar that is cheapest this year is the one absorbing the heaviest unpaid load.

Open source did not need to win an argument about freedom here. One engineer, at one company, measured what community-maintained code was actually worth against the alternative, and built that measurement into infrastructure rather than a slide deck. The 1991 Usenet post did not need a business case either. That is the same story this series has told since the first Usenet post, told now in a currency a chief financial officer reads without translation.

The mechanism inside AT&T's gateway is itself an open-source project. LiteLLM is a community-maintained routing layer that any organisation can inspect, fork or run without paying anyone for the privilege of choosing a model. AT&T did not commission it and does not own it. The company adopted it because a piece of infrastructure built in public, by contributors nobody at AT&T employs, turned out to be more useful, and cheaper, than any single vendor's closed alternative. That is the argument this series has made about Linux since its first episode, applied now to the layer above the kernel rather than the kernel itself. The model choice is not the point. The choice about the choice is, and it belongs to whoever controls the routing layer. AT&T controls its own.

The same logic runs through defence procurement, formally rather than informally. The United States Defense Federal Acquisition Regulation Supplement (DFARS), paired with the Defense Information Systems Agency's Security Technical Implementation Guides (STIGs) for Linux, does something structurally similar to AT&T's gateway. It converts a decision that used to be negotiated contract by contract, which hardening baseline, which component gets trusted, into a standing engineering requirement built into the acquisition itself. A system meets the STIG baseline or it does not. Nobody re-litigates it vendor by vendor. AT&T's routing layer and a STIG-compliant build image are doing the same kind of work at different altitudes. Both move a judgement about which underlying component to trust out of a once-a-year negotiation and into infrastructure that answers the question automatically, every time, at the moment it is asked.

At 2 Degrees you were handed the job of engineering down the cost of a new telco's infrastructure by hand, one line item, one negotiated contract, at a time. AT&T has now built that job into a piece of software and pointed it at every AI request the company sends. If you had had that gateway on your desk in 2009, which of the decisions you made by hand back then would you have handed to it first? And which one would you still insist on making yourself?

This is one of seven weekly series in The Hamberger Report. Subscribe on LinkedIn and the next one arrives in your feed.


The views expressed in this article are entirely my own, informed by more than 30 years of professional experience in architecture, security, and technology leadership in New Zealand. I write as director of Te Pono Limited; the views are personal and do not represent the position of any client, any government agency, or the New Zealand government. My commentary on legislation and policy is analytical, drawing on publicly available sources and my professional expertise in architecture, security, and AI governance, and it is politically neutral.


About the Author: Andreas Hamberger is a New Zealand-based enterprise architect and technology strategist. Over 30 years, he has moved from compiling kernels on a 486 to leading cloud, cyber, and AI transformation programmes across government, banking, transport, and aviation. He founded Yoper Linux, served as a technology specialist for Novell during the Linux Wars, and is the author of "Generative AI: Skynet or Heaven" and "Space Mafia." He can be reached at andreas@thehambergerreport.com.

A Concise History of Linux chronicles the operating system that changed the world and the lessons it holds for the AI era.


This article was produced with AI assistance under my direction. Research, drafting and images pass through a pipeline I built and govern: automated gates for source verification, forbidden language and political neutrality, and my own review before anything is published. The tools include Claude, Gemini and Openart. The frameworks, arguments and editorial judgements are mine and are the same discipline I apply to the AI systems I audit for clients. AI accelerated the work; the thinking, and the responsibility for it, are mine.


[1] PYMNTS. "AT&T Slashes AI Costs by Adopting Model Routers and Open Source." 20 August 2026. https://www.pymnts.com/news/artificial-intelligence/2026/att-slashes-ai-costs-by-adopting-model-routers-and-open-source/

[2] The Star (Malaysia), syndicating New York Times reporting by Eli Tan. "Corporate America is getting hooked on open-source AI." 7 September 2026. https://www.thestar.com.my/tech/tech-news/2026/09/07/corporate-america-is-getting-hooked-on-open-source-ai

[3] insideai.news. "AT&T Leads Corporate Shift to Open-Source AI as Usage Hits 58%." 6 September 2026. https://insideai.news/news/ai-in-business/open-source-ai-adoption/9750/

[4] Fierce Network. "Open models are driving AT&T's AI 'tokenomics' strategy." 2026. https://www.fierce-network.com/cloud/open-models-are-driving-atts-ai-tokenomics-strategy

[5] Phoronix. "Linux 7.3-rc4 Released: More Fixes Caught By LLMs, But Nothing Too Scary." 20 September 2026. https://www.phoronix.com/news/Linux-7.3-rc4

[6] LinuxCompatible.org. "Linux Kernel 7.3-rc4 Released: 449 Commits, Filesystems and WiFi Take Center Stage." 20 September 2026. https://www.linuxcompatible.org/story/linux-kernel-73rc4-released-449-commits-filesystems-and-wifi-take-center-stage

[7] OSTechNix. "Linux Kernel 7.3 RC4 Released: Bigger Than Usual With LLM Bug Fixes." 21 September 2026. https://ostechnix.com/linux-kernel-7-3-rc4-released/

[8] Phoronix. "Greg KH Forewarns Of 'Rough' Linux 7.3 Kernel Cycle Due To Continued AI Churn." 2 September 2026. https://www.phoronix.com/news/Linux-7.3-Rough-Cycle

Previous
Previous

Linux Kernel Security: The Kernel Went Weekly Because the Machines Started Finding Bugs

Next
Next

Fifteen Months Behind the Manufacturer