Nobody Decided Who Owns It: The Interface Question Inside Your AI
Nobody has settled who owns what your AI was trained on.
That should worry a board more than it usually does. Software has run this experiment before: a fight over whether an interface, an application programming interface, or API, the map of names and calls a piece of code exposes to the rest of the world, can be owned at all. It ran for eleven years, reached the Supreme Court of the United States, and ended without an answer to the one question that mattered most.
An interface is not the software itself. It is the promise the software makes to everything built on top of it: call this name, in this order, with these parameters, and this is what comes back. Build against that promise and your code works. Change what sits underneath and, if the promise holds, nobody outside has to know or care. That is what makes an interface valuable independently of the code behind it, and it is exactly what made one worth 8.8 billion dollars to fight over.
Where this starts
Go back to 2010. Oracle had just spent 7.4 billion dollars acquiring Sun Microsystems, a deal announced on 20 April 2009 and closed on 27 January 2010. Sun's own press release called Java the most important software Oracle had ever acquired, and it meant it. Sun had built Java as an open platform, "write once, run anywhere," and its business model was profiting from what got built on top of the language rather than restricting the language itself. That model had made Java close to ubiquitous, a language whose interfaces ran on billions of devices in part because Sun had never tried to lock them down.
[BOOK IMAGE: row 71 - Sun's Java platform before Oracle's acquisition]
Oracle inherited that openness along with the acquisition, and it chose a different relationship with it. On 12 August 2010, seven months after the deal closed, Oracle America filed suit against Google in the United States District Court for the Northern District of California. The claim was that Google had infringed Oracle's patents and copyrights by using parts of Java inside Android. Oracle sought 8.8 billion dollars.
What Google actually copied
Google had not copied Java wholesale. It wrote its own virtual machine, first called Dalvik and later ART, and its own implementing code underneath, the actual machinery that does the work. What it copied, for 37 of Java's API packages, was the declaring code: the names, parameters and structure that define the interface, the part a developer reads and writes against, not the machinery that answers the call. Litigation later quantified the amount at approximately 11,500 lines, roughly 0.4 per cent of the 2.86 million line Java SE platform.
[BOOK IMAGE: row 72 - Android, built on Java's interface]
Oracle's case rested on treating that 0.4 per cent as the part that mattered regardless of its size, because it was the part every Android developer's code depended on behaving exactly the way Java's did. Google's case rested on the opposite premise: that an interface, however central to a platform, is a functional method of operating a system rather than a creative expression, and functional methods sit outside what copyright protects.
Put the distinction a different way. An API is a contract: application code writes to a specification, and the specification is fulfilled by whatever sits underneath it. Google wrote its own underneath. What it reused was the specification's wording.
[BOOK IMAGE: row 73 - the API layer, declaring code against implementation]
That is not a small distinction. It is the distinction the entire case turned on, and it is the same distinction a foundation model blurs. A trained model exposes no clean seam between what it learned and how it answers; the two are fused into a single set of weights. There is no equivalent yet of "the 11,500 lines Google copied," a discrete, countable thing a court could point to. That absence does not make the underlying question go away. It only makes it harder to litigate.
Eleven years
The case ran twice through the district court and twice through the Federal Circuit before it reached the Supreme Court, four rulings and one long argument over eleven years. A first jury, in 2012, found that Google had infringed copyright but deadlocked on fair use; the district court went further than the jury and ruled that the APIs were not copyrightable at all, entering judgment for Google on 31 May 2012. The Federal Circuit reversed that in 2014, holding that the declaring code was copyrightable and sending the case back down to try fair use. A second jury, in 2016, found for Google on fair use, and the district court entered judgment for Google a second time, on 8 June 2016. The Federal Circuit reversed a second time too, on 27 March 2018, ruling that no reasonable jury could have found fair use. Only then did the case reach the Supreme Court.
[BOOK IMAGE: row 76 - eleven years, four rulings, one Supreme Court term]
On 5 April 2021, Justice Stephen Breyer delivered the majority opinion: a 6-2 ruling that Google's copying of the declaring code was fair use. Justice Thomas dissented, joined by Justice Alito. Justice Barrett took no part, having joined the Court after argument.
[BOOK IMAGE: row 74 - the Supreme Court's 6-2 fair use ruling]
What the Court did not decide
Here is the part that matters more than the verdict. The Court assumed, "purely for argument's sake," that an API could be copyrighted at all, and it addressed only whether Google's use of it was fair. It neither affirmed nor reversed the Federal Circuit's 2014 finding that APIs are copyrightable. That question, the one the 2014 ruling had actually decided and the 2021 ruling stepped around, is still, formally, open.
[BOOK IMAGE: row 75 - the API layer, open enough to build on, unresolved enough to litigate]
That gap is not a technicality. It means the eleven-year fight resolved the narrower question, whether this specific copying was fair, without resolving the broader one that nearly every software company eventually depends on somebody answering: can an interface be owned at all. Every later interface dispute inherits the same gap, unless a future case puts copyrightability itself back in front of a court willing to decide it.
Leave a question like that open and it does not disappear. It waits. As long as it stays unresolved, any company that controls a widely adopted interface keeps a theoretical claim against anyone who reimplements it, regardless of how that company behaves today. A fair use defence does not settle the matter in advance. It is decided case by case, and proving it is expensive: Oracle and Google spent more than a decade demonstrating exactly how expensive, in public, arguing the same question through two juries, two Federal Circuit panels, and one Supreme Court term before either side had a final answer.
The pattern is already repeating
The manuscript's own modern example is the cleanest test of whether the industry learned anything from that decade. Amazon Web Services' S3, short for Simple Storage Service, has become the de facto standard interface for cloud object storage. More than a hundred S3-compatible services now reimplement it, including MinIO, Backblaze B2, and Cloudflare R2. AWS has not sued a single one of them.
AWS has chosen restraint here, not certainty. Nothing in the law compels it, and nothing prevents AWS, or any successor that comes to control the interface, from deciding differently tomorrow. Every one of those hundred-plus services is built on the same theoretical exposure Google carried into court in 2010, just further from anybody testing it.
Reimplementation like this is not exotic. It is how markets avoid single-vendor lock-in: a second, third and fourth vendor can offer a compatible service precisely because customers depend on the interface, not on any one company's implementation of it. Every S3-compatible vendor is offering exactly that competitive discipline. None of them is offering it because the ownership question got answered. They are offering it because, so far, nobody with standing to sue has chosen to.
The board consequence
Training data ownership is the same unresolved fight, one layer up. When a foundation model is trained on a corpus nobody has fully cleared, the legal status of that corpus is no more settled than the legal status of an API was in 2010, and every organisation deploying the model has inherited an exposure it never priced. In thirty years advising boards on architecture and technology risk, I have rarely seen a risk register carrying a line for "the law under this vendor's training data is still open." It should.
Ask what would happen if the answer changed. A court finding that a widely used interface is copyrightable does not require the copying to have been malicious, only for it to have happened without permission and for the copyright holder to decide the exposure is worth pursuing. Oracle did not need Google to have acted in bad faith to sue it for 8.8 billion dollars. Nobody training a foundation model today needs the underlying content owner to have acted in bad faith either.
Fair use is a defence, not a guarantee, and a defence that must be proven case by case is a defence that costs money every time it is invoked, win or lose. The industry left one interface ownership question unresolved and paid for it with eleven years of litigation and a Supreme Court term that, in the end, ruled on the narrower question and left the larger one exactly where it found it. It is running the same experiment again, one layer up the stack, on infrastructure most boards have never been told to ask about, and the first company to test it in court will not be the last.
The open-source dimension of this fight predates the fight itself. Sun open-sourced Java's core platform under the GNU General Public License, announced at JavaOne in November 2006 and shipped as OpenJDK's first release in May 2007, three years before Oracle's acquisition of Sun even closed. The reference implementation of the very API family Oracle went on to litigate had already been public, GPL-licensed code for years before the lawsuit existed. That did not settle the interface question. It only proved that a company can open the implementation underneath an interface and still contest the interface itself. Declaring code and implementing code are as separable in a licence as they were in Oracle's own complaint, and an open licence on one layer says nothing at all about who controls the other.
The unresolved question sits exactly where the manuscript leaves it. The Supreme Court decided Google's copying was fair use and explicitly declined to decide whether an API is copyrightable at all. Fair use is assessed case by case and is expensive to establish; copyrightability, left open, means any sufficiently resourced holder of a widely adopted interface keeps a theoretical claim against anyone who reimplements it. Amazon Web Services' S3 application programming interface is the live example: more than a hundred S3-compatible services, including MinIO, Backblaze B2, and Cloudflare R2, now reimplement it, and AWS has not sued any of them. That restraint is a choice, renewed every year it is not withdrawn, and it is the same choice the training-data question now puts in front of every organisation building on a foundation model it did not train itself.
Has anyone at your organisation asked a vendor, in writing, what training-data ownership claim you are exposed to, and written down the answer they got?
The views expressed in this article are entirely my own, informed by more than 30 years of professional experience in architecture, security, and technology leadership in New Zealand. They do not represent the views of my employer, any government agency, or the New Zealand government. My commentary on legislation and policy is analytical, drawing on publicly available sources and my professional expertise in architecture, security, and AI governance. I follow the Public Service Commissioner's Code of Conduct for the Public Sector and social media guidance.
About the Author: Andreas Hamberger is a New Zealand-based enterprise architect and technology strategist. Over 30 years, he has moved from compiling kernels on a 486 to leading cloud, cyber, and AI transformation programmes across government, banking, transport, and aviation. He founded Yoper Linux, served as a technology specialist for Novell during the Linux Wars, and is the author of "Generative AI: Skynet or Heaven" and "Space Mafia." He can be reached at linux@linux.co.nz. Free as in Theft: The Hidden History of Open Source Software traces the openness, adoption, enclosure, and resistance cycle from the first shared source tapes to the AI licensing wars.
I use AI tools, including Sudowrite, Claude, Perplexity AI, DeepSeek AI, ChatGPT, Grok, Copilot, Openart and Gemini, as deliberate production tools, not ghostwriters. This is consistent with my position: AI amplifies human judgement; it does not replace it. The frameworks, arguments, and editorial decisions in this series are original work. AI accelerated the process. The thinking is mine.
References
[1] Hamberger, A. "Free as in Theft: The Hidden History of Open Source Software." Te Pono Limited, 2026. ISBN 978-0-473-78455-3. (Chapter 12.)
[2] Google LLC v. Oracle America, Inc., 593 U.S. 1 (2021), slip opinion. United States Supreme Court, argued 7 October 2020, decided 5 April 2021. https://www.supremecourt.gov/opinions/20pdf/18-956_d18f.pdf
[3] "Sun Releases Java Under GPL Open Source License." IT Jungle, 20 November 2006. https://www.itjungle.com/2006/11/20/tfh112006-story02-2/
[4] Wikipedia contributors. "OpenJDK." Wikipedia, The Free Encyclopedia. Accessed 17 August 2026. https://en.wikipedia.org/wiki/OpenJDK
[5] Electronic Frontier Foundation. "Oracle v. Google." Case page, procedural history 2012 to 2021. https://www.eff.org/cases/oracle-v-google

