Introduction
Earlier this year, an engineer at a leading global AI company forgot to add one line to a configuration file. By the next morning, the entire source code of the company’s flagship coding product had been copied across the internet. The company responded with a US Digital Millennium Copyright Act (DMCA) takedown blitz that wiped roughly 8,100 GitHub repositories (essentially online folders for software projects) overnight. The code is, of course, still everywhere. Welcome to copyright enforcement in 2026.
For Australian organisations whose value is locked up in software, content, designs and other copyright works, the question is no longer whether copyright is a meaningful protection. It is whether the enforcement tools built over the past 25 years can do anything useful once material has been absorbed into the training data of generative AI models. The short answer is: less than we would like.
What happened, and why it matters
The mechanics of the leak are simple. The product is distributed publicly in compiled form as a downloadable software package. A single line of build configuration was meant to exclude the product’s source map (a translation file that effectively contains the readable source code) from the published package. That line was missing. By the time the slip was spotted, the source had been downloaded, mirrored to file-sharing services and dissected in public.
The company filed DMCA takedown notices with GitHub, a cloud-based code sharing platform. GitHub processed them at scale, and approximately 8,100 repositories vanished overnight, including a Korean developer’s repository with around 30,000 stars (a measure of public approval). The mirrors continued on other platforms. The takedowns recovered nothing.
The leak was self-inflicted, there was no hack, no theft and no breach of confidence by an insider. Confidentiality protections largely fall away because the material was published, however inadvertently, by the rights holder (unless a basis in equity can still be established). What is left is copyright, and copyright in 2026 is a fairly blunt instrument when the “copy” you are chasing has already been read by every model on the internet.
The threshold question: Is AI-generated code even copyrightable?
The source code in that incident was written by humans, so this issue did not arise there. But for the next generation of software, the issue of copyright subsisting is a very real issue because the code is not being written by humans at all. “Vibe coding”, where a developer describes intent in natural language and an agentic AI system produces a working application end-to-end, is no longer experimental. It is the way meaningful volumes of working software are now being built.
Copyright in Australia requires a human author. The High Court in IceTV Pty Ltd v Nine Network Australia Pty Ltd (2009) 239 CLR 458 reaffirmed “the classical notion of an individual author”. The Full Federal Court applied that to computer-generated material in Telstra Corporation Limited v Phone Directories Company Pty Ltd (2010) 194 FCR 142, where Yates J observed that “in relation to works, an author is, under Australian law, a human author”. The same natural-person requirement runs across the IP regimes following Commissioner of Patents v Thaler [2022] FCAFC 62, where the Full Federal Court held an “inventor” under the Patents Act 1990 (Cth) must be a natural person; the High Court refused special leave. The US Copyright Office takes a materially similar position via its registration practice.
In the pre-AI era, copyright in code was the inevitable by-product of writing it. With AI-generated code, copyright is becoming a deliberate choice that requires documented evidence of meaningful human contribution.
A company that wants enforceable copyright in its codebase needs to show what humans actually did. A company that cannot, or does not bother, needs to make a conscious decision that copyright is not part of its protection strategy and rely on something else.
Why copyright is doing less work: The idea/expression dichotomy in the age of AI
A foundational principle of copyright is that it protects the particular expression of an idea, but not the idea itself. The plaintiff has to identify expressive material that has been copied, not functional ideas or methods. This is sensible doctrine as it permits independent creation and ensures copyright does not lock up ideas.
However, generative AI now exploits the gap with precision. Models are trained to absorb material at scale and produce new outputs that achieve the same functionality without reproducing identifiable original expression. A developer using an agentic AI tool can produce a functional equivalent of a software product in hours, with output that bears none of the textual hallmarks that would establish copying. The economic moat copyright provided, the time and skill cost of producing a non-infringing equivalent, has collapsed from weeks to hours.
The economic consequences are now visible in equity markets. Through 2025 and into 2026, listed software companies came under sustained valuation pressure, in a market correction that some commentators have called the “SaaS-pocalypse”. The investment thesis behind it is straightforward: if generative AI can produce a functional equivalent of an established product without infringing copyright (which, itself, is a highly contested issue), then the competitive moat that supported premium valuations no longer exists. The idea/expression dichotomy, designed to permit independent invention, has become a template for AI-assisted replication, and the market has begun to price that in.
The takedown apparatus shares the same blind spot. The rights-holder playbook is takedown notices under section 512 of the US Digital Millennium Copyright Act or, in Australia, the narrower safe-harbour regime in Part V Division 2AA of the Copyright Act 1968 (Cth). The mechanism is fast and asymmetric: the rights holder files, the platform pulls. But both regimes assume infringing content lives in a small number of identifiable locations, and operate by removing copies of expression once notified. A source archive can be mirrored to dozens of sites in minutes, posted to peer-to-peer networks, summarised on a hundred blogs, and absorbed into the training corpus of generative AI models. A takedown notice can address the copies; it cannot address what the models have already learned from them.
The cryptography community has been worrying for years about adversaries capturing encrypted data today to decrypt with a future quantum computer. Copyright has its own version. Material ingested into a model’s training data may never be cleanly recoverable from the model, even if every public copy is taken down. The model becomes, in effect, a permanent compressed archive of works it has read. Removing the source files does not unread the book.
The role copyright still plays
None of which means copyright has stopped doing useful work. It remains an effective tool against direct re-publishers, against bad-faith mirrors, and in damages claims where the infringement is clear and the defendant identifiable. More importantly, it remains the legal foundation for the licence-based control structures that the software industry runs on: open-source licences and end-user licence agreements rely on the underlying copyright to bind downstream users who are not in a contractual relationship with the rights holder. Without the underlying copyright, those licensing regimes lose much of their enforceability.
What has changed is the scope of the work copyright is doing. Before generative AI, copyright served two functions: it was the primary mechanism for preventing unauthorised use of software, and it was the legal foundation for licensing arrangements. It continues to provide that foundation, but it faces additional challenges in preventing unauthorised use where generative AI can quickly implement similar ideas or concepts in software.
Beyond just copyright: Choosing your IP regime
For functional software, copyright now needs to share the workload with the other IP regimes. Each has a different strength.
Patents
Patents protect underlying methods, system architectures and workflows. That is precisely the kind of subject matter generative AI now reproduces most easily and, as such, patents are an important part of your IP strategy. Patents are enforceable against independent development and reverse engineering in a way copyright is not, which is the gap generative AI exposes.
There are, of course, trade-offs. Patents can be expensive and slow. They also require patentable subject matter, novelty and inventive step thresholds to be satisfied. Further, Thaler means there has to be a human inventor. Finally, patents must generally be filed before public disclosure (subject to grace periods for self-disclosure in some jurisdictions): companies routinely shipping new features into a market should be filing protectively as a matter of course, not retrofitting after a leak or a suspected infringement.
Confidential information and trade secrets
The equitable action for breach of confidence, and the contractual frameworks that supplement it, have always required two things: that the information has the necessary quality of confidence, and that reasonable steps have been taken to maintain it.
Generative AI has changed the practical content of “reasonable steps”. The single fastest way to destroy a confidentiality claim in 2026 is to submit confidential code, business logic, system prompts or strategic documents to a public AI service whose terms permit retention or training. The threshold question is no longer just “who has the company’s confidential information?”, but “into which models has it been ingested?” Confidentiality frameworks that do not address AI use, both internally and in vendor contracts, are not doing their job.
So what is the lesson?
Copyright remains an effective tool against direct re-publishers, continues to provide the legal foundation for the licensing arrangements that the software industry depends on, and still has an indispensable place in IP strategy. But it cannot reliably, on its own, prevent leaked or published material from being absorbed into the global generative AI ecosystem, and it may not even subsist in code that the AI itself produced.
The lesson is to stop treating copyright as the default and start choosing IP regimes deliberately, asset by asset. Patents for software-implemented methods. Confidentiality for material that genuinely remains within the organisation, supported by AI-aware contractual restrictions. And, behind all of them, the operational and governance controls that prevent the leak in the first place. A layered IP strategy is required, not reliance on copyright alone.