AI Ate Your Art—Here's How to Fight Back (and Still Stay Open)
Somewhere in the weights of a large language model or an image-generation system, there's a trace of your work. Your illustrations, your blog posts, your photographs, your music. It was scraped, ingested, and used to train a commercial AI product—and nobody asked you first.
This isn't paranoia. It's documented. Lawsuits from visual artists against Stability AI and Midjourney, from authors against OpenAI and Meta, from Getty Images against multiple AI companies—they're all built on the same basic allegation: your creative work was taken without permission and used to build products generating enormous revenue.
So what do you actually do about it? And how do you protect yourself without retreating entirely from the open, collaborative web that makes creative community possible in the first place? Let's get into it.
First, Understand What's Actually Happening Legally
The honest answer is: US copyright law as it applies to AI training is genuinely unsettled right now. Courts are working through a series of cases that will shape the landscape for years, and the outcomes are not guaranteed to favor creators.
Here's the current state of play. AI companies have largely argued that training on publicly available data constitutes "fair use" under US copyright law—the same doctrine that allows criticism, commentary, and transformative works. Their position is that ingesting data to learn patterns isn't the same as copying or distributing the work.
Creators and their advocates push back hard on this. Fair use is a four-factor balancing test, not a blanket permission slip. When an AI company trains on millions of images to build a commercial product that directly competes with the artists it scraped, the "transformative" argument gets shakier. The commercial nature of the use, the scale of the copying, and the market harm to creators are all factors courts will weigh.
The US Copyright Office released guidance in 2023 clarifying that purely AI-generated works aren't eligible for copyright protection—but that guidance doesn't address the training data question directly. Legislation has been proposed but not passed. For now, the fight is happening in courtrooms, and the outcomes are mixed.
Bottom line for creators: You don't have a clear legal remedy yet, but the legal landscape is shifting. Staying informed matters.
Licensing: Your Most Practical Tool Right Now
While the courts sort things out, licensing is your most actionable lever. And this is where the conversation gets nuanced for creators who care about open culture.
Creative Commons licenses are foundational to the open web—they're how creators share work with the world while retaining meaningful rights. But they were designed in a pre-AI era, and some of them create unintended openings for AI training.
Here's a quick breakdown of what different CC licenses mean in the AI context:
CC0 (Public Domain Dedication): You've waived all rights. AI companies can and do use CC0 datasets freely. If you're contributing to open datasets, understand that CC0 means truly open—including for commercial AI training.
CC BY (Attribution): Allows commercial use with attribution. AI training likely qualifies as a "use" under this license, though attribution at scale is essentially meaningless in an ML context.
CC BY-NC (Non-Commercial): Restricts commercial use. This is a meaningful protection—if an AI company is building a commercial product, training on NC-licensed work is legally questionable. This is probably the most relevant license choice for creators who want to share openly but exclude commercial AI training.
CC BY-SA (ShareAlike): Requires derivatives to carry the same license. The "derivative" question for AI outputs is legally unresolved, but this license sends a clear signal about your intent.
For creators who want to be explicit, some are now using custom license addenda—language added to their terms of use that specifically prohibits AI training. This isn't a Creative Commons license; it's a supplemental restriction on how your work can be used. It's not foolproof (scrapers don't read licenses), but it creates a clearer legal record of your intent.
Practical Steps: What to Actually Do
1. Register your copyright. This sounds basic, but it matters. Copyright registration with the US Copyright Office is required before you can sue for statutory damages in federal court. Registration is inexpensive (around $65 for a single work online) and creates a timestamped public record of your authorship. For high-volume creators, group registration options exist.
2. Use opt-out tools where they exist. Some AI companies have created opt-out mechanisms—Google's Extended Opt-Out, for instance, and Spawning's Have I Been Trained tool, which lets you check whether your images are in common training datasets and request removal. These are imperfect and not legally binding, but they're worth using.
3. Add robots.txt and AI-specific directives to your website. The emerging standard includes directives like User-agent: GPTBot (OpenAI's crawler) and User-agent: Google-Extended in your robots.txt file. This doesn't stop all scrapers, but it signals your preferences and may matter legally.
4. Watermark strategically. Tools like Glaze and Nightshade, developed by researchers at the University of Chicago, allow visual artists to add imperceptible perturbations to images that disrupt how AI models learn from them. It's a technical countermeasure, not a legal one, but it's a meaningful act of resistance.
5. Join collective action. The Authors Guild, the Artist Rights Alliance, and organizations like Fight for the Future are actively lobbying for legislative protections. Individual creators have limited leverage; organized communities have much more.
Staying Open Without Getting Exploited
Here's the tension that's real and worth naming: many of us believe deeply in open access, shared knowledge, and the kind of collaborative creativity that Creative Commons licensing was designed to enable. The prospect of locking everything down to prevent AI scraping feels like it runs counter to those values.
But openness was never supposed to mean unconditional extraction. The Creative Commons framework was built on the idea of sharing with conditions—attribution, non-commercial use, share-alike requirements. Those conditions exist precisely because sharing without any framework can enable exploitation.
A world where individual creators share openly with each other and with the public, while commercial AI companies harvest that generosity to build proprietary products, is not an open ecosystem. It's an asymmetric one.
The goal isn't to close the creative web. It's to make sure the rules of openness apply to everyone—including the companies with the biggest servers.
What We Need From Policymakers
Individual action only goes so far. The structural solutions require policy:
- Mandatory disclosure of training datasets, so creators can know whether their work was used.
- Opt-in rather than opt-out frameworks for commercial AI training—flipping the default so that silence isn't consent.
- Collective licensing mechanisms, similar to how music performing rights organizations like ASCAP work, that allow creators to license AI training use collectively and receive compensation.
- Federal registration incentives that make copyright registration easier and cheaper for independent creators.
None of this will happen without pressure from organized creator communities. The good news is those communities are forming, getting louder, and getting better at navigating both the legal and technical dimensions of this fight.
Your work has value. The fact that a machine learned from it doesn't change that—it confirms it.