A 21-Million-Song "Raw Material Pool": How One Search Box Is Rewriting the Music Industry's AI Ledger
Updated:
In late June 2026, Grammy-winning singer SZA posted a Story on Instagram with a screenshot of search results: "just checked. AI was trained on 238 of my songs. I'm sure some were never released." A second line followed: "If you're a musician and you support this degenerate shit? You're DISGUSTING and there's NOTHING YOU COULD EVER SAY TO ME TO MAKE THIS OKAY." (Rolling Stone)
What let her "just check" was the AI Watchdog project run by Alex Reisner of The Atlantic — that June, the outlet formally extended its monitoring from books, papers and video into music, publishing the large music datasets circulating in the AI development community and launching a searchable database. Overnight, an industry-wide self-audit swept through global music, from trade associations and Grammy winners to bedroom musicians "not famous at all": the database catalogs the works of more than 250,000 artists, and singer Kehlani posted her own search results — 179 songs (Rolling Stone).

Why could one small search box do so much damage? Because it laid, for the first time, a murky account that had long hung over the industry directly in front of every creator: your work has long since entered the candidate raw-material pool that AI developers can freely grab and call on — and you knew nothing about it.
One Search Box Makes the Murky Account Visible
AI Watchdog launched in 2025 as an investigative project by The Atlantic, created specifically to trace the copyright provenance of AI training data. According to the original report's account of The Atlantic's investigation, the four music datasets made public this time include two giant catalogs — LAION-DISCO-12M (12 million tracks) and Sleeping-DISCO-9M (9 million tracks) — and two mid-sized ones, each holding more than 100,000 recordings: the Spotify Tracks Dataset and the Free Music Archive Dataset. Add them up: more than 21 million tracks.

Rolling Stone's reporting corroborates the largest of them: LAION-DISCO-12M, released in November 2024 by the German nonprofit LAION, holds more than 12 million songs scraped from YouTube, paired with metadata "to support basic machine learning research in foundation models, music information retrieval, and audio dataset analysis." The organization's other claim to fame: it built the training dataset behind the image-generation model Stable Diffusion (Rolling Stone).

Per the official statement relayed in the original report, LAION-DISCO-12M is "for academic research only," with commercial use explicitly forbidden. But in reality, a "research only" restriction runs into murky legal judgment and the diffusion momentum of open-source communities — once released, a dataset's downstream flow is no longer under its publisher's control. The other giant catalog, Sleeping-DISCO-9M, is more blunt about the problem: per the original report, it draws its core material from YouTube content and Genius lyrics, with developers using cloudscraper to bypass Cloudflare's scraping restrictions and match lyrics, metadata and YouTube links — a pretraining dataset built for generative music models.
The Spotify Tracks Dataset's trouble is provenance transparency: a batch of tracks scraped from Spotify, uploaded to the open-source community Hugging Face by an unidentified AI developer; Spotify has explicitly disavowed any connection. The relatively traceable one of the four is the Free Music Archive Dataset — compiled in 2016 by EPFL (Swiss Federal Institute of Technology in Lausanne) from the Free Music Archive, mostly Creative Commons-licensed, and long used as a benchmark dataset in music information retrieval; both Google and Stability AI have confirmed in model documentation that they used portions of it.

That said, AI Watchdog itself stresses two caveats: the four catalogs do not exhaust the sources of training data; and a song appearing in a dataset does not necessarily mean it was actually used to train a model. What the database proves is "availability," not "use." But another fact completes the puzzle: in legal filings responding to the RIAA lawsuit, Suno itself admitted its model was trained on "tens of millions of recordings" available online (Rolling Stone). How big the candidate pool is, and whether the water in it has been drunk — both sides of the account now match.
The Self-Audit Wave: Rage, Dark Humor, and "Seen by a Crawler"
In the first week after the search box went live, the global music scene's reaction can be summed up in two phrases: rage, and dark humor.
Darren Hayes, frontman of Australia's Savage Garden, wrote that every piece of work he had created or performed over the past 30 years had been stolen by AI software. Dean Ormston — chair of Australian collecting society APRA AMCOS and current chair of CISAC, the international confederation of authors' societies — said AI companies are stealing the works of Midnight Oil, Sia, Crowded House, Lorde and others (per the original report; Ormston's CISAC chairmanship per Billboard).

SZA pointed her spear at a specific person. After calling out superstar producer Diplo for "holding equity in Suno" and "trying to train it on the best and brightest black minds of writers and producers," she pushed the conflict into cultural exploitation: "We're 13% of the US population and influence the whole world with our voices and perspectives. I have yet to hear a white AI song… We have no protection in legislation, healthcare or creativity — we're the easiest to steal from." (Variety)

One widely circulated claim here needs untangling. The original report said Diplo "subsequently clarified that he does not hold Suno shares" — but per Variety's fact-check, whether Diplo holds Suno equity remains unsettled; what is confirmed is that he invested in another AI startup, Aaru. What's really worth noting is his position: in April he declared "there's no fighting AI," said he no longer needs human voices on his tracks because "I can get the best voice from AI," and told AI-averse musicians to "adapt or just give up and become an Uber driver until everyone has a Waymo" (Variety; the original report's "clarification" framing is corrected here). Down the same news cycle, Suno chief product officer Jack Brody's response was that Suno's training metadata does not include artists' names, that the model cannot replicate its training material, and that the platform keeps improving impersonation detection (Variety).
Producer Kenneth Blume — better known by his artist name, Kenny Beats — was less polite. On X he let fire: "You are true losers. I can't imagine going into work daily knowing you are stealing from countless struggling musicians." The last line: "Get fucked, every single one of you." (Rolling Stone)
The indie scene's reaction bordered on absurdist theater. Retro-electronic artist DJ Sabrina the Teenage DJ, after finding 22 of her songs in the dataset, closed the loop perfectly on Bluesky: "Did the people saying my music sounds like AI garbage ever consider it's because Suno used a dataset containing my 22 songs?" Titus Andronicus' frontman offered dark humor: "They took deep cuts from an album nobody listened to… good luck, guys." On Reddit, one user wrote: "Found my own song — and I'm not famous at all!" (per the original report)


That's the part that stings most. For top-tier artists, AI training means commercial value being invoked without permission; for mid-tier and grassroots musicians, it can feel like belated recognition — the work finally got seen, except by a crawler instead of an audience. And the creator's agency gets inverted: AI swallows the creator first, then the market turns around and judges the creator by AI's standards. Even a distribution-company executive admitted to Rolling Stone that he himself had been fooled by an AI "act," bursting into his team shouting "how is this kid not on our radar yet" (Rolling Stone).
The Majors Split: Settling on One Front, Buying Tickets on Another
Rage aside, what actually sets the rules of the game is capital's moves. And over this past year-plus, the three major labels' paths have split completely.
The starting point was June 2024: Universal, Sony and Warner, under the banner of the RIAA, jointly sued Suno and Udio for mass copyright infringement, accusing them of copying recordings without authorization to train their models.

On October 29, 2025, Universal broke ranks first: a package deal with Udio — settling the lawsuit and announcing a jointly built, licensed AI music platform. Under the deal, the revamped Udio is due to launch in 2026 using only licensed material from UMG artists who choose to participate; users can remix and riff inside a "walled garden," but tracks can't leave the platform. The same day the news broke, Udio pulled its users' download feature overnight, and paying subscribers erupted on Reddit; three days later Udio relented with a 48-hour download window (Billboard). On November 19, Warner followed with its own settlement and licensing deal with Udio (Billboard); that same day, Suno announced a $250 million raise valuing it at $2.45 billion (Billboard). At month's end, Warner went one further and settled with Suno: Suno must build a new model on licensed music and shut down the old one.
In other words, Warner became the only major to make peace with both Suno and Udio. Universal settled only with Udio; its lawsuit against Suno continues to this day. Sony settled with neither. Udio kept "buying its tickets": in January 2026 it signed Merlin, the alliance of independent labels (Billboard), and on April 9 it struck a partnership with indie publishing giant Kobalt (Billboard) — AI music companies litigating with one hand while queueing up to reconnect to the licensing system with the other.
But the new licensing deals did not automatically solve creators' share of the pie — they tore open a second rift. On June 5, 2026, the American Federation of Musicians (AFM), the union for session musicians, sued Universal and Warner: it accused the two majors of licensing recordings featuring its members' work to AI companies with no compensation, no credit, and even a refusal to tell the union which recordings had been licensed (Billboard).

On June 22, a global coalition of artists, songwriters and managers' groups published an open letter that put it even more bluntly: labels and publishers "rightly argue that AI companies need permission to train on their music catalog, but will not grant artists and songwriters the same rights." Many artists received notices that they would be "opted in by default, with little actual choice offered"; new contracts have begun embedding AI-rights clauses as a standard condition of signing (Billboard). In other words: on the question of "who authorized this," the labels answered it for themselves and left no answer sheet for the creators.
The Munich Ruling and the "Poisoned Tree": The Battle Is Not Over
If the majors' settlements made the first half of 2026 look like a shift "from confrontation to negotiation," a ruling on July 31 reminded everyone the battle was far from over — and it landed exactly one month after the original report was published.
The Munich Regional Court ruled that Suno had violated German copyright law by feeding large numbers of German works administered by GEMA — including Boney M.'s "Rasputin," Alphaville's "Forever Young" and Lou Bega's "Mambo No. 5" — into its model without a license and without payment, and ordered financial damages (not yet quantified; the ruling is appealable). It is one of the first major legal defeats for an AI music company anywhere. GEMA CEO Tobias Holzmüller's statement was unsparing: "The court made it clear today: AI models based on the theft of intellectual property are not protected by law." Suno responded that the ruling "rests on a fundamental mischaracterization of how Suno's technology works" and is evaluating an appeal (Billboard).
Across the ocean, the fight escalated in parallel. From April 2026, in their lawsuit against Suno, Universal and Sony argued that Suno's settlement with Warner "bears directly" on market harm, the heart of any fair-use analysis — "a market for such licenses exists, and Suno is now a repeat, paying participant in it." Suno, meanwhile, fought in court to keep Universal and Sony from seeing the Warner settlement's terms, arguing it was "shaped by litigation risk, not by the competitive forces that define a functioning market," and so proves nothing (Billboard).
On September 9, Suno released its v6 model, billed as built "in partnership with the music industry" — training material licensed from Warner, BMG, Believe and other partners, revenue settled by share, the old model shut down per the agreement, downloads restricted (Billboard). Warner CEO Robert Kyncl, in an internal memo, called it "a smart win" and framed Warner's strategy as "three Ls": licensing, litigation, legislation — all three at once, none optional (Billboard).
Nine days later, Universal and Sony's answer arrived: on September 18 the two majors jointly filed a new lawsuit against Suno, adding 61,000 songs to the docket and handing down a verdict on v6 — "fruit of the same poisoned tree." The reason: beyond licensed material, v6 was trained on the outputs and user interactions of the earlier models (so-called synthetic data), which does not fix the training problem but merely "launders" it. The new complaint turned Suno's own licensing deals into evidence: having become "a repeat, paying participant" in a licensing market, Suno can no longer credibly claim no such market exists (Billboard).
One opposing view deserves mention: Tori Noble, an attorney at the digital-rights nonprofit Electronic Frontier Foundation, backs Suno's argument here, holding that the Meta precedent "is probably the most correct outcome on the law" — "otherwise, the rightsholder could create a market for licenses that they don't have a right to license… almost like letting someone put up gates around a public park and charge for access." Likewise, in the book authors' case against Anthropic, at least one U.S. judge has already sided with the AI industry on fair use for training (Billboard). Whether AI training is fair use remains unsettled in the United States — which is precisely the question everyone is betting on.
Writing the Ledger: From "Training Royalties" to Attribution Tech
The legal war settles the past; ledger design decides the future. Once training data became a real business, the industry had to invent new bookkeeping.
The market's size is on the table: per figures cited in the original report from research firm Global Info Research, the global AI training dataset market was worth about $1.847 billion in 2025 and is projected to reach $11.458 billion by 2032, a 29.7% CAGR; another firm estimates up to $23.18 billion by 2034 (these figures could not be independently verified; treat with caution). Estimates differ, but the direction agrees: data is becoming the most expensive bargaining chip of the AI era.

The ledger's core difficulty: traditional copyright distribution is built on "one use, one payment," while AI training distills a work into a model once and lets it resurface in outputs indefinitely. So the new ledger has to split compensation apart — and that is now actually happening.
STIM, the Swedish collecting society mentioned in the original report, delivered the first finished homework on September 9, 2026: the industry's first collective-management AI music license, signed with AI music company Songfox and attribution-tech firm Sureel. Per Billboard's breakdown, the license splits compensation into three stages: an up-front payment for training, a share of the AI company's revenue, and a further share of revenue from AI music built on the licensed works; each work must be explicitly opted in, compositions and recordings split 50-50, with Sureel's attribution technology handling monitoring. STIM's management called it "a blueprint," explicitly analogizing the logic to the streaming transition — "pursue those who are stealing, but also offer a way to operate legally" (Billboard).

Industry and academia had been paving this road for a while. Per the original report, a paper led by Sony AI researchers, "Attribution-by-design" (October 2025), proposed distinguishing training-time attribution from inference-time attribution and building verifiable provenance tracking and royalty distribution; and the earlier "Computational Copyright" paper from University of Illinois Urbana-Champaign researchers suggested borrowing the distribution logic of Spotify and YouTube — using attribution technology to determine which training works influenced a given AI-generated track and designing revenue-share models accordingly. (Details of both papers follow the original report's account.) That STIM's license includes a role for an "attribution technology company" shows the papers' blueprint is turning into contract clauses.
Upstream, the change has already been written into record contracts. Billboard's reporting shows several labels and publishers adding AI-training clauses to new agreements — Sony's dance imprint B1 Recordings even wrote in "unlimited, exclusive rights" to use recordings "in models and systems of generative artificial intelligence… including AI training"; top music attorneys warn that labels could invoke pre-existing blanket licensing clauses (like those used for TikTok or Instagram) to feed works into AI training sets without individual artist consent (Billboard). The ledger is being rewritten at both ends: the new money AI companies pay labels on one end, and the renegotiation of old contracts between labels and creators on the other.
How Big Are the Losses? CISAC's Decade-Long Warning
What exactly are creators losing? The most systematic quantification comes from the study commissioned by CISAC, the international confederation of authors' societies, from strategy consultancy PMP (released November 2024): by 2028, generative AI will take about €4 billion a year from songwriters and rightsholders — 24% of the revenue they collect through collective management organizations — by which point the generative AI music market itself will be worth about $16 billion (Billboard). CISAC president and ABBA frontman Björn Ulvaeus put it plainly at the study's launch: "The success of AI isn't based on public content — it's based on copyrighted works. We need to negotiate a fair deal."
The same study, as relayed by the original report, also made structural projections: by 2028, AI music is expected to account for 20% of music streaming platforms' revenue and 60% of music-library revenue. What gets eroded first is likely not stars' new albums but the "background music" business — the functional music played in restaurants, shops and low-budget productions. Billboard's columnist offered a sharp reminder: background music keeps a large population of musicians fed — "if that business declines, will those musicians still be able to afford the studio bill?" (Billboard)
Perhaps the study's coldest line: "In an unchanged regulatory framework, creators will not benefit from the Gen AI revolution." What they face is twofold — the revenue lost when their works are used for training without authorization, and the structural substitution of their traditional income by AI-generated output. That is why SZA's rage ultimately converges on three concrete questions: Who authorized this? Who profits? If a song has already become part of an AI's capability, why is its creator still excluded from the distribution of the gains?
Conclusion: The Rules Are Growing Out of the Details
Legislation around artistic style is another front in this ledger-rewriting. The U.S. already has precedents: Tennessee's ELVIS Act of 2024 protects artists' voices from AI impersonation (Billboard), and the federal NO FAKES Act keeps moving through Congress (Billboard). Per the original report, in June 2026 U.S. lawmakers also introduced the CREATOR Act for visual artists, giving creators the right to block and claim damages when AI imitates their distinctive visual style commercially without permission — current copyright law protects specific works, not style (no independently verifiable English coverage of this bill was found; relayed here per the original report). Most of this legislation is still early-stage, but the direction is consistent: pushing "style-like rights" out of academic debate and into statute.

The platforms are moving too. On June 29, TIDAL announced it would identify and tag music found to be almost entirely AI-generated — tagged releases can stay on the platform but are no longer eligible for monetization; Spotify and Apple Music have rolled out similar identification features (Rolling Stone). And AI music has not stayed at the "slop" level: an AI-completed song built with Suno, "Let Me Be," spent six weeks at No. 1 on Billboard's U.S. Afrobeats Songs chart, and AI artist Xania Monet received a $3 million record-deal offer — though Kehlani publicly said, "I don't respect it" (Billboard).
So lay 2026's picture side by side: Munich's ruling gave "unlicensed use is infringement" its first judicial precedent; Universal and Sony used a new lawsuit to declare that settlement is not the end — v6 doesn't get a pass either; Warner's "three Ls" prove licensing, litigation and legislation can run in parallel; STIM's first collective AI license wrote training, subscription and output-stage revenue into a contract; and the AFM and the artists' coalition remind everyone that the labels' new money has not yet reached the creators.
Which circles back to the deepest layer of this affair: the question has long since moved from "whether to fight back" to "how to price and split." The core battlefield of the future is not under the courtroom's spotlight but in the quieter details of training, generation, attribution and distribution. What the AI era truly needs rewritten is not just copyright law but the music industry's new ledger — and right now, every page of that ledger is being written and fought over at the same time.
Sources
- 音乐先声 (Music First Sound): 《音乐圈版权保卫战:AI是如何洗劫音乐人的?》 (original article, in Chinese)
- Rolling Stone: Have Artists Reached Their Breaking Point With AI?
- Variety: SZA Slams 'Disgusting' Musicians Using AI, Says Platforms Like Suno Train on the 'Best and Brightest Black Minds of Writers and Producers'
- Billboard: UMG-Udio Deal FAQ: What Questions Remain About the AI Agreement That's Shaken Up the Music Biz?
- Billboard: Warner Music Settles With Udio, Signs Deal for Licensed AI Music Platform
- Billboard: Udio Strikes AI Licensing Deal With Merlin for Independent Labels
- Billboard: Udio and Kobalt Ink Partnership
- Billboard: Musicians Union Sues UMG, WMG Over AI Settlements: 'Refused to Provide Compensation'
- Billboard: Musicians Pen Letter, Warning About AI Music Deals: 'Innovation Cannot Be Used to Override Artists' Rights'
- Billboard: Suno Held Liable for Infringing German Song Copyrights in Landmark Court Ruling
- Billboard: Danish Rights Group Koda Sues Suno for 'Biggest Theft in Music History'
- Billboard: Licenses Versus Lawsuits: Why Suno's Parallel Legal Strategies May Pose a 'Conundrum'
- Billboard: UMG & Sony Hit Suno With New Lawsuit After Label-Backed Model: 'Fruit of the Same Poisoned Tree'
- Billboard: Suno Launches First AI Music Models 'in Partnership With the Music Industry'
- Billboard: Warner Music Group CEO Pens Memo About Suno Partnership: 'This Is a Smart Win'
- Billboard: As Some Labels Strike Deals With Suno, Artist Reps Sound Off: 'It's the Wild, Wild West'
- Billboard: STIM Explores AI Music Licensing Framework With Songfox Deal: 'This Is a Blueprint'
- Billboard: Will Generative AI Become a Creator Terminator?
- Billboard: CISAC Elects APRA AMCOS CEO Dean Ormston as New Chair
- Billboard: Tennessee Adopts ELVIS Act, Protecting Artists' Voices From AI Impersonation
- Billboard: NO FAKES Act Returns to Congress With Support From YouTube, OpenAI for AI Deepfake Bill
- Billboard: Kehlani Slams AI Artist Xania Monet Over $3 Million Record Deal Offer: 'I Don't Respect It'