AI’s “poisoned tree” problem
Mannn, the AI music news train doesn’t stop, does it?
K here’s an interesting question:
Can an AI company ever CLEAN UP its training data?
Because Universal and Sony are suing Suno… again.
But this lawsuit is different! And it could have some pretty enormous implications for where AI music goes next.
Quick context:
If you don’t know, Suno is one of the biggest generative AI music platforms in the world. You type in a prompt and, basically, it spits out a finished song.
The company has spent the last few years in the same giant argument surrounding basically every generative AI company:
What was the AI trained on?
Major labels have accused Suno of training its models on copyrighted recordings without permission. Suno has since launched its new v6 models, which were trained on a fresh dataset that includes music licensed from Warner Music Group and BMG.
Okay! So… problem solved?
Get licenses. Train the new model on authorized music. Move forward.
Except Universal and Sony are now arguing: “Not so fast.”
Their new lawsuit alleges that Suno’s v6 models were also trained using “user interactions” with Suno’s previous models… including outputs and preference data generated from those models.
And those previous models, the labels argue, were trained using their copyrighted recordings without permission.
Their phrase for this is pretty brutal: “The fruit of the same poisoned tree.”
And THAT is where this gets really interesting.
Because think about what they’re arguing: let’s say you build AI Model #1 using a bunch of copyrighted music you didn’t license. Millions of people use it. They generate millions of songs. They choose which generations they like. They regenerate things. They reject things. They favorite things, blah blah.
All of that activity potentially becomes incredibly valuable training data.
Then you build Model #2.
This time, you license the music! But you also train Model #2 using everything you learned from people interacting with Model #1.
So… Is Model #2 clean?
That’s basically the question. And the answer could matter WAY beyond Suno.
Because there’s an assumption floating around the AI industry that eventually all of this gets cleaned up. Deals get signed. Licenses get negotiated. Checks get written. New models get trained on authorized datasets. Everyone moves forward.
But… what if you can’t completely separate the new model from the old one?
What if the thing you’ve built isn’t just the music it trained on… but years of outputs, rankings, preferences, interactions and synthetic data that only exist because of the original model?!
Now we’re not just talking about training data. We’re talking about AI ancestry!
And that’s a much messier problem. Because you can replace a dataset. You can license a catalog. You can write a check.
But how exactly do you remove everything a machine has already learned… including everything humans subsequently taught it while using what it learned?
Ooof.
I don’t know how the courts are going to answer that.
And Suno hasn’t made its counterargument yet. But I think we’re entering a much more interesting phase of the AI copyright debate.
The first question was: “Did you have permission to train on this?”
The next one's gonna be: “Okay… but what did that training eventually become?”
And if courts decide that infringement can follow a model’s descendants?
Wheeeew.
Some AI companies may discover that getting licenses today doesn’t necessarily erase how they got here.
You can clean the dataset. The question is whether you can clean the bloodline.
K so what’s the Zag?
Stop judging AI tools ONLY by what they can do.
Start asking where they CAME from.
In other words: pay attention to provenance.
Right now, everyone wants to know which AI tool is the best. Which one makes the best songs? Which one sounds the most realistic? Which one saves me the most time?
But I think there’s another question worth asking before you build your workflow… or your BUSINESS… around one: Where did this thing COME from?
Because if the labels’ argument gains traction, the risk attached to an AI model may not disappear just because the company signs licensing deals later.
And that means provenance could eventually become a feature!
“Ethically trained” won’t just be marketing language. Licensed datasets. Transparent model lineage. Artist consent. Clear rights around outputs. Those things could become genuine competitive advantages… especially for professional creators who actually need to release, license and monetize what they make.
So yeah, by all means, pay attention to which AI tool makes the coolest shit.
But pay attention to the foundation underneath it, too.
Because in AI, we may be learning something artists have known forever:
Where something comes from… absolutely… matters.

