nonono, the difference is that they pirated it.
This is why I’m just back to pirating shit.
Especially if it’s from Amazon.
Digitizing a dvd shouldn't count as copyright infringement unless you are actively distributing the video or selling it. Bullshit.
The law overwhelmingly favors corporate entities over people. We've been over this more times than HBO Max changed their logo.
They can buy a one of a kind book from you, scan it, and shred the original and all they get is more military contracts
You are reposting a screen capture of a (probably AI generated) shortform video of a repost of a tweet screenshot.
Haha he thinks these AI companies are buying the books.
They are buying the books. Used booksellers are reporting a rash of unusual purchases: high volumes of book purchases with no regard for content or cost.
Why don't you just buy a few senators and do something about it, if you're so upset?
Yup. Donors get policies, voters get apologies.
Donors get policies, voters get ~~apologies~~ fined.
False, they never bought your book.
They're both infringement, and no-one in authority cares about either. This is kind of the situation we collectively negotiated post-Napster. Piracy is allowed to exist as long as distribution isn't being directly commercialized; it is only addressed in a piece-meal fashion by big publishers on a short-term basis for their critical first weeks.
Ripping a DVD isn't copyright infringement. If the playback program caches the output video 500ms in advance, is it infringement? If my computer has any DRAM, is it infringement? If the law says it is, the law needs to change.
"The law says Not P. If the law says P then the law is wrong." Thanks, councilor; we'll take your learned contribution under advisement.
But the LLM companies do make money off the content it was trained on, no?
Yeah? I can make money off the information I pirate. Nobody asks, "Where did you learn that knot? Did you pay for Ashley's Book Of Knots? Because if not, you're in trouble!"
This is just an example. I do have a pirated copy of ABoK but I also own an original print.
It's easy. Don't follow the rules.
That's an option, but it's got a pretty expensive penalty if you get caught. It's a lot of risk for a song / movie / game. By all means, do your thing, but you gotta be prepared for what could happen.
Yes of course. However, if all understand and do that, it's impossible to go after all.
The law has to follow the habit. So education is key, equal right for everyone.
That's the part many corrupt people have not understood. It all collapses at some point.
Someone should make an ai model that steals from other ai models: call it privateer.
Uh, they already did, and have been doing so for about 2 years now.
I agree with the conclusion, but that rationale is wrong. First, you can digitize a DVD. Second, it's not a double standard. You can grab copies of random stuff and jam it in an AI model too.
Our laws are written such that it's making a copy outside of reasonable use that's illegal, and AI training only makes a copy incidentally to what they're doing and then it's deleted. It's the same standard that makes viewing a photo on an artists website legal.
It's not bullshit because they're breaking the law, but because we need to refine the law to make it clear training an AI model isn't a reasonable usage anymore than a public broadcast of a DVD is a reasonable use.
Trying to shoehorn it into the existing laws will just create a nightmare of loopholes and complications.
AI training only makes a copy incidentally to what they’re doing and then it’s deleted. It’s the same standard that makes viewing a photo on an artists website legal.
That's not what the U.S. Copyright office says about training. They hold that it does implicate the copyright of reproduction. Meaning: If you train on a protected work without a license you are violating copyright, and if that's not a fair use then you are breaking the law.
Training ~ viewing might be an analogy used by "AI" brands, but it is not legal reality.
I'm not sure that's been extensively tested in courts. The document you referenced below appears to be as-yet not officially published, so I don't believe it actually qualifies as an official position yet, but the bigger issue is that it's untested in court.
This thread is a response to an AI court case where the ruling was that training on copy written works is fair use.
Alsup ruled in June that Anthropic made fair use of the authors' work to train Claude, but found that the company violated their rights by saving more than 7 million pirated books to a "central library" that would not necessarily be used for that purpose
Regardless, you do make good points and I think we agree that the end state is "they shouldn't be able to do that". I have concerns that using existing standards that take copying too literally results in some unintended ambiguity, and situations where AI training is incidentally blocked, but so is stuff like "opening a news article on a computer", which does the same things the copyright office report highlights as infringement.
I think we'd be in a much more agreeable place if we just legally state that a commercial AI tools training isn't fair use. That lets you have nuance like "search engine? It's a statistical model, but not generative: allowed. AI agent? Statistical model that's generating content as opposed to classification or ranking: not allowed".
I think we’d be in a much more agreeable place if we just legally state that a commercial AI tools training isn’t fair use.
Good luck getting any federal law changes through before 2028, at best. So, for at least a couple of years, we get to use the existing laws to address generative "AI"'s use of works still under copyright protection.
Any change in status before then we be policy changes by the U.S. Copyright Office, but they have already come down against training (mostly; the publications are long because there's a lot of nuance).
Alsup ruled in June that Anthropic made fair use of the authors’ work to train
But, Kadrey v. Meta Platforms, Inc. (Judge Chhabria, June 25, 2025) states that “in most cases,” training LLMs on copyrighted works without permission is likely infringing and not fair use and "this ruling does not stand for the proposition that [...] use of copyrighted materials to train its language models is lawful."
The courts are divided, but the copyright office is not.
The document you referenced below appears to be as-yet not officially published
Parts 1 and 2 are officially published.
On May 9, 2025, the Office released a pre-publication version of Part 3 [...] A final version of Part 3 will be published in the future, without any substantive changes expected in the analysis or conclusions.
Missed a major part: the AI companies are buying rare books, scanning them and then burning them. It's not like buying a DVD which can easily be digitally copied with hardware that most people have at home, it's so mucn worse than that.
Oh no no no no you got it all wrong, you have to digitize it and then burn the CD so that nobody can ever see it unless they have access to that digital copy.
Something something who else was it that burned books?
I'm not pirating, I'm building an AI model
Honestly, this sounds like the best test of copyright laws.
If corporations can train models on pirated data, and corporations are persons, then you, as a person, should be able to infringe copyright to train an AI model too.
Copyright laws are just state's monopoly on information, at this point.
The AI company doesn't even have to buy it.
remember folks: the entire problem is capitalism. it can literally all be solved by dismantling capitalism and …undismantling some guillotines.
It's not a Jellyfin library. It's a future AI training centre.

Man this community is just so sadly misguided. It's not illegal to copy even under current draconian laws - publishing is the illegal part. Not to mention that learning from copyrighted material is fair use and you can argue that LLMs are not learning but first you must come to this argument fairly rather than spewing this ludite shit - what's the point of this?
Fuck AI
"We did it, Patrick! We made a technological breakthrough!"
A place for all those who loathe AI to discuss things, post articles, and ridicule the AI hype. Proud supporter of working people. And proud booer of SXSW 2024.
AI, in this case, refers to LLMs, GPT technology, and anything listed as "AI" meant to increase market valuations.