529

‘Impossible’ to create AI tools like ChatGPT without copyrighted material, OpenAI says (www.theguardian.com)

submitted 2 years ago by L4s@lemmy.world to c/technology@lemmy.world

300 comments fedilink hide all child comments

‘Impossible’ to create AI tools like ChatGPT without copyrighted material, OpenAI says::Pressure grows on artificial intelligence firms over the content used to train their products

you are viewing a single comment's thread
view the rest of the comments

[-] lolcatnip@reddthat.com 8 points 2 years ago

I don't understand why people are defending AI companies

Because it's not just big companies that are affected; it's the technology itself. People saying you can't train a model on copyrighted works are essentially saying nobody can develop those kinds of models at all. A lot of people here are naturally opposed to the idea that the development of any useful technology should be effectively illegal.

[-] assassin_aragorn@lemmy.world 15 points 2 years ago

This is frankly very simple.

If the AI is trained on copyrighted material and doesn't pay for it, then the model should be freely available for everyone to use.
If the AI is trained on copyrighted material and pays a license for it, then the company can charge people for using the model.

If information should be free and copyright is stifling, then OpenAI shouldn't be able to charge for access. If information is valuable and should be paid for, then OpenAI should have paid for the training material.

OpenAI is trying to have it both ways. They don't want to pay for information, but they want to charge for information. They can't have one without the either.

[-] BURN@lemmy.world 11 points 2 years ago

You can make these models just fine using licensed data. So can any hobbyist.

You just can’t steal other people’s creations to make your models.

[-] lolcatnip@reddthat.com 6 points 2 years ago

Of course it sounds bad when you using the word "steal", but I'm far from convinced that training is theft, and using inflammatory language just makes me less inclined to listen to what you have to say.

[-] BURN@lemmy.world 10 points 2 years ago

Training is theft imo. You have to scrape and store the training data, which amounts to copyright violation based on replication. It’s an incredibly simple concept. The model isn’t the problem here, the training data is.

[-] lolcatnip@reddthat.com 1 points 2 years ago

Training is theft imo.

Then it appears we have nothing to discuss.

[-] dhork@lemmy.world 9 points 2 years ago

I am not saying you can't train on copyrighted works at all, I am saying you can't train on copyrighted works without permission. There are fair use exemptions for copyright, but training AI shouldn't apply. AI companies will have to acknowledge this and get permission (probably by paying money) before incorporating content into their models. They'll be able to afford it.

[-] lolcatnip@reddthat.com 3 points 2 years ago

What if I do it myself? Do I still need to get permission? And if so, why should I?

I don't believe the legality of doing something should depend on who's doing it.

[-] BURN@lemmy.world 5 points 2 years ago

Yes you would need permission. Just because you’re a hobbyist doesn’t mean you’re exempt from needing to follow the rules.

As soon as it goes beyond a completely offline, personal, non-replicatible project, it should be subject to the same copyright laws.

If you purely create a data agnostic AI model and share the code, there’s no problem, as you’re not profiting off of the training data. If you create an AI model that’s available for others to use, then you’d need to have the licensing rights to all of the training data.

this post was submitted on 09 Jan 2024

529 points (100.0% liked)

Technology

86309 readers

2746 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related news or articles.
Be excellent to each other!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
Check for duplicates before posting, duplicates may be removed
Accounts 7 days and younger will have their posts automatically removed.

Approved Bots

founded 3 years ago

MODERATORS

L3s@lemmy.world

enu@lemmy.world

technopagan@lemmy.world

L4s@lemmy.world

L3s@hackingne.ws