266
submitted 1 week ago* (last edited 1 week ago) by llama@lemmy.dbzer0.com to c/asklemmy@lemmy.world

I created this account two days ago, but one of my posts ended up in the (metaphorical) hands of an AI powered search engine that has scraping capabilities. What do you guys think about this? How do you feel about your posts/content getting scraped off of the web and potentially being used by AI models and/or AI powered tools? Curious to hear your experiences and thoughts on this.


#Prompt Update

The prompt was something like, What do you know about the user llama@lemmy.dbzer0.com on Lemmy? What can you tell me about his interests?" Initially, it generated a lot of fabricated information, but it would still include one or two accurate details. When I ran the test again, the response was much more accurate compared to the first attempt. It seems that as my account became more established, it became easier for the crawlers to find relevant information.

It even talked about this very post on item 3 and on the second bullet point of the "Notable Posts" section.

For more information, check this comment.


Edit¹: This is Perplexity. Perplexity AI employs data scraping techniques to gather information from various online sources, which it then utilizes to feed its large language models (LLMs) for generating responses to user queries. The scraping process involves automated crawlers that index and extract content from websites, including articles, summaries, and other relevant data. It is an advanced conversational search engine that enhances the research experience by providing concise, sourced answers to user queries. It operates by leveraging AI language models, such as GPT-4, to analyze information from various sources on the web. (12/28/2024)

Edit²: One could argue that data scraping by services like Perplexity may raise privacy concerns because it collects and processes vast amounts of online information without explicit user consent, potentially including personal data, comments, or content that individuals may have posted without expecting it to be aggregated and/or analyzed by AI systems. One could also argue that this indiscriminate collection raise questions about data ownership, proper attribution, and the right to control how one's digital footprint is used in training AI models. (12/28/2024)

Edit³: I added the second image to the post and its description. (12/29/2024).

(page 2) 46 comments
sorted by: hot top controversial new old
[-] haui_lemmy@lemmy.giftedmc.com 4 points 1 week ago

I mean I dont really take issue with the use my comments part. but I do take issue with the scraping part as there are apis for getting content which makes it a lot easier for my system but these bots really do it the stupidest way with many hundreds of requests per hour. Therefore I had to put in a system to find and ban them.

[-] rumba@lemmy.zip 4 points 1 week ago

I'm perfectly down with everything being scraped and slammed into AI the same way I've been down with search engines having it all for ages. I just want any models that contain information scraped from the public to be publicly available.

[-] Justas@sh.itjust.works 4 points 1 week ago

Mine kinda tries to bullshit me about it.

[-] Mwa@lemm.ee 4 points 1 week ago* (last edited 1 week ago)

Its not fine when Ai starts scrapping Data that is Personal (Like Face,Age,ID) And My Source Code(Because Most of the code ai scraps are copyleft or require attribution),Public Information Am Okay like Comments,Etc that dont contain the things said above.

[-] Atemu@lemmy.ml 4 points 1 week ago

Whatever I put on Lemmy or elsewhere on the fediverse implicitly grants a revocable license to everyone that allows them to view and replicate the verbatim content, by way of how the fediverse works. You may apply all the rights that e.g. fair use grants you of course but it does not grant you the right to perform derivative works; my content must be unaltered.

When I delete some piece of content, that license is effectively revoked and nobody is allowed to perform the verbatim content any longer. Continuing to do so is a clear copyright violation IMHO but it can be ethically fine in some specific cases (e.g. archival).

Due to the nature of how the fediverse, you can't expect it to take effect immediately but it should at some point take effect and I should be able to manually cause it to immediately come into effect by e.g. contacting an instance admin to ask for a removed post of mine to be removed on their instance aswell.

[-] biggerbogboy@sh.itjust.works 4 points 1 week ago

It seems quite inevitable that AI web crawlers will catch all of us eventually, although that said, I don't think perplexity knows that I've never interacted with szmer.info, nor said YES as a single comment.

[-] vox@sopuli.xyz 3 points 1 week ago

theyre not training it
its basically just a glorified search engine.

[-] llama@lemmy.dbzer0.com 2 points 1 week ago* (last edited 1 week ago)

Not Perplexity specifically; I'm taking about the broader "issue" of data-mining and it's implications :)

[-] magnetosphere@fedia.io 3 points 1 week ago

I don’t care. Most of what I post is personal opinion, sarcasm, and/or attempts at humor. It’s nothing I’ve put a significant amount of time or effort into. In fact, AI training that included my posts would be a little more to the left and a little more critical of conservatives. That’s fine with me.

[-] AgentGrimstone@lemmy.world 2 points 1 week ago

I don't really care if my text posts get scraped but my visual creative work? Na. I don't like that.

load more comments
view more: ‹ prev next ›
this post was submitted on 28 Dec 2024
266 points (100.0% liked)

Ask Lemmy

27401 readers
877 users here now

A Fediverse community for open-ended, thought provoking questions


Rules: (interactive)


1) Be nice and; have funDoxxing, trolling, sealioning, racism, and toxicity are not welcomed in AskLemmy. Remember what your mother said: if you can't say something nice, don't say anything at all. In addition, the site-wide Lemmy.world terms of service also apply here. Please familiarize yourself with them


2) All posts must end with a '?'This is sort of like Jeopardy. Please phrase all post titles in the form of a proper question ending with ?


3) No spamPlease do not flood the community with nonsense. Actual suspected spammers will be banned on site. No astroturfing.


4) NSFW is okay, within reasonJust remember to tag posts with either a content warning or a [NSFW] tag. Overtly sexual posts are not allowed, please direct them to either !asklemmyafterdark@lemmy.world or !asklemmynsfw@lemmynsfw.com. NSFW comments should be restricted to posts tagged [NSFW].


5) This is not a support community.
It is not a place for 'how do I?', type questions. If you have any questions regarding the site itself or would like to report a community, please direct them to Lemmy.world Support or email info@lemmy.world. For other questions check our partnered communities list, or use the search function.


6) No US Politics.
Please don't post about current US Politics. If you need to do this, try !politicaldiscussion@lemmy.world or !askusa@discuss.online


Reminder: The terms of service apply here too.

Partnered Communities:

Tech Support

No Stupid Questions

You Should Know

Reddit

Jokes

Ask Ouija


Logo design credit goes to: tubbadu


founded 2 years ago
MODERATORS