165
submitted 8 months ago by Blaze@reddthat.com to c/reddit@lemmy.world

cross-posted from: https://infosec.pub/post/8775123

Reddit said in a filing to the Securities and Exchange Commission that its users’ posts are “a valuable source of conversation data and knowledge” that has been and will continue to be an important mechanism for training AI and large language models. The filing also states that the company believes “we are in the early stages of monetizing our user base,” and proceeds to say that it will continue to sell users’ content to companies that want to train LLMs and that it will also begin “increased use of artificial intelligence in our advertising solutions.”

The long-awaited S-1 filing reveals much of what Reddit users knew and feared: That many of the changes the company has made over the last year in the leadup to an IPO are focused on exerting control over the site, sanitizing parts of the platform, and monetizing user data.

Posting here because of the privacy implications of all this, but I wonder if at some point there should be an "Enshittification" community :-)

you are viewing a single comment's thread
view the rest of the comments
[-] twinnie@feddit.uk 4 points 8 months ago

It’s probably time I deleted my account and scrubbed my data but I’ve been almost using it as a kind of online record. If I’ve solved a problem I’ve often recorded the answer on Reddit just so I can search and find it again at a later date.

Is there anyway of scrubbing my account but downloading all my data first?

[-] umbraroze@kbin.social 4 points 8 months ago

Reddit has an user data checkout feature (IIRC, check out the user settings or maybe reddit help pages to find it).

It's a bit crap though.

It takes a long time to process, especially if you happened to post in the era when the Reddit data infrastructure was horribly terrible instead of merely ordinarily terrible, and apparently this involves some handwork in the worst cases on behalf of the staff.

Some data may be missing or truncated. It doesn't give you data from privated/banned subreddits (which was a fun thing to discover because last time I tried to do this the blackouts were on), and even for legit stuff, long comments/posts may be truncated. Even so, I'm pretty sure that the dumps just straight up didn't have all of my posts from several years ago, even if those were on public subreddits. So you need to make sure the checked out data is sensible.

In conjunction to the official dumps, I recommend a few other tools, especially since the dumps aren't really magnificently usable on their own. One tool that I found personally invaluable is reddit-user-to-sqlite, which allows you to import Reddit data dumps and available live user data (I think it does this by scraping or something, I'm sure it worked despite the API being shut down) to sqlite database, and Datasette is a nice frontend for browsing the posts.

As for scrubbing, there's tools for that are supposed to work. I think.

[-] Anticorp@lemmy.world 3 points 8 months ago

They'll just un-delete it a few days later.

this post was submitted on 23 Feb 2024
165 points (100.0% liked)

Reddit

17657 readers
88 users here now

News and Discussions about Reddit

Welcome to !reddit. This is a community for all news and discussions about Reddit.

The rules for posting and commenting, besides the rules defined here for lemmy.world, are as follows:

Rules


Rule 1- No brigading.

**You may not encourage brigading any communities or subreddits in any way. **

YSKs are about self-improvement on how to do things.



Rule 2- No illegal or NSFW or gore content.

**No illegal or NSFW or gore content. **



Rule 3- Do not seek mental, medical and professional help here.

Do not seek mental, medical and professional help here. Breaking this rule will not get you or your post removed, but it will put you at risk, and possibly in danger.



Rule 4- No self promotion or upvote-farming of any kind.

That's it.



Rule 5- No baiting or sealioning or promoting an agenda.

Posts and comments which, instead of being of an innocuous nature, are specifically intended (based on reports and in the opinion of our crack moderation team) to bait users into ideological wars on charged political topics will be removed and the authors warned - or banned - depending on severity.



Rule 6- Regarding META posts.

Provided it is about the community itself, you may post non-Reddit posts using the [META] tag on your post title.



Rule 7- You can't harass or disturb other members.

If you vocally harass or discriminate against any individual member, you will be removed.

Likewise, if you are a member, sympathiser or a resemblant of a movement that is known to largely hate, mock, discriminate against, and/or want to take lives of a group of people, and you were provably vocal about your hate, then you will be banned on sight.



Rule 8- All comments should try to stay relevant to their parent content.



Rule 9- Reposts from other platforms are not allowed.

Let everyone have their own content.



:::spoiler Rule 10- Majority of bots aren't allowed to participate here.

founded 1 year ago
MODERATORS