24
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
this post was submitted on 21 Oct 2024
24 points (100.0% liked)
TechTakes
1489 readers
30 users here now
Big brain tech dude got yet another clueless take over at HackerNews etc? Here's the place to vent. Orange site, VC foolishness, all welcome.
This is not debate club. Unless it’s amusing debate.
For actually-good tech, you want our NotAwfulTech community
founded 2 years ago
MODERATORS
It's almost completely ineffective, sorry. It's certainly not as effective as exfiltrating weights via neighborly means.
On Glaze and Nightshade, my prior rant hasn't yet been invalidated and there's no upcoming mathematics which tilt the scales in favor of anti-training techniques. In general, scrapers for training sets are now augmented with alignment models, which test inputs to see how well the tags line up; your example might be rejected as insufficiently normal-cat-like.
I think that "force-feeding" is probably not the right metaphor. At scale, more effort goes into cleaning and tagging than into scraping; most of that "forced" input is destined to be discarded or retagged.
yeah this is the thing I’ve been thinking a lot about
fucking reCaptcha is literally mass-weaponising users for data filtration, and there is no good counter besides just not using reCaptcha (which is something one can’t easily pull off without things like regulatory action, massive reputational problems that make people gtfo, etc)
I have similar worries about cloudflare being such a massive chokepoint and using that position to enable “ai bot filter” services. feels extremely monopolistic, but ianal and I’m not entirely sure what the case grounds/structure on that would be (if any)
the only other viable strategy at the moment is fully breaking contact with any potential bad traffic systems, and that’s extremely fucking dire because that’s yet another nail in the coffin of the increasingly less open internet
The whole Cloudflare bot detection is so weird and eerie. I've had issues where I can't get past it presumably just because I'm using some in-application browser just to get a login cookie, but other times it just lets fucking curl through no questions asked.
Fucking what. I've heard of sites blocking curl and I've been able to get around it by copying user agent and sometimes cookies from the browser. Now I'm cursed with the knowledge that I could probably just scrape stuff from everywhere