73
Open source project fools AI scrapers with poisoned font
(www.theregister.com)
A nice place to discuss rumors, happenings, innovations, and challenges in the technology sphere. We also welcome discussions on the intersections of technology and society. If it’s technological news or discussion of technology, it probably belongs here.
Remember the overriding ethos on Beehaw: Be(e) Nice. Each user you encounter here is a person, and should be treated with kindness (even if they’re wrong, or use a Linux distro you don’t like). Personal attacks will not be tolerated.
Subcommunities on Beehaw:
This community's icon was made by Aaron Schneider, under the CC-BY-NC-SA 4.0 license.
very rose colored glasses. the extra work to identify such sites is trivial less than even the hashing approach anubis uses.
risks are minimal for data pipelining on the training side. you can bootstrap a classifier that errors towards reject to sweep 99% of the weirdness this thing is doing in a few days. we already have clean datasets we can use to baseline such systems.
on top of that its fairly easy to detect if a run is failing due to collapse. the capital costs are small.
Yes, rose coloured. But I'd rather be in an environment where there are a million ways to make the life of a scraping AI company hard so it's impossible for them to account for everything.
More options, are more better especially when they are otherwise meanlessly different for users.
Also there is evidence that poisoning an AI's training set does not require much data.