20
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
this post was submitted on 20 Jul 2026
20 points (100.0% liked)
TechTakes
2623 readers
25 users here now
Big brain tech dude got yet another clueless take over at HackerNews etc? Here's the place to vent. Orange site, VC foolishness, all welcome.
This is not debate club. Unless it’s amusing debate.
For actually-good tech, you want our NotAwfulTech community
founded 3 years ago
MODERATORS
Adversarial tokenmaxxing could probably be done by using non-English character sets in lieu of English letters (e.g. faux Cryllic) - for two examples from the Greek alphabet, alpha and omicron alone can easily substitute for A and O, respectively.
As a bonus, this would likely make the text look like complete gibberish to LLMs, potentially leaving them unable to process the document altogether. This would probably shaft anyone using screen readers, though.
EDIT: Turns out the demonstration's already caught on to this idea, didn't notice beforehand:
I feel like there's a version of this that has a slider for how aggressively you're willing to sacrifice readability, and you could probably get pretty decent results on the scale of 2x to 2.5x just using different encodings of the same basic glyph.
I tested this with the Navy Seal Copypasta (well, the first 400 letters of it, the demo's got a limit) - using just Cryllic characters got me a 3.01x increase, and turning on all four control types got a 4.53x increase.
Testing your own recent comment, Cryllic only got 3.38x, and all four controls got 5.11x.
Going from those two, the boost from homoglyphs alone is likely higher than you think - 3x to 3.5x, by my guess - pretty good for human readable text.
Tarpits like Iocaine and Nepenthes can easily sacrifice readability for token burn, so they can easily go higher.