17
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
this post was submitted on 28 Jun 2026
17 points (100.0% liked)
TechTakes
2627 readers
37 users here now
Big brain tech dude got yet another clueless take over at HackerNews etc? Here's the place to vent. Orange site, VC foolishness, all welcome.
This is not debate club. Unless it’s amusing debate.
For actually-good tech, you want our NotAwfulTech community
founded 3 years ago
MODERATORS
officially out of the loop here:
What is an LLM "system card", does it have any sort of scientific/technical validity, or is it just performance theater by the LLM vendor?
It's the white paper-ish thing they publish when launching new models. Here's the one about fable and mythos. About half of it (~150 pages) is discussing alignment and model welfare and another third of it is benchmarks, and the rest is mostly risk evaluation, i.e. how far along Claude is on its way to paperclipping everything.
There's also a Functional Decision Theory jumpscare at 6.3.6 that I haven't heard anyone mention yet, apparently Claude has a tendency to defer to Yud's half baked sham of a decision theory:
That’s not a card! That’s a book!!! If they can’t get this simple classification right, how am I supposed to trust their probabilistic text extruder?
Oh don’t worry, it’s slop-generated anyway. You can ask the LLM to summarize it for you.
Thanks! Does every LLM vendor publish them, or is it an Anthropic thing only?
And of course it's self-published by the vendor, so basically just PR.
The real question is if there're any sort of standards to what constitutes a system/model card, which I don't think so, as far as I can tell it just has to look like a publishable paper, openAI even uploads theirs to arxiv.
Otherwise yes, google returns a bunch of cards for a bunch of vendors, so it's safe to say it's a widespread practice.
IIRC the cards thing was originally from a Gebru paper https://arxiv.org/abs/1810.03993 but that dates from the "fairness" era and not the "safety" era. Hugging Face has "a" standard - https://huggingface.co/docs/hub/en/model-cards - but I don't think it's "the" standard.
More the later than the former... they are better than purely marketing focused stuff pushed out by the LLM companies, and if you dig through them and read between the lines you can occasionally sift out useful details. Like here is a pretty solid sneer digging through Mythos's 'system card' and pointing out all the ways it contradicts the hype and press headlines Anthropic was pushing.
But even so they have some big problems...