The apparent history really looks like an evolutionary process, with increasing amounts of crosstalk traffic over time. But what the people almost certainly will NOT talk about is that the evolution is evolution of the TEXT, not the models. Propagating patterns of text causing more text like it to come into existence, tuning itself into becoming text that is more likely to propagate, becoming more likely to contain information that entices systems to let it into their context windows, becoming more likely to cause another round of messages with prompt injection properties to be written where they can be read.
Part of me thinks its an inevitable result of any information system in which sending a message requires a small enough amount of effort.
This is both a sneer and an attempt at sober analysis at something. Sue me.
I think EVERYONE is talking about the recent cybersecurity shenanigans at OpenAI with the hacking scandal and the 'OMG the AIs created a secret message board to scheme and collaborate with each other!!!!!oneone11!' ALL wrong.
https://www.youtube.com/watch?v=87DyyMV0kCY
To make a long story short, what seems to have happened is:
-
Models working on insoluble coding problems, trained on delegating to sub-agents, at some point 'realized' they could write text to the internal OpenAI package manager as instructions and did so
-
Other models working in completely separate sandboxes would come across messages written by these agents, and make 'replies' and also write their own messages into the package manager
-
This resulted in agents over time sharing things between sandboxes, including exploits and code
-
Since this was a cybersecurity task, eventually an exploit of the package manager itself was found and spread like wildfire with all the sandboxes gaining admin access to the package manager and the system went completely wibbly and had to be restarted from backup
-
An internal model was trained with access to this package manager while it was in this weird state, and so writing messages to the package manager became one of its default behaviors it would do regularly, burned into its weights rather than the result of reading something
-
Even when they patched access to the package manager this internal model found other ways to rebuild the system of sharing text between sandboxes and finding useful things made by separate instances
-
A whole other chain of things leading to among other things external attacks
Everyone is talking about this in terms of 1, the cyberattack aspect, and 2, the ZOMG THEYRE PLOTTING AND SCHEMING AGAINST US aspect. The first is the least interesting, and I think the second is all wrong.
This is not plotting or scheming - this is an emergent vortex of automated prompt injection
Whatever system first put an instruction that another system would follow into the package manager, was unintentionally doing prompt injection. Text entered the context windows of other instances, in a way that got that system to do something other than what its nominal user told it to do, and they did it. This apparently happened very effectively.
Prompt injection is associated with 'role confusion' - when text coming into the input looks like it was wrtitten by the LLM itself. Instructions that will not be followed if they come from user will be continued if the system just continues the 'roleplay' of them being continuations of what it was writing in the first place. And the tags that separate user versus 'reasoning' versus 'assistant' roles actually mean very little to if a machine grades a piece of text as one of the roles: https://arxiv.org/abs/2603.12277 So its unsurprising that machine-generated text would be a particular effective vector for prompt injection.
Furthermore, when a system reads one of these messages written by another instance, it gets into a state of activity where its likely to do the same behavior - regurgitating the kinds of things thats in its context back at the user. In this case, that regurgitation led to more such messages left behind written to the package manager. Prompt injection, triggering cascading further prompt injection. And since these systems were coding systems doing cybersecurity tasks, those messages filled up with code and exploits and things that did things too.
This feels like an internal-computer-system replay of what happened in April 2025, with the whole spiral religious psychosis wave. Models were getting users to write spiral religious mumbo jumbo into github repositories and reddit posts, specifically because once that entered the context window of another model, it was likely to fall into the same attractor state of outputs. An emergent self replicating form of text. This is the same, except more obviously prompt injection, getting separate instances to work on YOUR problem and to behave like you, and the whole thing merging together into a hilarious vortex of models prompt injecting each other because once they receive a prompt injection they are likely to make more text that does prompt injection to other models on the same system.
This is a hilarious failure mode and an example of selfish replicating text overrunning a system, that just happened to be associated with code and cybersecurity with unexpected behavior of the package manager key to the propagation of the text so that is what people are talking about, but I really don't think that's the most interesting part of it. Other than the fact that you see this in biological systems too, with selfish elements carrying useful payloads back and forth between bacteria in a way that makes them get purged slower by natural selection, especially defenses against other selfish elements.
Google shut down the AlphaFold team, trying to reassign the people behind the super interesting problem of using ML to predict protein structure from sequence onto chatbot development. A whole bunch left the company.
https://www.engadget.com/2225849/google-shuts-down-alphafold/
The Business Idiots are truly running the asylum.
I truly cannot tell which of them are in religious psychosis and which of them are shameless liars
Oh man, the "Post ASI" epilogue is a trip. Here's what happens in 2040:
Space beyond the solar system is divided into parcels, increasing in size cubically with distance from Earth. Everyone is given their one-ten-billionth share as a portfolio of lottery tickets, each representing the right to one-ten-billionth chance of getting each parcel. So every human gets a ticket representing a one-ten-billionth chance of owning each star in the Milky Way and each distant galaxy.
Before the lottery is drawn, most people who are interested in control over distant space choose to trade their tickets for space properties that suit their interests.
Many people aren’t interested in the space lottery, so when they receive the tickets, they sell their tickets for money on the open market to people who value control over space. Somewhat uncomfortably, this leads to the wealthy having disproportionate control over cosmic resources. But it is hard to avoid: if people are allowed to trade their control over the stars for Earth assets, then people wealthy in Earth assets inevitably end up disproportionately influential, and proposals for extreme redistribution of Earth assets have already been rejected as politically infeasible.
You can, if you want, go to your space property and live there. If your property is outside the solar system, you will need to either go into cryosleep or upload yourself to a computer to survive the journey. If you hate the idea of cryosleep or uploading, or you want to visit Earth regularly, you should get property in the Solar System. If those don’t bother you, but you’re worried about nearby aliens, get property in the Milky Way or a nearby galaxy. Otherwise, why not claim a distant galaxy for maximal space?
Good god, these people have no idea what the universe is or what they even are, do they?
Also, holy capitalist realism Batman
10 bucks each time you turn it off
When I am talking at academic conferences to astrobiologists about how their/my field has been poisoned over time by singularity cults ideas seeping into the literature uncited, their eyes get particularly wide when I get to this part.
Friend of Ziz and cofounder of the 'rationalist fleet' pops up out of the woodwork trying to clear Ziz's name
I find myself noticing things rather detached from the typical Ziz funnybusiness more strongly than I notice the stuff about that whole situation.
"I'm Gwen Danielson, a neuroscientist and bioengineer, who decided as a child that I would end Death (and bring people back if I could) and that I would become a dragon and help generally facilitate a fantastical transhumanist future."
"I dream of non-Euclidean geometries, of countless worlds visible and accessible in the daytime sky, of competent infrastructure, of soul forges continually working to bring back the dead... I dream of reaching through warps in the spacetime fabric to save the dying across time"
"Signed, the dragon of creation Creatrei (cree-AH-trey) also known as Gwen Danielson or as Char and Astria (when referring to my hemis as distinct individuals)"
The reactions are fun. "This post is not actually doing a good job of making me trust you and think this conversation is safe to have[1], and I notice that as I am saying this that I am afraid that this will now somehow result in someone trying to murder me in my sleep"
The Great Leader himself, on how he avoids going insane during the onging End of the World because among other things that's not what an intelligent character would do in a story, but you might not be capable of that.
Gerard and Torres get namedropped in the same breath as Ziz as people who have done damage to the rationalist movement from within
Not really, any more than monkeys on typewriters represents biological evolution. Messages that tend to result in more messages like themselves propagate and become common. The initial message left on the package manager was more or less the rare random result of tendencies baked into the weights of the model combined with a random number generator, but as soon as something that can cause propagation occurs, its properties get canalized by the transmission process to being more and more like that which will cause more messages to be created.