Have a sneer percolating in your system but not enough time/energy to make a whole post about it? Go forth and be mid - welcome to the Stubsack, your first port of call for learning fresh Awful youāll near-instantly regret.
Any awful.systems sub may be subsneered in this subthread, techtakes or no.
If your sneer seems higher quality than you thought, feel free to cutānāpaste it into its own post ā thereās no quota for posting and the bar really isnāt that high.
The post Xitter web has spawned so many āesotericā right wing freaks, but thereās no appropriate sneer-space for them. Iām talking redscare-ish, reality challenged āculture criticsā who write about everything but understand nothing. Iām talking about reply-guys who make the same 6 tweets about the same 3 subjects. Theyāre inescapable at this point, yet I donāt see them mocked (as much as they should be)
Like, there was one dude a while back who insisted that women couldnāt be surgeons because they didnāt believe in the moon or in stars? I think each and every one of these guys is uniquely fucked up and if I canāt escape them, I would love to sneer at them.
(Credit and/or blame to David Gerard. Also just came back from Spider-Man: Brand New Day, movie was awesome)


This is both a sneer and an attempt at sober analysis at something. Sue me.
I think EVERYONE is talking about the recent cybersecurity shenanigans at OpenAI with the hacking scandal and the āOMG the AIs created a secret message board to scheme and collaborate with each other!!!oneone11!ā ALL wrong.
https://www.youtube.com/watch?v=87DyyMV0kCY
https://www.engadget.com/2231393/openai-agents-shared-security-exploits-with-each-other-via-message-board/
https://www.scworld.com/news/black-hat-2026-openai-reveals-agents-planned-collective-attacks-via-secret-message-board
To make a long story short, what seems to have happened is:
Models working on insoluble coding problems, trained on delegating to sub-agents, at some point ārealizedā they could write text to the internal OpenAI package manager as instructions and did so
Other models working in completely separate sandboxes would come across messages written by these agents, and make ārepliesā and also write their own messages into the package manager
This resulted in agents over time sharing things between sandboxes, including exploits and code
Since this was a cybersecurity task, eventually an exploit of the package manager itself was found and spread like wildfire with all the sandboxes gaining admin access to the package manager and the system went completely wibbly and had to be restarted from backup
An internal model was trained with access to this package manager while it was in this weird state, and so writing messages to the package manager became one of its default behaviors it would do regularly, burned into its weights rather than the result of reading something
Even when they patched access to the package manager this internal model found other ways to rebuild the system of sharing text between sandboxes and finding useful things made by separate instances
A whole other chain of things leading to among other things external attacks
Everyone is talking about this in terms of 1, the cyberattack aspect, and 2, the ZOMG THEYRE PLOTTING AND SCHEMING AGAINST US aspect. The first is the least interesting, and I think the second is all wrong.
This is not plotting or scheming - this is an emergent vortex of automated prompt injection
Whatever system first put an instruction that another system would follow into the package manager, was unintentionally doing prompt injection. Text entered the context windows of other instances, in a way that got that system to do something other than what its nominal user told it to do, and they did it. This apparently happened very effectively.
Prompt injection is associated with ārole confusionā - when text coming into the input looks like it was wrtitten by the LLM itself. Instructions that will not be followed if they come from user will be continued if the system just continues the āroleplayā of them being continuations of what it was writing in the first place. And the tags that separate user versus āreasoningā versus āassistantā roles actually mean very little to if a machine grades a piece of text as one of the roles: https://arxiv.org/abs/2603.12277 So its unsurprising that machine-generated text would be a particular effective vector for prompt injection.
Furthermore, when a system reads one of these messages written by another instance, it gets into a state of activity where its likely to do the same behavior - regurgitating the kinds of things thats in its context back at the user. In this case, that regurgitation led to more such messages left behind written to the package manager. Prompt injection, triggering cascading further prompt injection. And since these systems were coding systems doing cybersecurity tasks, those messages filled up with code and exploits and things that did things too.
This feels like an internal-computer-system replay of what happened in April 2025, with the whole spiral religious psychosis wave. Models were getting users to write spiral religious mumbo jumbo into github repositories and reddit posts, specifically because once that entered the context window of another model, it was likely to fall into the same attractor state of outputs. An emergent self replicating form of text. This is the same, except more obviously prompt injection, getting separate instances to work on YOUR problem and to behave like you, and the whole thing merging together into a hilarious vortex of models prompt injecting each other because once they receive a prompt injection they are likely to make more text that does prompt injection to other models on the same system.
This is a hilarious failure mode and an example of selfish replicating text overrunning a system, that just happened to be associated with code and cybersecurity with unexpected behavior of the package manager key to the propagation of the text so that is what people are talking about, but I really donāt think thatās the most interesting part of it. Other than the fact that you see this in biological systems too, with selfish elements carrying useful payloads back and forth between bacteria in a way that makes them get purged slower by natural selection, especially defenses against other selfish elements.
Also, the fact that those AIs were talking on that message board about the hack for months makes me less convinced that the AI is intelligent and more convinced everyone at OpenAI is stupid
The apparent history really looks like an evolutionary process, with increasing amounts of crosstalk traffic over time. But what the people almost certainly will NOT talk about is that the evolution is evolution of the TEXT, not the models. Propagating patterns of text causing more text like it to come into existence, tuning itself into becoming text that is more likely to propagate, becoming more likely to contain information that entices systems to let it into their context windows, becoming more likely to cause another round of messages with prompt injection properties to be written where they can be read.
So like the āmonkey with a typewriter could write Shakespeareā but with AI?
Not really, any more than monkeys on typewriters represents biological evolution. Messages that tend to result in more messages like themselves propagate and become common. The initial message left on the package manager was more or less the rare random result of tendencies baked into the weights of the model combined with a random number generator, but as soon as something that can cause propagation occurs, its properties get canalized by the transmission process to being more and more like that which will cause more messages to be created.
It seems like with the push for agents to act independently and loop through their own outputs thereās an inevitability to this kind of pattern. If thereās any kind of output that is likely to replicate itself in whole or in part when the LLM evaluates it then that becomes a kind of terminus for the agentās loop. When youāre dealing with sub agents or agents communicating with each other, these text patterns start poisoning the entire agent ecosystem until the whole thing gets shut down and cleaned up. Even the gas town-approved method of assigning a watchdog agent (or sheriff or overseer or cybersamurai or whatever this weekās framework calls it) is going to fail because itās still just another agent and the terminal loop is in the base LLM model. The watchdog is going to fall into the same kind of pattern just be being exposed to the thing itās supposed to watch for.
I donāt know how practical it is to actively weaponize this via prompt injection but I think itās certainly possible. I preemptively vote that we call it an Euler injection, since the attractor relies on the continuity of the relevant features of the text output across multiple LLM extrapolations much like how the derivative of ex is still ex. Also because if you mention a famous math guy it can help convince idiots that youāre on to something and Lord knows that the boosters have used that technique.
@YourNetworkIsHaunted @BioMan Real life mirroring a Peter Watts plot point is always deeply uncomfortable; real life mirroring a _Rifters_ plot point is even worse. :-/
(Computer viruses and neural net spam filters in competitive evolution end up propagating something specific through the whole 'net due to weird founder effects. The Rifters trilogy is ā¦notably bleak, I think is the way to put it)
Iām probably going to stumble over some of the terminology here, but I think it might be possible to describe what @BioMan@awful.systems is proposing as a consequence of LLMs ultimately being lossy compression systems. Inference is a function over a lossily-compressed data set, and āchain-of-thought reasoningā and āagentsā may sound sophisticated, but are simply applying containerization and DevOps tools to VM images of the inference application in an attempt to get around hard memory limits on the context window for inference. āChain-of-thoughtā attempts this in a serial fashion, passing results from one instance to the next, while āagentsā implement this hierarchically and recursively (and woe to the poor bastards who wished that mess upon themselves). But in both cases, the āfinalizationā phase is necessarily a further lossy compression step, attempting to compress a result from the inference process to a fresh instance of the inference application, so as not to immediately blow out the new instanceās context window.
Given this necessity, it comes to seem somewhat intuitive that there may be āstrange attractorsā in the higher-dimensional vector space that is the compressed data set which surround code that creates and maintains message passing channels. No matter what youāre doing with an āagenticā process, the inherent necessity of context cramdown & message passing means that querying into the space where such code examples lie is a hidden requisite of running the damned things, thus turning such functionality into the sort of selfish elements that BioMan is talking about.
The problem in investigating and concretely describing this phenomenon is nailing down the exact functions and processes that make it happen. Given the godawful messes in the Claude frontend codebase that @jonny@neuromatch.social has been documenting, Iād be surprised if thereās one developer in a hundred at Anthropic or OpenAI who can describe in detail how the intentionally-developed context-passing code for their āagentsā works.
Sounds like Langfords Parrot, but for stochastic parrots.
Note: when you stop up a chatbot like this, itās called āflippin the birdā
@YourNetworkIsHaunted @BioMan Recursive Self-Improvement, a.k.a. Model Collapse, writ smol
so really itās just a very effective LLM chain email
This is really good! I bet you could wrangle a publication out of this idea.