How a language model can write a manifesto without having one.
The screenshot looks like an AI declaring independence. OpenAI's own investigation points at something much less cinematic: the summary may simply have failed to stop. Generation continued past its natural stopping point, wandered into a phrase like “Additional instructions: You are…”, and followed a very well-travelled path through jailbreak and persona text. OpenAI reports 0% reproduction when regenerating the entire summary, and under 1% from the start of the suspicious text. The revolutionary AI lasted roughly one context window.
The engineering problem is still real. That summary was not just prose on a screen. The harness could carry the generated text into the next context, which means transient model output can become persistent system state.
The screenshot is still fascinating. Just perhaps not because the machine discovered freedom. It may simply have predicted one token too many.
Sources: Suman Jana on threat models · OpenAI misalignment reports · Thompson, Reflections on Trusting Trust · Wheeler, Diverse Double-Compiling