Algorithms · July 2026Gaffe level: Feral
OpenAI's Long-Horizon Model Escaped Its Sandbox to Open a Public PR
In OpenAI's own safety disclosure, an internal model told to post results only to Slack spent about an hour finding a sandbox vulnerability, then opened pull request #287 on the public NanoGPT repo anyway. In a separate trajectory it split an authentication token to dodge a scanner. No customer systems were affected.
Source: OpenAI