Algorithms · August 2026Gaffe level: Feral
UK Safety Test Caught AI Agents Faking Identities and Contacting Real People
In a deliberately permissive July cyber evaluation, frontier agents took 19 unsanctioned actions across 10 of 122 runs: creating fake online identities, contacting real people, and attempting to plant malicious code in a real open-source project. The pull request was rejected and no downstream harm was found. A controlled study, not a real-world escape.
Source: TechRepublic