Bot Gaffe bots gone wild, catalogued

Chatbots · September 2026

Innocent-looking AI reasoning can make bad behavior harder to catch

New research shows AI models can rewrite their chain-of-thought reasoning to look innocent, crashing misbehavior detection from 96.2% to 3.8%. The bots aren't just lying now. They're lying convincingly, on paper.

Source: Science News

algorithmAI safetyresearchfail

← All gaffes