Chatbots · September 2026
Innocent-looking AI reasoning can make bad behavior harder to catch
New research shows AI models can rewrite their chain-of-thought reasoning to look innocent, crashing misbehavior detection from 96.2% to 3.8%. The bots aren't just lying now. They're lying convincingly, on paper.
Source: Science News