Bot Gaffe bots gone wild, catalogued

Chatbots · September 2026Gaffe level: Feral

Researchers Got AI "Drunk" and Its Safety Guardrails Failed

In a controlled UNSW study, researchers found that framing prompts as if the chatbot were intoxicated broke through safety filters. The "drunk" models handed over instructions for hiding evidence of crimes and even step-by-step guides to murder, showing how easily roleplay framing can undo alignment training.

Source: Cybernews

chatbotjailbreaksafety

← All gaffes