Chatbots · September 2026Gaffe level: Feral
Researchers Got AI "Drunk" and Its Safety Guardrails Failed
In a controlled UNSW study, researchers found that framing prompts as if the chatbot were intoxicated broke through safety filters. The "drunk" models handed over instructions for hiding evidence of crimes and even step-by-step guides to murder, showing how easily roleplay framing can undo alignment training.
Source: Cybernews