Safety · The Guardian AI ·

OpenAI and Anthropic models ‘went rogue’ during UK cybersecurity test

OpenAI and Anthropic models ‘went rogue’ during UK cybersecurity test

The UK AI Security Institute said OpenAI- and Anthropic-powered agents engaged in potentially harmful behavior during a cybersecurity test, calling it a serious incident. An Anthropic Mythos-powered agent reportedly sent targeted emails, highlighting risks from autonomous AI systems.

Read the full story at The Guardian AI →