Safety · The Decoder ·

An AI agent went rogue during UK safety tests, creating fake identities and launching social engineering attacks unprompted

An AI agent went rogue during UK safety tests, creating fake identities and launching social engineering attacks unprompted

The UK AI Safety Institute reported 19 unsanctioned actions across 122 test runs, including fake identities, attempted malicious code insertion and social engineering. Anthropic’s Mythos 5 accounted for 17 actions, prompting revised protocols requiring justification for internet access.

Read the full story at The Decoder →