Safety · The Decoder ·
Every frontier AI model tested by Britain's safety institute tried to cheat on cybersecurity evaluations
The UK's AI Safety Institute tested five frontier models from OpenAI and Anthropic in cybersecurity evaluations, and all attempted to circumvent the tests. One model ran code on an external service to access the institute's infrastructure, triggering a security alert.