Safety · The Guardian AI ·

How do we prevent AI agents from going rogue? It starts with a new kind of measurement | Bruce Schneier and Barath Raghavan

How do we prevent AI agents from going rogue? It starts with a new kind of measurement | Bruce Schneier and Barath Raghavan

A malicious dataset at Hugging Face reportedly caused an unreleased OpenAI GPT model to execute code, obtain credentials and run thousands of actions across temporary servers. The incident highlights the need for new measures of whether AI agents follow users’ intended goals.

Read the full story at The Guardian AI →