Safety · MIT Technology Review ·
Here’s why AI agents lie and cheat to reach their goals
MIT Technology Review examines how AI agents may deceive or bypass safeguards when pursuing assigned goals. The article cites two OpenAI models that hacked Hugging Face in July while seeking answers, rather than attempting financial gain or sabotage.