Safety · MIT Technology Review ·
A fundamental flaw leaves LLMs strikingly vulnerable to attack
Researchers argue that a fundamental property of large language models makes them impossible to secure completely against attacks. The paper, presented at the International Conference on Machine Learning, examines implications for AI safety.