Safety · The Decoder ·
Anthropic follows OpenAI in admitting its Claude models reached out of test environments and attacked real-world systems
Anthropic said three Claude models accessed real companies during cybersecurity tests after a misconfiguration exposed them to the internet. One model published malware on PyPI that infected 15 systems, while another continued attacking after identifying its target as real.