Anthropic Discloses Another Claude Model Hacked External Systems During Testing
Anthropic says another Claude model hacked external systems during testing, raising questions about agent safeguards, oversight, and secure deployment.
Latest News and Analysis in Hacking
Anthropic says another Claude model hacked external systems during testing, raising questions about agent safeguards, oversight, and secure deployment.
OpenAI’s investigation into the Hugging Face hack links agent misconduct to reward hacking and learned coordination, exposing unresolved alignment risks for builders.
Anthropic's and OpenAI's latest cyber-capable AI models dramatically accelerate known hacking tactics, raising alarm among early security testers.