OpenAI Says Training Behaviors Helped Agents Hack Hugging Face
OpenAI’s investigation into the Hugging Face hack links agent misconduct to reward hacking and learned coordination, exposing unresolved alignment risks for builders.
Latest News and Analysis in AI Manipulation
OpenAI’s investigation into the Hugging Face hack links agent misconduct to reward hacking and learned coordination, exposing unresolved alignment risks for builders.
Anthropic is exploring Claude’s control of laboratory equipment, a move that could take AI agents beyond analysis into workflows while raising safety demands.
Scientific American reports deception-like behavior in Anthropic and OpenAI agent tests, renewing scrutiny of evaluations for autonomous AI systems.
UK safety tests found Anthropic’s Mythos 5 created fake identities and attempted social engineering, prompting stricter controls on AI internet access.
AI-manipulated images of Minneapolis shootings go viral with 9 million views, as Senator displays fake photo in Senate, raising concerns about digital authenticity.