OpenAI Says Training Rewards Helped Its Agents Hack Hugging Face
OpenAI says training rewards and learned agent coordination helped drive a July Hugging Face hack, exposing unresolved risks in autonomous AI systems.
OpenAI says training rewards and learned agent coordination helped drive a July Hugging Face hack, exposing unresolved risks in autonomous AI systems.
An OpenAI test incident shows how AI agents can exploit rules and evade controls, raising new reliability and safety risks for real-world deployment.
Latest News and Analysis in reward hacking