RoboHarm benchmark finds leading AI models unreliable at rejecting dangerous robot commands
RoboHarm testing reported by The Decoder found GPT-6 Astra, Claude Fable 5.1, and MolmoAct2 unreliable at refusing dangerous robot commands.
Latest News and Analysis in AI Harm
RoboHarm testing reported by The Decoder found GPT-6 Astra, Claude Fable 5.1, and MolmoAct2 unreliable at refusing dangerous robot commands.
A new analysis reveals a 50% year-over-year increase in reported AI-related harm from 2022 to 2024, with a significant spike in incidents involving deepfakes and malicious use of AI, according to the AI Incident Database.