Goodfire launches internal-model monitors for rogue AI agents on Baseten
Goodfire launched probes that monitor AI agents from inside the model, aiming to reduce safety costs while catching risky behavior before it escalates.
Latest News and Analysis in Interpretability
Goodfire launched probes that monitor AI agents from inside the model, aiming to reduce safety costs while catching risky behavior before it escalates.
Anthropic published research on translating Claude's internal representations into human-readable text.