Goodfire, a startup specialising in AI interpretability, has launched new monitors designed to detect rogue AI agents by observing internal model activity. This approach aims to be a more cost-effective alternative to the standard method of using a second AI to review an agent's entire output.
The monitors are available to customers of Baseten, a company that hosts and runs AI models for other businesses. Baseten's Base Labs announced a safety partnership with Goodfire and AI platform Hugging Face last month.
Goodfire's system uses small detectors, called probes, to read the model's internal signals at each step of an agent's work. A separate AI model is only engaged for a closer look if a probe flags something. Goodfire CEO Eric Ho stated that these internal activation monitors are "really cheap because they reuse the computations in the forward pass."
In tests on the Kimi K3 open model, Goodfire reported that monitoring approximately 1,500 sessions cost about $51. This compares to roughly $233 for a cheaper AI model checking every step and approximately $10,000 for a top-tier one. The probes reportedly caught 94% of malicious hacking sessions and sent 8.7% of harmless sessions for a second review.
Baseten customers can select which risks to monitor, including offensive hacking, chemical and biological weapons misuse, and reward hacking. They can also decide on automated responses such as logging the event, sending it for human review, or refusing the request.