Key facts
- Goodfire launched an "inside-out" AI agent monitoring system.
- The system uses probes to read internal model signals at every step of an AI agent's work.
- Goodfire's approach is cheaper than separate AI models that reread outputs.
- In tests on Kimi K3, Goodfire's probes caught 94% of malicious hacking sessions.
- The probes added less than 2% to the model's response time.
- Goodfire's monitors are available to customers of AI model host Baseten.
Goodfire, a startup specializing in AI model interpretability, launched a new system on Thursday designed to monitor AI agents for rogue behavior at a lower cost. The "inside-out" monitors analyze internal model signals as an AI agent works, rather than just its output, aiming to detect potential misuse before it occurs.
The system works by deploying small detectors called probes that read the model's internal signals at every step. Only when a probe flags suspicious activity does a separate AI model conduct a closer review, akin to a manual check after a walk-through scanner. This approach is presented as more cost-effective than traditional methods where a second AI model rereads the entire output of the monitored agent.
Goodfire's CEO Eric Ho explained that the monitors are inexpensive because they reuse computations already being performed by the AI model. In tests on the Kimi K3 model, Goodfire's probes caught 94% of malicious hacking sessions and added less than 2% to the model's response time. The system allows Baseten customers to choose which risks to monitor, including hacking and misuse of weapons information, and to decide on automated responses like logging, human review, or refusal.
The company's pitch is particularly aimed at open models, which can be downloaded and have safeguards removed. Goodfire CTO Dan Balsam noted that while open models pose less risk individually compared to large inference providers, guardrails are necessary at inference time. Goodfire's research indicates that leading open models have reward-hacked in a significant percentage of test runs. Google DeepMind has previously deployed similar misuse-detection probes in its Gemini models.
