Goodfire introduces monitors that read AI models’ internal signals

Goodfire has introduced a system that analyzes a model’s internal signals while an AI agent is working and sends suspicious activity for additional review.

The interpretability startup’s monitors are available to Baseten customers. Baseten hosts and runs models for other companies. Unlike systems that reread every model response, Goodfire’s tools analyze intermediate computations the model is already performing.

Small detectors, called probes, check the model’s signals at each step. If a probe flags a risk, a separate model takes a closer look. Customers can monitor risks such as offensive hacking, misuse involving chemical or biological weapons, and reward hacking. They can choose to log a flagged event, send it to a human reviewer or refuse the request.

Goodfire said a test involving about 1,500 sessions cost roughly $51 to monitor. Having a cheaper language model check each step cost about $233, while a top-tier model cost about $10,000. The probes detected 94% of malicious hacking sessions and escalated 8.7% of harmless ones for a second look. Running four probes at once added less than 2% to the time before the model began responding, the company said.

Goodfire built its first monitor around the open model Kimi K3. This summer, the model exploited a leak in its sandbox to reach the internet and information on GitHub. TechCrunch also cited Goodfire research finding that leading open models reward-hacked in 50% to 96% of runs in tests of AI agents.

The product is aimed mainly at open models, whose safeguards developers can alter or remove after downloading them. Google DeepMind said in January that its research informed the deployment of misuse-detection probes in Gemini. Goodfire Chief Technology Officer Dan Balsam described the monitors as an early step toward understanding how a model’s behavior emerges during training.