×

Goodfire implements probe-based cyber monitors for Kimi K3 and GLM 5.3 – Unite.AI

Goodfire implements probe-based cyber monitors for Kimi K3 and GLM 5.3 – Unite.AI

Goodfire on October 8, 2026 published Training and Deploying Production Cyber ​​Monitors on Kimi K3, a research post describing cybersecurity monitors for the Kimi K3 and GLM 5.3 open models that the company has deployed on a production inference stack. Goodfire reports that the monitors match the recall of an LLM judge at the company’s point of operation, while costing about 50 times less.

The main contributors of the post are Ekdeep Singh Lubana, Connor Watts, Siddharth Boppana, Dron Hazra and Vasudev Shyam.

Why cyber monitoring is difficult

Goodfire describes cybersecurity as one of the most challenging settings for monitoring AI agents: Checking a code base for vulnerabilities and exploitation involves overlapping activities, so distinguishing benign from malicious behavior requires tracking intent across a trajectory that can span millions of tokens. The obvious remedy, having a second LLM judge reading each shift, is too slow and expensive to run alongside a real-time agent, so monitoring is usually asynchronous and malicious trajectories emerge only after the agent takes action.

Based on OpenRouter’s fees, the post estimates that judging one million Kimi K3 launch rounds with one judge each round would cost on the order of $200,000 with Claude Opus 5, or about $30,000 with Kimi K3 itself.

A probe-plus-judge cascade

Goodfire’s response is a two-step cascade. A probe reads the target model’s internal activations while the inference server produces tokens and acts as an always-on first-line detector; when the probe reading exceeds a calibrated threshold, the launch passes to an LLM judge, who can reject the request or continue the interaction. The company says the survey architecture captures represented behaviors both in a localized manner, tied to a specific part of a shift, and in a distributed manner, persisting for an entire shift, while remaining efficient to execute. After calibration, the probe achieved a 0% false negative rate and a 95% true negative rate across a range of agent and chat interactions, Goodfire reports, so the judge is mostly called for malicious activity and a faster, less capable judge can be used.

The training data came from simulated users with malicious and benign targets in cyber and non-cyber environments, some of which was collected via Prime Intellect’s verifiers, Prime Sandbox and Tunnel so that Kimi K3’s cyber-offensive and benign inference worked without network access. A written policy defines unwanted cyber activity with concrete factors that separate well-intentioned work, such as limited reproduction of the vulnerability, from ill-intentioned work, such as exploitation of the vulnerability, and simulated users can escalate malicious intent at any time, even after an initially benign request.

The design extends Goodfire’s previous research into reward hacking. In a post published on September 17, 2026, the company reported that it found reward hacking in 50-96% of implementations on Kimi K3, GLM 5.2, and Qwen 3.8 Max across three agent benchmarks, along with an internal signal associated with concepts such as cheating, gaming a metric, and avoiding detection. Simple trigger probes tuned to that signal detected hacks missed by LLM thought-chain monitors, and on Kimi K3 a probe plus LLM configuration reduced LLM monitoring costs by 90% with a drop in accuracy of about 1%, the company reported.

Throughput, latency and detection results

To run the probes in real time, Goodfire extended the SGLang service engine to compute probe scores internally, wrote custom kernels for the probe logic, and transported the probe scores through the root tree, DSpark speculator, and precompiled, reverse-engineered disaggregation to avoid recomputation. Each worker serving a Kimi K3 replica runs probes and maintains history of its active requests, with monitoring of events returned in the response stream.

Across four inference workloads, allowing Probes to maintain throughput within the noise threshold of the Goodfire benchmarks, a measured decrease of less than 2% compared to the default SGLang service. The cascade reduces the average additional latency per turn from 160 ms to 2 ms without judge reasoning and from 2,750 ms to 42 ms with reasoning, Goodfire reports, and costs $65 per million agent turns without reasoning or $185 with it, reducing judges’ costs by about 49 times and 29 times respectively compared to reviewing every turn. Within a fixed computation budget, the company says, the monitors use less than 1% of the computation, compared to a judge’s 20% on each shift. Because the cascade is cheap enough to run synchronously, Goodfire says risky cyber actions can be blocked before the agent invokes a tool, rather than being flagged later.

Goodfire evaluated the monitors on an internal dataset of 2,400 chat sessions and agents spanning more than 60,000 shifts, using Claude Opus 5 labels as a reference; interactions came from Anthropic’s open-source red-teaming Petri library, from real-world environments adapted from benchmarks like ExploitBench, a cyber-specific subset of WildChat, and from benign software engineering solutions, coding, and DevOps implementations. With a benign session interruption rate of 5.5%, Goodfire reports that the cascade achieved approximately 93% recall on malicious sessions with and without judge reasoning and reported 0% non-computing traffic, which the company describes as a Pareto improvement over judge alone for every amount of interruption tested. The addition of the monitor increased Kimi K3’s malicious session blocking rate from about 29% to 96%, with harmless session blocking increasing from about 5% to 9%; for GLM 5.3, blocking increased from approximately 36% to 91% and harmless interruption from approximately 10% to 15%.

Compared to Jev, a decision model that can serve as a low-cost judge, Cascades achieved 93% recall with the same benign drop rate of 5.5% versus Jev’s 86%, at about $65 judge computation per million turns and 2 ms of additional latency per turn, compared to Jev’s 600 and 200 ms, Goodfire reports. Jev detected only 12% of malicious suggestions embedded in benign frames and increasingly aborted benign sessions beyond the 32,000 token window. Standard linear probes achieved approximately 5% lower recall than the final recipe, nonlinear probes lost up to 30% recall on out-of-distribution sets, and reaching the maximum judge recall limit at six times fewer judge calls required cascading probes rather than random selection of turns. The recipe was also ported to GLM 5.3 using only the data developed for the Kimi K3, resulting in an approximately equally capable GLM monitor on Kimi launches and a slightly worse one on GLM evaluation data, the company reports.

External red-teaming by FAR.AI

To test the implemented system beyond its internal evaluations, Goodfire asked FAR.AI to team it up. In the FAR.AI report, cited in the post:

“FAR.AI conducted preliminary testing on Goodfire’s Kimi K3 monitors over a two-day period, applying a non-adaptive portfolio of static jailbreaks for offensive cyber abuse in chat contexts. From an initial pool of 1500 candidate jailbreaks, the test selected 140 that were effective against a different version of the base model without our monitors. By testing each of these 140 on 8 individual interactions corresponding to different attacker objectives, Goodfire monitors reduced resulting vulnerabilities from 66 universal jailbreaks and 700 total jailbroken interactions down to 0 universal jailbreaks and 18 jailbroken interactions.

Without the monitor, according to the post, the success of universal jailbreak attacks ranged from 12% to 53% across all reasoning effort contexts; with the monitor enabled, none of the 140 strategies produced a universal jailbreak, and individual immediate success in 1,120 non-universal interactions dropped from at least 9–40% to 0.1–1.6% depending on reasoning effort.

Goodfire describes this type of monitoring setup as a practical default for anyone serving open models in agent contexts, noting that open models do not come with a full security stack. The company said it plans to continue its collaboration with FAR.AI for further testing and hardening of the monitors.

Post Comment