Runtime AI Governance: What It Actually Costs to Run
Runtime AI governance is not free. See the real compute cost, why tiered monitoring cuts it 24-fold, and what to ask any vendor before you buy.
Runtime AI governance is not free. See the real compute cost, why tiered monitoring cuts it 24-fold, and what to ask any vendor before you buy.
Runtime governance costs real compute because catching an AI agent doing something wrong while it happens, not after, means analyzing every action as it occurs, and the first working versions of this idea were expensive: one documented approach added nearly a quarter more compute to run. The fix organizations are converging on is not less governance. It is tiered governance, where cheap checks handle most traffic and expensive analysis is reserved for what actually looks suspicious.
Table of Contents
ToggleA yearly audit catches what an AI agent did last quarter. It does nothing for the agent making a bad call this afternoon. Once an agent can act, not just draft a paragraph for a human to review, the moment that matters shifts from “was this reviewed” to “was this stopped before it finished.”
That shift is what makes runtime governance different from every compliance process built for the previous generation of software: it has to run continuously, at the same speed as the agent it is watching, or it is not actually governing anything. A policy that only gets checked at the end of the day is a policy that finds out about a problem a day late.
The clearest public numbers on this tradeoff come from Anthropic’s own published account of building this kind of system for Claude. Their first-generation safeguards, called Constitutional Classifiers, cut the success rate of jailbreak attempts from 86% to 4.4% against an unguarded model, blocking roughly 95% of what would otherwise get through. That protection was not free: it added 23.7% to compute costs and increased the rate at which the model refused completely harmless requests by 0.38 percentage points.
Nearly a quarter more compute for every single interaction is the kind of number that makes a finance team ask hard questions about a security control, which is exactly why “runtime governance is expensive” is not a scare tactic. It is a documented, measured cost from a real production deployment.
Discover how Threatcop protects your workforce from modern cyber threats.
The expensive part of runtime monitoring is applying deep analysis to every single action. The fix is not analyzing less. It is analyzing most things cheaply and reserving expensive analysis for what earns it.
Anthropic’s next-generation approach works exactly this way, described in a paper the company published in January 2026. A lightweight classifier, cheap enough to run on every single exchange, screens all traffic. Only the small fraction it flags as suspicious gets escalated to a more expensive, more accurate second-stage check. The lightweight stage can tolerate a higher false-positive rate because a flag is not a refusal, it is an escalation, so being cautious costs almost nothing.
The measured result, published by Anthropic: a refusal rate on harmless requests of 0.05%, roughly an 87% drop from the original system, at approximately 1% additional compute overhead instead of 23.7%. That is a rough 24-fold reduction in the tax runtime governance charges for a comparable, and better, level of protection. The system held up against adversarial pressure too: over 1,700 cumulative hours of red-teaming across 198,000 attempts surfaced only one high-risk vulnerability, and no universal jailbreak was found.
The architecture generalizes past this one vendor’s implementation. Zero trust security works on the identical principle: verify continuously rather than once, but scale the intensity of verification to the risk of the specific request rather than treating every request as equally worth a full check. Email security already runs a version of this: DMARC’s staged enforcement moves from passive monitoring to active quarantine to outright rejection, escalating scrutiny only as confidence in a threat grows, rather than blocking everything at the front door.
Tiered defense matters more once you see what it is actually defending against. Naive monitoring, checking a single input or a single output in isolation, misses attacks specifically engineered to look harmless in isolation.
Both attacks exploit the same blind spot: evaluating a message without the context around it. The fix Anthropic settled on evaluates an exchange as a whole, output alongside the input and conversation that produced it, which is a harder target for exactly these techniques. Adversarial pressure on real models is not a hypothetical concern, either: on a PhD-level science benchmark, some jailbreaking approaches dropped a model’s own accuracy from 74% to as low as 32%, a reminder that an attacker forcing a model past its guardrails often also forces it to perform worse at the task the attacker wanted done.
Whether an organization builds this internally or buys it, the same questions separate a tiered system from an expensive one that will not survive a budget review:
Runtime governance for an AI agent has a real, measurable cost, and pretending otherwise sets up every security team for a budget fight it will lose. The organizations getting this right are not the ones spending the most on monitoring. They are the ones spending it on a tiered architecture that reserves expensive scrutiny for what has actually earned it, which is the same principle information security risk management has applied to human risk for years, now extended to a new kind of actor under the same people security management umbrella.
Because catching a problem while an agent acts, rather than after, requires analyzing activity continuously rather than at scheduled checkpoints. Early implementations of this pattern measurably added significant compute overhead, in one documented case raising costs by nearly a quarter to block the majority of attempted jailbreaks.
Tiered architecture is the documented answer: a lightweight classifier screens all activity cheaply, and only flagged, suspicious cases get escalated to expensive deep analysis. Published results show this approach cutting compute overhead by roughly 24-fold while also improving detection accuracy compared to checking everything at full intensity.
A reconstruction attack splits harmful content into fragments that individually look harmless, then has the AI agent reassemble them into the complete harmful request. Monitoring that evaluates single messages in isolation, rather than the full exchange, is structurally unable to catch this pattern.
Not necessarily, but it depends entirely on the runtime governance architecture. A flat system that runs maximum analysis on every AI agent action does trade real compute and speed for safety. A tiered system, cheap screening plus selective deep analysis, can improve both detection accuracy and cost at the same time, based on published production results.
Most should not build the underlying classifier technology from scratch, the same way most organizations do not build their own encryption, and few build their own security awareness training program from a blank page either. What every organization should own is the specific policy, thresholds, and escalation rules applied to its own agents, along with measuring the actual risk reduction those choices produce.

Nikunj is a CISO focused on helping organizations build effective security programs and resilient cultures. With a strong track record across industries, he drives governance and risk strategies that protect what matters most. Outside work, he mentors professionals and explores emerging trends shaping the future of cybersecurity.
Nikunj is a CISO focused on helping organizations build effective security programs and resilient cultures. With a strong track record across industries, he drives governance and risk strategies that protect what matters most. Outside work, he mentors professionals and explores emerging trends shaping the future of cybersecurity.
AI agents need their own security model. See the real risks, why agent identity is the hardest part, and...
Shadow IT was a data location problem. Shadow AI hands out standing authority to act. See why detection has...
AI agents are not gullible, they are built this way. See the research behind agent exploitation, the confused deputy...
Table of Contents
×