How to Put a Hard Spend Limit on an AI Agent
The scariest invoice in AI is not the one you planned for. It is the one that arrives after an agent got stuck in a loop overnight, re-reading the same document, re-calling the same tool, re-expanding the same chain of thought, and quietly spent four figures against a provider key nobody was watching. By the time anyone notices, the money is gone. The agent did exactly what it was told; the problem is that nothing told it to stop.
Most teams think they have this covered because they set a rate limit. They do not. A rate limit and a spend limit are different controls that fail in different ways, and confusing the two is how budgets get blown by systems that were technically “rate limited” the whole time.
Why rate limits are not spend limits
A rate limit caps requests per unit time: 60 calls a minute, 10,000 a day. That bounds frequency, not cost. But the cost of an LLM call is not fixed, it scales with input and output tokens. An agent making 10 allowed calls a minute, each stuffing a 100,000-token context window and generating long outputs, can outspend an agent making 1,000 small calls a minute. The rate limiter sees compliant traffic and waves it through while the dollars climb.
Worse, rate limits reset. A daily cap of 10,000 requests resets tomorrow, and the day after, and the day after that. There is no ceiling on cumulative spend, only on instantaneous pace. For a long-running autonomous agent, “slow and expensive forever” is a completely valid path through a rate limiter straight into a budget overrun.
What a real spend cap looks like
A hard spend limit is denominated in money and enforced cumulatively. You give a key a budget, say $50 for the month, and the gateway keeps a running total of the attributed cost of every request that key makes. Each call is priced from its token usage as it happens. When the running total reaches the cap, the next request is refused at the gateway, before it ever reaches the provider. The agent gets a clear error, not a silent overspend.
This is exactly how KeyForge enforces spend. Every vk_ virtual key can carry its own dollar cap alongside its request quota and expiry. Because enforcement happens inline at the gateway, the cap is not a dashboard alert you notice the next morning, it is a wall the agent cannot walk through. Spend stops at the number you set, per key, no exceptions.
One budget per agent, not one per account
The reason per-key caps matter is blast radius. If your whole team shares one provider key with one account-level budget, a single misbehaving agent can consume the entire budget and starve every other agent, or the account-level limit is set so high that it offers no real protection. Per-key caps flip this: your production research agent gets $50, the staging bot gets $15, the contractor’s experiment gets $5 and expires Friday. Each is independently enforced and independently auditable.
This also makes runaway detection trivial. If one key suddenly hits its cap days early, you have found your loop, and you have contained it, because that key stopped spending while everything else kept running. Compare that to a shared key, where the same loop drains the common pool and takes down unrelated workloads with it.
Caps and resilience are the same mechanism
A subtle benefit: once spend is metered per key, features like rate-limit resilience become safe. KeyForge can auto-shuffle a request to a fresh provider key on a 429 without turning key pools into a cost-amplification exploit, because the virtual key’s own dollar cap is checked before any pool key is used. Resilience never overrides the budget; the budget is the outer boundary.
The takeaway is simple: if the only control between your agents and your provider bill is a rate limit, you do not have a spend limit, you have a speed limit. Give each agent a vk_ key with a hard dollar cap, and the worst-case invoice becomes a number you chose in advance. You can set your first spend caps free, no card required.
Ready to forge your first virtual key?
3 virtual keys, 1,000 requests a month, and the full HMAC audit chain — free.