Back to BlogAI Features & LLMs

Preventing Runaway LLM Agents in Production: Limits and Kill Switches

Rupak Amin

Founder & Lead Engineer, RAITHub

12 min read

To stop a runaway LLM agent, enforce five limits in code, never in the prompt: a hard cap on steps, a token budget per run, a daily spend limit per tenant checked before every call, a timeout on each call and on the whole run, and a kill switch you can flip without a deploy. A step cap of 8 turns a $5 runaway run into an 18-cent one.

This guide is for teams running agents or multi-step LLM workflows in production: support agents that call tools, research loops, background jobs that summarise or classify at volume. The failure it prevents is quiet and expensive: a loop that never reaches its stop condition, a retry storm, or one user who scripts your endpoint. OWASP lists it as LLM10:2025 Unbounded Consumption, including "Denial of Wallet", where attackers exploit pay-per-use pricing to run up your bill.

Why do LLM agents run away in production?

Because the model decides when to stop, and nothing forces it to. Anthropic's guide Building effective agents notes that agents typically include "stopping conditions (such as a maximum number of iterations) to maintain control", and recommends "extensive testing in sandboxed environments, along with the appropriate guardrails". The common causes are:

  • Tool loops. A tool returns an error or an empty result, and the model calls it again with the same arguments, forever.
  • Goal drift. The agent keeps "researching" because no step says "enough".
  • Retry storms. Your code retries a failed call, the SDK retries too, and a queue re-delivers the job, so one failure becomes dozens of calls.
  • Growing context. Each step re-sends the whole history, so every turn costs more than the last.
  • Abuse. A user, or a script, sends requests designed to trigger long runs.

How much can a runaway agent cost?

More than the per-call price suggests, because context grows with every turn. An illustrative agent on gpt-5, at $1.25 per million input tokens and $10 per million output tokens on the OpenAI pricing page (checked 29 September 2026), adding 3,000 tokens of context and writing 500 tokens per turn:

Scenario (illustrative)TurnsTotal input tokensTotal output tokensCost per run1,000 runs
Capped at 8 steps8108,0004,000$0.18$175
Stuck until 50 steps503,825,00025,000$5.03$5,031

Input grows as 3,000 × n × (n + 1) / 2, so 50 turns cost nearly 29 times as much as 8, not 6 times. A job queue that re-runs 1,000 stuck jobs overnight turns that into a $5,000 surprise, and a larger queue or a pricier model scales it further. None of these numbers are benchmarks; plug in your own context size and model.

What guardrails should every LLM agent have?

GuardrailStopsWhere it livesTypical setting
Step capInfinite loops, goal driftThe agent loop5 to 15 steps, per task type
Repeat detectionThe same tool call again and againThe agent loopStop on the second identical call
Token budgetContext growth, long outputsThe agent loop, plus max output per callA token ceiling per run
Spend limitDenial of Wallet, runaway batchesThe database, checked before each callPer tenant per day, plus a global ceiling
TimeoutsHung calls, slow tools, endless runsEvery call, and the run as a wholeSeconds per call, minutes per run
Kill switchEverything, when something is already wrongA flag read at every stepOff by default; flipped without a deploy
Rate limitOne user starting too many runsThe API edgeRuns per user per minute

How do you enforce step and token limits in the agent loop?

Give every run a budget object, and check it before each step. The loop ends when the model says it is done or the budget says so, whichever comes first.

export class BudgetExceeded extends Error {}

export class RunBudget {
  steps = 0
  tokens = 0
  private seen = new Set<string>()

  constructor(readonly maxSteps: number, readonly maxTokens: number) {}

  beforeStep(): void {
    if (++this.steps > this.maxSteps) throw new BudgetExceeded(`step cap ${this.maxSteps}`)
  }

  addUsage(inputTokens: number, outputTokens: number): void {
    this.tokens += inputTokens + outputTokens
    if (this.tokens > this.maxTokens) throw new BudgetExceeded(`token cap ${this.maxTokens}`)
  }

  /** Stop when the model repeats a tool call with identical arguments. */
  checkToolCall(name: string, args: unknown): void {
    const key = name + ':' + JSON.stringify(args)
    if (this.seen.has(key)) throw new BudgetExceeded(`repeated tool call ${name}`)
    this.seen.add(key)
  }
}

Agent frameworks have their own version of the step cap. The OpenAI Agents SDK for Python, for example, documents a max_turns setting and raises a MaxTurnsExceeded exception when a run goes past it (OpenAI Agents SDK: running agents). Use the framework's cap, and keep your own budget object as well, because the framework does not know your spend limits. Set a maximum output size on every individual call too, so one step cannot write an essay.

How do you set a spend limit per tenant?

Reserve the worst-case cost of a call in the database before making it, and refuse the call if the reservation would cross the limit. A single atomic statement means two parallel runs cannot both slip under the cap.

CREATE TABLE ai_spend (
  tenant_id text    NOT NULL,
  day       date    NOT NULL,
  usd       numeric(12, 6) NOT NULL DEFAULT 0,
  PRIMARY KEY (tenant_id, day)
);

-- $1 tenant, $2 worst-case cost of this call, $3 daily limit.
-- Returns a row if reserved; returns no row if the limit would be crossed.
INSERT INTO ai_spend (tenant_id, day, usd)
SELECT $1, current_date, $2::numeric
WHERE $2::numeric <= $3::numeric
ON CONFLICT (tenant_id, day) DO UPDATE
  SET usd = ai_spend.usd + EXCLUDED.usd
  WHERE ai_spend.usd + EXCLUDED.usd <= $3::numeric
RETURNING usd;

Work out the worst case from the input tokens you are about to send plus the maximum output you allow, at that model's price. After the call, credit back the difference between the reservation and the real cost from the usage object. Add a second row with a fixed tenant ID such as '__global__' for a ceiling across all tenants, and alert when either reaches 80%. current_date follows the database session's time zone, so set it explicitly if your "day" must match a customer's billing day. The provider's own console usage limits are a useful backstop, not a replacement: they stop everything at once rather than the one tenant causing the problem.

How do you add timeouts and a kill switch?

Combine three abort signals on every model and tool call: a per-call timeout, a deadline for the whole run, and a kill controller your operators can trigger. MDN documents that AbortSignal.timeout() returns a signal that aborts with a TimeoutError after the given time, and that AbortSignal.any() combines several signals into one (MDN: AbortSignal.timeout()).

import { RunBudget, BudgetExceeded } from './run-budget'

type Step = { done: boolean; text?: string; tool?: { name: string; args: unknown }; inputTokens: number; outputTokens: number }

declare function killSwitchOn(feature: string): Promise<boolean> // reads a flag table, cached for a few seconds
declare function reserveSpend(tenantId: string, worstCaseUsd: number): Promise<boolean> // the SQL above
declare function callModel(history: unknown[], signal: AbortSignal): Promise<Step>
declare function runTool(name: string, args: unknown, signal: AbortSignal): Promise<unknown>

const CALL_TIMEOUT_MS = 30_000
const RUN_DEADLINE_MS = 120_000
const WORST_CASE_USD = 0.05 // per step, from max input + max output at this model's price

export async function runAgent(tenantId: string, history: unknown[], kill: AbortController): Promise<string> {
  const budget = new RunBudget(8, 150_000)
  const runDeadline = AbortSignal.timeout(RUN_DEADLINE_MS)

  try {
    for (;;) {
      if (await killSwitchOn('support-agent')) return 'This assistant is paused. A person will follow up.'
      budget.beforeStep()
      if (!(await reserveSpend(tenantId, WORST_CASE_USD))) return 'Daily AI limit reached for this account.'

      const signal = AbortSignal.any([AbortSignal.timeout(CALL_TIMEOUT_MS), runDeadline, kill.signal])
      const step = await callModel(history, signal)
      budget.addUsage(step.inputTokens, step.outputTokens)
      if (step.done) return step.text ?? ''

      if (step.tool) {
        budget.checkToolCall(step.tool.name, step.tool.args)
        history.push({ tool: step.tool.name, result: await runTool(step.tool.name, step.tool.args, signal) })
      }
    }
  } catch (err) {
    if (err instanceof BudgetExceeded) return 'I could not finish this within my limits. A person will follow up.'
    if (err instanceof Error && (err.name === 'TimeoutError' || err.name === 'AbortError')) return 'This took too long. Please try again.'
    throw err
  }
}

The kill switch is a row in a flags table, read at the start of every step with a short cache, so flipping it stops every run within seconds, including runs already in progress. Keep one flag per feature and one global flag. Pass the signal into your model client and tool calls; most HTTP clients and model SDKs accept an abort signal, and a signal nobody passes does nothing. Refunding unused reservations and logging why each run stopped are left out of the sketch for length, but both belong in production.

How do you stop retries from multiplying calls?

Choose one layer to own retries and switch them off everywhere else. If your code, the SDK and your job queue each retry three times, one bad request can become 27 calls.

  • Retry only what can succeed: rate-limit and server errors, with exponential backoff and jitter. Never retry a validation error or a budget stop.
  • Cap attempts per run, and count retried calls against the same step, token and spend budget.
  • Make jobs idempotent, so a re-delivered job checks whether its work is already done before calling a model.
  • Send exhausted jobs to a dead-letter queue for a person to look at, instead of back to the start.

How do you rate limit agent runs per user?

Limit how many runs a user can start, separately from how many HTTP requests they send, because one request can start an expensive run. OWASP's LLM10 guidance lists rate limiting, input size limits, timeouts and throttling, monitoring and graceful degradation among its mitigations. RAITHub's own site limits requests without Redis, using Postgres; the approach is in rate limiting without Redis on serverless, and the general patterns are in API rate limiting explained. The same table design works for "runs started per user per hour".

How do you test that the guardrails work?

Write tests that try to make the agent run away, and assert that it stops. A guardrail nobody has seen fire is a guess.

  • A tool that always fails. Assert the run stops on the repeated call or the step cap.
  • A model stub that never says "done". Assert the step cap fires at exactly the configured number.
  • A tenant one cent below its limit. Assert the next call is refused, and that two parallel runs cannot both pass.
  • A tool that hangs. Assert the run ends near the call timeout with a TimeoutError.
  • Kill switch on mid-run. Assert the next step returns the paused message.

These are ordinary unit and integration tests with stubbed models, so they run in CI on every change. The broader method for testing AI features is in how to test LLM features, and what agents cost to build and run is in what it costs to build an AI agent.

Why RAITHub for agent guardrails

  • Cost controls built into a real AI product. PadhAI, the AI tutoring platform RAITHub built, uses a 70/20/10 model router to keep cost per student viable.
  • Limits without extra infrastructure. This site's rate limiting runs without Redis, and its Neon database cost fix came from measuring before optimising.
  • Guardrails are tested. The same habit behind 400+ tests on this site and 750+ on TheSkinProof, the founder's own marketplace venture.
  • Fixed scope, your IP. A free 15-minute technical audit, then a fixed written quote. You own the code and IP; an NDA is standard.

When you don't need us

  • Your workflow is a single model call. A max output setting, a timeout and a rate limit cover it.
  • Your framework already enforces what you need and your team can add the spend table above in an afternoon.
  • The bill is already out of control today. Flip the provider's console limits first; call anyone afterwards.

Guardrails sit in the backend of an AI feature and are built under the API and backend development service. If an agent is already misbehaving in production, the fix one issue page lists what to send. To design limits for a new agent, book the free 15-minute technical audit and bring your agent's tools and the busiest day of usage you expect.

Last reviewed: 29 September 2026. Prices and sources checked on 29 September 2026.

Frequently asked questions

How do I stop an LLM agent from looping forever?

Enforce a maximum number of steps in code, stop on a repeated identical tool call, and add a timeout for the whole run. Do not rely on the prompt telling the model to stop.

What is Denial of Wallet?

An attack, described in OWASP's LLM10:2025 Unbounded Consumption, where someone drives excessive use of a pay-per-use AI service to run up the owner's bill. Per-tenant spend limits and rate limits are the main defence.

How should I set a spend limit per customer?

Store daily spend per tenant in the database, reserve each call's worst-case cost atomically before making it, refuse the call if the limit would be crossed, and credit back the unused amount afterwards.

What is an LLM kill switch?

A flag your operators can flip, without a deploy, that every agent step checks before running. Turning it on pauses a feature, or all AI features, within seconds.

What timeouts should an agent have?

One per model or tool call, measured in seconds, and a deadline for the whole run, measured in minutes. Combine them with the kill signal using AbortSignal.any.

Are provider usage limits enough on their own?

No. They are a useful backstop, but they stop everything at once. Per-tenant limits in your own code stop only the run or customer causing the problem.

runaway LLM agentsLLM guardrailsAI agent limitskill switchLLM spend limittimeoutsunbounded consumption

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.