Founder & Lead Engineer, RAITHub
To cut an OpenAI API bill: log cost per feature, cap output tokens (8 times the input price on gpt-5), put stable instructions first so caching bills them at a tenth, batch work that can wait at 50% off, cache repeated answers, and route easy requests to smaller models. The worked example below falls from $775 to $75 a month.
This guide is for teams whose AI feature works and whose invoice is now the problem. The levers are the same on every major provider; the prices quoted are OpenAI's, from the OpenAI pricing page, checked on 29 September 2026. The approach draws on PadhAI, the AI tutoring platform RAITHub built, whose 70/20/10 model router is how it keeps AI cost per student viable.
Where does an OpenAI API bill actually come from?
From tokens: input tokens you send, output tokens the model writes, and, on reasoning models, hidden reasoning tokens. The OpenAI reasoning guide says reasoning tokens are billed as output tokens, even though you never see them.
| Model | Input per 1M tokens | Cached input | Output |
|---|---|---|---|
| gpt-5-nano | $0.05 | $0.005 | $0.40 |
| gpt-5-mini | $0.25 | $0.025 | $2.00 |
| gpt-5 | $1.25 | $0.125 | $10.00 |
| gpt-5.5 | $5.00 | $0.50 | $30.00 |
Standard-tier list prices from the OpenAI pricing page above. Three things stand out. Output costs 8 times input on gpt-5. Cached input costs a tenth of fresh input. And gpt-5-nano costs one twenty-fifth of gpt-5 on every line. Each of those is a lever.
How do you find which feature is costing the most?
Log the usage object of every call with the feature name and tenant, then sum it. Most bills are dominated by one or two features, and you cannot cut what you have not measured.
import OpenAI from 'openai'
const openai = new OpenAI()
const PRICE = { 'gpt-5-mini': { input: 0.25, cached: 0.025, output: 2.0 } } as const
export async function tracked(feature: string, tenantId: string, instructions: string, input: string) {
const res = await openai.responses.create({
model: 'gpt-5-mini',
instructions, // stable text first, so it can be cached
input, // the part that changes per request
max_output_tokens: 800, // caps visible output and reasoning together
reasoning: { effort: 'low' },
})
const u = res.usage
if (u) {
const cached = u.input_tokens_details?.cached_tokens ?? 0
const p = PRICE['gpt-5-mini']
const costUsd = ((u.input_tokens - cached) * p.input + cached * p.cached + u.output_tokens * p.output) / 1_000_000
await logCost({ feature, tenantId, model: 'gpt-5-mini', ...u, costUsd }) // your own table
}
return res.output_text
}
declare function logCost(row: Record<string, unknown>): Promise<void>
The usage object reports input tokens, cached input tokens, output tokens and, under output_tokens_details, reasoning tokens. Two cautions from the reasoning guide: max_output_tokens covers reasoning and visible output together, so a cap set too low returns an incomplete answer, and supported reasoning.effort values vary by model. Keep a price table per model in code and update it when the pricing page changes.
Which cost levers save the most?
| Lever | What it saves | Effort | Catch |
|---|---|---|---|
| Cap output and reasoning | Output tokens, the priciest line | Low | Too low a cap truncates answers |
| Prompt caching | Cached input billed at 10% of the input price | Low: reorder the prompt | Only the repeated prefix, and only above a minimum length |
| Batch API | 50% off input and output | Medium: a job queue | Results arrive later, not in the request |
| Response caching | 100% of repeated calls | Medium | Only for non-personal, repeatable questions |
| Trim context | Input tokens on every call | Medium | Cut too much and quality drops |
| Route by difficulty | Up to 96% per request moved from gpt-5 to gpt-5-nano | Medium to high | Needs an evaluation set to route safely |
How does prompt caching reduce OpenAI costs?
When the start of a prompt matches a recent request, OpenAI bills those tokens at the cached rate. It is automatic; your job is to make the start of every prompt identical.
The OpenAI prompt caching guide says a prefix must reach the model's minimum cacheable length, 1,024 tokens on GPT-5.6 and later, and advises: "Put stable developer instructions and shared reference material first." It also describes cached prefixes staying eligible for about 30 minutes after last use on newer models. In practice:
- Instructions, examples and tool definitions go first, word for word the same on every call.
- Anything that varies, such as the user's name, today's date, the question or retrieved passages, goes after them.
- A timestamp at the top of the system prompt silently breaks caching for every request. Move it to the end.
- Check
cached_tokensin your logs. If it stays at zero, the prefix is changing.
When should you use the OpenAI Batch API?
For anything a user is not waiting on: nightly tagging, backfills, bulk summaries, evaluation runs. The OpenAI pricing page lists batch processing at 50% of standard rates.
{"custom_id": "ticket-1042", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "gpt-5-mini", "messages": [{"role": "system", "content": "Tag the ticket with one category code."}, {"role": "user", "content": "Refund not received after 10 days."}], "max_completion_tokens": 300}}
{"custom_id": "ticket-1043", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "gpt-5-mini", "messages": [{"role": "system", "content": "Tag the ticket with one category code."}, {"role": "user", "content": "Cannot log in with SSO."}], "max_completion_tokens": 300}}
Upload the JSONL file, create a batch job, and match results back by custom_id. Anthropic and Google list the same 50% batch discount on their pricing pages, so the pattern carries across providers.
How do you cache whole responses safely?
Hash the model, instructions and input into a key, and store the answer for a fixed time. A cache hit costs nothing. Scope the key by tenant, and never cache answers that contain personal data.
import { createHash } from 'node:crypto'
import type { Pool } from 'pg'
// CREATE TABLE llm_cache (key text PRIMARY KEY, output text NOT NULL, expires_at timestamptz NOT NULL);
export async function cachedAnswer(
db: Pool,
tenantId: string,
model: string,
instructions: string,
input: string,
call: () => Promise<string>,
): Promise<string> {
const key = createHash('sha256').update(JSON.stringify([tenantId, model, instructions, input.trim().toLowerCase()])).digest('hex')
const hit = await db.query('SELECT output FROM llm_cache WHERE key = $1 AND expires_at > now()', [key])
if (hit.rows[0]) return hit.rows[0].output
const output = await call()
await db.query(
"INSERT INTO llm_cache (key, output, expires_at) VALUES ($1, $2, now() + interval '7 days') " +
'ON CONFLICT (key) DO UPDATE SET output = EXCLUDED.output, expires_at = EXCLUDED.expires_at',
[key, output],
)
return output
}
Including the instructions and model in the key means a prompt change or model switch misses the cache instead of serving stale answers. Exact-match caching works well for FAQ-style questions, product descriptions and summaries of documents that have not changed.
How do you send less context without hurting quality?
Send the passages the model needs, not the whole document or the whole conversation. Input tokens are paid on every call, so context that grows with each turn grows the bill with it.
- Retrieve, don't paste. Retrieval-augmented generation (RAG) sends the few relevant passages instead of a full manual. The build is in how to add AI features to an existing SaaS.
- Summarise long chats. After a set number of turns, replace older turns with a short summary.
- Trim tool results. Return the fields the model needs, not the entire API payload.
- Ask for short output. "Reply in at most three sentences" plus a token cap does more than either alone.
How does model routing cut LLM costs?
Classify each request, and send it to the least expensive model that handles it well. Most traffic in most products is easy; paying premium prices for it is the largest avoidable cost.
PadhAI's 70/20/10 model router classifies query complexity and sends each query to a lightweight, mid-tier or premium model. The name describes the intended traffic split, and the router is how PadhAI keeps AI cost per student viable. RAITHub does not publish PadhAI's cost figures. Route only behind an evaluation set: run the golden set on the smaller model first, and send it only the request types where it passes. PadhAI's wider design is in how to build an AI tutor.
What does 50–80% savings look like in real numbers?
An illustrative workload: 100,000 calls a month on gpt-5, each with a 2,000-token stable prefix, 1,000 tokens of changing input and 400 output tokens. Prices from the table above.
| Step | What changes | Monthly cost | Saved so far |
|---|---|---|---|
| Baseline | 300M input at $1.25, 40M output at $10 | $775.00 | 0% |
| Cap output | Average output cut to 250 tokens: 25M at $10 | $625.00 | 19% |
| Prompt caching | 80% of prefix tokens hit the cache: 160M at $0.125, 140M at $1.25 | $445.00 | 43% |
| Route 70/20/10 | 70% to gpt-5-nano, 20% to gpt-5-mini, 10% stays on gpt-5 | $74.76 | 90% |
Worked examples, not benchmarks and not PadhAI figures. The routing line assumes 70% of requests pass your evaluation on the smallest model; if only half do, the saving is smaller. That is why the realistic range for most products is 50–80%, and why output caps and caching, which need no quality trade-off, come first. Moving any waiting work to the Batch API halves that portion again.
Why RAITHub for LLM cost work
- Cost designed in. RAITHub built PadhAI's 70/20/10 model router, alongside math verification and RAG, across 11 services.
- Measured, not guessed. A cost-per-feature log comes first, so every change is judged by the number it moved.
- Serverless cost habits. This site's own Neon database cost fix and rate limiting without Redis came from the same habit of measuring before optimising.
- Fixed scope. A free 15-minute technical audit, then a fixed written quote. You own the code and the IP, and an NDA is standard.
When you don't need us
- Your bill is small. If the AI line is under what an afternoon of engineering costs each month, reordering the prompt for caching and adding a token cap may be all you need.
- Your team can follow this list. None of these levers is exotic; a capable in-house engineer can apply them in order.
- You need a volume discount. That is a conversation with the provider's sales team, not an engineering job.
Cost work on the AI layer of a backend is covered by the API and backend development service. To find where your bill goes, book the free 15-minute technical audit and bring one month's usage export and the features that call the API.
Last reviewed: 29 September 2026. Prices and docs checked on 29 September 2026.
Frequently asked questions
What is the quickest way to reduce OpenAI API costs?
Cap output tokens and move stable instructions to the start of the prompt so they are cached. Neither needs a model change, and cached input costs a tenth of the normal input price.
Does OpenAI prompt caching happen automatically?
Yes, for prompts whose prefix reaches the model's minimum cacheable length, 1,024 tokens on GPT-5.6 and later. Keep the start of each prompt identical to benefit.
How much does the OpenAI Batch API save?
50% of standard rates, per OpenAI's pricing page, in exchange for results that arrive asynchronously rather than in the request.
Are reasoning tokens billed?
Yes. OpenAI's reasoning guide says reasoning tokens are billed as output tokens. Lower reasoning effort where quality allows, and cap total output with max_output_tokens.
Is switching to a smaller model safe?
Only for request types where the smaller model passes your evaluation set. Route those, and keep the rest on the larger model, as PadhAI's 70/20/10 router does by complexity.
Can I cache LLM responses?
Yes, for repeatable, non-personal questions. Key the cache by tenant, model, instructions and input, give it an expiry, and never cache answers containing personal data.
Related posts
Ready to discuss your project?
Book a free 15-minute technical audit with our engineering team.