LLM costs rarely become a problem gradually. They are negligible in development, then a feature ships, usage grows, and finance asks what the line item is.
The good news: most LLM spend has obvious waste in it, and removing it rarely costs quality.
Measure cost per feature first
You cannot optimise an aggregate. A single monthly total tells you nothing about which of eight features is responsible.
Log, for every call: feature name, model, input tokens, output tokens, cached tokens, latency, and whether the user kept the result. Aggregate daily by feature.
The result is almost always lopsided — a single feature dominating, often one nobody considered expensive. A background summarisation job re-running over unchanged content is a classic. Find that before tuning prompts.
Also track cost per successful outcome, not cost per call. A cheap model that needs three attempts is more expensive than an accurate one that needs one, and only the outcome metric shows it.
Model routing
The largest single lever. Frontier models are priced for frontier work; most requests are not frontier work.
Classify by difficulty and route:
Simple — classification, extraction, formatting, short rewrites → small fast model
Moderate — summarisation, routine drafting → mid-tier
Hard — multi-step reasoning, code generation, nuanced analysis → largest model
Routing can be a heuristic (input length, feature, user tier) before it is anything clever. Even a crude split moves the number substantially, because the simple bucket is usually the biggest by volume.
Two patterns worth adding:
Escalate on failure. Try the small model; if output fails validation, retry on the larger one. Most requests settle at the cheap tier.
Draft then refine. A small model drafts, a large one polishes. Cheaper than generating entirely at the top tier.
Prompt caching
If you have a long stable prefix — a system prompt, a schema, a document being discussed — prompt caching is close to free money. Cached input tokens are billed at a large discount.
To benefit, the prefix must be byte-identical. Practical implications:
[ system prompt ][ schema ][ examples ][ document ] ← stable, cacheable
[ conversation turns ][ current question ] ← varies
Put everything stable first and everything variable last. A timestamp or a user id injected near the top of the prompt invalidates the cache on every call — and this is a genuinely common mistake, because it looks harmless.
Controlling output length
Output tokens typically cost several times more than input tokens, so verbosity is expensive.
Set
max_tokensdeliberately per feature rather than leaving a large defaultAsk for the format you want: "Answer in at most three sentences", "Return JSON only, no prose"
Use structured output / tool schemas — they eliminate the explanatory padding around the data you actually wanted
Stop generating as soon as the UI has what it needs
A tone instruction like "be concise" is not decoration; it is a cost control.
Batching
For anything not user-facing — nightly summarisation, bulk classification, backfills — batch APIs offer a substantial discount in exchange for latency.
Ask which jobs genuinely need a synchronous answer. Usually only the interactive ones do, and moving the rest to batch is a large saving for very little work.
Caching at the application layer
Before the model call, ask whether you need it at all.
Exact-match cache. Hash the normalised input; serve identical requests from cache. Content-generation workloads repeat more than you expect.
Semantic cache. Embed the query and serve a cached answer above a similarity threshold. Effective for support-style questions asked many ways, but set the threshold conservatively — a wrong cache hit is worse than a cache miss.
Cache derived artefacts. Embeddings keyed by content hash, summaries keyed by document version. Recomputing unchanged content is the most common waste of all.
Do not call the model for deterministic work. Parsing dates, validating formats, simple arithmetic and lookups do not need an LLM. This sounds obvious; it appears in production code constantly.
When a smaller model is genuinely enough
The real question is not "which model is best" but "what is the smallest model that passes my evals".
Build a test set from real inputs with acceptable outputs. Run each candidate. Compare accuracy and cost. Often a small model matches on the actual distribution of work, and the frontier model was bought for a hard case that represents 2% of traffic — which routing handles.
Re-run that comparison periodically. Model pricing and capability move quickly, and a routing decision from a year ago is likely no longer optimal.
The order that works: measure per feature → cache what repeats → route by difficulty → cache the prompt prefix → constrain output → batch the asynchronous work. Each step is independent, and the first two usually pay for the rest.