Token economics
Every interaction with a large language model is billed in tokens — fragments of text, roughly four characters or three-quarters of a word in English. You pay for the tokens you send (the prompt, including any context and history) and the tokens the model generates (the completion). Understanding this single mechanic is the foundation of every cost decision you will make, because almost all AI cost reduces to: how many tokens, how often, on which model.
The trap is that token costs are invisible per request and enormous in aggregate. A prompt that is a few hundred tokens larger than it needs to be costs nothing noticeable once — and a fortune across millions of calls. This unit gives you the economics and the practical levers to control spend without degrading quality.
By the end of this unit- Explain how token-based pricing works and why input and output tokens are billed separately.
- Apply prompt, context, and output efficiency techniques that cut tokens without losing quality.
- Choose appropriate caching, batching, model-selection, and capacity strategies for a given workload.
Where the tokens — and the cost — come from
- The prompt, system message, retrieved context, and conversation history
- Grows with every document you ground on and every turn you retain
- Often the larger share in retrieval-grounded and long-conversation apps
- The biggest controllable lever in most workloads
- The text the model generates in reply
- Typically priced higher per token than input
- Controlled by asking for concise answers and capping max tokens
- Verbose-by-default responses quietly inflate this
- A 300-token saving per call is invisible once and material across millions
- Cost scales with users × frequency × tokens-per-call
- Small per-call inefficiencies become the dominant line item at scale
- This is why cost is a design constraint, not a post-launch cleanup
Prompt and context efficiency
The single largest cost lever in most production AI systems is the size of what you send. Retrieval-grounded applications routinely pad the prompt with far more context than the model needs to answer well. Trimming this is the highest-return optimisation available, and done correctly it improves quality as well as cost — a focused context produces a more focused answer.
Right-size the retrieved context
Retrieve only the passages relevant to the query, not whole documents. Tune how many chunks you pass and their size so the model receives enough to answer and no more. Over-retrieval is the most common source of bloated input cost, and it also dilutes the model's focus with irrelevant material.
Trim the system prompt and instructions
System prompts accumulate instructions over time, many redundant or never triggered. Because the system prompt is sent on every single call, every wasted token there is multiplied by your entire request volume. Periodically prune it to the instructions that genuinely change behaviour.
Manage conversation history deliberately
Sending the full conversation history on every turn makes each turn progressively more expensive. Summarise older turns, keep a sliding window of recent context, or retain only what is needed for continuity. Unmanaged history is a cost that compounds with conversation length.
Constrain the output
Ask for the format and length you actually need, and set a sensible max-tokens cap. A request that yields a three-sentence answer instead of three paragraphs costs less on the higher-priced output tokens and is often more useful. Specify conciseness explicitly rather than hoping for it.
Measure before you cut. Use your token telemetry to find which calls carry the largest input — they are almost always the over-retrieving, full-history ones. Optimising the heaviest 10% of calls usually delivers most of the saving, and lets you leave the rest alone.
A team wants to cut the cost of a retrieval-grounded assistant. Telemetry shows input tokens are five times output tokens. Where should they look first?
Caching, batching, and model selection
Beyond trimming individual requests, three patterns reduce cost structurally: avoid paying for the same work twice (caching), pay less for work that is not urgent (batching), and pay only for the capability you need (model selection). Each suits a different part of a workload.
Three structural levers
Caching — do not pay twice for the same work
Batching — pay less for work that can wait
Model selection — match capability to task
Combining the levers
Think of one AI workload you run today. Is any of it non-interactive enough to batch? Is it using a bigger model than the task needs? Often the largest savings are sitting in a workload that was never questioned after it first shipped.
Capacity and commitment
The final lever is how you buy capacity. Consumption (pay-as-you-go) pricing flexes with demand and is ideal for variable or unpredictable workloads. For steady, high-volume workloads, committing to capacity in advance — provisioned throughput or reserved capacity — trades flexibility for a lower effective rate and predictable performance. Choosing the wrong model for a workload's shape wastes money in either direction.
Consumption pricing for variable demand
Pay per token used, scaling up and down with no commitment. Right for pilots, spiky workloads, and anything whose volume you cannot yet predict. The flexibility is the value; you never pay for idle capacity. The trade-off is a higher per-unit rate and less predictable cost at high volume.
Provisioned throughput for steady, high volume
Reserve dedicated capacity for a predictable, sustained workload. You get stable, predictable performance and a lower effective cost per unit at high utilisation. The trade-off is that you pay for the reserved capacity whether or not you use it — so it only pays off when utilisation is genuinely high and steady.
Decide on evidence, not instinct
Use real usage data to choose. Run on consumption first, watch the token telemetry until the demand shape is clear, then commit capacity only for the portion of demand that is proven, steady, and high enough to justify it. Committing too early, on a guess, is a classic way to lock in cost for capacity you do not use.
Govern cost continuously
Budget per workload, alert on cost anomalies the way you alert on quality, and review spend against value regularly. Cost optimisation is not a one-off project; it is a standing part of operating AI in production, and it ties directly back to the monitoring discipline from Course 17.
Platform context for model choice and efficient deployment:
End of Unit 20
You should now be able to:
- Reason about AI cost through token economics — input, output, and aggregate volume.
- Cut tokens through context, prompt, history, and output efficiency without losing quality.
- Apply caching, batching, model selection, and the right capacity-commitment model to each workload.
Unit review
Why is per-request token inefficiency such a serious cost problem in production?
In a retrieval-grounded app where input tokens far exceed output, which optimisation usually yields the most saving?
An overnight job classifies a backlog of thousands of documents with no need for instant results. What is the most cost-effective approach?
When does committing to provisioned throughput rather than consumption pricing pay off?
End of module
You have completed Course 20: Cost Optimisation for AI Workloads — and the Deploy pillar's operations sequence. You can now manage quality, monitoring, incidents, scaling, and cost as one connected production discipline.