Section 01 · Unit introduction

Token economics

Every interaction with a large language model is billed in tokens — fragments of text, roughly four characters or three-quarters of a word in English. You pay for the tokens you send (the prompt, including any context and history) and the tokens the model generates (the completion). Understanding this single mechanic is the foundation of every cost decision you will make, because almost all AI cost reduces to: how many tokens, how often, on which model.

The trap is that token costs are invisible per request and enormous in aggregate. A prompt that is a few hundred tokens larger than it needs to be costs nothing noticeable once — and a fortune across millions of calls. This unit gives you the economics and the practical levers to control spend without degrading quality.

By the end of this unit
  • Explain how token-based pricing works and why input and output tokens are billed separately.
  • Apply prompt, context, and output efficiency techniques that cut tokens without losing quality.
  • Choose appropriate caching, batching, model-selection, and capacity strategies for a given workload.

Where the tokens — and the cost — come from

  • The prompt, system message, retrieved context, and conversation history
  • Grows with every document you ground on and every turn you retain
  • Often the larger share in retrieval-grounded and long-conversation apps
  • The biggest controllable lever in most workloads
  • The text the model generates in reply
  • Typically priced higher per token than input
  • Controlled by asking for concise answers and capping max tokens
  • Verbose-by-default responses quietly inflate this
  • A 300-token saving per call is invisible once and material across millions
  • Cost scales with users × frequency × tokens-per-call
  • Small per-call inefficiencies become the dominant line item at scale
  • This is why cost is a design constraint, not a post-launch cleanup
Token cost is invisible per request and decisive in aggregate. Optimise the request, and you optimise the bill a million times over.
Working principle · AI cost optimisation
Section 02

Prompt and context efficiency

The single largest cost lever in most production AI systems is the size of what you send. Retrieval-grounded applications routinely pad the prompt with far more context than the model needs to answer well. Trimming this is the highest-return optimisation available, and done correctly it improves quality as well as cost — a focused context produces a more focused answer.

Right-size the retrieved context

Retrieve only the passages relevant to the query, not whole documents. Tune how many chunks you pass and their size so the model receives enough to answer and no more. Over-retrieval is the most common source of bloated input cost, and it also dilutes the model's focus with irrelevant material.

Trim the system prompt and instructions

System prompts accumulate instructions over time, many redundant or never triggered. Because the system prompt is sent on every single call, every wasted token there is multiplied by your entire request volume. Periodically prune it to the instructions that genuinely change behaviour.

Manage conversation history deliberately

Sending the full conversation history on every turn makes each turn progressively more expensive. Summarise older turns, keep a sliding window of recent context, or retain only what is needed for continuity. Unmanaged history is a cost that compounds with conversation length.

Constrain the output

Ask for the format and length you actually need, and set a sensible max-tokens cap. A request that yields a three-sentence answer instead of three paragraphs costs less on the higher-priced output tokens and is often more useful. Specify conciseness explicitly rather than hoping for it.

Practitioner note

Measure before you cut. Use your token telemetry to find which calls carry the largest input — they are almost always the over-retrieving, full-history ones. Optimising the heaviest 10% of calls usually delivers most of the saving, and lets you leave the rest alone.

Knowledge check

A team wants to cut the cost of a retrieval-grounded assistant. Telemetry shows input tokens are five times output tokens. Where should they look first?

Section 03

Caching, batching, and model selection

Beyond trimming individual requests, three patterns reduce cost structurally: avoid paying for the same work twice (caching), pay less for work that is not urgent (batching), and pay only for the capability you need (model selection). Each suits a different part of a workload.

Three structural levers

Caching — do not pay twice for the same work
If many requests share the same large, stable prefix — a long system prompt, a fixed instruction set, the same grounding documents — prompt caching lets the model reuse that processed prefix at reduced cost rather than reprocessing it every call. For workloads with a big, repeated context and varying user queries, caching can cut input cost substantially. You can also cache full responses to identical queries where answers are stable.
Batching — pay less for work that can wait
Not all AI work is interactive. Bulk classification, summarisation of a backlog, or overnight enrichment jobs do not need an instant response. Batch processing submits many requests together for asynchronous completion, typically at a meaningfully lower price than real-time calls. Move every non-interactive workload onto the batch path and reserve real-time pricing for genuinely interactive use.
Model selection — match capability to task
Larger, more capable models cost more per token. Many tasks — classification, extraction, routing, simple summarisation — are handled well by smaller, cheaper models. Reserve the most capable models for the tasks that genuinely need them. A common pattern is to route requests: a small model handles the routine majority, escalating only hard cases to a larger model. This often cuts cost dramatically with no perceptible quality loss.
Combining the levers
These compound. A batch job, on a right-sized smaller model, with a cached stable prefix, can cost a fraction of the naive real-time, large-model, full-context equivalent — for identical output. The discipline is to ask of every workload: does this need to be real-time? Does it need the biggest model? Is its context repeated? The answers route it to the cheapest path that still meets the quality bar.
Reflect

Think of one AI workload you run today. Is any of it non-interactive enough to batch? Is it using a bigger model than the task needs? Often the largest savings are sitting in a workload that was never questioned after it first shipped.

Section 04

Capacity and commitment

The final lever is how you buy capacity. Consumption (pay-as-you-go) pricing flexes with demand and is ideal for variable or unpredictable workloads. For steady, high-volume workloads, committing to capacity in advance — provisioned throughput or reserved capacity — trades flexibility for a lower effective rate and predictable performance. Choosing the wrong model for a workload's shape wastes money in either direction.

Consumption pricing for variable demand

Pay per token used, scaling up and down with no commitment. Right for pilots, spiky workloads, and anything whose volume you cannot yet predict. The flexibility is the value; you never pay for idle capacity. The trade-off is a higher per-unit rate and less predictable cost at high volume.

Provisioned throughput for steady, high volume

Reserve dedicated capacity for a predictable, sustained workload. You get stable, predictable performance and a lower effective cost per unit at high utilisation. The trade-off is that you pay for the reserved capacity whether or not you use it — so it only pays off when utilisation is genuinely high and steady.

Decide on evidence, not instinct

Use real usage data to choose. Run on consumption first, watch the token telemetry until the demand shape is clear, then commit capacity only for the portion of demand that is proven, steady, and high enough to justify it. Committing too early, on a guess, is a classic way to lock in cost for capacity you do not use.

Govern cost continuously

Budget per workload, alert on cost anomalies the way you alert on quality, and review spend against value regularly. Cost optimisation is not a one-off project; it is a standing part of operating AI in production, and it ties directly back to the monitoring discipline from Course 17.

Continue on Microsoft Learn

Platform context for model choice and efficient deployment:

Get started with Azure AI Foundry ↗

Fine-tune models with Azure AI Foundry ↗

End of Unit 20

You should now be able to:

  • Reason about AI cost through token economics — input, output, and aggregate volume.
  • Cut tokens through context, prompt, history, and output efficiency without losing quality.
  • Apply caching, batching, model selection, and the right capacity-commitment model to each workload.
Section 05

Unit review

Question 1 of 4

Why is per-request token inefficiency such a serious cost problem in production?

Question 2 of 4

In a retrieval-grounded app where input tokens far exceed output, which optimisation usually yields the most saving?

Question 3 of 4

An overnight job classifies a backlog of thousands of documents with no need for instant results. What is the most cost-effective approach?

Question 4 of 4

When does committing to provisioned throughput rather than consumption pricing pay off?

End of module

You have completed Course 20: Cost Optimisation for AI Workloads — and the Deploy pillar's operations sequence. You can now manage quality, monitoring, incidents, scaling, and cost as one connected production discipline.

Craig Stanley Studio · Deploy — Architecture, Integration & Operations · Cost Optimisation for AI Workloads · Access by direct link only.