Section 01 · Unit introduction

From pilot to production

A pilot proves that a thing can work. Production proves that it keeps working when the load is ten times higher, the traffic is uneven, the model endpoint occasionally throttles, and someone in finance is watching the bill. The gap between the two is almost entirely architectural, and it is where well-received pilots quietly die.

This unit covers the four decisions that separate a demo from a dependable service: reliability, capacity and quota planning, performance, and cost. None of them is about making the AI smarter. They are about making the system around the AI predictable under load.

By the end of this unit
  • Design for reliability using retries, fallbacks, and regional resilience for Azure AI workloads.
  • Plan capacity using tokens-per-minute quota and the choice between standard and provisioned throughput.
  • Identify the main levers for controlling latency and cost as an AI workload scales.

What changes between pilot and production

  • One or two users, predictable traffic
  • The endpoint is always available and never throttles
  • Latency is "fine" because nobody is waiting in a queue
  • Cost is invisible because volume is tiny
  • Bursty, concurrent traffic with peaks far above the average
  • The endpoint returns 429 (rate-limited) under load
  • Tail latency matters — the slowest 5% of calls shape user perception
  • Token consumption is the dominant operating cost
  • Retry-with-backoff and a fallback path are mandatory, not optional
  • You must plan quota deliberately and consider provisioned throughput
  • You must measure tokens and latency, and tune prompts and models for both
  • You need cost attribution per use case to make decisions
The pilot answers "can it work?" Production answers "does it keep working when I am not watching?" Only the second question pays the bills.
Working principle · Productionising AI
Section 02

Reliability and resilience

Azure AI model endpoints are highly available but not infinitely so. Under load they will sometimes return a 429 (too many requests) or a transient 5xx. A production design treats these as expected events with a planned response, not as exceptions that surface to the user.

1. Retry with exponential backoff and jitter

On a 429 or transient 5xx, retry after a short, increasing delay with a small random component so that many clients do not retry in lockstep. Respect the Retry-After header when the service provides one. Cap the number of retries so a struggling endpoint does not amplify into a retry storm.

2. Provide a fallback path

Decide in advance what happens when retries are exhausted. Options include falling back to a second deployment (a different region or a smaller, cheaper model), serving a cached or templated response, or degrading the feature gracefully with a clear message. The worst outcome is an unhandled error reaching the user.

3. Design for regional resilience

For workloads that cannot tolerate a regional outage, deploy the same model in two regions and route between them. Microsoft's Global and Data Zone deployment options can also spread load across regions for managed throughput. Match the resilience investment to the actual availability requirement — not every workload needs multi-region.

Design note

A circuit breaker complements retries: if a downstream endpoint is failing consistently, stop calling it for a short cool-off period and serve the fallback immediately. This protects both your latency budget and the recovering endpoint. Retries handle the occasional blip; the circuit breaker handles the sustained outage.

Section 03

Capacity and quota planning

Azure OpenAI capacity is governed by quota expressed in tokens-per-minute (TPM), with an associated requests-per-minute limit. Getting this wrong is the most common cause of throttling in production. Plan it before launch, not after the first incident.

Standard versus provisioned throughput

Standard (pay-as-you-go) deployments
You are billed per token consumed and allocated TPM quota that you distribute across your deployments. This is ideal for variable or unpredictable traffic and for early production where volume is still being learned. The trade-off is that latency can vary under shared load and you must manage quota carefully to avoid throttling at peak.
Provisioned Throughput Units (PTUs)
You reserve dedicated capacity, measured in PTUs, for predictable, high-volume workloads. This gives stable, predictable latency and a known throughput ceiling, billed at a fixed rate rather than per token. It suits steady, high-traffic production workloads where consistent performance matters and volume is high enough to make the reservation economical. Many mature deployments use a hybrid: PTUs for the steady baseline and standard deployment to absorb spikes (spillover).
Estimating your TPM requirement
Estimate peak concurrent requests, multiply by the average tokens per request (prompt plus completion), and convert to per-minute. Add headroom for bursts — peak traffic is typically well above average. Crucially, count both input and output tokens; a long system prompt or large retrieved context can dwarf the user's question. Use Azure AI Foundry's metrics to measure real token usage in your pilot and refine the estimate before scaling.
Protecting against noisy neighbours
When several use cases share one deployment's quota, a single misbehaving caller can starve the others. Separate critical workloads onto their own deployments with reserved quota, and apply client-side rate limiting per use case so no single consumer can exhaust the shared pool. In a hub-and-spoke architecture, attribute quota to projects deliberately rather than letting it be first-come-first-served.
Reflect

For a workload you know, what is the average number of tokens per request — including the system prompt and any retrieved context, not just the user's input? If you are guessing, that is the first measurement to take. Capacity planning built on a guessed token count is a planning exercise built on sand.

Section 04

Performance and cost at scale

Latency and cost are two faces of the same set of decisions. Both are driven primarily by how many tokens you process and which model processes them. The good news is that the levers are concrete and largely within your control.

1. Right-size the model

The largest model is rarely the right default. A smaller, faster model often meets quality requirements at a fraction of the cost and latency. Route easy requests to a small model and reserve the large model for the cases that genuinely need it. Validate the smaller model's quality with an evaluation set rather than assuming it is inadequate.

2. Cut tokens, not corners

Tokens are the unit of both cost and latency. Trim verbose system prompts, retrieve only the most relevant context rather than stuffing everything in, and cap completion length where the use case allows. In RAG patterns, better retrieval often beats more retrieval — fewer, more relevant chunks lower cost and improve quality at once.

3. Cache and stream

Cache responses for repeated or near-identical requests to avoid paying for the same generation twice. Stream tokens to the user so perceived latency drops even when total generation time is unchanged — the user sees the answer begin immediately. Both improve the experience without touching model quality.

4. Attribute cost to use cases

Tag deployments and projects so you can answer "what did this use case cost last month". Without attribution you cannot make rational decisions about which features justify their spend. Cost governance at scale begins with being able to see where the money goes.

End of Unit 2

You should now be able to:

  • Add retries, fallbacks, and regional resilience to an Azure AI workload.
  • Plan TPM quota and choose between standard and provisioned throughput.
  • Apply the model, token, caching, and attribution levers to control latency and cost.
Section 05

Unit review

Question 1 of 4

An Azure OpenAI endpoint returns HTTP 429 under load. What is the correct production response?

Question 2 of 4

A workload has steady, high, predictable traffic and needs consistent latency. Which capacity option fits best?

Question 3 of 4

When estimating TPM quota, which tokens must you count?

Question 4 of 4

Costs are too high at scale. Which lever reduces both cost and latency at the same time?

End of module

You have completed Course 02: Designing for Scale. Next: Azure OpenAI Service Overview — models from deployment to consumption.

Craig Stanley Studio · Deploy — Architecture, Integration & Operations · Designing for Scale · Access by direct link only.