From pilot to production
A pilot proves that a thing can work. Production proves that it keeps working when the load is ten times higher, the traffic is uneven, the model endpoint occasionally throttles, and someone in finance is watching the bill. The gap between the two is almost entirely architectural, and it is where well-received pilots quietly die.
This unit covers the four decisions that separate a demo from a dependable service: reliability, capacity and quota planning, performance, and cost. None of them is about making the AI smarter. They are about making the system around the AI predictable under load.
By the end of this unit- Design for reliability using retries, fallbacks, and regional resilience for Azure AI workloads.
- Plan capacity using tokens-per-minute quota and the choice between standard and provisioned throughput.
- Identify the main levers for controlling latency and cost as an AI workload scales.
What changes between pilot and production
- One or two users, predictable traffic
- The endpoint is always available and never throttles
- Latency is "fine" because nobody is waiting in a queue
- Cost is invisible because volume is tiny
- Bursty, concurrent traffic with peaks far above the average
- The endpoint returns 429 (rate-limited) under load
- Tail latency matters — the slowest 5% of calls shape user perception
- Token consumption is the dominant operating cost
- Retry-with-backoff and a fallback path are mandatory, not optional
- You must plan quota deliberately and consider provisioned throughput
- You must measure tokens and latency, and tune prompts and models for both
- You need cost attribution per use case to make decisions
Reliability and resilience
Azure AI model endpoints are highly available but not infinitely so. Under load they will sometimes return a 429 (too many requests) or a transient 5xx. A production design treats these as expected events with a planned response, not as exceptions that surface to the user.
1. Retry with exponential backoff and jitter
On a 429 or transient 5xx, retry after a short, increasing delay with a small random component so that many clients do not retry in lockstep. Respect the Retry-After header when the service provides one. Cap the number of retries so a struggling endpoint does not amplify into a retry storm.
2. Provide a fallback path
Decide in advance what happens when retries are exhausted. Options include falling back to a second deployment (a different region or a smaller, cheaper model), serving a cached or templated response, or degrading the feature gracefully with a clear message. The worst outcome is an unhandled error reaching the user.
3. Design for regional resilience
For workloads that cannot tolerate a regional outage, deploy the same model in two regions and route between them. Microsoft's Global and Data Zone deployment options can also spread load across regions for managed throughput. Match the resilience investment to the actual availability requirement — not every workload needs multi-region.
A circuit breaker complements retries: if a downstream endpoint is failing consistently, stop calling it for a short cool-off period and serve the fallback immediately. This protects both your latency budget and the recovering endpoint. Retries handle the occasional blip; the circuit breaker handles the sustained outage.
Capacity and quota planning
Azure OpenAI capacity is governed by quota expressed in tokens-per-minute (TPM), with an associated requests-per-minute limit. Getting this wrong is the most common cause of throttling in production. Plan it before launch, not after the first incident.
Standard versus provisioned throughput
Standard (pay-as-you-go) deployments
Provisioned Throughput Units (PTUs)
Estimating your TPM requirement
Protecting against noisy neighbours
For a workload you know, what is the average number of tokens per request — including the system prompt and any retrieved context, not just the user's input? If you are guessing, that is the first measurement to take. Capacity planning built on a guessed token count is a planning exercise built on sand.
Performance and cost at scale
Latency and cost are two faces of the same set of decisions. Both are driven primarily by how many tokens you process and which model processes them. The good news is that the levers are concrete and largely within your control.
1. Right-size the model
The largest model is rarely the right default. A smaller, faster model often meets quality requirements at a fraction of the cost and latency. Route easy requests to a small model and reserve the large model for the cases that genuinely need it. Validate the smaller model's quality with an evaluation set rather than assuming it is inadequate.
2. Cut tokens, not corners
Tokens are the unit of both cost and latency. Trim verbose system prompts, retrieve only the most relevant context rather than stuffing everything in, and cap completion length where the use case allows. In RAG patterns, better retrieval often beats more retrieval — fewer, more relevant chunks lower cost and improve quality at once.
3. Cache and stream
Cache responses for repeated or near-identical requests to avoid paying for the same generation twice. Stream tokens to the user so perceived latency drops even when total generation time is unchanged — the user sees the answer begin immediately. Both improve the experience without touching model quality.
4. Attribute cost to use cases
Tag deployments and projects so you can answer "what did this use case cost last month". Without attribution you cannot make rational decisions about which features justify their spend. Cost governance at scale begins with being able to see where the money goes.
Go deeper on capacity, evaluation, and RAG at scale:
Build AI apps and agents with Azure AI Foundry ↗
Evaluate and monitor AI models in Azure AI Foundry ↗
Implement Retrieval Augmented Generation with Azure AI ↗
End of Unit 2
You should now be able to:
- Add retries, fallbacks, and regional resilience to an Azure AI workload.
- Plan TPM quota and choose between standard and provisioned throughput.
- Apply the model, token, caching, and attribution levers to control latency and cost.
Unit review
An Azure OpenAI endpoint returns HTTP 429 under load. What is the correct production response?
A workload has steady, high, predictable traffic and needs consistent latency. Which capacity option fits best?
When estimating TPM quota, which tokens must you count?
Costs are too high at scale. Which lever reduces both cost and latency at the same time?
End of module
You have completed Course 02: Designing for Scale. Next: Azure OpenAI Service Overview — models from deployment to consumption.