Section 01 · Unit introduction

The service in context

Azure OpenAI, now offered within Azure AI Foundry, gives you OpenAI's models — GPT-4 family, embeddings, and others — running inside your Azure tenant with enterprise controls: Microsoft Entra ID authentication, private networking, data residency, and content safety. The model capability is the same OpenAI you may know from elsewhere; what Azure adds is the governance, security, and commercial framework that an organisation needs to run it responsibly.

This unit walks the full path from choosing a model to consuming it in an application: the model catalogue, deployments, endpoints, how tokens are counted, how rate limits work, and how to keep the cost under control.

By the end of this unit
  • Navigate the model catalogue and create a deployment from a base model.
  • Explain the relationship between an endpoint, a deployment, and authentication.
  • Describe how tokens, rate limits, and quota drive both throughput and cost.

Azure OpenAI versus OpenAI directly

  • Authenticate with Microsoft Entra ID and managed identity, not just API keys
  • Role-based access control governs who can deploy and call models
  • Private endpoints keep traffic off the public internet
  • Choose the Azure region your deployment runs in
  • Your prompts and completions are not used to train the base models
  • Content and abuse-monitoring controls are configurable to your policy
  • Billed through your existing Azure agreement and cost-management tooling
  • Standard pay-as-you-go or reserved provisioned throughput
  • Cost attributable to subscriptions, resource groups, and tags
The model is the easy part. The reason organisations choose Azure OpenAI is everything around the model — identity, networking, residency, and billing they already trust.
Working principle · Enterprise AI adoption
Section 02

Model catalogue and deployments

In Azure AI Foundry, the model catalogue is where you browse and select models — from the GPT family to embedding models, plus open and partner models. Selecting a model is not the same as using it. To call a model, you create a deployment: a named instance of a model, in a region, with an assigned capacity. Your application then talks to the deployment, not to the model directly.

1. Choose a model from the catalogue

Match the model to the task. Chat and reasoning tasks use a GPT chat model; semantic search and RAG retrieval use an embedding model; some workloads combine both. Consider quality, latency, context window, and cost — the largest model is not always the right default. The catalogue shows the capabilities and the regions where each model is available.

2. Create a deployment

A deployment binds a model version to a region and a capacity allocation (TPM for standard, or PTUs for provisioned). You give it a deployment name — the identifier your application code references. Keeping deployment names stable lets you swap the underlying model version behind the same name during an upgrade.

3. Manage model versions and lifecycle

Models are versioned and have a lifecycle: new versions are released, older ones are retired on a published schedule. Plan upgrades deliberately — validate the new version against an evaluation set before switching, because behaviour can shift between versions. Auto-update settings can ease this, but for production you generally want controlled, tested upgrades rather than silent ones.

Design note

Treat the deployment name as a stable contract between your application and the platform. By pointing code at a deployment name rather than a model version, you can perform a model upgrade by changing what the name resolves to — ideally after validating the new version in a parallel deployment first.

Section 03

Endpoints, tokens, and rate limits

Once a deployment exists, your application reaches it through an endpoint — the resource's base URL — plus the deployment name and an authentication credential. The amount you can send through that endpoint, and what it costs, are both governed by tokens.

How consumption actually works

Endpoint and authentication
The endpoint is the HTTPS address of your Azure OpenAI resource. A request specifies the deployment to use and authenticates with either an API key or — preferred for production — a Microsoft Entra ID token via managed identity, so no secret is stored in the application. Role-based access control then determines whether that identity is permitted to call the model.
Tokens are the unit of everything
Text is broken into tokens — roughly four characters of English per token, though it varies. Both the input (your prompt, including any system message and retrieved context) and the output (the model's completion) are counted. Tokens drive cost (you are billed per token) and they drive capacity (quota is measured in tokens-per-minute). Counting tokens accurately is the foundation of both cost and capacity planning.
Rate limits: TPM and RPM
A standard deployment has a tokens-per-minute (TPM) limit and an associated requests-per-minute (RPM) limit, derived from the quota you assign it. Exceed either and the service returns HTTP 429. The fix is to plan adequate quota, distribute it sensibly across deployments, and handle 429s with backoff and a fallback in the client — the same reliability practices covered in Designing for Scale.
Context window versus rate limit
Do not confuse the two. The context window is the maximum number of tokens a single request can contain (input plus output) for a given model — exceed it and that one request fails. The rate limit (TPM/RPM) is how much you can send across many requests per minute. A request can be within the context window yet still be throttled because the per-minute limit is exhausted.
Reflect

For an application you are planning, do you know which authentication method it will use — API key or managed identity? If the answer is "API key, for now", note that production hardening will mean moving to Entra ID. It is far easier to design that in from the start than to retrofit it.

Section 04

Cost management

Azure OpenAI cost is dominated by token consumption, with the model choice and deployment type setting the per-token rate. Because cost flows through standard Azure billing, you have the full cost-management toolkit available — but only if you have set things up so the costs can be seen and attributed.

1. Choose the right billing model

Standard pay-as-you-go bills per token and suits variable or early-stage traffic. Provisioned throughput (PTUs) bills a fixed rate for reserved capacity and suits steady, high-volume workloads where predictable performance and cost matter more than paying only for what you use. Pick based on traffic shape, not on which sounds cheaper in the abstract.

2. Tag and attribute spend

Apply resource tags and use separate deployments or resource groups so Azure Cost Management can attribute spend to a team, project, or use case. Set budgets and alerts so a runaway loop or an unexpected spike is caught early rather than discovered on the invoice.

3. Reduce tokens where you can

Trim system prompts, retrieve only relevant context, cap completion length, and route simpler requests to smaller models. These reduce the bill directly. Caching repeated responses avoids paying twice for the same generation. The cheapest token is the one you never send.

End of Unit 3

You should now be able to:

  • Select a model and create a deployment with the right region and capacity.
  • Describe how endpoints, deployments, and authentication fit together.
  • Explain tokens, rate limits, and the levers for managing cost.
Section 05

Unit review

Question 1 of 4

What does an application actually call when it uses an Azure OpenAI model?

Question 2 of 4

Which authentication approach is preferred for a production Azure OpenAI application?

Question 3 of 4

A single request fails because it contains too many tokens for the model. Which limit has it hit?

Question 4 of 4

What is the most direct way to reduce Azure OpenAI cost for a chat workload?

End of module

You have completed Course 03: Azure OpenAI Service Overview. Next: Building With Azure AI — language, vision, speech, document intelligence, and RAG in practice.

Craig Stanley Studio · Deploy — Architecture, Integration & Operations · Azure OpenAI Service Overview · Access by direct link only.