The service in context
Azure OpenAI, now offered within Azure AI Foundry, gives you OpenAI's models — GPT-4 family, embeddings, and others — running inside your Azure tenant with enterprise controls: Microsoft Entra ID authentication, private networking, data residency, and content safety. The model capability is the same OpenAI you may know from elsewhere; what Azure adds is the governance, security, and commercial framework that an organisation needs to run it responsibly.
This unit walks the full path from choosing a model to consuming it in an application: the model catalogue, deployments, endpoints, how tokens are counted, how rate limits work, and how to keep the cost under control.
By the end of this unit- Navigate the model catalogue and create a deployment from a base model.
- Explain the relationship between an endpoint, a deployment, and authentication.
- Describe how tokens, rate limits, and quota drive both throughput and cost.
Azure OpenAI versus OpenAI directly
- Authenticate with Microsoft Entra ID and managed identity, not just API keys
- Role-based access control governs who can deploy and call models
- Private endpoints keep traffic off the public internet
- Choose the Azure region your deployment runs in
- Your prompts and completions are not used to train the base models
- Content and abuse-monitoring controls are configurable to your policy
- Billed through your existing Azure agreement and cost-management tooling
- Standard pay-as-you-go or reserved provisioned throughput
- Cost attributable to subscriptions, resource groups, and tags
Model catalogue and deployments
In Azure AI Foundry, the model catalogue is where you browse and select models — from the GPT family to embedding models, plus open and partner models. Selecting a model is not the same as using it. To call a model, you create a deployment: a named instance of a model, in a region, with an assigned capacity. Your application then talks to the deployment, not to the model directly.
1. Choose a model from the catalogue
Match the model to the task. Chat and reasoning tasks use a GPT chat model; semantic search and RAG retrieval use an embedding model; some workloads combine both. Consider quality, latency, context window, and cost — the largest model is not always the right default. The catalogue shows the capabilities and the regions where each model is available.
2. Create a deployment
A deployment binds a model version to a region and a capacity allocation (TPM for standard, or PTUs for provisioned). You give it a deployment name — the identifier your application code references. Keeping deployment names stable lets you swap the underlying model version behind the same name during an upgrade.
3. Manage model versions and lifecycle
Models are versioned and have a lifecycle: new versions are released, older ones are retired on a published schedule. Plan upgrades deliberately — validate the new version against an evaluation set before switching, because behaviour can shift between versions. Auto-update settings can ease this, but for production you generally want controlled, tested upgrades rather than silent ones.
Treat the deployment name as a stable contract between your application and the platform. By pointing code at a deployment name rather than a model version, you can perform a model upgrade by changing what the name resolves to — ideally after validating the new version in a parallel deployment first.
Endpoints, tokens, and rate limits
Once a deployment exists, your application reaches it through an endpoint — the resource's base URL — plus the deployment name and an authentication credential. The amount you can send through that endpoint, and what it costs, are both governed by tokens.
How consumption actually works
Endpoint and authentication
Tokens are the unit of everything
Rate limits: TPM and RPM
Context window versus rate limit
For an application you are planning, do you know which authentication method it will use — API key or managed identity? If the answer is "API key, for now", note that production hardening will mean moving to Entra ID. It is far easier to design that in from the start than to retrofit it.
Cost management
Azure OpenAI cost is dominated by token consumption, with the model choice and deployment type setting the per-token rate. Because cost flows through standard Azure billing, you have the full cost-management toolkit available — but only if you have set things up so the costs can be seen and attributed.
1. Choose the right billing model
Standard pay-as-you-go bills per token and suits variable or early-stage traffic. Provisioned throughput (PTUs) bills a fixed rate for reserved capacity and suits steady, high-volume workloads where predictable performance and cost matter more than paying only for what you use. Pick based on traffic shape, not on which sounds cheaper in the abstract.
2. Tag and attribute spend
Apply resource tags and use separate deployments or resource groups so Azure Cost Management can attribute spend to a team, project, or use case. Set budgets and alerts so a runaway loop or an unexpected spike is caught early rather than discovered on the invoice.
3. Reduce tokens where you can
Trim system prompts, retrieve only relevant context, cap completion length, and route simpler requests to smaller models. These reduce the bill directly. Caching repeated responses avoids paying twice for the same generation. The cheapest token is the one you never send.
Build hands-on familiarity with the service:
Get started with Azure AI Foundry ↗
Build AI apps and agents with Azure AI Foundry ↗
Fine-tune models with Azure AI Foundry ↗
End of Unit 3
You should now be able to:
- Select a model and create a deployment with the right region and capacity.
- Describe how endpoints, deployments, and authentication fit together.
- Explain tokens, rate limits, and the levers for managing cost.
Unit review
What does an application actually call when it uses an Azure OpenAI model?
Which authentication approach is preferred for a production Azure OpenAI application?
A single request fails because it contains too many tokens for the model. Which limit has it hit?
What is the most direct way to reduce Azure OpenAI cost for a chat workload?
End of module
You have completed Course 03: Azure OpenAI Service Overview. Next: Building With Azure AI — language, vision, speech, document intelligence, and RAG in practice.