Craig Stanley
Home / Capabilities / Foundry / Choosing a deployment type in Foundry

Choosing a deployment type in Foundry

Foundry deployment types decide where your prompts are processed and how you pay. A plain guide to Global, Data Zone and regional options.

· 3 min read · Craig Stanley
In short, explained

When you switch on an AI model in Foundry, you choose where in the world it does its thinking and whether you pay per use or rent space in advance.

Each Foundry model deployment has a type. The type sets where prompts are processed (anywhere, within a region like the EU, or in one Azure geography), how you pay (per token, reserved capacity, or cheaper batch jobs), and how steady the response times are.

Serverless API deployment types combine processing scope (Global, Data Zone US/EU/APAC, regional geography) with billing (Standard pay-per-token, Provisioned hourly reserved, Batch ~50% cheaper, 24-hour target). Data at rest stays in the resource's geography for all types. Global Standard is Microsoft's default; Provisioned reduces latency variance.

Two questions in one choice

Microsoft's documentation says the deployment type decides three things: where your data is processed, how you pay, and performance characteristics. I find it easier to think of it as two questions.

Where can my prompts and responses be processed? And do I want to pay per use, reserve capacity, or send big jobs to run later at a lower price?

Where it's processed

Microsoft gives three options.

ScopeWhat Microsoft says about processing
GlobalMay be processed in any Azure region where the model is deployed
Data ZoneProcessed only within a Microsoft-defined zone: US, EU or Asia Pacific
Regional (Standard or Regional Provisioned)Processed within the Azure geography you choose

For every type, Microsoft says data stored at rest stays in the designated Azure geography. Only the place where inference happens changes.

Two details are worth knowing. The EU data zone follows the Azure EU Data Boundary, which can include EFTA countries such as Norway and Switzerland, and Microsoft says it can add regions to a data zone without prior notice. For a UK organisation, the EU zone isn't the UK; if processing must stay in the UK, you'd need a regional deployment in a UK region, if the model you want is offered there.

How you pay

BillingHow it works
StandardPay per token
ProvisionedHourly reserved capacity, with lower latency variance
BatchAsynchronous jobs with a 24-hour target turnaround, which Microsoft says cost 50% less than Global Standard

Microsoft notes that Global and Data Zone Standard can show more latency variation at high, steady volumes, and recommends provisioned types where low variance matters. Global deployments also get new models and features first.

Microsoft's quick guide

Microsoft's own recommendation table starts with Global Standard as the default for "newest models, lowest price, broadest regions". It recommends Data Zone types to keep processing within a zone, regional types to keep it within a geography, provisioned types for predictable throughput, and batch for large jobs that aren't urgent.

Why this is a decision in itself

The deployment type is a trade-off between cost, model choice and where data goes. It's the kind of decision that's easy to make once in a hurry and hard to revisit. I'd write it down with the reason, so the next person knows whether "Global" was a choice or a default.

A worked example

This scenario is illustrative. A UK charity wants a model to summarise 20,000 archived case notes once, and then to help staff draft replies to new enquiries every day.

WorkloadConstraintChoice
One-off archive summariesNot urgent; data protection review allows EU processingData Zone Batch, if available for the model
Daily reply draftingStaff are waiting; UK processing required by policyRegional Standard in a UK region, if the model is offered there

The same model could end up with two deployments for two decisions. If the daily model isn't offered in a UK region, that's a reason to revisit the model choice before the processing rule.

What I'm still checking

Model availability differs by region and deployment type, and Microsoft keeps a separate region availability page for each category. I haven't checked which current models are offered as regional deployments in UK regions. That's the first thing I'd look up for a real project.

Sources

Read next

A question to take awayWhich of these do you already pay for and not use?

About me

Craig Stanley

Microsoft AI consultant and technical architect, based in Whitley Bay. Over the last few years I've delivered Microsoft 365 Copilot, Copilot Studio agents, Microsoft Foundry (formerly Azure AI Foundry) work and governance for UK public sector and financial services organisations.

What interests me is the decision underneath the tool: what it costs, what it risks, and whether a small, transparent model can make it better. I write the methods up here and on Substack so anyone can use them.

I write this site to learn in public: explaining each idea simply is how I check I understand it. Why I write this site.

Find me