The Azure AI service family
Generative models are not the whole of Azure AI. Alongside Azure OpenAI sits a family of task-specific services — language, vision, speech, and document intelligence — that solve well-defined problems reliably and cheaply. A capable solution often combines them: document intelligence to extract structured data from a scanned form, a language service to classify it, and a generative model to summarise the result. Knowing what each service does well keeps you from reaching for a large language model when a purpose-built service would be faster, cheaper, and more accurate.
This unit surveys those services, then goes deep on the pattern most production solutions need: retrieval-augmented generation (RAG) with Azure AI Search, which grounds a generative model in your own data.
By the end of this unit- Match a problem to the right Azure AI service rather than defaulting to a generative model.
- Describe the RAG pattern and the role Azure AI Search plays in grounding responses.
- Identify the steps that turn a RAG prototype into a production-grade solution.
Generative versus task-specific services
- Open-ended generation: drafting, summarising, answering in natural language
- Reasoning over context you supply at request time
- Tasks where the output format and content are flexible
- A well-defined, repeatable task with a known output shape
- Extracting fields from a form, transcribing speech, detecting objects
- When accuracy, cost, and latency on that narrow task all matter
- Document intelligence extracts → language service classifies → model summarises
- Speech-to-text transcribes → model answers questions about the transcript
- The task-specific service does the structured work; the model does the open-ended work
Language, vision, speech, documents
The Azure AI services group purpose-built capabilities behind consistent APIs and the same enterprise controls as Azure OpenAI. Here are the four you will reach for most often.
Azure AI Language
Text analytics: entity recognition, key-phrase extraction, sentiment, language detection, summarisation, and custom classification. Use it to structure unstructured text — routing support tickets, tagging documents, or extracting named entities — without invoking a generative model for every record.
Azure AI Vision
Image analysis: object and tag detection, optical character recognition (OCR) for text in images, and image captioning. Pair it with generative models for multimodal scenarios, or use it standalone where the task is a defined classification or extraction.
Azure AI Speech
Speech-to-text (transcription), text-to-speech (natural voices), and speech translation. Common in contact-centre and accessibility scenarios. A frequent pattern: transcribe a call with Speech, then use a generative model to summarise it and extract action items.
Azure AI Document Intelligence
Extracts structured data from documents — invoices, receipts, forms, contracts — including key-value pairs, tables, and layout. Prebuilt models cover common document types; custom models can be trained for your own forms. Far more reliable than asking a language model to parse a scanned PDF.
The deciding question is whether the task is open-ended or well-defined. Open-ended generation and reasoning are a generative model's strength. A repeatable extraction or classification with a known output shape is usually better served by a task-specific service — and the two compose cleanly in a pipeline.
RAG with Azure AI Search
A generative model knows only what it was trained on and what you give it at request time. Retrieval-augmented generation (RAG) is the pattern for grounding a model in your own current data: retrieve the most relevant content, supply it to the model as context, and have the model answer using that content. Azure AI Search is the retrieval engine in the Microsoft stack.
How the pattern works
Index your content
Retrieve at query time
Generate the grounded answer
Better retrieval beats more retrieval
Think of a question your users currently cannot get answered from your systems. What corpus would need to be indexed to answer it, and how current does that index need to be? RAG quality starts with the content you index and how fresh you keep it — the model is the last step, not the first.
From prototype to production-grade
A RAG prototype can be stood up in an afternoon. A production-grade one accounts for evaluation, freshness, safety, and the operational realities the prototype ignored. Walk these four before you ship.
1. Evaluate, do not eyeball
Build an evaluation set of representative questions with known good answers, and measure groundedness, relevance, and accuracy systematically. Azure AI Foundry's evaluation tooling supports this. Eyeballing a few demo questions hides the failure cases that real users will find on day one.
2. Keep the index fresh
Decide how the index stays current as source content changes — scheduled re-indexing, incremental updates, or change-driven pipelines. A RAG system answering from stale content is confidently wrong, which is worse than admitting it does not know.
3. Apply content safety and grounding guards
Use Azure AI Content Safety to filter harmful inputs and outputs, and instruct the model to refuse when the answer is not in the retrieved context rather than inventing one. Pass citations through so answers are verifiable. Safety and groundedness are production requirements, not polish.
4. Handle scale and failure
Apply the reliability and capacity practices from earlier units — retries, fallbacks, quota planning, cost attribution — to both the search service and the model. A RAG solution has two dependencies that can throttle or fail, and both need a planned response.
Go deeper on building and grounding Azure AI solutions:
Implement Retrieval Augmented Generation with Azure AI ↗
Build AI apps and agents with Azure AI Foundry ↗
Develop AI agents with Azure AI Foundry (Applied Skill) ↗
End of Unit 4
You should now be able to:
- Choose between a generative model and a task-specific Azure AI service.
- Explain the RAG pattern and Azure AI Search's role in it.
- List what turns a RAG prototype into a production-grade solution.
Unit review
You need to extract the line items and totals from thousands of scanned invoices. Which service fits best?
What is the core purpose of the RAG pattern?
In a RAG solution, what generally gives the biggest improvement in answer quality?
Which step is essential to move a RAG prototype to production?
End of module
You have completed Course 04: Building With Azure AI. Next: Azure AI in Production — monitoring, resilience, safe deployment, and cost governance at scale.