Section 01 · Unit introduction

Why pipelines decide AI quality

An AI workload is only as good as the data flowing into it. A retrieval system grounded on stale or duplicated documents will give confident wrong answers. A model trained on data with silent quality problems will learn those problems. The pipeline — the machinery that moves, cleans, and shapes data — is where AI quality is quietly won or lost.

On the Microsoft platform the main tools are Azure Data Factory, Azure Synapse Analytics, and Microsoft Fabric. They share a common vocabulary: ingest data, transform it through stages, apply quality gates, and run it on a dependable schedule. This unit covers the patterns that make those pipelines reliable.

By the end of this unit
  • Explain why pipeline quality, not model choice, often determines AI output quality.
  • Apply the medallion (bronze/silver/gold) layering pattern to an ingestion design.
  • Place quality gates and scheduling so that bad data is caught before it reaches an AI workload.

The Microsoft data-pipeline tools

  • Cloud data-integration service for orchestrating movement and transformation
  • Connects to a wide range of sources with copy activities and data flows
  • Strong for hybrid scenarios and lift-and-shift ETL
  • Pipelines, triggers, and integration runtimes are its core concepts
  • Unified analytics combining data integration, warehousing, and Spark
  • Pipelines reuse the Data Factory engine within the Synapse workspace
  • Good when you need warehousing and big-data processing together
  • Being superseded by Fabric for many new builds
  • The current unified SaaS analytics platform built on OneLake
  • Data Factory pipelines, dataflows, notebooks, and lakehouses in one place
  • The default choice for most new AI data foundations on the platform
  • Covered in depth in the next course
Nobody is impressed by a sophisticated model fed on dirty data. The pipeline is where the real reliability work happens.
Working principle · Data engineering for AI
Section 02

Ingestion and the medallion pattern

The medallion architecture is a simple, durable way to organise data as it moves from raw to ready. Data flows through three layers — bronze, silver, and gold — each with a clear contract about what it contains.

Bronze — raw, as it arrived

Land source data unchanged. Keep it exactly as the source produced it, with ingestion metadata such as load time and source identifier. Bronze is your record of truth and your ability to reprocess if a downstream bug is found. Never overwrite history here.

Silver — cleaned and conformed

Apply validation, deduplication, type enforcement, and joins. Silver data is trustworthy and consistent — the layer most transformations read from. This is where most quality gates live.

Gold — shaped for consumption

Aggregate and model the data for a specific purpose: a feature table for a model, a curated document set for retrieval, a reporting mart. Gold is purpose-built and may be regenerated freely from silver.

Design note

The discipline that makes medallion work is the one-way flow: bronze feeds silver, silver feeds gold, and you never write business logic that reads gold back into silver. When something breaks, you can rebuild any layer from the one before it, with bronze as the ultimate fallback.

Incremental versus full loads
Reloading everything every run is simple but does not scale. For large sources, capture only what changed since the last run using a watermark column or change-data-capture. Incremental loading is the difference between a pipeline that runs in minutes and one that times out.
Schema drift will happen
Sources add and rename columns without warning. Decide in advance how the pipeline reacts: fail loudly, quarantine the affected rows, or map the change. Silent acceptance of schema drift is how bad data reaches gold unnoticed.
Section 03

Transformation and quality gates

A quality gate is a checkpoint that data must pass before it is promoted to the next layer. Without gates, problems travel silently to the AI workload and surface as inexplicable bad answers weeks later.

Completeness checks

Are the required fields present and non-null at the rate you expect? A sudden jump in nulls usually means an upstream change, not a real shift in the world.

Validity and range checks

Do values fall within plausible bounds? Dates in the future, negative quantities, and impossible codes should be caught and quarantined, not averaged into a gold table.

Uniqueness and referential checks

Are keys unique where they should be? Do foreign keys resolve? Duplicates are especially corrosive to retrieval workloads, where the same chunk returned twice wastes context and skews relevance.

Volume and freshness checks

Is the row count in the expected range, and is the data recent enough? A run that loads a tenth of the usual volume has probably failed partway, even if no error was raised.

Reflect

Think about an AI workload you support or plan to build. If the upstream source silently halved its row count tomorrow, would anything in your pipeline notice before the AI started giving thin answers? If not, that is your first quality gate to build.

Knowledge check

A retrieval-augmented generation system has started returning the same passage twice in its grounding context. Where in the pipeline is the most likely cause?

Section 04

Scheduling and reliability

A pipeline that works once in a demo is not the deliverable. The deliverable is a pipeline that runs on schedule, recovers from transient failures, and tells someone when it cannot.

Triggers — schedule, tumbling window, and event
Data Factory and Fabric support scheduled triggers, tumbling-window triggers for time-sliced batches, and storage-event triggers that fire when a file lands. Match the trigger to the source: nightly batch, time-windowed processing, or react-on-arrival.
Retries and transient faults
Network blips and momentary throttling are normal in cloud pipelines. Configure retry policies on activities so a transient fault does not fail the whole run. Distinguish transient faults (retry) from data faults (quarantine and alert) — retrying bad data forever helps no one.
Idempotent reruns
When a run fails halfway and you rerun it, the result must be correct, not doubled. Design loads so that reprocessing the same window produces the same end state — using merge or upsert logic rather than blind append.
Monitoring and alerting
A failed pipeline that no one is told about is the worst outcome, because downstream consumers keep trusting stale data. Wire failures and quality-gate breaches to alerts that reach a named owner. Silent failure is the enemy.
Continue on Microsoft Learn

Build the foundations with these paths:

Get started with Microsoft Fabric ↗
Prepare data for analysis with Power BI ↗

End of Unit 13

You should now be able to:

  • Argue why pipeline quality drives AI output quality.
  • Structure ingestion using the bronze/silver/gold medallion pattern.
  • Place quality gates, triggers, retries, and alerting to make a pipeline production-ready.
Section 05

Unit review

Question 1 of 4

In the medallion pattern, what is the defining purpose of the bronze layer?

Question 2 of 4

Why prefer incremental loads over full reloads for a large source?

Question 3 of 4

A pipeline run loads only a tenth of the usual row count but raises no error. Which quality gate would catch this?

Question 4 of 4

What distinguishes a transient fault from a data fault, and why does it matter?

End of module

You have completed Course 13: Data Pipelines for AI. Next: Fabric and AI Together — using Microsoft Fabric and OneLake as the data foundation for AI applications.

Craig Stanley Studio · Deploy — Architecture, Integration & Operations · Data Pipelines for AI · Access by direct link only.