Section 01 · Unit introduction

Fabric as the data foundation

Microsoft Fabric is a unified, software-as-a-service analytics platform. It brings data integration, lakehouses, warehousing, notebooks, real-time intelligence, and Power BI together on a single foundation called OneLake. For AI work, Fabric matters because it removes the seams between where data lands, where it is shaped, and where it is served to an application.

Previous generations of analytics stacks scattered data across separate services that each kept their own copy. Fabric's premise is one logical data lake, one copy, many engines. That has direct consequences for how you build the data foundation behind an AI application.

By the end of this unit
  • Describe how OneLake, lakehouses, and Delta tables form a single data foundation for AI.
  • Explain how Fabric notebooks and real-time intelligence fit into an AI data workflow.
  • Outline how data shaped in Fabric is served to an AI application such as a retrieval workload or Copilot.

The Fabric building blocks

  • One logical data lake for the whole tenant — the "OneDrive for data"
  • Built on Delta Lake and the open Parquet format underneath
  • Every Fabric workload reads and writes the same copy of the data
  • Shortcuts let you reference data in other stores without copying it
  • Combines a data lake's files with a relational table layer
  • Tables are Delta; files area holds raw and semi-structured data
  • Queryable with Spark (notebooks) and SQL (the SQL analytics endpoint)
  • The natural home for the medallion layers on Fabric
  • Spark notebooks in Python, Scala, SQL, or R
  • Where transformation, feature engineering, and ML work happen
  • Read and write OneLake tables directly
  • Schedulable as part of a Data Factory pipeline
  • Ingest and query streaming and event data at low latency
  • Eventstreams, eventhouses, and KQL for high-volume telemetry
  • Feeds near-real-time signals into AI and alerting
  • Complements the batch lakehouse, not a replacement
One copy of the data, many engines over it. That is the idea behind Fabric — and it is what makes it a clean foundation for AI.
Working principle · Unified data foundation
Section 02

OneLake, lakehouses, and Delta

Three concepts do most of the work in a Fabric data foundation: OneLake as the single store, the lakehouse as the structured container within it, and Delta as the table format that makes lake data behave reliably.

OneLake — one store, no copies

OneLake is automatically provisioned for the tenant. Every workspace and every workload writes into it. Because there is one copy, a table written by a notebook is immediately queryable by SQL and visible to Power BI without an export or a sync.

Shortcuts — reference, don't duplicate

A shortcut points OneLake at data living elsewhere — another lakehouse, Azure Data Lake Storage, or Amazon S3 — so you can use it in Fabric without copying it. This is the antidote to the old habit of duplicating data into every system that needs it.

Delta tables — reliability over a lake

Delta Lake adds ACID transactions, schema enforcement, and time travel to files in the lake. That means a half-finished write does not leave a corrupt table, schema changes are controlled, and you can query the table as it stood at a previous point — invaluable for reproducing what an AI workload saw.

Why this matters for AI

Time travel and ACID guarantees let you answer "what data was the model grounded on last Tuesday?" with precision. Reproducibility is a core requirement for trustworthy AI, and Delta gives it to you at the storage layer rather than as a bolt-on.

Medallion layers as lakehouse tables
The bronze/silver/gold pattern from the previous course maps cleanly onto Fabric: bronze, silver, and gold become sets of Delta tables in one or more lakehouses, transformed by notebooks or dataflows between layers. The pattern does not change; Fabric just gives it one consistent home.
The SQL analytics endpoint
Every lakehouse exposes a read-only SQL endpoint over its Delta tables. Analysts and applications that speak SQL can query the same data the notebooks produced, without anyone moving it into a separate warehouse first.
Section 03

Notebooks and real-time intelligence

Notebooks are where the active work of preparing data for AI happens, and real-time intelligence is how Fabric handles the streaming signals that batch pipelines cannot.

Notebooks in the workflow

A Fabric notebook runs Spark over OneLake data. For an AI workload you might use one to chunk and clean documents before they are embedded, to engineer features for a model, or to run the transformation that promotes silver tables to a gold feature table. Notebooks slot into a Data Factory pipeline as a scheduled activity, so the exploratory work and the production run use the same code.

Real-time intelligence

Some AI signals cannot wait for the nightly batch — fraud indicators, equipment telemetry, live engagement. Real-Time Intelligence ingests event streams into an eventhouse and queries them with KQL at low latency. The results can feed an alerting flow or be landed into the lakehouse for the batch world to use later.

Reflect

For an AI use case you care about, which signals are genuinely real-time and which only feel urgent? Be honest. Most data is comfortably served by a batch lakehouse; reserve real-time intelligence for the signals that lose their value within minutes.

Knowledge check

A notebook writes a cleaned gold table to a lakehouse. A colleague wants to query it from a SQL-based reporting tool. What must they do to copy the data into a separate warehouse first?

Section 04

Feeding AI applications

Once data is shaped and trustworthy in a gold layer, the final step is serving it to the AI application. Fabric connects to AI in several complementary ways.

Grounding a retrieval workload
For a retrieval-augmented generation system, a notebook reads the curated gold documents, chunks them, generates embeddings, and writes them to a vector index. Because the source is a governed Delta table, you always know exactly what corpus the system is grounded on and when it last refreshed.
Copilot in Power BI over Fabric data
Power BI sits natively on Fabric, and Copilot in Power BI can answer questions and build reports over a well-modelled semantic layer. A clean gold model is what makes those AI answers reliable rather than confidently wrong.
Feature tables for model training and scoring
Gold feature tables produced by notebooks feed model training and batch scoring. Keeping features as governed Delta tables means training and inference read consistent, versioned data — avoiding the training-serving skew that quietly degrades model quality.
Governance travels with the data
Sensitivity labels, lineage, and permissions in Fabric apply to the data the AI consumes. Governing the foundation is far easier than trying to govern every downstream AI application separately. Govern once, at OneLake, and the AI inherits it.

End of Unit 14

You should now be able to:

  • Explain the one-copy-many-engines model of OneLake and why it suits AI.
  • Use lakehouses, Delta tables, notebooks, and real-time intelligence in an AI data workflow.
  • Describe how governed gold data is served to retrieval, Copilot, and model workloads.
Section 05

Unit review

Question 1 of 4

What does a OneLake shortcut let you do?

Question 2 of 4

Why is Delta's time-travel capability valuable for trustworthy AI?

Question 3 of 4

When is Real-Time Intelligence the right tool rather than the batch lakehouse?

Question 4 of 4

What is the governance advantage of building the AI data foundation on OneLake?

End of module

You have completed Course 14: Fabric and AI Together. Next: Testing AI Outputs — systematic evaluation, red-teaming, and quality gates for AI-generated content.

Craig Stanley Studio · Deploy — Architecture, Integration & Operations · Fabric and AI Together · Access by direct link only.