Fabric as the data foundation
Microsoft Fabric is a unified, software-as-a-service analytics platform. It brings data integration, lakehouses, warehousing, notebooks, real-time intelligence, and Power BI together on a single foundation called OneLake. For AI work, Fabric matters because it removes the seams between where data lands, where it is shaped, and where it is served to an application.
Previous generations of analytics stacks scattered data across separate services that each kept their own copy. Fabric's premise is one logical data lake, one copy, many engines. That has direct consequences for how you build the data foundation behind an AI application.
By the end of this unit- Describe how OneLake, lakehouses, and Delta tables form a single data foundation for AI.
- Explain how Fabric notebooks and real-time intelligence fit into an AI data workflow.
- Outline how data shaped in Fabric is served to an AI application such as a retrieval workload or Copilot.
The Fabric building blocks
- One logical data lake for the whole tenant — the "OneDrive for data"
- Built on Delta Lake and the open Parquet format underneath
- Every Fabric workload reads and writes the same copy of the data
- Shortcuts let you reference data in other stores without copying it
- Combines a data lake's files with a relational table layer
- Tables are Delta; files area holds raw and semi-structured data
- Queryable with Spark (notebooks) and SQL (the SQL analytics endpoint)
- The natural home for the medallion layers on Fabric
- Spark notebooks in Python, Scala, SQL, or R
- Where transformation, feature engineering, and ML work happen
- Read and write OneLake tables directly
- Schedulable as part of a Data Factory pipeline
- Ingest and query streaming and event data at low latency
- Eventstreams, eventhouses, and KQL for high-volume telemetry
- Feeds near-real-time signals into AI and alerting
- Complements the batch lakehouse, not a replacement
OneLake, lakehouses, and Delta
Three concepts do most of the work in a Fabric data foundation: OneLake as the single store, the lakehouse as the structured container within it, and Delta as the table format that makes lake data behave reliably.
OneLake — one store, no copies
OneLake is automatically provisioned for the tenant. Every workspace and every workload writes into it. Because there is one copy, a table written by a notebook is immediately queryable by SQL and visible to Power BI without an export or a sync.
Shortcuts — reference, don't duplicate
A shortcut points OneLake at data living elsewhere — another lakehouse, Azure Data Lake Storage, or Amazon S3 — so you can use it in Fabric without copying it. This is the antidote to the old habit of duplicating data into every system that needs it.
Delta tables — reliability over a lake
Delta Lake adds ACID transactions, schema enforcement, and time travel to files in the lake. That means a half-finished write does not leave a corrupt table, schema changes are controlled, and you can query the table as it stood at a previous point — invaluable for reproducing what an AI workload saw.
Time travel and ACID guarantees let you answer "what data was the model grounded on last Tuesday?" with precision. Reproducibility is a core requirement for trustworthy AI, and Delta gives it to you at the storage layer rather than as a bolt-on.
Medallion layers as lakehouse tables
The SQL analytics endpoint
Notebooks and real-time intelligence
Notebooks are where the active work of preparing data for AI happens, and real-time intelligence is how Fabric handles the streaming signals that batch pipelines cannot.
Notebooks in the workflow
A Fabric notebook runs Spark over OneLake data. For an AI workload you might use one to chunk and clean documents before they are embedded, to engineer features for a model, or to run the transformation that promotes silver tables to a gold feature table. Notebooks slot into a Data Factory pipeline as a scheduled activity, so the exploratory work and the production run use the same code.
Real-time intelligence
Some AI signals cannot wait for the nightly batch — fraud indicators, equipment telemetry, live engagement. Real-Time Intelligence ingests event streams into an eventhouse and queries them with KQL at low latency. The results can feed an alerting flow or be landed into the lakehouse for the batch world to use later.
For an AI use case you care about, which signals are genuinely real-time and which only feel urgent? Be honest. Most data is comfortably served by a batch lakehouse; reserve real-time intelligence for the signals that lose their value within minutes.
A notebook writes a cleaned gold table to a lakehouse. A colleague wants to query it from a SQL-based reporting tool. What must they do to copy the data into a separate warehouse first?
Feeding AI applications
Once data is shaped and trustworthy in a gold layer, the final step is serving it to the AI application. Fabric connects to AI in several complementary ways.
Grounding a retrieval workload
Copilot in Power BI over Fabric data
Feature tables for model training and scoring
Governance travels with the data
Go deeper with these official resources:
Get started with Microsoft Fabric ↗
Implement analytics in Microsoft Fabric (Applied Skill) ↗
Use Copilot in Power BI ↗
End of Unit 14
You should now be able to:
- Explain the one-copy-many-engines model of OneLake and why it suits AI.
- Use lakehouses, Delta tables, notebooks, and real-time intelligence in an AI data workflow.
- Describe how governed gold data is served to retrieval, Copilot, and model workloads.
Unit review
What does a OneLake shortcut let you do?
Why is Delta's time-travel capability valuable for trustworthy AI?
When is Real-Time Intelligence the right tool rather than the batch lakehouse?
What is the governance advantage of building the AI data foundation on OneLake?
End of module
You have completed Course 14: Fabric and AI Together. Next: Testing AI Outputs — systematic evaluation, red-teaming, and quality gates for AI-generated content.