Kirkpatrick in the AI era
The Kirkpatrick model describes four levels of learning evaluation: reaction (did learners like it?), learning (did they learn something?), behaviour (did their behaviour change?), and results (did the organisation benefit?). The model has not changed. What has changed is the complexity of measuring each level when both the learning and the work involve AI tools.
Level 2 (learning) is harder to measure when the skill is prompting and evaluation, not factual recall. Level 3 (behaviour) is harder to measure when the behaviour includes directing an AI that does part of the task. These complications are manageable — but they require deliberate design, not just a post-module survey.
By the end of this unit- Describe the specific evaluation challenges that arise at each Kirkpatrick level for AI-focused learning.
- Design a Level 2 assessment that measures prompting and evaluation skill, not just recall.
- Identify three ways Copilot Studio agents can be used to automate Level 1 and Level 2 evaluation.
The four levels — with AI complications
Level 1 — Reaction
AI-focused learning often produces a split reaction: learners who found it useful and engaged, and learners who are resistant to AI in general and bring that resistance to the learning. Standard reaction surveys conflate these two responses. Design for both.
Level 2 — Learning
Measuring whether someone can recall facts about Copilot is easy. Measuring whether they can prompt effectively, evaluate output critically, and iterate to a better result requires performance-based assessment — not multiple choice. This is the most commonly underdesigned level for AI learning.
Level 3 — Behaviour
The behaviour change question for AI learning is: are they working with Copilot differently than before? This requires observation or evidence from the work environment — not self-report. Microsoft 365 usage data, manager observation, and work product quality reviews are the primary sources.
Level 4 — Results
Attributing business results to AI adoption training is difficult but not impossible. Define your leading indicators before the programme starts: time saved on specific tasks, volume of AI-assisted outputs, error rates in AI-assisted decisions. Measure them before and after.
Measuring reaction to AI learning
The standard post-module survey ("please rate this course out of 5") is an inadequate Level 1 instrument for AI-focused learning. It measures satisfaction with the learning experience — which tells you little about whether the learning will transfer to a tool that many learners approach with anxiety or resistance.
What a useful Level 1 instrument measures for AI learning
Confidence — not just satisfaction
Intention — not just attitude
Friction — where the learning felt hard or wrong
One specific transfer intention
The Level 1 instrument should take no more than three minutes to complete. Four questions maximum. If you cannot fit the essentials into three minutes, you are measuring the wrong things.
Agents as evaluation tools
Copilot Studio agents can automate significant portions of the evaluation process — freeing evaluators to focus on the qualitative analysis that automation cannot do. Three uses are immediately practical.
Automated Level 1 follow-up
An agent connected to your LMS and calendar can send the Level 1 survey at the right moment after module completion (immediately for reaction; 30 days later for transfer intent follow-up) — without manual scheduling. It can aggregate responses and surface outliers for human review. The agent handles the logistics; you review the data.
Level 2 prompt submission and review
For performance-based Level 2 assessments (where learners submit a Copilot prompt as evidence of skill), an agent can receive the submission, run it against a defined rubric, and return structured feedback. The rubric elements — specificity, context inclusion, constraint setting, format instruction — can be assessed at scale without a human reviewing every submission. Human review is then reserved for borderline cases and qualitative insight.
Manager nudge and 30-day check-in
An agent can send a structured 30-day check-in to both the learner and their manager: "Here are the three things [Learner name] intended to do differently after completing [Module name]. Have you observed any of these changes?" This closes the feedback loop between learning intent and behavioural observation at a scale that would be impossible to manage manually.
Which of the three agent-automated evaluation tasks would have the most impact on your current evaluation practice? What would you need in place before deploying it?
The behaviour change question
Level 3 evaluation — did behaviour change? — is the most important level and the most rarely done well. For AI-focused learning, it is also the most technically tractable: Microsoft 365 generates usage data that can serve as a proxy for behavioural evidence.
Using Microsoft 365 data for Level 3 evaluation
- Copilot feature activation rate in the target cohort (before vs. after)
- Frequency of Copilot use in the specific applications covered by the module (Teams, Outlook, Word)
- Prompt length distribution — longer, more specific prompts are evidence of better delegation skill
- Manager-reported changes in work product quality or process efficiency
- Do not conclude that higher Copilot usage means better outcomes — usage without quality is a vanity metric
- Do not conclude that low post-training usage means the training failed — it may mean the learner has no use case for the features yet
- Do not use usage data alone without pairing it with a work product quality indicator or manager observation
You cannot measure change without a baseline. Before any AI training programme launches, capture:
- Current Copilot activation rate and usage frequency in the target cohort
- Manager assessment of current AI-related capability (simple 1–5 scale is sufficient)
- Learner self-assessment of confidence with each skill the training covers
These three baselines take less than 20 minutes to gather. Not gathering them makes post-training evaluation impossible.
Designing for Level 3 from the start
Level 3 evaluation cannot be designed retrospectively. The decisions that make it possible — defining the measurable behaviour indicators, establishing the baseline, building in the manager check-in — must be made before the training launches. Add a Level 3 design step to every AI learning project brief.
A training programme measures Copilot adoption rate six weeks after launch and reports a 40% increase as evidence of success. What is the most important limitation of this evaluation?
End of Unit 8
You should now be able to:
- Apply the Kirkpatrick four levels to AI-focused learning with an understanding of the specific complications each level presents.
- Design a Level 1 instrument that measures confidence and transfer intention, not just satisfaction.
- Identify three uses of Copilot Studio agents in the evaluation process.
- Build a Level 3 measurement plan before a programme launches, including baseline data collection.
Unit review
Which Kirkpatrick level is most commonly underdesigned for AI learning programmes?
What is the single best Level 1 question for predicting whether learning will transfer?
Which of the three agent-automated evaluation tasks provides the most direct evidence of behaviour change?
When must the Level 3 measurement plan be designed for it to be usable?
Series complete
You have completed the full Designing Learning for the AI Era series. You now have a complete framework for designing, producing, and evaluating learning in an organisation where Copilot, Cowork, and agents are part of everyday work.
Continue to the Resources section — prompts, Cowork skills, and agent blueprints to put this knowledge into practice immediately.