Frameworks

The Kirkpatrick model of training evaluation

The Kirkpatrick model evaluates training at four levels: reaction (did they like it), learning (did they gain anything), behaviour (are they doing the job differently) and results (did the outcome move). The levels are easy to name and hard to use, because most organisations measure level one, call it evaluation, and never find out whether the training worked.

Donald Kirkpatrick, an academic at the University of Wisconsin, set the four levels out in a series of articles for the journal of the American Society of Training Directors, later ASTD, in 1959, drawing on his doctoral research earlier that decade. He called them steps rather than a model; the branding came later, along with his 1994 book on evaluating training programmes. It survives because it asks the question sponsors actually care about: did anything change afterwards?

The four levels at a glance

LevelQuestion it answersTypical evidenceWhen to collect
1. ReactionWas it engaging and relevant?Post-session survey, facilitator observationImmediately after
2. LearningDid knowledge, skill or confidence shift?Before-and-after assessment, live demonstrationEnd of programme
3. BehaviourAre they doing the job differently?Manager and peer observation, work samplesSix weeks, then three months
4. ResultsDid the outcome it was bought for move?Measures you already collect: retention, quality, cycle time, safetyOne to two quarters later

Level 1: reaction deserves better questions

Level 1 measures the experience: engagement, relevance, whether the room stayed with the facilitator. It is not a measure of effectiveness and never was. It dominates because it is cheap, it happens while everyone is still in the room, and it produces a number for Friday's deck.

The fix is not to abandon it but to stop asking questions that only produce a warm glow. Ask what the person intends to do, and what will stop them.

High satisfaction with low intention to apply is the most useful result a level 1 survey can give you. The session worked as an event and failed as training.

Level 2: learning, measured before and after

Level 2 asks whether anything shifted: knowledge, skill, attitude, confidence, commitment. The critical design decision is a baseline. Without a before measure you cannot separate what the training added from what people already knew, and post-only assessments flatter the programme.

Post-only measurement does most damage on skill subjects. Someone can pass a quiz on active listening and still interrupt every second sentence in the next meeting. Watch the behaviour, do not test the vocabulary.

Level 3: behaviour, the level that decides everything

Level 3 asks whether the person does the job differently four, eight, twelve weeks later. This is where training either converts or evaporates, and it is the level most organisations skip, because it means going back to people after the invoice has been paid.

The uncomfortable truth of level 3 is that the training is rarely the binding constraint. People return from a good workshop to a manager who never mentions it and a system that still rewards the old behaviour. James and Wendy Kirkpatrick, who developed the New World Kirkpatrick Model from around 2009, named this well: required drivers, the coaching, accountability and reward that must exist at work for a behaviour to survive contact with it.

How to actually measure level 3

  1. Define the critical behaviours before the training runs. Two or three, observable, phrased as actions: "opens one-to-ones by asking rather than updating", not "communicates better".
  2. Ask the observer, not the participant. Self-report inflates. Managers, peers and direct reports see the change or its absence.
  3. Measure twice. At six weeks, then at three months. The first check catches early collapse, the second tells you whether it stuck.
  4. Record what got in the way. The obstacles list beats the score, because it is the part you can change next time.
  5. Build the drivers into the design, not after it. Manager briefings beforehand, a structured follow-up conversation after, one practice commitment per person.

This is also where everyday feedback quality does more work than any evaluation instrument. If nobody ever comments on the new behaviour, it has no reason to persist.

Level 4: results, and the attribution trap

Level 4 connects behaviour to the outcome the training was bought for: retention in a team that was leaking people, fewer escalations, faster onboarding, a safety record. Four rules keep it honest.

Design backwards, evaluate forwards

The most useful correction the Kirkpatricks made to their own model was the order of operations. Plan at level 4 and work down. Measure at level 1 and work up.

  1. Start with the result the sponsor wants and its leading indicators.
  2. Ask what people would have to do differently for that result to move. Those are the critical behaviours.
  3. Ask what they would have to know, do or believe to behave that way. That is the learning objective.
  4. Only then design the session, and only then decide what a good reaction would look like.

Programmes built the other way round, topic first and outcome invented afterwards, are why so much training evaluates well and changes nothing.

Where the model breaks down

Kirkpatrick is a checklist of questions, not a theory of how learning works. Four honest limitations:

  1. The causal chain is assumed, not proven. The model implies a good reaction leads to learning, then behaviour, then results. Research reviews going back decades have repeatedly failed to find a reliable link between satisfaction scores and later performance. People enjoy sessions that teach them nothing and resent sessions that change them.
  2. It ignores everything before the training. There is no level for needs analysis, participant selection, or whether this was ever a skills problem. Many programmes that fail at level 3 were solving a systems problem with a workshop.
  3. Levels 3 and 4 cost money, so they do not happen. In practice the model is honoured almost entirely at level 1. Any plan assuming unlimited follow-up capacity will be quietly abandoned.
  4. Attribution at level 4 is genuinely hard. Outside a controlled comparison you cannot separate the training's effect from everything else that quarter. Treat level 4 as evidence in an argument, not proof.

Used well, it still earns its place by forcing one conversation: what would have to be different in six months for this to have been worth doing? Ask that before you book the room and most of the evaluation design writes itself.

Kirkpatrick compared with the alternatives

ModelWhat it addsBest used when
Kirkpatrick four levelsThe base vocabulary: reaction, learning, behaviour, resultsYou need a shared language with sponsors
New World KirkpatrickRequired drivers, leading indicators, planning from level 4 downTransfer to the job is the problem
Phillips ROIA fifth level converting results into a financial returnFinance is the audience and costs are tracked
Brinkerhoff Success Case MethodInterviews with the most and least successful participantsYou want to know why it worked, not just whether
Kaufman's levelsExtends the frame to client and societal impactPublic sector or purpose-led programmes

Phillips inherits the attribution problem and adds a costing burden, so use it where outcome data is already reliable. Brinkerhoff is often the better second instrument: rather than averaging everyone, it studies the few for whom the training clearly worked and the few for whom it clearly did not. The gap between those groups is usually about the workplace, not the workshop.

A minimum viable evaluation plan

None of this replaces the thing evaluation is trying to detect. Behaviour change is visible long before a survey catches it, which is why experiential learning and evaluation fit together: put a team into a real challenge under real pressure and level 3 happens in front of you, unprompted. Who reverts to old habits when time gets short, who tries the new approach and drops it the moment it gets uncomfortable, who quietly coaches someone else through it. Read honestly, that tells you more about whether the training survives Monday than any post-course score. The four levels then give you the language to report it.

Tour De Force runs this as live, experiential training for teams worldwide — online and in person. Talk to us, or play Gamified learning appsThe Weekly Challenge to see the method in ten minutes.

Questions

The Kirkpatrick model of training evaluation FAQs

What are the four levels of the Kirkpatrick model?

Reaction, learning, behaviour and results. Reaction measures how participants experienced the training, learning measures the change in knowledge, skill, attitude or confidence, behaviour measures what people do differently at work weeks later, and results measures the business outcome the training was intended to affect.

Who created the Kirkpatrick model and when?

Donald Kirkpatrick, an academic at the University of Wisconsin, set the four steps out in a series of articles for the journal of the American Society of Training Directors, later ASTD, in 1959, based on his doctoral research. He expanded them into a book in 1994. James and Wendy Kirkpatrick later developed the New World Kirkpatrick Model from around 2009.

What is the difference between Kirkpatrick level 3 and level 4?

Level 3 measures behaviour: what individuals do differently in the job after training. Level 4 measures results: the organisational outcome that behaviour is supposed to produce, such as retention, quality or cycle time. Level 3 is about people, level 4 is about numbers, and level 3 is the one you cannot skip without losing the causal story.

What is the difference between the Kirkpatrick model and the Phillips ROI model?

Phillips adds a fifth level that converts level 4 results into a financial return by comparing monetised benefits with programme costs. It is useful when finance is the audience, but it inherits Kirkpatrick's attribution problem and adds a costing burden, so it works best where outcome data is already reliable and well tracked.

How do you measure behaviour change after training?

Define two or three observable behaviours before the training runs, then ask observers rather than participants. Managers, peers or direct reports report what they see at around six weeks and again at three months. Record the obstacles people hit as well as the score, because those obstacles are usually what you change next time.

Can you use the Kirkpatrick model for soft skills training?

Yes, but shift the weight to levels 2 and 3 and measure by demonstration rather than testing. Someone can pass a quiz on listening or feedback and behave no differently in a real conversation. Use observed practice for level 2 and manager or peer observation for level 3, with the critical behaviours agreed in advance.

Keep reading

Related reading

Get in touch

Build it with your team.

Book a 30-minute discovery call and we'll shape an experiential programme around your goals.

Book a discovery call