The Evaluation We Can Do On Our Own
I have spent the last few weeks marking Level 7 CIPD assignments in strategic learning and development. The submissions came from senior practitioners writing about their own organisations. The sectors, the scale and the problems were all different, and yet the answer to one question was almost identical in every case.
Asked how they measure the impact and transfer of learning, nearly every writer described the same practice. Reaction data is collected consistently, knowledge checks are carried out sometimes, and beyond that evaluation thins out and stops. One writer put it plainly, observing that evaluation concentrates where the evidence is easiest to collect.
This is not a knowledge problem. These are practitioners who can name Kirkpatrick's four levels without looking them up, and several of them explained in writing why a satisfied learner is not a changed one before going on to describe why their own organisation stops at reaction data anyway. The CIPD's Learning at Work 2023 survey found that 24% of respondents agreed their organisation assesses the impact of learning and development (Overton, 2023). If the constraint is not knowledge, it is worth asking what it actually is.
Evaluation stops where our authority stops
Levels 1 and 2 can be collected by learning and development on our own. We are in the room, we control the moment, we hand out the form or build the assessment, and the result belongs to us. Levels 3 and 4 cannot be collected that way. Behavioural change happens in a workplace we do not control, weeks after we have left it, and it is observed by line managers whose time we do not own, against performance data held by another business area. Each of those steps requires somebody else to do something they were probably never asked to do when the intervention was commissioned.
It is my view that this, rather than any failure of rigour, intent or time, is what explains where evaluation stops. Structure shapes behaviour and behaviour shapes culture, so when something is not working the system is the first place to look and not the people. The same CIPD survey found that 36% of respondents agreed line managers support their teams in transferring learning back into the workplace, and 29% agreed line managers are involved in assessing its impact (Overton, 2023). If transfer sits with the line manager and measurement sits with learning and development, there is a handover in the middle that nobody owns.
The model tells us what happened rather than why
There is a second problem underneath the first. Holton (1996), writing in Human Resource Development Quarterly, argued that Kirkpatrick's four levels are a taxonomy of outcomes rather than an evaluation model, in that they describe what happened but explain very little about why learning did or did not transfer.
This matters in practice. If you do the difficult thing, reach Level 3, observe the workplace and find that little has changed, the framework has given you a result but it has not told you whether the cause lies in the design, the learner, the line manager, the opportunity to practise, or an organisation that never made room for the new behaviour. You are left with a finding that you cannot act on.
Alliger and Janak (1989) had raised a related concern earlier, questioning the assumption that the four levels form a causal chain in which each level leads to the next. Alliger returned to the question with a different set of colleagues eight years later, and their meta-analysis found that utility reactions, meaning whether participants judged the session useful for their job, correlated with subsequent job performance at .18, compared with .07 for affective reactions (Alliger et al., 1997). Both figures rest on a small number of studies and I would not place weight on the precise values, but the direction is useful. Whether people enjoyed the day tells us very little, whereas whether they believed they could use it tells us rather more.
A narrow shelf
One further pattern was only visible because I was reading a whole cohort at once. Asked to compare two different approaches to measuring impact and transfer, most writers reached for a model and its own successor. Kirkpatrick was compared with Phillips, which adds return on investment as a fifth level, or with Thalheimer's LTEM, which was designed to supersede Kirkpatrick and within whose eight tiers the original four largely sit. These are two points on one lineage rather than two approaches, and where the things being compared share a parent the comparison has nowhere to go. That is why the recommendation which follows is so often to run both, which adds duplication rather than the balance the question was asking for.
Very few writers reached outside that family. Brinkerhoff's Success Case Method asks a different question, which is not how far up the tiers an organisation has climbed but who succeeded, who did not, and what conditions account for the difference between them (Brinkerhoff, 2003). It produces a diagnosis that a sponsor can act on. Where a cohort of strategic practitioners reaches instinctively for a single lineage, it is my view that this tells us something about the profession's shelf rather than about the individuals standing at it.
What I found when I did this myself
I should be clear that this is not a criticism made from the outside. When I led learning and development for a police service of around 10,000 officers and staff, the evaluation work that told me most was not a level climbed on a ladder.
An evaluation framework built into student officer firearms training identified a disparity in performance by gender, with female students performing 25% less effectively than male students. That was not a comfortable finding and it was not one I had gone looking for. It was available only because the measurement had been designed into the programme rather than requested after delivery, and it led to changes in training style and equipment that would not otherwise have been made.
On a separate piece of work, the design of a critical incident management course, I used the CIRO model (Warr, Bird and Rackham, 1970) rather than Kirkpatrick, because CIRO begins with context and requires an assessment of the performance deficiency before any intervention is designed. In both cases the evaluation was agreed at the point of commissioning, and it is my view that neither would have survived being raised at the end.
What I would take from this
Four things, and none of them is to try harder.
1. Agree the evaluation before agreeing the intervention. The measure, the owner and the date belong in the scoping conversation, while the sponsor still wants something from you. Asked afterwards, you are requesting a favour, whereas asked at the outset you are setting a condition of doing the work.
2. Treat transfer as something to be engineered rather than something that happens afterwards. Baldwin and Ford (1988) established that transfer is shaped by trainee characteristics, training design and the work environment, and Weinbauer-Heidel and Ibeschitz-Manderbach (2018) organise those same conditions into twelve levers that can be built deliberately. Where the conditions are engineered, less time is spent afterwards pleading for evidence that they worked.
3. Separate the evaluation of the learning from the evaluation of the system. When behaviour does not change, the content is rarely the cause. Opportunity to practise, line manager reinforcement and organisational tolerance for a new way of working usually are, and an evaluation that cannot distinguish between them will not support any useful action.
4. Be honest about proportionality and say so openly. Not every intervention warrants Level 4, and pretending otherwise is what makes the whole subject feel like a standing reproach. It is my view that deep evaluation should be reserved for the expensive, the strategic and the high risk, and that we should say publicly which interventions fall into which category, so that routine work is not quietly judged against a standard nobody intended to apply to it.
The uncomfortable part
There is a conclusion in this that I am still sitting with. If we consistently evaluate only what we can measure without anyone else's help, then our evaluation data describes the reach of our own function rather than the impact of learning.
Some of the reluctance to pursue Levels 3 and 4 is genuinely about access and resource, and I would not dismiss that. I am not convinced that all of it is. Reaction data is safe and tells us that people enjoyed the day, whereas transfer data can tell us that a programme we designed, defended and delivered made no measurable difference to anything. I would be interested to know which of those two others think is really keeping us at Level 1.
References
Alliger, G.M. and Janak, E.A. (1989) 'Kirkpatrick's levels of training criteria: thirty years later', Personnel Psychology, 42(2), pp. 331-342.
Alliger, G.M., Tannenbaum, S.I., Bennett, W., Traver, H. and Shotland, A. (1997) 'A meta-analysis of the relations among training criteria', Personnel Psychology, 50(2), pp. 341-358.
Baldwin, T.T. and Ford, J.K. (1988) 'Transfer of training: a review and directions for future research', Personnel Psychology, 41(1), pp. 63-105.
Brinkerhoff, R.O. (2003) The Success Case Method: Find Out Quickly What's Working and What's Not. San Francisco: Berrett-Koehler.
Holton, E.F. III (1996) 'The flawed four-level evaluation model', Human Resource Development Quarterly, 7(1), pp. 5-21.
Overton, L. (2023) Learning at Work 2023: survey report. London: Chartered Institute of Personnel and Development.
Warr, P., Bird, M. and Rackham, N. (1970) Evaluation of Management Training. London: Gower Press.
Weinbauer-Heidel, I. and Ibeschitz-Manderbach, M. (2018) What Makes Training Really Work: 12 Levers of Transfer Effectiveness. Hamburg: tredition.