The KPI That Measured the Wrong Thing
When the dashboard is green and the service is broken
Alex submitted the third query through the portal on a Thursday afternoon, copying in the reference numbers from the first two. The housing benefit calculation still didn’t match what Alex had been told over the phone three months earlier. The difference was not large — roughly forty pounds a month — but over four months it had accumulated into an overpayment notice that Alex was now expected to repay.
Each of the three queries had been acknowledged within twenty-four hours. Each had been assigned a case reference. Each had received a response. The most recent response explained, in clear language, how housing benefit calculations worked. It did not explain why Alex’s calculation differed from the figure given on the phone. The case had been marked as resolved.
The week after Alex received the overpayment notice, the council’s customer services director presented the quarterly performance report. Digital channel adoption was up. Average response time was down to 1.4 working days. Cost per interaction had fallen 28%. Case resolution rates were running at 94%. The dashboard was green across every metric. Nobody in that meeting had Alex’s name in front of them.
The metrics the council was reporting were accurate. An average response time of 1.4 days is faster — measurably, demonstrably faster — than what came before. A 28% reduction in cost per interaction is a real achievement. A 94% case resolution rate sounds like a service that is working well.
What the dashboard doesn’t count
What none of those numbers captured: whether the question had been answered. The 94% resolution rate counted cases that had received a response and been closed. It did not count whether the response had resolved the issue. A case could be opened, acknowledged, responded to with a standard letter, and closed as resolved — three times — without the person ever getting what they needed. That is not a failure in the data. It is a failure in what the data was designed to count.
The calls that ended without resolution. The cases where the automated response was technically accurate and substantively irrelevant to the specific question asked. The people who received four replies and gave up on the fifth attempt. None of that appears in the performance report. Not because it was hidden, but because the system was never designed to see it.
The people running the service have no signal that anything is wrong. The people funding it have no reason to ask. The people using it — the ones for whom it is not working — have no channel to make the failure visible, except to keep submitting queries that will each receive a response and remain unresolved.
The politics of choosing what to count
This is where Goodhart’s Law applies — the principle that the economist Marilyn Strathern summarised as: when a measure becomes a target, it ceases to be a good measure. Response time was a reasonable proxy for service quality until it became a target. Once it was a target, the system optimised for it. Cases were responded to quickly. Whether the responses were useful was a different question — and not one the target required anyone to ask.
KPI structures for technology services are not purely technical documents. They are political ones. The metrics chosen reflect what the organisation is prepared to be accountable for, and organisations, rationally, choose metrics they can improve. Response time can be reduced through better routing and automation. Case closure rates can be improved by defining “resolution” in ways that are achievable. Cost per interaction falls as volume grows. All of these things can be improved without making the service more useful to the person at the end of it.
The human outcome — did the person get what they needed? — is harder to measure because it requires asking them. That takes time, raises uncomfortable questions, and produces data that is more difficult to present as success. So it is left out. Not maliciously. Structurally. The governance body that approves the metrics does not require it. The reporting cycle that reviews them does not miss it. And the dashboard that summarises them has no row for it.
There is also a governance dimension to this. The people who approve the KPI structure are presented with a set of metrics that look complete because they cover all the dimensions the previous governance model required: speed, volume, cost, compliance. Nobody in that room has ever been asked to approve a metric for whether the service actually resolved the person’s problem. Nobody has been told that such a metric is missing. The governance process validates what it is shown. What it is not shown does not fail — it simply does not exist in the conversation.
This is how a service comes to be celebrated and broken at the same time. The celebration is for what is measured. The brokenness is in what is not. The two things co-exist perfectly until someone submits a third query and receives a fourth standard letter.
The cost that doesn’t appear in the report
The pattern repeats across services. A case that the digital channel cannot resolve eventually gets escalated to a phone call — the kind of contact the digital system was built to reduce. The query is answered. The discrepancy is corrected. The outcome the portal was supposed to deliver arrives via the channel the portal was supposed to replace, weeks later than it should have.
The quarterly performance report will show digital channel adoption at its highest level. What it will not show: the volume of digital cases subsequently resolved by phone; the number of people who returned to a human channel because the digital one had not worked; or the real cost of a failed digital interaction, which includes not just the cost of the digital transaction but the cost of the phone call that followed it.
The service looks more efficient than it is. The measurement system has been designed in a way that makes that possible — by measuring the transactions that occurred rather than the outcomes they were supposed to produce. When a governance body reviews a performance report and sees green across every metric, the natural response is to keep doing what is working. The investment flows toward the metrics that are already performing well. The things that are failing silently receive nothing, because they are not visible as failures.
That is not a data quality problem. It is a design choice. The absence of an outcome metric is itself a decision, with consequences that accumulate quietly in every case the service was not built to see. And because the consequences are not measured, they do not generate pressure to change the measurement. The organisation that cannot see its own failures cannot feel the cost of them, and cannot direct resources toward addressing them. The green dashboard is not just a reporting artefact. It is a protection against accountability — not deliberate, but effective.
My Opinion
I have watched organisations celebrate dashboards that tell them almost nothing about whether their services are working. The instinct to measure is right. The flaw is in what gets chosen. SLAs, KPIs, OKRs all share the same structural problem: they measure what the organisation can control. An automated acknowledgement meets the response SLA. The service still failed. I’ve witnessed a large outsourcing organisation hit their KPIs and SLAs every single month for response times…by putting an automated email response thanking the user for their query.
The organisations closing this gap are measuring outcomes, not transactions — asking users directly whether their problem was resolved. Some are going further, tracking social impact metrics: what actually changed for the person, beyond whether the case was closed. A number is a starting point, not a conclusion. Before accepting one, ask what human reality it represents. If you cannot answer that, the number is not yet telling you what you think it is.
The questions the system was never required to answer
The question a measurement structure should start with is whether the technology is actually helping the specific person using it — not whether it is processing their request efficiently. Those are different things. A system that processes twelve thousand requests per quarter and resolves sixty percent of them is not necessarily more helpful than one that processes eight thousand and resolves ninety. But a performance report that counts volume and response time will describe the first system as the better one.
The second question is whether the technology adds something real to the person’s capability, or only reduces costs for the organisation. A reduction in cost per interaction is not the same as an improvement in the quality of the interaction. Faster and cheaper are not synonyms for better. A measurement system that tracks only cost and speed cannot answer whether the technology is doing what it was built to do. It can only confirm that the technology is doing what it was built to do cheaply and quickly — which is a different question entirely, and a less important one.
The third question — and the one most completely absent from most KPI structures — is whether the system responds to the specific circumstances of the person in front of it. The 94% resolution rate is a population-level figure. It tells you nothing about the person in front of it. It tells you nothing about the six percent for whom the service did not work, or what those cases had in common, or whether they were disproportionately the people who needed the service most. A system that performs well on average while failing consistently at the edges is not performing well. It is performing adequately for people whose situations fit the expected pattern, and invisibly for everyone else. Those are not the same thing, and no dashboard that measures only the average will ever surface the difference.
The question behind the numbers
Think of a technology or service you use regularly — at work, or as a citizen or customer. What do you think is being measured about your experience of it? Does that match what you would want to be measured?
If you are responsible for reporting on technology performance in your organisation, when did you last ask whether the metrics you report capture the experience of the person at the end of the system — not just the performance of the system itself?
If you could add one metric to your organisation’s technology reporting — something currently unmeasured that would tell a more complete story — what would it be, and why is it not already there?
A system that hits every KPI while failing the people it serves: is that a success or a failure? And if the answer is complicated, what does that tell us about how we have defined success?
A service can hit every target on its dashboard and still fail the person who needed it. The performance report will describe it as operating at record efficiency. Both things are true. They are not the same truth, and the distance between them is where the accountability isn’t.
Authors Note
Alex is a fictional character. Their story is drawn from a combination of professional observation and personal proximity to real events. The experiences described are real. The person is not.
You’re reading The Next Evolution by Neil Catton, articles that explore the human world and the intersection of technology, they try and ask difficult questions - not to scare - but to inform. If someone forwarded this to you, you can subscribe free at neilcatton.substack.com.
Neil Catton is the author of The Next Evolution, The Cognitive Crucible and The Shadow System - available on Amazon, and writes at the intersection of technology, ethics, and human purpose.


