Why technical performance is not sufficient evidence that an AI deployment created real-world value.
Evaluating the model measures the model — not the system it entered.
The EVOLVRS Research series · Reads with The Measurement Gap · Research Note 01 · Essay 02
Two different questions
The two questions sound like one. They are not, and they are answered by two different measurements.
The first is internal: accuracy, latency, task completion, uptime, adoption — the model held against its own specification. The second is external: the real-world system before the model, and the same system after, held against its own baseline. A model can pass every internal test while the system around it is unchanged, or worse. Faster outputs that no one acts on. Higher adoption of a tool that shifts judgment onto a bottleneck elsewhere. Automation of a step that was never the constraint. Each is a technical success and a system non-event.
Model telemetry and system change are different readings. Only one of them answers the question a leader is actually asking.
The evidence gap
The numbers describe an economy that measures whether the model works and infers that the system improved.
KPMG Global AI Pulse (2,110 leaders) · BCG survey of 152 chief executives
In KPMG's Global AI Pulse, 95% of organizations report having an AI strategy while only 8% report established ROI.1 In BCG's survey of chief executives, more than half name the missing link between AI and the P&L as a key barrier, yet only 14% have clearly defined the P&L impact for all their AI initiatives.2 And 30% of chief data and analytics officers name measuring the impact of data, analytics, and AI on business outcomes as their single biggest challenge.3
This is not, at root, a returns problem. It is a measurement problem wearing a returns problem's clothes. Organizations cannot report the value because they never measured the system the technology was supposed to change.
The reading you cannot take later
The reason the second question is so often unanswerable is timing.
To know what a deployment changed, you need a reading of the system from before it — and the one moment you can take that reading is before you deploy. AI cannot establish yesterday's baseline today. Once the model is live, the pre-deployment system is gone, and "did it work?" collapses back onto the only measurement still available: the model's own telemetry.
This is why technical evaluation and real-world value keep getting conflated — not because anyone believes they are the same, but because the model's numbers are the only ones on the table. The discipline the moment requires is unglamorous. Decide what in the surrounding system the technology is meant to change. Measure it before. Measure it again after, on the same instrument.
Absent that, "the model worked" is a true statement about the model — and no statement at all about the system.