The most important number in enterprise AI still doesn’t exist — and that should worry you
Two weeks ago I told you the missing human receipt was becoming the real story. This week, after another round of digging, I’m even more convinced that the absence is not a research inconvenience. It’s a signal.
We now have named enterprises deploying agentic systems at meaningful scale. Citigroup’s Arc platform is generating more than 100,000 agentic development hours per week. DBS is publicly scaling agentic AI for corporate bankers and explicitly targeting major time reductions in credit memo work. Microsoft 365 Copilot is at 30 million paid seats, while Microsoft’s headcount is down and the company very carefully refuses to draw a straight causal line from AI to jobs. Then you zoom out and the pooled benchmarks start to rhyme: roughly 6.4 to 7.2 hours saved per knowledge worker per week across the usual benchmark set, while Digital Applied’s synthesis of Forrester data says only 41% of agent rollouts hit positive ROI inside year one.
What we still do not have is the one table that would settle the extension-versus-replacement argument in the real world: pre-and-post workload hours, headcount trajectory, and cognitive load, all together, from the same named company, for the same deployment. Not anecdotes. Not aggregate telemetry. Not a target slide. A receipt.
Here’s what I think is happening. The organizations deploying agents at scale are measuring the right things for procurement decks, board updates, and vendor negotiations — output, cycle time, ticket volume, cost per task, agent hours generated. They are not measuring, or at least not publishing, what happened to the humans standing next to the system. Did those people move up-stack into more valuable work? Did expectations just rise so the same team now produces more under more pressure? Did headcount flatten because the work got absorbed? Did cognitive load actually drop, or did it just mutate from doing the task to supervising a machine that does it unevenly?
That omission matters because my thesis has been clear from the beginning: agents are extensions of people, not replacements. I still believe that is the most likely macro pattern. But I can’t prove it with the data we have, and neither can the people confidently arguing the opposite. If Citigroup can tell you the machine-hours, but not the human redeployment story, that is a choice. If DBS can tell you the percentage of a banker’s day tied up in credit memos, but not yet the realized human delta after deployment, that is a choice too. Either the data is messy enough that nobody wants to publish it, or nobody built the instrumentation for the human side in the first place. I’m honestly not sure which possibility is worse.
And if you run a company, don’t miss the practical lesson here: if your vendor or internal team can show you agent output but not human outcome, you are not looking at transformation data. You are looking at a partial ledger.


