- Published
- Reading time
- 6 min
- Author
- Gireesh Malhotra
If you cannot name the measure before you build, you will be reconstructing the benefit afterwards. I have had to do that. It does not survive a second question.
I have sat on the wrong side of this conversation, so I will start there.
A quarterly review, a customer’s CFO on the call, an initiative we had all worked hard on and genuinely believed in. He asked what it had delivered. And what I had was a set of anecdotes — three teams who said the new process felt faster, a champion who was enthusiastic, a screenshot of a dashboard that had been built after the fact, and a benefit figure that had been estimated by the people who wanted the project to continue.
He did not argue with any of it. He asked one follow-up question — where the baseline came from — and the whole thing came apart in about forty seconds, because the baseline had been reconstructed from memory nine months after the process changed.
That was a bad afternoon, and it was entirely avoidable. Nobody had been dishonest. We had simply left the measurement until the point where it was needed, by which time it was too late to measure anything.
The window closes at go-live
Here is the thing about benefit measurement that took me too long to internalise: the baseline is only available before you change the process.
That sounds obvious written down. In practice it is the single most-skipped step in enterprise AI delivery, because at the point where you could capture the baseline you are busy with design, integration, and testing, and nobody feels the absence of a number they will not need for a year.
Then the process changes. The old path stops being exercised. The tickets that used to flow through it are gone. And the honest answer to “how long did this take before?” becomes “roughly, according to the team, about a day or so” — which is not a baseline, it is a recollection, and a CFO can tell the difference immediately.
So the discipline starts with one rule: the measure and its baseline are part of the build, in the same sprint as the use case. Not a follow-up workstream. Not a phase two. If it is not instrumented, it is not done, in the same way that an untested integration is not done.
What actually counts as a measure
Not everything that gets called a benefit will survive contact with finance. Over the last few years I have developed a fairly blunt hierarchy.
Strongest: a figure that already exists in a system of record and moves. Days sales outstanding. Invoice exception rate. Touchless invoice percentage. Order-to-cash cycle time. On-time delivery. These are numbers the business already reports, already trusts, and already argues about. If your use case moves one of them, you do not need to persuade anyone that the measure is legitimate — that argument was won years ago.
Reasonable: a rate or volume you can count cleanly on both sides of the change. Exceptions auto-resolved versus escalated. Requests handled without a human touch. Cases closed within SLA. These require you to have counted before, which is exactly why the instrumentation has to precede the change.
Weak: time saved per person, multiplied out. This is the most common benefit in AI business cases and the least durable. It requires a self-reported baseline, it assumes the recovered time goes somewhere valuable, and it invites the question nobody wants — “so are we reducing headcount?” If the answer is no, then finance quite reasonably books the saving at zero. I now avoid leading with this. It is a supporting number, not a headline.
Worthless in a review: sentiment. Adoption enthusiasm matters enormously for whether a programme survives internally. It is not a benefit. Do not put it in the value column; put it in the risk column, where it belongs, as evidence that the change is sticking.
The QBR as an engineering artifact
The reason I care about this so much is that in my world the quarterly business review is not a status meeting. It is where a customer decides whether to renew, expand, or quietly let something lapse. And a review built on instrumented measures is a completely different conversation from one built on narrative.
The structural difference is who is doing the work. In a narrative review, I am persuading. Every claim I make invites a counter-claim, the customer’s team is politely sceptical, and the meeting is a negotiation about interpretation. In an instrumented review, I am reading. The number came from their system, on a definition their team agreed, against a baseline captured before we started. Nobody has to take my word for anything, which means the discussion moves immediately past “did this work” to “where do we do it next” — which is the conversation I actually want, and the one expansion comes out of.
The Balanced Scorecard works the same way when it is used properly rather than decoratively. Its value is not the four quadrants. It is that it forces you to declare, in advance and in front of the customer, which measures you are willing to be judged on. That declaration is the commitment. Everything after it is arithmetic.
I have started treating the scorecard as a design input rather than a reporting output. If a proposed use case cannot be attached to something on the scorecard, that is worth knowing during design, when the use case can still be reshaped, rather than at the review, when it cannot.
The part of this I get wrong
The obvious risk in everything I have written is that measurement discipline quietly becomes a filter that only passes the measurable — and the measurable is not the same as the valuable.
Some of the most consequential work I have been involved in did not have a clean number attached. Retiring an integration that three people understood and two had left did not move a KPI; it removed a failure mode that would eventually have taken out a month-end close. Getting a definition of “available inventory” agreed between two departments produced no benefit line at all, and unblocked four subsequent use cases. Neither of those would have passed the hierarchy I laid out above, and both were clearly worth doing.
There is a subtler version of the same trap. Once a measure is declared and reported on, people optimise for it. Declare touchless invoice percentage and you will get a touchless invoice percentage — occasionally by narrowing what counts as an invoice. That is not fraud, it is ordinary human response to being measured, and it is the reason a value review needs at least one qualitative check alongside the numbers. I ask the process owner a version of the same question every quarter: is this actually better to work with, or does it just report better?
So the position I hold is not “only do measurable things.” It is: be explicit about which category each piece of work is in. Instrument what can be instrumented, rigorously and in advance. For the foundational work that cannot be, argue it on risk removed and options created, name it as such, and do not dress it up as a benefit figure — because a soft number in a hard column is what costs you credibility on all the other numbers.
The one-line test
Before a use case gets built, I want an answer to this sentence:
Twelve months from now, we will know this worked because __________ moved from __________ to __________, read from __________.
Four blanks. If they can all be filled in, the value conversation a year from now is already won. If they cannot, that is not a reason to stop — but it is a reason to be honest, in the room, about what kind of case you are going to be able to make.
- Value realization
- QBR
- Balanced Scorecard
- Customer success
Written by Gireesh Malhotra, Customer Success Senior Manager at SAP. Views are my own and do not represent my employer.
