Skip to content

Adoption

The Adoption Gap: Why Enterprise AI Stalls After the Demo

The pilot works. Everyone in the room is impressed. Six months later nothing has changed. Four things have to be true before an AI use case survives contact with a business process.

An empty corporate conference room after a meeting has ended, with a faint three-tier architecture sketch left behind on the glass writing wall.
Published
Reading time
6 min
Author
Gireesh Malhotra

A demo is a claim about capability. Production is a claim about accountability. Those are different arguments and they need different evidence.

I have sat in a version of the same meeting perhaps forty times in the last three years.

The setup does not vary much. A conference room, or more often a call with fourteen people on it. Someone from the customer’s innovation team runs a demo. An agent reads an invoice, matches it against a purchase order, catches a discrepancy in the freight line, drafts the exception note. It takes about eleven seconds. Three years ago that sequence took a person twenty minutes, and a second person to check the first one.

The room is impressed. I am impressed, and I have been doing this since 2000, so I would like to think I am hard to impress.

Then somebody — usually from finance, usually the most senior person on the call — asks the question that decides everything.

“What happens when it gets one wrong?”

And the temperature in the meeting changes.

I used to read that question as resistance. The innovation team certainly reads it that way; you can watch them deflate. I have come to think it is the opposite. That question is not an obstacle in front of the project. It is the project. Everything between a working demo and a production business process is contained in the honest answer to it, and the reason so much enterprise AI stalls at exactly that moment is that almost nobody walks into the room with the answer prepared.

The gap is not technical

Here is what I have found consistently: when an enterprise AI initiative dies, it very rarely dies because the model could not do the task.

It dies in the eleven months after the demo, quietly, from a specific and boring set of causes. The data the agent needed was scattered across three systems and one of them was a BW instance nobody had owned since 2019. No one could say what the agent was permitted to read, so security said no to all of it. The process owner in accounts payable was never in the room, and when she finally saw it she pointed out that the exception the agent flags is not actually the exception her team cares about. The business case was built on a headcount saving that HR was never going to sign, so when renewal came around there was no number to point at.

None of that is an AI problem. All of it is the work.

This is the part of the current moment I find genuinely interesting, and slightly funny. We have spent two years talking about model capability as though it were the scarce resource. In my portfolio it is the least scarce thing I have. Capability arrives on a release cycle now, whether the customer is ready or not. What is scarce is a landscape clean enough to ground an agent on, a governance answer a risk officer will accept, and a business case that survives a CFO in a bad mood.

Four things that have to be true

I now run this as a checklist before I let a customer put an agentic use case on a roadmap. Not because checklists are elegant, but because every item on it is something I have watched a program die from.

One: the agent has something trustworthy to ground on, and somebody owns it.

An agent’s output is bounded by what it can retrieve. If the data it needs is spread across a PI landscape and a BW estate that has drifted for six years, the agent will produce fluent, confident answers built on a version of reality nobody in the business recognizes — which is worse than no answer, because a wrong number that reads well gets forwarded.

This is why so much of my week is spent on work that does not look like AI work at all: moving integration onto BTP Integration Suite, moving BW onto Datasphere, getting a grounding layer in place through Business Data Cloud. Customers sometimes hear that as a delay tactic. It is the opposite. It is the AI project arriving early, wearing unglamorous clothes.

Two: the blast radius is defined before the scope is granted.

For any agent that can act rather than only suggest, I want three things written down: what it can touch, what the worst realistic outcome is if it is confidently wrong, and how we would know. If the worst case is a badly worded draft sitting in someone’s queue, grant it wide scope and move fast. If the worst case is a posted journal entry or a payment released to a vendor, the conversation is entirely different and it should be.

Autonomy is earned incrementally against evidence. It is not a starting condition, and the customers who try to start there are the ones who end up with a moratorium from their own risk committee eighteen months later.

Three: a specific human is in the loop, and it is not a committee.

“Human-in-the-loop” has become one of those phrases that sounds like governance without doing any. The version that works is much more specific: which named role reviews which class of decision, at what threshold, inside which system, and with how long to respond before the item escalates or expires.

If the answer to any part of that is vague, what you actually have is an agent operating unsupervised with a plausible-sounding paragraph in a slide deck about oversight.

Four: the value is instrumented before go-live, not reconstructed afterwards.

This is the one I have seen skipped most often, and it is the one that kills renewals. If you cannot state, before you build, which measure will move and where that measure is read from, then twelve months later you will be sitting in a value review assembling a narrative out of anecdotes. I have had to do that. It is a bad afternoon, and it does not survive a second question.

Decide the measure first. Instrument it in the same sprint as the use case. Then a quarterly review becomes evidence rather than persuasion, which is a much better position to argue an expansion from.

What I get wrong about this

I should be honest about the cost of what I have just described, because I have overcorrected with it more than once.

Governance discipline is not free, and it is not automatically virtuous. There is a real failure mode on my side of this table: you can hold a use case at the guardrail stage so long that the momentum leaks out of it, the sponsor moves to another role, and the capability that was genuinely exciting in March is a line item nobody defends in November. I have done that. I have watched a good use case die of caution, and caution left no fingerprints, so nobody wrote it down as a failure.

There is also a version of the innovation team’s argument that is simply correct, and the checklist above can be used to avoid hearing it. Sometimes the reason the agent’s output does not match what accounts payable cares about is not that the agent is wrong — it is that the process was assembled in 2011 around a constraint that no longer exists, and the honest move is to redesign it rather than automate a bad version of it faster. That is a harder conversation, it is not a customer success conversation on paper, and I have ducked it when the quarter was tight.

So the discipline has to cut both ways. The four items are a way of making risk explicit, not a way of making it someone else’s problem. If the answer to all four is genuinely unknown, the right response is a two-week exercise to find out, not a two-quarter delay dressed up as governance.

Where this leaves me

The gap between a working demo and a production process is not a gap in the technology. It is a gap in ownership, and it closes when one person is willing to hold both ends at once — the architecture and the business case — and to still be in the room at the quarterly review when it is time to prove the number.

That is the job I have. Eighteen years in, and across five platform shifts that each followed this same arc, it is the most interesting version of the job I have had. The capability is not the hard part anymore. Getting an organization to actually use it, safely, in a way it can defend to itself — that has never been harder or more worth doing.

The next time someone asks me what happens when the agent gets one wrong, I want to be able to answer in one sentence. That is the whole discipline, really. If I cannot, we are not ready, and the demo was never the point.

  • Enterprise AI
  • Adoption
  • SAP Joule
  • Change management

Written by Gireesh Malhotra, Customer Success Senior Manager at SAP. Views are my own and do not represent my employer.