The gap between a demo that dazzles and a system that survives real traffic is where nearly every enterprise AI program actually lives — and where nearly every vendor discussion is silent.
The most cited number in enterprise AI in 2025 came from the MIT NANDA initiative's State of AI in Business 2025 report, and it was ninety-five percent. That was the share of generative-AI pilots surveyed across three hundred announced enterprise deployments, and a hundred-and-fifty additional interviews with executives, that had produced no measurable financial return. The report was picked up by the Fortune business desk, then Axios, then everywhere. Within a month it had become the framing statistic — the number every board deck opened with, the number every skeptic used to argue the technology wasn't ready, the number every vendor tried to reframe. And within a month of that, the framing had eaten the substance. Nobody was reading the paragraphs around the number. They were reading the number.
The paragraphs are more useful than the number. What the MIT researchers actually found — the report is a public PDF, worth reading directly — was not that the models don't work. It was that the pilots don't turn into programs. The technology, on the customer-facing side, performed. What failed was the transition from a working prototype to a working production system, and the failure was consistent enough across companies that it had a shape.
The shape was this. A vendor demo produces enthusiasm. An internal champion sponsors a pilot. The pilot is scoped narrowly and staffed with a small team of the company's better engineers, and it works, because narrowly scoped prototypes staffed with good engineers usually work. The pilot's success is presented at a leadership offsite. Budget is approved for scale-up. The scale-up begins. And then, six to nine months in, three things arrive together: the data required at production scale turns out to be dirtier than the pilot's data, the eval framework built for the pilot cannot detect the failure modes real users produce, and the workflow the tool was meant to displace turns out to be more subtle than the workflow the pilot modeled. The program stalls, quietly, in leadership updates. A year in, it is a case study in what everyone else is failing to do.
This is not a new pattern. It is the same pattern that ate a decade of enterprise machine-learning projects before generative AI arrived — VentureBeat's oft-cited 87 percent number from 2019 was measuring the same phenomenon, on a different model class. What's new is the scale of the money involved, and the visibility of the failure. When AI programs stalled in 2019, they stalled inside data-science departments nobody outside the CTO org read email from. When they stall in 2026, they stall in front of the board, because the board approved the spend.
The three gaps between pilot and program
The first gap is data. The pilot ran on a hand-selected sample. The sample was clean because a human curated it. In production, the sample is whatever the operational systems produce — data with the missing columns, the inconsistent schemas, the historical migrations, the tenant-specific overrides, the fields that mean one thing in the CRM and something different in the ERP. This is the reason a large-model AI system that scored 92 percent on the pilot dataset returns useful answers on 61 percent of production queries — a gap not caused by the model, and not fixable by a bigger model. It is caused by the data.
Ravin Kohli, who works on enterprise deployment at a large integrator, framed it plainly in his commentary on the McKinsey State of AI tracking: "The AI part isn't the bottleneck. The data part is the bottleneck. The AI part is the thing that reveals the data bottleneck." That is close to right. It is also the reason "we need a data platform" becomes the standard follow-on ask six months into a stalled program — an ask that would have been more useful two months into the strategy phase, before anyone signed a vendor contract, and that is almost always more expensive to satisfy after the fact than before.
The second gap is evals. The pilot's success was measured by whether the demo worked and whether the pilot users liked it. Neither is an eval. An eval is a specific written definition of what "good enough to ship" means for the specific system, a golden dataset that exercises the failure modes the system is likely to encounter, and a repeatable measurement of the system's performance against that definition — offline, before every model change; online, on a sample of live traffic, continuously. Anthropic's public work on feature-steering evaluations and OpenAI's open-source Evals framework are two of the more thoughtful public discussions of what this looks like in practice. Almost no enterprise AI pilot has anything like it. Most enterprise AI pilots have a spreadsheet of prompts a product manager copied down after a testing session, and a Slack channel where users complain when something goes wrong. That is not an eval framework. It is anecdotal QA masquerading as one.
The third gap is the workflow the tool was meant to displace. Every AI initiative is, in the end, an intervention in an existing operational process. The pilot approximated the process. The production reality includes the process's edge cases, the informal workarounds employees developed over years to handle the exceptions the formal system doesn't cover, the second-order dependencies on the data the process produces, and the identity of the manager whose team's numbers depend on the process running the old way. Any of these is enough to stall a rollout. Together they are load-bearing. The pilot did not test any of them, because the pilot was scoped narrowly enough to sidestep them. The production system has to face all of them at once.
What the vendors don't tell you
Every AI vendor pitching a large enterprise in 2026 has an answer to the pilot-to-production gap. The answer is always some variant of "our platform handles that." The platform does not handle that. The platform provides tooling that can be used to handle that, when composed with data engineering the customer has not yet done, evaluation infrastructure the customer has not yet built, and workflow-integration work the vendor's sales engineers are not allowed to describe honestly on a discovery call. The vendor sells the platform as a solution because the vendor is selling a platform. The customer buys the platform as a solution because the customer wants a solution.
This is not a moral failure on either side. It is the ordinary geometry of enterprise software sales, and it produces the ordinary result: platforms sit half-configured for eighteen months, then get renewed anyway because writing off the sunk cost is a harder political move than paying for another year. The Gartner prediction that thirty percent of generative-AI projects would be abandoned after proof-of-concept by end of 2025 was almost certainly conservative; the number now looks closer to forty percent for programs that started in the first big wave, though the abandonment is often re-labeled as "de-prioritization" for face-saving reasons.
The vendor economics reinforce this. Every model provider is competing for the same enterprise wallet; every hyperscaler is bundling AI credits into cloud renewals; every integrator is booking hours against transformation budgets. Nobody in the vendor ecosystem has an incentive to tell a customer you don't need this yet — go fix your data first. The one party with that incentive is the customer's own executive team, and the executive team is the one whose bonus depends on the transformation narrative being intact for the year.
The failure mode that isn't in the deck
There is a fourth failure mode that gets less attention, because it is embarrassing to name. Many AI pilots succeed on their explicit metrics — completion rate, user satisfaction, cost per query — and still fail to produce program-level value, because the workflow they optimized was not on the critical path of the business. The pilot displaced work that wasn't the bottleneck. The customer was measuring what the tool was good at, not what the business needed done.
This is the reason "AI adoption" surveys read so oddly. A BCG survey published in mid-2025 found that 74 percent of companies had deployed AI in at least one function, that 76 percent of those deployments met their pilot-phase KPIs, and that only 26 percent had translated the deployment into measurable enterprise value. The gap is not, in most cases, technical. The gap is that the enterprise value case was never rigorously constructed, and the pilot's KPIs did not require it to be.
Building the value case is unglamorous work. It looks like a two-page memo that names the specific dollar or hour flow the initiative is intended to change, the size of that flow, the current baseline, the target baseline, the mechanism by which the AI intervention changes the baseline, and the assumption set that has to hold for the mechanism to work. Most AI programs never produce this memo, because producing it forces uncomfortable questions — starting with "is this actually the highest-leverage AI investment we could be making this year?" and ending with "if we can't defend this on a first principles basis, are we sure we should be doing it?"
What actually gets a pilot into production
The programs that make the transition — and there are some — tend to have four things in place before the pilot is greenlit, not after it stalls.
A named business problem, quantified. Not "we want to explore what AI can do for us." A specific workflow, with a current cost, and a target cost, and a plausible mechanism by which the AI intervention will move the current toward the target. If nobody on the executive team is willing to write this down before the pilot begins, the pilot will produce a demo, not a program.
A data readiness assessment, honest. Not "our data is fine." A written read of the specific data the initiative will touch, its actual quality (measured, not asserted), its lineage, its ownership, and the specific fixes required before AI can act on it. This is dull and expensive to produce, and it is the single highest-leverage thing an executive team can commission before signing a vendor contract.
An eval framework, from day one. A golden dataset of at least a few hundred examples representative of production traffic, an offline metric that ties to the business outcome, an online metric that runs continuously on a sample of production requests, and a specific owner on the team whose job is to maintain and report on both. A pilot without evals cannot become a program without evals — and building evals after the pilot is far more expensive than building them as part of it.
A named person owning the process the tool changes. Not "IT owns it." The manager whose team's numbers depend on the old workflow needs to be in the room when the new workflow is designed, and they need to be measured on the success of the new one before their old measurement structure disappears from under them. Adoption is a design problem, not a change-management problem, when the design accounts for the person whose incentives shift.
None of this is technically hard. All of it is politically hard, because it forces the executive team to be specific about what they are trying to change and how they will know if it worked — before the pilot's success gives them permission to postpone those questions.
The programs that translate into enterprise value do the specific, boring, discipline-imposing work of being clear about the value case before they start. The programs that produce ninety-five-percent failure rates do the reverse. They start with the technology, hope the value will follow, and discover a year in that the executive team never agreed on what "value" meant to begin with. The pilot was working. The program was never designed.
Adam Valine leads Veil Consulting's AI Strategy & Roadmap and Agent & LLM Systems Design practices. He is also AI Strategy Lead at Solutions Plus Consulting.
References
- MIT NANDA Initiative — The GenAI Divide: State of AI in Business 2025 (PDF) — The primary source for the "95 percent of pilots produce no measurable return" statistic; three hundred announced enterprise deployments plus one-hundred-fifty executive interviews.
- Fortune: MIT Report — 95% of Generative AI Pilots at Companies Are Failing (August 2025) — First widely-read business-press coverage that turned the NANDA number into a board-deck fixture.
- Axios: AI GenAI Pilots Failure — MIT Report (September 2025) — Follow-on coverage; useful context on how the framing propagated.
- BCG: Closing the AI Impact Gap (2025) — 74 percent of companies deployed AI in one function, 76 percent met pilot KPIs, only 26 percent translated the deployment into enterprise value. The value-case gap, quantified.
- VentureBeat: Why do 87% of data science projects never make it into production? (2019) — The pre-generative version of the same failure mode; useful as historical grounding.
- McKinsey QuantumBlack: State of AI — Longitudinal enterprise AI survey series; useful for tracking the shift in enterprise perception year over year.
- Anthropic Research: Evaluating Feature Steering — Public work on how a serious eval practice looks inside a frontier lab; useful as a template.
- OpenAI evals framework (open source, GitHub) — The open-source template most enterprise eval practices are downstream of.