Delivery
← All insights

Why 95% of AI Pilots Fail to Scale, and What to Do Instead

MIT found that 95% of enterprise AI pilots never reach production. The technology usually works. The organisation does not. Here is why, and how to fix it before you start.

Delivery

CTO, COO

Part of our guide: AI implementation and operating model

In July 2025, MIT's NANDA Initiative published research that should have stopped every AI programme in its tracks, drawing on over 300 publicly disclosed AI initiatives, interviews with 52 organisations and 153 survey responses. The research found that 95 percent of enterprise AI pilots failed to progress to scaled adoption, a rate of ninety-five percent rather than 50 percent or 70 percent.

A note on that number, added July 2026. We have since gone back to the primary source, and the figure deserves rather more caution than it usually gets. The report was labelled preliminary findings, was never peer reviewed and says of itself that its figures are "directionally accurate based on individual interviews rather than official company reporting". The 5 percent measures one narrow category, task-specific tools reaching production, and the same chart shows general-purpose tools reaching implementation around 83 percent of the time. Wharton's own 2025 research found three in four leaders reporting positive returns, and we set out what the number does and does not support, along with five other statistics boards routinely quote, in the AI statistics that survive checking. The argument below does not rest on the 95 percent figure, and a better-sampled version of the same conclusion comes from McKinsey's November 2025 survey. In that survey, 39 percent of firms attribute any enterprise-level EBIT impact to AI and just 7 percent report full scaling.

The reaction in most boardrooms was to question the technology, and the technology is rarely the problem: the pilots work, the models perform and the demos are impressive. What fails is everything around the technology, the operating model, the governance, the integration, the change management and the business case, and organisations are not built to sustain what their pilots produce.

The five failure patterns

Across the research, from MIT, Harvard Business Review and our own experience, the same patterns repeat.

No executive alignment. Seventy-three percent of failed projects lack clear executive alignment on what success looks like, and different stakeholders hold different expectations. The CTO thinks success is a working model, the CEO thinks success is revenue impact and the CFO thinks success is cost reduction. Without alignment before the pilot starts on the definition that applies, the pilot delivers something that nobody is satisfied with.

Pilot-to-production gap. The team that builds a pilot is rarely the team that runs a production capability, and the pilot is built by enthusiastic technologists with temporary resource. The production capability needs to be run by operational teams with permanent accountability, and if you do not design this handover before the pilot starts, there is nobody to hand over to.

No operating model. The pilot proves the technology works, and the questions that follow are who owns it in production, who monitors its outputs, who is accountable when it gets something wrong and what happens when the model needs retraining. These are operating model questions and they are almost never addressed during the pilot, and AI Operating Model Design sets out why this layer is critical.

Underinvestment in data. Forty-three percent of organisations cite data quality and readiness as their top obstacle to AI success, and the pilot uses a curated dataset while production requires real data. That data is messy, incomplete and inconsistent, and if the data foundation is not addressed, the capability that worked beautifully in the pilot fails in production.

Loss of sponsorship. Fifty-six percent of failed projects lose active C-suite sponsorship within six months, and the pilot takes longer than expected while other priorities emerge and the champion moves on. Without sustained senior commitment behind it, the pilot quietly dies.

What to do instead

The organisations that do reach production share common characteristics.

Design for production from day one. Do not build a pilot that needs to be rebuilt for production, and design the architecture, the governance and the operating model from the start. Designing that way takes slightly longer upfront and eliminates the pilot-to-production gap that kills most initiatives.

Align before you build. Before writing a line of code, get your leadership team aligned on what success looks like, in business terms, not technology terms. The alignment covers what metric improves, by how much, over what timeframe and who is accountable, and that conversation takes half a day while skipping it costs months.

Compress the delivery cycle. The longer the gap between investment and value, the higher the failure rate, and projects with sustained CEO involvement achieve 68 percent success rates versus 11 percent for those that lose sponsorship. The way to maintain sponsorship is to deliver visible value fast, and an 8-day sprint that produces a working capability keeps momentum alive where a 6-month pilot never will. See The 8-Day AI Sprint for how we structure this.

Build governance as a first-class deliverable. Governance is not what you add after the technology works, and it is what makes the technology work in production. Design ownership, accountability, monitoring and review cadence alongside the capability rather than after it, and AI Governance in Financial Services sets out what good governance looks like.

Invest in the business case, not just the technology. Projects with a clear, quantified business case, tied to specific metrics the board cares about, survive the inevitable moments when other priorities compete for attention. See How to Write an AI Strategy Your Board Will Back.

The strategic alternative

That failure rate, whatever the exact figure for your category, is not inevitable, and it is the natural consequence of an approach that puts technology before strategy, pilots before architecture and enthusiasm before governance. The alternative is to start with your business capabilities, design the operating model, build governance from day one and compress delivery into cycles short enough to maintain momentum.

Breathe gives you the strategic foundation, and Flow delivers the first capability with production-grade architecture and governance, and that combination is how you land in the group that reaches production rather than the pile that stalls.

Work with us

Flow: Build something real

We take the highest-priority opportunity and build a working AI capability, with governance, measurement and an operating model designed from day one: a real, deployable tool (not a proof of concept).