insights

Why Most AI POCs Never Make It to Production

Walk into almost any mid-sized company right now and ask "what's happening with AI?" and you'll get the same answer, dressed in different words. There was a pilot. It went well. Everyone was excited. And then... nothing. The Slack channel went quiet. The demo deck is still in someone's downloads folder. Nobody quite remembers whose job it was to take it further.

This isn't a failure of the technology. The models work. The demos are genuinely impressive — we've sat in plenty of them and they're not smoke and mirrors. The failure is somewhere else entirely, and it's a more interesting problem than "the AI wasn't good enough."

The pilot proves the idea. Production proves you can run it. The pilot ✓ Clean, hand-picked data ✓ One user, one happy path ✓ A human laughs off mistakes ✓ No real load, no real cost ✓ Built fast, slightly off-process ✓ Success = applause in a demo TAKES WEEKS THE GAP Production ✓ Messy, real-world data — every edge case ✓ Every user, every malformed input, at once ✓ A designed failure path — who sees it, then what ✓ Real load, real cost-per-call at real volume ✓ A named owner, security review, change control ✓ Success = a number, agreed before launch TAKES MONTHS — IF ANYONE OWNS IT Most AI investment doesn't fail in the pilot. It evaporates in the gap.
The pilot and the production system share a model. They don't share much else.

The gap nobody puts in the slide deck

A proof of concept has to do one thing: prove the concept. That's it. It needs to work once, for one user, on one dataset, in a controlled demo, ideally in front of people who already want to be impressed.

Production has to do a different thing entirely: keep working, for everyone, on data nobody cleaned up first, when the network is slow, when the input is malformed, when three people hit it at once, six months after the person who built it has moved to a different project.

Those are not the same problem with different scale. They're different problems that happen to share a model underneath. And the gap between them is where the vast majority of AI investment quietly evaporates.

We'd estimate — and this lines up with most of what we hear across the industry — that somewhere around three in four AI pilots never reach production in any meaningful form. Not because the pilot was bad. Because nobody budgeted for, or even named, the work that sits between "it worked in the demo" and "it works."

What the demo conveniently skips

Here's the part that's almost funny if you've sat through enough of these. A POC nearly always uses a clean, hand-picked dataset. Production data is never clean. It has duplicate customer records, three different date formats depending on which system it came from, fields that are blank for half of 2019 because someone changed the schema and didn't tell anyone, and at least one column where "N/A" and "n/a" and an empty string are all being used to mean three subtly different things.

A model trained or prompted against the curated version falls over against the real version, not because it's a worse model, but because the inputs it's actually going to see were never part of the plan.

Then there's the question of what happens when the model gets something wrong — and it will, regularly, because that's the nature of probabilistic systems. A POC handles this by having a human in the room who laughs it off. Production needs an actual answer to "what happens now." Does the system flag low-confidence outputs for review? Does it fail loudly or silently? Who's accountable when it confidently gives the wrong answer to a customer at 11pm on a Tuesday? Nobody designs this during the exciting part of the project, because it isn't exciting, and then it becomes the reason the whole thing stalls.

And underneath all of that sits the genuinely unglamorous stuff: logging, monitoring, retry logic, rate limits, cost controls (a model that costs four cents per demo call costs a frightening amount per month at real volume if nobody designed for it), access control, and a credible answer to what your compliance team is going to ask the first time they hear "we're sending customer data to a third-party API." None of this shows up in a thirty-minute pilot. All of it shows up in the first month of running for real.

The organisational problem hiding inside the technical one

Here's the bit that's easy to miss if you're only looking at the tech. A lot of POCs stall for reasons that have nothing to do with engineering at all.

The pilot usually gets built by whoever was curious enough and senior enough to get the budget — often a small, motivated team operating slightly outside the normal process, deliberately, because that's the only way anything moves fast inside a big organisation. That's a feature during the pilot phase. It becomes the exact reason it stalls afterward, because taking the thing into production means handing it to the people who run "normal," and normal has change-control boards, security review cycles, and a backlog that was already full before this showed up uninvited.

Nobody owns the handoff. The person who built the pilot isn't usually the person who runs production systems, and the people who run production systems weren't in the room when the pilot got approved, so they have no context, no investment in it, and a long list of reasons to be cautious about something they didn't choose. We've watched a genuinely good pilot sit untouched for four months purely because nobody could agree whose backlog it belonged on.

And there's a quieter version of this that's almost worse: the pilot succeeds technically and then nobody can agree what "success" actually means for shipping it. Did it save time? Whose time, measured how? Is a 90% accurate answer good enough if the other 10% has to go to a human anyway — and if so, did it actually save anything, or just move the work around? These should be questions asked before the pilot starts. They're almost always asked, if at all, after it's already built, when the answers are far more political and far less honest.

What actually separates the ones that make it

The pattern in the ones that do reach production isn't more impressive technology. It's that someone treated the production path as a real piece of work from day one, not an afterthought to be sorted out once the demo gets applause.

That means deciding, before a line of code gets written, what "good enough to ship" actually looks like in numbers, not vibes — and who signs off on it. It means designing the failure path alongside the success path: what happens when the model is wrong, who sees it, what they do about it. It means putting the pilot in front of the messiest data you can find, deliberately, as early as possible, instead of being delighted by how well it does on the tidy sample — because the tidy sample was never going to be the problem.

It also means picking, early, who owns this once the interesting part is over. Not symbolically. Actually owns it — has it on their plate, gets asked about it, is accountable when it breaks at 2am. Pilots without a named production owner from week one are, in our experience, the ones still sitting in the downloads folder a year later.

And it means being honest, uncomfortably early, about cost at scale. The economics of a model call in a demo and the economics of the same call multiplied by your actual transaction volume are not the same conversation, and finding that out in month four of a production rollout is a bad time to find it out.

The boring truth at the centre of this

None of the above is exciting. There's no part of "design the failure path" or "agree what good enough means in numbers" that gets applause in a steering committee meeting. That's exactly why it gets skipped, and exactly why skipping it is so expensive.

The companies getting real value out of AI right now aren't the ones with the flashiest pilots. They're the ones willing to spend time on the unglamorous middle — the data cleanup, the failure handling, the ownership question, the cost model — before they let anyone near a launch date. The pilot proves the idea works. Everything after that proves you can actually run it. Those are different jobs, and treating them as the same one is the most expensive mistake in AI right now, because it's invisible until the bill arrives.


Plexa Solutions helps companies take AI from an interesting pilot to something that actually runs — the unglamorous middle included. If you've got a POC stuck in the downloads folder, we'd be glad to take a look. plexasolutions.ie

Got a pilot stuck in the downloads folder?

We help companies take AI from an interesting demo to something that actually runs. Let's talk.

Book a scoping call