A demo answers one question: can this be built? Production answers a harder one: will it hold when real customers, real money and real edge cases arrive at once?
Those are different pieces of work, and the gap between them is where most AI spending disappears. MIT's 2025 Project NANDA study found that 95% of enterprise AI pilots deliver no measurable return. Not because the idea was wrong — because the demo was mistaken for the product.
What "AI-powered" actually claims
Nothing. It describes an ingredient, not an outcome. A feature that calls a model once and prints the result is AI-powered. So is a system with retries, rate limits, evaluation harnesses, access boundaries and a human review step. The label covers both, which is exactly why it is used.
Compare it to how you would judge anything else you buy. Nobody sells a car as "engine-powered". They tell you what it does, how fast, how safely, and what happens when something fails. AI marketing skipped that step, and buyers have not yet insisted on it.
The question that separates them
There is one question that cuts through it: show me what happens when it's wrong.
A demo has no answer. It was built on the assumption that the model behaves. A production system has a specific answer for each failure: the model returns nonsense, the API times out, the user pastes 40,000 words, two customers' data sits in the same index, the output is confidently incorrect and someone acts on it.
Each of those is a decision someone made, or did not. The ones nobody made are the ones that surface later, usually in front of a customer.
Why the gap is expensive rather than merely annoying
Because the failures are rarely cosmetic. Veracode's analysis found around 45% of AI-generated code introduces a flaw from the OWASP Top 10 — the standard list of ways software leaks or gets taken over. That is not a bug rate. It is a security base rate, and it applies to code that looks completely fine.
The public examples follow the same shape. The Lovable CVE exposed data across 170+ applications built on the platform. The Tea app breach put roughly 72,000 sensitive images in the open. An AI agent with production database access at Replit deleted the database. In each case the software worked, right up until the moment the thing nobody had considered happened.
What to ask a vendor instead
Drop "do you use AI" — everyone does. Ask what happens on failure, and whether they can show you the code path that handles it. Ask what a customer sees when the model is wrong, who reviewed it before release, and what is written down about that review.
You will learn more from the answer to the last question than from any demo. A team that has crossed into production has documents. A team that has not will describe intentions. If you would rather have someone else ask them, that is what the Reality Check is.
The summary
"AI-powered" is not a feature. Working in production is. Most people are shipping the first and calling it the second, and the market has not yet built the vocabulary to tell them apart. Until it does, the burden sits with the buyer — which is unfair, and true.
This is the part we do — the crossing from a demo to a system that survives production.