Skip to content
Technology Munch

Choosing  / Checklist

Real AI Product or Demonstration?

This field produces more impressive demonstrations than working products. The signals that distinguish them, and the questions vendors cannot dodge.

The gap between what an AI product shows and what it does is wider than in any other software category, because the technology genuinely produces spectacular results on selected examples and unreliable results on real ones.

What a demonstration hides

Selection. The example was chosen after trying many. The failures are not shown.

Preparation. Clean input, formatted as the system prefers.

Editing. A video of a five-minute task compressed to thirty seconds.

Best of several attempts.

A human in the loop who does not appear.

None of this is dishonest by the standards of software marketing. It does mean a demonstration tells you the ceiling and nothing about the floor.

The signals

No statement of what model is underneath. Companies doing real work generally say. Silence usually means an API wrapper.

Marketing entirely in outcomes. "Transform your workflow", "10x productivity". No description of what happens between input and output.

Benchmark claims with no methodology. Which benchmark, what conditions, single attempt or best of many, evaluated by whom.

No failure shown anywhere. Documentation, demonstrations, marketing. A product with real deployments has known limitations and mature vendors publish them.

"Accuracy" quoted without a task. Accuracy on what, measured how, against what baseline.

No free trial without a sales call. For a small team purchase, this is a signal about who the product is really sold to.

No export.

Testimonials without named organisations and roles.

A waitlist that never ends.

Pricing that is not published.

Three or more together is a reliable pattern.

The questions vendors cannot dodge

"What does it do when it does not know?" A real product has an answer — it flags uncertainty, it declines, it escalates. A demonstration has no answer because the question never arose.

"Show me a failure case." Mature vendors have these ready. It is the single most informative question you can ask.

"What is the error rate on our kind of data, and how was it measured?"

"What happens between my input and the model?"

"Can I run it on my own data during evaluation?" Reluctance here is decisive.

"Who has deployed this at our scale, and can I speak to them?"

"What does the human review process look like?" Any product making consequential outputs needs one, and vendors who have not thought about it have not deployed anywhere serious.

The evaluation that settles it

Your data, your people, your hard cases, over a realistic period.

Include the messy input. The badly scanned document, the inconsistent spreadsheet, the domain jargon. This is where products separate and where demonstrations never go.

Measure against what you do now, not against nothing. Many AI products are compared to a baseline of doing the task perfectly, when the real comparison is to a person doing it in twenty minutes with a 3 percent error rate.

Count your own hours. Integration effort is the most underestimated cost and it rarely appears in a comparison.

The category to be most careful with

Anything making consequential decisions about people — hiring, lending, assessment, risk scoring, monitoring.

These attract confident claims, they are hard to evaluate, the errors fall unevenly, and the regulatory position is tightening. The questions above apply with more force, and two more: what is the audit trail, and can a decision be explained to the person it affects.

A vendor without good answers to those two should not be selling into that space, and buying from them transfers the problem to you.

The pilot that tells you something

If a product survives the questions, the pilot should be designed so it can fail.

Define success before starting, in a sentence with a number in it. Not "evaluate the tool" — a threshold you would accept.

Use your own data, at realistic volume, including the awkward cases.

Have your people operate it, not the vendor's.

Run it long enough that novelty wears off. Two weeks measures enthusiasm; two months measures usefulness.

Count your own hours, honestly. Integration is the cost everyone underestimates.

Compare against what you do now, not against perfection.

Write the decision down with the reasoning, so that when the vendor returns in six months with a new version you can check whether anything relevant changed.

A pilot that cannot fail is a procurement formality, and it is what most AI pilots are.