Choosing / Procedure
Evaluating an AI Tool in 20 Minutes
A repeatable test that separates products doing real work from interfaces passing your text to someone else's model with a markup.
New AI products appear faster than anyone can assess them, and the marketing is uniformly excellent. A short structured test tells you more than an hour of reading the site.
Before you sign up
What model is underneath? If the product does not say, that is informative. Most tools are an interface over a major provider's API. That is legitimate — the value can be in the interface, the workflow, or the data — but it sets the price you should expect to pay.
What does it do that the underlying assistant does not? If the honest answer is "a nicer prompt", you are paying a markup for a prompt.
Where does your data go, and what is done with it? Look for a plain statement about training on your inputs, retention, and sub-processors. Vagueness here is the most reliable negative signal in the entire evaluation.
Is there a free tier or trial without a card? Products confident in their value usually let you see it.
How long has it existed, and who is behind it? This field is full of products that will not exist next year, and migrating away from one is real work.
The twenty-minute test
Use your own material, not the examples they provide.
One: the representative task. Give it something you actually do, of realistic size and messiness. Not a clean sample.
Two: the hard case. Something at the edge of what you would want it to handle. This is where products separate.
Three: the wrong input. Something outside its purpose. A good product declines clearly. A bad one produces confident nonsense, which tells you it has no idea what it is doing.
Four: the same input twice. How much does the output vary? High variance means you cannot rely on it in a workflow.
Five: check a fact. If it produces anything checkable — a figure, a citation, a claim — check it. This single step disqualifies a surprising number of products.
Six: the export. Can you get your work out, in a usable format, right now? A tool you cannot leave is a liability regardless of quality.
What to look for in the answers
Does it fail well? Saying "I cannot do this reliably" is worth more than a confident wrong answer, and almost no product does it.
Does it show its sources? For anything factual, unsourced output is unverifiable output.
Is the output in a form you can use, or does it need reformatting every time?
Does it handle your actual mess — inconsistent formatting, mixed languages, long documents, domain jargon?
The pricing questions
What counts against the limit? Messages, tokens, "credits", or something undefined. Credits that do not map to anything measurable are a warning.
What happens at the limit? Blocked, throttled, or silently downgraded to a weaker model. The last is common and rarely disclosed.
Does the price scale with your use or your team? Per-seat pricing for a tool used occasionally by many people is expensive in a way that is not obvious at signup.
Is there an annual lock-in? In a field moving this fast, a year is a long commitment.
Signals that a product is thin
No statement about which model it uses, anywhere.
Marketing entirely in outcomes — "10x your productivity" — with no description of mechanism.
Benchmark claims with no methodology.
A demonstration video that never shows a failure.
Testimonials without named organisations.
No export.
Pricing that requires a sales call for a small team.
A privacy policy that permits training on your content with no opt-out.
None is disqualifying alone. Three or more together usually means a wrapper.
The question that settles most cases
Could I do this with a general assistant and a saved prompt?
For a large share of AI products, the answer is yes. That does not make them worthless — a good interface, a saved workflow and team access have real value. But it tells you what you are buying and what it should cost.
If the answer is no — because the product has proprietary data, a genuine integration, a specialised model, or does real work between your input and the model — that is a product rather than a wrapper, and it is worth paying for.
A one-page record
Keeping a short written record of each evaluation is worth the five minutes, because in a fast-moving field you will otherwise re-evaluate the same product twice and forget why you rejected it.
Date and version tested. Both change quickly, and a rejection from a year ago may no longer hold.
What you tested it on, specifically, so a future comparison is like for like.
What it did well and badly, in a sentence each.
The pricing at the time and what counted against the limit.
The data policy as it stood.
The decision and the reason.
A trigger for reconsidering — a capability that would change the answer, a price point, a policy change.
Six lines. When the product relaunches with a large announcement, you can check whether anything relevant actually changed rather than starting the evaluation over.