Choosing
Judging a tool without a leaderboard
Rankings go stale in a quarter. What lasts is knowing how to test a product on your own work, what a free tier costs you, and which questions a vendor cannot dodge.
Evaluating an AI Tool in 20 Minutes
A repeatable test that separates products doing real work from interfaces passing your text to someone else's model with a markup.
Procedure
Most AI products are an interface over someone else's model. Sometimes that interface is worth the money and often it is a prompt.
Analysis
Rankings go stale within a quarter. The durable comparison is about interface, integration, data policy and how each one fails.
Procedure
Free AI access is funded somehow. Knowing which model of funding you are inside determines what you should put into the box.
Analysis
Specialised Tools vs General Ones
General assistants have absorbed most of what narrow AI products did. Four conditions still favour the specialist; the rest is cancellable.
Analysis
Local models are private, free per use, and slower and weaker than hosted ones. The gap has narrowed enough that the calculation has genuinely changed.
Analysis
Before subscribing to anything new, check what your existing tools now include. A significant amount of it arrived without announcement and is unused.
Reference
Benchmark scores are real measurements of narrow things. What they predict about your work, and the specific ways the numbers are made to look better.
Reference
Memory is the constraint, not processing speed. What each tier of hardware actually runs, and why the obvious upgrade is frequently the wrong one.
Reference
Real AI Product or Demonstration?
This field produces more impressive demonstrations than working products. The signals that distinguish them, and the questions vendors cannot dodge.
Checklist
10 notes in this section