Choosing / Procedure
Comparing Assistants Properly
Rankings go stale within a quarter. The durable comparison is about interface, integration, data policy and how each one fails.
Any article naming the best assistant is out of date by the time you read it. The models leapfrog each other continuously and the differences on general tasks have narrowed to the point where preference matters more than capability.
What does not change quickly is everything around the model.
The dimensions that stay relevant
Data policy. Does the provider train on your inputs by default? Can you turn it off? What is retained and for how long? Is there a business tier with contractual commitments? This varies substantially and it is the difference between a tool you can use for work and one you cannot.
Retrieval quality. Does it search, when, and does it show sources you can check? For factual work this matters more than raw model capability.
Context handling. How much material can you give it, and does it actually use the middle of a long document.
File handling. Which formats, how large, and does it read documents properly or extract text badly.
Code execution. Can it run code to check its own arithmetic and data work. This is a genuine accuracy improvement, not a convenience.
Where it lives. In your browser, your editor, your documents, your terminal. Integration determines whether you actually use it.
Export and history. Can you get your conversations out.
Failure behaviour. Does it decline clearly or fabricate confidently. This is a stable personality difference between products and it matters more than benchmark scores.
Running your own comparison
Trials are free. A structured hour is worth more than any review.
Assemble five representative tasks from your actual work. Include a long document, something factual and checkable, something creative, something structured, and something in your domain jargon.
Run all five through each candidate, same inputs.
Score for usefulness, not impressiveness. How much editing before you could use it.
Check the checkable one. Facts, figures, citations.
Note the failures specifically. How each one behaves when it does not know is the most informative part of the exercise.
Then check the data policy for whichever won.
What benchmarks do and do not tell you
Published benchmarks measure performance on standardised tasks, mostly mathematics, coding and multiple-choice knowledge.
They are real measurements and they correlate with capability.
They correlate weakly with your experience unless your work resembles the benchmark. A model that is better at competition mathematics may be no better at drafting your reports.
They are subject to contamination. Benchmark content leaks into training data, inflating scores without capability gains. Providers work to prevent this and it is difficult to verify.
They are chosen selectively by marketing. A provider quotes the benchmarks it leads on.
Human preference rankings — where people compare anonymous outputs — track general usefulness better than academic benchmarks, and they reward style as much as substance.
Use benchmarks to rule out the clearly weaker options, not to choose between the leaders.
The multi-tool position
Many people who use these tools heavily keep two.
The reason is failure decorrelation. When one produces something that seems wrong, asking another is the fastest available check. Different training and different tuning means they fail differently, and agreement between two is weak but real evidence.
The cost is two subscriptions, which for professional use is usually trivial against the time saved.
The alternative is one paid and one free tier for cross-checking, which covers most of the benefit.
What to re-examine, and when
Every six months is enough. More often is churn.
Trigger a re-examination when: your provider changes its data policy, a capability you specifically need appears elsewhere, pricing changes materially, or you notice yourself working around a limitation repeatedly.
Do not switch on a launch announcement. Wait for people to use the thing for a month. Launch benchmarks and lived experience diverge routinely.
Building the five-task set
The comparison is only as good as the tasks, and most people assemble them badly by picking things that are easy to score rather than things they actually do.
Take them from last week's work. Open your sent folder or your document history and pick five real items.
Include one long document task. Something over ten pages, with a question that requires reading the middle rather than the summary.
Include one checkable factual task, where you already know the answer.
Include one task in your domain vocabulary, with the jargon and the assumptions of your field intact.
Include one messy input — a badly formatted spreadsheet, a transcript with crosstalk, a scanned document.
Include one task where you care about voice.
Keep the set. Run it again in six months against whatever has launched. A stable benchmark of your own work is worth more than any published comparison, and it takes an hour to build once.