Writing / Reference
Transcription and Voice Tools
Automatic transcription became genuinely good, with specific failure patterns that matter if anyone relies on the transcript.
Speech recognition improved dramatically and quietly. For clear speech in a well-supported language it is now accurate enough that correcting a transcript is faster than typing one.
That is a real change, and the remaining failures are predictable.
Where it works well
Clear single-speaker audio in a major language. Dictation, voice notes, presentations.
Recorded interviews with decent microphones.
On-device transcription on modern phones and computers, which is both fast and private.
Where it degrades
Multiple speakers, especially overlapping. Speaker separation is the weakest part of most systems. Crosstalk produces attribution errors, and a quote assigned to the wrong person is worse than a missing quote.
Accents underrepresented in training data. Accuracy varies by accent in ways that map onto familiar inequalities. This is well documented and improving slowly.
Names and jargon. Proper nouns, technical terms, product names, acronyms. Consistently the largest error category in professional transcripts.
Poor audio. Room echo, distance from the microphone, background noise, phone quality.
Code-switching. Mixing languages within a sentence, which is normal for many speakers and handled badly.
Numbers and figures. Frequently transcribed wrongly in ways that look plausible.
The failure that matters
Transcription errors are usually plausible words rather than gibberish. A misheard name becomes a different real name. A figure becomes a different figure.
A reader cannot detect this from the transcript alone. If someone will rely on the text — a quote in an article, a decision from a meeting, a medical or legal record — the transcript needs checking against the audio at the points that matter.
Check names, numbers and quotations specifically. These are where errors concentrate and where they cost most.
Summaries of transcripts
Most meeting tools now produce a summary and action items on top of the transcript. This compounds two error sources: transcription errors feed into a summary that then discards the detail you would need to notice them.
Summaries are useful and are not a record. Keep the transcript, and keep the audio for anything consequential.
Action items are frequently wrong in a specific way: they attribute a task to whoever spoke last about it rather than whoever agreed to it. Worth checking before anyone relies on the list.
Privacy, which is the part people skip
Recording others generally requires consent, and the rules vary substantially by jurisdiction. In some places one party's consent suffices; in others everyone must agree. For calls that cross borders, the stricter rule is the safe assumption.
Announce recording. Beyond the legal position, a meeting tool that joins silently and transcribes is a trust problem.
Where does the audio go? Cloud transcription sends the recording to a provider. For confidential meetings, client conversations, medical or legal discussions, that is a disclosure. Check the terms and check whether transcripts are used for training.
On-device transcription avoids this entirely and is now good enough for many purposes. For sensitive material it is the right default.
Retention. Meeting tools accumulate a searchable archive of everything anyone said. That is an asset and a liability, and most organisations have not decided which.
Practical setup
Better audio beats better software. A cheap dedicated microphone improves accuracy more than switching tools.
Supply a vocabulary list where the tool allows it — names, terms, acronyms. This eliminates the largest error category.
Record separate tracks per speaker where you can, which solves the attribution problem entirely.
Keep the audio, at least until the transcript has served its purpose.
Decide the retention policy before you accumulate two years of recorded meetings.
Improving accuracy before the software runs
The largest gains come from the recording, not the tool, and they are cheap.
Get the microphone closer. Distance is the single biggest determinant of accuracy. A lapel microphone or a phone on the table beats a laptop across the room by a wide margin.
Record separate tracks per speaker where the platform allows it. This eliminates the attribution problem entirely, which is the failure that matters most.
Reduce the room. Soft furnishings, closed windows, away from air conditioning.
Supply a vocabulary list to the tool — names of participants, product names, acronyms, technical terms. Most services accept this and it removes the largest error category.
Ask people to say their name before speaking in the first few minutes, which helps speaker labelling considerably.
Record a clean sample of each voice at the start where the tool supports voice profiles.
None of this requires equipment beyond what most people have.