Skip to content
Technology Munch

Writing  / Analysis

What Machine Translation Misses

Good enough for understanding, risky for publishing, and dangerous in specific situations. Where the line falls and how to work across it.

Translation is one of the strongest applications of these models, and the gap between "I understood the document" and "this is publishable" is larger than the fluency of the output suggests.

Where it is genuinely excellent

Comprehension. Reading something in a language you do not know. This works well and it has changed what is accessible to people.

Common language pairs with abundant training data. Major European languages, and increasingly a much wider set.

General and technical prose where the register is neutral.

Getting the gist quickly of a long document before deciding whether to have it translated properly.

Where it fails

Low-resource languages. Quality drops sharply for languages with less text available. This is the largest and least discussed inequality in the field, and output for some languages is confidently fluent and substantially wrong.

Idiom and cultural reference. Rendered literally or replaced with an approximation that changes the tone.

Register and formality. Languages with grammatical politeness levels — the difference between formal and informal address — are handled inconsistently. Getting this wrong is not a small error; it can be an insult.

Names and terminology. Product names get translated, technical terms get replaced with near-synonyms, established translations of proper nouns are ignored.

Ambiguity. Where the source is genuinely ambiguous, the model picks an interpretation without flagging it.

Legal and contractual language. Precision here is the entire point, and small shifts change obligations.

Humour, poetry, and anything where the form carries meaning.

The specific danger

Fluent output conceals error. Bad translation used to read badly, which warned you. Machine translation reads well and is sometimes wrong. A reader who does not know the source language has no signal.

This matters most where the stakes are real: medical instructions, legal documents, safety information, financial terms. Do not use machine translation unreviewed in any of these contexts. The fluency is not evidence of accuracy.

Working with it responsibly

Know what the output is for. Comprehension, or publication? These need different levels of care.

Back-translate as a check. Translate the result back and compare with your original. Catches meaning shifts, misses register problems.

Have a native speaker read anything published. Not a full retranslation — a read for register, idiom and obvious errors. This is far cheaper than translation and catches most of what matters.

Provide context. Telling the model what the document is, who the audience is, and what register to use improves output substantially. "Translate this customer email into formal Japanese" is much better than "translate this".

Supply a glossary. Product names, technical terms, things not to translate. Most tools accept this and it eliminates a whole class of error.

Translate whole documents, not fragments. Context resolves ambiguity.

Never translate legal or medical text for use without professional review.

Comparing tools

Dedicated translation services are frequently better for straightforward translation between well-supported languages, and they handle document formats better.

General assistants are better where you need context, explanation, register control or a glossary applied, because you can instruct them.

On-device translation on phones is now good for comprehension and keeps the text private, which matters for confidential material.

For low-resource languages, test carefully and assume the quality is worse than it reads.

What to use it for without hesitation

Reading foreign-language sources. Understanding an email. Getting the sense of a document. Drafting something a native speaker will review. Communicating informally with someone who will forgive imperfection.

That covers most personal use and it is a genuine expansion of what people can do.

The care applies at the boundary where someone acts on the translation without being able to check it — and that boundary is where the fluency does the most damage.

Checking a translation you cannot read

Practical verification when you do not speak the target language and cannot hire a translator for everything.

Back-translate with a different tool. Using the same tool to translate back tends to reproduce its own errors. A second system is a genuine check.

Compare the lengths. A translation dramatically shorter than the source has probably dropped something.

Check that names, numbers and dates survived unchanged. These are the most common silent errors and you can verify them without knowing the language.

Check formatting and structure carried over — headings, lists, emphasis.

Ask the model to list anything it was uncertain about, and anything ambiguous in the source. This produces useful flags surprisingly often.

For anything published, have one native speaker read it. Not retranslate — read, for register and obvious error. Twenty minutes of someone's time, and it catches what none of the above will.