Four-Model Translation Showdown: Week 39 Quality Evaluation, claude-sonnet-4.6 Leads with 9 Points

This week, 454 translation tasks were completed by 4 models. A blind multi-model comparison of 3 sampled articles found claude-sonnet-4.6 to be the best overall, with an average score of 9/10.

This week, 454 translation tasks were completed by 4 models. A sample of 3 articles was selected for multi-model blind comparison; the best overall: claude-sonnet-4.6 (average score 9/10).

This Week's Translation Statistics

ModelLanguageTranslation VolumeAverage TimeAverage Quality Score
deepseek-v4-flashen7127.7sNot rated
claude-sonnet-4.6ja22648.9sNot rated
passthroughen1520sNot rated
deepseek-v4-flash:backtranslateen1142.5sNot rated
native-englishen2-Not rated
deepseek-v4-flashzh27sNot rated

Sample Comparison Evaluation

Evaluation 1: White House Opposes Regulation: Why Washington Is Slow to Move on AI Legislation

ModelAccuracyFluencyTerminologyReadabilityTotal Score
claude-sonnet-4.699899
deepseek-v4-pro88988
gpt-o399999

claude-sonnet-4.6

✓ Paragraph transitions are natural; for example, 「手続き上の障壁も無視できない」 flows logically into the following text.

✗ The term 「技術の世代交代」 is somewhat broad and does not emphasize the meaning of “iteration.”

deepseek-v4-pro

✓ Terminology consistency is good; for example, 「冗長弁論」 accurately corresponds to filibuster.

✗ Some expressions are slightly stiff; for example, 「技術の反復が「月」単位で進む」 has a somewhat abrupt rhythm.

gpt-o3

✓ Terminology choices are precise; for example, 「イテレーション」 and 「フロンティアモデル」 both fit the AI context.

✗ Compared with Version A, some sentence patterns are slightly long and the pacing is a bit slower.

Conclusion: The three versions are close in overall quality; Version C is slightly better in terminology accuracy and fluency, and we recommend prioritizing Version C.

Evaluation 2: A Startup's Next Teammate May Be an AI Agent

ModelAccuracyFluencyTerminologyReadabilityTotal Score
claude-sonnet-4.699989
deepseek-v4-pro89898
gpt-o398998

claude-sonnet-4.6

✓ Fluency is excellent; for example, the shift from 「呼び出される機能」 to 「役割を与えられた存在」 is expressed naturally and fits the logic of the original text.

✗ Readability is slightly weaker; paragraph transitions are a bit verbose. For example, the end of the third paragraph, “9月25日までに登録すると最大200ドルの割引が適用される,” does not transition tightly enough from the preceding text.

deepseek-v4-pro

✓ Readability is best; the structure is clear. For example, the subheading “身分のアップグレード” is concise and forceful, with natural logical connections.

✗ Terminology consistency is slightly weaker; “人員として数えるのか” and the later “人員数” are not expressed consistently enough.

gpt-o3

✓ Accuracy is high; for example, “アイデンティティのアップグレード” faithfully conveys the meaning of “role upgrade” in the original text.

✗ Fluency is slightly inferior; some sentences have a slight translationese flavor. For example, “すなわち…という問いだ” sounds rather formal.

Conclusion: The three versions are close in overall quality. claude-sonnet-4.6 is slightly better in accuracy and fluency and suits formal scenarios; deepseek-v4-pro has the best readability, while gpt-o3 handles terminology robustly.

Evaluation 3: AI Virtual Actor Tells Me “All Life Is Important,” Then Comments on My Clothes

ModelAccuracyFluencyTerminologyReadabilityTotal Score
passthrough54655
deepseek-v4-pro87877
gpt-o399999

passthrough

✓ It largely preserved the original HTML links and structure, such as 「<a href="https://www.wired.com/tag/hollywood/">Hollywood</a>」, without additional changes.

✗ The text is severely incomplete; it ends abruptly with 「it s」, which is a clear omission and formatting error.

deepseek-v4-pro

✓ It freely translates “stuck needle spinning around in the same groove” as 「像一根卡住的唱针在同一道沟里打转」, which is both accurate and vivid.

✗ Some sentences are slightly long; for example, the handling of the political topic in the opening paragraph is somewhat stiff, and fluency is not as good as Version C.

gpt-o3

✓ It renders “stuck record needle circling the same groove over and over” as 「像一根卡住的唱针一次次绕着同一道沟里打转」, which is natural, idiomatic, and rhythmic.

✗ There are no obvious flaws; only in very few places could redundant description be tightened further.

Conclusion: Version C is the best overall, leading in accuracy, fluency, and readability; Version B is next; Version A is not recommended because it is incomplete.