4 Major Models Translation Showdown: Week 26 Quality Review, claude-sonnet-4.6 Leads with Score 9

This week, 393 translation tasks were completed by 4 models. A multi-model blind evaluation of 3 sampled tasks showed claude-sonnet-4.6 achieved the best average score (9/10).

This week, 393 translation tasks were completed by 4 models. A blind comparison of 3 sampled tasks across multiple models was conducted. Top overall: claude-sonnet-4.6 (average score 9/10).

Weekly Translation Statistics

ModelLanguageTranslation VolumeAverage TimeAverage Quality Score
deepseek-v4-flashen7214sNot rated
claude-sonnet-4.6ja19633sNot rated
passthroughen1230sNot rated
native-englishen1-Not rated
deepseek-v4-flashzh110.7sNot rated

Sampled Comparative Evaluation

Evaluation 1: Hands-on with Siri AI: The New Evolution of Conversational Intelligent Assistants

ModelAccuracyFluencyTerminologyReadabilityTotal Score
passthrough26733
deepseek-v4-pro88888
gpt-o399999

passthrough

✓ Preserved some original English links and structure

✗ Severely incomplete, interspersed with large amounts of HTML code and untranslated content, e.g., truncated directly at "Since <a href="

deepseek-v4-pro

✓ Idiomatic translation of "From 'Ask and Answer' to 'Read Your Mind'" is natural and fluent

✗ JSON format traces appear at the beginning, breaking overall coherence

gpt-o3

✓ Paragraph transitions and citation handling are the clearest, e.g., "Siri AI is no longer just a tool for executing commands" is translated accurately and naturally

✗ A very few long sentences are slightly formal

Conclusion: gpt-o3 version is the best overall, deepseek-v4-pro is second, the passthrough version has no reference value

Evaluation 2: Claude Fable 5 and Mythos 5 Removed Globally on June 12 — Security Verification Requirements and Privacy Controversy Coexist

ModelAccuracyFluencyTerminologyReadabilityTotal Score
claude-sonnet-4.698999
deepseek-v4-pro87877
gpt-o399999

claude-sonnet-4.6

✓ Full text structure is complete, paragraph transitions are natural. "官方アナウンスによると" directly corresponds to the original official statement, logic is clear.

✗ Some long sentences are slightly verbose, e.g., "この要件が一部地域でのユーザー離れを直接引き起こした" could be more concise.

deepseek-v4-pro

✓ Terminology such as "脱獄プロンプト" is used consistently, and "販売中止" in a business context closely matches the original meaning of "下架."

✗ JSON format wraps the body text, and the ending is clearly truncated. "クリエイティブおよびプロトタイピングのシナリオでは" is incomplete, affecting readability.

gpt-o3

✓ Most natural language, "段階的リリース" and "連携経路を直接中断させた" are idiomatic and faithful, with high consistency of "提供停止" throughout the text.

✗ Also has an ending truncation issue, "クリエイティブ用途やプロトタイプ開発のシナリオで" is unfinished.

Conclusion: Versions A and C are close in overall quality and better than Version B. Version C is recommended for priority use due to slightly superior fluency and terminology consistency. Version B is not recommended due to formatting issues and truncation.

Evaluation 3: Leaked Financial Reports Show OpenAI Loses Billions Annually

ModelAccuracyFluencyTerminologyReadabilityTotal Score
passthrough910999
deepseek-v4-pro78787
gpt-o389898

passthrough

✓ Directly uses the original English report, language is natural and fluent. For example, "newly leaked financial documents show a company with quickly growing revenues" accurately conveys the contrast between growth and losses.

✗ Content is severely truncated, lacking complete paragraphs and subsequent analysis, resulting in incomplete information.

deepseek-v4-pro

✓ Clear structure, with added subheadings and citations, e.g., "Losses exceed expectations" makes for easy reading.

✗ Numerical deviations and excessive paraphrase occur. For example, the original text does not mention "40% revenue increase"—content has been added.

gpt-o3

✓ Consistent terminology and natural citation translation, e.g., "You either choose scale or you are out" preserves the original meaning well.

✗ Some expressions are slightly stiff, and the description of R&D spending share deviates slightly from the original details.

Conclusion: Version A is closest to the original report but incomplete. Version C has the best overall balance. Version B contains obvious numerical mistranslations and added content. It is recommended to use Version C or the complete version of A.