5-Model Translation Face-Off: Week 37 Quality Review — claude-sonnet-4.6 Leads with 9 Points

This week's 368 translation tasks were handled by 5 models. A 3-article multi-model blind review crowned claude-sonnet-4.6 the overall best, with an average score of 9/10.

This week, 368 translation tasks were completed by 5 models. A sample of 3 articles underwent multi-model blind comparison; overall best: claude-sonnet-4.6 (average score 9/10).

Weekly Translation Statistics

ModelLanguageVolumeAvg. TimeAvg. Quality Score
deepseek-v4-flashen7860.8sUnrated
claude-sonnet-4.6ja18349sUnrated
passthroughen1020sUnrated
native-englishen2-Unrated
deepseek-v4-flashzh248.6sUnrated
deepseek-v4-flash:backtranslateen125.9sUnrated

Sampled Comparative Evaluation

Evaluation 1: AfterQuery Sets YC's Fastest Unicorn Record, Valued at $3.2 Billion

ModelAccuracyFluencyTerminologyReadabilityTotal
passthrough98978
deepseek-v4-pro68787
gpt-o369797

passthrough

✓ Most faithful to the original. "This just five months after announcing its $30 million Series A at a $300 million valuation in April" directly mirrors the original's time and valuation comparison, with no added content.

✗ The final paragraph is clearly truncated — "However, rather than ensuring th" is cut off — undermining readability.

deepseek-v4-pro

✓ The headline is handled cleanly. "AfterQuery Sets YC Record for Fastest Unicorn, Valued at $3.2 Billion" gets straight to the core information.

✗ Adds a large amount of content not present in the original, such as "According to TechCrunch" and "oversubscribed within days of launching" — excessive paraphrasing.

gpt-o3

✓ Paragraph transitions flow smoothly, and the subheading "Why Is This Company So Popular With Investors?" makes the structure clearer.

✗ Also adds "oversubscribed within days of launch" and descriptions of investor preferences not found in the original, deviating from the source.

Conclusion: Version A is the most accurate, but its truncation hurts completeness; B and C are more fluent yet both contain considerable added content. Overall, A comes closest to the original text.

Evaluation 2: Grok 4's Code Execution Score Plunges 21.8 Points; Main Leaderboard Drops from 89.15 to 84.99

ModelAccuracyFluencyTerminologyReadabilityTotal
claude-sonnet-4.698999
deepseek-v4-pro87787
gpt-o398898

claude-sonnet-4.6

✓ The use of "素材制約" is highly consistent with the original's "素材约束," and the paragraph structure is clear; quoted passages such as "コード実行の問題で1問失敗するだけで" accurately correspond to the original meaning.

✗ The translation is clearly truncated at the end — "Grok 4を全体として過大評価または過" is unfinished — affecting overall completeness.

deepseek-v4-pro

✓ Clear structure with natural flow between headline and body; for example, "メインランキングはコード実行と材料制約の加重のみで構成される" is logically clear.

✗ The term "材料制約" is inconsistent with the original's "素材约束"; repeated use may undermine professionalism, and some sentences read stiffly.

gpt-o3

✓ The causal analysis flows smoothly; phrasing such as "問題のランダム抽出による変動か" is natural and faithful to the original logic.

✗ It also suffers from truncation, with content after the final "問題" incomplete; its term "材料制約" is also slightly weaker than Version A's.

Conclusion: Version A performs best on terminology accuracy and structural readability but loses points over the truncation; B and C are close overall, with C slightly better than B. Preference should go to A (once the truncation is fixed) or C.

Evaluation 3: WDCD Run #311: Grok 4 Leads with 93.2 Points as GLM-4.6 Sets New Decay Resistance Record

ModelAccuracyFluencyTerminologyReadabilityTotal
native-english99989
deepseek-v4-pro21111
gpt-o388978

native-english

✓ Terminology is accurate and professional; for example, "Winzheng Dynamic Contextual Decay (WDCD)" fully preserves the benchmark name and spells out its full form.

✗ The text is clearly truncated near the end — content after "At the opposite end of" is missing — affecting readability.

deepseek-v4-pro

✓ No notable strengths; it flat-out refused the translation task.

✗ It provided no translation whatsoever, only outputting the error message "未能找到中文文本" — a severe omission.

gpt-o3

✓ Terminology consistency is solid; for instance, "multi-turn commitment resistance" corresponds accurately to the source meaning.

✗ It also suffers from truncation — the final paragraph is clearly unfinished, ending abruptly at "original constrain" — affecting overall fluency.

Conclusion: Version A has the highest overall quality, Version B is a complete failure, and Version C is close to A but equally truncated.