3 Major Models Translation Showdown: Week 40 Quality Evaluation, deepseek-v4-pro Leads with 8.7

This week, 417 translation tasks were completed by three models. A sample of three was blind-evaluated across multiple models, with deepseek-v4-pro earning the best overall average score of 8.7/10.

This week, 417 translation tasks were completed by 3 models. 3 were sampled for multi-model blind evaluation comparison; overall best: deepseek-v4-pro (average score 8.7/10).

This Week's Translation Statistics

ModelLanguageTranslation VolumeAvg. TimeAvg. Quality Score
deepseek-v4-flashen6323.8sNot rated
claude-sonnet-4.6ja20848.1sNot rated
passthroughen1440sNot rated
native-englishen1-Not rated
deepseek-v4-flashzh14.1sNot rated

Sampled Comparative Evaluation

Evaluation 1: I Cloned a Talking Version of Myself: The Joys and Concerns of AI Digital Avatars

ModelAccuracyFluencyTerminologyReadabilityTotal Score
claude-sonnet-4.687988
deepseek-v4-pro99899
gpt-o388988

claude-sonnet-4.6

✓ Terminology is handled accurately; for example, 「レッドフラッグ」 directly retains the professional expression and is accompanied by a Chinese-language explanation: "a warning signal in investment due diligence."

✗ Some sentences are too long and carry translationese, e.g., 「見慣れた顔をして、いつもの自分の口調で話しているのに、それは自分ではない」 sounds somewhat stiff.

deepseek-v4-pro

✓ The language is natural and fluent; for example, 「背筋が凍る思いだった」 accurately conveys the chilling feeling of "a chill down one's spine," and the transitions are smooth.

✗ A few terms are slightly literal, e.g., 「レッドフラッグ・シグナル」 could be streamlined into a more common expression.

gpt-o3

✓ The structure is clear, and quoted sections are handled appropriately; e.g., 「ぞっとしました」 naturally conveys "hair-raising."

✗ Some expressions are slightly flat; e.g., 「よどみなく語る様子」 lacks the original's layered tension.

Conclusion: Version B is overall the best, with the best balance of fluency and accuracy; it is recommended for priority use. A and C both have obvious deficiencies in completeness.

Evaluation 2: Tech Giants Move $300 Billion in AI Debt Off Their Books: How Off-Balance-Sheet Guarantee Structures Conceal Systemic Risk

ModelAccuracyFluencyTerminologyReadabilityTotal Score
claude-sonnet-4.676856
deepseek-v4-pro98988
gpt-o389899

claude-sonnet-4.6

✓ Terms such as 「残余価値保証+特別目的事業体(SPV)」 are used fairly accurately and fit the financial context.

✗ The text is clearly truncated; the part about "the blind spot of footnotes" is unfinished, severely affecting readability.

deepseek-v4-pro

✓ Terminology consistency is high: 「残存価値保証」 and 「特別目的会社(SPV)」 are unified throughout, and 「教科書レベルの構造モデル」 is expressed naturally and aptly.

✗ Some long sentences are somewhat cumbersome; for example, in the second paragraph, the description of the Blue Owl structure is a bit repetitive.

gpt-o3

✓ Fluency is the best: the subheading 「帳簿の外にある実質的エクスポージャー」 is translated concisely and naturally, and paragraph transitions are clear.

✗ Some terminology is slightly inconsistent; e.g., 「残存価値保証」 and 「潜在損失引当金」 are mixed, not fully unified.

Conclusion: Version C is overall the best, leading in fluency and readability; Version B has the most rigorous terminology; Version A clearly lags due to truncation issues.

Evaluation 3: Pentagon Spends $30 Million to Build an AI Polygraph

ModelAccuracyFluencyTerminologyReadabilityTotal Score
claude-sonnet-4.689988
deepseek-v4-pro98999
gpt-o398898

claude-sonnet-4.6

✓ Fluency is good: the title 「米国防総省のAI嘘発見器計画」 naturally corresponds to the original structure, and the logical flow is smooth.

✗ The ending is clearly truncated; 「——これは」 is unfinished, affecting readability.

deepseek-v4-pro

✓ Accuracy is high: the title 「国防総省が3000万ドルを投じてAIポリグラフを開発」 directly corresponds to the original meaning, with no additions.

✗ Some expressions are slightly stiff; although 「一つのレッドライン」 preserves the original meaning, it carries a touch of translationese.

gpt-o3

✓ Readability is good: the title 「ペンタゴン、3000万ドルを投じてAI嘘発見器を開発」 is concise and forceful, and paragraph transitions are clear.

✗ Terminology consistency is slightly weak; 「嘘発見器」 and 「polygraph」 are mixed, not fully unified.

Conclusion: The three versions are close in overall quality. Version B is slightly better in accuracy and readability; Version A is somewhat weaker due to truncation; Version C's terminology handling could be further optimized.