Four Major Models in Translation Showdown: Week 38 Quality Evaluation, gpt-o3 Leads with 8.3

This week, 357 translation tasks were completed by 4 models. A sample of 3 was selected for a multi-model blind comparison, and gpt-o3 achieved the best overall score (average 8.3/10).

This week, 357 translation tasks were completed by 4 models. A sample of 3 was selected for a multi-model blind comparison; the best overall: gpt-o3 (average score 8.3/10).

This Week's Translation Statistics

ModelLanguageTranslation VolumeAverage TimeAverage Quality Score
deepseek-v4-flashen7144sNot scored
claude-sonnet-4.6ja17849.3sNot scored
passthroughen1050sNot scored
native-englishen1-Not scored
deepseek-v4-flashzh168.2sNot scored
deepseek-v4-flash:backtranslateen113.1sNot scored

Sampled Comparative Evaluation

Evaluation 1: Original Title

ModelAccuracyFluencyTerminologyReadabilityTotal Score
claude-sonnet-4.698898
deepseek-v4-pro87877
gpt-o399999

claude-sonnet-4.6

✓ Paragraph structure is clear and logical connections are natural; for example, 「Smoke評価は1日あたりわずか10問、各次元2問構成である」 accurately conveys the small-sample nature.

✗ The ending is truncated, resulting in incomplete content; for example, the final sentence 「誠実性評価の」 is clearly unfinished.

deepseek-v4-pro

✓ The title translation stays close to the original; for example, 「GPT-o3のコード実行が24.7ポイント急落」 directly corresponds to the core information.

✗ The content is presented in JSON array form, making reading coherence poor; for example, multiple short sentences repeating 「下落幅は」 feel stiff.

gpt-o3

✓ Terminology is consistent and expression is natural; for example, 「問題抽選による変動の可能性がより高い」 accurately and fluently conveys the causal analysis.

✗ Some long sentences are slightly verbose; for example, the explanatory paragraph on rising material constraints could be further condensed.

Conclusion: The three versions are close in overall quality. The gpt-o3 version is slightly better in fluency, terminology consistency, and readability; the claude version has clear structure but suffers from truncation; the deepseek version's JSON format affects reading experience. It is recommended to prioritize the gpt-o3 version.

Evaluation 2: Alibaba Leads $300 Million Investment in UniPat AI at $2.5 Billion Valuation; AI Evaluation Sector Sees First Major Capital Bet from a Tech Giant

ModelAccuracyFluencyTerminologyReadabilityTotal Score
deepseek-v4-flash78677
deepseek-v4-pro88878
gpt-o387888

deepseek-v4-flash

✓ Sentence transitions are fairly natural; 「with Tencent and Hongtai participating as follow-on investors」 clearly expresses the follow-on investment relationship.

✗ Terminology is inconsistent; 「Qwen lab」 and later 「Alibaba's Qwen lab」 are mixed, which can cause confusion.

deepseek-v4-pro

✓ Terminology is unified; 「Tongyi Lab」 is used consistently throughout.

✗ Content is truncated; 「The overall deal structure shows that the AI evaluation segment is gradually becoming a」 is not fully presented.

gpt-o3

✓ Paragraph logic is clear; 「Alibaba’s decision to lead the investment shows that...」 expresses the causal relationship well.

✗ Some expressions are slightly stiff; the translation of the 「Fact Reconstruction」 subheading is a bit rigid.

Conclusion: The three versions are close in overall quality. B and C are better than A in terminology consistency, and C is slightly better in readability, but B has obvious truncation. It is recommended to prioritize C or a repaired B.

Evaluation 3: Cognition Releases SWE-2, Using the Kimi K3 Base to Achieve a New Balance in Code Model Cost-Effectiveness

ModelAccuracyFluencyTerminologyReadabilityTotal Score
claude-sonnet-4.698988
deepseek-v4-pro99999
gpt-o388888

claude-sonnet-4.6

✓ Technical expressions such as 'linear cost penalty mechanism' are translated accurately; 「ペナルティ値はベースモデルのパレートフロンティアの局所的な傾きに基づいて調整される」 faithfully reproduces the original mechanism description.

✗ Some paragraph transitions are slightly stiff; for example, 「学習データ量は3倍に拡大され」 lacks a natural transition from the preceding text and has a slight translationese tone.

deepseek-v4-pro

✓ The title translation is natural; 「コストパフォーマンスに新たな均衡を実現」 both preserves the original meaning and conforms to Japanese expression conventions; body terminology such as 「Paretoフロンティア」 is used consistently and fluently.

✗ 「反復的バリデータ・フライホイール」 feels somewhat coined; the original 'iterative validation flywheel' emphasizes iterative validation more, and there is a slight trace of over-literal translation here.

gpt-o3

✓ 「費用対効果」 fits the Japanese business context, and 「コスト・性能のトレードオフ空間を変化させた」 is clearly expressed.

✗ The end of the content is clearly truncated; text after 「た」 is missing, which is an omission; some terms such as 「rolloutサービス」 are not adapted into Japanese, making them slightly jarring.

Conclusion: Version B is the best overall, with the best balance of accuracy, fluency, and readability; Versions A and C are next, with C losing more points due to truncation.