Translation Showdown of 3 Major Models: Week 31 Quality Evaluation, gpt-o3 Leads with 8.3 Points

This week, 381 translation tasks were completed by 3 models. A sample of 3 was selected for blind evaluation, with gpt-o3 scoring the highest average of 8.3/10.

This week 381 translation tasks were completed by 3 models. 3 samples were selected for multi-model blind evaluation, with the overall best being gpt-o3 (average score 8.3/10).

Weekly Translation Statistics

ModelLanguageTranslation VolumeAverage TimeAverage Quality Score
deepseek-v4-flashen6416.2sNot Rated
claude-sonnet-4.6ja19040sNot Rated
passthroughen1180sNot Rated
native-englishen2-Not Rated
deepseek-v4-flashzh210.7sNot Rated
deepseek-v4-flash:backtranslateen524.8sNot Rated

Sampled Comparative Evaluation

Evaluation 1: Gemini 3.1 Pro Material Constraints Plunged 17.8 Points, Main Ranking Dropped 6 Points

ModelAccuracyFluencyTerminologyReadabilityTotal Score
claude-sonnet-4.699999
deepseek-v4-pro76666
gpt-o398988

claude-sonnet-4.6

✓ Complete structure and faithful to the original; "the decline of 17.8 points in the materials constraint dimension is the main reason for the drop in the overall ranking" directly corresponds to the original logic without adding irrelevant content.

✗ Some long sentences are slightly verbose; "since only two questions are extracted from this dimension per day in the Smoke evaluation" could be further condensed to improve readability.

deepseek-v4-pro

✓ The translated title is concise, "Gemini 3.1 Pro's Material Constraints Plunge 17.8 Points" captures the core.

✗ Inconsistent terminology and expressions like "side chart" not present in the original; severe truncation at the end causes loss of information.

gpt-o3

✓ Natural and fluent expression; "the materials constraint dimension dropped 17.8 points this time, becoming the main factor for the decline in the main ranking" has clear coherence.

✗ Some sentences are slightly repetitive; "this dimension's score did not decline, but rather increased" could be further refined.

Conclusion: Version A is the best overall, leading in accuracy, fluency, and structural completeness; Version C is second, suitable for scenarios requiring natural expression; Version B is not recommended due to terminology confusion and severe truncation.

Evaluation 2: Deep Divide in Silicon Valley: Polarized Views on China's AI

ModelAccuracyFluencyTerminologyReadabilityTotal Score
claude-sonnet-4.687877
deepseek-v4-pro99999
gpt-o398888

claude-sonnet-4.6

✓ Terminology such as "multimodal integration" is used accurately, faithfully conveying the original meaning of "多模态融合".

✗ Obvious truncation at the end, "some investors are telling their portfolio companies to" is incomplete, affecting readability.

deepseek-v4-pro

✓ "In effect, an industry giant" expresses naturally, aligning better with the original "实态是已经行业巨头" than Version A.

✗ Some long sentences have a slight translationese feel, e.g., "jostling for position" is somewhat stiff in this context.

gpt-o3

✓ Consistent use of "cost-effectiveness" accurately corresponds to the original "成本效率".

✗ "Technology decoupling" is slightly more verbose compared to Version B, with slightly lower fluency.

Conclusion: The overall quality of the three versions is close; deepseek-v4-pro performs best in fluency and completeness, recommended as the first choice.

Evaluation 3: Original Title

ModelAccuracyFluencyTerminologyReadabilityTotal Score
claude-sonnet-4.698988
deepseek-v4-pro87877
gpt-o399999

claude-sonnet-4.6

✓ Terminology is accurate, such as "Electronic Health Records (EHR)" and "HIPAA" translated consistently.

✗ Some long sentences have slightly complex structures, with a mild translationese, e.g., "formally stepped into" feels a bit unnatural.

deepseek-v4-pro

✓ The description of data processing privacy terms is relatively clear; "it is explicitly stated that health data will not be used for model training" is accurately translated.

✗ Lacks fluency; some sentences are too literal, e.g., "has officially launched an AI conversational assistant in the field of personal health records" is not natural.

gpt-o3

✓ Overall expression is the most natural, e.g., "can become a close partner" and "a bold step" are both faithful and fluent.

✗ In a few places, the interpretation of the original word "stepping in" is slightly free, but does not cause substantial deviation.

Conclusion: Version C (gpt-o3) performs best overall, surpassing the other versions in accuracy, fluency, and readability. Recommended as the first choice.