Translation Showdown of 5 Models: Week 25 Quality Evaluation, passthrough Leads with 9 Points

This week, <strong>443</strong> translation tasks were completed by <strong>5</strong> models. A sample of <strong>3</strong> articles was evaluated in a multi-model blind comparison, with the overall best being <strong>passthrough</strong> (average score 9/10).

This week, 443 translation tasks were completed by 5 models. A sample of 3 articles was evaluated in a multi-model blind comparison, with the overall best being passthrough (average score 9/10).

Translation Statistics for This Week

ModelLanguageTranslation VolumeAverage TimeAverage Quality Score
deepseek-v4-flashen8515.2sNot rated
claude-sonnet-4.6ja21933.1sNot rated
passthroughen1290sNot rated
native-englishen5-Not rated
deepseek-v4-flashzh517.1sNot rated

Sample Comparison Evaluation

Evaluation 1: Sinclair to Test Whole-Body Rejuvenation Drug at XPrize

ModelAccuracyFluencyTerminologyReadabilityTotal Score
claude-sonnet-4.698988
deepseek-v4-pro89899
gpt-o398988

claude-sonnet-4.6

✓ Accurately conveyed Sinclair's explanation of the "information theory" from the original text, and faithfully retained the scientific details by quoting "Previously, we achieved short-term expression of Yamanaka factors in mice."

✗ The ending was obviously truncated: "bioavailability of oral drugs, ta..." was incomplete, affecting overall readability.

deepseek-v4-pro

✓ The language was natural and fluent; for example, the sentence "In mice, short-term expression of Yamanaka factors has been shown to rejuvenate the heart, liver, and brain" flowed smoothly.

✗ Some terminology translations were slightly simplified, e.g., "Epigenetic Reprogramming" was uniformly rendered as "EpigeneticReprogramming" without the original spacing.

gpt-o3

✓ Accurately captured the meaning of the original title, using "Sinclair to Test Whole-Body Rejuvenation Drug at XPrize" as the headline, with a clear structure.

✗ Some sentences were slightly verbose; for example, "While this method is safer and more accessible than traditional gene therapy, it also faces stricter regulatory review" was a bit stiff in logical transition.

Conclusion: The three versions are of similar overall quality, all faithful to the original with accurate terminology. Version B slightly excels in fluency and readability, while Version A is slightly weaker due to truncation.

Evaluation 2: Wrongfully Arrested: America's Oldest Police Facial Recognition Tool Fails

ModelAccuracyFluencyTerminologyReadabilityTotal Score
passthrough99889
deepseek-v4-pro67776
gpt-o378887

passthrough

✓ The greatest strength is faithfully preserving the original case details, such as "Robert Dillon, a 52-year-old commercial crabber from Fort Myers," fully presenting the character's background and arrest process.

✗ The biggest drawback is the inclusion of HTML tags and incomplete paragraphs; for example, the ending "His mug shot stayed online for nearly a year," was cut off, affecting overall coherence.

deepseek-v4-pro

✓ The biggest strength is the addition of subheading structures, such as "Case Details: A Faulty Match," which makes the logical hierarchy clearer.

✗ The biggest drawback is adding the time information "2025" not present in the original, which is an over-addition that could be misleading.

gpt-o3

✓ The greatest strength is the more natural phrasing of quotes; for example, "This tool is not a reliable method of identification" was translated fluently and in context.

✗ The biggest drawback is also adding the time detail "2025" not in the original, along with content truncation.

Conclusion: Version A has the highest overall quality, closest to the original meaning with natural language. Versions B and C both suffer from unfounded additions and truncation, and are not recommended.

Evaluation 3: NVIDIA Deepens AI Collaboration with Hyundai, Embodied Intelligent Robot Commercialization Accelerates

ModelAccuracyFluencyTerminologyReadabilityTotal Score
deepseek-v4-flash87977
deepseek-v4-pro98988
gpt-o399999

deepseek-v4-flash

✓ Terminology such as "NVIDIA's Omniverse and Isaac platforms" is used accurately and professionally.

✗ The ending suddenly truncates with "From an in," making the content incomplete.

deepseek-v4-pro

✓ The title translation "NVIDIA Deepens AI Collaboration with Hyundai" is concise and faithful to the original.

✗ Some long sentences, such as the end of the second paragraph, are slightly stiff with a mild translationese tone.

gpt-o3

✓ The phrasing "bringing embodied intelligence technology into real-world commercial deployment" is more natural and fluent.

✗ There is slight expansion compared to the original, such as the addition of "real-world," but the impact is minor.

Conclusion: Version C is the best overall, with the highest fluency and readability, followed by Version B. Version A is the weakest due to truncation.