3 Major Models Translation Showdown: Week 24 Quality Evaluation, passthrough Leads with a Score of 9

This week, <strong>2425</strong> translation tasks were completed by <strong>3</strong> models. <strong>3</strong> samples were selected for multi-model blind comparison, with the overall best being <strong>passthrough</strong> (average score 9/10).

This week, 2425 translation tasks were completed by 3 models. 3 samples were selected for multi-model blind comparison, with the overall best being passthrough (average score 9/10).

This Week's Translation Statistics

ModelLanguageTranslationsAverage TimeAverage Quality Score
deepseek-v4-flashen3411.9sNot Rated
claude-sonnet-4.6ja14738sNot Rated
native-englishen1-Not Rated
deepseek-v4-flashzh115.3sNot Rated
passthroughen22420sNot Rated

Sampling Comparison Evaluation

Evaluation 1: WordPress.com Adds AI Assistant: Edit Text, Adjust Style, Generate Images All in One

ModelAccuracyFluencyTerminologyReadabilityTotal Score
passthrough99989
deepseek-v4-pro67776
gpt-o368787

passthrough

✓ Faithfully presents the core content of the original, for example, "make this section feel more modern or spacious" directly corresponds to the natural language command function, with no extraneous additions.

✗ The text is truncated at the end, for example, the block themes link part is incomplete, affecting readability.

deepseek-v4-pro

✓ The structured list is clear, for example, using a subheading "Analysis of the AI Assistant's Powerful Features" to organize function descriptions.

✗ Adds a lot of content not present in the original, for example, at the beginning "In the field of content creation, the rapid advancement of AI technology is quietly changing how creators work," which is excessive paraphrasing and addition.

gpt-o3

✓ Paragraph transitions are smooth, for example, list items use bold subheadings to distinguish function modules, with clear logic.

✗ Also adds an introductory paragraph not present in the original, for example, "In the field of content creation...", deviating from the news announcement nature of the original title.

Conclusion: Version A is the best overall, with the highest accuracy and fidelity. The other two versions both have significant issues with additions and deviations from the original.

Evaluation 2: Google Adds AI Skills to Chrome to Help You Easily Save Workflows

ModelAccuracyFluencyTerminologyReadabilityTotal Score
passthrough99989
deepseek-v4-pro67787
gpt-o368787

passthrough

✓ Terminology is natural and accurate, for example, "Gemini AI into Chrome" directly corresponds to the professional product name without additional explanation.

✗ The end of the text is clearly truncated, for example, "For instance, Google suggests that if a user often asks Gemini to suggest vegan substitutions when looking at recipe websites, the" results in incomplete content.

deepseek-v4-pro

✓ Adds subheadings such as "Highlights of the 'Skills' Feature," making the paragraph hierarchy clearer.

✗ Added citation content not present in the original, for example, "The launch of this feature marks another important milestone in the deep integration of AI and browsers." — Technology Analyst, which is excessive addition.

gpt-o3

✓ The language is somewhat smoother, for example, "This feature is especially practical for users who frequently use AI for data processing or information retrieval" is naturally expressed.

✗ Also adds a citation paragraph not present in the original, for example, "The launch of this feature marks another important milestone in the deep integration of AI and browsers." — Technology analyst, deviating from the original.

Conclusion: Version A is the most faithful to the original with natural language, while versions B and C both have issues with obvious additions. It is recommended to prioritize Version A.

Evaluation 3: Meta May Hold 10% of AMD Shares: 6-Gigawatt Chip Deal Boosts AI Ambitions

ModelAccuracyFluencyTerminologyReadabilityTotal Score
passthrough99989
deepseek-v4-pro68777
gpt-o368777

passthrough

✓ Accurately and faithfully reproduces the original news facts, for example, "AMD's chief executive Lisa Su said that 'each gigawatt of compute is worth double-digit billions'" directly corresponds to the original without additions.

✗ The final sentence "The terms of the warrant a" is clearly truncated, affecting completeness.

deepseek-v4-pro

✓ Structurally adds subheadings such as "Deal Details: 6-Gigawatt Chip Supply and Equity Exchange," making the logical hierarchy clearer.

✗ Overly adds explanations not present in the original, such as "performance approaching NVIDIA's H100 and Blackwell series," which is excessive paraphrasing.

gpt-o3

✓ The subheading "Deal Details Emerge" makes paragraph transitions more natural.

✗ Also adds content not present in the original, such as "performance approaching NVIDIA's H100 and Blackwell series," deviating from the principle of fidelity.

Conclusion: Version A is the best overall, with the highest accuracy and fluency, suitable as a benchmark translation. Although versions B and C have slightly better readability, they contain significant excessive additions, making them not recommended for strict translation.