This week, 443 translation tasks were completed by 5 models. A sample of 3 articles was evaluated in a multi-model blind comparison, with the overall best being passthrough (average score 9/10).
Translation Statistics for This Week
| Model | Language | Translation Volume | Average Time | Average Quality Score |
|---|---|---|---|---|
| deepseek-v4-flash | en | 85 | 15.2s | Not rated |
| claude-sonnet-4.6 | ja | 219 | 33.1s | Not rated |
| passthrough | en | 129 | 0s | Not rated |
| native-english | en | 5 | - | Not rated |
| deepseek-v4-flash | zh | 5 | 17.1s | Not rated |
Sample Comparison Evaluation
Evaluation 1: Sinclair to Test Whole-Body Rejuvenation Drug at XPrize
| Model | Accuracy | Fluency | Terminology | Readability | Total Score |
|---|---|---|---|---|---|
| claude-sonnet-4.6 | 9 | 8 | 9 | 8 | 8 |
| deepseek-v4-pro | 8 | 9 | 8 | 9 | 9 |
| gpt-o3 | 9 | 8 | 9 | 8 | 8 |
claude-sonnet-4.6
✓ Accurately conveyed Sinclair's explanation of the "information theory" from the original text, and faithfully retained the scientific details by quoting "Previously, we achieved short-term expression of Yamanaka factors in mice."
✗ The ending was obviously truncated: "bioavailability of oral drugs, ta..." was incomplete, affecting overall readability.
deepseek-v4-pro
✓ The language was natural and fluent; for example, the sentence "In mice, short-term expression of Yamanaka factors has been shown to rejuvenate the heart, liver, and brain" flowed smoothly.
✗ Some terminology translations were slightly simplified, e.g., "Epigenetic Reprogramming" was uniformly rendered as "EpigeneticReprogramming" without the original spacing.
gpt-o3
✓ Accurately captured the meaning of the original title, using "Sinclair to Test Whole-Body Rejuvenation Drug at XPrize" as the headline, with a clear structure.
✗ Some sentences were slightly verbose; for example, "While this method is safer and more accessible than traditional gene therapy, it also faces stricter regulatory review" was a bit stiff in logical transition.
Conclusion: The three versions are of similar overall quality, all faithful to the original with accurate terminology. Version B slightly excels in fluency and readability, while Version A is slightly weaker due to truncation.
Evaluation 2: Wrongfully Arrested: America's Oldest Police Facial Recognition Tool Fails
| Model | Accuracy | Fluency | Terminology | Readability | Total Score |
|---|---|---|---|---|---|
| passthrough | 9 | 9 | 8 | 8 | 9 |
| deepseek-v4-pro | 6 | 7 | 7 | 7 | 6 |
| gpt-o3 | 7 | 8 | 8 | 8 | 7 |
passthrough
✓ The greatest strength is faithfully preserving the original case details, such as "Robert Dillon, a 52-year-old commercial crabber from Fort Myers," fully presenting the character's background and arrest process.
✗ The biggest drawback is the inclusion of HTML tags and incomplete paragraphs; for example, the ending "His mug shot stayed online for nearly a year," was cut off, affecting overall coherence.
deepseek-v4-pro
✓ The biggest strength is the addition of subheading structures, such as "Case Details: A Faulty Match," which makes the logical hierarchy clearer.
✗ The biggest drawback is adding the time information "2025" not present in the original, which is an over-addition that could be misleading.
gpt-o3
✓ The greatest strength is the more natural phrasing of quotes; for example, "This tool is not a reliable method of identification" was translated fluently and in context.
✗ The biggest drawback is also adding the time detail "2025" not in the original, along with content truncation.
Conclusion: Version A has the highest overall quality, closest to the original meaning with natural language. Versions B and C both suffer from unfounded additions and truncation, and are not recommended.
Evaluation 3: NVIDIA Deepens AI Collaboration with Hyundai, Embodied Intelligent Robot Commercialization Accelerates
| Model | Accuracy | Fluency | Terminology | Readability | Total Score |
|---|---|---|---|---|---|
| deepseek-v4-flash | 8 | 7 | 9 | 7 | 7 |
| deepseek-v4-pro | 9 | 8 | 9 | 8 | 8 |
| gpt-o3 | 9 | 9 | 9 | 9 | 9 |
deepseek-v4-flash
✓ Terminology such as "NVIDIA's Omniverse and Isaac platforms" is used accurately and professionally.
✗ The ending suddenly truncates with "From an in," making the content incomplete.
deepseek-v4-pro
✓ The title translation "NVIDIA Deepens AI Collaboration with Hyundai" is concise and faithful to the original.
✗ Some long sentences, such as the end of the second paragraph, are slightly stiff with a mild translationese tone.
gpt-o3
✓ The phrasing "bringing embodied intelligence technology into real-world commercial deployment" is more natural and fluent.
✗ There is slight expansion compared to the original, such as the addition of "real-world," but the impact is minor.
Conclusion: Version C is the best overall, with the highest fluency and readability, followed by Version B. Version A is the weakest due to truncation.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接