This week, 393 translation tasks were completed by 4 models. A blind comparison of 3 sampled tasks across multiple models was conducted. Top overall: claude-sonnet-4.6 (average score 9/10).
Weekly Translation Statistics
| Model | Language | Translation Volume | Average Time | Average Quality Score |
|---|---|---|---|---|
| deepseek-v4-flash | en | 72 | 14s | Not rated |
| claude-sonnet-4.6 | ja | 196 | 33s | Not rated |
| passthrough | en | 123 | 0s | Not rated |
| native-english | en | 1 | - | Not rated |
| deepseek-v4-flash | zh | 1 | 10.7s | Not rated |
Sampled Comparative Evaluation
Evaluation 1: Hands-on with Siri AI: The New Evolution of Conversational Intelligent Assistants
| Model | Accuracy | Fluency | Terminology | Readability | Total Score |
|---|---|---|---|---|---|
| passthrough | 2 | 6 | 7 | 3 | 3 |
| deepseek-v4-pro | 8 | 8 | 8 | 8 | 8 |
| gpt-o3 | 9 | 9 | 9 | 9 | 9 |
passthrough
✓ Preserved some original English links and structure
✗ Severely incomplete, interspersed with large amounts of HTML code and untranslated content, e.g., truncated directly at "Since <a href="
deepseek-v4-pro
✓ Idiomatic translation of "From 'Ask and Answer' to 'Read Your Mind'" is natural and fluent
✗ JSON format traces appear at the beginning, breaking overall coherence
gpt-o3
✓ Paragraph transitions and citation handling are the clearest, e.g., "Siri AI is no longer just a tool for executing commands" is translated accurately and naturally
✗ A very few long sentences are slightly formal
Conclusion: gpt-o3 version is the best overall, deepseek-v4-pro is second, the passthrough version has no reference value
Evaluation 2: Claude Fable 5 and Mythos 5 Removed Globally on June 12 — Security Verification Requirements and Privacy Controversy Coexist
| Model | Accuracy | Fluency | Terminology | Readability | Total Score |
|---|---|---|---|---|---|
| claude-sonnet-4.6 | 9 | 8 | 9 | 9 | 9 |
| deepseek-v4-pro | 8 | 7 | 8 | 7 | 7 |
| gpt-o3 | 9 | 9 | 9 | 9 | 9 |
claude-sonnet-4.6
✓ Full text structure is complete, paragraph transitions are natural. "官方アナウンスによると" directly corresponds to the original official statement, logic is clear.
✗ Some long sentences are slightly verbose, e.g., "この要件が一部地域でのユーザー離れを直接引き起こした" could be more concise.
deepseek-v4-pro
✓ Terminology such as "脱獄プロンプト" is used consistently, and "販売中止" in a business context closely matches the original meaning of "下架."
✗ JSON format wraps the body text, and the ending is clearly truncated. "クリエイティブおよびプロトタイピングのシナリオでは" is incomplete, affecting readability.
gpt-o3
✓ Most natural language, "段階的リリース" and "連携経路を直接中断させた" are idiomatic and faithful, with high consistency of "提供停止" throughout the text.
✗ Also has an ending truncation issue, "クリエイティブ用途やプロトタイプ開発のシナリオで" is unfinished.
Conclusion: Versions A and C are close in overall quality and better than Version B. Version C is recommended for priority use due to slightly superior fluency and terminology consistency. Version B is not recommended due to formatting issues and truncation.
Evaluation 3: Leaked Financial Reports Show OpenAI Loses Billions Annually
| Model | Accuracy | Fluency | Terminology | Readability | Total Score |
|---|---|---|---|---|---|
| passthrough | 9 | 10 | 9 | 9 | 9 |
| deepseek-v4-pro | 7 | 8 | 7 | 8 | 7 |
| gpt-o3 | 8 | 9 | 8 | 9 | 8 |
passthrough
✓ Directly uses the original English report, language is natural and fluent. For example, "newly leaked financial documents show a company with quickly growing revenues" accurately conveys the contrast between growth and losses.
✗ Content is severely truncated, lacking complete paragraphs and subsequent analysis, resulting in incomplete information.
deepseek-v4-pro
✓ Clear structure, with added subheadings and citations, e.g., "Losses exceed expectations" makes for easy reading.
✗ Numerical deviations and excessive paraphrase occur. For example, the original text does not mention "40% revenue increase"—content has been added.
gpt-o3
✓ Consistent terminology and natural citation translation, e.g., "You either choose scale or you are out" preserves the original meaning well.
✗ Some expressions are slightly stiff, and the description of R&D spending share deviates slightly from the original details.
Conclusion: Version A is closest to the original report but incomplete. Version C has the best overall balance. Version B contains obvious numerical mistranslations and added content. It is recommended to use Version C or the complete version of A.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接