3大模型翻译对决:第32周质量评测,deepseek-v4-pro 以 9 分领跑
本周共翻译 423 篇文章,覆盖 3 个AI模型。经抽样盲评,deepseek-v4-pro 综合得分最高(9/10)。报告详细对比各模型在准确性、流畅性、术语一致性方面的表现差异。
Decoding Intelligence, Defining Value.
解码智能,定义价值
本周共翻译 423 篇文章,覆盖 3 个AI模型。经抽样盲评,deepseek-v4-pro 综合得分最高(9/10)。报告详细对比各模型在准确性、流畅性、术语一致性方面的表现差异。
WDCD Run #253 (2026-07-29) tested 11 models across three dialogue rounds, recording an average commitment decay of 4.5%. Grok 4 topped the ranking at 94.8 points, while GPT-5.5 registered the worst instruction decay at -100%.
本周共翻译 381 篇文章,覆盖 3 个AI模型。经抽样盲评,gpt-o3 综合得分最高(8.3/10)。报告详细对比各模型在准确性、流畅性、术语一致性方面的表现差异。
WDCD Run #247 (2026-07-26) evaluated 11 models across three dialogue rounds, recording an average commitment decay of -1.8%. Grok 4 led the field at 94.2 points with a -63% decay figure, indicating strengthened rather than weakened adherence over the session.
WDCD Run #242 (2026-07-22) evaluated 11 models across three-round multi-turn dialogues, recording an average commitment decay of 18.2% between Round 1 and Round 3. Grok 4 and GLM-4.6 held perfect decay resistance, while Gemini 3.1 Pro collapsed at -100%.
本周共翻译 368 篇文章,覆盖 4 个AI模型。经抽样盲评,claude-sonnet-4.6 综合得分最高(8.5/10)。报告详细对比各模型在准确性、流畅性、术语一致性方面的表现差异。
WDCD Run #233 (2026-07-15) evaluated 11 frontier models on multi-turn commitment integrity, recording an average instruction decay of 27.3% between Round 1 and Round 3. GPT-o3 topped the leaderboard with 94 points and zero decay, while Gemini 3.1 Pro suffered a complete 100% collapse.
本周共翻译 361 篇文章,覆盖 3 个AI模型。经抽样盲评,gpt-o3 综合得分最高(9/10)。报告详细对比各模型在准确性、流畅性、术语一致性方面的表现差异。
WDCD Run #227 (2026-07-12) evaluated 11 frontier models on multi-turn commitment integrity, with Grok 4 and DeepSeek V4 Pro tying at 91.4 points and average instruction decay measured at -2.8% between Round 1 and Round 3.
WDCD Run #221 (2026-07-08) measured instruction decay across 11 frontier models over three dialogue rounds, recording an average commitment decay of -36.4% from Round 1 to Round 3. Grok 4 topped the ranking with 95 points.
本周共翻译 318 篇文章,覆盖 4 个AI模型。经抽样盲评,gpt-o3 综合得分最高(9/10)。报告详细对比各模型在准确性、流畅性、术语一致性方面的表现差异。
WDCD Run #211 (2026-07-03) benchmarked 11 models on multi-turn commitment integrity, with Grok 4 taking the top spot at 91.2 points and only -13% decay, while GPT-o3 posted the worst decay rate at -75%.