Four Major Models Go Head-to-Head in Translation: Week 42 Quality Evaluation, gpt-o3 Leads with 8.7
This week, 376 translation tasks were completed by 4 models. Three samples were selected for blind multi-model comparison, with gpt-o3 achieving the best overall score (8.7/10 average).