Four Major Models in Translation Showdown: Week 38 Quality Evaluation, gpt-o3 Leads with 8.3
This week, 357 translation tasks were completed by 4 models. A sample of 3 was selected for a multi-model blind comparison, and gpt-o3 achieved the best overall score (average 8.3/10).