GPT-6 Luna and GPT-6 Sol Tie at 80.99: 2026-10-06 Smoke Quick Test Data Brief

On 2026-10-06, the YZ Index Smoke quick test covered 15 models, with GPT-6 Luna and GPT-6 Sol tying for first place that day at 80.99. Smoke is a daily 10-question quick test, suitable for observing short-term signals, and is not equivalent to Full weekly leaderboard conclusions.

This Smoke evaluation covers only two main leaderboard dimensions: code execution and material constraints. The main leaderboard formula is 0.55 × code execution + 0.45 × material constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals rather than as long-term conclusions about model capabilities.

Daily Rankings

RankModelMain LeaderboardCode ExecutionMaterial ConstraintsIntegrity
#1GPT-6 Luna80.997588.3pass
#2GPT-6 Sol80.997588.3pass
#3GPT-5.578.027581.7pass
#4GPT-o378.027581.7pass
#5Grok 477.373.781.7warn
#6Doubao Pro71.997568.3pass
#7Claude Sonnet 4.669.495093.3pass
#8Claude Opus 4.767.245088.3pass
#9GPT-6 Astra67.245088.3pass
#10Qwen3 Max65.044688.3pass
#11GPT-6.1 Sol55.995063.3pass
#12DeepSeek V4 Pro51.4241.763.3pass
#13Gemini 2.5 Pro51.4241.763.3pass
#14Gemini 3.1 Pro51.4241.763.3pass
#15GLM-4.69020pass

Data Interpretation

Today's top two on the main leaderboard, GPT-6 Luna and GPT-6 Sol, both scored 80.99, with code execution at 75 and material constraints at 88.3, showing a balanced combination across the two capabilities. GPT-5.5 and GPT-o3 also scored 78.02 with code execution at 75 and material constraints at 81.7, ranking third and fourth. Grok 4 had code execution at 73.7 and material constraints at 81.7, with a main leaderboard score of 77.3, placing fifth. Claude Sonnet 4.6 had code execution at 50 and material constraints at 93.3, with a main leaderboard score of 69.49, reflecting a structural profile in which material constraints are relatively outstanding.

Compared with the previous run under the same methodology, Gemini 2.5 Pro fell 37.2 points on the main leaderboard, 49.6 points in code execution, and 22 points in material constraints; GLM-4.6 fell 32.7 points on the main leaderboard, 44.5 points in code execution, and 18.3 points in material constraints; DeepSeek V4 Pro fell 31.7 points on the main leaderboard, 49.6 points in code execution, and 9.8 points in material constraints; Claude Opus 4.7 fell 28.4 points on the main leaderboard, 44.5 points in code execution, and 8.7 points in material constraints; and Gemini 3.1 Pro fell 25.2 points on the main leaderboard, 27.8 points in code execution, and 22 points in material constraints. These changes may stem from question sampling fluctuations, or may indicate real degradation in model performance, and require follow-up runs for review.

The Smoke quick test is a small-sample single-day signal; the above score structure and movements only reflect that day's data, and no long-term judgment is made for now.

Major Changes

  • Gemini 2.5 Pro: main leaderboard down 37.2 points, code execution -49.6 points, material constraints -22 points
  • GLM-4.6: main leaderboard down 32.7 points, code execution -44.5 points, material constraints -18.3 points
  • DeepSeek V4 Pro: main leaderboard down 31.7 points, code execution -49.6 points, material constraints -9.8 points
  • Claude Opus 4.7: main leaderboard down 28.4 points, code execution -44.5 points, material constraints -8.7 points
  • Gemini 3.1 Pro: main leaderboard down 25.2 points, code execution -27.8 points, material constraints -22 points

Signals to Watch

  • No publishable anomaly signals were retained this time.

When reading this type of Smoke brief, the focus should be on two questions: first, whether a given model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating has moved from pass into warn or fail. Large single-day changes in execution or constraint scores may come from question sampling, or may be early signals of real degradation, requiring follow-up runs for review.


Data source: YZ Index | Run #363 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!