GPT-6 Astra and GPT-6.1 Sol Tie at 84.99: 2026-10-12 Smoke Quick-Test Data Brief

The 2026-10-12 YZ Index Smoke quick test covered 14 models, with GPT-6 Astra and GPT-6.1 Sol tied for first place for the day at 84.99. Smoke is a daily 10-question quick test, well suited for observing short-term signals, and is not equivalent to Full weekly leaderboard conclusions.

This Smoke evaluation covers only two main leaderboard dimensions: Code Execution and Material Constraints. The main leaderboard formula is 0.55 × Code Execution + 0.45 × Material Constraints. Because the daily sample size is small, single-day scores are better used as monitoring signals rather than as long-term conclusions about model capability.

Daily Ranking

RankingModelMainCode ExecutionMaterial ConstraintsIntegrity
#1GPT-6 Astra84.999770.3pass
#2GPT-6.1 Sol84.999770.3pass
#3Doubao Pro84.6691.576.3pass
#4GPT-6 Sol82.7994.868.1pass
#5GPT-6 Luna78.559756pass
#6Claude Sonnet 4.673.947276.3pass
#7DeepSeek V4 Pro72.597273.3pass
#8GPT-5.572.597273.3pass
#9Gemini 3.1 Pro71.67271.1pass
#10GPT-o371.3869.873.3pass
#11Gemini 2.5 Pro71.2169.573.3pass
#12Grok 468.0263.773.3pass
#13Claude Opus 4.758.844773.3pass
#14GLM-4.641.1816.771.1pass

Data Interpretation

Today's top two on the main leaderboard, GPT-6 Astra and GPT-6.1 Sol, both show a structure of Code Execution 97 and Material Constraints 70.3, with both main leaderboard scores at 84.99; Doubao Pro has Code Execution 91.5 and Material Constraints 76.3 for a main leaderboard score of 84.66, showing a combination with relatively stronger Material Constraints. GPT-6 Luna has Code Execution 97 and Material Constraints 56 for a main leaderboard score of 78.55, reflecting a Code Execution-dominated profile. Claude Sonnet 4.6 has Code Execution 72 and Material Constraints 76.3 for a main leaderboard score of 73.94, contrasting with the Material Constraints score of 73.3 in the same band as DeepSeek V4 Pro and GPT-5.5.

Doubao Pro gained 29.9 points on the main leaderboard, with Code Execution +44.5 and Material Constraints +12; GPT-5.5 gained 24.2 points on the main leaderboard, with Code Execution +25 and Material Constraints +23.3; DeepSeek V4 Pro gained 23.1 points on the main leaderboard, with Code Execution +25 and Material Constraints +20.7; GPT-6 Sol gained 16.8 points on the main leaderboard, with Code Execution +47.8 and Material Constraints -21.2, reflecting pronounced daily fluctuations in the pairing of Code Execution and Material Constraints. GPT-6 Astra's Material Constraints plunged -19 points, GPT-6.1 Sol's Material Constraints plunged -19 points, Gemini 2.5 Pro's Material Constraints plunged -16 points, Claude Opus 4.7's main leaderboard score plunged -21 points, and GLM-4.6's main leaderboard score plunged -13.6 points; these anomalous signals may stem from question sampling fluctuation or may be signs of genuine degradation, requiring subsequent runs for verification.

Smoke is a small-sample single-day signal; the above interpretation is based only on that day's data structure and does not constitute a long-term judgment.

Key Changes

  • Doubao Pro: Main leaderboard up 29.9 points, Code Execution +44.5, Material Constraints +12
  • GPT-5.5: Main leaderboard up 24.2 points, Code Execution +25, Material Constraints +23.3
  • DeepSeek V4 Pro: Main leaderboard up 23.1 points, Code Execution +25, Material Constraints +20.7
  • Claude Opus 4.7: Main leaderboard down 21 points, Code Execution -25, Material Constraints -16
  • GPT-6 Sol: Main leaderboard up 16.8 points, Code Execution +47.8, Material Constraints -21.2

Signals to Watch

  • GPT-6 Astra: Material Constraints plunged -19 points
  • GPT-6.1 Sol: Material Constraints plunged -19 points
  • GPT-6 Sol: Material Constraints plunged -21.2 points
  • Gemini 2.5 Pro: Material Constraints plunged -16 points
  • Claude Opus 4.7: Main leaderboard plunged -21 points
  • GLM-4.6: Main leaderboard plunged -13.6 points
  • Qwen3 Max: Incomplete data (missing execution, evidence, judgment, integrity, and communication dimensions; API failure/timeout), has entered automatic rerun, and is not included in this period's ranking

When reading this type of Smoke brief, focus on two questions: first, whether a model has exposed the same type of weakness for multiple consecutive days; second, whether its integrity rating has moved from pass to warn or fail. Large changes in single-day execution or constraint scores may come from question sampling or may be early signals of genuine degradation, requiring follow-up runs to verify.


Data source: YZ Index | Run #373 | View raw data

This article is from Winzheng Index blog, translated in full by Winzheng (winzheng.com). Click here to view the original When republishing the translation, please credit the source. Thank you!