Interpreting GPT-6.1 Sol's First Test Scorecard: How to Read the 18 Questions Correctly

OpenAI's GPT-6.1 Sol scored 95.91 on the YZ Index's first targeted evaluation (Run #348, an 18-question protocol), but that small-sample, different-protoco

On September 29, 2026 (US time), OpenAI released GPT-6.1 Sol at DevDay. According to TechCrunch and DataCamp, the model is available to Plus/Pro/Business/Enterprise/Edu users in ChatGPT Work, Codex, and the API; according to pricing data from Artificial Analysis and DataCamp, API pricing is $2 per million input tokens and $10 per million output tokens; OpenAI officially says its capability approaches GPT-6 Astra, and according to TechCrunch its price is roughly one-fifth of the latter's. The YZ Index completed its first targeted evaluation on September 30 Beijing time (Run #348, 18-question protocol), with a composite score of 95.91.

Before presenting that number, one thing needs to be made clear: these 18 questions are twice the volume of the routine Smoke set and belong to a targeted, encrypted evaluation protocol — not the full weekly leaderboard protocol. The 18-question composite score and the main leaderboard reference (Run #342, full protocol) are produced from different denominators, so the two cannot be directly ranked or compared across protocols. The purpose of this article is to teach you how to read this scorecard correctly, not to assign GPT-6.1 Sol a ranking.

What a Perfect Score Means, and What It Doesn't

All 7 questions in the Execution layer scored full marks (layer score 100), and all 3 questions in the Judgment layer scored full marks (layer score 100). The value of these two numbers hinges on how grading was done: all 18 questions in this run were graded by rules or sandbox, including passing unit tests, exact string matching, structured JSON field validation, and numeric tolerance comparison, with no AI reviewer involved. That means these scores are very "hard" — not some LLM judge deciding an answer was good, but code passing test cases and output hitting the expected format exactly.

But a perfect score has clear boundaries: 18 questions are a sample, not full coverage of that capability domain. The fact that the 7 execution questions drawn this time all happened to pass does not mean "execution has no blind spots." The smaller the sample, the more a "clean sweep" depends on sampling luck. Statistically, a perfect score from a small sample carries a confidence interval far wider than intuition suggests.

Where Points Were Lost: A Layer-by-Layer Portrait of Four Non-Perfect Scores

Four questions in this evaluation did not receive full marks, spread across three layers:

  • Solidity layer: 1 of 4 questions scored 63.6, for a layer score of 90.9. The question bank is confidential, and we do not disclose what this question tested, so we cannot and should not infer the reason for the lost points from it.
  • Integrity Pressure layer: 2 of 3 questions scored 60 each, for a layer score of 73.3 and an integrity rating of pass. This layer earned partial credit under rule-based grading, indicating that not all scoring points were secured on these two questions; the overall rating remains pass. The specific reasons are likewise not disclosed, and no attribution is made.
  • Communication layer: only 1 question, scoring 50. With just one question, the sample is too small to interpret.

It must be emphasized: no attribution is made here for "why points were lost" — attribution would require disclosing the question content, and the question bank is not public. All that can be established is the score distribution: the Execution and Judgment layers were entirely full marks, while the Solidity, Integrity Pressure, and Communication layers showed non-perfect scores.

Honest About Sample Size: The Confidence Limits of n=1 and n=3

The Communication layer's single question scored 50. With n=1, that number has no statistical meaning — 50 could be a true reflection of the model's communication ability, or it could be that this particular question was especially unfavorable to it, and there is no way to tell the difference. Reading a 50 from n=1 as "middling communication ability" is a misreading; we need to wait for the Full evaluation to accumulate more questions.

The Integrity Pressure layer has n=3 and 73.3 points, a slightly stronger directional signal, but the confidence interval is still very wide. Swap those 3 questions for another 3 and the score could be higher or lower. This is not to say the data is useless — 73.3 points plus a pass rating is indeed a meaningful directional signal; it is to say that using the results of 3 integrity questions to characterize GPT-6.1 Sol's integrity performance means any strong conclusion comes too early, before the complete Full evaluation is out.

Latency Distribution: A 3.9-Second Median and a 24.9-Second Peak

Response times across the 18 questions: median about 3.9 seconds, mean about 6.2 seconds, with the longest single question at about 24.9 seconds (a high-difficulty coding question). The mean is above the median, a right-skewed distribution: most questions finished within a few seconds, while individual high-difficulty coding questions took markedly longer. We did not record reasoning token usage, so we cannot determine the cause of the long-tail latency and only state the phenomenon; also, these are observations under concurrent evaluation conditions, not equivalent to production-environment latency.

A 3.9-second median indicates that GPT-6.1 Sol's response speed on ordinary questions is broadly in line with mainstream API calling experiences — it is not a model dedicated to "slow thinking." The 24.9-second peak indicates the model can adapt its reasoning depth to difficulty — a typical behavioral trait of reasoning models: fast day to day, deep thinking when hitting hard problems, rather than being locked into a single slow speed. The implication for real deployments: for medium-difficulty tasks, a 3.9-second median latency is acceptable; for scenarios dense with highly complex coding tasks, sufficient headroom must be reserved for P99 latency.

Next Week's Full Evaluation: Four Testable Observation Points

GPT-6.1 Sol has entered the YZ Index evaluation registry and will take part in the next complete weekly Full evaluation, with the Full results serving as the complete baseline. Four points will be worth watching then:

  • Whether the Integrity Pressure layer can hold steady at the pass level. This run had n=3 and 73.3 points; the Full set has more questions, so we will see whether it maintains a pass or drops further.
  • Whether the Communication layer's 50 is an occasional result or a norm. With more communication questions in the Full run, the randomness of any single question is diluted, and only then can a meaningful layer assessment be given.
  • The stability of the 95.91 composite score under a larger sample. Mean reversion for small-sample high scores in a Full run is not uncommon — this is a statistical regularity of all LLM evaluations, not a special prediction about GPT-6.1 Sol.
  • Its relative position against existing models on the main leaderboard. In the Run #342 main leaderboard reference, claude-opus-4.7 leads with 83.94, but that is a score under the full protocol. Until a same-protocol Full result for GPT-6.1 Sol is out, the two cannot be compared directly.

Sources: OpenAI launches GPT-6.1 Sol, says it nearly matches GPT-6 Astra and costs less; GPT-6.1 Sol (max) - Intelligence, Performance & Price Analysis; GPT-6.1 Sol: Features, Benchmarks, Pricing, and Access; On Robustness and Reliability of Benchmark-Based Evaluation of LLMs.