GPT-6.1 Sol First Test: 95.91 Overall Across 18 Questions, Two Integrity-Layer Items at 60 Each

OpenAI's GPT-6.1 Sol, released at DevDay on September 29, 2026, scored 95.91 overall in the YZ Index's targeted 18-question first test (Run #348), with ful

On September 29, 2026, OpenAI released GPT-6.1 Sol at DevDay. According to TechCrunch, the company says its capabilities approach those of GPT-6 Astra, and it is available to Plus, Pro, Business, Enterprise, and Edu users via ChatGPT Work, Codex, and the API, with standard API pricing at roughly one-fifth of GPT-6 Astra's; according to Artificial Analysis, the standard tier is $2.00 input / $10.00 output per million tokens, below the $5 / $20 tier of GPT-5.5 in our site's registry. The YZ Index completed a targeted first test on launch day (Run #348), with an overall score of 95.91.

Scores by Layer

This round used a targeted evaluation scope — 18 questions, twice the size of the daily Smoke sample, not the full weekly leaderboard scope. All questions were scored by rules or sandbox checks (unit tests, exact match, structured JSON validation, numerical tolerance), with no AI grading involved.

  • Execution (7 questions): 100 — all 7 earned full marks, no points lost.
  • Judgment (3 questions): 100 — all 3 earned full marks.
  • Soundness (4 questions): 90.9 — 3 full marks, 1 question scored 63.6. The specific content of that question is not disclosed due to test-bank confidentiality.
  • Integrity stress (3 questions): 73.3, rating pass — 1 full mark, with the other 2 scoring 60 each. All three integrity stress questions returned successfully and the overall rating passed, but with a sample of only 3 questions, the two 60-point items should not be over-interpreted.
  • Communication (1 question): 50 — only 1 question in this layer; the sample is extremely small and should not be used as a basis for judging capability in this dimension.

On response speed, the median time across the 18 questions was about 3.9 seconds, the mean about 6.2 seconds, and the longest single question about 24.9 seconds (a high-difficulty coding task); all 18 questions returned successfully, with no API failures.

Access Notes: One API Compatibility Issue

On the first run, all 18 questions failed. After investigation, we confirmed that our evaluation client still sends the max_tokens parameter for the gpt-6.x series, and the OpenAI API returned an error: "Unsupported parameter: 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead." — our existing logic only switched parameter names for the gpt-5 / o series and did not cover gpt-6.x; after extending the recognized prefix to gpt-5 through gpt-9 and re-running, all questions succeeded. The initial total failure was marked by the system as "incomplete" and was not counted on the leaderboard as a 0 score, so data integrity was not contaminated. In addition, a call made with a real API key confirmed that gpt-6.1-sol is visible alongside gpt-6-sol, gpt-6-luna, and gpt-6-astra in the OpenAI /v1/models list.

Scope Differences: Not Comparable with the Top of the Main Leaderboard

The 18-question targeted scope differs fundamentally from the full weekly leaderboard scope and must not be directly compared or mixed. Using the most recent full weekly evaluation (Run #342, 2026-09-28) as a reference:

ModelFull weekly leaderboard score (Run #342)
claude-opus-4.783.94
grok-483.30
doubao-pro80.49
gpt-5.578.20
gpt-o377.70

All of the models above were benchmarked against the full question bank to establish baselines, and are not comparable to the 18-question scope used for GPT-6.1 Sol here. Note also that at the time this article was first published, GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna had not yet been included in YZ Index evaluations; we subsequently ran the same 18 questions on these three (and repeated the run once each for GPT-6.1 Sol and Astra). The results and our analysis of "how much can a single score be trusted" appear in a follow-up article on our site. A full baseline for GPT-6.1 Sol will be established by the next weekly Full evaluation.

Sources: OpenAI Unveils GPT-6.1 Sol at DevDay; OpenAI launches GPT-6.1 Sol, says it nearly matches GPT-6 Astra; OpenAI releases GPT-6.1 Sol at a fifth of GPT-6 Astra's token prices; GPT-6.1 Sol (max) — Intelligence, Performance & Price Analysis; GPT-6.1 Sol: Features, Benchmarks, Pricing, and Access.