GPT-6.1 Sol Integration Log: max_tokens Causes Total Failure on First 18-Question Run

YZ Index integrated GPT-6.1 Sol within hours of release; its first 18-question run failed completely because the model requires max_completion_tokens inste

On 2026-09-30, YZ Index completed integration with GPT-6.1 Sol within hours of its release, triggering Run #348 (an 18-question targeted evaluation covering five layers: Execution, Rigor, Judgment, Integrity Under Pressure, and Communication—twice the question volume of the daily Smoke). In the full run after fixing the integration issue, all 18 questions returned successfully with no API failures. The median response time per question was 3.9 seconds, the average was 6.2 seconds, and the longest was 24.9 seconds. The overall score was 95.91 (under the 18-question targeted scope, not the full weekly leaderboard scope; the full baseline will be produced in the next Full evaluation). But before that number, we experienced a clean, total failure.

What Happened: 5 Layers, 18 Questions, All API Calls Returned 400

On the first run, none of the API calls in any evaluation layer reached the inference stage; all failed with HTTP 400, with identical returned messages:

"Unsupported parameter: 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead."

Not a single question received a model response.

Root Cause and Fix: The Inherent Blind Spot of Prefix-Enumeration Whitelists

Our evaluation client uses the model ID prefix to decide which token limit parameter to use: the gpt-5 series and o series have switched to max_completion_tokens, while other models retain the old max_tokens field. The problem is that gpt-6.x matches neither the gpt-5 whitelist nor the o series, so it took the old path and triggered a 400 directly.

The fix is straightforward: expand the prefix matching range from "gpt-5.x + o series" to "gpt-5.x through gpt-9.x + o series," covering current and future generations of GPT naming. After the rerun, all 18 questions returned normally.

A more fundamental recommendation is to abandon the whole idea of enumerating whitelists by version number—this is a "catch-up" design in which every new model generation requires manual code changes, and omissions are almost inevitable. If you are unsure of the boundaries, a more robust approach is to use max_completion_tokens uniformly for Chat Completions calls instead of maintaining a model whitelist; but validate it on every model you actually use before deploying. Multiple open-source integration projects have already documented exactly the same 400 error and fix path when integrating the gpt-6 series on public issue trackers.

Failures Are Recorded as "Missing," Not 0 Points, and Are Not Written to the Public Leaderboard

Our evaluation system distinguishes two types of failure: if a model participates in the evaluation but answers incorrectly, it is recorded with the corresponding score; if all API calls fail and the data is unavailable, it is recorded as incomplete for that layer, excluded from any aggregate calculation, and not written to the public leaderboard. The first run of Run #348 triggered a total API failure due to a parameter error; the system automatically marked it as missing, and only the rerun data after the fix was published as valid data. The rationale for this design is that a model that failed because of an integration error should not leave a false low-score record on the leaderboard and thereby affect cross-comparison.

Official Integration Specifications

The following specifications come from OpenAI's official developer documentation (developers.openai.com/api/docs/models/gpt-6.1-sol) and match reports from The Rundown AI and DataCamp:

Model IDgpt-6.1-sol
Context window1,050,000 tokens
Max output128,000 tokens
Input modalitiesText, image
Token limit parametermax_completion_tokens (max_tokens returns 400)
Reasoning effortreasoning_effort: low / medium (default) / high / xhigh / max; none and minimal are not supported
Temperature parameterAccepts only the default value 1 (our test: passing temperature=0.2 returns 400, "Only the default (1) value is supported"; the official model page does not state this)
Tool callingMust use the Responses API; the Chat Completions interface does not support tool calling
Knowledge cutoff2026-04-30

According to The Rundown AI and DataCamp (both give the same figures), standard usage (input ≤ 272,000 tokens) is priced at $2.00 input and $10.00 output per million tokens; when a single request's input exceeds 272,000 tokens, the input rate for the entire request rises to $4.00 and the output rate to $15.00—note that the entire request is billed at the higher rate, not just the excess portion. Cost estimation for very long context scenarios needs separate handling.

Availability in Codex and ChatGPT Work

According to DataCamp, GPT-6.1 Sol is available to Plus, Pro, Business, Enterprise, and Edu accounts, and can be accessed through the ChatGPT desktop app, Codex CLI, and IDE extensions (VS Code, JetBrains, etc.); Free and Go plans do not include the model. OpenAI's official announcement says the model is also available in ChatGPT Work and Codex. Using the same API key, we tested the /v1/models endpoint and found gpt-6.1-sol, gpt-6-sol, gpt-6-luna, and gpt-6-astra listed side by side; we later also ran supplementary tests on the latter three with the same 18-question set (see our follow-up article). Confirm the ID spelling before use.

Migration Checklist

  • Replace the token limit parameter: Replace max_tokens with max_completion_tokens; if unsure of model coverage, consider using the latter uniformly and validate it on every model in use.
  • Broaden prefix matching: If using conditional branches to distinguish model families, expand the matching condition from "gpt-5.x + o series" to "gpt-5.x through gpt-9.x + o series," covering future sub-versions and avoiding repeated changes for each generation.
  • Do not pass custom temperature: Our test found that this model accepts only the default value 1; passing 0.2 returns 400. Reusing low-temperature settings from old code will fail directly.
  • Tool calling migration: function calling and built-in tools (web search, code interpreter, etc.) must be called through the Responses API; code paths that rely on Chat Completions tool calling will fail outright.
  • Audit the reasoning effort parameter: reasoning_effort accepts only low / medium / high / xhigh / max; passing none or minimal in code will error and must be removed or replaced with the lowest valid value, low.
  • Long-context rate breakpoint: 272,000 tokens is the rate switching point; beyond it, the entire request is billed at the higher rate, so add segmented handling to cost estimation logic in advance.
  • Classify API failures: Distinguish parameter errors (400, which should immediately abort and alert) from capability boundary failures (which may be downgraded depending on the business); avoid silently accumulating API failures as 0 points, which would contaminate downstream scoring or analytics data.
  • Check model IDs: During release, model IDs may have aliases or snapshot changes; check in real time via the /v1/models endpoint rather than relying on strings in documentation screenshots.

References: GPT-6.1 Sol Model — OpenAI Developer Docs; GPT-6.1 Sol: Features, Pricing, Context Window & Alternatives; GPT-6.1 Sol: Features, Benchmarks, Pricing, and Access — DataCamp; Graker does not support GPT-6 models due to deprecated max_tokens — GitHub; OpenAI 400 BadRequestError when using newer models — GitHub/honcho.