On October 6, 2026, OpenAI officially announced in its developer community that the Decisions API is open for public beta to all developers. The API is powered by GPT-6 Luna and runs through the dedicated endpoint POST /v1/decisions. It supports text and images as input and returns answers in three predefined formats—predicate (a probability value between 0 and 1), choice (one selected from a fixed list), and score (the probability-weighted mean of ordered levels). According to an announcement from OpenAI's official developer account, its decision speed is up to 10 times faster than GPT-6 Luna completing comparable tasks through the Responses API.
The Decisions API first appeared on September 29, 2026, at OpenAI DevDay in the form of a limited-invitation preview, launching alongside products such as Dots continuously running agents, the GPT-6.1 Sol reasoning model, and Codex Security Cloud. According to eesel.ai, the API appeared on stage for less than a minute during the keynote. One week later, it moved from invitation-only to public beta open to all developers.
Why Separate "Decisions" Out on Their Own
To understand the value of this API, one must first see a long-overlooked structural contradiction in large language model architecture. Large models are designed to generate text word by word, but inside modern AI agent systems, the vast majority of model calls do not need to generate any content at all—they are essentially classification or routing judgments: "Is this customer service message a billing issue or a technical failure?" "Does this product photo show signs of damage?" "Which department should this request be routed to?"
The current standard practice is to feed these questions into a full language model; the model first forms a reasoning chain internally, then generates a passage of text as the answer. This is like asking an expert to write a memo just to get a yes or no. The Decisions API cuts off this unnecessary generation step. It accepts input and directly returns a typed answer; the model does not need to produce an output sequence token by token. This is exactly the source of the official claim of a 10x speedup—removing the generation step, rather than simply increasing compute.
The three output formats each have their use cases: predicate is used to determine whether a condition holds and returns a probability from 0 to 1, for example checking whether damage exists in an image; choice is used to select one from unordered categories, for example content classification or department routing; score is suited to ordered levels, for example severity rating, and returns the weighted mean of the probabilities of each level, which can fall between two levels. Multiple questions can be asked simultaneously in a single request, and the API returns a unified answers array. According to OpenAI developer documentation, each question must be given a unique name so it can be matched in the response.
Pricing Structure Reveals Strategic Intent
According to an announcement in the OpenAI developer community, pricing for the Decisions API charges only for input, at $0.10 per million tokens, with no output token or cache fees. This pricing model is relatively rare in OpenAI's product line, and its signaling significance is no less than the feature itself.
Compared with the billing structure of the Responses API—which usually counts both input and output tokens—the cost profile of the Decisions API is closer to that of the Embeddings API: high-frequency calls, extremely low per-call cost, and suitability for large-scale infrastructure deployment. For a system handling tens of thousands of customer service tickets per day, if it uses the full Responses API for routing classification, every ticket must bear generation-side token costs; after switching to the Decisions API, the cost structure becomes pure input computation, and the speedup means it can handle higher concurrency.
From a stakeholder perspective, the gains and losses for each party are as follows: for individual developers and small teams, the low input-only pricing lowers the barrier to introducing a classification layer into agent architectures; for enterprise users, predictable input-only billing makes cost budgeting easier, and classification accuracy becomes the only key variable; for the agent framework ecosystem (LangChain, Crew AI, etc.), an endpoint specifically optimized for routing and classification may become the default choice for standard routing layers within frameworks, and existing code that uses the Responses API for classification faces migration pressure; for competitors, this creates some pressure in terms of product form—currently Anthropic's Claude API and Google's Gemini API have no equivalent dedicated endpoint for this specific scenario, and general capabilities are not necessarily equivalent to dedicated efficiency.
The Historical Precedent of the Embeddings API
This is not the first time OpenAI has separated a specific cognitive function into a dedicated endpoint. When the Embeddings API became independent from the Completions API in 2022, it was likewise aimed at high-frequency operational scenarios that did not require text generation but only semantic understanding. The independence of the Embeddings API ultimately spawned a large number of vector retrieval applications—not because it was "smarter," but because it provided a more suitable speed and cost structure for a specific task. The Decisions API follows the same logic: isolate a specific cognitive operation, optimize it separately, and price it separately.
The difference is that embedding vectors are lossless intermediate representations, whereas classification judgments must bear accuracy risk. When the Embeddings API was launched, vector retrieval quality had clear metrics such as cosine similarity; the Decisions API currently has no officially published benchmark for classification accuracy, and during the public beta this is the greatest engineering uncertainty—it determines whether this endpoint is suitable only for coarse-grained routing or can also reliably support fine-grained multi-class judgments.
What to Watch Next
OpenAI's official announcement says it expects to move the Decisions API to general availability (GA) "within the coming weeks," and pricing adjustments at GA will be the first key signal. If the current low-cost input-only structure is maintained, it shows that OpenAI is indeed betting on a high-frequency infrastructure positioning and is willing to trade low margins for ecosystem penetration; if output billing is introduced or prices rise sharply, it shows that the pricing during the public beta is a customer acquisition strategy rather than a long-term positioning.
The second thing to watch is third-party accuracy testing. The narrative of a 10x speed advantage depends on one premise: classification quality is high enough that the application layer does not need to perform secondary verification. If testing shows that the Decisions API's accuracy on fine-grained multi-class classification is significantly lower than calling the full Responses API for the same tasks, then the speed advantage will be partially offset by error-correction costs at the application layer.
The actual performance of image input capabilities. Text routing scenarios already have substantial engineering practice, while image input means the Decisions API can directly process user-uploaded screenshots, product photos, and scanned documents without first going through a separate visual parsing step. If this path is stably usable in production environments, it will bring architectural possibilities different from pure text routing to scenarios such as customer service, quality inspection, and content moderation.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接