Evidence at a glance
From Parsing Prose to Receiving a Decision
OpenAI has opened the Decisions API in public beta. It is a hosted endpoint for text, images, or both: an application sends shared input and a set of questions, and receives typed answers that its code can branch on. It targets a familiar pattern—asking a general-purpose model for prose and then parsing that text into a label. The only supported model today is gpt-6-luna, called through POST /v1/decisions.
The change is not that the model has suddenly learned to classify. It is that classification is being offered through a dedicated interface. Developers have often had to manage prompts, output formatting, parsing failures, and retries themselves; Decisions moves some of that work into the request protocol and returns probabilities, choices, or scores directly. OpenAI says it is about ten times faster than the Responses API, which may change integration costs but does not, by itself, establish better judgment.
Three Answer Types, Three Kinds of Branch
A request has three parts: model, input, and questions. Each question has a unique name, a type, and instructions; the shared input can be text or can include an image, and the response maps answers back by name. This structure allows an application to submit one body of evidence and ask several related questions without making a separate call for every label.
The three types represent different operations, not merely different ways to phrase an answer. A predicate tests whether a condition holds and returns a probability from zero to one. A choice selects from caller-supplied options and includes per-option probabilities and a confidence field. A score rates the input against ordered levels and also returns probabilities and confidence. In the guide’s severity example, level probabilities of 0.1, 0.7, and 0.2 produce a weighted score of 1.1. That places the result around the second level; it is not a precisely measured physical quantity.
Typed Outputs Reduce Parsing, Not Model Uncertainty
Decisions has a narrower role than two other interface patterns. Use it for probabilities, selections from a fixed set, or ratings on ordered levels. OpenAI points developers to Structured Outputs when they need to fill a custom JSON schema or generate explanations, and to function calling when the model needs to request a tool call with arguments. Decisions is not a replacement for every form of structured output; it makes a limited set of decision operations explicit in the contract.
That contract can remove a layer of text parsing, but it cannot tell a team whether a probability of 0.8 is reliable enough or what action a score should trigger. The supplied materials include no endpoint accuracy or calibration data and no independent evaluation. OpenAI’s documentation recommends setting thresholds with labeled examples from the application itself. Probability and confidence should therefore be treated as signals to validate, not as guarantees of risk. If a threshold can deny service, block a transaction, or trigger another costly action, the local evaluation must account for the cost of mistakes.
Faster and Cheaper, but with a Narrow Deployment Choice
The speed evidence is currently dominated by provider claims. OpenAI says responses are about ten times faster than the Responses API. DevDay coverage cited a decision near 150 milliseconds, compared with about 1.6 seconds for a regular Luna call. Those figures help explain why teams may want to try the endpoint, but the materials do not specify independently reproducible test conditions, and they do not establish end-to-end latency for a particular application. Network time, input size, and downstream business logic still matter.
The pricing is more concrete: gpt-6-luna input costs $0.10 per million tokens, with no charges for output tokens, cache reads, or cache writes. Prompts above 272K tokens are billed at twice the input rate, and regional processing may add a premium. The model card lists a 1,050,000-token context window but does not disclose the parameter count. The API is hosted by OpenAI, with no open weights or self-hosting option. Minimum SDK versions are listed as Python 3.26.0, JavaScript 7.30.0, Go 3.73.0, Ruby 0.101.0, and Java 4.78.0. The materials also list Zero Data Retention and HIPAA for eligible customers, and data residency in the United States and Europe, including the EEA and Switzerland.
Useful for Leaner Decision Paths, Not for Outsourcing Accountability
For a technical lead, the right pilot is not a wholesale replacement of existing model calls. It is a task with established labels, options, or ordered levels whose output already drives a code branch. For example, a predicate can return a probability for whether a product photo shows a crack, tear, or dent, allowing the application to decide what to do next. Whether that decision should automatically block a product still depends on labeled examples and the cost of false positives and negatives. A pilot can measure end-to-end latency, whether parsing and retries decrease, and whether probabilities remain useful across the inputs that matter.
The constraints are equally concrete: one supported model, public beta, no independent evaluation, and hosted deployment only. Teams that require self-hosting, comparison across several models, or extensive long-context and regional processing should weigh those limits alongside the interface benefits. Decisions turns a model response into an answer software can consume. It does not turn the question of whether to act on that answer into a responsibility the API can take over.