Evidence at a glance
The Interface Changes Before the Control Flow
At DevDay on September 29, 2026, OpenAI announced Decisions API for developers who need a model to answer bounded questions and return results that software can process. On October 6, Simon Willison released `llm-openai-decisions 0.1a0`, bringing the API into his `llm` command-line tool. The API was in limited preview at the time, and the plugin was marked as a pre-release, so this is an early route to integration rather than a production approach with established maturity.
The central change is not that the model has suddenly learned how to make judgments. It is that an application can define the answer space in advance. A conventional approach might ask a model for prose and then extract a conclusion from it, or rely on the model to follow a required output format. Decisions API makes question types and structured answers part of the interface. That reduces free-text parsing and makes it easier to place a model in classification or request-routing flows. It does not remove incorrect judgments. It delivers them to software more directly.
Three Question Types, With Boundaries Set by the Caller
Decisions API offers three question types: `predicate`, `choice`, and `score`. A `predicate` tests whether a condition holds and returns a probability. A `choice` selects from options supplied by the developer and returns option probabilities and confidence. A `score` assigns a rating against predefined levels and returns corresponding probabilities and confidence. The common feature is that the caller defines what counts as a possible answer, and the model responds within that boundary instead of expanding into open-ended conversation.
This design has practical advantages for integration. A classifier can present a fixed set of labels, a router can define the candidate queues, and a scoring flow can use previously agreed rating levels. A score question requires at least two levels. With three levels, the score range is zero to two. Structured fields are easier for code to consume, but the labels and levels still have to be designed by product and engineering teams. The model does not decide whether those categories adequately cover the real business.
Request Design Determines Whether It Is a Workflow
A request can carry shared input along with multiple questions in a defined order, and the API returns answers in that order. OpenAI recommends putting independent questions in the same request. This lets the same text or set of images be checked against several unrelated criteria without rebuilding the input for each one.
A different rule applies when the second question depends on the first answer. Those questions should not be treated as a single set of parallel tasks. OpenAI recommends splitting dependent questions across multiple requests, with the application orchestrating what happens next. This boundary matters: an interface that accepts multiple questions does not automatically plan subsequent steps, and it is not a complete agent. Technical leads still need to decide which layer owns retries, branching, fallback behavior, and authorization for resulting actions.
The Image Example Shows an Input Path, Not Reliability
Willison demonstrated visual input with a photograph of a pelican. He selected the model with `llm -m openai-decisions/gpt-6-luna`, attached the image with `-a`, and asked with `-s` whether the image contained any mammals. The example returned a `predicate` result with probability `0.0`. It shows that an image can provide context for a bounded judgment and that the plugin can fit the API into a familiar command-line workflow. One example, however, cannot establish how the model performs on different images, edge cases, or production data.
The API reference says images must be supplied as data URLs and limits a request to 128 images. The plugin can be installed with `llm install llm-openai-decisions`. Willison says he had GPT-6 Astra read the new API documentation and built the plugin by taking inspiration from his existing `llm-typesafe` plugin. This makes experimentation easier for developers already using `llm`, but the plugin remains a wrapper around the call. The application still has to construct inputs, handle errors, validate results, and determine which judgments are allowed to trigger automated actions.
Pricing and Probability Still Need Independent Checks
The pricing comparison looks straightforward, but the billing basis is not fully clear. In his release post, Willison said Decisions charges for input and not for output, at an OpenAI rate of $0.10 per million input tokens. He also gave Jev's rate as $0.042 per million input tokens. On those figures, OpenAI's input price is about 2.4 times Jev's. Read only as a comparison of the numbers in that post, Decisions is not the cheaper option.
The available materials do not, however, establish that figure as independently verified Decisions-specific pricing. OpenAI's changelog lists a general GPT-6 Luna API output rate of $0.50 per million tokens, while the available official Decisions material does not separately state a dedicated price. The two figures may reflect different billing arrangements. Before integrating the API, teams should confirm how Decisions calls are actually charged, rather than putting the claim that output is free into a cost model without checking it. For high-volume use, input size, images, and request composition can all affect the bill.
Probability and confidence fields deserve the same caution. The materials say that the API returns these fields, but provide no evidence that the probabilities are calibrated or describe error distributions on business data. A team can first evaluate representative text and images for error types, threshold behavior, and cost. That evidence can help determine which results are suitable for automated routing and which require fallback or human review. Since the API was still in limited preview and the plugin was a 0.1a0 p