


Evidence at a glance
OpenAI Is Offering a Production Decision Model, Not a Model List
OpenAI published “A model guide for the GPT-6 family” on October 2, 2026 for teams using GPT-6 to prototype ideas, build features, work across code repositories, access databases, and coordinate external APIs. The subject is not one conversational model. It is a family consisting of GPT-6 Astra, GPT-6.1 Sol, and GPT-6 Luna, together with the way teams configure reasoning effort, response speed, tools, and context around them. The practical question is specific: how should different tasks inside the same application balance capability, time, and cost in a repeatable way?
The guide deserves attention from engineering leaders not because it repeats a claim about model strength, but because it makes model selection only one variable in a larger system design. In production, a failed answer can trigger a retry, human review, or a rollback of the surrounding workflow, so the price of one call is not the full cost. OpenAI therefore places task success, latency, and cost per successful task in the same evaluation frame. Model selection is becoming a runtime decision rather than a procurement decision.
The Routing Unit Shifts from the Application to the Task Step
The three models in the guide correspond to three broad workload patterns. GPT-6 Astra is intended for the hardest reasoning tasks and serves as the choice when maximum capability is required. GPT-6.1 Sol targets complex coding, research, and computer use. GPT-6 Luna is designed for focused, repeated work at scale, such as extracting invoice fields, classifying requests, and producing structured summaries. This division does not mean that an application must be bound to one model. It means that the system should route work after the task has been decomposed.
Reasoning effort adds another layer of routing. Low effort suits fact extraction and small edits, medium effort supports feature planning and option comparison, and high effort is intended for difficult debugging, deeper analysis, and careful review. Extra high or Max should be tested only when High is insufficient and retained only when the improvement is worth the added time and cost. The API can also change reasoning effort during a conversation without breaking the cache. A workflow can therefore start with a smaller budget and reserve expensive reasoning for the step that actually needs it.
Speed, Reasoning, and Price Are Separate Axes
This is where deployment teams can easily misread the guide. Higher reasoning effort does not automatically produce a faster user experience, and lowering reasoning effort is not necessarily the fastest or cheapest option. The guide treats response speed as a separate configuration dimension. The API can use Fast mode, and in some cases the faster Ultrafast mode, but faster modes carry a higher per-token price, while Ultrafast is currently available only for GPT-6 Astra. Speed optimization and intelligence budgeting must be measured separately.
Engineering leaders should therefore ask neither only which model is strongest nor which mode has the lowest response time. They should ask how much a complete task costs, how long it takes, and whether it succeeds on the first attempt. A cheaper model that fails often and triggers retries may cost more than one successful call from a more capable model. Conversely, a high-reasoning configuration that improves only rare edge cases may be a poor default if it slows every request. Useful routing combines task difficulty, success rate, latency, and cost.
Caching and Compaction Change the Cost Equation
For long workflows, cost optimization is no longer just about reducing output tokens. The guide recommends placing stable instructions and reference material before changing task details, keeping tool definitions consistent, and reusing shared context through prompt caching. Cached input tokens can cost up to 95 percent less than uncached input tokens, depending on the model, but that is not the actual discount for an entire workflow. Cache writes, long-context rates, and cache misses must be included in the estimate, and cache diagnostics are needed to explain why reuse breaks down.
Compaction addresses a different problem. In longer conversations or agent runs, it reduces the size of the context while preserving the state needed to continue. It is not simply deleting the history. It compresses the information required for continued operation into a smaller context. Together, caching and compaction move the optimization target from tokens used in one call to the context, calls, and retries required for one completed task. That is why the guide asks teams to measure cost per successful task before deployment instead of relying on nominal per-token prices.
Long-Running Tasks Need Interruptible Orchestration
Once GPT-6 is operating on code, websites, desktop applications, and external tools, a prompt is no longer a complete control plane. The guide recommends steering, asynchronous tools, and delegation so the model can continue independent work while a tool is running. It also recommends running independent tasks together so that one slow step does not block unrelated work. GPT-6.1 Sol supports multi-agent workflows, but that capability is still in beta and should not be treated as mature infrastructure by default.
Mid-task control has a defined limitation. Updated instructions can be sent through the Responses WebSocket, but new instructions are queued and do not cancel tools that are already running. Steering is therefore an additional control signal for a continuing task, not an emergency brake that takes effect everywhere immediately. A production system must decide in advance which actions can proceed independently and which require human confirmation. It also needs recovery paths for tool failures, changing permissions, and incorrect results. Without those boundaries, more automation also means a larger impact when control is lost.
The design also requires teams to align prompts, skills, and repository instructions. These materials must consistently state what the model should deliver, what it may do independently, what counts as complete, and which documents or tests matter for the current task. Code tests, browser verification, and high-risk changes should be observable stages in the workflow rather than judgments left entirely to the model. For an engineering team,