Three model cards: GPT-6 Astra, our most intelligent model for the best results, $10 input, $50 output, and $1 cached input; GPT-6.1 Sol, near-Astra intelligence for a fifth of the price, $2 input, $10 output, and $0.10 cached input; GPT-6
Source material Open source material ↗
Introducing GPT-6 Sol and Luna — Art card
Source material Open source material ↗
DevDay 2026 Recap — cover image (1:1)
Source material Open source material ↗

Evidence at a glance

Approx. GPT-6 AstraEvidence
Input 0.10 , 95%Evidence
Input GPT-6 Sol 50%Evidence
DeepSWE Astra, Approx.Evidence
DeepSWE GPT-6 Sol 6.4Evidence
GDP.pdf Opus 5.5,Evidence

Frontier-Like Capability Is No Longer Limited to Frontier Pricing

OpenAI has introduced GPT-6.1 Sol, an upgrade to GPT-6 Sol aimed at agentic coding, computer use, professional document work, and multi-step business workflows. It is not positioned as the new top-performing model. Instead, OpenAI says it delivers near-GPT-6 Astra intelligence at roughly one-fifth of Astra’s standard input and output token prices. Developers can call it through the OpenAI API as gpt-6.1-sol, within an ecosystem that also includes products such as Codex and ChatGPT Work.

The important change is not another model name on a leaderboard. It is a shift in the economics of model selection. Near-frontier capability has often been reserved for a small number of high-value requests because every extra reasoning step, tool call, and context replay carries a premium. Sol is designed to move that level of capability into agents that run continuously, repeatedly reading context, using tools, changing code, and continuing from intermediate results without paying the flagship rate at every turn.

The Real Lever Is Cached Context, Not Just the Per-Call Price

GPT-6.1 Sol is priced at $2 per million input tokens and $10 per million output tokens, while cached input costs $0.10 per million tokens. That is 95% below its standard input price and 50% below GPT-6 Sol’s cached-input price. For a one-off question, caching may look like a billing discount. For an agent, it is closer to an architectural primitive because the agent may need to carry project instructions, tool results, task history, and intermediate state across many requests.

That makes Sol’s value impossible to judge from the price of a single answer alone. A coding agent may first inspect a repository, propose a change, run tests, read failure logs, and revise the patch. Replaying the full context at every stage can dominate the bill. Cheap cached input creates room for this repeated context, but it does not remove output tokens, tool calls, reasoning effort, or retries. “One-fifth the price” is therefore a model-pricing comparison, not a guarantee that every long-running agent will cost one-fifth as much.

The Evaluations Point to a Default Model, Not a Universal Fallback

OpenAI’s evidence focuses on long-horizon execution rather than simple factual questions. On DeepSWE v1.1, which evaluates software-engineering tasks in real codebases, GPT-6.1 Sol matches GPT-6 Astra at roughly one-fifth of the cost and improves on GPT-6 Sol’s best result by 6.4 percentage points at lower reasoning effort. On GDP.pdf, which tests professional questions over complex PDFs from finance, healthcare, legal, and other domains, Sol scores above Opus 5.5 with fallbacks at less than half the per-task cost, while approaching Astra at roughly one-fifth the cost.

On AutomationBench, Sol scores 2.2 percentage points above Opus 5.5 at medium reasoning effort for about one-third of the cost, and improves by 4.8 points over GPT-6 Sol at the same setting. On the offline set of OSWorld 2.0, it beats GPT-6 Sol by seven points at maximum reasoning effort, comes within 2.1 points of Astra, and costs roughly one-seventh as much per task. Taken together, these results support a practical role: when work requires sustained coordination across code, documents, interfaces, and tools, Sol can be the default tier instead of routing every request to Astra.

Scientific Work Reveals the Boundary of the Capability Tiers

Sol does not erase the gap on every difficult task. Terminal-Bench Science 0.1 covers scientific workflows including data analysis, simulation, model fitting, and theorem proving. At maximum reasoning effort, GPT-6.1 Sol more than doubles GPT-6 Sol’s score and costs an average of $5.47 per task. Opus 5.5 costs $23.21 and Astra $23.80, yet Astra still leads the tested models with a score of 68.1%.

That boundary matters more than the phrase “near frontier,” because it shows that model tiers remain useful. Sol can make a large volume of demanding work cheaper, but Astra should remain the escalation path for scientific tasks that require deeper reasoning, complex experimental chains, or a higher success threshold. OpenAI also reports that Sol’s factual-error rate at low reasoning effort falls from 11.4% for GPT-6 Sol to 7.7%, with the gap to Astra no larger than 1.9 percentage points across the tested settings. Those results come from difficult prompts and should not be treated as a universal accuracy guarantee for everyday conversations.

The Engineering Choice Is Dynamic Routing, Not a Full Migration

Treating GPT-6.1 Sol as simply “a cheaper Astra” would lead to a poor deployment decision. A more robust architecture is to place Sol in the default tier for code generation and debugging, complex document questions, routine computer use, and cross-tool business workflows. Escalate to Astra when a task shows high risk, insufficient scientific reasoning, repeated failure, or a requirement for a higher success rate. The central change is not model replacement but making model selection part of task routing and budget control.

Teams should track cache hits, input and output tokens, reasoning effort, tool-call counts, retries, and end-to-end success for each workflow. Competitive pricing also needs a common accounting basis: the source notes that roughly 40% of Claude Fable 5.1 tasks triggered fallbacks, while the cited cost omits them. The actionable decision is to let Sol handle repeatable and observable work with a clear escalation path, measure savings on real long-horizon bills, and keep Astra for the steps where failure is more expensive than the model premium.