One Release Cycle, Four Different Trade-Offs
Anthropic, OpenAI, and Google DeepMind released Claude Fable 5.1, GPT-6 Astra, GPT-6.1 Sol, and Gemini 4 Argon within about 30 days. All four are frontier models aimed at demanding knowledge work and software development, but they differ in price, access, context capabilities, and benchmark results. The useful question is not which model wins every task, but which one can perform a given job at an acceptable cost and within the right access constraints.
This comparison matters because the benchmark gaps are less decisive than launch messaging might suggest, while actual bills and deployment conditions diverge more sharply. There is also a governance signal: OpenAI cancelled GPT-6.1 Astra on September 28 after it failed internal scope and authorization tests, according to the material. Model capability is not only about what a system can do, but also whether it can be kept within the work it is authorized to perform.
The Leaderboard Is Not a Single Capability Curve
The published comparisons do not show one model sweeping every category. In Google DeepMind's comparison, Astra scores 77.9% on DeepSWE v1.1, Argon 74.1%, and Fable 5.1 67.4%. But on FrontierSWE v2, Argon leads at 65.5%, ahead of Astra at 55.0%. Knowledge work is not a single dimension either: on the Vals Index, Astra scores 68.9%, Argon 63.1%, and Fable 5.1 65.8%. These are vendor-reported results, not an independent, common-lab ranking of every model.
The pattern shifts again on tasks involving real tool use. In the listed OSWorld 2.0 results, Astra scores 69.2% and Argon 72.6%; Fable 5.1 is absent from Google's table. The three reported Terminal-Bench 4.0 scores cluster between 57.4% and 58.2%. A separate set of figures from OpenAI says Sol is close to Astra on DeepSWE v1.1 and within 2.1 percentage points on the OSWorld 2.0 offline set. Since Sol is not in Google's comparison, these results cannot be combined into a fully like-for-like leaderboard.
The Same List Price Does Not Mean the Same Cost per Task
Per million tokens, Astra and Fable 5.1 both list at $10 for input and $50 for output. Sol lists at $2 and $10, as does Argon during its introductory period, after which Argon's price doubles. Long-context and cached-input pricing also affect the bill: cached input costs $1 per million tokens for Astra, $0.25 for Fable 5.1, and $0.10 for Sol. Argon's cached-input rate is also $0.10 during its introductory period. The material does not disclose Argon's long-prompt pricing.
Cost per task also depends on how many tokens a model uses to finish the job. Artificial Analysis reports $9.18 per task for Fable 5.1 and $4.72 for Astra, despite their identical list prices. That result reflects both token consumption and rates in that particular evaluation; it is not a fixed cost for every workload. The material also notes that Anthropic's newer tokenizer produces roughly 30% more tokens for the same text, so comparing bills across models requires looking beyond the per-million-token rate.
Agent Selection Must Account for the Call Pattern
Cached-input pricing matters especially for agents. Across multiple steps, an agent often resends system prompts, tool definitions, and conversation history, so the cost of repeatedly reading that input can accumulate. The material puts Astra's cache-read price at four times Fable 5.1's and ten times Sol's. For a workflow with frequent calls and relatively stable context, a lower cache rate may affect the total bill more directly than a small lead on a benchmark.
Context and output limits also determine whether a model can handle a particular job. Astra and Sol have context windows of 1.05 million tokens, while Fable 5.1 has 1 million; Argon's context window is not disclosed in the material. The more striking difference is maximum output per response: 1 million tokens for Argon, compared with 128,000 for each of the other three. That limit may suit tasks requiring an unusually long single response, but it does not show that Argon is more accurate or economical on every long task. Access routes matter too: Astra and Sol are available through the OpenAI API, ChatGPT, and Codex; Fable 5.1 is also available through Bedrock, Google Cloud, and Microsoft Foundry, while Argon was available only through the Fairwind Program during the period described.
Authorization Tests Add Another Gate to Capability Evaluation
GPT-6.1 Astra's cancellation should not be read simply as a failed launch, nor inflated into a complete account of model risk. The material says only that it failed internal scope and authorization tests; it does not disclose the tests, the reasons for failure, or their boundaries. That fact is enough to remind technical leads that for systems able to call tools, operate interfaces, or perform multi-step tasks, staying within authorized scope is a launch requirement, not an appendix to the capability scorecard.
For procurement or deployment, it helps to split the evaluation into two questions. First, test completion quality, call counts, and token consumption on the target workload. Then separately check tool permissions, scope limits, and behavior after a failure, while confirming whether prices are temporary and access routes are dependable. Many results in the material are vendor-reported, some models lack directly comparable figures, and Argon's introductory price will change. The soundest decision is therefore not to pick an all-purpose champion, but to run a small, like-for-like evaluation for the intended tasks and make authorization testing a condition of launch.