


Evidence at a glance
One Workbook Exposes the Real Bottleneck in Long Tasks
In a case published by OpenAI on September 28, 2026, Basis described a comparison involving GPT-6 Astra. Basis builds AI agents for accounting work, with the aim of automating much of the manual work accountants perform every day. The test asked GPT-6 Astra and GPT-5.6 Sol to complete a complicated tax workbook accurately and reliably. The workbook contained 50 tabs, and Astra took 50% less time than Sol, which Basis described as completing the task twice as fast.
The number does not primarily show that one tax answer can be generated more quickly. It suggests that a large share of the cost in a long workflow may come from choosing the wrong path and correcting it later. An agent can produce acceptable individual outputs and still waste substantial time if it misreads the goal at the beginning, then has to backtrack, redo work, or add missing explanations. Basis specifically observed better early decisions and less correction with Astra, offering a more concrete explanation for the time reduction than raw generation speed alone.
The Advantage Appears Before the Main Work Begins
Basis’s description of Astra is not simply that it reasons more powerfully. The model is said to understand more clearly what the user is actually trying to accomplish. It can use broader context to infer the goal, determine when to ask a question, identify assumptions that should be flagged, and carry out the user’s instructions. For a tax agent, these judgments often occur before the workbook is filled, calculations are made, or sources are retrieved, yet they determine which path the rest of the work will follow.
That shifts what an agent system should optimize. Teams often focus on prompt design, tool calls, and the quality of individual answers, then compensate for missing business requirements with extensive rules. Astra gives Basis a different possibility: if the model can infer requirements such as following a template, consulting primary tax sources, and checking its own work, the system may need fewer rules for each individual situation. Fewer rules do not make the system inherently simple. They move part of the control logic from external configuration into the model’s ability to interpret context.
Adaptive Reasoning Turns Capability into Cost Difference
Another mechanism described by Basis connects model capability more directly to operating cost. As a task progresses, Astra can adjust its reasoning effort according to the difficulty of the current step. It uses more computation for difficult steps and less for easier ones, while keeping its cache intact during the adjustment. For a workbook with 50 tabs, this means the system does not have to apply the same reasoning intensity to every part of the job.
The value is not that every step becomes faster. It is that expensive computation need not be spread evenly across simple steps. A long task may combine source discovery, formatting, judgment, review, and exception handling, each requiring a different level of computation. If the model can retain prior context while temporarily increasing effort at difficult points, it may reduce both latency and cost. The material does not disclose Astra’s specific pricing or the absolute cost per workbook, so the evidence establishes a mechanism and a direction, not a complete economic case.
The Internal Evaluation Gain Points to Controllable Behavior
Basis also reported an approximately 20% improvement in its internal evaluation scores with Astra. The evaluations do not look only at whether the final answer is correct. They also check whether the agent follows templates, consults primary sources for tax questions, and checks its own work. In other words, the evaluation asks whether a workflow behaves as the business expects, rather than whether a model can produce an impressive answer to an isolated question.
This distinction matters to technical leaders because the hardest part of enterprise deployment is rarely one prompt. It is the large set of boundary behaviors: when the agent should stop and ask the user, when it should expose an assumption, when it must cite a source, and when it should refuse to proceed without sufficient evidence. If Astra can infer these requirements from broader context, it may reduce the burden of maintaining rules for every case and improve coverage beyond situations explicitly included in internal tests. But fewer explicit rules also mean more behavior depends on model judgment. The evaluation system must therefore be specific enough to reveal when that dependency fails.
Twice as Fast Is Still Not a Deployment Decision
The most useful lesson from this case is to bring model selection back to the economics of a complete task. For a tax agent, a team can track completion time, rework, the timing of questions, the correctness of primary-source citations, template adherence, self-check results, and cost per task. Astra’s reported behavior suggests that early planning and adaptive reasoning should be part of those measures rather than searching for the answer only in model benchmark scores.
The evidence also has clear limits. The 50% time reduction comes from Basis’s test on one 50-tab tax workbook. It cannot be generalized to every tax workflow, and it does not establish a twofold improvement in accuracy, compliance, or end-to-end return on investment. The material does not disclose the billing details of Astra’s adaptive reasoning or provide an independent validation of tax citation accuracy. A responsible deployment can begin with a bounded long-running workflow, compare models using the same inputs and audit standards, and place speed gains alongside the cost of errors. The safest conclusion is not that a stronger model should replace the whole system. It is that the model should prove it can choose the right path earlier in the organization’s own process, while leaving enough evidence to explain when it does not.