Evidence at a glance
The Test Moves from Contract Review to Workflow Setup
On October 6, 2026, OpenAI announced a research collaboration with contracting software company Ironclad to train and evaluate agents on complex work in specialized software. The partners brought practical tasks from legal, commercial, and procurement workflows into the research. GPT-6 Astra is the first frontier model trained on Ironclad tasks. Rather than answering questions about an individual contract, it is asked to configure agreements, approvals, and reusable legal terms inside the software, then check whether the resulting setup meets the original requirements.
This has a different boundary from Ironclad’s AI Assist, introduced in 2023. That earlier tool helped users identify contract anomalies and suggested redlines or company-approved terms, but people still decided whether to accept the suggestions. The new research aims at multi-step execution. An agent must operate the software while retaining the rules and checking whether the workflow behaves correctly under different conditions. That shift matters to technical leaders because the hard part of professional automation is moving from “can it click?” to “can it carry the business rules through the whole process?”
Evaluate the Whole Task, Not a Single Action
Ironclad employees and OpenAI staff who use Ironclad first selected 11 tasks spanning legal, commercial, and procurement work. Examples include setting up nondisclosure agreements, creating procurement approval processes, and updating reusable terms according to a requester’s chosen jurisdiction. OpenAI estimates that an experienced user would take 30 to 40 minutes to complete each task on average. This gives a sense of the configuration work involved, but it is an estimate, not a measured duration across customer workflows.
Depending on complexity, each task was assessed against 8 to 50 criteria. The evaluation therefore looks beyond whether one action succeeded. A procurement process might require finance approval above a spending threshold, security review for certain requests, and legal review for nonstandard terms. The agent has to express those conditions in forms, templates, and approval rules, then check that requests on both sides of the threshold follow the right paths. Scoring the whole workflow can reveal requirements lost along a long sequence of actions, although the result still depends on which situations the task design and criteria cover.
A Software Practice Environment, with a Defined Data Boundary
Ironclad provided hosted environments of its software so the models could practice these tasks in the product itself. OpenAI created synthetic training tasks around representative workflows and used reinforcement learning and feedback to improve the models. Together, practitioner input about what counts as completion, repeatable tasks, and a place to practice make this more relevant to contracting than a generic test of interface actions.
The source and limits of the training material also matter. OpenAI says the synthetic tasks were based on public contracts in the U.S. Securities and Exchange Commission’s EDGAR database. The training and evaluation did not use OpenAI customer data, OpenAI’s internal contracts, or nonpublic Ironclad customer data and contracts. That addresses concerns about relying on private customer contracts for training, but it does not show that a model can automatically understand each company’s internal policies, authorization structures, or exceptions. Public contracts can provide examples of terms and document structures. Whether those patterns transfer to a company’s own workflows is a separate question that still needs testing.
Higher Scores and Lower Estimated Time Are Not Production Gains
Across the 11 research tasks, GPT-6 Astra scored an average of 55.0%, compared with 41.6% for GPT-5.6 Sol. OpenAI describes Astra’s average as 32% higher, while the direct difference between the scores is 13.4 percentage points. The reasoning settings also differed: Astra used Max, the setting on which it scored highest, and Sol used High, its highest-scoring setting. The result therefore compares each model in its best reported setting, rather than isolating model performance under an identical reasoning configuration.
The time figures are not observed customer results either. OpenAI estimates an average of 19.2 minutes per attempt for Astra and 37.0 minutes for Sol, describing Astra’s estimate as 48% lower. The estimates are based on model processing and generation speed. In one demonstration task, Astra took about 20 minutes to meet roughly 94% of the criteria, while Sol took about 32 minutes to meet roughly 85%. That is a single example, not the average across all 11 tasks. OpenAI also reports a 63.7% score for an internal model used during Astra’s development, but provides no name or further evaluation details. These figures indicate a direction of progress, not proof of hours saved in practice or reliable completion of complex contracting work.
Bring the Evaluation into Your Own Workflow Before Delegating
For technical leaders, the most reusable part of this research is the method of turning hard-to-automate business rules into repeatable evaluation tasks. Start with one approval workflow, specify spending thresholds, role permissions, exceptions, and final acceptance conditions, then record which criteria the agent misses, how much human intervention it needs, and whether real completion time changes. This makes “the model seems able to use the software” into an engineering problem that can be diagnosed. It also helps teams decide which steps might be automated and which should still require human confirmation.
The current evidence does not justify extending the Ironclad results to every contracting workflow. The public material does not say whether the research has entered Ironclad’s customer product. It also does not report an independent third-party review, explain how representative the 11 tasks are of customer work, or establish performance beyond those tasks. An average score of 55% means many criteria were still unmet, while the time advantage comes from a speed estimate. The safer current judgment is to treat these agents as supervised workflow assistants, validate them against local rules and exception paths, and expand their permissions only in response to observed results. That is a more defensible route than handing over contracting work as a whole.