OpenAI Ironclad GPT-6 Astra Study Targets Contract Work
OpenAI says a research collaboration with contracting software maker Ironclad helped GPT-6 Astra score 32% higher than its predecessor on 11 contracting tasks while taking about half the time.

OpenAI is trying a new way to make AI agents useful at specialized business software: bringing the software makers into the training process. On Tuesday, October 6, 2026, the company described a research collaboration with Ironclad, which sells AI contracting tools, as the first of a small number of such partnerships. The OpenAI Ironclad GPT-6 Astra results, published by OpenAI, show its newest frontier model completing more of a set of contracting tasks in less estimated time than the previous generation.
What happened
In a post on its website, OpenAI said it is partnering directly with a few software companies "that understand these workflows best" to identify challenging, high-value tasks and turn them into research problems for training and evaluating models. Ironclad is the first. Ironclad employees and people who use Ironclad at OpenAI helped researchers identify 11 tasks across legal, commercial and procurement work, such as setting up nondisclosure agreements, creating procurement approval processes and updating a reusable legal clause so that it reflects the jurisdiction a requester selects. OpenAI estimates each task would take an experienced user about 30 to 40 minutes.
Each task was graded against 8 to 50 criteria depending on its complexity. Ironclad also provided hosted environments of its product where models could practice. OpenAI said its researchers built synthetic training tasks around representative workflows and used reinforcement learning to improve the models through practice and feedback. The company said the simulated tasks were created from contracts publicly available in the SEC's EDGAR database after filtering out personal information, and that it did not use OpenAI customer data, its own internal contracts, or nonpublic Ironclad customer data for training or evaluation.
The headline numbers compare GPT-6 Astra, using its maximum reasoning setting, with GPT-5.6 Sol on high reasoning, the settings where each scored best. Across the 11 tasks, Astra's average rubric score was 55.0%, against 41.6% for Sol, a 32% relative improvement. Estimated average time per attempt fell from 37.0 minutes to 19.2 minutes, or 48% lower. An internal model used in Astra's development reached 63.7%, and OpenAI said it aims to bring those gains to future models. Crypto Briefing noted that OpenAI introduced GPT-6 Astra in late September as its most capable model for professional tasks, making the Ironclad results an early field report on that claim.
Why it matters
The OpenAI Ironclad GPT-6 Astra work addresses a known weakness of AI agents: keeping many business rules in view across a long, multi-step workflow. OpenAI gave an example of a legal operations team setting up software purchasing, where finance must approve spending above a threshold, security must review certain requests and legal must check nonstandard terms. "Getting individual steps right is not enough," the company wrote; the finished process must work across all the situations it was designed for.
The partnership model also matters commercially. Domain specialists define the tasks and the grading criteria, which gives AI developers a better signal about what "good" looks like, and gives software companies early influence over how frontier models handle their core use cases. "Agents need to understand the full contracting lifecycle, including how business workflows connect while preserving the controls teams rely on," said Sunita Verma, Ironclad's chief technology officer, in OpenAI's post.
The results deserve careful reading. OpenAI says they cover only the 11 research tasks, not all Ironclad workflows, and that the times are simulated estimates based on assumed processing speeds, not measured customer time savings. A 55% average rubric score is a clear improvement but is not a level most legal teams would accept for unsupervised work, a point Crypto Briefing also made. OpenAI itself said the collaboration "underscores why human oversight still matters." The benchmark was co-designed with a single company, so how well the gains carry over to other contracting processes, industries or jurisdictions remains open.
Called It has tracked how quickly AI agents are moving into regulated work, from Robinhood's previewed trading agents to the model race reflected in Google's Gemini 4 Argon launch.
What's next
OpenAI is inviting more software companies to apply, asking for a concrete task that today's agents fail at, evidence of the failure and a way to judge success, along with domain experts, a secure test environment and data that can be safely used for research. Expect similar announcements in other professional software categories if the approach produces measurable gains. For Ironclad, the next test is whether these model improvements show up in its own products. For buyers, the OpenAI Ironclad GPT-6 Astra figures are a research benchmark, not a guarantee of performance in their own contracting workflows.
This article is for information only and is not investment advice.