🔍 Read the full analysis: Your Software, OpenAI’s Agents, And Ironclad’s Fine Print on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
OpenAI described training and testing GPT-6 Astra inside Ironclad’s contract-management software on 11 legal, commercial and procurement tasks. Astra met an average of 55% of task criteria, while its estimated completion time was simulated; the results do not establish customer productivity gains or readiness for unsupervised contract work.
OpenAI said on October 6 that it trained and evaluated its GPT-6 Astra model on tasks inside Ironclad’s contract-management software, reporting that Astra met an average of 55% of the criteria set for 11 legal, commercial and procurement workflows. The project matters because it tests AI agents against rules and processes in a specialised business product, but OpenAI’s results do not show that the system is ready to handle contracts without human review.
OpenAI said the tasks were selected by Ironclad staff and OpenAI employees who use the product. They included setting up nondisclosure agreements, creating procurement approval processes and changing a reusable contract clause based on a requester’s chosen jurisdiction. OpenAI estimated that an experienced user would need 30 to 40 minutes for each task. The tasks were scored against rubrics containing 8 to 50 criteria, depending on their complexity.
Ironclad provided hosted copies of its product for models to practise in. OpenAI said it generated synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database and filtered those materials to remove personal information. It also said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data. The report describes GPT-6 Astra as the first frontier model trained this way.
OpenAI reported an average of 55.0% of criteria met for Astra, compared with 41.6% for GPT-5.6 Sol at high reasoning effort. The estimated time per attempt was 19.2 minutes for Astra and 37.0 minutes for Sol. An internal OpenAI model used during Astra’s development reached 63.7%, while Astra met about 94% of the criteria on one showcase task. These are rubric scores across the reported tasks, not percentages of tasks completed successfully.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Partial Contract Scores Matter
The results point to a possible way of teaching agents to work inside professional software: test them on concrete tasks, score their actions against business requirements, and use failures to guide model development. For vendors, the collaboration could help expose where agents struggle in their products and improve how those products support automated work. OpenAI also said it is inviting a small number of software companies to explore similar partnerships.
But a high average score on a rubric does not necessarily mean a workflow is safe to use. A procurement process may require Finance approval above a spending threshold, Security review for certain purchases and Legal review for nonstandard terms. Missing even one of those steps could undermine the process, even if the agent satisfies many of the other criteria. OpenAI’s report says that losing track of a business rule limits what a company can confidently delegate and that human oversight remains necessary.
The shift could also affect what makes business software valuable. If users increasingly give instructions through an agent rather than operating screens themselves, a vendor’s lasting value may depend more on its underlying business rules, records, audit trail and controls. That is an implication of the experiment, not a demonstrated outcome. Ironclad’s report frames its platform as important to maintaining those controls.
AI-powered contract management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the Ironclad Test Was Set Up
The report is about model development and evaluation in a specific software product, rather than a general announcement that an autonomous contract agent is available to customers. The work focused on 11 selected research tasks and used a rubric to measure how many stated criteria the model met. OpenAI’s comparison with GPT-5.6 Sol provides a result on those tasks; it does not establish performance across every Ironclad workflow or other companies’ systems.
The timing figures also need careful interpretation. OpenAI said the per-attempt times were simulated estimates based on assumed processing and generation speeds, not measurements of customer work or observed time saved in production. They cover the research tasks, not Ironclad workflows generally. The reported reduction from 37.0 to 19.2 minutes compares estimated model attempt times between the two models; it is not evidence that customers completed work faster.
OpenAI’s final section invites a limited number of software companies to work on tasks current agents cannot reliably complete. It asks prospective partners to provide a concrete example of a failure, people with deep knowledge of the work, a secure test environment and data that can safely be used for research. The post does not name additional partners or give a timetable for future projects.
“the controls teams rely on”
— Sunita Verma, Ironclad’s chief technology officer, as quoted in the report
As an affiliate, we earn on qualifying purchases.
Limits of the Reported Results
The report does not establish how Astra would perform on all tasks in Ironclad, on live customer work or in other contract-management products. It also does not give a breakdown showing which criteria the model missed across the 11 tasks, beyond the reported averages and showcase example. That makes it hard to assess whether the remaining errors involved minor details or mandatory approvals and controls.
OpenAI’s figures do not show measured customer time savings, and the report does not provide evidence that the model can safely complete these workflows without human review. The source material gives no deployment schedule, commercial terms or details about how any future partner projects would be governed. OpenAI’s data-use statements describe the material used for this research, but do not by themselves establish the terms of future collaborations.
As an affiliate, we earn on qualifying purchases.
Further Software Partnerships
OpenAI said it plans to work with a small number of software companies on tasks that current agents cannot reliably complete. It has not announced which companies will take part or when the work will begin. The next useful evidence would include task-level results, clearer reporting on failure types and independent or customer-based measurements of performance in real workflows.
For companies considering agents in systems that handle contracts, finances or customer records, the immediate issue is how to test them against required business rules. Buyers can ask which criteria failed, how approvals and exceptions are checked, who reviews an agent’s work and what records are kept of its actions. Until stronger evidence is available, the reported results support further testing—not treating the agent as a substitute for review.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI and Ironclad test?
They tested OpenAI models on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management product. Examples included setting up nondisclosure agreements and building procurement approval processes.
Does the 55% result mean Astra completed 55% of the tasks?
No. OpenAI said Astra met an average of 55% of the rubric criteria across the tasks. That is not the share of tasks completed, and it does not mean each workflow was usable or correct.
Did OpenAI show that Astra saves customers time?
No. The reported times were simulated estimates based on assumed processing and generation speeds. OpenAI said they were not measured customer time savings and covered the 11 research tasks.
Can companies use Astra to handle contracts without review?
The report does not establish that. OpenAI said human oversight remains necessary because agents may fail to preserve business rules during multi-step work. The reported averages are not evidence of readiness for unsupervised contract processing.
What happens after the Ironclad project?
OpenAI said it is inviting a small number of software companies to explore similar work. It has not named additional partners or provided a schedule for those projects.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
