AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Your Software, OpenAI’s Agents, And Ironclad’s Fine Print on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training and testing GPT-6 Astra inside Ironclad’s contract-management software on 11 legal, commercial and procurement tasks. Astra met an average of 55% of task criteria, while its estimated completion time was simulated; the results do not establish customer productivity gains or readiness for unsupervised contract work.

OpenAI said on October 6 that it trained and evaluated its GPT-6 Astra model on tasks inside Ironclad’s contract-management software, reporting that Astra met an average of 55% of the criteria set for 11 legal, commercial and procurement workflows. The project matters because it tests AI agents against rules and processes in a specialised business product, but OpenAI’s results do not show that the system is ready to handle contracts without human review.

OpenAI said the tasks were selected by Ironclad staff and OpenAI employees who use the product. They included setting up nondisclosure agreements, creating procurement approval processes and changing a reusable contract clause based on a requester’s chosen jurisdiction. OpenAI estimated that an experienced user would need 30 to 40 minutes for each task. The tasks were scored against rubrics containing 8 to 50 criteria, depending on their complexity.

Ironclad provided hosted copies of its product for models to practise in. OpenAI said it generated synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database and filtered those materials to remove personal information. It also said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data. The report describes GPT-6 Astra as the first frontier model trained this way.

OpenAI reported an average of 55.0% of criteria met for Astra, compared with 41.6% for GPT-5.6 Sol at high reasoning effort. The estimated time per attempt was 19.2 minutes for Astra and 37.0 minutes for Sol. An internal OpenAI model used during Astra’s development reached 63.7%, while Astra met about 94% of the criteria on one showcase task. These are rubric scores across the reported tasks, not percentages of tasks completed successfully.

At a glance
reportWhen: Published October 6; further partner wo…
The developmentOpenAI published a report on October 6 describing a collaboration with Ironclad to train and evaluate a frontier model on workflows in the contract-management product.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Partial Contract Scores Matter

The results point to a possible way of teaching agents to work inside professional software: test them on concrete tasks, score their actions against business requirements, and use failures to guide model development. For vendors, the collaboration could help expose where agents struggle in their products and improve how those products support automated work. OpenAI also said it is inviting a small number of software companies to explore similar partnerships.

But a high average score on a rubric does not necessarily mean a workflow is safe to use. A procurement process may require Finance approval above a spending threshold, Security review for certain purchases and Legal review for nonstandard terms. Missing even one of those steps could undermine the process, even if the agent satisfies many of the other criteria. OpenAI’s report says that losing track of a business rule limits what a company can confidently delegate and that human oversight remains necessary.

The shift could also affect what makes business software valuable. If users increasingly give instructions through an agent rather than operating screens themselves, a vendor’s lasting value may depend more on its underlying business rules, records, audit trail and controls. That is an implication of the experiment, not a demonstrated outcome. Ironclad’s report frames its platform as important to maintaining those controls.

Amazon

AI-powered contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Test Was Set Up

The report is about model development and evaluation in a specific software product, rather than a general announcement that an autonomous contract agent is available to customers. The work focused on 11 selected research tasks and used a rubric to measure how many stated criteria the model met. OpenAI’s comparison with GPT-5.6 Sol provides a result on those tasks; it does not establish performance across every Ironclad workflow or other companies’ systems.

The timing figures also need careful interpretation. OpenAI said the per-attempt times were simulated estimates based on assumed processing and generation speeds, not measurements of customer work or observed time saved in production. They cover the research tasks, not Ironclad workflows generally. The reported reduction from 37.0 to 19.2 minutes compares estimated model attempt times between the two models; it is not evidence that customers completed work faster.

OpenAI’s final section invites a limited number of software companies to work on tasks current agents cannot reliably complete. It asks prospective partners to provide a concrete example of a failure, people with deep knowledge of the work, a secure test environment and data that can safely be used for research. The post does not name additional partners or give a timetable for future projects.

“the controls teams rely on”

— Sunita Verma, Ironclad’s chief technology officer, as quoted in the report

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Reported Results

The report does not establish how Astra would perform on all tasks in Ironclad, on live customer work or in other contract-management products. It also does not give a breakdown showing which criteria the model missed across the 11 tasks, beyond the reported averages and showcase example. That makes it hard to assess whether the remaining errors involved minor details or mandatory approvals and controls.

OpenAI’s figures do not show measured customer time savings, and the report does not provide evidence that the model can safely complete these workflows without human review. The source material gives no deployment schedule, commercial terms or details about how any future partner projects would be governed. OpenAI’s data-use statements describe the material used for this research, but do not by themselves establish the terms of future collaborations.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Further Software Partnerships

OpenAI said it plans to work with a small number of software companies on tasks that current agents cannot reliably complete. It has not announced which companies will take part or when the work will begin. The next useful evidence would include task-level results, clearer reporting on failure types and independent or customer-based measurements of performance in real workflows.

For companies considering agents in systems that handle contracts, finances or customer records, the immediate issue is how to test them against required business rules. Buyers can ask which criteria failed, how approvals and exceptions are checked, who reviews an agent’s work and what records are kept of its actions. Until stronger evidence is available, the reported results support further testing—not treating the agent as a substitute for review.

Amazon

Contract review AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI and Ironclad test?

They tested OpenAI models on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management product. Examples included setting up nondisclosure agreements and building procurement approval processes.

Does the 55% result mean Astra completed 55% of the tasks?

No. OpenAI said Astra met an average of 55% of the rubric criteria across the tasks. That is not the share of tasks completed, and it does not mean each workflow was usable or correct.

Did OpenAI show that Astra saves customers time?

No. The reported times were simulated estimates based on assumed processing and generation speeds. OpenAI said they were not measured customer time savings and covered the 11 research tasks.

Can companies use Astra to handle contracts without review?

The report does not establish that. OpenAI said human oversight remains necessary because agents may fail to preserve business rules during multi-step work. The reported averages are not evidence of readiness for unsupervised contract processing.

What happens after the Ironclad project?

OpenAI said it is inviting a small number of software companies to explore similar work. It has not named additional partners or provided a schedule for those projects.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Microsoft Word For Windows 1.1A, Native X64 Port

Microsoft releases Word for Windows 1.1a with native support for x64 architecture, marking a significant update for performance and compatibility.

Top Features Of xAI’s Imagine Image 2.0 In Grok Quality Mode For AI Developers

xAI announced the release of Imagine Image 2.0 within Grok’s Quality Mode, but technical details and availability remain unconfirmed.

Windows 11’S Built-in Weather App Wastes More Than 1 GB Of RAM

Researchers find Windows 11’s built-in Weather app uses more than 1 GB of RAM, raising concerns about efficiency and system impact.

Swami Sarvapriyananda Travels To The US For A Private Meeting On Claude

India Today reported that Anthropic flew Swami Sarvapriyananda to the U.S. for a private meeting about Claude training; details remain unconfirmed.