What OpenAI’s Agent Training Inside Your Software Means For Ironclad
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: What OpenAI’s Agent Training Inside Your Software Means For Ironclad on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get movie nights delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI says it trained GPT-6 Astra in hosted copies of Ironclad’s contract-management software, using 11 legal, commercial and procurement tasks. Astra met an average 55% of each task’s evaluation criteria; OpenAI’s estimated completion times were simulated, not measured customer savings. The test shows both the potential for software-specific agent training and the limits that remain before high-stakes workflows can run without human review.

OpenAI said on October 6 that it trained its GPT-6 Astra model on legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. The test averaged 55% of evaluation criteria met across 11 tasks, a result that points to progress in training agents on real business workflows but falls short of showing they can handle contract work without human oversight.

The work asked models to complete tasks selected by Ironclad staff and OpenAI employees familiar with the product. The 11 examples included setting up nondisclosure agreements, building procurement approval processes and changing a reusable contract clause to reflect a requester’s selected jurisdiction. OpenAI estimated an experienced user would take 30 to 40 minutes on each task.

Tasks were evaluated against rubrics containing 8 to 50 criteria, depending on complexity. OpenAI reported that GPT-5.6 Sol, run at a high setting, met an average 41.6% of criteria, while GPT-6 Astra, at a maximum setting, met 55.0%. An internal OpenAI model used during Astra’s development reached 63.7%. On one highlighted task, Astra met about 94% of the criteria; that single result is not the average across the test.

OpenAI said Ironclad provided hosted product copies for model practice and that it created synthetic training tasks using publicly filed contracts from the SEC’s EDGAR database, filtered to remove personal information. The company said it used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data. It also reported estimated attempt times of 37 minutes for GPT-5.6 Sol and 19.2 minutes for Astra, while stating those figures were simulated estimates based on assumed processing and generation speeds—not measured customer time savings.

At a glance
reportWhen: Published October 6; further partner wo…
The developmentOpenAI published details of a joint test with Ironclad in which a frontier model practised contract workflows inside hosted copies of the company’s software.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Partial Scores Matter in Contracting

The test addresses a practical challenge for AI agents: completing multi-step tasks in specialised software while keeping track of business rules. For software vendors, training in a product may help reveal where agents fail and could make future models more capable within that product. For customers, the result is a reason to ask how performance is measured before allowing an agent to handle approvals, clauses or other consequential work.

A score of 55% of criteria met is not the same as completing 55% of tasks or delivering 55% of the value. A procurement workflow might need Finance approval above a spending threshold, Security review for certain requests and Legal review for nonstandard terms. Missing one required check can make the workflow unsafe, even if many other criteria are satisfied. The rubric average alone does not identify which rules were missed on each task.

The reported time estimates also do not establish a productivity gain. They are simulated, and the comparison is with an estimated experienced-user task time—not a trial measuring employees’ actual work. Until accuracy and review requirements are demonstrated in customer settings, the human remains responsible for checking the work. The experiment is better read as a research result than as evidence of deployment-ready contract automation.

Amazon

AI contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Test Was Set Up

OpenAI’s October 6 post described two different developments that day; the Ironclad announcement concerned training and evaluating an agent inside business software, rather than introducing a new agent framework called Ironclad. Ironclad is a contract-management software company. The test focused on whether a model could follow workflows and preserve rules in that product.

OpenAI said its evaluation design uses task-specific rubrics to check whether work meets stated requirements. The post also acknowledged that an agent can lose track of a business rule partway through a workflow, limiting what a software company can safely delegate. Ironclad CTO Sunita Verma stressed that agents must preserve “the controls teams rely on.” The post presented human oversight as necessary for the work described.

OpenAI is also inviting a small number of software companies to partner on tasks agents cannot yet complete reliably. The post says prospective partners should bring a concrete example of failure, people with detailed knowledge of the work, a secure test environment and data that can safely be used for research. It frames the Ironclad work as an example of this partnership approach.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Evaluation Does Not Show

The published results do not establish how Astra would perform across Ironclad’s broader range of workflows, on live customer work or under varied operating conditions. OpenAI reported an average rubric score and a result for one showcase task, but the supplied information does not give a full task-by-task account of which requirements were missed or how often errors affected consequential controls.

It is also unclear whether the model can maintain reliable performance over repeated use, how much human checking each task requires, or whether customers would see net time savings after review and correction. The time estimates cover the 11 research tasks and rely on simulation; they are not measurements of actual customer use. OpenAI’s stated data limits describe the materials used in this work, but do not establish the terms or results of future partnerships.

The broader commercial implications remain prospective. Agents could make software more useful, but the source does not show that customers are replacing product interfaces or that vendors’ business models are changing. Claims about those outcomes would go beyond the published test.

Amazon

AI-powered contract review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Further Software Partnerships Ahead

OpenAI says it plans to work with a small number of software companies on tasks current agents cannot reliably complete. The company has asked potential partners to provide concrete examples of failures, domain experts, secure test environments and research-safe data. The source does not give a timetable, name additional partners or announce a product release tied to the Ironclad test.

For customers considering agents in contract or procurement systems, the next useful evidence would include task-by-task results, the specific criteria agents miss, the level of human review needed and measured outcomes in real customer workflows. Until such evidence is available, the published figures describe a controlled research evaluation, not verified time savings or permission to automate high-stakes work without checks.

Amazon

contract clause editing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI and Ironclad test?

They tested whether OpenAI models could complete 11 legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. The tasks included setting up agreements and building approval workflows.

What does Astra’s 55% score mean?

It is the average share of evaluation rubric criteria met across the tasks. It does not mean Astra completed 55% of the tasks, and the average does not show which specific requirements were missed.

Did Astra cut contract-work time by about half?

OpenAI reported a simulated estimate of 19.2 minutes per Astra attempt, compared with 37 minutes for GPT-5.6 Sol. It said the figures rely on assumed processing and generation speeds and are not measured customer time savings.

What data did OpenAI say it used?

OpenAI said it created synthetic tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It said it did not use OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.

Can companies use agents for contract workflows without review?

The results do not support that conclusion. OpenAI’s average score was 55% of criteria met, and the post says human oversight remains necessary. The published test does not establish reliable, unsupervised performance in customer workflows.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

When a Content Network Starts Publishing to Itself

A growing trend sees content networks shifting from external distribution to internal publishing, transforming ecosystems and audience control.

Meet ‘Little Gen Z’

A new poll reveals growing political divergence within Gen Z, especially among men aged 18-22, indicating a split between older and younger members of the generation.

Why Business Leaders Are Obsessed With Resilience

Inevitable challenges push business leaders to prioritize resilience, unlocking the secrets to sustained success and what they can achieve by embracing it.

The runway.How enterprise-revenuelock becomes the load-bearing valuation argument.

OpenAI and Anthropic prepare for historic IPOs, emphasizing enterprise revenue lock as the core justification for their high valuations amidst profitability uncertainties.