OpenAI’s Software Agent Training: A Reader’s Guide To Ironclad’s Fine Print
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI’s Software Agent Training: A Reader’s Guide To Ironclad’s Fine Print on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training a frontier model in Ironclad’s contract-management software using 11 selected legal, commercial and procurement tasks. GPT-6 Astra met an average 55% of task criteria, while the reported time estimates are simulations—not measured customer productivity gains. OpenAI says it used public contract filings for synthetic tasks and no non-public Ironclad customer data.

OpenAI has described training a frontier model inside Ironclad’s contract-management software, using 11 legal, commercial and procurement tasks to test how well it follows business rules. OpenAI reported that GPT-6 Astra met an average 55% of the criteria across those tasks, while stressing that its time estimates are simulated and do not measure customer savings.

Ironclad employees and OpenAI staff who use the product selected tasks such as creating nondisclosure agreements, setting up procurement approval processes and updating a reusable contract clause based on a requester’s chosen jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes for each task. The work was assessed against rubrics containing 8 to 50 criteria, depending on complexity.

Ironclad supplied hosted copies of its software for model practice. OpenAI says it generated synthetic tasks from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. It says it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data. The reported average share of criteria met was 41.6% for GPT-5.6 Sol at high effort and 55.0% for GPT-6 Astra at maximum effort. An internal OpenAI model used during Astra’s development scored 63.7%.

The 55% figure is a share of rubric criteria met, not a share of tasks completed or a measure of safe deployment. OpenAI also reported a result of about 94% of criteria on one showcase task, but that single example does not establish typical performance. The company’s estimated attempt times fell from 37 minutes for GPT-5.6 Sol to 19.2 minutes for Astra; OpenAI says those figures use assumed processing and generation speeds and are simulated, not measured.

At a glance
reportWhen: Published October 6; the source does no…
The developmentOpenAI published details of a training collaboration with Ironclad, using the contract-management platform’s workflows to evaluate and train an AI model.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Contract Workflow Accuracy Matters

The results point to both the potential and the limits of training agents on the software people use for professional work. A model that can navigate a specialised product and follow company rules could reduce repetitive work. But in contracting and procurement, a missed requirement may undermine the whole workflow: skipping a required finance, security or legal approval is not made safe by getting other steps right.

For businesses, the reported average does not support treating these agents as independent operators. Human review remains central, particularly where approvals, contract terms or compliance obligations are involved. The useful question is not just how many criteria an agent meets, but which ones it misses and whether those errors can be caught before a workflow takes effect.

The arrangement also has implications for software providers. Training in a vendor’s product may improve an agent’s ability to use it, while also making the agent a potential route through which customers interact with that product. A vendor’s lasting value may depend increasingly on its business rules, records, controls and audit trail, rather than only on the interface agents learn to operate. That is an interpretation of the collaboration’s possible effects, not a result demonstrated by the reported tests.

Amazon

AI contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Tests Were Set Up

OpenAI’s October 6 post was one of two publications described in the source material; the other, involving 722 mathematics manuscripts, drew more attention. The Ironclad post focused instead on training models for multi-step work in specialised business software. OpenAI framed the project as a way to teach models to understand company rules, carry out workflows and check whether the result meets the task’s requirements.

The experiment used a limited set of 11 tasks selected by people familiar with the product and work. It was not described as a broad test of all Ironclad functions or a live deployment measuring customer outcomes. OpenAI says it is inviting a small number of software companies to work on tasks current agents cannot reliably complete. It asks potential partners to bring a concrete example of failure, knowledgeable staff, a secure test environment and data that can safely be used for research.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Scores Do Not Establish

The post does not establish how the model would perform across Ironclad’s full range of workflows, on different companies’ rules, or in live customer use. The 55% average does not show which criteria were missed on each task, and the approximately 94% result was for one showcase task. The source material also does not provide enough detail to assess how the task set was selected or how repeatable the scores are.

There is no measured evidence here of customer time saved, financial return or reduced error rates. OpenAI’s time figures are simulations, and the tasks are a small research set. It is also unclear when any partner work might lead to a customer-facing product, what deployment safeguards would apply, or whether other software companies will take part.

Amazon

AI-powered NDA creation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What OpenAI’s Partner Program Requires

OpenAI says it is seeking a small number of software-company partners to study difficult workflows that current agents do not reliably complete. The next steps described in the post are for partners to identify a concrete failure case, provide people with deep knowledge of the work, and supply a secure environment and research-eligible data. OpenAI has not provided a public timetable for further partnerships or product releases in the source material.

Companies considering agent use in contract or procurement systems can ask vendors for task-level evaluation results, including which requirements failed, how errors are detected and what human approvals remain mandatory. Further published tests would be needed to show whether performance extends beyond the 11 selected tasks and whether simulated speed estimates translate into verified customer outcomes.

Amazon

procurement approval workflow software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI and Ironclad test?

They tested model performance on 11 legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. The tasks included creating nondisclosure agreements and building procurement approval workflows.

Does GPT-6 Astra’s 55% score mean it completed 55% of the tasks?

No. OpenAI’s figure is the average share of rubric criteria met across the tasks. It is not the percentage of tasks completed and does not by itself show that results are safe to use without review.

Did OpenAI measure customer time savings?

No measured customer savings were reported. OpenAI says the 19.2-minute estimate for Astra is a simulation based on assumed processing and generation speeds, covering the research tasks rather than Ironclad workflows generally.

What data did OpenAI say it used?

OpenAI says it generated synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It says it did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data.

Can businesses use these results to deploy agents without oversight?

The reported results do not support that conclusion. OpenAI’s average score leaves criteria unmet, and the source material says human oversight remains important. Businesses would need to review task-level failures and approval controls before relying on agents for consequential workflows.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Steam App 1905180 Climbing The Steam Charts

The Steam app 1905180 has surged in popularity, climbing to rank 17 with a peak of over 30,000 players, signaling a notable trend on the platform.

What OpenAI’s Cyber Access Extension Means For Civilian Defense In Ukraine

OpenAI says it is extending cyber-related access in Ukraine for civilian defense, but has not detailed eligibility, safeguards or the tools involved.

Macintosh Surges In Global Coverage

Macintosh’s media coverage has surged, with 24 mentions in recent reports, marking a notable increase in global attention.

Alienware Surges In Global Coverage

Search interest in Alienware has spiked, with 20 mentions this week, indicating increased media and public attention. The cause remains unconfirmed.