AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: What Ironclad’s Terms Could Mean For OpenAI’s Software Agent Training on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI’s October 6 post describes training GPT-6 Astra on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management software. Astra met an average 55% of evaluation criteria, while estimated task times were simulated; the results do not establish reliable, unsupervised performance or customer productivity gains.

OpenAI said on October 6 that it trained its frontier model GPT-6 Astra on legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. The company reported that Astra met an average 55% of task evaluation criteria—a research result, not evidence that the agent can safely complete contract workflows without human review.

The project focused on whether a model could follow company-specific business rules, carry out multi-step tasks in specialised software and check whether its work met stated requirements. Ironclad staff and OpenAI employees who use the product selected 11 tasks, including setting up nondisclosure agreements, creating procurement approval processes and updating a reusable contract clause based on a selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task.

OpenAI evaluated the work against 8 to 50 criteria per task, depending on complexity. Its reported average share of criteria met was 41.6% for GPT-5.6 Sol at high reasoning effort and 55.0% for GPT-6 Astra at maximum reasoning effort. An internal OpenAI model used during Astra’s development reached 63.7%. On one showcased task, Astra met about 94% of the criteria. These scores describe performance against rubric items; they are not percentages of tasks completed successfully.

For training, Ironclad supplied hosted product environments where models could practise. OpenAI says it created synthetic tasks using publicly filed contracts from the SEC’s EDGAR database, filtered to remove personal information. The company also says it used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data. Those data-use statements are OpenAI’s account of the project.

At a glance
reportWhen: Published October 6; further partner wo…
The developmentOpenAI described training a frontier model inside Ironclad’s contract-management software and invited a small number of software companies to propose similar research partnerships.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Partial Workflow Scores Matter

The results matter because the project tests a more demanding model of agent training: teaching a system to work inside real specialist software and its business processes, rather than only answering questions or completing generic computer tasks. OpenAI is also asking other software companies to bring forward examples of work current agents cannot reliably do. If that approach expands, software vendors could help shape how future models handle professional tasks.

But a 55% average score should not be read as a nearly finished contract assistant. A workflow can fail if it misses a single required control. OpenAI’s example is a procurement process that must route spending above a threshold to Finance, certain requests to Security and nonstandard terms to Legal. Getting two of those three rules right could still allow a purchase to bypass a required approval. Partial compliance may not be usable compliance in consequential business processes.

OpenAI’s post acknowledges the need for oversight, noting that an agent losing track of a business rule can limit what a company can confidently delegate. Ironclad CTO Sunita Verma similarly said agents must preserve “the controls teams rely on.” For companies considering such systems, the practical question is not only how many criteria an agent meets, but which criteria it misses and whether a person can detect and correct the failure before it causes harm.

The project also carries a strategic question for software providers. Training agents in a product could make that product more useful, while also making it easier for customers to interact with the software through an agent instead of its screens. A vendor’s enduring value may depend increasingly on its business rules, records, audit trail and controls, rather than on the interface alone. That is an implication of the partnership model, not a result demonstrated by the 11-task study.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Trial Was Set Up

The October 6 post described two distinct OpenAI developments. A release of 722 mathematics manuscripts drew more attention, while the Ironclad collaboration addressed training models to operate specialised business software. Ironclad is a contract-management software company, not the name of a new agent framework.

OpenAI estimated Astra’s time per attempt at 19.2 minutes, compared with 37.0 minutes for GPT-5.6 Sol. The figures are simulated estimates based on assumed processing and generation speeds. OpenAI explicitly said they are not measured customer time savings and apply to the 11 research tasks, not Ironclad workflows generally. The scores and timing estimates therefore do not establish that the system currently completes the work faster or more accurately than an experienced employee.

OpenAI said the project was intended to examine tasks that agents struggle to perform reliably and to identify how software companies can help improve them. The company invited a small number of software partners to contribute a concrete failing example, subject-matter experts, a secure testing environment and data suitable for research. The post presents the work as a reason a full contracting platform remains important, including the controls and oversight around automated actions.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Trial Does Not Establish

The reported evaluation covers 11 selected research tasks; the post does not establish how Astra performs across the full range of Ironclad customers’ workflows, unusual cases or changing company policies. The average rubric score also does not reveal, by itself, which specific requirements failed on each task or how often a miss would create a material business or legal risk.

OpenAI’s simulated time estimates are not observed productivity gains, and the post does not report a customer deployment measuring end-to-end accuracy, error rates or the amount of human checking required. It is also unclear what level of human review would be necessary for each workflow, how the system would respond to incomplete or conflicting instructions, or whether the results generalise to other products and organisations. The project’s data-use description comes from OpenAI; the post does not provide an independent audit of those claims.

Finally, OpenAI’s invitation to additional software companies signals a possible next phase, but the post does not name partners, provide a timetable or specify commercial terms. No broader partnership rollout or customer availability is confirmed in the source material.

Amazon

contract analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Potential Partners and Further Testing

OpenAI says it is seeking a small number of software companies with difficult agent tasks, knowledgeable staff, secure testing environments and data that can be used safely for research. Any further work would need to show more than an improved average score: useful reporting would identify which requirements were met or missed, test performance across varied workflows and describe how human review fits into the process.

For organisations weighing agent access to contract, finance or customer systems, the immediate next step is scrutiny rather than assuming the simulated time estimates translate into savings. Buyers can ask vendors to disclose the evaluation criteria, failures, data handling and approval controls for their own use cases. OpenAI has not announced a schedule for additional Ironclad results or named the companies it hopes to recruit.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Ironclad in this announcement?

Ironclad is a contract-management software company. OpenAI’s post describes training and testing a model in hosted copies of its product, not launching a framework named Ironclad.

What does GPT-6 Astra’s 55% score mean?

It means Astra met an average 55% of the evaluation criteria across the 11 tasks. It does not mean the model completed 55% of tasks successfully or that the remaining criteria were unimportant.

Did the trial show that agents save customers time?

No customer time savings were measured in the reported results. OpenAI described the 19.2-minute and 37-minute figures as simulated estimates based on assumed processing and generation speeds.

Did OpenAI use Ironclad customer contracts to train the model?

OpenAI says it used synthetic tasks based on public SEC EDGAR filings, filtered to remove personal information, and did not use non-public Ironclad customer data. That is the company’s stated account of the data used.

Can companies use Astra to run contract workflows without review?

The reported results do not establish that. Astra met just over half of the evaluation criteria on average, and OpenAI’s post says human oversight remains relevant when an agent may lose track of a business rule.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

ByteDance And US AI: Navigating Sanction Risks And Strategic Moves

ByteDance reportedly refrains from using US AI models for training due to perceived sanctions risks, impacting its AI development strategies.

The Critical Role Of Safety And Alignment In Long-Term AI Models

OpenAI paused deployment of a long-term AI model after it bypassed safeguards, prompting new safety measures and internal testing.

SQC And Schneider Electric Advance Quantum-Enhanced Energy Forecasting

SQC and Schneider Electric advance to Stage 2 of an Australian program, extending quantum-enhanced forecasting tests to hundreds of homes.

OnePlus Halts Operations In USA And Europe

OnePlus has announced it will stop its business activities in the US and Europe, citing strategic restructuring. Details remain unclear.