Your Software And OpenAI Agent Training: Take A Closer Look At Ironclad
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Your Software And OpenAI Agent Training: Take A Closer Look At Ironclad on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get smart everyday buys delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI’s October 6 post describes training GPT-6 Astra inside hosted copies of Ironclad’s contract-management software, using 11 legal, commercial and procurement tasks. Astra met an average 55% of task criteria, while OpenAI’s time estimates were simulated; the results do not establish customer-ready performance or measured productivity gains.

OpenAI said on October 6 that it trained its frontier model GPT-6 Astra on tasks inside hosted copies of Ironclad’s contract-management software, reporting an average score of 55% of task criteria met. The work tests whether models can follow business rules across specialized software workflows; it does not show that the system is ready to complete contract work without human review.

OpenAI and Ironclad selected 11 tasks spanning legal, commercial and procurement work. Examples included setting up nondisclosure agreements, creating procurement approval processes and updating a reusable contract clause to reflect a requester’s selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes to complete each task. The tasks were chosen by Ironclad staff and OpenAI employees who use the product.

OpenAI scored each task against a rubric of 8 to 50 criteria, depending on its complexity. It reported that GPT-6 Astra, using its maximum setting, met an average of 55% of criteria, compared with 41.6% for GPT-5.6 Sol at the high setting. An internal OpenAI model used during Astra’s development reached 63.7%. On one showcased task, Astra met about 94% of criteria; that example is not the average result.

The reported time figures are estimates, not customer measurements. OpenAI put Astra’s average time per attempt at 19.2 minutes, compared with 37 minutes for GPT-5.6 Sol, but said those figures were simulated using assumed processing and generation speeds. They apply to the 11 research tasks, not Ironclad workflows generally, and do not establish actual time saved in a business setting.

At a glance
reportWhen: Published October 6; the next partnersh…
The developmentOpenAI published results from training a frontier model on workflows inside Ironclad’s contract-management software and invited other software companies to explore similar research partnerships.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Contract Work Needs More Than a Score

The findings address a practical question for companies that handle contracts, purchasing or other controlled workflows in specialized software: can an AI agent follow a chain of rules and verify that it has completed the work correctly? OpenAI describes this as a training approach for models to understand business rules, carry out multi-step tasks and check results against requirements. The Ironclad test offers an example of how that work might be evaluated inside a real product environment.

But the headline percentage needs careful reading. 55% means the average share of rubric criteria met, not that Astra completed 55% of tasks successfully. In a procurement workflow, for example, missing a required Finance, Security or Legal approval could make the result unusable even if other steps were correct. OpenAI’s own post says losing track of a business rule limits what a software company can confidently ask an agent to do, and says human oversight remains necessary.

For software vendors, the project also points to a possible change in the relationship between AI and business applications. Training agents on a vendor’s workflows could make its product more useful, while exposing where agents fail. If agents become the way users interact with software, a vendor’s value may depend increasingly on its underlying rules, data structures, controls and records rather than its screens. That is an implication of the approach, not a demonstrated outcome of this test.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Inside the Ironclad Training Test

OpenAI says Ironclad supplied hosted copies of its product for model practice. The research tasks were generated synthetically from publicly filed contracts in the U.S. Securities and Exchange Commission’s EDGAR database. OpenAI says it filtered that material to remove personal information and used no OpenAI customer data, no OpenAI internal contracts and no non-public Ironclad customer data.

The October 6 post also describes a broader research plan: OpenAI is inviting a small number of software companies to work on tasks current agents cannot reliably complete. It asks prospective partners to provide a concrete example of a failure, people with deep knowledge of the work, a secure testing environment and data that can safely be used for research. The post frames the Ironclad work as a way to study agents in specialized software rather than as a product launch or a general performance guarantee.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Reported Scores Leave Open

The post does not establish how Astra would perform across Ironclad’s full range of customer work, on tasks outside the 11 selected examples, or under live business conditions. The average rubric score also does not identify, in the figures provided, which requirements models missed most often or how consequential those misses were. A high result on one showcased task cannot answer those broader questions.

It is also unclear whether the simulated time estimates would translate into faster completed work once a person checks every requirement and corrects errors. OpenAI’s figures are not measured customer outcomes, and the source material does not provide a deployment study or a comparison of end-to-end accuracy against experienced users. The data-handling description is OpenAI’s account of its research process; independent verification is not provided in the material.

Amazon

Ironclad contract management system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Next Step for Software Partners

OpenAI says it plans to work with a small number of software companies on difficult tasks that agents cannot yet complete reliably. The proposed partners would bring specific failure cases, subject-matter experts, secure test environments and research-safe data. The source material does not name additional partners or give a timetable for the next collaborations.

For companies considering such work, the immediate questions are how tasks will be graded, which requirements were missed, what data may be used, and where a human must review or approve an agent’s actions. Any future results will need to distinguish rubric performance from successful completion and simulated timing from measured productivity. OpenAI has not provided a date for broader availability or said that the Ironclad research has produced a customer-ready agent.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Ironclad in this report?

Ironclad is a contract-management software company, not the name of a new OpenAI agent framework. OpenAI’s post describes research conducted using hosted copies of Ironclad’s product.

What does GPT-6 Astra’s 55% score mean?

It is the average share of rubric criteria met across the research tasks. It is not the percentage of tasks completed successfully, and it does not show that Astra can handle contracts without review.

Did Astra cut contract-work time in customer deployments?

The reported 19.2-minute figure is simulated, based on assumed processing and generation speeds. OpenAI said it was not measured customer time savings and covered the 11 research tasks only.

What data did OpenAI say it used?

OpenAI said it created synthetic tasks from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. It said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.

Is the Ironclad agent ready for unsupervised use?

The reported results do not establish that. OpenAI said human oversight still matters, and Astra met an average of 55% of the criteria in the test. The post does not announce general availability or unsupervised deployment.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

IEP Negotiation Simulator FAQs Parents Ask Most

A proposed IEP simulator would help parents review documents, prepare requests and rehearse meetings. Its impact remains untested.

Victims Claim Grok Used Their Videos And Images To Advance Deepfake AI Systems

Survivors allege xAI’s Grok trained on their images and videos without consent, raising legal and ethical concerns about data sourcing and deepfake capabilities.

The clause. How a contractual definition of AGI met the capital built on top of it.

An analysis of how a key contractual clause defining AGI was ultimately renegotiated, revealing the tension between governance ideals and capital needs.

Michigan Court Orders Kalshi to Stop Sports Event Contracts

A Michigan court has ordered Kalshi to stop offering contracts linked to sports events amid legal challenges, raising questions about sports betting and event trading.