AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI’s Agent Training In Your Software: Ironclad Terms To Keep In Mind on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training GPT-6 Astra on 11 legal, commercial and procurement tasks in hosted copies of contract-management software company Ironclad’s product. Astra met an average 55% of task rubric criteria, while estimated completion times were simulated rather than measured customer savings. OpenAI is inviting a small number of software companies to propose similar training partnerships.

OpenAI said on October 6 that it trained its frontier model GPT-6 Astra on legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. The results, which averaged 55% of task criteria met, describe a research exercise rather than a ready-to-deploy agent, and OpenAI is inviting a small number of software companies to explore similar partnerships.

OpenAI and Ironclad selected 11 tasks, including creating nondisclosure agreements, setting up procurement approval processes and changing a reusable contract clause based on a requester’s jurisdiction. OpenAI estimates that an experienced user would take 30 to 40 minutes on each task. The tasks were scored against rubrics of 8 to 50 criteria, depending on complexity; the score measures criteria met, not the percentage of tasks completed successfully.

OpenAI reported that GPT-6 Astra met an average 55.0% of criteria, compared with 41.6% for GPT-5.6 Sol in a high-effort setting. Astra’s estimated time per attempt was 19.2 minutes, compared with 37.0 minutes for Sol. An internal OpenAI model used during Astra’s development reached 63.7%. On one highlighted task, Astra met about 94% of the criteria. These figures apply to the research tasks described, not all Ironclad workflows.

For training, Ironclad provided hosted copies of its product where models could practise. OpenAI says it created synthetic tasks from publicly filed contracts in the US Securities and Exchange Commission’s EDGAR database, filtered to remove personal information. The company said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data. OpenAI also says its time figures are simulated estimates based on assumed processing and generation speeds, not measured customer time savings.

At a glance
reportWhen: Published October 6; partnership outrea…
The developmentOpenAI published details of training a frontier model on professional workflows inside Ironclad’s contract-management software and invited other software firms to partner on similar work.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Software Vendors May Partner

The project points to a model of AI development in which software companies supply specialized workflows, test environments and expertise to train agents for professional tasks. That could improve how models handle rules and multi-step work in particular products. OpenAI is asking prospective partners for specific examples of tasks agents cannot reliably complete, people who know the workflows, secure testing environments and data that can be used for research.

For customers, the reported results do not establish that an agent can safely take over contract or procurement work. A workflow can meet many rubric criteria and still miss one mandatory approval. OpenAI’s own example describes a process that may need Finance approval above a spending threshold, Security review for certain requests and Legal review for nonstandard terms. Missing a single required control can make the process unsuitable, regardless of its average score. The findings support continued human review, especially for work with legal or financial consequences.

The arrangement may also affect software vendors’ role. An agent that operates a product could make its underlying business rules, records, audit trail and controls more important than its screens. OpenAI’s post argues that a full contracting platform remains essential, but the reported experiment does not show how such partnerships would affect vendors’ products or customer relationships.

Amazon

contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Inside the Ironclad Training Trial

The October 6 publication included material that received less attention than OpenAI’s release of 722 mathematics manuscripts on the same day. Some AI news trackers inferred from the word “Ironclad” that OpenAI had introduced a hardened agent framework. In this case, Ironclad is the name of a contract-management software company; the post describes training in its product, not a new framework by that name.

The research focused on whether a model could follow business rules, complete multi-step work in specialized software and check that its results met the original requirements. OpenAI describes GPT-6 Astra as the first frontier model it trained this way. The company’s account is the source for the task design, data practices and results; the published figures should be read as reported research measurements, not an independent evaluation of customer deployments.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Scores Do Not Show

The reported 55% is an average share of criteria met, not a task success rate. OpenAI’s description does not establish how often the model completed an entire workflow without missing a requirement, or whether the results would hold across a broader range of contracts, users and business rules. The article also does not provide a customer deployment study or measured time savings.

It remains unclear how Astra would perform in live operations, what safeguards or approval steps a deployed system would require, and how errors would be tracked and corrected. OpenAI says the research used no non-public Ironclad customer data, but the source material does not detail all security arrangements for future partnerships. The effects on participating vendors’ products and business models are also not established.

Amazon

automated contract creation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

OpenAI Seeks More Software Partners

OpenAI says it is looking to work with a small number of software companies on tasks that current agents cannot reliably complete. Prospective partners are expected to provide concrete failure cases, subject-matter expertise, secure environments for testing and research-usable data. The source material does not give a partner list, schedule or details of any further agreements.

For businesses considering agents in contracts, finance or customer-record systems, the next practical step is to ask vendors how performance is measured and which requirements failed in testing. Buyers should seek task-level results, evidence of controls and human review arrangements rather than relying on an average rubric score or simulated time estimate. Further partnership announcements or deployment evidence would show whether the research translates into dependable production use.

Amazon

AI-powered procurement software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Ironclad in this announcement?

Ironclad is a contract-management software company. OpenAI’s post describes training work in hosted copies of its product; it does not announce an agent framework called Ironclad.

What does Astra’s 55% score mean?

It is the average share of rubric criteria met across the research tasks, not the percentage of tasks completed. OpenAI says tasks had between 8 and 50 criteria, depending on complexity.

Did the study prove that AI agents save customers time?

No. OpenAI described the 19.2-minute Astra figure as a simulated estimate based on assumed processing and generation speeds. It was not a measurement of time saved by customers using the software.

Did OpenAI use private customer contracts to train Astra?

OpenAI said it used synthetic tasks based on publicly filed contracts from the SEC’s EDGAR database, filtered to remove personal information. It said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.

Can companies deploy agents for contract work now?

The reported results do not establish that Astra is ready to handle live contract workflows without oversight. OpenAI’s average score and the possibility of missed approvals point to the need for human checks and clear safeguards; deployment readiness remains unproven by this research account.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Grok Bot And X.ai: Pioneering The Next Wave Of Artificial Intelligence

xAI has revealed Grok Bot with minimal details, leaving its functions, availability, and purpose unclear, sparking industry speculation.

Immich 3.0

Immich has released version 3.0, introducing new features and improvements. This update aims to enhance user experience and security in photo management.

The Frameworks Can’t See the Thing That Matters: A Year of AI-Enabled Cyber Threats

A new analysis reveals AI is transforming cyberattacks, making even less skilled actors more dangerous and challenging existing threat assessment methods.

Exploring ByteDance Seed’s SeedRealtime: A Pioneering Full-duplex AI That Sees, Hears, And Responds

ByteDance Seed introduces SeedRealtime, a native audio-visual, full-duplex AI capable of seeing, listening, and speaking in real time, but details remain limited.