🔍 Read the full analysis: OpenAI’s Agent Training In Your Software: Ironclad Terms To Keep In Mind on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
OpenAI described training GPT-6 Astra on 11 legal, commercial and procurement tasks in hosted copies of contract-management software company Ironclad’s product. Astra met an average 55% of task rubric criteria, while estimated completion times were simulated rather than measured customer savings. OpenAI is inviting a small number of software companies to propose similar training partnerships.
OpenAI said on October 6 that it trained its frontier model GPT-6 Astra on legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. The results, which averaged 55% of task criteria met, describe a research exercise rather than a ready-to-deploy agent, and OpenAI is inviting a small number of software companies to explore similar partnerships.
OpenAI and Ironclad selected 11 tasks, including creating nondisclosure agreements, setting up procurement approval processes and changing a reusable contract clause based on a requester’s jurisdiction. OpenAI estimates that an experienced user would take 30 to 40 minutes on each task. The tasks were scored against rubrics of 8 to 50 criteria, depending on complexity; the score measures criteria met, not the percentage of tasks completed successfully.
OpenAI reported that GPT-6 Astra met an average 55.0% of criteria, compared with 41.6% for GPT-5.6 Sol in a high-effort setting. Astra’s estimated time per attempt was 19.2 minutes, compared with 37.0 minutes for Sol. An internal OpenAI model used during Astra’s development reached 63.7%. On one highlighted task, Astra met about 94% of the criteria. These figures apply to the research tasks described, not all Ironclad workflows.
For training, Ironclad provided hosted copies of its product where models could practise. OpenAI says it created synthetic tasks from publicly filed contracts in the US Securities and Exchange Commission’s EDGAR database, filtered to remove personal information. The company said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data. OpenAI also says its time figures are simulated estimates based on assumed processing and generation speeds, not measured customer time savings.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Software Vendors May Partner
The project points to a model of AI development in which software companies supply specialized workflows, test environments and expertise to train agents for professional tasks. That could improve how models handle rules and multi-step work in particular products. OpenAI is asking prospective partners for specific examples of tasks agents cannot reliably complete, people who know the workflows, secure testing environments and data that can be used for research.
For customers, the reported results do not establish that an agent can safely take over contract or procurement work. A workflow can meet many rubric criteria and still miss one mandatory approval. OpenAI’s own example describes a process that may need Finance approval above a spending threshold, Security review for certain requests and Legal review for nonstandard terms. Missing a single required control can make the process unsuitable, regardless of its average score. The findings support continued human review, especially for work with legal or financial consequences.
The arrangement may also affect software vendors’ role. An agent that operates a product could make its underlying business rules, records, audit trail and controls more important than its screens. OpenAI’s post argues that a full contracting platform remains essential, but the reported experiment does not show how such partnerships would affect vendors’ products or customer relationships.
As an affiliate, we earn on qualifying purchases.
Inside the Ironclad Training Trial
The October 6 publication included material that received less attention than OpenAI’s release of 722 mathematics manuscripts on the same day. Some AI news trackers inferred from the word “Ironclad” that OpenAI had introduced a hardened agent framework. In this case, Ironclad is the name of a contract-management software company; the post describes training in its product, not a new framework by that name.
The research focused on whether a model could follow business rules, complete multi-step work in specialized software and check that its results met the original requirements. OpenAI describes GPT-6 Astra as the first frontier model it trained this way. The company’s account is the source for the task design, data practices and results; the published figures should be read as reported research measurements, not an independent evaluation of customer deployments.
As an affiliate, we earn on qualifying purchases.
What the Scores Do Not Show
The reported 55% is an average share of criteria met, not a task success rate. OpenAI’s description does not establish how often the model completed an entire workflow without missing a requirement, or whether the results would hold across a broader range of contracts, users and business rules. The article also does not provide a customer deployment study or measured time savings.
It remains unclear how Astra would perform in live operations, what safeguards or approval steps a deployed system would require, and how errors would be tracked and corrected. OpenAI says the research used no non-public Ironclad customer data, but the source material does not detail all security arrangements for future partnerships. The effects on participating vendors’ products and business models are also not established.
automated contract creation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
OpenAI Seeks More Software Partners
OpenAI says it is looking to work with a small number of software companies on tasks that current agents cannot reliably complete. Prospective partners are expected to provide concrete failure cases, subject-matter expertise, secure environments for testing and research-usable data. The source material does not give a partner list, schedule or details of any further agreements.
For businesses considering agents in contracts, finance or customer-record systems, the next practical step is to ask vendors how performance is measured and which requirements failed in testing. Buyers should seek task-level results, evidence of controls and human review arrangements rather than relying on an average rubric score or simulated time estimate. Further partnership announcements or deployment evidence would show whether the research translates into dependable production use.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Ironclad in this announcement?
Ironclad is a contract-management software company. OpenAI’s post describes training work in hosted copies of its product; it does not announce an agent framework called Ironclad.
What does Astra’s 55% score mean?
It is the average share of rubric criteria met across the research tasks, not the percentage of tasks completed. OpenAI says tasks had between 8 and 50 criteria, depending on complexity.
Did the study prove that AI agents save customers time?
No. OpenAI described the 19.2-minute Astra figure as a simulated estimate based on assumed processing and generation speeds. It was not a measurement of time saved by customers using the software.
Did OpenAI use private customer contracts to train Astra?
OpenAI said it used synthetic tasks based on publicly filed contracts from the SEC’s EDGAR database, filtered to remove personal information. It said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.
Can companies deploy agents for contract work now?
The reported results do not establish that Astra is ready to handle live contract workflows without oversight. OpenAI’s average score and the possibility of missed approvals point to the need for human checks and clear safeguards; deployment readiness remains unproven by this research account.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
