The Fine Print To Consider As OpenAI Trains Agents In Your Software
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Fine Print To Consider As OpenAI Trains Agents In Your Software on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training a frontier model in hosted copies of Ironclad’s contract-management software, using 11 legal, commercial and procurement tasks. GPT-6 Astra met an average 55% of task criteria, while its reported time savings were simulated estimates, not measured customer results.

OpenAI said on Oct. 6 that it trained its frontier model GPT-6 Astra in hosted copies of Ironclad’s contract-management software, testing whether an AI agent could carry out multi-step legal, commercial and procurement work. The results point to a possible new way of training agents on professional software, but Astra met an average of 55% of task criteria and the reported completion times were simulations, not measured customer savings.

OpenAI and Ironclad selected 11 tasks with staff who use the contract platform. Examples included setting up nondisclosure agreements, creating procurement approval processes and changing a reusable contract clause based on a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes for each task.

Each task was evaluated against 8 to 50 criteria, depending on its complexity. OpenAI reported that GPT-6 Astra met an average of 55% of those criteria, compared with 41.6% for GPT-5.6 Sol in a high-reasoning setting. An internal model used in Astra’s development reached 63.7%. Astra met about 94% of the criteria on one showcase task, but that example is not the overall result.

OpenAI reported estimated attempt times of 19.2 minutes for Astra and 37 minutes for GPT-5.6 Sol. The company’s footnote says those figures are simulated from assumed processing and generation speeds. They are not observed customer time savings and apply to the 11 research tasks, not to Ironclad workflows broadly.

At a glance
reportWhen: Published Oct. 6; further software-comp…
The developmentOpenAI published a report on Oct. 6 describing how it trained and evaluated GPT-6 Astra on workflows inside Ironclad’s contract-management software.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Partial Workflow Accuracy Matters

The reported 55% is the average share of criteria met, not the percentage of tasks completed successfully. That distinction matters in contract and procurement work, where missing one requirement can invalidate an otherwise polished result. For example, a procurement process may need Finance approval above a spending threshold, Security review for certain requests and Legal review for unusual terms. Missing one control can route a purchase incorrectly.

OpenAI’s report recognizes that an agent can lose track of a business rule during a task and says human oversight remains important. Ironclad CTO Sunita Verma likewise emphasized that agents must preserve “the controls teams rely on.” The test suggests progress on complex software tasks, but does not establish that agents are ready to handle consequential workflows without people checking their work.

For software vendors, the partnership model could help expose where agents struggle in real products and give vendors a role in shaping future capabilities. It may also change how customers use software: if agents increasingly act through a product, a vendor’s lasting value may depend less on its screens and more on its business rules, data, audit records and controls.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Ironclad Became a Test Environment

OpenAI’s post, titled “Advancing computer use with Ironclad,” describes training and evaluation inside a specialist software product rather than a general computer-use benchmark. Ironclad is a contract-management software company, not the name of a new agent framework. The tasks were drawn from legal, commercial and procurement workflows selected by people familiar with the product.

For practice, Ironclad supplied hosted copies of its product. OpenAI says it created synthetic training tasks from publicly filed contracts in the U.S. Securities and Exchange Commission’s EDGAR database, with personal information filtered out. The company says it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.

OpenAI is asking a small number of software companies to partner on similar research. It says prospective partners should bring a concrete task agents cannot reliably complete, people with detailed knowledge of the work, a secure test environment and data that can safely be used for research. The report presents this as a way to study agents in specialized software, not as evidence that the tested agent is ready for general deployment.

Amazon

AI workflow automation tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Test Does Not Establish

The results do not show how often Astra would complete a full workflow to a standard that a company could accept without review. The 55% figure averages rubric criteria, and the report’s headline metric does not by itself show which requirements were missed or how serious each miss was. A high score on one example does not settle performance across the other tasks.

It is also unclear whether the results would hold across different companies, software configurations, contract types or live customer environments. OpenAI’s time figures are simulated, and the test covered 11 selected research tasks. The material does not establish measured productivity gains, error rates in production or the level of human review needed for each workflow.

OpenAI says it excluded specified customer and internal data from training, but the report does not, in the supplied material, detail every security, retention or access arrangement for hosted testing. Companies considering agent use should seek task-level evaluation results and written information about data handling and controls rather than treating an aggregate score as a deployment guarantee.

Amazon

AI-powered contract review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Next Test Is Vendor Adoption

OpenAI says it is inviting a small number of software companies to work on tasks current agents cannot reliably finish. The next developments to watch are which vendors participate, what tasks they select and whether future reports include clearer task-by-task results, production testing and independently observed time savings.

For buyers, the immediate next step is to ask vendors how an agent’s work is evaluated and supervised before enabling it in systems that handle contracts, finances or customer records. Useful details include the specific criteria it failed, how approval controls are enforced, what actions require human sign-off and how activity is recorded. Until those answers are available, OpenAI’s Ironclad study is best read as a research demonstration with incomplete performance, not proof of hands-off automation.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI test with Ironclad?

OpenAI evaluated GPT-6 Astra on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management software.

Does the 55% score mean Astra completed 55% of the tasks?

No. It is the average share of rubric criteria met across the tasks, not the percentage of tasks completed. The figure also does not show, on its own, which requirements were missed.

Did Astra cut customer task times in half?

That has not been shown. OpenAI reported a simulated estimate of 19.2 minutes per attempt for Astra, compared with 37 minutes for GPT-5.6 Sol. The company says these were not measured customer time savings.

Did OpenAI use Ironclad customer contracts to train the model?

OpenAI says it used synthetic tasks based on publicly filed contracts in the SEC’s EDGAR database and did not use non-public Ironclad customer data, OpenAI customer data or OpenAI internal contracts.

Can companies rely on agents to run these workflows without review?

The report does not establish that. Astra met an average of 55% of task criteria, and OpenAI says human oversight remains important. Companies would need more task-specific evidence and clear controls before relying on agents for consequential work.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Exploring The Future Of AI Beyond Sentence Generation With Jev

TypeSafe’s Jev introduces a new AI model focusing on structured decisions for automation, moving beyond traditional text-based chatbots.

Singapore: Engineer the Transition

Singapore employs a calibrated, multi-instrument strategy to reskill workers and advance AI, emphasizing state capacity and continuous adaptation.

7 Best PC Routers for Prime Day Deals in 2026

Discover the best PC router deals for Prime Day 2026, including Wi-Fi 7, Wi-Fi 6, and budget options, with expert insights on features and suitability.

The Local-First Agentic Operator

Exploring how a single operator, empowered by agentic AI, now builds and manages diverse software portfolios without traditional organizational structures.