Skip to content
WattsUpNext
OpenAI logo displayed on a computer screen

AI

OpenAI's GPT-6 Astra scores 55% on Ironclad contract tasks

OpenAI trained its new model on contract work built with Ironclad. It beats the older model, but the speed gains are simulated and the tasks don't come from real customer deals.

Photo by Andrew Neel on Unsplash · Unsplash License · source

By WattsUpNext Desk · Edited by Juda B. Hur2 min read

OpenAI has started training its AI agents on Ironclad's contract software, and its newest model, GPT-6 Astra, got a little over half of a small contracting test right. If your week goes to contract drafts, approval chains and sign-offs in legal, sales or procurement, OpenAI wants an agent to take that work off your plate.

That also means the model missed on close to half of what it was graded on.

Ironcad's software helps companies draft, route and approve contracts, and OpenAI calls Astra its first frontier model trained on Ironclad tasks. On OpenAI's own research test, Astra averaged 55%, compared with roughly 42% for the older GPT-5.6 Sol. OpenAI also estimates it finishes each task in about half the time. Ironclad's chief technology officer, Sunita Verma, said: "Agents need to understand the full contracting lifecycle, including how business workflows connect while preserving the controls teams rely on".

Now the fine print. The time savings weren't timed in anyone's office. They're simulated estimates based on assumed processing speeds. The tasks came from public contracts in the SEC's EDGAR database with personal details stripped out, and OpenAI says it used no private Ironclad customer contracts and none of its own customer data. That's good for privacy. It also means the agent hasn't practiced on the messy, one-off deals real legal teams actually argue over.

The test is small, too. It covers 11 research tasks rather than the whole product, picked with help from Ironclad employees and people at OpenAI who use Ironclad, and nobody outside OpenAI has checked the results. Each task was scored against a rubric of 8 to 50 criteria, and the 55% is an average of those rubric scores, so the model earned partial credit rather than a simple pass or fail on each task. On a set that small, one or two tasks going differently could shift the 13 point gap noticeably. An unreleased internal OpenAI model did better on the same set, so the version you can use isn't the company's best on this test. Neither side has said anything about pricing or exclusivity.

OpenAI says Ironclad is the first of a small group of software makers it plans to work with, and other companies can apply. The result worth waiting for will come from a live pilot: how often a person had to step in and fix the agent's work.

Sources

  1. 1.OpenAI Says GPT-6 Astra Scored 55.0% on 11 Ironclad Tasks · FourWeekMBA
  2. 2.OpenAI partners with Ironclad to train and improve AI agents · Seeking Alpha via TradingView
  3. 3.Advancing computer use with Ironclad · OpenAI

Reported by the WattsUpNext desk from the sources linked below. Spot an error? Tell us at corrections@wattsupnext.com.

The WattsUpNext Brief

Get stories like this in your inbox.

One email each weekday, only the topics you choose. Real news, sourced, no fluff. Unsubscribe in one click.

What would you like to get?

Keep reading

Published Oct 8, 2026