The same model scored 62.7% and 99.9% on the same test
OpenAI's GPT-6 Astra posted two very different numbers on the same benchmark in the same week. The gap was not the model. It was the harness, and that distinction matters to anyone buying software.
EK Ra Sunya Team
EK Ra Sunya Inc.
Figures in this article are as published by ARC Prize on 3 September 2026. This was a moving story: OpenAI's headline number changed between the embargoed press draft and the live post. Check the primary source before quoting them.
On 3 September, OpenAI launched GPT-6 Astra and reported 99.9% on ARC-AGI-3, a benchmark built to be hard for machines and easy for people. Its predecessor, GPT-5.6 Sol, had managed 7.8%. A jump like that reads as a different species of software.
ARC Prize, who build the benchmark, ran the same model themselves. They got 62.7%.
Both numbers are real. Neither party is lying. The gap is the interesting part, and it is the most useful thing that happened in AI last week for anyone who buys software rather than builds it.
The harness is not the model
A harness is the scaffolding around a model during a test: what tools it can reach, what it remembers between requests, how its context gets managed when a conversation runs long.
ARC Prize's harness gives every model the same minimal interface and lets the model decide what to keep. The 99.9% came through a Provider Adapter, which preserves the model's reasoning state between requests and compacts longer conversations automatically. Two settings, changed for defensible reasons, and ARC Prize published both numbers side by side on the day.
Here is the number that should stop you. Under the Provider Adapter with reasoning effort set to none, Astra scored 96.7%. Under ARC Prize's standard harness at that same setting, it scored 35.2%. Same model, same reasoning budget, sixty-one points apart.
The scaffolding was worth more than the thinking.
What the honest comparison looks like
A 99.9%-versus-7.8% framing compares a number from one harness against a number from another. The comparison that holds is 62.7% against 7.8%, both measured on the standard harness. Claude Opus 5 scored 30.2% on that same harness, which is the other number worth having when someone tells you the field just jumped.
That is still a large improvement. It is a real result. It is not the result that got posted, and the difference between those two sentences is where procurement decisions go wrong.
ARC Prize published both numbers on launch day and said plainly they are not claiming AGI: "while we believe Astra represents meaningful progress towards generalization, we are not claiming that it is AGI." That is a benchmark maintainer behaving well under pressure, and it is worth noticing because most do not.
Why this lands on your desk
You are unlikely to buy a frontier model. You are very likely to buy software whose vendor quotes a benchmark: a CRM claiming a lead-scoring accuracy, an OCR tool claiming an extraction rate, an ERP claiming an implementation timeline.
Every one of those numbers was produced by a harness. On clean data. On documents the vendor chose. In a configuration nobody wrote down.
We see this on both sides. When we run a proof of concept for a client, our numbers come from our setup, on data we prepared, with an engineer watching. That is not dishonest, but it is not what week three of production looks like either. The gap is not deception. It is the difference between a lab and a Tuesday.
Four questions that survive contact with a sales deck
Ask who ran the test. A vendor's own number is a starting point, not a finding.
Ask what the setup was, in enough detail to repeat it. If the configuration cannot be described, the number cannot be reproduced, and a number nobody can reproduce is marketing.
Ask what the same test gives on your data. Not a demo dataset. Yours, with its missing fields and its inconsistent dates and the three years of history nobody cleaned.
Ask what it costs to run that way. ARC Prize's 62.7% run cost $26,098. The 99.9% run through the Provider Adapter cost $18,817. The better-scoring configuration was also the cheaper one, which is worth sitting with: the scaffolding was not buying accuracy at the price of efficiency, it was buying both.
None of this requires you to understand the model. It requires you to ask how the number was made, which is a question any buyer can ask and most do not.
What we do about it
When we build for a client, the number that matters is the one measured on their data in their environment, not the one from our demo. We would rather show a smaller figure we can defend in month six than a larger one that came from a setup we controlled.
That costs us the occasional deal against a vendor quoting a lab number. We think it is the right trade, because the alternative is a system that tests beautifully and disappoints in production, and that is a relationship you only get to damage once.
If you are evaluating software and the benchmark in the deck looks too clean, the useful question is not whether the vendor is honest. It is: who built the harness, and what happens when you swap in your data.
We build web and mobile applications, cloud infrastructure and business systems from Lalitpur, for clients in Nepal, Australia and elsewhere. If you want a second opinion on a vendor claim before you sign, get in touch.
Have a project in mind?
Let us help you build something scalable, fast, and built to last.
Start a conversationKeep reading
Bulk SMS and Messaging APIs: A Practical Guide for Nepali Businesses
OTPs, alerts, and campaigns at scale. What to look for in an SMS platform, and how we built EK SMS for reliable delivery.
SEO & GrowthGEO vs SEO in 2026: How to Rank in AI Answers and on Google
Generative Engine Optimization (GEO) is changing how businesses get found. Here is how we make sites that rank on Google and get cited by AI assistants.

