News

GPT-6 Astra scores 62.7% or 99.9% on ARC-AGI-3

ARC Prize's 3 September analysis says GPT-6 Astra scores 62.7% on ARC-AGI-3 Semi-Private under the Standard harness for about $26,098, and 99.9% under the Provider Adapter harness for about $18,817. Saturating the bench is not an AGI claim.

GPT-6 Astra scores 62.7% or 99.9% on ARC-AGI-3

The ARC Prize Foundation published an independent benchmark writeup on 3 September 2026, OpenAI's GPT-6 Astra on ARC-AGI-3. It reports two scores for the same model on ARC-AGI-3 Semi-Private: 62.7% under the Standard harness and 99.9% under the Provider Adapter harness.

This is a foundation blog with a results table, not a peer-reviewed paper and not a regulator finding. ARC Prize says saturating ARC-AGI-3 is not proof of AGI, and that it will publish both harness results going forward.

The number that does not travel with most launch headlines is the provider-neutral run. Astra (max) scored 62.7% on the Standard harness, where the model carries only the notes it chooses to keep, at a reported cost of about $26,098.

The 99.9% figure is Astra (high) on the Provider Adapter harness, which preserves opaque reasoning state between requests and uses compaction, at about $18,817. Across 167 paired solves, Provider Adapter runs were about 3.66x faster by elapsed time and used 49% fewer total tokens. Under that harness, Astra (max) used fewer actions than the human baseline on 96.0% of levels and 51.7% fewer actions per level on average.

OpenAI's own GPT-6 Astra launch post is the company announcement from the same day. Fortune (Emily Forlini) says an embargo draft listed ARC-AGI-3 at 98.6%, the live post later showed 99.99%, and other published metrics moved after launch, including a hallucination rate that went from 4.2% to 2% and back to 4.2%. OpenAI blamed a CMS bug and then an outage for a pull-and-republish, and told Fortune that harness, reasoning level, and other factors inform evals.

ARC Prize also notes Astra built compact algebraic and symbolic world models, plus domain-specific shorthand in its notes. That is a replay observation, not a claim the model generalized past the bench's closed-ended games. Other recent lab writeups, from Claude's Lean formalization of Fermat to OpenAI agents writing on a German wiki, are the same class of document: a research or incident post, not a verdict. The company is also in a commercial and legal week, from ChatGPT Ads at a $1 billion run rate to a Seattle Times and Newsday copyright complaint.

If you buy models, write eval memos, or brief a board on Astra this week, put 62.7% at $26,098 next to 99.9% at $18,817 and label the harness on each line. Treat the Provider Adapter score as the provider-tuned number, not the apples-to-apples one, and do not file this as an AGI declaration.

Subscribe to Techpresso

Free daily newsletter, read in 5 minutes.

Subscribe free