Hot take on GPT-6 Astra

An impressive system that can (to some unknown extent) build symbolic world models

Hot take on GPT-6 Astra

TL;DR

  • OpenAI's Astra achieves SOTA on ARC-AGI, scoring 63% on ARC-AGI-3 and surpassing human performance on 96% of its levels.
  • The system builds precise symbolic models of novel environments.
  • The robustness of Astra's symbolic modeling capability is the key unknown.
  • Success on ARC-AGI is not proof of AGI, and open-ended real-world tasks may present significant challenges.
  • Lack of transparency regarding how the system works raises concerns about its capabilities, limitations, and potential risks.
  • The system appears less monitorable than prior systems, which is a safety concern.
  • The author awaits further information and detailed examination of Astra's limitations.
@OpenAIachieves SOTA on ARC-AGI:\n\n- Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness\n- It surpasses human performance on 96% of ARC-AGI-3 levels\n- It builds the most precise symbolic model of novel environments we've seen\n\nOur analysis: ","username":"arcprize","name":"ARC Prize","profile_image_url":"https://pbs.substack.com/profile_images/1800565057493078016/XigYRdTo_normal.jpg","date":"2026-09-03T19:39:37.000Z","photos":[{"img_url":"https://pbs.substack.com/media/HRUN5czbMAASQkl.jpg","link_url":"https://t.co/GX77KsRNer"}],"quoted_tweet":{},"reply_count":51,"retweet_count":194,"like_count":1808,"impression_count":293356,"expanded_url":null,"video_url":null,"video_preview_media_key":null,"belowTheFold":false}" data-component-name="Twitter2ToDOM">

Hot take on OpenAI GPT-6 Astra*1, with a challenge to Greg Brockman’s claims about it being AGI toward the end:

  • Looks to be pretty impressive. Multiple reports suggest it is a genuine advance.

  • As someone who has campaigned for nearly a decade for (neuro)symbolic world models, often to exceptional hostility, it isextraordinarilyvindicating to see that a product from OpenAI explicitly creates and manipulate symbolic world models in the course of some of its most impressive computations.

  • What we don’t know is how robust that capability is. That is THE key question.

  • Success on ARC-AGI is great and impressive, but not —despite the name of the task—proof of AGI; I suspect we will see loads of problems with open-ended real world tasks. As with other recent models I would suspect best performance in verifiable domains.

  • And as a scientist, it’s disappointing that we don’t (yet?) know much about how the system actually works.

  • Without a clearer sense of what’s under the hood, I feel less confident about both what it can and can’t do, and what new risks we may encounter. I doubt the world is ready.

  • As ever, enthusiasts got an advance look; skeptics did not. That’s a sound marketing strategy, but it often turns out to be misleading. What we have often seen is initial enthusiasm that gets tempered over time. I suspect we will see that here as well.

  • The new system appears to belessmonitorablethan prior systems, which is not great from a safety perspective. One really doesn’t want more capability in conjunction with less monitorability. But also morealignable, not sure why.

  • Would be great to see whether Astra can make progress on any of the ten tasks that Miles Brundage and I bet on at the end of 2024. (No AI to date has succeeded on any, AFAIK.)

Subscribe now

1

This hot take is VERY tentative, pending more information about how the systems works and more detailed examination of what its limitations are.

Share