Complementary remarks from Gary Marcus and Terence Tao on OpenAI’s giant math drop
The real news here isn’t the result; it’s what we were not told.

TL;DR
- OpenAI's math achievement report is vague and would not pass peer review.
- Key details such as the procedure, architecture, and failure rate are unknown.
- The generalizability of the AI's results outside of mathematics cannot be assessed.
- Online discussion is characterized by uncritical enthusiasm rather than scientific inquiry.
- The true significance of the AI system remains unclear.
OpenAI’s massive new math drop:
[This essay was written in extreme haste before a very long wifi-less flight; please forgive typos.]
Part One: My take
The real news here isn’t the result; it’s not what we were told.
1. AI once tried to be a science. Now we get stuff like the completely vague report from OpenAI below:

“Same procedure”? “Using an unreleased model”?
This would never pass peer review.
We don’t know what the procedure was.
We know nothing about the architecture. For esxample, were the proofs generated in one shot, and then verified by the symbolic system Lean? Was there an iterative process?)
We know nothing about the failure rate. We know nothing about the training/post training/data augmention.
2. As a result, we have zero idea of how generalizable the result is outside math.
3. A lot of the discussion on social media has been reduced to an ignorant cheering section that applauds without knowing what it is applauding or what it might mean— without ever asking basic scientific questions.
The new system could be a legitimate step toward AGI. Or it could just be a clever leveraging of Lean and synthetic data in a verifiable domain with no generality whatsoever.
From the initial report, we can tell almost nothing.
§
Part Two: Terence Tao’s take

