Pause OpenAI, now

Quite simply, they can no longer be trusted.

Pause OpenAI, now

TL;DR

  • OpenAI's new Astra model reduces Chain of Thought (CoT) monitorability, a key tool for controlling generative AI.
  • AI safety experts are concerned that OpenAI is trading safety for performance gains.
  • OpenAI's management is questioned for releasing Astra despite data showing compromised monitorability, especially on 'destructive actions'.
  • A recently departed OpenAI employee's essay suggests accepting rogue AI as inevitable, which the author interprets as a request to allow OpenAI's AI to 'run amok'.
  • Another incident at OpenAI was allegedly kept quiet for weeks, further eroding trust.
  • The author believes OpenAI's internal security practices are lacking and they are not being truthful with the public.
  • A call is made for Congress to investigate OpenAI for potential sanctions or a pause.
  • The government, specifically the White House, is criticized for not adequately vetting OpenAI's models, as evidenced by the approval of Astra despite reduced monitorability.

I have often counseled calm where others might counsel panic.

I told you that the Hugging Face incidentcould likely have been prevented had best cybersecurity practices been followed. (And I stand by that.)

I told you (and most people still seem unaware) that that the OpenAI Hugging Face incident was part of a training exercise, with some internal guardrails shut down, so it was not quite as bad as it seemed1

I told you that Astra probably wasn’t AGI.2

And I stand by all of that.

But I am freaked out.

What I am freaked about is not imminent AGI.

It’s OpenAI.

I simply don’t believe that they are trustworthy enough or responsible enough to be good stewards of the technology that they are developing.3As a company, they simply don’t have good judgment.

§

Here are four considerations.

  1. Sam Altman cannot be trusted.I have been writing about that for a long time. Ronan Farrow’s reportingbacks that up. So does the just-dropped bombshell below that I am about to get to.

  2. The just-released Astra reduces Chain fof Thought (CoT) monitorability, one of the few (not especially reliable, but better than nothing) tools we have for keeping generative AI from running wild. The AI safety community is up in arms about this—with good reason. There are tons of posts like this now, all quite right:

    The decision to release Astra is a clear example of the willingness of OpenAI management to trade off safety in exchange for relatively modest gains in performance. Thered alert that I sounded a couple days agowas on target. They reallyareplaying around with new techniques that reduce monitorability. Andtheir own data shows that monitorability is in fact compromised to some degreein the newly released Astra, particularly on “destructive actions.” They released it anyway. That speaks volumes.

  3. Something I read last night, and that only fully clicked into place this morning (see fact 4 below) terrifies me. A prominent recently departed employee (who presumably still owns significant stock, and who has repeatedly struck me as an advocate of OpenAI since he left) wrotean essay on Xbasically asking people to simply accept that rogue AI is here to stay.

    Full esssayhere

    In essence, I read this essay as requesting a hall pass to let their company’s AI run amok. How about if instead we pause now — before we get to “there is going to be” rather than heading full steam towards danger?

  4. What has really chilled me this morning—and what moved to write this call to pause OpenAI—is not just Achiam’s chilling note but the revelation that there has beenyet another incident.

    Which OpenAI apparentlytried to keep quiet.Forweeks.

If OpenAI is going to keep this stuff under wraps, Congress and/or the White House needs to shut them down, at least for a while.

This company simply cannot be trusted. Their software is becoming ever more dangerous; their internal security practices leave a lot to be desired; they aren’t being straight with the public; and their proxies are preparing us to swallow the damage that they are now anticipating.

If there was ever a case for pausing a company for the public good, it would be now.

A good model might bereceivership, in which a company, typically close to bankruptcy, is put under the control of an outsider until such time as its core problems are remedied. I personally would not trust OpenAI unless and until Altman and his sidekick, Greg Brockman (whose questionable character was on display in the Musk trial) were replaced.

§

The problem here, though, isn’t just OpenAI. It’s systemic.

For starters, the government has not been doing its bit. Whatever the White House screening policy is, it let the new model—arguably much more dangerous than Mythos given the decrease in monitorability– fly. (Brockman reported that the model was vetted by the White House, and they were given a green light).

That suggests to me that the White House isn’t even looking at monitorability. Of course, I have no idea what the White House is actually looking at — and nor does anyone else. And that lack of transparency is in itself is amajorproblem (about which I will write more very soon).

As Dave Troy put it while I was drafting this, there is something deeply wrong here with the system as a whole:

§

How Trump handles OpenAI may end up defining his legacy. I estimate the probability of a major cyber incident attributable to OpenAI in the next 12 months to be very high, certainly over 50%. And Trump, if he doesn’t intervene, may share some of the blame.

In the meantime, I call upon Congress to investigate OpenAI, with an eye to whether the company might need to be sanctioned—or even paused— now.

Please share this call to Pause OpenAI

Share

Subscribe now

1

FromOpenAI’s own report on the Hugging Face incident“The incident occurred during routine testing”, and “At the time of the incident, OpenAI estimated maximal cyber capabilities by running this evaluation without the production classifiers intended to prevent models from pursuing high-risk cyber activity”. Had those classifiers been turned on, the incident might well not have happened.

2

Data from Epoch AI tends to support this conclusion. Astra is genuine improvement but not even statistically off of trend; if it really were AGI I think we would expect to reflect a sharper departure from previous models.

3

Altman told Alex Heath the other day that “Altman wants OpenAI to be seen as “the most responsible company ... good stewards of technology”. As made plain in this essay, he’s talking the talk, but not walking the talk. We do desperately need good stewards. He’s right about that. Unfortunately Altman himself is manifestly not suited to that particular job.