Red Alert: OpenAI is poised to cross an AI safety redline.
Making models harder to monitor is not what we need

TL;DR
- OpenAI is exploring new techniques to make AI models less transparent in their 'thinking'.
- These techniques could make monitoring AI behavior difficult or impossible.
- Better monitoring is considered crucial for AI safety, potentially preventing incidents.
- Chain of Thought monitoring, though imperfect, is identified as a vital tool for understanding LLMs.
- Researchers departing OpenAI express strong agreement on the safety concerns.
The Information just broke the scoop thatOpenAI is playing around with a new technique, in which models will reveal less of their “thinking”, making them harder to monitor.

As Zack Korman and I argued here a few days ago, better monitoring is one of the things that might have prevented the Hugging Face incident, by OpenAI’s own admission:

The new techniques they are exploring may make such monitoring difficult or impossible.
§
Last year, an all-star cast wrote a fascinating paper that feels deeply relevant now, calledChain of Thought Monitorability: A New and Fragile Opportunity for AI Safety, I fully agree with the highlighted bit:

They are exactly right. CoT monitoring is imperfect (asSubbarao Kambhampatiand others have shown), but it is one of the best threads we have for monitoring the giant black boxes that we call LLM. It is a slender thread, but sacrificing it thread for (small?) performance gain feels like a dangerous game.
Earlier tonight Steven Adler, of Guidelight.ai andone of the many researchers to have departed from OpenAI’s safety teams, said this, echoing Nathan Calvin:
I fully, 100% agree.
P.S. Bonus those who prefer Terminator references: