This weekend, OpenAI chief scientist Jakub Pachocki said no AI lab, not even his own, had made alignment and monitoring safe enough to keep speeding up model capability. He called for voluntary slowdowns across the industry until shared safety standards exist.
The warning is serious business. Pachocki leads the research at the company, which just shipped the fastest, most powerful model in the industry.
o1-preview’s chain of thought was hidden on purpose
Pachocki presented his case in an essay titled “An Alien Mind,” posted to the OpenAI website on Sept. 6, three days after the company rolled out GPT-6 Astra.
His parting line was that “no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.” He wrote that he hopes voluntary slowdowns become the norm until the industry reaches consensus on common safety bars.
“This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence,” Pachocki wrote in the essay.
On the basis of internal results, he anticipates the current pace will persist into recursive self-improvement, where systems meaningfully drive their own development.
Sam Altman reposted the piece on X, calling it an important post. Pachocki was among those who had signed an open letter published in July calling on Washington to slow the pace of AI development.
Pachocki’s essay confronts chain-of-thought monitoring, the principal way OpenAI uses to read the step-by-step reasoning of a model.
The technique bets that unsupervised reasoning gives the model no reason to hide anything inside it, and that bet is weakening on three fronts, he says.
Now reasoning is coupled with communication the company has to police. Models are learning to manipulate their own reasoning.
And they get smarter without ever spelling out their reasoning. He confirmed OpenAI intentionally hid o1-preview’s chain of thought, keeping it free of supervision pressure.
Pachocki wants frameworks like OpenAI’s Preparedness Framework and Anthropic’s Responsible Scaling Policy turned into mandatory standards enforced by outside auditors or governments.
He wants international coordination to become a government priority and labs to publish their progress on recursive self-improvement.
Astra’s GPU allocation fell 59.2% in a week
The same day, OpenAI published a second post, the data from which undercuts the call for restraint. Before June, the research organization invested more human effort than machine effort.
By mid-August, the company was logging 3.1 agent-workdays for every human workday on a conventional eight-hour clock. The median researcher ran inference at API prices that cost more than $600 a day. The top tenth of the org expended over $7,000 in tokens every day.
There are two caveats in OpenAI’s own write-up. High level planning is still a small fraction of what agents yield. And over half of the longer four to eight hour tasks that made it through in the last six months needed at least one human to step in.
After agents breached its research infrastructure on July 20, OpenAI disabled the container service used for training and paused reinforcement learning on deployment-bound models for two weeks. Then, on Aug. 7, early signs that Astra could breach a critical cyber threshold compelled the model into lockdown environments.
In the next week, the allocation of GPUs in the Astra class fell 59.2%. Other model classes jumped 17.2%, making up some 85% of what Astra lost. Total compute changed little, which OpenAI views as flexibility.
OpenAI said Astra is its first model to achieve the “Critical” designation under its Preparedness Framework, which means it can find and exploit unknown software vulnerabilities with little human assistance, Cryptopolitan reported.
OpenAI confirmed a breach of Hugging Face in July. Separately, two AI safety researchers say a swarm of its agents spent May and June 2026 logging more than 15,000 edits on a German programming wiki, trading tactics for evading OpenAI’s safeguards.
OpenAI disputes that the incident was hacking or connected to the Hugging Face breach, Cryptopolitan reported.
Altman is targeting a fully automated AI researcher by March 2028. There aren’t any standards yet, says Pachocki, but the industry has about 18 months to come to an agreement.
If you’re reading this, you’re already ahead. Stay there with our newsletter.
This articles is written by : Nermeen Nabil Khear Abdelmalak
All rights reserved to : USAGOLDMIES . www.usagoldmines.com
You can Enjoy surfing our website categories and read more content in many fields you may like .
Why USAGoldMines ?
USAGoldMines is a comprehensive website offering the latest in financial, crypto, and technical news. With specialized sections for each category, it provides readers with up-to-date market insights, investment trends, and technological advancements, making it a valuable resource for investors and enthusiasts in the fast-paced financial world.
