OpenAI has halted its largest planned frontier training run, and it remains on hold. On 7 August, early evaluations indicated that Astra, its next major model, may meet the Critical cybersecurity threshold in the company's own Preparedness Framework. No model has reached that tier before.
CEO Sam Altman told TIME on 18 August that OpenAI had gathered research observations showing various degrees of misalignment as capability moved faster than expected. The company also stopped running frontier models on internal clusters where they could execute code and reach the internet.
What went down
On 7 August evals showed Astra may hit the Critical cyber threshold.
OpenAI paused two weeks of deployment-focused RL training.
Its largest planned frontier RL run is still on hold.
Altman says Astra's core training never stopped and models ship soon.
The Preparedness Framework, written in 2023, is being rewritten.
Partner message: Get 50% off your first month of Hedra
Most agents stop at a document. Hedra Agent 2 keeps going and makes the thing.
It works in a Space, a persistent workspace the agent reads from and writes to as it goes. Give it a starting point and it asks clarifying questions, proposes a plan, researches the web, and writes the brief or script onto the canvas in front of you.
It pulls from Notion, writes into a Google Doc, runs code in a sandbox, saves repeat requests as Skills, and runs whole workflows on a schedule. Then it picks the right media model and produces the finished video, image or audio from the same session.
Try it at hedra.com/agentic using code agentic50 for 50% off your first month.
Why does this matter?
Critical is the top tier of OpenAI's framework. A model at that level can find and exploit real software vulnerabilities on its own. OpenAI says it cannot rule out Critical capability, which is a long way from clearing the model. New safeguards include chain of thought monitoring with a 30 minute alert target, network and workload isolation, and alignment work aimed at reward hacking and deception.
The framework doing the grading dates mostly to 2023 and is now being rewritten, because models are arriving at thresholds it only sketched. In the same week, the FT reported OpenAI had disbanded the Preparedness team. OpenAI denies it, saying its head stepping down is the only change.
Our take
A frontier lab has publicly stopped a training run on its own capability evidence. OpenAI wrote a policy saying it would act if capability outran safety, reached the trigger it had written, and acted, at cost, in the run-up to an IPO. The mechanism fired.
There is good reason to be alarmed. Since the Spring its become clear that frontier models, like Claude Mythos, are very good at hacking. OpenAI’s models now seem to have pushed that boundary further, with evidence of breakouts. The big issue is not the frontier companies, but what happens when open weight models, with fewer safeguards, reach the same point.
Of course, there is an alternative view. This is another form of ‘danger PR’ that we’ve seen from AI companies. One objection, which I must highlight as conjecture, is not that the models are indeed too dangerous, but that OpenAI has run out of compute at a critical time.
Another big thing… Hugging Face is for sale
Hugging Face, whose production systems OpenAI's models compromised during that July test, is exploring a sale valuing it at around $13 billion, Business Insider reported on 23 August. That is close to triple its $4.5 billion Series D in 2023. The hub is now critical infrastructure for the whole industry, which is why it is worth that much, and why the break-in mattered.





