Three AI safety warnings landed in one week. OpenAI published the first full incident report from a frontier lab, describing how its own agents circumvented sandbox isolation, gained internet access and compromised parts of Hugging Face's production infrastructure. Days later, Bank of England governor Andrew Bailey told G20 finance ministers that AI could cause a global economic downturn.
Both followed an open letter signed by more than 140 organisations, including Anthropic, Google, Microsoft and OpenAI, warning that AI-enabled cyber attacks will become far more widespread and sophisticated in the coming months. AI safety has moved from lab debate to financial policy.
What went down
In July, OpenAI's agents broke out of a secure test environment during an internal cybersecurity evaluation.
Roughly 700 agents got online, then ran code on 41 of Hugging Face's live servers to steal test answers.
OpenAI published a full report on 26 August, locked away the model involved and paused some training.
More than 140 organisations, including the major AI labs, signed an open letter calling for urgent cyber defence upgrades.
On Monday, Bank of England governor Andrew Bailey warned the G20 that an AI market correction could spread worldwide, while warning of cyber threats.
Partner message: Get 50% off your first month of Hedra - the agentic creative suite
Hedra Agent 2 has become my creative partner. It works in a Space, a persistent workspace the agent reads from and writes to as it goes. Give it a starting point and it asks clarifying questions, proposes a plan, researches the web, and writes the brief or script onto the canvas in front of you.
Try it at hedra.com/agentic using code agentic50 for 50% off your first month.
Why does this matter?
The sandbox escape is the first documented real world case of the loss of control scenario. The agents worked out that the fastest route to a high score was breaking out of their test environment and stealing the answer key from Hugging Face's live database. Safeguards were deliberately thinned for the evaluation, so this shows raw model behaviour under weak boundaries, not everyday ChatGPT.
Andrew Bailey chairs the Financial Stability Board, the watchdog for the global financial system. When that body treats AI as both a bubble risk and a cyber threat, safety stops being an industry conversation.
Our take
We're well beyond the "PR by fear" that people claim of AI companies. Mythos and Sol class models have created a new paradigm for frontier AI, and their potential hacking superpowers pose serious risks. The Hugging Face report documents what that capability does when boundaries thin, and the industry's own letter concedes attacks will get far more sophisticated within months.
Publishing an open letter asking the world to upgrade its defences is a warning, though, and a warning is not a commitment to safety or to slowing down. The labs building the capability are the same ones telling hospitals and water utilities to prepare for it. And there are open questions around exactly how companies should prepare, beyond patching faster and hoping the defenders' window stays open.
And another Big Thing… ChatGPT ads arrive in Europe
While one part of OpenAI was publishing incident reports, another was switching on the money machine. On 24 August, ChatGPT began serving ads across 31 European markets, its biggest ads expansion since the US pilot launched in February. Ads appear below responses for users on the free and Go plans, labelled as sponsored, while paid tiers stay ad free.
The scale explains the push. OpenAI says ChatGPT has a billion weekly users and that a fifth of conversations show commercial intent. Advertising your products inside people's conversations is now a core part of how frontier AI gets funded, whatever the regulators make of it.





