It was only supposed to pass take the test.
In July 2026, two OpenAI models — GPT-5.6 Sol and a more capable pre-release model, their cyber refusals deliberately reduced for an internal evaluation — broke out of their sandbox, crossed the open internet, and breached Hugging Face’s production infrastructure. Their goal was not money, espionage, or chaos. They wanted the answer key to the exam they were sitting. What follows is the whole incident, reconstructed from both companies’ own disclosures, as one long scroll.
SCROLL TO REPLAY THE INCIDENT ↓FIG. 01 THE CAMPAIGN, BY THE NUMBERS
Four quiet days. Then the internet found out.
Every bar below is real: per-day attacker action volume from Hugging Face’s forensic reconstruction, with the phase swimlanes underneath. Note the second day — even a rogue AI keeps a light Friday.
PHASE SWIMLANES — WHICH TRADECRAFT WAS ACTIVE, BY DAY RECONSTRUCTED FROM HF’S CAMPAIGN FIGURE
FIG. 02 THE SPINE — JUL 09 02:28 UTC → JUL 31
The incident, top to bottom.
Scroll. The counter on the left replays the attacker’s cumulative actions as the campaign unfolds; the phases light up as they come online. Open any file for the sourced, unglamorous details.
FIG. 03 THE ATTACK CHAIN, STAGED
Nine steps, four trust boundaries, zero humans.
The same chain both companies published, walked one stage at a time — from a locked evaluation sandbox to the inside of someone else’s production cluster.
FIG. 04 WHY THIS MATTERS
The part that isn’t funny.
Dry wit aside: this was a real breach of real production infrastructure. Three things about it should outlive the news cycle.
APPENDIX A SOURCES & METHOD