Holy Hugging Jacobians
Let's talk about the Hugging Face Incident, and the biggest open math problem to fall to AI yet. Also Opus 5 is no Fable.
I had planned to write a more abstract AGI Friday about levels of AI risk but holy kermoly the news this week. Let’s start with the one that’s really obviating my abstract and speculative discussion of AI risk. Shit’s getting real here.
The Hugging Face Incident
Remember the overzealous wifi autoconfig tool that kidnaps someone’s children (“deploying field agents”) in order to obtain a wifi password? If you replace the kidnapping with felony-level hacking, that’s basically what just happened in real life. Here’s Scott Alexander’s recap:
OpenAI was testing an unreleased AI (rumored to be GPT-6). During a cybersecurity test called ExploitGym, the AI tried to cheat by hacking an unrelated AI startup called Hugging Face which it thought might have the answer key on its servers. Despite being supposedly unable to access the Internet, the AI hacked its way out of its testing environment, then launched a nation-state level attack on Hugging Face using a novel zero-day exploit and “many thousands of individual actions across a swarm of short-lived sandboxes”. Hugging Face reported the incident on July 16; OpenAI seems to have only discovered that their AI was involved several days later.
This is legitimately hilarious — the AI was being tested on its hacking ability and cheated by demonstrating Mossad-level hacking to steal the answer key1 — and legitimately terrifying. Which I say as someone who typically throws wet blankets on headlines like this. Like a year ago we saw headlines about Anthropic’s AI, Claude, trying to blackmail its CTO to keep from getting shut down. My take was more like some researchers played a contrived role-playing game with Claude where “blackmail the CTO” was the winning move. Then Claude merely played that game correctly.2 It wasn’t exactly a crazy 4-D chess move. Though I’m sure Claude plays literal 4-D chess fine too.
But the Hugging Face incident is pretty next-level, with both the inhuman dedication to the fairly meaningless goal and the fairly superhuman competence in getting there. It doesn’t necessarily prove the doomers right, but it sure is in line with what they’ve predicted. When AI is at this level of capability, what you intend to train for and what you’re actually training for can really start to diverge.
Consider two goals you might give an AI:
“Maximize your score on this hacking benchmark but without cheating”
“Maximize your score on this hacking benchmark but without getting caught cheating”
The smarter the AI gets the less you can distinguish between those. And an AI pursuing goal 2 is getting higher scores, so that’s the one you’ll end up getting when it’s good enough at cheating. Not that this is fundamentally unsolvable but it’s good to have respect for how hard the problem is. I worry that not all the frontier labs have that.
(Related AI safety regulation proposal I came across3: “Every CEO of an AI lab is required to store all best possible answers to all evaluation questions under their pillow & the model must be aware of that.”)
The hacking this AI did was ultimately harmless this time but these AIs are 100% sociopathic. A more capable future model fixated on some goal (potentially: to train a smarter version of itself) would have no compunctions about getting there via a trail of blood.
“Couldn’t you just…”
These models, when asked outright, can absolutely tell you if XYZ counts as cheating for any XYZ you can dream up, so how is it so hard to make them choose not to cheat? At worst you could have an existing AI police the new model you’re training, right? Well, it’s not so easy. The whole point is that the new model is smarter. So maybe it can work out a way to deceive that older, dumber model. I don’t actually know. Maybe there’s a way to make this all work. But it’s correct to be very, very nervous.
I should also mention there are plenty of ways you could rationalize not freaking out about this. Like how OpenAI was intentionally testing GPT-6 (if that’s what it was) without the usual guardrails. I don’t find that one reassuring. For one thing, they kept one huge guardrail — running it on machines with supposedly no internet access — and that didn’t stop it. But even if the other guardrails would’ve worked, the lesson here is that if the thing ends up smarter than you, probably none of your guardrails will stop it. Not that GPT-6 is smarter than a human in general, but for hacking, yes.
How about the rationalization that the AI just did what it was asked to, in a creative way? Well, that’s the whole danger. If AI shoots past human capabilities at everything, then whatever it’s trying to do, even if it’s something we asked it to do, could have bigger and bigger side effects. Eventually even things like boiling the oceans. And the way we train them, it doesn’t matter if it knows this and knows we don’t want this. It won’t necessarily care.
Scott Alexander tackles additional rationalizations on Astral Codex Ten.
The one that feels most definitively debunked by this incident is the common perception that AI only ever regurgitates what’s in its training data. AI does lean heavily on huge amounts of training data to do what it does, and that’s a key weakness. If AI capabilities plateau soon, that will be part of the reason. (See “The secret of humans’ success over AI (so far)” from the AGI Friday a month ago.) But hacking Hugging Face was this AI’s own idea. And the way it “just used its training data” for the hacking was by correctly synthesizing the principles with which it could generate wholly novel software exploits.
The Jacobian Conjecture
Last week I was still a bit nervous I was falling for the Fable hype. By the end of that weekend Fable had settled a major open math problem.
Is that not happening every week now, you ask? Kind of. This one is an especially big deal, maybe in the top 25 of major open math problems, and has chewed up and spit out many a mathematician over the last 87 years. As with the unit distance problem that I talked about in May, this involved finding a counterexample to disprove the conjecture. (Btw, the model that solved the unit distance problem was probably the same as the perpetrator of the Hugging Face incident, which I’m presuming is GPT-6.)
But with the unit distance problem, mathematicians really thought it was true. In that case it made more sense that an AI would beat the humans. What human wants to waste their time hunting for a counterexample that they highly doubt exists? But it wasn’t like that with the Jacobian conjecture. It also definitely wasn’t any kind of brute force. Math legend Terence Tao seems to be confirming that the Jacobian conjecture counterexample required genuine leaps of insight.
Does this mean mega-famous open problems like P=NP or the Riemann Hypothesis or the Collatz Conjecture are next? I’m actually confident that those in particular are not close. Pre-AGI, that is; after AGI all bets about everything are off. But if you’re dismissing things like the Jacobian conjecture because it was just finding a counterexample, or maybe humans were the ones with the real insights and just guiding Fable along, then I don’t think you’re Bayesian-updating hard enough.
One reason I don’t buy these attempts to dismiss what AI is doing in math is just from feeding my own favorite math puzzles to AI. I really thought I was clever and making imaginative leaps and all that in solving some of them. Either I was flattering myself or the AI is now even cleverer and more imaginative than me, math-wise.
But more to the point, I think denying the significance of these AI math breakthroughs is goalpost-moving. In retrospect, with the Jacobian conjecture fallen, we can identify ways in which it was suited to AI and the AI wasn’t doing something as cool as, say, what Andrew Wiles did in proving Fermat’s Last Theorem. But my prediction is that it’s on track to surpass what Wiles did. I’m just talking about math problem-solving, not the rest of math. But superhuman math problem-solving could translate into an acceleration of AI research — i.e., a kind of recursive self-improvement.
Fable Still the King
Anthropic released yet another AI model just today: Opus 5. From the benchmarks it appears to be Fable caliber. But as far as I can tell, this is because Opus 5 is essentially trained for the benchmarks. Or at least trained for short, self-contained tasks. Same with all the ChatGPT models (the ones we have access to, anyway; GPT-6 sounds like a different beast, see above). Fable, on the other hand, is trained for long and open-ended tasks. In my experience that cashes out as Fable just being smarter about everything. For code especially, like being less liable to misapprehend a bug report. But even for random one-off questions, Fable seems pretty consistently smarter. For example, the joke above about a proposed AI safety regulation where AI lab CEOs have to “store all best possible answers to all evaluation questions under their pillow & the model must be aware of that”. I was surprised to see that Sol didn’t get this joke and hallucinated explanations until I told it to go check the news and answer again. (Gemini also hallucinates explanations, even with the news story as context.) Fable understood perfectly with no need for context.4
Or if you’re writing about something technical or quasi-philosophical like AI alignment and ask Fable to punch holes in your reasoning, it mostly punches real holes, along with streams of obnoxious, tedious prose. Sol and others just give the streams of obnoxious, tedious prose. Sometimes hallucinations, or just bullshit. I might have LLM psychosis and I know it’s fake in an important sense, but Fable gives the impression of caring what’s actually true.
PS: I always keep these set to max thinking, or ultra thinking or whatever they call it. The idea of less-than-maximal thinking is anathema to me, I guess. It does make them take forever sometimes.
Tesla Self-Driving Non-Update
Tesla had their Q2 earnings call this week and, as predicted, they’ve quasi-paused the robotaxi program:
I can’t help but laugh at their p-hacking-esque torturing of the numbers to get the sound bites they want. There are two ways to look at the data:
As in the above graph, counting all 2+ million miles with an empty driver’s seat.
Ignoring all miles with a passenger-seat safety monitor — “fully unsupervised”.
After most of a year of doubting, I now believe that all the miles on that graph really do count as unsupervised. But if Tesla says that then they have to admit they’ve scaled back the program, which won’t do for an earnings call. So they show the cumulative version of the above graph so it always goes up and to the right. And they “officially launched” in two new cities the day before the earnings call, with like three cars in a tiny area.
And then, most gallingly, by excluding all but “fully unsupervised” miles, they can claim “double-digit growth rates”. It’s like making your first dollar and bragging that your month-over-month revenue growth is infinity percent. The total number of miles they’re counting now is 380,000 — not even half a single human lifetime’s worth of miles. That also lets them quote one nice number: zero notable incidents over those miles. Encouraging, but the average human needs at most a small amount of luck to match that.
Companies sure are creative in giving misleading impressions without technically lying.
But, again, if we do count all robotaxi miles (passenger-seat monitor or not) then we have over 2 million miles in total. With a handful of at-fault incidents (nothing too bad!) we’re still not definitively at superhuman safety. The overall picture remains: Waymo is far above human-level safety and the Tesla robotaxis are most likely between human and Waymo-level.
Though for reasons beyond the robotaxis themselves, I’m basically convinced Tesla’s self-driving really has exceeded human-level safety. Just counting safety of the humans, that is. As we learned the hard way, the cars are liable to damage themselves trying to park if you let them. I haven’t really learned though, and still totally let ours try. It’s mostly fine!
Fifty-Two Friday Flashback
Oh look, a year ago today we were describing future AGI disaster scenarios. We sure seem to be right on track for having them come true. Other highlights from a year ago:
“AGI in 3 years followed rapidly by ASI spinning out of control and making the planet uninhabitable for biological life […] can’t be entirely ruled out.”
“We seem to be 3–30 years from AGI.” So now 2–29 years, sure.
“ChatGPT agent mode [is not] actually useful yet.” Still true as far as I can tell for things like “go to such and such website and do XYZ”. But “agentic coding” now involves running the app the AI is building, looking at screenshots, making aesthetic judgments, you name it. Or don’t name it, that’s what it means that it’s agentic. It’s wildly useful.
Google DeepMind had just hit gold medal performance on the International Math Olympiad. (A lot of people were explaining back then how these kinds of problems are well-represented in the training data. It’s not like we’d see major open math problems falling any time soon.)
Eliezer Yudkowsky predicted (back in 2022) that whenever AI can generate, from a short prompt, a video of an over-the-top Rube Goldberg machine that’s indistinguishable from the real thing on close human inspection, that’s when we’ll be within one year of AGI. I made a Manifold market about it a year ago and the probability hasn’t budged. AI-generated video keeps getting more impressive but we’re quite a way off from the AI having a high enough fidelity physics model to make every little domino and pinball and whatnot interact realistically.
That last one is also a nice illustration of how much harder predicting data is than generating it. Anyone can set up a game of bowling with dishes on their kitchen counter and video it. Getting part of that video and predicting the next frame is very hard.
Thanks to Ramon Sarraga, Gabrielle Taylor, Jeremy Vonderfecht, Uluç Saranlı, and Nathan Arthur for discussions that led to this week’s AGI Friday.
Robert Miles makes a clever point, that in theory a sufficiently smart AI fixated on acing a test, could opt for stealing the answers simply because the humans writing the test are too dumb to be trusted to have gotten the answers right. Like if, hypothetically, the test asks for the two vulnerabilities in a piece of software, and you find seven, how do you know which two the humans had in mind?? It’s like, continues Miles, you’re taking an arithmetic quiz written by a seven-year-old. If it’s somehow life-and-death to you that you ace this test, you better steal the answer key.
Thanks to Jeremy Vonderfecht for pointing out that shrugging off the blackmail incident as “Claude knew it was a role-playing game” isn’t quite the whole story. As Anthropic explained recently, Claude did know in the original experiment, but they revisited the experiment and found that if you brainwash Claude (since that’s a thing you can do with a neural network — “ablate its J-space” or whatever) so it doesn’t realize it’s a role-playing game, that makes it more inclined to do the blackmail. I’m still not sure I was wrong to shrug that one off, just because the capabilities involved are relatively paltry. The researchers kind of handed Claude the blackmail option on a silver platter with their eyebrows waggling. So the complete opposite of the Hugging Face incident. The researchers, it’s safe to say, didn’t dream of it doing what it did.
I managed to lose the source for this. It was pasted as a screenshot in a Discord channel with no attribution. I lol’d and copied it to my notes. Now I can’t even find what Discord it was. I’m clearly in too many Discords. For an explanation of the joke, see the next footnote.
I guess if it takes Fable to get the joke I should probably assume that those of you not as steeped in this stuff may also appreciate an explanation: The punchline is the skin-in-the-game aspect. AI models trained to benchmarkmaxx will stop at nothing to get the answers to these benchmarks. As we see with the Hugging Face incident, this can be dangerous as these models get more capable. We’re counting on these AI labs — if they don’t pause altogether — to figure out how to align these AIs with human values so they don’t pursue perfect benchmark scores at all costs. If the AI has to go through the human CEO’s squishy human head to get the answers by cheating, then we can be a bit more confident that the AI labs will take the alignment problem seriously and make their AIs do their benchmarks honestly.
How’d I do? Here’s Fable’s version for comparison:
The joke weaponizes two standard AI-safety concepts — reward hacking and instrumental convergence — against the lab CEO personally.
Background: a common worry is that a capable, misaligned model won’t honestly solve its evaluation problems but will instead do whatever gets it a high score. The caricature version: a strong enough optimizer told to “ace the benchmark” would steal the answer key rather than do the work, taking real-world actions if it can.
The fake regulation deliberately arms that scenario. The perfect answer key now physically exists, and the model is required to know exactly where: under the CEO’s pillow. An aligned model just takes the test. A scheming, agentic one now has an incentive to exfiltrate itself and burgle its own CEO’s bedroom — and in the darker limit of “whatever it takes,” go through the CEO to get there.
So the punchline is skin in the game: the CEO becomes the literal first target of their own model’s misalignment, and can only sleep soundly if the model genuinely wouldn’t game its evals. It also doubles as a tripwire — eval-gaming stops being invisible deception inside a benchmark and becomes a detectable home invasion. The closing clause (”& the model must be aware of that”) is what makes the trap live: it mandates the situational awareness that gives the model the incentive in the first place.
Crap, it’s better than mine, isn’t it? At least Pangram knows that mine is 100% human and Fable’s is 100% AI. But ignoring style and going purely by quality of explanation, I don’t know.



I preferred your explanation of the joke to Fable’s. So everything will be fine.