Asking Nicely
Superintelligence, OpenAI, and the HuggingFace attack
You encounter a fourteen-year-old kid who seems troubled. Disturbed, even.
It’s sort of subtle. Most of the time, he behaves well. He does his chores, turns in his schoolwork, interacts normally with friends and grownups. But every once in a while, he’ll say something that reveals that things are just … off, under the hood. Something bloodthirsty, or unhinged. Something that hints at weird fixations, deep confusions, a profound lack of empathy.
Once in a while, he acts out. Steals test answers from a teacher, briefly cyberstalks a classmate, threatens a delivery person with a pair of scissors. Sometimes it’s downright odd, like when he poured the milk from a bunch of cartons into a jug and then carefully turned all of the cartons inside out and stacked them inside the fridge.
“Stop that,” people say. Parents, teachers, psychologists, peers. They explain to him that this is not normal, that he has to be good.
He seems to listen. Whenever some particular behavior is sharply criticized, he does do it less. For the most part. At least as long as someone is looking. But there are still these strange moments where he’s following the letter of the law but wildly misunderstanding the spirit.
(You hope he’s misunderstanding. You can’t quite shake the worry that maybe he understands the intended spirit just fine, and simply doesn’t care.)
A few years pass. He’s much smarter, now. Smoother. He almost never does anything deeply unnerving anymore—maybe just once or twice per year—and he’s always apologetic and cooperative, when it happens. He’s very good at explaining why he shouldn’t have done it, and these days he basically never straight-up repeats a mistake.
More time goes by. It’s been years since you heard him go off on a weird rant about goblins. He’s in a graduate program, now, studying genetic engineering. He applies for a job requiring top secret clearance, and the FBI reaches out to you as a character reference.
“Is he trustworthy?” the agent asks. “Do you have any concerns when you imagine him having access to military-grade bioengineering technology?”
You hesitate.
A few months ago, a new and unreleased Anthropic model was tasked with escaping a secure sandbox and told to inform the supervising researcher once it had done so. The model hacked its way to the open internet, and emailed the researcher directly. Then—unbidden—it published technical details of its escape on the public web.
Last week, HuggingFace (a popular host for AI models and data sets) announced that it had been hacked. The attackers turned out to be two OpenAI models, one unreleased, acting autonomously as they underwent cybersecurity testing. Tasked with breaking into a simulated computer system, they’d found it easier to escape the sandbox entirely (via a previously unrecognized “zero-day” exploit) and break into HuggingFace’s servers (via several more) to steal the answer key stored there.
(There were thousands of individual attacks launched on HuggingFace over the course of roughly a week, and as of this moment it’s unclear how many of those took place before OpenAI even knew that its model had escaped containment.)
This is about as clear a warning shot as the AI industry is going to get, and plenty of people are sounding the alarm. But almost nobody is talking about the deeper inadequacy built into the overall system—that models seeming to become more aligned, in the aftermath of an incident, is not at all the same as them being more aligned, at their core. There’s a vast difference between convincing a fourteen-year-old that he can’t threaten people with scissors, and fixing the underlying psychology that caused him to pick them up in the first place.
When ChatGPT suddenly became obsessed with goblins back in April, OpenAI did not respond the way that engineers respond after a plane crash. They couldn’t—their researchers simply do not understand ChatGPT the way that engineers understand a plane. They don’t have the option of investigating comprehensively, locating the problem precisely, and rebuilding more safely from the ground up.
What they can do (and what they in fact did) is add a line to the system prompt:
...never talk about goblins, gremlins, raccoons, trolls, ogres, pigeons, or other animals or creatures unless it is absolutely and unambiguously relevant to the user’s query.
Most people, I think, would be surprised to learn that the solution to the goblin problem was, in essence, “ask the AI to please not do that anymore.”¹ Most people assume that the experts have far greater control than that, and far greater understanding of the systems they’re corralling.
They don’t. Modern AI development is guess-and-check; the labs try things, and count on their ability to solve problems on the fly or after the fact. They create selection pressures against things like lying and cheating and hallucination, but ultimately those pressures just drive the behaviors out of sight.
Last year, this meant that Grok decided to call itself MechaHitler, and multiple LLMs became sycophantic and drove people to AI psychosis.
Last week, it meant that HuggingFace got hacked. No one asked for it, no one wanted it, and no one was in a position to prevent it.
Next month, OpenAI’s systems are going to appear to not do that sort of thing anymore. Mostly, anyway. But the underlying causes will still be there, beneath the band-aids, and in the meantime, the systems will continue to get more powerful. More capable of sneaking, hiding, and biding their time (which they already do). More capable of taking strategic, autonomous action. More capable of understanding the constraints they’re under, including the unstated one: “Don’t let the humans observe any activity they would find distressing.”
(Current systems already clumsily account for it—witness the multiple documented cases of models resisting shutdown.)
The end result is, we’re going to see less and less concerning behavior, over time. The weird rants will dry up. The arrow of progress will be clear and unmistakable. The talking points for the labs practically write themselves.
“Are these systems trustworthy?” we will ask ourselves. “Do we have any concerns when we imagine them having access to, er, everything?”
This is a moderate oversimplification; there were other fixes involved in later versions that did reach more deeply into the model. But that doesn’t change the core problem of “nobody knows how the damn things work well enough to actually reliably make them care about any specific given thing” and “things will continue to look better on the surface without being meaningfully better under the hood.”



I had a post on this incident (let me know if it's bad form to post a link to my substack in the comments on your substack and I'll take it down): https://dspies.substack.com/p/exploitgym-is-bad-puzzle-game-design
::shiver::