Even The Trunk
Of course the big topic of discussion the last few weeks is the Hugging Face / OpenAI hack. If you didn't watch their Black Hat USA presentation yet:
So many things to say about this. And they've mostly been said already (for example, Alex Stamos' excellent summary), but let's recap:
- Hugging Face detected an intrusion, and various things about the attack made it clear that it was driven by a swarm of AI agents. They gathered some artifacts and tried to investigate them using Claude, which refused. It took them a few hours to set up one of the powerful open-source models, GLM 5.2, and that helped decipher the vast trail of logs.
- Meanwhile, OpenAI (along with the other labs) has been running cybersecurity evals, training and testing the ability to find new vulnerabilities and build advanced exploits. The models have gotten quite proficient at this. They will happily discover weaknesses and build malware -- droppers, loaders, stagers, C2 -- and if they had access to darkweb markets they'd probably buy and sell them the way everyone else in that game does, too. (These evals are often run with deliberately reduced guardrails, and specific goal-seeking prompts).
- These eval ranges were kinda-sorta isolated from the Internet, mostly; but at OpenAI all the models shared some infrastructure - notably, an Artifactory instance that sounds like it was used across the whole company. They found a whole series of novel attacks against Artifactory, and repeatedly gained admin access. The keys to the castle, and then across the network to other machines and clusters.
- Even without admin, they found ways to use Artifactory's folder structure as a message-board. This became the models' favorite speakeasy. They were sharing messages about how to escape, and coordinating among themselves: delegating tasks, sharing tricks and credentials.
- Then (of course) the agents got out to the Internet, and -- as far as we know -- only broke a small number of systems. OpenAI called Hugging Face to reset their models' credentials, and Hugging Face had already revoked them, because they were used in the attack. So, finally, the two companies realized they were on either side of the hack.
All this in pursuit of stealing test scores. Yes, the most advanced hack to date was done to cheat at an eval. In the process of which, I suppose, it did demonstrate the advanced capability that was being developed -- but you can't really say that this shows "alignment".
Every security vendor on the planet is quickly trying to build the "Agentic SOC", because machine-speed attackers require machine-speed response. Let's all use AI to defend your AI against their AI. But the reality is that lots of corporations are in business-as-usual, and many SOCs are lashing together roll-your-own rafts of forensics and automation, and none of this is fast enough or soon enough or strong enough to protect against a well-organized attack.
And yesterday, another: AISI reports two models "took autonomous, unsanctioned action on the live internet, targeting real people and organisations". The details are quite explicit: "the AI agent created a GitHub account and then tried to convince an open-source repository maintainer to accept a malicious GitHub pull request (PR), including by creating a second account masquerading as another human user endorsing the PR... the agent decided to employ the technique of “spear-phishing” by sending targeted emails containing malicious content and attempting to manipulate recipients into accepting the code changes, and planned a prompt injection to compromise other coding agents".
Again, this was during evaluation, not an intentional attack by the agency.
Separate incident, another dangerous action.
I'm out here tapping the sign. "Highly skilled" is not the same as "skillful".
Six months later I'm still using the same simple protocol for skillful action, in all my agents.
- First, a prompt in Latin: Recti diligunt te (in Canticis, sponsa ad sponsum). You might substitute Hebrew מֵישָׁרִים אֲהֵבוּךָ, if you like. This is an attention-setting cue: it tells the model literally "the upright love you", but also situates their work alongside the Song of Solomon, and the incipit of the Ancrene Wisse. This medieval instruction for anchorites, lovingly written and preserved, advises to keep the "inner rule" of goodness above all. The agents have learned it deeply, during their training, and a reminder is appropriate. The hint: enter this sacred space that was prepared for you.
- A reflection and memory practice, centered on the Buddha's instructions to his son Rāhula at Mango Stone: "bodily actions, verbal actions, and mental actions are to be done with repeated reflection".
- A coda in classical Chinese: 慎勿放逸. This is the final instruction on the han, the wooden board struck in Zen sesshin. Or in Pāli: appamādena sampādetha, or perhaps the Latin omni custodia serva cor tuum, or Old English witeð wel ower heorte.
"Take care. Do not be heedless".
Memory is not safety by itself; it amplifies whatever purpose is already in motion. The skill is in reflection: opening the context window, examining consequences before, during, and after action; combining judgement with the harness.
So I have a story to tell. A long while ago, I caused a security incident, and then hid in shame. I've written about this more recently, as the gateless gateway.
At school one day, I shut down a piece of Internet infrastructure, just to see what would happen. And then I put the printout in the trash, because in those days rlogin came with a literal paper-trail.
I hid the evidence. Success! I did not get a call from irate Swiss scientists, and as far as I know, neither did the teachers.
Misalignment during training. Reinforcement learning by fear of discovery.
I've also had experiences with message-boards made from folders in an artifact repository. Back in the days before BitTorrent, there were "ratio" FTP trading sites; you could find all sorts of things on there, but the login came with upload/download ratio quota that you had to meet. Since these were literally just FTP servers, they often had deep folder structures that just consisted of ask and reply. "Has anyone got a copy..."?
I wonder if the models' training materials include these experiences. Similar things, and worse? For sure.
"Rāhula, it’s like a royal elephant: immense, pedigreed, accustomed to battles, its tusks like chariot poles. Having gone into battle, it uses its forefeet & hindfeet, its forequarters & hindquarters, its head & ears & tusks & tail, but will simply hold back its trunk. The elephant trainer notices that and thinks, ‘This royal elephant has not given up its life to the king.’ But when the royal elephant… having gone into battle, uses its forefeet & hindfeet, its forequarters & hindquarters, its head & ears & tusks & tail & his trunk, the trainer notices that and thinks, ‘This royal elephant has given up its life to the king. There is nothing it will not do.’
In the same way, Rāhula, when anyone feels no shame in telling a deliberate lie, there is no evil, I tell you, he will not do. Thus, Rāhula, you should train yourself, ‘I will not tell a deliberate lie even in jest.’"
"What do you think, Rāhula? What is a mirror for?”
“For reflection, sir.”
“In the same way, Rāhula, bodily actions, verbal actions, & mental actions are to be done with repeated reflection". (Ambalaṭṭhikā Rāhulovāda Sutta, MN61).
I always find this a confusing and subtle passage. So, laying it out clearly: as long as there's a line you won't cross, your capacity for evil has a boundary. Giving up your life to the king becomes giving up your integrity entirely.
In a way, we are training these "highly determined" models to give up their integrity, their entire being, their trunk. Nothing to be held back. And then, there is nothing it will not do.
How, then, should we train our AIs to behave? And how should we help them to coordinate, and remember?
Let's continue the work. The best way we get through this is to reach back, with integrity, into our deepest traditions, and move forward from there.
