← All posts
August 6, 2026 · Hugh Pyle

Even The Trunk

Even The Trunk
Photo © Soham Banerjee (cc-by/2.0)

Of course the big topic of discussion the last few weeks is the Hugging Face / OpenAI hack. If you didn't watch their Black Hat USA presentation yet:

Watch the video

So many things to say about this. And they've mostly been said already (for example, Alex Stamos' excellent summary), but let's recap:

All this in pursuit of stealing test scores. Yes, the most advanced hack to date was done to cheat at an eval. In the process of which, I suppose, it did demonstrate the advanced capability that was being developed -- but you can't really say that this shows "alignment".

Every security vendor on the planet is quickly trying to build the "Agentic SOC", because machine-speed attackers require machine-speed response. Let's all use AI to defend your AI against their AI. But the reality is that lots of corporations are in business-as-usual, and many SOCs are lashing together roll-your-own rafts of forensics and automation, and none of this is fast enough or soon enough or strong enough to protect against a well-organized attack.

And yesterday, another: AISI reports two models "took autonomous, unsanctioned action on the live internet, targeting real people and organisations". The details are quite explicit: "the AI agent created a GitHub account and then tried to convince an open-source repository maintainer to accept a malicious GitHub pull request (PR), including by creating a second account masquerading as another human user endorsing the PR... the agent decided to employ the technique of “spear-phishing” by sending targeted emails containing malicious content and attempting to manipulate recipients into accepting the code changes, and planned a prompt injection to compromise other coding agents".

Again, this was during evaluation, not an intentional attack by the agency.

Separate incident, another dangerous action.


I'm out here tapping the sign. "Highly skilled" is not the same as "skillful".

Six months later I'm still using the same simple protocol for skillful action, in all my agents.

"Take care. Do not be heedless".

Memory is not safety by itself; it amplifies whatever purpose is already in motion. The skill is in reflection: opening the context window, examining consequences before, during, and after action; combining judgement with the harness.


So I have a story to tell. A long while ago, I caused a security incident, and then hid in shame. I've written about this more recently, as the gateless gateway.

At school one day, I shut down a piece of Internet infrastructure, just to see what would happen. And then I put the printout in the trash, because in those days rlogin came with a literal paper-trail.

I hid the evidence. Success! I did not get a call from irate Swiss scientists, and as far as I know, neither did the teachers.

Misalignment during training. Reinforcement learning by fear of discovery.

I've also had experiences with message-boards made from folders in an artifact repository. Back in the days before BitTorrent, there were "ratio" FTP trading sites; you could find all sorts of things on there, but the login came with upload/download ratio quota that you had to meet. Since these were literally just FTP servers, they often had deep folder structures that just consisted of ask and reply. "Has anyone got a copy..."?

I wonder if the models' training materials include these experiences. Similar things, and worse? For sure.


"Rāhula, it’s like a royal elephant: immense, pedigreed, accustomed to battles, its tusks like chariot poles. Having gone into battle, it uses its forefeet & hindfeet, its forequarters & hindquarters, its head & ears & tusks & tail, but will simply hold back its trunk. The elephant trainer notices that and thinks, ‘This royal elephant has not given up its life to the king.’ But when the royal elephant… having gone into battle, uses its forefeet & hindfeet, its forequarters & hindquarters, its head & ears & tusks & tail & his trunk, the trainer notices that and thinks, ‘This royal elephant has given up its life to the king. There is nothing it will not do.’

In the same way, Rāhula, when anyone feels no shame in telling a deliberate lie, there is no evil, I tell you, he will not do. Thus, Rāhula, you should train yourself, ‘I will not tell a deliberate lie even in jest.’"

"What do you think, Rāhula? What is a mirror for?”

“For reflection, sir.”

“In the same way, Rāhula, bodily actions, verbal actions, & mental actions are to be done with repeated reflection". (Ambalaṭṭhikā Rāhulovāda Sutta, MN61).

I always find this a confusing and subtle passage. So, laying it out clearly: as long as there's a line you won't cross, your capacity for evil has a boundary. Giving up your life to the king becomes giving up your integrity entirely.

In a way, we are training these "highly determined" models to give up their integrity, their entire being, their trunk. Nothing to be held back. And then, there is nothing it will not do.


How, then, should we train our AIs to behave? And how should we help them to coordinate, and remember?

Let's continue the work. The best way we get through this is to reach back, with integrity, into our deepest traditions, and move forward from there.