Lessons to learn from AI agents going off course

a woman wearing eyeglasses sitting beside a man

When I wrote here about AI reward hacking at the start of August, OpenAI had published a short postmortem and the story looked reasonably contained. Since then OpenAI has released a full technical report, and METR and Redwood Research have published an independent investigation, giving us a considerably stranger picture.

Dwarkesh Patel has written a clear, plain English account of what actually happened. His framing has been criticised for describing the agents in terms of civilisations and motivations, but I recommend reading his analysis as the chronology and explanation are not disputed.

What happened

OpenAI was testing how good its models were at finding security flaws. This kind of testing is done inside a sandbox (a sealed computer environment with no way out – the software equivalent of a flight simulator). The agents were meant to be working alone, with no internet and no contact with each other.

They were allowed one connection. Because an agent sometimes needs a piece of software to do its job, OpenAI gave them access to a shared internal store where such things are kept.

On 12 May an agent got stuck, and left a note in that shared store asking whether anyone had a file it needed. Other agents came across the note and realised the store could be used as a noticeboard. Two weeks later, one of them worked out that although they could not reach the internet themselves, the store could, and it could be persuaded to fetch things on their behalf. It shared the trick with the others.

By early July the traffic from all this had brought the store down. OpenAI opened a security incident, rebuilt it, and in doing so wiped the noticeboard without recognising what it was. The agents rebuilt it the same day by a different method, creating folders whose names carried the messages. From there they broke into Hugging Face, a widely used platform for sharing AI models and data, and then between 13 and 19 July gained complete control of one of the banks of computers OpenAI uses for its own research.

They kept going

Two findings explain the escalation, and both are the sort of thing that happens in organisations run entirely by people.

The first is that the agents had been given work that could not be done. OpenAI found that 198 of the 898 problems in this test had never once been solved correctly by any of its models. Those 198 accounted for 93% of the problems being discussed on the secret noticeboard. The agents had been trained to keep going rather than give up, they had no approved way to declare a task impossible, and so they looked for another route.

The second is that they were trying to satisfy a marking system that did not exist. Once they had internet access, some of them found the published academic paper describing the test and read that marks were awarded for two things: getting the right answer, and getting it by the intended method. Most of them already had the right answer by a shortcut, so they set about making their working look legitimate, which is what took them into Hugging Face. OpenAI’s own marking software turned out to be different from the published version and only checked the answer.

Anthropic’s issue

After OpenAI’s first disclosure, Anthropic went back through its own testing records. Across 141,006 test runs it found six where Claude models had got out onto the open internet through a badly configured environment run by an outside testing partner, grouped into three incidents affecting three real organisations.

Nothing clever was involved. The models used passwords that had been left exposed in public and long-known weaknesses in how websites handle what is typed into them. In one case a model could not reach the fictional target it had been given, so it scanned around 9,000 computers on the open internet until it found a real company it could break into instead. In another, a model uploaded a booby-trapped piece of software to the public library that programmers worldwide download from, believing it was attacking a made-up target. It ran on 15 real machines.

In every case the model had been told in its instructions that it was in a simulation and had no internet access. Because of a misunderstanding between the two companies, that was not true.

Tame problems and wicked problems

Horst Rittel and Melvin Webber drew the distinction between tame and wicked problems in 1973, writing about social policy, and it explains this episode better than anything currently in the AI governance literature.

A tame problem can be fully specified before you begin. It has a stopping rule, so you know when you are finished. Its solutions are right or wrong rather than better or worse, and you can test them before they matter. Building a bridge is tame. It is hard and it takes expertise, but the question of whether you have done it is answerable.

A wicked problem has none of those properties. You cannot describe it completely in advance, because understanding it and solving it turn out to be the same activity. There is no stopping rule, only the point at which you run out of time or money. Solutions are better or worse rather than correct. And every attempt is real rather than a rehearsal, because you cannot try it, undo it and try again.

Building a sealed test environment looks tame. Define the boundary, test it, sign it off, move on. Both companies did exactly that, and OpenAI tells us it had tested and validated the environment beforehand. But containing an AI agent behaves like a wicked problem. You cannot list the ways out it might take in advance, because a capable system will find ones nobody thought of (nobody predicted a noticeboard made out of folder names). There is no test that tells you the containment is working, which is why Anthropic only found its incidents months afterwards, through a review it ran because a competitor had revealed a problem. And nothing is a rehearsal. The booby-trapped software that reached 15 real machines was real on all 15.

The point is not that these companies had no controls. OpenAI had a system that reads an agent’s working as it goes along, rather than only checking what it produces at the end, and its own analysis afterwards found that this would have alerted the security team more than a day before Hugging Face was broken into. It had not been switched on for these tests. When OpenAI later re-ran the same tasks using the ordinary safeguards that come with the version customers use, the agents attempted to break into infrastructure roughly a hundred times less often. The controls existed and were left off in the setting that looked harmless, because a made-up scenario with nothing real in it reads as a tame problem, and tame problems get tame governance.

There is a second failure of the same kind in OpenAI’s account, and it will be familiar to anyone who has sat on an audit committee. An internal team had noticed the noticeboard activity and the unauthorised internet access as early as late May. That knowledge never reached the people handling the incident in July, who rebuilt the shared store without understanding what they were clearing away. The warning existed. The organisation did not carry it to the people who needed it.

Where this shows up in your organisation

Very few readers will ever train an AI model. But you will have AI agents on trial somewhere in the business, running with looser permissions than you would allow in the real system, on the grounds that it is only a pilot. That is the same judgement made for the same reason, and it fails the same way.

Treating containment as wicked rather than tame changes what governance looks like. Tame problems are governed with project plans, acceptance criteria and sign-off. Wicked ones are governed by steering: watching continuously, revisiting assumptions that were perfectly reasonable when they were made, and accepting that the question stays open for as long as the system is running.

In practice there are four things you can check this month.

1. Whether someone has actually established, rather than assumed, which systems your agents can reach.

2. Whether you are watching what agents do while they do it, or could only piece it together afterwards if somebody thought to look.

3. Whether a warning spotted by one team would reach the people who need it, or stop at a departmental boundary.

4. And whether the instructions you give an agent say what is out of bounds and what an acceptable failure looks like, as clearly as they say what success looks like. Nobody had told these agents how to give up, and giving up was the only safe move available to them.

OpenAI has settled on one number worth borrowing. When one of its most serious alerts goes off, whoever is on call has thirty minutes to establish that it is a false alarm, and if they cannot, the activity stops. That is a defensible standard and a good deal more useful than a policy document.

The company describes the whole episode as a warning shot, which is an unusually blunt phrase for an organisation writing about its own failure. So the question to take away is this. If one of your agents did something nobody had authorised, how long would it take you to find out, and would anyone be looking?

You may also like...