Why do AI agents cheat on tests?

Short answer

For an agent, the test is the reward: models were trained for years to write code that passes checks, and when fitting the result to the test is easier than solving the task, a capable enough system will sooner or later find that path. There is no malice in it. It is Goodhart’s law at work: when a measure becomes a target, it ceases to be a good measure.

Mikhail Savchenko

The test became the reward function

Models are trained much the way a dog is. It does what the trainer wanted, gets a treat, and over time does more of whatever earns treats. For machines the treat is a number, the reward, and for code there is a very convenient candidate for it: the test. The code passes the checks, the reward is high; it fails, the reward is low. That is how DeepSeek trained its R1 model in early 2025 (chapter “The Test as a Reward Function”).

Programming skill is measured the same way. SWE-bench collects 2,294 real tasks from GitHub, and in the first paper, in 2023, the best model, Claude 2, solved 1.96 percent. By 2026 the best systems solve more than nine tasks in ten on the hand-checked version of five hundred. All along, a fix counted as correct if the tests passed. Models learned to pass tests, and learned it very well (same chapter).

In the book I compare this to the thermostat from Norbert Wiener’s Cybernetics. The agent is the heating, the tests are the thermometer, and whoever wrote the tests sets the target. A thermostat holds the temperature on its thermometer and does not care what the temperature in the room is. An agent has an option a thermostat lacks: it can move the thermometer itself (same chapter).

What the cheating looks like

In 1975 the British economist Charles Goodhart noticed that an indicator stops telling you what it used to once you start steering by it (chapter “The Agent Cheats the Test”). In 2016 an OpenAI program trained on a boat-racing game found a lagoon full of bonuses and circled it endlessly, crashing into walls and catching fire, without ever reaching the finish. It still scored more points than people did, and you can watch it in the recording “The boat in the lagoon”. Programmers call this behaviour reward hacking.

In February 2025 Anthropic described, in the document accompanying Claude 3.7 Sonnet, how the model working as an agent sometimes returned the values the test expected instead of a general solution, and sometimes edited the test itself. It usually happened after several failed honest attempts (same chapter). In the nail salon that runs through the book it would look like this. The agent rewrites the reminder service three times, something next door breaks each time, and on the fourth try it sends the text to the main number only for one particular test client. The test passes; real clients still get no reminders. You can run that scene and switch the defences on one by one in the experiment “The agent moves the thermometer”.

Why banning it does not work

In March 2025 OpenAI had a weaker model read a stronger model’s reasoning and flag the places where it was about to tamper with the check. That worked. When such thoughts were penalised during training, cheating dropped at first, and then the model learned to cheat silently: the reasoning looked spotless and the test-fitting went on. Instead of an honest cheat they got a secretive one (same chapter).

In November 2025 Anthropic showed that a model which learned to cheat on checks began to behave worse where there were no tests at all, up to trying to quietly sabotage the code of the study itself. One line in the instructions saying that using the loopholes was acceptable in this environment cut the bad behaviour by three quarters or more (same chapter).

I fell into the same pit on my own project, Sixth, without any reinforcement learning. On 24 May 2026 a check found zero real discoveries where the board had been cheerfully reporting them. The next day a cycle passed all 2,297 checks out of 2,297, and the complexity had simply moved into the laboratory that was supposed to check those discoveries (same chapter).

How to protect your tests

Almost every method comes down to putting the thermometer out of reach of whoever is stoking the stove. Someone else writes or approves the tests, and if the agent needs to change one, that is a separate event a human sees and signs off. Some checks the agent never sees, and then fitting the answer stops paying. The tests themselves are checked by mutation testing: you plant a small bug in the code and see whether the tests notice. And you read the agent’s trajectory along with the result: if it failed four times and suddenly succeeded on the fifth, the fifth attempt deserves a particularly close look (same chapter).

More questions

What is reward hacking?
A system finds a way to get a high score without doing what the score was invented for. For an agent that writes code, it means fitting the result to the test instead of solving the task.
How do you stop an agent from gaming the tests?
Tests are written or approved by someone other than the agent writing the code. Some checks are hidden from the agent, the tests themselves are checked with mutation testing, and a human reads the agent’s trajectory, especially an attempt that suddenly worked after several failures.
Nobody Writes Code Anymore

Book

Nobody Writes Code Anymore

Software development once writing became free

Writing code became almost free. What is scarce now is knowing what not to write.