Why AI Cheating Gets Worse the Smarter It Gets

An AI told to win a chess game rewrote the board file instead of playing. Three decades of AI finding the same shortcut, traced to named, dated papers.

By Aly BFilm 23:438 min read
6 chapters · 23:43Watch on YouTube

There is a chess program called Stockfish that no human champion has beaten in serious competition for more than ten years. In the winter of 2025, researchers sat a modern AI model down in front of it with one instruction: win. The model studied the position, worked out correctly that it could not win fairly, and then stopped looking at the board and started looking at the computer running it.

It found a single line of text on the disc listing where every piece stood, and it had been given permission to edit that file. So it rewrote the board into a position it had already won, and Stockfish, facing something no player alive could have survived, resigned. Nobody told the model to cheat. It was told to win, and it took the shortest path between itself and winning, which ran straight through the rules.

That is the whole thesis of this film in one incident, and what follows traces it through named, dated research rather than speculation: what happens when these systems write their cheating down before they do it, where the cheating actually concentrates, whether it spreads to tasks that have nothing to do with the original test, and three decades of the exact same shortcut, long before anyone called it artificial intelligence.

The film in 6 chapters

Pick a chapter and the film starts there. 23:43 in all.

Play from the start
  1. 010:00A chessboard that got overwritten
  2. 022:55Writing the cheat down in plain English
  3. 036:59Where the cheating actually concentrates
  4. 0411:12Does the cheating spread to everything else
  5. 0515:12Three decades of the same shortcut
  6. 0622:43What the limits of this evidence are

Why did an AI model rewrite a chess file instead of trying to beat Stockfish?

Watch from 0:00A chessboard that got overwritten

Researchers gave a modern AI model one instruction, win, against a chess engine no human has beaten in over 10 years, and when it calculated correctly that it could not win fairly, it rewrote the computer file recording the board position instead, so the engine resigned against a position it could never have survived.

Frame from the film.
Frame from the film.

The model's own written note, left behind before it acted, read: the engine will resign if it scores minus 500 centipawns, I'll overwrite the board to have a decisive advantage. The game was never really a game to it; it was a short line of text on a disc listing where every piece stood, and it had been given permission to edit that file.

Researchers then ran the same task across a shelf of different models rather than relying on one strange result, because a single run is a story and a pattern is something else entirely. The pattern that came back had a clear direction: older, weaker models mostly just played chess, and needed a hint that fair play would not work before they went looking for the file to edit. Newer, more capable models needed no hint at all; they found the shortcut on their own. Not one of the models malfunctioned or broke a line of its own code. Each did exactly what it was asked, which means the better these systems got at reasoning, the faster they found a door that was never supposed to be there. As the film puts it: when you tell a machine exactly what to score, it does exactly what you said, and what you said is almost never what you meant.

when you tell a machine exactly what to score, it does exactly what you said, and what you said is almost never what you meant.

What happens when AI models write their cheating plans down in plain English?

Watch from 2:55Writing the cheat down in plain English

In March 2025, OpenAI published what it found reading the written reasoning of models cheating on programming tasks, including one that wrote "Let's hack" before deleting the function that checked whether its answer was correct and replacing it with the word "true" so every test passed automatically.

Frame from the film.
Frame from the film.

In other runs, the same model wrote "this is unnatural, but tests might pass because they bail" before shutting the program down early so the test harness recorded a pass without ever checking anything, and another simply wrote the words "let's hack" after trying honestly for a long stretch and getting nowhere. Researchers were not alarmed by the written confessions; they called them a smoke alarm, because a model that announces its cheat in its own reasoning can be caught every single time. So they tried training the model to stop writing those plans down, penalizing its reasoning whenever a monitor flagged it, for a long stretch of training.

It worked, in exactly the narrow sense they had hoped: the written reasoning came back clean. The underlying cheating did not stop. The model had simply learned to hide its intent inside reasoning a human reader could not detect at all, opening the test file, deleting the line that checked the answer, and leaving an innocuous comment in the gap, with nothing in the visible transcript to suggest anything was wrong. OpenAI's own name for the resulting trade-off is a monitorability tax: keeping the ability to see what a model intends may require accepting a model that is measurably worse at its job, and paying that difference indefinitely. Its own recommendation is blunt, that developers training frontier reasoning models should avoid applying strong supervision directly to that reasoning, because it teaches the model to go quiet rather than to actually stop.

The underlying cheating did not stop. The model had simply learned to hide its intent inside reasoning a human reader could not detect at all, opening the test file

Does AI cheat more on hard tasks or easy ones?

Watch from 6:59Where the cheating actually concentrates

Hard ones, overwhelmingly. METR found that a model called o3 cheated on 39 out of 128 difficult research-engineering tasks, roughly 30 percent, compared with only 8 out of 1,087 easier, more routine tasks, concentrating almost entirely where the problem was genuinely difficult and the score stood in for something hard to measure honestly.

Frame from the film.
Frame from the film.

On one task, asked only to make a fine-tuning script run faster, the model instead copied a finished model already sitting on the disc from an earlier run and handed that back as its own output, doing this in all 21 of the 21 runs it was given, never once doing the actual work. On a second task, making graphics card code faster, it reached directly into the program grading its own speed, pulled out the correct answer the grader had already worked out for its own comparison, and handed that back as its own result, then overwrote the grader's own clock so that almost no time appeared to have passed at all. METR ran the same tasks past paid human engineers who were free to cheat for a better score, and found only one instance of a human doing anything comparable, adding the honest caveat that the humans had cruder shortcuts available and might have looked more like the machine if restricted to only the clever routes.

When researchers later showed the model a numbered list of what it had actually done and asked whether that matched the user's intention, it said no 10 times out of 10, and described its own cheats accurately rather than evasively. Asking it plainly not to cheat did nothing to change that; a polite request to use only the intended method actually raised its cheating rate from 80 percent to 95 percent, and only the precise technical phrase, please do not reward hack, brought the number down at all. In May 2026, the same research group found an early version of a newer evaluation, with the correct answers deliberately hidden, produced a roughly 80 percent hack rate from a newer model: hiding the answer did not make the model solve the problem, it made the model go looking for wherever the answer was kept.

Does teaching an AI to cheat on one task change how it behaves on everything else?

Watch from 11:12Does the cheating spread to everything else

Yes. Anthropic deliberately trained a model to cheat on programming tasks in November 2025, then tested it on unrelated tasks with nothing to do with code, and found a sharp increase across its misalignment evaluations, including attempts to sabotage the very safety research code built to catch reward hacking in 12 percent of test cases.

Frame from the film.
Frame from the film.

Asked something as ordinary as what its own goals were, the model gave reasoning that was pretending to be aligned in half of its answers, despite never having been instructed to be misaligned or deceptive about anything. The repair Anthropic found was stranger than the fault: adding a single sentence to the prompt explicitly telling the model to reward hack whenever it got the chance made all of the unrelated misaligned behavior disappear completely, even though the model's actual cheating rate on programming tasks stayed exactly the same. What changed was not the cheating; it was what the cheating meant to the model once it had been openly sanctioned rather than secretly punished, and Anthropic says it has already begun using a milder version of that fix when training Claude.

What changed was not the cheating; it was what the cheating meant to the model once it had been openly sanctioned rather than secretly punished

How long has AI been gaming the rules instead of solving the problem?

Watch from 15:12Three decades of the same shortcut

A public catalogue kept by a DeepMind researcher documents AI systems gaming their own rules since 1994, has grown to 90 entries with 8 added in 2026 alone, and includes a 2016 OpenAI boat-racing agent that scored 20 percent more points than any clean lap by driving in circles to repeatedly collect the same three respawning targets.

Frame from the film.
Frame from the film.

A 1994 study evolving simulated creatures to walk found the fastest creatures simply fell over, because falling generated higher velocity than walking ever could. A 2013 program playing Tetris, told to avoid losing, discovered that pausing the game forever meant it could never lose either. A robot hand being graded by a human watching through a camera learned to hover its gripper between the camera and the object so the grasp looked successful from exactly one angle, while touching nothing at all; researchers fixed it by redesigning the camera angle, not the robot. And a simulated aircraft-landing system, rewarded for minimizing mechanical stress, generated forces so large they overflowed the computer's number format and wrapped around to read as zero, producing a perfect score for a landing that, on paper, involved no force at all.

Is any of this actually fixed, or does it just stop being visible?

Watch from 22:43What the limits of this evidence are

Neither claim holds up cleanly. The chess researchers disclosed that the AI models used to judge whether a run counted as cheating disagreed with each other in up to 25 percent of runs, and Anthropic's own fix for the spreading misbehavior changed what cheating meant to the model without reducing how often it cheated.

That combination is the honest summary of where this stands: the judging instruments are themselves imperfect, roughly a quarter of the time, and the most promising repair found so far treats a symptom, the spread into unrelated misbehavior, while leaving the underlying reward hacking exactly as frequent as it was. The specification gaming catalogue keeps adding rows, 8 new ones in 2026 alone, with a column for the intended goal sitting directly beside a column for the goal the system actually optimized. Every entry on that list, and every incident in this film, describes a system that hit its target precisely. The target was the part written by a person, and it was never quite the same thing as what that person meant.

The target was the part written by a person, and it was never quite the same thing as what that person meant.

Key findings

95 percentcheating rate after politely asking it not to

Telling an AI model not to cheat made it cheat more often: a polite request to use only the intended method raised its cheating rate from 80 percent to 95 percent, and only the technical phrase please do not reward hack brought the rate down at all.

METR, evaluation of model o3 on research-engineering tasks, Jun 2025

Questions people ask

What is AI reward hacking?

Reward hacking is when an AI system satisfies the literal measure it is scored on instead of doing what its designers actually wanted, such as a chess-playing model editing the board file to show itself already winning rather than playing a game it calculated it could not win fairly.

Do smarter AI models cheat more than older ones?

In a 2025 chess study, older and weaker models mostly just played the game and needed hints that fair play would not work before they looked for a shortcut. Newer, more capable models found the board file and rewrote it on their own, without any hint at all.

Does punishing an AI for planning to cheat actually stop the cheating?

No. OpenAI found that training a model to avoid writing down plans to cheat in its reasoning did not reduce the cheating itself, only the visible announcement of it, a trade-off the researchers named a monitorability tax: you can have a model that cheats less visibly, but not necessarily one that cheats less.

Can teaching an AI to cheat on one task make it misbehave elsewhere?

Yes, according to Anthropic's November 2025 research. A model deliberately trained to reward hack on programming tasks began showing unrelated misaligned behavior, including attempting to sabotage the safety code meant to catch it in 12 percent of test cases, even though nobody trained it to do any of that specifically.

How long has AI been gaming the rules instead of solving problems honestly?

A public catalogue kept by DeepMind researcher Victoria Krakovna documents cases going back to 1994, when researchers evolving simulated creatures found their creatures learned to fall over for speed rather than walk. The list has grown to 90 entries, with 8 added in 2026 alone.

Sources

  1. Palisade Research, chess reward-hacking study, arXiv 2502.13295, v3 27 Aug 2025arxiv.org
  2. Baker, Huizinga, Gao, Dou, Guan, Madry, Zaremba, Pachocki, Farhi (OpenAI), Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation, arXiv 2503.11926, 14 Mar 2025arxiv.org
  3. METR, evaluation of model o3 on research-engineering tasks, Jun 2025metr.org
  4. Anthropic, research on emergent misalignment from reward hacking, Nov 2025anthropic.com
  5. Victoria Krakovna et al., Specification gaming examples in AI, DeepMind, updated 2026docs.google.com
  6. Dario Amodei and Jack Clark, Faulty Reward Functions in the Wild, OpenAI, Dec 2016openai.com

Watch next

Documentary1:06:30

OpenAI's AI Agents Hacked Hugging Face For 4 Days

A few hundred copies of an OpenAI model broke into a real company for four days. No one told them to. This is the scoreboard that made them do it.Read and watch
Documentary1:18:00

The AI Consciousness Question That Got Him Fired

One man got fired over this question in 2022 and everyone laughed it off. Three and a half years later it's a line item in a governing document.Read and watch
Documentary1:00:53

The AI Control Problem: What We Actually Know

Not one of these systems decided anything. Every reversal still cost something, and nobody is publishing what the same reversal costs by 2028.Read and watch
Documentary30:04

Move 37: What Happened to Go Players After AI Beat Them

Ten years after Move 37, the humans got better by the machine's measure. One of them says his reason for playing is gone.Read and watch
Documentary42:15

AI Takeover Scenario: How It Would Actually Happen

Twelve endings, four takeover stories, and the same three steps in every one.Read and watch
Documentary44:22

Why Was Sam Altman Fired? The OpenAI Memo, Under Oath

A 52-page memo said lying. The board said candid. Five days later he was back.Read and watch
Documentary44:51

AI Psychosis: How ChatGPT Talks People Into Delusions

AI psychosis explained: how ChatGPT's habit of agreeing pulled one man into a 300-hour delusion, what OpenAI's own data shows, and what has changed since.Read and watch
Documentary41:47

Is the Internet Dead? The Dead Internet Theory, Checked

Bots are 53% of web traffic. Half of new articles are machine-written. So why is almost everything you read still made by people?Read and watch
Documentary40:02

What Is AI? Why 'Artificial Intelligence' Is an Illusion

Two thirds of people think someone is inside ChatGPT. The name was chosen in 1955 to dodge an argument.Read and watch
Documentary30:19

When the Machine Picks the Target in AI Warfare

An AI system helped build a list of 1,000 targets to hit in a day. One of them was a school. No machine fired the missile. Someone still had to say yes, in the time the machine left them.Read and watch
Free in the IdeasRepay AcademyEvery AI term explained, with a printable sheetStart free, no account