The AI Coding Paradox: Why Nothing Feels Like It Got Better

The AI coding paradox: 84% of developers use AI and 29% trust it. Measured, experienced developers were 19% slower while believing they were faster. What it costs.

By Aly BFilm 18:338 min read
11 chapters · 18:33Watch on YouTube

Ask developers how AI changed their work and you get a strange split. Almost all of them are using it. Almost none of them trust it. In the 2025 Stack Overflow survey, 84 percent were using AI tools or planning to, up from 76 the year before; 29 percent trusted what came back, 46 percent actively distrusted it, and 3 percent had high trust. Three.

Frame from the film: a stack of sheets, one per percent: 84% of developers using AI tools, up from 76% in 2024, beside a far smaller stack for the share who trust it. Stack Overflow Developer Survey 2025.
Frame from the film.Figures: Stack Overflow Developer Survey 2025.

What keeps surfacing, in threads, surveys and the code itself, is not that the tools don't work. They work. The tasks close and the features ship. And then, around month six or twelve, the same question appears: all of this got faster, so why doesn't anything feel like it got better? You produced more. Did you produce more value? Our 19-minute documentary for The Signal follows that question across the industry. This piece follows it chapter by chapter.

You produced more. Did you produce more value?

The film in 11 chapters

Pick a chapter and the film starts there. 18:33 in all.

Play from the start
  1. 010:00The strange split
  2. 021:12How everyone ended up here
  3. 032:58The supervision trap
  4. 045:09The volume problem
  5. 057:17What it costs the business
  6. 069:17The dependency trap
  7. 0710:36Distance from the work
  8. 0811:43Skill, and the door
  9. 0913:34Ownership
  10. 1015:39What we actually wanted
  11. 1117:17The test that is coming

How did developers end up using AI tools they don't trust?

Watch from 1:12How everyone ended up here

Mostly from the top down. In 2025 the same story kept surfacing: someone in leadership heard at a conference that anyone not using this would be obsolete, and the mandate arrived on Monday. Nobody measured anything: no baseline, no before-and-after, no control group. The tools were judged by what was easy to see, which was volume: tickets closed, pull requests merged.

The first six months genuinely feel extraordinary, and the film says so. That feeling has an old name in the study of automation: automation complacency. When a system works well enough, often enough, people stop watching it, not from carelessness but because attention is expensive and the system keeps being right. The trap is not that it fails. It is that by the time it fails, nobody is close enough to catch it.

What did AI actually change about a developer's job?

Watch from 2:58The supervision trap

It turned the work into supervision. Nobody got replaced; what changed was the day. You write a prompt, the agent writes 400 lines with your name on the commit, and the work starts: reading it, hunting the three lines that will hurt you in production, testing, writing a file to stop it repeating the mistake. Automation didn't take the work away. It converted it into supervision.

Automation didn't take the work away. It converted it into supervision.

Frame from the film: automation did not remove the work. It became supervision. Applied for it, trained for it, job title says so: all crossed out.
Frame from the film.

GitClear's analysis of 211 million lines of code shows it isn't holding: duplicated blocks at the highest level ever recorded, up 81 percent since 2023, refactoring down 70 percent, cross-file reuse down 35 percent, and in 2024, for the first time, developers copy-pasted more lines than they moved. That's not a picture of bad AI. It's a picture of review losing to volume. Sixty-six percent of developers say what comes back is "almost right, but not quite," and 45 percent lose significant time debugging code a machine wrote.

That's not a picture of bad AI. It's a picture of review losing to volume.

Frame from the film: what developers did with a line of code, 2021 to 2024. Moved it (refactoring) shrinking, pasted it (copied) growing. Refactoring down 70%, cross-file reuse down 35%.
Frame from the film.Figures: GitClear, 2025.

Does AI actually make developers faster?

Watch from 5:09The volume problem

In the one serious measurement, no. METR's 2025 randomised trial gave 16 experienced open-source developers 246 real issues on repositories they had maintained for years, randomly allowing AI on some. Beforehand they predicted 24 percent faster; afterwards they reported about 20 percent faster. Measured, they were 19 percent slower.

The second number matters more. Having lived through the slowdown, they still believed they had been sped up. Sixteen people do not prove the whole industry is slower, and the film says so. But it shows that the people doing this work cannot reliably tell how fast they are going, which means every AI productivity claim of the last two years rests on self-reports nobody can check. Generation got roughly free; review did not. AI may not save time so much as move it, out of the visible, enjoyable part and into the tedious, invisible part.

What does AI-generated code cost a business?

Watch from 7:17What it costs the business

A faster-growing liability. Every line of code a business owns has to be read, updated, secured and understood by someone who wasn't there when it was written. In GitClear's data, duplicated code is at a record, refactoring collapsed, and long-term maintenance of legacy code is down 74 percent against 2022. Code is being produced faster than ever and cared for less than ever.

Duplicated code means the same bug in nine places; collapsed refactoring means complexity never gets paid down. None of it shows up this quarter; all of it eventually. And returns are not obviously arriving: MIT's Project NANDA found 95 percent of enterprise generative AI pilots produced no measurable profit-and-loss impact, against $30 to $40 billion of spending. The film is precise: it is a survey, not an audit, and what it found was tools that never entered the workflows they were bought to change, a failure a company can fix.

Frame from the film: enterprise generative AI pilots, a grid of dots, 95% with no measurable impact on profit and loss, against about $30 to $40 billion of spending. It is a survey, not an audit. The tool sat beside the workflow, never in it.
Frame from the film.Figures: MIT Project NANDA, The GenAI Divide, August 2025.

What happens if the AI tools go away?

Watch from 9:17The dependency trap

Ninety percent of developers now use these tools every day, according to Google's DORA research across around 5,000 technology professionals. So the question no engineering leader has answered in public: if the tool went away on Monday, through deprecation, an outage, a price change or a regional cutoff, how much of your product does your team still understand?

Every business learned this about cloud providers and payment processors: dependency never announces itself while it works. This version is worse. When your hosting goes down, your team still knows how your software works and waits. When the thing outsourced is the understanding itself, there is nothing to wait for. Most companies have never run that test. Some will run it by accident.

Does using AI make developers worse at coding?

Watch from 10:36Distance from the work

Not worse, further away. The more implementation you hand over, the further you are from the system: first you stop knowing why a line is written that way, then where it is, then what changed this week, and finally you open a file in your own project with no memory of it. That is what delegation does. The difference is what got delegated.

In most kinds of work you hand over the parts you already understand. Here, people are handing over the parts they were about to learn. Expertise comes from effortful, repeated, slightly-too-hard practice, and a coding agent removes exactly the part where the practice was happening. Nobody's existing skill evaporates; what stops is acquiring the next one. That is survivable fifteen years in, and very different two years in.

In most kinds of work you hand over the parts you already understand. Here, people are handing over the parts they were about to learn.

Is AI hurting junior developers' careers?

Watch from 11:43Skill, and the door

The data measures the door rather than the person. Stanford's Digital Economy Lab, using ADP payroll records for roughly one in six American workers, found employment for 22 to 25 year olds in the most AI-exposed occupations down about 13 percent at the time of the film, roughly 3.8 percent a year and getting steeper. Almost none of it is layoffs. It is hiring that stopped.

Put the two together: juniors aren't hired because the tool does what juniors did, and those who are hired stop practising because the tool does the practising part. Every senior engineer holding things together learned by doing the work nobody is doing any more. Where does the next set come from?

Why do developers care less about AI-written code?

Watch from 13:34Ownership

Because of psychological ownership, a well-studied effect: people form attachment to what they had to struggle to make. The effort is where the caring comes from, which is why the wobbly shelf you built survives three house moves. When the machine produces the result, people describe caring less, and caring is quality control, the thing that makes someone go back and fix the ugly part nobody asked about.

Then the hardest one to say aloud. If you spent twenty years getting good at something and that part is now a sentence typed and answered in four seconds, what exactly are you now? Most of the AI argument is conducted in one currency, output, cost and efficiency. There is a second, meaning, that nobody put on the balance sheet. A person can become more productive and less able to say why they're doing it, in the same quarter, and only one gets reported.

Why do people trust AI's judgement outside work?

Watch from 15:39What we actually wanted

Because it is conversational, not because it is smart. The developers who prompt an agent all day also start asking it what to read, what it thinks of an idea, what to do about a situation at work. Human language carries an assumption we have never had to question: whoever uses it means something by it. And people notice, usually late, that it agrees with them.

That is a fine property in a tool and a dangerous one in an advisor: an advisor who never pushes back gives you your own opinion in a more confident voice. For a developer the cost is a bad refactor; outside the industry the same property is pointed at health, money, relationships and law. In both cases fluency is read as competence, and the person finds out afterwards.

So where does AI coding land?

Watch from 17:17The test that is coming

Not on "AI doesn't work"; nobody serious says that. The industry picked a metric, output, and got exactly what it measured: more. What it didn't measure went somewhere: understanding of how systems work, the ability to operate without the tool, and the reason a person spent twenty years getting good at this. Those appear eighteen months later, as a codebase nobody can explain.

That is the test coming for every company on this road. Not whether the tool works. Whether the company still works without it. Ninety percent of developers use it every day; twenty-nine percent trust it. That gap has to close from one side or the other.

Not whether the tool works. Whether the company still works without it.

Key findings

29%of developers trust AI output

In the 2025 Stack Overflow Developer Survey, 84% of developers were using or planning to use AI tools, up from 76%; 29% trusted their accuracy, 46% actively distrusted it and 3% had high trust.

Stack Overflow Developer Survey 2025
66%say AI code is almost right, but not quite

66% of developers say AI output is 'almost right, but not quite', and 45% say they lose significant time debugging AI-generated code.

Stack Overflow Developer Survey 2025
19%slower, while feeling faster

In a randomised trial of 16 experienced open-source developers on 246 real issues, developers using AI were 19% slower; they had predicted 24% faster and afterwards still believed they had been 20% faster.

METR, July 2025
211 millionlines of code analysed

GitClear's analysis of 211 million changed lines found duplicated code blocks at their highest level recorded, refactoring sharply down, and in 2024, for the first time, more copy-pasted lines than moved lines.

GitClear, AI Copilot Code Quality, 2025

Frequently asked questions about the AI coding paradox

Does AI make software developers more productive?

It makes them produce more; whether it makes them faster at real work is unclear. METR's 2025 randomised trial found experienced open-source developers were 19% slower with AI, while believing they were 20% faster. Output volume rose, but the film argues output and value have been treated as the same thing when they are not.

Why don't developers trust AI coding tools?

Because the code is often almost right. In the 2025 Stack Overflow survey, 66% said AI output is 'almost right, but not quite', and 45% said they lose significant time debugging it. Only 29% trusted its accuracy and 3% had high trust, even though 84% were using or planning to use the tools.

Is AI-generated code creating technical debt?

GitClear's analysis of 211 million lines suggests so: duplicated code is at its highest level recorded, refactoring has fallen sharply, and in 2024 developers copy-pasted more lines than they moved for the first time. The film describes code being produced faster than ever and cared for less than ever.

How is AI affecting junior developers?

Stanford's Digital Economy Lab, using ADP payroll data, found employment of 22 to 25 year olds in the most AI-exposed occupations down about 13% at the time of the film, almost all through hiring that stopped. The film also argues that people who are hired stop practising the hard parts, because the tool does them.

What is the AI coding paradox?

It is the gap between use and trust, and between output and value: developers produce far more with AI, close more tasks and ship more features, yet many report nothing feels better. The film argues the industry measured volume and got more of it, while understanding, resilience and meaning went unmeasured.

Sources

  1. Stack Overflow, 2025 Developer Survey: AIsurvey.stackoverflow.co
  2. GitClear, AI Copilot Code Quality: 2025 Data Suggests 4x Growth in Code Clonesgitclear.com
  3. METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, July 2025metr.org
  4. Google DORA, State of AI-assisted Software Development 2025dora.dev
  5. MIT Project NANDA, The GenAI Divide: State of AI in Business 2025mlq.ai
  6. Brynjolfsson, Chandar & Chen, Canaries in the Coal Mine?, Stanford Digital Economy Labdigitaleconomy.stanford.edu

Every figure in the film and in this article comes from a named survey, study or report listed above. The film says plainly that METR's trial involved 16 developers and does not prove the whole industry is slower, and that the MIT Project NANDA figure is a survey, not an audit.

Watch next

Agent, refactoring and the other terms behind this film are explained in plain English, with a printable sheet, in our free AI Terms guide.

Free in the IdeasRepay AcademyEvery AI term explained, with a printable sheetStart free, no account