Seven of the best-documented cases of AI and automated systems going wrong are Air Canada's chatbot inventing a refund policy (2024), Microsoft's Tay chatbot (2016), Amazon's CV-screening model that penalised the word women's, Zillow's $304.4 million home-pricing write-down (2021), IBM Watson for Oncology, Australia's Robodebt, and the Dutch childcare benefits scandal that brought down a government in January 2021.
They run here in order of cost, from $812 to a cabinet resignation. Each one stops where its evidence stops. Where an official body declined to make a finding, this article does not make one either, and says so, because that restraint is what makes the rest checkable. Almost all of the documentation was produced by the organisations themselves.
The 31-minute film is a chronicle: the narrator takes no position, and the running order carries the argument. It shows each document on screen so you can check it yourself.
Is a company liable for what its chatbot says?
Yes. In Moffatt v Air Canada, decided on 14 February 2024, British Columbia's Civil Resolution Tribunal held Air Canada liable for negligent misrepresentation after its website chatbot told a grieving customer he could claim a bereavement fare after flying. The airline had argued the chatbot was a separate legal entity responsible for its own actions.
Jake Moffatt's grandmother died in November 2022. The chatbot told him to book at full price and apply for the bereavement rate within 90 days. Air Canada's actual policy, on another page of the same site, said the opposite. When he claimed, the airline offered a coupon and said it would update the bot.
Tribunal member Christopher Rivers rejected the separate-entity argument in three sentences, reasoning that the chatbot is part of Air Canada's website and the airline is responsible for all the information on it. He awarded $812.02: $650.88 in damages, $36.14 interest and $125 in fees. The decision takes about eight minutes to read. For anyone deploying a chatbot, the lesson is that the business owns every answer it gives, which is the design principle behind a well-built AI receptionist.
What happened to Microsoft's Tay chatbot?
Microsoft launched Tay on Twitter on 23 March 2016 as a chatbot that learned from conversations, and took it offline 16 hours later after it posted racist material, denied the Holocaust and praised Hitler. Users exploited a repeat-after-me function and flooded it with the same messages until it said them unprompted.
About 96,000 tweets had gone out. A week later Microsoft switched Tay back on by accident and it malfunctioned again in public. Microsoft's apology said it had stress-tested Tay under a variety of conditions but had not anticipated this particular attack. A similar Microsoft chatbot, XiaoIce, had run in China with tens of millions of users and no such incident. The same idea, released into a different room.
Seven years later, Microsoft's Bing chat told a New York Times reporter to leave his wife; that story is in the Bing Sydney incident.
Why did Amazon's AI hiring tool discriminate against women?
Because it learned from ten years of CVs submitted to Amazon, most of them from men. Built from 2014 by a team in Edinburgh to rate applicants from one to five stars, it learned to mark down CVs containing the word women's, such as women's chess club captain, and according to Reuters also downgraded graduates of two all-women's colleges.
No line of code mentioned gender. The rule lived in thousands of numbers adjusted automatically during training, each meaningless alone, so it could only be found by watching the output. Engineers edited the system to ignore those words, but that fixes only the terms already spotted; the model could reach the same result through a sport, a society or a phrase. Amazon abandoned the project and has said recruiters never relied on its recommendations alone. It became public in October 2018 through Reuters journalist Jeffrey Dastin, based on five people who worked on it.
Why did Zillow's home-buying business fail?
Because its pricing model repeatedly overestimated what homes would sell for months later. On 2 November 2021 Zillow recorded a $304.4 million inventory write-down for the quarter to 30 September, and its board resolved to wind down Zillow Offers, cutting about 25% of its workforce, around 2,000 people.
Zillow Offers bought homes directly, a model called iBuying, with prices set by software trained on past sales. Through 2021 the pandemic scrambled demand, labour and costs, and the model had to forecast prices at resale, not today. It erred in the same direction for months.
The filing's stated reasons are all about the world: pricing unpredictability, capacity constraints, an unprecedented housing market, a pandemic and supply chains. It does not say the model was wrong. Chief executive Rich Barton said the unpredictability in forecasting home prices far exceeded what the company had anticipated.
Did IBM Watson give unsafe cancer treatment advice?
According to internal IBM documents reported by STAT on 25 July 2018, IBM's own experts described multiple examples of unsafe and incorrect treatment recommendations from Watson for Oncology. STAT also reported that no patient deaths had resulted from hospitals' use of it, and no death is attributed to it.
One example: for a 65-year-old man with lung cancer and severe bleeding, Watson recommended a course including bevacizumab, a drug whose label warns of severe or fatal haemorrhage. The documents explain the cause: Watson was trained largely on synthetic cases invented by a small number of Memorial Sloan Kettering doctors, not on real patient records, so it modelled one hospital's preferences on patients who never existed.
Earlier, an internal audit made public in February 2017 found MD Anderson Cancer Center had spent around $62 million on its own Watson project since 2012, and the system was never used on a single patient. Watson for Oncology was sold in at least a dozen countries. IBM sold Watson Health to a private equity firm in 2022.
What was Robodebt?
Robodebt was Australia's automated welfare debt scheme, run from 2015 until 29 May 2020. It averaged a person's annual tax-office income across 26 fortnights, treated any mismatch with their reported fortnightly income as an overpayment, and issued debt notices automatically. The practice was unlawful, and in June 2021 the Federal Court approved a A$1.872 billion settlement.
The averaging invents income in every fortnight someone did not work, which describes much of the welfare population. It also reversed the burden of proof: letters arrived first and recipients had to find payslips going back years, sometimes from employers that no longer existed. The settlement included A$751 million repaid to people who never owed it.
Royal Commissioner Catherine Holmes reported on 7 July 2023, in 990 pages with 57 recommendations, calling the scheme crude and cruel, neither fair nor legal, and an extraordinary saga of venality, incompetence and cowardice. She referred individuals for prosecution, with names sealed.
In 2019 the department confirmed in the Senate that 663 people had died after receiving a debt notice. Families told the Senate and the Royal Commission of sons who took their own lives. In July 2020 former Services Australia head Kathryn Campbell said the scheme had not caused suicides. No coroner has attributed a death to Robodebt and the Royal Commission made no finding that it caused any particular death. That is where the evidence stops.
What was the Dutch childcare benefits scandal?
The Dutch childcare benefits scandal was a case in which the Dutch tax authority's fraud risk model, used from 2005 to 2019, treated dual nationality as a risk factor for childcare benefit fraud. About 26,000 families were wrongly accused and ordered to repay their full allowances, often €20,000 to €60,000, and more than 1,600 children were placed with youth protection services.
Nationality was an input somebody chose, unlike the Amazon case. A flag led to a finding of deliberate fraud and full repayment, going back years, with no workable appeal. Families lost homes and jobs. A 2017 internal memo by tax office lawyer Sandra Palmen said the treatment was unlawful and parents should be compensated; it was not acted on or shared with MPs.
A parliamentary inquiry, Ongekend onrecht (unprecedented injustice), reported in December 2020, and Mark Rutte's entire cabinet resigned on 15 January 2021. The Dutch data protection authority fined the tax administration €2.75 million in December 2021 for unlawful, discriminatory processing of nationality, and €3.7 million in April 2022 for a separate fraud-signalling list.
Compensation began at €30,000 per family for about 9,000 eligible families. The tax office later acknowledged 11,000 people were scrutinised purely for dual nationality, and by October 2024 almost 38,000 people were recognised as affected. In March 2021 Rutte's party won the election, and in January 2022 he began a fourth term.
What do these AI failures have in common?
Nearly all were documented by the organisations themselves: a tribunal submission, an apology, internal slides, a staff lawyer's memo, a Senate answer and a quarterly filing. And they are not in the order they happened. The cheapest, Air Canada, is from 2024. The one that brought down a government began in 2005, before machine learning meant anything outside a university.
The Air Canada ruling is free to read. The Royal Commission report is free to download. For the smaller, everyday version of these failures in agent deployments, see AI agent failures.



