The Most Expensive Software Bugs in History: How One Line of Code Burned Billions
Bugs that blew up a rocket in 37 seconds, nearly sank a company in 45 minutes and crashed 8.5 million computers. No hackers involved; each was one small missed detail.

Contents 11
The most expensive software bugs in history usually came from a single line: a wrong unit, a number that did not fit, code that someone forgot to deploy to one server. This post walks through the bugs that blew up a rocket in 37 seconds, nearly sank a company in 45 minutes and sent 8.5 million computers to a blue screen on the same morning, based on the official reports.
In short:
- Ariane 5 was destroyed 37 seconds after lift-off because a 64-bit number did not fit into a 16-bit variable.
- Knight Capital lost more than $460 million in 45 minutes because old code was left on one of eight servers.
- The 2024 CrowdStrike update crashed 8.5 million Windows devices; losses for Fortune 500 companies were estimated at $5.4 billion.
- The next big date bug is waiting on 19 January 2038, and it is not fully solved yet.
What these stories have in common: none of them involves a malicious hacker. Each one is a small detail missed by good engineers who were trying to do their job. That is exactly what makes them unsettling.
1. Mars Climate Orbiter (1999): pounds versus newtons
NASA's Mars Climate Orbiter travelled for nine months and went silent on 23 September 1999, the moment it reached Mars. The investigation showed the cause was not as big as space, but as small as a physics homework mistake.
The ground software calculated thruster impulse in pound-force seconds (imperial units). The spacecraft's navigation software expected the same value in newton seconds (metric). Since one pound-force is about 4.45 newtons, every trajectory correction was systematically wrong. The orbiter approached far lower than planned and burned up in the Martian atmosphere.
Lesson: if the contract between two systems (API, data format, unit) is not written down and tested, both sides can work "correctly" and still produce a disaster.
2. Ariane 5 Flight 501 (1996): a 37-second firework
On 4 June 1996, the European Space Agency's new Ariane 5 rocket veered off course about 37 seconds into its maiden flight and self-destructed. It carried four Cluster science satellites. The loss is usually estimated at around $370 million.
According to ESA's Inquiry Board report, the inertial reference system tried to convert a 64-bit floating point value related to horizontal velocity into a 16-bit signed integer. Ariane 5 was much faster than its predecessor Ariane 4, so the value exceeded 32,767 and caused an overflow. The backup system ran exactly the same software and had failed for the same reason a few milliseconds earlier.
The most painful part: this calculation served no purpose after lift-off. The code was carried over from Ariane 4 as it was, and because it had worked fine there, it was never retested against the new rocket's flight profile.
Lesson: "it worked somewhere else" is not a test result. Reused code has to be tested again against the limits of its new environment.
3. Therac-25 (1985–1987): the operator who typed too fast
The darkest example of software bugs comes from a radiation therapy machine. Between 1985 and 1987, Therac-25 gave six patients doses hundreds of times higher than intended; at least three people died as a result.
Nancy Leveson and Clark Turner's classic investigation showed that the problem was a race condition. When an experienced operator corrected the treatment settings quickly, within a few seconds, the software prepared the high-energy beam without noticing the change. The screen only showed a cryptic "Malfunction 54" message, and operators could simply resume treatment.
Earlier models had hardware interlocks that physically prevented this kind of failure. In Therac-25, safety was left entirely to software.
Lesson: in life-critical systems, safety is never left to a single layer. Cryptic error messages are almost as dangerous as the bug itself.
4. The Patriot missile (1991): one tenth of a second
During the Gulf War, on 25 February 1991, a Patriot air defence battery in Dhahran, Saudi Arabia, failed to intercept an incoming Scud missile. The missile hit a US Army barracks and killed 28 soldiers.
The US Government Accountability Office (GAO) report explained why. The system counted time in tenths of a second. But 0.1 cannot be represented exactly in binary, just as 1/3 becomes 0.333… forever in decimal. Because the system used a 24-bit register, each tick added a tiny rounding error. The battery had been running for about 100 hours without a restart, and the accumulated drift had reached 0.34 seconds. For a missile travelling faster than 1.6 kilometres per second, that meant the radar was looking in the wrong place.
The software fix reached Dhahran the day after the attack.
Lesson: floating point numbers are not real numbers. In long-running systems, tiny errors grow silently.
5. Knight Capital (2012): on the edge of bankruptcy in 45 minutes
On the morning of 1 August 2012, the automated trading system of Knight Capital, one of the largest market makers in the US, went haywire. According to the US Securities and Exchange Commission (SEC), while trying to fill just 212 customer orders, the system sent more than 4 million orders into the market, traded 397 million shares, and the firm lost more than $460 million. Knight had to be rescued shortly afterwards and later merged with a rival.
The cause was a deployment mistake. New code for a new exchange programme was deployed to only seven of eight servers. The eighth server still contained an old function that had not been used for years, and the new code reused a flag that now triggered that old function. The system sent 97 warning e-mails before the market opened; nobody treated them as an alarm.
Lesson: dead code is not dead, it is only sleeping. Automated, repeatable deployments are always safer than "everyone copies files by hand".
6. left-pad (2016): 11 lines that broke the internet
On 22 March 2016, a developer unpublished all of his packages from npm after a naming dispute. One of them was left-pad, a package of roughly 11 lines that pads a string on the left.
The problem was that left-pad sat, directly or indirectly, in the dependency tree of huge projects such as Babel and React. JavaScript builds around the world suddenly failed. npm restored the package within hours and then introduced new rules that sharply restrict unpublishing packages.
Lesson: your application is not only the code you write. Every dependency is a risk outside your control; that is why lockfiles and dependency audits exist.
7. CrowdStrike (2024): the world's biggest blue screen
On the morning of 19 July 2024, flights were grounded, hospitals cancelled appointments, and banks and TV channels went offline. The cause was a routine content update that the security company CrowdStrike pushed to its Falcon software.
Microsoft estimated that the update crashed 8.5 million Windows devices. According to CrowdStrike's root cause analysis, the new template expected 21 input fields but the sensor supplied only 20. The software tried to read the missing 21st field, read outside its memory, and because it ran in the operating system kernel, it took the whole computer down with it. The faulty file was pulled after about 78 minutes, but most machines had to be recovered one by one by hand.
The insurance analytics firm Parametrix estimated the direct loss for US Fortune 500 companies alone at $5.4 billion.
| CrowdStrike (2024, Fortune 500) | 5400 million $ |
|---|---|
| Knight Capital (2012) | 460 million $ |
| Ariane 5 (1996) | 370 million $ |
| Mars Climate Orbiter (1999) | 125 million $ |
Kaynak: SEC, ESA, NASA and Parametrix reports
Lesson: staged rollouts (a small group of devices first, then everyone) are not a luxury, they are insurance. No matter how "small" an update looks, it should never go to every device at once.
The next big bug: 19 January 2038
Many systems store time as the number of seconds since 1 January 1970 in a signed 32-bit integer. The largest value that integer can hold is 2,147,483,647, and that second arrives on 19 January 2038 at 03:14:07 UTC. One second later the counter wraps to a negative number and those systems think it is 13 December 1901.
Modern operating systems and languages moved to 64-bit time long ago. But embedded devices, old database columns, industrial control systems and code that has not been touched in years are still at risk. The Y2K problem did not turn into a major crisis thanks to billions of dollars of preventive work. 2038 needs similar preparation, and it is close enough that much of the code written today will still be running.
5 practical rules from these bugs
The most expensive software bugs in history leave lessons every team can apply today:
- Test the boundaries. Largest value, smallest value, empty value, overflowing value. Ariane 5 and 2038 are two faces of the same bug.
- Write down units and contracts. Every field exchanged between systems needs a defined type, unit and format.
- Automate deployment. Every manual step is a candidate to become Knight Capital's eighth server.
- Roll out in stages and make rollback easy. Send updates to a small group first; when something breaks, go back with one click.
- Make alerts meaningful. "Malfunction 54" or 97 unread e-mails are no better than no alert at all.
At EngerekTech, these stories are exactly why we set up automated tests, CI/CD pipelines and staged releases from day one in our web, mobile and enterprise software projects.
Frequently asked questions
What is the most expensive software bug in history?
In terms of measurable direct losses, the 2024 CrowdStrike outage stands out: losses for US Fortune 500 companies alone were estimated at $5.4 billion. As a loss suffered by a single company in a very short time, Knight Capital's $460 million in 45 minutes is one of the most striking examples.
Why do software bugs cause such huge losses?
Because software can repeat the same mistake millions of times per second and on thousands of devices at once. A human error affects one transaction; an error in an automated system affects all of them.
Is the year 2038 problem a real danger?
Current computers and phones are largely safe because they store time as 64-bit numbers. The risk lies in old embedded devices, databases with 32-bit time columns and systems that have not been updated for a long time. These need to be audited and updated before 2038.
How can these kinds of bugs be prevented?
There is no single method; it takes layers: boundary tests, code review, automated deployment, staged rollouts, monitoring and meaningful alerts. Most bugs turn into disasters because they slip past several safeguards at once, not just one.


