Weekend computer crashes inspired the invention of error-correcting code
In the 1940s, mathematician Richard Hamming had to share computer time on Bell Labs punch-card mainframes. Over weekends, when operators were absent, machines frequently encountered error-prone read errors and aborted his calculations entirely. Frustrated by wasted compute time, Hamming argued that computers should not only detect errors but fix them automatically. In 1950, he published Hamming codes, enabling computers to silently correct digital data corruption without stopping.
The Monday Morning Problem at Bell Labs
In the late 1940s, Richard Hamming worked as a mathematician at Bell Telephone Laboratories, where he shared access to electromechanical relay computers. These early machines processed numerical calculations using punched paper cards. Because computing time was scarce, daytime hours were closely managed by human operators, while evenings and weekends were reserved for long, unattended batch jobs. Programmers loaded stacks of cards into the hopper on Friday afternoon, expecting completed calculations when they returned on Monday morning.
The relay machines were notoriously prone to physical and mechanical faults. Relay contacts could fail to close, paper cards could suffer minor misreads, or electrical noise could introduce transient glitches. During regular weekday shifts, human operators monitored the computers. When an error lamp lit up, an operator would inspect the failure, reset the machine, or reload the offending card to keep the process moving. Unattended weekend runs enjoyed no such oversight.
Whenever the computer encountered an error without an operator present, its built-in safety mechanism immediately aborted the program or dropped the entire job and moved on to the next one. Hamming repeatedly arrived on Monday mornings to discover that his programs had crashed within minutes of his Friday departure, wasting days of valuable computational time. Frustrated by this recurring failure mode, Hamming reasoned that if a machine had enough information to know that an error had occurred, it ought to have enough information to determine where the error was and fix it automatically.