Debardeleben and Blanchard’s Supercomputer Work Featured in WIRED Magazine
Researchers Nathan DeBardeleben and Sean Blanchard, of the High Performance Computing Design Group and the Ultra Scale Research Center, are working to ensure that next-generation supercomputers remain reliable despite the constant bombardment of cosmic radiation. Their research, recently featured in WIRED magazine, explores how tiny particles created by cosmic rays can disrupt some of the world’s most powerful computing systems—and how engineers can design computers to withstand them.
High-energy cosmic rays continuously strike Earth’s atmosphere, producing secondary particles known as neutrons. These neutrons can penetrate buildings and electronic systems, occasionally colliding with computer processors and memory chips. The result can be memory errors, corrupted calculations, or even complete system failures.
Scientists have known about this phenomenon for decades. The problem first came to light in the 1970s when legendary computer engineer Seymour Cray loaned one of his pioneering supercomputers to Los Alamos National Laboratory for evaluation. Researchers discovered that particles originating from cosmic radiation could interfere with the machine’s electronics, creating unpredictable errors that were difficult to diagnose.
As supercomputers have become dramatically larger and more powerful, understanding these radiation-induced errors has become increasingly important. Today’s high-performance computing systems perform billions of calculations every second, supporting critical scientific discoveries and national security missions. Even a single bit flip caused by a stray neutron can affect the accuracy of a simulation or interrupt a lengthy computation.
To better understand and prevent these failures, DeBardeleben and Blanchard subject new computing hardware to rigorous testing before it is deployed. Using specialized neutron beam facilities, they expose processors, memory modules, and other electronic components to neutron levels far greater than those found in everyday environments. This accelerated testing compresses years of cosmic-ray exposure into a much shorter period, allowing researchers to identify vulnerabilities long before the hardware is installed in production supercomputers.
Their work also focuses on making computing systems more resilient. One important strategy is the use of checkpoints, which periodically save a system’s state and data during long-running calculations. If an error occurs, the computation can restart from the most recent checkpoint rather than beginning again from scratch, reducing lost time and preserving valuable scientific results. In some cases, systems can even detect certain errors, intentionally pause, and recover before more serious failures occur.
Beyond laboratory testing, DeBardeleben and Blanchard install neutron detectors inside supercomputing facilities to continuously monitor radiation levels during normal operation. These measurements provide valuable insight into the environment experienced by computing hardware and help researchers improve models that predict system reliability and component lifetimes.
As supercomputers continue to grow in scale and capability, ensuring their reliability remains one of high-performance computing’s greatest challenges. Through advanced radiation testing, fault-tolerant system design, and continuous monitoring, DeBardeleben and Blanchard are helping ensure that the world’s most powerful computers can continue producing accurate, dependable results—even when confronted with particles arriving from deep space.
This research is supported by the New Mexico Consortium.
To read more about DeBardeleben and Blanchard’s research on protecting our supercomputers read the full WIRED article: Cosmic Ray Showers Crash Supercomputers. Here’s What to Do About It.
