This timely text presents a comprehensive overview of fault tolerance techniques for highperformance computing (HPC). The text opens with a detailed introduction to the concepts of checkpoint protocols and scheduling algorithms, prediction, replication, silent error detection and correction, together with some applicationspecific techniqu