Data centers (DC) are the backbone of the modern digital economy, critical to the U.S. economical growth, national security, public health, and enhanced data security and management. DCs are an energy intensive infrastructure, accounting for over 4% of total electricity use worldwide. As demand for AI and cloud computing grows, efficient cooling systems are critical to ensuring reliable and resilient DC operations. A critical failure in DC cooling systems can have catastrophic consequences, including total system shutdown, loss of data, and IT equipment. To prevent such catastrophic events, novel Fault Detection and Diagnostics (FDD) and mitigation techniques are essential. Currently, most FDD methods rely on conventional statistical techniques, machine learning models, or ad-hoc estimations. However, these methods are often limited in scope and may fail to detect rare or complex failure scenarios – particularly those arising from complex cascading events or malicious cyber-attacks. To tackle this challenge, this project develops a new FDD method based on failure and cyber-attack detection in supervisory control theory of discrete event systems. The intellectual merits of this project are: (1) new FDD methods for detecting and mitigating cascading faults and cyber-attacks resulting in resilient DC cooling system operation, (2) an open-source virtual testbed for evaluating performance of the proposed algorithms, and (3) a hardware-in-the-loop testbed to understand the challen