Traditional machine learning often involves collecting data from multiple sources, which can raise significant privacy concerns. One approach has emerged as a promising solution to solve this challenge by enabling models to be trained across many different sources without directly sharing private data. This approach has become particularly valuable in sensitive sectors such as medical diagnostics, where individual data privacy is legally protected. Despite these advancements, existing systems for training models across multiple sources lack standardized assessment tools, posing challenges to research reproducibility, validation, and trust. Without proper testing tools, organizations cannot verify that their privacy protections work as intended, creating barriers to adoption in critical areas like healthcare, finance, and national security. This project addresses this challenge by developing comprehensive testing tools that ensure privacy-preserving artificial intelligence systems work reliably, serving the national interest by enabling secure collaboration on AI development while protecting individual privacy, supporting American competitiveness in artificial intelligence technologies, and strengthening data security across critical infrastructure. This project designs, develops, and sustains FLTest, an interdisciplinary testbed that automates privacy and robustness evaluations in federated learning systems, addressing gaps often overlooked by traditional tools. The rese