This project’s goal is to improve the robustness of machine learning models used to detect malicious software, or “malware”. Computers are constantly vulnerable to malware, which can steal sensitive information, disrupt operations, or damage systems. Detecting and preventing malware is a significant challenge, and although recent advancements in deep learning have made it easier to identify potential malware threats, many deep learning models operate as "black boxes" that provide little information about how they work. This lack of information makes it hard to understand what about the software caused the model to flag it as malware or not. This in turn may make extra work for software developers who have to check legitimate software flagged as malware, or assess the risks posed by new, evolving malware that the models can’t yet detect. Therefore, improving the interpretability and reliability of malware detection systems is crucial to efficiently making software safer. This project aims to develop a robust framework for improving malware detection by identifying the key features learned by deep learning models, then generating new malware samples that might fool models with respect to those features. First, the project seeks to understand what features are learned and extracted by different neural network architectures, particularly those designed for image, sequence, and graph-based inputs. The goal is to gain insights into the representations and decision-making process