Knowledge acquisition—the ability of artificial intelligence (AI) systems to extract actionable insights from vast amounts of unstructured text—is critical for advancements in healthcare, education, and scientific discovery. While Large Language Models (LLMs) have shown impressive capabilities, their reliability depends heavily on massive, perfectly curated datasets, which are expensive and often unavailable in specialized domains. This CAREER project addresses this bottleneck by developing a new paradigm called “structure-aware weak supervision.” Instead of relying on perfect human annotations, the project enables AI systems to learn autonomously from incomplete, noisy, and ambiguous data by discovering and utilizing underlying semantic structures, such as concept hierarchies and retrieval pathways. By reducing the dependency on expensive labeled data, this research democratizes the development of highly accurate, domain-specific AI tools for resource-constrained environments, such as public health agencies and community organizations. The project also integrates these research outcomes into new undergraduate and graduate curricula, open-source educational toolkits, and targeted K-12 outreach programs designed to broaden participation in computing and teach the next generation how to build reliable, human-centered AI systems. This project proposes a unified framework for learning under weak supervision by bridging unstructured language data with structured, interpretable