Medical image segmentation is essential for clinical decision making and disease monitoring, yet current deep-learning approaches are limited by their reliance on imaging data alone and lack of contextual understanding. Vision-language models (VLMs) offer a promising alternative by generating textual annotations, but their dependence on manually crafted prompts and poor adaptation to segmentation tasks constrain their clinical utility. Moreover, these models struggle to generalize across imaging modalities and anatomical regions. This project develops an automated framework that generates segmentation-specific textual descriptions without human-created prompts, improving annotation efficiency and segmentation quality. It further integrates dynamic knowledge-graph reasoning to embed evolving medical expertise into the annotation process, enhancing adaptability across diverse imaging contexts. The approach aims to create robust and generalizable artificial intelligence (AI) tools for real-world clinical use. Broader-impact aspects of the project include the engagement of students through hands-on research, interdisciplinary collaboration, and open-access tools that advance science and education as well as clinical relevance that is ensured through close collaboration with medical experts, allowing the research to address real-world healthcare needs and support translational impact. The project introduces a multimodal framework that leverages the bidirectional relationship be