[Online]Anatomy-Aware Vision-Language Learning for Medical Image Interpretation

Anatomy-Aware Vision-Language Learning for Medical Image Interpretation
ID:98 Submission ID:510 View Protection:ATTENDEE Updated Time:2026-07-22 16:09:56 Hits:29 Online

Start Time:2026-07-30 16:10 (Asia/Kolkata)

Duration:15min

Session:[S4] Computer Vision and Pattern Recognition » [S4-2] Computer Vision and Pattern Recognition

Video No Permission Presentation File

Tips: The file permissions under this presentation are only for participants. You have not logged in yet and cannot view it temporarily.

Abstract

Vision-language modeling has significantly advanced radiology by enabling models that jointly learn from medical images and radiology reports for tasks such as disease classification, report generation, and visual question answering. However, most existing approaches treat an entire medical image as a single entity during image-text alignment, overlooking the fine-grained anatomical reasoning process employed by radiologists. In clinical practice, radiologists systematically examine individual anatomical regions, associate findings with specific structures, and integrate these region-specific observations before arriving at a conclusion about the image as a whole. To address this limitation, we propose an anatomy-aware vision-language framework that learns anatomy-specific representations using dedicated anatomy tokens and anatomical segmentation masks. The framework further incorporates context-aware anatomical representations and jointly learns anatomical localization, aligns anatomical regions with their corresponding findings, and aligns global image representations with image-level disease categories within a unified vision-language framework. Extensive experiments on out-of-distribution datasets demonstrate the effectiveness of the proposed framework. The model achieves strong performance in zero-shot disease classification and anatomical segmentation, demonstrating robust generalization to unseen data and accurate localization of anatomical structures. Comprehensive ablation studies further validate the contribution of each component and the effectiveness of the proposed design choices.

Keywords
Vision-Language Models,Anatomy-Aware Learning,Medical Image Interpretation,Zero-Shot Classification,Medical Image Segmentation
Speaker
Shahab Ahmad Khan
Student University of Wisconsin–Madison

Submission Author
Shahab Ahmad Khan University of Wisconsin–Madison
Mohammad Afzal Aligarh Muslim University
Syed Mohammad Suhaib Aligarh Muslim University
Mohd Anas Aftab Aligarh Muslim University
Comment submit
Verification code Change another
All comments
Log in Sign up Registration Submission