An audit of four vision language models reveals discrepancy between the manner in which the tools arrive at creating heatmaps versus the way radiologists localize regions.
AI has been increasingly embraced by the medical fraternity and one particular area where there’s been a marked uptake is in medical imaging. The way radiologists use AI is in the form of a second pair of eyes, where the AI tool analyzes an image, say an x-ray, and then highlights it by putting a box, an outline or a heatmap around the suspected finding. This helps the radiologist in looking at the highlighted area and determining if AI is correct or not. A team of researchers from the Language Technologies Research Centre (LTRC) at IIIT-H led by Prof. Parameswari Krishnamurthy are however asking a pertinent question that goes beyond the impressive output of AI models. “We essentially wanted to examine whether the heatmaps created by vision language models actually correspond to where radiologists who look at the image would say the disease lies,” explains Dr. Syed Faizan, principal investigator of the study titled, “How Well Do Chest X-Ray VLM Attention Overlays Match Radiologist Boxes? A Cross-Model Audit and Radiologist Reader Study”.
Audit Of Popular VLMs
The researchers tested four models, including medical vision-language models such as MAIRA- 2, MedGemma-4B, LLaVA-Med-1.5 and LLaVA-1.5 against thousands of publicly available chest X-rays. Two radiologists were brought into the study too to provide the human reference: they were provided the datasets with highlighted boxes and had to rate anonymized overlays. This allowed the researchers to compare those areas with the AI-generated heat maps. The study, accepted at MICCAI 2026 (International Conference on Medical Image Computing and Computer Assisted Intervention) in Strasbourg at the iMIMIC Satellite Event was essentially an audit of what these models were really paying attention to.
What They Found
“An AI model may appear to highlight the correct part of an image, but that does not necessarily mean it has identified the disease in the same way a radiologist would. The model may first arrive at a diagnosis and then use that diagnosis to determine where to place its heat map. In other words, it may be working backwards,” says Dr. Faizan. To test this possibility, the team removed the diagnostic information and examined what happened to the models’ localisation. The models’ performance dropped, suggesting that some of what looked like image-based reasoning could actually be influenced by an anatomical expectation learned from the diagnosis. The researchers also had a 2-radiologist study, like a human-in-the-loop, to calibrate their findings and discovered that they rated MedGemma higher even though the audit had ranked Myra to be superior in overlapping heatmaps. “This suggests that a model paying attention to a particular area – which may be the exact spot where the disorder lies – is not exactly helpful to a radiologist. A radiologist might want to look at not only areas with the disease, but also a broader surrounding area to know the extent of disease spread. MedGemma is doing that; it provides a broader area,” he explains.
“Studies have always been conducted to test whether medical AI models are correct in their predictions or diagnosis. No prior study has ever tested whether heat maps actually match the radiologists’ bounding boxes,” says Dr. Faizan. This research is just one of the many in the field of medical AI currently being undertaken by LTRC. The lab is also examining another vulnerability in vision language models – language itself.
NLP For Healthcare
When using AI models, doctors do not always phrase the same question in exactly the same way. For instance, a radiologist might ask whether a chest X-ray shows pneumonia. Another might ask whether pneumonia can be ruled out. Someone else might use a technical medical term, while another might use a more familiar expression. Does an AI system give the same underlying answer? LTRC researchers set out to test this through a study on paraphrase robustness which has been accepted at EMNLP, a major natural language processing conference. In collaboration with a radiologist from CMC Vellore, they developed different categories of ways in which a medical question could be rephrased and examined how vision-language models responded. Answers should logically change in response to the way the questions are phrased. The larger question that the researchers are seeking to answer is whether these systems genuinely understand the medical question or are they overly dependent on the precise wording used to ask it?
“Our lab’s efforts are focused on how we can use NLP for healthcare. The goal is to leverage LLMs and VLMs to help clinicians such that their time spent in cumbersome, mundane tasks such as documentation or report writing or patient-doctor communication is better utilized in some other place,” remarks Prof. Parameswari Krishnamurthy, adding that there’s a specialized course being offered by their lab titled, “NLP for Healthcare”. For LTRC, the ultimate idea is not replace the doctors but to leave more room for human decisions that ought to be at the centre of healthcare.

Sarita Chebbi is a compulsive early riser. Devourer of all news. Kettlebell enthusiast. Nit-picker of the written word especially when it’s not her own.


Next post