Srija Mukhopadhyay supervised by Dr. Manish Shrivastava received her Master of Science – Dual Degree in Computational Linguistics (CLD). Here’s a summary of her research work on When Models Miss the Plot: Evaluating Reliability in Multimodal Reasoning over Charts, Grids and Maps
Multimodal models are increasingly presented as general-purpose analysts of visual information. A user can ask a model to inspect a chart, compare several plots, read a thematic map, or choose an action from a visual state, and get seemingly correct answers. We study why such answers can look convincing without being reliable. The central argument of the thesis is that visual reasoning should be evaluated by an evidence standard. A model should not only produce the right answer, but instead it should produce the right answer because it used the relevant visual evidence, and it should continue to do so when that evidence becomes harder to extract, changes visual form, conflicts with prior knowledge, or is distributed across views. Aggregate accuracy is useful for comparison, but it hides the route by which an answer was produced. A correct answer may come from reading the display, but it may also come from question wording, dataset regularities, familiar rendering styles, or memorized world knowledge and these routes can be difficult to disentangle. We develop this argument through a sequence of diagnostic evaluations. ChartQA-Split and RobustCQA show that chart QA scores conceal large differences across chart and question complexity, that semantics-preserving perturbations can sharply change model performance, and that blank or irrelevant-image controls expose non-visual shortcuts. InterChart moves to multi-chart reasoning, where models must align and integrate evidence across related visualizations. MAPWise and MapIQ extend the evaluation to thematic maps, showing that models remain far below human readers, are sensitive to map design, and often fail when the rendered map conflicts with memorized geography. Finally, Rubicon moves from visual interpretation to visually grounded planning, showing that several models can recognize unsolvable states when asked directly but still lack the capability of implicitly detecting bottlenecks, thereby proving to be unreliable planners. Across charts, maps, and grid worlds, we notice that multimodal models have real but unevenly grounded capabilities. They often look competent under favorable conditions while failing when probed under more challenging scenarios. The thesis contributes new evaluation resources and argues that progress should be judged by through diagnostic evaluations that probe the evidence used by models, rather than by aggregate accuracy alone. We hope that this work will help to clarify the capabilities and limitations of multimodal models, and to guide future research towards more robust and reliable visual reasoning systems
June 2026

