Surabhi Jain supervised by Dr. Radhika Mamidi received her Master of Science – Dual Degree in Computer Science & Engineering (LCD) by Research. Here’s a summary of her research work on Characterizing Learning Behavior in Optical Recognition Models for Indic Script Recognition
Modern Optical character recognition (OCR) systems increasingly rely on deep visual sequence models, yet they are often evaluated mainly through final recognition accuracy. This leaves an important question underexplored: how do such models learn, stabilize, and exhibit persistent errors over the course of training? This question is especially relevant for Indic scripts, where visual text recognition is shaped by complex orthographic structures such as aksharas, modifiers, conjunct-like forms, and dense character compositions. This thesis studies the learning behavior of two representative OCR models, CRNN and PARSeq, using a five-language corpus covering Hindi, Bengali, Tamil, Telugu, and English. The main chapters present representative analyses, and the same qualitative trends were observed in the remaining languages. The experiments are conducted on cropped printed word images, allowing the analysis to focus on recognition behavior at the word level. Rather than comparing models only by final CER and WER, the thesis examines their behavior across four complementary dimensions: training data scale, sequence length, subword structure, and optimization geometry. The results show that learning behavior varies meaningfully across model family, data regime, sequence length, and script structure. Under the tested settings, CRNN is more sensitive to reductions in training data, while PARSeq remains comparatively more stable. Short words are learned faster than long words, showing that sequence length remains an important source of recognition difficulty. Frequent and structurally simpler subword patterns stabilize earlier than rarer and more compositionally involved patterns. Loss landscape visualizations further suggest that training moves from sharper and more irregular regions toward smoother and more stable basins, with fuller data settings reaching such regions more readily. Together, these findings show that final recognition scores alone do not fully characterize OCR model behavior. The analysis can support more informed model selection, especially when data availability, script complexity, and word length vary across recognition settings. The findings may also inform dataset design by indicating where additional examples are likely to be useful. By focusing on learning trajectories, this thesis offers a structured way to evaluate OCR models beyond final accuracy, with attention to how their behavior changes with data availability, word length, script structure, and training stability.
July 2026

