Anshul Sharma supervised by Dr. Radhika Mamidi received his Master of Science by Research in Computer Science Engineering (CSE). Here’s a summary of his research work on Beyond Abstracts: Scalable Evidence-Based Verification for Scientific Claims in English and Hindi
Verifying scientific claims is a critical yet challenging task in the era of automated misinformation detection, characterized by high-stakes technical jargon and the requirement for complex domain-specific reasoning. Despite its importance, the field suffers from a significant “data bottleneck”: existing datasets are often limited in scale, primarily English-centric, and frequently overlook the nuances of numerical reasoning and multilingual accessibility. This thesis addresses these gaps by proposing a unified framework that moves from large-scale English baselines to multilingual frontiers in Hindi. The first contribution of this work is the introduction of SciClaimHunt and SciClaimHunt Num, two large-scale datasets derived from scientific research papers. These datasets address the scale deficiency in English scientific NLP and provide a specialized subset for numeral-aware verification. We evaluate these through several baseline models, including Retrieval-Augmented Generation (RAG) and specialized scientific encoders, demonstrating that SciClaimHunt serves as a robust resource for training models that generalize across diverse scientific domains. The second contribution extends this frontier to morphologically rich, low-resource languages with the introduction of HindiSciVerify. Leveraging a semi-automatic pipeline grounded in NCERT textbooks and Parameter-Efficient Fine-Tuning (PEFT), we develop a ternary classification dataset (Support, Refute, Not Enough Information) for Hindi. By utilizing LoRA adapter merging and high-capacity LLMs like Llama-3, we demonstrate a computationally tractable pathway for specializing models for Hindi scientific prose. Our findings indicate that the proposed datasets and methodologies significantly improve the reliability of automated scientific fact-checking. By bridging the gap between English-scale heuristics and the technical nuances of Hindi, this thesis provides a foundational benchmark for future research in multilingual scientific NLP and offers a scalable blueprint for addressing the digital divide in automated fact verification.
July 2026

