Devisetti Sai Asrith supervised by Dr. Radhika Mamidi received his Master of Science – Dual Degree in Computer Science and Engineering (CSD). Here’s a summary of his research work on A Probabilistic Framework for Measuring Gender Bias in AI Generated Occupational Narratives: A Multilingual and Multi-Model Study
The rapid adoption of large language models (LLMs) in applications such as recruitment, content generation, and decision support has raised critical concerns regarding fairness and bias in automated text generation. This thesis presents a comprehensive investigation into gender and occupational bias in LLM-generated narratives, combining theoretical formulation, large-scale empirical analysis, and multilingual evaluation across three complementary studies. At the foundational level, this work formalizes bias in generative AI using a statistical framework, defining bias as the deviation between generated outputs and an ideal, unbiased representation. Building on this, a robust bias quantification pipeline is introduced, incorporating measures such as mean bias (MB), mean absolute bias (MAB), sentiment analysis, and probabilistic distribution distances. Using a large dataset of over 11,000 AI-generated job narratives, the study demonstrates that LLM outputs exhibit measurable and consistent gender bias, particularly in the association of male entities with high-intensity, high-stress, and challenging occupations. Extending beyond single-model analysis, the thesis further explores bias across multiple LLMs, including Gemini, Mistral, LLaMA, and OpenAI models, and introduces a structured evaluation framework based on narrative sentiment, syntactic agency, and a derived prestige score. The findings reveal that gender bias persists across all models, though its magnitude and direction vary depending on model architecture and prompt design. Notably, prompt framing is shown to significantly influence bias, with negative or incompetence-related prompts consistently amplifying male representation.
A key contribution of this work is its multilingual analysis of bias in English and Telugu, addressing a gap in existing research that predominantly focuses on high-resource languages. The results highlight that while bias is present across both languages, its manifestation differs structurally: English narratives show greater sentiment variability, whereas Telugu narratives exhibit higher agency representations. These findings underscore the importance of linguistic context in shaping model behavior. Overall, this thesis contributes a unified, scalable, and interpretable framework for analyzing bias in generative AI systems. By integrating statistical rigor, large-scale datasets, cross-model comparison, and multilingual evaluation, the work advances understanding of how bias emerges and propagates in LLM-generated text and provides a foundation for developing more fair, transparent, and inclusive AI systems.
June 2026

