[month] [year]

Utsav Shekhar

Utsav Shekhar supervised by Dr. Radhika Mamidi received his Master of Science – Dual Degree in Computational Linguistics (CLD). Here’s a summary of his research work on Voices of Dissent: A Multimodal Computational Framework for the Analysis of Protest Music

Music has historically functioned as a primary vehicle for political expression and transforming personal struggles into collective narratives of resistance. While the cultural and historical significance of protest music is well documented, its specific linguistic and acoustic signatures, and the quantifiable features that differentiate it from general music remain largely underexplored. This thesis presents Voices of Dissent, a comprehensive multimodal computational framework designed to analyze and characterize protest music through the dual lenses of natural language processing and audio analysis. To address Because of the scarcity of specialized resources, we first curated a novel, balanced multimodal dataset comprising 458 protest songs (sourced via Wikidata) and 370 non protest songs (matched by time period and genre diversity using GPT-4 inference). Our methodology follows a multi step approach: first, we extract linguistic and audio features including lexical diversity, rhyme density, sentiment polarity, spectral flux etc. to isolate the stylistic dimensions of dissent in lyrics and audio. Second, we evaluate a suite of transformer based models, to assess the performance of deep embeddings in capturing protest intent across text and audio modalities. A key technical contribution of this work is the application of source separation to decompose audio tracks into vocal and accompaniment stems. By conducting controlled mixing experiments, we quantify the relative influence of the ”message” (vocal delivery) versus the ”medium” (instrumental arrangement) in the perception of protest. Our results indicate that protest lyrics exhibit significantly higher repetition and more negative valence compared to non protest counterparts. Furthertext based models (91.10% F1-score) slightly outperform audio-based models (90.62% F1-score), though the high performance of both suggests that protest signals are deeply embedded in both modalities. To validate these computational findings on the audio side, we conduct a human annotation study involving 50 participants, revealing a strong correlation between automated feature extraction and human perception of musical energy, vocal roughness, and melodic disjunctness. This research provides empirical grounding for the study of protest music, offering a scalable framework for researchers in musicology, sociolinguistics, and digital activism to analyze the evolving language of resistance in the digital age. 

June 2026