Pratyaksh Gautam supervised by Dr. Vinoo Alluri received his Master of Science by Research – Dual Degree in Computational Linguistics (CLD). Here’s a summary of his Investigating Auditory Concepts in Deep Neural Networks
Deep neural networks have achieved strong performance across a wide range of auditory tasks, from speech recognition to music classification. Despite this success, their internal decision-making processes remain poorly understood, limiting both scientific insight and practical trust. In real-world settings, models that rely on opaque or spurious cues can behave unpredictably, making interpretability a central challenge in modern deep learning. This thesis addresses this gap by asking a central question: do auditory deep neural networks organize sound in a hierarchical manner similar to human auditory perception? To test this hypothesis, we conduct a large-scale probing analysis of several widely used audio architectures, including three CNN models (VGGish, CLAP, MobileNetV3), and an Audio Spectrogram Transformer (AST). Using linear probes across six tasks of increasing abstraction, we examine how different forms of auditory information are distributed across network depth. Across all convolutional architectures, we observe a clear and consistent hierarchy. Low-level tasks such as note name classification peak in early layers, while higher-level semantic tasks, including genre classification and speaker count estimation, depend on deeper representations. Although this hierarchy is less sharply delineated in the AST due to its global attention mechanism, the same overall progression remains evident. This suggests that hierarchical organization is a robust property of effective auditory models, even when architectural constraints differ. Beyond identifying this structure, we also show complex auditory concepts emerge within a deep neural network. Focusing on musical audio, we show that an AST trained solely for genre classification develops explicit representations of musical instruments in intermediate layers. Steering vector interventions further show that these instrument representations actively influence genre predictions, demonstrating that they are causally involved in the model’s decisions. These findings collectively establish a principled framework for analyzing, comparing, and causally testing internal representations in auditory deep neural networks, while also grounding modern audio models in theories of hierarchical auditory perception.
June 2026

