Harinie Sivaramasethu supervised by Prof. Anil Kumar Vuppala received her Master of Science – Dual Degree in Computational Linguistics (CLD) by Research. Here’s a summary of her research work on Towards Phonetic Abstraction Beyond Orthography in Multilingual Speech
Modern speech technology often relies on orthographic text such as transcripts, spellings, and characterlevel edit distances, even though human listeners judge speech by phonological and semantic content. For morphologically rich and low-resource languages, this orthographic bottleneck produces unreliable evaluation and weak pronunciation modelling. This thesis explores phonetic abstraction: the deliberate reduction of surface script to shared phonological structure before comparison or transfer.
Automatic Speech Recognition (ASR) evaluation for morphologically complex languages remains an open and practically significant problem. Standard metrics such as Word Error Rate (WER) and Character Error Rate (CER) operate purely at the lexical level, assigning uniform penalties to all substitutions regardless of whether the resulting output is intelligible to a human listener. In agglutinative and fusional languages with rich phonological inventories, such as Tamil, Telugu, Hindi, and Marathi, these metrics routinely misalign with human intelligibility.
This thesis proposes and evaluates a semantic-phonetic hybrid evaluation framework that combines a deterministic phonetic skeleton, derived by mapping consonant characters across languagespecific Indic scripts into shared articulatory equivalence classes, with contextual semantic embeddings from large pre-trained multilingual models. A linear combination of these two components, optimised via Ordinary Least Squares regression against expert human scores on 1,287 challenging transcription pairs, yields a metric that consistently outperforms string-edit and LLM based baselines across all evaluated languages.
Arabic, Persian, and Urdu share phonology but not computational resources. As a complementary contribution, the thesis also presents a Common Phone Space (CPS-APU) for Arabic, Persian, and Urdu, which applies cross-lingual phonological transfer to the problem of Grapheme-to-Phoneme conversion for low-resource Perso-Arabic languages. It is a three-stage Grapheme-to-Phoneme (G2P) pipeline (expert rules, corpus-mined rules, Large Language Model (LLM)-powered generation), and shows that cross-lingual transfer from high-resource Arabic improves G2P for low-resource Persian and Urdu.
Together, these works argue that underlying phonetic and phonological structures offer a robust computational resource that can be systematically exploited to overcome text-level limitations. Both contributions instantiate the same principle: related languages share phonological structure that can be
encoded as equivalence classes (evaluation-time bucketing in Indic scripts; CPS labels in Perso-Arabic
G2P). Phonetic abstraction is a transferable resource for building fair and effective speech technology
for the world’s linguistic diversity beyond surface orthography.
August 2026

