Tilak Purohit

dblp:274/5323 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
7since 2021 · last 2025
0009-0009-8580-7746ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Emotion information recovery potential of wav2vec2 network fine-tuned for speech recognition task
abstract
Fine-tuning has become a norm to achieve state-of-the-art performance when employing pre-trained networks like foundation models. These models are typically pre-trained on large-scale unannotated data using self-supervised learning (SSL) methods. The SSL-based pre-training on large-scale data enables the network to learn the inherent structure/properties of the data, providing it with capabilities in generalization and knowledge transfer for various downstream tasks. However, when fine-tuned for a specific task, these models become task-specific. Finetuning may cause distortions in the patterns learned by the network during pre-training. In this work, we investigate these distortions by analyzing the network’s information recovery capabilities by designing a study where speech emotion recognition is the target task and automatic speech recognition is an intermediary task. We show that the network recovers the task-specific information but with a shift in the decisions also through attention analysis, we demonstrate some layers do not recover the information fully.
Tilak Purohit, Mathew Magimai-Doss
ICASSP1
2025 Automatic Parkinson's disease detection from speech: Layer selection vs adaptation of foundation models
abstract
In this work, we investigate Speech Foundation Models (SFMs) for Parkinson’s Disease (PD) detection. We explore two main approaches: (1) using SFMs as frozen feature extractors and, (2) fine-tuning/adapting SFMs for PD detection. We propose a cross-validation-based layer selection methodology to identify the layer effective for PD detection. Additionally, we compare the performance of the layer selection scheme with full fine-tuning and, parameter-efficient fine-tuning (PEFT) using Low-Rank Adaptation (LoRA). Our results show that layer selection and LoRA-based fine-tuning can perform on par with full fine-tuning, providing a more parameter-efficient alternative. The highest accuracy was achieved by fine-tuning Whisper using LoRA.
Tilak Purohit, Barbara Ruvolo, Juan Rafael Orozco-Arroyave, Mathew Magimai-Doss
ICASSP1
2024 Cross-transfer Knowledge between Speech and Text Encoders to Evaluate Customer Satisfaction
L. Felipe Parra-Gallego, Tilak Purohit, Bogdan Vlasenko, Juan Rafael Orozco-Arroyave, Mathew Magimai-Doss
INTERSPEECH2
2023 Towards Learning Emotion Information from Short Segments of Speech
abstract
Conventionally, speech emotion recognition has been approached by utterance or turn-level modelling of input signals, either through extracting hand-crafted low-level descriptors, bag-of-audio-words features or by feeding long-duration signals directly to deep neural networks (DNNs). While this approach has been successful, there is a growing interest in modelling speech emotion information at the short segment level, at around 250ms-500ms (e.g. the 2021-22 MuSe Challenges). This paper investigates both hand-crafted feature-based and end-to-end raw waveform DNN approaches for modelling speech emotion information in such short segments. Through experimental studies on IEMOCAP corpus, we demonstrate that the end-to-end raw waveform modelling approach is more effective than using hand-crafted features for short-segment level modelling. Furthermore, through relevance signal-based analysis of the trained neural networks, we observe that the top performing end-to-end approach tends to emphasize cepstral information instead of spectral information (such as flux and harmonicity).
Tilak Purohit, Sarthak Yadav, Bogdan Vlasenko, S. Pavankumar Dubagunta, Mathew Magimai-Doss
ICASSP1
2023 A Study on the Importance of Formant Transitions for Stop-Consonant Classification in VCV Sequence
Siddarth Chandrasekar, Arvind Ramesh, Tilak Purohit, Prasanta Kumar Ghosh
INTERSPEECH3
2023 Implicit phonetic information modeling for speech emotion recognition
Tilak Purohit, Bogdan Vlasenko, Mathew Magimai-Doss
INTERSPEECH1
2021 Impact of Speaking Rate on the Source Filter Interaction in Speech: A Study
abstract
Source filter interaction (SFI) explains the drop in pitch caused due to the constriction in the vocal tract during voiced consonant production in a vowel-consonant-vowel (VCV) sequence. In this work, we examine how the drop in pitch alters when such a VCV sequence is spoken at three different speaking rates - slow, normal and fast. In the absence of electroglottograph (EGG) recording, a high resolution pitch contour is determined using a glottal closure instant (GCI) detector. For this, in this work, firstly, five different GCI detector and pitch estimation techniques are compared against EGG based pitch estimates on a small dataset where simultaneous EGG recordings are available. Yet Another GCI Algorithm (YAGA) is found to be the best choice among all. For examining the impact of speaking rate on SFI, VCV recordings from six subjects with five vowels (/a/, /e/, /i/, /o/, /u/) and five consonants (/b/, /d/, /g/, /v/, /z/) at three speaking rates are used. The study reveals a significant difference in the pitch drop values between slow and fast rates, with increasing pitch drop as speaking rate reduces. For slow speaking rate, vowel /o/ and /u/ tend to show higher pitch drop values compared to remaining vowels.
Tilak Purohit, M. V. Achuth Rao, Prasanta Kumar Ghosh
ICASSP1
2020 An Investigation of the Virtual Lip Trajectories During the Production of Bilabial Stops and Nasal at Different Speaking Rates
abstract
We propose a technique to estimate virtual upper lip (VUL) and virtual lower lip (VLL) trajectories during production of bilabial stop consonants (/p/, /b/) and nasal (/m/). A VUL (VLL) is a hypothetical trajectory below (above) the measured UL (LL) trajectory which could have been achieved by UL (LL) if UL and LL were not in contact with each other during bilabial stops and nasal. Maximum deviation of UL from VUL and its location as well as the range of VUL are used as features, denoted by VUL MD, VUL MDL, and VUL R, respectively. Similarly, VLL MD, VLL MDL, and VLL R are also computed. Analyses of these six features are carried out for /p/, /b/, and /m/ at slow, normal and fast rates based on electromagnetic articulograph (EMA) recordings of VCV stimuli spoken by ten subjects. While no significant differences were observed among /p/, /b/, and /m/ in every rate, all six features except VLL MD were found to drop significantly from slow to fast rates. These six features were also found to perform better in an automatic classification task between slow vs fast rates compared to five baseline features computed from UL and LL comprising their ranges, velocities and minimum distance from each other. © 2020 ISCA
Tilak Purohit, Prasanta Kumar Ghosh
INTERSPEECH1