Hemant Kumar Kathania

dblp:207/4498 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-6367-5203ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 A study on the layer-wise transferability of self-supervised learning features for children's speech processing tasks
Abhijit Sinha, Hemant Kumar Kathania, Mikko Kurimo
Speech Commun.2
2025 Beyond Traditional Speech Modifications : Utilizing Self Supervised Features for Enhanced Zero-Shot Children ASR
Abhijit Sinha, Hemant Kumar Kathania, Mikko Kurimo
INTERSPEECH2
2025 Zero-shot KWS for children's speech using layer-wise features from SSL models
Subham Kutum, Abhijit Sinha, Hemant Kumar Kathania, Sudarsana Reddy Kadiri, Mahesh Chandra Govil
Pattern Recognit. Lett.3
2025 Do all features matter? Layer-wise feature probing of self-supervised speech models for dysarthria severity classification
Paban Sapkota, Harsh Srivastava, Hemant Kumar Kathania, Shri Narayanan, Sudarsana Reddy Kadiri
Speech Commun.3
2025 ResEmoteNet: Bridging Accuracy and Loss Reduction in Facial Emotion Recognition
abstract
The human face is a silent communicator, expressing emotions and thoughts through it's facial expressions. With the advancements in computer vision in recent years, facial emotion recognition technology has made significant strides, enabling machines to decode the intricacies of facial cues. In this work, we propose ResEmoteNet, a novel deep learning architecture for facial emotion recognition designed with the combination of Convolutional, Squeeze-Excitation (SE) and Residual Networks. The inclusion of SE block selectively focuses on the important features of the human face, enhances the feature representation and suppresses the less relevant ones. This helps in reducing the loss and enhancing the overall model performance. We also integrate the SE block with three residual blocks that help in learning more complex representation of the data through deeper layers. We evaluated ResEmoteNet on four open-source databases: FER2013, RAF-DB, AffectNet-7 and ExpW, achieving accuracies of 79.79%, 94.76%, 72.39% and 75.67% respectively. The proposed network outperforms state-of-the-art models across all four databases.
Arnab Kumar Roy, Hemant Kumar Kathania, Adhitiya Sharma, Abhishek Dey, Sarfaraj Alam Ansari
IEEE Signal Process. Lett.2
2025 Can Layer-Wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?
abstract
Automatic Speech Recognition (ASR) systems often struggle to accurately process children's speech due to its distinct and highly variable acoustic and linguistic characteristics. While recent advancements in self-supervised learning (SSL) models have greatly enhanced the transcription of adult speech, accurately transcribing children's speech remains a significant challenge. This study investigates the effectiveness of layer-wise features extracted from state-of-the-art SSL pre-trained models - specifically, Wav2Vec2, HuBERT, Data2Vec, and WavLM in improving the performance of ASR for children's speech in zero-shot scenarios. A detailed analysis of features extracted from these models was conducted, integrating them into a simplified DNN-based ASR system using the Kaldi toolkit. The analysis identified the most effective layers for enhancing ASR performance on children's speech in a zero-shot scenario, where WSJCAM0 adult speech was used for training and PFSTAR children speech for testing. Experimental results indicated that Layer 22 of the Wav2Vec2 model achieved the lowest Word Error Rate (WER) of 5.15%, representing a 51.64% relative improvement over the direct zero-shot decoding using Wav2Vec2 (WER of 10.65%). Additionally, age group-wise analysis demonstrated consistent performance improvements with increasing age, along with significant gains observed even in younger age groups using the SSL features. Further experiments on the CMU Kids dataset confirmed similar trends, highlighting the generalizability of the proposed approach.
Abhijit Sinha, Hemant Kumar Kathania, Sudarsana Reddy Kadiri, Shri Narayanan
IEEE Signal Process. Lett.2
2024 Spectral warping based data augmentation for low resource children's speaker verification
abstract
Abstract In this paper, we present our effort to develop an automatic speaker verification (ASV) system for low resources children’s data. For the children’s speakers, very limited amount of speech data is available in majority of the languages for training the ASV system. Developing an ASV system under low resource conditions is a very challenging problem. To develop the robust baseline system, we merged out of domain adults’ data with children’s data to train the ASV system and tested with children’s speech. This kind of system leads to acoustic mismatches between training and testing data. To overcome this issue, we have proposed spectral warping based data augmentation. We modified adult speech data using spectral warping method (to simulate like children’s speech) and added it to the training data to overcome data scarcity and mismatch between adults’ and children’s speech. The proposed data augmentation gives 20.46% and 52.52% relative improvement (in equal error rate) for Indian Punjabi and British English speech databases, respectively. We compared our proposed method with well known data augmentation methods: SpecAugment, speed perturbation (SP) and vocal tract length perturbation (VTLP), and found that the proposed method performed best. The proposed spectral warping method is publicly available at https://github.com/kathania/Speaker-Verification-spectral-warping .
Hemant Kumar Kathania, Virender Kadyan, Sudarsana Reddy Kadiri, Mikko Kurimo
Multim. Tools Appl.1
2022 A formant modification method for improved ASR of children's speech
abstract
Differences in acoustic characteristics between children’s and adults’ speech degrade performance of automatic speech recognition systems when systems trained using adults’ speech are used to recognize children’s speech. This performance degradation is due to the acoustic mismatch between training and testing. One of the main sources of the acoustic mismatch is the difference in vocal tract resonances (formant frequencies) between adult and child speakers. The present study aims to reduce the mismatch in formant frequencies by modifying formants of children’s speech to better correspond to formants of adults’ speech. This is carried out by warping the linear prediction (LP) spectrum computed from children’s speech. The warped LP spectra computed in a frame-based manner from children’s speech are used with the corresponding LP residuals to synthesize speech whose formant structure is closer to that of adults’ speech. When used in testing of an ASR system trained using adults’ speech, the warping reduces the spectral mismatch in speech between training and testing and improves the system performance in recognition of children’s speech. Experiments were conducted using narrowband (8 kHz) and wideband (16 kHz) speech of adult and child speakers from the WSJCAM0 and PF_STAR databases, respectively, and by recognizing children’s speech using acoustic models trained with adults’ speech. The proposed method gave relative improvements of 24% and 11% for the DNN and TDNN acoustic models, respectively, for narrowband speech. For wideband speech, the technique gave relative improvements of 27% and 13% for the DNN and TDNN acoustic models, respectively. The performance of the proposed method was also compared to two speaker adaptation methods: vocal tract length normalization (VTLN) and speaking rate adaptation (SRA). This comparison showed the best recognition performance for the proposed method. We also combined the proposed method with VTLN and SRA, and found that the combined method gave a further reduction in WER. Moreover, our experiments carried out for noisy speech using various types of additive noise and signal-to-noise ratios showed that the proposed method performs well also for degraded speech.
Hemant Kumar Kathania, Sudarsana Reddy Kadiri, Paavo Alku, Mikko Kurimo
Speech Commun.1
2021 Vowel Non-Vowel Based Spectral Warping and Time Scale Modification for Improvement in Children's ASR
abstract
Acoustic differences between children’s and adults’ speech causes the degradation in the automatic speech recognition system performance when system trained on adults’ speech and tested on children’s speech. The key acoustic mismatch factors are formant, speaking rate, and pitch. In this paper, we proposed a linear prediction based spectral warping method by using the knowledge of vowel and non-vowel regions in speech signals to mitigate the formant frequencies differences between child and adult speakers. The proposed method gives 31% relative improvement over the baseline system. We have also investigated time scale modification using RTISILA and SOLAFS algorithms and found that our proposed method performs better. Combining the proposed method with RTISILA and SOLAFS results in a further error rate reduction. The final combined system gives 49% relative improvement compared to the baseline system.
Hemant Kumar Kathania, Mikko Kurimo
ICASSP1
2020 Study of Formant Modification for Children ASR
abstract
The performance of automatic speech recognition systems for children’s speech is known to suffer from the large variation and mismatch in the acoustic and linguistic attributes between children’s and adults’ speech. One of the various identified sources of mismatch is the difference in formant frequencies between adults and children. In this paper, we propose a formant modification method to mitigate differences between adults’ and children’s speech and to improve the performance of ASR for children. The explored technique gives a relative 27% improvement in system performance compared to a hybrid DNN-HMM baseline. We also compare the system performance with related speaker adaptation methods like vocal tract length normalization (VTLN) and speaking rate adaptation (SRA) and find that the proposed method gives improvements over them, as well. Combining the proposed method with VTLN and SRA results in a further reduction of WER. We also found that the proposed method performs well even for noisy speech.
Hemant Kumar Kathania, Sudarsana Reddy Kadiri, Paavo Alku, Mikko Kurimo
ICASSP1
2020 Data Augmentation Using Prosody and False Starts to Recognize Non-Native Children's Speech
abstract
This paper describes AaltoASR's speech recognition system for the INTERSPEECH 2020 shared task on Automatic Speech Recognition (ASR) for non-native children's speech. The task is to recognize non-native speech from children of various age groups given a limited amount of speech. Moreover, the speech being spontaneous has false starts transcribed as partial words, which in the test transcriptions leads to unseen partial words. To cope with these two challenges, we investigate a data augmentation-based approach. Firstly, we apply the prosody-based data augmentation to supplement the audio data. Secondly, we simulate false starts by introducing partial-word noise in the language modeling corpora creating new words. Acoustic models trained on prosody-based augmented data outperform the models using the baseline recipe or the SpecAugment-based augmentation. The partial-word noise also helps to improve the baseline language model. Our ASR system, a combination of these schemes, is placed third in the evaluation period and achieves the word error rate of 18.71%. Post-evaluation period, we observe that increasing the amounts of prosody-based augmented data leads to better performance. Furthermore, removing low-confidence-score words from hypotheses can lead to further gains. These two improvements lower the ASR error rate to 17.99%.
Hemant Kumar Kathania, Mittul Singh, Tamás Grósz, Mikko Kurimo
INTERSPEECH1
2020 Creating speaker independent ASR system through prosody modification based data augmentation
Syed Shahnawazuddin, Nagaraj Adiga, Hemant Kumar Kathania, B. Tarun Sai
Pattern Recognit. Lett.3
2018 Role of Prosodic Features on Children's Speech Recognition
abstract
In this paper, we have explored the role of combining prosodic variables with the existing acoustic features in the context of children's speech recognition using acoustic models trained on adults' speech. The explored acoustic features are Mel-frequency cepstral coefficients (MFCC) and perceptual linear prediction cepstral coefficients (PLPCC) while the considered prosodic variables are loudness, voice-intensity and voice-probability. An analysis presented in this paper shows that, given that the textual content remains the same, the considered prosodic variables exhibit very similar contours for adults' and children's speech. At the same time, the contours differ a lot when the context is different. Consequently, inclusion of prosodic information reduces the inter-speaker differences and increases the class discrimination. This subsequently improves the recognition performance. Further improvements are obtained by projecting the feature vectors obtained by combining the two features to a lower-dimensional subspace. The same has been experimentally verified in this study for mismatched speech recognition using deep neural network (DNN) based system. On combining MFCC (PLPCC) and prosodic features, a relative improvement of 16% (14%) is noted on decoding children's speech using adult data trained DNN models.
Hemant Kumar Kathania, Syed Shahnawazuddin, Nagaraj Adiga, Waquar Ahmad
ICASSP1
2018 Improving children's mismatched ASR using structured low-rank feature projection
Syed Shahnawazuddin, Hemant Kumar Kathania, Abhishek Dey, Rohit Sinha 0003
Speech Commun.2
2017 Improving Children's Speech Recognition Through Explicit Pitch Scaling Based on Iterative Spectrogram Inversion
Waquar Ahmad, Syed Shahnawazuddin, Hemant Kumar Kathania, Gayadhar Pradhan, Arun B. Samaddar
INTERSPEECH3
2017 Effect of Prosody Modification on Children's ASR
abstract
Transcribing children's speech using acoustic models trained on adults' speech is very challenging. In such conditions, a highly degraded recognition performance is reported due to large mismatch in the acoustic/linguistic attributes of the training and test data. The differences in pitch (or fundamental frequency) between the two groups of speakers is one among several mismatch factors. Another important mismatch factor is the difference in speaking rates. To overcome these two sources of mismatch, prosody modification is explored in this letter. Prosody modification is done by using glottal closure instants (GCIs) as anchoring points. The GCIs, in turn, are determined using zero-frequency filtering (ZFF). The ZFF-GCI-based prosody modification is fast and results in highly accurate scaling of pitch and speaking rate. The experimental evaluations studying the effect of prosody modification resulted in a relative improvement of 50% over the baseline.
Syed Shahnawazuddin, Nagaraj Adiga, Hemant Kumar Kathania
IEEE Signal Process. Lett.3