VLDB 2026 Research / reviewers in the wild / expert
Prasanta Kumar Ghosh
dblp:73/6634
· DBLP profile ↗
158ranked-venue papers
13as first author
52since 2021 · last 2026
0000-0002-1137-0838ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 146 · 12 first-author · 50 since 2021Artificial intelligence and machine learning · 97 · 6 first-author · 34 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Source and filter characteristics based transfer learning for dysarthria severity classification in amyotrophic lateral sclerosis
Tanuka Bhattacharjee, Yamini Belur, Atchayaram Nalini, Prasanta Kumar Ghosh |
Speech Commun. | 4 |
| 2025 | Bottleneck Transformer-Based Approach for Improved Automatic STOI Score Prediction
Amartyaveer, Murali Kadambi, Chandra Mohan Sharma, Anupam Mandal, Prasanta Kumar Ghosh |
ASRU | 5 |
| 2025 | MADASR 2.0: Multi-Lingual Multi-Dialect ASR Challenge in 8 Indian LanguagesabstractWe present MADASR 2.0, a challenge at ASRU 2025 aimed at advancing multilingual and multidialectal automatic speech recognition (ASR) in low-resource Indian languages. Building on the 2023 edition, it introduces a subset of the RESPIN corpus, over 1200 hours of read speech across 8 languages and 33 dialects, with test sets including both read and spontaneous speech. The challenge comprises four tracks varying by training data size and external resource usage, and supports auxiliary tasks like language and dialect identification. We detail the dataset, tasks, baselines, and submissions and analyse trends across tracks and speech styles. Results highlight the continued difficulty of spontaneous ASR, the benefits of multitask and transfer learning, and effective strategies for building dialect-aware ASR systems. MADASR 2.0 offers a standardised benchmark to support future research on inclusive and scalable ASR for linguistically diverse populations. Sumit Sharma 0016, Deekshitha G, Abhayjeet Singh, Amartyaveer, Sathvik Udupa, Sandhya Badiger, Sanjeev Khudanpur, Sunayana Sitaram, Srinivasan Umesh, Bhuvana Ramabhadran, Brian Kingsbury, Hema A. Murthy, Srikanth S. Narayanan, Howard Lakougna, Prasanta Kumar Ghosh |
ASRU | 16 |
| 2025 | Improving Dialect Identification in Indian Languages Using Multimodal Features from Dialect Informed ASRabstractDialect identification (DID) addresses the challenge of recog-nizing regional variations within a language. The current deep learning approaches focus on audio-only, text-only, or multi-task setups combining automatic speech recognition (ASR) with DID. This work introduces a novel multimodal architecture that leverages speech and text features to enhance DID performance. Our method integrates ASR-generated speech representations with text embeddings derived from ASR hypotheses using a RoBERTa-based encoder. Additionally, we perform a layer-wise analysis of the IndicWav2Vec model to identify the layers most effective for extracting dialectal features. We evaluate our approach on a subset of the RESPIN dataset featuring eight Indian languages and 33 dialects. Experimental results show that our proposed multimodal DID system achieves an average DID accuracy of 79.81%, consistently outperforming baseline methods. This study is the first to analyse comprehensively DID in Indian languages, providing new insights into their dialectal diversity. Amartyaveer, Sumit Sharma 0016, Sathvik Udupa, Sandhya Badiger, Abhayjeet Singh, Deekshitha G, Jesuraja Bandekar, Savitha Murthy, Prasanta Kumar Ghosh |
ICASSP | 10 |
| 2025 | Role of the Pretraining and the Adaptation data sizes for low-resource real-time MRI video segmentationabstractReal-time Magnetic Resonance Imaging (rtMRI) is frequently used in speech production studies as it provides a complete view of the vocal tract during articulation. This study investigates the effectiveness of rtMRI in analyzing vocal tract movements by employing the SegNet and UNet models for Air-Tissue Boundary (ATB) segmentation tasks. We conducted pretraining of a few base models using increasing numbers of subjects and videos, to assess performance on two datasets. First, consisting of unseen subjects with unseen videos from the same data source, achieving 0.33% and 0.91% (Pixel-wise Classification Accuracy (PCA) and Dice Coefficient respectively) better than its matched condition. Second, comprising unseen videos from a new data source, where we obtained an accuracy of 99.63% and 98.09% (PCA and Dice Coefficient respectively) of its matched condition performance. Here, matched condition performance refers to the performance of a model trained only on the test subjects which was set as a benchmark for the other models. Our findings highlight the significance of fine-tuning and adapting models with limited data. Notably, we demonstrated that effective model adaptation can be achieved with as few as 15 rtMRI frames from any new dataset. Masoud Thajudeen Tholan, Vinayaka Hegde, Prasanta Kumar Ghosh |
ICASSP | 4 |
| 2025 | Enhancing Acoustic-to-Articulatory Inversion with Multi-Target Pretraining for Low-Resource SettingsabstractAcoustic-to-Articulatory Inversion (AAI) estimates vocal tract articulator movements from speech, benefiting tasks like ASR, speech synthesis, and speaker verification. While deep learning-based methods (CNNs, RNNs, Transformers) have advanced AAI, recent studies show that Self-Supervised Learning (SSL) features further enhance performance, particularly in low-resource settings. However, SSL feature extractors introduce inference latency and computational overhead. To address this, we propose a novel pretraining method leveraging three target representations-Phoneme Labels, Articulatory Feature Labels, and Critical-articulator Labels-eliminating the need for an SSL extractor during inference. We evaluate our approach against both baseline and SSL-based models across various data conditions. Results demonstrate that our method consistently improves AAI performance, particularly in low-resource scenarios, while significantly reducing inference costs without sacrificing accuracy. Jesuraja Bandekar, Prasanta Kumar Ghosh |
INTERSPEECH | 2 |
| 2025 | Jointly Improving Dialect Identification and ASR in Indian Languages using Multimodal Feature FusionabstractAutomatic Speech Recognition (ASR) and Dialect Identification (DID) are crucial for Indian languages, many of which are low-resource and exhibit significant dialectal differences. Existing methods often optimize ASR or DID individually, resulting in performance trade-offs. In this work, we propose a multimodal framework that jointly improves ASR and DID. Our method employs a Bottleneck Encoder to extract dialectal features from Conformer-based speech representations and a RoBERTa encoder to process ASR-generated CTC embeddings. A gating mechanism merges these features, followed by an attention encoder to refine the representations. The learned embeddings are concatenated with Conformer outputs to enhance ASR features. Evaluated on eight Indian languages with thirty-three dialects, our method achieves an average DID accuracy of 81.63% and average CER and WER of 4.65% and 17.73%, respectively. These results highlight the effectiveness of our method for joint ASR-DID modeling. Amartyaveer, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2025 | An approach to measuring the performance of Automatic Speech Recognition(ASR) models in the context of Large Language Model(LLM) powered applications
Sujith Pulikodan, Sahapthan K, Prasanta Kumar Ghosh, Visruth Sanka, Nihar Desai |
INTERSPEECH | 3 |
| 2025 | A real-time MRI study on asymmetry in velum dynamics during VCV production with nasal sounds
Vaishnavi Chandwanshi, Shreya Shrikant Karkun, Aditya Anand Gupta, Prasanta Kumar Ghosh |
INTERSPEECH | 5 |
| 2025 | Boosting StoRM Convergence with Metric Guidance and Non-uniform State-Sampling for Optimal Dereverberation
Chandra Mohan Sharma, Arnab Kumar Roy, Anupam Mandal, Prasanta Kumar Ghosh, Prasanna Kumar Kr |
INTERSPEECH | 4 |
| 2025 | Comparison of Acoustic and Textual Features for Dysarthria Severity Classification in Amyotrophic Lateral Sclerosis
Y. S. Upendra Vishwanath, Tanuka Bhattacharjee, Deekshitha G, Sathvik Udupa, Chowdam Venkata Thirumala Kumar, Madassu Keerthipriya, Darshan Chikktimmegowda, Dipti Baskar, Yamini Belur, Seena Vengalil, Atchayaram Nalini, Prasanta Kumar Ghosh |
INTERSPEECH | 12 |
| 2025 | RESPIN-S1.0: A read speech corpus of 10000+ hours in dialects of nine Indian LanguagesabstractWe introduce RESPIN-S1.0, the largest publicly available dialect-rich read-speech corpus for Indian languages, comprising more than 10,000 hours of validated audio across nine major languages: Bengali, Bhojpuri, Chhattisgarhi, Hindi, Kannada, Magahi, Maithili, Marathi, and Telugu. Indian languages exhibit high dialectal variation and are spoken by populations that remain digitally underserved. Existing speech corpora typically represent only standard dialects and lack domain and linguistic diversity. RESPIN-S1.0 addresses this limitation by collecting speech across more than 38 dialects and two high-impact domains: agriculture and finance. Text data were composed by native dialect speakers and validated through a pipeline combining automated and manual checks. Over 200,000 unique sentences were recorded through a crowdsourced mobile platform and categorised into clean, semi-noisy, and noisy subsets based on transcription quality, with the clean portion alone exceeding 10,000 hours. Along with audio and transcriptions, RESPIN provides dialect-aware phonetic lexicons, speaker metadata, and reproducible train, development, and test splits. To benchmark performance, we evaluate multiple ASR models, including TDNN-HMM, E-Branchformer, Whisper, and wav2vec2-based self-supervised models, and find that fine-tuning on RESPIN significantly improves recognition accuracy over pretrained baselines. A subset of RESPIN-S1.0 has already supported community challenges such as the SLT Code Hackathon 2022 and MADASR@ASRU 2023 and 2025, releasing more than 1,200 hours publicly. This resource supports research in dialectal ASR, language identification, and related speech technologies, establishing a comprehensive benchmark for inclusive, dialect-rich ASR in multilingual low-resource settings. Dataset: https://spiredatasets.ee.iisc.ac.in/respincorpus Code: https://github.com/labspire/respin_baselines.git Abhayjeet Singh, Deekshitha G, Amartya Veer, Jesuraja Bandekar, Savitha Murthy, Sumit Sharma 0016, Sandhya Badiger, Sathvik Udupa, Amala Nagireddi, Srinivasa Raghavan K. M., Rohan Saxena, Jai Nanavati, Raoul Nanavati, Janani Sridharan, Arjun Singh Mehta, Ashish Seth, Sai Praneeth Reddy Mora, Prashanthi V, Gauri Date, Karthika P, Prasanta Kumar Ghosh |
NeurIPS | 22 |
| 2024 | Spectral Analysis of Vowels and Fricatives at Varied Levels of Dysarthria Severity for Amyotrophic Lateral SclerosisabstractDysarthria due to Amyotrophic Lateral Sclerosis (ALS) affects the acoustic characteristics of different speech sounds. The effects intensify with increasing severity leading to the collapse of the acoustic space of the affected individuals. With an aim to characterize such changes in the acoustic space, this paper studies the variations in band-specific and full-band spectral properties of 4 sustained vowels (/a/, /i/, /o/, /u/) and 3 sustained fricatives (/s/, /sh/, /f/) at different dysarthria severity levels. Effect of dysarthria on spectral features of these phonemes are not well explored. Statistical comparison of these features among different severities for the phonemes considered and among different vowels/fricatives for every severity level using speech data from 119 ALS and 40 healthy subjects indicate the followings. Though all band-specific and full-band features of the three fricatives and most of those features for the four vowels become statistically similar at high severity levels, certain features remain distinguishable. Spectral differences in 0-2 kHz band between /a/ and the other vowels and in the 2-6 kHz band between /a/ and /o/, /u/ persist through all severity levels. Moreover, properties of /f/ remain mostly unchanged with increasing dysarthria severity levels. Chowdam Venkata Thirumala Kumar, Tanuka Bhattacharjee, Seena Vengalil, Saraswati Nashi, Madassu Keerthipriya, Yamini Belur, Atchayaram Nalini, Prasanta Kumar Ghosh |
ICASSP | 8 |
| 2024 | An Unsupervised Segmentation of Vocal Breath SoundsabstractBreathing is essential to human survival, which carries information about a person’s physiological and psychological state. Mostly breath sound boundaries are marked manually before being used for any task such as classification, spectral analysis, etc., which is very tedious. Various techniques have been proposed to segment breath sounds recorded at the chest, and trachea but vocal breath sounds (VBS) are under-explored. An unsupervised algorithm for VBS segmentation has been proposed in this work. Each breath phase in continuous breaths has been modeled using triangles, where the end points of triangles representing breath boundaries are estimated using dynamic programming. Data from 60 subjects (31 healthy, 29 asthmatic patients) having 307 breaths have been used. The proposed method’s performance was found to be comparable with the manually marked boundaries. Comparable asthmatic versus healthy subject mean(standard deviation) classification accuracy using manually marked and predicted boundaries are 75%(±11%) and 72%(±15%),respectively are found. Dipanjan Gope, K. Uma Maheswari, Prasanta Kumar Ghosh |
ICASSP | 4 |
| 2024 | Articulatory synthesis using representations learnt through phonetic label-aware contrastive loss
Jesuraja Bandekar, Sathvik Udupa, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2024 | Exploring Syllable Discriminability during Diadochokinetic Task with Increasing Dysarthria Severity for Patients with Amyotrophic Lateral Sclerosis
Neelesh Samptur, Tanuka Bhattacharjee, Anirudh Chakravarty K, Seena Vengalil, Yamini Belur, Atchayaram Nalini, Prasanta Kumar Ghosh |
INTERSPEECH | 7 |
| 2024 | A comparative study of the impact of voiceless alveolar and palato-alveolar sibilants in English on lip aperture and protrusion during VCV production
Vaishnavi Chandwanshi, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2024 | Adapter pre-training for improved speech recognition in unseen domains using low resource adapter tuning of self-supervised models
Sathvik Udupa, Jesuraja Bandekar, Deekshitha G, Sandhya Badiger, Abhayjeet Singh Savitha Murthy, Priyanka Pai, Srinivasa Raghavan K. M., Raoul Nanavati, Prasanta Kumar Ghosh |
INTERSPEECH | 10 |
| 2024 | IndicMOS: Multilingual MOS Prediction for 7 Indian languages
Sathvik Udupa, Soumi Maiti, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2023 | Gated Multi Encoders and Multitask Objectives for Dialectal Speech Recognition in Indian LanguagesabstractIn this work, several methods have been proposed towards improving the performance of dialectal automatic speech recognition (ASR). A novel encoder architecture has been introduced that is suited for multi-dialect ASR training. Further, we propose Multi-Task Self-Supervised learning (SSL) fine-tuning using CTC and dialect identification. Additionally, the use of different language models (LM) to improve the performance of dialectal ASR has been investigated. Around 800 hours of Bengali and Bhojpuri data, released as a part of the MADASR ASRU challenge have been used to train these models. The work shows that the proposed multi-encoder ASR observes a relative reduction of 7.5% and 9% in WER in Bhojpuri and Bengali, respectively. Additionally, we also observe a 1-2% WER reduction in fine-tuning SSL, further improving performance in these languages. Moreover, we observe advantages in using dialect-specific LM decoding based on predicted dialect. Sathvik Udupa, Jesuraja Bandekar, Deekshitha G, Prasanta Kumar Ghosh, Sandhya Badiger, Abhayjeet Singh, Savitha Murthy, Priyanka Pai, Srinivasa Raghavan K. M., Raoul Nanavati |
ASRU | 5 |
| 2023 | Exploring the Role of Fricatives in Classifying Healthy Subjects and Patients with Amyotrophic Lateral Sclerosis and Parkinson's DiseaseabstractDysarthria due to Amyotrophic Lateral Sclerosis (ALS) and Parkinson’s Disease (PD) impairs sustained phoneme productions. Vowels and fricatives get affected differently owing to the differences in their production mechanisms. This paper examines three sustained voiceless fricatives - /s/, /sh/ and /f/, as compared to three sustained vowels - /a/, /i/ and /o/, for classifying patients with ALS/PD and Healthy Controls (HC). Fricatives are found to achieve higher classification accuracies than /a/ and /o/, though /i/ outperforms all. Patients seem to find it difficult to form constrictions while producing fricatives, or to proximally position the tongue and palate while uttering /i/, due to dysarthria. Unwanted voicing added to voiceless fricatives by the patient population further contributes towards the discrimination. Both source (related to vocal cord) and filter (related to vocal tract) cues of fricatives, on average, outperform those of vowels. Lastly, decision-level fusion of /i/-/s/-/sh/, with a pooled classifier for these three phonemes, achieves the highest mean ALS vs. HC classification accuracy of 83.35%, although in PD vs. HC case, fusion of multiple /i/ utterances performs the best with an accuracy of 80.03%. Tanuka Bhattacharjee, Yamini Belur, Atchayaram Nalini, Prasanta Kumar Ghosh |
ICASSP | 5 |
| 2023 | Static and Dynamic Source and Filter Cues for Classification of Amyotrophic Lateral Sclerosis Patients and Healthy SubjectsabstractDysarthria due to Amyotrophic Lateral Sclerosis (ALS) affects speech production. Even the elementary sustained vowel utterances get impaired. For these, the impairments can be in achieving vowel-specific articulatory configurations, reflected in static acoustic cues, and/or in sustaining a configuration for a prolonged duration, reflected in dynamic cues. Such cues can further be attributed to the vocal cord (source) and vocal tract (filter) involved in speech production. This paper analyzes the relative contributions of these static (captured through average spectral characteristics) and dynamic (captured through spectral variations over time) source and filter cues toward automatic classification of ALS patients and healthy subjects using sustained utterances of /a/, /i/, /o/ and /u/. Experiments with 80 ALS patients and 80 healthy subjects suggest that the source cues (static/dynamic) are not the primary discriminators. For /i/, the static filter cues achieve the highest mean classification accuracy of 76.66%, whereas, for /a/, /o/ and /u/, the dynamic filter attributes contribute the most attaining average accuracies of 66.29%, 73.03% and 70.27%, respectively. Hence, ALS patients seem to face difficulties in forming the front closed vocal tract structure of /i/, whereas, holding the target vocal tract shape for long appears to be the primary challenge in case of /a/, /o/ and /u/. Tanuka Bhattacharjee, Chowdam Venkata Thirumala Kumar, Yamini Belur, Atchayaram Nalini, Prasanta Kumar Ghosh |
ICASSP | 6 |
| 2023 | Lightweight, Multi-Speaker, Multi-Lingual Indic Text-to-SpeechabstractThe Lightweight, Multi-speaker, Multi-lingual Indic Text-to-Speech (LIMMITS’23) challenge is organized as part of the ICASSP 2023 signal processing grand challenge. LIMMITS’23 aims at the development of a lightweight, multi-speaker, multi-lingual Text to Speech (TTS) model using datasets in Marathi, Hindi, and Telugu. The challenge encourages the advancement of TTS in Indian Languages as well as the development of techniques involved in TTS data selection and model compression. The 3 tracks of LIMMITS’23 have provided an opportunity for various researchers and practitioners around the world to explore the state of the art in TTS research. Abhayjeet Singh, Amala Nagireddi, Deekshitha G, Jesuraja Bandekar, Roopa R., Sandhya Badiger, Sathvik Udupa, Prasanta Kumar Ghosh, Hema A. Murthy, Heiga Zen, Pranaw Kumar, Kamal Kant, Amol Bole, Bira Chandra Singh, Keiichi Tokuda, Mark Hasegawa-Johnson, Philipp Olbrich |
ICASSP | 8 |
| 2023 | Real-Time MRI Video Synthesis from Time Aligned Phonemes with Sequence-to-Sequence NetworksabstractReal-Time Magnetic resonance imaging (rtMRI) of the midsagittal plane of the mouth is of interest for speech production research. In this work, we focus on estimating utterance level rtMRI video from the spoken phoneme sequence. We obtain time-aligned phonemes from forced alignment, to obtain frame-level phoneme sequences which are aligned with rtMRI frames. We propose a sequence-tosequence learning model with a transformer phoneme encoder and convolutional frame decoder. We then modify the learning by using intermediary features obtained from sampling from a pretrained phoneme-conditioned variational autoencoder (CVAE). We train on 8 subjects in a subject-specific manner and demonstrate the performance with a subjective test. We also use an auxiliary task of air tissue boundary (ATB) segmentation to obtain the objective scores on the proposed models. We show that the proposed method is able to generate realistic rtMRI video for unseen utterances, and adding CVAE is beneficial for learning the sequence-to-sequence mapping for subjects where the mapping is hard to learn. Sathvik Udupa, Prasanta Kumar Ghosh |
ICASSP | 2 |
| 2023 | Improved Acoustic-to-Articulatory Inversion Using Representations from Pretrained Self-Supervised Learning ModelsabstractIn this work, we investigate the effectiveness of pretrained Self-Supervised Learning (SSL) features for learning the mapping for acoustic to articulatory inversion (AAI). Signal processing-based acoustic features such as MFCCs have been predominantly used for the AAI task with deep neural networks. With SSL features working well for various other speech tasks such as speech recognition, emotion classification, etc., we experiment with its efficacy for AAI. We train on SSL features with transformer neural networks-based AAI models of 3 different model complexities and compare its performance with MFCCs in subject-specific (SS), pooled and fine-tuned (FT) configurations with data from 10 subjects, and evaluate with correlation coefficient (CC) score on the unseen sentence test set. We find that acoustic feature reconstruction objective-based SSL features such as TERA and DeCoAR work well for AAI, with SS CCs of these SSL features reaching close to the best FT CCs of MFCC. We also find the results consistent across different model sizes. Sathvik Udupa, C. Siddarth, Prasanta Kumar Ghosh |
ICASSP | 3 |
| 2023 | Exploring a classification approach using quantised articulatory movements for acoustic to articulatory inversion
Jesuraja Bandekar, Sathvik Udupa, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2023 | Weakly supervised glottis segmentation in high-speed videoendoscopy using bounding box labels
Varun Belagali, M. V. Achuth Rao, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2023 | Transfer Learning to Aid Dysarthria Severity Classification for Patients with Amyotrophic Lateral Sclerosis
Tanuka Bhattacharjee, Anjali Jayakumar, Yamini Belur, Atchayaram Nalini, Prasanta Kumar Ghosh |
INTERSPEECH | 6 |
| 2023 | A Study on the Importance of Formant Transitions for Stop-Consonant Classification in VCV Sequence
Siddarth Chandrasekar, Arvind Ramesh, Tilak Purohit, Prasanta Kumar Ghosh |
INTERSPEECH | 4 |
| 2023 | An Investigation of Indian Native Language Phonemic Influences on L2 English Pronunciations
Shelly Jain, Priyanshi Pal, Anil Kumar Vuppala, Prasanta Kumar Ghosh, Chiranjeevi Yarra |
INTERSPEECH | 4 |
| 2023 | Classification of Multi-class Vowels and Fricatives From Patients Having Amyotrophic Lateral Sclerosis with Varied Levels of Dysarthria Severity
Chowdam Venkata Thirumala Kumar, Tanuka Bhattacharjee, Yamini Belur, Atchayaram Nalini, Prasanta Kumar Ghosh |
INTERSPEECH | 6 |
| 2023 | Do Vocal Breath Sounds Encode Gender Cues for Automatic Gender Classification?abstractThe acoustic features of continuous speech, such as pitch (F0) and formant frequencies (F1, F2) have been utilized for gender classification.However, non-speech signals including vocal breath sounds have not been explored due to the absence of gender-specific acoustic features.This study investigates if vocal breath sounds carry gender information and if they can be used for automatic gender classification.The study examines the use of data-driven and knowledge-based features from breath sounds, classifier complexity, and the importance of breath signal segment location and duration.Results from experiments on 54 minutes of male and 52 minutes of female breath sounds demonstrate that classifiers with low-complexity and knowledge-based features (MFCC statistics) perform similarly to high-complexity classifiers with data-driven features.Breath segments of around 3 seconds are found to be the most suitable choice regardless of location, eliminating the need for breath cycle boundary marking. Mohammad Shaique Solanki, Ashutosh Bharadwaj, Jeevan Kylash, Prasanta Kumar Ghosh |
INTERSPEECH | 4 |
| 2022 | The impact of cross language on acoustic-to-articulatory inversion and its influence on articulatory speech synthesisabstractEstimating articulatory representations (ARs) from acoustic features is known as acoustic-to-articulatory inversion (AAI). Various factors of input acoustic features impact the performance of AAI. In this work, we investigate the effect of unseen language on the AAI performance in both seen and unseen speaker conditions. We further perform experiments to analyze how these AAI predictions in unseen language and unseen speaker conditions, in turn, impact the articulatory speech synthesis, i.e., articulatory-to-acoustic forward mapping (AAF). We hypothesize that this investigation enables the exploration of alternative approaches to voice conversion across unseen languages using ARs. Experiments are performed on the AAF model trained using English ARs and evaluated on ARs from unseen speakers speaking different native Indian languages, namely, Hindi, Kannada, Telugu, and Tamil. Experiments reveal that, for AAI, there is a drop in performance due to the mismatch in language in both seen and unseen speaker evaluations. For AAF, subjective evaluations reveal that the synthesized speech quality of non-native (mismatched language) speech is comparable with that of English (matched language). Aravind Illa, Aanish Nair, Prasanta Kumar Ghosh |
ICASSP | 3 |
| 2022 | Dual Attention Pooling Network for Recording Device Classification Using Neutral and Whispered SpeechabstractIn this work, we proposed a method for recording device classification using the recorded speech signal. With the rapid increase in different mobile and professional recording devices, determining the source device has many applications in forensics and in further improving various speech-based applications. This paper proposes dual and single attention pooling-based convolutional neural networks (CNN) for recording device classification using neutral and whispered speech. Experiments using five recording devices with simultaneous direct recordings from 88 speakers speaking both in neutral and whisper and recordings from 21 mobile devices with simultaneous playback recordings reveal that the proposed dual attention pooling based CNN method performs better than the best baseline scheme. We show that we achieve a better performance in recording device classification with whispered speech recordings than corresponding neutral speech. We also demonstrate the importance of voiced/unvoiced speech and different frequency bands in classifying the recording devices. Abinay Reddy Naini, Bhavuk Singhal, Prasanta Kumar Ghosh |
ICASSP | 3 |
| 2022 | An Error Correction Scheme for Improved Air-Tissue Boundary in Real-Time MRI Video for Speech ProductionabstractThe best performance in Air-tissue boundary (ATB) segmentation of real-time Magnetic Resonance Imaging (rtMRI) videos in speech production is known to be achieved by a 3-dimensional convolutional neural network (3D-CNN) model. However, the evaluation of this model, as well as other ATB segmentation techniques reported in the literature, is done using Dynamic Time Warping (DTW) distance between the entire original and predicted contours. Such an evaluation measure may not capture local errors in the predicted contour. Careful analysis of predicted contours reveals errors in regions like the velum part of contour1 (ATB comprising of upper lip, hard palate, and velum) and tongue base section of contour2 (ATB covering jawline, lower lip, tongue base, and epiglottis), which are not captured in a global evaluation metric like DTW distance. In this work, we automatically detect such errors and propose a correction scheme for the same. We also propose two new evaluation metrics for ATB segmentation separately in contour1 and contour2 to explicitly capture two types of errors in these contours. The proposed detection and correction strategies result in an improvement of these two evaluation metrics by 61.8% and 61.4% for contour1 and by 67.8% and 28.4% for contour2. Traditional DTW distance, on the other hand, improves by 44.6% for contour1 and 4.0% for contour2. Anwesha Roy, Varun Belagali, Prasanta Kumar Ghosh |
ICASSP | 3 |
| 2022 | SegNet-Based Deep Representation Learning for Dysphagia ClassificationabstractSwallowing disorders, broadly known as Dysphagia, are difficulties in the process of swallowing food. Many currently available methods for classifying healthy and dysphagic swallows typically use hand-picked acoustic features. This article presents a SegNet-based method for classifying healthy and dysphagic swallow signals by learning mel-spectrogram features. Swallow sounds were recorded from a total of 24 subjects using a microphone based cervical auscultation (CA) system. Each subject swallowed multiple samples of water of volumes 5ml, 10ml and 15ml, and also performed multiple dry swallows. The experiments investigated the significance of temporal structures in the SegNet-learnt representations. The classification performance was evaluated at different model depths in order to identify the optimum feature time-scale that maximized the classification performance. The proposed method was found to be more robust to variations in the signatures of swallow signals across multiple volumes of water, against a baseline method across a single volume of water. The best performing model yielded a mean test F1-score of 80.13% (±4.62%) in a 5-fold cross validation setup. Siddharth Subramani, M. V. Achuth Rao, Anwesha Roy, Prasanna Suresh Hegde, Prasanta Kumar Ghosh |
ICASSP | 5 |
| 2022 | Gram Vaani ASR Challenge on spontaneous telephone speech recordings in regional variations of HindiabstractThis paper describes the corpus and baseline systems for the Gram Vaani Automatic Speech Recognition (ASR) challenge in regional variations of Hindi. The corpus for this challenge comprises the spontaneous telephone speech recordings collected by a social technology enterprise, Gram Vaani. The regional variations of Hindi together with spontaneity of speech, natural background and transcriptions with variable accuracy due to crowdsourcing make it a unique corpus for ASR on spontaneous telephonic speech. Around, 1108 hours of real-world spontaneous speech recordings, including 1000 hours of unlabelled training data, 100 hours of labelled training data, 5 hours of development data and 3 hours of evaluation data, have been released as a part of the challenge. The efficacy of both training and test sets are validated on different ASR systems in both traditional time-delay neural network-hidden Markov model (TDNN-HMM) frameworks and fully-neural end-to-end (E2E) setup. The word error rate (WER) and character error rate (CER) on eval set for a TDNN model trained on 100 hours of labelled data are 29.7 and 15.1, respectively. While, in E2E setup, WER and CER on eval set for a conformer model trained on 100 hours of data are 32.9 and 19.0, respectively. Anish Bhanushali, Grant Bridgman, Deekshitha G, Prasanta Kumar Ghosh, Pratik Kumar, Adithya Raj Kolladath, Nithya Ravi, Aaditeshwar Seth, Ashish Seth, Abhayjeet Singh, Vrunda N. Sukhadia, Srinivasan Umesh, Sathvik Udupa, Lodagala Durga Prasad |
INTERSPEECH | 4 |
| 2022 | Air tissue boundary segmentation using regional loss in real-time Magnetic Resonance Imaging video for speech productionabstractThe SegNet model has been shown to provide the best performance in air-tissue boundary (ATB) segmentation in real-time Magnetic Resonance Imaging (rtMRI) videos in seen subject conditions. The SegNet model uses overall binary cross entropy as the loss function. However, such a global loss function does not give enough emphasis on regions which are more prone to errors. In this work, together with global loss, we explore the use of regional loss functions which focus on areas of the contours which have been analysed as error prone in the past. Evaluation is done using global Dynamic Time Warping (DTW) distance as well as regional metrics. The regional metrics used are EVEL and VELrDTW for contour1, and ETB and TBrDTW for contour2. We show that using such combinations of regional and global losses improves the regional, as well as global, evaluation metrics. For the best combination of losses, the two regional metrics show an improvement of 37.2 and 25.3 for contour1 and 23.9 and 28.4 for contour2, over a baseline model which uses only global loss. Global DTW distance, on the other hand, improves by 11.2 for contour1 and 5.6 for contour2. Anwesha Roy, Varun Belagali, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2022 | Watch Me Speak: 2D Visualization of Human Mouth during Speech
C. Siddarth, Sathvik Udupa, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2022 | Streaming model for Acoustic to Articulatory Inversion with transformer networksabstractEstimating speech articulatory movements from speech acoustics is known as Acoustic to Articulatory Inversion (AAI). Recently, transformer-based AAI models have been shown to achieve state-of-art performance. However, in transformer networks, the attention is applied over the whole utterance, thereby needing to obtain the full utterance before the inference, which leads to high latency and is impractical for streaming AAI. To enable streaming during inference, evaluation could be performed on non-overlapping chucks instead of a full utterance. However, due to a mismatch of the attention receptive field during training and evaluation, there could be a drop in AAI performance. To overcome this scenario, in this work we perform experiments with different attention masks and use context from previous predictions during training. Experiments results revealed that using the random start mask attention with the context from previous predictions of transformer decoder performs better than the baseline results. Sathvik Udupa, Aravind Illa, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2022 | Automatic syllable stress detection under non-parallel label and data condition
Chiranjeevi Yarra, Prasanta Kumar Ghosh |
Speech Commun. | 2 |
| 2021 | Effect of Noise and Model Complexity on Detection of Amyotrophic Lateral Sclerosis and Parkinson's Disease Using Pitch and MFCCabstractDysarthria due to Amyotrophic Lateral Sclerosis (ALS) and Parkinson’s disease (PD) impacts both articulation and prosody in an individual’s speech. Complex deep neural networks exploit these cues for detection of ALS and PD. These are typically done using recordings in laboratory condition. This study aims to examine the robustness of these cues against background noise and model complexity, which has not been investigated before. We perform classification experiments with pitch and Mel-frequency cepstral coefficients (MFCC) using models of three different complexities and additive white Gaussian noise in four signal-to-noise-ratio (SNR) conditions. The findings are as follows: 1) In clean condition, pitch performs similar to MFCC across most model complexities considered, suggesting that one-dimensional pitch pattern provides discriminative cues for the classification to an extent equal to that of multi-dimensional MFCC, 2) Similar trend is observed in noisy cases when classifiers are trained and tested in matched noise and SNR conditions, 3) When the classifiers trained on clean data are applied in noisy cases, pitch based average classification accuracies are found to be 20.09% and 24.73% higher than those using MFCC for ALS vs. healthy and PD vs. healthy, respectively, suggesting robustness of pitch based classifier against noise and model complexity. Tanuka Bhattacharjee, Jhansi Mallela, Yamini Belur, Nalini Atchayarcmf, Pradeep Reddy, Dipanjan Gope, Prasanta Kumar Ghosh |
ICASSP | 8 |
| 2021 | Acoustic-to-Articulatory Inversion for Dysarthric Speech by Using Cross-Corpus Acoustic-Articulatory DataabstractIn this work, we focus on estimating articulatory movements from acoustic features, known as acoustic-to-articulatory inversion (AAI), for dysarthric patients with amyotrophic lateral sclerosis (ALS). Unlike healthy subjects, there are two potential challenges involved in AAI on dysarthric speech. Due to speech impairment, the pronunciation of dysarthric patients is unclear and inaccurate, which could impact the AAI performance. In addition, acoustic-articulatory data from dysarthric patients is limited due to the difficulty in the recording. These challenges motivate us to utilize cross-corpus acoustic-articulatory data. In this study, we propose an AAI model by conditioning speaker information using x-vectors at the input, and multi-target articulatory trajectory outputs for each corpus separately. Results reveal that the proposed AAI model shows relative improvements of the Pearson correlation coefficient (CC) by ~13.16% and ~16.45% over a randomly initialized baseline AAI model trained with only dysarthric corpus in seen and unseen conditions, respectively. In the seen conditions, the proposed AAI model outperforms the three baseline AAI models, that utilize the cross-corpus, by ~3.49%, ~6.46%, and ~4.03% in terms of CC. Sarthak Kumar Maharana, Aravind Illa, Renuka Mannem, Yamini Belur, Preetie Shetty, Preethish-Kumar Veeramani, Seena Vengalil, Kiran Polavarapu, Atchayaram Nalini, Prasanta Kumar Ghosh |
ICASSP | 10 |
| 2021 | Impact of Speaking Rate on the Source Filter Interaction in Speech: A StudyabstractSource filter interaction (SFI) explains the drop in pitch caused due to the constriction in the vocal tract during voiced consonant production in a vowel-consonant-vowel (VCV) sequence. In this work, we examine how the drop in pitch alters when such a VCV sequence is spoken at three different speaking rates - slow, normal and fast. In the absence of electroglottograph (EGG) recording, a high resolution pitch contour is determined using a glottal closure instant (GCI) detector. For this, in this work, firstly, five different GCI detector and pitch estimation techniques are compared against EGG based pitch estimates on a small dataset where simultaneous EGG recordings are available. Yet Another GCI Algorithm (YAGA) is found to be the best choice among all. For examining the impact of speaking rate on SFI, VCV recordings from six subjects with five vowels (/a/, /e/, /i/, /o/, /u/) and five consonants (/b/, /d/, /g/, /v/, /z/) at three speaking rates are used. The study reveals a significant difference in the pitch drop values between slow and fast rates, with increasing pitch drop as speaking rate reduces. For slow speaking rate, vowel /o/ and /u/ tend to show higher pitch drop values compared to remaining vowels. Tilak Purohit, M. V. Achuth Rao, Prasanta Kumar Ghosh |
ICASSP | 3 |
| 2021 | Source and Vocal Tract Cues for Speech-Based Classification of Patients with Parkinson's Disease and Healthy Subjects
Tanuka Bhattacharjee, Jhansi Mallela, Yamini Belur, Atchayaram Nalini, Pradeep Reddy, Dipanjan Gope, Prasanta Kumar Ghosh |
Interspeech | 8 |
| 2021 | MUCS 2021: Multilingual and Code-Switching ASR Challenges for Low Resource Indian LanguagesabstractRecently, there is increasing interest in multilingual automatic speech recognition (ASR) where a speech recognition system caters to multiple low resource languages by taking advantage of low amounts of labeled corpora in multiple languages. With multilingualism becoming common in today's world, there has been increasing interest in code-switching ASR as well. In code-switching, multiple languages are freely interchanged within a single sentence or between sentences. The success of low-resource multilingual and code-switching ASR often depends on the variety of languages in terms of their acoustics, linguistic characteristics as well as the amount of data available and how these are carefully considered in building the ASR system. In this challenge, we would like to focus on building multilingual and code-switching ASR systems through two different subtasks related to a total of seven Indian languages, namely Hindi, Marathi, Odia, Tamil, Telugu, Gujarati and Bengali. For this purpose, we provide a total of ~600 hours of transcribed speech data, comprising train and test sets, in these languages including two code-switched language pairs, Hindi-English and Bengali-English. We also provide a baseline recipe for both the tasks with a WER of 30.73% and 32.45% on the test sets of multilingual and code-switching subtasks, respectively. Anuj Diwan, Rakesh Vaideeswaran, Sanket Shah, Ankita Singh, Srinivasa Raghavan K. M., Shreya Khare, Vinit Unni, Saurabh Vyas, Akash Rajpuria, Chiranjeevi Yarra, Ashish R. Mittal, Prasanta Kumar Ghosh, Preethi Jyothi, Kalika Bali, Vivek Seshadri, Sunayana Sitaram, Samarth Bharadwaj, Jai Nanavati, Raoul Nanavati, Karthik Sankaranarayanan |
Interspeech | 12 |
| 2021 | DiCOVA Challenge: Dataset, Task, and Baseline System for COVID-19 Diagnosis Using AcousticsabstractThe DiCOVA challenge aims at accelerating research in diagnosing COVID-19 using acoustics (DiCOVA), a topic at the intersection of speech and audio processing, respiratory health diagnosis, and machine learning.This challenge is an open call for researchers to analyze a dataset of sound recordings, collected from COVID-19 infected and non-COVID-19 individuals, for a two-class classification.These recordings were collected via crowdsourcing from multiple countries, through a website application.The challenge features two tracks, one focusing on cough sounds, and the other on using a collection of breath, sustained vowel phonation, and number counting speech recordings.In this paper, we introduce the challenge and provide a detailed description of the task, and present a baseline system for the task. Ananya Muguli, Lancelot Pinto, Nirmala R., Neeraj Kumar Sharma 0001, Prashant Krishnan V, Prasanta Kumar Ghosh, Shrirama Bhat, Srikanth Raj Chetupalli, Sriram Ganapathy, Shreyas Ramoji, Viral Nanda |
Interspeech | 6 |
| 2021 | A Comparative Study of Different EMG Features for Acoustics-to-EMG MappingabstractElectromyography (EMG) signals have been extensively used to capture facial muscle movements while speaking since they are one of the most closely related bio-signals generated during speech production. In this work, we focus on speech acoustics to EMG prediction. We present a comparative study of ten different EMG signal-based features including Time Domain (TD) features existing in the literature to examine their effectiveness in speech acoustics to EMG inverse (AEI) mapping. We propose a novel feature based on the Hilbert envelope of the filtered EMG signal. The raw EMG signal is reconstructed from these features as well. For the AEI mapping, we use a bi-directional long short-term memory (BLSTM) network in a session-dependent manner. To estimate the raw EMG signal from the EMG features, we use a CNN-BLSTM model comprising of a convolution neural network (CNN) followed by BLSTM layers. AEI mapping performance using the BLSTM network reveals that the Hilbert envelope based feature is predicted from speech with the highest accuracy, among all the features. Therefore, it could be the most representative feature of the underlying muscle activation during speech production. The proposed Hilbert envelope feature, when used together with the existing TD features, improves the raw EMG signal reconstruction performance compared to using the TD features alone. Copyright © 2021 ISCA. Manthan Sharma, Navaneetha Gaddam, Tejas Umesh, Aditya Murthy, Prasanta Kumar Ghosh |
Interspeech | 5 |
| 2021 | Estimating Articulatory Movements in Speech Production with Transformer NetworksabstractWe estimate articulatory movements in speech production from different modalities - acoustics and phonemes. Acoustic-to articulatory inversion (AAI) is a sequence-to-sequence task. On the other hand, phoneme to articulatory (PTA) motion estimation faces a key challenge in reliably aligning the text and the articulatory movements. To address this challenge, we explore the use of a transformer architecture - FastSpeech, with explicit duration modelling to learn hard alignments between the phonemes and articulatory movements. We also train a transformer model on AAI. We use correlation coefficient (CC) and root mean squared error (rMSE) to assess the estimation performance in comparison to existing methods on both tasks. We observe 154%, 11.8% & 4.8% relative improvement in CC with subject-dependent, pooled and fine-tuning strategies, respectively, for PTA estimation. Additionally, on the AAI task, we obtain 1.5%, 3% and 3.1% relative gain in CC on the same setups compared to the state-of-the-art baseline. We further present the computational benefits of having transformer architecture as representation blocks. Sathvik Udupa, Anwesha Roy, Abhayjeet Singh, Aravind Illa, Prasanta Kumar Ghosh |
Interspeech | 5 |
| 2021 | Web Interface for Estimating Articulatory Movements in Speech Production from Acoustics and Text
Sathvik Udupa, Anwesha Roy, Abhayjeet Singh, Aravind Illa, Prasanta Kumar Ghosh |
Interspeech | 5 |
| 2021 | Noise Robust Pitch Stylization Using Minimum Mean Absolute Error CriterionabstractWe propose a pitch stylization technique in the presence of pitch halving and doubling errors. The technique uses an optimization criterion based on a minimum mean absolute error to make the stylization robust to such pitch estimation errors, particularly under noisy conditions. We obtain segments for the stylization automatically using dynamic programming. Experiments are performed at the frame level and the syllable level. At the frame level, the closeness of stylized pitch is analyzed with the ground truth pitch, which is obtained using a laryngograph signal, considering root mean square error (RMSE) measure. At the syllable level, the effectiveness of perceptual relevant embeddings in the stylized pitch is analyzed by estimating syllabic tones and comparing those with manual tone markings using the Levenshtein distance measure. The proposed approach performs better than a minimum mean squared error criterion based pitch stylization scheme at the frame level and a knowledge-based tone estimation scheme at the syllable level under clean and 20dB, 10dB and 0dB SNR conditions with five noises and four pitch estimation techniques. Among all the combinations of SNR, noise and pitch estimation techniques, the highest absolute RMSE and mean distance improvements are found to be 6.49Hz and 0.23, respectively. Copyright © 2021 ISCA. Chiranjeevi Yarra, Prasanta Kumar Ghosh |
Interspeech | 2 |
| 2021 | A deep neural network based correction scheme for improved air-tissue boundary prediction in real-time magnetic resonance imaging video
Renuka Mannem, Prasanta Kumar Ghosh |
Comput. Speech Lang. | 2 |
| 2020 | Voice based classification of patients with Amyotrophic Lateral Sclerosis, Parkinson's Disease and Healthy Controls with CNN-LSTM using transfer learningabstractIn this paper, we consider 2-class and 3-class classification problems for classifying patients with Amyotrophic Lateral Sclerosis (ALS), Parkinson's Disease (PD), and Healthy Controls (HC) using a CNNLSTM network. Classification performance is examined for three different tasks, namely, Spontaneous speech (SPON), Diadochokinetic rate (DIDK) and Sustained phoneme production (PHON). Experiments are conducted using speech data recorded from 60 ALS, 60 PD, and 60 HC subjects. Classifications using SVM and DNN are considered as baseline schemes. Classification accuracy of ALS and HC (indicated by ALS/HC) using CNN-LSTM has shown an improvement of 10.40%, 4.22% and 0.08% for PHON, SPON and DIDK tasks, respectively over the best of the baseline schemes. Furthermore, the CNN-LSTM network achieves the highest PD/HC classification accuracy of 88.5% for the SPON task and the highest 3-class (ALS/PD/HC) classification accuracy of 85.24% for the DIDK task. Experiments using transfer learning at low resource training data show that data from ALS benefits PD/HC classification and vice-versa. Experiments with fine-tuning weights of 3-class (ALS/PD/HC) classifier for 2-class classification (PD/HC or ALS/HC) gives an absolute improvement of 2% classification accuracy in SPON task when compared with randomly initialized 2-class classifier. Jhansi Mallela, Aravind Illa, Suhas B. N., Sathvik Udupa, Yamini Belur, Atchayaram Nalini, Pradeep Reddy, Dipanjan Gope, Prasanta Kumar Ghosh |
ICASSP | 10 |
| 2020 | Pseudo Likelihood Correction Technique for Low Resource Accented ASRabstractWith the availability of large data, ASRs perform well on native English but poorly for non-native English data. Training nonnative ASRs or adapting a native English ASR is often limited by the availability of data, particularly for low resource scenarios. A typical HMM-DNN based ASR decoding requires pseudo-likelihood of states given an acoustic observation, which changes significantly from native to non-native speech due to accent variation. In order to improve the performance of a native English ASR on non-native English data, we, in this work, propose a DNN-based pseudo-likelihood correction (PLC) technique, in which a non-native pseudo-likelihood vector is mapped to match its native counterpart. Instead of correcting all elements of a non-native pseudo-likelihood vector, a loss function is proposed to correct only top few of them. Experiments with one native and multiple Indian English corpora show an improvement of WER by ~12% and over ~5% using the proposed PLC technique unadapted and adapted native English ASR respectively, when recognition is performed on an Indian English corpus different from that used for both PLC and adaptation. Experiments with upto 2 hours of parallel native and non-native English data reveal that, PLC performs better than adaptation for all unseen cases considered. Avni Rajpal, M. V. Achuth Rao, Chiranjeevi Yarra, Ritu Aggarwal, Prasanta Kumar Ghosh |
ICASSP | 5 |
| 2020 | A Comparative Study of Estimating Articulatory Movements from Phoneme Sequences and Acoustic FeaturesabstractUnlike phoneme sequences, movements of speech articulators (lips, tongue, jaw, velum) and the resultant acoustic signal are known to encode not only the linguistic message but also carry para-linguistic information. While several works exist for estimating articulatory movement from acoustic signals, little is known to what extent articulatory movements can be predicted only from linguistic information, i.e., phoneme sequence. In this work, we estimate articulatory movements from three different input representations: R1) acoustic signal, R2) phoneme sequence, R3) phoneme sequence with timing information. While an attention network is used for estimating articulatory movement in the case of R2, BLSTM network is used for R1 and R3. Experiments with ten subjects’ acoustic-articulatory data reveal that the estimation techniques achieve an average correlation coefficient of 0.85, 0.81, and 0.81 in the case of R1, R2, and R3 respectively. This indicates that attention network, although uses only phoneme sequence (R2) without any timing information, results in an estimation performance similar to that using rich acoustic signal (R1), suggesting that articulatory motion is primarily driven by the linguistic message. The correlation coefficient is further improved to 0.88 when R1 and R3 are used together for estimating articulatory movements. Abhayjeet Singh, Aravind Illa, Prasanta Kumar Ghosh |
ICASSP | 3 |
| 2020 | Automatic Classification of Volumes of Water Using Swallow Sounds from Cervical AuscultationabstractThe signatures of swallowing vary depending on the volume of bolus swallowed. Among existing instrumental methods, cervical auscultation (CA) captures the acoustic signatures of the swallow sound. Although many features present in the literature can characterize volumes of swallow using CA, they require manual annotations of the different components in the sound. In this work, a rich set of acoustic features, the ComParE 2016 acoustic feature set (OS) is used to investigate whether several temporal, spectral, vocal and source features and their functionals provide cues for volume classification. Experiments are performed with CA data from 56 subjects, with dry swallow and swallows of 2ml, 5ml, and 10ml of water. Three types of classification namely, dry-vs-2ml, dry-vs-5ml and dry-vs-10ml are performed separately to analyze characteristic features. Experiments reveal that OS, which does not require annotations, performs better than the baseline features that require annotation. Within OS, the features unrelated to voice source yield a better performance than the features related to voice source. In this subset of features, MFCC, RASTA filtered audio spectrum and RMS energy are found to be consistently the top performing features across all three types of classifications. Siddharth Subramani, M. V. Achuth Rao, Divya Giridhar, Prasanna Suresh Hegde, Prasanta Kumar Ghosh |
ICASSP | 5 |
| 2020 | Automatic Identification of Speakers From Head Gestures in a NarrationabstractIn this work, we focus on quantifying speaker identity information encoded in the head gestures of speakers, while they narrate a story. We hypothesize that the head gestures over a long duration have speaker-specific patterns. To establish this, we consider a classification problem to identify speakers from head gestures. We represent every head orientation as a triplet of Euler angles and a sequence of head orientations as head gestures. We use a database having recordings from 24 speakers where the head movements are recorded using a motion capture device, with each subject narrating ten stories. We get the best speaker identification accuracy of 0.836 using head gestures over a duration of 40 seconds. Further, the accuracy increases by combining decisions from multiple 40 second windows when a recording is available with duration more than the window length. We achieve an average accuracy of 0.9875 on our database when the entire recording is used. Analysis of the speaker identification performance over 40 second windows across a recording reveals that the speaker-identity information is more prevalent in some parts of a story than others. Sanjeev Kadagathur Vadiraj, M. V. Achuth Rao, Prasanta Kumar Ghosh |
ICASSP | 3 |
| 2020 | Analysis of Acoustic Features for Speech Sound Based Classification of Asthmatic and Healthy SubjectsabstractNon-speech sounds (cough, wheeze) are typically known to perform better than speech sounds for asthmatic and healthy subject classification. In this work, we use sustained phonations of speech sounds, namely, /α:/, /i:/, /u:/, /eI/, /ou/, /s/, and /z/ from 47 asthmatic and 48 healthy controls. We consider INTERSPEECH 2013 Computational Paralinguistics Challenge baseline (ISCB) acoustic features for the classification task as they provide a rich set of characteristics of the speech sounds. Mel-frequency cepstral coefficients (MFCC) are used as the baseline features. The classification accuracy using ISCB improves over MFCC for all voiced speech sounds with the highest classification accuracy of 75.4% (18.28% better than baseline) for /ou/. The exhale achieves the highest classification accuracy of 77.8% (4.2% better than baseline). Comparable accuracies using speech sound /ou/ and non-speech exhale indicate the benefit of the rich acoustic features from ISCB. An analysis of 21 ISCB features groups using forward feature group selection shows that loudness and MFCC groups contribute the most in the case of /ou/, with interquartile range between 2ndand 3rdquartile of loudness feature being the best discriminator feature. Merugu Keerthana, Dipanjan Gope, Uma Maheswari Krishnaswamy, Prasanta Kumar Ghosh |
ICASSP | 5 |
| 2020 | Automatic Glottis Detection and Segmentation in Stroboscopic Videos Using Convolutional Networks
Divya Degala, M. V. Achuth Rao, Rahul Krishnamurthy, Pebbili Gopikishore, Veeramani Priyadharshini, Prakash T. K., Prasanta Kumar Ghosh |
INTERSPEECH | 7 |
| 2020 | Speaker Conditioned Acoustic-to-Articulatory Inversion Using x-VectorsabstractSpeech production involves the movement of various articulators, including tongue, jaw, and lips. Estimating the movement of the articulators from the acoustics of speech is known as acoustic-to-articulatory inversion (AAI). Recently, it has been shown that instead of training AAI in a speaker specific manner, pooling the acoustic-articulatory data from multiple speakers is beneficial. Further, additional conditioning with speaker specific information by one-hot encoding at the input of AAI along with acoustic features benefits the AAI performance in a closed-set speaker train and test condition. In this work, we carry out an experimental study on the benefit of using x-vectors for providing speaker specific information to condition AAI. Experiments with 30 speakers have shown that the AAI performance benefits from the use of x-vectors in a closed set seen speaker condition. Further, x-vectors also generalizes well for unseen speaker evaluation. Aravind Illa, Prasanta Kumar Ghosh |
INTERSPEECH | 2 |
| 2020 | Raw Speech Waveform Based Classification of Patients with ALS, Parkinson's Disease and Healthy Controls Using CNN-BLSTM
Jhansi Mallela, Aravind Illa, Yamini Belur, Atchayaram Nalini, Pradeep Reddy, Dipanjan Gope, Prasanta Kumar Ghosh |
INTERSPEECH | 8 |
| 2020 | Air-Tissue Boundary Segmentation in Real Time Magnetic Resonance Imaging Video Using 3-D Convolutional Neural NetworkabstractThe real-time Magnetic Resonance Imaging (rtMRI) is often used for speech production research as it captures the complete view of the vocal tract during speech. Air-tissue boundaries (ATBs) are the contours that trace the transition between high-intensity tissue region and low-intensity airway cavity region in an rtMRI video. The ATBs are used in several speech related applications. However, the ATB segmentation is a challenging task as the rtMRI frames have low resolution and low signal-to-noise ratio. Several works have been proposed in the past for ATB segmentation. Among these, the supervised algorithms have been shown to perform well compared to the unsupervised algorithms. However, the supervised algorithms have limited generalizability towards subjects not involved in training. In this work, we propose a 3-dimensional convolutional neural network (3D-CNN) which utilizes both spatial and temporal information from the rtMRI video for accurate ATB segmentation. The 3D-CNN model captures the vocal tract dynamics in an rtMRI video independent of the morphology of the subject leading to an accurate ATB segmentation for unseen subjects. In a leave-one-subject-out experimental setup, it is observed that the proposed approach provides ~32 relative improvement in the performance compared to the best (SegNet based) baseline approach. © 2020 ISCA Renuka Mannem, Navaneetha Gaddam, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2020 | Speech Rate Task-Specific Representation Learning from Acoustic-Articulatory DataabstractIn this work, speech rate is estimated using the task-specific representations which are learned from the acoustic-articulatory data, in contrast to generic representations which may not be optimal for the speech rate estimation. 1-D convolutional filters are used to learn speech rate specific acoustic representations from the raw speech. A convolutional dense neural network (CDNN) is used to estimate the speech rate from the learned representations. In practice, articulatory data is not directly available; thus, we use Acoustic-to-Articulatory Inversion (AAI) to derive the articulatory representations from acoustics. However, these pseudo-articulatory representations are also generic and not optimized for any task. To learn the speech-rate specific pseudo-articulatory representations, we propose a joint training of BLSTM-based AAI and CDNN using a weighted loss function that considers the losses corresponding to speech rate estimation and articulatory prediction. The proposed model yields an improvement in speech rate estimation by ~18.5 in terms of pearson correlation coefficient (CC) compared to the baseline CDNN model with generic articulatory representations as inputs. To utilize complementary information from articulatory features, we further perform experiments by concatenating task-specific acoustic and pseudo-articulatory representations, which yield an improvement in CC by ~2.5 compared to the baseline CDNN model. © 2020 ISCA Renuka Mannem, Hima Jyothi, Aravind Illa, Prasanta Kumar Ghosh |
INTERSPEECH | 4 |
| 2020 | Whisper Activity Detection Using CNN-LSTM Based Attention Pooling Network Trained for a Speaker Identification TaskabstractIn this work, we proposed a method to detect the whispered speech region in a noisy audio file called whisper activity detection (WAD). Due to the lack of pitch and noisy nature of whispered speech, it makes WAD a way more challenging task than standard voice activity detection (VAD). In this work, we proposed a Long-short term memory (LSTM) based whisper activity detection algorithm. However, this LSTM network is trained by keeping it as an attention pooling layer to a Convolutional neural network (CNN), which is trained for a speaker identification task. WAD experiments with 186 speakers, with eight noise types in seven different signal-to-noise ratio (SNR) conditions, show that the proposed method performs better than the best baseline scheme in most of the conditions. Particularly in the case of unknown noises and environmental conditions, the proposed WAD performs significantly better than the best baseline scheme. Another key advantage of the proposed WAD method is that it requires only a small part of the training data with annotation to fine-tune the post-processing parameters, unlike the existing baseline schemes requiring full training data annotated with the whispered speech regions. Copyright © 2020 ISCA Abinay Reddy Naini, Malla Satyapriya, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2020 | An Investigation of the Virtual Lip Trajectories During the Production of Bilabial Stops and Nasal at Different Speaking RatesabstractWe propose a technique to estimate virtual upper lip (VUL) and virtual lower lip (VLL) trajectories during production of bilabial stop consonants (/p/, /b/) and nasal (/m/). A VUL (VLL) is a hypothetical trajectory below (above) the measured UL (LL) trajectory which could have been achieved by UL (LL) if UL and LL were not in contact with each other during bilabial stops and nasal. Maximum deviation of UL from VUL and its location as well as the range of VUL are used as features, denoted by VUL MD, VUL MDL, and VUL R, respectively. Similarly, VLL MD, VLL MDL, and VLL R are also computed. Analyses of these six features are carried out for /p/, /b/, and /m/ at slow, normal and fast rates based on electromagnetic articulograph (EMA) recordings of VCV stimuli spoken by ten subjects. While no significant differences were observed among /p/, /b/, and /m/ in every rate, all six features except VLL MD were found to drop significantly from slow to fast rates. These six features were also found to perform better in an automatic classification task between slow vs fast rates compared to five baseline features computed from UL and LL comprising their ranges, velocities and minimum distance from each other. © 2020 ISCA Tilak Purohit, Prasanta Kumar Ghosh |
INTERSPEECH | 2 |
| 2020 | Coswara - A Database of Breathing, Cough, and Voice Sounds for COVID-19 DiagnosisabstractThe COVID-19 pandemic presents global challenges transcending boundaries of country, race, religion, and economy.The current gold standard method for COVID-19 detection is the reverse transcription polymerase chain reaction (RT-PCR) testing.However, this method is expensive, time-consuming, and violates social distancing.Also, as the pandemic is expected to stay for a while, there is a need for an alternate diagnosis tool which overcomes these limitations, and is deployable at a large scale.The prominent symptoms of COVID-19 include cough and breathing difficulties.We foresee that respiratory sounds, when analyzed using machine learning techniques, can provide useful insights, enabling the design of a diagnostic tool.Towards this, the paper presents an early effort in creating (and analyzing) a database, called Coswara, of respiratory sounds, namely, cough, breath, and voice.The sound samples are collected via worldwide crowdsourcing using a website application.The curated dataset is released as open access.As the pandemic is evolving, the data collection and analysis is a work in progress.We believe that insights from analysis of Coswara can be effective in enabling sound based technology solutions for point-of-care diagnosis of respiratory infection, and in the near future this can help to diagnose COVID-19. Neeraj Kumar Sharma 0001, Prashant Krishnan V, Shreyas Ramoji, Srikanth Raj Chetupalli, Nirmala R., Prasanta Kumar Ghosh, Sriram Ganapathy |
INTERSPEECH | 7 |
| 2020 | Attention and Encoder-Decoder Based Models for Transforming Articulatory Movements at Different Speaking RatesabstractWhile speaking at different rates, articulators (like tongue, lips) tend to move differently and the enunciations are also of different durations. In the past, affine transformation and DNN have been used to transform articulatory movements from neutral to fast(N2F) and neutral to slow(N2S) speaking rates [1]. In this work, we improve over the existing transformation techniques by modeling rate specific durations and their transformation using AstNet, an encoder-decoder framework with attention. In the current work, we propose an encoder-decoder architecture using LSTMs which generates smoother predicted articulatory trajectories. For modeling duration variations across speaking rates, we deploy attention network, which eliminates the needto align trajectories in different rates using DTW. We performa phoneme specific duration analysis to examine how well duration is transformed using the proposed AstNet. As the range of articulatory motions is correlated with speaking rate, we also analyze amplitude of the transformed articulatory movements at different rates compared to their original counterparts, to examine how well the proposed AstNet predicts the extent of articulatory movements in N2F and N2S. We observe that AstNet could model both duration and extent of articulatory movements better than the existing transformation techniques resulting in more accurate transformed articulatory trajectories. Abhayjeet Singh, Aravind Illa, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2020 | The impact of speaking rate on acoustic-to-articulatory inversion
Aravind Illa, Prasanta Kumar Ghosh |
Comput. Speech Lang. | 2 |
| 2020 | SFNet: A Computationally Efficient Source Filter Model Based Neural Speech SynthesisabstractRecently, neural speech synthesizers have achieved a high-quality synthesis for text-to-speech applications, but a real-time synthesis is possible only in the devices which have high memory and allow large computational complexity. In this work, we reduce the complexity of a speech synthesizer by reformulating the source-filter model of speech where the excitation signal is modeled as a sum of two signals. The first signal contains an impulse train that is computed from the pitch sequence. The second signal is modeled as white noise passed through a filter bank with frequency dependent gains. The parameters of the reformulated source-filter model are predicted using a neural network, referred to as SFNet. The network parameters are learnt by training the network using L1-error between the log Mel-spectrum of the predicted waveform and that of the ground-truth waveform. We demonstrate that there is a significant reduction in the memory and computational complexity compared to the state-of-the-art speaker independent neural speech synthesizer without any loss of the naturalness of the synthesized speech. M. V. Achuth Rao, Prasanta Kumar Ghosh |
IEEE Signal Process. Lett. | 2 |
| 2019 | Representation Learning Using Convolution Neural Network for Acoustic-to-articulatory InversionabstractRecent techniques employ end-to-end systems to learn relevant features for several speech related applications, including speech recognition, and speaker verification. In this work, we focus on the task of acoustic-to-articulatory inversion (AAI) for which we propose an end-to-end system that comprises a convolution neural network (CNN) and a bidirectional long short-term memory network (BLSTM). The aim of this work is to understand the nature of the features learnt by the end-to-end model and the importance of pre-emphasis in representation learning for AAI. Further, we propose a subject adaptation scheme to overcome the limitations of the availability of parallel acoustic-articulatory data to train an end-to-end AAI system. The AAI performance is evaluated with ~3.19 hours of acoustic-articulatory data collected from 8 subjects. Experiments reveal that, the frequency response of filters learnt by the CNN in the proposed system resembles those of the mel-scale, and hence, the performance of the proposed system (RMSE=1.47mm) is on par with that using mel-frequency cepstral coefficients (1.42mm) as features. Using pre-emphasis reduces RMSE by 0.13mm, and also the proposed adaptation scheme performs better than a subject-specific AAI model by an RMSE of 0.21mm despite of limited acoustic-articulatory data from a subject. Aravind Illa, Prasanta Kumar Ghosh |
ICASSP | 2 |
| 2019 | Air-tissue Boundary Segmentation in Real Time Magnetic Resonance Imaging Video Using a Convolutional Encoder-decoder NetworkabstractIn this paper, we propose a convolutional encoder-decoder network (CEDN) based approach for upper and lower Air-Tissue Boundary (ATB) segmentation within vocal tract in real-time magnetic resonance imaging (rtMRI) video frames. The output images from CEDN are processed using perimeter and moving average filters to generate smooth contours representing ATBs. Experiments are performed in both seen subject and unseen subject conditions to examine the generalizability of the CEDN based approach. The performance of the segmented ATBs is evaluated using Dynamic Time Warping distance between the ground truth contours and predicted contours. The proposed approach is compared with three baseline schemes - one grid-based unsupervised and two supervised schemes. Experiments with 5779 rtMRI images from four subjects show that the CEDN based approach performs better than the unsupervised baseline scheme by 8.5% for seen subjects case whereas it does better than the supervised baseline schemes only for lower ATB. For unseen subjects case, the proposed approach performs better than the supervised baseline schemes by 63.96%, 22.9% respectively whereas it performs worse than the unsupervised baseline scheme. However, the proposed approach outperforms the unsupervised baseline scheme when a minimum of 30 images from unseen subjects are used to adapt the trained CEDN model. Renuka Mannem, Prasanta Kumar Ghosh |
ICASSP | 2 |
| 2019 | Formant-gaps Features for Speaker Verification Using Whispered SpeechabstractIn this work, we propose a new feature based on formants for whispered speaker verification (SV) task, where neutral data is used for enrollment and whispered recordings are used for test. Such a mismatch between enrollment and test often degrades the performance of whispered SV systems due to the difference in acoustic characteristics of whispered and neutral speech. We hypothesize that the proposed formant and formant gap (F oG) features are more invariant to the modes of speech in capturing speaker specific information compared to traditional baseline features for SV including mel frequency cepstral coefficients (MFCC) and auditory-inspired amplitude modulation features (AAMF). Whispered SV experiments with 714 speakers comprising 29232 neutral and 22932 whispered recordings reveal that the equal error rate (EER) using the proposed features is lower than that using the best baseline features by ~3.79% (absolute). It was also observed that at least four whispered recordings during enrollment are required for the baseline features to perform at par with the proposed features. However, it was found that the best performing baseline features yield an EER for neutral SV task which is ~1.88% higher than that using the proposed features. Abinay Reddy Naini, M. V. Achuth Rao, Prasanta Kumar Ghosh |
ICASSP | 3 |
| 2019 | A Study on Robustness of Articulatory Features for Automatic Speech Recognition of Neutral and Whispered SpeechabstractTraditionally, automatic speech recognition (ASR) systems are trained on acoustic representations of neutral speech. As a result, their performance degrades when tested with whispered speech. In this work, we explore the robustness of articulatory features in ASR of neutral and whispered speech. We use acoustic, articulatory, and integrated acoustic and articulatory feature vectors in matched and mismatched train-test cases. The results suggest that the articulatory data is useful in ASR of both neutral and whispered speech, especially in the mismatched train-test cases. When we concatenate acoustic and articulatory feature vectors and deploy it to the mismatched train-test case where the model is trained with neutral speech and tested with whispered speech, a relative improvement in phone error rate of 27.2% is observed compared to when only acoustic features are used. This suggests that articulatory data contains information complementary to acoustic representations. A phone specific recognition error is also presented which illustrates phones where adding articulatory information gives maximum benefit. Gokul Srinivasan, Aravind Illa, Prasanta Kumar Ghosh |
ICASSP | 3 |
| 2019 | An Improved Air Tissue Boundary Segmentation Technique for Real Time Magnetic Resonance Imaging Video Using SegnetabstractThis paper presents an improved methodology for the segmentation of the Air-Tissue boundaries (ATBs) in the upper airway of the human vocal tract using Real-Time Magnetic Resonance Imaging (rtMRI) videos. Semantic segmentation is deployed in the proposed approach using a Deep learning architecture called SegNet. The network processes an input image to produce a binary output image of the same dimensions having classified each pixel as air cavity or tissue, following which contours are predicted. A Multi-dimensional least square smoothing technique is applied to smoothen the contours. To quantify the precision of predicted contours, Dynamic Time Warping (DTW) distance is calculated between the predicted contours and the manually annotated ground truth contour. Four fold experiments are conducted with four subjects from the USC-TIMIT corpus, which demonstrates that the proposed approach achieves a lower DTW distance of 1.02 and 1.09 for the upper and lower ATB compared to the best baseline scheme. The proposed SegNet based approach has an average pixel classification accuracy of 99.3% across all the subjects with only 2 rtMRI videos (~180 frames) per subject for training. C. A. Valliappan, Renuka Mannem, Girija Ramesan Karthik, Prasanta Kumar Ghosh |
ICASSP | 5 |
| 2019 | An Investigation on Speaker Specific Articulatory Synthesis with Speaker Independent Articulatory InversionabstractEstimating speech representations from articulatory movements is known as articulatory-to-acoustic forward (AAF) mapping. Typically this mapping is learned using directly measured articulatory movement in a subject-specific manner. Such AAF mapping has been shown to benefit the speech synthesis applications. In this work, we investigate the speaker similarity and naturalness of utterances generated by AAF which is driven by the articulatory movements from a subject (referred to as cross speaker) different from the speaker (target speaker) used for training AAF mapping. Experiments are performed with directly measured articulatory data from 9 speakers (8 target speakers and 1 cross speaker), which are recorded using Electromagnetic articulograph AG501. Experiments are also performed with articulatory features estimated using speaker independent acoustic-to-articulatory inversion (SI-AAI) model trained on 26 reference speakers. Objective evaluation on target speakers reveal that the articulatory features estimated from SI-AAI result in a lower Mel-cepstrum distortion compared to that using directly measured articulatory features. Further, listening tests reveal that the directly measured articulatory movements preserve the speaker similarity better than estimated ones. Although, for naturalness, articulatory movements predicted by SI-AAI perform better than the direct measurements. Aravind Illa, Prasanta Kumar Ghosh |
INTERSPEECH | 2 |
| 2019 | Acoustic and Articulatory Feature Based Speech Rate Estimation Using a Convolutional Dense Neural NetworkabstractIn this paper, we propose a speech rate estimation approach using a convolutional dense neural network (CDNN). The CDNN based approach uses the acoustic and articulatory features for speech rate estimation. The Mel Frequency Cepstral Coefficients (MFCCs) are used as acoustic features and the articulograms representing time-varying vocal tract profile are used as articulatory features. The articulogram is computed from a real-time magnetic resonance imaging (rtMRI) video in the midsagittal plane of a subject while speaking. However, in practice, the articulogram features are not directly available, unlike acoustic features from speech recording. Thus, we use an Acoustic-to-Articulatory Inversion method using a bidirectional long-short-term memory network which estimates the articulogram features from the acoustics. The proposed CDNN based approach using estimated articulatory features requires both acoustic and articulatory features during training but it requires only acoustic data during testing. Experiments are conducted using rtMRI videos from four subjects each speaking 460 sentences. The Pearson correlation coefficient is used to evaluate the speech rate estimation. It is found that the CDNN based approach gives a better correlation coefficient than the temporal and selected sub-band correlation (TCSSBC) based baseline scheme by 81.58 and 73.68 (relative) in seen and unseen subject conditions respectively. Renuka Mannem, Jhansi Mallela, Aravind Illa, Prasanta Kumar Ghosh |
INTERSPEECH | 4 |
| 2019 | Comparison of Speech Tasks and Recording Devices for Voice Based Automatic Classification of Healthy Subjects and Patients with Amyotrophic Lateral Sclerosis
Suhas B. N., Deep Patel, Nithin Rao Koluguri, Yamini Belur, Pradeep Reddy, Atchayaram Nalini, Dipanjan Gope, Prasanta Kumar Ghosh |
INTERSPEECH | 9 |
| 2019 | Whisper to Neutral Mapping Using Cosine Similarity Maximization in i-Vector Space for Speaker VerificationabstractIn this work, we propose a novel feature mapping (FM) from whispered to neutral speech features using a cosine similarity based objective function for speaker verification (SV) using whispered speech. Typically the performance of an SV system enrolled with neutral speech degrades significantly when tested using whispered speech, due to the differences between spectral characteristics of neutral and whispered speech. We hypothesize that FM from whispered Mel frequency cepstral coefficients (MFCC) to neutral MFCC by maximizing cosine similarity between neutral and whisper i-vectors yields better performance than the baseline method, which typically performs a direct FM between MFCC features by minimizing mean squared error (MSE). We also explored an affine transform between MFCC features using the proposed objective function. Whisper SV experiments with 1882 speakers reveal that the equal error rate (EER) using the proposed method is lower than that using the best baseline by ∼24% (relative). We show that the proposed FM system maintains the neutral SV performance, while improving the EER of whispered SV unlike baseline methods. We also show that the bias in the learned affine transform is corresponds to the glottal flow information, which is absent in the whispered speech. Abinay Reddy Naini, M. V. Achuth Rao, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2019 | ASR Inspired Syllable Stress Detection for Pronunciation Evaluation Without Using a Supervised Classifier and Syllable Level FeaturesabstractABSTRAKTBu məqalədə milli iqtisadiyyatın özünəməxsus xüsusiyyətlərini özündə cəmləşdirən DSÜT modelinin ekonometrik qiymətləndirməsi aparılır.Empirik qiymətləndirmə zamanı rüblük məlumatlara əsaslanılır və müşahidə sayının məhdud olması nəzərə alınıb Bayez metodlarına müraciət edilir.Qiymətləndirmələr bir sıra maraqlı məqamların ortaya çıxmasına şərait yaradır.İlk öncə məlum olur ki, milli iqtisadiyyat üçün qurulan əvvəlki yeni Keynezçi modellərdə istifadə edilən bir sıra vacib parametrlərin kalibrasiya praktikası əldə edilən empirik nəticələrlə uzlaşmır və ölkənin xüsusiyyətlərini əks etdirmir.İkincisi, aydın olur ki, əksər struktur parametrlər dövrü stabillik sərgiləsələr də, milli iqtisdiyyatı sarsan şokların strukturunda mühim dəyişiklər baş vermişdir.Bu tapıntı post-neft bumu dövründə proqnozlaşdırma işini çətinləşdirən amillərdən biri hesab oluna bilər.Üçüncüsü, pul kütləsi üzrə qiymətləndirilən parametrlərin identifikasiyasında problemlərin mövcud olduğu aşkardır.Bununla yanaşı, qiymətləndirilən model bir sıra adekvatlıq sınaqlarından uğurla keçir.Model siyasət qurumları tərəfindən müxtəlif ssenari analizlərinin aparılması və proqnozlaşdırma məqsədi üçün istifadə edilə bilər.Həmçinin, qurulan model ölkə iqtisadiyyatının özünəməxsusluğunu özündə ehtiva edən və ekonometrik qiymətləndirməsi aparılan ilk DSÜT modeli olması səbəbindən də maraq kəsb edir. Manoj Kumar Ramanathi, Chiranjeevi Yarra, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2019 | Low Resource Automatic Intonation Classification Using Gated Recurrent Unit (GRU) Networks Pre-Trained with Synthesized Pitch PatternsabstractSecond language learners of British English (BE) are typically trained to learn four intonation classes - Glide-up, Glide-down, Dive and Take-off. We predict the intonation class in a learner's utterance by modeling the temporal dependencies in the pitch patterns with gated recurrent unit (GRU) networks. For these, we pre-train the GRU network using a set of synthesized pitch patterns representing each intonation class. For the synthesis, we propose to obtain pitch patterns from the tone sequences representing each intonation class obtained from domain knowledge. Experiments are conducted on speech data collected from experts in a spoken English training material for teaching BE intonation. The absolute improvements in the unweighted average recall (UAR) using the proposed scheme with pre-training are found to be 4.14 and 6.01 respectively over the proposed approach without pre-training and the baseline scheme that uses hidden Markov models (HMMs). Atreyee Saha, Chiranjeevi Yarra, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2019 | An Improved Goodness of Pronunciation (GoP) Measure for Pronunciation Evaluation with DNN-HMM System Considering HMM Transition ProbabilitiesabstractGoodness of pronunciation (GoP) is typically formulated with Gaussian mixture model-hidden Markov model (GMM-HMM) based acoustic models considering HMM state transition probabilities (STPs) and GMM likelihoods of context dependent phonemes. On the other hand, deep neural network (DNN)HMM based acoustic models employed sub-phonemic (senone) posteriors instead of GMM likelihoods along with STPs. However, each senone is shared across many states; thus, there is no one-to-one correspondence between them. In order to circumvent this, most of the existing works have proposed modifications to the GoP formulation considering only posteriors neglecting the STPs. In this work, we derive a formulation for the GoP and it results in the formulation involving both senone posteriors and STPs. Further, we illustrate the steps to implement the proposed GoP formulation in Kaldi, a state-of-the-art automatic speech recognition toolkit. Experiments are conducted on English data collected from Indian speakers using acoustic models trained with native English data from LibriSpeech and Fisher-English corpora. The highest improvement in the correlation coefficient between the scores from the formulations and the expert ratings is found to be 14.89 (relative) better with the proposed approach compared to the best of the existing formulations that don't include STPs. Sweekar Sudhakara, Manoj Kumar Ramanathi, Chiranjeevi Yarra, Prasanta Kumar Ghosh |
INTERSPEECH | 4 |
| 2019 | SPIRE-fluent: A Self-Learning App for Tutoring Oral Fluency to Second Language English Learners
Chiranjeevi Yarra, Aparna Srinivasan, Sravani Gottimukkala, Prasanta Kumar Ghosh |
INTERSPEECH | 4 |
| 2019 | Dirichlet Latent Variable Model: A Dynamic Model Based on Dirichlet Prior for Audio ProcessingabstractWe propose a dynamic latent variable model for learning latent bases from time varying, non-negative data. We take a probabilistic approach to modeling the temporal dependence in data by introducing a dynamic Dirichlet prior—a Dirichlet distribution with dynamic parameters. This new distribution allows us to assure non-negativity and avoid intractability when sequential updates are performed (otherwise encountered in using Dirichlet prior). We refer to the proposed model as the Dirichlet latent variable model (DLVM). We develop an expectation maximization algorithm for the proposed model, and also derive a maximuma posterioriestimate of the parameters. Furthermore, we connect the proposed DLVM to two popular latent basis learning methods—probabilistic latent component analysis (PLCA) and non-negative matrix factorization (NMF). We show that 1) PLCA is a special case of our DLVM, and 2) DLVM can be interpreted as a dynamic version of NMF. The usefulness of DLVM is demonstrated for three audio processing applications—speaker source separation, denoising, and bandwidth expansion. To this end, a new algorithm for source separation is also proposed. Through extensive experiments on benchmark databases, we show that the proposed model outperforms several relevant existing methods in all three applications. Anurendra Kumar, Tanaya Guha, Prasanta Kumar Ghosh |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Glottal Inverse Filtering Using Probabilistic Weighted Linear PredictionabstractGlottal inverse filtering is a noninvasive method for getting the glottal flow estimate from the speech. In this paper, we propose a method for glottal inverse filtering based on probabilistic weighted linear prediction (PWLP) in which the speech is assumed to be the output of an all-pole filter with glottal flow as an excitation. First, we introduce a probabilistic interpretation of the WLP, and we propose a probabilistic temporal weighting as convolution of a binary vector and a fixed window. We construct the posterior distribution based on the PWLP likelihood and a Gaussian prior on the filter coefficients. The parameters are estimated using the Gibbs sampling. The experiments are performed using the Lijencrants-Fant (LF) model based synthetic data, a physical model based synthetic data of different vowels and real speech data. Results demonstrate that the proposed method outperforms the best of the existing state-of-the-art methods in terms of the normalized amplitude quotient by 0.035 and 0.12 for the LF model and physical model based synthetic data, respectively. The results based on real speech data show that the glottal flow estimated by the proposed method in the closed phase is flatter and has less formant ripple compared to existing state-of-the-art methods. We also show two key features of the proposed method: first, the proposed method does not need prior detection of glottal closure or opening instants. The temporal weights are learnt in a data-driven manner, which is often found to be high near the closed phase of the glottal cycle, second, the Gaussian prior helps in estimating the filter coefficients when the closed phase duration is small. M. V. Achuth Rao, Prasanta Kumar Ghosh |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Concatenative Articulatory Video Synthesis Using Real-Time MRI Data for Spoken Language TrainingabstractSpoken language training benefits from showing a video of native speakers' articulatory movements to train the second language learners. Typically, the articulatory video is prepared in conjunction with the audio which is collected simultaneously with the articulatory recording. Articulatory video recording requires specialized equipment and, hence, is expensive and time consuming. In this work, we propose a concatenative synthesis approach to obtain articulatory videos for an audio, which may not have a simultaneous articulatory recording. In the training stage of the proposed approach, we make a repository for phoneme specific articulatory image sequence from the available articulatory video. During testing, image sequences are selected from this repository to ensure a smooth transition across phonetic events. The selected image sequences are finally stitched to synthesize the articulatory video for the test audio. Articulatory videos are synthesized for 50 words randomly selected from the MRI-TIMIT database, not seen in the training data. Subjective evaluation on the quality of the synthesized videos using twelve subjects suggests that the videos are close to the original ones with a rating of 3.78 out of 5, where a score of 5 (1) indicates that there is no (great) difference in quality between the original and the synthesized videos. Urvish Desai, Chiranjeevi Yarra, Prasanta Kumar Ghosh |
ICASSP | 3 |
| 2018 | Comparison of Speech Tasks for Automatic Classification of Patients with Amyotrophic Lateral Sclerosis and Healthy SubjectsabstractIn this work, we consider the task of acoustic and articulatory feature based automatic classification of Amyotrophic Lateral Sclerosis (ALS) patients and healthy subjects using speech tasks. In particular, we compare the roles of different types of speech tasks, namely rehearsed speech, spontaneous speech and repeated words for this purpose. Simultaneous articulatory and speech data were recorded from 8 healthy controls and 8 ALS patients using AG501 for the classification experiments. In addition to typical acoustic and articulatory features, new articulatory features are proposed for classification. As classifiers, both Deep Neural Networks (DNN) and Support Vector Machines (SVM) are examined. Classification experiments reveal that the proposed articulatory features outperform other acoustic and articulatory features using both DNN and SVM classifier. However, SVM performs better than DNN classifier using the proposed feature. Among three different speech tasks considered, the rehearsed speech was found to provide the highest F-score of 1, followed by an F-score of 0.92 when both repeated words and spontaneous speech are used for classification. Aravind Illa, Deep Patel, Yamini Belur, Meera SS, N. Shivashankar, Preethish-Kumar Veeramani, Seena Vengalil, Kiran Polavarapu, Saraswati Nashi, Atchayaram Nalini, Prasanta Kumar Ghosh |
ICASSP | 11 |
| 2018 | Speech Enhancement Using Multiple Deep Neural NetworksabstractIn this work, we present a variant of multiple deep neural network (DNN) based speech enhancement method. We directly estimate clean speech spectrum as a weighted average of outputs from multiple DNNs. The weights are provided by a gating network. The multiple DNNs and the gating network are trained jointly. The objective function is set as the mean square logarithmic error between the target clean spectrum and the estimated spectrum. We conduct experiments using two and four DNNs using the TIMIT corpus with nine noise types (four seen noises and five unseen noises) taken from the AURORA database at four different signal-to-noise ratios (SNRs). We also compare the proposed method with a single DNN based speech enhancement scheme and existing multiple DNN schemes using segmental SNR, perceptual evaluation of speech quality (PESQ) and short-term objective intelligibility (STOI) as the evaluation metrics. These comparisons show the superiority of proposed method over baseline schemes in both seen and unseen noises. Specifically, we observe an absolute improvement of 0.07 and 0.04 in PESQ measure compared to single DNN when averaged over all noises and SNRs for seen and unseen noise cases respectively. Pavan Karjol, Ajay Kumar M, Prasanta Kumar Ghosh |
ICASSP | 3 |
| 2018 | Binaural Speech Source Localization Using Template Matching of Interaural Time Difference PatternsabstractIn this paper we present a template based algorithm for localizing speech sources from a binaural recording. Binaural recordings are associated with head related transfer functions (HRTFs) for each direction which are specific to the object, say head, in between the two microphones. So, using these HRTFs and time-frequency representations of the binaural signals, we learn direction specific two dimensional reference templates using histograms of interaural time difference (ITD) in each frequency subband. These are called ITD pattern templates (IPTs). Test templates are then compared with each of the reference IPTs. The reference IPT, that matches best with the test template, provides the estimated direction of arrival for the test speech source. Experimental results obtained using subject_003 from the CIPIC database show that IPT based localization performs better than existing methods where the ITD distribution is modeled using Gaussian mixture model. Given n time-frequency points, we also present a method with complexity O( n) to compute the IPT, thus making it computationally efficient. Girija Ramesan Karthik, Prasanta Kumar Ghosh |
ICASSP | 2 |
| 2018 | A Supervised Air-Tissue Boundary Segmentation Technique in Real-Time Magnetic Resonance Imaging Video Using a Novel Measure of Contrast and Dynamic ProgrammingabstractThis paper introduces a technique for the supervised segmentation of Air-Tissue Boundaries (ATBs) in the upper airway of the vocal tract in the real time magnetic resonance imaging (rtMRI) videos. The proposed technique uses a novel measure of contrast across a boundary using Fisher discriminant function. ATBs in all frames of an rtMRI video are jointly estimated by maximizing the proposed measure of contrast around the predicted ATBs and incorporating a smoothness constraint to ensure the ATBs in consecutive frames do not change drastically. Dynamic programming is used for this purpose. The accuracy of the proposed technique is evaluated separately for the upper and lower ATBs using the Dynamic Time Warping distance between the predicted and the ground truth contours. Experiments with rtMRI videos from four subjects show that the error in ATB prediction using the proposed technique is 8.99% less than that using a semi-supervised grid based segmentation approach. A key feature of the proposed approach is that it can reliably predict the ATB outside the vocal tract unlike those with the existing methods. Advait Koparkar, Prasanta Kumar Ghosh |
ICASSP | 2 |
| 2018 | Intonation tutor by SPIRE (In-SPIRE): An Online Tool for an Automatic Feedback to the Second Language Learners in Learning Intonation
Anand P. A, Chiranjeevi Yarra, N. K. Kausthubha, Prasanta Kumar Ghosh |
INTERSPEECH | 4 |
| 2018 | Low Resource Acoustic-to-articulatory Inversion Using Bi-directional Long Short Term MemoryabstractEstimating articulatory movements from speech acoustic features is known as acoustic-to-articulatory inversion (AAI). Large amount of parallel data from speech and articulatory motion is required for training an AAI model in a subject dependent manner, referred to as subject dependent AAI (SD-AM). Electromagnetic articulograph (EMA) is a promising technology to record such parallel data, but it is expensive, time consuming and tiring for a subject. In order to reduce the demand for parallel acoustic-articulatory data in the AAI task for a subject, we, in this work, propose a subject-adaptative AAI method (SA-AAI) from an existing AAI model which is trained using large amount of parallel data from a fixed set of subjects. Experiments are performed with 30 subjects' acoustic-articulatory data and AM is trained using BLSTM network to examine the amount of data needed from a new target subject for the SAAAI to achieve an AAI performance equivalent to that of SDAAI. Experimental results reveal that the proposed SA-AAI performs similar to that of the SD-AAI with-.62.5% less training data. Among different articulators, the SA-AAI performance for tongue articulators matches with the corresponding SD-AAI performance with only,-,12.5% of the data used for SD-AAI training. Aravind Illa, Prasanta Kumar Ghosh |
INTERSPEECH | 2 |
| 2018 | Speech Enhancement Using Deep Mixture of Experts Based on Hard Expectation MaximizationabstractWe consider the problem of deep mixture of experts based speech enhancement. The deep mixture of experts, where experts are considered as deep neural network (DNN), is difficult to train due to the network structure. In this work, we propose a pre -training method for individual DNN in deep mixture of experts. We use hard expectation maximization (EM) to pre -train the individual DNNs. After pre -training, we take a weighted combination of outputs of individual DNN experts and jointly train the whole system. We compare the proposed method with single DNN based speech enhancement scheme. Speech enhancement experiments, in four SNR conditions, show the superiority of the proposed method over the baseline scheme. The average improvements obtained for four seen noise cases over single DNN scheme are 0.08, 0.59 dB and 0.015 in terms of objective measures viz perceptual evaluation of speech quality (PESQ), segmental signal to noise ratio (seg SNR) and short time objective intelligibility (STOI) respectively. Pavan Karjol, Prasanta Kumar Ghosh |
INTERSPEECH | 2 |
| 2018 | Subband Weighting for Binaural Speech Source Localization
Girija Ramesan Karthik, Parth Suresh, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2018 | Automatic Glottis Localization and Segmentation in Stroboscopic Videos Using Deep Neural NetworkabstractExact analysis of the glottal vibration patten is vital for assessing
voice pathologies. One of the primary steps in this
analysis is automatic glottis segmentation, which, in turn, has
two main parts, namely, glottis localization and the glottis segmentation.
In this paper, we propose a deep neural network
(DNN) based automatic glottis localization and segmentation
scheme. We pose the problem as a classification problem where
colors of each pixel and its neighborhood is classified as belonging
to inside or outside the glottis region. We further process
the classification result to get the biggest cluster, which is
declared as the segmented glottis. The proposed algorithm is
evaluated on a dataset comprising of stroboscopic videos from
18 subjects where the glottis region is marked by the three
Speech Language Pathologists (SLPs). On average, the proposed
DNN based segmentation scheme achieves a localization
performance of 65.33% and segmentation DICE score of 0.74
(absolute), which is better than the baseline scheme by 22.66%
and 0.09 respectively. We also find that the DICE score obtained
by the DNN based segmentation scheme correlates well
with the average DICE score computed between annotation provided
by any two SLPs suggesting the robustness of the proposed
glottis segmentation scheme. M. V. Achuth Rao, Rahul Krishnamurthy, Pebbili Gopikishore, Veeramani Priyadharshini, Prasanta Kumar Ghosh |
INTERSPEECH | 5 |
| 2018 | Whispered Speech to Neutral Speech Conversion Using Bidirectional LSTMsabstractWe propose a bidirectional long short-term memory (BLSTM) based whispered speech to neutral speech conversion system that employs the STRAIGHT speech synthesizer. We use a BLSTM to map the spectral features of whispered speech to those of neutral speech. Three other BLSTMs are employed to predict the pitch, periodicity levels and the voiced/unvoiced phoneme decisions from the spectral features of whispered speech. We use objective measures to quantify the quality of the predicted spectral features and excitation parameters, using data recorded from six subjects, in a four fold setup. We find that the temporal smoothness of the spectral features predicted using the proposed BLSTM based system is statistically more compared to that predicted using deep neural network based baseline schemes. We also observe that while the performance of the proposed system is comparable to the baseline scheme for pitch prediction, it is superior in terms of classifying voicing decisions and predicting periodicity levels. From subjective evaluation via listening test, we find that the proposed method is chosen as the best performing scheme 26.61% (absolute) more often than the best baseline scheme. This reveals that the proposed method yields a more natural sounding neutral speech from whispered speech. Nisha Meenakshi, Prasanta Kumar Ghosh |
INTERSPEECH | 2 |
| 2018 | Reconstructing Neutral Speech from Tracheoesophageal Speech
Abinay Reddy Naini, M. V. Achuth Rao, Nisha Meenakshi, Prasanta Kumar Ghosh |
INTERSPEECH | 4 |
| 2018 | Automatic Visual Augmentation for Concatenation Based Synthesized Articulatory Videos from Real-time MRI Data for Spoken Language Training
Chandana Srinivasan, Chiranjeevi Yarra, Ritu Aggarwal, Sanjeev Kumar Mittal, N. K. Kausthubha, Raseena K. T, Astha Singh, Prasanta Kumar Ghosh |
INTERSPEECH | 8 |
| 2018 | Relating Articulatory Motions in Different Speaking RatesabstractMovements of articulators (e.g., tongue, lips and jaw) in different speaking rates are related in a complex manner. In this work, we examine the underlying function to transform articulatory movements involved in producing speech at a neutral speaking rate into those at fast and slow speaking rates (N2F and N2S). For this we use articulatory movement data collected from five subjects using an Electromagnetic articulograph at neutral, fast and slow speaking rates. As candidate transformation functions (TF), we use affine transformations with a diagonal matrix and a full matrix and a nonlinear function modeled by a deep neural network (DNN). Since the duration of an utterance in different speaking rates would typically be unequal, it is required to time align the articulatory movement trajectories, which, in turn, affects the TF learnt. Therefore, we propose an iterative algorithm to alternately optimize for the TF and the time alignments. Subject specific experiments reveal that while N2F transformation can be well described by an affine transformation with a full matrix, N2S transformation is better represented by a more complex nonlinear function modeled by a DNN. This could be because subjects exhibit gross articulatory movements during fast speech and hyper-articulate while producing slow speech. Astha Singh, Nisha Meenakshi, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2018 | Air-Tissue Boundary Segmentation in Real-Time Magnetic Resonance Imaging Video Using Semantic Segmentation with Fully Convolutional NetworksabstractIn this paper, we propose a new technique for the segmentation of the Air-Tissue Boundaries (ATBs) in the vocal tract from the real-time magnetic resonance imaging (rtMRI) videos of the upper airway in the midsagittal plane. The proposed technique uses the approach of semantic segmentation using the Deep learning architecture called Fully Convolutional Networks (FCN). The architecture takes an input image and produces images of the same size with air and tissue class labels at each pixel. These output images are post processed using morphological filling and image smoothing to predict realistic ATBs. The performance of the predicted contours is evaluated using Dynamic Time Warping (DTW) distance between the manually annotated ground truth contours and the predicted contours. Four fold experiments with four subjects from USC-TIMIT corpus (with 2900 training images in every fold) demonstrate that the proposed FCN based approach has 8.87% and 9.65% lesser average error than the baseline Maeda Grid based scheme, for the lower and upper ATBs respectively. In addition, the proposed FCN based rtMRI segmentation achieves an average pixel classification accuracy of 99.05% across all subjects. C. A. Valliappan, Renuka Mannem, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2018 | SPIRE-SST: An Automatic Web-based Self-learning Tool for Syllable Stress Tutoring (SST) to the Second Language Learners
Chiranjeevi Yarra, Anand P. A, N. K. Kausthubha, Prasanta Kumar Ghosh |
INTERSPEECH | 4 |
| 2018 | Optimal sensor placement in electromagnetic articulography recording for speech production study
Ashok Kumar Pattem, Aravind Illa, Amber Afshan, Prasanta Kumar Ghosh |
Comput. Speech Lang. | 4 |
| 2018 | PSFM - A Probabilistic Source Filter Model for Noise Robust Glottal Closure Instant DetectionabstractAccurate estimation of glottal closure instant (GCI) enables several pitch synchronous speech analysis, such as prosody modifications, glottal inverse filtering, and study of pathological speech. We propose a probabilistic source-filter model (PSFM) for voiced speech, where the source is modeled using the Bernoulli Gaussian distribution, which models the GCI locations and the all-pole filter coefficients are modeled using Gaussian distribution. The probability of GCIs at each speech sample is estimated using the Gibbs sampling. We propose a cost to estimate the exact GCI locations using the N-best dynamic programming. A key feature of the proposed PSFM is that it allows us to include the second-order statistics of the noise for estimating the GCI locations, thereby resulting in a noise robust GCI detection technique, although it has high computational complexity. Evaluation on archivable priority list actual-word database (APLAWD) database shows the proposed algorithm performs at par with the state-of-the-art GCI detection method on clean speech. However, when evaluated in noisy conditions using five types of noises at six different signal-to-noise ratio (SNR) levels, we observe that the proposed method performs better than the best of the existing GCI detection scheme, particularly at low SNR condition indicating the noise robustness of the proposed method. M. V. Achuth Rao, Prasanta Kumar Ghosh |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | A comparative study of acoustic-to-articulatory inversion for neutral and whispered speechabstractWhispered speech is known to have different characteristics in acoustics and articulation compared to neutral speech. In this study, we compare the accuracy with which the articulation can be recovered from the acoustics of both types of speech, individually. Acoustic-to-articulatory inversion (AAI) is performed with twelve articulatory features using the deep neural network (DNN) with data obtained from four subjects. We consider AAI in matched and mis-matched train-test conditions, where the speech types in training and test are identical and different respectively. Experiments in matched condition reveal that the AAI performance for whispered speech drops significantly compared to that for neutral speech, only for jaw, tongue tip and tongue body, consistently, for all four subjects. This indicates that the whispered speech encodes information about the rest of the articulators to a degree similar to that of the neutral speech. Experiments in the mis-matched condition show a consistent drop in the AAI performance compared to the matched condition. This drop in performance from matched to mis-matched condition is found be the highest for upper lip which indicates that the upper lip movement could be encoded differently in whispered speech compared to that in neutral speech. Aravind Illa, Nisha Meenakshi, Prasanta Kumar Ghosh |
ICASSP | 3 |
| 2017 | Automatic detection of syllable stress using sonority based prominence features for pronunciation evaluationabstractAutomatic syllable stress detection is useful in assessing and diagnosing the quality of the pronunciation of second language (L2) learners in an automated way. Typically, the syllable stress depends on three prominence measures - intensity level, duration, pitch - around the sound unit with the highest sonority in the respective syllable. Stress detection is often formulated as a binary classification task using cues from the feature contours representing the prominence measures. We observe that cues from a feature contour obtained by incorporating relative sonority levels in the prominence measures are more indicative of the syllable stress compared to those from the feature contours representing only the prominence measures. Based on this observation, we propose a new feature contour based on temporal correlation selected sub-band correlation with an optimal set of sub-bands, called sonorous sub-bands, to maximize the stress detection accuracy. Experiments on ISLE corpus show that, for German and Italian non-native English speakers, the syllable stress detection accuracies (87.53% and 86.26%) are higher when the proposed features are used compared to the baseline accuracies (85.81% and 83.17%) indicating the effectiveness of the sonority based prominence features. Chiranjeevi Yarra, Om Deshmukh, Prasanta Kumar Ghosh |
ICASSP | 3 |
| 2017 | An Information Theoretic Analysis of the Temporal Synchrony Between Head Gestures and Prosodic Patterns in Spontaneous SpeechabstractWe analyze the temporal co-ordination between head gestures and prosodic patterns in spontaneous speech in a data-driven manner. For this study, we consider head motion and speech data from 24 subjects while they tell a fixed set of five stories. The head motion, captured using a motion capture system, is converted to Euler angles and translations in X, Y and Z-directions to represent head gestures. Pitch and short-time energy in voiced segments are used to represent the prosodic patterns. To capture the statistical relationship between head gestures and prosodic patterns, mutual information (MI) is computed at various delays between the two using data from 24 subjects in six native languages. The estimated MI, averaged across all subjects, is found to be maximum when the head gestures lag the prosodic patterns by 30msec. This is found to be true when subjects tell stories in English as well as in their native language. We observe a similar pattern in the root mean squared error of predicting head gestures from prosodic patterns using Gaussian mixture model. These results indicate that there could be an asynchrony between head gestures and prosody during spontaneous speech where head gestures follow the corresponding prosodic patterns. Gaurav Fotedar, Prasanta Kumar Ghosh |
INTERSPEECH | 2 |
| 2017 | Subband Selection for Binaural Speech Source LocalizationabstractWe consider the task of speech source localization using binaural cues, namely interaural time and level difference (ITD & ILD). A typical approach is to process binaural speech using gammatone filters and calculate frame-level ITD and ILD in each subband. The ITD, ILD and their combination (ITLD) in each subband are statistically modelled using Gaussian mixture models for every direction during training. Given a binaural test-speech, the source is localized using maximum likelihood criterion assuming that the binaural cues in each subband are independent. We, in this work, investigate the robustness of each subband for localization and compare their performance against the full-band scheme with 32 gammatone filters. We propose a subband selection procedure using the training data where subbands are rank ordered based on their localization performance. Experiments on Subject 003 from the CIPIC database reveal that, for high SNRs, the ITD and ITLD of just one subband centered at 296Hz is sufficient to yield localization accuracy identical to that of the full-band scheme with a test-speech of duration 1sec. At low SNRs, in case of ITD, the selected subbands are found to perform better than the full-band scheme. Girija Ramesan Karthik, Prasanta Kumar Ghosh |
INTERSPEECH | 2 |
| 2017 | A Robust Voiced/Unvoiced Phoneme Classification from Whispered Speech Using the 'Color' of Whispered Phonemes and Deep Neural Network
Nisha Meenakshi, Prasanta Kumar Ghosh |
INTERSPEECH | 2 |
| 2017 | PRAV: A Phonetically Rich Audio Visual Corpus
Abhishek Narwekar, Prasanta Kumar Ghosh |
INTERSPEECH | 2 |
| 2017 | A Dual Source-Filter Model of Snore Audio for Snorer Group Classification
M. V. Achuth Rao, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2017 | Phoneme State Posteriorgram Features for Speech Based Automatic Classification of Speakers in Cold and Healthy ConditionabstractWe consider the problem of automatically detecting if a speaker is suffering from common cold from his/her speech. When a speaker has symptoms of cold, his/her voice quality changes compared to the normal one. We hypothesize that such a change in voice quality could be reflected in lower likelihoods from a model built using normal speech. In order to capture this, we compute a 120-dimensional posteriorgram feature in each frame using Gaussian mixture model from 120 states of 40 three-states phonetic hidden Markov models trained on approximately 16.4 hours of normal English speech. Finally, a fixed 5160-dimensional phoneme state posteriorgram (PSP) feature vector for each utterance is obtained by computing statistics from the posteriorgram feature trajectory. Experiments on the 2017-Cold sub-challenge data show that when the decisions from bag-of-Audio-words (BoAW) and end-To-end (e2e) are combined with those from PSP features with unweighted majority rule, the UAR on the development set becomes 69 which is 2.9 (absolute) better than the best of the UARs obtained by the baseline schemes. When the decisions from ComParE, BoAW and PSP features are combined with simple majority rule, it results in a UAR of 68.52 on the test set. Akshay Kalkunte Suresh, Srinivasa Raghavan K. M., Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2017 | Spectrogram Enhancement Using Multiple Window Savitzky-Golay (MWSG) Filter for Robust Bird Sound DetectionabstractBird sound detection from real-field recordings is essential for identifying bird species in bioacoustic monitoring. Variations in the recording devices, environmental conditions, and the presence of vocalizations from other animals make the bird sound detection very challenging. In order to overcome these challenges, we propose an unsupervised algorithm comprising two main stages. In the first stage, a spectrogram enhancement technique is proposed using a multiple window Savitzky-Golay (MWSG) filter. We show that the spectrogram estimate using MWSG filter is unbiased and has lower variance compared with its single window counterpart. It is known that bird sounds are highly structured in the time-frequency (T-F) plane. We exploit these cues of prominence of T-F activity in specific directions from the enhanced spectrogram, in the second stage of the proposed method, for bird sound detection. In this regard, we use a set of four moving average filters that when applied to the enhanced spectrogram, yield directional spectrograms that capture the direction specific information. We propose a thresholding scheme on the time varying energy profile computed from each of these directional spectrograms to obtain frame-level binary decisions of bird sound activity. These individual decisions are then combined to obtain the final decision. Experiments are performed with three different datasets, with varying recording and noise conditions. Frame level F-score is used as the evaluation metric for bird sound detection. We find that the proposed method, on average, achieves higher F-score ($10.24\%$ relative) compared to the best of the six baseline schemes considered in this work. Nithin Rao Koluguri, Nisha Meenakshi, Prasanta Kumar Ghosh |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Better acoustic normalization in subject independent acoustic-to-articulatory inversion: Benefit to recognitionabstractIn subject independent acoustic-to-articulatory inversion (SII), the training and test subjects are in general different, whereas subject dependent inversion (SDI) uses the same training and test subjects. Thus, acoustic normalization is used to compensate for the mismatch between the training and the test subjects in SII. We show that a better acoustic normalization not only results in better articulatory estimates using SII, but also improves the broad class phonetic recognition accuracy, when the articulatory features estimated from SII are used for recognition. Recognition experiments using male and female subjects from the MOCHA-TIMIT corpus also show that there is no significant difference between the recognition accuracy using the articulatory features obtained by the best acoustic normalization in SII and that obtained using SDI as well as directly measured articulatory features. Amber Afshan, Prasanta Kumar Ghosh |
ICASSP | 2 |
| 2016 | A robust speech rate estimation based on the activation profile from the selected acoustic unit dictionaryabstractA typical solution for the speech rate estimation consists of two stages, which involves first computing a short-time feature contour such that most of peaks of the contour correspond to the syllable nuclei followed by the detection of the peaks of the contour corresponding to the syllable nuclei. Temporal correlation selected subband correlation (TCSSBC) is often used as a feature contour for the speech rate estimation in which correlation within and across a few selected sub-band energies are computed. In this work, instead of a fixed set of sub-bands, we learn them in a data-driven manner using a dictionary learning approach. Similarly, instead of the energy contours, we use the activation profile from the learned dictionary elements. We found that the peaks detected from the data-driven approach significantly improve the speech rate estimation when combined with the traditional TCSSBC approach using a proposed peak-merging strategy. Experiments are performed separately using Switchboard, TIMIT and CTIMIT corpora. Except Switchboard, the correlation coefficient for the speech rate estimation using the proposed approach is found to be higher than those by the TCSSBC technique - 3.1% and 5.2% (relative) improvements for TIMIT and CTIMIT respectively. Supriya Nagesh, Chiranjeevi Yarra, Om Deshmukh, Prasanta Kumar Ghosh |
ICASSP | 4 |
| 2016 | Automatic Recognition of Social Roles Using Long Term Role Transitions in Small Group InteractionsabstractRecognition of social roles in small group interactions is challenging because of the presence of disfluency in speech, frequent overlaps between speakers, short speaker turns and the need for reliable data annotation. In this work, we consider the problem of recognizing four roles, namely Gatekeeper, Protagonist, Neutral, and Supporter in small group interactions in AMI corpus. In general, Gatekeeper and Protagonist roles occur less frequently compared to Neutral, and Supporter. In this work, we exploit role transitions across segments in a meeting by incorporating role transition probabilities and formulating the role recognition as a decoding problem over the sequence of segments in an interaction. Experiments are performed in a five fold cross validation setup using acoustic, lexical and structural features with precision, recall and F-score as the performance metrics. The results reveal that precision averaged across all folds and different feature combinations improves in the case of Gatekeeper and Protagonist by 13.64% and 12.75% when the role transition information is used which in turn improves the F-score for Gatekeeper by 6.58% while the F-scores for the rest of the roles do not change significantly. Gaurav Fotedar, Aditya Gaonkar P., Saikat Chatterjee, Prasanta Kumar Ghosh |
INTERSPEECH | 4 |
| 2016 | A Class-Specific Speech Enhancement for Phoneme Recognition: A Dictionary Learning ApproachabstractWe study the influence of using class-specific dictionaries for enhancement over class-independent dictionary in phoneme recognition of noisy speech. We hypothesize that, using class specific dictionaries would remove the noise more compared to a class-independent dictionary, thereby resulting in better phoneme recognition. Experiments are performed with speech data from TIMIT corpus and noise samples from NOISEX-92 database. Using KSVD, four types of dictionaries have been learned: class-independent, manner-of-articulation-class, place-of-articulation-class and 39 phoneme-class. Initially, a set of labels are obtained by recognizing the speech, enhanced using a class-independent dictionary. Using these approximate labels, the corresponding class-specific dictionaries are used to enhance each frame of the original noisy speech, and this enhanced speech is then recognized. Compared to the results obtained using the class-independent dictionary, the 39 phoneme class based dictionaries provide a relative phoneme recognition accuracy improvement of 5.5%, 3.7%, 2.4% and 2.2%, respectively for factory2, m109, leopard and babble noises, when averaged over 0, 5 and 10 dB SNRs. Nazreen P. M., A. G. Ramakrishnan, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2016 | Speaker verification based on the fusion of speech acoustics and inverted articulatory signals
Ming Li 0026, Jangwon Kim, Adam C. Lammert, Prasanta Kumar Ghosh, Vikram Ramanarayanan, Shri Narayanan |
Comput. Speech Lang. | 4 |
| 2016 | Information theoretic optimal vocal tract region selection from real time magnetic resonance images for broad phonetic class recognition
Abhay Prasad, Prasanta Kumar Ghosh |
Comput. Speech Lang. | 2 |
| 2016 | A mode-shape classification technique for robust speech rate estimation and syllable nuclei detectionabstractAcoustic feature based speech (syllable) rate estimation and syllable nuclei detection are important problems in automatic speech recognition (ASR), computer assisted language learning (CALL) and fluency analysis. A typical solution for both the problems consists of two stages. The first stage involves computing a short-time feature contour such that most of the peaks of the contour correspond to the syllabic nuclei. In the second stage, the peaks corresponding to the syllable nuclei are detected. In this work, instead of the peak detection, we perform a mode-shape classification, which is formulated as a supervised binary classification problem – mode-shapes representing the syllabic nuclei as one class and remaining as the other. We use the temporal correlation and selected sub-band correlation (TCSSBC) feature contour and the mode-shapes in the TCSSBC feature contour are converted into a set of feature vectors using an interpolation technique. A support vector machine classifier is used for the classification. Experiments are performed separately using Switchboard, TIMIT and CTIMIT corpora in a five-fold cross validation setup. The average correlation coefficients for the syllable rate estimation turn out to be 0.6761, 0.6928 and 0.3604 for three corpora respectively, which outperform those obtained by the best of the existing peak detection techniques. Similarly, the average F -scores (syllable level) for the syllable nuclei detection are 0.8917, 0.8200 and 0.7637 for three corpora respectively. Chiranjeevi Yarra, Om Deshmukh, Prasanta Kumar Ghosh |
Speech Commun. | 3 |
| 2016 | Cumulative Impulse Strength for Epoch ExtractionabstractAlgorithms for extracting epochs or glottal closure instants (GCIs) from voiced speech typically fall into two categories: i) ones which operate on linear prediction residual (LPR) and ii) those which operate directly on the speech signal. While the former class of algorithms (such as YAGA and DPI) tend to be more accurate, the latter ones (such as ZFR and SEDREAMS) tend to be more noise-robust. In this letter, a temporal measure termed the cumulative impulse strength is proposed for locating the impulses in a quasi-periodic impulse-sequence embedded in noise. Subsequently, it is applied for detecting the GCIs from the inverted integrated LPR using a recursive algorithm. Experiments on two large corpora of speech with simultaneous electroglottographic recordings demonstrate that the proposed method is more robust to additive noise than the state-of-the-art algorithms, despite operating on the LPR. Prathosh A. P., P. Sujith, A. G. Ramakrishnan, Prasanta Kumar Ghosh |
IEEE Signal Process. Lett. | 4 |
| 2015 | Estimation of the invariant and variant characteristics in speech articulation and its application to speaker identificationabstractSpeech articulation varies across speakers for producing a speech sound due to the differences in their vocal tract morphologies, though the speech motor actions are executed in terms of relatively invariant gestures [1]. While the invariant articulatory gestures are driven by the linguistic content of the spoken utterance, the component of speech articulation that varies across speakers reflects speaker-specific and other paralinguistic information. In this work, we present a formulation to decompose the speech articulation from multiple speakers into the variant and invariant aspects when they speak the same sentence. The variant component is found to be a better representation for discriminating speakers compared to the speech articulation which includes the invariant part. Experiments with real-time magnetic resonance imaging (rtMRI) videos of speech production from multiple speakers reveal that the variant component of speech articulation yields a better frame-level speaker identification accuracy compared to the speech articulation as well as acoustic features by 29.9% and 9.4% (absolute) respectively. Abhay Prasad, Vijitha Periyasamy, Prasanta Kumar Ghosh |
ICASSP | 3 |
| 2015 | A discriminative analysis within and across voiced and unvoiced consonants in neutral and whispered speech in multiple indian languagesabstractWhispered speech lacks the vocal chord vibration which is typically used to distinguish voiced and unvoiced consonants, making their discrimination a challenging task. In this work, we objectively and subjectively quantify the amount of discrimination between a voiced (V) consonant and its unvoiced (UV) counterpart using seven V-UV consonant pairs in six Indian languages, in neutral and whispered speech. We also quantify the extent to which the voicing characteristics in a consonant changes from neutral to whispered speech. Experiments using vowelconsonant-vowel (VCV) stimuli demonstrate that the V-UV discrimination reduces from neutral to whispered speech in a consonant specific manner with highest reduction for /g/-/k/ pair and least reduction for /z/-/s/ pair. Interestingly, this reduction in objectively measured discrimination does not directly correlate with the reduction in the V-UV classification accuracy obtained from subjective evaluation. Results from listening test show that the maximum and minimum reduction in the V-UV classification accuracy occur for /d3/-/tf/ and /v/-/f/ pairs when whispered. Whispered Tamil and Telugu VCV achieve the highest (85.71%) and lowest (58.93%) subjective V-UV classification accuracy respectively, demonstrating the variability in the production and perception whispered consonants across languages. Nisha Meenakshi, Prasanta Kumar Ghosh |
INTERSPEECH | 2 |
| 2015 | Estimation of the air-tissue boundaries of the vocal tract in the mid-sagittal plane from electromagnetic articulograph dataabstractElectromagnetic articulograph (EMA) provides movement data of sensors attached to a few flesh points on different speech articulators including lips, jaw, and tongue while a subject speaks. In this work, we quantify the amount of information these flesh points provide about the vocal tract (VT) shape in the mid-sagittal plane. VT shape is described by the air-tissue boundaries, which are obtained manually from the recordings by real-time magnetic resonance imaging (rtMRI) of a set of utterances spoken by a subject, from whom the EMA recordings of the same set of utterances are also available. We propose a two-stage approach for reconstructing the VT shape from the EMA data. The first stage involves a co-registration of the EMA data with the VT shape from the rtMRI frames. The second stage involves the estimation of the air-tissue boundaries from the co-registered EMA points. Co-registration is done by a spatio-temporal alignment of the VT shapes from the rtMRI frames and EMA sensor data, while radial basis function (RBF) network is used for estimating the air tissue boundaries (ATBs). Experiments with the EMA and rtMRI recordings of five sentences spoken by one male and one female speakers show that the VT shape in the mid-sagittal plane can be recovered from the EMA flesh points with an average reconstruction error of 2.55 mm and 2.75 mm respectively. Satyabrata Parida, Ashok Kumar Pattem, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2015 | Automatic classification of eating conditions from speech using acoustic feature selection and a set of hierarchical support vector machine classifiersabstractThe problem of automatic classification of seven types of eating conditions from speech is considered. Based on the confusion among different eating conditions from a seven class support vector machine (SVM) classifier, a hierarchical SVM classifier is designed. Experiments on the iHEARu-EAT database show that the hierarchical classifier results in a better classification accuracy compared to a seven class classifier. We also perform a feature selection for each of the classifiers in the hierarchical approach. This further improves the unweighted average recall (UAR) to 73.7% compared to an UAR of 60.9% obtained from the baseline scheme of a direct seven-way classification. Abhay Prasad, Prasanta Kumar Ghosh |
INTERSPEECH | 2 |
| 2015 | An error correction scheme for GCI detection algorithms using pitch smoothness criterionabstractDetection of error-free glottal closure instants (GCI) is a critical requirement for many applications including text-to-speech synthesis, causal anti-causal decomposition and voice morphing. Many existing GCI detection algorithms commit errors under certain conditions. In this paper, we propose a post processing scheme for correcting errors of any GCI detection algorithm. The proposed error correction scheme works on the principle that the fundamental frequency over a voiced segment is slowly varying. The error correction is thus formulated as an optimization problem such that the pitch contour from the corrected GCIs has the least high frequency components. The proposed error correction scheme is experimentally evaluated on speech corpus with simultaneous EGG recordings using three state-of-the-art GCI detection algorithms viz., Dynamic Plosion Index (DPI), Zero Frequency Resonator (ZFR), and Speech Event Detection using the Residual Excitation And a Mean-based Signal (SEDREAMS). It is found that the proposed error correction scheme improves the performance of the GCI detection in clean speech as well as noisy conditions at different SNRs. P. Sujith, Prathosh A. P., A. G. Ramakrishnan, Prasanta Kumar Ghosh |
INTERSPEECH | 4 |
| 2015 | Improved subject-independent acoustic-to-articulatory inversion
Amber Afshan, Prasanta Kumar Ghosh |
Speech Commun. | 2 |
| 2015 | Robust Whisper Activity Detection Using Long-Term Log Energy Variation of Sub-Band SignalabstractThe goal in the whisper activity detection (WAD) is to find the whispered speech segments in a given noisy recording of whispered speech. Since whispering lacks the periodic glottal excitation, it resembles an unvoiced speech. This noise-like nature of the whispered speech makes WAD a more challenging task compared to a typical voice activity detection (VAD) problem. In this paper, we propose a feature based on the long term variation of the logarithm of the short-time sub-band signal energy for WAD. We also propose an automatic sub-band selection algorithm to maximally discriminate noisy whisper from noise. Experiments with eight noise types in four different signal-to-noise ratio (SNR) conditions show that, for most of the noises, the performance of the proposed WAD scheme is significantly better than that of the existing VAD schemes and whisper detection schemes when used for WAD. Nisha Meenakshi, Prasanta Kumar Ghosh |
IEEE Signal Process. Lett. | 2 |
| 2015 | Multiple Spectral Peak Tracking for Heart Rate Monitoring from Photoplethysmography Signal During Intensive Physical ExerciseabstractWe propose a multiple initialization based spectral peak tracking (MISPT) technique for heart rate monitoring from photoplethysmography (PPG) signal. MISPT is applied on the PPG signal after removing the motion artifact using an adaptive noise cancellation filter. MISPT yields several estimates of the heart rate trajectory from the spectrogram of the denoised PPG signal which are finally combined using a novel measure called trajectory strength. Multiple initializations help in correcting erroneous heart rate trajectories unlike the typical SPT which uses only single initialization. Experiments on the PPG data from 12 subjects recorded during intensive physical exercise show that the MISPT based heart rate monitoring indeed yields a better heart rate estimate compared to the SPT with single initialization. On the 12 datasets MISPT results in an average absolute error of 1.11 BPM which is lower than 1.28 BPM obtained by the state-of-the-art online heart rate monitoring algorithm. Navaneet K. Lakshminarasimha Murthy, Pavan C. Madhusudana, Pradyumna Suresha, Vijitha Periyasamy, Prasanta Kumar Ghosh |
IEEE Signal Process. Lett. | 5 |
| 2014 | Multi-pitch tracking using Gaussian mixture model with time varying parameters and Grating Compression TransformabstractGrating Compression Transform (GCT) is a two-dimensional analysis of speech signal which has been shown to be effective in multi-pitch tracking in speech mixtures. Multi-pitch tracking methods using GCT apply Kalman filter framework to obtain pitch tracks which requires training of the filter parameters using true pitch tracks. We propose an unsupervised method for obtaining multiple pitch tracks. In the proposed method, multiple pitch tracks are modeled using time-varying means of a Gaussian mixture model (GMM), referred to as TVGMM. The TVGMM parameters are estimated using multiple pitch values at each frame in a given utterance obtained from different patches of the spectrogram using GCT. We evaluate the performance of the proposed method on all voiced speech mixtures as well as random speech mixtures having well separated and close pitch tracks. TVGMM achieves multi-pitch tracking with 51% and 53% multi-pitch estimates having error ≤ 20% for random mixtures and all-voiced mixtures respectively. TVGMM also results in lower root mean squared error in pitch track estimation compared to that by Kalman filtering. M. N. Abhijith, Prasanta Kumar Ghosh, K. Rajgopal |
ICASSP | 2 |
| 2014 | A sparse smoothing approach for Gaussian Mixture Model based Acoustic-to-Articulatory InversionabstractIt is well-known that the performance of the Gaussian Mixture Model (GMM) based Acoustic-to-Articulatory Inversion (AAI) improves by either incorporating smoothness constraint directly in the inversion criterion or smoothing (low-pass filtering) estimated articulator trajectories in a post-processing step, where smoothing is performed independently of the inversion. As the low-pass filtering is independent of inversion, the smoothed articulator trajectory samples no longer remain optimal as per the inversion criterion. In this work, we propose a sparse smoothing technique which constrains the smoothed articulator trajectory to be different from the estimated trajectory only at a sparse subset of samples while simultaneously achieving the required degree of smoothness. Inversion experiments on the articulatory database show that the sparse smoothing achieves an AAI performance similar to that using low-pass filtering but in sparse smoothing ~15% (on average) of the samples in the smoothed articulator trajectory remain identical to those in the estimated articulator trajectory thereby preserve their AAI optimality as opposed to 0% in low-pass filtering. Prasad Sudhakar, Laurent Jacques, Prasanta Kumar Ghosh |
ICASSP | 3 |
| 2014 | Maximum a-posteriori estimation of missing samples with continuity constraint in Electromagnetic Articulography dataabstractElectromagnetic Articulography (EMA) technique is used to record the kinematics of different articulators while one speaks. EMA data often contains missing segments due to sensor failure. In this work, we propose a maximum a-posteriori (MAP) estimation with continuity constraint to recover the missing samples in the articulatory trajectories recorded using EMA. In this approach, we combine the benefits of statistical MAP estimation as well as the temporal continuity of the articulatory trajectories. Experiments on articulatory corpus using different missing segment durations show that the proposed continuity constraint results in a 30% reduction in average root mean squared error in estimation over statistical estimation of missing segments without any continuity constraint. P. Sujith, Prasanta Kumar Ghosh |
ICASSP | 2 |
| 2014 | Classification of clean and noisy bilingual movie audio for speech-to-speech translation corpora designabstractIdentifying suitable sources of bilingual audio and text data is a crucial part of statistical Speech to Speech (S2S) research and development. Movies, often dubbed in other languages, offer a good source for this purpose; but not all data are directly usable because of noise and other audio condition differences. Hence, automatically selecting the bilingual audio data that are suitable for analysis, and training S2S systems for specific environments becomes crucial. In this work, we extract bilingual speech segments from movies and aim at classifying segments as clean speech or speech with background noise (i.e. music, babble noise etc.). We examine various features in solving this problem and our best performing method delivers accuracy up to 87% in discriminating clean and noisy speech in bilingual data. Andreas Tsiartas, Prasanta Kumar Ghosh, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 2 |
| 2014 | Comparison of speech quality with and without sensors in electromagnetic articulograph AG 501 recordingabstractIn the recordings using electromagnetic articulograph AG 501, sensors are glued to subject’s articulators such as jaw, lips and tongue and both speech and articulatory movements are simultaneously recorded. In this work, we study the effect of the presence of the sensors on the quality of speech spoken by the subject. This is done by recording when a subject speaks a set of 19 VCV stimuli while sensors are attached to subject’s articulators. For comparison we also record the same set of stimuli spoken by the same subject but with no sensors attached to subject’s articulators. Both subjective and objective comparisons are made on the recorded stimuli in these two settings. Subjective evaluation is carried out using 16 evaluators. Listening experiments with recordings from five subjects show that the recordings with sensors attached are significantly different from those without sensors attached in terms of human recognition score as well as on a perceptual difference measure. This is also supported in the objective comparison which computes dissimilarity measure using the spectral shape information. Index Terms: Electromagnetic Articulography, speech quality, listening test Nisha Meenakshi, Chiranjeevi Yarra, Yamini Belur, Prasanta Kumar Ghosh |
INTERSPEECH | 4 |
| 2014 | Selection of optimal vocal tract regions using real-time magnetic resonance imaging for robust voice activity detectionabstractReal time magnetic resonance imaging (rtMRI) enables direct video capture of the moving vocal tract concurrent with audio signal providing valuable data for speech research. We consider a multimodal approach to voice activity detection (VAD) in the rtMRI recording that uses audio signal as well as MRI image sequence. The degraded quality of the audio recorded in the scanner motivates this multimodal scheme for robust VAD. Optimal regions in the MRI image are selected for performing VAD with a novel algorithm. VAD experiments using rtMRI data of two male and two female subjects show that VAD performance using optimally selected regions from MRI images is comparable to that using only audio signal. The optimal regions turn out to be parts of jaw, velum, glottis and lips. VAD performance using audio signal and MRI image sequence together is found to be significantly better (∼14% absolute improvement in VAD accuracy) than that using the audio only when the audio is contaminated with additive noise at low SNR. Index Terms: voice activity detection, vocal tract imaging, optimal region selection. Abhay Prasad, Prasanta Kumar Ghosh, Shri Narayanan |
INTERSPEECH | 2 |
| 2014 | Sparse smoothing of articulatory features from Gaussian mixture model based acoustic-to-articulatory inversion: benefit to speech recognition
Prasad Sudhakar, Prasanta Kumar Ghosh |
INTERSPEECH | 2 |
| 2014 | Missing samples estimation in electromagnetic articulography data using equality constrained kalman smootherabstractElectromagnetic articulography (EMA) data provides the movement of sensors attached to different articulators of a subject when the subject is speaking. EMA data often contains missing segments due to sensor failure. In this work, we propose an equality constrained Kalman smoother to estimate the missing samples in the EMA data. We incorporate the dynamics of the articulatory movement for missing samples estimation by considering the EMA data vector as the observations from a linear dynamical system. The proposed approach gives 41% reduction on the root mean square error of the estimates compared to the minimum mean square error estimator which does not utilize the dynamics of the articulatory movement. When compared to the maximum a-posteriori estimation with continuity constraints (MAPC) which incorporates smoothness of the articulatory trajectory during estimation, the proposed approach gives an average performance improvement of 4.8%. P. Sujith, Prasanta Kumar Ghosh |
INTERSPEECH | 2 |
| 2013 | Spatial and temporal alignment of multimodal human speech production data: Real time imaging, flesh point tracking and audioabstractIn speech production research, the integration of articulatory data derived from multiple measurement modalities can provide rich description of vocal tract dynamics by overcoming the limited spatio-temporal representations offered by individual modalities. This paper presents a spatial and temporal alignment method between two promising modalities using a corpus of TIMIT sentences obtained from the same speaker: flesh point tracking from Electromagnetic Articulography (EMA) that offers high temporal resolution but sparse spatial information and real time Magnetic Resonance Imaging (MRI) that offers good spatial details but at lower temporal rates. Spatial alignment is done by using palate tracking of EMA, but distortion in MRI audio and articulatory data variability make temporal alignment challenging. This paper proposes a novel alignment technique using joint acoustic-articulatory features which combines dynamic time warping and automatic feature extraction from MRI images. Experimental results show that the temporal alignment obtained using this technique is better (12% relative) than that using acoustic feature only. Jangwon Kim, Adam C. Lammert, Prasanta Kumar Ghosh, Shri Narayanan |
ICASSP | 3 |
| 2013 | Information theoretic acoustic feature selection for acoustic-to-articulatory inversionabstractWe use mutual information as the criterion to rank the Mel frequency cepstral coefficients (MFCCs) and their derivatives according to the information they provide about different articulatory features in acoustic-to-articulatory (AtoA) inversion. It is found that just a small subset of the coefficients encodes maximal information about articulatory features and interestingly, this subset is articulatory feature specific. We use these subsets of MFCCs(+derivatives) in AtoA inversion using Gaussian mixture model (GMM) mapping. Inversion experiments with articulatory data support the information theoretic finding that the subsets of MFCCs(+derivatives) as selected by feature ranking method are sufficient to achieve an inversion performance similar to that obtained by a conventional full set of MFCCs(+derivatives). This drastically reduces the modeling complexity of the acoustic-articulatory map using GMM without degrading inversion performance significantly. Index Terms: Acoustic-to-articulatory inversion, mutual information, Gaussian mixture model. Prasanta Kumar Ghosh, Shri Narayanan |
INTERSPEECH | 1 |
| 2013 | Speaker verification based on fusion of acoustic and articulatory informationabstractWe propose a practical, feature-level fusion approach for com-bining acoustic and articulatory information in speaker ver-ification task. We find that concatenating articulation fea-tures obtained from the measured speech production data with conventional Mel-frequency cepstral coefficients (MFCCs) im-proves the overall speaker verification performance. However, since access to the measured articulatory data is impractical for real world speaker verification applications, we also ex-periment with estimated articulatory features obtained using acoustic-to-articulatory inversion technique. Specifically, we show that augmenting MFCCs with articulatory features ob-tained from subject-independent acoustic-to-articulatory inver-sion technique also significantly enhances the speaker verifi-cation performance. This performance boost could be due to the information about inter-speaker variation present in the es-timated articulatory features, especially at the mean and vari-ance level. Experimental results on the Wisconsin X-Ray Mi-crobeam database show that the proposed acoustic-estimated-articulatory fusion approach significantly outperforms the tra-ditional acoustic-only baseline, providing up to 10 % relative re-duction in Equal Error Rate (EER). We further show that we can achieve an additional 5 % relative reduction in EER after score-level fusion. Index Terms: speech production, speaker verification, articula-tion features, acoustic-to-articulatory inversion, biometrics Ming Li 0026, Jangwon Kim, Prasanta Kumar Ghosh, Vikram Ramanarayanan, Shri Narayanan |
INTERSPEECH | 3 |
| 2013 | Multi-band long-term signal variability features for robust voice activity detectionabstractIn this paper, we propose robust features for the problem of voice activity detection (VAD). In particular, we extend the long term signal variability (LTSV) feature to accommodate multiple spectral bands. The motivation of the multi-band approach stems from the non-uniform frequency scale of speech phonemes and noise characteristics. Our analysis shows that the multi-band approach offers advantages over the single band LTSV for voice activity detection. In terms of classification accuracy, we show 0.3%-61.2% relative improvement over the best accuracy of the baselines considered for 7 out 8 different noisy channels. Experimental results, and error analysis, are reported on the DARPA RATS corpora of noisy speech. Index Terms: noisy speech data, voice activity detection, robust feature extraction Andreas Tsiartas, Theodora Chaspari, Athanasios Katsamanis, Prasanta Kumar Ghosh, Ming Li 0026, Maarten Van Segbroeck, Alexandros Potamianos, Shri Narayanan |
INTERSPEECH | 4 |
| 2013 | High-quality bilingual subtitle document alignments with application to spontaneous speech translation
Andreas Tsiartas, Prasanta Kumar Ghosh, Panayiotis G. Georgiou, Shri Narayanan |
Comput. Speech Lang. | 2 |
| 2011 | A subject-independent acoustic-to-articulatory inversionabstractAcoustic-to-articulatory inversion is usually done in a subject-dependent manner, i.e., the inversion procedure may not work well if the parallel acoustic and articulatory training data is not available from the subjects in the test set. In this paper, we propose a subject-independent acoustic-to-articulatory inversion procedure; the proposed scheme requires acoustic-articulatory training data only from one subject and uses a generic acoustic model to perform acoustic-to-articulatory inversion for any arbitrary test subject. Experimental results on the MOCHA database show that the subject-independent inversion procedure can achieve an inversion accuracy close to the accuracy of the subject-dependent procedure especially for the lip aperture, tongue tip and tongue body articulatory trajectories. We also investigate various articulatory features to analyze the effectiveness of the proposed inversion procedure. Prasanta Kumar Ghosh, Shri Narayanan |
ICASSP | 1 |
| 2011 | Bilingual audio-subtitle extraction using automatic segmentation of movie audioabstractExtraction of bilingual audio and text data is crucial for designing Speech to Speech (S2S) systems. In this work, we propose an automatic method to segment multilingual audio streams from movies. In addition, the audio streams are aligned with the corresponding subtitles. We found that the proposed method gives 89% perfectly segmented bilingual audio and 6% partially segmented bilingual audio. In addition, the mapping of the audio to the corresponding subtitles has accuracy 91%. Andreas Tsiartas, Prasanta Kumar Ghosh, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 2 |
| 2011 | Overlapped speech detection using long-term spectro-temporal similarity in stereo recordingabstractThe problem of detecting overlapped speech in stereo recordings using close-talk microphones is important for a variety of applications including the identification of back-channels, interruptions etc. in a dyadic or multi-party interactions. For detecting overlapped speech, we propose a feature derived using the spectral similarity of two channels over a range of acoustic frames. During overlapped speech frames the proposed spectro-temporal similarity-based feature values decrease and during non-overlapped speech frames the feature values increase due to the presence of cross-talk. Thus the proposed feature helps to discriminate the overlapped speech frames from the non-overlapped ones. Using overlapped speech detection experiments on a dyadic interaction corpus, it is shown that the proposed feature provides a significant improvement ~26% absolute, in the accuracy of detecting the overlapped speech frames when used as an additional feature to the baseline feature obtained from the two channels' intensity profiles. Bo Xiao 0003, Prasanta Kumar Ghosh, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 2 |
| 2011 | Analysis of Inter-Articulator Correlation in Acoustic-to-Articulatory Inversion Using Generalized Smoothness CriterionabstractThe movements of the different speech articulators are known to be correlated to various degrees during speech production. In this paper, we investigate whether the inter-articulator correlation is preserved among the articulators estimated through acoustic-toarticulatory inversion using the generalized smoothness criterion (GSC). GSC estimates each articulator separately without explicitly using any correlation information between the articulators. Theoretical analysis of inter-articulator correlation in GSC reveals that the correlation between any two estimated articulators may not be identical to that between the corresponding measured articulatory trajectories; however, based on smoothness constraints provided by the real articulatory data, we found that, in practice, the correlation among articulators is approximately preserved in GSC based inversion. To validate the theoretical analysis on interarticulator correlation, we propose a modified version of GSC where correlations among articulators are explicitly imposed. We found that there is no significant benefit in inversion using such modified GSC, which further strengthens the conclusions drawn from the theoretical analysis of inter-articulator correlation. Index Terms: acoustic-to-articulatory inversion, interarticulation correlation, generalized smoothness criterion Prasanta Kumar Ghosh, Shri Narayanan |
INTERSPEECH | 1 |
| 2011 | A Multimodal Real-Time MRI Articulatory Corpus for Speech ResearchabstractWe present MRI-TIMIT: a large-scale database of synchronized audio and real-time magnetic resonance imaging (rtMRI) data for speech research. The database currently consists of speech data acquired from two male and two female speakers of Amer-ican English. Subjects ’ upper airways were imaged in the mid-sagittal plane while reading the same 460 sentence corpus used in the MOCHA-TIMIT corpus [1]. Accompanying acoustic recordings were phonemically transcribed using forced align-ment. Vocal tract tissue boundaries were automatically identi-fied in each video frame, allowing for dynamic quantification of each speaker’s midsagittal articulation. The database and com-panion toolset provide a unique resource with which to examine articulatory-acoustic relationships in speech production. Index Terms: speech production, speech corpora, real-time MRI, multi-modal database, large-scale phonetic tools Shri Narayanan, Erik Bresch, Prasanta Kumar Ghosh, Louis Goldstein, Athanasios Katsamanis, Adam C. Lammert, Michael I. Proctor, Vikram Ramanarayanan, Yinghua Zhu |
INTERSPEECH | 3 |
| 2011 | Joint source-filter optimization for robust glottal source estimation in the presence of shimmer and jitter
Prasanta Kumar Ghosh, Shri Narayanan |
Speech Commun. | 1 |
| 2011 | Robust Voice Activity Detection Using Long-Term Signal VariabilityabstractWe propose a novel long-term signal variability (LTSV) measure, which describes the degree of nonstationarity of the signal. We analyze the LTSV measure both analytically and empirically for speech and various stationary and nonstationary noises. Based on the analysis, we find that the LTSV measure can be used to discriminate noise from noisy speech signal and, hence, can be used as a potential feature for voice activity detection (VAD). We describe an LTSV-based VAD scheme and evaluate its performance under eleven types of noises and five types of signal-to-noise ratio (SNR) conditions. Comparison with standard VAD schemes demonstrates that the accuracy of the LTSV-based VAD scheme averaged over all noises and all SNRs is ~6% (absolute) better than that obtained by the best among the considered VAD schemes, namely AMR-VAD2. We also find that, at -10 dB SNR, the accuracies of VAD obtained by the proposed LTSV-based scheme and the best considered VAD scheme are 88.49% and 79.30%, respectively. This improvement in the VAD accuracy indicates the robustness of the LTSV feature for VAD at low SNR condition for most of the noises considered. Prasanta Kumar Ghosh, Andreas Tsiartas, Shri Narayanan |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | Robust voice activity detection in stereo recording with crosstalkabstractCrosstalk in a stereo recording occurs when the speech from one participant is leaked into the close-talking microphones of the other participants. This crosstalk causes degradation of the voice activity detection (VAD) performance on individual channels, in spite of the strength of the crosstalk signal being lower than that of the participant’s speech. To address this problem, we first detect speech using a standard VAD scheme on the merged signal obtained by adding the signals from two channels and then determine the target channel using a channel selection scheme. Although VAD is performed on a short-term frame basis, we found that the channel selection performance improves with long-term signal information. Experiments using stereo recordings of real conversations demonstrate that the VAD accuracy averaged over both channels improves by 22% (absolute) indicating the robustness of the proposed approach to crosstalk compared to the single channel VAD scheme. Prasanta Kumar Ghosh, Andreas Tsiartas, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 1 |
| 2010 | Bark Frequency Transform Using an Arbitrary Order Allpass FilterabstractWe propose an arbitrary order stable allpass filter structure for frequency transformation from Hertz to Bark scale. According to the proposed filter structure, the first order allpass filter is causal, but the second and higher order allpass filters are non-causal. We find that the accuracy of the transformation significantly improves when a second or higher order allpass filter is designed compared to a first order allpass filter. We also find that the RMS error of the transformation monotonically decreases by increasing the order of the allpass filter. Prasanta Kumar Ghosh, Shri Narayanan |
IEEE Signal Process. Lett. | 1 |
| 2009 | Robust word boundary detection in spontaneous speech using acoustic and lexical cuesabstractWe consider the problem of word boundary detection in spontaneous speech utterances. Acoustic features have been well explored in the literature in the context of word boundary detection; however, in spontaneous speech of Switchboard-I corpus, we found that the accuracy of word boundary detection using acoustic features is poor (F-score ~ 0.63). We propose a new feature - that captures lexical cues in the context of the word boundary detection problem. We show that including proposed lexical feature along with the usual acoustic features, the accuracy of the word boundary detection improves considerably (F-score ~ 0.81). We also demonstrate the robustness of our proposed feature in presence of different noise levels for additive white and pink noise. Andreas Tsiartas, Prasanta Kumar Ghosh, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 2 |
| 2009 | Estimation of articulatory gesture patterns from speech acousticsabstractWe investigated dynamic programming (DP) and statemodel (SM) approaches for estimating gestural scores from speech acoustics. We performed a word-identification task using the gestural pattern vector sequences estimated by each approach. For a set of 75 randomly chosen words, we obtained the best word-identification accuracy (66.67%) using the DP approach. This result implies that considerable support for lexical access during speech perception might be provided by such a method of recovering gestural information from acoustics. Index Terms: gestural patterns, acoustic to gesture inversion Prasanta Kumar Ghosh, Shri Narayanan, Pierre L. Divenyi, Louis Goldstein, Elliot Saltzman |
INTERSPEECH | 1 |
| 2009 | Context-driven automatic bilingual movie subtitle alignmentabstractMovie subtitle alignment is a potentially useful approach for deriving automatically parallel bilingual/multilingual spoken language data for automatic speech translation. In this paper, we consider the movie subtitle alignment task. We propose a distance metric between utterances of different languages based on lexical features derived from bilingual dictionaries. We use the dynamic time warping algorithm to obtain the best alignment. The best F-score of ! 0.713 is obtained using the proposed approach. Index Terms :s ubtitles alignment, dynamic time warping Andreas Tsiartas, Prasanta Kumar Ghosh, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 2 |
| 2009 | Pitch Contour Stylization Using an Optimal Piecewise Polynomial ApproximationabstractWe propose a dynamic programming (DP) based piecewise polynomial approximation of discrete data such that theL2norm of the approximation error is minimized. We apply this technique for the stylization of speech pitch contour. Objective evaluation verifies that the DP based technique indeed yields minimum mean square error (MSE) compared to other approximation methods. Subjective evaluation reveals that the quality of the synthesized speech using stylized pitch contour obtained by the DP method is almost identical to that of the original speech. Prasanta Kumar Ghosh, Shri Narayanan |
IEEE Signal Process. Lett. | 1 |
| 2008 | Automatic classification of question turns in spontaneous speech using lexical and prosodic evidenceabstractThe ability to identify speech acts reliably is desirable in any spoken language system that interacts with humans. Minimally, such a system should be capable of distinguishing between question-bearing turns and other types of utterances. However, this is a non-trivial task, since spontaneous speech tends to have incomplete syntactic, and even ungrammatical, structure and is characterized by disfluencies, repairs and other non-linguistic vocalizations that make simple rule based pattern learning difficult. In this paper, we present a system for identifying question-bearing turns in spontaneous multi-party speech (ICSI Meeting Corpus) using lexical and prosodic evidence. On a balanced test set, our system achieves an accuracy of 71.9% for the binary question vs. non-question classification task. Further, we investigate the robustness of our proposed technique to uncertainty in the lexical feature stream (e.g. caused by speech recognition errors). Our experiments indicate that classification accuracy of the proposed method is robust to errors in the text stream, dropping only about 0.8% for every 10% increase in word error rate (WER). Sankaranarayanan Ananthakrishnan, Prasanta Kumar Ghosh, Shri Narayanan |
ICASSP | 2 |
| 2007 | Speech Segmentation using Extrema-Based Signal Track Length MeasureabstractWe introduce a novel temporal feature of a signal, namely extrema-based signal track length (ESTL) for the problem of speech segmentation. We show that ESTL measure is sensitive to both amplitude and frequency of the signal. The short-time ESTL (ST_ESTL) shows a promising way to capture the significant segments of speech signal, where the segments correspond to acoustic units of speech having distinct temporal waveforms. We compare ESTL based segmentation with ML and STM methods and find that it is as good as spectral feature based segmentation, but with lesser computational complexity. Prasanta Kumar Ghosh |
ICASSP (4) | 1 |
| 2007 | Pitch period estimation using multipulse model and wavelet transform
Prasanta Kumar Ghosh, Antonio Ortega, Shri Narayanan |
INTERSPEECH | 1 |
| 2006 | Dynamic Programming Based Optimum Non-Uniform Samples For Speech Reconstruction and CodingabstractNon-uniform sampling of a signal is formulated as an optimization problem which minimizes the reconstruction signal error. Dynamic programming (DP) has been used to solve this problem efficiently for a finite duration signal. Further, the optimum samples are quantized to realize a speech coder. The quantizer and the DP based optimum search for non-uniform samples (DP-NUS) can be combined in a closed-loop manner, which provides distinct advantage over the open-loop formulation. The DP-NUS formulation provides a useful control over the trade-off between bitrate and performance (reconstruction error). It is shown that 5-10 dB SNR improvement is possible using DP-NUS compared to extrema sampling approach. In addition, the close-loop DP-NUS gives a 4-5 dB improvement in reconstruction error Prasanta Kumar Ghosh, Thippur V. Sreenivas |
ICASSP (1) | 1 |
| 2006 | Time-varying filter interpretation of Fourier transform and its variants
Prasanta Kumar Ghosh, Thippur V. Sreenivas |
Signal Process. | 1 |