Aravind Illa

dblp:201/7322 · DBLP profile ↗
← Back
21ranked-venue papers
8as first author
6since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 4 since 2021
YearPublicationVenuePosition
2022 The impact of cross language on acoustic-to-articulatory inversion and its influence on articulatory speech synthesis
abstract
Estimating articulatory representations (ARs) from acoustic features is known as acoustic-to-articulatory inversion (AAI). Various factors of input acoustic features impact the performance of AAI. In this work, we investigate the effect of unseen language on the AAI performance in both seen and unseen speaker conditions. We further perform experiments to analyze how these AAI predictions in unseen language and unseen speaker conditions, in turn, impact the articulatory speech synthesis, i.e., articulatory-to-acoustic forward mapping (AAF). We hypothesize that this investigation enables the exploration of alternative approaches to voice conversion across unseen languages using ARs. Experiments are performed on the AAF model trained using English ARs and evaluated on ARs from unseen speakers speaking different native Indian languages, namely, Hindi, Kannada, Telugu, and Tamil. Experiments reveal that, for AAI, there is a drop in performance due to the mismatch in language in both seen and unseen speaker evaluations. For AAF, subjective evaluations reveal that the synthesized speech quality of non-native (mismatched language) speech is comparable with that of English (matched language).
Aravind Illa, Aanish Nair, Prasanta Kumar Ghosh
ICASSP1
2022 Learning Under Label Noise for Robust Spoken Language Understanding systems
Aravind Illa, Sriram Venkatapathy, Subhrangshu Nandi, Pritam Varma, Anurag Dwarakanath, Aram Galstyan
INTERSPEECH3
2022 Streaming model for Acoustic to Articulatory Inversion with transformer networks
abstract
Estimating speech articulatory movements from speech acoustics is known as Acoustic to Articulatory Inversion (AAI). Recently, transformer-based AAI models have been shown to achieve state-of-art performance. However, in transformer networks, the attention is applied over the whole utterance, thereby needing to obtain the full utterance before the inference, which leads to high latency and is impractical for streaming AAI. To enable streaming during inference, evaluation could be performed on non-overlapping chucks instead of a full utterance. However, due to a mismatch of the attention receptive field during training and evaluation, there could be a drop in AAI performance. To overcome this scenario, in this work we perform experiments with different attention masks and use context from previous predictions during training. Experiments results revealed that using the random start mask attention with the context from previous predictions of transformer decoder performs better than the baseline results.
Sathvik Udupa, Aravind Illa, Prasanta Kumar Ghosh
INTERSPEECH2
2021 Acoustic-to-Articulatory Inversion for Dysarthric Speech by Using Cross-Corpus Acoustic-Articulatory Data
abstract
In this work, we focus on estimating articulatory movements from acoustic features, known as acoustic-to-articulatory inversion (AAI), for dysarthric patients with amyotrophic lateral sclerosis (ALS). Unlike healthy subjects, there are two potential challenges involved in AAI on dysarthric speech. Due to speech impairment, the pronunciation of dysarthric patients is unclear and inaccurate, which could impact the AAI performance. In addition, acoustic-articulatory data from dysarthric patients is limited due to the difficulty in the recording. These challenges motivate us to utilize cross-corpus acoustic-articulatory data. In this study, we propose an AAI model by conditioning speaker information using x-vectors at the input, and multi-target articulatory trajectory outputs for each corpus separately. Results reveal that the proposed AAI model shows relative improvements of the Pearson correlation coefficient (CC) by ~13.16% and ~16.45% over a randomly initialized baseline AAI model trained with only dysarthric corpus in seen and unseen conditions, respectively. In the seen conditions, the proposed AAI model outperforms the three baseline AAI models, that utilize the cross-corpus, by ~3.49%, ~6.46%, and ~4.03% in terms of CC.
Sarthak Kumar Maharana, Aravind Illa, Renuka Mannem, Yamini Belur, Preetie Shetty, Preethish-Kumar Veeramani, Seena Vengalil, Kiran Polavarapu, Atchayaram Nalini, Prasanta Kumar Ghosh
ICASSP2
2021 Estimating Articulatory Movements in Speech Production with Transformer Networks
abstract
We estimate articulatory movements in speech production from different modalities - acoustics and phonemes. Acoustic-to articulatory inversion (AAI) is a sequence-to-sequence task. On the other hand, phoneme to articulatory (PTA) motion estimation faces a key challenge in reliably aligning the text and the articulatory movements. To address this challenge, we explore the use of a transformer architecture - FastSpeech, with explicit duration modelling to learn hard alignments between the phonemes and articulatory movements. We also train a transformer model on AAI. We use correlation coefficient (CC) and root mean squared error (rMSE) to assess the estimation performance in comparison to existing methods on both tasks. We observe 154%, 11.8% & 4.8% relative improvement in CC with subject-dependent, pooled and fine-tuning strategies, respectively, for PTA estimation. Additionally, on the AAI task, we obtain 1.5%, 3% and 3.1% relative gain in CC on the same setups compared to the state-of-the-art baseline. We further present the computational benefits of having transformer architecture as representation blocks.
Sathvik Udupa, Anwesha Roy, Abhayjeet Singh, Aravind Illa, Prasanta Kumar Ghosh
Interspeech4
2021 Web Interface for Estimating Articulatory Movements in Speech Production from Acoustics and Text
Sathvik Udupa, Anwesha Roy, Abhayjeet Singh, Aravind Illa, Prasanta Kumar Ghosh
Interspeech4
2020 Voice based classification of patients with Amyotrophic Lateral Sclerosis, Parkinson's Disease and Healthy Controls with CNN-LSTM using transfer learning
abstract
In this paper, we consider 2-class and 3-class classification problems for classifying patients with Amyotrophic Lateral Sclerosis (ALS), Parkinson's Disease (PD), and Healthy Controls (HC) using a CNNLSTM network. Classification performance is examined for three different tasks, namely, Spontaneous speech (SPON), Diadochokinetic rate (DIDK) and Sustained phoneme production (PHON). Experiments are conducted using speech data recorded from 60 ALS, 60 PD, and 60 HC subjects. Classifications using SVM and DNN are considered as baseline schemes. Classification accuracy of ALS and HC (indicated by ALS/HC) using CNN-LSTM has shown an improvement of 10.40%, 4.22% and 0.08% for PHON, SPON and DIDK tasks, respectively over the best of the baseline schemes. Furthermore, the CNN-LSTM network achieves the highest PD/HC classification accuracy of 88.5% for the SPON task and the highest 3-class (ALS/PD/HC) classification accuracy of 85.24% for the DIDK task. Experiments using transfer learning at low resource training data show that data from ALS benefits PD/HC classification and vice-versa. Experiments with fine-tuning weights of 3-class (ALS/PD/HC) classifier for 2-class classification (PD/HC or ALS/HC) gives an absolute improvement of 2% classification accuracy in SPON task when compared with randomly initialized 2-class classifier.
Jhansi Mallela, Aravind Illa, Suhas B. N., Sathvik Udupa, Yamini Belur, Atchayaram Nalini, Pradeep Reddy, Dipanjan Gope, Prasanta Kumar Ghosh
ICASSP2
2020 A Comparative Study of Estimating Articulatory Movements from Phoneme Sequences and Acoustic Features
abstract
Unlike phoneme sequences, movements of speech articulators (lips, tongue, jaw, velum) and the resultant acoustic signal are known to encode not only the linguistic message but also carry para-linguistic information. While several works exist for estimating articulatory movement from acoustic signals, little is known to what extent articulatory movements can be predicted only from linguistic information, i.e., phoneme sequence. In this work, we estimate articulatory movements from three different input representations: R1) acoustic signal, R2) phoneme sequence, R3) phoneme sequence with timing information. While an attention network is used for estimating articulatory movement in the case of R2, BLSTM network is used for R1 and R3. Experiments with ten subjects’ acoustic-articulatory data reveal that the estimation techniques achieve an average correlation coefficient of 0.85, 0.81, and 0.81 in the case of R1, R2, and R3 respectively. This indicates that attention network, although uses only phoneme sequence (R2) without any timing information, results in an estimation performance similar to that using rich acoustic signal (R1), suggesting that articulatory motion is primarily driven by the linguistic message. The correlation coefficient is further improved to 0.88 when R1 and R3 are used together for estimating articulatory movements.
Abhayjeet Singh, Aravind Illa, Prasanta Kumar Ghosh
ICASSP2
2020 Speaker Conditioned Acoustic-to-Articulatory Inversion Using x-Vectors
abstract
Speech production involves the movement of various articulators, including tongue, jaw, and lips. Estimating the movement of the articulators from the acoustics of speech is known as acoustic-to-articulatory inversion (AAI). Recently, it has been shown that instead of training AAI in a speaker specific manner, pooling the acoustic-articulatory data from multiple speakers is beneficial. Further, additional conditioning with speaker specific information by one-hot encoding at the input of AAI along with acoustic features benefits the AAI performance in a closed-set speaker train and test condition. In this work, we carry out an experimental study on the benefit of using x-vectors for providing speaker specific information to condition AAI. Experiments with 30 speakers have shown that the AAI performance benefits from the use of x-vectors in a closed set seen speaker condition. Further, x-vectors also generalizes well for unseen speaker evaluation.
Aravind Illa, Prasanta Kumar Ghosh
INTERSPEECH1
2020 Raw Speech Waveform Based Classification of Patients with ALS, Parkinson's Disease and Healthy Controls Using CNN-BLSTM
Jhansi Mallela, Aravind Illa, Yamini Belur, Atchayaram Nalini, Pradeep Reddy, Dipanjan Gope, Prasanta Kumar Ghosh
INTERSPEECH2
2020 Speech Rate Task-Specific Representation Learning from Acoustic-Articulatory Data
abstract
In this work, speech rate is estimated using the task-specific representations which are learned from the acoustic-articulatory data, in contrast to generic representations which may not be optimal for the speech rate estimation. 1-D convolutional filters are used to learn speech rate specific acoustic representations from the raw speech. A convolutional dense neural network (CDNN) is used to estimate the speech rate from the learned representations. In practice, articulatory data is not directly available; thus, we use Acoustic-to-Articulatory Inversion (AAI) to derive the articulatory representations from acoustics. However, these pseudo-articulatory representations are also generic and not optimized for any task. To learn the speech-rate specific pseudo-articulatory representations, we propose a joint training of BLSTM-based AAI and CDNN using a weighted loss function that considers the losses corresponding to speech rate estimation and articulatory prediction. The proposed model yields an improvement in speech rate estimation by ~18.5 in terms of pearson correlation coefficient (CC) compared to the baseline CDNN model with generic articulatory representations as inputs. To utilize complementary information from articulatory features, we further perform experiments by concatenating task-specific acoustic and pseudo-articulatory representations, which yield an improvement in CC by ~2.5 compared to the baseline CDNN model. © 2020 ISCA
Renuka Mannem, Hima Jyothi, Aravind Illa, Prasanta Kumar Ghosh
INTERSPEECH3
2020 Attention and Encoder-Decoder Based Models for Transforming Articulatory Movements at Different Speaking Rates
abstract
While speaking at different rates, articulators (like tongue, lips) tend to move differently and the enunciations are also of different durations. In the past, affine transformation and DNN have been used to transform articulatory movements from neutral to fast(N2F) and neutral to slow(N2S) speaking rates [1]. In this work, we improve over the existing transformation techniques by modeling rate specific durations and their transformation using AstNet, an encoder-decoder framework with attention. In the current work, we propose an encoder-decoder architecture using LSTMs which generates smoother predicted articulatory trajectories. For modeling duration variations across speaking rates, we deploy attention network, which eliminates the needto align trajectories in different rates using DTW. We performa phoneme specific duration analysis to examine how well duration is transformed using the proposed AstNet. As the range of articulatory motions is correlated with speaking rate, we also analyze amplitude of the transformed articulatory movements at different rates compared to their original counterparts, to examine how well the proposed AstNet predicts the extent of articulatory movements in N2F and N2S. We observe that AstNet could model both duration and extent of articulatory movements better than the existing transformation techniques resulting in more accurate transformed articulatory trajectories.
Abhayjeet Singh, Aravind Illa, Prasanta Kumar Ghosh
INTERSPEECH2
2020 The impact of speaking rate on acoustic-to-articulatory inversion
Aravind Illa, Prasanta Kumar Ghosh
Comput. Speech Lang.1
2019 Representation Learning Using Convolution Neural Network for Acoustic-to-articulatory Inversion
abstract
Recent techniques employ end-to-end systems to learn relevant features for several speech related applications, including speech recognition, and speaker verification. In this work, we focus on the task of acoustic-to-articulatory inversion (AAI) for which we propose an end-to-end system that comprises a convolution neural network (CNN) and a bidirectional long short-term memory network (BLSTM). The aim of this work is to understand the nature of the features learnt by the end-to-end model and the importance of pre-emphasis in representation learning for AAI. Further, we propose a subject adaptation scheme to overcome the limitations of the availability of parallel acoustic-articulatory data to train an end-to-end AAI system. The AAI performance is evaluated with ~3.19 hours of acoustic-articulatory data collected from 8 subjects. Experiments reveal that, the frequency response of filters learnt by the CNN in the proposed system resembles those of the mel-scale, and hence, the performance of the proposed system (RMSE=1.47mm) is on par with that using mel-frequency cepstral coefficients (1.42mm) as features. Using pre-emphasis reduces RMSE by 0.13mm, and also the proposed adaptation scheme performs better than a subject-specific AAI model by an RMSE of 0.21mm despite of limited acoustic-articulatory data from a subject.
Aravind Illa, Prasanta Kumar Ghosh
ICASSP1
2019 A Study on Robustness of Articulatory Features for Automatic Speech Recognition of Neutral and Whispered Speech
abstract
Traditionally, automatic speech recognition (ASR) systems are trained on acoustic representations of neutral speech. As a result, their performance degrades when tested with whispered speech. In this work, we explore the robustness of articulatory features in ASR of neutral and whispered speech. We use acoustic, articulatory, and integrated acoustic and articulatory feature vectors in matched and mismatched train-test cases. The results suggest that the articulatory data is useful in ASR of both neutral and whispered speech, especially in the mismatched train-test cases. When we concatenate acoustic and articulatory feature vectors and deploy it to the mismatched train-test case where the model is trained with neutral speech and tested with whispered speech, a relative improvement in phone error rate of 27.2% is observed compared to when only acoustic features are used. This suggests that articulatory data contains information complementary to acoustic representations. A phone specific recognition error is also presented which illustrates phones where adding articulatory information gives maximum benefit.
Gokul Srinivasan, Aravind Illa, Prasanta Kumar Ghosh
ICASSP2
2019 An Investigation on Speaker Specific Articulatory Synthesis with Speaker Independent Articulatory Inversion
abstract
Estimating speech representations from articulatory movements is known as articulatory-to-acoustic forward (AAF) mapping. Typically this mapping is learned using directly measured articulatory movement in a subject-specific manner. Such AAF mapping has been shown to benefit the speech synthesis applications. In this work, we investigate the speaker similarity and naturalness of utterances generated by AAF which is driven by the articulatory movements from a subject (referred to as cross speaker) different from the speaker (target speaker) used for training AAF mapping. Experiments are performed with directly measured articulatory data from 9 speakers (8 target speakers and 1 cross speaker), which are recorded using Electromagnetic articulograph AG501. Experiments are also performed with articulatory features estimated using speaker independent acoustic-to-articulatory inversion (SI-AAI) model trained on 26 reference speakers. Objective evaluation on target speakers reveal that the articulatory features estimated from SI-AAI result in a lower Mel-cepstrum distortion compared to that using directly measured articulatory features. Further, listening tests reveal that the directly measured articulatory movements preserve the speaker similarity better than estimated ones. Although, for naturalness, articulatory movements predicted by SI-AAI perform better than the direct measurements.
Aravind Illa, Prasanta Kumar Ghosh
INTERSPEECH1
2019 Acoustic and Articulatory Feature Based Speech Rate Estimation Using a Convolutional Dense Neural Network
abstract
In this paper, we propose a speech rate estimation approach using a convolutional dense neural network (CDNN). The CDNN based approach uses the acoustic and articulatory features for speech rate estimation. The Mel Frequency Cepstral Coefficients (MFCCs) are used as acoustic features and the articulograms representing time-varying vocal tract profile are used as articulatory features. The articulogram is computed from a real-time magnetic resonance imaging (rtMRI) video in the midsagittal plane of a subject while speaking. However, in practice, the articulogram features are not directly available, unlike acoustic features from speech recording. Thus, we use an Acoustic-to-Articulatory Inversion method using a bidirectional long-short-term memory network which estimates the articulogram features from the acoustics. The proposed CDNN based approach using estimated articulatory features requires both acoustic and articulatory features during training but it requires only acoustic data during testing. Experiments are conducted using rtMRI videos from four subjects each speaking 460 sentences. The Pearson correlation coefficient is used to evaluate the speech rate estimation. It is found that the CDNN based approach gives a better correlation coefficient than the temporal and selected sub-band correlation (TCSSBC) based baseline scheme by 81.58 and 73.68 (relative) in seen and unseen subject conditions respectively.
Renuka Mannem, Jhansi Mallela, Aravind Illa, Prasanta Kumar Ghosh
INTERSPEECH3
2018 Comparison of Speech Tasks for Automatic Classification of Patients with Amyotrophic Lateral Sclerosis and Healthy Subjects
abstract
In this work, we consider the task of acoustic and articulatory feature based automatic classification of Amyotrophic Lateral Sclerosis (ALS) patients and healthy subjects using speech tasks. In particular, we compare the roles of different types of speech tasks, namely rehearsed speech, spontaneous speech and repeated words for this purpose. Simultaneous articulatory and speech data were recorded from 8 healthy controls and 8 ALS patients using AG501 for the classification experiments. In addition to typical acoustic and articulatory features, new articulatory features are proposed for classification. As classifiers, both Deep Neural Networks (DNN) and Support Vector Machines (SVM) are examined. Classification experiments reveal that the proposed articulatory features outperform other acoustic and articulatory features using both DNN and SVM classifier. However, SVM performs better than DNN classifier using the proposed feature. Among three different speech tasks considered, the rehearsed speech was found to provide the highest F-score of 1, followed by an F-score of 0.92 when both repeated words and spontaneous speech are used for classification.
Aravind Illa, Deep Patel, Yamini Belur, Meera SS, N. Shivashankar, Preethish-Kumar Veeramani, Seena Vengalil, Kiran Polavarapu, Saraswati Nashi, Atchayaram Nalini, Prasanta Kumar Ghosh
ICASSP1
2018 Low Resource Acoustic-to-articulatory Inversion Using Bi-directional Long Short Term Memory
abstract
Estimating articulatory movements from speech acoustic features is known as acoustic-to-articulatory inversion (AAI). Large amount of parallel data from speech and articulatory motion is required for training an AAI model in a subject dependent manner, referred to as subject dependent AAI (SD-AM). Electromagnetic articulograph (EMA) is a promising technology to record such parallel data, but it is expensive, time consuming and tiring for a subject. In order to reduce the demand for parallel acoustic-articulatory data in the AAI task for a subject, we, in this work, propose a subject-adaptative AAI method (SA-AAI) from an existing AAI model which is trained using large amount of parallel data from a fixed set of subjects. Experiments are performed with 30 subjects' acoustic-articulatory data and AM is trained using BLSTM network to examine the amount of data needed from a new target subject for the SAAAI to achieve an AAI performance equivalent to that of SDAAI. Experimental results reveal that the proposed SA-AAI performs similar to that of the SD-AAI with-.62.5% less training data. Among different articulators, the SA-AAI performance for tongue articulators matches with the corresponding SD-AAI performance with only,-,12.5% of the data used for SD-AAI training.
Aravind Illa, Prasanta Kumar Ghosh
INTERSPEECH1
2018 Optimal sensor placement in electromagnetic articulography recording for speech production study
Ashok Kumar Pattem, Aravind Illa, Amber Afshan, Prasanta Kumar Ghosh
Comput. Speech Lang.2
2017 A comparative study of acoustic-to-articulatory inversion for neutral and whispered speech
abstract
Whispered speech is known to have different characteristics in acoustics and articulation compared to neutral speech. In this study, we compare the accuracy with which the articulation can be recovered from the acoustics of both types of speech, individually. Acoustic-to-articulatory inversion (AAI) is performed with twelve articulatory features using the deep neural network (DNN) with data obtained from four subjects. We consider AAI in matched and mis-matched train-test conditions, where the speech types in training and test are identical and different respectively. Experiments in matched condition reveal that the AAI performance for whispered speech drops significantly compared to that for neutral speech, only for jaw, tongue tip and tongue body, consistently, for all four subjects. This indicates that the whispered speech encodes information about the rest of the articulators to a degree similar to that of the neutral speech. Experiments in the mis-matched condition show a consistent drop in the AAI performance compared to the matched condition. This drop in performance from matched to mis-matched condition is found be the highest for upper lip which indicates that the upper lip movement could be encoded differently in whispered speech compared to that in neutral speech.
Aravind Illa, Nisha Meenakshi, Prasanta Kumar Ghosh
ICASSP1