VLDB 2026 Research / reviewers in the wild / expert
S. R. Mahadeva Prasanna
dblp:32/5647 · also S. R. M. Prasanna
· DBLP profile ↗
116ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0002-8135-7938ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 84 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 83 · 5 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 2Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Frugal Voice: Offline Community-Governed Speech Models for the Soliga Tribe in KarnatakaabstractLow-resource and tribal languages in India are at acute risk of digital extinction because contemporary language technologies overwhelmingly target high-resource languages and centralized, cloud-hosted large models that are expensive to train, deploy, and adapt. Centralized automatic speech recognition (ASR) and text-to-speech (TTS) pipelines require large amounts of labelled data, stable connectivity, and significant compute resources, all of which are misaligned with the infrastructural realities and socio-economic conditions of many tribal communities in India. These constraints are particularly visible for the Soliga community, a Dravidian-language-speaking tribe in the Biligiri Rangaswamy Hills of Karnataka, whose language has no widely used script and for which digital resources remain scarce. We present a socio-technical framework and prototype for building low-cost voice models with and for the Soliga tribe using offline, edge-based federated learning. The framework combines community-governed data collection and consent processes, on-device training of compact ASR and keyword-spotting models, and an intermittent-connectivity aggregation protocol that runs on low-cost edge hardware. Our experiments show that, even with only 5 hours of Soliga speech, a centralized Wav2Vec 2.0 baseline achieves a word error rate (WER) of 37.95% and character error rate (CER) of 11.11%, and that a federated frugal architecture trained on the same small corpus can approach this performance (WER 44.20%, CER 15.80%) while keeping data local and under community control. Sarbani Banerjee Belur, Nandha Sathiaseelan, Prashant Bannulmath, Sunil Saumya, Shruti Maralappanavar, Deepak K. T, S. R. Mahadeva Prasanna, Arjuna Sathiaseelan |
COMPASS | 7 |
| 2025 | Cross-lingual Evaluation Of Hypernasality Using Wav2Vec2 FeaturesabstractHypernasality, a speech resonance disorder characterized by excessive nasal airflow, presents challenges in accurate detection across languages. Traditional assessments of hypernasality include perceptual evaluation and nasometry. More recently, objective acoustic measures based on formant analysis and other acoustic features have been proposed as proxies for hypernasality; however, these acoustic measures exhibit considerable variability, especially in multilingual contexts. In this study, we utilize the wav2vec2-large-xlsr-53 model, a cross-lingual speech representation framework, and evaluate hypernasality-related features within its transformer layers. We extracted features from each layer of the model’s transformer and trained a machine learning model to predict hypernasality ratings across three datasets: Americleft (English), New Mexico Cleft Palate Center (English), and All India Institute of Speech and Hearing (Kannada). Our analysis reveals that the 11th and 12th layer contextualized embeddings effectively model hypernasality cross-lingually, demonstrating significant within-language average correlations (0.78 and 0.78) and cross-lingual average correlations (0.68 and 0.67) between predicted and perceptual ratings. These findings suggest that the wav2vec2-large-xlsr-53 model’s intermediate layers effectively capture hypernasality cross-lingually. Krupaben Kothadia, Vikram C. M., Ajish K. Abraham, Pushpavathi M, S. R. Mahadeva Prasanna, Nancy Scherer, Kathy Chapman, Julie M. Liss, Visar Berisha |
ICASSP | 5 |
| 2025 | Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion RecognitionabstractIn this study, we investigate multimodal foundation models (MFMs) for emotion recognition from non-verbal sounds. We hypothesize that MFMs, with their joint pre-training across multiple modalities, will be more effective in non-verbal sounds emotion recognition (NVER) by better interpreting and differentiating subtle emotional cues that may be ambiguous in audio-only foundation models (AFMs). To validate our hypothesis, we extract representations from state-of-the-art (SOTA) MFMs and AFMs and evaluated them on benchmark NVER datasets. We also investigate the potential of combining selected foundation model (FM) representations to enhance NVER further inspired by research in speech recognition and audio deepfake detection. To achieve this, we propose a framework called MATA (Intra-Modality Alignment through Transport Attention). Through MATA coupled with the combination of MFMs: LanguageBind and ImageBind, we report the topmost performance with accuracies of 76.47%, 77.40%, 75.12% and F1-scores of 70.35%, 76.19%, 74.63% for ASVP-ESD, JNV, and VIVAE datasets against individual FMs and baseline fusion techniques and report SOTA on the benchmark datasets. Orchid Chetia Phukan, Mohd Mujtaba Akhtar, Girish, Swarup Ranjan Behera, Sishir Kalita, Arun Balaji Buduru, Rajesh Sharma 0002, S. R. Mahadeva Prasanna |
ICASSP | 8 |
| 2025 | Tone recognition in low-resource languages of North-East India: peeling the layers of SSL-based speech models
Parismita Gogoi, Sishir Kalita, Wendy Lalhminghlui, Viyazonuo Terhiija, Moakala Tzudir, Priyankoo Sarmah, S. R. Mahadeva Prasanna |
INTERSPEECH | 7 |
| 2025 | Leveraging AM and FM Rhythm Spectrograms for Dementia Classification and Assessment
Parismita Gogoi, Vishwanath Pratap Singh, Seema Khadirnaikar, Soma Siddhartha, Sishir Kalita, Jagabandhu Mishra, Md. Sahidullah, Priyankoo Sarmah, S. R. Mahadeva Prasanna |
INTERSPEECH | 9 |
| 2024 | Evaluating the Efficacy of Large Acoustic Model for Documenting Non-Orthographic Tribal Languages in IndiaabstractPre-trained Large Acoustic Models, when fine-tuned, have largely shown to improve the performances in various tasks related to spoken language technologies. However, their evaluation has been mostly on datasets that contain English or other widely spoken languages, and their potential for novel under-resourced languages is not fully known. In this work, four novel under-resourced tribal languages that do not have a standard writing system were introduced and the application of such large pre-trained models was assessed to document such languages using Automatic Speech Recognition and Direct Speech-to-Text Translation systems. The transcriptions for these tribal languages were generated by adapting scripts from those languages that held a prominent presence in the geographical regions where these tribal languages are spoken. The results from this study suggest a viable direction to document these languages in the electronic domain by using Spoken Language Technologies that incorporate LAMs. Additionally, this study helped in understanding the varying performances exhibited by the Large Acoustic Model between these four languages. This study not only informs the adoption of appropriate scripts for transliterating spoken-only languages based on the language family but also aids in making informed decisions in analyzing the behavior of particular Large Acoustic Model in linguistic contexts. Tonmoy Rajkhowa, Amartya Chowdhury, Hrishikesh Ravindra Karande, S. R. Mahadeva Prasanna |
LREC/COLING | 4 |
| 2024 | The Second DISPLACE Challenge: DIarization of SPeaker and LAnguage in Conversational Environments
Shareef Babu Kalluri, Prachi Singh, Pratik Roy Chowdhuri, Apoorva Kulkarni, Shikha Baghel, Pradyoth Hegde, Swapnil Sontakke, K. T. Deepak, S. R. Mahadeva Prasanna, Deepu Vijayasenan, Sriram Ganapathy |
INTERSPEECH | 9 |
| 2024 | TM-PATHVQA: 90000+ Textless Multilingual Questions for Medical Visual Question AnsweringabstractIn healthcare and medical diagnostics, Visual Question Answering (VQA) may emerge as a pivotal tool in scenarios where analysis of intricate medical images becomes critical for accurate diagnoses.Current text-based VQA systems limit their utility in scenarios where hands-free interaction and accessibility are crucial while performing tasks.A speech-based VQA system may provide a better means of interaction where information can be accessed while performing tasks simultaneously.To this end, this work implements a speech-based VQA system by introducing a Textless Multilingual Pathological VQA (TM-PathVQA) dataset, an expansion of the PathVQA dataset, containing spoken questions in English, German & French.This dataset comprises 98,397 multilingual spoken questions and answers based on 5,004 pathological images along with 70 hours of audio.Finally, this work benchmarks and compares TM-PathVQA systems implemented using various combinations of acoustic and visual features. Tonmoy Rajkhowa, Amartya Chowdhury, Sankalp Nagaonkar, Achyut Mani Tripathi, S. R. Mahadeva Prasanna |
INTERSPEECH | 5 |
| 2024 | Spectro-Temporally Compressed Source Features for Replay Attack DetectionabstractThe role of pitch-synchronous source features in the detection of replay attacks on speaker verification systems has been explored earlier. This work presents some advancements in the processing which enable the use of the entire source signal as well as better capture of the source dynamics. The resulting features are used to develop replay attack detection systems using both the Gaussian mixture model and ResNet-18 classifiers. The performances of these systems are evaluated on the physical attack subset of the ASVSpoof 2019 database. In a similar setup, the proposed features are noted to yield quite competitive performance compared to some of the existing replay attack detection features. Sarfaraz Jelil, Rohit Sinha 0003, S. R. Mahadeva Prasanna |
IEEE Signal Process. Lett. | 3 |
| 2024 | Cross-linguistic rhythm analysis of Mising and AssameseabstractThe objective of the current study is to explore a quantitative frequency domain technique to evaluate rhythm in spontaneous speech data of 19 native speakers of Mising and Assamese, two low-resourced languages spoken in Assam, North-East India. The concept of analyzing speech rhythm using amplitude modulation (AM) low-frequency (LF) spectrum, also known as rhythm formant analysis (RFA), is initially put forth by Gibbon and Li [ 17 ]. We propose three features from rhythm formants of the LF spectrum and also explore discrete cosine transform (DCT)–based characterization of the retrieved LF spectrum. We aim to distinguish the rhythm of Assamese and two Mising dialects, namely Pagro and Delu, with the aid of machine learning techniques fed with the derived features as input. We have observed that the features are efficient in classifying Assamese vs Pagro and Assamese vs Delu with an accuracy of 92.73% and 91.15%, respectively. The experimental analysis further reveals that Assamese is rhythmically closer to Delu than Pagro. Parismita Gogoi, Priyankoo Sarmah, S. R. Mahadeva Prasanna |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2024 | Implicit Self-Supervised Language Representation for Spoken Language DiarizationabstractThe use of spoken language diarization (LD) as a preprocessing system might be essential in a code-switched (CS) scenario. Furthermore, implicit frameworks are preferable to explicit ones, as implicit frameworks can be easily adapted to deal with low/zero resource languages. Inspired by speaker diarization literature, three frameworks based on (a) fixed segmentation, (b) change-point-based segmentation, and (c) end-to-end (E2E) are used in this study to perform LD. The initial exploration in the constructed text-to-speech female language diarization (TTSF-LD) dataset shows, that using the x-vector as implicit language representation with appropriate analysis window length achieves, comparable performance to explicit LD. The best implicit LD performance of 6.4% in terms of Jaccard error rate (JER) is achieved by using the E2E framework. However, using the natural Microsoft CS dataset, the performance of the E2E implicit LD degrades to 60.4% JER. The performance degradation is due to the inability of the x-vector representation to capture language-specific traits. To address this shortcoming, a self-supervised implicit language representation framework is used in this study. Compared to the x-vector representation, the self-supervised representation yields a relative improvement of 63.9%, achieving a JER of 21.8% when used in conjunction with the E2E framework. Jagabandhu Mishra, S. R. Mahadeva Prasanna |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | Exploration of Speech and Music Information for Movie Genre ClassificationabstractMovie genre prediction from trailers is mostly attempted in a multi-modal manner. However, the characteristics of movie trailer audio indicate that this modality alone might be highly effective in genre prediction. Movie trailer audio predominantly consists of speech and music signals in isolation or overlapping conditions. This work hypothesizes that the genre labels of movie trailers might relate to the composition of their audio component. In this regard, speech-music confidence sequences for the trailer audio are used as a feature. In addition, two other features previously proposed for discriminating speech-music are also adopted in the current task. This work proposes a time and channel Attention Convolutional Neural Network (ACNN) classifier for the genre classification task. The convolutional layers in ACNN learn the spatial relationships in the input features. The time and channel attention layers learn to focus on crucial timesteps and CNN kernel outputs, respectively. The Moviescope dataset is used to perform the experiments, and two audio-based baseline methods are employed to benchmark this work. The proposed feature set with the ACNN classifier improves the genre classification performance over the baselines. Moreover, decent generalization performance is obtained for genre prediction of movies with different cultural influences (EmoGDB). Mrinmoy Bhattacharjee, S. R. Mahadeva Prasanna, Prithwijit Guha |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | End to End Spoken Language Diarization with Wav2vec Embeddings
Jagabandhu Mishra, Jayadev N. Patil, Amartya Chowdhury, S. R. Mahadeva Prasanna |
INTERSPEECH | 4 |
| 2023 | Clean vs. Overlapped Speech-Music Detection Using Harmonic-Percussive Features and Multi-Task LearningabstractDetection of speech and music signals in isolated and overlapped conditions is an essential preprocessing step for many audio applications. Speech signals have wavy and continuous harmonics, while music signals exhibit horizontally linear and discontinuous harmonic patterns. Music signals also contain more percussive components than speech signals, manifested as vertical striations in the spectrograms. In case of speech music overlap, it might be challenging for automatic feature learning systems to extract class-specific horizontal and vertical striations from the combined spectrogram representation. A pre-processing step of separating the harmonic and percussive components before training might aid the classifier. Thus, this work proposes the use of harmonic-percussive source separation method to generate features for better detection of speech and music signals. Additionally, this work also explores the traditional and cascaded-information multi-task learning (MTL) frameworks to design better classifiers. MTL framework aids the training of the main task by employing simultaneous learning of several related auxiliary tasks. Results have been reported both on synthetically generated speech music overlapped signals and real recordings. Four state-of-the-art approaches are used for performance comparison. Experiments show that harmonic and percussive decomposition of spectrograms perform better as features. Moreover, the MTL-framework based classifiers further improve performances. Mrinmoy Bhattacharjee, S. R. Mahadeva Prasanna, Prithwijit Guha |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Prosodic Information in Dialect Identification of a Tonal Language: The case of Ao
Moakala Tzudir, Priyankoo Sarmah, S. R. Mahadeva Prasanna |
INTERSPEECH | 3 |
| 2022 | Speech/music classification using phase-based and magnitude-based features
Mrinmoy Bhattacharjee, S. R. Mahadeva Prasanna, Prithwijit Guha |
Speech Commun. | 2 |
| 2021 | Automatic Detection of Shouted Speech Segments in Indian News Debates
Shikha Baghel, Mrinmoy Bhattacharjee, S. R. Mahadeva Prasanna, Prithwijit Guha |
Interspeech | 3 |
| 2021 | Excitation Source Feature Based Dialect Identification in Ao - A Low Resource Language
Moakala Tzudir, Shikha Baghel, Priyankoo Sarmah, S. R. Mahadeva Prasanna |
Interspeech | 4 |
| 2021 | Enhancing the Intelligibility of Cleft Lip and Palate Speech Using Cycle-Consistent Adversarial NetworksabstractCleft lip and palate (CLP) refer to a congenital craniofacial condition that causes various speech-related disorders. As a result of structural and functional deformities, the affected subjects' speech intelligibility is significantly degraded, limiting the accessibility and usability of speech-controlled devices. Towards addressing this problem, it is desirable to improve the CLP speech intelligibility. Moreover, it would be useful during speech therapy. In this study, the cycle-consistent adversarial network (CycleGAN) method is exploited for improving CLP speech intelligibility. The model is trained on native Kannada-speaking childrens' speech data. The effectiveness of the proposed approach is also measured using automatic speech recognition performance. Further, subjective evaluation is performed, and those results also confirm the intelligibility improvement in the enhanced speech over the original. Protima Nomo Sudro, Rohan Kumar Das, Rohit Sinha 0003, S. R. Mahadeva Prasanna |
SLT | 4 |
| 2020 | Spectral Moment and Duration of Burst of Plosives in Speech of Children with Hearing Impairment and Typically Developing Children - A Comparative Study
Ajish K. Abraham, M. Pushpavathi, Narayan sreedevi, A. Navya, Vikram C. M., S. R. Mahadeva Prasanna |
INTERSPEECH | 6 |
| 2020 | VOP Detection in Variable Speech Rate Condition
Ayush Agarwal, Jagabandhu Mishra, S. R. Mahadeva Prasanna |
INTERSPEECH | 3 |
| 2020 | Lexical Tone Recognition in Mizo using Acoustic-Prosodic FeaturesabstractMizo is an under-studied Tibeto-Burman tonal language of the North-East India. Preliminary research findings have confirmed that four distinct tones of Mizo (High, Low, Rising and Falling) appear in the language. In this work, an attempt is made to automatically recognize four phonological tones in Mizo distinctively using acoustic-prosodic parameters as features. Six features computed from Fundamental Frequency (F0) contours are considered and two classifier models based on Support Vector Machine (SVM) & Deep Neural Network (DNN) are implemented for automatic tonerecognition task respectively. The Mizo database consists of 31950 iterations of the four Mizo tones, collected from 19 speakers using trisyllabic phrases. A four-way classification of tones is attempted with a balanced (equal number of iterations per tone category) dataset for each tone of Mizo. it is observed that the DNN based classifier shows comparable performance in correctly recognizing four phonological Mizo tones as of the SVM based classifier. Parismita Gogoi, Abhishek Dey, Wendy Lalhminghlui, Priyankoo Sarmah, S. R. Mahadeva Prasanna |
LREC | 5 |
| 2020 | Sinusoidal model-based hypernasality detection in cleft palate speech using CVCV sequence
Akhilesh Kumar Dubey, S. R. Mahadeva Prasanna, Samarendra Dandapat |
Speech Commun. | 2 |
| 2020 | Enhancement of cleft palate speech using temporal and spectral processing
Protima Nomo Sudro, S. R. Mahadeva Prasanna |
Speech Commun. | 2 |
| 2020 | Speech/Music Classification Using Features From Spectral PeaksabstractSpectrograms of speech and music contain distinct striation patterns. Traditional features represent various properties of the audio signal but do not necessarily capture such patterns. This work proposes to model such spectrogram patterns using a novel Spectral Peak Tracking (SPT) approach. Two novel time-frequency features for speech vs. music classification are proposed. The proposed features are extracted in two stages. First, SPT is performed to track a preset number of highest amplitude spectral peaks in an audio interval. In the second stage, the location and amplitudes of these peak traces are used to compute the proposed feature sets. The first feature involves the computation of mean and standard deviation of peak traces. The second feature is obtained as averaged component posterior probability vectors of Gaussian mixture models learned on the peak traces. Speech vs. music classification is performed by training various binary classifiers on these proposed features. Three standard datasets are used to evaluate the efficiency of the proposed features for speech/music classification. The proposed features are benchmarked against five baseline approaches. Finally, the best-proposed feature is combined with two contemporary deep-learning based features to show that such combinations can lead to more robust speech vs. music classification systems. Mrinmoy Bhattacharjee, S. R. Mahadeva Prasanna, Prithwijit Guha |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Vowel Onset Point Based Screening of Misarticulated Stops in Cleft Lip and Palate SpeechabstractThe presence of velopharyngeal dysfunction, dental occlusion, and mislearned articulation in individuals with cleft lip and palate (CLP) results in the production of misarticulated stop consonants. The present work considers vowel onset points (VOPs) as the anchor points, around which the consonant-vowel (CV) transition regions are segmented to analyze the difference between normal and misarticulated stops. VOPs are located using an epoch-synchronously computed feature called maximum weighted inner product. Spectro-temporal dynamics of CV transitions anchored around VOP are analyzed using two-dimensional discrete cosine transform (2D-DCT) coefficients, where 2D-DCT coefficients are derived from single pole filter (SPF) based time-frequency representation. The SPF-based 2D-DCT coefficients are used to train a support vector machine for the classification of normal and misarticulated stops, where the class of misarticulated stops includes weak, nasalized, palatal, velar, pharyngeal, glottal, and devoicing errors produced by CLP speakers. The performance of the proposed VOP detection algorithm is evaluated on a database containing CV units of normal and misarticulated stops, and the results are compared with the state-of-the-art VOP detection methods. The classification results obtained for the proposed SPF-based 2D-DCT coefficients are compared with the short-time Fourier transform-based 2D-DCT coefficients and Mel-frequency cepstral coefficients. Further, the performance of the proposed system is compared with the hidden Markov model-based goodness of pronunciation approach. Vikram C. M., S. R. Mahadeva Prasanna |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Glottal Instants Extraction from Speech Signal Using Generative Adversarial NetworkabstractThe Glottal Closure and Opening instants (GCIs and GOIs) form important events in excitation source signal. These instants represent closing and opening events of vocal folds while producing voiced speech signal. Estimation of such instants from speech signal is beneficial and several applications rely on accurate estimation of Closure and Opening instants. In this work, Electroglottographic like (EGG-like) signal is synthesized from speech signal using Generative Adversarial Network (GAN). The Glottal Closure and Opening instants are located using the derivative of EGG-like signal, which is essentially a difference EGG-like signal. The proposed method is evaluated on CMU-Arctic database, as the database consists of simultaneous recordings of speech and EGG signal, respectively. To evaluate the results, the locations obtained from synthesized EGG-like signal are com-pared with the reference difference EGG signal. The results are evaluated for both seen and unseen conditions. It is shown that the performance of GCI and GOI estimation is comparable to existing state-of-the-art methods. K. T. Deepak, Pavitra Kulkarni, Uma Mudenagudi, S. R. Mahadeva Prasanna |
ICASSP | 4 |
| 2019 | Synthesis of Handwriting Dynamics using Sinusoidal ModelabstractHandwriting production is a complex mechanism of fine motor control, associated with mainly two degrees of freedom in the horizontal and vertical directions. The relation between the horizontal and vertical velocities depends on the trajectory shape and its length. In this work, we explore the generation of handwriting velocities using two sinusoidal oscillations. The proposed method follows the motor equivalence theory and considers that the patterns are stored in the form of a sequence of corner shapes and its relative location in the letter. These points are referred to as the modulation points, where the parameters of the sinusoidal oscillations are modulated to generate required velocity profiles. Depending on the location and shape of the corners, the amplitude, phase, and frequency relations between the two underlying oscillations are changed. Accordingly, this paper presents an efficient method to synthesize the velocity profiles and hence the handwriting. Further, the shape variability in the synthesized data can also be introduced by modifying the position of the modulation points and its corner shapes. The quality of the synthesized handwriting is evaluated using both subjective and quantitative evaluation methods. Himakshi Choudhury, S. R. Mahadeva Prasanna |
ICDAR | 2 |
| 2019 | Exploration of CNN Features for Online Handwriting RecognitionabstractRecently, convolution neural network (CNN) has demonstrated its powerful ability in learning features particularly from image data. In this work, its capability of feature learning in online handwriting is explored, by constructing various CNN architectures. The developed CNNs can process online handwriting directly unlike the existing works that convert the online handwriting to an image to utilize the architecture. The first convolution layer accepts the sequence of (x; y) coordinates along the trace of the character as an input and outputs a convolved filtered signal. Thereafter, via alternating steps of convolution and Rectified Linear Unit layers, in a hierarchical fashion, we obtain a set of deep features that can be employed for classification. We utilize the proposed CNN features to develop a Support Vector Machine (SVM)-based character recognition system and an implicit-segmentation based large vocabulary word recognition system employing hidden Markov model (HMM) framework. To the best of our knowledge, this is the first work of its kind that applies CNN directly on the (x; y) coordinates of the online handwriting data. Experiments are carried out on two publicly available English online handwritten database: UNIPEN character and UNIPEN ICROW-03 word databases. The obtained results are promising over the reported works employing the point-based features. Subhasis Mandal, S. R. Mahadeva Prasanna, Suresh Sundaram 0001 |
ICDAR | 2 |
| 2019 | Hypernasality Severity Detection Using Constant Q Cepstral Coefficients
Akhilesh Kumar Dubey, S. R. Mahadeva Prasanna, Samarendra Dandapat |
INTERSPEECH | 2 |
| 2019 | SpeechMarker: A Voice Based Multi-Level Attendance Application
Sarfaraz Jelil, Rohan Kumar Das, S. R. Mahadeva Prasanna, Rohit Sinha 0003 |
INTERSPEECH | 4 |
| 2019 | Nasal Air Emission in Sibilant Fricatives of Cleft Lip and Palate Speech
Sishir Kalita, Protima Nomo Sudro, S. R. Mahadeva Prasanna, Samarendra Dandapat |
INTERSPEECH | 3 |
| 2019 | Modification of Devoicing Error in Cleft Lip and Palate Speech
Protima Nomo Sudro, S. R. Mahadeva Prasanna |
INTERSPEECH | 2 |
| 2019 | Emotion Recognition from Raw Speech using WavenetabstractThis paper proposes the use of Wavenet architecture to the task of speech emotion recognition using raw speech signals. In contrast to the conventional deep learning methods, Wavenet utilises dilation filters, residual blocks and skip connections to model the long-term dependencies in speech signals and eliminates the need for LSTMs for the same. Wavenet is trained for the classification task on two popular datasets- EMO-DB and IEMOCAP. Experimental findings along with reasoning have been presented and compared with the standard baseline CNN+LSTM network for the four basic human emotions- Angry, Happy, Neutral and Sad. Sandeep Kumar Pandey, Hanumant Singh Shekhawat, S. R. Mahadeva Prasanna |
TENCON | 3 |
| 2019 | Acoustic Correlates of Aspiration in Fricatives and NasalsabstractThis paper focuses on the phonetic analysis of Korean and Rabha fricatives and Angami nasals. Though aspirated consonants have been studied earlier, very few studies were found for the comparative study of aspirated fricatives and aspirated nasals. Previous literature has suggested the presence of aspirated fricatives. As there are limited studies on aspirated nasals, this paper tries to investigate the properties of aspirated nasals by comparing them with the aspiration in Korean aspirated fricative and analyses the feature that might distinguish between the aspirated and unaspirated counterparts of both the consonants. Features such as Intensity, Duration, Centre of Gravity (COG), F1 onset and Spectral tilt (H1-H2) are used to investigate whether there is a distinction between the aspirated and unaspirated fricatives and nasals. Results confirm that COG is a distinctive acoustic cue to discriminate the aspirated and unaspirated counterparts. Saswati Rabha, Priyankoo Sarmah, S. R. Mahadeva Prasanna |
TENCON | 3 |
| 2019 | Exploiting forced alignment of time-reversed data for improving HMM-based handwriting segmentation
Himakshi Choudhury, Subhasis Mandal, S. R. Mahadeva Prasanna |
Expert Syst. Appl. | 3 |
| 2019 | An improved discriminative region selection methodology for online handwriting recognition
Subhasis Mandal, S. R. Mahadeva Prasanna, Suresh Sundaram 0001 |
Int. J. Document Anal. Recognit. | 2 |
| 2019 | Representation of online handwriting using multi-component sinusoidal model
Himakshi Choudhury, S. R. Mahadeva Prasanna |
Pattern Recognit. | 2 |
| 2019 | Handwriting recognition using sinusoidal model parameters
Himakshi Choudhury, S. R. Mahadeva Prasanna |
Pattern Recognit. Lett. | 2 |
| 2019 | Detection of Nasalized Voiced Stops in Cleft Palate Speech Using Epoch-Synchronous FeaturesabstractThe presence of velopharyngeal dysfunction in individuals with cleft palate (CP) nasalizes the voiced stops. Due to this, voiced stops (/b/, /d/, /g/) tend to be perceive like nasal consonants (/m/, /n/, /ng/). In this work, a novel algorithm is proposed for the detection of nasalized voiced stops in CP speech using epoch-synchronous features. Speech regions corresponding to consonant and consonant-vowel transitions are segmented using the knowledge of glottal activity, syllable nucleus, low-frequency spectral dominance, and vowel onset point. The segmented regions are epoch-synchronously processed to analyze the spectral, spectro-temporal, excitation source, and periodicity characteristics of normal and nasalized voiced stops. Spectral and spectro temporal features are computed using single pole filter based time-frequency representation. The amplitude of Hilbert envelope of linear prediction residual, measured around the epoch is used to analyze the effect of nasalization on excitation source. Comparison of speech frames of successive inter-epoch intervals is carried out to analyze the periodicity characteristics. The proposed features are used to develop a support vector machine classifier for the classification of normal and nasalized voiced stops. Segmentation accuracy for the proposed knowledge based method is found to be better than the hidden Markov model based force-alignment approach. The detection rate of nasalized voiced stops is found to be high for the proposed epoch synchronous features than the conventional Mel-frequency cepstral coefficients. Vikram C. M., Nagaraj Adiga, S. R. Mahadeva Prasanna |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | DNN-HMM Based Large Vocabulary Online Handwritten Assamese Word Recognition SystemabstractIn this work, we consider recognizing online handwritten Assamese word (an Indic script) using the hybrid deep neural network - hidden Markov model (DNN-HMM) framework. The recognition task is generally challenging since Assamese handwriting is mixed with the discrete and cursive writing style, and further complicated due to the placing of vowel/consonant modifier above and/or below already written character in delayed order. Also, as it contains large character set, there is lack of common definition of the basic unit (BU) for automatic recognition. As a step in this direction, first, we select 173 BU capable of characterizing 20K most frequent Assamese words. The HMMs are created for these selected BUs. Next, we create a lexicon for word recognition where each word is represented by all probable "sequence of BUs" that constitute the word. The system is developed using the state-of-the-art Kaldi Automatic speech recognition toolkit, under large vocabulary word recognition framework and is evaluated on the lexicon of sizes 1K, 5k, 10K and 20K words. To the best of our knowledge, this paper is the first of its kind that considers online handwritten Assamese word instead of isolated characters. The experiments are conducted on locally collected Assamese word database, and a promising recognition performance is obtained, noting that the DNN-HMM framework outperforms the GMM-HMM system significantly Subhasis Mandal, Himakshi Choudhury, S. R. Mahadeva Prasanna, Suresh Sundaram 0001 |
ICFHR | 3 |
| 2018 | Exploring Sparse Representation for Improved Online Handwriting RecognitionabstractThis work studies the sparse representation based classification (SRC) framework for online handwriting recognition (HR) task. In this framework, first, an exemplar dictionary is created using the training samples from each of the classes in the chosen task. Subsequently, the test samples are sparse coded over exemplar dictionary for classification. In sparse coding, both l_0 - and l_1 -norm based greedy algorithms are studied. Further, for reducing the computational cost of the SRC-based HR approach, the learned exemplar dictionary has also been explored. The proposed SRC-based approach is demonstrated for character and limited vocabulary word recognition task and evaluated on three different corpora: the Assamese digit database, the UNIPEN English character database and the UNIPEN ICROW-03 English word database. The experimental results are promising over the reported works on these databases employing the hidden Markov model or the support vector machine. Subhasis Mandal, Syed Shahnawazuddin, Rohit Sinha 0003, S. R. Mahadeva Prasanna, Suresh Sundaram 0001 |
ICFHR | 4 |
| 2018 | Exploring Discriminative HMM States for Improved Recognition of Online HandwritingabstractIn this paper, we propose a novel approach for online handwriting recognition (HR) based on hidden Markov model (HMM). In a conventional HMM-based HR system, the input test sample is recognized by first measuring the log-likelihood score from each class-specific HMM, and then the class with the highest score is assigned as the recognized class. It is observed that, for a given test sample, the difference in log-likelihood scores of top-2 outputs (classes) is often less for faithful classification. The problem intensifies for those scripts that have a large set of similar shape characters such as the Indic script. To address this problem, first, we analyze the HMM states corresponding to the top-2 classes and identify a subset of states that most discriminate the two classes. Afterwards, the final recognition among the two classes is carried out by comparing the log-likelihood scores of these chosen states. Since the proposed methodology focuses only on the most discriminative states of the two classes, therefore it enhances the classification confidence as well as overall recognition accuracy with least added complexity. The proposal is demonstrated for character and limited vocabulary word recognition tasks and evaluated on the locally collected Assamese character and word databases. The experimental results are promising over the conventional HMM-based HR system. Subhasis Mandal, Himakshi Choudhury, S. R. Mahadeva Prasanna, Suresh Sundaram 0001 |
ICPR | 3 |
| 2018 | Glotto Vibrato Graph: A Device and Method for Recording, Analysis and Visualization of Glottal Activity
Kishalay Chakraborty, Senjam Shantirani Devi, Sanjeevan Devnath, S. R. Mahadeva Prasanna, Priyankoo Sarmah |
INTERSPEECH | 4 |
| 2018 | AGROASSAM: A Web Based Assamese Speech Recognition Application for Retrieving Agricultural Commodity Price and Weather Information
Abhishek Dey, Abhash Deka, Siddika Imani, Barsha Deka, Rohit Sinha 0003, S. R. Mahadeva Prasanna, Priyankoo Sarmah, K. Samudravijaya, S. R. Nirmala |
INTERSPEECH | 6 |
| 2018 | Robust Mizo Continuous Speech Recognition
Abhishek Dey, Biswajit Dev Sarma, Wendy Lalhminghlui, Lalnunsiami Ngente, Parismita Gogoi, Priyankoo Sarmah, S. R. Mahadeva Prasanna, Rohit Sinha 0003, S. R. Nirmala |
INTERSPEECH | 7 |
| 2018 | Pitch-Adaptive Front-end Feature for Hypernasality Detection
Akhilesh Kumar Dubey, S. R. Mahadeva Prasanna, Samarendra Dandapat |
INTERSPEECH | 2 |
| 2018 | Analysis of Breathiness in Contextual Vowel of Voiceless Nasals in Mizo
Pamir Gogoi, Sishir Kalita, Parismita Gogoi, Ratree Wayland, Priyankoo Sarmah, S. R. Mahadeva Prasanna |
INTERSPEECH | 6 |
| 2018 | Exploration of Compressed ILPR Features for Replay Attack Detection
Sarfaraz Jelil, Sishir Kalita, S. R. Mahadeva Prasanna, Rohit Sinha 0003 |
INTERSPEECH | 3 |
| 2018 | Self-similarity Matrix Based Intelligibility Assessment of Cleft Lip and Palate Speech
Sishir Kalita, S. R. Mahadeva Prasanna, Samarendra Dandapat |
INTERSPEECH | 2 |
| 2018 | Epoch Extraction from Pathological Children Speech Using Single Pole Filtering Approach
Vikram C. M., S. R. Mahadeva Prasanna |
INTERSPEECH | 2 |
| 2018 | Detection of Glottal Activity Errors in Production of Stop Consonants in Children with Cleft Lip and Palate
Vikram C. M., S. R. Mahadeva Prasanna, Ajish K. Abraham, Pushpavathi M, Girish K. S |
INTERSPEECH | 2 |
| 2018 | Estimation of Hypernasality Scores from Cleft Lip and Palate Speech
Vikram C. M., Ayush Tripathi, Sishir Kalita, S. R. Mahadeva Prasanna |
INTERSPEECH | 4 |
| 2018 | Spoken Keyword Detection Using Joint DTW-CNN
Vikram C. M., S. R. Mahadeva Prasanna |
INTERSPEECH | 3 |
| 2018 | Processing Transition Regions of Glottal Stop Substituted /S/ for Intelligibility Enhancement of Cleft Palate Speech
Protima Nomo Sudro, Sishir Kalita, S. R. Mahadeva Prasanna |
INTERSPEECH | 3 |
| 2018 | Speaker Identification Using Tensor Decomposition of Acoustic ModelsabstractThis paper explores speaker identification based on the speaker adaptation via multilinear decomposition of a speaker model. Tucker decomposition of the third order mean Tensor of training speaker yields three subspaces corresponding to each mode. The mean of the mixtures for speakers is expressed as a product of the mixture space and a weight matrix comprising of the other two spaces. We update only the mean of the mixtures for the adaptation stage. During testing, log likelihood is used to identify the scores for the test speakers.Experiments conducted on the TIMIT and NIST 2003 databases shows comparable performance with the baseline even in channel mismatch conditions of the development and enrollment dataset. Furthermore, using higher order tensors, we can easily adapt this problem to include noise and environment factors as well. Sandeep Kumar Pandey, Sarfaraz Jelil, S. R. Mahadeva Prasanna, Hanumant Singh Shekhawat |
TENCON | 3 |
| 2018 | GMM posterior features for improving online handwriting recognition
Subhasis Mandal, S. R. Mahadeva Prasanna, Suresh Sundaram 0001 |
Expert Syst. Appl. | 2 |
| 2018 | Analysis of the Hilbert Spectrum for Text-Dependent Speaker Verification
Rajib Sharma, Ramesh K. Bhukya, S. R. Mahadeva Prasanna |
Speech Commun. | 3 |
| 2018 | Significance of sonority information for voiced/unvoiced decision in speech synthesis
Bidisha Sharma, S. R. Mahadeva Prasanna |
Speech Commun. | 2 |
| 2017 | Zero Frequency Filter Based Analysis of Voice Disorders
Nagaraj Adiga, Vikram C. M., Keerthi Pullela, S. R. Mahadeva Prasanna |
INTERSPEECH | 4 |
| 2017 | Phase Modeling Using Integrated Linear Prediction Residual for Statistical Parametric Speech Synthesis
Nagaraj Adiga, S. R. Mahadeva Prasanna |
INTERSPEECH | 2 |
| 2017 | Spoof Detection Using Source, Instantaneous Frequency and Cepstral Features
Sarfaraz Jelil, Rohan Kumar Das, S. R. Mahadeva Prasanna, Rohit Sinha 0003 |
INTERSPEECH | 3 |
| 2017 | Hypernasality Severity Analysis in Cleft Lip and Palate Speech Using Vowel Space Area
Nikitha K., Sishir Kalita, Vikram C. M., M. Pushpavathi, S. R. Mahadeva Prasanna |
INTERSPEECH | 5 |
| 2017 | Acoustic Characterization of Word-Final Glottal Stops in Mizo and Assam Sora
Sishir Kalita, Wendy Lalhminghlui, Luke Horo, Priyankoo Sarmah, S. R. Mahadeva Prasanna, Samarendra Dandapat |
INTERSPEECH | 5 |
| 2017 | Indoor/Outdoor Audio Classification Using Foreground Speech Segmentation
Banriskhem K. Khonglah, K. T. Deepak, S. R. Mahadeva Prasanna |
INTERSPEECH | 3 |
| 2017 | IITG-Indigo System for NIST 2016 SRE ChallengeabstractOrientador : Juarez Brandão Lopes Nagendra Kumar 0004, Rohan Kumar Das, Sarfaraz Jelil, Dhanush B. K, H. Kashyap, K. Sri Rama Murty, Sriram Ganapathy, Rohit Sinha 0003, S. R. Mahadeva Prasanna |
INTERSPEECH | 9 |
| 2017 | Vowel Onset Point Detection Using Sonority Information
Bidisha Sharma, S. R. Mahadeva Prasanna |
INTERSPEECH | 2 |
| 2017 | Exploring kernel discriminant analysis for speaker verification with limited test data
Rohan Kumar Das, Akhil Babu Manam, S. R. Mahadeva Prasanna |
Pattern Recognit. Lett. | 3 |
| 2017 | Consonant-vowel unit recognition using dominant aperiodic and transition region detection
Biswajit Dev Sarma, S. R. Mahadeva Prasanna, Priyankoo Sarmah |
Speech Commun. | 2 |
| 2017 | Analysis of the Intrinsic Mode Functions for Speaker Information
Rajib Sharma, S. R. Mahadeva Prasanna, Ramesh K. Bhukya, Rohan Kumar Das |
Speech Commun. | 2 |
| 2017 | Empirical Mode Decomposition for adaptive AM-FM analysis of Speech: A Review
Rajib Sharma, Leandro Daniel Vignolo, Gastón Schlotthauer, Marcelo Alejandro Colominas, Hugo Leonardo Rufiner, S. R. Mahadeva Prasanna |
Speech Commun. | 6 |
| 2017 | Enhancement of Spectral Tilt in Synthesized SpeechabstractThe research in statistical parametric speech synthesis is towards improving naturalness and intelligibility. In this work, the deviation in spectral tilt of the natural and synthesized speech is analyzed and observed a large gap between the two. Furthermore, the same is analyzed for different classes of sounds, namely low-vowels, mid-vowels, high-vowels, semi-vowels, nasals, and found to be varying with category of sound units. Based on variation, a novel method for spectral tilt enhancement is proposed, where the amount of enhancement introduced is different for different classes of sound units. The proposed method yields improvement in terms of intelligibility, naturalness, and speaker similarity of the synthesized speech. Bidisha Sharma, S. R. Mahadeva Prasanna |
IEEE Signal Process. Lett. | 2 |
| 2017 | Epoch Extraction From Telephone Quality Speech Using Single Pole FilterabstractEpoch extraction from speech involves the suppression of vocal tract resonances, either by linear prediction based inverse filtering or filtering at very low frequency. Degradations due to channel effect and significant attenuation of low frequency components (<;300 Hz) create challenges for the epoch extraction from telephone quality speech. An epoch extraction method is proposed that considers the vertical striations present in the time-frequency representation of voiced speech as the representative candidates for the epochs. Time-frequency representation with better localized vertical striations is estimated using single pole filter based filter bank. The time marginal of time-frequency representation is computed to locate the epochs. The proposed algorithm is evaluated on the database of five speakers, which provide simultaneous speech and electroglottographic recordings. Telephone quality speech is simulated using G.191 software tools. The identification rate of the state-of-the-art methods degrades substantially for the telephone quality speech whereas that of the proposed method remains the same, comparable to that of clean speech. Vikram C. M., S. R. Mahadeva Prasanna |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Sonority Measurement Using System, Source, and Suprasegmental InformationabstractSonorant sounds are characterized by regions with a prominent formant structure, high energy, and high degree of periodicity. In this work, the vocal-tract system, excitation source, and suprasegmental features derived from the speech signal are analyzed to measure the sonority information present in each of them. Vocal-tract system information is extracted from the Hilbert envelope of the numerator of the group-delay function. It is derived from a zero-time-windowed speech signal that provides a better resolution of the formants. A 5-D feature set is computed from the estimated formants to measure the prominence of the spectral peaks. A feature representing strength of excitation is derived from the Hilbert envelope of linear prediction residual, which represents the source information. Correlation of speech over ten consecutive pitch periods is used as the suprasegmental feature representing periodicity information. The combination of evidence from the three different aspects of speech provides a better discrimination among different sonorant classes, compared to the baseline mel frequency cepstral coefficient features. The usefulness of the proposed sonority feature is demonstrated in the tasks of phoneme recognition and sonorant classification. Bidisha Sharma, S. R. Mahadeva Prasanna |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Source modeling for HMM based speech synthesis using integrated LP residualabstractIn this work, new method of source modeling for HMM based speech synthesis is proposed using integrated LP residual (ILPR). The nature of ILPR waveform resembles the glottal flow derivative signal and may keep the speaker characteristics in a better way. The ILPR signal is modeled in the frequency domain by dividing the spectrum into two bands to characterize harmonic and noise components of the voice speech segment. The harmonic components of ILPR signals below the maximum voiced frequency (fm) is modeled using mel-cepstral coefficients called as RMCEPs, whereas noise component above fmis modeled by pitch adaptive triangular noise envelope weighted by the strength of excitation (SoE). The RMCEPs and SoE are modeled on the HMM framework along with MCEPs and F0representing vocal tract information and fundamental frequency, respectively. The synthesized speech by the proposed source modeling reduces the buzziness and improves the speaker similarity compared to the conventional impulse / noise and mixed excitation source modeling and comparable with STRAIGHT based excitation. This is further reflected in both objective and subjective valuations. Nagaraj Adiga, S. R. Mahadeva Prasanna |
ICASSP | 2 |
| 2016 | Exploring Session Variability and Template Aging in Speaker Verification for Fixed Phrase Short Utterances
Rohan Kumar Das, Sarfaraz Jelil, S. R. Mahadeva Prasanna |
INTERSPEECH | 3 |
| 2016 | Analysis of Glottal Stop in Assam Sora Language
Sishir Kalita, Luke Horo, Priyankoo Sarmah, S. R. Mahadeva Prasanna, Samarendra Dandapat |
INTERSPEECH | 4 |
| 2016 | Spectral Enhancement of Cleft Lip and Palate Speech
Vikram C. M., Nagaraj Adiga, S. R. Mahadeva Prasanna |
INTERSPEECH | 3 |
| 2016 | Speech Synthesis in Noisy Environment by Enhancing Strength of Excitation and Formant Prominence
Bidisha Sharma, S. R. Mahadeva Prasanna |
INTERSPEECH | 2 |
| 2016 | Feature optimisation for stress recognition in speech
Leandro Daniel Vignolo, S. R. Mahadeva Prasanna, Samarendra Dandapat, Hugo Leonardo Rufiner, Diego H. Milone |
Pattern Recognit. Lett. | 2 |
| 2016 | Foreground Speech Segmentation and Enhancement Using Glottal Closure Instants and Mel Cepstral CoefficientsabstractIn this paper, the speech signal recorded from the desired speaker close to microphone in natural environment is regarded as foreground speech and rest of the interfering sources as background noise . The proposed paper exploits speech production features like glottal closure instants in time domain and vocal tract information in spectral domain to segment the desired speaker's speech and to further enhance it. The foreground speech is perceptually enhanced using the auditory perception feature in mel-frequency domain using mel-cepstral coefficients and its inversion using mel log spectrum approximation filter. The focus is on enhancing the production and perceptual features of foreground speech rather than relying on modeling the interfering sources. The speech data are collected in different natural environments from different speakers in order to evaluate the proposed method. The enhanced speech signals derived at three different stages of the proposed method are evaluated with state-of-the-art methods in terms of subjective and objective measures. The proposed method provides improved performance compared to the considered state-of-the-art methods. In terms of the proposed objective measure foreground to background Ratio, the enhancement approach presented in this paper gives an average improvement of 12 dB as opposed to existing spectral subtraction-based method which provides 3 dB. Moreover, subjective evaluation using 24 different subjects corroborates the objective test results. K. T. Deepak, S. R. Mahadeva Prasanna |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Speaker verification using Gaussian posteriorgrams on fixed phrase short utterances
Sarfaraz Jelil, Rohan Kumar Das, Rohit Sinha 0003, S. R. Mahadeva Prasanna |
INTERSPEECH | 4 |
| 2015 | Detection of mizo tones
Biswajit Dev Sarma, Priyankoo Sarmah, Wendy Lalhminghlui, S. R. Mahadeva Prasanna |
INTERSPEECH | 4 |
| 2015 | Detection of Glottal Activity Using Different Attributes of Source InformationabstractThe major activity during speech production is glottal activity and is earlier detected using strength of excitation (SoE). This work uses the normalized autocorrelation peak strength (NAPS) and higher order statistics (HOS) as additional features for detecting glottal activity. The three features, namely, SoE, NAPS, and HOS, are, respectively indicators of different attributes of glottal activity, namely, energy, periodicity, and asymmetrical nature of the resulting source signal. The effectiveness of these features is analyzed using the differential electroglottograph signal, zero-frequency filtered signal, and integrated linear prediction residual, as representatives of source signal. The combination of glottal activity information from the three features outperforms any single of them, demonstrating different information represented by each of these features. Nagaraj Adiga, S. R. Mahadeva Prasanna |
IEEE Signal Process. Lett. | 2 |
| 2014 | Combining source and system information for limited data speaker verificationabstractSpeaker verification using limited data is always a challenge for practical implementation as an application. An analysis on speaker verification studies for an i-vector based method using Mel-Frequency Cepstral Coefficient (MFCC) feature shows that the performance drops drastically as the duration of test data is reduced. This decrease in performance is due to insufficient phonetic coverage when we capture only the vocal tract feature. However the same can be improved if some source characteristics are taken into consideration. This paper attempts to improve the speaker verification performance using source characteristics. A recently proposed characterization of the voice source signal called the discrete cosine transform of the integrated linear prediction residual (DCTILPR) has been found to be useful as a speaker-specific feature. Speaker verification is performed over short test utterances in the NIST 2003 database using both the DCTILPR and MFCC features, and their score-level combination is found to give a significant performance improvement over the system using only the MFCC features. Rohan Kumar Das, S. Abhiram, S. R. Mahadeva Prasanna, A. G. Ramakrishnan |
INTERSPEECH | 3 |
| 2014 | Detection of vowel onset points in voiced aspirated sounds of indian languages
Biswajit Dev Sarma, S. R. Mahadeva Prasanna |
INTERSPEECH | 2 |
| 2014 | Analysis of Vocal Tract Constrictions using Zero Frequency FilteringabstractThis work proposes evidence using zero frequency filtering (ZFF) that gives an approximate measure of vocal tract constriction in terms of the low frequency component present in the speech signal. The vocal tract is completely closed in the case of voice bars and nasals and is wide open for low vowels. Intermediate cases are for high vowels, semivowels, laterals, voiced fricatives and other sounds. Vocal tract constriction affects the spectrum by reducing the first formant and attenuating the amplitude of the spectrum. The attenuation is relatively high in higher frequencies resulting in an increase in the low frequency component. The proposed method exploits the sinusoid like nature of ZFF signal (ZFFS) to obtain the evidence. Epoch synchronous analysis is performed and the ZFFS between successive epochs is compared with the corresponding speech segment using a cosine kernel. The low frequency dominant voiced regions match closely with ZFFS as compared to other regions and hence give higher value. This evidence when used as a feature gives relatively higher performance for the constricted phones in an HMM-based phoneme recognizer. Biswajit Dev Sarma, S. R. Mahadeva Prasanna |
IEEE Signal Process. Lett. | 2 |
| 2013 | The IITG speaker verification systems for NIST SRE 2012abstractIn this paper, we describe the speaker verification (SV) systems developed by Indian Institute of Technology Guwahati (IITG) for the NIST 2012 speaker recognition evaluations. The primary submission consists of five gender dependent SV systems combined at score level. Among the five systems two are based on sparse representation over learned and exemplar dictionaries, and the remaining are based on the generic i-vector and its variants obtained by vowel and non-vowel conditioning. The exemplar dictionary based system in particular exploits the new evaluation rule allowing the knowledge of all targets in each detection trial. The performance of the system is presented for the NIST SRE 2012 core task. Haris B. C., Gayadhar Pradhan, Rohit Sinha 0003, S. R. Mahadeva Prasanna |
ICASSP | 4 |
| 2013 | Significance of instants of significant excitation for source modeling
Nagaraj Adiga, S. R. Mahadeva Prasanna |
INTERSPEECH | 2 |
| 2013 | Detection of glottal opening instants using Hilbert envelopeabstractThe objective of this work is to develop an automatic method for estimating glottal opening instants (GOIs) using Hilbert envelope (HE). The GOIs are secondary major excitations after glottal closure instants (GCIs) during the production of voiced speech. The HE is defined as the magnitude of complex time function (CTF) of a given signal. The unipolar property of HE is exploited for picking the second largest peak present in a given glottal cycle and hypothesize as glottal opening instant (GOI). The electroglottogram (EGG) / speech signal is first passed through the zero frequency filtering (ZFF) method to extract GCIs. With the help of detected GCIs, the secondary peaks present in the HE of dEGG / residual are hypothesized as GOIs. The hypothesized GOIs are compared with secondary peaks estimated from the dEGG / residual. The GOIs hypothesized by the proposed method show less variance compared to peak picking from dEGG / residual. K. Ramesh 0002, S. R. Mahadeva Prasanna, D. Govind 0001 |
INTERSPEECH | 2 |
| 2013 | Speaker Verification by Vowel and Nonvowel Like SegmentationabstractThis work proposes methods for detecting vowel-like regions (VLRs) and non-vowel-like regions (non-VLRs) using excitation source information. The VLR onset and end points are hypothesized and used in an iterative algorithm for detecting the VLRs. Next, for detection of non-VLRs, the linear prediction (LP) residual samples in the VLRs are attenuated significantly to indirectly emphasize the residual samples in the non-VLRs. The modified LP residual samples excite the time varying all pole filter to reconstruct non-VLRs enhanced speech and used for detecting non-VLRs. The VLRs and non-VLRs are used independently during training and testing of a speaker verification (SV) system to reduce gross level mismatch due to sound units and achieve better compensation of degradation effects by applying different normalization to these two different energy regions. Finally, the scores are combined with higher weight on VLRs, which are more speaker specific. Experiments verify that the proposed approach provides improved performance for clean and degraded speech. On the NIST-2003 speaker recognition database, using VLRs and non-VLRs improves the equal error rate from 6.63% to 6% and from 2.29% to 1.89% for a GMM-UBM based and ani-vector based SV system, respectively. Gayadhar Pradhan, S. R. Mahadeva Prasanna |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Foreground Speech Segmentation using Zero Frequency Filtered Signal
K. T. Deepak, Biswajit Dev Sarma, S. R. Mahadeva Prasanna |
INTERSPEECH | 3 |
| 2011 | Study of robustness of zero frequency resonator method for extraction of fundamental frequencyabstractThe objective of this work is to develop and study the robustness of the zero frequency resonator (ZFR) based method for extraction of the fundamental frequency (F0) of speech signals. The proposed ZFR method for estimating F0consists of zero frequency filtering of the Hilbert envelope (HE) of the linear prediction (LP) residual of speech signal, followed by short-term spectrum analysis of the filtered output. The robustness of the proposed method is tested using speech signals collected in practical environments like distant, reverberant, telephone, mobile and multispeaker. Experimental results show that the proposed ZFR method estimates F0in majority of the cases. Bayya Yegnanarayana, S. R. Mahadeva Prasanna, Sunitha Guruprasad |
ICASSP | 2 |
| 2011 | Epoch Extraction in High Pass Filtered Speech Using Hilbert EnvelopeabstractHilbert envelope (HE) is defined as the magnitude of the analytic signal. This work proposes HE based zero frequency filtering (ZFF) approach for the extraction of epochs in high pass filtered speech. Epochs in speech correspond to instants of significant excitation like glottal closure instants. The ZFF method for epoch extraction is based on the signal energy around the impulse at zero frequency which seems to be significantly attenuated in case of high pass filtered speech. The low frequency nature of HE reinforces the signal energy around the impulse at zero frequency. This work therefore processes the HE of high pass filtered speech or its residual by zero frequency filtering for epoch extraction. The proposed approach shows significant improvement in performance for the high pass filtered speech compared to the conventional ZFF of speech. Index Terms: Epochs, zero frequency filtering, Hilbert envelope D. Govind 0001, S. R. Mahadeva Prasanna, Debadatta Pati |
INTERSPEECH | 2 |
| 2011 | Neutral to Target Emotion Conversion Using Source and Suprasegmental InformationabstractThis work uses instantaneous pitch and strength of excitation along with duration of syllable-like units as the parameters for emotion conversion. Instantaneous pitch and duration of the syllable-like units of the neutral speech are modified by the prosody modification of its linear prediction (LP) residual using the instants of significant excitation. The strength of excitation is modified by scaling the Hilbert envelope (HE) of the LP residual. The target emotion speech is then synthesized using the prosody and strength modified LP residual. The pitch, duration and strength modification factors for emotion conversion are derived using the syllable-like units of initial, middle and final regions from an emotion speech database having different speakers, texts and emotions. The effectiveness of the region wise modification of source and supra segmental features over the gross level modification is confirmed by the waveforms, spectrograms and subjective evaluations. Index Terms: Emotions, ZFF, strength of excitation, instantaneous pitch, duration D. Govind 0001, S. R. Mahadeva Prasanna, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2011 | Speaker recognition under limited data condition by noise addition
P. Krishnamoorthy, Haradagere Siddaramaiah Jayanna, S. R. Mahadeva Prasanna |
Expert Syst. Appl. | 3 |
| 2011 | Enhancement of noisy speech by temporal and spectral processing
P. Krishnamoorthy, S. R. Mahadeva Prasanna |
Speech Commun. | 2 |
| 2011 | Significance of Vowel-Like Regions for Speaker Verification Under Degraded ConditionsabstractVowel-like regions (VLRs) in speech includes vowels, semi-vowels, and diphthong sound units. VLR can be identified using a vowel-like region onset point (VLROP) event. By production, the VLR has impulse-like excitation and therefore information about the vocal tract system may be better manifested in them. Also, the VLR is a relatively high signal-to-noise ratio (SNR) region. Speaker information extracted from such a region may therefore be more speaker discriminative and relatively less affected by the degradations like noise, reverberation, and sensor mismatches. Due to this, better speaker modeling and reliable testing may be possible. In this paper, VLRs are detected using the knowledge of VLROPs during training and testing. Features from the VLRs are then used for training and testing the speaker models. As a result, significant improvement in the performance is reported for speaker verification under degraded conditions. S. R. Mahadeva Prasanna, Gayadhar Pradhan |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2010 | Analysis of instantaneous F0 contours from two speakers mixed signal using zero frequency filteringabstractInstantaneous fundamental frequency (F0) in voiced speech can be obtained from the sequence of epochs corresponding to the instants of significant excitation. The epoch sequence can be derived using the recently proposed epoch extraction method based on zero frequency filtering. The epoch extraction method is robust against additive noise degradation. But in a multispeaker mixed signal, the degradation is caused due to overlapping impulse-like excitations of two or more speakers. The feasibility of extracting the instantaneous F0contours from the two speaker mixed signal using zero frequency filtering is studied in this paper. The present study is based on deriving speaker-specific Hilbert Envelope (HE) signal which emphasizes peaks due to impulse-like excitation of one speaker and suppresses peaks due to other speaker. The epochs from this speaker-specific signal are obtained using the approach based on zero frequency filtering. The results of the proposed method is demonstrated for three different cases of mixed signals of two speakers data. Bayya Yegnanarayana, S. R. Mahadeva Prasanna |
ICASSP | 2 |
| 2010 | Analysis of excitation source information in emotional speechabstractThe objective of this work is to analyze the effect of emotions on the excitation source of speech production. The neutral, angry, happy, boredom and fear emotions are considered for the study. Initially the electroglottogram (EGG) and its derivative signals are compared across different emotions. The mean, standard deviation and contour of instantaneous pitch, and strength of excitation parameters are derived by processing the derivative of the EGG and also speech using zero-frequency filtering (ZFF) approach. The comparative study of these features across different emotions reveals that the effect of emotions on the excitation source is distinct and significant. The comparative study of the parameters from the derivative of EGG and speech waveform indicate that both cases have the same trend and range, inferring any of them may be used. Use of the computed parameters are found to be effective in the prosodic modification task. Index Terms: source, emotion, pitch, strength. S. R. Mahadeva Prasanna, D. Govind 0001 |
INTERSPEECH | 1 |
| 2009 | Vowel Onset Point Detection Using Source, Spectral Peaks, and Modulation Spectrum EnergiesabstractVowel onset point (VOP) is the instant at which the onset of vowel takes place during speech production. There are significant changes occurring in the energies of excitation source, spectral peaks, and modulation spectrum at the VOP. This paper demonstrates the independent use of each of these three energies in detecting the VOPs. Since each of these energies represents a different aspect of speech production, it may be possible that they contain complementary information about the VOP. The individual evidences are therefore combined for detecting the VOPs. The error rates measured as the ratio of missing and spurious to the total number of VOPs evaluated on the sentences taken from the TIMIT database are 6.92%, 8.8%, 6.13%, and 4.0% for source, spectral peaks, modulation spectrum, and combined information, respectively. The performance of the combined method for VOP detection is improved by 2.13% compared to the best performing individual VOP detection method. S. R. Mahadeva Prasanna, B. V. Sandeep Reddy, P. Krishnamoorthy |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | MRASTA and PLP in automatic speech recognition
S. R. Mahadeva Prasanna, Hynek Hermansky |
INTERSPEECH | 1 |
| 2007 | Determination of Instants of Significant Excitation in Speech Using Hilbert Envelope and Group Delay FunctionabstractThis letter proposes a time-effective method for determining the instants of significant excitation in speech signals. The instants of significant excitation correspond to the instants of glottal closure (epochs) in the case of voiced speech, and to some random excitations like onset of burst in the case of nonvoiced speech. The proposed method consists of two phases: the first phase determines the approximate epoch locations using the Hilbert envelope of the linear prediction residual of the speech signal. The second phase determines the accurate locations of the instants of significant excitation by computing the group delay around the approximate epoch locations derived from the first phase. The accuracy in determining the instants of significant excitation and the time complexity of the proposed method is compared with the group delay based approach. K. Sreenivasa Rao, S. R. Mahadeva Prasanna, Bayya Yegnanarayana |
IEEE Signal Process. Lett. | 2 |
| 2006 | Extraction of speaker-specific excitation information from linear prediction residual of speech
S. R. Mahadeva Prasanna, Cheedella S. Gupta, Bayya Yegnanarayana |
Speech Commun. | 1 |
| 2005 | Detection of vowel onset point events using excitation informationabstractThis paper proposes a method for the detection of Vowel Onset Point (VOP) events in speech using excitation information. VOP event is defined as the instant at which the onset of vowel takes place. For syllable-like units such as Consonant Vowel (CV) type, VOP event is the instant at which the consonant ends and the vowel begins. The speech signal is processed by the Linear Prediction (LP) analysis to extract the LP residual. The LP residual mostly contains the excitation information. The Hilbert envelope of the LP residual is derived using the analytic signal concept. A method is developed for detecting the VOP events using the Hilbert envelope of the LP residual and a modulated Gaussian window function. The performance of the proposed method is evaluated using reference VOP markings. The performance of the proposed method is also compared with the existing methods based on the vocal tract system features. The comparison shows that the excitation source also contains significant information about the VOP events. S. R. Mahadeva Prasanna, Bayya Yegnanarayana |
INTERSPEECH | 1 |
| 2005 | Speaker Localization Using Excitation Source Information in SpeechabstractThis paper presents the results of simulation and real room studies for localization of a moving speaker using information about the excitation source of speech production. The first step in localization is the estimation of time-delay from speech collected by a pair of microphones. Methods for time-delay estimation generally use spectral features that correspond mostly to the shape of vocal tract during speech production. Spectral features are affected by degradations due to noise and reverberation. This paper proposes a method for localizing a speaker using features that arise from the excitation source during speech production. Experiments were conducted by simulating different noise and reverberation conditions to compare the performance of the time-delay estimation and source localization using the proposed method with the results obtained using the spectrum-based generalized cross correlation (GCC) methods. The results show that the proposed method shows lower number of discrepancies in the estimated time-delays. The bias, variance and the root mean square error (RMSE) of the proposed method is consistently equal or less than the GCC methods. The location of a moving speaker estimated using the time-delays obtained by the proposed method are closer to the actual values, than those obtained by the GCC method. Vikas C. Raykar, Bayya Yegnanarayana, S. R. Mahadeva Prasanna, Ramani Duraiswami |
IEEE Trans. Speech Audio Process. | 3 |
| 2005 | Processing of reverberant speech for time-delay estimationabstractIn this paper, we present a method of extracting the time-delay between speech signals collected at two microphone locations. Time-delay estimation from microphone outputs is the first step for many sound localization algorithms, and also for enhancement of speech. For time-delay estimation, speech signals are normally processed using short-time spectral information (either magnitude or phase or both). The spectral features are affected by degradations in speech caused by noise and reverberation. Features corresponding to the excitation source of the speech production mechanism are robust to such degradations. We show that these source features can be extracted reliably from the speech signal. The time-delay estimate can be obtained using the features extracted even from short segments (50-100 ms) of speech from a pair of microphones. The proposed method for time-delay estimation is found to perform better than the generalized cross-correlation (GCC) approach. A method for enhancement of speech is also proposed using the knowledge of the time-delay and the information of the excitation source. Bayya Yegnanarayana, S. R. Mahadeva Prasanna, Ramani Duraiswami, Dmitry N. Zotkin |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Combining evidence from source, suprasegmental and spectral features for a fixed-text speaker verification systemabstractThis paper proposes a text-dependent (fixed-text) speaker verification system which uses different types of information for making a decision regarding the identity claim of a speaker. The baseline system uses the dynamic time warping (DTW) technique for matching. Detection of the end-points of an utterance is crucial for the performance of the DTW-based template matching. A method based on the vowel onset point (VOP) is proposed for locating the end-points of an utterance. The proposed method for speaker verification uses the suprasegmental and source features, besides spectral features. The suprasegmental features such as pitch and duration are extracted using the warping path information in the DTW algorithm. Features of the excitation source, extracted using the neural network models, are also used in the text-dependent speaker verification system. Although the suprasegmental and source features individually may not yield good performance, combining the evidence from these features seem to improve the performance of the system significantly. Neural network models are used to combine the evidence from multiple sources of information. Bayya Yegnanarayana, S. R. Mahadeva Prasanna, Jinu Mariam Zachariah, Cheedella S. Gupta |
IEEE Trans. Speech Audio Process. | 2 |
| 2004 | Extraction of pitch in adverse conditionsabstractThe paper proposes a method for the extraction of pitch in adverse conditions. The real environment, in which degradation is due to several unpredictable sources, like additive noise, reverberation and channel noise, is treated as an adverse condition. The proposed method is based on knowledge of glottal closure (GC) events. A GC event is the instant at which closure of vocal folds takes place within a pitch period. The Hilbert envelope of the linear prediction (LP) residual gives information about the location of GC events. Autocorrelation analysis is performed on the Hilbert envelope of the LP residual. The properties of the Hilbert envelope of the LP residual are exploited for the extraction of pitch from the autocorrelation sequence. The results of the proposed method are compared with the simple inverse filtering technique (SIFT) algorithm. The performance of the proposed algorithm is found to be superior, even in adverse conditions. S. R. Mahadeva Prasanna, Bayya Yegnanarayana |
ICASSP (1) | 1 |
| 2004 | Two-Stage Duration Model for Indian Languages Using Neural Networks
K. Sreenivasa Rao, S. R. Mahadeva Prasanna, Bayya Yegnanarayana |
ICONIP | 2 |
| 2004 | Enhancement of reverberant speech using excitation source informationabstractThis paper proposes a method for the enhancement of re-verberant speech using the knowledge of the excitation source of speech production. The degradation level in the reverberant speech is measured in terms of Speech-to-Reverberation component Ratio (SRR). From percep-tion and processing point of view high SRR regions are important. Hence the proposed method identifies and en-hances the speech in high SRR regions. The high SRR re-gions are identified using the Hilbert envelope of the Lin-ear Prediction (LP) residual, which contains information about the excitation source of speech production. The Hilbert envelope of the LP residual derived from the re-verberant speech is processed by the covariance analysis to derive the weight function. The LP residual of the re-verberant speech is multiplied with the weight function to enhance the excitations of speech in the high SRR re-gions. The speech signal synthesized from the modified LP residual is found to be less reverberant. 1. M. Chaitanya, S. R. Mahadeva Prasanna, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2003 | Tracking a moving speaker using excitation source informationabstractMicrophone arrays are widely used to detect, locate, and track a stationary or moving speaker. The first step is to estimate the time delay, between the speech signals received by a pair of microphones. Conventional methods like generalized crosscorrelation are based on the spectral content of the vocal tract system in the speech signal. The spectral content of the speech signal is affected due to degradations in the speech signal caused by noise and reverberation. However, features corresponding to the excitation source of speech are less affected by such degradations. This paper proposes a novel method to estimate the time delays using the excitation source information in speech. The estimated delays are used to get the position of the moving speaker. The proposed method is compared with the spectrumbased approach using real data from a microphone array setup. 1. Vikas C. Raykar, Ramani Duraiswami, Bayya Yegnanarayana, S. R. Mahadeva Prasanna |
INTERSPEECH | 4 |
| 2003 | Enhancement of speech in multispeaker environmentabstractIn this paper a method based on the excitation source information is proposed for enhancement of speech, degraded by speech from other speakers. Speech from multiple speakers is simultaneously collected over two spatially distributed microphones. Time-delay of each speaker with respect to the two microphones is estimated using the excitation source information. A weight function is derived for each speaker using the knowledge of the timedelay and the excitation source information. Linear prediction (LP) residuals of the microphone signals are processed separately using the weight functions. Speech signals are synthesized from the modified residuals. One speech signal per speaker is derived from each microphone signal. The synthesized speech signals of each speaker are combined to produce enhanced speech. Significant enhancement of the speech of one speaker relative to other was observed from the combined signal. Bayya Yegnanarayana, S. R. Mahadeva Prasanna, Mathew Magimai-Doss |
INTERSPEECH | 2 |
| 2002 | Linear and nonlinear compression of feature vectors for speech recognitionabstractIn this paper, we consider approaches for linear and nonlinear compression of feature vectors for recognition of utterances of syllable-like units in Indian languages. The distribution capturing ability of an autoassociative neural network model is exploited to derive the components for compressing the feature vectors. The nonlinear compression is accomplished by a five layer autoassociative neural network model. Linear compression is realized by principal component analysis. Both linear and nonlinear compressions are performed on each subgroup of the sound units separately. The results show that it is indeed possible to compress the feature vectors from 50 to 19 dimension without affecting the performance of the classifier. Suryakanth V. Gangashetty, S. R. Mahadeva Prasanna, Bayya Yegnanarayana |
ICASSP | 2 |
| 2002 | Detection of vowel onset point in speechabstractSound units in many languages are syllabic in nature, and frequently used syllables are of consonant-vowel (CV) type. Vowel onset point (VOP) is an important event in CV units. Knowledge of VOPs helps in many applications such as speech recognition, speaker recognition, speech enhancement, begin-end detection, segmentation of speech into vowel/nonvowel-like units and finding duration of vowels. In this paper we describe parameters or features useful for manually identifying the VOPs for different types of CV units. An automatic algorithm is proposed for detecting VOPs in continuous speech, which is motivated by the nature of production and perception of speech. Speech signal is a result of exciting a time varying vocal tract system with time varying excitation. Changes in the source and system characteristics around the VOP are both useful for the detection of VOPs. In this paper we use the changes in the source characteristics for detecting the VOPs. The performance of the proposed algorithm is evaluated using 25 sentences for which a total of 236 VOPs have been identified manually. It is found that 216 VOPs have been detected within a resolution of +/− 30 ms. Compared to the energy-based approach, VOP-based begin-end detection has significantly improved the performance in the case of a text-dependent speaker verification system. For a telephone database of 32 speakers consisting of 480 genuine S. R. Mahadeva Prasanna, Jinu Mariam Zachariah |
ICASSP | 1 |
| 2002 | Speech enhancement using excitation source informationabstractThis paper proposes an approach for processing speech from multiple microphones to enhance speech degraded by noise and reverberation. The approach is based on exploiting the features of the excitation source in speech production. In particular, the characteristics of voiced speech can be used to derive a coherently added signal from the linear prediction (LP) residuals of the degraded speech data from different microphones. A weight function is derived from the coherently added signal. For coherent addition the time-delay between a pair of microphones is estimated using the knowledge of the source information present in the LP residual. The enhanced speech is generated by exciting the time varying all-pole filter with the weighted LP residual. Bayya Yegnanarayana, S. R. Mahadeva Prasanna, K. Sreenivasa Rao |
ICASSP | 2 |