VLDB 2026 Research / reviewers in the wild / expert
Sudarsana Reddy Kadiri
dblp:140/2646 · also Sudarsana Kadiri
· DBLP profile ↗
58ranked-venue papers
20as first author
31since 2021 · last 2025
0000-0001-5806-3053ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 45 · 15 first-author · 22 since 2021Artificial intelligence and machine learning · 33 · 12 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Wavelet Scattering Network Features for Intensity Category Classification and Prediction of SPL from SpeechabstractSpeakers change vocal intensity in daily life to communicate over long distances and to express vocal emotions. Humans produce speech using different intensity categories (e.g. soft, normal and loud voice) and they can regulate intensity across a wide sound pressure level (SPL) range. Knowing the intensity category or the SPL of speech is beneficial in speech-based biomarking of health. Recent studies have explored the vocal intensity category classification and prediction of SPL from speech, which has been recorded without SPL calibration information and is presented on an arbitrary amplitude scale. Using speech signals in such scenario, this study investigates the wavelet scattering network (WSN) features in two tasks: (1) classification of speech into four intensity categories (soft, normal, loud, very loud) (multi-class classification task) and (2) prediction of SPL (regression task). In the former task, the WSN features showed absolute accuracy improvements of 4-14% compared to reference features. For the latter task, the WSN features improved the prediction of SPL by an average of 1-2 dB compared to the reference features. Manila Kodali, Sudarsana Reddy Kadiri, Shri Narayanan, Paavo Alku |
ICASSP | 2 |
| 2025 | Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence PredictionabstractBrain-computer interfaces (BCI) offer numerous human-centered application possibilities, particularly affecting people with neurological disorders. Text or speech decoding from brain activities is a relevant domain that could augment the quality of life for people with impaired speech perception. We propose a novel approach to enhance listened speech decoding from electroencephalography (EEG) signals by utilizing an auxiliary phoneme predictor that simultaneously decodes textual phoneme sequences. The proposed model architecture consists of three main parts: EEG module, speech module, and phoneme predictor. The EEG module learns to properly represent EEG signals into EEG embeddings. The speech module generates speech waveforms from the EEG embeddings. The phoneme predictor outputs the decoded phoneme sequences in text modality. Our proposed approach allows users to obtain decoded listened speech from EEG signals in both modalities (speech waveforms and textual phoneme sequences) simultaneously, eliminating the need for a concatenated sequential pipeline for each modality. The proposed approach also outperforms previous methods in both modalities. The source code and speech samples are publicly available1. Tiantian Feng, Aditya Kommineni, Sudarsana Reddy Kadiri, Shri Narayanan |
ICASSP | 4 |
| 2025 | Can Multimodal Foundation Models Help Analyze Child-Inclusive Autism Diagnostic Videos?
Aditya Kommineni, Digbalay Bose, Tiantian Feng, So Hyun Kim, Helen Tager-Flusberg, Somer Bishop, Catherine Lord, Sudarsana Reddy Kadiri, Shri Narayanan |
INTERSPEECH | 8 |
| 2025 | Zero-shot KWS for children's speech using layer-wise features from SSL models
Subham Kutum, Abhijit Sinha, Hemant Kumar Kathania, Sudarsana Reddy Kadiri, Mahesh Chandra Govil |
Pattern Recognit. Lett. | 4 |
| 2025 | Automatic classification of vocal intensity categories from amplitude-normalized speech signals by comparing acoustic features and classifier modelsabstractRegulation of vocal intensity is a fundamental phenomenon in speech communication. Speakers use different intensity categories (e.g., soft, normal, and loud voice) to generate different vocal emotions or to communicate in noisy conditions or over varying distances. Vocal intensity categories have been studied in fundamental research of speech, but much less is known about their automatic classification. This study investigates the classification of vocal intensity categories from speech signals in a scenario, where the original level information of speech is absent and the signal is presented on a normalized amplitude scale. Different acoustic features were studied together with machine learning (ML) and deep learning (DL) classifiers using two different labeling approaches. Speech signals recorded from 50 speakers reciting sentences in four intensity categories (soft, normal, loud, and very loud) were analyzed. Altogether 15 feature sets including different cepstral, spectral and handcrafted (eGeMAPS) features were compared. Three ML classifiers (support vector machine, random forest and AdaBoost), and four DL classifiers (deep neural network, convolutional neural network, recurrent neural network and bidirectional long short-term memory network) were compared. The best classification accuracy of 86.0% was obtained by combining the best performing cepstral and spectral features and using the bidirectional long short-term memory classifier. • Multi-class classification of vocal intensity categories is studied. • The classification is studied using amplitude-normalized speech signals. • Various labelling approaches, features and classifiers are compared. • DL models outperformed ML models, with BiLSTM achieving the best performance. Manila Kodali, Luna Ansari, Sudarsana Reddy Kadiri, Shri Narayanan, Paavo Alku |
Speech Commun. | 3 |
| 2025 | Do all features matter? Layer-wise feature probing of self-supervised speech models for dysarthria severity classification
Paban Sapkota, Harsh Srivastava, Hemant Kumar Kathania, Shri Narayanan, Sudarsana Reddy Kadiri |
Speech Commun. | 5 |
| 2025 | Towards robust heart failure detection in digital telephony environments by utilizing transformer-based codec inversionabstractThis study introduces the Codec Transformer Network (CTN) to enhance the reliability of automatic heart failure (HF) detection from coded telephone speech by addressing codec-related challenges in digital telephony. The study specifically addresses the codec mismatch between training and inference in HF detection. CTN is designed to map the mel-spectrogram representations of encoded speech signals back to their original, non-encoded forms, thereby recovering HF-related discriminative information. The effectiveness of CTN is demonstrated in conjunction with three HF detectors, based on Support Vector Machine, Random Forest, and K-Nearest Neighbors classifiers. The results show that CTN effectively retrieves the discriminative information between patients and controls, and performs comparably to or better than a baseline approach, based on multi-condition training. Saska Tirronen, Farhad Javanmardi, Hilla Pohjalainen, Sudarsana Reddy Kadiri, Mittapalle Kiran Reddy, Pyry Helkkula, Kasimir Kaitue, Mikko Minkkinen, Heli Tolppanen, Tuomo Nieminen, Paavo Alku |
Speech Commun. | 4 |
| 2025 | Can Layer-Wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?abstractAutomatic Speech Recognition (ASR) systems often struggle to accurately process children's speech due to its distinct and highly variable acoustic and linguistic characteristics. While recent advancements in self-supervised learning (SSL) models have greatly enhanced the transcription of adult speech, accurately transcribing children's speech remains a significant challenge. This study investigates the effectiveness of layer-wise features extracted from state-of-the-art SSL pre-trained models - specifically, Wav2Vec2, HuBERT, Data2Vec, and WavLM in improving the performance of ASR for children's speech in zero-shot scenarios. A detailed analysis of features extracted from these models was conducted, integrating them into a simplified DNN-based ASR system using the Kaldi toolkit. The analysis identified the most effective layers for enhancing ASR performance on children's speech in a zero-shot scenario, where WSJCAM0 adult speech was used for training and PFSTAR children speech for testing. Experimental results indicated that Layer 22 of the Wav2Vec2 model achieved the lowest Word Error Rate (WER) of 5.15%, representing a 51.64% relative improvement over the direct zero-shot decoding using Wav2Vec2 (WER of 10.65%). Additionally, age group-wise analysis demonstrated consistent performance improvements with increasing age, along with significant gains observed even in younger age groups using the SSL features. Further experiments on the CMU Kids dataset confirmed similar trends, highlighting the generalizability of the proposed approach. Abhijit Sinha, Hemant Kumar Kathania, Sudarsana Reddy Kadiri, Shri Narayanan |
IEEE Signal Process. Lett. | 3 |
| 2024 | Fine-tuning of Pre-trained Models for Classification of Vocal Intensity Category from Speech SignalsabstractSpeakers regulate vocal intensity on many occasions for example to be heard over a long distance or to express vocal emotions. Humans can regulate vocal intensity over a wide sound pressure level (SPL) range and therefore speech can be categorized into different vocal intensity categories. Recent machine learning experiments have studied classification of vocal intensity category from speech signals which have been recorded without SPL information and which are represented on arbitrary amplitude scales. By fine-tuning four pre-trained models (wav2vec2-BASE, wav2vec2-LARGE, HuBERT, audio speech transformers), this paper studies classification of speech into four intensity categories (soft, normal, loud, very loud), when speech is presented on such arbitrary amplitude scale. The fine-tuned model embeddings showed absolute improvements of 5% and 10-12% in accuracy compared to baselines for the target intensity category label and the SPL-based intensity category label, respectively. Manila Kodali, Sudarsana Reddy Kadiri, Paavo Alku |
INTERSPEECH | 2 |
| 2024 | Toward Fully-End-to-End Listened Speech Decoding from EEG SignalsabstractSpeech decoding from EEG signals is a challenging task, where brain activity is modeled to estimate salient characteristics of acoustic stimuli.We propose FESDE, a novel framework for Fully-End-to-end Speech Decoding from EEG signals.Our approach aims to directly reconstruct listened speech waveforms given EEG signals, where no intermediate acoustic feature processing step is required.The proposed method consists of an EEG module and a speech module along with a connector.The EEG module learns to better represent EEG signals, while the speech module generates speech waveforms from model representations.The connector learns to bridge the distributions of the latent spaces of EEG and speech.The proposed framework is both simple and efficient, by allowing single-step inference, and outperforms prior works on objective metrics.A fine-grained phoneme analysis is conducted to unveil model characteristics of speech decoding.The source code is available here: github.com/lee-jhwn/fesde. Aditya Kommineni, Tiantian Feng, Kleanthis Avramidis, Xuan Shi, Sudarsana Reddy Kadiri, Shri Narayanan |
INTERSPEECH | 6 |
| 2024 | MMSD-Net: Towards Multi-modal Stuttering Detection
Liangyu Nie, Sudarsana Reddy Kadiri, Ruchit Agrawal |
INTERSPEECH | 2 |
| 2024 | A comparison of data augmentation methods in voice pathology detectionabstractTo distinguish pathological voices from healthy voices, automatic voice pathology detection systems can be built using machine learning (ML) and deep learning (DL) techniques. To fully exploit such systems, large quantities of training data are typically required. The amount of training data is, however, small in the area of pathological voice, and therefore data augmentation (DA) becomes a potential technology to artificially increase the quantity of training data. This study presents a systematic comparison between various DA methods in the detection of pathological voice, including three time domain methods (noise addition, pitch shifting and time stretching), one time-frequency domain method (SpecAugment), and two vocoder-based methods (harmonic-to-noise ratio (HNR) modification and glottal pulse length modification). Detection systems were built using four popular spectral feature representations (static mel-frequency cepstral coefficients (MFCCs), dynamic MFCCs, spectrogram and mel-spectrogram). As classifiers, two widely used ML models (support vector machine (SVM) and random forest (RF)) and two DL models (long short-term memory (LSTM) network and convolutional neural network (CNN) with 1-dimensional (1-D) and 2-dimensional (2-D) architectures) were used. These systems were trained using a small number of training samples from two popular databases of pathological voice (HUPA and SVD) to find the best feature/classifier combination for each database. As a result, one ML-based detection system (mel-spectrogram/SVM for HUPA and SVD) and two DL-based detection systems (dynamic MFCCs/2-D CNN for HUPA and mel-spectrogram/2-D CNN for SVD) were selected for the comparison of the DA methods. The results show that by using DA in the system training, detection accuracy increased compared to the baseline systems that were trained without using DA. This improvement in accuracy was, however, clearly larger for the 2D-CNN system than for the SVM system. Furthermore, all six DA methods improved accuracy of the 2-D CNN system compared to the baseline system for both databases. The highest improvements were achieved using the time-frequency domain SpecAugment DA method, which improved accuracy by 1.5% and 3.8% (absolute) for the HUPA and SVD database, respectively. Farhad Javanmardi, Sudarsana Reddy Kadiri, Paavo Alku |
Comput. Speech Lang. | 2 |
| 2024 | Investigation of self-supervised pre-trained models for classification of voice quality from speech and neck surface accelerometer signalsabstractPrior studies in the automatic classification of voice quality have mainly studied the use of the acoustic speech signal as input. Recently, a few studies have been carried out by jointly using both speech and neck surface accelerometer (NSA) signals as inputs, and by extracting mel-frequency cepstral coefficients (MFCCs) and glottal source features. This study examines simultaneously-recorded speech and NSA signals in the classification of voice quality (breathy, modal, and pressed) using features derived from three self-supervised pre-trained models (wav2vec2-BASE, wav2vec2-LARGE, and HuBERT) and using a support vector machine (SVM) as well as convolutional neural networks (CNNs) as classifiers. Furthermore, the effectiveness of the pre-trained models is compared in feature extraction between glottal source waveforms and raw signal waveforms for both speech and NSA inputs. Using two signal processing methods (quasi-closed phase (QCP) glottal inverse filtering and zero frequency filtering (ZFF)), glottal source waveforms are estimated from both speech and NSA signals. The study has three main goals: (1) to study whether features derived from pre-trained models improve classification accuracy compared to conventional features (spectrogram, mel-spectrogram, MFCCs, i-vector, and x-vector), (2) to investigate which of the two modalities (speech vs. NSA) is more effective as input in the classification task with pre-trained model-based features, and (3) to evaluate whether the deep learning-based CNN classifier can enhance the classification accuracy in comparison to the SVM classifier. The results revealed that the use of the NSA input showed better classification performance compared to the speech signal. Between the features, the pre-trained model-based features showed better classification accuracies, both for speech and NSA inputs compared to the conventional features. The two classifiers performed equally well for all the pre-trained model-based features for both speech and NSA signals. It was also found that the HuBERT features performed better than the wav2vec2-BASE and wav2vec2-LARGE features for both speech and NSA inputs. In particular, when compared to the conventional features, the HuBERT features showed an absolute accuracy improvement of 3%–6% for speech and NSA signals in the classification of voice quality. Sudarsana Reddy Kadiri, Farhad Javanmardi, Paavo Alku |
Comput. Speech Lang. | 1 |
| 2024 | Automatic classification of the severity level of Parkinson's disease: A comparison of speaking tasks, features, and classifiersabstractAutomatic speech-based severity level classification of Parkinson’s disease (PD) enables objective assessment and earlier diagnosis. While many studies have been conducted on the binary classification task to distinguish speakers in PD from healthy controls (HCs), clearly fewer studies have addressed multi-class PD severity level classification problems. Furthermore, in studying the three main issues of speech-based classification systems—speaking tasks, features, and classifiers—previous investigations on the severity level classification have yielded inconclusive results due to the use of only a few, and sometimes just one, type of speaking task, feature, or classifier in each study. Hence, a systematic comparison is conducted in this study between different speaking tasks, features, and classifiers. Five speaking tasks (vowel task, sentence task, diadochokinetic (DDK) task, read text task, and monologue task), four features (phonation, articulation, prosody, and their fusion), and four classifier architectures (support vector machine (SVM), random forest (RF), multilayer perceptron (MLP), and AdaBoost) were compared. The classification task studied was a 3-class problem to classify PD severity level as healthy vs. mild vs. severe. Two MDS-UPDRS scales (MDS-UPDRS-III and MDS-UPDRS-S) were used for the ground truth severity level labels. The results showed that the use of the monologue task and the articulation and fusion of features improved classification accuracy significantly compared to the use of the other speaking tasks and features. The best classification systems resulted in a rate of accuracy of 58% (using the monologue task with the articulation features) for the MDS-UPDR-III scale and 56% (using the monologue task with fusion of features) for the MDS-UPDRS-S scale. Manila Kodali, Sudarsana Reddy Kadiri, Paavo Alku |
Comput. Speech Lang. | 2 |
| 2024 | Spectral warping based data augmentation for low resource children's speaker verificationabstractAbstract In this paper, we present our effort to develop an automatic speaker verification (ASV) system for low resources children’s data. For the children’s speakers, very limited amount of speech data is available in majority of the languages for training the ASV system. Developing an ASV system under low resource conditions is a very challenging problem. To develop the robust baseline system, we merged out of domain adults’ data with children’s data to train the ASV system and tested with children’s speech. This kind of system leads to acoustic mismatches between training and testing data. To overcome this issue, we have proposed spectral warping based data augmentation. We modified adult speech data using spectral warping method (to simulate like children’s speech) and added it to the training data to overcome data scarcity and mismatch between adults’ and children’s speech. The proposed data augmentation gives 20.46% and 52.52% relative improvement (in equal error rate) for Indian Punjabi and British English speech databases, respectively. We compared our proposed method with well known data augmentation methods: SpecAugment, speed perturbation (SP) and vocal tract length perturbation (VTLP), and found that the proposed method performed best. The proposed spectral warping method is publicly available at https://github.com/kathania/Speaker-Verification-spectral-warping . Hemant Kumar Kathania, Virender Kadyan, Sudarsana Reddy Kadiri, Mikko Kurimo |
Multim. Tools Appl. | 3 |
| 2024 | AVID: A speech database for machine learning studies on vocal intensityabstractVocal intensity, which is quantified typically with the sound pressure level (SPL), is a key feature of speech. To measure SPL from speech recordings, a standard calibration tone (with a reference SPL of 94 dB or 114 dB) needs to be recorded together with speech. However, most of the popular databases that are used in areas such as speech and speaker recognition have been recorded without calibration information by expressing speech on arbitrary amplitude scales. Therefore, information about vocal intensity of the recorded speech, including SPL, is lost. In the current study, we introduce a new open and calibrated speech/electroglottography (EGG) database named Aalto Vocal Intensity Database (AVID). AVID includes speech and EGG produced by 50 speakers (25 males, 25 females) who varied their vocal intensity in four categories (soft, normal, loud and very loud). Recordings were conducted using a constant mouth-to-microphone distance and by recording a calibration tone. The speech data was labelled sentence-wise using a total of 19 labels that support the utilisation of the data in machine learing (ML) -based studies of vocal intensity based on supervised learning. In order to demonstrate how the AVID data can be used to study vocal intensity, we investigated one multi-class classification task (classification of speech into soft, normal, loud and very loud intensity classes) and one regression task (prediction of SPL of speech). In both tasks, we deliberately warped the level of the input speech by normalising the signal to have its maximum amplitude equal to 1.0, that is, we simulated a scenario that is prevalent in current speech databases. The results show that using the spectrogram feature with the support vector machine classifier gave an accuracy of 82% in the multi-class classification of the vocal intensity category. In the prediction of SPL, using the spectrogram feature with the support vector regressor gave an mean absolute error of about 2 dB and a coefficient of determination of 92%. We welcome researchers interested in classification and regression problems to utilise AVID in the study of vocal intensity, and we hope that the current results could serve as baselines for future ML studies on the topic. Paavo Alku, Manila Kodali, Laura Laaksonen, Sudarsana Reddy Kadiri |
Speech Commun. | 4 |
| 2024 | Pre-trained models for detection and severity level classification of dysarthria from speechabstractAutomatic detection and severity level classification of dysarthria from speech enables non-invasive and effective diagnosis that helps clinical decisions about medication and therapy of patients. In this work, three pre-trained models (wav2vec2-BASE, wav2vec2-LARGE, and HuBERT) are studied to extract features to build automatic detection and severity level classification systems for dysarthric speech. The experiments were conducted using two publicly available databases (UA-Speech and TORGO). One machine learning-based model (support vector machine, SVM) and one deep learning-based model (convolutional neural network, CNN) was used as the classifier. In order to compare the performance of the wav2vec2-BASE, wav2vec2-LARGE, and HuBERT features, three popular acoustic feature sets, namely, mel-frequency cepstral coefficients (MFCCs), openSMILE and extended Geneva minimalistic acoustic parameter set (eGeMAPS) were considered. Experimental results revealed that the features derived from the pre-trained models outperformed the three baseline features. It was also found that the HuBERT features performed better than the wav2vec2-BASE and wav2vec2-LARGE features. In particular, when compared to the best-performing baseline feature (openSMILE), the HuBERT features showed in the detection problem absolute accuracy improvements that varied between 1.33% (the SVM classifier, the TORGO database) and 2.86% (the SVM classifier, the UA-Speech database). In the severity level classification problem, the HuBERT features showed absolute accuracy improvements that varied between 6.54% (the SVM classifier, the TORGO database) and 10.46% (the SVM classifier, the UA-Speech database) compared to the best-performing baseline feature (eGeMAPS). Farhad Javanmardi, Sudarsana Reddy Kadiri, Paavo Alku |
Speech Commun. | 2 |
| 2024 | Exploring the Impact of Fine-Tuning the Wav2vec2 Model in Database-Independent Detection of Dysarthric SpeechabstractMany acoustic features and machine learning models have been studied to build automatic detection systems to distinguish dysarthric speech from healthy speech. These systems can help to improve the reliability of diagnosis. However, speech recorded for diagnosis in real-life clinical conditions can differ from the training data of the detection system in terms of, for example, recording conditions, speaker identity, and language. These mismatches may lead to a reduction in detection performance in practical applications. In this study, we investigate the use of the wav2vec2 model as a feature extractor together with a support vector machine (SVM) classifier to build automatic detection systems for dysarthric speech. The performance of the wav2vec2 features is evaluated in two cross-database scenarios, language-dependent and language-independent, to study their generalizability to unseen speakers, recording conditions, and languages before and after fine-tuning the wav2vec2 model. The results revealed that the fine-tuned wav2vec2 features showed better generalization in both scenarios and gave an absolute accuracy improvement of 1.46%-8.65% compared to the non-fine-tuned wav2vec2 features. Farhad Javanmardi, Sudarsana Reddy Kadiri, Paavo Alku |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Wav2vec-Based Detection and Severity Level Classification of Dysarthria From SpeechabstractAutomatic detection and severity level classification of dysarthria directly from acoustic speech signals can be used as a tool in medical diagnosis. In this work, the pre-trained wav2vec 2.0 model is studied as a feature extractor to build detection and severity level classification systems for dysarthric speech. The experiments were carried out with the popularly used UA-speech database. In the detection experiments, the results revealed that the best performance was obtained using the embeddings from the first layer of the wav2vec model that yielded an absolute improvement of 1.23% in accuracy compared to the best performing baseline feature (spectrogram). In the studied severity level classification task, the results revealed that the embeddings from the final layer gave an absolute improvement of 10.62% in accuracy compared to the best baseline features (mel-frequency cepstral coefficients). Farhad Javanmardi, Saska Tirronen, Manila Kodali, Sudarsana Reddy Kadiri, Paavo Alku |
ICASSP | 4 |
| 2023 | Automatic Classification of Vocal Intensity Category from SpeechabstractRegulation of vocal intensity is a fundamental phenomenon in speech communication. Vocal intensity can be quantified using sound pressure level (SPL), which can be measured easily by recording a standard calibration signal with speech and by comparing the energy of the recorded speech signal with that of the calibration tone. Unfortunately, speech recordings are mostly conducted without the SPL calibration signal, and speech signals are saved to databases using arbitrary amplitude scales. Therefore, neither the SPL nor the intensity category (e.g. soft or loud phonation) of a saved speech signal can be determined afterwards. Even though the original level information of speech is lost when the signal is presented on arbitrary amplitude scales, the speech signal contains other acoustic cues of vocal intensity. In the current study, we study machine learning and deep learning -based methods in automatic classification of vocal intensity category when the input speech is expressed using an arbitrary amplitude scale. A new gender-balanced database consisting of speech produced in four vocal intensity categories (soft, normal, loud, and very loud) was first recorded. Support vector machine and deep neural network (DNN) models were used to develop automatic classification systems using spectrograms, mel-spectrograms, and mel-frequency cepstral coefficients as features. The DNN classifier using the mel-spectrogram showed the best classification accuracy of about 90%. The database is made publicly available at https://bit.ly/3tLPGRx. Manila Kodali, Sudarsana Reddy Kadiri, Laura Laaksonen, Paavo Alku |
ICASSP | 2 |
| 2023 | Utilizing Wav2Vec In Database-Independent Voice Disorder DetectionabstractAutomatic detection of voice disorders from acoustic speech signals can help to improve reliability of medical diagnosis. However, the real-life environment in which speech signals are recorded for diagnosis can be different from the environment in which the detection system’s training data was originally collected. This mismatch between the recording conditions can decrease detection performance in practical scenarios. In this work, we propose to use a pre-trained wav2vec 2.0 model as a feature extractor to build automatic detection systems for voice disorders. The embeddings from the first layers of the context network contain information about phones, and these features are useful in voice disorder detection. We evaluate the performance of the wav2vec features in single-database and crossdatabase scenarios to study their generalizability to unseen speakers and recording conditions. The results indicate that the wav2vec features generalize better than popular spectral and cepstral baseline features. Saska Tirronen, Farhad Javanmardi, Manila Kodali, Sudarsana Reddy Kadiri, Paavo Alku |
ICASSP | 4 |
| 2023 | Severity Classification of Parkinson's Disease from Speech using Single Frequency Filtering-based FeaturesabstractDeveloping objective methods for assessing the severity of Parkinson's disease (PD) is crucial for improving the diagnosis and treatment. This study proposes two sets of novel features derived from the single frequency filtering (SFF) method: (1) SFF cepstral coefficients (SFFCC) and (2) MFCCs from the SFF (MFCC-SFF) for the severity classification of PD. Prior studies have demonstrated that SFF offers greater spectro-temporal resolution compared to the short-time Fourier transform. The study uses the PC-GITA database, which includes speech of PD patients and healthy controls produced in three speaking tasks (vowels, sentences, text reading). Experiments using the SVM classifier revealed that the proposed features outperformed the conventional MFCCs in all three speaking tasks. The proposed SFFCC and MFCC-SFF features gave a relative improvement of 5.8% and 2.3% for the vowel task, 7.0% & 1.8% for the sentence task, and 2.4% and 1.1% for the read text task, in comparison to MFCC features. Sudarsana Reddy Kadiri, Manila Kodali, Paavo Alku |
INTERSPEECH | 1 |
| 2023 | Classification of Vocal Intensity Category from Speech using the Wav2vec2 and Whisper EmbeddingsabstractIn speech communication, talkers regulate vocal intensity resulting in speech signals of different intensity categories (e.g., soft, loud). Intensity category carries important information about the speaker's health and emotions. However, many speech databases lack calibration information, and therefore sound pressure level cannot be measured from the recorded data. Machine learning, however, can be used in intensity category classification even though calibration information is not available. This study investigates pre-trained model embeddings (Wav2vec2 and Whisper) in classification of vocal intensity category (soft, normal, loud, and very loud) from speech signals expressed using arbitrary amplitude scales. We use a new database consisting of two speaking tasks (sentence and paragraph). Support vector machine is used as a classifier. Our results show that the pre-trained model embeddings outperformed three baseline features, providing improvements of up to 7%(absolute) in accuracy. Manila Kodali, Sudarsana Reddy Kadiri, Paavo Alku |
INTERSPEECH | 2 |
| 2023 | Refining a deep learning-based formant tracker using linear prediction methodsabstractIn this study, formant tracking is investigated by refining the formants tracked by an existing data-driven tracker, DeepFormants, using the formants estimated in a model-driven manner by linear prediction (LP)-based methods. As LP-based formant estimation methods, conventional covariance analysis (LP-COV) and the recently proposed quasi-closed phase forward–backward (QCP-FB) analysis are used. In the proposed refinement approach, the contours of the three lowest formants are first predicted by the data-driven DeepFormants tracker, and the predicted formants are replaced frame-wise with local spectral peaks shown by the model-driven LP-based methods. The refinement procedure can be plugged into the DeepFormants tracker with no need for any new data learning. Two refined DeepFormants trackers were compared with the original DeepFormants and with five known traditional trackers using the popular vocal tract resonance (VTR) corpus. The results indicated that the data-driven DeepFormants trackers outperformed the conventional trackers and that the best performance was obtained by refining the formants predicted by DeepFormants using QCP-FB analysis. In addition, by tracking formants using VTR speech that was corrupted by additive noise, the study showed that the refined DeepFormants trackers were more resilient to noise than the reference trackers. In general, these results suggest that LP-based model-driven approaches, which have traditionally been used in formant estimation, can be combined with a modern data-driven tracker easily with no further training to improve the tracker’s performance. Paavo Alku, Sudarsana Reddy Kadiri, Dhananjaya Gowda |
Comput. Speech Lang. | 2 |
| 2023 | Analysis of Instantaneous Frequency Components of Speech Signals for Epoch ExtractionabstractThe major impulse-like excitation in the speech signal is due to abrupt closure of the vocal folds, which takes place at the glottal closure instant (GCI) or epoch in each cycle. GCIs are used in many areas of speech science and technology, such as in prosody modification, voice source analysis, formant extraction and speech synthesis. It is difficult to observe these discontinuities (corresponding to GCIs) in the speech signal because of the superimposed time-varying response of the vocal tract system. This paper examines the phase part of different frequency components of the speech signal to extract epochs. Three analysis methods to decompose the speech signal into different frequency components are considered. These methods are the short-time Fourier transform (STFT), narrow bandpass filtering (NBPF), and single frequency filtering (SFF). The locations of the discontinuities in the speech signal are obtained from the instantaneous frequency (IF) (i.e., the time derivative of the phase) of each of the frequency components. A method for automatic detection of epochs using the amplitude weighted IF is proposed. Performance of the proposed epoch detection method is compared with four state-of-the-art methods in clean and telephone quality speech. The performance of the proposed method is comparable with the performance of the existing epoch detection methods for clean speech but better for telephone quality speech. Sudarsana Reddy Kadiri, Paavo Alku, Bayya Yegnanarayana |
Comput. Speech Lang. | 1 |
| 2022 | Comparing 1-dimensional and 2-dimensional spectral feature representations in voice pathology detection using machine learning and deep learning classifiersabstractThis work was supported by the Academy of Finland (grant number 313390). The computational resources were provided by Aalto ScienceIT. Farhad Javanmardi, Sudarsana Reddy Kadiri, Manila Kodali, Paavo Alku |
INTERSPEECH | 2 |
| 2022 | Convolutional Neural Networks for Classification of Voice Qualities from Speech and Neck Surface Accelerometer SignalsabstractThis work was supported by the Academy of Finland (grant number 313390). The computational resources were provided by Aalto ScienceIT. Sudarsana Reddy Kadiri, Farhad Javanmardi, Paavo Alku |
INTERSPEECH | 1 |
| 2022 | Wav2vec2-based Paralinguistic Systems to Recognise Vocalised Emotions and StutteringabstractWith the rapid advancement in automatic speech recognition and natural language understanding, a complementary field (paralinguistics) emerged, focusing on the non-verbal content of speech. The ACM Multimedia 2022 Computational Paralinguistics Challenge introduced several exciting tasks of this field. In this work, we focus on tackling two Sub-Challenges using modern, pre-trained models called wav2vec2. Our experimental results demonstrated that wav2vec2 is an excellent tool for detecting the emotions behind vocalisations and recognising different types of stutterings. Albeit they achieve outstanding results on their own, our results demonstrated that wav2vec2-based systems could be further improved by ensembling them with other models. Our best systems outperformed the competition baselines by a considerable margin, achieving an unweighted average recall of 44.0 (absolute improvement of 6.6% over baseline) on the Vocalisation Sub-Challenge and 62.1 (absolute improvement of 21.7% over baseline) on the Stuttering Sub-Challenge. Tamás Grósz, Dejan Porjazovski, Yaroslav Getman, Sudarsana Reddy Kadiri, Mikko Kurimo |
ACM Multimedia | 4 |
| 2022 | A formant modification method for improved ASR of children's speechabstractDifferences in acoustic characteristics between children’s and adults’ speech degrade performance of automatic speech recognition systems when systems trained using adults’ speech are used to recognize children’s speech. This performance degradation is due to the acoustic mismatch between training and testing. One of the main sources of the acoustic mismatch is the difference in vocal tract resonances (formant frequencies) between adult and child speakers. The present study aims to reduce the mismatch in formant frequencies by modifying formants of children’s speech to better correspond to formants of adults’ speech. This is carried out by warping the linear prediction (LP) spectrum computed from children’s speech. The warped LP spectra computed in a frame-based manner from children’s speech are used with the corresponding LP residuals to synthesize speech whose formant structure is closer to that of adults’ speech. When used in testing of an ASR system trained using adults’ speech, the warping reduces the spectral mismatch in speech between training and testing and improves the system performance in recognition of children’s speech. Experiments were conducted using narrowband (8 kHz) and wideband (16 kHz) speech of adult and child speakers from the WSJCAM0 and PF_STAR databases, respectively, and by recognizing children’s speech using acoustic models trained with adults’ speech. The proposed method gave relative improvements of 24% and 11% for the DNN and TDNN acoustic models, respectively, for narrowband speech. For wideband speech, the technique gave relative improvements of 27% and 13% for the DNN and TDNN acoustic models, respectively. The performance of the proposed method was also compared to two speaker adaptation methods: vocal tract length normalization (VTLN) and speaking rate adaptation (SRA). This comparison showed the best recognition performance for the proposed method. We also combined the proposed method with VTLN and SRA, and found that the combined method gave a further reduction in WER. Moreover, our experiments carried out for noisy speech using various types of additive noise and signal-to-noise ratios showed that the proposed method performs well also for degraded speech. Hemant Kumar Kathania, Sudarsana Reddy Kadiri, Paavo Alku, Mikko Kurimo |
Speech Commun. | 2 |
| 2021 | Glottal features for classification of phonation type from speech and neck surface accelerometer signalsabstractGlottal source characteristics vary between phonation types due to the tension of laryngeal muscles with the respiratory effort. Previous studies in the classification of phonation type have mainly used speech signals recorded by microphone. Recently, two studies were published in the classification of phonation type using neck surface accelerometer (NSA) signals. However, there are no previous studies comparing the use of the acoustic speech signal vs. the NSA signal as input in classifying phonation type. Therefore, the current study investigates simultaneously recorded speech and NSA signals in the classification of three phonation types (breathy, modal, pressed). The general goal is to understand which of the two signals (speech vs. NSA) is more effective in the classification task. We hypothesize that by using the same feature set for both signals, classification accuracy is higher for the NSA signal, which is more closely related to the physical vibration of the vocal folds and less affected by the vocal tract compared to the acoustical speech signal. Glottal source waveforms were computed using two signal processing methods, quasi-closed phase (QCP) glottal inverse filtering and zero frequency filtering (ZFF), and a group of time-domain and frequency-domain scalar features were computed from the obtained waveforms. In addition, the study investigated the use of mel-frequency cepstral coefficients (MFCCs) derived from the glottal source waveforms computed by QCP and ZFF. Classification experiments with support vector machine classifiers revealed that the NSA signal showed better discrimination of the phonation types compared to the speech signal when the same feature set was used. Furthermore, it was observed that the glottal features showed complementary information with the conventional MFCC features resulting in the best classification accuracy both for the NSA signal (86.9%) and the speech signal (80.6%). Sudarsana Reddy Kadiri, Paavo Alku |
Comput. Speech Lang. | 1 |
| 2021 | Extraction and Utilization of Excitation Information of Speech: A ReviewabstractSpeech production can be regarded as a process where a time-varying vocal tract system (filter) is excited by a time-varying excitation. In addition to its linguistic message, the speech signal also carries information about, for example, the gender and age of the speaker. Moreover, the speech signal includes acoustical cues about several speaker traits, such as the emotional state and the state of health of the speaker. In order to understand the production of these acoustical cues by the human speech production mechanism and utilize this information in speech technology, it is necessary to extract features describing both the excitation and the filter of the human speech production mechanism. While the methods to estimate and parameterize the vocal tract system are well established, the excitation appears less studied. This article provides a review of signal processing approaches used for the extraction of excitation information from speech. This article highlights the importance of excitation information in the analysis and classification of phonation type and vocal emotions, in the analysis of nonverbal laughter sounds, and in studying pathological voices. Furthermore, recent developments of deep learning techniques in the context of extraction and utilization of the excitation information are discussed. Sudarsana Reddy Kadiri, Paavo Alku, Bayya Yegnanarayana |
Proc. IEEE | 1 |
| 2020 | Comparison of Glottal Closure Instants Detection Algorithms for Emotional SpeechabstractIn production of voiced speech, epochs or glottal closure instants (GCIs) refer to the instants of significant excitation of the vocal tract. Extraction of GCIs is used as a pre-processing stage in many areas of speech technology, such as in prosody modification, speech synthesis and voice source analysis. In the past decades, several GCI detection algorithms have been developed and most of them provide excellent results for speech signals produced using modal (normal) type of phonation. There are, however, no studies comparing multiple state-of-the-art GCI detection methods in emotional speech. In this paper, we compare six GCI detection algorithms using emotional speech and known evaluation metrics. We use the Berlin EMO-DB acted emotional speech database which contains seven emotions and simultaneous electroglottography (EGG) recordings as ground truth. The results show that all six GCI detection algorithms give best performance in processing speech of neutral emotion and that the performance degrade particularly in emotions of high arousal (anger and joy). To improve the performance of GCI detection in emotional speech, the study underlines the importance of local average pitch period estimates. Sudarsana Reddy Kadiri, Paavo Alku, Bayya Yegnanarayana |
ICASSP | 1 |
| 2020 | Study of Formant Modification for Children ASRabstractThe performance of automatic speech recognition systems for children’s speech is known to suffer from the large variation and mismatch in the acoustic and linguistic attributes between children’s and adults’ speech. One of the various identified sources of mismatch is the difference in formant frequencies between adults and children. In this paper, we propose a formant modification method to mitigate differences between adults’ and children’s speech and to improve the performance of ASR for children. The explored technique gives a relative 27% improvement in system performance compared to a hybrid DNN-HMM baseline. We also compare the system performance with related speaker adaptation methods like vocal tract length normalization (VTLN) and speaking rate adaptation (SRA) and find that the proposed method gives improvements over them, as well. Combining the proposed method with VTLN and SRA results in a further reduction of WER. We also found that the proposed method performs well even for noisy speech. Hemant Kumar Kathania, Sudarsana Reddy Kadiri, Paavo Alku, Mikko Kurimo |
ICASSP | 2 |
| 2020 | Learning Filterbanks from Raw Waveform for Accent ClassificationabstractMost of the applications in speech use mel-frequency spectral coefficients (MFSC) as features as they match the human perceptual mechanism, where the emphasis is given to vocal tract characteristics. But in accent classification, mel-scale distribution of filters may not always be the best representations, e.g., pitch accented languages where the emphasis should be on vocal source information too. Motivated by this, we use end-to-end classification of accents directly from waveforms which will reduce the effort of designing features specific to each corpus. The convolution neural network (CNN) model architecture is designed in such a way that the initial layers exhibit similar operation as in MFSC by initializing the weights using time approximate of MFSC. The entire network along with initial layers is trained to learn accent classification. We observed that learning directly from waveform improved the performance of accent classification when compared to CNN trained on hand-engineered features by 10.94% UAR on the test dataset of common voice corpus. Analyzing the filters after learning, we observed changes in distribution and bandwidths of center frequencies. We further observed the importance of appropriately initializing CNN filters. Rashmi Kethireddy, Sudarsana Reddy Kadiri, Suryakanth V. Gangashetty |
IJCNN | 2 |
| 2020 | Parkinson's Disease Detection from Speech Using Single Frequency Filtering Cepstral CoefficientsabstractParkinson's disease (PD) is a progressive deterioration of the human central nervous system. Detection of PD (discriminating patients with PD from healthy subjects) from speech is a useful approach due to its non-invasive nature. This study proposes to use novel cepstral coefficients derived from the single frequency filtering (SFF) method, called as single frequency filtering cepstral coefficients (SFFCCs) for the detection of PD. SFF has been shown to provide higher spectro-temporal resolution compared to the short-time Fourier transform. The current study uses the PC-GITA database, which consists of speech from speakers with PD and healthy controls (50 males, 50 females). Our proposed detection system is based on the i-vectors derived from SFFCCs using SVM as a classifier. In the detection of PD, better performance was achieved when the i-vectors were computed from the proposed SFFCCs compared to the popular conventional MFCCs. Furthermore, we investigated the effect of temporal variations by deriving the shifted delta cepstral (SDC) coefficients using SFFCCs. These experiments revealed that the i-vectors derived from the proposed SFFCCs+SDC features gave an absolute improvement of 9% compared to the i-vectors derived from the baseline MFCCs+SDC features, indicating the importance of temporal variations in the detection of PD. Sudarsana Reddy Kadiri, Rashmi Kethireddy, Paavo Alku |
INTERSPEECH | 1 |
| 2020 | Determination of glottal closure instants from clean and telephone quality speech signals using single frequency filtering
Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
Comput. Speech Lang. | 1 |
| 2020 | Analysis and classification of phonation types in speech and singing voice
Sudarsana Reddy Kadiri, Paavo Alku, Bayya Yegnanarayana |
Speech Commun. | 1 |
| 2020 | Detection of glottal closure instant and glottal open region from speech signals using spectral flatness measure
Sudarsana Reddy Kadiri, RaviShankar Prasad, Bayya Yegnanarayana |
Speech Commun. | 1 |
| 2020 | Time-Varying Quasi-Closed-Phase Analysis for Accurate Formant Tracking in Speech SignalsabstractIn this paper, we propose a new method for the accurate estimation and tracking of formants in speech signals using time-varying quasi-closed-phase (TVQCP) analysis. Conventional formant tracking methods typically adopt a two-stage estimateand-track strategy wherein an initial set of formant candidates are estimated using short-time analysis (e.g., 10-50 ms), followed by a tracking stage based on dynamic programming or a linear state-space model. One of the main disadvantages of these approaches is that the tracking stage, however good it may be, cannot improve upon the formant estimation accuracy of the first stage. The proposed TVQCP method provides a single-stage formant tracking that combines the estimation and tracking stages into one. TVQCP analysis combines three approaches to improve formant estimation and tracking: (1) it uses temporally weighted quasi-closed-phase analysis to derive closed-phase estimates of the vocal tract with reduced interference from the excitation source, (2) it increases the residual sparsity by using the L1 optimization and (3) it uses time-varying linear prediction analysis overlong time windows (e.g., 100-200 ms) to impose a continuity constraint on the vocal tract model and hence on the formant trajectories. Formant tracking experiments with a wide variety of synthetic and natural speech signals show that the proposed TVQCP method performs better than conventional and popular formant tracking tools, such as Wavesurfer and Praat (based on dynamic programming), the KARMA algorithm (based on Kalman filtering), and DeepFormants (based on deep neural networks trained in a supervised manner). Matlab scripts for the proposed method can be found at: Dhananjaya Gowda, Sudarsana Reddy Kadiri, Brad H. Story, Paavo Alku |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | A Quantitative Comparison of Epoch Extraction Algorithms for Telephone SpeechabstractTelephone speech is one of the degradations involved in building speech systems in practical environments. The potential use of the speech systems depends on the speech analysis algorithms that can handle different acoustic variations and degradations often found in the human speech communication. Detection of epochs/glottal closure instants (GCIs) is typically required in such analysis stages. In this paper, the effect of telephone channel speech on the accuracy of detection of epochs using state-of-art epoch extraction methods is investigated. Epoch is the instant of significant excitation to the vocal tract system in voiced speech. Most of the existing epoch extraction algorithms are shown to perform excellently well on the speech data collected under lab environment. The efficiency of these algorithms for the analysis of telephone quality speech is quantitatively studied and the strengths and weaknesses of the methods are discussed here. The methods are evaluated on six large databases containing speech and simultaneous EGG recordings as the ground truth. The state-of-art epoch extraction algorithms considered in this study for comparison are: ZFF, YAGA, DYPSA, SEDREAMS, SE-VQ and MMF. The performance of the algorithms is evaluated in terms of both reliability and accuracy measures. Sudarsana Reddy Kadiri |
ICASSP | 1 |
| 2019 | Mel-Frequency Cepstral Coefficients of Voice Source Waveforms for Classification of Phonation Types in SpeechabstractVoice source characteristics in different phonation types vary due to the tension of laryngeal muscles along with the respiratory effort. This study investigates the use of mel-frequency cepstral coefficients (MFCCs) derived from voice source waveforms for classification of phonation types in speech. The cepstral coefficients are computed using two source waveforms: (1) glottal flow waveforms estimated by the quasi-closed phase (QCP) glottal inverse filtering method and (2) approximate voice source waveforms obtained using the zero frequency filtering (ZFF) method. QCP estimates voice source waveforms based on the source-filter decomposition while ZFF yields source waveforms without explicitly computing the source-filter decomposition. Experiments using MFCCs computed from the two source waveforms show improved accuracy in classification of phonation types compared to the existing voice source features and conventional MFCC features. Further, it is observed that the proposed features have complimentary information to the existing features. Sudarsana Reddy Kadiri, Paavo Alku |
INTERSPEECH | 1 |
| 2019 | Spectral and temporal manipulations of SFF envelopes for enhancement of speech intelligibility in noise
Nivedita Chennupati, Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
Comput. Speech Lang. | 2 |
| 2018 | Detection of Glottal Closure Instants in Degraded Speech Using Single Frequency Filtering Analysis
Gunnam Aneeja, Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2018 | Breathy to Tense Voice Discrimination using Zero-Time Windowing Cepstral Coefficients (ZTWCCs)
Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
INTERSPEECH | 1 |
| 2018 | Analysis and Detection of Phonation Modes in Singing Voice using Excitation Source Features and Single Frequency Filtering Cepstral Coefficients (SFFCC)
Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
INTERSPEECH | 1 |
| 2018 | Estimation of Fundamental Frequency from Singing Voice Using Harmonics of Impulse-like Excitation Source
Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
INTERSPEECH | 1 |
| 2018 | Discriminating Nasals and Approximants in English Language Using Zero Time Windowing
RaviShankar Prasad, Sudarsana Reddy Kadiri, Suryakanth V. Gangashetty, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2018 | Significance of phase in single frequency filtering outputs of speech signals
Nivedita Chennupati, Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
Speech Commun. | 2 |
| 2017 | Speech polarity detection using strength of impulse-like excitation extracted from speech epochsabstractIn this paper, we address the issue of speech polarity detection using strength of impulse-like excitation around epoch. The correct detection of speech polarity is a crucial step for many speech processing algorithms to extract suitable information. Occurrence of errors in the detection of speech polarity could have an impact on the performance of speech systems. Automatic detection of speech polarity has become an important preliminary step for many speech processing algorithms. We propose a method based on the knowledge of impulse-like excitation of speech production mechanism. The impulse-like excitation is reflected across all frequencies including the zero frequency (0 Hz). Using the slope around zero crossings of the zero frequency filtered signal, an automatic speech polarity detection method is proposed. Performance of the proposed method is demonstrated on 8 different speech corpora. The proposed method is compared with the three existing techniques such as gradient of the spurious glottal waveforms (GSGW), oscillating moments-based polarity detection (OMPD) and residual excitation skewness (RESKEW). From the experimental results, it is observed that the performance of the proposed method is comparable or better than the existing methods for the experiments considered. Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
ICASSP | 1 |
| 2017 | SFF Anti-Spoofer: IIIT-H Submission for Automatic Speaker Verification Spoofing and Countermeasures Challenge 2017
K. N. R. K. Raju Alluri, Sivanand Achanta, Sudarsana Reddy Kadiri, Suryakanth V. Gangashetty, Anil Kumar Vuppala |
INTERSPEECH | 3 |
| 2017 | Detection of Replay Attacks Using Single Frequency Filtering Cepstral Coefficients
K. N. R. K. Raju Alluri, Sivanand Achanta, Sudarsana Reddy Kadiri, Suryakanth V. Gangashetty, Anil Kumar Vuppala |
INTERSPEECH | 3 |
| 2017 | Locating Burst Onsets Using SFF Envelope and Phase Information
Bhanu Teja Nellore, RaviShankar Prasad, Sudarsana Reddy Kadiri, Suryakanth V. Gangashetty, Bayya Yegnanarayana |
INTERSPEECH | 3 |
| 2017 | Epoch extraction from emotional speech using single frequency filtering approachabstractEpochs are instants of significant excitation of the vocal tract system during production of voiced speech. Existing methods for epoch extraction provide good results on neutral speech. But effectiveness of these methods has not been examined carefully for analysis of emotional speech, where the emotion characteristics are embedded mainly in the source component of the signal. Performance of the state-of-art epoch extraction methods on emotional speech data may be affected due to large variations in the pitch period. An approach, which exploits the nature of impulse-like excitation in the speech signal, instead of the pitch period information, is explored in this paper. The approach uses single frequency filtering (SFF) analysis of speech signals, which provides the high temporal resolution of some features of excitation source (such as impulse-like events) and high spectral resolution for some features of spectrum (such as harmonics and resonances). The Berlin emotional speech database (EMO-DB), which contains the simultaneous electroglottograph (EGG) recordings is used as the ground truth. For comparison, several epoch extraction methods are evaluated in terms of both reliability and accuracy measures for six different emotion categories and neutral speech. The results indicate that the performance of the proposed SFF-based methods for emotional speech is comparable to the results for neutral speech, and is better than the results from many of the standard methods. Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
Speech Commun. | 1 |
| 2016 | Robust Estimation of Fundamental Frequency Using Single Frequency Filtering Approach
Vishala Pannala, Gunnam Aneeja, Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
INTERSPEECH | 3 |
| 2015 | Analysis of singing voice for epoch extraction using Zero Frequency Filtering methodabstractEpoch is the instant of significant excitation of the vocal tract system during the production of voiced speech. Estimation of epochs or Glottal closure instants (GCIs) is a well studied topic in the speech analysis. From the recent studies on GCI detection from singing voice with state-of-art methods proposed for speech, there exist a clear gap in accuracy between speech and singing voice. This is because of source-filter interaction in singing voice compared to speech. Performance of existing algorithms deteriorates as most of the techniques depends on the ability to model the vocal tract system in order to emphasize the excitation characteristics in the residual. The objective of this paper is to analyze the singing voice for the estimation of epochs by studying the characteristics of the source-filter interaction and the effect of wider range of pitch using the Zero Frequency Filtering (ZFF) method. It is observed that high source-filter interaction can be captured in the form of the impulse-like excitation by passing the signal through three ideal digital resonators having poles at zero frequency, and the effect of wider range of pitch can be controlled by processing short segment (0.4-0.5 sec) signal. Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
ICASSP | 1 |
| 2015 | Analysis of excitation source features of speech for emotion recognition
Sudarsana Reddy Kadiri, P. Gangamohan, Suryakanth V. Gangashetty, Bayya Yegnanarayana |
INTERSPEECH | 1 |
| 2014 | Excitation source features for discrimination of anger and happy emotionsabstractStudies on the emotion recognition task indicate that there is confusion in discrimination among higher activation states like ‘anger’ and ‘happy’. In this study, features related to excitation source of speech are examined for discriminating ‘anger’ and ‘happy’ emotions. The objective is to explore the features which are independent of lexical content, language, channel and speaker. The features like strength of excitation from zero frequency filtering method and spectral band magnitude energies from short-time spectral analysis are used. Experimental results show that these features can discriminate ‘anger’ and ‘happy’ emotion states to a good extent. Index Terms: Emotion recognition, zero frequency filtering method, KL distance measure. P. Gangamohan, Sudarsana Reddy Kadiri, Suryakanth V. Gangashetty, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2013 | Analysis of emotional speech at subsegmental level
P. Gangamohan, Sudarsana Reddy Kadiri, Bayya Yegnanarayana |
INTERSPEECH | 2 |