VLDB 2026 Research / reviewers in the wild / expert
Dhananjaya Gowda
dblp:255/7153 · also Dhananjaya N. Gowda, N. Dhananjaya
· DBLP profile ↗
48ranked-venue papers
16as first author
16since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 42 · 13 first-author · 14 since 2021Artificial intelligence and machine learning · 35 · 12 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | zFLoRA: Zero-Latency Fused Low-Rank AdaptersabstractLarge language models (LLMs) are increasingly deployed with task-specific adapters catering to multiple downstream applications.In such a scenario, the additional compute associated with these apparently insignificant number of adapter parameters (typically less than 1% of the base model) turns out to be disproportionately significant during inference time (upto 2.5x times that of the base model).In this paper, we propose a new zero-latency fused low-rank adapter (zFLoRA) that introduces zero or negligible latency overhead on top of the base model.Experimental results on LLMs of size 1B, 3B and 7B show that zFLoRA compares favorably against the popular supervised fine-tuning benchmarks including lowrank adapters (LoRA) as well as full fine-tuning (FFT).Experiments are conducted on 18 different tasks across three different categories namely commonsense reasoning, math reasoning and summary-dialogue.Latency measurements made on NPU (Samsung Galaxy S25+) as well as GPU (NVIDIA H100) platforms show that the proposed zFLoRA adapters introduce zero to negligible latency overhead. Dhananjaya Gowda, Seoha Song, Harshith Goka, Jun-Hyun Lee |
EMNLP | 1 |
| 2024 | Data Driven Grapheme-to-Phoneme Representations for a Lexicon-Free Text-to-SpeechabstractGrapheme-to-Phoneme (G2P) is an essential first step in any modern, high-quality Text-to-Speech (TTS) system. Most of the current G2P systems rely on carefully hand-crafted lexicons developed by experts. This poses a two-fold problem. Firstly, the lexicons are generated using a fixed phoneme set, usually, ARPABET or IPA, which might not be the most optimal way to represent phonemes for all languages. Secondly, the man-hours required to produce such an expert lexicon are very high. In this paper, we eliminate both of these issues by using recent advances in self-supervised learning to obtain data-driven phoneme representations instead of fixed representations. We compare our lexicon-free approach against strong baselines that utilize a well-crafted lexicon. Furthermore, we show that our data-driven lexicon-free method performs as good or even marginally better than the conventional rule-based or lexicon-based neural G2Ps in terms of Mean Opinion Score (MOS) while using no prior language lexicon or phoneme set, i.e. no linguistic expertise. Abhinav Garg, Jiyeon Kim, Sushil Khyalia, Chanwoo Kim 0001, Dhananjaya Gowda |
ICASSP | 5 |
| 2023 | A Transformer-Based E2E SLU Model for Improved Semantic ParsingabstractSpoken Language Understanding (SLU) is an essential part of voice and speech assistant tools. End-to-End (E2E) SLU models attempt to automatically extract semantic meanings from the speech signal without the need for an intermediate transcription of speech. However, SLU is a challenging task mainly due to the lack of labeled, in-domain, and multilingual datasets. The Spoken Task-Oriented Semantic Parsing (STOP) dataset tries to address this problem and is the most extensive public dataset for the SLU task. This paper demonstrates our contribution to the Spoken Language Understanding Grand Challenge at ICASSP 2023. The fundamental idea of the proposed model is to utilize the pre-trained HuBERT model as an encoder alongside a transformer decoder with layer-drop and ensemble learning. The combination of HuBERT large encoder and a base transformer decoder obtained the best results, with an Exact Match (EM) accuracy of 75.05% on the STOP dataset. Ensemble decoding improved the accuracy to 75.92%. Othman Istaiteh, Yasmeen Kussad, Yahya Daqour, Maria Habib, Mohammad Habash, Dhananjaya Gowda |
ICASSP | 6 |
| 2023 | Self-Supervised Accent Learning for Under-Resourced Accents Using Native Language DataabstractIn this paper, we propose a novel method to improve the accuracy of an English speech recognizer for a target accent using the corresponding native language data. Collecting labeled data for all accents of English to train an end-to-end neural speech recognizer for English is a difficult and expensive task. Also, finding a pool of representative English speakers for any arbitrary accent to collect unlabeled data can be a difficult task. However, collecting unlabeled speech data for any native language is a much simpler task. It is important to note that the accents of most non-native English speakers are heavily biased by the co-articulation of sounds in their own native language. In view of this, we propose to use unlabeled native language data to learn self-supervised representations during the pre-training stage. The pre-trained model is then fine-tuned using limited labeled English data for the target accent. Experiments using native language data to pre-train an English recognizer followed by fine-tuning using target accented English show significant improvements in word error rates on four different accents (Great Britain, Korean, Chinese, Spanish). Mehul Kumar, Jiyeon Kim, Dhananjaya Gowda, Abhinav Garg, Chanwoo Kim 0001 |
ICASSP | 3 |
| 2023 | Mitigating the Exposure Bias in Sentence-Level Grapheme-to-Phoneme (G2P) Transduction
Eunseop Yoon, Hee Suk Yoon, Dhananjaya Gowda, SooHwan Eom, Daehyeok Kim, John B. Harvill, Heting Gao, Mark Hasegawa-Johnson, Chanwoo Kim 0001, Chang Dong Yoo |
INTERSPEECH | 3 |
| 2023 | Refining a deep learning-based formant tracker using linear prediction methodsabstractIn this study, formant tracking is investigated by refining the formants tracked by an existing data-driven tracker, DeepFormants, using the formants estimated in a model-driven manner by linear prediction (LP)-based methods. As LP-based formant estimation methods, conventional covariance analysis (LP-COV) and the recently proposed quasi-closed phase forward–backward (QCP-FB) analysis are used. In the proposed refinement approach, the contours of the three lowest formants are first predicted by the data-driven DeepFormants tracker, and the predicted formants are replaced frame-wise with local spectral peaks shown by the model-driven LP-based methods. The refinement procedure can be plugged into the DeepFormants tracker with no need for any new data learning. Two refined DeepFormants trackers were compared with the original DeepFormants and with five known traditional trackers using the popular vocal tract resonance (VTR) corpus. The results indicated that the data-driven DeepFormants trackers outperformed the conventional trackers and that the best performance was obtained by refining the formants predicted by DeepFormants using QCP-FB analysis. In addition, by tracking formants using VTR speech that was corrupted by additive noise, the study showed that the refined DeepFormants trackers were more resilient to noise than the reference trackers. In general, these results suggest that LP-based model-driven approaches, which have traditionally been used in formant estimation, can be combined with a modern data-driven tracker easily with no further training to improve the tracker’s performance. Paavo Alku, Sudarsana Reddy Kadiri, Dhananjaya Gowda |
Comput. Speech Lang. | 3 |
| 2022 | Prototypical speaker-interference loss for target voice separation using non-parallel audio samples
Seongkyu Mun, Dhananjaya Gowda, Dokyun Lee, Chanwoo Kim 0001 |
INTERSPEECH | 2 |
| 2022 | Multi-stage Progressive Compression of Conformer Transducer for On-device Speech RecognitionabstractThe smaller memory bandwidth in smart devices prompts development of smaller Automatic Speech Recognition (ASR) models. To obtain a smaller model, one can employ the model compression techniques. Knowledge distillation (KD) is a popular model compression approach that has shown to achieve smaller model size with relatively lesser degradation in the model performance. In this approach, knowledge is distilled from a trained large size teacher model to a smaller size student model. Also, the transducer based models have recently shown to perform well for on-device streaming ASR task, while the conformer models are efficient in handling long term dependencies. Hence in this work we employ a streaming transducer architecture with conformer as the encoder. We propose a multi-stage progressive approach to compress the conformer transducer model using KD. We progressively update our teacher model with the distilled student model in a multi-stage setup. On standard LibriSpeech dataset, our experimental results have successfully achieved compression rates greater than 60% without significant degradation in the performance compared to the larger teacher model. Jash Rathod, Nauman Dawalatabad, Shatrughan Singh, Dhananjaya Gowda |
INTERSPEECH | 4 |
| 2021 | Two-Pass End-to-End ASR Model CompressionabstractSpeech recognition on smart devices is challenging owing to the small memory footprint. Hence small size ASR models are desirable. With the use of popular transducer-based models, it has become possible to practically deploy streaming speech recognition models on small devices [1]. Recently, the two-pass model [2] combining RNN-T and LAS modules has shown exceptional performance for streaming on-device speech recognition. In this work, we propose a simple and effective approach to reduce the size of the two-pass model for memory-constrained devices. We employ a popular knowledge distillation approach in three stages using the Teacher-Student training technique. In the first stage, we use a trained RNN-T model as a teacher model and perform knowledge distillation to train the student RNN-T model. The second stage uses the shared encoder and trains a LAS rescorer for student model using the trained RNN-T+LAS teacher model. Finally, we perform deep-finetuning for the student model with a shared RNN-T encoder, RNN-T decoder, and LAS rescorer. Our experimental results on standard LibriSpeech dataset show that our system can achieve a high compression rate of 55% without significant degradation in the WER compared to the two-pass teacher model. Nauman Dawalatabad, Tushar Vatsal, Shatrughan Singh, Dhananjaya Gowda, Chanwoo Kim 0001 |
ASRU | 6 |
| 2021 | HiTNet: Byte-to-BPE Hierarchical Transcription Network for End-to-End Speech RecognitionabstractIn this paper, we propose a new byte to byte-pair-encoding (BPE) Hierarchical Transcription Network (HiTNet) architecture for end-to-end (e2e) automatic speech recognition (ASR). The proposed HiTNet architecture simultaneously encodes as well as decodes information hierarchically at different levels of linguistic granularity such as bytes and BPE. In general this idea can be extended to any levels of granularity including phonemes or graphemes or bytes (character to sub-character in some languages), to sub-words or byte-pair encodings (BPE), to words, and so on. Existing hierarchical e2e ASR models primarily encode the acoustic information in an hierarchical manner governed by weaker linguistic constraints at each level. The language information at each level is neither embedded or used explicitly, nor is the information decoded at each level passed on to the next stage. The proposed architecture primarily decodes information in an hierarchical manner utilizing the linguistic information at each level explicitly, while at the same time utilizing the hierarchically encoded acoustic information at each level. Experiments with a two-level byte-to-BPE (b2B) hierarchical transcription show that the proposed architecture significantly reduces the word error rates of both the byte and BPE decoders compared to baseline byte and BPE based attention encoder-decoder models. Dhananjaya Gowda, Abhinav Garg, Jiyeon Kim, Mehul Kumar, Nauman Dawalatabad, Aman Maghan, Shatrughan Singh, Chanwoo Kim 0001 |
ASRU | 1 |
| 2021 | Voice to Action: Spoken Language Understanding for Memory-Constrained SystemsabstractSpoken Language Understanding (SLU) is the task of extracting semantic information from speech. Traditional architectures of SLU involve a two-stage pipeline consisting of an Automatic Speech Recognition (ASR) module that converts speech into text hypotheses, followed by a Natural Language Understanding (NLU) module that extracts semantic information from the text hypotheses. The increasing demand for light-weight modules, especially in on-device applications, promulgated end-to-end SLU, where a single module directly converts speech into semantic information. This work pro-poses a low resource, lightweight end-to-end SLU model that treats SLU as a sequence-to-sequence task and is optimized similar to the ASR task. The proposed model utilizes transfer learning from a heavy-resource, general domain ASR model to achieve an 10.8% absolute improvement in the Word Error Rate (WER) over the base model. We then utilize Knowledge Distillation to achieve a further 31.9% reduction in the model size with 22.7% relative WER improvement. In addition, we compress our models by more than 4 times using Low-Rank Approximation (LRA) method with minimum degradation in accuracy. Aditya Jayasimha, Aman Maghan, Shatrughan Singh, Dhananjaya Gowda, Chanwoo Kim 0001 |
ASRU | 5 |
| 2021 | Semi-Supervised Transfer Learning for Language Expansion of End-to-End Speech Recognition Models to Low-Resource Languages
Jiyeon Kim, Mehul Kumar, Dhananjaya Gowda, Abhinav Garg, Chanwoo Kim 0001 |
ASRU | 3 |
| 2021 | A Comparison of Streaming Models and Data Augmentation Methods for Robust Speech RecognitionabstractIn this paper, we present a comparative study on the robustness of two different online streaming speech recognition models: Monotonic Chunkwise Attention (MoChA) and Recurrent Neural Network-Transducer (RNN-T). We explore three recently proposed data augmentation techniques, namely, multi-conditioned training using an acoustic simulator, Vocal Tract Length Perturbation (VTLP) for speaker variability, and SpecAugment. Experimental results show that unidirectional models are in general more sensitive to noisy examples in the training set. It is observed that the final performance of the model depends on the proportion of training examples processed by data augmentation techniques. MoChA models generally perform better than RNN-T models. However, we observe that training of MoChA models seems to be more sensitive to various factors such as the characteristics of training sets and the incorporation of additional augmentations techniques. On the other hand, RNN-T models perform better than MoChA models in terms of latency, inference time, and the stability of training. Additionally, RNN-T models are generally more robust against noise and reverberation. All these advantages make RNN-T models a better choice for streaming on-device speech recognition compared to MoChA models. Jiyeon Kim, Mehul Kumar, Dhananjaya Gowda, Abhinav Garg, Chanwoo Kim 0001 |
ASRU | 3 |
| 2021 | Comparative Study of Different Tokenization Strategies for Streaming End-to-End ASRabstractMost End-to-End Automatic Speech Recognition (ASR) models use character-based vocabularies: characters, sub-words (BPE), or words. While these work well for training a monolingual model, there are certain limitations when ap-plying these to a multilingual model. Rare characters from character-rich languages like Korean can easily result in large vocabulary size, limiting the model's compactness. Repre-senting text at the level of bytes has also been proposed. However, a byte sequence representation of text is often much longer, which increases the decoding time and makes it computationally expensive for on-device use. Byte-based sub-words (BBPE) are proposed in neural machine translation for word representation but are still unexplored in the ASR domain. In this work, we conduct an empirical study comparing the above three tokenization strategies across three metrics: Word Error Rate (WER), model size, and the decoding time, which are critical for an on-device ASR. We did extensive experiments for both monolingual and bilingual, with languages belonging to same (English and Spanish) and different (English and Korean) language families. Our exper-iments show that BBPE and BPE models yield a similar WER for English and Spanish. While for a character-rich language like Korean, we get 26% and 14% relative WER improvement with BBPE monolingual and bilingual models, respectively. In contrast, the byte models trade-off small model size and a fixed vocabulary at the cost of high xRT. Among all three, we found the BBPE strategy to be the most flexible and optimal for most cases. Aman Maghan, Dhananjaya Gowda, Shatrughan Singh, Chanwoo Kim 0001 |
ASRU | 4 |
| 2021 | Neural Utterance Confidence Measure for RNN-Transducers and Two Pass ModelsabstractIn this paper, we propose methods to compute confidence score on the predictions made by an end-to-end speech recognition model in a 2-pass framework. We use RNN-Transducer for a streaming model, and an attention-based decoder for the second pass model. We use neural technique to compute the confidence score, and experiment with various combinations of features from RNN-Transducer and second pass models. The neural confidence score model is trained as a binary classification task to accept or reject a prediction made by speech recognition model. The model is evaluated in a distributed speech recognition environment, and performs significantly better when features from second pass model are used as compared to the features from streaming model. Dhananjaya Gowda, Kwangyoun Kim, Shatrughan Singh, Chanwoo Kim 0001 |
ICASSP | 3 |
| 2021 | Streaming End-to-End Speech Recognition with Jointly Trained Neural Feature EnhancementabstractIn this paper, we present a streaming end-to-end speech recognition model based on Monotonic Chunkwise Attention (MoCha) jointly trained with enhancement layers. Even though the MoCha attention enables streaming speech recognition with recognition accuracy comparable to a full attention-based approach, training this model is sensitive to various factors such as the difficulty of training examples, hyper-parameters, and so on. Because of these issues, speech recognition accuracy of a MoCha-based model for clean speech drops significantly when a multi-style training approach is applied. Inspired by Curriculum Learning [1], we introduce two training strategies: Gradual Application of Enhanced Features (GAEF) and Gradual Reduction of Enhanced Loss (GREL). With GAEF, the model is initially trained using clean features. Subsequently, the portion of outputs from the enhancement layers gradually increases. With GREL, the portion of the Mean Squared Error (MSE) loss for the enhanced output gradually reduces as training proceeds. In experimental results on the LibriSpeech corpus and noisy far-field test sets, the proposed model with GAEF-GREL training strategies shows significantly better results than the conventional multi-style training approach. Chanwoo Kim 0001, Abhinav Garg, Dhananjaya Gowda, Seongkyu Mun |
ICASSP | 3 |
| 2020 | Hierarchical Multi-Stage Word-to-Grapheme Named Entity Corrector for Automatic Speech Recognition
Abhinav Garg, Dhananjaya Gowda, Shatrughan Singh, Chanwoo Kim 0001 |
INTERSPEECH | 3 |
| 2020 | Streaming On-Device End-to-End ASR System for Privacy-Sensitive Voice-Typing
Abhinav Garg, Gowtham P. Vadisetti, Dhananjaya Gowda, Sichen Jin, Aditya Jayasimha, Youngho Han, Jiyeon Kim, Junmo Park, Kwangyoun Kim, Young-Yoon Lee, Kyungbo Min, Chanwoo Kim 0001 |
INTERSPEECH | 3 |
| 2020 | Utterance Invariant Training for Hybrid Two-Pass End-to-End Speech Recognition
Dhananjaya Gowda, Kwangyoun Kim, Hejung Yang, Abhinav Garg, Jiyeon Kim, Mehul Kumar, Sichen Jin, Shatrughan Singh, Chanwoo Kim 0001 |
INTERSPEECH | 1 |
| 2020 | Utterance Confidence Measure for End-to-End Speech Recognition with Applications to Distributed Speech Recognition Scenarios
Dhananjaya Gowda, Abhinav Garg, Shatrughan Singh, Chanwoo Kim 0001 |
INTERSPEECH | 3 |
| 2020 | Time-Varying Quasi-Closed-Phase Analysis for Accurate Formant Tracking in Speech SignalsabstractIn this paper, we propose a new method for the accurate estimation and tracking of formants in speech signals using time-varying quasi-closed-phase (TVQCP) analysis. Conventional formant tracking methods typically adopt a two-stage estimateand-track strategy wherein an initial set of formant candidates are estimated using short-time analysis (e.g., 10-50 ms), followed by a tracking stage based on dynamic programming or a linear state-space model. One of the main disadvantages of these approaches is that the tracking stage, however good it may be, cannot improve upon the formant estimation accuracy of the first stage. The proposed TVQCP method provides a single-stage formant tracking that combines the estimation and tracking stages into one. TVQCP analysis combines three approaches to improve formant estimation and tracking: (1) it uses temporally weighted quasi-closed-phase analysis to derive closed-phase estimates of the vocal tract with reduced interference from the excitation source, (2) it increases the residual sparsity by using the L1 optimization and (3) it uses time-varying linear prediction analysis overlong time windows (e.g., 100-200 ms) to impose a continuity constraint on the vocal tract model and hence on the formant trajectories. Formant tracking experiments with a wide variety of synthetic and natural speech signals show that the proposed TVQCP method performs better than conventional and popular formant tracking tools, such as Wavesurfer and Praat (based on dynamic programming), the KARMA algorithm (based on Kalman filtering), and DeepFormants (based on deep neural networks trained in a supervised manner). Matlab scripts for the proposed method can be found at: Dhananjaya Gowda, Sudarsana Reddy Kadiri, Brad H. Story, Paavo Alku |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Improved Multi-Stage Training of Online Attention-Based Encoder-Decoder ModelsabstractIn this paper, we propose a refined multi-stage multi-task training strategy to improve the performance of onlineattention-based encoder-decoder (AED) models. A three-stage training based on three levels of architectural granularity namely, character encoder, byte pair encoding (BPE) based encoder, and attention decoder, is proposed. Also, multi-task learning based on two-levels of linguistic granularity namely, character and BPE, is used. We explore different pre-training strategies for the encoders including transfer learning from a bidirectional encoder. Our encoder-decoder models with online attention show ~35% and ~10% relative improvement over their baselines for smaller and bigger models, respectively. Our models achieve a word error rate (WER) of 5.04% and 4.48% on the Librispeech test-clean data for the smaller and bigger models respectively after fusion with long short-term memory (LSTM) based external language model (LM). Abhinav Garg, Dhananjaya Gowda, Kwangyoun Kim, Mehul Kumar, Chanwoo Kim 0001 |
ASRU | 2 |
| 2019 | Attention Based On-Device Streaming Speech Recognition with Large Speech CorpusabstractIn this paper, we present a new on-device automatic speech recognition (ASR) system based on monotonic chunk-wise attention (MoChA) models trained with large (> 10K hours) corpus. We attained around 90% of a word recognition rate for general domain mainly by using joint training of connectionist temporal classifier (CTC) and cross entropy (CE) losses, minimum word error rate (MWER) training, layer-wise pretraining and data augmentation methods. In addition, we compressed our models by more than 3.4 times smaller using an iterative hyper low-rank approximation (LRA) method while minimizing the degradation in recognition accuracy. The memory footprint was further reduced with 8-bit quantization to bring down the final model size to lower than 39 MB. For on-demand adaptation, we fused the MoChA models with statistical n-gram models, and we could achieve a relatively 36% improvement on average in word error rate (WER) for target domains including the general domain. Kwangyoun Kim, Seokyeong Jung, Jungin Lee, Myoungji Han, Chanwoo Kim 0001, Kyungmin Lee, Dhananjaya Gowda, Junmo Park, Sichen Jin, Young-Yoon Lee, Jinsu Yeo |
ASRU | 7 |
| 2019 | Power-Law Nonlinearity with Maximally Uniform Distribution Criterion for Improved Neural Network Training in Automatic Speech RecognitionabstractIn this paper, we describe the Maximum Uniformity of Distribution (MUD) algorithm with the power-law nonlinearity. In this approach, we hypothesize that neural network training will become more stable if feature distribution is not too much skewed. We propose two different types of MUD approaches: power function-based MUD and histogram-based MUD. In these approaches, we first obtain the mel filterbank coefficients and apply nonlinearity functions for each filterbank channel. With the power function-based MUD, we apply a power-function based nonlinearity where power function coefficients are chosen to maximize the likelihood assuming that nonlinearity outputs follow the uniform distribution. With the histogram-based MUD, the empirical Cumulative Density Function (CDF) from the training database is employed to transform the original distribution into a uniform distribution. In MUD processing, we do not use any prior knowledge (e.g. logarithmic relation) about the energy of the incoming signal and the perceived intensity by a human. Experimental results using an end-to-end speech recognition system demonstrate that power-function based MUD shows better result than the conventional Mel Filterbank Cepstral Coefficients (MFCCs). On the LibriSpeech database, we could achieve 4.02 % WER on test-clean and 13.34 % WER on test-other without using any Language Models (LMs). The major contribution of this work is that we developed a new algorithm for designing the compressive nonlinearity in a data-driven way, which is much more flexible than the previous approaches and may be extended to other domains as well. Chanwoo Kim 0001, Mehul Kumar, Kwangyoun Kim, Dhananjaya Gowda |
ASRU | 4 |
| 2019 | End-to-End Training of a Large Vocabulary End-to-End Speech Recognition SystemabstractIn this paper, we present an end-to-end training framework for building state-of-the-art end-to-end speech recognition systems. Our training system utilizes a cluster of Central Processing Units (CPUs) and Graphics Processing Units (GPUs). The entire data reading, large scale data augmentation, neural network parameter updates are all performed “on-the-fly”. We use vocal tract length perturbation [1] and an acoustic simulator [2] for data augmentation. The processed features and labels are sent to the GPU cluster. The Horovod allreduce approach is employed to train neural network parameters. We evaluated the effectiveness of our system on the standard Librispeech corpus [3] and the 10,000-hr anonymized Bixby English dataset. Our end-to-end speech recognition system built using this training infrastructure showed a 2.44 % WER on test-clean of the LibriSpeech test set after applying shallow fusion with a Transformer language model (LM). For the proprietary English Bixby open domain test set, we obtained a WER of 7.92 % using a Bidirectional Full Attention (BFA) end-to-end model after applying shallow fusion with an RNN-LM. When the monotonic chunckwise attention (MoCha) based approach is employed for streaming speech recognition, we obtained a WER of 9.95 % on the same Bixby open domain test set. Chanwoo Kim 0001, Minkyoo Shin, Shatrughan Singh, Larry Heck, Dhananjaya Gowda, Kwangyoun Kim, Mehul Kumar, Jiyeon Kim, Kyungmin Lee, Abhinav Garg, Eunhyang Kim |
ASRU | 5 |
| 2019 | Multi-Task Multi-Resolution Char-to-BPE Cross-Attention Decoder for End-to-End Speech Recognition
Dhananjaya Gowda, Abhinav Garg, Kwangyoun Kim, Mehul Kumar, Chanwoo Kim 0001 |
INTERSPEECH | 1 |
| 2019 | Improved Vocal Tract Length Perturbation for a State-of-the-Art End-to-End Speech Recognition System
Chanwoo Kim 0001, Minkyu Shin, Abhinav Garg, Dhananjaya Gowda |
INTERSPEECH | 4 |
| 2018 | Speaker recognition from whispered speech: A tutorial survey and an application of time-varying linear prediction
Ville Vestman, Dhananjaya Gowda, Md. Sahidullah, Paavo Alku, Tomi Kinnunen |
Speech Commun. | 2 |
| 2017 | Time-Varying Autoregressions for Speaker Verification in Reverberant ConditionsabstractAutomatic speaker verification (ASV) systems are vulnerable to spoofing attacks using speech generated by voice conversion and speech synthesis techniques. Commonly, a countermeasure (CM) system is integrated with an ASV system for improved protection against spoofing attacks. But integration of the two systems is challenging and often leads to increased false rejection rates. Furthermore, the performance of CM severely degrades if in-domain development data are unavailable. In this study, therefore, we propose a solution that uses two separate background models — one from human speech and another from spoofed data. During test, the ASV score for an input utterance is computed as the difference of the log-likelihood against the target model and the combination of the log-likelihoods against two background models. Evaluation experiments are conducted using the joint ASV and CM protocol of ASVspoof 2015 corpus consisting of text-independent ASV tasks with short utterances. Our proposed system reduces error rates in the presence of spoofing attacks by using out-of-domain spoofed data for system development, while maintaining the performance for zero-effort imposter attacks compared to the baseline system. Ville Vestman, Dhananjaya Gowda, Md. Sahidullah, Paavo Alku, Tomi Kinnunen |
INTERSPEECH | 2 |
| 2016 | Quasi closed phase analysis of speech signals using time varying weighted linear prediction for accurate formant trackingabstractRecent research on temporally weighted linear prediction shows that quasi closed phase (QCP) analysis of speech signals provides better modeling of the vocal tract and the glottal source. Quasi closed phase analysis gives more weightage on the closed phase of the glottal cycle, at the same time deemphasizing the region around the instant of significant excitation which is often poorly predicted. However, all the traditional analysis techniques including the QCP analysis is performed over short intervals of time. They do not impose any continuity constraints either on the vocal tract system or the glottal source. Such constraints are often imposed at a later stage to either smooth or track the estimated features over time. Time varying linear prediction (TVLP) provides a framework for modeling speech with a long-term continuity constraint imposed on the vocal tract shape. In this paper, we propose a new method for accurate modeling and tracking of the vocal tract resonances by integrating the advantages of a QCP analysis with that of TVLP. Formant tracking experiments show consistent improvement in performance over traditional LP or TVLP methods under a variety of conditions including different voice types and over a wide range of fundamental frequency. Dhananjaya Gowda, Manu Airaksinen, Paavo Alku |
ICASSP | 1 |
| 2016 | Time-Varying Quasi-Closed-Phase Weighted Linear Prediction Analysis of Speech for Accurate Formant Detection and Tracking
Dhananjaya Gowda, Paavo Alku |
INTERSPEECH | 1 |
| 2016 | Whispered Speech Detection Using Fusion of Group-Delay-Based Subband Modulation Spectrum and Correntropy FeaturesabstractIn this letter, we propose a novel fusion feature for detection of whispered speech in noisy environment using a group-delay-based instantaneous spectrum analysis. The fusion feature involves two individual components, namely, subband modulation spectrum (SMS)-based features and subband correntropy (SCE) features, both extracted from the instantaneous spectrum. The instantaneous spectrum estimation involves zero-time windowing for improved temporal resolution and group-delay computation for improved spectral resolution, as compared to the traditional discrete-Fourier-transform-based spectrum estimation. The SMS features capture the spectral representation of the subband energy time trajectories, while the SCE features model the fluctuations in the subband energy time trajectories. The SMS captures both the short-term as well as long-term spectral characteristics of whispered speech and is known to provide good separation between speech and noise components. The correntropy features help capture the dynamics of the vocal tract system to discriminate noisy whisper from noise. Whisper speech detection experiments using support vector machine models and the proposed features indicate promising performance under low signal-to-noise conditions. Jinfang Wang, Yongqiang Shang, Shuangshuang Jiang, Dhananjaya Gowda |
IEEE Signal Process. Lett. | 4 |
| 2015 | AM-FM based filter bank analysis for estimation of spectro-temporal envelopes and its application for speaker recognition in noisy reverberant environments
Dhananjaya Gowda, Rahim Saeidi, Paavo Alku |
INTERSPEECH | 1 |
| 2014 | On the role of missing data imputation and NMF feature enhancement in building synthetic voices using reverberant speech
Dhananjaya Gowda, Heikki Kallasjoki, Reima Karhila, Cristian Contan, Kalle J. Palomäki, Mircea Giurgiu, Mikko Kurimo |
INTERSPEECH | 1 |
| 2013 | Analysis of breathy, modal and pressed phonation based on low frequency spectral density
Dhananjaya Gowda, Mikko Kurimo |
INTERSPEECH | 1 |
| 2013 | Robust formant detection using group delay function and stabilized weighted linear prediction
Dhananjaya Gowda, Jouni Pohjalainen, Mikko Kurimo, Paavo Alku |
INTERSPEECH | 1 |
| 2013 | Spectro-temporal analysis of speech signals using zero-time windowing and group delay function
Bayya Yegnanarayana, Dhananjaya Gowda |
Speech Commun. | 2 |
| 2012 | Effect of Tongue Tip Trilling on the Glottal Excitation SourceabstractRecent studies have indicated changes in the glottal excitation source characteristics apart from vocal tract resonances due to tongue tip trilling. In this paper we study the significance of changing vocal tract system and the associated glottal excitation source characteristics due to trilling, from perception point of view. These studies are made by generating speech signal by either retaining the features of the vocal tract system or of the glottal excitation source of trill sounds. Experiments are conducted to understand the perceptual significance of the excitation source characteristics on production of different trill sounds. Speech sounds of sustained trill and approximant pair, and apical trills produced by four different places of articulation are considered. Features of the vocal tract system are extracted using linear prediction analysis, and those of the source by zero frequency filtering. Vinay Kumar Mittal, Dhananjaya Gowda, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2011 | Acoustic-phonetic information from excitation source for refining manner hypotheses of a phone recognizerabstractReliable acoustic-phonetic (AP) information derived from the speech signal can be used to detect and correct errors in the output of a phone recognizer. In this paper, limited acoustic-phonetic information derived primarily by processing the excitation source information in the speech signal is used to improve the performance of detection of manner of articulation from a baseline phone recognition system. A context-independent HMM-based monophone system without any language information is used as the baseline system for this purpose. The performance of the phone recognizer in terms of its ability to detect the manners of articulation is studied. The errors in the hypothesis of the manner of articulation of phones are corrected using AP information such as voicing, voice bar and frication. It is shown that significant improvement can be achieved by using simple or limited AP information. Dhananjaya Gowda, Bayya Yegnanarayana, Suryakanth V. Gangashetty |
ICASSP | 1 |
| 2011 | Decomposition of speech signals for analysis of aperiodic components of excitationabstractThe motivation for this study is the need for careful analysis of aperiodicity of the excitation component in expressive voices. The paper proposes analysis methods which can preserve the excitation information corresponding to sequence of impulse-like excitation with variable strengths. To analyze the details of the excitation source characteristics, the epochs and the strength of the excitation at the epochs are obtained using the output of an ideal zero-frequency digital resonator. The vocal tract system characteristics are derived from the signal between two successive epochs using the numerator of the group delay function. The spectrogram of the zero-frequency filtered signal and the group delay spectrum correspond to characteristics of the excitation and the vocal tract system, respectively. Decomposition of the speech signal into these two components bring out the features of excitation and vocal tract system, which can be used to explain the perception of expressive voices in terms of features of aperiodicity, pitch, harmonics and sub-harmonics. The decomposition method is illustrated using examples from linguistically significant glottalized sounds (glottal stops and ejectives), singing voices and Noh voice. Bayya Yegnanarayana, Anand Joseph Xavier Medabalimi, Suryakanth V. Gangashetty, Dhananjaya Gowda |
ICASSP | 4 |
| 2011 | Exploring Bessel Features for Detection of Glottal Closure InstantsabstractFor voiced speech, the most significant excitation takes place around the instant of glottal closure. Glottal closure instants (GCI) information is useful for accurate speech analysis. In particular accurate spectrum analysis is performed by considering the speech in the intervals of glottal closure. In this paper we propose an approach for detection of GCI by exploring Bessel feature, and the use of AM-FM signal. Using appropriate range of Bessel coefficients, the narrow band, band limited signal is obtained for the given signal.The bandlimited signal is considered as AM-FM signal. The signal is band limited for 0300 Hz to remove effect of formants. Amplitude envelope (AE) function of the AM-FM signal model has been estimated by the discrete energy separation algorithm (DESA). The performance of the method is demonstrated using CMU-Arctic database. The corresponding electro-glottograph (EGG) signals are used as a reference for the validation of the detected GCI locations. Chetana Prakash, Dhananjaya Gowda, Suryakanth V. Gangashetty |
INTERSPEECH | 2 |
| 2010 | Voiced/Nonvoiced Detection Based on Robustness of Voiced EpochsabstractIn this paper, a new method for voiced/nonvoiced detection based on epoch extraction is proposed. Zero-frequency filtered speech signal is used to extract the instants of significant excitation (or epochs). The robustness of the method to extract epochs in the voiced regions, even with small amount of additive white noise, is used to distinguish voiced epochs from random instants detected in nonvoiced regions. The main feature of the proposed method is that it uses the strength of glottal activity as against using the periodicity of the signal. Performance of the proposed algorithm is studied on TIMIT and CMU ARCTIC databases, for two different noise types, white and vehicle noise from the NOISEX database, at different signal-to-noise ratios (SNRs). The proposed method performs similar or better than the popular normalized crosscorrelation based voiced/nonvoiced detection used in the open source utilitywavesurfer, especially at lower SNRs. Dhananjaya Gowda, Bayya Yegnanarayana |
IEEE Signal Process. Lett. | 1 |
| 2008 | Video Shot Segmentation Using Late Fusion TechniqueabstractIn this paper, a new method for detecting shot boundaries in video sequences using a late fusion technique is proposed. The method uses color histogram as the feature, and processes each bin separately for detecting shot boundaries. The decisions from individual bins are combined later for hypothesizing the presence of shot boundaries. The method provides a certain degree of robustness against illumination and camera/object motion, as it ignores small changes in the bins. While the early fusion techniques rely on the extent of change in color information, the proposed technique relies on the number of significant changes. Experimental results successfully validate the new method and show that it can effectively detect both abrupt and gradual transitions. C. Krishna Mohan, Dhananjaya Gowda, Bayya Yegnanarayana |
ICMLA | 2 |
| 2008 | Features for automatic detection of voice bars in continuous speechabstractIn this paper we propose features for automatic detection of voice bar, which is an essential component of voiced stop consonants, in continuous speech. The acoustic-phonetic and production based knowledge such as, the presence of voicing, low strength of excitation compared to other voiced phones and a predominant low-frequency spectral energy, are mapped onto a set of acoustic features that can be automatically extracted from the signal. The usefulness of the proposed features in the detection of voice bars is studied using a knowledge-based as well as a neural network based approach. The performance of the proposed features and approaches is studied on phones from databases of two languages, namely English and Hindi. Dhananjaya Gowda, Bayya Yegnanarayana |
INTERSPEECH | 1 |
| 2008 | Analysis of glottal stops in speech signalsabstractDuring production of glottal stops the glottal vibration has un-equal cycles and is caused by laryngealization. While one can perceive the features of laryngealization in the speech, it is dif-ficult to analyse the signal to detect these source features from the standard spectrum-based analysis methods. In this paper we propose methods to extract the voice source vibration charac-teristics, and show that in the region of glottal stop, the pitch periods will be irregular, and the crosscorrelation coefficient of the signal in successive pitch periods will be low. This analy-sis enable us to locate the regions of glottal stops in continuous speech, and also help us to study the characteristics of creaky voice. Index Terms: glottal stop, laryngealization, creaky voice, glot-tal vibration, voice source, pitch period. Bayya Yegnanarayana, Hussien Seid Worku, Dhananjaya Gowda |
INTERSPEECH | 4 |
| 2008 | Speaker change detection in casual conversations using excitation source features
Dhananjaya Gowda, Bayya Yegnanarayana |
Speech Commun. | 1 |
| 2004 | Speaker Segmentation Based on Subsegmental Features and Neural Network Models
Dhananjaya Gowda, Sunitha Guruprasad, Bayya Yegnanarayana |
ICONIP | 1 |
| 2003 | AANN models for speaker recognition based on difference cepstralsabstractThis paper presents a novel method for representing speaker characteristics present in the speech signal, by the way of deemphasizing the linguistic content of the signal. Cepstral coefficients that are widely employed as features for automatic speaker recognition task, contain considerable speech information in addition to the speaker information, and hence do not highlight the latter. The proposed method is based on using the difference between all-pole spectra due to higher order and lower order of linear prediction analysis. Distribution of the feature vectors in the multi-dimensional feature space is captured by employing autoassociative neural network models. A speaker recognition system is developed using the proposed method of feature extraction, whose performance is evaluated against that of the system based on cepstral coefficients. The complementary nature of evidence due to the proposed feature is also examined, so as to improve the overall system performance. Sunitha Guruprasad, Dhananjaya Gowda, Bayya Yegnanarayana |
IJCNN | 2 |