EDBT 2026 Demo / reviewers in the wild / expert
Ganesh Sivaraman
dblp:50/5286
· DBLP profile ↗
22ranked-venue papers
7as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Investigating voiced and unvoiced regions of speech for audio deepfake detectionabstractDeep neural network based deepfake detection systems have achieved high levels of accuracy on benchmark datasets and competitions. However, most models lack interpretability. It is challenging to extract reasoning from the network that can convince the human evaluator to trust the decision. Humans often rely on acoustic cues like unnatural pitch jitter, robotic intonation, acoustic artifacts, and unnatural sounding fricatives to judge the quality of the synthetic audio. This study explores the role played by the voiced and unvoiced regions of speech in discriminating synthetic from bonafide speech. A measure of signal periodicity is used to analyze speech into voiced and unvoiced components. Then, the graph attention based AASIST detection system is trained independently on each component. This work compares the accuracy of deepfake detection system using voiced and unvoiced components and analyzes the results on the MLAAD dataset. Our results show that unvoiced regions are particularly more effective in distinguishing synthetic (deepfake) speech from bonafide, and achieves an equal error rate of 6.62%. When combined with voice regions through score-level fusion, the overall performance improves further, yielding a 5.82% EER, a relative improvement of 49% over the baseline system that uses the full audio. Ganesh Sivaraman, Hemlata Tak, Elie Khoury 0001 |
ICASSP | 1 |
| 2025 | Pindrop it! Audio and Visual Deepfake Countermeasures for Robust Detection and Fine-Grained LocalizationabstractThe field of visual and audio generation is burgeoning with new state-of-the-art methods. This rapid proliferation of new techniques underscores the need for robust solutions for detecting synthetic content in videos. In particular, when fine-grained alterations via localized manipulations are performed in visual, audio, or both domains, these subtle modifications add challenges to the detection algorithms. This paper presents solutions for the problems of deepfake video classification and localization. The methods were submitted to the ACM 1M Deepfakes Detection Challenge, achieving the best performance in the temporal localization task and a top four ranking in the classification task for the TestA split of the evaluation dataset. Nicholas Klein, Hemlata Tak, James Fullwood, Krishna Regmi, Leonidas Spinoulas, Ganesh Sivaraman, Elie Khoury 0001 |
ACM Multimedia | 6 |
| 2022 | Unsupervised Model Adaptation for End-to-End ASRabstractEnd-to-end (E2E) Automatic Speech Recognition (ASR) systems are widely applied in various devices and communication domains. However, state-of-the-art ASR systems are known to underperform when there is a mismatch in the training and test domains. As a result, acoustic models deployed in production are often adapted to the target domain to improve accuracy. This paper proposes a method to perform unsupervised model adaptation for E2E ASR using first-pass transcriptions of adaptation data produced by the baseline ASR model itself. The paper proposes two transcription confidence measures that can be used to select an optimal in-domain adaptation set. Experiments were performed using the Quartznet ASR architecture on the HarperValleyBank corpus. Results show that the unsupervised adaptation technique with the confidence measure based data selection results in a 8% absolute reduction in word error rate on the HarperValleyBank test set. The proposed method can be applied to any E2E ASR system and is suitable for model adaptation on call center audio with little to no manual transcription. Ganesh Sivaraman, Ricardo Casal, Matt Garland, Elie Khoury 0001 |
ICASSP | 1 |
| 2022 | Acoustic To Articulatory Speech Inversion Using Multi-Resolution Spectro-Temporal Representations Of Speech SignalsabstractMulti-resolution spectro-temporal features of a speech signal represent how the brain perceives sounds by tuning cortical cells to different spectral and temporal modulations. These features produce a higher dimensional representation of the speech signals. The purpose of this paper is to evaluate how well the auditory cortex representation of speech signals contribute to estimate articulatory features of those corresponding signals. Since obtaining articulatory features from acoustic features of speech signals has been a challenging topic of interest for different speech communities, we investigate the possibility of using this multi-resolution representation of speech signals as acoustic features. We used U. of Wisconsin X-ray Microbeam (XRMB) database of clean speech signals to train a feed-forward deep neural network (DNN) to estimate articulatory trajectories of six tract variables. The optimal set of multi-resolution spectro-temporal features to train the model were chosen using appropriate scale and rate vector parameters to obtain the best performing model. Experiments achieved a correlation of 0.675 with ground-truth tract variables. We compared the performance of this speech inversion system with prior experiments conducted using Mel Frequency Cepstral Coefficients (MFCCs). Rahil Parikh, Nadee Seneviratne, Ganesh Sivaraman, Shihab A. Shamma, Carol Y. Espy-Wilson |
INTERSPEECH | 3 |
| 2022 | Confidence Measure for Automatic Age Estimation From Speech
Amruta Saraf, Ganesh Sivaraman, Elie Khoury 0001 |
INTERSPEECH | 2 |
| 2022 | Acoustic-to-articulatory Speech Inversion with Multi-task LearningabstractMulti-task learning (MTL) frameworks have proven to be effective in diverse speech related tasks like automatic speech recognition (ASR) and speech emotion recognition. This paper proposes a MTL framework to perform acoustic-to-articulatory speech inversion by simultaneously learning an acoustic to phoneme mapping as a shared task. We use the Haskins Production Rate Comparison (HPRC) database which has both the electromagnetic articulography (EMA) data and the corresponding phonetic transcriptions. Performance of the system was measured by computing the correlation between estimated and actual tract variables (TVs) from the acoustic to articulatory speech inversion task. The proposed MTL based Bidirectional Gated Recurrent Neural Network (RNN) model learns to map the input acoustic features to nine TVs while outperforming the baseline model trained to perform only acoustic to articulatory inversion. Yashish M. Siriwardena, Ganesh Sivaraman, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2022 | A Zero-Shot Approach to Identifying Children's Speech in Automatic Gender ClassificationabstractDetecting whether a speech utterance belongs to an adult male, adult female or a child category, also known as male-female-child (MFC) classification is particularly challenging due to two main reasons - paucity of children's speech data, and high variability in children's speech due to developmental changes. It is difficult to obtain speech datasets with children's voices due to privacy reasons. This paper explores a zero-shot learning approach to MFC classification. Different algorithms are explored to create artificial childlike voices from adult voices. Methods such as pitch shifting, Vocal Tract Length Perturbation, and Segmental Warping Perturbation are used to create synthetic childlike speech for the MFC classification task. Speaker embeddings extracted from a DNN based speaker recognition system are used as features for MFC classification. Compared to a pitch frequency based baseline MFC classifier, the proposed method improves the child classification accuracy by 47%. Amruta Saraf, Ganesh Sivaraman, Elie Khoury 0001 |
SLT | 2 |
| 2021 | Proxima: accelerating the integration of machine learning in atomistic simulationsabstractAtomistic-scale simulations are prominent scientific applications that require the repetitive execution of a computationally expensive routine to calculate a system's potential energy. Prior work shows that these expensive routines can be replaced with a machine-learned surrogate approximation to accelerate the simulation at the expense of the overall accuracy. The exact balance of speed and accuracy depends on the specific configuration of the surrogate-modeling workflow and the science itself, and prior work leaves it up to the scientist to find a configuration that delivers the required accuracy for their science problem. Unfortunately, due to the underlying system dynamics, it is rare that a single surrogate configuration presents an optimal accuracy/latency trade-off for the entire simulation. In practice, scientists must choose conservative configurations so that accuracy is always acceptable, forgoing possible acceleration. As an alternative, we propose Proxima, a systematic and automated method for dynamically tuning a surrogate-modeling configuration in response to real-time feedback from the ongoing simulation. Proxima estimates the uncertainty of applying a surrogate approximation in each step of an iterative simulation. Using this information, the specific surrogate configuration can be adjusted dynamically to ensure maximum speedup while sustaining a required accuracy metric. We evaluate Proxima using a Monte Carlo sampling application and find that Proxima respects a wide range of user-defined accuracy goals while achieving speedups of 1.02--5.5X relative to a standard Yuliana Zamora, Logan T. Ward, Ganesh Sivaraman, Ian T. Foster, Henry Hoffmann |
ICS | 3 |
| 2019 | Pindrop Labs' Submission to the First Multi-Target Speaker Detection and Identification Challenge
Elie Khoury 0001, Khaled Lakhdhar, Andrew Vaughan, Ganesh Sivaraman, Parav Nagarsheth |
INTERSPEECH | 4 |
| 2019 | Multi-Corpus Acoustic-to-Articulatory Speech Inversion
Nadee Seneviratne, Ganesh Sivaraman, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2019 | Articulatory and bottleneck features for speaker-independent ASR of dysarthric speech
Emre Yilmaz 0001, Vikramjit Mitra, Ganesh Sivaraman, Horacio Franco |
Comput. Speech Lang. | 3 |
| 2018 | Smoothing Model Predictions Using Adversarial Training Procedures for Speech Based Emotion RecognitionabstractTraining discriminative classifiers involves learning a conditional distribution p(yi|xi), given a set of feature vectors xiand the corresponding labels yi, i=1...N. For a classifier to be generalizable and not overfit to training data, the resulting conditional distribution p(yi|xi) is desired to be smoothly varying over the inputs xi. Adversarial training procedures enforce this smoothness using manifold regularization techniques. Manifold regularization makes the model's output distribution more robust to local perturbation added to a datapoint xi. In this paper, we experiment with the application of adversarial training procedures to increase the accuracy of a deep neural network based emotion recognition system using speech cues. Specifically, we investigate two training procedures: (i) adversarial training where we determine the adversarial direction based on the given labels for the training data and, (ii) virtual adversarial training where we determine the adversarial direction based only on the output distribution of the training data. We demonstrate the efficacy of adversarial training procedures by performing a k-fold cross validation experiment on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) and a cross-corpus performance analysis on three separate corpora. Results show improvement over a purely supervised approach, as well as better generalization capability to cross-corpus settings. Saurabh Sahu, Rahul Gupta 0001, Ganesh Sivaraman, Carol Y. Espy-Wilson |
ICASSP | 3 |
| 2018 | Noise Robust Acoustic to Articulatory Speech Inversion
Nadee Seneviratne, Ganesh Sivaraman, Vikramjit Mitra, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2018 | Speech Synthesis in the Wild
Ganesh Sivaraman, Parav Nagarsheth, Elie Khoury 0001 |
INTERSPEECH | 1 |
| 2017 | Joint modeling of articulatory and acoustic spaces for continuous speech recognition tasksabstractArticulatory information can effectively model variability in speech and can improve speech recognition performance under varying acoustic conditions. Learning speaker-independent articulatory models has always been challenging, as speaker-specific information in the articulatory and acoustic spaces increases the complexity of the speech-to-articulatory space inverse modeling, which is already an ill-posed problem due to its inherent nonlinearity and non-uniqueness. This paper investigates using deep neural networks (DNN) and convolutional neural networks (CNNs) for mapping speech data into its corresponding articulatory space. Our results indicate that the CNN models perform better than their DNN counterparts for speech inversion. In addition, we used the inverse models to generate articulatory trajectories from speech for three different standard speech recognition tasks. To effectively model the articulatory features' temporal modulations while retaining the acoustic features' spatiotemporal signatures, we explored a joint modeling strategy to simultaneously learn both the acoustic and articulatory spaces. The results from multiple speech recognition tasks indicate that articulatory features can improve recognition performance when the acoustic and articulatory spaces are jointly learned with one common objective function. Vikramjit Mitra, Ganesh Sivaraman, Chris Bartels, Hosung Nam, Wen Wang 0001, Carol Y. Espy-Wilson, Dimitra Vergyri, Horacio Franco |
ICASSP | 2 |
| 2017 | Adversarial Auto-Encoders for Speech Based Emotion RecognitionabstractRecently, generative adversarial networks and adversarial autoencoders have gained a lot of attention in machine learning community due to their exceptional performance in tasks such as digit classification and face recognition. They map the autoencoder's bottleneck layer output (termed as code vectors) to different noise Probability Distribution Functions (PDFs), that can be further regularized to cluster based on class information. In addition, they also allow a generation of synthetic samples by sampling the code vectors from the mapped PDFs. Inspired by these properties, we investigate the application of adversarial autoencoders to the domain of emotion recognition. Specifically, we conduct experiments on the following two aspects: (i) their ability to encode high dimensional feature vector representations for emotional utterances into a compressed space (with a minimal loss of emotion class discriminability in the compressed space), and (ii) their ability to regenerate synthetic samples in the original feature space, to be later used for purposes such as training emotion recognition classifiers. We demonstrate the promise of adversarial autoencoders with regards to these aspects on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) corpus and present our analysis. Saurabh Sahu, Rahul Gupta 0001, Ganesh Sivaraman, Wael Abd-Almageed, Carol Y. Espy-Wilson |
INTERSPEECH | 3 |
| 2017 | Analysis of Acoustic-to-Articulatory Speech Inversion Across Different Accents and Languages
Ganesh Sivaraman, Carol Y. Espy-Wilson, Martijn Wieling 0001 |
INTERSPEECH | 1 |
| 2017 | Hybrid convolutional neural networks for articulatory and acoustic information based speech recognition
Vikramjit Mitra, Ganesh Sivaraman, Hosung Nam, Carol Y. Espy-Wilson, Elliot Saltzman, Mark K. Tiede |
Speech Commun. | 2 |
| 2016 | Vocal Tract Length Normalization for Speaker Independent Acoustic-to-Articulatory Speech Inversion
Ganesh Sivaraman, Vikramjit Mitra, Hosung Nam, Mark K. Tiede, Carol Y. Espy-Wilson |
INTERSPEECH | 1 |
| 2015 | Analysis of coarticulated speech using estimated articulatory trajectories
Ganesh Sivaraman, Vikramjit Mitra, Mark K. Tiede, Elliot Saltzman, Louis Goldstein, Carol Y. Espy-Wilson |
INTERSPEECH | 1 |
| 2014 | Articulatory features from deep neural networks and their role in speech recognitionabstractThis paper presents a deep neural network (DNN) to extract articulatory information from the speech signal and explores different ways to use such information in a continuous speech recognition task. The DNN was trained to estimate articulatory trajectories from input speech, where the training data is a corpus of synthetic English words generated by the Haskins Laboratories' task-dynamic model of speech production. Speech parameterized as cepstral features were used to train the DNN, where we explored different cepstral features to observe their role in the accuracy of articulatory trajectory estimation. The best feature was used to train the final DNN system, where the system was used to predict articulatory trajectories for the training and testing set of Aurora-4, the noisy Wall Street Journal (WSJ0) corpus. This study also explored the use of hidden variables in the DNN pipeline as a potential acoustic feature candidate for speech recognition and the results were encouraging. Word recognition results from Aurora-4 indicate that the articulatory features from the DNN provide improvement in speech recognition performance when fused with other standard cepstral features; however when tried by themselves, they failed to match the baseline performance. Vikramjit Mitra, Ganesh Sivaraman, Hosung Nam, Carol Y. Espy-Wilson, Elliot Saltzman |
ICASSP | 2 |
| 2001 | System Software For Digital Television ApplicationsabstractInteractive Television is fast becoming a necessity as it converges the popular web browsing and the standard television systems better. This paper discusses the underlying system - Operating system and Java Runtime Environment - for the Digital TV. A review of the needed system capabilities for Digital TV, a probable solution of the underlying system, and future improvisation of the system are dealt herewith. Ganesh Sivaraman, Pablo César, Petri Vuorimaa |
ICME | 1 |