EDBT 2026 Demo / reviewers in the wild / expert
Saurabh Kataria 0001
dblp:22/6491-1
· DBLP profile ↗
15ranked-venue papers
8as first author
10since 2021 · last 2024
0000-0001-6311-2155ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Time-Domain Speech Super-Resolution With GAN Based Modeling for Telephony Speaker VerificationabstractAutomatic Speaker Verification(ASV) technology has become commonplace in virtual assistants. However, its performance suffers when there is a mismatch between the train and test domains. Mixed bandwidth training, i.e., pooling training data from both domains, is a preferred choice for developing a universal model that works for both narrowband and wideband domains. We propose complementing this technique by performing neural upsampling of narrowband signals, also known as bandwidth extension. We aim to discover and analyze high-performing time-domain Generative Adversarial Network (GAN) based models to improve our downstream state-of-the-art ASV system. We choose GANs since they 1) are powerful for learning conditional distribution and 2) allow flexibleplug-inusage as a pre-processor during the training of downstream tasks (ASV) with data augmentation. Prior work mainly focused on feature-domain bandwidth extension and limited experimental setups. We address these limitations by 1) using time-domain extension models, 2) reporting results on three real test sets, 3) extending training data, and 4) devising new test-time schemes. We compare supervised (conditional GAN) and unsupervised GANs (CycleGAN) and demonstrate an average relative improvement in the equal error rate of 8.6% and 7.7%, respectively. For further analysis, we study changes in the visual quality of the spectrogram, audio perceptual quality, t-SNE embeddings, and ASV score distributions. We show that our bandwidth extension leads to phenomena such as a shift of telephone (test) embeddings towards wideband (train) signals, a negative correlation of perceptual quality with downstream performance, and condition-independent score calibration. Saurabh Kataria 0001, Jesús Villalba 0001, Laureano Moro-Velázquez, Piotr Zelasko, Najim Dehak |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Advances in Language Recognition in Low Resource African Languages: The JHU-MIT Submission for NIST LRE22
Jesús Villalba 0001, Jonas Borgstrom, Maliha Jahan, Saurabh Kataria 0001, L. Paola García-Perera, Pedro A. Torres-Carrasquillo, Najim Dehak |
INTERSPEECH | 4 |
| 2023 | Self-FiLM: Conditioning GANs with self-supervised representations for bandwidth extension based speaker recognitionabstractSpeech super-resolution/Bandwidth Extension (BWE) can improve downstream tasks like Automatic Speaker Verification (ASV).We introduce a simple novel technique called Self-FiLM to inject self-supervision into existing BWE models via Feature-wise Linear Modulation.We hypothesize that such information captures domain/environment information, which can give zero-shot generalization.Self-FiLM Conditional GAN (CGAN) gives 18% relative improvement in Equal Error Rate and 8.5% in minimum Decision Cost Function using state-ofthe-art ASV system on SRE21 test.We further by 1) deep feature loss from time-domain models and 2) re-training of data2vec 2.0 models on naturalistic wideband (VoxCeleb) and telephone data (SRE Superset etc.).Lastly, we integrate selfsupervision with CycleGAN to present a completely unsupervised solution that matches the semi-supervised performance. Saurabh Kataria 0001, Jesús Villalba 0001, Laureano Moro-Velázquez, Thomas Thebaud, Najim Dehak |
INTERSPEECH | 1 |
| 2022 | Defense against Adversarial Attacks on Hybrid Speech Recognition System using Adversarial Fine-tuning with Denoiser
Sonal Joshi, Saurabh Kataria 0001, Yiwen Shao, Piotr Zelasko, Jesús Villalba 0001, Sanjeev Khudanpur, Najim Dehak |
INTERSPEECH | 2 |
| 2022 | AdvEst: Adversarial Perturbation Estimation to Classify and Detect Adversarial Attacks against Speaker IdentificationabstractAdversarial attacks pose a severe security threat to the state-ofthe-art speaker identification systems, thereby making it vital to propose countermeasures against them.Building on our previous work that used representation learning to classify and detect adversarial attacks, we propose an improvement to it using Ad-vEst, a method to estimate adversarial perturbation.First, we prove our claim that training the representation learning network using adversarial perturbations as opposed to adversarial examples (consisting of the combination of clean signal and adversarial perturbation) is beneficial because it eliminates nuisance information.At inference time, we use a time-domain denoiser to estimate the adversarial perturbations from adversarial examples.Using our improved representation learning approach to obtain attack embeddings (signatures), we evaluate their performance for three applications: known attack classification, attack verification, and unknown attack detection.We show that common attacks in the literature (Fast Gradient Sign Method (FGSM), Projected Gradient Descent (PGD), Carlini-Wagner (CW) with different Lp threat models) can be classified with an accuracy of ∼ 96%.We also detect unknown attacks with an equal error rate (EER) of ∼9%, which is absolute improvement of ∼12% from our previous work. Sonal Joshi, Saurabh Kataria 0001, Jesús Villalba 0001, Najim Dehak |
INTERSPEECH | 2 |
| 2022 | Joint domain adaptation and speech bandwidth extension using time-domain GANs for speaker verificationabstractSpeech systems developed for a particular choice of acoustic domain and sampling frequency do not translate easily to others.The usual practice is to learn domain adaptation and bandwidth extension models independently.Contrary to this, we propose to learn both tasks together.Particularly, we learn to map narrowband conversational telephone speech to wideband microphone speech.We developed parallel and non-parallel learning solutions which utilize both paired and unpaired data.First, we first discuss joint and disjoint training of multiple generative models for our tasks.Then, we propose a two-stage learning solution where we use a pre-trained domain adaptation system for pre-processing in bandwidth extension training.We evaluated our schemes on a Speaker Verification downstream task.We used the JHU-MIT experimental setup for NIST SRE21, which comprises SRE16, SRE-CTS Superset and SRE21.Our results provide the first evidence that learning both tasks is better than learning just one.On SRE16, our best system achieves 22% relative improvement in Equal Error Rate w.r.t. a direct learning baseline and 8% w.r.t. a strong bandwidth expansion system. Saurabh Kataria 0001, Jesús Villalba 0001, Laureano Moro-Velázquez, Najim Dehak |
INTERSPEECH | 1 |
| 2022 | Chunking Defense for Adversarial Attacks on ASR
Yiwen Shao, Jesús Villalba 0001, Sonal Joshi, Saurabh Kataria 0001, Sanjeev Khudanpur, Najim Dehak |
INTERSPEECH | 4 |
| 2021 | Perceptual Loss Based Speech Denoising with an Ensemble of Audio Pattern Recognition and Self-Supervised ModelsabstractDeep learning based speech denoising still suffers from the challenge of improving perceptual quality of enhanced signals. We introduce a generalized framework called Perceptual Ensemble Regularization Loss (PERL) built on the idea of perceptual losses. Perceptual loss discourages distortion to certain speech properties and we analyze it using six large-scale pre-trained models: speaker classification, acoustic model, speaker embedding, emotion classification, and two self-supervised speech encoders (PASE+, wav2vec 2.0). We first build a strong baseline (w/o PERL) using Conformer Transformer Networks on the popular enhancement benchmark called VCTK-DEMAND. Using auxiliary models one at a time, we find acoustic event and self-supervised model PASE+ to be most effective. Our best model (PERL-AE) only uses acoustic event model (utilizing AudioSet) to outperform state-of-the-art methods on major perceptual metrics. To explore if denoising can leverage full framework, we use all networks but find that our seven-loss formulation suffers from the challenges of Multi-Task Learning. Finally, we report a critical observation that state-of-the-art Multi-Task weight learning methods cannot outperform hand tuning, perhaps due to challenges of domain mismatch and weak complementarity of losses. Saurabh Kataria 0001, Jesús Villalba 0001, Najim Dehak |
ICASSP | 1 |
| 2021 | Deep Feature CycleGANs: Speaker Identity Preserving Non-Parallel Microphone-Telephone Domain Adaptation for Speaker VerificationabstractWith the increase in the availability of speech from varied domains, it is imperative to use such out-of-domain data to improve existing speech systems. Domain adaptation is a prominent pre-processing approach for this. We investigate it for adapt microphone speech to the telephone domain. Specifically, we explore CycleGAN-based unpaired translation of microphone data to improve the x-vector/speaker embedding network for Telephony Speaker Verification. We first demonstrate the efficacy of this on real challenging data and then, to improve further, we modify the CycleGAN formulation to make the adaptation task-specific. We modify CycleGAN's identity loss, cycle-consistency loss, and adversarial loss to operate in the deep feature space. Deep features of a signal are extracted from an auxiliary (speaker embedding) network and, hence, preserves speaker identity. Our 3D convolution-based Deep Feature Discriminators (DFD) show relative improvements of 5-10% in terms of equal error rate. To dive deeper, we study a challenging scenario of pooling (adapted) microphone and telephone data with data augmentations and telephone codecs. Finally, we highlight the sensitivity of CycleGAN hyper-parameters and introduce a parameter called probability of adaptation. Saurabh Kataria 0001, Jesús Villalba 0001, Piotr Zelasko, Laureano Moro-Velázquez, Najim Dehak |
Interspeech | 1 |
| 2021 | Multi-Channel Speaker Verification for Single and Multi-Talker SpeechabstractTo improve speaker verification in real scenarios with interference speakers, noise, and reverberation, we propose to bring together advancements made in multi-channel speech features.Specifically, we combine spectral, spatial, and directional features, which includes inter-channel phase difference, multichannel sinc convolutions, directional power ratio features, and angle features.To maximally leverage supervised learning, our framework is also equipped with multi-channel speech enhancement and voice activity detection.On all simulated, replayed, and real recordings, we observe large and consistent improvements at various degradation levels.On real recordings of multi-talker speech, we achieve a 36% relative reduction in equal error rate w.r.t.single-channel baseline.We find the improvements from speaker-dependent directional features more consistent in multi-talker conditions than clean.Lastly, we investigate if the learned multi-channel speaker embedding space can be made more discriminative through a contrastive loss-based fine-tuning.With a simple choice of Triplet loss, we observe a further 8.3% relative reduction in EER. Saurabh Kataria 0001, Shixiong Zhang 0001, Dong Yu 0001 |
Interspeech | 1 |
| 2020 | Feature Enhancement with Deep Feature Losses for Speaker VerificationabstractSpeaker Verification still suffers from the challenge of generalization to novel adverse environments. We leverage on the recent advancements made by deep learning based speech enhancement and propose a feature-domain supervised denoising based solution. We propose to use Deep Feature Loss which optimizes the enhancement network in the hidden activation space of a pre-trained auxiliary speaker embedding network. We experimentally verify the approach on simulated and real data. A simulated testing setup is created using various noise types at different SNR levels. For evaluation on real data, we choose BabyTrain corpus which consists of children recordings in uncontrolled environments. We observe consistent gains in every condition over the state-of-the-art augmented Factorized-TDNN x-vector system. On BabyTrain corpus, we observe relative gains of 10.38% and 12.40% in minDCF and EER respectively. Saurabh Kataria 0001, Phani S. Nidadavolu, Jesús Villalba 0001, Nanxin Chen, L. Paola García-Perera, Najim Dehak |
ICASSP | 1 |
| 2020 | Unsupervised Feature Enhancement for Speaker VerificationabstractThe task of making speaker verification systems robust to adverse scenarios remains a challenging and an active area of research. We developed an unsupervised feature enhancement approach in log-filter bank space with the end goal of improving speaker verification performance. We experimented with using both real speech recorded in adverse environments and degraded speech obtained by simulation to train the enhancement systems. The effectiveness of this approach was shown by testing on several real, simulated noisy, and reverberant test sets. The approach yielded significant improvements on both real and simulated sets when data augmentation was not used in speaker verification pipeline. We also experimented with training the x-vector and PLDA systems with enhanced augmented features instead of augmented features and observed better performance on real test conditions (4.2% relative improvement in minDCF on SRI). Phani S. Nidadavolu, Saurabh Kataria 0001, Jesús Villalba 0001, L. Paola García-Perera, Najim Dehak |
ICASSP | 2 |
| 2019 | Low-Resource Domain Adaptation for Speaker Recognition Using Cycle-GansabstractCurrent speaker recognition technology provides great performance with the x-vector approach. However, performance decreases when the evaluation domain is different from the training domain, an issue usually addressed with domain adaptation approaches. Recently, unsupervised domain adaptation using cycle-consistent Generative Adversarial Networks (CycleGAN) has received a lot of attention. Cycle-GAN learn mappings between features of two domains given non-parallel data. We investigate their effectiveness in low resource scenario i.e. when limited amount of target domain data is available for adaptation, a case unexplored in previous works. We experiment with two adaptation tasks: microphone to telephone and a novel reverberant to clean adaptation with the end goal of improving speaker recognition performance. Number of speakers present in source and target domains are 7000 and 191 respectively. By adding noise to the target domain during CycleGAN training, we were able to achieve better performance compared to the adaptation system whose CycleGAN was trained on a larger target data. On reverberant to clean adaptation task, our models improved EER by 18.3% relative on VOiCES dataset compared to a system trained on clean data. They also slightly improved over the state-of-the-art Weighted Prediction Error (WPE) de-reverberation algorithm. Phani S. Nidadavolu, Saurabh Kataria 0001, Jesús Villalba 0001, Najim Dehak |
ASRU | 2 |
| 2017 | Hearing in a shoe-box: Binaural source position and wall absorption estimation using virtually supervised learningabstractThis paper introduces a new framework for supervised sound source localization referred to as virtually-supervised learning. An acoustic shoe-box room simulator is used to generate a large number of binaural single-source audio scenes. These scenes are used to build a dataset of spatial binaural features annotated with acoustic properties such as the 3D source position and the walls' absorption coefficients. A probabilistic high- to low-dimensional regression framework is used to learn a mapping from these features to the acoustic properties. Results indicate that this mapping successfully estimates the azimuth and elevation of new sources, but also their range and even the walls' absorption coefficients solely based on binaural signals. Results also reveal that incorporating random-diffusion effects in the data significantly improves the estimation of all parameters. Saurabh Kataria 0001, Clément Gaultier, Antoine Deleforge |
ICASSP | 1 |
| 2015 | Representation and modeling of spherical harmonics manifold for source localizationabstractSource localization has been studied in the spatial domain using differential geometry in earlier work. However, parameters of the sensor array manifold have hitherto not been investigated for source localization in spherical harmonics domain. The objective of this work is to represent and model the manifold surface using differential geometry. The system model for source localization over a spherical harmonic manifold is first formulated. Subsequently, the manifold parameters are modeled in the spherical harmonics domain. Source localization methods using MUSIC and MVDR over the spherical harmonics manifold are developed. Experiments on source localization using a spherical microphone array indicate high resolution in noise. Arun Parthasarathy, Saurabh Kataria 0001, Lalan Kumar, Rajesh M. Hegde |
ICASSP | 2 |