VLDB 2026 Research / reviewers in the wild / expert
Anurag Chowdhury
dblp:192/4596
· DBLP profile ↗
12ranked-venue papers
8as first author
6since 2021 · last 2025
0000-0001-8758-988XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 4 since 2021Security and privacy · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Hybrid Approach to Combining Role Diarization with ASR for Professional Conversations
Bongjun Kim, Mark C. Fuhs, Anurag Chowdhury, Deblin Bagchi, Monika Woszczyna |
INTERSPEECH | 4 |
| 2024 | Investigating Confidence Estimation Measures for Speaker Diarization
Anurag Chowdhury, Abhinav Misra, Mark C. Fuhs, Monika Woszczyna |
INTERSPEECH | 1 |
| 2024 | Transcription-Free Fine-Tuning of Speech Separation Models for Noisy and Reverberant Multi-Speaker Automatic Speech RecognitionabstractOne solution to automatic speech recognition (ASR) of overlapping speakers is to separate speech and then perform ASR on the separated signals.Commonly, the separator produces artefacts which often degrade ASR performance.Addressing this issue typically requires reference transcriptions to jointly train the separation and ASR networks.This is often not viable for training on real-world in-domain audio where reference transcript information is not always available.This paper proposes a transcription-free method for joint training using only audio signals.The proposed method uses embedding differences of pre-trained ASR encoders as a loss with a proposed modification to permutation invariant training (PIT) called guided PIT (GPIT).The method achieves a 6.4% improvement in word error rate (WER) measures over a signal-level loss and also shows enhancement improvements in perceptual measures such as short-time objective intelligibility (STOI). William Ravenscroft, George Close, Stefan Goetze, Thomas Hain, Mohammad Soleymanpour, Anurag Chowdhury, Mark C. Fuhs |
INTERSPEECH | 6 |
| 2022 | Domain Adaptation for Speaker Recognition in Singing and Spoken VoiceabstractIn this work, we study the effect of speaking style and audio condition variability between the spoken and singing voice on speaker recognition performance. Furthermore, we also explore the utility of domain adaptation for bridging the gap between multiple speaking styles (singing versus spoken) and improving overall speaker recognition performance. In that regard, we first extend a publicly available singing voice dataset, JukeBox, with corresponding spoken voice data and refer to it as JukeBox-V2. Next, we use domain adaptation for developing a speaker recognition method robust to varying speaking styles and audio conditions. Finally, we analyze the speech embeddings of domain-adapted models to explain their generalizability across varying speaking styles and audio conditions. Anurag Chowdhury, Austin Cozzo, Arun Ross |
ICASSP | 1 |
| 2022 | Deducing health cues from biometric dataabstractMedical diagnosis involves the expert opinion of trained health care professionals based on causal inference from medical data. While medical data are typically collected using specialized medical-grade sensors, similar data characteristics useful for medical diagnosis are sometimes present in biometric data (e.g., face images, ocular images, and speech signals). In this paper, we explore the biometrics and medical literature to study the following questions. 1) What kind of health cues are embedded in the commonly utilized forms of audio-visual biometric data? 2) How can these health cues be gleaned from the biometric data, and what kind of diseases can it help diagnose? 3) What are some of the implications of using biometric data for medical diagnosis? Arun Ross, Sudipta Banerjee, Anurag Chowdhury |
Comput. Vis. Image Underst. | 3 |
| 2021 | DEEPTALK: Vocal Style Encoding for Speaker Recognition and Speech SynthesisabstractAutomatic speaker recognition algorithms typically characterize speech audio using short-term spectral features that encode the physiological and anatomical aspects of speech production. Such algorithms do not fully capitalize on speaker-dependent characteristics present in behavioral speech features. In this work, we propose a prosody encoding network called DeepTalk for extracting vocal style features directly from raw audio data. The DeepTalk method outperforms several state-of-the-art speaker recognition systems across multiple challenging datasets. The speaker recognition performance is further improved by combining DeepTalk with a state-of-the-art physiological speech feature-based speaker recognition system. We also integrate DeepTalk into a current state-of-the-art speech synthesizer to generate synthetic speech. A detailed analysis of the synthetic speech shows that the DeepTalk captures F0 contours essential for vocal style modeling. Furthermore, DeepTalk-based synthetic speech is shown to be almost indistinguishable from real speech in the context of speaker recognition. Anurag Chowdhury, Arun Ross, Prabu David |
ICASSP | 1 |
| 2020 | JukeBox: A Multilingual Singer Recognition DatasetabstractA text-independent speaker recognition system relies on successfully encoding speech factors such as vocal pitch, intensity, and timbre to achieve good performance. A majority of such systems are trained and evaluated using spoken voice or everyday conversational voice data. Spoken voice, however, exhibits a limited range of possible speaker dynamics, thus constraining the utility of the derived speaker recognition models. Singing voice, on the other hand, covers a broader range of vocal and ambient factors and can, therefore, be used to evaluate the robustness of a speaker recognition system. However, a majority of existing speaker recognition datasets only focus on the spoken voice. In comparison, there is a significant shortage of labeled singing voice data suitable for speaker recognition research. To address this issue, we assemble \textit{JukeBox} - a speaker recognition dataset with multilingual singing voice audio annotated with singer identity, gender, and language labels. We use the current state-of-the-art methods to demonstrate the difficulty of performing speaker recognition on singing voice using models trained on spoken voice alone. We also evaluate the effect of gender and language on speaker recognition performance, both in spoken and singing voice data. The complete \textit{JukeBox} dataset can be accessed at http://iprobe.cse.msu.edu/datasets/jukebox.html. Anurag Chowdhury, Austin Cozzo, Arun Ross |
INTERSPEECH | 1 |
| 2020 | Can a CNN Automatically Learn the Significance of Minutiae Points for Fingerprint Matching?abstractMost automated fingerprint recognition systems use minutiae points for comparing fingerprints. In the parlance of Computer Vision, minutiae can be viewed as handcrafted features, i.e., features that have been proposed by human experts for the task of fingerprint recognition. In this work, we raise the following question: Can a machine learning system automatically determine the significance of minutiae points for fingerprint matching? To this effect, a patch-based Siamese Convolutional Neural Network (CNN), which does not explicitly rely on the extraction of minutiae points, is designed and trained from scratch. The purpose of this network is to learn the most effective features for matching fingerprint images. The features learned by this network are analyzed using Gradient-weighted Class Activation Mapping (Grad-CAM) to determine if they correlate with the locations of minutiae points. Our experiments suggest that the proposed network automatically learns to focus on minutiae points, when available, for fingerprint matching. Thus, an automated learner without any explicit domain knowledge establishes the significance of minutiae points for fingerprint matching. Anurag Chowdhury, Simon Kirchgasser, Andreas Uhl, Arun Ross |
WACV | 1 |
| 2020 | Security in smart cities: A brief review of digital forensic schemes for biometric data
Arun Ross, Sudipta Banerjee, Anurag Chowdhury |
Pattern Recognit. Lett. | 3 |
| 2020 | Fusing MFCC and LPC Features Using 1D Triplet CNN for Speaker Recognition in Severely Degraded Audio SignalsabstractSpeaker recognition algorithms are negatively impacted by the quality of the input speech signal. In this work, we approach the problem of speaker recognition from severely degraded audio data by judiciously combining two commonly used features: Mel Frequency Cepstral Coefficients (MFCC) and Linear Predictive Coding (LPC). Our hypothesis rests on the observation that MFCC and LPC capture two distinct aspects of speech, viz., speech perception and speech production. A carefully crafted 1D Triplet Convolutional Neural Network (1D-Triplet-CNN) is used to combine these two features in a novel manner, thereby enhancing the performance of speaker recognition in challenging scenarios. Extensive evaluation on multiple datasets, different types of audio degradations, multi-lingual speech, varying length of audio samples, etc. convey the efficacy of the proposed approach over existing speaker recognition methods, including those based on iVector and xVector. Anurag Chowdhury, Arun Ross |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2018 | MSU-AVIS dataset: Fusing Face and Voice Modalities for Biometric Recognition in Indoor Surveillance VideosabstractIndoor video surveillance systems often use the face modality to establish the identity of a person of interest. However, the face image may not offer sufficient discriminatory information in many scenarios due to substantial variations in pose, illumination, expression, resolution and distance between the subject and the camera. In such cases, the inclusion of an additional biometric modality can benefit the recognition process. In this regard, we consider the fusion of voice and face modalities for enhancing the recognition accuracy. The main contribution of this work is assembling a multimodal (face and voice), semi-constrained, indoor video surveillance dataset referred to as the MSU Audio-Video Indoor Surveillance (MSU-AVIS) dataset. We use a consumer-grade camera with a built-in microphone to acquire data for this purpose. We use current state-of-art deep-learning based methods to perform face and speaker recognition on the collected dataset for establishing baseline performance. We also explore multiple fusion schemes to combine face and speaker recognition to perform effective person recognition on audio-video surveillance data. Experiments convey the efficacy of the proposed multimodal fusion scheme (face and voice) over unimodal approaches in surveillance scenarios. The collected dataset is being made available for research purposes. Anurag Chowdhury, Yousef Atoum, Luan Tran, Xiaoming Liu 0002, Arun Ross |
ICPR | 1 |
| 2017 | Extracting sub-glottal and Supra-glottal features from MFCC using convolutional neural networks for speaker identification in degraded audio signalsabstractWe present a deep learning based algorithm for speaker recognition from degraded audio signals. We use the commonly employed Mel-Frequency Cepstral Coefficients (MFCC) for representing the audio signals. A convolutional neural network (CNN) based on 1D filters, rather than 2D filters, is then designed. The filters in the CNN are designed to learn inter-dependency between cepstral coefficients extracted from audio frames of fixed temporal expanse. Our approach aims at extracting speaker dependent features, like Sub-glottal and Supra-glottal features, of the human speech production apparatus for identifying speakers from degraded audio signals. The performance of the proposed method is compared against existing baseline schemes on both synthetically and naturally corrupted speech data. Experiments convey the efficacy of the proposed architecture for speaker recognition. Anurag Chowdhury, Arun Ross |
IJCB | 1 |