VLDB 2026 Research / reviewers in the wild / expert
Rohan Kumar Das
dblp:158/4101
· DBLP profile ↗
55ranked-venue papers
11as first author
26since 2021 · last 2025
0000-0002-1332-3357ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 49 · 10 first-author · 23 since 2021Artificial intelligence and machine learning · 32 · 8 first-author · 10 since 2021Security and privacy · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-modal Speech Enhancement with Limited Electromyography ChannelsabstractSpeech enhancement (SE) aims to improve the clarity, intelligibility, and quality of speech signals for various speech enabled applications. However, air-conducted (AC) speech is highly susceptible to ambient noise, particularly in low signal-to-noise ratio (SNR) and non-stationary noise environments. Incorporating multi-modal information has shown promise in enhancing speech in such challenging scenarios. Electromyography (EMG) signals, which capture muscle activity during speech production, offer noise-resistant properties beneficial for SE in adverse conditions. Most previous EMG-based SE methods required 35 EMG channels, limiting their practicality. To address this, we propose a novel method that considers only 8-channel EMG signals with acoustic signals using a modified SEMamba network with added cross-modality modules. Our experiments demonstrate substantial improvements in speech quality and intelligibility over traditional approaches, especially in extremely low SNR settings. Notably, compared to the SE (AC) approach, our method achieves a significant PESQ gain of 0.235 under matched low SNR conditions and 0.527 under mismatched conditions, highlighting its robustness. Fuyuan Feng, Longting Xu, Rohan Kumar Das |
ICASSP | 3 |
| 2025 | UCIL: An Unsupervised Class Incremental Learning Approach for Sound Event DetectionabstractThis work explores class-incremental learning (CIL) for sound event detection (SED), advancing adaptability towards real-world scenarios. CIL’s success in domains like computer vision inspired our SED-tailored method, addressing the unique challenges of diverse and complex audio environments. Our approach employs an independent unsupervised learning frame-work with a distillation loss function to integrate new sound classes while preserving the SED model consistency across incremental tasks. We further enhance this framework with a sample selection strategy for unlabeled data and a balanced exemplar update mechanism, ensuring varied and illustrative sound representations. Evaluating various continual learning methods on the DCASE 2023 Task 4 dataset, our research offers insights into each method’s applicability for real-world SED systems that can have newly added sound classes. The findings also delineate future directions of CIL in dynamic audio settings. Yang Xiao 0019, Rohan Kumar Das |
ICASSP | 2 |
| 2025 | Exploring Text-Queried Sound Event Detection with Audio Source SeparationabstractIn sound event detection (SED), overlapping sound events pose a significant challenge, as certain events can be easily masked by background noise or other events, resulting in poor detection performance. To address this issue, we propose the text-queried SED (TQ-SED) framework. Specifically, we first pre-train a language-queried audio source separation (LASS) model to separate the audio tracks corresponding to different events from the input audio. Then, multiple target SED branches are employed to detect individual events. AudioSep is a state-of-the-art LASS model, but has limitations in extracting dynamic audio information because of its pure convolutional structure for separation. To address this, we integrate a dual-path recurrent neural network block into the model. We refer to this structure as AudioSep-DP, which achieves the first place in DCASE 2024 Task 9 on language-queried audio source separation (objective single model track). Experimental results show that TQ-SED can significantly improve the SED performance, with an improvement of 7.22% on F1 score over the conventional framework. Additionally, we setup comprehensive experiments to explore the impact of model complexity. The source code and pre-trained model are released at https://github.com/apple-yinhan/TQ-SED. Han Yin, Jisheng Bai, Yang Xiao 0019, Hui Wang 0030, Yafeng Chen, Rohan Kumar Das, Chong Deng |
ICASSP | 7 |
| 2025 | Where's That Voice Coming? Continual Learning for Sound Source LocalizationabstractSound source localization (SSL) is essential for many speech-processing applications. Deep learning models have achieved high performance, but often fail when the training and inference environments differ. Adapting SSL models to dynamic acoustic conditions faces a major challenge: catastrophic forgetting. In this work, we propose an exemplar-free continual learning strategy for SSL (CL-SSL) to address such a forgetting phenomenon. CL-SSL applies task-specific sub-networks to adapt across diverse acoustic environments while retaining previously learned knowledge. It also uses a scaling mechanism to limit parameter growth, ensuring consistent performance across incremental tasks. We evaluated CL-SSL on simulated data with varying microphone distances and real-world data with different noise levels. The results demonstrate CL-SSL’s ability to maintain high accuracy with minimal parameter increase, offering an efficient solution for SSL applications. Yang Xiao 0019, Rohan Kumar Das |
ICME | 2 |
| 2025 | Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource LanguagesabstractAutomatic speech recognition (ASR) for dysarthric speech remains challenging due to data scarcity, particularly in non-English languages. To address this, we fine-tune a voice conversion model on English dysarthric speech (UASpeech) to encode both speaker characteristics and prosodic distortions, then apply it to convert healthy non-English speech (FLEURS) into non-English dysarthric-like speech. The generated data is then used to fine-tune a multilingual ASR model, Massively Multilingual Speech (MMS), for improved dysarthric speech recognition. Evaluation on PC-GITA (Spanish), EasyCall (Italian), and SSNCE (Tamil) demonstrates that VC with both speaker and prosody conversion significantly outperforms the off-the-shelf MMS performance and conventional augmentation techniques such as speed and tempo perturbation. Objective and subjective analyses of the generated data further confirm that the generated speech simulates dysarthric characteristics. Chin-Jou Li, Eunjung Yeo, Kwanghee Choi, Paula Andrea Pérez-Toro, Masao Someki, Rohan Kumar Das, Zhengjun Yue, Juan Rafael Orozco-Arroyave, Elmar Nöth, David R. Mortensen |
INTERSPEECH | 6 |
| 2025 | TF-Mamba: A Time-Frequency Network for Sound Source Localization
Yang Xiao 0019, Rohan Kumar Das |
INTERSPEECH | 2 |
| 2025 | Listen, Analyze, and Adapt to Learn New Attacks: An Exemplar-Free Class Incremental Learning Method for Audio Deepfake Source Tracing
Yang Xiao 0019, Rohan Kumar Das |
INTERSPEECH | 2 |
| 2025 | AdaKWS: Towards Robust Keyword Spotting with Test-Time Adaptation
Yang Xiao 0019, Tianyi Peng, Yanghao Zhou, Rohan Kumar Das |
INTERSPEECH | 4 |
| 2025 | EnvSDD: Benchmarking Environmental Sound Deepfake Detection
Han Yin, Yang Xiao 0019, Rohan Kumar Das, Jisheng Bai, Haohe Liu, Wenwu Wang 0001, Mark D. Plumbley |
INTERSPEECH | 3 |
| 2025 | XLSR-Mamba: A Dual-Column Bidirectional State Space Model for Spoofing Attack DetectionabstractTransformers and their variants have achieved great success in speech processing. However, their multi-head self-attention mechanism is computationally expensive. Therefore, one novel selective state space model, Mamba, has been proposed as an alternative. Building on its success in automatic speech recognition, we apply Mamba for spoofing attack detection. Mamba is well-suited for this task as it can capture the artifacts in spoofed speech signals by handling long-length sequences. However, Mamba's performance may suffer when it is trained with limited labeled data. To mitigate this, we propose combining a new structure of Mamba based on a dual-column architecture with self-supervised learning, using the pre-trained wav2vec 2.0 model. The experiments show that our proposed approach achieves competitive results and faster inference on the ASVspoof 2021 LA and DF datasets, and on the more challenging In-the-Wild dataset, it emerges as the strongest candidate for spoofing attack detection. Yang Xiao 0019, Rohan Kumar Das |
IEEE Signal Process. Lett. | 2 |
| 2025 | Nes2Net: A Lightweight Nested Architecture for Foundation Model Driven Speech Anti-SpoofingabstractSpeech foundation models have significantly advanced various speech-related tasks by providing exceptional representation capabilities. However, their high-dimensional output features often create a mismatch with downstream task models, which typically require lower-dimensional inputs. A common solution is to apply a dimensionality reduction (DR) layer, but this approach increases parameter overhead, computational costs, and risks losing valuable information. To address these issues, we propose Nested Res2Net (Nes2Net), a lightweight back-end architecture designed to directly process high-dimensional features without DR layers. The nested structure enhances multi-scale feature extraction, improves feature interaction, and preserves high-dimensional information. We first validate Nes2Net on CtrSVDD, a singing voice deepfake detection dataset, and report a 22% performance improvement and an 87% back-end computational cost reduction over the state-of-the-art baseline. Additionally, extensive testing across four diverse datasets: ASVspoof 2021, ASVspoof 5, PartialSpoof, and In-the-Wild, covering fully spoofed speech, adversarial attacks, partial spoofing, and real-world scenarios, consistently highlights Nes2Net’s superior robustness and generalization capabilities. The code package and pre-trained models are available at https://github.com/Liu-Tianchi/Nes2Net. Tianchi Liu 0004, Duc-Tuan Truong, Rohan Kumar Das, Kong-Aik Lee, Haizhou Li 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Enhancing Real-World Active Speaker Detection With Multi-Modal Extraction Pre-TrainingabstractAudio-visual active speaker detection (AV-ASD) aims to identify which visible face is speaking in a scene with one or more persons. Most existing AV-ASD methods prioritize capturing speech-lip correspondence. However, there is a noticeable gap in addressing the challenges from real-world AV-ASD scenarios. Due to the presence of low-quality noisy videos in such cases, AV-ASD systems without a selective listening ability are short of effectively filtering out disruptive voice components from mixed audio inputs. In this paper, we propose a Multi-modal Speech Extraction-to-Detection framework named ‘MuSED’, which is pre-trained with audio-visual target speech extraction to learn the denoising ability, then it is fine-tuned with the AV-ASD task. Meanwhile, to better capture the multi-modal information and deal with real-world problems such as missing modality, MuSED is modelled on the time domain directly and integrates the multi-modal plus-and-minus augmentation strategy. Our experiments demonstrate that MuSED substantially outperforms the state-of-the-art AV-ASD methods and achieves 95.6% mAP on the AVA-ActiveSpeaker dataset, 98.3% AP on the ASW dataset, and 97.9% F1 on the Columbia AV-ASD dataset, respectively. We will publicly release the code in due course. Ruijie Tao, Xinyuan Qian 0001, Rohan Kumar Das, Xiaoxue Gao, Haizhou Li 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Adaptive-Avg-Pooling Based Attention Vision Transformer for Face Anti-SpoofingabstractTraditional vision transformer consists of two parts: transformer encoder and multi-layer perception (MLP). The former plays the role of feature learning to obtain better representation, while the latter plays the role of classification. Here, the MLP is constituted of two fully connected (FC) layers, average value computing, FC layer and softmax layer. However, due to the use of average value computing module, some useful information may get lost, which we plan to preserve by the use of alternative framework. In this work, we propose a novel vision transformer referred to as adaptive-avg-pooling based attention vision transformer (AAViT) that uses modules of adaptive average pooling and attention to replace the module of average value computing. We explore the proposed AAViT for the studies on face anti-spoofing using Replay-Attack database. The experiments show that the AAViT outperforms vision transformer in face anti-spoofing by producing a reduced equal error rate. In addition, we found that the proposed AAViT can perform much better than some commonly used neural networks such as ResNet and some other known systems on the Replay-Attack corpus. Fangfan Chen, Rohan Kumar Das, Shunsi Zhang |
ICASSP | 3 |
| 2024 | How Do Neural Spoofing Countermeasures Detect Partially Spoofed Audio?
Tianchi Liu 0004, Lin Zhang 0054, Rohan Kumar Das, Ruijie Tao, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2024 | A Synopsis of FAME 2024 Challenge: Associating Faces with Voices in Multilingual Environments
Muhammad Saad Saeed, Shah Nawaz, Marta Moscati, Rohan Kumar Das, Muhammad Salman Tahir, Muhammad Zaigham Zaheer, Muhammad Irzam Liaqat, Muhammad Haris Khan, Karthik Nandakumar, Muhammad Haroon Yousaf, Markus Schedl |
ACM Multimedia | 4 |
| 2023 | A Multi-Task Learning Framework for Sound Event Detection using High-level Acoustic Characteristics of Sounds
Tanmay Khandelwal, Rohan Kumar Das |
INTERSPEECH | 2 |
| 2023 | Self-Supervised Training of Speaker Encoder With Multi-Modal Diverse Positive PairsabstractWe study a novel neural speaker encoder and its training strategies for speaker recognition without using any identity labels. The speaker encoder is trained to extract a fixed dimensional speaker embedding from a spoken utterance of variable length. Contrastive learning is a typical self-supervised learning technique. However, the contrastive learning of the speaker encoder depends very much on the sampling strategy of positive and negative pairs. It is common that we sample a positive pair of segments from the same utterance. Unfortunately, such a strategy, denoted as poor-man's positive pairs (PPP), lacks the necessary diversity. In this work, we propose a multi-modal contrastive learning technique with novel sampling strategies. By cross-referencing between speech and face data, we find diverse positive pairs (DPP) for contrastive learning, thus improving the robustness of speaker encoder. We train the speaker encoder on the VoxCeleb2 dataset without any speaker labels, and achieve an equal error rate (EER) of 2.89%, 3.17% and 6.27% under the proposed progressive clustering strategy, and an EER of 1.44%, 1.77% and 3.27% under the two-stage learning strategy with pseudo labels, on the three test sets of VoxCeleb1. This novel solution outperforms the state-of-the-art self-supervised learning methods by a large margin, at the same time, achieves comparable results with the supervised learning counterpart. We also evaluate our self-supervised learning technique on the LRS2 and LRW datasets, where speaker information is unavailable. All experiments suggest that the proposed neural architecture and sampling strategies are robust across datasets. Ruijie Tao, Kong-Aik Lee, Rohan Kumar Das, Ville Hautamäki, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | MFA: TDNN with Multi-Scale Frequency-Channel Attention for Text-Independent Speaker Verification with Short UtterancesabstractThe time delay neural network (TDNN) represents one of the state-of-the-art of neural solutions to text-independent speaker verification. However, they require a large number of filters to capture the speaker characteristics at any local frequency region. In addition, the performance of such systems may degrade under short utterance scenarios. To address these issues, we propose a multi-scale frequency-channel attention (MFA), where we characterize speakers at different scales through a novel dual-path design which consists of a convolutional neural network and TDNN. We evaluate the proposed MFA on the VoxCeleb database and observe that the proposed framework with MFA can achieve state-of-the-art performance while reducing parameters and computation complexity. Further, the MFA mechanism is found to be effective for speaker verification with short test utterances. Tianchi Liu 0004, Rohan Kumar Das, Kong-Aik Lee, Haizhou Li 0001 |
ICASSP | 2 |
| 2022 | Self-Supervised Speaker Recognition with Loss-Gated LearningabstractIn self-supervised learning for speaker recognition, pseudo labels are useful as the supervision signals. It is a known fact that a speaker recognition model doesn’t always benefit from pseudo labels due to their unreliability. In this work, we observe that a speaker recognition network tends to model the data with reliable labels faster than those with unreliable labels. This motivates us to study a loss-gated learning (LGL) strategy, which extracts the reliable labels through the fitting ability of the neural network during training. With the proposed LGL, our speaker recognition model obtains a 46.3% performance gain over the system without it. Further, the proposed self-supervised speaker recognition with LGL trained on the VoxCeleb2 dataset without any labels achieves an equal error rate of 1.66% on the VoxCeleb1 original test set. Ruijie Tao, Kong-Aik Lee, Rohan Kumar Das, Ville Hautamäki, Haizhou Li 0001 |
ICASSP | 3 |
| 2022 | Neural Acoustic-Phonetic Approach for Speaker Verification With Phonetic Attention MaskabstractTraditional acoustic-phonetic approach makes use of both spectral and phonetic information when comparing the voice of speakers. While phonetic units are not equally informative, the phonetic context of speech plays an important role in speaker verification (SV). In this paper, we propose a neural acoustic-phonetic approach that learns to dynamically assign differentiated weights to spectral features for SV. Such differentiated weights form a phonetic attention mask (PAM). The neural acoustic-phonetic framework consists of two training pipelines, one for SV and another for speech recognition. Through the PAM, we leverage the phonetic information for SV. We evaluate the proposed neural acoustic-phonetic framework on the RSR2015 database Part III corpus, that consists of random digit strings. We show that the proposed framework with PAM consistently outperforms baseline with an equal error rate reduction of 13.45% and 10.20% for female and male data, respectively. Tianchi Liu 0004, Rohan Kumar Das, Kong-Aik Lee, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 2 |
| 2021 | Data Augmentation with Signal Companding for Detection of Logical Access AttacksabstractThe recent advances in voice conversion (VC) and text-to-speech (TTS) make it possible to produce natural sounding speech that poses threat to automatic speaker verification (ASV) systems. To this end, research on spoofing countermeasures has gained attention to protect ASV systems from such attacks. While the advanced spoofing countermeasures are able to detect known nature of spoofing attacks, they are not that effective under unknown attacks. In this work, we propose a novel data augmentation technique using a-law and mu-law based signal companding. We believe that the proposed method has an edge over traditional data augmentation by adding small perturbation or quantization noise. The studies are conducted on ASVspoof 2019 logical access corpus using light convolutional neural network based system. We find that the proposed data augmentation technique based on signal companding outperforms the state-of-the-art spoofing countermeasures showing ability to handle unknown nature of attacks. Rohan Kumar Das, Haizhou Li 0001 |
ICASSP | 1 |
| 2021 | Diagnosis of COVID-19 Using Auditory Acoustic CuesabstractCOVID-19 can be pre-screened based on symptoms and confirmed using other laboratory tests.The cough or speech from patients are also studied in the recent time for detection of COVID-19 as they are indicators of change in anatomy and physiology of the respiratory system.Along this direction, the diagnosis of COVID-19 using acoustics (DiCOVA) challenge aims to promote such research by releasing publicly available cough/speech corpus.We participated in the Track-1 of the challenge, which deals with COVID-19 detection using cough sounds from individuals.In this challenge, we use a few novel auditory acoustic cues based on long-term transform, equivalent rectangular bandwidth spectrum and gammatone filterbank.We evaluate these representations using logistic regression, random forest and multilayer perceptron classifiers for detection of COVID-19.On the blind test set, we obtain an area under the ROC curve (AUC) of 83.49% for the best system submitted to the challenge.It is worth noting that the submitted system ranked among the top few systems on the leaderboard and outperformed the challenge baseline by a large margin. Rohan Kumar Das, Maulik C. Madhavi, Haizhou Li 0001 |
Interspeech | 1 |
| 2021 | Is Someone Speaking?: Exploring Long-term Temporal Features for Audio-visual Active Speaker DetectionabstractActive speaker detection (ASD) seeks to detect who is speaking in a visual scene of one or more speakers. The successful ASD depends on accurate interpretation of short-term and long-term audio and visual information, as well as audio-visual interaction. Unlike the prior work where systems make decision instantaneously using short-term features, we propose a novel framework, named TalkNet, that makes decision by taking both short-term and long-term features into consideration. TalkNet consists of audio and visual temporal encoders for feature representation, audio-visual cross-attention mechanism for inter-modality interaction, and a self-attention mechanism to capture long-term speaking evidence. The experiments demonstrate that TalkNet achieves 3.5% and 2.2% improvement over the state-of-the-art systems on the AVA-ActiveSpeaker dataset and Columbia ASD dataset, respectively. Code has been made available at: https://github.com/TaoRuijie/TalkNet_ASD. Ruijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian 0001, Zheng Shou 0001, Haizhou Li 0001 |
ACM Multimedia | 3 |
| 2021 | Enhancing the Intelligibility of Cleft Lip and Palate Speech Using Cycle-Consistent Adversarial NetworksabstractCleft lip and palate (CLP) refer to a congenital craniofacial condition that causes various speech-related disorders. As a result of structural and functional deformities, the affected subjects' speech intelligibility is significantly degraded, limiting the accessibility and usability of speech-controlled devices. Towards addressing this problem, it is desirable to improve the CLP speech intelligibility. Moreover, it would be useful during speech therapy. In this study, the cycle-consistent adversarial network (CycleGAN) method is exploited for improving CLP speech intelligibility. The model is trained on native Kannada-speaking childrens' speech data. The effectiveness of the proposed approach is also measured using automatic speech recognition performance. Further, subjective evaluation is performed, and those results also confirm the intelligibility improvement in the enhanced speech over the original. Protima Nomo Sudro, Rohan Kumar Das, Rohit Sinha 0003, S. R. Mahadeva Prasanna |
SLT | 2 |
| 2021 | Graph Fourier Transform Based Audio Zero-WatermarkingabstractThe frequent exchange of multimedia information in the present era projects an increasing demand for copyright protection. In this work, we propose a novel audio zerowatermarking technology based on graph Fourier transform for enhancing the robustness with respect to copyright protection. In this approach, the combined shift operator is used to construct the graph signal, upon which the graph Fourier analysis is performed. The selected maximum absolute graph Fourier coefficients representing the characteristics of the audio segment are then encoded into a feature binary sequence using K-means algorithm. Finally, the resultant feature binary sequence is XORed with the watermark binary sequence to realize the embedding of the zero-watermark. The experimental studies show that the proposed approach performs more effectively in resisting common or synchronization attacks than the existing state ofthe-art methods. Longting Xu, Daiyu Huang, Syed Faham Ali Zaidi, Abdul Rauf 0004, Rohan Kumar Das |
IEEE Signal Process. Lett. | 5 |
| 2021 | Modified Magnitude-Phase Spectrum Information for Spoofing DetectionabstractMost of the existing feature representations for spoofing countermeasures consider information either from the magnitude or phase spectrum. We hypothesize that both magnitude and phase spectra can be beneficial for spoofing detection (SD) when collectively used to capture the signal artifacts. In this work, we propose a novel feature referred to as modified magnitude-phase spectrum (MMPS) to capture both magnitude and phase information from the speech signal. The constant-Q transform is used to obtain the magnitude and phase information in terms of MMPS, which can be denoted as CQT-MMPS. We then use this information for the proposal of a handcrafted feature, namely, constant-Q modified octave coefficients (CQMOC). To evaluate the proposed CQT-MMPS and CQMOC features, three classic anti-spoofing models are adopted, including the Gaussian mixture model (GMM), the light CNN (LCNN) and the ResNet. Additionally, since there is usually no prior knowledge about the spoofing kind in real-world applications, two novel methods referred to as three-class classifiers with maximum spoofing-score (TCMS) and multi-task learning (MTL) are designed for unknown-kind SD (UKSD). The experimental results on ASVspoof 2019 corpus show that CQMOC outperforms most of the commonly-used handcrafted features, and the CQT-based MMPS performs better than the magnitude-phase spectrum and the commonly-used log power spectrum. Further, the MMPS-based systems can achieve comparable or even better performance when compared with the state-of-the-art systems. We find that the newly-designed TCMS and MTL methods outperform the combination-based method for UKSD and meanwhile, generalize much better than the respective-kind-based methods in cross-spoofing-kind evaluation scenarios. Rohan Kumar Das, Yanmin Qian |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | On the Importance of Vocal Tract Constriction for Speaker Characterization: The Whispered Speech StudyabstractCharacterizing speakers under stressed condition is a challenge because speakers deviate from the normal speech production process. Whispered speech is one among them that is produced by abducting the vocal folds to pass the air out of mouth. During this process, the airflow thus passed is influenced by the physiological structure of vocal tract, that can be described by the level of constriction in vocal tract during speech production. We believe such information may be helpful for capturing speaker traits in such scenarios. The vocal tract constriction evidence in combination with mel frequency cepstral coefficients is investigated for speaker recognition experiments. The studies are conducted on CHAINS corpus that includes both neutral and whispered speech. We find that considering the constriction evidence from vocal tract improves the speaker characterization for whispered speech and under different mismatched scenarios. Rohan Kumar Das, Haizhou Li 0001 |
ICASSP | 1 |
| 2020 | Assessing the Scope of Generalized Countermeasures for Anti-SpoofingabstractMost of the research on anti-spoofing countermeasures are specific to a type of spoofing attacks, where models are trained on data of a particular nature, either synthetic or replay. However, one does not have such leverage as there is no prior knowledge about the kind of spoofing attack in practice. Therefore, there is a requirement to assess the scope of generalized countermeasures for anti-spoofing. The ASVspoof 2019 challenge covers both synthetic as well as replay attacks, which makes the database suitable for such study. In this work, we consider widely popular constant-Q cepstral coefficient features along with two other promising front-ends that capture long-term signal characteristics to assess their scope as generalized countermeasures. Additionally, a comprehensive study is made across different editions of ASVspoof corpora to highlight the need of robust generalized countermeasures in unseen conditions. Rohan Kumar Das, Haizhou Li 0001 |
ICASSP | 1 |
| 2020 | End-to-End Code-Switching TTS with Cross-Lingual Language ModelabstractCode-switching text-to-speech (TTS) aims to enable a system to speak two languages with a single voice and in the same utterance. In this paper, we propose to incorporate cross-lingual word embedding into an end-to-end TTS system, to improve the voice rendering. The cross-lingual word embedding, generated from a pre-trained cross-lingual language model, is able to encode words of two languages in the same embedding space, therefore, allows words across languages to share each other's contextual information, which is useful for the voice rendering of code-switching content. To investigate the effectiveness of this idea, we conduct studies on two multi-speaker monolingual corpora, namely, THCHS30 Mandarin and LibriTTS English database. The evaluation results show that our proposed framework outperforms the baseline systems when presented with code-switching text input, and achieves state-of-the-art performance. Xuehao Zhou, Xiaohai Tian, Grandee Lee, Rohan Kumar Das, Haizhou Li 0001 |
ICASSP | 4 |
| 2020 | Speaker-Utterance Dual Attention for Speaker and Utterance VerificationabstractIn this paper, we study a novel technique that exploits the interaction between speaker traits and linguistic content to improve both speaker verification and utterance verification performance. We implement an idea of speaker-utterance dual attention (SUDA) in a unified neural network. The dual attention refers to an attention mechanism for the two tasks of speaker and utterance verification. The proposed SUDA features an attention mask mechanism to learn the interaction between the speaker and utterance information streams. This helps to focus only on the required information for respective task by masking the irrelevant counterparts. The studies conducted on RSR2015 corpus confirm that the proposed SUDA outperforms the framework without attention mask as well as several competitive systems for both speaker and utterance verification. Tianchi Liu 0004, Rohan Kumar Das, Maulik C. Madhavi, Shengmei Shen, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2020 | The Attacker's Perspective on Automatic Speaker Verification: An OverviewabstractSecurity of automatic speaker verification (ASV) systems is compromised by various spoofing attacks.While many types of non-proactive attacks (and their defenses) have been studied in the past, attacker's perspective on ASV, represents a far less explored direction.It can potentially help to identify the weakest parts of ASV systems and be used to develop attackeraware systems.We present an overview on this emerging research area by focusing on potential threats of adversarial attacks on ASV, spoofing countermeasures, or both.We conclude the study with discussion on selected attacks and leveraging from such knowledge to improve defense mechanisms against adversarial attacks. Rohan Kumar Das, Xiaohai Tian, Tomi Kinnunen, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2020 | The INTERSPEECH 2020 Far-Field Speaker Verification ChallengeabstractThe INTERSPEECH 2020 Far-Field Speaker Verification Challenge (FFSVC 2020) addresses three different research problems under well-defined conditions: far-field text-dependent speaker verification from single microphone array, far-field textindependent speaker verification from single microphone array, and far-field text-dependent speaker verification from distributed microphone arrays.All three tasks pose a cross-channel challenge to the participants.To simulate the real-life scenario, the enrollment utterances are recorded from close-talk cellphone, while the test utterances are recorded from the far-field microphone arrays.In this paper, we describe the database, the challenge, and the baseline system, which is based on a ResNetbased deep speaker network with cosine similarity scoring.For a given utterance, the speaker embeddings of different channels are equally averaged as the final embedding.The baseline system achieves minDCFs of 0.62, 0.66, and 0.64 and EERs of 6.27%, 6.55%, and 7.18% for task 1, task 2, and task 3, respectively. Xiaoyi Qin, Ming Li 0026, Hui Bu, Wei Rao 0002, Rohan Kumar Das, Shri Narayanan, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2020 | Audio-Visual Speaker Recognition with a Cross-Modal Discriminative NetworkabstractAudio-visual speaker recognition is one of the tasks in the recent 2019 NIST speaker recognition evaluation (SRE).Studies in neuroscience and computer science all point to the fact that vision and auditory neural signals interact in the cognitive process.This motivated us to study a cross-modal network, namely voice-face discriminative network (VFNet) that establishes the general relation between human voice and face.Experiments show that VFNet provides additional speaker discriminative information.With VFNet, we achieve 16.54% equal error rate relative reduction over the score level fusion audio-visual baseline on evaluation set of 2019 NIST SRE. Ruijie Tao, Rohan Kumar Das, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2020 | Light Convolutional Neural Network with Feature Genuinization for Detection of Synthetic Speech AttacksabstractModern text-to-speech (TTS) and voice conversion (VC) systems produce natural sounding speech that questions the security of automatic speaker verification (ASV).This makes detection of such synthetic speech very important to safeguard ASV systems from unauthorized access.Most of the existing spoofing countermeasures perform well when the nature of the attacks is made known to the system during training.However, their performance degrades in face of unseen nature of attacks.In comparison to the synthetic speech created by a wide range of TTS and VC methods, genuine speech has a more consistent distribution.We believe that the difference between the distribution of synthetic and genuine speech is an important discriminative feature between the two classes.In this regard, we propose a novel method referred to as feature genuinization that learns a transformer with convolutional neural network (CNN) using the characteristics of only genuine speech.We then use this genuinization transformer with a light CNN classifier.The ASVspoof 2019 logical access corpus is used to evaluate the proposed method.The studies show that the proposed feature genuinization based LCNN system outperforms other state-ofthe-art spoofing countermeasures, depicting its effectiveness for detection of synthetic speech attacks. Zhenzong Wu, Rohan Kumar Das, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2020 | Significance of Subband Features for Synthetic Speech DetectionabstractIn text-to-speech or voice conversion based synthetic speech detection, it is a common practice that spectral information over the entire frequency band is used for feature representation. We propose a new method, referred to as subband transform, that characterizes the signals by subband. It is found that subband transform captures the artifacts in synthetic speech more effectively than full band transform. We propose equal subband transform, octave subband transform, and mel subband transform for three novel features, namely, constant-Q equal subband transform (CQ-EST), constant-Q octave subband transform (CQ-OST) and discrete Fourier mel subband transform (DF-MST). We evaluate the three features on the ASVspoof 2015, noisy ASVspoof 2015 and ASVspoof 2019 logical access corpora. The experiments show that the proposed CQ-EST feature achieves an average equal error rate of 0.056% on ASVspoof 2015 evaluation set. The study observes that the features based on subband transform outperform those based on full band transform under both clean and noisy conditions. In addition, the tandem detection cost function of CQ-OST can reach 0.188 on ASVspoof 2019 logical access evaluation set. Rohan Kumar Das, Haizhou Li 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2019 | Long Range Acoustic and Deep Features Perspective on ASVspoof 2019abstractTo secure automatic speaker verification (ASV) systems from intruders, robust countermeasures for spoofing attack detection are required. The ASVspoof series of challenge provides a shared anti-spoofing task. The recent edition, ASVspoof 2019, focuses on attacks by both synthetic and replay speech that are referred to as logical and physical access attacks, respectively. In the ASVspoof 2019 submission, we considered novel countermeasures based on long range acoustic features, that are unique in many ways as they are derived using octave power spectrum and subbands, as opposed to the commonly used linear power spectrum. During the post-challenge study, we further investigate the use of deep features that enhances the discriminative ability between genuine and spoofed speech. In this paper, we summarize the findings from the perspective of long range acoustic and deep features for spoof detection. We make a comprehensive analysis on the nature of different kinds of spoofing attacks and system development. Rohan Kumar Das, Haizhou Li 0001 |
ASRU | 1 |
| 2019 | A Modularized Neural Network with Language-Specific Output Layers for Cross-Lingual Voice ConversionabstractThis paper presents a cross-lingual voice conversion framework that adopts a modularized neural network. The modularized neural network has a common input structure that is shared for both languages, and two separate output modules, one for each language. The idea is motivated by the fact that phonetic systems of languages are similar because humans share a common vocal production system, but acoustic renderings, such as prosody and phonotactic, vary a lot from language to language. The modularized neural network is trained to map Phonetic PosteriorGram (PPG) to acoustic features for multiple speakers. It is conditioned on a speaker i-vector to generate the desired target voice. We validated the idea between English and Mandarin languages in objective and subjective tests. In addition, mixed-lingual PPG derived from a unified English-Mandarin acoustic model is proposed to capture the linguistic information from both languages. It is found that our proposed modularized neural network significantly outperforms the baseline approaches in terms of speech quality and speaker individuality, and mixed-lingual PPG representation further improves the conversion performance. Yi Zhou 0020, Xiaohai Tian, Emre Yilmaz 0001, Rohan Kumar Das, Haizhou Li 0001 |
ASRU | 4 |
| 2019 | Cross-lingual Voice Conversion with Bilingual Phonetic Posteriorgram and Average ModelingabstractThis paper presents a cross-lingual voice conversion approach using bilingual Phonetic PosteriorGram (PPG) and average modeling. The proposed approach makes use of bilingual PPGs to represent speaker-independent features of speech signals from different languages in the same feature space. In particular, a bilingual PPG is formed by stacking two monolingual PPG vectors, which are extracted from two monolingual speech recognition systems. The conversion model is trained to learn the relationship between bilingual PPGs and the corresponding acoustic features. To leverage the linguistic and acoustic information from other speakers in different languages, an average model is trained with multiple speakers in both source and target languages. I-vector is utilized as an additional input feature of the average model for network adaptation. Experiments are performed for intralingual and cross-lingual voice conversion between English and Mandarin speakers. Both objective and subjective evaluations demonstrate the effectiveness of our proposed approach. Yi Zhou 0020, Xiaohai Tian, Haihua Xu 0001, Rohan Kumar Das, Haizhou Li 0001 |
ICASSP | 4 |
| 2019 | Instantaneous Phase and Long-Term Acoustic Cues for Orca Activity Detection
Rohan Kumar Das, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2019 | Long Range Acoustic Features for Spoofed Speech Detection
Rohan Kumar Das, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2019 | SpeechMarker: A Voice Based Multi-Level Attendance Application
Sarfaraz Jelil, Rohan Kumar Das, S. R. Mahadeva Prasanna, Rohit Sinha 0003 |
INTERSPEECH | 3 |
| 2019 | I4U Submission to NIST SRE 2018: Leveraging from a Decade of Shared ExperiencesabstractThe I4U consortium was established to facilitate a joint entry to NIST speaker recognition evaluations (SRE). The latest edition of such joint submission was in SRE 2018, in which the I4U submission was among the best-performing systems. SRE'18 also marks the 10-year anniversary of I4U consortium into NIST SRE series of evaluation. The primary objective of the current paper is to summarize the results and lessons learned based on the twelve sub-systems and their fusion submitted to SRE'18. It is also our intention to present a shared view on the advancements, progresses, and major paradigm shifts that we have witnessed as an SRE participant in the past decade from SRE'08 to SRE'18. In this regard, we have seen, among others, a paradigm shift from supervector representation to deep speaker embedding, and a switch of research challenge from channel compensation to domain adaptation. Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Hitoshi Yamamoto, Koji Okabe, Ville Vestman, Jing Huang 0019, Guo-Hong Ding, Hanwu Sun, Anthony Larcher, Rohan Kumar Das, Haizhou Li 0001, Mickael Rouvier, Pierre-Michel Bousquet, Wei Rao 0002, Qing Wang 0039, Fahimeh Bahmaninezhad, Héctor Delgado, Massimiliano Todisco |
INTERSPEECH | 11 |
| 2019 | A Unified Framework for Speaker and Utterance Verification
Tianchi Liu 0004, Maulik C. Madhavi, Rohan Kumar Das, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2019 | Multi-Level Adaptive Speech Activity Detector for Speech in Naturalistic Environments
Bidisha Sharma, Rohan Kumar Das, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2019 | On the Importance of Audio-Source Separation for Singer Identification in Polyphonic Music
Bidisha Sharma, Rohan Kumar Das, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2019 | Robust Sound Recognition: A Neuromorphic Approach
Jibin Wu, Zihan Pan, Malu Zhang, Rohan Kumar Das, Yansong Chua, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2019 | Extraction of Octave Spectra Information for Spoofing Attack DetectionabstractThis article focuses on extracting information from the octave power spectra of long-term constant-Q transform (CQT) for spoofing attack detection. A novel framework based on multi-level transform (MLT) is proposed that can capture the relevant information from octave power spectra using level by level in a multi-level manner. We then derive a novel feature referred to as constant-Q multi-level coefficient (CMC) based on proposed MLT. The proposed feature is evaluated on synthetic as well as replay speech detection studies on ASVspoof 2015 and ASVspoof 2017 version 2.0 database, respectively. We find the proposed CMC feature outperforms the conventional constant-Q cepstral coefficient based long-term feature obtained from linear power spectrum after uniform resampling. This depicts the usefulness of MLT to extract salient artifacts from octave power spectrum. Further, the proposed CMC feature performs better than the existing the well known other state-of-the-art systems for spoofing attack detection that showcases its importance. Rohan Kumar Das, Nina Zhou |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Generative X-Vectors for Text-Independent Speaker VerificationabstractSpeaker verification (SV) systems using deep neural network embeddings, so-called the x-vector systems, are becoming popular due to its good performance superior to the i-vector systems. The fusion of these systems provides improved performance benefiting both from the discriminatively trained x-vectors and generative i-vectors capturing distinct speaker characteristics. In this paper, we propose a novel method to include the complementary information of i-vector and x-vector, that is called generative x-vector. The generative x-vector utilizes a transformation model learned from the i-vector and x-vector representations of the background data. Canonical correlation analysis is applied to derive this transformation model, which is later used to transform the standard x-vectors of the enrollment and test segments to the corresponding generative x-vectors. The SV experiments performed on the NIST SRE 2010 dataset demonstrate that the system using generative x-vectors provides considerably better performance than the baseline i-vector and x-vector systems. Furthermore, the generative x-vectors outperform the fusion of i-vector and x-vector systems for long-duration utterances, while yielding comparable results for short-duration utterances. Longting Xu, Rohan Kumar Das, Emre Yilmaz 0001, Haizhou Li 0001 |
SLT | 2 |
| 2017 | Spoof Detection Using Source, Instantaneous Frequency and Cepstral Features
Sarfaraz Jelil, Rohan Kumar Das, S. R. Mahadeva Prasanna, Rohit Sinha 0003 |
INTERSPEECH | 2 |
| 2017 | IITG-Indigo System for NIST 2016 SRE ChallengeabstractOrientador : Juarez Brandão Lopes Nagendra Kumar 0004, Rohan Kumar Das, Sarfaraz Jelil, Dhanush B. K, H. Kashyap, K. Sri Rama Murty, Sriram Ganapathy, Rohit Sinha 0003, S. R. Mahadeva Prasanna |
INTERSPEECH | 2 |
| 2017 | Exploring kernel discriminant analysis for speaker verification with limited test data
Rohan Kumar Das, Akhil Babu Manam, S. R. Mahadeva Prasanna |
Pattern Recognit. Lett. | 1 |
| 2017 | Analysis of the Intrinsic Mode Functions for Speaker Information
Rajib Sharma, S. R. Mahadeva Prasanna, Ramesh K. Bhukya, Rohan Kumar Das |
Speech Commun. | 4 |
| 2016 | Exploring Session Variability and Template Aging in Speaker Verification for Fixed Phrase Short Utterances
Rohan Kumar Das, Sarfaraz Jelil, S. R. Mahadeva Prasanna |
INTERSPEECH | 1 |
| 2015 | Speaker verification using Gaussian posteriorgrams on fixed phrase short utterances
Sarfaraz Jelil, Rohan Kumar Das, Rohit Sinha 0003, S. R. Mahadeva Prasanna |
INTERSPEECH | 2 |
| 2014 | Combining source and system information for limited data speaker verificationabstractSpeaker verification using limited data is always a challenge for practical implementation as an application. An analysis on speaker verification studies for an i-vector based method using Mel-Frequency Cepstral Coefficient (MFCC) feature shows that the performance drops drastically as the duration of test data is reduced. This decrease in performance is due to insufficient phonetic coverage when we capture only the vocal tract feature. However the same can be improved if some source characteristics are taken into consideration. This paper attempts to improve the speaker verification performance using source characteristics. A recently proposed characterization of the voice source signal called the discrete cosine transform of the integrated linear prediction residual (DCTILPR) has been found to be useful as a speaker-specific feature. Speaker verification is performed over short test utterances in the NIST 2003 database using both the DCTILPR and MFCC features, and their score-level combination is found to give a significant performance improvement over the system using only the MFCC features. Rohan Kumar Das, S. Abhiram, S. R. Mahadeva Prasanna, A. G. Ramakrishnan |
INTERSPEECH | 1 |