Raghuveer Peri

dblp:251/8508 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0002-1010-065XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 8 since 2021
YearPublicationVenuePosition
2025 Knowledge Distillation From Ensemble for Spoken Language Identification
abstract
Spoken language identification (LID) has seen substantial performance gains with the rise of large-scale models. However, these models are often computationally expensive and impractical for many real-world applications. In this work, we propose a novel knowledge distillation from ensemble framework to address this challenge. By distilling an ensemble of large LID models into a single, more efficient student, we achieve comparable or even superior performance while reducing computational cost by 67%. Our approach yields a student model with less than 10% the size of a 200M+ parameter teacher ensemble, yet outperforming a 140M parameter teacher by 13% relative. Additionally, combining our distillation technique with decoupled knowledge distillation leads to substantial gains (50% relative), especially for confusable and low-resource languages in the FLEURS dataset.
Raghuveer Peri, Seyed Omid Sadjadi, Daniel Garcia-Romero, Srikanth Vishnubhotla, Kyu J. Han
ICASSP1
2025 Defending Speech-enabled LLMs Against Adversarial Jailbreak Threats
Antonios Alexos, Raghuveer Peri, Sai Muralidhar Jayanthi, Metehan Cekic, Srikanth Vishnubhotla, Kyu J. Han, Srikanth Ronanki
INTERSPEECH2
2024 SWAN: SubWord Alignment Network for HMM-free word timing estimation in end-to-end automatic speech recognition
Woo Hyun Kang, Srikanth Vishnubhotla, Rudolf Braun, Yogesh Virkar, Raghuveer Peri, Kyu J. Han
INTERSPEECH5
2023 A study of bias mitigation strategies for speaker recognition
Raghuveer Peri, Krishna Somandepalli, Shri Narayanan
Comput. Speech Lang.1
2022 Scene Representation Learning from Videos Using Self-Supervised and Weakly-Supervised Techniques
abstract
Holistic understanding of videos requires the recognition of the overall scene beyond detecting foreground activity and objects. It provides valuable information for various video understanding tasks such as video summarization, scene change detection and content filtering. While significant effort has been put into developing models for scene classification in images (e.g. Places365), video-level scene recognition is relatively nascent. The scope of this paper is to address this problem of going from image representations to video for scene classification. In particular, we compare self-supervised deep learning methods on video scene recognition task using the HVU dataset. Starting from strong image level scene representations, with triplets based contrastive loss, we train a video-level scene classifier. We propose triplet sampling strategies that aid the self-supervision. We compare the self-supervised techniques against the image level scene representations, as well as a weakly supervised classifier trained on image labels. We observe that the models learned using self-supervised method outperform both baselines (with statistical significance), showing that we are able to retain the representative power of the video-level scene representations compared to a competitive image-level scene recognition model trained on Places365, while showing benefits over weakly supervised techniques.
Raghuveer Peri, Srinivas Parthasarathy, Shiva Sundaram
ICIP1
2022 User-Level Differential Privacy against Attribute Inference Attack of Speech Emotion Recognition on Federated Learning
abstract
Many existing privacy-enhanced speech emotion recognition (SER) frameworks focus on perturbing the original speech data through adversarial training within a centralized machine learning setup. However, this privacy protection scheme can fail since the adversary can still access the perturbed data. In recent years, distributed learning algorithms, especially federated learning (FL), have gained popularity to protect privacy in machine learning applications. While FL provides good intuition to safeguard privacy by keeping the data on local devices, prior work has shown that privacy attacks, such as attribute inference attacks, are achievable for SER systems trained using FL. In this work, we propose to evaluate the user-level differential privacy (UDP) in mitigating the privacy leaks of the SER system in FL. UDP provides theoretical privacy guarantees with privacy parameters $\epsilon$ and $\delta$. Our results show that the UDP can effectively decrease attribute information leakage while keeping the utility of the SER system with the adversary accessing one model update. However, the efficacy of the UDP suffers when the FL system leaks more model updates to the adversary. We make the code publicly available to reproduce the results in https://github.com/usc-sail/fed-ser-leakage.
Tiantian Feng, Raghuveer Peri, Shri Narayanan
INTERSPEECH2
2021 A Computational Tool to Study Vocal Participation of Women in UN-ITU Meetings
abstract
International organizations such as the United Nations drive policies that impact our everyday lives. Diverse representation of people and ideas in the decision making process of such bodies is critical to ensure that the policies work for everyone. One aspect of the representation is the partipants' expressed gender. In this work, we focus on analyzing meetings at the International Telecommunication Union (ITU). These meetings include a moderator who mediates the proceedings between delegates from across the world speaking in different languages. For the purpose of quantifying the participation of delegates, we propose a scalable, human-in-the-loop system to first identify the moderator's speech and estimate the speaking time with respect to gender for all the speakers. Our proposed system includes three main audio modules: speech activity detection, gender identification and moderator verification using a human-labelled speech probe. We then estimate percentage of speaking time controlled for the moderator's speech. We present detailed and multilingual performance evaluation of the component systems using state-of-the-art technologies for these tasks. Finally, we examine the vocal participation of female delegates in the 2018 ITU Plenipotentiary Conference spanning for 18 days and about 108 hours of audio recordings.
Rajat Hebbar, Krishna Somandepalli, Raghuveer Peri, Ruchir Travadi, Tracy Tuplin, Fernando Rivera, Shri Narayanan
CBMI3
2021 Adversarial Defense for Deep Speaker Recognition Using Hybrid Adversarial Training
abstract
Deep neural network based speaker recognition systems can easily be deceived by an adversary using minuscule imperceptible perturbations to the input speech samples. These adversarial attacks pose serious security threats to the speaker recognition systems that use speech biometric. To address this concern, in this work, we propose a new defense mechanism based on a hybrid adversarial training (HAT) setup. In contrast to existing works on countermeasures against adversarial attacks in deep speaker recognition that only use class-boundary information by supervised cross-entropy (CE) loss, we propose to exploit additional information from supervised and unsupervised cues to craft diverse and stronger perturbations for adversarial training. Specifically, we employ multi-task objectives using CE, feature-scattering (FS), and margin losses to create adversarial perturbations and include them for adversarial training to enhance the robustness of the model. We conduct speaker recognition experiments on the Librispeech dataset, and compare the performance with state-of-the-art projected gradient descent (PGD)-based adversarial training which employs only CE objective. The proposed HAT improves adversarial accuracy by absolute 3.29% and 3.18% for PGD and Carlini-Wagner (CW) attacks respectively, while retaining high accuracy on benign examples.
Monisankha Pal, Arindam Jati, Raghuveer Peri, Chin-Cheng Hsu, Wael Abd-Almageed, Shri Narayanan
ICASSP3
2021 Disentanglement for Audio-Visual Emotion Recognition Using Multitask Setup
abstract
Deep learning models trained on audio-visual data have been successfully used to achieve state-of-the-art performance for emotion recognition. In particular, models trained with multitask learning have shown additional performance improvements. However, such multitask models entangle information between the tasks, encoding the mutual dependencies present in label distributions in the real world data used for training. This work explores the disentanglement of multimodal signal representations for the primary task of emotion recognition and a secondary person identification task. In particular, we developed a multitask framework to extract low-dimensional embeddings that aim to capture emotion specific information, while containing minimal information related to person identity. We evaluate three different techniques for disentanglement and report results of up to 13% disentanglement while maintaining emotion recognition performance.
Raghuveer Peri, Srinivas Parthasarathy, Charles Bradshaw, Shiva Sundaram
ICASSP1
2021 Adversarial attack and defense strategies for deep speaker recognition systems
Arindam Jati, Chin-Cheng Hsu, Monisankha Pal, Raghuveer Peri, Wael Abd-Almageed, Shri Narayanan
Comput. Speech Lang.4
2021 Temporal Dynamics of Workplace Acoustic Scenes: Egocentric Analysis and Prediction
abstract
Identification of the acoustic environment from an audio recording, also known as acoustic scene classification, is an active area of research. In this paper, we study dynamically-changing background acoustic scenes from the egocentric perspective of an individual in a workplace. In a novel data collection setup, wearable sensors were deployed on individuals to collect audio signals within a built environment, while Bluetooth-based hubs continuously tracked the individual's location which represents the acoustic scene at a certain time. The data of this paper come from 170 hospital workers gathered continuously during work shifts for a 10 week period. In the first part of our study, we investigate temporal patterns in the egocentric sequence of acoustic scenes encountered by an employee, and the association of those patterns with factors such as job-role and daily routine of the individual. Motivated by evidence of multifaceted effects of ambient sounds on human psychology, we also analyze the association of the temporal dynamics of the perceived acoustic scenes with particular behavioral traits of the individual. Experiments reveal rich temporal patterns in the acoustic scenes experienced by the individuals during their work shifts, and a strong association of those patterns with various constructs related to job-roles and behavior of the employees. In the second part of our study, we employ deep learning models to predict the temporal sequence of acoustic scenes from the egocentric audio signal. We propose a two-stage framework where a recurrent neural network is trained on top of the latent acoustic representations learned by a segment-level neural network. The experimental results show the efficacy of the proposed system in predicting sequence of acoustic scenes, highlighting the existence of underlying temporal patterns in the acoustic scenes experienced in workplace.
Arindam Jati, Amrutha Nadarajan, Raghuveer Peri, Karel Mundnich, Tiantian Feng, Benjamin Girault, Shri Narayanan
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Meta-Learning With Latent Space Clustering in Generative Adversarial Network for Speaker Diarization
abstract
The performance of most speaker diarization systems with x-vector embeddings is both vulnerable to noisy environments and lacks domain robustness. Earlier work on speaker diarization using generative adversarial network (GAN) with an encoder network (ClusterGAN) to project input x-vectors into a latent space has shown promising performance on meeting data. In this paper, we extend the ClusterGAN network to improve diarization robustness and enable rapid generalization across various challenging domains. To this end, we fetch the pre-trained encoder from the ClusterGAN and fine tune it by using prototypical loss (meta-ClusterGAN or MCGAN) under the meta-learning paradigm. Experiments are conducted on CALLHOME telephonic conversations, AMI meeting data, DIHARD-II (dev set) which includes challenging multi-domain corpus, and two child-clinician interaction corpora (ADOS, BOSCC) related to the autism spectrum disorder domain. Extensive analyses of the experimental data are done to investigate the effectiveness of the proposed ClusterGAN and MCGAN embeddings over x-vectors. The results show that the proposed embeddings with normalized maximum eigengap spectral clustering (NME-SC) back-end consistently outperform the Kaldi state-of-the-art x-vector diarization system. Finally, we employ embedding fusion with x-vectors to provide further improvement in diarization performance. We achieve a relative diarization error rate (DER) improvement of 6.67% to 53.93% on the aforementioned datasets using the proposed fused embeddings over x-vectors. Besides, the MCGAN embeddings provide better performance in the number of speakers estimation and short speech segment diarization compared to x-vectors and ClusterGAN on telephonic conversations.
Monisankha Pal, Manoj Kumar 0007, Raghuveer Peri, Tae Jin Park, So Hyun Kim, Catherine Lord, Somer Bishop, Shri Narayanan
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Speaker Diarization Using Latent Space Clustering in Generative Adversarial Network
abstract
In this work, we propose deep latent space clustering for speaker diarization using generative adversarial network (GAN) back-projection with the help of an encoder network. The proposed diarization system is trained jointly with GAN loss, latent variable recovery loss, and a clustering-specific loss. It uses x-vector speaker embeddings at the input, while the latent variables are sampled from a combination of continuous random variables and discrete one-hot encoded variables using the original speaker labels. We benchmark our proposed system on the AMI meeting corpus, and two child-clinician interaction corpora (ADOS and BOSCC) from the autism diagnosis domain. ADOS and BOSCC contain diagnostic and treatment outcome sessions respectively obtained in clinical settings for verbal children and adolescents with autism. Experimental results show that our proposed system significantly outperform the state-of-the-art x-vector based diarization system on these databases. Further, we perform embedding fusion with x-vectors to achieve a relative diarization error rate (DER) improvement of 31%, 36% and 49% on AMI eval, ADOS and BOSCC corpora respectively, when compared to the x-vector baseline using oracle speech segmentation.
Monisankha Pal, Manoj Kumar 0007, Raghuveer Peri, Tae Jin Park, So Hyun Kim, Catherine Lord, Somer Bishop, Shri Narayanan
ICASSP3
2020 Robust Speaker Recognition Using Unsupervised Adversarial Invariance
abstract
In this paper, we address the problem of speaker recognition in challenging acoustic conditions using a novel method to extract robust speaker-discriminative speech representations. We adopt a recently proposed unsupervised adversarial invariance architecture to train a network that maps speaker embeddings extracted using a pretrained model onto two lower dimensional embedding spaces. The embedding spaces are learnt to disentangle speaker-discriminative information from all other information present in the audio recordings, without supervision about the acoustic conditions. We analyze the robustness of the proposed embeddings to various sources of variability present in the signal for speaker verification and unsupervised clustering tasks on a large-scale speaker recognition corpus. Our analyses show that the proposed system substantially outperforms the baseline in a variety of challenging acoustic scenarios. Furthermore, for the task of speaker diarization on a real-world meeting corpus, our system shows a relative improvement of 36% in the diarization error rate compared to the state-of-the-art baseline.
Raghuveer Peri, Monisankha Pal, Arindam Jati, Krishna Somandepalli, Shri Narayanan
ICASSP1
2019 Multi-Task Discriminative Training of Hybrid DNN-TVM Model for Speaker Verification with Noisy and Far-Field Speech
Arindam Jati, Raghuveer Peri, Monisankha Pal, Tae Jin Park, Naveen Kumar 0004, Ruchir Travadi, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH2
2019 The Second DIHARD Challenge: System Description for USC-SAIL Team
Tae Jin Park, Manoj Kumar 0007, Nikolaos Flemotomos, Monisankha Pal, Raghuveer Peri, Rimita Lahiri, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH5