EDBT 2026 Demo / reviewers in the wild / expert
Bhiksha Raj
dblp:60/3996 · also Bhiksha Ramakrishnan
· DBLP profile ↗
255ranked-venue papers
16as first author
94since 2021 · last 2026
0000-0003-0038-5513ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 182 · 14 first-author · 54 since 2021Artificial intelligence and machine learning · 151 · 9 first-author · 70 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Security and privacy · 2Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring features for membership inference in ASR model auditing
Francisco Teixeira, Karla Pizzi, Raphaël Olivier, Alberto Abad, Bhiksha Raj, Isabel Trancoso |
Comput. Speech Lang. | 5 |
| 2026 | Introduction to the Special Issue on Evaluations of Large Language Models Part 2
Jindong Wang 0001, Linyi Yang, Sunayana Sitaram, Qiang Yang 0001, Bhiksha Raj |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2025 | Audio Entailment: Assessing Deductive Reasoning for Audio UnderstandingabstractRecent literature uses language to build foundation models for audio. These Audio-Language Models (ALMs) are trained on a vast number of audio-text pairs and show remarkable performance in tasks including Text-to-Audio Retrieval, Captioning, and Question Answering. However, their ability to engage in more complex open-ended tasks, like Interactive Question-Answering, requires proficiency in logical reasoning- a skill not yet benchmarked. We introduce the novel task of Audio Entailment to evaluate an ALM's deductive reasoning ability. This task assesses whether a text description (hypothesis) of audio content can be deduced from an audio recording (premise), with potential conclusions being entailment, neutral, or contradiction, depending on the sufficiency of the evidence. We create two datasets for this task with audio recordings sourced from two audio captioning datasets-AudioCaps and Clotho-and hypotheses generated using Large Language Models (LLMs). We benchmark state-of-the-art ALMs and find deficiencies in logical reasoning with both zero-shot and linear probe evaluations. Finally, we propose "caption-before-reason", an intermediate step of captioning that improves the Zero-Shot and linear-probe performance of ALMs by an absolute 6% and 3%, respectively. Soham Deshmukh, Hazim T. Bukhari, Benjamin Elizalde, Hannes Gamper, Rita Singh, Bhiksha Raj |
AAAI | 7 |
| 2025 | CoLMbo: Speaker Language Model for Descriptive ProfilingabstractSpeaker recognition systems are often limited to classification tasks and struggle to generate detailed speaker characteristics or provide context-rich descriptions. These models primarily extract embeddings for speaker identification but fail to capture demographic attributes such as dialect, gender, and age in a structured manner. This paper introduces CoLMbo, a Speaker Language Model (SLM) that addresses these limitations by integrating a speaker encoder with prompt-based conditioning. This allows for the creation of detailed captions based on speaker embeddings. CoLMbo utilizes userdefined prompts to adapt dynamically to new speaker characteristics and provides customized descriptions, including regional dialect variations and age-related traits. This innovative approach not only enhances traditional speaker profiling but also excels in zero-shot scenarios across diverse datasets, marking a significant advancement in the field of speaker recognition. The code is available at: https://github.com/massabaali7/CoLMbo Massa Baali, Syed Abdul Hannan, Purusottam Samal, Karanveer Singh, Soham Deshmukh, Rita Singh, Bhiksha Raj |
ASRU | 8 |
| 2025 | SoftVQ-VAE: Efficient 1-Dimensional Continuous TokenizerabstractEfficient image tokenization with high compression ratios remains a critical challenge for training generative models. We present SoftVQ-VAE, a continuous image tokenizer that leverages soft categorical posteriors to aggregate multiple codewords into each latent token, substantially increasing the representation capacity of the latent space. When applied to Transformer-based architectures, our approach compresses 256×256 and 512×512 images using as few as 32 or 64 1-dimensional tokens. Not only does SoftVQ-VAE show consistent and high-quality reconstruction, more importantly, it also achieves state-of-the-art and significantly faster image generation results across different denoising-based generative models. Remarkably, SoftVQ-VAE improves inference throughput by up to 18x for generating 256×256 images and 55x for 512×512 images while achieving competitive FID scores of 1.78 and 2.21 for SiT-XL. It also improves the training efficiency of the generative models by reducing the number of training iterations by 2.3x while maintaining comparable performance. With its fully-differentiable design and semantic-rich latent space, our experiment demonstrates that SoftVQ-VAE achieves efficient tokenization without compromising generation quality, paving the way for more efficient generative models. Code and model are released1. Hao Chen 0102, Ze Wang 0008, Xiang Li 0106, Ximeng Sun, Fangyi Chen, Jiang Liu 0014, Jindong Wang 0001, Bhiksha Raj, Zicheng Liu 0001, Emad Barsoum |
CVPR | 8 |
| 2025 | FALCON: Fairness Learning via Contrastive Attention Approach to Continual Semantic Scene UnderstandingabstractContinual Learning in semantic scene segmentation aims to continually learn new unseen classes in dynamic environments while maintaining previously learned knowledge. Prior studies focused on modeling the catastrophic forgetting and background shift challenges in continual learning. However, fairness, another major challenge that causes unfair predictions leading to low performance among major and minor classes, still needs to be well addressed. In addition, prior methods have yet to model the unknown classes well, thus resulting in producing non-discriminative features among unknown classes. This work presents a novel Fairness Learning via Contrastive Attention Approach to continual learning in semantic scene understanding. In particular, we first introduce a new Fairness Contrastive Clustering loss to address the problems of catastrophic forgetting and fairness. Then, we propose an attention-based visual grammar approach to effectively model the background shift problem and unknown classes, producing better feature representations for different unknown classes. Through our experiments, our proposed approach achieves State-of-the-Art (SoTA) performance on different continual learning benchmarks, i.e., ADE20K, Cityscapes, and Pascal VOC. It promotes the fairness of the continual semantic segmentation model. Thanh-Dat Truong, Utsav Prabhu, Bhiksha Raj, Jackson David Cothren, Khoa Luu |
CVPR | 3 |
| 2025 | PhoniTale: Phonologically Grounded Mnemonic Generation for Typologically Distant Language PairsabstractVocabulary acquisition poses a significant challenge for second-language (L2) learners, especially when learning typologically distant languages such as English and Korean, where phonological and structural mismatches complicate vocabulary learning.Recently, large language models (LLMs) have been used to generate keyword mnemonics by leveraging similar keywords from a learner's first language (L1) to aid in acquiring L2 vocabulary.However, most methods still rely on direct IPA-based phonetic matching or employ LLMs without phonological guidance.In this paper, we present PHONI-TALE, a novel cross-lingual mnemonic generation system that performs IPA-based phonological adaptation and syllable-aware alignment to retrieve L1 keyword sequence and uses LLMs to generate verbal cues.We evaluate PHONI-TALE through automated metrics and a shortterm recall test with human participants, comparing its output to human-written and prior automated mnemonics.Our findings show that PHONITALE consistently outperforms previous automated approaches and achieves quality comparable to human-written mnemonics. Sana Kang, Myeongseok Gwon, Su Young Kwon, Jaewook Lee 0006, Andrew S. Lan, Bhiksha Raj, Rita Singh |
EMNLP | 6 |
| 2025 | Tessellated Linear Model for Age Prediction from VoiceabstractVoice biometric tasks, such as age estimation require modelling the often complex relationship between voice features and the biometric variable. While deep learning models can handle such complexity, they typically require large amounts of accurately labeled data to perform well. Such data are often scarce for biometric tasks such as voice-based age prediction. On the other hand, simpler models like linear regression can work with smaller datasets but often fail to generalize to the underlying non-linear patterns present in the data. In this paper we propose the Tessellated Linear Model (TLM), a piecewise linear approach that combines the simplicity of linear models with the capacity of non-linear functions. TLM tessellates the feature space into convex regions and fits a linear model within each region. We optimize the tessellation and the linear models using a hierarchical greedy partitioning. We evaluated TLM on the TIMIT dataset on the task of age prediction from voice, where it outperformed state-of-the-art deep learning models. The source code will be made publicly available1. Dareen Alharthi, Mahsa Zamani, Bhiksha Raj, Rita Singh |
ICASSP | 3 |
| 2025 | Revisiting Acoustic Features for Robust ASRabstractAutomatic Speech Recognition (ASR) systems must be robust to the myriad types of noises present in real-world environments including environmental noise, room impulse response, special effects as well as attacks by malicious actors (adversarial attacks). Recent works seek to improve accuracy and robustness by developing novel Deep Neural Networks (DNNs) and curating diverse training datasets for them, while using relatively simple acoustic features. While this approach improves robustness to the types of noise present in the training data, it confers limited robustness against unseen noises and negligible robustness to adversarial attacks. In this paper, we revisit the approach of earlier works that developed acoustic features inspired by biological auditory perception that could be used to perform accurate and robust ASR. In contrast, Specifically, we evaluate the ASR accuracy and robustness of several biologically inspired acoustic features. In addition to several features from prior works, such as gammatone filterbank features (GammSpec), we also propose two new acoustic features called frequency masked spectrogram (FreqMask) and difference of gammatones spectrogram (DoGSpec) to simulate the neuro-psychological phenomena of frequency masking and lateral suppression. Experiments on diverse models and datasets show that (1) DoGSpec achieves significantly better robustness than the highly popular log mel spectrogram (LogMelSpec) with minimal accuracy degradation, and (2) GammSpec achieves better accuracy and robustness to non-adversarial noises from the Speech Robust Bench benchmark, but it is outperformed by DoGSpec against adversarial attacks. Muhammad A. Shah, Bhiksha Raj |
ICASSP | 2 |
| 2025 | Toward Material-Agnostic System Identification From Videos
Chunjiang Liu, Charles Herrmann, Junhwa Hur, Yinxiao Li, Ming-Hsuan Yang 0001, Bhiksha Raj, Min Xu 0009 |
ICCV | 9 |
| 2025 | ImageFolder: Autoregressive Image Generation with Folded TokensabstractImage tokenizers are crucial for visual generative models, \eg, diffusion models (DMs) and autoregressive (AR) models, as they construct the latent representation for modeling. Increasing token length is a common approach to improve image reconstruction quality. However, tokenizers with longer token lengths are not guaranteed to achieve better generation quality. There exists a trade-off between reconstruction and generation quality regarding token length. In this paper, we investigate the impact of token length on both image reconstruction and generation and provide a flexible solution to the tradeoff. We propose \textbf{ImageFolder}, a semantic tokenizer that provides spatially aligned image tokens that can be folded during autoregressive modeling to improve both efficiency and quality. To enhance the representative capability without increasing token length, we leverage dual-branch product quantization to capture different contexts of images. Specifically, semantic regularization is introduced in one branch to encourage compacted semantic information while another branch is designed to capture pixel-level details. Extensive experiments demonstrate the superior quality of image generation and shorter token length with ImageFolder tokenizer. Xiang Li 0106, Hao Chen 0102, Jason Kuen, Jiuxiang Gu, Bhiksha Raj |
ICLR | 6 |
| 2025 | ADIFF: Explaining audio difference using natural languageabstractUnderstanding and explaining differences between audio recordings is crucial for fields like audio forensics, quality assessment, and audio generation. This involves identifying and describing audio events, acoustic scenes, signal characteristics, and their emotional impact on listeners. This paper stands out as the first work to comprehensively study the task of explaining audio differences and then propose benchmark, baselines for the task. First, we present two new datasets for audio difference explanation derived from the AudioCaps and Clotho audio captioning datasets. Using Large Language Models (LLMs), we generate three levels of difference explanations: (1) concise descriptions of audio events and objects, (2) brief sentences about audio events, acoustic scenes, and signal properties, and (3) comprehensive explanations that include semantics and listener emotions. For the baseline, we use prefix tuning where audio embeddings from two audio files are used to prompt a frozen language model. Our empirical analysis and ablation studies reveal that the naive baseline struggles to distinguish perceptually similar sounds and generate detailed tier 3 explanations. To address these limitations, we propose ADIFF, which introduces a cross-projection module, position captioning, and a three-step training process to enhance the model’s ability to produce detailed explanations. We evaluate our model using objective metrics and human evaluation and show our model enhancements lead to significant improvements in performance over naive baseline and SoTA Audio-Language Model (ALM) Qwen Audio. Lastly, we conduct multiple ablation studies to study the effects of cross-projection, language model parameters, position captioning, third stage fine-tuning, and present our findings. Our benchmarks, findings, and strong baseline pave the way for nuanced and human-like explanations of audio differences. Soham Deshmukh, Rita Singh, Bhiksha Raj |
ICLR | 4 |
| 2025 | Speech Robust Bench: A Robustness Benchmark For Speech RecognitionabstractAs Automatic Speech Recognition (ASR) models become ever more pervasive, it is important to ensure that they make reliable predictions under corruptions present in the physical and digital world. We propose Speech Robust Bench (SRB), a comprehensive benchmark for evaluating the robustness of ASR models to diverse corruptions. SRB is composed of 114 input perturbations which simulate an heterogeneous range of corruptions that ASR models may encounter when deployed in the wild. We use SRB to evaluate the robustness of several state-of-the-art ASR models and observe that model size and certain modeling choices such as the use of discrete representations, or self-training appear to be conducive to robustness. We extend this analysis to measure the robustness of ASR models on data from various demographic subgroups, namely English and Spanish speakers, and males and females. Our results revealed noticeable disparities in the model's robustness across subgroups. We believe that SRB will significantly facilitate future research towards robust ASR models, by making it easier to conduct comprehensive and comparable robustness evaluations. Muhammad A. Shah, David Solans Noguero, Mikko A. Heikkilä, Bhiksha Raj, Nicolas Kourtellis |
ICLR | 4 |
| 2025 | Unsupervised Disentanglement of Content and Style via Variance-Invariance ConstraintsabstractWe contribute an unsupervised method that effectively learns disentangled content and style representations from sequences of observations. Unlike most disentanglement algorithms that rely on domain-specific labels or knowledge, our method is based on the insight of domain-general statistical differences between content and style --- content varies more among different fragments within a sample but maintains an invariant vocabulary across data samples, whereas style remains relatively invariant within a sample but exhibits more significant variation across different samples. We integrate such inductive bias into an encoder-decoder architecture and name our method after V3 (variance-versus-invariance). Experimental results show that V3 generalizes across multiple domains and modalities, successfully learning disentangled content and style representations, such as pitch and timbre from music audio, digit and color from images of hand-written digits, and action and character appearance from simple animations. V3 demonstrates strong disentanglement performance compared to existing unsupervised methods, along with superior out-of-distribution generalization and few-shot learning capabilities compared to supervised counterparts. Lastly, symbolic-level interpretability emerges in the learned content codebook, forging a near one-to-one alignment between machine representation and human knowledge. Ziyu Wang 0008, Bhiksha Raj, Gus Xia |
ICLR | 3 |
| 2025 | Scalable Benchmarking and Robust Learning for Noise-Free Ego-Motion and 3D Reconstruction from Noisy VideoabstractWe aim to redefine robust ego-motion estimation and photorealistic 3D reconstruction by addressing a critical limitation: the reliance on noise-free data in existing models. While such sanitized conditions simplify evaluation, they fail to capture the unpredictable, noisy complexities of real-world environments. Dynamic motion, sensor imperfections, and synchronization perturbations lead to sharp performance declines when these models are deployed in practice, revealing an urgent need for frameworks that embrace and excel under real-world noise.
To bridge this gap, we tackle three core challenges: scalable data generation, comprehensive benchmarking, and model robustness enhancement. First, we introduce a scalable noisy data synthesis pipeline that generates diverse datasets simulating complex motion, sensor imperfections, and synchronization errors. Second, we leverage this pipeline to create Robust-Ego3D, a benchmark rigorously designed to expose noise-induced performance degradation, highlighting the limitations of current learning-based methods in ego-motion accuracy and 3D reconstruction quality. Third, we propose Correspondence-guided Gaussian Splatting (CorrGS), a novel method that progressively refines an internal clean 3D representation by aligning noisy observations with rendered RGB-D frames from clean 3D map, enhancing geometric alignment and appearance restoration through visual correspondence.
Extensive experiments on synthetic and real-world data demonstrate that CorrGS consistently outperforms prior state-of-the-art methods, particularly in scenarios involving rapid motion and dynamic illumination. We will release our code and benchmark to advance robust 3D vision, setting a new standard for ego-motion estimation and high-fidelity reconstruction in noisy environments. Xiaohao Xu, Tianyi Zhang 0014, Shibo Zhao, Xiang Li 0106, Bhiksha Raj, Matthew Johnson-Roberson, Sebastian A. Scherer, Xiaonan Huang |
ICLR | 8 |
| 2025 | Masked Autoencoders Are Effective Tokenizers for Diffusion ModelsabstractRecent advances in latent diffusion models have demonstrated their effectiveness for high-resolution image synthesis. However, the properties of the latent space from tokenizer for better learning and generation of diffusion models remain under-explored. Theoretically and empirically, we find that improved generation quality is closely tied to the latent distributions with better structure, such as the ones with fewer Gaussian Mixture modes and more discriminative features. Motivated by these insights, we propose MAETok, an autoencoder (AE) leveraging mask modeling to learn semantically rich latent space while maintaining reconstruction fidelity.
Extensive experiments validate our analysis, demonstrating that the variational form of autoencoders is not necessary, and a discriminative latent space from AE alone enables state-of-the-art performance on ImageNet generation using only 128 tokens. MAETok achieves significant practical improvements, enabling a gFID of 1.69 with 76× faster training and 31× higher inference throughput for 512×512 generation. Our findings show that the structure of the latent space, rather than variational constraints, is crucial for effective diffusion models. Code and trained models will be released. Hao Chen 0102, Yujin Han, Fangyi Chen, Xiang Li 0106, Yidong Wang 0003, Jindong Wang 0001, Ze Wang 0008, Zicheng Liu 0001, Difan Zou, Bhiksha Raj |
ICML | 10 |
| 2025 | uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data RegimesabstractAbdul Waheed, Karima Kadaoui, Bhiksha Raj, Muhammad Abdul-Mageed. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Karima Kadaoui, Bhiksha Raj, Muhammad Abdul-Mageed |
NAACL (Long Papers) | 3 |
| 2025 | Mellow: a small audio language model for reasoningabstractMultimodal Audio-Language Models (ALMs) can understand and reason over both audio and text. Typically, reasoning performance correlates with model size, with the best results achieved by models exceeding 8 billion parameters. However, no prior work has explored enabling small audio-language models to perform reasoning tasks, despite the potential applications for edge devices. To address this gap, we introduce Mellow, a small Audio-Language Model specifically designed for reasoning. Mellow achieves state-of-the-art performance among existing small audio-language models and surpasses several larger models in reasoning capabilities. For instance, Mellow scores 52.11 on MMAU, comparable to SoTA Qwen2 Audio (which scores 52.5) while using 50 times fewer parameters and being trained on 60 times less data (audio hrs). To train Mellow, we introduce ReasonAQA, a dataset designed to enhance audio-grounded reasoning in models. It consists of a mixture of existing datasets (30\% of the data) and synthetically generated data (70\%). The synthetic dataset is derived from audio captioning datasets, where Large Language Models (LLMs) generate detailed and multiple-choice questions focusing on audio events, objects, acoustic scenes, signal properties, semantics, and listener emotions. To evaluate Mellow’s reasoning ability, we benchmark it on a diverse set of tasks, assessing on both in-distribution and out-of-distribution data, including audio understanding, deductive reasoning, and comparative reasoning. Finally, we conduct extensive ablation studies to explore the impact of projection layer choices, synthetic data generation methods, and language model pretraining on reasoning performance. Our training dataset, findings, and baseline pave the way for developing small ALMs capable of reasoning. Soham Deshmukh, Satvik Dixit, Rita Singh, Bhiksha Raj |
NeurIPS | 4 |
| 2025 | On Fairness of Unified Multimodal Large Language Model for Image GenerationabstractUnified multimodal large language models (U-MLLMs) have demonstrated impressive performance in end-to-end visual understanding and generation tasks. However, compared to generation-only systems (e.g., Stable Diffusion), the unified architecture of U-MLLMs introduces new risks of propagating demographic stereotypes. In this paper, we benchmark several state-of-the-art U-MLLMs and show that they exhibit significant gender and race biases in the generated outputs. To diagnose the source of these biases, we propose a locate-then-fix framework: we first audit the vision and language components — using techniques such as linear probing and controlled generation — and find that the language model appears to be a primary origin of the observed generative bias. Moreover, we observe a ``partial alignment'' phenomenon, where the U-MLLMs exhibit less bias in understanding tasks yet produce substantially biased images. To address this, we introduce a novel \emph{balanced preference loss} that enforces uniform generation probabilities across demographics by leveraging a synthetically balanced dataset. Extensive experiments show that our approach significantly reduces demographic bias while preserving semantic fidelity and image quality. Our findings underscore the need for targeted debiasing strategies in unified multimodal systems and introduce a practical approach to mitigate biases. Hao Chen 0102, Jindong Wang 0001, Bhiksha Raj |
NeurIPS | 5 |
| 2025 | Directed-Tokens: A Robust Multi-Modality Alignment Approach to Large Language-Vision ModelsabstractLarge multimodal models (LMMs) have gained impressive performance due to their outstanding capability in various understanding tasks. However, these models still suffer from some fundamental limitations related to robustness and generalization due to the alignment and correlation between visual and textual features. In this paper, we introduce a simple but efficient learning mechanism for improving the robust alignment between visual and textual modalities by solving shuffling problems. In particular, the proposed approach can improve reasoning capability, visual understanding, and cross-modality alignment by introducing two new tasks: reconstructing the image order and the text order into the LMM's pre-training and fine-tuning phases. In addition, we propose a new directed-token approach to capture visual and textual knowledge, enabling the capability to reconstruct the correct order of visual inputs. Then, we introduce a new Image-to-Response Guided loss to further improve the visual understanding of the LMM in its responses. The proposed approach consistently achieves state-of-the-art (SoTA) performance compared with prior LMMs on academic task-oriented and instruction-following LMM benchmarks. Thanh-Dat Truong, Huu-Thien Tran, Tran Thai Son, Bhiksha Raj, Khoa Luu |
NeurIPS | 4 |
| 2025 | Autoregressive Temporal Modeling for Advanced Tracking-by-DiffusionabstractObject tracking is a widely studied computer vision task with video and instance analysis applications. While paradigms such as tracking-by-regression , -detection , -attention have advanced the field, generative modeling offers new potential. Although some studies explore the generative process in instance-based understanding tasks, they rely on prediction refinement in the coordinate space rather than the visual domain. Instead, this paper presents Tracking-by-Diffusion , a novel paradigm for object tracking in video, leveraging visual generative models via the perspective of autoregressive models. This paradigm demonstrates broad applicability across point, box, and mask modalities while uniquely enabling textual guidance. We present DIFTracker, a framework that utilizes iterative latent variable diffusion models to redefine tracking as a next-frame reconstruction task. Our approach uniquely combines spatial and temporal dependencies in video data, offering a unified solution that encompasses existing tracking paradigms within a single Inversion-Reconstruction process. DIFTracker operates online and auto-regressively, enabling flexible instance-based video understanding. It allows us to overcome difficulties in variable-length video understanding encountered by video-inflated models and perform superior performance on seven benchmarks across five modalities. This paper not only introduces a new perspective on visual autoregressive modeling in understanding sequential visual data, specifically videos, but also provides robust theoretical validations and demonstrates broader applications in visual tracking and computer vision. Pha A. Nguyen, Rishi Madhok, Bhiksha Raj, Khoa Luu |
Int. J. Comput. Vis. | 3 |
| 2025 | Impact of Noisy Supervision in Foundation Model LearningabstractFoundation models are usually pre-trained on large-scale datasets and then adapted to different downstream tasks through tuning. This pre-training and then fine-tuning paradigm has become a standard practice in deep learning. However, the large-scale pre-training datasets, often inaccessible or too expensive to handle, can contain label noise that may adversely affect the generalization of the model and pose unexpected risks. This paper stands out as the first work to comprehensively understand and analyze the nature of noise in pre-training datasets and then effectively mitigate its impacts on downstream tasks. Specifically, through extensive experiments of fully-supervised and image-text contrastive pre-training on synthetic noisy ImageNet-1 K, YFCC15 M, and CC12 M datasets, we demonstrate that, while slight noise in pre-training can benefit in-domain (ID) performance, where the training and testing data share a similar distribution, it always deteriorates out-of-domain (OOD) performance, where training and testing distributions are significantly different. These observations are agnostic to scales of pre-training datasets, pre-training noise types, model architectures, pre-training objectives, downstream tuning methods, and downstream applications. We empirically ascertain that the reason behind this is that the pre-training noise shapes the feature space differently. We then propose a tuning method (NMTune) to affine the feature space to mitigate the malignant effect of noise and improve generalization, which is applicable in both parameter-efficient and black-box tuning manners, considering one may not be able to access or fully fine-tune the pre-trained models. We additionally conduct extensive experiments on popular vision and language models, including APIs, which are supervised and self-supervised pre-trained on realistic noisy data for evaluation. Our analysis and results demonstrate the importance of this novel and fundamental research direction, which we term as Noisy Model Transfer Learning. Hao Chen 0102, Ran Tao 0013, Hongxin Wei, Xing Xie 0001, Masashi Sugiyama, Bhiksha Raj, Jindong Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | Introduction to the Special Issue on Evaluations of Large Language Models: Part 1abstractNo abstract available. Jindong Wang 0001, Linyi Yang, Sunayana Sitaram, Qiang Yang 0001, Bhiksha Raj |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2024 | Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?abstractReference summaries for abstractive speech summarization require human annotation, which can be performed by listening to an audio recording or by reading textual transcripts of the recording.In this paper, we examine whether summaries based on annotators listening to the recordings differ from those based on annotators reading transcripts.Using existing intrinsic evaluation based on human evaluation, automatic metrics, LLM-based evaluation, and a retrieval-based reference-free method.We find that summaries are indeed different based on the source modality, and that speechbased summaries are more factually consistent and information-selective than transcript-based summaries.Meanwhile, transcript-based summaries are impacted by recognition errors in the source, and expert-written summaries are more informative and reliable.We make all the collected data and analysis code public 1 to facilitate the reproduction of our work and advance research in this area. Roshan S. Sharma, Suwon Shon, Mark Lindsey, Hira Dhamyal, Bhiksha Raj |
ACL (1) | 5 |
| 2024 | QDFormer: Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic DecompositionabstractAudiovisual segmentation (AVS) is a challenging task that aims to segment visual objects in videos according to their associated acoustic cues. With multiple sound sources and background disturbances involved, establishing robust correspondences between audio and visual contents poses unique challenges due to (1) complex entanglement across sound sources and (2) frequent changes in the occurrence of distinct sound events. Assuming sound events occur in- dependently, the multi-source semantic space can be rep- resented as the Cartesian product of single-source sub- spaces. We are motivated to decompose the multi-source audio semantics into single-source semantics for more ef- fective interactions with visual content. We propose a se- mantic decomposition method based on product quanti- zation, where the multi-source semantics can be decom- posed and represented by several disentangled and noise- suppressed single-source semantics. Furthermore, we in- troduce a global-to-local quantization mechanism, which distills knowledge from stable global (clip-level) features into local (frame-level) ones, to handle frequent changes in audio semantics. Extensive experiments demonstrate that our semantically decomposed audio representation signifi- cantly improves AVS performance, e.g., +21.2% mIoU on the challenging AVS-Semantic benchmark with ResNet50 backbone. Xiang Li 0106, Jinglu Wang, Xiaohao Xu, Xiulian Peng, Rita Singh, Yan Lu 0001, Bhiksha Raj |
CVPR | 7 |
| 2024 | Synergistic Global-Space Camera and Human Reconstruction from VideosabstractRemarkable strides have been made in reconstructing static scenes or human bodies from monocular videos. Yet, the two problems have largely been approached independently, without much synergy. Most visual SLAM methods can only reconstruct camera trajectories and scene structures up to scale, while most HMR methods reconstruct human meshes in metric scale but fall short in reasoning with cameras and scenes. This work introduces Synergistic Camera and Human Reconstruction (SynCHMR) to marry the best of both worlds. Specifically, we design Human-aware Metric SLAM to reconstruct metric-scale camera poses and scene point clouds using camera-frame HMR as a strong prior, addressing depth, scale, and dynamic ambiguities. Conditioning on the dense scene recovered, we further learn a Scene-aware SMPL Denoiser to enhance world-frame HMR by incorporating spatio-temporal coherency and dynamic scene constraints. Together, they lead to consistent reconstructions of camera trajectories, human meshes, and dense scene point clouds in a common world frame. Tuanfeng Y. Wang, Bhiksha Raj, Min Xu 0009, Jimei Yang, Chun-Hao Paul Huang |
CVPR | 3 |
| 2024 | R2-Bench: Benchmarking the Robustness of Referring Perception Models Under Perturbations
Xiang Li 0106, Jinglu Wang, Xiaohao Xu, Rita Singh, Kashu Yamazaki, Hao Chen 0102, Xiaonan Huang, Bhiksha Raj |
ECCV (9) | 9 |
| 2024 | Importance of Negative Sampling in Weak Label LearningabstractWeak-label learning is a challenging task that requires learning from data "bags" containing positive and negative instances, but only the bag labels are known. The pool of negative instances is usually larger than positive instances, thus making selecting the most informative negative instance critical for performance. Such a selection strategy for negative instances from each bag is an open problem that has not been well studied for weak-label learning. In this paper, we study several sampling strategies that can measure the usefulness of negative instances for weak-label learning and select them accordingly. We test our method on CIFAR-10 and AudioSet datasets and show that it improves the weak-label classification performance and reduces the computational cost compared to random sampling methods. Our work reveals that negative instances are not all equally irrelevant, and selecting them wisely can benefit weak-label learning. Ankit Shah 0001, Fuyu Tang, Zelin Ye, Rita Singh, Bhiksha Raj |
ICASSP | 5 |
| 2024 | Training Audio Captioning Models without AudioabstractAutomated Audio Captioning (AAC) is the task of generating natural language descriptions given an audio stream. A typical AAC system requires manually curated training data of audio segments and corresponding text caption annotations. The creation of these audio-caption pairs is costly, resulting in general data scarcity for the task. In this work, we address this major limitation and propose an approach to train AAC systems using only text. Our approach leverages the multimodal space of contrastively trained audio-text models, such as CLAP. During training, a decoder generates captions conditioned on the pretrained CLAP text encoder. During inference, the text encoder is replaced with the pretrained CLAP audio encoder. To bridge the modality gap between text and audio embeddings, we propose the use of noise injection or a learnable adapter, during training. We find that the proposed text-only framework performs competitively with stateof-the-art models trained with paired audio, showing that efficient text-to-audio transfer is possible. Finally, we showcase both stylized audio captioning and caption enrichment while training without audio or human-created text captions. Soham Deshmukh, Benjamin Elizalde, Dimitra Emmanouilidou, Bhiksha Raj, Rita Singh, Huaming Wang |
ICASSP | 4 |
| 2024 | Prompting Audios Using Acoustic Properties for Emotion RepresentationabstractEmotions lie on a continuum, but current models treat emotions as a finite valued discrete variable. This representation does not capture the diversity in the expression of emotion. To better represent emotions we propose the use of natural language descriptions (or prompts). In this work, we address the challenge of automatically generating these prompts and training a model to better learn emotion representations from audio and prompt pairs. We use acoustic properties that are correlated to emotion like pitch, intensity, speech rate, and articulation rate to automatically generate prompts i.e. ‘acoustic prompts’. We use a contrastive learning objective to map speech to their respective acoustic prompts. We evaluate our model on Emotion Audio Retrieval and Speech Emotion Recognition. Our results show that the acoustic prompts significantly improve the model’s performance in EAR, in various Precision@K metrics. In SER, we observe a 3.8% relative accuracy improvement on the Ravdess dataset. Hira Dhamyal, Benjamin Elizalde, Soham Deshmukh, Huaming Wang, Bhiksha Raj, Rita Singh |
ICASSP | 5 |
| 2024 | Dynamic-Superb: Towards a Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark For SpeechabstractText language models have shown remarkable zero-shot capability in generalizing to unseen tasks when provided with well-formulated instructions. However, existing studies in speech processing primarily focus on limited or specific tasks. Moreover, the lack of standardized benchmarks hinders a fair comparison across different approaches. Thus, we present Dynamic-SUPERB, a benchmark designed for building universal speech models capable of leveraging instruction tuning to perform multiple tasks in a zero-shot fashion. To achieve comprehensive coverage of diverse speech tasks and harness instruction tuning, we invite the community to collaborate and contribute, facilitating the dynamic growth of the benchmark. To initiate, Dynamic-SUPERB features 55 evaluation instances by combining 33 tasks and 22 datasets. This spans a broad spectrum of dimensions, providing a comprehensive platform for evaluation. Additionally, we propose several approaches to establish benchmark baselines. These include the utilization of speech models, text language models, and the multimodal encoder. Evaluation results indicate that while these baselines perform reasonably on seen tasks, they struggle with unseen ones. We release all materials to the public and welcome researchers to collaborate on the project, advancing technologies in the field together1. Chien-Yu Huang, Ke-Han Lu, Shih-Heng Wang, Chi-Yuan Hsiao, Chun-Yi Kuan, Siddhant Arora, Kai-Wei Chang 0001, Jiatong Shi, Yifan Peng 0003, Roshan S. Sharma, Shinji Watanabe 0001, Bhiksha Raj, Shady Shehata, Hung-yi Lee |
ICASSP | 13 |
| 2024 | AugSumm: Towards Generalizable Speech Summarization Using Synthetic Labels from Large Language ModelsabstractAbstractive speech summarization (SSUM) aims to generate humanlike summaries from speech. Given variations in information captured and phrasing, recordings can be summarized in multiple ways. Therefore, it is more reasonable to consider a probabilistic distribution of all potential summaries rather than a single summary. However, conventional SSUM models are mostly trained and evaluated with a single ground-truth (GT) human-annotated deterministic summary for every recording. Generating multiple human references would be ideal to better represent the distribution statistically, but is impractical because annotation is expensive. We tackle this challenge by proposing AugSumm, a method to leverage large language models (LLMs) as a proxy for human annotators to generate augmented summaries for training and evaluation. First, we explore prompting strategies to generate synthetic summaries from ChatGPT. We validate the quality of synthetic summaries using multiple metrics including human evaluation, where we find that summaries generated using AugSumm are perceived as more valid to humans. Second, we develop methods to utilize synthetic summaries in training and evaluation. Experiments on How2 demonstrate that pre-training on synthetic summaries and fine-tuning on GT summaries improves ROUGE-L by 1 point on both GT and AugSumm-based test sets. AugSumm summaries are available at https://github.com/Jungjee/AugSumm. Jee-Weon Jung, Roshan S. Sharma, Bhiksha Raj, Shinji Watanabe 0001 |
ICASSP | 4 |
| 2024 | Fixed Inter-Neuron Covariability Induces Adversarial RobustnessabstractVulnerability to adversarial perturbations is a major flaw of Deep Neural Networks (DNNs) that brings their reliability in real-world scenarios into question. However, biological perception, which DNNs emulate, is highly robust to such perturbations. This suggests that DNNs diverge from biological perception and develop vulnerability to adversarial perturbations. One point of divergence is that the activity of biological neurons is correlated and the structure of this correlation tends to be static across stimuli and over time, even if it hampers performance and learning. Conversely, we observe that small perturbations of the inputs can significantly change the inter-neuron correlations. We hypothesize that constraining the DNN’s neurons to respect a fixed inter-neuron correlation structure would improve its robustness. To test this hypothesis, we develop a Self-Consistent Activation (SCA) layer, which consists of neurons whose activations are optimized to conform to a fixed, but learned, correlation pattern. We train models with our SCA layer on image and audio recognition tasks and evaluate their accuracy under AutoAttack, an ensemble of white and black box adversarial attacks. Models with an SCA layer achieved up to 6% (abs) higher accuracy than traditional MLPs, without being trained on adversarially perturbed data. The code is available at https://github.com/ahmedshah1494/SCA-Layer. Muhammad A. Shah, Bhiksha Raj |
ICASSP | 2 |
| 2024 | Improving Continual Learning of Acoustic Scene Classification via Mutual Information OptimizationabstractContinual learning, which aims to incrementally accumulate knowledge, has been an increasingly significant but challenging research topic for deep models that are prone to catastrophic forgetting. In this paper, we propose a novel replay-based continual learning approach in the context of class-incremental learning in acoustic scene classification, to classify audio recordings into an expanding set of classes that characterize the acoustic scenes. Our approach is improving both the modeling and memory selection mechanism via mutual information optimization in continual learning. By regarding incremental classes of acoustic scenes as different tasks, our model is expected to learn both task-agnostic and task-specific knowledge by replaying representative and informative samples. This optimization also enables the model to utilize past knowledge effectively and learn from new information during continual learning. We demonstrate that our approach has a superior performance compared to existing methods on multiple datasets and continual learning evaluation metrics. Muqiao Yang, Umberto Cappellazzo, Bhiksha Raj |
ICASSP | 4 |
| 2024 | uSee: Unified Speech Enhancement And Editing with Conditional Diffusion ModelsabstractSpeech enhancement aims to improve the quality of speech signals in terms of quality and intelligibility, and speech editing refers to the process of editing the speech according to specific user needs. In this paper, we propose a Unified Speech Enhancement and Editing (uSee) model with conditional diffusion models to handle various tasks at the same time in a generative manner. Specifically, by providing multiple types of conditions including self-supervised learning embeddings and proper text prompts to the score-based diffusion model, we can enable controllable generation of the unified speech enhancement and editing model to perform corresponding actions on the source speech. Our experiments show that our proposed uSee model can achieve superior performance in both speech denoising and dereverberation compared to other related generative speech enhancement models, and can perform speech editing given desired environmental sound text description, signal-to-noise ratios (SNR), and room impulse responses (RIR). Demos of the generated speech are available at https://muqiaoy.github.io/usee. Muqiao Yang, Yong Xu 0004, Zhongweiyang Xu, Heming Wang, Bhiksha Raj, Dong Yu 0001 |
ICASSP | 6 |
| 2024 | Understanding and Mitigating the Label Noise in Pre-training on Downstream TasksabstractPre-training on large-scale datasets and then fine-tuning on downstream tasks have become a standard practice in deep learning. However, pre-training data often contain label noise that may adversely affect the generalization of the model. This paper aims to understand the nature of noise in pre-training datasets and to mitigate its impact on downstream tasks. More specifically, through extensive experiments of supervised pre-training models on synthetic noisy ImageNet-1K and YFCC15M datasets, we demonstrate that while slight noise in pre-training can benefit in-domain (ID) transfer performance, where the training and testing data share the same distribution, it always deteriorates out-of-domain (OOD) performance, where training and testing data distribution are different. We empirically verify that the reason behind is noise in pre-training shapes the feature space differently. We then propose a light-weight black-box tuning method (NMTune) to affine the feature space to mitigate the malignant effect of noise and improve generalization on both ID and OOD tasks, considering one may not be able to fully fine-tune or even access the pre-trained models. We conduct practical experiments on popular vision and language models that are pre-trained on noisy data for evaluation of our approach. Our analysis and results show the importance of this interesting and novel research direction, which we term Noisy Model Learning. Hao Chen 0102, Jindong Wang 0001, Ankit Shah 0001, Ran Tao 0013, Hongxin Wei, Xing Xie 0001, Masashi Sugiyama, Bhiksha Raj |
ICLR | 8 |
| 2024 | A General Framework for Learning from Weak SupervisionabstractWeakly supervised learning generally faces challenges in applicability to various scenarios with diverse weak supervision and in scalability due to the complexity of existing algorithms, thereby hindering the practical deployment. This paper introduces a general framework for learning from weak supervision (GLWS) with a novel algorithm. Central to GLWS is an Expectation-Maximization (EM) formulation, adeptly accommodating various weak supervision sources, including instance partial labels, aggregate statistics, pairwise observations, and unlabeled data. We further present an advanced algorithm that significantly simplifies the EM computational demands using a Non-deterministic Finite Automaton (NFA) along with a forward-backward algorithm, which effectively reduces time complexity from quadratic or factorial often required in existing solutions to linear scale. The problem of learning from arbitrary weak supervision is therefore converted to the NFA modeling of them. GLWS not only enhances the scalability of machine learning models but also demonstrates superior performance and versatility across 11 weak supervision scenarios. We hope our work paves the way for further advancements and practical deployment in this field. Hao Chen 0102, Jindong Wang 0001, Lei Feng 0006, Xiang Li 0106, Yidong Wang 0003, Xing Xie 0001, Masashi Sugiyama, Rita Singh, Bhiksha Raj |
ICML | 9 |
| 2024 | Completing Visual Objects via Bridging Generation and SegmentationabstractThis paper presents a novel approach to object completion, with the primary goal of reconstructing a complete object from its partially visible components. Our method, named MaskComp, delineates the completion process through iterative stages of generation and segmentation. In each iteration, the object mask is provided as an additional condition to boost image generation, and, in return, the generated images can lead to a more accurate mask by fusing the segmentation of images. We demonstrate that the combination of one generation and one segmentation stage effectively functions as a mask denoiser. Through alternation between the generation and segmentation stages, the partial object mask is progressively refined, providing precise shape guidance and yielding superior object completion results. Our experiments demonstrate the superiority of MaskComp over existing approaches, e.g., ControlNet and Stable Diffusion, establishing it as an effective solution for object completion. Xiang Li 0106, Yinpeng Chen, Chung-Ching Lin, Hao Chen 0102, Kai Hu 0010, Rita Singh, Bhiksha Raj, Zicheng Liu 0001 |
ICML | 7 |
| 2024 | Fashion Image Retrieval with Occlusion
Jimin Sohn, Haeji Jung, Zhiwen Yan, Vibha Masti, Bhiksha Raj |
ICPR (21) | 6 |
| 2024 | SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios
Hazim T. Bukhari, Soham Deshmukh, Hira Dhamyal, Bhiksha Raj, Rita Singh |
INTERSPEECH | 4 |
| 2024 | PAM: Prompting Audio-Language Models for Audio Quality Assessment
Soham Deshmukh, Dareen Alharthi, Benjamin Elizalde, Hannes Gamper, Mahmoud Al Ismail, Rita Singh, Bhiksha Raj, Huaming Wang |
INTERSPEECH | 7 |
| 2024 | Domain Adaptation for Contrastive Audio-Language Models
Soham Deshmukh, Rita Singh, Bhiksha Raj |
INTERSPEECH | 3 |
| 2024 | DeWinder: Single-Channel Wind Noise Reduction using Ultrasound SensingabstractThe quality of audio recordings in outdoor environments is often degraded by the presence of wind. Mitigating the impact of wind noise on the perceptual quality of single-channel speech remains a significant challenge due to its non-stationary characteristics. Prior work in noise suppression treats wind noise as a general background noise without explicit modeling of its characteristics. In this paper, we leverage ultrasound as an auxiliary modality to explicitly sense the airflow and characterize the wind noise. We propose a multi-modal deep-learning framework to fuse the ultrasonic Doppler features and speech signals for wind noise reduction. Our results show that DeWinder can significantly improve the noise reduction capabilities of state-of-the-art speech enhancement models. Kuang Yuan, Swarun Kumar, Bhiksha Raj |
INTERSPEECH | 4 |
| 2024 | AutoPRM: Automating Procedural Supervision for Multi-Step Reasoning via Controllable Question DecompositionabstractZhaorun Chen, Zhuokai Zhao, Zhihong Zhu, Ruiqi Zhang, Xiang Li, Bhiksha Raj, Huaxiu Yao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zhaorun Chen, Zhuokai Zhao, Zhihong Zhu 0001, Xiang Li 0106, Bhiksha Raj, Huaxiu Yao |
NAACL-HLT | 6 |
| 2024 | Imprecise Label Learning: A Unified Framework for Learning with Various Imprecise Label ConfigurationsabstractLearning with reduced labeling standards, such as noisy label, partial label, and supplementary unlabeled data, which we generically refer to as imprecise label, is a commonplace challenge in machine learning tasks. Previous methods tend to propose specific designs for every emerging imprecise label configuration, which is usually unsustainable when multiple configurations of imprecision coexist.
In this paper, we introduce imprecise label learning (ILL), a framework for the unification of learning with various imprecise label configurations. ILL leverages expectation-maximization (EM) for modeling the imprecise label information, treating the precise labels as latent variables. Instead of approximating the correct labels for training, it considers the entire distribution of all possible labeling entailed by the imprecise information. We demonstrate that ILL can seamlessly adapt to partial label learning, semi-supervised learning, noisy label learning, and, more importantly, a mixture of these settings, with closed-form learning objectives derived from the unified EM modeling. Notably, ILL surpasses the existing specified techniques for handling imprecise labels, marking the first practical and unified framework with robust and effective performance across various challenging settings. We hope our work will inspire further research on this topic, unleashing the full potential of ILL in wider scenarios where precise labels are expensive and complicated to obtain. Hao Chen 0102, Ankit Shah 0001, Jindong Wang 0001, Ran Tao 0013, Yidong Wang 0003, Xiang Li 0106, Xing Xie 0001, Masashi Sugiyama, Rita Singh, Bhiksha Raj |
NeurIPS | 10 |
| 2024 | Slight Corruption in Pre-training Data Makes Better Diffusion ModelsabstractDiffusion models (DMs) have shown remarkable capabilities in generating realistic high-quality images, audios, and videos.
They benefit significantly from extensive pre-training on large-scale datasets, including web-crawled data with paired data and conditions, such as image-text and image-class pairs.
Despite rigorous filtering, these pre-training datasets often inevitably contain corrupted pairs where conditions do not accurately describe the data.
This paper presents the first comprehensive study on the impact of such corruption in pre-training data of DMs.
We synthetically corrupt ImageNet-1K and CC3M to pre-train and evaluate over $50$ conditional DMs.
Our empirical findings reveal that various types of slight corruption in pre-training can significantly enhance the quality, diversity, and fidelity of the generated images across different DMs, both during pre-training and downstream adaptation stages.
Theoretically, we consider a Gaussian mixture model and prove that slight corruption in the condition leads to higher entropy and a reduced 2-Wasserstein distance to the ground truth of the data distribution generated by the corruptly trained DMs.
Inspired by our analysis, we propose a simple method to improve the training of DMs on practical datasets by adding condition embedding perturbations (CEP).
CEP significantly improves the performance of various DMs in both pre-training and downstream tasks.
We hope that our study provides new insights into understanding the data and pre-training processes of DMs. Hao Chen 0102, Yujin Han, Diganta Misra, Xiang Li 0106, Kai Hu 0010, Difan Zou, Masashi Sugiyama, Jindong Wang 0001, Bhiksha Raj |
NeurIPS | 9 |
| 2024 | EAGLE: Efficient Adaptive Geometry-based Learning in Cross-view UnderstandingabstractUnsupervised Domain Adaptation has been an efficient approach to transferring the semantic segmentation model across data distributions. Meanwhile, the recent Open-vocabulary Semantic Scene understanding based on large-scale vision language models is effective in open-set settings because it can learn diverse concepts and categories. However, these prior methods fail to generalize across different camera views due to the lack of cross-view geometric modeling. At present, there are limited studies analyzing cross-view learning. To address this problem, we introduce a novel Unsupervised Cross-view Adaptation Learning approach to modeling the geometric structural change across views in Semantic Scene Understanding. First, we introduce a novel Cross-view Geometric Constraint on Unpaired Data to model structural changes in images and segmentation masks across cameras. Second, we present a new Geodesic Flow-based Correlation Metric to efficiently measure the geometric structural changes across camera views. Third, we introduce a novel view-condition prompting mechanism to enhance the view-information modeling of the open-vocabulary segmentation network in cross-view adaptation learning. The experiments on different cross-view adaptation benchmarks have shown the effectiveness of our approach in cross-view modeling, demonstrating that we achieve State-of-the-Art (SOTA) performance compared to prior unsupervised domain adaptation and open-vocabulary semantic segmentation methods. Thanh-Dat Truong, Utsav Prabhu, Dongyi Wang, Bhiksha Raj, Susan Gauch, Jeyamkondan Subbiah, Khoa Luu |
NeurIPS | 4 |
| 2024 | Metric from Human: Zero-shot Monocular Metric Depth Estimation via Test-time AdaptationabstractMonocular depth estimation (MDE) is fundamental for deriving 3D scene structures from 2D images. While state-of-the-art monocular relative depth estimation (MRDE) excels in estimating relative depths for in-the-wild images, current monocular metric depth estimation (MMDE) approaches still face challenges in handling unseen scenes. Since MMDE can be viewed as the composition of MRDE and metric scale recovery, we attribute this difficulty to scene dependency, where MMDE models rely on scenes observed during supervised training for predicting scene scales during inference. To address this issue, we propose to use humans as landmarks for distilling scene-independent metric scale priors from generative painting models. Our approach, Metric from Human (MfH), bridges from generalizable MRDE to zero-shot MMDE in a generate-and-estimate manner. Specifically, MfH generates humans on the input image with generative painting and estimates human dimensions with an off-the-shelf human mesh recovery (HMR) model. Based on MRDE predictions, it propagates the metric information from painted humans to the contexts, resulting in metric depth estimations for the original input. Through this annotation-free test-time adaptation, MfH achieves superior zero-shot performance in MMDE, demonstrating its strong generalization ability. Hengwei Bian, Kaihua Chen, Pengliang Ji, Liao Qu, Shao-yu Lin, Weichen Yu, Haoran Li 0024, Hao Chen 0102, Jun Shen 0001, Bhiksha Raj, Min Xu 0009 |
NeurIPS | 11 |
| 2024 | PDAF: A Phonetic Debiasing Attention Framework For Speaker VerificationabstractSpeaker verification systems are crucial for authenticating identity through voice. Traditionally, these systems focus on comparing feature vectors, overlooking the speech’s content. However, this paper challenges this by highlighting the importance of phonetic dominance, a measure of the frequency or duration of phonemes, as a crucial cue in speaker verification. A novel Phoneme-Debiasing Attention Framework (PDAF) is introduced, integrating with existing attention frameworks to mitigate biases caused by phonetic dominance. PDAF adjusts the weighting for each phoneme and influences feature extraction, allowing for a more nuanced analysis of speech. This approach paves the way for more accurate and reliable identity authentication through voice. Furthermore, by employing various weighting strategies, we evaluate the influence of phonetic features on the efficacy of the speaker verification system. Massa Baali, Abdulhamid Aldoobi, Hira Dhamyal, Rita Singh, Bhiksha Raj |
SLT | 5 |
| 2024 | ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs For Audio, Music, and SpeechabstractNeural codecs have become crucial to recent speech and audio generation research. In addition to signal compression capabilities, discrete codecs have also been found to enhance downstream training efficiency and compatibility with autoregressive language models. However, as extensive downstream applications are investigated, challenges have arisen in ensuring fair comparisons across diverse applications. To address these issues, we present a new open-source platform ESPnet-Codec, which is built on ESPnet and focuses on neural codec training and evaluation. ESPnet-Codec offers various recipes in audio, music, and speech for training and evaluation using several widely adopted codec models. Together with ESPnet-Codec, we present VERSA, a standalone evaluation toolkit, which provides a comprehensive evaluation of codec performance over 20 audio evaluation metrics. Notably, we demonstrate that ESPnet-Codec can be integrated into six ESPnet tasks, supporting diverse applications. Jiatong Shi, Jinchuan Tian, Yihan Wu 0008, Jee-Weon Jung, Jia Qi Yip, Yoshiki Masuyama, Yuning Wu 0001, Yuxun Tang, Massa Baali, Dareen Alharthi, Ruifan Deng, Tejes Srivastava, Alexander H. Liu, Bhiksha Raj, Qin Jin, Ruihua Song, Shinji Watanabe 0001 |
SLT | 17 |
| 2024 | A closer look at reinforcement learning-based automatic speech recognition
Fan Yang 0080, Muqiao Yang, Xiang Li 0106, Zhiyuan Zhao 0002, Bhiksha Raj, Rita Singh |
Comput. Speech Lang. | 6 |
| 2023 | Panoramic Video Salient Object Detection with Ambisonic Audio GuidanceabstractVideo salient object detection (VSOD), as a fundamental computer vision problem, has been extensively discussed in the last decade. However, all existing works focus on addressing the VSOD problem in 2D scenarios. With the rapid development of VR devices, panoramic videos have been a promising alternative to 2D videos to provide immersive feelings of the real world. In this paper, we aim to tackle the video salient object detection problem for panoramic videos, with their corresponding ambisonic audios. A multimodal fusion module equipped with two pseudo-siamese audio-visual context fusion (ACF) blocks is proposed to effectively conduct audio-visual interaction. The ACF block equipped with spherical positional encoding enables the fusion in the 3D context to capture the spatial correspondence between pixels and sound sources from the equirectangular frames and ambisonic audios. Experimental results verify the effectiveness of our proposed components and demonstrate that our method achieves state-of-the-art performance on the ASOD60K dataset. Xiang Li 0003, Haoyuan Cao, Shijie Zhao 0001, Li Zhang 0006, Bhiksha Raj |
AAAI | 6 |
| 2023 | VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph CaptioningabstractVideo Paragraph Captioning aims to generate a multi-sentence description of an untrimmed video with multiple temporal event locations in a coherent storytelling. Following the human perception process, where the scene is effectively understood by decomposing it into visual (e.g. human, animal) and non-visual components (e.g. action, relations) under the mutual influence of vision and language, we first propose a visual-linguistic (VL) feature. In the proposed VL feature, the scene is modeled by three modalities including (i) a global visual environment; (ii) local visual main agents; (iii) linguistic scene elements. We then introduce an autoregressive Transformer-in-Transformer (TinT) to simultaneously capture the semantic coherence of intra- and inter-event contents within a video. Finally, we present a new VL contrastive loss function to guarantee the learnt embedding features are consistent with the captions semantics. Comprehensive experiments and extensive ablation studies on the ActivityNet Captions and YouCookII datasets show that the proposed Visual-Linguistic Transformer-in-Transform (VLTinT) outperforms previous state-of-the-art methods in terms of accuracy and diversity. The source code is made publicly available at: https://github.com/UARK-AICV/VLTinT. Kashu Yamazaki, Viet-Khoa Vo-Ho, Quang Sang Truong, Bhiksha Raj, T. Hoang Ngan Le |
AAAI | 4 |
| 2023 | Espnet-Summ: Introducing a Novel Large Dataset, Toolkit, and a Cross-Corpora Evaluation of Speech Summarization SystemsabstractSpeech summarization has garnered significant interest and progressed rapidly over the past few years. In particular, end-to-end models have recently emerged as a competitive alternative to cascade systems for abstractive video summarization. This paper aims to establish progress in this rapidly evolving research field, by introducing ESPNet-SUMM, a new open-source toolkit that facilitates a comprehensive comparison of end-to-end and cascade speech summarization models on 4 different speech summarization tasks spanning diverse applications. Experiments demonstrate that end-to-end models perform better for larger corpora with shorter inputs. This work also introduces Interview, the largest public open-domain multiparty interview corpus with $4400 \mathrm{~h}$ of conversations between radio hosts and guests. Finally, this work explores the use of multiple datasets to improve end-to-end summarization, and experiments demonstrate the benefit of multi-style training over fine-tuning. 1 Roshan S. Sharma, Takatomo Kano, Ruchira Sharma, Siddhant Arora, Shinji Watanabe 0001, Atsunori Ogawa, Marc Delcroix, Rita Singh, Bhiksha Raj |
ASRU | 10 |
| 2023 | FREDOM: Fairness Domain Adaptation Approach to Semantic Scene UnderstandingabstractAlthough Domain Adaptation in Semantic Scene Segmentation has shown impressive improvement in recent years, the fairness concerns in the domain adaptation have yet to be well defined and addressed. In addition, fairness is one of the most critical aspects when deploying the segmentation models into human-related real-world applications, e.g., autonomous driving, as any unfair predictions could influence human safety. In this paper, we propose a novel Fairness Domain Adaptation (FREDOM) approach to semantic scene segmentation. In particular, from the proposed formulated fairness objective, a new adaptation framework will be introduced based on the fair treatment of class distributions. Moreover, to generally model the context of structural dependency, a new conditional structural constraint is introduced to impose the consistency of predicted segmentation. Thanks to the proposed Conditional Structure Network, the self-attention mechanism has sufficiently modeled the structural information of segmentation. Through the ablation studies, the proposed method has shown the performance improvement of the segmentation models and promoted fairness in the model predictions. The experimental results on the two standard benchmarks, i.e., SYNTHIA$\rightarrow$Cityscapes and GTA5$\rightarrow$Cityscapes, have shown that our method achieved State-of-the-Art (SOTA) performance11The implementation of FREDOM is available at https://github.com/uark-cviu/FREDOM Thanh-Dat Truong, T. Hoang Ngan Le, Bhiksha Raj, Jackson David Cothren, Khoa Luu |
CVPR | 3 |
| 2023 | Token Prediction as Implicit Classification to Identify LLM-Generated TextabstractThis paper introduces a novel approach for identifying the possible large language models (LLMs) involved in text generation.Instead of adding an additional classification layer to a base LM, we reframe the classification task as a next-token prediction task and directly fine-tune the base LM to perform it.We utilize the Text-to-Text Transfer Transformer (T5) model as the backbone for our experiments.We compared our approach to the more direct approach of utilizing hidden states for classification.Evaluation shows the exceptional performance of our method in the text classification task, highlighting its simplicity and efficiency.Furthermore, interpretability studies on the features extracted by our model reveal its ability to differentiate distinctive writing styles among various LLMs even in the absence of an explicit classifier.We also collected a dataset named OpenLLMText, containing approximately 340k text samples from human and LLMs, including GPT3. Hao Kang, Vivian Zhai, Liangze Li, Rita Singh, Bhiksha Raj |
EMNLP | 6 |
| 2023 | Towards Noise-Tolerant Speech-Referring Video Object Segmentation: Bridging Speech and TextabstractLinguistic communication is prevalent in Human-Computer Interaction (HCI).Speech (spoken language) serves as a convenient yet potentially ambiguous form due to noise and accents, exposing a gap compared to text.In this study, we investigate the prominent HCI task, Referring Video Object Segmentation (R-VOS), which aims to segment and track objects using linguistic references.While text input is well-investigated, speech input is underexplored.Our objective is to bridge the gap between speech and text, enabling the adaptation of existing text-input R-VOS models to accommodate noisy speech input effectively.Specifically, we propose a method to align the semantic spaces between speech and text by incorporating two key modules: 1) Noise-Aware Semantic Adjustment (NSA) for clear semantics extraction from noisy speech; and 2) Semantic Jitter Suppression (SJS) enabling R-VOS models to tolerate noisy queries.Comprehensive experiments conducted on the challenging AVOS benchmarks reveal that our proposed method outperforms state-of-the-art approaches. Xiang Li 0106, Jinglu Wang, Xiaohao Xu, Muqiao Yang, Fan Yang 0080, Rita Singh, Bhiksha Raj |
EMNLP | 8 |
| 2023 | An Approach to Ontological Learning from Weak LabelsabstractOntologies encompass a formal representation of knowledge through the definition of concepts or properties of a domain, and the relationships between those concepts. In this work, we seek to investigate whether using this ontological information will improve learning from weakly labeled data, which are easier to collect since it requires only the presence or absence of an event to be known. We use the AudioSet ontology and dataset, which contains audio clips weakly labeled with the ontology concepts and the ontology providing the "Is A" relations between the concepts. We first re-implemented the model proposed by [1] with modifications to fit the multi-label scenario and then expand on that idea by using a Graph Convolutional Network (GCN) to model the ontology information to learn the concepts. We find that the baseline Twin Neural Network (TNN) does not perform better by incorporating ontology information in the weak and multi-label scenario, but that the GCN does capture the ontology knowledge better for weak, multi-labeled data. We also investigate how different modules can tolerate noises introduced from weak labels and better incorporate ontology information. Our best TNN-GCN model achieves mAP=0.45 and AUC=0.87 for lower-level concepts and mAP=0.72 and AUC=0.86 for higher-level concepts, which is an improvement over the baseline TNN but about the same as our models that do not use ontology information. Ankit Shah 0001, Larry Tang 0003, Po Hao Chou, Yi Yu Zheng, Ziqian Ge, Bhiksha Raj |
ICASSP | 6 |
| 2023 | Privacy-Preserving Automatic Speaker DiarizationabstractAutomatic Speaker Diarization (ASD) is an enabling technology with numerous applications, which deals with recordings of multiple speakers, raising special concerns in terms of privacy. In fact, in remote settings, where recordings are shared with a server, clients relinquish not only the privacy of their conversation, but also of all the information that can be inferred from their voices. However, to the best of our knowledge, the development of privacy-preserving ASD systems has been overlooked thus far. In this work, we tackle this problem using a combination of two cryptographic techniques, Secure Multiparty Computation (SMC) and Secure Modular Hashing, and apply them to the two main steps of a cascaded ASD system: speaker embedding extraction and agglomerative hierarchical clustering. Our system is able to achieve a reasonable trade-off between performance and efficiency, presenting real-time factors of 1.1 and 1.6, for two different SMC security settings. Francisco Teixeira, Alberto Abad, Bhiksha Raj, Isabel Trancoso |
ICASSP | 3 |
| 2023 | Paaploss: A Phonetic-Aligned Acoustic Parameter Loss for Speech EnhancementabstractDespite rapid advancement in recent years, current speech enhancement models often produce speech that differs in perceptual quality from real clean speech. We propose a learning objective that formalizes differences in perceptual quality, by using domain knowledge of acoustic-phonetics. We identify temporal acoustic parameters – such as spectral tilt, spectral flux, shimmer, etc. – that are non-differentiable, and we develop a neural network estimator that can accurately predict their time-series values across an utterance. We also model phoneme-specific weights for each feature, as the acoustic parameters are known to show different behavior in different phonemes. We can add this criterion as an auxiliary loss to any model that produces speech, to optimize speech outputs to match the values of clean speech in these features. Experimentally we show that it improves speech enhancement workflows in both time-domain and time-frequency domain, as measured by standard evaluation metrics. We also provide an analysis of phoneme-dependent improvement on acoustic parameters, demonstrating the additional interpretability that our method provides. This analysis can suggest which features are currently the bottleneck for improvement. Muqiao Yang, Joseph Konan, David Bick, Yunyang Zeng, Anurag Kumar 0003, Shinji Watanabe 0001, Bhiksha Raj |
ICASSP | 8 |
| 2023 | TAPLoss: A Temporal Acoustic Parameter Loss for Speech EnhancementabstractSpeech enhancement models have greatly progressed in recent years, but still show limits in perceptual quality of their speech outputs. We propose an objective for perceptual quality based on temporal acoustic parameters. These are fundamental speech features that play an essential role in various applications, including speaker recognition and paralinguistic analysis. We provide a differentiable estimator for four categories of low-level acoustic descriptors involving: frequency-related parameters, energy or amplitude-related parameters, spectral balance parameters, and temporal features. Un-like prior work that looks at aggregated acoustic parameters or a few categories of acoustic parameters, our temporal acoustic parameter (TAP) loss enables auxiliary optimization and improvement of many fine-grained speech characteristics in enhancement workflows. We show that adding TAPLoss as an auxiliary objective in speech enhancement produces speech with improved perceptual quality and intelligibility. We use data from the Deep Noise Suppression 2020 Challenge to demonstrate that both time-domain models and time-frequency domain models can benefit from our method. Yunyang Zeng, Joseph Konan, David Bick, Muqiao Yang, Anurag Kumar 0003, Shinji Watanabe 0001, Bhiksha Raj |
ICASSP | 8 |
| 2023 | Robust Referring Video Object Segmentation with Cyclic Structural ConsensusabstractReferring Video Object Segmentation (R-VOS) is a challenging task that aims to segment an object in a video based on a linguistic expression. Most existing R-VOS methods have a critical assumption: the object referred to must appear in the video. This assumption, which we refer to as "semantic consensus", is often violated in real-world scenarios, where the expression may be queried against false videos. In this work, we highlight the need for a robust R-VOS model that can handle semantic mismatches. Accordingly, we propose an extended task called Robust R-VOS (R2-VOS), which accepts unpaired video-text inputs. We tackle this problem by jointly modeling the primary R-VOS problem and its dual (text reconstruction). A structural text-to-text cycle constraint is introduced to discriminate semantic consensus between video-text pairs and impose it in positive pairs, thereby achieving multi-modal alignment from both positive and negative pairs. Our structural constraint effectively addresses the challenge posed by linguistic diversity, overcoming the limitations of previous methods that relied on the point-wise constraint. A new evaluation dataset, R2-Youtube-VOS is constructed to measure the model robustness. Our model achieves state-of-the-art performance on R-VOS benchmarks, Ref-DAVIS17 and Ref-Youtube-VOS, and also our R2-Youtube-VOS dataset. Xiang Li 0106, Jinglu Wang, Xiaohao Xu, Xiao Li 0030, Bhiksha Raj, Yan Lu 0001 |
ICCV | 5 |
| 2023 | Pairwise Similarity Learning is SimPLEabstractIn this paper, we focus on a general yet important learning problem, pairwise similarity learning (PSL). PSL subsumes a wide range of important applications, such as open-set face recognition, speaker verification, image retrieval and person re-identification. The goal of PSL is to learn a pairwise similarity function assigning a higher similarity score to positive pairs (i.e., a pair of samples with the same label) than to negative pairs (i.e., a pair of samples with different label). We start by identifying a key desideratum for PSL, and then discuss how existing methods can achieve this desideratum. We then propose a surprisingly simple proxy-free method, called SimPLE, which requires neither feature/proxy normalization nor angular margin and yet is able to generalize well in open-set recognition. We apply the proposed method to three challenging PSL tasks: open-set face recognition, image retrieval and speaker verification. Comprehensive experimental results on large-scale benchmarks show that our method performs significantly better than current state-of-the-art methods. Our project page is available at simple.is.tue.mpg.de. Yandong Wen, Weiyang Liu, Yao Feng 0001, Bhiksha Raj, Rita Singh, Adrian Weller, Michael J. Black, Bernhard Schölkopf |
ICCV | 4 |
| 2023 | SoftMatch: Addressing the Quantity-Quality Tradeoff in Semi-supervised Learning
Hao Chen 0102, Ran Tao 0013, Yidong Wang 0003, Jindong Wang 0001, Bernt Schiele, Xing Xie 0001, Bhiksha Raj, Marios Savvides |
ICLR | 8 |
| 2023 | FreeMatch: Self-adaptive Thresholding for Semi-supervised Learning
Yidong Wang 0003, Hao Chen 0102, Qiang Heng, Wenxin Hou, Zhen Wu 0002, Jindong Wang 0001, Marios Savvides, Takahiro Shinozaki, Bhiksha Raj, Bernt Schiele, Xing Xie 0001 |
ICLR | 10 |
| 2023 | How Many Perturbations Break This Model? Evaluating Robustness Beyond Adversarial AccuracyabstractRobustness to adversarial attacks is typically evaluated with adversarial accuracy. While essential, this metric does not capture all aspects of robustness and in particular leaves out the question of how many perturbations can be found for each point. In this work, we introduce an alternative approach, adversarial sparsity, which quantifies how difficult it is to find a successful perturbation given both an input point and a constraint on the direction of the perturbation. We show that sparsity provides valuable insight into neural networks in multiple ways: for instance, it illustrates important differences between current state-of-the-art robust models them that accuracy analysis does not, and suggests approaches for improving their robustness. When applying broken defenses effective against weak attacks but not strong ones, sparsity can discriminate between the totally ineffective and the partially effective defenses. Finally, with sparsity we can measure increases in robustness that do not affect accuracy: we show for example that data augmentation can by itself increase adversarial robustness, without using adversarial training. Raphaël Olivier, Bhiksha Raj |
ICML | 2 |
| 2023 | BASS: Block-wise Adaptation for Speech Summarization
Roshan S. Sharma, Siddhant Arora, Kenneth Zheng, Shinji Watanabe 0001, Rita Singh, Bhiksha Raj |
INTERSPEECH | 6 |
| 2023 | There is more than one kind of robustness: Fooling Whisper with adversarial examples
Raphaël Olivier, Bhiksha Raj |
INTERSPEECH | 2 |
| 2023 | The Hidden Dance of Phonemes and Visage: Unveiling the Enigmatic Link between Phonemes and Facial Features
Liao Qu, Xianwei Zou, Xiang Li 0106, Yandong Wen, Rita Singh, Bhiksha Raj |
INTERSPEECH | 6 |
| 2023 | Rethinking Voice-Face Correlation: A Geometry ViewabstractPrevious works on voice-face matching and voice-guided face synthesis demonstrate strong correlations between voice and face, but mainly rely on coarse semantic cues such as gender, age, and emotion. In this paper, we aim to investigate the capability of reconstructing the 3D facial shape from voice from a geometry perspective without any semantic information. We propose a voice-anthropometric measurement (AM)-face paradigm, which identifies predictable facial AMs from the voice and uses them to guide 3D face reconstruction. By leveraging AMs as a proxy to link the voice and face geometry, we can eliminate the influence of unpredictable AMs and make the face geometry tractable. Our approach is evaluated on our proposed dataset with ground-truth 3D face scans and corresponding voice recordings, and we find significant correlations between voice and specific parts of the face geometry, such as the nasal cavity and cranium. Our work offers a new perspective on voice-face correlation and can serve as a good empirical study for anthropometry science. Xiang Li 0106, Yandong Wen, Muqiao Yang, Jinglu Wang, Rita Singh, Bhiksha Raj |
ACM Multimedia | 6 |
| 2023 | PaintSeg: Painting Pixels for Training-free SegmentationabstractThe paper introduces PaintSeg, a new unsupervised method for segmenting objects without any training. We propose an adversarial masked contrastive painting (AMCP) process, which creates a contrast between the original image and a painted image in which a masked area is painted using off-the-shelf generative models. During the painting process, inpainting and outpainting are alternated, with the former masking the foreground and filling in the background, and the latter masking the background while recovering the missing part of the foreground object. Inpainting and outpainting, also referred to as I-step and O-step, allow our method to gradually advance the target segmentation mask toward the ground truth without supervision or training. PaintSeg can be configured to work with a variety of prompts, e.g. coarse masks, boxes, scribbles, and points. Our experimental results demonstrate that PaintSeg outperforms existing approaches in coarse mask-prompt, box-prompt, and point-prompt segmentation tasks, providing a training-free solution suitable for unsupervised segmentation. Code: https://github.com/lxa9867/PaintSeg. Xiang Li 0106, Chung-Ching Lin, Yinpeng Chen, Zicheng Liu 0001, Jinglu Wang, Rita Singh, Bhiksha Raj |
NeurIPS | 7 |
| 2023 | Weakly-Supervised Audio-Visual SegmentationabstractAudio-visual segmentation is a challenging task that aims to predict pixel-level masks for sound sources in a video.
Previous work applied a comprehensive manually designed architecture with countless pixel-wise accurate masks as supervision. However, these pixel-level masks are expensive and not available in all cases.
In this work, we aim to simplify the supervision as the instance-level annotation, $\textit{i.e.}$, weakly-supervised audio-visual segmentation.
We present a novel Weakly-Supervised Audio-Visual Segmentation framework, namely WS-AVS, that can learn multi-scale audio-visual alignment with multi-scale multiple-instance contrastive learning for audio-visual segmentation.
Extensive experiments on AVSBench demonstrate the effectiveness of our WS-AVS in the weakly-supervised audio-visual segmentation of single-source and multi-source scenarios. Shentong Mo, Bhiksha Raj |
NeurIPS | 2 |
| 2023 | Training on Foveated Images Improves Robustness to Adversarial AttacksabstractDeep neural networks (DNNs) have been shown to be vulnerable to adversarial attacks
-- subtle, perceptually indistinguishable perturbations of inputs that change the response of the model. In the context of vision, we hypothesize that an important contributor to the robustness of human visual perception is constant exposure to low-fidelity visual stimuli in our peripheral vision. To investigate this hypothesis, we develop RBlur, an image transform that simulates the loss in fidelity of peripheral vision by blurring the image and reducing its color saturation based on the distance from a given fixation point. We show that compared to DNNs trained on the original images, DNNs trained on images transformed by RBlur are substantially more robust to adversarial attacks, as well as other, non-adversarial, corruptions, achieving up to 25% higher accuracy on perturbed data. Muhammad A. Shah, Aqsa Kashaf, Bhiksha Raj |
NeurIPS | 3 |
| 2023 | Fairness Continual Learning Approach to Semantic Scene Understanding in Open-World EnvironmentsabstractContinual semantic segmentation aims to learn new classes while maintaining the information from the previous classes. Although prior studies have shown impressive progress in recent years, the fairness concern in the continual semantic segmentation needs to be better addressed. Meanwhile, fairness is one of the most vital factors in deploying the deep learning model, especially in human-related or safety applications. In this paper, we present a novel Fairness Continual Learning approach to the semantic segmentation problem.
In particular, under the fairness objective, a new fairness continual learning framework is proposed based on class distributions.
Then, a novel Prototypical Contrastive Clustering loss is proposed to address the significant challenges in continual learning, i.e., catastrophic forgetting and background shift. Our proposed loss has also been proven as a novel, generalized learning paradigm of knowledge distillation commonly used in continual learning. Moreover, the proposed Conditional Structural Consistency loss further regularized the structural constraint of the predicted segmentation. Our proposed approach has achieved State-of-the-Art performance on three standard scene understanding benchmarks, i.e., ADE20K, Cityscapes, and Pascal VOC, and promoted the fairness of the segmentation model. Thanh-Dat Truong, Hoang-Quan Nguyen, Bhiksha Raj, Khoa Luu |
NeurIPS | 3 |
| 2023 | AOE-Net: Entities Interactions Modeling with Adaptive Attention Mechanism for Temporal Action Proposals Generation
Viet-Khoa Vo-Ho, Sang Truong, Kashu Yamazaki, Bhiksha Raj, Minh-Triet Tran, T. Hoang Ngan Le |
Int. J. Comput. Vis. | 4 |
| 2023 | SphereFace Revived: Unifying Hyperspherical Face RecognitionabstractThis paper addresses the deep face recognition problem under an open-set protocol, where ideal face features are expected to have smaller maximal intra-class distance than minimal inter-class distance under a suitably chosen metric space. To this end, hyperspherical face recognition, as a promising line of research, has attracted increasing attention and gradually become a major focus in face recognition research. As one of the earliest works in hyperspherical face recognition, SphereFace explicitly proposed to learn face embeddings with large inter-class angular margin. However, SphereFace still suffers from severe training instability which limits its application in practice. In order to address this problem, we introduce a unified framework to understand large angular margin in hyperspherical face recognition. Under this framework, we extend the study of SphereFace and propose an improved variant with substantially better training stability - SphereFace-R. Specifically, we propose two novel ways to implement the multiplicative margin, and study SphereFace-R under three different feature normalization schemes (no feature normalization, hard feature normalization and soft feature normalization). We also propose an implementation strategy - "characteristic gradient detachment" - to stabilize training. Extensive experiments on SphereFace-R show that it is consistently better than or competitive with state-of-the-art methods. Weiyang Liu, Yandong Wen, Bhiksha Raj, Rita Singh, Adrian Weller |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | SphereFace2: Binary Classification is All You Need for Deep Face Recognition
Yandong Wen, Weiyang Liu, Adrian Weller, Bhiksha Raj, Rita Singh |
ICLR | 4 |
| 2022 | Positional Encoding for Capturing Modality Specific Cadence for Emotion Detection
Hira Dhamyal, Bhiksha Raj, Rita Singh |
INTERSPEECH | 2 |
| 2022 | Recent improvements of ASR models in the face of adversarial attacksabstractLike many other tasks involving neural networks, Speech Recognition models are vulnerable to adversarial attacks.However recent research has pointed out differences between attacks and defenses on ASR models compared to image models.Improving the robustness of ASR models requires a paradigm shift from evaluating attacks on one or a few models to a systemic approach in evaluation.We lay the ground for such research by evaluating on various architectures a representative set of adversarial attacks: targeted and untargeted, optimization and speech processing-based, white-box, black-box and targeted attacks.Our results show that the relative strengths of different attack algorithms vary considerably when changing the model architecture, and that the results of some attacks are not to be blindly trusted.They also indicate that training choices such as self-supervised pretraining may significantly impact robustness by enabling transferable perturbations.We release our source code as a package that should help future research in evaluating their attacks and defenses. Raphaël Olivier, Bhiksha Raj |
INTERSPEECH | 2 |
| 2022 | Towards End-to-End Private Automatic Speaker RecognitionabstractThe development of privacy-preserving automatic speaker verification systems has been the focus of a number of studies with the intent of allowing users to authenticate themselves without risking the privacy of their voice. However, current privacy-preserving methods assume that the template voice representations (or speaker embeddings) used for authentication are extracted locally by the user. This poses two important issues: first, knowledge of the speaker embedding extraction model may create security and robustness liabilities for the authentication system, as this knowledge might help attackers in crafting adversarial examples able to mislead the system; second, from the point of view of a service provider the speaker embedding extraction model is arguably one of the most valuable components in the system and, as such, disclosing it would be highly undesirable. In this work, we show how speaker embeddings can be extracted while keeping both the speaker's voice and the service provider's model private, using Secure Multiparty Computation. Further, we show that it is possible to obtain reasonable trade-offs between security and computational cost. This work is complementary to those showing how authentication may be performed privately, and thus can be considered as another step towards fully private automatic speaker recognition. Francisco Teixeira, Alberto Abad, Bhiksha Raj, Isabel Trancoso |
INTERSPEECH | 3 |
| 2022 | Improving Speech Enhancement through Fine-Grained Speech CharacteristicsabstractWhile deep learning based speech enhancement systems have made rapid progress in improving the quality of speech signals, they can still produce outputs that contain artifacts and can sound unnatural.We propose a novel approach to speech enhancement aimed at improving perceptual quality and naturalness of enhanced signals by optimizing for key characteristics of speech.We first identify key acoustic parameters that have been found to correlate well with voice quality (e.g.jitter, shimmer, and spectral flux) and then propose objective functions which are aimed at reducing the difference between clean speech and enhanced speech with respect to these features.The full set of acoustic features is the extended Geneva Acoustic Parameter Set (eGeMAPS), which includes 25 different attributes associated with perception of speech.Given the non-differentiable nature of these feature computation, we first build differentiable estimators of the eGeMAPS and then use them to fine-tune existing speech enhancement systems.Our approach is generic and can be applied to any existing deep learning based enhancement systems to further improve the enhanced speech signals.Experimental results conducted on the Deep Noise Suppression (DNS) Challenge dataset shows that our approach can improve the state-of-the-art deep learning based enhancement systems. Muqiao Yang, Joseph Konan, David Bick, Anurag Kumar 0003, Shinji Watanabe 0001, Bhiksha Raj |
INTERSPEECH | 6 |
| 2022 | USB: A Unified Semi-supervised Learning Benchmark for ClassificationabstractSemi-supervised learning (SSL) improves model generalization by leveraging massive unlabeled data to augment limited labeled samples. However, currently, popular SSL evaluation protocols are often constrained to computer vision (CV) tasks. In addition, previous work typically trains deep neural networks from scratch, which is time-consuming and environmentally unfriendly. To address the above issues, we construct a Unified SSL Benchmark (USB) for classification by selecting 15 diverse, challenging, and comprehensive tasks from CV, natural language processing (NLP), and audio processing (Audio), on which we systematically evaluate the dominant SSL methods, and also open-source a modular and extensible codebase for fair evaluation of these SSL methods. We further provide the pre-trained versions of the state-of-the-art neural models for CV tasks to make the cost affordable for further tuning. USB enables the evaluation of a single SSL algorithm on more tasks from multiple domains but with less cost. Specifically, on a single NVIDIA V100, only 39 GPU days are required to evaluate FixMatch on 15 tasks in USB while 335 GPU days (279 GPU days on 4 CV datasets except for ImageNet) are needed on 5 CV tasks with TorchSSL. Yidong Wang 0003, Hao Chen 0102, Wang Sun, Ran Tao 0013, Wenxin Hou, Linyi Yang, Zhi Zhou 0007, Lan-Zhe Guo, Heli Qi, Zhen Wu 0002, Yufeng Li 0008, Satoshi Nakamura 0001, Wei Ye 0004, Marios Savvides, Bhiksha Raj, Takahiro Shinozaki, Bernt Schiele, Jindong Wang 0001, Xing Xie 0001, Yue Zhang 0004 |
NeurIPS | 17 |
| 2021 | Point3D: tracking actions as moving points with 3D CNNs
Shentong Mo, Jingfei Xia, Xiaoqing Tan, Bhiksha Raj |
BMVC | 4 |
| 2021 | Sequential Randomized Smoothing for Adversarially Robust Speech RecognitionabstractWhile Automatic Speech Recognition has been shown to be vulnerable to adversarial attacks, defenses against these attacks are still lagging.Existing, naive defenses can be partially broken with an adaptive attack.In classification tasks, the Randomized Smoothing paradigm has been shown to be effective at defending models.However, it is difficult to apply this paradigm to ASR tasks, due to their complexity and the sequential nature of their outputs.Our paper overcomes some of these challenges by leveraging speech-specific tools like enhancement and ROVER voting to design an ASR model that is robust to perturbations.We apply adaptive versions of stateof-the-art attacks, such as the Imperceptible ASR attack, to our model, and show that our strongest defense is robust to all attacks that use inaudible noise, and can only be broken with very high distortion. Raphaël Olivier, Bhiksha Raj |
EMNLP (1) | 2 |
| 2021 | The in-the-Wild Speech Medical CorpusabstractAutomatic detection of speech affecting (SA) diseases has received significant attention, particularly in clinical scenarios. However, the same task in in-the-wild conditions is often neglected, in part, due to the lack of appropriate datasets.In this work, we present the in-the-Wild Speech Medical (WSM) Corpus, a collection of in-the-wild videos, featuring subjects potentially affected by a SA disease - specifically, depression or Parkinson’s disease. The WSM Corpus contains a total 928 videos, and over 131 hours of speech. Each video is accompanied by a crowdsourced annotation for perceived age/gender, and self-reported health status of the speaker. The WSM Corpus is balanced over all the labels.In this work we present a detailed description of the collection, and annotation processes of the WSM corpus. Furthermore, we present present several baseline systems for the detection of SA diseases using speech alone, thus motivating the use of this type of in-the-wild data in paralinguistic audiovisual tasks. Maria Joana Correia, Francisco Teixeira, Catarina Botelho, Isabel Trancoso, Bhiksha Raj |
ICASSP | 5 |
| 2021 | High-Frequency Adversarial Defense for Speech and AudioabstractRecent work suggests that adversarial examples are enabled by high-frequency components in the dataset. In the speech domain where spectrograms are used extensively, masking those components seems like a sound direction for defenses against attacks. We explore a smoothing approach based on additive noise masking in priority high frequencies. We show that this approach is much more robust than the naive noise filtering approach, and a promising research direction. We successfully apply our defense on a Librispeech speaker identification task, and on the UrbanSound8K audio classification dataset. Raphaël Olivier, Bhiksha Raj, Muhammad A. Shah |
ICASSP | 2 |
| 2021 | Towards Adversarial Robustness Via Compact Feature RepresentationsabstractDeep Neural Networks (DNNs), while providing state-of-the-art performance in a wide variety of tasks, have been shown to be vulnerable to adversarial attacks. Recent studies have posited that this vulnerability arises because DNNs operate over a grossly overspecified input space with very sparse human supervision due to which they tend to learn spurious features that humans would ignore. These spurious features provide an attack vector for the adversary because perturbing these features would not alter the human’s decision but may alter the model’s prediction. In this paper we explore hypothesis that reducing the size of the model’s feature representation while maintaining its generalizability would discard spurious features while retaining perceptually relevant ones. We find that after the size of the feature representation has been reduced the models exhibit increased adversarial robustness, while suffering only a minimal loss in accuracy. In addition to being more robust, models with compact feature representations have the benefit of being more resource efficient. Muhammad A. Shah, Raphaël Olivier, Bhiksha Raj |
ICASSP | 3 |
| 2021 | FoolHD: Fooling Speaker Identification by Highly Imperceptible Adversarial DisturbancesabstractSpeaker identification models are vulnerable to carefully designed adversarial perturbations of their input signals that induce misclassification. In this work, we propose a white-box steganography-inspired adversarial attack that generates imperceptible adversarial perturbations against a speaker identification model. Our approach, FoolHD, uses a Gated Convolutional Autoencoder that operates in the DCT domain and is trained with a multi-objective loss function, to generate and conceal the adversarial perturbation within the original audio files. In addition to hindering speaker identification performance, this multi-objective loss accounts for human perception through a frame-wise cosine similarity between MFCC feature vectors extracted from the original and adversarial audio files. We validate the effectiveness of FoolHD with a 250-speaker identification x-vector network, trained using VoxCeleb, in terms of accuracy, success rate, and imperceptibility. Our results show that FoolHD generates highly imperceptible adversarial audio files (average PESQ scores above 4.30), while achieving a success rate of 99.6% and 99.2% in misleading the speaker identification model, for untargeted and targeted settings, respectively. Ali Shahin Shamsabadi, Francisco Teixeira, Alberto Abad, Bhiksha Raj, Andrea Cavallaro, Isabel Trancoso |
ICASSP | 4 |
| 2021 | Contrast and Order Representations for Video Self-supervised LearningabstractThis paper studies the problem of learning self-supervised representations on videos. In contrast to image modality that only requires appearance information on objects or scenes, video needs to further explore the relations between multiple frames/clips along the temporal dimension. However, the recent proposed contrastive-based self-supervised frameworks do not grasp such relations explicitly since they simply utilize two augmented clips from the same video and compare their distance without referring to their temporal relation. To address this, we present a contrast-and-order representation (CORP) framework for learning self-supervised video representations that can automatically capture both the appearance information within each frame and temporal information across different frames. In particular, given two video clips, our model first predicts whether they come from the same input video, and then predict the temporal ordering of the clips if they come from the same video. We also propose a novel decoupling attention method to learn symmetric similarity (contrast) and anti-symmetric patterns (order). Such design involves neither extra parameters nor computation, but can speed up the learning process and improve accuracy compared to the vanilla multi-head attention. We extensively validate the representation ability of our learned video features for the downstream action recognition task on Kinetics-400 and Something-something V2. Our method outperforms previous state-of-the-arts by a significant margin. Kai Hu 0010, Jie Shao 0006, Bhiksha Raj, Marios Savvides |
ICCV | 4 |
| 2021 | The Right to Talk: An Audio-Visual Transformer ApproachabstractTurn-taking has played an essential role in structuring the regulation of a conversation. The task of identifying the main speaker (who is properly taking his/her turn of speaking) and the interrupters (who are interrupting or reacting to the main speaker’s utterances) remains a challenging task. Although some prior methods have partially addressed this task, there still remain some limitations. Firstly, a direct association of Audio and Visual features may limit the correlations to be extracted due to different modalities. Secondly, the relationship across temporal segments helping to maintain the consistency of localization, separation and conversation contexts is not effectively exploited. Finally, the interactions between speakers that usually contain the tracking and anticipatory decisions about transition to a new speaker is usually ignored. Therefore, this work introduces a new Audio-Visual Transformer approach to the problem of localization and highlighting the main speaker in both audio and visual channels of a multi-speaker conversation video in the wild. The proposed method exploits different types of correlations presented in both visual and audio signals. The temporal audio-visual relationships across spatial-temporal space are anticipated and optimized via the self-attention mechanism in a Transformer structure. Moreover, a newly collected dataset is introduced for the main speaker detection. To the best of our knowledge, it is one of the first studies that is able to automatically localize and highlight the main speaker in both visual and audio channels in multi-speaker conversation videos. Thanh-Dat Truong, Chi Nhan Duong, The De Vu, Bhiksha Raj, T. Hoang Ngan Le, Khoa Luu |
ICCV | 5 |
| 2021 | Self-Supervised 3D Face Reconstruction via Conditional EstimationabstractWe present a conditional estimation (CEST) framework to learn 3D facial parameters from 2D single-view images by self-supervised training from videos. CEST is based on the process of analysis by synthesis, where the 3D facial parameters (shape, reflectance, viewpoint, and illumination) are estimated from the face image, and then recombined to reconstruct the 2D face image. In order to learn semantically meaningful 3D facial parameters without explicit access to their labels, CEST couples the estimation of different 3D facial parameters by taking their statistical dependency into account. Specifically, the estimation of any 3D facial parameter is not only conditioned on the given image, but also on the facial parameters that have already been derived. Moreover, the reflectance symmetry and consistency among the video frames are adopted to improve the disentanglement of facial parameters. Together with a novel strategy for incorporating the reflectance symmetry and consistency, CEST can be efficiently trained with in-the-wild video clips. Both qualitative and quantitative experiments demonstrate the effectiveness of CEST. Yandong Wen, Weiyang Liu, Bhiksha Raj, Rita Singh |
ICCV | 3 |
| 2021 | Improving Weakly Supervised Sound Event Detection with Self-Supervised Auxiliary TasksabstractWhile multitask and transfer learning has shown to improve the performance of neural networks in limited data settings, they require pretraining of the model on large datasets beforehand. In this paper, we focus on improving the performance of weakly supervised sound event detection in low data and noisy settings simultaneously without requiring any pretraining task. To that extent, we propose a shared encoder architecture with sound event detection as a primary task and an additional secondary decoder for a self-supervised auxiliary task. We empirically evaluate the proposed framework for weakly supervised sound event detection on a remix dataset of the DCASE 2019 task 1 acoustic scene data with DCASE 2018 Task 2 sounds event data under 0, 10 and 20 dB SNR. To ensure we retain the localisation information of multiple sound events, we propose a two-step attention pooling mechanism that provides a time-frequency localisation of multiple audio events in the clip. The proposed framework with two-step attention outperforms existing benchmark models by 22.3%, 12.8%, 5.9% on 0, 10 and 20 dB SNR respectively. We carry out an ablation study to determine the contribution of the auxiliary task and two-step attention pooling to the SED performance improvement. Soham Deshmukh, Bhiksha Raj, Rita Singh |
Interspeech | 2 |
| 2021 | Masked Proxy Loss for Text-Independent Speaker VerificationabstractOpen-set speaker recognition can be regarded as a metric learning problem, which is to maximize inter-class variance and minimize intra-class variance. Supervised metric learning can be categorized into entity-based learning and proxy-based learning. Most of the existing metric learning objectives like Contrastive, Triplet, Prototypical, GE2E, etc all belong to the former division, the performance of which is either highly dependent on sample mining strategy or restricted by insufficient label information in the mini-batch. Proxy-based losses mitigate both shortcomings, however, fine-grained connections among entities are either not or indirectly leveraged. This paper proposes a Masked Proxy (MP) loss which directly incorporates both proxy-based relationships and pair-based relationships. We further propose Multinomial Masked Proxy (MMP) loss to leverage the hardness of speaker pairs. These methods have been applied to evaluate on VoxCeleb test set and reach state-of-the-art Equal Error Rate(EER). Jiachen Lian, Aiswarya Vinod Kumar, Hira Dhamyal, Bhiksha Raj, Rita Singh |
Interspeech | 4 |
| 2021 | Detection and Evaluation of Human and Machine Generated Speech in Spoofing Attacks on Automatic Speaker Verification SystemsabstractAutomatic speaker verification (ASV) systems utilize the biometric information in human speech to verify the speaker's identity. The techniques used for performing speaker verification are often vulnerable to malicious attacks that attempt to induce the ASV system to return wrong results, allowing an impostor to bypass the system and gain access. Attackers use a multitude of spoofing techniques for this, such as voice conversion, audio replay, speech synthesis, etc. In recent years, easily available tools to generate deepfaked audio have increased the potential threat to ASV systems. In this paper, we compare the potential of human impersonation (voice disguise) based attacks with attacks based on machinegenerated speech, on black-box and white-box ASV systems. We also study countermeasures by using features that capture the unique aspects of human speech production, under the hypothesis that machines cannot emulate many of the finelevel intricacies of the human speech production mechanism. We show that fundamental frequency sequence-related entropy, spectral envelope, and aperiodic parameters are promising candidates for robust detection of deepfaked speech generated by unknown methods. Yang Gao 0029, Jiachen Lian, Bhiksha Raj, Rita Singh |
SLT | 3 |
| 2020 | Deriving Compact Feature Representations Via Annealed ContractionabstractIt is common practice to use pretrained image recognition models to compute feature representations for visual data. The size of the feature representations can have a noticeable impact on the complexity of the models that use these representations, and by extension on their deployablity and scalability. Therefore it would be beneficial to have compact visual representations that carry as much information as their high-dimensional counterparts. To this end we propose a technique that shrinks a layer by an iterative process in which neurons are removed from the and network is fine tuned. Using this technique we are able to remove 99% of the neurons from the penultimate layer of AlexNet and VGG16, while suffering less than 5% drop in accuracy on CIFAR10, Caltech101 and Caltech256. We also show that our method can reduce the size of AlexNet by 95% while only suffering a 4% reduction in accuracy on Caltech101. Muhammad Ahmed Shah, Bhiksha Raj |
ICASSP | 2 |
| 2020 | Artificial Creative Intelligence: Breaking the Imitation Barrier
Rowland Chen, Roger B. Dannenberg, Bhiksha Raj, Rita Singh |
ICCC | 3 |
| 2020 | Exploiting Non-Linear Redundancy for Neural Model CompressionabstractDeploying deep learning models with millions, even billions, of parameters is challenging given real world memory, power and compute constraints. In an effort to make these models more practical, in this paper, we propose a novel model compression approach that exploits linear dependence between the activations in a layer to eliminate entire structural units (neurons/convolutional filters). Our approach also adjusts the weights of the layer in a manner that is provably lossless while training if the removed neuron was perfectly predictable. We combine this approach with an annealing algorithm that may be applied during training, or even on a trained model, and demonstrate, using popular datasets, that our technique can reduce the parameters of VGG and AlexNet by more than 97% on CIFAR-10, 85% on Caltech-256, and 19% on ImageNet at less than 2% loss in accuracy. Furthermore, we provide theoretical results showing that in overparametrized, locally linear (ReLU) neural networks where redundant features exist, and with correct hyperparameter selection, our method is indeed able to capture and suppress those dependencies. Muhammad Ahmed Shah, Raphaël Olivier, Bhiksha Raj |
ICPR | 3 |
| 2020 | Optimal Strategies For Comparing Covariates To Solve Matching ProblemsabstractMany machine learning tasks can be posed as matching problems in which we are given a “probe” entry that we expect matches some of the entries in our “gallery”. The general solution to these problems is to retrieve matching entries based on statistical dependencies between the probe and the gallery data that are learned using complex models. Often, however, there are other common covariates to the probe and gallery data which might be easily inferred and may explain some of the statistical dependencies between the two. In this paper we present a probabilistic framework to derive optimal matching strategies based only on covariate features for three broad tasks, namely N-way classification, pairwise verification and ranking. We use canonical metrics to determine the maximum performance that can be expected if only covariate features are used and determine the marginal gain of using complex models. We find that covariate matching achieves an EER within 10% of a CNN in the verification task, and an MAP within 22% of the a DNN based model in the ranking task. Muhammad A. Shah, Raphaël Olivier, Bhiksha Raj |
ICPR | 3 |
| 2020 | Hierarchical Routing Mixture of ExpertsabstractIn regression tasks, the data distribution is often too complex to be fitted by a single model. In contrast, partition-based models are developed where data is divided and fitted by local models. These models partition the input space and do not leverage the input-output dependency of multimodal-distributed data, and strong local models are needed to make good predictions. Addressing these problems, we propose a binary tree-structured hierarchical routing mixture of experts (HRME) model that has classifiers as non-leaf node experts and simple regression models as leaf node experts. The classifier nodes jointly soft-partition the input-output space based on the natural separateness of multimodal data. This enables simple leaf experts to be effective for prediction. Further, we develop a probabilistic framework for the HRME model and propose a recursive Expectation-Maximization (EM) based algorithm to learn both the tree structure and the expert models. Experiments on a collection of regression tasks validate our method's effectiveness compared to various other regression models. Wenbo Zhao 0006, Yang Gao 0029, Shahan Ali Memon, Bhiksha Raj, Rita Singh |
ICPR | 4 |
| 2020 | The Phonetic Bases of Vocal Expressed Emotion: Natural versus ActedabstractCan vocal emotions be emulated? This question has been a recurrent concern of the speech community, and has also been vigorously investigated. It has been fueled further by its link to the issue of validity of acted emotion databases. Much of the speech and vocal emotion research has relied on acted emotion databases as valid proxies for studying natural emotions. To create models that generalize to natural settings, it is crucial to work with valid prototypes -- ones that can be assumed to reliably represent natural emotions. More concretely, it is important to study emulated emotions against natural emotions in terms of their physiological, and psychological concomitants. In this paper, we present an on-scale systematic study of the differences between natural and acted vocal emotions. We use a self-attention based emotion classification model to understand the phonetic bases of emotions by discovering the most 'attended' phonemes for each class of emotions. We then compare these attended-phonemes in their importance and distribution across acted and natural classes. Our tests show significant differences in the manner and choice of phonemes in acted and natural speech, concluding moderate to low validity and value in using acted speech databases for emotion classification tasks. Hira Dhamyal, Shahan Ali Memon, Bhiksha Raj, Rita Singh |
INTERSPEECH | 3 |
| 2020 | Hide and Speak: Towards Deep Neural Networks for Speech SteganographyabstractSteganography is the science of hiding a secret message within an ordinary public message, which is referred to as Carrier. Traditionally, digital signal processing techniques, such as least significant bit encoding, were used for hiding messages. In this paper, we explore the use of deep neural networks as steganographic functions for speech data. We showed that steganography models proposed for vision are less suitable for speech, and propose a new model that includes the short-time Fourier transform and inverse-short-time Fourier transform as differentiable layers within the network, thus imposing a vital constraint on the network outputs. We empirically demonstrated the effectiveness of the proposed method comparing to deep learning based on several speech datasets and analyzed the results quantitatively and qualitatively. Moreover, we showed that the proposed approach could be applied to conceal multiple messages in a single carrier using multiple decoders or a single conditional decoder. Lastly, we evaluated our model under different channel distortions. Qualitative experiments suggest that modifications to the carrier are unnoticeable by human listeners and that the decoded messages are highly intelligible. Felix Kreuk, Yossi Adi, Bhiksha Raj, Rita Singh, Joseph Keshet |
INTERSPEECH | 3 |
| 2020 | Automatic In-the-wild Dataset Annotation with Deep Generalized Multiple Instance LearningabstractThe automation of the diagnosis and monitoring of speech affecting diseases in real life situations, such as Depression or Parkinson’s disease, depends on the existence of rich and large datasets that resemble real life conditions, such as those collected from in-the-wild multimedia repositories like YouTube. However, the cost of manually labeling these large datasets can be prohibitive. In this work, we propose to overcome this problem by automating the annotation process, without any requirements for human intervention. We formulate the annotation problem as a Multiple Instance Learning (MIL) problem, and propose a novel solution that is based on end-to-end differentiable neural networks. Our solution has the additional advantage of generalizing the MIL framework to more scenarios where the data is stil organized in bags but does not meet the MIL bag label conditions. We demonstrate the performance of the proposed method in labeling the in-the-Wild Speech Medical (WSM) Corpus, using simple textual cues extracted from videos and their metadata. Furthermore we show what is the contribution of each type of textual cues for the final model performance, as well as study the influence of the size of the bags of instances in determining the difficulty of the learning problem Maria Joana Correia, Isabel Trancoso, Bhiksha Raj |
LREC | 3 |
| 2020 | Sherlock: A Crowd-sourced System For Automatic Tagging Of Indoor Floor PlansabstractHaving knowledge of the users' indoor location and the semantics of their environment can facilitate the development of many indoor context-aware applications. For such applications, an accurate indoor map is often needed. While current techniques are capable of producing such maps, these maps are not labeled and hence are of limited utility for many applications. To address this shortcoming, we propose Sherlock, a crowdsourced system for automatically tagging indoor floor plans. Sherlock leverages the myriad of sensors embedded in modern smartphones to intelligently gather audio and visual data, and upload it to the Sherlock Server. At the Sherlock Server, acoustic monitoring and object recognition techniques are used to classify these data samples. The classification scores of current and past samples are then aggregated in a probabilistic framework to determine the confidence with which we can apply as label to a given space. We evaluate Sherlock on a dataset of more than 11,000 audio recordings and 1,200 images, that we collected in three different university campuses. In our evaluation, the confidence for the true label generally outstripped the confidence for all other labels and, in some cases, even reached as high as 100% with as little as 30 data samples. Muhammad Ahmed Shah, Khaled A. Harras, Bhiksha Raj |
MASS | 3 |
| 2020 | Is normalization indispensable for training deep neural network?abstractNormalization operations are widely used to train deep neural networks, and they can improve both convergence and generalization in most tasks. The theories for normalization's effectiveness and new forms of normalization have always been hot topics in research. To better understand normalization, one question can be whether normalization is indispensable for training deep neural network? In this paper, we study what would happen when normalization layers are removed from the network, and show how to train deep neural networks without normalization layers and without performance degradation. Our proposed method can achieve the same or even slightly better performance in a variety of tasks: image classification in ImageNet, object detection and segmentation in MS-COCO, video classification in Kinetics, and machine translation in WMT English-German, etc. Our study may help better understand the role of normalization layers and can be a competitive alternative to normalization layers. Codes are available. Jie Shao 0006, Kai Hu 0010, Changhu Wang, Xiangyang Xue 0001, Bhiksha Raj |
NeurIPS | 5 |
| 2019 | In-the-Wild End-to-End Detection of Speech Affecting DiseasesabstractSpeech is a complex bio-signal that has the potential to provide a rich bio-marker for health. It enables the development of non-invasive routes to early diagnosis and monitoring of speech affecting diseases, such as the ones studied in this work: Depression, and Parkinson's Disease. However, the major limitation of current speech based diagnosis and monitoring tools is the lack of large and diverse datasets. Existing datasets are small, and collected under very controlled conditions. As such, there is an upper bound in the complexity of the models that can be trained using these datasets. There is also limited applicability in real life scenarios where the channel and noise conditions, among others, are impossible to control. In this work, we show that datasets collected from in-the-wild sources, such as collections of vlogs, can contribute to improve the performance of diagnosis tools both in controlled and in-the-wild conditions, even though the data are noisier. Moreover, we show that it is possible to successfully move away from hand-crafted features (i.e. features that are computed based on predefined algorithms, that based on human expertise) and adopt end-to-end modeling paradigms, such as CNN-LSTMs, that extract data driven features from the raw spectrograms of the speech signal, and capture temporal information from the speech signals. Maria Joana Correia, Isabel Trancoso, Bhiksha Raj |
ASRU | 3 |
| 2019 | Optimizing Neural Network Embeddings Using a Pair-Wise Loss for Text-Independent Speaker VerificationabstractThis paper proposes a new loss function called the “quartet” loss for the better optimization of the neural networks for matching tasks. For such tasks, where neural network embeddings are the key component, the optimization of the network for better embeddings is critical. The embeddings are required to be class discriminative, resulting in minimal inter-class variation and maximal intra-class variation even for unseen classes for better generalization of the network. The quartet loss explicitly computes the distance metric between pairs of inputs and increases the gap between the similarity score distributions between the same class pairs and the different class pairs. We evaluate on the speaker verification task and demonstrate the performance of the loss on our proposed neural network. Hira Dhamyal, Tianyan Zhou, Bhiksha Raj, Rita Singh |
ASRU | 3 |
| 2019 | Cross Modal Audio Search and Retrieval with Joint Embeddings Based on Text and AudioabstractExisting audio search engines use one of two approaches: matching text-text or audio-audio pairs. In the former, text queries are matched to semantically similar words in an index of audio metadata to retrieve corresponding audio clips or segments, while in the latter, audio signals are directly used to retrieve acoustically-similar recordings from an audio database. However, independent treatment of text and audio has precluded information exchange between the two modalities. This is a problem because similarity in language does not always imply similarity in acoustics, and vice versa. Moreover, independent modeling can be error prone especially for ad hoc, user-generated recordings, which are noisy in both audio and their associated textual labels. To overcome this limitation, we propose a framework that learns joint embeddings from a shared lexico-acoustic space, where vectors from either modality can be mapped together and compared directly. Thus, we improve semantic knowledge and enable the use of either text or audio queries to search and retrieve audio. Our results break new ground for a cross-modal audio search engine, and further exploration of lexico-acoustic spaces. Benjamin Elizalde, Shuayb Zarar, Bhiksha Raj |
ICASSP | 3 |
| 2019 | Time Signal Classification Using Random Convolutional FeaturesabstractIn this paper we present a transformation to convert time signals into a randomized low-dimensional vectors such that the inner product between these new features provides information about the similarity of the signals. We show that the described inner product approximates a cross-correlation based kernel. This is very useful at the moment of use Kernel Machines, such as Non-linear Support Vector Machines. Indeed, this allows to apply simpler and faster linear methods on the generated random features. Our proposed scheme improves computational storage and time cost over the direct kernel approach, while performing the classification performance with minimal loss. We support our statements by providing theoretical guarantees as well as empirical evaluation across different data sets. Abelino Jiménez, Bhiksha Raj |
ICASSP | 2 |
| 2019 | Human Behaviour Recognition Using Wifi Channel State InformationabstractThe following topics are dealt with: learning (artificial intelligence); neural nets; speech recognition; feature extraction; convolutional neural nets; acoustic signal processing; speech processing; optimisation; speaker recognition; recurrent neural nets. Daanish Ali Khan, Saquib Razak, Bhiksha Raj, Rita Singh |
ICASSP | 3 |
| 2019 | Disjoint Mapping Network for Cross-modal Matching of Voices and Faces
Yandong Wen, Mahmoud Al Ismail, Weiyang Liu, Bhiksha Raj, Rita Singh |
ICLR (Poster) | 4 |
| 2019 | Learning Sound Events from Webly Labeled DataabstractIn the last couple of years, weakly labeled learning has turned out to be an exciting approach for audio event detection. In this work, we introduce webly labeled learning for sound events which aims to remove human supervision altogether from the learning process. We first develop a method of obtaining labeled audio data from the web (albeit noisy), in which no manual labeling is involved. We then describe methods to efficiently learn from these webly labeled audio recordings. In our proposed system, WeblyNet, two deep neural networks co-teach each other to robustly learn from webly labeled data, leading to around 17% relative improvement over the baseline method. The method also involves transfer learning to obtain efficient representations. Anurag Kumar 0003, Ankit Shah 0001, Alex Hauptmann 0001, Bhiksha Raj |
IJCAI | 4 |
| 2019 | Neural Regression TreesabstractRegression-via-Classification (RvC) is the process of converting a regression problem to a classification one. Current approaches for RvC use ad-hoc discretization strategies and are suboptimal. We propose a neural regression tree model for RvC. In this model, we employ a joint optimization framework where we learn optimal discretization thresholds while simultaneously optimizing the features for each node in the tree. We empirically show the validity of our model by testing it on two challenging regression tasks where we establish the state of the art. Shahan Ali Memon, Wenbo Zhao 0006, Bhiksha Raj, Rita Singh |
IJCNN | 3 |
| 2019 | Face Reconstruction from Voice using Generative Adversarial NetworksabstractVoice profiling aims at inferring various human parameters from their speech, e.g. gender, age, etc. In this paper, we address the challenge posed by a subtask of voice profiling - reconstructing someone's face from their voice. The task is designed to answer the question: given an audio clip spoken by an unseen person, can we picture a face that has as many common elements, or associations as possible with the speaker, in terms of identity? To address this problem, we propose a simple but effective computational framework based on generative adversarial networks (GANs). The network learns to generate faces from voices by matching the identities of generated faces to those of the speakers, on a training set. We evaluate the performance of the network by leveraging a closely related task - cross-modal matching. The results show that our model is able to generate faces that match several biometric characteristics of the speaker, and results in matching accuracies that are much better than chance. The code is publicly available in https://github.com/cmu-mlsp/reconstructingfacesfrom_voices Yandong Wen, Bhiksha Raj, Rita Singh |
NeurIPS | 2 |
| 2019 | Sound Event Detection in the DCASE 2017 ChallengeabstractEach edition of the challenge on Detection and Classification of Acoustic Scenes and Events (DCASE) contained several tasks involving sound event detection in different setups. DCASE 2017 presented participants with three such tasks, each having specific datasets and detection requirements: Task 2, in which target sound events were very rare in both training and testing data, Task 3 having overlapping events annotated in real-life audio, and Task 4, in which only weakly labeled data were available for training. In this paper, we present three tasks, including the datasets and baseline systems, and analyze the challenge entries for each task. We observe the popularity of methods using deep neural networks, and the still widely used mel frequency-based representations, with only few approaches standing out as radically different. Analysis of the systems behavior reveals that task-specific optimization has a big role in producing good performance; however, often this optimization closely follows the ranking metric, and its maximization/minimization does not result in universally good performance. We also introduce the calculation of confidence intervals based on a jackknife resampling procedure to perform statistical analysis of the challenge results. The analysis indicates that while the 95% confidence intervals for many systems overlap, there are significant differences in performance between the top systems and the baseline for all tasks. Annamaria Mesaros, Aleksandr Diment, Benjamin Elizalde, Toni Heittola, Emmanuel Vincent 0001, Bhiksha Raj, Tuomas Virtanen |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2018 | Framework for Evaluation of Sound Event Detection in Web VideosabstractThe largest source of sound events is web videos. Most videos lack sound event labels at segment level, however, a significant number of them do respond to text queries, from a match found using metadata by search engines. In this paper we explore the extent to which a search query can be used as the true label for detection of sound events in videos. We present a framework for large-scale sound event recognition on web videos. The framework crawls videos using search queries corresponding to 78 sound event labels drawn from three datasets. The datasets are used to train three classifiers, and we obtain a prediction on 3.7 million web video segments. We evaluated performance using the search query as true label and compare it with human labeling. Both types of ground truth exhibited close performance, to within 10%, and similar performance trend with increasing number of evaluated segments. Hence, our experiments show potential for using search query as a preliminary true label for sound event recognition in web videos. Rohan Badlani, Ankit Shah 0001, Benjamin Elizalde, Anurag Kumar 0003, Bhiksha Raj |
ICASSP | 5 |
| 2018 | Voice Impersonation Using Generative Adversarial NetworksabstractVoice impersonation is not the same as voice transformation, although the latter is an essential element of it. In voice impersonation, the resultant voice must convincingly convey the impression of having been naturally produced by the target speaker, mimicking not only the pitch and other perceivable signal qualities, but also the style of the target speaker. In this paper, we propose a novel neural-network based speech quality- and style-mimicry framework for the synthesis of impersonated voices. The framework is built upon a fast and accurate generative adversarial network model. Given spectrographic representations of source and target speakers' voices, the model learns to mimic the target speaker's voice quality and style, regardless of the linguistic content of either's voice, generating a synthetic spectrogram from which the time-domain signal is reconstructed using the Griffin-Lim method. In effect, this model reframes the well-known problem of style-transfer for images as the problem of style-transfer for speech signals, while intrinsically addressing the problem of durational variability of speech sounds. Experiments demonstrate that the model can generate extremely convincing samples of impersonated speech. It is even able to impersonate voices across different genders effectively. Results are qualitatively evaluated using standard procedures for evaluating synthesized voices. Yang Gao 0029, Rita Singh, Bhiksha Raj |
ICASSP | 3 |
| 2018 | Acoustic Scene Classification Using Discrete Random Hashing for Laplacian Kernel MachinesabstractState of the art acoustic scene classification techniques often employ features of large dimensionality, which are then used to train and perform inferences with kernel machines such as Support Vector Machines. However, the complexity of computing the non-linear kernel matrix for these methods increases with the dimensionality of the features and the size of the dataset. In this work, we introduce a new scheme that hashes features, which combined with a linear function approximates a non-linear Laplacian kernel. Each hash typically has lower dimensionality than the input features and each component is represented by one bit instead of floating values. Hence, allowing efficient computation of the kernel matrix using XOR operations rather than dot-products. Our scheme is demonstrated mathematically and tested in the 2017 DCASE: Acoustic Scene Classification. The hashes reduce up to six powers of two the feature representation with minimal loss of accuracy. Abelino Jiménez, Benjamin Elizalde, Bhiksha Raj |
ICASSP | 3 |
| 2018 | Content-Based Representations of Audio Using Siamese Neural NetworksabstractIn this paper, we focus on the problem of content-based retrieval for audio, which aims to retrieve all semantically similar audio recordings for a given audio clip query. This problem is similar to the problem of query by example of audio, which aims to retrieve media samples from a database, which are similar to the user-provided example. We propose a novel approach which encodes the audio into a vector representation using Siamese Neural Networks. The goal is to obtain an encoding similar for files belonging to the same audio class, thus allowing retrieval of semantically similar audio. Using simple similarity measures such as those based on simple euclidean distance and cosine similarity we show that these representations can be very effectively used for retrieving recordings similar in audio content. Pranay Manocha, Rohan Badlani, Anurag Kumar 0003, Ankit Shah 0001, Benjamin Elizalde, Bhiksha Raj |
ICASSP | 6 |
| 2018 | A Corrective Learning Approach for Text-Independent Speaker VerificationabstractWe present a conceptually plausible approach for text-independent speaker verification (TISV) which treats speech recordings as a collection of segments providing incremental evidence. This approach, called corrective learning, gradually improves an initial prediction of speaker identity based on incoming speech and the latest prediction. Specifically, we propose deep corrective learning networks (CLNets) that explicitly learn a mapping from a new speech segment and the current predictions, to a correction. Intuitively, the predictions eventually converge to the ground truth after several corrections. Trained on NIST SRE datasets, CLNets outperform current CNN and the i-vector baselines. Moreover, CLNets and i-vectors are complementary, and their fusion leads to significant performance improvements compared to what can be achieved by each of them individually. Yandong Wen, Tianyan Zhou, Rita Singh, Bhiksha Raj |
ICASSP | 4 |
| 2018 | Interactive Evaluation of Classifiers Under Limited ResourcesabstractIn this paper, we propose strategies to estimate the accuracy of classifiers on a dataset when resource limitations restrict the number of instances for which true labels can be obtained. Our target scenarios include situations where the classifier output labels, but no scores, e.g. when the "classifier" is not an automated classifier but an inexpert human labeller who only outputs labels. Our objective is to optimally select a subset of the data to obtain true labels for, such that they provide the best estimate of classifier accuracy. We use techniques based on stratified sampling to address this problem. However, stratified sampling poses two challenges: i) how best to stratify the data, and ii) how to allocate samples among the strata. We propose a method of stratifying data and then present two novel interactive algorithms to approximate optimal allocation of samples to the strata. Our proposed methods for stratification and allocation are seen to outperform other popular approaches to the problem. Sabit Hassan, Shaden Shaar, Bhiksha Raj, Saquib Razak |
ICMLA | 3 |
| 2018 | Mining Multimodal Repositories for Speech Affecting Diseases
Maria Joana Correia, Bhiksha Raj, Isabel Trancoso, Francisco Teixeira |
INTERSPEECH | 2 |
| 2018 | Classifier Risk Estimation Under Limited Labeling Resources
Anurag Kumar 0003, Bhiksha Raj |
PAKDD (1) | 2 |
| 2018 | Querying Depression VlogsabstractSpeech based diagnosis-aid tools for depression typically depend on few and small datasets, that are expensive to collect. The limited availability of training data poses a limitation to the quality that these systems can achieve. An unexplored alternative for large scale source of data are vlogs collected from online multimedia repositories. Along with the automation of the mining process, it is necessary to automate the labeling process too.In this work, we propose a framework to automatically label a corpus of in-the-wild vlogs of possibly depressed subjects, and we estimate the quality of the predicted labels, without ever having access to a ground truth for the majority of the corpus. The framework uses a small subset to train a model and estimate the labels for the remainder of the corpus. Then, using the predicted labels, we train a noisy model and attempt to reconstruct the labels of the original labeled subset. We hypothesize that the quality of the estimated labels for the unlabelled subset of the corpus is correlated to the quality of the label reconstruction of the labeled subset.The results of the bi-modal experiment using in-the-wild data are compared to the ones obtained using controlled data. Maria Joana Correia, Bhiksha Raj, Isabel Trancoso |
SLT | 2 |
| 2017 | SphereFace: Deep Hypersphere Embedding for Face RecognitionabstractThis paper addresses deep face recognition (FR) problem under open-set protocol, where ideal face features are expected to have smaller maximal intra-class distance than minimal inter-class distance under a suitably chosen metric space. However, few existing algorithms can effectively achieve this criterion. To this end, we propose the angular softmax (A-Softmax) loss that enables convolutional neural networks (CNNs) to learn angularly discriminative features. Geometrically, A-Softmax loss can be viewed as imposing discriminative constraints on a hypersphere manifold, which intrinsically matches the prior that faces also lie on a manifold. Moreover, the size of angular margin can be quantitatively adjusted by a parameter m. We further derive specific m to approximate the ideal feature criterion. Extensive analysis and experiments on Labeled Face in the Wild (LFW), Youtube Faces (YTF) and MegaFace Challenge 1 show the superiority of A-Softmax loss in FR tasks. Weiyang Liu, Yandong Wen, Zhiding Yu, Ming Li 0026, Bhiksha Raj |
CVPR | 5 |
| 2017 | Discovering sound concepts and acoustic relations in textabstractIn this paper we describe approaches for discovering acoustic concepts and relations in text. The first major goal is to be able to identify text phrases which contain a notion of audibility and can be termed as a sound or an acoustic concept. We also propose a method to define an acoustic scene through a set of sound concepts. We use pattern matching and parts of speech tags to generate sound concepts from large scale text corpora. We use dependency parsing and LSTM recurrent neural network to predict a set of sound concepts for a given acoustic scene. These methods are not only helpful in creating an acoustic knowledge base but in the future can also directly help acoustic event and scene detection research. Anurag Kumar 0003, Bhiksha Raj, Ndapandula Nakashole |
ICASSP | 2 |
| 2017 | Privacy preserving Distance computation using somewhat-trusted third partiesabstractA critically important component of most signal processing procedures is that of computing the distance between signals. In multiparty processing applications where these signals belong to different parties, this introduces privacy challenges. The signals may themselves be private, and the parties to the computation may not be willing to expose them. Solutions proposed to the problem in the literature generally invoke homomorphic encryption schemes, secure multi-party computation, or other cryptographic methods which introduce significant computational complexity into the proceedings, often to the point of making more complex computations requiring repeated computations unfeasible. Other solutions invoke third parties, making unrealistic assumptions about their trustworthiness. In this paper we propose an alternate approach, also based on third party computation, but without assuming as much trust in the third party. Individual participants to the computation “secure” their data through a proposed secure hashing scheme with shared keys, prior to sharing it with the third party. The hashing ensures that the third party cannot recover any information about the individual signals or their statistics, either from analysis of individual computations or their long-term aggregate patterns. We provide theoretical proof of these properties and empirical demonstration of the feasibility of the computation. Abelino Jiménez, Bhiksha Raj |
ICASSP | 2 |
| 2017 | Supervised monaural source separation based on autoencodersabstractIn this paper, we propose a new supervised monaural source separation based on autoencoders. We employ the autoencoder for the dictionary training such that the nonlinear network can encode the target source with high expressiveness. The dictionary is trained by each target source without the mixture signal, which makes the system independent from the context where the dictionaries will be used. In separation process, the decoder portions of the trained autoencoders are used as dictionaries to find the activations in a iterative manner such that a summation of the decoder outputs approximates the original mixture. The results of the instruments source separation experiments revealed that the separation performance of the proposed method was superior to that of the NMF. Keiichi Osako, Yuki Mitsufuji, Rita Singh, Bhiksha Raj |
ICASSP | 4 |
| 2017 | Audio event and scene recognition: A unified approach using strongly and weakly labeled dataabstractIn this paper we propose a novel learning framework called Supervised and Weakly Supervised Learning where the goal is to learn simultaneously from weakly and strongly labeled data. Strongly labeled data can be simply understood as fully supervised data where all labeled instances are available. In weakly supervised learning only data is weakly labeled which prevents one from directly applying supervised learning methods. Our proposed framework is motivated by the fact that a small amount of strongly labeled data can give considerable improvement over only weakly supervised learning. The primary problem domain focus of this paper is acoustic event and scene detection in audio recordings. We first propose a naive formulation for leveraging labeled data in both forms. We then propose a more general framework for Supervised and Weakly Supervised Learning (SWSL). Based on this general framework, we propose a graph based approach for SWSL. Our main method is based on manifold regularization on graphs in which we show that the unified learning can be formulated as a constraint optimization problem which can be solved by iterative concave-convex procedure (CCCP). Our experiments show that our proposed framework can address several concerns of audio content analysis using weakly labeled data. Bhiksha Raj, Anurag Kumar 0003 |
IJCNN | 1 |
| 2017 | Audio Content Based Geotagging in MultimediaabstractIn this paper we propose methods to extract geographically relevant information in a multimedia recording using its audio. Our method primarily is based on the fact that urban acoustic environment consists of a variety of sounds. Hence, location information can be inferred from the composition of sound events/classes present in the audio. More specifically, we adopt matrix factorization techniques to obtain semantic content of recording in terms of different sound classes. These semantic information are then combined to identify the location of recording. Anurag Kumar 0003, Benjamin Elizalde, Bhiksha Raj |
INTERSPEECH | 3 |
| 2017 | Hidden Markov Model Variational Autoencoder for Acoustic Unit DiscoveryabstractVariational Autoencoders (VAEs) have been shown to provide efficient neural-network-based approximate Bayesian inference for observation models for which exact inference is intractable. Its extension, the so-called Structured VAE (SVAE) allows inference in the presence of both discrete and continuous latent variables. Inspired by this extension, we developed a VAE with Hidden Markov Models (HMMs) as latent models. We applied the resulting HMM-VAE to the task of acoustic unit discovery in a zero resource scenario. Starting from an initial model based on variational inference in an HMM with Gaussian Mixture Model (GMM) emission probabilities, the accuracy of the acoustic unit discovery could be significantly improved by the HMM-VAE. In doing so we were able to demonstrate for an unsupervised learning task what is well-known in the supervised learning case: Neural networks provide superior modeling power compared to GMMs. Janek Ebbers, Jahn Heymann, Lukas Drude, Thomas Glarner, Reinhold Häb-Umbach, Bhiksha Raj |
INTERSPEECH | 6 |
| 2016 | The relationship of voice onset time and Voice Offset Time to physical ageabstractIn a speech signal, Voice Onset Time (VOT) is the period between the release of a plosive and the onset of vocal cord vibrations in the production of the following sound. Voice Offset Time (VOFT), on the other hand, is the period between the end of a voiced sound and the release of the following plosive. Traditionally, VOT has been studied across multiple disciplines and has been related to many factors that influence human speech production, including physical, physiological and psychological characteristics of the speaker. The mechanism of extraction of VOT has however been largely manual, and studies have been carried out over small ensembles of individuals under very controlled conditions, usually in clinical settings. Studies of VOFT follow similar trends, but are more limited in scope due to the inherent difficulty in the extraction of VOFT from speech signals. In this paper we use a structured-prediction based mechanism for the automatic computation of VOT and VOFT. We show that for specific combinations of plosives and vowels, these are re-latable to the physical age of the speaker. The paper also highlights the ambiguities in the prediction of age from VOT and VOFT, and consequently in the use of these measures in forensic analysis of voice. Rita Singh, Joseph Keshet, Deniz Gençaga, Bhiksha Raj |
ICASSP | 4 |
| 2016 | Weakly supervised scalable audio content analysisabstractAudio Event Detection is an important task for content analysis of multimedia data. Most of the current works on detection of audio events is driven through supervised learning approaches. We propose a weakly supervised learning framework which can make use of the tremendous amount of web multimedia data with significantly reduced annotation effort and expense. Specifically, we use several multiple instance learning algorithms to show that audio event detection through weak labels is feasible. We also propose a novel scalable multiple instance learning algorithm and show that its competitive with other multiple instance learning algorithms for audio event detection tasks. Anurag Kumar 0003, Bhiksha Raj |
ICME | 2 |
| 2016 | Viral Spread via Entertainment and Voice-Messaging Among Telephone Users in IndiaabstractWe explore how development-related, voice-based, information services could organically spread among low-literate masses in the developing world. We report lessons learned from a remote deployment of "Polly" in India (from the US) to spread job-related information. Polly is an entertainment driven, voice-based service, available over simple phones that is aimed at familiarizing people with speech interfaces and mass-dissemination of development related information to low-literate users. In 2012, Polly had become viral in Pakistan and successfully spread recorded newspaper job ads to thousands of mobile phone users. Remotely deployed in India, Polly did not take off immediately as it did in Pakistan. Instead, it initially entered a six-month long phase of fluctuating, intermittent activity. We experimented with various forms of seeding and it eventually transitioned into a viral phase, with sustained transmission that continued for five months but without (exponential) growth. Finally, interface adjustments in response to user feedback enabling plain-voice asynchronous voice-messaging resulted in an abrupt exponential and viral growth amassing 10,349 phone calls by 1,613 users over a span of seven days. Of these, 299 users also transitioned to the job service. User feedback and surveys suggest possible reasons for each phase. We study the challenges of remote deployment and the interplay of user interface; language of the system; seeding mechanisms and active response to user feedback towards the uptake of the service. We also report a detailed comparison of viral spread in the two countries. Agha Ali Raza, Rajat Kulshreshtha, Spandana Gella, Sean Olin Blagsvedt, Maya Chandrasekaran, Bhiksha Raj, Ronald Rosenfeld |
ICTD | 6 |
| 2016 | On the Appropriateness of Complex-Valued Neural Networks for Speech EnhancementabstractAlthough complex-valued neural networks (CVNNs) â?? networks which can operate with complex arithmetic â?? have been around for a while, they have not been given reconsideration since the breakthrough of deep network architectures. This paper presents a critical assessment whether the novel tool set of deep neural networks (DNNs) should be extended to complex-valued arithmetic. Indeed, with DNNs making inroads in speech enhancement tasks, the use of complex-valued input data, specifically the short-time Fourier transform coefficients, is an obvious consideration. In particular when it comes to performing tasks that heavily rely on phase information, such as acoustic beamforming, complex-valued algorithms are omnipresent. In this contribution we recapitulate backpropagation in CVNNs, develop complex-valued network elements, such as the split-rectified non-linearity, and compare real- and complex-valued networks on a beamforming task. We find that CVNNs hardly provide a performance gain and conclude that the effort of developing the complex-valued counterparts of the building blocks of modern deep or recurrent neural networks can hardly be justified. Lukas Drude, Bhiksha Raj, Reinhold Häb-Umbach |
INTERSPEECH | 2 |
| 2016 | Audio Event Detection using Weakly Labeled DataabstractAcoustic event detection is essential for content analysis and description of multimedia recordings. The majority of current literature on the topic learns the detectors through fully-supervised techniques employing strongly labeled data. However, the labels available for majority of multimedia data are generally weak and do not provide sufficient detail for such methods to be employed. In this paper we propose a framework for learning acoustic event detectors using only weakly labeled data. We first show that audio event detection using weak labels can be formulated as an Multiple Instance Learning problem. We then suggest two frameworks for solving multiple-instance learning, one based on support vector machines, and the other on neural networks. The proposed methods can help in removing the time consuming and expensive process of manually annotating data to facilitate fully supervised learning. Moreover, it can not only detect events in a recording but can also provide temporal locations of events in the recording. This helps in obtaining a complete description of the recording and is notable since temporal information was never known in the first place in weakly labeled data. Anurag Kumar 0003, Bhiksha Raj |
ACM Multimedia | 2 |
| 2016 | Adaptation of SVM for MIL for inferring the polarity of movies and movie reviewsabstractPolarity detection is a research topic of major interest, with many applications including detecting the polarity of product reviews. However, in some cases, the polarity of the product reviews might not be available while the polarity of the product itself might be, prohibiting the use of any form of fully supervised learning technique. This scenario, while different, is close to that of multiple instance learning (MIL). In this work we propose two new adaptations of support vector machines (SVM) for MIL, θ-MIL, to suit this new scenario, and infer the polarity of products and product reviews. We perform experiments on the proposed methods using the IMDb movie review corpus, and compare the performance of the proposed methods to the traditional SVM for MIL approach. Although we make weaker assumptions about the data, the proposed methods achieve a comparable performance to the SVM for MIL in accurately detecting the polarity of movies and movie reviews. Maria Joana Correia, Isabel Trancoso, Bhiksha Raj |
SLT | 3 |
| 2016 | Learning Model-Based Sparsity via Projected Gradient DescentabstractSeveral convex formulation methods have been proposed previously for statistical estimation with structured sparsity as the prior. These methods often require a carefully tuned regularization parameter, often a cumbersome or heuristic exercise. Furthermore, the estimate that these methods produce might not belong to the desired sparsity model, albeit accurately approximating the true parameter. Therefore, greedy-type algorithms could often be more desirable in estimating structured-sparse parameters. So far, these greedy methods have mostly focused on linear statistical models. In this paper, we study the projected gradient descent with a non-convex structured-sparse parameter model as the constraint set. Should the cost function have a stable model-restricted Hessian, the algorithm produces an approximation for the desired minimizer. As an example, we elaborate on application of the main results to estimation in generalized linear models. Sohail Bahmani, Petros Boufounos, Bhiksha Raj |
IEEE Trans. Inf. Theory | 3 |
| 2015 | Efficient autism spectrum disorder prediction with eye movement: A machine learning frameworkabstractWe propose an autism spectrum disorder (ASD) prediction system based on machine learning techniques. Our work features the novel development and application of machine learning methods over traditional ASD evaluation protocols. Specifically, we are interested in discovering the latent patterns that possibly indicate the symptom of ASD underneath the observations of eye movement. A group of subjects (either ASD or non-ASD) are shown with a set of aligned human face images, with eye gaze locations on each image recorded sequentially. An image-level feature is then extracted from the recorded eye gaze locations on each face image. Such feature extraction process is expected to capture discriminative eye movement patterns related to ASD. In this work, we propose a variety of feature extraction methods, seeking to evaluate their prediction performance comprehensively. We further propose an ASD prediction framework in which the prediction model is learned on the labeled features. At testing stage, a test subject is also asked to view the face images with eye gaze locations recorded. The learned model predicts the image-level labels and a threshold is set to determine whether the test subject potentially has ASD or not. Despite the inherent difficulty of ASD prediction, experimental results indicates statistical significance of the predicted results, showing promising perspective of this framework. Wenbo Liu 0002, Zhiding Yu, Xiaobing Zou, Bhiksha Raj, Ming Li 0026 |
ACII | 5 |
| 2015 | Beyond Gaussian Pyramid: Multi-skip Feature Stacking for action recognitionabstractMost state-of-the-art action feature extractors involve differential operators, which act as highpass filters and tend to attenuate low frequency action information. This attenuation introduces bias to the resulting features and generates ill-conditioned feature matrices. The Gaussian Pyramid has been used as a feature enhancing technique that encodes scale-invariant characteristics into the feature space in an attempt to deal with this attenuation. However, at the core of the Gaussian Pyramid is a convolutional smoothing operation, which makes it incapable of generating new features at coarse scales. In order to address this problem, we propose a novel feature enhancing technique called Multi-skIp Feature Stacking (MIFS), which stacks features extracted using a family of differential filters parameterized with multiple time skips and encodes shift-invariance into the frequency space. MIFS compensates for information lost from using differential operators by recapturing information at coarse scales. This recaptured information allows us to match actions at different speeds and ranges of motion. We prove that MIFS enhances the learnability of differential-based features exponentially. The resulting feature matrices from MIFS have much smaller conditional numbers and variances than those from conventional methods. Experimental results show significantly improved performance on challenging action recognition and event detection tasks. Specifically, our method exceeds the state-of-the-arts on Hollywood2, UCF101 and UCF50 datasets and is comparable to state-of-the-arts on HMDB51 and Olympics Sports datasets. MIFS can also be used as a speedup strategy for feature extraction with minimal or no accuracy cost. Zhen-Zhong Lan, Ming Lin 0002, Xuanchong Li, Alex Hauptmann 0001, Bhiksha Raj |
CVPR | 5 |
| 2015 | A novel ranking method for multiple classifier systemsabstractWe introduce an unsupervised optimization method for optimal fusion of multiple classifiers in retrieval problems. The method is based on a ranking loss called the “clarity” index, which does not depend on the label of the test instances. The technique optimizes the weights with which individual classifier scores must be combined to maximize this clarity. Our method is instance-specific; the weights are optimized individually for each test instance. The proposed schema can also be used for instance-specific ranking of classifiers. We also show that the method is highly tolerant to the introduction of noise in classifier outputs. Anurag Kumar 0003, Bhiksha Raj |
ICASSP | 2 |
| 2015 | Reducing communication overhead in distributed learning by an order of magnitude (almost)abstractLarge-scale distributed learning plays an ever-more increasing role in modern computing. However, whether using a compute cluster with thousands of nodes, or a single multi-GPU machine, the most significant bottleneck is that of communication. In this work, we explore the effects of applying quantization and encoding to the parameters of distributed models. We show that, for a neural network, this can be done - without slowing down the convergence, or hurting the generalization of the model. In fact, in our experiments we were able to reduce the communication overhead by nearly an order of magnitude - while actually improving the generalization accuracy. Anders Øland, Bhiksha Raj |
ICASSP | 2 |
| 2015 | Privacy-preserving Query-by-Example Speech SearchabstractThis paper investigates a new privacy-preserving paradigm for the task of Query-by-Example Speech Search using Secure Binary Embeddings, a hashing method that converts vector data to bit strings through a combination of random projections followed by banded quantization. The proposed method allows performing spoken query search in an encrypted domain, by analyzing ciphered information computed from the original recordings. Unlike other hashing techniques, the embeddings allow the computation of the distance between vectors that are close enough, but are not perfect matches. This paper shows how these hashes can be combined with Dynamic Time Warping based on posterior derived features to perform secure speech search. Experiments performed on a sub-set of the Speech-Dat Portuguese corpus showed that the proposed privacy-preserving system obtains similar results to its non-private counterpart. José Portelo, Alberto Abad, Bhiksha Raj, Isabel Trancoso |
ICASSP | 3 |
| 2015 | Locality constrained transitive distance clustering on speech data
Wenbo Liu 0002, Zhiding Yu, Bhiksha Raj, Ming Li 0026 |
INTERSPEECH | 3 |
| 2014 | Iterative Bayesian word segmentation for unsupervised vocabulary discovery from phoneme latticesabstractIn this paper we present an algorithm for the unsupervised segmentation of a lattice produced by a phoneme recognizer into words. Using a lattice rather than a single phoneme string accounts for the uncertainty of the recognizer about the true label sequence. An example application is the discovery of lexical units from the output of an error-prone phoneme recognizer in a zero-resource setting, where neither the lexicon nor the language model (LM) is known. We propose a computationally efficient iterative approach, which alternates between the following two steps: First, the most probable string is extracted from the lattice using a phoneme LM learned on the segmentation result of the previous iteration. Second, word segmentation is performed on the extracted string using a word and phoneme LM which is learned alongside the new segmentation. We present results on lattices produced by a phoneme recognizer on the WSJ-CAM0 dataset. We show that our approach delivers superior segmentation performance than an earlier approach found in the literature, in particular for higher-order language models. Jahn Heymann, Oliver Walter, Reinhold Häb-Umbach, Bhiksha Raj |
ICASSP | 4 |
| 2014 | Active-set newton algorithm for non-negative sparse coding of audioabstractWe propose a new algorithm to efficiently obtain non-negative sparse representations for audio. The spectrum of an audio signal is represented as a sparse linear combination of atoms taken from an overcomplete dictionary. The algorithm is based on minimizing the generalized Kullback-Leibler divergence between an observed magnitude spectrum and a non-negative linear combination of atoms, plus an ℓ1regularization term. The proposed method consists of an active-set method that iteratively updates a set of active atoms that have non-zero weights, using a Newton step where the weights of the active atoms are updated. The proposed method was evaluated using mixtures of two speakers, and it was shown to yield more than 10 times faster convergence in comparison to an established algorithm based on multiplicative update rules. Moreover, the ℓ1regularization was found to decrease the computation time and to improve the source separation performance. Tuomas Virtanen, Bhiksha Raj, Jort F. Gemmeke, Hugo Van hamme |
ICASSP | 2 |
| 2014 | Post-masking: a hybrid approach to array processing for speech recognition
Amir R. Moghimi, Bhiksha Raj, Richard M. Stern |
INTERSPEECH | 2 |
| 2013 | Unsupervised word segmentation from noisy inputabstractIn this paper we present an algorithm for the unsupervised segmentation of a character or phoneme lattice into words. Using a lattice at the input rather than a single string accounts for the uncertainty of the character/phoneme recognizer about the true label sequence. An example application is the discovery of lexical units from the output of an error-prone phoneme recognizer in a zero-resource setting, where neither the lexicon nor the language model is known. Recently a Weighted Finite State Transducer (WFST) based approach has been published which we show to suffer from an issue: language model probabilities of known words are computed incorrectly. Fixing this issue leads to greatly improved precision and recall rates, however at the cost of increased computational complexity. It is therefore practical only for single input strings. To allow for a lattice input and thus for errors in the character/phoneme recognizer, we propose a computationally efficient suboptimal two-stage approach, which is shown to significantly improve the word segmentation performance compared to the earlier WFST approach. Jahn Heymann, Oliver Walter, Reinhold Häb-Umbach, Bhiksha Raj |
ASRU | 4 |
| 2013 | A hierarchical system for word discovery exploiting DTW-based initializationabstractDiscovering the linguistic structure of a language solely from spoken input asks for two steps: phonetic and lexical discovery. The first is concerned with identifying the categorical subword unit inventory and relating it to the underlying acoustics, while the second aims at discovering words as repeated patterns of subword units. The hierarchical approach presented here accounts for classification errors in the first stage by modelling the pronunciation of a word in terms of subword units probabilistically: a hidden Markov model with discrete emission probabilities, emitting the observed subword unit sequences. We describe how the system can be learned in a completely unsupervised fashion from spoken input. To improve the initialization of the training of the word pronunciations, the output of a dynamic time warping based acoustic pattern discovery system is used, as it is able to discover similar temporal sequences in the input data. This improved initialization, using only weak supervision, has led to a 40% reduction in word error rate on a digit recognition task. Oliver Walter, Timo Korthals, Reinhold Häb-Umbach, Bhiksha Raj |
ASRU | 4 |
| 2013 | Unsupervised hierarchical structure induction for deeper semantic analysis of audioabstractCurrent audio analysis techniques rely on fairly shallow analysis of audio content, using symbols or patterns extracted directly from the observed acoustics. We hypothesize that the observed acoustics actually map to semantics in a hierarchical manner, and that the higher levels of this hierarchy correspond to increasingly higher-level semantics. In this paper, we present a model for deeper analysis of the observed acoustics, that induces a probabilistic tree structure depending on estimated constituent identities and contexts. Audio characterization using the deeper structure outperforms the standard shallow-feature based characterizations. Sourish Chaudhuri, Bhiksha Raj |
ICASSP | 2 |
| 2013 | Optimization of the DET curve in speaker verification under noisy conditionsabstractThe increasing need for secure authentication systems has motivated recent interest in effective algorithms for Speaker Verification (SV). In particular, there is increasing need for noise robust algorithms for SV, which will allow SV systems to operate successfully in real conditions, which are typically noisy. L. Paola García-Perera, Bhiksha Raj, Juan A. Nolazco-Flores |
ICASSP | 2 |
| 2013 | Speaker tracking with spherical microphone arraysabstractIn prior work, we investigated the application of a spherical microphone array to a distant speech recognition task. In that work, the relative positions of a fixed loud speaker and the spherical array required for beamforming were measured with an optical tracking device. In the present work, we investigate how these relative positions can be determined automatically for real, human speakers based solely on acoustic evidence. We first derive an expression for the complex pressure field of a plane wave scattering from a rigid sphere. We then use this theoretical field as the predicted observation in an extended Kalman filter whose state is the speaker's current position, the direction of arrival of the plane wave. By minimizing the squared-error between the predicted pressure field and that actually recorded, we are able to infer the position of the speaker. John W. McDonough, Ken'ichi Kumatani, Takayuki Arakawa, Kazumasa Yamamoto, Bhiksha Raj |
ICASSP | 5 |
| 2013 | Ensemble approach in speaker verification
L. Paola García-Perera, Bhiksha Raj, Juan A. Nolazco-Flores |
INTERSPEECH | 2 |
| 2013 | Discriminatively trained dependency language modeling for conversational speech recognition
Benjamin Lambert, Bhiksha Raj, Rita Singh |
INTERSPEECH | 2 |
| 2013 | Secure binary embeddings of front-end factor analysis for privacy preserving speaker verificationabstractRemote speaker verification services typically rely on the system to have access to the users recordings, or features derived from them, and also a model of the users voice. This conventional scheme raises several privacy concerns. In this work, we address this privacy problem in the context of a speaker verification system using a factor analysis based front-end extractor, the so-called i-vectors. Speaker verification without exposing speaker data is achieved by transforming speaker i-vectors to bit strings in a way that allows the computation of approximate distances, instead of exact ones. The key to the transformation uses a hashing scheme known as Secure Binary Embeddings. Then, a modified SVM kernel permits operating on the i-vector hashes. Experiments on sub-sets of NIST SRE 2008 showed that the secure system yielded similar results as its non-private counterpart. José Portelo, Alberto Abad, Bhiksha Raj, Isabel Trancoso |
INTERSPEECH | 3 |
| 2013 | Greedy sparsity-constrained optimization
Sohail Bahmani, Bhiksha Raj, Petros Boufounos |
J. Mach. Learn. Res. | 2 |
| 2013 | Privacy-Preserving Speaker Verification and Identification Using Gaussian Mixture ModelsabstractSpeech being a unique characteristic of an individual is widely used in speaker verification and speaker identification tasks in applications such as authentication and surveillance respectively. In this article, we present frameworks for privacy-preserving speaker verification and speaker identification systems, where the system is able to perform the necessary operations without being able to observe the speech input provided by the user. In a speech-based authentication setting, this privacy constraint protect against an adversary who can break into the system and use the speech models to impersonate legitimate users. In surveillance applications, we require the system to first identify if the speech recording belongs to a suspect while preserving the privacy constraints. This prevents the system from listening in on conversations of innocent individuals. In this paper we formalize the privacy criteria for the speaker verification and speaker identification problems and construct Gaussian mixture model-based protocols. We also report experiments with a prototype implementation of the protocols on a standardized dataset for execution time and accuracy. Manas A. Pathak, Bhiksha Raj |
IEEE Trans. Speech Audio Process. | 2 |
| 2013 | Active-Set Newton Algorithm for Overcomplete Non-Negative Representations of AudioabstractThis paper proposes a computationally efficient algorithm for estimating the non-negative weights of linear combinations of the atoms of large-scale audio dictionaries, so that the generalized Kullback-Leibler divergence between an audio observation and the model is minimized. This linear model has been found useful in many audio signal processing tasks, but the existing algorithms are computationally slow when a large number of atoms is used. The proposed algorithm is based on iteratively updating a set of active atoms, with the weights updated using the Newton method and the step size estimated such that the weights remain non-negative. Algorithm convergence evaluations on representing audio spectra that are mixtures of two speakers show that with all the tested dictionary sizes the proposed method reaches a much lower value of the divergence than can be obtained by conventional algorithms, and is up to 8 times faster. A source separation evaluation revealed that when using large dictionaries, the proposed method produces a better separation quality in less time. Tuomas Virtanen, Jort F. Gemmeke, Bhiksha Raj |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | An Unsupervised Dynamic Bayesian Network Approach to Measuring Speech Style Accommodation
Mahaveer Jain, John W. McDonough, Gahgene Gweon, Bhiksha Raj, Carolyn P. Rosé |
EACL | 4 |
| 2012 | Spectrographic seam patterns for discriminative word spottingabstractThis paper presents a novel method for deriving patterns for classification of speech sounds. In contrast to conventional methods that attempt to capture time-frequency patterns as represented by spectral envelopes or peaks, our method captures patterns of high-energy tracks, or seams, of maximum “whiteness” across frequency in spectrograms. Our hypothesis is that these seams could potentially carry relatively invariant signatures of underlying sounds. We present a method to derive feature vectors from seam patterns for discriminative word spotting. We show experimentally that spectrographic seam patterns are indeed distinctive for different spoken words, and are effective for word spotting. Shubhranshu Barnwal, Kamal Sahni, Rita Singh, Bhiksha Raj |
ICASSP | 4 |
| 2012 | Audio event detection from acoustic unit occurrence patternsabstractIn most real-world audio recordings, we encounter several types of audio events. In this paper, we develop a technique for detecting signature audio events, that is based on identifying patterns of occurrences of automatically learned atomic units of sound, which we call Acoustic Unit Descriptors or AUDs. Experiments show that the methodology works as well for detection of individual events and their boundaries in complex recordings. Anurag Kumar 0003, Pranay Dighe, Rita Singh, Sourish Chaudhuri, Bhiksha Raj |
ICASSP | 5 |
| 2012 | Privacy-preserving speaker verification as password matchingabstractWe present a text-independent privacy-preserving speaker verification system that functions similar to conventional password-based authentication. Our privacy constraints require that the system does not observe the speech input provided by the user, as this can be used by an adversary to impersonate the user in the same system or elsewhere. We represent the speech input using supervectors and apply locality sensitive hashing (LSH) to transform these into bit strings, where two supervectors, and therefore inputs, are likely to be similar if they map to the same string. This transformation, therefore, reduces the problem of identifying nearest neighbors to string comparison. The users then apply a cryptographic hash function to the strings obtained from their enrollment and verification data, thereby obfuscating it from the server, who can only check if two hashed strings match without being able to reconstruct their content. We present execution time and accuracy experiments with the system on the YOHO dataset, and observe that the system achieves acceptable accuracy with minimal computational overhead needed to satisfy the privacy constraints. Manas A. Pathak, Bhiksha Raj |
ICASSP | 2 |
| 2012 | Attacking a privacy preserving music matching algorithmabstractSecure multi-party computation based techniques are often used to perform audio database search tasks, such as music matching, with privacy. However, in spite of the security of individual components of the matching schemes, the overall scheme may still not be secure. This paper explains how such flaws may occur, using a privacy preserving music matching problem as a template, and provides a solution, and analyzes the resulting tradeoff between privacy and computational complexity. Although the paper focus on a music matching application, the principles can be easily adapted to perform other tasks, such as speaker verification and keyword spotting. José Portelo, Bhiksha Raj, Isabel Trancoso |
ICASSP | 2 |
| 2012 | Exploiting Temporal Sequence Structure for Semantic Analysis of Multimedia
Sourish Chaudhuri, Rita Singh, Bhiksha Raj |
INTERSPEECH | 3 |
| 2012 | Plagiarism Detection in Polyphonic Music using Monaural Signal SeparationabstractGiven the large number of new musical tracks released each year, automated approaches to plagiarism detection are essential to help us track potential violations of copyright.Most current approaches to plagiarism detection are based on musical similarity measures, which typically ignore the issue of polyphony in music.We present a novel feature space for audio derived from compositional modelling techniques, commonly used in signal separation, that provides a mechanism to account for polyphony without incurring an inordinate amount of computational overhead.We employ this feature representation in conjunction with traditional audio feature representations in a classification framework which uses an ensemble of distance features to characterize pairs of songs as being plagiarized or not.Our experiments on a database of about 3000 musical track pairs show that the new feature space characterization produces significant improvements over standard baselines. Soham De, Indradyumna Roy, Tarunima Prabhakar, Kriti Suneja, Sourish Chaudhuri, Rita Singh, Bhiksha Raj |
INTERSPEECH | 7 |
| 2012 | Microphone Array Post-filter based on Spatially-Correlated Noise Measurements for Distant Speech RecognitionabstractThis paper presents a new microphone-array post-filtering algorithm for distant speech recognition (DSR). Conventionally, post-filtering methods assume static noise field models, and using this assumption, employ a Wiener filter mechanism for estimating the noise parameters. In contrast to this, we show how we can build the Wiener post-filter based on actual noise observations without any noise-field assumption. The algorithm is framed within a state-of-the-art beamforming technique, namely maximum negentropy (MN) beamforming with super directivity. We investigate the effectiveness of the proposed post-filter on DSR through experiments on noisy data collected in a car under different acoustic conditions. Experiments show that the new post-filtering mechanism is able to achieve up to 20 % relative reduction of word error rates (WER) under the represented noise conditions, as compared to a single distant microphone. In contrast, super-directive (SD) beamforming followed by Zelinski post-filtering achieves a relative WER reduction of only up to 11%. Other post-filters evaluated perform similarly in comparison to the proposed post-filter. Ken'ichi Kumatani, Bhiksha Raj, Rita Singh, John W. McDonough |
INTERSPEECH | 2 |
| 2012 | Privacy-Preserving Speaker Authentication
Manas A. Pathak, José Portelo, Bhiksha Raj, Isabel Trancoso |
ISC | 3 |
| 2012 | Unsupervised Structure Discovery for Semantic Analysis of AudioabstractApproaches to audio classification and retrieval tasks largely rely on detection-based discriminative models. We submit that such models make a simplistic assumption in mapping acoustics directly to semantics, whereas the actual process is likely more complex. We present a generative model that maps acoustics in a hierarchical manner to increasingly higher-level semantics. Our model has 2 layers with the first being generic sound units with no clear semantic associations, while the second layer attempts to find patterns over the generic sound units. We evaluate our model on a large-scale retrieval task from TRECVID 2011, and report significant improvements over standard baselines. Sourish Chaudhuri, Bhiksha Raj |
NIPS | 2 |
| 2012 | Optimization of the DET curve in speaker verificationabstractSpeaker verification systems are, in essence, statistical pattern detectors which can trade off false rejections for false acceptances. Any operating point characterized by a specific tradeoff between false rejections and false acceptances may be chosen. Training paradigms in speaker verification systems however either learn the parameters of the classifier employed without actually considering this tradeoff, or optimize the parameters for a particular operating point exemplified by the ratio of positive and negative training instances supplied. In this paper we investigate the optimization of training paradigms to explicitly consider the tradeoff between false rejections and false acceptances, by minimizing the area under the curve of the detection error tradeoff curve. To optimize the parameters, we explicitly minimize a mathematical characterization of the area under the detection error tradeoff curve, through generalized probabilistic descent. Experiments on the NIST 2008 database show that for clean signals the proposed optimization approach is at least as effective as conventional learning. On noisy data, verification performance obtained with the proposed approach is considerably better than that obtained with conventional learning methods. L. Paola García-Perera, Juan A. Nolazco-Flores, Bhiksha Raj, Richard M. Stern |
SLT | 3 |
| 2012 | The Markov selection model for concurrent speech recognition
Paris Smaragdis, Bhiksha Raj |
Neurocomputing | 2 |
| 2012 | Learning-Based Auditory Encoding for Robust Speech RecognitionabstractThis paper describes an approach to the optimization of the nonlinear component of a physiologically motivated feature extraction system for automatic speech recognition. Most computational models of the peripheral auditory system include a sigmoidal nonlinear function that relates the log of signal intensity to output level, which we represent by a set of frequency dependent logistic functions. The parameters of these rate-level functions are estimated to maximize the a posteriori probability of the correct class in training data. The performance of this approach was verified by the results of a series of experiments conducted with the CMU S phinx-III speech recognition system on the DARPA Resource Management, Wall Street Journal databases, and on the AURORA 2 database. In general, it was shown that feature extraction that incorporates the learned rate-nonlinearity, combined with a complementary loudness compensation function, results in better recognition accuracy in the presence of background noise than traditional MFCC feature extraction without the optimized nonlinearity when the system is trained on clean speech and tested in noise. We also describe the use of lattice structure that constraints the training process, enabling training with much more complicated acoustic models. Yu-Hsiang Bosco Chiu, Bhiksha Raj, Richard M. Stern |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Large Margin Gaussian Mixture Models with Differential PrivacyabstractAs increasing amounts of sensitive personal information is aggregated into data repositories, it has become important to develop mechanisms for processing the data without revealing information about individual data instances. The differential privacy model provides a framework for the development and theoretical analysis of such mechanisms. In this paper, we propose an algorithm for learning a discriminatively trained multiclass Gaussian mixture model-based classifier that preserves differential privacy using a large margin loss function with a perturbed regularization term. We present a theoretical upper bound on the excess risk of the classifier introduced by the perturbation. Manas A. Pathak, Bhiksha Raj |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2011 | Maximum kurtosis beamforming with a subspace filter for distant speech recognitionabstractThis paper presents a new beamforming method for distant speech recognition (DSR). The dominant mode subspace is considered in order to efficiently estimate the active weight vectors for maximum kurtosis (MK) beamforming with the generalized sidelobe canceler (GSC). We demonstrated in [1], [2], [3] that the beamforming method based on the maximum kurtosis criterion can remove reverberant and noise effects without signal cancellation encountered in the conventional beamforming algorithms. The MK beamforming algorithm, however, required a relatively large amount of data for reliably estimating the active weight vector because it relies on a numerical optimization algorithm. In order to achieve efficient estimation, we propose to cascade the subspace (eigenspace) filter [4, Section 6.8] with the active weight vector. The subspace filter can decompose the output of the blocking matrix into directional signals and ambient noise components. Then, the ambient noise components are averaged and would be subtracted from the beamformer's output, which leads to reliable estimation as well as significant computational reduction. We show the effectiveness of our method through a set of distant speech recognition experiments on real microphone array data captured in the real environment. Our new beamforming algorithm provided the best recognition performance among conventional beamforming techniques, a word error rate (WER) of 5.3%, which is comparable to the WER of 4.2% obtained with a close-talking microphone. Moreover, it achieved better recognition performance with a fewer amounts of adaptation data than the conventional MK beamformer. Ken'ichi Kumatani, John W. McDonough, Bhiksha Raj |
ASRU | 3 |
| 2011 | An iterative least-squares technique for dereverberationabstractSome recent dereverberation approaches that have been effective for automatic speech recognition (ASR) applications, model reverberation as a linear convolution operation in the spectral domain, and derive a factorization to decompose spectra of reverberated speech in to those of clean speech and room-response filter. Typically, a general non-negative matrix factorization (NMF) framework is employed for this. In this work we present an alternative to NMF and propose an iterative least-squares deconvolution technique for spectral factorization. We propose an efficient algorithm for this and experimentally demonstrate it's effectiveness in improving ASR performance. The new method results in 40-50% relative reduction in word error rates over standard baselines on artificially reverberated speech. Kshitiz Kumar, Bhiksha Raj, Rita Singh, Richard M. Stern |
ICASSP | 2 |
| 2011 | Gammatone sub-band magnitude-domain dereverberation for ASRabstractWe present an algorithm for dereverberation of speech signals for automatic speech recognition (ASR) applications. Often ASR systems are presented with speech that has been recorded in environments that include noise and reverberation. The performance of ASR systems degrades with increasing levels of noise and reverberation. While many algorithms have been proposed for robust ASR in noisy environments, reverberation is still a challenging problem. In this paper, we present ' an approach for dereverberation that models reverberation as a convolution operation in the speech spectral domain. Using a least-squares error criterion we decompose reverberated spectra into clean spectra convolved with a filter. We incorporate non-negativity and sparsity of the speech spectra as constraints within a non-negative matrix factorization (NMF) frame work to achieve the decomposition. In ASR experiments where the system is trained with unreverberated and reverberated speech, we show that the proposed approach can provide upto 40% and 19% relative reduction respectively in performance. Kshitiz Kumar, Rita Singh, Bhiksha Raj, Richard M. Stern |
ICASSP | 3 |
| 2011 | Privacy preserving probabilistic inference with Hidden Markov ModelsabstractAlice possesses a sample of private data from which she wishes to obtain some probabilistic inference. Bob possesses Hidden Markov Models (HMMs) for this purpose, but he wants the model parameters to remain private. This paper develops a framework that enables Alice and Bob to collaboratively compute the so-called forward algorithm for HMMs while satisfying their privacy constraints. This is achieved using a public-key additively homomorphic cryptosystem. Our framework is asymmetric in the sense that a larger computational overhead is incurred by Bob who has higher computational resources at his disposal, compared with Alice who has limited computing resources. Practical issues such as the encryption of probabilities and the effect of finite precision on the accuracy of probabilistic inference are considered. The protocol is implemented in software and used for secure keyword recognition. Manas A. Pathak, Shantanu Rane, Wei Sun 0008, Bhiksha Raj |
ICASSP | 4 |
| 2011 | A paired test for recognizer selection with untranscribed dataabstractTraditionally, the use of untranscribed speech has been restricted to unsupervised or semi-supervised training of acoustic models. Comparison of recognizers has required labeled data. In this paper we show how recognizers may be rank-ordered in terms of their performance using only a large quantity of untranscribed data, given a third "reference" recognizer. We develop statistical tests for comparing recognizers in this scenario. The accuracy of the reference system need not be known. Also, while the accuracy of the reference system affects the amount of data required, with enough data it only needs to perform better than chance. We show through detailed experiments that the rank ordering predicted from untranscribed data is indeed correct. Bhiksha Raj, Rita Singh, James Baker |
ICASSP | 1 |
| 2011 | Unsupervised Learning of Acoustic Unit Descriptors for Audio Content Representation and Classification
Sourish Chaudhuri, Mark Harvilla, Bhiksha Raj |
INTERSPEECH | 3 |
| 2011 | A Paradigm for Limited Vocabulary Speech Recognition Based on Redundant Spectro-Temporal Feature Sets
Sourish Chaudhuri, Bhiksha Raj, Tony Ezzat |
INTERSPEECH | 2 |
| 2011 | Privacy Preserving Speaker Verification Using Adapted GMMsabstractIn this paper we present an adapted UBM-GMM based privacy preserving speaker verification (PPSV) system, where the system is not able to observe the speech data provided by the user and the user does not observe the models trained by the system. These privacy criteria are important in order to prevent an adversary having unauthorized access to the user’s client device from impersonating a user and also from another adversary who can break into the verification system can learn about the user’s speech patterns to impersonate the user in another system. We present protocols for speaker enrollment and verification which preserve privacy according to these requirements and report experiments with a prototype implementation on the YOHO dataset. Index Terms: speaker verification, secure biometrics 1. Manas A. Pathak, Bhiksha Raj |
INTERSPEECH | 2 |
| 2011 | Phoneme-Dependent NMF for Speech Enhancement in Monaural MixturesabstractThe problem of separating speech signals out of monaural mix-tures (with other non-speech or speech signals) has become in-creasingly popular in recent times. Among the various solutions proposed, the most popular methods are based on compositional models such as non-negative matrix factorization (NMF) and latent variable models. Although these techniques are highly effective they largely ignore the inherently phonetic nature of speech. In this paper we present a phoneme-dependent NMF-based algorithm to separate speech from monaural mixtures. Experiments performed on speech mixed with music indicate that the proposed algorithm can result in significant improve-ment in separation performance, over conventional NMF-based separation. Index terms: Monaural signal separation, speech enhance-ment, restoration, Non-negative matrix factorization. 1. Bhiksha Raj, Rita Singh, Tuomas Virtanen |
INTERSPEECH | 1 |
| 2011 | A Comparison of Latent Variable Models For Conversation Analysis
Sourish Chaudhuri, Bhiksha Raj |
SIGDIAL Conference | 2 |
| 2011 | Preface
Martin Heckmann, Bhiksha Raj, Paris Smaragdis |
Speech Commun. | 2 |
| 2010 | A hybrid physical and statistical dynamic articulatory framework incorporating analysis-by-synthesis for improved phone classificationabstractIn this paper, we present a dynamic articulatory model for phone classification. The model integrates real articulatory information derived from ElectroMagnetic Articulograph (EMA) data into its inner states. It maps from the articulatory space to the acoustic one using an adapted vocal tract model for each speaker and a physiologically-motivated articulatory synthesis approach. We apply the analysis-by-synthesis paradigm in a statistical fashion. We first present a fast approach for deriving analysis-by-synthesis distortion features. Next, the distortion between the speech synthesized from the articulatory states and the incoming speech signal is used to compute the output observation probabilities of the Hidden Markov Model (HMM) used for classification. Experiments with the novel framework show improvements over baseline in phone classification accuracy. Ziad Al Bawab, Bhiksha Raj, Richard M. Stern |
ICASSP | 2 |
| 2010 | Learning-based auditory encoding for robust speech recognitionabstractThis paper describes ways of speeding up the optimization process for learning physiologically-motivated components of a feature computation module directly from data. During training, word lattices generated by the speech decoder and conjugate gradient descent were included to train the parameters of logistic functions in a fashion that maximizes the a posteriori probability of the correct class in the training data. These functions represent the rate-level nonlinearities found in most mammalian auditory systems. Experiments conducted using the CMU SPHINX-III system on the DARPA Resource Management and Wall Street Journal tasks show that the use of discriminative training to estimate the shape of the rate-level nonlinearity provides better recognition accuracy in the presence of background noise than traditional procedures which do not employ learning. More importantly, the inclusion of conjugate gradient descent optimization and a word lattice to reduce the number of hypotheses considered greatly increases the training speed, which makes training with much more complicated models possible. Yu-Hsiang Bosco Chiu, Bhiksha Raj, Richard M. Stern |
ICASSP | 2 |
| 2010 | Latent-variable decomposition based dereverberation of monaural and multi-channel signalsabstractWe present an algorithm to dereverberate single- and multi-channel audio recordings. The proposed algorithm models the magnitude spectrograms of clean audio signals as histograms drawn from a multinomial process. Spectrograms of reverberated signals are obtained as histograms of draws from the PDF of the sum of two random variables, one representing the spectrogram of clean speech and the second the frequency decomposition of the room response. The spectrogram of the clean signal is computed as a maximum-likelihood estimate from the spectrogram of reverberant speech using an EM algorithm. Experimental evaluations show that the proposed algorithm is able to greatly reduce the reverberation effects in even highly reverberant signals captured in auditoria and other open spaces. Rita Singh, Bhiksha Raj, Paris Smaragdis |
ICASSP | 2 |
| 2010 | Ultrasonic sensing for robust speech recognitionabstractIn this paper, we present our work using ultrasonic sensing of speech for digit recognition. First, a set of spectral ultrasonic features are developed and tuned in order to achieve optimal performance for the digit recognition task. Using these features, we demonstrate an overall accuracy of 33.00% on a digit recognition task using HMMs with recordings from 6 speakers. The results indicate that ultrasonic sensing of speech is viable, but that further work is needed to achieve word accuracies that match those of audio. Finally, experimental results are presented which demonstrate that fusing information from ultrasound and audio sources show marginal improvements over audio-only performances. Sundararajan Srinivasan, Bhiksha Raj, Tony Ezzat |
ICASSP | 2 |
| 2010 | Synthesizing speech from Doppler signalsabstractIt has long been considered a desirable goal to be able to construct an intelligible speech signal merely by observing the talker in the act of speaking. Past methods at performing this have been based on camera-based observations of the talker's face, combined with statistical methods that infer the speech signal from the facial motion captured by the camera. Other methods have included synthesis of speech from measurements taken by electro-myelo graphs and other devices that are tethered to the talker - an undesirable setup. In this paper we present a new device for synthesizing speech from characterizations of facial motion associated with speech - a Doppler sonar. Facial movement is characterized through Doppler frequency shifts in a tone that is incident on the talker's face. These frequency shifts are used to infer the underlying speech signal. The setup is farfield and untethered, with the sonar acting from the distance of a regular desktop microphone. Preliminary experimental evaluations show that the mechanism is very promising - we are able to synthesize reasonable speech signals, comparable to those obtained from tethered devices such as EMGs. Arthur R. Toth, Kaustubh Kalgaonkar, Bhiksha Raj, Tony Ezzat |
ICASSP | 3 |
| 2010 | Spectrogram dimensionality reductionwith independence constraintsabstractWe present an algorithm to find a low-dimensional decomposition of a spectrogram by formulating this as a regularized non-negative matrix factorization (NMF) problem with a regularization term chosen to encourage independence. This algorithm provides a better decomposition than standard NMF when the underlying sources are independent. It is directly applicable to non-square matrices, and it makes better use of additional observation streams than previous nonnegative ICA algorithms. Kevin W. Wilson, Bhiksha Raj |
ICASSP | 2 |
| 2010 | Creating a linguistic plausibility dataset with non-expert annotatorsabstractWe describe the creation of a linguistic plausibility dataset that contains annotated examples of language judged to be linguistically plausible, implausible, and every-thing in between. To create the dataset we randomly generate sentences and have them annotated by crowd sourcing over the Amazon Mechanical Turk. Obtaining inter-annotator agreement is a difficult problem because linguistic plausibility is highly subjective. The annotations obtained depend, among other factors, on the manner in which annotators are questioned about the plausibility of sentences. We describe our experiments on posing a number of different questions to the annotators, in order to elicit the responses with greatest agreement, and present several methods for analyzing the resulting responses. The generated dataset and annotations are being made available to public. Benjamin Lambert, Rita Singh, Bhiksha Raj |
INTERSPEECH | 3 |
| 2010 | Non-negative matrix factorization based compensation of music for automatic speech recognitionabstractThis paper proposes to use non-negative matrix factorization based speech enhancement in robust automatic recognition of mixtures of speech and music. We represent magnitude spectra of noisy speech signals as the non-negative weighted linear combination of speech and noise spectral basis vectors, that are obtained from training corpora of speech and music. We use overcomplete dic-tionaries consisting of random exemplars of the training data. The method is tested on theWall Street Journal large vocabulary speech corpus which is artificially corrupted with polyphonic music from the RWC music database. Various music styles and speech-to-music ratios are evaluated. The proposed methods are shown to produce a consistent, significant improvement on the recog-nition performance in the comparison with the baseline method. Audio demonstrations of the enhanced signals are available at Bhiksha Raj, Tuomas Virtanen, Sourish Chaudhuri, Rita Singh |
INTERSPEECH | 1 |
| 2010 | Ungrounded independent non-negative factor analysisabstractWe describe an algorithm that performs regularized non-negative matrix factorization (NMF) to find independent components in nonnegative data. Previous techniques proposed for this purpose require the data to be grounded, with support that goes down to 0 along each dimension. In our work, this requirement is eliminated. Based on it, we present a technique to find a low-dimensional decomposition of spectrograms by casting it as a problem of discovering independent non-negative components from it. The algorithm itself is implemented as regularized non-negative matrix factorization (NMF). Unlike other ICA algorithms, this algorithm computes the mixing matrix rather than an unmixing matrix. This algorithm provides a better decomposition than standard NMF when the underlying sources are independent. It makes better use of additional observation streams than previous nonnegative ICA algorithms. Index Terms — matrix decomposition, ICA 1. Bhiksha Raj, Kevin W. Wilson, Alexander Krueger, Reinhold Häb-Umbach |
INTERSPEECH | 1 |
| 2010 | The use of sense in unsupervised training of acoustic models for ASR systemsabstractIn unsupervised training of ASR systems, no annotated data are assumed to exist. Word-level annotations for training audio are generated iteratively using an ASR system. At each iteration a subset of data judged as having the most reliable transcriptions is selected to train the next set of acoustic models. Data selection however remains a difficult problem, particularly when the error rate of the recognizer providing the initial annotation is very high. In this paper we propose an iterative algorithm that uses a combination of likelihoods and a simple model of sense to select data. We show that the algorithm is effective for unsupervised training of acoustic models, particularly when the initial annotation is highly erroneous. Experiments conducted on Fisher-1 data using initial models from Switchboard, and a vocabulary and LM derived from the Google N-grams, show that performance on a selected held-out test set from Fisher data improves more with iterations relative to likelihood-based data selection. Rita Singh, Benjamin Lambert, Bhiksha Raj |
INTERSPEECH | 3 |
| 2010 | Multiparty Differential Privacy via Aggregation of Locally Trained ClassifiersabstractAs increasing amounts of sensitive personal information finds its way into data repositories, it is important to develop analysis mechanisms that can derive aggregate information from these repositories without revealing information about individual data instances. Though the differential privacy model provides a framework to analyze such mechanisms for databases belonging to a single party, this framework has not yet been considered in a multi-party setting. In this paper, we propose a privacy-preserving protocol for composing a differentially private aggregate classifier using classifiers trained locally by separate mutually untrusting parties. The protocol allows these parties to interact with an untrusted curator to construct additive shares of a perturbed aggregate classifier. We also present a detailed theoretical analysis containing a proof of differential privacy of the perturbed aggregate classifier and a bound on the excess risk introduced by the perturbation. We verify the bound with an experimental evaluation on a real dataset. Manas A. Pathak, Shantanu Rane, Bhiksha Raj |
NIPS | 3 |
| 2009 | Word Particles Applied to Information Retrieval
Evandro B. Gouvêa, Bhiksha Raj |
ECIR | 2 |
| 2009 | A joint decoding algorithm for multiple-example-based addition of words to a pronunciation lexiconabstractWe propose an algorithm that enables joint Viterbi decoding of multiple independent audio recordings of a word to derive its pronunciation. Experiments show that this method results in better pronunciation estimation and word recognition accuracy than that obtained either with a single example of the word or using conventional approaches to pronunciation estimation using multiple examples. Dhananjay Bansal, Nishanth Ulhas Nair, Rita Singh, Bhiksha Raj |
ICASSP | 4 |
| 2009 | One-handed gesture recognition using ultrasonic Doppler sonarabstractThis paper presents a new device based on ultrasonic sensors to recognize one-handed gestures. The device uses three ultrasonic receivers and a single transmitter. Gestures are characterized through the Doppler frequency shifts they generate in reflections of an ultrasonic tone emitted by the transmitter. We show that this setup can be used to classify simple one-handed gestures with high accuracy. The ultrasonic doppler based device is very inexpensive - $20 USD for the whole setup including the acquisition system, and computationally efficient as compared to most traditional devices (e.g. video). These gestures, could potentially be used to control and drive a device. Kaustubh Kalgaonkar, Bhiksha Raj |
ICASSP | 2 |
| 2009 | Deriving vocal tract shapes from electromagnetic articulograph data via geometric adaptation and matchingabstractIn this paper, we present our efforts towards deriving vocal tract shapes from ElectroMagnetic Articulograph data (EMA) via geometric adaptation and matching. We describe a novel approach for adapting Maeda’s geometric model of the vocal tract to one speaker in the MOCHA database. We show how we can rely solely on the EMA data for adaptation. We present our search technique for the vocal tract shapes that best fit the given EMA data. We then describe our approach of synthesizing speech from these shapes. Results on Mel-cepstral distortion reflect improvement in synthesis over the approach we used before without adaptation. Index Terms: MOCHA EMA data, Maeda Model, vocal tract adaptation, articulatory model fitting, articulatory synthesis Ziad Al Bawab, Lorenzo Turicchia, Richard M. Stern, Bhiksha Raj |
INTERSPEECH | 4 |
| 2009 | Towards fusion of feature extraction and acoustic model training: a top down process for robust speech recognitionabstractThis paper presents a strategy to learn physiologicallymotivated components in a feature computation module discriminatively, directly from data, in a manner that is inspired by the presence of efferent processes in the human auditory system. In our model a set of logistic functions which represent the rate-level nonlinearities found in most mammal hearing system are put in as part of the feature extraction process. The parameters of these rate-level functions are estimated to maximize the a posteriori probability of the correct class in the training data. The estimated feature computation is observed to be robust against environmental noise. Experiments conducted with the CMU Sphinx-III on the DARPA Resource Management task show that the discriminatively estimated rate-nonlinearity results in better performance in the presence of background noise than traditional procedures which separate the feature extraction and model training into two distinct parts without feed back from the latter to the former. Index Terms: automatic speech recognition, discriminative training, auditory model, data analysis Yu-Hsiang Bosco Chiu, Bhiksha Raj, Richard M. Stern |
INTERSPEECH | 2 |
| 2009 | Signal separation for robust speech recognition based on phase difference information obtained in the frequency domainabstractIn this paper, we present a new two-microphone approach that improves speech recognition accuracy when speech is masked by other speech. The algorithm improves on previous systems that have been successful in separating signals based on differences in arrival time of signal components from two microphones. The present algorithm differs from these efforts in that the signal selection takes place in the frequency domain. We observe that additional smoothing of the phase estimates over time and frequency is needed to support adequate speech recognition performance. We demonstrate that the algorithm described in this paper provides better recognition accuracy than timedomain-based signal separation algorithms, and at less than 10 percent of the computation cost. Index Terms: Robust speech recognition, signal separation, time delay analysis, phase difference analysis Chanwoo Kim 0001, Kshitiz Kumar, Bhiksha Raj, Richard M. Stern |
INTERSPEECH | 3 |
| 2009 | A Sparse Non-Parametric Approach for Single Channel Separation of Known SoundsabstractIn this paper we present an algorithm for separating mixed sounds from a monophonic recording. Our approach makes use of training data which allows us to learn representations of the types of sounds that compose the mixture. In contrast to popular methods that attempt to extract com- pact generalizable models for each sound from training data, we employ the training data itself as a representation of the sources in the mixture. We show that mixtures of known sounds can be described as sparse com- binations of the training data itself, and in doing so produce significantly better separation results as compared to similar systems based on compact statistical models. Paris Smaragdis, Madhusudana V. S. Shashanka, Bhiksha Raj |
NIPS | 3 |
| 2008 | Recognizing talking faces from acoustic Doppler reflectionsabstractFace recognition algorithms typically deal with the classification of static images of faces that are obtained using a camera. In this paper we propose a new sensing mechanism based on the Doppler effect to capture the patterns of motion of talking faces. We incident an ultrasonic tone on subjects' faces and capture the reflected signal. When the subject talks, different parts of their face move with different velocities in a characteristic manner. Each of these velocities imparts a different Doppler shift to the reflected ultrasonic signal. Thus, the set of frequencies in the reflected ultrasonic signal is characteristic of the subject. We show that even using a simple feature computation scheme to characterize the spectrum of the reflected signal, and a simple GMM based Bayesian classifier, we are able to recognize talkers with an accuracy of over 90%. Interestingly, we are also able to identify the gender of the talker with an accuracy of over 90%. Kaustubh Kalgaonkar, Bhiksha Raj |
FG | 2 |
| 2008 | Analysis-by-synthesis features for speech recognitionabstractWe present a framework for speech recognition that accounts for hidden articulatory information. We model the articulatory space using a codebook of articulatory configurations geometrically derived from EMA measurements available in the MOCHA database. The articulatory parameter set we derive is in the form of Maeda parameters. In turn, these parameters are used in a physiologically motivated articulatory speech synthesizer based on the model by Sondhi and Schroeter. We use the distortion between the speech synthesized from each of the articulatory configurations and the original speech as features for recognition. We setup a segmented phoneme recognition task on the MOCHA database using Gaussian mixture models (GMMs). Improvements are achieved when combining the probability scores generated using the distortion features with the scores using acoustic features. Ziad Al Bawab, Bhiksha Raj, Richard M. Stern |
ICASSP | 2 |
| 2008 | Ultrasonic Doppler sensor for speaker recognitionabstractIn this paper we present a novel use of an acoustic Doppler sonar for multi-modal speaker identification. An ultrasonic emitter directs a 40 kHz tone toward the speaker. Reflections from the speaker's face are recorded as the speaker talks. The frequency of the tone is modified by the velocity of the facial structures it is reflected by. The received ultrasonic signal thus contains an entire spectrum of frequencies representing the set of all velocities of facial components. The pattern of frequencies in the reflected signal is observed to be typical of the speaker. The captured ultrasonic signal is synchronously analyzed with the corresponding voice signal to extract specific characteristics that can be used to identify the speaker. Experiments show that the information this can result in significant improvements in speaker identification accuracy both under clean conditions and in noise. Kaustubh Kalgaonkar, Bhiksha Raj |
ICASSP | 2 |
| 2008 | Sparse and shift-invariant feature extraction from non-negative dataabstractIn this paper we describe a technique that allows the extraction of multiple local shift-invariant features from analysis of non-negative data of arbitrary dimensionality. Our approach employs a probabilistic latent variable model with sparsity constraints. We demonstrate its utility by performing feature extraction in a variety of domains ranging from audio to images and video. Paris Smaragdis, Bhiksha Raj, Madhusudana V. S. Shashanka |
ICASSP | 2 |
| 2008 | Speech denoising using nonnegative matrix factorization with priorsabstractWe present a technique for denoising speech using nonnegative matrix factorization (NMF) in combination with statistical speech and noise models. We compare our new technique to standard NMF and to a state-of-the-art Wiener filter implementation and show improvements in speech quality across a range of interfering noise types. Kevin W. Wilson, Bhiksha Raj, Paris Smaragdis, Ajay Divakaran |
ICASSP | 2 |
| 2008 | Regularized non-negative matrix factorization with temporal dependencies for speech denoisingabstractWe present a tecchnique for denoising speech using temporally regularized nonnegative matrix factorization (NMF). In previ-ous work [1], we used a regularized NMF update to impose structure within each audio frame. In this paper, we add frame-to-frame regularization across time and show that this additional regularization can also improve our speech denoising results. We evaluate our algorithm on a range of nonstationary noise types and outperform a state-of-the-art Wiener filter implemen-tation. Index Terms: speech enhancement, source separation, speech modeling, speech processing Kevin W. Wilson, Bhiksha Raj, Paris Smaragdis |
INTERSPEECH | 2 |
| 2007 | Acoustic Doppler sonar for gait recoginationabstractA person's gait is a characteristic that might be employed to identify him/her automatically. Conventionally, automatic for gait-based identification of subjects employ video and image processing to characterize gait. In this paper we present an Acoustic Doppler Sensor(ADS) based technique for the characterization of gait. The ADS is very inexpensive sensor that can be built using off-the-shelf components, for under $20 USD at today's prices. We show that remarkably good gait recognition is possible with the ADS sensor. Kaustubh Kalgaonkar, Bhiksha Raj |
AVSS | 2 |
| 2007 | Sensor and Data Systems, Audio-Assisted Cameras and Acoustic Doppler SensorsabstractIn this chapter we present two technologies for sensing and surveillance -audio-assisted cameras and acoustic Doppler sensors for gait recognition. Kaustubh Kalgaonkar, Paris Smaragdis, Bhiksha Raj |
CVPR | 3 |
| 2007 | Bandwidth Expansionwith a pólya URN ModelabstractWe present a new statistical technique for the estimation of the high frequency components (4-8 kHz) of speech signals from narrow-band (0-4 kHz) signals. The magnitude spectra of broadband speech are modelled as the outcome of a Polya Urn process, that represents the spectra as the histogram of the outcome of several draws from a mixture multinomial distribution over frequency indices. The multinomial distributions that compose this process are learnt from a corpus of broadband (0-8 kHz) speech. To estimate high-frequency components of narrow-band speech, its spectra are also modelled as the outcome of draws from a mixture-multinomial process that is composed of the learnt multinomials, where the counts of the indices of higher frequencies have been obscured. The obscured high-frequency components are then estimated as the expected number of draws of their indices from the mixture-multinomial. Experiments conducted on bandlimited signals derived from the WSJ corpus show that the proposed procedure is able to accurately estimate the high frequency components of these signals. Bhiksha Raj, Rita Singh, Madhusudana V. S. Shashanka, Paris Smaragdis |
ICASSP (4) | 1 |
| 2007 | Sparse Overcomplete Decomposition for Single Channel Speaker SeparationabstractWe present an algorithm for separating multiple speakers from a mixed single channel recording. The algorithm is based on a model proposed by Raj and Smaragdis (2005). The idea is to extract certain characteristic spectra-temporal basis functions from training data for individual speakers and decompose the mixed signals as linear combinations of these learned bases. In other words, their model extracts a compact code of basis functions that can explain the space spanned by spectral vectors of a speaker. In our model, we generate a sparse-distributed code where we have more basis functions than the dimensionality of the space. We propose a probabilistic framework to achieve sparsity. Experiments show that the resulting sparse code better captures the structure in data and hence leads to better separation. Madhusudana V. S. Shashanka, Bhiksha Raj, Paris Smaragdis |
ICASSP (2) | 2 |
| 2007 | Probabilistic deduction of symbol mappings for extension of lexiconsabstractThis paper proposes a statistical mapping-based technique for guessing pronunciations of novel words from their spellings. The technique is based on the automatic determination and utilization of unidirectional mappings between n-tuples of characters and n-tuples of phonemes, and may be viewed as a statistical extension of analogy-based pronunciation guessing algorithms. 1. Rita Singh, Evandro B. Gouvêa, Bhiksha Raj |
INTERSPEECH | 3 |
| 2007 | Sparse Overcomplete Latent Variable Decomposition of Counts DataabstractAn important problem in many fields is the analysis of counts data to extract meaningful latent components. Methods like Probabilistic Latent Semantic Analysis (PLSA) and Latent Dirichlet Allocation (LDA) have been proposed for this purpose. However, they are limited in the number of components they can extract and also do not have a provision to control the expressiveness" of the extracted components. In this paper, we present a learning formulation to address these limitations by employing the notion of sparsity. We start with the PLSA framework and use an entropic prior in a maximum a posteriori formulation to enforce sparsity. We show that this allows the extraction of overcomplete sets of latent components which better characterize the data. We present experimental evidence of the utility of such representations." Madhusudana V. S. Shashanka, Bhiksha Raj, Paris Smaragdis |
NIPS | 2 |
| 2007 | Ultrasonic Doppler Sensor for Voice Activity DetectionabstractThis letter describes a robust voice activity detector using an ultrasonic Doppler sonar device. An ultrasonic beam is incident on the talker's face. Facial movements result in Doppler frequency shifts in the reflected signal that are sensed by an ultrasonic sensor. Speech-related facial movements result in identifiable patterns in the spectrum of the received signal that can be used to identify speech activity. These sensors are not affected by even high levels of ambient audio noise. Unlike most other non-acoustic sensors, the device need not be taped to a talker. A simple yet robust method of extracting the voice activity information from the ultrasonic Doppler signal is developed and presented in this letter. The algorithm is seen to be very effective and robust to noise, and it can be implemented in real time. Kaustubh Kalgaonkar, Rongquiang Hu, Bhiksha Raj |
IEEE Signal Process. Lett. | 3 |
| 2007 | Soft Mask Methods for Single-Channel Speaker SeparationabstractThe problem of single-channel speaker separation attempts to extract a speech signal uttered by the speaker of interest from a signal containing a mixture of acoustic signals. Most algorithms that deal with this problem are based on masking, wherein unreliable frequency components from the mixed signal spectrogram are suppressed, and the reliable components are inverted to obtain the speech signal from speaker of interest. Most current techniques estimate this mask in a binary fashion, resulting in a hard mask. In this paper, we present two techniques to separate out the speech signal of the speaker of interest from a mixture of speech signals. One technique estimates all the spectral components of the desired speaker. The second technique estimates a soft mask that weights the frequency subbands of the mixed signal. In both cases, the speech signal of the speaker of interest is reconstructed from the complete spectral descriptions obtained. In their native form, these algorithms are computationally expensive. We also present fast factored approximations to the algorithms. Experiments reveal that the proposed algorithms can result in significant enhancement of individual speakers in mixed recordings, consistently achieving better performance than that obtained with hard binary masks. Aarthi M. Reddy, Bhiksha Raj |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Latent Dirichlet Decomposition for Single Channel Speaker SeparationabstractWe present an algorithm for the separation of multiple speakers from mixed single-channel recordings by latent variable decomposition of the speech spectrogram. We model each magnitude spectral vector in the short-time Fourier transform of a speech signal as the outcome of a discrete random process that generates frequency bin indices. The distribution of the process is modeled as a mixture of multinomial distributions, such that the mixture weights of the component multinomials vary from analysis window to analysis window. The component multinomials are assumed to be speaker specific and are learned from training signals for each speaker. We model the prior distribution of the mixture weights for each speaker as a Dirichlet distribution. The distributions representing magnitude spectral vectors for the mixed signal are decomposed into mixtures of the multinomials for all component speakers. The frequency distribution, i.e the spectrum for each speaker, is reconstructed from this decomposition Bhiksha Raj, Madhusudana V. S. Shashanka, Paris Smaragdis |
ICASSP (5) | 1 |
| 2006 | An integrated approach to improve speech recognition rate for non-native speakers
Yunbin Deng, Xiaokun Li, Chiman Kwan, Roger Xu, Bhiksha Raj, Richard M. Stern, David Williamson |
INTERSPEECH | 5 |
| 2006 | An acoustic Doppler-Based Front End for Hands Free spoken User InterfacesabstractTwo major problems facing hands-free spoken user interfaces are first, to be able to identify accurately when it is addressed and second to effectively determine the start and end points of utterances so that proper sections of speech can be used for further processing and recognition tasks. Accurate determination of end points not only aids accurate noise estimation but also improves speech recognition accuracy. This paper presents an acoustic Doppler sensor based front end for a hands-free spoken user device which provides an efficient and effective solution to both the aforementioned problems. Kaustubh Kalgaonkar, Bhiksha Raj |
SLT | 2 |
| 2005 | A Companding Front End for Noise-Robust Automatic Speech RecognitionabstractFeature computation modules for automatic speech recognition (ASR) systems have long been modeled on the human auditory system. Most current ASR systems model the critical band response and equal loudness characteristics of the auditory system. It has been postulated that more detailed models of the human auditory system can lead to more noise-robust speech recognition. An auditory phenomenon that is of particular relevance to robustness is simultaneous masking, whereby dominant frequencies suppress adjacent weaker frequencies. In this paper, we present a companding-based model that mimics simultaneous masking in the front end of a speech recognizer. In an automotive digits recognition task, the front end improves word error rate by 4.0% (25% relative to Mel cepstra) at -5 dB SNR at the cost of a 1.7% increase at 15 dB SNR. Jethran Guinness, Bhiksha Raj, Bent Schmidt-Nielsen, Lorenzo Turicchia, Rahul Sarpeshkar |
ICASSP (1) | 2 |
| 2005 | A Comparison Between Spoken Queries and Menu-Based Interfaces for In-car Digital Music Selection
Clifton Forlines, Bent Schmidt-Nielsen, Bhiksha Raj, Kent Wittenburg |
INTERACT | 3 |
| 2005 | Bandwidth expansion of narrowband speech using non-negative matrix factorizationabstractIn this paper, we present a novel technique for the estimation of the high frequency components (4-8kHz) of speech signals from narrow-band (0-4 kHz) signals using convolutive Non-Negative Matrix Factorisation (NMF). The proposed technique utilizes a brief recording of simultaneous broad band and narrow band signals from a target speaker to learn a set of broad-band non-negative bases for the speaker. The low-frequency components of these bases are used to determine how the high-frequency components must be combined in order to reconstruct the high-frequency components of new narrow-band signals from the speaker. Experiments reveal that the technique is able to reconstruct broadband sppech that is perceptually virtually indistinguishable from true broadband recordings. Dhananjay Bansal, Bhiksha Raj, Paris Smaragdis |
INTERSPEECH | 2 |
| 2005 | Recognizing speech from simultaneous speakersabstractIn this paper we present and evaluate factored methods for recognition of simultaneous speech from multiple speakers in single-channel recordings. Factored methods decompose the problem of jointly recognizing the speech from each of the speakers by separately recognizing the speech from each speaker. In order to achieve this, the signal components of the target speaker in each case must be enhanced in some manner. We do this in two ways: using an NMF-based speaker separation algorithm that generates separated spectra for each speaker, and a mask estimation method that generates spectral masks for each speaker that must be used in conjunction with a missing-feature method that can recognize speech from partial spectral data. Experiments on synthetic mixtures of signals from the Wall Street Journal corpus show that both approaches can greatly improve the recognition of the individual signals in the mixture. Bhiksha Raj, Rita Singh, Paris Smaragdis |
INTERSPEECH | 1 |
| 2004 | On tracking noise with linear dynamical system modelsabstractThis paper investigates the use of higher-order autoregressive vector predictors for tracking the noise in noisy speech signals. The autoregressive predictors form the state equation of a linear dynamical system that models the spectral dynamics of the noise process. Experiments show that the use of such models to track noise can lead to large gains in recognition performance on speech compensated for the estimated noise. However, predictors of order greater than 1 are not observed to improve the performance beyond that obtained with a first-order predictor. We analyze and explain why this is so. Bhiksha Raj, Rita Singh, Richard M. Stern |
ICASSP (1) | 1 |
| 2004 | A minimum mean squared error estimator for single channel speaker separationabstractThe problem of separating out the signals for multiple speakers from a single mixed recording has received considerable atten- tio ni n recent times. Most current techniques are based on the principle of masking :i n order the separate out the signal for any speaker, frequency components that are not believed to be- long to that speaker are suppressed. The signals for the speaker is reconstructed fro mt hepartial spectral information that re- mains. In this paper we present a different kind of technique - one that attempts to estimate all spectral components for the desired speaker. Separated signals are derived from the com- plete spectral descriptions so obtained. Experiments show that this method results in superior reconstruction to masking based methods. form representations of the various speakers by hidden Markov models (HMMs). The parameters of the HMM for any speaker are learnt from training data recorded from the speaker. In addi- tion, Roweis assumes that the log energy in any frequency band of the mixed signal at any time can be attributed to only one of the speakers. This log-max assumption is justified by two observations. First, when two or more speakers speak simulta- neously, at any time, any given frequency band is usually domi- nated by a single speaker. Second, in any given frequency band the disparity in the energy levels of the dominant speaker and the other speakers is such that the logarithm of the sum of the energies of the individual speakers can be well approximated by the logarithm of the energy of the dominant speaker. In or- der to reconstruct the signal for any speaker, Roweis estimates the mask for that speaker, i.e. the identity of the time-frequency locations where the speaker dominates. The entire signal is re- constructed entirely from the masked spectrum for the speaker, i.e. fro mt he spectral components identified by the mask. The results achieved with this method are remarkably good. Hershey et. al. (6) augment audio recordings with visual features, such as lip and facial movement, in order to enhance the separation. Additionally, the ys eparate the signal into mul- tiple frequency bands, which are then processed independently. As in Roweis' algorithm, the signals for the individual speakers are reconstructed from masked spectra. Aarthi M. Reddy, Bhiksha Raj |
INTERSPEECH | 2 |
| 2004 | Spokenquery: an alternate approach to chosing items with speechabstractA majority of spoken user interfaces deal with the task of retrieving an element from a list. Conventionally, spoken UIs deal with such tasks through hierarchies of menus or dialogs, that navigate users through a series of steps, each of which present them with a limited set of choices. In a recent paper [2] we presented an alternative approach to such UIs, termed SpokenQuery, that recasts the problem of selection from lists as one of retrieval, and demonstrated that it could result in significantly lowered cognitive load on the user. In this paper, we examine varius aspects of retrieval from spoken queries, and UIs based on such retrieval, and demonstrate that in addition to reducing the cognitive load on the user, the system is effective for searching large databases, is robust to environment noise, and is effective as a UI. Joseph Woelfel, Jan C. van Gemert, Bhiksha Raj, David Wong 0006 |
INTERSPEECH | 4 |
| 2004 | Reconstruction of missing features for robust speech recognition
Bhiksha Raj, Michael L. Seltzer, Richard M. Stern |
Speech Commun. | 1 |
| 2004 | A Bayesian classifier for spectrographic mask estimation for missing feature speech recognition
Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
Speech Commun. | 2 |
| 2004 | Likelihood-maximizing beamforming for robust hands-free speech recognitionabstractSpeech recognition performance degrades significantly in distant-talking environments, where the speech signals can be severely distorted by additive noise and reverberation. In such environments, the use of microphone arrays has been proposed as a means of improving the quality of captured speech signals. Currently, microphone-array-based speech recognition is performed in two independent stages: array processing and then recognition. Array processing algorithms, designed for signal enhancement, are applied in order to reduce the distortion in the speech waveform prior to feature extraction and recognition. This approach assumes that improving the quality of the speech waveform will necessarily result in improved recognition performance and ignores the manner in which speech recognition systems operate. In this paper a new approach to microphone-array processing is proposed in which the goal of the array processing is not to generate an enhanced output waveform but rather to generate a sequence of features which maximizes the likelihood of generating the correct hypothesis. In this approach, called likelihood-maximizing beamforming, information from the speech recognition system itself is used to optimize a filter-and-sum beamformer. Speech recognition experiments performed in a real distant-talking environment confirm the efficacy of the proposed approach. Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
IEEE Trans. Speech Audio Process. | 2 |
| 2003 | Multi-channel source separation by factorial HMMsabstractWe present a new speaker-separation algorithm for separating signals with known statistical characteristics from mixed multi-channel recordings. Speaker separation has conventionally been treated as a problem of blind source separation (BSS). This approach does not utilize any knowledge of the statistical characteristics of the signals to be separated, relying mainly on the independence between the various signals to separate them. We present an algorithm that utilizes detailed statistical information about the signals to be separated, represented in the form of hidden Markov models (HMM). We treat the signal separation problem as one of beamforming, where each signal is extracted using a filter-and-sum array. The filters are estimated to maximize the likelihood of the summed output, measured on the HMM for the desired signal. This is done by iteratively estimating the best state sequence through the HMM from a factorial HMM (FHMM) that is the cross-product of the HMMs for the multiple signals, using the current output of the array, and estimating the filters to maximize the likelihood of that state sequence. Experiments show that the proposed method can cleanly extract a background speaker who is 20 dB below the foreground speaker in a two-speaker mixture, when the HMMs for the signals are constructed from knowledge of the utterance transcriptions. Manuel Reyes-Gomez, Bhiksha Raj, Daniel P. W. Ellis |
ICASSP (1) | 2 |
| 2003 | Lossless compression of language model structure and word identifiersabstractVery large reductions in language model memory requirements have recently been reported for large vocabulary continuous speech recognition applications through the pruning and quantization of the floating-point components of the language model: the probabilities and back-off weights. In this paper that work is extended through the compression of the integer components: the word identifiers and storage structures. A novel algorithm is presented for converting ordered lists of monotonically increasing integer values (such as are commonly found in language models) into variable-bit width tree structures such that the most memory efficient configuration is obtained for each original list. By applying this new technique together with the techniques reported previously we obtain an 86% reduction in language model size to 10Mb for no increase in word error rate on the DARPA Hub4 1998 task and a 0.5% absolute increase on the Hub4 1997 task. Bhiksha Raj, Edward W. D. Whittaker |
ICASSP (1) | 1 |
| 2003 | Tracking noise via dynamical systems with a continuum of statesabstractWe model noise as a sequence of states of a dynamical system with a continuum of states. Observations generated by such a system are assumed to be related to the state of the system by a functional relation which models clean speech as the corrupting influence on noise. We show how the closed-form representation of such a dynamical system can be rendered tractable and solved iteratively by dynamically sampling the state space, resulting in an estimated noise sequence (sequence of states), which can then be removed from the noisy speech signal by standard methods. Experiments on speech corrupted by various noises show that the proposed algorithm performs better than our best previous algorithm, VTS, which assumes that the noise is stationary. Rita Singh, Bhiksha Raj |
ICASSP (1) | 2 |
| 2003 | Design of the CMU sphinx-4 decoderabstractSphinx-4 is an open source HMM-based speech recognition system written in the Java ™ programming language. The design of the Sphinx-4 decoder incorporates several new features in response to current demands on HMM-based large vocabulary systems. Some new design aspects include graph construction for multilevel parallel decoding with multiple feature streams without the use of compound HMMs, the incorporation of a generalized search algorithm that subsumes Viterbi decoding as a special case, token stack decoding for efficient maintenance of multiple paths during search, design of a generalized language HMM graph from grammars and language models of multiple standard formats, that can potentially toggle between flat search structure, tree search structure, etc. This paper describes a few of these design aspects, and reports some preliminary performance measures for speed and accuracy. 1. Paul Lamere, Philip Kwok, William Walker, Evandro B. Gouvêa, Rita Singh, Bhiksha Raj |
INTERSPEECH | 6 |
| 2003 | Classification with free energy at raised temperaturesabstractIn this paper we describe a generalized classification method for HMM-based speech recognition systems, that uses free energy as a discriminant function rather than conventional probabilities. The discriminant function incorporates a single adjustable temperature parameter T. The computation of free energy can be motivated using an entropy regularization, where the entropy grows monotonically with the temperature. In the resulting generalized classification scheme, the values of T = 0 and T = 1 give the conventional Viterbi and forward algorithms, respectively, as special cases. We show experimentally that if the test data are mismatched with the classifier, classification at temperatures higher than one can lead to significant improvements in recognition performance. The temperature parameter is far more effective in improving performance on mismatched data than a variance scaling factor, which is another apparent single adjustable parameter that has a very similar analytical form. 1. Rita Singh, Manfred K. Warmuth, Bhiksha Raj, Paul Lamere |
INTERSPEECH | 3 |
| 2003 | Classifier-based non-linear projection for adaptive endpointing of continuous speech
Bhiksha Raj, Rita Singh |
Comput. Speech Lang. | 1 |
| 2003 | Speech-recognizer-based filter optimization for microphone array processingabstractConventional microphone array processing schemes used for speech recognition enhance the output waveform using optimization criteria that are independent of the recognition system. We present a new filter-and-sum array processing algorithm in which the filter parameters are calibrated to maximize recognizer likelihoods. The proposed method provides significant improvement in recognition accuracy over conventional methods. Michael L. Seltzer, Bhiksha Raj |
IEEE Signal Process. Lett. | 2 |
| 2002 | Speech recognizer-based microphone array processing for robust hands-free speech recognitionabstractWe present a new array processing algorithm for microphone array speech recognition. Conventionally, the goal of array processing is to take distorted signals captured by the array and generate a cleaner output waveform. However, speech recognition systems operate on a set of features derived from the waveform, rather than the waveform itself. The goal of an array processor used in conjunction with a recognition system is to generate a waveform which produces a set of recognition features which maximize die likelihood for the words that are spoken, rather than to minimize the waveform distortion. We propose a new array processing algorithm which maximizes the likelihood of the recognition features. This is accomplished through the use of a new objective function which utilizes information from the recognition system itself, obtained in an unsupervised manner, to optimize the parameters of a filter-and-sum array processor. Using the proposed method, improvements in word error rate of up to 36% over conventional methods are achieved on real microphone array tasks in a wide range of environments. Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
ICASSP | 2 |
| 2002 | The MERL SpokenQuery information retrieval system a system for retrieving pertinent documents from a spoken queryabstractThis paper describes some key concepts developed and used in the design of a spoken-query based information retrieval system developed at the Mitsubishi Electric Research Labs (MERL). Innovations in the system include automatic inclusion of signature terms of documents in the recognizer's vocabulary, the use of uncertainty vectors to represent spoken queries, and a method of indexing that accommodates the usage of uncertainty vectors. This paper describes these techniques and includes experimental results that demonstrate their effectiveness. Bhiksha Raj |
ICME (2) | 2 |
| 2002 | Automatic generation of subword units for speech recognition systemsabstractLarge vocabulary continuous speech recognition (LVCSR) systems traditionally represent words in terms of smaller subword units. Both during training and during recognition, they require a mapping table, called the dictionary, which maps words into sequences of these subword units. The performance of the LVCSR system depends critically on the definition of the subword units and the accuracy of the dictionary. In current LVCSR systems, both these components are manually designed. While manually designed subword units generalize well, they may not be the optimal units of classification for the specific task or environment for which an LVCSR system is trained. Moreover, when human expertise is not available, it may not be possible to design good subword units manually. There is clearly a need for data-driven design of these LVCSR components. In this paper, we present a complete probabilistic formulation for the automatic design of subword units and dictionary, given only the acoustic data and their transcriptions. The proposed framework permits easy incorporation of external sources of information, such as the spellings of words in terms of a nonideographic script. Rita Singh, Bhiksha Raj, Richard M. Stern |
IEEE Trans. Speech Audio Process. | 2 |
| 2001 | Speech in Noisy Environments: robust automatic segmentation, feature extraction, and hypothesis combinationabstractThe first evaluation for Speech in Noisy Environments (SPINE1) was conducted by the Naval Research Labs (NRL) in August, 2000. The purpose of the evaluation was to test existing core speech recognition technologies for speech in the presence of varying types and levels of noise. In this case the noises were taken from military settings. Among the strategies used by Carnegie Mellon University's successful systems designed for this task were session-adaptive segmentation, robust mel-scale filtering for the computation of cepstra, the use of parallel front-end features and noise-compensation algorithms, and parallel hypotheses combination through word-graphs. This paper describes the motivations behind the design decisions taken for these components, supported by observations and experiments. Rita Singh, Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
ICASSP | 3 |
| 2001 | A boosting approach for confidence scoringabstractIn this paper we present the application of a boosting classification algorithm to confidence scoring. We derive feature vectors from speech recognition lattices and feed them into a boosting classifier. This classifier combines hundreds of very simple `weak learners' and derives classification rules that can reduce the confidence error rate by up to 34%. We compare our results to those obtained using two other standard classification techniques, Support Vector Machines (SVMs) and Classification and Regression Trees (CART), and show significant improvements. Furthermore, the nature of the boosting algorithm allows us to combine the best single classifier and improve its performance. We present experimental results on real world corpora derived from our SpeechBot Web index http://www.speechbot.com and from the HUB4 DARPA evaluation sets. We believe these results have wide applicability to audio indexing and to acoustic and language modeling adaptation where word confidence scores can be used in iterative adaptation schemes. 1. Pedro J. Moreno 0001, Beth Logan, Bhiksha Raj |
INTERSPEECH | 3 |
| 2001 | Calibration of microphone arrays for improved speech recognitionabstractWe present a new microphone array calibration algorithm specifically designed for speech recognition. Currently, microphone-array-based speech recognition is performed in two independent stages: array processing, and then recognition. Array processing algorithms designed for speech enhancement are used to process the waveforms before recognition. These systems make the assumption that the best array processing methods will result in the best recognition performance. However, recognition systems interpret a set of features extracted from the speech waveform, not the waveform itself. In our calibration method, the filter parameters of a filter-and-sum array processing scheme are optimized to maximize the likelihood of the recognition features extracted from the resulting output signal. By incorporating the speech recognition system into the design of the array processing algorithm, we are able to achieve improvements in word error rate of up to 37 % over conventional array processing methods on both simulated and actual microphone array data. 1. Michael L. Seltzer, Bhiksha Raj |
INTERSPEECH | 2 |
| 2001 | Quantization-based language model compressionabstractThis paper describes two techniques for reducing the size of statistical back-off gram language models in computer memory. Language model compression is achieved through a combination of quantizing language model probabilities and back-off weights and the pruning of parameters that are determined to be unnecessary after quantization. The recognition performance of the original and compressed language models is evaluated across three different language models and two different recognition tasks. The results show that the language models can be compressed by up to 60% of their original size with no significant loss in recognition performance. Moreover, the techniques that are described provide a principled method with which to compress language models further while minimising degradation in recognition performance. Edward W. D. Whittaker, Bhiksha Raj |
INTERSPEECH | 2 |
| 2001 | Comparison of width-wise and length-wise language model compressionabstractIn this paper we investigate the extent to which Katz backoff language models can be compressed through a combination of parameter quantization (width-wise compression) and parameter pruning (length-wise compression) methods while preserving performance. We compare the compression and performance that is achieved using entropy-based pruning against that achieved using only parameter quantization. We then compare combinations of both methods. It is shown that a broadcast news language model can be compressed by up to 83 % to only 12.6Mb with no loss in performance on a broadcast news task. Compressing the language model further by quantization to 10.3Mb resulted in only a 0.4 % degradation in word error rate which is better than can be achieved through entropy-based pruning alone. 1. Edward W. D. Whittaker, Bhiksha Raj |
INTERSPEECH | 2 |
| 2000 | Automatic generation of phone sets and lexical transcriptionsabstractLarge vocabulary automatic speech recognition systems model words as sequences of a small set of basic sub-word units (the phoneset), which the systems are trained to classify. All words in the system's vocabulary are transcribed in terms of this set in a dictionary. The phoneset and dictionary are specific to a language and are typically designed manually. The system's performance is critically dependent on the quality of the phoneset and the accuracy of the dictionary. The authors attempt to generate the phoneset and dictionary automatically, using only the training data and their transcriptions. We treat this as a joint optimization problem with a maximum a posteriori solution for the dictionary and a maximum likelihood solution for the phoneset and its acoustic models. Experiments with the DARPA Resource Management corpus show that the automatically generated phoneset and dictionary result in recognition accuracies close to those obtained using manually designed ones. Rita Singh, Bhiksha Raj, Richard M. Stern |
ICASSP | 2 |
| 2000 | Reconstruction of damaged spectrographic features for robust speech recognitionabstractWe present two missing-feature based algorithms that recover noise-corrupted regions of spectrographic representations of speech for noise-robust speech recognition. These algorithms modify the incoming feature vector without any changes to the speech recognition system, in contrast to previously-described approaches. The first approach clusters the feature vectors representing clean speech. Missing data are recovered by estimating the spectral cluster in each analysis frame based on the uncorrupted feature values. The second approach uses MAP procedures to estimate the values of missing data elements based on their correlations with the features that are present. Both methods take into account bounds on the clean spectrogram implied by the noisy spectrogram. Large improvements in recognition accuracy are observed when these methods are used on speech corrupted by non-stationary noise when the locations of the corrupt regions of the spectrogram are known. We also present a new method of estimating the locations of corrupt regions in spectrograms that treats the problem of identifying these regions as one of Bayesian classification. This method, when used along with the best method to reconstruct them, results in recognition accuracies comparable with the best previous data compensation algorithm on speech corrupted by white noise. It also provides significant improvement on speech corrupted by music when the global SNR of the corrupted signal is known a priori. 1. Bhiksha Raj, Michael L. Seltzer, Richard M. Stern |
INTERSPEECH | 1 |
| 2000 | Classifier-based mask estimation for missing feature methods of robust speech recognitionabstractMissing feature methods of noise compensation for speech recognition operate by removing components of a spectrographic representation of speech that are considered to be corrupt, as indicated by a low signal-to-noise ratio. Recognition is either performed directly on the incomplete spectrograms or the missing components are reconstructed prior to recognition. These methods require a spectrographic mask which accurately labels the reliable and corrupt regions of the spectrogram. Current methods of mask estimation rely on assumptions about the corrupting noise such as stationarity. This is a significant drawback since the missing feature methods themselves have no such restrictions. We present a new mask estimation technique that uses a Bayesian classifier to determine the reliability of spectrographic elements. Features were designed that make no assumptions about the corrupting noise signal, but rather exploit characteristics of the speech signal itself. Missing feature compensation experiments were performed on speech corrupted by a variety of noises. In all cases, classifier-based mask estimation resulted in significantly better recognition accuracy than conventional mask estimation methods. 1. Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
INTERSPEECH | 2 |
| 2000 | Structured redefinition of sound units by merging and splitting for improved speech recognitionabstractThe performance of speech recognition systems degrades when the basic sound units used are poorly defined or inconsistently used. Several attempts have been made to improve dictionaries automatically, either by redefining pronunciations of words in terms of existing sound units, or by redefining the sound units themselves completely. The problem with these approaches is that, while the former is limited by the sound units used, the latter discards all human information that has been incorporated into an expert-designed recognition dictionary. In this paper we propose a new merging-andsplitting algorithm that attempts to redefine the basic sound units used in the dictionary, while maintaining the expert knowledge built into a manually designed dictionary. Sound units from an existing dictionary are merged based on their inherent confusability, as measured by a Monte-Carlo based metric, and subsequently split to maximize the likelihood of the training data. Experiments with the Resource Management database indicate that this approach results in an improvement in recognition accuracy when context-independent models are used for recognition. When context-dependent models are used, the improvement observed is reduced. 1. Rita Singh, Bhiksha Raj, Richard M. Stern |
INTERSPEECH | 2 |
| 1999 | Automatic clustering and generation of contextual questions for tied states in hidden Markov modelsabstractMost current automatic speech recognition systems based on HMMs cluster or tie together subsets of the subword units with which speech is represented. This tying improves the recognition accuracy when systems are trained with limited data, and is performed by classifying the sub-phonetic units using a series of binary tests based on speech production, called "linguistic questions". This paper describes a new method for automatically determining the best combinations of subword units to form these questions. The hybrid algorithm proposed clusters state distributions of context-independent phones to obtain questions for triphonetic contexts. Experiments confirm that the questions thus generated can replace manually generated questions and can provide improved recognition accuracy. Automatic generation of questions has the additional important advantage of extensibility to languages for which the phonetic structure is not well understood by the system designer, and can be effectively used in situations where the subword units are not phonetically motivated. Rita Singh, Bhiksha Raj, Richard M. Stern |
ICASSP | 2 |
| 1999 | Domain adduced state tying for cross-domain acoustic modellingabstractHomophone words is one of the specific problems of Automatic Speech Recognition (ASR) in French. Moreover, this phenomenon is particularly high for some inflections like the singular/plural inflection (72% of the 40.7K lemma of our 240K word dictionary have inflected forms which are homophonic). In order to take into account worddependencies spanning over a variable number of words, it is interesting to merge local language models, like 3-gram or 3-class models, with largespan models. We present in this paper two kinds of models : a phrase-based model, using phrases obtained from a training corpus by means of a finite-state parser; a homophone cache-based model, using derivation of constraints from word histories stored in a cache memory. Rita Singh, Bhiksha Raj, Richard M. Stern |
EUROSPEECH | 2 |
| 1998 | Inference of missing spectrographic features for robust speech recognitionabstractTwo types of algorithms are introduced that recover missing time-frequency regions of log-spectral representations of speech. These compensation algorithms modify the incoming feature vector without any changes to the speech recognition system, in contrast to previously-described approaches. The first approach clusters the log-spectral vectors representing clean speech. Missing data are recovered by estimating the spectral cluster in each analysis frame on the basis of the feature values that are present. The second approach uses MAP procedures to estimate the values of missing data elements based on their correlation with the features that are present. Greatest recognition accuracy was obtained using the correlation-based approach, presumably because of its ability to exploit the temporal as well as spectral structure of speech. The recognition accuracy provided by these algorithms approaches but does not exceed that obtained by traditional marginalization. Nevertheless, it is believed that these algorithms provide greater computational efficiency and enable greater flexibility in recognition system structure. 1. Bhiksha Raj, Rita Singh, Richard M. Stern |
ICSLP | 1 |
| 1998 | Data-driven environmental compensation for speech recognition: A unified approach
Pedro J. Moreno 0001, Bhiksha Raj, Richard M. Stern |
Speech Commun. | 2 |
| 1997 | The effects of background music on speech recognition accuracyabstractRecognition of broadcast data, such as TV and radio programs is a topic of great interest. One of the problems with such data is the frequent presence of background music that degrades the performance of speech recognition systems. In this paper we examine the effects of different kinds of music on automatic speech recognition systems by comparing the effects of music with the relatively well-known effects of white noise on these systems. We also examine the extent to which compensation algorithms that have been successfully applied to noisy speech are also helpful in improving recognition accuracy for speech that is corrupted by music. It is hoped that these experimental comparisons will lead to a better understanding of how to compensate for the effects of background music. 1. Bhiksha Raj, Vipul N. Parikh, Richard M. Stern |
ICASSP | 1 |
| 1996 | A vector Taylor series approach for environment-independent speech recognitionabstractIn this paper we introduce a new analytical approach to environment compensation for speech recognition. Previous attempts at solving analytically the problem of noisy speech recognition have either used an overly-simplified mathematical description of the effects of noise on the statistics of speech or they have relied on the availability of large environment-specific adaptation sets. Some of the previous methods required the use of adaptation data that consists of simultaneously-recorded or "stereo" recordings of clean and degraded speech. In this work we introduce the use of a vector Taylor series (VTS) expansion to characterize efficiently and accurately the effects on speech statistics of unknown additive noise and unknown linear filtering in a transmission channel. The VTS approach is computationally efficient. It can be applied either to the incoming speech feature vectors, or to the statistics representing these vectors. In the first case the speech is compensated and then recognized; in the second case HMM statistics are modified using the VTS formulation. Both approaches use only the actual speech segment being recognized to compute the parameters required for environmental compensation. We evaluate the performance of two implementations of VTS algorithms using the CMU SPHINX-II system on the 100-word alphanumeric CENSUS database and on the 1993 5000-word ARPA Wall Street Journal database. Artificial white Gaussian noise is added to both databases. The VTS approaches provide significant improvements in recognition accuracy compared to previous algorithms. Pedro J. Moreno 0001, Bhiksha Raj, Richard M. Stern |
ICASSP | 2 |
| 1996 | Cepstral compensation by polynomial approximation for environment-independent speech recognition
Bhiksha Raj, Evandro B. Gouvêa, Pedro J. Moreno 0001, Richard M. Stern |
ICSLP | 1 |
| 1995 | Multivariate-Gaussian-based cepstral normalization for robust speech recognitionabstractWe introduce a new family of environmental compensation algorithms called multivariate gaussian based cepstral normalization (RATZ). RATZ assumes that the effects of unknown noise and filtering on speech features can be compensated by corrections to the mean and variance of components of Gaussian mixtures, and an efficient procedure for estimating the correction factors is provided. The RATZ algorithm can be implemented to work with or without the use of "stereo" development data that had been simultaneously recorded in the training and testing environments. "Blind" RATZ partially overcomes the loss of information that would have been provided by stereo training through the use of a more accurate description of how noisy environments affect clean speech. We evaluate the performance of the two RATZ algorithms using the CMU SPHINX-II system on the alphanumeric census database and compare their performance with that of previous environmental-robustness developed at CMU. Pedro J. Moreno 0001, Bhiksha Raj, Evandro B. Gouvêa, Richard M. Stern |
ICASSP | 2 |
| 1995 | A unified approach for robust speech recognitionabstractThere are two major structural approaches to robust speech recognition.In the first approach to the problem, compensation is performed by modifying the incoming cepstral stream using ML or MMSE methods to estimate parameters characterizing environmental degradation, from direct frame-by-frame comparisons between speech recorded in high-quality and degraded acoustical environments, or by signal processing techniques such as spectral subtraction.The second approach tackles the problem by modifying the statistics of the internal representation of speech cepstra in the classifier to make them more closely resemble the statistics of degraded speech.This paper attempts to unify these approaches to robust speech recognition by presenting three techniques that share the same basic assumptions and internal structure but differ in whether they modify the incoming speech cepstra or whether they modify the classifier statistics.We present SNR-dependent multi-vaRiate gAussian-based cepsTral normaliZation (SNR-RATZ) and SNR-based Blind RATZ (SNR-BRATZ), which modify incoming cepstra, along with STAR (STAtistical Re-estimation), which modifies the internal statistics of the classifier.The algorithms were tested using the SPHINX-II speech recognition system on the CENSUS database, a database of strings of letters and numbers to which unknown added and unknown linear filtering was introduced artificially.While all the algorithms showed good performance, STAR was observed to provide lower error rates as SNR decreases than any of the algorithms that modify incoming cepstra. Pedro J. Moreno 0001, Bhiksha Raj, Richard M. Stern |
EUROSPEECH | 2 |