VLDB 2026 Research / reviewers in the wild / expert
Nima Mesgarani
dblp:23/1353
· DBLP profile ↗
72ranked-venue papers
10as first author
29since 2021 · last 2026
0000-0002-2987-759XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 9 first-author · 21 since 2021Artificial intelligence and machine learning · 40 · 6 first-author · 16 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech SynthesisabstractDiffusion-based text-to-speech (TTS) systems have made remarkable progress in zero-shot speech synthesis, yet optimizing all components for perceptual metrics remains challenging. Prior work with DMOSpeech demonstrated direct metric optimization for speech generation components, but duration prediction remained unoptimized. This paper presents DMOSpeech 2, which extends metric optimization to the duration predictor through a reinforcement learning approach. The proposed system implements a novel duration policy framework using group relative preference optimization (GRPO) with speaker similarity and word error rate as reward signals. By optimizing this previously unoptimized component, DMOSpeech 2 creates a more complete metric-optimized synthesis pipeline. Additionally, this paper introduces teacher-guided sampling, a hybrid approach leveraging a teacher model for initial denoising steps before transitioning to the student model, significantly improving output diversity while maintaining efficiency. Comprehensive evaluations demonstrate superior performance across all metrics compared to previous systems, while reducing sampling steps by half without quality degradation. These advances represent a significant step toward speech synthesis systems with metric optimization across multiple components. Yinghao Aaron Li, Xilin Jiang, Cheng Niu, Kaifeng Xu, Juntong Song, Nima Mesgarani |
AAAI | 7 |
| 2025 | AAD-LLM: Neural Attention-Driven Auditory Scene UnderstandingabstractXilin Jiang, Sukru Samet Dindar, Vishal Choudhari, Stephan Bickel, Ashesh Mehta, Guy M McKhann, Daniel Friedman, Adeen Flinker, Nima Mesgarani. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xilin Jiang, Sukru Samet Dindar, Vishal Choudhari, Stephan Bickel, Ashesh D. Mehta, Guy M. McKhann II, Daniel Friedman, Adeen Flinker, Nima Mesgarani |
ACL (1) | 9 |
| 2025 | Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech RepresentationsabstractTransformer-based speech language models (SLMs) have significantly improved neural speech recognition and understanding.While existing research has examined how well SLMs encode shallow acoustic and phonetic features, the extent to which SLMs encode nuanced syntactic and conceptual features remains unclear.By drawing parallels with linguistic competence assessments for large language models, this study is the first to systematically evaluate the presence of contextual syntactic and semantic features across SLMs for self-supervised learning (S3M), automatic speech recognition (ASR), speech compression (codec), and as the encoder for auditory large language models (AudioLLMs).Through minimal pair designs and diagnostic feature analysis across 71 tasks spanning diverse linguistic levels, our layer-wise and time-resolved analysis uncovers that 1) all speech encode grammatical features more robustly than conceptual ones.2) Despite never seeing text, S3M match or surpass ASR encoders on every linguistic level, demonstrating that rich grammatical and even conceptual knowledge can arise purely from audio.3) S3M representations peak mid-network and then crash in the final layers, whereas ASR and AudioLLM encoders maintain or improve, reflecting how pre-training objectives reshape late-layer content.4) Temporal probing further shows that S3Ms encode grammatical cues 500 ms before a word begins, whereas AudioLLMs distribute evidence more evenly-indicating that objectives shape not only where but also when linguistic information is most salient.Together, these findings establish the first largescale map of contextual syntax and semantics in speech models and highlight both the promise and the limits of current SLM training paradigms. * Equal contribution. Linyang He, Qiaolin Wang, Xilin Jiang, Nima Mesgarani |
EMNLP | 4 |
| 2025 | Decoding the Unintelligible: Neural Speech Tracking in Low Signal-to-Noise RatiosabstractUnderstanding speech in noisy environments is challenging for both human listeners and speech technologies, with significant implications for hearing aid design and communication systems. Auditory attention decoding (AAD) aims to decode the attended talker from neural signals to enhance their speech and improve perception. However, whether this decoding remains reliable under severely degraded listening conditions remains unclear. In this study, we investigated selective neural tracking of the attended speaker under adverse listening conditions. Using EEG recordings in a multi-talker speech perception task with varying SNR, participants' task performance—quantified through a repeated-word detection task—was analyzed as a proxy for perceptual accuracy and attentional focus, while neural responses were used to decode the attended talker. Despite substantial degradation in task performance, we found that neural tracking of attended speech persists, suggesting that the brain retains sufficient information for decoding. These findings demonstrate that even in highly challenging conditions, AAD remains feasible, offering a potential avenue for improving speech perception in brain-informed audio technologies, such as hearing aids, that leverage AAD to enhance listening experiences in real-world noisy environments. Xiaomin He, Vinay S. Raghavan, Nima Mesgarani |
ICASSP | 3 |
| 2025 | Dual-path Mamba: Short and Long-term Bidirectional Selective Structured State Space Models for Speech SeparationabstractTransformers have been the most successful architecture for various speech modeling tasks, including speech separation. However, the self-attention mechanism in transformers with quadratic complexity is inefficient in computation and memory. Recent models incorporate new layers and modules along with transformers for better performance but also introduce extra model complexity. In this work, we replace transformers with Mamba, a selective state space model, for speech separation. We propose dual-path Mamba, which models short-term and long-term forward and backward dependency of speech signals using selective state spaces. Our experimental results on the WSJ0-2mix data show that our dual-path Mamba models of comparably smaller sizes outperform state-of-the-art RNN model DPRNN, CNN model WaveSplit, and transformer model Sepformer. Our largest model reaches an SI-SNRi of 23.4 dB. Code available: https://github.com/xi-j/Mamba-TasNet Xilin Jiang, Cong Han 0001, Nima Mesgarani |
ICASSP | 3 |
| 2025 | Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and SynthesisabstractIt is too early to conclude that Mamba is a better alternative to transformers for speech before comparing Mamba with transformers in terms of both performance and efficiency in multiple speech-related tasks. To reach this conclusion, we propose and evaluate three models for three tasks: Mamba-TasNet for speech separation, ConMamba for speech recognition, and VALL-M for speech synthesis. We compare them with transformers of similar sizes in performance, memory, and speed. Our Mamba or Mamba-transformer hybrid models show comparable or higher performance than their transformer counterparts: Sepformer, Conformer, and VALL-E. They are more efficient than transformers in memory and speed for speech longer than a threshold duration, inversely related to the resolution of a speech token. Mamba for separation is the most efficient, and Mamba for recognition is the least. Further, we show that Mamba is not more efficient than transformer for speech shorter than the threshold duration and performs worse in models that require joint modeling of text and speech, such as cross or masked attention of two inputs. Therefore, we argue that the superiority of Mamba or transformer depends on particular problems and models. Code available1 Xilin Jiang, Yinghao Aaron Li, Adrian Nicolas Florea, Cong Han 0001, Nima Mesgarani |
ICASSP | 5 |
| 2025 | Neuro2Semantic: A Transfer Learning Framework for Semantic Reconstruction of Continuous Language from Human Intracranial EEG
Siavash Shams, Richard J. Antonello, Gavin Mischler, Stephan Bickel, Ashesh D. Mehta, Nima Mesgarani |
INTERSPEECH | 6 |
| 2025 | StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style DiffusionabstractYinghao Aaron Li, Xilin Jiang, Cong Han, Nima Mesgarani. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yinghao Aaron Li, Xilin Jiang, Cong Han 0001, Nima Mesgarani |
NAACL (Long Papers) | 4 |
| 2025 | Far from the Shallow: Brain-Predictive Reasoning Embedding through Residual DisentanglementabstractUnderstanding how the human brain progresses from processing simple linguistic inputs to performing high-level reasoning is a fundamental challenge in neuroscience. While modern large language models (LLMs) are increasingly used to model neural responses to language, their internal representations are highly "entangled," mixing information about lexicon, syntax, meaning, and reasoning. This entanglement biases conventional brain encoding analyses toward linguistically shallow features (e.g., lexicon and syntax), making it difficult to isolate the neural substrates of cognitively deeper processes. Here, we introduce a residual disentanglement method that computationally isolates these components. By first probing an LM to identify feature-specific layers, our method iteratively regresses out lower-level representations to produce four nearly orthogonal embeddings for lexicon, syntax, meaning, and, critically, reasoning. We used these disentangled embeddings to model intracranial (ECoG) brain recordings from neurosurgical patients listening to natural speech. We show that: 1) This isolated reasoning embedding exhibits unique predictive power, accounting for variance in neural activity not explained by other linguistic features and even extending to the recruitment of visual regions beyond classical language areas. 2) The neural signature for reasoning is temporally distinct, peaking later (~350-400ms) than signals related to lexicon, syntax, and meaning, consistent with its position atop a processing hierarchy. 3) Standard, non-disentangled LLM embeddings can be misleading, as their predictive success is primarily attributable to linguistically shallow features, masking the more subtle contributions of deeper cognitive processing. Our work provides compelling neural evidence for an abstract reasoning computation during language comprehension and offers a robust framework for mapping distinct cognitive functions from artificial models to the human brain. Linyang He, Tianjun Zhong, Richard J. Antonello, Gavin Mischler, Micah Goldblum, Nima Mesgarani |
NeurIPS | 6 |
| 2025 | ZeroSep: Separate Anything in Audio with Zero TrainingabstractAudio source separation is fundamental for machines to understand complex acoustic environments and underpins numerous audio applications. Current supervised deep learning approaches, while powerful, are limited by the need for extensive, task-specific labeled data and struggle to generalize to the immense variability and open-set nature of real-world acoustic scenes. Inspired by the success of generative foundation models, we investigate whether pre-trained text-guided audio diffusion models can overcome these limitations. We make a surprising discovery: zero-shot source separation can be achieved purely through a pre-trained text-guided audio diffusion model under the right configuration. Our method, named ZeroSep, works by inverting the mixed audio into the diffusion model's latent space and then using text conditioning to guide the denoising process to recover individual sources. Without any task-specific training or fine-tuning, ZeroSep repurposes the generative diffusion model for a discriminative separation task and inherently supports open-set scenarios through its rich textual priors. ZeroSep is compatible with a variety of pre-trained text-guided audio diffusion backbones and delivers strong separation performance on multiple separation benchmarks, surpassing even supervised methods. Chao Huang 0033, Yuesheng Ma, Junxuan Huang, Susan Liang, Yunlong Tang 0002, Jing Bi 0002, Nima Mesgarani, Chenliang Xu |
NeurIPS | 8 |
| 2024 | Exploring Self-supervised Contrastive Learning of Spatial Sound Event RepresentationabstractIn this study, we present a simple multi-channel framework for contrastive learning (MC-SimCLR) to encode 'what' and 'where' of spatial audios. MC-SimCLR learns joint spectral and spatial representations from unlabeled spatial audios, thereby enhancing both event classification and sound localization in downstream tasks. At its core, we propose a multi-level data augmentation pipeline that augments different levels of audio features, including waveforms, Mel spectrograms, and generalized cross-correlation (GCC) features. In addition, we introduce simple yet effective channel-wise augmentation methods to randomly swap the order of the microphones and mask Mel and GCC channels. By using these augmentations, we find that linear layers on top of the learned representation significantly outperform supervised models in terms of both event classification accuracy and localization error. We also perform a comprehensive analysis of the effect of each augmentation method and a comparison of the fine-tuning performance using different amounts of labeled data. Xilin Jiang, Cong Han 0001, Yinghao Aaron Li, Nima Mesgarani |
ICASSP | 4 |
| 2024 | SSAMBA: Self-Supervised Audio Representation Learning With Mamba State Space ModelabstractTransformers have revolutionized deep learning across various tasks, including audio representation learning, due to their powerful modeling capabilities. However, they often suffer from quadratic complexity in both GPU memory usage and computational inference time, affecting their efficiency. Recently, state space models (SSMs) like Mamba have emerged as a promising alternative, offering a more efficient approach by avoiding these complexities. Given these advantages, we explore the potential of SSM-based models in audio tasks. In this paper, we introduce Self-Supervised Audio Mamba (SSAMBA), the first self-supervised, attention-free, and SSM-based model for audio representation learning. SSAMBA leverages the bidirectional Mamba to capture complex audio patterns effectively. We incorporate a self-supervised pretraining framework that optimizes both discriminative and generative objectives, enabling the model to learn robust audio representations from large-scale, unlabeled datasets. We evaluated SSAMBA on various tasks such as audio classification, keyword spotting, speaker identification, and emotion recognition. Our results demonstrate that SSAMBA outperforms the Self-Supervised Audio Spectrogram Transformer (SSAST) in most tasks. Notably, SSAMBA is approximately 92.7% faster in batch inference speed and 95.4% more memory-efficient than SSAST for the tiny model size with an input token size of 22k. These efficiency gains, combined with superior performance, underscore the effectiveness of SSAMBA’s architectural innovation, making it a compelling choice for a wide range of audio processing applications. Code at https://github.com/SiavashShams/ssamba. Siavash Shams, Sukru Samet Dindar, Xilin Jiang, Nima Mesgarani |
SLT | 4 |
| 2024 | Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify And Understand Speaker in Spoken DialogueabstractIn recent years, we have observed a rapid advancement in speech language models (SpeechLLMs), catching up with humans’ listening and reasoning abilities. SpeechLLMs have demonstrated impressive spoken dialog question-answering (SQA) performance in benchmarks like Gaokao, the English listening test of the college entrance exam in China, which seemingly requires understanding both the spoken content and voice characteristics of speakers in a conversation. However, after carefully examining Gaokao’s questions, we find the correct answers to many questions can be inferred from the conversation transcript alone, i.e. without speaker segmentation and identification. Our evaluation of state-of-the-art models Qwen-Audio and WavLLM on both Gaokao and our proposed “What Do You Like?” dataset shows a significantly higher accuracy in these context-based questions than in identity-critical questions, which can only be answered reliably with correct speaker identification. The results and analysis suggest that when solving SQA, the current SpeechLLMs exhibit limited speaker awareness from the audio and behave similarly to an LLM reasoning from the conversation transcription without sound. We propose that tasks focused on identity-critical questions could offer a more accurate evaluation framework of SpeechLLMs in SQA. Junkai Wu, Xulin Fan, Bo-Ru Lu, Xilin Jiang, Nima Mesgarani, Mark Hasegawa-Johnson, Mari Ostendorf |
SLT | 5 |
| 2023 | Online Binaural Speech Separation Of Moving Speakers With A Wavesplit NetworkabstractBinaural speech separation in real-world scenarios often involves moving speakers. Most current speech separation methods use utterance-level permutation invariant training (u-PIT) for training. In inference time, however, the order of outputs can be inconsistent over time particularly in long-form speech separation. This situation which is referred to as the speaker swap problem is even more problematic when speakers constantly move in space and therefore poses a challenge for consistent placement of speakers in output channels. Here, we describe a real-time binaural speech separation model based on a Wavesplit network to mitigate the speaker swap problem for moving speaker separation. Our model computes a speaker embedding for each speaker at each time frame from the mixed audio, aggregates embeddings using online clustering, and uses cluster centroids as speaker profiles to track each speaker throughout the long duration. Experimental results on reverberant, long-form moving multitalker speech separation show that the proposed method is less prone to speaker swap and achieves comparable performance with u-PIT based models with ground truth tracking in both separation accuracy and preserving the interaural cues. Cong Han 0001, Nima Mesgarani |
ICASSP | 2 |
| 2023 | Phoneme-Level Bert for Enhanced Prosody of Text-To-Speech with Grapheme PredictionsabstractLarge-scale pre-trained language models have been shown to be helpful in improving the naturalness of text-to-speech (TTS) models by enabling them to produce more naturalistic prosodic patterns. However, these models are usually word-level or sup-phoneme-level and jointly trained with phonemes, making them inefficient for the downstream TTS task where only phonemes are needed. In this work, we propose a phoneme-level BERT (PL-BERT) with a pretext task of predicting the corresponding graphemes along with the regular masked phoneme predictions. Subjective evaluations show that our phoneme-level BERT encoder has significantly improved the mean opinion scores (MOS) of rated naturalness of synthesized speech compared with the state-of-the-art (SOTA) StyleTTS baseline on out-of-distribution (OOD) texts. Yinghao Aaron Li, Cong Han 0001, Xilin Jiang, Nima Mesgarani |
ICASSP | 4 |
| 2023 | DeCoR: Defy Knowledge Forgetting by Predicting Earlier Audio CodesabstractLifelong audio feature extraction involves learning new sound classes incrementally, which is essential for adapting to new data distributions over time.However, optimizing the model only on new data can lead to catastrophic forgetting of previously learned tasks, which undermines the model's ability to perform well over the long term.This paper introduces a new approach to continual audio representation learning called DeCoR.Unlike other methods that store previous data, features, or models, DeCoR indirectly distills knowledge from an earlier model to the latest by predicting quantization indices from a delayed codebook.We demonstrate that DeCoR improves acoustic scene classification accuracy and integrates well with continual self-supervised representation learning.Our approach introduces minimal storage and computation overhead, making it a lightweight and efficient solution for continual learning. Xilin Jiang, Yinghao Aaron Li, Nima Mesgarani |
INTERSPEECH | 3 |
| 2023 | StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language ModelsabstractIn this paper, we present StyleTTS 2, a text-to-speech (TTS) model that leverages style diffusion and adversarial training with large speech language models (SLMs) to achieve human-level TTS synthesis. StyleTTS 2 differs from its predecessor by modeling styles as a latent random variable through diffusion models to generate the most suitable style for the text without requiring reference speech, achieving efficient latent diffusion while benefiting from the diverse speech synthesis offered by diffusion models. Furthermore, we employ large pre-trained SLMs, such as WavLM, as discriminators with our novel differentiable duration modeling for end-to-end training, resulting in improved speech naturalness. StyleTTS 2 surpasses human recordings on the single-speaker LJSpeech dataset and matches it on the multispeaker VCTK dataset as judged by native English speakers. Moreover, when trained on the LibriTTS dataset, our model outperforms previous publicly available models for zero-shot speaker adaptation. This work achieves the first human-level TTS on both single and multispeaker datasets, showcasing the potential of style diffusion and adversarial training with large SLMs. The audio demos and source code are available at https://styletts2.github.io/. Yinghao Aaron Li, Cong Han 0001, Vinay S. Raghavan, Gavin Mischler, Nima Mesgarani |
NeurIPS | 5 |
| 2022 | Styletts-VC: One-Shot Voice Conversion by Knowledge Transfer From Style-Based TTS ModelsabstractOne-shot voice conversion (VC) aims to convert speech from any source speaker to an arbitrary target speaker with only a few seconds of reference speech from the target speaker. This relies heavily on disentangling the speaker's identity and speech content, a task that still remains challenging. Here, we propose a novel approach to learning disentangled speech representation by transfer learning from style-based text-to-speech (TTS) models. With cycle consistent and adversarial training, the style-based TTS models can perform transcription-guided one-shot VC with high fidelity and similarity. By learning an additional mel-spectrogram encoder through a teacher-student knowledge transfer and novel data augmentation scheme, our approach results in disentangled speech representation without needing the input text. The subjective evaluation shows that our approach can significantly outperform the previous state-of-the-art one-shot voice conversion models in both naturalness and similarity. Yinghao Aaron Li, Cong Han 0001, Nima Mesgarani |
SLT | 3 |
| 2021 | Speaker and Direction Inferred Dual-Channel Speech SeparationabstractMost speech separation methods, trying to separate all channel sources simultaneously, are still far from having enough generalization capabilities for real scenarios where the number of input sounds is usually uncertain and even dynamic. In this work, we employ ideas from auditory attention with two ears and propose a speaker and direction inferred speech separation network (dubbed SDNet) to solve the cocktail party problem. Specifically, our SDNet first parses out the respective perceptual representations with their speaker and direction characteristics from the mixture of the scene in a sequential manner. Then, the perceptual representations are utilized to attend to each corresponding speech. Our model generates more precise perceptual representations with the help of spatial features and successfully deals with the problem of the unknown number of sources and the selection of outputs. The experiments on standard fully-overlapped speech separation benchmarks, WSJ0-2mix, WSJ0-3mix, and WSJ0-2&3mix, show the effectiveness, and our method achieves SDR improvements of 25.31 dB, 17.26 dB, and 21.56 dB under anechoic settings. Our codes will be released at https://github.com/aispeech-lab/SDNet. Chenxing Li, Jiaming Xu 0001, Nima Mesgarani, Bo Xu 0002 |
ICASSP | 3 |
| 2021 | Rethinking The Separation Layers In Speech Separation NetworksabstractModules in all existing speech separation networks can be categorized into single-input-multi-output (SIMO) modules and single-input-single-output (SISO) modules. SIMO modules generate more outputs than input, and SISO modules keep the numbers of input and output the same. While the majority of separation models only contain SIMO architectures, it has also been shown that certain two-stage separation systems integrated with a post-enhancement SISO module can improve the separation quality. Why performance improvements can be achieved by incorporating the SISO modules? Are SIMO modules always necessary? In this paper, we empirically examine those questions by designing models with varying configurations in the SIMO and SISO modules. We show that comparing with the standard SIMO-only design, a mixed SIMO-SISO design with a same model size is able to improve the separation performance especially under low-overlap conditions. We further validate the necessity of SIMO modules and show that SISO-only models are still able to perform separation without sacrificing the performance. The observations allow us to rethink the model design paradigm and present different views on how the separation is performed. Yi Luo 0004, Zhuo Chen 0006, Cong Han 0001, Chenda Li, Tianyan Zhou, Nima Mesgarani |
ICASSP | 6 |
| 2021 | Ultra-Lightweight Speech Separation Via Group CommunicationabstractModel size and complexity remain the biggest challenges in the deployment of speech enhancement and separation systems on low-resource devices such as earphones and hearing aids. Although methods such as compression, distillation and quantization can be applied to large models, they often come with a cost on the model performance. In this paper, we provide a simple model design paradigm that explicitly designs ultra-lightweight models without sacrificing the performance. Motivated by the sub-band frequency-LSTM (F-LSTM) architectures, we introduce the group communication (GroupComm), where a feature vector is split into smaller groups and a small processing block is used to perform inter-group communication. Unlike standard F-LSTM models where the sub-band outputs are concatenated, an ultra-small module is applied on all the groups in parallel, which allows a significant decrease on the model size. Experiment results show that comparing with a strong baseline model which is already lightweight, GroupComm can achieve on par performance with 35.6 times fewer parameters and 2.3 times fewer operations. Yi Luo 0004, Cong Han 0001, Nima Mesgarani |
ICASSP | 3 |
| 2021 | Continuous Speech Separation Using Speaker Inventory for Long Recording
Cong Han 0001, Yi Luo 0004, Chenda Li, Tianyan Zhou, Keisuke Kinoshita, Shinji Watanabe 0001, Marc Delcroix, Hakan Erdogan, John R. Hershey, Nima Mesgarani, Zhuo Chen 0006 |
Interspeech | 10 |
| 2021 | Binaural Speech Separation of Moving Speakers With Preserved Spatial Cues
Cong Han 0001, Yi Luo 0004, Nima Mesgarani |
Interspeech | 3 |
| 2021 | StarGANv2-VC: A Diverse, Unsupervised, Non-Parallel Framework for Natural-Sounding Voice ConversionabstractWe present an unsupervised non-parallel many-to-many voice conversion (VC) method using a generative adversarial network (GAN) called StarGAN v2. Using a combination of adversarial source classifier loss and perceptual loss, our model significantly outperforms previous VC models. Although our model is trained only with 20 English speakers, it generalizes to a variety of voice conversion tasks, such as any-to-many, cross-lingual, and singing conversion. Using a style encoder, our framework can also convert plain reading speech into stylistic speech, such as emotional and falsetto speech. Subjective and objective evaluation experiments on a non-parallel many-to-many voice conversion task revealed that our model produces natural sounding voices, close to the sound quality of state-of-the-art text-to-speech (TTS) based voice conversion methods without the need for text labels. Moreover, our model is completely convolutional and with a faster-than-real-time vocoder such as Parallel WaveGAN can perform real-time voice conversion. Yinghao Aaron Li, Ali Zare, Nima Mesgarani |
Interspeech | 3 |
| 2021 | Empirical Analysis of Generalized Iterative Speech Separation Networks
Yi Luo 0004, Cong Han 0001, Nima Mesgarani |
Interspeech | 3 |
| 2021 | Implicit Filter-and-Sum Network for End-to-End Multi-Channel Speech Separation
Nima Mesgarani |
Interspeech | 2 |
| 2021 | Understanding Adaptive, Multiscale Temporal Integration In Deep Speech Recognition SystemsabstractNatural signals such as speech are hierarchically structured across many different timescales, spanning tens (e.g., phonemes) to hundreds (e.g., words) of milliseconds, each of which is highly variable and context-dependent. While deep neural networks (DNNs) excel at recognizing complex patterns from natural signals, relatively little is known about how DNNs flexibly integrate across multiple timescales. Here, we show how a recently developed method for studying temporal integration in biological neural systems – the temporal context invariance (TCI) paradigm – can be used to understand temporal integration in DNNs. The method is simple: we measure responses to a large number of stimulus segments presented in two different contexts and estimate the smallest segment duration needed to achieve a context invariant response. We applied our method to understand how the popular DeepSpeech2 model learns to integrate across time in speech. We find that nearly all of the model units, even in recurrent layers, have a compact integration window within which stimuli substantially alter the response and outside of which stimuli have little effect. We show that training causes these integration windows to shrink at early layers and expand at higher layers, creating a hierarchy of integration windows across the network. Moreover, by measuring integration windows for time-stretched/compressed speech, we reveal a transition point, midway through the trained network, where integration windows become yoked to the duration of stimulus structures (e.g., phonemes or words) rather than absolute time. Similar phenomena were observed in a purely recurrent and purely convolutional network although structure-yoked integration was more prominent in the recurrent network. These findings suggest that deep speech recognition systems use a common motif to encode the hierarchical structure of speech: integrating across short, time-yoked windows at early layers and long, structure-yoked windows at later layers. Our method provides a straightforward and general-purpose toolkit for understanding temporal integration in black-box machine learning models. Menoua Keshishian, Samuel Norman-Haignere, Nima Mesgarani |
NeurIPS | 3 |
| 2021 | Distortion-Controlled Training for end-to-end Reverberant Speech Separation with Auxiliary Autoencoding LossabstractThe performance of speech enhancement and separation systems in anechoic environments has been significantly advanced with the recent progress in end-to-end neural network architectures. However, the performance of such systems in reverberant environments is yet to be explored. A core problem in reverberant speech separation is about the training and evaluation metrics. Standard time-domain metrics may introduce unexpected distortions during training and fail to properly evaluate the separation performance due to the presence of the reverberations. In this paper, we first introduce the "equal-valued contour" problem in reverberant separation where multiple outputs can lead to the same performance measured by the common metrics. We then investigate how "better" outputs with lower target-specific distortions can be selected by auxiliary autoencoding training (A2T). A2T assumes that the separation is done by a linear operation on the mixture signal, and it adds an loss term on the autoencoding of the direct-path target signals to ensure that the distortion introduced on the direct-path signals is controlled during separation. Evaluations on separation signal quality and speech recognition accuracy show that A2T is able to control the distortion on the direct-path signals and improve the recognition accuracy. Yi Luo 0004, Cong Han 0001, Nima Mesgarani |
SLT | 3 |
| 2021 | Group Communication With Context Codec for Lightweight Source SeparationabstractDespite the recent progress on neural network architectures for speech separation, the balance between the model size, model complexity and model performance is still an important and challenging problem for the deployment of such models to low-resource platforms. In this paper, we propose two simple modules, group communication and context codec, that can be easily applied to a wide range of architectures to jointly decrease the model size and complexity without sacrificing the performance. A group communication module splits a high-dimensional feature into groups of low-dimensional features and captures the inter-group dependency. A separation module with a significantly smaller model size can then be shared by all the groups. A context codec module, containing a context encoder and a context decoder, is designed as a learnable downsampling and upsampling module to decrease the length of a sequential feature processed by the separation module. The combination of the group communication and the context codec modules is referred to as the GC3 design. Experimental results show that applying GC3 on multiple network architectures for speech separation can achieve on-par or better performance with as small as 2.5% model size and 17.6% model complexity, respectively. Yi Luo 0004, Cong Han 0001, Nima Mesgarani |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Real-Time Binaural Speech Separation with Preserved Spatial CuesabstractDeep learning speech separation algorithms have achieved great success in improving the quality and intelligibility of separated speech from mixed audio. Most previous methods focused on generating a single-channel output for each of the target speakers, hence discarding the spatial cues needed for the localization of sound sources in space. However, preserving the spatial information is important in many applications that aim to accurately render the acoustic scene such as in hearing aids and augmented reality (AR). Here, we propose a speech separation algorithm that preserves the interaural cues of separated sound sources and can be implemented with low latency and high fidelity, therefore enabling a real-time modification of the acoustic scene. Based on the time-domain audio separation network (TasNet), a single-channel time-domain speech separation system that can be implemented in real-time, we propose a multi-input-multi-output (MIMO) end-to-end extension of TasNet that takes binaural mixed audio as input and simultaneously separates target speakers in both channels. Experimental results show that the proposed end-to-end MIMO system is able to significantly improve the separation performance and keep the perceived location of the modified sources intact in various acoustic scenes. Cong Han 0001, Yi Luo 0004, Nima Mesgarani |
ICASSP | 3 |
| 2020 | End-to-end Microphone Permutation and Number Invariant Multi-channel Speech SeparationabstractAn important problem in ad-hoc microphone speech separation is how to guarantee the robustness of a system with respect to the locations and numbers of microphones. The former requires the system to be invariant to different indexing of the microphones with the same locations, while the latter requires the system to be able to process inputs with varying dimensions. Conventional optimization-based beamforming techniques satisfy these requirements by definition, while for deep learning-based end-to-end systems those constraints are not fully addressed. In this paper, we propose transform-average-concatenate (TAC), a simple design paradigm for channel permutation and number invariant multi-channel speech separation. Based on the filter-and-sum network (FaSNet), a recently proposed end-to-end time-domain beamforming system, we show how TAC significantly improves the separation performance across various numbers of microphones in noisy reverberant separation tasks with ad-hoc arrays. Moreover, we show that TAC also significantly improves the separation performance with fixed geometry array configuration, further proving the effectiveness of the proposed paradigm in the general problem of multi-microphone speech separation. Yi Luo 0004, Zhuo Chen 0006, Nima Mesgarani, Takuya Yoshioka |
ICASSP | 3 |
| 2020 | Separating Varying Numbers of Sources with Auxiliary Autoencoding LossabstractMany recent source separation systems are designed to separate a fixed number of sources out of a mixture. In the cases where the source activation patterns are unknown, such systems have to either adjust the number of outputs or to identify invalid outputs from the valid ones. Iterative separation methods have gain much attention in the community as they can flexibly decide the number of outputs, however (1) they typically rely on long-term information to determine the stopping time for the iterations, which makes them hard to operate in a causal setting; (2) they lack a fault tolerance mechanism when the estimated number of sources is different from the actual number. In this paper, we propose a simple training method, the auxiliary autoencoding permutation invariant training (A2PIT), to alleviate the two issues. A2PIT assumes a fixed number of outputs and uses auxiliary autoencoding loss to force the invalid outputs to be the copies of the input mixture, and detects invalid outputs in a fully unsupervised way during inference phase. Experiment results show that A2PIT is able to improve the separation performance across various numbers of speakers and effectively detect the number of speakers in a mixture. Yi Luo 0004, Nima Mesgarani |
INTERSPEECH | 2 |
| 2019 | FaSNet: Low-Latency Adaptive Beamforming for Multi-Microphone Audio ProcessingabstractBeamforming has been extensively investigated for multi-channel audio processing tasks. Recently, learning-based beamforming methods, sometimes called neural beamformers, have achieved significant improvements in both signal quality (e.g. signal-to-noise ratio (SNR)) and speech recognition (e.g. word error rate (WER)). Such systems are generally non-causal and require a large context for robust estimation of inter-channel features, which is impractical in applications requiring low-latency responses. In this paper, we propose filter-and-sum network (FaSNet), a time-domain, filter-based beamforming approach suitable for low-latency scenarios. FaSNet has a two-stage system design that first learns frame-level time-domain adaptive beamforming filters for a selected reference channel, and then calculate the filters for all remaining channels. The filtered outputs at all channels are summed to generate the final output. Experiments show that despite its small model size, FaSNet is able to outperform several traditional oracle beamformers with respect to scale-invariant signal-to-noise ratio (SI-SNR) in reverberant speech enhancement and separation tasks. Moreover, when trained with a frequency-domain objective function on the CHiME-3 dataset, FaSNet achieves 14.3% relative word error rate reduction (RWERR) compared with the baseline model. These results show the efficacy of FaSNet particularly in reverberant and noisy signal conditions. Yi Luo 0004, Cong Han 0001, Nima Mesgarani, Enea Ceolini, Shih-Chii Liu |
ASRU | 3 |
| 2019 | Augmented Time-frequency Mask Estimation in Cluster-based Source Separation AlgorithmsabstractTime-frequency mask estimation with various clustering approaches has proven effective in solving the audio source separation problem. In this framework, the time-frequency bins of the mixture spectrogram are represented in a high-dimensional embedding space, where various methods can be applied to group the embedded points to calculate either hard or soft source assignments and subsequently the time-frequency masks. However, the mismatch between the assignment algorithm during the training and inference phases in majority of the current approaches leads to a suboptimal solution, because the assignment objective that is used during the training (e.g. ideal binary mask) is not the same as the one used during the inference phase (e.g. k-means clustering). We propose a method to reduce the mismatch between these two conditions where the source embedding is trained such that the source assignment during training and inference phases results in similar outcomes. Our results show that matching the source assignment during training- and inference-phase results in more accurate and consistent mask estimation in the inference phase which significantly improves the source separation accuracy for various hard and soft clustering methods. Yi Luo 0004, Nima Mesgarani |
ICASSP | 2 |
| 2019 | Online Deep Attractor Network for Real-time Single-channel Speech SeparationabstractSpeaker-independent speech separation is a challenging audio processing problem. In recent years, several deep learning algorithms have been proposed to address this problem. The majority of these methods use noncausal implementation which limits their application in real-time scenarios such as in wearable hearing devices and low-latency telecommunication. In this paper, we propose the Online Deep Attractor Network (ODANet), an extension to the Deep Attractor Network (DANet) which is causal and enables real-time speech separation. In contrast with DANet that estimates the global attractor point for each speaker using the entire utterance, ODANet estimates the attractors for each time step and tracks them using a dynamic weighting function with only causal information. This not only solves the speaker tracking problem, but also allows ODANet to generate more stable embeddings across time. Experimental results show that ODANet can achieve a similar separation accuracy as the noncausal DANet in both two speaker and three speaker speech separation problems, which makes it a suitable candidate for applications that require robust real-time speech processing. Cong Han 0001, Yi Luo 0004, Nima Mesgarani |
ICASSP | 3 |
| 2019 | Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech SeparationabstractSingle-channel, speaker-independent speech separation methods have recently seen great progress. However, the accuracy, latency, and computational cost of such methods remain insufficient. The majority of the previous methods have formulated the separation problem through the time-frequency representation of the mixed signal, which has several drawbacks, including the decoupling of the phase and magnitude of the signal, the suboptimality of time-frequency representation for speech separation, and the long latency of the entire system. To address these shortcomings, we propose a fully-convolutional time-domain audio separation network (Conv-TasNet), a deep learning framework for end-to-end time-domain speech separation. Conv-TasNet uses a linear encoder to generate a representation of the speech waveform optimized for separating individual speakers. Speaker separation is achieved by applying a set of weighting functions (masks) to the encoder output. The modified encoder representations are then inverted back to the waveforms using a linear decoder. The masks are found using a temporal convolutional network (TCN) consisting of stacked 1-D dilated convolutional blocks, which allows the network to model the long-term dependencies of the speech signal while maintaining a small model size. The proposed Conv-TasNet system significantly outperforms previous time-frequency masking methods in separating two- and three-speaker mixtures. Additionally, Conv-TasNet surpasses several ideal time-frequency magnitude masks in two-speaker speech separation as evaluated by both objective distortion measures and subjective quality assessment by human listeners. Finally, Conv-TasNet has a significantly smaller model size and a much shorter minimum latency, making it a suitable solution for both offline and real-time speech separation applications. This study therefore represents a major step toward the realization of speech separation systems for real-world speech processing technologies. Yi Luo 0004, Nima Mesgarani |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | TaSNet: Time-Domain Audio Separation Network for Real-Time, Single-Channel Speech SeparationabstractRobust speech processing in multi-talker environments requires effective speech separation. Recent deep learning systems have made significant progress toward solving this problem, yet it remains challenging particularly in real-time, short latency applications. Most methods attempt to construct a mask for each source in time-frequency representation of the mixture signal which is not necessarily an optimal representation for speech separation. In addition, time-frequency decomposition results in inherent problems such as phase/magnitude decoupling and long time window which is required to achieve sufficient frequency resolution. We propose Time-domain Audio Separation Network (TasNet) to overcome these limitations. We directly model the signal in the time-domain using an encoder-decoder framework and perform the source separation on nonnegative encoder outputs. This method removes the frequency decomposition step and reduces the separation problem to estimation of source masks on encoder outputs which is then synthesized by the decoder. Our system outperforms the current state-of-the-art causal and noncausal speech separation algorithms, reduces the computational cost of speech separation, and significantly reduces the minimum required latency of the output. This makes TasNet suitable for applications where low-power, real-time implementation is desirable such as in hearable and telecommunication devices. Yi Luo 0004, Nima Mesgarani |
ICASSP | 2 |
| 2018 | Lip2Audspec: Speech Reconstruction from Silent Lip Movements VideoabstractIn this study, we propose a deep neural network for reconstructing intelligible speech from silent lip movement videos. We use auditory spectrogram as spectral representation of speech and its corresponding sound generation method resulting in a more natural sounding reconstructed speech. Our proposed network consists of an autoencoder to extract bottleneck features from the auditory spectrogram which is then used as target to our main lip reading network comprising of CNN, LSTM and fully connected layers. Our experiments show that the autoencoder is able to reconstruct the original auditory spectrogram with a 98% correlation and also improves the quality of reconstructed speech from the main lip reading network. Our model, trained jointly on different speakers is able to extract individual speaker characteristics and gives promising results of reconstructing intelligible speech with superior word recognition accuracy. Hassan Akbari, Himani Arora, Liangliang Cao, Nima Mesgarani |
ICASSP | 4 |
| 2018 | Real-time Single-channel Dereverberation and Separation with Time-domain Audio Separation Network
Yi Luo 0004, Nima Mesgarani |
INTERSPEECH | 2 |
| 2018 | Music Source Activity Detection and Separation Using Deep Attractor Network
Rajath Kumar, Yi Luo 0004, Nima Mesgarani |
INTERSPEECH | 3 |
| 2018 | Speech Processing in the Human Brain Meets Deep Learning
Nima Mesgarani |
INTERSPEECH | 1 |
| 2018 | Speaker-Independent Speech Separation With Deep Attractor NetworkabstractDespite the recent success of deep learning for many speech processing tasks, single-microphone, speaker-independent speech separation remains challenging for two main reasons. The first reason is the arbitrary order of the target and masker speakers in the mixture (permutation problem), and the second is the unknown number of speakers in the mixture (output dimension problem). We propose a novel deep learning framework for speech separation that addresses both of these issues. We use a neural network to project the time-frequency representation of the mixture signal into a high-dimensional embedding space. A reference point (attractor) is created in the embedding space to represent each speaker which is defined as the centroid of the speaker in the embedding space. The time-frequency embeddings of each speaker are then forced to cluster around the corresponding attractor point which is used to determine the time-frequency assignment of the speaker. We propose three methods for finding the attractors for each source in the embedding space and compare their advantages and limitations. The objective function for the network is standard signal reconstruction error which enables end-to-end operation during both training and test phases. We evaluated our system using the Wall Street Journal dataset (WSJ0) on two and three speaker mixtures and report comparable or better performance than other state-of-the-art deep learning methods for speech separation. Yi Luo 0004, Zhuo Chen 0006, Nima Mesgarani |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Deep attractor network for single-microphone speaker separationabstractDespite the overwhelming success of deep learning in various speech processing tasks, the problem of separating simultaneous speakers in a mixture remains challenging. Two major difficulties in such systems are the arbitrary source permutation and unknown number of sources in the mixture. We propose a novel deep learning framework for single channel speech separation by creating attractor points in high dimensional embedding space of the acoustic signals which pull together the time-frequency bins corresponding to each source. Attractor points in this study are created by finding the centroids of the sources in the embedding space, which are subsequently used to determine the similarity of each bin in the mixture to each source. The network is then trained to minimize the reconstruction error of each source by optimizing the embeddings. The proposed model is different from prior works in that it implements an end-to-end training, and it does not depend on the number of sources in the mixture. Two strategies are explored in the test time, K-means and fixed attractor points, where the latter requires no post-processing and can be implemented in real-time. We evaluated our system on Wall Street Journal dataset and show 5.49% improvement over the previous state-of-the-art methods. Zhuo Chen 0006, Yi Luo 0004, Nima Mesgarani |
ICASSP | 3 |
| 2017 | NAPLib: An open source toolbox for real-time and offline Neural Acoustic ProcessingabstractIn this paper, we introduce the Neural Acoustic Processing Library (NAPLib), a toolbox containing novel processing methods for real-time and offline analysis of neural activity in response to speech. Our method divides the speech signal and resultant neural activity into segmental units (e.g., phonemes), allowing for fast and efficient computations that can be implemented in real-time. NAPLib contains a suite of tools that characterize various properties of the neural representation of speech, which can be used for functionality such as characterizing electrode tuning properties, brain mapping and brain computer interfaces. The library is general and applicable to both invasive and non-invasive recordings, including electroencephalography (EEG), electrocorticography (ECoG) and magnetoecnephalography (MEG). In this work, we describe the structure of NAPLib, as well as demonstrating its use in both EEG and ECoG. We believe NAPLib provides a valuable tool to both clinicians and researchers who are interested in the representation of speech in the brain. Bahar Khalighinejad, Tasha Nagamine, Ashesh D. Mehta, Nima Mesgarani |
ICASSP | 4 |
| 2017 | Deep clustering and conventional networks for music separation: Stronger togetherabstractDeep clustering is the first method to handle general audio separation scenarios with multiple sources of the same type and an arbitrary number of sources, performing impressively in speaker-independent speech separation tasks. However, little is known about its effectiveness in other challenging situations such as music source separation. Contrary to conventional networks that directly estimate the source signals, deep clustering generates an embedding for each time-frequency bin, and separates sources by clustering the bins in the embedding space. We show that deep clustering outperforms conventional networks on a singing voice separation task, in both matched and mismatched conditions, even though conventional networks have the advantage of end-to-end training for best signal approximation, presumably because its more flexible objective engenders better regularization. Since the strengths of deep clustering and conventional network architectures appear complementary, we explore combining them in a single hybrid network trained via an approach akin to multi-task learning. Remarkably, the combination significantly outperforms either of its components. Yi Luo 0004, Zhuo Chen 0006, John R. Hershey, Jonathan Le Roux, Nima Mesgarani |
ICASSP | 5 |
| 2017 | Understanding the Representation and Computation of Multilayer Perceptrons: A Case Study in Speech RecognitionabstractDespite the recent success of deep learning, the nature of the transformations they apply to the input features remains poorly understood. This study provides an empirical framework to study the encoding properties of node activations in various layers of the network, and to construct the exact function applied to each data point in the form of a linear transform. These methods are used to discern and quantify properties of feed-forward neural networks trained to map acoustic features to phoneme labels. We show a selective and nonlinear warping of the feature space, achieved by forming prototypical functions to account for the possible variation of each class. This study provides a joint framework where the properties of node activations and the functions implemented by the network can be linked together. Tasha Nagamine, Nima Mesgarani |
ICML | 2 |
| 2016 | Analyzing distributional learning of phonemic categories in unsupervised deep neural networks
Okko Johannes Räsänen, Tasha Nagamine, Nima Mesgarani |
CogSci | 3 |
| 2016 | Synaptic depression in deep neural networks for speech processingabstractA characteristic property of biological neurons is their ability to dynamically change the synaptic efficacy in response to variable input conditions. This mechanism, known as synaptic depression, significantly contributes to the formation of normalized representation of speech features. Synaptic depression also contributes to the robust performance of biological systems. In this paper, we describe how synaptic depression can be modeled and incorporated into deep neural network architectures to improve their generalization ability. We observed that when synaptic depression is added to the hidden layers of a neural network, it reduces the effect of changing background activity in the node activations. In addition, we show that when synaptic depression is included in a deep neural network trained for phoneme classification, the performance of the network improves under noisy conditions not included in the training phase. Our results suggest that more complete neuron models may further reduce the gap between the biological performance and artificial computing, resulting in networks that better generalize to novel signal conditions. Minda Yang, Nima Mesgarani |
ICASSP | 4 |
| 2016 | Adaptation of Neural Networks Constrained by Prior Statistics of Node Co-Activations
Tasha Nagamine, Zhuo Chen 0006, Nima Mesgarani |
INTERSPEECH | 3 |
| 2016 | On the Role of Nonlinear Transformations in Deep Neural Network Acoustic Models
Tasha Nagamine, Michael L. Seltzer, Nima Mesgarani |
INTERSPEECH | 3 |
| 2015 | Exploring how deep neural networks form phonemic categoriesabstractDeep neural networks (DNNs) have become the dominant technique for acoustic-phonetic modeling due to their markedly improved performance over other models. Despite this, little is understood about the computation they implement in creating phonemic categories from highly variable acoustic signals. In this paper, we analyzed a DNN trained for phoneme recognition and characterized its representational properties, both at the single node and population level in each layer. At the single node level, we found strong selectivity to distinct phonetic features in all layers. Node selectivity to specific manners and places of articulation appeared from the first hidden layer and became more explicit in deeper layers. Furthermore, we found that nodes with similar phonetic feature selectivity were differentially activated to different exemplars of these features. Thus, each node becomes tuned to a particular acoustic manifestation of the same feature, providing an effective representational basis for the formation of invariant phonemic categories. This study reveals that phonetic features organize the activations in different layers of a DNN, a result that mirrors the recent findings of feature encoding in the human auditory system. These insights may provide better understanding of the limitations of current models, leading to new strategies to improve their performance. Index Terms: Deep neural networks, deep learning, automatic speech recognition. Tasha Nagamine, Michael L. Seltzer, Nima Mesgarani |
INTERSPEECH | 3 |
| 2015 | Speech reconstruction from human auditory cortex with deep neural networks
Minda Yang, Sameer A. Sheth, Catherine A. Schevon, Guy M. McKhann II, Nima Mesgarani |
INTERSPEECH | 5 |
| 2014 | Principal components of auditory spectro-temporal receptive fields
Nagaraj Mahajan, Nima Mesgarani, Hynek Hermansky |
INTERSPEECH | 2 |
| 2013 | Developing a speaker identification system for the DARPA RATS projectabstractThis paper describes the speaker identification (SID) system developed by the Patrol team for the first phase of the DARPA RATS (Robust Automatic Transcription of Speech) program, which seeks to advance state of the art detection capabilities on audio from highly degraded communication channels. We present results using multiple SID systems differing mainly in the algorithm used for voice activity detection (VAD) and feature extraction. We show that (a) unsupervised VAD performs as well supervised methods in terms of downstream SID performance, (b) noise-robust feature extraction methods such as CFCCs out-perform MFCC front-ends on noisy audio, and (c) fusion of multiple systems provides 24% relative improvement in EER compared to the single best system when using a novel SVM-based fusion algorithm that uses side information such as gender, language, and channel id. Oldrich Plchot, Spyridon Matsoukas, Pavel Matejka, Najim Dehak, Jeff Z. Ma, Sandro Cumani, Ondrej Glembek, Hynek Hermansky, Sri Harish Reddy Mallidi, Nima Mesgarani, Richard M. Schwartz, Mehdi Soufifar, Zheng-Hua Tan, Samuel Thomas 0001, Bing Zhang 0004, Xinhui Zhou |
ICASSP | 10 |
| 2012 | The UMD-JHU 2011 speaker recognition systemabstractIn recent years, there have been significant advances in the field of speaker recognition that has resulted in very robust recognition systems. The primary focus of many recent developments have shifted to the problem of recognizing speakers in adverse conditions, e.g in the presence of noise/reverberation. In this paper, we present the UMD-JHU speaker recognition system applied on the NIST 2010 SRE task. The novel aspects of our systems are: 1) Improved performance on trials involving different vocal effort via the use of linear-scale features; 2) Expected improved recognition performance in the presence of reverberation and noise via the use of frequency domain perceptual linear predictor and cortical features; 3) A new discriminative kernel partial least squares (KPLS) framework that complements state-of-the-art back-end systems JFA and PLDA to aid in better overall recognition; and 4) Acceleration of JFA, PLDA and KPLS back-ends via distributed computing. The individual components of the system and the fused system are compared against a baseline JFA system and results reported by SRI and MIT-LL on SRE2010. Daniel Garcia-Romero, Xinhui Zhou, Dmitry N. Zotkin, Balaji Vasan Srinivasan, Yuancheng Luo, Sriram Ganapathy, Samuel Thomas 0001, Sridhar Krishna Nemala, Garimella S. V. S. Sivaram, Majid Mirbagheri, Sri Harish Reddy Mallidi, Thomas Janu, Padmanabhan Rajan, Nima Mesgarani, Mounya Elhilali, Hynek Hermansky, Shihab A. Shamma, Ramani Duraiswami |
ICASSP | 14 |
| 2012 | Speech and speaker separation in human auditory cortex
Nima Mesgarani, Edward Chang |
INTERSPEECH | 1 |
| 2012 | Developing a Speech Activity Detection System for the DARPA RATS ProgramabstractThis paper describes the speech activity detection (SAD) system developed by the Patrol team for the first phase of the DARPA RATS (Robust Automatic Transcription of Speech) program, which seeks to advance state of the art detection capabilities on audio from highly degraded communication channels. We present two approaches to SAD, one based on Gaussian mixture models, and one based on multi-layer perceptrons. We show that significant gains in SAD accuracy can be obtained by careful design of acoustic front end, feature normalization, incorporation of long span features via data-driven dimensionality reducing transforms, and channel dependent modeling. We also present a novel technique for normalizing detection scores from different systems for the purpose of system combination. Tim Ng, Bing Zhang 0004, Long Nguyen 0001, Spyridon Matsoukas, Xinhui Zhou, Nima Mesgarani, Karel Veselý, Pavel Matejka |
INTERSPEECH | 6 |
| 2012 | Acoustic and Data-driven Features for Robust Speech Activity DetectionabstractIn this paper we evaluate different features for speech activity detection (SAD). Several signal processing techniques are used to derive acoustic features that capture attributes of speech useful in differentiating speech segments in noise. The acoustic features include short-term spectral features, long-term modulation features both derived using Frequency Domain Linear Prediction (FDLP), and joint spectro-temporal features extracted using 2D filters on a cortical representation of speech. Posteriors of speech and non-speech from a trained multi-layer perceptron are also used as data-driven features for this task. These feature extraction techniques form part of an elaborate feature extraction front-end where information spanning several hundreds of milliseconds of the signal are used along with heteroscedastic linear discriminant analysis for dimensionality reduction. Processed feature outputs from the proposed front-end are used to train SAD systems based on Gaussian mixture models for processing of speech from multiple languages transmitted over noisy radio communication channels under the ongoing DARPA Robust Automatic Transcription of Speech (RATS) program. The proposed front-end performs significantly better than standard acoustic feature extraction techniques in these noisy conditions. Samuel Thomas 0001, Sri Harish Reddy Mallidi, Thomas Janu, Hynek Hermansky, Nima Mesgarani, Xinhui Zhou, Shihab A. Shamma, Tim Ng, Bing Zhang 0004, Long Nguyen 0001, Spyridon Matsoukas |
INTERSPEECH | 5 |
| 2012 | Automatic intelligibility assessment of pathologic speech in head and neck cancer based on auditory-inspired spectro-temporal modulationsabstractOral, head and neck cancer represents 3% of all cancers in the United States and is the 6th most common cancer worldwide. Depending on the tumor size, location and staging, patients are treated by radical surgery, radiology, chemotherapy or a combination of those treatments. As a result, their anatomical structures for speech are impaired and this leads to some negative impact on their speech intelligibility. As a part of the INTERSPEECH 2012 speaker trait Pathology sub-challenge, this study explored the use of auditory-inspired spectro-temporal modulation features for automatic speech intelligibility assessment of those pathologic speech. The averaged spectro-temporal modulations of speech considered as either intelligible or non-intelligible in the challenge database were analyzed and it was found that the non-intelligible speech tends to have its modulation amplitude peaks shift towards a smaller rate and scale. Based on SVM and GMM, variants of spectro-temporal modulation features were tested on the speaker trait challenge problem and the resulting performances on both the development and the test datasets are comparable to the baseline performance. Xinhui Zhou, Daniel Garcia-Romero, Nima Mesgarani, Maureen Stone 0001, Carol Y. Espy-Wilson, Shihab A. Shamma |
INTERSPEECH | 3 |
| 2011 | Speech processing with a cortical representation of audioabstractNeurophysiological studies in the primary auditory cortex have recently demonstrated a rich diversity of responses that provide an explicit multidimensional representation of phonemic acoustic features (Mesgarani 2008). Specifically, distinct subsets of cortical neurons are activated by articulatory gestures and dynamics that are characteristic of different phonemes. Here we use a computational cortical model to illustrate how these phonetic features appear in such a multiresolution representation. We also review how this representation has been successfully applied in variety of speech processing tasks including robust speech discrimination, speech enhancement and phoneme recognition. Nima Mesgarani, Shihab A. Shamma |
ICASSP | 1 |
| 2011 | Adaptive Stream Fusion in Multistream Recognition of SpeechabstractA new method to deal with variable distortions of speech during the operation of the system is proposed. First, multiple processing streams are formed by extracting different spectral and temporal modulation components from the speech signal. Information in each stream is used to estimate posterior probabilities of phonemes. Initial values for a weighted integration of these individual estimates are found by normalized cross-correlation of the estimates with the actual phoneme labels on the training data. A statistical model of the final estimated posterior probabilities is used to characterize the system performance. During the operation, the weights in the linear fusion are adapted using particle filtering to optimize the performance. Results on phoneme recognition from noisy speech indicate the effectiveness of the proposed method. Index Terms: multistream speech recognition, spectrotemporal modulations Nima Mesgarani, Samuel Thomas 0001, Hynek Hermansky |
INTERSPEECH | 1 |
| 2010 | Nonlinear filtering of spectrotemporal modulations in speech enhancementabstractA monaural noise-suppression algorithm is proposed that nonlinearly manipulates the spectrotemporal modulations of speech as represented in a model of auditory cortical processing. A distinctive aspect of this approach is its consideration of the non-stationary dynamic behavior of speech that is captured using nonlinear filters, thus achieving excellent perceptual quality in the presence of many types of additive noise. Subjective tests demonstrate a significant improvement in the quality of the enhanced speech over the noisy samples even at SNRs as low as -5 dB. Majid Mirbagheri, Nima Mesgarani, Shihab A. Shamma |
ICASSP | 2 |
| 2010 | A multistream multiresolution framework for phoneme recognitionabstractSpectrotemporal representation of speech has already shown promising results in speech processing technologies, however, many inherent issues of such representation, such as high dimensionality have limited their use in speech and speaker recognition. Multistream framework fits very well to such representation where different regions can be separately mapped into posterior probabilities of classes before merging. In this study, we investigated the effective ways of forming streams out of this representation for robust phoneme recognition. We also investigated multiple ways of fusing the posteriors of different streams based on their individual confidence or interactions between them. We observed 8.6% relative improvement in clean and 4 % in noise. We developed a simple yet effective linear combination technique that provides intuitive understanding of stream combinations and how even systematic errors can be learned to reduce confusions. Index Terms: speech recognition, spectrotemporal modulations, multistream Nima Mesgarani, Samuel Thomas 0001, Hynek Hermansky |
INTERSPEECH | 1 |
| 2010 | A phoneme recognition framework based on auditory spectro-temporal receptive fieldsabstractWe propose to incorporate features derived using spectrotemporal receptive fields (STRFs) of neurons in the auditory cortex for phoneme recognition. Each of these STRFs is tuned to different auditory frequencies, scales and modulation rates. We select different sets of STRFs which are specific for phonemes in different broad phonetic classes (BPC) of sounds. These STRFs are then used as spectro-temporal filters on spectrograms of speech to extract features for phoneme recognition. For the phoneme recognition task on the TIMIT database, the proposed features show a relative improvement of about 5% over conventional feature extraction techniques. Samuel Thomas 0001, Kailash Patil, Sriram Ganapathy, Nima Mesgarani, Hynek Hermansky |
INTERSPEECH | 4 |
| 2010 | The use of spike-based representations for hardware audition systemsabstractHumans are able to process speech and other sounds effectively in adverse environments, hearing through noise, reverberation, and interference from other speakers. To date, machines have been unable to match human performance. One profound difference between biological and engineering systems comes at the input stage. In machines, an acoustic signal is typically chopped into short equally spaced segments in time. In biological systems, the cochlea outputs asynchronous spikes that respond in real-time to acoustic inputs. In this paper we describe a spiking cochlea implementation and recent experiments in both speaker and speech recognition that use spikes as input. Shih-Chii Liu, Nima Mesgarani, John G. Harris, Hynek Hermansky |
ISCAS | 2 |
| 2010 | Data-Driven and Feedback Based Spectro-Temporal Features for Speech RecognitionabstractThis paper proposes novel data-driven and feedback based discriminative spectro-temporal filters for feature extraction in automatic speech recognition (ASR). Initially a first set of spectro-temporal filters are designed to separate each phoneme from the rest of the phonemes. A hybrid Hidden Markov Model/Multilayer Perceptron (HMM/MLP) phoneme recognition system is trained on the features derived using these filters. As a feedback to the feature extraction stage, top confusions of this system are identified, and a second set of filters are designed specifically to address these confusions. Phoneme recognition experiments on TIMIT show that the features derived from the combined set of discriminative filters outperform conventional speech recognition features, and also contain significant complementary information. Garimella S. V. S. Sivaram, Sridhar Krishna Nemala, Nima Mesgarani, Hynek Hermansky |
IEEE Signal Process. Lett. | 3 |
| 2009 | Discriminant spectrotemporal features for phoneme recognitionabstractWe propose discriminant methods for deriving twodimensional spectrotemporal features for phoneme recognition that are estimated to maximize the separation between the representations of phoneme classes. The linearity of the filters results in their intuitive interpretation enabling us to investigate the working principles of the system and to improve its performance by locating the sources of error. Two methods for the estimation of filters are proposed: Regularized Least Square (RLS) and Modified Linear Discriminant Analysis (MLDA). Both methods reach a comparable improvement over the baseline condition demonstrating the advantage of the discriminant spectrotemporal filters. Nima Mesgarani, Garimella S. V. S. Sivaram, Sridhar Krishna Nemala, Mounya Elhilali, Hynek Hermansky |
INTERSPEECH | 1 |
| 2007 | Representation of Phonemes in Primary Auditory Cortex: How the Brain Analyzes SpeechabstractMany transformations inspired by the auditory system have improved the performance of automatic speech recognition (ASR) systems. However, humans perform substantially better than today's ASR systems, suggesting that ASR systems can further benefit from understanding how the brain represents speech. To learn about the cortical representation of speech, we measured the neural responses in the primary auditory cortex to sentences from the TIMIT database. Here we examine how individual phonemes activate different subsets of auditory neurons, reflecting the diversity of neural tuning properties. We find that neurons with different spectro-temporal tuning provide an explicit multidimensional representation of articulatory features independent of speaker and context. This representation that matches the human perception could provide a framework for ASR in adverse conditions. Nima Mesgarani, Stephen V. David, Shihab A. Shamma |
ICASSP (4) | 1 |
| 2006 | Discriminating speech and non-speech with regularized least squaresabstractWe consider the task of discriminating speech and non-speech in noisy environments. Previously, Mesgarani et. al [1] achieved state-of-the-art performance using a cortical representation of sound in conjunction with a feature reduction algorithm and a nonlinear support vector machine classifier. In the present work, we show that we can achieve the same or better accuracy by using a linear regularized least squares classifier directly on the highdimensional cortical representation; the new system is substantially simpler conceptually and computationally. We select the regularization constant automatically, yielding a parameter-free learning system. Intriguingly, we find that optimal classifiers for noisy data can be trained on clean data using heavy regularization. Index Terms: speech detection, discriminative methods, regularization. 1. Ryan Rifkin, Nima Mesgarani |
INTERSPEECH | 2 |
| 2006 | Discrimination of speech from nonspeech based on multiscale spectro-temporal ModulationsabstractWe describe a content-based audio classification algorithm based on novel multiscale spectro-temporal modulation features inspired by a model of auditory cortical processing. The task explored is to discriminate speech from nonspeech consisting of animal vocalizations, music, and environmental sounds. Although this is a relatively easy task for humans, it is still difficult to automate well, especially in noisy and reverberant environments. The auditory model captures basic processes occurring from the early cochlear stages to the central cortical areas. The model generates a multidimensional spectro-temporal representation of the sound, which is then analyzed by a multilinear dimensionality reduction technique and classified by a support vector machine (SVM). Generalization of the system to signals in high level of additive noise and reverberation is evaluated and compared to two existing approaches (Scheirer and Slaney, 2002 and Kingsbury et al., 2002). The results demonstrate the advantages of the auditory model over the other two systems, especially at low signal-to-noise ratios (SNRs) and high reverberation. Nima Mesgarani, Malcolm Slaney, Shihab A. Shamma |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Speech Enhancement Based on Filtering the Spectrotemporal ModulationsabstractA monaural noise suppression algorithm is proposed based on filtering the spectrotemporal modulations of noisy speech. The modulations are estimated from a multiscale representation of the signal spectrogram generated by a model of sound processing in the auditory system. A significant advantage of this method is its ability to suppress noise that has distinctive modulation patterns, despite being spectrally overlapping with the speech. The performance of the algorithm is evaluated using subjective and objective tests and compared to the optimal smoothing and minimum statistics approach (Martin (2001)). The results demonstrate the efficacy of the spectrotemporal filtering approach in the conditions examined. Nima Mesgarani, Shihab A. Shamma |
ICASSP (1) | 1 |
| 2004 | Speech discrimination based on multiscale spectro-temporal modulationsabstractA novel approach for content based audio classification is presented based on multiscale spectro-temporal modulation features extracted using a model of auditory cortex. The task is to discriminate speech from non-speech which consists of animal vocalizations, music and environmental sounds. Generalization of the system to signals in high level of additive noise and reverberation is evaluated and compared to two existing approaches. The results demonstrate the advantages of the auditory model over the other two systems, especially at low SNR and high reverberation. Nima Mesgarani, Shihab A. Shamma, Malcolm Slaney |
ICASSP (1) | 1 |