EDBT 2026 Demo / reviewers in the wild / expert
Mounya Elhilali
dblp:08/2568
· DBLP profile ↗
45ranked-venue papers
4as first author
15since 2021 · last 2025
0000-0003-2597-738XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 14 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion TransformerabstractIn this paper, we introduce SoloAudio, a novel diffusion-based generative model for target sound extraction (TSE). Our approach trains latent diffusion models on audio, replacing the previous U-Net backbone with a skip-connected Transformer that operates on latent features. SoloAudio supports both audio-oriented and language-oriented TSE by utilizing a CLAP model as the feature extractor for target sounds. Furthermore, SoloAudio leverages synthetic audio generated by state-of-the-art text-to-audio models for training, demonstrating strong generalization to out-of-domain data and unseen sound events. We evaluate this approach on the FSD Kaggle 2018 mixture dataset and real data from AudioSet, where SoloAudio achieves the state-of-the-art results on both in-domain and out-of-domain data, and exhibits impressive zero-shot and few-shot capabilities. Source code1and demos2are released. Helin Wang, Jiarui Hai, Yen-Ju Lu, Karan Thakkar, Mounya Elhilali, Najim Dehak |
ICASSP | 5 |
| 2025 | EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
Jiarui Hai, Yong Xu 0004, Hao Zhang 0112, Chenxing Li, Helin Wang, Mounya Elhilali, Dong Yu 0001 |
INTERSPEECH | 6 |
| 2024 | DPM-TSE: A Diffusion Probabilistic Model for Target Sound ExtractionabstractCommon target sound extraction (TSE) approaches primarily relied on discriminative approaches in order to separate the target sound while minimizing interference from the unwanted sources, with varying success in separating the target from the background. This study introduces DPM-TSE, a generative method based on diffusion probabilistic modeling (DPM) for Target Sound Extraction (TSE), to achieve both cleaner target renderings as well as improved separability from unwanted sounds. The technique also tackles the noise floor of DPM by introducing a correction method for noise schedules and sample steps. This approach is evaluated using both objective and subjective quality metrics on the FSD Kaggle 2018 dataset. The results show that DPM-TSE has a significant improvement in perceived quality in terms of target extraction and purity. Jiarui Hai, Helin Wang, Dongchao Yang, Karan Thakkar, Najim Dehak, Mounya Elhilali |
ICASSP | 6 |
| 2024 | Biomimetic Mappings for Active Sonar Object Recognition in ClutterabstractSONAR technology plays a pivotal role in terrain exploration and specifically identification of objects of interest. However, it grapples with a recurring challenge of clutter and noise which limits the performance of target recognition models. The challenge of noisy observations renders the choice of robust signal representations critical. Inspired by mammalian representation in the midbrain of echolocating bats, the present study evaluates the robustness of a decomposition of echo measurements that matches the statistics of natural vocalizations. This representation is contrasted with equally rich generic mappings as well as digital sonar images based on time-frequency representations. The study shows the clear advantage of the naturally optimized representation for object recognition in presence of background noise and clutter, and further underscores the potential of bio-inspired approaches in advancing SONAR technology. Sangwook Park 0002, Angeles Salles, Kathryne Allen, Cynthia F. Moss, Mounya Elhilali |
ICASSP | 5 |
| 2024 | Investigating Self-Supervised Deep Representations for EEG-Based Auditory Attention DecodingabstractAuditory Attention Decoding (AAD) algorithms play a crucial role in isolating desired sound sources within challenging acoustic environments directly from brain activity. Although recent research has shown promise in AAD using shallow representations such as auditory envelope and spectrogram, there has been limited exploration of deep Self-Supervised (SS) representations on a larger scale. In this study, we undertake a comprehensive investigation into the performance of linear decoders across 12 deep and 2 shallow representations, applied to EEG data from multiple studies spanning 57 subjects and multiple languages. Our experimental results consistently reveal the superiority of deep features for AAD at decoding background speakers, regardless of the datasets and analysis windows. This result indicates possible nonlinear encoding of unattended signals in the brain that are revealed using deep nonlinear features. Additionally, we analyze the impact of different layers of SS representations and window sizes on AAD performance. These findings underscore the potential for enhancing EEG-based AAD systems through the integration of deep feature representations. Karan Thakkar, Jiarui Hai, Mounya Elhilali |
ICASSP | 3 |
| 2024 | DreamVoice: Text-Guided Voice Conversion
Jiarui Hai, Karan Thakkar, Helin Wang, Zengyi Qin, Mounya Elhilali |
INTERSPEECH | 5 |
| 2023 | Boosting Modality Representation With Pre-Trained Models and Multi-Task Training for Multimodal Sentiment AnalysisabstractSentiment analysis has traditionally leveraged information from text data. More recently, it has become increasingly clear that multimodal data provides a rich space to drastically boost interpretation of human sentiments by harnessing information across multiple modalities. In this study, we incorporate pre-trained feature extractors and propose a multitask training strategy to improve modality representations for Multimodal Sentiment Analysis (MSA). The experimental results on the CH-SIMS v2 dataset demonstrate the superior performance of the proposed system compared to existing state-of-the-art methods, validating the effectiveness of our proposed approach. Furthermore, our framework reduces reliance on textual data, achieving competitive outcomes even when utilizing only auditory and visual modalities. Jiarui Hai, Yu-Jeh Liu, Mounya Elhilali |
ASRU | 3 |
| 2023 | Cross-Referencing Self-Training Network for Sound Event Detection in Audio MixturesabstractSound event detection is an important facet of audio tagging that aims to identify sounds of interest and define both the sound category and time boundaries for each sound event in a continuous recording. With advances in deep neural networks, there has been tremendous improvement in the performance of sound event detection systems, although at the expense of costly data collection and labeling efforts. In fact, current state-of-the-art methods employ supervised training methods that leverage large amounts of data samples and corresponding labels in order to facilitate identification of sound category and time stamps of events. As an alternative, the current study proposes a semi-supervised method for generating pseudo-labels from unsupervised data using a student-teacher scheme that balances self-training and cross-training. Additionally, this paper explores post-processing which extracts sound intervals from network prediction, for further improvement in sound event detection performance. The proposed approach is evaluated on sound event detection task for the DCASE2020 challenge. The results of these methods on both "validation" and "public evaluation" sets of DESED database show significant improvement compared to the state-of-the art systems in semi-supervised learning. Sangwook Park 0002, David K. Han, Mounya Elhilali |
IEEE Trans. Multim. | 3 |
| 2022 | Temporal Contrastive-Loss for Audio Event DetectionabstractTemporal coherence is a feature-binding mechanism that ensures features that evolve together in time belong to the same object or event. Coherence has been extensively studied in biological systems, demonstrating how our brain leverages this mechanism to perform complex tasks in real environments and facilitate segregation of complex sensory signals (or wholes) into individual objects (or parts), following Gestalt principles. Although intuitive and computationally tractable, these concepts have rarely been leveraged in audio technologies. Audio event detection is an application that specifically deals with identifying sound events in an audio recording; hence is a natural avenue to explore principles of temporal coherence. In this study, we propose coherence-based learning, formulated as a contrastive loss, to train event detection models whereby embeddings driven by acoustic events are coherently constrained to maximize discriminability across events. This approach results in improved detection performance with no additional computational cost and a very small overhead during the training procedure. Sandeep Kothinti, Mounya Elhilali |
ICASSP | 2 |
| 2022 | Time-Balanced Focal Loss for Audio Event DetectionabstractSound Event Detection (SED) tackles the challenge of identifying sound events in an audio recording by delimiting both their temporal boundaries as well as sound category. With recent advances in deep learning, current systems are able to leverage availability of large datasets to train sophisticated and highly effective SED models. Nonetheless, sound sources and acoustic characteristics of different classes vary greatly in their prevalence as well as representation in labeled datasets. The challenge with data imbalance in the case of SED stems not only from the representation (number of samples) across classes but also the natural asymmetry in time duration across different events varying from short transient events such as the clacking of dishes to more sustained events such as vacuuming. This variability results in an inherent disproportional representation of effective training samples. To address this compounded imbalance issue, this work proposes a balanced focal learning function that introduces a novel time-sensitive classwise weight. The proposed loss is applied to SED in the context of DCASE2021 challenge, and reports a notable improvement over the baseline, particularly in the case of shorter sound events. Sangwook Park 0002, Mounya Elhilali |
ICASSP | 2 |
| 2022 | Temporal coding with magnitude-phase regularization for sound event detection
Sangwook Park 0002, Sandeep Kothinti, Mounya Elhilali |
INTERSPEECH | 3 |
| 2021 | Self-Training for Sound Event Detection in Audio MixturesabstractSound event detection (SED) takes on the task of identifying presence of specific sound events in a complex audio recording. SED has tremendous implications in video analytics, smart speaker algorithms and audio tagging. Recent advances in deep learning have afforded remarkable advances in performance of SED systems; albeit at the cost of extensive labeling efforts to train supervised methods using fully described sound class labels and timestamps. In order to address limitations in availability of training data, this work proposes a self-training technique to leverage unlabeled datasets in supervised learning using pseudo label estimation. This approach proposes a dual-term objective function: a classification loss for the original labels and expectation loss for pseudo labels. The proposed self training technique is applied to sound event detection in the context of the DCASE 2020 challenge, and reports a notable improvement over the baseline system for this task. The self-training approach is particularly effective in extending the labeled database with concurrent sound events. Sangwook Park 0002, Ashwin Bellur, David K. Han, Mounya Elhilali |
ICASSP | 4 |
| 2021 | Adaptive Listening to Everyday Soundscapes
Mounya Elhilali |
Interspeech | 1 |
| 2021 | Design and Comparative Performance of a Robust Lung Auscultation System for Noisy Clinical SettingsabstractChest auscultation is a widely used clinical tool for respiratory disease detection. The stethoscope has undergone a number of transformative enhancements since its invention, including the introduction of electronic systems in the last two decades. Nevertheless, stethoscopes remain riddled with a number of issues that limit their signal quality and diagnostic capability, rendering both traditional and electronic stethoscopes unusable in noisy or non-traditional environments (e.g., emergency rooms, rural clinics, ambulatory vehicles). This work outlines the design and validation of an advanced electronic stethoscope that dramatically reduces external noise contamination through hardware redesign and real-time, dynamic signal processing. The proposed system takes advantage of an acoustic sensor array, an external facing microphone, and on-board processing to perform adaptive noise suppression. The proposed system is objectively compared to six commercially-available acoustic and electronic devices in varying levels of simulated noisy clinical settings and quantified using two metrics that reflect perceptual audibility and statistical similarity, normalized covariance measure (NCM) and magnitude squared coherence (MSC). The analyses highlight the major limitations of current stethoscopes and the significant improvements the proposed system makes in challenging settings by minimizing both distortion of lung sounds and contamination by ambient noise. Ian McLane, Dimitra Emmanouilidou, James E. West, Mounya Elhilali |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | Electronic Stethoscope Filtering Mimics the Perceived Sound Characteristics of Acoustic StethoscopeabstractElectronic stethoscopes offer several advantages over conventional acoustic stethoscopes, including noise reduction, increased amplification, and ability to store and transmit sounds. However, the acoustical characteristics of electronic and acoustic stethoscopes can differ significantly, introducing a barrier for clinicians to transition to electronic stethoscopes. This work proposes a method to process lung sounds recorded by an electronic stethoscope, such that the sounds are perceived to have been captured by an acoustic stethoscope. The proposed method calculates an electronic-to-acoustic stethoscope filter by measuring the difference between the average frequency responses of an acoustic and an electronic stethoscope to multiple lung sounds. To validate the method, a change detection experiment was conducted with 51 medical professionals to compare filtered electronic, unfiltered electronic, and acoustic stethoscope lung sounds. Participants were asked to detect when transitions occurred in sounds comprising several sections of the three types of recordings. Transitions between the filtered electronic and acoustic stethoscope sections were detected, on average, by chance (sensitivity index equal to zero) and also detected significantly less than transitions between the unfiltered electronic and acoustic stethoscope sections ( ), demonstrating the effectiveness of the method to filter electronic stethoscopes to mimic an acoustic stethoscope. This processing could incentivize clinicians to adopt electronic stethoscopes by providing a means to shift between the sound characteristics of acoustic and electronic stethoscopes in a single device, allowing for a faster transition to new technology and greater appreciation for the electronic sound quality. Valerie E. Rennoll, Ian McLane, Dimitra Emmanouilidou, James E. West, Mounya Elhilali |
IEEE J. Biomed. Health Informatics | 5 |
| 2020 | Bio-Mimetic Attentional Feedback in Music Source SeparationabstractAttention plays a vital role in helping us navigate our acoustic surroundings. It guides sensory processing to sift through the cacophony of sounds in everyday scenes and modulates the representation of targets sounds relative to distractors. While its conceptual role is well established, there are competing theories as to how attentional feedback operates in the brain and how its mechanistic underpinnings can be incorporated into computational systems. These interpretations differ in the manner in which attentional feedback operates as an information bottleneck to aid perception. One interpretation is that attention adapts the sensory mapping itself to encode only the target cues. An alternative interpretation is that attention behaves as a gain modulator that enhances the target cues after they are encoded. Further, the theory of temporal coherence states that attention seeks to bind temporally coherent features relative to anchor features as determined by prior knowledge of target objects. In this work, we study these competing theories within a deep-network framework for the task of music source separation. We show that these theories complement each other, and when employed together, yield state of the art performance in music source separation. We further show that systems with attentional mechanisms can be made to scale to mismatched conditions by retuning only the attentional modules with minimal data. Ashwin Bellur, Mounya Elhilali |
ICASSP | 2 |
| 2020 | Synthesizing Engaging Music Using Dynamic Models of Statistical SurprisalabstractSynthesis of music content generally leverages the underlying statistical structure of music to develop generative models, able to create new musical expressions within the same genre. In this work, we explore the statistical structure of a musical corpus and its effect on modulating the attention of listeners. The study specifically explores listeners' engagement to newly synthesized music and tests the hypothesis that maximizing statistical surprisal would result in increased auditory salience. The study employs a dynamical statistical model to estimate melodic line surprisal and develops an optimization procedure using parametrized codebooks to synthesize musical segments that maximize statistical surprisal. A behavioral experiment with a dichotic listening task is designed to probe salience of the synthesized melodies against original melodies by measuring listeners' engagement in a continuous-fashion. Results indicate that we can control the salience of sounds by manipulating the statistical surprisal, guided by the complexity of the temporal structure of the musical corpus. This work suggests that future work in automated music synthesis could leverage statistical models of music beyond musical aesthetics to also manipulate the degree of engagement. Sandeep Kothinti, Benjamin Skerritt-Davis, Aditya Nair, Mounya Elhilali |
ICASSP | 4 |
| 2020 | Ensemble modeling of auditory streaming reveals potential sources of bistability across the perceptual hierarchyabstractPerceptual bistability-the spontaneous, irregular fluctuation of perception between two interpretations of a stimulus-occurs when observing a large variety of ambiguous stimulus configurations. This phenomenon has the potential to serve as a tool for, among other things, understanding how function varies across individuals due to the large individual differences that manifest during perceptual bistability. Yet it remains difficult to interpret the functional processes at work, without knowing where bistability arises during perception. In this study we explore the hypothesis that bistability originates from multiple sources distributed across the perceptual hierarchy. We develop a hierarchical model of auditory processing comprised of three distinct levels: a Peripheral, tonotopic analysis, a Central analysis computing features found more centrally in the auditory system, and an Object analysis, where sounds are segmented into different streams. We model bistable perception within this system by applying adaptation, inhibition and noise into one or all of the three levels of the hierarchy. We evaluate a large ensemble of variations of this hierarchical model, where each model has a different configuration of adaptation, inhibition and noise. This approach avoids the assumption that a single configuration must be invoked to explain the data. Each model is evaluated based on its ability to replicate two hallmarks of bistability during auditory streaming: the selectivity of bistability to specific stimulus configurations, and the characteristic log-normal pattern of perceptual switches. Consistent with a distributed origin, a broad range of model parameters across this hierarchy lead to a plausible form of perceptual bistability. David Frank Little, Joel S. Snyder, Mounya Elhilali |
PLoS Comput. Biol. | 3 |
| 2020 | Amphibian Sounds Generating Network Based on Adversarial LearningabstractThis letter proposes a generative network based on adversarial learning for synthesizing short-time audio streams and investigates the effectiveness of data augmentation for amphibian call sounds classification. Based on Fourier analysis, the generator is designed by a multi-layer perceptron composed of frequency basis learning layers and an output layer, and a discriminator is constructed by a convolutional neural network. Additionally, regularization on weights is introduced to train the networks with practical data that includes some disturbances. Synthetic audio streams are evaluated by quantitative comparison using inception score, and classification results are compared for real versus synthetic data. In conclusion, the proposed generative network is shown to produce realistic sounds and therefore useful for data augmentation. Sangwook Park 0002, Mounya Elhilali, David K. Han, Hanseok Ko |
IEEE Signal Process. Lett. | 2 |
| 2020 | Audio Object Classification Using Distributed Beliefs and AttentionabstractOne of the unique characteristics of human hearing is its ability to recognize acoustic objects even in presence of severe noise and distortions. In this work, we explore two mechanisms underlying this ability: 1) redundant mapping of acoustic waveforms along distributed latent representations and 2) adaptive feedback based on prior knowledge to selectively attend to targets of interest. We propose a bio-mimetic account of acoustic object classification by developing a novel distributed deep belief network validated for the task of robust acoustic object classification using the UrbanSound database. The proposed distributed belief network (DBN) encompasses an array of independent sub-networks trained generatively to capture different abstractions of natural sounds. A supervised classifier then performs a readout of this distributed mapping. The overall architecture not only matches the state of the art system for acoustic object classification but leads to significant improvement over the baseline in mismatched noisy conditions (31.4% relative improvement in 0dB conditions). Furthermore, we incorporate mechanisms of attentional feedback that allows the DBN to deploy local memories of sounds targets estimated at multiple views to bias network activation when attending to a particular object. This adaptive feedback results in further improvement of object classification in unseen noise conditions (relative improvement of 54% over the baseline in 0dB conditions). Ashwin Bellur, Mounya Elhilali |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Joint Acoustic and Class Inference for Weakly Supervised Sound Event DetectionabstractSound event detection is a challenging task, especially for scenes with multiple simultaneous events. While event classification methods tend to be fairly accurate, event localization presents additional challenges, especially when large amounts of labeled data are not available. Task4 of the 2018 DCASE challenge presents an event detection task that requires accuracy in both segmentation and recognition of events while providing only weakly labeled training data. Supervised methods can produce accurate event labels but are limited in event segmentation when training data lacks event timestamps. On the other hand, unsupervised methods that model the acoustic properties of the audio can produce accurate event boundaries but are not guided by the characteristics of event classes and sound categories. We present a hybrid approach that combines an acoustic-driven event boundary detection and a supervised label inference using a deep neural network. This framework leverages benefits of both unsupervised and supervised methodologies and takes advantage of large amounts of unlabeled data, making it ideal for large-scale weakly la-beled event detection. Compared to a baseline system, the proposed approach delivers a 15% absolute improvement in F-score, demonstrating the benefits of the hybrid bottom-up, top-down approach. Sandeep Kothinti, Keisuke Imoto, Debmalya Chakrabarty, Gregory Sell, Shinji Watanabe 0001, Mounya Elhilali |
ICASSP | 6 |
| 2019 | A Study of a Cross-Language Perception Based on Cortical Analysis Using Biomimetic STRFs
Sangwook Park 0002, David K. Han, Mounya Elhilali |
INTERSPEECH | 3 |
| 2019 | A Gestalt inference model for auditory scene segregationabstractOur current understanding of how the brain segregates auditory scenes into meaningful objects is in line with a Gestaltism framework. These Gestalt principles suggest a theory of how different attributes of the soundscape are extracted then bound together into separate groups that reflect different objects or streams present in the scene. These cues are thought to reflect the underlying statistical structure of natural sounds in a similar way that statistics of natural images are closely linked to the principles that guide figure-ground segregation and object segmentation in vision. In the present study, we leverage inference in stochastic neural networks to learn emergent grouping cues directly from natural soundscapes including speech, music and sounds in nature. The model learns a hierarchy of local and global spectro-temporal attributes reminiscent of simultaneous and sequential Gestalt cues that underlie the organization of auditory scenes. These mappings operate at multiple time scales to analyze an incoming complex scene and are then fused using a Hebbian network that binds together coherent features into perceptually-segregated auditory objects. The proposed architecture successfully emulates a wide range of well established auditory scene segregation phenomena and quantifies the complimentary role of segregation and binding cues in driving auditory scene segregation. Debmalya Chakrabarty, Mounya Elhilali |
PLoS Comput. Biol. | 2 |
| 2018 | Sensory Mapping Adaptation Under Multiple Task ScenariosabstractDemands on auditory perception change constantly with natural changes in everyday acoustic environments. Mechanisms, such as attentional feedback, direct the brain to adapt processing of the incoming signal to maximize its ability to detect the presence of a sound of interest or enhance its representation. These top-down feedback processes induce adaptation of the spectrotemporal representation of incoming sounds in a manner that enhances our ability to perform the desired task. In this work, we propose a computational model to implement and study sensory mapping adaptation under different task-demands. We propose a common processing framework to examine how sensory mapping adaptation manifests under different task -driven conditions like speech enhancement and robust speech activity detection. Objective measures of speech enhancement and discrimination are used to quantify the impact of the adaptation under different contexts and its impact on performance outcomes. Ashwin Bellur, Mounya Elhilali |
ICASSP | 2 |
| 2018 | Detecting change in stochastic sound sequencesabstractOur ability to parse our acoustic environment relies on the brain's capacity to extract statistical regularities from surrounding sounds. Previous work in regularity extraction has predominantly focused on the brain's sensitivity to predictable patterns in sound sequences. However, natural sound environments are rarely completely predictable, often containing some level of randomness, yet the brain is able to effectively interpret its surroundings by extracting useful information from stochastic sounds. It has been previously shown that the brain is sensitive to the marginal lower-order statistics of sound sequences (i.e., mean and variance). In this work, we investigate the brain's sensitivity to higher-order statistics describing temporal dependencies between sound events through a series of change detection experiments, where listeners are asked to detect changes in randomness in the pitch of tone sequences. Behavioral data indicate listeners collect statistical estimates to process incoming sounds, and a perceptual model based on Bayesian inference shows a capacity in the brain to track higher-order statistics. Further analysis of individual subjects' behavior indicates an important role of perceptual constraints in listeners' ability to track these sensory statistics with high fidelity. In addition, the inference model facilitates analysis of neural electroencephalography (EEG) responses, anchoring the analysis relative to the statistics of each stochastic stimulus. This reveals both a deviance response and a change-related disruption in phase of the stimulus-locked response that follow the higher-order statistics. These results shed light on the brain's ability to process stochastic sound sequences. Benjamin Skerritt-Davis, Mounya Elhilali |
PLoS Comput. Biol. | 2 |
| 2017 | Feedback-Driven Sensory Mapping Adaptation for Robust Speech Activity DetectionabstractParsing natural acoustic scenes using computational methodologies poses many challenges. Given the rich and complex nature of the acoustic environment, data mismatch between train and test conditions is a major hurdle in data-driven audio processing systems. In contrast, the brain exhibits a remarkable ability at segmenting acoustic scenes with relative ease. When tackling challenging listening conditions that are often faced in everyday life, the biological system relies on a number of principles that allow it to effortlessly parse its rich soundscape. In the current study, we leverage a key principle employed by the auditory system: its ability to adapt the neural representation of its sensory input in a high-dimensional space. We propose a framework that mimics this process in a computational model for robust speech activity detection. The system employs a 2-D Gabor filter bank whose parameters are retuned offline to improve the separability between the feature representation of speech and nonspeech sounds. This retuning process, driven by feedback from statistical models of speech and nonspeech classes, attempts to minimize the misclassification risk of mismatched data, with respect to the original statistical models. We hypothesize that this risk minimization procedure results in an emphasis of unique speech and nonspeech modulations in the high-dimensional space. We show that such an adapted system is indeed robust to other novel conditions, with a marked reduction in equal error rates for a variety of databases with additive and convolutive noise distortions. We discuss the lessons learned from biology with regard to adapting to an ever-changing acoustic environment and the impact on building truly intelligent audio processing systems. Ashwin Bellur, Mounya Elhilali |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Abnormal sound event detection using temporal trajectories mixturesabstractDetection of anomalous sound events in audio surveillance is a challenging task when applied to realistic settings. Part of the difficulty stems from properly defining the `normal' behavior of a crowd or an environment (e.g. airport, train station, sport field). By successfully capturing the heterogeneous nature of sound events in an acoustic environment, we can use it as a reference against which anomalous behavior can be detected in continuous audio recordings. The current study proposes a methodology for representing sound classes using a hierarchical network of convolutional features and mixture of temporal trajectories (MTT). The framework couples unsupervised and supervised learning and provides a robust scheme for detection of abnormal sound events in a subway station. The results reveal the strength of the proposed representation in capturing non-trivial commonalities within a single sound class and variabilities across different sound classes as well as high degree of robustness in noise. Debmalya Chakrabarty, Mounya Elhilali |
ICASSP | 2 |
| 2015 | A Framework for Speech Activity Detection Using Adaptive Auditory Receptive FieldsabstractOne of the hallmarks of sound processing in the brain is the ability of the nervous system to adapt to changing behavioral demands and surrounding soundscapes. It can dynamically shift sensory and cognitive resources to focus on relevant sounds. Neurophysiological studies indicate that this ability is supported by adaptively retuning the shapes of cortical spectro-temporal receptive fields (STRFs) to enhance features of target sounds while suppressing those of task-irrelevant distractors. Because an important component of human communication is the ability of a listener to dynamically track speech in noisy environments, the solution obtained by auditory neurophysiology implies a useful adaptation strategy for speech activity detection (SAD). SAD is an important first step in a number of automated speech processing systems, and performance is often reduced in highly noisy environments. In this paper, we describe how task-driven adaptation is induced in an ensemble of neurophysiological STRFs, and show how speech-adapted STRFs reorient themselves to enhance spectro-temporal modulations of speech while suppressing those associated with a variety of nonspeech sounds. We then show how an adapted ensemble of STRFs can better detect speech in unseen noisy environments compared to an unadapted ensemble and a noise-robust baseline. Finally, we use a stimulus reconstruction task to demonstrate how the adapted STRF ensemble better captures the spectrotemporal modulations of attended speech in clean and noisy conditions. Our results suggest that a biologically plausible adaptation framework can be applied to speech processing systems to dynamically adapt feature representations for improving noise robustness. Michael A. Carlin, Mounya Elhilali |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Segregating Complex Sound Sources through Temporal CoherenceabstractA new approach for the segregation of monaural sound mixtures is presented based on the principle of temporal coherence and using auditory cortical representations. Temporal coherence is the notion that perceived sources emit coherently modulated features that evoke highly-coincident neural response patterns. By clustering the feature channels with coincident responses and reconstructing their input, one may segregate the underlying source from the simultaneously interfering signals that are uncorrelated with it. The proposed algorithm requires no prior information or training on the sources. It can, however, gracefully incorporate cognitive functions and influences such as memories of a target source or attention to a specific set of its attributes so as to segregate it from its background. Aside from its unusual structure and computational innovations, the proposed model provides testable hypotheses of the physiological mechanisms of this ubiquitous and remarkable perceptual ability, and of its psychophysical manifestations in navigating complex sensory environments. Lakshmi Krishnan, Mounya Elhilali, Shihab A. Shamma |
PLoS Comput. Biol. | 2 |
| 2013 | Task-driven attentional mechanisms for auditory scene recognitionabstractHow do humans attend to and pick out relevant auditory objects amongst all other sounds in the environment? Based on neurophysiological findings we propose two task oriented attentional mechanisms acting as Bayesian priors which act on two separate levels of processing: a sensory mapping stage and object representation stage. The former sensory stage is modeled as a high dimensional mapping which captures the spectrotemporal nuances and cues of auditory objects. The latter object representation stage then captures the statistical distribution of the different classes of acoustic scenes. This scheme shows a relative improvement in performance by 81% compared to a baseline system. Kailash Patil, Mounya Elhilali |
ICASSP | 2 |
| 2013 | Sustained Firing of Model Central Auditory Neurons Yields a Discriminative Spectro-temporal Representation for Natural SoundsabstractThe processing characteristics of neurons in the central auditory system are directly shaped by and reflect the statistics of natural acoustic environments, but the principles that govern the relationship between natural sound ensembles and observed responses in neurophysiological studies remain unclear. In particular, accumulating evidence suggests the presence of a code based on sustained neural firing rates, where central auditory neurons exhibit strong, persistent responses to their preferred stimuli. Such a strategy can indicate the presence of ongoing sounds, is involved in parsing complex auditory scenes, and may play a role in matching neural dynamics to varying time scales in acoustic signals. In this paper, we describe a computational framework for exploring the influence of a code based on sustained firing rates on the shape of the spectro-temporal receptive field (STRF), a linear kernel that maps a spectro-temporal acoustic stimulus to the instantaneous firing rate of a central auditory neuron. We demonstrate the emergence of richly structured STRFs that capture the structure of natural sounds over a wide range of timescales, and show how the emergent ensembles resemble those commonly reported in physiological studies. Furthermore, we compare ensembles that optimize a sustained firing code with one that optimizes a sparse code, another widely considered coding strategy, and suggest how the resulting population responses are not mutually exclusive. Finally, we demonstrate how the emergent ensembles contour the high-energy spectro-temporal modulations of natural sounds, forming a discriminative representation that captures the full range of modulation statistics that characterize natural sound ensembles. These findings have direct implications for our understanding of how sensory systems encode the informative components of natural stimuli and potentially facilitate multi-sensory integration. Michael A. Carlin, Mounya Elhilali |
PLoS Comput. Biol. | 2 |
| 2013 | A Multistream Feature Framework Based on Bandpass Modulation Filtering for Robust Speech RecognitionabstractThere is strong neurophysiological evidence suggesting that processing of speech signals in the brain happens along parallel paths which encode complementary information in the signal. These parallel streams are organized around a duality of slow vs. fast: Coarse signal dynamics appear to be processed separately from rapidly changing modulations both in the spectral and temporal dimensions. We adapt such duality in a multistream framework for robust speaker-independent phoneme recognition. The scheme presented here centers around a multi-path bandpass modulation analysis of speech sounds with each stream covering an entire range of temporal and spectral modulations. By performing bandpass operations along the spectral and temporal dimensions, the proposed scheme avoids the classic feature explosion problem of previous multistream approaches while maintaining the advantage of parallelism and localized feature analysis. The proposed architecture results in substantial improvements over standard and state-of-the-art feature schemes for phoneme recognition, particularly in presence of nonstationary noise, reverberation and channel distortions. Sridhar Krishna Nemala, Kailash Patil, Mounya Elhilali |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | The UMD-JHU 2011 speaker recognition systemabstractIn recent years, there have been significant advances in the field of speaker recognition that has resulted in very robust recognition systems. The primary focus of many recent developments have shifted to the problem of recognizing speakers in adverse conditions, e.g in the presence of noise/reverberation. In this paper, we present the UMD-JHU speaker recognition system applied on the NIST 2010 SRE task. The novel aspects of our systems are: 1) Improved performance on trials involving different vocal effort via the use of linear-scale features; 2) Expected improved recognition performance in the presence of reverberation and noise via the use of frequency domain perceptual linear predictor and cortical features; 3) A new discriminative kernel partial least squares (KPLS) framework that complements state-of-the-art back-end systems JFA and PLDA to aid in better overall recognition; and 4) Acceleration of JFA, PLDA and KPLS back-ends via distributed computing. The individual components of the system and the fused system are compared against a baseline JFA system and results reported by SRI and MIT-LL on SRE2010. Daniel Garcia-Romero, Xinhui Zhou, Dmitry N. Zotkin, Balaji Vasan Srinivasan, Yuancheng Luo, Sriram Ganapathy, Samuel Thomas 0001, Sridhar Krishna Nemala, Garimella S. V. S. Sivaram, Majid Mirbagheri, Sri Harish Reddy Mallidi, Thomas Janu, Padmanabhan Rajan, Nima Mesgarani, Mounya Elhilali, Hynek Hermansky, Shihab A. Shamma, Ramani Duraiswami |
ICASSP | 15 |
| 2012 | Multilevel speech intelligibility for robust speaker recognitionabstractIn the real world, natural conversational speech is an amalgam of speech segments, silences and environmental/ background and channel effects. Labeling the different regions of an acoustic signal according to their information levels would greatly benefit all automatic speech processing tasks. In the current work, we propose a novel segmentation approach based on a perception-based measure of speech intelligibility. Unlike segmentation approaches based on various forms of voice-activity detection (VAD), the proposed parsing approach exploits higher-level perceptual information about signal intelligibility levels. This labeling information is integrated into a novel multilevel framework for automatic speaker recognition task. The system processes the input acoustic signal along independent streams reflecting various levels of intelligibility and then fusing the decision scores from the multiple steams according to their intelligibility contribution. Our results show that the proposed system achieves significant improvements over standard baseline and VAD-based approaches, and attains a performance similar to the one obtained with oracle speech segmentation information. Sridhar Krishna Nemala, Mounya Elhilali |
ICASSP | 2 |
| 2012 | A model of attention-driven scene analysisabstractParsing complex acoustic scenes involves an intricate interplay between bottom-up, stimulus-driven salient elements in the scene with top-down, goal-directed, mechanisms that shift our attention to particular parts of the scene. Here, we present a framework for exploring the interaction between these two processes in a simulated cocktail party setting. The model shows improved digit recognition in a multi-talker environment with a goal of tracking the source uttering the highest value. This work highlights the relevance of both data-driven and goal-driven processes in tackling real multi-talker, multi-source sound analysis. Malcolm Slaney, Trevor Agus, Shih-Chii Liu, Emine Merve Kaya, Mounya Elhilali |
ICASSP | 5 |
| 2012 | Robust phoneme recognition based on biomimetic speech contours
Michael A. Carlin, Kailash Patil, Sridhar Krishna Nemala, Mounya Elhilali |
INTERSPEECH | 4 |
| 2012 | Goal-Oriented Auditory Scene Recognition
Kailash Patil, Mounya Elhilali |
INTERSPEECH | 2 |
| 2012 | Music in Our Ears: The Biological Bases of Musical Timbre PerceptionabstractTimbre is the attribute of sound that allows humans and other animals to distinguish among different sound sources. Studies based on psychophysical judgments of musical timbre, ecological analyses of sound's physical characteristics as well as machine learning approaches have all suggested that timbre is a multifaceted attribute that invokes both spectral and temporal sound features. Here, we explored the neural underpinnings of musical timbre. We used a neuro-computational framework based on spectro-temporal receptive fields, recorded from over a thousand neurons in the mammalian primary auditory cortex as well as from simulated cortical neurons, augmented with a nonlinear classifier. The model was able to perform robust instrument classification irrespective of pitch and playing style, with an accuracy of 98.7%. Using the same front end, the model was also able to reproduce perceptual distance judgments between timbres as perceived by human listeners. The study demonstrates that joint spectro-temporal features, such as those observed in the mammalian primary auditory cortex, are critical to provide the rich-enough representation necessary to account for perceptual judgments of timbre by human listeners, as well as recognition of musical instruments. Kailash Patil, Daniel Pressnitzer, Shihab A. Shamma, Mounya Elhilali |
PLoS Comput. Biol. | 4 |
| 2011 | Multistream Bandpass Modulation Features for Robust Speech Recognition
Sridhar Krishna Nemala, Kailash Patil, Mounya Elhilali |
INTERSPEECH | 3 |
| 2010 | A joint acoustic and phonological approach to speech intelligibility assessmentabstractWhile current models of speech intelligibility rely on intricate acoustic analyses of speech attributes, they are limited by the lack of any linguistic information; hence failing to capture natural variability of speech sounds and confining their applicability to average intelligibility assessments. Another important limitation is that the existing models rely on the use of reference clean speech templates (or average profiles). In this work, we propose a novel approach to speech intelligibility by combining a biologically-inspired acoustic analysis of peripheral and cortical processing with phonological statistical models of speech using a hybrid GMM-SVM system. The model results in a novel scheme for speech intelligibility assessment without the use of reference clean speech templates, and the model predictions strongly correlate with scores obtained from human listeners under a variety of realistic listening environments. We further show that the proposed model enables local level tracking of intelligibility and also generalizes well to multiple speech corpora. Sridhar Krishna Nemala, Mounya Elhilali |
ICASSP | 2 |
| 2010 | Sparse coding for speech recognitionabstractThis paper proposes a novel feature extraction technique for speech recognition based on the principles of sparse coding. The idea is to express a spectro-temporal pattern of speech as a linear combination of an overcomplete set of basis functions such that the weights of the linear combination are sparse. These weights (features) are subsequently used for acoustic modeling. We learn a set of overcomplete basis functions (dictionary) from the training set by adopting a previously proposed algorithm which iteratively minimizes the reconstruction error and maximizes the sparsity of weights. Furthermore, features are derived using the learned basis functions by applying the well established principles of compressive sensing. Phoneme recognition experiments show that the proposed features outperform the conventional features in both clean and noisy conditions. Garimella S. V. S. Sivaram, Sridhar Krishna Nemala, Mounya Elhilali, Trac D. Tran, Hynek Hermansky |
ICASSP | 3 |
| 2009 | Discriminant spectrotemporal features for phoneme recognitionabstractWe propose discriminant methods for deriving twodimensional spectrotemporal features for phoneme recognition that are estimated to maximize the separation between the representations of phoneme classes. The linearity of the filters results in their intuitive interpretation enabling us to investigate the working principles of the system and to improve its performance by locating the sources of error. Two methods for the estimation of filters are proposed: Regularized Least Square (RLS) and Modified Linear Discriminant Analysis (MLDA). Both methods reach a comparable improvement over the baseline condition demonstrating the advantage of the discriminant spectrotemporal filters. Nima Mesgarani, Garimella S. V. S. Sivaram, Sridhar Krishna Nemala, Mounya Elhilali, Hynek Hermansky |
INTERSPEECH | 4 |
| 2008 | Information-bearing components of speech intelligibility under babble-noise and bandlimiting distortionsabstractPerformance of speech technologies can benefit greatly from a deeper appreciation of the nature of the information- bearing features in continuous speech. To explore these features, we focus here on the role of the spectral and temporal modulations in maintaining the intelligibility of speech as it becomes severely degraded by low-pass filtering and additive babble noise. These modulations are estimated using a biological model of auditory processing which approximates the representation of sound in the cortex. Intelligibility of the noisy speech is computed directly from this model via the spectro-temporal modulation index (STMI), and the validity of this metric is confirmed by a detailed comparison with results of psychoacoustic tests. Our analysis reveals quantitatively why certain types of noise are more disruptive to speech intelligibility than others (e.g., babble vs. white noise). It also highlights the important contribution of both spectral and temporal modulations in accurately predicting the intelligibility of speech under adverse conditions. Mounya Elhilali, Shihab A. Shamma |
ICASSP | 1 |
| 2006 | A Biologically-Inspired Approach to the Cocktail Party ProblemabstractThough seemingly effortless, our auditory system engages in complex processes and transformations which enable us to segregate speech and other sounds in cocktail party settings. This paper presents a computational approach to modelling monaural auditory scene analysis, where we attempt to account for perceptual and neuronal findings of receptive field selectivity and adaptation in the auditory cortex. The model introduces a biologically-inspired scheme of dynamic segregation of auditory streams, based on unsupervised clustering and the statistical theory of Kalman prediction. Our method demonstrates its ability to emulate known percepts reported by human subjects in auditory streaming and sound organization tests, and yields successful results in segregating speech from concurrent speaker and music interferences Mounya Elhilali, Shihab A. Shamma |
ICASSP (5) | 1 |
| 2003 | A spectro-temporal modulation index (STMI) for assessment of speech intelligibility
Mounya Elhilali, Taishih Chi, Shihab A. Shamma |
Speech Commun. | 1 |