Shihab A. Shamma

dblp:02/4347 · DBLP profile ↗
← Back
37ranked-venue papers
0as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 6 since 2021Artificial intelligence and machine learning · 16 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2025 Sparse high-dimensional decomposition of non-primary auditory cortical receptive fields
abstract
Characterizing neuronal responses to natural stimuli remains a central goal in sensory neuroscience. In auditory cortical neurons, the stimulus selectivity of elicited spiking activity is summarized by a spectrotemporal receptive field (STRF) that relates neuronal responses to the stimulus spectrogram. Though effective in characterizing primary auditory cortical responses, STRFs of non-primary auditory neurons can be quite intricate, reflecting their mixed selectivity. The complexity of non-primary STRFs hence impedes understanding how acoustic stimulus representations are transformed along the auditory pathway. Here, we focus on the relationship between ferret primary auditory cortex (A1) and a secondary region, dorsal posterior ectosylvian gyrus (PEG). We propose estimating receptive fields in PEG with respect to a well-established high-dimensional computational model of primary-cortical stimulus representations. These "cortical receptive fields" (CortRF) are estimated greedily to identify the salient primary-cortical features modulating spiking responses and in turn related to corresponding spectrotemporal features. Hence, they provide biologically plausible hierarchical decompositions of STRFs in PEG. Such CortRF analysis was applied to PEG neuronal responses to speech and temporally orthogonal ripple combination (TORC) stimuli and, for comparison, to A1 neuronal responses. CortRFs of PEG neurons captured their selectivity to more complex spectrotemporal features than A1 neurons; moreover, CortRF models were more predictive of PEG (but not A1) responses to speech. Our results thus suggest that secondary-cortical stimulus representations can be computed as sparse combinations of primary-cortical features that facilitate encoding natural stimuli. Thus, by adding the primary-cortical representation, we can account for PEG single-unit responses to natural sounds better than bypassing it and considering as input the auditory spectrogram. These results confirm with explicit details the presumed hierarchical organization of the auditory cortex.
Shoutik Mukherjee, Behtash Babadi, Shihab A. Shamma
PLoS Comput. Biol.3
2023 Investigating the cortical tracking of speech and music with sung speech
abstract
International audience
Giorgia Cantisani, Amirhossein Chalehchaleh, Giovanni M. Di Liberto, Shihab A. Shamma
INTERSPEECH4
2023 Learning to Compute the Articulatory Representations of Speech with the MIRRORNET
Yashish M. Siriwardena, Carol Y. Espy-Wilson, Shihab A. Shamma
INTERSPEECH3
2022 Harmonicity Plays a Critical Role in DNN Based Versus in Biologically-Inspired Monaural Speech Segregation Systems
abstract
Recent advancements in deep learning have led to drastic improvements in speech segregation models. Despite their success and growing applicability, few efforts have been made to analyze the underlying principles that these networks learn to perform segregation. Here we analyze the role of harmonicity on two state-of-the-art Deep Neural Networks (DNN)-based models- Conv-TasNet and DPT-Net [1],[2]. We evaluate their performance with mixtures of natural speech versus slightly manipulated inharmonic speech, where harmonics are slightly frequency jittered. We find that performance deteriorates significantly if one source is even slightly harmonically jittered, e.g., an imperceptible 3% harmonic jitter degrades performance of Conv-TasNet from 15.4 dB to 0.70 dB. Training the model on inharmonic speech does not remedy this sensitivity, instead resulting in worse performance on natural speech mixtures, making inharmonicity a powerful adversarial factor in DNN models. Furthermore, additional analyses reveal that DNN algorithms deviate markedly from biologically inspired algorithms [3] that rely primarily on timing cues and not harmonicity to segregate speech.
Rahil Parikh, Ilya Kavalerov, Carol Y. Espy-Wilson, Shihab A. Shamma
ICASSP4
2022 The Mirrornet : Learning Audio Synthesizer Controls Inspired by Sensorimotor Interaction
abstract
Experiments to understand the sensorimotor neural interactions in the human cortical speech system support the existence of a bidirectional flow of interactions between the auditory and motor regions. Their key function is to enable the brain to ‘learn’ how to control the vocal tract for speech production. This idea is the impetus for the recently proposed "MirrorNet", a constrained autoencoder architecture. In this paper, the MirrorNet is applied to learn, in an unsupervised manner, the controls of a specific audio synthesizer (DIVA) to produce melodies only from their auditory spectrograms. The results demonstrate how the MirrorNet discovers the synthesizer parameters to generate the melodies that closely resemble the original and those of unseen melodies, and even determine the best set parameters to approximate renditions of complex piano melodies generated by a different synthesizer. This generalizability of the MirrorNet illustrates its potential to discover from sensory data the controls of arbitrary motor-plants.
Yashish M. Siriwardena, Guilhem Marion, Shihab A. Shamma
ICASSP3
2022 An Empirical Analysis on the Vulnerabilities of End-to-End Speech Segregation Models
abstract
International audience
Rahil Parikh, Gaspar Rochette, Carol Y. Espy-Wilson, Shihab A. Shamma
INTERSPEECH4
2022 Acoustic To Articulatory Speech Inversion Using Multi-Resolution Spectro-Temporal Representations Of Speech Signals
abstract
Multi-resolution spectro-temporal features of a speech signal represent how the brain perceives sounds by tuning cortical cells to different spectral and temporal modulations. These features produce a higher dimensional representation of the speech signals. The purpose of this paper is to evaluate how well the auditory cortex representation of speech signals contribute to estimate articulatory features of those corresponding signals. Since obtaining articulatory features from acoustic features of speech signals has been a challenging topic of interest for different speech communities, we investigate the possibility of using this multi-resolution representation of speech signals as acoustic features. We used U. of Wisconsin X-ray Microbeam (XRMB) database of clean speech signals to train a feed-forward deep neural network (DNN) to estimate articulatory trajectories of six tract variables. The optimal set of multi-resolution spectro-temporal features to train the model were chosen using appropriate scale and rate vector parameters to obtain the best performing model. Experiments achieved a correlation of 0.675 with ground-truth tract variables. We compared the performance of this speech inversion system with prior experiments conducted using Mel Frequency Cepstral Coefficients (MFCCs).
Rahil Parikh, Nadee Seneviratne, Ganesh Sivaraman, Shihab A. Shamma, Carol Y. Espy-Wilson
INTERSPEECH4
2016 Dynamic Reweighting of Auditory Modulation Filters
abstract
Sound waveforms convey information largely via amplitude modulations (AM). A large body of experimental evidence has provided support for a modulation (bandpass) filterbank. Details of this model have varied over time partly reflecting different experimental conditions and diverse datasets from distinct task strategies, contributing uncertainty to the bandwidth measurements and leaving important issues unresolved. We adopt here a solely data-driven measurement approach in which we first demonstrate how different models can be subsumed within a common 'cascade' framework, and then proceed to characterize the cascade via system identification analysis using a single stimulus/task specification and hence stable task rules largely unconstrained by any model or parameters. Observers were required to detect a brief change in level superimposed onto random level changes that served as AM noise; the relationship between trial-by-trial noisy fluctuations and corresponding human responses enables targeted identification of distinct cascade elements. The resulting measurements exhibit a dynamic complex picture in which human perception of auditory modulations appears adaptive in nature, evolving from an initial lowpass to bandpass modes (with broad tuning, Q∼1) following repeated stimulus exposure.
Eva R. M. Joosten, Shihab A. Shamma, Christian Lorenzi, Peter Neri
PLoS Comput. Biol.2
2014 A State-Space Model for Decoding Auditory Attentional Modulation from MEG in a Competing-Speaker Environment
Sahar Akram, Jonathan Z. Simon, Shihab A. Shamma, Behtash Babadi
NIPS3
2014 Segregating Complex Sound Sources through Temporal Coherence
abstract
A new approach for the segregation of monaural sound mixtures is presented based on the principle of temporal coherence and using auditory cortical representations. Temporal coherence is the notion that perceived sources emit coherently modulated features that evoke highly-coincident neural response patterns. By clustering the feature channels with coincident responses and reconstructing their input, one may segregate the underlying source from the simultaneously interfering signals that are uncorrelated with it. The proposed algorithm requires no prior information or training on the sources. It can, however, gracefully incorporate cognitive functions and influences such as memories of a target source or attention to a specific set of its attributes so as to segregate it from its background. Aside from its unusual structure and computational innovations, the proposed model provides testable hypotheses of the physiological mechanisms of this ubiquitous and remarkable perceptual ability, and of its psychophysical manifestations in navigating complex sensory environments.
Lakshmi Krishnan, Mounya Elhilali, Shihab A. Shamma
PLoS Comput. Biol.3
2013 Speech enhancement using convolutive nonnegative matrix factorization with cosparsity regularization
Majid Mirbagheri, Yanbo Xu, Sahar Akram, Shihab A. Shamma
INTERSPEECH4
2012 The UMD-JHU 2011 speaker recognition system
abstract
In recent years, there have been significant advances in the field of speaker recognition that has resulted in very robust recognition systems. The primary focus of many recent developments have shifted to the problem of recognizing speakers in adverse conditions, e.g in the presence of noise/reverberation. In this paper, we present the UMD-JHU speaker recognition system applied on the NIST 2010 SRE task. The novel aspects of our systems are: 1) Improved performance on trials involving different vocal effort via the use of linear-scale features; 2) Expected improved recognition performance in the presence of reverberation and noise via the use of frequency domain perceptual linear predictor and cortical features; 3) A new discriminative kernel partial least squares (KPLS) framework that complements state-of-the-art back-end systems JFA and PLDA to aid in better overall recognition; and 4) Acceleration of JFA, PLDA and KPLS back-ends via distributed computing. The individual components of the system and the fused system are compared against a baseline JFA system and results reported by SRI and MIT-LL on SRE2010.
Daniel Garcia-Romero, Xinhui Zhou, Dmitry N. Zotkin, Balaji Vasan Srinivasan, Yuancheng Luo, Sriram Ganapathy, Samuel Thomas 0001, Sridhar Krishna Nemala, Garimella S. V. S. Sivaram, Majid Mirbagheri, Sri Harish Reddy Mallidi, Thomas Janu, Padmanabhan Rajan, Nima Mesgarani, Mounya Elhilali, Hynek Hermansky, Shihab A. Shamma, Ramani Duraiswami
ICASSP17
2012 Dimension reduction in regression using Gaussian Mixture Models
abstract
Linear-Nonlinear regression models play a fundamental role in characterizing nonlinear systems. In this paper, we propose a method to estimate the linear transform in such models equivalent to a subspace of a small dimension in the input space that is relevant for eliciting response. The novel aspect of this work is the formulation of the mutual information between the transformed inputs and output as a closed-form function of the parameters of their joint density in the form of Gaussian Mixture Models and we subsequently maximize this measure to find relevant dimensions. Instead of a commonly used mutual information measure based on Kullback-Leibler divergence, we use a measure called Quadratic Euclidean Mutual Information. Through experiments on both synthesized data and real MEG recordings, the effectiveness of the proposed method is demonstrated.
Majid Mirbagheri, Yanbo Xu, Shihab A. Shamma
ICASSP3
2012 An Auditory Inspired Multimodal Framework for Speech Enhancement
Majid Mirbagheri, Sahar Akram, Shihab A. Shamma
INTERSPEECH3
2012 Acoustic and Data-driven Features for Robust Speech Activity Detection
abstract
In this paper we evaluate different features for speech activity detection (SAD). Several signal processing techniques are used to derive acoustic features that capture attributes of speech useful in differentiating speech segments in noise. The acoustic features include short-term spectral features, long-term modulation features both derived using Frequency Domain Linear Prediction (FDLP), and joint spectro-temporal features extracted using 2D filters on a cortical representation of speech. Posteriors of speech and non-speech from a trained multi-layer perceptron are also used as data-driven features for this task. These feature extraction techniques form part of an elaborate feature extraction front-end where information spanning several hundreds of milliseconds of the signal are used along with heteroscedastic linear discriminant analysis for dimensionality reduction. Processed feature outputs from the proposed front-end are used to train SAD systems based on Gaussian mixture models for processing of speech from multiple languages transmitted over noisy radio communication channels under the ongoing DARPA Robust Automatic Transcription of Speech (RATS) program. The proposed front-end performs significantly better than standard acoustic feature extraction techniques in these noisy conditions.
Samuel Thomas 0001, Sri Harish Reddy Mallidi, Thomas Janu, Hynek Hermansky, Nima Mesgarani, Xinhui Zhou, Shihab A. Shamma, Tim Ng, Bing Zhang 0004, Long Nguyen 0001, Spyridon Matsoukas
INTERSPEECH7
2012 Automatic intelligibility assessment of pathologic speech in head and neck cancer based on auditory-inspired spectro-temporal modulations
abstract
Oral, head and neck cancer represents 3% of all cancers in the United States and is the 6th most common cancer worldwide. Depending on the tumor size, location and staging, patients are treated by radical surgery, radiology, chemotherapy or a combination of those treatments. As a result, their anatomical structures for speech are impaired and this leads to some negative impact on their speech intelligibility. As a part of the INTERSPEECH 2012 speaker trait Pathology sub-challenge, this study explored the use of auditory-inspired spectro-temporal modulation features for automatic speech intelligibility assessment of those pathologic speech. The averaged spectro-temporal modulations of speech considered as either intelligible or non-intelligible in the challenge database were analyzed and it was found that the non-intelligible speech tends to have its modulation amplitude peaks shift towards a smaller rate and scale. Based on SVM and GMM, variants of spectro-temporal modulation features were tested on the speaker trait challenge problem and the resulting performances on both the development and the test datasets are comparable to the baseline performance.
Xinhui Zhou, Daniel Garcia-Romero, Nima Mesgarani, Maureen Stone 0001, Carol Y. Espy-Wilson, Shihab A. Shamma
INTERSPEECH6
2012 Music in Our Ears: The Biological Bases of Musical Timbre Perception
abstract
Timbre is the attribute of sound that allows humans and other animals to distinguish among different sound sources. Studies based on psychophysical judgments of musical timbre, ecological analyses of sound's physical characteristics as well as machine learning approaches have all suggested that timbre is a multifaceted attribute that invokes both spectral and temporal sound features. Here, we explored the neural underpinnings of musical timbre. We used a neuro-computational framework based on spectro-temporal receptive fields, recorded from over a thousand neurons in the mammalian primary auditory cortex as well as from simulated cortical neurons, augmented with a nonlinear classifier. The model was able to perform robust instrument classification irrespective of pitch and playing style, with an accuracy of 98.7%. Using the same front end, the model was also able to reproduce perceptual distance judgments between timbres as perceived by human listeners. The study demonstrates that joint spectro-temporal features, such as those observed in the mammalian primary auditory cortex, are critical to provide the rich-enough representation necessary to account for perceptual judgments of timbre by human listeners, as well as recognition of musical instruments.
Kailash Patil, Daniel Pressnitzer, Shihab A. Shamma, Mounya Elhilali
PLoS Comput. Biol.3
2012 Functional Connectivity and Tuning Curves in Populations of Simultaneously Recorded Neurons
abstract
How interactions between neurons relate to tuned neural responses is a longstanding question in systems neuroscience. Here we use statistical modeling and simultaneous multi-electrode recordings to explore the relationship between these interactions and tuning curves in six different brain areas. We find that, in most cases, functional interactions between neurons provide an explanation of spiking that complements and, in some cases, surpasses the influence of canonical tuning curves. Modeling functional interactions improves both encoding and decoding accuracy by accounting for noise correlations and features of the external world that tuning curves fail to capture. In cortex, modeling coupling alone allows spikes to be predicted more accurately than tuning curve models based on external variables. These results suggest that statistical models of functional interactions between even relatively small numbers of neurons may provide a useful framework for examining neural coding.
Ian H. Stevenson, Brian M. London, Emily R. Oby, Nicholas A. Sachs, Jacob Reimer, Bernhard Englitz, Stephen V. David, Shihab A. Shamma, Timothy J. Blanche, Kenji Mizuseki, Amin Zandvakili, Nicholas G. Hatsopoulos, Lee E. Miller, Konrad P. Kording
PLoS Comput. Biol.8
2011 Linear versus mel frequency cepstral coefficients for speaker recognition
abstract
Mel-frequency cepstral coefficients (MFCC) have been dominantly used in speaker recognition as well as in speech recognition. However, based on theories in speech production, some speaker characteristics associated with the structure of the vocal tract, particularly the vocal tract length, are reflected more in the high frequency range of speech. This insight suggests that a linear scale in frequency may provide some advantages in speaker recognition over the mel scale. Based on two state-of-the-art speaker recognition back-end systems (one Joint Factor Analysis system and one Probabilistic Linear Discriminant Analysis system), this study compares the performances between MFCC and LFCC (Linear frequency cepstral coefficients) in the NIST SRE (Speaker Recognition Evaluation) 2010 extended-core task. Our results in SRE10 show that, while they are complementary to each other, LFCC consistently outperforms MFCC, mainly due to its better performance in the female trials. This can be explained by the relatively shorter vocal tract in females and the resulting higher formant frequencies in speech. LFCC benefits more in female speech by better capturing the spectral characteristics in the high frequency region. In addition, our results show some advantage of LFCC over MFCC in reverberant speech. LFCC is as robust as MFCC in the babble noise, but not in the white noise. It is concluded that LFCC should be more widely used, at least for the female trials, by the mainstream of the speaker recognition community.
Xinhui Zhou, Daniel Garcia-Romero, Ramani Duraiswami, Carol Y. Espy-Wilson, Shihab A. Shamma
ASRU5
2011 Speech processing with a cortical representation of audio
abstract
Neurophysiological studies in the primary auditory cortex have recently demonstrated a rich diversity of responses that provide an explicit multidimensional representation of phonemic acoustic features (Mesgarani 2008). Specifically, distinct subsets of cortical neurons are activated by articulatory gestures and dynamics that are characteristic of different phonemes. Here we use a computational cortical model to illustrate how these phonetic features appear in such a multiresolution representation. We also review how this representation has been successfully applied in variety of speech processing tasks including robust speech discrimination, speech enhancement and phoneme recognition.
Nima Mesgarani, Shihab A. Shamma
ICASSP2
2010 Nonlinear filtering of spectrotemporal modulations in speech enhancement
abstract
A monaural noise-suppression algorithm is proposed that nonlinearly manipulates the spectrotemporal modulations of speech as represented in a model of auditory cortical processing. A distinctive aspect of this approach is its consideration of the non-stationary dynamic behavior of speech that is captured using nonlinear filters, thus achieving excellent perceptual quality in the presence of many types of additive noise. Subjective tests demonstrate a significant improvement in the quality of the enhanced speech over the noisy samples even at SNRs as low as -5 dB.
Majid Mirbagheri, Nima Mesgarani, Shihab A. Shamma
ICASSP3
2008 Information-bearing components of speech intelligibility under babble-noise and bandlimiting distortions
abstract
Performance of speech technologies can benefit greatly from a deeper appreciation of the nature of the information- bearing features in continuous speech. To explore these features, we focus here on the role of the spectral and temporal modulations in maintaining the intelligibility of speech as it becomes severely degraded by low-pass filtering and additive babble noise. These modulations are estimated using a biological model of auditory processing which approximates the representation of sound in the cortex. Intelligibility of the noisy speech is computed directly from this model via the spectro-temporal modulation index (STMI), and the validity of this metric is confirmed by a detailed comparison with results of psychoacoustic tests. Our analysis reveals quantitatively why certain types of noise are more disruptive to speech intelligibility than others (e.g., babble vs. white noise). It also highlights the important contribution of both spectral and temporal modulations in accurately predicting the intelligibility of speech under adverse conditions.
Mounya Elhilali, Shihab A. Shamma
ICASSP2
2007 Representation of Phonemes in Primary Auditory Cortex: How the Brain Analyzes Speech
abstract
Many transformations inspired by the auditory system have improved the performance of automatic speech recognition (ASR) systems. However, humans perform substantially better than today's ASR systems, suggesting that ASR systems can further benefit from understanding how the brain represents speech. To learn about the cortical representation of speech, we measured the neural responses in the primary auditory cortex to sentences from the TIMIT database. Here we examine how individual phonemes activate different subsets of auditory neurons, reflecting the diversity of neural tuning properties. We find that neurons with different spectro-temporal tuning provide an explicit multidimensional representation of articulatory features independent of speaker and context. This representation that matches the human perception could provide a framework for ASR in adverse conditions.
Nima Mesgarani, Stephen V. David, Shihab A. Shamma
ICASSP (4)3
2007 Temporal Symmetry in Primary Auditory Cortex: Implications for Cortical Connectivity
abstract
Neurons in primary auditory cortex (AI) in the ferret (Mustela putorius) that are well described by their spectrotemporal response field (STRF) are found also to have a distinctive property that we call temporal symmetry. For temporally symmetric neurons, every temporal cross-section of the STRF (impulse response) is given by the same function of time, except for a scaling and a Hilbert rotation. This property held in 85% of neurons (123 out of 145) recorded from awake animals and in 96% of neurons (70 out of 73) recorded from anesthetized animals. This property of temporal symmetry is highly constraining for possible models of functional neural connectivity within and into AI. We find that the simplest models of functional thalamic input, from the ventral medial geniculate body (MGB), into the entry layers of AI are ruled out because they are incompatible with the constraints of the observed temporal symmetry. This is also the case for the simplest models of functional intracortical connectivity. Plausible models that do generate temporal symmetry, from both thalamic and intracortical inputs, are presented. In particular, we propose that two specific characteristics of the thalamocortical interface may be responsible. The first is a temporal mismatch between the fast dynamics of the thalamus and the slow responses of the cortex. The second is that all thalamic inputs into a cortical module (or a cluster of cells) must be restricted to one point of entry (or one cell in the cluster). This latter property implies a lack of correlated horizontal interactions across cortical modules during the STRF measurements. The implications of these insights in the auditory system, and comparisons with similar properties in the visual system, are explored.
Jonathan Z. Simon, Didier A. Depireux, David J. Klein, Jonathan B. Fritz, Shihab A. Shamma
Neural Comput.5
2006 A Biologically-Inspired Approach to the Cocktail Party Problem
abstract
Though seemingly effortless, our auditory system engages in complex processes and transformations which enable us to segregate speech and other sounds in cocktail party settings. This paper presents a computational approach to modelling monaural auditory scene analysis, where we attempt to account for perceptual and neuronal findings of receptive field selectivity and adaptation in the auditory cortex. The model introduces a biologically-inspired scheme of dynamic segregation of auditory streams, based on unsupervised clustering and the statistical theory of Kalman prediction. Our method demonstrates its ability to emulate known percepts reported by human subjects in auditory streaming and sound organization tests, and yields successful results in segregating speech from concurrent speaker and music interferences
Mounya Elhilali, Shihab A. Shamma
ICASSP (5)2
2006 Spectrum restoration from multiscale auditory phase singularities by generalized projections
abstract
We examine the encoding of acoustic spectra by parameters derived from singularities found in their multiscale auditory representations. The multiscale representation is a wavelet transform of an auditory version of the spectrum, formulated based on findings of perceptual experiments and physiological research in the auditory cortex. The multiscale representation of a spectral pattern usually contains well-defined singularities in its phase function that reflect prominent features of the underlying spectrum such as its relative peak locations and amplitudes. Properties (locations and strength) of these singularities are examined and employed to reconstruct the original spectrum by using an iterative projection algorithm. Although the singularities form a nonconvex set, simulations demonstrate that a well-chosen initial pattern usually converges on a good approximation of the input spectrum. Perceptually intelligible speech can be resynthesized from the reconstructed auditory spectrograms, and hence these singularities can potentially serve as efficient features in speech compression. Besides, the singularities are very noise-robust which makes them useful features in various applications such as vowel recognition and speaker identification
Taishih Chi, Shihab A. Shamma
IEEE Trans. Speech Audio Process.2
2006 Discrimination of speech from nonspeech based on multiscale spectro-temporal Modulations
abstract
We describe a content-based audio classification algorithm based on novel multiscale spectro-temporal modulation features inspired by a model of auditory cortical processing. The task explored is to discriminate speech from nonspeech consisting of animal vocalizations, music, and environmental sounds. Although this is a relatively easy task for humans, it is still difficult to automate well, especially in noisy and reverberant environments. The auditory model captures basic processes occurring from the early cochlear stages to the central cortical areas. The model generates a multidimensional spectro-temporal representation of the sound, which is then analyzed by a multilinear dimensionality reduction technique and classified by a support vector machine (SVM). Generalization of the system to signals in high level of additive noise and reverberation is evaluated and compared to two existing approaches (Scheirer and Slaney, 2002 and Kingsbury et al., 2002). The results demonstrate the advantages of the auditory model over the other two systems, especially at low signal-to-noise ratios (SNRs) and high reverberation.
Nima Mesgarani, Malcolm Slaney, Shihab A. Shamma
IEEE Trans. Speech Audio Process.3
2005 Speech Enhancement Based on Filtering the Spectrotemporal Modulations
abstract
A monaural noise suppression algorithm is proposed based on filtering the spectrotemporal modulations of noisy speech. The modulations are estimated from a multiscale representation of the signal spectrogram generated by a model of sound processing in the auditory system. A significant advantage of this method is its ability to suppress noise that has distinctive modulation patterns, despite being spectrally overlapping with the speech. The performance of the algorithm is evaluated using subjective and objective tests and compared to the optimal smoothing and minimum statistics approach (Martin (2001)). The results demonstrate the efficacy of the spectrotemporal filtering approach in the conditions examined.
Nima Mesgarani, Shihab A. Shamma
ICASSP (1)2
2004 Speech discrimination based on multiscale spectro-temporal modulations
abstract
A novel approach for content based audio classification is presented based on multiscale spectro-temporal modulation features extracted using a model of auditory cortex. The task is to discriminate speech from non-speech which consists of animal vocalizations, music and environmental sounds. Generalization of the system to signals in high level of additive noise and reverberation is evaluated and compared to two existing approaches. The results demonstrate the advantages of the auditory model over the other two systems, especially at low SNR and high reverberation.
Nima Mesgarani, Shihab A. Shamma, Malcolm Slaney
ICASSP (1)2
2003 Pitch and timbre manipulations using cortical representation of sound
abstract
The sound received at the ears is processed by humans using signal processing that separates the signal along intensity, pitch and timbre dimensions. Conventional Fourier-based signal processing, while endowed with fast algorithms, is unable to represent a signal easily along the lines of these attributes. We use a recently proposed cortical representation (Elhilali, M. et al., Speech Communications, 2002) to represent and manipulate sound. We briefly overview algorithms for obtaining, manipulating and inverting cortical representation of a sound and describe algorithms for manipulating signal pitch and timbre separately. The algorithms are first used to create the sound of an instrument between a "guitar" and a "trumpet". Applications to creating maximally separable sounds in auditory user interfaces are discussed.
Dmitry N. Zotkin, Shihab A. Shamma, Powen Ru, Ramani Duraiswami, Larry Davis 0001
ICASSP (5)2
2003 Pitch and timbre manipulations using cortical representation of sound
abstract
The sound receiver at the ears is processed by humans using signal processing that separate the signal along intensity, pitch and timbre dimensions. Conventional Fourier-based signal processing, while endowed with fast algorithms, is unable to easily represent signal along these attributes. In this paper we use a cortical representation to represent the manipulate sound. We briefly overview algorithms for obtaining, manipulating and inverting cortical representation of sound and describe algorithms for manipulating signal pitch and timbre separately. The algorithms are first used to create sound of an instrument between a guitar and a trumpet. Applications to creating maximally separable sounds in auditory user interfaces are discussed.
Dmitry N. Zotkin, Shihab A. Shamma, Powen Ru, Ramani Duraiswami, Larry Davis 0001
ICME2
2003 A spectro-temporal modulation index (STMI) for assessment of speech intelligibility
Mounya Elhilali, Taishih Chi, Shihab A. Shamma
Speech Commun.3
1999 A dendritic model of coincidence detection in the avian brainstem
Jonathan Z. Simon, Catherine E. Carr, Shihab A. Shamma
Neurocomputing3
1995 Spectral shape analysis in the central auditory system
abstract
A model of spectral shape analysis in the central auditory system is developed based on neurophysiological mappings in the primary auditory cortex and on results from psychoacoustical experiments in human subjects. The model suggests that the auditory system analyzes an input spectral pattern along three independent dimensions: a logarithmic frequency axis, a local symmetry axis, and a local spectral bandwidth axis. It is shown that this representation is equivalent to performing an affine wavelet transform of the spectral pattern and preserving both the magnitude (a measure of the scale or local bandwidth of the spectrum) and phase (a measure of the local symmetry of the spectrum). Such an analysis is in the spirit of the cepstral analysis commonly used in speech recognition systems, the major difference being that the double Fourier-like transformation that the auditory system employs is carried out in a local fashion. Examples of such a representation for various speech and synthetic signals are discussed, together with its potential significance and applications for speech and audio processing.>
Kuansan Wang, Shihab A. Shamma
IEEE Trans. Speech Audio Process.2
1994 Self-normalization and noise-robustness in early auditory representations
abstract
A common sequence of operations in the early stages of most sensory systems is a multiscale transform followed by a compressive nonlinearity. The authors explore the contribution of these operations to the formation of robust and perceptually significant representation in the early auditory system. It is shown that auditory representation of the acoustic spectrum is effectively a self-normalized spectral analysis, i.e., the auditory system computes a spectrum divided by a smoothed version of itself. Such a self-normalization induces significant effects such as spectral shape enhancement and robustness against scaling and noise corruption. Examples using synthesized signals and a natural speech vowel are presented to illustrate these results. Furthermore, the characteristics of auditory representation are discussed in the context of several psychoacoustical findings, together with the possible benefits of this model for various engineering applications.>
Kuansan Wang, Shihab A. Shamma
IEEE Trans. Speech Audio Process.2
1993 Noise robustness in the auditory representation of speech signals
Kuansan Wang, Shihab A. Shamma, William J. Byrne
ICASSP (2)2
1992 Auditory representations of acoustic signals
abstract
An analytically tractable framework is presented to describe mechanical and neural processing in the early stages of the auditory system. Algorithms are developed to assess the integrity of the acoustic spectrum at all processing stages. The algorithms employ wavelet representations, multiresolution processing, and the method of convex projections to construct a close replica of the input stimulus. Reconstructions using natural speech sounds demonstrate minimal loss of information along the auditory pathway. Close inspection of the final auditory patterns reveals spectral enhancements and noise suppression that have close perceptual correlates. The functional significance of the various auditory processing stages is discussed in light of the model, together with their potential applications in automatic speech recognition and low bit-rate data compression.>
Kuansan Wang, Shihab A. Shamma
IEEE Trans. Inf. Theory3