EDBT 2026 Demo / reviewers in the wild / expert
Afsaneh Asaei
dblp:46/1308
· DBLP profile ↗
34ranked-venue papers
15as first author
0since 2021 · last 2020
0000-0002-1917-601XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 12 first-authorArtificial intelligence and machine learning · 15 · 7 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Speech recognition and synthesis · 100% | |
| Computer graphics and multimedia
1 paper |
Audio and music processing · 100% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
keyword spotting |
0.3 | 1 | 2018 | Sparse Subspace Modeling for Query by Example Spoken Term Detection · IEEE ACM Trans. Audio Speech Lang. Process. 2018 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › keyword spotting
query-by-example spoken term detection |
0.3 | 1 | 2018 | Sparse Subspace Modeling for Query by Example Spoken Term Detection · IEEE ACM Trans. Audio Speech Lang. Process. 2018 |
Natural language and speech › Speech recognition and synthesis
speech production |
0.3 | 1 | 2017 | Perceptual Information Loss due to Impaired Speech Production · IEEE ACM Trans. Audio Speech Lang. Process. 2017 |
Natural language and speech › Speech recognition and synthesis › speech coding
low-bit-rate speech coding |
0.2 | 1 | 2016 | Composition of Deep and Spiking Neural Networks for Very Low Bit Rate Speech Coding · IEEE ACM Trans. Audio Speech Lang. Process. 2016 |
Natural language and speech › Speech recognition and synthesis › speech coding
neural speech coding |
0.2 | 1 | 2016 | Composition of Deep and Spiking Neural Networks for Very Low Bit Rate Speech Coding · IEEE ACM Trans. Audio Speech Lang. Process. 2016 |
Natural language and speech › Speech recognition and synthesis
speech coding |
0.2 | 1 | 2016 | Composition of Deep and Spiking Neural Networks for Very Low Bit Rate Speech Coding · IEEE ACM Trans. Audio Speech Lang. Process. 2016 |
Audio and music processing › source separation
reverberant source separation |
0.2 | 1 | 2014 | Structured Sparsity Models for Reverberant Speech Separation · IEEE ACM Trans. Audio Speech Lang. Process. 2014 |
Audio and music processing
room acoustics |
0.2 | 1 | 2014 | Structured Sparsity Models for Reverberant Speech Separation · IEEE ACM Trans. Audio Speech Lang. Process. 2014 |
Audio and music processing › room acoustics
room geometry inference |
0.2 | 1 | 2014 | Structured Sparsity Models for Reverberant Speech Separation · IEEE ACM Trans. Audio Speech Lang. Process. 2014 |
Audio and music processing › source separation
speech separation |
0.2 | 1 | 2014 | Structured Sparsity Models for Reverberant Speech Separation · IEEE ACM Trans. Audio Speech Lang. Process. 2014 |
Information theory
information measures |
0.1 | 1 | 2017 | Perceptual Information Loss due to Impaired Speech Production · IEEE ACM Trans. Audio Speech Lang. Process. 2017 |
Audio and music processing
sound source localization |
0.1 | 1 | 2014 | Structured Sparsity Models for Reverberant Speech Separation · IEEE ACM Trans. Audio Speech Lang. Process. 2014 |
Methods — techniques the papers use, named apart from their topics
deep neural network · 0.8information-theoretic framework · 0.6template matching · 0.3sparse coding · 0.3dynamic programming · 0.3spiking neural network · 0.2phonological features · 0.2structured sparsity · 0.2image model · 0.2convex optimization · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | On quantifying the quality of acoustic models in hybrid DNN-HMM ASR
Pranay Dighe, Afsaneh Asaei, Hervé Bourlard |
Speech Commun. | 2 |
| 2019 | Low-rank and sparse subspace modeling of speech for DNN based acoustic modeling
Pranay Dighe, Afsaneh Asaei, Hervé Bourlard |
Speech Commun. | 2 |
| 2018 | Phonological Posterior Hashing for Query by Example Spoken Term DetectionabstractState of the art query by example spoken term detection (QbE-STD) systems in zero-resource conditions rely on representation of speech in terms of sequences of class-conditional posterior probabilities estimated by deep neural network (DNN). The posteriors are often used for pattern matching or dynamic time warping (DTW). Exploiting posterior probabilities as speech representation propounds diverse advantages in a classification system. One key property of the posterior representations is that they admit a highly effective hashing strategy that enables indexing a large audio archive in divisions for reducing the search complexity. Moreover, posterior indexing leads to a compressed representation and enables pronunciation dewarping and partial detection with no need for DTW. We exploit these characteristics of the posterior space in the context of redundant hash addressing for query-by-example spoken term detection (QbE-STD). We evaluate the QbE-STD system on AMI corpus and demonstrate that tremendous speedup and superior accuracy is achieved compared to the state-of-the-art pattern matching solution based on DTW. The system has the potential to enable massively large scale spoken query detection. Afsaneh Asaei, Dhananjay Ram, Hervé Bourlard |
INTERSPEECH | 1 |
| 2018 | Far-Field ASR Using Low-Rank and Sparse Soft Targets from Parallel DataabstractFar-field automatic speech recognition (ASR) of conversational speech is often considered to be a very challenging task due to the poor quality of alignments available for training the DNN acoustic models. A common way to alleviate this problem is to use clean alignments obtained from parallelly recorded close-talk speech data. In this work, we advance the parallel data approach by obtaining enhanced low-rank and sparse soft targets from a close-talk ASR system and using them for training more accurate far-field acoustic models. Specifically, we (i) exploit eigenposteriors and Compressive Sensing dictionaries to learn low-dimensional senone subspaces in DNN posterior space, and (ii) enhance close-talk DNN posteriors to achieve high quality soft targets for training far-field DNN acoustic models. We show that the enhanced soft targets encode the structural and temporal interrelationships among senone classes which are easily accessible in the DNN posterior space of close-talk speech but not in its noisy far-field counterpart. We exploit enhanced soft targets to improve the mapping of far-field acoustics to close-talk senone classes. The experiments are performed on AMI meeting corpus where our approach improves DNN based acoustic modeling by 4.4% absolute (~8% rel.) reduction in WER as compared to a system which doesn't use parallel data. Finally, the approach is also validated on state-of-the-art recurrent and time delay neural network architectures. Pranay Dighe, Afsaneh Asaei, Hervé Bourlard |
SLT | 2 |
| 2018 | Phonetic subspace features for improved query by example spoken term detection
Dhananjay Ram, Afsaneh Asaei, Hervé Bourlard |
Speech Commun. | 2 |
| 2018 | Sparse Subspace Modeling for Query by Example Spoken Term DetectionabstractThis paper focuses on the problem of query by example spoken term detection (QbE-STD) in zero-resource scenario. Current state-of-the-art approaches to tackle this problem rely on dynamic programming based template matching techniques using phone posterior features extracted at the output of a deep neural network. Previously, it has been shown that the space of phone posteriors is highly structured, as a union of low-dimensional subspaces. To exploit the temporal and sparse structure of the speech data, we investigate here three different QbE-STD systems based on sparse model recovery. More specifically, we use query examples to model the query subspace using dictionary for sparse coding. Reconstruction errors calculated using sparse representation of feature vectors are then used to characterize the underlying subspaces. The first approach uses these reconstruction errors in a dynamic programming framework to detect the spoken query, resulting in a much faster search compared to standard template matching. The other two methods aim at merging template matching and sparsity-based approaches to further improve the performance. The first one proposes to regularize the template matching local distances using sparse reconstruction errors. The second approach aims at using the sparse reconstruction errors to rescore (improve) the template matching likelihood. Experiments on two different databases (AMI and MediaEval) show that the proposed hybrid systems perform better than a highly competitive QbE-STD baseline system. Dhananjay Ram, Afsaneh Asaei, Hervé Bourlard |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Low-rank and sparse soft targets to learn better DNN acoustic modelsabstractConventional deep neural networks (DNN) for speech acoustic modeling rely on Gaussian mixture models (GMM) and hidden Markov model (HMM) to obtain binary class labels as the targets for DNN training. Subword classes in speech recognition systems correspond to context-dependent tied states or senones. The present work addresses some limitations of GMM-HMM senone alignments for DNN training. We hypothesize that the senone probabilities obtained from a DNN trained with binary labels can provide more accurate targets to learn better acoustic models. However, DNN outputs bear inaccuracies which are exhibited as high dimensional unstructured noise, whereas the informative components are structured and low-dimensional. We exploit principal component analysis (PCA) and sparse coding to characterize the senone subspaces. Enhanced probabilities obtained from low-rank and sparse reconstructions are used as soft-targets for DNN acoustic modeling, that also enables training with untranscribed data. Experiments conducted on AMI corpus shows 4.6% relative reduction in word error rate. Pranay Dighe, Afsaneh Asaei, Hervé Bourlard |
ICASSP | 2 |
| 2017 | Exploiting Eigenposteriors for Semi-Supervised Training of DNN Acoustic Models with Sequence DiscriminationabstractLIDIAP Pranay Dighe, Afsaneh Asaei, Hervé Bourlard |
INTERSPEECH | 2 |
| 2017 | Perceptual Information Loss due to Impaired Speech ProductionabstractPhonological classes define articulatory-free and articulatory-bound phone attributes. Deep neural network is used to estimate the probability of phonological classes from the speech signal. In theory, a unique combination of phone attributes form a phoneme identity. Probabilistic inference of phonological classes thus enables estimation of their compositional phoneme probabilities. A novel information theoretic framework is devised to quantify the information conveyed by each phone attribute, and assess the speech production quality for perception of phonemes. As a use case, we hypothesize that disruption in speech production leads to information loss in phone attributes, and thus confusion in phoneme identification. We quantify the amount of information loss due to dysarthric articulation recorded in the TORGO database. A novel information measure is formulated to evaluate the deviation from an ideal phone attribute production leading us to distinguish healthy production from pathological speech. Afsaneh Asaei, Milos Cernak, Hervé Bourlard |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2016 | Exploiting low-dimensional structures to enhance DNN based acoustic modeling in speech recognitionabstractWe propose to model the acoustic space of deep neural network (DNN) class-conditional posterior probabilities as a union of low-dimensional subspaces. To that end, the training posteriors are used for dictionary learning and sparse coding. Sparse representation of the test posteriors using this dictionary enables projection to the space of training data. Relying on the fact that the intrinsic dimensions of the posterior subspaces are indeed very small and the matrix of all posteriors belonging to a class has a very low rank, we demonstrate how low-dimensional structures enable further enhancement of the posteriors and rectify the spurious errors due to mismatch conditions. The enhanced acoustic modeling method leads to improvements in continuous speech recognition task using hybrid DNN-HMM (hidden Markov model) framework in both clean and noisy conditions, where upto 15.4% relative reduction in word error rate (WER) is achieved. Pranay Dighe, Gil Luyet, Afsaneh Asaei, Hervé Bourlard |
ICASSP | 3 |
| 2016 | Phonetic and Phonological Posterior Search Space Hashing Exploiting Class-Specific Sparsity StructuresabstractThis paper shows that exemplar-based speech processing using class-conditional posterior probabilities admits a highly effective search strategy relying on posteriors' intrinsic sparsity structures. The posterior probabilities are estimated for phonetic and phonological classes using deep neural network (DNN) computational framework. Exploiting the class-specific sparsity leads to a simple quantized posterior hashing procedure to reduce the search space of posterior exemplars. To that end, small number of quantized posteriors are regarded as representatives of the posterior space and used as hash keys to index subsets of neighboring exemplars. The $k$ nearest neighbor ($k$NN) method is applied for posterior based classification problems. The phonetic posterior probabilities are used as exemplars for phonetic classification whereas the phonological posteriors are used as exemplars for automatic prosodic event detection. Experimental results demonstrate that posterior hashing improves the efficiency of $k$NN classification drastically. This work encourages the use of posteriors as discriminative exemplars appropriate for large scale speech classification tasks. Afsaneh Asaei, Gil Luyet, Milos Cernak, Hervé Bourlard |
INTERSPEECH | 1 |
| 2016 | Sound Pattern Matching for Automatic Prosodic Event DetectionabstractLIDIAP Milos Cernak, Afsaneh Asaei, Pierre-Edouard Honnet, Philip N. Garner, Hervé Bourlard |
INTERSPEECH | 2 |
| 2016 | Low-Rank Representation of Nearest Neighbor Posterior Probabilities to Enhance DNN Based Acoustic ModelingabstractLIDIAP Gil Luyet, Pranay Dighe, Afsaneh Asaei, Hervé Bourlard |
INTERSPEECH | 3 |
| 2016 | Subspace Detection of DNN Posterior Probabilities via Sparse Representation for Query by Example Spoken Term DetectionabstractWe cast the query by example spoken term detection (QbE-STD) problem as subspace detection where query and background subspaces are modeled as union of low-dimensional subspaces. The speech exemplars used for subspace modeling are class-conditional posterior probabilities estimated using deep neural network (DNN). The query and background training exemplars are exploited to model the underlying low-dimensional subspaces through dictionary learning for sparse representation. Given the dictionaries characterizing the query and background subspaces, QbE-STD is performed based on the ratio of the two corresponding sparse representation reconstruction errors. The proposed subspace detection method can be formulated as the generalized likelihood ratio test for composite hypothesis testing. The experimental evaluation demonstrate that the proposed method is able to detect the query given a single example and performs significantly better than a highly competitive QbE-STD baseline system based on template matching. Dhananjay Ram, Afsaneh Asaei, Hervé Bourlard |
INTERSPEECH | 2 |
| 2016 | Computational methods for underdetermined convolutive speech localization and separation via model-based sparse component analysis
Afsaneh Asaei, Hervé Bourlard, Mohammad Javad Taghizadeh, Volkan Cevher |
Speech Commun. | 1 |
| 2016 | On structured sparsity of phonological posteriors for linguistic parsing
Milos Cernak, Afsaneh Asaei, Hervé Bourlard |
Speech Commun. | 2 |
| 2016 | Sparse modeling of neural network posterior probabilities for exemplar-based speech recognition
Pranay Dighe, Afsaneh Asaei, Hervé Bourlard |
Speech Commun. | 2 |
| 2016 | Composition of Deep and Spiking Neural Networks for Very Low Bit Rate Speech CodingabstractMost current very low bit rate (VLBR) speech coding systems use hidden Markov model (HMM) based speech recognition and synthesis techniques. This allows transmission of information (such as phonemes) segment by segment; this decreases the bit rate. However, an encoder based on a phoneme speech recognition may create bursts of segmental errors; these would be further propagated to any suprasegmental (such as syllable) information coding. Together with the errors of voicing detection in pitch parametrization, HMM-based speech coding leads to speech discontinuities and unnatural speech sound artifacts. In this paper, we propose a novel VLBR speech coding framework based on neural networks (NNs) for end-to-end speech analysis and synthesis without HMMs. The speech coding framework relies on a phonological (subphonetic) representation of speech. It is designed as a composition of deep and spiking NNs: a bank of phonological analyzers at the transmitter, and a phonological synthesizer at the receiver. These are both realized as deep NNs, along with a spiking NN as an incremental and robust encoder of syllable boundaries for coding of continuous fundamental frequency (F0). A combination of phonological features defines much more sound patterns than phonetic features defined by HMM-based speech coders; this finer analysis/synthesis code contributes to smoother encoded speech. Listeners significantly prefer the NN-based approach due to fewer discontinuities and speech artifacts of the encoded speech. A single forward pass is required during the speech encoding and decoding. The proposed VLBR speech coding operates at a bit rate of approximately 360 bits/s. Milos Cernak, Alexandros Lazaridis, Afsaneh Asaei, Philip N. Garner |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | On application of non-negative matrix factorization for ad hoc microphone array calibration from incomplete noisy distancesabstractWe propose to use non-negative matrix factorization (NMF) to estimate the unknown pairwise distances and reconstruct a distance matrix for microphone array position calibration. We develop new multiplicative update rules for NMF with incomplete input matrix that take into account the symmetry of the distance matrix. Additionally, we develop a convex matrix completion method which is related to an l2-regularized symmetric NMF. Thorough experiments demonstrate that the proposed methods lead to substantial improvement over the state-of-the-art techniques in a wide range of signal-to-noise and unknown-distance ratios. The convex symmetric matrix completion method was found to be the most robust method with less computational cost. Afsaneh Asaei, Nasser Mohammadiha, Mohammad Javad Taghizadeh, Simon Doclo, Hervé Bourlard |
ICASSP | 1 |
| 2015 | Robust microphone placement for source localization from noisy distance measurementsabstractWe propose a novel algorithm to design an optimum array geometry for source localization inside an enclosure. We assume a square-law decay propagation model for the sound acquisition so that the additive noise on the measured source-microphone distances is proportional to the distances regardless of the noise distribution. We formulate the source localization as an instance of the “Generalized Trust Region Subproblem” (GTRS) whose solution gives the location of the source. We show that by suitable selection of the microphone locations, one can tremendously decrease the noise-sensitivity of the resulting solution. In particular, by minimizing the noise-sensitivity of the source location in terms of sensor positions, we find the optimal noise-robust array geometry for the enclosure. Simulation results are provided to show the efficiency of the proposed algorithm. Mohammad Javad Taghizadeh, Saeid Haghighatshoar, Afsaneh Asaei, Philip N. Garner, Hervé Bourlard |
ICASSP | 3 |
| 2015 | Novel GCC-PHAT model in diffuse sound field for microphone array pairwise distance based calibrationabstractWe propose a novel formulation of the generalized cross correlation with phase transform (GCC-PHAT) for a pair of microphones in diffuse sound field. This formulation elucidates the links between the microphone distances and the GCC-PHAT output. Hence, it leads to a new model that enables estimation of the pairwise distances by optimizing over the distances best matching the GCC-PHAT observations. Furthermore, the relation of this model to the coherence function is elaborated along with the dependency on the signal bandwidth. The experiments conducted on real data recordings demonstrate the theories and support the effectiveness of the proposed method. José F. Velasco, Mohammad Javad Taghizadeh, Afsaneh Asaei, Hervé Bourlard, Carlos Julian Martín-Arguedas, Javier Macías Guarasa, Daniel Pizarro-Perez |
ICASSP | 3 |
| 2015 | On compressibility of neural network phonological features for low bit rate speech codingabstractPhonological features extracted by neural network have shown interesting potential for low bit rate speech vocoding. The span of phonological features is wider than the span of phonetic features, and thus fewer frames need to be transmitted. Moreover, the binary nature of phonological features enables a higher compression ratio at minor quality cost. In this paper, we study the compressibility and structured sparsity of the phonological features. We propose a compressive sampling framework for speech coding and sparse reconstruction for decoding prior to synthesis. Compressive sampling is found to be a principled way for compression in contrast to the conventional pruning approach; it leads to $50$\\% reduction in the bit-rate for better or equal quality of the decoded speech. Furthermore, exploiting the structured sparsity and binary characteristic of these features have shown to enable very low bit-rate coding at 700 bps with negligible quality loss; this coding scheme imposes no latency. If we consider a latency of $256$~ms for supra-segmental structures, the rate of $250-350$~bps is achieved. Afsaneh Asaei, Milos Cernak, Hervé Bourlard |
INTERSPEECH | 1 |
| 2015 | Sparse modeling of posterior exemplars for keyword detectionabstractSparse representation has been shown to be a powerful modeling framework for classification and detection tasks.In this paper, we propose a new keyword detection algorithm based on sparse representation of the posterior exemplars.The posterior exemplars are phone conditional probabilities obtained from a deep neural network.This method relies on the concept that a keyword exemplar lies in a low-dimensional subspace which can be represented as a sparse linear combination of the training exemplars.The training exemplars are used to learn a dictionary for sparse representation of the keywords and background classes.Given this dictionary, the sparse representation of a test exemplar is used to detect the keywords.The experimental results demonstrate the potential of the proposed sparse modeling approach and it compares favorably with the state-of-the-art HMM-based framework on Numbers'95 database. Dhananjay Ram, Afsaneh Asaei, Pranay Dighe, Hervé Bourlard |
INTERSPEECH | 2 |
| 2015 | Ad hoc microphone array calibration: Euclidean distance matrix completion algorithm and theoretical guarantees
Mohammad Javad Taghizadeh, Reza Parhizkar, Philip N. Garner, Hervé Bourlard, Afsaneh Asaei |
Signal Process. | 5 |
| 2014 | Model-based sparse component analysis for reverberant speech localizationabstractIn this paper, the problem of multiple speaker localization via speech separation based on model-based sparse recovery is studies. We compare and contrast computational sparse optimization methods incorporating harmonicity and block structures as well as autoregressive dependencies underlying spectrographic representation of speech signals. The results demonstrate the effectiveness of block sparse Bayesian learning framework incorporating autoregressive correlations to achieve a highly accurate localization performance. Furthermore, significant improvement is obtained using ad-hoc microphones for data acquisition set-up compared to the compact microphone array. Afsaneh Asaei, Hervé Bourlard, Mohammad Javad Taghizadeh, Volkan Cevher |
ICASSP | 1 |
| 2014 | Posterior-based sparse representation for automatic speech recognitionabstractPosterior features have been shown to yield very good performance in multiple contexts including speech recognition, spoken term detection, and template matching. These days, posterior features are usually estimated at the output of a neural network. More recently, sparse representation has also been shown to potentially provide additional advantages to improve discrimination and robustness. One possible instance of this, is referred to as exemplar-based sparse representation. The present work investigates how to exploit sparse modelling together with posterior space properties to further improve speech recognition features. In that context, we leverage exemplar-based sparse representation, and propose a novel approach to project phone posterior features into a new, high-dimensional, sparse feature space. In fact, exploiting the properties of posterior spaces, we generate, new, high-dimensional, linguistically inspired (sub-phone and words), posterior distributions. Validation experiments are performed on the Phonebook (isolated words) and HIWIRE (continuous speech) databases, which support the effectiveness of the proposed approach for speech recognition tasks. Sara Bahaadini, Afsaneh Asaei, David Imseng, Hervé Bourlard |
INTERSPEECH | 2 |
| 2014 | Structured Sparsity Models for Reverberant Speech SeparationabstractWe tackle the speech separation problem through modeling the acoustics of the reverberant chambers. Our approach exploits structured sparsity models to perform speech recovery and room acoustic modeling from recordings of concurrent unknown sources. The speakers are assumed to lie on a two-dimensional plane and the multipath channel is characterized using the image model. We propose an algorithm for room geometry estimation relying on localization of the early images of the speakers by sparse approximation of the spatial spectrum of the virtual sources in a free-space model. The images are then clustered exploiting the low-rank structure of the spectro-temporal components belonging to each source. This enables us to identify the early support of the room impulse response function and its unique map to the room geometry. To further tackle the ambiguity of the reflection ratios, we propose a novel formulation of the reverberation model and estimate the absorption coefficients through a convex optimization exploiting joint sparsity model formulated upon spatio-spectral sparsity of concurrent speech representation. The acoustic parameters are then incorporated for separating individual speech signals through either structured sparse recovery or inverse filtering the acoustic channels. The experiments conducted on real data recordings of spatially stationary sources demonstrate the effectiveness of the proposed approach for speech separation and recognition. Afsaneh Asaei, Mohammad Golbabaee, Hervé Bourlard, Volkan Cevher |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2012 | Computational methods for structured sparse component analysis of convolutive speech mixturesabstractWe cast the under-determined convolutive speech separation as sparse approximation of the spatial spectra of the mixing sources. In this framework we compare and contrast the major practical algorithms for structured sparse recovery of speech signal. Specific attention is paid to characterization of the measurement matrix. We first propose how it can be identified using the Image model of multipath effect where the acoustic parameters are estimated by localizing a speaker and its images in a free space model. We further study the circumstances in which the coherence of the projections induced by microphone array design tend to affect the recovery performance. Afsaneh Asaei, Mike E. Davies 0001, Hervé Bourlard, Volkan Cevher |
ICASSP | 1 |
| 2011 | Model-based compressive sensing for multi-party distant speech recognitionabstractWe leverage the recent algorithmic advances in compressive sensing, and propose a novel source separation algorithm for efficient recovery of convolutive speech mixtures in spectro-temporal domain. Compared to the common sparse component analysis techniques, our approach fully exploits structured sparsity models to obtain substantial improvement over the existing state-of-the-art. We evaluate our method for separation and recognition of a target speaker in a multi-party scenario. Our results provide compelling evidence of the effectiveness of sparse recovery formulations in speech recognition. Afsaneh Asaei, Hervé Bourlard, Volkan Cevher |
ICASSP | 1 |
| 2011 | Multi-Party Speech Recovery Exploiting Structured Sparsity ModelsabstractWe study the sparsity of spectro-temporal representation of speech in reverberant acoustic conditions. This study motivates the use of structured sparsity models for efficient speech recov-ery. We formulate the underdetermined convolutive speech sep-aration in spectro-temporal domain as the sparse signal recovery where we leverage model-based recovery algorithms. To tackle the ambiguity of the real acoustics, we exploit the Image Model of the enclosures to estimate the room impulse response func-tion through a structured sparsity constraint optimization. The experiments conducted on real data recordings demonstrate the effectiveness of the proposed approach for multi-party speech applications. Index Terms: speech sparsity, structured sparsity models, un-derdetermined convolutive speech separation, Image Model Afsaneh Asaei, Mohammad Javad Taghizadeh, Hervé Bourlard, Volkan Cevher |
INTERSPEECH | 1 |
| 2010 | Analysis of phone posterior feature space exploiting class-specific sparsity and MLP-based similarity measureabstractClass posterior distributions have recently been used quite successfully in Automatic Speech Recognition (ASR), either for frame or phone level classification or as acoustic features, which can be further exploited (usually after some “ad hoc” transformations) in different classifiers (e.g., in Gaussian Mixture based HMMs). In the present paper, we show preliminary results showing that it may be possible to perform speech recognition without explicit subword unit (phone) classification or likelihood estimation, simply answering the question whether two acoustic (posterior) vectors belong to the same subword unit class or not. In this paper, we first exhibit specific properties of the posterior acoustic space before showing how those properties can be exploited to reach very high performance in deciding (based on an appropriate, trained, distance metric, and hypothesis testing approaches) whether two posterior vectors belong to the same class or not. Performance as high as 90% correct decision rates are reported on the TIMIT database, before reporting kNN phone classification rates. Afsaneh Asaei, Benjamin Picart, Hervé Bourlard |
ICASSP | 1 |
| 2010 | Sparse component analysis for speech recognition in multi-speaker environmentabstractSparse Component Analysis is a relatively young technique that relies upon a representation of signal occupying only a small part of a larger space. Mixtures of sparse components are disjoint in that space. As a particular application of sparsity of speech signals, we investigate the DUET blind source separation algorithm in the context of speech recognition for multi-party recordings. We show how DUET can be tuned to the particular case of speech recognition with interfering sources, and evaluate the limits of performance as the number of sources increases. We show that the separated speech fits a common metric for sparsity, and conclude that sparsity assumptions lead to good performance in speech separation and hence ought to benefit other aspects of the speech recognition chain. Afsaneh Asaei, Hervé Bourlard, Philip N. Garner |
INTERSPEECH | 1 |
| 2009 | Verified speaker localization utilizing voicing level in split-bands
Afsaneh Asaei, Mohammad Javad Taghizadeh, Marjan Bahrololum, Mohammed Ghanbari 0001 |
Signal Process. | 1 |
| 2008 | Far-field continuous speech recognition system based on speaker Localization and sub-band BeamformingabstractThis paper proposes a distant speech recognition system based on a novel speaker localization and beamforming (SRLB) algorithm. To localize the speaker an algorithm based on steered response power by utilizing harmonic structures of speech signal is proposed. This new scheme has the ability of speaker verification by fundamental frequency variation; therefore it can be utilized in the design of a speech recognition system for verified speakers. Then the performance of the Farsi speech recognition engine is evaluated under notorious conditions of noise and reverberation. Simulation results and tests on real data shows that by utilizing proposed localization scheme, recognition accuracy improves by 28% in high noise and reverberant conditions compared to the accuracy of single channel recognition. The capability of this algorithm in localizing a verified speaker improves system robustness to speech noises and reduces recognition errors up to %52 in the presence of speech noise. Afsaneh Asaei, Mohammad Javad Taghizadeh, Hossein Sameti |
AICCSA | 1 |