Xiaofei Li 0001

dblp:99/5152-1 · DBLP profile ↗
← Back
38ranked-venue papers
14as first author
13since 2021 · last 2026
0000-0003-0393-9905ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 9 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 6 since 2021Systems, architecture and hardware · 4 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 CRT-IFPD: Chinese Remainder Theorem-Based Interfrequency Phase Difference for Absolute Acoustic Ranging Without Time Synchronization
Liu Yang 0021, Xiaofei Li 0001, Guangyao Liu, Weimeng Cui, Naizheng Jia, Can Xue, Peixuan Hu, Zhi Wang 0003
IEEE Trans. Ind. Informatics2
2025 Mamba for Streaming ASR Combined with Unimodal Aggregation
abstract
This paper works on streaming automatic speech recognition (ASR). Mamba, a recently proposed state space model, has demonstrated the ability to match or surpass Transformers in various tasks while benefiting from a linear complexity advantage. We explore the efficiency of Mamba encoder for streaming ASR and propose an associated lookahead mechanism for leveraging controllable future information. Additionally, a streaming-style unimodal aggregation (UMA) method is implemented, which automatically detects token activity and streamingly triggers token output, and meanwhile aggregates feature frames for better learning token representation. Based on UMA, an early termination (ET) method is proposed to further reduce recognition latency. Experiments conducted on two Mandarin Chinese datasets demonstrate that the proposed model achieves competitive ASR performance in terms of both recognition accuracy and latency. Code will be open-sourced1.
Xiaofei Li 0001
ICASSP2
2024 Diffusion-Based Adversarial Purification for Speaker Verification
abstract
Recently, automatic speaker verification (ASV) based on deep learning is easily contaminated by adversarial attacks, which is a new type of attack that injects imperceptible perturbations to audio signals so as to make ASV produce wrong decisions. This poses a significant threat to the security and reliability of ASV systems. To address this issue, we propose a Diffusion-Based Adversarial Purification (DAP) method that enhances the robustness of ASV systems against such adversarial attacks. Our method leverages a conditional denoising diffusion probabilistic model to effectively purify the adversarial examples and mitigate the impact of perturbations. DAP first introduces controlled noise into adversarial examples, and then performs a reverse denoising process to reconstruct clean audio. Experimental results demonstrate the efficacy of the proposed DAP in enhancing the security of ASV and meanwhile minimizing the distortion of the purified audio signals.
Yibo Bai, Xiaofei Li 0001, Xuelong Li 0001
IEEE Signal Process. Lett.2
2024 Multichannel Long-Term Streaming Neural Speech Enhancement for Static and Moving Speakers
abstract
In this work, we extend our previously proposed offline SpatialNet for long-term streaming multichannel speech enhancement in both static and moving speaker scenarios. SpatialNet exploits spatial information, such as the spatial/steering direction of speech, for discriminating between target speech and interferences, and achieved outstanding performance. The core of SpatialNet is a narrow-band self-attention module used for learning the temporal dynamic of spatial vectors. Towards long-term streaming speech enhancement, we propose to replace the offline self-attention network with online networks that have linear inference complexity w.r.t signal length and meanwhile maintain the capability of learning long-term information. Three variants are developed based on i) masked self-attention, ii) Retention, a self-attention variant with linear inference complexity, and iii) Mamba, a structured-state-space-based RNN-like network. Moreover, we investigate the length extrapolation ability of different networks, namely test on signals that are much longer than training signals, and propose a short-signal training plus long-signal fine-tuning strategy, which largely improves the length extrapolation ability of the networks within limited training time. Overall, the proposed online SpatialNet achieves outstanding speech enhancement performance for long audio streams, and for both static and moving speakers.
Changsheng Quan, Xiaofei Li 0001
IEEE Signal Process. Lett.2
2024 Self-Supervised Audio Teacher-Student Transformer for Both Clip-Level and Frame-Level Tasks
abstract
Self-supervised learning (SSL) has emerged as a popular approach for learning audio representations. One goal of audio self-supervised pre-training is to transfer knowledge to downstream audio tasks, generally including clip-level and frame-level tasks. While frame-level tasks are important for fine-grained acoustic scene/event understanding, prior studies primarily evaluate on clip-level downstream tasks. In order to tackle both clip-level and frame-level tasks, this paper proposes Audio Teacher-Student Transformer (ATST), with a clip-level version (named ATST-Clip) and a frame-level version (named ATST-Frame), responsible for learning clip-level and frame-level representations, respectively. Both methods use a Transformer encoder and a teacher-student training scheme. We have carefully designed a view creation strategy for ATST-Clip and ATST-Frame. Specifically, ATST-Clip uses segment-wise data augmentations, and ATST-Frame integrates frame-wise data augmentations and masking. Experimental results show that our ATST-Frame model obtains state-of-the-art (SOTA) performances on most of the clip-level and frame-level downstream tasks. Especially, it outperforms other models by a large margin on the frame-level sound event detection task. In addition, the performance can be further improved by combining the two models through knowledge distillation.
Nian Shao, Xiaofei Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2024 SpatialNet: Extensively Learning Spatial Information for Multichannel Joint Speech Separation, Denoising and Dereverberation
abstract
This work proposes a neural network to extensively exploit spatial information for multichannel joint speech separation, denoising and dereverberation, named SpatialNet. In the short-time Fourier transform (STFT) domain, the proposed network performs end-to-end speech enhancement. It is mainly composed of interleaved narrow-band and cross-band blocks to respectively exploit narrow-band and cross-band spatial information. The narrow-band blocks process frequencies independently, and use self-attention mechanism and temporal convolutional layers to respectively perform spatial-feature-based speaker clustering and temporal smoothing/filtering. The cross-band blocks process frames independently, and use full-band linear layer and frequency convolutional layers to respectively learn the correlation between all frequencies and adjacent frequencies. Experiments are conducted on various simulated and real datasets, and the results show that 1) the proposed network achieves the state-of-the-art performance on almost all tasks; 2) the proposed network suffers little from the spectral generalization problem; and 3) the proposed network is indeed performing speaker clustering (demonstrated by attention maps).
Changsheng Quan, Xiaofei Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 IPDnet: A Universal Direct-Path IPD Estimation Network for Sound Source Localization
abstract
Extracting direct-path spatial feature is crucial for sound source localization in adverse acoustic environments. This paper proposes IPDnet, a neural network that estimates direct-path inter-channel phase difference (DP-IPD) of sound sources from microphone array signals. The estimated DP-IPD can be easily translated to source location based on the known microphone array geometry. First, a full-band and narrow-band fusion network is adopted for DP-IPD estimation, in which combined narrow-band and full-band layers are responsible for estimating the raw DP-IPD information in one frequency band and capturing the frequency correlations of DP-IPD, respectively. Second, a new multi-track DP-IPD learning target is proposed for the localization of a flexible number of sound sources. Third, the network is extended to handle variable microphone arrays. This version of IPDnet is trained with a large set of different microphone arrays, and then it is able to infer the source locations using new microphone arrays not seen at training time. Experiments with multiple number of moving speakers are conducted on both simulated and real-world data, which show that the full-band and narrow-band fusion network and the proposed multi-track DP-IPD learning target together achieve excellent sound source localization performance. Moreover, the proposed variable-array model generalizes well to unseen microphone arrays.
Bing Yang 0004, Xiaofei Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2024 Self-Supervised Learning of Spatial Acoustic Representation With Cross-Channel Signal Reconstruction and Multi-Channel Conformer
abstract
Supervised learning methods have shown effectiveness in estimating spatial acoustic parameters such as time difference of arrival, direct-to-reverberant ratio and reverberation time. However, they still suffer from the simulation-to-reality generalization problem due to the mismatch between simulated and real-world acoustic characteristics and the deficiency of annotated real-world data. To this end, this work proposes a self-supervised method that takes full advantage of unlabeled data for spatial acoustic parameter estimation. First, a new pretext task, i.e. cross-channel signal reconstruction (CCSR), is designed to learn a universal spatial acoustic representation from unlabeled multi-channel microphone signals. We mask partial signals of one channel and ask the model to reconstruct them, which makes it possible to learn spatial acoustic information from unmasked signals and extract source information from the other microphone channel. An encoder-decoder structure is used to disentangle the two kinds of information. By fine-tuning the pre-trained spatial encoder with a small annotated dataset, this encoder can be used to estimate spatial acoustic parameters. Second, a novel multi-channel audio Conformer (MC-Conformer) is adopted as the encoder model architecture, which is suitable for both the pretext and downstream tasks. It is carefully designed to be able to capture the local and global characteristics of spatial acoustics exhibited in the time-frequency domain. Experimental results of five acoustic parameter estimation tasks on both simulated and real-world data show the effectiveness of the proposed method. To the best of our knowledge, this is the first self-supervised learning method in the field of spatial acoustic representation learning and multi-channel audio signal processing.
Bing Yang 0004, Xiaofei Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 SRP-DNN: Learning Direct-Path Phase Difference for Multiple Moving Sound Source Localization
abstract
Multiple moving sound source localization in real-world scenarios remains a challenging issue due to interaction between sources, time-varying trajectories, distorted spatial cues, etc. In this work, we propose to use deep learning techniques to learn competing and time-varying direct-path phase differences for localizing multiple moving sound sources. A causal convolutional recurrent neural network is designed to extract the direct-path phase difference sequence from signals of each microphone pair. To avoid the assignment ambiguity and the problem of uncertain output-dimension encountered when simultaneously predicting multiple targets, the learning target is designed in a weighted sum format, which encodes source activity in the weight and direct-path phase differences in the summed value. The learned direct-path phase differences for all microphone pairs can be directly used to construct the spatial spectrum according to the formulation of steered response power (SRP). This deep neural network (DNN) based SRP method is referred to as SRP-DNN. The locations of sources are estimated by iteratively detecting and removing the dominant source from the spatial spectrum, in which way the interaction between sources is reduced. Experimental results on both simulated and real-world data show the superiority of the proposed method in the presence of noise and reverberation.
Bing Yang 0004, Hong Liu 0008, Xiaofei Li 0001
ICASSP3
2021 Fullsubnet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement
abstract
This paper proposes a full-band and sub-band fusion model, named as FullSubNet, for single-channel real-time speech enhancement. Full-band and sub-band refer to the models that input full-band and sub-band noisy spectral feature, output full-band and sub-band speech target, respectively. The sub-band model processes each frequency independently. Its input consists of one frequency and several context frequencies. The output is the prediction of the clean speech target for the corresponding frequency. These two types of models have distinct characteristics. The full-band model can capture the global spectral context and the long-distance cross-band dependencies. However, it lacks the ability to modeling signal stationarity and attending the local spectral pattern. The sub-band model is just the opposite. In our proposed FullSubNet, we connect a pure full-band model and a pure sub-band model sequentially and use practical joint training to integrate these two types of models' advantages. We conducted experiments on the DNS challenge (INTERSPEECH 2020) dataset to evaluate the proposed method. Experimental results show that full-band and sub-band information are complementary, and the FullSubNet can effectively integrate them. Besides, the performance of the FullSubNet also exceeds that of the top-ranked methods in the DNS Challenge (INTERSPEECH 2020).
Xiangdong Su, Radu Horaud, Xiaofei Li 0001
ICASSP4
2021 Supervised Direct-Path Relative Transfer Function Learning for Binaural Sound Source Localization
abstract
Direct-path relative transfer function (DP-RTF) refers to the ratio between the direct-path acoustic transfer functions of two channels. Though DP-RTF fully encodes the sound directional cues and serves as a reliable localization feature, it is often erroneously estimated in the presence of noise and reverberation. This paper proposes a supervised DP-RTF learning method with deep neural networks for robust binaural sound source localization. To exploit the complementarity of single-channel spectrogram and dual-channel difference information, we first recover the direct-path magnitude spectrogram from the contaminated one using a monaural enhancement network, and then predict the DP-RTF from the dual-channel (enhanced-) intensity and phase cues using a binaural enhancement network. In addition, a weighted-matching softmax training loss is designed to promote the predicted DP-RTFs to be concentrated for the same direction and separated for different directions. Finally, the direction of arrival (DOA) of source is estimated by matching the predicted DP-RTF with the ground truths of candidate directions. Experimental results show the superiority of our method for DOA estimation in the environments with various levels of noise and reverberation.
Bing Yang 0004, Xiaofei Li 0001, Hong Liu 0008
ICASSP2
2021 AcousticFusion: Fusing Sound Source Localization to Visual SLAM in Dynamic Environments
abstract
Dynamic objects in the environment, such as people and other agents, lead to challenges for existing simultaneous localization and mapping (SLAM) approaches. To deal with dynamic environments, computer vision researchers usually apply some learning-based object detectors to remove these dynamic objects. However, these object detectors are computationally too expensive for mobile robot on-board processing. In practical applications, these objects output noisy sounds that can be effectively detected by on-board sound source localization. The directional information of the sound source object can be efficiently obtained by direction of sound arrival (DoA) estimation, but the depth estimation is difficult. Therefore, in this paper, we propose a novel audio-visual fusion approach that fuses sound source direction into the RGB-D image and thus removes the effect of dynamic obstacles on the multi-robot SLAM system. Experimental results of multirobot SLAM in different dynamic environments show that the proposed method uses very small computational resources to obtain very stable self-localization results.
Tianwei Zhang 0002, Huayan Zhang, Xiaofei Li 0001, Tin Lun Lam, Sethu Vijayakumar
IROS3
2021 Learning Deep Direct-Path Relative Transfer Function for Binaural Sound Source Localization
abstract
Direct-path relative transfer function (DP-RTF) refers to the ratio between the direct-path acoustic transfer functions of two microphone channels. Though DP-RTF fully encodes the sound spatial cues and serves as a reliable localization feature, it is often erroneously estimated in the presence of noise and reverberation. This paper proposes to learn DP-RTF with deep neural networks for robust binaural sound source localization. A DP-RTF learning network is designed to regress the binaural sensor signals to a real-valued representation of DP-RTF. It consists of a branched convolutional neural network module to separately extract the inter-channel magnitude and phase patterns, and a convolutional recurrent neural network module for joint feature learning. To better explore the speech spectra to aid the DP-RTF estimation, a monaural speech enhancement network is used to recover the direct-path spectrograms from the noisy ones. The enhanced spectrograms are stacked onto the noisy spectrograms to act as the input of the DP-RTF learning network. We train one unique DP-RTF learning network using many different binaural arrays to enable the generalization of DP-RTF learning across arrays. This way avoids time-consuming training data collection and network retraining for a new array, which is very useful in practical application. Experimental results on both simulated and real-world data show the effectiveness of the proposed method for direction of arrival (DOA) estimation in the noisy and reverberant environment, and a good generalization ability to unseen binaural arrays.
Bing Yang 0004, Hong Liu 0008, Xiaofei Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 A Covert Ultrasonic Phone-to-Phone Communication Scheme
Liming Shi, Limin Yu, Kaizhu Huang, Xu Zhu 0001, Zhi Wang 0003, Xiaofei Li 0001, Wenwu Wang 0001, Xinheng Wang 0001
CollaborateCom (1)6
2020 Online Monaural Speech Enhancement Using Delayed Subband LSTM
abstract
International audience
Xiaofei Li 0001, Radu Horaud
INTERSPEECH1
2019 Audio-Visual Variational Fusion for Multi-Person Tracking with Robots
abstract
Robust multi-person tracking with robots opens the door to analysing engagement and social signals in real-world environments. Multi-person scenarios are charaterised by (i) a time-varying number of people, (ii) intermittent auditory (\eg speech turns) and visual cues (\eg person appearing/disappearing) and (iii) impact of the robot actions in perception. The various sensors (cameras and microphones) available for perception, provide a rich flow of information of intermittent and complementary nature. How to jointly exploit these cues to tackle the multi-person tracking problem with an autonomous system has been an intense research line of the Perception Team in the past few years. In this demo we want to present our, now mature, achievements in the field, and demonstrate two robotic systems able to track multiple persons using auditory and visual cues, when they are available. We will bring the two robots and the necessary computing resources with us, as well as the required presentation materials to discuss the models, methods and tools supporting this technology with the attendants.
Xavier Alameda-Pineda, Soraya Arias, Yutong Ban, Guillaume Delorme 0002, Laurent Girin, Radu Horaud, Xiaofei Li 0001, Bastien Mourgue, Guillaume Sarrazin
ACM Multimedia7
2019 Audio-Noise Power Spectral Density Estimation Using Long Short-Term Memory
abstract
We propose a method using a long short-term memory (LSTM) network to estimate the noise power spectral density (PSD) of single-channel audio signals represented in the short-time Fourier transform (STFT) domain. An LSTM network common to all frequency bands is trained, which processes each frequency band individually by mapping the noisy STFT magnitude sequence to its corresponding noise PSD sequence. Unlike deep-learning-based speech-enhancement methods, which learn the full-band spectral structure of speech segments, the proposed method exploits the sub-band STFT magnitude evolution of noise with long time dependence, in the spirit of the unsupervised noise estimators described in the literature. Speaker- and speech-independent experiments with different types of noise show that the proposed method outperforms the unsupervised estimators, and it generalizes well to noise types that are not present in the training set.
Xiaofei Li 0001, Simon Leglaive, Laurent Girin, Radu Horaud
IEEE Signal Process. Lett.1
2019 Multichannel Speech Separation and Enhancement Using the Convolutive Transfer Function
abstract
This paper addresses the problem of speech separation and enhancement from multichannel convolutive and noisy mixtures, assuming known mixing filters. We propose to perform speech separation and enhancement in the short-time Fourier transform domain using the convolutive transfer function (CTF) approximation. Compared to time-domain filters, the CTF has much less taps. Consequently, it requires less computational cost and sometimes is more robust against the filter perturbations. We propose three methods: 1) for the multisource case, the multichannel inverse filtering method, i.e., the multiple input/output inverse theorem (MINT), is exploited in the CTF domain; 2) a beamforming-like multichannel inverse filtering method applying the single-source MINT and using power minimization, which is suitable whenever the source CTFs are not all known; and 3) a basis pursuit method, where the sources are recovered by minimizing their ℓ1-norm to impose spectral sparsity, while the ℓ2-norm fitting cost between microphone signals and mixing model is constrained to be lower than a tolerance. The noise can be reduced by setting this tolerance at the noise power level. Experiments under various acoustic conditions are carried out to evaluate and compare the three proposed methods. Comparison with four baseline methods-beamforming-based, two time-domain inverse filters, and time-domain Lasso-shows the applicability of the proposed methods.
Xiaofei Li 0001, Laurent Girin, Sharon Gannot, Radu Horaud
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 Multichannel Online Dereverberation Based on Spectral Magnitude Inverse Filtering
abstract
This paper addresses the problem of multichannel online dereverberation. The proposed method is carried out in the short-time Fourier transform (STFT) domain, and for each frequency band independently. In the STFT domain, the time-domain room impulse response is approximately represented by the convolutive transfer function (CTF). The multichannel CTFs are adaptively identified based on the cross-relation method, and using the recursive least square criterion. Instead of the complex-valued CTF convolution model, we use a nonnegative convolution model between the STFT magnitude of the source signal and the CTF magnitude, which is just a coarse approximation of the former model, but is shown to be more robust against the CTF perturbations. Based on this nonnegative model, we propose an online STFT magnitude inverse filtering method. The inverse filters of the CTF magnitude are formulated based on the multiple-input/output inverse theorem, and adaptively estimated based on the gradient descent criterion. Finally, the inverse filtering is applied to the STFT magnitude of the microphone signals, obtaining an estimate of the STFT magnitude of the source signal. Experiments regarding both speech enhancement and automatic speech recognition are conducted, which demonstrate that the proposed method can effectively suppress reverberation, even for the difficult case of a moving speaker.
Xiaofei Li 0001, Laurent Girin, Sharon Gannot, Radu Horaud
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 Multiple Sound Source Counting and Localization Based on TF-Wise Spatial Spectrum Clustering
abstract
This paper addresses the problem of multiple sound source counting and localization in adverse acoustic environments, using microphone array recordings. The proposed time-frequency (TF) wise spatial spectrum clustering based method contains two stages. First, given the received sensor signals, the spatial correlation matrix is computed and denoised in the TF domain. The TF-wise spatial spectrum is estimated based on the signal subspace information, and further enhanced by an exponential transform, which can increase the reliability of the source presence possibility reflected by spatial spectrum. Second, to jointly count and localize sound sources, the enhanced TF-wise spatial spectra are divided into several clusters with each cluster corresponding to one source. Sources are successively detected by searching the significant peaks of the remaining global spatial spectrum, which is formed using unassigned spatial spectra. After each new source detection, spatial spectra are reassigned to detected sources according to the dominance association between them. The interaction between sources is reduced by iteratively performing new source detection and spatial spectrum assignment. Experiments on both simulated data and real-world data demonstrate the superiority of the proposed method for multiple sound source counting and localization in the environment with different levels of noise and reverberation.
Bing Yang 0004, Hong Liu 0008, Xiaofei Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2018 Accounting for Room Acoustics in Audio-Visual Multi-Speaker Tracking
abstract
Multiple-speaker tracking is a crucial task for many applications. In real-world scenarios, exploiting the complementarity between auditory and visual data enables to track people outside the visual field of view. However, practical methods must be robust to changes in acoustic conditions, e.g. reverberation. We investigate how to combine state-of-the-art audio-source localization techniques with Bayesian multi-person tracking. Our experiments demonstrate that the performance of the proposed system is not affected by changes in the acoustic environment.
Yutong Ban, Xiaofei Li 0001, Xavier Alameda-Pineda, Laurent Girin, Radu Horaud
ICASSP2
2018 Multisource Mint Using Convolutive Transfer Function
abstract
The multichannel inverse filtering method, i.e. multiple input/output inverse theorem (MINT), is widely used. However, it is usually performed in the time domain, and based on the long room impulse responses, thus it has a high computational complexity and a large number of near-common zeros. In this paper, we propose to perform MINT in the short-time Fourier transform (STFT) domain, in which the time-domain filter is approximated by the convolutive transfer function. The oversampled STFT is used to avoid frequency aliasing, which however leads to a common zero region in the subband frequency response due to the frequency response of the STFT window. A new inverse filtering target function concerning the STFT window is proposed to overcome this problem. In addition, unlike most studies using MINT for single source dereverberation, the multisource MINT is proposed for both source separation and dereverberation.
Xiaofei Li 0001, Sharon Gannot, Laurent Girin, Radu Horaud
ICASSP1
2018 Audio-Visual Speaker Diarization Based on Spatiotemporal Bayesian Fusion
abstract
Speaker diarization consists of assigning speech signals to people engaged in a dialogue. An audio-visual spatiotemporal diarization model is proposed. The model is well suited for challenging scenarios that consist of several participants engaged in multi-party interaction while they move around and turn their heads towards the other participants rather than facing the cameras and the microphones. Multiple-person visual tracking is combined with multiple speech-source localization in order to tackle the speech-to-person association problem. The latter is solved within a novel audio-visual fusion method on the following grounds: binaural spectral features are first extracted from a microphone pair, then a supervised audio-visual alignment technique maps these features onto an image, and finally a semi-supervised clustering method assigns binaural spectral features to visible persons. The main advantage of this method over previous work is that it processes in a principled way speech signals uttered simultaneously by multiple persons. The diarization itself is cast into a latent-variable temporal graphical model that infers speaker identities and speech turns, based on the output of an audio-visual association process, executed at each time slice, and on the dynamics of the diarization variable itself. The proposed formulation yields an efficient exact inference procedure. A novel dataset, that contains audio-visual training data as well as a number of scenarios involving several participants engaged in formal and informal dialogue, is introduced. The proposed method is thoroughly tested and benchmarked with respect to several state-of-the art diarization algorithms.
Israel D. Gebru, Sileye O. Ba, Xiaofei Li 0001, Radu Horaud
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 Multichannel Identification and Nonnegative Equalization for Dereverberation and Noise Reduction Based on Convolutive Transfer Function
abstract
This paper addresses the problems of blind multichannel identification and equalization for joint speech dereverberation and noise reduction. The time-domain cross-relation method is hardly applicable for blind room impulse response identification due to the near-common zeros of the long impulse responses. We extend the cross-relation method to the short-time Fourier transform (STFT) domain, in which the time-domain impulse response is approximately represented by the convolutive transfer function (CTF) with much less coefficients. For the oversampled STFT, CTFs suffer from the common zeros caused by the nonflat frequency response of the STFT window. To overcome this, we propose to identify CTFs using the STFT framework with oversampled signals and critically sampled CTFs, which is a good tradeoff between the frequency aliasing of the signals and the common zeros problem of CTFs. The identified complex-valued CTFs are not accurate enough for multichannel equalization due to the frequency aliasing of the CTFs. Hence, we only use the CTF magnitudes, which leads to a nonnegative multichannel equalization method based on a nonnegative convolution model between the STFT magnitude of the source signal and the CTF magnitude. Compared with the complex-valued convolution model, this nonnegative convolution model is shown to be more robust against the CTF perturbations. To recover the STFT magnitude of the source signal and to reduce the additive noise, the l2-norm fitting error between the STFT magnitude of the microphone signals and the nonnegative convolution is constrained to be less than a noise power related tolerance. Meanwhile, the l1-norm of the STFT magnitude of the source signal is minimized to impose the sparsity.
Xiaofei Li 0001, Sharon Gannot, Laurent Girin, Radu Horaud
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 Audio source separation based on convolutive transfer function and frequency-domain lasso optimization
abstract
This paper addresses the problem of under-determined convolutive audio source separation in a semi-oracle configuration where the mixing filters are assumed to be known. We propose a separation procedure based on the convolutive transfer function (CTF), which is a more appropriate model for strongly reverberant signals than the widely-used multiplicative transfer function approximation. In the short-time Fourier transform domain, source signals are estimated by minimizing the mixture fitting cost using Lasso optimization, with a ℓ1-norm regularization to exploit the spectral sparsity of source signals. Experiments show that the proposed method achieves satisfactory performance on highly reverberant speech mixtures, with a much lower computational cost compared to time-domain dual techniques.
Xiaofei Li 0001, Laurent Girin, Radu Horaud
ICASSP1
2017 Multiple-Speaker Localization Based on Direct-Path Features and Likelihood Maximization With Spatial Sparsity Regularization
abstract
This paper addresses the problem of multiple-speaker localization in noisy and reverberant environments, using binaural recordings of an acoustic scene. A complex-valued Gaussian mixture model (CGMM) is adopted, whose components correspond to all the possible candidate source locations defined on a grid. After optimizing the CGMM-based objective function, given an observed set of complex-valued binaural features, both the number of sources and their locations are estimated by selecting the CGMM components with the largest weights. An entropy-based penalty term is added to the likelihood to impose sparsity over the set of CGMM component weights. This favors a small number of detected speakers with respect to the large number of initial candidate source locations. In addition, the direct-path relative transfer function (DP-RTF) is used to build robust binaural features. The DP-RTF, recently proposed for single-source localization, encodes interchannel information corresponding to the direct path of sound propagation and is thus robust to reverberations. In this paper, we extend the DP-RTF estimation to the case of multiple sources. In the short-time Fourier transform domain, a consistency test is proposed to check whether a set of consecutive frames is associated with the same source or not. Reliable DP-RTF features are selected from the frames that pass the consistency test to be used for source localization. Experiments carried out using both simulation data and real data recorded with a robotic head confirm the efficiency of the proposed multisource localization method.
Xiaofei Li 0001, Laurent Girin, Radu Horaud, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 Binaural Sound Localization Based on Reverberation Weighting and Generalized Parametric Mapping
abstract
Binaural sound source localization is an important technique for speech enhancement, video conferencing, and human-robot interaction, etc. However, in realistic scenarios, the reverberation and environmental noise would degrade the precision of sound direction estimation. Therefore, reliable sound localization is essential to practical applications. To deal with these disturbances, this paper presents a novel binaural sound source localization approach based on reverberation weighting and generalized parametric mapping. First, the reverberation weighting as a preprocessing stage, is used to separately suppress the early and late reverberation, while preserving interaural cues. Then, two binaural cues, i.e., interaural time and intensity differences, are extracted from the frequency-domain representations of dereverberated binaural signals for the online localization. Their corresponding templates are established using the training data. Furthermore, the generalized parametric mapping is proposed to build a generalized parametric model for describing relationships between azimuth and binaural cues analytically. Finally, a two-step sound localization process is introduced to refine azimuth estimation based on the generalized parametric model and template matching. Experiments in both simulated and real scenarios validate that the proposed method can achieve better localization performance compared to state-of-the-art methods.
Hong Liu 0008, Jie Zhang 0042, Xiaofei Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 Non-stationary noise power spectral density estimation based on regional statistics
abstract
Estimating the noise power spectral density (PSD) is essential for single channel speech enhancement algorithms. In this paper, we propose a noise PSD estimation approach based on regional statistics. The proposed regional statistics consist of four features representing the statistics of the past and present periodograms in a short-time period. We show that these features are efficient in characterizing the statistical difference between noise PSD and noisy speech PSD. We therefore propose to use these features for estimating the speech presence probability (SPP). The noise PSD is recursively estimated by averaging past spectral power values with a time-varying smoothing parameter controlled by the SPP. The proposed method exhibits good tracking capability for non-stationary noise, even for abruptly increasing noise level.
Xiaofei Li 0001, Laurent Girin, Sharon Gannot, Radu Horaud
ICASSP1
2016 Reverberant sound localization with a robot head based on direct-path relative transfer function
abstract
This paper addresses the problem of sound-source localization (SSL) with a robot head, which remains a challenge in real-world environments. In particular we are interested in locating speech sources, as they are of high interest for human-robot interaction. The microphone-pair response corresponding to the direct-path sound propagation is a function of the source direction. In practice, this response is contaminated by noise and reverberations. The direct-path relative transfer function (DP-RTF) is defined as the ratio between the direct-path acoustic transfer function (ATF) of the two microphones, and it is an important feature for SSL. We propose a method to estimate the DP-RTF from noisy and reverberant signals in the short-time Fourier transform (STFT) domain. First, the convolutive transfer function (CTF) approximation is adopted to accurately represent the impulse response of the microphone array, and the first coefficient of the CTF is mainly composed of the direct-path ATF. At each frequency, the frame-wise speech auto- and cross-power spectral density (PSD) are obtained by spectral subtraction. Then a set of linear equations is constructed by the speech auto- and cross-PSD of multiple frames, in which the DP-RTF is an unknown variable, and is estimated by solving the equations. Finally, the estimated DP-RTFs are concatenated across frequencies and used as a feature vector for SSL. Experiments with a robot, placed in various reverberant environments, show that the proposed method outperforms two state-of-the-art methods.
Xiaofei Li 0001, Laurent Girin, Fabien Badeig, Radu Horaud
IROS1
2016 Estimation of the Direct-Path Relative Transfer Function for Supervised Sound-Source Localization
abstract
This paper addresses the problem of sound-source localization of a single speech source in noisy and reverberant environments. For a given binaural microphone setup, the binaural response corresponding to the direct-path propagation of a single source is a function of the source direction. In practice, this response is contaminated by noise and reverberations. The direct-path relative transfer function (DP-RTF) is defined as the ratio between the direct-path acoustic transfer function of the two channels. We propose a method to estimate the DP-RTF from the noisy and reverberant microphone signals in the short-time Fourier transform (STFT) domain. First, the convolutive transfer function approximation is adopted to accurately represent the impulse response of the sensors in the STFT domain. Second, the DP-RTF is estimated by using the auto- and cross-power spectral densities at each frequency and over multiple frames. In the presence of stationary noise, an interframe spectral subtraction algorithm is proposed, which enables to achieve the estimation of noise-free auto- and cross-power spectral densities. Finally, the estimated DP-RTFs are concatenated across frequencies and used as a feature vector for the localization of speech source. Experiments with both simulated and real data show that the proposed localization method performs well, even under severe adverse acoustic conditions, and outperforms state-of-the-art localization methods under most of the acoustic conditions.
Xiaofei Li 0001, Laurent Girin, Radu Horaud, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 A Novel Lip Descriptor for Audio-Visual Keyword Spotting Based on Adaptive Decision Fusion
abstract
Keyword spotting remains a challenge when applied to real-world environments with dramatically changing noise. In recent studies, audio-visual integration methods have demonstrated superiorities since visual speech is not influenced by acoustic noise. However, for visual speech recognition, individual utterance mannerisms can lead to confusion and false recognition. To solve this problem, a novel lip descriptor is presented involving both geometry-based and appearance-based features in this paper. Specifically, a set of geometry-based features is proposed based on an advanced facial landmark localization method. In order to obtain robust and discriminative representation, a spatiotemporal lip feature is put forward concerning similarities among textons and mapping the feature to intra-class subspace. Moreover, a parallel two-step keyword spotting strategy based on decision fusion is proposed in order to make the best use of audio-visual speech and adapt to diverse noise conditions. Weights generated using a neural network combine acoustic and visual contributions. Experimental results on the OuluVS dataset and PKU-AV dataset demonstrate that the proposed lip descriptor shows competitive performance compared to the state of the art. Additionally, the proposed audio-visual keyword spotting (AV-KWS) method based on decision-level fusion significantly improves the noise robustness and attains better performance than feature-level fusion, which is also capable of adapting to various noisy conditions.
Hong Liu 0008, Xiaofei Li 0001, Ting Fan, Xuewu Zhang 0003
IEEE Trans. Multim.3
2015 Estimation of relative transfer function in the presence of stationary noise based on segmental power spectral density matrix subtraction
abstract
This paper addresses the problem of relative transfer function (RTF) estimation in the presence of stationary noise. We propose an RTF identification method based on segmental power spectral density (PSD) matrix subtraction. First multiple channel microphone signals are divided into segments corresponding to speech-plus-noise activity and noise-only. Then, the subtraction of two segmental PSD matrices leads to an almost noise-free PSD matrix by reducing the stationary noise component and preserving non-stationary speech component. This noise-free PSD matrix is used for single speaker RTF identification by eigenvalue decomposition. Experiments are performed in the context of sound source localization to evaluate the efficiency of the proposed method.
Xiaofei Li 0001, Laurent Girin, Radu Horaud, Sharon Gannot
ICASSP1
2015 A Distributed Architecture for Interacting with NAO
abstract
One of the main applications of the humanoid robot NAO - a small robot companion - is human-robot interaction (HRI). NAO is particularly well suited for HRI applications because of its design, hardware specifications, programming capabilities, and affordable cost. Indeed, NAO can stand up, walk, wander, dance, play soccer, sit down, recognize and grasp simple objects, detect and identify people, localize sounds, understand some spoken words, engage itself in simple and goal-directed dialogs, and synthesize speech. This is made possible due to the robot's 24 degree-of-freedom articulated structure (body, legs, feet, arms, hands, head, etc.), motors, cameras, microphones, etc., as well as to its on-board computing hardware and embedded software, e.g., robot motion control. Nevertheless, the current NAO configuration has two drawbacks that restrict the complexity of interactive behaviors that could potentially be implemented. Firstly, the on-board computing resources are inherently limited, which implies that it is difficult to implement sophisticated computer vision and audio signal analysis algorithms required by advanced interactive tasks. Secondly, programming new robot functionalities currently implies the development of embedded software, which is a difficult task in its own right necessitating specialized knowledge. The vast majority of HRI practitioners may not have this kind of expertise and hence they cannot easily and quickly implement their ideas, carry out thorough experimental validations, and design proof-of-concept demonstrators. We have developed a distributed software architecture that attempts to overcome these two limitations. Broadly speaking, NAO's on-board computing resources are augmented with external computing resources. The latter is a computer platform with its CPUs, GPUs, memory, operating system, libraries, software packages, internet access, etc. This configuration enables easy and fast development in Matlab, C, C++, or Python. Moreover, it allows the user to combine on-board libraries (motion control, face detection, etc.) with external toolboxes, e.g., OpenCv.
Fabien Badeig, Quentin Pelorson, Soraya Arias, Vincent Drouard, Israel D. Gebru, Xiaofei Li 0001, Georgios Evangelidis 0002, Radu Horaud
ICMI6
2013 A two-layer probabilistic model based on time-delay compensation for binaural sound localization
abstract
Interaural Intensity Difference (IID) and Interaural Time Difference (ITD) are two important cues for robot acoustic localization both in Artificial Intelligence (AI) and Human-Robot Interaction (HRI) areas. However, it is a challenge job to localize a sound source accurately and swiftly only by two acoustic sensors. In this paper, a time-delay compensation based two-layer probabilistic model is presented for binaural sound source localization. In the first layer, a weighting function of Generalized Cross Correlation (GCC) named PHAT-ργ is used in low-frequency to obtain the prior time-delay. And in this layer a crude estimate of azimuth can also be acquired. At the same time, the probability of all possible time-delay lags can be achieved from the training data. In the Second layer, a new improved algorithm of IID based on time-delay compensation(named IIDδτ) is introduced to refine the probability of the azimuth and the elevation. Lastly, localization result is obtained by Bayes-Rule method. Comparing with three state-of-art algorithms, experimental results show that the proposed method has higher accuracy and costs less time for sound source localization.
Hong Liu 0008, Zhuo Fu, Xiaofei Li 0001
ICRA3
2013 Sound Source Localization for HRI Using FOC-Based Time Difference Feature and Spatial Grid Matching
abstract
In human-robot interaction (HRI), speech sound source localization (SSL) is a convenient and efficient way to obtain the relative position between a speaker and a robot. However, implementing a SSL system based on TDOA method encounters many problems, such as noise of real environments, the solution of nonlinear equations, switch between far field and near field. In this paper, fourth-order cumulant spectrum is derived, based on which a time delay estimation (TDE) algorithm that is available for speech signal and immune to spatially correlated Gaussian noise is proposed. Furthermore, time difference feature of sound source and its spatial distribution are analyzed, and a spatial grid matching (SGM) algorithm is proposed for localization step, which handles some problems that geometric positioning method faces effectively. Valid feature detection algorithm and a decision tree method are also suggested to improve localization performance and reduce computational complexity. Experiments are carried out in real environments on a mobile robot platform, in which thousands of sets of speech data with noise collected by four microphones are tested in 3D space. The effectiveness of our TDE method and SGM algorithm is verified.
Xiaofei Li 0001, Hong Liu 0008
IEEE Trans. Cybern.1
2012 Time Delay Estimation for Speech Signal Based on FOC-Spectrum
abstract
Higher-order statistics can be used for time delay estimation (TDE) to suppress spatially correlated Gaussian noise, since the higher-order cumulant of Gaussian signal is always zero. However, third-order statistics is invalid for those signals with zero skewness, speech signal as a typical one. In this paper, the fourth-order cumulant (FOC) spectrum is derived, based on which a TDE algorithm that is valid for speech signal and immune to spatially correlated Gaussian noise is proposed. This method can estimate the time delay between two sensor signals or simultaneously estimate the time delays between one sensor signal and other three. In addition, just like generalized cross correlation method, this spectrum domain algorithm is more robust than time domain FOC-based TDE algorithm, especially for speech signal due to its periodicity. Experiments verify the effectiveness of this TDE method for speech signal with spatially correlated Gaussian noise.
Hong Liu 0008, Xiaofei Li 0001
INTERSPEECH2
2011 Sound source localization for mobile robot based on time difference feature and space grid matching
abstract
Auditory is a convenient and efficient way for Human-Robot Interaction, however implementing a sound source localization system based on TDOA method encounters many problems, such as noise of real environments, and resolution of nonlinear equations, switch between far field and near field and lack of microphones for geometric positioning localization method. In this paper, a new spectral weighting GCC-PHAT method is proposed to deal with noise. Furthermore, the time difference feature of sound source and its spatial distribution are analyzed. Based on prosperities of the distribution, a space grid matching (SGM) algorithm is proposed for localization step, which handles those problems that geometric positioning method faces effectively. Decision tree and valid feature detection algorithm are also proposed to reduce computational complexity and improve performance. Experiments are achieved in real environments on a mobile robot platform, in which 2016 sets of speech data are tested using four microphones in 3D space. More than 95% azimuth localization rate with error less than 5 degrees and approximate 90% horizontal distance localization rate are obtained.
Xiaofei Li 0001, Hong Liu 0008, Xuesong Yang
IROS1
2010 A selection method of speech vocabulary for human-robot speech interaction
abstract
Speech is the most natural and efficient way for Human-Robot Interaction (HRI), although speech recognition systems face some challenges on a mobile robot platform due to the wide range of users and varied noisy environments. This paper proposes a selection method of speech vocabulary for HRI, which can choose the most robust sub-vocabulary from the predefined isolated word vocabulary. We define a new concept, called Word Robustness, to represent the robustness of a word to speaker-independent and noise. The algorithm for computing Word Robustness is given based on Hidden Markov Model (HMM), then the most robust sub-vocabulary can be selected based on it. For convenience, this method makes use of selecting vocabulary to avoid other procedures such as speaker-adaption. Experiments were achieved based on an isolated word recognition system using a speech database which includes more than ten thousands of speech signals recorded in quiet laboratory and noisy environment respectively. Several sub-sets of vocabulary were selected for robot control based on Word Robustness. The best speaker-independent word recognition rate is 95.19% in noisy environments. Experimental results demonstrate the effectiveness of Word Robustness and the selection method.
Hong Liu 0008, Xiaofei Li 0001
SMC2