EDBT 2026 Demo / reviewers in the wild / expert
Qiang Fu 0001
dblp:17/1352-1
· DBLP profile ↗
21ranked-venue papers
0as first author
8since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 8 since 2021Artificial intelligence and machine learning · 10 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual ConformerabstractIn recent years, neural network-based Wake Word Spotting achieves good performance on clean audio samples but struggles in noisy environments. Audio-Visual Wake Word Spotting (AVWWS) receives lots of attention because visual lip movement information is not affected by complex acoustic scenes. Previous works usually use simple addition or concatenation for multi-modal fusion. The inter-modal correlation remains relatively under-explored. In this paper, we propose a novel module called Frame-Level Cross-Modal Attention (FLCMA) to improve the performance of AVWWS systems. This module can help model multi-modal information at the frame-level through synchronous lip movements and speech signals. We train the end-to-end FLCMA based Audio-Visual Conformer and further improve the performance by fine-tuning pre-trained uni-modal models for the AVWWS task. The proposed system achieves a new state-of-the-art result (4.57% WWS score) on the far-field MISP dataset. Haoxu Wang, Ming Cheng 0005, Qiang Fu 0001, Ming Li 0026 |
ICASSP | 3 |
| 2023 | The WHU-Alibaba Audio-Visual Speaker Diarization System for the MISP 2022 ChallengeabstractThis paper describes the system developed by the WHU-Alibaba team for the Multimodal Information Based Speech Processing (MISP) 2022 Challenge. We extend the Sequence-to-Sequence Target-Speaker Voice Activity Detection framework to simultaneously detect multiple speakers’ voice activities from audio-visual signals. The final system achieves a diarization error rate (DER) of 8.82% on the evaluation set of the competition database, which ranks 1st in the speaker diarization track of the MISP 2022, ICASSP Signal Processing Grand Challenge. Ming Cheng 0005, Haoxu Wang, Qiang Fu 0001, Ming Li 0026 |
ICASSP | 4 |
| 2023 | The DKU Post-Challenge Audio-Visual Wake Word Spotting System for the 2021 MISP Challenge: Deep AnalysisabstractThis paper further explores our previous wake word spotting system ranked 2-nd in Track 1 of the MISP Challenge 2021. First, we investigate a robust unimodal approach based on 3D and 2D convolution and adopt the simple attention module (SimAM) for our system to improve performance. Second, we explore different combinations of data augmentation methods for better performance. Finally, we study the fusion strategies, including score-level, cascaded and neural fusion. Our proposed multimodal system leverages multimodal features and uses the complementary visual information to mitigate the performance degradation of audio-only systems in complex acoustic scenarios. Our system obtains a false reject rate of 2.15% and a false alarm rate of 3.44% in the evaluation set of the competition database, which achieves the new state-of-the-art performance by 21% relative improvement compared to previous systems. Related resource can be found at: https://github.com/Mashiro009/DKU_WWS_MISP. Haoxu Wang, Ming Cheng 0005, Qiang Fu 0001, Ming Li 0026 |
ICASSP | 3 |
| 2023 | Small Footprint Multi-channel Network for Keyword Spotting with Centroid Based Awareness
Dianwen Ng, Yang Xiao 0019, Jia Qi Yip, Biao Tian 0002, Qiang Fu 0001, Chng Eng Siong, Bin Ma 0001 |
INTERSPEECH | 6 |
| 2022 | Convmixer: Feature Interactive Convolution with Curriculum Learning for Small Footprint and Noisy Far-Field Keyword SpottingabstractBuilding efficient architecture in neural speech processing is paramount to success in keyword spotting deployment. However, it is very challenging for lightweight models to achieve noise robustness with concise neural operations. In a real-world application, the user environment is typically noisy and may contain reverberations. We proposed a novel feature interactive convolutional model with merely 100K parameters to tackle this under the noisy far-field condition. The interactive unit is proposed in place of the attention module that promotes the flow of information with more efficient computations. Moreover, curriculum-based multi-condition training is adopted to attain better noise robustness. Our model achieves 98.2% top-1 accuracy on Google Speech Command V2-12 and is competitive against large transformer models under the designed noise condition. Dianwen Ng, Yunqi Chen, Biao Tian 0002, Qiang Fu 0001, Chng Eng Siong |
ICASSP | 4 |
| 2022 | NN3A: Neural Network Supported Acoustic Echo Cancellation, Noise Suppression and Automatic Gain Control for Real-Time CommunicationsabstractAcoustic echo cancellation (AEC), noise suppression (NS) and automatic gain control (AGC) are three often required modules for real-time communications (RTC). This paper proposes a neural network supported algorithm for RTC, namely NN3A, which incorporates an adaptive filter and a multi-task model for residual echo suppression, noise reduction and near-end speech activity detection. The proposed algorithm is shown to outperform both a method using separate models and an end-to-end alternative. It is further shown that there exists a trade-off in the model between residual suppression and near-end speech distortion, which could be balanced by a novel loss weighting function. Several practical aspects of training the joint model are also investigated to push its performance to limit. Yueyue Na, Biao Tian 0002, Qiang Fu 0001 |
ICASSP | 4 |
| 2021 | Weighted Recursive Least Square Filter and Neural Network Based Residual ECHO Suppression for the AEC-ChallengeabstractThis paper presents a real-time Acoustic Echo Cancellation (AEC) algorithm submitted to the AEC-Challenge. The algorithm consists of three modules: Generalized Cross-Correlation with PHAse Transform (GCC-PHAT) based time delay compensation, weighted Recursive Least Square (wRLS) based linear adaptive filtering and neural network based residual echo suppression. The wRLS filter is derived from a novel semi-blind source separation perspective. The neural network model predicts a Phase-Sensitive Mask (PSM) based on the aligned reference and the linear filter output. The algorithm achieved a mean subjective score of 4.00 and ranked 2nd in the AEC-Challenge. Yueyue Na, Zhang Liu 0006, Biao Tian 0002, Qiang Fu 0001 |
ICASSP | 5 |
| 2021 | Joint Online Multichannel Acoustic Echo Cancellation, Speech Dereverberation and Source SeparationabstractThis paper presents a joint source separation algorithm that simultaneously reduces acoustic echo, reverberation and interfering sources.Target speeches are separated from the mixture by maximizing independence with respect to the other sources.It is shown that the separation process can be decomposed into cascading sub-processes that separately relate to acoustic echo cancellation, speech dereverberation and source separation, all of which are solved using the auxiliary function based independent component/vector analysis techniques, and their solving orders are exchangeable.The cascaded solution not only leads to lower computational complexity but also better separation performance than the vanilla joint algorithm. Yueyue Na, Zhang Liu 0006, Biao Tian 0002, Qiang Fu 0001 |
Interspeech | 5 |
| 2020 | A Visual-Pilot Deep Fusion for Target Speech Separation in Multitalker Noisy EnvironmentabstractSeparating the target speech in multi-talker noisy environment is a challenging problem for audio-only source separation algorithms. The major problem behind is that the separated speech from the same talker can switch among the outputs across consecutive segments, causing the talker permutation issue. In this paper, we deploy face tracking and propose the low-dimension hand-crafted visual features and the low-cost deep fusion architectures to separate the unseen but visible target sources in multi-talker noisy environment. It is shown that our approach is not only capable of addressing the talker permutation issue but also producing additional separation improvement in challenging mixtures such as the same-gender overlapping ones on the public dataset. We also show that the significant improvement of the target speech recognition is achieved on the simulated real-world dataset. Our training is independent of the number of visible sources providing flexibility in deployment. Zhang Liu 0006, Yueyue Na, Biao Tian 0002, Qiang Fu 0001 |
ICASSP | 6 |
| 2020 | A Semi-Blind Source Separation Approach for Speech Dereverberation
Yueyue Na, Zhang Liu 0006, Biao Tian 0002, Qiang Fu 0001 |
INTERSPEECH | 6 |
| 2016 | A Robust Dual-Microphone Speech Source Localization Algorithm for Reverberant EnvironmentsabstractSpeech source localization (SSL) using a microphone array \naims to estimate the direction-of-arrival (DOA) of the speech \nsource. However, its performance often degrades rapidly in reverberant \nenvironments. In this paper, a novel dual-microphone \nSSL algorithm is proposed to address this problem. First, the \ntime-frequency regions dominated by direct sound are extracted \nby tracking the envelopes of speech, reverberation and background \nnoise. The time-difference-of-arrival (TDOA) is then \nestimated by considering only these reliable regions. Second, \na bin-wise de-aliasing strategy is introduced to make better use \nof the DOA information carried at high frequencies, where the \nspatial resolution is higher and there is typically less corruption \nby diffuse noise. Our experiments show that when compared \nwith other widely-used algorithms, the proposed algorithm produces \nmore reliable performance in realistic reverberant environments. Yanmeng Guo, Xiaofei Wang 0007, Chao Wu 0011, Qiang Fu 0001, Ning Ma 0002, Guy J. Brown |
INTERSPEECH | 4 |
| 2016 | Adaptive Group Sparsity for Non-Negative Matrix Factorization with Application to Unsupervised Source Separation
Xiaofei Wang 0007, Qiang Fu 0001, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2016 | A DNN-HMM Approach to Non-Negative Matrix Factorization Based Speech Enhancement
Xiaofei Wang 0007, Qiang Fu 0001, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2016 | Cross Array and Rank-1 MUSIC Algorithm for Acoustic Highway Lane DetectionabstractA vehicle emits sound as it travels along the road, which can be used as a kind of robust feature for traffic monitoring. In this paper, an acoustic-based lane detection approach is introduced for a multilane traffic monitoring system. First, a microphone array is designed according to a typical Chinese highway configuration. The design is based on the cross-array structure, and the cross-correlation matrix from the two subarrays in the selected working frequency band is calculated for the subsequent traffic monitoring operations. Then, a cross section across the road is constructed by beamforming, in which the single-source assumption can be applied, and the passing vehicle azimuth is detected by the proposed rank-1 Multiple Signal Classification (MUSIC) algorithm. Finally, a Parzen-window-based technique is proposed to estimate the vehicle azimuth probability density function (pdf) from the individual azimuth observations. Lane centers and boundaries can be revealed from the peak and valley patterns of the estimated pdf. A prototype traffic monitoring system is developed, and several lane detection approaches are compared in both simulated and real-world environments in the developed system framework. The experimental results exhibit the efficiency of the proposed approach. Yueyue Na, Yanmeng Guo, Qiang Fu 0001, Yonghong Yan 0002 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2015 | A reverberation robust target speech detection method using dual-microphone in distant-talking scene
Xiaofei Wang 0007, Yanmeng Guo, Chao Wu 0011, Qiang Fu 0001, Yonghong Yan 0002 |
Speech Commun. | 4 |
| 2014 | A robust step-size control algorithm for frequency domain acoustic echo cancellationabstractThe presence of near-end interferences and echo path changes make it essential for an adaptive filter to vary its learning rate according to corresponding conditions. In this paper, a robust step-size control algorithm which is based on the optimization of the square of the bin-wise a posteriori error is proposed. To prevent the adaptive filter from diverging in the presence of interferences, constraints on the filter update are applied. The learning rate expression is derived and then we extend the method to multidelay block frequency domain adaptive filter (MDF) so as to meet the demand of low delay in practical application. An updating strategy for the constraints is proposed as well. Experiments are carried out to demonstrate the superiority of the proposed approach, especially in double-talk and echo path change situations. Index Terms: Acoustic echo cancellation, step-size control, robust filtering. Chao Wu 0011, Kaiyu Jiang, Yanmeng Guo, Qiang Fu 0001, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2014 | Acoustic Echo Control with Frequency-Domain Stage-Wise RegressionabstractThis letter introduces frequency domain stage-wise regression to acoustic echo control. By approximating the echo path as concatenate short segments in frequency domain, simple regression in the frequency domain can be carried out to estimate the echo contributed by consecutive far-end signal blocks stage by stage. A non-stationarity controlled smoothing factor is proposed alongside the regression procedure to mitigate the increasing variance of estimation when no significant echo but only near-end background noise is present. Experiments are carried out to demonstrate the superiority of the proposed approach, especially in unstable environment. Kaiyu Jiang, Chao Wu 0011, Yanmeng Guo, Qiang Fu 0001, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 4 |
| 2012 | A two-microphone based voice activity detection for distant-talking speech in wide range of direction of arrivalabstractIn this paper, a two-microphone based voice activity detection (VAD) algorithm is proposed to detect the distant-talking speech coming randomly from a wide range of direction of arrival (DOA). The long-term information of inter-channel phase difference (LTIPD) is introduced as a target speech existence measure, which describes the concentration degree of DOA estimations on a sound source with harmonic structure. The proposed algorithm performs robustly on distant-talking speech recorded in several real environments. Yanmeng Guo, Qiang Fu 0001, Yonghong Yan 0002 |
ICASSP | 3 |
| 2010 | Speech enhancement using improved generalized sidelobe canceller in frequency domain with multi-channel postfiltering
Qiang Fu 0001, Yonghong Yan 0002 |
INTERSPEECH | 2 |
| 2008 | Cochannel speech separation using multi-pitch estimation and model based voiced sequential groupingabstractIn this paper, a new cochannel speech separation algorithm us-ing multi-pitch extraction and speaker model based sequential grouping is proposed. After auditory segmentation based on on-set and offset analysis, robust multi-pitch estimation algorithm is performed on each segment and the corresponding voiced portions are segregated. Then speaker pair model based on support vector machine (SVM) is employed to determine the optimal sequential grouping alignments and group the speaker homogeneous segments into pure speaker streams. Systematic evaluation on the speech separation challenge database shows significant improvement over the baseline performance. Index Terms: Auditory scene analysis, cochannel speech, multi-pitch estimation, sequential grouping Ming Li 0026, Chuan Cao, Ping Lu 0009, Qiang Fu 0001, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2008 | A frequency domain approach for speech enhancement with directionality using compact microphone array
Qiang Fu 0001, Yonghong Yan 0002 |
INTERSPEECH | 2 |