Yueyue Na

dblp:118/4924 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
5since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2022 Joint Ego-Noise Suppression and Keyword Spotting on Sweeping Robots
abstract
Keyword spotting is necessary for triggering human-machine speech interaction. It is a challenging task especially in low signal-to-noise ratio and moving scenarios, such as on a sweeping robot with strong ego-noise. This paper proposes a novel approach for joint ego-noise suppression and keyword detection. The keyword detection model accepts outputs from multi-look adaptive beamformers. The noise covariance matrix in the beamformer is in turn updated using the keyword absence probability given by the model, forming an end-to-end loop-back. The keyword model also adopts a multi-channel feature fusion using self-attention, and a hidden Markov model for online decoding. The performance of the proposed approach is verified on real-word datasets recorded on a sweeping robot.
Yueyue Na, Liang Wang 0018
ICASSP1
2022 NN3A: Neural Network Supported Acoustic Echo Cancellation, Noise Suppression and Automatic Gain Control for Real-Time Communications
abstract
Acoustic echo cancellation (AEC), noise suppression (NS) and automatic gain control (AGC) are three often required modules for real-time communications (RTC). This paper proposes a neural network supported algorithm for RTC, namely NN3A, which incorporates an adaptive filter and a multi-task model for residual echo suppression, noise reduction and near-end speech activity detection. The proposed algorithm is shown to outperform both a method using separate models and an end-to-end alternative. It is further shown that there exists a trade-off in the model between residual suppression and near-end speech distortion, which could be balanced by a novel loss weighting function. Several practical aspects of training the joint model are also investigated to push its performance to limit.
Yueyue Na, Biao Tian 0002, Qiang Fu 0001
ICASSP2
2022 Personalized Acoustic Echo Cancellation for Full-duplex Communications
abstract
Deep neural networks (DNNs) have shown promising results for acoustic echo cancellation (AEC). But the DNN-based AEC models let through all near-end speakers including the interfering speech. In light of recent studies on personalized speech enhancement, we investigate the feasibility of personalized acoustic echo cancellation (PAEC) in this paper for full-duplex communications, where background noise and interfering speakers may coexist with acoustic echoes. Specifically, we first propose a novel backbone neural network termed as gated temporal convolutional neural network (GTCNN) that outperforms state-of-the-art AEC models in performance. Speaker embeddings like d-vectors are further adopted as auxiliary information to guide the GTCNN to focus on the target speaker. A special case in PAEC is that speech snippets of both parties on the call are enrolled. Experimental results show that auxiliary information from either the near-end speaker or the far-end speaker can improve the DNN-based AEC performance. Nevertheless, there is still much room for improvement in the utilization of the finite-dimensional speaker embeddings.
Yukai Jv, Yihui Fu, Yueyue Na, Lei Xie 0001
INTERSPEECH5
2021 Weighted Recursive Least Square Filter and Neural Network Based Residual ECHO Suppression for the AEC-Challenge
abstract
This paper presents a real-time Acoustic Echo Cancellation (AEC) algorithm submitted to the AEC-Challenge. The algorithm consists of three modules: Generalized Cross-Correlation with PHAse Transform (GCC-PHAT) based time delay compensation, weighted Recursive Least Square (wRLS) based linear adaptive filtering and neural network based residual echo suppression. The wRLS filter is derived from a novel semi-blind source separation perspective. The neural network model predicts a Phase-Sensitive Mask (PSM) based on the aligned reference and the linear filter output. The algorithm achieved a mean subjective score of 4.00 and ranked 2nd in the AEC-Challenge.
Yueyue Na, Zhang Liu 0006, Biao Tian 0002, Qiang Fu 0001
ICASSP2
2021 Joint Online Multichannel Acoustic Echo Cancellation, Speech Dereverberation and Source Separation
abstract
This paper presents a joint source separation algorithm that simultaneously reduces acoustic echo, reverberation and interfering sources.Target speeches are separated from the mixture by maximizing independence with respect to the other sources.It is shown that the separation process can be decomposed into cascading sub-processes that separately relate to acoustic echo cancellation, speech dereverberation and source separation, all of which are solved using the auxiliary function based independent component/vector analysis techniques, and their solving orders are exchangeable.The cascaded solution not only leads to lower computational complexity but also better separation performance than the vanilla joint algorithm.
Yueyue Na, Zhang Liu 0006, Biao Tian 0002, Qiang Fu 0001
Interspeech1
2020 A Visual-Pilot Deep Fusion for Target Speech Separation in Multitalker Noisy Environment
abstract
Separating the target speech in multi-talker noisy environment is a challenging problem for audio-only source separation algorithms. The major problem behind is that the separated speech from the same talker can switch among the outputs across consecutive segments, causing the talker permutation issue. In this paper, we deploy face tracking and propose the low-dimension hand-crafted visual features and the low-cost deep fusion architectures to separate the unseen but visible target sources in multi-talker noisy environment. It is shown that our approach is not only capable of addressing the talker permutation issue but also producing additional separation improvement in challenging mixtures such as the same-gender overlapping ones on the public dataset. We also show that the significant improvement of the target speech recognition is achieved on the simulated real-world dataset. Our training is independent of the number of visible sources providing flexibility in deployment.
Zhang Liu 0006, Yueyue Na, Biao Tian 0002, Qiang Fu 0001
ICASSP3
2020 A Semi-Blind Source Separation Approach for Speech Dereverberation
Yueyue Na, Zhang Liu 0006, Biao Tian 0002, Qiang Fu 0001
INTERSPEECH2
2016 Cross Array and Rank-1 MUSIC Algorithm for Acoustic Highway Lane Detection
abstract
A vehicle emits sound as it travels along the road, which can be used as a kind of robust feature for traffic monitoring. In this paper, an acoustic-based lane detection approach is introduced for a multilane traffic monitoring system. First, a microphone array is designed according to a typical Chinese highway configuration. The design is based on the cross-array structure, and the cross-correlation matrix from the two subarrays in the selected working frequency band is calculated for the subsequent traffic monitoring operations. Then, a cross section across the road is constructed by beamforming, in which the single-source assumption can be applied, and the passing vehicle azimuth is detected by the proposed rank-1 Multiple Signal Classification (MUSIC) algorithm. Finally, a Parzen-window-based technique is proposed to estimate the vehicle azimuth probability density function (pdf) from the individual azimuth observations. Lane centers and boundaries can be revealed from the peak and valley patterns of the estimated pdf. A prototype traffic monitoring system is developed, and several lane detection approaches are compared in both simulated and real-world environments in the developed system framework. The experimental results exhibit the efficiency of the proposed approach.
Yueyue Na, Yanmeng Guo, Qiang Fu 0001, Yonghong Yan 0002
IEEE Trans. Intell. Transp. Syst.1
2012 Kernel and spectral methods for solving the permutation problem in frequency domain BSS
abstract
In frequency domain blind source separation (FDBSS), separated frequency bin data in the same source must be grouped together before outputting the final result, which is the well-known permutation problem. Clustering techniques are broadly used in solving the permutation problem, however, some challenges still exist, for example, elongated datasets should be handled, and constraint from the background knowledge must be considered. Inspired by various successful applications of kernel and spectral clustering methods in machine learning and data mining community, we try to solve the permutation problem by these methods. In this paper, the weighted kernel k-means algorithm is modified according to the specific requirement of the permutation problem, and the spectral interpretation of the kernel approach is also investigated. In addition, we propose several kernel construction approaches to improving the permutation performance. Different experiments are carried out on a uniform platform, and show better performance of the proposed approach.
Yueyue Na, Jian Yu 0001
IJCNN1