Biao Tian 0002

dblp:47/3845-2 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
5since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021
YearPublicationVenuePosition
2023 Small Footprint Multi-channel Network for Keyword Spotting with Centroid Based Awareness
Dianwen Ng, Yang Xiao 0019, Jia Qi Yip, Biao Tian 0002, Qiang Fu 0001, Chng Eng Siong, Bin Ma 0001
INTERSPEECH5
2022 Convmixer: Feature Interactive Convolution with Curriculum Learning for Small Footprint and Noisy Far-Field Keyword Spotting
abstract
Building efficient architecture in neural speech processing is paramount to success in keyword spotting deployment. However, it is very challenging for lightweight models to achieve noise robustness with concise neural operations. In a real-world application, the user environment is typically noisy and may contain reverberations. We proposed a novel feature interactive convolutional model with merely 100K parameters to tackle this under the noisy far-field condition. The interactive unit is proposed in place of the attention module that promotes the flow of information with more efficient computations. Moreover, curriculum-based multi-condition training is adopted to attain better noise robustness. Our model achieves 98.2% top-1 accuracy on Google Speech Command V2-12 and is competitive against large transformer models under the designed noise condition.
Dianwen Ng, Yunqi Chen, Biao Tian 0002, Qiang Fu 0001, Chng Eng Siong
ICASSP3
2022 NN3A: Neural Network Supported Acoustic Echo Cancellation, Noise Suppression and Automatic Gain Control for Real-Time Communications
abstract
Acoustic echo cancellation (AEC), noise suppression (NS) and automatic gain control (AGC) are three often required modules for real-time communications (RTC). This paper proposes a neural network supported algorithm for RTC, namely NN3A, which incorporates an adaptive filter and a multi-task model for residual echo suppression, noise reduction and near-end speech activity detection. The proposed algorithm is shown to outperform both a method using separate models and an end-to-end alternative. It is further shown that there exists a trade-off in the model between residual suppression and near-end speech distortion, which could be balanced by a novel loss weighting function. Several practical aspects of training the joint model are also investigated to push its performance to limit.
Yueyue Na, Biao Tian 0002, Qiang Fu 0001
ICASSP3
2021 Weighted Recursive Least Square Filter and Neural Network Based Residual ECHO Suppression for the AEC-Challenge
abstract
This paper presents a real-time Acoustic Echo Cancellation (AEC) algorithm submitted to the AEC-Challenge. The algorithm consists of three modules: Generalized Cross-Correlation with PHAse Transform (GCC-PHAT) based time delay compensation, weighted Recursive Least Square (wRLS) based linear adaptive filtering and neural network based residual echo suppression. The wRLS filter is derived from a novel semi-blind source separation perspective. The neural network model predicts a Phase-Sensitive Mask (PSM) based on the aligned reference and the linear filter output. The algorithm achieved a mean subjective score of 4.00 and ranked 2nd in the AEC-Challenge.
Yueyue Na, Zhang Liu 0006, Biao Tian 0002, Qiang Fu 0001
ICASSP4
2021 Joint Online Multichannel Acoustic Echo Cancellation, Speech Dereverberation and Source Separation
abstract
This paper presents a joint source separation algorithm that simultaneously reduces acoustic echo, reverberation and interfering sources.Target speeches are separated from the mixture by maximizing independence with respect to the other sources.It is shown that the separation process can be decomposed into cascading sub-processes that separately relate to acoustic echo cancellation, speech dereverberation and source separation, all of which are solved using the auxiliary function based independent component/vector analysis techniques, and their solving orders are exchangeable.The cascaded solution not only leads to lower computational complexity but also better separation performance than the vanilla joint algorithm.
Yueyue Na, Zhang Liu 0006, Biao Tian 0002, Qiang Fu 0001
Interspeech4
2020 A Visual-Pilot Deep Fusion for Target Speech Separation in Multitalker Noisy Environment
abstract
Separating the target speech in multi-talker noisy environment is a challenging problem for audio-only source separation algorithms. The major problem behind is that the separated speech from the same talker can switch among the outputs across consecutive segments, causing the talker permutation issue. In this paper, we deploy face tracking and propose the low-dimension hand-crafted visual features and the low-cost deep fusion architectures to separate the unseen but visible target sources in multi-talker noisy environment. It is shown that our approach is not only capable of addressing the talker permutation issue but also producing additional separation improvement in challenging mixtures such as the same-gender overlapping ones on the public dataset. We also show that the significant improvement of the target speech recognition is achieved on the simulated real-world dataset. Our training is independent of the number of visible sources providing flexibility in deployment.
Zhang Liu 0006, Yueyue Na, Biao Tian 0002, Qiang Fu 0001
ICASSP5
2020 A Semi-Blind Source Separation Approach for Speech Dereverberation
Yueyue Na, Zhang Liu 0006, Biao Tian 0002, Qiang Fu 0001
INTERSPEECH5