Fei Jiang 0007

dblp:80/4567-7 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
5since 2021 · last 2023
0000-0002-7685-0242ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2023 Low-Complexity Acoustic Echo Cancellation with Neural Kalman Filtering
abstract
The Kalman filter has been adopted in acoustic echo cancellation due to its robustness to double-talk, fast convergence, and good steady-state performance. The performance of Kalman filter is closely related to the estimation accuracy of the state noise covariance and the observation noise covariance. The estimation error may lead to unacceptable results, especially when the echo path suffers abrupt changes, the tracking performance of the Kalman filter could be degraded significantly. In this paper, we propose the neural Kalman filtering (NKF), which uses neural networks to implicitly model the covariance of the state noise and observation noise and to output the Kalman gain in real-time. Experimental results on both synthetic test sets and real-recorded test sets show that, the proposed NKF has superior convergence and re-convergence performance while ensuring low near-end speech degradation compared with the state-of-the-art model-based methods. Moreover, the model size of the proposed NKF is merely 5.3 K and the RTF is as low as 0.09, which indicates that it can be deployed in low-resource platforms.
Fei Jiang 0007, Xuefei Fang, Muyong Cao
ICASSP2
2022 Music Source Separation With Generative Flow
abstract
Fully-supervised models for source separation are trained on parallel mixture-source data and are currently state-of-the-art. However, such parallel data is often difficult to obtain, and it is cumbersome to adapt trained models to mixtures with new sources. Source-only supervised models, in contrast, only require individual source data for training. In this paper, we first leverage flow-based generators to train individual music source priors and then use these models, along with likelihood-based objectives, to separate music mixtures. We show that in singing voice separation and music separation tasks, our proposed method is competitive with a fully-supervised approach. We also demonstrate that we can flexibly add new types of sources, whereas fully-supervised approaches would require retraining of the entire model.
Jordan Darefsky, Fei Jiang 0007, Anton Selitskiy, Zhiyao Duan
IEEE Signal Process. Lett.3
2021 An Empirical Study on Channel Effects for Synthetic Voice Spoofing Countermeasure Systems
abstract
Spoofing countermeasure (CM) systems are critical in speaker verification; they aim to discern spoofing attacks from bona fide speech trials.In practice, however, acoustic condition variability in speech utterances may significantly degrade the performance of CM systems.In this paper, we conduct a cross-dataset study on several state-of-the-art CM systems and observe significant performance degradation compared with their singledataset performance.Observing differences of average magnitude spectra of bona fide utterances across the datasets, we hypothesize that channel mismatch among these datasets is one important reason.We then verify it by demonstrating a similar degradation of CM systems trained on original but evaluated on channel-shifted data.Finally, we propose several channel robust strategies (data augmentation, multi-task learning, adversarial learning) for CM systems, and observe a significant performance improvement on cross-dataset experiments.
You Zhang 0001, Fei Jiang 0007, Zhiyao Duan
Interspeech3
2021 Y-Vector: Multiscale Waveform Encoder for Speaker Embedding
abstract
State-of-the-art text-independent speaker verification systems typically use cepstral features or filter bank energies as speech features.Recent studies attempted to extract speaker embeddings directly from raw waveforms and have shown competitive results.In this paper, we propose a novel multi-scale waveform encoder that uses three convolution branches with different time scales to compute speech features from the waveform.These features are then processed by squeeze-and-excitation blocks, a multi-level feature aggregator, and a time delayed neural network (TDNN) to compute speaker embedding.We show that the proposed embeddings outperforms existing raw-waveformbased speaker embeddings on speaker verification by a large margin.A further analysis of the learned filters shows that the multi-scale encoder attends to different frequency bands at its different scales while resulting in a more flat overall frequency response than any of the single-scale counterparts.
Fei Jiang 0007, Zhiyao Duan
Interspeech2
2021 One-Class Learning Towards Synthetic Voice Spoofing Detection
abstract
Human voices can be used to authenticate the identity of the speaker, but the automatic speaker verification (ASV) systems are vulnerable to voice spoofing attacks, such as impersonation, replay, text-to-speech, and voice conversion. Recently, researchers developed anti-spoofing techniques to improve the reliability of ASV systems against spoofing attacks. However, most methods encounter difficulties in detecting unknown attacks in practical use, which often have different statistical distributions from known attacks. Especially, the fast development of synthetic voice spoofing algorithms is generating increasingly powerful attacks, putting the ASV systems at risk of unseen attacks. In this work, we propose an anti-spoofing system to detect unknown synthetic voice spoofing attacks (i.e., text-to-speech or voice conversion) using one-class learning. The key idea is to compact the bona fide speech representation and inject an angular margin to separate the spoofing attacks in the embedding space. Without resorting to any data augmentation methods, our proposed system achieves an equal error rate (EER) of 2.19% on the evaluation set of ASVspoof 2019 Challenge logical access scenario, outperforming all existing single systems (i.e., those without model ensemble).
You Zhang 0001, Fei Jiang 0007, Zhiyao Duan
IEEE Signal Process. Lett.2
2020 Speaker Attractor Network: Generalizing Speech Separation to Unseen Numbers of Sources
abstract
Most existing speech separation research focuses on improving the separation performance under consistent source number conditions between training and testing. In real-world applications, however, the source number may be different from that in training sets. In this letter, we address this problem by thoroughly improving the deep attractor network in terms of the network architecture and learning objectives so that it can well generalize to separating an unseen number of sources. Experimental results show that, compared with existing models, the proposed method significantly improves the separation performance when generalizing to an unseen number of speakers, and can separate up to five speakers even the model is only trained on two-speaker mixtures.
Fei Jiang 0007, Zhiyao Duan
IEEE Signal Process. Lett.1