Xinlei Ren

dblp:252/4895 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2022 Multi-Stage and Multi-Loss Training for Fullband Non-Personalized and Personalized Speech Enhancement
abstract
Deep learning-based wideband (16kHz) speech enhancement approaches have surpassed traditional methods. This work further extends the existing wideband systems to enable full-band (48kHz) speech enhancement while simultaneously ensuring automatic speech recognition compatibility and optionally, personalized speech enhancement. As shown in the evaluation results, this is achieved by employing a multi-stage and multi-loss training architecture that incorporates the recently proposed two-step structure, ASR loss produced by a back-end ASR encoder, and the speaker extraction network.
Lianwu Chen, Chenglin Xu, Xinlei Ren, Xiguang Zheng
ICASSP4
2022 L3DAS22 Challenge: Learning 3D Audio Sources in a Real Office Environment
abstract
The L3DAS22 Challenge is aimed at encouraging the development of machine learning strategies for 3D speech enhancement and 3D sound localization and detection in office-like environments. This challenge improves and extends the tasks of the L3DAS21 edition1. We generated a new dataset, which maintains the same general characteristics of L3DAS21 datasets, but with an extended number of data points and adding constrains that improve the baseline model’s efficiency and overcome the major difficulties encountered by the participants of the previous challenge. We updated the baseline model of Task 1, using the architecture that ranked first in the previous challenge edition. We wrote a new supporting API, improving its clarity and ease-of-use. In the end, we present and discuss the results submitted by all participants. L3DAS22 Challenge website: www.l3das.com/icassp2022.
Eric Guizzo, Christian Marinoni, Marco Pennese, Xinlei Ren, Xiguang Zheng, Bruno S. Masiero, Aurelio Uncini, Danilo Comminiello
ICASSP4
2022 A Two-Step Backward Compatible Fullband Speech Enhancement System
abstract
Speech enhancement methods based on deep learning have surpassed traditional methods. While many of these new approaches are operating on the wideband (16kHz) sample rate, a new fullband (48kHz) speech enhancement system is proposed in this paper. Compared to the existing full-band systems that utilize perceptually motivated features to train the fullband speech enhancement with a single network structure, the proposed system is a two-step system ensuring good fullband speech enhancement quality while backward compatible to the existing wideband systems.
Lianwu Chen, Xiguang Zheng, Xinlei Ren
ICASSP4
2022 Impairment Representation Learning for Speech Quality Assessment
Lianwu Chen, Xinlei Ren, Xiguang Zheng
INTERSPEECH2
2021 A Causal U-Net Based Neural Beamforming Network for Real-Time Multi-Channel Speech Enhancement
Xinlei Ren, Lianwu Chen, Xiguang Zheng
Interspeech1
2021 Low-Delay Speech Enhancement Using Perceptually Motivated Target and Loss
Xinlei Ren, Xiguang Zheng, Lianwu Chen
Interspeech2