Xiguang Zheng

dblp:89/10816 · DBLP profile ↗
← Back
25ranked-venue papers
7as first author
17since 2021 · last 2025
0000-0003-3103-9090ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 6 first-author · 17 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 A Metric for Predicting the Quality of Ambisonic Spatial Audio Reproduced Using Spatially Interpolated or Extrapolated Room Impulse Responses
abstract
In virtual reality (VR), sound sources are convolved with room impulse responses (RIRs) to create immersive and dynamic audio experiences. Assessing the quality of spatial audio synthesis in VR is challenging. Subjective listening tests are accurate, but they are time-consuming and costly. This paper introduces a novel objective quality metric to predict the listening quality (LQ) and localization accuracy (LA) of Ambisonic spatial audio at new positions using spatially interpolated or extrapolated first-order Ambisonic (FOA) RIRs based on known nearby FOA RIRs. The LQ and LA scores are computed based on the Kolmogorov-Smirnov test, to measure the similarity between segments of direct sound and reflections of reference and synthesized FOA RIRs. Results show that these scores strongly correlate with subjective test results, proving the reliability of the proposed method. A major advantage is that it predicts spatial audio quality directly from synthesized FOA RIRs, without requiring convolved signals or being affected by specific sound sources, providing a more practical and efficient solution for development.
Hualin Ren, Christian H. Ritz, Jiahong Zhao, Xiguang Zheng, Daeyoung Jang
ICASSP4
2024 BAE-Net: a Low Complexity and High Fidelity Bandwidth-Adaptive Neural Network for Speech Super-Resolution
abstract
Speech bandwidth extension (BWE) has demonstrated promising performance in enhancing the perceptual speech quality in real communication systems. Most existing BWE researches primarily focus on fixed upsampling ratios, disregarding the fact that the effective bandwidth of captured audio may fluctuate frequently due to various capturing devices and transmission conditions. In this paper, we propose a streaming adaptive bandwidth extension solution dubbed BAE-Net, which is suitable to handle the low-resolution speech with unknown and varying effective bandwidth. To address the challenges of recovering both the high-frequency magnitude and phase components of the speech content blindly, we devise a dual-stream architecture that incorporates the magnitude inpainting and phase refinement. For potential applications on edge devices, this paper also introduces BAE-NET-lite, which is a lightweight, streaming and efficient framework. Quantitative results demonstrate the superiority of BAE-Net in terms of performance and computational efficiency when compared with existing state-of-the-art BWE methods.
Guochen Yu, Xiguang Zheng, Runqiang Han, Chengshi Zheng
ICASSP2
2024 Uncertainty-Aware Mean Opinion Score Prediction
Hui Wang 0075, Shiwan Zhao, Jiaming Zhou 0001, Xiguang Zheng, Haoqin Sun, Xuechen Wang
INTERSPEECH4
2023 KAQ: A Non-Intrusive Stacking Framework for Mean Opinion Score Prediction with Multi-Task Learning
abstract
Non-intrusive speech quality assessment aims to replace the time-consuming subjective evaluation metric, i.e., mean opinion score (MOS), by designing a trainable model to predict the MOS. In this work, we propose a non-intrusive stacking framework, named KAQ, to automatically achieve MOS prediction. KAQ benefits from the stacking algorithm to harness the capabilities of multiple well-performing machine learning models. And these models are trained to predict not only MOS but also intelligibility score via a multi-task learning. The experimental results show that KAQ achieves significantly better performance than the baselines in the 2023 VoiceMOS challenge and also wins second place in terms of mean square error at the utterance and system levels.
Chenglin Xu, Xiguang Zheng
ASRU2
2023 A Low-Latency Deep Hierarchical Fusion Network for Fullband Acoustic Echo Cancellation
abstract
This paper describes our submission to the fourth Acoustic Echo Cancellation (AEC) Challenge, which is part of ICASSP 2023 Signal Processing Grand Challenge. The proposed system is developed based on our earlier system submitted to the ICASSP 2022 AEC challenge with significant latency and network structure improvement, while achieving better subjective results.
Runqiang Han, Xiguang Zheng
ICASSP4
2023 Intermediate-Task Learning with Pretrained Model for Synthesized Speech MOS Prediction
abstract
Mean Opinion Scores (MOS) prediction has been a crucial task for speech quality assessment. Recently, methods based on pre-training and fine-tuning achieved the start-of-the-art performance for MOS prediction. While these two-stage methods can improve the generalization ability of the model, there still exists a large gap between the pre-training and the downstream MOS prediction objectives in the input distribution and the output label space. In this paper, we proposed a three-stage intermediate task training scheme, then we tailored two possible intermediate tasks (classification and contrastive learning) for the speech MOS prediction task. The experimental results show that the proposed models achieve the SOTA results compared with the existing two-stage systems in most of the metrics for in-domain and out-of-domain datasets with less computational complexity.
Hui Wang 0075, Xiguang Zheng
ICME2
2023 RAMP: Retrieval-Augmented MOS Prediction via Confidence-based Dynamic Weighting
abstract
Automatic Mean Opinion Score (MOS) prediction is crucial to evaluate the perceptual quality of the synthetic speech. While recent approaches using pre-trained self-supervised learning (SSL) models have shown promising results, they only partly address the data scarcity issue for the feature extractor. This leaves the data scarcity issue for the decoder unresolved and leading to suboptimal performance. To address this challenge, we propose a retrieval-augmented MOS prediction method, dubbed {\bf RAMP}, to enhance the decoder's ability against the data scarcity issue. A fusing network is also proposed to dynamically adjust the retrieval scope for each instance and the fusion weights based on the predictive confidence. Experimental results show that our proposed method outperforms the existing methods in multiple scenarios.
Hui Wang 0075, Shiwan Zhao, Xiguang Zheng
INTERSPEECH3
2022 Multi-Stage and Multi-Loss Training for Fullband Non-Personalized and Personalized Speech Enhancement
abstract
Deep learning-based wideband (16kHz) speech enhancement approaches have surpassed traditional methods. This work further extends the existing wideband systems to enable full-band (48kHz) speech enhancement while simultaneously ensuring automatic speech recognition compatibility and optionally, personalized speech enhancement. As shown in the evaluation results, this is achieved by employing a multi-stage and multi-loss training architecture that incorporates the recently proposed two-step structure, ASR loss produced by a back-end ASR encoder, and the speaker extraction network.
Lianwu Chen, Chenglin Xu, Xinlei Ren, Xiguang Zheng
ICASSP5
2022 L3DAS22 Challenge: Learning 3D Audio Sources in a Real Office Environment
abstract
The L3DAS22 Challenge is aimed at encouraging the development of machine learning strategies for 3D speech enhancement and 3D sound localization and detection in office-like environments. This challenge improves and extends the tasks of the L3DAS21 edition1. We generated a new dataset, which maintains the same general characteristics of L3DAS21 datasets, but with an extended number of data points and adding constrains that improve the baseline model’s efficiency and overcome the major difficulties encountered by the participants of the previous challenge. We updated the baseline model of Task 1, using the architecture that ranked first in the previous challenge edition. We wrote a new supporting API, improving its clarity and ease-of-use. In the end, we present and discuss the results submitted by all participants. L3DAS22 Challenge website: www.l3das.com/icassp2022.
Eric Guizzo, Christian Marinoni, Marco Pennese, Xinlei Ren, Xiguang Zheng, Bruno S. Masiero, Aurelio Uncini, Danilo Comminiello
ICASSP5
2022 A Two-Step Backward Compatible Fullband Speech Enhancement System
abstract
Speech enhancement methods based on deep learning have surpassed traditional methods. While many of these new approaches are operating on the wideband (16kHz) sample rate, a new fullband (48kHz) speech enhancement system is proposed in this paper. Compared to the existing full-band systems that utilize perceptually motivated features to train the fullband speech enhancement with a single network structure, the proposed system is a two-step system ensuring good fullband speech enhancement quality while backward compatible to the existing wideband systems.
Lianwu Chen, Xiguang Zheng, Xinlei Ren
ICASSP3
2022 A Deep Hierarchical Fusion Network for Fullband Acoustic Echo Cancellation
abstract
Deep learning based wideband (16kHz) acoustic echo cancellation (AEC) approaches have surpassed traditional methods. This work proposes a deep hierarchical fusion (DHF) network with intra-network and inter-network fusion to further improve the wideband AEC performance. Meanwhile, this work extends the existing wideband systems to enable fullband (48kHz) AEC while simultaneously ensuring automatic speech recognition compatibility by incorporating with an ASR loss. The proposed system has ranked 2nd place in ICASSP 2022’s AEC Challenge.
Runqiang Han, Lianwu Chen, Xiguang Zheng
ICASSP5
2022 Multi-Scale Temporal-Frequency Attention for Music Source Separation
abstract
In recent years, deep neural networks (DNNs) based approaches have achieved the start-of-the-art performance for music source separation (MSS). Although previous methods have addressed the large receptive field modeling using various methods, the temporal and frequency correlations of the music spectrogram with repeated patterns have not been explicitly explored for the MSS task. In this paper, a temporal-frequency attention module is proposed to model the spectrogram correlations along both temporal and frequency dimensions. Moreover, a multi-scale attention is proposed to effectively capture the correlations for music signal. The experimental results on MUSDB18 dataset show that the proposed method outperforms the existing state-of-the-art systems with 9.51 dB signal-to-distortion ratio (SDR) on separating the vocal stems, which is the primary practical application of MSS.
Lianwu Chen, Xiguang Zheng
ICME2
2022 Impairment Representation Learning for Speech Quality Assessment
Lianwu Chen, Xinlei Ren, Xiguang Zheng
INTERSPEECH4
2022 End-to-End Multi-Loss Training for Low Delay Packet Loss Concealment
Xiguang Zheng
INTERSPEECH2
2021 A Causal U-Net Based Neural Beamforming Network for Real-Time Multi-Channel Speech Enhancement
Xinlei Ren, Lianwu Chen, Xiguang Zheng
Interspeech4
2021 Low-Delay Speech Enhancement Using Perceptually Motivated Target and Loss
Xinlei Ren, Xiguang Zheng, Lianwu Chen
Interspeech3
2021 Towards Blind Audio Quality Assessment using a Convolutional-Recurrent Neural Network
abstract
Blind estimation of audio quality is desired for practical applications since the original reference audio signal is sometimes unavailable. The subjective quality degradation of the audio signals can be caused by the low bitrate compression and multiple compression during the content submission and distribution stages. Existing methods have been proposed to classify the audio signals by the encoding bitrates with the informed audio codec name as well as the estimated MDCT framing grid and window type sequences during the encoding stage. In this work, a convolutional-recurrent neural network is proposed to perform blind AAC bitrates classification. Compared to the existing methods, the proposed method can perform the AAC bitrate classification directly from the MDCT coefficients without any prior knowledge of the encoding framing grid and window type sequences. The proposed method is further extended to perform multi-codec bitrate-related perceptual audio quality classification, which has not been extensively studied in the existing literature. For the AAC bitrate classification task, the evaluation results show the proposed method can achieve similar accuracy for most of the bitrates without using any framing grid information compared to the existing methods. For the multi-codec bitrate-related perceptual audio quality classification task, the proposed method can achieve 95% accuracy to classify three perceptual classes for the double compressed unseen data using three common audio codecs.
Xiguang Zheng
QoMEX1
2016 Encoding and communicating navigable speech soundfields
Xiguang Zheng, Christian H. Ritz, Jiangtao Xi
Multim. Tools Appl.1
2015 Encoding Multiple Audio Objects Using Intra-Object Sparsity
abstract
Preserving audio scenes in the form of audio objects has become common in recent years. Object-based audio techniques provide more flexibility for personalized rendering as well as a more accurate audio object trajectory. For encoding and transmitting multiple audio objects in a lossy manner, a new compression framework for multiple simultaneously occurring audio objects is presented in this work. The proposed encoding approach is based on the intra-object sparsity (approximate k-sparsity). After establishing a quantitative measure of approximate k-sparsity, statistical analysis is employed to validate the proposed intra-object sparsity of audio objects. By exploring this intra-object sparsity, multiple simultaneously occurring audio objects are compressed into a mono downmix signal with side information. This downmix signal can be further compressed by legacy audio codecs. Meanwhile, the side information is transmitted in a lossless manner. The objective and subjective evaluations revealed that the proposed compression framework achieved better perceptual quality compared to an existing technique where up to eight audio objects are considered. The subjective evaluations also confirmed that the proposed approach is able to achieve scalable transmission according to the bandwidth while preserving the perceptual quality of both the individual audio objects and the spatial audio scenes.
Mao-shen Jia, Changchun Bao, Xiguang Zheng, Christian H. Ritz
IEEE ACM Trans. Audio Speech Lang. Process.4
2013 A psychoacoustic-based analysis-by-synthesis scheme for jointly encoding multiple audio objects into independent mixtures
abstract
Perceptually accurate representation of audio objects obtained from multi-track audio signals is desired for applications such as interactive soundfield rendering and browsing. Presented in this work is a scalable psychoacoustic analysis-by-synthesis approach to extract the perceptually dominant time-frequency audio objects from a multi-track audio signal. The proposed compression framework exploits sparsity in the perceptual time-frequency domain where up to eight audio objects can be efficiently encoded using only two audio mixtures with side information representing the origin of the time-frequency instances in the mixture signals. The proposed approach, judged by both objective and subjective tests, results in superior audio quality compared to existing techniques when encoding more than 5 audio objects.
Xiguang Zheng, Christian H. Ritz, Jiangtao Xi
ICASSP1
2013 Collaborative Blind Source Separation Using Location Informed Spatial Microphones
abstract
This letter presents a new Collaborative Blind Source Separation (CBSS) technique that uses a pair of location informed coincident microphone arrays to jointly separate simultaneous speech sources based on time-frequency source localization estimates from each microphone recording. While existing BSS approaches are based on localization estimates of sparse time-frequency components, the proposed approach can also recover non-sparse (overlapping) time-frequency components. The proposed method has been evaluated using up to three simultaneous speech sources under both anechoic and reverberant conditions. Results from objective and subjective measures of the perceptual quality of the separated speech show that the proposed approach significantly outperforms existing BSS approaches.
Xiguang Zheng, Christian H. Ritz, Jiangtao Xi
IEEE Signal Process. Lett.1
2013 A General Compression Approach to Multi-Channel Three-Dimensional Audio
abstract
This paper presents a technique for low bit rate compression of three-dimensional (3D) audio produced by multiple loudspeaker channels. The approach is based on the time-frequency analysis of the localization of spatial sound sources within the 3D space as rendered by a multi-channel audio signal (in this case 16 channels). This analysis results in the derivation of a stereo downmix signal representing the original 16 channels. Alternatively, a mono-downmix signal with side information representing the location of sound sources within the 3D spatial scene can also be derived. The resulting downmix signals are then compressed with a traditional audio coder, resulting in a representation of the 3D soundfield at bit rates comparable with existing stereo audio coders while maintaining the perceptual quality produced from separate encoding of each channel.
Christian H. Ritz, Ian S. Burnett, Xiguang Zheng
IEEE Trans. Speech Audio Process.4
2013 Encoding Navigable Speech Sources: A Psychoacoustic-Based Analysis-by-Synthesis Approach
abstract
This paper presents a psychoacoustic-based analysis-by-synthesis approach for compressing navigable speech sources. The approach targets multi-party teleconferencing applications, where selective reproduction of individual speech sources is desired. Based on exploiting sparsity of speech in the perceptual time-frequency domain, multiple speech signals are encoded into one mono mixture signal, which can be further compressed using a standard speech codec. Using side information indicating the active speech source for each time frequency instant enables flexible decoding and reproduction. Objective results highlight the importance of considering perception when exploiting the sparse nature of speech in the time-frequency domain. Results show that this sparsity, as measured by the preserved energy level of perceptually important time-frequency components extracted from mixtures of speech signals, is similar in both anechoic and reverberant environments. The proposed approach is applied to a series of simulated and real reverberant speech recordings, where the resulting speech mixtures are compressed using a standard speech codec operating at 32 kbps. The perceptual quality, as judged both by objective and subjective evaluations, outperforms a simple sparsity approach that does not consider perception as well as the approach that encodes each source separately. While the perceptual quality of individual speech sources is maintained, subjective tests also confirm the approach maintains the perceptual quality of the spatialized speech scene.
Xiguang Zheng, Christian H. Ritz, Jiangtao Xi
IEEE Trans. Speech Audio Process.1
2012 Encoding navigable speech sources: An analysis by synthesis approach
abstract
This paper pressents an analysis-by-synthesis coding architecture for compressing navigable speech sources. The proposed coding scheme encodes multiple overlapped speech sources recorded, for example, during a multi-participant meeting or teleconference, into a mono or stereo mixture signal that can be compressed with an existing speech coder. The individual speech sources can be separated from the received compressed mixture, which allows the listener to determine the active sources and their spatial locations at the reproduction site. The approach was applied to the compression of a series of speech soundfields created from multiple clean speech sentences and real meeting recordings, where each sound-field contained four participants with up to three simultaneous speech sources. At a total bit rate of 48 kbps, the perceptual quality of each decoded speech source, as judged by subjective listening tests, was found to be significantly better than either a non-a-by-s approach or separate encoding of each source at the same overall total bit rate. Subjective listening tests also confirm that the quality of the spatialised speech scene is maintained as well.
Xiguang Zheng, Christian H. Ritz, Jiangtao Xi
ICASSP1
2011 Compression of navigable speech soundfield zones
abstract
This paper presents a new coding architecture for the compression of navigable speech soundfield zones. The proposed coding scheme encodes multiple speech soundfields, each representing different spatial zones, into a mono or stereo sound-field mixture signal that can be compressed with an existing speech or audio coder. The resulting compressed signals can be decoded back to individual soundfield zones. Objective and subjective testing results show that the approach successfully compresses up to 3 speech soundfields (each consisting of 4 individual speakers) at a bit rate of 48 kbps whilst maintaining the perceptual quality of each decoded soundfield zone.
Xiguang Zheng, Christian H. Ritz
MMSP1