Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zhongweiyang Xu

dblp:324/2576 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2025
0009-0009-3477-0170ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Audio and music processing · 100%
Artificial intelligence
1 paper
Generative modeling · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing › source separation
speech separation
0.912025
ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion Prior · ICML 2025
Machine learning › Generative modeling
diffusion model
0.312025
ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion Prior · ICML 2025
Machine learning › Generative modeling › diffusion model › diffusion sampling
diffusion posterior sampling
0.312025
ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion Prior · ICML 2025

Methods — techniques the papers use, named apart from their topics

relative transfer function estimation · 1.7optimization · 1.7diffusion posterior sampling · 1.7
YearPublicationVenuePosition
2025 ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion Prior
abstract
Blind Speech Separation (BSS) aims to separate multiple speech sources from audio mixtures recorded by a microphone array. The problem is challenging because it is a blind inverse problem, i.e., the microphone array geometry, the room impulse response (RIR), and the speech sources, are all unknown. We propose ArrayDPS to solve the BSS problem in an unsupervised, array-agnostic, and generative manner. The core idea builds on diffusion posterior sampling (DPS), but unlike DPS where the likelihood is tractable, ArrayDPS must approximate the likelihood by formulating a separate optimization problem. The solution to the optimization approximates room acoustics and the relative transfer functions between microphones. These approximations, along with the diffusion priors, iterate through the ArrayDPS sampling process and ultimately yield separated voice sources. We only need a simple single-speaker speech diffusion model as a prior, along with the mixtures recorded at the microphones; no microphone array information is necessary. Evaluation results show that ArrayDPS outperforms all baseline unsupervised methods while being comparable to supervised methods in terms of SDR. Audio demos and codes are provided at: https://arraydps.github.io/ArrayDPSDemo/ and https://github.com/ArrayDPS/ArrayDPS.
Zhongweiyang Xu, Xulin Fan, Xilin Jiang, Romit Roy Choudhury
ICML1
2024 SPATIALCODEC: Neural Spatial Speech Coding
abstract
In this work, we address the challenge of encoding speech captured by a microphone array using deep learning techniques with the aim of preserving and accurately reconstructing crucial spatial cues embedded in multi-channel recordings. We propose a neural spatial audio coding framework that achieves a high compression ratio, leveraging single-channel neural sub-band codec and SpatialCodec. Our approach encompasses two phases: (i) a neural sub-band codec is designed to encode the reference channel with low bit rates, and (ii), a SpatialCodec captures relative spatial information for accurate multi-channel reconstruction at the decoder end. In addition, we also propose novel evaluation metrics to assess the spatial cue preservation: (i) spatial similarity, which calculates cosine similarity on a spatially intuitive beamspace, and (ii), beamformed audio quality. Our system shows superior spatial performance compared with high bitrate baselines and black-box neural architecture. Demos are available at https://xzwy.github.io/SpatialCodecDemo. Codes and models are available at https://github.com/XZWY/SpatialCodec.
Zhongweiyang Xu, Yong Xu 0004, Vinay Kothapally, Heming Wang, Muqiao Yang, Dong Yu 0001
ICASSP1
2024 uSee: Unified Speech Enhancement And Editing with Conditional Diffusion Models
abstract
Speech enhancement aims to improve the quality of speech signals in terms of quality and intelligibility, and speech editing refers to the process of editing the speech according to specific user needs. In this paper, we propose a Unified Speech Enhancement and Editing (uSee) model with conditional diffusion models to handle various tasks at the same time in a generative manner. Specifically, by providing multiple types of conditions including self-supervised learning embeddings and proper text prompts to the score-based diffusion model, we can enable controllable generation of the unified speech enhancement and editing model to perform corresponding actions on the source speech. Our experiments show that our proposed uSee model can achieve superior performance in both speech denoising and dereverberation compared to other related generative speech enhancement models, and can perform speech editing given desired environmental sound text description, signal-to-noise ratios (SNR), and room impulse responses (RIR). Demos of the generated speech are available at https://muqiaoy.github.io/usee.
Muqiao Yang, Yong Xu 0004, Zhongweiyang Xu, Heming Wang, Bhiksha Raj, Dong Yu 0001
ICASSP4
2024 FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses
Zhongweiyang Xu, Ali Aroudi, Ke Tan 0001, Ashutosh Pandey 0004, Jung-Suk Lee, Buye Xu, Francesco Nesta
INTERSPEECH1
2023 Dual-Path Cross-Modal Attention for Better Audio-Visual Speech Extraction
abstract
Audiovisual target speaker extraction is the task of separating, from an audio mixture, the speaker whose face is visible in an accompanying video. Published approaches typically upsample the video or downsample the audio, then fuse the two streams using concatenation, multiplication, or cross-modal attention. This paper proposes, instead, to use a dual-path attention architecture in which the audio chunk length is comparable to the duration of a video frame. Audio is transformed by intra-chunk attention, concatenated to video features, then transformed by inter-chunk attention. Because of residual connections, the audio and video features remain logically distinct across multiple network layers, therefore dual-path audiovisual feature fusion can be performed repeatedly across multiple layers. When given 2-5-speaker mixtures constructed from the challenging LRS3 test set, results are about 7dB better than Con-vTasNet or AV-ConvTasNet, with the performance gap widening slightly as the number of speakers increases.
Zhongweiyang Xu, Xulin Fan, Mark Hasegawa-Johnson
ICASSP1