EDBT 2026 Demo / reviewers in the wild / expert
Zhongweiyang Xu
dblp:324/2576
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2025
0009-0009-3477-0170ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
1 paper |
Audio and music processing · 100% | |
| Artificial intelligence
1 paper |
Generative modeling · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing › source separation
speech separation |
0.9 | 1 | 2025 | ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion Prior · ICML 2025 |
Machine learning › Generative modeling
diffusion model |
0.3 | 1 | 2025 | ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion Prior · ICML 2025 |
Machine learning › Generative modeling › diffusion model › diffusion sampling
diffusion posterior sampling |
0.3 | 1 | 2025 | ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion Prior · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
relative transfer function estimation · 1.7optimization · 1.7diffusion posterior sampling · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion PriorabstractBlind Speech Separation (BSS) aims to separate multiple speech sources from audio mixtures
recorded by a microphone array. The problem is
challenging because it is a blind inverse problem,
i.e., the microphone array geometry, the room impulse response (RIR), and the speech sources, are
all unknown. We propose ArrayDPS to solve the
BSS problem in an unsupervised, array-agnostic,
and generative manner. The core idea builds on
diffusion posterior sampling (DPS), but unlike
DPS where the likelihood is tractable, ArrayDPS
must approximate the likelihood by formulating
a separate optimization problem. The solution to the optimization approximates room acoustics
and the relative transfer functions between microphones. These approximations, along with the
diffusion priors, iterate through the ArrayDPS
sampling process and ultimately yield separated
voice sources. We only need a simple single-speaker speech diffusion model as a prior, along
with the mixtures recorded at the microphones; no
microphone array information is necessary. Evaluation results show that ArrayDPS outperforms
all baseline unsupervised methods while being
comparable to supervised methods in terms of
SDR. Audio demos and codes are provided at:
https://arraydps.github.io/ArrayDPSDemo/ and
https://github.com/ArrayDPS/ArrayDPS. Zhongweiyang Xu, Xulin Fan, Xilin Jiang, Romit Roy Choudhury |
ICML | 1 |
| 2024 | SPATIALCODEC: Neural Spatial Speech CodingabstractIn this work, we address the challenge of encoding speech captured by a microphone array using deep learning techniques with the aim of preserving and accurately reconstructing crucial spatial cues embedded in multi-channel recordings. We propose a neural spatial audio coding framework that achieves a high compression ratio, leveraging single-channel neural sub-band codec and SpatialCodec. Our approach encompasses two phases: (i) a neural sub-band codec is designed to encode the reference channel with low bit rates, and (ii), a SpatialCodec captures relative spatial information for accurate multi-channel reconstruction at the decoder end. In addition, we also propose novel evaluation metrics to assess the spatial cue preservation: (i) spatial similarity, which calculates cosine similarity on a spatially intuitive beamspace, and (ii), beamformed audio quality. Our system shows superior spatial performance compared with high bitrate baselines and black-box neural architecture. Demos are available at https://xzwy.github.io/SpatialCodecDemo. Codes and models are available at https://github.com/XZWY/SpatialCodec. Zhongweiyang Xu, Yong Xu 0004, Vinay Kothapally, Heming Wang, Muqiao Yang, Dong Yu 0001 |
ICASSP | 1 |
| 2024 | uSee: Unified Speech Enhancement And Editing with Conditional Diffusion ModelsabstractSpeech enhancement aims to improve the quality of speech signals in terms of quality and intelligibility, and speech editing refers to the process of editing the speech according to specific user needs. In this paper, we propose a Unified Speech Enhancement and Editing (uSee) model with conditional diffusion models to handle various tasks at the same time in a generative manner. Specifically, by providing multiple types of conditions including self-supervised learning embeddings and proper text prompts to the score-based diffusion model, we can enable controllable generation of the unified speech enhancement and editing model to perform corresponding actions on the source speech. Our experiments show that our proposed uSee model can achieve superior performance in both speech denoising and dereverberation compared to other related generative speech enhancement models, and can perform speech editing given desired environmental sound text description, signal-to-noise ratios (SNR), and room impulse responses (RIR). Demos of the generated speech are available at https://muqiaoy.github.io/usee. Muqiao Yang, Yong Xu 0004, Zhongweiyang Xu, Heming Wang, Bhiksha Raj, Dong Yu 0001 |
ICASSP | 4 |
| 2024 | FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses
Zhongweiyang Xu, Ali Aroudi, Ke Tan 0001, Ashutosh Pandey 0004, Jung-Suk Lee, Buye Xu, Francesco Nesta |
INTERSPEECH | 1 |
| 2023 | Dual-Path Cross-Modal Attention for Better Audio-Visual Speech ExtractionabstractAudiovisual target speaker extraction is the task of separating, from an audio mixture, the speaker whose face is visible in an accompanying video. Published approaches typically upsample the video or downsample the audio, then fuse the two streams using concatenation, multiplication, or cross-modal attention. This paper proposes, instead, to use a dual-path attention architecture in which the audio chunk length is comparable to the duration of a video frame. Audio is transformed by intra-chunk attention, concatenated to video features, then transformed by inter-chunk attention. Because of residual connections, the audio and video features remain logically distinct across multiple network layers, therefore dual-path audiovisual feature fusion can be performed repeatedly across multiple layers. When given 2-5-speaker mixtures constructed from the challenging LRS3 test set, results are about 7dB better than Con-vTasNet or AV-ConvTasNet, with the performance gap widening slightly as the number of speakers increases. Zhongweiyang Xu, Xulin Fan, Mark Hasegawa-Johnson |
ICASSP | 1 |