Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yiwen Wang 0009

dblp:00/4918-9 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Audio and music processing · 100%
Artificial intelligence
1 paper
Generative modeling · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing › source separation › speech separation
single-channel speech separation
0.712023
PGSS: Pitch-Guided Speech Separation · AAAI 2023
Audio and music processing › source separation
speech separation
0.712023
PGSS: Pitch-Guided Speech Separation · AAAI 2023
Machine learning › Generative modeling › generative adversarial network
conditional GAN
0.212023
PGSS: Pitch-Guided Speech Separation · AAAI 2023
Machine learning › Generative modeling
generative adversarial network
0.212023
PGSS: Pitch-Guided Speech Separation · AAAI 2023

Methods — techniques the papers use, named apart from their topics

pitch extraction · 1.3conditional generative adversarial network · 1.3adversarial training · 1.3
YearPublicationVenuePosition
2025 Cross-attention Inspired Selective State Space Models for Target Sound Extraction
abstract
The Transformer model, particularly its cross-attention module, is widely used for feature fusion in target sound extraction which extracts the signal of interest based on given clues. Despite its effectiveness, this approach suffers from low computational efficiency. Recent advancements in state space models, notably the latest work Mamba, have shown comparable performance to Transformer-based methods while significantly reducing computational complexity in various tasks. However, Mamba’s applicability in target sound extraction is limited due to its inability to capture dependencies between different sequences as the cross-attention does. In this paper, we propose CrossMamba for target sound extraction, which leverages the hidden attention mechanism of Mamba to compute dependencies between the given clues and the audio mixture. The calculation of Mamba can be divided to the query, key and value. We utilize the clue to generate the query and the audio mixture to derive the key and value, adhering to the principle of the cross-attention mechanism in Transformers. Experimental results from two representative target sound extraction methods validate the efficacy of the proposed CrossMamba.
Donghang Wu, Yiwen Wang 0009, Xihong Wu, Tianshu Qu
ICASSP2
2025 Position also matters! Separating Same Instruments in String Quartet using Timbral and Positional Cues
Yuetonghui Xu, Yiwen Wang 0009, Xihong Wu
INTERSPEECH2
2024 TSE-PI: Target Sound Extraction under Reverberant Environments with Pitch Information
Yiwen Wang 0009, Xihong Wu
INTERSPEECH1
2023 PGSS: Pitch-Guided Speech Separation
abstract
Monaural speech separation aims to separate concurrent speakers from a single-microphone mixture recording. Inspired by the effect of pitch priming in auditory scene analysis (ASA) mechanisms, a novel pitch-guided speech separation framework is proposed in this work. The prominent advantage of this framework is that both the permutation problem and the unknown speaker number problem existing in general models can be avoided by using pitch contours as the primary means to guide the target speaker. In addition, adversarial training is applied, instead of a traditional time-frequency mask, to improve the perceptual quality of separated speech. Specifically, the proposed framework can be divided into two phases: pitch extraction and speech separation. The former aims to extract pitch contour candidates for each speaker from the mixture, modeling the bottom-up process in ASA mechanisms. Any pitch contour can be selected as the condition in the second phase to separate the corresponding speaker, where a conditional generative adversarial network (CGAN) is applied. The second phase models the effect of pitch priming in ASA. Experiments on the WSJ0-2mix corpus reveal that the proposed approaches can achieve higher pitch extraction accuracy and better separation performance, compared to the baseline models, and have the potential to be applied to SOTA architectures.
Xiang Li 0072, Yiwen Wang 0009, Xihong Wu, Jing Chen 0019
AAAI2
2023 TT-Net: Dual-Path Transformer Based Sound Field Translation in the Spherical Harmonic Domain
abstract
In the current method for the sound field translation tasks based on spherical harmonic (SH) analysis, the solution based on the additive theorem usually faces the problem of singular values caused by large matrix condition numbers. The influence of different distances and frequencies of the spherical radial function on the stability of the translation matrix will affect the accuracy of the SH coefficients at the selected point. Due to the problems mentioned above, we propose a neural network scheme based on the dual-path transformer. More specifically, the dual-path network is constructed by the selfattention module along the two dimensions of the frequency and order axes. The transform-average-concatenate layer and upscaling layer are introduced in the network, which provides solutions for multiple sampling points and upscaling. Numerical simulation results indicate that both the working frequency range and the distance range of the translation are extended. More accurate higher-order SH coefficients are obtained with the proposed dual-path network.
Yiwen Wang 0009, Zijian Lan, Xihong Wu, Tianshu Qu
ICASSP1