EDBT 2026 Demo / reviewers in the wild / expert
Yulun Wu 0002
dblp:218/8680-2
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Music Denoising with Channel Attention and Multi-Scale Sequence EncodingabstractMusic denoising aims to suppress unwanted noise while preserving the fidelity of musical signals. In real-world recordings such as pub performances, street busking, and concerts, smartphone audio is often degraded by complex and non-stationary noise sources (e.g., cheering, chatter, and environmental sounds), making robust and efficient denoising particularly challenging. While recent neural architectures developed for audio separation and enhancement have shown promise, their direct application to music denoising remains underexplored, especially under realistic recording conditions. To address this gap, we propose MusicTLE-U, a lightweight U-shaped denoising architecture built on a band-split backbone and augmented with two complementary modules: TLE (Temporal-LSTM-ECA) and MSSE (Multi-Scale Sequence Encoding). The TLE module improves temporal modeling by combining fully connected layers, an LSTM, and Efficient Channel Attention in a compact manner, while MSSE captures hierarchical global context through downsampling, multi-head attention, and upsampling operations. Experiments conducted on real-world noisy music mixtures demonstrate that MusicTLE-U achieves competitive or improved performance in objective evaluation metrics compared to strong baseline models, while maintaining computational efficiency. To support reproducibility, we release the training and evaluation protocol along with code and audio examples. Seungmin Ha, Wei Li 0012, Yulun Wu 0002 |
ICMR | 3 |
| 2025 | KCE-Unet: A novel music denoising method with KANConv ECA UnetabstractDuring concerts, people often spontaneously record memorable moments with their phones. However, these recordings are frequently accompanied by noise, such as cheering and applause, which diminishes the playback experience. In this paper, we introduce a novel task specifically designed for denoising music in concert environments, a challenge that has been largely overlooked in previous research. To support this task, we created a new concert denoising dataset that includes songs performed in various major languages at concerts, with noise segments like cheering and applause. Building on this, we propose KANConv ECA Unet (KCE-Unet), a method that combines the U-Net network, efficient channel attention (ECA), and the recently proposed KAN network to flexibly remove noise in the mid-to-high frequency range of spectrograms. Extensive experiments demonstrate that our method outperforms previous models in denoising performance and effectively restore disrupted musical structures. Yulun Wu 0002, Ganghui Ru, Yi Yu 0001, Wei Li 0012 |
ICASSP | 2 |
| 2025 | HingeNet: A Harmonic-Aware Fine-Tuning Approach for Beat TrackingabstractFine-tuning pre-trained foundation models has made significant progress in music information retrieval. However, applying these models to beat tracking tasks remains unexplored as the limited annotated data renders conventional fine-tuning methods ineffective. To address this challenge, we propose HingeNet, a novel and general parameter-efficient fine-tuning method specifically designed for beat tracking tasks. HingeNet is a lightweight and separable network, visually resembling a hinge, designed to tightly interface with pre-trained foundation models by using their intermediate feature representations as input. This unique architecture grants HingeNet broad generalizability, enabling effective integration with various pre-trained foundation models. Furthermore, considering the significance of harmonics in beat tracking, we introduce harmonic-aware mechanism during the fine-tuning process to better capture and emphasize the harmonic structures in musical signals. Experiments on benchmark datasets demonstrate that HingeNet achieves state-of-the-art performance in beat and downbeat tracking. Ganghui Ru, Jieying Wang, Yulun Wu 0002, Yi Yu 0001, Nannan Jiang, Wei Wang 0414, Wei Li 0012 |
ICME | 4 |
| 2025 | BeatFM: Improving Beat Tracking with Pre-trained Music Foundation ModelabstractBeat tracking is a widely researched topic in music information retrieval. However, current beat tracking methods face challenges due to the scarcity of labeled data, which limits their ability to generalize across diverse musical styles and accurately capture complex rhythmic structures. To overcome these challenges, we propose a novel beat tracking paradigm BeatFM, which introduces a pre-trained music foundation model and leverages its rich semantic knowledge to improve beat tracking performance. Pre-training on diverse music datasets endows music foundation models with a robust understanding of music, thereby effectively addressing these challenges. To further adapt it for beat tracking, we design a plug-and-play multi-dimensional semantic aggregation module, which is composed of three parallel sub-modules, each focusing on semantic aggregation in the temporal, frequency, and channel domains, respectively. Extensive experiments demonstrate that our method achieves state-of-the-art performance in beat and downbeat tracking across multiple benchmark datasets. Ganghui Ru, Jieying Wang, Yulun Wu 0002, Yi Yu 0001, Nannan Jiang, Wei Wang 0414, Wei Li 0012 |
ICME | 4 |
| 2024 | Mertech: Instrument Playing Technique Detection Using Self-Supervised Pretrained Model with Multi-Task FinetuningabstractInstrument playing techniques (IPTs) constitute a pivotal component of musical expression. However, the development of automatic IPT detection methods suffers from limited labeled data and inherent class imbalance issues. In this paper, we propose to apply a self-supervised learning model pre-trained on large-scale unlabeled music data and finetune it on IPT detection tasks. This approach addresses data scarcity and class imbalance challenges. Recognizing the significance of pitch in capturing the nuances of IPTs and the importance of onset in locating IPT events, we investigate multi-task finetuning with pitch and onset detection as auxiliary tasks. Additionally, we apply a post-processing approach for event-level prediction, where an IPT activation initiates an event only if the onset output confirms an onset in that frame. Our method outperforms prior approaches in both frame-level and event-level metrics across multiple IPT benchmark datasets. Further experiments demonstrate the efficacy of multi-task finetuning on each IPT class.1 Dichucheng Li, Yinghao Ma, Weixing Wei, Qiuqiang Kong, Yulun Wu 0002, Mingjin Che, Emmanouil Benetos, Wei Li 0012 |
ICASSP | 5 |
| 2024 | MS-SENet: Enhancing Speech Emotion Recognition Through Multi-Scale Feature Fusion with Squeeze-and-Excitation BlocksabstractSpeech Emotion Recognition (SER) has become a growing focus of research in human-computer interaction. Spatiotemporal features play a crucial role in SER, yet current research lacks comprehensive spatiotemporal feature learning. This paper focuses on addressing this gap by proposing a novel approach. In this paper, we employ Convolutional Neural Network (CNN) with varying kernel sizes for spatial and temporal feature extraction. Additionally, we introduce Squeeze-and-Excitation (SE) modules to capture and fuse multi-scale features, facilitating effective information fusion for improved emotion recognition and a deeper understanding of the temporal evolution of speech emotion. Moreover, we employ skip-connections and Spatial Dropout (SD) layers to prevent overfitting and increase the model’s depth. Our method outperforms the previous state-of-the-art method, achieving an average UAR and WAR improvement of 1.62% and 1.32%, respectively, across six benchmark SER datasets. Further experiments demonstrated that our method can fully extract spatiotemporal features in low-resource conditions. Mengbo Li, Yuanzhong Zheng, Dichucheng Li, Yulun Wu 0002, Yaoxuan Wang, Haojun Fei |
ICASSP | 4 |
| 2024 | Harmonic Frequency-Separable Transformer for Instrument-Agnostic Music TranscriptionabstractAutomatic Music Transcription (AMT) aims to convert music audio into symbolic representations. Recently, transformer-based methods have been successfully applied to instrument-agnostic music transcription. This allows transcription models can no longer focus on specific characteristics for an instrument class. However, these transformer-based methods designs for AMT were mainly motivated by other research fields and uses additional large-scale datasets, without considering the intrinsic features and patterns of the music signals. In this paper, we propose the Harmonic Frequency-Separable Transformer (HFSFormer), providing effective prior information based on music knowledge for instrument-agnostic transcription. The HFSFormer can capture the harmonic structure of music and separate time-frequency representations to decouple multiple pitches and different timbres, which can better explicitly model the note’s onset/offset and pitch. Experimental results show that our proposed method outperforms state-of-the-art peers on public datasets while having an order of magnitude fewer parameters. Yulun Wu 0002, Weixing Wei, Dichucheng Li, Mengbo Li, Yi Yu 0001, Yongwei Gao, Wei Li 0012 |
ICME | 1 |
| 2023 | Frame-Level Multi-Label Playing Technique Detection Using Multi-Scale Network and Self-Attention MechanismabstractInstrument playing technique (IPT) is a key element of musical presentation. However, most of the existing works for IPT detection only concern monophonic music signals, yet little has been done to detect IPTs in polyphonic instrumental solo pieces with overlapping IPTs or mixed IPTs. In this paper, we formulate it as a frame-level multi-label classification problem and apply it to Guzheng, a Chinese plucked string instrument. We create a new dataset, Guzheng Tech99, containing Guzheng recordings and onset, offset, pitch, IPT annotations of each note. Because different IPTs vary a lot in their lengths, we propose a new method to solve this problem using multi-scale network and self-attention. The multi-scale network extracts features from different scales, and the self-attention mechanism applied to the feature maps at the coarsest scale further enhances the long-range feature extraction. Our approach outperforms existing works by a large margin, indicating its effectiveness in IPT detection. Dichucheng Li, Mingjin Che, Wenwu Meng, Yulun Wu 0002, Yi Yu 0001, Wei Li 0012 |
ICASSP | 4 |
| 2023 | MFAE: Masked frame-level autoencoder with hybrid-supervision for low-resource music transcriptionabstractAutomantic Music Transcription (AMT) is an essential topic in music information retrieval (MIR), and it aims to transcribe audio recordings into symbolic representations. Recently, large-scale piano datasets with high-quality notations have been proposed for high-resolution piano transcription, which resulted in domain-specific AMT models achieved state-of- the-art results. However, those methods are hardly generalized to other ’low-resource’ instruments (such as guitar, cello, clarinet, etc.) transcription. In this paper, we propose a hybrid-supervised framework, the masked frame-level autoencoder (MFAE), to solve this issue. The proposed MFAE reconstructs the frame-level features of low-resource data to understand generic representations of low-resource instruments and improves low-resource transcription performance. Experimental results on several low- resource datasets (MAPS, MusicNet, and Guitarset) show that our framework achieves state-of-the-art performance in note-wise scores (Note F1 83.4%\64.1%\86.7%, Note-with-offset F1 59.8%\41.4%\71.6%). Moreover, our framework can be well generalized to various genres of instrument transcription, both in data-plentiful and data-limited scenarios. Yulun Wu 0002, Yi Yu 0001, Wei Li 0012 |
ICME | 1 |
| 2022 | Multimodal Music Emotion Recognition with Hierarchical Cross-Modal Attention NetworkabstractComputational music emotion recognition is to recognize the emotional content in music tracks. In computational music emotion recognition studies, researchers have paid close attention to the audio content of the music tracks. Although lyrics content and music context contribute greatly to the perceived emotion, these kinds of emotional information are usually ignored. Based on this finding, we propose a multimodal music emotion recognition method jointly predicting the valence and arousal values by combining the audio, lyrics, track name, and artist of a given track. Audio features, lyrics features and context features are extracted separately and fused by a cross-modal attention mechanism, forming a hierarchical structure. Our proposed model outperforms two baselines by a large margin and achieves state-of-the-art performance on two public datasets. Ganghui Ru, Yi Yu 0001, Yulun Wu 0002, Dichucheng Li, Wei Li 0012 |
ICME | 4 |