Dichucheng Li

dblp:327/8743 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0008-7817-8948ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Decomposed Retrieval-Edit-Rerank Framework for Chord Generation
abstract
Chord generation is an inherently constrained creative task that requires balancing stylistic diversity with music-theoretic feasibility. Existing approaches typically entangle candidate generation and constraint enforcement within a single model, making the diversity–feasibility trade-off difficult to control and interpret.
Qiqi He, Dichucheng Li, Xiaoheng Sun
ICMR2
2025 Piano Transcription by Hierarchical Language Modeling with Pretrained Roll-based Encoders
abstract
Automatic Music Transcription (AMT), aiming to get musical notes from raw audio, typically uses frame-level systems with piano-roll outputs or language model (LM)-based systems with note-level predictions. However, frame-level systems require manual thresholding, while the LM-based systems struggle with long sequences. In this paper, we propose a hybrid method combining pre-trained roll-based encoders with an LM decoder to leverage the strengths of both methods. Besides, our approach employs a hierarchical prediction strategy, first predicting onset and pitch, then velocity, and finally offset. The hierarchical prediction strategy reduces computational costs by breaking down long sequences into different hierarchies. Evaluated on two benchmark roll-based encoders, our method outperforms traditional piano-roll outputs 0.01 and 0.022 in onset-offset-velocity F1 score, demonstrating its potential as a performance-enhancing plug-in for arbitrary roll-based music transcription encoder. We release the code of this work at https://github.com/yongyizang/AMT_train
Dichucheng Li, Yongyi Zang, Qiuqiang Kong
ICASSP1
2024 Mertech: Instrument Playing Technique Detection Using Self-Supervised Pretrained Model with Multi-Task Finetuning
abstract
Instrument playing techniques (IPTs) constitute a pivotal component of musical expression. However, the development of automatic IPT detection methods suffers from limited labeled data and inherent class imbalance issues. In this paper, we propose to apply a self-supervised learning model pre-trained on large-scale unlabeled music data and finetune it on IPT detection tasks. This approach addresses data scarcity and class imbalance challenges. Recognizing the significance of pitch in capturing the nuances of IPTs and the importance of onset in locating IPT events, we investigate multi-task finetuning with pitch and onset detection as auxiliary tasks. Additionally, we apply a post-processing approach for event-level prediction, where an IPT activation initiates an event only if the onset output confirms an onset in that frame. Our method outperforms prior approaches in both frame-level and event-level metrics across multiple IPT benchmark datasets. Further experiments demonstrate the efficacy of multi-task finetuning on each IPT class.1
Dichucheng Li, Yinghao Ma, Weixing Wei, Qiuqiang Kong, Yulun Wu 0002, Mingjin Che, Emmanouil Benetos, Wei Li 0012
ICASSP1
2024 MS-SENet: Enhancing Speech Emotion Recognition Through Multi-Scale Feature Fusion with Squeeze-and-Excitation Blocks
abstract
Speech Emotion Recognition (SER) has become a growing focus of research in human-computer interaction. Spatiotemporal features play a crucial role in SER, yet current research lacks comprehensive spatiotemporal feature learning. This paper focuses on addressing this gap by proposing a novel approach. In this paper, we employ Convolutional Neural Network (CNN) with varying kernel sizes for spatial and temporal feature extraction. Additionally, we introduce Squeeze-and-Excitation (SE) modules to capture and fuse multi-scale features, facilitating effective information fusion for improved emotion recognition and a deeper understanding of the temporal evolution of speech emotion. Moreover, we employ skip-connections and Spatial Dropout (SD) layers to prevent overfitting and increase the model’s depth. Our method outperforms the previous state-of-the-art method, achieving an average UAR and WAR improvement of 1.62% and 1.32%, respectively, across six benchmark SER datasets. Further experiments demonstrated that our method can fully extract spatiotemporal features in low-resource conditions.
Mengbo Li, Yuanzhong Zheng, Dichucheng Li, Yulun Wu 0002, Yaoxuan Wang, Haojun Fei
ICASSP3
2024 Improving Drum Source Separation with Temporal-Frequency Statistical Descriptors
abstract
Drum Source Separation (DSS) aims to separate drum mixtures into individual drum sounds, such as kick and snare. Deep neural network methods have been successfully applied for source separation. However, due to the limited size of existing datasets and the strongly overlap of drums in frequency and time, these methods still have certain shortcomings. To address these challenges, we construct a large drum sound dataset and propose a novel training objective to improve performance of DSS task. The training objective leverages three temporal-frequency statistical descriptors (spectral centroid, spectral spread, and spectral flux) to separate drum sources. Our experimental results demonstrate that our method can make a SDR improvement of 0.98 dB on UNet and 1.07 dB on MERT. Furthermore, our method achieves consistent improvements in low-resource and cross-dataset scenarios. Our code and dataset are available at https://github.com/150042/Drum-Separation-TF.
Dichucheng Li, Xinlu Liu, Yongwei Gao, Wei Li 0012
ICME4
2024 Harmonic Frequency-Separable Transformer for Instrument-Agnostic Music Transcription
abstract
Automatic Music Transcription (AMT) aims to convert music audio into symbolic representations. Recently, transformer-based methods have been successfully applied to instrument-agnostic music transcription. This allows transcription models can no longer focus on specific characteristics for an instrument class. However, these transformer-based methods designs for AMT were mainly motivated by other research fields and uses additional large-scale datasets, without considering the intrinsic features and patterns of the music signals. In this paper, we propose the Harmonic Frequency-Separable Transformer (HFSFormer), providing effective prior information based on music knowledge for instrument-agnostic transcription. The HFSFormer can capture the harmonic structure of music and separate time-frequency representations to decouple multiple pitches and different timbres, which can better explicitly model the note’s onset/offset and pitch. Experimental results show that our proposed method outperforms state-of-the-art peers on public datasets while having an order of magnitude fewer parameters.
Yulun Wu 0002, Weixing Wei, Dichucheng Li, Mengbo Li, Yi Yu 0001, Yongwei Gao, Wei Li 0012
ICME3
2023 Frame-Level Multi-Label Playing Technique Detection Using Multi-Scale Network and Self-Attention Mechanism
abstract
Instrument playing technique (IPT) is a key element of musical presentation. However, most of the existing works for IPT detection only concern monophonic music signals, yet little has been done to detect IPTs in polyphonic instrumental solo pieces with overlapping IPTs or mixed IPTs. In this paper, we formulate it as a frame-level multi-label classification problem and apply it to Guzheng, a Chinese plucked string instrument. We create a new dataset, Guzheng Tech99, containing Guzheng recordings and onset, offset, pitch, IPT annotations of each note. Because different IPTs vary a lot in their lengths, we propose a new method to solve this problem using multi-scale network and self-attention. The multi-scale network extracts features from different scales, and the self-attention mechanism applied to the feature maps at the coarsest scale further enhances the long-range feature extraction. Our approach outperforms existing works by a large margin, indicating its effectiveness in IPT detection.
Dichucheng Li, Mingjin Che, Wenwu Meng, Yulun Wu 0002, Yi Yu 0001, Wei Li 0012
ICASSP1
2022 Multimodal Music Emotion Recognition with Hierarchical Cross-Modal Attention Network
abstract
Computational music emotion recognition is to recognize the emotional content in music tracks. In computational music emotion recognition studies, researchers have paid close attention to the audio content of the music tracks. Although lyrics content and music context contribute greatly to the perceived emotion, these kinds of emotional information are usually ignored. Based on this finding, we propose a multimodal music emotion recognition method jointly predicting the valence and arousal values by combining the audio, lyrics, track name, and artist of a given track. Audio features, lyrics features and context features are extracted separately and fused by a cross-modal attention mechanism, forming a hierarchical structure. Our proposed model outperforms two baselines by a large margin and achieves state-of-the-art performance on two public datasets.
Ganghui Ru, Yi Yu 0001, Yulun Wu 0002, Dichucheng Li, Wei Li 0012
ICME5