EDBT 2026 Demo / reviewers in the wild / expert
Kangjie Dong
dblp:407/9211
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0004-0657-3069ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
2 papers |
Audio and music processing · 79% Image and video processing · 21% | |
| Artificial intelligence
1 paper |
Transfer learning and domain adaptation · 87% Learning paradigms · 13% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing
music information retrieval |
1.9 | 2 | 2026 | Every Little Bit Helps: Exploring Better Utilization of Unlabeled Data for Semi-supervised Singing Melody Extraction Using Multi-bands Diffusion Model · AAAI 2026 DUDA: A Two-stage Decoupling Unsupervised Domain Adaptation Framework for Semi-supervised Singing Melody Extraction from Polyphonic Music · ACM Multimedia 2025 |
Audio and music processing › music information retrieval › singing information processing
singing melody extraction |
1.9 | 2 | 2026 | Every Little Bit Helps: Exploring Better Utilization of Unlabeled Data for Semi-supervised Singing Melody Extraction Using Multi-bands Diffusion Model · AAAI 2026 DUDA: A Two-stage Decoupling Unsupervised Domain Adaptation Framework for Semi-supervised Singing Melody Extraction from Polyphonic Music · ACM Multimedia 2025 |
Image and video processing
data augmentation |
1.0 | 1 | 2026 | Every Little Bit Helps: Exploring Better Utilization of Unlabeled Data for Semi-supervised Singing Melody Extraction Using Multi-bands Diffusion Model · AAAI 2026 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › low-resource domain adaptation
semi-supervised domain adaptation |
0.9 | 1 | 2025 | DUDA: A Two-stage Decoupling Unsupervised Domain Adaptation Framework for Semi-supervised Singing Melody Extraction from Polyphonic Music · ACM Multimedia 2025 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation |
0.9 | 1 | 2025 | DUDA: A Two-stage Decoupling Unsupervised Domain Adaptation Framework for Semi-supervised Singing Melody Extraction from Polyphonic Music · ACM Multimedia 2025 |
Machine learning › Learning paradigms › semi-supervised learning
consistency regularization |
0.3 | 1 | 2025 | DUDA: A Two-stage Decoupling Unsupervised Domain Adaptation Framework for Semi-supervised Singing Melody Extraction from Polyphonic Music · ACM Multimedia 2025 |
Methods — techniques the papers use, named apart from their topics
domain adaptation · 1.7decoupled feature alignment · 1.7consistency regularization · 1.7multi-band augmentation · 1.0diffusion model · 1.0confidence estimation · 1.0channel cross attention · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Every Little Bit Helps: Exploring Better Utilization of Unlabeled Data for Semi-supervised Singing Melody Extraction Using Multi-bands Diffusion ModelabstractSemi-supervised singing melody extraction (SSME) is one of the key tasks in the field of music information retrieval (MIR). Recently, several SSME methods have been proposed and achieved remarkable successes. However, existing methods are still facing two critical issues: firstly, there is a lack of an effective data augmentation method for SSME, which results in insufficient utilization of unlabeled data. Secondly, existing SSME methods discards too much unlabeled data in the stage of consistency regularization, which hinders the further improvements of SSME task. In this paper, we present \emph{ELH-SME}, a novel framework that better utilizes the unlabeled musical data for SSME task. Specifically, our proposed ELH-SME framework consists of three modules: (1) we first propose a diffusion-based multi-bands augmentation (DMA) method to increase the amounts of training data. The proposed DMA methods employs a diffusion model to generate perturbation at the specific frequency bands in an end-to-end manner, thereby avoiding sharply perturbations to the spectrogram. (2) To improve the utilization rate of unlabeled data, we suggest a global-class confidence (GCC) module. During the phase of consistency regularization, we consider both the global-wise and class-wise confidence values, improving the utilization rate of unlabeled data. (3) To further improve the utilization of unlabeled data, we also propose to enhance the representation capability of unlabeled data by extracting channel-level features from labeled data via channel cross attention (CCA). We evaluate our proposed framework on several well-known public available datasets, and the conducted experiments demonstrate the effectiveness of our method. Shuai Yu 0002, Xiaoliang He, Kangjie Dong, Yi Yu 0001 |
AAAI | 3 |
| 2025 | A Mamba-based Network for Semi-supervised Singing Melody Extraction Using Confidence Binary RegularizationabstractSinging melody extraction (SME) is a key task in the field of music information retrieval. However, existing methods are facing several limitations: firstly, prior models use transformers to capture the contextual dependencies, which requires quadratic computation resulting in low efficiency in the inference stage. Secondly, prior works typically rely on frequency-supervised methods to estimate the fundamental frequency (f0), which ignores that the musical performance is actually based on notes. Thirdly, transformers typically require large amounts of labeled data to achieve optimal performances, but the SME task lacks of sufficient annotated data. To address these issues, in this paper, we propose a mamba-based network, called SpectMamba, for semi-supervised singing melody extraction using confidence binary regularization. In particular, we begin by introducing vision mamba to achieve computational linear complexity. Then, we propose a novel note-f0 decoder that allows the model to better mimic the musical performance. Further, to alleviate the scarcity of the labeled data, we introduce a confidence binary regularization (CBR) module to leverage the unlabeled data by maximizing the probability of the correct classes. The proposed method is evaluated on several public datasets and the conducted experiments demonstrate the effectiveness of our proposed method1. Xiaoliang He, Kangjie Dong, Jingkai Cao, Shuai Yu 0002, Wei Li 0012, Yi Yu 0001 |
ICASSP | 2 |
| 2025 | Ultra Lightweight Singing Melody Extraction via Combination of Convolution and MLPabstractSinging melody extraction serves as an important foundation in the realm of music information retrieval (MIR). Although fully convolutional neural networks (CNNs) are commonly employed for singing melody extraction, they are constrained by inductive biases and face challenges in establishing long range dependency. Transformer-based networks have better performance, but the computational load is high. Recently, many multi-layer perceptron (MLP) architectures have been applied for a variety of computer vision tasks, demonstrating competitive performance. However, its potential ability in the task of singing melody extraction remains to be further explored. In this paper, we propose the lightweight convolutional MLP (LcMLP), an ultra lightweight model without sacrificing the performance. Firstly, we improve the original MLP-Mixer. We change the sequential MLPs to parallel ones and add some skip connections. Secondly, we propose a multi-level convolution fusion module that facilitates the interaction of features at various depths in MLP-Mixer. We conducted extensive experiments on several well-known public datasets, and our model demonstrates significant advantages in inference speed and computational load, while also achieving competitive performance. Kangjie Dong, Qiubo Huang, Shuai Yu 0002, Wei Li 0012 |
ICASSP | 2 |
| 2025 | DUDA: A Two-stage Decoupling Unsupervised Domain Adaptation Framework for Semi-supervised Singing Melody Extraction from Polyphonic MusicabstractSemi-supervised singing melody extraction (SSME) is one of the key tasks in the field of music information retrieval (MIR). However, there are two critical issues that remain to be addressed in data limited scenarios. Firstly, the prior unsupervised domain adaptation methods for SSME typically rely on learning domain-agnostic features at holistic level, which ignores the associations between holistic information (i.e., fundamental frequency) and fine-grained information (i.e., tone and octave). Secondly, the fine-grained information can be utilized to judge the availability of unlabeled data, which is ignored by prior methods. There is a lack of a consistency regularization method that utilizes fine-grained information to validate the availability of unlabeled data. To address these issues, in this paper, we propose a novel two-stage decoupling unsupervised domain adaptation framework for semi-supervised singing melody extraction, termed as DUDA. Specifically, in the first stage, we decouple the holistic information into fine-grained information: tone and octave, and narrow the domain gap at the tone and octave level, respectively. This enables the model to align the tone-octave information between source and target domains for better feature distribution. Then, we leverage the learned domain-agnostic fine-grained features as additional information to obtain domain-agnostic holistic features. We also suggest to align intra-domain, inter-domain, and sample-level features to further improve the performances. In the second stage, we propose a novel tone-octave consistency regularization method by leveraging the extracted fine-grained information to judge the availability of unlabeled data. We evaluate our proposed framework on several well-known public datasets, and the conducted experiments demonstrate the effectiveness of our method. Shuai Yu 0002, Xiaoliang He, Kangjie Dong, Yi Yu 0001 |
ACM Multimedia | 3 |