Xiaoliang He

dblp:70/2495 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Audio and music processing · 84% Image and video processing · 16%
Artificial intelligence
2 papers
Transfer learning and domain adaptation · 50% Learning paradigms · 29% Efficient and distributed learning · 22%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing
music information retrieval
2.632026
Every Little Bit Helps: Exploring Better Utilization of Unlabeled Data for Semi-supervised Singing Melody Extraction Using Multi-bands Diffusion Model · AAAI 2026
DUDA: A Two-stage Decoupling Unsupervised Domain Adaptation Framework for Semi-supervised Singing Melody Extraction from Polyphonic Music · ACM Multimedia 2025
HKDSME: Heterogeneous Knowledge Distillation for Semi-supervised Singing Melody Extraction Using Harmonic Supervision · ACM Multimedia 2024
Audio and music processing › music information retrieval › singing information processing
singing melody extraction
2.632026
Every Little Bit Helps: Exploring Better Utilization of Unlabeled Data for Semi-supervised Singing Melody Extraction Using Multi-bands Diffusion Model · AAAI 2026
DUDA: A Two-stage Decoupling Unsupervised Domain Adaptation Framework for Semi-supervised Singing Melody Extraction from Polyphonic Music · ACM Multimedia 2025
HKDSME: Heterogeneous Knowledge Distillation for Semi-supervised Singing Melody Extraction Using Harmonic Supervision · ACM Multimedia 2024
Image and video processing
data augmentation
1.012026
Every Little Bit Helps: Exploring Better Utilization of Unlabeled Data for Semi-supervised Singing Melody Extraction Using Multi-bands Diffusion Model · AAAI 2026
Machine learning › Transfer learning and domain adaptation › domain adaptation › low-resource domain adaptation
semi-supervised domain adaptation
0.912025
DUDA: A Two-stage Decoupling Unsupervised Domain Adaptation Framework for Semi-supervised Singing Melody Extraction from Polyphonic Music · ACM Multimedia 2025
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.912025
DUDA: A Two-stage Decoupling Unsupervised Domain Adaptation Framework for Semi-supervised Singing Melody Extraction from Polyphonic Music · ACM Multimedia 2025
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.812024
HKDSME: Heterogeneous Knowledge Distillation for Semi-supervised Singing Melody Extraction Using Harmonic Supervision · ACM Multimedia 2024
Machine learning › Learning paradigms
semi-supervised learning
0.812024
HKDSME: Heterogeneous Knowledge Distillation for Semi-supervised Singing Melody Extraction Using Harmonic Supervision · ACM Multimedia 2024
Machine learning › Learning paradigms › semi-supervised learning
consistency regularization
0.312025
DUDA: A Two-stage Decoupling Unsupervised Domain Adaptation Framework for Semi-supervised Singing Melody Extraction from Polyphonic Music · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

domain adaptation · 1.7decoupled feature alignment · 1.7consistency regularization · 1.7semi-supervised learning · 1.5knowledge distillation · 1.5confidence-guided loss · 1.5multi-band augmentation · 1.0diffusion model · 1.0confidence estimation · 1.0channel cross attention · 1.0
YearPublicationVenuePosition
2026 Every Little Bit Helps: Exploring Better Utilization of Unlabeled Data for Semi-supervised Singing Melody Extraction Using Multi-bands Diffusion Model
abstract
Semi-supervised singing melody extraction (SSME) is one of the key tasks in the field of music information retrieval (MIR). Recently, several SSME methods have been proposed and achieved remarkable successes. However, existing methods are still facing two critical issues: firstly, there is a lack of an effective data augmentation method for SSME, which results in insufficient utilization of unlabeled data. Secondly, existing SSME methods discards too much unlabeled data in the stage of consistency regularization, which hinders the further improvements of SSME task. In this paper, we present \emph{ELH-SME}, a novel framework that better utilizes the unlabeled musical data for SSME task. Specifically, our proposed ELH-SME framework consists of three modules: (1) we first propose a diffusion-based multi-bands augmentation (DMA) method to increase the amounts of training data. The proposed DMA methods employs a diffusion model to generate perturbation at the specific frequency bands in an end-to-end manner, thereby avoiding sharply perturbations to the spectrogram. (2) To improve the utilization rate of unlabeled data, we suggest a global-class confidence (GCC) module. During the phase of consistency regularization, we consider both the global-wise and class-wise confidence values, improving the utilization rate of unlabeled data. (3) To further improve the utilization of unlabeled data, we also propose to enhance the representation capability of unlabeled data by extracting channel-level features from labeled data via channel cross attention (CCA). We evaluate our proposed framework on several well-known public available datasets, and the conducted experiments demonstrate the effectiveness of our method.
Shuai Yu 0002, Xiaoliang He, Kangjie Dong, Yi Yu 0001
AAAI2
2025 A Mamba-based Network for Semi-supervised Singing Melody Extraction Using Confidence Binary Regularization
abstract
Singing melody extraction (SME) is a key task in the field of music information retrieval. However, existing methods are facing several limitations: firstly, prior models use transformers to capture the contextual dependencies, which requires quadratic computation resulting in low efficiency in the inference stage. Secondly, prior works typically rely on frequency-supervised methods to estimate the fundamental frequency (f0), which ignores that the musical performance is actually based on notes. Thirdly, transformers typically require large amounts of labeled data to achieve optimal performances, but the SME task lacks of sufficient annotated data. To address these issues, in this paper, we propose a mamba-based network, called SpectMamba, for semi-supervised singing melody extraction using confidence binary regularization. In particular, we begin by introducing vision mamba to achieve computational linear complexity. Then, we propose a novel note-f0 decoder that allows the model to better mimic the musical performance. Further, to alleviate the scarcity of the labeled data, we introduce a confidence binary regularization (CBR) module to leverage the unlabeled data by maximizing the probability of the correct classes. The proposed method is evaluated on several public datasets and the conducted experiments demonstrate the effectiveness of our proposed method1.
Xiaoliang He, Kangjie Dong, Jingkai Cao, Shuai Yu 0002, Wei Li 0012, Yi Yu 0001
ICASSP1
2025 DUDA: A Two-stage Decoupling Unsupervised Domain Adaptation Framework for Semi-supervised Singing Melody Extraction from Polyphonic Music
abstract
Semi-supervised singing melody extraction (SSME) is one of the key tasks in the field of music information retrieval (MIR). However, there are two critical issues that remain to be addressed in data limited scenarios. Firstly, the prior unsupervised domain adaptation methods for SSME typically rely on learning domain-agnostic features at holistic level, which ignores the associations between holistic information (i.e., fundamental frequency) and fine-grained information (i.e., tone and octave). Secondly, the fine-grained information can be utilized to judge the availability of unlabeled data, which is ignored by prior methods. There is a lack of a consistency regularization method that utilizes fine-grained information to validate the availability of unlabeled data. To address these issues, in this paper, we propose a novel two-stage decoupling unsupervised domain adaptation framework for semi-supervised singing melody extraction, termed as DUDA. Specifically, in the first stage, we decouple the holistic information into fine-grained information: tone and octave, and narrow the domain gap at the tone and octave level, respectively. This enables the model to align the tone-octave information between source and target domains for better feature distribution. Then, we leverage the learned domain-agnostic fine-grained features as additional information to obtain domain-agnostic holistic features. We also suggest to align intra-domain, inter-domain, and sample-level features to further improve the performances. In the second stage, we propose a novel tone-octave consistency regularization method by leveraging the extracted fine-grained information to judge the availability of unlabeled data. We evaluate our proposed framework on several well-known public datasets, and the conducted experiments demonstrate the effectiveness of our method.
Shuai Yu 0002, Xiaoliang He, Kangjie Dong, Yi Yu 0001
ACM Multimedia2
2025 Multi-Source Domain Generalization for Machine Remaining Useful Life Prediction via Risk Minimization-Based Test-Time Adaptation
abstract
Due to unnecessary access to the target data during training, domain generalization (DG) has received great attention in remaining useful life (RUL) prediction for rotating machines. However, existing methods often fail to estimate the ubiquitous adaptation gap, which intractably minimizes the generalization risk. In this study, a novel multi-source DG method is proposed for cross-domain RUL prediction, which considers adaptation gap and performs test-time adaptation to minimize the risk of generalization. Initially, multi-head domain-specific regressors are pretrained to learn the hypothesis from multi-source domains separately. Afterward, the test-time model selection and ensemble is utilized to collaboratively minimize adaptation gap, wherein two strategies of domain similarity and predictive indicator are presented to dynamically integrate the optimal regressor adapted to target domain. Meanwhile, the multioutputs integrated pseudo-labels are used to retrain and optimize the model. Experimental studies indicate that the proposed approach is promising with a maximum 13.05% improvement on prediction performance.
Chun Su, Xiaoliang He, Mingjiang Xie, Zhigang Tian
IEEE Trans. Ind. Informatics3
2024 RevNet: A Review Network with Group Aggregation Fusion for Singing Melody Extraction
abstract
Singing melody extraction (SME) is a critical task in the field of music information retrieval (MIR). Recently, deep learning based methods have achieved remarkable successes for singing melody extraction. However, most of the existing models are based on stacked convolution layers to progressively obtain task-specific features. Such an architecture has two limitations: 1) in the training stage, when the global semantic feature is obtained, the global semantic feature will be directly used to make predictions. There is a lack of a process that makes SME models learn knowledge from training errors in an in-depth way. 2) there exist semantic gaps between features from different levels in the prior SME models. The inconsistent features from different levels may cause suboptimal performances. To address the above mentioned problem, in this paper, we propose a review network (RevNet) with group aggregation fusion for singing melody extraction. Specifically, the proposed network is based on an encoder-decoder network, which consists of two modules: review module and group aggregation fusion (GAF) module. The review module aims to make the model be able to learn training errors interactively. We design multiple review modules to iterately review the training errors. The design of this module is like a review process to force the model to learn knowledge from prior prediction errors. The GAF module aims to fuse the features from different levels and makes multi-level features complementary. A set of dilated convolution operations are performed on our designed grouped high-level and low-level features. Moreover, to explicitly eliminate the difference between multi-level features, the feature maps from different levels are allowed to be directly supervised by the ground truth. We conduct experiments on several public datasets and the promising results demonstrate the effectiveness of our proposed method.
Shuai Yu 0002, Xiaoliang He
ICME2
2024 HKDSME: Heterogeneous Knowledge Distillation for Semi-supervised Singing Melody Extraction Using Harmonic Supervision
abstract
Singing melody extraction is a key task in the field of music information retrieval (MIR). However, decades of research works have uncovered two difficult issues. First, binary classification on frequency-domain audio features (e.g., spectrogram) is regarded as the primary method, which ignores the potential associations of musical information at different frequency bins, as well as their varying significance for output decisions. Second, the existing semi-supervised singing melody extraction models ignore the accuracy of the generated pseudo labels by semi-supervised models, which largely limits the further improvements of the model. To solve the two issues, in this paper, we propose a heterogeneous knowledge distillation framework for semi-supervised singing melody extraction using harmonic supervision, termed as HKDSME. We begin by proposing a four-class classification paradigm for determining the results of singing melody extraction using harmonic supervision. This enables the model to capture more information regarding melodic relations in spectrograms. To improve the accuracy issue of pseudo labels, we then build a semi-supervised method by leveraging the extracted harmonics as a consistent regularization. Different from previous methods, it judges the availability of unlabeled data in terms of the inner positional relations of extracted harmonics. To further build a light-weight semi-supervised model, we propose a heterogeneous knowledge distillation (HKD) module, which enables the prior knowledge to transfer between heterogeneous models. We also propose a novel confidence guided loss, which incorporates with the proposed HKD module to reduce the wrong pseudo labels. We evaluate our proposed method using several well-known public available datasets, and the findings demonstrate the efficacy of our proposed method.
Shuai Yu 0002, Xiaoliang He, Ke Chen 0021, Yi Yu 0001
ACM Multimedia2