Shuai Yu 0002

dblp:120/8707-2 · DBLP profile ↗
← Back
19ranked-venue papers
14as first author
15since 2021 · last 2026
0000-0003-1847-563XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 10 first-author · 13 since 2021Artificial intelligence and machine learning · 8 · 6 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Every Little Bit Helps: Exploring Better Utilization of Unlabeled Data for Semi-supervised Singing Melody Extraction Using Multi-bands Diffusion Model
abstract
Semi-supervised singing melody extraction (SSME) is one of the key tasks in the field of music information retrieval (MIR). Recently, several SSME methods have been proposed and achieved remarkable successes. However, existing methods are still facing two critical issues: firstly, there is a lack of an effective data augmentation method for SSME, which results in insufficient utilization of unlabeled data. Secondly, existing SSME methods discards too much unlabeled data in the stage of consistency regularization, which hinders the further improvements of SSME task. In this paper, we present \emph{ELH-SME}, a novel framework that better utilizes the unlabeled musical data for SSME task. Specifically, our proposed ELH-SME framework consists of three modules: (1) we first propose a diffusion-based multi-bands augmentation (DMA) method to increase the amounts of training data. The proposed DMA methods employs a diffusion model to generate perturbation at the specific frequency bands in an end-to-end manner, thereby avoiding sharply perturbations to the spectrogram. (2) To improve the utilization rate of unlabeled data, we suggest a global-class confidence (GCC) module. During the phase of consistency regularization, we consider both the global-wise and class-wise confidence values, improving the utilization rate of unlabeled data. (3) To further improve the utilization of unlabeled data, we also propose to enhance the representation capability of unlabeled data by extracting channel-level features from labeled data via channel cross attention (CCA). We evaluate our proposed framework on several well-known public available datasets, and the conducted experiments demonstrate the effectiveness of our method.
Shuai Yu 0002, Xiaoliang He, Kangjie Dong, Yi Yu 0001
AAAI1
2025 A Mamba-based Network for Semi-supervised Singing Melody Extraction Using Confidence Binary Regularization
abstract
Singing melody extraction (SME) is a key task in the field of music information retrieval. However, existing methods are facing several limitations: firstly, prior models use transformers to capture the contextual dependencies, which requires quadratic computation resulting in low efficiency in the inference stage. Secondly, prior works typically rely on frequency-supervised methods to estimate the fundamental frequency (f0), which ignores that the musical performance is actually based on notes. Thirdly, transformers typically require large amounts of labeled data to achieve optimal performances, but the SME task lacks of sufficient annotated data. To address these issues, in this paper, we propose a mamba-based network, called SpectMamba, for semi-supervised singing melody extraction using confidence binary regularization. In particular, we begin by introducing vision mamba to achieve computational linear complexity. Then, we propose a novel note-f0 decoder that allows the model to better mimic the musical performance. Further, to alleviate the scarcity of the labeled data, we introduce a confidence binary regularization (CBR) module to leverage the unlabeled data by maximizing the probability of the correct classes. The proposed method is evaluated on several public datasets and the conducted experiments demonstrate the effectiveness of our proposed method1.
Xiaoliang He, Kangjie Dong, Jingkai Cao, Shuai Yu 0002, Wei Li 0012, Yi Yu 0001
ICASSP4
2025 Ultra Lightweight Singing Melody Extraction via Combination of Convolution and MLP
abstract
Singing melody extraction serves as an important foundation in the realm of music information retrieval (MIR). Although fully convolutional neural networks (CNNs) are commonly employed for singing melody extraction, they are constrained by inductive biases and face challenges in establishing long range dependency. Transformer-based networks have better performance, but the computational load is high. Recently, many multi-layer perceptron (MLP) architectures have been applied for a variety of computer vision tasks, demonstrating competitive performance. However, its potential ability in the task of singing melody extraction remains to be further explored. In this paper, we propose the lightweight convolutional MLP (LcMLP), an ultra lightweight model without sacrificing the performance. Firstly, we improve the original MLP-Mixer. We change the sequential MLPs to parallel ones and add some skip connections. Secondly, we propose a multi-level convolution fusion module that facilitates the interaction of features at various depths in MLP-Mixer. We conducted extensive experiments on several well-known public datasets, and our model demonstrates significant advantages in inference speed and computational load, while also achieving competitive performance.
Kangjie Dong, Qiubo Huang, Shuai Yu 0002, Wei Li 0012
ICASSP4
2025 DUDA: A Two-stage Decoupling Unsupervised Domain Adaptation Framework for Semi-supervised Singing Melody Extraction from Polyphonic Music
abstract
Semi-supervised singing melody extraction (SSME) is one of the key tasks in the field of music information retrieval (MIR). However, there are two critical issues that remain to be addressed in data limited scenarios. Firstly, the prior unsupervised domain adaptation methods for SSME typically rely on learning domain-agnostic features at holistic level, which ignores the associations between holistic information (i.e., fundamental frequency) and fine-grained information (i.e., tone and octave). Secondly, the fine-grained information can be utilized to judge the availability of unlabeled data, which is ignored by prior methods. There is a lack of a consistency regularization method that utilizes fine-grained information to validate the availability of unlabeled data. To address these issues, in this paper, we propose a novel two-stage decoupling unsupervised domain adaptation framework for semi-supervised singing melody extraction, termed as DUDA. Specifically, in the first stage, we decouple the holistic information into fine-grained information: tone and octave, and narrow the domain gap at the tone and octave level, respectively. This enables the model to align the tone-octave information between source and target domains for better feature distribution. Then, we leverage the learned domain-agnostic fine-grained features as additional information to obtain domain-agnostic holistic features. We also suggest to align intra-domain, inter-domain, and sample-level features to further improve the performances. In the second stage, we propose a novel tone-octave consistency regularization method by leveraging the extracted fine-grained information to judge the availability of unlabeled data. We evaluate our proposed framework on several well-known public datasets, and the conducted experiments demonstrate the effectiveness of our method.
Shuai Yu 0002, Xiaoliang He, Kangjie Dong, Yi Yu 0001
ACM Multimedia1
2024 MCSSME: Multi-Task Contrastive Learning for Semi-supervised Singing Melody Extraction from Polyphonic Music
abstract
Singing melody extraction is an important task in the field of music information retrieval (MIR). The development of data-driven models for this task have achieved great successes. However, the existing models have two major limitations: firstly, most of the existing singing melody extraction models have formulated this task as a pixel-level prediction task. The lack of labeling data has limited the model for further improvements. Secondly, the generalization of the existing models are prone to be disturbed by the music genres. To address the issues mentioned above, in this paper, we propose a multi-Task contrastive learning framework for semi-supervised singing melody extraction, termed as MCSSME. Specifically, to deal with data scarcity limitation, we propose a self-consistency regularization (SCR) method to train the model on the unlabeled data. Transformations are applied to the raw signal of polyphonic music, which makes the network to improve its representation capability via recognizing the transformations. We further propose a novel multi-task learning (MTL) approach to jointly learn singing melody extraction and classification of transformed data. To deal with generalization limitation, we also propose a contrastive embedding learning, which strengthens the intra-class compactness and inter-class separability. To improve the generalization on different music genres, we also propose a domain classification method to learn task-dependent features by mapping data from different music genres to shared subspace. MCSSME evaluates on a set of well-known public melody extraction datasets with promising performances. The experimental results demonstrate the effectiveness of the MCSSME framework for singing melody extraction from polyphonic music using very limited labeled data scenarios.
Shuai Yu 0002
AAAI1
2024 A Scalable Sparse Transformer Model for Singing Melody Extraction
abstract
Extracting the melody of a singing voice is an essential task within the realm of music information retrieval (MIR). Recently, transformer based models have drawn great attention in the field of MIR. However, due to the expensive computation cost and extensive parameters, it is difficult to train and deploy a transformer-based model for practical singing melody extraction. In this paper, we propose a simple yet effective scalable sparse transformer for singing melody extraction. To be specific, we first propose to employ a sparse transformer to reduce computation cost and the amount of parameters. Then, we proposed to scale the self-attention region of the sparse transformer in the spectrogram to obtain more accurate performance. Moreover, we propose to combine a scalable sparse transformer (S2Former) with CNN-based model to extract global and local features in the spectrogram. The proposed scalable transformer model can achieve a better balance between a standard transformer and a sparse transformer. To better fuse the features from transformer and CNN, we further propose a transformer-CNN fusion (TCF) module to combine significant features from transformer and CNN. The proposed model obtains state-of-the-art results on several public datasets. The conducted experiments confirm the effectiveness of the model we proposed.
Shuai Yu 0002, Yi Yu 0001, Wei Li 0012
ICASSP1
2024 RevNet: A Review Network with Group Aggregation Fusion for Singing Melody Extraction
abstract
Singing melody extraction (SME) is a critical task in the field of music information retrieval (MIR). Recently, deep learning based methods have achieved remarkable successes for singing melody extraction. However, most of the existing models are based on stacked convolution layers to progressively obtain task-specific features. Such an architecture has two limitations: 1) in the training stage, when the global semantic feature is obtained, the global semantic feature will be directly used to make predictions. There is a lack of a process that makes SME models learn knowledge from training errors in an in-depth way. 2) there exist semantic gaps between features from different levels in the prior SME models. The inconsistent features from different levels may cause suboptimal performances. To address the above mentioned problem, in this paper, we propose a review network (RevNet) with group aggregation fusion for singing melody extraction. Specifically, the proposed network is based on an encoder-decoder network, which consists of two modules: review module and group aggregation fusion (GAF) module. The review module aims to make the model be able to learn training errors interactively. We design multiple review modules to iterately review the training errors. The design of this module is like a review process to force the model to learn knowledge from prior prediction errors. The GAF module aims to fuse the features from different levels and makes multi-level features complementary. A set of dilated convolution operations are performed on our designed grouped high-level and low-level features. Moreover, to explicitly eliminate the difference between multi-level features, the feature maps from different levels are allowed to be directly supervised by the ground truth. We conduct experiments on several public datasets and the promising results demonstrate the effectiveness of our proposed method.
Shuai Yu 0002, Xiaoliang He
ICME1
2024 Is Translation Helpful? An Exploration of Cross-Lingual Transfer in Low-Resource Dialog Generation
abstract
Cross-lingual transfer is important for developing high-quality chatbots in multiple languages to address the imbalanced distribution of language resources. A typical approach of cross-lingual transfer, which has been proved effective on classification tasks, is to leverage machine translation (MT) systems to utilize either the training corpus or models from high-resource languages. In this work, we investigate whether it is helpful to utilize MT for cross-lingual transfer in dialog generation tasks. We collect a benchmark dataset for the low-resource scenario, assuming access to limited Chinese dialog data in the movie domain and large amounts of English dialog from multiple domains. Experiments show that leveraging English dialog corpora can improve the naturalness, relevance, and cross-domain transferability in Chinese. However, directly using English corpora in their original form is better than translating them into Chinese. As the topics and wording habits in dialogs are strongly culture-dependent, translating them can reinforce the bias from high-resource languages. To avoid this issue and also reduce the embedding mismatch, we propose to use embedding freezing and post-alignment, which align words from different languages as translation would, but without introducing translation biases. Experiments show that embedding freezing and post-alignment can further improve generation performance. The results of the analysis together with the collected benchmark dataset are presented to draw attention to this area and support future research.
Lei Shen 0002, Shuai Yu 0002
IJCNN2
2024 HKDSME: Heterogeneous Knowledge Distillation for Semi-supervised Singing Melody Extraction Using Harmonic Supervision
abstract
Singing melody extraction is a key task in the field of music information retrieval (MIR). However, decades of research works have uncovered two difficult issues. First, binary classification on frequency-domain audio features (e.g., spectrogram) is regarded as the primary method, which ignores the potential associations of musical information at different frequency bins, as well as their varying significance for output decisions. Second, the existing semi-supervised singing melody extraction models ignore the accuracy of the generated pseudo labels by semi-supervised models, which largely limits the further improvements of the model. To solve the two issues, in this paper, we propose a heterogeneous knowledge distillation framework for semi-supervised singing melody extraction using harmonic supervision, termed as HKDSME. We begin by proposing a four-class classification paradigm for determining the results of singing melody extraction using harmonic supervision. This enables the model to capture more information regarding melodic relations in spectrograms. To improve the accuracy issue of pseudo labels, we then build a semi-supervised method by leveraging the extracted harmonics as a consistent regularization. Different from previous methods, it judges the availability of unlabeled data in terms of the inner positional relations of extracted harmonics. To further build a light-weight semi-supervised model, we propose a heterogeneous knowledge distillation (HKD) module, which enables the prior knowledge to transfer between heterogeneous models. We also propose a novel confidence guided loss, which incorporates with the proposed HKD module to reduce the wrong pseudo labels. We evaluate our proposed method using several well-known public available datasets, and the findings demonstrate the efficacy of our proposed method.
Shuai Yu 0002, Xiaoliang He, Ke Chen 0021, Yi Yu 0001
ACM Multimedia1
2023 A neural harmonic-aware network with gated attentive fusion for singing melody extraction
Shuai Yu 0002, Yi Yu 0001, Xiaoheng Sun, Wei Li 0012
Neurocomputing1
2022 Tonet: Tone-Octave Network for Singing Melody Extraction from Polyphonic Music
abstract
Singing melody extraction is an important problem in the field of music information retrieval. Existing methods typically rely on frequency-domain representations to estimate the sung frequencies. However, this design does not lead to human-level performance in the perception of melody information for both tone (pitch-class) and octave. In this paper, we propose TONet1, a plug-and-play model that improves both tone and octave perceptions by leveraging a novel input representation and a novel network architecture. First, we present an improved input representation, the Tone-CFP, that explicitly groups harmonics via a rearrangement of frequency-bins. Second, we introduce an encoder-decoder architecture that is designed to obtain a salience feature map, a tone feature map, and an octave feature map. Third, we propose a tone-octave fusion mechanism to improve the final salience feature map. Experiments are done to verify the capability of TONet with various baseline backbone models. Our results show that tone-octave fusion with Tone-CFP can significantly improve the singing voice extraction performance across various datasets – with substantial gains in octave and tone accuracy.
Ke Chen 0021, Shuai Yu 0002, Cheng-i Wang, Wei Li 0012, Taylor Berg-Kirkpatrick, Shlomo Dubnov
ICASSP2
2022 Hierarchical Graph-Based Neural Network for Singing Melody Extraction
abstract
Singing melody extraction from polyphonic music is a critical and challenging task in music information retrieval (MIR). However, due to the interfere of the accompaniment and the background noise, it is key and challenging to obtain a global semantic representation that discriminates the singing melody line. To address this issue, we consider the two aspects that regards to obtaining the global semantic representation: the global relationships in the spectrum and the relationships between channels. In this paper, we propose a novel hierarchical graph-based network for singing melody extraction. In particular, according to its characteristics of the spectrum, we first model the spectrum into graph structure, a two-layer graph convolution network is used to obtain the global semantic representation in the spectrum. Then to capture the relationships between channels, channel-wise graph convolution module is devised to capture and reasoning the relationship between channels. The conducted experiments demonstrate the effectiveness of the proposed network.
Shuai Yu 0002, Wei Li 0012
ICASSP1
2022 A Glance-and-Gaze Network for Respiratory Sound Classification
abstract
A plethora of great successes has been achieved by the existing convolutional neural networks (CNN) for respiratory sound classification. Nevertheless, simultaneously capturing both the local and global features can never be an easy task due to the limitation of a CNN’s structure. In this contribution, we propose a novel glance-and-gaze network to address the aforementioned issue. The glance block aims to learn global information, while the gaze block is responsible for learning local patterns and suppressing the noises that attenuates the final performance. In the proposed method, both the global and local information can be extracted. Moreover, the spectral and temporal representations can be learnt via a feature fusion module. Experimental results on the largest public respiratory sound database demonstrate that the proposed model outperforms the state-of-the-art methods.
Shuai Yu 0002, Yiwei Ding, Kun Qian 0003, Bin Hu 0001, Wei Li 0012, Björn W. Schuller
ICASSP1
2021 Frequency-Temporal Attention Network for Singing Melody Extraction
abstract
Musical audio is generally composed of three physical properties: frequency, time and magnitude. Interestingly, human auditory periphery also provides neural codes for each of these dimensions to perceive music. Inspired by these intrinsic characteristics, a frequency-temporal attention network is proposed to mimic human auditory for singing melody extraction. In particular, the proposed model contains frequency-temporal attention modules and a selective fusion module corresponding to these three physical properties. The frequency attention module is used to select the same activation frequency bands as did in cochlear and the temporal attention module is responsible for analyzing temporal patterns. Finally, the selective fusion module is suggested to recalibrate magnitudes and fuse the raw information for prediction. In addition, we propose to use another branch to simultaneously predict the presence of singing voice melody. The experimental results show that the proposed model outperforms existing state-of-the-art methods1.
Shuai Yu 0002, Xiaoheng Sun, Yi Yu 0001, Wei Li 0012
ICASSP1
2021 HANME: Hierarchical Attention Network for Singing Melody Extraction
abstract
Singing melody extraction in polyphonic musical audio is a very critical and challenging task in music information retrieval (MIR). Contextual frame-level information has proven its effectiveness in this task. However, existing works assign equal weight to each contextual frame, which may hinder the further improvement of the performance. To this end, we propose a hierarchical attention network for singing melody extraction (HANME) to extract the discriminative attention-aware features and alleviate the workload of the convolutional recurrent neural network (CRNN) for extracting local spatial and temporal features. Specifically, the first attention layer learns the context vector based on local spatial features extracted by residual convolutional neural network (CNN), and the second attention layer learns the temporal context vector based on long-term features extracted by Bidirectional Gated Recurrent Units (BiGRU). Due to the scarcity of labeled training data, we further propose a partial parameter adaptation approach to address the imbalance distribution of the labels for this task. We use the RWC dataset and part of vocal tracks of the MedleyDB dataset for training the model and evaluate the performance on the ADC2004, MIREX 05 and MedleyDB datasets. The experimental study demonstrates the superiority of our method compared with other state-of-the-art ones.
Shuai Yu 0002, Yi Yu 0001, Wei Li 0012
IEEE Signal Process. Lett.1
2019 WAIS: Word Attention for Joint Intent Detection and Slot Filling
abstract
Attention-based recurrent neural network models for joint intent detection and slot filling have achieved a state-of-the-art performance. Most previous works exploited semantic level information to calculate the attention weights. However, few works have taken the importance of word level information into consideration. In this paper, we propose WAIS, word attention for joint intent detection and slot filling. Considering that intent detection and slot filling have a strong relationship, we further propose a fusion gate that integrates the word level information and semantic level information together for jointly training the two tasks. Extensive experiments show that the proposed model has robust superiority over its competitors and sets the state-of-the-art.
Sixuan Chen, Shuai Yu 0002
AAAI2
2019 NAIRS: A Neural Attentive Interpretable Recommendation System
abstract
In this paper, we develop a neural attentive interpretable recommendation system, named NAIRS. A self-attention network, as a key component of the system, is designed to assign attention weights to interacted items of a user. This attention mechanism can distinguish the importance of the various interacted items in contributing to a user profile. %, and it also provides interpretable recommendations. Based on the user profiles obtained by the self-attention network, NAIRS offers personalized high-quality recommendation. Moreover, it develops visual cues to interpret recommendations. This demo application with the implementation of NAIRS enables users to interact with a recommendation system, and it persistently collects training data to improve the system. The demonstration and experimental results show the effectiveness of NAIRS.
Shuai Yu 0002, Min Yang 0007, Baocheng Li, Qiang Qu 0001, Jialie Shen 0001
WSDM1
2019 Contextual-boosted deep neural collaborative filtering model for interpretable recommendation
Shuai Yu 0002, Min Yang 0007, Qiang Qu 0001, Ying Shen 0001
Expert Syst. Appl.1
2018 ACJIS: A Novel Attentive Cross Approach For Joint Intent Detection And Slot Filling
abstract
Intent detection and slot filling are two important tasks in Spoken Language Understanding. The Condition Random Fields (CRF) was introduced for the tasks pretty much the same fashion to deep neural networks. Recently, attention based encoder-decoder models have shown promising results for joint intent detection and slot filling tasks in spoken language understanding and dialog systems. However, the two tasks are often trained separately. In this paper, we propose ACJIS, a novel Attentive Cross approach for Joint Intent detection and Slot filling. We introduce a cross attention approach to enhance the modeling power on capturing the meaning of word at both tagging level and word level. In order to utilize the information from the two tasks, we leverage multi-task learning to train the model. Our model generates state-of-the-art results on the bench-mark ATIS task. The proposed model also achieves significant gains over the attention based RNN modeling approach for intent detection and slot filling respectively.
Shuai Yu 0002, Lei Shen 0002, Jiansong Chen
IJCNN1