Xiaoheng Sun

dblp:272/6462 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-0711-1577ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Decomposed Retrieval-Edit-Rerank Framework for Chord Generation
abstract
Chord generation is an inherently constrained creative task that requires balancing stylistic diversity with music-theoretic feasibility. Existing approaches typically entangle candidate generation and constraint enforcement within a single model, making the diversity–feasibility trade-off difficult to control and interpret.
Qiqi He, Dichucheng Li, Xiaoheng Sun
ICMR3
2023 Tg-Critic: A Timbre-Guided Model For Reference-Independent Singing Evaluation
abstract
Automatic singing evaluation independent of reference melody is a challenging task due to its subjective and multi-dimensional nature. As an essential attribute of singing voices, vocal timbre has a non-negligible effect and influence on human perception of singing quality. However, no research has been done to include timbre information explicitly in singing evaluation models. In this paper, a data-driven model TG-Critic is proposed to introduce timbre embeddings as one of the model inputs to guide the evaluation of singing quality. The trunk structure of TG-Critic is designed as a multi-scale network to summarize the contextual information from constant-Q transform features in a high-resolution way. Furthermore, an automatic annotation method is designed to construct a large three-class singing evaluation dataset with low human-effort. The experimental results show that the proposed model outperforms the existing state-of-the-art models in most cases.
Xiaoheng Sun, Yuejie Gao, Hanyao Lin
ICASSP1
2023 A neural harmonic-aware network with gated attentive fusion for singing melody extraction
Shuai Yu 0002, Yi Yu 0001, Xiaoheng Sun, Wei Li 0012
Neurocomputing3
2022 Deepchorus: A Hybrid Model of Multi-Scale Convolution And Self-Attention for Chorus Detection
abstract
Chorus detection is a challenging problem in musical signal processing as the chorus often repeats more than once in popular songs, usually with rich instruments and complex rhythm forms. Most of the existing works focus on the receptiveness of chorus sections based on some explicit features such as loudness and occurrence frequency. These pre-assumptions for chorus limit the generalization capacity of these methods, causing misdetection on other repeated sections such as verse. To solve the problem, in this paper we propose an end-to-end chorus detection model DeepChorus, reducing the engineering effort and the need for prior knowledge. The proposed model includes two main structures: i) a Multi-Scale Network to derive preliminary representations of chorus segments, and ii) a Self-Attention Convolution Network to further process the features into probability curves representing chorus presence. To obtain the final results, we apply an adaptive threshold to binarize the original curve. The experimental results show that DeepChorus outperforms existing state-of-the-art methods in most cases.
Qiqi He, Xiaoheng Sun, Yi Yu 0001, Wei Li 0012
ICASSP2
2022 GIO: A Timbre-informed Approach for Pitch Tracking in Highly Noisy Environments
abstract
As one of the fundamental tasks in music and speech signal processing, pitch tracking has been attracting attention for decades. While a human can focus on the voiced pitch even in highly noisy environments, most existing automatic pitch tracking systems show unsatisfactory performance encountering noise. To mimic human auditory, a data-driven model named GIO is proposed in this paper, in which timbre information is introduced to guide pitch tracking. The proposed model takes two inputs: a short audio segment to extract pitch from and a timbre embedding derived from the speaker's or singer's voice. In experiments, we use a music artist classification model to extract timbre embedding vectors. A dual-branch structure and a two-step training method are designed to enable the model to predict voice presence. The experimental results show that the proposed model gains a significant improvement in noise robustness and outperforms existing state-of-the-art methods with fewer parameters.
Xiaoheng Sun, Xia Liang, Qiqi He, Bilei Zhu, Zejun Ma 0001
ICMR1
2021 An Hrnet-Blstm Model With Two-Stage Training For Singing Melody Extraction
abstract
Well-labeled datasets available for melody extraction are scarce, which limits the further advancement of deep learning based methods. To overcome this problem, we propose to use a pitch refinement method to refine the semitone-level pitch sequences decoded from massive melody MIDI files to generate a large number of fundamental frequency (F0) values for model training. Since the refined pitch values used for the first round of training contain errors, a small set of well-labeled data is used for a second round of training. A high-resolution network (HRNet), initially developed for human pose estimation, is introduced for melody extraction. It considers multi-resolution feature learning, making the resulting representation semantically richer. Subsequently, a bidirectional long short-term memory (BLSTM) layer is used to exploit the temporal information of melody. In addition, a new loss function where the unvoiced frames only contribute to voicing detection but not to pitch classification is also proposed to alleviate the class imbalance problem. Experiment results on three public datasets show that the proposed system outperforms four state-of-the-art algorithms in most cases.
Yongwei Gao, Xingjian Du, Bilei Zhu, Xiaoheng Sun, Wei Li 0012, Zejun Ma 0001
ICASSP4
2021 Frequency-Temporal Attention Network for Singing Melody Extraction
abstract
Musical audio is generally composed of three physical properties: frequency, time and magnitude. Interestingly, human auditory periphery also provides neural codes for each of these dimensions to perceive music. Inspired by these intrinsic characteristics, a frequency-temporal attention network is proposed to mimic human auditory for singing melody extraction. In particular, the proposed model contains frequency-temporal attention modules and a selective fusion module corresponding to these three physical properties. The frequency attention module is used to select the same activation frequency bands as did in cochlear and the temporal attention module is responsible for analyzing temporal patterns. Finally, the selective fusion module is suggested to recalibrate magnitudes and fuse the raw information for prediction. In addition, we propose to use another branch to simultaneously predict the presence of singing voice melody. The experimental results show that the proposed model outperforms existing state-of-the-art methods1.
Shuai Yu 0002, Xiaoheng Sun, Yi Yu 0001, Wei Li 0012
ICASSP2
2020 Residual Attention Based Network for Automatic Classification of Phonation Modes
abstract
Phonation mode is an essential characteristic of singing style as well as an important expression of performance. It can be classified into four categories, called neutral, breathy, pressed and flow. Previous studies used voice quality features and feature engineering for classification. While deep learning has achieved significant progress in other fields of music information retrieval (MIR), there are few attempts in the classification of phonation modes. In this study, a Residual Attention based network is proposed for automatic classification of phonation modes. The network consists of a convolutional network performing feature processing and a soft mask branch enabling the network focus on a specific area. In comparison experiments, the models with proposed network achieve better results in three of the four datasets than previous works, among which the highest classification accuracy is 94.58%, 2.29% higher than the baseline.
Xiaoheng Sun, Yiliang Jiang, Wei Li 0012
ICME1