Xianke Wang

dblp:297/7999 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2024
0000-0001-7037-9800ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2024 Piano Transcription with Harmonic Attention
abstract
Automatic Music Transcription (AMT) aims to convert music audio into digital sheet music. Piano transcription is a popular but challenging subtask of AMT. For every piano pitch, the harmonic structure is fixed in the frequency domain, while the Transformer based on self-attention has great potential to extract features in the long sequence. In this paper, we propose piano harmonic attention, a mask self-attention, for better capturing harmonic features. The mask matrix is designed with the harmonic prior to pre-modeling the harmonic structure during calculating attention scores. To verify its effectiveness, we append the harmonic attention-based Transformer after every convolutional neural network block of the High-resolution piano transcription system. The evaluation results on the MAESTRO dataset show that the proposed model achieves comprehensive improvements over the baseline, with a note F1 score of 97.33%, which is comparable to the state-of-the-art system.
Ruimin Wu, Xianke Wang, Wei Xu 0038, Wenqing Cheng
ICASSP2
2024 CNN-Transformer Ensemble: Advancing Visual Piano Transcription with Global and Local Features
abstract
Piano transcription is a significant problem in music information retrieval, which aims to infer the note sequence from recorded music signals. Recently, visual piano transcription heavily relies on Convolutional Neural Networks (CNNs) for local feature detection, facing challenges in capturing global features of the piano keyboard. Therefore, we propose a novel visual piano transcription model that integrates CNN and transformer branches. The CNN branch has inductive bias and extracts local features of the piano key, while the transformer branch aggregates global features over the keyboard. The proposed model integrates local features and global features at different resolutions, enhancing the performance of the visual transcription model. Finally, our model achieves a remarkable F1-score of 92.35% on the OMAPS2 dataset and attains state-of-the-art results on other datasets. This substantiates the model’s innovative approach and its potential to advance visual piano transcription within music information retrieval.
Xianke Wang, Ruimin Wu, Wei Xu 0038
IJCNN3
2024 A Two-Stage Audio-Visual Fusion Piano Transcription Model Based on the Attention Mechanism
abstract
Piano transcription is a significant problem in the field of music information retrieval, aiming to obtain symbolic representations of music from captured audio or visual signals. Previous research has mainly focused on single-modal transcription methods using either audio or visual information, yet there is a small number of studies based on audio-visual fusion. To leverage the complementary advantages of both modalities and achieve higher transcription accuracy, we propose a two-stage audio-visual fusion piano transcription model based on the attention mechanism, utilizing both audio and visual information from the piano performance. In the first stage, we propose an audio model and a visual model. The audio model utilizes frequency domain sparse attention to capture harmonic relationships in the frequency domain, while the visual model includes both CNN and Transformer branches to merge local and global features at different resolutions. In the second stage, we employ cross-attention to learn the correlations between different modalities and the temporal relationships of the sequences. Experimental results on the OMAPS2 dataset show that our model achieves an F1-score of 98.60%, demonstrating significant improvement compared with the single-modal transcription models.
Xianke Wang, Ruimin Wu, Wei Xu 0038, Wenqing Cheng
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 A Two-stage Progressive Neural Network for Acoustic Echo Cancellation
abstract
Recent studies in deep learning based acoustic echo cancellation proves the benefits of introducing a linear echo cancellation module. However, the convergence problem and potential target speech distortion impose an additional learning burden for the neural network. In this paper, we propose a two-stage progressive neural network consisting of a coarse-stage and a fine-stage module. For the coarse-stage, a light-weighted network module is designed to suppress partial echo and potential noise, where a voice activity detection path is used to enhance the learned features. For the fine-stage, a larger network is employed to deal with the more complex echo path and restore the near-end speech. We have conducted extensive experiments to verify the proposed method, and the results show that the proposed two-stage method provides a superior performance to other state-of-the-art methods.
Zhuangqi Chen, Xianjun Xia, Xianke Wang, Yanhong Leng, Roberto Togneri, Yijian Xiao, Piao Ding, Shenyi Song, Pingjian Zhang
INTERSPEECH4
2023 MusicYOLO: A Vision-Based Framework for Automatic Singing Transcription
abstract
Automatic singing transcription (AST), which refers to the process of inferring the onset, offset, and pitch from the singing audio, is of great significance in music information retrieval. Most AST models use the convolutional neural network to extract spectral features and predict the onset and offset moments separately. The frame-level probabilities are inferred first, and then the note-level transcription results are obtained through post-processing. In this paper, a new AST framework called MusicYOLO is proposed, which obtains the note-level transcription results directly. The onset/offset detection is based on the object detection model YOLOX, and the pitch labeling is completed by a spectrogram peak search. Compared with previous methods, the MusicYOLO detects note objects rather than isolated onset/offset moments, thus greatly enhancing the transcription performance. On the sight-singing vocal dataset (SSVD) established in this paper, the MusicYOLO achieves an 84.60% transcription F1-score, which is the state-of-the-art method.
Xianke Wang, Wei Xu 0038, Wenqing Cheng
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 A Multi-Stage Automatic Evaluation System for Sight-Singing
abstract
Sight-singing exercises are a fundamental part of music education. In this paper, we present an objective and complete automatic evaluation system for sight-singing, which has two critical stages: note transcription and note alignment. In the first stage, we use an onset detector based on the convolutional recurrent neural network (CRNN) for note segmentation and the pitch extractor described in (Kimet al.2018) for note labeling. In the second stage, an alignment algorithm based on relative pitch modeling is proposed. Due to the lack of datasets for sight-singing note alignment and the overall system evaluation, we construct the sight-singing vocal dataset (SSVD). Each module of the system and the entire system are tested on this dataset. The onset detector achieves an F-measure of 90.61%, and the stages of note transcription and note alignment achieve an F-measure of 88.42% and 94.79%, respectively. In addition, we propose an objective criterion for the sight-singing evaluation system. Based on this criterion, our automatic sight-singing system achieves an F-measure of 77.95% on the SSVD dataset.
Xianke Wang, Wei Xu 0038, Wenqing Cheng
IEEE Trans. Multim.2
2022 Musicyolo: A Sight-Singing Onset/Offset Detection Framework Based on Object Detection Instead of Spectrum Frames
abstract
In this paper, we propose MusicYOLO based on object detection to detect the onset and offset in singing for the first time. The onset of the vocal is not as stable and clear as that of musical instruments, which makes the frame-based onset/offset detection methods often not work well. Compared with the previous onset/offset detection methods, MusicYOLO detects the whole note object in the spectrogram image instead of transient frame features around onset/offset, improving the onset/offset detection performance significantly. The experiment results show that the MusicYOLO framework has obtained a 94.16% F1 score of onset detection and a 91.35% F1 score of offset detection on the ISMIR2014 dataset, which proves that MusicYOLO is the state-of-the-art onset/offset detection framework for singing situation.
Xianke Wang, Wei Xu 0038, Wenqing Cheng
ICASSP1
2022 SingMaster: A Sight-singing Evaluation System of "Shoot and Sing" Based on Smartphone
abstract
Based on the smart phone, this paper integrates OMR (Optical Music Recognition) with sight-singing evaluation, and develops a "shoot and sing" practice APP called SingMaster. This system is mainly composed of three modules: OMR, evaluation and user interface. The OMR module converts the score photographed in the real scene into a note reference sequence. The sight-sing evaluation module first completes the note transcription of the sound spectrum through onset detection and pitch extraction, then aligns the transcribed note sequence with the reference sequence, and performs the evaluation. Finally, the evaluation results are visually fed back to the practitioners through the user interface module. It can provide guidance for practitioners at any time, any place and on any score instead of a real teacher.
Wei Xu 0038, Lijie Luo, Xianke Wang
ACM Multimedia5
2021 Transition-Aware: A More Robust Approach for Piano Transcription
abstract
Piano transcription is a classic problem in music information retrieval. More and more transcription methods based on deep learning have been proposed in recent years. In 2019, Google Brain published a larger piano transcription dataset, MAESTRO. On this dataset, Onsets and Frames transcription approach proposed by Hawthorne achieved a stunning onset F1 score of 94.73%. Unlike the annotation method of Onsets and Frames, Transition-aware model presented in this paper annotates the attack process of piano signals called atack transition in multiple frames, instead of only marking the onset frame. In this way, the piano signals around onset time are taken into account, enabling the detection of piano onset more stable and robust. Transition-aware achieves a higher transcription F1 score than Onsets and Frames on MAESTRO dataset and MAPS dataset, reducing many extra note detection errors. This indicates that Transition-aware approach has better generalization ability on different datasets.
Xianke Wang, Wei Xu 0038, Juanting Liu, Wenqing Cheng
DAFx1
2021 An Audio-Visual Fusion Piano Transcription Approach Based on Strategy
abstract
Piano transcription is a fundamental problem in the field of music information retrieval. At present, a large number of transcriptional studies are mainly based on audio or video, yet there is a small number of discussion based on audio-visual fusion. In this paper, a piano transcription model based on strategy fusion is proposed, in which the transcription results of the video model are used to assist audio transcription. Due to the lack of datasets currently used for audio-visual fusion, the OMAPS data set is proposed in this paper. Meanwhile, our strategy fusion model achieves a 92.07% F1 score on OMAPS dataset. The transcription model based on feature fusion is also compared with the one based on strategy fusion. The experiment results show that the transcription model based on strategy fusion achieves better results than the one based on feature fusion.
Xianke Wang, Wei Xu 0038, Juanting Liu, Wenqing Cheng
DAFx1