EDBT 2026 Demo / reviewers in the wild / expert
Xingjian Du
dblp:220/3082
· DBLP profile ↗
13ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0001-7777-1178ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NotaGen: Advancing Musicality in Symbolic Music Generation with Large Language Model Training ParadigmsabstractWe introduce NotaGen, a symbolic music generation model aiming to explore the potential of producing high-quality classical sheet music. Inspired by the success of Large Language Models (LLMs), NotaGen adopts pre-training, fine-tuning, and reinforcement learning paradigms (henceforth referred to as the LLM training paradigms). It is pre-trained on 1.6M pieces of music in ABC notation, and then fine-tuned on approximately 9K high-quality classical compositions conditioned on "period-composer-instrumentation" prompts. For reinforcement learning, we propose the CLaMP-DPO method, which further enhances generation quality and controllability without requiring human annotations or predefined rewards. Our experiments demonstrate the efficacy of CLaMP-DPO in symbolic music generation models with different architectures and encoding schemes. Furthermore, subjective A/B tests show that NotaGen outperforms baseline models against human compositions, greatly advancing musical aesthetics in symbolic music generation. Yashan Wang, Shangda Wu, Jianhuai Hu, Xingjian Du, Yueqi Peng, Yongxin Huang, Shuai Fan 0013, Feng Yu 0027, Maosong Sun 0001 |
IJCAI | 4 |
| 2024 | Joint Music and Language Attention Models for Zero-Shot Music TaggingabstractMusic tagging is a task to predict the tags of music recordings. However, previous music tagging research primarily focuses on close-set music tagging tasks which can not be generalized to new tags. In this work, we propose a zero-shot music tagging system modeled by a joint music and language attention (JMLA) model to address the open-set music tagging problem. The JMLA model consists of an audio encoder modeled by a pretrained masked autoencoder and a decoder modeled by a Falcon7B. We introduce preceiver resampler to convert arbitrary length audio into fixed length embeddings. We introduce dense attention connections between encoder and decoder layers to improve the information flow between the encoder and decoder layers. We collect a large-scale music and description dataset from the internet. We propose to use ChatGPT to convert the raw descriptions into formalized and diverse descriptions to train the JMLA models. Our proposed JMLA system achieves a zero-shot audio tagging accuracy of 64.82% on the GTZAN dataset, outperforming previous zero-shot systems and achieves comparable results to previous systems on the FMA and the MagnaTagATune datasets. Xingjian Du, Zhesong Yu, Jiaju Lin, Bilei Zhu, Qiuqiang Kong |
ICASSP | 1 |
| 2024 | ByteHum: Fast and Accurate Query-by-Humming in the WildabstractQuery by Humming (QBH) is a practically meaningful task, while most existing methods struggle to scale to real-life applications due to the complex preprocessing for building the database and the limited search speed. In this paper, we propose the ByteHum system, a fast and efficient humming retrieval system which is capable of searching against large-scale databases built on raw song audios without the need for extensive preprocessing. ByteHum employs a convolutional neural network to extract features from raw audio, and utilizes a source-separated cover song identification dataset for weakly supervised training of the feature extractor. We explore the use of unsupervised domain adaptation techniques to enhance the performance of our weakly supervised model on the QBH task. Furthermore, to evaluate QBH systems’ performance on non-manually processed databases in the wild, we annotate original recordings for three existing QBH benchmark sets. Our experimental results demonstrate that ByteHum significantly outperforms existing QBH systems in terms of speed and accuracy under both classical and unconstrained settings. Xingjian Du, Pei Zou, Xia Liang, Minghang Chu, Bilei Zhu |
ICASSP | 1 |
| 2023 | Bytecover3: Accurate Cover Song Identification On Short QueriesabstractDeep learning based methods have become a paradigm for cover song identification (CSI) in recent years, where the ByteCover systems have achieved state-of-the-art results on all the mainstream datasets of CSI. However, with the burgeon of short videos, many real-world applications require matching short music excerpts to full-length music tracks in the database, which is still under-explored and waiting for an industrial-level solution. In this paper, we upgrade the previous ByteCover systems to ByteCover3 that utilizes local features to further improve the identification performance of short music queries. ByteCover3 is designed with a local alignment loss (LAL) module and a two-stage feature retrieval pipeline, allowing the system to perform CSI in a more precise and efficient way. We evaluated ByteCover3 on multiple datasets with different benchmark settings, where ByteCover3 beat all the compared methods including its previous versions. Xingjian Du, Xia Liang, Huidong Liang, Bilei Zhu, Zejun Ma 0001 |
ICASSP | 1 |
| 2023 | Graph contrastive learning with implicit augmentations
Huidong Liang, Xingjian Du, Bilei Zhu, Zejun Ma 0001, Ke Chen 0021, Junbin Gao |
Neural Networks | 2 |
| 2022 | Zero-Shot Audio Source Separation through Query-Based Learning from Weakly-Labeled DataabstractDeep learning techniques for separating audio into different sound sources face several challenges. Standard architectures require training separate models for different types of audio sources. Although some universal separators employ a single model to target multiple sources, they have difficulty generalizing to unseen sources. In this paper, we propose a three-component pipeline to train a universal audio source separator from a large, but weakly-labeled dataset: AudioSet. First, we propose a transformer-based sound event detection system for processing weakly-labeled training data. Second, we devise a query-based audio separation model that leverages this data for model training. Third, we design a latent embedding processor to encode queries that specify audio targets for separation, allowing for zero-shot generalization. Our approach uses a single model for source separation of multiple sound types, and relies solely on weakly-labeled data for training. In addition, the proposed audio separator can be used in a zero-shot setting, learning to separate types of audio sources that were never seen in training. To evaluate the separation performance, we test our model on MUSDB18, while training on the disjoint AudioSet. We further verify the zero-shot performance by conducting another experiment on audio source types that are held-out from training. The model achieves comparable Source-to-Distortion Ratio (SDR) performance to current supervised models in both cases. Ke Chen 0021, Xingjian Du, Bilei Zhu, Zejun Ma 0001, Taylor Berg-Kirkpatrick, Shlomo Dubnov |
AAAI | 2 |
| 2022 | HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and DetectionabstractAudio classification is an important task of mapping audio samples into their corresponding labels. Recently, the transformer model with self-attention mechanisms has been adopted in this field. However, existing audio transformers require large GPU memories and long training time, meanwhile relying on pretrained vision models to achieve high performance, which limits the model’s scalability in audio tasks. To combat these problems, we introduce HTS-AT: an audio transformer with a hierarchical structure to reduce the model size and training time. It is further combined with a token-semantic module to map final outputs into class featuremaps, thus enabling the model for the audio event detection (i.e. localization in time). We evaluate HTS-AT on three datasets of audio classification where it achieves new state-of-the-art (SOTA) results on AudioSet and ESC50, and equals the SOTA on Speech Command V2. It also achieves better performance in event localization than the previous CNN-based models. Moreover, HTS-AT requires only 35% model parameters and 15% training time of the previous audio transformer. These results demonstrate the high performance and high efficiency of HTS-AT. Ke Chen 0021, Xingjian Du, Bilei Zhu, Zejun Ma 0001, Taylor Berg-Kirkpatrick, Shlomo Dubnov |
ICASSP | 2 |
| 2022 | Bytecover2: Towards Dimensionality Reduction of Latent Embedding for Efficient Cover Song IdentificationabstractConvolutional neural network (CNN)-based methods have dominated the recent research of cover song identification (CSI). A typical example is the ByteCover system we proposed, which has achieved state-of-the-art results on all the mainstream datasets of CSI. In this paper, we propose an up-graded version of ByteCover, termed ByteCover2, which further improves ByteCover in both identification performance and efficiency. Compared with ByteCover, ByteCover2 is designed with an additional PCA-FC module, which integrates the capability of principal component analysis (PCA) and fully-connected (FC) neural network for dimensionality reduction of the audio embedding, allowing ByteCover2 to perform CSI in a more precise and efficient way. We evaluated ByteCover2 on multiple datasets in different dimension sizes and training settings, where ByteCover2 beat all the compared methods including ByteCover, even with a dimension size of 128, which is 15 times smaller than that of ByteCover. Xingjian Du, Ke Chen 0021, Bilei Zhu, Zejun Ma 0001 |
ICASSP | 1 |
| 2021 | Bytecover: Cover Song Identification Via Multi-Loss TrainingabstractWe present in this paper ByteCover, which is a new feature learning method for cover song identification (CSI). Byte-Cover is built based on the classical ResNet model, and two major improvements are designed to further enhance the capability of the model for CSI. In the first improvement, we introduce the integration of instance normalization (IN) and batch normalization (BN) to build IBN blocks, which are major components of our ResNet-IBN model. With the help of the IBN blocks, our CSI model can learn features that are invariant to the changes of musical attributes such as key, tempo, timbre and genre, while preserving the version information. In the second improvement, we employ the BN-Neck method to allow a multi-loss training and encourage our method to jointly optimize a classification loss and a triplet loss, and by this means, the inter-class discrimination and intra-class compactness of cover songs, can be ensured at the same time. A set of experiments demonstrated the effectiveness and efficiency of ByteCover on multiple datasets, and in the Da-TACOS dataset, ByteCover outperformed the best competitive system by 18.0%. Xingjian Du, Zhesong Yu, Bilei Zhu, Xiaoou Chen, Zejun Ma 0001 |
ICASSP | 1 |
| 2021 | Singing Melody Extraction from Polyphonic Music based on Spectral Correlation ModelingabstractConvolutional neural network (CNN) based methods have achieved state-of-the-art performance for singing melody extraction from polyphonic music. However, most of these methods focus on the learning of local features, while relationships among spectral components locating far apart are often neglected. In this paper, we explore the idea of modeling spectral correlation explicitly for melody extraction. Specifically, we present a spectral correlation module (SCM) that can learn to model the relationships among all frequency bands in a time-frequency representation, thus allowing the encoding of global spectral information into a conventional CNN. Furthermore, we propose to integrate center frequencies with the input feature map of SCM to improve the performance. We implement a light-weight model comprised of SCM blocks to verify the efficacy of our system. Our system achieves a state-of-the-art overall accuracy of 83.5% on the MedleyDB dataset. Xingjian Du, Bilei Zhu, Qiuqiang Kong, Zejun Ma 0001 |
ICASSP | 1 |
| 2021 | An Hrnet-Blstm Model With Two-Stage Training For Singing Melody ExtractionabstractWell-labeled datasets available for melody extraction are scarce, which limits the further advancement of deep learning based methods. To overcome this problem, we propose to use a pitch refinement method to refine the semitone-level pitch sequences decoded from massive melody MIDI files to generate a large number of fundamental frequency (F0) values for model training. Since the refined pitch values used for the first round of training contain errors, a small set of well-labeled data is used for a second round of training. A high-resolution network (HRNet), initially developed for human pose estimation, is introduced for melody extraction. It considers multi-resolution feature learning, making the resulting representation semantically richer. Subsequently, a bidirectional long short-term memory (BLSTM) layer is used to exploit the temporal information of melody. In addition, a new loss function where the unvoiced frames only contribute to voicing detection but not to pitch classification is also proposed to alleviate the class imbalance problem. Experiment results on three public datasets show that the proposed system outperforms four state-of-the-art algorithms in most cases. Yongwei Gao, Xingjian Du, Bilei Zhu, Xiaoheng Sun, Wei Li 0012, Zejun Ma 0001 |
ICASSP | 2 |
| 2021 | Attention-Based Cross-Modal Fusion for Audio-Visual Voice Activity Detection in Musical Video StreamsabstractMany previous audio-visual voice-related works focus on speech, ignoring the singing voice in the growing number of musical video streams on the Internet. For processing diverse musical video data, voice activity detection is a necessary step. This paper attempts to detect the speech and singing voices of target performers in musical video streams using audio-visual information. To integrate information of audio and visual modalities, a multi-branch network is proposed to learn audio and image representations, and the representations are fused by attention based on semantic similarity to shape the acoustic representations through the probability of anchor vocalization. Experiments show the proposed audio-visual multi-branch network far outperforms the audio-only model in challenging acoustic environments, indicating the cross-modal information fusion based on semantic correlation is sensible and successful. Yuanbo Hou, Zhesong Yu, Xia Liang, Xingjian Du, Bilei Zhu, Zejun Ma 0001, Dick Botteldooren |
Interspeech | 4 |
| 2021 | Speech Enhancement with Weakly Labelled Data from AudioSetabstractSpeech enhancement is a task to improve the intelligibility and perceptual quality of degraded speech signal.Recently, neural networks based methods have been applied to speech enhancement.However, many neural network based methods require noisy and clean speech pairs for training.We propose a speech enhancement framework that can be trained with large-scale weakly labelled AudioSet dataset.Weakly labelled data only contain audio tags of audio clips, but not the onset or offset times of speech.We first apply pretrained audio neural networks (PANNs) to detect anchor segments that contain speech or sound events in audio clips.Then, we randomly mix two detected anchor segments containing speech and sound events as a mixture, and build a conditional source separation network using PANNs predictions as soft conditions for speech enhancement.In inference, we input a noisy speech signal with the one-hot encoding of "Speech" as a condition to the trained system to predict enhanced speech.Our system achieves a PESQ of 2.28 and an SSNR of 8.75 dB on the VoiceBank-DEMAND dataset, outperforming the previous SEGAN system of 2.16 and 7.73 dB respectively. Qiuqiang Kong, Haohe Liu, Xingjian Du, Yuxuan Wang 0002 |
Interspeech | 3 |