Bilei Zhu

dblp:88/8693 · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
15since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Joint Music and Language Attention Models for Zero-Shot Music Tagging
abstract
Music tagging is a task to predict the tags of music recordings. However, previous music tagging research primarily focuses on close-set music tagging tasks which can not be generalized to new tags. In this work, we propose a zero-shot music tagging system modeled by a joint music and language attention (JMLA) model to address the open-set music tagging problem. The JMLA model consists of an audio encoder modeled by a pretrained masked autoencoder and a decoder modeled by a Falcon7B. We introduce preceiver resampler to convert arbitrary length audio into fixed length embeddings. We introduce dense attention connections between encoder and decoder layers to improve the information flow between the encoder and decoder layers. We collect a large-scale music and description dataset from the internet. We propose to use ChatGPT to convert the raw descriptions into formalized and diverse descriptions to train the JMLA models. Our proposed JMLA system achieves a zero-shot audio tagging accuracy of 64.82% on the GTZAN dataset, outperforming previous zero-shot systems and achieves comparable results to previous systems on the FMA and the MagnaTagATune datasets.
Xingjian Du, Zhesong Yu, Jiaju Lin, Bilei Zhu, Qiuqiang Kong
ICASSP4
2024 ByteHum: Fast and Accurate Query-by-Humming in the Wild
abstract
Query by Humming (QBH) is a practically meaningful task, while most existing methods struggle to scale to real-life applications due to the complex preprocessing for building the database and the limited search speed. In this paper, we propose the ByteHum system, a fast and efficient humming retrieval system which is capable of searching against large-scale databases built on raw song audios without the need for extensive preprocessing. ByteHum employs a convolutional neural network to extract features from raw audio, and utilizes a source-separated cover song identification dataset for weakly supervised training of the feature extractor. We explore the use of unsupervised domain adaptation techniques to enhance the performance of our weakly supervised model on the QBH task. Furthermore, to evaluate QBH systems’ performance on non-manually processed databases in the wild, we annotate original recordings for three existing QBH benchmark sets. Our experimental results demonstrate that ByteHum significantly outperforms existing QBH systems in terms of speed and accuracy under both classical and unconstrained settings.
Xingjian Du, Pei Zou, Xia Liang, Minghang Chu, Bilei Zhu
ICASSP6
2024 MINT: Boosting Audio-Language Model via Multi-Target Pre-Training and Instruction Tuning
Yifei Xin, Zhesong Yu, Bilei Zhu, Lu Lu 0015, Zejun Ma 0001
INTERSPEECH4
2023 Bytecover3: Accurate Cover Song Identification On Short Queries
abstract
Deep learning based methods have become a paradigm for cover song identification (CSI) in recent years, where the ByteCover systems have achieved state-of-the-art results on all the mainstream datasets of CSI. However, with the burgeon of short videos, many real-world applications require matching short music excerpts to full-length music tracks in the database, which is still under-explored and waiting for an industrial-level solution. In this paper, we upgrade the previous ByteCover systems to ByteCover3 that utilizes local features to further improve the identification performance of short music queries. ByteCover3 is designed with a local alignment loss (LAL) module and a two-stage feature retrieval pipeline, allowing the system to perform CSI in a more precise and efficient way. We evaluated ByteCover3 on multiple datasets with different benchmark settings, where ByteCover3 beat all the compared methods including its previous versions.
Xingjian Du, Xia Liang, Huidong Liang, Bilei Zhu, Zejun Ma 0001
ICASSP5
2023 Graph contrastive learning with implicit augmentations
Huidong Liang, Xingjian Du, Bilei Zhu, Zejun Ma 0001, Ke Chen 0021, Junbin Gao
Neural Networks3
2022 Zero-Shot Audio Source Separation through Query-Based Learning from Weakly-Labeled Data
abstract
Deep learning techniques for separating audio into different sound sources face several challenges. Standard architectures require training separate models for different types of audio sources. Although some universal separators employ a single model to target multiple sources, they have difficulty generalizing to unseen sources. In this paper, we propose a three-component pipeline to train a universal audio source separator from a large, but weakly-labeled dataset: AudioSet. First, we propose a transformer-based sound event detection system for processing weakly-labeled training data. Second, we devise a query-based audio separation model that leverages this data for model training. Third, we design a latent embedding processor to encode queries that specify audio targets for separation, allowing for zero-shot generalization. Our approach uses a single model for source separation of multiple sound types, and relies solely on weakly-labeled data for training. In addition, the proposed audio separator can be used in a zero-shot setting, learning to separate types of audio sources that were never seen in training. To evaluate the separation performance, we test our model on MUSDB18, while training on the disjoint AudioSet. We further verify the zero-shot performance by conducting another experiment on audio source types that are held-out from training. The model achieves comparable Source-to-Distortion Ratio (SDR) performance to current supervised models in both cases.
Ke Chen 0021, Xingjian Du, Bilei Zhu, Zejun Ma 0001, Taylor Berg-Kirkpatrick, Shlomo Dubnov
AAAI3
2022 HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection
abstract
Audio classification is an important task of mapping audio samples into their corresponding labels. Recently, the transformer model with self-attention mechanisms has been adopted in this field. However, existing audio transformers require large GPU memories and long training time, meanwhile relying on pretrained vision models to achieve high performance, which limits the model’s scalability in audio tasks. To combat these problems, we introduce HTS-AT: an audio transformer with a hierarchical structure to reduce the model size and training time. It is further combined with a token-semantic module to map final outputs into class featuremaps, thus enabling the model for the audio event detection (i.e. localization in time). We evaluate HTS-AT on three datasets of audio classification where it achieves new state-of-the-art (SOTA) results on AudioSet and ESC50, and equals the SOTA on Speech Command V2. It also achieves better performance in event localization than the previous CNN-based models. Moreover, HTS-AT requires only 35% model parameters and 15% training time of the previous audio transformer. These results demonstrate the high performance and high efficiency of HTS-AT.
Ke Chen 0021, Xingjian Du, Bilei Zhu, Zejun Ma 0001, Taylor Berg-Kirkpatrick, Shlomo Dubnov
ICASSP3
2022 Bytecover2: Towards Dimensionality Reduction of Latent Embedding for Efficient Cover Song Identification
abstract
Convolutional neural network (CNN)-based methods have dominated the recent research of cover song identification (CSI). A typical example is the ByteCover system we proposed, which has achieved state-of-the-art results on all the mainstream datasets of CSI. In this paper, we propose an up-graded version of ByteCover, termed ByteCover2, which further improves ByteCover in both identification performance and efficiency. Compared with ByteCover, ByteCover2 is designed with an additional PCA-FC module, which integrates the capability of principal component analysis (PCA) and fully-connected (FC) neural network for dimensionality reduction of the audio embedding, allowing ByteCover2 to perform CSI in a more precise and efficient way. We evaluated ByteCover2 on multiple datasets in different dimension sizes and training settings, where ByteCover2 beat all the compared methods including ByteCover, even with a dimension size of 128, which is 15 times smaller than that of ByteCover.
Xingjian Du, Ke Chen 0021, Bilei Zhu, Zejun Ma 0001
ICASSP4
2022 S3T: Self-Supervised Pre-Training with Swin Transformer For Music Classification
abstract
In this paper, we propose S3T, a self-supervised pre-training method with Swin Transformer for music classification, aiming to learn meaningful music representations from massive easily accessible unlabeled music data. S3T introduces a momentum-based paradigm, MoCo, with Swin Transformer as its feature extractor to music time-frequency domain. For better music representations learning, S3T contributes a music data augmentation pipeline and two specially designed pre-processors. To our knowledge, S3T is the first method combining the Swin Transformer with a self-supervised learning method for music classification. We evaluate S3T on music genre classification and music tagging tasks with linear classifiers trained on learned representations. Experimental results show that S3T outperforms the previous self-supervised method (CLMR) by 12.5 percents top-1 accuracy and 4.8 percents PR-AUC on two tasks respectively, and also surpasses the task-specific state-of-the-art supervised methods. Besides, S3T shows advances in label efficiency using only 10% labeled data exceeding CLMR on both tasks with 100% labeled data.
Chen Zhang 0020, Bilei Zhu, Zejun Ma 0001
ICASSP3
2022 GIO: A Timbre-informed Approach for Pitch Tracking in Highly Noisy Environments
abstract
As one of the fundamental tasks in music and speech signal processing, pitch tracking has been attracting attention for decades. While a human can focus on the voiced pitch even in highly noisy environments, most existing automatic pitch tracking systems show unsatisfactory performance encountering noise. To mimic human auditory, a data-driven model named GIO is proposed in this paper, in which timbre information is introduced to guide pitch tracking. The proposed model takes two inputs: a short audio segment to extract pitch from and a timbre embedding derived from the speaker's or singer's voice. In experiments, we use a music artist classification model to extract timbre embedding vectors. A dual-branch structure and a two-step training method are designed to enable the model to predict voice presence. The experimental results show that the proposed model gains a significant improvement in noise robustness and outperforms existing state-of-the-art methods with fewer parameters.
Xiaoheng Sun, Xia Liang, Qiqi He, Bilei Zhu, Zejun Ma 0001
ICMR4
2021 Bytecover: Cover Song Identification Via Multi-Loss Training
abstract
We present in this paper ByteCover, which is a new feature learning method for cover song identification (CSI). Byte-Cover is built based on the classical ResNet model, and two major improvements are designed to further enhance the capability of the model for CSI. In the first improvement, we introduce the integration of instance normalization (IN) and batch normalization (BN) to build IBN blocks, which are major components of our ResNet-IBN model. With the help of the IBN blocks, our CSI model can learn features that are invariant to the changes of musical attributes such as key, tempo, timbre and genre, while preserving the version information. In the second improvement, we employ the BN-Neck method to allow a multi-loss training and encourage our method to jointly optimize a classification loss and a triplet loss, and by this means, the inter-class discrimination and intra-class compactness of cover songs, can be ensured at the same time. A set of experiments demonstrated the effectiveness and efficiency of ByteCover on multiple datasets, and in the Da-TACOS dataset, ByteCover outperformed the best competitive system by 18.0%.
Xingjian Du, Zhesong Yu, Bilei Zhu, Xiaoou Chen, Zejun Ma 0001
ICASSP3
2021 Singing Melody Extraction from Polyphonic Music based on Spectral Correlation Modeling
abstract
Convolutional neural network (CNN) based methods have achieved state-of-the-art performance for singing melody extraction from polyphonic music. However, most of these methods focus on the learning of local features, while relationships among spectral components locating far apart are often neglected. In this paper, we explore the idea of modeling spectral correlation explicitly for melody extraction. Specifically, we present a spectral correlation module (SCM) that can learn to model the relationships among all frequency bands in a time-frequency representation, thus allowing the encoding of global spectral information into a conventional CNN. Furthermore, we propose to integrate center frequencies with the input feature map of SCM to improve the performance. We implement a light-weight model comprised of SCM blocks to verify the efficacy of our system. Our system achieves a state-of-the-art overall accuracy of 83.5% on the MedleyDB dataset.
Xingjian Du, Bilei Zhu, Qiuqiang Kong, Zejun Ma 0001
ICASSP2
2021 An Hrnet-Blstm Model With Two-Stage Training For Singing Melody Extraction
abstract
Well-labeled datasets available for melody extraction are scarce, which limits the further advancement of deep learning based methods. To overcome this problem, we propose to use a pitch refinement method to refine the semitone-level pitch sequences decoded from massive melody MIDI files to generate a large number of fundamental frequency (F0) values for model training. Since the refined pitch values used for the first round of training contain errors, a small set of well-labeled data is used for a second round of training. A high-resolution network (HRNet), initially developed for human pose estimation, is introduced for melody extraction. It considers multi-resolution feature learning, making the resulting representation semantically richer. Subsequently, a bidirectional long short-term memory (BLSTM) layer is used to exploit the temporal information of melody. In addition, a new loss function where the unvoiced frames only contribute to voicing detection but not to pitch classification is also proposed to alleviate the class imbalance problem. Experiment results on three public datasets show that the proposed system outperforms four state-of-the-art algorithms in most cases.
Yongwei Gao, Xingjian Du, Bilei Zhu, Xiaoheng Sun, Wei Li 0012, Zejun Ma 0001
ICASSP3
2021 Rule-Embedded Network for Audio-Visual Voice Activity Detection in Live Musical Video Streams
abstract
Detecting anchor’s voice in live musical streams is an important preprocessing step for music and speech signal processing. Existing approaches to voice activity detection (VAD) primarily rely on audio, however, audio-based VAD is difficult to effectively focus on the target voice in noisy environments. This paper proposes a rule-embedded network to fuse the audio-visual (A-V) inputs for better detection of the target voice. The core role of the rule in the model is to coordinate the relation between the bi-modal information and use visual representations as a mask to filter out the information of non-target sound. Experiments show that: 1) with the help of cross-modal fusion using the proposed rule, the detection results of the A-V branch outperform that of the audio branch in the same model framework; 2) the performance of the bimodal A-V model far outperforms that of audio-only models, indicating that the incorporation of both audio and visual signals is highly beneficial for VAD. To attract more attention to the cross-modal music and audio signal processing, a new live musical video corpus with frame-level labels is introduced.
Yuanbo Hou, Bilei Zhu, Zejun Ma 0001, Dick Botteldooren
ICASSP3
2021 Attention-Based Cross-Modal Fusion for Audio-Visual Voice Activity Detection in Musical Video Streams
abstract
Many previous audio-visual voice-related works focus on speech, ignoring the singing voice in the growing number of musical video streams on the Internet. For processing diverse musical video data, voice activity detection is a necessary step. This paper attempts to detect the speech and singing voices of target performers in musical video streams using audio-visual information. To integrate information of audio and visual modalities, a multi-branch network is proposed to learn audio and image representations, and the representations are fused by attention based on semantic similarity to shape the acoustic representations through the probability of anchor vocalization. Experiments show the proposed audio-visual multi-branch network far outperforms the audio-only model in challenging acoustic environments, indicating the cross-modal information fusion based on semantic correlation is sensible and successful.
Yuanbo Hou, Zhesong Yu, Xia Liang, Xingjian Du, Bilei Zhu, Zejun Ma 0001, Dick Botteldooren
Interspeech5
2019 Vocal Melody Extraction via DNN-based Pitch Estimation and Salience-based Pitch Refinement
abstract
Data-driven methods for melody extraction from polyphonic music generally require large amounts of labeled data for model training. However, musical data with annotations of melody fundamental frequency (F0) are rare and hard to obtain. To overcome this limitation, in this paper we propose to use melody MIDI files, which are more massively available, as the sources of labels to train a deep neural network (DNN) model for melody extraction. For each testing audio, the pitch sequence estimated by DNN is comprised of note numbers quantized at semitone level, and their resolution is relatively low. Therefore, we further propose a salience-based method to refine the pitch estimate of DNN to a higher resolution of 10 cents. Experimental results on three public datasets indicate that our method outperforms four state-of-the-art melody extraction methods in most cases.
Yongwei Gao, Bilei Zhu, Wei Li 0012, Ke Li 0015, Yongjian Wu 0001, Feiyue Huang
ICASSP2
2017 Fusing transcription results from polyphonic and monophonic audio for singing melody transcription in polyphonic music
abstract
This paper presents a new system for singing melody transcription from polyphonic songs. Instead of operating solely on polyphonic audio of each song to be processed (as most existing systems do), our system takes as inputs additionally multiple monophonic recordings of people singing the song. To transcribe the singing melody in a song, our system first tracks the singing pitch from polyphonic audio of the song by using a deep neural network (DNN)-based method, and then uses the estimated pitch series as reference to select the pitch sequences extracted from the multiple monophonic singing recordings. The selected monophonic pitch sequences, as well as the DNN pitch series from the polyphonic audio, are then transcribed separately, and their transcriptions results are fused to form the final note sequence. Experimental results show that, by introducing monophonic singings into transcription, the performance of singing melody transcription from polyphonic songs can be significantly improved.
Bilei Zhu, Fuzhang Wu, Ke Li 0015, Yongjian Wu 0001, Feiyue Huang, Yunsheng Wu
ICASSP1
2015 Latent time-frequency component analysis: A novel pitch-based approach for singing voice separation
abstract
Monaural singing voice separation has aroused considerable attention. Many pitch-based methods have been proposed to address this task, but generally have limited performance. The most crucial difficulties lie in the inaccurate judgment on voiced pitches and the failed recognition on unvoiced singing sounds. In this paper, we propose a novel algorithm based on the latent component analysis of time-frequency representation to overcome these difficulties. Specifically, the time-frequency (T-F) representations of the song are firstly decomposed into components, and each component approximately originates from a single sound source. We then construct non-overlapping T-F segments with these components, to complete the omitted useful singing voice information. Extensive experiments on the MIR-1K public dataset shows the effectiveness of the proposed algorithm.
Wei Li 0012, Bilei Zhu
ICASSP3
2015 Towards Solving the Bottleneck of Pitch-based Singing Voice Separation
abstract
Singing voice separation from accompaniment in monaural music recordings is a crucial technique in music information retrieval. A majority of existing algorithms are based on singing pitch detection, and take the detected pitch as the cue to identify and separate the harmonic structure of the singing voice. However, as a key yet undependable premise, vocal pitch detection makes the separation performance of these algorithms rather limited. To overcome the inherent weakness of pitch-based inference algorithms, two novel methods based on non-negative matrix factorization (NMF) are devised in this paper. The first one combines NMF with the distribution regularities of vocals under different time frequency resolutions, so that many vocal unrelated portions are eliminated and the singing voice is hence enhanced. In consequence, the accuracy of vocal pitch detection is significantly improved. The second method applies NMF to decompose the spectrogram into non-overlapping and indivisible segments, which can be used as another cue besides the pitch to help identify the vocal harmonic structure. The two proposed methods are integrated into the framework of pitch-based inference. Extensive testing on the MIR-1K public dataset shows that both of them are rather effective, and the overall performances outperform other state-of-the-art singing separation algorithms.
Bilei Zhu, Wei Li 0012
ACM Multimedia1
2013 Multi-Stage Non-Negative Matrix Factorization for Monaural Singing Voice Separation
abstract
Separating singing voice from music accompaniment can be of interest for many applications such as melody extraction, singer identification, lyrics alignment and recognition, and content-based music retrieval. In this paper, a novel algorithm for singing voice separation in monaural mixtures is proposed. The algorithm consists of two stages, where non-negative matrix factorization (NMF) is applied to decompose the mixture spectrograms with long and short windows respectively. A spectral discontinuity thresholding method is devised for the long-window NMF to select out NMF components originating from pitched instrumental sounds, and a temporal discontinuity thresholding method is designed for the short-window NMF to pick out NMF components that are from percussive sounds. By eliminating the selected components, most pitched and percussive elements of the music accompaniment are filtered out from the input sound mixture, with little effect on the singing voice. Extensive testing on the MIR-1K public dataset of 1000 short audio clips and the Beach-Boys dataset of 14 full-track real-world songs showed that the proposed algorithm is both effective and efficient.
Bilei Zhu, Wei Li 0012, Ruijiang Li, Xiangyang Xue 0001
IEEE Trans. Speech Audio Process.1
2012 On the music content authentication
abstract
Digital audio has been ubiquitous over the past decade. Since it can be easily modified by editing tools, there has been a strong need to protect its content for secure multimedia applications. Existing audio authentication algorithms are mainly focused on either human speech or general audio with music as part of the test data, while special research on music authentication has been somewhat neglected. In this article, we propose a novel algorithm to protect the integrity and authenticity of music signals. Its main contributions include: (1) Music is segmented into beat-based frames, which not only endows the authentication units with more semantic meaning but also perfectly resolves the challenging synchronization problem; (2) Robust hashes are generated from Chroma-based mid-level audio feature which can appropriately characterize the music content, and integrated with an encryption procedure to ensure the security against malicious block-wise vector quantization attack; (3) Fuzzy logic is adopted to make the authentication decision in light of three measures defined on bit errors, coinciding with the inherent blurred nature of authentication. Experiments exhibit good discriminative ability between admissible and malicious operations.
Wei Li 0012, Bilei Zhu, Zhurong Wang
ACM Multimedia2
2010 Robust hashing for music copyright protection by combining beat segmentation and chroma
abstract
Time-scale modification and pitching shifting are two recognized challenging attacks to music copyright protection. To resist them simultaneously, a novel robust hashing method is proposed by combining the strength of music beat segmentation and chroma-based music feature. These two measures are aimed at solving the problem of desynchronization and frequency shifting respectively. Moreover, two layers of scrambling are performed to ensure the security. Experiments exhibit remarkable robustness against various attacks including pitch [email protected]%, time-scale [email protected]%, and [email protected]/10 etc.
Wei Li 0012, Zhurong Wang, Bilei Zhu, Xiangyang Xue 0001
ACM Multimedia3
2010 A novel audio fingerprinting method robust to time scale modification and pitch shifting
abstract
A novel audio fingerprinting method that is highly robust to Time Scale Modification (TSM) and pitch shifting is proposed. Instead of simply employing spectral or tempo-related features, our system is based on computer-vision techniques. We transform each 1-D audio signal into a 2-D image and treat TSM and pitch shifting of the audio signal as stretch and translation of the corresponding image. Robust local descriptors are extracted from the image and matched against those of the reference audio signals. Experimental results show that our system is highly robust to various audio distortions, including the challenging TSM and pitch shifting.
Bilei Zhu, Wei Li 0012, Zhurong Wang, Xiangyang Xue 0001
ACM Multimedia1