Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Xiaoou Chen

dblp:97/1334 · DBLP profile ↗
← Back
26ranked-venue papers
0as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 1 since 2021Artificial intelligence and machine learning · 3Databases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
4 papers
Audio and music processing · 95% Multimedia analysis and retrieval · 5%
Artificial intelligence
2 papers
Representation and self-supervised learning · 72% Deep learning architectures and training · 28%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing › music information retrieval › music similarity
cover song identification
0.722019
Temporal Pyramid Pooling Convolutional Neural Network for Cover Song Identification · IJCAI 2019
Effective Music Feature NCP: Enhancing Cover Song Recognition with Music Transcription · SIGIR 2017
Audio and music processing
music information retrieval
0.722019
Temporal Pyramid Pooling Convolutional Neural Network for Cover Song Identification · IJCAI 2019
Effective Music Feature NCP: Enhancing Cover Song Recognition with Music Transcription · SIGIR 2017
Audio and music processing › audio representation learning
music representation learning
0.412019
Temporal Pyramid Pooling Convolutional Neural Network for Cover Song Identification · IJCAI 2019
Machine learning › Representation and self-supervised learning › representation learning › sequence representation learning
audio representation learning
0.312017
Audio Feature Learning with Triplet-Based Embedding Network · AAAI 2017
Audio and music processing › music information retrieval
music feature extraction
0.312017
Effective Music Feature NCP: Enhancing Cover Song Recognition with Music Transcription · SIGIR 2017
Audio and music processing
music transcription
0.312017
Effective Music Feature NCP: Enhancing Cover Song Recognition with Music Transcription · SIGIR 2017
Machine learning › Deep learning architectures and training
convolutional neural network
0.112019
Temporal Pyramid Pooling Convolutional Neural Network for Cover Song Identification · IJCAI 2019
Multimedia analysis and retrieval
audio retrieval
0.112006
Audio similarity measure by graph modeling and matching · ACM Multimedia 2006
Multimedia analysis and retrieval › audio retrieval
audio similarity
0.112006
Audio similarity measure by graph modeling and matching · ACM Multimedia 2006

Methods — techniques the papers use, named apart from their topics

temporal pyramid pooling · 0.8convolutional neural network · 0.8triplet network · 0.6metric learning · 0.6music transcription · 0.3optimal matching · 0.1graph matching · 0.1dynamic programming · 0.1
YearPublicationVenuePosition
2021 Bytecover: Cover Song Identification Via Multi-Loss Training
abstract
We present in this paper ByteCover, which is a new feature learning method for cover song identification (CSI). Byte-Cover is built based on the classical ResNet model, and two major improvements are designed to further enhance the capability of the model for CSI. In the first improvement, we introduce the integration of instance normalization (IN) and batch normalization (BN) to build IBN blocks, which are major components of our ResNet-IBN model. With the help of the IBN blocks, our CSI model can learn features that are invariant to the changes of musical attributes such as key, tempo, timbre and genre, while preserving the version information. In the second improvement, we employ the BN-Neck method to allow a multi-loss training and encourage our method to jointly optimize a classification loss and a triplet loss, and by this means, the inter-class discrimination and intra-class compactness of cover songs, can be ensured at the same time. A set of experiments demonstrated the effectiveness and efficiency of ByteCover on multiple datasets, and in the Da-TACOS dataset, ByteCover outperformed the best competitive system by 18.0%.
Xingjian Du, Zhesong Yu, Bilei Zhu, Xiaoou Chen, Zejun Ma 0001
ICASSP4
2020 Similarity Learning For Cover Song Identification Using Cross-Similarity Matrices of Multi-Level Deep Sequences
abstract
In recent years, several deep learning models have been proposed for cover song identification and they have been designed to learn fixed-length feature vectors for music tracks. However, the aspect of temporal progression of music, which is important for measuring the melody similarity between two tracks, is not well represented by fixed-length vectors. In this paper, we propose a new Siamese network architecture for music melody similarity metric learning. The architecture consists of two parts. One part is a network for learning the deep sequence representation of music tracks, and the other is a similarity estimation network which takes as input the cross-similarity matrices calculated from the deep sequences of a pair of tracks. The two networks are jointly trained and optimized to achieve high melody similarity prediction accuracy. Experiments conducted on several public datasets demonstrate the superiority of the proposed architecture.
Chaoya Jiang, Deshun Yang, Xiaoou Chen
ICASSP3
2020 Learning a Representation for Cover Song Identification Using Convolutional Neural Network
abstract
Cover song identification is a challenging task in the field of Music Information Retrieval (MIR) due to complex musical variations between query tracks and cover versions. Previous works typically utilize hand-crafted features and alignment algorithms for the task. More recently, further breakthroughs are achieved by employing neural network approaches. In this paper, we propose a novel Convolutional Neural Network (CNN) towards cover song identification. We train the network through classification criteria. Having been trained, the network is used to extract music representation for cover song identification. A training scheme is designed to train robust models against tempo changes. Experimental results show that our approach outperforms state-of-the-art methods on several public datasets with low time complexity.
Zhesong Yu, Xiaoshuo Xu, Xiaoou Chen, Deshun Yang
ICASSP3
2020 Learn A Robust Representation For Cover Song Identification Via Aggregating Local And Global Music Temporal Context
abstract
Recently, deep learning models have been proposed for cover song identification and designed to learn fixed-length feature vectors for music recordings. However, the aspect of the temporal progression of music, which is important for measuring the melody similarity between two recordings, is not well exploited in those models. In this paper, we propose a new Siamese architecture to learn deep representations for cover song identification where Dilated Temporal Pyramid Convolution is used to exploit the local temporal context and Temporal Self-Attention to exploit the global temporal context in music recordings. In addition to the traditional block which calculates the similarity between a pair of recordings, we add a classification block to classify the recordings to their respective cliques. By combining the regression loss and the classification loss, our model can leam more robust and discriminative latent representations. The representations extracted by our model show substantial superiority to existing hand-crafted features and learned deep features. Experimental results show that our approach far outperforms the state-of the-art methods on several public datasets.
Chaoya Jiang, Deshun Yang, Xiaoou Chen
ICME3
2020 Gen-Res-Net: A Novel Generative Model for Singing Voice Separation
Congzhou Tian, Deshun Yang, Xiaoou Chen
MMM (1)4
2020 A Distinct Synthesizer Convolutional TasNet for Singing Voice Separation
Congzhou Tian, Deshun Yang, Xiaoou Chen
MMM (1)3
2019 Temporal Pyramid Pooling Convolutional Neural Network for Cover Song Identification
abstract
Cover song identification is an important problem in the field of Music Information Retrieval. Most existing methods rely on hand-crafted features and sequence alignment methods, and further breakthrough is hard to achieve. In this paper, Convolutional Neural Networks (CNNs) are used for representation learning toward this task. We show that they could be naturally adapted to deal with key transposition in cover songs. Additionally, Temporal Pyramid Pooling is utilized to extract information on different scales and transform songs with different lengths into fixed-dimensional representations. Furthermore, a training scheme is designed to enhance the robustness of our model. Extensive experiments demonstrate that combined with these techniques, our approach is robust against musical variations existing in cover songs and outperforms state-of-the-art methods on several datasets with low time complexity.
Zhesong Yu, Xiaoshuo Xu, Xiaoou Chen, Deshun Yang
IJCAI3
2018 Effective Cover Song Identification Based on Skipping Bigrams
abstract
So far, few cover song identification systems that utilize index techniques achieve great success. In this paper, we propose a novel approach based on skipping bigrams that could be used for effective index. By applying Vector Quantization, our algorithm encodes signals into code sequences. Then, the bigram histograms of code sequences are used to represent the original recordings and measure their similarities. Through Vector Quantization and skipping bigrams, our model shows great robustness against speed and structure variations in cover songs. Experimental results demonstrate that our model achieves better performance than recent methods and is less computationally demanding.
Xiaoshuo Xu, Xiaoou Chen, Deshun Yang
ICASSP2
2018 Key-Invariant Convolutional Neural Network Toward Efficient Cover Song Identification
abstract
Cover song identification has long been a challenging task due to key, timbre and structure variations in different renditions of a song. Previous research mostly involves handcrafted features and sequence alignment methods, where further breakthroughs can hardly be achieved. In this paper, we utilize a supervised deep learning method to learn an effective feature extractor for cover song identification. Regarding it as a classification problem, we propose a key-invariant convolutional neural network robust against key transposition for classification. Having been trained, the network is used to extract representations of music, which could be used to measure the similarity between songs. Besides, the representations are highly sparse; effective algorithms can be devised to accelerate the computation. Experimental results show that our model achieves high precision with low time cost on several datasets. Especially, our method outperforms the state-of-the-art approach on one of the datasets.
Xiaoshuo Xu, Xiaoou Chen, Deshun Yang
ICME2
2018 Triplet Convolutional Network for Music Version Identification
Xiaoyu Qi, Deshun Yang, Xiaoou Chen
MMM (1)3
2018 Efficient Two-Layer Model Towards Cover Song Identification
Xiaoshuo Xu, Xiaoou Chen, Deshun Yang
MMM (2)3
2018 Robust Polyphonic Sound Event Detection by Using Multi Frame Size Denoising Autoencoder
abstract
Over the past few years, lots of research has been done on polyphonic sound event detection. A main problem with sound event detection is that the detection performance sharply degrades in the presence of noise. As denoising autoencoder reportedly has superior performance in noisy environments, this paper proposes to use denoising autoencoder, which is trained by multi frame size information of audio signals, to extract robust features in a task of polyphonic sound event detection under noisy conditions. Performance of the extracted feature is evaluated by polyphonic sound event detection experiments with different noise levels, and compared with that of baseline features including Mel-band Energy (Mel), Log mel-band Energy (Logmel) and mel-frequency cepstral coefficients (MFCC). The experiemntal results show that the proposed feature has the best robustness among all features and achieves the best detection effect under noisy conditions.
Jianchao Zhou, Xiaoou Chen, Deshun Yang
MMSP2
2017 Audio Feature Learning with Triplet-Based Embedding Network
abstract
We propose a triplet-based network for audio feature learning for version identification. Existing methods use hand-crafted features for a music as a whole while we learn features by a triplet-based neural network on segment-level, focusing on the most similar parts between music versions. We conduct extensive experiments and demonstrate our merits.
Xiaoyu Qi, Deshun Yang, Xiaoou Chen
AAAI3
2017 Effective Music Feature NCP: Enhancing Cover Song Recognition with Music Transcription
abstract
Chroma is a widespread feature for cover song recognition, as it is robust against non-tonal components and independent of timbre and specific instruments. However, Chroma is derived from spectrogram, thus it provides a coarse approximation representation of musical score. In this paper, we proposed a similar but more effective feature Note Class Profile (NCP) derived with music transcription techniques. NCP is a multi-dimensional time serie, each column of which denotes the energy distribution of 12 note classes. Experimental results on benchmark datasets demonstrated its superior performance over existing music features. In addition, NCP feature can be enhanced further with the development of music transcription techniques. The source code can be found in github1.
Xiaoou Chen, Deshun Yang, Xiaoshuo Xu
SIGIR2
2016 Two-layer large-scale cover song identification system based on music structure segmentation
abstract
This paper focuses on cover song identification over a large-scale dataset. Identifying all covers of a query song from music collection is a challenging task since covers vary in multiple aspects, such as tempo, key, and structure. For the large-scale dataset, cover song identification is more challenging and few works have been published. Previous works usually use a single representation for a whole song, such as 2D Fourier transform and chord profiles, which cannot reflect the property that covers are largely determined by a local similarity. To address this problem, we propose a novel cover song identification method based on music structure segmentation. The proposed structural method identifies cover songs on section level instead of song level. The experimental results show that the structural method improves the mean average precision of 2D Fourier transform method from 9.5% to 12.1%. In addition, we also propose a two-layer cover song identification system to improve the efficiency.
Kang Cai, Deshun Yang, Xiaoou Chen
MMSP3
2016 Robust sound event classification by using denoising autoencoder
abstract
Over the last decade, a lot of research has been done on sound event classification. But a main problem with sound event classification is that the performance sharply degrades in the presence of noise. As spectrogram-based image features and denoising auto encoder reportedly have superior performance in noisy conditions, this paper proposes a new robust feature called denoising auto encoder image feature (DIF) for sound event classification which is an image feature extracted from an image-like representation produced by denoising auto encoder. Performance of the feature is evaluated by a classification experiment using a SVM classifier on audio examples with different noise levels, and compared with that of baseline features including mel-frequency cepstral coefficients (MFCC) and spectrogram image feature. The proposed DIF demonstrates better performance under noise-corrupted conditions.
Jianchao Zhou, Liqun Peng, Xiaoou Chen, Deshun Yang
MMSP3
2015 Music identification based on music word model
abstract
For music identification, conventional bag of audio words model methods generally compute a histogram for a piece of music, which ignores the temporal characteristic of music and has a negative influence on the accuracy. In addition, they are usually based on DFT spectrogram, which cannot represent music as well as Constant Q (CQ) spectrogram. To address the above problems, we propose a two-layer representation method based on a set of music words for music identification. Firstly, music words are learned from the CQ spectrogram as typical patterns. Then, based on the obtained music words, a piece of music can be represented as word sequence and word histogram. We can reduce the number of possible similar candidates effectively with the histogram similarity measure, and the final result is determined by the sequence similarity measure. Based on the distribution of music word frequency, a low frequency word filter strategy is devised to increase the identification speed, which is essential for large systems such as a million song library. Experiments demonstrate the effectiveness and efficiency of our proposed method.
Wanyi Yang, Deshun Yang, Xiaoou Chen, Haiqian He
ICME3
2015 Audio event recognition based on DBN features from multiple filter-bank representations
abstract
In the audio event classification or detection research field, the representation of the audio itself is important. Many researchers tried to apply Deep Belief Network (DBN) to learn new representations of the audio. The mel filter-bank feature, which is obtained based on mel scale, is commonly used as the low level representation of the audio in the pre-processing procedure of DBN. However, the mel bands used in mel filter-bank feature may not be sufficient for the comprehensive representation of the diverse audio events in the real world and then it will make it difficult for DBN to learn good audio features. In this paper, two steps are taken to explore and tackle the problem. In the first step, we conduct a comparison of the effects among different arrangements of frequency bands to DBN feature learning in the audio event recognition. Here the arrangements of frequency bands include mel bands, bark bands, linear bands and pyramid bands. In the second step, in order to utilize the different classification capabilities of the DBN features on different audio events, we adopt the Adaboost algorithm to fuse them. We conduct the experiments on real datasets collected from findsound website, and the results verifies that our proposed audio event classification system, which uses diverse features selected by Adaboost from all sets of DBN features, outperforms the one using only one kind of DBN feature set.
Xiaoou Chen, Deshun Yang
MMSP2
2014 Multi-person tracking-by-detection with local particle filtering and global occlusion handling
abstract
This paper presents a detection-based method for tracking an uncertain number of persons in complex scenarios with frequent occlusions. Frame-by-frame data association based particle filters are adopted to track targets in occlusion-free regions. When occlusion is detected, the associated trackers are deactivated and they are re-activated when the tracked persons are re-identified after occlusion. The re-identification problem is solved by global data association. And the association cost matrix only integrates information collected from the frames after occlusion to avoid tracking failure caused by false detections during occlusion. Furthermore, we improve the particle initialization by motion prediction and automatically configured dynamic model. Experimental results show that the proposed algorithm effectively reduces id switches and lost trajectories which happen frequently in local filtering methods. In the meantime, the algorithm is suitable for time-critical applications.
Yaowen Guan, Xiaoou Chen, Deshun Yang, Yuqian Wu
ICME2
2014 Efficient music identification by utilizing space-saving audio fingerprinting system
abstract
Audio fingerprints can be used to implement an efficient music identification system on a million-song library, but the system requires huge amount of memory to hold the fingerprints and indexes. Therefore, for a large-scale music library, memory imposes a restriction on the speed of music identification. In this paper, we propose an efficient music identification system which utilizes a kind of space-saving audio fingerprints. For saving space, original fingerprints are sub-sampled and only one quarter of the original data is reserved. In this way, memory requirement is much decreased and the search speed is significantly increased while the robustness and reliability are well preserved. Extensive experiments have been conducted to compare our system with a previous method, and the results show that our system requires much less memory, runs faster and achieves comparable accuracy for a large-scale database.
Xiaoou Chen, Deshun Yang
ICME2
2014 Smoke Detection Based on a Semi-supervised Clustering Model
Haiqian He, Liqun Peng, Deshun Yang, Xiaoou Chen
MMM (2)4
2010 Enriching music mood annotation by semantic association reasoning
abstract
Mood annotation of music is challenging as it concerns not only audio content but also extra-musical information. It is a representative research topic about how to traverse the well-known semantic gap. In this paper, we propose a new music-mood-specific ontology. Novel ontology-based semantic reasoning methods are applied to effectively bridge content-based information with web-based resources. Also, the system can automatically discover closely relevant semantics for music mood and thus a novel weighting method is proposed for mood propagation. Experiments show that the proposed method outperforms purely content-based methods and significantly enhances the mood prediction accuracy. Furthermore, evaluations show the system's accuracy could be promisingly increased with the enrichment of metadata.
Xavier Anguera Miró, Xiaoou Chen, Deshun Yang
ICME3
2009 Learning element similarity matrix for semi-structured document analysis
Jianwu Yang, William Kwok-Wai Cheung, Xiaoou Chen
Knowl. Inf. Syst.3
2006 Audio similarity measure by graph modeling and matching
abstract
This paper proposes a new approach for the similarity measure and ranking of audio clips by graph modeling and matching. Instead of using frame-based or salient-based features to measure the acoustical similarity of audio clips, segment-based similarity is proposed. The novelty of our approach lies in two aspects: segment-based representation, and the similarity measure and ranking based on four kinds of similarity factors. In segmentbased representation, segments not only capture the change property of audio clip, but also keep and present the change relation and temporal order of audio features. In the similarity measure and ranking, four kinds of similarity factors: acoustical, granularity, temporal order and interference are progressively and jointly measured by optimal matching and dynamic programming, which guarantee the comprehensive and sufficient similarity measure between two audio clips. The experimental result shows that the proposed approach is better than some existing methods in terms of retrieval and ranking capabilities.
Yuxin Peng 0001, Chong-Wah Ngo, Cuihua Fang, Xiaoou Chen, Jianguo Xiao
ACM Multimedia4
2005 Integrating Element and Term Semantics for Similarity-Based XML Document Clustering
abstract
Structured link vector model (SLVM) is a recently proposed document representation that takes into account both structural and semantic information for measuring XML document similarity. Its formulation includes an element similarity matrix for capturing the semantic similarity between XML elements - the structural components of XML documents. In this paper, instead of applying heuristics to define the similarity matrix, we proposed to learn the matrix using pair wise similar training data in an iterative manner. In addition, we extended SLVM to SLVM-LSI by incorporating term semantics into SLVM using latent semantic indexing, with the element similarity related properties of the original SLVM preserved. For performance evaluation, we applied SLVM-LSI to similarity-based clustering of two XML datasets and the proposed SLVM-LSI was found to significantly outperform the conventional vector space model and the edit-distance based methods. The similarity matrix, obtained as a byproduct via the learning, can provide higher level knowledge about the semantic relationship between the XML elements.
Jianwu Yang, William Kwok-Wai Cheung, Xiaoou Chen
Web Intelligence3
2002 A Semi-Structured Document Model for Text Mining
Jianwu Yang, Xiaoou Chen
J. Comput. Sci. Technol.2