Li Su 0004

dblp:05/365-4 · DBLP profile ↗
← Back
32ranked-venue papers
4as first author
8since 2021 · last 2024
0000-0003-4275-8832ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 MOSA: Music Motion With Semantic Annotation Dataset for Cross-Modal Music Processing
abstract
In cross-modal music processing, translation between visual, auditory, and semantic content opens up new possibilities as well as challenges. The construction of such a transformative scheme depends upon a benchmark corpus with a comprehensive data infrastructure. In particular, the assembly of a large-scale cross-modal dataset presents major challenges. In this paper, we present the MOSA (Music mOtion with Semantic Annotation) dataset, which contains high quality 3-D motion capture data, aligned audio recordings, and note-by-note semantic annotations of pitch, beat, phrase, dynamic, articulation, and harmony for 742 professional music performances by 23 professional musicians, comprising more than 30 hours and 570 K notes of data. To our knowledge, this is the largest cross-modal music dataset with note-level annotations to date. To demonstrate the usage of the MOSA dataset, we present several innovative cross-modal music information retrieval (MIR) and musical content generation tasks, including the detection of beats, downbeats, phrases, and expressive contents from audio, video and motion data, and the generation of musicians' body motion from given music audio. The dataset and codes are available alongside this publication (https://github.com/yufenhuang/MOSA-Music-mOtion-and-Semantic-Annotation-dataset).
Yu-Fen Huang, Nikki Moran, Simon Coleman, Jon Kelly, Shun-Hwa Wei, Po-Yin Chen, Yun-Hsin Huang, Tsung-Ping Chen, Yu-Chia Kuo, Yu-Chi Wei, Chih-Hsuan Li, Da-Yu Huang, Hsuan-Kai Kao, Ting-Wei Lin, Li Su 0004
IEEE ACM Trans. Audio Speech Lang. Process.15
2023 Zero-Shot Singing Voice Synthesis from Musical Score
abstract
Zero-shot singing voice synthesis (SVS), the task to synthesize the singing voice of an arbitrary target singer, has gained increasing attentions in the past few years. Several recently proposed systems have demonstrated promising results on this task. However, these systems require detailed musical features at the frame level as the musical content. To deal with this issue, we propose a model that performs zero-shot SVS with only musical score as the musical content condition. To help model training, we build an acoustic encoder that extracts linguistic features from audio, and train it with the lyrics transcription objective. The output of the acoustic encoder serves as an alternative to the musical score, allowing the SVS model to learn from weakly labeled data. Results suggest that the proposed method outperforms baseline semi-supervised method in both subjective and objective tests.
Jun-You Wang, Hung-yi Lee, Jyh-Shing Roger Jang, Li Su 0004
ASRU4
2023 Adapting Pretrained Speech Model for Mandarin Lyrics Transcription and Alignment
abstract
The tasks of automatic lyrics transcription and lyrics alignment have witnessed significant performance improvements in the past few years. However, most of the previous works only focus on English in which large-scale datasets are available. In this paper, we address lyrics transcription and alignment of polyphonic Mandarin pop music in a low-resource setting. To deal with the data scarcity issue, we adapt pretrained Whisper model and fine-tune it on a monophonic Mandarin singing dataset. With the use of data augmentation and source separation model, results show that the proposed method achieves a character error rate of less than 18% on a Mandarin polyphonic dataset for lyrics transcription, and a mean absolute error of 0.071 seconds for lyrics alignment. Our results demonstrate the potential of adapting a pretrained speech model for lyrics transcription and alignment in low-resource scenarios.
Jun-You Wang, Chon-In Leong, Li Su 0004, Jyh-Shing Roger Jang
ASRU4
2023 Decoding Musical Pitch from Human Brain Activity with Automatic Voxel-Wise Whole-Brain FMRI Feature Selection
abstract
Decoding models seek to infer stimulus or task information from neural activity and play a central role in brain-computer interfaces. However, the high spatial resolution of fMRI means that the number of available features far exceeds the number of trials in a typical experiment. Although a common approach is to restrict features to a priori-defined regions of interest, related information present in other brain regions are consequently omitted. Here, we propose a two-stage thresholding approach that automatically pools relevant voxels from the whole-brain to enhance decoding performance. Testing on an fMRI dataset of 20 subjects, we show that our approach significantly improves regression performance in decoding musical pitch value by 2-fold compared to restricting voxels to the auditory cortex. We further examine properties of the selected voxels, and compare performance between random forest and convolutional neural network decoders.
Vincent K. M. Cheung, Yueh-Po Peng, Jing-Hua Lin, Li Su 0004
ICASSP4
2021 Improving Automatic Drum Transcription Using Large-Scale Audio-to-Midi Aligned Data
abstract
One of the major challenges in Automatic Drum Transcription (ADT) research is the lack of large-scale labeled dataset featuring audio with polyphonic mixtures; this limitation around data availability greatly impedes the progress of data-driven approaches in the context of ADT. To tackle this issue, we propose a semi-automatic way of compiling a labeled dataset using the audio-to-MIDI alignment technique. The resulting dataset consists of 1565 polyphonic mixtures of music with audio-aligned MIDI ground truth. To validate the quality and generality of this dataset, an ADT model based on Convolutional Neural Network (CNN) is trained and evaluated on several publicly available datasets. The evaluation results suggest that our proposed model, which is trained solely on the compiled dataset, compares favorably with the state-of-the-art ADT systems. The result also implies the possibility of leveraging audio-to-MIDI alignment in creating datasets for a broader range of audio related tasks.
I-Chieh Wei, Chih-Wei Wu, Li Su 0004
ICASSP3
2021 Positioning Left-hand Movement in Violin Performance: A System and User Study of Fingering Pattern Generation
abstract
Organizing fingerings, i.e., choosing which fingers to press on which positions and strings, is a crucial step for playing the violin. As the violin fingering comprises several components, the mapping of a musical phrase to the corresponding fingering arrangement is not unique, and it requires comprehensive musical knowledge for organizing adequate fingerings. In this paper, we study the human-machine cooperative approach to the generation of violin fingering, aiming to build an intelligent system which can provide multiple generation paths and yield adaptable fingering arrangements. For this sake, we compile a new dataset with fingering annotations of multiple versions of performance, propose a deep neural network with conditions on the left-hand movement for fingering generation, and conduct an in-depth user study for detailed responses. Result shows that the proposed system can yield various fingering arrangements according to different performance requirements, though a single generation may not satisfy all the requirements at a time. This highlights the importance of multi-path and human-in-the-loop architecture for violin fingering generation.
Yi-Hsin Jen, Tsung-Ping Chen, Shih-Wei Sun, Li Su 0004
IUI4
2021 ReconVAT: A Semi-Supervised Automatic Music Transcription Framework for Low-Resource Real-World Data
abstract
Most of the current supervised automatic music transcription (AMT) models lack the ability to generalize. This means that they have trouble transcribing real-world music recordings from diverse musical genres that are not presented in the labelled training data. In this paper, we propose a semi-supervised framework, ReconVAT, which solves this issue by leveraging the huge amount of available unlabelled music recordings. The proposed ReconVAT uses reconstruction loss and virtual adversarial training. When combined with existing U-net models for AMT, ReconVAT achieves competitive results on common benchmark datasets such as MAPS and MusicNet. For example, in the few-shot setting for the string part version of MusicNet, ReconVAT achieves F1-scores of 61.0% and 41.6% for the note-wise and note-with-offset-wise metrics respectively, which translates into an improvement of 22.2% and 62.5% compared to the supervised baseline model. Our proposed framework also demonstrates the potential of continual learning on new data, which could be useful in real-world applications whereby new data is constantly available.
Kin Wai Cheuk, Dorien Herremans, Li Su 0004
ACM Multimedia3
2021 Actions Speak Louder than Listening: Evaluating Music Style Transfer based on Editing Experience
Wei Tsung Lu, Meng-Hsuan Wu, Yuh-Ming Chiu, Li Su 0004
ACM Multimedia4
2020 Body Movement Generation for Expressive Violin Performance Applying Neural Networks
abstract
Generating body movements based on given music audio recordings is an emerging research topic. This problem remains challenging particularly for string instruments, considering the fact that the relationship between the musical note sequences and the body movement sequences in string instruments does not have an one-to-one correspondence and is highly context-dependent. In this paper, we take a divide-and-rule approach to tackle the multifaceted characteristics of musical movement, and propose a framework for generating violinists' body movements. Both objective and subjective evaluations show that the proposed framework improves the stability as well as the perceptual quality of the generation outputs by using the task-specific models for bowing and expressive movement. To the best of our knowledge, this work represents the first attempt to generate violinists' body movements considering music expression.
Jun-Wei Liu, Hung-Yi Lin, Yu-Fen Huang, Hsuan-Kai Kao, Li Su 0004
ICASSP5
2020 Temporally Guided Music-to-Body-Movement Generation
abstract
This paper presents a neural network model to generate virtual violinist's 3-D skeleton movements from music audio. Improved from the conventional recurrent neural network models for generating 2-D skeleton data in previous works, the proposed model incorporates an encoder-decoder architecture, as well as the self-attention mechanism to model the complicated dynamics in body movement sequences. To facilitate the optimization of self-attention model, beat tracking is applied to determine effective sizes and boundaries of the training examples. The decoder is accompanied with a refining network and a bowing attack inference mechanism to emphasize the right-hand behavior and bowing attack timing. Both objective and subjective evaluations reveal that the proposed model outperforms the state-of-the-art methods. To the best of our knowledge, this work represents the first attempt to generate 3-D violinists? body movements considering key features in musical body movement.
Hsuan-Kai Kao, Li Su 0004
ACM Multimedia2
2020 Crossing You in Style: Cross-modal Style Transfer from Music to Visual Arts
abstract
Music-to-visual style transfer is a challenging yet important cross-modal learning problem in the practice of creativity. Its major difference from the traditional image style transfer problem is that the style information is provided by music rather than images. Assuming that musical features can be properly mapped to visual contents through semantic links between the two domains, we solve the music-to-visual style transfer problem in two steps: music visualization and style transfer. The music visualization network utilizes an encoder-generator architecture with a conditional generative adversarial network to generate image-based music representations from music data. This network is integrated with an image style transfer method to accomplish the style transfer process. Experiments are conducted on WikiArt-IMSLP, a newly compiled dataset including Western music recordings and paintings listed by decades. By utilizing such a label to learn the semantic connection between paintings and music, we demonstrate that the proposed framework can generate diverse image style representations from a music piece, and these representations can unveil certain art forms of the same era. Subjective testing results also emphasize the role of the era label in improving the perceptual quality on the compatibility between music and visual content.
Cheng-Che Lee, Wan-Yi Lin, Yen-Ting Shih, Pei-Yi Kuo, Li Su 0004
ACM Multimedia5
2020 A Human-Computer Duet System for Music Performance
abstract
Virtual musicians have become a remarkable phenomenon in the contemporary multimedia arts. However, most of the virtual musicians nowadays have not been endowed with abilities to create their own behaviors, or to perform music with human musicians. In this paper, we firstly create a virtual violinist, who can collaborate with a human pianist to perform chamber music automatically without any intervention. The system incorporates the techniques from various fields, including real-time music tracking, pose estimation, and body movement generation. In our system, the virtual musician's behavior is generated based on the given music audio alone, and such a system results in a low-cost, efficient and scalable way to produce human and virtual musicians' co-performance. The proposed system has been validated in public concerts. Objective quality assessment approaches and possible ways to systematically improve the system are also discussed.
Yuen-Jen Lin, Hsuan-Kai Kao, Yih-Chih Tseng, Ming Tsai, Li Su 0004
ACM Multimedia5
2020 Multi-Instrument Automatic Music Transcription With Self-Attention-Based Instance Segmentation
abstract
Multi-instrument automatic music transcription (AMT) is a critical but less investigated problem in the field of music information retrieval (MIR). With all the difficulties faced by traditional AMT research, multi-instrument AMT needs further investigation on high-level music semantic modeling, efficient training methods for multiple attributes, and a clear problem scenario for system performance evaluation. In this article, we propose a multi-instrument AMT method, with signal processing techniques specifying pitch saliency, novel deep learning techniques, and concepts partly inspired by multi-object recognition, instance segmentation, and image-to-image translation in computer vision. The proposed method is flexible for all the sub-tasks in multi-instrument AMT, including multi-instrument note tracking, a task that has rarely been investigated before. State-of-the-art performance is also reported in the sub-task of multi-pitch streaming.
Yu-Te Wu, Berlin Chen, Li Su 0004
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Play as You Like: Timbre-Enhanced Multi-Modal Music Style Transfer
abstract
Style transfer of polyphonic music recordings is a challenging task when considering the modeling of diverse, imaginative, and reasonable music pieces in the style different from their original one. To achieve this, learning stable multi-modal representations for both domain-variant (i.e., style) and domaininvariant (i.e., content) information of music in an unsupervised manner is critical. In this paper, we propose an unsupervised music style transfer method without the need for parallel data. Besides, to characterize the multi-modal distribution of music pieces, we employ the Multi-modal Unsupervised Image-to-Image Translation (MUNIT) framework in the proposed system. This allows one to generate diverse outputs from the learned latent distributions representing contents and styles. Moreover, to better capture the granularity of sound, such as the perceptual dimensions of timbre and the nuance in instrument-specific performance, cognitively plausible features including mel-frequency cepstral coefficients (MFCC), spectral difference, and spectral envelope, are combined with the widely-used mel-spectrogram into a timbreenhanced multi-channel input representation. The Relativistic average Generative Adversarial Networks (RaGAN) is also utilized to achieve fast convergence and high stability. We conduct experiments on bilateral style transfer tasks among three different genres, namely piano solo, guitar solo, and string quartet. Results demonstrate the advantages of the proposed method in music style transfer with improved sound quality and in allowing users to manipulate the output.
Chien-Yu Lu, Min-Xin Xue, Chia-Che Chang, Che-Rung Lee, Li Su 0004
AAAI5
2019 A Streamlined Encoder/decoder Architecture for Melody Extraction
abstract
Melody extraction in polyphonic musical audio is important for music signal processing. In this paper, we propose a novel streamlined encoder/decoder network that is designed for the task. We make two technical contributions. First, drawing inspiration from a state-of-the-art model for semantic pixel-wise segmentation, we pass through the pooling indices between pooling and un-pooling layers to localize the melody in frequency. We can achieve result close to the state-of-the-art with much fewer convolutional layers and simpler convolution modules. Second, we propose a way to use the bottleneck layer of the network to estimate the existence of a melody line for each time frame, and make it possible to use a simple argmax function instead of ad-hoc thresholding to get the final estimation of the melody line. Our experiments on both vocal melody extraction and general melody extraction validate the effectiveness of the proposed model.
Tsung-Han Hsieh, Li Su 0004, Yi-Hsuan Yang
ICASSP2
2019 Polyphonic Music Transcription with Semantic Segmentation
abstract
The multi-instrument transcription task refers to joint recognition of instrument and pitch of every event in polyphonic music signals generated by one or more classes of music instruments. In this paper, we leverage multi-object semantic segmentation techniques to solve this problem. We design a time-frequency representation, which has multiple channels to jointly represent the harmonic structure and pitch saliency of a pitch activation. The transcription task therefore becomes a pixel-wise multi-task classification problem including pitch activity detection and instrument recognition. Experiments on both single- and multi-instrument data verify the competitiveness of the proposed method.
Yu-Te Wu, Berlin Chen, Li Su 0004
ICASSP3
2018 Singing Voice Correction Using Canonical Time Warping
abstract
Expressive singing voice correction is an appealing but challenging problem. A robust time-warping algorithm which synchronizes two singing recordings can provide a promising solution. We thereby propose to address the problem by canonical time warping (CTW) which aligns amateur singing recordings to professional ones. A new pitch contour is generated given the alignment information, and a pitch-corrected singing is synthesized back through the vocoder. The objective evaluation shows that CTW is robust against pitch-shifting and time-stretching effects, and the subjective test demonstrates that CTW prevails the other methods including DTW and the commercial auto-tuning software. Finally, we demonstrate the applicability of the proposed method in a practical, real-world scenario.
Yin-Jyun Luo, Ming-Tso Chen, Tai-Shih Chi, Li Su 0004
ICASSP4
2018 Vocal Melody Extraction Using Patch-Based CNN
abstract
A patch-based convolutional neural network (CNN) model presented in this paper for vocal melody extraction in polyphonic music is inspired from object detection in image processing. The input of the model is a novel time-frequency representation which enhances the pitch contours and suppresses the harmonic components of a signal. This succinct data representation and the patch-based CNN model enable an efficient training process with limited labeled data. Experiments on various datasets show excellent speed and competitive accuracy comparing to other deep learning approaches.
Li Su 0004
ICASSP1
2018 Automatic Music Transcription Leveraging Generalized Cepstral Features and Deep Learning
abstract
Spectral features are limited in modeling musical signals with multiple concurrent pitches due to the challenge to suppress the interference of the harmonic peaks from one pitch to another. In this paper, we show that using multiple features represented in both the frequency and time domains with deep learning modeling can reduce such interference. These features are derived systematically from conventional pitch detection functions that relate to one another through the discrete Fourier transform and a nonlinear scaling function. Neural networks modeled with these features outperform state-of-the-art methods while using less training data.
Yu-Te Wu, Berlin Chen, Li Su 0004
ICASSP3
2018 Online Music Performance Tracking Using Parallel Dynamic Time Warping
abstract
Resource allocation is a critical issue in the implementation of a portable, low-latency and efficient system for interactive experience. To address this, we propose the parallel dynamic time warping (PDTW), that utilizes multi-thread computation on a modern multi-core system, for online audio-to-audio music alignment. We also discuss the evaluation methodology that benchmarks the trade-offs among latency, accuracy of alignment, and computing resource. The proposed system utilizes multiple DTW alignment processes and parallel computing technique to reduce the processing time, as well as to improve both the alignment accuracy and robustness to tempo variation. Evaluation on several datasets with artificial time stretching exhibits the capacity of the system in enriched concert experience.
I-Chieh Wei, Li Su 0004
MMSP2
2018 Monaural Source Separation Using Ramanujan Subspace Dictionaries
abstract
Most source separation algorithms are implemented as spectrogram decomposition. In contrast, time-domain source separation is less investigated since there is a lack of an efficient signal representation that facilitates decomposing oscillatory components of a signal directly in the time domain. In this letter, we utilize the Ramanujan subspace and the nested periodic subspace to address this issue, by constructing a parametric dictionary that emphasizes period information with less redundancy. Methods including iterative subspace projection and convolutional sparse coding can decompose a mixture into signals with distinct oscillation periods according to the dictionary. Experiments on score-informed source separation show that the proposed method is competitive to the state-of-the-art, frequency-domain approaches when the provided pitch information and the signal parameters are the same.
Hsueh-Wei Liao, Li Su 0004
IEEE Signal Process. Lett.2
2017 Polyphonic piano note transcription with non-negative matrix factorization of differential spectrogram
abstract
Automatic music transcription is usually approached by using a time-frequency (TF) representation such as the short-time Fourier transform (STFT) spectrogram or the constant-Q transform. In this paper, we propose a novel yet simple TF representation that capitalizes the effectiveness of spectral flux features in highlighting note onset times. We refer to this representation as the differential spectrogram and investigate its usefulness for note-level piano transcription using two different non-negative matrix factorization (NMF) algorithms. Experiments on the MAPS ENSTDkCl dataset validate the advantages of the differential spectrogram over the STFT spectrogram for this task. Moreover, by adapting a state-of-the-art convolutional NMF algorithm with the differential spectrogram, we can achieve even better accuracy than the state-of-the-art on this dataset. Our analysis shows that the new representation suppresses unwanted TF patterns and performs particularly well in improving the recall rate.
Lufei Gao, Li Su 0004, Yi-Hsuan Yang, Tan Lee
ICASSP2
2017 Multi-pitch streaming of interwoven streams
abstract
In this paper, we discuss the multipitch streaming (MPS) problem for a multi-source audio signal having interweaving pitch contours. We propose two approaches to tackle this challenge, one relates to a feature extracted from the energy levels distributed in multi-channel recordings for better characterization of the source, and the other uses particle swarm optimization (PSO) to enlarge the search space and alleviate the initialization problem in constrained clustering of the features representing different sources. Experiments on music and speech samples having highly interweaving pitch contours are presented to assess its effectiveness.
Chih Yi Kuan, Li Su 0004, Yu-Hao Chin, Jia-Ching Wang
ICASSP2
2017 Automatic conversion of Pop music into chiptunes for 8-bit pixel art
abstract
In this paper, we propose an audio mosaicing method that converts Pop songs into a specific music style called “chiptune,” or “8-bit music.” The goal is to reproduce Pop songs by using the sound of the chips on the old game consoles in 1980s/1990s. The proposed method goes through a procedure that first analyzes the pitches of an incoming Pop song in the frequency domain, and then synthesizes the song with template waveforms in the time domain to make it sound like 8-bit music. Because a Pop song is usually composed of the vocal melody and the instrumental accompaniment, in the analysis stage we use a singing voice separation algorithm to separate the vocals from the instruments, and then apply different pitch detection algorithms to transcribe the two separated sources. We validate through a subjective listening test that the proposed method creates much better 8-bit music than existing nonnegative matrix factorization based methods can do. Moreover, we find that synthesis in the time domain is important for this task.
Shih-Yang Su, Cheng-Kai Chiu, Li Su 0004, Yi-Hsuan Yang
ICASSP3
2016 Monaural Music Source Separation Using Convolutional Sparse Coding
abstract
We present a comprehensive performance study of a new time-domain approach for estimating the components of an observed monaural audio mixture. Unlike existing time-frequency approaches that use the product of a set of spectral templates and their corresponding activation patterns to approximate the spectrogram of the mixture, the proposed approach uses the sum of a set of convolutions of estimated activations with prelearned dictionary filters to approximate the audio mixture directly in the time domain. The approximation problem can be solved by an efficient convolutional sparse coding algorithm. The effectiveness of this approach for source separation of musical audio has been demonstrated in our prior work, but under rather restricted and controlled conditions, requiring the musical score of the mixture being informed a priori and little mismatch between the dictionary filters and the source signals. In this paper, we report an evaluation that considers wider, and more practical, experimental settings. This includes the use of an audio-based multipitch estimation algorithm to replace the musical score, and an external dataset of audio single notes to construct the dictionary filters. Our result shows that the proposed approach remains effective with a larger dictionary, and compares favorably with the state-of-the-art nonnegative matrix factorization approach. However, in the absence of the score and in the case of a small dictionary, our approach may not be better.
Ping-Keng Jao, Li Su 0004, Yi-Hsuan Yang, Brendt Wohlberg
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 Vocal activity informed singing voice separation with the iKala dataset
abstract
A new algorithm is proposed for robust principal component analysis with predefined sparsity patterns. The algorithm is then applied to separate the singing voice from the instrumental accompaniment using vocal activity information. To evaluate its performance, we construct a new publicly available iKala dataset that features longer durations and higher quality than the existing MIR-1K dataset for singing voice separation. Part of it will be used in the MIREX Singing Voice Separation task. Experimental results on both the MIR-1K dataset and the new iKala dataset confirmed that the more informed the algorithm is, the better the separation results are.
Tak-Shing Chan, Tzu-Chun Yeh, Zhe-Cheng Fan, Hung-Wei Chen, Li Su 0004, Yi-Hsuan Yang, Jyh-Shing Roger Jang
ICASSP5
2015 Musical Onset Detection Using Constrained Linear Reconstruction
abstract
This letter presents a multi-frame extension of the well-known spectral flux method for unsupervised musical onset detection. Instead of comparing only the spectral content of two frames, the proposed method takes into account a wider temporal context to evaluate the dissimilarity between a given frame and its previous frames. More specifically, the dissimilarity is measured by using the previous frames to obtain a linear reconstruction of the given frame, and then calculating the rectified, l2-norm reconstruction error. Evaluation on a dataset comprising 2,169 onset events of 12 instruments shows that this simple idea works fairly well. When a non-negativity constraint is imposed in the linear reconstruction, the proposed method can outperform the state-of-the-art unsupervised method SuperFlux by 2.9% in F-score. Moreover, the proposed method is particularly effective for instruments with soft onsets, such as violin, cello, and ney. The proposed method is efficient, easy to implement, and is applicable to scenarios of online onset detection.
Che-Yuan Liang, Li Su 0004, Yi-Hsuan Yang
IEEE Signal Process. Lett.2
2015 Combining Spectral and Temporal Representations for Multipitch Estimation of Polyphonic Music
abstract
Due to the difficulty of creating pitch-labeled training data that cover the rich diversity found in music signals, unsupervised feature-based approaches derived from signal processing and feature design remain critical for multipitch estimation (MPE) of polyphonic music. While a large number of feature representations have been proposed in the literature, an effective means of combining different domains of features for MPE is still needed. In this paper, we propose a novel approach, referred to as combined frequency and periodicity (CFP), that detects pitches according to the agreement of a harmonic series in the frequency domain and a subharmonic series in the lag (quefrency) domain. This approach nicely aggregates the complementary advantages of the two feature domains in different frequency ranges, and improves the robustness of the pitch detection function to the interference of the overtones of simultaneous pitches. We report a comprehensive evaluation that compares CFP against three state-of-the-art approaches using three MPE datasets and four symphonies. The evaluation is characteristic of the coverage and complexity of music (in terms of instrument type and degree of polyphony). In addition, we also evaluate the performance of the MPE approaches when a number of audio degradations are applied. Results show that the proposed unsupervised method performs consistently well across the types of Western polyphonic music considered, and is robust to audio degradations such as high-pass filtering and MP3 compression.
Li Su 0004, Yi-Hsuan Yang
IEEE ACM Trans. Audio Speech Lang. Process.1
2014 Sparse cepstral codes and power scale for instrument identification
abstract
This paper presents a novel feature representation called sparse cepstral codes for instrument identification. We first motivate the approach by discussing why cepstrum is suitable for instrument identification. Then we propose the use of sparse coding and power normalization to derive compact codes that better represent the information of the cepstrum. Our evaluation on both uni-source and multi-source instrument identification tasks show that the proposed feature leads to significantly better accuracy than existing methods. We further show that cepstrum obtained from power-scaled spectrum can do better than typical cepstrum especially in multi-source signal. The proposed system achieves 0.955 F-score in uni-source dataset and 0.688 F-score in multi-source dataset.
Li-Fan Yu, Li Su 0004, Yi-Hsuan Yang
ICASSP2
2014 Sparse modeling of magnitude and phase-derived spectra for playing technique classification
abstract
Computational modeling of musical timbre is important for a variety of music information retrieval applications. While considerable progress has been made to recognize musical genres and instruments, relatively little attention has been paid to modeling playing techniques, which affect timbre in more subtle ways. In this paper, we contribute to this area of research by systematically evaluating various audio features and processing methods for multi-class playing technique classification, considering up to nine distinct playing techniques of bowed string instruments. Specifically, a collection of 6,759 chamber-recorded single notes of four bowed string instruments and a collection of 33 real-world solo violin recordings are used in the evaluation. Our evaluation shows that using sparse features extracted from the magnitude spectra and phase derivatives including group delay function (GDF) and instantaneous frequency deviation (IFD) leads to significantly better performance than using a combination of state-of-the-art temporal, spectral, cepstral and harmonic feature descriptors. For playing technique classification of violin singe notes, the former approach attains 0.915 macro-average F-score under a tenfold cross validation setting, while the latter only attains 0.835. Moreover, sparse modeling of magnitude and phase-derived spectra also performs well for single-note joint instrument-technique classification (F-score 0.770) and for playing technique classification of real-world violin solos (F-score 0.547). We find that phase information is particularly important in discriminating playing techniques with subtle differences, such as playing with different bowing positions (i.e., normal, sul tasto, and sul ponticello). A systematic investigation of the effect of parameters such as window sizes, hop factors, window types for phase-derived features is also reported to provide more insights.
Li Su 0004, Hsin-Ming Lin, Yi-Hsuan Yang
IEEE ACM Trans. Audio Speech Lang. Process.1
2014 A Systematic Evaluation of the Bag-of-Frames Representation for Music Information Retrieval
abstract
There has been an increasing attention on learning feature representations from the complex, high-dimensional audio data applied in various music information retrieval (MIR) problems. Unsupervised feature learning techniques, such as sparse coding and deep belief networks have been utilized to represent music information as a term-document structure comprising of elementary audio codewords. Despite the widespread use of such bag-of-frames (BoF) model, few attempts have been made to systematically compare different component settings. Moreover, whether techniques developed in the text retrieval community are applicable to audio codewords is poorly understood. To further our understanding of the BoF model, we present in this paper a comprehensive evaluation that compares a large number of BoF variants on three different MIR tasks, by considering different ways of low-level feature representation, codebook construction, codeword assignment, segment-level and song-level feature pooling, tf-idf term weighting, power normalization, and dimension reduction. Our evaluations lead to the following findings: 1) modeling music information by two levels of abstraction improves the result for difficult tasks such as predominant instrument recognition, 2) tf-idf weighting and power normalization improve system performance in general, 3) topic modeling methods such as latent Dirichlet allocation does not work for audio codewords.
Li Su 0004, Chin-Chia Michael Yeh, Jen-Yu Liu, Ju-Chiang Wang, Yi-Hsuan Yang
IEEE Trans. Multim.1
2013 Dual-layer bag-of-frames model for music genre classification
abstract
This paper concerns the development of a music dictionary-based model for summarizing local feature descriptors computed over time. Comparing to a holistic representation, this text-like, bag-of-frames representation better captures the rich and time-varying information of music. However, the dictionary used in classical bag-of-frames model only captures frame-level elements of the music; thus, there exists a semantic gap between the dictionary element and commonly seen music description. In order to reduce the gap, a new feature representation called dual-layer bag-of-frames is proposed in this paper. It models the music with a two layer structure, where the first-layer dictionary captures the frame-level characteristics, and the second-layer dictionary captures the segment-level semantics. This hierarchical structure resembles the alphabet-word-document structure of text. Our result demonstrates that the proposed dual-layer bag-of-frames feature achieves state-of-the-art accuracy of music genre classification. The classification accuracy for the GTZAN benchmark reaches 86.7% with dictionary trained from GTZAN, and 83.6% with dictionary trained from another data set USPOP.
Chin-Chia Michael Yeh, Li Su 0004, Yi-Hsuan Yang
ICASSP2