EDBT 2026 Demo / reviewers in the wild / expert
Yi-Hsuan Yang
dblp:64/6611
· DBLP profile ↗
126ranked-venue papers
18as first author
16since 2021 · last 2025
0000-0002-2724-6161ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 90 · 12 first-author · 11 since 2021Artificial intelligence and machine learning · 38 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 12 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Systems, architecture and hardware · 1Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DDSP Guitar Amp: Interpretable Guitar Amplifier ModelingabstractNeural network models for guitar amplifier emulation, while being effective, often demand high computational cost and lack interpretability. Drawing ideas from physical amplifier design, this paper aims to address these issues with a new differentiable digital signal processing (DDSP)-based model, called "DDSP guitar amp," that models the four components of a guitar amp (i.e., preamp, tone stack, power amp, and output transformer) using specific DSP-inspired designs. With a set of time- and frequency-domain metrics, we demonstrate that DDSP guitar amp achieves performance comparable with that of black-box baselines while requiring less than 10% of the computational operations per audio sample, thereby holding greater potential for usages in real-time applications. Yen-Tung Yeh, Yu-Hua Chen, Yuan-Chiao Cheng, Jui-Te Wu, Jun-Jie Fu, Yi-Fan Yeh, Yi-Hsuan Yang |
ICASSP | 7 |
| 2025 | MuseControlLite: Multifunctional Music Generation with Lightweight ConditionersabstractWe propose MuseControlLite, a lightweight mechanism designed to fine-tune text-to-music generation models for precise conditioning using various time-varying musical attributes and reference audio signals. The key finding is that positional embeddings, which have been seldom used by text-to-music generation models in the conditioner for text conditions, are critical when the condition of interest is a function of time. Using melody control as an example, our experiments show that simply adding rotary positional embeddings to the decoupled cross-attention layers increases control accuracy from 56.6% to 61.1%, while requiring 6.75 times fewer trainable parameters than state-of-the-art fine-tuning mechanisms, using the same pre-trained diffusion Transformer model of Stable Audio Open. We evaluate various forms of musical attribute control, audio inpainting, and audio outpainting, demonstrating improved controllability over MusicGen-Large and Stable Audio Open ControlNet at a significantly lower fine-tuning cost, with only 85M trainable parameters. Source code, model checkpoints, and demo examples are available at: https://MuseControlLite.github.io/web/ Fang-Duo Tsai, Shih-Lun Wu, Weijaw Lee, Sheng-Ping Yang, Bo-Rui Chen, Yi-Hsuan Yang |
ICML | 7 |
| 2025 | METEOR: Melody-aware Texture-controllable Symbolic Music Re-Orchestration via Transformer VAEabstractRe-orchestration is the process of adapting a music piece for a different set of instruments. By altering the original instrumentation, the orchestrator often modifies the musical texture while preserving a recognizable melodic line and ensures that each part is playable within the technical and expressive capabilities of the chosen instruments. In this work, we propose METEOR, a model for generating Melody-aware Texture-controllable re-Orchestration with a Transformer-based variational auto-encoder (VAE). This model performs symbolic instrumental and textural music style transfers with a focus on melodic fidelity and controllability. We allow bar- and track-level controllability of the accompaniment with various textural attributes while keeping a homophonic texture. With both subjective and objective evaluations, we show that our model outperforms style transfer models on a re-orchestration task in terms of generation quality and controllability. Moreover, it can be adapted for a lead sheet orchestration task as a zero-shot learning model, achieving performance comparable to a model specifically trained for this task. Dinh-Viet-Toan Le, Yi-Hsuan Yang |
IJCAI | 2 |
| 2024 | PiCoGen: Generate Piano Covers with a Two-stage ApproachabstractCover song generation stands out as a popular way of music making in the music-creative community. In this study, we introduce Piano Cover Generation (PiCoGen), a two-stage approach for automatic cover song generation that transcribes the melody line and chord progression of a song given its audio recording, and then uses the resulting lead sheet as the condition to generate a piano cover in the symbolic domain. This approach is advantageous in that it does not required paired data of covers and their original songs for training. Compared to an existing approach that demands such paired data, our evaluation shows that PiCoGen demonstrates competitive or even superior performance across songs of different musical genres. Chih-Pin Tan, Shuen-Huei Guan, Yi-Hsuan Yang |
ICMR | 3 |
| 2023 | Compose & Embellish: Well-Structured Piano Performance Generation via A Two-Stage ApproachabstractEven with strong sequence models like Transformers, generating expressive piano performances with long-range musical structures remains challenging. Meanwhile, methods to compose well-structured melodies or lead sheets (melody + chords), i.e., simpler forms of music, gained more success. Observing the above, we devise a two-stage Transformer-based framework that Composes a lead sheet first, and then Embellishes it with accompaniment and expressive touches. Such a factorization also enables pretraining on non-piano data. Our objective and subjective experiments show that Compose & Embellish shrinks the gap in structureness between a current state of the art and real performances by half, and improves other musical aspects such as richness and coherence as well. Shih-Lun Wu, Yi-Hsuan Yang |
ICASSP | 2 |
| 2023 | Local Periodicity-Based Beat Tracking for Expressive Classical Piano MusicabstractTo model the periodicity of beats, state-of-the-art beat tracking systems use “post-processing trackers” (PPTs) that rely on several empirically determined global assumptions for tempo transition, which work well for music with a steady tempo. For expressive classical music, however, these assumptions can be too rigid. With two large datasets of Western classical piano music, namely the Aligned Scores and Performances (ASAP) dataset and a dataset of Chopin's Mazurkas (Maz-5), we report on experiments showing the failure of existing PPTs to cope with local tempo changes, thus calling for new methods. In this paper, we propose a new local periodicity-based PPT, called predominant local pulse-based dynamic programming (PLPDP) tracking, that allows for more flexible tempo transitions. Specifically, the new PPT incorporates a method called “predominant local pulses” (PLP) in combination with a dynamic programming (DP) component to jointly consider the locally detected periodicity and beat activation strength at each time instant. Accordingly, PLPDP accounts for the local periodicity, rather than relying on a global tempo assumption. Compared to existing PPTs, PLPDP particularly enhances the recall values at the cost of a lower precision, resulting in an overall improvement of F1-score for beat tracking in ASAP (from 0.473 to 0.493) and Maz-5 (from 0.595 to 0.838). Ching-Yu Chiu, Meinard Müller, Matthew E. P. Davies, Alvin Wen-Yu Su, Yi-Hsuan Yang |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | MuseMorphose: Full-Song and Fine-Grained Piano Music Style Transfer With One Transformer VAEabstractTransformers and variational autoencoders (VAE) have been extensively employed for symbolic (e.g., MIDI) domain music generation. While the former boast an impressive capability in modeling long sequences, the latter allow users to willingly exert control over different parts (e.g., bars) of the music to be generated. In this paper, we are interested in bringing the two together to construct a single model that exhibits both strengths. The task is split into two steps. First, we equip Transformer decoders with the ability to accept segment-level, time-varying conditions during sequence generation. Subsequently, we combine the developed and tested in-attention decoder with a Transformer encoder, and train the resulting MuseMorphose model with the VAE objective to achieve style transfer of long pop piano pieces, in which users can specify musical attributes including rhythmic intensity and polyphony (i.e., harmonic fullness) they desire, down to the bar level. Experiments show that MuseMorphose outperforms recurrent neural network (RNN) based baselines on numerous widely-used metrics for style transfer tasks. Shih-Lun Wu, Yi-Hsuan Yang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Theme Transformer: Symbolic Music Generation With Theme-Conditioned TransformerabstractAttention-based Transformer models have been increasingly employed for automatic music generation. To condition the generation process of such a model with a user-specified sequence, a popular approach is to take that conditioning sequence as a priming sequence and ask a Transformer decoder to generate a continuation. However, thisprompt-based conditioningcannot guarantee that the conditioning sequence would develop or even simply repeat itself in the generated continuation. In this paper, we propose an alternative conditioning approach, calledtheme-based conditioning, that explicitly trains the Transformer to treat the conditioning sequence as a thematic material that has to manifest itself multiple times in its generation result. This is achieved with two main technical contributions. First, we propose a deep learning-based approach that uses contrastive representation learning and clustering to automatically retrieve thematic materials from music pieces in the training data. Second, we propose a novel gated parallel attention module to be used in a sequence-to-sequence (seq2seq) encoder/decoder architecture to more effectively account for a given conditioning thematic material in the generation process of the Transformer decoder. We report on objective and subjective evaluations of variants of the proposed Theme Transformer and the conventional prompt-based baseline, showing that our best model can generate, to some extent, polyphonic pop piano music with repetition and plausible variations of a given condition. Yi-Jen Shih, Shih-Lun Wu, Frank Zalkow, Meinard Müller, Yi-Hsuan Yang |
IEEE Trans. Multim. | 5 |
| 2022 | Towards Automatic Transcription of Polyphonic Electric Guitar Music: A New Dataset and a Multi-Loss Transformer ModelabstractIn this paper, we propose a new dataset named EGDB, that contains transcriptions of the electric guitar performance of 240 tablatures rendered with different tones. Moreover, we benchmark the performance of two well-known transcription models proposed originally for the piano on this dataset, along with a multi-loss Transformer model that we newly propose. Our evaluation on this dataset and a separate set of real-world recordings demonstrate the influence of timbre on the accuracy of guitar sheet transcription, the potential of using multiple losses for Transformers, as well as the room for further improvement for this task. Yu-Hua Chen, Wen-Yi Hsiao, Tsu-Kuang Hsieh, Jyh-Shing Roger Jang, Yi-Hsuan Yang |
ICASSP | 5 |
| 2022 | Automatic DJ Transitions with Differentiable Audio Effects and Generative Adversarial NetworksabstractA central task of a Disc Jockey (DJ) is to create a mixset of music with seamless transitions between adjacent tracks. In this paper, we explore a data-driven approach that uses a generative adversarial network to create the song transition by learning from real-world DJ mixes. The generator uses two differentiable digital signal processing components, an equalizer (EQ) and a fader, to mix two tracks selected by a data generation pipeline. The generator has to set the parameters of the EQs and fader in such a way that the resulting mix resembles real mixes created by human DJ, as judged by the discriminator counterpart. Result of a listening test shows that the model can achieve competitive results compared with a number of baselines. Bo-Yu Chen, Wei-Han Hsu, Wei-Hsiang Liao 0001, Marco A. Martínez Ramírez, Yuki Mitsufuji, Yi-Hsuan Yang |
ICASSP | 6 |
| 2022 | KaraSinger: Score-Free Singing Voice Synthesis with VQ-VAE Using Mel-SpectrogramsabstractIn this paper, we propose a novel neural network model called KaraSinger for a less-studied singing voice synthesis (SVS) task named score-free SVS, in which the prosody and melody are spontaneously decided by machine. KaraSinger comprises a vector-quantized variational autoencoder (VQ-VAE) that compresses the Mel-spectrograms of singing audio to sequences of discrete codes, and a language model (LM) that learns to predict the discrete codes given the corresponding lyrics. For the VQ-VAE part, we employ a Connectionist Temporal Classification (CTC) loss to encourage the discrete codes to carry phoneme-related information. For the LM part, we use location-sensitive attention for learning a robust alignment between the input phoneme sequence and the output discrete code. We keep the architecture of both the VQ-VAE and LM light-weight for fast training and inference speed. We validate the effectiveness of the proposed design choices using a proprietary collection of 550 English pop songs sung by multiple amateur singers. The result of a listening test shows that KaraSinger achieves high scores in intelligibility, musicality, and the overall quality. Chien-Feng Liao, Jen-Yu Liu, Yi-Hsuan Yang |
ICASSP | 3 |
| 2022 | An Analysis Method for Metric-Level Switching in Beat TrackingabstractFor expressive music, the tempo may change over time, posing challenges to tracking the beats by an automatic model. The model may first tap to the correct tempo, but then may fail to adapt to a tempo change, or switch between several incorrect but perceptually plausible ones (e.g., half- or double-tempo). Existing evaluation metrics for beat tracking do not reflect such behaviors, as they typically assume a fixed relationship between the reference beats and estimated beats. In this letter, we propose a new performance analysis method, called annotation coverage ratio (ACR), that accounts for a variety of possible metric-level switching behaviors of beat trackers. The idea is to derive sequences of modified reference beats of all metrical levels for every two consecutive reference beats, and compare every sequence of modified reference beats to the subsequences of estimated beats. We show via experiments on three datasets of different genres the usefulness of ACR when being utilized alongside existing metrics, and discuss the new insights that can be gained. Ching-Yu Chiu, Meinard Müller, Matthew E. P. Davies, Alvin Wen-Yu Su, Yi-Hsuan Yang |
IEEE Signal Process. Lett. | 5 |
| 2021 | Compound Word Transformer: Learning to Compose Full-Song Music over Dynamic Directed HypergraphsabstractTo apply neural sequence models such as the Transformers to music generation tasks, one has to represent a piece of music by a sequence of tokens drawn from a finite set of pre-defined vocabulary. Such a vocabulary usually involves tokens of various types. For example, to describe a musical note, one needs separate tokens to indicate the note’s pitch, duration, velocity (dynamics), and placement (onset time) along the time grid. While different types of tokens may possess different properties, existing models usually treat them equally, in the same way as modeling words in natural languages. In this paper, we present a conceptually different approach that explicitly takes into account the type of the tokens, such as note types and metric types. And, we propose a new Transformer decoder architecture that uses different feed-forward heads to model tokens of different types. With an expansion-compression trick, we convert a piece of music to a sequence of compound words by grouping neighboring tokens, greatly reducing the length of the token sequences. We show that the resulting model can be viewed as a learner over dynamic directed hypergraphs. And, we employ it to learn to compose expressive Pop piano music of full-song length (involving up to 10K individual tokens per song), both conditionally and unconditionally. Our experiment shows that, compared to state-of-the-art models, the proposed model converges 5 to 10 times faster at training (i.e., within a day on a single GPU with 11 GB memory), and with comparable quality in the generated music Wen-Yi Hsiao, Jen-Yu Liu, Yin-Cheng Yeh, Yi-Hsuan Yang |
AAAI | 4 |
| 2021 | Relative Positional Encoding for Transformers with Linear ComplexityabstractRecent advances in Transformer models allow for unprecedented sequence lengths, due to linear space and time complexity. In the meantime, relative positional encoding (RPE) was proposed as beneficial for classical Transformers and consists in exploiting lags instead of absolute positions for inference. Still, RPE is not available for the recent linear-variants of the Transformer, because it requires the explicit computation of the attention matrix, which is precisely what is avoided by such methods. In this paper, we bridge this gap and present Stochastic Positional Encoding as a way to generate PE that can be used as a replacement to the classical additive (sinusoidal) PE and provably behaves like RPE. The main theoretical contribution is to make a connection between positional encoding and cross-covariance structures of correlated Gaussian processes. We illustrate the performance of our approach on the Long-Range Arena benchmark and on music generation. Antoine Liutkus, Ondrej Cífka, Shih-Lun Wu, Umut Simsekli, Yi-Hsuan Yang, Gaël Richard |
ICML | 5 |
| 2021 | Drum-Aware Ensemble Architecture for Improved Joint Musical Beat and Downbeat TrackingabstractThis letter presents a novel system architecture that integrates blind source separation with joint beat and downbeat tracking in musical audio signals. The source separation module segregates the percussive and non-percussive components of the input signal, over which beat and downbeat tracking are performed separately and then the results are aggregated with a learnable fusion mechanism. This way, the system can adaptively determine how much the tracking result for an input signal should depend on the input's percussive or non-percussive components. Evaluation on four testing sets that feature different levels of presence of drum sounds shows that the new architecture consistently outperforms the widely-adopted baseline architecture that does not employ source separation. Ching-Yu Chiu, Alvin Wen-Yu Su, Yi-Hsuan Yang |
IEEE Signal Process. Lett. | 3 |
| 2021 | Leveraging Affective Hashtags for Ranking Music RecommendationsabstractMood and emotion play an important role when it comes to choosing musical tracks to listen to. In the field of music information retrieval and recommendation, emotion is considered contextual information that is hard to capture, albeit highly influential. In this study, we analyze the connection between users` emotional states and their musical choices. Particularly, we perform a large-scale study based on two data sets containing 560,000 and 90,000 #nowplaying tweets, respectively. We extract affective contextual information from hashtags contained in these tweets by applying an unsupervised sentiment dictionary approach. Subsequently, we utilize a state-of-the-art network embedding method to learn latent feature representations of users, tracks and hashtags. Based on both the affective information and the latent features, a set of eight ranking methods is proposed. We find that relying on a ranking approach that incorporates the latent representations of users and tracks allows for capturing a user's general musical preferences well (regardless of used hashtags or affective information). However, for capturing context-specific preferences (a more complex and personal ranking task), we find that ranking strategies that rely on affective information and that leverage hashtags as context information outperform the other ranking strategies. Eva Zangerle, Chih-Ming Chen 0003, Ming-Feng Tsai, Yi-Hsuan Yang |
IEEE Trans. Affect. Comput. | 4 |
| 2020 | A Comparative Study of Western and Chinese Classical Music Based on Soundscape ModelsabstractWhether literally or suggestively, the concept of soundscape is alluded in both modern and ancient music. In this study, we examine whether we can analyze and compare Western and Chinese classical music based on soundscape models. We addressed this question through a comparative study. Specifically, corpora of Western classical music excerpts (WCMED) and Chinese classical music excerpts (CCMED) were curated and annotated with emotional valence and arousal through a crowdsourcing experiment. We used a sound event detection (SED) and soundscape emotion recognition (SER) models with transfer learning to predict the perceived emotion of WCMED and CCMED. The results show that both SER and SED models could be used to analyze Chinese and Western classical music. The fact that SER and SED work better on Chinese classical music emotion recognition provides evidence that certain similarities exist between Chinese classical music and soundscape recordings, which permits transferability between machine learning models. Jianyu Fan, Yi-Hsuan Yang, Kui Dong, Philippe Pasquier |
ICASSP | 2 |
| 2020 | Addressing The Confounds Of Accompaniments In Singer IdentificationabstractIdentifying singers is an important task with many applications. However, the task remains challenging due to many issues. One major issue is related to the confounding factors from the background instrumental music that is mixed with the vocals in music production. A singer identification model may learn to extract non-vocal related features from the instrumental part of the songs, if a singer only sings in certain musical contexts (e.g., genres). The model cannot therefore generalize well when the singer sings in unseen contexts. In this paper, we attempt to address this issue. Specifically, we employ open-unmix, an open source tool with state-of-the-art performance in source separation, to separate the vocal and instrumental tracks of music. We then investigate two means to train a singer identification model: by learning from the separated vocal only, or from an augmented set of data where we "shuffle-and-remix" the separated vocal tracks and instrumental tracks of different songs to artificially make the singers sing in different contexts. We also incorporate melodic features learned from the vocal melody contour for better performance. Evaluation results on a benchmark dataset called the artist20 shows that this data augmentation method greatly improves the accuracy of singer identification. Tsung-Han Hsieh, Kai-Hsiang Cheng, Zhe-Cheng Fan, Yu-Ching Yang, Yi-Hsuan Yang |
ICASSP | 5 |
| 2020 | Speech-To-Singing Conversion in an Encoder-Decoder FrameworkabstractIn this paper our goal is to convert a set of spoken lines into sung ones. Unlike previous signal processing based methods, we take a learning based approach to the problem. This allows us to automatically model various aspects of this transformation, thus overcoming dependence on specific inputs such as high quality singing templates or phoneme-score synchronization information. Specifically, we propose an encoder-decoder framework for our task. Given time-frequency representations of speech and a target melody contour, we learn encodings that enable us to synthesize singing that preserves the linguistic content and timbre of the speaker while adhering to the target melody. We also propose a multi-task learning based objective to improve lyric intelligibility. We present a quantitative and qualitative analysis of our framework. Jayneel Parekh, Preeti Rao, Yi-Hsuan Yang |
ICASSP | 3 |
| 2020 | Score and Lyrics-Free Singing Voice Generation
Jen-Yu Liu, Yu-Hua Chen, Yin-Cheng Yeh, Yi-Hsuan Yang |
ICCC | 4 |
| 2020 | Unconditional Audio Generation with Generative Adversarial Networks and Cycle RegularizationabstractIn a recent paper, we have presented a generative adversarial network (GAN)-based model for unconditional generation of the mel-spectrograms of singing voices. As the generator of the model is designed to take a variable-length sequence of noise vectors as input, it can generate mel-spectrograms of variable length. However, our previous listening test shows that the quality of the generated audio leaves room for improvement. The present paper extends and expands that previous work in the following aspects. First, we employ a hierarchical architecture in the generator to induce some structure in the temporal dimension. Second, we introduce a cycle regularization mechanism to the generator to avoid mode collapse. Third, we evaluate the performance of the new model not only for generating singing voices, but also for generating speech voices. Evaluation result shows that new model outperforms the prior one both objectively and subjectively. We also employ the model to unconditionally generate sequences of piano and violin music and find the result promising. Audio examples, as well as the code for implementing our model, will be publicly available online upon paper publication. Jen-Yu Liu, Yu-Hua Chen, Yin-Cheng Yeh, Yi-Hsuan Yang |
INTERSPEECH | 4 |
| 2020 | Speech-to-Singing Conversion Based on Boundary Equilibrium GANabstractThis paper investigates the use of generative adversarial network (GAN)-based models for converting a speech signal into a singing one, without reference to the phoneme sequence underlying the speech.This is achieved by viewing speech-to-singing conversion as a style transfer problem.Specifically, given a speech input, and the F0 contour of the target singing output, the proposed model generates the spectrogram of a singing signal with a progressive-growing encoder/decoder architecture.Moreover, the model uses a boundary equilibrium GAN loss term such that it can learn from both paired and unpaired data.The spectrogram is finally converted into wave with a separate GAN-based vocoder.Our quantitative and qualitative analysis show that the proposed model generates singing voices with much higher naturalness than an existing non adversarially-trained baseline. Da-Yi Wu, Yi-Hsuan Yang |
INTERSPEECH | 2 |
| 2020 | Pop Music Transformer: Beat-based Modeling and Generation of Expressive Pop Piano CompositionsabstractA great number of deep learning based models have been recently proposed for automatic music composition. Among these models, the Transformer stands out as a prominent approach for generating expressive classical piano performance with a coherent structure of up to one minute. The model is powerful in that it learns abstractions of data on its own, without much human-imposed domain knowledge or constraints. In contrast with this general approach, this paper shows that Transformers can do even better for music modeling, when we improve the way a musical score is converted into the data fed to a Transformer model. In particular, we seek to impose a metrical structure in the input data, so that Transformers can be more easily aware of the beat-bar-phrase hierarchical structure in music. The new data representation maintains the flexibility of local tempo changes, and provides hurdles to control the rhythmic and harmonic structure of music. With this approach, we build a Pop Music Transformer that composes Pop piano music with better rhythmic structure than existing Transformer models. Yu-Siang Huang, Yi-Hsuan Yang |
ACM Multimedia | 2 |
| 2020 | Mixing-Specific Data Augmentation Techniques for Improved Blind Violin/Piano Source SeparationabstractBlind music source separation has been a popular and active subject of research in both the music information retrieval and signal processing communities. To counter the lack of available multi-track data for supervised model training, a data augmentation method that creates artificial mixtures by combining tracks from different songs has been shown useful in recent works. Following this light, we examine further in this paper extended data augmentation methods that consider more sophisticated mixing settings employed in the modern music production routine, the relationship between the tracks to be combined, and factors of silence. As a case study, we consider the separation of violin and piano tracks in a violin piano ensemble, evaluating the performance in terms of common metrics, namely SDR, SIR, and SAR. In addition to examining the effectiveness of these new data augmentation methods, we also study the influence of the amount of training data. Our evaluation shows that the proposed mixing-specific data augmentation methods can help improve the performance of a deep learning-based model for source separation, especially in the case of small training data. Ching-Yu Chiu, Wen-Yi Hsiao, Yin-Cheng Yeh, Yi-Hsuan Yang, Alvin Wen-Yu Su |
MMSP | 4 |
| 2020 | Fast Tensor Factorization for Large-Scale Context-Aware Recommendation from Implicit FeedbackabstractThis paper presents a fast Tensor Factorization (TF) algorithm for context-aware recommendation from implicit feedback. For such a recommendation problem, the observed data indicate the (positive) association between users and items in some given contexts. For better accuracy, it has been shown essential to include unobserved data that indicate the negative user-item-context associations. As such unobserved data greatly outnumber the observed ones, for efficiency existing algorithms usually use only a small part of the unobserved data for model training. We show in this paper that it is possible, and beneficial, to use all the unobserved data in training a TF based context-aware recommender system. This is achieved by two technical innovations. First, we scrutinize the matrix computation of the closed-form solution and accelerate the computation by memorizing the repetitive computation. Second, we further boost the generalization and scalability by dropping out some pairwise interactions when updating user, item or context factors to prevent overfitting and to reduce the training time. The resulting whole-data based learning algorithm, referred to as DropTF in the paper, is efficient and scale well. Our evaluation on two small benchmark datasets and a million-scale large dataset demonstrates improved accuracy over some existing algorithms for context-aware recommendation. Szu-Yu Chou, Jyh-Shing Roger Jang, Yi-Hsuan Yang |
IEEE Trans. Big Data | 3 |
| 2020 | Backpropagation With $N$ -D Vector-Valued Neurons Using Arbitrary Bilinear ProductsabstractVector-valued neural learning has emerged as a promising direction in deep learning recently. Traditionally, training data for neural networks (NNs) are formulated as a vector of scalars; however, its performance may not be optimal since associations among adjacent scalars are not modeled. In this article, we propose a new vector neural architecture called the Arbitrary BIlinear Product NN (ABIPNN), which processes information as vectors in each neuron, and the feedforward projections are defined using arbitrary bilinear products. Such bilinear products can include circular convolution, 7-D vector product, skew circular convolution, reversed-time circular convolution, or other new products that are not seen in the previous work. As a proof-of-concept, we apply our proposed network to multispectral image denoising and singing voice separation. Experimental results show that ABIPNN obtains substantial improvements when compared to conventional NNs, suggesting that associations are learned during training. Zhe-Cheng Fan, Tak-Shing Chan, Yi-Hsuan Yang, Jyh-Shing Roger Jang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | PerformanceNet: Score-to-Audio Music Generation with Multi-Band Convolutional Residual NetworkabstractMusic creation is typically composed of two parts: composing the musical score, and then performing the score with instruments to make sounds. While recent work has made much progress in automatic music generation in the symbolic domain, few attempts have been made to build an AI model that can render realistic music audio from musical scores. Directly synthesizing audio with sound sample libraries often leads to mechanical and deadpan results, since musical scores do not contain performance-level information, such as subtle changes in timing and dynamics. Moreover, while the task may sound like a text-to-speech synthesis problem, there are fundamental differences since music audio has rich polyphonic sounds. To build such an AI performer, we propose in this paper a deep convolutional model that learns in an end-to-end manner the score-to-audio mapping between a symbolic representation of music called the pianorolls and an audio representation of music called the spectrograms. The model consists of two subnets: the ContourNet, which uses a U-Net structure to learn the correspondence between pianorolls and spectrograms and to give an initial result; and the TextureNet, which further uses a multi-band residual network to refine the result by adding the spectral texture of overtones and timbre. We train the model to generate music clips of the violin, cello, and flute, with a dataset of moderate size. We also present the result of a user study that shows our model achieves higher mean opinion score (MOS) in naturalness and emotional expressivity than a WaveNet-based model and two off-the-shelf synthesizers. We open our source code at https://github.com/bwang514/PerformanceNet Bryan Wang, Yi-Hsuan Yang |
AAAI | 2 |
| 2019 | Learning to Match Transient Sound Events Using Attentional Similarity for Few-shot Sound RecognitionabstractIn this paper, we introduce a novel attentional similarity module for the problem of few-shot sound recognition. Given a few examples of an unseen sound event, a classifier must be quickly adapted to recognize the new sound event without much fine-tuning. The proposed attentional similarity module can be plugged into any metric-based learning method for few-shot learning, allowing the resulting model to especially match related short sound events. Extensive experiments on two datasets show that the proposed module consistently improves the performance of five different metric-based learning methods for few-shot sound recognition. The relative improvement ranges from +4.1% to +7.7% for 5-shot 5-way accuracy for the ESC-50 dataset, and from +2.1% to +6.5% for noiseESC-50. Qualitative results demonstrate that our method contributes in particular to the recognition of transient sound events. Szu-Yu Chou, Kai-Hsiang Cheng, Jyh-Shing Roger Jang, Yi-Hsuan Yang |
ICASSP | 4 |
| 2019 | A Streamlined Encoder/decoder Architecture for Melody ExtractionabstractMelody extraction in polyphonic musical audio is important for music signal processing. In this paper, we propose a novel streamlined encoder/decoder network that is designed for the task. We make two technical contributions. First, drawing inspiration from a state-of-the-art model for semantic pixel-wise segmentation, we pass through the pooling indices between pooling and un-pooling layers to localize the melody in frequency. We can achieve result close to the state-of-the-art with much fewer convolutional layers and simpler convolution modules. Second, we propose a way to use the bottleneck layer of the network to estimate the existence of a melody line for each time frame, and make it possible to use a simple argmax function instead of ad-hoc thresholding to get the final estimation of the melody line. Our experiments on both vocal melody extraction and general melody extraction validate the effectiveness of the proposed model. Tsung-Han Hsieh, Li Su 0004, Yi-Hsuan Yang |
ICASSP | 3 |
| 2019 | Multitask Learning for Frame-level Instrument RecognitionabstractFor many music analysis problems, we need to know the presence of instruments for each time frame in a multi-instrument musical piece. However, such a frame-level instrument recognition task remains difficult, mainly due to the lack of labeled datasets. To address this issue, we present in this paper a large-scale dataset that contains synthetic polyphonic music with frame-level pitch and instrument labels. Moreover, we propose a simple yet novel network architecture to jointly predict the pitch and instrument for each frame. With this multitask learning method, the pitch information can be leveraged to predict the instruments, and also the other way around. And, by using the so-called pianoroll representation of music as the main target output of the model, our model also predicts the instruments that play each individual note event. We validate the effectiveness of the proposed method for frame-level instrument recognition by comparing it with its single-task ablated versions and three state-of-the-art methods. We also demonstrate the result of the proposed method for multi-pitch streaming with real-world music. For reproducibility, we will share the code to crawl the data and to implement the proposed model at: https://github.com/biboamy/ instrument-streaming. Yun-Ning Hung, Yi-Hsuan Yang |
ICASSP | 3 |
| 2019 | Demonstration of PerformanceNet: A Convolutional Neural Network Model for Score-to-Audio Music GenerationabstractWe present in this paper PerformacnceNet, a neural network model we proposed recently to achieve score-to-audio music generation. The model learns to convert a music piece from the symbolic domain to the audio domain, assigning performance-level attributes such as changes in velocity automatically to the music and then synthesizing the audio. The model is therefore not just a neural audio synthesizer, but an AI performer that learns to interpret a musical score in its own way. The code and sample outputs of the model can be found online at https://github.com/bwang514/PerformanceNet. Yu-Hua Chen, Bryan Wang, Yi-Hsuan Yang |
IJCAI | 3 |
| 2019 | Musical Composition Style Transfer via Disentangled Timbre RepresentationsabstractMusic creation involves not only composing the different parts (e.g., melody, chords) of a musical work but also arranging/selecting the instruments to play the different parts. While the former has received increasing attention, the latter has not been much investigated. This paper presents, to the best of our knowledge, the first deep learning models for rearranging music of arbitrary genres. Specifically, we build encoders and decoders that take a piece of polyphonic musical audio as input, and predict as output its musical score. We investigate disentanglement techniques such as adversarial training to separate latent factors that are related to the musical content (pitch) of different parts of the piece, and that are related to the instrumentation (timbre) of the parts per short-time segment. By disentangling pitch and timbre, our models have an idea of how each piece was composed and arranged. Moreover, the models can realize “composition style transfer” by rearranging a musical piece without much affecting its pitch content. We validate the effectiveness of the models by experiments on instrument activity detection and composition style transfer. To facilitate follow-up research, we open source our code at https://github.com/biboamy/instrument-disentangle. Yun-Ning Hung, I-Tung Chiang, Yi-Hsuan Yang |
IJCAI | 4 |
| 2019 | Dilated Convolution with Dilated GRU for Music Source SeparationabstractStacked dilated convolutions used in Wavenet have been shown effective for generating high-quality audios. By replacing pooling/striding with dilation in convolution layers, they can preserve high-resolution information and still reach distant locations. Producing high-resolution predictions is also crucial in music source separation, whose goal is to separate different sound sources while maintain the quality of the separated sounds. Therefore, in this paper, we use stacked dilated convolutions as the backbone for music source separation. Although stacked dilated convolutions can reach wider context than standard convolutions do, their effective receptive fields are still fixed and might not be wide enough for complex music audio signals. To reach even further information at remote locations, we propose to combine a dilated convolution with a modified GRU called Dilated GRU to form a block. A Dilated GRU receives information from k-step before instead of the previous step for a fixed k. This modification allows a GRU unit to reach a location with fewer recurrent steps and run faster because it can execute in parallel partially. We show that the proposed model with a stack of such blocks performs equally well or better than the state-of-the-art for separating both vocals and accompaniment. Jen-Yu Liu, Yi-Hsuan Yang |
IJCAI | 2 |
| 2019 | Deep Cyclic Group NetworksabstractWe propose a new network architecture called deep cyclic group network (DCGN) that uses the cyclic group algebra for convolutional vector-neuron learning. The input to DCGN is a three-way tensor, where the mode-3 dimension corresponds to the dimensionality of the input data, e.g., three for RGB images. To handle vector-valued inputs, we replace scalar multiplication with circular convolution for the feedforward and backpropagation processes. As a result, every feature map and kernel map is a three-way tensor with the same mode-3 dimension as the input data. This way, DCGN may capture more of the relations among different data dimensions, especially for regression tasks where the target output has the same dimensionality as the input data. Moreover, DCGN can deal with input data of arbitrary dimensions, a property that existing architectures such as deep complex networks and deep quaternion networks (DQN) lack. Experiments show that DCGN indeed performs better than convolutional neural networks and DQN for two regression tasks, namely color image inpainting and multispectral image denoising. Zhe-Cheng Fan, Tak-Shing Chan, Yi-Hsuan Yang, Jyh-Shing Roger Jang |
IJCNN | 3 |
| 2019 | Multi-label Few-shot Learning for Sound Event RecognitionabstractFew-shot classification aims to generalize the concept from seen classes to unseen novel classes using only a few examples. Although significant progress in few-shot classification has been made, most approaches focus on a standard multi-class scenario and are based on learning single-label embedding of the labeled examples to classify the unlabeled examples. Besides, we note that state-of-the-art methods in few-shot learning mostly adopt a metric-based architecture and the the so-called episode training strategy. While this approach works nicely for multiclass classification, it is hard to apply it to the multi-label scenario because of the complexity of forming an episode. In this paper, we propose a One-vs.-Rest episode selection strategy to mitigate this issue and apply the strategy to the multi-label few-shot problem. Experiments conducted using the large-scale data found in the AudioSet show that the models with our training strategy extract the semantic features under the multi-label setting. Kai-Hsiang Cheng, Szu-Yu Chou, Yi-Hsuan Yang |
MMSP | 3 |
| 2019 | Collaborative Similarity Embedding for Recommender SystemsabstractWe present collaborative similarity embedding (CSE), a unified framework that exploits comprehensive collaborative relations available in a user-item bipartite graph for representation learning and recommendation. In the proposed framework, we differentiate two types of proximity relations: direct proximity and k-th order neighborhood proximity. While learning from the former exploits direct user-item associations observable from the graph, learning from the latter makes use of implicit associations such as user-user similarities and item-item similarities, which can provide valuable information especially when the graph is sparse. Moreover, for improving scalability and flexibility, we propose a sampling technique that is specifically designed to capture the two types of proximity relations. Extensive experiments on eight benchmark datasets show that CSE yields significantly better performance than state-of-the-art recommendation methods. Chih-Ming Chen 0003, Chuan-Ju Wang, Ming-Feng Tsai, Yi-Hsuan Yang |
WWW | 4 |
| 2019 | Weakly-Supervised Visual Instrument-Playing Action Detection in VideosabstractMusic videos are one of the most popular types of video streaming services, and instrument playing is among the most common scenes in such videos. In order to understand the instrument-playing scenes in the videos, it is important to know what instruments are played, when they are played, and where the playing actions occur in the scene. While audio-based recognition of instruments has been widely studied, the visual aspect of music instrument playing remains largely unaddressed in the literature. One of the main obstacles is the difficulty in collecting annotated data of the action locations for training-based methods. To address this issue, we propose a weakly supervised framework to find when and where the instruments are played in the videos. We propose using two auxiliary models: 1) a sound model and 2) an object model to provide supervision for training the instrument-playing action model. The sound model provides temporal supervisions, while the object model provides spatial supervisions. They together can simultaneously provide temporal and spatial supervisions. The resulting model only needs to analyze the visual part of a music video to deduce which, when, and where instruments are played. We found that the proposed method significantly improves localization accuracy. We evaluate the result of the proposed method temporally and spatially on a small dataset (a total of 5400 frames) that we manually annotated. Jen-Yu Liu, Yi-Hsuan Yang, Shyh-Kang Jeng |
IEEE Trans. Multim. | 2 |
| 2018 | MuseGAN: Multi-track Sequential Generative Adversarial Networks for Symbolic Music Generation and AccompanimentabstractGenerating music has a few notable differences from generating images and videos. First, music is an art of time, necessitating a temporal model. Second, music is usually composed of multiple instruments/tracks with their own temporal dynamics, but collectively they unfold over time interdependently. Lastly, musical notes are often grouped into chords, arpeggios or melodies in polyphonic music, and thereby introducing a chronological ordering of notes is not naturally suitable. In this paper, we propose three models for symbolic multi-track music generation under the framework of generative adversarial networks (GANs). The three models, which differ in the underlying assumptions and accordingly the network architectures, are referred to as the jamming model, the composer model and the hybrid model. We trained the proposed models on a dataset of over one hundred thousand bars of rock music and applied them to generate piano-rolls of five tracks: bass, drums, guitar, piano and strings. A few intra-track and inter-track objective metrics are also proposed to evaluate the generative results, in addition to a subjective user study. We show that our models can generate coherent music of four bars right from scratch (i.e. without human inputs). We also extend our models to human-AI cooperative music generation: given a specific track composed by human, we can generate four additional tracks to accompany it. All code, the dataset and the rendered audio samples are available at https://salu133445.github.io/musegan/. Hao-Wen Dong, Wen-Yi Hsiao, Li-Chia Yang, Yi-Hsuan Yang |
AAAI | 4 |
| 2018 | Generating Music Medleys via Playing Music Puzzle GamesabstractGenerating music medleys is about finding an optimal permutation of a given set of music clips. Toward this goal, we propose a self-supervised learning task, called the music puzzle game, to train neural network models to learn the sequential patterns in music. In essence, such a game requires machines to correctly sort a few multisecond music fragments. In the training stage, we learn the model by sampling multiple non-overlapping fragment pairs from the same songs and seeking to predict whether a given pair is consecutive and is in the correct chronological order. For testing, we design a number of puzzle games with different difficulty levels, the most difficult one being music medley, which requiring sorting fragments from different songs. On the basis of state-of-the-art Siamese convolutional network, we propose an improved architecture that learns to embed frame-level similarity scores computed from the input fragment pairs to a common space, where fragment pairs in the correct order can be more easily identified. Our result shows that the resulting model, dubbed as the similarity embedding network (SEN), performs better than competing models across different games, including music jigsaw puzzle, music sequencing, and music medley. Example results can be found at our project website, https://remyhuang.github.io/DJnet. Yu-Siang Huang, Szu-Yu Chou, Yi-Hsuan Yang |
AAAI | 3 |
| 2018 | Modeling Multi-way Relations with Hypergraph EmbeddingabstractHypergraph is a data structure commonly used to represent connections and relations between multiple objects. Embedding a hypergraph into a low-dimensional space and representing each vertex as a vector is useful in various tasks such as visualization, classification, and link prediction. However, most hypergraph embedding or learning algorithms reduce multi-way relations to pairwise ones, which turn hypergraphs into graphs and lose a lot of information. Inspired by Laplacian tensors of uniform hypergraphs, we propose in this paper a novel method that incorporates multi-way relations into an optimization problem. We design an objective that is applicable to both uniform and non-uniform hypergraphs with the constraint of having non-negative embedding vectors. For scalability, we apply negative sampling and use constrained stochastic gradient descent to solve the optimization problem. We test our method in a context-aware recommendation task on a real-world dataset. Experimental results show that our method outperforms a few well-known graph and hypergraph embedding methods. Chia-An Yu, Ching-Lun Tai, Tak-Shing Chan, Yi-Hsuan Yang |
CIKM | 4 |
| 2018 | Seethevoice: Learning from Music to Visual Storytelling of ShotsabstractTypes of shots in the language of film are considered the key elements used by a director for visual storytelling. In filming a musical performance, manipulating shots could stimulate desired effects such as manifesting the emotion or deepening the atmosphere. However, while the visual storytelling technique is often employed in creating professional recordings of a live concert, audience recordings of the same event often lack such sophisticated manipulations. Thus it would be useful to have a versatile system that can perform video mashup to create a refined video from such amateur clips. To this end, we propose to translate the music into a near-professional shot (type) sequence by learning the relation between music and visual storytelling of shots. The resulting shot sequence can then be used to better portray the visual storytelling of a song and guide the concert video mashup process. Our method introduces a novel probabilistic-based fusion approach, named as multi-resolution fused recurrent neural networks (MF-RNNs) with film-language, which integrates multi-resolution fused RNNs and a film-language model for boosting the translation performance. The results from objective and subjective experiments demonstrate that MF-RNNs with film-language can generate an appealing shot sequence with better viewing experience. Wen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Yi-Hsuan Yang, Hsin-Min Wang, Hsiao-Rong Tyan, Hong-Yuan Mark Liao |
ICME | 4 |
| 2018 | Cross-Cultural Music Emotion Recognition by Adversarial Discriminative Domain AdaptationabstractAnnotation of the perceived emotion of a music piece is required for an automatic music emotion recognition system. Most music emotion datasets are developed for Western pop songs. The problem is that a music emotion recognizer trained on such datasets may not work well for non-Western pop songs due to the differences in acoustic characteristics and emotion perception that are inherent to cultural background. The problem was also found in cross-cultural and cross-dataset studies; however, little has been done to learn how to adapt a model pre-trained on a source music genre to a target music genre of interest. In this paper, we propose to address the problem by an unsupervised adversarial domain adaptation method. It employs neural network models to make the target music indistinguishable from the source music in a learned feature representation space. Because emotion perception is multifaceted, three types of input feature representations related to timbre, pitch, and rhythm are considered for performance evaluation. The results show that the proposed method effectively improves the prediction of the valence of Chinese pop songs from a model trained for Western pop songs. Yi-Wei Chen, Yi-Hsuan Yang, Homer H. Chen |
ICMLA | 2 |
| 2018 | Lead Sheet Generation and Arrangement by Conditional Generative Adversarial NetworkabstractResearch on automatic music generation has seen great progress due to the development of deep neural networks. However, the generation of multi-instrument music of arbitrary genres still remains a challenge. Existing research either works on lead sheets or multi-track piano-rolls found in MIDIs, but both musical notations have their limits. In this work, we propose a new task called lead sheet arrangement to avoid such limits. A new recurrent convolutional generative model for the task is proposed, along with three new symbolic-domain harmonic features to facilitate learning from unpaired lead sheets and MIDIs. Our model can generate lead sheets and their arrangements of eightbar long. Source code and audio samples of the generated result can be found at the project webpage: https://liuhaumin. github.io/Leadsheet Arrangement. Hao-Min Liu, Yi-Hsuan Yang |
ICMLA | 2 |
| 2018 | Denoising Auto-Encoder with Recurrent Skip Connections and Residual Regression for Music Source SeparationabstractConvolutional neural networks with skip connections have shown good performance in music source separation. In this work, we propose a denoising Auto-encoder with Recurrent skip Connections (ARC). We use 1D convolution along the temporal axis of the time-frequency feature map in all layers of the fully-convolutional network. The use of 1D convolution makes it possible to apply recurrent layers to the intermediate outputs of the convolution layers. In addition, we also propose an enhancement network and a residual regression method to further improve the separation result. The recurrent skip connections, the enhancement module, and the residual regression all improve the separation quality. The ARC model with residual regression achieves 5.74 siganl-to-distoration ratio (SDR) in vocals with MUSDB (used in SiSEC 2018). We also evaluate the ARC model alone on the older dataset DSD100 (used in SiSEC 2016) and it achieves 5.91 SDR in vocals. Jen-Yu Liu, Yi-Hsuan Yang |
ICMLA | 2 |
| 2018 | Learning to Recognize Transient Sound Events using Attentional SupervisionabstractMaking sense of the surrounding context and ongoing events through not only the visual inputs but also acoustic cues is critical for various AI applications. This paper presents an attempt to learn a neural network model that recognizes more than 500 different sound events from the audio part of user generated videos (UGV). Aside from the large number of categories and the diverse recording conditions found in UGV, the task is challenging because a sound event may occur only for a short period of time in a video clip. Our model specifically tackles this issue by combining a main subnet that aggregates information from the entire clip to make clip-level predictions, and a supplementary subnet that examines each short segment of the clip for segment-level predictions. As the labeled data available for model training are typically on the clip level, the latter subnet learns to pay attention to segments selectively to facilitate attentional segment-level supervision. We call our model the M&mnet, for it leverages both “M”acro (clip-level) supervision and “m”icro (segment-level) supervision derived from the macro one. Our experiments show that M&mnet works remarkably well for recognizing sound events, establishing a new state-of-theart for DCASE17 and AudioSet data sets. Qualitative analysis suggests that our model exhibits strong gains for short events. In addition, we show that the micro subnet is computationally light and we can use multiple micro subnets to better exploit information in different temporal scales. Szu-Yu Chou, Jyh-Shing Roger Jang, Yi-Hsuan Yang |
IJCAI | 3 |
| 2018 | Predicting the Probability Density Function of Music Emotion Using Emotion Space MappingabstractComputationally modeling the affective content of music has been intensively studied in recent years because of its wide applications in music retrieval and recommendation. Although significant progress has been made, this task remains challenging due to the difficulty in properly characterizing the emotion of a music piece. Music emotion perceived by people is subjective by nature and thus complicates the process of collecting the emotion annotations as well as developing the predictive model. Instead of assuming people can reach a consensus on the emotion of music, in this work we propose a novel machine learning approach that characterizes the music emotion as a probability distribution in the valence-arousal (VA) emotion space, not only tackling the subjectivity but also precisely describing the emotions of a music piece. Specifically, we represent the emotion of a music piece as a probability density function (PDF) in the VA space via kernel density estimation from human annotations. To associate emotion with the audio features extracted from music pieces, we learn the combination coefficients by optimizing some objective functions of audio features, and then predict the emotion of an unseen piece by linearly combining the PDFs of the training pieces with the coefficients. Several algorithms for learning the coefficients are studied. Evaluations on the NTUMIR and MediaEval 2013 datasets validate the effectiveness of the proposed methods in predicting the probability distributions of emotion from audio features. We also demonstrate how to use the proposed approach in emotion-based music retrieval. Yu-Hao Chin, Jia-Ching Wang, Ju-Chiang Wang, Yi-Hsuan Yang |
IEEE Trans. Affect. Comput. | 4 |
| 2018 | Coherent Deep-Net Fusion To Classify Shots In Concert VideosabstractVarying types of shots is a fundamental element in the language of film, commonly used by a visual storytelling director. The technique is often used in creating professional recordings of a live concert, but meanwhile may not be appropriately applied in audience recordings of the same event. Such variations could cause the task of classifying shots in concert videos, professional or amateur, very challenging. To achieve more reliable shot classification, we propose a novel probabilistic-based approach, named as coherent classification net (CC-Net), by addressing three crucial issues. First, we focus on learning more effective features by fusing the layer-wise outputs extracted from a deep convolutional neural network (CNN), pretrained on a large-scale data set for object recognition. Second, we introduce a frame-wise classification scheme, the error weighted deep cross-correlation model (EW-Deep-CCM), to boost the classification accuracy. Specifically, the deep neural network-based cross-correlation model (deep-CCM) is constructed to not only model the extracted feature hierarchies of CNN independently, but also relate the statistical dependencies of paired features from different layers. Then, a Bayesian error weighting scheme for a classifier combination is adopted to explore the contributions from individual Deep-CCM classifiers to enhance the accuracy of shot classification in each image frame. Third, we feed the frame-wise classification results to a linear-chain conditional random field module to refine the shot predictions by taking into account the global and temporal regularities. We provide extensive experimental results on a data set of live concert videos to demonstrate the advantage of the proposed CC-Net over existing popular fusion approaches for shot classification. Jen-Chun Lin, Wen-Li Wei, Tyng-Luh Liu, Yi-Hsuan Yang, Hsin-Min Wang, Hsiao-Rong Tyan, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 4 |
| 2017 | Low-Rank Matrix Completion over Finite Abelian Group Algebras for Context-Aware RecommendationabstractThe incorporation of contextual information is an important part of context-aware recommendation. Many context-aware recommendation systems adopt tensor completion to include contextual information. However, the symmetries between dimensions of a tensor induce an unreasonable assumption that users, items and contexts should be treated equally in recommender systems. In this paper, we address this by using matrices over finite abelian group algebra (AGA) to model context-aware interactions between users and items. Specifically, we formulate context-aware recommendation as a low-rank matrix completion problem over AGA (MC-AGA) and derive a new algorithm using the inexact augmented Lagrange multiplier method. We then test MC-AGA on two real-world datasets: one containing implicit feedback and one with explicit feedback. Experiment results show that MC-AGA outperforms not only existing tensor completion algorithms but also recommendation systems with other context-aware representations. Chia-An Yu, Tak-Shing Chan, Yi-Hsuan Yang |
CIKM | 3 |
| 2017 | Polyphonic piano note transcription with non-negative matrix factorization of differential spectrogramabstractAutomatic music transcription is usually approached by using a time-frequency (TF) representation such as the short-time Fourier transform (STFT) spectrogram or the constant-Q transform. In this paper, we propose a novel yet simple TF representation that capitalizes the effectiveness of spectral flux features in highlighting note onset times. We refer to this representation as the differential spectrogram and investigate its usefulness for note-level piano transcription using two different non-negative matrix factorization (NMF) algorithms. Experiments on the MAPS ENSTDkCl dataset validate the advantages of the differential spectrogram over the STFT spectrogram for this task. Moreover, by adapting a state-of-the-art convolutional NMF algorithm with the differential spectrogram, we can achieve even better accuracy than the state-of-the-art on this dataset. Our analysis shows that the new representation suppresses unwanted TF patterns and performs particularly well in improving the recall rate. Lufei Gao, Li Su 0004, Yi-Hsuan Yang, Tan Lee |
ICASSP | 3 |
| 2017 | Automatic conversion of Pop music into chiptunes for 8-bit pixel artabstractIn this paper, we propose an audio mosaicing method that converts Pop songs into a specific music style called “chiptune,” or “8-bit music.” The goal is to reproduce Pop songs by using the sound of the chips on the old game consoles in 1980s/1990s. The proposed method goes through a procedure that first analyzes the pitches of an incoming Pop song in the frequency domain, and then synthesizes the song with template waveforms in the time domain to make it sound like 8-bit music. Because a Pop song is usually composed of the vocal melody and the instrumental accompaniment, in the analysis stage we use a singing voice separation algorithm to separate the vocals from the instruments, and then apply different pitch detection algorithms to transcribe the two separated sources. We validate through a subjective listening test that the proposed method creates much better 8-bit music than existing nonnegative matrix factorization based methods can do. Moreover, we find that synthesis in the time domain is important for this task. Shih-Yang Su, Cheng-Kai Chiu, Li Su 0004, Yi-Hsuan Yang |
ICASSP | 4 |
| 2017 | Weakly-supervised audio event detection using event-specific Gaussian filters and fully convolutional networksabstractAudio event detection aims at discovering the elements inside an audio clip. In addition to labeling the clips with the audio events, we want to find out the temporal locations of these events. However, creating clearly annotated training data can be time-consuming. Therefore, we provide a model based on convolutional neural networks that relies only on weakly-supervised data for training. These data can be directly obtained from online platforms, such as Freesound, with the clip-level labels assigned by the uploaders. The structure of our model is extended to a fully convolutional networks, and an event-specific Gaussian filter layer is designed to advance its learning ability. Besides, this model is able to detect frame-level information, e.g., the temporal position of sounds, even when it is trained merely with clip-level labels. Ting-Wei Su, Jen-Yu Liu, Yi-Hsuan Yang |
ICASSP | 3 |
| 2017 | Deep-net fusion to classify shots in concert videosabstractVarying types of shots is a fundamental element in the language of film, commonly used by a visual storytelling director to convey the emotion, ideas, and art. To classify such types of shots from images, we present a new framework that facilitates the intriguing task by addressing two key issues. We first focus on learning more effective features by fusing the layer-wise outputs extracted from a deep convolutional neural network (CNN), pre-trained on a large-scale dataset for object recognition. We then introduce a probabilistic fusion model, termed as error weighted deep cross-correlation model (EW-Deep-CCM), to boost the classification accuracy. Specifically, the deep neural network-based cross-correlation model (Deep-CCM) is constructed to not only model the extracted feature hierarchies of CNN independently but also relate the statistical dependencies of paired features from different layers. Then, a Bayesian error weighting scheme for classifier combination is adopted to explore the contributions from individual Deep-CCM classifiers to enhance the accuracy of shot classification. We provide extensive experimental results on a dataset of live concert videos to demonstrate the advantage of the proposed EW-Deep-CCM over existing popular fusion approaches. The video demos can be found at https://sites.google.com/site/ewdeepccm2/demo. Wen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Yi-Hsuan Yang, Hsin-Min Wang, Hsiao-Rong Tyan, Hong-Yuan Mark Liao |
ICASSP | 4 |
| 2017 | Revisiting the problem of audio-based hit song prediction using convolutional neural networksabstractBeing able to predict whether a song can be a hit has important applications in the music industry. Although it is true that the popularity of a song can be greatly affected by external factors such as social and commercial influences, to which degree audio features computed from musical signals (whom we regard as internal factors) can predict song popularity is an interesting research question on its own. Motivated by the recent success of deep learning techniques, we attempt to extend previous work on hit song prediction by jointly learning the audio features and prediction models using deep learning. Specifically, we experiment with a convolutional neural network model that takes the primitive mel-spectrogram as the input for feature learning, a more advanced JYnet model that uses an external song dataset for supervised pre-training and auto-tagging, and the combination of these two models. We also consider the inception model to characterize audio information in different scales. Our experiments suggest that deep structures are indeed more accurate than shallow structures in predicting the popularity of either Chinese or Western Pop songs in Taiwan. We also use the tags predicted by JYnet to gain insights into the result of different models. Li-Chia Yang, Szu-Yu Chou, Jen-Yu Liu, Yi-Hsuan Yang |
ICASSP | 4 |
| 2017 | Conditional preference nets for user and item cold start problems in music recommendationabstractA great amount of data is usually needed for a recommender system to learn the associations between users and items. However, in practical applications, new users and new items emerge everyday, and the system has to react to them promptly. The ability to recommend proper items to new users affects the users' first impression and accordingly the retention rate, whereas recommending new items to proper users contributes to the freshness of the recommendation. In this paper, we propose a deep learning model called the conditional preference nets (CPN) to deal with both new users and items under the same model framework. CPN employs an introductory user survey to learn about new users, and content features automatically extracted from items for the item side. Through a new idea called the preference vectors and an existing content embedding technique, the same model can capitalize the observed associations between known (i.e. old) users and items, thereby benefiting from the cumulative knowledge of user behavior gained over time. We validate the superiority of CPN over prior arts using the Million Song Dataset. We also demonstrate how CPN allows a user to pick either genres, artists, or the combination of them in the introductory survey. Szu-Yu Chou, Li-Chia Yang, Yi-Hsuan Yang, Jyh-Shing Roger Jang |
ICME | 3 |
| 2017 | The mood of Chinese Pop music: Representation and recognitionabstractMusic mood recognition (MMR) has attracted much attention in music information retrieval research, yet there are few MMR studies that focus on non‐Western music. In addition, little has been done on connecting the 2 most adopted music mood representation models: categorical and dimensional. To bridge these gaps, we constructed a new data set consisting of 818 Chinese Pop (C‐Pop) songs, 3 complete sets of mood annotations in both representations, as well as audio features corresponding to 5 distinct categories of musical characteristics. The mood space of C‐Pop songs was analyzed and compared to that of Western Pop songs. We also explored the relationship between categorical and dimensional annotations and the results revealed that one set of annotations could be reliably predicted by the other. Classification and regression experiments were conducted on the data set, providing benchmarks for future research on MMR of non‐Western music. Based on these analyses, we reflect and discuss the implications of the findings to MMR research. Xiao Hu 0001, Yi-Hsuan Yang |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2017 | Informed Group-Sparse Representation for Singing Voice SeparationabstractSinging voice separation attempts to separate the vocal and instrumental parts of a music recording, which is a fundamental problem in music information retrieval. Recent work on singing voice separation has shown that the low-rank representation and informed separation approaches are both able to improve separation quality. However, low-rank optimizations are computationally inefficient due to the use of singular value decompositions. Therefore, in this letter, we propose a new linear-time algorithm called informed group-sparse representation, and use it to separate the vocals from music using pitch annotations as side information. Experimental results on the iKala dataset confirm the efficacy of our approach, suggesting that the music accompaniment follows a group-sparse structure given a pretrained instrumental dictionary. We also show how our work can be easily extended to accommodate multiple dictionaries using the DSD100 dataset. Tak-Shing Chan, Yi-Hsuan Yang |
IEEE Signal Process. Lett. | 2 |
| 2017 | Cross-Dataset and Cross-Cultural Music Mood Prediction: A Case on Western and Chinese Pop SongsabstractIn music mood prediction, regression models are built to predict values on several mood-representing dimensions such as valence (level of pleasure) and arousal (level of energy). Many studies have shown that music mood is generally predictable based on music acoustic features, but these experiments were mostly conducted on datasets with homogeneous music. Little research has been done to explore the generalizability of mood regression models cross datasets, especially those with music in different cultures. In the increasingly global market of music listening, generalizable models are highly desirable for automated processing, searching and managing music collections with heterogeneous characteristics. In this study, we evaluated mood regression models built on fifteen acoustic features in five mood-related musical aspects, with a focus on cross-dataset generalizability. Specifically, three distinct datasets were involved in a series of five experiments to examine the effects of dataset size, reliability of annotations and cultural backgrounds of music and annotators on mood regression performances and model generalizability. The results reveal that the size of the training dataset and the annotation reliability of the testing dataset affect mood regression performances. When both factors are controlled, regression models are generalizable between datasets sharing a common cultural background of music or annotators. Xiao Hu 0001, Yi-Hsuan Yang |
IEEE Trans. Affect. Comput. | 2 |
| 2017 | Component Tying for Mixture Model Adaptation in Personalization of Music Emotion RecognitionabstractPersonalizing a music emotion recognition model is needed because the perception of music emotion is highly subjective, but it is a time-consuming process. In this paper, we consider how to expedite the personalization process that begins with a general model trained offline using a general user base and progressively adapts the model to a music listener using the emotion annotations of the listener. Specifically, we focus on reducing the number of user annotations needed for the personalization. We investigate and evaluate four component tying methods: single group tying, quadrantwise tying, hierarchical tying, and random tying. These methods aim to exploit the available annotations by identifying related model parameters on-the-fly and updating them jointly. In the evaluation, we use the AMG1608 dataset, which contains the clip-level valence-arousal emotion ratings of 1608 30-s music clips annotated by 665 listeners. Also, we use the acoustic emotion Gaussians model as the general model that uses a mixture of Gaussian components to learn the mapping between the acoustic feature space and the emotion space. The results show that the model adaptation with component tying requires only 10-20 personal annotations to obtain the same level of prediction accuracy as the baseline model adaptation method that uses 50 personal annotations without component tying. Yu-An Chen, Ju-Chiang Wang, Yi-Hsuan Yang, Homer H. Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Introduction to Intelligent Music Systems and ApplicationsabstractIntelligent technologies have become an essential part of music systems and applications. This is evidenced by today's omnipresence of digital online music stores and streaming services, which rely on music recommenders, automatic playlist generators, and music browsing interfaces. A large amount of research leading to intelligent music applications deals with the extraction of musical and acoustic information directly from the audio signal using signal processing techniques. Other strategies exploit contextual aspects of music, not present in the signal, for example, community meta-data and trails of user interaction, as found, for instance, on social media platforms. In this editorial, we discuss the notion of “intelligent music system” and give an overview of the papers selected to this special issue. Markus Schedl, Yi-Hsuan Yang, Perfecto Herrera |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2016 | Event Localization in Music Auto-taggingabstractIn music auto-tagging, people develop models to automatically label a music clip with attributes such as instruments, styles or acoustic properties. Many of these tags are actually descriptors of local events in a music clip, rather than a holistic description of the whole clip. Localizing such tags in time can potentially innovate the way people retrieve and interact with music, but little work has been done to date due to the scarcity of labeled data with granularity specific enough to the frame level. Most labeled data for training a learning-based model for music auto-tagging are in the clip level, providing no cues when and how long these attributes appear in a music clip. To bridge this gap, we propose in this paper a convolutional neural network (CNN) architecture that is able to make accurate frame-level predictions of tags in unseen music clips by using only clip-level annotations in the training phase. Our approach is motivated by recent advances in computer vision for localizing visual objects, but we propose new designs of the CNN architecture to account for the temporal information of music and the variable duration of such local tags in time. We report extensive experiments to gain insights into the problem of event localization in music, and validate through experiments the effectiveness of the proposed approach. In addition to quantitative evaluations, we also present qualitative analyses showing the model can indeed learn certain characteristics of music tags. Jen-Yu Liu, Yi-Hsuan Yang |
ACM Multimedia | 2 |
| 2016 | Query-based Music Recommendations via Preference EmbeddingabstractA common scenario considered in recommender systems is to predict a user's preferences on unseen items based on his/her preferences on observed items. A major limitation of this scenario is that a user might be interested in different things each time when using the system, but there is no way to allow the user to actively alter or adjust the recommended results. To address this issue, we propose the idea of "query-based recommendation" that allows a user to specify his/her search intention while exploring new items, thereby incorporating the concept of information retrieval into recommendation systems. Moreover, the idea is more desirable when the user intention can be expressed in different ways. Take music recommendation as an example: the proposed system allows a user to explore new song tracks by specifying either a track, an album, or an artist. To enable such heterogeneous queries in a recommender system, we present a novel technique called "Heterogeneous Preference Embedding" to encode user preference and query intention into low-dimensional vector spaces. Then, with simple search methods or similarity calculations, we can use the encoded representation of queries to generate recommendations. This method is fairly flexible and it is easy to add other types of information when available. Evaluations on three music listening datasets confirm the effectiveness of the proposed method over the state-of-the-art matrix factorization and network embedding methods. Chih-Ming Chen 0003, Ming-Feng Tsai, Yu-Ching Lin, Yi-Hsuan Yang |
RecSys | 4 |
| 2016 | Addressing Cold Start for Next-song RecommendationabstractThe cold start problem arises in various recommendation applications. In this paper, we propose a tensor factorization-based algorithm that exploits content features extracted from music audio to deal with the cold start problem for the emerging application next-song recommendation. Specifically, the new algorithm learns sequential behavior to predict the next song that a user would be interested in based on the last song the user just listened to. A unique characteristic of the algorithm is that it learns and updates the mapping between the audio feature space and the item latent space each time during the iterations of the factorization process. This way, the content features can be better exploited in forming the latent features for both users and items, leading to more effective solutions for cold-start recommendation. Evaluation on a large-scale music recommendation dataset shows that the recommendation result of the proposed algorithm exhibits not only higher accuracy but also better novelty and diversity, suggesting its applicability in helping a user explore new items in next-item recommendation. Our implementation is available at https://github.com/fearofchou/ALMM. Szu-Yu Chou, Yi-Hsuan Yang, Jyh-Shing Roger Jang, Yu-Ching Lin |
RecSys | 2 |
| 2016 | Complex and Quaternionic Principal Component Pursuit and Its Application to Audio SeparationabstractRecently, the principal component pursuit has received increasing attention in signal processing research ranging from source separation to video surveillance. So far, all existing formulations are real-valued and lack the concept of phase, which is inherent in inputs such as complex spectrograms or color images. Thus, in this letter, we extend principal component pursuit to the complex and quaternionic cases to account for the missing phase information. Specifically, we present both complex and quaternionic proximity operators for the ℓ1- and trace-norm regularizers. These operators can be used in conjunction with proximal minimization methods such as the inexact augmented Lagrange multiplier algorithm. The new algorithms are then applied to the singing voice separation problem, which aims to separate the singing voice from the instrumental accompaniment. Results on the iKala and MSD100 datasets confirmed the usefulness of phase information in principal component pursuit. Tak-Shing Chan, Yi-Hsuan Yang |
IEEE Signal Process. Lett. | 2 |
| 2016 | Monaural Music Source Separation Using Convolutional Sparse CodingabstractWe present a comprehensive performance study of a new time-domain approach for estimating the components of an observed monaural audio mixture. Unlike existing time-frequency approaches that use the product of a set of spectral templates and their corresponding activation patterns to approximate the spectrogram of the mixture, the proposed approach uses the sum of a set of convolutions of estimated activations with prelearned dictionary filters to approximate the audio mixture directly in the time domain. The approximation problem can be solved by an efficient convolutional sparse coding algorithm. The effectiveness of this approach for source separation of musical audio has been demonstrated in our prior work, but under rather restricted and controlled conditions, requiring the musical score of the mixture being informed a priori and little mismatch between the dictionary filters and the source signals. In this paper, we report an evaluation that considers wider, and more practical, experimental settings. This includes the use of an audio-based multipitch estimation algorithm to replace the musical score, and an external dataset of audio single notes to construct the dictionary filters. Our result shows that the proposed approach remains effective with a larger dictionary, and compares favorably with the state-of-the-art nonnegative matrix factorization approach. However, in the absence of the score and in the case of a small dictionary, our approach may not be better. Ping-Keng Jao, Li Su 0004, Yi-Hsuan Yang, Brendt Wohlberg |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Vocal activity informed singing voice separation with the iKala datasetabstractA new algorithm is proposed for robust principal component analysis with predefined sparsity patterns. The algorithm is then applied to separate the singing voice from the instrumental accompaniment using vocal activity information. To evaluate its performance, we construct a new publicly available iKala dataset that features longer durations and higher quality than the existing MIR-1K dataset for singing voice separation. Part of it will be used in the MIREX Singing Voice Separation task. Experimental results on both the MIR-1K dataset and the new iKala dataset confirmed that the more informed the algorithm is, the better the separation results are. Tak-Shing Chan, Tzu-Chun Yeh, Zhe-Cheng Fan, Hung-Wei Chen, Li Su 0004, Yi-Hsuan Yang, Jyh-Shing Roger Jang |
ICASSP | 6 |
| 2015 | The AMG1608 dataset for music emotion recognitionabstractAutomated recognition of musical emotion from audio signals has received considerable attention recently. To construct an accurate model for music emotion prediction, the emotion-annotated music corpus has to be of high quality. It is desirable to have a large number of songs annotated by numerous subjects to characterize the general emotional response to a song. Due to the need for personalization of the music emotion prediction model to address the subjective nature of emotion perception, it is also important to have a large number of annotations per subject for training and evaluating a personalization method. In this paper, we discuss the deficiency of existing datasets and present a new one. The new dataset, which is publically available to the research community, is composed of 1608 30-second music clips annotated by 665 subjects. Furthermore, 46 subjects annotated more than 150 songs, making this dataset the largest of its kind to date. Yu-An Chen, Yi-Hsuan Yang, Ju-Chiang Wang, Homer H. Chen |
ICASSP | 2 |
| 2015 | Informed monaural source separation of music based on convolutional sparse codingabstractMonaural source separation is a challenging problem that has many important applications in music information retrieval. In this paper, we focus on the score-informed variant of this problem. While non-negative matrix factorization and some other approaches have been shown effective, few existing approaches have properly taken the phase information into account. There are unnatural sound in the separation result, as the phase of each source signal is considered equivalent to the phase of the mixed signal. To remedy this, we propose to perform source separation directly in the time domain using a convolutional sparse coding (CSC) approach. Evaluation on the Bach10 dataset shows that, when the instrument, pitch and onset/offset time are informed, the source to distortion ratio of the separation result reaches 8.59 dB, which is 2.02 dB higher than a state-of-the-art system called Soundprism. Ping-Keng Jao, Yi-Hsuan Yang, Brendt Wohlberg |
ICASSP | 2 |
| 2015 | Evaluating music recommendation in a real-world setting: On data splitting and evaluation metricsabstractEvaluation is important to assess the performance of a computer system in fulfilling a certain user need. In the context of recommendation, researchers usually evaluate the performance of a recommender system by holding out a random subset of observed ratings and calculating the accuracy of the system in reproducing such ratings. This evaluation strategy, however, does not consider the fact that in a real-world setting we are actually given the observed ratings of the past and have to predict for the future. There might be new songs, which create the cold-start problem, and the users' musical preference might change over time. Moreover, the user satisfaction of a recommender system may be related to factors other than accuracy. In light of these observations, we propose in this paper a novel evaluation framework that uses various time-based data splitting methods and evaluation metrics to assess the performance of recommender systems. Using millions of listening records collected from a commercial music streaming service, we compare the performance of collaborative filtering (CF) and content-based (CB) models with low-level audio features and semantic audio descriptors. Our evaluation shows that the CB model with semantic descriptors obtains a better trade-off among accuracy, novelty, diversity, freshness and popularity, and can nicely deal with the cold-start problems of new songs. Szu-Yu Chou, Yi-Hsuan Yang, Yu-Ching Lin |
ICME | 2 |
| 2015 | eMosic: Mobile Media Pushing through Social Emotion SensingabstractNo abstract available. Jheng-Wei Peng, Shih-Wei Sun, Wen-Huang Cheng, Yi-Hsuan Yang |
ACM Multimedia | 4 |
| 2015 | ASM'15: The 1st International Workshop on Affect and Sentiment in MultimediaabstractNo abstract available. Mohammad Soleymani 0001, Yi-Hsuan Yang, Yu-Gang Jiang 0001, Shih-Fu Chang |
ACM Multimedia | 2 |
| 2015 | Music Annotation and Retrieval using Unlabeled Exemplars: Correlation and Sparse CodesabstractTagging music signals with semantic labels such as genres, moods and instruments is important for content-based music retrieval and recommendation. While considerable effort has been made, automatic music annotation is still considered challenging due to the difficulty of extracting good audio features that capture the characteristics of different tags. To address this issue, we present in this letter two exemplar-based approaches that represent the content of a music clip by referring to a large set of unlabeled audio exemplars. The first approach represents a music clip by the set of audio exemplars that is highly correlated with the short-time feature vectors of the clip, whereas the second approach represents a music clip as sparse linear combinations of its short-time feature vectors over the audio exemplars. Music annotation is then performed by learning the relevance of the audio examples to different tags using labeled data. These two approaches effectively capitalize the availability of unlabeled data to explore the commonality of music signals to find out tag-specific acoustic patterns, without domain knowledge and feature design. Evaluation on the CAL10k music genre tagging dataset for tag-based music retrieval shows that, with thousands of unlabeled audio examples randomly drawn from the Million Song Dataset, the proposed approaches lead to remarkably higher precision rates than existing approaches. Ping-Keng Jao, Yi-Hsuan Yang |
IEEE Signal Process. Lett. | 2 |
| 2015 | Musical Onset Detection Using Constrained Linear ReconstructionabstractThis letter presents a multi-frame extension of the well-known spectral flux method for unsupervised musical onset detection. Instead of comparing only the spectral content of two frames, the proposed method takes into account a wider temporal context to evaluate the dissimilarity between a given frame and its previous frames. More specifically, the dissimilarity is measured by using the previous frames to obtain a linear reconstruction of the given frame, and then calculating the rectified, l2-norm reconstruction error. Evaluation on a dataset comprising 2,169 onset events of 12 instruments shows that this simple idea works fairly well. When a non-negativity constraint is imposed in the linear reconstruction, the proposed method can outperform the state-of-the-art unsupervised method SuperFlux by 2.9% in F-score. Moreover, the proposed method is particularly effective for instruments with soft onsets, such as violin, cello, and ney. The proposed method is efficient, easy to implement, and is applicable to scenarios of online onset detection. Che-Yuan Liang, Li Su 0004, Yi-Hsuan Yang |
IEEE Signal Process. Lett. | 3 |
| 2015 | Guest Editorial: Challenges and Perspectives for Affective Analysis in MultimediaabstractThe articles in this special section focus on new areas of development in the multimedia industry. Mohammad Soleymani 0001, Yi-Hsuan Yang, Go Irie, Alan Hanjalic |
IEEE Trans. Affect. Comput. | 2 |
| 2015 | Modeling the Affective Content of Music with a Gaussian Mixture ModelabstractModeling the association between music and emotion has been considered important for music information retrieval and affective human computer interaction. This paper presents a novel generative model called acoustic emotion Gaussians (AEG) for computational modeling of emotion. Instead of assigning a music excerpt with a deterministic (hard) emotion label, AEG treats the affective content of music as a (soft) probability distribution in the valence-arousal space and parameterizes it with a Gaussian mixture model (GMM). In this way, the subjective nature of emotion perception is explicitly modeled. Specifically, AEG employs two GMMs to characterize the audio and emotion data. The fitting algorithm of the GMM parameters makes the model learning process transparent and interpretable. Based on AEG, a probabilistic graphical structure for predicting the emotion distribution from music audio data is also developed. A comprehensive performance study over two emotion-labeled datasets demonstrates that AEG offers new insights into the relationship between music and emotion (e.g., to assess the “affective diversity” of a corpus) and represents an effective means of emotion modeling. Readers can easily implement AEG via the publicly available codes. As the AEG model is generic, it holds the promise of analyzing any signal that carries affective or other highly subjective information. Ju-Chiang Wang, Yi-Hsuan Yang, Hsin-Min Wang, Shyh-Kang Jeng |
IEEE Trans. Affect. Comput. | 2 |
| 2015 | Combining Spectral and Temporal Representations for Multipitch Estimation of Polyphonic MusicabstractDue to the difficulty of creating pitch-labeled training data that cover the rich diversity found in music signals, unsupervised feature-based approaches derived from signal processing and feature design remain critical for multipitch estimation (MPE) of polyphonic music. While a large number of feature representations have been proposed in the literature, an effective means of combining different domains of features for MPE is still needed. In this paper, we propose a novel approach, referred to as combined frequency and periodicity (CFP), that detects pitches according to the agreement of a harmonic series in the frequency domain and a subharmonic series in the lag (quefrency) domain. This approach nicely aggregates the complementary advantages of the two feature domains in different frequency ranges, and improves the robustness of the pitch detection function to the interference of the overtones of simultaneous pitches. We report a comprehensive evaluation that compares CFP against three state-of-the-art approaches using three MPE datasets and four symphonies. The evaluation is characteristic of the coverage and complexity of music (in terms of instrument type and degree of polyphony). In addition, we also evaluate the performance of the MPE approaches when a number of audio degradations are applied. Results show that the proposed unsupervised method performs consistently well across the types of Western polyphonic music considered, and is robust to audio degradations such as high-pass filtering and MP3 compression. Li Su 0004, Yi-Hsuan Yang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Quantitative Study of Music Listening Behavior in a Smartphone ContextabstractContext-based services have attracted increasing attention because of the prevalence of sensor-rich mobile devices such as smartphones. The idea is to recommend information that a user would be interested in according to the user’s surrounding context. Although remarkable progress has been made to contextualize music playback, relatively little research has been made using a large collection of real-life listening records collected in situ . In light of this fact, we present in this article a quantitative study of the personal, situational, and musical factors of musical preference in a smartphone context, using a new dataset comprising the listening records and self-report context annotation of 48 participants collected over 3wk via an Android app. Although the number of participants is limited and the population is biased towards students, the dataset is unique in that it is collected in a daily context, with sensor data and music listening profiles recorded at the same time. We investigate 3 core research questions evaluating the strength of a rich set of low-level and high-level audio features for music usage auto-tagging (i.e., music preference in different user activities), the strength of time-domain and frequency-domain sensor features for user activity classification, and how user factors such as personality traits are correlated with the predictability of music usage and user activity, using a closed set of 8 activity classes. We provide an in-depth discussion of the main findings of this study and their implications for the development of context-based music services for smartphones. Yi-Hsuan Yang, Yuan-Ching Teng |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2014 | Linear regression-based adaptation of music emotion recognition models for personalizationabstractPersonalization techniques can be applied to address the subjectivity issue of music emotion recognition, which is important for music information retrieval. However, achieving satisfactory accuracy in personalized music emotion recognition for a user is difficult because it requires an impractically huge amount of annotations from the user. In this paper, we adopt a probabilistic framework for valence-arousal music emotion modeling and propose an adaptation method based on linear regression to personalize a background model in an online learning fashion. We also incorporate a component-tying strategy to enhance the model flexibility. Comprehensive experiments are conducted to test the performance of the proposed method on three datasets, including a new one created specifically in this work for personalized music emotion recognition. Our results demonstrate the effectiveness of the proposed method. Yu-An Chen, Ju-Chiang Wang, Yi-Hsuan Yang, Homer H. Chen |
ICASSP | 3 |
| 2014 | Modified lasso screening for audio word-based music classification using large-scale dictionaryabstractRepresenting music information using audio codewords has led to state-of-the-art performance on various music classifcation benchmarks. Comparing to conventional audio descriptors, audio words offer greater fexibility in capturing the nuance of music signals, in that each codeword can be viewed as a quantization of the music universe and that the quantization goes finer as the size of the dictionary (i.e., audio codebook) increases. In practice, however, the high computational cost of codeword assignment might discourage the use of a large dictionary. This paper presents two modifications of a LASSO screening technique developed in the compressive sensing field to speed up the codeword assignment process. The first modification exploits the repetitive nature of music signals, whereas the second one relaxes a screening constraint that is specific to reconstruction but not for classifcation. Our experiments show that the proposed method enables the use of a dictionary of 10,000 codewords with runtime close to the case of using a dictionary of 1,000 codewords. Moreover, using the larger dictionary significantly improves the mean average precision (MAP) from 0.219 to 0.246 for tagging thousands of tracks with 147 possible genre tags. Ping-Keng Jao, Chin-Chia Michael Yeh, Yi-Hsuan Yang |
ICASSP | 3 |
| 2014 | Improving music auto-tagging by intra-song instance baggingabstractBagging is one the most classic ensemble learning techniques in the machine learning literature. The idea is to generate multiple subsets of the training data via bootstrapping (random sampling with replacement), and then aggregate the output of the models trained from each subset via voting or averaging. As music is a temporal signal, we propose and study two bagging methods in this paper: the inter-song instance bagging that bootstraps song-level features, and the intra-song instance bagging that draws bootstrapping samples directly from short-time features for each training song. In particular, we focus on the latter method, as it better exploits the temporal information of music signals. The bagging methods result in surprisingly effective models for music auto-tagging: incorporating the idea to a simple linear support vector machine (SVM) based system yields accuracies that are comparable or even superior to state-of-the-art, possibly more sophisticated methods for three different datasets. As the bagging method is a meta algorithm, it holds the promise of improving other MIR systems. Chin-Chia Michael Yeh, Ju-Chiang Wang, Yi-Hsuan Yang, Hsin-Min Wang |
ICASSP | 3 |
| 2014 | Sparse cepstral codes and power scale for instrument identificationabstractThis paper presents a novel feature representation called sparse cepstral codes for instrument identification. We first motivate the approach by discussing why cepstrum is suitable for instrument identification. Then we propose the use of sparse coding and power normalization to derive compact codes that better represent the information of the cepstrum. Our evaluation on both uni-source and multi-source instrument identification tasks show that the proposed feature leads to significantly better accuracy than existing methods. We further show that cepstrum obtained from power-scaled spectrum can do better than typical cepstrum especially in multi-source signal. The proposed system achieves 0.955 F-score in uni-source dataset and 0.688 F-score in multi-source dataset. Li-Fan Yu, Li Su 0004, Yi-Hsuan Yang |
ICASSP | 3 |
| 2014 | LJ2M dataset: Toward better understanding of music listening behavior and user moodabstractRecent years have witnessed a growing interest in modeling user behaviors in multimedia research, emphasizing the need to consider human factors such as preference, activity, and emotion in system development and evaluation. Following this research line, we present in this paper the LiveJournal two-million post (LJ2M) dataset to foster research on user-centered music information retrieval. The new dataset is characterized by the great diversity of real-life listening contexts where people and music interact. It contains blog articles from the social blogging website LiveJournal, along with tags self-reporting a user's emotional state while posting and the musical track that the user considered as the best match for the post. More importantly, the data are contributed by users spontaneously in their daily lives, instead of being collected in a controlled environment. Therefore, it offers new opportunities to understand the interrelationship among the personal, situational, and musical factors of music listening. As an example application, we present research investigating the interaction between the affective context of the listener and the affective content of music, using audio-based music emotion recognition techniques and a psycholinguistic tool. The study offers insights into the role of music in mood regulation and demonstrates how LJ2M can contribute to studies on real-world music listening behavior. Jen-Yu Liu, Sung-Yen Liu, Yi-Hsuan Yang |
ICME | 3 |
| 2014 | Towards time-varying music auto-tagging based on CAL500 expansionabstractMusic auto-tagging refers to automatically assigning semantic labels (tags) such as genre, mood and instrument to music so as to facilitate text-based music retrieval. Although significant progress has been made in recent years, relatively little research has focused on semantic labels that are time-varying within a track. Existing approaches and datasets usually assume that different fragments of a track share the same tag labels, disregarding the tags that are time-varying (e.g., mood) or local in time (e.g., instrument solo). In this paper, we present a new dataset dedicated to time-varying music auto-tagging. The dataset, called CAL500exp, is an enriched version of the well-known CAL500 dataset used for conventional track-level tagging. Given the tag set of CAL500, eleven subjects with strong music background were recruited to annotate the time-varying tag labels. A new user interface for annotation is developed to reduce the subject's annotation effort yet increase the quality of labels. Moreover, we present an empirical evaluation that demonstrates the performance improvement CAL500exp brings about for time-varying music auto-tagging. By providing more accurate and consistent descriptions of music content in a finer granularity, CAL500exp may open new opportunities to understand and to model the temporal context of musical semantics. Shuo-Yang Wang, Ju-Chiang Wang, Yi-Hsuan Yang, Hsin-Min Wang |
ICME | 3 |
| 2014 | Music Driven Human Motion Manipulation for Characters in a VideoabstractMultimedia content creation and manipulation have garnered attention in recent days due to the desires of personalization. As a content producing application, we propose a novel idea that requires the fusion of video and audio intelligence. The system is composed of at least three core techniques: 1) the capability to process the video sequence to have access to the geometric and appearance information pertaining to meaningful and representative targets, 2) a systematic way to reliably classify and identify important emotions from the music, 3) effective approaches to manipulate the video targets according to the extracted music emotions. In this paper, we report preliminary results of the proposed system. Specifically, we introduce the employed framework to manipulate the magnitude and speed of music conducting gestures of a video sequence of human skeleton according to the emotion intensity and tempo of an arbitrary music excerpt, using state-of-the-art inverse kinematics and music information retrieval techniques. We present the details of the prototype system and validate its effectiveness with a video demonstrating how we can manipulate the music conducting gestures according to the proposed manipulation rules. Che-Hua Yeh, Yi-Hsuan Yang, Ming-Hsu Chang, Hong-Yuan Mark Liao |
ISM | 2 |
| 2014 | Emotional Analysis of Music: A Comparison of MethodsabstractMusic as a form of art is intentionally composed to be emotionally expressive. The emotional features of music are invaluable for music indexing and recommendation. In this paper we present a cross-comparison of automatic emotional analysis of music. We created a public dataset of Creative Commons licensed songs. Using valence and arousal model, the songs were annotated both in terms of the emotions that were expressed by the whole excerpt and dynamically with 1 Hz temporal resolution. Each song received 10 annotations on Amazon Mechanical Turk and the annotations were averaged to form a ground truth. Four different systems from three teams and the organizers were employed to tackle this problem in an open challenge. We compare their performances and discuss the best practices. While the effect of a larger feature set was not very apparent in the static emotion estimation, the combination of a comprehensive feature set and a recurrent neural network that models temporal dependencies has largely outperformed the other proposed methods for dynamic music emotion estimation. Mohammad Soleymani 0001, Anna Aljanaki, Yi-Hsuan Yang, Michael N. Caro, Florian Eyben, Konstantin Markov, Björn W. Schuller, Remco C. Veltkamp, Felix Weninger, Frans Wiering |
ACM Multimedia | 3 |
| 2014 | AWtoolbox: Characterizing Audio Information Using Audio WordsabstractThis paper presents the AWtoolbox, an open-source software designed for extracting the audio word (AW) representation of audio signals. The toolbox comes with a graphical user interface that helps a user design custom AW extraction pipelines and various algorithms for feature encoding, dictionary learning, result rectification, pooling, normalization and others. This paper also reports a benchmark comparing eight AW representations computed by the toolbox against state-of-the-art low-level and mid-level timbre, rhythmic and tonal descriptors of music and sound. The evaluation result shows that sparse coding (SC) based AW representation leads to very competitive performances across the three tested sound and music classification tasks. AWtoolbox is available for download at http://mac.citi.sinica.edu.tw/awtoolbox. Chin-Chia Michael Yeh, Ping-Keng Jao, Yi-Hsuan Yang |
ACM Multimedia | 3 |
| 2014 | Sparse modeling of magnitude and phase-derived spectra for playing technique classificationabstractComputational modeling of musical timbre is important for a variety of music information retrieval applications. While considerable progress has been made to recognize musical genres and instruments, relatively little attention has been paid to modeling playing techniques, which affect timbre in more subtle ways. In this paper, we contribute to this area of research by systematically evaluating various audio features and processing methods for multi-class playing technique classification, considering up to nine distinct playing techniques of bowed string instruments. Specifically, a collection of 6,759 chamber-recorded single notes of four bowed string instruments and a collection of 33 real-world solo violin recordings are used in the evaluation. Our evaluation shows that using sparse features extracted from the magnitude spectra and phase derivatives including group delay function (GDF) and instantaneous frequency deviation (IFD) leads to significantly better performance than using a combination of state-of-the-art temporal, spectral, cepstral and harmonic feature descriptors. For playing technique classification of violin singe notes, the former approach attains 0.915 macro-average F-score under a tenfold cross validation setting, while the latter only attains 0.835. Moreover, sparse modeling of magnitude and phase-derived spectra also performs well for single-note joint instrument-technique classification (F-score 0.770) and for playing technique classification of real-world violin solos (F-score 0.547). We find that phase information is particularly important in discriminating playing techniques with subtle differences, such as playing with different bowing positions (i.e., normal, sul tasto, and sul ponticello). A systematic investigation of the effect of parameters such as window sizes, hop factors, window types for phase-derived features is also reported to provide more insights. Li Su 0004, Hsin-Ming Lin, Yi-Hsuan Yang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | A Systematic Evaluation of the Bag-of-Frames Representation for Music Information RetrievalabstractThere has been an increasing attention on learning feature representations from the complex, high-dimensional audio data applied in various music information retrieval (MIR) problems. Unsupervised feature learning techniques, such as sparse coding and deep belief networks have been utilized to represent music information as a term-document structure comprising of elementary audio codewords. Despite the widespread use of such bag-of-frames (BoF) model, few attempts have been made to systematically compare different component settings. Moreover, whether techniques developed in the text retrieval community are applicable to audio codewords is poorly understood. To further our understanding of the BoF model, we present in this paper a comprehensive evaluation that compares a large number of BoF variants on three different MIR tasks, by considering different ways of low-level feature representation, codebook construction, codeword assignment, segment-level and song-level feature pooling, tf-idf term weighting, power normalization, and dimension reduction. Our evaluations lead to the following findings: 1) modeling music information by two levels of abstraction improves the result for difficult tasks such as predominant instrument recognition, 2) tf-idf weighting and power normalization improve system performance in general, 3) topic modeling methods such as latent Dirichlet allocation does not work for audio codewords. Li Su 0004, Chin-Chia Michael Yeh, Jen-Yu Liu, Ju-Chiang Wang, Yi-Hsuan Yang |
IEEE Trans. Multim. | 5 |
| 2013 | Singing voice timbre classification of Chinese popular musicabstractSinging voice plays an important role in the listening experience of music. In this paper, we propose to classify popular music by the timbre quality of the singing voice. Specifically, we adopt six singing voice timbre classes as the taxonomy and build a new data set, KKTIC, that contains the expert annotations of 387 Chinese popular songs. To build an automatic classifier, we resort to signal processing and machine learning techniques and extract a number of singing voice-related features such as vibrato and harmonic-to-noise ratio. We also propose the use of vocal segment detection and singing voice separation as preprocessing steps. Our evaluation identifies the relevant acoustic features and validates the importance of these preprocessing steps. The accuracy in timbre classification reaches 79.84% in a five-fold stratified cross validation. Cheng-Ya Sha, Yi-Hsuan Yang, Yu-Ching Lin, Homer H. Chen |
ICASSP | 2 |
| 2013 | Dual-layer bag-of-frames model for music genre classificationabstractThis paper concerns the development of a music dictionary-based model for summarizing local feature descriptors computed over time. Comparing to a holistic representation, this text-like, bag-of-frames representation better captures the rich and time-varying information of music. However, the dictionary used in classical bag-of-frames model only captures frame-level elements of the music; thus, there exists a semantic gap between the dictionary element and commonly seen music description. In order to reduce the gap, a new feature representation called dual-layer bag-of-frames is proposed in this paper. It models the music with a two layer structure, where the first-layer dictionary captures the frame-level characteristics, and the second-layer dictionary captures the segment-level semantics. This hierarchical structure resembles the alphabet-word-document structure of text. Our result demonstrates that the proposed dual-layer bag-of-frames feature achieves state-of-the-art accuracy of music genre classification. The classification accuracy for the GTZAN benchmark reaches 86.7% with dictionary trained from GTZAN, and 83.6% with dictionary trained from another data set USPOP. Chin-Chia Michael Yeh, Li Su 0004, Yi-Hsuan Yang |
ICASSP | 3 |
| 2013 | Towards real-time music auto-tagging using sparse featuresabstractUnsupervised feature learning algorithms such as sparse coding and deep belief networks have been shown a viable alternative to hand-crafted feature design for music information retrieval. Nevertheless, such algorithms are usually computationally expensive. This paper investigates techniques to accelerate sparse feature extraction and music classification. To study the trade-off between computational efficiency and accuracy, we compare state-of-the-art, dense audio features with sparse features computed using 1) sparse coding with a random dictionary, 2) randomized clustering forest, and 3) an extension of randomized clustering forest to temporal signals. For classifier training and prediction, we compare support vector machines with linear or non-linear kernel functions. We conduct evaluation on music auto-tagging for 140 genre/style tags using a subset of 7,799 songs of the CAL10k data set. Our result leads to an 11-fold speed increase with 3.45% accuracy loss comparing to dense features. With the proposed sparse features, the feature extraction and auto-tagging operations can be finished in 1 second per song, with 0.1302 tagging accuracy in mean average precision. Yi-Hsuan Yang |
ICME | 1 |
| 2013 | Using emotional context from article for contextual music recommendationabstractThis paper proposes a context-aware approach that recommends music to a user based on the user's emotional state predicted from the article the user writes. We analyze the association between user-generated text and music by using a real-world dataset with user, text, music tripartite information collected from the social blogging website LiveJournal. The audio information represents various perceptual dimensions of music listening, including danceability, loudness, mode, and tempo; the emotional text information consists of bag-of-words and three dimensional affective states within an article: valence, arousal and dominance. To combine these factors for music recommendation, a factorization machine-based approach is taken. Our evaluation shows that the emotional context information mined from user-generated articles does improve the quality of recommendation, comparing to either the collaborative filtering approach or the content-based approach. Chih-Ming Chen 0003, Ming-Feng Tsai, Jen-Yu Liu, Yi-Hsuan Yang |
ACM Multimedia | 4 |
| 2013 | Music Recommendation Based on Multiple Contextual Similarity InformationabstractThis paper proposes a music recommendation approach based on various similarity information via Factorization Machines (FM). We introduce the idea of similarity, which has been widely studied in the filed of information retrieval, and incorporate multiple feature similarities into the FM framework, including content-based and context-based similarities. The similarity information not only captures the similar patterns from the referred objects, but enhances the convergence speed and accuracy of FM. In addition, in order to avoid the noise within large similarity of features, we also adopt the grouping FM as an extended method to model the problem. In our experiments, a music-recommendation dataset is used to assess the performance of the proposed approach. The datasets is collected from an online blogging Web site, which includes user listening history, user profiles, social information, and music information. Our experimental results show that, with various types of feature similarities the performance of music recommendation can be enhanced significantly. Furthermore, via the grouping technique, the performance can be improved significantly in terms of Mean Average Precision, compared to the traditional collaborative filtering approach. Chih-Ming Chen 0003, Ming-Feng Tsai, Jen-Yu Liu, Yi-Hsuan Yang |
Web Intelligence | 4 |
| 2013 | Automatic highlights extraction for drama video using music emotion and human face features
Keng-Sheng Lin, Ann Lee 0002, Yi-Hsuan Yang, Cheng-Te Lee, Homer H. Chen |
Neurocomputing | 3 |
| 2013 | Quantitative Study of Music Listening Behavior in a Social and Affective ContextabstractA scientific understanding of emotion experience requires information on the contexts in which the emotion is induced. Moreover, as one of the primary functions of music is to regulate the listener's mood, the individual's short-term music preference may reveal the emotional state of the individual. In light of these observations, this paper presents the first scientific study that exploits the online repository of social data to investigate the connections between a blogger's emotional state, user context manifested in the blog articles, and the content of the music titles the blogger attached to the post. A number of computational models are developed to evaluate the accuracy of different content or context cues in predicting emotional state, using 40,000 pieces of music listening records collected from the social blogging website LiveJournal. Our study shows that it is feasible to computationally model the latent structure underlying music listening and mood regulation. The average area under the receiver operating characteristic curve (AUC) for the content-based and context-based models attains 0.5462 and 0.6851, respectively. The association among user mood, music emotion, and individual's personality is also identified. Yi-Hsuan Yang, Jen-Yu Liu |
IEEE Trans. Multim. | 1 |
| 2012 | Supervised dictionary learning for music genre classificationabstractThis paper concerns the development of a music codebook for summarizing local feature descriptors computed over time. Comparing to a holistic representation, this text-like representation better captures the rich and time-varying information of music. We systematically compare a number of existing codebook generation techniques and also propose a new one that incorporates labeled data in the dictionary learning process. Several aspects of the encoding system such as local feature extraction and codeword encoding are also analyzed. Our result demonstrates the superiority of sparsity-enforced dictionary learning over conventional VQ-based or exemplar-based methods. With the new supervised dictionary learning algorithm and the optimal settings inferred from the performance study, we achieve state-of-the-art accuracy of music genre classification using just the log-power spectrogram as the local feature descriptor. The classification accuracies for benchmark datasets GTZAN and IS-MIR2004Genre are 84.7% and 90.8%, respectively. Chin-Chia Michael Yeh, Yi-Hsuan Yang |
ICMR | 2 |
| 2012 | Bilingual analysis of song lyrics and audio wordsabstractThanks to the development of music audio analysis, state-of-the-art techniques can now detect musical attributes such as timbre, rhythm, and pitch with certain level of reliability and effectiveness. An emerging body of research has begun to model the high-level perceptual properties of music listening, including the mood and the preferable listening context of a music piece. Towards this goal, we propose a novel text-like feature representation that encodes the rich and time-varying information of music using a composite of features extracted from the song lyrics and audio signals. In particular, we investigate dictionary learning algorithms to optimize the generation of local feature descriptors and also probabilistic topic models to group semantically relevant text and audio words. This text-like representation leads to significant improvement in automatic mood classification over conventional audio features. Jen-Yu Liu, Chin-Chia Michael Yeh, Yi-Hsuan Yang, Yuan-Ching Teng |
ACM Multimedia | 3 |
| 2012 | The acousticvisual emotion guassians model for automatic generation of music videoabstractThis paper presents a novel content-based system that utilizes the perceived emotion of multimedia content as a bridge to connect music and video. Specifically, we propose a novel machine learning framework, called Acousticvisual Emotion Guassians (AVEG), to jointly learn the tripartite relationship among music, video, and emotion from an emotion-annotated corpus of music videos. For a music piece (or a video sequence), the AVEG model is applied to predict its emotion distribution in a stochastic emotion space from the corresponding low-level acoustic (resp. visual) features. Finally, music and video are matched by measuring the similarity between the two corresponding emotion distributions, based on a distance measure such as KL divergence. Ju-Chiang Wang, Yi-Hsuan Yang, I-Hong Jhuo, Yen-Yu Lin, Hsin-Min Wang |
ACM Multimedia | 2 |
| 2012 | The acoustic emotion gaussians model for emotion-based music annotation and retrievalabstractOne of the most exciting but challenging endeavors in music research is to develop a computational model that comprehends the affective content of music signals and organizes a music collection according to emotion. In this paper, we propose a novel acoustic emotion Gaussians (AEG) model that defines a proper generative process of emotion perception in music. As a generative model, AEG permits easy and straightforward interpretations of the model learning processes. To bridge the acoustic feature space and music emotion space, a set of latent feature classes, which are learned from data, is introduced to perform the end-to-end semantic mappings between the two spaces. Based on the space of latent feature classes, the AEG model is applicable to both automatic music emotion annotation and emotion-based music retrieval. To gain insights into the AEG model, we also provide illustrations of the model learning process. A comprehensive performance study is conducted to demonstrate the superior accuracy of AEG over its predecessors, using two emotion annotated music corpora MER60 and MTurk. Our results show that the AEG model outperforms the state-of-the-art methods in automatic music emotion annotation. Moreover, for the first time a quantitative evaluation of emotion-based music retrieval is reported. Ju-Chiang Wang, Yi-Hsuan Yang, Hsin-Min Wang, Shyh-Kang Jeng |
ACM Multimedia | 2 |
| 2012 | On sparse and low-rank matrix decomposition for singing voice separationabstractOver recent years there has been a growing interest in finding ways to transform signals/matrices into sparse or low-rank representations, i.e., representations which are sparse in support or of low redundancy. Such decompositions are proving to be particularly powerful for a variety of signal processing and compression problems. In this paper, we investigate the application of this technique to the challenging task of singing voice/accompaniment separation for popular music. The vocal part is modeled as a sparse signal, whereas the instrumental part is considered to be low-rank. In addition, to better account for the particular properties of music, two new algorithms are proposed to improve the decomposition, including the incorporation of harmonicity priors and a back-end drum removal procedure. Evaluations on the MIR-1K benchmark dataset show that the proposed algorithms outperform the state-of-the-art by 0.01-2.41 db. Yi-Hsuan Yang |
ACM Multimedia | 1 |
| 2012 | Machine Recognition of Music Emotion: A ReviewabstractThe proliferation of MP3 players and the exploding amount of digital music content call for novel ways of music organization and retrieval to meet the ever-increasing demand for easy and effective information access. As almost every music piece is created to convey emotion, music organization and retrieval by emotion is a reasonable way of accessing music information. A good deal of effort has been made in the music information retrieval community to train a machine to automatically recognize the emotion of a music signal. A central issue of machine recognition of music emotion is the conceptualization of emotion and the associated emotion taxonomy. Different viewpoints on this issue have led to the proposal of different ways of emotion annotation, model training, and result visualization. This article provides a comprehensive review of the methods that have been proposed for music emotion recognition. Moreover, as music emotion recognition is still in its infancy, there are many open issues. We review the solutions that have been proposed to address these issues and conclude with suggestions for further research. Yi-Hsuan Yang, Homer H. Chen |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2012 | Multipitch Estimation of Piano Music by Exemplar-Based Sparse RepresentationabstractPitch, together with other midlevel music features such as rhythm and timbre, holds the promise of bridging the semantic gap between low-level features and high-level semantics for music understanding. This paper investigates the pitch estimation of a piano music signal by exemplar-based sparse representation. A note exemplar is a segment of a piano note, stored in the dictionary. We first describe how to represent a segment of the piano music signal as a linear combination of a small number of note exemplars from a large note exemplar dictionary and then show how the sparse representation problem can be solved by -regularized minimization. The proposed approach incorporates tuning factor estimation, note candidate selection, and hidden-Markov-model-based smoothing into the estimation process to improve accuracy. Unlike previous approaches, the proposed approach does not require retraining for a new piano. Instead, only a dozen notes of the new piano are needed. This feature is computationally attractive and avoids intense manual labeling. The system performance is evaluated using 70 classical music recordings of two real pianos under different recording conditions. The results show that the proposed system outperforms four state-of-the-art systems. Cheng-Te Lee, Yi-Hsuan Yang, Homer H. Chen |
IEEE Trans. Multim. | 2 |
| 2011 | Unsupervised auxiliary visual words discovery for large-scale image object retrievalabstractImage object retrieval-locating image occurrences of specific objects in large-scale image collections-is essential for manipulating the sheer amount of photos. Current solutions, mostly based on bags-of-words model, suffer from low recall rate and do not resist noises caused by the changes in lighting, viewpoints, and even occlusions. We propose to augment each image with auxiliary visual words (AVWs), semantically relevant to the search targets. The AVWs are automatically discovered by feature propagation and selection in textual and visual image graphs in an unsupervised manner. We investigate variant optimization methods for effectiveness and scalability in large-scale image collections. Experimenting in the large-scale consumer photos, we found that the the proposed method significantly improves the traditional bag-of-words (111% relatively). Meanwhile, the selection process can also notably reduce the number of features (to 1.4%) and can further facilitate indexing in large-scale image object retrieval. Yin-Hsi Kuo, Hsuan-Tien Lin, Wen-Huang Cheng, Yi-Hsuan Yang, Winston H. Hsu |
CVPR | 4 |
| 2011 | Automatic transcription of piano music by sparse representation of magnitude spectraabstractAssuming that the waveforms of piano notes are pre-stored and that the magnitude spectrum of a piano signal segment can be represented as a linear combination of the magnitude spectra of the pre-stored piano waveforms, we formulate the automatic transcription of polyphonic piano music as a sparse representation problem. First, the note candidates of the piano signal segment are found by using heuristic rules. Then, the sparse representation problem is solved by l1-regularized minimization, followed by temporal smoothing the frame-level results based on hidden Markov models. Evaluation against three state-of-the-art systems using ten classical music recordings of a real piano is performed to show the performance improvement of the proposed system. Cheng-Te Lee, Yi-Hsuan Yang, Homer H. Chen |
ICME | 2 |
| 2011 | Automatic highlights extraction for drama video using music emotion and human face featuresabstractThe rich emotion part of a drama video is often the center of attraction to the viewer. Emotion-based highlights extraction is useful for applications such as drama video retrieval and automatic trailer generation. In this paper, we propose a system that uses music emotion and human face as features for automatic extraction of the emotion highlights of a drama video. These high-level audiovisual features are used because music invokes emotion response from the viewer and characters express emotion on their faces. To avoid the interference of speech signal and environmental noise, a novel two-stage music emotion recognition scheme is developed. We first detect the presence of incidental music in a drama video using an audio fingerprint technique, and then perform emotion recognition on the noise-free music available from the album of the incidental music. This simple but effective approach greatly improves the accuracy of music emotion recognition. Besides the conventional subjective evaluation, we propose a new metric for quantitative performance evaluation of highlights extraction. Evaluation results are provided to illustrate the performance of the system. Keng-Sheng Lin, Ann Lee 0002, Yi-Hsuan Yang, Cheng-Te Lee, Homer H. Chen |
MMSP | 3 |
| 2011 | Ranking-Based Emotion Recognition for Music Organization and RetrievalabstractDetermining the emotion of a song that best characterizes the affective content of the song is a challenging issue due to the difficulty of collecting reliable ground truth data and the semantic gap between human's perception and the music signal of the song. To address this issue, we represent an emotion as a point in the Cartesian space with valence and arousal as the dimensions and determine the coordinates of a song by the relative emotion of the song with respect to other songs. We also develop an RBF-ListNet algorithm to optimize the ranking-based objective function of our approach. The cognitive load of annotation, the accuracy of emotion recognition, and the subjective quality of the proposed approach are extensively evaluated. Experimental results show that this ranking-based approach simplifies emotion annotation and enhances the reliability of the ground truth. The performance of our algorithm for valence recognition reaches 0.326 in Gamma statistic. Yi-Hsuan Yang, Homer H. Chen |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | Prediction of the Distribution of Perceived Music Emotions Using Discrete SamplesabstractTypically, a machine learning model of automatic music emotion recognition is trained to learn the relationship between music features and perceived emotion values. However, simply assigning an emotion value to a clip in the training phase does not work well because the perceived emotion of a clip varies from person to person. To resolve this problem, we propose a novel approach that represents the perceived emotion of a clip as a probability distribution in the emotion plane. In addition, we develop a methodology that predicts the emotion distribution of a clip by estimating the emotion mass at discrete samples of the emotion plane. We also develop model fusion algorithms to integrate different perceptual dimensions of music listening and to enhance the modeling of emotion perception. The effectiveness of the proposed approach is validated through an extensive performance study. An averageR2statistics of 0.5439 for emotion prediction is achieved. We also show how this approach can be applied to enhance our understanding of music emotion. Yi-Hsuan Yang, Homer H. Chen |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | Exploiting online music tags for music emotion classificationabstractThe online repository of music tags provides a rich source of semantic descriptions useful for training emotion-based music classifier. However, the imbalance of the online tags affects the performance of emotion classification. In this paper, we present a novel data-sampling method that eliminates the imbalance but still takes the prior probability of each emotion class into account. In addition, a two-layer emotion classification structure is proposed to harness the genre information available in the online repository of music tags. We show that genre-based grouping as a precursor greatly improves the performance of emotion classification. On the average, the incorporation of online genre tags improves the performance of emotion classification by a factor of 55% over the conventional single-layer system. The performance of our algorithm for classifying 183 emotion classes reaches 0.36 in example-based f-score. Yu-Ching Lin, Yi-Hsuan Yang, Homer H. Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2010 | A technical demonstration of large-scale image object retrieval by efficient query evaluation and effective auxiliary visual feature discoveryabstractIn this demonstration, we present a real-time system that addresses three essential issues of large-scale image object retrieval: 1) image object retrieval-facilitating pseudo-objects in inverted indexing and novel object-level pseudo-relevance feedback for retrieval accuracy; 2) time efficiency-boosting the time efficiency and memory usage of object-level image retrieval by a novel inverted indexing structure and efficient query evaluation; 3) recall rate improvement--mining semantically relevant auxiliary visual features through visual and textual clusters in an unsupervised and scalable (i.e., MapReduce) manner. We are able to search over one-million image collection in respond to a user query in 121ms, with significantly better accuracy (+99%) than the traditional bag-of-words model. Yin-Hsi Kuo, Yi-Lun Wu, Kuan-Ting Chen, Yi-Hsuan Yang, Tzu-Hsuan Chiu, Winston H. Hsu |
ACM Multimedia | 4 |
| 2009 | Music emotion rankingabstractContent-based retrieval has emerged as a promising approach to information access. In this paper, we propose an approach to music emotion ranking. Specifically, we rank music in terms of arousal and valence and represent each song as a point in the 2D emotion space. Novel ranking-based methods for annotation, learning, and evaluation of music emotion recognition are developed and tested on a moderately large-scale database composed of 1240 pop songs. Results are provided to show the feasibility of the proposed approach. Yi-Hsuan Yang, Homer H. Chen |
ICASSP | 1 |
| 2009 | Exploiting genre for music emotion classificationabstractGenre and emotion have been applied to content-based music retrieval and organization; however, the intrinsic correlation between them has not been explored. In this paper we present a statistical association analysis to examine such intrinsic correlation and propose a two-layer scheme that exploits the correlation for emotion classification. Significant improvement of classification accuracy over the traditional single-layer scheme is obtained. Yu-Ching Lin, Yi-Hsuan Yang, Homer H. Chen, I-Bin Liao, Yeh-Chin Ho |
ICME | 2 |
| 2009 | Clustering for music search resultsabstractClustering for better representation of the diversity of text or image search results has been studied extensively. In this paper, we extend this methodology to the novel domain of music search. We conduct empirical evaluation of different clustering algorithms, audio feature representations, and the incorporation of lyrics for music clustering. Our evaluation shows the fusion of audio and text features yields the best clustering accuracy. Yi-Hsuan Yang, Yu-Ching Lin, Homer H. Chen |
ICME | 1 |
| 2009 | Multimodal Structure Segmentation and Analysis of Music using Audio and Textual InformationabstractIn this paper, we present a multimodal approach to structure segmentation of music with applications to audio content analysis and music information retrieval. In particular, since lyrics contain rich information about the semantic structure of a song, our approach incorporates lyrics to overcome the existing difficulties associated with large acoustic variation in music. We further design a constrained clustering algorithm for music segmentation and evaluate its performance on commercial recordings. Experimental results show that our method can effectively detect the boundaries and the types of semantic structure of music segments. Heng Tze Cheng, Yi-Hsuan Yang, Yu-Ching Lin, Homer H. Chen |
ISCAS | 2 |
| 2009 | Canonical image selection and efficient image graph construction for large-scale flickr photosabstractEfficient image search clustering is prominent for image search engines for exponentially growing photo collections. In this work, we propose an image search clustering approach which selects multiple canonical images from image search results and constructs image clusters in real time on an image sub-graph for the search results. The efficiency is achieved with the help of offline-computed image context graphs by distributed computing methods. Extending our prior works, we demonstrate the results of the proposed canonical image selection and preliminary outcomes of large-scale image graph construction in this proposal. We experiment in Flickr550 dataset, containing 540,321 Flickr photos. Liang-Chi Hsieh, Kuan-Ting Chen, Chien-Hsing Chiang, Yi-Hsuan Yang, Guan-Long Wu, Chun-Sung Ferng, Hsiu-Wen Hsueh, Angela Charng-Rurng Tsai, Winston H. Hsu |
ACM Multimedia | 4 |
| 2009 | Personalized music emotion recognitionabstractIn recent years, there has been a dramatic proliferation of research on information retrieval based on highly subjective concepts such as emotion, preference and aesthetic. Such retrieval methods are fascinating but challenging since it is difficult to built a general retrieval model that performs equally well to everyone. In this paper, we propose two novel methods, bag-of-users model and residual modeling, to accommodate the individual differences for emotion-based music retrieval. The proposed methods are intuitive and generally applicable to other information retrieval tasks that involve subjective perception. Evaluation result shows the effectiveness of the proposed methods. Yi-Hsuan Yang, Yu-Ching Lin, Homer H. Chen |
SIGIR | 1 |
| 2009 | Online Reranking via Ordinal Informative Concepts for Context Fusion in Concept Detection and Video SearchabstractTo exploit the co-occurrence patterns of semantic concepts while keeping the simplicity of context fusion, a novel reranking approach is proposed in this paper. The approach, called ordinal reranking, adjusts the ranking of an initial search (or detection) list based on the co-occurrence patterns obtained by using ranking functions such as ListNet. Ranking functions are by nature more effective than classification-based reranking methods in mining ordinal relationships. In addition, the ordinal reranking is free of thead hocthresholding for noisy binary labels and requires no extra offline learning or training data. To select informative concepts for reranking, we also propose a new concept selection measurement,wc-tf-idf, which considers the underlying ordinal information of ranking lists and is thus more effective than the feature selection algorithms for classification. Being largely unsupervised, the reranking approach to context fusion can be applied equally well to concept detection and video search. While being extremely efficient, ordinal reranking outperforms existing methods by up to 40% in mean average precision (MAP) for the baseline text-based search and 12% for the baseline concept detection over TRECVID 2005 video search and concept detection benchmark. Yi-Hsuan Yang, Winston H. Hsu, Homer H. Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | Smooth Control of Adaptive Media Playout for Video StreamingabstractClient-side data buffering is a common technique to deal with media playout interruptions of streaming video caused by network jitters and packet losses of best-effort networks. However, stronger playout interruption protection inevitably amounts to larger data buffering and results in more memory requirements and longer playout delay. Adaptive media playout (AMP), also a client-side technique, can reduce the buffer requirement and avoid buffer outage but at the expense of visual quality degradation because of the fluctuation of playout speed. In this paper, we propose a novel AMP scheme to keep the video playout as smooth as possible while adapting to the channel condition. The triggering of the playout control is based on buffer variation rather than buffer fullness. Experimental results show that our AMP scheme surpasses conventional schemes in unfriendly network conditions. Unlike previous schemes that are tuned for a specific range of packet loss and network instability, the proposed AMP scheme maintains consistent performance across a wide range of network conditions. Ya-Fan Su, Yi-Hsuan Yang, Meng-Ting Lu, Hsin-Hsi Chen |
IEEE Trans. Multim. | 2 |
| 2008 | Video search reranking via online ordinal rerankingabstractTo exploit co-occurrence patterns among features and target semantics while keeping the simplicity of the keyword-based visual search, a novel reranking methods is proposed. The approach, ordinal reranking, reranks an initial search list by utilizing the co-occurrence patterns via the ranking functions such as ListNet. Ranking functions are by nature more effective than classification-based reranking methods in mining ordinal relationships. In addition, ordinal reranking is ease of the ad-hoc thresholding for noisy binary labels and requires no extra off-line learning or training data. When evaluated in TRECVID search benchmark, ordinal reranking, while being extremely efficient, outperforms existing methods and offers 35.6% relative improvement over the text-based search baseline in nearly real time. Yi-Hsuan Yang, Winston H. Hsu |
ICME | 1 |
| 2008 | Automatic chord recognition for music classification and retrievalabstractAs one of the most important mid-level features of music, chord contains rich information of harmonic structure that is useful for music information retrieval. In this paper, we present a chord recognition system based on the N-gram model. The system is time-efficient, and its accuracy is comparable to existing systems. We further propose a new method to construct chord features for music emotion classification and evaluate its performance on commercial song recordings. Experimental results demonstrate the advantage of using chord features for music classification and retrieval. Heng Tze Cheng, Yi-Hsuan Yang, Yu-Ching Lin, I-Bin Liao, Homer H. Chen |
ICME | 2 |
| 2008 | Interactive content presentation based on expressed emotion and physiological feedbackabstractIn this technical demonstration, we showcase an interactive content presentation (ICP) system that integrates media-expressed-emotion-based composition, user-perceived preference feedback, and interactive digital art creation. ICP harmonizes the browsing of multimedia contents by presenting them in the form of music videos (photos, blog articles with accompanied music) based on their expressed emotion similarity. ICP facilitates content browsing by automatically and dynamically selecting the media to be played next in real time, responding to user's preference feedback measured from physiological signals. In addition, ICP enhances the enjoyments of content browsing by incorporating interactive digital art creation. ICP achieves these goals by properly integrating recent researches on media-expressed emotion classification,cross-media composition, and physiological signal processing. Tien-Lin Wu, Hsuan-Kai Wang, Murphy Chien-Chang Ho, Yuan-Pin Lin, Ting-Ting Hu, Ming-Fang Weng, Li-Wei Chan 0001, Changhua Yang, Yi-Hsuan Yang, Yi-Ping Hung, Yung-Yu Chuang, Hsin-Hsi Chen, Homer H. Chen, Jyh-Horng Chen, Shyh-Kang Jeng |
ACM Multimedia | 9 |
| 2008 | Keyword-based concept search on consumer photos by web-based kernel functionabstractIn light of the strong demands for semantic search over large-scale consumer photos, which generally lack reliable user-provided annotations, we investigate the feasibility and challenges entailed by the new paradigm, concept search - retrieving visual objects by large-scale automatic concept detectors with keywords. We investigate the problem in three folds: (1) the effective concept mapping and selection methods over large-scale concept ontology; (2) the quality and feasibility of the pre-trained concept detectors applying on cross-domain consumer data (i.e., Flickr photos); (3) the search quality by fusing automatic concepts and user-annotated data (tags). Through experiments over large-scale benchmarks, TRECVID and Flickr550, we confirm the effectiveness of concept search in the proposed framework, where the semantic mapping by web-based kernel function over Google snippets significantly outperforms conventional WordNet-like methods both in accuracy and efficiency. Po Tun Wu, Yi-Hsuan Yang, Kuan-Ting Chen, Winston H. Hsu, Tien-Hsu Lee, Chun Jen Lee |
ACM Multimedia | 2 |
| 2008 | Mr. Emo: music retrieval in the emotion planeabstractThis technical demo presents a novel emotion-based music retrieval platform, called Mr. Emo, for organizing and browsing music collections. Unlike conventional approaches which quantize emotions into classes, Mr. Emo defines emotions by two continuous variables arousal and valence and employs regression algorithms to predict them. Associated with arousal and valence values (AV values), each music sample becomes a point in the arousal-valence emotion plane, so a user can easily retrieve music samples of certain emotion(s) by specifying a point or a trajectory in the emotion plane. Being content centric and functionally powerful, such emotion-based retrieval complements traditional keyword- or artist-based retrieval. The demo shows the effectiveness and novelty of music retrieval in the emotion plane. Yi-Hsuan Yang, Yu-Ching Lin, Heng Tze Cheng, Homer H. Chen |
ACM Multimedia | 1 |
| 2008 | ContextSeer: context search and recommendation at query time for shared consumer photosabstractThe advent of media-sharing sites like Flickr has drastically increased the volume of community-contributed multimedia resources on the web. However, due to their magnitudes, these collections are increasingly difficult to understand, search and navigate. To tackle these issues, a novel search system, ContextSeer, is developed to improve search quality (by reranking) and recommend supplementary information (i.e., search-related tags and canonical images) by leveraging the rich context cues, including the visual content, high-level concept scores, time and location metadata. First, we propose an ordinal reranking algorithm to enhance the semantic coherence of text-based search result by mining contextual patterns in an unsupervised fashion. A novel feature selection method, wc-tf-idf is also developed to select informative context cues. Second, to represent the diversity of search result, we propose an efficient algorithm cannoG to select multiple canonical images without clustering. Finally, ContextSeer enhances the search experience by further recommending relevant tags. Besides being effective and unsupervised, the proposed methods are efficient and can be finished at query time, which is vital for practical online applications. To evaluate ContextSeer, we have collected 0.5 million consumer photos from Flickr and manually annotated a number of queries by pooling to form a new benchmark, Flickr550. Ordinal reranking achieves significant performance gains both in Flcikr550 and TRECVID search benchmarks. Through a subjective test, cannoG expresses its representativeness and excellence for recommending multiple canonical images. Yi-Hsuan Yang, Po Tun Wu, Ching-Wei Lee, Kuan Hung Lin, Winston H. Hsu, Homer H. Chen |
ACM Multimedia | 1 |
| 2008 | A Regression Approach to Music Emotion RecognitionabstractContent-based retrieval has emerged in the face of content explosion as a promising approach to information access. In this paper, we focus on the challenging issue of recognizing the emotion content of music signals, or music emotion recognition (MER). Specifically, we formulate MER as a regression problem to predict the arousal and valence values (AV values) of each music sample directly. Associated with the AV values, each music sample becomes a point in the arousal-valence plane, so the users can efficiently retrieve the music sample by specifying a desired point in the emotion plane. Because no categorical taxonomy is used, the regression approach is free of the ambiguity inherent to conventional categorical approaches. To improve the performance, we apply principal component analysis to reduce the correlation between arousal and valence, and RReliefF to select important features. An extensive performance study is conducted to evaluate the accuracy of the regression approach for predicting AV values. The best performance evaluated in terms of theR2statistics reaches 58.3% for arousal and 28.1% for valence by employing support vector machine as the regressor. We also apply the regression approach to detect the emotion variation within a music selection and find the prediction accuracy superior to existing works. A group-wise MER scheme is also developed to address the subjectivity issue of emotion perception. Yi-Hsuan Yang, Yu-Ching Lin, Ya-Fan Su, Homer H. Chen |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | Music Emotion Classification: A Regression ApproachabstractTypical music emotion classification (MEC) approaches categorize emotions and apply pattern recognition methods to train a classifier. However, categorized emotions are too ambiguous for efficient music retrieval. In this paper, we model emotions as continuous variables composed of arousal and valence values (AV values), and formulate MEC as a regression problem. The multiple linear regression, support vector regression, and AdaBoost.RT are adopted to evaluate the prediction accuracy. Since the regression approach is inherently continuous, it is free of the ambiguity problem existing in its categorical counterparts. Yi-Hsuan Yang, Yu-Ching Lin, Ya-Fan Su, Homer H. Chen |
ICME | 1 |
| 2006 | Smooth Playout Control for Video Streaming over Error-Prone ChannelsabstractThe quality of media streaming over best-effort networks suffers from network delays and packet losses. The latter is more profound for wireless video. To enhance the QoS of streaming services, adaptive media playout (AMP) has been developed to adjust the playout interval. With AMP, the risk of delay and buffer underflow is reduced. However, the smoothness of playback is not guaranteed. In this paper, we propose a novel AMP control that enables smooth playout and meanwhile maintains reliable visual quality. Our AMP control adjusts the playout interval based on an estimation of channel quality, so it is more adaptive than conventional AMP controls that are based on buffer fullness. Experimental results are provided to justify our approach. Even at 20% packet loss rate, the proposed AMP control is still able to provide smooth and reliable playback Yi-Hsuan Yang, Meng-Ting Lu, Homer H. Chen |
ISM | 1 |
| 2006 | Music emotion classification: a fuzzy approachabstractDue to the subjective nature of human perception, classification of the emotion of music is a challenging problem. Simply assigning an emotion class to a song segment in a deterministic way does not work well because not all people share the same feeling for a song. In this paper, we consider a different approach to music emotion classification. For each music segment, the approach determines how likely the song segment belongs to an emotion class. Two fuzzy classifiers are adopted to provide the measurement of the emotion strength. The measurement is also found useful for tracking the variation of music emotions in a song. Results are shown to illustrate the effectiveness of the approach. Yi-Hsuan Yang, Chia Chu Liu, Homer H. Chen |
ACM Multimedia | 1 |