EDBT 2026 Demo / reviewers in the wild / expert
Ju-Chiang Wang
dblp:33/8541
· DBLP profile ↗
29ranked-venue papers
9as first author
10since 2021 · last 2026
0009-0002-8265-4229ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 8 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MuseDance: A Diffusion-based Music-Driven Image Animation SystemabstractImage animation is a rapidly developing area in multimodal research, with a focus on generating videos from reference images. While much of the work has emphasized generic video generation guided by text, music-driven dance image animation remains underexplored. In this paper, we introduce MuseDance, an end-to-end model that animates reference images using both music and text inputs. By integrating music as a conditioning modality, MuseDance generates personalized videos that not only adhere to textual descriptions but also synchronize character movements with the rhythm and dynamics of the music. Unlike existing methods, MuseDance eliminates the need for explicit motion guidance, such as pose sequences or depth maps, reducing the complexity of video generation while enhancing accessibility and flexibility. To support further research in this field, we present a new multimodal dataset comprising of 3,122 dance videos, each paired with the corresponding background music and text descriptions. Our approach leverages diffusion-based methods to achieve robust generalization, precise control, and temporal consistency, setting a new benchmark for the task of music-driven image animation. The dataset of this work is available at https://github.com/Dongzhikang/musedance. Zhikang Dong, Weituo Hao, Ju-Chiang Wang, Pawel Polak |
WACV | 3 |
| 2025 | MQAD: A Large-Scale Question Answering Dataset for Training Music Large Language ModelsabstractQuestion-answering (QA) is a natural approach for humans to understand a piece of music audio. However, for machines, accessing a large-scale dataset covering diverse aspects of music is crucial, yet challenging, due to the scarcity of publicly available music data of this type. This paper introduces MQAD, a music QA dataset built on the Million Song Dataset (MSD), encompassing a rich array of musical features - including beat, chord, key, structure, instrument, and genre — across 270,000 tracks, featuring nearly 3 million diverse questions and captions. MQAD distinguishes itself by offering detailed time-varying musical information such as chords and sections, enabling exploration into the inherent structure of music within a song. To compile MQAD, our methodology leverages specialized Music Information Retrieval (MIR) models to extract higher-level musical features and Large Language Models (LLMs) to generate natural language QA pairs. Then, we leverage a multimodal LLM that integrates the LLaMA2 and Whisper architectures, along with novel subjective metrics to assess the performance of MQAD. In experiments, our model trained on MQAD demonstrates advancements over conventional music audio captioning approaches. The dataset and codes are at https://github.com/oyzh888/MQAD. Zhihao Ouyang, Ju-Chiang Wang, Daiyu Zhang, Shangjie Li, Quan Lin |
ICASSP | 2 |
| 2024 | Music Source Separation With Band-Split Rope TransformerabstractMusic source separation (MSS) aims to separate a music recording into multiple musically distinct stems, such as vocals, bass, drums, and more. Recently, deep learning approaches such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs) have been used, but the improvement is still limited. In this paper, we propose a novel frequency-domain approach (called BS-RoFormer) based on a Band-Split RoPE Transformer architecture. BS-RoFormer relies on a band-split module to project the input complex spectrogram into subband-level representations, and then arranges a stack of hierarchical Transformers to model the inner-band as well as inter-band sequences for multi-band mask estimation. To facilitate training the model for MSS, we propose to use the Rotary Position Embedding (RoPE). The BS-RoFormer system trained on MUSDB18HQ and 500 extra songs ranked the first place in the Music Separation contest of Sound Demixing Challenge (SDX’23). Benchmarking a smaller version of BS-RoFormer on MUSDB18HQ, we achieve state-of-the-art result without extra training data, with 9.80 dB of average SDR. Wei Tsung Lu, Ju-Chiang Wang, Qiuqiang Kong, Yun-Ning Hung |
ICASSP | 2 |
| 2024 | STEMGEN: A Music Generation Model That ListensabstractEnd-to-end generation of musical audio using deep learning techniques has seen an explosion of activity recently. However, most models concentrate on generating fully mixed music in response to abstract conditioning information. In this work, we present an alternative paradigm for producing music generation models that can listen and respond to musical context. We describe how such a model can be constructed using a non-autoregressive, transformer-based model architecture and present a number of novel architectural and sampling improvements. We train the described architecture on both an open-source and a proprietary dataset. We evaluate the produced models using standard quality metrics and a new approach based on music information retrieval descriptors. The resulting model reaches the audio quality of state-of-the-art text-conditioned models, as well as exhibiting strong musical coherence with its context. Julian Parker, Janne Spijkervet, Katerina Kosta, Furkan Yesiler, Boris Kuznetsov, Ju-Chiang Wang, Matt Avent, Jitong Chen |
ICASSP | 6 |
| 2023 | SingNet: a real-time Singing Voice beat and Downbeat Tracking SystemabstractSinging voice beat and downbeat tracking posses several applications in automatic music production, analysis and manipulation. Among them, some require real-time processing, such as live performance processing and auto-accompaniment for singing inputs. This task is challenging owing to the non-trivial rhythmic and harmonic patterns in singing signals. For real-time processing, it introduces further constraints such as inaccessibility to future data and the impossibility to correct the previous results that are inconsistent with the latter ones. In this paper, we introduce the first system that tracks the beats and downbeats of singing voices in real-time. Specifically, we propose a novel dynamic particle filtering approach that incorporates offline historical data to correct the online inference by using a variable number of particles. We evaluate the performance on two datasets: GTZAN with the separated vocal tracks, and an in-house dataset with the original vocal stems. Experimental result demonstrates that our proposed approach outperforms the baseline by 3–5%. Mojtaba Heydari, Ju-Chiang Wang, Zhiyao Duan |
ICASSP | 2 |
| 2023 | Multitrack Music Transcription with a Time-Frequency PerceiverabstractMultitrack music transcription aims to transcribe a music audio input into the musical notes of multiple instruments simultaneously. It is a very challenging task that typically requires a more complex model to achieve satisfactory result. In addition, prior works mostly focus on transcriptions of regular instruments, however, neglecting vocals, which are usually the most important signal source if present in a piece of music. In this paper, we propose a novel deep neural network architecture, Perceiver TF, to model the time-frequency representation of audio input for multitrack transcription. Perceiver TF augments the Perceiver architecture by introducing a hierarchical expansion with an additional Transformer layer to model temporal coherence. Accordingly, our model inherits the benefits of Perceiver that posses better scalability, allowing it to well handle transcriptions of many instruments in a single model. In experiments, we train a Perceiver TF to model 12 instrument classes as well as vocal in a multi-task learning manner. Our result demonstrates that the proposed system outperforms the state-of-the-art counterparts (e.g., MT3 and SpecTNT) on various public datasets. Wei Tsung Lu, Ju-Chiang Wang, Yun-Ning Hung |
ICASSP | 2 |
| 2022 | Modeling Beats and Downbeats with a Time-Frequency TransformerabstractTransformer is a successful deep neural network (DNN) architecture that has shown its versatility not only in natural language processing but also in music information retrieval (MIR). In this paper, we present a novel Transformer-based approach to tackle beat and downbeat tracking. This approach employs SpecTNT (Spectral- Temporal Transformer in Transformer), a variant of Transformer that models both spectral and temporal dimensions of a time-frequency input of music audio. A SpecTNT model uses a stack of blocks, where each consists of two levels of Transformer encoders. The lower-level (or spectral) encoder handles the spectral features and enables the model to pay attention to harmonic components of each frame. Since downbeats indicate bar boundaries and are often accompanied by harmonic changes, this step may help downbeat modeling. The upper-level (or temporal) encoder aggregates useful local spectral information to pay attention to beat/downbeat positions. We also propose an architecture that combines SpecTNT with a state-of- the-art model, Temporal Convolutional Networks (TCN), to further improve the performance. Extensive experiments demonstrate that our approach can significantly outperform TCN in downbeat tracking while maintaining comparable result in beat tracking. Yun-Ning Hung, Ju-Chiang Wang, Xuchen Song, Wei Tsung Lu, Minz Won |
ICASSP | 2 |
| 2022 | To Catch A Chorus, Verse, Intro, or Anything Else: Analyzing a Song with Structural FunctionsabstractConventional music structure analysis algorithms aim to divide a song into segments and to group them with abstract labels (e.g., ‘A’, ‘B’, and ‘C’). However, explicitly identifying the function of each segment (e.g., ‘verse’ or ‘chorus’) is rarely attempted, but has many applications. We introduce a multi-task deep learning framework to model these structural semantic labels directly from audio by estimating "verseness," "chorusness," and so forth, as a function of time. We propose a 7-class taxonomy (i.e., intro, verse, chorus, bridge, outro, instrumental, and silence) and provide rules to consolidate annotations from four disparate datasets. We also propose to use a spectral-temporal Transformer-based model, called SpecTNT, which can be trained with an additional connectionist temporal localization (CTL) loss. In cross-dataset evaluations using four public datasets, we demonstrate the effectiveness of the SpecTNT model and CTL loss, and obtain strong results overall: the proposed system outperforms state-of-the-art chorus-detection and boundary-detection methods at detecting choruses and boundaries, respectively. Ju-Chiang Wang, Yun-Ning Hung, Jordan B. L. Smith |
ICASSP | 1 |
| 2021 | Modeling the Compatibility of Stem Tracks to Generate Music MashupsabstractA music mashup combines audio elements from two or more songs to create a new work. To reduce the time and effort required to make them, researchers have developed algorithms that predict the compatibility of audio elements. Prior work has focused on mixing unaltered excerpts, but advances in source separation enable the creation of mashups from isolated stems (e.g., vocals, drums, bass, etc.). In this work, we take advantage of separated stems not just for creating mashups, but for training a model that predicts the mutual compatibility of groups of excerpts, using self-supervised and semi-supervised methods. Specifically, we first produce a random mashup creation pipeline that combines stem tracks obtained via source separation, with key and tempo automatically adjusted to match, since these are prerequisites for high-quality mashups. To train a model to predict compatibility, we use stem tracks obtained from the same song as positive examples, and random combinations of stems with key and/or tempo unadjusted as negative examples. To improve the model and use more data, we also train on "average" examples: random combinations with matching key and tempo, where we treat them as unlabeled data as their true compatibility is unknown. To determine whether the combined signal or the set of stem signals is more indicative of the quality of the result, we experiment on two model architectures and train them using semi-supervised learning technique. Finally, we conduct objective and subjective evaluations of the system, comparing them to a standard rule-based system. Ju-Chiang Wang, Jordan B. L. Smith, Xuchen Song, Yuxuan Wang 0002 |
AAAI | 2 |
| 2021 | Supervised Chorus Detection for Popular Music Using Convolutional Neural Network and Multi-Task LearningabstractThis paper presents a novel supervised approach to detecting the chorus segments in popular music. Traditional approaches to this task are mostly unsupervised, with pipelines designed to target some quality that is assumed to define "chorusness," which usually means seeking the loudest or most frequently repeated sections. We propose to use a convolutional neural network with a multi-task learning objective, which simultaneously fits two temporal activation curves: one indicating "chorusness" as a function of time, and the other the location of the boundaries. We also propose a post-processing method that jointly takes into account the chorus and boundary predictions to produce binary output. In experiments using three datasets, we compare our system to a set of public implementations of other segmentation and chorus-detection algorithms, and find our approach performs significantly better. Ju-Chiang Wang, Jordan B. L. Smith, Jitong Chen, Xuchen Song, Yuxuan Wang 0002 |
ICASSP | 1 |
| 2018 | Predicting the Probability Density Function of Music Emotion Using Emotion Space MappingabstractComputationally modeling the affective content of music has been intensively studied in recent years because of its wide applications in music retrieval and recommendation. Although significant progress has been made, this task remains challenging due to the difficulty in properly characterizing the emotion of a music piece. Music emotion perceived by people is subjective by nature and thus complicates the process of collecting the emotion annotations as well as developing the predictive model. Instead of assuming people can reach a consensus on the emotion of music, in this work we propose a novel machine learning approach that characterizes the music emotion as a probability distribution in the valence-arousal (VA) emotion space, not only tackling the subjectivity but also precisely describing the emotions of a music piece. Specifically, we represent the emotion of a music piece as a probability density function (PDF) in the VA space via kernel density estimation from human annotations. To associate emotion with the audio features extracted from music pieces, we learn the combination coefficients by optimizing some objective functions of audio features, and then predict the emotion of an unseen piece by linearly combining the PDFs of the training pieces with the coefficients. Several algorithms for learning the coefficients are studied. Evaluations on the NTUMIR and MediaEval 2013 datasets validate the effectiveness of the proposed methods in predicting the probability distributions of emotion from audio features. We also demonstrate how to use the proposed approach in emotion-based music retrieval. Yu-Hao Chin, Jia-Ching Wang, Ju-Chiang Wang, Yi-Hsuan Yang |
IEEE Trans. Affect. Comput. | 3 |
| 2017 | Component Tying for Mixture Model Adaptation in Personalization of Music Emotion RecognitionabstractPersonalizing a music emotion recognition model is needed because the perception of music emotion is highly subjective, but it is a time-consuming process. In this paper, we consider how to expedite the personalization process that begins with a general model trained offline using a general user base and progressively adapts the model to a music listener using the emotion annotations of the listener. Specifically, we focus on reducing the number of user annotations needed for the personalization. We investigate and evaluate four component tying methods: single group tying, quadrantwise tying, hierarchical tying, and random tying. These methods aim to exploit the available annotations by identifying related model parameters on-the-fly and updating them jointly. In the evaluation, we use the AMG1608 dataset, which contains the clip-level valence-arousal emotion ratings of 1608 30-s music clips annotated by 665 listeners. Also, we use the acoustic emotion Gaussians model as the general model that uses a mixture of Gaussian components to learn the mapping between the acoustic feature space and the emotion space. The results show that the model adaptation with component tying requires only 10-20 personal annotations to obtain the same level of prediction accuracy as the baseline model adaptation method that uses 50 personal annotations without component tying. Yu-An Chen, Ju-Chiang Wang, Yi-Hsuan Yang, Homer H. Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | The AMG1608 dataset for music emotion recognitionabstractAutomated recognition of musical emotion from audio signals has received considerable attention recently. To construct an accurate model for music emotion prediction, the emotion-annotated music corpus has to be of high quality. It is desirable to have a large number of songs annotated by numerous subjects to characterize the general emotional response to a song. Due to the need for personalization of the music emotion prediction model to address the subjective nature of emotion perception, it is also important to have a large number of annotations per subject for training and evaluating a personalization method. In this paper, we discuss the deficiency of existing datasets and present a new one. The new dataset, which is publically available to the research community, is composed of 1608 30-second music clips annotated by 665 subjects. Furthermore, 46 subjects annotated more than 150 songs, making this dataset the largest of its kind to date. Yu-An Chen, Yi-Hsuan Yang, Ju-Chiang Wang, Homer H. Chen |
ICASSP | 3 |
| 2015 | A histogram density modeling approach to music emotion recognitionabstractMusic emotion recognition is concerned with developing predictive models that comprehend the affective content of musical signals. Recently, a growing number of attempts has been made to model the music emotion as a probability distribution in the valence-arousal (VA) space to better account for the subjectivity. In this paper, we present a novel histogram density modeling approach that models the emotion distribution by a 2-D histogram over the quantized VA space and learns a set of latent histograms to predict the emotion probability density of a song from audio. The proposed model is free from parametric distribution assumptions over the VA space, easy to implement, and extremely fast to train. We also extend our model to deal with the temporal dynamics of time-varying emotion labels. Comprehensive performance study on two larger-scale datasets demonstrates that our approach achieves comparable performance to the state-of-the-art ones, but with much better training and testing efficiency. Ju-Chiang Wang, Hsin-Min Wang, Gert R. G. Lanckriet |
ICASSP | 1 |
| 2015 | Modeling the Affective Content of Music with a Gaussian Mixture ModelabstractModeling the association between music and emotion has been considered important for music information retrieval and affective human computer interaction. This paper presents a novel generative model called acoustic emotion Gaussians (AEG) for computational modeling of emotion. Instead of assigning a music excerpt with a deterministic (hard) emotion label, AEG treats the affective content of music as a (soft) probability distribution in the valence-arousal space and parameterizes it with a Gaussian mixture model (GMM). In this way, the subjective nature of emotion perception is explicitly modeled. Specifically, AEG employs two GMMs to characterize the audio and emotion data. The fitting algorithm of the GMM parameters makes the model learning process transparent and interpretable. Based on AEG, a probabilistic graphical structure for predicting the emotion distribution from music audio data is also developed. A comprehensive performance study over two emotion-labeled datasets demonstrates that AEG offers new insights into the relationship between music and emotion (e.g., to assess the “affective diversity” of a corpus) and represents an effective means of emotion modeling. Readers can easily implement AEG via the publicly available codes. As the AEG model is generic, it holds the promise of analyzing any signal that carries affective or other highly subjective information. Ju-Chiang Wang, Yi-Hsuan Yang, Hsin-Min Wang, Shyh-Kang Jeng |
IEEE Trans. Affect. Comput. | 1 |
| 2014 | Linear regression-based adaptation of music emotion recognition models for personalizationabstractPersonalization techniques can be applied to address the subjectivity issue of music emotion recognition, which is important for music information retrieval. However, achieving satisfactory accuracy in personalized music emotion recognition for a user is difficult because it requires an impractically huge amount of annotations from the user. In this paper, we adopt a probabilistic framework for valence-arousal music emotion modeling and propose an adaptation method based on linear regression to personalize a background model in an online learning fashion. We also incorporate a component-tying strategy to enhance the model flexibility. Comprehensive experiments are conducted to test the performance of the proposed method on three datasets, including a new one created specifically in this work for personalized music emotion recognition. Our results demonstrate the effectiveness of the proposed method. Yu-An Chen, Ju-Chiang Wang, Yi-Hsuan Yang, Homer H. Chen |
ICASSP | 2 |
| 2014 | Improving music auto-tagging by intra-song instance baggingabstractBagging is one the most classic ensemble learning techniques in the machine learning literature. The idea is to generate multiple subsets of the training data via bootstrapping (random sampling with replacement), and then aggregate the output of the models trained from each subset via voting or averaging. As music is a temporal signal, we propose and study two bagging methods in this paper: the inter-song instance bagging that bootstraps song-level features, and the intra-song instance bagging that draws bootstrapping samples directly from short-time features for each training song. In particular, we focus on the latter method, as it better exploits the temporal information of music signals. The bagging methods result in surprisingly effective models for music auto-tagging: incorporating the idea to a simple linear support vector machine (SVM) based system yields accuracies that are comparable or even superior to state-of-the-art, possibly more sophisticated methods for three different datasets. As the bagging method is a meta algorithm, it holds the promise of improving other MIR systems. Chin-Chia Michael Yeh, Ju-Chiang Wang, Yi-Hsuan Yang, Hsin-Min Wang |
ICASSP | 2 |
| 2014 | Towards time-varying music auto-tagging based on CAL500 expansionabstractMusic auto-tagging refers to automatically assigning semantic labels (tags) such as genre, mood and instrument to music so as to facilitate text-based music retrieval. Although significant progress has been made in recent years, relatively little research has focused on semantic labels that are time-varying within a track. Existing approaches and datasets usually assume that different fragments of a track share the same tag labels, disregarding the tags that are time-varying (e.g., mood) or local in time (e.g., instrument solo). In this paper, we present a new dataset dedicated to time-varying music auto-tagging. The dataset, called CAL500exp, is an enriched version of the well-known CAL500 dataset used for conventional track-level tagging. Given the tag set of CAL500, eleven subjects with strong music background were recruited to annotate the time-varying tag labels. A new user interface for annotation is developed to reduce the subject's annotation effort yet increase the quality of labels. Moreover, we present an empirical evaluation that demonstrates the performance improvement CAL500exp brings about for time-varying music auto-tagging. By providing more accurate and consistent descriptions of music content in a finer granularity, CAL500exp may open new opportunities to understand and to model the temporal context of musical semantics. Shuo-Yang Wang, Ju-Chiang Wang, Yi-Hsuan Yang, Hsin-Min Wang |
ICME | 2 |
| 2014 | A Systematic Evaluation of the Bag-of-Frames Representation for Music Information RetrievalabstractThere has been an increasing attention on learning feature representations from the complex, high-dimensional audio data applied in various music information retrieval (MIR) problems. Unsupervised feature learning techniques, such as sparse coding and deep belief networks have been utilized to represent music information as a term-document structure comprising of elementary audio codewords. Despite the widespread use of such bag-of-frames (BoF) model, few attempts have been made to systematically compare different component settings. Moreover, whether techniques developed in the text retrieval community are applicable to audio codewords is poorly understood. To further our understanding of the BoF model, we present in this paper a comprehensive evaluation that compares a large number of BoF variants on three different MIR tasks, by considering different ways of low-level feature representation, codebook construction, codeword assignment, segment-level and song-level feature pooling, tf-idf term weighting, power normalization, and dimension reduction. Our evaluations lead to the following findings: 1) modeling music information by two levels of abstraction improves the result for difficult tasks such as predominant instrument recognition, 2) tf-idf weighting and power normalization improve system performance in general, 3) topic modeling methods such as latent Dirichlet allocation does not work for audio codewords. Li Su 0004, Chin-Chia Michael Yeh, Jen-Yu Liu, Ju-Chiang Wang, Yi-Hsuan Yang |
IEEE Trans. Multim. | 4 |
| 2013 | Non-reference audio quality assessment for online live music recordingsabstractImmensely popular video sharing websites such as YouTube have become the most important sources of music information for Internet users and the most prominent platform for sharing live music. The audio quality of this huge amount of live music recordings, however, varies significantly due to factors such as environmental noise, location, and recording device. However, most video search engines do not take audio quality into consideration when retrieving and ranking results. Given the fact that most users prefer live music videos with better audio quality, we propose the first automatic, non-reference audio quality assessment framework for live music video search online. We first construct two annotated datasets of live music recordings. The first dataset contains 500 human-annotated pieces, and the second contains 2,400 synthetic pieces systematically generated by adding noise effects to clean recordings. Then, we formulate the assessment task as a ranking problem and try to solve it using a learning-based scheme. To validate the effectiveness of our framework, we perform both objective and subjective evaluations. Results show that our framework significantly improves the ranking performance of live music recording retrieval and can prove useful for various real-world music applications. Ju-Chiang Wang, Jingli Cai, Zhiyan Duan, Hsin-Min Wang, Ye Wang 0007 |
ACM Multimedia | 2 |
| 2012 | Playing with tagging: A real-time tagging music playerabstractVisualizing audio signals during playback has long been a fundamental function of music players. However, most visual effects are generated by audio signal processing directly and render meaningless or incomprehensible displays to users. In this paper, we present an intelligent music player called the Playing with Tagging (PWT) music player. By integrating a real-time music tagger, the PWT player can display dynamic tag distributions via a set of tag bars that move in sync with the music. To synchronize the tag distributions, the music tagger must be able to online recognize the music tags. We utilize a Gaussian mixture model (GMM) as an auditory feature encoding reference and a mixture of tag-based aspect models (TBAMs) to predict the tag distribution for a short sliding chunk of the music played. To evaluate the real-time tagging function, we simulate tag prediction on short music chunks. The results of experiments on the MajorMiner dataset demonstrate the potential and effectiveness of the proposed music tagging method. Ju-Chiang Wang, Hsin-Min Wang, Shyh-Kang Jeng |
ICASSP | 1 |
| 2012 | The acousticvisual emotion guassians model for automatic generation of music videoabstractThis paper presents a novel content-based system that utilizes the perceived emotion of multimedia content as a bridge to connect music and video. Specifically, we propose a novel machine learning framework, called Acousticvisual Emotion Guassians (AVEG), to jointly learn the tripartite relationship among music, video, and emotion from an emotion-annotated corpus of music videos. For a music piece (or a video sequence), the AVEG model is applied to predict its emotion distribution in a stochastic emotion space from the corresponding low-level acoustic (resp. visual) features. Finally, music and video are matched by measuring the similarity between the two corresponding emotion distributions, based on a distance measure such as KL divergence. Ju-Chiang Wang, Yi-Hsuan Yang, I-Hong Jhuo, Yen-Yu Lin, Hsin-Min Wang |
ACM Multimedia | 1 |
| 2012 | The acoustic emotion gaussians model for emotion-based music annotation and retrievalabstractOne of the most exciting but challenging endeavors in music research is to develop a computational model that comprehends the affective content of music signals and organizes a music collection according to emotion. In this paper, we propose a novel acoustic emotion Gaussians (AEG) model that defines a proper generative process of emotion perception in music. As a generative model, AEG permits easy and straightforward interpretations of the model learning processes. To bridge the acoustic feature space and music emotion space, a set of latent feature classes, which are learned from data, is introduced to perform the end-to-end semantic mappings between the two spaces. Based on the space of latent feature classes, the AEG model is applicable to both automatic music emotion annotation and emotion-based music retrieval. To gain insights into the AEG model, we also provide illustrations of the model learning process. A comprehensive performance study is conducted to demonstrate the superior accuracy of AEG over its predecessors, using two emotion annotated music corpora MER60 and MTurk. Our results show that the AEG model outperforms the state-of-the-art methods in automatic music emotion annotation. Moreover, for the first time a quantitative evaluation of emotion-based music retrieval is reported. Ju-Chiang Wang, Yi-Hsuan Yang, Hsin-Min Wang, Shyh-Kang Jeng |
ACM Multimedia | 1 |
| 2011 | Cost-sensitive stacking for audio tag annotation and retrievalabstractAudio tags correspond to keywords that people use to de scribe different aspects of a music clip, such as the genre, mood, and instrumentation. Since social tags are usually as signed by people with different levels of musical knowledge, they inevitably contain noisy information. By treating the tag counts as costs, we can model the audio tagging problem as a cost-sensitive classification problem. In addition, tag correlation is another useful information for automatic audio tagging since some tags often co-occur. By considering the co-occurrences of tags, we can model the audio tagging problem as a multi-label classification problem. To exploit the tag count and correlation information jointly, we formulate the audio tagging task as a novel cost-sensitive multi-label (CSML) learning problem. The results of audio tag annotation and retrieval experiments demonstrate that the new approach outperforms our MIREX 2009 winning method. Hung-Yi Lo, Ju-Chiang Wang, Hsin-Min Wang, Shou-De Lin |
ICASSP | 2 |
| 2011 | Query by multi-tags with multi-level preferences for content-based music retrievalabstractThis paper presents a novel content-based music retrieval system that accepts a query containing multiple tags with multiple levels of preference (denoted as an MTML query) to retrieve music from an untagged music database. We select a limited number of popular music tags to form the tag space and design an interface for users to input queries by operating the scroll bars. To effect MTML content-based music retrieval, we introduce a tag-based music aspect model that jointly models the auditory features and tag-based text features of a song. Two indexing methods and their corresponding matching methods, namely pseudo song-based matching and tag co-occurrence pattern-based matching, are incorporated into the pre-learned tag-based music aspect model. Finally, we evaluate the proposed system on the Major Miner dataset. The results demonstrate the potential of using MTML queries to retrieve music from an untagged music database. Ju-Chiang Wang, Meng-Sung Wu, Hsin-Min Wang, Shyh-Kang Jeng |
ICME | 1 |
| 2011 | Colorizing tags in tag cloud: a novel query-by-tag music search systemabstractThis paper presents a novel content-based query-by-tag music search system for an untagged music database. We design a new tag query interface that allows users to input multiple tags with multiple levels of preference (denoted as an MTML query) by colorizing desired tags in a web-based tag cloud interface. When a user clicks and holds the left mouse button (or presses and holds his/her finger on a touch screen) on a desired tag, the color of the tag will change cyclically according to a color map (from dark blue to bright red), which represents the level of preference (from 0 to 1). In this way, the user can easily organize and check the query of multiple tags with multiple levels of preference through the colored tags. To effect the MTML content-based music retrieval, we introduce a probabilistic fusion model (denoted as GMFM), which consists of two mixture models, namely a Gaussian mixture model and a multinomial mixture model. GMFM can jointly model the auditory features and tag labels of a song. Two indexing methods and their corresponding matching methods, namely pseudo song-based matching and tag affinity-based matching, are incorporated into the pre-learned GMFM. We evaluate the proposed system on the MajorMiner and CAL-500 datasets. The experimental results demonstrate the effectiveness of GMFM and the potential of using MTML queries to search music from an untagged music database. Ju-Chiang Wang, Yu-Chin Shih, Meng-Sung Wu, Hsin-Min Wang, Shyh-Kang Jeng |
ACM Multimedia | 1 |
| 2011 | Cost-Sensitive Multi-Label Learning for Audio Tag Annotation and RetrievalabstractAudio tags correspond to keywords that people use to describe different aspects of a music clip. With the explosive growth of digital music available on the Web, automatic audio tagging, which can be used to annotate unknown music or retrieve desirable music, is becoming increasingly important. This can be achieved by training a binary classifier for each tag based on the labeled music data. Our method that won the MIREX 2009 audio tagging competition is one of this kind of methods. However, since social tags are usually assigned by people with different levels of musical knowledge, they inevitably contain noisy information. By treating the tag counts as costs, we can model the audio tagging problem as a cost-sensitive classification problem. In addition, tag correlation information is useful for automatic audio tagging since some tags often co-occur. By considering the co-occurrences of tags, we can model the audio tagging problem as a multi-label classification problem. To exploit the tag count and correlation information jointly, we formulate the audio tagging task as a novel cost-sensitive multi-label (CSML) learning problem and propose two solutions to solve it. The experimental results demonstrate that the new approach outperforms our MIREX 2009 winning method. Hung-Yi Lo, Ju-Chiang Wang, Hsin-Min Wang, Shou-De Lin |
IEEE Trans. Multim. | 2 |
| 2010 | Background music identification through content filtering and min-hash matchingabstractA novel framework for background music identification is proposed in this paper. Given a piece of audio signals that mixes background music with speech/noise, we identify the music part with source music data. Conventional methods that take the whole audio signals for identification are inappropriate in terms of efficiency and accuracy. In our framework, the audio content is filtered through speech center cancellation and noise removal to extract clear music segments. To identify these music segments, we use a compact feature representation and efficient similarity measurement based on the min-hash theory. The results of experiments on the RWC music database show a promising direction. Chih-Yi Chiu, Dimitrios Bountouridis, Ju-Chiang Wang, Hsin-Min Wang |
ICASSP | 3 |
| 2010 | Homogeneous segmentation and classifier ensemble for audio tag annotation and retrievalabstractAudio tags describe different types of musical information such as genre, mood, and instrument. This paper aims to automatically annotate audio clips with tags and retrieve relevant clips from a music database by tags. Given an audio clip, we divide it into several homogeneous segments by using an audio novelty curve, and then extract audio features from each segment with respect to various musical information, such as dynamics, rhythm, timbre, pitch, and tonality. The features in frame-based feature vector sequence format are further represented by their mean and standard deviation such that they can be combined with other segment-based features to form a fixed-dimensional feature vector for a segment. We train an ensemble classifier, which consists of SVM and AdaBoost classifiers, for each tag. For the audio annotation task, the individual classifier outputs are transformed into calibrated probability scores such that probability ensemble can be employed. For the audio retrieval task, we propose using ranking ensemble. We participated in the MIREX 2009 audio tag classification task and our system was ranked first in terms of F-measure and the area under the ROC curve given a tag. Hung-Yi Lo, Ju-Chiang Wang, Hsin-Min Wang |
ICME | 2 |