Xavier Serra

dblp:01/2680 · DBLP profile ↗
← Back
66ranked-venue papers
1as first author
24since 2021 · last 2026
0000-0003-1395-2345ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 40 · 1 first-author · 16 since 2021Artificial intelligence and machine learning · 16 · 5 since 2021Databases, data management, data science and information retrieval · 9 · 1 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Difficulty-aware score generation for piano sight-reading
abstract
• Generative model for piano sight-reading exercises controlled by difficulty level. • Scores are generated in MusicXML, ensuring readability and playability. • An auxiliary loss improves difficulty control and prevents conditioning collapse. • Expert pianists rated the pieces as clear, natural, and suitable for teaching. • The approach advances towards personalized music learning. Sight-reading is a core skill in music education. It refers to the ability to play a written piece of music correctly the first time it is seen. Developing this skill requires frequent practice with completely new musical excerpts that match the difficulty level of the student. However, creating new sight-reading exercises at a specific difficulty level requires significant time and expert knowledge. As a result, students and teachers often rely on pieces from the existing piano literature, even though sight-reading exams typically use compositions written specifically for the exam. Generative music systems provide a promising approach for creating new sight-reading material with explicit control over performance difficulty. In this work, we frame the creation of sight-reading exercises as a symbolic music generation task that produces piano scores with controllable difficulty. Existing approaches typically rely on control tokens to guide generation, but we show that this strategy does not result in piano scores with reliably controlled difficulty. To address this issue, we introduce an auxiliary difficulty prediction objective using synthetic difficulty labels produced by an expert-based system, enabling scalable training. Our method improves difficulty conditioning accuracy from 69.3% to 92.9% compared to a baseline that conditions generation solely on difficulty control tokens, and reduces mean squared error from 0.30 to 0.09. A user study with expert pianists shows that the generated scores are rated comparable or superior to human-written material in readability and naturalness, while maintaining appropriate playability across difficulty levels. These results represent a step toward the generation of music exercises for a variety of educational applications.
Pedro Ramoneda, Masahiro Suzuki 0002, Akira Maezawa, Xavier Serra
Expert Syst. Appl.4
2025 Efficient and Fast Generative-Based Singing Voice Separation using a Latent Diffusion Model
abstract
Extracting individual elements from music mixtures is a valuable tool for music production and practice. While neural networks optimized to mask or transform mixture spectrograms into the individual source(s) have been the leading approach, the source overlap and correlation in music signals poses an inherent challenge. Also, accessing all sources in the mixture is crucial to train these systems, while complicated. Attempts to address these challenges in a generative fashion exist, however, the separation performance and inference efficiency remain limited. In this work, we study the potential of diffusion models to advance toward bridging this gap, focusing on generative singing voice separation relying only on corresponding pairs of isolated vocals and mixtures for training. To align with creative workflows, we leverage latent diffusion: the system generates samples encoded in a compact latent space, and subsequently decodes these into audio. This enables efficient optimization and faster inference. Our system is trained using only open data. We outperform existing generative separation systems, and level the compared non-generative systems on a list of signal quality measures and on interference removal. We provide a noise robustness study on the latent encoder, providing insights on its potential for the task. We release a modular toolkit for further research on the topic.1
Genís Plaja-Roglans, Yun-Ning Hung, Xavier Serra, Igor Pereira
IJCNN3
2025 Improving Singing Voice Transcription Generalization with AI Generated Accompaniments
Miguel Pérez-Francisco, Holger Kirchhoff, Peter Grosche, Xavier Serra
MMM (2)4
2025 Enhanced television broadcast monitoring with source separation-assisted audio fingerprinting: A case study
abstract
Abstract Music identification is crucial for distributing royalties in the music industry. This problem is solved using Audio fingerprinting (AFP) algorithms. However, these methods often struggle in real-world scenarios such as TV broadcasting, when music is in the background, masked by other sounds such as speech. While prior research has focused on improving AFP robustness to pitch and tempo variations, less attention has been given to enhancing robustness for background music identification. In this work, we assess whether source separation systems improve background music identification by recovering the music signal in these recordings. We present the first extensive study comprising 13 source separation algorithms and five AFP models. We evaluate them on a public dataset of TV recordings, assessing both music identification performance and computational cost. Our results show that source separation substantially improves peak-based AFP identifications, particularly when music is in the background. Additionally, this finding extends to foreground music, making the approach versatile for various music identification tasks, such as query-by-example. Deep learning-based model NeuralFP* (tailored for background music identification) shows no substantial benefit from adding a separation model as preprocessing. This reproducible study provides a comprehensive evaluation framework, offering valuable insights into using source separation methods to improve music identification in real-world contexts.
Guillem Cortès-Sebastià, Marius Miron, Emilio Molina, Alex Ciurana, Xavier Serra
Multim. Tools Appl.5
2025 Velocity2DMs: A Contextual Modeling Approach to Dynamics Marking Prediction in Piano Performance
abstract
Expressive dynamics in music performance are subjective and context-dependent, yet most symbolic models treat Dynamics Markings (DMs) as static with fixed MIDI velocities. This paper proposes a method for predicting DMs in piano performance by combining MusicXML score information with performance MIDI data through a novel tokenization scheme and an adapted RoBERTa-based Masked Language Model (MLM). Our approach focuses on contextual aggregated MIDI velocities and corresponding DMs, accounting for subjective interpretations of pianists. Note-level features are serialized and translated into a sequence of tokens to predict both constant (e.g.,mp,ff) and non-constant DMs (e.g.,crescendo,fp). Evaluation across three expert performance datasets shows that the model effectively learns dynamics transitions from contextual note blocks and generalizes beyond constant markings. This is the first study to model both constant and non-constant dynamics in a unified framework using contextual sequence learning. The results suggest promising applications for expressive music analysis, performance modeling, and computer-assisted music education.
Hyon Kim, Emmanouil Benetos, Xavier Serra
IEEE Signal Process. Lett.3
2024 WikiMuTe: A Web-Sourced Dataset of Semantic Descriptions for Music Audio
Benno Weck, Holger Kirchhoff, Peter Grosche, Xavier Serra
MMM (5)4
2024 Combining piano performance dimensions for score difficulty classification
abstract
Predicting the difficulty of playing a musical score is essential for structuring and exploring score collections. Despite its importance for music education, the automatic difficulty classification of piano scores is not yet solved, mainly due to the lack of annotated data and the subjectiveness of the annotations. This paper aims to advance the state-of-the-art in score difficulty classification with two major contributions. To address the lack of data, we present Can I Play It? (CIPI) dataset, a machine-readable piano score dataset with difficulty annotations obtained from the renowned classical music publisher Henle Verlag. The dataset is created by matching public domain scores with difficulty labels from Henle Verlag, then reviewed and corrected by an expert pianist. As a second contribution, we explore various input representations from score information to pre-trained ML models for piano fingering and expressiveness inspired by the musicology definition of performance. We show that combining the outputs of multiple classifiers performs better than the classifiers on their own, pointing to the fact that the representations capture different aspects of difficulty. In addition, we conduct numerous experiments that lay a foundation for score difficulty classification and create a basis for future research. Our best-performing model reports a 39.5% balanced accuracy and 1.1 median square error across the nine difficulty levels proposed in this study. Code, dataset, and models are made available for reproducibility.
Pedro Ramoneda, Dasaem Jeong, Vsevolod Eremenko, Nazif Can Tamer, Marius Miron, Xavier Serra
Expert Syst. Appl.6
2023 Pre-Training Strategies Using Contrastive Learning and Playlist Information for Music Classification and Similarity
abstract
In this work, we investigate an approach that relies on contrastive learning and music metadata as a weak source of supervision to train music representation models. Recent studies show that contrastive learning can be used with editorial metadata (e.g., artist or album name) to learn audio representations that are useful for different classification tasks. In this paper, we extend this idea to using playlist data as a source of music similarity information and investigate three approaches to generate anchor and positive track pairs. We evaluate these approaches by fine-tuning the pre-trained models for music multi-label classification tasks (genre, mood, and instrument tagging) and music similarity. We find that creating anchor and positive track pairs by relying on co-occurrences in playlists provides better music similarity and competitive classification results compared to choosing tracks from the same artist as in previous works. Additionally, our best pre-training approach based on playlists provides superior classification performance for most datasets.
Pablo Alonso-Jiménez, Xavier Favory, Hadrien Foroughmand, Grigoris Bourdalas, Xavier Serra, Thomas Lidy, Dmitry Bogdanov
ICASSP5
2023 Flowgrad: Using Motion for Visual Sound Source Localization
abstract
Most recent work in visual sound source localization relies on semantic audio-visual representations learned in a self-supervised manner and, by design, excludes temporal information present in videos. While it proves to be effective for widely used benchmark datasets, the method falls short for challenging scenarios like urban traffic. This work introduces temporal context into the state-of-the-art methods for sound source localization in urban scenes using optical flow to encode motion information. An analysis of the strengths and weaknesses of our methods helps us better understand the problem of visual sound source localization and sheds light on open challenges for audio-visual scene understanding. The code and pretrained models are publicly available at https://github.com/rrrajjjj/flowgrad
Rajsuryan Singh, Pablo Zinemanas, Xavier Serra, Juan Pablo Bello, Magdalena Fuentes
ICASSP3
2023 TAPE: An End-to-End Timbre-Aware Pitch Estimator
abstract
Pitch estimation of a target musical source within a multi-source polyphonic signal is of great interest for music performance analysis. One possible approach for extracting the pitch of a target source is to first perform source separation and then estimate the pitch of the separated track. However, as we will show, this typically leads to poor results. As an alternative to this approach, we introduce a timbre-aware pitch estimator (TAPE), which estimates the pitch of a target source in an end-to-end manner without the need for an explicit source separation step. Opposed to existing approaches that assume the predominance of a lead voice, our approach builds upon other cues that only rely on the timbral characteristics. Our results on real violin–piano duets show that, without any pre-processing step, TAPE trained on synthetic mixes outperforms the sequential procedure of source separation and pitch estimation under many settings, even if the target source is not predominant.
Nazif Can Tamer, Yigitcan Özer, Meinard Müller, Xavier Serra
ICASSP4
2023 Data Leakage in Cross-Modal Retrieval Training: A Case Study
abstract
The recent progress in text-based audio retrieval was largely propelled by the release of suitable datasets. Since the manual creation of such datasets is a laborious task, obtaining data from online resources can be a cheap solution to create large-scale datasets. We study the recently proposed SoundDesc benchmark dataset, which was automatically sourced from the BBC Sound Effects web page. In our analysis, we find that SoundDesc contains several duplicates that cause leakage of training data to the evaluation data. This data leakage ultimately leads to overly optimistic retrieval performance estimates in previous benchmarks. We propose new training, validation, and testing splits for the dataset that we make available online. To avoid weak contamination of the test data, we pool audio files that share similar recording setups. In our experiments, we find that the new splits serve as a more challenging benchmark.
Benno Weck, Xavier Serra
ICASSP2
2023 Multilabel Prototype Generation for data reduction in K-Nearest Neighbour classification
abstract
Prototype Generation (PG) methods are typically considered for improving the efficiency of the k-Nearest Neighbour (kNN) classifier when tackling high-size corpora. Such approaches aim at generating a reduced version of the corpus without decreasing the classification performance when compared to the initial set. Despite their large application in multiclass scenarios, very few works have addressed the proposal of PG methods for the multilabel space. In this regard, this work presents the novel adaptation of four multiclass PG strategies to the multilabel case. These proposals are evaluated with three multilabel kNN-based classifiers, 12 corpora comprising a varied range of domains and corpus sizes, and different noise scenarios artificially induced in the data. The results obtained show that the proposed adaptations are capable of significantly improving—both in terms of efficiency and classification performance—the only reference multilabel PG work in the literature as well as the case in which no PG method is applied, also presenting statistically superior robustness in noisy scenarios. Moreover, these novel PG strategies allow prioritising either the efficiency or efficacy criteria through its configuration depending on the target scenario, hence covering a wide area in the solution space not previously filled by other works.
Jose J. Valero-Mas, Antonio Javier Gallego 0001, Pablo Alonso-Jiménez, Xavier Serra
Pattern Recognit.4
2022 An Overview of Automatic Piano Performance Assessment within the Music Education Context
abstract
Comunicació presentada a la 14th International Conference on Computer Supported Education, celebrada del 22 a 24 d'abril de 2021 de manera virtual.
Hyon Kim, Pedro Ramoneda, Marius Miron, Xavier Serra
CSEDU (1)4
2022 Urban Sound & Sight: Dataset And Benchmark For Audio-Visual Urban Scene Understanding
abstract
Automatic audio-visual urban traffic understanding is a growing area of research with many potential applications of value to industry, academia, and the public sector. Yet, the lack of well-curated resources for training and evaluating models to research in this area hinders their development. To address this we present a curated audio-visual dataset, Urban Sound & Sight (Urbansas), developed for investigating the detection and localization of sounding vehicles in the wild. Urbansas consists of 12 hours of unlabeled data along with 3 hours of manually annotated data, including bounding boxes with classes and unique id of vehicles, and strong audio labels featuring vehicle types and indicating off-screen sounds. We discuss the challenges presented by the dataset and how to use its annotations for the localization of vehicles in the wild through audio models.
Magdalena Fuentes, Bea Steers, Pablo Zinemanas, Martín Rocamora, Luca Bondi, Julia Wilkins, Qianyi Shi, Yao Hou, Samarjit Das, Xavier Serra, Juan Pablo Bello
ICASSP10
2022 Score Difficulty Analysis for Piano Performance Education based on Fingering
abstract
In this paper, we introduce score difficulty classification as a sub-task of music information retrieval (MIR), which may be used in music education technologies, for personalised curriculum generation, and score retrieval. We introduce a novel dataset for our task, Mikrokosmos-difficulty, containing 147 piano pieces in symbolic representation and the corresponding difficulty labels derived by its composer Béla Bartók and the publishers. As part of our methodology, we propose piano technique feature representations based on different piano fingering algorithms. We use these features as input for two classifiers: a Gated Recurrent Unit neural network (GRU) with attention mechanism and gradient-boosted trees trained on score segments. We show that for our dataset fingering based features perform better than a simple baseline considering solely the notes in the score. Furthermore, the GRU with attention mechanism classifier surpasses the gradient-boosted trees. Our proposed models are interpretable and are capable of generating difficulty feedback both locally, on short term segments, and globally, for whole pieces. Code, datasets, models, and an online demo are made available for reproducibility.
Pedro Ramoneda, Nazif Can Tamer, Vsevolod Eremenko, Xavier Serra, Marius Miron
ICASSP4
2022 Automatic Piano Fingering from Partially Annotated Scores using Autoregressive Neural Networks
abstract
Piano fingering is a creative and highly individualised task acquired by musicians progressively in their first music education years. Pianists must learn to choose the order of fingers to play the piano keys because scores do not have engraved finger and hand movements as other technique elements. Numerous research efforts have been conducted for automatic piano fingering based on a previous dataset composed of 150 score excerpts fully annotated by multiple expert annotators. However, most piano sheets include partial annotations for problematic finger and hand movements. We introduce a novel dataset for the task, the ThumbSet dataset, containing 2523 pieces with partial and noisy annotations of piano fingering crowdsourced from non-expert annotators. As part of our methodology, we propose two autoregressive neural networks with beam search decoding for modelling automatic piano fingering as a sequence-to-sequence learning problem, considering the correlation between output finger labels. We design the first model with the exact pitch representation of previous proposals. The second model uses graph neural networks to more effectively represent polyphony, whose treatment has been a common issue across previous studies. Finally, we finetune the models on the existing expert annotations dataset. The evaluation shows that (1) we are able to achieve high performance when training on the ThumbSet dataset and that (2) the proposed models outperform the state-of-the-art hidden Markov models and recurrent neural network baselines. Code, dataset, models, and results are made available to enhance the task reproducibility, including a new framework for evaluation.
Pedro Ramoneda, Dasaem Jeong, Eita Nakamura, Xavier Serra, Marius Miron
ACM Multimedia4
2022 FSD50K: An Open Dataset of Human-Labeled Sound Events
abstract
Most existing datasets for sound event recognition (SER) are relatively small and/or domain-specific, with the exception of AudioSet, based on over 2 M tracks from YouTube videos and encompassing over 500 sound classes. However, AudioSet is not an open dataset as its official release consists of pre-computed audio features. Downloading the original audio tracks can be problematic due to YouTube videos gradually disappearing and usage rights issues. To provide an alternative benchmark dataset and thus foster SER research, we introduceFSD50K, an open dataset containing over 51 k audio clips totalling over 100 h of audio manually labeled using 200 classes drawn from the AudioSet Ontology. The audio clips are licensed under Creative Commons licenses, making the dataset freely distributable (including waveforms). We provide a detailed description of the FSD50K creation process, tailored to the particularities of Freesound data, including challenges encountered and solutions adopted. We include a comprehensive dataset characterization along with discussion of limitations and key factors to allow its audio-informed usage. Finally, we conduct sound event classification experiments to provide baseline systems as well as insight on the main factors to consider when splitting Freesound audio data for SER. Our goal is to develop a dataset to be widely adopted by the community as a new open benchmark for SER research.
Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, Xavier Serra
IEEE ACM Trans. Audio Speech Lang. Process.5
2021 Break the Loop: Gender Imbalance in Music Recommenders
abstract
As recommender systems play an important role in everyday life, there is an increasing pressure that such systems are fair. Besides serving diverse groups of users, recommenders need to represent and serve item providers fairly as well. In interviews with music artists, we identified that gender fairness is one of the artists' main concerns. They emphasized that female artists should be given more exposure in music recommendations. We analyze a widely-used collaborative filtering approach with two public datasets-enriched with gender information-to understand how this approach performs with respect to the artists' gender. To achieve gender balance, we propose a progressive re-ranking method that is based on the insights from the interviews. For the evaluation, we rely on a simulation of feedback loops and provide an in-depth analysis using state-of-the-art performance measures and metrics concerning gen-der fairness.
Andres Ferraro, Xavier Serra, Christine Bauer 0001
CHIIR2
2021 Loopnet: Musical Loop Synthesis Conditioned on Intuitive Musical Parameters
abstract
Loops, seamlessly repeatable musical segments, are a cornerstone of modern music production. Contemporary artists often mix and match various sampled or pre-recorded loops based on musical criteria such as rhythm, harmony and timbral texture to create compositions. Taking such criteria into account, we present LoopNet, a feed-forward generative model for creating loops conditioned on intuitive parameters. We leverage Music Information Retrieval (MIR) models as well as a large collection of public loop samples in our study and use the Wave-U-Net architecture to map control parameters to audio. We also evaluate the quality of the generated audio and propose intuitive controls for composers to map the ideas in their minds to an audio loop.
Pritish Chandna, António Ramires, Xavier Serra, Emilia Gómez
ICASSP3
2021 Learning Contextual Tag Embeddings for Cross-Modal Alignment of Audio and Tags
abstract
Self-supervised audio representation learning offers an attractive alternative for obtaining generic audio embeddings, capable to be employed into various downstream tasks. Published approaches that consider both audio and words/tags associated with audio do not employ text processing models that are capable to generalize to tags unknown during training. In this work we propose a method for learning audio representations using an audio autoencoder (AAE), a general word embed-dings model (WEM), and a multi-head self-attention (MHA) mechanism. MHA attends on the output of the WEM, pro-viding a contextualized representation of the tags associated with the audio, and we align the output of MHA with the out-put of the encoder of AAE using a contrastive loss. We jointly optimize AAE and MHA and we evaluate the audio representations (i.e. the output of the encoder of AAE) by utilizing them in three different downstream tasks, namely sound, music genre, and music instrument classification. Our results show that employing multi-head self-attention with multiple heads in the tag-based network can induce better learned audio representations.
Xavier Favory, Konstantinos Drossos, Tuomas Virtanen, Xavier Serra
ICASSP4
2021 Melon Playlist Dataset: A Public Dataset for Audio-Based Playlist Generation and Music Tagging
abstract
One of the main limitations in the field of audio signal processing is the lack of large public datasets with audio representations and high-quality annotations due to restrictions of copyrighted commercial music. We present Melon Playlist Dataset, a public dataset of mel-spectrograms for 649,091 tracks and 148,826 associated playlists annotated by 30,652 different tags. All the data is gathered from Melon, a popular Korean streaming service. The dataset is suitable for music information retrieval tasks, in particular, auto-tagging and automatic playlist continuation. Even though the latter can be addressed by collaborative filtering approaches, audio provides opportunities for research on track suggestions and building systems resistant to the cold-start problem, for which we provide a baseline. Moreover, the playlists and the annotations included in the Melon Playlist Dataset make it suitable for metric learning and representation learning.
Andres Ferraro, Yuntae Kim, Soohyeon Lee, Biho Kim, Namjun Jo, Semi Lim, Suyon Lim, Jungtaek Jang, Sehwan Kim, Xavier Serra, Dmitry Bogdanov
ICASSP10
2021 Unsupervised Contrastive Learning of Sound Event Representations
abstract
Self-supervised representation learning can mitigate the limitations in recognition tasks with few manually labeled data but abundant unlabeled data—a common scenario in sound event research. In this work, we explore unsupervised contrastive learning as a way to learn sound event representations. To this end, we propose to use the pretext task of contrasting differently augmented views of sound events. The views are computed primarily via mixing of training examples with unrelated backgrounds, followed by other data augmentations. We analyze the main components of our method via ablation experiments. We evaluate the learned representations using linear evaluation, and in two in-domain downstream sound event classification tasks, namely, using limited manually labeled data, and using noisy labeled data. Our results suggest that unsupervised contrastive pre-training can mitigate the impact of data scarcity and increase robustness against noisy labels.
Eduardo Fonseca, Diego Ortego, Kevin McGuinness, Noel E. O'Connor, Xavier Serra
ICASSP5
2021 Multimodal Metric Learning for Tag-Based Music Retrieval
abstract
Tag-based music retrieval is crucial to browse large-scale mu-sic libraries efficiently. Hence, automatic music tagging has been actively explored, mostly as a classification task, which has an inherent limitation: a fixed vocabulary. On the other hand, metric learning enables flexible vocabularies by using pretrained word embeddings as side information. Also, met-ric learning has proven its suitability for cross-modal retrieval tasks in other domains (e.g., text-to-image) by jointly learning a multimodal embedding space. In this paper, we investigate three ideas to successfully introduce multimodal metric learning for tag-based music retrieval: elaborate triplet sampling, acoustic and cultural music information, and domain-specific word embeddings. Our experimental results show that the proposed ideas enhance the retrieval system quantitatively and qualitatively. Furthermore, we release the MSD500: a subset of the Million Song Dataset (MSD) containing 500 cleaned tags, 7 manually annotated tag categories, and user taste profiles.
Minz Won, Sergio Oramas, Oriol Nieto, Fabien Gouyon, Xavier Serra
ICASSP5
2021 What Is Fair? Exploring the Artists' Perspective on the Fairness of Music Streaming Platforms
Andres Ferraro, Xavier Serra, Christine Bauer 0001
INTERACT (2)2
2020 Performance Assessment Technologies for the Support of Musical Instrument Learning
abstract
Comunicació presentada a: CSEDU 2020 The 12th International Conference on Computer Supported Education, celebrada del 2 al 4 de maig de 2020, en línia.
Vsevolod Eremenko, Alia Morsi, Jyoti Narang, Xavier Serra
CSEDU (1)4
2020 Tensorflow Audio Models in Essentia
abstract
Essentia is a reference open-source C++/Python library for audio and music analysis. In this work, we present a set of algorithms that employ TensorFlow in Essentia, allow predictions with pre-trained deep learning models, and are designed to offer flexibility of use, easy extensibility, and real-time inference. To show the potential of this new interface with TensorFlow, we provide a number of pre-trained state-of-the-art music tagging and classification CNN models. We run an extensive evaluation of the developed models. In particular, we assess the generalization capabilities in a cross-collection evaluation utilizing both external tag datasets as well as manual annotations tailored to the taxonomies of our models.
Pablo Alonso-Jiménez, Dmitry Bogdanov, Jordi Pons, Xavier Serra
ICASSP4
2020 Neural Percussive Synthesis Parameterised by High-Level Timbral Features
abstract
We present a deep neural network-based methodology for synthesising percussive sounds with control over high-level timbral characteristics of the sounds. This approach allows for intuitive control of a synthesizer, enabling the user to shape sounds without extensive knowledge of signal processing. We use a feedforward convolutional neural network-based architecture, which is able to map input parameters to the corresponding waveform. We propose two datasets to evaluate our approach on both a restrictive context, and in one covering a broader spectrum of sounds. The timbral features used as parameters are taken from recent literature in signal processing. We also use these features for evaluation and validation of the presented model, to ensure that changing the input parameters produces a congruent waveform with the desired characteristics. Finally, we evaluate the quality of the output sound using a subjective listening test. We provide sound examples and the system's source code for reproducibility.
António Ramires, Pritish Chandna, Xavier Favory, Emilia Gómez, Xavier Serra
ICASSP5
2020 Data-Driven Harmonic Filters for Audio Representation Learning
abstract
We introduce a trainable front-end module for audio representation learning that exploits the inherent harmonic structure of audio signals. The proposed architecture, composed of a set of filters, compels the subsequent network to capture harmonic relations while preserving spectro-temporal locality. Since the harmonic structure is known to have a key role in human auditory perception, one can expect these harmonic filters to yield more efficient audio representations. Experimental results show that a simple convolutional neural network back-end with the proposed front-end outperforms state-of-the-art baseline methods in automatic music tagging, keyword spotting, and sound event tagging tasks.
Minz Won, Sanghyuk Chun, Oriol Nieto, Xavier Serra
ICASSP4
2020 Search Result Clustering in Collaborative Sound Collections
abstract
The large size of nowadays' online multimedia databases makes retrieving their content a difficult and time-consuming task. Users of online sound collections typically submit search queries that express a broad intent, often making the system return large and unmanageable result sets. Search Result Clustering is a technique that organises search-result content into coherent groups, which allows users to identify useful subsets in their results. Obtaining coherent and distinctive clusters that can be explored with a suitable interface is crucial for making this technique a useful complement of traditional search engines. In our work, we propose a graph-based approach using audio features for clustering diverse sound collections obtained when querying large online databases. We propose an approach to assess the performance of different features at scale, by taking advantage of the metadata associated with each sound. This analysis is complemented with an evaluation using ground-truth labels from manually annotated datasets. We show that using a confidence measure for discarding inconsistent clusters improves the quality of the partitions. After identifying the most appropriate features for clustering, we conduct an experiment with users performing a sound design task, in order to evaluate our approach and its user interface. A qualitative analysis is carried out including usability questionnaires and semi-structured interviews. This provides us with valuable new insights regarding the features that promote efficient interaction with the clusters.
Xavier Favory, Frederic Font, Xavier Serra
ICMR3
2020 Exploring Longitudinal Effects of Session-based Recommendations
abstract
Session-based recommendation is a problem setting where the task of a recommender system is to make suitable item suggestions based only on a few observed user interactions in an ongoing session. The lack of long-term preference information about individual users in such settings usually results in a limited level of personalization, where a small set of popular items may be recommended to many users. This repeated exposure of such a subset of the items through the recommendations may in turn lead to a reinforcement effect over time, and to a system which is not able to help users discover new content anymore to the desirable extent.
Andres Ferraro, Dietmar Jannach, Xavier Serra
RecSys3
2019 Learning Sound Event Classifiers from Web Audio with Noisy Labels
abstract
As sound event classification moves towards larger datasets, issues of label noise become inevitable. Web sites can supply large volumes of user-contributed audio and metadata, but inferring labels from this metadata introduces errors due to unreliable inputs, and limitations in the mapping. There is, however, little research into the impact of these errors. To foster the investigation of label noise in sound event classification we present FSDnoisy18k, a dataset containing 42.5 hours of audio across 20 sound classes, including a small amount of manually-labeled data and a larger quantity of real-world noisy data. We characterize the label noise empirically, and provide a CNN baseline system. Experiments suggest that training with large amounts of noisy data can outperform training with smaller amounts of carefully-labeled data. We also show that noise-robust loss functions can be effective in improving performance in presence of corrupted labels.
Eduardo Fonseca, Manoj Plakal, Daniel P. W. Ellis, Frederic Font, Xavier Favory, Xavier Serra
ICASSP6
2019 Randomly Weighted CNNs for (Music) Audio Classification
abstract
The computer vision literature shows that randomly weighted neural networks perform reasonably as feature extractors. Following this idea, we study how non-trained (randomly weighted) convolutional neural networks perform as feature extractors for (music) audio classification tasks. We use features extracted from the embeddings of deep architectures as input to a classifier - with the goal to compare classification accuracies when using different randomly weighted architectures. By following this methodology, we run a comprehensive evaluation of the current architectures for audio classification, and provide evidence that the architectures alone are an important piece for resolving (music) audio problems using deep neural networks.
Jordi Pons, Xavier Serra
ICASSP2
2019 Training Neural Audio Classifiers with Few Data
abstract
We investigate supervised learning strategies that improve the training of neural network audio classifiers on small annotated collections. In particular, we study whether (i) a naive regularization of the solution space, (ii) prototypical networks, (iii) transfer learning, or (iv) their combination, can foster deep learning models to better leverage a small amount of training examples. To this end, we evaluate (i-iv) for the tasks of acoustic event recognition and acoustic scene classification, considering from 1 to 100 labeled examples per class. Results indicate that transfer learning is a powerful strategy in such scenarios, but prototypical networks show promising results when one does not count with external or validation data.
Jordi Pons, Joan Serrà, Xavier Serra
ICASSP3
2019 End-to-End Music Source Separation: Is it Possible in the Waveform Domain?
abstract
Most of the currently successful source separation techniques use the magnitude spectrogram as input, and are therefore by default omitting part of the signal: the phase. To avoid omitting potentially useful information, we study the viability of using end-to-end models for music source separation --- which take into account all the information available in the raw audio signal, including the phase. Although during the last decades end-to-end music source separation has been considered almost unattainable, our results confirm that waveform-based models can perform similarly (if not better) than a spectrogram-based deep learning model. Namely: a Wavenet-based model we propose and Wave-U-Net can outperform DeepConvSep, a recent spectrogram-based deep learning model.
Francesc Lluís, Jordi Pons, Xavier Serra
INTERSPEECH3
2018 A Wavenet for Speech Denoising
abstract
Most speech processing techniques use magnitude spectrograms as front-end and are therefore by default discarding part of the signal: the phase. In order to overcome this limitation' we propose an end-to-end learning method for speech denoising based on Wavenet. The proposed model adaptation retains Wavenet's powerful acoustic modeling capabilities, while significantly reducing its time-complexity by eliminating its autoregressive nature. Specifically, the model makes use of non-causal, dilated convolutions and predicts target fields instead of a single target sample. The discriminative adaptation of the model we propose, learns in a supervised fashion via minimizing a regression loss. These modifications make the model highly parallelizable during both training and inference. Both quantitative and qualitative evaluations indicate that the proposed method is preferred over Wiener filtering, a common method based on processing the magnitude spectrogram.
Dario Rethage, Jordi Pons, Xavier Serra
ICASSP3
2018 Singing Voice Phoneme Segmentation by Hierarchically Inferring Syllable and Phoneme Onset Positions
abstract
Comunicació presentada a: Interspeech 2018, celebrada del 2 al 6 de setembre de 2018 a Hyderabad, India.
Rong Gong, Xavier Serra
INTERSPEECH2
2018 Multi criteria biased randomized method for resource allocation in distributed systems: Application in a volunteer computing system
Javier Panadero, Jésica de Armas, Xavier Serra, Joan Manuel Marquès
Future Gener. Comput. Syst.3
2017 Designing efficient architectures for modeling temporal features with convolutional neural networks
abstract
Many researchers use convolutional neural networks with small rectangular filters for music (spectrograms) classification. First, we discuss why there is no reason to use this filters setup by default and second, we point that more efficient architectures could be implemented if the characteristics of the music features are considered during the design process. Specifically, we propose a novel design strategy that might promote more expressive and intuitive deep learning architectures by efficiently exploiting the representational capacity of the first layer - using different filter shapes adapted to fit musical concepts within the first layer. The proposed architectures are assessed by measuring their accuracy in predicting the classes of the Ballroom dataset. We also make available1the used code (together with the audio-data) so that this research is fully reproducible.
Jordi Pons, Xavier Serra
ICASSP2
2017 Sound and Music Recommendation with Knowledge Graphs
abstract
The Web has moved, slowly but steadily, from a collection of documents towards a collection of structured data. Knowledge graphs have then emerged as a way of representing the knowledge encoded in such data as well as a tool to reason on them in order to extract new and implicit information. Knowledge graphs are currently used, for example, to explain search results, to explore knowledge spaces, to semantically enrich textual documents, or to feed knowledge-intensive applications such as recommender systems. In this work, we describe how to create and exploit a knowledge graph to supply a hybrid recommendation engine with information that builds on top of a collections of documents describing musical and sound items. Tags and textual descriptions are exploited to extract and link entities to external graphs such as WordNet and DBpedia, which are in turn used to semantically enrich the initial data. By means of the knowledge graph we build, recommendations are computed using a feature combination hybrid approach. Two explicit graph feature mappings are formulated to obtain meaningful item feature representations able to catch the knowledge embedded in the graph. Those content features are further combined with additional collaborative information deriving from implicit user feedback. An extensive evaluation on historical data is performed over two different datasets: a dataset of sounds composed of tags, textual descriptions, and user’s download information gathered from Freesound.org and a dataset of songs that mixes song textual descriptions with tags and user’s listening habits extracted from Songfacts.com and Last.fm, respectively. Results show significant improvements with respect to state-of-the-art collaborative algorithms in both datasets. In addition, we show how the semantic expansion of the initial descriptions helps in achieving much better recommendation quality in terms of aggregated diversity and novelty.
Sergio Oramas, Vito Ostuni, Tommaso Di Noia, Xavier Serra, Eugenio Di Sciascio
ACM Trans. Intell. Syst. Technol.4
2016 Discovering rāga motifs by characterizing communities in networks of melodic patterns
abstract
Ra̅ga motifs are the main building blocks of the melodic structures in Indian art music. Therefore, the discovery and characterization of such motifs is fundamental for the computational analysis of this music. We propose an approach for discovering ra̅ga motifs from audio music collections. First, we extract melodic patterns from a collection of 44 hours of audio comprising 160 recordings belonging to 10 ra̅gas. Next, we characterize these patterns by performing a network analysis, detecting non-overlapping communities, and exploiting the topological properties of the network to determine a similarity threshold. With that, we select a number of motif candidates that are representative of a ra̅ga, the ra̅ga motifs. For a formal evaluation we perform listening tests with 10 professional musicians. The results indicate that, on an average, the selected melodic phrases correspond to ra̅ga motifs with 85% positive ratings. This opens up the possibilities for many musically-meaningful computational tasks in Indian art music, including human-interpretable ra̅ga recognition, semantic-based music discovery, or pedagogical tools.
Sankalp Gulati, Joan Serrà, Vignesh Ishwar, Xavier Serra
ICASSP4
2016 Phrase-based rĀga recognition using vector space modeling
abstract
Automatic raga recognition is one of the fundamental computational tasks in Indian art music. Motivated by the way seasoned listeners identify ragas, we propose a raga recognition approach based on melodic phrases. Firstly, we extract melodic patterns from a collection of audio recordings in an unsupervised way. Next, we group similar patterns by exploiting complex networks concepts and techniques. Drawing an analogy to topic modeling in text classification, we then represent audio recordings using a vector space model. Finally, we employ a number of classification strategies to build a predictive model for raga recognition. To evaluate our approach, we compile a music collection of over 124 hours, comprising 480 recordings and 40 ragas. We obtain 70% accuracy with the full 40-raga collection, and up to 92% accuracy with its 10-raga subset. We show that phrase-based raga recognition is a successful strategy, on par with the state of the art, and sometimes outperforms it. A by-product of our approach, which arguably is as important as the task of raga recognition, is the identification of raga-phrases. These phrases can be used as a dictionary of semantically-meaningful melodic units for several computational tasks in Indian art music.
Sankalp Gulati, Joan Serrà, Vignesh Ishwar, Sertan Sentürk, Xavier Serra
ICASSP5
2016 A generalized Bayesian model for tracking long metrical cycles in acoustic music signals
abstract
Most musical phenomena involve repetitive structures that enable listeners to track meter, i.e. the tactus or beat, the longer over-arching measure or bar, and possibly other related layers. Meters with long measure duration, sometimes lasting more than a minute, occur in many music cultures, e.g. from India, Turkey, and Korea. However, current meter tracking algorithms, which were devised for cycles of a few seconds length, cannot process such structures accurately. We present a novel generalization to an existing Bayesian model for meter tracking that overcomes this limitation. The proposed model is evaluated on a set of Indian Hindustani music recordings, and we document significant performance increase over the previous models. The presented model opens the way for computational analysis of performances with long metrical cycles, and has important applications in music studies as well as in commercial applications that involve such musics.
Ajay Srinivasamurthy, Andre Holzapfel, A. Taylan Cemgil, Xavier Serra
ICASSP4
2016 Improving Audio Retrieval through Loudness Profile Categorization
abstract
The increasing popularity of audio content sharing in online platforms requires the development of techniques to better organize and retrieve this data. In this paper we look at how to improve similarity search through content categorization in the context of Freesound, a popular online sound sharing site. We focus on organization based on morphological description. In particular, we propose to improve search results by incorporating information about query sound's loudness profile. This is performed within a thresholding based framework and can be generalized to structure information about the temporal evolution of other sound attributes. We perform a subjective evaluation to demonstrate the practical relevance of our method.
Sanjeel Parekh, Frederic Font, Xavier Serra
ISM3
2016 ELMD: An Automatically Generated Entity Linking Gold Standard Dataset in the Music Domain
Sergio Oramas, Luis Espinosa Anke, Mohamed Sordo, Horacio Saggion, Xavier Serra
LREC5
2016 Information extraction for knowledge base construction in the music domain
Sergio Oramas, Luis Espinosa Anke, Mohamed Sordo, Horacio Saggion, Xavier Serra
Data Knowl. Eng.5
2015 An evaluation of methodologies for melodic similarity in audio recordings of Indian art music
abstract
We perform a comparative evaluation of methodologies for computing similarity between short-time melodic fragments of audio recordings of Indian art music. We experiment with 560 different combinations of procedures and parameter values. These include the choices made for the sampling rate of the melody representation, pitch quantization levels, normalization techniques and distance measures. The dataset used for evaluation consists of 157 and 340 annotated melodic fragments of Carnatic and Hindustani music recordings, respectively. Our results indicate that melodic fragment similarity is particularly sensitive to distance measures and normalization techniques. Sampling rates do not have a significant impact for Hindustani music, but can significantly degrade the performance for Carnatic music. Overall, the performed evaluation provides a better understanding of the processing steps and parameter settings for melodic similarity in Indian art music. Importantly, it paves the way for developing unsupervised melodic pattern discovery approaches, whose evaluation is a challenging and, many times, ill-defined task.
Sankalp Gulati, Joan Serrà, Xavier Serra
ICASSP3
2015 Analysis of the Impact of a Tag Recommendation System in a Real-World Folksonomy
abstract
Collaborative tagging systems have emerged as a successful solution for annotating contributed resources to online sharing platforms, facilitating searching, browsing, and organizing their contents. To aid users in the annotation process, several tag recommendation methods have been proposed. It has been repeatedly hypothesized that these methods should contribute to improving annotation quality and reducing the cost of the annotation process. It has been also hypothesized that these methods should contribute to the consolidation of the vocabulary of collaborative tagging systems. However, to date, no empirical and quantitative result supports these hypotheses. In this work, we deeply analyze the impact of a tag recommendation system in the folksonomy of Freesound, a real-world and large-scale online sound sharing platform. Our results suggest that tag recommendation effectively increases vocabulary sharing among users of the platform. In addition, tag recommendation is shown to contribute to the convergence of the vocabulary as well as to a partial increase in the quality of annotations. However, according to our analysis, the cost of the annotation process does not seem to be effectively reduced. Our work is relevant to increase our understanding about the nature of tag recommendation systems and points to future directions for the further development of those systems and their analysis.
Frederic Font, Joan Serrà, Xavier Serra
ACM Trans. Intell. Syst. Technol.3
2014 A supervised approach to hierarchical metrical cycle tracking from audio music recordings
abstract
A supervised approach to metrical cycle tracking from audio is presented, with a main focus on tracking the tala, the hierarchical cyclic metrical structure in Carnatic music. Given the tala of a piece, we aim to estimate the aksara (lowest metrical pulse), the aksara period, and the sama (first pulse of the tala cycle). Starting with percussion enhanced audio, we estimate the aksara pulse period from a tempogram computed using an onset detection function. A novelty function is computed using a self similarity matrix constructed using frame level audio features. These are then used to estimate possible aksara and sama candidates, followed by a candidate selection based on periodicity constraints, which leads to the final estimates. The algorithm is tested on an annotated collection of 176 pieces spanning four different talas. Though applied to Carnatic music, the framework presented is general and can be extended to other music cultures with cyclical metrical structures.
Ajay Srinivasamurthy, Xavier Serra
ICASSP2
2014 A study of instrument-wise onset detection in Beijing Opera percussion ensembles
abstract
Note onset detection and instrument recognition are two of the most investigated tasks in Music Information Retrieval (MIR). Various detection methods have been proposed in previous research for western music, with less focus on other music cultures of the world. In this paper, we focus on onset detection for percussion instruments in Beijing Opera, a major genre of Chinese traditional music. A dataset of individual audio samples of four primary percussion instruments is used to obtain the spectral bases for each instrument. With these bases, we separate the input percussion ensemble recordings into its spectral sources and their activations using a Non-negative Matrix Factorization (NMF) based algorithm. A simple onset detection conducted on each NMF activation presents satisfactory overall detection rates, and provides us valuable implications and suggestions for future development of drum transcription and percussion pattern analysis in Beijing Opera.
Mi Tian 0001, Ajay Srinivasamurthy, Mark B. Sandler, Xavier Serra
ICASSP4
2014 Class-based tag recommendation and user-based evaluation in online audio clip sharing
Frederic Font, Joan Serrà, Xavier Serra
Knowl. Based Syst.3
2013 ESSENTIA: an open-source library for sound and music analysis
abstract
We present Essentia 2.0, an open-source C++ library for audio analysis and audio-based music information retrieval released under the Affero GPL license. It contains an extensive collection of reusable algorithms which implement audio input/output functionality, standard digital signal processing blocks, statistical characterization of data, and a large set of spectral, temporal, tonal and high-level music descriptors. The library is also wrapped in Python and includes a number of predefined executable extractors for the available music descriptors, which facilitates its use for fast prototyping and allows setting up research experiments very rapidly. Furthermore, it includes a Vamp plugin to be used with Sonic Visualiser for visualization purposes. The library is cross-platform and currently supports Linux, Mac OS X, and Windows systems. Essentia is designed with a focus on the robustness of the provided music descriptors and is optimized in terms of the computational cost of the algorithms. The provided functionality, specifically the music descriptors included in-the-box and signal processing algorithms, is easily expandable and allows for both research experiments and development of large-scale industrial applications.
Dmitry Bogdanov, Nicolas Wack, Emilia Gómez, Sankalp Gulati, Perfecto Herrera, Oscar Mayor, Gerard Roma, Justin Salamon, José Ricardo Zapata, Xavier Serra
ACM Multimedia10
2013 Freesound technical demo
abstract
Freesound is an online collaborative sound database where people with diverse interests share recorded sound samples under Creative Commons licenses. It was started in 2005 and it is being maintained to support diverse research projects and as a service to the overall research and artistic community. In this demo we want to introduce Freesound to the multimedia community and show its potential as a research resource. We begin by describing some general aspects of Freesound, its architecture and functionalities, and then explain potential usages that this framework has for research applications.
Frederic Font, Gerard Roma, Xavier Serra
ACM Multimedia3
2013 Folksonomy-Based Tag Recommendation for Collaborative Tagging Systems
abstract
Collaborative tagging has emerged as a common solution for labelling and organising online digital content. However, collaborative tagging systems typically suffer from a number of issues such as tag scarcity or ambiguous labelling. As a result, the organisation and browsing of tagged content is far from being optimal. In this work the authors present a general scheme for building a folksonomy-based tag recommendation system to help users tagging online content resources. Based on this general scheme, the authorse describe eight tag recommendation methods and extensively evaluate them with data coming from two real-world large-scale datasets of tagged images and sound clips. Their results show that the proposed methods can effectively recommend relevant tags, given a set of input tags and tag co-occurrence information. Moreover, the authors show how novel strategies for selecting the appropriate number of tags to be recommended can significantly improve methods performances. Approaches such as the one presented here can be useful to obtain more comprehensive and coherent descriptions of tagged resources, thus allowing a better organisation, browsing and reuse of online content. Moreover, they can increase the value of folksonomies as reliable sources for knowledge-mining.
Frederic Font, Joan Serrà, Xavier Serra
Int. J. Semantic Web Inf. Syst.3
2013 Evaluation in Music Information Retrieval
Julián Urbano, Markus Schedl, Xavier Serra
J. Intell. Inf. Syst.3
2012 Characterization and exploitation of community structure in cover song networks
Joan Serrà, Massimiliano Zanin, Perfecto Herrera, Xavier Serra
Pattern Recognit. Lett.4
2012 Predictability of Music Descriptor Time Series and its Application to Cover Song Detection
abstract
Intuitively, music has both predictable and unpredictable components. In this paper, we assess this qualitative statement in a quantitative way using common time series models fitted to state-of-the-art music descriptors. These descriptors cover different musical facets and are extracted from a large collection of real audio recordings comprising a variety of musical genres. Our findings show that music descriptor time series exhibit a certain predictability not only for short time intervals, but also for mid-term and relatively long intervals. This fact is observed independently of the descriptor, musical facet and time series model we consider. Moreover, we show that our findings are not only of theoretical relevance but can also have practical impact. To this end we demonstrate that music predictability at relatively long time intervals can be exploited in a real-world application, namely the automatic identification of cover songs (i.e., different renditions or versions of the same musical piece). Importantly, this prediction strategy yields a parameter-free approach for cover song identification that is substantially faster, allows for reduced computational storage and still maintains highly competitive accuracies when compared to state-of-the-art systems.
Joan Serrà, Holger Kantz, Xavier Serra, Ralph G. Andrzejak
IEEE Trans. Speech Audio Process.3
2012 A Rule-Based Evolutionary Approach to Music Performance Modeling
abstract
We describe an evolutionary approach to one of the most challenging problems in computer music: modeling how skilled musicians manipulate sound properties such as timing and amplitude in order to express their view of the emotional content of musical pieces. Starting with a collection of audio recordings of real performances, we apply a sequential-covering genetic algorithm in order to obtain computational models for different aspects of expressive performance. We use these models to automatically synthesize performances with the timing and energy expressiveness that characterizes the music generated by a professional musician. The reported results indicate that evolutionary computation is an appropriate technique for solving the problem considered. Specifically, our evolutionary algorithm provides a number of potential advantages over other supervised learning algorithms, such as a method for non-deterministically obtaining models capturing different possible interpretations of a musical piece.
Rafael Ramírez 0001, Esteban Maestre, Xavier Serra
IEEE Trans. Evol. Comput.3
2011 Extending Sound Sample Descriptions through the Extraction of Community Knowledge
Frederic Font, Xavier Serra
UMAP2
2011 Unifying Low-Level and High-Level Music Similarity Measures
abstract
Measuring music similarity is essential for multimedia retrieval. For music items, this task can be regarded as obtaining a suitable distance measurement between songs defined on a certain feature space. In this paper, we propose three of such distance measures based on the audio content: first, a low-level measure based on tempo-related description; second, a high-level semantic measure based on the inference of different musical dimensions by support vector machines. These dimensions include genre, culture, moods, instruments, rhythm, and tempo annotations. Third, a hybrid measure which combines the above-mentioned distance measures with two existing low-level measures: a Euclidean distance based on principal component analysis of timbral, temporal, and tonal descriptors, and a timbral distance based on single Gaussian Mel-frequency cepstral coefficient (MFCC) modeling. We evaluate our proposed measures against a number of baseline measures. We do this objectively based on a comprehensive set of music collections, and subjectively based on listeners' ratings. Results show that the proposed methods achieve accuracies comparable to the baseline approaches in the case of the tempo and classifier-based measures. The highest accuracies are obtained by the hybrid distance. Furthermore, the proposed classifier-based approach opens up the possibility to explore distance measures that are based on semantic notions.
Dmitry Bogdanov, Joan Serrà, Nicolas Wack, Perfecto Herrera, Xavier Serra
IEEE Trans. Multim.5
2010 Indexing music by mood: design and integration of an automatic content-based annotator
Cyril Laurier, Owen Meyers, Joan Serrà, Martin Blech, Perfecto Herrera, Xavier Serra
Multim. Tools Appl.6
2010 Automatic performer identification in commercial monophonic Jazz performances
Rafael Ramírez 0001, Esteban Maestre, Xavier Serra
Pattern Recognit. Lett.3
2009 What/when causal expectation modelling applied to audio signals
abstract
A causal system to represent a stream of music into musical events, and to generate further expected events, is presented. Starting from an auditory front-end that extracts low-level (i.e. MFCC) and mid-level features such as onsets and beats, an unsupervised clustering process builds and maintains a set of symbols aimed at representing musical stream events using both timbre and time descriptions. The time events are represented using inter-onset intervals relative to the beats. These symbols are then processed by an expectation module using Predictive Partial Match, a multiscale technique based on N-grams. To characterise the ability of the system to generate an expectation that matches both ground truth and system transcription, we introduce several measures that take into account the uncertainty associated with the unsupervised encoding of the musical sequence. The system is evaluated using a subset of the ENST-drums database of annotated drum recordings. We compare three approaches to combine timing (when) and timbre (what) expectation. In our experiments, we show that the induced representation is useful for generating expectation patterns in a causal way.
Amaury Hazan, Ricard Marxer, Paul Brossier, Hendrik Purwins, Perfecto Herrera, Xavier Serra
Connect. Sci.6
2008 Chroma Binary Similarity and Local Alignment Applied to Cover Song Identification
abstract
We present a new technique for audio signal comparison based on tonal subsequence alignment and its application to detect cover versions (i.e., different performances of the same underlying musical piece). Cover song identification is a task whose popularity has increased in the music information retrieval (MIR) community along in the past, as it provides a direct and objective way to evaluate music similarity algorithms. This paper first presents a series of experiments carried out with two state-of-the-art methods for cover song identification. We have studied several components of these (such as chroma resolution and similarity, transposition, beat tracking or dynamic time warping constraints), in order to discover which characteristics would be desirable for a competitive cover song identifier. After analyzing many cross-validated results, the importance of these characteristics is discussed, and the best performing ones are finally applied to the newly proposed method. Multiple evaluations of this one confirm a large increase in identification accuracy when comparing it with alternative state-of-the-art approaches.
Joan Serrà, Emilia Gómez, Perfecto Herrera, Xavier Serra
IEEE Trans. Speech Audio Process.4
2008 FOAFing the music: Bridging the semantic gap in music recommendation
Òscar Celma, Xavier Serra
J. Web Semant.2
2007 State of the Art and Future Directions in Musical Sound Synthesis
abstract
Sound synthesis and processing has been the most active research topic in the field of sound and music computing for more than 40 years. Quite a number of the early research results are now standard components of many audio and music devices and new technologies are continuously being developed and integrated into new products. Through the years there have been important changes. For example, most of the abstract algorithms that were the focus of work in the 70s and 80s are considered obsolete. Then the 1990s saw the emergence of computational approaches that aimed either at capturing the characteristics of a sound source, known as physical models, or at capturing the perceptual characteristics of the sound signal, generally referred to as spectral or signal models. More recent trends include the combination of physical and spectral models and the corpus-based concatenative methods. But the field faces major challenges that might revolutionize the standard paradigms and applications of sound synthesis. In this article, we will first place the sound synthesis topic within its research context, then we will highlight some of the current trends, and finally we will attempt to identify some challenges for the future.
Xavier Serra
MMSP1
2007 Performance-Based Interpreter Identification in Saxophone Audio Recordings
abstract
We propose a novel approach to the task of identifying performers from their playing styles. We investigate how skilled musicians (Jazz saxophone players in particular) express and communicate their view of the musical and emotional content of musical pieces and how to use this information in order to automatically identify performers. We study deviations of parameters such as pitch, timing, amplitude and timbre both at an inter-note level and at an intra-note level. Our approach to performer identification consists of establishing a performer dependent mapping of inter-note features (essentially a "score" whether or not the score physically exists) to a repertoire of inflections characterized by intra-note features. We present a successful performer identification case study
Rafael Ramírez 0001, Esteban Maestre, Antonio Pertusa, Emilia Gómez, Xavier Serra
IEEE Trans. Circuits Syst. Video Technol.5