VLDB 2026 Research / reviewers in the wild / expert
Joan Serrà
dblp:67/3884
· DBLP profile ↗
54ranked-venue papers
18as first author
16since 2021 · last 2026
0000-0003-1303-6558ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 21 · 12 first-author · 4 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Emergent, not Immanent: A Baradian Reading of Explainable AIabstractExplainable AI (XAI) is frequently positioned as a technical problem of revealing the inner workings of an AI model. This position is affected by unexamined onto-epistemological assumptions: meaning is treated as immanent to the model, the explainer is positioned outside the system, and a causal structure is presumed recoverable through computational techniques. In this paper, we draw on Barad’s agential realism to develop an alternative onto-epistemology of XAI. We propose that interpretations are material-discursive performances that emerge from situated entanglements of the AI model with humans, context, and the interpretative apparatus. To develop this position, we read a comprehensive set of XAI methods through agential realism and reveal the assumptions and limitations that underpin several of these methods. We then articulate the framework’s ethical dimension and propose design directions for XAI interfaces that support emergent interpretation, using a speculative text-to-music interface as a case study. Fabio Morreale, Joan Serrà, Yuki Mitsufuji |
CHI | 2 |
| 2025 | Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration With Improved IntelligibilityabstractSpeech restoration aims at restoring full-band speech with high quality and intelligibility, considering a diverse set of distortions. MaskSR is a recently proposed generative model for this task. As other models of its kind, MaskSR attains high quality but, as we show, intelligibility can be substantially improved. We do so by boosting the speech encoder component of MaskSR with predictions of semantic representations of the target speech, using a pre-trained self-supervised teacher model. Then, a masked language model is conditioned on the learned semantic features to predict acoustic tokens that encode low level spectral details of the target speech. We show that, with the same MaskSR model capacity and inference time, the proposed model, MaskSR2, significantly reduces the word error rate, a typical metric for intelligibility. MaskSR2 also achieves competitive word error rate among other models, while providing superior quality. An ablation study shows the effectiveness of various semantic representations. Xiaoyu Liu 0003, Joan Serrà, Santiago Pascual |
ICASSP | 3 |
| 2025 | Sequential Contrastive Audio-Visual LearningabstractContrastive learning has emerged as a powerful technique in audio-visual representation learning, leveraging the natural co-occurrence of audio and visual modalities in webscale video datasets. However, conventional contrastive audio-visual learning (CAV) methodologies often rely on aggregated representations derived through temporal aggregation, neglecting the intrinsic sequential nature of the data. This oversight raises concerns regarding the ability of standard approaches to capture and utilize fine-grained information within sequences. In response to this limitation, we propose sequential contrastive audiovisual learning (SCAV), which contrasts examples based on their non-aggregated representation space using multidimensional sequential distances. Audio-visual retrieval experiments with the VGGSound and Music datasets demonstrate the effectiveness of SCAV, with up to 3.5× relative improvements in recall against traditional aggregation-based contrastive learning and other previously proposed methods, which utilize more parameters and data. We also show that models trained with SCAV exhibit a significant degree of flexibility regarding the metric employed for retrieval, allowing us to use a hybrid retrieval approach that is both effective and efficient. Ioannis Tsiamas, Santiago Pascual, Chunghsin Yeh, Joan Serrà |
ICASSP | 4 |
| 2025 | Supervised Contrastive Learning from Weakly-Labeled Audio Segments for Musical Version MatchingabstractDetecting musical versions (different renditions of the same piece) is a challenging task with important applications. Because of the ground truth nature, existing approaches match musical versions at the track level (e.g., whole song). However, most applications require to match them at the segment level (e.g., 20s chunks). In addition, existing approaches resort to classification and triplet losses, disregarding more recent losses that could bring meaningful improvements. In this paper, we propose a method to learn from weakly annotated segments, together with a contrastive loss variant that outperforms well-studied alternatives. The former is based on pairwise segment distance reductions, while the latter modifies an existing loss following decoupling, hyper-parameter, and geometric considerations. With these two elements, we do not only achieve state-of-the-art results in the standard track-level evaluation, but we also obtain a breakthrough performance in a segment-level evaluation. We believe that, due to the generality of the challenges addressed here, the proposed methods may find utility in domains beyond audio or musical version matching. Joan Serrà, Recep Oguz Araz, Dmitry Bogdanov, Yuki Mitsufuji |
ICML | 1 |
| 2025 | A Comprehensive Real-World Assessment of Audio Watermarking Algorithms: Will They Survive Neural Codecs?
Yigitcan Özer, Woosung Choi, Joan Serrà, Mayank Kumar Singh, Wei-Hsiang Liao 0001, Yuki Mitsufuji |
INTERSPEECH | 3 |
| 2024 | Masked Generative Video-to-Audio Transformers with Enhanced Synchronicity
Santiago Pascual, Chunghsin Yeh, Ioannis Tsiamas, Joan Serrà |
ECCV (87) | 4 |
| 2024 | GASS: Generalizing Audio Source Separation with Large-Scale DataabstractUniversal source separation targets at separating the audio sources of an arbitrary mix, removing the constraint to operate on a specific domain like speech or music. Yet, the potential of universal source separation is limited because most existing works focus on mixes with predominantly sound events, and small training datasets also limit its potential for supervised learning. Here, we study a single general audio source separation (GASS) model trained to separate speech, music, and sound events in a supervised fashion with a large-scale dataset. We assess GASS models on a diverse set of tasks. Our strong in-distribution results show the feasibility of GASS models, and the competitive out-of-distribution performance in sound event and speech separation shows its generalization abilities. Yet, it is challenging for GASS models to generalize for separating out-of-distribution cinematic and music content. We also fine-tune GASS models on each dataset and consistently outperform the ones without pre-training. All fine-tuned models (except the music separation one) obtain state-of-the-art results in their respective benchmarks. Jordi Pons, Xiaoyu Liu 0003, Santiago Pascual, Joan Serrà |
ICASSP | 4 |
| 2023 | Quantitative Evidence on Overlooked Aspects of Enrollment Speaker Embeddings for Target Speaker SeparationabstractSingle channel target speaker separation (TSS) aims at extracting a speaker’s voice from a mixture of multiple talkers given an enrollment utterance of that speaker. A typical deep learning TSS framework consists of an upstream model that obtains enrollment speaker embeddings and a downstream model that performs the separation conditioned on the embeddings. In this paper, we look into several important but overlooked aspects of the enrollment embeddings, including the suitability of the widely used speaker identification embeddings, the introduction of the log-mel filterbank and self-supervised embeddings, and the embeddings’ cross-dataset generalization capability. Our results show that the speaker identification embeddings could lose relevant information due to a sub-optimal metric, training objective, or common pre-processing. In contrast, both the filterbank and the self-supervised embeddings preserve the integrity of the speaker information, but the former consistently out-performs the latter in a cross-dataset evaluation. The competitive separation and generalization performance of the previously over-looked filterbank embedding is consistent across our study, which calls for future research on better upstream features. Xiaoyu Liu 0003, Joan Serrà |
ICASSP | 3 |
| 2023 | Full-Band General Audio Synthesis with Score-Based DiffusionabstractRecent works have shown the capability of deep generative models to tackle general audio synthesis from a single label, producing a variety of impulsive, tonal, and environmental sounds. Such models operate on band-limited signals and, as a result of an autoregressive approach, they are typically conformed by pre-trained latent encoders and/or several cascaded modules. In this work, we propose a diffusion-based generative model for general audio synthesis, named DAG, which deals with full-band signals end-to-end in the waveform domain. Results show the superiority of DAG over existing label-conditioned generators in terms of both quality and diversity. More specifically, when compared to the state of the art, the band-limited and full-band versions of DAG achieve relative improvements that go up to 40 and 65%, respectively. We believe DAG is flexible enough to accommodate different conditioning schemas while providing good quality synthesis. Santiago Pascual, Gautam Bhattacharya, Chunghsin Yeh, Jordi Pons, Joan Serrà |
ICASSP | 5 |
| 2023 | Adversarial Permutation Invariant Training for Universal Sound SeparationabstractUniversal sound separation consists of separating mixes with arbitrary sounds of different types, and permutation invariant training (PIT) is used to train source agnostic models that do so. In this work, we complement PIT with adversarial losses but find it challenging with the standard formulation used in speech source separation. We overcome this challenge with a novel I-replacement context-based adversarial loss, and by training with multiple discriminators. Our experiments show that by simply improving the loss (keeping the same model and dataset) we obtain a non-negligible improvement of 1.4 dB SI-SNRIin the reverberant FUSS dataset. We also find adversarial PIT to be effective at reducing spectral holes, ubiquitous in mask-based separation models, which highlights the potential relevance of adversarial losses for source separation. Emilian Postolache, Jordi Pons, Santiago Pascual, Joan Serrà |
ICASSP | 4 |
| 2022 | On Loss Functions and Evaluation Metrics for Music Source SeparationabstractWe investigate which loss functions provide better separations via benchmarking an extensive set of those for music source separation. To that end, we first survey the most representative audio source separation losses we identified, to later consistently benchmark them in a controlled experimental setup. We also explore using such losses as evaluation metrics, via cross-correlating them with the results of a subjective test. Based on the observation that the standard signal-to-distortion ratio metric can be misleading in some scenarios, we study alternative evaluation metrics based on the considered losses. Enric Gusó, Jordi Pons, Santiago Pascual, Joan Serrà |
ICASSP | 4 |
| 2022 | Assessing Algorithmic Biases for Musical Version IdentificationabstractVersion identification (VI) systems now offer accurate and scalable solutions for detecting different renditions of a musical composition, allowing the use of these systems in industrial applications and throughout the wider music ecosystem. Such use can have an important impact on various stakeholders regarding recognition and financial benefits, including how royalties are circulated for digital rights management. In this work, we take a step toward acknowledging this impact and consider VI systems as socio-technical systems rather than isolated technologies. We propose a framework for quantifying performance disparities across 5 systems and 6 relevant side attributes: gender, popularity, country, language, year, and prevalence. We also consider 3 main stakeholders for this particular information retrieval use case: the performing artists of query tracks, those of reference (original) tracks, and the composers. By categorizing the recordings in our dataset using such attributes and stakeholders, we analyze whether the considered VI systems show any implicit biases. We find signs of disparities in identification performance for most of the groups we include in our analyses. We also find that learning- and rule-based systems behave differently for some attributes, which suggests an additional dimension to consider along with accuracy and scalability when evaluating VI systems. Lastly, we share our dataset to encourage VI researchers to take these aspects into account while building new systems. Furkan Yesiler, Marius Miron, Joan Serrà, Emilia Gómez |
WSDM | 3 |
| 2021 | Upsampling Artifacts in Neural Audio SynthesisabstractA number of recent advances in neural audio synthesis rely on up-sampling layers, which can introduce undesired artifacts. In computer vision, upsampling artifacts have been studied and are known as checkerboard artifacts (due to their characteristic visual pattern). However, their effect has been overlooked so far in audio processing. Here, we address this gap by studying this problem from the audio signal processing perspective. We first show that the main sources of upsampling artifacts are: (i) the tonal and filtering artifacts introduced by problematic upsampling operators, and (ii) the spectral replicas that emerge while upsampling. We then compare different upsampling layers, showing that nearest neighbor upsamplers can be an alternative to the problematic (but state-of-the-art) transposed and subpixel convolutions which are prone to introduce tonal artifacts. Jordi Pons, Santiago Pascual, Giulio Cengarle, Joan Serrà |
ICASSP | 4 |
| 2021 | SESQA: Semi-Supervised Learning for Speech Quality AssessmentabstractAutomatic speech quality assessment is an important, transversal task whose progress is hampered by the scarcity of human annotations, poor generalization to unseen recording conditions, and a lack of flexibility of existing approaches. In this work, we tackle these problems with a semi-supervised learning approach, combining available annotations with programmatically generated data, and using 3 different optimization criteria together with 5 complementary auxiliary tasks. Our results show that such a semi-supervised approach can cut the error of existing methods by more than 36%, while providing additional benefits in terms of reusable features or auxiliary outputs. Improvement is further corroborated with an out-of-sample test showing promising generalization capabilities. Joan Serrà, Jordi Pons, Santiago Pascual |
ICASSP | 1 |
| 2021 | Automatic Multitrack Mixing With A Differentiable Mixing Console Of Neural Audio EffectsabstractApplications of deep learning to automatic multitrack mixing are largely unexplored. This is partly due to the limited available data, coupled with the fact that such data is relatively unstructured and variable. To address these challenges, we propose a domain-inspired model with a strong inductive bias for the mixing task. We achieve this with the application of pre-trained sub-networks and weight sharing, as well as with a sum/difference stereo loss function. The proposed model can be trained with a limited number of examples, is permutation invariant with respect to the input ordering, and places no limit on the number of input sources. Furthermore, it produces human-readable mixing parameters, allowing users to manually adjust or refine the generated mix. Results from a perceptual evaluation involving audio engineers indicate that our approach generates mixes that outperform baseline approaches. To the best of our knowledge, this work demonstrates the first approach in learning multitrack mixing conventions from real-world data at the waveform level, without knowledge of the underlying mixing parameters. Christian J. Steinmetz, Jordi Pons, Santiago Pascual, Joan Serrà |
ICASSP | 4 |
| 2021 | Investigating the Efficacy of Music Version Retrieval Systems for Setlist IdentificationabstractThe setlist identification (SLI) task addresses a music recognition use case where the goal is to retrieve the metadata and times-tamps for all the tracks played in live music events. Due to various musical and non-musical changes in live performances, developing automatic SLI systems is still a challenging task that, despite its industrial relevance, has been under-explored in the academic literature. In this paper, we propose an end-to-end workflow that identifies relevant metadata and timestamps of live music performances using a version identification system. We compare 3 of such systems to investigate their suitability for this particular task. For developing and evaluating SLI systems, we also contribute a new dataset that contains 99.5 h of concerts with annotated metadata and timestamps, along with the corresponding reference set. The dataset is categorized by audio qualities and genres to analyze the performance of SLI systems in different use cases. Our approach can identify 68% of the annotated segments, with values ranging from 35% to 77% based on the genre. Finally, we evaluate our approach against a database of 56.8 k songs to illustrate the effect of expanding the reference set, where we can still identify 56% of the annotated segments. Furkan Yesiler, Emilio Molina, Joan Serrà, Emilia Gómez |
ICASSP | 3 |
| 2020 | Accurate and Scalable Version Identification Using Musically-Motivated EmbeddingsabstractThe version identification (VI) task deals with the automatic detection of recordings that correspond to the same underlying musical piece. Despite many efforts, VI is still an open problem, with much room for improvement, specially with regard to combining accuracy and scalability. In this paper, we present MOVE, a musically-motivated method for accurate and scalable version identification. MOVE achieves state-of-the-art performance on two publicly-available benchmark sets by learning scalable embeddings in an Euclidean distance space, using a triplet loss and a hard triplet mining strategy. It improves over previous work by employing an alternative input representation, and introducing a novel technique for temporal content summarization, a standardized latent space, and a data augmentation strategy specifically designed for VI. In addition to the main results, we perform an ablation study to highlight the importance of our design choices, and study the relation between embedding dimensionality and model performance. Furkan Yesiler, Joan Serrà, Emilia Gómez |
ICASSP | 2 |
| 2020 | Input Complexity and Out-of-distribution Detection with Likelihood-based Generative Models
Joan Serrà, David Álvarez 0004, Vicenç Gómez, Olga Slizovskaia, José F. Núñez, Jordi Luque |
ICLR | 1 |
| 2020 | Experience: advanced network operations in (Un)-connected remote communitiesabstractThe Internet Para Todos program is working to provide sustainable mobile broadband to 100 M unconnected people in Latin America. In this paper we present our commercial deployment in thousands remote small communities and describe the unique experience of maintaining this infrastructure. We describe the challenges related to managing operations containing the cost in these extreme geographical conditions. We also analyze operational data to understand outage patterns and present typical operational issues in this unique remote community environment. Finally, we present an extension of the operations support system (OSS) leveraging advanced analytics and machine learning with the goal of optimizing network maintenance while reducing costs. Diego Perino, Joan Serrà, Andra Lutu, Ilias Leontiadis |
MobiCom | 3 |
| 2019 | Training Neural Audio Classifiers with Few DataabstractWe investigate supervised learning strategies that improve the training of neural network audio classifiers on small annotated collections. In particular, we study whether (i) a naive regularization of the solution space, (ii) prototypical networks, (iii) transfer learning, or (iv) their combination, can foster deep learning models to better leverage a small amount of training examples. To this end, we evaluate (i-iv) for the tasks of acoustic event recognition and acoustic scene classification, considering from 1 to 100 labeled examples per class. Results indicate that transfer learning is a powerful strategy in such scenarios, but prototypical networks show promising results when one does not count with external or validation data. Jordi Pons, Joan Serrà, Xavier Serra |
ICASSP | 2 |
| 2019 | Learning Problem-Agnostic Speech Representations from Multiple Self-Supervised TasksabstractLearning good representations without supervision is still an open issue in machine learning, and is particularly challenging for speech signals, which are often characterized by long sequences with a complex hierarchical structure. Some recent works, however, have shown that it is possible to derive useful speech representations by employing a self-supervised encoder-discriminator approach. This paper proposes an improved self-supervised method, where a single neural encoder is followed by multiple workers that jointly solve different self-supervised tasks. The needed consensus across different tasks naturally imposes meaningful constraints to the encoder, contributing to discover general representations and to minimize the risk of learning superficial ones. Experiments show that the proposed approach can learn transferable, robust, and problem-agnostic features that carry on relevant information from the speech signal, such as speaker identity, phonemes, and even higher-level features such as emotional cues. In addition, a number of design choices make the encoder easily exportable, facilitating its direct usage or adaptation to different problems. Santiago Pascual, Mirco Ravanelli, Joan Serrà, Antonio Bonafonte, Yoshua Bengio |
INTERSPEECH | 3 |
| 2019 | Towards Generalized Speech Enhancement with Generative Adversarial NetworksabstractThe speech enhancement task usually consists of removing additive noise or reverberation that partially mask spoken utterances, affecting their intelligibility. However, little attention is drawn to other, perhaps more aggressive signal distortions like clipping, chunk elimination, or frequency-band removal. Such distortions can have a large impact not only on intelligibility, but also on naturalness or even speaker identity, and require of careful signal reconstruction. In this work, we give full consideration to this generalized speech enhancement task, and show it can be tackled with a time-domain generative adversarial network (GAN). In particular, we extend a previous GAN-based speech enhancement system to deal with mixtures of four types of aggressive distortions. Firstly, we propose the addition of an adversarial acoustic regression loss that promotes a richer feature extraction at the discriminator. Secondly, we also make use of a two-step adversarial training schedule, acting as a warm up-and-fine-tune sequence. Both objective and subjective evaluations show that these two additions bring improved speech reconstructions that better match the original speaker identity and naturalness. Santiago Pascual, Joan Serrà, Antonio Bonafonte |
INTERSPEECH | 2 |
| 2019 | Blow: a single-scale hyperconditioned flow for non-parallel raw-audio voice conversionabstractEnd-to-end models for raw audio generation are a challenge, specially if they have to work with non-parallel data, which is a desirable setup in many situations. Voice conversion, in which a model has to impersonate a speaker in a recording, is one of those situations. In this paper, we propose Blow, a single-scale normalizing flow using hypernetwork conditioning to perform many-to-many voice conversion between raw audio. Blow is trained end-to-end, with non-parallel data, on a frame-by-frame basis using a single speaker identifier. We show that Blow compares favorably to existing flow-based architectures and other competitive baselines, obtaining equal or better performance in both objective and subjective evaluations. We further assess the impact of its main components with an ablation study, and quantify a number of properties such as the necessary amount of training data or the preference for source or target speakers. Joan Serrà, Santiago Pascual, Carlos Segura |
NeurIPS | 1 |
| 2019 | Time-domain speech enhancement using generative adversarial networks
Santiago Pascual, Joan Serrà, Antonio Bonafonte |
Speech Commun. | 2 |
| 2018 | There goes Wally: Anonymously sharing your location gives you awayabstractWith current technology, a number of entities have access to user mobility traces at different levels of spatio-temporal granularity. At the same time, users frequently reveal their location through different means, including geo-tagged social media posts and mobile app usage. Such leaks are often bound to a pseudonym or a fake identity in an attempt to preserve one's privacy. In this work, we investigate how large-scale mobility traces can de-anonymize anonymous location leaks. By mining the country-wide mobility traces of tens of millions of users, we aim to understand how many location leaks are required to uniquely match a trace, how spatio-temporal obfuscation decreases the matching quality, and how the location popularity and time of the leak influence de-anonymization. We also study the mobility characteristics of those individuals whose anonymous leaks are more prone to identification. Finally, by extending our matching methodology to full traces, we show how large-scale human mobility is highly unique. Our quantitative results have implications for the privacy of users' traces, and may serve as a guideline for future policies regarding the management and publication of mobility data. Apostolos Pyrgelis, Nicolas Kourtellis, Ilias Leontiadis, Joan Serrà, Claudio Soriente |
IEEE BigData | 4 |
| 2018 | Language and Noise Transfer in Speech Enhancement Generative Adversarial NetworkabstractSpeech enhancement deep learning systems usually require large amounts of training data to operate in broad conditions or real applications. This makes the adaptability of those systems into new, low resource environments an important topic. In this work, we present the results of adapting a speech enhancement generative adversarial network by fine-tuning the generator with small amounts of data. We investigate the minimum requirements to obtain a stable behavior in terms of several objective metrics in two very different languages: Catalan and Korean. We also study the variability of test performance to unseen noise as a function of the amount of different types of noise available for training. Results show that adapting a pre-trained English model with 10 min of data already achieves a comparable performance to having two orders of magnitude more data. They also demonstrate the relative stability in test performance with respect to the number of training noise types. Santiago Pascual, Maruchan Park, Joan Serrà, Antonio Bonafonte, Kang-Hun Ahn |
ICASSP | 3 |
| 2018 | Overcoming Catastrophic Forgetting with Hard Attention to the TaskabstractCatastrophic forgetting occurs when a neural network loses the information learned in a previous task after training on subsequent tasks. This problem remains a hurdle for artificial intelligence systems with sequential learning capabilities. In this paper, we propose a task-based hard attention mechanism that preserves previous tasks’ information without affecting the current task’s learning. A hard attention mask is learned concurrently to every task, through stochastic gradient descent, and previous masks are exploited to condition such learning. We show that the proposed mechanism is effective for reducing catastrophic forgetting, cutting current rates by 45 to 80%. We also show that it is robust to different hyperparameter choices, and that it offers a number of monitoring capabilities. The approach features the possibility to control both the stability and compactness of the learned knowledge, which we believe makes it also attractive for online learning or network compression applications. Joan Serrà, Didac Suris, Marius Miron, Alexandros Karatzoglou |
ICML | 1 |
| 2018 | MobInsight: A Framework Using Semantic Neighborhood Features for Localized Interpretations of Urban MobilityabstractCollective urban mobility embodies the residents’ local insights on the city. Mobility practices of the residents are produced from their spatial choices , which involve various considerations such as the atmosphere of destinations, distance, past experiences, and preferences. The advances in mobile computing and the rise of geo-social platforms have provided the means for capturing the mobility practices; however, interpreting the residents’ insights is challenging due to the scale and complexity of an urban environment and its unique context. In this article, we present MobInsight, a framework for making localized interpretations of urban mobility that reflect various aspects of the urbanism. MobInsight extracts a rich set of neighborhood features through holistic semantic aggregation , and models the mobility between all-pairs of neighborhoods . We evaluate MobInsight with the mobility data of Barcelona and demonstrate diverse localized and semantically rich interpretations. Souneil Park, Joan Serrà, Enrique Frías-Martínez, Nuria Oliver |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2017 | Effect of acoustic conditions on algorithms to detect Parkinson's disease from speechabstractAutomatic detection of Parkinson's disease (PD) from speech is a basic step towards computer-aided tools supporting the diagnosis and monitoring of the disease. Although several methods have been proposed, their applicability to real-world situations is still unclear. In particular, the effect of acoustic conditions is not well understood. In this paper, the effects on the accuracy of five different methods to detect PD from speech are evaluated. Among the considered conditions, background noise produces the worst effect, while dynamic compression or some speech codecs can even have a marginal positive impact. We also consider, for the first time in this context, the problem of mismatches, i.e., when train/test acoustic conditions are different, and observe a high negative impact on all considered methods. Overall, this study is a step forward in performing a continuous monitoring of the neurological state of the patients in non-controlled acoustic conditions. Juan Camilo Vásquez-Correa, Joan Serrà, Juan Rafael Orozco-Arroyave, Jesús Francisco Vargas-Bonilla, Elmar Nöth |
ICASSP | 2 |
| 2017 | The Good, the Bad, and the KPIs: How to Combine Performance Metrics to Better Capture Underperforming Sectors in Mobile NetworksabstractMobile network operators collect a humongous amount of network measurements. Among those, sector Key Performance Indicators (KPIs) are used to monitor the radio access, i.e., the "last mile" of mobile networks. Thresholding mechanisms and synthetic combinations of KPIs are used to assess the network health, and rank sectors to identify the underperforming ones. It follows that the available monitoring methodologies heavily rely on the fine grained tuning of thresholds and weights, currently established through domain knowledge of both vendors and operators. In this paper, we study how to bridge sector KPIs to reflect Quality of Experience (QoE) groundtruth measurements, namely throughput, latency and video streaming stall events. We leverage one month of data collected in the operational network of mobile network operator serving more than 10 million subscribers. We extensively investigate up to which extent adopted methodologies efficiently capture QoE. Moreover, we challenge the current state of the art by presenting data-driven approaches based on Particle Swarm Optimization (PSO) metaheuristics and random forest regression algorithms, to better assess sector performance. Results show that the proposed methodologies outperforms state of the art solution improving the correlation with respect to the baseline by a factor of 3, and improving visibility on underperforming sectors. Our work opens new areas for research in monitoring solutions for enriching the quality and accuracy of the network performance indicators collected at the network edge. Ilias Leontiadis, Joan Serrà, Alessandro Finamore, Giorgos Dimopoulos, Konstantina Papagiannaki |
ICDE | 2 |
| 2017 | Hot or Not? Forecasting Cellular Network Hot Spots Using Sector Performance IndicatorsabstractTo manage and maintain large-scale cellular networks, operators need to know which sectors underperform at any given time. For this purpose, they use the so-called hot spot score, which is the result of a combination of multiple network measurements and reflects the instantaneous overall performance of individual sectors. While operators have a good understanding of the current performance of a network and its overall trend, forecasting the performance of each sector over time is a challenging task, as it is affected by both regular and non-regular events, triggered by human behavior and hardware failures. In this paper, we study the spatio-temporal patterns of the hot spot score and uncover its regularities. Based on our observations, we then explore the possibility to use recent measurements' history to predict future hot spots. To this end, we consider tree-based machine learning models, and study their performance as a function of time, amount of past data, and prediction horizon. Our results indicate that, compared to the best baseline, tree-based models can deliver up to 14% better forecasts for regular hot spots and 153% better forecasts for non-regular hot spots. The latter brings strong evidence that, for moderate horizons, forecasts can be made even for sectors exhibiting isolated, non-regular behavior. Overall, our work provides insight into the dynamics of cellular sectors and their predictability. It also paves the way for more proactive network operations with greater forecasting horizons. Joan Serrà, Ilias Leontiadis, Alexandros Karatzoglou, Konstantina Papagiannaki |
ICDE | 1 |
| 2017 | SEGAN: Speech Enhancement Generative Adversarial NetworkabstractCurrent speech enhancement techniques operate on the spectral domain and/or exploit some higher-level feature. The majority of them tackle a limited number of noise conditions and rely on first-order statistics. To circumvent these issues, deep networks are being increasingly used, thanks to their ability to learn complex functions from large example sets. In this work, we propose the use of generative adversarial networks for speech enhancement. In contrast to current techniques, we operate at the waveform level, training the model end-to-end, and incorporate 28 speakers and 40 different noise conditions into the same model, such that model parameters are shared across them. We evaluate the proposed model using an independent, unseen test set with two speakers and 20 alternative noise conditions. The enhanced samples confirm the viability of the proposed model, and both objective and subjective evaluations confirm the effectiveness of it. With that, we open the exploration of generative architectures for speech enhancement, which may progressively incorporate further speech-centric design choices to improve their performance. Santiago Pascual, Antonio Bonafonte, Joan Serrà |
INTERSPEECH | 3 |
| 2017 | Getting Deep Recommenders Fit: Bloom Embeddings for Sparse Binary Input/Output NetworksabstractRecommendation algorithms that incorporate techniques from deep learning are becoming increasingly popular. Due to the structure of the data coming from recommendation domains (i.e., one-hot-encoded vectors of item preferences), these algorithms tend to have large input and output dimensionalities that dominate their overall size. This makes them difficult to train, due to the limited memory of graphical processing units, and difficult to deploy on mobile devices with limited hardware. To address these difficulties, we propose Bloom embeddings, a compression technique that can be applied to the input and output of neural network models dealing with sparse high-dimensional binary-coded instances. Bloom embeddings are computationally efficient, and do not seriously compromise the accuracy of the model up to 1/5 compression ratios. In some cases, they even improve over the original accuracy, with relative increases up to 12%. We evaluate Bloom embeddings on 7 data sets and compare it against 4 alternative methods, obtaining favorable results. We also discuss a number of further advantages of Bloom embeddings, such as 'on-the-fly' constant-time operation, zero or marginal space requirements, training time speedups, or the fact that they do not require any change to the core model architecture or training configuration. Joan Serrà, Alexandros Karatzoglou |
RecSys | 1 |
| 2016 | Discovering rāga motifs by characterizing communities in networks of melodic patternsabstractRa̅ga motifs are the main building blocks of the melodic structures in Indian art music. Therefore, the discovery and characterization of such motifs is fundamental for the computational analysis of this music. We propose an approach for discovering ra̅ga motifs from audio music collections. First, we extract melodic patterns from a collection of 44 hours of audio comprising 160 recordings belonging to 10 ra̅gas. Next, we characterize these patterns by performing a network analysis, detecting non-overlapping communities, and exploiting the topological properties of the network to determine a similarity threshold. With that, we select a number of motif candidates that are representative of a ra̅ga, the ra̅ga motifs. For a formal evaluation we perform listening tests with 10 professional musicians. The results indicate that, on an average, the selected melodic phrases correspond to ra̅ga motifs with 85% positive ratings. This opens up the possibilities for many musically-meaningful computational tasks in Indian art music, including human-interpretable ra̅ga recognition, semantic-based music discovery, or pedagogical tools. Sankalp Gulati, Joan Serrà, Vignesh Ishwar, Xavier Serra |
ICASSP | 2 |
| 2016 | Phrase-based rĀga recognition using vector space modelingabstractAutomatic raga recognition is one of the fundamental computational tasks in Indian art music. Motivated by the way seasoned listeners identify ragas, we propose a raga recognition approach based on melodic phrases. Firstly, we extract melodic patterns from a collection of audio recordings in an unsupervised way. Next, we group similar patterns by exploiting complex networks concepts and techniques. Drawing an analogy to topic modeling in text classification, we then represent audio recordings using a vector space model. Finally, we employ a number of classification strategies to build a predictive model for raga recognition. To evaluate our approach, we compile a music collection of over 124 hours, comprising 480 recordings and 40 ragas. We obtain 70% accuracy with the full 40-raga collection, and up to 92% accuracy with its 10-raga subset. We show that phrase-based raga recognition is a successful strategy, on par with the state of the art, and sometimes outperforms it. A by-product of our approach, which arguably is as important as the task of raga recognition, is the identification of raga-phrases. These phrases can be used as a dictionary of semantically-meaningful melodic units for several computational tasks in Indian art music. Sankalp Gulati, Joan Serrà, Vignesh Ishwar, Sertan Sentürk, Xavier Serra |
ICASSP | 2 |
| 2016 | Ranking and significance of variable-length similarity-based time series motifs
Joan Serrà, Isabel Serra, Alvaro Corral, Josep Lluís Arcos |
Expert Syst. Appl. | 1 |
| 2016 | Particle swarm optimization for time series motif discovery
Joan Serrà, Josep Lluís Arcos |
Knowl. Based Syst. | 1 |
| 2015 | An evaluation of methodologies for melodic similarity in audio recordings of Indian art musicabstractWe perform a comparative evaluation of methodologies for computing similarity between short-time melodic fragments of audio recordings of Indian art music. We experiment with 560 different combinations of procedures and parameter values. These include the choices made for the sampling rate of the melody representation, pitch quantization levels, normalization techniques and distance measures. The dataset used for evaluation consists of 157 and 340 annotated melodic fragments of Carnatic and Hindustani music recordings, respectively. Our results indicate that melodic fragment similarity is particularly sensitive to distance measures and normalization techniques. Sampling rates do not have a significant impact for Hindustani music, but can significantly degrade the performance for Carnatic music. Overall, the performed evaluation provides a better understanding of the processing steps and parameter settings for melodic similarity in Indian art music. Importantly, it paves the way for developing unsupervised melodic pattern discovery approaches, whose evaluation is a challenging and, many times, ill-defined task. Sankalp Gulati, Joan Serrà, Xavier Serra |
ICASSP | 2 |
| 2015 | Analysis of the Impact of a Tag Recommendation System in a Real-World FolksonomyabstractCollaborative tagging systems have emerged as a successful solution for annotating contributed resources to online sharing platforms, facilitating searching, browsing, and organizing their contents. To aid users in the annotation process, several tag recommendation methods have been proposed. It has been repeatedly hypothesized that these methods should contribute to improving annotation quality and reducing the cost of the annotation process. It has been also hypothesized that these methods should contribute to the consolidation of the vocabulary of collaborative tagging systems. However, to date, no empirical and quantitative result supports these hypotheses. In this work, we deeply analyze the impact of a tag recommendation system in the folksonomy of Freesound, a real-world and large-scale online sound sharing platform. Our results suggest that tag recommendation effectively increases vocabulary sharing among users of the platform. In addition, tag recommendation is shown to contribute to the convergence of the vocabulary as well as to a partial increase in the quality of annotations. However, according to our analysis, the cost of the annotation process does not seem to be effectively reduced. Our work is relevant to increase our understanding about the nature of tag recommendation systems and points to future directions for the further development of those systems and their analysis. Frederic Font, Joan Serrà, Xavier Serra |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2014 | Class-based tag recommendation and user-based evaluation in online audio clip sharing
Frederic Font, Joan Serrà, Xavier Serra |
Knowl. Based Syst. | 2 |
| 2014 | An empirical evaluation of similarity measures for time series classification
Joan Serrà, Josep Lluís Arcos |
Knowl. Based Syst. | 1 |
| 2014 | Unsupervised Music Structure Annotation by Time Series Structure Features and Segment SimilarityabstractAutomatically inferring the structural properties of raw multimedia documents is essential in today's digitized society. Given its hierarchical and multi-faceted organization, musical pieces represent a challenge for current computational systems. In this article, we present a novel approach to music structure annotation based on the combination of structure features with time series similarity. Structure features encapsulate both local and global properties of a time series, and allow us to detect boundaries between homogeneous, novel, or repeated segments. Time series similarity is used to identify equivalent segments, corresponding to musically meaningful parts. Extensive tests with a total of five benchmark music collections and seven different human annotations show that the proposed approach is robust to different ground truth choices and parameter settings. Moreover, we see that it outperforms previous approaches evaluated under the same framework. Joan Serrà, Meinard Müller, Peter Grosche, Josep Lluís Arcos |
IEEE Trans. Multim. | 1 |
| 2013 | Towards cover group thumbnailingabstractIn this paper we investigate whether we can extract the commonalities shared by a group of cover songs or versions of the same musical piece. As a main contribution, we introduce the concept of cover group thumbnail, which is the most representative, essential subsequence for an entire group of versions. Opposed to previous approaches, we jointly consider all versions of a given song to compute a single cover group template, which then shows a high degree of robustness against version-specific aspects. To compute such a template, we introduce a modification of a recent audio thumbnailing technique. To evaluate the reliability of our conceptual contribution, we consider the task of template-based version identification, where we show comparable accuracies to existing systems. Peter Grosche, Meinard Müller, Joan Serrà |
ACM Multimedia | 3 |
| 2013 | Folksonomy-Based Tag Recommendation for Collaborative Tagging SystemsabstractCollaborative tagging has emerged as a common solution for labelling and organising online digital content. However, collaborative tagging systems typically suffer from a number of issues such as tag scarcity or ambiguous labelling. As a result, the organisation and browsing of tagged content is far from being optimal. In this work the authors present a general scheme for building a folksonomy-based tag recommendation system to help users tagging online content resources. Based on this general scheme, the authorse describe eight tag recommendation methods and extensively evaluate them with data coming from two real-world large-scale datasets of tagged images and sound clips. Their results show that the proposed methods can effectively recommend relevant tags, given a set of input tags and tag co-occurrence information. Moreover, the authors show how novel strategies for selecting the appropriate number of tags to be recommended can significantly improve methods performances. Approaches such as the one presented here can be useful to obtain more comprehensive and coherent descriptions of tagged resources, thus allowing a better organisation, browsing and reuse of online content. Moreover, they can increase the value of folksonomies as reliable sources for knowledge-mining. Frederic Font, Joan Serrà, Xavier Serra |
Int. J. Semantic Web Inf. Syst. | 2 |
| 2012 | Unsupervised Detection of Music Boundaries by Time Series Structure FeaturesabstractLocating boundaries between coherent and/or repetitive segments of a time series is a challenging problem pervading many scientific domains. In this paper we propose an unsupervised method for boundary detection, combining three basic principles: novelty, homogeneity, and repetition. In particular, the method uses what we call structure features, a representation encapsulating both local and global properties of a time series. We demonstrate the usefulness of our approach in detecting music structure boundaries, a task that has received much attention in recent years and for which exist several benchmark datasets and publicly available annotations. We find our method to significantly outperform the best accuracies published so far. Importantly, our boundary approach is generic, thus being applicable to a wide range of time series beyond the music and audio domains. Joan Serrà, Meinard Müller, Peter Grosche, Josep Lluís Arcos |
AAAI | 1 |
| 2012 | A Competitive Measure to Assess the Similarity between Two Time Series
Joan Serrà, Josep Lluís Arcos |
ICCBR | 1 |
| 2012 | Characterization and exploitation of community structure in cover song networks
Joan Serrà, Massimiliano Zanin, Perfecto Herrera, Xavier Serra |
Pattern Recognit. Lett. | 1 |
| 2012 | Predictability of Music Descriptor Time Series and its Application to Cover Song DetectionabstractIntuitively, music has both predictable and unpredictable components. In this paper, we assess this qualitative statement in a quantitative way using common time series models fitted to state-of-the-art music descriptors. These descriptors cover different musical facets and are extracted from a large collection of real audio recordings comprising a variety of musical genres. Our findings show that music descriptor time series exhibit a certain predictability not only for short time intervals, but also for mid-term and relatively long intervals. This fact is observed independently of the descriptor, musical facet and time series model we consider. Moreover, we show that our findings are not only of theoretical relevance but can also have practical impact. To this end we demonstrate that music predictability at relatively long time intervals can be exploited in a real-world application, namely the automatic identification of cover songs (i.e., different renditions or versions of the same musical piece). Importantly, this prediction strategy yields a parameter-free approach for cover song identification that is substantially faster, allows for reduced computational storage and still maintains highly competitive accuracies when compared to state-of-the-art systems. Joan Serrà, Holger Kantz, Xavier Serra, Ralph G. Andrzejak |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | Nonlinear audio recurrence analysis with application to genre classificationabstractIn this paper we apply nonlinear signal analysis to a music information retrieval task. More concretely, we apply the concept of recurrence plots and recurrence histograms to extract information from music audio frames. We evaluate the effectiveness of this approach with a typical genre classification framework and compare it against a baseline obtained from standard spectrum-based descriptors. The accuracy reached by the histogram-based descriptors alone does not surpass the one achieved by the spectral-based descriptors. However, we show that the combination of both descriptor sources results in consistent improvements up to 5 absolute percent points. This high lights the potential of nonlinear signal analysis for quantitative music description. In particular, it suggests that the information resulting from this approach is complementary to the information obtained through tile commonly used spectral representation. Joan Serrà, Carlos A. de los Santos, Ralph G. Andrzejak |
ICASSP | 1 |
| 2011 | Unifying Low-Level and High-Level Music Similarity MeasuresabstractMeasuring music similarity is essential for multimedia retrieval. For music items, this task can be regarded as obtaining a suitable distance measurement between songs defined on a certain feature space. In this paper, we propose three of such distance measures based on the audio content: first, a low-level measure based on tempo-related description; second, a high-level semantic measure based on the inference of different musical dimensions by support vector machines. These dimensions include genre, culture, moods, instruments, rhythm, and tempo annotations. Third, a hybrid measure which combines the above-mentioned distance measures with two existing low-level measures: a Euclidean distance based on principal component analysis of timbral, temporal, and tonal descriptors, and a timbral distance based on single Gaussian Mel-frequency cepstral coefficient (MFCC) modeling. We evaluate our proposed measures against a number of baseline measures. We do this objectively based on a comprehensive set of music collections, and subjectively based on listeners' ratings. Results show that the proposed methods achieve accuracies comparable to the baseline approaches in the case of the tempo and classifier-based measures. The highest accuracies are obtained by the hybrid distance. Furthermore, the proposed classifier-based approach opens up the possibility to explore distance measures that are based on semantic notions. Dmitry Bogdanov, Joan Serrà, Nicolas Wack, Perfecto Herrera, Xavier Serra |
IEEE Trans. Multim. | 2 |
| 2010 | Indexing music by mood: design and integration of an automatic content-based annotator
Cyril Laurier, Owen Meyers, Joan Serrà, Martin Blech, Perfecto Herrera, Xavier Serra |
Multim. Tools Appl. | 3 |
| 2009 | From Low-Level to High-Level: Comparative Study of Music Similarity MeasuresabstractStudying the ways to recommend music to a user is a central task within the music information research community. From a content-based point of view, this task can be regarded as obtaining a suitable distance measurement between songs defined on a certain feature space. We propose two such distance measures. First, a low-level measure based on tempo-related aspects, and second, a high-level semantic measure based on regression by support vector machines of different groups of musical dimensions such as genre and culture, moods and instruments, or rhythm and tempo. We evaluate these distance measures against a number of state-of-the-art measures objectively, based on 17 ground truth musical collections, and subjectively, based on 12 listeners’ ratings. Results show that, in spite of being conceptually different, the proposed methods achieve comparable or even higher performance than the considered baseline approaches. Furthermore, they open up the possibility to explore distance metrics that are based on truly semantic notions. Dmitry Bogdanov, Joan Serrà, Nicolas Wack, Perfecto Herrera |
ISM | 2 |
| 2008 | Audio cover song identification based on tonal sequence alignmentabstractNowadays, the term cover song (or simply cover) can mean any new version, performance, rendition, or recording of a previously recorded track. Cover song identification is a task that has received increased popularity in the Music Information Retrieval (MIR) community in recent years, as it provides a direct and objective way for evaluating music similarity. In this paper, we propose a new method for determining the similarity between tonal sequences and, therefore, for identifying cover songs. This is based on a novel chroma similarity measure, and on a newly developed dynamic programming local alignment technique. Results confirm that the performance of the proposed system is significantly superior to other state-of-the-art approaches (more than 57% better). Joan Serrà, Emilia Gómez |
ICASSP | 1 |
| 2008 | Chroma Binary Similarity and Local Alignment Applied to Cover Song IdentificationabstractWe present a new technique for audio signal comparison based on tonal subsequence alignment and its application to detect cover versions (i.e., different performances of the same underlying musical piece). Cover song identification is a task whose popularity has increased in the music information retrieval (MIR) community along in the past, as it provides a direct and objective way to evaluate music similarity algorithms. This paper first presents a series of experiments carried out with two state-of-the-art methods for cover song identification. We have studied several components of these (such as chroma resolution and similarity, transposition, beat tracking or dynamic time warping constraints), in order to discover which characteristics would be desirable for a competitive cover song identifier. After analyzing many cross-validated results, the importance of these characteristics is discussed, and the best performing ones are finally applied to the newly proposed method. Multiple evaluations of this one confirm a large increase in identification accuracy when comparing it with alternative state-of-the-art approaches. Joan Serrà, Emilia Gómez, Perfecto Herrera, Xavier Serra |
IEEE Trans. Speech Audio Process. | 1 |