VLDB 2026 Research / reviewers in the wild / expert
Johan Pauwels
dblp:21/4933
· DBLP profile ↗
19ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-5805-7144ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sparse Contrastive Learning for Content-Based Cold Item RecommendationabstractItem cold-start is a pervasive challenge for collaborative filtering (CF) recommender systems. Existing methods often train cold-start models by mapping auxiliary item content, such as images or text descriptions, into the embedding space of a CF model. However, such approaches can be limited by the fundamental information gap between CF signals and content features. In this work, we propose to avoid this limitation with purely content-based modeling of cold items, i.e. without alignment with CF user or item embeddings. We instead frame cold-start prediction in terms of item-item similarity, training a content encoder to project into a latent space where similarity correlates with user preferences. We define our training objective as a sparse generalization of sampled softmax loss with the α-entmax family of activation functions, which allows for sharper estimation of item relevance by zeroing gradients for uninformative negatives. We then describe how this Sampled Entmax for Cold-start (SEMCo) training regime can be extended via knowledge distillation, and show that it outperforms existing cold-start methods and standard sampled softmax in ranking accuracy. We also discuss the advantages of purely content-based modeling, particularly in terms of equity of item outcomes. Gregor Meehan, Johan Pauwels |
SIGIR | 2 |
| 2026 | Leveraging Artist Catalogs for Cold-Start Music RecommendationabstractThe item cold-start problem poses a fundamental challenge for music recommendation: newly added tracks lack the interaction history that collaborative filtering (CF) requires. Existing approaches often address this problem by learning mappings from content features such as audio, text, and metadata to the CF latent space. However, previous works either omit artist information or treat it as just another input modality, missing the fundamental hierarchy of artists and items. Since most new tracks come from artists with previous history available, we frame cold-start track recommendation as 'semi-cold' by leveraging the rich collaborative signal that exists at the artist level. We show that artist-aware methods can more than double Recall and NDCG compared to content-only baselines, and propose ACARec, an attention-based architecture that generates CF embeddings for new tracks by attending over the artist's existing catalog. We show that our approach has notable advantages in predicting user preferences for new tracks, especially for new artist discovery and more accurate estimation of cold item popularity. Yan-Martin Tamm, Gregor Meehan, Vojtech Nekl, Vojtech Vancura, Rodrigo Alves, Johan Pauwels, Anna Aljanaki |
UMAP | 6 |
| 2025 | Predicting Moral Values in Lyrics Through AudioabstractThis paper introduces the task of music morality recognition-predicting moral values in song lyrics using only audio features, which can be considered a form of music tagging with a set of new and well-defined tags grounded in social and cultural psychology research. Unlike previous research focused on lyrics analysis alone, this approach examines how musical elements correlate with moral content in associated lyrics. We used human-annotated lyrics and a set of experiments with XGBoost classifiers to establish a baseline recognition performance. Despite working with small and imbalanced data, we found audio features often outperformed a state of the art language model fine-tuned to detect moral content in lyrics, with some moral values being more reliably predicted than others in line with related work. SAGE and SHAP analyses revealed that specific timbral, harmonic, and melodic features playa prominent role in audio-lyrics moral associations. These findings advance our understanding of musical semantics and have potential applications in music and multimedia recommender systems and healthcare interventions. We provide a public repository containing all code and data used in this study. Charalampos Saitis, Ben Heyderman, Vjosa Preniqi, Kyriaki Kalimeri, Johan Pauwels |
CBMI | 5 |
| 2025 | Evaluating Contrastive Methodologies for Music Representation Learning Using Playlist DataabstractRecent research shows that weakly supervised contrastive pre-training holds significant promise in learning improved representations of musical audio. Several such works use metadata (e.g. artist names or genre tags) or consumption data (e.g. playlists or user listening history) for cross-modal supervision in their representation learning pipeline. However, methodological differences inhibit direct comparison of the results reported in these studies. In this work, we implement six of these contrastive pre-training regimes under a common framework based around playlist data, benchmarking against from-scratch and self-supervised baselines and systematically evaluating performance in downstream music tagging and playlist continuation. Furthermore, we combine existing methods to produce a novel hybrid approach which shows consistently strong performance by leveraging multiple data modes. Finally, we demonstrate that mixup data augmentation, previously only used in the self-supervised scenario, also has significant downstream benefits in weakly supervised music representation learning. Gregor Meehan, Johan Pauwels |
ICASSP | 2 |
| 2025 | Learning Music Audio Representations With Limited DataabstractLarge deep-learning models for music, including those focused on learning general-purpose music audio representations, are often assumed to require substantial training data to achieve high performance. If true, this would pose challenges in scenarios where audio data or annotations are scarce, such as for underrepresented music traditions, non-popular genres, and personalized music creation and listening. Understanding how these models behave in limited-data scenarios could be crucial for developing techniques to tackle them.In this work, we investigate the behavior of several music audio representation models under limited-data learning regimes. We consider music models with various architectures, training paradigms, and input durations, and train them on data collections ranging from 5 to 8,000 minutes long. We evaluate the learned representations on various music information retrieval tasks and analyze their robustness to noise. We show that, under certain conditions, representations from limited-data and even random models perform comparatively to ones from large-dataset models, though handcrafted features outperform all learned representations in some tasks. Christos Plachouras, Emmanouil Benetos, Johan Pauwels |
ICASSP | 3 |
| 2025 | Position Paper: Towards a Unified Representation Evaluation Framework Beyond Downstream TasksabstractDownstream probing has been the dominant method for evaluating model representations, an important process given the increasing prominence of self-supervised learning and foundation models. However, downstream probing primarily assesses the availability of task-relevant information in the model’s latent space, overlooking attributes such as equivariance, invariance, and disentanglement, which contribute to the interpretability, adaptability, and utility of representations in real-world applications. While some attempts have been made to measure these qualities in representations, no unified evaluation framework with modular, generalizable, and interpretable metrics exists.In this paper, we argue for the importance of representation evaluation beyond downstream probing. We introduce a standardized protocol to quantify informativeness, equivariance, invariance, and disentanglement of factors of variation in model representations. We use it to evaluate representations from a variety of models in the image and speech domains using different architectures and pretraining approaches on identified controllable factors of variation. We find that representations from models with similar downstream performance can behave substantially differently with regard to these attributes. This hints that the respective mechanisms underlying their downstream performance are functionally different, prompting new research directions to understand and improve representations. Christos Plachouras, Julien Guinot, György Fazekas, Elio Quinton, Emmanouil Benetos, Johan Pauwels |
IJCNN | 6 |
| 2025 | On Inherited Popularity Bias in Cold-Start Item RecommendationabstractCollaborative filtering (CF) recommender systems struggle with making predictions on unseen, or 'cold', items. Systems designed to address this challenge are often trained with supervision from warm CF models in order to leverage collaborative and content information from the available interaction data. However, since they learn to replicate the behavior of CF methods, cold-start models may therefore also learn to imitate their predictive biases. In this paper, we show that cold-start systems can inherit popularity bias, a common cause of recommender system unfairness arising when CF models overfit to more popular items, thereby maximizing user-oriented accuracy but neglecting rarer items. We demonstrate that cold-start recommenders not only mirror the popularity biases of warm models, but are in fact affected more severely: because they cannot infer popularity from interaction data, they instead attempt to estimate it based solely on content features. This leads to significant over-prediction of certain cold items with similar content to popular warm items, even if their ground truth popularity is very low. Through experiments on three multimedia datasets, we analyze the impact of this behavior on three generative cold-start methods. We then describe a simple post-processing bias mitigation method that, by using embedding magnitude as a proxy for predicted popularity, can produce more balanced recommendations with limited harm to user-oriented cold-start accuracy. Gregor Meehan, Johan Pauwels |
RecSys | 2 |
| 2025 | Temporal Considerations in DJ Mix Information Retrieval and Generation (Short Paper)
Alexander J. Williams, Gregor Meehan, Stefan Lattner, Johan Pauwels, Mathieu Barthet |
TIME | 4 |
| 2024 | Musician-AI partnership mediated by emotionally-aware smart musical instrumentsabstractThe integration of emotion recognition capabilities within musical instruments can spur the emergence of novel art formats and services for musicians. This paper proposes the concept of emotionally-aware smart musical instruments, a class of musical devices embedding an artificial intelligence agent able to recognize the emotion contained in the musical signal. This spurs the emergence of novel services for musicians. Two prototypes of emotionally-aware smart piano and smart electric guitar were created, which embedded a recognition method for happiness, sadness, relaxation, aggressiveness and combination thereof. A user study, conducted with eleven pianists and eleven electric guitarists, revealed the strengths and limitations of the developed technology. On average musicians appreciated the proposed concept, who found its value in various musical activities. Most of participants tended to justify the system with respect to erroneous or partially erroneous classifications of the emotions they expressed, reporting to understand the reasons why a given output was produced. Some participants even seemed to trust more the system than their own judgments. Conversely, other participants requested to improve the accuracy, reliability and explainability of the system in order to achieve a higher degree of partnership with it. Our results suggest that, while desirable, perfect prediction of the intended emotion is not an absolute requirement for music emotion recognition to be useful in the construction of smart musical instruments. Luca Turchet, Domenico Stefani, Johan Pauwels |
Int. J. Hum. Comput. Stud. | 3 |
| 2023 | On the Relevance of the Differences Between HRTF Measurement Setups for Machine LearningabstractAs spatial audio is enjoying a surge in popularity, data-driven machine learning techniques that have been proven successful in other domains are increasingly used to process head-related transfer function measurements. However, these techniques require much data, whereas the existing datasets are ranging from tens to the low hundreds of datapoints. It therefore becomes attractive to combine multiple of these datasets, although they are measured under different conditions. In this paper, we first establish the common ground between a number of datasets, then we investigate potential pitfalls of mixing datasets. We perform a simple experiment to test the relevance of the remaining differences between datasets when applying machine learning techniques. Finally, we pinpoint the most relevant differences. Johan Pauwels, Lorenzo Picinali |
ICASSP | 1 |
| 2023 | "Give me happy pop songs in C major and with a fast tempo": A vocal assistant for content-based queries to online music repositoriesabstractThis paper presents an Internet of Musical Things system devised to support recreational music-making, improvisation, composition, and music learning via vocal queries to an online music repository. The system involves a commercial voice-based interface and the Jamendo cloud-based repository of Creative Commons music content. Thanks to the system the user can query the Jamendo music repository by six content-based features and each combination thereof: mood, genre, tempo, chords, key and tuning. Such queries differ from the conventional methods for music retrieval, which are based on the piece’s title and the artist’s name. These features were identified following a survey with 112 musicians, which preliminary validated the concept underlying the proposed system. A user study with 20 musicians showed that the system was deemed usable, able to provide a satisfactory user experience, and useful in a variety of musical activities. Differences in the participants’ needs were identified, which highlighted the need for personalization mechanisms based on the expertise level of the user. Importantly, the system was seen as a concrete solution to physical encumbrances that arise from the concurrent use of the instrument and devices providing interactive media resources. Finally, the system offers benefits to visually-impaired musicians. Luca Turchet, Carlo Zanotto, Johan Pauwels |
Int. J. Hum. Comput. Stud. | 3 |
| 2022 | Music Emotion Recognition: Intention of Composers-Performers Versus Perception of Musicians, Non-Musicians, and Listening MachinesabstractThis paper investigates to which extent state of the art machine learning methods are effective in classifying emotions in the context of individual musical instruments, and how their performances compare with musically trained and untrained listeners. To address these questions we created a novel dataset of 391 classical and acoustic guitar excerpts annotated along four emotions (aggressiveness, relaxation, happiness and sadness) with three emotion intensity levels (low, medium, high), according to the intended emotion of 30 professional guitarists acting as both composers and performers. A first experiment investigated listeners’ perception involving 8 professional guitarists and 8 non-musicians. Results showed that the emotions intended by a composer-performer are not always well recognized by listeners, and in general not with the same intensity. Listeners’ identification accuracy was proportional to the intensity with which an emotion was expressed. Emotions were better recognized by musicians than by listeners without musical background. Such differences between the two groups were found for different intensity levels of the intended emotions. A second experiment investigated machine listening performance based on a transfer learning method. To compare machine and human identification accuracies fairly, we derived a fifth, “ambivalent” category from the machine listening output categories (i.e., excerpts rated with more than one predominant emotion). Results showed that the machine perception of emotions matched or even exceeded musicians’ performance for all emotions except “relaxation”. The differences between the intended and human-perceived emotions, as well as those due to musical training, suggest that a device or application involving a music emotion recognition system should take into account the characteristics of the users (in particular their musical expertise) as well as their roles (e.g., composers, performers, listeners). For developers this translates into the use of datasets annotated by different categories of annotators, whose role and musical expertise will match the characteristics of the end users. Such results are particularly relevant to the creation of emotionally-aware smart musical instruments. Luca Turchet, Johan Pauwels |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | The WASABI Dataset: Cultural, Lyrics and Audio Analysis Metadata About 2 Million Popular Commercially Released Songs
Michel Buffa, Elena Cabrio, Michael Fell, Fabien Gandon, Alain Giboin, Romain Hennequin, Franck Michel, Johan Pauwels, Guillaume Pellerin, Maroua Tikat, Marco Winckler |
ESWC | 8 |
| 2020 | Cloud-smart Musical Instrument Interactions: Querying a Large Music Collection with a Smart GuitarabstractLarge online music databases under Creative Commons licenses are rarely recorded by well-known artists, therefore conventional metadata-based search is insufficient in their adaptation to instrument players’ needs. The emerging class of smart musical instruments (SMIs) can address this challenge. Thanks to direct internet connectivity and embedded processing, SMIs can send requests to repositories and reproduce the response for improvisation, composition, or learning purposes. We present a smart guitar prototype that allows retrieving songs from large online music databases using criteria different from conventional music search, which were derived from interviewing 30 guitar players. We investigate three interaction methods coupled with four search criteria (tempo, chords, key and tuning) exploiting intelligent capabilities in the instrument: (i) keywords-based retrieval using an embedded touchscreen; (ii) cloud-computing where recorded content is transmitted to a server that extracts relevant audio features; (iii) edge-computing where the guitar detects audio features and sends the request directly. Overall, the evaluation of these methods with beginner, intermediate, and expert players showed a strong appreciation for the direct connectivity of the instrument with an online database and the approach to the search based on the actual musical content rather than conventional textual criteria, such as song title or artist name. Luca Turchet, Johan Pauwels, Carlo Fischione, György Fazekas |
ACM Trans. Internet Things | 2 |
| 2017 | Improved template based chord recognition using the CRP featureabstractThe task of chord recognition in music signals is often based upon pattern matching in chromagrams. Many variants of chroma exist and quality of chord recognition is related to the feature employed. Chroma Reduced Pitch (CRP) features are interesting in this context as they were designed to improve timbre invariance for the purpose of query retrieval. Their reapplication to chord recognition, however, has not been successful in previous studies. We consider that the default parametrisation of CRP attenuates some tonal information, as well as timbral, and consider alternatives to this default. We also provide a variant of a recently proposed compositional chroma feature, adapted for music pieces, rather than one instrument. Experiments described show improved results compared to existing features. Ken O'Hanlon, Sebastian Ewert, Johan Pauwels, Mark B. Sandler |
ICASSP | 3 |
| 2013 | Evaluating automatically estimated chord sequencesabstractIn this paper, we perform an in-depth evaluation of a large number of algorithms for chord estimation that have been submitted to the MIREX competitions in 2010, 2011 and 2012. Therefore we first present a rigorous scheme to describe evaluation methods in a sound, unambiguous way that extends previous work specifically to take into account the large variance in chord estimation vocabularies and to perform evaluations on select sets of chords. Then we take a look at the evaluation metrics used so far and propose some alternative ones. Finally, we use these different methods to get a deeper insight into the strengths of each of the competing algorithms and show that the choice of evaluation measure greatly influences the ranking. Johan Pauwels, Geoffroy Peeters |
ICASSP | 1 |
| 2013 | Segmenting music through the joint estimation of keys, chords and structural boundariesabstractIn this paper, we introduce a new approach to music structure segmentation that is based on the joint estimation of structural segments, keys and chords in one probabilistic framework. More precisely, the boundaries of a structure segment are determined by detecting key changes and by utilizing the difference in prior probability of chord transitions according to their position in a structural segment. In contrast to many of the recent approaches to structural segmentation, this system does not work with self-similarity matrices, although it has been designed to integrate this kind of approach into the framework at a later stage. However, just the current version of the system, using only the estimated harmony, is already producing encouraging results, especially with respect to the precise localization of the boundaries. Johan Pauwels, Geoffroy Peeters |
ACM Multimedia | 1 |
| 2011 | Improving the key extraction performance of a simultaneous local key and chord estimation systemabstractSinificant improvements of a previously developed key and chord extraction system are proposed. The major improvement is the introduction of a separate acoustic model, designed to verify local key hypotheses. The conducted experimental evaluation shows that the presented system improves the state of the art in local key estimation. Our experimental study further demonstrates that the chord estimation performance is already quite robust, whereas the key estimation performance still happens to be sensitive to a number of factors. In particular, we present figures that illustrate the significant impact of the embedded musicological model and the duration of the processed excerpt on the key estimation accuracy. Johan Pauwels, Jean-Pierre Martens, Marc Leman |
ICME | 1 |
| 2008 | A novel chroma representation of polyphonic music based on multiple pitch tracking techniquesabstractIt is common practice to map the frequency content of music onto a chroma representation, but there exist many different schemes for constructing such a representation. In this paper, a new scheme is proposed. It comprises a detection of salient frequencies, a conversion of salient frequencies to notes, a psychophysically motivated weighting of harmonics in support of a note, a restriction of harmonic relations between different notes and a restriction of the deviations from a predefined pitch scale (e.g. the equally tempered western scale). A large-scale experimental evaluation has confirmed that the novel chroma representation more closely matches manual chord labels than the representations generated by six other tested schemes. Therefore, the new chroma representation is expected to improve applications such as song similarity matching and chord detection and labeling. Matthias Varewyck, Johan Pauwels, Jean-Pierre Martens |
ACM Multimedia | 2 |