VLDB 2026 Research / reviewers in the wild / expert
Geoffroy Peeters
dblp:45/1754
· DBLP profile ↗
45ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0001-5255-3019ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AudioCAN: Enhanced Few-Shot Audio Classification via Energy-Guided Temporal Cross AttentionabstractDespite the growing research interest and practical potential of few-shot audio classification, its efficacy remains limited by the temporal sparsity of target sound events and the presence of non-target acoustic interference within audio samples. In this work, we leverage the cross attention mechanism to address these challenges. While the Cross Attention Network (CAN) has demonstrated superior performance in few-shot image classification by highlighting semantically relevant regions between support and query features, we propose AudioCAN by extending CAN to audio domain through two-fold, audio-specific modifications. In particular, we first reformulate the original 2D spatial cross attention to 1D temporal version to prioritize the time frames containing discriminative acoustic information, thereby mitigating interference from non-target sounds. Secondly, we introduce an energy-guided masking strategy to synthesize a pseudo-query set from the support set via data augmentation, which then serves as a guidance for optimizing the 1D temporal cross attention during training. Experimental results on several few-shot audio classification benchmarks demonstrate that AudioCAN achieves state-of-the-art performance in 5-way-1-shot settings while remaining highly competitive in 5-way-5-shot configurations. Xuanyu Zhuang, Geoffroy Peeters, Gaël Richard |
IEEE Signal Process. Lett. | 2 |
| 2025 | Twenty-Five Years of MIR Research: Achievements, Practices, Evaluations, and Future ChallengesabstractIn this paper, we trace the evolution of Music Information Retrieval (MIR) over the past 25 years. While MIR gathers all kinds of research related to music informatics, a large part of it focuses on signal processing techniques for music data, fostering a close relationship with the IEEE Audio and Acoustic Signal Processing Technical Commitee. In this paper, we reflect the main research achievements of MIR along the three EDICS related to music analysis, processing and generation. We then review a set of successful practices that fuel the rapid development of MIR research. One practice is the annual research benchmark, the Music Information Retrieval Evaluation eXchange, where participants compete on a set of research tasks. Another practice is the pursuit of reproducible and open research. The active engagement with industry research and products is another key factor for achieving large societal impacts and motivating younger generations of students to join the field. Last but not the least, the commitment to diversity, equity and inclusion ensures MIR to be a vibrant and open community where various ideas, methodologies, and career pathways collide. We finish by providing future challenges MIR will have to face. Geoffroy Peeters, Zafar Rafii, Magdalena Fuentes, Zhiyao Duan, Emmanouil Benetos, Juhan Nam, Yuki Mitsufuji |
ICASSP | 1 |
| 2025 | Masked Latent Prediction and Classification for Self-Supervised Audio Representation LearningabstractRecently, self-supervised learning methods based on masked latent prediction have proven to encode input data into powerful representations. However, during training, the learned latent space can be further transformed to extract higher-level information that could be more suited for down-stream classification tasks. Therefore, we propose a new method: MAsked latenT Prediction And Classification (MATPAC), which is trained with two pretext tasks solved jointly. As in previous work, the first pretext task is a masked latent prediction task, ensuring a robust input representation in the latent space. The second one is unsupervised classification, which utilises the latent representations of the first pretext task to match probability distributions between a teacher and a student. We validate the MATPAC method by comparing it to other state-of-the-art proposals and conducting ablations studies. MATPAC reaches state-of-the-art self-supervised learning results on reference audio classification datasets such as OpenMIC, GTZAN, ESC-50 and US8K and outperforms comparable supervised methods’ results for musical auto-tagging on Magna-tag-a-tune. Aurian Quelennec, Pierre Chouteau, Geoffroy Peeters, Slim Essid |
ICASSP | 3 |
| 2025 | Zero-shot Musical Stem Retrieval with Joint-Embedding Predictive ArchitecturesabstractIn this paper, we tackle the task of musical stem retrieval. Given a musical mix, it consists in retrieving a stem that would fit with it, i.e., that would sound pleasant if played together. To do so, we introduce a new method based on Joint-Embedding Predictive Architectures, where an encoder and a predictor are jointly trained to produce latent representations of a context and predict latent representations of a target. In particular, we design our predictor to be conditioned on arbitrary instruments, enabling our model to perform zero-shot stem retrieval. In addition, we discover that pretraining the encoder using contrastive learning drastically improves the model’s performance.We validate the retrieval performances of our model using the MUSDB18 and MoisesDB datasets. We show that it significantly out-performs previous baselines on both datasets, showcasing its ability to support more or less precise (and possibly unseen) conditioning. We also evaluate the learned embeddings on a beat tracking task, demonstrating that they retain temporal structure and local information. Alain Riou, Antonin Gagneré, Gaëtan Hadjeres, Stefan Lattner, Geoffroy Peeters |
ICASSP | 5 |
| 2024 | Invariant Audio Prints for Music Indexing and AlignmentabstractThis work deals with music indexing and alignment using audio codes designed to be representative of the music content and robust to sound modifications. First, based on properties of the Fourier Transform and of the logarithm, highdimensional audio descriptors are designed. Then, a dimension reduction is learned with criteria based on sound discrimination and invariance to transformations. Finally, a binarization is computed to derive codes (integers). This last process allows a fast searching for large catalogs with a hash table, and a Hamming distance on codes makes possible the time alignment using an adapted “Dynamic Time Warping”. The contributions of this paper are tested for two different tasks. The goal of the first task is to identify the segments of music medleys with the audio indexing process, and to accurately find the corresponding original time positions. The goal of the second task is to measure the accuracy of the time-alignment with synthesized MIDI files, where the tempo continuously varies, and with modified pitches and instruments. Additionally, the audio indexing is also tested for these data, in order to exhibit some properties of the used audio prints. Rémi Mignot, Geoffroy Peeters |
CBMI | 2 |
| 2024 | Adapting Pitch-Based Self Supervised Learning Models for Tempo EstimationabstractTempo estimation is the task of estimating the periodicity of the dominant rhythm pulse of a music audio signal. It has therefore a close relationship with dominant pitch estimation. Recently, both tasks have been addressed in a Self-Supervised Learning (SSL) fashion so as to leverage unlabelled data for training. In this work, we study the applicability of two successful pitch-based SSL models, SPICE and PESTO, for the purpose of tempo estimation. Both successfully exploit Siamese networks with a pitch-shifting view generation between the two branches. To apply these models for tempo estimation, we represent the audio signal by the Constant-Q transform (CQT) of its onset-strength-function and adapt their view generation using time-stretching (instead of pitch shifting), which is efficiently implemented by shifting the CQT. In a large experiment, we show that simply adapting PESTO in this way yields superior results than the previous SSL approach to tempo estimation for most datasets used in the reference benchmark. Further, since PESTO is light-weight, requiring only a few training data, we study a new learning scheme where the downstream datasets are processed directly in a SSL fashion (without access to labels) showing that this is an interesting alternative further improving the performance for some datasets. Antonin Gagneré, Slim Essid, Geoffroy Peeters |
ICASSP | 3 |
| 2024 | Blind Estimation of Audio Effects Using an Auto-Encoder Approach and Differentiable Digital Signal ProcessingabstractBlind Estimation of Audio Effects (BE-AFX) aims at estimating the audio effects (AFXs) applied to an original, unprocessed audio sample solely based on the processed audio sample. To train such a system traditional approaches optimize a loss between ground truth and estimated AFX parameters. This involves knowing the exact implementation of the AFXs used for the process. In this work, we propose an alternative solution that eliminates the requirement for knowing this implementation. Instead, we introduce an auto-encoder approach, which optimizes an audio quality metric. We explore, suggest, and compare various implementations of commonly used mastering AFXs, using differential signal processing or neural approximations. Our findings demonstrate that our auto-encoder approach yields superior estimates of the audio quality produced by a chain of AFXs, compared to the traditional parameter-based approach, even if the latter provides a more accurate parameter estimation. Côme Peladeau, Geoffroy Peeters |
ICASSP | 2 |
| 2024 | On The Choice of the Optimal Temporal Support for Audio Classification with Pre-Trained EmbeddingsabstractCurrent state-of-the-art audio analysis systems rely on pre-trained embedding models, often used off-the-shelf as (frozen) feature extractors. Choosing the best one for a set of tasks is the subject of many recent publications. However, one aspect often overlooked in these works is the influence of the duration of audio input considered to extract an embedding, which we refer to as Temporal Support (TS). In this work, we study the influence of the TS for well-established or emerging pre-trained embeddings, chosen to represent different types of architectures and learning paradigms. We conduct this evaluation using both musical instrument and environmental sound datasets, namely OpenMIC, TAU Urban Acoustic Scenes 2020 Mobile, and ESC-50. We especially highlight that Audio Spectrogram Transformer-based systems (PaSST and BEATs) remain effective with smaller TS, which therefore allows for a drastic reduction in memory and computational cost. Moreover, we show that by choosing the optimal TS we reach competitive results across all tasks. In particular, we improve the state-of-the-art results on OpenMIC, using BEATs and PaSST without any fine-tuning. Aurian Quelennec, Michel Olvera, Geoffroy Peeters, Slim Essid |
ICASSP | 3 |
| 2024 | Unsupervised Harmonic Parameter Estimation Using Differentiable DSP and Spectral Optimal TransportabstractIn neural audio signal processing, pitch conditioning has been used to enhance the performance of synthesizers. However, jointly training pitch estimators and synthesizers is a challenge when using standard audio-to-audio reconstruction loss, leading to reliance on external pitch trackers. To address this issue, we propose using a spectral loss function inspired by optimal transportation theory that minimizes the displacement of spectral energy. We validate this approach through an unsupervised autoencoding task that fits a harmonic template to harmonic signals. We jointly estimate the fundamental frequency and amplitudes of harmonics using a lightweight encoder and reconstruct the signals using a differentiable harmonic synthesizer. The proposed approach offers a promising direction for improving unsupervised parameter estimation in neural audio applications. Bernardo Torres, Geoffroy Peeters, Gaël Richard |
ICASSP | 2 |
| 2023 | Cosmopolite Sound Monitoring (CoSMo): A Study of Urban Sound Event Detection Systems Generalizing to Multiple CitiesabstractMeasuring noise in cities and automatically identifying the corresponding sound sources are a crucial challenge for policymakers. Indeed, such information helps addressing noise pollution and improving the well-being of urban dwellers. In recent years, researchers have provided annotated datasets recorded in two major cities to foster the development of urban sound event detection (SED) systems. This paper presents an in-depth study of the behaviour of state-of-the-art SED systems well suited to our problem, combining three far-field real recordings datasets which can be used jointly during training. In our evaluation, we highlight the performance gaps existing between simple and hard recording examples based on the salience of sound events and the polyphony of the recordings. We provide new proximity annotations for this analysis. We evaluate the ability of urban SED systems to generalize across cities with varying degrees of training supervision. We show that such generalization is hindered mostly by the difficulties current urban SED systems have to detect sound events with low salience along with sound events in highly polyphonic soundscapes. Florian Angulo, Slim Essid, Geoffroy Peeters, Christophe Mietlicki |
ICASSP | 3 |
| 2023 | Learning Interpretable Filters In Wav-UNet For Speech EnhancementabstractDue to their performances, deep neural networks have emerged as a major method in nearly all modern audio processing applications. Deep neural networks can be used to estimate some parameters or hyperparameters of a model, or in some cases the entire model in an end-to-end fashion. Although deep learning can lead to state of the art performances, they also suffer from inherent weaknesses as they usually remain complex and non interpretable to a large extent. For instance, the internal filters used in each layers are chosen in an adhoc manner with only a loose relation with the nature of the processed signal. We propose in this paper an approach to learn interpretable filters within a specific neural architecture which allow to better understand the behaviour of the neural network and to reduce its complexity. We validate the approach on a task of speech enhancement and show that the gain in interpretability does not degrade the performance of the model. Félix Mathieu, Thomas Courtat, Gaël Richard, Geoffroy Peeters |
ICASSP | 4 |
| 2023 | Video-to-Music Recommendation Using Temporal Alignment of SegmentsabstractWe study cross-modal recommendation of musictracks to be used as soundtracks for videos. This problem is known as the music supervision task. We build on a self-supervised system that learns a content association between music and video. In addition to the adequacy of content, adequacy of structure is crucial in music supervision to obtain relevant recommendations. We propose a novel approach to significantly improve the system’s performance using structure-aware recommendation. The core idea is to consider not only the full audio-video clips, but rather shorter segments for training and inference. We find that using semantic segments and ranking the tracks according to sequence alignment costs significantly improves the results. We investigate the impact of different ranking metrics and segmentation methods. Laure Prétet, Gaël Richard, Clément Souchier, Geoffroy Peeters |
IEEE Trans. Multim. | 4 |
| 2022 | Phase Shifted Bedrosian Filterbank: An Interpretable Audio Front-End for Time-Domain Audio Source SeparationabstractThe use of a parameterized encoders or audio front-ends has shown promises in improving the interpretability of time domain single-channel source separation models such as Conv-TasNet. This type of filters also allows a potential reduction of the computational cost since larger encoder filters can be used. In this work, we propose to build a new parameterization of such encoder filter-bank which allows gaining interpretability while keeping flexibility. Based on the Hilbert transform and the Bedrosian theorem, we propose to build phase-shifted set of filters by modulating sinusoids through freely learned low pass filters. We show that the use of these filters allows to keep the same performances when using small filters and even improve them when using large filters. Félix Mathieu, Thomas Courtat, Gaël Richard, Geoffroy Peeters |
ICASSP | 4 |
| 2022 | Lyrics segmentation via bimodal text-audio representationabstractAbstract Song lyrics contain repeated patterns that have been proven to facilitate automated lyrics segmentation, with the final goal of detecting the building blocks (e.g., chorus, verse) of a song text. Our contribution in this article is twofold. First, we introduce a convolutional neural network (CNN)-based model that learns to segment the lyrics based on their repetitive text structure. We experiment with novel features to reveal different kinds of repetitions in the lyrics, for instance based on phonetical and syntactical properties. Second, using a novel corpus where the song text is synchronized to the audio of the song, we show that the text and audio modalities capture complementary structure of the lyrics and that combining both is beneficial for lyrics segmentation performance. For the purely text-based lyrics segmentation on a dataset of 103k lyrics, we achieve an F-score of 67.4%, improving on the state of the art (59.2% F-score). On the synchronized text–audio dataset of 4.8k songs, we show that the additional audio features improve segmentation performance to 75.3% F-score, significantly outperforming the purely text-based approaches. Michael Fell, Yaroslav Nechaev, Gabriel Meseguer-Brocal, Elena Cabrio, Fabien Gandon, Geoffroy Peeters |
Nat. Lang. Eng. | 6 |
| 2022 | Comparing Deep Models and Evaluation Strategies for Multi-Pitch Estimation in Music RecordingsabstractExtracting pitch information from music recordings is a challenging but important problem in music signal processing. Frame-wise transcription or multi-pitch estimation aims for detecting the simultaneous activity of pitches in polyphonic music recordings and has recently seen major improvements thanks to deep-learning techniques, with a variety of proposed model architectures. In this paper, we compare different architectures based on convolutional neural networks, the U-net structure, and self-attention components. We propose several modifications to these architectures including self-attention modules for skip connections, recurrent layers to replace the self-attention, and a multi-task strategy with simultaneous prediction of the degree of polyphony. We compare variants of these architectures in different sizes for multi-pitch estimation, focusing on Western classical music beyond the piano-solo scenario using the MusicNet and Schubert Winterreise datasets. Our experiments indicate that most architectures yield competitive results and that larger model variants seem to be beneficial. However, we find that these results substantially depend on randomization effects and the particular choice of the training–test split, which questions the claim of superiority for particular architectures given only small improvements. We therefore investigate the influence of dataset splits in the presence of several movements of a work cycle (cross-version evaluation) and propose a best-practice evaluation strategy for MusicNet, which weakens the influence of individual test tracks and suppresses overfitting to specific works and recording conditions. A final cross-dataset evaluation suggests that improvements on one specific dataset do not necessarily generalize to other scenarios, thus emphasizing the need for further high-quality multi-pitch datasets in order to reliably measure progress in music transcription tasks. Christof Weiß, Geoffroy Peeters |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | The Jazz Ontology: A semantic model and large-scale RDF repositories for jazzabstractJazz is a musical tradition that is just over 100 years old; unlike in other Western musical traditions, improvisation plays a central role in jazz. Modelling the domain of jazz poses some ontological challenges due to specificities in musical content and performance practice, such as band lineup fluidity and importance of short melodic patterns for improvisation. This paper presents the Jazz Ontology – a semantic model that addresses these challenges. Additionally, the model also describes workflows for annotating recordings with melody transcriptions and for pattern search. The Jazz Ontology incorporates existing standards and ontologies such as FRBR and the Music Ontology. The ontology has been assessed by examining how well it supports describing and merging existing datasets and whether it facilitates novel discoveries in a music browsing application. The utility of the ontology is also demonstrated in a novel framework for managing jazz related music information. This involves the population of the Jazz Ontology with the metadata from large scale audio and bibliographic corpora (the Jazz Encyclopedia and the Jazz Discography). The resulting RDF datasets were merged and linked to existing Linked Open Data resources. These datasets are publicly available and are driving an online application that is being used by jazz researchers and music lovers for the systematic study of jazz. Polina Proutskova, Daniel Wolff, György Fazekas, Klaus Frieler, Frank Höger, Olga Velichkina, Gabriel Solis, Tillman Weyde, Martin Pfleiderer, Hélène C. Crayencour, Geoffroy Peeters, Simon Dixon |
J. Web Semant. | 11 |
| 2021 | Cross-Modal Music-Video Recommendation: A Study of Design ChoicesabstractIn this work, we study music/video cross-modal recommendation, i.e. recommending a music track for a video or vice versa. We rely on a self-supervised learning paradigm to learn from a large amount of unlabelled data. We rely on a self-supervised learning paradigm to learn from a large amount of unlabelled data. More precisely, we jointly learn audio and video embeddings by using their co-occurrence in music-video clips. In this work, we build upon a recent video-music retrieval system (the VM-NET), which originally relies on an audio representation obtained by a set of statistics computed over handcrafted features. We demonstrate here that using audio representation learning such as the audio embeddings provided by the pre-trained MuSimNet, OpenL3, MusicCNN or by AudioSet, largely improves recommendations. We also validate the use of the cross-modal triplet loss originally proposed in the VM-NET compared to the binary cross-entropy loss commonly used in self-supervised learning. We perform all our experiments using the Music Video Dataset (MVD). Laure Prétet, Gaël Richard, Geoffroy Peeters |
IJCNN | 3 |
| 2020 | A Prototypical Triplet Loss for Cover DetectionabstractAutomatic cover detection - the task of finding in an audio dataset all covers of a query track - has long been a challenging theoretical problem in MIR community. It also became a practical need for music composers societies requiring to detect automatically if an audio excerpt embeds musical content belonging to their catalog. In a recent work, we addressed this problem with a convolutional neural network mapping each track's dominant melody to an embedding vector, and trained to minimize cover pairs distance in the embeddings space, while maximizing it for non-covers. We showed in particular that training this model with enough works having five or more covers yields state-of-the-art results. This however does not reflect the realistic use case, where music catalogs typically contain works with zero or at most one or two covers. We thus introduce here a new test set incorporating these constraints, and propose two contributions to improve our model's accuracy under these stricter conditions: we replace dominant melody with multi-pitch representation as input data, and describe a novel prototypical triplet loss designed to improve covers clustering. We show that these changes improve results significantly for two concrete use cases, large dataset lookup and live songs identification. Guillaume Doras, Geoffroy Peeters |
ICASSP | 2 |
| 2020 | Audio-Based Auto-Tagging With Contextual Tags for MusicabstractMusic listening context such as location or activity has been shown to greatly influence the users' musical tastes. In this work, we study the relationship between user context and audio content in order to enable context-aware music recommendation agnostic to user data. For that, we propose a semi-automatic procedure to collect track sets which leverages playlist titles as a proxy for context labelling. Using this, we create and release a dataset of ~50k tracks labelled with 15 different contexts. Then, we present benchmark classification results on the created dataset using an audio auto-tagging model. As the training and evaluation of these models are impacted by missing negative labels due to incomplete annotations, we propose a sample-level weighted cross entropy loss to account for the confidence in missing labels and show improved context prediction results. Karim M. Ibrahim, Jimena Royo-Letelier, Elena V. Epure, Geoffroy Peeters, Gaël Richard |
ICASSP | 4 |
| 2020 | Learning to Rank Music Tracks Using Triplet LossabstractMost music streaming services rely on automatic recommendation algorithms to exploit their large music catalogs. These algorithms aim at retrieving a ranked list of music tracks based on their similarity with a target music track. In this work, we propose a method for direct recommendation based on the audio content without explicitly tagging the music tracks. To that aim, we propose several strategies to perform triplet mining from ranked lists. We train a Convolutional Neural Network to learn the similarity via triplet loss. These different strategies are compared and validated on a large-scale experiment against an auto-tagging based approach. The results obtained highlight the efficiency of our system, especially when associated with an Auto-pooling layer. Laure Prétet, Gaël Richard, Geoffroy Peeters |
ICASSP | 3 |
| 2020 | Confidence-based Weighted Loss for Multi-label Classification with Missing LabelsabstractThe problem of multi-label classification with missing labels (MLML) is a common challenge that is prevalent in several domains, e.g. image annotation and auto-tagging. In multi-label classification, each instance may belong to multiple class labels simultaneously. Due to the nature of the dataset collection and labelling procedure, it is common to have incomplete annotations in the dataset, i.e. not all samples are labelled with all the corresponding labels. However, the incomplete data labelling hinders the training of classification models. MLML has received much attention from the research community. However, in cases where a pre-trained model is fine-tuned on an MLML dataset, there has been no straightforward approach to tackle the missing labels, specifically when there is no information about which are the missing ones. In this paper, we propose a weighted loss function to account for the confidence in each label/sample pair that can easily be incorporated to fine-tune a pre-trained model on an incomplete dataset. Our experiment results show that using the proposed loss function improves the performance of the model as the ratio of missing labels increases. Karim M. Ibrahim, Elena V. Epure, Geoffroy Peeters, Gaël Richard |
ICMR | 3 |
| 2018 | Fast and Adaptive Blind Audio Source Separation Using Recursive Levenberg-Marquardt SynchrosqueezingabstractThis paper revisits the Degenerate Unmixing Estimation Technique (duet) for blind audio separation of an arbitrary number of sources given two mixtures through a recursively computed and adaptive time-frequency representation. Recently, synchrosqueezing was introduced as a promising signal disentangling method which allows to compute reversible and sharpen time-frequency representations. Thus, it can be used to reduce overlaps between the sources in the time-frequency plane and to improve the sources' sparsity which is often exploited by source separation techniques. Furthermore, synchrosqueezing can also be extended using the Levenberg-Marquardt algorithm to allow a user to adjust the energy concentration of a time-frequency representation which can be efficiently implemented without the FFT algorithm. Hence, we show that our approach can improve the quality of the source separation process while remaining suitable for real-time applications. Dominique Fourer, Geoffroy Peeters |
ICASSP | 2 |
| 2018 | Local AM/FM Parameters Estimation: Application to Sinusoidal Modeling and Blind Audio Source SeparationabstractThis letter extends our recently introduced method which was designed to estimate instantaneous frequency and chirp rate of linearly modulated signals. Indeed, we derive several new estimators related to our previous ones which provide in the time-frequency plane all the signal parameters of the investigated model: amplitude, frequency, and their local modulations (AM/FM). Our estimators are first introduced and compared in terms of statistical efficiency with theoretical bounds and with other state-of-the-art estimators. Then, they are used to improve spectral analysis applied to audio sinusoidal modeling. Finally, they lead to a new source separation technique based on coherent amplitude and frequency modulation that is evaluated on real-world music signals. Dominique Fourer, François Auger, Geoffroy Peeters |
IEEE Signal Process. Lett. | 3 |
| 2017 | Objective characterization of audio signal quality: Applications to music collection descriptionabstractIn this paper, we propose a set of audio features to describe the quality of an audio signal. Audio quality is here considered as being modified by the chain of processes/effects applied to the individual instrument tracks to obtain the final mix of a musical piece. Thus, the quality also depends on the mastering processes applied to the final mix or the signal degradation caused by MP3 compression. To evaluate our proposal, we created a large set of artificial mixes and also used real-world studio mixes. Using unsupervised and supervised classification methods, we show that our proposed audio features can detect the processing chain. Since this processing chain applied in professional studio has evolved over the years, we use our audio features to directly predict the decade during which a music track was recorded. Dominique Fourer, Geoffroy Peeters |
ICASSP | 2 |
| 2017 | Multimodal speaker clustering in full length movies
Ioannis Kapsouras, Anastasios Tefas, Nikos Nikolaidis 0001, Geoffroy Peeters, Elie-Laurent Benaroya, Ioannis Pitas |
Multim. Tools Appl. | 4 |
| 2014 | A Pitch Salience Function Derived from Harmonic Frequency Deviations for Polyphonic Music Analysis
Alessio Degani, Riccardo Leonardi, Pierangelo Migliorati, Geoffroy Peeters |
DAFx | 4 |
| 2014 | The Modulation Scale Spectrum and its Application to Rhythm-Content Description
Ugo Marchand, Geoffroy Peeters |
DAFx | 2 |
| 2014 | 2D/3D AudioVisual content analysis & descriptionabstractIn this paper, we propose a way of using the Audio-Visual Description Profile (AVDP) of the MPEG-7 standard for 2D or stereo video and multichannel audio content description. Our aim is to provide means of using AVDP in such a way, that 3D video and audio content can be correctly and consistently described. Since AVDP semantics do not include ways for dealing with 3D audiovisual content, a new semantic framework within AVDP is proposed and examples of using AVDP to describe the results of analysis algorithms on stereo video and multichannel audio content are presented. Ioannis Pitas, Konstantinos Papachristou, Nikos Nikolaidis 0001, Marco Liuni, Elie-Laurent Benaroya, Geoffroy Peeters, Axel Röbel, Antje Linnemann, Mohan Liu, Sebastian Gerke |
MMSP | 6 |
| 2013 | Multiple hypotheses at multiple scales for audio novelty computation within musicabstractNovelty-based segmentation of audio signals has proven good performances for the estimation of boundaries of structural sections within music pieces. However, boundaries are detected only if structural sections satisfy the condition of sufficient acoustic inner-homogeneity. While this constraint is very restrictive and not representative of all musical contents, we propose in this paper to extend the detection of acoustic novelty to transitions between homogeneous and non-homogeneous sections and vice versa. Moreover, the length of the considered sections for the boundary detection is crucial, we also introduce a multi-scale novelty approach that allows to capture boundaries between sections of different temporal scales in a same segmentation. Evaluation of the combination of these two methods proves convincing results for temporal segmentation of music pieces. Embedding the algorithm in a music structure segmentation system, we show that performances can be consistently improved for this task. Florian Kaiser, Geoffroy Peeters |
ICASSP | 2 |
| 2013 | Evaluating automatically estimated chord sequencesabstractIn this paper, we perform an in-depth evaluation of a large number of algorithms for chord estimation that have been submitted to the MIREX competitions in 2010, 2011 and 2012. Therefore we first present a rigorous scheme to describe evaluation methods in a sound, unambiguous way that extends previous work specifically to take into account the large variance in chord estimation vocabularies and to perform evaluations on select sets of chords. Then we take a look at the evaluation metrics used so far and propose some alternative ones. Finally, we use these different methods to get a deeper insight into the strengths of each of the competing algorithms and show that the choice of evaluation measure greatly influences the ranking. Johan Pauwels, Geoffroy Peeters |
ICASSP | 2 |
| 2013 | AudioPrint: An efficient audio fingerprint system based on a novel cost-less synchronization schemeabstractThis paper presents the latest improvements on AudioPrint: the IRCAM audio fingerprint system. Cosine filters are introduced in the short-term spectral analysis, in order to compensate the effect of pitch shifting, and a simple solution is proposed for the determination of the frame positions, robust to audio degradations, with nearly no additional cost. We then show that both contributions significantly improve the Audio-Print system, with evaluations both on a free corpus, made publicly available, and a real-world corpus of broadcast radio streams. Mathieu Ramona, Geoffroy Peeters |
ICASSP | 2 |
| 2013 | Segmenting music through the joint estimation of keys, chords and structural boundariesabstractIn this paper, we introduce a new approach to music structure segmentation that is based on the joint estimation of structural segments, keys and chords in one probabilistic framework. More precisely, the boundaries of a structure segment are determined by detecting key changes and by utilizing the difference in prior probability of chord transitions according to their position in a structural segment. In contrast to many of the recent approaches to structural segmentation, this system does not work with self-similarity matrices, although it has been designed to integrate this kind of approach into the framework at a later stage. However, just the current version of the system, using only the estimated harmony, is already producing encouraging results, especially with respect to the precise localization of the boundaries. Johan Pauwels, Geoffroy Peeters |
ACM Multimedia | 2 |
| 2013 | A professionally annotated and enriched multimodal data set on popular musicabstractThis paper presents the MusiClef data set, a multimodal data set of professionally annotated music. It includes editorial metadata about songs, albums, and artists, as well as MusicBrainz identifiers to facilitate linking to other data sets. In addition, several state-of-the-art audio features are provided. Different sets of annotations and music context data -- collaboratively generated user tags, web pages about artists and albums, and the annotation labels provided by music experts -- are included too. Versions of this data set were used in the MusiClef evaluation campaigns in 2011 and 2012 for auto-tagging tasks. We report on the motivation for the data set, on its composition, on related sets, and on the evaluation campaigns in which versions of the set were already used. These campaigns likewise represent one use case, i.e. music auto-tagging, of the data set. The complete data set is publicly available for download at http://www.cp.jku.at/musiclef. Markus Schedl, Nicola Orio, Cynthia C. S. Liem, Geoffroy Peeters |
MMSys | 4 |
| 2012 | Singer verification: Singer model .vs. song modelabstractThis paper proposes a method to verify the singer identity of a given song. The query song is modeled as a GMM learned on the features extracted from sustained sung notes of the song. Each note is described by the shape its spectral envelope and by the temporal variations in frequency and amplitude of its fundamental frequency. The singer identity is verified with two approaches: the model of the query song is compared to a singer-based GMM or compared to the GMM of another song performed by the same singer. The comparison is done using a dissimilarity measurement given by the Kullback Leibler divergence. When the two types of features are combined, the proposed approach verifies the singer identity of a given a cappella song with an error rate lower than 8% when the whole song is considered and an error rate lower than 10% when a short excerpt of the song (i.e. 15 consecutive sustained notes) is considered. Lise Regnier, Geoffroy Peeters |
ICASSP | 2 |
| 2012 | Local Key Estimation From an Audio Signal Relying on Harmonic and Metrical StructuresabstractIn this paper, we present a method for estimating the progression of musical key from an audio signal. We address the problem of local key finding by investigating the possible combination and extension of different previously proposed approaches for global key estimation. In this work, key progression is estimated from the chord progression. Specifically, we introduce key dependency on the harmonic and the metrical structures. A contribution of our work is that we address the problem of finding an analysis window length for local key estimation that is adapted to the intrinsic music content of the analyzed piece by introducing information related to the metrical structure in our model. Key estimation is not performed on empirically chosen segments but on segments that are expressed in relationship with the tempo period. We evaluate and analyze our results on two databases of different styles. We systematically analyze the influence of various parameters to determine factors important to our model, we study the relationships between the various musical attributes that are taken into account in our work, and we provide case study examples. Hélène Papadopoulos, Geoffroy Peeters |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Audio identification based on spectral modeling of bark-bands energy and synchronization through onset detectionabstractIn this paper, we present for the first time the fingerprint IRCAM system for audio identification in streams. The baseline system relies on a double-nested Short Time Fourier Transform. The first STFT computes the energies of a filter-bank, that are then modelled over 2 s, using a second STFT. We then present recent improvements of our system: first the inclusion of perceptual scales for amplitude and frequency (Bark bands), then the synchronization of stream and database frames using an onset detection system. The performance of these improvements is tested on a large set of real audio streams. We compare our results with the results of re-implementations of the two state-of-the-art systems of Philips and Shazam. Mathieu Ramona, Geoffroy Peeters |
ICASSP | 2 |
| 2011 | Drum extraction from polyphonic music based on a spectro-temporal model of percussive soundsabstractIn this paper, we present a new algorithm for removing drums from a polyphonic audio signal. The aim of this algorithm is to discard time/frequency bins which present a percussive magnitude evolution, according to a pre-defined parametric model. Special care is taken to reduce the irrelevant removal of frequency modulated signal such as the ones produced by the singing voice. Performance evaluation is carried out using objective measures commonly used by the community. Compared with four state-of-the-art algorithms, the proposed algorithm shows competitive performances at a low computational cost. François Rigaud, Mathieu Lagrange, Axel Röbel, Geoffroy Peeters |
ICASSP | 4 |
| 2011 | Joint Estimation of Chords and Downbeats From an Audio SignalabstractWe present a new technique for joint estimation of the chord progression and the downbeats from an audio file. Musical signals are highly structured in terms of harmony and rhythm. In this paper, we intend to show that integrating knowledge of mutual dependencies between chords and metric structure allows us to enhance the estimation of these musical attributes. For this, we propose a specific topology of hidden Markov models that enables modelling chord dependence on metric structure. This model allows us to consider pieces with complex metric structures such as beat addition, beat deletion or changes in the meter. The model is evaluated on a large set of popular music songs from the Beatles that present various metric structures. We compare a semi-automatic model in which the beat positions are annotated, with a fully automatic model in which a beat tracker is used as a front-end of the system. The results show that the downbeat positions of a music piece can be estimated in terms of its harmonic structure and that conversely the chord progression estimation benefits from considering the interaction between the metric and the harmonic structures. Hélène Papadopoulos, Geoffroy Peeters |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Spectral and Temporal Periodicity Representations of Rhythm for the Automatic Classification of Music Audio SignalabstractIn this paper, we study the spectral and temporal periodicity representations that can be used to describe the characteristics of the rhythm of a music audio signal. A continuous-valued energy-function representing the onset positions over time is first extracted from the audio signal. From this function we compute at each time a vector which represents the characteristics of the local rhythm. Four feature sets are studied for this vector. They are derived from the amplitude of the discrete Fourier transform (DFT), the auto-correlation function (ACF), the product of the DFT and the ACF interpolated on a hybrid lag/frequency axis and the concatenated DFT and ACF coefficients. Then the vectors are sampled at some specific frequencies, which represent various ratios of the local tempo. The ability of these periodicity representations to describe the rhythm characteristics of an audio item is evaluated through a classification task. In this, we test the use of the periodicity representations alone, combined with tempo information and combined with a proposed set of rhythm features. The evaluation is performed using annotated and estimated tempo. We show that using such simple periodicity representations allows achieving high recognition rates at least comparable to previously published results. Geoffroy Peeters |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | Simultaneous Beat and Downbeat-Tracking Using a Probabilistic Framework: Theory and Large-Scale EvaluationabstractThis paper deals with the simultaneous estimation of beat and downbeat location in an audio-file. We propose a probabilistic framework in which the time of the beats and their associated beat-position-inside-a-bar roles; hence, the downbeats, are considered as hidden states and are estimated simultaneously using signal observations. For this, we propose a “reverse” Viterbi algorithm which decodes hidden states over beat-numbers. A beat-template is used to derive the beat observation probabilities. For this task, we propose the use of a machine-learning method, the Linear Discriminant Analysis, to estimate the most discriminative beat-templates. We propose two functions to derive the beat-position-inside-a-bar observation probability: the variation over time of chroma vectors and the spectral balance. We then perform a large-scale evaluation of beat and downbeat-tracking using six test-sets. In this, we study the influence of the various parameters of our method, compare this method to our previous beat and downbeat-tracking algorithms, and compare our results to state-of-the-art results on two test-sets for which results have been published. We finally discuss the results obtained by our system in the MIREX-09 and MIREX-10 contests for which our system ranked among the first for the “McKinney Collection” test-set. Geoffroy Peeters, Hélène Papadopoulos |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | Partial clustering using a time-varying frequency model for singing voice detectionabstractWe propose a new method to group partials produced by each instrument of a polyphonic audio mixture. This method works for pitched and harmonic instruments and is specially adapted to singing voice. In our approach, we model time-varying frequencies of partials as a slowly varying frequency plus a sinusoidal modulation. The parameters obtained with this model plus some common Auditory Scene Analysis principles are used to define a similarity measure between partials. This multi-criterion based measure is then used to build the input similarity matrix of a clustering algorithm. Clusters obtained are groups of harmonically related partials. We evaluate the ability of our method to group partials per source when one of the sources is a singing voice. We show that partial clustering is a promising approach for singing voice detection and separation. Lise Regnier, Geoffroy Peeters |
ICASSP | 2 |
| 2010 | Sound Indexing Using Morphological DescriptionabstractSound sample indexing usually deals with the recognition of the source/cause that has produced the sound. For abstract sounds, sound effects, unnatural, or synthetic sounds, this cause is usually unknown or unrecognizable. An efficient description of these sounds has been proposed by Schaeffer under the name morphological description. Part of this description consists in describing a sound by identifying the temporal evolution of its acoustic properties to a set of profiles. In this paper, we consider three morphological descriptions: dynamic profiles (ascending, descending, ascending/descending, stable, impulsive), melodic profiles (up, down, stable, up/down, down/up) and complex-iterative sound description (non-iterative, iterative, grain, repetition). We study the automatic indexing of a sound into these profiles. Because this automatic indexing is difficult using standard audio features, we propose new audio features to perform this task. The dynamic profiles are estimated by modeling the loudness over-time of a sound by a second-order B-spline model and derive features from this model. The melodic profiles are estimated by tracking over time the perceptual filter which has the maximum excitation. A function is derived from this track which is then modeled using a second-order B-spline model. The features are again derived from the B-spline model. The description of complex-iterative sounds is obtained by estimating the amount of repetition and the period of the repetition. These are obtained by computing an audio similarity function derived from an Mel frequency cepstral coefficients (MFCC) similarity matrix. The proposed audio features are then tested for automatic classification. We consider three classification tasks corresponding to the three profiles. In each case, the results are compared with the ones obtained using standard audio features. Geoffroy Peeters, Emmanuel Deruty |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | Singing voice detection in music tracks using direct voice vibrato detectionabstractIn this paper we investigate the problem of locating singing voice in music tracks. As opposed to most existing methods for this task, we rely on the extraction of the characteristics specific to singing voice. In our approach we suppose that the singing voice is characterized by harmonicity, formants, vibrato and tremolo. In the present study we deal only with the vibrato and tremolo characteristics. For this, we first extract sinusoidal partials from the musical audio signal . The frequency modulation (vibrato) and amplitude modulation (tremolo) of each partial are then studied to determine if the partial corresponds to singing voice and hence the corresponding segment is supposed to contain singing voice. For this we estimate for each partial the rate (frequency of the modulations) and the extent (amplitude of modulation) of both vibrato and tremolo. A partial selection is then operated based on these values. A second criteria based on harmonicity is also introduced. Based on this, each segment can be labelled as singing or non-singing. Post-processing of the segmentation is then applied in order to remove short-duration segments. The proposed method is then evaluated on a large manually annotated test-set. The results of this evaluation are compared to the one obtained with a usual machine learning approach (MFCC and SFM modeling with GMM). The proposed method achieves very close results to the machine learning approach : 76.8% compared to 77.4% F-measure (frame classification). This result is very promising, since both approaches are orthogonal and can then be combined. Lise Regnier, Geoffroy Peeters |
ICASSP | 2 |
| 2008 | Simultaneous estimation of chord progression and downbeats from an audio fileabstractHarmony and metrical structure are some of the most important attributes of Western tonal music. In this paper, we present a new method for simultaneously estimating the chord progression and the downbeats from an audio file. For this, we propose a specific topology of hidden Markov models that allows us to model chords dependency on metrical structure. The model is evaluated on a dataset of 66 popular music songs from the Beatles and shows improvement over the state of the art. Hélène Papadopoulos, Geoffroy Peeters |
ICASSP | 2 |
| 2006 | Music Pitch Representation by Periodicity Measures Based on Combined Temporal and Spectral RepresentationsabstractPeriodicity estimation of an audio signal, for applications such as pitch, multiple pitch or tempo estimation is often problematic due to the presence of multiple harmonics in the audio signal producing octave errors. While pitch models or rhythm models can be used, they remain often dedicated to a specific problem. In this paper, we propose a straightforward approach for periodicity estimation based on the combination of a spectral representation and a temporal representation. This method allows a better emphasis on the frequencies corresponding to the various pitches. We show the ability of this representation to adequately estimate pitch and visualize signals with multiple pitch content Geoffroy Peeters |
ICASSP (5) | 1 |