Mathieu Lagrange

dblp:87/3503 · DBLP profile ↗
← Back
34ranked-venue papers
13as first author
7since 2021 · last 2025
0000-0002-1253-4427ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 9 first-author · 5 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 2 since 2021Computer networks · 2 · 1 since 2021
YearPublicationVenuePosition
2025 S-KEY: Self-supervised Learning of Major and Minor Keys from Audio
abstract
STONE, the current method in self-supervised learning for tonality estimation in music signals, cannot distinguish relative keys, such as C major versus A minor. In this article, we extend the neural network architecture and learning objective of STONE to perform self-supervised learning of major and minor keys (S-KEY). Our main contribution is an auxiliary pretext task to STONE, formulated using transposition-invariant chroma features as a source of pseudo-labels. S-KEY matches the supervised state of the art in tonality estimation on FMAKv2 and GTZAN datasets while requiring no human annotation and having the same parameter budget as STONE. We build upon this result and expand the training set of S-KEY to a million songs, thus showing the potential of large-scale self-supervised learning in music information retrieval.
Yuexuan Kong, Gabriel Meseguer-Brocal, Vincent Lostanlen, Mathieu Lagrange, Romain Hennequin
ICASSP4
2025 Understanding Equivariant Self-Supervised Learning in Musical Pitch Class Space
abstract
STONE, which stands for self-supervised tonality estimator, has recently demonstrated the practical feasibility of recognizing key signatures in music signals given little or no human annotation. In this article, we revisit STONE from a more theoretical standpoint. We show that cross-power spectral density (CPSD) defines a differentiable measure of harmonic discrepancy between key signature profiles (KSP). Having set the CPSD frequency to seven cycles per octave, we offer a geometric interpretation of this discrepancy via the circle of fifths and conduct an algebraic study to prove that all its local minima are global. We rely on the equivariance property of deep convolutional networks to prove that the STONE loss function is invariant to circular frequency shifts of the constant-Q transform. We conclude by identifying a phenomenon of spontaneous symmetry breaking in STONE: since modes of limited transposition (e.g., augmented, diminished, whole-tone) are multistable in the circle of fifths, the associated CPSD gradient is driven towards more “tonal” (i.e., asymmetric) scales, such as diatonic or pentatonic.
Vincent Lostanlen, Yuexuan Kong, Gabriel Meseguer-Brocal, Mathieu Lagrange, Romain Hennequin
IEEE Signal Process. Lett.4
2024 Emvd Dataset: a Dataset of Extreme Vocal Distortion Techniques Used in Heavy Metal
abstract
In this paper, we introduce the Extreme Metal Vocals Dataset, which comprises a collection of recordings of extreme vocal techniques performed within the realm of heavy metal music. The dataset consists of 760 audio excerpts of 1 second to 30 seconds long, totaling about 100 min of audio material, roughly composed of 60 minutes of distorted voices and 40 minutes of clear voice recordings. These vocal recordings are from 27 different singers and are provided without accompanying musical instruments or post-processing effects. The distortion taxonomy within this dataset encompasses four distinct distortion techniques and three vocal effects, all performed in different pitch ranges. Performance of a state-of-the-art deep learning model is evaluated for two different classification tasks related to vocal techniques, demonstrating the potential of this resource for the audio processing community.
Modan Tailleur, Julien Pinquier, Laurent Millot, Corsin Vogel, Mathieu Lagrange
CBMI5
2024 Learning to Solve Inverse Problems for Perceptual Sound Matching
abstract
Perceptual sound matching (PSM) aims to find the input parameters to a synthesizer so as to best imitate an audio target. Deep learning for PSM optimizes a neural network to analyze and reconstruct prerecorded samples. In this context, our article addresses the problem of designing a suitable loss function when the training set is generated by a differentiable synthesizer. Our main contribution is perceptual–neural–physical loss (PNP), which aims at addressing a tradeoff between perceptual relevance and computational efficiency. The key idea behind PNP is to linearize the effect of synthesis parameters upon auditory features in the vicinity of each training sample. The linearization procedure is massively parallelizable, can be precomputed, and offers a 100-fold speedup during gradient descent compared to differentiable digital signal processing (DDSP). We demonstrate PNP on two datasets of nonstationary sounds: an AM/FM arpeggiator and a physical model of rectangular membranes. We show that PNP is able to accelerate DDSP with joint time–frequency scattering transform (JTFS) as auditory feature while preserving its perceptual fidelity. Additionally, we evaluate the impact of other design choices in PSM: parameter rescaling, pretraining, auditory representation, and gradient clipping. We report state-of-the-art results on both datasets and find that PNP-accelerated JTFS has greater influence on PSM performance than any other design choice.
Vincent Lostanlen, Mathieu Lagrange
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Perceptual-Neural-Physical Sound Matching
abstract
Sound matching algorithms seek to approximate a target waveform by parametric audio synthesis. Deep neural networks have achieved promising results in matching sustained harmonic tones. However, the task is more challenging when targets are nonstationary and inharmonic, e.g., percussion. We attribute this problem to the inadequacy of loss function. On one hand, mean square error in the parametric domain, known as "P-loss", is simple and fast but fails to accommodate the differing perceptual significance of each parameter. On the other hand, mean square error in the spectrotemporal domain, known as "spectral loss", is perceptually motivated and serves in differentiable digital signal processing (DDSP). Yet, spectral loss is a poor predictor of pitch intervals and its gradient may be computationally expensive; hence a slow convergence. Against this conundrum, we present Perceptual-Neural-Physical loss (PNP). PNP is the optimal quadratic approximation of spectral loss while being as fast as P-loss during training. We instantiate PNP with physical modeling synthesis as decoder and joint time–frequency scattering transform (JTFS) as spectral representation. We demonstrate its potential on matching synthetic drum sounds in comparison with other loss functions.
Vincent Lostanlen, Mathieu Lagrange
ICASSP3
2023 Explainable audio Classification of Playing Techniques with Layer-wise Relevance Propagation
abstract
Deep convolutional networks (convnets) in the time–frequency domain can learn an accurate and fine-grained categorization of sounds. For example, in the context of music signal analysis, this categorization may correspond to a taxonomy of playing techniques: vibrato, tremolo, trill, and so forth. However, convnets lack an explicit connection with the neurophysiological underpinnings of musical timbre perception. In this article, we propose a data-driven approach to explain audio classification in terms of physical attributes in sound production. We borrow from current literature in "explainable AI" (XAI) to study the predictions of a convnet which achieves an almost perfect score on a challenging task: i.e., the classification of five comparable real-world playing techniques from 30 instruments spanning seven octaves. Mapping the signal into the carrier-modulation domain using scattering transform, we decompose the networks’ predictions over this domain with layer-wise relevance propagation. We find that regions highly-relevant to the predictions localized around the physical attributes with which the playing techniques are performed.
Changhong Wang 0002, Vincent Lostanlen, Mathieu Lagrange
ICASSP3
2023 The Internet of Sounds: Convergent Trends, Insights, and Future Directions
abstract
Current sound-based practices and systems developed in both academia and industry point to convergent research trends that bring together the field of Sound and Music Computing with that of the Internet of Things. This paper proposes a vision for the emerging field of the Internet of Sounds (IoS), which stems from such disciplines. The IoS relates to the network of Sound Things, i.e., devices capable of sensing, acquiring, processing, actuating, and exchanging data serving the purpose of communicating sound-related information. In the IoS paradigm, which merges under a unique umbrella the emerging fields of the Internet of Musical Things and the Internet of Audio Things, heterogeneous devices dedicated to musical and non-musical tasks can interact and cooperate with one another and with other things connected to the Internet to facilitate sound-based services and applications that are globally available to the users. We survey the state of the art in this space, discuss the technological and non-technological challenges ahead of us and propose a comprehensive research agenda for the field.
Luca Turchet, Mathieu Lagrange, Cristina Rottondi, György Fazekas, Nils Peters, Jan Østergaard, Frederic Font, Tom Bäckström, Carlo Fischione
IEEE Internet Things J.2
2020 Privacy Aware Acoustic Scene Synthesis Using Deep Spectral Feature Inversion
abstract
Gathering information about the acoustic environment of urban areas is now possible and studied in many major cities in the world. Part of the research is to find ways to inform the citizen about its sound environment while ensuring her privacy.We study in this paper how this application can be cast into a feature inversion problem. We argue that considering deep learning techniques to solve this problem allows us to produce sound sketches that are representative and privacy aware. Experiments done considering the dcase2017 dataset shows that the proposed learning based approach achieves state of the art performance when compared to blind inversion approaches.
Félix Gontier, Mathieu Lagrange, Catherine Lavandier, Jean-François Petiot
ICASSP2
2020 Bandwidth Extension of Musical Audio Signals With No Side Information Using Dilated Convolutional Neural Networks
abstract
Bandwidth extension has a long history in audio processing. While speech processing tools do not rely on side information, production-ready bandwidth extension tools of general audio signals rely on side information that has to be transmitted alongside the bitstream of the low frequency part, mostly because polyphonic music has a more complex and less predictable spectral structure than speech.This paper studies the benefit of considering a dilated fully convolutional neural network to perform the bandwidth extension of musical audio signals with no side information on the magnitude spectra. Experimental evaluation using two public datasets, medley-solos-db and gtzan, respectively of monophonic and polyphonic music demonstrate that the proposed architecture achieves state of the art performance.
Mathieu Lagrange, Félix Gontier
ICASSP1
2020 The Internet of Audio Things: State of the Art, Vision, and Challenges
abstract
The Internet of Audio Things (IoAuT) is an emerging research field positioned at the intersection of the Internet of Things, sound and music computing, artificial intelligence, and human-computer interaction. The IoAuT refers to the networks of computing devices embedded in physical objects (Audio Things) dedicated to the production, reception, analysis, and understanding of audio in distributed environments. Audio Things, such as nodes of wireless acoustic sensor networks, are connected by an infrastructure that enables multidirectional communication, both locally and remotely. In this article, we first review the state of the art of this field, then we present a vision for the IoAuT and its motivations. In the proposed vision, the IoAuT enables the connection of digital and physical domains by means of appropriate information and communication technologies, fostering novel applications and services based on auditory information. The ecosystems associated with the IoAuT include interoperable devices and services that connect humans and machines to support human-human and human-machines interactions. We discuss the challenges and implications of this field, which lead to future research directions on the topics of privacy, security, design of Audio Things, and methods for the analysis and representation of audio-related information.
Luca Turchet, György Fazekas, Mathieu Lagrange, Hossein Shokri Ghadikolaei, Carlo Fischione
IEEE Internet Things J.3
2018 Detection and Classification of Acoustic Scenes and Events: Outcome of the DCASE 2016 Challenge
abstract
Public evaluation campaigns and datasets promote active development in target research areas, allowing direct comparison of algorithms. The second edition of the challenge on detection and classification of acoustic scenes and events (DCASE 2016) has offered such an opportunity for development of the state-of-the-art methods, and succeeded in drawing together a large number of participants from academic and industrial backgrounds. In this paper, we report on the tasks and outcomes of the DCASE 2016 challenge. The challenge comprised four tasks: acoustic scene classification, sound event detection in synthetic audio, sound event detection in real-life audio, and domestic audio tagging. We present each task in detail and analyze the submitted systems in terms of design and performance. We observe the emergence of deep learning as the most popular classification method, replacing the traditional approaches based on Gaussian mixture models and support vector machines. By contrast, feature representations have not changed substantially throughout the years, as mel frequency-based representations predominate in all tasks. The datasets created for and used in DCASE 2016 are publicly available and are a valuable resource for further research.
Annamaria Mesaros, Toni Heittola, Emmanouil Benetos, Peter Foster, Mathieu Lagrange, Tuomas Virtanen, Mark D. Plumbley
IEEE ACM Trans. Audio Speech Lang. Process.5
2017 Polyphonic Sound Event Tracking Using Linear Dynamical Systems
abstract
In this paper, a system for polyphonic sound event detection and tracking is proposed, based on spectrogram factorization techniques and state space models. The system extends probabilistic latent component analysis (PLCA) and is modeled around a four-dimensional spectral template dictionary of frequency, sound event class, exemplar index, and sound state. In order to jointly track multiple overlapping sound events over time, the integration of linear dynamical systems (LDS) within the PLCA inference is proposed. The system assumes that the PLCA sound event activation is the (noisy) observation in an LDS, with the latent states corresponding to the true event activations. LDS training is achieved using fully observed data, making use of ground truth-informed event activations produced by the PLCA-based model. Several LDS variants are evaluated, using polyphonic datasets of office sounds generated from an acoustic scene simulator, as well as real and synthesized monophonic datasets for comparative purposes. Results show that the integration of LDS tracking within PLCA leads to an improvement of +8.5-10.5% in terms of frame-based F-measure as compared to the use of the PLCA model alone. In addition, the proposed system outperforms several state-of-the-art methods for the task of polyphonic sound event detection.
Emmanouil Benetos, Grégoire Lafay, Mathieu Lagrange, Mark D. Plumbley
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 Detection of overlapping acoustic events using a temporally-constrained probabilistic model
abstract
In this paper, a system for overlapping acoustic event detection is proposed, which models the temporal evolution of sound events. The system is based on probabilistic latent component analysis, supporting the use of a sound event dictionary where each exemplar consists of a succession of spectral templates. The temporal succession of the templates is controlled through event class-wise Hidden Markov Models (HMMs). As input time/frequency representation, the Equivalent Rectangular Bandwidth (ERB) spectrogram is used. Experiments are carried out on polyphonic datasets of office sounds generated using an acoustic scene simulator, as well as real and synthesized monophonic datasets for comparative purposes. Results show that the proposed system outperforms several state-of-the-art methods for overlapping acoustic event detection on the same task, using both frame-based and event-based metrics, and is robust to varying event density and noise levels.
Emmanouil Benetos, Grégoire Lafay, Mathieu Lagrange, Mark D. Plumbley
ICASSP3
2016 A Morphological Model for Simulating Acoustic Scenes and Its Application to Sound Event Detection
abstract
This paper introduces a model for simulating environmental acoustic scenes that abstracts temporal structures from audio recordings. This model allows us to explicitly control key morphological aspects of the acoustic scene and to isolate their impact on the performance of the system under evaluation. Thus, more information can be gained on the behavior of an evaluated system, providing guidance for further improvements. To demonstrate its potential, this model is employed to evaluate the performance of nine state of the art sound event detection systems submitted to the IEEE DCASE 2013 Challenge. Results indicate that the proposed scheme is able to successfully build datasets useful for evaluating important aspects of the performance of sound event detection systems, such as their robustness to new recording conditions and to varying levels of background audio.
Grégoire Lafay, Mathieu Lagrange, Mathias Rossignol, Emmanouil Benetos, Axel Röbel
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 On automatic drum transcription using non-negative matrix deconvolution and itakura saito divergence
abstract
This paper presents an investigation into the detection and classification of drum sounds in polyphonic music and drum loops using non-negative matrix deconvolution (NMD) and the Itakura Saito divergence. The Itakura Saito divergence has recently been proposed as especially appropriate for decomposing audio spectra due to the fact that it is scale invariant, but it has not yet been widely adopted. The article studies new contributions for audio event detection methods using the Itakura Saito divergence that improve efficiency and numerical stability, and simplify the generation of target pattern sets. A new approach for handling background sounds is proposed and moreover, a new detection criteria based on estimating the perceptual presence of the target class sources is introduced. Experimental results obtained for drum detection in polyphonic music and drum soli demonstrate the beneficial effects of the proposed extensions.
Axel Röbel, Jordi Pons, Marco Liuni, Mathieu Lagrange
ICASSP4
2015 Detection and Classification of Acoustic Scenes and Events
abstract
For intelligent systems to make best use of the audio modality, it is important that they can recognize not just speech and music, which have been researched as specific tasks, but also general sounds in everyday environments. To stimulate research in this field we conducted a public research challenge: the IEEE Audio and Acoustic Signal Processing Technical Committee challenge on Detection and Classification of Acoustic Scenes and Events (DCASE). In this paper, we report on the state of the art in automatically classifying audio scenes, and automatically detecting and classifying audio events. We survey prior work as well as the state of the art represented by the submissions to the challenge from various research groups. We also provide detail on the organization of the challenge, so that our experience as challenge hosts may be useful to those organizing challenges in similar domains. We created new audio datasets and baseline systems for the challenge; these, as well as some submitted systems, are publicly available under open licenses, to serve as benchmarks for further research in general-purpose machine listening.
Dan Stowell, Dimitrios Giannoulis, Emmanouil Benetos, Mathieu Lagrange, Mark D. Plumbley
IEEE Trans. Multim.4
2013 Representing environmental sounds using the separable scattering transform
abstract
Environmental sounds are an interesting subject of study for machine audition because of their wide variety of acoustical characteristics and their central presence in our everyday life. They are perceived effortlessly in the human auditory system whereas state-of-the-art computational systems are far from reaching the same efficiency. In this paper we propose a novel representation of such sounds based on the scattering transform which has the property of stability to time-warping deformations and invariance to time-shift useful for classifications tasks. This representation is compared to several state-of-the-art approaches for the task of quantifying similarity between environmental sounds.
Carlo Bauge, Mathieu Lagrange, Joakim Andén, Stéphane Mallat
ICASSP2
2013 Uncertainty-based learning of acoustic models from noisy data
Alexey Ozerov, Mathieu Lagrange, Emmanuel Vincent 0001
Comput. Speech Lang.2
2012 A regressive boosting approach to automatic audio tagging based on soft annotator fusion
abstract
Automatic tagging of music has mostly been treated as a classification problem. In this framework, the association of a tag to a song is characterized in a “hard” fashion: the tag is either relevant or not. Yet, the relevance of a tag to a song is not always evident. Indeed, during the ground-truth annotation process, several annotators may express doubts, or disagree with each other. In this paper, we propose to fuse annotators' decisions in a way to keep information about this uncertainty. This fusion provides us continuous scores, that are used for training a regressive boosting algorithm. Our experiments show that regression with this soft ground truth leads to a more accurate learning, and better predictions, compared to traditionally used binary classification.
Rémi Foucard, Slim Essid, Mathieu Lagrange, Gaël Richard
ICASSP3
2012 Cluster aware normalization for enhancing audio similarity
abstract
An important task in Music Information Retrieval is content-based similarity retrieval in which given a query music track, a set of tracks that are similar in terms of musical content are retrieved. A variety of audio features that attempt to model different aspects of the music have been proposed. In most cases the resulting audio feature vector used to represent each music track is high dimensional. It has been observed that high dimensional music similarity spaces exhibit some anomalies: hubs which are tracks that are similar to many other tracks, and orphans which are tracks that are not similar to most other tracks. These anomalies are an artifact of the high dimensional representation rather than actually based on the musical content. In this work we describe a distance normalization method that is shown to reduce the number of hubs and orphans. It is based on post-processing the similarity matrix that encodes the pair-wise track similarities and utilizes clustering to adapt the distance normalization to the local structure of the feature space.
Mathieu Lagrange, Luis Gustavo Martins, George Tzanetakis
ICASSP1
2011 Adaptive N-normalization for enhancing music similarity
abstract
The N-Normalization is an efficient method for normalizing a given similarity computed among multimedia objects. It can be considered for clustering and kernel enhancement. However, most approaches to N-Normalization parametrize the method arbitrarily in an ad-hoc manner. In this paper, we show that the optimal parameterization is tightly related to the geometry of the problem at hand. For that purpose, we propose a method for estimating an optimal parameterization given only the associated pair-wise similarities computed from any specific dataset. This allows us to normalize the similarity in a meaningful manner. More specifically, the proposed method allows us to improve retrieval performance as well as minimize unwanted phenomena such as hubs and orphans.
Mathieu Lagrange, George Tzanetakis
ICASSP1
2011 Drum extraction from polyphonic music based on a spectro-temporal model of percussive sounds
abstract
In this paper, we present a new algorithm for removing drums from a polyphonic audio signal. The aim of this algorithm is to discard time/frequency bins which present a percussive magnitude evolution, according to a pre-defined parametric model. Special care is taken to reduce the irrelevant removal of frequency modulated signal such as the ones produced by the singing voice. Performance evaluation is carried out using objective measures commonly used by the community. Compared with four state-of-the-art algorithms, the proposed algorithm shows competitive performances at a low computational cost.
François Rigaud, Mathieu Lagrange, Axel Röbel, Geoffroy Peeters
ICASSP2
2010 Multimodal similarity between musical streams for cover version detection
abstract
Expressing the similarity between musical streams is a challenging task as it involves the understanding of many factors which are most often blended into one information channel: the audio stream. Consequently, separating the musical audio stream into its main melody and its accompaniment may prove as being useful to root the similarity computation on a more robust and expressive representation. In this paper, we show that considering the mixture, an estimation of its main melody and its accompaniment as modalities allows us to propose new ways of defining the similarity between musical streams. In the context of the detection of cover version, we show that highest performance is achieved by jointly considering the mixture and the estimated accompaniment. As demonstrated by the experiments carried out using two different evaluation databases, this scheme allows the scoring system to focus more on the chord progression by considering the accompaniment while being robust to the potential separation errors by also considering the mixture.
Rémi Foucard, Jean-Louis Durrieu, Mathieu Lagrange, Gaël Richard
ICASSP3
2010 Robust similarity metrics between audio signals based on asymmetrical spectral envelope matching
abstract
In this paper, a new type of metric that defines the similarity between musical audio signals is proposed. Based on the spectral flatness criterion, those metrics achieve low computational cost and low sensitivity to acoustical degradations. Validation is performed by studying the ability of the proposed metric to determine whether two audio signals have been played by the same musical instrument. For this task, proposed metrics are shown to overcome metrics based on the comparison of standard spectral features especially when the request and the records of the database are of different acoustical properties.
Mathieu Lagrange, Roland Badeau, Gaël Richard
ICASSP1
2010 Spectral similarity metrics for sound source formation based on the common variation cue
Mathieu Lagrange, Martin Raspaud
Multim. Tools Appl.1
2010 Explicit modeling of temporal dynamics within musical signals for acoustical unit similarity
Mathieu Lagrange, Martin Raspaud, Roland Badeau, Gaël Richard
Pattern Recognit. Lett.1
2010 Analysis/Synthesis of Sounds Generated by Sustained Contact Between Rigid Objects
abstract
This paper introduces an analysis/synthesis scheme for the reproduction of sounds generated by sustained contact between rigid bodies. This scheme is rooted in a Source/Filter decomposition of the sound where the filter is described as a set of poles and the source is described as a set of impulses representing the energy transfer between the interacting objects. Compared to single impacts, sustained contact interactions like rolling and sliding make the estimation of the parameters of the Source/Filter model challenging because of two issues. First, the objects are almost continuously interacting. The second is that the source is generally unknown and has therefore to be modeled in a generic way. In an attempt to tackle those issues, the proposed analysis/synthesis scheme combines advanced analysis techniques for the estimation of the filter parameters and a flexible model of the source. It allows the modeling of a wide range of sounds. Examples are presented for objects of various shapes and sizes, rolling or sliding over plates of different materials. In order to demonstrate the versatility of the approach, the system is also considered for the modeling of sounds produced by percussive musical instruments.
Mathieu Lagrange, Gary P. Scavone, Philippe Depalle
IEEE Trans. Speech Audio Process.1
2008 A Computationally Efficient Scheme for Dominant Harmonic Source Separation
abstract
The leading voice is an important feature of musical pieces and can often be considered as the dominant harmonic source. We propose in this paper a new scheme for the purpose of efficient dominant harmonic source separation. This is achieved by considering an harmonicity cue which is first compared with state-of-the-art cues using a generic evaluation methodology. The proposed separation scheme is then compared to a generic computational auditory scene analysis framework. Computational speed-up and performance comparison is done using source separation and music information retrieval tasks.
Mathieu Lagrange, Luis Gustavo Martins, George Tzanetakis
ICASSP1
2008 MarsyasX: multimedia dataflow processing with implicit patching
abstract
The design and implementation of multimedia signal processing systems is challenging especially when efficiency and real-time performance is desired. In many modern applications, software systems must be able to handle multiple flows of various types of multimedia data such as audio and video. Researchers frequently have to rely on a combination of different software tools for each modality to assemble proof-of-concept systems that are inefficient, brittle and hard to maintain. Marsyas is a software framework originally developed to address these issues in the domain of audio processing. In this paper we describe MarsyasX, a new open-source cross-modal analysis framework that aims at a broader score of applications. It follows a dataflow architecture where complex networks of processing objects can be assembled to form systems that can handle multiple and different types of multimedia flows with expressiveness and efficiency.
Luís F. Teixeira 0001, Luis Gustavo Martins, Mathieu Lagrange, George Tzanetakis
ACM Multimedia3
2008 Normalized Cuts for Predominant Melodic Source Separation
abstract
The predominant melodic source, frequently the singing voice, is an important component of musical signals. In this paper, we describe a method for extracting the predominant source and corresponding melody from ldquoreal-worldrdquo polyphonic music. The proposed method is inspired by ideas from computational auditory scene analysis. We formulate predominant melodic source tracking and formation as a graph partitioning problem and solve it using the normalized cut which is a global criterion for segmenting graphs that has been used in computer vision. Sinusoidal modeling is used as the underlying representation. A novel harmonicity cue which we term harmonically wrapped peak similarity is introduced. Experimental results supporting the use of this cue are presented. In addition, we show results for automatic melody extraction using the proposed approach.
Mathieu Lagrange, Luis Gustavo Martins, Jennifer Murdoch, George Tzanetakis
IEEE Trans. Speech Audio Process.1
2007 Sound Source Tracking and Formation using Normalized Cuts
abstract
The goal of computational auditory scene analysis (CASA) is to create computer systems that can take as input a mixture of sounds and form packages of acoustic evidence such that each package most likely has arisen from a single sound source. We formulate sound source tracking and formation as a graph partitioning problem and solve it using the normalized cut which is a global criterion for segmenting graphs that has been used in computer vision. It measures both the total dissimilarity between the different groups as well as the total similarity within groups. We describe how this formulation can be used with sinusoidal modeling, a common technique for sound analysis, manipulation and synthesis. Several examples showing the potential of this approach are provided.
Mathieu Lagrange, George Tzanetakis
ICASSP (1)1
2007 Enhancing the Tracking of Partials for the Sinusoidal Modeling of Polyphonic Sounds
abstract
This paper addresses the problem of tracking partials, i.e., determining the evolution over time of the parameters of a given number of sinusoids with respect to the analyzed audio stream. We first show that the minimal frequency difference heuristic generally used to identify continuities between local maxima of successive short-time spectra can be successfully generalized using the linear prediction formalism to handle modulated sounds such as musical tones with vibrato. The spectral properties of the evolutions in time of the parameters of the partials are next studied to ensure that the parameters of the partials effectively satisfy the slow time-varying constraint of the sinusoidal model. These two improvements are combined in a new algorithm designed for the sinusoidal modeling of polyphonic sounds. The comparative tests show that onsets/offsets of sinusoids as well as closely spaced sinusoids are better identified and stochastic components are better avoided.
Mathieu Lagrange, Sylvain Marchand, Jean-Bernard Rault
IEEE Trans. Speech Audio Process.1
2005 Tracking partials for the sinusoidal modeling of polyphonic sounds
abstract
The paper proposes to improve further the tracking of partials in a polyphonic context. Spectral characteristics of the controlling parameters (amplitude and frequency) are taken into account to ensure that these parameters evolve slowly with time. The resulting algorithm tracks closely-spaced sinusoids better and is able to avoid most of the spectral data belonging to noise. As a consequence, the proposed algorithm extracts a more meaningful sinusoidal representation from polyphonic recordings.
Mathieu Lagrange, Sylvain Marchand, Jean-Bernard Rault
ICASSP (3)1
2004 Using linear prediction to enhance the tracking of partials [musical audio processing]
abstract
In this article, we present an enhanced algorithm, of low complexity, for the tracking of partials in the context of sinusoidal modeling. By considering the past evolution of each partial in the time/frequency and time/amplitude planes to predict its future evolutions, this algorithm allows a better discrimination between sinusoidal and noisy components and an easier cancellation of sudden changes in the evolutions of the partials.
Mathieu Lagrange, Sylvain Marchand, Jean-Bernard Rault
ICASSP (4)1