Gaël Richard

dblp:34/1310 · DBLP profile ↗
← Back
151ranked-venue papers
7as first author
30since 2021 · last 2026
0000-0002-4960-0010ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 101 · 5 first-author · 18 since 2021Artificial intelligence and machine learning · 54 · 2 first-author · 13 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 AudioCAN: Enhanced Few-Shot Audio Classification via Energy-Guided Temporal Cross Attention
abstract
Despite the growing research interest and practical potential of few-shot audio classification, its efficacy remains limited by the temporal sparsity of target sound events and the presence of non-target acoustic interference within audio samples. In this work, we leverage the cross attention mechanism to address these challenges. While the Cross Attention Network (CAN) has demonstrated superior performance in few-shot image classification by highlighting semantically relevant regions between support and query features, we propose AudioCAN by extending CAN to audio domain through two-fold, audio-specific modifications. In particular, we first reformulate the original 2D spatial cross attention to 1D temporal version to prioritize the time frames containing discriminative acoustic information, thereby mitigating interference from non-target sounds. Secondly, we introduce an energy-guided masking strategy to synthesize a pseudo-query set from the support set via data augmentation, which then serves as a guidance for optimizing the 1D temporal cross attention during training. Experimental results on several few-shot audio classification benchmarks demonstrate that AudioCAN achieves state-of-the-art performance in 5-way-1-shot settings while remaining highly competitive in 5-way-5-shot configurations.
Xuanyu Zhuang, Geoffroy Peeters, Gaël Richard
IEEE Signal Process. Lett.3
2025 iKnow-audio: Integrating Knowledge Graphs with Audio-Language Models
abstract
Contrastive Language-Audio Pretraining (CLAP) models learn by aligning audio and text in a shared embedding space, enabling powerful zero-shot recognition.However, their performance is highly sensitive to prompt formulation and language nuances, and they often inherit semantic ambiguities and spurious correlations from noisy pretraining data.While prior work has explored prompt engineering, adapters, and prefix tuning to address these limitations, the use of structured prior knowledge remains largely unexplored.We present iKnow-audio, a framework that integrates knowledge graphs with audio-language models to provide robust semantic grounding.iKnow-audio builds on the Audio-centric Knowledge Graph (AKG), which encodes ontological relations comprising semantic, causal, and taxonomic connections reflective of everyday sound scenes and events.By training knowlege graph embedding models on the AKG and refining CLAP predictions through this structured knowledge, iKnow-audio improves disambiguation of acoustically similar sounds and reduces reliance on prompt engineering.Comprehensive zero-shot evaluations across six benchmark datasets demonstrate consistent gains over baseline CLAP, supported by embedding-space analyses that highlight improved relational grounding.Resources are publicly available at
Michel Olvera, Changhong Wang 0002, Paraskevas Stamatiadis, Gaël Richard, Slim Essid
EMNLP4
2025 F-StrIPE: Fast Structure-Informed Positional Encoding for Symbolic Music Generation
abstract
While music remains a challenging domain for generative models like Transformers, recent progress has been made by exploiting suitable musically-informed priors. One technique to leverage information about musical structure in Transformers is inserting such knowledge into the positional encoding (PE) module. However, Transformers carry a quadratic cost in sequence length. In this paper, we propose F-StrIPE, a structure-informed PE scheme that works in linear complexity. Using existing kernel approximation techniques based on random features, we show that F-StrIPE is a generalization of Stochastic Positional Encoding (SPE). We illustrate the empirical merits of F-StrIPE using melody harmonization for symbolic music.
Manvi Agarwal, Changhong Wang 0002, Gaël Richard
ICASSP3
2025 A Hybrid Model for Weakly-Supervised Speech Dereverberation
abstract
This paper introduces a new training strategy to improve speech dereverberation systems using minimal acoustic information and reverberant (wet) speech. Most existing algorithms rely on paired dry/wet data, which is difficult to obtain, or on target metrics that may not adequately capture reverberation characteristics and can lead to poor results on non-target metrics. Our approach uses limited acoustic information, like the reverberation time (RT60), to train a dereverberation system. The system’s output is resynthesized using a generated room impulse response and compared with the original reverberant speech, providing a novel reverberation matching loss replacing the standard target metrics. During inference, only the trained dereverberation model is used. Experimental results demonstrate that our method achieves more consistent performance across various objective metrics used in speech dereverberation than the state-of-the-art.
Louis Bahrman, Mathieu Fontaine 0002, Gaël Richard
ICASSP3
2025 Learning Source Disentanglement in Neural Audio Codec
abstract
Neural audio codecs have significantly advanced audio compression by efficiently converting continuous audio signals into discrete tokens. These codecs preserve high-quality sound and enable sophisticated sound generation through generative models trained on these tokens. However, existing neural codec models are typically trained on large, undifferentiated audio datasets, neglecting the essential discrepancies between sound domains like speech, music, and environmental sound effects. This oversight complicates data modeling and poses additional challenges to the controllability of sound generation. To tackle these issues, we introduce the Source-Disentangled Neural Audio Codec (SD-Codec), a novel approach that combines audio coding and source separation. By jointly learning audio resynthesis and separation, SD-Codec explicitly assigns audio signals from different domains to distinct codebooks, sets of discrete representations. Experimental results indicate that SD-Codec not only maintains competitive resynthesis quality but also, supported by the separation results, demonstrates successful disentanglement of different sources in the latent space, thereby enhancing interpretability in audio codec and providing potential finer control over the audio generation process.
Xiaoyu Bie, Gaël Richard
ICASSP3
2025 Investigating the Sensitivity of Pre-trained Audio Embeddings to Common Effects
abstract
In recent years, foundation models have significantly advanced data-driven systems across various domains. Yet, their underlying properties, especially when functioning as feature extractors, remain under-explored. In this paper, we investigate the sensitivity to audio effects of audio embeddings extracted from widely-used foundation models, including OpenL3, PANNs, and CLAP. We focus on audio effects as the source of sensitivity due to their prevalent presence in large audio datasets. By applying parameterized audio effects (gain, low-pass filtering, reverberation, and bitcrushing), we analyze the correlation between the deformation trajectories and the effect strength in the embedding space. We propose to quantify the dimensionality and linearizability of the deformation trajectories induced by audio effects using canonical correlation analysis. We find that there exists a direction along which the embeddings move monotonically as the audio effect strength increases, but that the subspace containing the displacements is generally high-dimensional. This shows that pre-trained audio embeddings do not globally linearize the effects. Our empirical results on instrument classification downstream tasks confirm that projecting out the estimated deformation directions cannot generally improve the robustness of pre-trained embeddings to audio effects.
Victor Deng, Changhong Wang 0002, Gaël Richard, Brian McFee
ICASSP3
2025 Multiple Choice Learning for Efficient Speech Separation with Many Speakers
abstract
Training speech separation models in the supervised setting raises a permutation problem: finding the best assignation between the model predictions and the ground truth separated signals. This inherently ambiguous task is customarily solved using Permutation Invariant Training (PIT). In this article, we instead consider using the Multiple Choice Learning (MCL) framework, which was originally introduced to tackle ambiguous tasks. We demonstrate experimentally on the popular WSJ0-mix and LibriMix benchmarks that MCL matches the performances of PIT, while being computationally advantageous. This opens the door to a promising research direction, as MCL can be naturally extended to handle a variable number of speakers, or to tackle speech separation in the unsupervised setting.
David Perera, François Derrida, Théo Mariotte, Gaël Richard, Slim Essid
ICASSP4
2025 AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder
abstract
This article introduces AnCoGen, a novel method that leverages a masked autoencoder to unify the analysis, control, and generation of speech signals within a single model. AnCoGen can analyze speech by estimating key attributes, such as speaker identity, pitch, content, loudness, signal-to-noise ratio, and clarity index. In addition, it can generate speech from these attributes and allow precise control of the synthesized speech by modifying them. Extensive experiments demonstrated the effectiveness of AnCoGen across speech analysis-resynthesis, pitch estimation, pitch modification, and speech enhancement. Code and audio examples are available online1.
Samir Sadok, Simon Leglaive, Laurent Girin, Gaël Richard, Xavier Alameda-Pineda
ICASSP4
2024 Structure-Informed Positional Encoding for Music Generation
abstract
Music generated by deep learning methods often suffers from a lack of coherence and long-term organization. Yet, multi-scale hierarchical structure is a distinctive feature of music signals. To leverage this information, we propose a structure-informed positional encoding framework for music generation with Transformers. We design three variants in terms of absolute, relative and non-stationary positional information. We comprehensively test them on two symbolic music generation tasks: next-timestep prediction and accompaniment generation. As a comparison, we choose multiple baselines from the literature and demonstrate the merits of our methods using several musically-motivated evaluation metrics. In particular, our methods improve the melodic and structural consistency of the generated pieces.
Manvi Agarwal, Changhong Wang 0002, Gaël Richard
ICASSP3
2024 SpecDiff-GAN: A Spectrally-Shaped Noise Diffusion GAN for Speech and Music Synthesis
abstract
Generative adversarial network (GAN) models can synthesize high-quality audio signals while ensuring fast sample generation. However, they are difficult to train and are prone to several issues including mode collapse and divergence. In this paper, we introduce SpecDiff-GAN, a neural vocoder based on HiFi-GAN, which was initially devised for speech synthesis from mel spectrogram. In our model, the training stability is enhanced by means of a forward diffusion process which consists in injecting noise from a Gaussian distribution to both real and fake samples before inputting them to the discriminator. We further improve the model by exploiting a spectrally-shaped noise distribution with the aim to make the discriminator's task more challenging. We then show the merits of our proposed model for speech and music synthesis on several datasets. Our experiments confirm that our model compares favorably in audio quality and efficiency compared to several baselines.
Teysir Baoueb, Haocheng Liu, Mathieu Fontaine 0002, Jonathan Le Roux, Gaël Richard
ICASSP5
2024 GLA-GRAD: A Griffin-Lim Extended Waveform Generation Diffusion Model
abstract
Diffusion models are receiving a growing interest for a variety of signal generation tasks such as speech or music synthesis. WaveGrad, for example, is a successful diffusion model that conditionally uses the mel spectrogram to guide a diffusion process for the generation of high-fidelity audio. However, such models face important challenges concerning the noise diffusion process for training and inference, and they have difficulty generating high-quality speech for speakers that were not seen during training. With the aim of minimizing the conditioning error and increasing the efficiency of the noise diffusion process, we propose in this paper a new scheme called GLA-Grad, which consists in introducing a phase recovery algorithm such as the Griffin-Lim algorithm (GLA) at each step of the regular diffusion process. Furthermore, it can be directly applied to an already-trained waveform generation model, without additional training or fine-tuning. We show that our algorithm outperforms state-of-the-art diffusion models for speech generation, especially when generating speech for a previously unseen target speaker.
Haocheng Liu, Teysir Baoueb, Mathieu Fontaine 0002, Jonathan Le Roux, Gaël Richard
ICASSP5
2024 A Fully Differentiable Model for Unsupervised Singing Voice Separation
abstract
A novel model was recently proposed by Schulze-Forster et al. in [1] for unsupervised music source separation. This model allows to tackle some of the major shortcomings of existing source separation frameworks. Specifically, it eliminates the need for isolated sources during training, performs efficiently with limited data, and can handle homogeneous sources (such as singing voice). But, this model relies on an external multipitch estimator and incorporates an Ad hoc voice assignment procedure. In this paper, we propose to extend this framework and to build a fully differentiable model by integrating a multipitch estimator and a novel differentiable assignment module within the core model. We show the merits of our approach through a set of experiments, and we highlight in particular its potential for processing diverse and unseen data.
Gaël Richard, Pierre Chouteau, Bernardo Torres
ICASSP1
2024 Unsupervised Harmonic Parameter Estimation Using Differentiable DSP and Spectral Optimal Transport
abstract
In neural audio signal processing, pitch conditioning has been used to enhance the performance of synthesizers. However, jointly training pitch estimators and synthesizers is a challenge when using standard audio-to-audio reconstruction loss, leading to reliance on external pitch trackers. To address this issue, we propose using a spectral loss function inspired by optimal transportation theory that minimizes the displacement of spectral energy. We validate this approach through an unsupervised autoencoding task that fits a harmonic template to harmonic signals. We jointly estimate the fundamental frequency and amplitudes of harmonics using a lightweight encoder and reconstruct the signals using a differentiable harmonic synthesizer. The proposed approach offers a promising direction for improving unsupervised parameter estimation in neural audio applications.
Bernardo Torres, Geoffroy Peeters, Gaël Richard
ICASSP3
2024 Winner-takes-all learners are geometry-aware conditional density estimators
abstract
Winner-takes-all training is a simple learning paradigm, which handles ambiguous tasks by predicting a set of plausible hypotheses. Recently, a connection was established between Winner-takes-all training and centroidal Voronoi tessellations, showing that, once trained, hypotheses should quantize optimally the shape of the conditional distribution to predict. However, the best use of these hypotheses for uncertainty quantification is still an open question. In this work, we show how to leverage the appealing geometric properties of the Winner-takes-all learners for conditional density estimation, without modifying its original training scheme. We theoretically establish the advantages of our novel estimator both in terms of quantization and density estimation, and we demonstrate its competitiveness on synthetic and real-world datasets, including audio data.
Victor Letzelter, David Perera, Cédric Rommel, Mathieu Fontaine 0002, Slim Essid, Gaël Richard, Patrick Pérez
ICML6
2024 Speech dereverberation constrained on room impulse response characteristics
Louis Bahrman, Mathieu Fontaine 0002, Jonathan Le Roux, Gaël Richard
INTERSPEECH4
2024 Annealed Multiple Choice Learning: Overcoming limitations of Winner-takes-all with annealing
abstract
We introduce Annealed Multiple Choice Learning (aMCL) which combines simulated annealing with MCL. MCL is a learning framework handling ambiguous tasks by predicting a small set of plausible hypotheses. These hypotheses are trained using the Winner-takes-all (WTA) scheme, which promotes the diversity of the predictions. However, this scheme may converge toward an arbitrarily suboptimal local minimum, due to the greedy nature of WTA. We overcome this limitation using annealing, which enhances the exploration of the hypothesis space during training. We leverage insights from statistical physics and information theory to provide a detailed description of the model training trajectory. Additionally, we validate our algorithm by extensive experiments on synthetic datasets, on the standard UCI benchmark, and on speech separation.
David Perera, Victor Letzelter, Théo Mariotte, Adrien Cortés, Mickaël Chen, Slim Essid, Gaël Richard
NeurIPS7
2024 Tackling Interpretability in Audio Classification Networks With Non-negative Matrix Factorization
abstract
This paper tackles two major problem settings for interpretability of audio processing networks,post-hocandby-designinterpretation. For post-hoc interpretation, we aim to interpret decisions of a network in terms of high-level audio objects that are also listenable for the end-user. This is extended to present an inherently interpretable model with high performance. To this end, we propose a novel interpreter design that incorporates non-negative matrix factorization (NMF). In particular, an interpreter is trained to generate a regularized intermediate embedding from hidden layers of a target network, learnt as time-activations of a pre-learnt NMF dictionary. Our methodology allows us to generate intuitive audio-based interpretations that explicitly enhance parts of the input signal most relevant for a network's decision. We demonstrate our method's applicability on a variety of classification tasks, including multi-label data for real-world audio and music.
Jayneel Parekh, Sanjeel Parekh, Pavlo Mozharovskyi, Gaël Richard, Florence d'Alché-Buc
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Learning Interpretable Filters In Wav-UNet For Speech Enhancement
abstract
Due to their performances, deep neural networks have emerged as a major method in nearly all modern audio processing applications. Deep neural networks can be used to estimate some parameters or hyperparameters of a model, or in some cases the entire model in an end-to-end fashion. Although deep learning can lead to state of the art performances, they also suffer from inherent weaknesses as they usually remain complex and non interpretable to a large extent. For instance, the internal filters used in each layers are chosen in an adhoc manner with only a loose relation with the nature of the processed signal. We propose in this paper an approach to learn interpretable filters within a specific neural architecture which allow to better understand the behaviour of the neural network and to reduce its complexity. We validate the approach on a task of speech enhancement and show that the gain in interpretability does not degrade the performance of the model.
Félix Mathieu, Thomas Courtat, Gaël Richard, Geoffroy Peeters
ICASSP3
2023 Resilient Multiple Choice Learning: A learned scoring scheme with application to audio scene analysis
abstract
We introduce Resilient Multiple Choice Learning (rMCL), an extension of the MCL approach for conditional distribution estimation in regression settings where multiple targets may be sampled for each training input. Multiple Choice Learning is a simple framework to tackle multimodal density estimation, using the Winner-Takes-All (WTA) loss for a set of hypotheses. In regression settings, the existing MCL variants focus on merging the hypotheses, thereby eventually sacrificing the diversity of the predictions. In contrast, our method relies on a novel learned scoring scheme underpinned by a mathematical framework based on Voronoi tessellations of the output space, from which we can derive a probabilistic interpretation. After empirically validating rMCL with experiments on synthetic data, we further assess its merits on the sound source localization problem, demonstrating its practical usefulness and the relevance of its interpretation.
Victor Letzelter, Mathieu Fontaine 0002, Mickaël Chen, Patrick Pérez, Slim Essid, Gaël Richard
NeurIPS6
2023 Unsupervised Music Source Separation Using Differentiable Parametric Source Models
abstract
Supervised deep learning approaches to underdetermined audio source separation achieve state-of-the-art performance but require a dataset of mixtures along with their corresponding isolated source signals. Such datasets can be extremely costly to obtain for musical mixtures. This raises a need for unsupervised methods. We propose a novel unsupervised model-based deep learning approach to musical source separation. Each source is modelled with a differentiable parametric source-filter model. A neural network is trained to reconstruct the observed mixture as a sum of the sources by estimating the source models' parameters given their fundamental frequencies. At test time, soft masks are obtained from the synthesized source signals. The experimental evaluation on a vocal ensemble separation task shows that the proposed method outperforms learning-free methods based on nonnegative matrix factorization and a supervised deep learning baseline. Integrating domain knowledge in the form of source models into a data-driven method leads to high data efficiency: the proposed approach achieves good separation quality even when trained on less than three minutes of audio. This work makes powerful deep learning based separation usable in scenarios where training data with ground truth is expensive or nonexistent.
Kilian Schulze-Forster, Gaël Richard, Liam Kelley, Clement S. J. Doire, Roland Badeau
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Video-to-Music Recommendation Using Temporal Alignment of Segments
abstract
We study cross-modal recommendation of musictracks to be used as soundtracks for videos. This problem is known as the music supervision task. We build on a self-supervised system that learns a content association between music and video. In addition to the adequacy of content, adequacy of structure is crucial in music supervision to obtain relevant recommendations. We propose a novel approach to significantly improve the system’s performance using structure-aware recommendation. The core idea is to consider not only the full audio-video clips, but rather shorter segments for training and inference. We find that using semantic segments and ranking the tracks according to sequence alignment costs significantly improves the results. We investigate the impact of different ranking metrics and segmentation methods.
Laure Prétet, Gaël Richard, Clément Souchier, Geoffroy Peeters
IEEE Trans. Multim.2
2022 Rate-Distortion Theoretic Generalization Bounds for Stochastic Learning Algorithms
abstract
Understanding generalization in modern machine learning settings has been one of the major challenges in statistical learning theory. In this context, recent years have witnessed the development of various generalization bounds suggesting different complexity notions such as the mutual information between the data sample and the algorithm output, compressibility of the hypothesis space, and the fractal dimension of the hypothesis space. While these bounds have illuminated the problem at hand from different angles, their suggested complexity notions might appear seemingly unrelated, thereby restricting their high-level impact. In this study, we prove novel generalization bounds through the lens of rate-distortion theory, and explicitly relate the concepts of mutual information, compressibility, and fractal dimensions in a single mathematical framework. Our approach consists of (i) defining a generalized notion of compressibility by using source coding concepts, and (ii) showing that the ’compression error rate’ can be linked to the generalization error both in expectation and with high probability. We show that in the ’lossless compression’ setting, we recover and improve existing mutual information-based bounds, whereas a ’lossy compression’ scheme allows us to link generalization to the rate-distortion dimension - a particular notion of fractal dimension. Our results bring a more unified perspective on generalization and open up several future research directions.
Milad Sefidgaran, Amin Gohari, Gaël Richard, Umut Simsekli
COLT3
2022 Phase Shifted Bedrosian Filterbank: An Interpretable Audio Front-End for Time-Domain Audio Source Separation
abstract
The use of a parameterized encoders or audio front-ends has shown promises in improving the interpretability of time domain single-channel source separation models such as Conv-TasNet. This type of filters also allows a potential reduction of the computational cost since larger encoder filters can be used. In this work, we propose to build a new parameterization of such encoder filter-bank which allows gaining interpretability while keeping flexibility. Based on the Hilbert transform and the Bedrosian theorem, we propose to build phase-shifted set of filters by modulating sinusoids through freely learned low pass filters. We show that the use of these filters allows to keep the same performances when using small filters and even improve them when using large filters.
Félix Mathieu, Thomas Courtat, Gaël Richard, Geoffroy Peeters
ICASSP3
2022 Listen to Interpret: Post-hoc Interpretability for Audio Networks with NMF
abstract
This paper tackles post-hoc interpretability for audio processing networks. Our goal is to interpret decisions of a trained network in terms of high-level audio objects that are also listenable for the end-user. To this end, we propose a novel interpreter design that incorporates non-negative matrix factorization (NMF). In particular, a regularized interpreter module is trained to take hidden layer representations of the targeted network as input and produce time activations of pre-learnt NMF components as intermediate outputs. Our methodology allows us to generate intuitive audio-based interpretations that explicitly enhance parts of the input signal most relevant for a network's decision. We demonstrate our method's applicability on popular benchmarks, including a real-world multi-label classification task.
Jayneel Parekh, Sanjeel Parekh, Pavlo Mozharovskyi, Florence d'Alché-Buc, Gaël Richard
NeurIPS5
2021 Neuro-Steered Music Source Separation With EEG-Based Auditory Attention Decoding And Contrastive-NMF
abstract
We propose a novel informed music source separation paradigm, which can be referred to as neuro-steered music source separation. More precisely, the source separation process is guided by the user’s selective auditory attention decoded from his/her EEG response to the stimulus. This high-level prior information is used to select the desired instrument to isolate and to adapt the generic source separation model to the observed signal. To this aim, we leverage the fact that the attended instrument’s neural encoding is substantially stronger than the one of the unattended sources left in the mixture. This "contrast" is extracted using an attention decoder and used to inform a source separation model based on non-negative matrix factorization named Contrastive-NMF. The results are promising and show that the EEG information can automatically select the desired source to enhance and improve the separation quality.
Giorgia Cantisani, Slim Essid, Gaël Richard
ICASSP3
2021 Self-Supervised VQ-VAE for One-Shot Music Style Transfer
abstract
Neural style transfer, allowing to apply the artistic style of one image to another, has become one of the most widely showcased computer vision applications shortly after its introduction. In contrast, related tasks in the music audio domain remained, until recently, largely untackled. While several style conversion methods tailored to musical signals have been proposed, most lack the ‘one-shot’ capability of classical image style transfer algorithms. On the other hand, the results of existing one-shot audio style transfer methods on musical inputs are not as compelling. In this work, we are specifically interested in the problem of one-shot timbre transfer. We present a novel method for this task, based on an extension of the vector-quantized variational autoencoder (VQ-VAE), along with a simple self-supervised learning strategy designed to obtain disentangled representations of timbre and pitch. We evaluate the method using a set of objective metrics and show that it is able to outperform selected baselines.
Ondrej Cífka, Alexey Ozerov, Umut Simsekli, Gaël Richard
ICASSP4
2021 Relative Positional Encoding for Transformers with Linear Complexity
abstract
Recent advances in Transformer models allow for unprecedented sequence lengths, due to linear space and time complexity. In the meantime, relative positional encoding (RPE) was proposed as beneficial for classical Transformers and consists in exploiting lags instead of absolute positions for inference. Still, RPE is not available for the recent linear-variants of the Transformer, because it requires the explicit computation of the attention matrix, which is precisely what is avoided by such methods. In this paper, we bridge this gap and present Stochastic Positional Encoding as a way to generate PE that can be used as a replacement to the classical additive (sinusoidal) PE and provably behaves like RPE. The main theoretical contribution is to make a connection between positional encoding and cross-covariance structures of correlated Gaussian processes. We illustrate the performance of our approach on the Long-Range Arena benchmark and on music generation.
Antoine Liutkus, Ondrej Cífka, Shih-Lun Wu, Umut Simsekli, Yi-Hsuan Yang, Gaël Richard
ICML6
2021 Cross-Modal Music-Video Recommendation: A Study of Design Choices
abstract
In this work, we study music/video cross-modal recommendation, i.e. recommending a music track for a video or vice versa. We rely on a self-supervised learning paradigm to learn from a large amount of unlabelled data. We rely on a self-supervised learning paradigm to learn from a large amount of unlabelled data. More precisely, we jointly learn audio and video embeddings by using their co-occurrence in music-video clips. In this work, we build upon a recent video-music retrieval system (the VM-NET), which originally relies on an audio representation obtained by a set of statistics computed over handcrafted features. We demonstrate here that using audio representation learning such as the audio embeddings provided by the pre-trained MuSimNet, OpenL3, MusicCNN or by AudioSet, largely improves recommendations. We also validate the use of the cross-modal triplet loss originally proposed in the VM-NET compared to the binary cross-entropy loss commonly used in self-supervised learning. We perform all our experiments using the Music Video Dataset (MVD).
Laure Prétet, Gaël Richard, Geoffroy Peeters
IJCNN2
2021 Heavy Tails in SGD and Compressibility of Overparametrized Neural Networks
abstract
Neural network compression techniques have become increasingly popular as they can drastically reduce the storage and computation requirements for very large networks. Recent empirical studies have illustrated that even simple pruning strategies can be surprisingly effective, and several theoretical studies have shown that compressible networks (in specific senses) should achieve a low generalization error. Yet, a theoretical characterization of the underlying causes that make the networks amenable to such simple compression schemes is still missing. In this study, focusing our attention on stochastic gradient descent (SGD), our main contribution is to link compressibility to two recently established properties of SGD: (i) as the network size goes to infinity, the system can converge to a mean-field limit, where the network weights behave independently [DBDFŞ20], (ii) for a large step-size/batch-size ratio, the SGD iterates can converge to a heavy-tailed stationary distribution [HM20, GŞZ21]. Assuming that both of these phenomena occur simultaneously, we prove that the networks are guaranteed to be '$\ell_p$-compressible', and the compression errors of different pruning techniques (magnitude, singular value, or node pruning) become arbitrarily small as the network size increases. We further prove generalization bounds adapted to our theoretical framework, which are consistent with the observation that the generalization error will be lower for more compressible networks. Our theory and numerical study on various neural networks show that large step-size/batch-size ratios introduce heavy tails, which, in combination with overparametrization, result in compressibility.
Melih Barsbey, Milad Sefidgaran, Murat A. Erdogdu, Gaël Richard, Umut Simsekli
NeurIPS4
2021 Phoneme Level Lyrics Alignment and Text-Informed Singing Voice Separation
abstract
The goal of singing voice separation is to recover the vocals signal from music mixtures. State-of-the-art performance is achieved by deep neural networks trained in a supervised fashion. Since training data are scarce and music signals are extremely diverse, it remains challenging to achieve high separation quality across various recording and mixing conditions as well as music styles. In this paper, we investigate to which extent the separation can be improved when lyrics transcripts are used as additional information. To this end, we propose a joint approach to phoneme level lyrics alignment and text-informed singing voice separation. It is based on DTW-attention, a new monotonic attention mechanism including a differentiable approximation of dynamic time warping. Experimental results show that the method can align phonemes with mixed singing voice with high precision given accurate transcripts. It also achieves competitive results on challenging word level alignment test sets using less training data than state-of-the-art methods. Sequential alignment and informed separation lead to improved separation quality according to objective measures. Text information helps preserving spectral phoneme properties in the separated voice signals.
Kilian Schulze-Forster, Clement S. J. Doire, Gaël Richard, Roland Badeau
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Speech Intelligibility Enhancement by Equalization for in-Car Applications
abstract
In this paper, we propose a speech intelligibility enhancement method for typical in-car applications in noisy environments. While traditional speech enhancement algorithms aim at increasing the Signal to Noise Ratio (SNR), the goal here is to increase intelligibility by applying dedicated voice transformation techniques without changing the original SNR. The proposed method consists in an adaptive equalizer which reallocates the energy of frequency bands to maximize the Speech Intelligibility Index (SII) under the constraint of a fixed perceived loudness. The validation of the algorithm is carried out by means of a perceptual test derived from the Hearing in Noise Test (HINT) using four typical in-car noises of different driving conditions. The results obtained demonstrate the merit of the algorithm for low-frequency noises, that correspond to usual driving conditions, but also show the limit of the algorithm on noises with a spectrum more spread out induced by rain.
Enguerrand Gentet, Bertrand David 0002, Sébastien Denjean, Gaël Richard, Vincent Roussarie
ICASSP4
2020 Neutral to Lombard Speech Conversion with Deep Learning
abstract
In this paper, we propose several approaches for neutral to Lombard speech conversion. We study in particular the influence of different recurrent neural network architectures where their main hyper-parameters are carefully selected using a bandit-based approach. We also apply the Continuous Wavelet Transform (CWT) as a multi-resolution analysis framework to better model temporal dependencies of the different features selected. The speech conversion results obtained are validated by means of objective evaluations which highlight in particular the interest of the wavelet transform for the learning process.
Enguerrand Gentet, Bertrand David 0002, Sébastien Denjean, Gaël Richard, Vincent Roussarie
ICASSP4
2020 Audio-Based Auto-Tagging With Contextual Tags for Music
abstract
Music listening context such as location or activity has been shown to greatly influence the users' musical tastes. In this work, we study the relationship between user context and audio content in order to enable context-aware music recommendation agnostic to user data. For that, we propose a semi-automatic procedure to collect track sets which leverages playlist titles as a proxy for context labelling. Using this, we create and release a dataset of ~50k tracks labelled with 15 different contexts. Then, we present benchmark classification results on the created dataset using an audio auto-tagging model. As the training and evaluation of these models are impacted by missing negative labels due to incomplete annotations, we propose a sample-level weighted cross entropy loss to account for the confidence in missing labels and show improved context prediction results.
Karim M. Ibrahim, Jimena Royo-Letelier, Elena V. Epure, Geoffroy Peeters, Gaël Richard
ICASSP5
2020 Learning to Rank Music Tracks Using Triplet Loss
abstract
Most music streaming services rely on automatic recommendation algorithms to exploit their large music catalogs. These algorithms aim at retrieving a ranked list of music tracks based on their similarity with a target music track. In this work, we propose a method for direct recommendation based on the audio content without explicitly tagging the music tracks. To that aim, we propose several strategies to perform triplet mining from ranked lists. We train a Convolutional Neural Network to learn the similarity via triplet loss. These different strategies are compared and validated on a large-scale experiment against an auto-tagging based approach. The results obtained highlight the efficiency of our system, especially when associated with an Auto-pooling layer.
Laure Prétet, Gaël Richard, Geoffroy Peeters
ICASSP2
2020 Joint Phoneme Alignment and Text-Informed Speech Separation on Highly Corrupted Speech
abstract
Speech separation quality can be improved by exploiting textual information. However, this usually requires text-to-speech alignment at phoneme level. Classical alignment methods are made for rather clean speech and do not work as well on corrupted speech. We propose to perform text-informed speech-music separation and phoneme alignment jointly using recurrent neural networks and the attention mechanism. We show that it leads to benefits for both tasks. In experiments, phoneme transcripts are used to improve the perceived quality of separated speech over a non-informed baseline. Moreover, our novel phoneme alignment method based on the attention mechanism achieves state-of-the-art alignment accuracy on clean and on heavily corrupted speech.
Kilian Schulze-Forster, Clement S. J. Doire, Gaël Richard, Roland Badeau
ICASSP3
2020 Audio-Based Detection of Explicit Content in Music
abstract
We present a novel automatic system for performing explicit content detection directly on the audio signal. Our modular approach uses an audio-to-character recognition model, a keyword spotting model associated with a dictionary of carefully chosen keywords, and a Random Forest classification model for the final decision. To the best of our knowledge, this is the first explicit content detection system based on audio only. We demonstrate the individual relevance of our modules on a set of sub-tasks and compare our approach to a lyrics-informed oracle and an end-to-end naive architecture. The results obtained are encouraging with a F1-score of 67% on a industrial scale explicit content dataset.
Andrea Vaglio, Romain Hennequin, Manuel Moussallam, Gaël Richard, Florence d'Alché-Buc
ICASSP4
2020 The POTUS Corpus, a Database of Weekly Addresses for the Study of Stance in Politics and Virtual Agents
abstract
One of the main challenges in the field of Embodied Conversational Agent (ECA) is to generate socially believable agents. The common strategy for agent behaviour synthesis is to rely on dedicated corpus analysis. Such a corpus is composed of multimedia files of socio-emotional behaviors which have been annotated by external observers. The underlying idea is to identify interaction information for the agent’s socio-emotional behavior by checking whether the intended socio-emotional behavior is actually perceived by humans. Then, the annotations can be used as learning classes for machine learning algorithms applied to the social signals. This paper introduces the POTUS Corpus composed of high-quality audio-video files of political addresses to the American people. Two protagonists are present in this database. First, it includes speeches of former president Barack Obama to the American people. Secondly, it provides videos of these same speeches given by a virtual agent named Rodrigue. The ECA reproduces the original address as closely as possible using social signals automatically extracted from the original one. Both are annotated for social attitudes, providing information about the stance observed in each file. It also provides the social signals automatically extracted from Obama’s addresses used to generate Rodrigue’s ones.
Thomas Janssoone, Kevin Bailly, Gaël Richard, Chloé Clavel
LREC3
2020 Confidence-based Weighted Loss for Multi-label Classification with Missing Labels
abstract
The problem of multi-label classification with missing labels (MLML) is a common challenge that is prevalent in several domains, e.g. image annotation and auto-tagging. In multi-label classification, each instance may belong to multiple class labels simultaneously. Due to the nature of the dataset collection and labelling procedure, it is common to have incomplete annotations in the dataset, i.e. not all samples are labelled with all the corresponding labels. However, the incomplete data labelling hinders the training of classification models. MLML has received much attention from the research community. However, in cases where a pre-trained model is fine-tuned on an MLML dataset, there has been no straightforward approach to tackle the missing labels, specifically when there is no information about which are the missing ones. In this paper, we propose a weighted loss function to account for the confidence in each label/sample pair that can easily be incorporated to fine-tune a pre-trained model on an incomplete dataset. Our experiment results show that using the proposed loss function improves the performance of the model as the ratio of missing labels increases.
Karim M. Ibrahim, Elena V. Epure, Geoffroy Peeters, Gaël Richard
ICMR4
2020 Groove2Groove: One-Shot Music Style Transfer With Supervision From Synthetic Data
abstract
Style transfer is the process of changing the style of an image, video, audio clip or musical piece so as to match the style of a given example. Even though the task has interesting practical applications within the music industry, it has so far received little attention from the audio and music processing community. In this article, we present Groove2Groove, a one-shot style transfer method for symbolic music, focusing on the case of accompaniment styles in popular music and jazz. We propose an encoder-decoder neural network for the task, along with a synthetic data generation scheme to supply it with parallel training examples. This synthetic parallel data allows us to tackle the style transfer problem using end-to-end supervised learning, employing powerful techniques used in natural language processing. We experimentally demonstrate the performance of the model on style transfer using existing and newly proposed metrics, and also explore the possibility of style interpolation.
Ondrej Cífka, Umut Simsekli, Gaël Richard
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Weakly Supervised Representation Learning for Audio-Visual Scene Analysis
abstract
Audio-visual (AV) representation learning is an important task from the perspective of designing machines with the ability to understand complex events. To this end, we propose a novel multimodal framework that instantiates multiple instance learning. Specifically, we develop methods that identify events and localize corresponding AV cues in unconstrained videos. Importantly, this is done using weak labels where only video-level event labels are known without any information about their location in time. We show that the learnt representations are useful for performing several tasks such as event/object classification, audio event detection, audio source separation and visual object localization. An important feature of our method is its capacity to learn from unsynchronized audio-visual events. We also demonstrate our framework's ability to separate out the audio source of interest through a novel use of nonnegative matrix factorization. State-of-the-art classification results, with a F1-score of 65.0, are achieved on DCASE 2017 smart cars challenge data with promising generalization to diverse object types such as musical instruments. Visualizations of localized visual regions and audio segments substantiate our system's efficacy, especially when dealing with noisy situations where modality-specific cues appear asynchronously.
Sanjeel Parekh, Slim Essid, Alexey Ozerov, Ngoc Q. K. Duong, Patrick Pérez, Gaël Richard
IEEE ACM Trans. Audio Speech Lang. Process.6
2019 Non-Asymptotic Analysis of Fractional Langevin Monte Carlo for Non-Convex Optimization
abstract
Recent studies on diffusion-based sampling methods have shown that Langevin Monte Carlo (LMC) algorithms can be beneficial for non-convex optimization, and rigorous theoretical guarantees have been proven for both asymptotic and finite-time regimes. Algorithmically, LMC-based algorithms resemble the well-known gradient descent (GD) algorithm, where the GD recursion is perturbed by an additive Gaussian noise whose variance has a particular form. Fractional Langevin Monte Carlo (FLMC) is a recently proposed extension of LMC, where the Gaussian noise is replaced by a heavy-tailed $\alpha$-stable noise. As opposed to its Gaussian counterpart, these heavy-tailed perturbations can incur large jumps and it has been empirically demonstrated that the choice of $\alpha$-stable noise can provide several advantages in modern machine learning problems, both in optimization and sampling contexts. However, as opposed to LMC, only asymptotic convergence properties of FLMC have been yet established. In this study, we analyze the non-asymptotic behavior of FLMC for non-convex optimization and prove finite-time bounds for its expected suboptimality. Our results show that the weak-error of FLMC increases faster than LMC, which suggests using smaller step-sizes in FLMC. We finally extend our results to the case where the exact gradients are replaced by stochastic gradients and show that similar results hold in this setting as well.
Thanh Huy Nguyen 0001, Umut Simsekli, Gaël Richard
ICML3
2019 First Exit Time Analysis of Stochastic Gradient Descent Under Heavy-Tailed Gradient Noise
abstract
Stochastic gradient descent (SGD) has been widely used in machine learning due to its computational efficiency and favorable generalization properties. Recently, it has been empirically demonstrated that the gradient noise in several deep learning settings admits a non-Gaussian, heavy-tailed behavior. This suggests that the gradient noise can be modeled by using $\alpha$-stable distributions, a family of heavy-tailed distributions that appear in the generalized central limit theorem. In this context, SGD can be viewed as a discretization of a stochastic differential equation (SDE) driven by a L\'{e}vy motion, and the metastability results for this SDE can then be used for illuminating the behavior of SGD, especially in terms of `preferring wide minima'. While this approach brings a new perspective for analyzing SGD, it is limited in the sense that, due to the time discretization, SGD might admit a significantly different behavior than its continuous-time limit. Intuitively, the behaviors of these two systems are expected to be similar to each other only when the discretization step is sufficiently small; however, to the best of our knowledge, there is no theoretical understanding on how small the step-size should be chosen in order to guarantee that the discretized system inherits the properties of the continuous-time system. In this study, we provide formal theoretical analysis where we derive explicit conditions for the step-size such that the metastability behavior of the discrete-time system is similar to its continuous-time limit. We show that the behaviors of the two systems are indeed similar for small step-sizes and we identify how the error depends on the algorithm and problem parameters. We illustrate our results with simulations on a synthetic model and neural networks.
Thanh Huy Nguyen 0001, Umut Simsekli, Mert Gürbüzbalaban, Gaël Richard
NeurIPS4
2019 Independent-Variation Matrix Factorization With Application to Energy Disaggregation
abstract
Matrix factorization techniques have proven to be useful in many unsupervised learning applications. Such techniques have been recently applied to Non Intrusive Load Monitoring (NILM), the process of breaking down the total electric consumption of a building into consumptions of individual appliances. While several studies addressed the NILM problem for small-scale buildings, only few studies considered the problem for large buildings, where the signals exhibit significantly different behavior. To overcome the unaddressed difficulties of processing high frequency current signals that are measured in large buildings, we propose a novel technique called Independent-Variation Matrix Factorization (IVMF), which expresses an observation matrix as the product of two matrices: the signature and the activation. Motivated by the nature of the current signals, it uses a regularization term on the temporal variations of the activation matrix and a positivity constraint, and the columns of the signature matrix are constrained to lie in a specific set. To solve the resulting optimization problem, we rely on an alternating minimization strategy involving dual optimization and quasi-Newton algorithms. The algorithm is tested against Independent Component Analysis (ICA) and Semi Nonnegative Matrix Factorization (SNMF) on a synthetic source separation problem and on a realistic NILM application for large commercial buildings. We show that IVMF outperforms competing methods and is particularly appropriate to recover positive sources that have a strong temporal dependency and sources whose variations are independent from each other.
Simon Henriet, Umut Simsekli, Sergio Dos Santos, Benoit Fuentes, Gaël Richard
IEEE Signal Process. Lett.5
2018 Alpha-Stable Low-Rank Plus Residual Decomposition for Speech Enhancement
abstract
In this study, we propose a novel probabilistic model for separating clean speech signals from noisy mixtures by decomposing the mixture spectra into a structured speech part and a more flexible residual part. The main novelty in our model is that it uses a family of heavy-tailed distributions, so called the α-stable distributions, for modeling the residual signal. We develop an expectation-maximization algorithm for parameter estimation and a Monte Carlo scheme for posterior estimation of the clean speech. Our experiments show that the proposed method outperforms relevant factorization-based algorithms by a significant margin.
Umut Simsekli, Halil Erdogan, Simon Leglaive, Antoine Liutkus, Roland Badeau, Gaël Richard
ICASSP6
2018 Asynchronous Stochastic Quasi-Newton MCMC for Non-Convex Optimization
abstract
Recent studies have illustrated that stochastic gradient Markov Chain Monte Carlo techniques have a strong potential in non-convex optimization, where local and global convergence guarantees can be shown under certain conditions. By building up on this recent theory, in this study, we develop an asynchronous-parallel stochastic L-BFGS algorithm for non-convex optimization. The proposed algorithm is suitable for both distributed and shared-memory settings. We provide formal theoretical analysis and show that the proposed method achieves an ergodic convergence rate of ${\cal O}(1/\sqrt{N})$ ($N$ being the total number of iterations) and it can achieve a linear speedup under certain conditions. We perform several experiments on both synthetic and real datasets. The results support our theory and show that the proposed algorithm provides a significant speedup over the recently proposed synchronous distributed L-BFGS algorithm.
Umut Simsekli, Çagatay Yildiz, Thanh Huy Nguyen 0001, A. Taylan Cemgil, Gaël Richard
ICML5
2018 Efficient Bayesian Model Selection in PARAFAC via Stochastic Thermodynamic Integration
abstract
Parallel factor analysis (PARAFAC) is one of the most popular tensor factorization models. Even though it has proven successful in diverse application fields, the performance of PARAFAC usually hinges up on the rank of the factorization, which is typically specified manually by the practitioner. In this study, we develop a novel parallel and distributed Bayesian model selection technique for rank estimation in large-scale PARAFAC models. The proposed approach integrates ideas from the emerging field of stochastic gradient Markov Chain Monte Carlo, statistical physics, and distributed stochastic optimization. As opposed to the existing methods, which are based on some heuristics, our method has a clear mathematical interpretation, and has significantly lower computational requirements, thanks to data subsampling and parallelization. We provide formal theoretical analysis on the bias induced by the proposed approach. Our experiments on synthetic and large-scale real datasets show that our method is able to find the optimal model order while being significantly faster than the state-of-the-art.
Thanh Huy Nguyen 0001, Umut Simsekli, Gaël Richard, A. Taylan Cemgil
IEEE Signal Process. Lett.3
2018 Hybrid Projective Nonnegative Matrix Factorization With Drum Dictionaries for Harmonic/Percussive Source Separation
abstract
One of the most general models of music signals considers that such signals can be represented as a sum of two distinct components: a tonal part that is sparse in frequency and temporally stable and a transient (or percussive) part that is composed of short-term broadband sounds. In this paper, we propose a novel hybrid method built upon nonnegative matrix factorization (NMF) that decomposes the time frequency representation of an audio signal into such two components. The tonal part is estimated by a sparse and orthogonal nonnegative decomposition, and the transient part is estimated by a straightforward NMF decomposition constrained by a pre-learned dictionary of smooth spectra. The optimization problem at the heart of our method remains simple with very few hyperparameters and can be solved thanks to simple multiplicative update rules. The extensive benchmark on a large and varied music database against four state of the art harmonic/percussive source separation algorithms demonstrate the merit of the proposed approach.
Clement Laroche, Matthieu Kowalski, Hélène Papadopoulos, Gaël Richard
IEEE ACM Trans. Audio Speech Lang. Process.4
2018 Student's t Source and Mixing Models for Multichannel Audio Source Separation
abstract
This paper presents a Bayesian framework for under-determined audio source separation in multichannel reverberant mixtures. We model the source signals as Student's t latent random variables in a time-frequency domain. The specific structure of musical signals in this domain is exploited by means of a nonnegative matrix factorization model. Conversely, we design the mixing model in the time domain. In addition to leading to an exact representation of the convolutive mixing process, this approach allows us to develop simple probabilistic priors for the mixing filters. Indeed, as those filters correspond to room responses they exhibit a simple characteristic structure in the time domain that can be used to guide their estimation. We also rely on the Student's t distribution for modeling the impulse response of the mixing filters. From this model, we develop a variational inference algorithm in order to perform source separation. The experimental evaluation demonstrates the potential of this approach for separating multichannel reverberant mixtures.
Simon Leglaive, Roland Badeau, Gaël Richard
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 Overlapping sound event detection with supervised Nonnegative Matrix Factorization
abstract
In this paper we propose a supervised Nonnegative Matrix Factorization (NMF) model for overlapping sound event detection in real life audio. We start by highlighting the usefulness of non-euclidean NMF to learn representations for detecting and classifying acoustic events in a multi-label setting. Then, we propose to learn a classifier and the NMF decomposition in a joint optimization problem. This is done with a general β-divergence version of the nonnegative task-driven dictionary learning model. An experimental evaluation is performed on the development set of the DCASE 2016 task3 challenge. The proposed supervised NMF-based system improves performance over the baseline and the submitted systems.
Victor Bisot, Slim Essid, Gaël Richard
ICASSP3
2017 Drum extraction in single channel audio signals using multi-layer Non negative Matrix Factor Deconvolution
abstract
In this paper, we propose a supervised multilayer factorization method designed for harmonic/percussive source separation and drum extraction. Our method decomposes the audio signals in sparse orthogonal components which capture the harmonic content, while the drum is represented by an extension of non negative matrix factorization which is able to exploit time-frequency dictionaries to take into account non stationary drum sounds. The drum dictionaries represent various real drum hits and the decomposition has more physical sense and allows for a better interpretation of the results. Experiments on real music data for a harmonic/percussive source separation task show that our method outperforms other state of the art algorithms. Finally, our method is very robust to non stationary harmonic sources that are usually poorly decomposed by existing methods.
Clement Laroche, Hélène Papadopoulos, Matthieu Kowalski, Gaël Richard
ICASSP4
2017 Multichannel audio source separation: Variational inference of time-frequency sources from time-domain observations
abstract
A great number of methods for multichannel audio source separation are based on probabilistic approaches in which the sources are modeled as latent random variables in a Time-Frequency (TF) domain. For reverberant mixtures, it is common to approximate the time-domain convolutive mixing process as being instantaneous in the short-term Fourier transform domain, under a short mixing filters assumption. The TF latent sources are then inferred from the TF mixture observations. In this paper we propose to infer the TF latent sources from the time-domain observations. This approach allows us to exactly model the convolutive mixing process. The inference procedure relies on a variational expectation-maximization algorithm. In significant reverberation conditions, our approach leads to a signal-to-distortion ratio improvement of 5.5 dB compared with the usual TF approximation of the convolutive mixing process.
Simon Leglaive, Roland Badeau, Gaël Richard
ICASSP3
2017 Alpha-stable multichannel audio source separation
abstract
In this paper, we focus on modeling multichannel audio signals in the short-time Fourier transform domain for the purpose of source separation. We propose a probabilistic model based on a class of heavy-tailed distributions, in which the observed mixtures and the latent sources are jointly modeled by using a certain class of multivariate alpha-stable distributions. As opposed to the conventional Gaussian models, where the observations are constrained to lie just within a few standard deviations from the mean, the proposed heavy-tailed model allows us to account for spurious data or important uncertainties in the model. We develop a Monte Carlo Expectation-Maximization algorithm for inferring the sources from the proposed model. We show that our approach leads to significant performance improvements in audio source separation under corrupted mixtures and in spatial audio object coding.
Simon Leglaive, Umut Simsekli, Antoine Liutkus, Roland Badeau, Gaël Richard
ICASSP5
2017 Motion informed audio source separation
abstract
In this paper we tackle the problem of single channel audio source separation driven by descriptors of the sounding object's motion. As opposed to previous approaches, motion is included as a soft-coupling constraint within the nonnegative matrix factorization framework. The proposed method is applied to a multimodal dataset of instruments in string quartet performance recordings where bow motion information is used for separation of string instruments. We show that the approach offers better source separation result than an audio-based baseline and the state-of-the-art multimodal-based approaches on these very challenging music mixtures.
Sanjeel Parekh, Slim Essid, Alexey Ozerov, Ngoc Q. K. Duong, Patrick Pérez, Gaël Richard
ICASSP6
2017 Supervised group nonnegative matrix factorisation with similarity constraints and applications to speaker identification
abstract
This paper presents supervised feature learning approaches for speaker identification that rely on nonnegative matrix factorisation. Recent studies have shown that group nonnegative matrix factorisation and task-driven supervised dictionary learning can help performing effective feature learning for audio classification problems. This paper proposes to integrate a recent method that relies on group nonnegative matrix factorisation into a task-driven supervised framework for speaker identification. The goal is to capture both the speaker variability and the session variability while exploiting the discriminative learning aspect of the task-driven approach. Results on a subset of the ESTER corpus prove that the proposed approach can be competitive with I-vectors.
Romain Serizel, Victor Bisot, Slim Essid, Gaël Richard
ICASSP4
2017 Parallelized Stochastic Gradient Markov Chain Monte Carlo algorithms for non-negative matrix factorization
abstract
Stochastic Gradient Markov Chain Monte Carlo (SG-MCMC) methods have become popular in modern data analysis problems due to their computational efficiency. Even though they have proved useful for many statistical models, the application of SG-MCMC to non-negative matrix factorization (NMF) models has not yet been extensively explored. In this study, we develop two parallel SG-MCMC algorithms for a broad range of NMF models. We exploit the conditional independence structure of the NMF models and utilize a stratified sub-sampling approach for enabling parallelization. We illustrate the proposed algorithms on an image restoration task and report encouraging results.
Umut Simsekli, Alain Durmus, Roland Badeau, Gaël Richard, Eric Moulines, A. Taylan Cemgil
ICASSP4
2017 Reassigned time-frequency representations of discrete time signals and application to the Constant-Q Transform
Sébastien Fenet, Roland Badeau, Gaël Richard
Signal Process.3
2017 Speech intelligibility improvement in car noise environment by voice transformation
Karan Nathwani, Gaël Richard, Bertrand David 0002, Pierre Prablanc, Vincent Roussarie
Speech Commun.2
2017 Feature Learning With Matrix Factorization Applied to Acoustic Scene Classification
abstract
In this paper, we study the usefulness of various matrix factorization methods for learning features to be used for the specific acoustic scene classification (ASC) problem. A common way of addressing ASC has been to engineer features capable of capturing the specificities of acoustic environments. Instead, we show that better representations of the scenes can be automatically learned from time–frequency representations using matrix factorization techniques. We mainly focus on extensions including sparse, kernel-based, convolutive and a novel supervised dictionary learning variant of principal component analysis and nonnegative matrix factorization. An experimental evaluation is performed on two of the largest ASC datasets available in order to compare and discuss the usefulness of these methods for the task. We show that the unsupervised learning methods provide better representations of acoustic scenes than the best conventional hand-crafted features on both datasets. Furthermore, the introduction of a novel nonnegative supervised matrix factorization model and deep neural networks trained on spectrograms, allow us to reach further improvements.
Victor Bisot, Romain Serizel, Slim Essid, Gaël Richard
IEEE ACM Trans. Audio Speech Lang. Process.4
2017 Robust Downbeat Tracking Using an Ensemble of Convolutional Networks
abstract
In this paper, we present a novel state-of-the-art system for automatic downbeat tracking from music signals. The audio signal is first segmented in frames which are synchronized at the tatum level of the music. We then extract different kind of features based on harmony, melody, rhythm, and bass content to feed convolutional neural networks that are adapted to take advantage of the characteristics of each feature. This ensemble of neural networks is combined to obtain one downbeat likelihood per tatum. The downbeat sequence is finally decoded with a flexible and efficient temporal model which takes advantage of the assumed metrical continuity of a song. We then perform an evaluation of our system on a large base of nine datasets, compare its performance to four other published algorithms and obtain a significant increase of 16.8% points compared to the second-best system, for altogether a moderate cost in test and training. The influence of each step of the method is studied to show its strengths and shortcomings.
Simon Durand, Juan Pablo Bello, Bertrand David 0002, Gaël Richard
IEEE ACM Trans. Audio Speech Lang. Process.4
2017 Introduction to the Special Section on Sound Scene and Event Analysis
abstract
The papers in this special section are devoted to the growing field of acoustic scene classification and acoustic event recognition. Machine listening systems still have difficulties to reach the ability of human listeners in the analysis of realistic acoustic scenes. If sustained research efforts have been made for decades in speech recognition, speaker identification and to a lesser extent in music information retrieval, the analysis of other types of sounds, such as environmental sounds, is the subject of growing interest from the community and is targeting an ever increasing set of audio categories. This problem appears to be particularly challenging due to the large variety of potential sound sources in the scene, which may in addition have highly different acoustic characteristics, especially in bioacoustics. Furthermore, in realistic environments, multiple sources are often present simultaneously, and in reverberant conditions.
Gaël Richard, Tuomas Virtanen, Juan Pablo Bello, Nobutaka Ono, Hervé Glotin
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 Acoustic scene classification with matrix factorization for unsupervised feature learning
abstract
In this paper we study the use of unsupervised feature learning for acoustic scene classification (ASC). The acoustic environment recordings are represented by time-frequency images from which we learn features in an unsupervised manner. After a set of preprocessing and pooling steps, the images are decomposed using matrix factorization methods. By decomposing the data on a learned dictionary, we use the projection coefficients as features for classification. An experimental evaluation is done on a large ASC dataset to study popular matrix factorization methods such as Principal Component Analysis (PCA) and Non-negative Matrix Factorization (NMF) as well as some of their extensions including sparse, kernel based and convolutive variants. The results show the compared variants lead to significant improvement compared to the state-of-the-art results in ASC.
Victor Bisot, Romain Serizel, Slim Essid, Gaël Richard
ICASSP4
2016 Feature adapted convolutional neural networks for downbeat tracking
abstract
We define a novel system for the automatic estimation of downbeat positions from audio music signals. New rhythm and melodic features are introduced and feature adapted convolutional neural networks are used to take advantage of their specificity. Indeed, invariance to melody transposition, chroma data augmentation and length-specific rhythmic patterns prove to be useful to learn downbeat likelihood. After the data is segmented in tatums, complementary features related to melody, rhythm and harmony are extracted and the likelihood of a tatum being at a downbeat position is computed with the aforementioned neural networks. The downbeat sequence is then extracted with a flexible temporal hidden Markov model. We then show the efficiency and robustness of our approach with a comparative evaluation conducted on 9 datasets.
Simon Durand, Juan Pablo Bello, Bertrand David 0002, Gaël Richard
ICASSP4
2016 Formant shifting for speech intelligibility improvement in car noise environment
abstract
In this paper, we propose a novel approach aiming at improving the intelligibility of speech in the context of in-car applications. Speech produced in noisy environments is subject to the Lombard effect which gathers a number of voice transformation effects compared to the speech produced in calm environments. To improve intelligibility of in car speech (radio, message alerts, ...), we propose to modify the original speech signal by incorporating one of the important Lombard effect, namely the shift of the lower formant center frequencies away from the competing noise regions. The proposed approach exploits traditional Linear Prediction analysis and overlap and add synthesis. We explore several modification strategies and the merit of each modification is evaluated using both objective and subjective tests. It is in particular shown that the improvement of speech intelligibility in car noise is significantly improved for a majority of listeners.
Karan Nathwani, Morgane Daniel, Gaël Richard, Bertrand David 0002, Vincent Roussarie
ICASSP3
2016 Group nonnegative matrix factorisation with speaker and session variability compensation for speaker identification
abstract
This paper presents a feature learning approach for speaker identification that is based on nonnegative matrix factorisation. Recent studies have shown that with such models, the dictionary atoms can represent well the speaker identity. The approaches proposed so far focused only on speaker variability and not on session variability. However, this later point is a crucial aspect in the success of the I-vector approach that is now the state-of-the-art in speaker identification. This paper proposes a method that relies on group nonnegative matrix factorisation and that is inspired by the I-vector training procedure. By doing so the proposed approach intends to capture both the speaker variability and the session variability. Results on a small corpus prove that the proposed approach can be competitive with I-vectors.
Romain Serizel, Slim Essid, Gaël Richard
ICASSP3
2016 Stochastic thermodynamic integration: Efficient Bayesian model selection via stochastic gradient MCMC
abstract
Model selection is a central topic in Bayesian machine learning, which requires the estimation of the marginal likelihood of the data under the models to be compared. During the last decade, conventional model selection methods have lost their charm as they have high computational requirements. In this study, we propose a computationally efficient model selection method by integrating ideas from Stochastic Gradient Markov Chain Monte Carlo (SG-MCMC) literature and statistical physics. As opposed to conventional methods, the proposed method has very low computational needs and can be implemented almost without modifying existing SG-MCMC code. We provide an upper-bound for the bias of the proposed method. Our experiments show that, our method is 40 times as fast as the baseline method on finding the optimal model order in a matrix factorization problem.
Umut Simsekli, Roland Badeau, Gaël Richard, A. Taylan Cemgil
ICASSP3
2016 Machine listening techniques as a complement to video image analysis in forensics
abstract
Video is now one of the major sources of information for forensics. However, video documents can be originating from various recording devices (CCTV, mobile devices, etc.) with inconsistent quality and can sometimes be recorded in challenging light or motion conditions. Therefore, the amount of information that can be extracted relying solely on video image can vary to a great extent. Most of the videos however generally include audio recording as well. Machine listening can then become a valuable complement to video image analysis in challenging scenarios. In this paper, the authors present a brief overview of some machine listening techniques and their application to the analysis of video documents for forensics. The applicability of these techniques to forensics problems is then discussed in the light of machine listening system performances.
Romain Serizel, Victor Bisot, Slim Essid, Gaël Richard
ICIP4
2016 Stochastic Quasi-Newton Langevin Monte Carlo
abstract
Recently, Stochastic Gradient Markov Chain Monte Carlo (SG-MCMC) methods have been proposed for scaling up Monte Carlo computations to large data problems. Whilst these approaches have proven useful in many applications, vanilla SG-MCMC might suffer from poor mixing rates when random variables exhibit strong couplings under the target densities or big scale differences. In this study, we propose a novel SG-MCMC method that takes the local geometry into account by using ideas from Quasi-Newton optimization methods. These second order methods directly approximate the inverse Hessian by using a limited history of samples and their gradients. Our method uses dense approximations of the inverse Hessian while keeping the time and memory complexities linear with the dimension of the problem. We provide a formal theoretical analysis where we show that the proposed method is asymptotically unbiased and consistent with the posterior expectations. We illustrate the effectiveness of the approach on both synthetic and real datasets. Our experiments on two challenging applications show that our method achieves fast convergence rates similar to Riemannian approaches while at the same time having low computational requirements similar to diagonal preconditioning approaches.
Umut Simsekli, Roland Badeau, A. Taylan Cemgil, Gaël Richard
ICML4
2016 Using Temporal Association Rules for the Synthesis of Embodied Conversational Agents with a Specific Stance
Thomas Janssoone, Chloé Clavel, Kevin Bailly, Gaël Richard
IVA4
2016 Stochastic Gradient Richardson-Romberg Markov Chain Monte Carlo
abstract
Stochastic Gradient Markov Chain Monte Carlo (SG-MCMC) algorithms have become increasingly popular for Bayesian inference in large-scale applications. Even though these methods have proved useful in several scenarios, their performance is often limited by their bias. In this study, we propose a novel sampling algorithm that aims to reduce the bias of SG-MCMC while keeping the variance at a reasonable level. Our approach is based on a numerical sequence acceleration method, namely the Richardson-Romberg extrapolation, which simply boils down to running almost the same SG-MCMC algorithm twice in parallel with different step sizes. We illustrate our framework on the popular Stochastic Gradient Langevin Dynamics (SGLD) algorithm and propose a novel SG-MCMC algorithm referred to as Stochastic Gradient Richardson-Romberg Langevin Dynamics (SGRRLD). We provide formal theoretical analysis and show that SGRRLD is asymptotically consistent, satisfies a central limit theorem, and its non-asymptotic bias and the mean squared-error can be bounded. Our results show that SGRRLD attains higher rates of convergence than SGLD in both finite-time and asymptotically, and it achieves the theoretical accuracy of the methods that are based on higher-order integrators. We support our findings using both synthetic and real data experiments.
Alain Durmus, Umut Simsekli, Eric Moulines, Roland Badeau, Gaël Richard
NIPS5
2016 Fusion Methods for Speech Enhancement and Audio Source Separation
abstract
A wide variety of audio source separation techniques exist and can already tackle many challenging industrial issues. However, in contrast with other application domains, fusion principles were rarely investigated in audio source separation despite their demonstrated potential in classification tasks. In this paper, we propose a general fusion framework which takes advantage of the diversity of existing separation techniques in order to improve separation quality. We obtain new source estimates by summing the individual estimates given by different separation techniques weighted by a set of fusion coefficients. We investigate three alternative fusion methods which are based on standard nonlinear optimization, Bayesian model averaging, or deep neural networks. Experiments conducted for both speech enhancement and singing voice extraction demonstrate that all the proposed methods outperform traditional model selection. The use of deep neural networks for the estimation of time-varying coefficients notably leads to large quality improvements, up to 3 dB in terms of signal-to-distortion ratio compared to model selection.
Xabier Jaureguiberry, Emmanuel Vincent 0001, Gaël Richard
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 Multichannel Audio Source Separation With Probabilistic Reverberation Priors
abstract
Incorporating prior knowledge about the sources and/or the mixture is a way to improve under-determined audio source separation performance. A great number of informed source separation techniques concentrate on taking priors on the sources into account, but fewer works have focused on constraining the mixing model. In this paper, we address the problem of underdetermined multichannel audio source separation in reverberant conditions. We target a semi-informed scenario where some room parameters are known. Two probabilistic priors on the frequency response of the mixing filters are proposed. Early reverberation is characterized by an autoregressive model while according to statistical room acoustics results, late reverberation is represented by an autoregressive moving average model. Both reverberation models are defined in the frequency domain. They aim to transcribe the temporal characteristics of the mixing filters into frequency-domain correlations. Our approach leads to a maximum a posteriori estimation of the mixing filters which is achieved thanks to the expectation-maximization algorithm. We experimentally show the superiority of this approach compared with a maximum likelihood estimation of the mixing filters.
Simon Leglaive, Roland Badeau, Gaël Richard
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Downbeat tracking with multiple features and deep neural networks
abstract
In this paper, we introduce a novel method for the automatic estimation of downbeat positions from music signals. Our system relies on the computation of musically inspired features capturing important aspects of music such as timbre, harmony, rhythmic patterns, or local similarities in both timbre and harmony. It then uses several independent deep neural networks to learn higher-level representations. The downbeat sequences are finally obtained thanks to a temporal decoding step based on the Viterbi algorithm. The comparative evaluation conducted on varied datasets demonstrates the efficiency and robustness across different music styles of our approach.
Simon Durand, Juan Pablo Bello, Bertrand David 0002, Gaël Richard
ICASSP4
2015 Multipitch estimation using a PLCA-based model: Impact of partial user annotation
abstract
In this paper one investigates the merit of partial user annotation for music transcription using a PLCA-based model. The original algorithm, called Blind Harmonic Adaptive Decomposition (BHAD), provides an estimation of the polyphonic pitch content of the input signal in an entirely unsupervised manner. In this paper, one studies how the performance of the BHAD algorithm can be further improved by involving a user by means of a partial annotation. This user input allows for a better model initialisation with adapted or learned spectral envelope models. Furthermore, it is studied how a fine control of the convergence rate of some parameters can better exploit this additional information. It is then shown that this partial annotation can bring an improvement of up to 3% on the transcription of the remaining file.
Camila de Andrade Scatolini, Gaël Richard, Benoit Fuentes
ICASSP2
2015 Late Reverberation Synthesis: From Radiance Transfer to Feedback Delay Networks
abstract
In room acoustic modeling, feedback delay networks (FDN) are known to efficiently model late reverberation due to their capacity to generate exponentially decaying dense impulses. However, this method relies on a careful tuning of the different synthesis parameters, either estimated from a pre-recorded impulse response from the real acoustic scene, or set manually from experience. In this paper, we present a new method, which still inherits the efficiency of the FDN structure, but aims at linking the parameters of the FDN directly to the geometry setting. This relation is achieved by studying the sound energy exchange between each delay line using the acoustic radiance transfer method (RTM). Experimental results show that the late reverberation modeled by this method is in good agreement with the virtual geometry setting.
Hequn Bai, Gaël Richard, Laurent Daudet
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Enhancing downbeat detection when facing different music styles
abstract
This paper focuses on the automatic rhythm analysis of musical audio at the bar level. We propose a novel approach for robust downbeat detection. It uses well-chosen complementary features, inspired by musical considerations. In particular, a note accentuation model and a detection of pattern changes are introduced. We estimate the time signature by examining the similarity of frames at the beat level. The features are selected through a linear SVM model or a weighted sum. The whole system is evaluated on five different datasets of various musical styles and shows improvement over the state of the art.
Simon Durand, Bertrand David 0002, Gaël Richard
ICASSP3
2014 Single channel reverberation suppression based on sparse linear prediction
abstract
Reverberation degrades speech intelligibility in telecommunications as well as it increases the word error rate in automatic speech recognition tasks. Several dereverberation methods have been proposed recently in order to counter these effects. In the single microphone case, the dereverberation problem is underdetermined and reverberation suppression approaches are preferred. In this paper we propose a novel method for single channel reverberation suppression. Late reverberation is estimated in the time-frequency domain as a sparse linear combination of previous frames. The predictors associated to the model are determined in a Lasso framework and a spectral subtraction filter is designed to produce the enhanced signal. This model does not require any additional information about the room acoustics and it is well suited for real-time applications. The method has state-of-the-art performance in terms of both reverberation suppression and spectral distortion.
Nicolás López, Yves Grenier, Gaël Richard, Ivan Bourmeyster
ICASSP3
2014 Gesture recognition using a NMF-based representation of motion-traces extracted from depth silhouettes
abstract
We present a novel approach that classifies full-body human gestures using original spatio-temporal features obtained by applying non-negative matrix factorisation (NMF) to an extended depth silhouette representation. This extended representation, the motion-trace representation, incorporates temporal dimensions as it is built by superimposition of consecutive depth silhouettes. From this representation, a dictionary of local motion features is learned using NMF. Thus the projection of these local motion feature components on the incoming motion-traces results in a compact spatio-temporal feature representation. Those new features are then exploited using hidden Markov models for gesture recognition. Our experiments on a gesture dataset show that our approach outperforms more traditional methods that use pose features or decomposition techniques such as principal component analysis.
Aymeric Masurelle, Slim Essid, Gaël Richard
ICASSP3
2014 Multiple-order non-negative matrix factorization for speech enhancement
abstract
Amongst the speech enhancement techniques, statistical models based on Non-negative Matrix Factorization (NMF) have received great attention. In a single channel configuration, NMF is used to describe the spectral content of both the speech and noise sources. As the number of components can have a crucial influence on separation quality, we here propose to investigate model order selection based on the variational Bayesian approximation to the marginal likelihood of models of different orders. To go further, we propose to use model averaging to combine several single-order NMFs and we show that a straightforward application of model averaging principles is inefficient as it turned out to be equivalent to model selection. We thus introduce a parameter to control the entropy of the model order distribution which makes the averaging effective. We also show that our probabilistic model nicely extends to a multiple-order NMF model where several NMFs are jointly estimated and averaged. Experiments are conducted on real data from the CHiME challenge and give an interesting insight on the entropic parameter and model order priors. Separation results are also promising as model averaging outperforms single-order model selection. Finally, our multiple-order NMF shows an interesting gain in computation time.
Xabier Jaureguiberry, Emmanuel Vincent 0001, Gaël Richard
INTERSPEECH3
2014 Blind Denoising with Random Greedy Pursuits
abstract
Denoising methods require some assumptions about the signal of interest and the noise. While most denoising procedures require some knowledge about the noise level, which may be unknown in practice, here we assume that the signal expansion in a given dictionary has a distribution that is more heavy-tailed than the noise. We show how this hypothesis leads to a stopping criterion for greedy pursuit algorithms which is independent from the noise level. Inspired by the success of ensemble methods in machine learning, we propose a strategy to reduce the variance of greedy estimates by averaging pursuits obtained from randomly subsampled dictionaries. We call this denoising procedure Blind Random Pursuit Denoising (BIRD). We offer a generalization to multidimensional signals, with a structured sparse model (S-BIRD). The relevance of this approach is demonstrated on synthetic and experimental MEG signals where, without any parameter tuning, BIRD outperforms state-of-the-art algorithms even when they are informed by the noise level. Code is available to reproduce all experiments.
Manuel Moussallam, Alexandre Gramfort, Laurent Daudet, Gaël Richard
IEEE Signal Process. Lett.4
2013 Low bitrate informed source separation of realistic mixtures
abstract
Demixing consists in recovering the sounds that compose a multichannel mix. Important applications include karaoke or respatialization. Several approaches to this problem have been proposed in a coding/decoding framework, which are denoted either as spatial audio object coding or informed source separation. They assume that the constituent sounds are available at an encoding stage and used to compute a side-information transmitted to the end-user. At a decoding stage, only the mixtures and the side information are used to recover the sources. Here, we propose an advanced model, which encompasses many practical scenarios and permits to reach bitrates as low as 0:5kbps/source. First, the sources may be mono or multichannel. Second, the mixing process is assumed to be diffuse, generalizing the usual linear-instantaneous or convolutive cases and permitting professional mixes to be processed. Third, the signals to be recovered may either be the original sources or their spatial images.
Antoine Liutkus, Roland Badeau, Gaël Richard
ICASSP3
2013 An Overview on Perceptually Motivated Audio Indexing and Classification
abstract
An audio indexing system aims at describing audio content by identifying, labeling, or categorizing different acoustic events. Since the resulting audio classification and indexing is meant for direct human consumption, it is highly desirable that it produces perceptually relevant results. This can be obtained by integrating specific knowledge of the human auditory system in the design process to various extent. In this paper, we highlight some of the important concepts used in audio classification and indexing that are perceptually motivated or that exploit some principles of perception. In particular, we discuss several different strategies to integrate human perception, including: 1) the use of generic audition models; 2) the use of perceptually relevant features for the analysis stage that are perceptually justified either as a component of a hearing model or as being correlated with a perceptual dimension of sound similarity; and 3) the involvement of the user in the audio indexing or classification task. In this paper, we also illustrate some of the recent trends in semantic audio retrieval that approximate higher level perceptual processing and cognitive aspects of human audio recognition capabilities, including affect-based audio retrieval.
Gaël Richard, Shiva Sundaram, Shri Narayanan
Proc. IEEE1
2013 Parametric Audio Coding With Exponentially Damped Sinusoids
abstract
Sinusoidal modeling is one of the most popular techniques for low bitrate audio coding. Usually, the sinusoidal parameters (amplitude, pulsation and phase of each sinusoidal component) are kept constant within a time segment. An alternative model, the so-called Exponentially-Damped Sinusoidal (EDS) model, includes an additional damping parameter for each sinusoidal component to better represent the signal characteristics. It was however never shown that the EDS model could be efficient for perceptual audio coding. To that aim, we propose in this paper an efficient analysis/synthesis framework with dynamic time-segmentation on transients and psychoacoustic modeling, and an asymptotically optimal entropy-constrained quantization method for the four sinusoid parameters (e.g., including damping). We then apply this coding technique to real audio excerpts for a given entropy target corresponding to a low bitrate (20 kbits/s), and compare this method with a classical sinusoidal coding scheme using a constant-amplitude sinusoidal model and the perceptually weighted Matching Pursuit algorithm. Subjective listening tests show that the EDS model is more efficient on audio samples with fast transient content, and similar to the classical model for more stationary audio samples.
Olivier Derrien, Roland Badeau, Gaël Richard
IEEE Trans. Speech Audio Process.3
2013 Harmonic Adaptive Latent Component Analysis of Audio and Application to Music Transcription
abstract
Recently, new methods for smart decomposition of time-frequency representations of audio have been proposed in order to address the problem of automatic music transcription. However those techniques are not necessarily suitable for notes having variations of both pitch and spectral envelope over time. The HALCA (Harmonic Adaptive Latent Component Analysis) model presented in this article allows considering those two kinds of variations simultaneously. Each note in a constant-Q transform is locally modeled as a weighted sum of fixed narrowband harmonic spectra, spectrally convolved with some impulse that defines the pitch. All parameters are estimated by means of the expectation-maximization (EM) algorithm, in the framework of Probabilistic Latent Component Analysis. Interesting priors over the parameters are also introduced in order to help the EM algorithm converging towards a meaningful solution. We applied this model for automatic music transcription: the onset time, duration and pitch of each note in an audio file are inferred from the estimated parameters. The system has been evaluated on two different databases and obtains very promising results.
Benoit Fuentes, Roland Badeau, Gaël Richard
IEEE Trans. Speech Audio Process.3
2013 Learning Optimal Features for Polyphonic Audio-to-Score Alignment
abstract
This paper addresses the design of feature functions for the matching of a musical recording to the symbolic representation of the piece (the score). These feature functions are defined as dissimilarity measures between the audio observations and template vectors corresponding to the score. By expressing the template construction as a linear mapping from the symbolic to the audio representation, one can learn the feature functions by optimizing the linear transformation. In this paper, we explore two different learning strategies. The first one uses a best-fit criterion (minimum divergence), while the second one exploits a discriminative framework based on a Conditional Random Fields model (maximum likelihood criterion). We evaluate the influence of the feature functions in an audio-to-score alignment task, on a large database of popular and classical polyphonic music. The results show that with several types of models, using different temporal constraints, the learned mappings have the potential to outperform the classic heuristic mappings. Several representations of the audio observations, along with several distance functions are compared in this alignment task. Our experiments elect the symmetric Kullback-Leibler divergence. Moreover, both the spectrogram and a CQT-based representation turn out to provide very accurate alignments, detecting more than 97% of the onsets with a precision of 100 ms with our most complex system.
Cyril Joder, Slim Essid, Gaël Richard
IEEE Trans. Speech Audio Process.3
2013 Coding-Based Informed Source Separation: Nonnegative Tensor Factorization Approach
abstract
Informed source separation (ISS) aims at reliably recovering sources from a mixture. To this purpose, it relies on the assumption that the original sources are available during an encoding stage. Given both sources and mixture, a side-information may be computed and transmitted along with the mixture, whereas the original sources are not available any longer. During a decoding stage, both mixture and side-information are processed to recover the sources. ISS is motivated by a number of specific applications including active listening and remixing of music, karaoke, audio gaming, etc. Most ISS techniques proposed so far rely on a source separation strategy and cannot achieve better results than oracle estimators. In this study, we introduce Coding-based ISS (CISS) and draw the connection between ISS and source coding. CISS amounts to encode the sources using not only a model as in source coding but also the observation of the mixture. This strategy has several advantages over conventional ISS methods. First, it can reach any quality, provided sufficient bandwidth is available as in source coding. Second, it makes use of the mixture in order to reduce the bitrate required to transmit the sources, as in classical ISS. Furthermore, we introduce Nonnegative Tensor Factorization as a very efficient model for CISS and report rate-distortion results that strongly outperform the state of the art.
Alexey Ozerov, Antoine Liutkus, Roland Badeau, Gaël Richard
IEEE Trans. Speech Audio Process.4
2012 A regressive boosting approach to automatic audio tagging based on soft annotator fusion
abstract
Automatic tagging of music has mostly been treated as a classification problem. In this framework, the association of a tag to a song is characterized in a “hard” fashion: the tag is either relevant or not. Yet, the relevance of a tag to a song is not always evident. Indeed, during the ground-truth annotation process, several annotators may express doubts, or disagree with each other. In this paper, we propose to fuse annotators' decisions in a way to keep information about this uncertainty. This fusion provides us continuous scores, that are used for training a regressive boosting algorithm. Our experiments show that regression with this soft ground truth leads to a more accurate learning, and better predictions, compared to traditionally used binary classification.
Rémi Foucard, Slim Essid, Mathieu Lagrange, Gaël Richard
ICASSP4
2012 Probabilistic model for main melody extraction using Constant-Q transform
abstract
Dimension reduction techniques such as Nonnegative Tensor Factorization are now classical for both source separation and estimation of multiple fundamental frequencies in audio mixtures. Still, few studies jointly addressed these tasks so far, mainly because separation is often based on the Short Term Fourier Transform (STFT) whereas recent music analysis algorithms are rather based on the Constant-Q Transform (CQT). The CQT is practical for pitch estimation because a pitch shift amounts to a translation of the CQT representation, whereas it produces a scaling of the STFT. Conversely, no simple inversion of the CQT was available until recently, preventing it from being used for source separation. Benefiting from advances both in the inversion of the CQT and in statistical modeling, we show how recent techniques designed for music analysis can also be used for source separation with encouraging results, thus opening the path to many crossovers between separation and analysis.
Benoit Fuentes, Antoine Liutkus, Roland Badeau, Gaël Richard
ICASSP4
2012 A probabilistic approach to simultaneous extraction of beats and downbeats
abstract
This paper focuses on the automatic extraction of beat structure from a musical piece. A novel statistical approach to modeling beat sequences based on the application of Hidden Markov Models (HMM) is introduced. The resulting beat labels are obtained by running the Viterbi decoder and subsequent lattice rescoring. For the observation vectors we propose a new feature set that is based on the impulsive and harmonic components of the reassigned spectrogram. Different components of observation vectors have been investigated for their efficiency. The main advantage of the proposed approach is the absence of imposed deterministic rules. All the parameters are learned from the training data, and the experimental results show the efficiency of the proposed schema.
Maksim Khadkevich, Thomas Fillon, Gaël Richard, Maurizio Omologo
ICASSP3
2012 Adaptive filtering for music/voice separation exploiting the repeating musical structure
abstract
The separation of the lead vocals from the background accompaniment in audio recordings is a challenging task. Recently, an efficient method called REPET (REpeating Pattern Extraction Technique) has been proposed to extract the repeating background from the non-repeating foreground. While effective on individual sections of a song, REPET does not allow for variations in the background (e.g. verse vs. chorus), and is thus limited to short excerpts only. We overcome this limitation and generalize REPET to permit the processing of complete musical tracks. The proposed algorithm tracks the period of the repeating structure and computes local estimates of the background pattern. Separation is performed by soft time-frequency masking, based on the deviation between the current observation and the estimated background pattern. Evaluation on a dataset of 14 complete tracks shows that this method can perform at least as well as a recent competitive music/voice separation method, while being computationally efficient.
Antoine Liutkus, Zafar Rafii, Roland Badeau, Bryan Pardo, Gaël Richard
ICASSP5
2012 Random time-frequency subdictionary design for sparse representations with greedy algorithms
abstract
Sparse signal approximation can be used to design efficient low bit-rate coding schemes. It heavily relies on the ability to design appropriate dictionaries and corresponding decomposition algorithms. The size of the dictionary, and therefore its resolution, is a key parameter that handles the tradeoff between sparsity and tractability. This work proposes the use of a non adaptive random sequence of subdictionaries in a greedy decomposition process, thus browsing a larger dictionary space in a probabilistic fashion with no additional projection cost nor parameter estimation. This technique leads to very sparse decompositions, at a controlled computational complexity. Experimental evaluation is provided as proof of concept for low bit rate compression of audio signals.
Manuel Moussallam, Laurent Daudet, Gaël Richard
ICASSP3
2012 Informed source separation through spectrogram coding and data embedding
Antoine Liutkus, Jonathan Pinel, Roland Badeau, Laurent Girin, Gaël Richard
Signal Process.5
2012 Matching Pursuits with random sequential subdictionaries
Manuel Moussallam, Laurent Daudet, Gaël Richard
Signal Process.3
2012 Multiclass Feature Selection With Kernel Gram-Matrix-Based Criteria
abstract
Feature selection has been an important issue in recent decades to determine the most relevant features according to a given classification problem. Numerous methods have emerged that take into account support vector machines (SVMs) in the selection process. Such approaches are powerful but often complex and costly. In this paper, we propose new feature selection methods based on two criteria designed for the optimization of SVM: kernel target alignment and kernel class separability. We demonstrate how these two measures, when fully expressed, can build efficient and simple methods, easily applicable to multiclass problems and iteratively computable with minimal memory requirements. An extensive experimental study is conducted both on artificial and real-world datasets to compare the proposed methods to state-of-the-art feature selection algorithms. The results demonstrate the relevance of the proposed methods both in terms of performance and computational cost.
Mathieu Ramona, Gaël Richard, Bertrand David 0002
IEEE Trans. Neural Networks Learn. Syst.2
2011 Entropy-constrained quantization of exponentially damped sinusoids parameters
abstract
Sinusoidal modeling is traditionally one of the most popular techniques for low bitrate audio coding. Usually, the sinusoidal parameters are kept constant within a time segment but the exponentially damped sinusoidal (EDS) model is also an efficient alternative. However, the inclusion of an additional damping parameter calls for a specific quantization scheme. In this paper, we propose an asymptotically optimal entropy-constrained quantization method for amplitude, phase and damping parameters. We show that this scheme is nearly optimal in terms of rate-distortion trade-off. We also show that damping consumes the smallest part of the total entropy of quantization indexes, which suggests that the EDS model is truly efficient for audio coding.
Olivier Derrien, Roland Badeau, Gaël Richard
ICASSP3
2011 Adaptive harmonic time-frequency decomposition of audio using shift-invariant PLCA
abstract
Numerous methods have been developed for the time-frequency analysis and smart decomposition of audio signals. However, these techniques are not consistently suitable for real music signals where each note presents continuous variations of both pitch and spectral envelope. This paper presents a new model for analyzing the harmonic structures of an audio signal that can jointly handle those two types of variations. Each note in a constant-Q transform is modeled as a weighted sum of narrowband parametric spectra, and positive deconvolution is performed to estimate the model parameters, in the framework of probabilistic latent component analysis. The algorithm has been tested in a task of monopitch estimation. The very promising results highlight the reliability and the robustness of the model.
Benoit Fuentes, Roland Badeau, Gaël Richard
ICASSP3
2011 Hidden Discrete Tempo Model: A tempo-aware timing model for audio-to-score alignment
abstract
In this paper, we present the Hidden Discrete Tempo Model, an effective Dynamic Bayesian Network for audio to score matching. Its main feature is an explicit modeling of tempo, which directly in fluences the timing model of the musical performance. Thanks to a discretization of the tempo set, it allows for an efficient decoding by the Viterbi algorithm, and facilitates the introduction of features which directly depend on the local tempo. We take advantage of this property by using the cyclic tempogram descriptor in addition to chroma vectors and onset detection features. Experiment run on both classical piano and pop music show the very high accuracy of this model for audio to score alignment, as well as the usefulness of die tempo feature used.
Cyril Joder, Slim Essid, Gaël Richard
ICASSP3
2011 Audio Signal Representations for Factorization in the Sparse Domain
abstract
In this paper, a new class of audio representations is introduced, together with a corresponding fast decomposition algorithm. The main feature of these representations is that they are both sparse and approximately shift-invariant, which allows similarity search in a sparse domain. The common sparse support of detected similar patterns is then used to factorize their representations. The potential of this method for simultaneous structural analysis and compressing tasks is illustrated by preliminary experiments on simple musical data.
Manuel Moussallam, Laurent Daudet, Gaël Richard
ICASSP3
2011 Combining monaural source separation with Long Short-Term Memory for increased robustness in vocalist gender recognition
abstract
We present a novel and unique combination of algorithms to detect the gender of the leading vocalist in recorded popular music. Building on our previous successful approach that enhanced the harmonic parts by means of Non-Negative Matrix Factorization (NMF) for increased accuracy, we integrate on the one hand a new source separation algorithm specifically tailored to extracting the leading voice from monaural recordings. On the other hand, we introduce Bidirectional Long Short-Term Memory Recurrent Neural Networks (BLSTM-RNNs) as context-sensitive classifiers for this scenario, which have lately led to great success in Music Information Retrieval tasks. Through a combination of leading voice separation and BLSTM networks, as opposed to a baseline approach using Hidden Naive Bayes on the original recordings, the accuracy of simultaneous detection of vocal presence and vocalist gender on beat level is improved by up to 10% absolute. Furthermore, using this technique we achieve 91.6% accuracy in determining the gender of the predominant vocalist on song level, which is 4% absolute above our previous best result.
Felix Weninger, Jean-Louis Durrieu, Florian Eyben, Gaël Richard, Björn W. Schuller
ICASSP4
2011 An audio-driven virtual dance-teaching assistant
abstract
This work addresses the Huawei/3Dlife Grand challenge proposing a set of audio tools for a virtual dance-teaching assistant. These tools are meant to help the dance student develop a sense of rhythm to correctly synchronize his/her movements and steps to the musical timing of the choreographies to be executed. They consist of three main components, namely a music (beat) analysis module, a source separation and remastering module and a dance step segmentation module. These components enable to create augmented tutorial videos highlighting the rhythmic information using, for instance, a synthetic dance teacher voice, but also videos highlighting the steps executed by a student to help in the evaluation of his/her performance.
Slim Essid, Yves Grenier, Mounira Maazaoui, Gaël Richard, Robin Tournemenne
ACM Multimedia4
2011 Tutorial on multimedia music signal processing
abstract
No abstract available.
Gaël Richard
ACM Multimedia1
2011 A Conditional Random Field Framework for Robust and Scalable Audio-to-Score Matching
abstract
In this paper, we introduce the use of conditional random fields (CRFs) for the audio-to-score alignment task. This framework encompasses the statistical models which are used in the literature and allows for more flexible dependency structures. In particular, it allows observation functions to be computed from several analysis frames. Three different CRF models are proposed for our task, for different choices of tradeoff between accuracy and complexity. Three types of features are used, characterizing the local harmony, note attacks and tempo. We also propose a novel hierarchical approach, which takes advantage of the score structure for an approximate decoding of the statistical model. This strategy reduces the complexity, yielding a better overall efficiency than the classic beam search method used in HMM-based models. Experiments run on a large database of classical piano and popular music exhibit very accurate alignments. Indeed, with the best performing system, more than 95% of the note onsets are detected with a precision finer than 100 ms. We additionally show how the proposed framework can be modified in order to be robust to possible structural differences between the score and the musical performance.
Cyril Joder, Slim Essid, Gaël Richard
IEEE ACM Trans. Audio Speech Lang. Process.3
2010 Robust frequency-based Audio Fingerprinting
abstract
Pure frequency-based audio fingerprint systems have the capacity of handling very short fingerprints while being highly robust to perturbations such as additive noise or compression. However, these approaches are often complex and fail to identify time stretched signals. We propose in this paper two extensions of an existing system and test the robustness of the overall system in different conditions. It is shown that the search strategy adopted allows for a clear reduction of complexity with very limited degradation of performances and that the new system is robust to additive noise and speed changes up to 5%.
Elsa Dupraz, Gaël Richard
ICASSP2
2010 Multimodal similarity between musical streams for cover version detection
abstract
Expressing the similarity between musical streams is a challenging task as it involves the understanding of many factors which are most often blended into one information channel: the audio stream. Consequently, separating the musical audio stream into its main melody and its accompaniment may prove as being useful to root the similarity computation on a more robust and expressive representation. In this paper, we show that considering the mixture, an estimation of its main melody and its accompaniment as modalities allows us to propose new ways of defining the similarity between musical streams. In the context of the detection of cover version, we show that highest performance is achieved by jointly considering the mixture and the estimated accompaniment. As demonstrated by the experiments carried out using two different evaluation databases, this scheme allows the scoring system to focus more on the chord progression by considering the accompaniment while being robust to the potential separation errors by also considering the mixture.
Rémi Foucard, Jean-Louis Durrieu, Mathieu Lagrange, Gaël Richard
ICASSP4
2010 A comparative study of tonal acoustic features for a symbolic level music-to-score alignment
abstract
In this paper we review the acoustic features used for music-to-score alignment and study their influence on the performance in a challenging alignment task, where the audio data is polyphonic and may contain percussion. Furthermore, as we aim at using “real world” scores, we follow an approach which does exploit the rhythm information (considered unreliable) and test its robustness to score errors. We use a unified framework to handle different state-of-the-art features, and propose a simple way to exploit either a model of the feature values, or an audio synthesis of a musical score, in an audio-to-score alignment system. We confirm that chroma vectors drawn from representations using a logarithmic frequency scale are the most efficient features, and lead to a good precision, even with a simple alignment strategy. Robustness tests also show that the relative performance of the features do not depend on possible musical score degradations.
Cyril Joder, Slim Essid, Gaël Richard
ICASSP3
2010 Robust similarity metrics between audio signals based on asymmetrical spectral envelope matching
abstract
In this paper, a new type of metric that defines the similarity between musical audio signals is proposed. Based on the spectral flatness criterion, those metrics achieve low computational cost and low sensitivity to acoustical degradations. Validation is performed by studying the ability of the proposed metric to determine whether two audio signals have been played by the same musical instrument. For this task, proposed metrics are shown to overcome metrics based on the comparison of standard spectral features especially when the request and the records of the database are of different acoustical properties.
Mathieu Lagrange, Roland Badeau, Gaël Richard
ICASSP3
2010 Robust visual features for the multimodal identification of unregistered speakers in TV talk-shows
abstract
In this paper we propose a novel multimodal method for identifying unregistered speakers in a TV talk-show using a semi-supervised learning approach based on Support Vector Machines. Our study highlights the fact that specific visual features prove to be very efficient for this particular type of video content which is edited from multi-camera recordings. These visual features, motivated by prior knowledge on the approach followed by the TV director in choosing the appropriate shots, are found to bring a significant improvement in identification accuracy when used together with classic audio Mel-frequency cepstral coefficients (+8% compared to various baseline systems, in particular a standard audio only system).
Félicien Vallet, Slim Essid, Jean Carrive, Gaël Richard
ICIP4
2010 A conditional random field viewpoint of symbolic audio-to-score matching
abstract
We present a new approach of symbolic audio-to-score alignment, with the use of Conditional Random Fields (CRFs). Unlike Hidden Markov Models, these graphical models allow the calculation of state conditional probabilities to be made on the basis of several audio frames. The CRF models that we propose exploit this property to take into account the rhythmic information of the musical score. Assuming that the tempo is locally constant, they confront the neighborhood of each frame with several tempo hypotheses.
Cyril Joder, Slim Essid, Gaël Richard
ACM Multimedia3
2010 Explicit modeling of temporal dynamics within musical signals for acoustical unit similarity
Mathieu Lagrange, Martin Raspaud, Roland Badeau, Gaël Richard
Pattern Recognit. Lett.4
2010 Source/Filter Model for Unsupervised Main Melody Extraction From Polyphonic Audio Signals
abstract
Extracting the main melody from a polyphonic music recording seems natural even to untrained human listeners. To a certain extent it is related to the concept of source separation, with the human ability of focusing on a specific source in order to extract relevant information. In this paper, we propose a new approach for the estimation and extraction of the main melody (and in particular the leading vocal part) from polyphonic audio signals. To that aim, we propose a new signal model where the leading vocal part is explicitly represented by a specific source/filter model. The proposed representation is investigated in the framework of two statistical models: a Gaussian Scaled Mixture Model (GSMM) and an extended Instantaneous Mixture Model (IMM). For both models, the estimation of the different parameters is done within a maximum-likelihood framework adapted from single-channel source separation techniques. The desired sequence of fundamental frequencies is then inferred from the estimated parameters. The results obtained in a recent evaluation campaign (MIREX08) show that the proposed approaches are very promising and reach state-of-the-art performances on all test sets.
Jean-Louis Durrieu, Gaël Richard, Bertrand David 0002, Cédric Févotte
IEEE Trans. Speech Audio Process.2
2010 Audio Signal Representations for Indexing in the Transform Domain
abstract
Indexing audio signals directly in the transform domain can potentially save a significant amount of computation when working on a large database of signals stored in a lossy compression format, without having to fully decode the signals. Here, we show that the representations used in standard transform-based audio codecs (e.g., MDCT for AAC, or hybrid PQF/MDCT for MP3) have a sufficient time resolution for some rhythmic features, but a poor frequency resolution, which prevents their use in tonality-related applications. Alternatively, a recently developed audio codec based on a sparse multi-scale MDCT transform has a good resolution both for time- and frequency-domain features. We show that this new audio codec allows efficient transform-domain audio indexing for three different applications, namely beat tracking, chord recognition, and musical genre classification. We compare results obtained with this new audio codec and the two standard MP3 and AAC codecs, in terms of performance and computation time.
Emmanuel Ravelli, Gaël Richard, Laurent Daudet
IEEE Trans. Speech Audio Process.2
2009 An iterative approach to monaural musical mixture de-soloing
abstract
In this article, we introduce a novel approach for monaural source separation with the specific aim to separate a polyphonic musical recording into two main sources: a main instrument (or melody) track and an accompaniment track. To that aim, we propose to model the power spectral densities (PSDs) of both contributions with a source/filter model for the main instrument while retaining a model emphasizing temporal repetitions of the musical background. We show that improved source separation performances can be obtained by a two-step estimation strategy where the model parameters are re-estimated in a second stage by adequately exploiting the main melody line estimated in a first stage. The experiments conducted on several monaural signal databases show that our system achieves state-of-the-art performances compared to other unsupervised source separation algorithms.
Jean-Louis Durrieu, Gaël Richard, Bertrand David 0002
ICASSP2
2009 Incorporating prior knowledge on the digital media creation process into audio classifiers
abstract
In the process of music content creation, a wide range of typical audio effects such as reverberation, equalization or dynamic compression are very commonly used. Despite the fact that such effects have a clear impact on the audio features, they are rarely taken into account when building an automatic audio classifier. In this paper, it is shown that the incorporation of prior knowledge of the digital media creation chain can clearly improve the robustness of the audio classifiers, which is demonstrated on a task of musical instrument recognition. The proposed system is based on a robust feature selection strategy, on a novel use of the virtual support vector machines technique and a specific equalization used to normalize the signals to be classified. The robustness of the proposed system is experimentally evidenced using a rather large and varied sound database.
Maxime Lardeur, Slim Essid, Gaël Richard, Martin Haller, Thomas Sikora
ICASSP3
2009 Temporal Integration for Audio Classification With Application to Musical Instrument Classification
abstract
Nowadays, it appears essential to design automatic indexing tools which provide meaningful and efficient means to describe the musical audio content. There is in fact a growing interest for music information retrieval (MIR) applications amongst which the most popular are related to music similarity retrieval, artist identification, musical genre or instrument recognition. Current MIR-related classification systems usually do not take into account the mid-term temporal properties of the signal (over several frames) and lie on the assumption that the observations of the features in different frames are statistically independent. The aim of this paper is to demonstrate the usefulness of the information carried by the evolution of these characteristics over time. To that purpose, we propose a number of methods for early and late temporal integration and provide an in-depth experimental study on their interest for the task of musical instrument recognition on solo musical phrases. In particular, the impact of the time horizon over which the temporal integration is performed will be assessed both for fixed and variable frame length analysis. Also, a number of proposed alignment kernels will be used for late temporal integration. For all experiments, the results are compared to a state of the art musical instrument recognition system.
Cyril Joder, Slim Essid, Gaël Richard
IEEE Trans. Speech Audio Process.3
2008 Singer melody extraction in polyphonic signals using source separation methods
abstract
We propose a new approach for singer melody extraction, based on blind source separation techniques. The short time Fourier transform (STFT) of the singer signal is modelled by a Gaussian mixture model (GMM) explicitly coupled with a generative source/filter model. We then introduce a simplification of this general GMM and approximate the STFT of the music signal using Non-negative Matrix Factorization (NMF) techniques. The melody line is extracted from the explicit source component of the model thanks to a Viterbi algorithm. The results are very promising and comparable or better than those of state-of-the-art systems.
Jean-Louis Durrieu, Gaël Richard, Bertrand David 0002
ICASSP2
2008 Vocal detection in music with support vector machines
abstract
We propose a statistical learning approach for the automatic detection of vocal regions in a polyphonic musical signal. A support vector model, based on a large feature set, is employed to discriminate accompanied singing voice from pure instrumental regions. We propose a temporal smoothing of the posterior probabilities with a hidden Markov model that helps adapting the segmentation sequence to the precision of the manual annotation. Quantitative results on a copyright- free public musical corpus show a classification accuracy of 82%.
Mathieu Ramona, Gaël Richard, Bertrand David 0002
ICASSP2
2008 Fear-type emotion recognition for future audio-based surveillance systems
Chloé Clavel, Ioana Vasilescu, Laurence Devillers, Gaël Richard, Thibaut Ehrette
Speech Commun.4
2008 A New Model-Based Algorithm for Optimizing the MPEG-AAC in MS-Stereo
abstract
In this paper, a new model-based algorithm for optimizing the MPEG-advanced audio coder (AAC) in MS-stereo mode is presented. This algorithm is an extension to stereo signals of prior work on a statistical model of quantization noise. Traditionally, MS-stereo coding approaches replace the left (l) and right (R) channels by the middle (M) and sides (S) channels, each channel being independently processed, almost like a monophonic signal. In contrast, our method proposes a global approach for coding both channels in the same process. A model for the quantization error allows us to tune the quantizers on channels M and S with respect to a distortion constraint on the reconstructed channels L and R as they will appear in the decoder. This approach leads to a more efficient perceptual noise-shaping and avoids using complex psychoacoustic models built on the M and S channels. Furthermore, it provides a straightforward scheme to choose between LR and MS modes in each subband for each frame. Subjective listening tests prove that the coding efficiency at a medium bitrate (96 kbits/s for both channels) is significantly better with our algorithm than with the standard algorithm, without increase of complexity.
Olivier Derrien, Gaël Richard
IEEE Trans. Speech Audio Process.2
2008 Transcription and Separation of Drum Signals From Polyphonic Music
abstract
The purpose of this article is to present new advances in music transcription and source separation with a focus on drum signals. A complete drum transcription system is described, which combines information from the original music signal and a drum track enhanced version obtained by source separation. In addition to efficient fusion strategies to take into account these two complementary sources of information, the transcription system integrates a large set of features, optimally selected by feature selection. Concurrently, the problem of drum track extraction from polyphonic music is tackled both by proposing a novel approach based on harmonic/noise decomposition and time/frequency masking and by improving an existing Wiener filtering-based separation method. The separation and transcription techniques presented are thoroughly evaluated on a large public database of music signals. A transcription accuracy between 64.5% and 80.3% is obtained, depending on the drum instrument, for well-balanced mixes, and the efficiency of our drum separation algorithms is illustrated in a comprehensive benchmark.
Olivier Gillet, Gaël Richard
IEEE Trans. Speech Audio Process.2
2008 Instrument-Specific Harmonic Atoms for Mid-Level Music Representation
abstract
Several studies have pointed out the need for accurate mid-level representations of music signals for information retrieval and signal processing purposes. In this paper, we propose a new mid-level representation based on the decomposition of a signal into a small number of sound atoms or molecules bearing explicit musical instrument labels. Each atom is a sum of windowed harmonic sinusoidal partials whose relative amplitudes are specific to one instrument, and each molecule consists of several atoms from the same instrument spanning successive time windows. We design efficient algorithms to extract the most prominent atoms or molecules and investigate several applications of this representation, including polyphonic instrument recognition and music visualization.
Pierre Leveau, Emmanuel Vincent 0001, Gaël Richard, Laurent Daudet
IEEE Trans. Speech Audio Process.3
2008 Union of MDCT Bases for Audio Coding
abstract
This paper investigates the use of sparse overcomplete decompositions for audio coding. Audio signals are decomposed over a redundant union of modified discrete cosine transform (MDCT) bases having eight different scales. This approach produces a sparser decomposition than the traditional MDCT-based orthogonal transform and allows better coding efficiency at low bitrates. Contrary to state-of-the-art low bitrate coders, which are based on pure parametric or hybrid representations, our approach is able to provide transparency. Moreover, we use a bitplane encoding approach, which provides a fine-grain scalable coder that can seamlessly operate from very low bitrates up to transparency. Objective evaluation, as well as listening tests, show that the performance of our coder is significantly better than a state-of-the-art transform coder at very low bitrates and has similar performance at high bitrates. We provide a link to test soundfiles and source code to allow better evaluation and reproducibility of the results.
Emmanuel Ravelli, Gaël Richard, Laurent Daudet
IEEE Trans. Speech Audio Process.2
2007 Conjugate Gradient Algorithms for Minor Subspace Analysis
abstract
We introduce a conjugate gradient method for estimating and tracking the minor eigenvector of a data correlation matrix. This new algorithm is less computationally demanding and converges faster than other methods derived from the conjugate gradient approach. It can also be applied in the context of minor subspace tracking, as a pre-processing step for the YAST algorithm, in order to enhance its performance. Simulations show that the resulting algorithm converges much faster than existing minor subspace trackers.
Roland Badeau, Bertrand David 0002, Gaël Richard
ICASSP (3)3
2007 Blind Signal Decompositions for Automatic Transcription of Polyphonic Music: NMF and K-SVD on the Benchmark
abstract
This paper investigates on the behavior of two blind signal decomposition algorithms, non negative matrix factorization (NMF) and non negative K-SVD (NKSVD), in a polyphonic music transcription task. State-of-the-art transcription systems are based on a frame-by-frame, low-level approach; blind systems could be an alternative to them. Two raw but effective audio-to-MIDI systems are proposed and evaluated. Performances are similar, but in favor of NMF, which is more robust to initialization, choice of the order and computationally less costly.
Nancy Bertin, Roland Badeau, Gaël Richard
ICASSP (1)3
2007 Detection and Analysis of Abnormal Situations Through Fear-Type Acoustic Manifestations
abstract
Recent work on emotional speech processing has demonstrated the interest to consider the information conveyed by the emotional component in speech to enhance the understanding of human behaviors. But to date, there has been little integration of emotion detection systems in effective applications. The present research focuses on the development of a fear-type emotions recognition system to detect and analyze abnormal situations for surveillance applications. The Fear vs. Neutral classification gets a mean accuracy rate at 70.3%. It corresponds to quite optimistic results given the diversity of fear manifestations illustrated in the data. More specific acoustic models are built inside the fear class by considering the context of emergence of the emotional manifestations, i.e. the type of the threat during which they occur, and which has a strong influence on fear acoustic manifestations. The potential use of these models for a threat type recognition system is also investigated. Such information about the situation can indeed be useful for surveillance systems.
Chloé Clavel, Laurence Devillers, Gaël Richard, Ioana Vasilescu, Thibaut Ehrette
ICASSP (4)3
2007 Combined Supervised and Unsupervised Approaches for Automatic Segmentation of Radiophonic Audio Streams
abstract
Speech/music discrimination is one of the most studied topics in the domain of audio data segmentation. In this paper, we propose and evaluate a novel method that includes feature selection and a combined supervised and unsupervised strategy for audio streams segmentation. A number of alternatives solutions for each component are assessed and the optimized system is compared to the approaches proposed in the framework of the ESTER campaign.
Gaël Richard, Mathieu Ramona, Slim Essid
ICASSP (2)1
2007 On the Correlation of Automatic Audio and Visual Segmentations of Music Videos
abstract
The study of the associations between audio and video content has numerous important applications in the fields of information retrieval and multimedia content authoring. In this work, we focus on music videos which exhibit a broad range of structural and semantic relationships between the music and the video content. To identify such relationships, a two-level automatic structuring of the music and the video is achieved separately. Note onsets are detected from the music signal, along with section changes. The latter is achieved by a novel algorithm which makes use of feature selection and statistical novelty detection approaches based on kernel methods. The video stream is independently segmented to detect changes in motion activity, as well as shot boundaries. Based on this two-level segmentation of both streams, four audio–visual correlation measures are computed. The usefulness of these correlation measures is illustrated by a query by video experiment on a 100 music video database, which also exhibits interesting genre dependencies.
Olivier Gillet, Slim Essid, Gaël Richard
IEEE Trans. Circuits Syst. Video Technol.3
2006 Yast Algorithm for Minor Subspace Tracking
abstract
This paper introduces a new algorithm for tracking the minor subspace of the correlation matrix associated with time series. This algorithm is shown to have a better convergence rate than existing methods. Moreover, it guarantees the orthonormality of the subspace weighting matrix at each iteration, and reaches a linear complexity
Roland Badeau, Bertrand David 0002, Gaël Richard
ICASSP (3)3
2006 Hrhatrac Algorithm for Spectral Line Tracking of Musical Signals
abstract
HRHATRAC combines the last improvements regarding the fast subspace tracking algorithms with a gradient update for adapting the signal poles estimates. It leads to a line spectral tracker which is able to robustly estimate the frequencies, even in a noisy context, when the lines are close to each other and when a modulation occurs. HRHATRAC is also successfully applied in this paper to a piano note recording
Bertrand David 0002, Roland Badeau, Gaël Richard
ICASSP (3)3
2006 Hierarchical Classification of Musical Instruments on Solo Recordings
abstract
We propose a study on the use of hierarchical taxonomies for musical instrument recognition on solo recordings. Both a natural taxonomy (inspired by instrument families) and a taxonomy inferred automatically by means of hierarchical clustering are examined. They are used to build a hierarchical classification scheme based on support vector machine classifiers and an efficient selection of features from a wide set of candidate descriptors. The classification results found with each taxonomy are compared and analysed. The automatic taxonomy is found to perform slightly better than the "natural" one. However, our analysis of the confusion matrices related to these taxonomies suggest that both are limited. In fact, it shows that it could be more advantageous to utilise taxonomies such that the instruments which are commonly confused are put in distinct decision nodes
Slim Essid, Gaël Richard, Bertrand David 0002
ICASSP (5)2
2006 Comparing Audio and Video Segmentations for Music Videos Indexing
abstract
Music videos are good examples of multimedia documents in which the structures of the audio and video streams are highly correlated. This paper presents a system that matches these structures and extracts audio-visual correlation measures. The audio and video streams are independently segmented at two-levels: shots (sections for audio) and events. Audio segmentation is performed at the event level by detecting onsets, and at the section level by a novelty detection algorithm identifying instrumentation changes. Video segmentation is performed at the event level by detecting changes in the motion intensity descriptor, and at the shot level by using a classical histogram-based shot detection algorithm. Audio-visual correlation measures are computed on the extracted structures. Possible applications include audio/video stream resynchronization, video retrieval from audio content, or classification of music videos by genre
Olivier Gillet, Gaël Richard
ICASSP (5)2
2006 Fear-type emotions of the SAFE Corpus: annotation issues
Chloé Clavel, Ioana Vasilescu, Laurence Devillers, Thibaut Ehrette, Gaël Richard
LREC5
2006 A new quantization optimization algorithm for the MPEG advanced audio coder using a statistical subband model of the quantization noise
abstract
In this paper, an improvement of the quantization optimization algorithm for the MPEG Advanced Audio Coder (AAC) is presented. This algorithm, given a bit-rate constraint, minimizes the perceived distortion generated by the signal compression. The distortion can be related to the quantization error level over frequency subbands through an auditory model. Thus, optimizing the quantization requires knowledge of the rate-distortion function for each subband. When this function can be modeled in a simple way, the algorithm can take a one-loop recursive structure. However, in the MPEG AAC, the rate-distortion function is hard to characterize, since AAC makes use of nonlinear quantizers and variable length entropy coders. As a result, the standard algorithm makes use of two nested loops with a local decoder, in order to measure the error level rather than predicting its value. We first describe a partial subband modeling of the rate-distortion function of interest in the MPEG AAC. Then, using a statistical approach, we find a relationship between the error level and the so-called quantization "scale-factor" and propose a new algorithm that is basically similar to a classical one loop "bit allocation" process. Finally, we describe the complete algorithm and show that it is more efficient than the standard one
Olivier Derrien, Pierre Duhamel, Maurice Charbit, Gaël Richard
IEEE Trans. Speech Audio Process.4
2006 Instrument recognition in polyphonic music based on automatic taxonomies
abstract
We propose a new approach to instrument recognition in the context of real music orchestrations ranging from solos to quartets. The strength of our approach is that it does not require prior musical source separation. Thanks to a hierarchical clustering algorithm exploiting robust probabilistic distances, we obtain a taxonomy of musical ensembles which is used to efficiently classify possible combinations of instruments played simultaneously. Moreover, a wide set of acoustic features is studied including some new proposals. In particular, signal to mask ratios are found to be useful features for audio classification. This study focuses on a single music genre (i.e., jazz) but combines a variety of instruments among which are percussion and singing voice. Using a varied database of sound excerpts from commercial recordings, we show that the segmentation of music with respect to the instruments played can be achieved with an average accuracy of 53%.
Slim Essid, Gaël Richard, Bertrand David 0002
IEEE Trans. Speech Audio Process.2
2006 Musical instrument recognition by pairwise classification strategies
abstract
Musical instrument recognition is an important aspect of music information retrieval. In this paper, statistical pattern recognition techniques are utilized to tackle the problem in the context of solo musical phrases. Ten instrument classes from different instrument families are considered. A large sound database is collected from excerpts of musical phrases acquired from commercial recordings translating different instrument instances, performers, and recording conditions. More than 150 signal processing features are studied including new descriptors. Two feature selection techniques, inertia ratio maximization with feature space projection and genetic algorithms are considered in a class pairwise manner whereby the most relevant features are fetched for each instrument pair. For the classification task, experimental results are provided using Gaussian mixture models (GMMs) and support vector machines (SVMs). It is shown that higher recognition rates can be reached with pairwise optimized subsets of features in association with SVM classification using a radial basis function kernel
Slim Essid, Gaël Richard, Bertrand David 0002
IEEE Trans. Speech Audio Process.2
2005 Yet another subspace tracker
abstract
The paper introduces a new algorithm for tracking the dominant subspace of the correlation matrix associated with time series. This algorithm greatly outperforms many well-known subspace trackers in terms of subspace estimation. Moreover, it guarantees the orthonormality of the subspace weighting matrix at each iteration, and reaches the lowest complexity found in the literature.
Roland Badeau, Bertrand David 0002, Gaël Richard
ICASSP (4)3
2005 Instrument recognition in polyphonic music
abstract
We propose a method for the recognition of musical instruments in polyphonic music excerpted from commercial recordings. By exploiting some cues on the common structures of musical ensembles, we show that it is possible to recognize up to 4 instruments playing concurrently. The system associates a hierarchical classification tree with a class-pairwise feature selection technique and Gaussian mixture models to discriminate possible combinations of instruments. Successful identification is achieved over short-time windows, enabling the system to be employed for segmentation purposes.
Slim Essid, Gaël Richard, Bertrand David 0002
ICASSP (3)2
2005 Automatic transcription of drum sequences using audiovisual features
abstract
The transcription of a musical performance from the audio signal is often problematic, either because it requires the separation of complex sources, or simply because some important high-level music information cannot be directly extracted from the audio signal. We propose a novel multimodal approach for the transcription of drum sequences using audiovisual features. The transcription is performed by support vector machine (SVM) classifiers, and three different information fusion strategies are evaluated. A correct recognition rate of 85.8% can be achieved for a detailed taxonomy and a fully automated transcription.
Olivier Gillet, Gaël Richard
ICASSP (3)2
2005 Iterative algorithms for multichannel equalization in sound reproduction systems
abstract
A fast iterative algorithm, with computation based on the fast Fourier transform (FFT), is presented. It can be used to control a sound field at several control points with a loudspeaker array from multiple reference signals. It designs an equalizer able to invert long FIR filters and which achieves better performance than traditional FFT-based deconvolution methods with an equal number of coefficients in the inverse filters.
Mathieu Franck Guillaume, Yves Grenier, Gaël Richard
ICASSP (3)3
2005 Extracting note onsets from musical recordings
abstract
Automatic temporal segmentation of music signals into note onsets is central for a large number of audio applications. In this paper, we present a variation of a previously existing note onset detection method, based on the so-called spectral energy flux. The proposed algorithm has a lower computational cost and incorporates a more accurate estimation of the frequency content derivative, yielding better results for a wide range of music signals. The performance of the system was validated using a database of musical recordings containing 670 note onsets. This database was hand-labeled and cross validated by three annotators. Comparisons to previous work are also presented along with possible directions of future research.
Miguel A. Alonso-Arévalo, Gaël Richard, Bertrand David 0002
ICME2
2005 Events Detection for an Audio-Based Surveillance System
abstract
The present research deals with audio events detection in noisy environments for a multimedia surveillance application. In surveillance or homeland security most of the systems aiming to automatically detect abnormal situations are only based on visual clues while, in some situations, it may be easier to detect a given event using the audio information. This is in particular the case for the class of sounds considered in this paper, sounds produced by gun shots. The automatic shot detection system presented is based on a novelty detection approach which offers a solution to detect abnormality (abnormal audio events) in continuous audio recordings of public places. We specifically focus on the robustness of the detection against variable and adverse conditions and the reduction of the false rejection rate which is particularly important in surveillance applications. In particular, we take advantage of potential similarity between the acoustic signatures of the different types of weapons by building a hierarchical classification system
Chloé Clavel, Thibaut Ehrette, Gaël Richard
ICME3
2005 Drum Loops Retrieval from Spoken Queries
Olivier Gillet, Gaël Richard
J. Intell. Inf. Syst.2
2004 Selecting the modeling order for the ESPRIT high resolution method: an alternative approach
abstract
High resolution methods, such as the ESPRIT (estimation of signal parameters by rotational invariance techniques) algorithm, perform an accurate representation of a harmonic signal as a sum of exponentially damped sinusoids. However, in coding applications, the signal must be represented with a minimum number of parameters. Unfortunately, it is well known that applying the ESPRIT algorithm with an under-estimated model order generates biased frequency estimates. We propose a new method for selecting an appropriate modeling order, which minimizes this bias. This approach was applied to both synthetic and musical signals and outperformed the classical information theoretic criteria.
Roland Badeau, Bertrand David 0002, Gaël Richard
ICASSP (2)3
2004 Automatic transcription of drum loops
abstract
Recent efforts in audio indexing and retrieval in music databases mostly focus on melody. If this is appropriate for polyphonic music signals, specific approaches are needed for systems dealing with percussive audio signals such as those produced by drums, tabla or djembe. Most studies of drum signal transcription focus on sounds taken in isolation. In this paper, we propose several methods for drum loop transcription where the drums signals dataset reflects the variability encountered in modern audio recordings (real and natural drum kits, audio effects, simultaneous instruments, etc.). The approaches described are based on hidden Markov models (HMM) and support vector machines (SVM). Promising results are obtained with a 83.9% correct recognition rate for a simplified taxonomy.
Olivier Gillet, Gaël Richard
ICASSP (4)2
2003 Sliding window orthonormal PAST algorithm
abstract
This paper introduces an orthonormal version of the sliding-window projection approximation subspace tracker (PAST). The new algorithm guarantees the orthonormality of the signal subspace basis at each iteration. Moreover, it has the same complexity as the original PAST algorithm, and like the more computationally demanding natural power (NP) method, it satisfies a global convergence property, and reaches an excellent tracking performance.
Roland Badeau, Karim Abed-Meraim, Gaël Richard, Bertrand David 0002
ICASSP (5)3
2003 Adaptive ESPRIT algorithm based on the PAST subspace tracker
abstract
The Estimation of Signal Parameters via Rotational Invariance Techniques (ESPRIT) algorithm is a subspace-based analysis method used in source localization or frequency estimation, originally designed in a block signal processing context. In other respects, the Projection Approximation Subspace Tracker (PAST) is a fast and robust subspace tracking method. This paper introduces a new frequency estimation and tracking algorithm, which relies on the PAST subspace tracker and a fast adaptive implementation of the ESPRIT algorithm.
Roland Badeau, Gaël Richard, Bertrand David 0002
ICASSP (6)2
2000 SPEECHDAT-CAR. A Large Speech Database for Automotive Environments
Asunción Moreno, Børge Lindberg, Christoph Draxler, Gaël Richard, Khalid Choukri, Stephan Euler, Jeffrey Allen
LREC4
1999 Compensating for variable recording conditions in frontal face authentication algorithms
abstract
This paper addresses the problem of compensating for variable recording conditions such as changes in illumination, scale differences, and varying face position. It is well known that the performance of any face authentication/recognition algorithm deteriorates significantly in the presence of the aforementioned conditions as well as the expression variations. The use of simple and powerful pre-processing techniques aiming at compensating for variable recording conditions prior to the application of any authentication algorithm is proposed. It is shown that such an approach overcomes indeed the image variations and guarantees an almost stable performance for the Morphological Dynamic Link Architecture developed within the European research project M2VTS.
Anastasios Tefas, Yann Menguy, Constantine Kotropoulos, Gaël Richard, Ioannis Pitas, Philip Lockwood
ICASSP4
1999 The speechdat-car multilingual speech databases for in-car applications: some first validation results
abstract
International audience
Henk van den Heuvel, Jérôme Boudy, Robrecht Comeyne, Stephan Euler, Asunción Moreno, Gaël Richard
EUROSPEECH6
1997 Voice mimic system using an articulatory codebook for estimation of vocal tract shape
abstract
VOICEMIMICSYSTEMUSINGANARTICULATORYCODEBOOKFORESTIMATIONOFVOCALTRACTSHAPES. Chennoukh, D. Sinder, G. Richard* and J.L. FlanaganCenter for Computer Aids for Industrial Pro ductivity (CAIP), Rutgers University,Piscataway, NJ 08855-1390, USA*Matra-Communication, rue J.P. Timbaud, 78392 Bois d'Arcy,FranceTel.+1 908-445-0080, FAX: +1 908 445-4775, E-mail:[email protected] mimic systems using articulatory co deb o oks re-quireaninitialestimateofthevo caltractshap einthe vicinity of the global optimum.For this purp ose,we need to gather a large set of corresp onding articu-latory and acoustic data in the articulatory co deb o ok.Thus, searching and accessing the co deb o ok b ecomesa dicult task.In this pap er, the design of an artic-ulatory co deb o ok is presented where an acoustic net-work sub-samples the acoustic space such that vo caltract mo del shap es are ordered and clustered in thenetwork according toacousticparameters.Anotherissue addressed in this pap er concerns estimating thetra jectory of vo cal tract shap es as they change withtime.Sincetheinversemappingfromacousticpa-rameters to mo del shap e do es not have a unique so-lution, several vo cal tract shap e variations are p ossi-ble.Therefore, a dynamic optimization of tra jectorieshas b een develop ed.This optimization uses dynamicprop erties of each articulatory parameter to estimatethe next p osition.1.INTRODUCTIONThestudyofsp eechp erceptionandpro duc-tion has b een enhanced in the last two decades by thedevelopment of computers capable of large amountofcomputation.As a result, Stevens' study towards anarticulatorymo delforsp eechrecognition-synthesisb ecomesmorefeasiblethanitwasintheearlysix-ties([9]).However, an incomplete understanding ofsp eechpro ductionandtheacousticsofpre-ventedusfromachievingStevens'goal.Thegoalwastomimicinputsp eechsignalsbyrecognition-synthesis using a mo del of the vo cal tract area func-tionthatcanmimicthesp eechsignalswithoutun-derstanding their structure or meaning.An early attempt at creating a complete computersimulation of articulatory mo del sp eech co ding usingan optimization technique was rep orted by Flanaganetal.([4]).Thesimulationiscalled\voice mimic.The voice mimic attempts to provide an articulatorydescription of the vo cal tract that corresp onds to anarbitrary natural sp eech input and to generate a syn-thetic signal that, within p erceptual accuracy, dupli-catesthenaturalone.Centraltoe ortisinverse mapping from an acoustic signal to an articu-latory description.However, acoustic-to-articulatorymappings are non-unique and, given a cost function,the optimization techniques converge only to a lo calextremum that may b e near the vicinity of the initialparameters.Therefore, one needs to cho ose accuratestartup parameters to initialize the optimization pro-cedure.Schro eter andSondhi([8]),whocontinuedalong the same lines of Flanagan et al.'s study, usedan articulatory co deb o ok prop osed earlier byAtal etal.([1]).Since a co deb o ok is used to obtain the rstestimates of the vo cal tract shap e that may pro ducea given combination of acoustic parameters, it mustbedesignedsuchthatitspansthenatural articula-tory space of a sp eaker.Furthermore, sampling of thespace must b e ne enough so that an acoustic entryalways exists very close to the global optimum.Suchco deb o oks require a large set of matching pairs of vo-cal tract and acoustic parameters.The complexityofsearching a large co deb o ok for all p ossible vo cal tractmo del shap es b ecomes an issue.For this reason, thevoice mimic system needs, in addition to a go o d artic-ulatory co deb o ok, an ecient pro cedure for accessingthe co deb o ok ([6],[7]).The numb er and p osition of the co deb o ok vectorsa ect the p erformance of the voice mimic system ac-cording to two compromising problems.On one hand,increasing the size of the co deb o ok increases the dif- cultyoftheaccesstaskand,onotherhand,reductionofthissizecomplicatestheinverseprob-lemsolution.Inthesecond sectionofthispap er,anew design of the articulatory co deb o ok is presentedfor which the inversion of the articulatory-to-acousticmapping is pro cessed during the building of the co de-b o ok.Thisco deb o okdesignallowsreal-time accessto the set of acoustically equivalent shap es, regardlessthe size of the co deb o ok.Sincetheinversemappingfromacousticparam-eterstomo delshap edo esnothaveauniquesolu-tion,severalvo caltractshap eariationsarep ossi-ble.Schro eter and Sondhi([7]) prop osed the use ofdynamicprogrammingtoestimatetheoptimaltra-jectory of the vo cal tract mo del shap e variation path.The dynamic programming requires a delay of severaldata frames for the sp eech output ([8]).In the thirdsection, a metho d is prop osed where the articulatoryparameters are estimated within one frame.Section
Samir Chennoukh, Daniel J. Sinder, Gaël Richard, James L. Flanagan
EUROSPEECH3
1996 Analysis/synthesis and modification of the speech aperiodic component
Gaël Richard, Christophe d'Alessandro
Speech Commun.1
1995 Numerical simulations of fluid flow in the vocal tract
Gaël Richard, D. Snider, H. Duncan, Qiguang Lin, James L. Flanagan, Stephen E. Levinson, Donald Davis, Scott Slimon
EUROSPEECH1
1993 A speech formant synthesizer based on harmonic + random formant-waveforms representations
abstract
This paper describes a new type of speech synthe sizer a parametric concatenation PACO speech syn thesizer which is suitable both for formant synthe sis and concatenation synthesis This synthesizer is based on a hybrid quasi harmonic and random formant waveforms model of the speech signal The synthesizer can be controlled by acoustic parameters formant pa rameters and voice source parameters expressed in fre quency domain These acoustic parameters are con verted into sinusoidal and formant waveforms parame ters The keypoint of this method is that spectral am plitudes are set according to a parallel formant model both on sinusoidal waveforms and random formant waveforms whereas the spectral phases are set accord ing to a serial formant model for sinusoidal waveforms and are randomly distributed for formant waveforms This approach avoids the phase interference problems inherent to parallel synthesis while keeping the ad vantage of formant amplitudes control An automatic analysis synthesis system is also proposed for segments coding Our model has been successfully implemented both as a formant synthesis system and in a concate nation synthesis Text To Speech system keywords Speech synthesis formant and concate nation synthesis harmonic representation random for mant waveforms
Sophie Grau, Christophe d'Alessandro, Gaël Richard
EUROSPEECH3