Axel Röbel

dblp:34/3702 · also Axel Roebel · DBLP profile ↗
← Back
59ranked-venue papers
8as first author
7since 2021 · last 2025
0000-0001-6136-4391ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 44 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 29 · 6 first-author · 5 since 2021
YearPublicationVenuePosition
2025 MusicGen-Stem: Multi-stem music generation and edition through autoregressive modeling
abstract
While most music generation models generate a mixture of stems (in mono or stereo), we propose to train a multi-stem generative model with 3 stems (bass, drums and other) that learn the musical dependencies between them. To do so, we train one specialized compression algorithm per stem to tokenize the music into parallel streams of tokens. Then, we leverage recent improvements in the task of music source separation to train a multi-stream text-to-music language model on a large dataset. Finally, thanks to a particular conditioning method, our model is able to edit bass, drums or other stems on existing or generated songs as well as doing iterative composition (e.g. generating bass on top of existing drums). This gives more flexibility in music generation algorithms and it is to the best of our knowledge the first open-source multi-stem autoregressive music generation model that can perform good quality generation and coherent source editing. Code and model weights will be released and samples are available on simonrouard.github.io/musicgenstem.
Simon Rouard, Robin San-Roman, Yossi Adi, Axel Röbel
ICASSP4
2024 Small-E: Small Language Model with Linear Attention for Efficient Speech Synthesis
abstract
Recent advancements in text-to-speech (TTS) powered by language models have showcased remarkable capabilities in achieving naturalness and zero-shot voice cloning. Notably, the decoder-only transformer is the prominent architecture in this domain. However, transformers face challenges stemming from their quadratic complexity in sequence length, impeding training on lengthy sequences and resource-constrained hardware. Moreover they lack specific inductive bias with regards to the monotonic nature of TTS alignments. In response, we propose to replace transformers with emerging recurrent architectures and introduce specialized cross-attention mechanisms for reducing repeating and skipping issues. Consequently our architecture can be efficiently trained on long samples and achieve state-of-the-art zero-shot voice cloning against baselines of comparable size. Our implementation and demos are available at https://github.com/theodorblackbird/lina-speech.
Théodor Lemerle, Nicolas Obin, Axel Röbel
INTERSPEECH3
2023 Analysis and Transformation of Voice Level in Singing Voice
abstract
We introduce a neural auto-encoder that transforms the musical dynamic in recordings of singing voice via changes in voice level. Since most recordings of singing voice are not annotated with voice level we propose a means to estimate the voice level from the signal’s timbre using a neural voice level estimator. We introduce the recording factor that relates the voice level to the recorded signal power as a proportionality constant. This unknown constant depends on the recording conditions and the post-processing and may thus be different for each recording (but is constant across each recording). We provide two approaches to estimate the voice level without knowing the recording factor. The unknown recording factor can either be learned alongside the weights of the voice level estimator, or a special loss function based on the scalar product can be used to only match the contour of the recorded signal’s power. The voice level models are used to condition a previously introduced bottleneck auto-encoder that disentangles its input, the mel-spectrogram, from the voice level. We evaluate the voice level models on recordings annotated with musical dynamic and by their ability to provide useful information to the auto-encoder. A perceptive test is carried out that evaluates the perceived change in voice level in transformed recordings and the synthesis quality. The perceptive test confirms that changing the conditional input changes the perceived voice level accordingly thus suggesting that the proposed voice level models encode information about the true voice level.
Frederik Bous, Axel Röbel
ICASSP2
2022 Production Strategies of Vocal Attitudes
abstract
Humans have an impressive ability to communicate precise social intentions and desires with their voice - through vocal attitudes. Previous studies have shown how isolated acoustic features such as pitch can convey social attitudes, but have mostly worked with single attitudes and have not controlled for inter-speaker variability. Thus, the vocal behaviours used to produce social attitudes remain mostly unknown. That is the aim of the current study, to uncover the anatomic production strategies that speakers use to communicate vocal attitudes. To do this, we analysed recordings from N=20 French speakers producing dominant, friendly, seductive and distant speech. For each of these attitudes, we investigated their vocal fold behaviour, vocal tract actuation and phonetic speech structure, with the support of deep alignment methods, and compared them with group statistics. We notably produced high-level representations of speakers' articulation (e.g. Vowel Space Density) and speech rhythm. Our results reveal speakers' prototypical strategies to produce vocal attitudes, and highlight how vocal behaviours can communicate social signals. We expect these results to provide an objective validation method for deep voice attitude conversions.
Léane Salais, Pablo Arias 0003, Clément Le Moine, Victor Rosi, Yann Teytaut, Nicolas Obin, Axel Röbel
INTERSPEECH7
2022 A study on constraining Connectionist Temporal Classification for temporal audio alignment
abstract
International audience
Yann Teytaut, Baptiste Bouvier, Axel Röbel
INTERSPEECH3
2021 Speaker Attentive Speech Emotion Recognition
abstract
International audience
Clément Le Moine, Nicolas Obin, Axel Röbel
Interspeech3
2021 Phoneme-to-Audio Alignment with Recurrent Neural Networks for Speaking and Singing Voice
abstract
International audience
Yann Teytaut, Axel Röbel
Interspeech2
2020 GCI Detection from Raw Speech Using a Fully-Convolutional Network
abstract
Glottal Closure Instants (GCI) detection consists in automatically detecting temporal locations of most significant excitation of the vocal tract from the speech signal. It is used in many speech analysis and processing applications, and various algorithms have been proposed for this purpose. Recently, new approaches using convolutional neural networks have emerged, with encouraging results. Following this trend, we propose a simple approach that performs a mapping from the speech waveform to a target signal from which the GCIs are obtained by peak-picking. However, the ground truth GCIs used for training and evaluation are usually extracted from EGG signals, which are not perfectly reliable and often not available. To overcome this problem, we propose to train our network on high-quality synthetic speech with perfect ground truth. The performances of the proposed algorithm are compared with three other state-of-the-art approaches using publicly available datasets, and the impact of using controlled synthetic or real speech signals in the training stage is investigated. The experimental results demonstrate that the proposed method obtains similar or better results than other state-of-the-art algorithms and that using large synthetic datasets with many speakers offers a better generalization ability than using a smaller database of real speech and EGG signals.
Luc Ardaillon, Axel Röbel
ICASSP2
2020 Sound Texture Synthesis Using RI Spectrograms
abstract
This article introduces a new parametric synthesis method for sound textures based on existing works in visual and sound texture synthesis. Starting from a base sound signal, an optimization process is performed until the cross-correlations between the feature-maps of several untrained 2D Convolutional Neural Networks (CNN) resemble those of an original sound texture. We use compressed Real-Imaginary (RI) spectrograms as input to the CNN: this time-frequency representation is the stacking of the real and imaginary part of the Short Time Fourier Transform (STFT) and thus implicitly contains both the magnitude and phase information, allowing for convincing syntheses of various audio events. The optimization is however performed directly on the time signal to avoid any STFT consistency issue. The results of an online perceptual evaluation are also detailed, and show that this method achieves results that are more realistic-sounding than existing parametric methods on a wide array of textures.
Hugo Caracalla, Axel Röbel
ICASSP2
2020 Realistic Transformation of Facial and Vocal Smiles in Real-Time Audiovisual Streams
abstract
Research in affective computing and cognitive science has shown the importance of emotional facial and vocal expressions during human-computer and human-human interactions. But, while models exist to control the display and interactive dynamics of emotional expressions, such as smiles, in embodied agents, these techniques can not be applied to video interactions between humans. In this work, we propose an audiovisual smile transformation algorithm able to manipulate an incoming video stream in real-time to parametrically control the amount of smile seen on the user's face and heard in their voice, while preserving other characteristics such as the user's identity or the timing and content of the interaction. The transformation is composed of separate audio and visual pipelines, both based on a warping technique informed by real-time detection of audio and visual landmarks. Taken together, these two parts constitute a unique audiovisual algorithm which, in addition to providing simultaneous real-time transformations of a real person's face and voice, allows to investigate the integration of both modalities of smiles in real-world social interactions.
Pablo Arias 0003, Catherine Soladié, Oussema Bouafif, Axel Röbel, Renaud Séguier, Jean-Julien Aucouturier
IEEE Trans. Affect. Comput.4
2019 Sequence-to-sequence Modelling of F0 for Speech Emotion Conversion
abstract
Voice interfaces are becoming wildly popular and driving demand for more advanced speech synthesis and voice transformation systems. Current text-to-speech methods produce realistic sounding voices, but they lack the emotional expressivity that listeners expect, given the context of the interaction and the phrase being spoken. Emotional voice conversion is a research domain concerned with generating expressive speech from neutral synthesised speech or natural human voice. This research investigated the effectiveness of using a sequence-to-sequence (seq2seq) encoder-decoder based model to transform the intonation of a human voice from neutral to expressive speech, with some preliminary introduction of linguistic conditioning. A subjective experiment conducted on the task of speech emotion recognition by listeners successfully demonstrated the effectiveness of the proposed sequence-to-sequence models to produce convincing voice emotion transformations. In particular, conditioning the model on the position of the syllable in the phrase significantly improved recognition rates.
Carl Robinson, Nicolas Obin, Axel Röbel
ICASSP3
2019 Fully-Convolutional Network for Pitch Estimation of Speech Signals
abstract
International audience
Luc Ardaillon, Axel Röbel
INTERSPEECH2
2018 Binaural Localization of Multiple Sound Sources by Non-Negative Tensor Factorization
abstract
This paper presents non-negative factorization of audio signals for the binaural localization of multiple sound sources within realistic and unknown sound environments. Non-negative tensor factorization (NTF) provides a sparse representation of multichannel audio signals in time, frequency, and space that can be exploited in computational audio scene analysis and robot audition for the separation and localization of sound sources. In the proposed formulation, each sound source is represented by means of spectral dictionaries, temporal activation, and its distribution within each channel (here, left and right ears). This distribution, being dependent on the frequency, can be interpreted as an explicit estimation of the Head-Related Transfer Function (HRTF) of a binaural head which can then be converted into the estimated sound source position. Moreover, the semisupervised formulation of the non-negative factorization allows us to integrate prior knowledge about some sound sources of interest whose dictionaries can be learned in advance, whereas the remaining sources are considered as background sound, which remains unknown and is estimated on the fly. The proposed NTF-based sound source localization is applied here to binaural sound source localization of multiple speakers within realistic sound environments.
Elie-Laurent Benaroya, Nicolas Obin, Marco Liuni, Axel Röbel, Wilson Raumel, Sylvain Argentieri
IEEE ACM Trans. Audio Speech Lang. Process.4
2017 A Mouth Opening Effect Based on Pole Modification for Expressive Singing Voice Transformation
abstract
International audience
Luc Ardaillon, Axel Röbel
INTERSPEECH2
2016 A source/filter model with adaptive constraints for NMF-based speech separation
abstract
This paper introduces a constrained source/filter model for semi-supervised speech separation based on non-negative matrix factorization (NMF). The objective is to inform NMF with prior knowledge about speech, providing a physically meaningful speech separation. To do so, a source/filter model (indicated as Instantaneous Mixture Model or IMM) is integrated in the NMF. Furthermore, constraints are added to the IMM-NMF, in order to control the NMF behaviour during separation, and to enforce its physical meaning. In particular, a speech specific constraint - based on the source/filter coherence of speech - and a method for the automatic adaptation of constraints' weights during separation are presented. Also, the proposed source/filter model is semi-supervised: during training, one filter basis is estimated for each phoneme of a speaker; during separation, the estimated filter bases are then used in the constrained source/filter model. An experimental evaluation for speech separation was conducted on the TIMIT speakers database mixed with various environmental background noises from the QUT-NOISE database. This evaluation showed that the use of adaptive constraints increases the performance of the source/filter model for speaker-dependent speech separation, and compares favorably to fully-supervised speech separation.
Damien Bouvier, Nicolas Obin, Marco Liuni, Axel Röbel
ICASSP4
2016 Simple multi frame analysis methods for estimation of amplitude spectral envelope estimation in singing voice
abstract
In the state of the art, a single frame of DFT transform is commonly used as a basis for building amplitude spectral envelopes. Multiple Frame Analysis (MFA) has already been suggested for envelope estimation, but often with excessive complexity. In this paper, two MFA-based methods are presented: one simplifying an existing Least Square (LS) solution, and another one based on a simple linear interpolation. In the context of singing voice we study sustained segments with vibrato, because these ones are obviously critical for singing voice synthesis. They also provide a convenient context to study, prior to extension of this work in more general contexts. Numerical and perceptual experiments show clear improvements of the two methods described compared to the state of the art and encourage further studies in this research direction.
Gilles Degottex, Luc Ardaillon, Axel Röbel
ICASSP3
2016 Expressive Control of Singing Voice Synthesis Using Musical Contexts and a Parametric F0 Model
abstract
International audience
Luc Ardaillon, Celine Chabot-Canet, Axel Röbel
INTERSPEECH3
2016 Evaluation of Singing Synthesis: Methodology and Case Study with Concatenative and Performative Systems
abstract
International audience
Lionel Feugère, Christophe d'Alessandro, Samuel Delalez, Luc Ardaillon, Axel Röbel
INTERSPEECH5
2016 Multi-Frame Amplitude Envelope Estimation for Modification of Singing Voice
abstract
Singing voice synthesis benefits from very high quality estimation of the resonances and anti-resonances of the vocal tract filter (VTF), i.e., an amplitude spectral envelope. In the state of the art, a single frame of DFT transform is commonly used as a basis for building spectral envelopes. Even though multiple frame analysis (MFA) has already been suggested for envelope estimation, it is not yet used in concrete applications. Indeed, even though existing attempts have shown very interesting results, we will demonstrate that they are either over complicated or fail to satisfy the high accuracy that is necessary for singing voice. In order to allow future applications of MFA, this article aims to improve the theoretical understanding and advantages of MFA-based methods. The use of singing voice signals is very beneficial for studying MFA methods due to the fact that the VTF configuration can be relatively stable and, at the same time, the vibrato creates a regular variation that is easy to model. By simplifying and extending previous works, we also suggest and describe two MFA-based methods. To better understand the behaviors of the envelope estimates, we designed numerical measurements to assess single frame analysis and MFA methods using synthetic signals. With listening tests, we also designed two proofs of concept using pitch scaling and conversion of timbre. Both evaluations show clear and positive results for MFA-based methods, thus, encouraging this research direction for future applications.
Gilles Degottex, Luc Ardaillon, Axel Röbel
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 A Morphological Model for Simulating Acoustic Scenes and Its Application to Sound Event Detection
abstract
This paper introduces a model for simulating environmental acoustic scenes that abstracts temporal structures from audio recordings. This model allows us to explicitly control key morphological aspects of the acoustic scene and to isolate their impact on the performance of the system under evaluation. Thus, more information can be gained on the behavior of an evaluated system, providing guidance for further improvements. To demonstrate its potential, this model is employed to evaluate the performance of nine state of the art sound event detection systems submitted to the IEEE DCASE 2013 Challenge. Results indicate that the proposed scheme is able to successfully build datasets useful for evaluating important aspects of the performance of sound event detection systems, such as their robustness to new recording conditions and to varying levels of background audio.
Grégoire Lafay, Mathieu Lagrange, Mathias Rossignol, Emmanouil Benetos, Axel Röbel
IEEE ACM Trans. Audio Speech Lang. Process.5
2016 A Montage Approach to Sound Texture Synthesis
abstract
Sound texture synthesis has applications in creating audio scenes for film and video games. In this paper, a novel algorithm for sound texture synthesis is presented. The goal of this algorithm is to produce new examples of a given sampled texture, the synthesized textures being of any desired duration. The algorithm is based on a montage approach to synthesis in that the original sample is cut into small pieces, referred to as atoms, and these atoms are concatenated together in a new sequence, preserving certain structures of the original texture. The sequence modelling of the atoms has two levels: atoms are concatenated to create segments and segments are concatenated, based on their history, to create textures. This approach deals with problems of repetition associated with sampling based sound texture synthesis techniques. Listening tests show that the results of the synthesis are very promising for a broad range of textures, including quasi-periodic and more random textures.
Seán O'Leary, Axel Röbel
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Similarity Search of Acted Voices for Automatic Voice Casting
abstract
This paper presents a large-scale similarity search of professionally acted voices for computer-aided voice casting. The proposed voice casting system explores Gaussian mixture model-based acoustic models and multilabel recognition of perceived paralinguistic content (speaker states and speaker traits, e.g., age/gender, voice quality, emotion) for the voice casting of professionally acted voices. First, acoustic models (universal background model, super-vector, i-vector) are constructed to model the acoustic space of voices, from which the similarity between voices can be measured directly in the acoustic space. Second, multiple binary classification of speaker traits and states is added to the acoustic models in order to represent the vocal signature of a voice, which is then used to measure the similarity between voices in the paralinguistic space. Finally, a similarity search is processed in order to determine the set of target actors that are the most similar to the voice of a source actor. In a subjective experiment conducted in the real-context of cross-language voice casting, the multilabel scoring system significantly outperforms the acoustic scoring system. This constitutes a proof of concept for the role of perceived para-linguistic categories in the perception of voice similarity.
Nicolas Obin, Axel Röbel
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 One-formant vocal tract modeling for glottal pulse shape estimation
abstract
This work considers the task of estimating the source and filter from human voice signals. Since the energy of voiced sound concentrates on discrete frequencies, a notable challenge with this task would be that higher pitches in the signal can make the harmonically related frequency response samples of the vocal tract filter an incomplete representation. In view of this, we propose to model the magnitude and phase response of the first formant as an alternative to the minimum phase property of the vocal tract filter. In particular, the magnitude response of the vocal tract filter sampled at the first three partials only, is sufficient for determining the phase response of the first formant. We verified our new method with glottal pulse shape parameter estimation experiments conducted on the CMU Arctic dataset, which showed that single-formant filter is an adequate alternative to minimum-phase filter in vocal tract modeling for glottal pulse shape estimation.
Yu-Ren Chien, Axel Röbel
ICASSP2
2015 The role of glottal source parameters for high-quality transformation of perceptual age
abstract
The intuitive control of voice transformation (e.g., age/sex, emotions) is useful to extend the expressive repertoire of a voice. This paper explores the role of glottal source parameters for the control of voice transformation. First, the SVLN speech synthesizer (Separation of the Vocal-tract with the Liljencrants-fant model plus Noise) is used to represent the glottal source parameters (and thus, voice quality) during speech analysis and synthesis. Then, a simple statistical method is presented to control speech parameters during voice transformation: a GMM is used to model the speech parameters of a voice, and regressions are then used to adapt the GMMs statistics (mean and variance) to a control parameter (e.g., age/sex, emotions). A subjective experiment conducted on the control of perceptual age proves the importance of the glottal source parameters for the control of voice transformation, and shows the efficiency of the statistical model to control voice parameters while preserving a high-quality of the voice transformation.
Xavier Favory, Nicolas Obin, Gilles Degottex, Axel Röbel
ICASSP4
2015 On automatic drum transcription using non-negative matrix deconvolution and itakura saito divergence
abstract
This paper presents an investigation into the detection and classification of drum sounds in polyphonic music and drum loops using non-negative matrix deconvolution (NMD) and the Itakura Saito divergence. The Itakura Saito divergence has recently been proposed as especially appropriate for decomposing audio spectra due to the fact that it is scale invariant, but it has not yet been widely adopted. The article studies new contributions for audio event detection methods using the Itakura Saito divergence that improve efficiency and numerical stability, and simplify the generation of target pattern sets. A new approach for handling background sounds is proposed and moreover, a new detection criteria based on estimating the perceptual presence of the target class sources is introduced. Experimental results obtained for drum detection in polyphonic music and drum soli demonstrate the beneficial effects of the proposed extensions.
Axel Röbel, Jordi Pons, Marco Liuni, Mathieu Lagrange
ICASSP1
2015 A multi-layer F0 model for singing voice synthesis using a b-spline representation with intuitive controls
abstract
In singing voice, the fundamental frequency (F0) carries not only melody, but also music style, personal expressivity and other characteristics specific to voice production mechanism. The F0 modeling is therefore critical for a natural-sounding and expressive synthesis. In addition, for artistic purposes, composers also need to have control over expressive parameters of the F0 curve, which is missing in many current approaches. This paper presents a novel parametric F0 model for singing voice synthesis with intuitive control of expressive parameters. The proposed approach considers the various F0 variations of the singing voice as separate layers using B-splines to model the melodic component. This model has been implemented in a concatenative singing voice synthesis system and its perceived naturalness has been evaluated through listening tests. The validity of each layer is first evaluated independently, and the full model is then compared to real F0 curves from professional singers. The results of these tests suggest that the model is suitable to produce natural and expressive F0 contours.
Luc Ardaillon, Gilles Degottex, Axel Röbel
INTERSPEECH3
2015 On glottal source shape parameter transformation using a novel deterministic and stochastic speech analysis and synthesis system
abstract
In this paper we present a flexible deterministic plus stochastic model (DSM) approach for parametric speech analysis and synthesis with high quality. The novelty of the proposed speech processing system lies in its extended means to estimate the unvoiced stochastic component and to robustly handle the transformation of the glottal excitation source. It is therefore well suited as speech system within the context of Voice Transformation and Voice Conversion. The system is evaluated in the context of a voice quality transformation on natural human speech. The voice quality of a speech phrase is altered by means of resynthesizing the deterministic component with different pulse shapes of the glottal excitation source. A subjective listening test suggests that the speech processing system is able to successfully synthesize and arise to a listener the perceptual sensation of different voice quality characteristics. Additionally, improvements of the speech synthesis quality compared to a baseline method are demonstrated.
Stefan Huber 0003, Axel Röbel
INTERSPEECH2
2014 A Two Level Montage Approach to Sound Texture Synthesis with Treatment of Unique Events
Seán O'Leary, Axel Röbel
DAFx2
2014 Online NON-negative Tensor Deconvolution for source detection in 3DTV audio
abstract
The following article describes research on source detection in multi channel (3DTV) audio streams. The problem is extremely complex due to the fact that multiple layers can be present in scenes (background music, ambience, commentator). In this work a new algorithm is developed that exploits the information from the different audio channels to detect, and possibly localize and separate independent audio sources. An algorithm based on online Non-negative Tensor Deconvolution is realized, to deal with sound sources with time dependent positions in the channel matrix. The evaluation is made on 3DTV 5.1 film soundtracks and on synthetic mixes of 3DTV 5.1 audio with target sounds from a sound effects database: a significant improvement of the detection performance is shown, compared with other decomposition techniques.
Yuki Mitsufuji, Marco Liuni, Alex Baker, Axel Röbel
ICASSP4
2014 On automatic voice casting for expressive speech: Speaker recognition vs. speech classification
abstract
This paper presents the first large-scale automatic voice casting system, and explores the adaptation of speaker recognition techniques to measure voice similarities. The proposed system is based on the representation of a voice by classes (e.g., age/gender, voice quality, emotion). First, a multi-label system is used to classify speech into classes. Then, the output probabilities for each class are concatenated to form a vector that represents the vocal signature of a speech recording. Finally, a similarity search is performed on the vocal signatures to determine the set of target actors that are the most similar to a speech recording of a source actor. In a subjective experiment conducted in the real-context of voice casting for video games, the multi-label system clearly outperforms standard speaker recognition systems. This indicates evidence that speech classes successfully capture the principal directions that are used in the perception of voice similarity.
Nicolas Obin, Axel Röbel, Grégoire Bachman
ICASSP2
2014 2D/3D AudioVisual content analysis & description
abstract
In this paper, we propose a way of using the Audio-Visual Description Profile (AVDP) of the MPEG-7 standard for 2D or stereo video and multichannel audio content description. Our aim is to provide means of using AVDP in such a way, that 3D video and audio content can be correctly and consistently described. Since AVDP semantics do not include ways for dealing with 3D audiovisual content, a new semantic framework within AVDP is proposed and examples of using AVDP to describe the results of analysis algorithms on stereo video and multichannel audio content are presented.
Ioannis Pitas, Konstantinos Papachristou, Nikos Nikolaidis 0001, Marco Liuni, Elie-Laurent Benaroya, Geoffroy Peeters, Axel Röbel, Antje Linnemann, Mohan Liu, Sebastian Gerke
MMSP7
2014 On the use of voice descriptors for glottal source shape parameter estimation
Stefan Huber 0003, Axel Röbel
Comput. Speech Lang.2
2013 Sound source separation based on non-negative tensor factorization incorporating spatial cue as prior knowledge
abstract
This paper concerns a new method of source separation that uses a spatial cue given by a user or from accompanying images to extract a target sound. The algorithm is based on non-negative tensor factorization (NTF), which decomposes multichannel spectrograms into three matrices. The components of one of the three matrices represent spatial information and are associated with the spatial cue, thus indicating which bins of the spectrogram should be given preference. When a spatial cue is available, this method has a great advantage over conventional PARAFAC-NTF in terms of both computational costs and separation quality, as measured by evaluation metrics such as SDR, SIR and SAR.
Yuki Mitsufuji, Axel Röbel
ICASSP2
2013 Syll-O-Matic: An adaptive time-frequency representation for the automatic segmentation of speech into syllables
abstract
This paper introduces novel paradigms for the segmentation of speech into syllables. The main idea of the proposed method is based on the use of a time-frequency representation of the speech signal, and the fusion of intensity and voicing measures through various frequency regions for the automatic selection of pertinent information for the segmentation. The time-frequency representation is used to exploit the speech characteristics depending on the frequency region. In this representation, intensity profiles are measured to provide information into various frequency regions, and voicing profiles are measured to determine the frequency regions that are pertinent for the segmentation. The proposed method outperforms conventional methods for the detection of syllable landmark and boundaries on the TIMIT database of American-English, and provides a promising paradigm for the segmentation of speech into syllables.
Nicolas Obin, Francois Lamare, Axel Röbel
ICASSP3
2013 Mixed source model and its adapted vocal tract filter estimate for voice transformation and synthesis
Gilles Degottex, Pierre Lanchantin, Axel Röbel, Xavier Rodet
Speech Commun.3
2013 Automatic Adaptation of the Time-Frequency Resolution for Sound Analysis and Re-Synthesis
abstract
We present an algorithm for sound analysis and re-synthesis with local automatic adaptation of time-frequency resolution. The reconstruction formula we propose is highly efficient, and gives a good approximation of the original signal from analyses with different time-varying resolutions within complementary frequency bands: this is a typical case where perfect reconstruction cannot in general be achieved with fast algorithms, which provides an error to be minimized. We provide a theoretical upper bound for the reconstruction error of our method, and an example of automatic adaptive analysis and re-synthesis of a music sound.
Marco Liuni, Axel Röbel, Ewa Matusiak, Marco Romito, Xavier Rodet
IEEE Trans. Speech Audio Process.2
2012 Analysis and modification of excitation source characteristics for singing voice synthesis
abstract
The present article investigates into the use of the LF glottal pulse model for singing synthesis and transformation. A recent estimator of the LF-glottal pulse shape parameter (rd) is used to analyze a small collection of professional singing examples and the results are discussed in the context of recent findings relating the rd shape parameter to other speech signal parameters (intensity and vibrato). We propose a rd shape parameter model for vibrato rendering and present an algorithm that allows modifying the glottal pulse shape parameter of a given speech signal and is used to enhance the vibrato generation in a speech to singing transformation system.
Axel Röbel, Stefan Huber 0003, Xavier Rodet, Gilles Degottex
ICASSP1
2012 Glottal source shape parameter estimation using phase minimization variants
abstract
The glottal shape parameter Rd provides a one-dimensional parameterisation of the Liljencrants-Fant (LF) model which describes the deterministic component of the glottal source. In this paper we first propose to estimate the Rd parameter by means of extending a state-of-the-art method based on the phase minimization criterion. Then we propose an adaption of the standard Rd parameter regression which enables us to coherently assess the normal and the upper Rd range. By evaluating the confusion matrices depicting the error surfaces of the involved different Rd parameter estimation methods and by objective measurement tests we verify the overall improvement of one new method compared to the state-of-the-art baseline approach.
Stefan Huber 0003, Axel Röbel, Gilles Degottex
INTERSPEECH2
2011 Function of Phase-Distortion for glottal model estimation
abstract
In voice analysis, the parameters estimation of a glottal model, an analytic description of the deterministic component of the glottal source, is a challenging question to assess voice quality in clinical use or to model voice production for speech transformation and synthesis using a priori constraints. In this paper, we first describe the Function of Phase-Distortion (FPD) which allows to characterize the shape of the periodic pulses of the glottal source independently of other features of the glottal source. Then, using the FPD, we de scribe two methods to estimate a shape parameter of the Liljencrants-Fant glottal model. By comparison with state of the art methods using Electro-Giotto-Graphic signals, we show that the one of these method outperform the compared methods.
Gilles Degottex, Axel Röbel, Xavier Rodet
ICASSP2
2011 Pitch transposition and breathiness modification using a glottal source model and its adapted vocal-tract filter
abstract
The transformation of the voiced segments of a speech recording has many applications such as expressivity synthesis or voice conversion. This paper addresses the pitch transposition and the modification of breathiness by means of an analytic description of the deterministic component of the voice source, a glottal model. Whereas this model is dedicated to voice production, most of the current methods can be applied to any pseudo-periodic signals. Using the described method, the synthesized voice is thus expected to better preserve some naturalness compared to a more generic method. Using preference tests, it is shown that this method is preferred for important pitch transposition (e.g. one octave) compared to two state of the art methods. Additionally, it is shown that the breathiness of two male utterances can be controlled.
Gilles Degottex, Axel Röbel, Xavier Rodet
ICASSP2
2011 Rényi information measures for spectral change detection
abstract
Change detection within an audio stream is an important task in several domains, such as classification and segmentation of a sound or of a music piece, as well as indexing of broadcast news or surveillance applications. In this paper we propose two novel methods for spectral change detection without any assumption about the input sound: they are both based on the evaluation of information measures applied to a time-frequency representation of the signal, and in particular to the spectrogram. The class of measures we consider, the Rényi entropies, are obtained by extending the Shannon entropy definition: a biasing of the spectrogram coefficients is realized through the dependence of such measures on a parameter, which allows refined results compared to those obtained with standard divergences. These methods provide a low computational cost and are well-suited as a support for higher level analysis, segmentation and classification algorithms.
Marco Liuni, Axel Röbel, Marco Romito, Xavier Rodet
ICASSP2
2011 Drum extraction from polyphonic music based on a spectro-temporal model of percussive sounds
abstract
In this paper, we present a new algorithm for removing drums from a polyphonic audio signal. The aim of this algorithm is to discard time/frequency bins which present a percussive magnitude evolution, according to a pre-defined parametric model. Special care is taken to reduce the irrelevant removal of frequency modulated signal such as the ones produced by the singing voice. Performance evaluation is carried out using objective measures commonly used by the community. Compared with four state-of-the-art algorithms, the proposed algorithm shows competitive performances at a low computational cost.
François Rigaud, Mathieu Lagrange, Axel Röbel, Geoffroy Peeters
ICASSP3
2011 Phase Minimization for Glottal Model Estimation
abstract
In glottal source analysis, the phase minimization criterion has already been proposed to detect excitation instants. As shown in this paper, this criterion can also be used to estimate the shape parameter of a glottal model (ex. Liljencrants-Fant model) and not only its time position. Additionally, we show that the shape parameter can be estimated independently of the glottal model position. The reliability of the proposed methods is evaluated with synthetic signals and compared to that of the IAIF and minimum/maximum-phase decomposition methods. The results of the methods are evaluated according to the influence of the fundamental frequency and noise. The estimation of a glottal model is useful for the separation of the glottal source and the vocal-tract filter and therefore can be applied in voice transformation, synthesis, and also in clinical context or for the study of the voice production.
Gilles Degottex, Axel Röbel, Xavier Rodet
IEEE Trans. Speech Audio Process.2
2010 Joint estimate of shape and time-synchronization of a glottal source model by phase flatness
abstract
A new method is proposed to jointly estimate the shape parameter of a glottal model and its time position in a voiced segment. We show that, the idea of phase flatness (or phase minimization) used in the most robust Glottal Closure Instant detection methods can be generalized to estimate the shape of the glottal model. In this paper we evaluate the proposed method using synthetic signals. The reliability related to fundamental frequency and noise is evaluated. The estimation of the glottal source is useful for voice analysis (ex. separation of glottal source and vocal-tract filter), voice transformation and synthesis.
Gilles Degottex, Axel Röbel, Xavier Rodet
ICASSP2
2010 Shape-invariant speech transformation with the phase vocoder
abstract
This paper proposes a new phase vocoder based method for shape invariant real-time modification of speech signals. The performance of the method with respect voiced and unvoiced signal components as well as the control of the voiced/unvoiced balance of the transformed speech signals will be discussed. The algorithm has been compared in perceptual tests with implementations of PSOLA, and HNM algorithms demonstrating a very satisfying performance. Due to the fact that the quality of transformed signals is remaining acceptable over a wide range of transformation parameters the algorithm is especially suited for real-time gender and age transformations.
Axel Röbel
INTERSPEECH1
2010 Dynamic Spectral Envelope Modeling for Timbre Analysis of Musical Instrument Sounds
abstract
We present a computational model of musical instrument sounds that focuses on capturing the dynamic behavior of the spectral envelope. A set of spectro-temporal envelopes belonging to different notes of each instrument are extracted by means of sinusoidal modeling and subsequent frequency interpolation, before being subjected to principal component analysis. The prototypical evolution of the envelopes in the obtained reduced-dimensional space is modeled as a nonstationary Gaussian Process. This results in a compact representation in the form of a set of prototype curves in feature space, or equivalently of prototype spectro-temporal envelopes in the time-frequency domain. Finally, the obtained models are successfully evaluated in the context of two music content analysis tasks: classification of instrument samples and detection of instruments in monaural polyphonic mixtures.
Juan José Burred, Axel Röbel, Thomas Sikora
IEEE Trans. Speech Audio Process.2
2010 Multiple Fundamental Frequency Estimation and Polyphony Inference of Polyphonic Music Signals
abstract
This paper presents a frame-based system for estimating multiple fundamental frequencies (F0s) of polyphonic music signals based on the short-time Fourier transform (STFT) representation. To estimate the number of sources along with their F0s, it is proposed to estimate the noise level beforehand and then jointly evaluate all the possible combinations among pre-selected F0 candidates. Given a set of F0 hypotheses, their hypothetical partial sequences are derived, taking into account where partial overlap may occur. A score function is used to select the plausible sets of F0 hypotheses. To infer the best combination, hypothetical sources are progressively combined and iteratively verified. A hypothetical source is considered valid if it either explains more energy than the noise, or improves significantly the envelope smoothness once the overlapping partials are treated. The proposed system has been submitted to Music Information Retrieval Evaluation eXchange (MIREX) 2007 and 2008 contests where the accuracy has been evaluated with respect to the number of sources inferred and the precision of the F0s estimated. The encouraging results demonstrate its competitive performance among the state-of-the-art methods.
Chunghsin Yeh, Axel Röbel, Xavier Rodet
IEEE Trans. Speech Audio Process.2
2009 Polyphonic musical instrument recognition based on a dynamic model of the spectral envelope
abstract
We propose a new method for detecting the musical instruments that are present in single-channel mixtures. Such a task is of interest for audio and multimedia content analysis and indexing applications. The approach is based on grouping sinusoidal trajectories according to common onsets, and comparing each group's overall amplitude evolution with a set of pre-trained probabilistic templates describing the temporal evolution of the spectral envelopes of a given set of instruments. Classification is based on either an Euclidean or a probabilistic definition of timbral similarity, both of which are compared with respect to detection accuracy.
Juan José Burred, Axel Röbel, Thomas Sikora
ICASSP2
2009 Applying improved spectral modeling for High Quality voice conversion
abstract
In this work, accurate spectral envelope estimation is applied to voice conversion in order to achieve high-quality timbre conversion. True-envelope based estimators allow model order selection leading to an adaptation of the spectral features to the characteristics of the speaker. Optimal residual signals can also be computed following a local adaptation of the model order in terms of the F0. A new perceptual criteria is proposed to measure the impact of the spectral conversion error. The proposed envelope models show improved spectral conversion performance as well as increased converted-speech quality when compared to linear prediction.
Fernando Villavicencio, Axel Röbel, Xavier Rodet
ICASSP2
2009 The expected amplitude of overlapping partials of harmonic sounds
abstract
In analyzing polyphonic signals, the handling of overlapping partials is one important problem. The assumptions usually made for partial overlaps are the additivity of the linear spectrum or that of the power spectrum. In this study, the expected amplitude of two overlapping partials is derived based on the assumption that the partials overlap at the same frequency and the phase is uniformly distributed. An overlap chain rule algorithm is proposed to estimate the amplitude for the case that more than two partials overlap. The proposed algorithm has demonstrated its better accuracy over the usual two model assumptions.
Chunghsin Yeh, Axel Röbel
ICASSP2
2008 Extending efficient spectral envelope modeling to Mel-frequency based representation
abstract
In this work we consider the problem of spectral envelope estimation using spectra with perceptually warped frequency axis. The goal of this work is the reduction of the order of the spectral envelope model which will facilitate the use of these envelopes for training of voice conversion systems. We adapt the true-envelope estimator to Mel-frequency representations and adapt a recently proposed cepstral model order selection criterion taking into account the distortion of the frequency axis. We evaluate the modified order selection procedure using a perceptual framework for the evaluation of envelope estimation errors. The experimental evaluation carried out with real speech confirms our modifications. The results demonstrate that the Mel frequency based true envelope estimator achieves superior envelope estimation with significantly reduced model order.
Fernando Villavicencio, Axel Röbel, Xavier Rodet
ICASSP2
2007 All-Pole Spectral Envelope Modelling with Order Selection for Harmonic Signals
abstract
We present a study into all-pole spectral envelope estimation for the case of harmonic signals. We address the problem of the selection of the model order and propose to make use of the fact that the spectral envelope is sampled by means of the harmonic structure to derive a reasonable choice for an appropriate model order. The experimental investigation uses synthetic ARMA featured signals with varying fundamental frequency and differing model structure to evaluate the performance of the selected all-pole models. The experimental results confirm the relation between optimal model order and the fundamental frequency.
Fernando Villavicencio, Axel Röbel, Xavier Rodet
ICASSP (1)2
2007 Speech to chant transformation with the phase vocoder
Axel Röbel, Joshua Fineberg
INTERSPEECH1
2007 On cepstral and all-pole based spectral envelope modeling with unknown model order
Axel Röbel, Fernando Villavicencio, Xavier Rodet
Pattern Recognit. Lett.1
2006 Improving Lpc Spectral Envelope Extraction Of Voiced Speech By True-Envelope Estimation
abstract
n this work we address the problem of all pole spectral envelope estimation for speech signals. The currently widely used all pole spectral envelope model suffers from well-known systematic errors and more severely from model order mismatch. We will propose a procedure to first establish a band limited interpolation of the observed spectrum using a recently rediscovered true envelope estimator and then using the band limited envelope to derive an all pole envelope model named TE-LPC . The band-limited envelope that is used to derive the all pole envelope model reduces the problem of the unknown all pole model order. For the experimental investigation we propose a new perceptually motivated residual spectral peak flatness measure. The experimental results demonstrate that the proposed method significantly increases the spectral flatness for the perceptually especially important low order harmonics of voiced utterances
Fernando Villavicencio, Axel Röbel, Xavier Rodet
ICASSP (1)2
2006 Adaptive additive modeling with continuous parameter trajectories
abstract
This paper investigates the estimation of time varying amplitude and phase trajectories of sinusoidal signal components. The new algorithm adaptively optimizes the parameters of a smoothly connected piecewise polynomial trajectory model. A mathematical analysis is presented that relates the user-selected meta parameters of the trajectory model (polynomial order, segment size, and smoothness at the junctions) to the analysis properties of the adaptive algorithm. It reveals new insights into the relationships between the meta parameters and the resulting time/frequency resolution of the estimate. Moreover, it is shown that for efficient optimization, the phase trajectory needs to be represented in a specific form. A new approach to address the bias/variance tradeoff of the polynomial phase trajectory model by means of regularization is presented and a complete adaptive analysis/synthesis system for sinusoidal sound components is proposed. The adaptive analysis system is investigated by means of simple tracking experiments to demonstrate the effect of the smoothness constraints and compare the results with a standard short-time Fourier transformation (STFT) base frequency estimation technique and known Cramer-Rao bounds. The potential of the adaptive strategy for the modeling of sinusoidal transients is discussed and it is shown that it achieves similar transient quality as a previously proposed method, however, with considerably lower model error. Two examples for modeling real-world signals are discussed
Axel Röbel
IEEE Trans. Speech Audio Process.1
2005 Multiple fundamental frequency estimation of polyphonic music signals
abstract
The article is concerned with the estimation of fundamental frequencies, or F0s, in polyphonic music. We propose a new method for jointly evaluating multiple F0 hypotheses based on three physical principles, harmonicity, spectral smoothness and synchronous amplitude evolution, within a single source. Based on the generative quasiharmonic model, a set of hypothetical partial sequences is derived and an optimal assignment of the observed peaks to the hypothetical sources and noise is performed. The hypothetical partial sequences are then evaluated by a score function which formulates the guiding principles in a mathematical manner. The algorithm has been tested on a large collection of artificially mixed polyphonic samples and the results show the competitive performance of the proposed method.
Chunghsin Yeh, Axel Röbel, Xavier Rodet
ICASSP (3)2
1996 Neural Network Modeling of Speech and Music Signals
Axel Röbel
NIPS1
1994 Dynamic pattern selection for faster learning and controlled generalization of neural networks
Axel Röbel
ESANN1