Nicolas Obin

dblp:57/7823 · DBLP profile ↗
← Back
33ranked-venue papers
14as first author
9since 2021 · last 2025
0000-0002-5236-5306ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 28 · 12 first-author · 8 since 2021Artificial intelligence and machine learning · 23 · 10 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 TranSTYLer: Multimodal behavioural style transfer for facial and body gestures generation
abstract
This paper addresses the challenge of transferring the behaviour expressivity style of a virtual agent to another one while preserving behaviour shape as they carry communicative meaning. Behaviour expressivity style is viewed here as the qualitative properties of behaviours. We propose TranSTYLer , a multimodal transformer-based model that synthesises the multimodal behaviours of a source speaker with the style of a target speaker. We assume that behaviour expressivity style is encoded across various modalities of communication, including text, speech, body gestures, and facial expressions. The model employs a style-content disentanglement schema to ensure that the transferred style does not interfere with the meaning conveyed by the source’s behaviours. Our approach eliminates the need for style labels and allows the generalisation of styles not seen during the training phase. We train our model on the PATS corpus , which we extended to include dialogue acts and 2D facial landmarks. Objective and subjective evaluations show that our model outperforms state-of-the-art models in style transfer for both seen and unseen styles during training. To tackle the issues of style and content leakage that may arise, we propose a methodology to assess the degree to which behaviour and gestures associated with the target style are successfully transferred while ensuring the preservation of the ones related to the source content.
Mireille Fares, Catherine Pelachaud, Nicolas Obin
Speech Commun.3
2024 Auditory Cortex-Inspired Spectral Attention Modulation for Binaural Sound Localization in HRTF Mismatch
abstract
In applications like noise cancellation and virtual reality, precise sound source localization is crucial. Existing data-driven binaural systems offer high performance in adverse conditions such as noise and reverberation but face limitations with real-time operation and performance degradation in HRTF mismatch scenarios. Our work introduces a compact Vision Transformer tailored to address these issues, with a primary focus on horizontal speech localization. Inspired by the auditory cortex, our model uniquely incorporates spectral attention mechanisms using encoded speech representations. This architecture enhances generalization on the azimuth plane under mismatched HRTFs. Our empirical results show a marked improvement over conventional DNN, CNN-based and Transformer-based models, both in noisy and noise-free environments. Significantly, the proposed model maintains high accuracy in localizing adjacent azimuths, ideal for real-world applications.
Waradon Phokhinanan, Nicolas Obin, Sylvain Argentieri
ICASSP2
2024 BWSNET: Automatic Perceptual Assessment of Audio Signals
abstract
This paper introduces BWSNet, a model that can be trained from raw human judgements obtained through a Best-Worst scaling (BWS) experiment. It maps sound samples into an embedded space that represents the perception of a studied attribute. To this end, we propose a set of cost functions and constraints, interpreting trial-wise ordinal relations as distance comparisons in a metric learning task. We tested our proposal on data from two BWS studies investigating the perception of speech social attitudes and timbral qualities. For both datasets, our results show that the structure of the latent space is faithful to human judgements.
Clément Le Moine Veillon, Victor Rosi, Pablo Arias Sarah, Léane Salais, Nicolas Obin
ICASSP5
2024 Small-E: Small Language Model with Linear Attention for Efficient Speech Synthesis
abstract
Recent advancements in text-to-speech (TTS) powered by language models have showcased remarkable capabilities in achieving naturalness and zero-shot voice cloning. Notably, the decoder-only transformer is the prominent architecture in this domain. However, transformers face challenges stemming from their quadratic complexity in sequence length, impeding training on lengthy sequences and resource-constrained hardware. Moreover they lack specific inductive bias with regards to the monotonic nature of TTS alignments. In response, we propose to replace transformers with emerging recurrent architectures and introduce specialized cross-attention mechanisms for reducing repeating and skipping issues. Consequently our architecture can be efficiently trained on long samples and achieve state-of-the-art zero-shot voice cloning against baselines of comparable size. Our implementation and demos are available at https://github.com/theodorblackbird/lina-speech.
Théodor Lemerle, Nicolas Obin, Axel Röbel
INTERSPEECH2
2024 2D or not 2D: How Does the Dimensionality of Gesture Representation Affect 3D Co-Speech Gesture Generation?
abstract
Co-speech gestures are fundamental for communication. The advent of recent deep learning techniques has facilitated the creation of lifelike, synchronous co-speech gestures for Embodied Conversational Agents. "In-the-wild" datasets, aggregating video content from platforms like YouTube via human pose detection technologies, provide a feasible solution by offering 2D skeletal sequences aligned with speech. Concurrent developments in lifting models enable the conversion of these 2D sequences into 3D gesture databases. However, it is important to note that the 3D poses estimated from the 2D extracted poses are, in essence, approximations of the ground-truth, which remains in the 2D domain. This distinction raises questions about the impact of gesture representation dimensionality on the quality of generated motions. Our study examines the effect of using either 2D or 3D joint coordinates as training data on the performance of speech-to-gesture deep generative models.
Teo Guichoux, Laure Soulier, Nicolas Obin, Catherine Pelachaud
IVA3
2023 Zero-Shot Style Transfer for Multimodal Data-Driven Gesture Synthesis
abstract
We propose a multimodal speech driven approach to generate 2D upper-body gestures for virtual agents, in the communicative style of different speakers, seen or unseen by our model during training. Upper-body gestures of a source speaker are generated based on the content of his/her multimodal data - speech acoustics and text semantics. The synthesized source speaker's gestures are conditioned on the multimodal style representation of the target speaker. Our approach is zero-shot, and can generalize the style transfer to new unseen speakers, without any additional training. An objective evaluation is conducted to validate our approach.
Mireille Fares, Catherine Pelachaud, Nicolas Obin
FG3
2023 Binaural Sound Localization in Noisy Environments Using Frequency-Based Audio Vision Transformer (FAViT)
abstract
International audience
Waradon Phokhinanan, Nicolas Obin, Sylvain Argentieri
INTERSPEECH2
2022 Production Strategies of Vocal Attitudes
abstract
Humans have an impressive ability to communicate precise social intentions and desires with their voice - through vocal attitudes. Previous studies have shown how isolated acoustic features such as pitch can convey social attitudes, but have mostly worked with single attitudes and have not controlled for inter-speaker variability. Thus, the vocal behaviours used to produce social attitudes remain mostly unknown. That is the aim of the current study, to uncover the anatomic production strategies that speakers use to communicate vocal attitudes. To do this, we analysed recordings from N=20 French speakers producing dominant, friendly, seductive and distant speech. For each of these attitudes, we investigated their vocal fold behaviour, vocal tract actuation and phonetic speech structure, with the support of deep alignment methods, and compared them with group statistics. We notably produced high-level representations of speakers' articulation (e.g. Vowel Space Density) and speech rhythm. Our results reveal speakers' prototypical strategies to produce vocal attitudes, and highlight how vocal behaviours can communicate social signals. We expect these results to provide an objective validation method for deep voice attitude conversions.
Léane Salais, Pablo Arias 0003, Clément Le Moine, Victor Rosi, Yann Teytaut, Nicolas Obin, Axel Röbel
INTERSPEECH6
2021 Speaker Attentive Speech Emotion Recognition
abstract
International audience
Clément Le Moine, Nicolas Obin, Axel Röbel
Interspeech2
2019 Sequence-to-sequence Modelling of F0 for Speech Emotion Conversion
abstract
Voice interfaces are becoming wildly popular and driving demand for more advanced speech synthesis and voice transformation systems. Current text-to-speech methods produce realistic sounding voices, but they lack the emotional expressivity that listeners expect, given the context of the interaction and the phrase being spoken. Emotional voice conversion is a research domain concerned with generating expressive speech from neutral synthesised speech or natural human voice. This research investigated the effectiveness of using a sequence-to-sequence (seq2seq) encoder-decoder based model to transform the intonation of a human voice from neutral to expressive speech, with some preliminary introduction of linguistic conditioning. A subjective experiment conducted on the task of speech emotion recognition by listeners successfully demonstrated the effectiveness of the proposed sequence-to-sequence models to produce convincing voice emotion transformations. In particular, conditioning the model on the position of the syllable in the phrase significantly improved recognition rates.
Carl Robinson, Nicolas Obin, Axel Röbel
ICASSP2
2018 Binaural Localization of Multiple Sound Sources by Non-Negative Tensor Factorization
abstract
This paper presents non-negative factorization of audio signals for the binaural localization of multiple sound sources within realistic and unknown sound environments. Non-negative tensor factorization (NTF) provides a sparse representation of multichannel audio signals in time, frequency, and space that can be exploited in computational audio scene analysis and robot audition for the separation and localization of sound sources. In the proposed formulation, each sound source is represented by means of spectral dictionaries, temporal activation, and its distribution within each channel (here, left and right ears). This distribution, being dependent on the frequency, can be interpreted as an explicit estimation of the Head-Related Transfer Function (HRTF) of a binaural head which can then be converted into the estimated sound source position. Moreover, the semisupervised formulation of the non-negative factorization allows us to integrate prior knowledge about some sound sources of interest whose dictionaries can be learned in advance, whereas the remaining sources are considered as background sound, which remains unknown and is estimated on the fly. The proposed NTF-based sound source localization is applied here to binaural sound source localization of multiple speakers within realistic sound environments.
Elie-Laurent Benaroya, Nicolas Obin, Marco Liuni, Axel Röbel, Wilson Raumel, Sylvain Argentieri
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 A source/filter model with adaptive constraints for NMF-based speech separation
abstract
This paper introduces a constrained source/filter model for semi-supervised speech separation based on non-negative matrix factorization (NMF). The objective is to inform NMF with prior knowledge about speech, providing a physically meaningful speech separation. To do so, a source/filter model (indicated as Instantaneous Mixture Model or IMM) is integrated in the NMF. Furthermore, constraints are added to the IMM-NMF, in order to control the NMF behaviour during separation, and to enforce its physical meaning. In particular, a speech specific constraint - based on the source/filter coherence of speech - and a method for the automatic adaptation of constraints' weights during separation are presented. Also, the proposed source/filter model is semi-supervised: during training, one filter basis is estimated for each phoneme of a speaker; during separation, the estimated filter bases are then used in the constrained source/filter model. An experimental evaluation for speech separation was conducted on the TIMIT speakers database mixed with various environmental background noises from the QUT-NOISE database. This evaluation showed that the use of adaptive constraints increases the performance of the source/filter model for speaker-dependent speech separation, and compares favorably to fully-supervised speech separation.
Damien Bouvier, Nicolas Obin, Marco Liuni, Axel Röbel
ICASSP2
2016 Similarity Search of Acted Voices for Automatic Voice Casting
abstract
This paper presents a large-scale similarity search of professionally acted voices for computer-aided voice casting. The proposed voice casting system explores Gaussian mixture model-based acoustic models and multilabel recognition of perceived paralinguistic content (speaker states and speaker traits, e.g., age/gender, voice quality, emotion) for the voice casting of professionally acted voices. First, acoustic models (universal background model, super-vector, i-vector) are constructed to model the acoustic space of voices, from which the similarity between voices can be measured directly in the acoustic space. Second, multiple binary classification of speaker traits and states is added to the acoustic models in order to represent the vocal signature of a voice, which is then used to measure the similarity between voices in the paralinguistic space. Finally, a similarity search is processed in order to determine the set of target actors that are the most similar to the voice of a source actor. In a subjective experiment conducted in the real-context of cross-language voice casting, the multilabel scoring system significantly outperforms the acoustic scoring system. This constitutes a proof of concept for the role of perceived para-linguistic categories in the perception of voice similarity.
Nicolas Obin, Axel Röbel
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 The role of glottal source parameters for high-quality transformation of perceptual age
abstract
The intuitive control of voice transformation (e.g., age/sex, emotions) is useful to extend the expressive repertoire of a voice. This paper explores the role of glottal source parameters for the control of voice transformation. First, the SVLN speech synthesizer (Separation of the Vocal-tract with the Liljencrants-fant model plus Noise) is used to represent the glottal source parameters (and thus, voice quality) during speech analysis and synthesis. Then, a simple statistical method is presented to control speech parameters during voice transformation: a GMM is used to model the speech parameters of a voice, and regressions are then used to adapt the GMMs statistics (mean and variance) to a control parameter (e.g., age/sex, emotions). A subjective experiment conducted on the control of perceptual age proves the importance of the glottal source parameters for the control of voice transformation, and shows the efficiency of the statistical model to control voice parameters while preserving a high-quality of the voice transformation.
Xavier Favory, Nicolas Obin, Gilles Degottex, Axel Röbel
ICASSP2
2015 Real-time audio-to-score alignment of singing voice based on melody and lyric information
abstract
International audience
Rong Gong, Philippe Cuvillier, Nicolas Obin, Arshia Cont
INTERSPEECH3
2015 Symbolic Modeling of Prosody: From Linguistics to Statistics
abstract
The assignment of prosodic events (accent and phrasing) from the text is crucial in text-to-speech synthesis systems. This paper addresses the combination of linguistic and metric constraints for the assignment of prosodic events in text-to-speech synthesis. First, a linguistic processing chain is used to provide a rich linguistic description of a text. Then, a novel statistical representation based on a hierarchical HMM (HHMM) is used to model the prosodic structure of a text: the root layer represents the text, each intermediate layer a sequence of intermediate phrases, the pre-terminal layer the sequence of accents, and the terminal layer the sequence of linguistic contexts. For each intermediate layer, a segmental HMM and information fusion are used to fuse the linguistic and metric constraints for the segmentation of a text into phrases. A set of experiments conducted on multi-speaker databases with various speaking styles reports that: the rich linguistic representation improves drastically the assignment of prosodic events, and the fusion of linguistic and metric constraints significantly improves over standard methods for the segmentation of a text into phrases. These constitute substantial advances that can be further used to model the speech prosody of a speaker, a speaking style, and emotions for text-to-speech synthesis.
Nicolas Obin, Pierre Lanchantin
IEEE ACM Trans. Audio Speech Lang. Process.1
2014 On automatic voice casting for expressive speech: Speaker recognition vs. speech classification
abstract
This paper presents the first large-scale automatic voice casting system, and explores the adaptation of speaker recognition techniques to measure voice similarities. The proposed system is based on the representation of a voice by classes (e.g., age/gender, voice quality, emotion). First, a multi-label system is used to classify speech into classes. Then, the output probabilities for each class are concatenated to form a vector that represents the vocal signature of a speech recording. Finally, a similarity search is performed on the vocal signatures to determine the set of target actors that are the most similar to a speech recording of a source actor. In a subjective experiment conducted in the real-context of voice casting for video games, the multi-label system clearly outperforms standard speaker recognition systems. This indicates evidence that speech classes successfully capture the principal directions that are used in the perception of voice similarity.
Nicolas Obin, Axel Röbel, Grégoire Bachman
ICASSP1
2014 Phase distortion statistics as a representation of the glottal source: application to the classification of voice qualities
abstract
The representation of the glottal source is of paramount importance for describing para-linguistic information carried through the voice quality (e.g., emotions, mood, attitude). However, some existing representations of the glottal source are based on analytical glottal models, which assume strong a priori constraints on the shape of the glottal pulses. Thus, these representations are restricted to limited number of voices. Recent progresses in the estimation of the glottal models revealed that the Phase Distortion (PD) of the signal carries most of the information about the glottal pulses. This paper introduces a flexible representation of the glottal source - based on the short-term modelling of the phase distortion. This representation is not constrained by a specific analytical model, and thus can be used to describe a larger variety of expressive voices. We address the efficiency of this representation for the recognition of various voice qualities, with comparison to MFCC and standard glottal source representations.
Gilles Degottex, Nicolas Obin
INTERSPEECH2
2014 Rhapsodie: a Prosodic-Syntactic Treebank for Spoken French
Anne Lacheret, Sylvain Kahane, Julie Beliao, Anne Dister, Kim Gerdes, Jean-Philippe Goldman, Nicolas Obin, Paola Pietrandrea, Atanas Tchobanov
LREC7
2013 Syll-O-Matic: An adaptive time-frequency representation for the automatic segmentation of speech into syllables
abstract
This paper introduces novel paradigms for the segmentation of speech into syllables. The main idea of the proposed method is based on the use of a time-frequency representation of the speech signal, and the fusion of intensity and voicing measures through various frequency regions for the automatic selection of pertinent information for the segmentation. The time-frequency representation is used to exploit the speech characteristics depending on the frequency region. In this representation, intensity profiles are measured to provide information into various frequency regions, and voicing profiles are measured to determine the frequency regions that are pertinent for the segmentation. The proposed method outperforms conventional methods for the detection of syllable landmark and boundaries on the TIMIT database of American-English, and provides a promising paradigm for the segmentation of speech into syllables.
Nicolas Obin, Francois Lamare, Axel Röbel
ICASSP1
2012 Accentual Transfer from Swiss-German to French. A Study of "Français Fédéral"
abstract
cote interne IRCAM: Avanzi12c
Mathieu Avanzi, Pauline Dubosson, Sandra Schwab, Nicolas Obin
INTERSPEECH4
2012 Towards Glottal Source Controllability in Expressive Speech Synthesis
abstract
In order to obtain more human like sounding humanmachine interfaces we must first be able to give them expressive capabilities in the way of emotional and stylistic features so as to closely adequate them to the intended task. If we want to replicate those features it is not enough to merely replicate the prosodic information of fundamental frequency and speaking rhythm. The proposed additional layer is the modification of the glottal model, for which we make use of the GlottHMM parameters. This paper analyzes the viability of such an approach by verifying that the expressive nuances are captured by the aforementioned features, obtaining 95% recognition rates on styled speaking and 82% on emotional speech. Then we evaluate the effect of speaker bias and recording environment on the source modeling in order to quantify possible problems when analyzing multi-speaker databases. Finally we propose a speaking styles separation for Spanish based on prosodic features and check its perceptual significance.
Jaime Lorenzo-Trueba, Roberto Barra-Chicote, Tuomo Raitio, Nicolas Obin, Paavo Alku, Junichi Yamagishi, Juan Manuel Montero-Martínez
INTERSPEECH4
2012 Cries and Whispers - Classification of Vocal Effort in Expressive Speech
abstract
The expansion of the video games industry raises innovative and challenging issues for speech technologies, e.g. the development of automatic content-based speech processing and speech recognition systems in the context of video games post-production and voice casting. This paper presents a large-scale study on the classification of vocal effort in expressive speech for video games. Changes in vocal effort conduct to substantial modifications in the configuration of voice production mechanisms. In particular, registers of vocal effort affect especially voice quality which reflects qualitative modifications of the source excitation characteristics. This study introduces robust source characteristics to measure various types of voice quality (e.g., breathy, creaky, tense) for the classification of vocal effort into whispered, normal, and shouted speech. The system is evaluated in the real scenario of video games production with the complete speech recordings of a massive role-playing video game. The proposed features significantly improve the classification from 81.1% to 87% over conventional MFCCs. These advancements confirm the role of the source and voice quality for the description of changes in vocal effort.
Nicolas Obin
INTERSPEECH1
2012 On the generalization of Shannon entropy for speech recognition
abstract
This paper introduces an entropy-based spectral representation as a measure of the degree of noisiness in audio signals, complementary to the standard MFCCs for audio and speech recognition. The proposed representation is based on the Rényi entropy, which is a generalization of the Shannon entropy. In audio signal representation, Rényi entropy presents the advantage of focusing either on the harmonic content (prominent amplitude within a distribution) or on the noise content (equal distribution of amplitudes). The proposed representation outperforms all other noisiness measures - including Shannon and Wiener entropies - in a large-scale classification of vocal effort (whispered-soft/normal/loud-shouted) in the real scenario of multi-language massive role-playing video games. The improvement is around 10% in relative error reduction, and is particularly significant for the recognition of noisy speech - i.e., whispery/breathy speech. This confirms the role of noisiness for speech recognition, and will further be extended to the classification of voice quality for the design of an automatic voice casting system in video games.
Nicolas Obin, Marco Liuni
SLT1
2011 Toward a Continuous Modeling of French Prosodic Structure: Using Acoustic Features to Predict Prominence Location and Prominence Degree
abstract
International audience
Mathieu Avanzi, Nicolas Obin, Anne Lacheret, Bernard Victorri
INTERSPEECH2
2011 Reformulating Prosodic Break Model into Segmental HMMs and Information Fusion
abstract
In this paper, a method for prosodic break modelling based on segmental-HMMs and Dempster-Shafer fusion for speech synthesis is presented, and the relative importance of linguistic and metric constraints in prosodic break modelling is assessed 1 . A context-dependent segmental-HMM is used to explicitly model the linguistic and the metric constraints. Dempster-Shafer fusion is used to balance the relative importance of the linguistic and the metric constraints into the segmental-HMM. A linguistic processing chain based on surface and deep syntactic parsing is additionally used to extract linguistic informations of different nature. An objective evaluation proved evidence that the optimal combination of the linguistic and the metric constraints significantly outperforms both the conventional HMM (linguistic information only) and segmental-HMM (equal balance of linguistic and metric constraints), and confirmed that the linguistic constraint is prior to the metric. Index Terms: speech prosody, prosodic break, segmentalHMM, Dempster-Shafer fusion.
Nicolas Obin, Pierre Lanchantin, Anne Lacheret, Xavier Rodet
INTERSPEECH1
2011 Discrete/Continuous Modelling of Speaking Style in HMM-Based Speech Synthesis: Design and Evaluation
abstract
This paper assesses the ability of a HMM-based speech synthesis systems to model the speech characteristics of various speaking styles 1 .A discrete/continuous HMM is presented to model the symbolic and acoustic speech characteristics of a speaking style.The proposed model is used to model the average characteristics of a speaking style that is shared among various speakers, depending on specific situations of speech communication.The evaluation consists of an identification experiment of 4 speaking styles based on delexicalized speech, and compared to a similar experiment on natural speech.The comparison is discussed and reveals that discrete/continuous HMM consistently models the speech characteristics of a speaking style.
Nicolas Obin, Pierre Lanchantin, Anne Lacheret, Xavier Rodet
INTERSPEECH1
2011 Stylization and Trajectory Modelling of Short and Long Term Speech Prosody Variations
abstract
In this paper, a unified trajectory model based on the stylization and the modelling of f0 variations simultaneously over various temporal domains is proposed 1 . The syllable is used as the minimal temporal domain for the description of speech prosody, and short-term and long-termf0 variations are stylized and modelled simultaneously over various temporal domains. During the training, a context-dependent model is estimated according to the joint stylized f0 contours over the syllable and a set of long-term temporal domains. During the synthesis, f0 variations are determined using the long-term variations as trajectory constraints. In a subjective evaluation in speech synthesis, the stylization and trajectory modelling of short and long term speech prosody variations is shown to consistently model speech prosody and to outperform the conventional short-term modelling.
Nicolas Obin, Anne Lacheret, Xavier Rodet
INTERSPEECH1
2010 Expectations for discourse genre identification: a prosodic study
abstract
Speech can be divided into discourse genres based on the contextual environment it occurs in (e.g.political speech, sport commentary speech, etc.).The present study investigated whether listeners can distinguish between speech from different discourse genres on the basis of acoustic prosodic cues only 1 .In a perception experiment with delexicalized speech 70 listeners with varying experience in French (native speakers, nonnative speakers, and non-speakers) were asked to identify four different types of discourse genres (church service, political, journal, and sport commentary).Results revealed a fair identification ability with a significant increase in performance with increasing experience in French.Identification confusion was used to cluster discourse genres according to their perceptual similarity.
Nicolas Obin, Volker Dellwo, Anne Lacheret, Xavier Rodet
INTERSPEECH1
2010 HMM-based prosodic structure model using rich linguistic context
abstract
This paper presents a study on the use of deep syntactical features to improve prosody modeling. A French linguistic processing chain based on linguistic preprocessing, morpho- syntactical labeling, and deep syntactical parsing is used in order to extract syntactical features from an input text. These features are used to define more or less high-level syntactical feature sets. Such feature sets are compared on the basis of a HMM-based prosodic structure model. High-level syntactical features are shown to significantly improve the performance of the model (up to 21% error reduction combined with 19% BIC reduction).
Nicolas Obin, Xavier Rodet, Anne Lacheret
INTERSPEECH1
2009 A multi-level context-dependent prosodic model applied to durational modeling
abstract
International audience
Nicolas Obin, Xavier Rodet, Anne Lacheret
INTERSPEECH1
2008 French prominence: A probabilistic framework
abstract
Identification of prosodic phenomena is of first importance in prosodic analysis and modeling. In this paper, we introduce a new method for automatic prosodic phenomena labelling. The authors set their approach of prosodic phenomena in the framework of prominence. The proposed method for automatic prominence labelling is based on well-known machine learning techniques in a three step procedure: (i) a feature extraction step in which we propose a framework for systematic and multi-level speech acoustic feature extraction, (ii) a feature selection step for identifying the more relevant prominence acoustic correlates, and (iii) a modelling step in which a gaussian mixture model is used for predicting prominence. This model shows robust performance on read speech (84%).
Nicolas Obin, Xavier Rodet, Anne Lacheret
ICASSP1
2008 A method for automatic and dynamic estimation of discourse genre typology with prosodic features
abstract
ISBN: 978-1-61567-378-0 Special Session: Prosody of Spontaneous Speech I, Wednesday 24th September 2008.
Nicolas Obin, Anne Lacheret, Christophe Veaux, Xavier Rodet, Anne-Catherine Simon
INTERSPEECH1