VLDB 2026 Research / reviewers in the wild / expert
Yves Laprie
dblp:62/2792
· DBLP profile ↗
63ranked-venue papers
14as first author
5since 2021 · last 2025
0000-0002-2379-6481ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 59 · 14 first-author · 5 since 2021Artificial intelligence and machine learning · 49 · 10 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Complete Reconstruction of the Tongue Contour Through Acoustic to Articulatory Inversion Using Real-Time MRI DataabstractAcoustic articulatory inversion is a major processing challenge, with a wide range of applications from speech synthesis to feedback systems for language learning and rehabilitation.In recent years, deep learning methods have been applied to the inversion of less than a dozen geometrical positions corresponding to sensors glued to easily accessible articulators. It is therefore impossible to know the shape of the whole tongue from root to tip. In this work, we use high-quality real-time MRI data to track the contour of the tongue. The data used to drive the inversion are therefore the unstructured speech signal and the tongue contours. Several architectures relying on a Bi-LSTM including or not an autoencoder to reduce the dimensionality of the latent space, using or not the phonetic segmentation have been explored. The results show that the tongue contour can be recovered with a median accuracy of 2.21 mm (or 1.37 pixels) taking a context of 1 MFCC frame (static, delta and double-delta cepstral features). Sofiane Azzouz, Pierre-André Vuissoz, Yves Laprie |
ICASSP | 3 |
| 2025 | Reconstruction of the Complete Vocal Tract Contour Through Acoustic to Articulatory Inversion Using Real-Time MRI DataabstractAcoustic to articulatory inversion has often been limited to a small part of the vocal tract because the data are generally EMA (ElectroMagnetic Articulography) data requiring sensors to be glued to easily accessible articulators. The presented acoustic to articulation model focuses on the inversion of the entire vocal tract from the glottis, the complete tongue, the velum, to the lips. It relies on a realtime dynamic MRI database of more than 3 hours of speech. The data are the denoised speech signal and the automatically segmented articulator contours. Several bidirectional LSTM-based approaches have been used, either inverting each articulator individually or inverting all articulators simultaneously. To our knowledge, this is the first complete inversion of the vocal tract. The average RMSE precision on the test set is 1.65 mm to be compared with the pixel size which is 1.62 mm. Sofiane Azzouz, Pierre-André Vuissoz, Yves Laprie |
INTERSPEECH | 3 |
| 2022 | Autoencoder-Based Tongue Shape Estimation During Continuous SpeechabstractVocal tract shape estimation is a necessary step for articulatory speech synthesis.However, the literature on the topic is scarce, and most current methods lack adequacy to many physical constraints related to speech production.This study proposes an alternative approach to the task to solve specific issues faced in the previous work, especially those related to critical articulators.We present an autoencoder-based method for tongue shape estimation during continuous speech.An autoencoder is trained to learn the data's encoding and serves as an auxiliary network for the principal one, which maps phonemes to the shapes.Instead of predicting the exact points in the target curve, the neural network learns how to predict the curve's main components, i.e., the autoencoder's representation.We show how this approach allows imposing critical articulators' constraints, controlling the tongue shape through the latent space, and generating a smooth output without relying on any postprocessing method. Vinicius Ribeiro, Yves Laprie |
INTERSPEECH | 2 |
| 2022 | Automatic generation of the complete vocal tract shape from the sequence of phonemes to be articulated
Vinicius Ribeiro, Karyna Isaieva, Justine Leclere, Pierre-André Vuissoz, Yves Laprie |
Speech Commun. | 5 |
| 2021 | Towards the Prediction of the Vocal Tract Shape from the Sequence of Phonemes to be ArticulatedabstractInternational audience Vinicius Ribeiro, Karyna Isaieva, Justine Leclere, Pierre-André Vuissoz, Yves Laprie |
Interspeech | 5 |
| 2020 | Using Silence MR Image to Synthesise Dynamic MRI Vocal Tract Data of CVabstractInternational audience Ioannis K. Douros, Ajinkya Kulkarni, Chrysanthi Dourou, Jacques Felblinger, Karyna Isaieva, Pierre-André Vuissoz, Yves Laprie |
INTERSPEECH | 8 |
| 2019 | A Multimodal Real-Time MRI Articulatory Corpus of French for Speech ResearchabstractInternational audience Ioannis K. Douros, Jacques Felblinger, Jens Frahm, Karyna Isaieva, Arun A. Joseph, Yves Laprie, Freddy Odille, Anastasiia Tsukanova, Dirk Voit, Pierre-André Vuissoz |
INTERSPEECH | 6 |
| 2019 | Towards a Method of Dynamic Vocal Tract Shapes Generation by Combining Static 3D and Dynamic 2D MRI Speech DataabstractInternational audience Ioannis K. Douros, Anastasiia Tsukanova, Karyna Isaieva, Pierre-André Vuissoz, Yves Laprie |
INTERSPEECH | 5 |
| 2017 | Towards confidence measures on fundamental frequency estimationsabstractThe fundamental frequency is one of the prosodic parameters, and many algorithms have been developed for estimating the fundamental frequency of speech signals. Most of them provide good results on good quality speech signals, but their performance degrades when dealing with noisy signals. Moreover, although some provide a probability for the voicing decision, none of them indicate how reliable the estimated fundamental frequency is. In this paper, we investigate the computation of a confidence (or reliability) measure on the estimated fundamental frequency values. A neural network based approach is proposed for computing the posterior probability that the estimated fundamental frequency is correct. Experiments are conducted on the PTDB-TUG pitch-tracking database, using three fundamental frequency estimation algorithms. Boyuan Deng, Denis Jouvet, Yves Laprie, Ingmar Steiner, Aghilas Sini |
ICASSP | 3 |
| 2017 | Glottal Opening and Strategies of Production of FricativesabstractInternational audience Benjamin Elie, Yves Laprie |
INTERSPEECH | 2 |
| 2017 | End-to-End Acoustic Feedback in Language Learning for Correcting Devoiced French Final-FricativesabstractInternational audience Sucheta Ghosh, Camille Fauth, Yves Laprie, Aghilas Sini |
INTERSPEECH | 3 |
| 2016 | A glottal chink model for the synthesis of voiced fricativesabstractThis paper presents a simulation framework that enables a glottal chink model to be integrated into a time-domain continuous speech synthesizer along with self-oscillating vocal folds. The glottis is then made up of two main separated components: a self-oscillating part and a constantly open chink. This feature allows the simulation of voiced fricatives, thanks to a self-oscillating model of the vocal folds to generate the voiced source, and the glottal opening that is necessary to generate the frication noise. Numerical simulations show the accuracy of the model to simulate voiced fricative, and also phonetic assimilation, such as sonorization and devoicing. The simulation framework is also used to show that the phonatory/articulatory space for generating voiced fricatives is different according to the desired sound: for instance, the minimal glottal opening for generating frication noise is shorter for /z/ than for /3/. Benjamin Elie, Yves Laprie |
ICASSP | 2 |
| 2016 | L1-L2 Interference: The Case of Final Devoicing of French Voiced Fricatives in Final Position by German LearnersabstractInternational audience Sucheta Ghosh, Camille Fauth, Aghilas Sini, Yves Laprie |
INTERSPEECH | 4 |
| 2016 | The IFCASL Corpus of French and German Non-native and Native Read Speech
Jürgen Trouvain, Anne Bonneau, Vincent Colotte, Camille Fauth, Dominique Fohr, Denis Jouvet, Jeanin Jügler, Yves Laprie, Odile Mella, Bernd Möbius, Frank Zimmerer |
LREC | 8 |
| 2016 | Extension of the single-matrix formulation of the vocal tract: Consideration of bilateral channels and connection of self-oscillating models of the vocal folds with a glottal chink
Benjamin Elie, Yves Laprie |
Speech Commun. | 2 |
| 2014 | Designing a Bilingual Speech Corpus for French and German Language Learners: a Two-Step Process
Camille Fauth, Anne Bonneau, Frank Zimmerer, Jürgen Trouvain, Bistra Andreeva, Vincent Colotte, Dominique Fohr, Denis Jouvet, Jeanin Jügler, Yves Laprie, Odile Mella, Bernd Möbius |
LREC | 10 |
| 2013 | Articulatory copy synthesis from cine x-ray filmsabstractThis paper deals with articulatory copy synthesis from X-ray films. The underlying articulatory synthesizer uses an aerodynamic and an acoustic simulation using target area functions, F0 and transition patterns from one area function to the next as input data. The articulators, tongue in particular, have been delineated by hand or semi-automatically from the X-ray films. A specific attention has been paid on the determination of the centerline of the vocal tract from the image and on the coordination between glottal area and vocal tract constrictions since both aspects strongly impact on the acoustics. Experiments show that good quality speech can be resynthesized even if the interval between two images is 40\,ms. The same approach could be easily applied to cine MRI data. Yves Laprie, Matthieu Loosvelt, Shinji Maeda, Rudolph Sock, Fabrice Hirsch |
INTERSPEECH | 1 |
| 2013 | Vowel and prosodic factor dependent variations of vocal-tract lengthabstractInternational audience Shinji Maeda, Yves Laprie |
INTERSPEECH | 2 |
| 2009 | Registration of multimodal data for estimating the parameters of an articulatory modelabstractBeing able to animate a speech production model with articulatory data would open applications in many domains. In this paper, we first consider the problem of acquiring articulatory data from non invasive image and sensor modalities: dynamic ultrasound (US) images, stereovision 3D data, electromagnetic sensors and MRI. We here especially focus on automatic registration methods which enable the fusion of the articulatory features in a common frame. We then derive articulatory parameters by fitting these features with Maeda's model. To our knowledge, it is the first attempt to derive articulatory parameters from features automatically extracted and registered between the modalities. Results prove the soundness of the approach and the reliability of the fused articulatory data. Michael Aron, Asterios Toutios, Marie-Odile Berger, Erwan Kerrien, Brigitte Wrobel-Dautcourt, Yves Laprie |
ICASSP | 6 |
| 2009 | Articulatory modeling based on semi-polar coordinates and guided PCA techniqueabstractInternational audience Yves Laprie, Julie Busset, Fabrice Hirsch |
INTERSPEECH | 2 |
| 2009 | An evaluation of formant tracking methods on an Arabic databaseabstractIn this paper we present a formant database of Arabic used to evaluate our new automatic formant tracking algorithm based on Fourier ridges detection. In this method we have introduced a continuity constraint based on the computation of centres of gravity for a set of formant candidates. This leads to connect a frame of speech to its neighbours and thus improves the robustness of tracking. The formant trajectories obtained by the algorithm proposed are compared to those of the hand edited formant database and those given by Praat with LPC data. Imen Jemaa, Oussama Rekhis, Kaïs Ouni, Yves Laprie |
INTERSPEECH | 4 |
| 2009 | A robust variational method for the acoustic-to-articulatory problemabstractInternational audience Blaise Potard, Yves Laprie |
INTERSPEECH | 2 |
| 2009 | Efficient likelihood evaluation and dynamic Gaussian selection for HMM-based speech recognition
Ghazi Bouselmi, Yves Laprie, Jean Paul Haton |
Comput. Speech Lang. | 3 |
| 2008 | Dynamic Gaussian selection technique for speeding up HMM-based continuous speech recognitionabstractA fast likelihood computation approach called dynamic Gaussian selection (DGS) is proposed for HMM-based continuous speech recognition. DGS approach is a one-pass search technique which generates a dynamic shortlist of Gaussians for each state during the procedure of likelihood computation. The shortlist consists of the Gaussians which make prominent contribution to the likelihood. In principle, DGS is an extension of the technique of Partial Distance Elimination, and it requires almost no additional memory for the storage of Gaussian shortlists. DGS algorithm has been implemented by modifying the likelihood computation module in HTK 3.4 system. Results from experiments on TIMIT and HIWIRE corpora indicate that this approach can speed up the likelihood computation significantly without introducing apparent additional recognition error. Ghazi Bouselmi, Dominique Fohr, Yves Laprie |
ICASSP | 4 |
| 2007 | Acquisition and synchronization of multimodal articulatory dataabstractInternational audience Michael Aron, Nicolas Ferveur, Erwan Kerrien, Marie-Odile Berger, Yves Laprie |
INTERSPEECH | 5 |
| 2007 | Compact representations of the articulatory-to-acoustic mappingabstractArticulatory codebooks are very often used to represent the articulatory-to-acoustic mapping. They thus need to be com-pact while offering a very good acoustic precision. This paper presents a method of articulatory codebook construction more general than that of Ouni [1] in the sense that the articulatory-to-acoustic mapping is approximated by multivariable polyno-mials. The second major contribution concerns the subdivision process which finds out the most efficient subdivision, i.e. that which minimizes the size of the codebook while guarantying a very good acoustic precision. Experiments carried out show that the size of the codebook can be divided by a factor of 20, and simultaneously, the acous-tic precision can improved by a factor of 2 by using second order polynomials together with this new construction strategy. Index Terms: acoustic-to-articulatory inversion, codebook, polynomial interpolation. Blaise Potard, Yves Laprie |
INTERSPEECH | 2 |
| 2007 | A phonetic concatenative approach of labial coarticulationabstractInternational audience Vincent Robert, Yves Laprie, Anne Bonneau |
INTERSPEECH | 2 |
| 2005 | An elitist approach for extracting automatically well-realized speech sounds with high confidenceabstractThis paper presents an ëlitist approach\" for extracting automatically well-realized speech sounds with high confidence. The elitist approach uses a speech recognition system based on Hidden Markov Models (HMM). The HMM are trained on speech sounds which are systematically well-detected in an iterative procedure. The results show that, by using the HMM models defined in the training phase, the speech recognizer detects reliably specific speech sounds with a small rate of errors. Jean-Baptiste Maj, Anne Bonneau, Dominique Fohr, Yves Laprie |
INTERSPEECH | 4 |
| 2005 | Using phonetic constraints in acoustic-to-articulatory inversionabstractThe goal of this work is to recover articulatory information from the speech signal by acoustic-to-articulatory inversion. One of the main difficulties with inversion is that the problem is underdetermined and inversion methods generally offer no guarantee on the phonetical realism of the inverse solutions. A way to adress this issue is to use additional phonetic constraints. Knowledge of the phonetic caracteristics of French vowels enable the derivation of reasonable articulatory domains in the space of Maeda parameters: given the formants frequencies (F1,F2,F3) of a speech sample, and thus the vowel identity, an articulatory can be derived. The space of formants frequencies is partitioned into vowels, using either speaker-specific data or generic information on formants. Then, to each articulatory vector can be associated a phonetic score varying with the distance to the ideal domain associated with the corresponding vowel. Inversion experiments were conducted on isolated vowels and vowel-to-vowel transitions. Articulatory parameters were compared with those obtained without using these constraints and those measured from X-ray data. Blaise Potard, Yves Laprie |
INTERSPEECH | 2 |
| 2005 | Strategies of labial coarticulationabstractInternational audience Vincent Robert, Brigitte Wrobel-Dautcourt, Yves Laprie, Anne Bonneau |
INTERSPEECH | 3 |
| 2004 | A concurrent curve strategy for formant trackingabstractColloque avec actes et comité de lecture. internationale. Yves Laprie |
INTERSPEECH | 1 |
| 2002 | A copy synthesis method to pilot the klatt synthesiserabstractColloque avec actes et comité de lecture. internationale. Yves Laprie, Anne Bonneau |
INTERSPEECH | 1 |
| 2002 | Introduction of constraints in an acoustic-to-articulatory inversion method based on a hypercubic articulatory tableabstractOur acoustic to articulatory inversion method exploits an original articulatory table structured in the form of a hypercube hierarchy. The articulatory space is decomposed into regions where the articulatory-to-acoustic mapping is linear. Each region is represented by a hypercube. The inversion procedure retrieves articulatory vectors corresponding to an acoustic entry from the hypercube table. A dynamic procedure is used to recover the best articulatory trajectory according to a minimum articulatory effort criterion. The inversion ensures that inverse articulatory parameters generate original formant trajectories with a very good precision, but not that they are realistic from a phonetic point of view. This papers shows how additional simple articulatory constraints can be incorporated in the inversion process. Constraints are implemented in the form of bonus attached to the points which verify the constraints imposed. This enables the inversion to be guided towards more realistic inverse articulatory trajectories. Yves Laprie, Slim Ouni |
INTERSPEECH | 1 |
| 2001 | Suppression of phasiness for time-scale modifications of speech signals based on a shape invariance propertyabstractTime-scale modifications of speech signals, based on frequency-domain techniques, are hampered by two important artifacts which are "phasiness" and "transient smearing". They correspond to the destruction of the shape of the original signal, i.e. the de-synchronization between the phases of the frequency components. This paper describes an algorithm that preserves the shape invariance of speech signals in the context of a phase vocoder. Phases are corrected at the onset of each voiced region. Modified signals, even for large expansion factors, are of high quality and free from transient smearing or phasiness. A demonstration is proposed in the web page: http://www.loria.fr/-jdm/PhaseVocoder/index.html. Joseph Di Martino, Yves Laprie |
ICASSP | 2 |
| 2001 | Perceptual experiments on enhanced and slowed down speech sentences for second language acquisitionabstractColloque avec actes et comité de lecture. internationale. Vincent Colotte, Yves Laprie, Anne Bonneau |
INTERSPEECH | 2 |
| 2001 | Burst segmentation and evaluation of acoustic cuesabstractColloque avec actes et comité de lecture. internationale. Yves Laprie, Anne Bonneau |
INTERSPEECH | 1 |
| 2001 | Exploring the null space of the acoustic-to- articulatory inversion using a hypercube codebookabstractColloque avec actes et comité de lecture. internationale. Slim Ouni, Yves Laprie |
INTERSPEECH | 2 |
| 2000 | Automatic enhancement of speech intelligibilityabstractThis paper presents a speech signal transformation which slows down speech signals selectively and enhances some important acoustic cues. This transformation can be used not only for hearing aids but also for second language acquisition by facilitating oral comprehension. Selective slowing down relies on the use of the TD-PSOLA synthesis method. An automatic pitch marking algorithm was designed to apply this method automatically. The strategy used to control slowing down exploits a spectral variation function which locates rapid spectral changes. The enhancement simply consists of amplifying stop bursts and unvoiced fricatives. These acoustic cues are detected automatically through the examination of energy criteria. This approach was evaluated in the context of second language acquisition, more precisely by evaluating improvements in oral comprehension. Transformations triggered properly, i.e. the signal regions modified are those which were expected to be modified. Experiments show that the oral comprehension is improved. Vincent Colotte, Yves Laprie |
ICASSP | 2 |
| 2000 | Improving acoustic-to-articulatory inversion by using hypercube codebooksabstractColloque avec actes et comité de lecture. internationale. Slim Ouni, Yves Laprie |
INTERSPEECH | 2 |
| 1999 | Design of hypercube codebooks for the acoustic-to-articulatory inversion respecting the non-linearities of the articulatory-to-acoustic mapping
Slim Ouni, Yves Laprie |
EUROSPEECH | 2 |
| 1999 | An efficient F0 determination algorithm based on the implicit calculation of the autocorrelation of the temporal excitation signalabstractColloque avec actes et comité de lecture. Joseph Di Martino, Yves Laprie |
EUROSPEECH | 2 |
| 1998 | A variational approach for estimating vocal tract shapes from the speech signalabstractThis paper present's a novel approach to recovering articulatory trajectories from the speech signal using a variational calculus method and Maeda's (1979) articulatory model. The acoustic-to-articulatory mapping is generally assessed by a double criterion: the acoustic proximity of results to acoustic data and the smoothness of articulatory trajectories. Most of the existing methods are unable to exploit the two criteria simultaneously or at least at the same level. On the other hand, our variational calculus approach combines the two criteria simultaneously and ensures the global acoustic and articulatory consistency without further optimization. This method gives rise to an iterative process which optimizes a startup solution given by an improved lookup algorithm. Codebooks generated with an articulatory model show nonuniform sampling of the acoustic space due to nonlinearities of the acoustic-to-articulatory mapping. We therefore designed an improved lookup algorithm building realistic articulatory trajectories which are not necessarily defined throughout the speech signal. Yves Laprie, Bruno Mathieu |
ICASSP | 1 |
| 1998 | The effect of modifying formant amplitudes on the perception of French vowels generated by copy synthesis
Anne Bonneau, Yves Laprie |
ICSLP | 2 |
| 1997 | Adaptation of Maeda's model for acoustic to articulatory inversion
Bruno Mathieu, Yves Laprie |
EUROSPEECH | 2 |
| 1996 | Tracking articulators in X-ray images with minimal user interaction: Example of the tongue extractionabstractVocal tract X-ray image sequences are used to study articulatory phenomena and to design approximate articulatory models. The purpose of this paper is to describe an automatic tracking tool for extracting the contours of the tongue which is the most important articulator. Tracking the tongue in X ray images is an arduous task because it appears as a weak contour and above all because it is immersed among a lot of contours (teeth, palate, ...). Hence we have developed a robust algorithm that makes a snake based method and a motion based method cooperate. Significant results show the strength of the approach. Marie-Odile Berger, Yves Laprie |
ICIP (2) | 2 |
| 1996 | A new search algorithm in segmentation lattices of speech signals
Jean-Luc Husson, Yves Laprie |
ICSLP | 2 |
| 1996 | Extraction of tongue contours in x-ray images with minimal user interactionabstractIn spite of the development of new imaging techniques, X-ray images still keep a prominent place to studying articulatory phenomena.Indeed, they are still unsurpassed to obtain an overall view of the moving vocal tract.However X-ray images require that articulators contours are extracted by hand which is a tedious task.This paper describes an approach towards the automatization of the tongue contour extraction.The "Snake" method introduced in computer vision to extract contours is unable alone to achieve the task.Therefore we make "Snakes" cooperate with an optical flow method applied where contours are not sufficiently isolated from spurious contours.Our experiments have shown that the tongue is tracked successfully when it is visible, and that interaction with the user remains necessary when the tongue is obscured. Yves Laprie, Marie-Odile Berger |
ICSLP | 1 |
| 1996 | A new method for speech delexicalization, and its application to the perception of French prosody
Vincent Pagel, Noëlle Carbonell, Yves Laprie |
ICSLP | 3 |
| 1996 | Using decision trees to construct optimal acoustic cues
Sandrine Robbe-Reiter, Anne Bonneau, Sylvie Coste-Marquis, Yves Laprie |
ICSLP | 4 |
| 1996 | Cooperation of regularization and speech heuristics to control automatic formant tracking
Yves Laprie, Marie-Odile Berger |
Speech Commun. | 1 |
| 1995 | Influence of a prior knowledge of the vocalic context on stop burst perception
Anne Bonneau, Linda Djezzar, Yves Laprie |
EUROSPEECH | 3 |
| 1994 | A new paradigm for reliable automatic formant trackingabstractA new algorithm for tracking formants automatically is presented. From rough formant hypotheses a regularization method is used to provide formant trajectories both close to the spectrogram edge lines and sufficiently regular. The formant hypotheses are obtained by labelling edge lines of LPC or cepstrally smoothed spectrograms in terms of formants. Speech knowledge, in the form of admissible domains for F1, F2, F3, F1 vs. F2, F2 vs. F3 and of formant level is used to obtain consistent labellings. The advantage of this algorithm is that it provides reliable formant trajectories especially in regions where a knowledge of the formant transitions give important information about the location of articulation for consonants. We present very encouraging tracking results on a corpus of sentences consisting of stops and vowels. > Yves Laprie, Marie-Odile Berger |
ICASSP (2) | 1 |
| 1994 | Knowledge-Based Techniques in Acoustic-Phonetic Decoding of Speech: Interest and LimitationsabstractA major step in the process of speech understanding is the acoustic-phonetic decoding which can be defined as the automatic mapping of the continuous speech wave into a set of predetermined linguistic units such as phones, diphones, syllables, etc. This paper relates to the approach of this problem which consists in exploiting an explicit description of all kinds of available knowledge about the speech communication phenomena, in the general framework of an artificial intelligence knowledge-based system. We will first recall the main difficulties of acoustic-phonetic decoding with a practical example. We will then present the APHODEX system that we have been designing for the past eight years, in terms of software architecture and of knowledge representation and reasoning. The practical evaluation of this system will then be carried out at the different levels of feature extraction, segmentation and labelling. Finally, we will discuss the limitations of our approach and present the ongoing effort to overcome these limitations, especially through the use of abductive reasoning. Dominique Fohr, Jean Paul Haton, Yves Laprie |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 1993 | Perception of French stop bursts, implications for stop identificationabstractAn experiment is presented here concerning perception of release bursts in a corpus of natural tokens of French /p,t,k/ in CV context. The focus of the study was the perceptual role of spectral characteristics of the release burst; therefore tokens have the same duration of nearly 25 ms so that neither VOT nor formant transitions may influence listeners. Two training sessions appeared to be necessary to reach the utmost of listeners' ability to identify stops. Results show that the recognition rate are approximately the same for /p/ (89%), /t/ (87%) and /k/ (86%). More precisely, identification of /k/ in the context of a front vowel is not so high (75%) and is often confused with /t/; /k/ in the context of a back vowel is very well identified (98%). /t/in the context of /y/ or /OE/ is confused (22%) with /k/. These very high identification rates militate in favour of acoustical invariants, even if, in some cases (for example in the context of a back vowel), the knowledge of the subsequ... Anne Bonneau, Linda Djezzar, Yves Laprie |
EUROSPEECH | 3 |
| 1992 | A Model for Hypothetical Reasoning Applied to Speech Recognition
Anne Bonneau, François Charpillet, Sylvie Coste-Marquis, Jean Paul Haton, Yves Laprie, Pierre Marquis |
ECAI | 5 |
| 1992 | Global active method for automatic formant tracking guided by local processingabstractIn a formant tracking algorithm, the combination of both local viewpoint and global aspect of formant trajectories is difficult. The authors present an approach which incorporates local processing to build elementary tracks and an active method which generates rough formant track hypotheses and makes the resulting trajectories move towards the formant trajectories.> Marie-Odile Berger, Yves Laprie |
ICPR (3) | 2 |
| 1992 | Two level acoustic cues for consistent stop identificationabstractExtrait de : Proc. Intern. Conf. on Spoken Language Processing, Banff (Alberta, Canada), October 1992 Anne Bonneau, Sylvie Coste-Marquis, Linda Djezzar, Yves Laprie |
ICSLP | 4 |
| 1992 | Active models for regularizing formant trajectoriesabstractExtrait de : Proceedings International Conference on Spoken Language Processing, Banff (Alberta, Canada), October 1992, pages 815-818 Yves Laprie, Marie-Odile Berger |
ICSLP | 1 |
| 1991 | Phonetic triplets in acoustic-phonetic decoding of continuous speechabstractA knowledge-based approach which stores knowledge in the form of contextual prototypes called triplets is presented. A triplet consists of an acoustic description using acoustic events (burst features, formant trajectories, etc.) and a component representing the acoustic correlates used by a human expert. Attention is given to vowel and plosive centered triplets and the corresponding matching algorithm relying on the acoustic description. Representing knowledge in the form of prototypes allows the recognition system to compare reference triplets to each other. The same acoustic comparisons can be made between triplet instances. It is shown how a relaxation algorithm using this type of acoustic comparison (which can be viewed as constraints) allows the system to increase the consistency of global triplet labeling for the sentence to be decoded.> Yves Laprie |
ICASSP | 1 |
| 1990 | Optimum spectral peak track interpretation in terms of formantsabstractPublie dans : Proceedings ICSLP90 (International conference on spoken language processing), Kobe (Japan), November 1990 Yves Laprie |
ICSLP | 1 |
| 1990 | Phonetic triplets in knowledge based approach of acoustic-phonetic decodingabstractPublie dans : Proceedings ICSLP90 (International conference on spoken language processing), Kobe (Japan), November 1990 Yves Laprie, Jean Paul Haton, Jean-Marie Pierrel |
ICSLP | 1 |
| 1989 | Snorri: an interactive tool for speech analysisabstractPublie dans : Proceedings EUROSPEECH 89 (European conference on speech communication and technology), Paris, September 1989 Dominique Fohr, Yves Laprie |
EUROSPEECH | 2 |
| 1989 | Formant tracking adapted to acoustic-phonetic decodingabstractPublie dans : Proceedings EUROSPEECH 89 (European conference on speech communication and technology), Paris, September 1989 Yves Laprie |
EUROSPEECH | 1 |