Denis Jouvet

dblp:95/6584 · DBLP profile ↗
← Back
84ranked-venue papers
18as first author
10since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 78 · 18 first-author · 9 since 2021Artificial intelligence and machine learning · 51 · 10 first-author · 8 since 2021
YearPublicationVenuePosition
2023 Self-supervised learning with Diffusion-based multichannel speech enhancement for speaker verification under noisy conditions
abstract
Proceedings of Interspeech 2023
Sandipana Dowerah, Ajinkya Kulkarni, Romain Serizel, Denis Jouvet
INTERSPEECH4
2022 Are disentangled representations all you need to build speaker anonymization systems?
abstract
International audience
Pierre Champion, Anthony Larcher, Denis Jouvet
INTERSPEECH3
2022 Analysis of expressivity transfer in non-autoregressive end-to-end multispeaker TTS systems
abstract
International audience
Ajinkya Kulkarni, Vincent Colotte, Denis Jouvet
INTERSPEECH3
2022 Barlow Twins self-supervised learning for robust speaker recognition
abstract
International audience
Mohammad MohammadAmini, Driss Matrouf, Jean-François Bonastre, Sandipana Dowerah, Romain Serizel, Denis Jouvet
INTERSPEECH6
2022 Adapting Language Models When Training on Privacy-Transformed Data
abstract
In recent years, voice-controlled personal assistants have revolutionized the interaction with smart devices and mobile applications. The collected data are then used by system providers to train language models (LMs). Each spoken message reveals personal information, hence removing private information from the input sentences is necessary. Our data sanitization process relies on recognizing and replacing named entities by other words from the same class. However, this may harm LM training because privacy-transformed data is unlikely to match the test distribution. This paper aims to fill the gap by focusing on the adaptation of LMs initially trained on privacy-transformed sentences using a small amount of original untransformed data. To do so, we combine class-based LMs, which provide an effective approach to overcome data sparsity in the context of n-gram LMs, and neural LMs, which handle longer contexts and can yield better predictions. Our experiments show that training an LM on privacy-transformed data result in a relative 11% word error rate (WER) increase compared to training on the original untransformed data, and adapting that model on a limited amount of original untransformed data leads to a relative 8% WER improvement over the model trained solely on privacy-transformed data.
M. A. Tugtekin Turan, Dietrich Klakow, Emmanuel Vincent 0001, Denis Jouvet
LREC4
2022 Joint Optimization of Diffusion Probabilistic-Based Multichannel Speech Enhancement with Far-Field Speaker Verification
abstract
Smart devices using speaker verification are getting equipped with multiple microphones, improving spatial ambiguity and directivity. However, unlike other speech-based applications, the performance of speaker verification degrades in far-field scenarios due to the adverse effects of a noisy environment and room reverberation. This paper presents a novel diffusion probabilistic models-based multichannel speech enhancement as a front-end for the ECAPA-TDNN speaker verification system in a far-field noisy-reverberant scenario. The proposed approach incorporates a two-stage training approach. In the first stage, we individually train the speech enhancement and speaker verification modules. In the second stage, we combined both modules and trained them jointly. We use similarity-preserving knowledge distillation loss that guides the network to produce similar activation for enhanced signals like clean signals. Joint optimization achieved the best results on synthetic and VOiCES datasets.
Sandipana Dowerah, Romain Serizel, Denis Jouvet, Mohammad MohammadAmini, Driss Matrouf
SLT3
2021 On the Invertibility of a Voice Privacy System Using Embedding Alignment
abstract
This paper explores various attack scenarios on a voice anonymization system using embeddings alignment techniques. We use Wasserstein-Procrustes (an algorithm initially designed for unsupervised translation) or Procrustes analysis to match two sets of$x$-vectors, before and after voice anonymization, to mimic this transformation as a rotation function. We compute the optimal rotation and compare the results of this approximation to the official Voice Privacy Challenge results. We show that a complex system like the baseline of the Voice Privacy Challenge can be approximated by a rotation, estimated using a limited set of$x$-vectors. This paper studies the space of solutions for voice anonymization within the specific scope of rotations. Rotations being reversible, the proposed method can recover up to 62% of the speaker identities from anonymized embeddings.
Pierre Champion, Thomas Thebaud, Gaël Le Lan, Anthony Larcher, Denis Jouvet
ASRU5
2021 Modeling and Training Strategies for Language Recognition Systems
abstract
International audience
Raphaël Duroselle, Md. Sahidullah, Denis Jouvet, Irina Illina
Interspeech3
2021 Language Recognition on Unknown Conditions: The LORIA-Inria-MULTISPEECH System for AP20-OLR Challenge
abstract
International audience
Raphaël Duroselle, Md. Sahidullah, Denis Jouvet, Irina Illina
Interspeech3
2021 Duration modelling and evaluation for Arabic statistical parametric speech synthesis
Imene Zangar, Zied Mnasri, Vincent Colotte, Denis Jouvet
Multim. Tools Appl.4
2020 Metric Learning Loss Functions to Reduce Domain Mismatch in the x-Vector Space for Language Recognition
abstract
International audience
Raphaël Duroselle, Denis Jouvet, Irina Illina
INTERSPEECH2
2020 Kaldi-Web: An Installation-Free, On-Device Speech Recognition System
Mathieu Hu, Laurent Pierron, Emmanuel Vincent 0001, Denis Jouvet
INTERSPEECH4
2020 Transfer Learning of the Expressivity Using FLOW Metric Learning in Multispeaker Text-to-Speech Synthesis
abstract
International audience
Ajinkya Kulkarni, Vincent Colotte, Denis Jouvet
INTERSPEECH3
2020 Correlation Between Prosody and Pragmatics: Case Study of Discourse Markers in French and English
abstract
This paper investigates the prosodic characteristics of French and English discourse markers according to their pragmatic meaning in context. The study focusses on three French discourse markers (alors ['so'], bon ['well'], and donc ['so']) and three English markers (now, so, and well). Hundreds of occurrences of discourse markers were automatically extracted from French and English speech corpora and manually annotated with pragmatic functions labels. The paper compares the pro-sodic characteristics of discourse markers in different speech styles and in two languages. The first comparison is carried out with respect to two different speech styles in French: spontaneous speech vs. prepared speech. The other comparison of the prosodic characteristics is conducted between two languages, French vs. English, on the prepared speech. Results show that some pragmatic functions of discourse markers bring about specific prosodic behaviour in terms of presence and position of pauses, and their F0 articulation in their immediate context. Moreover, similar pragmatic functions frequently share similar prosodic characteristics, even across languages.
Lou Lee, Denis Jouvet, Katarina Bartkova, Yvon Keromnes, Mathilde Dargnat
INTERSPEECH2
2020 Achieving Multi-Accent ASR via Unsupervised Acoustic Model Adaptation
abstract
Current automatic speech recognition (ASR) systems trained on native speech often perform poorly when applied to non-native or accented speech. In this work, we propose to compute x-vector-like accent embeddings and use them as auxiliary inputs to an acoustic model trained on native data only in order to improve the recognition of multi-accent data comprising native, non-native, and accented speech. In addition, we leverage untranscribed accented training data by means of semi-supervised learning. Our experiments show that acoustic models trained with the proposed accent embeddings outperform those trained with conventional i-vector or x-vector speaker embeddings, and achieve a 15% relative word error rate (WER) reduction on non-native and accented speech w.r.t. acoustic models trained with regular spectral features only. Semi-supervised training using just 1 hour of untranscribed speech per accent yields an additional 15% relative WER reduction w.r.t. models trained on native data only.
M. A. Tugtekin Turan, Emmanuel Vincent 0001, Denis Jouvet
INTERSPEECH3
2017 Towards confidence measures on fundamental frequency estimations
abstract
The fundamental frequency is one of the prosodic parameters, and many algorithms have been developed for estimating the fundamental frequency of speech signals. Most of them provide good results on good quality speech signals, but their performance degrades when dealing with noisy signals. Moreover, although some provide a probability for the voicing decision, none of them indicate how reliable the estimated fundamental frequency is. In this paper, we investigate the computation of a confidence (or reliability) measure on the estimated fundamental frequency values. A neural network based approach is proposed for computing the posterior probability that the estimated fundamental frequency is correct. Experiments are conducted on the PTDB-TUG pitch-tracking database, using three fundamental frequency estimation algorithms.
Boyuan Deng, Denis Jouvet, Yves Laprie, Ingmar Steiner, Aghilas Sini
ICASSP2
2016 The IFCASL Corpus of French and German Non-native and Native Read Speech
Jürgen Trouvain, Anne Bonneau, Vincent Colotte, Camille Fauth, Dominique Fohr, Denis Jouvet, Jeanin Jügler, Yves Laprie, Odile Mella, Bernd Möbius, Frank Zimmerer
LREC6
2015 Discriminative uncertainty estimation for noise robust ASR
abstract
We consider the problem of uncertainty estimation for noiserobust ASR. Existing uncertainty estimation techniques improve ASR accuracy but they still exhibit a gap compared to the use of oracle uncertainty. This comes partly from the highly non-linear feature transformation and from additional assumptions such as Gaussian distribution and independence between frequency bins in the spectral domain. In this paper, we propose a method to rescale the estimated feature-domain full uncertainty covariance matrix in a statedependent fashion according to a discriminative criterion. The state-dependent and feature index-dependent scaling factors are learned from development data. Experimental evaluation on Track 1 of the 2nd CHiME challenge data shows that discriminative rescaling leads to better results than generative rescaling. Moreover, discriminative rescaling of the Wiener uncertainty estimator leads to 12% relative word error rate reduction compared to discriminative rescaling of the alternative estimator in [1].
Dung T. Tran, Emmanuel Vincent 0001, Denis Jouvet
ICASSP3
2015 Nonparametric Uncertainty Estimation and Propagation for Noise Robust ASR
abstract
We consider the framework of uncertainty propagation for automatic speech recognition (ASR) in highly nonstationary noise environments. Uncertainty is considered as the variance of speech distortion. Yet, its accurate estimation in the spectral domain and its propagation to the feature domain remain difficult. Existing methods typically rely on a single uncertainty estimator and propagator fixed by mathematical approximation. In this paper, we propose a new paradigm where we seek to learn more powerful mappings to predict uncertainty from data. We investigate two such possible mappings: linear fusion of multiple uncertainty estimators/propagators and nonparametric uncertainty estimation/propagation. In addition, a procedure to propagate the estimated spectral-domain uncertainty to the static Mel frequency cepstral coefficients (MFCCs), to the log-energy, and to their first- and second-order time derivatives is proposed. This results in a full uncertainty covariance matrix over both static and dynamic MFCCs. Experimental evaluation on Tracks 1 and 2 of the 2nd CHiME Challenge resulted in up to 29% and 28% relative keyword error rate reduction with respect to speech enhancement alone.
Dung T. Tran, Emmanuel Vincent 0001, Denis Jouvet
IEEE ACM Trans. Audio Speech Lang. Process.3
2014 Extension of uncertainty propagation to dynamic MFCCS for noise robust ASR
abstract
Uncertainty propagation has been successfully employed for speech recognition in nonstationary noise environments. The uncertainty about the features is typically represented as a diagonal covariance matrix for static features only. We present a framework for estimating the uncertainty over both static and dynamic features as a full covariance matrix. The estimated covariance matrix is then multiplied by scaling coefficients optimized on development data. We achieve 21% relative error rate reduction on the 2nd CHiME Challenge with respect to conventional decoding without uncertainty, that is five times more than the reduction achieved with diagonal uncertainty covariance for static features only.
Dung T. Tran, Emmanuel Vincent 0001, Denis Jouvet
ICASSP3
2014 Fusion of multiple uncertainty estimators and propagators for noise robust ASR
abstract
Uncertainty decoding has been successfully used for speech recognition in highly nonstationary noise environments. Yet, accurate estimation of the uncertainty on the denoised signals and propagation to the features remain difficult. In this work, we propose to fuse the uncertainty estimates obtained from different uncertainty estimators and propagators by linear combination. The fusion coefficients are optimized by minimizing a measure of divergence with oracle estimates on development data. Using the Kullback-Leibler divergence, we obtain 18% relative error rate reduction on the 2nd CHiME Challenge with respect to conventional decoding, that is about twice as much as the reduction achieved by the best single uncertainty estimator and propagator.
Dung T. Tran, Emmanuel Vincent 0001, Denis Jouvet
ICASSP3
2014 Component structuring and trajectory modeling for speech recognition
abstract
International audience
Arseniy Gorin, Denis Jouvet
INTERSPEECH2
2014 About combining forward and backward-based decoders for selecting data for unsupervised training of acoustic models
abstract
This paper introduces the combination of speech decoders for selecting automatically transcribed speech data for unsupervised training or adaptation of acoustic models. Here, the combination relies on the use of a forward-based and a backward-based decoder. Best performance is achieved when selecting automatically transcribed data (speech segments) that have the same word hypotheses when processed by the Sphinx forward-based and the Julius backward-based transcription systems, and this selection process outperforms confidence measure based selection. Results are reported and discussed for adaptation and for full training from scratch, using data resulting from various selection processes, whether alone or in addition to the baseline manually transcribed data. Overall, selecting automatically transcribed speech segments that have the same word hypotheses when processed by the Sphinx forward-based and Julius backward-based recognizers, and adding this automatically transcribed and selected data to the manually transcribed data leads to significant word error rate reductions on the ESTER2 data when compared to the baseline system trained only on manually transcribed speech data.
Denis Jouvet, Dominique Fohr
INTERSPEECH1
2014 Hybrid language models for speech transcription
abstract
This paper analyzes the use of hybrid language models for automatic speech transcription.The goal is to later use such an approach as a support for helping communication with deaf people, and to run it on an embedded decoder on a portable device, which introduces constraints on the model size.The main linguistic units considered for this task are the words and the syllables.Various lexicon sizes are studied by setting thresholds on the word occurrence frequencies in the training data, the less frequent words being therefore syllabified.A recognizer using this kind of language model can output between 62% and 96% of words (with respect to the thresholds on the word occurrence frequencies; the other recognized lexical units are syllables).By setting different thresholds on the confidence measures associated to the recognized words, the most reliable word hypotheses can be identified, and they have correct recognition rates between 70% and 92%.
Luiza Orosanu, Denis Jouvet
INTERSPEECH2
2014 Designing a Bilingual Speech Corpus for French and German Language Learners: a Two-Step Process
Camille Fauth, Anne Bonneau, Frank Zimmerer, Jürgen Trouvain, Bistra Andreeva, Vincent Colotte, Dominique Fohr, Denis Jouvet, Jeanin Jügler, Yves Laprie, Odile Mella, Bernd Möbius
LREC8
2013 Combining forward-based and backward-based decoders for improved speech recognition performance
abstract
Combining outputs of speech recognizers is a known way of increasing speech recognition performance. The ROVER approach handles efficiently such combinations. In this paper we show that the best performance is not achieved by combining the outputs of the best set of recognizers, but rather by combining outputs of recognizers that rely on different processing components, and in particular on a different order (backward vs. forward) for processing speech frames. Indeed, much better speech recognition results were obtained by combining outputs of sphinx-based recognizers with outputs of Julius-based recognizers than by combining the same number of outputs from only sphinx-based recognizers, even if the individual sphinx-based systems led to better results than the individual Julius-based recognizers. Further experiments have also been conducted using sphinx-based tools for processing speech frames in reverse order (i.e. backward in time). The results clearly show that combining forward-based and backward-based decoders provide significant improvement with respect to a combination of forward only or backward only decoders. Experiments have been conducted on the ESTER2 and ETAPE speech corpora. Overall, combining sphinx-based and Julius-based systems led to 18.6% word error rate on ESTER2 test data, and 24.5% word error rate on ETAPE test data.
Denis Jouvet, Dominique Fohr
INTERSPEECH1
2013 Comparison of approaches for an efficient phonetic decoding
abstract
This article analyzes the phonetic decoding performance obtained with different choices of linguistic units.The context is to later use such an approach as a support for helping communication with deaf people, and to run it on an embedded decoder on a portable terminal, which introduces constrains on the model size.As a first step, this paper presents and analyses the performance of various approaches.Two baseline systems are considered, one relying on a large vocabulary speech recognizer, and another one relying on a phonetic n-gram language model.Then syllable-based lexicons and language models are investigated.Various lexicon sizes are studied by setting thresholds on their frequency of occurrences in the training data.Evaluations are conducted on the ESTER and ETAPE speech corpora.Keeping only the most frequent syllables leads to a limited-size lexicon and language model, which nevertheless provides good phonetic decoding performance.The phone error rate is only 4% worse (absolute) than the phone error rate obtained with the large vocabulary recognizer, and much better than the phone error rate obtained with the phone n-gram language model.
Luiza Orosanu, Denis Jouvet
INTERSPEECH2
2012 Evaluating grapheme-to-phoneme converters in automatic speech recognition context
abstract
This paper deals with the evaluation of grapheme-to-phoneme (G2P) converters in a speech recognition context. The precision and recall rates are investigated as potential measures of the quality of the multiple generated pronunciation variants. Very different results are obtained whether or not we take into account the frequency of occurrence of the words. Since G2P systems are rarely evaluated on a speech recognition performance basis, the originality of this paper consists in using a speech recognition system to evaluate the G2P pronunciation variants. The results show that the training process is quite robust to some errors in the pronunciation lexicon, whereas pronunciation lexicon errors are harmful in the decoding process. Noticeable speech recognition performance improvements are achieved by combining two different G2P converters, one based on conditional random fields and the other on joint multigram models, as well as by checking the pronunciation variants of the most frequent words.
Denis Jouvet, Dominique Fohr, Irina Illina
ICASSP1
2012 Classification margin for improved class-based speech recognition performance
abstract
This paper investigates class-based speech recognition, and more precisely the impact of the selection of the training samples for each class on the final speech recognition performance. Increasing the number of recognition classes should lead to more specific models, and thus to better recognition performance, providing the trained model parameters are reliable. However, when the number of classes increases, the amount of training data for each class gets smaller, and may lead to unreliable parameters. The experiments described in the paper show that taking into account a classification margin tolerance helps associating more training data to each class, and improves the overall speech recognition performance.
Denis Jouvet, Nicolas Vinuesa
ICASSP1
2012 Class-based speech recognition using a maximum dissimilarity criterion and a tolerance classification margin
abstract
One of the difficult problems of Automatic Speech Recognition (ASR) is dealing with the acoustic signal variability. Much state-of-the-art research has demonstrated that splitting data into classes and using a model specific to each class provides better results. However, when the dataset is not large enough and the number of classes increases, there is less data for adapting the class models and the performance degrades. This work extends and combines previous research on un-supervised splits of datasets to build maximally separated classes and the introduction of a tolerance classification margin for a better training of the class model parameters. Experiments, carried out on the French radio broadcast ESTER2 data, show an improvement in recognition results compared to the ones obtained previously. Finally, we demonstrate that combining the decoding results from different class models leads to even more significant improvements.
Arseniy Gorin, Denis Jouvet
SLT2
2012 Combining criteria for the detection of incorrect entries of non-native speech in the context of foreign language learning
abstract
This article analyzes the detection of incorrect entries of non-native speech in the context of foreign language learning. The purpose is to detect and reject incorrect entries (i.e. those for which the speech signal does not correspond at all to the associated text) while being tolerant to the mispronunciations of non-native speech. The proposed approach exploits the comparison between two text-to-speech alignments : one constrained by the text which is being checked, with another one unconstrained, corresponding to a phonetic decoding. Several comparison criteria are described and combined via a logistic regression function. The article analyzes the influence of different settings, such as the impact of non-native pronunciation variants, the impact of learning the decision functions on native or on non-native speech, as well as the impact of combining various comparison criteria. The performance evaluations are conducted both on native and on non-native speech.
Luiza Orosanu, Denis Jouvet, Dominique Fohr, Irina Illina, Anne Bonneau
SLT2
2011 Grapheme-to-Phoneme Conversion Using Conditional Random Fields
abstract
International audience
Irina Illina, Dominique Fohr, Denis Jouvet
INTERSPEECH3
2011 About Handling Boundary Uncertainty in a Speaking Rate Dependent Modeling Approach
abstract
International audience
Denis Jouvet, Dominique Fohr, Irina Illina
INTERSPEECH1
2010 Detailed pronunciation variant modeling for speech transcription
abstract
International audience
Denis Jouvet, Dominique Fohr, Irina Illina
INTERSPEECH1
2008 Modeling inter-speaker variability in speech recognition
abstract
This paper details a method for taking into account variability influence in HMM-based speech recognition. The set of Gaussian components of the mixtures represents the entire acoustic space covered for all possible variability values. For each utterance to be recognized, the corresponding variability value is estimated and used to weight and/or constrain dynamically the acoustic space for each pdf. To do that, the weight coefficients of the Gaussian mixtures are set dependent on the variability value. As an example, the variability considered is the inter-speaker variability, and is handled through speaker classes. Taking into account for each utterance the four speaker classes that best match with the utterance signal leads to a significant word error rate reduction on a continuous speech recognition task, as compared to standard speaker-independent modeling.
Gwenael Cloarec, Denis Jouvet
ICASSP2
2007 On using units trained on foreign data for improved multiple accent speech recognition
Katarina Bartkova, Denis Jouvet
Speech Commun.2
2007 Automatic speech recognition and speech variability: A review
Mohamed Benzeghiba, Renato De Mori, Olivier Deroo, Stéphane Dupont, Teodora Erbes, Denis Jouvet, Luciano Fissore, Pietro Laface, Alfred Mertins, Christophe Ris, Richard Rose, Vivek Tyagi, Christian Wellekens
Speech Commun.6
2007 Introduction to the Special Issue on Intrinsic Speech Variations
Renato De Mori, Olivier Deroo, Stéphane Dupont, Denis Jouvet, Luciano Fissore, Pietro Laface, Alfred Mertins, Christian Wellekens
Speech Commun.4
2006 Using Multilingual Units for Improved Modeling of Pronunciation Variants
abstract
Standard speech modeling generally implies the combination of models of the phonemes of the current language with a description of possible pronunciation variants of the vocabulary words. When dealing with foreign accent, this standard native speech modeling is not adequate. In fact many variabilities have to be taken into account as the acoustic realization of the sounds by non-native speakers does not always match with native models and some phonemes may be replaced by others. By introducing models of phonemes estimated from speech data of other languages, and adding extra pronunciation variants through phonological rules, speech recognition performance improvements were achieved on non-native speech. In this study, a selection of the most frequently used variants is proposed, which relies on the frequency of usage of the various models associated to each phoneme on a development set. Although this selection process is rather simple it provides significant performance improvement
Katarina Bartkova, Denis Jouvet
ICASSP (5)2
2006 Automatic Speech Recognition and Intrinsic Speech Variation
abstract
This paper briefly reviews state of the art related to the topic of speech variability sources in automatic speech recognition systems. It focuses on some variations within the speech signal that make the ASR task difficult. The variations detailed in the paper are intrinsic to the speech and affect the different levels of the ASR processing chain. For different sources of speech variation, the paper summarizes the current knowledge and highlights specific feature extraction or modeling weaknesses and current trends
Mohamed Benzeghiba, Renato De Mori, Olivier Deroo, Stéphane Dupont, Teodora Erbes, Denis Jouvet, Luciano Fissore, Pietro Laface, Alfred Mertins, Christophe Ris, Richard Rose, Vivek Tyagi, Christian Wellekens
ICASSP (5)6
2004 Sequential clustering algorithm for Gaussian mixture initialization
abstract
A simple sequential algorithm for deriving initial values for Gaussian mixture parameters used in HMM-based speech recognition is presented. The proposed algorithm sequentially clusters the training frames, in the order in which they are available and according to the density to which they are associated. This frame-density association results from a frame-state alignment of the training data performed with a single-Gaussian model, which is good enough for such a force-alignment task. The models obtained with the proposed sequential clustering procedure provide good speech recognition performance when compared to models obtained with the usual Gaussian splitting procedure.
Ronaldo O. Messina, Denis Jouvet
ICASSP (1)2
2004 Context dependent "long units" for speech recognition
abstract
It is expected that longer-than-phoneme units such as syllables or multi-phone units can deal with sources of performance degradation, such as pronunciation variation or coarticulation, better than phoneme-sized units like triphones. The possible number of contextual realizations of those “long units” (LU) is very high, causing an explosion of the number of parameters to be estimated. As the training data are limited, the usual solution is to share parameters between different units to improve parameter estimation. Another problem is how to provide a model for a unit (syllable/multiphones) that was not present during training (unseen unit). In this paper we evaluate and compare syllable and automatically derived multi-phone units. We introduce a method called “contextual factorization” to share parameters between different models and we propose a figure of merit to decide which decomposition of an unseen syllable is the most appropriate. Performance is improved comparing to a triphone based system.
Denis Jouvet, Ronaldo O. Messina
INTERSPEECH1
2003 About improving recognition of spontaneously uttered French city-names
abstract
This paper deals with the recognition of French city-names over the telephone. This recognition task, critical in many applications, involves a 40,000 city-name vocabulary, ranging from short monosyllabic words to long official compound-names. Data collected from a field experiment are analyzed, and several ways of improving speech recognition performance are investigated. This includes a careful checking of the pronunciation lexicon, acceptation of shorter forms (common names), adaptation of the acoustic models and introduction of specific noise models as well as a few frequent words and expressions to facilitate out-of-vocabulary data rejection. Experiments show that all these techniques help improving the overall recognition performances and nicely combine together.
Denis Jouvet, Katarina Bartkova, Lionel Delphin-Poulat, Alexandre Ferrieux, Xavier Lamming, Jean Monné, Christophe Raix
ICASSP (1)1
2002 Prosodic parameter for speaker identification
Katarina Bartkova, David Le Gac, Delphine Charlet, Denis Jouvet
INTERSPEECH4
2002 Evaluation of a noise-robust DSR front-end on Aurora databases
Duncan Macho, Laurent Mauuary, Bernhard Noé, Yan Ming Cheng, Douglas Ealey, Denis Jouvet, Holly Kelleher, David Pearce 0002, Fabien Saadoun
INTERSPEECH6
2001 On combining recognizers for improved recognition of spelled names
abstract
This paper deals with the recognition of spelled names over the telephone. Two recognition approaches are recalled. One is based on a forward-backward algorithm in which the spelling lexicon is handled by the A* algorithm in the backward pass. The other is a 2-step approach, which relies on a discrete HMM-based retrieval procedure. Both approaches integrate a rejection test. Combinations of the two approaches are investigated in this paper. First, a sequential combination is presented. The 2-step approach is used only when the forward-backward approach does not yield an answer because of memory limitations. This sequential combination, evaluated on field data collected from a vocal directory service, takes the best of both approaches. Results are presented for the recognition of valid spelled names as well as for the rejection of incorrect data. Finally, a detailed analysis of the recognition results of the 2 approaches shows that a comparison of the 2 recognition results leads to an efficient reliability criterion.
Denis Jouvet, S. Droguet
ICASSP1
2001 On combining confidence measures for improved rejection of incorrect data
abstract
In this paper, techniques for combining confidence measures are proposed and evaluated. Confidence measures are useful for rejecting incorrect data, which is an important issue in speech recognition based interactive systems. Many ways of computing individual confidence measures have already been investigated. A detailed analysis of various confidence measures shows that they behave differently for what concerns rejection of incorrect data on various field data subsets (substitution errors, out-of-vocabulary data & noise tokens) collected from a vocal directory task. Two combination methods are then presented. One combines confidence measures by means of a neural network and the other through logistic regression. Evaluations shows that both combination techniques are efficient, and both take the best of the various individual confidence measures involved on each data subset.
Delphine Charlet, Guy Mercier, Denis Jouvet
INTERSPEECH3
2001 Noise reduction for noise robust feature extraction for distributed speech recognition
abstract
This paper describes the noise robust feature extraction meth ods developed by France Telecom and Alcatel for the noise robust front-end standardisation of ETSI Aurora.It is shown that both noise reduction methods give a substantial im provement when compared to a standard MFCC feature ex traction algorithm for speech recognition in noisy environ ments.In addition, blind equalisation and feature vector se lection were used for further improvement of recognition performance.Results are discussed for the ETSI Aurora 2 task and the SDC-Italian task as well.It was found that the combi nation of noise reduction with the proposed methods is capa ble to achieve around 50% reduction of the error rate.In the context of the open ETSI Aurora standardisation, two propos als were submitted based on these methods, they achieved the best results among all the proposals.
Bernhard Noé, Jürgen Sienel, Denis Jouvet, Laurent Mauuary, Johan de Veth, Lou Boves, Febe de Wet
INTERSPEECH3
2001 Feature vector selection to improve ASR robustness in noisy conditions
abstract
\n Contains fulltext :\n 75051.pdf (author's version ) (Open Access)\n
Johan de Veth, Laurent Mauuary, Bernhard Noé, Febe de Wet, Jürgen Sienel, Lou Boves, Denis Jouvet
INTERSPEECH7
2000 Detecting the end of spellings using statistics on recognized letter sequences for spelled names recognition
abstract
This paper addresses the problem of end-of-speech detection for continuously spelled names. In order to reduce errors due to the premature detection of the end-of-speech resulting from a hesitation or from a long pause between some letters, we propose to detect prefixes of names. In this case, the recognition system will wait for extra speech in order to obtain a complete spelled name. The recognizer must also deal with incorrect data. Consequently a ternary decision is made when checking a recognized hypothesis: complete spelled name, prefix of a name or incorrect data which is rejected. The decision is made using three sequences of letters that decode the speech input under different syntactical constraints and the associated scores. We evaluated the approach with speech data collected from a vocal server in operation. About 83% of the prefixes are detected correctly, 2% are confused with complete spelled names and 15% are rejected.
Stephan Hanel, Denis Jouvet
ICASSP2
2000 Confidence measure and incremental adaptation for the rejection of incorrect data
abstract
This paper deals with the problem of incorrect data rejection in a large vocabulary directory task. Two different strategies are investigated to improve the rejection of noises and OOV data. An incremental adaptation algorithm is first proposed to adapt word models and a garbage model to field data. The second method consists in post-processing the recogniser hypotheses by computing for each of them a confidence measure based on frame level likelihood ratios. Both methods yield a noticeable reduction in the false alarm rate on noises and OOV data. Their combination leads to a further false alarm rate reduction.
Nicolas Moreau, Delphine Charlet, Denis Jouvet
ICASSP3
2000 An alternative normalization scheme in HMM-based text-dependent speaker verification
Delphine Charlet, Denis Jouvet, O. Collin
Speech Commun.2
1999 Hypothesis dependent threshold setting for improved out-of-vocabulary data rejection
abstract
An efficient rejection procedure is necessary to reject out-of-vocabulary words and noise tokens that occur in voice activated vocal services. Garbage or filler models are very useful for such a task. However, a post-processing of the recognized hypothesis, based on a likelihood ratio statistic test, can refine the decision and improve performance. These tests can be applied either on acoustic parameters or on phonetic or prosodic parameters that are not taken into account by the HMM-based decoder. This paper focuses on the post-processing procedure and shows that making the likelihood ratio decision threshold dependent on the recognized hypothesis largely improves the efficiency of the rejection procedure. Models and anti-models are one of the key-points of such an approach. Their training and usage are also discussed, as well as the contextual modeling involved. Finally results are reported on a field database collected from a 2000-word directory task using various phonetic and prosodic parameters.
Denis Jouvet, Katarina Bartkova, Guy Mercier
ICASSP1
1999 Selective prosodic post-processing for improving recognition of French telephone numbers
abstract
This study describes a selective prosodic postprocessing procedure for improving the recognition of telephone numbers in French. The aim of the post-processing procedure is to recover recognition errors made by an HMM based ASR system. Instead of a global post-processing, this paper proposes a selective one. Post-processing is carried out only on some recognised numbers and only if its associated frequent confusion is also present in the N-best candidates. In such a case the discrimination between the solutions is carried out by checking the duration of a specific segment in a pertinent prosodic position. On the different data used in this study, about 23 % of the substitution errors are considered as being possible to recover with the selective duration post-processing and of this amount about 40 % of the errors are actually recovered. Key-words: speech recognition, post-processing, phone duration.
Katarina Bartkova, Denis Jouvet
EUROSPEECH2
1999 Recognition of spelled names over the telephone and rejection of data out of the spelling lexicon
abstract
This paper deals with the recognition of spelled names over the telephone. It introduces an efficient way of handling the spelling grammar, that is the lexicon of the allowed spelled names. The proposed approach is based on a forward-backward algorithm. The constraints on the sequences of letters are derived from the lexicon and are used by the A* algorithm in the backward pass. This forward-backward approach is compared to a 2-pass approach, which relies on a discrete HMM based retrieval procedure. The rejection of incorrect data is also investigated, based on the comparison of a lexicon constrained solution with an unconstrained decoding. The approaches are compared on field data collected from a vocal directory service. Results are presented for the recognition of valid spelled names and for the rejection of incorrect data (non-spelling and noise tokens and spellings not in the lexicon). The results show the efficiency of the proposed forward-backward procedure.
Denis Jouvet, Jean Monné
EUROSPEECH1
1999 Use of a confidence measure based on frame level likelihood ratios for the rejection of incorrect data
Nicolas Moreau, Denis Jouvet
EUROSPEECH2
1999 Derivation of the optimal set of phonetic transcriptions for a word from its acoustic realizations
Houda Mokbel, Denis Jouvet
Speech Commun.2
1998 An algorithm for maximum likelihood estimation of hidden Markov models with unknown state-tying
abstract
For speech recognition based on hidden Markov modeling, parameter-tying, which consists of constraining some of the parameters of the model to share the same value, has emerged as a standard practice. An original algorithm is proposed that makes it possible to jointly estimate both the shared model parameters and the tying characteristics, using the maximum likelihood criterion. The proposed algorithm is based on a previously introduced extension of the classic expectation-maximization (EM) framework. The convergence properties of this class of algorithms are analyzed in detail. The method is evaluated on an isolated word recognition task using hidden Markov models (HMMs) with Gaussian observation densities and tying at the state level. Finally, the extension of this method to the case of mixture observation densities with tying at the mixture component level is discussed.
Olivier Cappé, Chafic Mokbel, Denis Jouvet, Eric Moulines
IEEE Trans. Speech Audio Process.3
1997 Adapting PSN recognition models to the GSM environment by using spectral transformation
abstract
In this work, environment adaptation is studied in order to transform PSN speaker independent isolated words HMM to the GSM environment. Linear multiple regression (LMR) transformations associated with groups of HMM densities are used to adapt the densities. Both mean vectors and covariance matrices of the densities are adapted. It has been shown that a small amount of GSM data are sufficient to transform the PSN HMM in order to match the GSM environment and to achieve a performance equivalent to those of an HMM trained with a large amount of GSM data. The number of groups of Gaussian densities seems to have a small influence on the results. However, the minimum number of groups depends on the vocabulary size. Finally, this technique is compared to the Bayesian adaptation and the results show that similar performance can be obtained with both methods.
Thierry Soulas, Chafic Mokbel, Denis Jouvet, Jean Monné
ICASSP3
1997 Usefulness of phonetic parameters in a rejection procedure of an HMM-based speech recognition system
Katarina Bartkova, Denis Jouvet
EUROSPEECH2
1997 Design and analysis of a German telephone speech database for phoneme based training
abstract
Based on the Sotscheck text corpus, we developped a new corpus that was specifically optimised for training phoneme-based recognition systems. Particular attention was payed on good coverage of phone transitions. Even though the resulting corpus is only slightly enlarged, it shows an increased phonetic coverage while maintaining a good phonetic balance. Results of phonetic statistical analysis and of experiments for training an allophonebased recognizer are reported here.
Stefan Feldes, Bernhard Kaspar, Denis Jouvet
EUROSPEECH3
1997 Automatic derivation of multiple variants of phonetic transcriptions from acoustic signals
abstract
This paper deals with two methods for automatically finding multiple phonetic transcriptions of words, given sample utterances of the words and an inventory of context-dependent subword units. The two approaches investigated are based on an analysis of theN -best phonetic decoding of the available utterances. In the set of transcriptions resulting from theN -best decoding of all the utterances, the first method selects theK most frequent variants (Frequency Criterion) , while the second method selects the K most likely ones (Maximum Likelihood Criterion). Experiments carried out on speaker-independent recognition showed that the performance obtained with the ”Maximum Likelihood Criterion” is not much different from that obtained with manual transcriptions. In the case of speaker-dependent speech recognition, the estimate of the 3 most likely transcription variants of each word, yields promising results.
Houda Mokbel, Denis Jouvet
EUROSPEECH2
1997 Optimizing feature set for speaker verification
Delphine Charlet, Denis Jouvet
Pattern Recognit. Lett.2
1997 Towards improving ASR robustness for PSN and GSM telephone applications
Chafic Mokbel, Laurent Mauuary, Lamia Karray, Denis Jouvet, Jean Monné, Jacques Simonin, Katarina Bartkova
Speech Commun.4
1996 Bayesian adaptation of speech recognizers to field speech data
C. G. Miglietta, Chafic Mokbel, Denis Jouvet, Jean Monné
ICSLP3
1996 Parameter tying for flexible speech recognition
Jacques Simonin, S. Bodin, Denis Jouvet, Katarina Bartkova
ICSLP3
1996 Deconvolution of telephone line effects for speech recognition
Chafic Mokbel, Denis Jouvet, Jean Monné
Speech Commun.2
1995 On using a priori segmentation of the speech signal in an N-best solutions post-processing
abstract
This paper proposes a new approach to the incorporation of automatic a-priori segmentation into an HMM based speech recognizer. The approach used for the post-processing of N-best solutions is based on stochastic modelling of the number of speech signal stationarity changes which occur within the phonetic segments of each solution. The objective of this post-processing is to validate the presence of stationarity zones in the speech signal. This particular validation cannot be exploited using a centisecond approach. The signal stationarity changes are detected using an "a priori" segmentation algorithm. Two phonetic models are calculated for each phonetic segment. One corresponds to correct solutions and the other one corresponds to incorrect solutions. These two models are used simultaneously in order to compute a post-processing score for each solution. In the initial set of experiments, which was conducted on telephone databases, the use of this method resulted in a 9% error rate reduction on the "Number" database, and a 15% error rate reduction on the "Digit" database.
Thierry Moudenc, Denis Jouvet, Jean Monné
ICASSP2
1995 Error analysis on field data and improved garbage HMM modelling
Katarina Bartkova, Dominique Dubois, Denis Jouvet, Jean Monné
EUROSPEECH3
1995 Blind equalization using adaptive filtering for improving speech recognition over telephone
Chafic Mokbel, Denis Jouvet, Jean Monné
EUROSPEECH2
1995 Improving recognition performances on field data with an a-priori segmentation of the speech signal
Thierry Moudenc, Denis Jouvet, Jean Monné
EUROSPEECH2
1995 Operational and experimental French telecommunication services using CNET speech recognition and text-to-speech synthesis
Christel Sorin, Denis Jouvet, Christian Gagnoulet, Dominique Dubois, D. Sadek, M. Toularhoat
Speech Commun.2
1994 Structure of allophonic models and reliable estimation of the contextual parameters
Denis Jouvet, Katarina Bartkova, A. Stouff
ICSLP1
1994 Compensation of telephone line effects for robust speech recognition
Chafic Mokbel, R. Paches-Leal, Denis Jouvet, Jean Monné
ICSLP3
1993 Speaker-independent spelling recognition over the telephone
Denis Jouvet, A. Laine, Jean Monné, Christian Gagnoulet
ICASSP (2)1
1993 Application of the n-best solutions algorithm to speaker-independent spelling recognition over the telephone
Denis Jouvet, M. N. Lokbani, Jean Monné
EUROSPEECH1
1993 Segmental post-processing of the n-best solutions in a speech recognition system
M. N. Lokbani, Denis Jouvet, Jean Monné
EUROSPEECH2
1993 On-line adaptation of a speech recognizer to variations in telephone line conditions
Chafic Mokbel, Jean Monné, Denis Jouvet
EUROSPEECH3
1991 On the modelization of allophones in an HMM based speech recognition system
Denis Jouvet, Katarina Bartkova, Jean Monné
EUROSPEECH1
1991 Automatic adjustments of the structure of Markov models for speech recognition applications
Denis Jouvet, Laurent Mauuary, Jean Monné
EUROSPEECH1
1991 MAIRIEVOX: A voice-activated information system
Christian Gagnoulet, Denis Jouvet, J. Damay
Speech Commun.2
1989 An acoustic-phonetic decoder an automatic segmentation algorithm
V. Le Maire, Régine André-Obrecht, Denis Jouvet
EUROSPEECH3
1986 A new network-based speaker-independent connected-word recognition system
abstract
This paper describes a network-based speaker-independent connected-word recognition system based on Markov modelling and compares the results obtained using different basic-units (words, phonemes, diphones). The whole knowledge of the application, including syntactical and lexical descriptions and also phonological rules, is used to create a single integrated network; and the probabilistic density functions of the Markov chain are Gaussian multivariate with diagonal matrix. The recognition tests were performed on numbers from 0 up to 999 in French, recorded from 26 speakers. For whole-word basic-units, the recognition performances (percentage of numbers correctly recognized) increased up to 90% when the number of pdf per word increased up to 15. The same recognition rate was obtained with phoneme basic-units having only 5 pdf each, thus saving computations, and more than 95% was achieved using sub-phonemic description.
Denis Jouvet, Jean Monné, Dominique Dubois
ICASSP1
1984 One-pass syntax-directed connected-word recognition in a time-sharing environment
abstract
This paper describes a syntax-directed real-time connected-word recognition system and compares the performance of different recognition algorithms. This is a speaker dependent system based on whole word template matching, and the recognition is done using a one-pass dynamic time-warping algorithm. The symmetric local constraints of Sakoe and Chiba are used for connected word recognition, and with proper normalization of the accumulated distances, result in better performance than the Itakura local constraints. By allowing a noise template to start and end the sentences during the recognition process, most of the errors due to a bad endpoint detection are eliminated. When memory and computational resources are limited, vector quantization allows more templates for each word in the dictionary, and as a consequence, the recognition performance increases. By carefully implementing the algorithm on standard hardware (VAX 11/780 and FPS AP-120B), real-time recognition is achieved for vocabularies of up to 100 templates, or up to 250 templates if vector quantization is used.
Denis Jouvet, Richard M. Schwartz
ICASSP1