EDBT 2026 Demo / reviewers in the wild / expert
Mickael Rouvier
dblp:02/8760 · also Mickaël Rouvier
· DBLP profile ↗
46ranked-venue papers
14as first author
18since 2021 · last 2025
0000-0003-3541-3385ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 10 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 12 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Benchmark of French ASR Systems Based on Error SeverityabstractAutomatic Speech Recognition (ASR) transcription errors are commonly assessed using metrics that compare them with a reference transcription, such as Word Error Rate (WER), which measures spelling deviations from the reference, or semantic score-based metrics. However, these approaches often overlook what is understandable to humans when interpreting transcription errors. To address this limitation, a new evaluation is proposed that categorizes errors into four levels of severity, further divided into subtypes, based on objective linguistic criteria, contextual patterns, and the use of content words as the unit of analysis. This metric is applied to a benchmark of 10 state-of-the-art ASR systems on French language, encompassing both HMM-based and end-to-end models. Our findings reveal the strengths and weaknesses of each system, identifying those that provide the most comfortable reading experience for users. Antoine Tholly, Jane Wottawa, Mickael Rouvier, Richard Dufour |
COLING | 3 |
| 2024 | How Important Is Tokenization in French Medical Masked Language Models?abstractSubword tokenization has become the prevailing standard in the field of natural language processing (NLP) over recent years, primarily due to the widespread utilization of pre-trained language models. This shift began with Byte-Pair Encoding (BPE) and was later followed by the adoption of SentencePiece and WordPiece. While subword tokenization consistently outperforms character and word-level tokenization, the precise factors contributing to its success remain unclear. Key aspects such as the optimal segmentation granularity for diverse tasks and languages, the influence of data sources on tokenizers, and the role of morphological information in Indo-European languages remain insufficiently explored. This is particularly pertinent for biomedical terminology, characterized by specific rules governing morpheme combinations. Despite the agglutinative nature of biomedical terminology, existing language models do not explicitly incorporate this knowledge, leading to inconsistent tokenization strategies for common terms. In this paper, we seek to delve into the complexities of subword tokenization in French biomedical domain across a variety of NLP tasks and pinpoint areas where further enhancements can be made. We analyze classical tokenization algorithms, including BPE and SentencePiece, and introduce an original tokenization strategy that integrates morpheme-enriched word segmentation into existing tokenization methods. Yanis Labrak, Adrien Bazoge, Béatrice Daille, Mickael Rouvier, Richard Dufour |
LREC/COLING | 4 |
| 2024 | DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical DomainabstractThe biomedical domain has sparked a significant interest in the field of Natural Language Processing (NLP), which has seen substantial advancements with pre-trained language models (PLMs). However, comparing these models has proven challenging due to variations in evaluation protocols across different models. A fair solution is to aggregate diverse downstream tasks into a benchmark, allowing for the assessment of intrinsic PLMs qualities from various perspectives. Although still limited to few languages, this initiative has been undertaken in the biomedical field, notably English and Chinese. This limitation hampers the evaluation of the latest French biomedical models, as they are either assessed on a minimal number of tasks with non-standardized protocols or evaluated using general downstream tasks. To bridge this research gap and account for the unique sensitivities of French, we present the first-ever publicly available French biomedical language understanding benchmark called DrBenchmark. It encompasses 20 diversified tasks, including named-entity recognition, part-of-speech tagging, question-answering, semantic textual similarity, or classification. We evaluate 8 state-of-the-art pre-trained masked language models (MLMs) on general and biomedical-specific data, as well as English specific MLMs to assess their cross-lingual capabilities. Our experiments reveal that no single model excels across all tasks, while generalist models are sometimes still competitive. Yanis Labrak, Adrien Bazoge, Oumaima El Khettari, Mickael Rouvier, Pacôme Constant dit Beaufils, Natalia Grabar, Béatrice Daille, Solen Quiniou, Emmanuel Morin, Pierre-Antoine Gourraud, Richard Dufour |
LREC/COLING | 4 |
| 2024 | A Zero-shot and Few-shot Study of Instruction-Finetuned Large Language Models Applied to Clinical and Biomedical TasksabstractThe recent emergence of Large Language Models (LLMs) has enabled significant advances in the field of Natural Language Processing (NLP). While these new models have demonstrated superior performance on various tasks, their application and potential are still underexplored, both in terms of the diversity of tasks they can handle and their domain of application. In this context, we evaluate four state-of-the-art instruction-tuned LLMs (ChatGPT, Flan-T5 UL2, Tk-Instruct, and Alpaca) on a set of 13 real-world clinical and biomedical NLP tasks in English, including named-entity recognition (NER), question-answering (QA), relation extraction (RE), and more. Our overall results show that these evaluated LLMs approach the performance of state-of-the-art models in zero- and few-shot scenarios for most tasks, particularly excelling in the QA task, even though they have never encountered examples from these tasks before. However, we also observe that the classification and RE tasks fall short of the performance achievable with specifically trained models designed for the medical field, such as PubMedBERT. Finally, we note that no single LLM outperforms all others across all studied tasks, with some models proving more suitable for certain tasks than others. Yanis Labrak, Mickael Rouvier, Richard Dufour |
LREC/COLING | 2 |
| 2024 | RoboVox: A Single/Multi-channel Far-field Speaker Recognition Benchmark for a Mobile RobotabstractIn this paper, we introduce a new far-field speaker recognition benchmark called RoboVox. RoboVox is a French corpus recorded by a mobile robot. The files are recorded from different distances under severe acoustical conditions with the presence of several types of noise and reverberation. In addition to noise and reverberation, the robot’s internal noise acts as an extra additive noise. RoboVox can be used for both single-channel and multi-channel speaker recognition. In the evaluation protocols, we are considering both cases. The obtained results demonstrate a significant decline in performance in far-filed speaker recognition and urge the community to further research in this domain Mohammad MohammadAmini, Driss Matrouf, Mickael Rouvier, Jean-François Bonastre, Romain Serizel, Théophile Gonos |
LREC/COLING | 3 |
| 2024 | Synvox2: Towards A Privacy-Friendly Voxceleb2 DatasetabstractThe success of deep learning in speaker recognition relies heavily on the use of large datasets. However, the data-hungry nature of deep learning methods has already being questioned on account the ethical, privacy, and legal concerns that arise when using large-scale datasets of natural speech collected from real human speakers. For example, the widely-used VoxCeleb2 dataset for speaker recognition is no longer accessible from the official website. To mitigate these concerns, this work presents an initiative to generate a privacyfriendly synthetic VoxCeleb2 dataset that ensures the quality of the generated speech in terms of privacy, utility, and fairness. We also discuss the challenges of using synthetic data for the downstream task of speaker verification. Xiaoxiao Miao, Xin Wang 0037, Erica Cooper, Junichi Yamagishi, Nicholas W. D. Evans, Massimiliano Todisco, Jean-François Bonastre, Mickael Rouvier |
ICASSP | 8 |
| 2024 | Zero-Shot End-To-End Spoken Question Answering In Medical Domain
Yanis Labrak, Adel Moumen, Richard Dufour, Mickael Rouvier |
INTERSPEECH | 4 |
| 2024 | LeBenchmark 2.0: A standardized, replicable and enhanced framework for self-supervised representations of French speech
Titouan Parcollet, Solène Evain, Marcely Zanon Boito, Adrien Pupier, Salima Mdhaffar, Hang Le 0001, Sina Alisamir, Natalia A. Tomashenko, Marco Dinarelli, Shucong Zhang, Alexandre Allauzen, Maximin Coavoux, Yannick Estève, Mickael Rouvier, Jérôme Goulian, Benjamin Lecouteux, François Portet, Solange Rossato, Fabien Ringeval, Didier Schwab, Laurent Besacier |
Comput. Speech Lang. | 15 |
| 2024 | Open-Source Conversational AI with SpeechBrain 1.0abstractSpeechBrain is an open-source Conversational AI toolkit based on PyTorch, focused particularly on speech processing tasks such as speech recognition, speech enhancement, speaker recognition, text-to-speech, and much more. It promotes transparency and replicability by releasing both the pre-trained models and the complete recipes of code and algorithms required for training them. This paper presents SpeechBrain 1.0, a significant milestone in the evolution of the toolkit, which now has over 200 recipes for speech, audio, and language processing tasks, and more than 100 models available on Hugging Face. SpeechBrain 1.0 introduces new technologies to support diverse learning modalities, Large Language Model (LLM) integration, and advanced decoding strategies, along with novel models, tasks, and modalities. It also includes a new benchmark repository, offering researchers a unified platform for evaluating models across diverse tasks. Mirco Ravanelli, Titouan Parcollet, Adel Moumen, Sylvain de Langen, Cem Subakan, Peter Plantinga, Yingzhi Wang 0002, Pooneh Mousavi, Luca Della Libera, Artem Ploujnikov, Francesco Paissan, Davide Borra, Mohamed Salah Zaïem, Zeyu Zhao 0004, Shucong Zhang, Georgios Karakasidis, Sung-Lin Yeh, Pierre Champion, Aku Rouhe, Rudolf Braun, Florian Mai, Juan Zuluaga-Gomez, Seyed Mahed Mousavi, Andreas Nautsch, Xuechen Liu 0001, Sangeet Sagar, Jarod Duret, Salima Mdhaffar, Gaëlle Laperrière, Mickael Rouvier, Renato De Mori, Yannick Estève |
J. Mach. Learn. Res. | 31 |
| 2023 | DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domainsabstractYanis Labrak, Adrien Bazoge, Richard Dufour, Mickael Rouvier, Emmanuel Morin, Béatrice Daille, Pierre-Antoine Gourraud. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yanis Labrak, Adrien Bazoge, Richard Dufour, Mickael Rouvier, Emmanuel Morin, Béatrice Daille, Pierre-Antoine Gourraud |
ACL (1) | 4 |
| 2023 | Jeffreys Divergence-Based Regularization of Neural Network Output Distribution Applied to Speaker RecognitionabstractA new loss function for speaker recognition with deep neural network is proposed, based on Jeffreys Divergence. Adding this divergence to the cross-entropy loss function allows to maximize the target value of the output distribution while smoothing the non-target values. This objective function provides highly discriminative features. Beyond this effect, we propose a theoretical justification of its effectiveness and try to understand how this loss function affects the model, in particular the impact on dataset types (i.e. in-domain or out-of-domain w.r.t the training corpus). Our experiments show that Jeffreys loss consistently outperforms the state-of-the-art for speaker recognition, especially on out-of-domain data, and helps limit false alarms. Pierre-Michel Bousquet, Mickael Rouvier |
ICASSP | 2 |
| 2023 | Improving training datasets for resource-constrained speaker recognition neural networksabstractInternational audience Pierre-Michel Bousquet, Mickael Rouvier |
INTERSPEECH | 2 |
| 2022 | Reliability criterion based on learning-phase entropy for speaker recognition with neural networkabstractInternational audience Pierre-Michel Bousquet, Mickael Rouvier, Jean-François Bonastre |
INTERSPEECH | 2 |
| 2022 | Qualitative Evaluation of Language Model Rescoring in Automatic Speech RecognitionabstractInternational audience Thibault Bañeras-Roux, Mickael Rouvier, Jane Wottawa, Richard Dufour |
INTERSPEECH | 2 |
| 2022 | Speech Resources in the Tamasheq LanguageabstractIn this paper we present two datasets for Tamasheq, a developing language mainly spoken in Mali and Niger. These two datasets were made available for the IWSLT 2022 low-resource speech translation track, and they consist of collections of radio recordings from daily broadcast news in Niger (Studio Kalangou) and Mali (Studio Tamani). We share (i) a massive amount of unlabeled audio data (671 hours) in five languages: French from Niger, Fulfulde, Hausa, Tamasheq and Zarma, and (ii) a smaller 17 hours parallel corpus of audio recordings in Tamasheq, with utterance-level translations in the French language. All this data is shared under the Creative Commons BY-NC-ND 3.0 license. We hope these resources will inspire the speech community to develop and benchmark models using the Tamasheq language. Marcely Zanon Boito, Fethi Bougares, Florentin Barbier, Souhir Gahbiche-Braham, Loïc Barrault, Mickael Rouvier, Yannick Estève |
LREC | 6 |
| 2022 | Far-Field Speaker Recognition Benchmark Derived From The DiPCo CorpusabstractIn this paper, we present a far-field speaker verification benchmark derived from the publicly-available DiPCo corpus. This corpus comprise three different tasks that involve enrollment and test conditions with single- and/or multi-channels recordings. The main goal of this corpus is to foster research in far-field and multi-channel text-independent speaker verification. Also, it can be used for other speaker recognition tasks such as dereverberation, denoising and speech enhancement. In addition, we release a Kaldi and SpeechBrain system to facilitate further research. And we validate the evaluation design with a single-microphone state-of-the-art speaker recognition system (i.e. ResNet-101). The results show that the proposed tasks are very challenging. And we hope these resources will inspire the speech community to develop new methods and systems for this challenging domain. Mickael Rouvier, Mohammad MohammadAmini |
LREC | 1 |
| 2022 | On the Use of Semantically-Aligned Speech Representations for Spoken Language UnderstandingabstractIn this paper we examine the use of semantically-aligned speech representations for end-to-end spoken language understanding (SLU). We employ the recently-introduced SAMU-XLSR model, which is designed to generate a single embedding that captures the semantics at the utterance level, semantically aligned across different languages. This model combines the acoustic frame-level speech representation learning model (XLS-R) with the Language Agnostic BERT Sentence Embedding (LaBSE) model. We show that the use of the SAMU-XLSR model instead of the initial XLS-R model improves significantly the performance in the framework of end-to-end SLU. Finally, we present the benefits of using this model towards language portability in SLU. Gaëlle Laperrière, Valentin Pelloin, Mickael Rouvier, Themos Stafylakis, Yannick Estève |
SLT | 3 |
| 2021 | Studying Squeeze-and-Excitation Used in CNN for Speaker VerificationabstractIn speaker verification, the extraction of voice representations is mainly based on the Residual Neural Network (ResNet) architecture. ResNet is built upon convolution layers which learn filters to capture local spatial patterns along all the input, then generate feature maps that jointly encode the spatial and channel information. Unfortunately, all feature maps in a convolution layer are learnt independently (the convolution layer does not exploit the dependencies between feature maps) and locally. This problem has first been tackled in image processing. A channel attention mechanism, called squeeze-and-excitation (SE), has recently been proposed in convolution layers and applied to speaker verification. This mechanism re-weights the information extracted across features maps. In this paper, we first propose an original qualitative study about the influence and the role of the SE mechanism applied to the speaker verification task at different stages of the ResNet, and then evaluate several SE architectures. We finally propose to improve the SE approach with a new pooling variant based on the concatenation of mean- and standard-deviation-pooling. Results showed that applying SE only on the first stages of the ResNet allows to better capture speaker information for the verification task, and that significant discrimination gains on Voxceleb1-E, Voxceleb1-H and SITW evaluation tasks have been noted using the proposed pooling variant. Mickael Rouvier, Pierre-Michel Bousquet |
ASRU | 1 |
| 2019 | On Robustness of Unsupervised Domain Adaptation for Speaker RecognitionabstractInternational audience Pierre-Michel Bousquet, Mickael Rouvier |
INTERSPEECH | 2 |
| 2019 | I4U Submission to NIST SRE 2018: Leveraging from a Decade of Shared ExperiencesabstractThe I4U consortium was established to facilitate a joint entry to NIST speaker recognition evaluations (SRE). The latest edition of such joint submission was in SRE 2018, in which the I4U submission was among the best-performing systems. SRE'18 also marks the 10-year anniversary of I4U consortium into NIST SRE series of evaluation. The primary objective of the current paper is to summarize the results and lessons learned based on the twelve sub-systems and their fusion submitted to SRE'18. It is also our intention to present a shared view on the advancements, progresses, and major paradigm shifts that we have witnessed as an SRE participant in the past decade from SRE'08 to SRE'18. In this regard, we have seen, among others, a paradigm shift from supervector representation to deep speaker embedding, and a switch of research challenge from channel compensation to domain adaptation. Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Hitoshi Yamamoto, Koji Okabe, Ville Vestman, Jing Huang 0019, Guo-Hong Ding, Hanwu Sun, Anthony Larcher, Rohan Kumar Das, Haizhou Li 0001, Mickael Rouvier, Pierre-Michel Bousquet, Wei Rao 0002, Qing Wang 0039, Fahimeh Bahmaninezhad, Héctor Delgado, Massimiliano Todisco |
INTERSPEECH | 13 |
| 2017 | Duration Mismatch Compensation Using Four-Covariance Model and Deep Neural Network for Speaker VerificationabstractInternational audience Pierre-Michel Bousquet, Mickael Rouvier |
INTERSPEECH | 2 |
| 2017 | Acoustic Pairing of Original and Dubbed Voices in the Context of Video Game LocalizationabstractInternational audience Adrien Gresse, Mickael Rouvier, Richard Dufour, Vincent Labatut, Jean-François Bonastre |
INTERSPEECH | 2 |
| 2017 | The I4U Mega Fusion and Collaboration for NIST Speaker Recognition Evaluation 2016abstract18th Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, Stockholm, Sweden, 20-24 August 2017 Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Anthony Larcher, Andreas Nautsch, Themos Stafylakis, Gang Liu 0001, Mickael Rouvier, Wei Rao 0002, Federico Alegre, Man-Wai Mak, Achintya Kumar Sarkar, Héctor Delgado, Rahim Saeidi, Hagai Aronowitz, Aleksandr Sizov, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Bin Ma 0001, Ville Vestman, Md. Sahidullah, M. Halonen, Anssi Kanervisto, Gaël Le Lan, Fahimeh Bahmaninezhad, Sergey Isadskiy, Christian Rathgeb, Christoph Busch 0001, Georgios Tzimiropoulos, Q. Qian, Q. Zhao, J. Xue, R. Jin, T. Zhao, Pierre-Michel Bousquet, Moez Ajili, Waad Ben Kheder, Driss Matrouf, Zhi Hao Lim, Chenglin Xu, Haihua Xu 0001, Chng Eng Siong, Benoit G. B. Fauve, Kaavya Sriskandaraja, Vidhyasaharan Sethu, W. W. Lin, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Massimiliano Todisco, Nicholas W. D. Evans, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Eliathamby Ambikairajah |
INTERSPEECH | 9 |
| 2016 | Investigation of speaker embeddings for cross-show speaker diarizationabstractThis paper proposes to investigate speaker embeddings, a representation extracted from hidden layers of deep neural networks trained on a speaker identification task, on cross-show diarization. The new representation brings an improvement over i-vectors, and we show that while shallow hidden layers give best results on the single-show condition, deeper layers yield better performance on cross-show diarization. This confirms that deep representations model higher level features which help generalizing to different acoustic conditions. Experiments, conducted on the French corpus of REPERE, show that the deep speaker embeddings technique decreases DER by 0.82 points. Mickael Rouvier, Benoît Favre |
ICASSP | 1 |
| 2015 | Multimodal embedding fusion for robust speaker role recognition in video broadcastabstractPerson role recognition in video broadcasts consists in classifying people into roles such as anchor, journalist, guest, etc. Existing approaches mostly consider one modality, either audio (speaker role recognition) or image (shot role recognition), firstly because of the non-synchrony between both modalities, and secondly because of the lack of a video corpus annotated in both modalities. Deep Neural Networks (DNN) approaches offer the ability to learn simultaneously feature representations (embeddings) and classification functions. This paper presents a multimodal fusion of audio, text and image embeddings spaces for speaker role recognition in asynchronous data. Monomodal embeddings are trained on exogenous data and fine-tuned using a DNN on 70 hours of French Broadcasts corpus for the target task. Experiments on the REPERE corpus show the benefit of the embeddings level fusion compared to the monomodal embeddings systems and to the standard late fusion method. Mickael Rouvier, Sebastien Delecraz, Benoît Favre, Meriem Bendris, Frédéric Béchet |
ASRU | 1 |
| 2015 | "speech is silver, but silence is golden": improving speech-to-speech translation performance by slashing users inputabstractSpeech-to-speech translation is a challenging task mixing two of the most ambitious Natural Language Processing challenges: Machine Translation (MT) and Automatic Speech Recognition (ASR). Recent advances in both fields have led to operational systems achieving good performance when used in matching conditions with those of ASR and MT models training. Regardless of the quality of these models, errors are inevitable due to some technical limitations of the systems (e.g. closed vocabulary) and intrinsic ambiguities of spoken languages. However all ASR and MT errors don’t have the same impact on the usability of a given speech-to-speech dialog system: some can be very benign, unconsciously corrected by users, some can damage the understanding between users and eventually lead the dialog to a failure. We present in this paper a strategy focusing on ASR error segments that have a high negative impact on MT performance. We propose a method that consists firstly in automatically detecting these erroneous segments then secondly estimating their impact on MT. We show that removing such segments prior to translation can lead to a significant decrease in translation error rate, even without any correction strategy. Frédéric Béchet, Benoît Favre, Mickael Rouvier |
INTERSPEECH | 3 |
| 2015 | Audio-Based Video Genre IdentificationabstractThis paper presents investigations about the automatic identification of video genre by audio channel analysis. Genre refers to editorial styles such commercials, movies, sports... We propose and evaluate some methods based on both low and high level descriptors, in cepstral or time domains, but also by analyzing the global structure of the document and the linguistic contents. Then, the proposed features are combined and their complementarity is evaluated. On a database composed of single-stories web-videos, the best audio-only based system performs 9% of Classification Error Rate (CER). Finally, we evaluate the complementarity of the proposed audio features and video features that are classically used for Video Genre Identification (VGI). Results demonstrate the complementarity of the modalities for genre recognition, the final audio-video system reaching 6% CER. Mickael Rouvier, Stanislas Oger, Georges Linarès, Driss Matrouf, Bernard Mérialdo, Yingbo Li |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | Reranked aligners for interactive transcript correctionabstractClarification dialogs can help address ASR errors in speech-to-speech translation systems and other interactive applications. We propose to use variants of Levenshtein alignment for merging an er-rorful utterance with a targeted rephrase of an error segment. ASR errors that might harm the alignment are addressed through phonetic matching, and a word embedding distance is used to account for the use of synonyms outside targeted segments. These features lead to a relative improvement of 30% of word error rate on sentences with ASR errors compared to not performing the clarification. Twice as many utterances are completely corrected compared to using basic word alignment. Furthermore, we generate a set of potential merges and train a neural network on crowd-sourced rephrases in order to select the best merger, leading to 24% more instances completely corrected. The system is deployed in the framework of the BOLT project. Benoît Favre, Mickael Rouvier, Frédéric Béchet |
ICASSP | 2 |
| 2014 | Multimodal understanding for person recognition in video broadcastsabstractInternational audience Frédéric Béchet, Meriem Bendris, Delphine Charlet, Géraldine Damnati, Benoît Favre, Mickael Rouvier, Rémi Auguste, Benjamin Bigot, Richard Dufour, Corinne Fredouille, Georges Linarès, Jean Martinet, Grégory Senay, Pierre Tirilly |
INTERSPEECH | 6 |
| 2014 | Speaker adaptation of DNN-based ASR with i-vectors: does it actually adapt models to speakers?abstractDeep neural networks (DNN) are currently very successful for acoustic modeling in ASR systems. One of the main challenges with DNNs is unsupervised speaker adaptation from an initial speaker clustering, because DNNs have a very large number of parameters. Recently, a method has been proposed to adapt DNNs to speakers by combining speaker-specific information (in the form of i-vectors computed at the speaker-cluster level) with fMLLR-transformed acoustic features. In this paper we try to gain insight on what kind of adaptation is performed on DNNs when stacking i-vectors with acoustic features and what information exactly is carried by i-vectors. We observe on REPERE corpus that DNNs trained on i-vector features concatenated with fMLLR-transformed acoustic features lead to a gain of 0.7 points. The experiments shows that using ivector stacking in DNN acoustic models is not only performing speaker adaptation, but also adaptation to acoustic conditions. Mickael Rouvier, Benoît Favre |
INTERSPEECH | 1 |
| 2014 | Joint decoding of complementary utterancesabstractErrors in open-domain ASR can be corrected by asking the speaker to rephrase targeted segments in utterances where they have been detected. The utterance merging problem consists in generating a better transcript from the utterance where errors have been detected and a clarification utterance. We introduce an alignment-decoding algorithm for jointly processing the two utterances and benefit from the complementary information they contain. The algorithm aligns word lattices in the WFST framework with a probabilistic cost model. Results on the BOLT-BC speech-to-speech translation task show an improvement of 2.84 points of accuracy compared to aligning the one best without joint decoding. Mickael Rouvier, Benoît Favre, Frédéric Béchet |
SLT | 1 |
| 2013 | An open-source state-of-the-art toolbox for broadcast news diarizationabstractInternational audience Mickael Rouvier, Grégor Dupuy, Paul Gay, Elie Khoury 0001, Téva Merlin, Sylvain Meignier |
INTERSPEECH | 1 |
| 2012 | Low latency combination of parallelized single-pass LVCSR systemsabstractInternational audience Fethi Bougares, Mickael Rouvier, Yannick Estève, Georges Linarès |
INTERSPEECH | 2 |
| 2012 | Subspace Gaussian Mixture Models Based on Noise Compensation for Speech Recognition
Mohamed Bouallegue, Driss Matrouf, Georges Linarès, Mickael Rouvier |
INTERSPEECH | 4 |
| 2012 | I-vectors and ILP clustering adapted to cross-show speaker diarizationabstractInternational audience Grégor Dupuy, Mickael Rouvier, Sylvain Meignier, Yannick Estève |
INTERSPEECH | 2 |
| 2011 | Subspace Gaussian Mixture Models for vectorial HMM-states representationabstractIn this paper we present a vectorial representation of the HMM states that is inspired by the Subspace Gaussian Mixture Models paradigm (SGMM). This vectorial representation of states will make possible a large number of applications, such as HMM-states clustering and graphical visualization. Thanks to this representation, the Hidden Markov Model (HMM) states can be seen as sets of points in multi-dimensional space and then can be studied using statistical data analysis techniques. In this paper, we show how this representation can be obtained and used for tying states of an HHM-based automatic speech recognition system without any use of linguistic or phonetic knowledge. In experiments, this approach achieves significant and stable gain, while conserving the classical approach based on decision trees. We also show how it can be used for graphical visualization, which can be useful in other domains like phonetics or clinical phonetics. Mohamed Bouallegue, Driss Matrouf, Mickael Rouvier, Georges Linarès |
ASRU | 3 |
| 2011 | Factor analysis based session variability compensation for Automatic Speech RecognitionabstractIn this paper we propose a new feature normalization based on Factor Analysis (FA) for the problem of acoustic variability in Automatic Speech Recognition (ASR). The FA paradigm was previously used in the field of ASR, in order to model the usefull information: the HMM state dependent acoustic information. In this paper, we propose to use the FA paradigm to model the useless information (speaker- or channel-variability) in order to remove it from acoustic data frames. The transformed training data frames are then used to train new HMM models using the standard training algorithm. The transformation is also applied to the test data before the decoding process. With this approach we obtain, on french broadcast news, an absolute WER reduction of 1.3%. Mickael Rouvier, Mohamed Bouallegue, Driss Matrouf, Georges Linarès |
ASRU | 1 |
| 2011 | Speaker Role Recognition Using Question Detection and Characterization
Thierry Bazillon, Benjamin Maza, Mickael Rouvier, Frédéric Béchet, Alexis Nasr |
INTERSPEECH | 3 |
| 2011 | Static and dynamic video summariesabstractCurrently there are a lot of algorithms for video summarization; however most of them only represent visual information. In this paper, we propose two approaches for the construction of the summary using both video and text. One approach focuses on static summaries, where the summary is a set of selected keyframes and keywords, to be displayed in a fixed area. The second approach addresses dynamic summaries where video segments are selected based on both their visual and textual content to compose a new video sequence of predefined duration. Our approaches rely on an existing summarization algorithm, Video Maximal Marginal Relevance (Video-MMR), and its extension Text Video Maximal Marginal Relevance (TV-MMR) proposed by us. We describe the details of those approaches and present experimental results. Yingbo Li, Bernard Mérialdo, Mickael Rouvier, Georges Linarès |
ACM Multimedia | 3 |
| 2011 | Modeling nuisance variabilities with factor analysis for GMM-based audio pattern classification
Driss Matrouf, Florian Verdet, Mickael Rouvier, Jean-François Bonastre, Georges Linarès |
Comput. Speech Lang. | 3 |
| 2010 | Transcription-based video genre classificationabstractIn this paper, we present a new method for video genre identification based on the linguistic content analysis. This approach relies on the analysis of the most frequent words in the video transcriptions provided by an automatic speech recognition system. Experiments are conducted on a corpus composed of cartoons, movies, news, commercials, documentary, sport and music. On this 7-genre identification task, the proposed transcription-based method obtains up to 80% of correct identification. Finally, this rate is increased to 95% by combining the proposed linguistic-level features with low-level acoustic features. Stanislas Oger, Mickael Rouvier, Georges Linarès |
ICASSP | 2 |
| 2010 | On-the-fly video genre classification by combination of audio featuresabstractVideo genre identification methods are frequently based on image or motion analysis, which are relatively time-consuming processes. Since such approaches are tractable by batch processing, as-soon-as-possible identification requires faster methods. In this paper, we investigate the use of audio-only methods for on-the-fly video classification. We propose to use several acoustic feature streams and we evaluate various combination schemes at the frame or at the score level. Results are compared to those obtained by humans, according to the listening duration. Although the system based on model combination slightly outperforms the humans on very soon detection. The latter remain significantly more accurate on long sessions. Mickael Rouvier, Georges Linarès, Driss Matrouf |
ICASSP | 1 |
| 2010 | A language-identification inspired method for spontaneous speech detectionabstractInternational audience Mickael Rouvier, Richard Dufour, Georges Linarès, Yannick Estève |
INTERSPEECH | 1 |
| 2009 | Robust audio-based classification of video genreabstractInternational audience Mickael Rouvier, Georges Linarès, Driss Matrouf |
INTERSPEECH | 1 |
| 2009 | Factor analysis for audio-based video genre classificationabstractStatistical classifiers operate on features that generally include both useful and useless information. These two types of information are difficult to separate in the feature domain. Recently, a new paradigm based on a Latent Factor Analysis (LFA) proposed a model decomposition into usefull and useless components. This method was successfully applied to speaker and language recognition tasks. In this paper, we study the use of LFA for video genre classification by using only the audio channel. We propose a classification method based on short-term cep-stral features and Gaussian Mixture Models (GMM) or Support Vector Machine (SVM) classifiers, that are combined with Factor Analysis (FA). Experiments are conducted on a corpus composed of 5 types of video (musics, commercials, cartoons, movies and news). The relative classification error reduction obtained by using the best factor analysis configuration with respect to the baseline system, Gaussian Mixture Model Universal Background Model (GMM-UBM), is about 56%, corresponding to a correct identification rate of about 90%. Mickael Rouvier, Driss Matrouf, Georges Linarès |
INTERSPEECH | 1 |
| 2008 | On-the-fly term spotting by phonetic filtering and request-driven decodingabstractThis paper addresses the problem of on-the-fly term spotting in continuous speech streams. We propose a 2-level architecture in which recall and accuracy are sequentially optimized. The first level uses a cascade of phonetic filters to select the speech segments which probably contain the targeted terms. The second level performs a request-driven decoding of the selected speech segments. The results show good performance of the proposed system on broadcast news data : the best configuration reaches a F-measure of about 94% while respecting the on-the-fly processing constraint. Mickael Rouvier, Georges Linarès, Benjamin Lecouteux |
SLT | 1 |