EDBT 2026 Demo / reviewers in the wild / expert
Irina Illina
dblp:57/2152
· DBLP profile ↗
70ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0003-2598-4643ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 8 first-author · 5 since 2021Artificial intelligence and machine learning · 52 · 6 first-author · 14 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM-Based Data Generation and Clinical Skills Evaluation for Low-Resource French OSCEsabstractInternational audience Tian Huang, Tom Bourgeade, Irina Illina |
LREC | 3 |
| 2026 | A French OSCE Dialogue Dataset and Controllable Virtual Patient System for Clinical TrainingabstractThe clinical and communication skills of medical students are commonly assessed through Objective Structured Clinical Examinations (OSCEs), which consist of brief scenario-driven simulations of doctor-patient interactions. However, training is often limited by the low availability of human standardized patients, motivating the development of realistic virtual patients (VPs). To address this gap, we introduce a French OSCE dialogue dataset comprising 240 real student–patient training interactions. We build upon it a controllable LLM-based pipeline to generate synthetic OSCE dialogues. The pipeline integrates modular components, such as retrieval-based grounding and a reflection loop, to ensure patient fidelity, coherence, and realism. Additionally, we propose a multi-level evaluation framework assessing patient simulation quality, student performance, and linguistic quality, using an LLM-as-a-Judge approach. Experiments suggest that controllability modules generally improve patient fidelity and student evaluation consistency. Finally, we also implement an interactive prototype in which students can practice with a VP and receive automatic feedback. Doria Bonzi, Tom Bourgeade, Fabrice Lefèvre, Irina Illina |
SIGDIAL | 4 |
| 2025 | Mixture of LoRA Experts for Low-Resourced Multi-Accent Automatic Speech Recognition
Raphaël Bagat, Irina Illina, Emmanuel Vincent 0001 |
INTERSPEECH | 2 |
| 2025 | Efficient One-shot Compression via Low-Rank Local Feature DistillationabstractYaya Sy, Christophe Cerisara, Irina Illina. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yaya Sy, Christophe Cerisara, Irina Illina |
NAACL (Long Papers) | 3 |
| 2024 | Training RNN language models on uncertain ASR hypotheses in limited data scenarios
Imran A. Sheikh, Emmanuel Vincent 0001, Irina Illina |
Comput. Speech Lang. | 3 |
| 2022 | Domain Classification-based Source-specific Term Penalization for Domain Adaptation in Hate-speech DetectionabstractState-of-the-art approaches for hate-speech detection usually exhibit poor performance in out-of-domain settings. This occurs, typically, due to classifiers overemphasizing source-specific information that negatively impacts its domain invariance. Prior work has attempted to penalize terms related to hate-speech from manually curated lists using feature attribution methods, which quantify the importance assigned to input terms by the classifier when making a prediction. We, instead, propose a domain adaptation approach that automatically extracts and penalizes source-specific terms using a domain classifier, which learns to differentiate between domains, and feature-attribution scores for hate-speech classes, yielding consistent improvements in cross-domain evaluation. Tulika Bose, Nikolaos Aletras, Irina Illina, Dominique Fohr |
COLING | 3 |
| 2022 | Placing M-Phasis on the Plurality of Hate: A Feature-Based Corpus of Hate OnlineabstractEven though hate speech (HS) online has been an important object of research in the last decade, most HS-related corpora over-simplify the phenomenon of hate by attempting to label user comments as “hate” or “neutral”. This ignores the complex and subjective nature of HS, which limits the real-life applicability of classifiers trained on these corpora. In this study, we present the M-Phasis corpus, a corpus of ~9k German and French user comments collected from migration-related news articles. It goes beyond the “hate”-“neutral” dichotomy and is instead annotated with 23 features, which in combination become descriptors of various types of speech, ranging from critical comments to implicit and explicit expressions of hate. The annotations are performed by 4 native speakers per language and achieve high (0.77 <= k <= 1) inter-annotator agreements. Besides describing the corpus creation and presenting insights from a content, error and domain analysis, we explore its data characteristics by training several classification baselines. Dana Ruiter, Liane Reiners, Ashwin Geet D'Sa, Thomas Kleinbauer, Dominique Fohr, Irina Illina, Dietrich Klakow, Christian Schemer, Angeliki Monnier |
LREC | 6 |
| 2022 | Transformer versus LSTM Language Models trained on Uncertain ASR Hypotheses in Limited Data ScenariosabstractIn several ASR use cases, training and adaptation of domain-specific LMs can only rely on a small amount of manually verified text transcriptions and sometimes a limited amount of in-domain speech. Training of LSTM LMs in such limited data scenarios can benefit from alternate uncertain ASR hypotheses, as observed in our recent work. In this paper, we propose a method to train Transformer LMs on ASR confusion networks. We evaluate whether these self-attention based LMs are better at exploiting alternate ASR hypotheses as compared to LSTM LMs. Evaluation results show that Transformer LMs achieve 3-6% relative reduction in perplexity on the AMI scenario meetings but perform similar to LSTM LMs on the smaller Verbmobil conversational corpus. Evaluation on ASR N-best rescoring shows that LSTM and Transformer LMs trained on ASR confusion networks do not bring significant WER reductions. However, a qualitative analysis reveals that they are better at predicting less frequent words. Imran A. Sheikh, Emmanuel Vincent 0001, Irina Illina |
LREC | 3 |
| 2022 | Identification of Multiword Expressions in Tweets for Hate Speech DetectionabstractMultiword expression (MWE) identification in tweets is a complex task due to the complex linguistic nature of MWEs combined with the non-standard language use in social networks. MWE features were shown to be helpful for hate speech detection (HSD). In this article, we present joint experiments on these two related tasks on English Twitter data: first we focus on the MWE identification task, and then we observe the influence of MWE-based features on the HSD task. For MWE identification, we compare the performance of two systems: lexicon-based and deep neural networks-based (DNN). We experimentally evaluate seven configurations of a state-of-the-art DNN system based on recurrent networks using pre-trained contextual embeddings from BERT. The DNN-based system outperforms the lexicon-based one thanks to its superior generalisation power, yielding much better recall. For the HSD task, we propose a new DNN architecture for incorporating MWE features. We confirm that MWE features are helpful for the HSD task. Moreover, the proposed DNN architecture beats previous MWE-based HSD systems by 0.4 to 1.1 F-measure points on average on four Twitter HSD corpora. Nicolas Zampieri, Carlos Ramisch, Irina Illina, Dominique Fohr |
LREC | 3 |
| 2021 | Distributed Speech Separation in Spatially Unconstrained Microphone ArraysabstractSpeech separation with several speakers is a challenging task because of the non-stationarity of the speech and the strong signal similarity between interferent sources. Current state-of-the-art solutions can separate well the different sources using sophisticated deep neural networks which are very tedious to train. When several microphones are available, spatial information can be exploited to design much simpler algorithms to discriminate speakers. We propose a distributed algorithm that can process spatial information in a spatially unconstrained microphone array. The algorithm relies on a convolutional recurrent neural network that can exploit the signal diversity from the distributed nodes. In a typical case of a meeting room, this algorithm can capture an estimate of each source in a first step and propagate it over the microphone array in order to increase the separation performance in a second step. We show that this approach performs even better when the number of sources and nodes increases. We also study the influence of a mismatch in the number of sources between the training and testing conditions. Nicolas Furnon, Romain Serizel, Irina Illina, Slim Essid |
ICASSP | 3 |
| 2021 | Modeling and Training Strategies for Language Recognition SystemsabstractInternational audience Raphaël Duroselle, Md. Sahidullah, Denis Jouvet, Irina Illina |
Interspeech | 4 |
| 2021 | Language Recognition on Unknown Conditions: The LORIA-Inria-MULTISPEECH System for AP20-OLR ChallengeabstractInternational audience Raphaël Duroselle, Md. Sahidullah, Denis Jouvet, Irina Illina |
Interspeech | 4 |
| 2021 | BERT-Based Semantic Model for Rescoring N-Best Speech Recognition ListabstractInternational audience Dominique Fohr, Irina Illina |
Interspeech | 2 |
| 2021 | Multiword Expression Features for Automatic Hate Speech Detection
Nicolas Zampieri, Irina Illina, Dominique Fohr |
NLDB | 2 |
| 2021 | DNN-Based Mask Estimation for Distributed Speech Enhancement in Spatially Unconstrained Microphone ArraysabstractDeep neural network (DNN)-based speech enhancement algorithms in microphone arrays have now proven to be efficient solutions to speech understanding and speech recognition in noisy environments. However, in the context of ad-hoc microphone arrays, many challenges remain and raise the need for distributed processing. In this paper, we propose to extend a previously introduced distributed DNN-based time-frequency mask estimation scheme that can efficiently use spatial information in form of so-called compressed signals which are pre-filtered target estimations. We study the performance of this algorithm named Tango under realistic acoustic conditions and investigate practical aspects of its optimal application. We show that the nodes in the microphone array cooperate by taking profit of their spatial coverage in the room. We also propose to use the compressed signals not only to convey the target estimation but also the noise estimation in order to exploit the acoustic diversity recorded throughout the microphone array. Nicolas Furnon, Romain Serizel, Slim Essid, Irina Illina |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | DNN-based Distributed Multichannel Mask Estimation for Speech Enhancement in Microphone ArraysabstractMultichannel processing is widely used for speech enhancement but several limitations appear when trying to deploy these solutions in the real world. Distributed sensor arrays that consider several devices with a few microphones is a viable solution which allows for exploiting the multiple devices equipped with microphones that we are using in our everyday life. In this context, we propose to extend the distributed adaptive node-specific signal estimation approach to a neural network framework. At each node, a local filtering is performed to send one signal to the other nodes where a mask is estimated by a neural network in order to compute a global multichannel Wiener filter. In an array of two nodes, we show that this additional signal can be leveraged to predict the masks and leads to better speech enhancement performance than when the mask estimation relies only on the local signals. Nicolas Furnon, Romain Serizel, Irina Illina, Slim Essid |
ICASSP | 3 |
| 2020 | Metric Learning Loss Functions to Reduce Domain Mismatch in the x-Vector Space for Language RecognitionabstractInternational audience Raphaël Duroselle, Denis Jouvet, Irina Illina |
INTERSPEECH | 3 |
| 2020 | On Semi-Supervised LF-MMI Training of Acoustic Models with Limited DataabstractThis work investigates semi-supervised training of acoustic models (AM) with the lattice-free maximum mutual information (LF-MMI) objective in practically relevant scenarios with a limited amount of labeled in-domain data. An error detection driven semi-supervised AM training approach is proposed, in which an error detector controls the hypothesized transcriptions or lattices used as LF-MMI training targets on additional unlabeled data. Under this approach, our first method uses a single error-tagged hypothesis whereas our second method uses a modified supervision lattice. These methods are evaluated and compared with existing semi-supervised AM training methods in three different matched or mismatched, limited data setups. Word error recovery rates of 28 to 89% are reported. Imran A. Sheikh, Emmanuel Vincent 0001, Irina Illina |
INTERSPEECH | 3 |
| 2019 | VoiceHome-2, an extended corpus for multichannel speech processing in real homes
Nancy Bertin, Ewen Camberlein, Romain Lebarbenchon, Emmanuel Vincent 0001, Sunit Sivasankaran, Irina Illina, Frédéric Bimbot |
Speech Commun. | 6 |
| 2018 | Dynamic Extension of ASR Lexicon Using Wikipedia DataabstractDespite recent progress in developing Large Vocabulary Continuous Speech Recognition Systems (LVCSR), these systems suffer from-Of-Vocabulary words (OOV). In many cases, the OOV words are Proper Nouns (PNs). The correct recognition of PNs is essential for broadcast news, audio indexing, etc. In this article, we address the problem of OOV PN retrieval in the framework of broadcast news LVCSR. We focused on dynamic (document dependent) extension of LVCSR lexicon. To retrieve relevant OOV PNs, we propose to use a very large multipurpose text corpus: Wikipedia. This corpus contains a huge number of PNs. These PNs are grouped in semantically similar classes using word embedding. We use a two-step approach: first, we select OOV PN pertinent classes with a multi-class Deep Neural Network (DNN). Secondly, we rank the OOVs of the selected classes. The experiments on French broadcast news show that the Bi-GRU model outperforms other studied models. Speech recognition experiments demonstrate the effectiveness of the proposed methodology. Badr Abdullah, Irina Illina, Dominique Fohr |
SLT | 2 |
| 2018 | DNN Uncertainty Propagation Using GMM-Derived Uncertainty Features for Noise Robust ASRabstractThe uncertainty decoding framework is known to improve the deep neural network (DNN)-based automatic speech recognition (ASR) performance in noisy environments. It operates by estimating the statistical uncertainty about the input features and propagating it to the output senone posteriors by sampling. Unfortunately, this approximate propagation scheme limits the performance improvement. In this letter, we exploit the fact that uncertainty propagation can be achieved in closed form for Gaussian mixture acoustic models (GMMs). We introduce new GMM-derived (GMMD) uncertainty features for the robust DNN-based acoustic model training and decoding. The GMMD features are computed as the difference between the GMM log-likelihoods obtained with versus without uncertainty. They are concatenated with conventional acoustic features and used as inputs to the DNN. We evaluate the resulting ASR performance on the CHiME-2 and CHiME-3 datasets. The proposed features are shown to improve the performance on both datasets, both for the conventional decoding and for the uncertainty decoding with different uncertainty estimation/propagation techniques. Karan Nathwani, Emmanuel Vincent 0001, Irina Illina |
IEEE Signal Process. Lett. | 3 |
| 2017 | Consistent DNN uncertainty training and decoding for robust ASRabstractWe consider the problem of robust automatic speech recognition (ASR) in noisy conditions. The performance improvement brought by speech enhancement is often limited by residual distortions of the enhanced features, which can be seen as a form of statistical uncertainty. Uncertainty estimation and propagation methods have recently been proposed to improve the ASR performance with deep neural network (DNN) acoustic models. However, the performance is still limited due to the use of uncertainty only during decoding. In this paper, we propose a consistent approach to account for uncertainty in the enhanced features during both training and decoding. We estimate the variance of the distortions using a DNN uncertainty estimator that operates directly in the feature maximum likelihood linear regression (fMLLR) domain and we then sample the uncertain features using the unscented transform (UT). We report the resulting ASR performance on the CHiME-2 and CHiME-3 datasets for different uncertainty estimation/propagation techniques. The proposed DNN uncertainty training method brings 4% and 8% relative improvement on these two datasets, respectively, compared to a competitive fMLLR-domain DNN acoustic modeling baseline. Karan Nathwani, Emmanuel Vincent 0001, Irina Illina |
ASRU | 3 |
| 2017 | Topic segmentation in ASR transcripts using bidirectional RNNS for change detectionabstractTopic segmentation methods are mostly based on the idea of lexical cohesion, in which lexical distributions are analysed across the document and segment boundaries are marked in areas of low cohesion. We propose a novel approach for topic segmentation in speech recognition transcripts by measuring lexical cohesion using bidirectional Recurrent Neural Networks (RNN). The bidirectional RNNs capture context in the past and the following set of words. The past and following contexts are compared to perform topic change detection. In contrast to existing works based on sequence and discriminative models for topic segmentation, our approach does not use a segmented corpus nor (pseudo) topic labels for training. Our model is trained using news articles obtained from the internet. Evaluation on ASR transcripts of French TV broadcast news programs demonstrates the effectiveness of our proposed approach. Imran A. Sheikh, Dominique Fohr, Irina Illina |
ASRU | 3 |
| 2017 | Discriminative importance weighting of augmented training data for acoustic model trainingabstractDNN based acoustic models require a large amount of training data. Parametric data augmentation techniques such as adding noise, reverberation, or changing the speech rate, are often employed to boost the dataset size and the ASR performance. The choice of augmentation techniques and the associated parameters has been handled heuristically so far. In this work we propose an algorithm to automatically weight data perturbed using a variety of augmentation techniques and/or parameters. The weights are learned in a discriminative fashion so as to minimize the frame error rate using the standard gradient descent algorithm in an iterative manner. Experiments were performed using the CHiME-3 dataset. Data augmentation was done by adding noise at different SNRs. A relative WER improvement of 15% was obtained with the proposed data weighting algorithm compared to the unweighted augmented dataset. Interestingly, the resulting distribution of SNRs in the weighted training set differs significantly from that of the test set. Sunit Sivasankaran, Emmanuel Vincent 0001, Irina Illina |
ICASSP | 3 |
| 2017 | A combined evaluation of established and new approaches for speech recognition in varied reverberation conditions
Sunit Sivasankaran, Emmanuel Vincent 0001, Irina Illina |
Comput. Speech Lang. | 3 |
| 2017 | Modelling Semantic Context of OOV Words in Large Vocabulary Continuous Speech RecognitionabstractThe diachronic nature of broadcast news data leads to the problem of out-of-vocabulary (OOV) words in large vocabulary continuous speech recognition (LVCSR) systems. Analysis of OOV words reveals that a majority of them are proper names (PNs). However, PNs are important for automatic indexing of audio-video content and for obtaining reliable automatic transcriptions. In this paper, we focus on the problem of OOV PNs in diachronic audio documents. To enable the recovery of the PNs missed by the LVCSR system, relevant OOV PNs are retrieved by exploiting the semantic context of the LVCSR transcriptions. For retrieval of OOV PNs, we explore topic and semantic context derived from latent Dirichlet allocation (LDA) topic models, continuous word vector representations and the neural bag-of-words (NBOW) model which is capable of learning task specific word and context representations. We propose a neural bag-of-weighted words (NBOW2) model which learns to assign higher weights to words that are important for retrieval of an OOV PN. With experiments on French broadcast news videos, we show that the NBOW and NBOW2 models outperform the methods based on raw embeddings from LDA and Skip-gram models. Combining the NBOW and NBOW2 models gives a faster convergence during training. Second pass speech recognition experiments, in which the LVCSR vocabulary and language model are updated with the retrieved OOV PNs, demonstrate the effectiveness of the proposed context models. Imran A. Sheikh, Dominique Fohr, Irina Illina, Georges Linarès |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | A French Corpus for Distant-Microphone Speech Processing in Real HomesabstractInternational audience Nancy Bertin, Ewen Camberlein, Emmanuel Vincent 0001, Romain Lebarbenchon, Stéphane Peillon, Éric Lamande, Sunit Sivasankaran, Frédéric Bimbot, Irina Illina, Ariane Tom, Sylvain Fleury, Eric Jamet |
INTERSPEECH | 9 |
| 2016 | Improved Neural Bag-of-Words Model to Retrieve Out-of-Vocabulary Words in Speech RecognitionabstractInternational audience Imran A. Sheikh, Irina Illina, Dominique Fohr, Georges Linarès |
INTERSPEECH | 2 |
| 2016 | How Diachronic Text Corpora Affect Context based Retrieval of OOV Proper Names for Audio News
Imran A. Sheikh, Irina Illina, Dominique Fohr |
LREC | 2 |
| 2016 | Dynamic adjustment of language models for automatic speech recognition using word similarityabstractOut-of-vocabulary (OOV) words can pose a particular problem for automatic speech recognition (ASR) of broadcast news. The language models (LMs) of ASR systems are typically trained on static corpora, whereas new words (particularly new proper nouns) are continually introduced in the media. Additionally, such OOVs are often content-rich proper nouns that are vital to understanding the topic. In this work, we explore methods for dynamically adding OOVs to language models by adapting the n-gram language model used in our ASR system. We propose two strategies: the first relies on finding in-vocabulary (IV) words similar to the OOVs, where word embeddings are used to define similarity. Our second strategy leverages a small contemporary corpus to estimate OOV probabilities. The models we propose yield improvements in perplexity over the baseline; in addition, the corpus-based approach leads to a significant decrease in proper noun error rate over the baseline in recognition experiments. Anna Currey, Irina Illina, Dominique Fohr |
SLT | 2 |
| 2015 | Different word representations and their combination for proper name retrieval from diachronic documentsabstractThis paper deals with the problem of high-quality transcription systems for very large vocabulary automatic speech recognition (ASR). We investigate the problem of automatic retrieval of out-of-vocabulary (OOV) proper names (PNs). We want to take into account the temporal, syntactic and semantic context of words. Nowadays, Artificial Neural Networks (NN) are widely used in natural language processing: continuous space representations of words is learned automatically from unstructured text data. To model the latent topics at document level, Latent Dirichlet Allocation (LDA) has been successful. In this paper, we propose OOV PN retrieval using (1) temporal versus topic context modeling; (2) different word representation spaces for word-level and document-level context modeling; (3) combinations of retrieval results. Experimental evaluation on broadcast news data shows that the proposed method combinations lead to better results. This confirms the complementarity of methods. Irina Illina, Dominique Fohr |
ASRU | 1 |
| 2015 | Robust ASR using neural network based speech enhancement and feature simulationabstractWe consider the problem of robust automatic speech recognition (ASR) in the context of the CHiME-3 Challenge. The proposed system combines three contributions. First, we propose a deep neural network (DNN) based multichannel speech enhancement technique, where the speech and noise spectra are estimated using a DNN based regressor and the spatial parameters are derived in an expectation-maximization (EM) like fashion. Second, a conditional restricted Boltzmann machine (CRBM) model is trained using the obtained enhanced speech and used to generate simulated training and development datasets. The goal is to increase the similarity between simulated and real data, so as to increase the benefit of multicondition training. Finally, we make some changes to the ASR backend. Our system ranked 4th among 25 entries. Sunit Sivasankaran, Aditya Arie Nugraha, Emmanuel Vincent 0001, Juan Andres Morales-Cordovilla, Siddharth Dalmia, Irina Illina, Antoine Liutkus |
ASRU | 6 |
| 2015 | Experience of an International Collaborative Project with First Year Programming StudentsabstractThis paper describes an Erasmus Intensive Programme that used international collaboration as a novel pedagogical approach to teaching programming skills to first-year students in a blended learning context using a mixture of virtual environment and intensive teaching. The experience and outcomes of the Programme are evaluated from the viewpoints of the students and instructors and conclusions are drawn on the value and conduct of international student collaborations. James H. Paterson, Markku Karhu, Walter Cazzola, Irina Illina, Robert Law, Dario Machiodi, Marisa Maximiano, Catarina Silva 0001 |
COMPSAC | 4 |
| 2015 | OOV Proper Name retrieval using topic and lexical context modelsabstractRetrieving Proper Names (PNs) specific to an audio document can be useful for vocabulary selection and OOV recovery in speech recognition, as well as in keyword spotting and audio indexing tasks. We propose methods to infer and retrieve OOV PNs relevant to an audio news document by using probabilistic topic models trained over diachronic text news. LVCSR hypothesis on the audio news document is analysed for latent topics, which is then used to retrieve relevant OOV PNs. Using an LDA topic model we obtain Recall up to 0.87 and Mean Average Precision (MAP) of 0.26 with only top 10% of the retrieved OOV PNs. We further propose methods to re-score and retrieve rare OOV PNs, and a lexical context model to improve the target OOV PN rankings assigned by the topic model, which may be biased due to prominence of certain news events. Re-scoring rare OOV PNs improves Recall whereas the lexical context model improves MAP. Imran A. Sheikh, Irina Illina, Dominique Fohr, Georges Linarès |
ICASSP | 2 |
| 2015 | Continuous word representation using neural networks for proper name retrieval from diachronic documentsabstractDeveloping high-quality transcription systems for very large vocabulary corpora is a challenging task. Proper names are usually key to understanding the information contained in a document. One approach for increasing the vocabulary coverage of a speech transcription system is to automatically retrieve new proper names from contemporary diachronic text documents. In recent years, neural networks have been successfully applied to a variety of speech recognition tasks. In this paper, we investigate whether neural networks can enhance word representation in vector space for the vocabulary extension of a speech recognition system. This is achieved by using high-quality word vector representation of words from large amounts of unstructured text data proposed by Mikolov. This model allows to take into account lexical and semantic word relationships. Proposed methodology is evaluated in the context of broadcast news transcription. Obtained recall and ASR proper name error rate is compared to that obtained using cosine-based vector space methodology. Experimental results show a good ability of the proposed model to capture semantic and lexical information Dominique Fohr, Irina Illina |
INTERSPEECH | 2 |
| 2015 | Study of entity-topic models for OOV proper name retrievalabstractRetrieving Proper Names (PNs) relevant to an audio document
can improve speech recognition and content based audio-video
indexing. Latent Dirichlet Allocation (LDA) topic model has
been used to retrieve Out-Of-Vocabulary (OOV) PNs relevant to
an audio document with good recall rates. However, retrieval of
OOV PNs using LDA is affected by two issues, which we study
in this paper: (1)Word Frequency Bias (less frequent OOV PNs
are ranked lower); (2) Loss of Specificity (the reduced topic
space representation loses lexical context). Entity-Topic models
have been proposed as extensions of LDA to specifically learn
relations between words, entities (PNs) and topics. We study
OOV PN retrieval with Entity-Topic models and show that they
are also affected by word frequency bias and loss of specificity.
We evaluate our proposed methods for rare OOV PN re-ranking
and lexical context re-ranking for LDA as well as for Entity-
Topic models. The results show an improvement in both Recall
and the Mean Average Precision. Imran A. Sheikh, Irina Illina, Dominique Fohr |
INTERSPEECH | 2 |
| 2012 | Evaluating grapheme-to-phoneme converters in automatic speech recognition contextabstractThis paper deals with the evaluation of grapheme-to-phoneme (G2P) converters in a speech recognition context. The precision and recall rates are investigated as potential measures of the quality of the multiple generated pronunciation variants. Very different results are obtained whether or not we take into account the frequency of occurrence of the words. Since G2P systems are rarely evaluated on a speech recognition performance basis, the originality of this paper consists in using a speech recognition system to evaluate the G2P pronunciation variants. The results show that the training process is quite robust to some errors in the pronunciation lexicon, whereas pronunciation lexicon errors are harmful in the decoding process. Noticeable speech recognition performance improvements are achieved by combining two different G2P converters, one based on conditional random fields and the other on joint multigram models, as well as by checking the pronunciation variants of the most frequent words. Denis Jouvet, Dominique Fohr, Irina Illina |
ICASSP | 3 |
| 2012 | Combining criteria for the detection of incorrect entries of non-native speech in the context of foreign language learningabstractThis article analyzes the detection of incorrect entries of non-native speech in the context of foreign language learning. The purpose is to detect and reject incorrect entries (i.e. those for which the speech signal does not correspond at all to the associated text) while being tolerant to the mispronunciations of non-native speech. The proposed approach exploits the comparison between two text-to-speech alignments : one constrained by the text which is being checked, with another one unconstrained, corresponding to a phonetic decoding. Several comparison criteria are described and combined via a logistic regression function. The article analyzes the influence of different settings, such as the impact of non-native pronunciation variants, the impact of learning the decision functions on native or on non-native speech, as well as the impact of combining various comparison criteria. The performance evaluations are conducted both on native and on non-native speech. Luiza Orosanu, Denis Jouvet, Dominique Fohr, Irina Illina, Anne Bonneau |
SLT | 4 |
| 2011 | Grapheme-to-Phoneme Conversion Using Conditional Random FieldsabstractInternational audience Irina Illina, Dominique Fohr, Denis Jouvet |
INTERSPEECH | 1 |
| 2011 | About Handling Boundary Uncertainty in a Speaking Rate Dependent Modeling ApproachabstractInternational audience Denis Jouvet, Dominique Fohr, Irina Illina |
INTERSPEECH | 3 |
| 2010 | Detailed pronunciation variant modeling for speech transcriptionabstractInternational audience Denis Jouvet, Dominique Fohr, Irina Illina |
INTERSPEECH | 3 |
| 2010 | A wavelet-based parameterization for speech/music discrimination
E. Didiot, Irina Illina, Dominique Fohr, Odile Mella |
Comput. Speech Lang. | 2 |
| 2009 | Detection of OOV words by combining acoustic confidence measures with linguistic featuresabstractThis paper describes the design of an out-of-vocabulary words (OOV) detector. Such a system is assumed to detect segments that correspond to OOV words (words that are not included in the lexicon) in the output of a LVCSR system. The OOV detector uses acoustic confidence measures that are derived from several systems: a word recognizer constrained by a lexicon, a phone recognizer constrained by a grammar and a phone recognizer without constraints. On top of that it also uses some linguistic features. The experimental results on a French broadcast news transcription task showed that for our approach precision equals recall at 35%. Frederik Stouten, Dominique Fohr, Irina Illina |
ASRU | 3 |
| 2008 | Multi-accent and accent-independent non-native speech recognitionabstractInternational audience Ghazi Bouselmi, Dominique Fohr, Irina Illina |
INTERSPEECH | 3 |
| 2008 | Foreign accent identification based on prosodic parametersabstractInternational audience Marina Piat, Dominique Fohr, Irina Illina |
INTERSPEECH | 3 |
| 2008 | On-line Stochastic Matching compensation for non-stationary noise
Vincent Barreaud, Irina Illina, Dominique Fohr |
Comput. Speech Lang. | 2 |
| 2007 | Combined acoustic and pronunciation modelling for non-native speech recognitionabstractInternational audience Ghazi Bouselmi, Dominique Fohr, Irina Illina |
INTERSPEECH | 3 |
| 2006 | Fully Automated Non-Native Speech Recognition Using Confusion-Based Acoustic Model Integration and Graphemic ConstraintsabstractThis paper presents a fully automated approach for the recognition of non-native speech based on acoustic model modification. For a native language (L1) and a spoken language (L2), pronunciation variants of the phones of L2 are automatically extracted from an existing non-native database as a confusion matrix with sequences of phones of L1. This is done using LTs and L2's ASR systems. This confusion concept deals with the problem of non existence of match between some L2 and L1 phones. The confusion matrix is then used to modify the acoustic models (HMMs) of L2 phones by integrating corresponding L1 phone models as alternative HMM paths. We introduce graphemic constraints in the confusion extraction process: the phonetic confusion is established for each couple of 'L2-phone' and the grapheme(s) corresponding to that phone. We claim that pronunciation errors may depend on the graphemes related to each phone. The modified ASR system achieved an improvement between 32% and 40% (relative, L1=French and L2=English) in WER on the French non-native database used for testing. The introduction of graphemic constraints in the phonetic confusion allowed further improvements Ghazi Bouselmi, Dominique Fohr, Irina Illina, Jean Paul Haton |
ICASSP (1) | 3 |
| 2006 | Multilingual non-native speech recognition using phonetic confusion-based acoustic model modification and graphemic constraintsabstractIn this paper we present an automated approach for non-native speech recognition. We introduce a new phonetic confusion concept that associates sequences of native language (NL) phones to spoken language (SL) phones. Phonetic confusion rules are automatically extracted from a non-native speech database for a given NL and SL using both NL's and SL's ASR systems. These rules are used to modify the acoustic models (HMMs) of SL's ASR by adding acoustic models of NL's phones according to these rules. As pronunciation errors that non-native speakers produce depend on the writing of the words, we have also used graphemic constraints in the phonetic confusion extraction process. In the lexicon, the phones in words' pronunciations are linked to the corresponding graphemes (characters) of the word. In this way, the phonetic confusion is established between couples of (SL phones, graphemes) and sequences of NL phones. We evaluated our approach on French, Italian, Spanish and Greek non-native speech databases. The spoken language is English. The modified ASR system achieved significant improvements ranging from 20.3% to 43.2% (relative) in sentence error rate and from 26.6% to 50.0% in WER. Ghazi Bouselmi, Dominique Fohr, Irina Illina, Jean Paul Haton |
INTERSPEECH | 3 |
| 2006 | A wavelet-based parameterization for speech/music segmentationabstractThe problem of speech/music discrimination is a challenging research problem which significantly impacts Automatic Speech Recognition (ASR) performance. This paper proposes new features for the Speech/Music discrimination task. We propose to use a decomposition of the audio signal based on wavelets, which allows a good analysis of non stationary signal like speech or music. We compute different energy types in each frequency band obtained from wavelet decomposition. Two class/non-class classifiers are used : one for speech/non-speech, one for music/non-music. On the different test corpora, the proposed wavelet approach gives better results than the MFCC one. For instance, we have a significant relative improvements of the error rate of 71.6% on the ``Scheirer'' corpus for the speech/music discrimination task. E. Didiot, Irina Illina, Odile Mella, Dominique Fohr, Jean Paul Haton |
INTERSPEECH | 2 |
| 2005 | Fully automated non-native speech recognition using confusion-based acoustic model integrationabstractThis paper presents a fully automated approach for the recognition of non-native speech based on acoustic model modification. For a native language (L1) and a spoken language (L2), pronunciation variants of the phones of L2 are automatically extracted from an existing non-native database as a confusion matrix with sequences of phones of L1. This is done using L1's and L2's ASR systems. This confusion concept deals with the problem of non existence of match between some L2 and L1 phones. The confusion matrix is then used to modify the acoustic models (HMMs) of L2 phones by integrating corresponding L1 phone models as alternative HMM paths. In this way, no lexicon modification is carried. The modified ASR system achieved an improvement between 32% and 40% (relative, L1=French and L2=English) in WER on the French non-native database used for testing. Ghazi Bouselmi, Dominique Fohr, Irina Illina, Jean Paul Haton |
INTERSPEECH | 3 |
| 2004 | Exploiting models intrinsic robustness for noisy speech recognitionabstractColloque avec actes et comité de lecture. internationale. Christophe Cerisara, Dominique Fohr, Odile Mella, Irina Illina |
INTERSPEECH | 4 |
| 2004 | The automatic news transcription system: ANTS, some real time experimentsabstractThis paper presents the recent development of ANTS, the Automatic News Transcription System of LORIA. This system was designed in the framework of ESTER, the French broadcast radio news transcription task evaluation. After describing its different components and some segmentation and recognition results on the ESTER database, we present a number of experiments focusing on the real-time version of ANTS. We evaluate the system with different number of Gaussians, sizes of vocabulary and decoder settings. The non real time version of ANTS yields an overall error rate of 36 % that can be compared to 44.5 % for the real time version. Dominique Fohr, Odile Mella, Christophe Cerisara, Irina Illina |
INTERSPEECH | 4 |
| 2004 | Experiments on the accuracy of phone models and liaison processing in a French broadcast news transcription systemabstractColloque avec actes et comité de lecture. internationale. Dominique Fohr, Odile Mella, Irina Illina, Christophe Cerisara |
INTERSPEECH | 3 |
| 2004 | Hidden factor dynamic Bayesian networks for speech recognitionabstractColloque avec actes et comité de lecture. nationale. Filip Korkmazsky, Murat Deviren, Dominique Fohr, Irina Illina |
INTERSPEECH | 4 |
| 2004 | Using linear interpolation to improve histogram equalization for speech recognitionabstractColloque avec actes et comité de lecture. internationale. Filip Korkmazsky, Dominique Fohr, Irina Illina |
INTERSPEECH | 3 |
| 2003 | On-line frame-synchronous compensation of non-stationary noiseabstractW present a frame-synchronous noise compensation algorithm that uses a stochastic matching approach to cope with time-varying unknown noise. This method proposes to estimate a simple mapping function in parallel with Viterbi alignment. The technique is entirely general since no assumption is made on the nature, level and variation of noise. Our algorithm is evaluated on the VODIS database recorded in a moving car. For various tasks, our technique outperforms significantly classical methods. For instance, using the affine transformation the proposed algorithm gives an error rate improvement of 13.3 % compared to parallel model combination (PMC), 15.5 % on spectral subtraction (SS) and 27.8 % on frame-synchronous mean cepstre removal (MCR) for the numbers recognition task in real noise. Vincent Barreaud, Irina Illina, Dominique Fohr |
ICASSP (1) | 2 |
| 2003 | Combining EigenVoices and structural MLLR for speaker adaptationabstractThis paper considers the problem of speaker adaptation of acoustic models in speech recognition. We have investigated four different possible methods which integrate the concepts of both Structural Maximum Likelihood Linear Regression (SMLLR) and EigenVoices-based technique (EV) to adapt the Gaussian means of the speaker independent models for a new speaker. The experiments were evaluated using the speech recognition engine ESPERE on the data of the corpus Resource Management. They show that all of the proposed methods can improve the performances of an automatic speech recognition system (ASRS) in supervised batch adaptation as efficiently as SMLLR and EigenVoices-based techniques whatever the amount of adaptation data is available. For an unsupervised incremental adaptation, only the approach SMLLR + SEV gives the best results. Fabrice Lauri, Irina Illina, Dominique Fohr |
ICASSP (1) | 2 |
| 2003 | Structural state-based frame synchronous compensationabstractColloque avec actes et comité de lecture. internationale. Vincent Barreaud, Irina Illina, Dominique Fohr, Filip Korkmazsky |
INTERSPEECH | 2 |
| 2003 | Robust speech recognition to non-stationary noise based on model-driven approaches
Christophe Cerisara, Irina Illina |
INTERSPEECH | 2 |
| 2003 | Using genetic algorithms for rapid speaker adaptationabstractThis paper proposes two new approaches to rapid speaker adaptation of acoustic models by using genetic algorithms. Whereas conventional speaker adaptation techniques yield adapted models which represent local optimum solutions, genetic algorithms are capable to provide multiple optimal solutions, thereby delivering potentially more robust adapted models. We have investigated two different strategies of application of the genetic algorithm in the framework of speaker adaptation of acoustic models. The first approach (GA) consists in using a genetic algorithm to adapt the set of Gaussian means to a new speaker. The second approach (GA+ EV ) uses the genetic algorithm to enrich the set of speaker-dependant systems employed by the EigenVoices. Experiments with the Resource Management corpus show that, with one adaptation utterance, GA can improve the performances of a speaker-independant system as efficiently as EigenVoices. The method GA + EV outperforms EigenVoices. Fabrice Lauri, Irina Illina, Dominique Fohr, Filip Korkmazsky |
INTERSPEECH | 2 |
| 2002 | Tree-structured maximum a posteriori adaptation for a segment-based speech recognition systemabstractColloque avec actes et comité de lecture. internationale. Irina Illina |
INTERSPEECH | 1 |
| 1998 | Environment normalization training and environment adaptation using mixture stochastic trajectory model
Irina Illina, Mohamed Afify, Yifan Gong 0001 |
Speech Commun. | 1 |
| 1998 | Assessing the importance of the segmentation probability in segment-based speech recognition
Jan P. Verhasselt, Irina Illina, Jean-Pierre Martens, Yifan Gong 0001, Jean Paul Haton |
Speech Commun. | 2 |
| 1997 | Elimination of trajectory folding phenomenon: HMM, trajectory mixture HMM and mixture stochastic trajectory modelabstractIn this paper, a study of topology of hidden Markov model (HMM) used in speech recognition is addressed. Our main contribution is the introduction of the notion of trajectory folding phenomenon of HMM. In complex phonetic contexts and in speaker-variability, this phenomenon degrades the discriminability of HMM. The goal of this paper is to give some explanation and experimental evidence suggesting the existence of this phenomenon. The systems eliminating (partially or entirely) the trajectory folding are HMM with a special topology, called trajectory mixture HMM (TMHMM), and a mixture stochastic trajectory model (MSTM), proposed recently. HMM, TMHMM and MSTM have been tested on a 1011 words vocabulary, speaker dependent and multi-speaker continuous French speech recognition task. With similar number of model parameters, TMHMM and MSTM cuts down the error rate produced by the HMM, which confirms our hypothesis. Irina Illina, Yifan Gong 0001 |
ICASSP | 1 |
| 1997 | The importance of segmentation probability in segment based speech recognizersabstractIn segment based recognizers, variable length speech segments are mapped to the basic speech units (phones, diphones, ...). We address the acoustical modeling of these basic units in the framework of segmental posterior distribution models (SPDM). The joint posterior probability of a unit sequence u_ and a segmentation s_, Pr(u_,s_|X_) can be written as the product of the segmentation probability Pr(s_|X_) and the unit classification probability Pr(u_|s_,X_), where X_ is the sequence of acoustic observation parameter vectors. In particular, we point out the role of the segmentation probability and demonstrate that it does improve the recognition accuracy. We present evidence for this in two different tasks (speaker dependent continuous word recognition in French and speaker independent phone recognition in American English) in combination with two different unit classification models. Jan P. Verhasselt, Irina Illina, Jean-Pierre Martens, Yifan Gong 0001, Jean Paul Haton |
ICASSP | 2 |
| 1997 | Speaker normalization training for mixture stochastic trajectory modelabstractIn this paper we are interested in speaker and environment adaptation techniques for speaker independent (SI) continuous speech recognition. These techniques are used to reduce mismatch between training and the testing conditions, using a small amount of adaptation data. In addition to reducing this mismatch during the adaptation, we propose to reduce the variation due to speakers or environments during the training itself in the context of Speaker Normalisation (SN) approach, using MLLR transformation. SN also includes a combination of the context-dependent, phone dependent and broad phonetic class dependent information. The use of linear regression to model broad phonetic class dependent information assures our model to be used in the case that the adaptation data or training data is not given for some phonetic symbols. SN is developed for Mixture Stochastic Trajectory Model, a segment based model. The approach can be used for speaker, gender or environment normalization. We show the performance of SN compared to SI recognition and to MLLR speaker adaptation, through experiments on continuous speech recognition. Irina Illina, Yifan Gong 0001 |
EUROSPEECH | 1 |
| 1996 | Modelling long term variability information in mixture stochastic trajectory framework
Yifan Gong 0001, Irina Illina, Jean Paul Haton |
ICSLP | 2 |
| 1996 | Stochastic trajectory model with state-mixture for continuous speech recognition
Irina Illina, Yifan Gong 0001 |
ICSLP | 1 |
| 1996 | Improvement in n-best search for continuous speech recognition
Irina Illina, Yifan Gong 0001 |
ICSLP | 1 |