EDBT 2026 Demo / reviewers in the wild / expert
Olivier Siohan
dblp:66/58
· DBLP profile ↗
76ranked-venue papers
21as first author
12since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 67 · 19 first-author · 11 since 2021Artificial intelligence and machine learning · 46 · 11 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Large Scale Self-Supervised Pretraining for Active Speaker DetectionabstractIn this work we investigate the impact of a large-scale self-supervised pretraining strategy for active speaker detection (ASD) on an unlabeled dataset consisting of over 125k hours of YouTube videos. When compared to a baseline trained from scratch on much smaller in-domain labeled datasets we show that with pretraining we not only have a more stable supervised training due to better audio-visual features used for initialization, but also improve the ASD mean average precision by 23% on a challenging dataset collected with Google Nest Hub Max devices capturing real user interactions. Otavio Braga, Keith Johnson, Alice Chuang, Yunfan Ye, Olivier Siohan |
ICASSP | 6 |
| 2024 | Conformer is All You Need for Visual Speech RecognitionabstractVisual speech recognition models extract visual features in a hierarchical manner. At the lower level, there is a visual front-end with a limited temporal receptive field that processes the raw pixels depicting the lips or faces. At the higher level, there is an encoder that attends to the embeddings produced by the front-end over a large temporal receptive field. Previous work has focused on improving the visual front-end of the model to extract more useful features for speech recognition. Surprisingly, our work shows that complex visual front-ends are not necessary. Instead of allocating resources to a sophisticated visual front-end, we find that a linear visual front-end paired with a larger Conformer encoder results in lower latency, more efficient memory usage, and improved WER performance. We achieve a new state-of-the-art of 12.8% WER for visual speech recognition on the TED LRS3 dataset, which rivals the performance of audio-only models from just four years ago. Oscar Chang, Hank Liao, Dmitriy Serdyuk, Ankit Shahy, Olivier Siohan |
ICASSP | 5 |
| 2023 | Revisiting the Entropy Semiring for Neural Speech Recognition
Oscar Chang, Dongseong Hwang, Olivier Siohan |
ICLR | 3 |
| 2023 | Cascaded encoders for fine-tuning ASR models on overlapped speech
Richard Rose, Oscar Chang, Olivier Siohan |
INTERSPEECH | 3 |
| 2022 | Best of Both Worlds: Multi-Task Audio-Visual Automatic Speech Recognition and Active Speaker DetectionabstractUnder noisy conditions, automatic speech recognition (ASR) can greatly benefit from the addition of visual signals coming from a video of the speaker’s face. However, when multiple candidate speakers are visible this traditionally requires solving a separate problem, namely active speaker detection (ASD), which entails selecting at each moment in time which of the visible faces corresponds to the audio. Recent work has shown that we can solve both problems simultaneously by employing an attention mechanism over the competing video tracks of the speakers’ faces, at the cost of sacrificing some accuracy on active speaker detection. This work closes this gap in active speaker detection accuracy by presenting a single model that can be jointly trained with a multi-task loss. By combining the two tasks during training we reduce the ASD classification accuracy by approximately 25%, while simultaneously improving the ASR performance when compared to the multi-person baseline trained exclusively for ASR. Otavio Braga, Olivier Siohan |
ICASSP | 2 |
| 2022 | End-to-End multi-talker audio-visual ASR using an active speaker attention moduleabstractThis paper presents a new approach for end-to-end audio-visual multi-talker speech recognition.The approach, referred to here as the visual context attention model (VCAM), is important because it uses the available video information to assign decoded text to one of multiple visible faces.This essentially resolves the label ambiguity issue associated with most multi-talker modeling approaches which can decode multiple label strings but cannot assign the label strings to the correct speakers.This is implemented as a transformer-transducer based end-to-end model and evaluated using a two speaker audio-visual overlapping speech dataset created from YouTube videos.It is shown in the paper that the VCAM model improves performance with respect to previously reported audio-only and audio-visual multi-talker ASR systems. Richard Rose, Olivier Siohan |
INTERSPEECH | 2 |
| 2022 | Transformer-Based Video Front-Ends for Audio-Visual Speech Recognition for Single and Muti-Person Video
Dmitriy Serdyuk, Otavio Braga, Olivier Siohan |
INTERSPEECH | 3 |
| 2021 | Action Item Detection in Meetings Using Pretrained TransformersabstractDetecting the action items that were agreed upon during a meeting has important practical applications. But, extremely sparse positive labels, noisy annotations, and small datasets limited the improvement in this task via hand-crafted features techniques. Given the breakthrough performance demonstrated by pretrained transformer-based models in a wide variety of NLP tasks, such as BERT [1] and ETC [2], we revisit this task using these modern techniques. We empirically show how these modelling techniques advance the state-of-the-art on action item detection for the ICSI simulated meeting corpus [3] by 75% and establish a baseline for the action item detection from the AMI [4] meeting corpus. We show that these models are competitive on the related ICSI MRDA classification problem. In order to push the performance even further, we re-evaluate the task definition for action item detection, drawing upon the similarities with the span boundary detection realm. We propose the use of Generalized Hamming Distance ghd as an alternative evaluation. We hope to motivate further interest into the action item detection task by the community. Kishan Sachdeva, Joshua Maynez, Olivier Siohan |
ASRU | 3 |
| 2021 | Audio-Visual Speech Recognition is Worth $32\times 32\times 8$ VoxelsabstractAudio-visual automatic speech recognition (AV-ASR) intro-duces the video modality into the speech recognition process, often by relying on information conveyed by the motion of the speaker's mouth. The use of the video signal requires extracting visual features, which are then combined with the acoustic features to build an AV-ASR system [1]. This is tra-ditionally done with some form of 3D convolutional network (e.g. VGG) as widely used in the computer vision community. Recently, image transformers [2] have been introduced to ex-tract visual features useful for image classification tasks. In this work, we propose to replace the 3D convolutional visual front-end with a video transformer front-end. We train our systems on a large-scale dataset composed of YouTube videos and evaluate performance on the publicly available LRS3-TED set, as well as on a large set of YouTube videos. On a lip-reading task, the transformer-based front-end shows superior performance compared to a strong convolutional baseline. On an AV-ASR task, the transformer front-end performs as well as (or better than) the convolutional baseline. Fine-tuning our model on the LRS3-TED training set matches previous state of the art. Thus, we experimentally show the viability of the convolution-free model for AV-ASR. Dmitriy Serdyuk, Otavio Braga, Olivier Siohan |
ASRU | 3 |
| 2021 | A Closer Look at Audio-Visual Multi-Person Speech Recognition and Active Speaker SelectionabstractAudio-visual automatic speech recognition is a promising approach to robust ASR under noisy conditions. However, up until recently it had been traditionally studied in isolation assuming the video of a single speaking face matches the audio, and selecting the active speaker at inference time when multiple people are on screen was put aside as a separate problem. As an alternative, recent work has proposed to address the two problems simultaneously with an attention mechanism, baking the speaker selection problem directly into a fully differentiable model. One interesting finding was that the attention indirectly learns the association between the audio and the speaking face even though this correspondence is never explicitly provided at training time. In the present work we further investigate this connection and examine the interplay between the two problems. With experiments involving over 50 thousand hours of public YouTube videos as training data, we first evaluate the accuracy of the attention layer on an active speaker selection task. Secondly, we show under closer scrutiny that an end-to-end model performs at least as well as a considerably larger two-step system that utilizes a hard decision boundary under various noise conditions and number of parallel face tracks. Otavio Braga, Olivier Siohan |
ICASSP | 2 |
| 2021 | Bridging the Gap Between Streaming and Non-Streaming ASR Systems by Distilling Ensembles of CTC and RNN-T Models
Thibault Doutre, Wei Han 0002, Chung-Cheng Chiu, Ruoming Pang, Olivier Siohan, Liangliang Cao |
Interspeech | 5 |
| 2021 | End-to-End Audio-Visual Speech Recognition for Overlapping Speech
Richard Rose, Olivier Siohan, Anshuman Tripathi, Otavio Braga |
Interspeech | 2 |
| 2020 | End-to-End Multi-Person Audio/Visual Automatic Speech RecognitionabstractTraditionally, audio-visual automatic speech recognition has been studied under the assumption that the speaking face on the visual signal is the face matching the audio. However, in a more realistic setting, when multiple faces are potentially on screen one needs to decide which face to feed to the A/V ASR system. The present work takes the recent progress of A/V ASR one step further and considers the scenario where multiple people are simultaneously on screen (multi-person A/V ASR). We propose a fully differentiable A/V ASR model that is able to handle multiple face tracks in a video. Instead of relying on two separate models for speaker face selection and audiovisual ASR on a single face track, we introduce an attention layer to the ASR encoder that is able to soft-select the appropriate face video track. Experiments carried out on an A/V system trained on over 30k hours of YouTube videos illustrate that the proposed approach can automatically select the proper face tracks with minor WER degradation compared to an oracle selection of the speaking face while still showing benefits of employing the visual signal instead of the audio alone. Otavio Braga, Takaki Makino, Olivier Siohan, Hank Liao |
ICASSP | 3 |
| 2019 | Recurrent Neural Network Transducer for Audio-Visual Speech RecognitionabstractThis work presents a large-scale audio-visual speech recognition system based on a recurrent neural network transducer (RNN-T) architecture. To support the development of such a system, we built a large audio-visual (A/V) dataset of segmented utterances extracted from YouTube public videos, leading to 31k hours of audio-visual training content. The performance of an audio-only, visual-only, and audio-visual system are compared on two large-vocabulary test sets: a set of utterance segments from public YouTube videos called YTDEV18 and the publicly available LRS3-TED set. To highlight the contribution of the visual modality, we also evaluated the performance of our system on the YTDEV18 set artificially corrupted with background noise and overlapping speech. To the best of our knowledge, our system significantly improves the state-of-the-art on the LRS3-TED set. Takaki Makino, Hank Liao, Yannis M. Assael, Brendan Shillingford, Basilio Garcia, Otavio Braga, Olivier Siohan |
ASRU | 7 |
| 2017 | Acoustic Modeling for Google Home
Bo Li 0028, Tara N. Sainath, Arun Narayanan, Joe Caroselli, Michiel Bacchiani, Ananya Misra, Izhak Shafran, Hasim Sak, Golan Pundak, Kean K. Chin, Khe Chai Sim, Ron J. Weiss, Kevin W. Wilson, Ehsan Variani, Chanwoo Kim 0001, Olivier Siohan, Mitch Weintraub, Erik McDermott, Richard Rose, Matt Shannon |
INTERSPEECH | 16 |
| 2017 | Annealed f-Smoothing as a Mechanism to Speed up Neural Network Training
Tara N. Sainath, Vijayaditya Peddinti, Olivier Siohan, Arun Narayanan |
INTERSPEECH | 3 |
| 2017 | CTC Training of Multi-Phone Acoustic Models for Speech Recognition
Olivier Siohan |
INTERSPEECH | 1 |
| 2016 | Sequence training of multi-task acoustic models using meta-state labelsabstractIn this paper, we describe a multi-task learning approach for acoustic modeling where the multiple output layers are used to predict context-dependent (CD) states from different state inventories. Unlike the traditional multitask learning approach which defines a primary and secondary output layers but discards the secondary output after training, we propose to use all output layers for recognition. This can be achieved by designing a decoding network operating on tuples of CD states and combining the scores of the different outputs during search. To support training such models using a sequence-based criterion, we propose to replace the multiple output layers with a single layer encoding the CD state tuples as "meta-states". Experimental results are given on a large Voice Search task evaluated on children's speech. Olivier Siohan |
ICASSP | 1 |
| 2016 | Selection and combination of hypotheses for dialectal speech recognitionabstractWhile research has often shown that building dialect-specific Automatic Speech Recognizers is the optimal approach to dealing with dialectal variations of the same language, we have observed that dialect-specific recognizers do not always output the best recognitions. Often enough, another dialectal recognizer outputs a better recognition than the dialect-specific one. In this paper, we present two methods to select and combine the best decoded hypothesis from a pool of dialectal recognizers. We follow a Machine Learning approach and extract features from the Speech Recognition output along with Word Embeddings and use Shallow Neural Networks for classification. Our experiments using Dictation and Voice Search data from the main four Arabic dialects show good WER improvements for the hypothesis selection scheme, reducing the WER by 2.1 to 12.1% depending on the test set, and promising results for the hypotheses combination scheme. Victor Soto, Olivier Siohan, Mohamed G. Elfeky, Pedro J. Moreno 0001 |
ICASSP | 2 |
| 2016 | Automatic optimization of data perturbation distributions for multi-style training in speech recognitionabstractSpeech recognition performance using deep neural network based acoustic models is known to degrade when the acoustic environment and the speaker population in the target utterances are significantly different from the conditions represented in the training data. To address these mismatched scenarios, multi-style training (MTR) has been used to perturb utterances in an existing uncorrupted and potentially mismatched training speech corpus to better match target domain utterances. This paper addresses the problem of determining the distribution of perturbation levels for a given set of perturbation types that best matches the target speech utterances. An approach is presented that, given a small set of utterances from a target domain, automatically identifies an empirical distribution of perturbation levels that can be applied to utterances in an existing training set. Distributions are estimated for perturbation types that include acoustic background environments, reverberant room configurations, and speaker related variation like frequency and temporal warping. The end goal is for the resulting perturbed training set to characterize the variability in the target domain and thereby optimize ASR performance. An experimental study is performed to evaluate the impact of this approach on ASR performance when the target utterances are taken from a simulated far-field acoustic environment. Mortaza Doulaty, Richard Rose, Olivier Siohan |
SLT | 3 |
| 2015 | Multitask learning and system combination for automatic speech recognitionabstractIn this paper we investigate the performance of an ensemble of convolutional, long short-term memory deep neural networks (CLDNN) on a large vocabulary speech recognition task. To reduce the computational complexity of running multiple recognizers in parallel, we propose instead an early system combination approach which requires the construction of a static decoding network encoding the multiple context-dependent state inventories from the distinct acoustic models. To further reduce the computational load, the hidden units of those models can be shared while keeping the output layers distinct, leading to a multitask training formulation. However in contrast to the traditional multitask training, our formulation uses all predicted outputs leading to a multitask system combination strategy. Results are presented on a Voice Search task designed for children and outperform our current production system. Olivier Siohan, David Rybach |
ASRU | 1 |
| 2015 | Exemplar-based large vocabulary speech recognition using k-nearest neighborsabstractThis paper describes a large scale exemplar-based acoustic modeling approach for large vocabulary continuous speech recognition. We construct an index of labeled training frames using high-level features extracted from the bottleneck layer of a deep neural network as indexing features. At recognition time, each test frame is turned into a query and a set of k-nearest neighbor frames is retrieved from the index. This set is further filtered using majority voting and the remaining frames are used to derive an estimate of the context-dependent state posteriors of the query, which can then be used for recognition. Using an approximate nearest neighbor search approach based on asymmetric hashing, we are able to construct an index on over 25,000 hours of training data. We present both frame classification and recognition experiments on a Voice Search task. Yanbo Xu, Olivier Siohan, David Simcha, Sanjiv Kumar, Hank Liao |
ICASSP | 2 |
| 2015 | Large vocabulary automatic speech recognition for childrenabstractRecently, Google launched YouTube Kids, a mobile application for children, that uses a speech recognizer built specifically for recognizing children’s speech. In this paper we present techniques we explored to build such a system. We describe the use of a neural network classifier to identify matched acoustic training data, filtering data for language modeling to reduce the chance of producing offensive results. We also compare long short-term memory (LSTM) recurrent networks to convolutional, LSTM, deep neural networks (CLDNN). We found that a CLDNN acoustic model outperforms an LSTM across a variety of different conditions, but does not specifically model child speech relatively better than adult. Overall, these findings allow us to build a successful, state-of-the-art large vocabulary speech recognizer for both children and adults. Hank Liao, Golan Pundak, Olivier Siohan, Melissa K. Carroll, Noah Coccaro, Qi-Ming Jiang, Tara N. Sainath, Andrew W. Senior, Françoise Beaufays, Michiel Bacchiani |
INTERSPEECH | 3 |
| 2014 | Training data selection based on context-dependent state matchingabstractIn this paper we construct a data set for semi-supervised acoustic model training by selecting spoken utterances from a massive collection of anonymized Google Voice Search utterances. Semi-supervised training usually retains high-confidence utterances which are presumed to have an accurate hypothesized transcript, a necessary condition for successful training. Selecting high confidence utterances can however restrict the diversity of the resulting data set. We propose to introduce a constraint enforcing that the distribution of the context-dependent state symbols obtained by running forced alignment of the hypothesized transcript matches a reference distribution estimated from a curated development set. The quality of the obtained training set is illustrated on large scale Voice Search recognition experiments and outperforms random selection of high-confidence utterances. Olivier Siohan |
ICASSP | 1 |
| 2014 | A big data approach to acoustic model training corpus selectionabstractDeep neural networks (DNNs) have recently become the state of the art technology in speech recognition systems. In this pa-per we propose a new approach to constructing large high qual-ity unsupervised sets to train DNN models for large vocabulary speech recognition. The core of our technique consists of two steps. We first redecode speech logged by our production rec-ognizer with a very accurate (and hence too slow for real-time usage) set of speech models to improve the quality of ground truth transcripts used for training alignments. Using confidence scores, transcript length and transcript flattening heuristics de-signed to cull salient utterances from three decades of speech per language, we then carefully select training data sets consist-ing of up to 15K hours of speech to be used to train acoustic models without any reliance on manual transcription. We show that this approach yields models with approximately 18K con-text dependent states that achieve 10 % relative improvement in large vocabulary dictation and voice-search systems for Brazil-ian Portuguese, French, Italian and Russian languages. Index Terms: large unsupervised training sets, data selection, Olga Kapralova, John Alex, Eugene Weinstein, Pedro J. Moreno 0001, Olivier Siohan |
INTERSPEECH | 5 |
| 2013 | ivector-based acoustic data selectionabstractThis paper presents a data selection approach where spoken ut-terances are selected in a sequential fashion from a large out-of-domain data set to match the utterance distribution of an in-domain data set. We propose to represent each utterance by its iVector [1], a low dimensional vector indicating the coordi-nate of that utterance in a subspace acoustic model. We show that the distribution of iVectors can characterize a data set and enables distinguishing subsets of utterances from different do-mains. Last, we present experimental speech recognition results based on a system trained on a data set constructed by the pro-posed algorithm and a comparison with random data selection. Index Terms: speech recognition, data selection, acoustic mod-eling Olivier Siohan, Michiel Bacchiani |
INTERSPEECH | 1 |
| 2010 | Decision tree state clustering with word and syllable featuresabstractIn large vocabulary continuous speech recognition, decision trees are widely used to cluster triphone states. In addition to commonly used phonetically based questions, others have proposed additional questions such as phone position within word or syllable. This paper examines using the word or syllable context itself as a feature in the decision tree, providing an elegant way of introducing word- or syllable-specific models into the system. Positive results are reported on two state-of-the-art systems: voicemail transcription and a search by voice tasks across av ariety of acoustic model and training set sizes. Index Terms :d ecision tree state clustering, large vocabulary continuous speech recognition, tagged clustering. Hank Liao, Christopher Alberti, Michiel Bacchiani, Olivier Siohan |
INTERSPEECH | 4 |
| 2009 | An audio indexing system for election video materialabstractIn the 2008 presidential election race in the United States, the prospective candidates made extensive use of YouTube to post video material. We developed a scalable system that transcribes this material and makes the content searchable (by indexing the meta-data and transcripts of the videos) and allows the user to navigate through the video material based on content. The system is available as an iGoogle gadget1as well as a Labs product (labs.google.com/gaudi). Given the large exposure, special emphasis was put on the scalability and reliability of the system. This paper describes the design and implementation of this system. Christopher Alberti, Michiel Bacchiani, Ari Bezman, Ciprian Chelba, Anastassia Drofa, Hank Liao, Pedro J. Moreno 0001, Ted Power, Arnaud Sahuguet, Maria Shugrina, Olivier Siohan |
ICASSP | 11 |
| 2007 | The IBM 2007 speech transcription system for European parliamentary speechesabstractTC-STAR is an European Union funded speech to speech translation project to transcribe, translate and synthesize European Parliamentary Plenary Speeches (EPPS). This paper describes IBM's English speech recognition system submitted to the TC-STAR 2007 Evaluation. Language model adaptation based on clustering and data selection using relative entropy minimization provided significant gains in the 2007 evaluation. The additional advances over the 2006 system that we present in this paper include unsupervised training of acoustic and language models; a system architecture that is based on cross-adaptation across complementary systems and system combination through generation of an ensemble of systems using randomized decision tree state-tying. These advances reduced the error rate by 30% relative over the best-performing system in the TC-STAR 2006 evaluation on the 2006 English development and evaluation test sets, and produced one of the best performing systems on the 2007 evaluation in English with a word error rate of 7.1%. Bhuvana Ramabhadran, Olivier Siohan, Abhinav Sethy |
ASRU | 2 |
| 2007 | Gaussian Mixture Language Models for Speech RecognitionabstractWe propose a Gaussian mixture language model for speech recognition. Two potential benefits of using this model are smoothing unseen events, and ease of adaptation. It is shown how this model can be used alone or in conjunction with a a conventional N-gram model to calculate word probabilities. An interesting feature of the proposed technique is that many methods developed for acoustic models can be easily ported to GMLM. We developed two implementations of the proposed model for large vocabulary Arabic speech recognition with results comparable to conventional N-gram. Mohamed Afify, Olivier Siohan, Ruhi Sarikaya |
ICASSP (4) | 2 |
| 2007 | Vocabulary independent spoken term detectionabstractWe are interested in retrieving information from speech data like broadcast news, telephone conversations and roundtable meetings. Today, most systems use large vocabulary continuous speech recognition tools to produce word transcripts; the transcripts are indexed and query terms are retrieved from the index. However, query terms that are not part of the recognizer's vocabulary cannot be retrieved, and the recall of the search is affected. In addition to the output word transcript, advanced systems provide also phonetic transcripts, against which query terms can be matched phonetically. Such phonetic transcripts suffer from lower accuracy and cannot be an alternative to word transcripts.We present a vocabulary independent system that can handle arbitrary queries, exploiting the information provided by having both word transcripts and phonetic transcripts. A speech recognizer generates word confusion networks and phonetic lattices. The transcripts are indexed for query processing and ranking purpose.The value of the proposed method is demonstrated by the relative high performance ofour system, which received the highest overall ranking for US English speech data in the recent NIST Spoken Term Detection evaluation. Jonathan Mamou, Bhuvana Ramabhadran, Olivier Siohan |
SIGIR | 3 |
| 2007 | Comments on Vocal Tract Length Normalization Equals Linear Transformation in Cepstral SpaceabstractThe bilinear transformation (BT) is used for vocal tract length normalization (VTLN) in speech recogniton systems. We prove two properties of the bilinear mapping that motivated the band-diagonal transform proposed in M. Afify and O. Siohan, (ldquoConstrained maximum likelihood linear regression for speaker adaptation,rdquo in Proc. ICSLP, Beijing, China, Oct. 2000.) This is in contrast to what is stated in M. Pitz and H. Ney, (ldquoVocal tract length normalization equals linear transformation in cepstral space,rdquo IEEE Transactions on Speech and Audio Processing, vol. 13, no. 5, pp 930-944, September 2005) that the transform of Afify and Siohan was motivated by empirical observations. Mohamed Afify, Olivier Siohan |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Automated Quality Monitoring in the Call Center with ASR and Maximum EntropyabstractThis paper describes an automated system for assigning quality scores to recorded call center conversations. The system combines speech recognition, pattern matching, and maximum entropy classification to rank calls according to their measured quality. Calls at both end of the spectrum are flagged as "interesting" and made available for further human monitoring. In this process, pattern matching on the ASR transcript is used to answer a set of standard quality control questions such as "did the agent use courteous words and phrases," and to generate a question-based score. This is interpolated with the probability of a call being "bad," as determined by maximum entropy operating on a set of ASR-derived features such as "maximum silence length" and the occurrence of selected n-gram word sequences. The system is trained on a set of calls with associated manual evaluation forms. We present precision and recall results from IBM's North American Help Desk indicating that for a given amount of listening effort, this system triples the number of bad calls that are identified, over the current policy of randomly sampling calls Geoffrey Zweig, Olivier Siohan, George Saon, Bhuvana Ramabhadran, Daniel Povey, Lidia Mangu, Brian Kingsbury |
ICASSP (1) | 2 |
| 2006 | The IBM 2006 speech transcription system for european parliamentary speechesabstractTC-STAR is an European Union funded speech to speech translation project to transcribe, translate and synthesize European Parliamentary Plenary Speeches (EPPS). This paper describes IBM’s English and Spanish speech recognition systems submitted to the TC-STAR 2006 Evaluation. The technical advances in this submission include two different algorithms for automatic segmentation and speaker clustering of the input audio; a system architecture that is based on cross-adaptation across these two segmentation schemes and system combination through generation of an ensemble of systems using randomized decision tree state-tying; automatic punctuation of the speech recognition output; and the incorporation of an additional 35 hours of in-domain EPPS acoustic training data. These advances reduced the error rate by 30% relative over the best-performing system in the TC-STAR 2005 Evaluation on the 2006 English development test set, and produced one of the best performing systems on the 2006 evaluation in English with a word error rate of 8.3%. Index Terms: speech recognition, automatic segmentation, crossadaptation, randomized decision trees, TC-STAR. Bhuvana Ramabhadran, Olivier Siohan, Lidia Mangu, Geoffrey Zweig, Martin Westphal, Henrik Schulz 0002, Alvaro Soneiro |
INTERSPEECH | 2 |
| 2006 | Automated Quality Monitoring for Call Centers using Speech and NLP Technologies
Geoffrey Zweig, Olivier Siohan, George Saon, Bhuvana Ramabhadran, Daniel Povey, Lidia Mangu, Brian Kingsbury |
HLT-NAACL | 2 |
| 2005 | Contructing Ensembles of ASR Systems Using Randomized Decision TreesabstractBuilding multiple automatic speech recognition (ASR) systems and combining their outputs using voting techniques such as ROVER is an effective technique for lowering the overall word error rate. A successful system combination approach requires the construction of multiple systems with complementary errors, or the combination will not outperform any of the individual systems. In general, this is achieved empirically, for example by building systems on different input features. In this paper, we present a systematic approach for building multiple ASR systems in which the decision tree state-tying procedure that is used to specify context-dependent acoustic models is randomized. Experiments carried out on two large vocabulary recognition tasks, MALACH and DARPA EARS, illustrate the effectiveness of the approach. Olivier Siohan, Bhuvana Ramabhadran, Brian Kingsbury |
ICASSP (1) | 1 |
| 2005 | Fast vocabulary-independent audio search using path-based graph indexingabstractClassical audio retrieval techniques consist in transcribing audio documents using a large vocabulary speech recognition system and indexing the resulting transcripts. However, queries that are not part of the recognizer’s vocabulary or have a large probability of getting misrecognized can significantly impair the performance of the retrieval system. Instead, we propose a fast vocabulary independent audio search approach that operates on phonetic lattices and is suitable for any query. However, indexing phonetic lattices so that any arbitrary phone sequence query can be processed efficiently is a challenge, as the choice of the indexing unit is unclear. We propose an inverted index structure on lattices that uses paths as indexing features. The approach is inspired by a general graph indexing method that defines an automatic procedure to select a small number of paths as indexing features, keeping the index size small while allowing fast retrieval of the lattices matching a given query. The effectiveness of the proposed approach is illustrated on broadcast news and Switchboard databases. Olivier Siohan, Michiel Bacchiani |
INTERSPEECH | 1 |
| 2005 | A new verification-based fast-match for large vocabulary continuous speech recognitionabstractAcoustic fast-match is a popular way to accelerate the search in large-vocabulary continuous-speech recognition, where an efficient method is used to identify poorly scoring phonemes and discard them from detailed evaluation. In this paper we view acoustic fast-match as a verification problem, and hence develop an efficient likelihood ratio test, similar to other verification scenarios, to perform the fast match. Various aspects of the test like the design of alternate hypothesis models and the setting of phoneme look-ahead durations and decision thresholds are studied, resulting in an efficient implementation. The proposed fast-match is tested in a large vocabulary speech recognition task and it is demonstrated that depending on the decision threshold, it leads to 20-30% improvement in speed without any loss in recognition accuracy. In addition, it significantly outperforms a similar test based on using likelihoods only, which fails, in our setting, to bring any improvement in speed-accuracy trade-off. In a larger set of experiments with varying acoustic and task conditions, similar improvements are observed for the fast-match with the same model and setting. This indicates the robustness of the proposed technique. The gains due to the proposed method are obtained within a highly efficient 2-pass search strategy and similar or even higher gains are expected in other alternative search architectures. Mohamed Afify, Hui Jiang 0001, Olivier Siohan |
IEEE Trans. Speech Audio Process. | 4 |
| 2004 | Use of metadata to improve recognition of spontaneous speech and named entitiesabstractWith improved recognition accuracies for LVCSR tasks, it has become possible to search large collections of spontaneous speech for a variety of information. The MALACH corpus of Holocaust testimonials is one such collection, in which we are interested in automatically transcribing and retrieving portions that are relevant to named entities such as people, places, and organizations. Since the testimonials were gathered from thousands of people in countries throughout Europe, an extremely large number of potential named entities are possible, and this causes a well-known dilemma: increasing the size of the vocabulary allows for more of these words to be recognized, but also increases confusability, and can harm recognition performance. However, the MALACH corpus, like many other collections, includes side information or metadata that can be exploited to provide prior information on exactly which named entities are likely to appear. This paper proposes a method that capitalizes on this prior information to reduce named-entity recognition errors by over 50 % relative, and simultaneously decrease the overall word error rate by 7 % relative. The metadata we use derives from a pre-interview questionaire that includes the names of friends, relatives, places visited, membership of organizations, synonyms of place names, and similar information. By augmenting the lexicon and language model with this information on a speaker-by-speaker basis, we are able to exploit the textual information that is already available in the corpus to facilitate much improved speech recognition. 1. Bhuvana Ramabhadran, Olivier Siohan, Geoffrey Zweig |
INTERSPEECH | 2 |
| 2004 | Speech recognition error analysis on the English MALACH corpusabstractThis paper presents an analysis of the word recognition error rate on an English subset of the MALACH corpus. The MALACH project is an NSF-funded research program related to the development of multilingual access to large audio archives. The archive of interest is a large collection of testimonies from 52,000 survivors, liberators, rescuers and witnesses of the Nazi Holocaust, assembled by the Shoah Visual History Foundation. This data has some unique characteristics that make it quite unusual in the speech recognition community such as elderly speech, noisy conditions, heavily accented speech. Hence, it is a challenging task for automatic speech recognition (ASR). This paper attempts to identify the factors affecting the ASR performance on that task. It was found that the signal-to-noise ratio and syllable Olivier Siohan, Bhuvana Ramabhadran, Geoffrey Zweig |
INTERSPEECH | 1 |
| 2004 | Sequential estimation with optimal forgetting for robust speech recognitionabstractMismatch is known to degrade the performance of speech recognition systems. In real life applications we often encounter nonstationary mismatch sources. A general way to compensate for slowly time varying mismatch is by using sequential algorithms with forgetting. The choice of the forgetting factor is usually performed empirically on some development data, and no optimality criterion is used. In this paper we introduce a framework for obtaining optimal forgetting factor. In sequential algorithms, a recursion is usually used to calculate the required parameters so as to optimize a certain performance measure. To obtain optimal forgetting, we develop a recursion to calculate the forgetting factor that optimizes the same performance criterion as done in the original recursion. When combined together the two recursions result in a sequential algorithm that simultaneously optimizes the desired parameters and the forgetting factor. The proposed method is applied in conjunction with a sequential noise estimation algorithm, but the same principle can be extended to a wide range of sequential algorithms. The algorithm is extensively evaluated for different speech recognition tasks: the 5K Wall Street Journal task corrupted by different types of artificially added noise, a command and digit database recorded in a noisy car environment, and a 20K Japanese broadcast news task corrupted by field noise. In all situations it was found that the sequential algorithm performs as well as or better than batch estimation. In addition, the proposed optimal forgetting algorithm performs as well as the best hand tuned forgetting factor. This results in a continuously adaptive compensation technique without the need of any manual adjustment. Mohamed Afify, Olivier Siohan |
IEEE Trans. Speech Audio Process. | 2 |
| 2003 | Combining neighboring filter channels to improve quantile based histogram equalizationabstractA mismatch between the training data and the test condition of an automatic speech recognition system usually deteriorates the recognition performance. Quantile based histogram equalization can increase the system's robustness by approximating the cumulative density function of the current signal and then reducing an eventual mismatch based on this estimate. In a first step each output of the mel scaled filter bank can be transformed independent from the others. This paper describes an improved version of the algorithm that combines neighboring filter channels. On several databases recorded in real car environment the recognition error rates could be significantly reduced with this new approach. Florian Hilger, Hermann Ney, Olivier Siohan, Frank K. Soong |
ICASSP (1) | 3 |
| 2003 | Hierarchical class n-gram language models: towards better estimation of unseen events in speech recognition
Imed Zitouni, Olivier Siohan |
INTERSPEECH | 2 |
| 2002 | A discriminative training criterion and an associated EM learning algorithmabstractIn this paper we propose a new discriminative training criterion. The criterion is based on a probabilistic interpretation of the minimum classification error (MCE) training and on a modified objective function. A sequential expectation-maximization (EM) based algorithm is derived for optimizing the proposed criterion as an alternative to the well known generalized probabilistic descent (GPD) algorithm. Several variants of the training algorithm are also suggested. The superiority of all variants is shown on a oral/nasal vowel classification task and on a speech/music discrimination task in broadcast news, compared to a maximum likelihood-based system, Mohamed Afify, Olivier Siohan |
ICASSP | 2 |
| 2002 | A dynamic in-search discriminative training approach for large vocabulary speech recognitionabstractIn this paper, we propose a dynamic in-search discriminative training approach of a large-scale HMM model for large vocabulary speech recognition. A previously proposed data selection method is used to choose competing hypotheses dynamically during Viterbi beam search procedure. Particularly, all active word-ending paths are examined during search with reference transcription to identify competing tokens for different HMM's. Then HMMs are re-estimated based on an GPD-based discriminative training to minimize total number of possible error tokens among all collected competing tokens. In this way, recognition errors, e.g., word error rate, in training data can be reduced indirectly. The proposed approach is flexible enough to run in a batch or incremental mode. Also, the method can efficiently be implemented to process large amount of training data and update a large-scale state-tied HMM: set for large vocabulary recognition tasks. Some preliminary results on DARPA communicator task show the new discriminative training method can improve recognition performance over our best ML-trained system. Hui Jiang 0001, Olivier Siohan, Frank K. Soong |
ICASSP | 2 |
| 2002 | Towards knowledge-based features for HMM based large vocabulary automatic speech recognitionabstractThis paper describes an attempt to design a knowledge-based large vocabulary speech recognition system. Our motivation is to replace features based on the short-term spectra, such as Mel-frequency cepstral coefficients (MFCC), by features that explicitly represent some of the distinctive features of the speech signal. However, rather than attempting to compute acoustic correlates of these distinctive features, we have engineered an approach where neural networks are trained to map short-term spectral features to the posterior probability of some distinctive features. These probabilities are then used as features in a large vocabulary tied-state HMM-based recognizer. Experimental results on the Wall Street Journal Task show that such a system, while not outperforming a MFCC-based system, generates very different error patterns. After combining the results of a base-line MFCC system with the results of several systems based on the proposed approach, we were able to obtain reductions in word error rates of 19% and 10 % on the 5K and 20K tasks respectively over our best MFCC-based systems. Benoit Launay, Olivier Siohan, Arun C. Surendran |
ICASSP | 2 |
| 2002 | Bell labs approach to Aurora evaluation on connected digit recognitionabstractABSTRACTIn this paper we study various front-endfeatures, mod-eling and adaptation algorithms on the Aurora 3 databases,including auditory, moment, and AM-FMmodulation fea-tures, context-dependentdigit models, segmental K-meanstraining, discriminative training, and model adaptations.The evaluation results on Aurora 3 are presented with abrief summary of our Aurora 2 results.1. INTRODUCTIONThe Aurora evaluation is for researchers to test their algo-rithms on noise robustness and compare results measuredon the same databases. So far, there are two tasks on theAurora evaluation, Aurora 2 and 3, both are for connecteddigit recognition. While the Aurora 2 databases use thecontrolled experiments by adding noise digitally to cleanEnglish digit strings [1], the Aurora 3 databases are col-lected in a real-worldcar environment in 4 languages. Inthis paper, we report our evaluation results on two of thelanguages, Spanish and German.2. BELL LABS APPROACHESIn this section, we present our baseline system then describethe different feature sets that have been used for this eval-uation. Alternative training strategies and acoustic modeladaptation techniques are also reviewed.A. Context-DependentModel: Similar to last year ap-proach [1], we have decided to use context-dependent(CD)digit models, together with Bell Labs recognition engine asbackend. This contrasts with the officialAurora backendthat is based on whole-worddigit models and the HTK en-gine. The officialbackend setup typically leads to poorerresults, especially in larger databases, and we believe that abetter baseline is beneficialto properly study the effect ofdifferent front-endson the finalrecognition performance.Last year, we investigated several approaches to buildCD digit models. Given the limited amount of trainingdata, especially in the Aurora3 databases, it is required torely on some tying techniques to build CD digit models.The Head-Body-Tail digit model structure (HBT) assumesthat CD digit models are built by concatenating a left-context-dependentunit (head) with a context-independentunit(body)followedbyaright-context-dependentunit(tail).For example, assuming that the lexicon contains 10 digitsplus a silence model, each digit model consists of a set of 1body, 11 heads and 11 tails (representing all left/right con-texts) [2]. We typically model each head and tail with a3-stateHMM, while a 4-stateHMM is used for each body.Most of the experiments done this year have been based onthe HBT structure. CD digit models can also be built as tri-phone modelsusing a decision tree. This is the approach weintroduced last year [1], and some of this year experimentshave been carried out using this model topology.B. Auditory Feature: The new auditory front-endinour recognition system was developed to mimic the robusthuman hearing in adverse acoustic environments [3, 4]. Inthe front-end,efficientsignal processing functions were im-plemented to satisfy both real-timeand computation costrequirements. Based on the analysis of the outer and mid-dle ear, a transfer function was constructed to replace thecommonly used preemphasis filter, and then a new set ofdigital auditory filters,which simulate auditory filteringinthe cochlea, replaces those used in the MFCC and PLP.The auditory feature extraction procedure consists of: anouter-middle-eartransfer function, FFT, frequency conver-sion from linear to the Bark scale, auditory filtering,non-linearity,and discrete cosine transform (DCT). In our previ-ous study[3], the feature has been evaluated in two tasks:connected-digit and large vocabulary, continuous speechrecognitionundervariousnoiseconditions,usingbothhand-setand hands-freedatainlandlineand wirelesstransmissionwith additive car and babble noise. Compared with theLPCC, MFCC, MEL-LPCC,and PLP features, the audi-tory feature achieved significantperformance improvement Jingdong Chen, Dimitris Dimitriadis, Hui Jiang 0001, Tor André Myrvoll, Olivier Siohan, Frank K. Soong |
INTERSPEECH | 6 |
| 2002 | Backoff hierarchical class n-gram language modelling for automatic speech recognition systems
Imed Zitouni, Olivier Siohan, Hong-Kwang Jeff Kuo |
INTERSPEECH | 2 |
| 2002 | Structural maximum a posteriori linear regression for fast HMM adaptation
Olivier Siohan, Tor André Myrvoll |
Comput. Speech Lang. | 1 |
| 2002 | Upper and lower bounds on the mean of noisy speech: application to minimax classificationabstractIn this paper, we derive upper and lower bounds on the mean of speech corrupted by additive noise. The bounds are derived in the log spectral domain. Also approximate bounds on the first and second order time derivatives are developed. It is also shown how to transform these bounds to the mel frequency cepstral coefficient (MFCC) domain. The proposed bounds are used to define the mismatch neighborhood for minimax classification. It is shown that this parametric neighborhood works quite well for artificially added noise and for a real-life mismatch scenario (moving car environment) which does not fully conform with the theoretical conditions used to derive the bounds. In contrast to traditional neighborhood structure for minimax classification, no empirical tuning of the bounds is required. It is believed that the applicability of the derived bounds is not limited to a minimax setting and can be potentially used to develop various compensation scenarios in the log spectral domain. Mohamed Afify, Olivier Siohan |
IEEE Trans. Speech Audio Process. | 2 |
| 2001 | Sequential noise estimation with optimal forgetting for robust speech recognitionabstractMismatch is known to degrade the performance of speech recognition systems. In real life applications mismatch is usually nonstationary, and a general way to compensate for slowly time varying mismatch is by using sequential algorithms with forgetting. The choice of forgetting factor is usually performed empirically on some development data, and no optimality criterion is used. We introduce a framework for obtaining the optimal forgetting factor. The proposed method is applied in conjunction with a sequential noise estimation algorithm, but can be extended to sequential bias or affine transformation estimation. Speech recognition experiments conducted first under a controlled scenario on the 5K Wall Street Journal task corrupted by different noise types, then under a real-life scenario on speech recorded in a noisy car environment validate the proposed method. Mohamed Afify, Olivier Siohan |
ICASSP | 2 |
| 2001 | Evaluating the Aurora connected digit recognition task - a bell labs approach
Mohamed Afify, Hui Jiang 0001, Filipp Korkmazskiy, Olivier Siohan, Frank K. Soong, Arun C. Surendran |
INTERSPEECH | 6 |
| 2001 | Minimax classification with parametric neighborhoods for noisy speech recognition
Mohamed Afify, Olivier Siohan |
INTERSPEECH | 2 |
| 2001 | An auditory system-based feature for robust speech recognitionabstractAn auditory feature extraction algorithm for robust speech recognition in adverse acoustic environments is presented. The feature computation is comprised of an outer-middle-ear transfer function, FFT, frequency conversion from linear to the Bark scale, auditory filtering, nonlinearity, and discrete cosine transform. The feature is evaluated in two tasks: connected-digit recognition and large vocabulary continuous speech recognition. The tested data were under various noise conditions, including handset and hands-free speech data in landline and wireless communications with additive car and babble noise. Compared with the LPCC, MFCC, MEL-LPCC, and PLPfeatures, the proposed feature has an average 20 % to 30 % string error rate reduction on the connected-digit task, and 8 % to 14 % word error rate reduction on the Wall Street Journal task in various additive noise conditions. 1. Frank K. Soong, Olivier Siohan |
INTERSPEECH | 3 |
| 2001 | A new verification-based fast match approach to large vocabulary speech recognition
Mohamed Afify, Hui Jiang 0001, Olivier Siohan |
INTERSPEECH | 4 |
| 2001 | A real-time Japanese broadcast news closed-captioning systemabstractThis paper describes a collaboration between Bell Labs and NHK (Japan Broadcasting Corp.) STRL to develop a real-time large vocabulary speech recognition system for live closed-captioning of NHK news programs. Bell Labs broadcast news recognition engine consists of a two-pass decoder using bigram language models (LM) and right biphone models during the first pass, and trigram LM with within-word triphone models in the second pass. Various pruning strategies are used to achieve real time decoding, together with a noise compensation procedure aimed at improving recognition on noisy segments of the program. The system operates in a real-time mode and delivers less than 2% of word error rate (WER) on studio news conditions and about 5% of WER on noisy news and reporter speech when evaluated on a real broadcast news program. Olivier Siohan, Akio Ando, Mohamed Afify, Hui Jiang 0001, Kazuo Onoe, Frank K. Soong, Qiru Zhou |
INTERSPEECH | 1 |
| 2001 | Joint maximum a posteriori adaptation of transformation and HMM parametersabstractModel adaptation techniques are an efficient way to reduce the mismatch that typically occurs between the training and test condition of any speech recognizer. Adaptation techniques can usually be divided into two families of approaches. On one hand, direct model adaptation attempts to directly reestimate the model parameters, for example using MAP adaptation. Since direct adaptation only reestimates model parameters of the corresponding units appearing in the adaptation data, a large amount of such data is needed to observe any significant improvement in performance. However, nice asymptotic properties are usually observed, meaning that the performance improves as the amount of adaptation data increases. On the other hand, indirect model adaptation applies a general transformation on some clusters of model parameters. Because each individual model is transformed, the approach is quite effective when a small amount of adaptation data is available. However, as the amount of adaptation data increases, the performance improvement quickly saturates. We propose to jointly estimate model parameters and transformation parameters using a single estimation criterion based on Bayesian statistics. We show that by providing a prior distribution for the model parameters and the transformation parameters, it is possible to jointly estimate these two sets of parameters using maximum a posteriori estimation (MAP). Experimental evaluation on nonnative speaker and channel adaptation illustrates the effectiveness of the proposed approach. Olivier Siohan, Cristina Chesta |
IEEE Trans. Speech Audio Process. | 1 |
| 2000 | Multiple classifiers by constrained minimizationabstractThe paper describes an approach to combining multiple classifiers in order to improve classification accuracy. Since individual classifiers in the ensemble should somehow be uncorrelated to yield higher classification accuracy than a single classifier, we propose to train classifiers by minimizing the correlation between their classification errors. A simple combination strategy for three classifiers is then proposed and its achievable error rate is analyzed and compared to individual single classifier performance. The proposed approach has been evaluated on artificial data and a nasal/oral vowel classification task. Theoretical analyses and experimental results illustrate the effectiveness of the proposed approach. Partha Niyogi, Jean-Benoît Pierrot, Olivier Siohan |
ICASSP | 3 |
| 2000 | Joint maximum a posteriori estimation of transformation and hidden Markov model parametersabstractModel adaptation techniques can usually be divided into indirect and direct approaches. On one hand, indirect or transformation-based techniques assume that a general transformation shared amongst different acoustic units is applied to clusters of model parameters. Such approaches (e.g. MLLR-maximum likelihood linear regression) are quite efficient when the amount of adaptation data is limited, but have poor asymptotic properties as the amount of adaptation data increases. On the other hand, direct adaptation approaches, like maximum a posteriori (MAP) estimation have nice asymptotic properties but provide only a moderate improvement when the amount of adaptation data is small. In this work, we jointly optimize a direct and indirect adaptation to take advantage of both approaches. Contrary to published approaches where direct and indirect adaptation are performed one after the other with a very loose interaction and no joint estimation criterion, we propose to estimate a MLLR-like transformation as well as the HMM mean vectors simultaneously, using a MAP estimation criterion. The optimal interaction between the direct and indirect adaptation associated with the prior knowledge provided by the MAP criterion leads to improvement over MLLR and MAP for all size of adaptation data evaluated. Olivier Siohan, Cristina Chesta |
ICASSP | 1 |
| 2000 | Constrained maximum likelihood linear regression for speaker adaptation
Mohamed Afify, Olivier Siohan |
INTERSPEECH | 2 |
| 2000 | Extended maximum a posterior linear regression (EMAPLR) model adaptation for speech recognition
Wu Chou, Olivier Siohan, Tor André Myrvoll |
INTERSPEECH | 2 |
| 2000 | A high-performance auditory feature for robust speech recognitionabstractAn auditory feature extraction algorithm for robust speech recognition in adverse acoustic environments is proposed. Based on the analysis of human auditory system, the feature extraction algorithm consists of several modules: FFT, outer-middle-ear transfer function, frequency conversion from linear to Bark scales, auditory filtering, nonlinearity, and discrete cosine transform. Three recognition experiments have been conducted on connected digit recognition in wireless and land-line communications using handsets and handsfree microphones. Compared to LPCC and MFCC features, the proposed feature has shown 11% to 23% error-rate reductions on average in handset and hands-free acoustic environments in the experiments. Frank K. Soong, Olivier Siohan |
INTERSPEECH | 3 |
| 2000 | Structural maximum a-posteriori linear regression for unsupervised speaker adaptation
Tor André Myrvoll, Olivier Siohan, Wu Chou |
INTERSPEECH | 2 |
| 2000 | Small group speaker identification with common password phrases
Aaron E. Rosenberg, Olivier Siohan, Sarangarajan Parthasarathy |
Speech Commun. | 2 |
| 1999 | Background model design for flexible and portable speaker verification systemsabstractMost state-of-the-art speaker verification systems need a user model built from samples of the customer speech, and a speaker independent (SI) background model with high acoustic resolution. These systems rely heavily on the availability of speaker independent databases along with a priori knowledge about acoustic rules of the utterance, and depend on the consistency of acoustic conditions under which the SI models were trained. These constraints may be a burden in practical and portable devices such as palm-top computers or wireless handsets which place a premium on computation and memory, and where the user is free to choose any password utterance in any language, under any acoustic condition. In this paper, we present a novel and reliable approach to background model design when only the enrollment data is available. Preliminary results are provided to demonstrate the effectiveness of such systems. Olivier Siohan, Arun C. Surendran |
ICASSP | 1 |
| 1999 | Maximum a posteriori linear regression for hidden Markov model adaptationabstractIn the past few years, transformation-based model adaptation techniques have been widely used to help reducing acoustic mismatch between training and testing conditions of automatic speech recognizers. The estimation of the transformation parameters is usually carried out using estimation paradigms based on classical statistics such as maximum likelihood, mainly because of their conceptual and computational simplicity. However, it appears necessary to introduce some constraints on the possible values of the transformation parameters to avoid getting unreasonable estimates that might perturb the underlying structure of the acoustic space. In this paper, we propose to introduce such constraints using Bayesian statistics, where a prior distribution of the transformation parameters is used. A Bayesian counterpart of the well known maximum likelihood linear regression (MLLR) adaption is formulated based on maximum a posteriori (MAP) estimation. Supervised, unsupervised and incremental non-native speaker adaptation experiments are carried out to compare the proposed MAPLR approach to MLLR. Experimental results show that MAPLR outperforms MLLR. Cristina Chesta, Olivier Siohan |
EUROSPEECH | 2 |
| 1998 | Speaker verification using minimum verification error trainingabstractWe propose a minimum verification error (MVE) training scenario to design and adapt an HMM-based speaker verification system. By using the discriminative training paradigm, we show that customer and background models can be jointly estimated so that the expected number of verification errors (false accept and false reject) on the training corpus are minimized. An experimental evaluation of a fixed password speaker verification task over the telephone network was carried out. The evaluation shows that MVE training/adaptation performs as well as MLE training and MAP adaptation when the performance is measured by the average individual equal error rate (based on a posteriori threshold assignment). After model adaptation, both approaches lead to an individual equal error-rate close to 0.6%. However, experiments performed with a priori dynamic threshold assignment show that MVE adapted models exhibit false rejection and false acceptance rates 45% lower than the MAP adapted models, and therefore lead to the design of a more robust system for practical applications. Aaron E. Rosenberg, Olivier Siohan, Sarangarajan Parthasarathy |
ICASSP | 2 |
| 1998 | Speaker identification using minimum classification error trainingabstractWe use a minimum classification error (MCE) training paradigm to build a speaker identification system. The training is optimized at the string level for a text-dependent speaker identification task. Experiments performed on a small set speaker identification task show that MCE training can reduce closed-set identification errors by up to 20-25% over a baseline system trained using maximum likelihood estimation. Further experiments suggest that additional improvement can be obtained by using some additional training data from speakers outside the set of registered speakers, leading to an overall reduction of the closed-set identification errors by about 35%. Olivier Siohan, Aaron E. Rosenberg, Sarangarajan Parthasarathy |
ICASSP | 1 |
| 1997 | Iterative noise and channel estimation under the stochastic matching algorithm frameworkabstractIn this letter, we introduce an unsupervised iterative algorithm to adapt HMMs trained using clean speech in order to recognize speech corrupted by an additive and a convolutional noise. Both types of noise are considered as stochastic processes that can be modeled using HMMs and can be estimated by applying Sankar's stochastic matching (SM) algorithm successively in the cepstral and in the linear spectral domain. These estimates are derived directly from the given test speech signal and the set of clean speech models, and lead to the estimation of a new set of HMMs that maximize the likelihood of the test signal. Olivier Siohan |
IEEE Signal Process. Lett. | 1 |
| 1996 | A semi-continuous stochastic trajectory model for phoneme-based continuous speech recognitionabstractWe propose a model of phoneme-based speech unit, called semi-continuous stochastic trajectory model (SC-STM), which generalizes our stochastic trajectory models (STM). As STMs, the SC-STMs focus on the modeling of speech segments (called trajectories) in their parameter space, and can therefore handle segmental information, which is critical for large vocabulary continuous speech recognition. Compared to the STMs, the SC-STMs improve the resolution of the trajectory modeling, while keeping a moderate number of free parameters by sharing state probability density functions. The SC-STM can therefore maintain a good trade-off between detailed acoustic modeling and limited training data. We tested the idea on a 2010 words, speaker-dependent, continuous speech database. Preliminary results show that SC-STM gives a word accuracy close to that of STM, without using heuristic techniques that enhanced STM. Olivier Siohan, Yifan Gong 0001 |
ICASSP | 1 |
| 1996 | Comparative experiments of several adaptation approaches to noisy speech recognition using stochastic trajectory models
Olivier Siohan, Yifan Gong 0001, Jean Paul Haton |
Speech Commun. | 1 |
| 1995 | On the robustness of linear discriminant analysis as a preprocessing step for noisy speech recognitionabstractThis paper addresses the problem of speech recognition in a noisy environment by finding a robust speech parametric space. The framework of linear discriminant analysis (LDA) is used to derive an efficient speech parametric space for noisy speech recognition, from a classical static+dynamic MFCC space. We first show that the derived LDA space can lead to a higher discrimination than the MFCC related space, even at low signal-to-noise ratio (SNR). Then, we test the robustness of the LDA space to variations between the training and testing SNR. Experiments are performed on a continuous speech recognition task, where speech is degraded with various noise sources: Gaussian noise, F16, Lynx helicopter, autobus, hair dryer. It was found that LDA is highly sensitive to SNR variations for white noise (Gaussian, hair dryer), while remaining quite efficient for the others. Olivier Siohan |
ICASSP | 1 |
| 1995 | Noise adaptation using linear regression for continuous noisy speech recognition
Olivier Siohan, Yifan Gong 0001, Jean Paul Haton |
EUROSPEECH | 1 |
| 1994 | A comparison of three noisy speech recognition approaches
Olivier Siohan, Yifan Gong 0001, Jean Paul Haton |
ICSLP | 1 |
| 1993 | A Bayesian approach to phone duration adaptation for lombard speech recognition
Olivier Siohan, Yifan Gong 0001, Jean Paul Haton |
EUROSPEECH | 1 |
| 1992 | Minimization of speech alignment error by iterative transformation for speaker adaptationabstractExtrait de : Proc. International Conference on Spoken Language Processing, Banff (Alberta, Canada), October 1992 Yifan Gong 0001, Olivier Siohan, Jean Paul Haton |
ICSLP | 2 |