Sibo Tong

dblp:180/2576 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
6since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-author · 5 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2023 Hierarchical Attention-Based Contextual Biasing For Personalized Speech Recognition Using Neural Transducers
abstract
Although end-to-end (E2E) automatic speech recognition (ASR) systems excel in general tasks, they frequently struggle with accurately recognizing personal rare words. Leveraging contextual information to bias the internal states of E2E ASR model has proven to be an effective solution. However most existing work focuses on biasing for a single domain and it is still challenging to expand such contextualization mechanisms to many domains. To address this limitation, in this work we propose a hierarchical attention architecture to scale contextual biasing to a wide range of domains simultaneously. Given multiple catalogs of contextual information, the high-level attention determines which source of catalog to focus on and the low-level attention learns to attend to the most relevant entity within the focused catalog. Experiments on diverse domains demonstrate the proposed architecture results in $35 \%$ to $60 \%$ relative WER improvements on personal rare words and outperforms existing approaches.
Sibo Tong, Philip Harding, Simon Wiesler
ASRU1
2023 Slot-Triggered Contextual Biasing For Personalized Speech Recognition Using Neural Transducers
abstract
End-to-end (E2E) automatic speech recognition (ASR) models have been found to perform well on general transcription tasks but often fail to correctly recognize words that occur infrequently in the training data. Personalization is important for a variety of tasks, including virtual assistants where recall of infrequently observed words such as contact names, song titles and place names is critical. In these cases contextual information is often available which can be used to bias the E2E ASR model. Contextual biasing (CB) has been shown to be effective for this task, however most existing work focuses on biasing for a single domain and so in this work we focus on the application of biasing to multiple domains. We propose a method whereby the E2E ASR model is trained to emit opening and closing tags around slot content which are used to both selectively enable biasing and decide which catalog to use for biasing. Our method is shown to not only efficiently scale to multiple slots, but also further improves accuracy on slot content.
Sibo Tong, Philip Harding, Simon Wiesler
ICASSP1
2023 Selective Biasing with Trie-based Contextual Adapters for Personalised Speech Recognition using Neural Transducers
Philip Harding, Sibo Tong, Simon Wiesler
INTERSPEECH2
2023 Model-Internal Slot-triggered Biasing for Domain Expansion in Neural Transducer ASR Models
Yiting Lu, Philip Harding, Kanthashree Mysore Sathyendra, Sibo Tong, Xuandi Fu, Feng-Ju Chang, Simon Wiesler, Grant P. Strimel
INTERSPEECH4
2023 Effective Training of Attention-based Contextual Biasing Adapters with Synthetic Audio for Personalised ASR
Burin Naowarat, Philip Harding, Pasquale D'Alterio, Sibo Tong, Bashar Awwad Shiekh Hasan
INTERSPEECH4
2021 A Bayesian Approach to Recurrence in Neural Networks
abstract
We begin by reiterating that common neural network activation functions have simple Bayesian origins. In this spirit, we go on to show that Bayes's theorem also implies a simple recurrence relation; this leads to a Bayesian recurrent unit with a prescribed feedback formulation. We show that introduction of a context indicator leads to a variable feedback that is similar to the forget mechanism in conventional recurrent units. A similar approach leads to a probabilistic input gate. The Bayesian formulation leads naturally to the two pass algorithm of the Kalman smoother or forward-backward algorithm, meaning that inference naturally depends upon future inputs as well as past ones. Experiments on speech recognition confirm that the resulting architecture can perform as well as a bidirectional recurrent network with the same number of parameters as a unidirectional one. Further, when configured explicitly bidirectionally, the architecture can exceed the performance of a conventional bidirectional recurrence.
Philip N. Garner, Sibo Tong
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Lattice-Free Maximum Mutual Information Training of Multilingual Speech Recognition Systems
abstract
Multilingual acoustic model training combines data from multiple languages to train an automatic speech recognition system.Such a system is beneficial when training data for a target language is limited.Lattice-Free Maximum Mutual Information (LF-MMI) training performs sequence discrimination by introducing competing hypotheses through a denominator graph in the cost function.The standard approach to train a multilingual model with LF-MMI is to combine the acoustic units from all languages and use a common denominator graph.The resulting model is either used as a feature extractor to train an acoustic model for the target language or directly fine-tuned.In this work, we propose a scalable approach to train the multilingual acoustic model using a typical multitask network for the LF-MMI framework.A set of language-dependent denominator graphs is used to compute the cost function.The proposed approach is evaluated under typical multilingual ASR tasks using GlobalPhone and BABEL datasets.Relative improvements up to 13.2% in WER are obtained when compared to the corresponding monolingual LF-MMI baselines.The implementation is made available as a part of the Kaldi speech recognition toolkit.
Srikanth R. Madikeri, Banriskhem K. Khonglah, Sibo Tong, Petr Motlícek, Hervé Bourlard, Daniel Povey
INTERSPEECH3
2019 An Investigation of Multilingual ASR Using End-to-end LF-MMI
abstract
The end-to-end lattice-free maximum mutual information (LF-MMI) approach has recently been shown to be beneficial for automatic speech recognition (ASR) in general. More specifically, its end-to-end nature and use of context independent phone labels make it attractive for multilingual ASR. We show that end-to-end LF-MMI is indeed competitive on a low-resourced multilingual task, comfortably outperforming a connectionist temporal classification (CTC) baseline. We further investigate the feasibility of biphone contexts, being a candidate compromise between the context independent approach and the triphone contexts that usually perform well. We show that biphones do not initially perform well, but can do so after language adaptive training, concluding that biphones carry language variability but are promising for multilingual ASR.
Sibo Tong, Philip N. Garner, Hervé Bourlard
ICASSP1
2019 Analyzing Uncertainties in Speech Recognition Using Dropout
abstract
The performance of Automatic Speech Recognition (ASR) systems is often measured using Word Error Rates (WER) which requires time-consuming and expensive manually transcribed data. In this paper, we use state-of-the-art ASR systems based on Deep Neural Networks (DNN) and propose a novel framework which uses "Dropout" at the test time to model uncertainty in prediction hypotheses. We systematically exploit this uncertainty to estimate WER without the need for explicit transcriptions. In addition, we show that the predictive uncertainty can also be used to accurately localize the errors made by the ASR system. We study the performance of our approach on Switchboard database where it predicts WER accurately within a range of 2.6% and 5.0% for HMM-DNN and Connectionist Temporal Classification (CTC) ASR systems, respectively.
Apoorv Vyas, Pranay Dighe, Sibo Tong, Hervé Bourlard
ICASSP3
2019 Unbiased Semi-Supervised LF-MMI Training Using Dropout
Sibo Tong, Apoorv Vyas, Philip N. Garner, Hervé Bourlard
INTERSPEECH1
2018 Nasal Speech Sounds Detection Using Connectionist Temporal Classification
abstract
Phone attributes, known also as distinctive or phonological features, belong to important classification of the speech sounds used in automatic speech processing. Training of conventional phone attribute detectors (classifiers), either based on acoustic measurements or deep learning approaches, requires decent phone boundary segmentation. This paper proposes a solution to train a phone attribute detector without phone alignment using an end-to-end phone attribute modeling based on the connectionist temporal classification. Experiments, performed for the nasal phone attribute on the LibriSpeech database, confirm that the proposed system outperforms conventional deep neural network detector, trained even on the same training data. Further improvements are observed with more training data. Conventional complex system that consists of feature extraction, phone force-alignment and deep neural network training is replaced by a more simpler Python package based on PyTorch, released as open-source.
Milos Cernak, Sibo Tong
ICASSP2
2018 Fast Language Adaptation Using Phonological Information
abstract
Phoneme-based multilingual connectionist temporal classification (CTC) model is easily extensible to a new language by concatenating parameters of the new phonemes to the output layer. In the present paper, we improve cross-lingual adaptation in the context of phoneme-based CTC models by using phonological information. A universal (IPA) phoneme classifier is first trained on phonological features generated from a phonological attribute detector. When adapting the multilingual CTC to a new, never seen, language, phonological attributes of the unseen phonemes are derived based on phonology and fed into the phoneme classifier. Posteriors given by the classifier are used to initialize the parameters of the unseen phonemes when extending the multilingual CTC output layer to the target language. Adaptation experiments show that the proposed initialization approaches further improve the cross-lingual adaptation on CTC models and yield significant improvements over Deep Neural Network / Hidden Markov Model (DNN/HMM)-based adaptation using limited data.
Sibo Tong, Philip N. Garner, Hervé Bourlard
INTERSPEECH1
2018 Cross-lingual adaptation of a CTC-based multilingual acoustic model
Sibo Tong, Philip N. Garner, Hervé Bourlard
Speech Commun.1
2017 An Investigation of Deep Neural Networks for Multilingual Speech Recognition Training and Adaptation
abstract
Different training and adaptation techniques for multilingual Automatic Speech Recognition (ASR) are explored in the context of hybrid systems, exploiting Deep Neural Networks (DNN) and Hidden Markov Models (HMM). In multilingual DNN training, the hidden layers (possibly extracting bottleneck features) are usually shared across languages, and the output layer can either model multiple sets of language-specific senones or one single universal IPA-based multilingual senone set. Both architectures are investigated, exploiting and comparing different language adaptive training (LAT) techniques originating from successful DNN-based speaker-adaptation. More specifically, speaker adaptive training methods such as Cluster Adaptive Training (CAT) and Learning Hidden Unit Contribution (LHUC) are considered. In addition, a language adaptive output architecture for IPA-based universal DNN is also studied and tested. Experiments show that LAT improves the performance and adaptation on the top layer further improves the accuracy. By combining state-level minimum Bayes risk (sMBR) sequence training with LAT, we show that a language adaptively trained IPA-based universal DNN outperforms a monolingually sequence trained model.
Sibo Tong, Philip N. Garner, Hervé Bourlard
INTERSPEECH1
2016 A comparative study of robustness of deep learning approaches for VAD
abstract
Voice activity detection (VAD) is an important step for real-world automatic speech recognition (ASR) systems. Deep learning approaches, such as DNN, RNN or CNN, have been widely used in model-based VAD. Although they have achieved success in practice, they are developed on different VAD tasks separately. Whilst VAD performance under noisy conditions, especially with unseen noise or very low SNR, are of great interest, there has no robustness comparison of different deep learning approaches so far. In this paper, to learn the robustness property, VAD models based on DNN, LSTM and CNN are thoroughly compared at both frame and segment level under various noisy conditions on Aurora 4, a commonly used speech corpus with rich noises. To improve the robustness of deep learning based VAD models, a new noise-aware training (NAT) approach is also proposed. Experiments show that LSTM-based VAD is most robust but the performance degrades dramatically in the conditions with unseen noise or diverse SNR. By incorporating NAT, significant performance gains can be obtained in these conditions.
Sibo Tong, Kai Yu 0004
ICASSP1