VLDB 2026 Research / reviewers in the wild / expert
Thomas Niesler
dblp:195/9239 · also Thomas R. Niesler
· DBLP profile ↗
51ranked-venue papers
8as first author
9since 2021 · last 2024
0000-0002-7341-1017ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 33 · 4 first-author · 6 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Automatic Partitioning of a Code-Switched Speech Corpus Using Mixed-Integer ProgrammingabstractDefining training, development and test set partitions for speech corpora is usually accomplished by hand. However, for the dataset under investigation, which contains a large number of speakers, eight different languages and code-switching between all the languages, this style of partitioning is not feasible. Therefore, we view the partitioning task as a resource allocation problem and propose to solve it automatically and optimally by the application of mixed-integer linear programming. Using this approach, we are able to partition a new 41.6-hour multilingual corpus of code-switched speech into training, development and testing partitions while maintaining a fixed number of speakers and a specific amount of code-switched speech in the development and test partitions. For this newly partitioned corpus, we present baseline speech recognition results using a state-of-the-art multilingual transformer model (Wav2Vec2-XLS-R) and show that the exclusion of very short utterances (<1s) results in substantially improved speech recognition performance. Joshua Jansen van Vueren, Febe de Wet, Thomas Niesler |
LREC/COLING | 3 |
| 2024 | A Transformer-Based Voice Activity Detector
Biswajit Karan, Joshua Jansen van Vueren, Febe de Wet, Thomas Niesler |
INTERSPEECH | 4 |
| 2023 | Improving Under-Resourced Code-Switched Speech Recognition: Large Pre-trained Models or Architectural Interventions
Joshua Jansen van Vueren, Thomas Niesler |
INTERSPEECH | 2 |
| 2022 | TB or not TB? Acoustic cough analysis for tuberculosis classificationabstractIn this work, we explore recurrent neural network architectures for tuberculosis (TB) cough classification.In contrast to previous unsuccessful attempts to implement deep architectures in this domain, we show that a basic bidirectional long shortterm memory network (BiLSTM) can achieve improved performance.In addition, we show that by performing greedy feature selection in conjunction with a newly-proposed attention-based architecture that learns patient invariant features, substantially better generalisation can be achieved compared to a baseline and other considered architectures.Furthermore, this attention mechanism allows an inspection of the temporal regions of the audio signal considered to be important for classification to be performed.Finally, we develop a neural style transfer technique to infer idealised inputs which can subsequently be analysed.We find distinct differences between the idealised power spectra of TB and non-TB coughs, which provide clues about the origin of the features in the audio signal. Geoffrey T. Frost, Grant Theron, Thomas Niesler |
INTERSPEECH | 3 |
| 2022 | Code-Switched Language Modelling Using a Code Predictive Lstm in Under-Resourced South African LanguagesabstractWe present a new LSTM language model architecture for code-switched speech incorporating a neural structure that explicitly models language switches. Experimental evaluation of this code predictive model for four under-resourced South African languages shows consistent improvements in perplexity as well as perplexity specifically over code-switches compared to an LSTM baseline. Substantial reductions in absolute speech recognition word error rates (0.5%-1.2%) as well as errors specifically at code-switches (0.6%-2.3%) are also achieved during n-best rescoring. When used for both data augmentation and n-best rescoring, our code predictive model reduces word error rate by a further 0.8%-2.6% absolute and consistently outperforms a baseline LSTM. The similar and consistent trends observed across all four language pairs allows us to conclude that explicit modelling of language switches by a dedicated language model component is a suitable strategy for code-switched speech recognition. Joshua Jansen van Vueren, Thomas Niesler |
SLT | 2 |
| 2022 | Code-switched automatic speech recognition in five South African languages
Astik Biswas, Emre Yilmaz 0001, Ewald van der Westhuizen, Febe de Wet, Thomas Niesler |
Comput. Speech Lang. | 5 |
| 2022 | Feature learning for efficient ASR-free keyword spotting in low-resource languages
Ewald van der Westhuizen, Herman Kamper, Raghav Menon, John A. Quinn, Thomas Niesler |
Comput. Speech Lang. | 5 |
| 2021 | Deep Neural Network Based Cough Detection Using Bed-Mounted Accelerometer MeasurementsabstractWe have performed cough detection based on measurements from an accelerometer attached to the patient’s bed. This form of monitoring is less intrusive than body-attached accelerometer sensors, and sidesteps privacy concerns encountered when using audio for cough detection. For our experiments, we have compiled a manually-annotated dataset containing the acceleration signals of approximately 6000 cough and 68000 non-cough events from 14 adult male patients in a tuberculosis clinic. As classifiers, we have considered convolutional neural networks (CNN), long-short-term-memory (LSTM) networks, and a residual neural network (Resnet50). We find that all classifiers are able to distinguish between the acceleration signals due to coughing and those due to other activities including sneezing, throat-clearing and movement in the bed with high accuracy. The Resnet50 performs the best, achieving an area under the ROC curve (AUC) exceeding 0.98 in cross-validation experiments. We conclude that high-accuracy cough monitoring based only on measurements from the accelerometer in a consumer smartphone is possible. Since the need to gather audio is avoided and therefore privacy is inherently protected, and since the accelerometer is attached to the bed and not worn, this form of monitoring may represent a more convenient and readily accepted method of long-term patient cough monitoring. Madhurananda Pahar, Igor D. S. Miranda, Andreas H. Diacon, Thomas Niesler |
ICASSP | 4 |
| 2021 | A Hybrid CNN-BiLSTM Voice Activity DetectorabstractThis paper presents a new hybrid architecture for voice activity detection (VAD) incorporating both convolutional neural network (CNN) and bidirectional long short-term memory (BiLSTM) layers trained in an end-to-end manner. In addition, we focus specifically on optimising the computational efficiency of our architecture in order to deliver robust performance in difficult in-the-wild noise conditions in a severely under-resourced setting. Nested k-fold cross-validation was used to explore the hyperparameter space, and the trade-off between optimal parameters and model size is discussed. The performance effect of a BiLSTM layer compared to a unidirectional LSTM layer was also considered. We compare our systems with three established baselines on the AVA-Speech dataset. We find that significantly smaller models with near optimal parameters perform on par with larger models trained with optimal parameters. BiLSTM layers were shown to improve accuracy over unidirectional layers by ≈2% absolute on average. With an area under the curve (AUC) of 0.951, our system outperforms all baselines, including a much larger ResNet system, particularly in difficult noise conditions. Nicholas Wilkinson, Thomas Niesler |
ICASSP | 2 |
| 2020 | Interactive Image Exploration for Visually Impaired Readers using Audio-augmented Touch GesturesabstractTechnologies such as text to speech and hardware braille displays provide alternative representations of electronic text. As a result, electronic documents provide increased accessibility to print-disabled users. However, non-textual graphical information in electronic documents remain largely inaccessible to the blind population. In this study, we explore audio-visual sensory substitution as a means of rendering graphical information present in electronic documents to blind users. To achieve this, we have extended the audio rendering approach used by the well-established vOICe algorithm to allow interactive and localised exploration of an image by means of gestures and the touch screen of a standard commercially-available tablet. The effectiveness of our approach was evaluated in a set of user trials that required six sighted and six blind subjects to identify elements of scenes consisting of a number of geometrical shapes and emoticons. Our results show that both groups of subjects were more successful at identifying shapes using the interactive algorithm than with the baseline vOICe algorithm to a highly statistically significant degree. Furthermore, the results indicate that this improvement is greatest for the most complex scenes. We conclude that, by introducing an interactive touch interface, the vOICe algorithm can be successfully extended to allow interactive exploration and interpretation of diagrams, thereby improving accessibility to material such as scientific publications. Rynhardt Kruger, Febe de Wet, Thomas Niesler |
IV | 3 |
| 2020 | Semi-supervised Development of ASR Systems for Multilingual Code-switched Speech in Under-resourced LanguagesabstractThis paper reports on the semi-supervised development of acoustic and language models for under-resourced, code-switched speech in five South African languages. Two approaches are considered. The first constructs four separate bilingual automatic speech recognisers (ASRs) corresponding to four different language pairs between which speakers switch frequently. The second uses a single, unified, five-lingual ASR system that represents all the languages (English, isiZulu, isiXhosa, Setswana and Sesotho). We evaluate the effectiveness of these two approaches when used to add additional data to our extremely sparse training sets. Results indicate that batch-wise semi-supervised training yields better results than a non-batch-wise approach. Furthermore, while the separate bilingual systems achieved better recognition performance than the unified system, they benefited more from pseudolabels generated by the five-lingual system than from those generated by the bilingual systems. Astik Biswas, Emre Yilmaz 0001, Febe de Wet, Ewald van der Westhuizen, Thomas Niesler |
LREC | 5 |
| 2020 | Practical evaluation of carrier sensing for a LoRa wildlife monitoring network
Morgan O'Kennedy, Thomas Niesler, Riaan Wolhuter, Nathalie Mitton |
Networking | 2 |
| 2019 | Improving Automatically Induced Lexicons for Highly Agglutinating Languages Using Data-Driven Morphological Segmentation
Wiehan Agenbag, Thomas Niesler |
INTERSPEECH | 2 |
| 2019 | Improved Low-Resource Somali Speech Recognition by Semi-Supervised Acoustic and Language Model TrainingabstractWe present improvements in automatic speech recognition (ASR) for Somali, a currently extremely under-resourced language. This forms part of a continuing United Nations (UN) effort to employ ASR-based keyword spotting systems to support humanitarian relief programmes in rural Africa. Using just 1.57 hours of annotated speech data as a seed corpus, we increase the pool of training data by applying semi-supervised training to 17.55 hours of untranscribed speech. We make use of factorised time-delay neural networks (TDNN-F) for acoustic modelling, since these have recently been shown to be effective in resource-scarce situations. Three semi-supervised training passes were performed, where the decoded output from each pass was used for acoustic model training in the subsequent pass. The automatic transcriptions from the best performing pass were used for language model augmentation. To ensure the quality of automatic transcriptions, decoder confidence is used as a threshold. The acoustic and language models obtained from the semi-supervised approach show significant improvement in terms of WER and perplexity compared to the baseline. Incorporating the automatically generated transcriptions yields a 6.55\% improvement in language model perplexity. The use of 17.55 hour of Somali acoustic data in semi-supervised training shows an improvement of 7.74\% relative over the baseline. Astik Biswas, Raghav Menon, Ewald van der Westhuizen, Thomas Niesler |
INTERSPEECH | 4 |
| 2019 | Semi-Supervised Acoustic Model Training for Five-Lingual Code-Switched ASRabstractThis paper presents recent progress in the acoustic modelling of under-resourced code-switched (CS) speech in multiple South African languages. We consider two approaches. The first constructs separate bilingual acoustic models corresponding to language pairs (English-isiZulu, English-isiXhosa, English-Setswana and English-Sesotho). The second constructs a single unified five-lingual acoustic model representing all the languages (English, isiZulu, isiXhosa, Setswana and Sesotho). For these two approaches we consider the effectiveness of semi-supervised training to increase the size of the very sparse acoustic training sets. Using approximately 11 hours of untranscribed speech, we show that both approaches benefit from semi-supervised training. The bilingual TDNN-F acoustic models also benefit from the addition of CNN layers (CNN-TDNN-F), while the five-lingual system does not show any significant improvement. Furthermore, because English is common to all language pairs in our data, it dominates when training a unified language model, leading to improved English ASR performance at the expense of the other languages. Nevertheless, the five-lingual model offers flexibility because it can process more than two languages simultaneously, and is therefore an attractive option as an automatic transcription system in a semi-supervised training pipeline. Astik Biswas, Emre Yilmaz 0001, Febe de Wet, Ewald van der Westhuizen, Thomas Niesler |
INTERSPEECH | 5 |
| 2019 | Feature Exploration for Almost Zero-Resource ASR-Free Keyword Spotting Using a Multilingual Bottleneck Extractor and Correspondence AutoencodersabstractWe compare features for dynamic time warping (DTW) when used to bootstrap keyword spotting (KWS) in an almost zero-resource setting. Such quickly-deployable systems aim to support United Nations (UN) humanitarian relief efforts in parts of Africa with severely under-resourced languages. Our objective is to identify acoustic features that provide acceptable KWS performance in such environments. As supervised resource, we restrict ourselves to a small, easily acquired and independently compiled set of isolated keywords. For feature extraction, a multilingual bottleneck feature (BNF) extractor, trained on well-resourced out-of-domain languages, is integrated with a correspondence autoencoder (CAE) trained on extremely sparse in-domain data. On their own, BNFs and CAE features are shown to achieve a more than 2% absolute performance improvement over baseline MFCCs. However, by using BNFs as input to the CAE, even better performance is achieved, with a more than 11% absolute improvement in ROC AUC over MFCCs and more than twice as many top-10 retrievals for two evaluated languages, English and Luganda. We conclude that integrating BNFs with the CAE allows both large out-of-domain and sparse in-domain resources to be exploited for improved ASR-free keyword spotting. Raghav Menon, Herman Kamper, Ewald van der Westhuizen, John A. Quinn, Thomas Niesler |
INTERSPEECH | 5 |
| 2019 | Automatic sub-word unit discovery and pronunciation lexicon induction for ASR with application to under-resourced languages
Wiehan Agenbag, Thomas Niesler |
Comput. Speech Lang. | 2 |
| 2019 | Synthesised bigrams using word embeddings for code-switched ASR of four South African language pairs
Ewald van der Westhuizen, Thomas Niesler |
Comput. Speech Lang. | 2 |
| 2018 | Multilingual Neural Network Acoustic Modelling for ASR of Under-Resourced English-isiZulu Code-Switched SpeechabstractThe 19th Annual Conference of the International Speech Communication Association (Interspeech), 2 september 2018 Astik Biswas, Febe de Wet, Ewald van der Westhuizen, Emre Yilmaz 0001, Thomas Niesler |
INTERSPEECH | 5 |
| 2018 | Fast ASR-free and Almost Zero-resource Keyword Spotting Using DTW and CNNs for Humanitarian MonitoringabstractWe use dynamic time warping (DTW) as supervision for training a convolutional neural network (CNN) based keyword spotting system using a small set of spoken isolated keywords. The aim is to allow rapid deployment of a keyword spotting system in a new language to support urgent United Nations (UN) relief programmes in parts of Africa where languages are extremely under-resourced and the development of annotated speech resources is infeasible. First, we use 1920 recorded keywords (40 keyword types, 34 minutes of speech) as exemplars in a DTW-based template matching system and apply it to untranscribed broadcast speech. Then, we use the resulting DTW scores as targets to train a CNN on the same unlabelled speech. In this way we use just 34 minutes of labelled speech, but leverage a large amount of unlabelled data for training. While the resulting CNN keyword spotter cannot match the performance of the DTW-based system, it substantially outperforms a CNN classifier trained only on the keywords, improving the area under the ROC curve from 0.54 to 0.64. Because our CNN system is several orders of magnitude faster at runtime than the DTW system, it represents the most viable keyword spotter on this extremely limited dataset. Raghav Menon, Herman Kamper, John A. Quinn, Thomas Niesler |
INTERSPEECH | 4 |
| 2018 | Building a Unified Code-Switching ASR System for South African LanguagesabstractWe present our first efforts towards building a single multilingual automatic speech recognition (ASR) system that can process code-switching (CS) speech in five languages spoken within the same population. This contrasts with related prior work which focuses on the recognition of CS speech in bilingual scenarios. Recently, we have compiled a small five-language corpus of South African soap opera speech which contains examples of CS between 5 languages occurring in various contexts such as using English as the matrix language and switching to other indigenous languages. The ASR system presented in this work is trained on 4 corpora containing English-isiZulu, English-isiXhosa, English-Setswana and English-Sesotho CS speech. The interpolation of multiple language models trained on these language pairs enables the ASR system to hypothesize mixed word sequences from these 5 languages. We evaluate various state-of-the-art acoustic models trained on this 5-lingual training data and report ASR accuracy and language recognition performance on the development and test sets of the South African multilingual soap opera corpus. Emre Yilmaz 0001, Astik Biswas, Ewald van der Westhuizen, Febe de Wet, Thomas Niesler |
INTERSPEECH | 5 |
| 2018 | A First South African Corpus of Multilingual Code-switched Soap Opera Speech
Ewald van der Westhuizen, Thomas Niesler |
LREC | 2 |
| 2017 | Radio-browsing for developmental monitoring in UgandaabstractWe consider the extraction of information from broadcast radio speech in Uganda for the purposes of informing relief and development programmes by the United Nations. Although Internet penetration in Uganda is low, mobile phones are ubiquitous and have made radio a vibrant medium for interactive public discussion. Vulnerable groups make use of radio to discuss issues related to, for example, agriculture, health, governance and gender by means of phone-in or text-in talk shows. We have compiled corpora and developed a radio-browsing system for Ugandan English and for two indigenous languages, Luganda and Acholi. The systems employ automatic speech recognisers using HMM/GMM, SGMM and DNN/HMM acoustic models as keyword spotters. We present the first results indicating promising performance of the radio-browsing system. Raghav Menon, Armin Saeb, Hugh Cameron, William Kibira, John A. Quinn, Thomas Niesler |
ICASSP | 6 |
| 2017 | Very Low Resource Radio Browsing for Agile Developmental and Humanitarian Monitoring
Armin Saeb, Raghav Menon, Hugh Cameron, William Kibira, John A. Quinn, Thomas Niesler |
INTERSPEECH | 6 |
| 2017 | Synthesising isiZulu-English Code-Switch Bigrams Using Word Embeddings
Ewald van der Westhuizen, Thomas Niesler |
INTERSPEECH | 2 |
| 2017 | Virtual reality assisted microscopy data visualization and colocalization analysisabstractBACKGROUND: Confocal microscopes deliver detailed three-dimensional data and are instrumental in biological analysis and research. Usually, this three-dimensional data is rendered as a projection onto a two-dimensional display. We describe a system for rendering such data using a modern virtual reality (VR) headset. Sample manipulation is possible by fully-immersive hand-tracking and also by means of a conventional gamepad. We apply this system to the specific task of colocalization analysis, an important analysis tool in biological microscopy. We evaluate our system by means of a set of user trials. RESULTS: The user trials show that, despite inaccuracies which still plague the hand tracking, this is the most productive and intuitive interface. The inaccuracies nevertheless lead to a perception among users that productivity is low, resulting in a subjective preference for the gamepad. Fully-immersive manipulation was shown to be particularly effective when defining a region of interest (ROI) for colocalization analysis. CONCLUSIONS: Virtual reality offers an attractive and powerful means of visualization for microscopy data. Fully immersive interfaces using hand tracking show the highest levels of intuitiveness and consequent productivity. However, current inaccuracies in hand tracking performance still lead to a disproportionately critical user perception. Rensu P. Theart, Ben Loos, Thomas Niesler |
BMC Bioinform. | 3 |
| 2016 | The Effect of Postlexical Deletion on Automatic Speech Recognition in Fast Spontaneously Spoken Zulu
Ewald van der Westhuizen, Thomas Niesler |
INTERSPEECH | 2 |
| 2015 | Unconstrained Speech Segmentation using Deep Neural Networks
Van Zyl van Vuuren, Louis ten Bosch, Thomas Niesler |
ICPRAM (1) | 3 |
| 2015 | Automatic segmentation and clustering of speech using sparse coding and metaheuristic search
Wiehan Agenbag, Thomas Niesler |
INTERSPEECH | 2 |
| 2014 | Capitalising on North American speech resources for the development of a South African English large vocabulary speech recognition system
Herman Kamper, Febe de Wet, Thomas Hain, Thomas Niesler |
Comput. Speech Lang. | 4 |
| 2013 | A dynamic programming framework for neural network-based automatic speech segmentation
Van Zyl van Vuuren, Louis ten Bosch, Thomas Niesler |
INTERSPEECH | 3 |
| 2012 | Multi-accent acoustic modelling of South African English
Herman Kamper, Félicien Jeje Muamba Mukanya, Thomas Niesler |
Speech Commun. | 3 |
| 2011 | Multi-Accent Speech Recognition of Afrikaans, Black and White Varieties of South African EnglishabstractIn this paper we investigate speech recognition performance of systems employing several accent-specific recognisers in parallel for the simultaneous recognition of multiple accents. We compare these systems with oracle systems, in which test utterances are presented to matching accent-specific recognisers, and with accent-independent systems, in which acoustic and language model training data are pooled. Our investigation is based on Afrikaans (AE), Black (BE) and White (EE) accents of South African English. We find that, when accent is classified on a per-utterance basis, parallel systems outperform oracle systems for the AE+EE accent pair while the opposite is observed for BE+EE. When accent identification is carried out on a per-speaker basis, oracle or better performance is obtained for both accent pairs. Furthermore, parallel systems based on multi-accent acoustic modelling, which allows selective crossaccent sharing of acoustic training data, outperform parallel systems using accent-specific acoustic models. The former also yields better performance than accent-independent recognition, which uses pooled acoustic and language models. Herman Kamper, Thomas Niesler |
INTERSPEECH | 2 |
| 2011 | A Study on the Perception of Tone and Intonation in SesothoabstractThis paper presents a study on the perception of Sesotho, a Southern African tonal language, employing a set of recorded minimal pairs, whose F0 contours were analyzed in a previous study using the Fujisaki model and resynthesized. Sequences of prosodically modified stimuli were produced to examine the effect of these modifications on word identification, statement/question distinction, as well as focus identification. With few exceptions, results regarding word identification are in line with our expectations. F0 modifications even seem to override vowel differences between words when both affect its meaning. With respect to the statement/question distinction, shortening of the penultimate syllable, higher speech rate and increased phrase command magnitude Ap all increase the probability of an utterance to be perceived as a question. The focus experiment, however, produced inconclusive results, possibly due to its complex setting. Hansjörg Mixdorff, Lehlohonolo Mohasi, Malillo Machobane, Thomas Niesler |
INTERSPEECH | 4 |
| 2011 | Automatic conversion between pronunciations of different English accents
Linsen Loots, Thomas Niesler |
Speech Commun. | 2 |
| 2009 | Data-driven phonetic comparison and conversion between south african, british and american English pronunciationsabstractWe analyse pronunciations in American, British and South African English pronunciation dictionaries. Three analyses are perfomed. First the accuracy is determined with which decision tree based grapheme-to-phoneme (G2P) conversion can be applied to each accent. It is found that there is little difference between the accents in this regard. Secondly, pronunciations are compared by performing pairwise alignments between the accents. Here we find that South African English pronunciation most closely matches British English. Finally, we apply decision trees to the conversion of pronunciations from one accent to another. We find that pronunciations of unknown words can be more accurately determined from a known pronunciation in a different accent than by means of G2P methods. This has important implications for the development of pronunciation dictionaries in less-resourced varieties of English, and hence also for the development of ASR systems. Linsen Loots, Thomas Niesler |
INTERSPEECH | 2 |
| 2009 | The effect of code-mixing on accent identification accuracy
Thomas Niesler, Febe de Wet |
Comput. Speech Lang. | 1 |
| 2009 | Automatic assessment of oral language proficiency and listening comprehension
Febe de Wet, Christa van der Walt, Thomas Niesler |
Speech Commun. | 3 |
| 2007 | Automatic large-scale oral language proficiency assessmentabstractPlease help us populate SUNScholar with the post print version of this article. It can be e-mailed to: [email protected] Febe de Wet, Christa van der Walt, Thomas Niesler |
INTERSPEECH | 3 |
| 2007 | Language-dependent state clustering for multilingual acoustic modelling
Thomas Niesler |
Speech Commun. | 1 |
| 2004 | The African Speech Technology Project: An Assessment
Justus C. Roux, Philippa H. Louw, Thomas Niesler |
LREC | 3 |
| 2003 | Pervasive unsupervised adaptation for lecture speech transcriptionabstractUnsupervised adaptation has evolved as a popular approach for tuning the acoustic models of speaker-independent speech recognition systems to specific speakers, speaker groups or channel conditions while making use of only untranscribed data. This study focuses on procedures for unsupervised adaptation of other probabilistic models that are involved in state-of-the-art speech recognizers and on the joint adaptation of multiple knowledge sources. In particular, we outline and evaluate approaches for adapting both the language model and the pronunciation model (lexicon) without supervision. Initial experiments on off-line lecture speech transcription achieved small but promising word error rate improvements with each approach applied separately. The experimental results on the joint application of acoustic, language and pronunciation model adaptation indicate that the individually achievable performance improvements are additive. Daniel Willett, Thomas Niesler, Erik McDermott, Yasuhiro Minami, Shigeru Katagiri |
ICASSP (1) | 2 |
| 2002 | Unsupervised language model adaptation for lecture speech transcriptionabstractUnsupervised adaptation methods have been applied successfully to the acoustic models of speech recognition systems for some time. Relatively little work has been carried out in the area of unsupervised language model adaptation however. The work presented here uses the output of a speech recogniser to adapt the backoff n-gram language model used in the decoding process. We report results for two different methods of language model adaptation, and find that best results are obtained when these two are used in conjunction with one another. The adaptation methods are applied to a Japanese large vocabulary transcription task, for which improvements both in perplexity and word error-rate are achieved. Thomas Niesler, Daniel Willett |
INTERSPEECH | 1 |
| 1999 | The 1998 HTK system for transcription of conversational telephone speechabstractThis paper describes the 1998 HTK large vocabulary speech recognition system for conversational telephone speech as used in the NIST 1998 Hub5E evaluation. Front-end and language modelling experiments conducted using various training and test sets from both the Switchboard and Callhome English corpora are presented. Our complete system includes reduced bandwidth analysis, side-based cepstral feature normalisation, vocal tract length normalisation (VTLN), triphone and quinphone hidden Markov models (HMMs) built using speaker adaptive training (SAT), maximum likelihood linear regression (MLLR) speaker adaptation and a confidence score based system combination. A detailed description of the complete system together with experimental results for each stage of our multi-pass decoding scheme is presented. The word error rate obtained is almost 20% better than our 1997 system on the development set. Thomas Hain, Philip C. Woodland, Thomas Niesler, Edward W. D. Whittaker |
ICASSP | 3 |
| 1999 | Improvements in accuracy and speed in the HTK broadcast news transcription system
Philip C. Woodland, J. J. Odell, Thomas Hain, Gareth L. Moore, Thomas Niesler, Andreas Tuerk, Edward W. D. Whittaker |
EUROSPEECH | 5 |
| 1999 | Variable-length categoryn-gram language models
Thomas Niesler, Philip C. Woodland |
Comput. Speech Lang. | 1 |
| 1998 | Comparison of part-of-speech and automatically derived category-based language models for speech recognitionabstractThis paper compares various category-based language models when used in conjunction with a word-based trigram by means of linear interpolation. Categories corresponding to parts-of-speech as well as automatically clustered groupings are considered. The category-based model employs variable-length n-grams and permits each word to belong to multiple categories. Relative word error rate reductions of between 2 and 7% over the baseline are achieved in N-best rescoring experiments on the Wall Street Journal corpus. The largest improvement is obtained with a model using automatically determined categories. Perplexities continue to decrease as the number of different categories is increased, but improvements in the word error rate reach an optimum. Thomas Niesler, Edward W. D. Whittaker, Philip C. Woodland |
ICASSP | 1 |
| 1998 | Experiments in broadcast news transcriptionabstractThis paper presents the development of the HTK broadcast news transcription system. Previously we have used data type specific modelling based on adapted Wall Street Journal trained HMMs. However, we are now experimenting with data for which no manual pre-classification or segmentation is available and therefore automatic techniques are required and compatible acoustic modelling strategies adopted. An approach for automatic audio segmentation and classification is described and evaluated as well as extensions to our previous work on segment clustering. A number of recognition experiments are presented that compare datatype specific and non-specific models; differing amounts of training data; the use of gender-dependent modelling and the effects of automatic data-type classification. It is shown that robust segmentation into a small number of audio types is possible and that models trained on a wide variety of data types can yield good performance. Philip C. Woodland, Thomas Hain, Sue Tranter, Thomas Niesler, Andreas Tuerk, Steve J. Young |
ICASSP | 4 |
| 1997 | Modelling word-pair relations in a category-based language modelabstractA new technique for modelling word occurrence correlations within a word-category based language model is presented. Empirical observations indicate that the conditional probability of a word given its category, rather than maintaining the constant value normally assumed, exhibits an exponential decay towards a constant as a function of an appropriately defined measure of separation between the correlated words. Consequently, a functional dependence of the probability upon this separation is postulated, and methods for determining both the related word pairs as well as the function parameters are developed. Experiments using the LOB, Switchboard and Wall Street Journal corpora indicate that this formulation captures the transient nature of the conditional probability effectively, and leads to reductions in perplexity of between 8 and 22%, where the largest improvements are delivered by correlations of words with themselves (self-triggers), and the reductions increase with the size of the training corpus. Thomas Niesler, Philip C. Woodland |
ICASSP | 1 |
| 1996 | A variable-length category-based n-gram language modelabstractA language model based on word-category n-grams and ambiguous category membership with n increased selectively to trade compactness for performance is presented. The use of categories leads intrinsically to a compact model with the ability to generalise to unseen word sequences, and diminishes the sparseness of the training data, thereby making larger n feasible. The language model implicitly involves a statistical tagging operation, which may be used explicitly to assign category assignments to untagged text. Experiments on the LOB corpus show the optimal model-building strategy to yield improved results with respect to conventional n-gram methods, and when used as a tagger, the model is seen to perform well in relation to a standard benchmark. Thomas Niesler, Philip C. Woodland |
ICASSP | 1 |
| 1996 | Combination of word-based and category-based language modelsabstractA language model combining word-based and category-based ngrams within a backoff framework is presented.Word n-grams conveniently capture sequential relations between particular words, while the category-model, which is based on part-of-speech classific ations and allows ambiguous category membership, is able to generalise to unseen word sequences and therefore appropriate in backoff situations.Experiments on the LOB, Switchboard and WSJ0 corpora demonstrate that the technique greatly improves language model perplexities for sparse training sets, and offers significantly improved complexity versus performance tradeoffs when compared with standard trigram models. Thomas Niesler, Philip C. Woodland |
ICSLP | 1 |