EDBT 2026 Demo / reviewers in the wild / expert
Philip N. Garner
dblp:42/7533 · also Philip Neil Garner
· DBLP profile ↗
81ranked-venue papers
13as first author
15since 2021 · last 2025
0000-0002-0814-1348ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 69 · 11 first-author · 11 since 2021Artificial intelligence and machine learning · 49 · 8 first-author · 12 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Bayesian Interpretation of Adaptive Low-Rank AdaptationabstractMotivated by the sensitivity-based importance score of the adaptive low-rank adaptation (AdaLoRA), we utilize more theoretically supported metrics, including the signal-to-noise ratio (SNR), along with the Improved Variational Online Newton (IVON) optimizer, for adaptive parameter budget allocation. The resulting Bayesian counterpart not only has matched or surpassed the performance of using the sensitivity-based importance metric but is also a faster alternative to AdaLoRA with Adam. Our theoretical analysis reveals a significant connection between the two metrics, providing a Bayesian perspective on the efficacy of sensitivity as an importance score. Furthermore, our findings suggest that the magnitude, rather than the variance, is the primary indicator of the importance of parameters. Philip N. Garner |
ICASSP | 2 |
| 2025 | Joint Fine-tuning and Conversion of Pretrained Speech and Language Models towards Linear ComplexityabstractArchitectures such as Linformer and Mamba have recently emerged as competitive linear time replacements for transformers. However, corresponding large pretrained models are often unavailable, especially in non-text domains. To remedy this, we present a Cross-Architecture Layerwise Distillation (CALD) approach that jointly converts a transformer model to a linear time substitute and fine-tunes it to a target task. We also compare several means to guide the fine-tuning to optimally retain the desired inference capability from the original model. The methods differ in their use of the target model and the trajectory of the parameters. In a series of empirical studies on language processing, language modeling, and speech processing, we show that CALD can effectively recover the result of the original model, and that the guiding strategy contributes to the result. Some reasons for the variation are suggested. Mutian He 0001, Philip N. Garner |
ICLR | 2 |
| 2025 | Exploring auditory feedback mechanisms in speech recognitionabstractFor many years, automatic speech recognition (ASR) has been built on compressed filter-bank features understood to be a rough model of the cochlea. However, recent understanding, evidenced by oto-acoustic emissions, is that the cochlea is composed of driven oscillators. The Hopf mechanism arising from an oscillator model explains the well known cube-root compression. A bifurcation arises from an inner feedback loop from the outer to inner hair cells. Further, larger feedback loops exist along the efferent path of the auditory nerve, particularly from the olivary complex. The means of combination of these signals is less well understood. In the present study, to the extent to which current compute power allows, we investigate how to incorporate the Hopf mechanism and the olivocochlear feedback mechanisms into ASR. Results show that adding this larger feedback loop appears beneficial for ASR. The results currently have modest implications for ASR, however, perhaps more importantly, such technology could be used to make inference about the biological mechanism. Louise Coppieters de Gibson, Philip N. Garner |
INTERSPEECH | 2 |
| 2024 | Bayesian Parameter-Efficient Fine-Tuning for Overcoming Catastrophic ForgettingabstractWe are motivated primarily by the adaptation of text-to-speech synthesis models; however we argue that more generic parameter-efficient fine-tuning (PEFT) is an appropriate framework to do such adaptation. Nevertheless, catastrophic forgetting remains an issue with PEFT, damaging the pre-trained model's inherent capabilities. We demonstrate that existing Bayesian learning techniques can be applied to PEFT to prevent catastrophic forgetting as long as the parameter shift of the fine-tuned layers can be calculated differentiably. In a principled series of experiments on language modeling and speech synthesis tasks, we utilize established Laplace approximations, including diagonal and Kronecker-factored approaches, to regularize PEFT with the low-rank adaptation (LoRA) and compare their performance in pre-training knowledge preservation. Our results demonstrate that catastrophic forgetting can be overcome by our methods without degrading the fine-tuning performance, and using the Kronecker-factored approximation produces a better preservation of the pre-training knowledge than the diagonal ones. Philip N. Garner |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Vulnerability of Automatic Identity Recognition to Audio-Visual DeepfakesabstractThe task of deepfakes detection is far from being solved by speech or vision researchers. Several publicly available databases of fake synthetic video and speech were built to aid the development of detection methods. However, existing databases typically focus on visual or voice modalities and provide no proof that their deepfakes can in fact impersonate any real person. In this paper, we present the first realistic audio-visual database of deepfakes SWAN-DF, where lips and speech are well synchronized and video have high visual and audio qualities. We took the publicly available SWAN dataset of real videos with different identities to create audio-visual deepfakes using several models from DeepFaceLab and blending techniques for face swapping and HiFiVC, DiffVC, YourTTS, and FreeVC models for voice conversion. From the publicly available speech dataset LibriTTS, we also created a separate database of only audio deepfakes LibriTTS-DF using several latest text to speech methods: YourTTS, Adaspeech, and TorToiSe. We demonstrate the vulnerability of a state of the art speaker recognition system, such as ECAPA-TDNN-based model from SpeechBrain, to the synthetic voices. Similarly, we tested face recognition system based on the MobileFaceNet architecture to several variants of our visual deepfakes. The vulnerability assessment show that by tuning the existing pretrained deepfake models to specific identities, one can successfully spoof the face and speaker recognition systems in more than 90% of the time and achieve a very realistic looking and sounding fake video of a given person. Pavel Korshunov, Philip N. Garner, Sébastien Marcel |
IJCB | 3 |
| 2023 | Can ChatGPT Detect Intent? Evaluating Large Language Models for Spoken Language Understanding
Mutian He 0001, Philip N. Garner |
INTERSPEECH | 2 |
| 2022 | Bayesian Recurrent Units and the Forward-Backward AlgorithmabstractUsing Bayes's theorem, we derive a unit-wise recurrence as well as a backward recursion similar to the forward-backward algorithm. The resulting Bayesian recurrent units can be integrated as recurrent neural networks within deep learning frameworks, while retaining a probabilistic interpretation from the direct correspondence with hidden Markov models. Whilst the contribution is mainly theoretical, experiments on speech recognition indicate that adding the derived units at the end of state-of-the-art recurrent architectures can improve the performance at a very low cost in terms of trainable parameters. Alexandre Bittar, Philip N. Garner |
INTERSPEECH | 2 |
| 2022 | Low-Level Physiological Implications of End-to-End Learning for Speech RecognitionabstractCurrent speech recognition architectures perform very well from the point of view of machine learning, hence user interaction. This suggests that they are emulating the human biological system well. We investigate whether the inference can be inverted to provide insights into that biological system; in particular the hearing mechanism. Using SincNet, we confirm that end-to-end systems do learn well known filterbank structures. However, we also show that wider band-width filters are important in the learned structure. Whilst some benefits can be gained by initialising both narrow and wide-band filters, physiological constraints suggest that such filters arise in mid-brain rather than the cochlea. We show that standard machine learning architectures must be modified to allow this process to be emulated neurally. Louise Coppieters de Gibson, Philip N. Garner |
INTERSPEECH | 2 |
| 2022 | Conversational Speech Recognition Needs Data? Experiments with Austrian GermanabstractConversational speech represents one of the most complex of automatic speech recognition (ASR) tasks owing to the high inter-speaker variation in both pronunciation and conversational dynamics. Such complexity is particularly sensitive to low-resourced (LR) scenarios. Recent developments in self-supervision have allowed such scenarios to take advantage of large amounts of otherwise unrelated data. In this study, we characterise an (LR) Austrian German conversational task. We begin with a non-pre-trained baseline and show that fine-tuning of a model pre-trained using self-supervision leads to improvements consistent with those in the literature; this extends to cases where a lexicon and language model are included. We also show that the advantage of pre-training indeed arises from the larger database rather than the self-supervision. Further, by use of a leave-one-conversation out technique, we demonstrate that robustness problems remain with respect to inter-speaker and inter-conversation variation. This serves to guide where future research might best be focused in light of the current state-of-the-art. Julian Linke, Philip N. Garner, Gernot Kubin, Barbara Schuppler |
LREC | 2 |
| 2022 | Investigating a neural all pass warp in modern TTS applicationsabstractWe present a neural implementation of the all pass warp (APW) previously used for vocal tract length normalisation. This includes an efficient back-propagation, which can easily be integrated in modern neural network frameworks. The APW offers a low-dimensional control to alter the spectrum, which by design generalises over different speakers. We investigate the APW in two tasks required for future dialogue or translation agents, and provide a fairly thorough literature review for both: (1) Zero-shot speaker adaptation to allow keeping the source speaker identity with very small amounts of data. Experiments show increased speaker similarity and prove that the APW increases the generalisability of a multi-speaker model. (2) Emotional speech synthesis to translate or produce affective cues. To the best of our knowledge this is the first attempt on emotional speech synthesis with an APW. While the APW is not able to increase expressiveness or audio quality, our analysis shows that the warping correlates with the level of valence in the emotion. This work should enable future research on emotion translation during machine translation. Bastian Schnell, Philip N. Garner |
Speech Commun. | 2 |
| 2021 | An Evaluation Benchmark for Automatic Speech Recognition of German-English Code-SwitchingabstractCode-switching arises when a (typically multilingual) speaker changes language during an utterance. This linguistic phenomenon causes problems for automatic speech recognition as the models are typically monolingual. In this work, we present a code-switching evaluation scenario for German-English that is created by resegmenting the German Spoken Wikipedia Corpus. Since these articles span a wide variety of (often technical) topics, they include a lot of borrowing and code-switching phenomena. The resulting corpus consists of around 34 hours of intra-sentential switches. We investigate end-to-end approaches using both monolingual and multilingual automatic speech recognition as well as language modeling to address the code-switching scenario. Results suggest that multilingual sequence-to-sequence approaches are to be preferred for code-switching thanks to the power of the attention mechanism. The segments are made available to the community as a benchmark. Abbas Khosravani, Philip N. Garner, Alexandros Lararidis |
ASRU | 2 |
| 2021 | Learning to Translate Low-Resourced Swiss German Dialectal Speech into Standard German TextabstractFor a low-resourced language like Swiss German with no standard orthography and a significant variation in its written form, spoken language resources are more likely to come with translations than transcriptions. Moreover, the desired output of an automatic transcription system for Swiss German multi-dialectal speech is Standard German. This, in turn, is due to many applications that include our TV Box voice assistant and broadcast media. It follows that a translation is usually required as Swiss German and Standard German have mismatches on all linguistic levels. Unfortunately, there are not enough parallel text corpora available for training a proper translation system, nor enough in-domain speech translation (ST) data for training an ST system. We aim at investigating an end-to-end approach for multi-dialect Swiss German ST using transfer learning. Our ST model is based on an encoder-decoder architecture where we initialize the encoder with a cross-lingual speech representation model which is adapted to in-domain Swiss German speech data. We demonstrate that training the decoder on an out-of-domain ST corpus by preserving the encoder unit and then fine-tuning on in-domain ST data can be more effective than a cascade or vanilla direct ST. Abbas Khosravani, Philip N. Garner, Alexandros Lazaridis |
ASRU | 2 |
| 2021 | A Bayesian Interpretation of the Light Gated Recurrent UnitabstractWe summarise previous work showing that the basic sigmoid activation function arises as an instance of Bayes’s theorem, and that recurrence follows from the prior. We derive a layerwise recurrence without the assumptions of previous work, and show that it leads to a standard recurrence with modest modifications to reflect use of log-probabilities. The resulting architecture closely resembles the Li-GRU which is the current state of the art for ASR. Although the contribution is mainly theoretical, we show that it is able to outperform the state of the art on the TIMIT and AMI datasets. Alexandre Bittar, Philip N. Garner |
ICASSP | 2 |
| 2021 | Modeling Dialectal Variation for Swiss German Automatic Speech Recognition
Abbas Khosravani, Philip N. Garner, Alexandros Lazaridis |
Interspeech | 2 |
| 2021 | A Bayesian Approach to Recurrence in Neural NetworksabstractWe begin by reiterating that common neural network activation functions have simple Bayesian origins. In this spirit, we go on to show that Bayes's theorem also implies a simple recurrence relation; this leads to a Bayesian recurrent unit with a prescribed feedback formulation. We show that introduction of a context indicator leads to a variable feedback that is similar to the forget mechanism in conventional recurrent units. A similar approach leads to a probabilistic input gate. The Bayesian formulation leads naturally to the two pass algorithm of the Kalman smoother or forward-backward algorithm, meaning that inference naturally depends upon future inputs as well as past ones. Experiments on speech recognition confirm that the resulting architecture can perform as well as a bidirectional recurrent network with the same number of parameters as a unidirectional one. Further, when configured explicitly bidirectionally, the architecture can exceed the performance of a conventional bidirectional recurrence. Philip N. Garner, Sibo Tong |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | A $t$-Distribution Based Operator for Enhancing Out of Distribution Robustness of Neural Network ClassifiersabstractNeural Network (NN) classifiers can assign extreme probabilities to samples that have not appeared during training (out-of-distribution samples) resulting in erroneous and unreliable predictions. One of the causes for this unwanted behaviour lies in the use of the standard softmax operator which pushes the posterior probabilities to be either zero or unity hence failing to model uncertainty. The statistical derivation of the softmax operator relies on the assumption that the distributions of the latent variables for a given class are Gaussian with known variance. However, it is possible to use different assumptions in the same derivation and attain from other families of distributions as well. This allows derivation of novel operators with more favourable properties. Here, a novel operator is proposed that is derived using t-distributions which are capable of providing a better description of uncertainty. It is shown that classifiers that adopt this novel operator can be more robust to out of distribution samples, often outperforming NNs that use the standard softmax operator. These enhancements can be reached with minimal changes to the NN architecture. Niccolò Antonello, Philip N. Garner |
IEEE Signal Process. Lett. | 2 |
| 2019 | An End-to-end Network to Synthesize Intonation Using a Generalized Command Response ModelabstractThe generalized command response (GCR) model represents intonation as a superposition of muscle responses to spike command signals. We have previously shown that the spikes can be predicted by a two-stage system, consisting of a recurrent neural network and a post-processing procedure, but the responses themselves were fixed dictionary atoms. We propose an end-to-end neural architecture that replaces the dictionary atoms with trainable second-order recurrent elements analogous to recursive filters. We demonstrate gradient stability under modest conditions, and show that the system can be trained by imposing temporal sparsity constraints. Subjective listening tests demonstrate that the system can synthesize intonation with high naturalness, comparable to state-of-the-art acoustic models, and retains the physiological plausibility of the GCR model. François Marelli, Bastian Schnell, Hervé Bourlard, Thierry Dutoit, Philip N. Garner |
ICASSP | 5 |
| 2019 | Empirical Evaluation and Combination of Punctuation Prediction Models Applied to Broadcast NewsabstractNatural language processing techniques are dependent upon punctuation to work well. When their input is taken from speech recognition, it is necessary to reconstruct the punctuation; in particular sentence boundaries. We define a range of features from low level acoustics to those with high level lexical semantics, including deep and recurrent models; these in turn are representative of a broad range of approaches used by previous authors for punctuation prediction. We combine the features using a gradient boosting machine that is also capable of indicating the relative importance of each feature. In an empirical study, we show that features from different semantic levels are in fact complementary, that combining statistical and deep learning methods yields better prediction results, and that generalization across different speaking styles is difficult to achieve without adaptation. Our best model achieves an F-Measure of 82.8 on a challenging broadcast news dataset. Alexandre Nanchen, Philip N. Garner |
ICASSP | 2 |
| 2019 | An Investigation of Multilingual ASR Using End-to-end LF-MMIabstractThe end-to-end lattice-free maximum mutual information (LF-MMI) approach has recently been shown to be beneficial for automatic speech recognition (ASR) in general. More specifically, its end-to-end nature and use of context independent phone labels make it attractive for multilingual ASR. We show that end-to-end LF-MMI is indeed competitive on a low-resourced multilingual task, comfortably outperforming a connectionist temporal classification (CTC) baseline. We further investigate the feasibility of biphone contexts, being a candidate compromise between the context independent approach and the triphone contexts that usually perform well. We show that biphones do not initially perform well, but can do so after language adaptive training, concluding that biphones carry language variability but are promising for multilingual ASR. Sibo Tong, Philip N. Garner, Hervé Bourlard |
ICASSP | 2 |
| 2019 | Self-Attention for Speech Emotion RecognitionabstractSpeech Emotion Recognition (SER) has been shown to benefit from many of the recent advances in deep learning, including recurrent based and attention based neural network architectures as well. Nevertheless, performance still falls short of that of humans. In this work, we investigate whether SER could benefit from the self-attention and global windowing of the transformer model. We show on the IEMOCAP database that this is indeed the case. Finally, we investigate whether using the distribution of, possibly conflicting, annotations in the training data, as soft targets could outperform a majority voting. We prove that this performance increases with the agreement level of the annotators. Lorenzo Tarantino, Philip N. Garner, Alexandros Lazaridis |
INTERSPEECH | 2 |
| 2019 | Unbiased Semi-Supervised LF-MMI Training Using Dropout
Sibo Tong, Apoorv Vyas, Philip N. Garner, Hervé Bourlard |
INTERSPEECH | 3 |
| 2018 | A Neural Model to Predict Parameters for a Generalized Command Response Model of IntonationabstractThe Generalised Command Response (GCR) model is a time-local model of intonation that has been shown to lend itself to (cross-language) transfer of emphasis. In order to generalise the model to longer prosodic sequences, we show that it can be driven by a recurrent neural network emulating a spiking neural network. We show that a loss function for error backpropagation can be formulated analogously to that of the Spike Pattern Association Neuron (SPAN) method for spiking networks. The resulting system is able to generate prosody comparable to a state-of-the-art deep neural network implementation, but potentially retaining the transfer capabilities of the GCR model. Bastian Schnell, Philip N. Garner |
INTERSPEECH | 2 |
| 2018 | Fast Language Adaptation Using Phonological InformationabstractPhoneme-based multilingual connectionist temporal classification (CTC) model is easily extensible to a new language by concatenating parameters of the new phonemes to the output layer. In the present paper, we improve cross-lingual adaptation in the context of phoneme-based CTC models by using phonological information. A universal (IPA) phoneme classifier is first trained on phonological features generated from a phonological attribute detector. When adapting the multilingual CTC to a new, never seen, language, phonological attributes of the unseen phonemes are derived based on phonology and fed into the phoneme classifier. Posteriors given by the classifier are used to initialize the parameters of the unseen phonemes when extending the multilingual CTC output layer to the target language. Adaptation experiments show that the proposed initialization approaches further improve the cross-lingual adaptation on CTC models and yield significant improvements over Deep Neural Network / Hidden Markov Model (DNN/HMM)-based adaptation using limited data. Sibo Tong, Philip N. Garner, Hervé Bourlard |
INTERSPEECH | 2 |
| 2018 | Context-Aware Attention Mechanism for Speech Emotion RecognitionabstractIn this work, we study the use of attention mechanisms to enhance the performance of the state-of-the-art deep learning model in Speech Emotion Recognition (SER). We introduce a new Long Short-Term Memory (LSTM)-based neural network attention model which is able to take into account the temporal information in speech during the computation of the attention vector. The proposed LSTM-based model is evaluated on the IEMOCAP dataset using a 5-fold cross-validation scheme and achieved 68.8% weighted accuracy on 4 classes, which outperforms the state-of-the-art models. Gaetan Ramet, Philip N. Garner, Michael Baeriswyl, Alexandros Lazaridis |
SLT | 2 |
| 2018 | Intonation modelling using a muscle model and perceptually weighted matching pursuit
Pierre-Edouard Honnet, Branislav Gerazov, Aleksandar Gjoreski, Philip N. Garner |
Speech Commun. | 4 |
| 2018 | Cross-lingual adaptation of a CTC-based multilingual acoustic model
Sibo Tong, Philip N. Garner, Hervé Bourlard |
Speech Commun. | 2 |
| 2017 | An Investigation of Deep Neural Networks for Multilingual Speech Recognition Training and AdaptationabstractDifferent training and adaptation techniques for multilingual Automatic Speech Recognition (ASR) are explored in the context of hybrid systems, exploiting Deep Neural Networks (DNN) and Hidden Markov Models (HMM). In multilingual DNN training, the hidden layers (possibly extracting bottleneck features) are usually shared across languages, and the output layer can either model multiple sets of language-specific senones or one single universal IPA-based multilingual senone set. Both architectures are investigated, exploiting and comparing different language adaptive training (LAT) techniques originating from successful DNN-based speaker-adaptation. More specifically, speaker adaptive training methods such as Cluster Adaptive Training (CAT) and Learning Hidden Unit Contribution (LHUC) are considered. In addition, a language adaptive output architecture for IPA-based universal DNN is also studied and tested. Experiments show that LAT improves the performance and adaptation on the top layer further improves the accuracy. By combining state-level minimum Bayes risk (sMBR) sequence training with LAT, we show that a language adaptively trained IPA-based universal DNN outperforms a monolingually sequence trained model. Sibo Tong, Philip N. Garner, Hervé Bourlard |
INTERSPEECH | 2 |
| 2016 | Sound Pattern Matching for Automatic Prosodic Event DetectionabstractLIDIAP Milos Cernak, Afsaneh Asaei, Pierre-Edouard Honnet, Philip N. Garner, Hervé Bourlard |
INTERSPEECH | 4 |
| 2016 | PhonVoc: A Phonetic and Phonological Vocoding ToolkitabstractWe present the PhonVoc toolkit, a cascaded deep neural network (DNN) composed of speech analyser and synthesizer that use a shared phonetic and/or phonological speech representation.The free toolkit is distributed as open-source software under a BSD 3-Clause License, available at https://github.com/idiap/phonvoc with the pre-trained US English analysis and synthesis DNNs, and thus it is ready for immediate use.In a broader context, the toolkit implements training and testing of the analysis by synthesis heuristic model.It is thus designed for the wider speech community working in acoustic phonetics, laboratory phonology, and parametric speech coding.The toolkit interprets the phonetic posterior probabilities as a sequential scheme, whereas the phonological posterior-class probabilities are considered as a parallel via K different phonological classes.A case study is presented on a LibriSpeech database and a LibriVox US English native female speaker.The phonetic and phonological vocoding yield comparable performance, improving speech quality by merging the phonetic and phonological speech representation. Milos Cernak, Philip N. Garner |
INTERSPEECH | 2 |
| 2016 | The SIWIS Database: A Multilingual Speech Database with Acted EmphasisabstractWe describe here a collection of speech data of bilingual and trilingual speakers of English, French, German and Italian. In the context of speech to speech translation (S2ST), this database is designed for several purposes and studies: training CLSA systems (cross-language speaker adaptation), conveying emphasis through S2ST systems, and evaluating TTS systems. More precisely, 36 speakers judged as accentless (22 bilingual and 14 trilingual speakers) were recorded for a set of 171 prompts in two or three languages, amounting to a total of 24 hours of speech. These sets of prompts include 100 sentences from news, 25 sentences from Europarl, the same 25 sentences with one acted emphasised word, 20 semantically unpredictable sentences, and finally a 240-word long text. All in all, it yielded 64 bilingual session pairs of the six possible combinations of the four languages. The database is freely available for non-commercial use and scientific research purposes. Jean-Philippe Goldman, Pierre-Edouard Honnet, Robert A. J. Clark, Philip N. Garner, Maria Ivanova, Alexandros Lazaridis, Tiago Macedo, Beat Pfister, Manuel Sam Ribeiro, Eric Wehrli, Junichi Yamagishi |
INTERSPEECH | 4 |
| 2016 | Probabilistic Amplitude Demodulation Features in Speech Synthesis for Improving ProsodyabstractAbstract Amplitude demodulation (AM) is a signal decomposition technique by which a signal can be decomposed to a product of two signals, i.e, a quickly varying carrier and a slowly varying modulator. In this work, the probabilistic amplitude demodulation (PAD) features are used to improve prosody in speech synthesis. The PAD is applied iteratively for generating syllable and stress amplitude modulations in a cascade manner. The PAD features are used as a secondary input scheme along with the standard text-based input features in statistical parametric speech syn- thesis. Specifically, deep neural network (DNN)-based speech synthesis is used to evaluate the importance of these features. Objective evaluation has shown that the proposed system using the PAD features has improved mainly prosody modelling; it outperforms the baseline system by approximately 5% in terms of relative reduction in root mean square error (RMSE) of the fundamental frequency (F0). The significance of this improvement is validated by subjective evaluation of the overall speech quality, achieving 38.6% over 19.5% preference score in respect to the baseline system, in an ABX test. Alexandros Lazaridis, Milos Cernak, Philip N. Garner |
INTERSPEECH | 3 |
| 2016 | Composition of Deep and Spiking Neural Networks for Very Low Bit Rate Speech CodingabstractMost current very low bit rate (VLBR) speech coding systems use hidden Markov model (HMM) based speech recognition and synthesis techniques. This allows transmission of information (such as phonemes) segment by segment; this decreases the bit rate. However, an encoder based on a phoneme speech recognition may create bursts of segmental errors; these would be further propagated to any suprasegmental (such as syllable) information coding. Together with the errors of voicing detection in pitch parametrization, HMM-based speech coding leads to speech discontinuities and unnatural speech sound artifacts. In this paper, we propose a novel VLBR speech coding framework based on neural networks (NNs) for end-to-end speech analysis and synthesis without HMMs. The speech coding framework relies on a phonological (subphonetic) representation of speech. It is designed as a composition of deep and spiking NNs: a bank of phonological analyzers at the transmitter, and a phonological synthesizer at the receiver. These are both realized as deep NNs, along with a spiking NN as an incremental and robust encoder of syllable boundaries for coding of continuous fundamental frequency (F0). A combination of phonological features defines much more sound patterns than phonetic features defined by HMM-based speech coders; this finer analysis/synthesis code contributes to smoother encoded speech. Listeners significantly prefer the NN-based approach due to fewer discontinuities and speech artifacts of the encoded speech. A single forward pass is required during the speech encoding and decoding. The proposed VLBR speech coding operates at a bit rate of approximately 360 bits/s. Milos Cernak, Alexandros Lazaridis, Afsaneh Asaei, Philip N. Garner |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2015 | Phonological vocoding using artificial neural networksabstractWe investigate a vocoder based on artificial neural networks using a phonological speech representation. Speech decomposition is based on the phonological encoders, realised as neural network classifiers, that are trained for a particular language. The speech reconstruction process involves using a Deep Neural Network (DNN) to map phonological features posteriors to speech parameters - line spectra and glottal signal parameters - followed by LPC resynthesis. This DNN is trained on a target voice without transcriptions, in a semi-supervised manner. Both encoder and decoder are based on neural networks and thus the vocoding is achieved using a simple fast forward pass. An experiment with French vocoding and a target male voice trained on 21 hour long audio book is presented. An application of the phonological vocoder to low bit rate speech coding is shown, where transmitted phonological posteriors are pruned and quantized. The vocoder with scalar quantization operates at 1 kbps, with potential for lower bit-rate. Milos Cernak, Blaise Potard, Philip N. Garner |
ICASSP | 3 |
| 2015 | Atom decomposition-based intonation modellingabstractCurrent statistical parametric text-to-speech (TTS) synthesis methods allow production of neutral speech with acceptable quality. However, prosody is often qualified as unsatisfactory and sounding too flat. In this paper, we address intonation modelling for TTS based on physiological aspects of prosody production. A set of gamma distribution shaped atoms is defined and then intonation decomposition is performed using a matching pursuit algorithm. Some preliminary experiments show that this model allows easy extraction of physiologically meaningful atoms that could be used to generate intonation in a TTS system. Pierre-Edouard Honnet, Branislav Gerazov, Philip N. Garner |
ICASSP | 3 |
| 2015 | Robust microphone placement for source localization from noisy distance measurementsabstractWe propose a novel algorithm to design an optimum array geometry for source localization inside an enclosure. We assume a square-law decay propagation model for the sound acquisition so that the additive noise on the measured source-microphone distances is proportional to the distances regardless of the noise distribution. We formulate the source localization as an instance of the “Generalized Trust Region Subproblem” (GTRS) whose solution gives the location of the source. We show that by suitable selection of the microphone locations, one can tremendously decrease the noise-sensitivity of the resulting solution. In particular, by minimizing the noise-sensitivity of the source location in terms of sensor positions, we find the optimal noise-robust array geometry for the enclosure. Simulation results are provided to show the efficiency of the proposed algorithm. Mohammad Javad Taghizadeh, Saeid Haghighatshoar, Afsaneh Asaei, Philip N. Garner, Hervé Bourlard |
ICASSP | 4 |
| 2015 | Weighted correlation based atom decomposition intonation modellingabstractIntonation modelling is an integral part of text-to-speech systems from their very beginnings. This has led to the proliferation of various intonation models, each with its own relative strengths and weaknesses. Only a few of these intonation models are based on physiology, despite the advantage that such models are language independent. We propose a new intonation model inspired by the physiology of intonation production, which is based on decomposing the F0 contour into elementary atoms. The model, named the Weighted Correlation Atom Decomposition model (WCAD), is a generalisation of the command response (CR) model and has the advantage of having a simple parameter extraction method. The decomposition process follows a matching pursuit approach based on using the perceptually relevant weighted correlation as a cost function. The results have affirmed the plausibility of using the WCAD model to model F0 contours across different languages and speakers. The results have also shown that the WCAD model has good comparative performance to the CR model, giving it practical importance. Branislav Gerazov, Pierre-Edouard Honnet, Aleksandar Gjoreski, Philip N. Garner |
INTERSPEECH | 4 |
| 2015 | Ad hoc microphone array calibration: Euclidean distance matrix completion algorithm and theoretical guarantees
Mohammad Javad Taghizadeh, Reza Parhizkar, Philip N. Garner, Hervé Bourlard, Afsaneh Asaei |
Signal Process. | 3 |
| 2015 | Incremental Syllable-Context Phonetic VocodingabstractCurrent very low bit rate speech coders are, due to complexity limitations, designed to work off-line. This paper investigates incremental speech coding that operates real-time and incrementally (i.e., encoded speech depends only on already-uttered speech without the need of future speech information). Since human speech communication is asynchronous (i.e., different information flows being simultaneously processed), we hypothesized that such an incremental speech coder should also operate asynchronously. To accomplish this task, we describe speech coding that reflects the human cortical temporal sampling that packages information into units of different temporal granularity, such as phonemes and syllables, in parallel. More specifically, a phonetic vocoder-cascaded speech recognition and synthesis systems-extended with syllable-based information transmission mechanisms is investigated. There are two main aspects evaluated in this work, the synchronous and asynchronous coding. Synchronous coding refers to the case when the phonetic vocoder and speech generation process depend on the syllable boundaries during encoding and decoding respectively. On the other hand, asynchronous coding refers to the case when the phonetic encoding and speech generation processes are done independently of the syllable boundaries. Our experiments confirmed that the asynchronous incremental speech coding performs better, in terms of intelligibility and overall speech quality, mainly due to better alignment of the segmental and prosodic information. The proposed vocoding operates at an uncompressed bit rate of 213 bits/sec and achieves an average communication delay of 243 ms. Milos Cernak, Philip N. Garner, Alexandros Lazaridis, Petr Motlícek, Xingyu Na |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Stress and accent transmission in HMM-based syllable-context very low bit rate speech codingabstractLIDIAP Milos Cernak, Alexandros Lazaridis, Philip N. Garner, Petr Motlícek |
INTERSPEECH | 3 |
| 2014 | Automatic speech recognition and translation of a Swiss German dialect: WalliserdeutschabstractWalliserdeutsch is a Swiss German dialect spoken in the south west of Switzerland. To investigate the potential of automatic speech processing of Walliserdeutsch, a small database was collected based mainly on broadcast news from a local radio station. Experiments suggest that automatic speech recognition is feasible: use of another (Swiss German) database shows that the small data size lends itself to bootstrapping from other data; use of Kullback-Leibler HMM suggests that phoneme mapping techniques can compensate for a grapheme-based dictionary. Experiments also indicate that statistical machine translation is feasible; the difficulty of small data size is offset by the close proximity to (high) German. Philip N. Garner, David Imseng, Thomas Meyer 0003 |
INTERSPEECH | 1 |
| 2014 | Enhanced diffuse field model for ad hoc microphone array calibration
Mohammad Javad Taghizadeh, Philip N. Garner, Hervé Bourlard |
Signal Process. | 2 |
| 2014 | Using out-of-language data to improve an under-resourced speech recognizer
David Imseng, Petr Motlícek, Hervé Bourlard, Philip N. Garner |
Speech Commun. | 4 |
| 2013 | Impact of deep MLP architecture on different acoustic modeling techniques for under-resourced speech recognitionabstractPosterior based acoustic modeling techniques such as Kullback-Leibler divergence based HMM (KL-HMM) and Tandem are able to exploit out-of-language data through posterior features, estimated by a Multi-Layer Perceptron (MLP). In this paper, we investigate the performance of posterior based approaches in the context of under-resourced speech recognition when a standard three-layer MLP is replaced by a deeper five-layer MLP. The deeper MLP architecture yields similar gains of about 15% (relative) for Tandem, KL-HMM as well as for a hybrid HMM/MLP system that directly uses the posterior estimates as emission probabilities. The best performing system, a bilingual KL-HMM based on a deep MLP, jointly trained on Afrikaans and Dutch data, performs 13% better than a hybrid system using the same bilingual MLP and 26% better than a subspace Gaussian mixture system only trained on Afrikaans data. David Imseng, Petr Motlícek, Philip N. Garner, Hervé Bourlard |
ASRU | 3 |
| 2013 | On the (UN)importance of the contextual factors in HMM-based speech synthesis and codingabstractThis paper presents an evaluation of the contextual factors of HMMbased speech synthesis and coding systems.Two experimental setups are proposed that are based on successive context addition from phonetic to full-context.The aim was to investigate the impact of the individual contextual factors on the speech quality.In that sense important and unimportant (i.e., not having significant impact on speech quality, also called weak) contextual factors were identified.The results imply that in speech coding the improvement in quality can be achieved just with reconstruction of syllable contexts.The sentence and utterance contexts are unimportant on the decoder side, and it is not necessary to deal with them.Although in speech coding the wider context was not necessary, in speech synthesis current syllable and utterance contexts are more important over others (previous and next word/phrase contexts). Milos Cernak, Petr Motlícek, Philip N. Garner |
ICASSP | 3 |
| 2013 | Accent adaptation using Subspace Gaussian Mixture ModelsabstractThis paper investigates employment of Subspace Gaussian Mixture Models (SGMMs) for acoustic model adaptation towards different accents for English speech recognition. The SGMMs comprise globally-shared and state-specific parameters which can efficiently be employed for various kinds of acoustic parameter tying. Research results indicate that well-defined sharing of acoustic model parameters in SGMMs can significantly outperform adapted systems based on conventional HMM/GMMs. Furthermore, SGMMs rapidly achieve target acoustic models with small amounts of data. Experiments performed with US and UK English versions of the Wall Street Journal (WSJ) corpora indicate that SGMMs lead to approximately 20% and 8% relative improvements with respect to speaker-independent and speaker-adapted acoustic models respectively over conventional HMM/GMMs. Finally, we demonstrate that SGMMs adapted only with 1.5 hours can reach performance of HMM/GMMs trained with 18 hours. Petr Motlícek, Philip N. Garner, Namhoon Kim, Jeongmi Cho |
ICASSP | 2 |
| 2013 | Syllable-based pitch encoding for low bit rate speech coding with recognition/synthesis architectureabstractCurrent HMM-based low bit rate speech coding systems work with phonetic vocoders.Pitch contour coding (on frame or phoneme level) is usually fairly orthogonal to other speech coding parameters.We make an assumption in our work that the speech signal contains supra-segmental cues.Hence, we present encoding of the pitch on the syllable level, used in the framework of a recognition/synthesis speech coder with phonetic vocoder.The results imply that high accuracy pitch contour reconstruction with negligible speech quality degradation is possible.The proposed pitch encoding technique operates on 30-35 bits per second. Milos Cernak, Xingyu Na, Philip N. Garner |
INTERSPEECH | 3 |
| 2013 | Crosslingual tandem-SGMM: exploiting out-of-language data for acoustic model and feature level adaptationabstractRecent studies have shown that speech recognizers may benefit from data in languages other than the target language through efficient acoustic model- or feature-level adaptation. Crosslingual Tandem-Subspace Gaussian Mixture Models (SGMM) are successfully able to combine acoustic model- and feature-level adaptation techniques. More specifically, we focus on under-resourced languages (Afrikaans in our case) and perform feature-level adaptation through the estimation of phone class posterior features with a Multilayer Perceptron that was trained on data from a similar language with large amounts of available speech data (Dutch in our case). The same Dutch data can also be exploited on an acoustic model-level by training globally-shared SGMM parameters in a crosslingual way. The two adaptation techniques are indeed complementary and result in a crosslingual Tandem-SGMM system that yields relative improvement of about 22% compared to a standard speech recognizer on an Afrikaans phoneme recognition task. Interestingly, eventual score-level combination of the individual SGMM systems yields additional 3% relative improvement. Petr Motlícek, David Imseng, Philip N. Garner |
INTERSPEECH | 3 |
| 2013 | A Simple Continuous Pitch Estimation AlgorithmabstractRecent work in text to speech synthesis has pointed to the benefit of using a continuous pitch estimate; that is, one that records pitch even when voicing is not present. Such an approach typically requires interpolation. The purpose of this letter is to show that a continuous pitch estimation is available from a combination of otherwise well known techniques. Further, in the case of an autocorrelation based estimate, the continuous requirement negates the need for other heuristics to correct for common errors. An algorithm is suggested, illustrated, and demonstrated using a parametric vocoder. Philip N. Garner, Milos Cernak, Petr Motlícek |
IEEE Signal Process. Lett. | 1 |
| 2013 | Applying Multi- and Cross-Lingual Stochastic Phone Space Transformations to Non-Native Speech RecognitionabstractIn the context of hybrid HMM/MLP Automatic Speech Recognition (ASR), this paper describes an investigation into a new type of stochastic phone space transformation, which maps “source” phone (or phone HMM state) posterior probabilities (as obtained at the output of a Multilayer Perceptron/MLP) into “destination” phone (HMM phone state) posterior probabilities. The resulting stochastic matrix transformation can be used within the same language to automatically adapt to different phone formats (e.g., IPA) or across languages. Additionally, as shown here, it can also be applied successfully to non-native speech recognition. In the same spirit as MLLR adaptation, or MLP adaptation, the approach proposed here is directly mapping posterior distributions, and is trained by optimizing on a small amount of adaptation data a Kullback-Leibler based cost function, along a modified version of an iterative EM algorithm. On a non-native English database (HIWIRE), and comparing with multiple setups (monophone and triphone mapping, MLLR adaptation) we show that the resulting posterior mapping yields state-of-the-art results using very limited amounts of adaptation data in mono-, cross- and multi-lingual setups. We also show that “universal” phone posteriors, trained on a large amount of multilingual data, can be transformed to English phone posteriors, resulting in an ASR system that significantly outperforms a system trained on English data only. Finally, we demonstrate that the proposed approach outperforms alternative data-driven, as well as a knowledge-based, mapping techniques. David Imseng, Hervé Bourlard, John Dines, Philip N. Garner, Mathew Magimai-Doss |
IEEE Trans. Speech Audio Process. | 4 |
| 2012 | Using KL-divergence and multilingual information to improve ASR for under-resourced languagesabstractSetting out from the point of view that automatic speech recognition (ASR) ought to benefit from data in languages other than the target language, we propose a novel Kullback-Leibler (KL) divergence based method that is able to exploit multilingual information in the form of universal phoneme posterior probabilities conditioned on the acoustics. We formulate a means to train a recognizer on several different languages, and subsequently recognize speech in a target language for which only a small amount of data is available. Taking the Greek SpeechDat(II) data as an example, we show that the proposed formulation is sound, and show that it is able to out-perform a current state-of-the-art HMM/GMM system. We also use a hybrid Tandem-like system to further understand the source of the benefit. David Imseng, Hervé Bourlard, Philip N. Garner |
ICASSP | 3 |
| 2012 | Combining vocal tract length normalization with hierarchial linear transformationsabstractRecent research has demonstrated the effectiveness of vocal tract length normalization (VTLN) as a rapid adaptation technique for statistical parametric speech synthesis. VTLN produces speech with naturalness preferable to that of MLLR-based adaptation techniques, being much closer in quality to that generated by the original average voice model. However with only a single parameter, VTLN captures very few speaker specific characteristics when compared to linear transform based adaptation techniques. This paper proposes that the merits of VTLN can be combined with those of linear transform based adaptation in a hierarchial Bayesian framework, where VTLN is used as the prior information. A novel technique for propagating the gender information from the VTLN prior through constrained structural maximum a posteriori linear regression (CSMAPLR) adaptation is presented. Experiments show that the resulting transformation has improved speech quality with better naturalness, intelligibility and improved speaker similarity. Lakshmi Babu Saheer, Junichi Yamagishi, Philip N. Garner, John Dines |
ICASSP | 3 |
| 2012 | Comparing different acoustic modeling techniques for multilingual boostingabstractIn this paper, we explore how different acoustic modeling techniques can benefit from data in languages other than the target language. We propose an algorithm to perform decision tree state clustering for the recently proposed Kullback-Leibler divergence based hidden Markov models (KL-HMM) and compare it to subspace Gaussian mixture modeling (SGMM). KL-HMM can exploit multilingual information in the form of universal phoneme posterior features and SGMM benefits from a universal background model that can be trained on multilingual data. Taking the Greek SpeechDat(II) data as an example, we show that KL-HMM performs best for small amounts of target language data. David Imseng, John Dines, Petr Motlícek, Philip N. Garner, Hervé Bourlard |
INTERSPEECH | 4 |
| 2012 | Combining cepstral normalization and cochlear implant-like speech processing for microphone array-based speech recognitionabstractThis paper investigates the combination of cepstral normalization and cochlear implant-like speech processing for microphone array-based speech recognition. Testing speech signals are recorded by a circular microphone array and are subsequently processed with superdirective beamforming and McCowan post-filtering. Training speech signals, from the multichannel overlapping Number corpus (MONC), are clean and not overlapping. Cochlear implant-like speech processing, which is inspired from the speech processing strategy in cochlear implants, is applied on the training and testing speech signals. Cepstral normalization, including cepstral mean and variance normalization (CMN and CVN), are applied on the training and testing cepstra. Experiments show that implementing either cepstral normalization or cochlear implant-like speech processing helps in reducing the WERs of microphone array-based speech recognition. Combining cepstral normalization and cochlear implant-like speech processing reduces further the WERs, when there is overlapping speech. Train/test mismatches are measured using the Kullback-Leibler divergence (KLD), between the global probability density functions (PDFs) of training and testing cepstral vectors. This measure reveals a train/test mismatch reduction when either cepstral normalization or cochlear implant-like speech processing is used. It reveals also that combining these two processing reduces further the train/test mismatches as well as the WERs. Cong-Thanh Do, Mohammad Javad Taghizadeh, Philip N. Garner |
SLT | 3 |
| 2012 | MediaParl: Bilingual mixed language accented speech databaseabstractMediaParl is a Swiss accented bilingual database containing recordings in both French and German as they are spoken in Switzerland. The data were recorded at the Valais Parliament. Valais is a bilingual Swiss canton with many local accents and dialects. Therefore, the database contains data with high variability and is suitable to study multilingual, accented and non-native speech recognition as well as language identification and language switch detection. We also define monolingual and mixed language automatic speech recognition and language identification tasks and evaluate baseline systems. The database is publicly available for download. David Imseng, Hervé Bourlard, Holger Caesar, Philip N. Garner, Gwénolé Lecorvé, Alexandre Nanchen |
SLT | 4 |
| 2012 | Transcribing Meetings With the AMIDA SystemsabstractIn this paper, we give an overview of the AMIDA systems for transcription of conference and lecture room meetings. The systems were developed for participation in the Rich Transcription evaluations conducted by the National Institute for Standards and Technology in the years 2007 and 2009 and can process close talking and far field microphone recordings. The paper first discusses fundamental properties of meeting data with special focus on the AMI/AMIDA corpora. This is followed by a description and analysis of improved processing and modeling, with focus on techniques specifically addressing meeting transcription issues such as multi-room recordings or domain variability. In 2007 and 2009, two different strategies of systems building were followed. While in 2007 we used our traditional style system design based on cross adaptation, the 2009 systems were constructed semi-automatically, supported by improved decoders and a new method for system representation. Overall these changes gave a 6%-13% relative reduction in word error rate compared to our 2007 results while at the same time requiring less training material and reducing the real-time factor by five times. The meeting transcription systems are available at www.webasr.org. Thomas Hain, Lukás Burget, John Dines, Philip N. Garner, Frantisek Grézl, Asmaa El Hannani, Marijn Huijbregts, Martin Karafiát, Mike Lincoln, Vincent Wan |
IEEE Trans. Speech Audio Process. | 4 |
| 2012 | Vocal Tract Length Normalization for Statistical Parametric Speech SynthesisabstractVocal tract length normalization (VTLN) has been successfully used in automatic speech recognition for improved performance. The same technique can be implemented in statistical parametric speech synthesis for rapid speaker adaptation during synthesis. This paper presents an efficient implementation of VTLN using expectation maximization and addresses the key challenges faced in implementing VTLN for synthesis. Jacobian normalization, high-dimensionality features and truncation of the transformation matrix are a few challenges presented with the appropriate solutions. Detailed evaluations are performed to estimate the most suitable technique for using VTLN in speech synthesis. Evaluating VTLN in the framework of speech synthesis is also not an easy task since the technique does not work equally well for all speakers. Speakers have been selected based on different objective and subjective criteria to demonstrate the difference between systems. The best method for implementing VTLN is confirmed to be use of the lower order features for estimating warping factors. Lakshmi Babu Saheer, John Dines, Philip N. Garner |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | Improving Non-Native ASR Through Stochastic Multilingual Phoneme Space TransformationsabstractWe propose a stochastic phoneme space transformation technique that allows the conversion of conditional source phoneme posterior probabilities (conditioned on the acoustics) into target phoneme posterior probabilities. The source and target phonemes can be in any language and phoneme format such as the International Phonetic Alphabet. The novel technique makes use of a Kullback-Leibler divergence based hidden Markov model and can be applied to non-native and accented speech recognition or used to adapt systems to under-resourced languages. In this paper, and in the context of hybrid HMM/MLP recognizers, we successfully apply the proposed approach to non-native English speech recognition on the HIWIRE dataset. David Imseng, Hervé Bourlard, John Dines, Philip N. Garner, Mathew Magimai-Doss |
INTERSPEECH | 4 |
| 2011 | A Just-in-Time Document Retrieval System for Dialogues or Monologues
Andrei Popescu-Belis, Majid Yazdani, Alexandre Nanchen, Philip N. Garner |
SIGDIAL Conference | 4 |
| 2011 | Cepstral normalisation and the signal to noise ratio spectrum in automatic speech recognition
Philip N. Garner |
Speech Commun. | 1 |
| 2010 | Automatic temporal alignment of AV data with confidence estimationabstractIn this paper, we propose a new approach for the automatic audio-based temporal alignment with confidence estimation of audio-visual data, recorded by different cameras, camcorders or mobile phones during social events. All recorded data is temporally aligned based on ASR-related features with a common master track, recorded by a reference camera, and the corresponding confidence of alignment is estimated. The core of the algorithm is based on perceptual time-frequency analysis with a precision of 10 ms. The results show correct alignment in 99% of cases for a real life dataset and surpass the performance of cross correlation while keeping lower system requirements. Danil Korchagin, Philip N. Garner, John Dines |
ICASSP | 2 |
| 2010 | VTLN adaptation for statistical speech synthesisabstractThe advent of statistical speech synthesis has enabled the unification of the basic techniques used in speech synthesis and recognition. Adaptation techniques that have been successfully used in recognition systems can now be applied to synthesis systems to improve the quality of the synthesized speech. The application of vocal tract length normalization (VTLN) for synthesis is explored in this paper. VTLN based adaptation requires estimation of a single warping factor, which can be accurately estimated from very little adaptation data and gives additive improvements over CMLLR adaptation. The challenge of estimating accurate warping factors using higher order features is solved by initializing warping factor estimation with the values calculated from lower order features. Lakshmi Babu Saheer, Philip N. Garner, John Dines |
ICASSP | 2 |
| 2010 | Sparse component analysis for speech recognition in multi-speaker environmentabstractSparse Component Analysis is a relatively young technique that relies upon a representation of signal occupying only a small part of a larger space. Mixtures of sparse components are disjoint in that space. As a particular application of sparsity of speech signals, we investigate the DUET blind source separation algorithm in the context of speech recognition for multi-party recordings. We show how DUET can be tuned to the particular case of speech recognition with interfering sources, and evaluate the limits of performance as the number of sources increases. We show that the separated speech fits a common metric for sparsity, and conclude that sparsity assumptions lead to good performance in speech separation and hence ought to benefit other aspects of the speech recognition chain. Afsaneh Asaei, Hervé Bourlard, Philip N. Garner |
INTERSPEECH | 3 |
| 2010 | Tracter: a lightweight dataflow frameworkabstractTracter is introduced as a dataflow framework particularly use-ful for speech recognition. It is designed to work on-line in real-time as well as off-line, and is the feature extraction means for the Juicer transducer based decoder. This paper places Tracter in context amongst the dataflow literature and other commer-cial and open source packages. Some design aspects and capa-bilities are discussed. Finally, a fairly large processing graph incorporating voice activity detection and feature extraction is presented as an example of Tracter’s capabilites. Index Terms: Dataflow, speech recognition, open source. 1. Philip N. Garner, John Dines |
INTERSPEECH | 1 |
| 2010 | The AMIDA 2009 meeting transcription systemabstractWe present the AMIDA 2009 system for participation in the NIST RT’2009 STT evaluations. Systems for close-talking, far field and speaker attributed STT conditions are described. Im- provements to our previous systems are: segmentation and diar- isation; stacked bottle-neck posterior feature extraction; fMPE training of acoustic models; adaptation on complete meetings; improvements to WFST decoding; automatic optimisation of decoders and system graphs. Overall these changes gave a 6- 13% relative reduction in word error rate while at the same time reducing the real-time factor by a factor of five and using con- siderably less data for acoustic model training. Thomas Hain, Lukás Burget, John Dines, Philip N. Garner, Asmaa El Hannani, Marijn Huijbregts, Martin Karafiát, Mike Lincoln, Vincent Wan |
INTERSPEECH | 4 |
| 2010 | Hands free audio analysis from home entertainmentabstractIn this paper, we describe a system developed for hands free audio analysis for a living room environment. It comprises detection and localisation of the verbal and paralinguistic events, which can augment the behaviour of virtual director and improve the overall experience of interactions between spatially separated families and friends. The results show good performance in reverberant environments and fulfil real-time requirements. Index Terms: real-time audio processing, direction of arrival, speech meta-data Danil Korchagin, Philip N. Garner, Petr Motlícek |
INTERSPEECH | 2 |
| 2010 | English spoken term detection in multilingual recordingsabstractThis paper investigates the automatic detection of English spoken terms in a multi-language scenario over real lecture recordings. Spoken Term Detection (STD) is based on an LVCSR where the output is represented in the form of word lattices. The lattices are then used to search the required terms. Processed lectures are mainly composed of English, French and Italian recordings where the language can also change within one recording. Therefore, the English STD system uses an Out-Of-Language (OOL) detection module to filter out non-English input segments. OOL detection is evaluated w.r.t. various confidence measures estimated from word lattices. Experimental studies of OOL detection followed by English STD are performed on several hours of multilingual recordings. Significant improvement of OOL+STD over a stand-alone STD system is achieved (relatively more than 50% in EER). Finally, an additional modality (text slides in the form of PowerPoint presentations) is exploited to improve STD. Petr Motlícek, Fabio Valente, Philip N. Garner |
INTERSPEECH | 3 |
| 2009 | SNR features for automatic speech recognitionabstractWhen combined with cepstral normalisation techniques, the features normally used in Automatic Speech Recognition are based on Signal to Noise Ratio (SNR). We show that calculating SNR from the outset, rather than relying on cepstral normalisation to produce it, gives features with a number of practical and mathematical advantages over power-spectral based ones. In a detailed analysis, we derive Maximum Likelihood and Maximum a-Posteriori estimates for SNR based features, and show that they can outperform more conventional ones, especially when subsequently combined with cepstral variance normalisation. We further show anecdotal evidence that SNR based features lend themselves well to noise estimates based on low-energy envelope tracking. Philip N. Garner |
ASRU | 1 |
| 2009 | Real-time ASR from meetingsabstractThe AMI(DA) system is a meeting room speech recognition system that has been developed and evaluated in the context of the NIST Rich Text (RT) evaluations. Recently, the "Distant Access" requirements of the AMIDA project have necessitated that the system operate in real-time. Another more difficult requirement is that the system fit into a live meeting transcription scenario. We describe an infrastructure that has allowed the AMI(DA) system to evolve into one that fulfils these extra requirements. We emphasise the components that address the live and real-time aspects. Philip N. Garner, John Dines, Thomas Hain, Asmaa El Hannani, Martin Karafiát, Danil Korchagin, Mike Lincoln, Vincent Wan, Le Zhang 0002 |
INTERSPEECH | 1 |
| 2009 | Beamforming With a Maximum Negentropy CriterionabstractIn this paper, we address a beamforming application based on the capture of far-field speech data from a single speaker in a real meeting room. After the position of the speaker is estimated by a speaker tracking system, we construct a subband-domain beamformer in generalized sidelobe canceller (GSC) configuration. In contrast to conventional practice, we then optimize the active weight vectors of the GSC so as to obtain an output signal with maximum negentropy (MN). This implies the beamformer output should be as non-Gaussian as possible. For calculating negentropy, we consider the Gamma and the generalized Gaussian (GG) pdfs. After MN beamforming, Zelinski postfiltering is performed to further enhance the speech by removing residual noise. Our beamforming algorithm can suppress noise and reverberation without the signal cancellation problems encountered in the conventional beamforming algorithms. We demonstrate this fact through a set of acoustic simulations. Moreover, we show the effectiveness of our proposed technique through a series of far-field automatic speech recognition experiments on the Multi-Channel Wall Street Journal Audio Visual Corpus (MC-WSJ-AV), a corpus of data captured with real far-field sensors, in a realistic acoustic environment, and spoken by real speakers. On the MC-WSJ-AV evaluation data, the delay-and-sum beamformer with postfiltering achieved a word error rate (WER) of 16.5%. MN beamforming with the Gamma pdf achieved a 15.8% WER, which was further reduced to 13.2% with the GG pdf, whereas the simple delay-and-sum beamformer provided a WER of 17.8%. To the best of our knowledge, no lower error rates at present have been reported in the literature on this automatic speech recognition (ASR) task. Ken'ichi Kumatani, John W. McDonough, Barbara Rauch 0001, Dietrich Klakow, Philip N. Garner, Weifeng Li 0001 |
IEEE Trans. Speech Audio Process. | 5 |
| 2008 | Filter bank design based on minimization of individual aliasing terms for minimum mutual information subband adaptive beamformingabstractThis paper presents new filter bank design methods for sub- band adaptive beamforming. In this work, we design analysis and synthesis prototypes for modulated filter banks so as to minimize each aliasing term individually. We then drive the total response error to null by constraining these prototypes to be Nyquist(M) filters. Thereafter those modulated filter banks are applied to a speech separation system which extracts a target speech signal. In our system, speech signals are first transformed into the subband domain with our filter banks, and the subband components are then processed with a beamforming algorithm. Following beamforming, post-filtering and binary masking are further performed to remove residual noises. We show that our filter banks can suppress the residual aliasing distortion more than conventional ones. Furthermore, we demonstrate the effectiveness of our design techniques through a set of automatic speech recognition experiments on the multi-channel speech data from the PASCAL Speech Separation Challenge. The experimental results prove that our beamforming system with the proposed filter banks achieves the best recognition performance, a 39.6 % word error rate (WER), with half the amount of computation of that of the conventional filter banks while the perfect reconstruction filter banks provided a 44.4 % WER. Ken'ichi Kumatani, John W. McDonough, S. Schachl, Dietrich Klakow, Philip N. Garner, Weifeng Li 0001 |
ICASSP | 5 |
| 2008 | Silence models in weighted finite-state transducersabstractWe investigate the effects of different silence modelling strategies in Weighted Finite-State Transducers for Automatic Speech Recognition. We show that the choice of silence models, and the way they are included in the transducer, can have a significant effect on the size of the resulting transducer; we present a means to prevent particularly large silence overheads. Our conclusions include that context-free silence modelling fits well with transducer based grammars, whereas modelling silence as a monophone and a context has larger overheads. Index Terms: speech recognition, weighted finite-state transducer, silence model Philip N. Garner |
INTERSPEECH | 1 |
| 2008 | Maximum kurtosis beamforming with the generalized sidelobe cancellerabstractThis paper presents an adaptive beamforming application based on the capture of far-field speech data from a real single speaker in a real meeting room. After the position of a speaker is estimated by a speaker tracking system, we construct a subbanddomain beamformer in generalized sidelobe canceller (GSC) configuration. In contrast to conventional practice, we then optimize the active weight vectors of the GSC so that the distribution of an output signal is as non-Gaussian as possible. We consider kurtosis in order to measure the degree of non-Gaussianity. Our beamforming algorithms can suppress noise and reverberation without the signal cancellation problems encountered in conventional beamforming algorithms. We demonstrate the effectiveness of our proposed techniques through a series of farfield automatic speech recognition experiments on the Multi-Channel Wall Street Journal Audio Visual Corpus (MC-WSJ-AV). The beamforming algorithm proposed here achieved a 13.6 % WER, whereas the simple delay-and-sum beamformer provided a WER of 17.8%. Index Terms: far-field speech recognition, microphone array, beamforming 1. Ken'ichi Kumatani, John W. McDonough, Barbara Rauch 0001, Philip N. Garner, Weifeng Li 0001, John Dines |
INTERSPEECH | 4 |
| 2004 | A differential spectral voice activity detectorabstractThe voice activity detection (VAD) problem is placed into a decision theoretic framework, and the Gaussian VAD model of Sohn et al. (1998, 1999) is then shown to fit well with the framework. It is argued that the Gaussian model can be made more robust to correlation and expected spectral shapes of speech and noise by using a differential spectral representation. Such a model is formulated theoretically. The differential spectral VAD is then shown by experiment to compare favourably with the basic Gaussian VAD in a speech recognition setting, especially for noisy environments. Philip N. Garner, Toshiaki Fukada, Yasuhiro Komori |
ICASSP (1) | 1 |
| 2001 | SpokenContent representation in MPEG-7abstractThe words spoken in an audio-visual document form an obvious and intuitive metadata component. This component is essential to ensure comprehensive coverage of audio-visual content by the MPEG-7 standard. With manual transcription prohibitively costly, such metadata will typically be derived from automatic speech recognition systems. The errors inherent in the output of such extraction tools cause particular difficulties for robust retrieval, as well as for interoperability in heterogeneous databases. We describe a structure comprising a probabilistic combined word and phone lattice along with an explanatory metadata header and detail how this structure avoids or ameliorates these problems. Jason P. A. Charlesworth, Philip N. Garner |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2000 | Representation and linking mechanisms for audio in MPEG-7
Adam T. Lindsay, Savitha Srinivasan, Jason P. A. Charlesworth, Philip N. Garner, Werner Kriechbaum |
Signal Process. Image Commun. | 4 |
| 1998 | On the robust incorporation of formant features into hidden Markov models for automatic speech recognitionabstractA formant analyser is interpreted probabilistically via a noisy channel model. This leads to a robust method of incorporating formant features into hidden Markov models for automatic speech recognition. Recognition equations follow trivially, and Baum-Welch style re-estimation equations are derived. Experimental results are presented which provide empirical proof of convergence, and demonstrate the effectiveness of the technique in achieving recognition performance advantages by including formant features rather than only using cepstrum features. Philip N. Garner, Wendy J. Holmes |
ICASSP | 1 |
| 1997 | A keyword selection strategy for dialogue move recognition and multi-class topic identificationabstractThe concept of usefulness for keyword selection in topic identification problems is reformulated and extended to the multi-class domain. The derivation is shown to be a generalisation of that for the two class problem. The technique is applied to both multinomial and Poisson based estimates of word probability, and shown to outperform or compare favourably to various information theoretic techniques classifying dialogue moves in the map task corpus, and reports in the LOB corpus. Philip N. Garner, Aidan Hemsworth |
ICASSP | 1 |
| 1997 | Using formant frequencies in speech recognitionabstractFormant frequencies have rarely been used as acoustic features for speech recognition, in spite of their phonetic significance. For some speech sounds one or more of the formants may be so badly defined that it is not useful to attempt a frequency measurement. Also, it is often difficult to decide which formant labels to attach to particular spectral peaks. This paper describes a new method of formant analysis which includes techniques to overcome both of the above difficulties. Using the same data and HMM model structure, results are compared between a recognizer using conventional cepstrum features and one using three formant frequencies, combined with fewer cepstrum features to represent general spectral trends. For the same total number of features, results show that including formant features can offer increased accuracy over using cepstrum features only. John N. Holmes, Wendy J. Holmes, Philip N. Garner |
EUROSPEECH | 3 |
| 1997 | On topic identification and dialogue move recognition
Philip N. Garner |
Comput. Speech Lang. | 1 |
| 1996 | Source position estimation using radial basis functionsabstractThe problem of bearing estimation using radar focal plane arrays is addressed. Motivated by a requirement for a compact integrated solution in the sensor focal plane, a basis function approach is adopted. The development of a solution that is robust to sensor noise leads to a ridge regression-type estimate for the model parameters. The approach is assessed using radially symmetric kernel functions on simulated radar data. Andrew R. Webb, Philip N. Garner |
ICPR | 2 |
| 1996 | A theory of word frequencies and its application to dialogue move recognitionabstractDialogue move recognition is taken as being representative of a class of spoken language applications where inference about high level semantic meaning is required from lower level acoustic, phonetic or word based features.Topic identication is another such application.In the particular case of inference from words, the multinomial distribution is shown to be inadequate for modelling word frequencies, and the multivariate Poisson is a more reasonable choice.Zipf's law is used to model a prior distribution.This more rigorous mathematical formulation is shown to improve dialogue move classication both subjectively and quantitatively. Philip N. Garner, Sue Browning, Roger K. Moore, Martin J. Russell |
ICSLP | 1 |