VLDB 2026 Research / reviewers in the wild / expert
Jon Barker
dblp:50/2212 · also Jon P. Barker
· DBLP profile ↗
103ranked-venue papers
20as first author
26since 2021 · last 2027
0000-0002-1684-5660ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 85 · 14 first-author · 21 since 2021Artificial intelligence and machine learning · 65 · 15 first-author · 15 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | The first Clarity Enhancement Challenge: Developing hearing aid algorithms for speech-in-noiseabstractHearing aid users frequently struggle to understand speech in noisy environments, negatively impacting their quality of life. Inspired by recent progress in speech technology through community-driven machine learning challenges, the Clarity project launched the first-ever Clarity Enhancement Challenge (CEC1). This challenge specifically addressed speech-in-noise enhancement for hearing aids, uniquely combining objective and subjective intelligibility evaluations to assess performance. Participants developed algorithms aimed at improving speech intelligibility in a simulated domestic environment featuring a target speaker and a stationary noise interferer—either competing speech or domestic appliances. Competitors were provided with an open-source dataset, comprising a novel 40-speaker British English corpus, realistic domestic noise samples, and a baseline hearing aid model with basic signal processing. This paper describes the design and outcomes of CEC1. Thirteen entries were evaluated objectively using the Modified Binaural Short-Time Objective Intelligibility metric (MBSTOI) and subjectively by a listening panel of hearing-impaired individuals. The majority of systems employed deep neural networks (DNNs), classical beamforming, or a combination of both. Results showed significant intelligibility gains over the baseline, particularly for systems combining adaptive beamforming with neural network-based noise reduction. However, algorithms optimised directly for MBSTOI scores did not always translate to real-world listening benefits, highlighting the critical importance of perceptual evaluation in assessing intelligibility. These findings underscore the potential of machine-learning-driven approaches for enhancing hearing aid performance and set a foundation for future challenges addressing dynamic and more realistic auditory scenarios. Simone Graetzer, Michael A. Akeroyd, Jon Barker, Trevor J. Cox, John F. Culling, Jennifer Firth, Graham Naylor, Eszter Porter, Rhoddy Viveros Muñoz |
Comput. Speech Lang. | 3 |
| 2026 | Raw acoustic-articulatory multimodal dysarthric speech recognitionabstractAutomatic speech recognition (ASR) for dysarthric speech is challenging. The acoustic characteristics of dysarthric speech are highly variable and there are often fewer distinguishing cues between phonetic tokens. Multimodal ASR utilises the data from other modalities to facilitate the task when a single acoustic modality proves insufficient. Articulatory information, which encapsulates knowledge about the speech production process, may constitute such a complementary modality. Although multimodal acoustic-articulatory ASR has received increasing attention recently, incorporating real articulatory data is under-explored for dysarthric speech recognition. This paper investigates the effectiveness of multimodal acoustic modelling using real dysarthric speech articulatory information in combination with acoustic features, especially raw signal representations which are more informative than classic features, leading to learning representations tailored to dysarthric ASR. In particular, various raw acoustic-articulatory multimodal dysarthric speech recognition systems are developed and compared with similar systems with hand-crafted features. Furthermore, the difference between dysarthric and typical speech in terms of articulatory information is systematically analysed by using a statistical space distribution indicator called Maximum Articulator Motion Range (MAMR). Additionally, we used mutual information analysis to investigate the robustness and phonetic information content of the articulatory features, offering insights that support feature selection and the ASR results. Experimental results on the widely used TORGO dysarthric speech dataset show that combining the articulatory and raw acoustic features at the empirically found optimal fusion level achieves a notable performance gain, leading to up to 7.6% and 12.8% relative word error rate (WER) reduction for dysarthric and typical speech, respectively. Zhengjun Yue, Erfan Loweimi, Zoran Cvetkovic, Jon Barker, Heidi Christensen |
Comput. Speech Lang. | 4 |
| 2025 | Objective and subjective evaluation of speech enhancement methods in the UDASE task of the 7th CHiME challengeabstractSupervised models for speech enhancement are trained using artificially generated mixtures of clean speech and noise signals. However, the synthetic training conditions may not accurately reflect real-world conditions encountered during testing. This discrepancy can result in poor performance when the test domain significantly differs from the synthetic training domain. To tackle this issue, the UDASE task of the 7th CHiME challenge aimed to leverage real-world noisy speech recordings from the test domain for unsupervised domain adaptation of speech enhancement models. Specifically, this test domain corresponds to the CHiME-5 dataset, characterized by real multi-speaker and conversational speech recordings made in noisy and reverberant domestic environments, for which ground-truth clean speech signals are not available. In this paper, we present the objective and subjective evaluations of the systems that were submitted to the CHiME-7 UDASE task, and we provide an analysis of the results. This analysis reveals a limited correlation between subjective ratings and several supervised nonintrusive performance metrics recently proposed for speech enhancement. Conversely, the results suggest that more traditional intrusive objective metrics can be used for in-domain performance evaluation using the reverberant LibriCHiME-5 dataset developed for the challenge. The subjective evaluation indicates that all systems successfully reduced the background noise, but always at the expense of increased distortion. Out of the four speech enhancement methods evaluated subjectively, only one demonstrated an improvement in overall quality compared to the unprocessed noisy speech, highlighting the difficulty of the task. The tools and audio material created for the CHiME-7 UDASE task are shared with the community. Simon Leglaive, Matthieu Fraticelli, Hend Elghazaly, Léonie Borne, Mostafa Sadeghi, Scott Wisdom, Manuel Pariente, John R. Hershey, Daniel Pressnitzer, Jon Barker |
Comput. Speech Lang. | 10 |
| 2024 | The 2nd Clarity Prediction Challenge: A Machine Learning Challenge for Hearing Aid Intelligibility PredictionabstractThis paper reports on the design and outcomes of the 2nd Clarity Prediction Challenge (CPC2) for predicting the intelligibility of hearing aid processed signals heard by individuals with a hearing impairment. The challenge was designed to promote new approaches for estimating the intelligibility of hearing aid signals that can be used in future hearing aid algorithm development. It extends an earlier round (CPC1, 2022) in a number of critical directions, including a larger dataset coming from new speech intelligibility listening experiments, a greater degree of variability in the test materials, and a design that requires prediction systems to generalise to unseen algorithms and listeners. This paper provides a full description of the new publicly available CPC2 dataset, the CPC2 challenge design, and the baseline systems. The challenge attracted 12 systems from 9 research teams. The systems are reviewed, their performance is analysed and conclusions are presented, with reference to the progress made since the earlier CPC1 challenge. In particular, it is seen how reference-free, non-intrusive systems based on pre-trained large acoustic models can perform well in this context. Jon Barker, Michael A. Akeroyd, Will Bailey, Trevor J. Cox, John F. Culling, Jennifer Firth, Simone Graetzer, Graham Naylor |
ICASSP | 1 |
| 2024 | Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users Using Intermediate ASR Features and Human Memory ModelsabstractNeural networks have been successfully used for non-intrusive speech intelligibility prediction. Recently, the use of feature representations sourced from intermediate layers of pre-trained self-supervised and weakly-supervised models has been found to be particularly useful for this task. This work combines the use of Whisper ASR decoder layer representations as neural network input features with an exemplar-based, psychologically motivated model of human memory to predict human intelligibility ratings for hearing-aid users. Substantial performance improvement over an established intrusive HASPI baseline system is found, including on enhancement systems and listeners unseen in the training data, with a root mean squared error of 25.3 compared with the baseline of 28.7. Rhiannon Mogridge, George Close, Robert Sutherland, Thomas Hain, Jon Barker, Stefan Goetze, Anton Ragni |
ICASSP | 5 |
| 2024 | Leveraging Bitstream Metadata for Fast, Accurate, Generalized Compressed Video Quality EnhancementabstractVideo compression is a central feature of the modern internet powering technologies from social media to video conferencing. While video compression continues to mature, for many compression settings, quality loss is still noticeable. These settings nevertheless have important applications to the efficient transmission of videos over bandwidth constrained or otherwise unstable connections. In this work, we develop a deep learning architecture capable of restoring detail to compressed videos which leverages the underlying structure and motion information embedded in the video bitstream. We show that this improves restoration accuracy compared to prior compression correction methods and is competitive when compared with recent deep-learning-based video compression methods on rate-distortion while achieving higher throughput. Furthermore, we condition our model on quantization data which is readily available in the bit-stream. This allows our single model to handle a variety of different compression quality settings which required an ensemble of models in prior work. Max Ehrlich, Jon Barker, Namitha Padmanabhan, Larry Davis 0001, Andrew Tao, Bryan Catanzaro, Abhinav Shrivastava |
WACV | 2 |
| 2023 | The 2nd Clarity Enhancement Challenge for Hearing Aid Speech Intelligibility Enhancement: Overview and OutcomesabstractThis paper reports on the design and outcomes of the 2nd Clarity Enhancement Challenge (CEC2), a challenge for stimulating novel approaches to hearing-aid speech intelligibility enhancement. The challenge was for a listener attending to a target speaker in a noisy, domestic environment. The challenge extends the previous edition, CEC1, in a number of key respects: scenes have multiple interferers including speech, noise and music; ambisonics are used to model listener head movement; target speaker identity is provided to encourage speaker extraction approaches. Systems are evaluated both via the HASPI intelligibility metric and with listening tests using a panel of hearing-impaired listeners. The paper reviews the 18 systems that were submitted describing them in terms of their enhancement and amplification stages. HASPI is seen to be a good predictor of listener performance. The top system, using carefully engineered neural approaches, produces highly intelligible signals for complex scenes with SNRs down to -12 dB while obeying the challenges 5 ms latency constraint. Michael A. Akeroyd, Will Bailey, Jon Barker, Trevor J. Cox, John F. Culling, Simone Graetzer, Graham Naylor, Zuzanna Podwinska, Zehai Tu |
ICASSP | 3 |
| 2023 | Overview of the 2023 ICASSP SP Clarity Challenge: Speech Enhancement for Hearing AidsabstractThis paper reports on the design and outcomes of the ICASSP SP Clarity Challenge: Speech Enhancement for Hearing Aids. The scenario was a listener attending to a target speaker in a noisy, domestic environment. There were multiple interferers and head rotation by the listener. The challenge extended the second Clarity Enhancement Challenge (CEC2) by fixing the amplification stage of the hearing aid; evaluating with a combined metric for speech intelligibility and quality; and providing two evaluation sets, one based on simulation and the other on real-room measurements. Five teams improved on the baseline system for the simulated evaluation set, but the performance on the measured evaluation set was much poorer. Investigations are on-going to determine the exact cause of the mismatch between the simulated and measured data sets. The presence of transducer noise in the measurements, lower order Ambisonics harming the ability for systems to exploit binaural cues and the differences between real and simulated room impulse responses are suggested causes. Trevor J. Cox, Jon Barker, Will Bailey, Simone Graetzer, Michael A. Akeroyd, John F. Culling, Graham Naylor |
ICASSP | 2 |
| 2022 | Improved Simulation of Realistically-Spatialised Simultaneous Speech Using Multi-Camera Analysis in The Chime-5 DatasetabstractRoom simulation is an essential tool in the development of distant microphone ASR and source separation. However, most commonly used simulated datasets adopt uninformed and potentially unrealistic speaker location distributions. In earlier work, we analysed a 50-hour audio-visual dataset of multiparty recordings made in real homes to estimate typical angular separations between speakers. We now refine and extend this work using a multi-camera analysis to estimate full 2-D speaker location distributions. Results show that commonly used simulated datasets use unrealistically large angular separations, but unrealistically small ranges for target to interferer distance ratios. We generate more realistically distributed datasets and use them to re-evaluate state-of-the-art source separation and ASR approaches. Results suggest that imposing realistic angular separation distributions makes datasets more challenging, however, the pattern when using realistic distance ratios is more complicated and can depend on room size. Jack Deadman, Jon Barker |
ICASSP | 2 |
| 2022 | Auditory-Based Data Augmentation for end-to-end Automatic Speech RecognitionabstractEnd-to-end models have achieved significant improvement on automatic speech recognition. One common method to improve performance of these models is expanding the data-space through data augmentation. Meanwhile, human auditory inspired front-ends have also demonstrated improvement for automatic speech recognisers. In this work, a well-verified auditory-based model, which can simulate various hearing abilities, is investigated for the purpose of data augmentation for end-to-end speech recognition. By introducing the auditory model into the data augmentation process, end-to-end systems are encouraged to ignore variation from the signal that cannot be heard and thereby focus on robust features for speech recognition. Two mechanisms in the auditory model, spectral smearing and loudness recruitment, are studied on the LibriSpeech dataset with a transformer-based end-to-end model. The results show that the proposed augmentation methods can bring statistically significant improvement on the performance of the state-of-the-art SpecAugment. Zehai Tu, Jack Deadman, Ning Ma 0002, Jon Barker |
ICASSP | 4 |
| 2022 | Multi-Modal Acoustic-Articulatory Feature Fusion For Dysarthric Speech RecognitionabstractBuilding automatic speech recognition (ASR) systems for speakers with dysarthria is a very challenging task. Although multi-modal ASR has received increasing attention recently, incorporating real articulatory data with acoustic features has not been widely explored in the dysarthric speech community. This paper investigates the effectiveness of multi-modal acoustic modelling for dysarthric speech recognition using acoustic features along with articulatory information. The proposed multi-stream architectures consist of convolutional, recurrent and fully-connected layers allowing for bespoke per-stream pre-processing, fusion at the optimal level of abstraction and post-processing. We study the optimal fusion level/scheme as well as training dynamics in terms of cross-entropy and WER using the popular TORGO dysarthric speech database. Experimental results show that fusing the acoustic and articulatory features at the empirically found optimal level of abstraction achieves a remarkable performance gain, leading to up to 4.6% absolute (9.6% relative) WER reduction for speakers with dysarthria. Zhengjun Yue, Erfan Loweimi, Zoran Cvetkovic, Heidi Christensen, Jon Barker |
ICASSP | 5 |
| 2022 | The 1st Clarity Prediction Challenge: A machine learning challenge for hearing aid intelligibility prediction
Jon Barker, Michael A. Akeroyd, Trevor J. Cox, John F. Culling, Jennifer Firth, Simone Graetzer, Holly Griffiths, Lara Harris, Graham Naylor, Zuzanna Podwinska, Eszter Porter, Rhoddy Viveros Muñoz |
INTERSPEECH | 1 |
| 2022 | Modelling Turn-taking in Multispeaker Parties for Realistic Data Simulation
Jack Deadman, Jon Barker |
INTERSPEECH | 2 |
| 2022 | Exploiting Hidden Representations from a DNN-based Speech Recogniser for Speech Intelligibility Prediction in Hearing-impaired ListenersabstractAn accurate objective speech intelligibility prediction algorithms is of great interest for many applications such as speech enhancement for hearing aids.Most algorithms measures the signal-to-noise ratios or correlations between the acoustic features of clean reference signals and degraded signals.However, these hand-picked acoustic features are usually not explicitly correlated with recognition.Meanwhile, deep neural network (DNN) based automatic speech recogniser (ASR) is approaching human performance in some speech recognition tasks.This work leverages the hidden representations from DNN-based ASR as features for speech intelligibility prediction in hearingimpaired listeners.The experiments based on a hearing aid intelligibility database show that the proposed method could make better prediction than a widely used short-time objective intelligibility (STOI) based binaural measure. Zehai Tu, Ning Ma 0002, Jon Barker |
INTERSPEECH | 3 |
| 2022 | Unsupervised Uncertainty Measures of Automatic Speech Recognition for Non-intrusive Speech Intelligibility PredictionabstractNon-intrusive intelligibility prediction is important for its application in realistic scenarios, where a clean reference signal is difficult to access. The construction of many non-intrusive predictors require either ground truth intelligibility labels or clean reference signals for supervised learning. In this work, we leverage an unsupervised uncertainty estimation method for predicting speech intelligibility, which does not require intelligibility labels or reference signals to train the predictor. Our experiments demonstrate that the uncertainty from state-of-the-art end-to-end automatic speech recognition (ASR) models is highly correlated with speech intelligibility. The proposed method is evaluated on two databases and the results show that the unsupervised uncertainty measures of ASR models are more correlated with speech intelligibility from listening results than the predictions made by widely used intrusive methods. Zehai Tu, Ning Ma 0002, Jon Barker |
INTERSPEECH | 3 |
| 2022 | Dysarthric Speech Recognition From Raw Waveform with Parametric CNNsabstractRaw waveform acoustic modelling has recently received increasing attention. Compared with the task-blind hand-crafted features which may discard useful information, representations directly learned from the raw waveform are task-specific and potentially include all task-relevant information. In the context of automatic dysarthric speech recognition (ADSR), raw waveform acoustic modelling is under-explored owing to data scarcity. Parametric convolutional neural networks (CNNs) can compensate for this problem due to having notably fewer parameters and requiring less training data in comparison with conventional non-parametric CNNs. In this paper, we explore the usefulness of raw waveform acoustic modelling using various parametric CNNs for ADSR. We investigate the properties of the learned filters and monitor the training dynamics of various models. Furthermore, we study the effectiveness of data augmentation and multi-stream acoustic modelling through combining the non-parametric and parametric CNNs fed by hand-crafted and raw waveform features. Experimental results on the TORGO dysarthric database show that the parametric CNNs significantly outperform the non-parametric CNNs, reaching up to 36.2% and 12.6% WERs (up to 3.4% and 1.1% absolute error reduction) for dysarthric and typical speech, respectively. Multi-stream acoustic modelling further improves the performance resulting in up to 33.2% and 10.3% WERs for dysarthric and typical speech, respectively. Zhengjun Yue, Erfan Loweimi, Heidi Christensen, Jon Barker, Zoran Cvetkovic |
INTERSPEECH | 4 |
| 2022 | On monoaural speech enhancement for automatic recognition of real noisy speech using mixture invariant trainingabstractIn this paper, we explore an improved framework to train a monoaural neural enhancement model for robust speech recognition. The designed training framework extends the existing mixture invariant training criterion to exploit both unpaired clean speech and real noisy data. It is found that the unpaired clean speech is crucial to improve quality of separated speech from real noisy speech. The proposed method also performs remixing of processed and unprocessed signals to alleviate the processing artifacts. Experiments on the single-channel CHiME-3 real test sets show that the proposed method improves significantly in terms of speech recognition performance over the enhancement system trained either on the mismatched simulated data in a supervised fashion or on the matched real data in an unsupervised fashion. Between 16% and 39% relative WER reduction has been achieved by the proposed system compared to the unprocessed signal using end-to-end and hybrid acoustic models without retraining on distorted data. Jisi Zhang, Catalin Zorila, Rama Sanand Doddipatla, Jon Barker |
INTERSPEECH | 4 |
| 2022 | SNuC: The Sheffield Numbers Spoken Language CorpusabstractWe present SNuC, the first published corpus of spoken alphanumeric identifiers of the sort typically used as serial and part numbers in the manufacturing sector. The dataset contains recordings and transcriptions of over 50 native British English speakers, speaking over 13,000 multi-character alphanumeric sequences and totalling almost 20 hours of recorded speech. We describe requirements taken into account in the designing the corpus and the methodology used to construct it. We present summary statistics describing the corpus contents, as well as a preliminary investigation into errors in spoken alphanumeric identifiers. We validate the corpus by showing how it can be used to adapt a deep learning neural network based ASR system, resulting in improved recognition accuracy on the task of spoken alphanumeric identifier recognition. Finally, we discuss further potential uses for the corpus and for the tools developed to construct it. Emma Barker, Jon Barker, Robert J. Gaizauskas, Ning Ma 0002, Monica Lestari Paramita |
LREC | 2 |
| 2022 | Acoustic Modelling From Raw Source and Filter Components for Dysarthric Speech RecognitionabstractAcoustic modelling for automatic dysarthric speech recognition (ADSR) is a challenging task. Data deficiency is a major problem and substantial differences between typical and dysarthric speech complicate the transfer learning. In this paper, we aim at building acoustic models using the raw magnitude spectra of the source and filter components for ADSR. The proposed multi-stream models consist of convolutional, recurrent and fully-connected layers allowing for pre-processing various information streams and fusing them at an optimal level of abstraction. We demonstrate that such a multi-stream processing leverages information encoded in the vocal tract and excitation components and leads to normalising nuisance factors such as speaker attributes and speaking style. This leads to a better handling of dysarthric speech that exhibits large inter- and intra-speaker variabilities and results in a notable performance gain. Furthermore, we analyse the learned convolutional filters and visualise the outputs of different layers after dimensionality reduction to demonstrate how the speaker-related attributes are normalised along the pipeline. We also compare the proposed multi-stream model with various systems based on MFCC, FBank, raw waveform and i-vector, and, study the training dynamics as well as usefulness of the feature normalisation and data augmentation via speed perturbation. On the widely used TORGO and UASpeech dysarthric speech corpora, the proposed approach leads to a competitive performance of up to 35.3% and 30.3% WERs for dysarthric speech, respectively. Zhengjun Yue, Erfan Loweimi, Heidi Christensen, Jon Barker, Zoran Cvetkovic |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | The use of Voice Source Features for Sung Speech RecognitionabstractIn this paper, we ask whether vocal source features (pitch, shimmer, jitter, etc) can improve the performance of automatic sung speech recognition, arguing that conclusions previously drawn from spoken speech studies may not be valid in the sung speech domain. We first use a parallel singing/speaking corpus (NUS-48E) to illustrate differences in sung vs spoken voicing characteristics including pitch range, syllables duration, vibrato, jitter and shimmer. We then use this analysis to inform speech recognition experiments on the sung speech DSing corpus, using a state of the art acoustic model and augmenting conventional features with various voice source parameters. Experiments are run with three standard (increasingly large) training sets, DSing1 (15.1 hours), DSing3 (44.7 hours) and DS-ing30 (149.1 hours). Pitch combined with degree of voicing produces a significant decrease in WER from 38.1% to 36.7% when training with DSing1 however smaller decreases in WER observed when training with the larger more varied DSing3 and DSing30 sets were not seen to be statistically significant. Voicing quality characteristics did not improve recognition performance although analysis suggests that they do contribute to an improved discrimination between voiced/unvoiced phoneme pairs. Gerardo Roa Dabike, Jon Barker |
ICASSP | 2 |
| 2021 | DHASP: Differentiable Hearing Aid Speech ProcessingabstractHearing aids are expected to improve speech intelligibility for listeners with hearing impairment. An appropriate amplification fitting tuned for the listener’s hearing disability is critical for good performance. The developments of most prescriptive fittings are based on data collected in subjective listening experiments, which are usually expensive and time-consuming. In this paper, we explore an alternative approach to finding the optimal fitting by introducing a hearing aid speech processing framework, in which the fitting is optimised in an automated way using an intelligibility objective function based on the HASPI physiological auditory model. The framework is fully differentiable, thus can employ the back-propagation algorithm for efficient, data-driven optimisation. Our initial objective experiments show promising results for noise-free speech amplification, where the automatically optimised processors outperform one of the well recognised hearing aid prescriptions. Zehai Tu, Ning Ma 0002, Jon Barker |
ICASSP | 3 |
| 2021 | Time-Domain Speech Extraction with Spatial Information and Multi Speaker Conditioning MechanismabstractIn this paper, we present a novel multi-channel speech extraction system to simultaneously extract multiple clean individual sources from a mixture in noisy and reverberant environments. The proposed method is built on an improved multi-channel time-domain speech separation network which employs speaker embeddings to identify and extract multiple targets without label permutation ambiguity. To efficiently inform the speaker information to the extraction model, we propose a new speaker conditioning mechanism by designing an additional speaker branch for receiving external speaker embeddings. Experiments on 2-channel WHAMR! data show that the proposed system improves by 9% relative the source separation performance over a strong multi-channel baseline, and it increases the speech recognition accuracy by more than 16% relative over the same baseline. Jisi Zhang, Catalin Zorila, Rama Sanand Doddipatla, Jon Barker |
ICASSP | 4 |
| 2021 | Clarity-2021 Challenges: Machine Learning Challenges for Advancing Hearing Aid ProcessingabstractIn recent years, rapid advances in speech technology have been made possible by machine learning challenges such as CHiME, REVERB, Blizzard, and Hurricane. In the Clarity project, the machine learning approach is applied to the problem of hearing aid processing of speech-in-noise, where current technology in enhancing the speech signal for the hearing aid wearer is often ineffective. The scenario is a (simulated) cuboid-shaped living room in which there is a single listener, a single target speaker and a single interferer, which is either a competing talker or domestic noise. All sources are static, the target is always within ±30° azimuth of the listener and at the same elevation, and the interferer is an omnidirectional point source at the same elevation. The target speech comes from an open source 40-speaker British English speech database collected for this purpose. This paper provides a baseline description of the round one Clarity challenges for both enhancement (CEC1) and prediction (CPC1). To the authors’ knowledge, these are the first machine learning challenges to consider the problem of hearing aid speech signal processing. Simone Graetzer, Jon Barker, Trevor J. Cox, Michael A. Akeroyd, John F. Culling, Graham Naylor, Eszter Porter, Rhoddy Viveros Muñoz |
Interspeech | 2 |
| 2021 | Optimising Hearing Aid Fittings for Speech in Noise with a Differentiable Hearing Loss ModelabstractThis is a repository copy of Optimising hearing aid fittings for speech in noise with a differentiable hearing loss model. Zehai Tu, Ning Ma 0002, Jon Barker |
Interspeech | 3 |
| 2021 | Parental Spoken Scaffolding and Narrative Skills in Crowd-Sourced Storytelling Samples of Young ChildrenabstractA novel crowdsourcing project to gather children's storytelling based language samples using a mobile app was undertaken across the United Kingdom. Parents' scaffolding of children's narratives was observed in many of the samples. This study was designed to examine the relationship of scaffolding and young children's narrative language ability in a story retell context which is analysed at the macro-structural (total macro-structure score), the micro-structural (mean length of utterances in morphemes) and verbal productivity (total number of utterances) levels. Young children with and without scaffolding were statistically compared. The interaction between the level of scaffolding support, the grammar complexity and the narrative structure was explored. A bidirectional relationship was observed between scaffolding and young children's narrative language ability. Young children with better performance were observed to receive less scaffolding from parents. Scaffolding was shown to support early narrative development of young children and was more able to benefit those with low-level grammatical complexity skills. It is crucial to encourage parental scaffolding to be well-attuned to the child's narrative ability. Zhengjun Yue, Jon Barker, Heidi Christensen, Cristina McKean, Elaine Ashton, Yvonne Wren, Swapnil Gadgil, Rebecca Bright |
Interspeech | 2 |
| 2021 | Teacher-Student MixIT for Unsupervised and Semi-Supervised Speech SeparationabstractIn this paper, we introduce a novel semi-supervised learning framework for end-to-end speech separation. The proposed method first uses mixtures of unseparated sources and the mixture invariant training (MixIT) criterion to train a teacher model. The teacher model then estimates separated sources that are used to train a student model with standard permutation invariant training (PIT). The student model can be fine-tuned with supervised data, i.e., paired artificial mixtures and clean speech sources, and further improved via model distillation. Experiments with single and multi channel mixtures show that the teacher-student training resolves the over-separation problem observed in the original MixIT method. Further, the semisupervised performance is comparable to a fully-supervised separation system trained using ten times the amount of supervised data. Jisi Zhang, Catalin Zorila, Rama Sanand Doddipatla, Jon Barker |
Interspeech | 4 |
| 2020 | Source Domain Data Selection for Improved Transfer Learning Targeting Dysarthric Speech RecognitionabstractThis paper presents an improved transfer learning framework applied to robust personalised speech recognition models for speakers with dysarthria. As the baseline of transfer learning, a state-of-the-art CNN-TDNN-F ASR acoustic model trained solely on source domain data is adapted onto the target domain via neural network weight adaptation with the limited available data from target dysarthric speakers. Results show that linear weights in neural layers play the most important role for an improved modelling of dysarthric speech evaluated using UASpeech corpus, achieving averaged 11.6% and 7.6% relative recognition improvement in comparison to the conventional speaker-dependent training and data combination, respectively. To further improve the transferability towards target domain, we propose an utterance-based data selection of the source domain data based on the entropy of posterior probability, which is analysed to statistically obey a Gaussian distribution. Compared to a speaker-based data selection via dysarthria similarity measure, this allows for a more accurate selection of the potentially beneficial source domain data for either increasing the target domain training pool or constructing an intermediate domain for incremental transfer learning, resulting in a further absolute recognition performance improvement of nearly 2% added to transfer learning baseline for speakers with moderate to severe dysarthria. Feifei Xiong, Jon Barker, Zhengjun Yue, Heidi Christensen |
ICASSP | 2 |
| 2020 | Exploring Appropriate Acoustic and Language Modelling Choices for Continuous Dysarthric Speech RecognitionabstractThere has been much recent interest in building continuous speech recognition systems for people with severe speech impairments, e.g., dysarthria. However, the datasets that are commonly used are typically designed for tasks other than ASR development, or they contain only isolated words. As such, they contain much overlap in the prompts read by the speakers. Previous ASR evaluations have often neglected this, using language models (LMs) trained on non-disjoint training and test data, potentially producing unrealistically optimistic results. In this paper, we investigate the impact of LM design using the widely used TORGO database. We combine state-of-the-art acoustic models with LMs trained with data originating from LibriSpeech. Using LMs with varying vocabulary size, we examine the trade-off between the out-of-vocabulary rate and recognition confusions for speakers with varying degrees of dysarthria. It is found that the optimal LM complexity is highly speaker dependent, highlighting the need to design speaker-dependent LMs alongside speaker-dependent acoustic models when considering atypical speech. Zhengjun Yue, Feifei Xiong, Heidi Christensen, Jon Barker |
ICASSP | 4 |
| 2020 | On End-to-end Multi-channel Time Domain Speech Separation in Reverberant EnvironmentsabstractThis paper introduces a new method for multi-channel time domain speech separation in reverberant environments. A fully-convolutional neural network structure has been used to directly separate speech from multiple microphone recordings, with no need of conventional spatial feature extraction. To reduce the influence of reverberation on spatial feature extraction, a dereverberation pre-processing method has been applied to further improve the separation performance. A spatialized version of wsj0-2mix dataset has been simulated to evaluate the proposed system. Both source separation and speech recognition performance of the separated signals have been evaluated objectively. Experiments show that the proposed fully-convolutional network improves the source separation metric and the word error rate (WER) by more than 13% and 50% relative, respectively, over a reference system with conventional features. Applying dereverberation as pre-processing to the proposed system can further reduce the WER by 29% relative using an acoustic model trained on clean and reverberated data. Jisi Zhang, Catalin Zorila, Rama Sanand Doddipatla, Jon Barker |
ICASSP | 4 |
| 2020 | Simulating Realistically-Spatialised Simultaneous Speech Using Video-Driven Speaker Detection and the CHiME-5 Dataset
Jack Deadman, Jon Barker |
INTERSPEECH | 2 |
| 2020 | Autoencoder Bottleneck Features with Multi-Task Optimisation for Improved Continuous Dysarthric Speech RecognitionabstractAutomatic recognition of dysarthric speech is a very challenging research problem where performances still lag far behind those achieved for typical speech. The main reason is the lack of suitable training data to accommodate for the large mismatch seen between dysarthric and typical speech. Only recently has focus moved from single-word tasks to exploring continuous speech ASR needed for dictation and most voice-enabled interfaces. This paper investigates improvements to dysarthric continuous ASR. In particular, we demonstrate the effectiveness of using unsupervised autoencoder-based bottleneck (AE-BN) feature extractor trained on out-of-domain (OOD) LibriSpeech data. We further explore multi-task optimisation techniques shown to benefit typical speech ASR. We propose a 5-fold cross-training setup on the widely used TORGO dysarthric database. A setup we believe is more suitable for this low-resource data domain. Results show that adding the proposed AE-BN features achieves an average absolute (word error rate) WER improvement of 2.63% compared to the baseline system. A further reduction of 2.33% and 0.65% absolute WER is seen when applying monophone regularisation and joint optimisation techniques, respectively. In general, the ASR system employing monophone regularisation trained on AE-BN features exhibits the best performance. Zhengjun Yue, Heidi Christensen, Jon Barker |
INTERSPEECH | 3 |
| 2019 | Phonetic Analysis of Dysarthric Speech Tempo and Applications to Robust Personalised Dysarthric Speech RecognitionabstractImproving the accuracy of personalised speech recognition for speakers with dysarthria is a challenging research field. In this paper, we explore an approach that non-linearly modifies speech tempo to reduce mismatch between typical and atypical speech. Speech tempo analysis at the phonetic level is accomplished using a forced-alignment process from traditional GMM-HMM in automatic speech recognition (ASR). Estimated tempo adjustments are applied directly to the acoustic features rather than to the time-domain signals. Two approaches are considered: i) adjusting dysarthric speech towards typical speech for input into ASR systems trained with typical speech, and ii) adjusting typical speech towards dysarthric speech for data augmentation in personalised dysarthric ASR training. Experimental results show that the latter strategy with data augmentation is more effective, resulting in a nearly 7% absolute improvement in comparison to baseline speaker-dependent trained system evaluated using UASpeech corpus. Consistent recognition performance improvements are observed across speakers, with greatest benefit in cases of moderate and severe dysarthria. Feifei Xiong, Jon Barker, Heidi Christensen |
ICASSP | 2 |
| 2019 | Automatic Lyric Transcription from Karaoke Vocal Tracks: Resources and a Baseline SystemabstractAutomatic sung speech recognition is a relatively understudied topic that has been held back by a lack of large and freely available datasets. This has recently changed thanks to the release of the DAMP Sing! dataset, a 1100 hour karaoke dataset originating from the social music-making company, Smule. This paper presents work undertaken to define an easily replicable, automatic speech recognition benchmark for this data. In particular, we describe how transcripts and alignments have been recovered from Karaoke prompts and timings; how suitable training, development and test sets have been defined with varying degrees of accent variability; and how language models have been developed using lyric data from the LyricWikia website. Initial recognition experiments have been performed using factored-layer TDNN acoustic models with lattice-free MMI training using Kaldi. The best WER is 19.60% - a new state-of-the-art for this type of data. The paper concludes with a discussion of the many challenging problems that remain to be solved. Dataset definitions and Kaldi scripts have been made available so that the benchmark is easily replicable. Gerardo Roa Dabike, Jon Barker |
INTERSPEECH | 2 |
| 2018 | SDC-Net: Video Prediction Using Spatially-Displaced Convolution
Fitsum A. Reda, Guilin Liu, Kevin J. Shih, Robert Kirby 0001, Jon Barker, David Tarjan, Andrew Tao, Bryan Catanzaro |
ECCV (7) | 5 |
| 2018 | Exploring the Use of Group Delay for Generalised VTS Based Noise CompensationabstractIn earlier work we studied the effect of statistical normalisation for phase-based features and observed it leads to a significant robustness improvement. This paper explores the extension of the generalised Vector Taylor Series (gVTS) noise compensation approach to the group delay (GD) domain. We discuss the problems it presents, propose some solutions and derive the corresponding formulae. Furthermore, the effects of additive and channel noise in the GD domain were studied. It was observed that the GD of the noisy observation is a convex combination of the GDs of the clean signal and the additive noise and also in the expected sense, channel GD tends to zero. Experiments on Aurora-4 showed that, despite training only on the clean speech, the proposed features provide average WER reductions of 0.8% absolute and 4.1% relative compared to an MFCC-based system trained on the multi-style data. Combining the gVTS with a bottleneck DNN-based system led to average absolute (relative) WER improvements of 6.0% (23.5%) when training on clean data and 2.5% (13.8%) when using multi-style training with additive noise. Erfan Loweimi, Jon Barker, Thomas Hain |
ICASSP | 2 |
| 2018 | The Fifth 'CHiME' Speech Separation and Recognition Challenge: Dataset, Task and BaselinesabstractThe CHiME challenge series aims to advance robust automatic speech recognition (ASR) technology by promoting research at the interface of speech and language processing, signal processing , and machine learning. This paper introduces the 5th CHiME Challenge, which considers the task of distant multi-microphone conversational ASR in real home environments. Speech material was elicited using a dinner party scenario with efforts taken to capture data that is representative of natural conversational speech and recorded by 6 Kinect microphone arrays and 4 binaural microphone pairs. The challenge features a single-array track and a multiple-array track and, for each track, distinct rankings will be produced for systems focusing on robustness with respect to distant-microphone capture vs. systems attempting to address all aspects of the task including conversational language modeling. We discuss the rationale for the challenge and provide a detailed description of the data collection procedure, the task, and the baseline systems for array synchronization, speech enhancement, and conventional and end-to-end ASR. Jon Barker, Shinji Watanabe 0001, Emmanuel Vincent 0001, Jan Trmal |
INTERSPEECH | 1 |
| 2018 | DNN Driven Speaker Independent Audio-Visual Mask Estimation for Speech SeparationabstractHuman auditory cortex excels at selectively suppressing background noise to focus on a target speaker. The process of selective attention in the brain is known to contextually exploit the available audio and visual cues to better focus on target speaker while filtering out other noises. In this study, we propose a novel deep neural network (DNN) based audiovisual (AV) mask estimation model. The proposed AV mask estimation model contextually integrates the temporal dynamics of both audio and noise-immune visual features for improved mask estimation and speech separation. For optimal AV features extraction and ideal binary mask (IBM) estimation, a hybrid DNN architecture is exploited to leverages the complementary strengths of a stacked long short term memory (LSTM) and convolution LSTM network. The comparative simulation results in terms of speech quality and intelligibility demonstrate significant performance improvement of our proposed AV mask estimation model as compared to audio-only and visual-only mask estimation approaches for both speaker dependent and independent scenarios. Mandar Gogate, Ahsan Adeel, Ricard Marxer, Jon Barker, Amir Hussain 0001 |
INTERSPEECH | 4 |
| 2018 | On the Usefulness of the Speech Phase Spectrum for Pitch ExtractionabstractMost frequency domain techniques for pitch extraction such as cepstrum, harmonic product spectrum (HPS) and summation residual harmonics (SRH) operate on the magnitude spectrum and turn it into a function in which the fundamental frequency emerges as argmax. In this paper, we investigate the extension of these three techniques to the phase and group delay (GD) domains. Our extensions exploit the observation that the bin at which F (magnitude) becomes maximum, for some monotonically increasing function F, is equivalent to bin at which F (phase) has maximum negative slope and F (group delay) has the maximum value. To extract the pitch track from speech phase spectrum, these techniques were coupled with the source-filter model in the phase domain that we proposed in earlier publications and a novel voicing detection algorithm proposed here. The accuracy and robustness of the phase-based pitch extraction techniques are illustrated and compared with their magnitude-based counterparts using six pitch evaluation metrics. On average, it is observed that the phase spectrum can be successfully employed in pitch tracking with comparable accuracy and robustness to the speech magnitude spectrum. Erfan Loweimi, Jon Barker, Thomas Hain |
INTERSPEECH | 2 |
| 2018 | The impact of the Lombard effect on audio and visual speech recognition systemsabstractWhen producing speech in noisy backgrounds talkers reflexively adapt their speaking style in ways that increase speech-in-noise intelligibility. This adaptation, known as the Lombard effect, is likely to have an adverse effect on the performance of automatic speech recognition systems that have not been designed to anticipate it. However, previous studies of this impact have used very small amounts of data and recognition systems that lack modern adaptation strategies. This paper aims to rectify this by using a new audio-visual Lombard corpus containing speech from 54 different speakers – significantly larger than any previously available – and modern state-of-the-art speech recognition techniques. The paper is organised as three speech-in-noise recognition studies. The first examines the case in which a system is presented with Lombard speech having been exclusively trained on normal speech. It was found that the Lombard mismatch caused a significant decrease in performance even if the level of the Lombard speech was normalised to match the level of normal speech. However, the size of the mismatch was highly speaker-dependent thus explaining conflicting results presented in previous smaller studies. The second study compares systems trained in matched conditions (i.e., training and testing with the same speaking style). Here the Lombard speech affords a large increase in recognition performance. Part of this is due to the greater energy leading to a reduction in noise masking, but performance improvements persist even after the effect of signal-to-noise level difference is compensated. An analysis across speakers shows that the Lombard speech energy is spectro-temporally distributed in a way that reduces energetic masking, and this reduction in masking is associated with an increase in recognition performance. The final study repeats the first two using a recognition system training on visual speech. In the visual domain, performance differences are not confounded by differences in noise masking. It was found that in matched-conditions Lombard speech supports better recognition performance than normal speech. The benefit was consistently present across all speakers but to a varying degree. Surprisingly, the Lombard benefit was observed to a small degree even when training on mismatched non-Lombard visual speech, i.e., the increased clarity of the Lombard speech outweighed the impact of the mismatch. The paper presents two generally applicable conclusions: i) systems that are designed to operate in noise will benefit from being trained on well-matched Lombard speech data, ii) the results of speech recognition evaluations that employ artificial speech and noise mixing need to be treated with caution: they are overly-optimistic to the extent that they ignore a significant source of mismatch but at the same time overly-pessimistic in that they do not anticipate the potential increased intelligibility of the Lombard speaking style. Ricard Marxer, Jon Barker, Najwa Alghamdi, Steve C. Maddock |
Speech Commun. | 2 |
| 2017 | Statistical normalisation of phase-based feature representation for robust speech recognitionabstractIn earlier work we have proposed a source-filter decomposition of speech through phase-based processing. The decomposition leads to novel speech features that are extracted from the filter component of the phase spectrum. This paper analyses this spectrum and the proposed representation by evaluating statistical properties at various points along the parametrisation pipeline. We show that speech phase spectrum has a bell-shaped distribution which is in contrast to the uniform assumption that is usually made. It is demonstrated that the uniform density (which implies that the corresponding sequence is least-informative) is an artefact of the phase wrapping and not an original characteristic of this spectrum. In addition, we extend the idea of statistical normalisation usually applied for the magnitudebased features into the phase domain. Based on the statistical structure of the phase-based features, which is shown to be super-gaussian in the clean condition, three normalisation schemes, namely, Gaussianisation, Laplacianisation and table-based histogram equalisation have been applied for improving the robustness. Speech recognition experiments using Aurora-2 show that applying an optimal normalisation scheme at the right stage of the feature extraction process can produce average relative WER reductions of up to 18.6% across the 0-20 dB SNR conditions. Erfan Loweimi, Jon Barker, Thomas Hain |
ICASSP | 2 |
| 2017 | Channel Compensation in the Generalised Vector Taylor Series Approach to Robust ASRabstractVector Taylor Series (VTS) is a powerful technique for robust ASR but, in its standard form, it can only be applied to log-filter bank and MFCC features. In earlier work, we presented a generalised VTS (gVTS) that extends the applicability of VTS to front-ends which employ a power transformation non-linearity. gVTS was shown to provide performance improvements in both clean and additive noise conditions. This paper makes two novel contributions. Firstly, while the previous gVTS formulation assumed that noise was purely additive, we now derive gVTS formulae for the case of speech in the presence of both additive noise and channel distortion. Second, we propose a novel iterative method for estimating the channel distortion which utilises gVTS itself and converges after a few iterations. Since the new gVTS blindly assumes the existence of both additive noise and channel effects, it is important not to introduce extra distortion when either are absent. Experimental results conducted on LVCSR Aurora-4 database show that the new formulation passes this test. In the presence of channel noise only, it provides relative WER reductions of up to 30% and 26%, compared with previous gVTS and multi-style training with cepstral mean normalisation, respectively. Erfan Loweimi, Jon Barker, Thomas Hain |
INTERSPEECH | 2 |
| 2017 | Robust Source-Filter Separation of Speech Signal in the Phase DomainabstractIn earlier work we proposed a framework for speech source-filter separation that employs phase-based signal processing. This paper presents a further theoretical investigation of the model and optimisations that make the filter and source representations less sensitive to the effects of noise and better matched to downstream processing. To this end, first, in computing the Hilbert transform, the log function is replaced by the generalised logarithmic function. This introduces a tuning parameter that adjusts both the dynamic range and distribution of the phase-based representation. Second, when computing the group delay, a more robust estimate for the derivative is formed by applying a regression filter instead of using sample differences. The effectiveness of these modifications is evaluated in clean and noisy conditions by considering the accuracy of the fundamental frequency extracted from the estimated source, and the performance of speech recognition features extracted from the estimated filter. In particular, the proposed filter-based front-end reduces Aurora-2 WERs by 6.3% (average 0-20 dB) compared with previously reported results. Furthermore, when tested in a LVCSR task (Aurora-4) the new features resulted in 5.8% absolute WER reduction compared to MFCCs without performance loss in the clean/matched condition. Erfan Loweimi, Jon Barker, Oscar Saz-Torralba, Thomas Hain |
INTERSPEECH | 2 |
| 2017 | Binary Mask Estimation Strategies for Constrained Imputation-Based Speech EnhancementabstractInternational audience Ricard Marxer, Jon Barker |
INTERSPEECH | 2 |
| 2017 | Multi-microphone speech recognition in everyday environments
Jon Barker, Ricard Marxer, Emmanuel Vincent 0001, Shinji Watanabe 0001 |
Comput. Speech Lang. | 1 |
| 2017 | The third 'CHiME' speech separation and recognition challenge: Analysis and outcomes
Jon Barker, Ricard Marxer, Emmanuel Vincent 0001, Shinji Watanabe 0001 |
Comput. Speech Lang. | 1 |
| 2017 | An analysis of environment, microphone and data simulation mismatches in robust speech recognition
Emmanuel Vincent 0001, Shinji Watanabe 0001, Aditya Arie Nugraha, Jon Barker, Ricard Marxer |
Comput. Speech Lang. | 4 |
| 2017 | The impact of automatic exaggeration of the visual articulatory features of a talker on the intelligibility of spectrally distorted speech
Najwa Alghamdi, Steve C. Maddock, Jon Barker, Guy J. Brown |
Speech Commun. | 3 |
| 2016 | Language Effects in Noise-Induced Word MisperceptionsabstractInternational audience María Luisa García Lecumberri, Jon Barker, Ricard Marxer, Martin Cooke |
INTERSPEECH | 2 |
| 2016 | Use of Generalised Nonlinearity in Vector Taylor Series Noise Compensation for Robust Speech RecognitionabstractDesigning good normalisation to counter the effect of environmental distortions is one of the major challenges for automatic speech recognition (ASR). The Vector Taylor series (VTS) method is a powerful and mathematically well principled technique that can be applied to both the feature and model domains to compensate for both additive and convolutional noises. One of the limitations of this approach, however, is that it is tied to MFCC (and log-filterbank) features and does not extend to other representations such as PLP, PNCC and phase-based front-ends that use power transformation rather than log compression. This paper aims at broadening the scope of the VTS method by deriving a new formulation that assumes a power transformation is used as the non-linearity during feature extraction. It is shown that the conventional VTS, in the log domain, is a special case of the new extended framework. In addition, the new formulation introduces one more degree of freedom which makes it possible to tune the algorithm to better fit the data to the statistical requirements of the ASR back-end. Compared with MFCC and conventional VTS, the proposed approach provides up to 12.2% and 2.0% absolute performance improvements on average, in Aurora-4 tasks, respectively. Erfan Loweimi, Jon Barker, Thomas Hain |
INTERSPEECH | 2 |
| 2016 | Multichannel Spatial Clustering for Robust Far-Field Automatic Speech Recognition in Mismatched Conditions
Michael I. Mandel, Jon Barker |
INTERSPEECH | 2 |
| 2016 | Misperceptions Arising from Speech-in-Babble Interactions
Máté Attila Tóth, Martin Cooke, Jon Barker |
INTERSPEECH | 3 |
| 2015 | The third 'CHiME' speech separation and recognition challenge: Dataset, task and baselinesabstractThe CHiME challenge series aims to advance far field speech recognition technology by promoting research at the interface of signal processing and automatic speech recognition. This paper presents the design and outcomes of the 3rd CHiME Challenge, which targets the performance of automatic speech recognition in a real-world, commercially-motivated scenario: a person talking to a tablet device that has been fitted with a six-channel microphone array. The paper describes the data collection, the task definition and the baseline systems for data simulation, enhancement and recognition. The paper then presents an overview of the 26 systems that were submitted to the challenge focusing on the strategies that proved to be most successful relative to the MVDR array processing and DNN acoustic modeling reference system. Challenge findings related to the role of simulated data in system training and evaluation are discussed. Jon Barker, Ricard Marxer, Emmanuel Vincent 0001, Shinji Watanabe 0001 |
ASRU | 1 |
| 2015 | Exploiting synchrony spectra and deep neural networks for noise-robust automatic speech recognitionabstractThis paper presents a novel system that exploits synchrony spectra and deep neural networks (DNNs) for automatic speech recognition (ASR) in challenging noisy environments. Synchrony spectra measure the extent to which each frequency channel in an auditory model is entrained to a particular pitch period, and they are used together with F0 estimates either in a DNN for time-frequency (T-M) mask estimation or to augment the input features for a DNN-based ASR system. The proposed approach was evaluated in the context of the CHiME 3 Challenge. Our experiments show that the synchrony spectra features work best when augmenting the input features to the DNN-based ASR system. Compared to the CHiME-3 baseline system, our best system provides a word error rate (WER) reduction of more than 14% absolute and achieved a WER of 18.56% on the evaluation test set. Ning Ma 0002, Ricard Marxer, Jon Barker, Guy J. Brown |
ASRU | 3 |
| 2015 | The effect of cochlear implant processing on speaker intelligibility: a perceptual study and computer model
Jon Barker, Guy J. Brown |
INTERSPEECH | 2 |
| 2015 | Source-filter separation of speech signal in the phase domainabstractDeconvolution of the speech excitation (source) and vocal tract (filter) components through log-magnitude spectral processing is well-established and has led to the well-known cepstral features used in a multitude of speech processing tasks. This paper presents a novel source-filter decomposition based on processing in the phase domain. We show that separation between source and filter in the log-magnitude spectra is far from perfect, leading to loss of vital vocal tract information. It is demonstrated that the same task can be better performed by trend and fluctuation analysis of the phase spectrum of the minimum-phase component of speech, which can be computed via the Hilbert transform. Trend and fluctuation can be separated through low-pass filtering of the phase, using additivity of vocal tract and source in the phase domain. This results in separated signals which have a clear relation to the vocal tract and excitation components. The effectiveness of the method is put to test in a speech recognition task. The vocal tract component extracted in this way is used as the basis of a feature extraction algorithm for speech recognition on the Aurora-2 database. The recognition results shows upto 8.5% absolute improvement in comparison with MFCC features on average (0-20dB). Erfan Loweimi, Jon Barker, Thomas Hain |
INTERSPEECH | 2 |
| 2015 | A framework for the evaluation of microscopic intelligibility modelsabstractInternational audience Ricard Marxer, Martin Cooke, Jon Barker |
INTERSPEECH | 3 |
| 2014 | Speech pre-enhancement using a discriminative microscopic intelligibility model
Maryam M. Al Dabel, Jon Barker |
INTERSPEECH | 2 |
| 2013 | The second 'CHiME' speech separation and recognition challenge: An overview of challenge systems and outcomesabstractDistant-microphone automatic speech recognition (ASR) remains a challenging goal in everyday environments involving multiple background sources and reverberation. This paper reports on the results of the 2nd ‘CHiME’ Challenge, an initiative designed to analyse and evaluate the performance of ASR systems in a real-world domestic environment. We discuss the rationale for the challenge and provide a summary of the datasets, tasks and baseline systems. The paper overviews the systems that were entered for the two challenge tracks: small-vocabulary with moving talker and medium-vocabulary with stationary talker. We present a summary of the challenge findings including novel results produced by challenge system combination. Possible directions for future challenges are discussed. Emmanuel Vincent 0001, Jon Barker, Shinji Watanabe 0001, Jonathan Le Roux, Francesco Nesta, Marco Matassoni |
ASRU | 2 |
| 2013 | The second 'chime' speech separation and recognition challenge: Datasets, tasks and baselinesabstractDistant-microphone automatic speech recognition (ASR) remains a challenging goal in everyday environments involving multiple background sources and reverberation. This paper is intended to be a reference on the 2nd `CHiME' Challenge, an initiative designed to analyze and evaluate the performance of ASR systems in a real-world domestic environment. Two separate tracks have been proposed: a small-vocabulary task with small speaker movements and a medium-vocabulary task without speaker movements. We discuss the rationale for the challenge and provide a detailed description of the datasets, tasks and baseline performance results for each track. Emmanuel Vincent 0001, Jon Barker, Shinji Watanabe 0001, Jonathan Le Roux, Francesco Nesta, Marco Matassoni |
ICASSP | 2 |
| 2013 | Special issue on speech separation and recognition in multisource environments
Jon Barker, Emmanuel Vincent 0001 |
Comput. Speech Lang. | 1 |
| 2013 | The PASCAL CHiME speech separation and recognition challenge
Jon Barker, Emmanuel Vincent 0001, Ning Ma 0002, Heidi Christensen, Phil D. Green |
Comput. Speech Lang. | 1 |
| 2013 | A hearing-inspired approach for distant-microphone speech recognition in the presence of multiple sources
Ning Ma 0002, Jon Barker, Heidi Christensen, Phil D. Green |
Comput. Speech Lang. | 2 |
| 2013 | Speech Spectral Envelope Enhancement by HMM-Based Analysis/ResynthesisabstractWe propose a speech enhancement-by-resynthesis framework whose strength lies in a common statistical speech model that is shared by the analysis and synthesis stages. First, a spectro-temporal analysis is performed and masked spectro-temporal regions are identified using a noise model. Then, HMM synthesis is used to reconstruct the spectral envelope in masked regions in a manner which is conditioned on the reliable regions, preventing the resynthesis from regressing to the training data mean. As a demonstration we enhance noise-corrupted speech utterances from a small vocabulary corpus for which good statistical models are available. Perceptual evaluation of speech quality and log spectral distances demonstrate considerable performance improvements over baseline approaches that do not exploit strong speech knowledge. The letter is accompanied by audio examples. José L. Carmona, Jon Barker, Ángel M. Gómez, Ning Ma 0002 |
IEEE Signal Process. Lett. | 2 |
| 2013 | MMSE-Based Missing-Feature Reconstruction With Temporal Modeling for Robust Speech RecognitionabstractThis paper addresses the problem of feature compensation in the log-spectral domain by using the missing-data (MD) approach to noise robust speech recognition, that is, the log-spectral features can be either almost unaffected by noise or completely masked by it. First, a general MD framework based on minimum mean square error (MMSE) estimation is introduced which exploits the correlation across frequency bands to reconstruct the missing features. This framework allows the derivation of different MD imputation approaches and, in particular, a novel technique taking advantage of truncated Gaussian distributions is presented. While the proposed technique provides excellent results at high and medium signal-to-noise ratios (SNRs), its performance diminishes at low SNRs where very few reliable features are available. The reconstruction technique is therefore extended to exploit temporal constraints using two different approaches. In the first approach, time-frequency patches of speech containing a number of consecutive frames are modeled using a Gaussian mixture model (GMM). In the second one, the sequential structure of speech is alternatively modeled by a hidden Markov model (HMM). The proposed techniques are evaluated on Aurora-2 and Aurora-4 databases using both oracle and estimated masks. In both cases, the proposed techniques outperform the recognition performance obtained by the baseline system and other related techniques. Also, the introduction of a temporal modeling turns out to be very effective in reconstructing spectra at low SNRs. In particular, HMMs show the highest capability of accounting for time correlations and, therefore, achieve the best results. José A. González 0001, Antonio M. Peinado, Ning Ma 0002, Ángel M. Gómez, Jon Barker |
IEEE Trans. Speech Audio Process. | 5 |
| 2012 | Combining missing-data reconstruction and uncertainty decoding for robust speech recognitionabstractThis paper proposes a novel approach for noise-robust speech recognition which combines a missing-data (MD) derived spectral reconstruction technique and uncertainty decoding based on the weighted Viterbi algorithm (WVA). First, the noisy feature vectors are compensated by using a novel MD imputation technique based on the integration of truncated Gaussian pdfs. Although the proposed MD estimator has both the advantages of MD techniques and the use of cepstral features, it may still be affected by a number of uncertainty sources. In order to deal with these uncertainties, WVA-based uncertainty decoding is proposed. Our experiments on the Aurora-2 and Aurora-4 tasks show that the proposed MD estimator outperforms other MD imputation techniques. Also, we show that the combination of MD imputation with WVA provides better results than the combination with other uncertainty processing techniques such as the use of evidence pdfs for the estimated features. José A. González 0001, Antonio M. Peinado, Ángel M. Gómez, Ning Ma 0002, Jon Barker |
ICASSP | 5 |
| 2012 | Coupling identification and reconstruction of missing features for noise-robust automatic speech recognitionabstractThe standard missing feature imputation approach to noiserobust automatic speech recognition requires that a single foreground/background segmentation mask is identified prior to reconstruction. This paper presents a novel imputation approach which more closely couples the identification and reconstruction of missing features by using a probabilistic framework based on the speech fragment decoding technique. Using fragment decoding, the most joint-likely state sequence and segmentation hypothesis is identified with which the missing data region is imputed. Crucially, however, imputation can exploit the speech state sequence recovered by the fragment decoding. Further, using N -best decodings allows the clean spectrogram to be estimated as a weighted combination of reconstructions which provides some allowance for uncertainty in the estimates. Experiments on the PASCAL CHiME Challenge task show that system performance is highly dependent on the complexity of the speech models used for segmentation and imputation, and by exploiting the temporal constraint of speech the system significantly outperforms those that ignore the constraint. Ning Ma 0002, Jon Barker |
INTERSPEECH | 2 |
| 2012 | Indication of slowly moving ground targets in non-Gaussian clutter using multi-channel synthetic aperture radarabstractThe problem of how best to maximise the ratio of mean target intensity to mean background intensity for slowly moving targets in sets of multi-channel synthetic aperture radar images is discussed for highly non-Gaussian background clutter. The problem is formulated as a direct maximisation of target-to-clutter ratio thus giving a true maximisation of that ratio. Complex-valued weights derived using generalised eigensystem theory are used to maximise the ratio of quadratic forms representing the mean intensity of the target and background derived from their coherence matrices. For two to four channels it is shown that when the target is highly coherent an optimum steering vector is a discrete Fourier transform. For more than two channels it is shown that the optimal solution is only valid within a subspace of the whole parameter space defined by the correlation parameters of the background clutter. Images from a publically released ground moving target indicator dataset are filtered using the results for three channels. The method outperform a standard space-time adaptive processing algorithm in suppressing the stationary background urban clutter image intensity relative to the image intensity because of a known slowly moving ground vehicle. Moreover, the steering vector is much simpler to implement. Brian Barber, Jon Barker |
IET Signal Process. | 2 |
| 2012 | Combining Speech Fragment Decoding and Adaptive Noise Floor ModelingabstractThis paper presents a novel noise-robust automatic speech recognition (ASR) system that combines aspects of the noise modeling and source separation approaches to the problem. The combined approach has been motivated by the observation that the noise backgrounds encountered in everyday listening situations can be roughly characterized as a slowly varying noise floor in which there are embedded a mixture of energetic but unpredictable acoustic events. Our solution combines two complementary techniques. First, an adaptive noise floor model estimates the degree to which high-energy acoustic events are masked by the noise floor (represented by a soft missing data mask). Second, a fragment decoding system attempts to interpret the high-energy regions that are not accounted for by the noise floor model. This component uses models of the target speech to decide whether fragments should be included in the target speech stream or not. Our experiments on the CHiME corpus task show that the combined approach performs significantly better than systems using either the noise model or fragment decoding approach alone, and substantially outperforms multicondition training. Ning Ma 0002, Jon Barker, Heidi Christensen, Phil D. Green |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | A pitch based noise estimation technique for robust speech recognition with Missing DataabstractThis paper presents a noise estimation technique based on knowledge of pitch information for robust speech recognition. In the first stage the noise is estimated by means of extrapolating the noise from frames where speech is believed to be absent. These frames are detected with a proposed pitch based VAD (Voice Activity Detector). In the second stage the noise estimation is revised in voiced frames using harmonic tunnelling technique. The tunnelling noise estimation is used at high SNRs as an upper bound of the noise rather than a suitable estimation. A spectrogram MD (Missing Data) recognition system is chosen to evaluate the proposed noise estimation. The proposed system is compared in Aurora-2 with other similar techniques like cepstral SS (Spectral Subtraction). Juan Andres Morales-Cordovilla, Ning Ma 0002, Victoria E. Sánchez, José L. Carmona, Antonio M. Peinado, Jon Barker |
ICASSP | 6 |
| 2011 | Crowdsourcing for Word Recognition in NoiseabstractAccess to large samples of listeners is an appealing prospect for speech perception researchers, but lack of control over key factors such as listeners ’ linguistic backgrounds and quality of stimulus delivery is a formidable barrier to the application of crowdsourcing. We describe the outcome of a web-based listening experiment designed to discover consistent confusions amongst words presented in noise, alongside an identical task carried out using traditional laboratory methods. Web listeners were graded according based on information they provided as well as via their responses to tokens recognised robustly by a majority of participants. While overall word identification scores even for the best-performing web subset were well below those obtained in the laboratory, word confusions with high levels of cross-listener agreement were obtained nevertheless, suggesting that focused application of crowdsourcing in speech perception can provide useful data for scientific analysis. Index Terms: speech perception, noise, web experiment 1. Martin Cooke, Jon Barker, María Luisa García Lecumberri, Krzysztof Wasilewski |
INTERSPEECH | 2 |
| 2011 | Binaural Cues for Fragment-Based Speech Recognition in Reverberant Multisource EnvironmentsabstractThis paper addresses the problem of speech recognition using distant binaural microphones in reverberant multisource noise conditions. Our scheme employs a two stage fragment decoding approach: first spectro-temporal acoustic source fragments are identified using signal level cues, and second, a hypothesisdriven stage simultaneously searches for the most probable speech/background fragment labelling and the corresponding acoustic model state sequence. The paper reports the first successful attempt to use binaural localisation cues within this framework. By integrating binaural cues and acoustic models in a consistent probabilistic framework, the decoder is able to derive significant recognition performance benefits from fragment location estimates despite their inherent unreliability. Ning Ma 0002, Jon Barker, Heidi Christensen, Phil D. Green |
INTERSPEECH | 2 |
| 2010 | The CHiME corpus: a resource and a challenge for computational hearing in multisource environmentsabstractWe present a new corpus designed for noise-robust speech processing research, CHiME. Our goal was to produce material which is both natural (derived from reverberant domestic environments with many simultaneous and unpredictable sound sources) and controlled (providing an enumerated range of SNRs spanning 20 dB). The corpus includes around 40 hours of background recordings from a head and torso simulator positioned in a domestic setting, and a comprehensive set of binaural impulse responses collected in the same environment. These have been used to add target utterances from the Grid speech recognition corpus into the CHiME domestic setting. Data has been mixed in a manner that produces a controlled and yet natural range of SNRs over which speech separation, enhancement and recognition algorithms can be evaluated. The paper motivates the design of the corpus, and describes the collection and post-processing of the data. We also present a set of baseline recognition results. Heidi Christensen, Jon Barker, Ning Ma 0002, Phil D. Green |
INTERSPEECH | 2 |
| 2010 | Speech fragment decoding techniques for simultaneous speaker identification and speech recognition
Jon Barker, Ning Ma 0002, André Coy, Martin Cooke |
Comput. Speech Lang. | 1 |
| 2009 | A speech fragment approach to localising multiple speakers in reverberant environmentsabstractSound source localisation cues are severely degraded when multiple acoustic sources are active in the presence of reverberation. We present a binaural system for localising simultaneous speakers which exploits the fact that in a speech mixture there exist spectro-temporal regions or dasiafragmentspsila, where the energy is dominated by just one of the speakers. A fragment-level localisation model is proposed that integrates the localisation cues within a fragment using a weighted mean. The weights are based on local estimates of the degree of reverberation in a given spectro-temporal cell. The paper investigates different weight estimation approaches based variously on, i) an established model of the perceptual precedence effect; ii) a measure of interaural coherence between the left and right ear signals; iii) a data-driven approach trained in matched acoustic conditions. Experiments with reverberant binaural data with two simultaneous speakers show appropriate weighting can improve frame-based localisation performance by up to 24%. Heidi Christensen, Ning Ma 0002, Stuart N. Wrigley, Jon Barker |
ICASSP | 4 |
| 2009 | Using location cues to track speaker changes from mobile, binaural microphonesabstractThis paper presents initial developments towards computational hearing models that move beyond stationary microphone assumptions. We present a particle filtering based system for using localisation cues to track speaker changes in meeting recordings. Recording are made using in-ear binaural microphones worn by a listener whose head is constantly moving. Tracking speaker changes requires simultaneously inferring the perceiver’s head orientation, as any change in relative spatial angle to a source can be caused by either the source moving or the microphones moving. In real applications, such as robotics, there may be access to external estimates of the perceiver’s position. We investigate the effect of simulating varying degrees of measurement noise in an external perceiver position estimate. We show that only limited self-position knowledge is needed to greatly improve the reliability with which we can decode the acoustic localisation cues in the meeting scenario. Index Terms: speaker change tracking, binaural hearing, particle filtering, active listening Heidi Christensen, Jon Barker |
INTERSPEECH | 2 |
| 2009 | Energetic and Informational Masking Effects in an Audiovisual Speech Recognition SystemabstractThe paper presents a robust audiovisual speech recognition technique called audiovisual speech fragment decoding. The technique addresses the challenge of recognizing speech in the presence of competing nonstationary noise sources. It employs two stages. First, an acoustic analysis decomposes the acoustic signal into a number of spectro-temporall fragments. Second, audiovisual speech models are used to select fragments belonging to the target speech source. The approach is evaluated on a small vocabulary simultaneous speech recognition task in conditions that promote two contrasting types of masking:energeticmaskingcaused by the energy of the masker utterance swamping that of the target, andinformationalmasking, caused by similarity between the target and masker making it difficult to selectively attend to the correct source. Results show that the system is able to use the visual cues to reduce the effects of both types of masking. Further, whereas recovery fromenergeticmaskingmay require detailed visual information (i.e., sufficient to carry phonetic content), release frominformationalmaskingcan be achieved using very crude visual representations that encode little more than the timing of mouth opening and closure. Jon Barker, Xu Shao |
IEEE Trans. Speech Audio Process. | 1 |
| 2008 | The CAVA corpus: synchronised stereoscopic and binaural datasets with head movementsabstractThis paper describes the acquisition and content of a new multi-modal database. Some tools for making use of the data streams are also presented. The Computational Audio-Visual Analysis (CAVA) database is a unique collection of three synchronised data streams obtained from a binaural microphone pair, a stereoscopic camera pair and a head tracking device. All recordings are made from the perspective of a person; i.e. what would a human with natural head movements see and hear in a given environment. The database is intended to facilitate research into humans' ability to optimise their multi-modal sensory input and fills a gap by providing data that enables human centred audio-visual scene analysis. It also enables 3D localisation using either audio, visual, or audio-visual cues. A total of 50 sessions, with varying degrees of visual and auditory complexity, were recorded. These range from seeing and hearing a single speaker moving in and out of field of view, to moving around a 'cocktail party' style situation, mingling and joining different small groups of people chatting. Elise Arnaud, Heidi Christensen, Yan-Chen Lu, Jon Barker, Vasil Khalidov, Miles E. Hansard, Bertrand Holveck, Hervé Mathieu, Ramya Narasimha, Elise Taillant, Florence Forbes, Radu Horaud |
ICMI | 4 |
| 2008 | Stream weight estimation for multistream audio-visual speech recognition in a multispeaker environment
Xu Shao, Jon Barker |
Speech Commun. | 2 |
| 2007 | Integrating pitch and localisation cues at a speech fragment levelabstractThis paper proposes a novel speech-fragment based approach for processing binaural data to improve the estimation of speech source locations in reverberant, multi-speaker recordings. The technique employs two stages. First, a robust multipitch tracking algorithm is used to locate local spectro-temporal ‘speech fragments ’ – regions where the energy in the mixture is dominated by a single speech source. Second, robust localisation estimates are formed by integrating interaural time difference cues over each speech fragment. The technique is applied to the analysis of more than five hours of two-party meetings that have been constructed from a mixture of binaural mannequin recordings. It is shown that estimating location at the speech fragment level produces better results than conventional location-estimate smoothing techniques leading to a an increase in relative frame accuracy rate of more than 35%. Index Terms: binaural localisation, pitch cues, speech fragment integration Heidi Christensen, Ning Ma 0002, Stuart N. Wrigley, Jon Barker |
INTERSPEECH | 4 |
| 2007 | Applying word duration constraints by using unrolled HMMsabstractConventional HMMs have weak duration constraints. In noisy conditions, the mismatch between corrupted speech signals and models trained on clean speech may cause the decoder to produce word matches with unrealistic durations. This paper presents a simple way to incorporate word duration constraints by unrolling HMMs to form a lattice where word duration probabilities can be applied directly to state transitions. The expanded HMMs are compatible with conventional Viterbi decoding. Experiments on connected-digit recognition show that when using explicit duration constraints the decoder generates word matches with more reasonable durations, and word error rates are significantly reduced across a broad range of noise conditions. Ning Ma 0002, Jon Barker, Phil D. Green |
INTERSPEECH | 2 |
| 2007 | Modelling speaker intelligibility in noise
Jon Barker, Martin Cooke |
Speech Commun. | 1 |
| 2007 | An automatic speech recognition system based on the scene analysis account of auditory perception
André Coy, Jon Barker |
Speech Commun. | 2 |
| 2007 | Exploiting correlogram structure for robust speech recognition with multiple speech sources
Ning Ma 0002, Phil D. Green, Jon Barker, André Coy |
Speech Commun. | 3 |
| 2006 | Speech Separation Based on The Statistics of Binaural Auditory FeaturesabstractA computational auditory scene analysis (CASA) system is described, in which sound separation according to spatial location is combined with the 'missing data' approach for automatic speech recognition. Time-frequency masks for the missing data recognizer are derived from the statistics of interaural time and level differences; these masks identify acoustic features that constitute reliable evidence of the target speech signal. It is demonstrated that this approach yields good performance in a challenging environment, in which a target voice is contaminated by another talker and reverberation. The ability of the system to generalize to source-receiver configurations that were not encountered during training is discussed Guy J. Brown, Sue Harding, Jon Barker |
ICASSP (5) | 3 |
| 2006 | Recognition of Reverberant Speech using Full Cepstral Features and Spectral Missing DataabstractWe describe a novel approach to feature combination within the missing data (MD) framework for automatic speech recognition, and show its application to reverberated speech. Likelihoods from a spectral MD classifier are combined with those from a full cepstral feature vector-based recogniser. Even though the performance of the cepstral recogniser is substantially below that of the MD recogniser, the combined recogniser performs better in all conditions. We also describe improvements to the generation of time-frequency masks for the MD recogniser. Our system is compared with a previous approach based on a hybrid MLP-HMM recogniser with MSG and PLP feature vectors. The proposed system has a substantial performance advantage in the most reverberated conditions Kalle J. Palomäki, Guy J. Brown, Jon Barker |
ICASSP (1) | 3 |
| 2006 | Recent advances in speech fragment decoding techniquesabstractThis paper addresses the problem of recognising speech in the presence of a competing speaker. We employ a speech fragment decoding technique that treats segregation and recognition as coupled problems. Data-driven techniques are used to segment a spectro-temporal representation into a set of spectro-temporal fragments, such that each fragment is dominated by one or other of the speech sources. A speech fragment decoder is used which employs missing data techniques and clean speech models to simultaneously search for the set of fragments and the word sequence that best matches the target speaker model. The paper reports recent advances in this technique, and presents an evaluation based on artificially mixed speech utterances. The fragment decoder produces significantly lower error rates than a conventional recogniser, and mimics the pattern of human performance whereby performance increases as the target-masker ratio is reduced below -3 dB. Index Terms: speech recognition, speech separation, simultaneous speech, auditory scene analysis, noise robustness. Jon Barker, André Coy, Ning Ma 0002, Martin Cooke |
INTERSPEECH | 1 |
| 2006 | A multipitch tracker for monaural speech segmentationabstractThis paper presents a novel algorithm for forming coherent harmonic fragments from a mixture of speech sources. A multiple pitch detection algorithm is used to produce pitch candidates which are tracked using a pair of parallel HMMs. One novel aspect of the technique is that it systematically models pitch doubling and halving errors, thereby facilitating the identification of smooth pitch segments even in the absence of the fundamental frequency. The system does not face the problem of incorrect source assignment that can occur when sources have similar fundamental frequency or are harmonically related. An evaluation of the technique shows that the algorithm’s emphasis on tracking coherent segments leads to the formation of speech fragments with high coherence, indicating a more reliable segmentation of the harmonic speech regions. Index Terms: multiple pitch detection, source separation André Coy, Jon Barker |
INTERSPEECH | 2 |
| 2006 | Audio-visual speech recognition in the presence of a competing speakerabstractThis paper examines the problem of estimating stream weights for a multistream audio-visual speech recogniser in the context of a simultaneous speaker task. The task is challenging because signalto-noise ratio (SNR) cannot be readily inferred from the acoustics alone. The method proposed employs artificial neural networks (ANNs) to estimate the SNR from HMM state-likelihoods. SNR is converted to stream weight using a mapping optimised on development data. The method produces an audio-visual recognition performance better than that of both the audio-only and the videoonly baselines across a wide range of SNRs. The performance using SNR estimates based on audio state-likelihoods is compared to that obtained using both audio and visual likelihoods. Although the audio-visual SNR estimator outperforms the audio-only SNR estimator, the recognition performance benefit is small. Ideas for making fuller use of the visual information are discussed. Index Terms: audio-visual speech recognition, multistream, stream weighting, SNR estimation, artificial neural networks Xu Shao, Jon Barker |
INTERSPEECH | 2 |
| 2006 | Mask estimation for missing data speech recognition based on statistics of binaural interactionabstractThis paper describes a perceptually motivated computational auditory scene analysis (CASA) system that combines sound separation according to spatial location with the "missing data" approach for robust speech recognition in noise. Missing data time-frequency masks are created using probability distributions based on estimates of interaural time and level differences (ITD and ILD) for mixed utterances in reverberated conditions; these masks indicate which regions of the spectrum constitute reliable evidence of the target speech signal. A number of experiments compare the relative efficacy of the binaural cues when used individually and in combination. We also investigate the ability of the system to generalize to acoustic conditions not encountered during training. Performance on a continuous digit recognition task using this method is found to be good, even in a particularly challenging environment with three concurrent male talkers. Sue Harding, Jon Barker, Guy J. Brown |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Tracking Facial Markers with an Adaptive Marker Collocation ModelabstractThe paper presents a robust and computationally low-cost technique for tracking facial markers which exploits an adaptive marker collocation model to recover from tracking errors. Marker collocation statistics are estimated during periods where the markers are successfully tracked, and employed to estimate the position of missing markers during periods where the tracker fails to locate the full set. Evaluation experiments have been conducted on a small audio-visual corpus of connected digits in which the speaker was recorded with small white markers affixed to easily locatable points on the chin, lips, and nose. It is demonstrated that use of the marker collocation model makes the tracker robust in the face of marker occlusion. Jon Barker |
ICASSP (2) | 1 |
| 2005 | Recognising Speech in the Presence of a Competing Speaker using a 'Speech Fragment Decoder'abstractThis paper addresses the problem of recognising speech in the presence of a competing speech source. A novel two stage approach is described. A spectral representation is first divided into a set of spectro-temporal fragments where each fragment is believed to be due to a single acoustic source. An unknown subset of these will be due to the target speech source. The standard ASR search is then extended to find the most likely combination of speech model sequence and fragment subset. The technique is tested with a fragment generation stage using pitch information to locate harmonic energy components, and image processing techniques to segment the inharmonic regions of the spectrogram. The system achieves an accuracy of 65.1% on a 0 dB simultaneous connected digit sequence task with cross-gender mixtures. Extension of the technique to handle matched-gender utterances is discussed. André Coy, Jon Barker |
ICASSP (1) | 2 |
| 2005 | Mask Estimation Based on Sound Localisation for Missing Data Speech RecognitionabstractThis paper describes a perceptually motivated computational auditory scene analysis (CASA) system that combines sound separation according to spatial location with 'missing data' techniques for robust speech recognition in noise. Missing data time-frequency masks are produced using cross-correlation to estimate interaural time difference (ITD) and hence spatial azimuth; this is used to determine which regions of the signal constitute reliable evidence of the target speech signal. Three experiments are performed that compare the effects of different reverberation surfaces, localisation methods and azimuth separations on recognition accuracy, together with the effects of two post-processing techniques (morphological operations and supervised learning) for improving mask estimation. Both post-processing techniques greatly improve performance; the best performance occurs using a learned mapping. Sue Harding, Jon Barker, Guy J. Brown |
ICASSP (1) | 2 |
| 2005 | Soft harmonic masks for recognising speech in the presence of a competing speakerabstractThe paper addresses the problem of recognising speech in the presence of a competing speaker. It uses a two stage ‘Speech Fragment Decoding ’ system. The system works by first segmenting a spectro-temporal representation of the mixture into a number of fragments, such that each fragment is dominated by a single source. An ASR search is then extended to find the combination of speech model sequence and fragment subset that best fits a set of clean speech models. This paper extends previous work by combining ‘Speech Fragment Decoding ’ with soft missing data techniques to better handle spectro-temporal regions that cannot be confidently ascribed to either foreground or background. Recognition experiments are performed on a connected digit task using 0 db mixtures of simultaneous mixedgender speakers. The incorporation of soft decisions leads to an increase in system performance from 66.9 % to 72.2%. 1. André Coy, Jon Barker |
INTERSPEECH | 2 |
| 2005 | Binaural feature selection for missing data speech recognitionabstractThe ‘missing data ’ approach for robust speech recognition uses masks indicating which regions of an acoustic mixture provide reliable evidence of the target to be recognised. Binaural cues for spatial location were used to determine missing data masks for signals consisting of utterances from three concurrent male speakers in reverberant conditions, by deriving probability distributions from estimates of interaural time and level differences (ITD and ILD) for the mixed signals. In such a system, a decision must be made about whether the acoustic features used for decoding are selected from the left or right ear, or a combination of the two. Here, features were selected from the “better ear ” (as determined by a simple heuristic) within whole time frames, or within individual time-frequency elements. A combination of left and right ear features gave better recognition performance than using either ear alone, and the best results were obtained when selecting features within individual time-frequency elements. 1. Sue Harding, Jon Barker, Guy J. Brown |
INTERSPEECH | 2 |
| 2005 | Decoding speech in the presence of other sources
Jon Barker, Martin Cooke, Daniel P. W. Ellis |
Speech Commun. | 1 |
| 2004 | Techniques for handling convolutional distortion with 'missing data' automatic speech recognition
Kalle J. Palomäki, Guy J. Brown, Jon Barker |
Speech Commun. | 3 |
| 2002 | Missing data speech recognition in reverberant conditionsabstractIn this study we describe an auditory processing front-end for missing data speech recognition, which is robust in the presence of reverberation. The model attempts to identify time-frequency regions that are not badly contaminated by reverberation and have strong speech energy. This is achieved by applying reverberation masking. Subsequently, reliable time-frequency regions are passed to a ‘missing data’ speech recogniser for classification. We demonstrate that the model improves recognition performance in three different virtual rooms where reverberation time T60 varies from 0.7 sec to 2.7 sec. We also discuss the advantages of our approach over RASTA and modulation filtered spectrograms. Kalle J. Palomäki, Guy J. Brown, Jon Barker |
ICASSP | 3 |
| 2001 | Robust ASR based on clean speech models: an evaluation of missing data techniques for connected digit recognition in noiseabstractIn this study, techniques for classification with missing or unreliable data are applied to the problem of noise-robustness in Automatic Speech Recognition (ASR). The techniques described make minimal assumptions about any noise background and rely instead on what is known about clean speech. A system is evaluated using the Aurora 2 connected digit recognition task. Using models trained on clean speech we obtain a 65% relative improvement over the Aurora clean training baseline system, a performance comparable with the Aurora baseline for multicondition training. Jon Barker, Martin Cooke, Phil D. Green |
INTERSPEECH | 1 |
| 2000 | Decoding speech in the presence of other sound sourcesabstractConventional speech recognition is notoriously vulnerable to additive noise, and even the best compensation methods are defeated if the noise is nonstationary.To address this problem, we propose a new integration of bottom-up techniques to identify 'coherent fragments' of spectro-temporal energy (based on local features), with the top-down hypothesis search of conventional speech recognition, extended to search also across possible assignments of each fragment as speech or interference.Initial tests demonstrate the feasibility of this approach, and achieve a reduction in word error rate of more than 25% relative at 5 dB SNR over stationary noise missing data recognition. Jon Barker, Martin Cooke, Daniel P. W. Ellis |
INTERSPEECH | 1 |
| 2000 | Soft decisions in missing data techniques for robust automatic speech recognitionabstractIn previous work we have developed the theory and demonstrated the promise of the Missing Data approach to robust Automatic Speech Recognition. This technique is based on hard decisions as to whether each time-frequency "pixel" is either reliable or unreliable. In this paper we replace these discrete decisions with soft estimates of the probability that each "pixel" is reliable. We adapt the probability calculation to use these estimates as weighting factors for the complementary reliable/unreliable interpretations for each feature vector component. Experiments using the TIDigits connected digit recognition task demonstrate that this technique affords significant performance improvements at low SNRs. 1. INTRODUCTION In previous work [2, 5, 6] we have developed the theory and demonstrated the promise of the Missing Data approach to robust Automatic Speech Recognition. In this technique, spectral-temporal regions uncontaminated by noise are identified and CDHMM recognition methods are ... Jon Barker, Ljubomir Josifovski, Martin Cooke, Phil D. Green |
INTERSPEECH | 1 |
| 1999 | Is the sine-wave speech cocktail party worth attending?
Jon Barker, Martin Cooke |
Speech Commun. | 1 |
| 1998 | Acoustic confidence measures for segmenting broadcast newsabstractIn this paper we define an acoustic confidence measure based on the estimates of local posterior probabilities produced by a HMM/ANN large vocabulary continuous speech recognition system. We use this measure to segment continuous audio into regions where it is and is not appropriate to expend recognition effort. The segmentation is computationally inexpensive and provides reductions in both overall word error rate and decoding time. The technique is evaluated using material from the Broadcast News corpus. 1. Jon Barker, Gethin Williams, Steve Renals |
ICSLP | 1 |
| 1997 | Modelling the recognition of spectrally reduced speechabstractProgress in robust automatic speech recognition may benefit from a fuller account of the mechanisms and representations used by listeners in processing distorted speech. This paper reports on a number of studies which consider how recognisers trained on clean speech can be adapted to cope with a particular form of spectral distortion, namely reduction of clean speech to sine-wave replicas. Using the Resource Management corpus, the first set of recognition experiments confirm the high information content of sine-wave replicas by demonstrating that such tokens can be recognised at levels approaching those for natural speech if matched conditions apply during training. Further recognition tests show that sine-wave speech can be recognised using natural speech models if a spectral peak representation is employed in concert with occluded speech recognition techniques. 1. INTRODUCTION Clean speech and speech with additive noise have been the primary conditions employed in most ASR studies.... Jon Barker, Martin Cooke |
EUROSPEECH | 1 |