Ning Ma 0002

dblp:60/3634-2 · DBLP profile ↗
← Back
47ranked-venue papers
17as first author
12since 2021 · last 2025
0000-0002-4112-3109ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 13 first-author · 9 since 2021Artificial intelligence and machine learning · 29 · 14 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 How Private are Language Models in Abstractive Summarization?
abstract
In sensitive domains such as medical and legal, protecting sensitive information is critical, with protective laws strictly prohibiting the disclosure of personal data.This poses challenges for sharing valuable data such as medical reports and legal cases summaries.While language models (LMs) have shown strong performance in text summarization, it is still an open question to what extent they can provide privacypreserving summaries from non-private source documents.In this paper, we perform a comprehensive study of privacy risks in LM-based summarization across two closed-and four open-weight models of different sizes and families.We experiment with both prompting and fine-tuning strategies for privacy-preservation across a range of summarization datasets including medical and legal domains.Our quantitative and qualitative analysis, including human evaluation, shows that LMs frequently leak personally identifiable information in their summaries, in contrast to human-generated privacy-preserving summaries, which demonstrate significantly higher privacy protection levels.These findings highlight a substantial gap between current LM capabilities and expert human expert performance in privacy-sensitive summarization tasks. 1
Anthony Hughes, Nikolaos Aletras, Ning Ma 0002
EMNLP3
2024 Acoustic Effects of Facial Feminisation Surgery on Speech and Singing: A Case Study
abstract
Transfeminine people may undergo facial feminisation surgery, a term covering a range of procedures that aim to alter the appearance of facial features, thereby potentially changing characteristics of the vocal tract. Effects of facial feminisation surgery on the voice are relatively understudied, however, so, little information on the vocal effects of these surgeries is available to people considering undergoing these procedures. In this single-case study, we present an acoustic analysis of speech and singing data collected from a transgender singer before and after facial feminisation surgery, alongside an examination of longitudinal interview data from the participant. Our quantitative results suggest facial feminisation surgery can have an impact on the voice, and our qualitative analysis suggests this may not only be as a result of the altered characteristics of the vocal tract, but also as a result of the altered social context. Several issues for future research are identified.
Cliodhna Hughes, Guy J. Brown, Ning Ma 0002, Nicola Dibben
INTERSPEECH3
2023 Robust Binaural Sound Localisation with Temporal Attention
abstract
Despite there being clear evidence for attentional effects in biological spatial hearing, relatively few machine hearing systems exploit attention in binaural sound localisation. This paper addresses this issue by proposing a novel binaural machine hearing system with temporal attention for robust localisation of sound sources in noisy and reverberant conditions. A convolutional neural network is employed to extract noise-robust localisation features, which are similar to interaural phase difference, directly from phase spectra of the left and right ears for each frame. A temporal attention layer operates on top of these frame-level features by incorporating outputs of a temporal mask estimation module that indicate target dominance within each frame. The combined features are then exploited by fully connected layers, which map them to the corresponding source azimuth. Both the temporal mask estimation module and the sound localisation module are trained jointly in a multi-task learning manner. Our evaluation shows that the proposed system is able to accurately estimate the azimuth of a sound source in various reverberant and noisy conditions.
Ning Ma 0002, Guy J. Brown
ICASSP2
2023 Obstructive sleep apnea screening with breathing sounds and respiratory effort: a multimodal deep learning approach
abstract
Obstructive sleep apnea (OSA) is a chronic and prevalent condition with well-established comorbidities. Due to limited diagnostic resources and high cost, a significant OSA population lives undiagnosed, and accurate and low-cost methods to screen for OSA are needed. We propose a novel screening method based on breathing sounds recorded with a smartphone and respiratory effort. Whole night recordings are divided into 30-s segments, each of which is classified for the presence or absence of OSA events by a multimodal deep neural network. Data fusion techniques were investigated and evaluated based on the apnea-hypopnea index estimated from whole night recordings. Real-world recordings made during home sleep apnea testing from 103 participants were used to develop and evaluate the proposed system. The late fusion system achieved the best sensitivity and specificity when screening for severe OSA, at 0.93 and 0.92, respectively. This offers the prospect of inexpensive OSA screening at home.
Hector E. Romero, Ning Ma 0002, Guy J. Brown, Sam Johnson
INTERSPEECH2
2022 Auditory-Based Data Augmentation for end-to-end Automatic Speech Recognition
abstract
End-to-end models have achieved significant improvement on automatic speech recognition. One common method to improve performance of these models is expanding the data-space through data augmentation. Meanwhile, human auditory inspired front-ends have also demonstrated improvement for automatic speech recognisers. In this work, a well-verified auditory-based model, which can simulate various hearing abilities, is investigated for the purpose of data augmentation for end-to-end speech recognition. By introducing the auditory model into the data augmentation process, end-to-end systems are encouraged to ignore variation from the signal that cannot be heard and thereby focus on robust features for speech recognition. Two mechanisms in the auditory model, spectral smearing and loudness recruitment, are studied on the LibriSpeech dataset with a transformer-based end-to-end model. The results show that the proposed augmentation methods can bring statistically significant improvement on the performance of the state-of-the-art SpecAugment.
Zehai Tu, Jack Deadman, Ning Ma 0002, Jon Barker
ICASSP3
2022 Exploiting Hidden Representations from a DNN-based Speech Recogniser for Speech Intelligibility Prediction in Hearing-impaired Listeners
abstract
An accurate objective speech intelligibility prediction algorithms is of great interest for many applications such as speech enhancement for hearing aids.Most algorithms measures the signal-to-noise ratios or correlations between the acoustic features of clean reference signals and degraded signals.However, these hand-picked acoustic features are usually not explicitly correlated with recognition.Meanwhile, deep neural network (DNN) based automatic speech recogniser (ASR) is approaching human performance in some speech recognition tasks.This work leverages the hidden representations from DNN-based ASR as features for speech intelligibility prediction in hearingimpaired listeners.The experiments based on a hearing aid intelligibility database show that the proposed method could make better prediction than a widely used short-time objective intelligibility (STOI) based binaural measure.
Zehai Tu, Ning Ma 0002, Jon Barker
INTERSPEECH2
2022 Unsupervised Uncertainty Measures of Automatic Speech Recognition for Non-intrusive Speech Intelligibility Prediction
abstract
Non-intrusive intelligibility prediction is important for its application in realistic scenarios, where a clean reference signal is difficult to access. The construction of many non-intrusive predictors require either ground truth intelligibility labels or clean reference signals for supervised learning. In this work, we leverage an unsupervised uncertainty estimation method for predicting speech intelligibility, which does not require intelligibility labels or reference signals to train the predictor. Our experiments demonstrate that the uncertainty from state-of-the-art end-to-end automatic speech recognition (ASR) models is highly correlated with speech intelligibility. The proposed method is evaluated on two databases and the results show that the unsupervised uncertainty measures of ASR models are more correlated with speech intelligibility from listening results than the predictions made by widely used intrusive methods.
Zehai Tu, Ning Ma 0002, Jon Barker
INTERSPEECH2
2022 SNuC: The Sheffield Numbers Spoken Language Corpus
abstract
We present SNuC, the first published corpus of spoken alphanumeric identifiers of the sort typically used as serial and part numbers in the manufacturing sector. The dataset contains recordings and transcriptions of over 50 native British English speakers, speaking over 13,000 multi-character alphanumeric sequences and totalling almost 20 hours of recorded speech. We describe requirements taken into account in the designing the corpus and the methodology used to construct it. We present summary statistics describing the corpus contents, as well as a preliminary investigation into errors in spoken alphanumeric identifiers. We validate the corpus by showing how it can be used to adapt a deep learning neural network based ASR system, resulting in improved recognition accuracy on the task of spoken alphanumeric identifier recognition. Finally, we discuss further potential uses for the corpus and for the tools developed to construct it.
Emma Barker, Jon Barker, Robert J. Gaizauskas, Ning Ma 0002, Monica Lestari Paramita
LREC4
2022 Acoustic Screening for Obstructive Sleep Apnea in Home Environments Based on Deep Neural Networks
abstract
Obstructive sleep apnea (OSA) is a chronic and prevalent condition with well-established comorbidities. However, many severe cases remain undiagnosed due to poor access to polysomnography (PSG), the gold standard for Obstructive sleep apnea (OSA) diagnosis. Accurate home-based methods to screen for OSA are needed, which can be applied inexpensively to high-risk subjects to identify those that require PSG to fully assess their condition. A number of methods that analyse speech or breathing sounds to screen for OSA have been previously investigated. However, these methods have constraints that limit their use in home environments (e.g., they require specialised equipment, are not robust to background noise, are obtrusive or depend on tightly controlled conditions). This paper proposes a novel method to screen for OSA, which analyses sleep breathing sounds recorded with a smartphone at home. Audio recordings made over a whole night are divided into segments, each of which is classified for the presence or absence of OSA by a deep neural network. The apnea-hypopnea index estimated from the segments predicted as containing evidence of OSA is then used to screen for the condition. Audio recordings made during home sleep apnea testing from 103 participants for 1 or 2 nights were used to develop and evaluate the proposed system. When screening for moderate OSA the acoustics based system achieved a sensitivity of 0.79 and a specificity of 0.80. The sensitivity and specificity when screening for severe OSA were 0.78 and 0.93, respectively. The system is suitable for implementation on consumer smartphones.
Hector E. Romero, Ning Ma 0002, Guy J. Brown, Elizabeth A. Hill
IEEE J. Biomed. Health Informatics2
2021 Exploiting Non-Negative Matrix Factorization for Binaural Sound Localization in the Presence of Directional Interference
abstract
This study presents a novel solution to the problem of binaural localization of a speaker in the presence of interfering directional noise and reverberation. Using a state-of-the-art binaural localization algorithm based on a deep neural network (DNN), we propose adding a source separation stage based on non-negative matrix factorization (NMF) to improve the localization performance in conditions with interfering sources. The separation stage is coupled with the localization stage and is optimized with respect to a broad range of different acoustic conditions, emphasizing a robust and generalizable solution. The machine listening system is shown to greatly benefit from the NMF-based separation stage at low target-to-masker ratios (TMRs) for a variety of noise types, especially for non-stationary noise. It is also demonstrated that training the NMF algorithm on anechoic speech provides better performance than using reverberant speech, and that optimizing the source separation stage using a localization metric rather than a source separation metric substantially increases the system performance.
Ingvi Örnolfsson, Torsten Dau, Ning Ma 0002, Tobias May
ICASSP3
2021 DHASP: Differentiable Hearing Aid Speech Processing
abstract
Hearing aids are expected to improve speech intelligibility for listeners with hearing impairment. An appropriate amplification fitting tuned for the listener’s hearing disability is critical for good performance. The developments of most prescriptive fittings are based on data collected in subjective listening experiments, which are usually expensive and time-consuming. In this paper, we explore an alternative approach to finding the optimal fitting by introducing a hearing aid speech processing framework, in which the fitting is optimised in an automated way using an intelligibility objective function based on the HASPI physiological auditory model. The framework is fully differentiable, thus can employ the back-propagation algorithm for efficient, data-driven optimisation. Our initial objective experiments show promising results for noise-free speech amplification, where the automatically optimised processors outperform one of the well recognised hearing aid prescriptions.
Zehai Tu, Ning Ma 0002, Jon Barker
ICASSP2
2021 Optimising Hearing Aid Fittings for Speech in Noise with a Differentiable Hearing Loss Model
abstract
This is a repository copy of Optimising hearing aid fittings for speech in noise with a differentiable hearing loss model.
Zehai Tu, Ning Ma 0002, Jon Barker
Interspeech2
2020 Snorer Diarisation Based On Deep Neural Network Embeddings
abstract
Acoustic analysis of sleep breathing sounds using a smartphone at home provides a much less obtrusive means of screening for sleep-disordered breathing (SDB) than assessment in a sleep clinic. However, application in a home environment is confounded by the problem that a bed partner may also be present and snore. This paper proposes a novel acoustic analysis system for snorer diarisation, a concept extrapolated from speaker diarisation research, which allows screening for SDB of both the user and the bed partner using a single smartphone. The snorer diarisation system involves three steps. First, a deep neural network (DNN) is employed to estimate the number of concurrent snorers in short segments of monaural audio recordings. Second, the identified snore segments are clustered using snorer embeddings, a feature representation that allows different snorers to be discriminated. Finally, a snore transcription is automatically generated for each snorer by combining consecutive snore segments. The system is evaluated on both synthetic snore mixtures and real two-snorer recordings. The results show that it is possible to accurately screen a subject and their bed partner for SDB in the same session from recordings of a single smartphone.
Hector E. Romero, Ning Ma 0002, Guy J. Brown
ICASSP2
2019 Deep Learning Features for Robust Detection of Acoustic Events in Sleep-disordered Breathing
abstract
Sleep-disordered breathing (SDB) is a serious and prevalent condition, and acoustic analysis via consumer devices (e.g. smartphones) offers a low-cost solution to screening for it. We present a novel approach for the acoustic identification of SDB sounds, such as snoring, using bottleneck features learned from a corpus of whole-night sound recordings. Two types of bottleneck features are described, obtained by applying a deep autoencoder to the output of an auditory model or a short-term autocorrelation analysis. We investigate two architectures for snore sound detection: a tandem system and a hybrid system. In both cases, a `language model' (LM) was incorporated to exploit information about the sequence of different SDB events. Our results show that the proposed bottleneck features give better performance than conventional mel-frequency cepstral coefficients, and that the tandem system outperforms the hybrid system given the limited amount of labelled training data available. The LM made a small improvement to the performance of both classifiers.
Hector E. Romero, Ning Ma 0002, Guy J. Brown, Amy V. Beeston, Madina Hasan
ICASSP2
2019 End-to-end Binaural Sound Localisation from the Raw Waveform
abstract
A novel end-to-end binaural sound localisation approach is proposed which estimates the azimuth of a sound source directly from the waveform. Instead of employing hand-crafted features commonly employed for binaural sound localisation, such as the interaural time and level difference, our end-to-end system approach uses a convolutional neural network (CNN) to extract specific features from the waveform that are suitable for localisation. Two systems are proposed which differ in the initial frequency analysis stage. The first system is auditory-inspired and makes use of a gammatone filtering layer, while the second system is fully data-driven and exploits a trainable convolutional layer to perform frequency analysis. In both systems, a set of dedicated convolutional kernels are then employed to search for specific localisation cues, which are coupled with a localisation stage using fully connected layers. Localisation experiments using binaural simulation in both anechoic and reverberant environments show that the proposed systems outperform a state-of-the-art deep neural network system. Furthermore, our investigation of the frequency analysis stage in the second system suggests that the CNN is able to exploit different frequency bands for localisation according to the characteristics of the reverberant environment.
Paolo Vecchiotti, Ning Ma 0002, Stefano Squartini, Guy J. Brown
ICASSP2
2018 Robust Binaural Localization of a Target Sound Source by Combining Spectral Source Models and Deep Neural Networks
abstract
Despite there being a clear evidence for top-down (e.g., attentional) effects in biological spatial hearing, relatively few machine hearing systems exploit the top-down model-based knowledge in sound localization. This paper addresses this issue by proposing a novel framework for the binaural sound localization that combines the model-based information about the spectral characteristics of sound sources and deep neural networks (DNNs). A target source model and a background source model are first estimated during a training phase using spectral features extracted from sound signals in isolation. When the identity of the background source is not available, a universal background model can be used. During testing, the source models are used jointly to explain the mixed observations and improve the localization process by selectively weighting source azimuth posteriors output by a DNN-based localization system. To address the possible mismatch between the training and testing, a model adaptation process is further employed the on-the-fly during testing, which adapts the background model parameters directly from the noisy observations in an iterative manner. The proposed system, therefore, combines the model-based and data-driven information flow within a single computational framework. The evaluation task involved localization of a target speech source in the presence of an interfering source and room reverberation. Our experiments show that by exploiting the model-based information in this way, the sound localization performance can be improved substantially under various noisy and reverberant conditions.
Ning Ma 0002, José A. González 0001, Guy J. Brown
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 Improving audio-visual speech recognition using deep neural networks with dynamic stream reliability estimates
abstract
Audio-visual speech recognition is a promising approach to tackling the problem of reduced recognition rates under adverse acoustic conditions. However, finding an optimal mechanism for combining multi-modal information remains a challenging task. Various methods are applicable for integrating acoustic and visual information in Gaussian-mixture-model-based speech recognition, e.g., via dynamic stream weighting. The recent advances of deep neural network (DNN)-based speech recognition promise improved performance when using audio-visual information. However, the question of how to optimally integrate acoustic and visual information remains. In this paper, we propose a state-based integration scheme that uses dynamic stream weights in DNN-based audio-visual speech recognition. The dynamic weights are obtained from a time-variant reliability estimate that is derived from the audio signal. We show that this state-based integration is superior to early integration of multi-modal features, even if early integration also includes the proposed reliability estimate. Furthermore, the proposed adaptive mechanism is able to outperform a fixed weighting approach that exploits oracle knowledge of the true signal-to-noise ratio.
Hendrik Meutzner, Ning Ma 0002, Robert M. Nickel, Christopher Schymura, Dorothea Kolossa
ICASSP2
2017 Exploiting Deep Neural Networks and Head Movements for Robust Binaural Localization of Multiple Sources in Reverberant Environments
abstract
This paper presents a novel machine-hearing system that exploits deep neural networks (DNNs) and head movements for robust binaural localization of multiple sources in reverberant environments. DNNs are used to learn the relationship between the source azimuth and binaural cues, consisting of the complete cross-correlation function (CCF) and interaural level differences (ILDs). In contrast to many previous binaural hearing systems, the proposed approach is not restricted to localization of sound sources in the frontal hemifield. Due to the similarity of binaural cues in the frontal and rear hemifields, front-back confusions often occur. To address this, a head movement strategy is incorporated in the localization model to help reduce the front-back errors. The proposed DNN system is compared to a Gaussian-mixture-model-based system that employs interaural time differences (ITDs) and ILDs as localization features. Our experiments show that the DNN is able to exploit information in the CCF that is not available in the ITD cue, which together with head movements substantially improves localization accuracies under challenging acoustic scenarios, in which multiple talkers and room reverberation are present.
Ning Ma 0002, Tobias May, Guy J. Brown
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 Robust audiovisual speech recognition using noise-adaptive linear discriminant analysis
abstract
Automatic speech recognition (ASR) has become a widespread and convenient mode of human-machine interaction, but it is still not sufficiently reliable when used under highly noisy or reverberant conditions. One option for achieving far greater robustness is to include another modality that is unaffected by acoustic noise, such as video information. Currently the most successful approaches for such audiovisual ASR systems, coupled hidden Markov models (HMMs) and turbo decoding, both allow for slight asynchrony between audio and video features, and significantly improve recognition rates in this way. However, both typically still neglect residual errors in the estimation of audio features, so-called observation uncertainties. This paper compares two strategies for adding these observation uncertainties into the decoder, and shows that significant recognition rate improvements are achievable for both coupled HMMs and turbo decoding.
Steffen Zeiler, Robert Nicheli, Ning Ma 0002, Guy J. Brown, Dorothea Kolossa
ICASSP3
2016 A Robust Dual-Microphone Speech Source Localization Algorithm for Reverberant Environments
abstract
Speech source localization (SSL) using a microphone array \naims to estimate the direction-of-arrival (DOA) of the speech \nsource. However, its performance often degrades rapidly in reverberant \nenvironments. In this paper, a novel dual-microphone \nSSL algorithm is proposed to address this problem. First, the \ntime-frequency regions dominated by direct sound are extracted \nby tracking the envelopes of speech, reverberation and background \nnoise. The time-difference-of-arrival (TDOA) is then \nestimated by considering only these reliable regions. Second, \na bin-wise de-aliasing strategy is introduced to make better use \nof the DOA information carried at high frequencies, where the \nspatial resolution is higher and there is typically less corruption \nby diffuse noise. Our experiments show that when compared \nwith other widely-used algorithms, the proposed algorithm produces \nmore reliable performance in realistic reverberant environments.
Yanmeng Guo, Xiaofei Wang 0007, Chao Wu 0011, Qiang Fu 0001, Ning Ma 0002, Guy J. Brown
INTERSPEECH5
2016 Speech Localisation in a Multitalker Mixture by Humans and Machines
abstract
Speech localisation in multitalker mixtures is affected by the listener’s expectations about the spatial arrangement of the sound sources. This effect was investigated via experiments with human listeners and a machine system, in which the task was to localise a female-voice target among four spatially distributed male-voice maskers. Two configurations were used: either the masker locations were fixed or the locations varied from trial-to-trial. The machine system uses deep neural networks (DNNs) to learn the relationship between binaural cues and source azimuth, and exploits top-down knowledge about the spectral characteristics of the target source. Performance was examined in both anechoic and reverberant conditions. Our experiments show that the machine system outperformed listeners in some conditions. Both the machine and listeners were able to make use of a priori knowledge about the spatial configuration of the sources, but the effect for headphone listening was smaller than that previously reported for listening in a real room.
Ning Ma 0002, Guy J. Brown
INTERSPEECH1
2015 Exploiting synchrony spectra and deep neural networks for noise-robust automatic speech recognition
abstract
This paper presents a novel system that exploits synchrony spectra and deep neural networks (DNNs) for automatic speech recognition (ASR) in challenging noisy environments. Synchrony spectra measure the extent to which each frequency channel in an auditory model is entrained to a particular pitch period, and they are used together with F0 estimates either in a DNN for time-frequency (T-M) mask estimation or to augment the input features for a DNN-based ASR system. The proposed approach was evaluated in the context of the CHiME 3 Challenge. Our experiments show that the synchrony spectra features work best when augmenting the input features to the DNN-based ASR system. Compared to the CHiME-3 baseline system, our best system provides a word error rate (WER) reduction of more than 14% absolute and achieved a WER of 18.56% on the evaluation test set.
Ning Ma 0002, Ricard Marxer, Jon Barker, Guy J. Brown
ASRU1
2015 A machine-hearing system exploiting head movements for binaural sound localisation in reverberant conditions
abstract
This paper is concerned with machine localisation of multiple active speech sources in reverberant environments using two (binaural) microphones. Such conditions typically present a problem for `classical' binaural models. Inspired by the human ability to utilise head movements, the current study investigated the influence of different head movement strategies on binaural sound localisation. A machine-hearing system that exploits a multi-step head rotation strategy for sound localisation was found to produce the best performance in simulated reverberant acoustic space. This paper also reports the public release of a free binaural room impulse responses (BRIRs) database that allows the simulation of head rotation used in this study.
Ning Ma 0002, Tobias May, Hagen Wierstorf, Guy J. Brown
ICASSP1
2015 Robust localisation of multiple speakers exploiting head movements and multi-conditional training of binaural cues
abstract
This paper addresses the problem of localising multiple competing speakers in the presence of room reverberation, where sound sources can be positioned at any azimuth on the horizontal plane. To reduce the amount of front-back confusions which can occur due to the similarity of interaural time differences (ITDs) and interaural level differences (ILDs) in the front and rear hemifield, a machine hearing system is presented which combines supervised learning of binaural cues using multi-conditional training (MCT) with a head movement strategy. A systematic evaluation showed that this approach substantially reduced the amount of front-back confusions in challenging acoustic scenarios. Moreover, the system was able to generalise to a variety of different acoustic conditions not seen during training.
Tobias May, Ning Ma 0002, Guy J. Brown
ICASSP2
2015 Exploiting top-down source models to improve binaural localisation of multiple sources in reverberant environments
abstract
Relatively few systems for machine hearing exploit top-down \ninformation in source localisation, despite there being clear \nevidence for top-down (e.g., attentional) effects in biological \nspatial hearing. This paper addresses this issue by proposing \na framework for binaural sound localisation that exploits top- \ndown knowledge about the source spectral characteristics in the \nacoustic scene. Information from source models is used to im- \nprove the localisation process by selectively weighting binaural \ncues. The system therefore combines top-down and bottom- \nup information flow within a single computational framework. \nOur experiments show that by exploiting source models in this \nway, sound localisation performance can be improved substan- \ntially under challenging conditions in which multiple sources \nand room reverberation are present.
Ning Ma 0002, Guy J. Brown, José A. González 0001
INTERSPEECH1
2015 Exploiting deep neural networks and head movements for binaural localisation of multiple speakers in reverberant conditions
abstract
This paper presents a novel machine-hearing system that exploits deep neural networks (DNNs) and head movements for binaural localisation of multiple speakers in reverberant conditions.DNNs are used to map binaural features, consisting of the complete crosscorrelation function (CCF) and interaural level differences (ILDs), to the source azimuth.Our approach was evaluated using a localisation task in which sources were located in a full 360-degree azimuth range.As a result, front-back confusions often occurred due to the similarity of binaural features in the front and rear hemifields.To address this, a head movement strategy was incorporated in the DNN-based model to help reduce the front-back errors.Our experiments show that, compared to a system based on a Gaussian mixture model (GMM) classifier, the proposed DNN system substantially reduces localisation errors under challenging acoustic scenarios in which multiple speakers and room reverberation are present.
Ning Ma 0002, Guy J. Brown, Tobias May
INTERSPEECH1
2013 The PASCAL CHiME speech separation and recognition challenge
Jon Barker, Emmanuel Vincent 0001, Ning Ma 0002, Heidi Christensen, Phil D. Green
Comput. Speech Lang.3
2013 A hearing-inspired approach for distant-microphone speech recognition in the presence of multiple sources
Ning Ma 0002, Jon Barker, Heidi Christensen, Phil D. Green
Comput. Speech Lang.1
2013 Speech Spectral Envelope Enhancement by HMM-Based Analysis/Resynthesis
abstract
We propose a speech enhancement-by-resynthesis framework whose strength lies in a common statistical speech model that is shared by the analysis and synthesis stages. First, a spectro-temporal analysis is performed and masked spectro-temporal regions are identified using a noise model. Then, HMM synthesis is used to reconstruct the spectral envelope in masked regions in a manner which is conditioned on the reliable regions, preventing the resynthesis from regressing to the training data mean. As a demonstration we enhance noise-corrupted speech utterances from a small vocabulary corpus for which good statistical models are available. Perceptual evaluation of speech quality and log spectral distances demonstrate considerable performance improvements over baseline approaches that do not exploit strong speech knowledge. The letter is accompanied by audio examples.
José L. Carmona, Jon Barker, Ángel M. Gómez, Ning Ma 0002
IEEE Signal Process. Lett.4
2013 MMSE-Based Missing-Feature Reconstruction With Temporal Modeling for Robust Speech Recognition
abstract
This paper addresses the problem of feature compensation in the log-spectral domain by using the missing-data (MD) approach to noise robust speech recognition, that is, the log-spectral features can be either almost unaffected by noise or completely masked by it. First, a general MD framework based on minimum mean square error (MMSE) estimation is introduced which exploits the correlation across frequency bands to reconstruct the missing features. This framework allows the derivation of different MD imputation approaches and, in particular, a novel technique taking advantage of truncated Gaussian distributions is presented. While the proposed technique provides excellent results at high and medium signal-to-noise ratios (SNRs), its performance diminishes at low SNRs where very few reliable features are available. The reconstruction technique is therefore extended to exploit temporal constraints using two different approaches. In the first approach, time-frequency patches of speech containing a number of consecutive frames are modeled using a Gaussian mixture model (GMM). In the second one, the sequential structure of speech is alternatively modeled by a hidden Markov model (HMM). The proposed techniques are evaluated on Aurora-2 and Aurora-4 databases using both oracle and estimated masks. In both cases, the proposed techniques outperform the recognition performance obtained by the baseline system and other related techniques. Also, the introduction of a temporal modeling turns out to be very effective in reconstructing spectra at low SNRs. In particular, HMMs show the highest capability of accounting for time correlations and, therefore, achieve the best results.
José A. González 0001, Antonio M. Peinado, Ning Ma 0002, Ángel M. Gómez, Jon Barker
IEEE Trans. Speech Audio Process.3
2012 Combining missing-data reconstruction and uncertainty decoding for robust speech recognition
abstract
This paper proposes a novel approach for noise-robust speech recognition which combines a missing-data (MD) derived spectral reconstruction technique and uncertainty decoding based on the weighted Viterbi algorithm (WVA). First, the noisy feature vectors are compensated by using a novel MD imputation technique based on the integration of truncated Gaussian pdfs. Although the proposed MD estimator has both the advantages of MD techniques and the use of cepstral features, it may still be affected by a number of uncertainty sources. In order to deal with these uncertainties, WVA-based uncertainty decoding is proposed. Our experiments on the Aurora-2 and Aurora-4 tasks show that the proposed MD estimator outperforms other MD imputation techniques. Also, we show that the combination of MD imputation with WVA provides better results than the combination with other uncertainty processing techniques such as the use of evidence pdfs for the estimated features.
José A. González 0001, Antonio M. Peinado, Ángel M. Gómez, Ning Ma 0002, Jon Barker
ICASSP4
2012 Log-spectral feature reconstruction based on an occlusion model for noise robust speech recognition
José A. González 0001, Antonio M. Peinado, Ángel M. Gómez, Ning Ma 0002
INTERSPEECH4
2012 Coupling identification and reconstruction of missing features for noise-robust automatic speech recognition
abstract
The standard missing feature imputation approach to noiserobust automatic speech recognition requires that a single foreground/background segmentation mask is identified prior to reconstruction. This paper presents a novel imputation approach which more closely couples the identification and reconstruction of missing features by using a probabilistic framework based on the speech fragment decoding technique. Using fragment decoding, the most joint-likely state sequence and segmentation hypothesis is identified with which the missing data region is imputed. Crucially, however, imputation can exploit the speech state sequence recovered by the fragment decoding. Further, using N -best decodings allows the clean spectrogram to be estimated as a weighted combination of reconstructions which provides some allowance for uncertainty in the estimates. Experiments on the PASCAL CHiME Challenge task show that system performance is highly dependent on the complexity of the speech models used for segmentation and imputation, and by exploiting the temporal constraint of speech the system significantly outperforms those that ignore the constraint.
Ning Ma 0002, Jon Barker
INTERSPEECH1
2012 Combining Speech Fragment Decoding and Adaptive Noise Floor Modeling
abstract
This paper presents a novel noise-robust automatic speech recognition (ASR) system that combines aspects of the noise modeling and source separation approaches to the problem. The combined approach has been motivated by the observation that the noise backgrounds encountered in everyday listening situations can be roughly characterized as a slowly varying noise floor in which there are embedded a mixture of energetic but unpredictable acoustic events. Our solution combines two complementary techniques. First, an adaptive noise floor model estimates the degree to which high-energy acoustic events are masked by the noise floor (represented by a soft missing data mask). Second, a fragment decoding system attempts to interpret the high-energy regions that are not accounted for by the noise floor model. This component uses models of the target speech to decide whether fragments should be included in the target speech stream or not. Our experiments on the CHiME corpus task show that the combined approach performs significantly better than systems using either the noise model or fragment decoding approach alone, and substantially outperforms multicondition training.
Ning Ma 0002, Jon Barker, Heidi Christensen, Phil D. Green
IEEE Trans. Speech Audio Process.1
2011 A pitch based noise estimation technique for robust speech recognition with Missing Data
abstract
This paper presents a noise estimation technique based on knowledge of pitch information for robust speech recognition. In the first stage the noise is estimated by means of extrapolating the noise from frames where speech is believed to be absent. These frames are detected with a proposed pitch based VAD (Voice Activity Detector). In the second stage the noise estimation is revised in voiced frames using harmonic tunnelling technique. The tunnelling noise estimation is used at high SNRs as an upper bound of the noise rather than a suitable estimation. A spectrogram MD (Missing Data) recognition system is chosen to evaluate the proposed noise estimation. The proposed system is compared in Aurora-2 with other similar techniques like cepstral SS (Spectral Subtraction).
Juan Andres Morales-Cordovilla, Ning Ma 0002, Victoria E. Sánchez, José L. Carmona, Antonio M. Peinado, Jon Barker
ICASSP2
2011 Binaural Cues for Fragment-Based Speech Recognition in Reverberant Multisource Environments
abstract
This paper addresses the problem of speech recognition using distant binaural microphones in reverberant multisource noise conditions. Our scheme employs a two stage fragment decoding approach: first spectro-temporal acoustic source fragments are identified using signal level cues, and second, a hypothesisdriven stage simultaneously searches for the most probable speech/background fragment labelling and the corresponding acoustic model state sequence. The paper reports the first successful attempt to use binaural localisation cues within this framework. By integrating binaural cues and acoustic models in a consistent probabilistic framework, the decoder is able to derive significant recognition performance benefits from fragment location estimates despite their inherent unreliability.
Ning Ma 0002, Jon Barker, Heidi Christensen, Phil D. Green
INTERSPEECH1
2010 The CHiME corpus: a resource and a challenge for computational hearing in multisource environments
abstract
We present a new corpus designed for noise-robust speech processing research, CHiME. Our goal was to produce material which is both natural (derived from reverberant domestic environments with many simultaneous and unpredictable sound sources) and controlled (providing an enumerated range of SNRs spanning 20 dB). The corpus includes around 40 hours of background recordings from a head and torso simulator positioned in a domestic setting, and a comprehensive set of binaural impulse responses collected in the same environment. These have been used to add target utterances from the Grid speech recognition corpus into the CHiME domestic setting. Data has been mixed in a manner that produces a controlled and yet natural range of SNRs over which speech separation, enhancement and recognition algorithms can be evaluated. The paper motivates the design of the corpus, and describes the collection and post-processing of the data. We also present a set of baseline recognition results.
Heidi Christensen, Jon Barker, Ning Ma 0002, Phil D. Green
INTERSPEECH3
2010 Speech fragment decoding techniques for simultaneous speaker identification and speech recognition
Jon Barker, Ning Ma 0002, André Coy, Martin Cooke
Comput. Speech Lang.2
2009 A speech fragment approach to localising multiple speakers in reverberant environments
abstract
Sound source localisation cues are severely degraded when multiple acoustic sources are active in the presence of reverberation. We present a binaural system for localising simultaneous speakers which exploits the fact that in a speech mixture there exist spectro-temporal regions or dasiafragmentspsila, where the energy is dominated by just one of the speakers. A fragment-level localisation model is proposed that integrates the localisation cues within a fragment using a weighted mean. The weights are based on local estimates of the degree of reverberation in a given spectro-temporal cell. The paper investigates different weight estimation approaches based variously on, i) an established model of the perceptual precedence effect; ii) a measure of interaural coherence between the left and right ear signals; iii) a data-driven approach trained in matched acoustic conditions. Experiments with reverberant binaural data with two simultaneous speakers show appropriate weighting can improve frame-based localisation performance by up to 24%.
Heidi Christensen, Ning Ma 0002, Stuart N. Wrigley, Jon Barker
ICASSP2
2009 Modelling the prepausal lengthening effect for speech recognition: a dynamic Bayesian network approach
abstract
Speech has a property that the speech unit preceding a speech pause tends to lengthen. This work presents the use of a dynamic Bayesian network to model the prepausal lengthening effect for robust speech recognition. Specifically, we introduce two distributions to model inter-state transitions in prepausal and non-prepausal words, respectively. The selection of the transition distributions depends on a random variable whose value is influenced by whether a pause will appear between the current and the following word. Two experiments are presented here. The first one considers pauses hypothesised during speech decoding. The second one employs an extra component for speech/non-speech determination. By modelling the prepausal lengthening effect we achieve a 5.5% relative reduction in word error rate on the 500-word task of the SVitchboard corpus.
Ning Ma 0002, Chris D. Bartels, Jeff A. Bilmes, Phil D. Green
ICASSP1
2008 A 'speechiness' measure to improve speech decoding in the presence of other sound sources
abstract
When speech is corrupted by other sound sources certain spectro-temporal regions will be dominated by speech energy and others by the noise. Listeners are able to exploit these cues to achieve robust speech perception in adverse conditions. Inspired by this perception process a ‘speech fragment decoding’ technique has shown promising robustness when handling multiple sound sources. This paper proposes an approach to estimating ‘speechiness’ – a degree of confidence that a spectrotemporal region is dominated by speech energy – using the modulation spectrogram. This additional knowledge is employed to steer the decoder towards selecting more reliable speech evidence in noise. Experiments show that the speechiness measure is capable of improving recognition accuracies in various noise conditions at 0 dB global signal-to-noise ratio.
Ning Ma 0002, Phil D. Green
INTERSPEECH1
2007 Integrating pitch and localisation cues at a speech fragment level
abstract
This paper proposes a novel speech-fragment based approach for processing binaural data to improve the estimation of speech source locations in reverberant, multi-speaker recordings. The technique employs two stages. First, a robust multipitch tracking algorithm is used to locate local spectro-temporal ‘speech fragments ’ – regions where the energy in the mixture is dominated by a single speech source. Second, robust localisation estimates are formed by integrating interaural time difference cues over each speech fragment. The technique is applied to the analysis of more than five hours of two-party meetings that have been constructed from a mixture of binaural mannequin recordings. It is shown that estimating location at the speech fragment level produces better results than conventional location-estimate smoothing techniques leading to a an increase in relative frame accuracy rate of more than 35%. Index Terms: binaural localisation, pitch cues, speech fragment integration
Heidi Christensen, Ning Ma 0002, Stuart N. Wrigley, Jon Barker
INTERSPEECH2
2007 Applying word duration constraints by using unrolled HMMs
abstract
Conventional HMMs have weak duration constraints. In noisy conditions, the mismatch between corrupted speech signals and models trained on clean speech may cause the decoder to produce word matches with unrealistic durations. This paper presents a simple way to incorporate word duration constraints by unrolling HMMs to form a lattice where word duration probabilities can be applied directly to state transitions. The expanded HMMs are compatible with conventional Viterbi decoding. Experiments on connected-digit recognition show that when using explicit duration constraints the decoder generates word matches with more reasonable durations, and word error rates are significantly reduced across a broad range of noise conditions.
Ning Ma 0002, Jon Barker, Phil D. Green
INTERSPEECH1
2007 Exploiting correlogram structure for robust speech recognition with multiple speech sources
Ning Ma 0002, Phil D. Green, Jon Barker, André Coy
Speech Commun.1
2006 Recent advances in speech fragment decoding techniques
abstract
This paper addresses the problem of recognising speech in the presence of a competing speaker. We employ a speech fragment decoding technique that treats segregation and recognition as coupled problems. Data-driven techniques are used to segment a spectro-temporal representation into a set of spectro-temporal fragments, such that each fragment is dominated by one or other of the speech sources. A speech fragment decoder is used which employs missing data techniques and clean speech models to simultaneously search for the set of fragments and the word sequence that best matches the target speaker model. The paper reports recent advances in this technique, and presents an evaluation based on artificially mixed speech utterances. The fragment decoder produces significantly lower error rates than a conventional recogniser, and mimics the pattern of human performance whereby performance increases as the target-masker ratio is reduced below -3 dB. Index Terms: speech recognition, speech separation, simultaneous speech, auditory scene analysis, noise robustness.
Jon Barker, André Coy, Ning Ma 0002, Martin Cooke
INTERSPEECH3
2006 Exploiting dendritic autocorrelogram structure to identify spectro-temporal regions dominated by a single sound source
abstract
Autocorrelograms exhibit tree-like structures whose spines are located at a delay of 1/F0. This paper exploits the dendritic autocorrelogram structure for the identification of spectro-temporal regions dominated by a single periodic sound source in monaural acoustic mixtures. Each frame of the mixture is first segmented into different sound sources in the autocorrelogram domain. Local pitch estimates are formed for each source and used as a cue for temporal integration. A confidence score is computed for each time-frequency pixel in the grouped regions to determine its probability of belonging to the group. The system is evaluated using simultaneous speech in a coherence measuring experiment and also employed within an ASR system where it produces improved results for the Interspeech 2006 Speech Separation Challenge. Index Terms: speech separation, correlogram, multipitch tracking.
Ning Ma 0002, Phil D. Green, André Coy
INTERSPEECH1
2005 Context-dependent word duration modelling for robust speech recognition
abstract
Conventional hidden Markov models (HMMs) have weak duration constraints. This may cause the decoder to produce word matches with unrealistic durations in noisy situations. This paper describes techniques for modelling context-dependent word duration cues and incorporating them directly in a multi-stack decoding algorithm. The proposed model is capable of penalising duration constraints of a word depending on its context. Experiments on connected digit recognition show that the new system can significantly improve recognition performance at different noise levels. 1.
Ning Ma 0002, Phil D. Green
INTERSPEECH1