EDBT 2026 Demo / reviewers in the wild / expert
Patrick A. Naylor
dblp:24/5472
· DBLP profile ↗
133ranked-venue papers
5as first author
21since 2021 · last 2025
0000-0001-8546-8013ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 93 · 3 first-author · 16 since 2021Artificial intelligence and machine learning · 47 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Physiologically-Informed Feature Analysis of Acquired Speech Disorders for Stroke Assessment
Giulia Sanguedolce, Jón Guðnason, Dragos-Cristian Gruia, Emilie D'Olne, Fatemeh Geranmayeh, Patrick A. Naylor |
INTERSPEECH | 6 |
| 2024 | Speech Enhancement in Hearing Aids Using Target Speech Presence Estimation Based on a Delayed Remote Microphone SignalabstractSpeech enhancement in hearing aids (HAs) can take advantage of a wireless remote microphone (RM) having a better signal-to-noise ratio than the HA microphones. However, using the RM effectively is complicated by the time delay between the acoustic and wireless signals. Methods in the literature assume an instantaneous transmission of the RM signal, which is never the case in practice. Hence, we propose a practically operational method to use an RM with HAs in the presence of wireless transmission delays. Specifically, we use the delayed target voice activity state from the RM signal, to derive an expression for target speech presence probability (SPP) at the local HA microphone signals. The proposed method uses this target SPP mask as a post-filter that follows a local multichannel Wiener filter. Through simulations, we demonstrate that the proposed method improves speech quality and intelligibility metrics, especially in very noisy acoustic environments, compared to a standard approach which relies solely on HA microphones. Vasudha Sathyapriyan, Michael Syskind Pedersen, Mike Brookes, Jan Østergaard, Patrick A. Naylor, Jesper Jensen 0001 |
ICASSP | 5 |
| 2024 | Binaural Speech Enhancement Using Deep Complex Convolutional Transformer NetworksabstractStudies have shown that in noisy acoustic environments, providing binaural signals to the user of an assistive listening device may improve speech intelligibility and spatial awareness. This paper presents a binaural speech enhancement method using a complex convolutional neural network with an encoder-decoder architecture and a complex multi-head attention transformer. The model is trained to estimate individual complex ratio masks in the time-frequency domain for the left and right-ear channels of binaural hearing devices. The model is trained using a novel loss function that incorporates the preservation of spatial information along with speech intelligibility improvement and noise reduction. Simulation results for acoustic scenarios with a single target speaker and isotropic noise of various types show that the proposed method improves the estimated binaural speech intelligibility and preserves the binaural cues better in comparison with several baseline algorithms. Vikas Tokala, Eric Grinstein, Mike Brookes, Simon Doclo, Jesper Jensen 0001, Patrick A. Naylor |
ICASSP | 6 |
| 2024 | XANE: eXplainable Acoustic Neural Embeddings
Sri Harsha Dumpala, Dushyant Sharma, Chandramouli Shama Sastry, Stanislav Yu. Kruchinin, James Fosburgh, Patrick A. Naylor |
INTERSPEECH | 6 |
| 2024 | When Whisper Listens to Aphasia: Advancing Robust Post-Stroke Speech RecognitionabstractDespite recent advancements in Automatic Speech Recogni tion (ASR), its accuracy remains low for pathological speech, thereby limiting AI-based healthcare interventions in such set tings. This work addresses this challenge by fine-tuning Whis per, an ASR known for its ability to capture high-dimensional features in healthy speech. Using our comprehensive dataset of patients with stroke, we fine-tuned Whisper and significantly reduced Word Error Rate (WER), surpassing previous work on severe aphasia. To demonstrate its generalisability, we tested the model on a separate database, AphasiaBank, and observed a lower WER despite variations in dialect, linguistics, and test protocols. Our result on the AphasiaBank was superior to pre vious ASRs trained on this database, confirming the generalis ability of our approach. These outcomes not only address ASR limitations in impaired speech but also establish the foundations for standardised and versatile AI solutions for remote speech monitoring for timely diagnosis and intervention. Index Terms: Speech Recognition, Fine-tuning, Pathological Speech Giulia Sanguedolce, Sophie Brook, Dragos-Cristian Gruia, Patrick A. Naylor, Fatemeh Geranmayeh |
INTERSPEECH | 4 |
| 2023 | Graph Neural Networks for Sound Source Localization on Distributed Microphone NetworksabstractDistributed Microphone Arrays (DMAs) present many challenges with respect to centralized microphone arrays. An important requirement of applications on these arrays is handling a variable number of input channels. We consider the use of Graph Neural Networks (GNNs) as a solution to this challenge. We present a localization method using the Relation Network GNN, which we show shares many similarities to classical signal processing algorithms for Sound Source Localization (SSL). We apply our method for the task of SSL and validate it experimentally using an unseen number of microphones. We test different feature extractors and show that our approach significantly outperforms classical baselines. Eric Grinstein, Mike Brookes, Patrick A. Naylor |
ICASSP | 3 |
| 2023 | The MBSTOI Binaural Intelligibility Metric Using a Close-Talking Microphone ReferenceabstractIntelligibility metrics are a fast way to determine how comprehensible a target signal is in a noisy situation. Most metrics however rely on having a clean reference signal for computation and are not adapted to live recordings. In this paper the deep correlation modified binaural short time objective intelligibility metric (Dcor-MBSTOI) is evaluated with a single-channel close-talking microphone signal as the reference. This reference signal inevitably contains some background noise and crosstalk from non-target sources. It is found that intelligibility is overestimated when using the close-talking microphone signal directly but that this overestimation can be eliminated by applying speech enhancement to the reference signal. Pierre Guiraud, Alastair H. Moore, Rebecca R. Vos, Patrick A. Naylor, Mike Brookes |
ICASSP | 4 |
| 2023 | Subspace Hybrid Beamforming for Head-Worn Microphone ArraysabstractA two-stage multi-channel speech enhancement method is proposed which consists of a novel adaptive beamformer, Hybrid Minimum Variance Distortionless Response (MVDR), Isotropic-MVDR (Iso), and a novel multi-channel spectral Principal Components Analysis (PCA) denoising. In the first stage, the Hybrid-MVDR performs multiple MVDRs using a dictionary of pre-defined noise field models and picks the minimum-power outcome, which benefits from the robustness of signal-independent beamforming and the performance of adaptive beamforming. In the second stage, the outcomes of Hybrid and Iso are jointly used in a two-channel PCA-based denoising to remove the ‘musical noise’ produced by Hybrid beamformer. On a dataset of real ‘cocktail-party’ recordings with head-worn array, the proposed method outperforms the baseline superdirective beamformer in noise suppression (fwSegSNR, SDR, SIR, SAR) and speech intelligibility (STOI) with similar speech quality (PESQ) improvement. Sina Hafezi, Alastair H. Moore, Pierre Guiraud, Patrick A. Naylor, Jacob Donley, Vladimir Tourbabin, Thomas Lunner |
ICASSP | 4 |
| 2023 | Two-Stage Voice Anonymization for Enhanced Privacy
Francesco Nespoli, Daniel Barreda, Jörg Bitzer, Patrick A. Naylor |
INTERSPEECH | 4 |
| 2023 | Using a single-channel reference with the MBSTOI binaural intelligibility metricabstractIn order to assess the intelligibility of a target signal in a noisy environment, intrusive speech intelligibility metrics are typically used. They require a clean reference signal to be available which can be difficult to obtain especially for binaural metrics like the modified binaural short time objective intelligibility metric (MBSTOI). We here present a hybrid version of MBSTOI that incorporates a deep learning stage that allows the metric to be computed with only a single-channel clean reference signal. The models presented are trained on simulated data containing target speech, localised noise, diffuse noise, and reverberation. The hybrid output metrics are then compared directly to MBSTOI to assess performances. Results show the performance of our single channel reference vs MBSTOI. The outcome of this work offers a fast and flexible way to generate audio data for machine learning (ML) and highlights the potential for low level implementation of ML into existing tools. Pierre Guiraud, Alastair H. Moore, Rebecca R. Vos, Patrick A. Naylor, Mike Brookes |
Speech Commun. | 4 |
| 2023 | Signal Compaction Using Polynomial EVD for Spherical Array Processing With ApplicationsabstractMulti-channel signals captured by spatially separated sensors often contain a high level of data redundancy. A compact signal representation enables more efficient storage and processing, which has been exploited for data compression, noise reduction, and speech and image coding. This paper focuses on the compact representation of speech signals acquired by spherical microphone arrays. A polynomial matrix eigenvalue decomposition (PEVD) can spatially decorrelate signals over a range of time lags and is known to achieve optimum multi-channel data compaction. However, the complexity of PEVD algorithms scales at best cubically with the number of channel signals, e.g., the number of microphones comprised in a spherical array used for processing. In contrast, the spherical harmonic transform (SHT) provides a compact spatial representation of the 3-dimensional sound field measured by spherical microphone arrays, referred to as eigenbeam signals, at a cost that rises only quadratically with the number of microphones. Yet, the SHT's spatially orthogonal basis functions cannot completely decorrelate sound field components over a range of time lags. In this work, we propose to exploit the compact representation offered by the SHT to reduce the number of channels used for subsequent PEVD processing. In the proposed framework for signal representation, we show that the diagonality factor improves by up to 7 dB over the microphone signal representation with a significantly lower computation cost. Moreover, when applying this framework to speech enhancement and source separation, the proposed method improves metrics known as short-time objective intelligibility (STOI) and source-to-distortion ratio (SDR) by up to 0.2 and 20 dB, respectively. Vincent W. Neo, Christine Evers, Stephan Weiss 0001, Patrick A. Naylor |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Spatial Processing Front-End for Distant ASR Exploiting Self-Attention Channel CombinatorabstractWe present a novel multi-channel front-end based on channel shortening with the Weighted Prediction Error (WPE) method followed by a fixed MVDR beamformer used in combination with a recently proposed self-attention-based channel combination (SACC) scheme, for tackling the distant ASR problem. We show that the proposed system used as part of a ContextNet based end-to-end (E2E) ASR system outperforms leading ASR systems as demonstrated by a 21.6% reduction in relative WER on a multi-channel LibriSpeech playback dataset. We also show how dereverberation prior to beam-forming is beneficial and compare the WPE method with a modified neural channel shortening approach. An analysis of the non-intrusive estimate of the signal C50 confirms that the 8 channel WPE method provides significant dereverberation of the signals (13.6 dB improvement). We also show how the weights of the SACC system allow the extraction of accurate spatial information which can be beneficial for other speech processing applications like diarization. Dushyant Sharma, Rong Gong, James Fosburgh, Stanislav Yu. Kruchinin, Patrick A. Naylor, Ljubomir Milanovic |
ICASSP | 5 |
| 2022 | Relative Acoustic Features for Distance Estimation in Smart-Homes
Francesco Nespoli, Daniel Barreda, Patrick A. Naylor |
INTERSPEECH | 3 |
| 2022 | An audio enhancement system to improve intelligibility for social-awareness in HRIabstractAbstract Improving the ability to interact through voice with a robot is still a challenge especially in real environments where multiple speakers coexist. This work has evaluated a proposal based on improving the intelligibility of the voice information that feeds an existing ASR service in the network and in conditions similar to those that could occur in a care centre for the elderly. The results indicate the feasibility and improvement of a proposal based on the use of an embedded microphone array and the use of a simple beamforming and masking technique. The system has been evaluated with 12 people and results obtained for time responsiveness indicate that the system would allow natural interaction with voice. It is shown to be necessary to incorporate a system to properly employ the masking algorithm, through the intelligent and stable estimation of the interfering signals. In addition, this approach allows to fix as sources of interest other speakers not located in the vicinity of the robot. Antonio Martinez-Colon, Raquel Viciana-Abad, José Manuel Pérez-Lorenzo, Christine Evers, Patrick A. Naylor |
Multim. Tools Appl. | 5 |
| 2022 | A Compact Noise Covariance Matrix Model for MVDR BeamformingabstractAcoustic beamforming is routinely used to improve the SNR of the received signal in applications such as hearing aids, robot audition, augmented reality, teleconferencing, source localisation and source tracking. The beamformer can be made adaptive by using an estimate of the time-varying noise covariance matrix in the spectral domain to determine an optimised beam pattern in each frequency bin that is specific to the acoustic environment and that can respond to temporal changes in it. However, robust estimation of the noise covariance matrix remains a challenging task especially in non-stationary acoustic environments. This paper presents a compact model of the signal covariance matrix that is defined by a small number of parameters whose values can be reliably estimated. The model leads to a robust estimate of the noise covariance matrix which can, in turn, be used to construct a beamformer. The performance of beamformers designed using this approach is evaluated for a spherical microphone array under a range of conditions using both simulated and measured room impulse responses. The proposed approach demonstrates consistent gains in intelligibility and perceptual quality metrics compared to the static and adaptive beamformers used as baselines. Alastair H. Moore, Sina Hafezi, Rebecca R. Vos, Patrick A. Naylor, Mike Brookes |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Multichannel Overlapping Speaker Segmentation Using Multiple Hypothesis Tracking Of Acoustic And Spatial FeaturesabstractAn essential part of any diarization system is the task of speaker segmentation which is important for many applications including speaker indexing and automatic speech recognition (ASR) in multi-speaker environments. Segmentation of overlapping speech has recently been a key focus of this work. In this paper we explore the use of a new multimodal approach for overlapping speaker segmentation that tracks both the fundamental frequency (F0) of the speaker and the speaker’s direction of arrival (DOA) simultaneously. Our proposed multiple hypothesis tracking system, which simultaneously tracks both features, shows an improvement in segmentation performance when compared to tracking these features separately. An illustrative example of overlapping speech demonstrates the effectiveness of our proposed system. We also undertake a statistical analysis on 12 meetings from the AMI corpus and show an improvement in the HIT rate of 14.1% on average against a commonly used deep learning bidirectional long short term memory network (BLSTM) approach. Aidan O. T. Hogg, Christine Evers, Patrick A. Naylor |
ICASSP | 3 |
| 2021 | Processing Pipelines for Efficient, Physically-Accurate Simulation of Microphone Array Signals in Dynamic Sound ScenesabstractMultichannel acoustic signal processing is predicated on the fact that the interchannel relationships between the received signals can be exploited to infer information about the acoustic scene. Recently there has been increasing interest in algorithms which are applicable in dynamic scenes, where the source(s) and/or microphone array may be moving. Simulating such scenes has particular challenges which are exacerbated when real-time, listener-in-the-loop evaluation of algorithms is required. This paper considers candidate pipelines for simulating the array response to a set of point/image sources in terms of their accuracy, scalability and continuity. A new approach, in which the filter kernels are obtained using principal component analysis from time-aligned impulse responses, is proposed. When the number of filter kernels is ≤40 the new approach achieves more accurate simulation than competing methods. Alastair H. Moore, Rebecca R. Vos, Patrick A. Naylor, Mike Brookes |
ICASSP | 3 |
| 2021 | Polynomial Matrix Eigenvalue Decomposition of Spherical Harmonics for Speech EnhancementabstractSpeech enhancement algorithms using polynomial matrix eigenvalue decomposition (PEVD) have been shown to be effective for noisy and reverberant speech. However, these algorithms do not scale well in complexity with the number of channels used in the processing. For a spherical microphone array sampling an order-limited sound field, the spherical harmonics provide a compact representation of the microphone signals in the form of eigenbeams. We propose a PEVD algorithm that uses only the lower dimension eigenbeams for speech enhancement at a significantly lower computation cost. The proposed algorithm is shown to significantly reduce complexity while maintaining full performance. Informal listening examples have also indicated that the processing does not introduce any noticeable artefacts. Vincent W. Neo, Christine Evers, Patrick A. Naylor |
ICASSP | 3 |
| 2021 | Overlapping Speaker Segmentation Using Multiple Hypothesis Tracking of Fundamental FrequencyabstractThis paper demonstrates how the harmonic structure of voiced speech can be exploited to segment multiple overlapping speakers in a speaker diarization task. We explore how a change in the speaker can be inferred from a change in pitch. We show that voiced harmonics can be useful in detecting when more than one speaker is talking, such as during overlapping speaker activity. A novel system is proposed to track multiple harmonics simultaneously, allowing for the determination of onsets and end-points of a speaker's utterance in the presence of an additional active speaker. This system is bench-marked against a segmentation system from the literature that employs a bidirectional long short term memory network (BLSTM) approach and requires training. Experimental results highlight that the proposed approach outperforms the BLSTM baseline approach by 12.9% in terms of HIT rate for speaker segmentation. We also show that the estimated pitch tracks of our system can be used as features to the BLSTM to achieve further improvements of 1.21% in terms of coverage and 2.45% in terms of purity. Aidan O. T. Hogg, Christine Evers, Alastair H. Moore, Patrick A. Naylor |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Enhancement of Noisy Reverberant Speech Using Polynomial Matrix Eigenvalue DecompositionabstractSpeech enhancement is important for applications such as telecommunications, hearing aids, automatic speech recognition and voice-controlled systems. Enhancement algorithms aim to reduce interfering noise and reverberation while minimizing any speech distortion. In this work for speech enhancement, we propose to use polynomial matrices to model the spatial, spectral and temporal correlations between the speech signals received by a microphone array and polynomial matrix eigenvalue decomposition (PEVD) to decorrelate in space, time and frequency simultaneously. We then propose a blind and unsupervised PEVD-based speech enhancement algorithm. Simulations and informal listening examples involving diverse reverberant and noisy environments have shown that our method can jointly suppress noise and reverberation, thereby achieving speech enhancement without introducing processing artefacts into the enhanced signal. Vincent W. Neo, Christine Evers, Patrick A. Naylor |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Speech Enhancement Based on Modulation-Domain Parametric Multichannel Kalman FilteringabstractRecently we presented a modulation-domain multichannel Kalman filtering (MKF) algorithm for speech enhancement, which jointly exploits the inter-frame modulation-domain temporal evolution of speech and the inter-channel spatial correlation to estimate the clean speech signal. The goal of speech enhancement is to suppress noise while keeping the speech undistorted, and a key problem is to achieve the best trade-off between speech distortion and noise reduction. In this paper, we extend the MKF by presenting a modulation-domain parametric MKF (PMKF) which includes a parameter that enables flexible control of the speech enhancement behaviour in each time-frequency (TF) bin. Based on the decomposition of the MKF cost function, a new cost function for PMKF is proposed, which uses the controlling parameter to weight the noise reduction and speech distortion terms. An optimal PMKF gain is derived using a minimum mean squared error (MMSE) criterion. We analyse the performance of the proposed MKF, and show its relationship to the speech distortion weighted multichannel Wiener filter (SDW-MWF). To evaluate the impact of the controlling parameter on speech enhancement performance, we further propose PMKF speech enhancement systems in which the controlling parameter is adaptively chosen in each TF bin. Experiments on a publicly available head-related impulse response (HRIR) database in different noisy and reverberant conditions demonstrate the effectiveness of the proposed method. Wei Xue 0002, Alastair H. Moore, Mike Brookes, Patrick A. Naylor |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | PEVD-Based Speech Enhancement in Reverberant EnvironmentsabstractThe enhancement of noisy speech is important for applications involving human-to-human interactions, such as telecommunications and hearing aids, as well as human-to-machine interactions, such as voice-controlled systems and robot audition. In this work, we focus on reverberant environments. It is shown that, by exploiting the lack of correlation between speech and the late reflections, further noise reduction can be achieved. This is verified using simulations involving actual acoustic impulse responses and noise from the ACE corpus. The simulations show that even without using a noise estimator, our proposed method simultaneously achieves noise reduction, and enhancement of speech quality and intelligibility, in reverberant environments over a wide range of SNRs. Furthermore, informal listening examples highlight that our approach does not introduce any significant processing artefacts such as musical noise. Vincent W. Neo, Christine Evers, Patrick A. Naylor |
ICASSP | 3 |
| 2020 | The LOCATA Challenge: Acoustic Source Localization and TrackingabstractThe ability to localize and track acoustic events is a fundamental prerequisite for equipping machines with the ability to be aware of and engage with humans in their surrounding environment. However, in realistic scenarios, audio signals are adversely affected by reverberation, noise, interference, and periods of speech inactivity. In dynamic scenarios, where the sources and microphone platforms may be moving, the signals are additionally affected by variations in the source-sensor geometries. In practice, approaches to sound source localization and tracking are often impeded by missing estimates of active sources, estimation errors, as well as false estimates. The aim of the LOCAlization and TrAcking (LOCATA) Challenge is an open-access framework for the objective evaluation and benchmarking of broad classes of algorithms for sound source localization and tracking. This article provides a review of relevant localization and tracking algorithms and, within the context of the existing literature, a detailed evaluation and dissemination of the LOCATA submissions. The evaluation highlights achievements in the field, open challenges, and identifies potential future directions. Christine Evers, Heinrich W. Löllmann, Heinrich Mellmann, Alexander Schmidt 0004, Hendrik Barfuss, Patrick A. Naylor, Walter Kellermann |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2020 | End-to-End Classification of Reverberant Rooms Using DNNsabstractReverberation is present in our workplaces, our homes, concert halls and theatres. This article investigates how deep learning can use the effect of reverberation on speech to classify a recording in terms of the room in which it was recorded. Existing approaches in the literature rely on domain expertise to manually select acoustic parameters as inputs to classifiers. Estimation of these parameters from reverberant speech is adversely affected by estimation errors, impacting the classification accuracy. In order to overcome the limitations of previously proposed methods, this paper shows how DNNs can perform the classification by operating directly on reverberant speech spectra and a CRNN with an attention-mechanism is proposed for the task. The relationship is investigated between the reverberant speech representations learned by the DNNs and acoustic parameters. For evaluation, AIRs are used from the ACE-challenge dataset that were measured in 7 real rooms. The classification accuracy of the CRNN classifier in the experiments is 78% when using 5 hours of training data and 90% when using 10 hours. Constantinos Papayiannis, Christine Evers, Patrick A. Naylor |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Speaker Change Detection Using Fundamental Frequency with Application to Multi-talker SegmentationabstractThis paper shows that time varying pitch properties can be used advantageously within the segmentation step of a multi-talker diarization system. First a study is conducted to verify that changes in pitch are strong indicators of changes in the speaker. It is then highlighted that an individual's pitch is smoothly varying and, therefore, can be predicted by means of a Kalman filter. Subsequently it is shown that if the pitch is not predictable then this is most likely due to a change in the speaker. Finally, a novel system is proposed that uses this approach of pitch prediction for speaker change detection. This system is then evaluated against a commonly used MFCC segmentation system. The proposed system is shown to increase the speaker change detection rate from 43.3% to 70.5% on meetings in the AMI corpus. Therefore, there are two equally weighted contributions in this paper: 1. We address the question of whether a change in pitch is a reliable estimator of a speaker change in multi-talk meeting audio. 2. We develop a method to extract such speaker changes and test them on a widely available meeting corpus. Aidan O. T. Hogg, Christine Evers, Patrick A. Naylor |
ICASSP | 3 |
| 2019 | Second Order Sequential Best Rotation Algorithm with Householder Reduction for Polynomial Matrix Eigenvalue DecompositionabstractThe Second-order Sequential Best Rotation (SBR2) algorithm, used for Eigenvalue Decomposition (EVD) on para-Hermitian polynomial matrices typically encountered in wideband signal processing applications like multichannel Wiener filtering and channel coding, involves a series of delay and rotation operations to achieve diagonalisation. In this paper, we proposed the use of Householder transformations to reduce polynomial matrices to tridiagonal form before zeroing the dominant element with rotation. Similar to performing Householder reduction on conventional matrices, our method enables SBR2 to converge in fewer iterations with smaller order of polynomial matrix factors because more off-diagonal Frobenius-norm (F-norm) could be transferred to the main diagonal at every iteration. A reduction in the number of iterations by 12.35% and 0.1% improvement in reconstruction error is achievable. Vincent W. Neo, Patrick A. Naylor |
ICASSP | 2 |
| 2019 | Joint Acoustic Localization and Dereverberation Through Plane Wave Decomposition and Sparse RegularizationabstractAcoustic source localization and dereverberation are formulated jointly as an inverse problem. The inverse problem consists of the approximation of the sound field measured by a set of microphones. The recorded sound pressure is matched with that of a particular acoustic model based on a collection of plane waves arriving from different directions at the microphone positions. In order to achieve meaningful results, spatial and spatio-spectral sparsity can be promoted in the weight signals controlling the plane waves. The large-scale optimization problem resulting from the inverse problem formulation is solved using a first order optimization algorithm combined with a weighted overlap-add procedure. It is shown that once the weight signals capable of effectively approximating the sound field are obtained, they can be readily used to localize a moving sound source in terms of direction of arrival (DOA) and to perform dereverberation in a highly reverberant environment. Results from simulation experiments and from real measurements show that the proposed algorithm is robust against both localized and diffuse noise exhibiting a noise reduction in the dereverberated signals. Niccolò Antonello, Enzo De Sena, Marc Moonen, Patrick A. Naylor, Toon van Waterschoot |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Noise Covariance Matrix Estimation for Rotating Microphone ArraysabstractThe noise covariance matrix computed between the signals from a microphone array is used in the design of spatial filters and beamformers with applications in noise suppression and dereverberation. This paper specifically addresses the problem of estimating the covariance matrix associated with a noise field when the array is rotating during desired source activity, as is common in head-mounted arrays. We propose a parametric model that leads to an analytical expression for the microphone signal covariance as a function of the array orientation and array manifold. An algorithm for estimating the model parameters during noise-only segments is proposed and the performance shown to be improved, rather than degraded, by array rotation. The stored model parameters can then be used to update the covariance matrix to account for the effects of any array rotation that occurs when the desired source is active. The proposed method is evaluated in terms of the Frobenius norm of the error in the estimated covariance matrix and of the noise reduction performance of a minimum variance distortionless response beamformer. In simulation experiments the proposed method achieves 18 dB lower error in the estimated noise covariance matrix than a conventional recursive averaging approach and results in noise reduction which is within 0.05 dB of an oracle beamformer using the ground truth noise covariance matrix. Alastair H. Moore, Wei Xue 0002, Patrick A. Naylor, Mike Brookes |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Joint Source Localization and Dereverberation by Sound Field Interpolation Using Sparse RegularizationabstractIn this paper, source localization and dereverberation are formulated jointly as an inverse problem. The inverse problem consists in the interpolation of the sound field measured by a set of microphones by matching the recorded sound pressure with that of a particular acoustic model. This model is based on a collection of equivalent sources creating either spherical or plane waves. In order to achieve meaningful results, spatial, spatio-temporal and spatio-spectral sparsity can be promoted in the signals originating from the equivalent sources. The inverse problem consists of a large-scale optimization problem that is solved using a first order matrix-free optimization algorithm. It is shown that once the equivalent source signals capable of effectively interpolating the sound field are obtained, they can be readily used to localize a speech sound source in terms of Direction of Arrival (DOA) and to perform dereverberation in a highly reverberant environment. Niccolò Antonello, Enzo De Sena, Marc Moonen, Patrick A. Naylor, Toon van Waterschoot |
ICASSP | 4 |
| 2018 | Room Identification Using Frequency Dependence of Spectral Decay StatisticsabstractA method for room identification is proposed based on the reverberation properties of multichannel speech recordings. The approach exploits the dependence of spectral decay statistics on the reverberation time of a room. The average negative-side variance within 1/3-octave bands is proposed as the identifying feature and shown to be effective in a classification experiment. However, negative-side variance is also dependent on the direct-to-reverberant energy ratio. The resulting sensitivity to different spatial configurations of source and microphones within a room are mitigated using a novel reverberation enhancement algorithm. A classification experiment using speech convolved with measured impulse responses and contaminated with environmental noise demonstrates the effectiveness of the proposed method, achieving 79% correct identification in the most demanding condition compared to 40% using unenhanced signals. Alastair H. Moore, Patrick A. Naylor, Mike Brookes |
ICASSP | 2 |
| 2018 | Multichannel Kalman Filtering for Speech EhnancementabstractThe use of spatial information in multichannel speech enhancement methods is well established but information associated with the temporal evolution of speech is less commonly exploited. Speech signals can be modelled using an autoregressive process in the time-frequency modulation domain, and Kalman filtering based speech enhancement algorithms have been developed for single-channel processing. In this paper, a multichannel Kalman filter (MKF) for speech enhancement is derived that jointly considers the multichannel spatial information and the temporal correlations of speech. We model the temporal evolution of speech in the modulation domain and, by incorporating the spatial information, an optimal MKF gain is derived in the short-time Fourier transform domain. We also show that the proposed MKF becomes a conventional multichannel Wiener filter if the temporal information is discarded. Experiments using the signals generated from a public head-related impulse response database demonstrate the effectiveness of the proposed method in comparison to other techniques. Wei Xue 0002, Alastair H. Moore, Mike Brookes, Patrick A. Naylor |
ICASSP | 4 |
| 2018 | Acoustic Analysis and Assessment of the Knee in Osteoarthritis During WalkingabstractWe examine the relation between the sounds emitted by the knee joint during walking and its condition, with particular focus on osteoarthritis, and investigate their potential for noninvasive detection of knee pathology. We present a comparative analysis of several features and evaluate their discriminant power for the task of normal-abnormal signal classification. We statistically evaluate the feature distributions using the two-sample Kolmogorov-Smirnov test and the Bhattacharyya distance. We propose the use of 11 statistics to describe the distributions and test with several classifiers. In our experiments with 249 normal and 297 abnormal acoustic signals from 40 knees, a Support Vector Machine with linear kernel gave the best results with an error rate of 13.9%. Costas Yiallourides, Alastair H. Moore, Edouard Auvinet, Catherine Van Der Straeten, Patrick A. Naylor |
ICASSP | 5 |
| 2018 | DoA Reliability for Distributed Acoustic TrackingabstractDistributed acoustic tracking estimates the trajectories of source positions using an acoustic sensor network. As it is often difficult to estimate the source-sensor range from individual nodes, the source positions have to be inferred from the direction-of-arrival (DoA) estimates. Due to reverberation and noise, the sound field becomes increasingly diffuse with increasing source-sensor distance, leading to a decreased Direction of Arrival (DoA)-estimation accuracy. To distinguish between accurate and uncertain DoA estimates, this letter proposes to incorporate the coherent-to-diffuse ratio as a measure of DoA reliability for single-source tracking. It is shown that the source positions, therefore, can be probabilistically triangulated by exploiting the spatial diversity of all nodes. Christine Evers, Emanuël A. P. Habets, Sharon Gannot, Patrick A. Naylor |
IEEE Signal Process. Lett. | 4 |
| 2018 | Acoustic SLAMabstractAn algorithm is presented that enables devices equipped with microphones, such as robots, to move within their environment in order to explore, adapt to, and interact with sound sources of interest. Acoustic scene mapping creates a three-dimensional (3D) representation of the positional information of sound sources across time and space. In practice, positional source information is only provided by Direction-of-Arrival (DoA) estimates of the source directions; the source-sensor range is typically difficult to obtain. DoA estimates are also adversely affected by reverberation, noise, and interference, leading to errors in source location estimation and consequent false DoA estimates. Moreover, many acoustic sources, such as human talkers, are not continuously active, such that periods of inactivity lead to missing DoA estimates. Withal, the DoA estimates are specified relative to the observer's sensor location and orientation. Accurate positional information about the observer therefore is crucial. This paper proposes Acoustic Simultaneous Localization and Mapping (aSLAM), which uses acoustic signals to simultaneously map the 3D positions of multiple sound sources while passively localizing the observer within the scene map. The performance of aSLAM is analyzed and evaluated using a series of realistic simulations. Results are presented to show the impact of the observer motion and sound source localization accuracy. Christine Evers, Patrick A. Naylor |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Modulation-Domain Multichannel Kalman Filtering for Speech EnhancementabstractCompared with single-channel speech enhancement methods, multichannel methods can utilize spatial information to design optimal filters. Although some filters adaptively consider second-order signal statistics, the temporal evolution of the speech spectrum is usually neglected. By using linear prediction (LP) to model the inter-frame temporal evolution of speech, single-channel Kalman filtering (KF) based methods have been developed for speech enhancement. In this paper, we derive a multichannel KF (MKF) that jointly uses both interchannel spatial correlation and interframe temporal correlation for speech enhancement. We perform LP in the modulation domain, and by incorporating the spatial information, derive an optimal MKF gain in the short-time Fourier transform domain. We show that the proposed MKF reduces to the conventional multichannel Wiener filter if the LP information is discarded. Furthermore, we show that, under an appropriate assumption, the MKF is equivalent to a concatenation of the minimum variance distortion response beamformer and a single-channel modulation-domain KF and therefore present an alternative implementation of the MKF. Experiments conducted on a public head-related impulse response database demonstrate the effectiveness of the proposed method. Wei Xue 0002, Alastair H. Moore, Mike Brookes, Patrick A. Naylor |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | Source tracking using moving microphone arrays for robot auditionabstractIntuitive spoken dialogues are a prerequisite for human-robot interaction. In many practical situations, robots must be able to identify and focus on sources of interest in the presence of interfering speakers. Techniques such as spatial filtering and blind source separation are therefore often used, but rely on accurate knowledge of the source location. In practice, sound emitted in enclosed environments is subject to reverberation and noise. Hence, sound source localization must be robust to both diffuse noise due to late reverberation, as well as spurious detections due to early reflections. For improved robustness against reverberation, this paper proposes a novel approach for sound source tracking that constructively exploits the spatial diversity of a microphone array installed in a moving robot. In previous work, we developed speaker localization approaches using expectation-maximization (EM) approaches and using Bayesian approaches. In this paper we propose to combine the EM and Bayesian approach in one framework for improved robustness against reverberation and noise. Christine Evers, Yuval Dorfan, Sharon Gannot, Patrick A. Naylor |
ICASSP | 4 |
| 2017 | Multiple source localization using Estimation Consistency in the Time-Frequency domainabstractThe extraction of multiple Direction-of-Arrival (DoA) information from estimated spatial spectra can be challenging when such spectra are noisy or the sources are adjacent. Smoothing or clustering techniques are typically used to remove the effect of noise or irregular peaks in the spatial spectra. As we will explain and show in this paper, the smoothing-based techniques require prior knowledge of minimum angular separation of the sources and the clustering-based techniques fail on noisy spatial spectrum. A broad class of localization techniques give direction estimates in each Time Frequency (TF) bin. Using this information as input, a novel technique for obtaining robust localization of multiple simultaneous sources is proposed using Estimation Consistency (EC) in the TF domain. The method is evaluated in the context of spherical microphone arrays. This technique does not require prior knowledge of the sources and by removing the noise in the estimated spatial spectrum makes clustering a reliable and robust technique for multiple DoA extraction from estimated spatial spectra. The results indicate that the proposed technique has the strongest robustness to separation with up to 10° median error for 5° to 180° separation for 2 and 3 sources, compared to the baseline and the state-of-the-art techniques. Sina Hafezi, Alastair H. Moore, Patrick A. Naylor |
ICASSP | 3 |
| 2017 | Measuring, modelling and predicting perceived reverberationabstractThis paper investigates the relationship between the perceived level of reverberation and parameters measured from the room impulse response (RIR), as well as the design of an instrumental measure that predicts this perceived level. We first present the results of an experimental listening test conducted to assess the level of perceived reverberation in speech captured by a single microphone, before analysing the gathered data to assess the influence of parameters such as the reverberation time (T60) or the direct-to-reverberant ratio (DRR). Secondly, we use the results of this analysis to improve the signal based reverberation decay tail (RDT) measure, previously proposed by the authors to predict the perceived level of reverberation. The accuracy of the proposed measure is evaluated in terms of correlation with the subjective scores and compared to the performance of predictors using parameters extracted from the RIR. Results show that the proposed modifications to the RDT does improve its accuracy. Though still slightly outperformed by measures based on parameters of the RIR, we believe the proposed measure to be useful in scenarios in which the RIR or its parameters are unknown. Hamza A. Javed, Benjamin Cauchi, Simon Doclo, Patrick A. Naylor, Stefan Goetze |
ICASSP | 4 |
| 2017 | Improving the perceptual quality of ideal binary masked speechabstractIt is known that applying a time-frequency binary mask to very noisy speech can improve its intelligibility but results in poor perceptual quality. In this paper we propose a new approach to applying a binary mask that combines the intelligibility gains of conventional binary masking with the perceptual quality gains of a classical speech enhancer. The binary mask is not applied directly as a time-frequency gain as in most previous studies. Instead, the mask is used to supply prior information to a classical speech enhancer about the probability of speech presence in different time-frequency regions. Using an oracle ideal binary mask, we show that the proposed method results in a higher predicted quality than other methods of applying a binary mask whilst preserving the improvements in predicted intelligibility. Leo Lightburn, Enzo De Sena, Alastair H. Moore, Patrick A. Naylor, Mike Brookes |
ICASSP | 4 |
| 2017 | Robust spherical harmonic domain interpolation of spatially sampled array manifoldsabstractAccurate interpolation of the array manifold is an important first step for the acoustic simulation of rapidly moving microphone arrays. Spherical harmonic domain interpolation has been proposed and well studied in the context of head-related transfer functions but has focussed on perceptual, rather than numerical, accuracy. In this paper we analyze the effect of measurement noise on spatial aliasing. Based on this analysis we propose a method for selecting the truncation orders for the forward and reverse spherical Fourier transforms given only the noisy samples in such a way that the interpolation error is minimized. The proposed method achieves up to 1.7 dB improvement over the baseline approach. Alastair H. Moore, Mike Brookes, Patrick A. Naylor |
ICASSP | 3 |
| 2017 | Discriminative feature domains for reverberant acoustic environmentsabstractSeveral speech processing and audio data-mining applications rely on a description of the acoustic environment as a feature vector for classification. The discriminative properties of the feature domain play a crucial role in the effectiveness of these methods. In this work, we consider three environment identification tasks and the task of acoustic model selection for speech recognition. A set of acoustic parameters and Machine Learning algorithms for feature selection are used and an analysis is performed on the resulting feature domains for each task. In our experiments, a classification accuracy of 100% is achieved for the majority of tasks and the Word Error Rate is reduced by 20.73 percentage points for Automatic Speech Recognition when using the resulting domains. Experimental results indicate a significant dissimilarity in the parameter choices for the composition of the domains, which highlights the importance of the feature selection process for individual applications. Constantinos Papayiannis, Christine Evers, Patrick A. Naylor |
ICASSP | 3 |
| 2017 | Channel estimation for crosstalk cancellation in wireless acoustic networksabstractIn this paper we deal with the estimation of the room impulse response (RIR) between each loudspeaker and each microphone of a wireless acoustic network of two nodes when used to implement a crosstalk canceller. The nodes of the network are commercial devices connected via standard wireless links, presenting low computational requirements and non-ideal synchronization between them. Moreover, the nodes can exchange information, but they cannot share their signals due to the high throughput and perfect synchronism that would be required. The proposed scheme adaptively estimates the global impulse response between the source signals and the recorded signal at each node of the network, and afterwards estimates the corresponding RIRs between each loudspeaker and the node's microphone. This scheme does not need any additional synchronism between loudspeakers. Simulations show that proportionate-type affine projection algorithms obtain good performance for order N = 4, being their cost affordable in commercial devices. Gema Piñero, Patrick A. Naylor |
ICASSP | 2 |
| 2017 | Frequency-domain under-modelled blind system identification based on cross power spectrum and sparsity regularizationabstractIn room acoustics, under-modelled multichannel blind system identification (BSI) aims to estimate the early part of the room impulse responses (RIRs), and it can be widely used in applications such as speaker localization, room geometry identification and beamforming based speech dereverberation. In this paper we extend our recent study on under-modelled BSI from the time domain to the frequency domain, such that the RIRs can be updated frame-wise and the efficiency of Fast Fourier Transform (FFT) is exploited to reduce the computational complexity. Analogous to the cross-correlation based criterion in the time domain, a frequency-domain cross power spectrum based criterion is proposed. As the early RIRs are usually sparse, the RIRs are estimated by jointly maximizing the cross power spectrum based criterion in the frequency domain and minimizing the l1-norm sparsity measure in the time domain. A two-stage LMS updating algorithm is derived to achieve joint optimization of these two targets. The experimental results in different under-modelled scenarios demonstrate the effectiveness of the proposed method. Wei Xue 0002, Mike Brookes, Patrick A. Naylor |
ICASSP | 3 |
| 2017 | A dynamic programming approach for automatic stride detection and segmentation in acoustic emission from the kneeabstractWe study the acquisition and analysis of sounds generated by the knee during walking with particular focus on the effects due to osteoarthritis. Reliable contact instant estimation is essential for stride synchronous analysis. We present a dynamic programming based algorithm for automatic estimation of both the initial contact instants (ICIs) and last contact instants (LCIs) of the foot to the floor. The technique is designed for acoustic signals sensed at the patella of the knee. It uses the phase-slope function to generate a set of candidates and then finds the most likely ones by minimizing a cost function that we define. ICIs are identified with an RMS error of 13.0% for healthy and 14.6% for osteoarthritic knees and LCIs with an RMS error of 16.0% and 17.0% respectively. Costas Yiallourides, Victoria Manning-Eid, Alastair H. Moore, Patrick A. Naylor |
ICASSP | 4 |
| 2017 | Speech enhancement for robust automatic speech recognition: Evaluation using a baseline system and instrumental measuresabstractAutomatic speech recognition in everyday environments must be robust to significant levels of reverberation and noise. One strategy to achieve such robustness is multi-microphone speech enhancement. In this study, we present results of an evaluation of different speech enhancement pipelines using a state-of-the-art ASR system for a wide range of reverberation and noise conditions. The evaluation exploits the recently released ACE Challenge database which includes measured multichannel acoustic impulse responses from 7 different rooms with reverberation times ranging from 0.33 to 1.34 s. The reverberant speech is mixed with ambient, fan and babble noise recordings made with the same microphone setups in each of the rooms. In the first experiment, performance of the ASR without speech processing is evaluated. Results clearly indicate the deleterious effect of both noise and reverberation. In the second experiment, different speech enhancement pipelines are evaluated with relative word error rate reductions of up to 82%. Finally, the ability of selected instrumental metrics to predict ASR performance improvement is assessed. The best performing metric, Short-Time Objective Intelligibility Measure, is shown to have a Pearson correlation coefficient of 0.79, suggesting that it is a useful predictor of algorithm performance in these tests. Alastair H. Moore, Pablo Peso Parada, Patrick A. Naylor |
Comput. Speech Lang. | 3 |
| 2017 | Room Impulse Response Interpolation Using a Sparse Spatio-Temporal Representation of the Sound FieldabstractRoom Impulse Responses (RIRs) are typically measured using a set of microphones and a loudspeaker. When RIRs spanning a large volume are needed, many microphone measurements must be used to spatially sample the sound field. In order to reduce the number of microphone measurements, RIRs can be spatially interpolated. In the present study, RIR interpolation is formulated as an inverse problem. This inverse problem relies on a particular acoustic model capable of representing the measurements. Two different acoustic models are compared: the plane wave decomposition model and a novel time-domain model, which consists of a collection of equivalent sources creating spherical waves. These acoustic models can both approximate any reverberant sound field created by a far-field sound source. In order to produce an accurate RIR interpolation, sparsity regularization is employed when solving the inverse problem. In particular, by combining different acoustic models with different sparsity promoting regularizations, spatial sparsity, spatio-spectral sparsity, and spatio-temporal sparsity are compared. The inverse problem is solved using a matrix-free large-scale optimization algorithm. Simulations show that the best RIR interpolation is obtained when combining the novel time-domain acoustic model with the spatio-temporal sparsity regularization, outperforming the results of the plane wave decomposition model even when far fewer microphone measurements are available. Niccolò Antonello, Enzo De Sena, Marc Moonen, Patrick A. Naylor, Toon van Waterschoot |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | Single-Channel Online Enhancement of Speech Corrupted by Reverberation and NoiseabstractThis paper proposes an online single-channel speech enhancement method designed to improve the quality of speech degraded by reverberation and noise. Based on an autoregressive model for the reverberation power and on a hidden Markov model for clean speech production, a Bayesian filtering formulation of the problem is derived and online joint estimation of the acoustic parameters and mean speech, reverberation, and noise powers is obtained in mel-frequency bands. From these estimates, a real-valued spectral gain is derived and spectral enhancement is applied in the short-time Fourier transform (STFT) domain. The method yields state-of-the-art performance and greatly reduces the effects of reverberation and noise while improving speech quality and preserving speech intelligibility in challenging acoustic environments. Clement S. J. Doire, Mike Brookes, Patrick A. Naylor, Christopher M. Hicks, Dave Betts, Mohammad A. Dmour, Søren Holdt Jensen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Augmented Intensity Vectors for Direction of Arrival Estimation in the Spherical Harmonic DomainabstractPseudointensity vectors (PIVs) provide a means of direction of arrival (DOA) estimation for spherical microphone arrays using only the zeroth and the first-order spherical harmonics. An augmented intensity vector (AIV) is proposed which improves the accuracy of PIVs by exploiting higher order spherical harmonics. We compared DOA estimation using our proposed AIVs against PIVs, steered response power (SRP) and subspace methods where the number of sources, their angular separation, the reverberation time of the room and the sensor noise level are varied. The results show that the proposed approach outperforms the baseline methods and performs at least as accurately as the state-of-theart method with strong robustness to reverberation, sensor noise, and number of sources. In the single and multiple source scenarios tested, which include realistic levels of reverberation and noise, the proposed method had average error of 1.5° and 2°, respectively. Sina Hafezi, Alastair H. Moore, Patrick A. Naylor |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Direction of Arrival Estimation in the Spherical Harmonic Domain Using Subspace Pseudointensity VectorsabstractDirection of arrival (DOA) estimation is a fundamental problem in acoustic signal processing. It is used in a diverse range of applications, including spatial filtering, speech dereverberation, source separation and diarization. Intensity vector-based DOA estimation is attractive, especially for spherical sensor arrays, because it is computationally efficient. Two such methods are presented that operate on a spherical harmonic decomposition of a sound field observed using a spherical microphone array. The first uses pseudointensity vectors (PIVs) and works well in acoustic environments where only one sound source is active at any time. The second uses subspace pseudointensity vectors (SSPIVs) and is targeted at environments where multiple simultaneous soures and significant levels of reverberation make the problem more challenging. Analytical models are used to quantify the effects of an interfering source, diffuse noise, and sensor noise on PIVs and SSPIVs. The accuracy of DOA estimation using PIVs and SSPIVs is compared against the state of the art in simulations including realistic reverberation and noise for single and multiple, stationary and moving sources. Finally, robust performance of the proposed methods is demonstrated by using speech recordings in a real acoustic environment. Alastair H. Moore, Christine Evers, Patrick A. Naylor |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | On Perceptual Audio Compression with Side Information at the DecoderabstractDue to the distributed structure of many modern audio transmission setups, it is likely to have an observation at the receiver which is correlated with the desired source at the transmitter. This observation could be used as side information to reduce the transmission rate using distributed source coding. How to integrate distributed source coding into the perceptual audio compression procedure is thus a fundamental question. In this paper, we take a completely analytical approach to this problem, in particular to the rate-distortion trade-off and the corresponding coding schemes. We then interpret the results from an audio coding perspective. The main result is that, to upgrade a regular perceptual audio coder to a distributed coder, one needs to revise the perceptual masking curve. The revised masking curve models the availability of the side information as an extra masking effect, yielding lower rates. Interestingly, this means that at least conceptually, the distributed coding scenario could be integrated into the audio coder with minor changes, and without destructing the original coder. Adel Zahedi, Jan Østergaard, Søren Holdt Jensen, Patrick A. Naylor, Soren Bech |
DCC | 4 |
| 2016 | Perceptual and instrumental evaluation of the perceived level of reverberationabstractPerceptual measures are usually considered more reliable than instrumental measures for evaluating the perceived level of reverberation. However, such measures are costly in both time and money, and, due to variations in stimuli or assessors, the resulting data is not always statistically significant. Therefore, an efficient perceptual measure of the perceived level of reverberation is needed. We compare the use of a multiple stimuli test with the use of pairwise comparison for the evaluation of the perceived level of reverberation. The results suggest that using multiple stimuli is preferable to pairwise comparison as long as the number of conditions to be compared is not too large. Additionally, we use the results from the conducted perceptual measurements to examine the reliability of existing instrumental measures of the perceived level of reverberation. Our observations show which instrumental measures are effective in highlighting differences between RIR characteristics and which ones have to be preferred if one aims at predicting the level of reverberation perceived by a human assessor. Benjamin Cauchi, Hamza A. Javed, Timo Gerkmann, Simon Doclo, Stefan Goetze, Patrick A. Naylor |
ICASSP | 6 |
| 2016 | Acoustic simultaneous localization and mapping (A-SLAM) of a moving microphone array and its surrounding speakersabstractAcoustic scene mapping creates a representation of positions of audio sources such as talkers within the surrounding environment of a microphone array. By allowing the array to move, the acoustic scene can be explored in order to improve the map. Furthermore, the spatial diversity of the kinematic array allows for estimation of the source-sensor distance in scenarios where source directions of arrival are measured. As sound source localization is performed relative to the array position, mapping of acoustic sources requires knowledge of the absolute position of the microphone array in the room. If the array is moving, its absolute position is unknown in practice. Hence, Simultaneous Localization and Mapping (SLAM) is required in order to localize the microphone array position and map the surrounding sound sources. In realistic environments, microphone arrays receive a convolutive mixture of direct-path speech signals, noise and reflections due to reverberation. A key challenge of Acoustic SLAM (a-SLAM) is robustness against reverberant clutter measurements and missing source detections. This paper proposes a novel bearing-only a-SLAM approach using a Single-Cluster Probability Hypothesis Density filter. Results demonstrate convergence to accurate estimates of the array trajectory and source positions. Christine Evers, Alastair H. Moore, Patrick A. Naylor |
ICASSP | 3 |
| 2016 | 3D acoustic source localization in the spherical harmonic domain based on optimized grid searchabstractAn approach for 3D source localization using a spherical microphone array is proposed that gives improved accuracy compared to intensity-based methods. First order spherical harmonics are first used to obtain an initial approximate localization result and then the initial result is improved based on an optimized grid search in the local vicinity using the method of least squares and high-order spherical harmonics. We show that this approach outperforms the first-order approach and shows strong robustness to reverberation and noise. The worst average error of 3 degrees was found in our experiments in the presence of realistic reverberation and noise. Sina Hafezi, Alastair H. Moore, Patrick A. Naylor |
ICASSP | 3 |
| 2016 | Spherical microphone array acoustic rake receiversabstractSeveral signal independent acoustic rake receivers are proposed for speech dereverberation using spherical microphone arrays. The proposed rake designs take advantage of multipaths, by separately capturing and combining early reflections with the direct path. We investigate several approaches in combining reflections with the direct path source signal, including the development of beam patterns that point nulls at all preceding reflections. The proposed designs are tested in experimental simulations and their dereverberation performances evaluated using objective measures. For the tested configuration, the proposed designs achieve higher levels of dereverberation compared to conventional signal independent beamforming systems; achieving up to 3.6 dB improvement in the direct-to-reverberant ratio over the plane-wave decomposition beamformer. Hamza A. Javed, Alastair H. Moore, Patrick A. Naylor |
ICASSP | 3 |
| 2016 | A data-driven non-intrusive measure of speech quality and intelligibility
Dushyant Sharma, Yu Wang 0027, Patrick A. Naylor, Mike Brookes |
Speech Commun. | 3 |
| 2016 | Estimation of Room Acoustic Parameters: The ACE ChallengeabstractReverberation time (T60) and Direct-to-reverberant ratio (DRR) are important parameters which together can characterize sound captured by microphones in nonanechoic rooms. These parameters are important in speech processing applications such as speech recognition and dereverberation. The values of T60and DRR can be estimated directly from the acoustic impulse response (AIR) of the room. In practice, the AIR is not normally available, in which case these parameters must be estimated blindly from the observed speech in the microphone signal. The acoustic characterization of environments (ACE) challenge aimed to determine the state-of-the-art in blind acoustic parameter estimation and also to stimulate research in this area. A summary of the ACE challenge, and the corpus used in the challenge is presented together with an analysis of the results. Existing algorithms were submitted alongside novel contributions, the comparative results for which are presented in this paper. The challenge showed that T60estimation is a mature field where analytical approaches dominate whilst DRR estimation is a less mature field where machine learning approaches are currently more successful. James Eaton, Nikolay D. Gaubitch, Alastair H. Moore, Patrick A. Naylor |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2016 | A Single-Channel Non-Intrusive C50 Estimator Correlated With Speech Recognition PerformanceabstractSeveral intrusive measures of reverberation can be computed from measured and simulated room impulse responses, over the full frequency band or for each individual mel-frequency subband. It is initially shown that full-band clarity index C50is the most correlated measure on average with reverberant speech recognition performance. This corroborates previous findings but now for the dataset to be used in this study. We extend the previous findings to show that C50also exhibits the highest mutual information on average. Motivated by these extended findings, a nonintrusive room acoustic (NIRA) estimation method is proposed to estimate C50from only the reverberant speech signal. The NIRA method is a data-driven approach based on computing a number of features from the speech signal and it employs these features to train a model used to perform the estimation. The choice of features and learning techniques are explored in this work using an evaluation set which comprises approximately 100 000 different reverberant signals (around 93 h of speech) including reverberation from measured and simulated room impulse responses. The feature importance of each feature with respect to the estimation of the target C50is analysed following two different approaches. In both cases, the newly chosen set of features shows high importance for the target. The best C50estimator provides a root-mean-square deviation around 3 dB on average for all reverberant test environments. Pablo Peso Parada, Dushyant Sharma, Jose Lainez, Daniel Barreda, Toon van Waterschoot, Patrick A. Naylor |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2015 | Coding and Enhancement in Wireless Acoustic Sensor NetworksabstractWe formulate a new problem which bridges between source coding and enhancement in wireless acoustic sensor networks. We consider a network of wireless microphones, each of which encoding its own measurement under a covariance matrix distortion constraint and sending it to a fusion center. To process the data at the center, we use a recent spatio-temporal prediction filter. We assume that a weighted sum-rate for the network is specified. The problem is to allocate optimal distortion matrices to the nodes in order to achieve a maximum output SNR at the fusion center after processing the received data, while the weighted sum-rate for the network is no more than the specified value. We formulate this problem as an optimization problem for which we derive a set of equalities imposed on the solution by studying the KKT conditions. In particular, for the special case of scalar sources with two microphones and a sum-rate constraint, we derive the distortion allocation in closed form and will show that if the given sum-rate is higher than a critical value, the stationary points from the KKT conditions lead to distortion allocations which maximize the output SNR of the filter. Adel Zahedi, Jan Østergaard, Søren Holdt Jensen, Patrick A. Naylor, Soren Bech |
DCC | 4 |
| 2015 | Single-channel blind estimation of reverberation parametersabstractThe reverberation of an acoustic channel can be characterised by two frequency-dependent parameters: the reverberation time and the direct-to-reverberant energy ratio. This paper presents an algorithm for blindly determining these parameters from a single-channel speech signal. The algorithm uses an extended Kalman filter to estimate the parameters together with a hidden semi-Markov model to identify intervals of speech activity. Clement S. J. Doire, Mike Brookes, Patrick A. Naylor, Dave Betts, Christopher M. Hicks, Mohammad A. Dmour, Søren Holdt Jensen |
ICASSP | 3 |
| 2015 | Direct-to-Reverberant Ratio estimation using a null-steered beamformerabstractReverberation affects the quality and intelligibility of distant speech recorded in a room. Direct-to-Reverberant Ratio (DRR) is a useful measure for assessing the acoustic configuration and can be used to inform dereverberation algorithms. We describe a novel DRR estimation algorithm applicable where the signal was recorded with two or more microphones, such as mobile communications devices and laptops. The method uses a null-steered beamformer. In simulations the proposed method yields accurate DRR estimates to within ±4 dB across a wide variety of room sizes, reverberation times and source-receiver distances. It is also shown that the proposed method is more robust to background noise than a baseline approach. The best estimation accuracy is obtained in the region from -5 to 5 dB which is a relevant range for portable devices. James Eaton, Alastair H. Moore, Patrick A. Naylor, Jan Skoglund |
ICASSP | 3 |
| 2015 | Speaker change detection and speaker diarization using spatial informationabstractIn this paper, we present a novel speaker change detection and speaker diarization algorithm using spatial information in the form of features derived from estimated Room Impulse Response (RIR)s. A blind system identification approach is used to obtain an estimate of the RIRs, from which the C5 feature is derived and used in the labeling algorithm. Experimental results using 2 speakers for different locations within a fixed room show that our approach achieves a higher hit rate in the speaker change detection task and a lower variance in the diarization error rate when compared with a baseline algorithm. Mathieu Hu, Dushyant Sharma, Simon Doclo, Mike Brookes, Patrick A. Naylor |
ICASSP | 5 |
| 2015 | Audio coding in wireless acoustic sensor networks
Adel Zahedi, Jan Østergaard, Søren Holdt Jensen, Soren Bech, Patrick A. Naylor |
Signal Process. | 5 |
| 2014 | Distributed Remote Vector Gaussian Source Coding for Wireless Acoustic Sensor NetworksabstractIn this paper, we consider the problem of remote vector Gaussian source coding for a wireless acoustic sensor network. Each node receives messages from multiple nodes in the network and decodes these messages using its own measurement of the sound field as side information. The node's measurement and the estimates of the source resulting from decoding the received messages are then jointly encoded and transmitted to a neighbouring node in the network. We show that for this distributed source coding scenario, one can encode a so-called conditional sufficient statistic of the sources instead of jointly encoding multiple sources. We focus on the case where node measurements are in form of noisy linearly mixed combinations of the sources and the acoustic channel mixing matrices are invertible. For this problem, we derive the rate-distortion function for vector Gaussian sources and under covariance distortion constraints. Adel Zahedi, Jan Østergaard, Søren Holdt Jensen, Patrick A. Naylor, Soren Bech |
DCC | 4 |
| 2014 | Noise-robust detection of peak-clipping in decoded speechabstractClipping is a commonplace problem in voice telecommunications and detection of clipping is useful in a range of speech processing applications. We analyse and evaluate the performance of three previously presented algorithms for clipping detection in decoded speech in high levels of ambient noise. We identify a baseline method which is well known for clipping detection, determine experimentally the optimized operation parameter for the baseline approach, and use this in our experiments. Our results indicate that the new algorithms outperform the baseline except at extreme levels of clipping and negative signal-to-noise ratios. James Eaton, Patrick A. Naylor |
ICASSP | 2 |
| 2014 | Non-intrusive estimation of the level of reverberation in speechabstractWe show corroborating evidence that, among a set of common acoustic parameters, the clarity index C50provides a measure of reverberation that is well correlated with speech recognition accuracy. We also present a data driven method for non-intrusive C50parameter estimation from a single channel speech signal. The method extracts a number of features from the speech signal and uses a binary regression tree, trained on appropriate training data, to estimate the C50. Evaluation is carried out using speech utterances convolved with real and simulated room impulse responses, and additive babble noise. The new method outperforms a baseline approach in our evaluation. Pablo Peso Parada, Dushyant Sharma, Patrick A. Naylor |
ICASSP | 3 |
| 2014 | Distributed remote vector gaussian source coding with covariance distortion constraintsabstractIn this paper, we consider a distributed remote source coding problem, where a sequence of observations of source vectors is available at the encoder. The problem is to specify the optimal rate for encoding the observations subject to a covariance matrix distortion constraint and in the presence of side information at the decoder. For this problem, we derive lower and upper bounds on the rate-distortion function (RDF) for the Gaussian case, which in general do not coincide. We then provide some cases, where the RDF can be derived exactly. We also show that previous results on specific instances of this problem can be generalized using our results. We finally show that if the distortion measure is the mean squared error, or if it is replaced by a certain mutual information constraint, the optimal rate can be derived from our main result. Adel Zahedi, Jan Østergaard, Søren Holdt Jensen, Patrick A. Naylor, Soren Bech |
ISIT | 4 |
| 2014 | Noise Reduction in the Spherical Harmonic Domain Using a Tradeoff Beamformer and Narrowband DOA EstimatesabstractIn noise reduction, a common approach is to use a microphone array with a beamformer that combines the individual microphone signals to extract a desired speech signal. The beamformer weights usually depend on the statistics of the noise and desired speech signals, which cannot be directly observed and must be estimated. Estimators based on the speech presence probability (SPP) seek to update the statistics estimates only when desired speech is known to be absent or present. However, they do not normally distinguish between desired and undesired speech sources. In this contribution, an algorithm is proposed to distinguish between these two types of sources using additional spatial information, by estimating a desired speech presence probability based on the combination of a multichannel SPP and a direction of arrival (DOA) based probability. The DOA-based probability is computed using DOA estimates for each time-frequency bin. The estimated statistics are then used to compute the weights of a spherical harmonic domain tradeoff beamformer, which achieves a balance between noise reduction and speech distortion. The performance evaluation demonstrates the effectiveness of the proposed approach at suppressing both background noise and spatially coherent noise. A number of audio examples and sample spectrograms are also provided. Daniel P. Jarrett, Maja Taseska, Emanuël A. P. Habets, Patrick A. Naylor |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2014 | Robust Multichannel Dereverberation using Relaxed Multichannel Least SquaresabstractA novel approach is proposed for robust multichannel dereverberation in the presence of system identification error (SIEs), based on channel shortening. A mathematical link is derived between the well known multiple-input/output inverse theorem (MINT) algorithm and channel shortening. The relaxed multichannel least squares (RMCLS) algorithm is then proposed as an efficient realization within the channel shortening paradigm and is shown through experimental results to outperform MINT in the presence of SIEs. While the RMCLS is robust to SIEs, the coloration of the output cannot be controlled. Two extensions to RMCLS are proposed to control the level of coloration and the performances of both extensions are evaluated comparatively. It is shown that both substantially maintain the dereverberation performance and robustness to SIEs obtained from RMCLS while effectively controlling the level of coloration introduced. Felicia Lim, Wancheng Zhang, Emanuël A. P. Habets, Patrick A. Naylor |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2013 | Noise-robust reverberation time estimation using spectral decay distributions with reduced computational costabstractReverberation Time (T60) is an important measure of the acoustic properties of a room. It can provide information about the acoustic environment, the intelligibility, and quality of speech recorded in the room, and help improve the performance of speech processing algorithms with reverberant speech. Where the acoustic impulse response of the room is not available, the T60must be estimated non-intrusively from reverberant speech. State-of-the-art non-intrusive T60estimators have been shown to be strongly biased in the presence of noise. We describe a novel T60estimation algorithm based on spectral decay distributions that provides robustness to additive noise for a range of realistic noise types for signal-to-noise ratios in the range 0 to 35 dB and T60s between 200 and 950 ms. The proposed method also has much reduced computational cost. James Eaton, Nikolay D. Gaubitch, Patrick A. Naylor |
ICASSP | 3 |
| 2013 | Spherical harmonic domain noise reduction using an MVDR beamformer and DOA-based second-order statistics estimationabstractMost beamformers used for noise reduction rely on the accurate estimation of the second-order statistics of the noise, and in some cases, of the desired signal. Speech presence probability (SPP) based statistics estimators seek to update the estimates only when speech is absent/present, however, when used with a fixed a priori SPP, they cannot distinguish between a coherent desired source and coherent noise sources. We propose to distinguish between desired and noise sources by estimating the second-order statistics with a direction of arrival dependent a priori SPP, which we then use to compute the weights of a spherical harmonic domain minimum variance distortionless response filter. Daniel P. Jarrett, Emanuël A. P. Habets, Patrick A. Naylor |
ICASSP | 3 |
| 2013 | Robust low-complexity multichannel equalization for dereverberationabstractMultichannel equalization of acoustic impulse responses (AIRs) is an important approach for dereverberation. Since AIRs are inevitably estimated with system identification error (SIE), it is necessary to develop equalization designs that are robust to such SIE, in order for dereverberation processing to be beneficial. We present here a novel subband equalizer employing the relaxed multichannel least squares (RMCLS) algorithm in each subband. We show that this new structure brings improved performance in dereverberation as well as a reduction in computational load by up to a factor of more than 90 in our experiments. We then develop a novel controller for the dereverberation processing in subbands that guarantees robustness to even very severe SIEs by backing off dereverberation in any subband with excessively high levels of SIEs. Felicia Lim, Patrick A. Naylor |
ICASSP | 2 |
| 2013 | Blind System Identification Using Sparse Learning for TDOA Estimation of Room ReflectionsabstractLocalization of early room reflections can be achieved by estimating the time-differences-of-arrival (TDOAs) of reflected waves between elements of a microphone array. For an unknown source, we propose to apply sparse blind system identification (BSI) methods to identify the acoustic impulse responses, from which the TDOAs of temporally sparse reflections are estimated. The proposed time- and frequency-domain adaptive algorithms based on crossrelation formulation are regularized by incorporating an l1-norm sparseness constraint, which is realized using a split Bregman method. These algorithms are shown to outperform standard crossrelation-based BSI techniques when estimating TDOAs of reflections in the presence of background noise. Konrad Kowalczyk, Emanuël A. P. Habets, Walter Kellermann, Patrick A. Naylor |
IEEE Signal Process. Lett. | 4 |
| 2013 | TDOA-Based Speed of Sound Estimation for Air Temperature and Room Geometry InferenceabstractSpatially distributed acoustic sensors find increasingly many new applications in speech-based human-machine interfaces. One well researched topic is the localization of sound sources from Time Differences Of Arrival (TDOAs) measurements. Typically, the propagation speed of sound is considered a known constant. However due to temperature variations its value is known only up to some uncertainty. This paper exploits TDOA-based localization techniques in order to estimate accurately the actual speed of sound. Experimental results using both simulated and real data demonstrate the feasibility of the proposed method. Furthermore, the practical validation of this work considers two distinct experiments that are aimed at inferring information about enclosed sound fields. The first experiment concerns the calculation of the air temperature from the estimated speed of sound. The second experiment highlights the effects of temperature variations on the inference of the physical location of reflective boundaries of the acoustic enclosure. In the latter case it is shown that the position estimates of the reflective surfaces in a room can be improved when the correct propagation speed is first estimated using this method. Paolo Annibale, Jason Filos, Patrick A. Naylor, Rudolf Rabenstein |
IEEE Trans. Speech Audio Process. | 3 |
| 2013 | Blind Channel Magnitude Response Estimation in Speech Using Spectrum ClassificationabstractWe present an algorithm for blind estimation of the magnitude response of an acoustic channel from single microphone observations of a speech signal. The algorithm employs channel robust RASTA filtered Mel-frequency cepstral coefficients as features to train a Gaussian mixture model based classifier and average clean speech spectra are associated with each mixture; these are then used to blindly estimate the acoustic channel magnitude response from speech that has undergone spectral modification due to the channel. Experimental results using a variety of simulated and measured acoustic channels and additive babble noise, car noise and white Gaussian noise are presented. The results demonstrate that the proposed method is able to estimate a variety of channel magnitude responses to within an Itakura distance of dI ≤0.5 for SNR ≥10 dB. Nikolay D. Gaubitch, Mike Brookes, Patrick A. Naylor |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Non intrusive codec identification algorithmabstractWe present a non-intrusive data driven method for codec detection and identification in the presence of background noise. The method uses a number of speech features which are then used to train a CART classifier. We demonstrate the performance of the method using several different noise types over a wide range of SNRs. Our results show that we can identify a codec and its bit rate to an accuracy of 92% and we are able to detect the presence of a codec with an accuracy of 97% at -5 dB SNR. Dushyant Sharma, Patrick A. Naylor, Nikolay D. Gaubitch, Mike Brookes |
ICASSP | 2 |
| 2012 | An insight into common filtering in noisy SIMO blind system identificationabstractThe effect of additive sensor noise on single-input-multiple-output (SIMO) blind system identification (BSI) algorithms based upon cross-relation (CR) error is investigated. Previous studies have shown that additive noise in the observed signal results in systems comprising the true estimated channels convolved with an erroneous ‘common filter’, and additionally that identification and removal of this filter significantly improves estimation error. However, the source of the common filter remained an open question. This paper explains the common filter through a first-order perturbation analysis of the CR matrix, showing that it be estimated from the perturbation and the eigenvectors of the noiseless CR matrix. The analysis given in this paper provides a new insight into the effect of noise on SIMO BSI algorithms and forms the first step towards an overall noise robust solution. Mark R. P. Thomas, Nikolay D. Gaubitch, Emanuël A. P. Habets, Patrick A. Naylor |
ICASSP | 4 |
| 2012 | Descriptive Vocabulary Development for Degraded SpeechabstractThis paper presents the development of a compact vocabulary for describing the audible characteristics of degraded speech. An experiment was conducted with 51 English-speaking subjects who were tasked with assigning one of a list of given text descriptors to 220 degradation conditions. Exploratory data analysis using hierarchical clustering resulted in a compact vocabulary of 10 classes, which was further validated by a bootstrap cluster analysis. Dushyant Sharma, Gaston Hilkhuysen, Patrick A. Naylor, Nikolay D. Gaubitch, Mark A. Huckvale, Mike Brookes |
INTERSPEECH | 3 |
| 2012 | Data-driven voice source waveform analysis and synthesis
Jón Guðnason, Mark R. P. Thomas, Daniel P. W. Ellis, Patrick A. Naylor |
Speech Commun. | 4 |
| 2012 | Inference of Room Geometry From Acoustic Impulse ResponsesabstractAcoustic scene reconstruction is a process that aims to infer characteristics of the environment from acoustic measurements. We investigate the problem of locating planar reflectors in rooms, such as walls and furniture, from signals obtained using distributed microphones. Specifically, localization of multiple two- dimensional (2-D) reflectors is achieved by estimation of the time of arrival (TOA) of reflected signals by analysis of acoustic impulse responses (AIRs). The estimated TOAs are converted into elliptical constraints about the location of the line reflector, which is then localized by combining multiple constraints. When multiple walls are present in the acoustic scene, an ambiguity problem arises, which we show can be addressed using the Hough transform. Additionally, the Hough transform significantly improves the robustness of the estimation for noisy measurements. The proposed approach is evaluated using simulated rooms under a variety of different controlled conditions where the floor and ceiling are perfectly absorbing. Results using AIRs measured in a real environment are also given. Additionally, results showing the robustness to additive noise in the TOA information are presented, with particular reference to the improvement achieved through the use of the Hough transform. Fabio Antonacci, Jason Filos, Mark R. P. Thomas, Emanuël A. P. Habets, Augusto Sarti, Patrick A. Naylor, Stefano Tubaro |
IEEE Trans. Speech Audio Process. | 6 |
| 2012 | Detection of Glottal Closure Instants From Speech Signals: A Quantitative ReviewabstractThe pseudo-periodicity of voiced speech can be exploited in several speech processing applications. This requires however that the precise locations of the glottal closure instants (GCIs) are available. The focus of this paper is the evaluation of automatic methods for the detection of GCIs directly from the speech waveform. Five state-of-the-art GCI detection algorithms are compared using six different databases with contemporaneous electroglottographic recordings as ground truth, and containing many hours of speech by multiple speakers. The five techniques compared are the Hilbert Envelope-based detection (HE), the Zero Frequency Resonator-based method (ZFR), the Dynamic Programming Phase Slope Algorithm (DYPSA), the Speech Event Detection using the Residual Excitation And a Mean-based Signal (SEDREAMS) and the Yet Another GCI Algorithm (YAGA). The efficacy of these methods is first evaluated on clean speech, both in terms of reliabililty and accuracy. Their robustness to additive noise and to reverberation is also assessed. A further contribution of the paper is the evaluation of their performance on a concrete application of speech processing: the causal-anticausal decomposition of speech. It is shown that for clean speech, SEDREAMS and YAGA are the best performing techniques, both in terms of identification rate and accuracy. ZFR and SEDREAMS also show a superior robustness to additive noise and reverberation. Thomas Drugman, Mark R. P. Thomas, Jón Guðnason, Patrick A. Naylor, Thierry Dutoit |
IEEE Trans. Speech Audio Process. | 4 |
| 2012 | A Speech Distortion and Interference Rejection Constraint BeamformerabstractSignals captured by a set of microphones in a speech communication system are mixtures of desired and undesired signals and ambient noise. Existing beamformers can be divided into those that preserve or distort the desired signal. Beamformers that preserve the desired signal are, for example, the linearly constrained minimum variance (LCMV) beamformer that is supposed, ideally, to reject the undesired signal and reduce the ambient noise power, and the minimum variance distortionless response (MVDR) beamformer that reduces the interference-plus-noise power. The multichannel Wiener filter, on the other hand, reduces the interference-plus-noise power without preserving the desired signal. In this paper, a speech distortion and interference rejection constraint (SDIRC) beamformer is derived that minimizes the ambient noise power subject to specific constraints that allow a tradeoff between speech distortion and interference-plus-noise reduction on the one hand, and undesired signal and ambient noise reductions on the other hand. Closed-form expressions for the performance measures of the SDIRC beamformer are derived and the relations to the aforementioned beamformers are derived. The performance evaluation demonstrates the tradeoffs that can be made using the SDIRC beamformer. Emanuël A. P. Habets, Jacob Benesty, Patrick A. Naylor |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | A Forced Spectral Diversity Algorithm for Speech Dereverberation in the Presence of Near-Common ZerosabstractBlind identification of single-input multiple-output (SIMO) systems is not normally possible if common zeros exist in the channels. Studies of measured acoustic SIMO systems show that near-common zeros occur in such systems as encountered in the speech dereverberation task. We therefore introduce a method to add additional diversity to the SIMO system to be identified which we term forced spectral diversity (FSD) and we show that its use leads to an identification-equalization approach that gives improved dereverberation. As part of this work, we show the link between channel diversity and the effect of common zeros. We also define and discuss in more detail the concept and impact of near-common zeros. The proposed algorithm is presented specifically for a two-channel system where such near-common zeros exist. Andy W. H. Khong, Patrick A. Naylor |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Estimation of Glottal Closing and Opening Instants in Voiced Speech Using the YAGA AlgorithmabstractAccurate estimation of glottal closing instants (GCIs) and opening instants (GOIs) is important for speech processing applications that benefit from glottal-synchronous processing including pitch tracking, prosodic speech modification, speech dereverberation, synthesis and study of pathological voice. We propose the Yet Another GCI/GOI Algorithm (YAGA) to detect GCIs from speech signals by employing multiscale analysis, the group delay function, andN-best dynamic programming. A novel GOI detector based upon the consistency of the candidates' closed quotients relative to the estimated GCIs is also presented. Particular attention is paid to the precise definition of the glottal closed phase, which we define as the analysis interval that produces minimum deviation from an all-pole model of the speech signal with closed-phase linear prediction (LP). A reference algorithm analyzing both electroglottograph (EGG) and speech signals is described for evaluation of the proposed speech-based algorithm. In addition to the development of a GCI/GOI detector, an important outcome of this work is in demonstrating that GOIs derived from the EGG signal are not necessarily well-suited to closed-phase LP analysis. Evaluation of YAGA against the APLAWD and SAM databases show that GCI identification rates of up to 99.3% can be achieved with an accuracy of 0.3 ms and GOI detection can be achieved equally reliably with an accuracy of 0.5 ms. Mark R. P. Thomas, Jón Guðnason, Patrick A. Naylor |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | Simulating room impulse responses for spherical microphone arraysabstractA method is proposed for simulating the sound pressure signals on a spherical microphone array in a reverberant enclosure. The method employs spherical harmonic decomposition and takes into account scattering from a solid sphere. An analysis shows that the error in the decomposition can be made arbitrarily small given a sufficient number of spherical harmonics. Daniel P. Jarrett, Emanuël A. P. Habets, Mark R. P. Thomas, Patrick A. Naylor |
ICASSP | 4 |
| 2011 | A proportionate adaptive algorithm with variable partitioned block length for acoustic echo cancellationabstractDue to the nature of an acoustic enclosure, the early part (i.e., direct path and early reflections) of the acoustic echo path is often sparse while the late reverberant part of the acoustic path is normally dispersive. In order to account for this structure within the acoustic impulse response when performing acoustic echo cancellation, we propose an adaptive filter that consists of two time-domain partition blocks, with adaptive block partitioning, such that different adaptive algorithms can be used for each block. Specifically, the improved proportionate normalized least-mean-square (IPNLMS) algorithm is used. Simulation results show that the proposed variable length partitioned block IPNLMS (VLPB-IPNLMS) algorithm works well in both sparse and dispersive circumstances and in practical applications involving time-varying systems. Pradeep Loganathan, Emanuël A. P. Habets, Patrick A. Naylor |
ICASSP | 3 |
| 2010 | An online quasi-Newton algorithm for blind SIMO identificationabstractIn the last decade various time- and frequency-domain algorithms were derived to blindly identify acoustic systems. One of these algorithms is the multichannel Newton (MCN) algorithm, which is also the basis of the well known normalized multichannel frequency-domain least-mean-square (NMCFLMS) algorithm. A major drawback of the MCN is that it requires the computation and inversion of a Hessian matrix, which involves extensive computation making it unsuitable for online applications. In this paper, we therefore derive and investigate an efficient online multichannel quasi-Newton (MCQN) algorithm that updates the inverse of the Hessian by analyzing successive gradient vectors. The new MCQN is shown to exhibit similar performance to MCN but with much reduced complexity. Emanuël A. P. Habets, Patrick A. Naylor |
ICASSP | 2 |
| 2010 | Performance analysis of IPNLMS for identification of time-varying systemsabstractThe tracking performance of adaptive filters is crucially important in practical applications involving time-varying systems. We present an analysis of the tracking performance for IPNLMS, one of the best known and best performing algorithms originally targeted at sparse system identification. We then validate our analytic results in practical simulations for echo cancellation for sparse and dispersive time-varying unknown echo path systems. These results show the analysis to be highly accurate in all the cases studied. Pradeep Loganathan, Emanuël A. P. Habets, Patrick A. Naylor |
ICASSP | 3 |
| 2010 | Voice source estimation for artificial bandwidth extension of telephone speechabstractArtificial bandwidth extension (ABWE) of speech signals aims to estimate wideband speech (50 Hz – 7 kHz) from narrowband signals (300 Hz – 3.4 kHz). Applying the source-filter model of speech, many existing algorithms estimate vocal tract filter parameters independently of the source signal. However, many current methods for extending the narrowband voice source signal are limited to straightforward signal processing techniques which are only effective for high-band estimation. This paper presents a method for ABWE that employs novel data-driven modelling and an existing spectral mirroring technique to estimate the wideband source signal in both the high and low extension bands. A state-of-the-art Hidden Markov Model-based estimator evaluates the temporal and spectral envelopes in the missing frequency bands, with which the ABWE speech signal is synthesized. Informal listening tests comparing two existing source estimation techniques and two permutations of the proposed approach show an improvement in the perceived bandwidth of speech signals, in particular towards low frequencies. Subjective tests on the same data show a preference for the proposed techniques over the existing methods under test. Mark R. P. Thomas, Jón Guðnason, Patrick A. Naylor, Bernd Geiser, Peter Vary |
ICASSP | 3 |
| 2010 | A System-Identification-Error-Robust Method for equalization of multichannel acoustic systemsabstractIn hands-free communications, speech received by a microphone is distorted by room reverberation that can reduce the intelligibility of speech. An approach to dereverberation is firstly to estimate the impulse responses of the acoustic channels between the speaker and the microphones and secondly to design a multichannel equalization system based on the estimated impulse responses. Traditional equalization techniques are designed without the consideration of estimation errors that are commonly introduced by the system identification process. In this work, a System-Identification-Error-Robust Equalization Method (SIEREM) for the equalization of multichannel room acoustic systems is presented. Experimental results for dereverberation using SIEREM applied to estimates of single-input multiple-output acoustic systems with known level of estimation errors show that the proposed equalization design significantly outperforms existing methods in the presence of both synthetic and real system identification errors. Wancheng Zhang, Emanuël A. P. Habets, Patrick A. Naylor |
ICASSP | 3 |
| 2010 | Introduction to the Special Issue on Processing Reverberant Speech: Methodologies and ApplicationsabstractThe 17 papers in this special issue focus on the methodologies and applications of processing reverberant speech. The issue highlights some major aspects of the recent progress in the field. Tomohiro Nakatani, Walter Kellermann, Patrick A. Naylor, Masato Miyoshi, Biing-Hwang Juang |
IEEE Trans. Speech Audio Process. | 3 |
| 2009 | Blind system identification for speech dereverberation with Forced Spectral DiversityabstractThe common zeros problem for blind system identification (BSI) is well known. It degrades the performance of classic BSI algorithms and therefore imposes the limit on the performance of subsequent speech dereverberation. The effect of near-common zeros has recently been studied in terms of channel diversity and the degradation in performance of BSI and multichannel equalization algorithms has been shown. We now introduce a novel approach to improve channel diversity which we refer to as Forced Spectral Diversity (FSD). The FSD concept uses a combination of spectral shaping filters and effective channel undermodelling. Simulation results show that the proposed approach achieves improved performance with reduced complexity for multichannel BSI in a room acoustics example. Andy W. H. Khong, Patrick A. Naylor |
ICASSP | 3 |
| 2009 | Data-driven voice soruce waveform modellingabstractThis paper presents a data-driven approach to the modelling of voice source waveforms. The voice source is a signal that is estimated by inverse-filtering speech signals with an estimate of the vocal tract filter. It is used in speech analysis, synthesis, recognition and coding to decompose a speech signal into its source and vocal tract filter components. Existing approaches parameterize the voice source signal with physically- or mathematically-motivated models. Though the models are well-defined, estimation of their parameters is not well understood and few are capable of reproducing the large variety of voice source waveforms. Here we present a data-driven approach to classify types of voice source waveforms based upon their mel frequency cepstrum coefficients with Gaussian mixture modelling. A set of ldquoprototyperdquo waveform classes is derived from a weighted average of voice source cycles from real data. An unknown speech signal is then decomposed into its prototype components and resynthesized. Results indicate that with sixteen voice source classes, low resynthesis errors can be achieved. Mark R. P. Thomas, Jón Guðnason, Patrick A. Naylor |
ICASSP | 3 |
| 2009 | Voice source waveform analysis and synthesis using principal component analysis and Gaussian mixture modellingabstractThe paper presents a voice source waveform modeling techniques based on principal component analysis (PCA) and Gaussian mixture modeling (GMM). The voice source is obtained by inverse-filtering speech with the estimated vocal tract filter. This decomposition is useful in speech analysis, synthesis, recognition and coding. Here, a data-driven approach is presented for signal decomposition and classification based on the principal components of the voice source. The principal components are analyzed and the 'prototype' voice source signals corresponding to the Gaussian mixture means are examined. We show how an unknown signal can be decomposed into its components and/or prototypes and resynthesized. We show how the techniques are suited for both low bitrate or high quality analysis/synthesis schemes. Jón Guðnason, Mark R. P. Thomas, Patrick A. Naylor, Daniel P. W. Ellis |
INTERSPEECH | 3 |
| 2009 | Equalization of Multichannel Acoustic Systems in Oversampled SubbandsabstractEqualization of room transfer functions (RTFs) is an important topic with several applications in acoustic signal processing. RTFs are often modeled as finite-impulse response filters, characterized by orders of thousands of taps and non-minimum phase. In practice, only approximate estimates of the actual RTFs are available due to measurement noise, limited estimation accuracy, and temporal variation of source-receiver position. These issues make equalization a difficult problem. In this paper, we discuss multichannel equalization with focus on inexact RTF estimates. We present a multichannel method for the equalization filter design utilizing decimated and oversampled subbands, where the full-band acoustic impulse response is decomposed into equivalent subband filters prior to equalization. This technique is not only more computationally efficient but also more robust to impulse response inaccuracies compared with the full-band counterpart. Nikolay D. Gaubitch, Patrick A. Naylor |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | A Class of Sparseness-Controlled Algorithms for Echo CancellationabstractIn the context of acoustic echo cancellation (AEC), it is shown that the level of sparseness in acoustic impulse responses can vary greatly in a mobile environment. When the response is strongly sparse, convergence of conventional approaches is poor. Drawing on techniques originally developed for network echo cancellation (NEC), we propose a class of AEC algorithms that can not only work well in both sparse and dispersive circumstances, but also adapt dynamically to the level of sparseness using a new sparseness-controlled approach. Simulation results, using white Gaussian noise (WGN) and speech input signals, show improved performance over existing methods. The proposed algorithms achieve these improvement with only a modest increase in computational complexity. Pradeep Loganathan, Andy W. H. Khong, Patrick A. Naylor |
IEEE Trans. Speech Audio Process. | 3 |
| 2009 | The SIGMA Algorithm: A Glottal Activity Detector for Electroglottographic SignalsabstractAccurate estimation of glottal closure instants (GCIs) and opening instants (GOIs) is important for speech processing applications that benefit from glottal-synchronous processing. The majority of existing approaches detect GCIs by comparing the differentiated EGG signal to a threshold and are able to provide accurate results during voiced speech. More recent algorithms use a similar approach across multiple dyadic scales using the stationary wavelet transform. All existing approaches are however prone to errors around the transition regions at the end of voiced segments of speech. This paper describes a new method for EGG-based glottal activity detection which exhibits high accuracy over the entirety of voiced segments, including, in particular, the transition regions, thereby giving significant improvement over existing methods. Following a stationary wavelet transform-based preprocessor, detection of excitation due to glottal closure is performed using a group delay function and then true and false detections are discriminated by Gaussian mixture modeling. GOI detection involves additional processing using the estimated GCIs. The main purpose of our algorithm is to provide a ground-truth for GCIs and GOIs. This is essential in order to evaluate algorithms that estimate GCIs and GOIs from the speech signal only, and is also of high value in the analysis of pathological speech where knowledge of GCIs and GOIs is often needed. We compare our algorithm with two previous algorithms against a hand-labeled database. Evaluation has shown an average GCI hit rate of 99.47% and GOI of 99.35%, compared to 96.08 and 92.54 for the best-performing existing algorithm. Mark R. P. Thomas, Patrick A. Naylor |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Temporal selective dereverberation of noisy speech using one microphoneabstractReverberant speech can be described as sounding distant with noticeable coloration and echo. These detrimental perceptual effects are caused by early and late reflections, respectively, and reduces the fidelity and intelligibility of speech. It is well-known that the echo density of the reflections increases with time. Therefore, the temporal structure of early and late reflections differs. In this paper, we combine two different dereverberation techniques that were recently developed to suppress early and late reverberation separately. First, late reverberation is suppressed using a spectral processing technique that is based on a statistical reverberation model. Secondly, early reverberation and residual late reverberation are suppressed using a linear prediction (LP) residual processing technique. In addition, an objective measure based on the kurtosis of the LP residual is proposed to measure the coloration caused by early reflections. Experimental results demonstrate the beneficial use of the new single microphone system that reduces echo and coloration with little speech distortion. Emanuël A. P. Habets, Nikolay D. Gaubitch, Patrick A. Naylor |
ICASSP | 3 |
| 2008 | A lowcomplexity fast converging partial update adaptive algorithm employing variable step-size for acoustic echo cancellationabstractPartial update adaptive algorithms have been proposed as a means of reducing complexity for adaptive filtering. The MMax tap-selection is one of the most popular tap-selection algorithms. It is well known that the performance of such partial update algorithm reduces with reducing number of filter coefficients selected for adaptation. We propose a low complexity and fast converging adaptive algorithm that exploits the MMax tap-selection. We achieve fast convergence with low complexity by deriving a variable step-size for the MMax normalized least-mean-square (MMax-NLMS) algorithm using its mean square deviation. Simulation results verify that the proposed algorithm achieves higher rate of convergence with lower computational complexity compared to the NLMS algorithm. Andy W. H. Khong, Woon-Seng Gan, Patrick A. Naylor, Mike Brookes |
ICASSP | 3 |
| 2008 | Frequency domain selective tap adaptive algorithms for sparse system identificationabstractWe propose a new low complexity and fast converging frequency-domain adaptive algorithm for sparse system identification. This is achieved by exploiting the MMax and SP tap-selection criteria for complexity reduction and fast convergence respectively. We incorporate these tap-selection techniques into the multi-delay filtering (MDF) algorithm in order to reduce the delay inherent in frequency-domain algorithms. We illustrate two such approaches and discuss the tradeoff between convergence performance and computational complexity for these approaches. Simulation results show an improvement in convergence rate for the proposed algorithm over MDF with reduced complexity. The proposed algorithm achieves a convergence performance close to that of the recently proposed but substantially more complex improved proportionate MDF algorithm. Andy W. H. Khong, Milos Doroslovacki, Patrick A. Naylor |
ICASSP | 4 |
| 2008 | Algorithms for identifying clusters of near-common zeros in multichannel blind system identification and equalizationabstractBlind system identification (BSI) and equalization algorithms have been applied to multichannel systems with high order such as found in acoustic impulse responses. Studies on the performance of such algorithms in the presence of near-common zeros have been limited to low order systems. In this work, we propose two high order clustering algorithms which efficiently extract clusters of near-common zeros within a specified pairwise distance in the z-plane. Using these algorithms, we then quantify the number of common zeros that exist in acoustic systems. In addition, we show how these algorithms can be applied to study of BSI and equalization algorithms in the presence of near-common zeros for such acoustic systems. Andy W. H. Khong, Patrick A. Naylor |
ICASSP | 3 |
| 2008 | Blind estimation of reverberation time based on the distribution of signal decay ratesabstractThe reverberation time is one of the most prominent acoustic characteristics of an enclosure. Its value can be used to predict speech intelligibility, and is used by speech enhancement techniques to suppress reverberation. The reverberation time is usually obtained by analysing the decay rate of (i) the energy decay curve that is observed when a noise source is switched off, and (ii) the energy decay curve of the room impulse response. Estimating the reverberation time using only the observed reverberant speech signal, i.e., blind estimation, is required for speech evaluation and enhancement techniques. Recently, (semi) blind methods have been developed. Unfortunately, these methods are not very accurate when the source consists of a human speaker, and unnatural speech pauses are required to detect and/or track the decay. In this paper we extract and analyse the decay rate of the energy envelope blindly from the observed reverberation speech signal in the short-time Fourier transform domain. We develop a method to estimate the reverberation time using a property of the distribution of the decay rates. Experimental results using simulated and real reverberant speech signals demonstrate the performance of the new method. Jimi Yung-Chuan Wen, Emanuël A. P. Habets, Patrick A. Naylor |
ICASSP | 3 |
| 2008 | Computationally efficient equalization of room impulse responses robust to system estimation errorsabstractEqualization techniques for room impulse responses (RIRs) are important in acoustic signal processing applications such as speech dereverberation. In practice, only approximate estimates of the RIRs are available and the inverse filters designed from these estimates may cause significant distortion in the equalized signal. A second issue is that existing equalizer design algorithms are computationally expensive. We here propose regularized subband equalizer design algorithm. Both the computational complexity and the robustness of the equalizer design to system estimation errors are improved. An analysis of the computational complexity and simulation examples are provided to support our study. Wancheng Zhang, Nikolay D. Gaubitch, Patrick A. Naylor |
ICASSP | 3 |
| 2008 | Multimicrophone speech dereverberation using spatiotemporal and spectral processingabstractSpeech signals acquired in a reverberant room with microphones positioned at a distance from the talker are degraded in quality due to reverberation and measurement noise. Therefore, enhancement of reverberant speech is important in hands-free telecommunications applications. The perceptual effects of reverberation can be linked to the room impulse response (RIR) between the talker and the microphone and are characterized by: (i) colouration, due to the strong early reflections and (ii) a distant 'echoey' quality due to the decaying tail of the RIR. Accordingly, we present a two-stage multimicrophone method for speech dereverberation. First, spatiotemporal averaging is performed on the linear prediction residual, which primarily reduces the effects of the early reflections. Secondly, a spectral subtraction method is employed to reduce late reverberation. Simulation results with measured RIRs and additive white Gaussian noise illustrate the performance of this method and show that the combined approach performs better than each of the two stages individually. Nikolay D. Gaubitch, Emanuël A. P. Habets, Patrick A. Naylor |
ISCAS | 3 |
| 2008 | A Class of Frobenius Norm-Based Algorithms Using Penalty Term and Natural Gradient for Blind Signal SeparationabstractWe consider the blind signal separation (BSS) problem of instantaneous mixtures using penalty term and natural gradient. A class of Frobenius norm-based algorithms consisting of the offline/block processing (BP), online processing (OP) algorithms, and their normalized versions is proposed for separating nonstationary and nonwhite signals. The BP and OP algorithms, respectively, suitable for blind separation with offline and online data, are derived by using the nonstationarity and nonwhiteness of signals and the natural gradient method in conjunction with an appropriate penalty term. Associated with almost all algorithms employing a gradient method is a gradient noise problem. We thus develop, from BP and OP, their normalized versions in which the update of an unknown demixing matrix is based on the minimal disturbance principle. We show that the resulting updates are in the same direction as those of the original algorithms but with a scaling factor whose upper bound is unity. Algorithms using the nonstationarity and nonwhiteness properties have been proposed before but, due to the use of logarithms in their derivation, they are not capable of separating signals that are not persistently active and require regularization parameters to mitigate the problem. In this paper, the superior performance of the proposed algorithms to the previously proposed logarithm-based algorithms with and without regularization when separating nonpersistently active source signals is presented through some illustrative numerical experiments. Uttachai Manmontri, Patrick A. Naylor |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Misalignment Performance of Selective Tap Adaptive Algorithms for System Identification of Time-Varying Unknown SystemsabstractSelective tap algorithms have been proposed as a means of reducing complexity for adaptive filtering. MMax tap selection has been employed in many algorithms due to its straightforward implementation. This paper formulates the analysis of two MMax-based algorithms under time-varying unknown system conditions as are often found in practical applications. The steady-state misalignment for the MMax normalized least mean square and the MMax recursive least squares algorithms are derived and their performance is compared to that of their respective full-update algorithms. The tradeoff between computational complexity and misalignment performance is also shown for the MMax normalized least mean square case. Patrick A. Naylor, Andy W. H. Khong, Mike Brookes |
ICASSP (1) | 1 |
| 2007 | Selective-Tap Adaptive Filtering With Performance Analysis for Identification of Time-Varying SystemsabstractSelective-tap algorithms employing the MMax tap selection criterion were originally proposed for low-complexity adaptive filtering. The concept has recently been extended to multichannel adaptive filtering and applied to stereophonic acoustic echo cancellation. This paper first briefly reviews least mean square versions of MMax selective-tap adaptive filtering and then introduces new recursive least squares and affine projection MMax algorithms. We subsequently formulate an analysis of the MMax algorithms for time-varying system identification by modeling the unknown system using a modified Markov process. Analytical results are derived for the tracking performance of MMax selective tap algorithms for normalized least mean square, recursive least squares, and affine projection algorithms. Simulation results are shown to verify the analysis. Andy W. H. Khong, Patrick A. Naylor |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Estimation of Glottal Closure Instants in Voiced Speech Using the DYPSA AlgorithmabstractWe present the Dynamic Programming Projected Phase-Slope Algorithm (DYPSA) for automatic estimation of glottal closure instants (GCIs) in voiced speech. Accurate estimation of GCIs is an important tool that can be applied to a wide range of speech processing tasks including speech analysis, synthesis and coding. DYPSA is automatic and operates using the speech signal alone without the need for an EGG signal. The algorithm employs the phase-slope function and a novel phase-slope projection technique for estimating GCI candidates from the speech signal. The most likely candidates are then selected using a dynamic programming technique to minimize a cost function that we define. We review and evaluate three existing methods of GCI estimation and compare the new DYPSA algorithm to them. Results are presented for the APLAWD and SAM databases for which 95.7% and 93.1% of GCIs are correctly identified Patrick A. Naylor, Anastasis Kounoudes, Jón Guðnason, Mike Brookes |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Proportionate Frequency Domain Adaptive Algorithms for Blind Channel IdentificationabstractWe present fast-converging adaptive blind channel identification algorithms for acoustic room impulse responses. These new algorithms exploit the fast-convergence of the improved proportionate normalized least-mean-square (IPNLMS) algorithm and address the problem of delay inherent in frequency domain algorithms by employing the multi-delay filter (MDF) structure. Simulation results for both speech and white Gaussian noise show that the proposed algorithms outperform current frequency domain blind channel estimation algorithms Rehan Ahmad, Andy W. H. Khong, Patrick A. Naylor |
ICASSP (5) | 3 |
| 2006 | Noise Robust Adaptive Blind Channel Identification Using Spectral ConstraintsabstractA class of adaptive blind channel identification algorithms were proposed recently and were demonstrated to be able to successfully identify various types of channels when the observed signals are free from significant levels of measurement noise. In this paper, we provide a study of the effects of noise on these algorithms and show that they misconverge even at moderate values of SNR. We introduce a spectral constraint into the adaptation rule and show that the robustness to noise can be considerably improved. Simulation results are presented for the new algorithm, which demonstrate a significant performance improvement in terms of normalized projection misalignment Nikolay D. Gaubitch, Md. Kamrul Hasan 0001, Patrick A. Naylor |
ICASSP (5) | 3 |
| 2006 | Effect of Interchannel Coherence on Conditioning and Misalignment Performance for Stereo Acoustic ECHO CancellationabstractIt is well known that the performance in terms of misalignment of adaptive algorithms, in general, is dependent on the conditioning of the input signal covariance matrix. For two-channel (stereophonic) adaptive algorithms, this performance is further degraded by the high interchannel coherence between the two input signals. In this paper, we establish the relationship between interchannel coherence of the two input signals and condition of the corre- corresonding covariance matrix for stereo acoustic echo cancellation application. We further show how this relationship affects the misalignment performance of a two-channel frequency-domain adaptive algorithm. We provide simulation results for both WGN and speech input to verify our mathematical analysis. Andy W. H. Khong, Jacob Benesty, Patrick A. Naylor |
ICASSP (5) | 3 |
| 2006 | Blind Signal Separation Using a Criterion Based on Principle of Minimal DisturbanceabstractThe concept underlying most on-line gradient-based algorithms for blind signal separation (BSS) is that the unknown demixing matrix is adjusted with an appropriate step-size in the direction of the gradient computed at each sample instant. Associated with these algorithms is a gradient noise problem. In this paper, we develop, from the on-line processing (OP) algorithm derived using the nonstationarity and nonwhiteness properties, a normalized algorithm in which the update of the demixing matrix is based on the minimal disturbance principle. We show that the resulting updates are in the same direction as those of the original algorithm but with a scaling factor whose upper bound is unity. We evaluate the convergence speed and robustness to gradient noise of the new algorithm Uttachai Manmontri, Patrick A. Naylor |
ICASSP (5) | 2 |
| 2006 | Adaptive algorithms for sparse echo cancellation
Patrick A. Naylor, Jingjing Cui 0005, Mike Brookes |
Signal Process. | 1 |
| 2006 | Generalized Optimal Step-Size for Blind Multichannel LMS System IdentificationabstractThe choice of step-size in adaptive blind channel identification using the multichannel least mean squares (MCLMS) algorithm is critical and controls its convergence rate, stability, and sensitivity to noise. In this letter, we derive the expression for an optimal step-size in the Wiener sense and investigate its properties. An implementation technique for the Wiener solution of the self-adaptive step-size is presented, and it is shown that significant performance improvements are obtained compared to existing approaches in the presence of noise Nikolay D. Gaubitch, Md. Kamrul Hasan 0001, Patrick A. Naylor |
IEEE Signal Process. Lett. | 3 |
| 2006 | Stereophonic acoustic echo cancellation: analysis of the misalignment in the frequency domainabstractThe performance in terms of misalignment of adaptive algorithms, in general, is dependent on the conditioning of the input signal covariance matrix. The performance of two-channel adaptive algorithms is further degraded by the high interchannel coherence between the two input signals. In this letter, we establish the relationship between interchannel coherence of the two input signals and condition of the corresponding covariance matrix for stereo acoustic echo cancellation application. We show how this relationship affects the misalignment of a frequency-domain adaptive algorithm. We provide simulation results for both white Gaussian noise and speech input to verify our mathematical analysis. Andy W. H. Khong, Jacob Benesty, Patrick A. Naylor |
IEEE Signal Process. Lett. | 3 |
| 2006 | A quantitative assessment of group delay methods for identifying glottal closures in voiced speechabstractMeasures based on the group delay of the LPC residual have been used by a number of authors to identify the time instants of glottal closure in voiced speech. In this paper, we discuss the theoretical properties of three such measures and we also present a new measure having useful properties. We give a quantitative assessment of each measure's ability to detect glottal closure instants evaluated using a speech database that includes a direct measurement of glottal activity from a Laryngograph/EGG signal. We find that when using a fixed-length analysis window, the best measures can detect the instant of glottal closure in 97% of larynx cycles with a standard deviation of 0.6 ms and that in 9% of these cycles an additional excitation instant is found that normally corresponds to glottal opening. We show that some improvement in detection rate may be obtained if the analysis window length is adapted to the speech pitch. If the measures are applied to the preemphasized speech instead of to the LPC residual, we find that the timing accuracy worsens but the detection rate improves slightly. We assess the computational cost of evaluating the measures and we present new recursive algorithms that give a substantial reduction in computation in all cases. Mike Brookes, Patrick A. Naylor, Jón Guðnason |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Stereophonic acoustic echo cancellation employing selective-tap adaptive algorithmsabstractStereophonic acoustic echo cancellation has generated much interest in recent years due to the nonuniqueness and misalignment problems that are caused by the strong interchannel signal coherence. In this paper, we introduce a novel adaptive filtering approach to reduce interchannel coherence which is based on a selective-tap updating procedure. This tap-selection technique is then applied to the normalized least-mean-square, affine projection and recursive least squares algorithms for stereophonic acoustic echo cancellation. Simulation results for the proposed algorithms have shown a significant improvement in convergence rate compared with existing techniques. Andy W. H. Khong, Patrick A. Naylor |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Performance analysis of dynamic acoustic source separation in reverberant roomsabstractWe study the effect of reverberation and source movement on the performance of blind source separation and deconvolution (BSSD) algorithms. Using the model of statistical room acoustics we derive theoretical performance measures for a class of unmixing algorithms when these are used in a reverberant room. We specifically investigate the cases: 1) where separation of only direct paths is performed and 2) the case where unmixing of the full reverberant paths is attempted. We develop closed-form performance measures that are dependent on the geometry used and the chosen unmixing system. Using these measures allows us to draw general conclusions on the robustness to source movement of typical BSSD algorithms. Results indicate that performance of systems that show very good separation in static reverberant environments is significantly reduced when sources move, with performance degrading to that of simple direct-path separation Fotios Talantzis, Darren B. Ward, Patrick A. Naylor |
IEEE Trans. Speech Audio Process. | 3 |
| 2005 | A family of selective-tap algorithms for stereo acoustic echo cancellationabstractThe use of adaptive filters employing tap-selection for stereophonic acoustic echo cancellation (SAEC) is investigated. We propose to employ subsampling of the tap-input vector, that is intrinsic to partial update schemes, to improve the conditioning of the tap-input autocorrelation matrix hence improving convergence. We investigate the effect of MMax tap-selection on the convergence rate for the single channel case by proposing a new measure which is then used as an optimization parameter in the development of our tap-selection scheme in the two channel case. The resultant exclusive maximum tap-selection is then applied to two channel NLMS, AP and RLS algorithms. Although our main motivation is not the reduction of complexity of SAEC, the proposed tap-selection nevertheless brings significant computation savings in additional to an improved rate of convergence over algorithms using only a nonlinear preprocessor. Andy W. H. Khong, Patrick A. Naylor |
ICASSP (3) | 2 |
| 2005 | Blind identification using second-order statistics: a nonstationarity and nonwhiteness approachabstractWe consider an approach to the blind identification problem of instantaneous mixtures using second-order statistics through the nonstationarity and nonwhiteness properties of signals. We propose the use of natural gradient learning to form off-line/block processing (BP) and on-line processing (OP) algorithms suitable respectively for blind identification with batch data and on-line data and show that the proposed algorithms can be considered as a class of algorithms offering quasi-uniform performance. The identifiability conditions are presented which provide a key insight into these algorithms. The paper shows simulation results and concludes with some connections of the proposed algorithms to other existing algorithms. Uttachai Manmontri, Patrick A. Naylor |
ICASSP (5) | 2 |
| 2005 | Selective-tap adaptive algorithms in the solution of the nonuniqueness problem for stereophonic acoustic echo cancellationabstractWe investigate stereophonic acoustic echo cancellation in which solutions for the system can be nonunique and propose the use of selective-tap adaptive filters to address this problem. The main concept is to employ tap selection to optimize jointly for minimum interchannel coherence and maximum L/sub 2/-norm of the subselected tap-input vectors. The exclusive maximum (XM) tap-selection approach is proposed and applied to normalized least-mean squares (NLMS) and recursive least-squares (RLS) algorithm. We propose an approach for solving the nonuniqueness problem employing XM tap selection in combination with a nonlinear preprocessor. Simulation results show a significant improvement in convergence rate compared with existing techniques. Andy W. H. Khong, Patrick A. Naylor |
IEEE Signal Process. Lett. | 2 |
| 2005 | Corrections to "Selective-Tap Adaptive Algorithms in the Solution of the Nonuniqueness Problem for Stereophonic Acoustic Echo Cancellation"
Andy W. H. Khong, Patrick A. Naylor |
IEEE Signal Process. Lett. | 2 |
| 2004 | An improved IPNLMS algorithm for echo cancellation in packet-switched networksabstractWe present an improved adaptive echo cancellation algorithm designed for use with sparse echo path impulse responses such as arise from packet-switched networks. The new approach implicitly segments the impulse response into 'active' and 'inactive' regions, and employs different proportionate updating in each region. An efficient partial updating scheme is then formulated for the new algorithm. Evaluation results are presented to compare the new algorithm against three existing methods in terms of convergence and computational complexity. The results show that the new algorithm outperforms the best existing technique and has lower complexity. Jingjing Cui 0005, Patrick A. Naylor, David T. Brown |
ICASSP (4) | 2 |
| 2004 | Expected performance of a family of blind source separation algorithms in a reverberant roomabstractUsing statistical room acoustics, we investigate the performance of blind source separation and deconvolution (BSSD) algorithms when used in a reverberant room. We focus on the case where one of the sources moves, and examine the relative impact of source movement and room reverberation on the expected performance. We derive theoretical expressions, and verify these through image model simulations. Fotios Talantzis, Darren B. Ward, Patrick A. Naylor |
ICASSP (4) | 3 |
| 2003 | A short-sort M-Max NLMS partial-update adaptive filter with applications to echo cancellationabstractPartial-update algorithms reduce adaptive filter complexity by updating only a subset of taps at each iteration. However, they suffer a processing overhead in tap selection that can substantially reduce the computational advantages of partial-update schemes. Short-sort M-Max NLMS (SM-NLMS) addresses this problem by having the advantages of other partial-update schemes but with very low computational overhead in tap selection. SM-NLMS uses a low-complexity short-sort procedure to perform tap selection and updates the selection periodically. We show a performance analysis based on contraction mapping for SM-NLMS using a time-varying unknown system and quantify its characteristics. Simulation results and the performance analysis show that SM-NLMS performs almost as well as NLMS but with substantially lower computational cost involved in tap selection and updating compared to other schemes. The straightforward structure and low complexity of SM-NLMS make it well suited to real-time and high-density applications such as echo cancellation and equalization. Patrick A. Naylor, Warren Sherliker |
ICASSP (5) | 1 |
| 2003 | I/Q mismatch compensation in zero-IF OFDM receivers with application to DABabstractThis work addresses the I/Q mismatch problem which arises due to analog component tolerances in zero-IF OFDM receivers. The application of this work is to digital audio broadcasting (DAB). The approach is to employ a decision directed LMS-based frequency-adaptive equalizer, applied to OFDM demodulated carriers. The equalizer is trained initially using the phase reference symbols in each frame and is switched to a decision directed mode when the error reaches a defined level. In the decision directed mode, the output is taken prior to the threshold decision device in order to present to the subsequent Viterbi decoder a signal suitable for soft decisions. Simulation results demonstrate convergence within six frames; the particular convergence rate depending on choice of step-size and number of carriers per group employed. Even with large phase and amplitude imbalances, 25/spl deg/ and 5 dB respectively, in a DAB system the compensation algorithm attains a reduction in raw BER from 10/sup -1.5/ to 10/sup -5/ with a 15 dB SNR. Alexander R. Wright, Patrick A. Naylor |
ICASSP (2) | 2 |
| 2002 | The DYPSA algorithm for estimation of glottal closure instants in voiced speechabstractWe present the DYPSA algorithm for automatic and reliable estimation of glottal closure instants (GCIs) in voiced speech. Reliable GCI estimation is essential for closed-phase speech analysis, from which can be derived features of the vocal tract and, separately, the voice source. It has been shown that such features can be used with significant advantages in applications such as speaker recognition. DYPSA is automatic and operates using the speech signal alone without the need for an EGG or Laryngograph signal. It incorporates a new technique for estimating GCI candidates and employs dynamic programming to select the most likely candidates according to a defined cost function. We review and evaluate three existing methods and compare our new algorithm to them. Results for DYPSA show GCI detection accuracy to within ±0.25ms on 87% of the test database and fewer than 1% false alarms and misses. Anastasis Kounoudes, Patrick A. Naylor, Mike Brookes |
ICASSP | 2 |
| 2001 | Dynamic structures for non-uniform subband adaptive filteringabstractSubband adaptive filters suffer degraded performance when high input energy occurs at frequencies coincident with subband boundaries. This is seen as increased error in critically sampled systems and as reduced asymptotic convergence speed in oversampled systems. To address this problem a dynamic frequency decomposition scheme is presented which aims to control the frequency of subband boundaries such that they avoid spectral regions of high input energy. An efficient structure for this is described, which maintains the low-complexity advantage of subband systems. Simulation results show reductions in MSE of around 5-10 dB in the critical case and convergence improvement in the oversampled case, in addition to increased robustness to coloured inputs in both cases. Amere Oakman, Patrick A. Naylor |
ICASSP | 2 |
| 1998 | Application of the leaky extended LMS (XLMS) algorithm in stereophonic acoustic echo cancellation
Tetsuya Hoya, Y. Loke, Jonathon A. Chambers, Patrick A. Naylor |
Signal Process. | 4 |
| 1998 | Subband adaptive filtering for acoustic echo control using allpass polyphase IIR filterbanksabstractAdaptive filtering in subbands is an attractive alternative to full-band schemes in many applications because of the potential for faster convergence and lower computational cost. However, the analysis of a signal into a subband representation and the synthesis back into its original full-band form carries three main penalties. These are that (1) the subsampling process often introduces aliasing, (2) the subband analysis and synthesis processes carry a computational overhead, thereby reducing the gain in efficiency, and (3) the subband analysis and synthesis processes introduce delay into the signal path. A subband scheme is presented that aims to minimize these penalties, thereby allowing the potential advantages of the subband approach to be more fully realized. The scheme is based on infinite impulse response (IIR) filterbanks, formed from allpass polyphase filters, which exhibit very high quality filtering compared to typical finite impulse response (FIR) implementations, have relatively low complexity, introduce a limited degree of phase distortion and have low delay. The scheme, in conjunction with normalized least mean squares (NLMS) adaptive filters, is tested in an acoustic echo control application and shown to give better convergence, lower delay, and lower computational cost than a comparable FIR subband scheme. Patrick A. Naylor, Oguz Tanrikulu, Anthony G. Constantinides |
IEEE Trans. Speech Audio Process. | 1 |
| 1997 | Voice activity detection using source separation techniquesabstractA novel Voice Activity Detector is presented that is based on Source Separation techniques applied to single sensor signals. It offers very accurate estimation of the endpoints in very low Signal to Noise ratio conditions, while maintaining low complexity. Since the procedure is totally iterative, it is suitable for use in real-time applications and is capable of operating in dynamically adapting situations. Results are presented for both White Gaussian and Car Engine background noise. The performance of the new technique is compared with that of the GSM Voice Activity Detector. 1. Introduction Voice Activity Detection (VAD) is important in many areas of speech processing technology, such as noise reduction, voice recognition, speech coding etc, and has been extensively studied ([7], [5], [1]). Most of the existing techniques focus on relatively mild noise conditions (small positive SNR, for example the conditions found in an office environment). The work presented in this paper focus... Nikos Doukas, Patrick A. Naylor, Tania Stathaki |
EUROSPEECH | 2 |
| 1995 | Finite-precision design and implementation of all-pass polyphase networks for echo cancellation in sub-bandsabstractAll-pass polyphase networks (APN) are particularly attractive for acoustical echo cancellation (AEC) arranged in sub-bands. They provide lower inter-band aliasing, delay and computational complexity than their FIR counterparts. Moreover, APNs achieve higher echo return loss enhancement (ERLE) performance and faster convergence than full-band processing. In the paper, the finite precision implementation of APNs is addressed. A procedure is presented for re-optimising the all-pass coefficients of the prototype low-pass filter for finite precision operation. Robust finite precision implementation of a prototype low-pass filter is discussed. The results of a set of AEC experiments are reported with full and 16-bit precision implementation. Oguz Tanrikulu, Buyurman Baykal, Anthony G. Constantinides, Jonathon A. Chambers, Patrick A. Naylor |
ICASSP | 5 |
| 1993 | Polyphase allpass IIR structures for sub-band acoustic echo cancellationabstractThe advantages and current limitations of sub-band approaches to echo cancellation are reviewed. Polyphase allpass IIR halfband decimators and interpolators are presented. Their performance is compared to QMF structures in NLMS sub-band echo cancellation for hands-free telephone signals recorded in a car. The polyphase allpass IIR case is shown to give around 2 dB more ERLE with one fifth of the number of multiplies compared to direct form QMF. J. E. Hart, Patrick A. Naylor, Oguz Tanrikulu |
EUROSPEECH | 2 |
| 1988 | Speech production modelling with variable glottal reflection coefficientabstractRecognition and synthesis of speech require accurate models of speech production. Inherent in linear predictive modelling is the assumption that the formant frequencies and bandwidths of speech do not change within the analysis frame-usually at least a larynx cycle. It is shown that there can be significant differences in formant characteristics between the closed and open glottis phases of the larynx cycle. A development of the lossless tube model is presented in which the formant variations between closed and open glottis phases are modelled by a time varying glottal reflection coefficient. A numerically based method for estimating the parameters of such a model is outlined. The results of analysis and re-synthesis of female voiced speech obtained using this model are compared to those obtained using conventional LPC.> David Michael Brookes, Patrick A. Naylor |
ICASSP | 2 |