Mike Brookes

dblp:41/4618 · DBLP profile ↗
← Back
64ranked-venue papers
2as first author
9since 2021 · last 2024
0000-0001-7105-4936ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 51 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 19 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2024 Speech Enhancement in Hearing Aids Using Target Speech Presence Estimation Based on a Delayed Remote Microphone Signal
abstract
Speech enhancement in hearing aids (HAs) can take advantage of a wireless remote microphone (RM) having a better signal-to-noise ratio than the HA microphones. However, using the RM effectively is complicated by the time delay between the acoustic and wireless signals. Methods in the literature assume an instantaneous transmission of the RM signal, which is never the case in practice. Hence, we propose a practically operational method to use an RM with HAs in the presence of wireless transmission delays. Specifically, we use the delayed target voice activity state from the RM signal, to derive an expression for target speech presence probability (SPP) at the local HA microphone signals. The proposed method uses this target SPP mask as a post-filter that follows a local multichannel Wiener filter. Through simulations, we demonstrate that the proposed method improves speech quality and intelligibility metrics, especially in very noisy acoustic environments, compared to a standard approach which relies solely on HA microphones.
Vasudha Sathyapriyan, Michael Syskind Pedersen, Mike Brookes, Jan Østergaard, Patrick A. Naylor, Jesper Jensen 0001
ICASSP3
2024 Binaural Speech Enhancement Using Deep Complex Convolutional Transformer Networks
abstract
Studies have shown that in noisy acoustic environments, providing binaural signals to the user of an assistive listening device may improve speech intelligibility and spatial awareness. This paper presents a binaural speech enhancement method using a complex convolutional neural network with an encoder-decoder architecture and a complex multi-head attention transformer. The model is trained to estimate individual complex ratio masks in the time-frequency domain for the left and right-ear channels of binaural hearing devices. The model is trained using a novel loss function that incorporates the preservation of spatial information along with speech intelligibility improvement and noise reduction. Simulation results for acoustic scenarios with a single target speaker and isotropic noise of various types show that the proposed method improves the estimated binaural speech intelligibility and preserves the binaural cues better in comparison with several baseline algorithms.
Vikas Tokala, Eric Grinstein, Mike Brookes, Simon Doclo, Jesper Jensen 0001, Patrick A. Naylor
ICASSP3
2023 Graph Neural Networks for Sound Source Localization on Distributed Microphone Networks
abstract
Distributed Microphone Arrays (DMAs) present many challenges with respect to centralized microphone arrays. An important requirement of applications on these arrays is handling a variable number of input channels. We consider the use of Graph Neural Networks (GNNs) as a solution to this challenge. We present a localization method using the Relation Network GNN, which we show shares many similarities to classical signal processing algorithms for Sound Source Localization (SSL). We apply our method for the task of SSL and validate it experimentally using an unseen number of microphones. We test different feature extractors and show that our approach significantly outperforms classical baselines.
Eric Grinstein, Mike Brookes, Patrick A. Naylor
ICASSP2
2023 The MBSTOI Binaural Intelligibility Metric Using a Close-Talking Microphone Reference
abstract
Intelligibility metrics are a fast way to determine how comprehensible a target signal is in a noisy situation. Most metrics however rely on having a clean reference signal for computation and are not adapted to live recordings. In this paper the deep correlation modified binaural short time objective intelligibility metric (Dcor-MBSTOI) is evaluated with a single-channel close-talking microphone signal as the reference. This reference signal inevitably contains some background noise and crosstalk from non-target sources. It is found that intelligibility is overestimated when using the close-talking microphone signal directly but that this overestimation can be eliminated by applying speech enhancement to the reference signal.
Pierre Guiraud, Alastair H. Moore, Rebecca R. Vos, Patrick A. Naylor, Mike Brookes
ICASSP5
2023 Epoch-Based Spectrum Estimation for Speech
Jón Guðnason, Guolin Fang, Mike Brookes
INTERSPEECH3
2023 Using a single-channel reference with the MBSTOI binaural intelligibility metric
abstract
In order to assess the intelligibility of a target signal in a noisy environment, intrusive speech intelligibility metrics are typically used. They require a clean reference signal to be available which can be difficult to obtain especially for binaural metrics like the modified binaural short time objective intelligibility metric (MBSTOI). We here present a hybrid version of MBSTOI that incorporates a deep learning stage that allows the metric to be computed with only a single-channel clean reference signal. The models presented are trained on simulated data containing target speech, localised noise, diffuse noise, and reverberation. The hybrid output metrics are then compared directly to MBSTOI to assess performances. Results show the performance of our single channel reference vs MBSTOI. The outcome of this work offers a fast and flexible way to generate audio data for machine learning (ML) and highlights the potential for low level implementation of ML into existing tools.
Pierre Guiraud, Alastair H. Moore, Rebecca R. Vos, Patrick A. Naylor, Mike Brookes
Speech Commun.5
2022 A Compact Noise Covariance Matrix Model for MVDR Beamforming
abstract
Acoustic beamforming is routinely used to improve the SNR of the received signal in applications such as hearing aids, robot audition, augmented reality, teleconferencing, source localisation and source tracking. The beamformer can be made adaptive by using an estimate of the time-varying noise covariance matrix in the spectral domain to determine an optimised beam pattern in each frequency bin that is specific to the acoustic environment and that can respond to temporal changes in it. However, robust estimation of the noise covariance matrix remains a challenging task especially in non-stationary acoustic environments. This paper presents a compact model of the signal covariance matrix that is defined by a small number of parameters whose values can be reliably estimated. The model leads to a robust estimate of the noise covariance matrix which can, in turn, be used to construct a beamformer. The performance of beamformers designed using this approach is evaluated for a spherical microphone array under a range of conditions using both simulated and measured room impulse responses. The proposed approach demonstrates consistent gains in intelligibility and perceptual quality metrics compared to the static and adaptive beamformers used as baselines.
Alastair H. Moore, Sina Hafezi, Rebecca R. Vos, Patrick A. Naylor, Mike Brookes
IEEE ACM Trans. Audio Speech Lang. Process.5
2021 Processing Pipelines for Efficient, Physically-Accurate Simulation of Microphone Array Signals in Dynamic Sound Scenes
abstract
Multichannel acoustic signal processing is predicated on the fact that the interchannel relationships between the received signals can be exploited to infer information about the acoustic scene. Recently there has been increasing interest in algorithms which are applicable in dynamic scenes, where the source(s) and/or microphone array may be moving. Simulating such scenes has particular challenges which are exacerbated when real-time, listener-in-the-loop evaluation of algorithms is required. This paper considers candidate pipelines for simulating the array response to a set of point/image sources in terms of their accuracy, scalability and continuity. A new approach, in which the filter kernels are obtained using principal component analysis from time-aligned impulse responses, is proposed. When the number of filter kernels is ≤40 the new approach achieves more accurate simulation than competing methods.
Alastair H. Moore, Rebecca R. Vos, Patrick A. Naylor, Mike Brookes
ICASSP4
2021 Speech Enhancement Based on Modulation-Domain Parametric Multichannel Kalman Filtering
abstract
Recently we presented a modulation-domain multichannel Kalman filtering (MKF) algorithm for speech enhancement, which jointly exploits the inter-frame modulation-domain temporal evolution of speech and the inter-channel spatial correlation to estimate the clean speech signal. The goal of speech enhancement is to suppress noise while keeping the speech undistorted, and a key problem is to achieve the best trade-off between speech distortion and noise reduction. In this paper, we extend the MKF by presenting a modulation-domain parametric MKF (PMKF) which includes a parameter that enables flexible control of the speech enhancement behaviour in each time-frequency (TF) bin. Based on the decomposition of the MKF cost function, a new cost function for PMKF is proposed, which uses the controlling parameter to weight the noise reduction and speech distortion terms. An optimal PMKF gain is derived using a minimum mean squared error (MMSE) criterion. We analyse the performance of the proposed MKF, and show its relationship to the speech distortion weighted multichannel Wiener filter (SDW-MWF). To evaluate the impact of the controlling parameter on speech enhancement performance, we further propose PMKF speech enhancement systems in which the controlling parameter is adaptively chosen in each TF bin. Experiments on a publicly available head-related impulse response (HRIR) database in different noisy and reverberant conditions demonstrate the effectiveness of the proposed method.
Wei Xue 0002, Alastair H. Moore, Mike Brookes, Patrick A. Naylor
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Modulation-Domain Kalman Filtering for Monaural Blind Speech Denoising and Dereverberation
abstract
We describe a monaural speech enhancement algorithm based on modulation-domain Kalman filtering to blindly track the time-frequency log-magnitude spectra of speech and reverberation. We propose an adaptive algorithm that performs blind joint denoising and dereverberation, while accounting for the inter-frame speech dynamics, by estimating the posterior distribution of the speech log-magnitude spectrum given the log-magnitude spectrum of the noisy reverberant speech. The Kalman filter update step models the non-linear relations between the speech, noise, and reverberation log spectra. The Kalman filtering algorithm uses a signal model that takes into account the reverberation parameters of the reverberation time T60and the direct-to-reverberant energy ratio (DRR) and also estimates and tracks T60and the DRR in every frequency bin to improve the estimation of the speech log spectrum. The proposed algorithm is evaluated in terms of speech quality, speech intelligibility, and dereverberation performance for a range of reverberation parameters and reverberant speech to noise ratios, in different noises, and is also compared to competing denoising and dereverberation techniques. Experimental results using noisy reverberant speech demonstrate the effectiveness of the enhancement algorithm.
Nikolaos Dionelis, Mike Brookes
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 Noise Covariance Matrix Estimation for Rotating Microphone Arrays
abstract
The noise covariance matrix computed between the signals from a microphone array is used in the design of spatial filters and beamformers with applications in noise suppression and dereverberation. This paper specifically addresses the problem of estimating the covariance matrix associated with a noise field when the array is rotating during desired source activity, as is common in head-mounted arrays. We propose a parametric model that leads to an analytical expression for the microphone signal covariance as a function of the array orientation and array manifold. An algorithm for estimating the model parameters during noise-only segments is proposed and the performance shown to be improved, rather than degraded, by array rotation. The stored model parameters can then be used to update the covariance matrix to account for the effects of any array rotation that occurs when the desired source is active. The proposed method is evaluated in terms of the Frobenius norm of the error in the estimated covariance matrix and of the noise reduction performance of a minimum variance distortionless response beamformer. In simulation experiments the proposed method achieves 18 dB lower error in the estimated noise covariance matrix than a conventional recursive averaging approach and results in noise reduction which is within 0.05 dB of an oracle beamformer using the ground truth noise covariance matrix.
Alastair H. Moore, Wei Xue 0002, Patrick A. Naylor, Mike Brookes
IEEE ACM Trans. Audio Speech Lang. Process.4
2018 Room Identification Using Frequency Dependence of Spectral Decay Statistics
abstract
A method for room identification is proposed based on the reverberation properties of multichannel speech recordings. The approach exploits the dependence of spectral decay statistics on the reverberation time of a room. The average negative-side variance within 1/3-octave bands is proposed as the identifying feature and shown to be effective in a classification experiment. However, negative-side variance is also dependent on the direct-to-reverberant energy ratio. The resulting sensitivity to different spatial configurations of source and microphones within a room are mitigated using a novel reverberation enhancement algorithm. A classification experiment using speech convolved with measured impulse responses and contaminated with environmental noise demonstrates the effectiveness of the proposed method, achieving 79% correct identification in the most demanding condition compared to 40% using unenhanced signals.
Alastair H. Moore, Patrick A. Naylor, Mike Brookes
ICASSP3
2018 Multichannel Kalman Filtering for Speech Ehnancement
abstract
The use of spatial information in multichannel speech enhancement methods is well established but information associated with the temporal evolution of speech is less commonly exploited. Speech signals can be modelled using an autoregressive process in the time-frequency modulation domain, and Kalman filtering based speech enhancement algorithms have been developed for single-channel processing. In this paper, a multichannel Kalman filter (MKF) for speech enhancement is derived that jointly considers the multichannel spatial information and the temporal correlations of speech. We model the temporal evolution of speech in the modulation domain and, by incorporating the spatial information, an optimal MKF gain is derived in the short-time Fourier transform domain. We also show that the proposed MKF becomes a conventional multichannel Wiener filter if the temporal information is discarded. Experiments using the signals generated from a public head-related impulse response database demonstrate the effectiveness of the proposed method in comparison to other techniques.
Wei Xue 0002, Alastair H. Moore, Mike Brookes, Patrick A. Naylor
ICASSP3
2018 Phase-Aware Single-Channel Speech Enhancement With Modulation-Domain Kalman Filtering
abstract
We present a speech enhancement algorithm that performs modulation-domain Kalman filtering to track the speech phase using circular statistics, along with the spectral log-amplitudes of speech and noise. In the proposed algorithm, the speech phase posterior is used to create an enhanced speech phase spectrum for the signal reconstruction of speech. The Kalman filter prediction step separately models the temporal inter-frame correlation of the speech and noise spectral log-amplitudes and of the speech phase, while the Kalman filter update step models their nonlinear relations under the assumption that speech and noise add in the complex short-time Fourier transform domain. The phase-sensitive enhancement algorithm is evaluated with speech quality and intelligibility metrics, using a variety of noise types over a range of SNRs. Instrumental measures predict that tracking the speech log-spectrum and phase with modulation-domain Kalman filtering leads to consistent improvements in speech quality, over both conventional enhancement algorithms and other algorithms that perform modulation-domain Kalman filtering.
Nikolaos Dionelis, Mike Brookes
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Model-Based Speech Enhancement in the Modulation Domain
abstract
This paper presents an algorithm for modulation-domain speech enhancement using a Kalman filter. The proposed estimator jointly models the estimated dynamics of the spectral amplitudes of speech and noise to obtain an MMSE estimation of the speech amplitude spectrum with the assumption that the speech and noise are additive in the complex domain. In order to include the dynamics of noise amplitudes with those of speech amplitudes, we propose a statistical “Gaussring” model that comprises a mixture of Gaussians whose centers lie in a circle on the complex plane. The performance of the proposed algorithm is evaluated using the perceptual evaluation of speech quality measure, segmental SNR measure, and short-time objective intelligibility measure. For speech quality measures, the proposed algorithm is shown to give a consistent improvement over a wide range of SNRs when compared to competitive algorithms. Speech recognition experiments also show that the Gaussring-model-based algorithm performs well for two types of noise.
Yu Wang 0027, Mike Brookes
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Modulation-Domain Multichannel Kalman Filtering for Speech Enhancement
abstract
Compared with single-channel speech enhancement methods, multichannel methods can utilize spatial information to design optimal filters. Although some filters adaptively consider second-order signal statistics, the temporal evolution of the speech spectrum is usually neglected. By using linear prediction (LP) to model the inter-frame temporal evolution of speech, single-channel Kalman filtering (KF) based methods have been developed for speech enhancement. In this paper, we derive a multichannel KF (MKF) that jointly uses both interchannel spatial correlation and interframe temporal correlation for speech enhancement. We perform LP in the modulation domain, and by incorporating the spatial information, derive an optimal MKF gain in the short-time Fourier transform domain. We show that the proposed MKF reduces to the conventional multichannel Wiener filter if the LP information is discarded. Furthermore, we show that, under an appropriate assumption, the MKF is equivalent to a concatenation of the minimum variance distortion response beamformer and a single-channel modulation-domain KF and therefore present an alternative implementation of the MKF. Experiments conducted on a public head-related impulse response database demonstrate the effectiveness of the proposed method.
Wei Xue 0002, Alastair H. Moore, Mike Brookes, Patrick A. Naylor
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 Identifying a multiple plane plenoptic function from a swiped image
abstract
Blur in images, caused by camera motion with an open shutter, is usually thought of as a problem. The algorithm described in this paper shows instead that it is possible to use the blur caused by the integration of light rays at different locations along a moving camera trajectory to extract information about the light rays that are present within the scene. Retrieving the light rays present within a scene from different viewpoints is equivalent to retrieving the plenoptic function of the scene. In this paper, we focus on a specific case in which the blurred image of a scene, containing fronto-parallel planes with uniform unknown textures, is analysed to recreate the plenoptic function. The image is captured by a digital single lens camera with shutter open, moving in a straight line between two points, resulting in a swiped image. We estimate the EPI from this blurred image, and the EPI can be used to generate unblurred images for a given camera location.
Michael Lawson, Mike Brookes, Pier Luigi Dragotti
ICASSP2
2017 Improving the perceptual quality of ideal binary masked speech
abstract
It is known that applying a time-frequency binary mask to very noisy speech can improve its intelligibility but results in poor perceptual quality. In this paper we propose a new approach to applying a binary mask that combines the intelligibility gains of conventional binary masking with the perceptual quality gains of a classical speech enhancer. The binary mask is not applied directly as a time-frequency gain as in most previous studies. Instead, the mask is used to supply prior information to a classical speech enhancer about the probability of speech presence in different time-frequency regions. Using an oracle ideal binary mask, we show that the proposed method results in a higher predicted quality than other methods of applying a binary mask whilst preserving the improvements in predicted intelligibility.
Leo Lightburn, Enzo De Sena, Alastair H. Moore, Patrick A. Naylor, Mike Brookes
ICASSP5
2017 Robust spherical harmonic domain interpolation of spatially sampled array manifolds
abstract
Accurate interpolation of the array manifold is an important first step for the acoustic simulation of rapidly moving microphone arrays. Spherical harmonic domain interpolation has been proposed and well studied in the context of head-related transfer functions but has focussed on perceptual, rather than numerical, accuracy. In this paper we analyze the effect of measurement noise on spatial aliasing. Based on this analysis we propose a method for selecting the truncation orders for the forward and reverse spherical Fourier transforms given only the noisy samples in such a way that the interpolation error is minimized. The proposed method achieves up to 1.7 dB improvement over the baseline approach.
Alastair H. Moore, Mike Brookes, Patrick A. Naylor
ICASSP2
2017 Frequency-domain under-modelled blind system identification based on cross power spectrum and sparsity regularization
abstract
In room acoustics, under-modelled multichannel blind system identification (BSI) aims to estimate the early part of the room impulse responses (RIRs), and it can be widely used in applications such as speaker localization, room geometry identification and beamforming based speech dereverberation. In this paper we extend our recent study on under-modelled BSI from the time domain to the frequency domain, such that the RIRs can be updated frame-wise and the efficiency of Fast Fourier Transform (FFT) is exploited to reduce the computational complexity. Analogous to the cross-correlation based criterion in the time domain, a frequency-domain cross power spectrum based criterion is proposed. As the early RIRs are usually sparse, the RIRs are estimated by jointly maximizing the cross power spectrum based criterion in the frequency domain and minimizing the l1-norm sparsity measure in the time domain. A two-stage LMS updating algorithm is derived to achieve joint optimization of these two targets. The experimental results in different under-modelled scenarios demonstrate the effectiveness of the proposed method.
Wei Xue 0002, Mike Brookes, Patrick A. Naylor
ICASSP2
2017 Single-Channel Online Enhancement of Speech Corrupted by Reverberation and Noise
abstract
This paper proposes an online single-channel speech enhancement method designed to improve the quality of speech degraded by reverberation and noise. Based on an autoregressive model for the reverberation power and on a hidden Markov model for clean speech production, a Bayesian filtering formulation of the problem is derived and online joint estimation of the acoustic parameters and mean speech, reverberation, and noise powers is obtained in mel-frequency bands. From these estimates, a real-valued spectral gain is derived and spectral enhancement is applied in the short-time Fourier transform (STFT) domain. The method yields state-of-the-art performance and greatly reduces the effects of reverberation and noise while improving speech quality and preserving speech intelligibility in challenging acoustic environments.
Clement S. J. Doire, Mike Brookes, Patrick A. Naylor, Christopher M. Hicks, Dave Betts, Mohammad A. Dmour, Søren Holdt Jensen
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Speech enhancement using an MMSE spectral amplitude estimator based on a modulation domain Kalman filter with a Gamma prior
abstract
In this paper, we propose a minimum mean square error spectral estimator for clean speech spectral amplitudes that uses a Kalman filter to model the temporal dynamics of the spectral amplitudes in the modulation domain. Using a two-parameter Gamma distribution to model the prior distribution of the speech spectral amplitudes, we derive closed form expressions for the posterior mean and variance of the spectral amplitudes as well as for the associated update step of the Kalman filter. The performance of the proposed algorithm is evaluated on the TIMIT core test set using the perceptual evaluation of speech quality (PESQ) measure and segmental SNR measure and is shown to give a consistent improvement over a wide range of SNRs when compared to competitive algorithms.
Yu Wang 0027, Mike Brookes
ICASSP2
2016 A weighted STOI intelligibility metric based on mutual information
abstract
It is known that the information required for the intelligibility of a speech signal is distributed non-uniformly in time. In this paper we propose WSTOI, a modified version of STOI, a speech intelligibility metric. With WSTOI the contribution of each time-frequency cell is weighted by an estimate of its intelligibility content. This estimate is equal to the mutual information between two hypothetical signals at either end of a simplified model of human communication. Listening tests show that the modification improves the prediction accuracy of STOI at all performance levels on both long and short utterances. An improvement was observed across all tested noise types and suppression algorithms.
Leo Lightburn, Mike Brookes
ICASSP2
2016 A data-driven non-intrusive measure of speech quality and intelligibility
Dushyant Sharma, Yu Wang 0027, Patrick A. Naylor, Mike Brookes
Speech Commun.4
2015 Single-channel blind estimation of reverberation parameters
abstract
The reverberation of an acoustic channel can be characterised by two frequency-dependent parameters: the reverberation time and the direct-to-reverberant energy ratio. This paper presents an algorithm for blindly determining these parameters from a single-channel speech signal. The algorithm uses an extended Kalman filter to estimate the parameters together with a hidden semi-Markov model to identify intervals of speech activity.
Clement S. J. Doire, Mike Brookes, Patrick A. Naylor, Dave Betts, Christopher M. Hicks, Mohammad A. Dmour, Søren Holdt Jensen
ICASSP2
2015 Speaker change detection and speaker diarization using spatial information
abstract
In this paper, we present a novel speaker change detection and speaker diarization algorithm using spatial information in the form of features derived from estimated Room Impulse Response (RIR)s. A blind system identification approach is used to obtain an estimate of the RIRs, from which the C5 feature is derived and used in the labeling algorithm. Experimental results using 2 speakers for different locations within a fixed room show that our approach achieves a higher hit rate in the speaker change detection task and a lower variance in the diarization error rate when compared with a baseline algorithm.
Mathieu Hu, Dushyant Sharma, Simon Doclo, Mike Brookes, Patrick A. Naylor
ICASSP4
2015 SOBM - a binary mask for noisy speech that optimises an objective intelligibility metric
abstract
It is known that the intelligibility of noisy speech can be improved by applying a binary-valued gain mask to a time-frequency representation of the speech. We present the SOBM, an oracle binary mask that maximises STOI, an objective speech intelligibility metric. We show how to determine the SOBM for a deterministic noise signal and also for a stochastic noise signal with a known power spectrum. We demonstrate that applying the SOBM to noisy speech results in a higher predicted intelligibility than is obtained with other masks and show that the stochastic version is robust to mismatch errors in SNR and noise spectrum.
Leo Lightburn, Mike Brookes
ICASSP2
2014 Mask-based enhancement for very low quality speech
abstract
We propose a mask-based enhancer for very low quality speech that is able to preserve important cues in a noise-robust manner by identifying the time-frequency regions that contain significant speech energy. We use a classifier to estimate a time-frequency mask from an input feature set that provides information about the energy distribution of both voiced and unvoiced speech. We evaluate the enhancer on a range of noisy speech signals and demonstrate that it yields consistent improvements in an objective intelligibility measure.
Sira Gonzalez, Mike Brookes
ICASSP2
2014 Speech enhancement usinga modulation domain Kalman filter post-processor with a Gaussian Mixture noise model
abstract
We propose a speech enhancement algorithm that applies a Kalman filter in the modulation domain to the output of a conventional enhancer operating in the time-frequency domain. We show that the prediction residual signal of the spectral amplitude errors at the output of the baseline MMSE enhancer do not follow a Gaussian distribution. Accordingly, the Kalman filter used in our enhancement algorithm combines a colored noise model with a Gaussian mixture model of the residual noise. We evaluate the performance of the speech enhancement algorithm on the core TIMIT test set and demonstrate that it gives consistent performance improvements over the baseline enhancer and over a previously proposed Kalman filter post-processor.
Yu Wang 0027, Mike Brookes
ICASSP2
2014 Wide-baseline image change detection
abstract
We present a fully automated method for the detection of changes within a scene between a reference and a sample image whose viewing angles differ by up to 30°. We also describe an extension to the SIFT technique that allows extracted feature points to be matched over wider viewing angles. Matched correspondences between reference and sample images are used to construct a Delaunay triangulation and changes are detected by comparing triangles after affine compensation using a dense SIFT metric. False positives are reduced by using a novel technique introduced as local plane matching (LPM) to match mean-shift segments in unmatched areas using the homographies of local planes to compensate for perspective distortions. The method is shown to achieve pixel-level equal error rates of 5% at a 10° azimuth view angle difference.
Ziggy Jones, Mike Brookes, Pier Luigi Dragotti, David M. Benton
ICIP2
2014 Tilted layer-based modeling for enhanced light-field processing and image based rendering
abstract
Image based rendering is an attractive approach for novel view synthesis due to its low complexity requirements and potential for photorealistic results. However for successful rendering, geometric priors about the structure of the scene are necessary. In this paper we present a tilted layer model approximation of the plenoptic function which gives improved modeling of scenes where the objects are not fronto-parallel to the camera views while preserving occlusion ordering. The framework is extended to the case where camera positions are not constrained to a single plane but can lie on multiple planes. Results on the Middlebury dataset and simulated scenes show that better rendering results can be obtained compared with the state-of-the-art using a fronto-parallel layer model, or alternatively similar results can be obtained with a more compact layer representation of the scene.
James Pearson, Marco Visentini Scarzanella, Mike Brookes, Pier Luigi Dragotti
ICIP3
2014 PEFAC - A Pitch Estimation Algorithm Robust to High Levels of Noise
abstract
We present PEFAC, a fundamental frequency estimation algorithm for speech that is able to identify voiced frames and estimate pitch reliably even at negative signal-to-noise ratios. The algorithm combines a normalization stage, to remove channel dependency and to attenuate strong noise components, with a harmonic summing filter applied in the log-frequency power spectral domain, the impulse response of which is chosen to sum the energy of the fundamental frequency harmonics while attenuating smoothly-varying noise components. Temporal continuity constraints are applied to the selected pitch candidates and a voiced speech probability is computed from the likelihood ratio of two classifiers, one for voiced speech and one for unvoiced speech/silence. We compare the performance of our algorithm with that of other widely used algorithms and demonstrate that it performs well in both high and low levels of additive noise.
Sira Gonzalez, Mike Brookes
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 On the Spectrum of the Plenoptic Function
abstract
The plenoptic function is a powerful tool to analyze the properties of multi-view image data sets. In particular, the understanding of the spectral properties of the plenoptic function is essential in many computer vision applications, including image-based rendering. In this paper, we derive for the first time an exact closed-form expression of the plenoptic spectrum of a slanted plane with finite width and use this expression as the elementary building block to derive the plenoptic spectrum of more sophisticated scenes. This is achieved by approximating the geometry of the scene with a set of slanted planes and evaluating the closed-form expression for each plane in the set. We then use this closed-form expression to revisit uniform plenoptic sampling. In this context, we derive a new Nyquist rate for the plenoptic sampling of a slanted plane and a new reconstruction filter. Through numerical simulations, on both real and synthetic scenes, we show that the new filter outperforms alternative existing filters.
Christopher Gilliam, Pier Luigi Dragotti, Mike Brookes
IEEE Trans. Image Process.3
2013 Speech active level estimation in noisy conditions
abstract
We present a new method for speech active level estimation which combines a novel algorithm based on voiced speech energy extraction with the standardized ITU-T Recommendation P.56. At poor signal-to-noise ratios, the algorithm estimates the active level by identifying intervals of voiced speech and summing the energy of the pitch harmonics in the time-frequency domain while rejecting that of the noise. We compare the performance of our method with that of ITU-T P.56 on the TIMIT database and demonstrate that it performs exceptionally well in both high and low levels of additive noise.
Sira Gonzalez, Mike Brookes
ICASSP2
2013 Speech enhancement using a robust Kalman filter post-processor in the modulation domain
abstract
We propose a speech enhancement algorithm that applies a Kalman filter in the modulation domain to the output of a conventional enhancer operating in the time-frequency domain. The speech model required by the Kalman filter is obtained by performing linear predictive analysis in each frequency bin of the modulation domain signal. We show, however, that the corresponding speech synthesis filter can have a very high gain at low frequencies and may approach instability. To improve the stability of the synthesis filter, we propose two alternative methods of limiting its low frequency gain. We evaluate the performance of the speech enhancement algorithm on the core TIMIT test set and demonstrate that it gives consistent performance improvements over the baseline enhancer.
Yu Wang 0027, Mike Brookes
ICASSP2
2013 Blind Channel Magnitude Response Estimation in Speech Using Spectrum Classification
abstract
We present an algorithm for blind estimation of the magnitude response of an acoustic channel from single microphone observations of a speech signal. The algorithm employs channel robust RASTA filtered Mel-frequency cepstral coefficients as features to train a Gaussian mixture model based classifier and average clean speech spectra are associated with each mixture; these are then used to blindly estimate the acoustic channel magnitude response from speech that has undergone spectral modification due to the channel. Experimental results using a variety of simulated and measured acoustic channels and additive babble noise, car noise and white Gaussian noise are presented. The results demonstrate that the proposed method is able to estimate a variety of channel magnitude responses to within an Itakura distance of dI ≤0.5 for SNR ≥10 dB.
Nikolay D. Gaubitch, Mike Brookes, Patrick A. Naylor
IEEE Trans. Speech Audio Process.2
2013 Plenoptic Layer-Based Modeling for Image Based Rendering
abstract
Image based rendering is an attractive alternative to model based rendering for generating novel views because of its lower complexity and potential for photo-realistic results. To reduce the number of images necessary for alias-free rendering, some geometric information for the 3D scene is normally necessary. In this paper, we present a fast automatic layer-based method for synthesizing an arbitrary new view of a scene from a set of existing views. Our algorithm takes advantage of the knowledge of the typical structure of multiview data to perform occlusion-aware layer extraction. In addition, the number of depth layers used to approximate the geometry of the scene is chosen based on plenoptic sampling theory with the layers placed non-uniformly to account for the scene distribution. The rendering is achieved using a probabilistic interpolation approach and by extracting the depth layer information on a small number of key images. Numerical results demonstrate that the algorithm is fast and yet is only 0.25 dB away from the ideal performance achieved with the ground-truth knowledge of the 3D geometry of the scene of interest. This indicates that there are measurable benefits from following the predictions of plenoptic theory and that they remain true when translated into a practical system for real world data.
James Pearson, Mike Brookes, Pier Luigi Dragotti
IEEE Trans. Image Process.2
2012 Image based rendering with depth cameras: How many are needed?
abstract
Image based rendering is a technique for producing arbitrary viewpoints of a scene using multiple images instead of exact object models. The recent emergence of low-price, fast, and reliable cameras for measuring depth makes possible the augmentation of traditional color images with depth images. This combination promises to improve the rendering quality of an arbitrary viewpoint and thus have a great impact on IBR. A key issue is to understand, for any particular scene of interest, how many depth images and how many color images are necessary in order to obtain good rendering results. In this paper, using a framework akin to the plenoptic function, we perform a spectral analysis of multi-view depth images in order to determine the relationship between the number of depth and color images required. Our analysis is then validated using both synthetic and real images.
Christopher Gilliam, James Pearson, Mike Brookes, Pier Luigi Dragotti
ICASSP3
2012 Non intrusive codec identification algorithm
abstract
We present a non-intrusive data driven method for codec detection and identification in the presence of background noise. The method uses a number of speech features which are then used to train a CART classifier. We demonstrate the performance of the method using several different noise types over a wide range of SNRs. Our results show that we can identify a codec and its bit rate to an accuracy of 92% and we are able to detect the presence of a codec with an accuracy of 97% at -5 dB SNR.
Dushyant Sharma, Patrick A. Naylor, Nikolay D. Gaubitch, Mike Brookes
ICASSP4
2012 Sibilant Speech Detection in Noise
abstract
We present an algorithm for identifying the location of sibilant phones in noisy speech. Our algorithm does not attempt to identify sibilant onsets and offsets directly but instead detects a sustained increase in power over the en-tire duration of a sibilant phone. The normalized esti-mate of the sibilant power in each of 14 frequency bands forms the input to two Gaussian mixture models that are trained on sibilant and non-sibilant frames respectively. The likelihood ratio of the two models is then used to classify each frame. We evaluate the performance of our algorithm on the TIMIT database and demonstrate that the classification accuracy is over 80 % at 0 dB signal to noise ratio for additive white noise. Index Terms: sibilant speech, spectrographic mask esti-mation, speech classification, speech segregation 1.
Sira Gonzalez, Mike Brookes
INTERSPEECH2
2012 Descriptive Vocabulary Development for Degraded Speech
abstract
This paper presents the development of a compact vocabulary for describing the audible characteristics of degraded speech. An experiment was conducted with 51 English-speaking subjects who were tasked with assigning one of a list of given text descriptors to 220 degradation conditions. Exploratory data analysis using hierarchical clustering resulted in a compact vocabulary of 10 classes, which was further validated by a bootstrap cluster analysis.
Dushyant Sharma, Gaston Hilkhuysen, Patrick A. Naylor, Nikolay D. Gaubitch, Mark A. Huckvale, Mike Brookes
INTERSPEECH6
2011 Accurate non-iterative depth layer extraction algorithm for image based rendering
abstract
Image based rendering is an attractive alternative for generating novel views compared to model based rendering due to its lower complexity and potential for photo-realistic results. We present a fast unsupervised method for synthesising arbitrary viewpoints of a scene from a set of existing views. Our novel improvements include optimising the placement of depth layers to take advantage of the composition of real world scenes and hierarchically building our simple geometric model to maximise its accuracy.
James Pearson, Pier Luigi Dragotti, Mike Brookes
ICASSP3
2011 Adaptive plenoptic sampling
abstract
The plenoptic function enables Image-based rendering (IBR) to be viewed in terms of sampling and reconstruction. Thus the spatial sampling rate can be determined through spectral analysis of the plenoptic function. In this paper we present a method of non-uniformly sampling a scene, with a smoothly varying surface, given a finite number of samples. This method approximates such a scene with a set of slanted planes subject to the constraint of finite number of samples. We use the recent spectral analysis of a single slanted plane to determine a piecewise constant spatial sampling rate for the scene. Finally, we show that this sampling rate results in a non-uniform sampling scheme that reconstructs the plenoptic function beyond that of uniform sampling.
Christopher Gilliam, Pier Luigi Dragotti, Mike Brookes
ICIP3
2010 A closed-form expression for the bandwidth of the plenoptic function under finite field of view constraints
abstract
The plenoptic function enables Image-based rendering (IBR) to be viewed in terms of sampling and reconstruction. Thus the spatial sampling rate can be determined through spectral analysis of the plenoptic function. In this paper we examine the bandwidth of the plenoptic function when both the field of view and the scene width are finite. This analysis is carried out on two planar Lambertian scenes, a fronto-parallel plane and a slanted plane, and in both cases the texture is bandlimited. We derive an exact closed-form expression for the plenoptic spectrum of a slanted plane with sinusoidal texture. We show that in both cases the finite constraints lead to band-unlimited spectra. By determining the essential bandwidth, we derive a sampling curve that gives an adequate camera spacing for a given distance between the scene and the camera line.
Christopher Gilliam, Pier Luigi Dragotti, Mike Brookes
ICIP3
2009 Adaptive layer extraction for image based rendering
abstract
Image based rendering is a promising way to produce arbitrary views of a scene using images instead of object models. However, depth variations and occlusions cause blurring in the rendered images. The solution is to use some geometrical information in order to steer the interpolation filters according to the depth. The level of detail of this geometry is often predetermined. In this paper, we present a method for extracting depth layers in the presence of occlusions for image based rendering. Moreover, we show how the layer extraction can be made to estimate depth layers in an adaptive manner, based on the spectral analysis of the plenoptic function. The rendering system therefore automatically adapts the number of depth layers based on the scene and the spacing of the sample cameras.
Jesse Berent, Pier Luigi Dragotti, Mike Brookes
MMSP3
2008 FEUDOR: Feature Extraction Using Distinctive Octagonal Regions
abstract
This paper introduces a novel feature extraction algorithm called FEUDOR. The features extracted by this method are octagonal homogeneous regions that have different mean square difference compared to their surrounding area. By using integral images we have implemented this algorithm effi-ciently. We have shown that the repeatability score of FEUDOR under vari-ous image transformations is comparable and in some cases better than other existing algorithms. 1
Ario Emaminejad, Mike Brookes
BMVC2
2008 Voice source cepstrum coefficients for speaker identification
abstract
We propose a novel feature set for speaker recognition that is based on the voice source signal. The feature extraction process uses closed-phase LPC analysis to estimate the vocal tract transfer function. The LPC spectrum envelope is converted to cepstrum coefficients which are used to derive the voice source features. Unlike approaches based on inverse-filtering, our procedure is robust to LPC analysis errors and low-frequency phase distortion. We have performed text-independent closed-set speaker identification experiments on the TIMIT and the YOHO databases using a standard Gaussian mixture model technique. Compared to using mel- frequency cepstrum coefficients, the misclassification rate for the TIMIT database reduced from 1.51% to 0.16% when combined with the proposed voice source features. For the YOHO database the mis- classification rate decreased from 13.79% to 10.07%. The new feature vector also compares favourably to other proposed voice source feature sets.
Jón Guðnason, Mike Brookes
ICASSP2
2008 A lowcomplexity fast converging partial update adaptive algorithm employing variable step-size for acoustic echo cancellation
abstract
Partial update adaptive algorithms have been proposed as a means of reducing complexity for adaptive filtering. The MMax tap-selection is one of the most popular tap-selection algorithms. It is well known that the performance of such partial update algorithm reduces with reducing number of filter coefficients selected for adaptation. We propose a low complexity and fast converging adaptive algorithm that exploits the MMax tap-selection. We achieve fast convergence with low complexity by deriving a variable step-size for the MMax normalized least-mean-square (MMax-NLMS) algorithm using its mean square deviation. Simulation results verify that the proposed algorithm achieves higher rate of convergence with lower computational complexity compared to the NLMS algorithm.
Andy W. H. Khong, Woon-Seng Gan, Patrick A. Naylor, Mike Brookes
ICASSP4
2007 Perceptual Gain Function for Eigenspectral Domain Speech Enhancement
abstract
The goals of speech enhancement are to improve its perceptual aspects most commonly by applying a gain function to the noisy signal coefficients in a transform domain. This gain function is normally chosen to provide a good trade-off between suppressing the noise and avoiding speech distortion. In this paper, we identify some desired gain characteristics for better sounding enhancement and propose a method for choosing the gain transfer function based on perceptual criteria. We implement our approach in the eigenspectral domain and compare our results with those from selected eigenspectral-based transfer functions.
Vinesh Bhunjun, Mike Brookes
ICASSP (4)2
2007 Resolving Near-Carrier Spectral Infinities Due to 1/f Phase Noise in Oscillators
abstract
In this paper, we derive an expression for the near-carrier power spectral density of an oscillator having 1/f phase noise. Motivated by empirical metrics such as the Allan variance, we develop a rigorous mathematical analysis and derive a closed-form expression for the oscillator autocorrelation function in the case of exactly 1/f phase noise that is smoothed using a rectangular time window. We show that this smoothed 1/f phase noise results in a finite variance noise process and preserves oscillator stationarity. Furthermore, in agreement with experimental data, we explain how a quadratic and a logarithmic term appear in the autocorrelation function and establish the relationship between the logarithmic term and the 1/f characteristics of the oscillator random process.
Arsenia Chorti, Mike Brookes
ICASSP (3)2
2007 The Effect of Calibration Errors on Source Localization with Microphone Arrays
abstract
Source localization employing time-differences-of-arrival has been employed for many applications. The accuracy of source localization is limited by the errors in the time differences of arrival estimation as well as microphone position calibration errors. Because a microphone position error will affect multiple time differences of arrival, correlation between these quantities will be introduced. This work presents a new mathematical framework in which we quantify the localization performance of a microphone array in which the microphone positions are subject to such errors.
Andy W. H. Khong, Mike Brookes
ICASSP (1)2
2007 Misalignment Performance of Selective Tap Adaptive Algorithms for System Identification of Time-Varying Unknown Systems
abstract
Selective tap algorithms have been proposed as a means of reducing complexity for adaptive filtering. MMax tap selection has been employed in many algorithms due to its straightforward implementation. This paper formulates the analysis of two MMax-based algorithms under time-varying unknown system conditions as are often found in practical applications. The steady-state misalignment for the MMax normalized least mean square and the MMax recursive least squares algorithms are derived and their performance is compared to that of their respective full-update algorithms. The tradeoff between computational complexity and misalignment performance is also shown for the MMax normalized least mean square case.
Patrick A. Naylor, Andy W. H. Khong, Mike Brookes
ICASSP (1)3
2007 Estimation of Glottal Closure Instants in Voiced Speech Using the DYPSA Algorithm
abstract
We present the Dynamic Programming Projected Phase-Slope Algorithm (DYPSA) for automatic estimation of glottal closure instants (GCIs) in voiced speech. Accurate estimation of GCIs is an important tool that can be applied to a wide range of speech processing tasks including speech analysis, synthesis and coding. DYPSA is automatic and operates using the speech signal alone without the need for an EGG signal. The algorithm employs the phase-slope function and a novel phase-slope projection technique for estimating GCI candidates from the speech signal. The most likely candidates are then selected using a dynamic programming technique to minimize a cost function that we define. We review and evaluate three existing methods of GCI estimation and compare the new DYPSA algorithm to them. Results are presented for the APLAWD and SAM databases for which 95.7% and 93.1% of GCIs are correctly identified
Patrick A. Naylor, Anastasis Kounoudes, Jón Guðnason, Mike Brookes
IEEE Trans. Speech Audio Process.4
2006 Adaptive algorithms for sparse echo cancellation
Patrick A. Naylor, Jingjing Cui 0005, Mike Brookes
Signal Process.3
2006 A quantitative assessment of group delay methods for identifying glottal closures in voiced speech
abstract
Measures based on the group delay of the LPC residual have been used by a number of authors to identify the time instants of glottal closure in voiced speech. In this paper, we discuss the theoretical properties of three such measures and we also present a new measure having useful properties. We give a quantitative assessment of each measure's ability to detect glottal closure instants evaluated using a speech database that includes a direct measurement of glottal activity from a Laryngograph/EGG signal. We find that when using a fixed-length analysis window, the best measures can detect the instant of glottal closure in 97% of larynx cycles with a standard deviation of 0.6 ms and that in 9% of these cycles an additional excitation instant is found that normally corresponds to glottal opening. We show that some improvement in detection rate may be obtained if the analysis window length is adapted to the speech pitch. If the measures are applied to the preemphasized speech instead of to the LPC residual, we find that the timing accuracy worsens but the detection rate improves slightly. We assess the computational cost of evaluating the measures and we present new recursive algorithms that give a substantial reduction in computation in all cases.
Mike Brookes, Patrick A. Naylor, Jón Guðnason
IEEE Trans. Speech Audio Process.1
2005 Automatic recognition of MSTAR targets using radar shadow and superresolution features
abstract
Automatic target recognition from high range resolution radar profiles remains an important and challenging problem. In this paper, we present a novel feature set for this task that combines a representation of the target's radar shadow with a noise-robust superresolution characterisation of the target scattering centres derived from the MUSIC algorithm. Using an HMM to represent aspect dependence, we demonstrate that the inclusion of the shadow features results in a significant improvement in recognition performance. We evaluate our proposed feature set on a closed-set identification task using targets from the MSTAR database and show that it results in lower recognition error rates than previously published methods using the same data.
Jingjing Cui 0005, Jón Guðnason, Mike Brookes
ICASSP (5)3
2004 Multiple Light Source Detection
abstract
This paper presents the V2R algorithm, a novel method for multiple light source detection using a Lambertian sphere as a calibration object. The algorithm segments the image of the sphere into regions that are each illuminated by a single virtual light and subtracts the virtual lights of adjacent regions to estimate the light source vectors. The algorithm uses all pixels within a region to form a robust estimate of the corresponding virtual light. The circumstances under which the light source detection problem lacks a unique solution are discussed in detail and the way in which the V2R algorithm resolves the ambiguity is explained. The V2R algorithm includes novel procedures for identifying the critical lines that bound the regions, for estimating the light source vectors, and for identifying opposite light pairs. Experiments are performed on synthetic and real images and the performance of the V2R algorithm is compared to that of a recent algorithm from the literature. The experimental results demonstrate that the proposed algorithm is robust and that it gives substantially improved accuracy.
Christos-Savvas Bouganis, Mike Brookes
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Class-based Multiple Light Detection: An Application to Faces
abstract
Multiple light detection approaches have limited applicability in real life scenarios due to the need of certain calibration objects. We propose a novel approach, the “class-based ” image-based multiple light detection, that relaxes the above assumption to a “class ” of calibration objects. We formulate it as follows: Given a set of images of objects belonging to the same class, similar 3D shape and reflectance properties, and illuminated under point light sources, the purpose is to determine the light distribution of a new object of that class. This paper concentrates on the class of human faces. Six algorithms are proposed and their performance is evaluated with real images. Experiments show that a good performance is achieved for up to three lights using a small database of faces. 1
Christos-Savvas Bouganis, Mike Brookes
BMVC2
2003 Robust multi-body segmentation
abstract
Good correspondences are a key to a correct 3D reconstruction of a scene, especially in the presence of multiple independent objects. In this paper a novel segmentation algorithm is presented that decouples the outlier rejection from the object segmentation. It is shown that the proposed outlier rejection scheme provides a dense set of correspondences across the image and eliminates gross outliers. These correspondences are subsequently used for the segmentation of objects by enforcing constraints of rigid motion. Simple additional constraints are incorporated in the segmentation process to ensure further stability under non-optimal conditions. The algorithm requires only two pictures of the scene and most of the computation can be parallelised easily, making the algorithm highly suitable for real-time hardware processing. The performance of the algorithm is illustrated on real images and the strengths and weaknesses of the approach are discussed. 1
Andreas Dante, Mike Brookes, Anthony G. Constantinides
BMVC2
2003 Soft decisions for DQPSK demodulation for the Viterbi decoding of the convolutional codes
abstract
The conventional soft decision algorithm for DQPSK uses only the differential angle between consecutive DQPSK symbols. However it is possible to improve the accuracy of the soft decision bits by taking the amplitude information of the two DQPSK symbols into account. This paper introduces novel soft decision algorithms based on this approach which give a performance improvement compared to the conventional methods equivalent to an SNR gain of up to 5 dB.
Thushara C. Hewavithana, Mike Brookes
ICASSP (4)2
2003 Precise real-time outlier removal from motion vector fields for 3D reconstruction
abstract
Finding the correct correspondences in an image sequence is a significant task for deriving 3D structure from motion. Most research has concentrated on extracting and matching salient feature points for correspondence. Block-matching has largely been disregarded due to its significant number of correspondence-outliers and its complexity. However, nowadays real-time hardware is available to obtain block-motion vectors. We present a fast method to filter out more than 99.7% of all outliers and show that the obtained correspondences can be used to derive the 3D scene depth of real image sequences.
Andreas Dante, Mike Brookes
ICIP (1)2
2002 Distribution based classification using Gaussian Mixture Models
abstract
A central task in classification is a measure of similarity between a dataset and a class that is characterised by a probability density function. The Bhattacharyya distance and the Kullback-Liebler divergence measure have been successful in comparing two multivariate normal density functions but their use is impracticable when the data is modelled using complex distributions such as Gaussian Mixture Models. The similarity is computed by combining the Bhattacharyya distances between corresponding mixtures in the reference and the test data model. In this paper we compare the performance of the Likelihood Ratio Test to a novel technique that defines a similarity measure between data and reference models having Gaussian Mixture probability density functions. When fitting a Gaussian Mixture Model to the test dataset our procedure ensures a one to one correspondence between the mixtures of the dataset and those of the reference model. This procedure has been tested using experiments, with both synthetic data and a Speaker Verification evaluation database. The performance was assessed using Detection Error Trade-off curves and demonstrates that the new measure performs significantly better than Likelihood Ratio Test.
Jón Guðnason, Mike Brookes
ICASSP2
2002 The DYPSA algorithm for estimation of glottal closure instants in voiced speech
abstract
We present the DYPSA algorithm for automatic and reliable estimation of glottal closure instants (GCIs) in voiced speech. Reliable GCI estimation is essential for closed-phase speech analysis, from which can be derived features of the vocal tract and, separately, the voice source. It has been shown that such features can be used with significant advantages in applications such as speaker recognition. DYPSA is automatic and operates using the speech signal alone without the need for an EGG or Laryngograph signal. It incorporates a new technique for estimating GCI candidates and employs dynamic programming to select the most likely candidates according to a defined cost function. We review and evaluate three existing methods and compare our new algorithm to them. Results for DYPSA show GCI detection accuracy to within ±0.25ms on 87% of the test database and fewer than 1% false alarms and misses.
Anastasis Kounoudes, Patrick A. Naylor, Mike Brookes
ICASSP3
1999 Modelling energy flow in the vocal tract with applications to glottal closure and opening detection
abstract
The pitch-synchronous analysis that is used in several areas of speech processing often requires robust detection of the instants of glottal closure and opening. In this paper we derive expressions for the flow of acoustic energy in the lossless-tube model of the vocal tract and show how linear predictive analysis may be used to estimate the waveform of acoustic input power at the glottis. We demonstrate that this signal may be used to identify the instants of glottal closure and opening during voiced speech and contrast it with the LPC residual signal that previous authors have used for this purpose.
Mike Brookes, Han Pin Loke
ICASSP1