VLDB 2026 Research / reviewers in the wild / expert
Trevor J. Cox
dblp:67/3034
· DBLP profile ↗
20ranked-venue papers
1as first author
6since 2021 · last 2027
0000-0002-4075-7564ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | The first Clarity Enhancement Challenge: Developing hearing aid algorithms for speech-in-noiseabstractHearing aid users frequently struggle to understand speech in noisy environments, negatively impacting their quality of life. Inspired by recent progress in speech technology through community-driven machine learning challenges, the Clarity project launched the first-ever Clarity Enhancement Challenge (CEC1). This challenge specifically addressed speech-in-noise enhancement for hearing aids, uniquely combining objective and subjective intelligibility evaluations to assess performance. Participants developed algorithms aimed at improving speech intelligibility in a simulated domestic environment featuring a target speaker and a stationary noise interferer—either competing speech or domestic appliances. Competitors were provided with an open-source dataset, comprising a novel 40-speaker British English corpus, realistic domestic noise samples, and a baseline hearing aid model with basic signal processing. This paper describes the design and outcomes of CEC1. Thirteen entries were evaluated objectively using the Modified Binaural Short-Time Objective Intelligibility metric (MBSTOI) and subjectively by a listening panel of hearing-impaired individuals. The majority of systems employed deep neural networks (DNNs), classical beamforming, or a combination of both. Results showed significant intelligibility gains over the baseline, particularly for systems combining adaptive beamforming with neural network-based noise reduction. However, algorithms optimised directly for MBSTOI scores did not always translate to real-world listening benefits, highlighting the critical importance of perceptual evaluation in assessing intelligibility. These findings underscore the potential of machine-learning-driven approaches for enhancing hearing aid performance and set a foundation for future challenges addressing dynamic and more realistic auditory scenarios. Simone Graetzer, Michael A. Akeroyd, Jon Barker, Trevor J. Cox, John F. Culling, Jennifer Firth, Graham Naylor, Eszter Porter, Rhoddy Viveros Muñoz |
Comput. Speech Lang. | 4 |
| 2024 | The 2nd Clarity Prediction Challenge: A Machine Learning Challenge for Hearing Aid Intelligibility PredictionabstractThis paper reports on the design and outcomes of the 2nd Clarity Prediction Challenge (CPC2) for predicting the intelligibility of hearing aid processed signals heard by individuals with a hearing impairment. The challenge was designed to promote new approaches for estimating the intelligibility of hearing aid signals that can be used in future hearing aid algorithm development. It extends an earlier round (CPC1, 2022) in a number of critical directions, including a larger dataset coming from new speech intelligibility listening experiments, a greater degree of variability in the test materials, and a design that requires prediction systems to generalise to unseen algorithms and listeners. This paper provides a full description of the new publicly available CPC2 dataset, the CPC2 challenge design, and the baseline systems. The challenge attracted 12 systems from 9 research teams. The systems are reviewed, their performance is analysed and conclusions are presented, with reference to the progress made since the earlier CPC1 challenge. In particular, it is seen how reference-free, non-intrusive systems based on pre-trained large acoustic models can perform well in this context. Jon Barker, Michael A. Akeroyd, Will Bailey, Trevor J. Cox, John F. Culling, Jennifer Firth, Simone Graetzer, Graham Naylor |
ICASSP | 4 |
| 2023 | The 2nd Clarity Enhancement Challenge for Hearing Aid Speech Intelligibility Enhancement: Overview and OutcomesabstractThis paper reports on the design and outcomes of the 2nd Clarity Enhancement Challenge (CEC2), a challenge for stimulating novel approaches to hearing-aid speech intelligibility enhancement. The challenge was for a listener attending to a target speaker in a noisy, domestic environment. The challenge extends the previous edition, CEC1, in a number of key respects: scenes have multiple interferers including speech, noise and music; ambisonics are used to model listener head movement; target speaker identity is provided to encourage speaker extraction approaches. Systems are evaluated both via the HASPI intelligibility metric and with listening tests using a panel of hearing-impaired listeners. The paper reviews the 18 systems that were submitted describing them in terms of their enhancement and amplification stages. HASPI is seen to be a good predictor of listener performance. The top system, using carefully engineered neural approaches, produces highly intelligible signals for complex scenes with SNRs down to -12 dB while obeying the challenges 5 ms latency constraint. Michael A. Akeroyd, Will Bailey, Jon Barker, Trevor J. Cox, John F. Culling, Simone Graetzer, Graham Naylor, Zuzanna Podwinska, Zehai Tu |
ICASSP | 4 |
| 2023 | Overview of the 2023 ICASSP SP Clarity Challenge: Speech Enhancement for Hearing AidsabstractThis paper reports on the design and outcomes of the ICASSP SP Clarity Challenge: Speech Enhancement for Hearing Aids. The scenario was a listener attending to a target speaker in a noisy, domestic environment. There were multiple interferers and head rotation by the listener. The challenge extended the second Clarity Enhancement Challenge (CEC2) by fixing the amplification stage of the hearing aid; evaluating with a combined metric for speech intelligibility and quality; and providing two evaluation sets, one based on simulation and the other on real-room measurements. Five teams improved on the baseline system for the simulated evaluation set, but the performance on the measured evaluation set was much poorer. Investigations are on-going to determine the exact cause of the mismatch between the simulated and measured data sets. The presence of transducer noise in the measurements, lower order Ambisonics harming the ability for systems to exploit binaural cues and the differences between real and simulated room impulse responses are suggested causes. Trevor J. Cox, Jon Barker, Will Bailey, Simone Graetzer, Michael A. Akeroyd, John F. Culling, Graham Naylor |
ICASSP | 1 |
| 2022 | The 1st Clarity Prediction Challenge: A machine learning challenge for hearing aid intelligibility prediction
Jon Barker, Michael A. Akeroyd, Trevor J. Cox, John F. Culling, Jennifer Firth, Simone Graetzer, Holly Griffiths, Lara Harris, Graham Naylor, Zuzanna Podwinska, Eszter Porter, Rhoddy Viveros Muñoz |
INTERSPEECH | 3 |
| 2021 | Clarity-2021 Challenges: Machine Learning Challenges for Advancing Hearing Aid ProcessingabstractIn recent years, rapid advances in speech technology have been made possible by machine learning challenges such as CHiME, REVERB, Blizzard, and Hurricane. In the Clarity project, the machine learning approach is applied to the problem of hearing aid processing of speech-in-noise, where current technology in enhancing the speech signal for the hearing aid wearer is often ineffective. The scenario is a (simulated) cuboid-shaped living room in which there is a single listener, a single target speaker and a single interferer, which is either a competing talker or domestic noise. All sources are static, the target is always within ±30° azimuth of the listener and at the same elevation, and the interferer is an omnidirectional point source at the same elevation. The target speech comes from an open source 40-speaker British English speech database collected for this purpose. This paper provides a baseline description of the round one Clarity challenges for both enhancement (CEC1) and prediction (CPC1). To the authors’ knowledge, these are the first machine learning challenges to consider the problem of hearing aid speech signal processing. Simone Graetzer, Jon Barker, Trevor J. Cox, Michael A. Akeroyd, John F. Culling, Graham Naylor, Eszter Porter, Rhoddy Viveros Muñoz |
Interspeech | 3 |
| 2019 | Generalisation in Environmental Sound Classification: The 'Making Sense of Sounds' Data Set and ChallengeabstractHumans are able to identify a large number of environmental sounds and categorise them according to high-level semantic categories, e.g. urban sounds or music. They are also capable of generalising from past experience to new sounds when applying these categories. In this paper we report on the creation of a data set that is structured according to the top-level of a taxonomy derived from human judgements and the design of an associated machine learning challenge, in which strong generalisation abilities are required to be successful. We introduce a baseline classification system, a deep convolutional network, which showed strong performance with an average accuracy on the evaluation data of 80.8%. The result is discussed in the light of two alternative explanations: An unlikely accidental category bias in the sound recordings or a more plausible true acoustic grounding of the high-level categories. Christian Kroos, Oliver Bones, Yin Cao, Lara Harris, Philip J. B. Jackson, William J. Davies, Wenwu Wang 0001, Trevor J. Cox, Mark D. Plumbley |
ICASSP | 8 |
| 2019 | Background Adaptation for Improved Listening Experience in BroadcastingabstractThe intelligibility of speech in noise can be improved by modifying the speech. But with object-based audio, there is the possibility of altering the background sound while leaving the speech unaltered. This may prove a less intrusive approach, affording good speech intelligibility without overly compromising the perceived sound quality. In this study, the technique of spectral weighting was applied to the background. The frequency-dependent weightings for adaptation were learnt by maximising a weighted combination of two perceptual objective metrics for speech intelligibility and audio quality. The balance between the two objective metrics was determined by the perceptual relationship between intelligibility and quality. A neural network was trained to provide a fast solution for real-time processing. Tested in a variety of background sounds and speech-to-background ratios (SBRs), the proposed method led to a large intelligibility gain over the unprocessed baseline. Compared to an approach using constant weightings, the proposed method was able to dynamically preserve the overall audio quality better with respect to SBR changes. Trevor J. Cox, Bruno Fazenda, Qingju Liu, Wenwu Wang 0001 |
ICASSP | 2 |
| 2018 | A non-intrusive method for estimating binaural speech intelligibility from noise-corrupted signals captured by a pair of microphonesabstractA non-intrusive method is introduced to predict binaural speech intelligibility in noise directly from signals captured using a pair of microphones. The approach combines signal processing techniques in blind source separation and localisation, with an intrusive objective intelligibility measure (OIM). Therefore, unlike classic intrusive OIMs, this method does not require a clean reference speech signal and knowing the location of the sources to operate. The proposed approach is able to estimate intelligibility in stationary and fluctuating noises, when the noise masker is presented as a point or diffused source, and is spatially separated from the target speech source on a horizontal plane. The performance of the proposed method was evaluated in two rooms. When predicting subjective intelligibility measured as word recognition rate, this method showed reasonable predictive accuracy with correlation coefficients above 0.82, which is comparable to that of a reference intrusive OIM in most of the conditions. The proposed approach offers a solution for fast binaural intelligibility prediction, and therefore has practical potential to be deployed in situations where on-site speech intelligibility is a concern. Qingju Liu, Wenwu Wang 0001, Trevor J. Cox |
Speech Commun. | 4 |
| 2018 | An Audio-Visual System for Object-Based Audio: From Recording to ListeningabstractObject-based audio is an emerging representation for audio content, where content is represented in a reproduction-format-agnostic way and, thus, produced once for consumption on many different kinds of devices. This affords new opportunities for immersive, personalized, and interactive listening experiences. This paper introduces an end-to-end object-based spatial audio pipeline, from sound recording to listening. A high-level system architecture is proposed, which includes novel audio-visual interfaces to support object-based capture and listener-tracked rendering, and incorporates a proposed component for objectification, that is, recording content directly into an object-based form. Text-based and extensible metadata enable communication between the system components. An open architecture for object rendering is also proposed. The system's capabilities are evaluated in two parts. First, listener-tracked reproduction of metadata automatically estimated from two moving talkers is evaluated using an objective binaural localization model. Second, object-based scene capture with audio extracted using blind source separation (to remix between two talkers) and beamforming (to remix a recording of a jazz group) is evaluated with perceptually motivated objective and subjective experiments. These experiments demonstrate that the novel components of the system add capabilities beyond the state of the art. Finally, we discuss challenges and future perspectives for object-based audio workflows. Philip Coleman, Andreas Franck, Jon Francombe, Qingju Liu, Teófilo Emídio de Campos, Richard J. Hughes 0001, Dylan Menzies, Marcos F. Simón Gálvez, James Woodcock, Philip J. B. Jackson, Frank Melchior, Chris Pike, Filippo Maria Fazi, Trevor J. Cox, Adrian Hilton 0001 |
IEEE Trans. Multim. | 15 |
| 2016 | Perception and automated assessment of audio quality in user generated content: An improved modelabstractTechnology to record sound, available in personal devices such as smartphones or video recording devices, is now ubiquitous. However, the production quality of the sound on this user-generated content is often very poor: distorted, noisy, with garbled speech or indistinct music. Our interest lies in the causes of the poor recording, especially what happens between the sound source and the electronic signal emerging from the microphone, and finding an automated method to warn the user of such problems. Typical problems, such as distortion, wind noise, microphone handling noise and frequency response, were tested. A perceptual model has been developed from subjective tests on the perceived quality of such errors and data measured from a training dataset composed of various audio files. It is shown that perceived quality is associated with distortion and frequency response, with wind and handling noise being just slightly less important. In addition, the contextual content of the audio sample was found to modulate perceived quality at similar levels to degradations such as wind and rendering those introduced by handling noise negligible. Bruno Fazenda, Paul Kendrick, Trevor J. Cox, Francis F. Li, Iain Jackson |
QoMEX | 3 |
| 2016 | Evaluating a distortion-weighted glimpsing metric for predicting binaural speech intelligibility in roomsabstractA distortion-weighted glimpse proportion metric (BiDWGP) for predicting binaural speech intelligibility were evaluated in simulated anechoic and reverberant conditions, with and without a noise masker. The predictive performance of BiDWGP was compared to four reference binaural intelligibility metrics, which were extended from the Speech Intelligibility Index (SII) and the Speech Transmission Index (STI). In the anechoic sound field, BiDWGP demonstrated high accuracy in predicting binaural intelligibility for individual maskers (ρ ≥ 0.95) and across maskers (ρ ≥ 0.94). The reference metrics however performed less well in across-masker prediction (0.54 ≤ ρ ≤ 0.86) despite reasonable accuracy for individual maskers. In reverberant rooms, BiDWGP was more stable in all test conditions (ρ ≥ 0.87) than the reference metrics, which showed different predictive patterns: the binaural STIs were more robust for the stationary than for the fluctuating noise masker, whilst the binaural SII displayed the opposite behaviour. The study shows that the new BiDWGP metric can provide similar or even more robust predictive power than the current standard metrics. Richard J. Hughes 0001, Bruno Fazenda, Trevor J. Cox |
Speech Commun. | 4 |
| 2015 | A glimpse-based approach for predicting binaural intelligibility with single and multiple maskers in anechoic conditionsabstractA distortion-weighted glimpsing metric developed for estimating monaural speech intelligibility is extended to predict binaural speech intelligibility in noise. Two aspects of binaural listen- ing, the better ear effect and the binaural advantage, are taken into account in the new metric, which predicts intelligibility using monaural target and masker signals and their location, and is therefore able to provide intelligibility estimates in situations where binaural signals are not readily available. Perceptual listening experiments were conducted to evaluate the predictive power of the proposed metric for speech in the presence of single and multiple maskers in anechoic conditions, for a range of source/masker azimuth combinations. The binaural metric is highly correlated (ρ > 0.9) with listeners’ performance in all conditions tested, but overestimates intelligibility somewhat in conditions where multiple maskers are present and the target speech source location is unknown. Martin Cooke, Bruno Fazenda, Trevor J. Cox |
INTERSPEECH | 4 |
| 2013 | Wind-induced microphone noise detection - automatically monitoring the audio quality of field recordingsabstractWind-induced microphone noise is one of the most common problems leading to poor audio quality in recordings. A wind-noise detector could alert the operator of a recording device to the presence of wind noise so that appropriate action can be taken. This paper presents a single channel algorithm which, within the presence of other sounds, detects and classifies wind noise according to level. A large training database is formed from a wind noise simulator which generates an audio stream based on time histories of real wind velocities. A Support Vector Machine detects and classifies according to wind noise level in 25 ms frames which may contain other sounds. Statistical and temporal data from the detector over a sequence of frames is then used to provide estimates for the average wind noise level. The detector is successfully demonstrated on a number of devices with non-simulated data. Paul Kendrick, Trevor J. Cox, Francis F. Li, Bruno Fazenda, Iain Jackson |
ICME | 2 |
| 2008 | A combined blind source separation and adaptive noise cancellation scheme with potential application in blind acoustic parameter extraction
Yonggang Zhang 0001, Jonathon A. Chambers, Paul Kendrick, Trevor J. Cox, Francis F. Li |
Neurocomputing | 4 |
| 2007 | A New Variable Step-Size LMS Algorithm with Robustness to Nonstationary NoiseabstractA new variable step-size least-mean-square (VSSLMS) algorithm is presented in this paper for applications in which the desired response contains nonstationary noise with high variance. The step size of the proposed VSSLMS algorithm is controlled by the normalized square Euclidean norm of the averaged gradient vector, and is henceforth referred to as the NSVSSLMS algorithm. As shown by the analysis and simulation results, the proposed algorithm has both fast convergence rate and robustness to high-variance noise signals, and performs better than Greenburg's sum method, which is a robust algorithm for applications with nonstationary noise. Yonggang Zhang 0001, Jonathon A. Chambers, Wenwu Wang 0001, Paul Kendrick, Trevor J. Cox |
ICASSP (3) | 5 |
| 2007 | A New Variable Tap-Length LMS Algorithm to Model an Exponential Decay Impulse ResponseabstractThis letter proposes a new variable tap-length least-mean-square (LMS) algorithm for applications in which the unknown filter impulse response sequence has an exponential decay envelope. The algorithm is designed to minimize the mean-square deviation (MSD) between the optimal and adaptive filter weight vectors at each iteration. Simulation results show the proposed algorithm has a faster convergence rate as compared with the fixed tap-length LMS algorithm and is robust to the initial tap-length choice. Yonggang Zhang 0001, Jonathon A. Chambers, Saeid Sanei, Paul Kendrick, Trevor J. Cox |
IEEE Signal Process. Lett. | 5 |
| 2006 | Room Acoustic Parameter Extraction from Music SignalsabstractA new method, employing machine learning techniques and a modified low frequency envelope spectrum estimator, for estimating important room acoustic parameters including Reverberation Time (RT) and Early Decay Time (EDT) from received music signals has been developed. It overcomes drawbacks found in applying music signals directly to the envelope spectrum detector developed for the estimation of RT from speech signals. The octave band music signal is first separated into sub bands corresponding to notes on the equal temperament scale and the level of each note normalised before applying an envelope spectrum detector. A typical artificial neural network is then trained to map these envelope spectra onto RT or EDT. Significant improvements in estimation accuracy were found and further investigations confirmed that the non-stationary nature of music envelopes is a major technical challenge hindering accurate parameter extraction from music and the proposed method to some extent circumvents the difficulty. Paul Kendrick, Trevor J. Cox, Yonggang Zhang 0001, Jonathon A. Chambers, Francis F. Li |
ICASSP (5) | 2 |
| 2006 | Acoustic Parameter Extraction from Occupied Rooms Utilizing Blind Source Separation
Yonggang Zhang 0001, Jonathon A. Chambers, Paul Kendrick, Trevor J. Cox, Francis F. Li |
KES (3) | 4 |
| 2003 | A neural network for blind identification of speech transmission indexabstractA hybrid neural network model is proposed to determine the speech transmission index of a transmission channel from transmitted speech signals without resort to prior knowledge of original speech. It comprises a Hilbert transform pre-processor, a PCA network for speech feature extraction and a multilayer back-propagation network for nonlinear mapping and case generalization. The developed method utilizes naturally occurring speech signals as probe stimuli, reduces measurement channels from two to one and hence facilitates speech transmission channel assessments under in-use conditions. Francis F. Li, Trevor J. Cox |
ICASSP (2) | 2 |