Philip Harding

dblp:18/10650 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
5since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 4 since 2021
YearPublicationVenuePosition
2023 Hierarchical Attention-Based Contextual Biasing For Personalized Speech Recognition Using Neural Transducers
abstract
Although end-to-end (E2E) automatic speech recognition (ASR) systems excel in general tasks, they frequently struggle with accurately recognizing personal rare words. Leveraging contextual information to bias the internal states of E2E ASR model has proven to be an effective solution. However most existing work focuses on biasing for a single domain and it is still challenging to expand such contextualization mechanisms to many domains. To address this limitation, in this work we propose a hierarchical attention architecture to scale contextual biasing to a wide range of domains simultaneously. Given multiple catalogs of contextual information, the high-level attention determines which source of catalog to focus on and the low-level attention learns to attend to the most relevant entity within the focused catalog. Experiments on diverse domains demonstrate the proposed architecture results in $35 \%$ to $60 \%$ relative WER improvements on personal rare words and outperforms existing approaches.
Sibo Tong, Philip Harding, Simon Wiesler
ASRU2
2023 Slot-Triggered Contextual Biasing For Personalized Speech Recognition Using Neural Transducers
abstract
End-to-end (E2E) automatic speech recognition (ASR) models have been found to perform well on general transcription tasks but often fail to correctly recognize words that occur infrequently in the training data. Personalization is important for a variety of tasks, including virtual assistants where recall of infrequently observed words such as contact names, song titles and place names is critical. In these cases contextual information is often available which can be used to bias the E2E ASR model. Contextual biasing (CB) has been shown to be effective for this task, however most existing work focuses on biasing for a single domain and so in this work we focus on the application of biasing to multiple domains. We propose a method whereby the E2E ASR model is trained to emit opening and closing tags around slot content which are used to both selectively enable biasing and decide which catalog to use for biasing. Our method is shown to not only efficiently scale to multiple slots, but also further improves accuracy on slot content.
Sibo Tong, Philip Harding, Simon Wiesler
ICASSP2
2023 Selective Biasing with Trie-based Contextual Adapters for Personalised Speech Recognition using Neural Transducers
Philip Harding, Sibo Tong, Simon Wiesler
INTERSPEECH1
2023 Model-Internal Slot-triggered Biasing for Domain Expansion in Neural Transducer ASR Models
Yiting Lu, Philip Harding, Kanthashree Mysore Sathyendra, Sibo Tong, Xuandi Fu, Feng-Ju Chang, Simon Wiesler, Grant P. Strimel
INTERSPEECH2
2023 Effective Training of Attention-based Contextual Biasing Adapters with Synthetic Audio for Personalised ASR
Burin Naowarat, Philip Harding, Pasquale D'Alterio, Sibo Tong, Bashar Awwad Shiekh Hasan
INTERSPEECH2
2017 Estimating acoustic speech features in low signal-to-noise ratios using a statistical framework
Philip Harding, Ben P. Milner
Comput. Speech Lang.1
2015 Reconstruction-based speech enhancement from robust acoustic features
abstract
This paper proposes a method of speech enhancement where a clean speech signal is reconstructed from a sinusoidal model of speech production and a set of acoustic speech features. The acoustic features are estimated from noisy speech and comprise, for each frame, a voicing classification (voiced, unvoiced or non-speech), fundamental frequency (for voiced frames) and spectral envelope. Rather than using different algorithms to estimate each parameter, a single statistical model is developed. This comprises a set of acoustic models and has similarity to the acoustic modelling used in speech recognition. This allows noise and speaker adaptation to be applied to acoustic feature estimation to improve robustness. Objective and subjective tests compare reconstruction-based enhancement with other methods of enhancement and show the proposed method to be highly effective at removing noise.
Philip Harding, Ben P. Milner
Speech Commun.1
2012 Enhancing Speech by Reconstruction from Robust Acoustic Features
Philip Harding, Ben P. Milner
INTERSPEECH1
2012 On the use of Machine Learning Methods for Speech and Voicing Classification
Philip Harding, Ben P. Milner
INTERSPEECH1
2011 Speech Enhancement by Reconstruction from Cleaned Acoustic Features
Philip Harding, Ben P. Milner
INTERSPEECH1