EDBT 2026 Demo / reviewers in the wild / expert
Zehai Tu
dblp:277/4011
· DBLP profile ↗
11ranked-venue papers
6as first author
10since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Can We "Cherry-Pick"? Investigating Multiple Renditions from a Generative Speech Synthesis ModelabstractGenerative Speech Models (GSMs) have seen a surge in popularity due to their ability to generate diverse and high-quality speech. Evaluating models that generate many different renditions for a given input sentence presents a new challenge. Listening tests are still the gold standard for evaluating synthetic speech, but current paradigms only consider a single arbitrary rendition: this fails to give a complete picture of best/typical/worst-rendition performance. We propose a general framework for evaluating and deploying generative speech models. This involves selecting amongst renditions using a sequence of filtering or ranking steps, each using either an objective or subjective (listening) method. The framework is not tied to a particular generative model, and so could be applied to any such model. In this paper, we provide a demonstration of a simple version of this framework which would apply to use-cases where best-rendition performance matters. We explore the concept of "cherry-picking", and ask the question "Is there a rendition that is consistently preferred above all others by listeners?". In a subjective listening test, participants ranked several renditions of the same sentence, from which we measured the prevalence of exceptional renditions. We find that there is indeed a preferred rendition in many, but not all cases. Our framework is flexible. In particular, the use of listeners is optional. In future, they could be replaced with model-based objective measures, for example. Adaeze Adigwe, Sarenne Wallbridge, Zehai Tu, Catherine Lai |
ICASSP | 3 |
| 2025 | Enabling Beam Search for Language Model-Based Text-to-Speech SynthesisabstractTokenising continuous speech into sequences of discrete tokens and modelling them with language models (LMs) has led to significant success in text-to-speech (TTS) synthesis. Despite these models can generate speech with high quality and naturalness, their synthesised samples can still suffer from artefacts, mispronunciation, word repeating, etc. In this paper, we argue these undesirable properties could partly be caused by the randomness of sampling-based strategies during the autoregressive decoding of LMs. Therefore, we look at maximization-based decoding approaches and propose Temporal Repetition Aware Diverse Beam Search (TRAD-BS) to find the most probable sequences of the generated speech tokens. Experiments with two recent LM-based TTS models demonstrate that our proposed maximisation-based decoding strategy generates speech with fewer mispronunciations and improved speaker consistency1. Zehai Tu, Guangyan Zhang, Yiting Lu, Adaeze Adigwe, Yiwen Guo |
ICASSP | 1 |
| 2025 | Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching ModelabstractThis paper introduces PFlow-VC, a conditional flow matching voice conversion model that leverages fine-grained discrete pitch tokens and target speaker prompt information for expressive voice conversion (VC). Previous VC works primarily focus on speaker conversion, with further exploration needed in enhancing expressiveness (such as prosody and emotion) for timbre conversion. Unlike previous methods, we adopt a simple and efficient approach to enhance the style expressiveness of voice conversion models. Specifically, we pretrain a self-supervised pitch VQVAE model to discretize speaker-irrelevant pitch information and leverage a masked pitch-conditioned flow matching model for Mel-spectrogram synthesis, which provides in-context pitch modeling capabilities for the speaker conversion model, effectively improving the voice style transfer capacity. Additionally, we improve timbre similarity by combining global timbre embeddings with time-varying timbre tokens. Experiments on unseen LibriTTS test-clean and emotional speech dataset ESD show the superiority of the PFlow-VC model in both timbre conversion and style transfer. Audio samples are available on the demo page https://speechai-demo.github.io/PFlow-VC/. Jialong Zuo, Shengpeng Ji, Minghui Fang 0002, Ziyue Jiang 0001, Xize Cheng, Qian Yang 0006, Wenrui Liu 0003, Guangyan Zhang, Zehai Tu, Yiwen Guo, Zhou Zhao 0001 |
ICASSP | 9 |
| 2024 | Energy-Based Models for Speech SynthesisabstractRecently there has been a lot of interest in non-autoregressive (non-AR) models for speech synthesis, such as FastSpeech 2 and diffusion models. Unlike AR models, these models do not have autoregressive dependencies among outputs which makes inference efficient. This paper expands the range of available non-AR models with another member called energy-based models (EBMs). The paper describes how noise contrastive estimation, which relies on the comparison between positive and negative samples, can be used to train EBMs. It proposes a number of strategies for generating effective negative samples, including using high-performing AR models. It also describes how sampling from EBMs can be performed using Langevin Markov Chain Monte-Carlo (MCMC). The use of Langevin MCMC enables to draw connections between EBMs and currently popular diffusion models. Experiments on LJSpeech dataset show that the proposed approach offers improvements over Tacotron 2. Wanli Sun, Zehai Tu, Anton Ragni |
ICASSP | 2 |
| 2023 | The 2nd Clarity Enhancement Challenge for Hearing Aid Speech Intelligibility Enhancement: Overview and OutcomesabstractThis paper reports on the design and outcomes of the 2nd Clarity Enhancement Challenge (CEC2), a challenge for stimulating novel approaches to hearing-aid speech intelligibility enhancement. The challenge was for a listener attending to a target speaker in a noisy, domestic environment. The challenge extends the previous edition, CEC1, in a number of key respects: scenes have multiple interferers including speech, noise and music; ambisonics are used to model listener head movement; target speaker identity is provided to encourage speaker extraction approaches. Systems are evaluated both via the HASPI intelligibility metric and with listening tests using a panel of hearing-impaired listeners. The paper reviews the 18 systems that were submitted describing them in terms of their enhancement and amplification stages. HASPI is seen to be a good predictor of listener performance. The top system, using carefully engineered neural approaches, produces highly intelligible signals for complex scenes with SNRs down to -12 dB while obeying the challenges 5 ms latency constraint. Michael A. Akeroyd, Will Bailey, Jon Barker, Trevor J. Cox, John F. Culling, Simone Graetzer, Graham Naylor, Zuzanna Podwinska, Zehai Tu |
ICASSP | 9 |
| 2022 | Auditory-Based Data Augmentation for end-to-end Automatic Speech RecognitionabstractEnd-to-end models have achieved significant improvement on automatic speech recognition. One common method to improve performance of these models is expanding the data-space through data augmentation. Meanwhile, human auditory inspired front-ends have also demonstrated improvement for automatic speech recognisers. In this work, a well-verified auditory-based model, which can simulate various hearing abilities, is investigated for the purpose of data augmentation for end-to-end speech recognition. By introducing the auditory model into the data augmentation process, end-to-end systems are encouraged to ignore variation from the signal that cannot be heard and thereby focus on robust features for speech recognition. Two mechanisms in the auditory model, spectral smearing and loudness recruitment, are studied on the LibriSpeech dataset with a transformer-based end-to-end model. The results show that the proposed augmentation methods can bring statistically significant improvement on the performance of the state-of-the-art SpecAugment. Zehai Tu, Jack Deadman, Ning Ma 0002, Jon Barker |
ICASSP | 1 |
| 2022 | Exploiting Hidden Representations from a DNN-based Speech Recogniser for Speech Intelligibility Prediction in Hearing-impaired ListenersabstractAn accurate objective speech intelligibility prediction algorithms is of great interest for many applications such as speech enhancement for hearing aids.Most algorithms measures the signal-to-noise ratios or correlations between the acoustic features of clean reference signals and degraded signals.However, these hand-picked acoustic features are usually not explicitly correlated with recognition.Meanwhile, deep neural network (DNN) based automatic speech recogniser (ASR) is approaching human performance in some speech recognition tasks.This work leverages the hidden representations from DNN-based ASR as features for speech intelligibility prediction in hearingimpaired listeners.The experiments based on a hearing aid intelligibility database show that the proposed method could make better prediction than a widely used short-time objective intelligibility (STOI) based binaural measure. Zehai Tu, Ning Ma 0002, Jon Barker |
INTERSPEECH | 1 |
| 2022 | Unsupervised Uncertainty Measures of Automatic Speech Recognition for Non-intrusive Speech Intelligibility PredictionabstractNon-intrusive intelligibility prediction is important for its application in realistic scenarios, where a clean reference signal is difficult to access. The construction of many non-intrusive predictors require either ground truth intelligibility labels or clean reference signals for supervised learning. In this work, we leverage an unsupervised uncertainty estimation method for predicting speech intelligibility, which does not require intelligibility labels or reference signals to train the predictor. Our experiments demonstrate that the uncertainty from state-of-the-art end-to-end automatic speech recognition (ASR) models is highly correlated with speech intelligibility. The proposed method is evaluated on two databases and the results show that the unsupervised uncertainty measures of ASR models are more correlated with speech intelligibility from listening results than the predictions made by widely used intrusive methods. Zehai Tu, Ning Ma 0002, Jon Barker |
INTERSPEECH | 1 |
| 2021 | DHASP: Differentiable Hearing Aid Speech ProcessingabstractHearing aids are expected to improve speech intelligibility for listeners with hearing impairment. An appropriate amplification fitting tuned for the listener’s hearing disability is critical for good performance. The developments of most prescriptive fittings are based on data collected in subjective listening experiments, which are usually expensive and time-consuming. In this paper, we explore an alternative approach to finding the optimal fitting by introducing a hearing aid speech processing framework, in which the fitting is optimised in an automated way using an intelligibility objective function based on the HASPI physiological auditory model. The framework is fully differentiable, thus can employ the back-propagation algorithm for efficient, data-driven optimisation. Our initial objective experiments show promising results for noise-free speech amplification, where the automatically optimised processors outperform one of the well recognised hearing aid prescriptions. Zehai Tu, Ning Ma 0002, Jon Barker |
ICASSP | 1 |
| 2021 | Optimising Hearing Aid Fittings for Speech in Noise with a Differentiable Hearing Loss ModelabstractThis is a repository copy of Optimising hearing aid fittings for speech in noise with a differentiable hearing loss model. Zehai Tu, Ning Ma 0002, Jon Barker |
Interspeech | 1 |
| 2020 | Acoustic Feature Extraction with Interpretable Deep Neural Network for Neurodegenerative Related Disorder ClassificationabstractSpeech-based automatic approaches for detecting neurodegenerative disorders (ND) and mild cognitive impairment (MCI) have received more attention recently due to being non-invasive and potentially more sensitive than current pen-and-paper tests. The performance of such systems is highly dependent on the choice of features in the classification pipeline. In particular for acoustic features, arriving at a consensus for a best feature set has proven challenging. This paper explores using deep neural network for extracting features directly from the speech signal as a solution to this. Compared with hand-crafted features, more information is present in the raw waveform, but the feature extraction process becomes more complex and less interpretable which is often undesirable in medical domains. Using a SincNet as a first layer allows for some analysis of learned features. We propose and evaluate the Sinc-CLA (with SincNet, Convolutional, Long Short-Term Memory and Attention layers) as a task-driven acoustic feature extractor for classifying MCI, ND and healthy controls (HC). Experiments are carried out on an in-house dataset. Compared with the popular hand-crafted feature sets, the learned task-driven features achieve a superior classification accuracy. The filters of the SincNet is inspected and acoustic differences between HC, MCI and ND are found. Yilin Pan, Bahman Mirheidari, Zehai Tu, Ronan O'Malley, Traci Walker, Annalena Venneri, Markus Reuber, Daniel Blackburn, Heidi Christensen |
INTERSPEECH | 3 |