VLDB 2026 Research / reviewers in the wild / expert
Adaeze Adigwe
dblp:222/2877 · also Adaeze O. Adigwe
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
0009-0006-2258-3181ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Can self-supervised speech models predict the perceived acceptability of prosodic variation?abstractThough producing an appropriate prosodic realisation of text is a one-to-many problem, modern speech generation often focuses on identifying the “best” or “most likely” output, overlooking acceptable variation across realisations. How listeners perceive such variation–and whether models capture it–is unaccounted for in current evaluation paradigms. In this study, we present exploratory analyses of whether self-supervised models encode acceptable prosodic variation. Using a new dataset of relative acceptability ratings across carefully controlled, high-quality synthetic utterances, we show that SSL representations contain information predictive of such judgments. By introducing a novel method for deriving probability-based uncertainty from autoregressive speech models, we examine whether this information is available in an unsupervised setting, highlighting the complexity of prosodic perception and the value of more human-centric evaluation paradigms. Sarenne Wallbridge, Adaeze Adigwe, Peter Bell 0001 |
ASRU | 2 |
| 2025 | Can We "Cherry-Pick"? Investigating Multiple Renditions from a Generative Speech Synthesis ModelabstractGenerative Speech Models (GSMs) have seen a surge in popularity due to their ability to generate diverse and high-quality speech. Evaluating models that generate many different renditions for a given input sentence presents a new challenge. Listening tests are still the gold standard for evaluating synthetic speech, but current paradigms only consider a single arbitrary rendition: this fails to give a complete picture of best/typical/worst-rendition performance. We propose a general framework for evaluating and deploying generative speech models. This involves selecting amongst renditions using a sequence of filtering or ranking steps, each using either an objective or subjective (listening) method. The framework is not tied to a particular generative model, and so could be applied to any such model. In this paper, we provide a demonstration of a simple version of this framework which would apply to use-cases where best-rendition performance matters. We explore the concept of "cherry-picking", and ask the question "Is there a rendition that is consistently preferred above all others by listeners?". In a subjective listening test, participants ranked several renditions of the same sentence, from which we measured the prevalence of exceptional renditions. We find that there is indeed a preferred rendition in many, but not all cases. Our framework is flexible. In particular, the use of listeners is optional. In future, they could be replaced with model-based objective measures, for example. Adaeze Adigwe, Sarenne Wallbridge, Zehai Tu, Catherine Lai |
ICASSP | 1 |
| 2025 | Enabling Beam Search for Language Model-Based Text-to-Speech SynthesisabstractTokenising continuous speech into sequences of discrete tokens and modelling them with language models (LMs) has led to significant success in text-to-speech (TTS) synthesis. Despite these models can generate speech with high quality and naturalness, their synthesised samples can still suffer from artefacts, mispronunciation, word repeating, etc. In this paper, we argue these undesirable properties could partly be caused by the randomness of sampling-based strategies during the autoregressive decoding of LMs. Therefore, we look at maximization-based decoding approaches and propose Temporal Repetition Aware Diverse Beam Search (TRAD-BS) to find the most probable sequences of the generated speech tokens. Experiments with two recent LM-based TTS models demonstrate that our proposed maximisation-based decoding strategy generates speech with fewer mispronunciations and improved speaker consistency1. Zehai Tu, Guangyan Zhang, Yiting Lu, Adaeze Adigwe, Yiwen Guo |
ICASSP | 4 |
| 2024 | What do people hear? Listeners' Perception of Conversational Speech
Adaeze Adigwe, Sarenne Wallbridge |
INTERSPEECH | 1 |
| 2022 | Strategies for developing a Conversational Speech Dataset for Text-To-Speech SynthesisabstractThere have been many efforts to improve the quality of speech synthesis systems in conversational AI. Although state-of-the-art systems are capable of producing natural-sounding speech, the generated speech often lacks prosodic variation and is not always suited to the task. In this paper, we examine dialogue data collection methods to use as training data for our acoustic models. We collect speech using three different setups: (1) Random read-aloud sentences; (2) Performed dialogues; (3) Semi-Spontaneous dialogues. We analyze prosodic and textual properties of the data collected in these setups and make some recommendations to collect data for speech synthesis in conversational AI settings. Adaeze Adigwe, Esther Klabbers |
INTERSPEECH | 1 |
| 2022 | Data-augmented cross-lingual synthesis in a teacher-student frameworkabstractCross-lingual synthesis can be defined as the task of letting a speaker generate fluent synthetic speech in another language.This is a challenging task, and resulting speech can suffer from reduced naturalness, accented speech, and/or loss of essential voice characteristics.Previous research shows that many models appear to have insufficient generalization capabilities to perform well on every of these cross-lingual aspects.To overcome these generalization problems, we propose to apply the teacher-student paradigm to cross-lingual synthesis.While a teacher model is commonly used to produce teacher forced data, we propose to also use it to produce augmented data of unseen speaker-language pairs, where the aim is to retain essential speaker characteristics.Both sets of data are then used for student model training, which is trained to retain the naturalness and prosodic variation present in the teacher forced data, while learning the speaker identity from the augmented data.Some modifications to the student model are proposed to make the separation of teacher forced and augmented data more straightforward.Results show that the proposed approach improves the retention of speaker characteristics in the speech, while managing to retain high levels of naturalness and prosodic variation. Marcel de Korte, Jaebok Kim, Aki Kunikoshi, Adaeze Adigwe, Esther Klabbers |
INTERSPEECH | 4 |
| 2022 | Annotation of Communicative Functions of Short Feedback Tokens in SwitchboardabstractThere has been a lot of work on predicting the timing of feedback in conversational systems. However, there has been less focus on predicting the prosody and lexical form of feedback given their communicative function. Therefore, in this paper we present our preliminary annotations of the communicative functions of 1627 short feedback tokens from the Switchboard corpus and an analysis of their lexical realizations and prosodic characteristics. Since there is no standard scheme for annotating the communicative function of feedback we propose our own annotation scheme. Although our work is ongoing, our preliminary analysis revealed lexical tokens such as “yeah” are ambiguous and therefore lexical forms alone are not indicative of the function. Both the lexical form and prosodic characteristics need to be taken into account in order to predict the communicative function. We also found that feedback functions have distinguishable prosodic characteristics in terms of duration, mean pitch, pitch slope, and pitch range. Carol Figueroa, Adaeze Adigwe, Magalie Ochs, Gabriel Skantze |
LREC | 2 |