VLDB 2026 Research / reviewers in the wild / expert
Sarenne Wallbridge
dblp:292/3188 · also Sarenne Carrol Wallbridge
· DBLP profile ↗
9ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0002-1401-4492ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Can self-supervised speech models predict the perceived acceptability of prosodic variation?abstractThough producing an appropriate prosodic realisation of text is a one-to-many problem, modern speech generation often focuses on identifying the “best” or “most likely” output, overlooking acceptable variation across realisations. How listeners perceive such variation–and whether models capture it–is unaccounted for in current evaluation paradigms. In this study, we present exploratory analyses of whether self-supervised models encode acceptable prosodic variation. Using a new dataset of relative acceptability ratings across carefully controlled, high-quality synthetic utterances, we show that SSL representations contain information predictive of such judgments. By introducing a novel method for deriving probability-based uncertainty from autoregressive speech models, we examine whether this information is available in an unsupervised setting, highlighting the complexity of prosodic perception and the value of more human-centric evaluation paradigms. Sarenne Wallbridge, Adaeze Adigwe, Peter Bell 0001 |
ASRU | 1 |
| 2025 | Can We "Cherry-Pick"? Investigating Multiple Renditions from a Generative Speech Synthesis ModelabstractGenerative Speech Models (GSMs) have seen a surge in popularity due to their ability to generate diverse and high-quality speech. Evaluating models that generate many different renditions for a given input sentence presents a new challenge. Listening tests are still the gold standard for evaluating synthetic speech, but current paradigms only consider a single arbitrary rendition: this fails to give a complete picture of best/typical/worst-rendition performance. We propose a general framework for evaluating and deploying generative speech models. This involves selecting amongst renditions using a sequence of filtering or ranking steps, each using either an objective or subjective (listening) method. The framework is not tied to a particular generative model, and so could be applied to any such model. In this paper, we provide a demonstration of a simple version of this framework which would apply to use-cases where best-rendition performance matters. We explore the concept of "cherry-picking", and ask the question "Is there a rendition that is consistently preferred above all others by listeners?". In a subjective listening test, participants ranked several renditions of the same sentence, from which we measured the prevalence of exceptional renditions. We find that there is indeed a preferred rendition in many, but not all cases. Our framework is flexible. In particular, the use of listeners is optional. In future, they could be replaced with model-based objective measures, for example. Adaeze Adigwe, Sarenne Wallbridge, Zehai Tu, Catherine Lai |
ICASSP | 2 |
| 2025 | Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning
Sarenne Wallbridge, Christoph Minixhofer, Catherine Lai, Peter Bell 0001 |
INTERSPEECH | 1 |
| 2024 | What do people hear? Listeners' Perception of Conversational Speech
Adaeze Adigwe, Sarenne Wallbridge |
INTERSPEECH | 2 |
| 2023 | Do dialogue representations align with perception? An empirical studyabstractThere has been a surge of interest regarding the alignment of large-scale language models with human language comprehension behaviour.The majority of this research investigates comprehension behaviours from reading isolated, written sentences.We propose studying the perception of dialogue, focusing on an intrinsic form of language use: spoken conversations.Using the task of predicting upcoming dialogue turns, we ask whether turn plausibility scores produced by state-of-the-art language models correlate with human judgements.We find a strong correlation for some but not all models: masked language models produce stronger correlations than autoregressive models.In doing so, we quantify human performance on the response selection task for open-domain spoken conversation.To the best of our knowledge, this is the first such quantification.We find that response selection performance can be used as a coarse proxy for the strength of correlation with human judgements, however humans and models make different response selection mistakes.The model which produces the strongest correlation also outperforms human response selection performance.Through ablation studies, we show that pre-trained language models provide a useful basis for turn representations; however, finegrained contextualisation, inclusion of dialogue structure information, and fine-tuning towards response selection all boost response selection accuracy by over 30 absolute points. Sarenne Wallbridge, Peter Bell 0001, Catherine Lai |
EACL | 1 |
| 2023 | Information Value: Measuring Utterance Predictability as Distance from Plausible AlternativesabstractWe present information value, a measure which quantifies the predictability of an utterance relative to a set of plausible alternatives.We introduce a method to obtain interpretable estimates of information value using neural text generators, and exploit their psychometric predictive power to investigate the dimensions of predictability that drive human comprehension behaviour.Information value is a stronger predictor of utterance acceptability in written and spoken dialogue than aggregates of token-level surprisal and it is complementary to surprisal for predicting eye-tracked reading times. 1 Mario Giulianelli, Sarenne Wallbridge, Raquel Fernández |
EMNLP | 2 |
| 2023 | Quantifying the perceptual value of lexical and non-lexical channels in speechabstractSpeech is a fundamental means of communication that can be seen to provide two channels for transmitting information: the lexical channel of \textit{which} words are said, and the non-lexical channel of \textit{how} they are spoken. Both channels shape listener expectations of upcoming communication; however, directly quantifying their relative effect on expectations is challenging. Previous attempts require spoken variations of lexically-equivalent dialogue turns or conspicuous acoustic manipulations. This paper introduces a generalised paradigm to study the value of non-lexical information in dialogue across unconstrained lexical content.By quantifying the perceptual value of the non-lexical channel with both accuracy and entropy reduction, we show that non-lexical information produces a consistent effect on expectations of upcoming dialogue: even when it leads to poorer discriminative turn judgements than lexical content alone, it yields higher consensus among participants. Sarenne Wallbridge, Peter Bell 0001, Catherine Lai |
INTERSPEECH | 1 |
| 2022 | Investigating perception of spoken dialogue acceptability through surprisalabstractSurprisal is used throughout computational psycholinguistics to model a range of language processing behaviour. There is growing evidence that language model (LM) estimates of surprisal correlate with human performance on a range of written language comprehension tasks. Although communicative interaction is arguably the primary form of language use, most studies of surprisal are based on monological, written data. Towards the goal of understanding perception in spontaneous, natural language, we present an exploratory investigation into whether the relationship between human comprehension behaviour and LM-estimated surprisal holds when applied to dialogue, considering both written dialogue, and the lexical component of spoken dialogue. We use a novel judgement task of dialogue utterance acceptability to ask two questions: “How well can people make predictions about written dialogue and transcripts of spoken dialogue?” and “Does surprisal correlate with these acceptability judgements?”. We demonstrate that people can make accurate predictions about upcoming dialogue and that their ability differs between spoken transcripts and written conversation. We investigate the relationship between global and local operationalisations of surprisal and human acceptability judgements, finding a combination of both to provide the most predictive power Sarenne Wallbridge, Catherine Lai, Peter Bell 0001 |
INTERSPEECH | 1 |
| 2021 | It's Not What You Said, it's How You Said it: Discriminative Perception of Speech as a Multichannel Communication SystemabstractPeople convey information extremely effectively through spoken interaction using multiple channels of information transmission: the lexical channel of what is said, and the non-lexical channel of how it is said. We propose studying human perception of spoken communication as a means to better understand how information is encoded across these channels, focusing on the question 'What characteristics of communicative context affect listener's expectations of speech?'. To investigate this, we present a novel behavioural task testing whether listeners can discriminate between the true utterance in a dialogue and utterances sampled from other contexts with the same lexical content. We characterize how perception - and subsequent discriminative capability - is affected by different degrees of additional contextual information across both the lexical and non-lexical channel of speech. Results demonstrate that people can effectively discriminate between different prosodic realisations, that non-lexical context is informative, and that this channel provides more salient information than the lexical channel, highlighting the importance of the non-lexical channel in spoken interaction. Sarenne Wallbridge, Peter Bell 0001, Catherine Lai |
Interspeech | 1 |