Iona Gessinger

dblp:212/6441 · DBLP profile ↗
← Back
16ranked-venue papers
9as first author
11since 2021 · last 2026
0000-0001-5333-9794ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 7 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Under the hood: Phonemic Restoration in transformer-based automatic speech recognition
abstract
This study investigates how the automatic speech recognition (ASR) models wav2vec 2.0 large-960h-lv60-self and Whisper large-v3 perform when segment-level signal perturbations (added noise, noisy gaps, and two types of silent gaps) are introduced in English words and pseudowords. We probed the speech embeddings throughout their encoder transformer layers to examine how they encode articulatory features (place and manner of articulation, and voicing). We found that wav2vec 2.0 was more successful than Whisper at restoring perturbed segments across conditions. For wav2vec 2.0 embeddings, classification accuracy was higher in words than in pseudowords. The articulatory features encoding of both ASR models was least disturbed by added noise, and most disturbed by noisy gaps, with silent gaps falling in between. Coarticulatory cues improved classification of articulatory features and classification accuracy increased from early to late layers for both models. Among the examined target sounds, [n] stood out from [m], , and [l], as it was classified particularly well under all conditions. We compare ASR performance to the Phonemic Restoration Effect in human speech perception and discuss potential reasons for the performance differences between the two ASR models. This approach aims to foster a better understanding of otherwise opaque systems.
Iona Gessinger, Erfan A. Shams, Julie Carson-Berndsen
Comput. Speech Lang.1
2025 ChatGPT and me: First-time and experienced users' perceptions of ChatGPT's communicative ability as a dialogue partner
Iona Gessinger, Katie Seaborn, Madeleine Steeds, Benjamin R. Cowan
Int. J. Hum. Comput. Stud.1
2025 The Partner Modelling Questionnaire: A Validated Self-Report Measure of Perceptions toward Machines as Dialogue Partners
abstract
Recent work has looked to understand user perceptions of speech agent capabilities as dialogue partners (termed partner models), and how this affects user interaction. Yet, partner model effects are currently inferred from language production as no metrics are available to quantify these subjective perceptions more directly. Through three phases of work, we develop and validate the Partner Modelling Questionnaire (PMQ): an 18-item self-report semantic differential scale designed to reliably measure people’s partner models of non-embodied speech interfaces. Through confirmatory factor analysis, we confirm that the PMQ scale consists of three factors: communicative competence and dependability, human-likeness in communication and communicative flexibility. Our studies show that the measure consistently demonstrates good internal reliability, strong test-retest reliability over 4- and 12-week intervals, and predictable convergent/divergent validity. Based on our findings, we discuss the multidimensional nature of partner models, while identifying key future research avenues that the development of the PMQ facilitates. Notably, this includes the need to identify the activation, sensitivity, and dynamism of partner models in speech interface interaction.
Philip R. Doyle, Iona Gessinger, Justin Edwards, Leigh Clark, Odile Dumbleton, Diego Garaialde, Daniel J. Rough, Anna Bleakley, Holly P. Branigan, Benjamin R. Cowan
ACM Trans. Comput. Hum. Interact.2
2024 Cooking With Agents: Designing Context-aware Voice Interaction
abstract
Voice Agents (VAs) are touted as being able to help users in complex tasks such as cooking and interacting as a conversational partner to provide information and advice while the task is ongoing. Through conversation analysis of 7 cooking sessions with a commercial VA, we identify challenges caused by a lack of contextual awareness leading to irrelevant responses, misinterpretation of requests, and information overload. Informed by this, we evaluated 16 cooking sessions with a wizard-led context-aware VA. We observed more fluent interaction between humans and agents, including more complex requests, explicit grounding within utterances, and complex social responses. We discuss reasons for this, the potential for personalisation, and the division of labour in VA communication and proactivity. Then, we discuss the recent advances in generative models and the VAs interaction challenges. We propose limited context awareness in VAs as a step toward explainable, explorable conversational interfaces.
Razan Jaber, Sabrina Zhong, Sanna Kuoppamäki, Aida Hosseini, Iona Gessinger, Duncan P. Brumby, Benjamin R. Cowan, Donald McMillan
CHI5
2024 The Use of Modifiers and f0 in Remote Referential Communication with Human and Computer Partners
Iona Gessinger, Bistra Andreeva, Benjamin R. Cowan
INTERSPEECH1
2024 PhoneViz: exploring alignments at a glance
Margot Masson, Erfan A. Shams, Iona Gessinger, Julie Carson-Berndsen
INTERSPEECH3
2024 Are Articulatory Feature Overlaps Shrouded in Speech Embeddings?
Erfan A. Shams, Iona Gessinger, Patrick Cormac English, Julie Carson-Berndsen
INTERSPEECH2
2023 Cross-linguistic Emotion Perception in Human and TTS Voices
Iona Gessinger, Michelle Cohn, Benjamin R. Cowan, Georgia Zellou, Bernd Möbius
INTERSPEECH1
2023 Audience design and egocentrism in reference production during human-computer dialogue
Paola Peña, Philip R. Doyle, Justin Edwards, Diego Garaialde, Daniel J. Rough, Anna Bleakley, Leigh Clark, Anita Tobar Henriquez, Holly P. Branigan, Iona Gessinger, Benjamin R. Cowan
Int. J. Hum. Comput. Stud.10
2022 Cross-Cultural Comparison of Gradient Emotion Perception: Human vs. Alexa TTS Voices
Iona Gessinger, Michelle Cohn, Georgia Zellou, Bernd Möbius
INTERSPEECH1
2021 Phonetic accommodation to natural and synthetic voices: Behavior of groups and individuals in speech shadowing
abstract
The present study investigates whether native speakers of German phonetically accommodate to natural and synthetic voices in a shadowing experiment. We aim to determine whether this phenomenon, which is frequently found in HHI, also occurs in HCI involving synthetic speech. The examined features pertain to different phonetic domains: allophonic variation, schwa epenthesis, realization of pitch accents, word-based temporal structure and distribution of spectral energy. On the individual level, we found that the participants converged to varying subsets of the examined features, while they maintained their baseline behavior in other cases or, in rare instances, even diverged from the model voices. This shows that accommodation with respect to one particular feature may not predict the behavior with respect to another feature. On the group level, the participants of the natural condition converged to all features under examination, however very subtly so for schwa epenthesis. The synthetic voices, while partly reducing the strength of effects found for the natural voices, triggered accommodating behavior as well. The predominant pattern for all voice types was convergence during the interaction followed by divergence after the interaction.
Iona Gessinger, Eran Raveh, Ingmar Steiner, Bernd Möbius
Speech Commun.1
2020 Differences in Gradient Emotion Perception: Human vs. Alexa Voices
Michelle Cohn, Eran Raveh, Kristin Predeck, Iona Gessinger, Bernd Möbius, Georgia Zellou
INTERSPEECH4
2020 Phonetic Accommodation of L2 German Speakers to the Virtual Language Learning Tutor Mirabella
Iona Gessinger, Bernd Möbius, Bistra Andreeva, Eran Raveh, Ingmar Steiner
INTERSPEECH1
2019 Phonetic Accommodation in a Wizard-of-Oz Experiment: Intonation and Segments
Iona Gessinger, Bernd Möbius, Bistra Andreeva, Eran Raveh, Ingmar Steiner
INTERSPEECH1
2019 Three's a Crowd? Effects of a Second Human on Vocal Accommodation with a Voice Assistant
Eran Raveh, Ingo Siegert, Ingmar Steiner, Iona Gessinger, Bernd Möbius
INTERSPEECH4
2017 Shadowing Synthesized Speech - Segmental Analysis of Phonetic Convergence
Iona Gessinger, Eran Raveh, Sébastien Le Maguer, Bernd Möbius, Ingmar Steiner
INTERSPEECH1