Richard Cave

dblp:239/9334 · also Richard J. N. Cave · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0002-6410-8200ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021
YearPublicationVenuePosition
2025 Exploring the Usability of Gaze-based Mobile Communication in Ghana
abstract
In Ghana, people with communication challenges could benefit from gaze-based Augmented and Assistive Communication devices (AACs), widely used in countries with greater resources.However, there is limited evidence about the potential of such devices by people with communication disabilities in the Global South.Our study sought to evaluate the usability, identifying barriers and facilitators of adoption of a freely available Android-based eye-gaze AAC application called Look to Speak.The study included training of 10 local speech and language therapists and 15 people with communication difficulties.Our findings highlight how, despite some initial successes and the positive opinions of clients, caregivers and speech and language therapists the Look to Speak application largely failed to deliver substantial communication benefits to individual users.This was due to a combination of factors including the high cognitive load, design flaws of the application -such as the lack of optimization of the selection process depending on the chosen interaction mode, and lack of wheelchairs with adequate postural support, which are necessary for users to be able to successfully utilise the application.We contribute insights surrounding the mismatch between expectations and reality of gaze-base AACs, and considerations about the broader ecosystem required to support adoption and impact of such technologies in Ghana.
Victoria Austin, Gifty Ayoka, Giulia Barbareschi, Richard Cave, Catherine Holloway
ASSETS4
2025 A Cookbook for Community-driven Data Collection of Impaired Speech in Low-Resource Languages
Sumaya Ahmed Salihs, Isaac Wiafe, Jamal-Deen Abdulai, Elikem Doe Atsakpo, Gifty Ayoka, Richard Cave, Akon O. Ekpezu, Catherine Holloway, Katrin Tomanek, Fiifi Baffoe Payin Winful
INTERSPEECH6
2024 Enhancing Communication Equity: Evaluation of an Automated Speech Recognition Application in Ghana
abstract
In Ghana people who struggle to articulate speech as a result of different conditions experience barriers in interacting with others due to difficulties in being understood. Automatic speech recognition software can be used to help listeners understand people with communication difficulties. However, studies have not looked at the practical feasibility of these technologies beyond the Global North. We present a novel user study examining the introduction of one such technology, Google Project Relate, to Ghana. This freely available mobile application can create personalised speech recognition models in English for non-standard speech to support communication. Our user study spans the training of local speech and language therapists and 20 people with communication difficulties. We utilise the Technology Amplification Theory to contribute insights on the need for technological adaptations, awareness and support to reduce differential gaps of access, capacity and motivation to expand the reach of these technologies rather than exacerbating inequalities.
Gifty Ayoka, Giulia Barbareschi, Richard Cave, Catherine Holloway
CHI3
2024 Large Language Models As A Proxy For Human Evaluation In Assessing The Comprehensibility Of Disordered Speech Transcription
abstract
Automatic Speech Recognition (ASR) systems, despite significant advances in recent years, still have much room for improvement particularly in the recognition of disordered speech. Even so, erroneous transcripts from ASR models can help people with disordered speech be better understood, especially if the transcription doesn’t significantly change the intended meaning. Evaluating the efficacy of ASR for this use case requires a methodology for measuring the impact of transcription errors on the intended meaning and comprehensibility. Human evaluation is the gold standard for this, but it can be laborious, slow, and expensive. In this work, we tune and evaluate large language models for this task and find them to be a much better proxy for human evaluators than other metrics commonly used. We further present a case-study using the presented approach to assess the quality of personalized ASR models to make model deployment decisions and correctly set user expectations for model quality as part of our trusted tester program.
Katrin Tomanek, Jimmy Tobin, Subhashini Venugopalan, Richard Cave, Katie Seaver, Jordan R. Green, Rus Heywood
ICASSP4
2024 Learnings from curating a trustworthy, well-annotated, and useful dataset of disordered English speech
abstract
Project Euphonia, a Google initiative, is dedicated to improving automatic speech recognition (ASR) of disordered speech. A central objective of the project is to create a large, high-quality, and diverse speech corpus. This report describes the project's latest advancements in data collection and annotation methodologies, such as expanding speaker diversity in the database, adding human-reviewed transcript corrections and audio quality tags to 350K (of the 1.2M total) audio recordings, and amassing a comprehensive set of metadata (including more than 40 speech characteristic labels) for over 75% of the speakers in the database. We report on the impact of transcript corrections on our machine-learning (ML) research, inter-rater variability of assessments of disordered speech patterns, and our rationale for gathering speech metadata. We also consider the limitations of using automated off-the-shelf annotation methods for assessing disordered speech.
Pan-Pan Jiang, Jimmy Tobin, Katrin Tomanek, Robert L. MacDonald, Katie Seaver, Richard Cave, Marilyn A. Ladewig, Rus Heywood, Jordan R. Green
INTERSPEECH6
2023 "The less I type, the better": How AI Language Models can Enhance or Impede Communication for AAC Users
abstract
Users of augmentative and alternative communication (AAC) devices sometimes find it difficult to communicate in real time with others due to the time it takes to compose messages. AI technologies such as large language models (LLMs) provide an opportunity to support AAC users by improving the quality and variety of text suggestions. However, these technologies may fundamentally change how users interact with AAC devices as users transition from typing their own phrases to prompting and selecting AI-generated phrases. We conducted a study in which 12 AAC users tested live suggestions from a language model across three usage scenarios: extending short replies, answering biographical questions, and requesting assistance. Our study participants believed that AI-generated phrases could save time, physical and cognitive effort when communicating, but felt it was important that these phrases reflect their own communication style and preferences. This work identifies opportunities and challenges for future AI-enhanced AAC devices.
Stephanie Valencia, Richard Cave, Krystal Kallarackal, Katie Seaver, Michael Terry, Shaun K. Kane
CHI2
2023 An Analysis of Degenerating Speech Due to Progressive Dysarthria on ASR Performance
abstract
Although personalized automatic speech recognition (ASR) models have recently been improved to recognize even severely impaired speech, model performance may degrade over time for persons with degenerating speech. The aims of this study were to (1) analyze the change of performance of ASR over time in individuals with degrading speech, and (2) explore mitigation strategies to optimize recognition throughout disease progression. Speech was recorded by four individuals with degrading speech due to amyotrophic lateral sclerosis (ALS). Word error rates (WER) across recording sessions were computed for three ASR models: Unadapted Speaker Independent (U-SI), Adapted Speaker Independent (A-SI), and Adapted Speaker Dependent (A-SD or personalized). The performance of all models degraded significantly over time as speech became more impaired, but the A-SD model improved markedly when updated with recordings from the severe stages of speech progression. Recording additional utterances early in the disease before significant speech degradation did not improve the performance of A-SD models. This emphasizes the importance of continuous recording (and model retraining) when providing personalized models for individuals with progressive speech impairments.
Katrin Tomanek, Katie Seaver, Pan-Pan Jiang, Richard Cave, Lauren Harrell, Jordan R. Green
ICASSP4
2023 Speech Intelligibility Classifiers from 550k Disordered Speech Samples
abstract
We developed dysarthric speech intelligibility classifiers on 551,176 disordered speech samples contributed by a diverse set of 468 speakers, with a range of self-reported speaking disorders and rated for their overall intelligibility on a five-point scale. We trained three models following different deep learning approaches and evaluated them on ~ 94K utterances from 100 speakers. We further found the models to generalize well (without further training) on the TORGO database[1] (100% accuracy), UASpeech[2] (0.93 correlation), ALS-TDI PMP[3] (0.81 AUC) datasets as well as on a dataset of realistic unprompted speech we gathered (106 dysarthric and 76 control speakers, ~ 2300 samples). We share our model1to advance research in this domain.
Subhashini Venugopalan, Jimmy Tobin, Samuel J. Yang, Katie Seaver, Richard Cave, Pan-Pan Jiang, Neil Zeghidour, Rus Heywood, Jordan R. Green, Michael P. Brenner
ICASSP5
2021 Automatic Speech Recognition of Disordered Speech: Personalized Models Outperforming Human Listeners on Short Phrases
Jordan R. Green, Robert L. MacDonald, Pan-Pan Jiang, Julie Cattiau, Rus Heywood, Richard Cave, Katie Seaver, Marilyn A. Ladewig, Jimmy Tobin, Michael P. Brenner, Philip C. Nelson, Katrin Tomanek
Interspeech6
2021 Disordered Speech Data Collection: Lessons Learned at 1 Million Utterances from Project Euphonia
Robert L. MacDonald, Pan-Pan Jiang, Julie Cattiau, Rus Heywood, Richard Cave, Katie Seaver, Marilyn A. Ladewig, Jimmy Tobin, Michael P. Brenner, Philip C. Nelson, Jordan R. Green, Katrin Tomanek
Interspeech5