Keith Vertanen

dblp:82/5076 · DBLP profile ↗
← Back
47ranked-venue papers
22as first author
15since 2021 · last 2025
0000-0002-7814-2450ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 33 · 13 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 7 first-author · 3 since 2021Artificial intelligence and machine learning · 10 · 7 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Adapting Design Workshops for Autistic Adults
abstract
Autism is a neurodevelopmental disability that impacts one's social communication and interaction. When left unsupported, this can increase the amount of loneliness felt by autistic people. Communication technology, such as AAC, can be helpful in supporting social communication, especially when co-designed with autistic people. We conducted a series of design workshops to co-design a new AAC system specifically supporting social communication. In this paper, we focus on the accessibility issues that were identified when running our workshops and provide recommendations on how to improve the process. We found that it is critical to build support for information processing time into the workshops, include a variety of AAC stakeholders, and create a shared vocabulary between the workshop participants to make design workshops more accessible to autistic adults.
Blade Frisch, Keith Vertanen
ASSETS2
2024 Leveraging Large Pretrained Models for Line-by-Line Spoken Program Recognition
abstract
Spoken programming languages significantly differ from natural English due to the inherent variability in speech patterns among programmers and the wide range of programming constructs. In this paper, we employ Wav2Vec 2.0 to enhance the accuracy of transcribing spoken programming languages like Java. Adapting a model with just one hour of spoken programs that had prior exposure to a substantial amount of natural English-labeled data, we achieve a word error rate (WER) of 8.7%, surpassing the high 28.4% WER of a model trained solely on natural English. Decoding with a domain-specific N-gram model and subsequently rescoring the N-best list with a fine-tuned large language model tailored to the programming domain resulted in a WER of 5.5% on our test set.
Sadia Nowrin, Keith Vertanen
ICASSP2
2024 StegoType: Surface Typing from Egocentric Cameras
abstract
Text input is a critical component of any general purpose computing system, yet efficient and natural text input remains a challenge in AR and VR. Headset based hand-tracking has recently become pervasive among consumer VR devices and affords the opportunity to enable touch typing on virtual keyboards. We present an approach for decoding touch typing on uninstrumented flat surfaces using only egocentric camera-based hand-tracking as input. While egocentric hand-tracking accuracy is limited by issues like self occlusion and image fidelity, we show that a sufficiently diverse training set of hand motions paired with typed text can enable a deep learning model to extract signal from this noisy input. Furthermore, by carefully designing a closed-loop data collection process, we can train an end-to-end text decoder that accounts for natural sloppy typing on virtual keyboards. We evaluate our work with a user study (n=18) showing a mean online throughput of 42.4 WPM with an uncorrected error rate (UER) of 7% with our method compared to a physical keyboard baseline of 74.5 WPM at 0.8% UER, showing progress towards unlocking productivity and high throughput use cases in AR/VR.
Fadi Botros, Yangyang Shi, Pinhao Guo, Bradford J. Snow, Linguang Zhang, Jingming Dong, Keith Vertanen, Shugao Ma, Robert Wang 0002
UIST8
2023 A Usability Study of Nomon: A Flexible Interface for Single-Switch Users
abstract
Many individuals with severe motor impairments communicate via a single switch—which might be activated by a blink, facial movement, or puff of air. These switches are commonly used as input to scanning systems that allow selection from a 2D grid of options. Nomon is an alternative interface that provides a more flexible layout, not confined to a grid. Previous work suggests that, even when options appear in a grid, Nomon may be faster and easier to use than scanning systems. However, previous work primarily tested Nomon with non–motor-impaired individuals, and evaluation with potential end-users was limited to a single motor-impaired participant. We provide a usability study following seven participants with motor impairments and compare their performance with Nomon against a row-column scanning system. Most participants were faster with Nomon in a picture selection task, while entry rates varied more in a text-entry task. However, we found participants had to click more times per selection using Nomon, motivating future research into mitigating this increased click load. All but one participant preferred using Nomon; most reported it felt faster and had better predictive text.
Nicholas Bonaker, Emli-Mari Nel, Keith Vertanen, Tamara Broderick
ASSETS3
2023 Language Model Personalization for Improved Touchscreen Typing
abstract
Touchscreen keyboards rely on language modeling to auto-correct noisy typing and to offer word predictions. While language models can be pre-trained on huge amounts of text, they may fail to capture a user's unique writing style. Using a recently released email personalization dataset, we show improved performance compared to a unigram cache by adapting to a user's text via language models based on prediction by partial match (PPM) and recurrent neural networks. On simulated noisy touchscreen typing of 44 users, our best model increased keystroke savings by 9.9% relative and reduced word error rate by 36% relative compared to a static background language model.
Jiban Adhikary, Keith Vertanen
INTERSPEECH2
2023 FlexType: Flexible Text Input with a Small Set of Input Gestures
abstract
In many situations, it may be impractical or impossible to enter text by selecting precise locations on a physical or touchscreen keyboard. We present an ambiguous keyboard with four character groups that has potential applications for eyes-free text entry, as well as text entry using a single switch or a brain-computer interface. We develop a procedure for optimizing these character groupings based on a disambiguation algorithm that leverages a long-span language model. We produce both alphabetically-constrained and unconstrained character groups in an offline optimization experiment and compare them in a longitudinal user study. Our results did not show a significant difference between the constrained and unconstrained character groups after four hours of practice. As expected, participants had significantly more errors with the unconstrained groups in the first session, suggesting a higher barrier to learning the technique. We therefore recommend the alphabetically-constrained character groups, where participants were able to achieve an average entry rate of 12.0 words per minute with a 2.03% character error rate using a single hand and with no visual feedback.
Dylan Gaines, Mackenzie M. Baker, Keith Vertanen
IUI3
2023 Understanding Adoption Barriers to Dwell-Free Eye-Typing: Design Implications from a Qualitative Deployment Study and Computational Simulations
abstract
Eye-typing is a slow and cumbersome text entry method typically used by individuals with no other practical means of communication. As an alternative, prior HCI research has proposed dwell-free eye-typing as a potential improvement that eliminates time-consuming and distracting dwell-timeouts. However, it is rare that such research ideas are translated into working products. This paper reports on a qualitative deployment study of a product that was developed to allow users access to a dwell-free eye-typing research solution. This allowed us to understand how such a research solution would work in practice, as part of users’ current communication solutions in their own homes. Based on interviews and observations, we discuss a number of design issues that currently act as barriers preventing widespread adoption of dwell-free eye-typing. The study findings are complemented with computational simulations in a range of conditions that were inspired by the findings in the deployment study. These simulations serve to both contextualize the qualitative findings and to explore quantitative implications of possible interface redesigns. The combined analysis gives rise to a set of design implications for enabling wider adoption of dwell-free eye-typing in practice.
Per Ola Kristensson, Morten Mjelde, Keith Vertanen
IUI3
2022 Exploring Motor-impaired Programmers' Use of Speech Recognition
abstract
Typing programs can be difficult or impossible for programmers with motor impairments. Programming by voice can be a promising alternative. In this research, we explored the perceptions of motor-impaired programmers with regard to programming by voice. We learned that leveraging existing voice-based programming platforms to speak code can be more complicated than it needs to be. The interviewees expressed their frustration with long hours of memorizing unnatural commands in order to enter code by voice. In addition, we found a preference for being able to speak code in a flexible manner without requiring strict adherence to a grammar.
Sadia Nowrin, Patricia Ordóñez 0002, Keith Vertanen
ASSETS3
2022 Understanding How People with Visual Impairments Take Selfies: Experiences and Challenges
abstract
Selfies are a pervasive form of communication in social media. While there has been some work on systems that guide people with visual impairments (PVI) in taking photos, nearly all has focused on using the camera on the back of the device. We do not know whether and how PVI take selfies. The aim of our work is to understand (1) PVI selfie-taking experiences and challenges, (2) what information do PVI need when taking selfies, and (3) what modalities do PVI prefer (e.g., tactile, verbal, or non-verbal audio) to support selfie-taking. To address this gap, we conducted interviews with 10 PVI. Our findings show that current selfie-taking applications do not provide enough assistance to meet the needs of PVI. We contribute design guidelines that researchers and designers can implement for creating accessible selfie-taking applications.
Ricardo E. Gonzalez, Paul Vermette, Cheng Zhang 0022, Keith Vertanen, Shiri Azenkot
ASSETS5
2022 A Performance Evaluation of Nomon: A Flexible Interface for Noisy Single-Switch Users
abstract
Some individuals with motor impairments communicate using a single switch — such as a button click, air puff, or blink. Row-column scanning provides a method for choosing items arranged in a grid using a single switch. An alternative, Nomon, allows potential selections to be arranged arbitrarily rather than requiring a grid (as desired for gaming, drawing, etc.) — and provides an alternative probabilistic selection method. While past results suggest that Nomon may be faster and easier to use than row-column scanning, no work has yet quantified performance of the two methods over longer time periods or in tasks beyond writing. In this paper, we also develop and validate a webcam-based switch that allows a user without a motor impairment to approximate the response times of a motor-impaired single switch user; although the approximation is not a replacement for testing with single-switch users, it allows us to better initialize, calibrate, and evaluate our method. Over 10 sessions with the webcam switch, we found users typed faster and more easily with Nomon than with row-column scanning. The benefits of Nomon were even more pronounced in a picture-selection task. Evaluation and feedback from a motor-impaired switch user further supports the promise of Nomon.
Nicholas Bonaker, Emli-Mari Nel, Keith Vertanen, Tamara Broderick
CHI3
2021 Accelerating Text Communication via Abbreviated Sentence Input
abstract
Jiban Adhikary, Jamie Berger, Keith Vertanen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jiban Adhikary, Jamie Berger, Keith Vertanen
ACL/IJCNLP (1)3
2021 Enhancing the Composition Task in Text Entry Studies: Eliciting Difficult Text and Improving Error Rate Calculation
abstract
Participants in text entry studies usually copy phrases or compose novel messages. A composition task mimics actual user behavior and can allow researchers to better understand how a system might perform in reality. A problem with composition is that participants may gravitate towards writing simple text, that is, text containing only common words. Such simple text is insufficient to explore all factors governing a text entry method, such as its error correction features. We contribute to enhancing composition tasks in two ways. First, we show participants can modulate the difficulty of their compositions based on simple instructions. While it took more time to compose difficult messages, they were longer, had more difficult words, and resulted in more use of error correction features. Second, we compare two methods for obtaining a participant’s intended text, comparing both methods with a previously proposed crowdsourced judging procedure. We found participant-supplied references were more accurate.
Dylan Gaines, Per Ola Kristensson, Keith Vertanen
CHI3
2021 Typing on Midair Virtual Keyboards: Exploring Visual Designs and Interaction Styles
Jiban Adhikary, Keith Vertanen
INTERACT (4)2
2021 Mining, analyzing, and modeling text written on mobile devices
abstract
Abstract We present a method for mining the web for text entered on mobile devices. Using searching, crawling, and parsing techniques, we locate text that can be reliably identified as originating from 300 mobile devices. This includes 341,000 sentences written on iPhones alone. Our data enables a richer understanding of how users type “in the wild” on their mobile devices. We compare text and error characteristics of different device types, such as touchscreen phones, phones with physical keyboards, and tablet computers. Using our mined data, we train language models and evaluate these models on mobile test data. A mixture model trained on our mined data, Twitter, blog, and forum data predicts mobile text better than baseline models. Using phone and smartwatch typing data from 135 users, we demonstrate our models improve the recognition accuracy and word predictions of a state-of-the-art touchscreen virtual keyboard decoder. Finally, we make our language models and mined dataset available to other researchers.
Keith Vertanen, Per Ola Kristensson
Nat. Lang. Eng.1
2021 Text Entry in Virtual Environments using Speech and a Midair Keyboard
abstract
Entering text in virtual environments can be challenging, especially without auxiliary input devices. We investigate text input in virtual reality using hand-tracking and speech. Our system visualizes users' hands in the virtual environment, allowing typing on an auto-correcting midair keyboard. It also supports speaking a sentence and then correcting errors by selecting alternative words proposed by a speech recognizer. We conducted a user study in which participants wrote sentences with and without speech. Using only the keyboard, users wrote at 11 words-per-minute at a 1.2% error rate. Speaking and correcting sentences was faster and more accurate at 28 words-per-minute and a 0.5% error rate. Participants achieved this performance despite half of sentences containing an uncommon out-of-vocabulary word (e.g. proper name). For sentences with only in-vocabulary words, performance using speech and midair keyboard corrections was faster at 36 words-per-minute with a low 0.3% error rate.
Jiban Adhikary, Keith Vertanen
IEEE Trans. Vis. Comput. Graph.2
2019 VelociWatch: Designing and Evaluating a Virtual Keyboard for the Input of Challenging Text
abstract
Virtual keyboard typing is typically aided by an auto-correct method that decodes a user's noisy taps into their intended text. This decoding process can reduce error rates and possibly increase entry rates by allowing users to type faster but less precisely. However, virtual keyboard decoders sometimes make mistakes that change a user's desired word into another. This is particularly problematic for challenging text such as proper names. We investigate whether users can guess words that are likely to cause auto-correct problems and whether users can adjust their behavior to assist the decoder. We conduct computational experiments to decide what predictions to offer in a virtual keyboard and design a smartwatch keyboard named VelociWatch. Novice users were able to use the features of VelociWatch to enter challenging text at 17 words-per-minute with a corrected error rate of 3%. Interestingly, they wrote slightly faster and just as accurately on a simpler keyboard with limited correction options. Our finding suggest users may be able to type difficult words on a smartwatch simply by tapping precisely without the use of auto-correct.
Keith Vertanen, Dylan Gaines, Crystal Fletcher, Alex M. Stanage, Robbie Watling, Per Ola Kristensson
CHI1
2018 The Impact of Word, Multiple Word, and Sentence Input on Virtual Keyboard Decoding Performance
abstract
Entering text on non-desktop computing devices is often done via an onscreen virtual keyboard. Input on such keyboards normally consists of a sequence of noisy tap events that specify some amount of text, most commonly a single word. But is single word-at-a-time entry the best choice? This paper compares user performance and recognition accuracy of word-at-a-time, phrase-at-a-time, and sentence-at-a-time text entry on a smartwatch keyboard. We evaluate the impact of differing amounts of input in both text copy and free composition tasks. We found providing input of an entire sentence significantly improved entry rates from 26 wpm to 32 wpm while keeping character error rates below 4%. In offline experiments with more processing power and memory, sentence input was recognized with a much lower 2.0% error rate. Our findings suggest virtual keyboards can enhance performance by encouraging users to provide more input per recognition event.
Keith Vertanen, Crystal Fletcher, Dylan Gaines, Jacob Gould, Per Ola Kristensson
CHI1
2018 Fast and Precise Touch-Based Text Entry for Head-Mounted Augmented Reality with Variable Occlusion
abstract
We present the VISAR keyboard: An augmented reality (AR) head-mounted display (HMD) system that supports text entry via a virtualised input surface. Users select keys on the virtual keyboard by imitating the process of single-hand typing on a physical touchscreen display. Our system uses a statistical decoder to infer users’ intended text and to provide error-tolerant predictions. There is also a high-precision fall-back mechanism to support users in indicating which keys should be unmodified by the auto-correction process. A unique advantage of leveraging the well-established touch input paradigm is that our system enables text entry with minimal visual clutter on the see-through display, thus preserving the user’s field-of-view. We iteratively designed and evaluated our system and show that the final iteration of the system supports a mean entry rate of 17.75wpm with a mean character error rate less than 1%. This performance represents a 19.6% improvement relative to the state-of-the-art baseline investigated: A gaze-then-gesture text entry technique derived from the system keyboard on the Microsoft HoloLens. Finally, we validate that the system is effective in supporting text entry in a fully mobile usage scenario likely to be encountered in industrial applications of AR HMDs.
John J. Dudley, Keith Vertanen, Per Ola Kristensson
ACM Trans. Comput. Hum. Interact.2
2017 Towards Improving Predictive AAC using Crowdsourced Dialogues and Partner Context
abstract
Augmentative and Alternative Communication (AAC) devices typically rely on a language model to help make predictions or disambiguate user input. We investigate how to improve predictions in two-sided conversational dialogues. We collect and share a new corpus of crowdsourced everyday dialogues. We show how language models based on recurrent neural networks outperform N-gram models on these dialogues. We demonstrate further gains are possible using text obtained from an AAC user's communication partner, even when that text is partial or contains errors.
Keith Vertanen
ASSETS1
2017 Efficient Typing on a Visually Occluded Physical Keyboard
abstract
The rise of affordable head-mounted displays (HMDs) has raised questions about how to best design user interfaces for this technology. This paper focuses on the use of HMDs for home and office applications that require substantial text input. A physical keyboard is a familiar and effective text input device in normal desktop computing. But without additional camera technology, an HMD occludes all visual feedback about a user's hand position over the keyboard. We describe a system that assists HMD users in typing on a physical keyboard. Our system has a virtual keyboard assistant that provides visual feedback inside the HMD about a user's actions on the physical keyboard. It also provides powerful automatic correction of typing errors by extending a state-of-the-art touchscreen decoder. In a study with 24 participants, we found our virtual keyboard assistant enabled users to type more accurately on a visually-occluded keyboard. We found users wearing an HMD could type at over 40 words-per-minute while obtaining an error rate of less than 5%.
James W. Walker, Bochao Li, Keith Vertanen, Scott A. Kuhl
CHI3
2015 VelociTap: Investigating Fast Mobile Text Entry using Sentence-Based Decoding of Touchscreen Keyboard Input
abstract
We present VelociTap: a state-of-the-art touchscreen keyboard decoder that supports a sentence-based text entry approach. VelociTap enables users to seamlessly choose from three word-delimiter actions: pushing a space key, swiping to the right, or simply omitting the space key and letting the decoder infer spaces automatically. We demonstrate that VelociTap has a significantly lower error rate than Google's keyboard while retaining the same entry rate. We show that intermediate visual feedback does not significantly affect entry or error rates and we find that using the space key results in the most accurate results. We also demonstrate that enabling flexible word-delimiter options does not incur an error rate penalty. Finally, we investigate how small we can make the keyboard when using VelociTap. We show that novice users can reach a mean entry rate of 41 wpm on a 40 mm wide smartwatch-sized keyboard at a 3% character error rate.
Keith Vertanen, Haythem Memmi, Justin Emge, Shyam Reyal, Per Ola Kristensson
CHI1
2014 Phoneme-based predictive text entry interface
abstract
Phoneme-based text entry provides an alternative typing method for nonspeaking individuals who often experience difficulties in orthographic spelling. In this paper, we investigate the application of rate enhancement strategies to improve the user performance of phoneme-based text entry systems. We have developed a phoneme-based predictive typing system, which employs statistical language modeling techniques to dynamically reduce the phoneme search space and offer accurate word predictions. Results of a case study with a nonspeaking participant demonstrated that our rate enhancement strategies led to improved text entry speed and error rates.
Ha Trinh, Annalu Waller, Keith Vertanen, Per Ola Kristensson, Vicki L. Hanson
ASSETS3
2014 Speech dasher: a demonstration of text input using speech and approximate pointing
abstract
Speech Dasher is a novel text entry interface in which users first speak their desired text and then use the zooming interface Dasher to confirm and correct the recognition result. After several hours of practice, users wrote using Speech Dasher at 40 (corrected) words per minute. They did this using only speech and the direction of their gaze (obtained via an eye tracker). Despite an initial recognition word error rate of 22%, users corrected virtually all recognition errors.
Keith Vertanen, David J. C. MacKay
ASSETS1
2014 An evaluation of Dasher with a high-performance language model as a gaze communication method
abstract
Dasher is a promising fast assistive gaze communication method. However, previous evaluations of Dasher have been inconclusive. Either the studies have been too short, involved too few participants, suffered from sampling bias, lacked a control condition, used an inappropriate language model, or a combination of the above. To rectify this, we report results from two new evaluations of Dasher carried out using a Tobii P10 assistive eye-tracker machine. We also present a method of modifying Dasher so that it can use a state-of-the-art long-span statistical language model. Our experimental results show that compared to a baseline eye-typing method, Dasher resulted in significantly faster entry rates (12.6 wpm versus 6.0 wpm in Experiment 1, and 14.2 wpm versus 7.0 wpm in Experiment 2). These faster entry rates were possible while maintaining error rates comparable to the baseline eye-typing method. Participants' perceived physical demand, mental demand, effort and frustration were all significantly lower for Dasher. Finally, participants significantly rated Dasher as being more likeable, requiring less concentration and being more fun.
Daniel J. Rough, Keith Vertanen, Per Ola Kristensson
AVI2
2014 Uncertain text entry on mobile devices
abstract
Users often struggle to enter text accurately on touchscreen keyboards. To address this, we present a flexible decoder for touchscreen text entry that combines probabilistic touch models with a language model. We investigate two different touch models. The first touch model is based on a Gaussian Process regression approach and implicitly models the inherent uncertainty of the touching process. The second touch model allows users to explicitly control the uncertainty via touch pressure. Using the first model we show that the character error rate can be reduced by up to 7% over a baseline method, and by up to 1.3% over a leading commercial keyboard. Using the second model we demonstrate that providing users with control over input certainty reduces the amount of text users have to correct manually and increases the text entry rate.
Daryl Weir, Henning Pohl, Simon Rogers, Keith Vertanen, Per Ola Kristensson
CHI4
2014 The inviscid text entry rate and its application as a grand goal for mobile text entry
abstract
We introduce the concept of the inviscid text entry rate: the point when the user's creativity is the bottleneck rather than the text entry method. We then apply the inviscid text entry rate to define a grand goal for mobile text entry. Via a proxy measure we estimate the population mean of the sufficiently inviscid entry rate to be 67 wpm. We then compare existing mobile text entry methods against this estimate and find that the vast majority of text entry methods in the literature are substantially slower. This analysis suggests the mobile text entry field needs to focus on methods that can viably approach the inviscid entry rate.
Per Ola Kristensson, Keith Vertanen
Mobile HCI2
2014 Complementing text entry evaluations with a composition task
abstract
A common methodology for evaluating text entry methods is to ask participants to transcribe a predefined set of memorable sentences or phrases. In this article, we explore if we can complement the conventional transcription task with a more externally valid composition task. In a series of large-scale crowdsourced experiments, we found that participants could consistently and rapidly invent high quality and creative compositions with only modest reductions in entry rates. Based on our series of experiments, we provide a best-practice procedure for using composition tasks in text entry evaluations. This includes a judging protocol which can be performed either by the experimenters or by crowdsourced workers on a microtask market. We evaluated our composition task procedure using a text entry method unfamiliar to participants. Our empirical results show that the composition task can serve as a valid complementary text entry evaluation method.
Keith Vertanen, Per Ola Kristensson
ACM Trans. Comput. Hum. Interact.1
2013 A collection of conversational AAC-like communications
abstract
We contribute a public test set of everyday conversational communications. The communications were written in response to ten hypothetical situations given to workers on the crowdsourcing site Amazon Mechanical Turk. After quality control, our public dataset consists of 1,506 unique communications. These communications can be used to help design and evaluate text-based predictive communication aids. The collection also provides a common public test set for research into predictive conversational text entry.
Keith Vertanen
ASSETS1
2013 The feasibility of eyes-free touchscreen keyboard typing
abstract
Typing on a touchscreen keyboard is very difficult without being able to see the keyboard. We propose a new approach in which users imagine a Qwerty keyboard somewhere on the device and tap out an entire sentence without any visual reference to the keyboard and without intermediate feedback about the letters or words typed. To demonstrate the feasibility of our approach, we developed an algorithm that decodes blind touchscreen typing with a character error rate of 18.5%. Our decoder currently uses three components: a model of the keyboard topology and tap variability, a point transformation algorithm, and a long-span statistical language model. Our initial results demonstrate that our proposed method provides fast entry rates and promising error rates. On one-third of the sentences, novices' highly noisy input was successfully decoded with no errors.
Keith Vertanen, Haythem Memmi, Per Ola Kristensson
ASSETS1
2013 Improving two-thumb text entry on touchscreen devices
abstract
We study the design of split keyboards for fast text entry with two thumbs on mobile touchscreen devices. The layout of KALQ was determined through first studying how users should grip a device with two hands. We then assigned letters to keys computationally, using a model of two-thumb tapping. KALQ minimizes thumb travel distance and maximizes alternation between thumbs. An error correction algorithm was added to help address linguistic and motor errors. Users reached a rate of 37 words per minute (with a 5% error rate) after a training program.
Antti Oulasvirta, Anna Reichel, Wenbin Li 0003, Yan Zhang 0001, Myroslav Bachynskyi, Keith Vertanen, Per Ola Kristensson
CHI6
2012 iSCAN: a phoneme-based predictive communication aid for nonspeaking individuals
abstract
The high incidence of literacy deficits among people with severe speech impairments (SSI) has been well documented. Without literacy skills, people with SSI are unable to effectively use orthographic-based communication systems to generate novel linguistic items in spontaneous conversation. To address this problem, phoneme-based communication systems have been proposed which enable users to create spoken output from phoneme sequences. In this paper, we investigate whether prediction techniques can be employed to improve the usability of such systems. We have developed iSCAN, a phoneme-based predictive communication system, which offers phoneme prediction and phoneme-based word prediction. A pilot study with 16 able-bodied participants showed that our predictive methods led to a 108.4% increase in phoneme entry speed and a 79.0% reduction in phoneme error rate. The benefits of the predictive methods were also demonstrated in a case study with a cerebral palsied participant. Moreover, results of a comparative evaluation conducted with the same participant after 16 sessions using iSCAN indicated that our system outperformed an orthographic-based predictive communication device that the participant has used for over 4 years.
Ha Trinh, Annalu Waller, Keith Vertanen, Per Ola Kristensson, Vicki L. Hanson
ASSETS3
2012 The potential of dwell-free eye-typing for fast assistive gaze communication
abstract
We propose a new research direction for eye-typing which is potentially much faster: dwell-free eye-typing. Dwell-free eye-typing is in principle possible because we can exploit the high redundancy of natural languages to allow users to simply look at or near their desired letters without stopping to dwell on each letter. As a first step we created a system that simulated a perfect recognizer for dwell-free eye-typing. We used this system to investigate how fast users can potentially write using a dwell-free eye-typing interface. We found that after 40 minutes of practice, users reached a mean entry rate of 46 wpm. This indicates that dwell-free eye-typing may be more than twice as fast as the current state-of-the-art methods for writing by gaze. A human performance model further demonstrates that it is highly unlikely traditional eye-typing systems will ever surpass our dwell-free eye-typing performance estimate.
Per Ola Kristensson, Keith Vertanen
ETRA2
2012 Spelling as a Complementary Strategy for Speech Recognition
abstract
We compare a variety of strategies for incorporating spelling to create more robust voice-only speech in-terfaces. These strategies use different combinations of speaking the word, spelling the word, and spelling the word using a phonetic alphabet. For correcting a single recognition error, spelling the word or speaking and spelling the word reduced error rates substantially. Phonetic-spelling was very accurate with error rates on a 5K task approaching zero. Most importantly, multiple input strategies can be used simultaneously with only a modest degradation in performance compared to allow-ing only a single input strategy. Thus our work shows that spelling-based input strategies offer the potential of a simple, natural and effective way for users to both avoid and correct recognition errors. Index Terms: speech recognition, error correction 1.
Keith Vertanen, Per Ola Kristensson
INTERSPEECH1
2012 Performance comparisons of phrase sets and presentation styles for text entry evaluations
abstract
We empirically compare five different publicly-available phrase sets in two large-scale (N = 225 and N = 150) crowdsourced text entry experiments. We also investigate the impact of asking participants to memorize phrases before writing them versus allowing participants to see the phrase during text entry. We find that asking participants to memorize phrases increases entry rates at the cost of slightly increased error rates. This holds for both a familiar and for an unfamiliar text entry method. We find statistically significant differences between some of the phrase sets in terms of both entry and error rates. Based on our data, we arrive at a set of recommendations for choosing suitable phrase sets for text entry evaluations.
Per Ola Kristensson, Keith Vertanen
IUI2
2011 The Imagination of Crowds: Conversational AAC Language Modeling using Crowdsourcing and Large Data Sources
Keith Vertanen, Per Ola Kristensson
EMNLP1
2011 Asynchronous Multimodal Text Entry Using Speech and Gesture Keyboards
abstract
We propose reducing errors in text entry by combining speech and gesture keyboard input. We describe a merge model that combines recognition results in an asynchronous and flexible manner. We collected speech and gesture data of users entering both short email sentences and web search queries. By merging recognition results from both modalities, word error rate was reduced by 53% relative for email sentences and 29% relative for web searches. For email utterances with speech errors, we investigated providing gesture keyboard corrections of only the erroneous words. Without the user explicitly indicating the incorrect words, our model was able to reduce the word error rate by 44% relative. Copyright © 2011 ISCA.
Per Ola Kristensson, Keith Vertanen
INTERSPEECH2
2011 A versatile dataset for text entry evaluations based on genuine mobile emails
abstract
Mobile text entry methods are typically evaluated by having study participants copy phrases. However, currently there is no available phrase set that has been composed by mobile users. Instead researchers have resorted to using invented phrases that probably suffer from low external validity. Further, there is no available phrase set whose phrases have been verified to be memorable. In this paper we present a collection of mobile email sentences written by actual users on actual mobile devices. We obtained our sentences from emails written by Enron employees on their BlackBerry mobile devices. We provide empirical data on how easy the sentences were to remember and how quickly and accurately users could type these sentences on a full-sized keyboard. Using this empirical data, we construct a series of phrase sets we suggest for use in text entry evaluations.
Keith Vertanen, Per Ola Kristensson
Mobile HCI1
2010 Intelligently Aiding Human-Guided Correction of Speech Recognition
abstract
Correcting recognition errors is often necessary in a speech interface. These errors not only reduce users' overall entry rate, but can also lead to frustration. While making fewer recognition errors is undoubtedly helpful, facilities for supporting user-guided correction are also critical. We explore how to better support user corrections using Parakeet — a continuous speech recognition system for mobile touch-screen devices. Parakeet's interface is designed for easy error correction on a handheld device. Users correct errors by selecting alternative words from a word confusion network and by typing on a predictive software keyboard. Our interface design was guided by computational experiments and used a variety of information sources to aid the correction process. In user studies, participants were able to write text effectively despite sometimes high initial recognition error rates. Using Parakeet as an example, we discuss principles we found were important for building an effective speech correction interface.
Keith Vertanen, Per Ola Kristensson
AAAI1
2010 Speech dasher: fast writing using speech and gaze
abstract
Speech Dasher allows writing using a combination of speech and a zooming interface. Users first speak what they want to write and then they navigate through the space of recognition hypotheses to correct any errors. Speech Dasher's model combines information from a speech recognizer, from the user, and from a letter-based language model. This allows fast writing of anything predicted by the recognizer while also providing seamless fallback to letter-by-letter spelling for words not in the recognizer's predictions. In a formative user study, expert users wrote at 40 (corrected) words per minute. They did this despite a recognition word error rate of 22%. Furthermore, they did this using only speech and the direction of their gaze (obtained via an eye tracker).
Keith Vertanen, David J. C. MacKay
CHI1
2010 Getting it right the second time: Recognition of spoken corrections
abstract
We investigate ways to improve recognition accuracy on spoken corrections. We show that a variety of simple techniques can greatly improve the accuracy on corrections. We further develop a flexible merge model that improves accuracy by combining information from the original recognition and the spoken correction. Our merge model operates on word confusion networks and can easily incorporate prior beliefs about the recognition events (e.g. which words are likely correct or incorrect). By combining all of our techniques, the percentage of correctly recognized spoken corrections increased from 21% to 53%.
Keith Vertanen, Per Ola Kristensson
SLT1
2009 Automatic selection of recognition errors by respeaking the intended text
abstract
We investigate how to automatically align spoken corrections with an initial speech recognition result. Such automatic alignment would enable one-step voice-only correction in which users simply respeak their intended text. We present three new models for automatically aligning corrections: a 1-best model, a word confusion network model, and a revision model. The revision model allows users to alter what they intended to write even when the initial recognition was completely correct. We evaluate our models with data gathered from two user studies. We show that providing just a single correct word of context dramatically improves alignment success from 65% to 84%. We find that a majority of users provide such context without being explicitly instructed to do so. We find that the revision model is superior when users modify words in their initial recognition, improving alignment success from 73% to 83%. We show how our models can easily incorporate prior information about correction location and we show that such information aids alignment success. Last, we observe that users speak their intended text faster and with fewer re-recordings than if they are forced to speak misrecognized text.
Keith Vertanen, Per Ola Kristensson
ASRU1
2009 Recognition and correction of voice web search queries
abstract
In this work we investigate how to recognize and correct voice web search queries. We describe our corpus of web search queries and show how it was used to improve recognition accuracy. We show that using a search-specific vocabulary with automatically generated pronunciations is superior to using a vocabulary limited to a fixed pronunciation dictionary. We conducted a formative user study to investigate recognition and correction aspects of voice search in a mobile context. In the user study, we found that despite a word error rate of 48%, users were able to speak and correct search queries in about 18 seconds. Users did this while walking around using a mobile touch-screen device. Index Terms: speech recognition, voice search, error correction, mobile web search
Keith Vertanen, Per Ola Kristensson
INTERSPEECH1
2009 Parakeet: a continuous speech recognition system for mobile touch-screen devices
abstract
We present Parakeet, a system for continuous speech recognition on mobile touch-screen devices. The design of Parakeet was guided by computational experiments and validated by a user study. Participants had an average text entry rate of 18 words-per-minute (WPM) while seated indoors and 13 WPM while walking outdoors. In an expert pilot study, we found that speech recognition has the potential to be a highly competitive mobile text entry method, particularly in an actual mobile setting where users are walking around while entering text.
Keith Vertanen, Per Ola Kristensson
IUI1
2009 Parakeet: a demonstration of speech recognition on a mobile touch-screen device
abstract
We demonstrate Parakeet -- a continuous speech recognition system for mobile touch-screen devices. Parakeet's interface is designed to make correcting errors easy on a handheld device while on the move. Users correct errors using a touch-screen to either select alternative words from a word confusion network or by typing on a predictive software keyboard. Our interface design was guided by computational experiments. We conducted a user study to validate our design. We found novices entered text at 18 WPM while seated indoors and 13 WPM while walking outdoors.
Keith Vertanen, Per Ola Kristensson
IUI1
2008 On the benefits of confidence visualization in speech recognition
abstract
In a typical speech dictation interface, the recognizer's best-guess is displayed as normal, unannotated text. This ignores potentially useful information about the recognizer's confidence in its recognition hypothesis. Using a confidence measure (which itself may sometimes be inaccurate), we investigated providing visual feedback about low-confidence portions of the recognition using shaded, red underlining. An evaluation showed, compared to a baseline without underlining, underlining low-confidence areas did not increase user's speed or accuracy in detecting errors. However, we found that when recognition errors were correctly underlined, they were discovered significantly more often than baseline. Conversely, when errors failed to be underlined, they were discovered less often. Our results indicate confidence visualization can be effective --- but only if the confidence measure has high accuracy. Further, since our results show that users tend to trust confidence visualization, designers should be careful in its application if a high accuracy confidence measure is not available.
Keith Vertanen, Per Ola Kristensson
CHI1
2008 Combining open vocabulary recognition and word confusion networks
abstract
A limitation of most speech recognizers is that they only recognize words from a fixed vocabulary. In this paper, we explore a technique for addressing this deficiency using automatically derived units made up of letters and phones. We show how these units can be used for letter-to-phone conversion and open-vocabulary recognition. We further show how these units can be merged to form novel words while maintaining a word lattice structure. This allows creation of a word confusion network containing both in- and out-of-vocabulary (OOV) words. Experiments show these open vocabulary confusion networks improve recognition accuracy. They also allow open vocabulary recognition to be used in concert with a convenient confusion network result representation.
Keith Vertanen
ICASSP1
2006 Speech and speech recognition during dictation corrections
abstract
A natural way to correct errors made while dictating to a computer is to respeak portions of the original sentence. But often spoken corrections are themselves misrecognized, costing the user time and testing their patience. To better understand how users behave while correcting, I created a simulated dictation interface and fooled users into believing they were correcting errors by respeaking. I found that users not only hyperarticulate during corrections, but they do so preemptively before any misrecognition. Depending on the recognizer, hyperarticulation was found to cause relatively minor changes in error rate. The correction of isolated words or phrases was more troublesome, causing substantial recognition problems for an HTK recognizer. Dragon Naturally Speaking, on the other hand, performed slightly better on hyperarticulated speech and only degraded slightly on isolated corrections. Index Terms: speech recognition, error correction, dictation, hyperarticulation, correcting by respeaking
Keith Vertanen
INTERSPEECH1