Sharon L. Oviatt

dblp:o/SharonLOviatt · DBLP profile ↗
← Back
69ranked-venue papers
44as first author
6since 2021 · last 2022
0000-0003-4664-1412ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 36 · 24 first-author · 4 since 2021Artificial intelligence and machine learning · 26 · 16 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 14 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2022 A Critical Review of Multimodal-multisensor Analytics for Anxiety Assessment
abstract
Recently, interest has grown in the assessment of anxiety that leverages human physiological and behavioral data to address the drawbacks of current subjective clinical assessments. Complex experiences of anxiety vary on multiple characteristics, including triggers, responses, duration and severity, and impact differently on the risk of anxiety disorders. This article reviews the past decade of studies that objectively analyzed various anxiety characteristics related to five common anxiety disorders in adults utilizing features of cardiac, electrodermal, blood pressure, respiratory, vocal, posture, movement, and eye metrics. Its originality lies in the synthesis and interpretation of consistently discovered heterogeneous predictors of anxiety and multimodal-multisensor analytics based on them. We reveal that few anxiety characteristics have been evaluated using multimodal-multisensor metrics, and many of the identified predictive features are confounded. As such, objective anxiety assessments are not yet complete or precise. That said, few multimodal-multisensor systems evaluated indicate an approximately 11.73% performance gain compared to unimodal systems, highlighting a promising powerful tool. We suggest six high-priority future directions to address the current gaps and limitations in infrastructure, basic knowledge, and application areas. Action in these directions will expedite the discovery of rich, accurate, continuous, and objective assessments and their use in impactful end-user applications.
Hashini Senaratne, Sharon L. Oviatt, Kirsten Ellis, Glenn Melvin
ACM Trans. Comput. Heal.2
2021 Advances in Multimodal Behavioral Analytics for Early Dementia Diagnosis: A Review
abstract
Clinical diagnosis of dementia is typically delayed and limited in accuracy, despite assessing cognitive impairments through neurological exams, brain imaging, and functional tests such as Activities of Daily Living. Recent advances in digital health and multimodal behavioral analytics are beginning to provide more sensitive, objective, unobtrusive and continuous assessment of functional abilities while people remain in a familiar setting such as their home, at work, or in the community. These new techniques analyze natural behaviors like speech, language, gait, eye gaze, hand movements, and facial expressions. This review compares existing clinical assessment methods with emerging behavioral analytic techniques that offer powerful capabilities for earlier and more precise diagnosis of dementia. It summarizes state-of-the-art multimodal behavioral analytics research for dementia diagnosis, including predictive features present in different human behaviors and the performance advantages of combining them into multimodal diagnostic systems. The many behavioral predictors documented in the literature are interpreted as deriving from six common cognitive deficits that are well known hallmarks of dementia. The review also discusses long-term trends in multimodal behavioral analytics research, and the five main areas requiring future work to realize the promise of earlier, more accurate, and widely accessible dementia diagnostic systems.
Chathurika Jayangani Palliya Guruge, Sharon L. Oviatt, Pari Delir Haghighi, Elizabeth Pritchard
ICMI2
2021 Technology as Infrastructure for Dehumanization: : Three Hundred Million People with the Same Face
abstract
Based on humanistic psychology, this paper reflects on how and why technology has led to increasing dehumanization within society. The literature on human autonomy, or individuals’ ability to exercise agency and control over their own lives, emphasizes that erosion of autonomy adversely impacts human behavior and health—including demotivating people, heightening their anxiety and apathy, undermining interpersonal relations, reducing self-efficacy and learning, and damaging mental and physical health in substantial ways. This paper describes concrete examples of how technology is accelerating loss of human autonomy, which often occurs during invasive surveillance and covert manipulation during user-technology interactions. These examples highlight abuses of multimodal-multisensor technologies, especially during education. An analysis is provided of the psychosocial context that encourages people to use technology in abusive ways, which directly fuels the growing schism between those who control and those who are disempowered victims. Five specific directions are outlined for how our community can design future technology using counter tactics, practices, and policies that strengthen rather than eroding autonomy and well-being.
Sharon L. Oviatt
ICMI1
2021 A Multimodal Dataset and Evaluation for Feature Estimators of Temporal Phases of Anxiety
abstract
Vicious cycles of anxiety responses underlie the onset of increasingly prevalent and highly impairing anxiety disorders and also contribute to their maintenance. Our goal is to evaluate whether different anxiety responses are evident in temporal patterns of physiological and behavioral features. Consequently, we established a rich multimodal-multisensor dataset of cardiac, electrodermal, movement, posture, and speech measures from 95 young adults during two anxiety experiments that induce social anxiety and bug-phobic anxiety. A subset of this dataset is publicly available at “Anxiety Phases Dataset” Figshare repository. We adopted a generalized mixed model approach and found that 10 out of 14 feature trajectories modeled for high- and low-anxiety groups differ significantly at 0.001 level in magnitude, creating at least two temporal phases in both groups. Further differences in magnitude, duration and the number of phases were observed for responses of confrontation, safety behaviors, escape, and avoidance in the high-anxiety group. Our findings contribute to the long-term aim of designing multimodal systems that have great potential to reduce the impacts of anxiety disorders and improve therapy.
Hashini Senaratne, Levin Kuhlmann, Kirsten Ellis, Glenn Melvin, Sharon L. Oviatt
ICMI5
2021 A Taxonomy of Social Errors in Human-Robot Interaction
abstract
Robotic applications have entered various aspects of our lives, such as health care and educational services. In such Human-robot Interaction (HRI), trust and mutual adaption are established and maintained through a positive social relationship between a user and a robot. This social relationship relies on the perceived competence of a robot on the social-emotional dimension. However, because of technical limitations and user heterogeneity, current HRI is far from error-free, especially when a system leaves controlled lab environments and is applied to in-the-wild conditions. Errors in HRI may either degrade a user’s perception of a robot’s capability in achieving a task (defined as performance errors in this work) or degrade a user’s perception of a robot’s socio-affective competence (defined as social errors in this work). The impact of these errors and effective strategies to handle such an impact remains an open question. We focus on social errors in HRI in this work. In particular, we identify the major attributes of perceived socio-affective competence by reviewing human social interaction studies and HRI error studies. This motivates us to propose a taxonomy of social errors in HRI. We then discuss the impact of social errors situated in three representative HRI scenarios. This article provides foundations for a systematic analysis of the social-emotional dimension of HRI. The proposed taxonomy of social errors encourages the development of user-centered HRI systems, designed to offer positive and adaptive interaction experiences and improved interaction outcomes.
Leimin Tian, Sharon L. Oviatt
ACM Trans. Hum. Robot Interact.2
2021 I Know What You Know: What Hand Movements Reveal about Domain Expertise
abstract
This research investigates whether students’ level of domain expertise can be detected during authentic learning activities by analyzing their physical activity patterns. More expert students reduced their manual activity by a substantial 50%, which was evident in fine-grained signal analyses and total rate of gesturing. The quality of experts’ discrete hand movements also averaged shorter in distance, briefer in duration, and slower in velocity than those of non-experts. Interestingly, experts adapted by nearly eliminating gestures on easier problems, while selectively increasing them on harder ones. They also strategically produced 62% more iconic gestures, which serve to retain spatial information in working memory while extracting inferences required to solve problems correctly. These findings highlight the close relation between hand movements and mental state and, more specifically, that hand movements provide an unusually clear window on students’ level of domain expertise. Embodied Cognition and Limited Resource theories only partially account for the present findings, which specify future directions for theoretical work.
Sharon L. Oviatt, Jionghao Lin, Abishek Sriramulu
ACM Trans. Interact. Intell. Syst.1
2020 LSTM-DNN based Approach for Pain Intensity and Protective Behaviour Prediction
abstract
This paper proposes an approach for pain intensity recognition and protective behaviour prediction task from body movements as a part of the EmoPain challenge. The given dataset consists of body part based sensor data for both the tasks. The proposed network is a lightweight LSTM-DNN model, which takes the angle, angle energy and sEMG features as input and predicts pain intensity level and protective behaviour as output. The performance of LSTM, Bi-LSTM, attention-LSTM and LSTM-DNN models are compared for this problem on the same dataset. In order to enhance the model’s discriminating power, joint training of all the models are performed, combining respective task labels with exercise type as an additional label. The experiments show that the proposed approach is effective and outperforms the baseline on the validation set by a margin of 35.00% for pain intensity prediction and 47.72% for protective behaviour prediction, respectively.
Shreya Ghosh 0001, Jyoti Joshi, Sharon L. Oviatt
FG4
2019 An Explainable Deep Fusion Network for Affect Recognition Using Physiological Signals
abstract
Affective computing is an emerging research area which provides insights on human's mental state through human-machine interaction. During the interaction process, bio-signal analysis is essential to detect human affective changes. Currently, machine learning methods to analyse bio-signals are the state of the art to detect the affective states, but most empirical works mainly deploy traditional machine learning methods rather than deep learning models due to the need for explainability. In this paper, we propose a deep learning model to process multimodal-multisensory bio-signals for affect recognition. It supports batch training for different sampling rate signals at the same time, and our results show significant improvement compared to the state of the art. Furthermore, the results are interpreted at the sensor- and signal- level to improve the explainaibility of our deep learning model.
Jionghao Lin, Shirui Pan, Cheng Siong Lee, Sharon L. Oviatt
CIKM4
2019 Dynamic Adaptive Gesturing Predicts Domain Expertise in Mathematics
abstract
Embodied Cognition theorists believe that mathematics thinking is embodied in physical activity, like gesturing while explaining math solutions. This research asks the question whether expertise in mathematics can be detected by analyzing students’ rate and type of manual gestures. The results reveal several unique findings, including that math experts reduced their total rate of gesturing by 50%, compared with non-experts. They also dynamically increased their rate of gesturing on harder problems. Although experts reduced their rate of gesturing overall, they selectively produced 62% more iconic gestures. Iconic gestures are strategic because they assist with retaining spatial information in working memory, so that inferences can be extracted to support correct problem solving. The present results on representation-level gesture patterns are convergent with recent findings on signal-level handwriting, while also contributing a causal understanding of how and why experts adapt their manual activity during problem solving.
Abishek Sriramulu, Jionghao Lin, Sharon L. Oviatt
ICMI3
2018 Ten Opportunities and Challenges for Advancing Student-Centered Multimodal Learning Analytics
abstract
This paper presents a summary and critical reflection on ten major opportunities and challenges for advancing the field of multimodal learning analytics (MLA). It identifies emerging technology trends likely to disrupt learning analytics, challenges involved in forging viable participatory design partnerships, and impending issues associated with the control of data and privacy. Trends in health care analytics provide one attractive model for how new infrastructure can enable the collection of largerscale and more diverse datasets, and how end-user analytics can be designed to empower individuals and expand market adoption.
Sharon L. Oviatt
ICMI1
2018 Dynamic Handwriting Signal Features Predict Domain Expertise
abstract
As commercial pen-centric systems proliferate, they create a parallel need for analytic techniques based on dynamic writing. Within educational applications, recent empirical research has shown that signal-level features of students’ writing, such as stroke distance, pressure and duration, are adapted to conserve total energy expenditure as they consolidate expertise in a domain. The present research examined how accurately three different machine-learning algorithms could automatically classify users’ domain expertise based on signal features of their writing, without any content analysis. Compared with an unguided machine-learning classification accuracy of 71%, hybrid methods using empirical-statistical guidance correctly classified 79–92% of students by their domain expertise level. In addition to improved accuracy, the hybrid approach contributed a causal understanding of prediction success and generalization to new data. These novel findings open up opportunities to design new automated learning analytic systems and student-adaptive educational technologies for the rapidly expanding sector of commercial pen systems.
Sharon L. Oviatt, Kevin Hang, Jianlong Zhou, Kun Yu 0001, Fang Chen 0001
ACM Trans. Interact. Intell. Syst.1
2016 Multimodal learning analytics data challenges
abstract
This is a proposal for organizing a Multimodal Learning Analytics (MLA) data challenge as part of the workshop offering of the Learning Analytics and Knowledge (LAK) conference. It explains the motivation of the event, its objectives, target groups, expected format, organization, dissemination strategy and schedule.
Xavier Ochoa 0001, Marcelo Worsley, Nadir Weibel, Sharon L. Oviatt
LAK4
2015 Spoken Interruptions Signal Productive Problem Solving and Domain Expertise in Mathematics
abstract
Prevailing social norms prohibit interrupting another person when they are speaking. In this research, simultaneous speech was investigated in groups of students as they jointly solved math problems and peer tutored one another. Analyses were based on the Math Data Corpus, which includes ground-truth performance coding and speech transcriptions. Simultaneous speech was elevated 120-143% during the most productive phase of problem solving, compared with matched intervals. It also was elevated 18-37% in students who were domain experts, compared with non-experts. Qualitative analyses revealed that experts differed from non-experts in the function of their interruptions. Analysis of these functional asymmetries produced nine key behaviors that were used to identify the dominant math expert in a group with 95-100% accuracy in three minutes. This research demonstrates that overlapped speech is a marker of group problem-solving progress and domain expertise. It provides valuable information for the emerging field of learning analytics.
Sharon L. Oviatt, Kevin Hang, Jianlong Zhou, Fang Chen 0001
ICMI1
2014 Written Activity, Representations and Fluency as Predictors of Domain Expertise in Mathematics
abstract
The emerging field of multimodal learning analytics evaluates natural communication modalities (digital pen, speech, images) to identify domain expertise, learning, and learning-oriented precursors. Using the Math Data Corpus, this research investigated students' digital pen input as small groups collaborated on solving math problems. Compared with non-experts, findings indicated that domain experts have an opposite pattern of accelerating total written activity as problem difficulty increases, a lower written and spoken disfluency rate, and they express different content--including a higher ratio of nonlinguistic symbolic representations and structured diagrams to elemental marks. Implications are discussed for developing reliable multimodal learning analytics systems that incorporate digital pen input to automatically track the consolidation of domain expertise. This includes prediction based on a combination of activity patterns, fluency, and content analysis. New MMLA systems are expected to have special utility on cell phones, which already have multimodal interfaces and are the dominant educational platform worldwide.
Sharon L. Oviatt, Adrienne Cohen
ICMI1
2013 ICMI 2013 grand challenge workshop on multimodal learning analytics
abstract
Advances in learning analytics are contributing new empirical findings, theories, methods, and metrics for understanding how students learn. It also contributes to improving pedagogical support for students' learning through assessment of new digital tools, teaching strategies, and curricula. Multimodal learning analytics (MMLA)[1] is an extension of learning analytics and emphasizes the analysis of natural rich modalities of communication across a variety of learning contexts. This MMLA Grand Challenge combines expertise from the learning sciences and machine learning in order to highlight the rich opportunities that exist at the intersection of these disciplines. As part of the Grand Challenge, researchers were asked to predict: (1) which student in a group was the dominant domain expert, and (2) which problems that the group worked on would be solved correctly or not. Analyses were based on a combination of speech, digital pen and video data. This paper describes the motivation for the grand challenge, the publicly available data resources and results reported by the challenge participants. The results demonstrate that multimodal prediction of the challenge goals: (1) is surprisingly reliable using rich multimodal data sources, (2) can be accomplished using any of the three modalities explored, and (3) need not be based on content analysis.
Louis-Philippe Morency, Sharon L. Oviatt, Stefan Scherer, Nadir Weibel, Marcelo Worsley
ICMI2
2013 Interfaces for thinkers: computer input capabilities that support inferential reasoning
abstract
Recent research has revealed that basic computer input capabilities can substantially facilitate or impede people's ability to produce ideas and solve problems correctly. This research asks: What type of interface provides best support for inferential reasoning in both low- and high-performing students' Students' ability to make accurate inferences about science and everyday reasoning tasks was compared while they used: (1) non digital pen and paper, (2) a digital pen and paper interface, (3) pen tablet interface, and (4) graphical tablet interface. Correct inferences averaged 10.5% higher when using a digital pen interface, compared with the tablet interfaces. Further analyses revealed that overgeneralization and redundancy errors were more common when using the tablet interfaces and among low performers. Implications are discussed for designing more effective computational thinking tools.
Sharon L. Oviatt
ICMI1
2013 Problem solving, domain expertise and learning: ground-truth performance results for math data corpus
abstract
Problem solving, domain expertise, and learning are analyzed for the Math Data Corpus, which involves multimodal data on collaborating student groups as they solve math problems together across sessions. Compared with non-expert students, domain experts contributed more group solutions, solved more problems correctly and took less time. These differences between experts and non-experts were accentuated on harder problems. A cumulative expertise metric validated that expert and non-expert students represented distinct non overlapping populations, a finding that replicated across sessions. Group performance also improved 9.4% across sessions, due mainly to learning by expert students. These findings satisfy ground-truth conditions for developing prediction techniques that aim to identify expertise based on multimodal communication and behavior patterns. Together with the Math Data Corpus, these results contribute valuable resources for supporting data-driven grand challenges on multimodal learning analytics, which aim to develop new techniques for predicting expertise early, reliably, and objectively. as well as learning-oriented precursors.
Sharon L. Oviatt
ICMI1
2013 Written and multimodal representations as predictors of expertise and problem-solving success in mathematics
abstract
One aim of multimodal learning analytics is to analyze rich natural communication modalities to identify domain expertise and learning rapidly and reliably. In this research, written and multimodal representations are analyzed from the Math Data Corpus, which involves multimodal data (digital pen, speech, images) on collaborating students as they solve math problems. Findings reveal that in 96-97% of cases the correctness of a group's solution was predictable in advance based on students' written work content. In addition, a linear regression revealed that 65% of the variance in individual students' domain expertise rankings could be accounted for based on their written work content. A multimodal content analysis based on both written and spoken input correctly predicted the dominant domain expert in a group 100% of the time, exceeding unimodal prediction rates. Further analysis revealed a reversal between experts and non-experts in the percentage of time that a match versus mismatch was present between their oral and written answer contributions, with non-experts demonstrating higher mismatches. Implications are discussed for developing reliable multimodal learning analytics systems that incorporate digital pen input to automatically track consolidation of domain expertise.
Sharon L. Oviatt, Adrienne Cohen
ICMI1
2013 Multimodal learning analytics: description of math data corpus for ICMI grand challenge workshop
abstract
This paper provides documentation on dataset resources for establishing a new research area called multimodal learning analytics (MMLA). Research on this topic has the potential to transform the future of educational practice and technology, as well as computational techniques for advancing data analytics. The Math Data Corpus includes high-fidelity time-synchronized multimodal data recordings (speech, digital pen, images) on collaborating groups of students as they work together to solve mathematics problems that vary in difficulty level. The Math Data Corpus resources include initial coding of problem segmentation, problem-solving correctness, and representational content of students' writing. These resources are made available to participants in the data-driven grand challenge for the Second International Workshop on Multimodal Learning Analytics. The primary goal of this event is to analyze coherent signal, activity, and lexical patterns that can identify domain expertise and change in domain expertise early, reliably, and objectively, as well as learning-oriented precursors. An additional aim is to build an international research community in the emerging area of multimodal learning analytics by organizing a series of workshops that bring together multidisciplinary scientists to work on MMLA topics.
Sharon L. Oviatt, Adrienne Cohen, Nadir Weibel
ICMI1
2012 The impact of interface affordances on human ideation, problem solving, and inferential reasoning
abstract
This article presents two studies investigating how computer interface affordances influence basic cognition, including ideational fluency, problem solving, and inferential reasoning. In one study comparing interfaces with different input capabilities, students expressed 56% more nonlinguistic representations (diagrams, symbols, numbers) when using pen interfaces. A linear regression confirmed that nonlinguistic communication directly mediated a substantial increase (38.5%) in students' ability to produce appropriate science ideas. In contrast, students expressed 41% more linguistic content when using a keyboard-based interface, which mediated a drop in science ideation. A follow-up study pursued the question of how interfaces that prime nonlinguistic communication so effectively facilitate cognition. This study examined the relation between students' expression of nonlinguistic representations and their inference accuracy when using analogous digital and non-digital pen tools. Perhaps surprisingly, the digital pen interface stimulated construction of more diagrams, more correct Venn diagrams, and more accurate domain inferences. Students' construction of multiple diagrams to represent a problem also directly suppressed overgeneralization errors, which were the most common inference failure. These research results reveal that computer interfaces have communications affordances which elicit communication patterns that can substantially stimulate or impede basic cognition. Implications are discussed for designing new digital tools for thinking, with an emphasis on nonlinguistic and especially spatial representations that are most poorly supported by current keyboard-based interfaces.
Sharon L. Oviatt, Adrienne Cohen, Andrea Miller, Kumi Hodge, Ariana Mann
ACM Trans. Comput. Hum. Interact.1
2011 Designing Interfaces that Stimulate Ideational Fluency in Science
Sharon L. Oviatt, Adrienne Cohen, Ariana Mann
CogSci1
2008 Implicit user-adaptive system engagement in speech and pen interfaces
abstract
As emphasis is placed on developing mobile, educational, and other applications that minimize cognitive load on users, it is becoming more essential to explore interfaces based on implicit engagement techniques so users can remain focused on their tasks. In this research, data were collected with 12 pairs of students who solved complex math problems using a tutorial system that they engaged over 100 times per session entirely implicitly via speech amplitude or pen pressure cues. Results revealed that users spontaneously, reliably, and substantially adapted these forms of communicative energy to designate and repair an intended interlocutor in a computer-mediated group setting. Furthermore, this behavior was harnessed to achieve system engagement accuracies of 75-86%, with accuracies highest using speech amplitude. However, students had limited awareness of their own adaptations. Finally, while continually using these implicit engagement techniques, students maintained their performance level at solving complex mathematics problems throughout a one-hour session.
Sharon L. Oviatt, Colin Swindells, Alexander M. Arthur
CHI1
2008 A high-performance dual-wizard infrastructure for designing speech, pen, and multimodal interfaces
abstract
The present paper reports on the design and performance of a novel dual-Wizard simulation infrastructure that has been used effectively to prototype next-generation adaptive and implicit multimodal interfaces for collaborative groupwork. This high-fidelity simulation infrastructure builds on past development of single-wizard simulation tools for multiparty multimodal interactions involving speech, pen, and visual input [1]. In the new infrastructure, a dual-wizard simulation environment was developed that supports (1) real-time tracking, analysis, and system adaptivity to a user's speech and pen paralinguistic signal features (e.g., speech amplitude, pen pressure), as well as the semantic content of their input. This simulation also supports (2) transparent user training to adapt their speech and pen signal features in a manner that enhances the reliability of system functioning, i.e., the design of mutually-adaptive interfaces. To accomplish these objectives, this new environment also is capable of handling (3) dynamic streaming digital pen input. We illustrate the performance of the simulation infrastructure during longitudinal empirical research in which a user-adaptive interface was designed for implicit system engagement based exclusively on users' speech amplitude and pen pressure [2]. While using this dual-wizard simulation method, the wizards responded successfully to over 3,000 user inputs with 95-98% accuracy and a joint wizard response time of less than 1.0 second during speech interactions and 1.65 seconds during pen interactions. Furthermore, the interactions they handled involved naturalistic multiparty meeting data in which high school students were engaged in peer tutoring, and all participants believed they were interacting with a fully functional system. This type of simulation capability enables a new level of flexibility and sophistication in multimodal interface design, including the development of implicit multimodal interfaces that place minimal cognitive load on users during mobile, educational, and other applications.
Phil Cohen 0001, Colin Swindells, Sharon L. Oviatt, Alexander M. Arthur
ICMI3
2007 Implicit user-adaptive system engagement in speech, pen and multimodal interfaces
abstract
The present research contributes new empirical research, theory, and prototyping toward developing implicit user-adaptive techniques for system engagement based exclusively on speech amplitude and pen pressure. The results reveal that people will spontaneously adapt their communicative energy level reliably, substantially, and in different modalities to designate and repair an intended interlocutor in a computer-mediated group setting. Furthermore, this sole behavior can be harnessed to achieve system engagement accuracies in the 75 - 86 % range. In short, there was a high level of correct system engagement based exclusively on implicit cues in users' energy level during communication.
Sharon L. Oviatt
ASRU1
2006 Prototyping novel collaborative multimodal systems: simulation, data collection and analysis tools for the next decade
abstract
To support research and development of next-generation multimodal interfaces for complex collaborative tasks, a comprehensive new infrastructure has been created for collecting and analyzing time-synchronized audio, video, and pen-based data during multi-party meetings. This infrastructure needs to be unobtrusive and to collect rich data involving multiple information sources of high temporal fidelity to allow the collection and annotation of simulation-driven studies of natural human-human-computer interactions. Furthermore, it must be flexibly extensible to facilitate exploratory research. This paper describes both the infrastructure put in place to record, encode, playback and annotate the meeting-related media data, and also the simulation environment used to prototype novel system concepts.
Alexander M. Arthur, Rebecca Lunsford, Matt Wesson, Sharon L. Oviatt
ICMI4
2006 Human perception of intended addressee during computer-assisted meetings
abstract
Recent research aims to develop new open-microphone engagement techniques capable of identifying when a speaker is addressing a computer versus human partner, including during computer-assisted group interactions. The present research explores: (1) how accurately people can judge whether an intended interlocutor is a human versus computer, (2) which linguistic, acoustic-prosodic, and visual information sources they use to make these judgments, and (3) what type of systematic errors are present in their judgments. Sixteen participants were asked to determine a speaker's intended addressee based on actual videotaped utterances matched on illocutionary force, which were played back as: (1) lexical transcriptions only, (2) audio-only, (3) visual-only, and (4) audio-visual information. Perhaps surprisingly, people's accuracy in judging human versus computer addressees did not exceed chance levels with lexical-only content (46%). As predicted, accuracy improved significantly with audio (58%), visual (57%), and especially audio-visual information (63%). Overall, accuracy in detecting human interlocutors was significantly worse than judging computer ones, and specifically worse when only visual information was present because speakers often looked at the computer when addressing peers. In contrast, accuracy in judging computer interlocutors was significantly better whenever visual information was present than with audio alone, and it yielded the highest accuracy levels observed (86%). Questionnaire data also revealed that speakers' gaze, peers' gaze, and tone of voice were considered the most valuable information sources. These results reveal that people rely on cues appropriate for interpersonal interactions in determining computer- versus human-directed speech during mixed human-computer interactions, even though this degrades their accuracy. Future systems that process actual rather than expected communication patterns potentially could be designed that perform better than humans.
Rebecca Lunsford, Sharon L. Oviatt
ICMI2
2006 Toward open-microphone engagement for multiparty interactions
abstract
There currently is considerable interest in developing new open-microphone engagement techniques for speech and multimodal interfaces that perform robustly in complex mobile and multiparty field environments. State-of-the-art audio-visual open-microphone engagement systems aim to eliminate the need for explicit user engagement by processing more implicit cues that a user is addressing the system, which results in lower cognitive load for the user. This is an especially important consideration for mobile and educational interfaces due to the higher load required by explicit system engagement. In the present research, longitudinal data were collected with six triads of high-school students who engaged in peer tutoring on math problems with the aid of a simulated computer assistant. Results revealed that amplitude was 3.25dB higher when users addressed a computer rather than human peer when no lexical marker of intended interlocutor was present, and 2.4dB higher for all data. These basic results were replicated for both matched and adjacent utterances to computer versus human partners. With respect to dialogue style, speakers did not direct a higher ratio of commands to the computer, although such dialogue differences have been assumed in prior work. Results of this research reveal that amplitude is a powerful cue marking a speaker's intended addressee, which should be leveraged to design more effective microphone engagement during computer-assisted multiparty interactions.
Rebecca Lunsford, Sharon L. Oviatt, Alexander M. Arthur
ICMI2
2006 Human-centered design meets cognitive load theory: designing interfaces that help people think
abstract
Historically, the development of computer systems has been primarily a technology-driven phenomenon, with technologists believing that "users can adapt" to whatever they build. Human-centered design advocates that a more promising and enduring approach is to model users' natural behavior to begin with so that interfaces can be designed that are more intuitive, easier to learn, and freer of performance errors. In this paper, we illustrate different user-centered design principles and specific strategies, as well as their advantages and the manner in which they enhance users' performance. We also summarize recent research findings from our lab comparing the performance characteristics of different educational interfaces that were based on user-centered design principles. One theme throughout our discussion is human-centered design that minimizes users' cognitive load, which effectively frees up mental resources for performing better while also remaining more attuned to the world around them.
Sharon L. Oviatt
ACM Multimedia1
2006 Quiet interfaces that help students think
abstract
As technical as we have become, modern computing has not permeated many important areas of our lives, including mathematics education which still involves pencil and paper. In the present study, twenty high school geometry students varying in ability from low to high participated in a comparative assessment of math problem solving using existing pencil and paper work practice (PP), and three different interfaces: an Anoto-based digital stylus and paper interface (DP), pen tablet interface (PT), and graphical tablet interface (GT). Cognitive Load Theory correctly predicted that as interfaces departed more from familiar work practice (GT > PT > DP), students would experience greater cognitive load such that performance would deteriorate in speed, attentional focus, meta-cognitive control, correctness of problem solutions, and memory. In addition, low-performing students experienced elevated cognitive load, with the more challenging interfaces (GT, PT) disrupting their performance disproportionately more than higher performers. The present results indicate that Cognitive Load Theory provides a coherent and powerful basis for predicting the rank ordering of users' performance by type of interface. In the future, new interfaces for areas like education and mobile computing could benefit from designs that minimize users' load so performance is more adequately supported.
Sharon L. Oviatt, Alexander M. Arthur, Julia Cohen
UIST1
2005 Individual differences in multimodal integration patterns: what are they and why do they exist?
abstract
Techniques for information fusion are at the heart of multimodal system design. To develop new user-adaptive approaches for multimodal fusion, the present research investigated the stability and underlying cause of major individual differences that have been documented between users in their multimodal integration pattern. Longitudinal data were collected from 25 adults as they interacted with a map system over six weeks. Analyses of 1,100 multimodal constructions revealed that everyone had a dominant integration pattern, either simultaneous or sequential, which was 95-96% consistent and remained stable over time. In addition, coherent behavioral and linguistic differences were identified between these two groups. Whereas performance speed was comparable, sequential integrators made only half as many errors and excelled during new or complex tasks. Sequential integrators also had more precise articulation (e.g., fewer disfluencies), although their speech rate was no slower. Finally, sequential integrators more often adopted terse and direct command-style language, with a smaller and less varied vocabulary, which appeared focused on achieving error-free communication. These distinct interaction patterns are interpreted as deriving from fundamental differences in reflective-impulsive cognitive style. Implications of these findings are discussed for the design of adaptive multimodal systems with substantially improved performance characteristics.
Sharon L. Oviatt, Rebecca Lunsford, Rachel Coulston
CHI1
2005 Audio-visual cues distinguishing self- from system-directed speech in younger and older adults
abstract
In spite of interest in developing robust open-microphone engagement techniques for mobile use and natural field contexts, there currently are no reliable techniques available. One problem is the lack of empirically-grounded models as guidance for distinguishing how users' audio-visual activity actually differs systematically when addressing a computer versus human partner. In particular, existing techniques have not been designed to handle high levels of user self talk as a source of "noise," and they typically assume that a user is addressing the system only when facing it while speaking. In the present research, data were collected during two related studies in which adults aged 18-89 interacted multimodally using speech and pen with a simulated map system. Results revealed that people engaged in self talk prior to addressing the system over 30% of the time, with no decrease in younger adults' rate of self talk compared with elders. Speakers' amplitude was lower during 96% of their self talk, with a substantial 26 dBr amplitude separation observed between self- and system-directed speech. The magnitude of speaker's amplitude separation ranged from approximately 10-60 dBr and diminished with age, with 79% of the variance predictable simply by knowing a person's age. In contrast to the clear differentiation of intended addressee revealed by amplitude separation, gaze at the system was not a reliable indicator of speech directed to the system, with users looking at the system over 98% of the time during both self- and system-directed speech. Results of this research have implications for the design of more effective open-microphone engagement for mobile and pervasive systems.
Rebecca Lunsford, Sharon L. Oviatt, Rachel Coulston
ICMI2
2004 When do we interact multimodally?: cognitive load and multimodal communication patterns
abstract
Mobile usage patterns often entail high and fluctuating levels of difficulty as well as dual tasking. One major theme explored in this research is whether a flexible multimodal interface supports users in managing cognitive load. Findings from this study reveal that multimodal interface users spontaneously respond to dynamic changes in their own cognitive load by shifting to multimodal communication as load increases with task difficulty and communicative complexity. Given a flexible multimodal interface, users' ratio of multimodal (versus unimodal) interaction increased substantially from 18.6% when referring to established dialogue context to 77.1% when required to establish a new context, a +315% relative increase. Likewise, the ratio of users' multimodal interaction increased significantly as the tasks became more difficult, from 59.2% during low difficulty tasks, to 65.5% at moderate difficulty, 68.2% at high and 75.0% at very high difficulty, an overall relative increase of +27%. Analysis of users' task-critical errors and response latencies across task difficulty levels increased systematically and significantly as well, corroborating the manipulation of cognitive processing load. The adaptations seen in this study reflect users' efforts to self-manage limitations on working memory when task complexity increases. This is accomplished by distributing communicative information across multiple modalities, which is compatible with a cognitive load theory of multimodal interaction. The long-term goal of this research is the development of an empirical foundation for proactively guiding flexible and adaptive multimodal system design.
Sharon L. Oviatt, Rachel Coulston, Rebecca Lunsford
ICMI1
2004 Toward adaptive conversational interfaces: Modeling speech convergence with animated personas
abstract
The design of robust interfaces that process conversational speech is a challenging research direction largely because users' spoken language is so variable. This research explored a new dimension of speaker stylistic variation by examining whether users' speech converges systematically with the text-to-speech (TTS) heard from a software partner. To pursue this question, a study was conducted in which twenty-four 7 to 10-year-old children conversed with animated partners that embodied different TTS voices. An analysis of children's amplitude, durational features, and dialogue response latencies confirmed that they spontaneously adapt several basic acoustic-prosodic features of their speech 10--50%, with the largest adaptations involving utterance pause structure and amplitude. Children's speech adaptations were relatively rapid, bidirectional, and dynamically readaptable when introduced to new partners, and generalized across different types of users and TTS voices. Adaptations also occurred consistently, with 70--95% of children converging with their partner's TTS, although individual differences in magnitude of adaptation were evident. In the design of future conversational systems, users' spontaneous convergence could be exploited to guide their speech within system processing bounds, thereby enhancing robustness. Adaptive system processing could yield further significant performance gains. The long-term goal of this research is the development of predictive models of human-computer communication to guide the design of new conversational interfaces.
Sharon L. Oviatt, Courtney Darves, Rachel Coulston
ACM Trans. Comput. Hum. Interact.1
2004 Introduction to mobile and adaptive conversational interfaces
abstract
introduction Introduction to mobile and adaptive conversational interfaces Share on Authors: Sharon Oviatt Oregon Health & Science University, Beaverton, OR Oregon Health & Science University, Beaverton, ORView Profile , Stephanie Seneff Massachusetts Institute of Technology, Cambridge, MA Massachusetts Institute of Technology, Cambridge, MAView Profile Authors Info & Claims ACM Transactions on Computer-Human InteractionVolume 11Issue 3September 2004 pp 237–240https://doi.org/10.1145/1017494.1017495Online:01 September 2004Publication History 3citation1,848DownloadsMetricsTotal Citations3Total Downloads1,848Last 12 Months10Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Sharon L. Oviatt, Stephanie Seneff
ACM Trans. Comput. Hum. Interact.1
2003 Toward a theory of organized multimodal integration patterns during human-computer interaction
abstract
As a new generation of multimodal systems begins to emerge, one dominant theme will be the integration and synchronization requirements for combining modalities into robust whole systems. In the present research, quantitative modeling is presented on the organization of users' speech and pen multimodal integration patterns. In particular, the potential malleability of users' multimodal integration patterns is explored, as well as variation in these patterns during system error handling and tasks varying in difficulty. Using a new dual-wizard simulation method, data was collected from twelve adults as they interacted with a map-based task using multimodal speech and pen input. Analyses based on over 1600 multimodal constructions revealed that users' dominant multimodal integration pattern was resistant to change, even when strong selective reinforcement was delivered to encourage switching from a sequential to simultaneous integration pattern, or vice versa. Instead, both sequential and simultaneous integrators showed evidence of entrenching further in their dominant integration patterns (i.e., increasing either their inter-modal lag or signal overlap) over the course of an interactive session, during system error handling, and when completing increasingly difficult tasks. In fact, during error handling these changes in the co-timing of multimodal signals became the main feature of hyper-clear multimodal language, with elongation of individual signals either attenuated or absent. Whereas Behavioral/Structuralist theory cannot account for these data, it is argued that Gestalt theory provides a valuable framework and insights into multimodal interaction. Implications of these findings are discussed for the development of a coherent theory of multimodal integration during human-computer interaction, and for the design of a new class of adaptive multimodal interfaces.
Sharon L. Oviatt, Rachel Coulston, Stefanie Tomko, Benfang Xiao, Rebecca Lunsford, R. Matthews Wesson, Lesley Carmichael
ICMI1
2003 Modeling multimodal integration patterns and performance in seniors: toward adaptive processing of individual differences
abstract
Multimodal interfaces are designed with a focus on flexibility, although very few currently are capable of adapting to major sources of user, task, or environmental variation. The development of adaptive multimodal processing techniques will require empirical guidance from quantitative modeling on key aspects of individual differences, especially as users engage in different types of tasks in different usage contexts. In the present study, data were collected from fifteen 66- to 86-year-old healthy seniors as they interacted with a map-based flood management system using multimodal speech and pen input. A comprehensive analysis of multimodal integration patterns revealed that seniors were classifiable as either simultaneous or sequential integrators, like children and adults. Seniors also demonstrated early predictability and a high degree of consistency in their dominant integration pattern. However, greater individual differences in multimodal integration generally were evident in this population. Perhaps surprisingly, during sequential constructions seniors' intermodal lags were no longer in average and maximum duration than those of younger adults, although both of these groups had longer maximum lags than children. However, an analysis of seniors' performance did reveal lengthy latencies before initiating a task, and high rates of self talk and task-critical errors while completing spatial tasks. All of these behaviors were magnified as the task difficulty level increased. Results of this research have implications for the design of adaptive processing strategies appropriate for seniors' applications, especially for the development of temporal thresholds used during multimodal fusion. The long-term goal of this research is the design of high-performance multimodal systems that adapt to a full spectrum of diverse users, supporting tailored and robust future systems.
Benfang Xiao, Rebecca Lunsford, Rachel Coulston, R. Matthews Wesson, Sharon L. Oviatt
ICMI5
2003 User-centered modeling and evaluation of multimodal interfaces
abstract
Historically, the development of computer interfaces has been a technology-driven phenomenon. However, new multimodal interfaces are composed of recognition-based technologies that must interpret human speech, gesture, gaze, movement patterns, and other complex natural behaviors, which involve highly automatized skills that are not under full conscious control. As a result, it now is widely acknowledged that multimodal interface design requires modeling of the modality-centered behavior and integration patterns upon which multimodal systems aim to build. This paper summarizes research on the cognitive science foundations of multimodal interaction, and on the essential role that user-centered modeling has played in prototyping, guiding, and evaluating the design of next-generation multimodal interfaces. In particular, it discusses the properties of different modalities and the information content they carry, the unique features of multimodal language and its processability, as well as when users are likely to interact multimodally and how their multimodal input is integrated and synchronized. It also reviews research on typical performance and linguistic efficiencies associated with multimodal interaction, and on the user-centered reasons why multimodal interaction minimizes errors and expedites error handling. In addition, this paper describes the important role that selective methodologies and evaluation metrics have played in shaping next-generation multimodal systems, and it concludes by highlighting future directions for designing a new class of adaptive multimodal-multisensor interfaces.
Sharon L. Oviatt
Proc. IEEE1
2002 Amplitude convergence in children²s conversational speech with animated personas
abstract
During interpersonal conversation, both children and adults adapt the basic acoustic-prosodic features of their speech to converge with those of their conversational partner. In this study, 7-to-10year-old children interacted with a conversational interface in which animated characters used text-to-speech output (TTS) to answer questions about marine biology. Analysis of children’s speech to different animated characters revealed a 29% average change in energy when they spoke to an extroverted loud software partner (E), compared with an introverted soft-spoken one (I). The majority, or 77% of children, adapted their amplitude toward their partner’s TTS voice. These adaptations were bi-directional, with increases in amplitude observed during I to E condition shifts, and decreases during E to I shifts. Finally, these results generalized across different user groups and TTS voices. Implications are discussed for guiding children’s speech to remain within system processing bounds, and for the future development of robust and adaptive conversational interfaces.
Rachel Coulston, Sharon L. Oviatt, Courtney Darves
INTERSPEECH2
2002 Adaptation of users² spoken dialogue patterns in a conversational interface
abstract
The design of robust new interfaces that process conversational speech is a challenging research direction largely because users’ spoken language is so variable, which is especially true of children. The present research explored whether children’s response latencies before initiating a conversational turn converge with those heard in the text-to-speech (TTS) of a computer partner. A study was conducted in which twenty-four 7-to-10-year-old children conversed with animated characters that responded with different types of TTS voices during an educational software application. Analyses confirmed that, while interacting with opposite TTS voices, children’s average response latencies adapted 18.4% in the direction of their computer partner’s speech. These adaptations were dynamic, bi-directional, and generalized across different types of users and TTS voices. The long-term goal of this research is the predictive modeling of human-computer communication patterns to guide the design of well synchronized, robust, and adaptive conversational interfaces.
Courtney Darves, Sharon L. Oviatt
INTERSPEECH2
2002 Multimodal integration patterns in children
abstract
Multimodal interfaces are designed with a focus on flexibility, although very few multimodal systems currently are capable of adapting to major sources of user or environmental variation. The development of adaptive multimodal processing techniques will require empirical guidance on modeling key aspects of individual differences. In the present study, we collected data from 24 7-to-10-year-old children as they interacted using speech and pen input with an educational software prototype. A comprehensive analysis of children's multimodal integration patterns revealed that they were classifiable as either simultaneous or sequential integrators, although they more often integrated signals simultaneously than adults. During their sequential constructions, intermodal lags also ranged faster than those of adult users. The high degree of consistency and early predictability of children's integration patterns were similar to previously reported adult data. These results have implications for the development of temporal thresholds and adaptive multimodal processing strategies for children's applications. The long-term goal of this research is life-span modeling of users' integration and synchronization patterns, which will be needed to design future high-performance adaptive multimodal systems.
Benfang Xiao, Cynthia Girand, Sharon L. Oviatt
INTERSPEECH3
2002 From members to teams to committee-a robust approach to gestural and multimodal recognition
abstract
When building a complex pattern recognizer with high-dimensional input features, a number of selection uncertainties arise. Traditional approaches to resolving these uncertainties typically rely either on the researcher's intuition or performance evaluation on validation data, both of which result in poor generalization and robustness on test data. This paper describes a novel recognition technique called members to teams to committee (MTC), which is designed to reduce modeling uncertainty. In particular, the MTC posterior estimator is based on a coordinated set of divide-and-conquer estimators that derive from a three-tiered architectural structure corresponding to individual members, teams, and the overall committee. Basically, the MTC recognition decision is determined by the whole empirical posterior distribution, rather than a single estimate. This paper describes the application of the MTC technique to handwritten gesture recognition and multimodal system integration and presents a comprehensive analysis of the characteristics and advantages of the MTC approach.
Lizhong Wu, Sharon L. Oviatt, Phil Cohen 0001
IEEE Trans. Neural Networks2
2000 Multimodal signal processing in naturalistic noisy environments
abstract
When a system must process spoken language in natural environments that involve different types and levels of noise, the problem of supporting robust recognition is a very difficult one. In the present studies, over 2,600 multimodal utterances were collected during both mobile and stationary use of a multimodal pen/voice system. The results confirmed that multimodal signal processing supports significantly improved robustness over spoken language processing alone, with the largest improvement during mobile use. The multimodal architecture decreased the spoken language error rate by 19-35%. In addition, data collected on a command-by-command basis while users were mobile emphasized the adverse impact of users' Lombard adaptation on system processing, even when a noise-canceling microphone was used. Implications of these findings are discussed for improving the reliability and stability of spoken language processing in mobile environments.
Sharon L. Oviatt
INTERSPEECH1
2000 Multimodal interface research: a science without borders
abstract
Multimodal research represents "Science without Borders" because it requires combining expertise from different component technologies, academic disciplines, and cultural/international perspectives. It also is rapidly erasing borders as it promotes the increased accessibility of computing for diverse and non-specialist users, and for field and mobile usage environments. This paper reviews two studies that highlight recent advances within the field. It also draws parallels between the multimodal areas of speech/pen and speech/lip movement research. Finally, it indicates new research challenges that will require additional bold "border crossings" in the near future. In the medical community, there is an international group called Physicians without Borders that many of you undoubtedly are familiar with (URL: http://www.dwb.org/). Physicians without Borders//Medecins sans Frontiers is an organization of volunteer medical personnel who respond to medical needs and emergencies around the w...
Sharon L. Oviatt
INTERSPEECH1
2000 Talking to thimble jellies: children²s conversational speech with animated characters
Sharon L. Oviatt
INTERSPEECH1
2000 Multimodal system processing in mobile environments
abstract
One major goal of multimodal system design is to support more robust performance than can be achieved with a unimodal recognition technology, such as a spoken language system.In recent years, the multimodal literatures on speech and pen input and speech and lip movements have begun developing relevant performance criteria and demonstrating a reliability advantage for multimodal architectures.In the present studies, over 2,600 utterances processed by a multimodal pen/voice system were collected during both mobile and stationary use.A new data collection infrastructure was developed, including instrumentation worn by the user while roaming, a researcher field station, and a multimodal data logger and analysis tool tailored for mobile research.Although speech recognition as a stand-alone failed more often during mobile system use, the results confirmed that a more stable multimodal architecture decreased this error rate by 19-35%.Furthermore, these findings were replicated across different types of microphone technology.In large part this performance gain was due to significant levels of mutual disambiguation in the multimodal architecture, with higher levels occurring in the noisy mobile environment.Implications of these findings are discussed for expanding computing to support more challenging usage contexts in a robust manner.
Sharon L. Oviatt
UIST1
2000 Designing the User Interface for Multimodal Speech and Pen-Based Gesture Applications: State-of-the-Art Systems and Future Research Directions
abstract
The growing interest in multimodal interface design is inspired in large part by the goals of supporting more transparent, flexible, efficient, and powerfully expressive means of human-computer interaction than in the past. Multimodal interfaces are expected to support a wider range of diverse applications, be usable by a broader spectrum of the average population, and function more reliably under realistic and challenging usage conditions. In this article, we summarize the emerging architectural approaches for interpreting speech and pen-based gestural input in a robust manner-including early and late fusion approaches, and the new hybrid symbolic-statistical approach. We also describe a diverse collection of state-of-the-art multimodal systems that process users' spoken and gestural input. These applications range from map-based and virtual reality systems for engaging in simulations and training, to field medic systems for mobile use in noisy environments, to web-based transactions and standard text-editing applications that will reshape daily computing and have a significant commercial impact. To realize successful multimodal systems of the future, many key research challenges remain to be addressed. Among these challenges are the development of cognitive theories to guide multimodal system design, and the development of effective natural language processing, dialogue processing, and error-handling techniques. In addition, new multimodal systems will be needed that can function more robustly and adaptively, and with support for collaborative multiperson use. Before this new class of systems can proliferate, toolkits also will be needed to promote software development for both simulated and functioning systems.
Sharon L. Oviatt, Phil Cohen 0001, Lizhong Wu, Lisbeth Duncan, Bernhard Suhm, Josh Bers, Thomas G. Holzman, Terry Winograd, James A. Landay, Jim Larson, David L. Ferro
Hum. Comput. Interact.1
1999 Mutual Disambiguation of Recognition Errors in a Multimodel Architecture
abstract
As a new generation of multimodal/media systems begins to define itself, researchers are attempting to learn how to combine different modes into strategically integrated whole systems. In theory, well designed multimodal systems should be able to integrate complementary modalities in a manner that supports mutual disambiguation (MD) of errors and leads to more robust performance. In this study, over 2,000 multimodal utterances by both native and accented speakers of English were processed by a multimodal system, and then logged and analyzed. The results confirmed that multimodal systems can indeed support significant levels of MD, and also higher levels of MD for the more challenging accented users. As a result, although speech recognition as a stand-alone performed far more poorly for accented speakers, their multimodal recognition rates did not differ from those of native speakers. Implications are discussed for the development of future multimodal architectures that can perform in a more robust and stable manner than individual recognition technologies. Also discussed is the design of interfaces that support diversity in tangible ways, and that function well under challenging real-world usage conditions,
Sharon L. Oviatt
CHI1
1999 Multimodal Integration - A Statistical View
abstract
We present a statistical approach to developing multimodal recognition systems and, in particular, to integrating the posterior probabilities of parallel input signals involved in the multimodal system. We first identify the primary factors that influence multimodal recognition performance by evaluating the multimodal recognition probabilities. We then develop two techniques, an estimate approach and a learning approach, which are designed to optimize accurate recognition during the multimodal integration process. We evaluate these methods using Quickset, a speech/gesture multimodal system, and report evaluation results based on an empirical corpus collected with Quickset. From an architectural perspective, the integration technique presented offers enhanced robustness. It also is premised on more realistic assumptions than previous multimodal systems using semantic fusion. From a methodological standpoint, the evaluation techniques that we describe provide a valuable tool for evaluating multimodal systems.
Lizhong Wu, Sharon L. Oviatt, Phil Cohen 0001
IEEE Trans. Multim.2
1998 STAMP: a suite of tools for analyzing multimodal system processing
abstract
In this paper we describe a new automated suite of tools for capturing and analyzing data on multimodal systems called STAMP. STAMP is
Josh Clow, Sharon L. Oviatt
ICSLP2
1998 The efficiency of multimodal interaction: a case study
abstract
This paper reports 1 on a case study comparison of a directmanipulation-based graphical user interface (GUI) with the QuickSet pen/voice multimodal interface for supporting the task of military force “laydown. ” In this task, a user places military units and “control measures, ” such as various types of lines, obstacles, objectives, etc., on a map. A military expert designed his own scenario and entered it via both interfaces. Usage of QuickSet led to a speed improvement of 3.2 to 8.7fold, depending on the kind of object being created. These results suggest that there may be substantial efficiency advantages to using multimodal interaction over GUIs for mapbased tasks. 1.
Phil Cohen 0001, Michael Johnston, David McGee, Sharon L. Oviatt, Josh Clow, Ira A. Smith
ICSLP4
1998 The CHAM model of hyperarticulate adaptation during human-computer error resolution
Sharon L. Oviatt
ICSLP1
1998 Referential features and linguistic indirection in multimodal language
abstract
The present report outlines differences between multimodal and unimodal communication patterns in linguistic features associated with ease of dialogue tracking and ambiguity resolution. A simulation method was used to collect data while participants used spoken, pen-based, or multimodal input during spatial tasks with a dynamic system. Users' linguistic constructions were analyzed for differences in the rates of reference, co-reference, definite and indefinite referring expressions, and deictic terms. Differences also were summarized in the prevalence of linguistic indirection. Results indicate that spoken language contains substantially higher levels of referring and co-referring expressions and also linguistic indirection, compared with multimodal language communicated by the same users completing the same task. In contrast, multimodal language not only has fewer referential expressions and relatively little anaphora, it also specifically lacks the regular use of determiners observed in spoken definite and indefinite noun phrases. In addition, multimodal language is distinct in its high levels of deictic reference. Implications of these findings are discussed for the relative ease of natural language processing for speech-only versus multimodal systems.
Sharon L. Oviatt, Karen Kuhn
ICSLP1
1998 Predicting hyperarticulate speech during human-computer error resolution
Sharon L. Oviatt, Margaret MacEachern, Gina-Anne Levow
Speech Commun.1
1997 Unification-based Multimodal Integration
abstract
Recent empirical research has shown conclusive advantages of multimodal interaction over speech-only interaction for map-based tasks. This paper describes a multimodal language processing architecture which supports interfaces allowing simultaneous input from speech and gesture recognition. Integration of spoken and gestural input is driven by unification of typed feature structures representing the semantic contributions of the different modes. This integration method allows the component modalities to mutually compensate for each others' errors. It is implemented in Quick-Set, a multimodal (pen/voice) system that enables users to set up and control distributed interactive simulations.
Michael Johnston, Phil Cohen 0001, David McGee, Sharon L. Oviatt, James A. Pittman, Ira A. Smith
ACL4
1997 Integration and Synchronization of Input Modes during Multimodal Human-Computer Interaction
abstract
Our ability to develop robust multimodal systems will depend on knowledge of the natural integration patterns that typify people's combined use of different input modes. To provide a foundation for theory and design, the present research analyzed multimodal interaction while people spoke and wrote to a simulated dynamic map system. Task analysis revealed that multimodal interaction occurred most frequently during spatial location commands, and with intermediate frequency during selection commands. In addition, microanalysis of input signals identified sequential, simultaneous, point-and-speak, and compound integration patterns, as well as data on the temporal precedence of modes and on inter-modal lags. In synchronizing input streams, the temporal precedence of writing over speech was a major theme, with pen input conveying location information first in a sentence. Linguistic analysis also revealed that the spoken and written modes consistently supplied complementary semantic information, rather than redundant. One long-term goal of this research is the development of predictive models of natural modality integration to guide the design of emerging multimodal architectures. Keywords multimodal interaction, integration and synchronization, speech and pen input, dynamic interactive maps, spatial location information, predictive modeling
Sharon L. Oviatt, Antonella De Angeli, Karen Kuhn
CHI1
1997 QuickSet: Multimodal Interaction for Distributed Applications
abstract
This paper presents an emerging application of multimodal interface research to distributed applications.We have developed the QuickSet prototype, a pen/voice system running on a hand-held PC, communicating via wireless LAN through an agent architecture to a number of systems, including NRaD's' LeatherNet system, a distributed interactive training simulator built for the US Marine Corps.The paper describes the overall system architecture, a novel multimodal integration strategy offering mutual compensation among modalities, and provides examples of multimodal simulation setup.Finally, we discuss our applications experience and evaluation.
Phil Cohen 0001, Michael Johnston, David McGee, Sharon L. Oviatt, Jay Pittman, Ira A. Smith, Josh Clow
ACM Multimedia4
1997 Mulitmodal Interactive Maps: Designing for Human Performance
Sharon L. Oviatt
Hum. Comput. Interact.1
1997 Introduction to This Special Issue on Multimodal Interfaces
Sharon L. Oviatt, Wolfgang Wahlster
Hum. Comput. Interact.1
1996 Multimodal Interfaces for Dynamic Interactive Maps
abstract
Dynamic interactive maps with transparent but power-ful human interface capabilities are beginning to emerge for a variety of geographical information systems, in-cluding ones situated on portables for travelers, stu-dents, business and service people, and others working in field settings. In the present research, interfaces sup-porting spoken, pen-based, and multimodal input were analyze for their potential effectiveness in interacting with this new generation of map systems. Input modal-ity (speech, writing, multimodal) and map display for-mat (highly versus minimally structured) were varied in a within-subject factorial design as people completed re-alistic tasks with a simulated map system. The results identified a constellation of performance difficulties asso-ciated with speech-only map interactions, including ele-vated performance errors, spontaneous disfluencies, and lengthier task completion time-- problems that declined substantially when people could interact multimodally with the map. These performance advantages also mir-rored a strong user preference to interact multimodally. The error-proneness and unacceptability of speech-only input to maps was attributed in large part to people's difficulty generating spoken descriptions of spatial loca-tion. Analyses also indicated that map display format can be used to minimize performance errors and dis-fluencies, and map interfaces that guide users ' speech toward brevity can nearly eliminate disfiuencies. Impli-cations of this research are discussed for the design of high-performance multimodal interfaces for future map systems.
Sharon L. Oviatt
CHI1
1996 Modeling hyperarticulate speech during human-computer error resolution
Sharon L. Oviatt, Gina-Anne Levow, Margaret MacEachern, Karen Kuhn
ICSLP1
1996 Error resolution during multimodal human-computer interaction
Sharon L. Oviatt, Robert VanGent
ICSLP1
1995 Predicting spoken disfluencies during human-computer interaction
Sharon L. Oviatt
Comput. Speech Lang.1
1995 The challenge of spoken language systems: research directions for the nineties
abstract
A spoken language system combines speech recognition, natural language processing and human interface technology. It functions by recognizing the person's words, interpreting the sequence of words to obtain a meaning in terms of the application, and providing an appropriate response back to the user. Potential applications of spoken language systems range from simple tasks, such as retrieving information from an existing database (traffic reports, airline schedules), to interactive problem solving tasks involving complex planning and reasoning (travel planning, traffic routing), to support for multilingual interactions. We examine eight key areas in which basic research is needed to produce spoken language systems: (1) robust speech recognition; (2) automatic training and adaptation; (3) spontaneous speech; (4) dialogue models; (5) natural language response generation; (6) speech synthesis and speech generation; (7) multilingual systems; and (8) interactive multimodal systems. In each area, we identify key research challenges, the infrastructure needed to support research, and the expected benefits. We conclude by reviewing the need for multidisciplinary research, for development of shared corpora and related resources, for computational support and far rapid communication among researchers. The successful development of this technology will increase accessibility of computers to a wide range of users, will facilitate multinational communication and trade, and will create new research specialties and jobs in this rapidly expanding area.>
Ronald A. Cole, Lynette Hirschman, Les E. Atlas, Mary E. Beckman, Alan Biermann, Marcia A. Bush, Mark A. Clements, Jordan Cohen, Oscar Garcia, Brian A. Hanson, Hynek Hermansky, Steve Levinson, Kathy McKeown, Nelson Morgan, David G. Novick, Mari Ostendorf, Sharon L. Oviatt, Patti Price, Harvey F. Silverman, Judy Spitz, Alex Waibel, Clifford J. Weinstein, Stephen A. Zahorian, Victor Zue
IEEE Trans. Speech Audio Process.17
1994 Interface techniques for minimizing disfluent input to spoken language systems
abstract
This research examines spontaneous spoken disfluencies during human-computer interaction, presents a predictive model accounting for their occurrence, and outlines interface techniques for minimizing disfluent input.Data were collected during two empirical studies in which people spoke or wrote to a highly interactive simulated system.The studies were based on a within-subject factorial design in which input modalit y and presentation format were varied.Two separate factors were found to be associated with an increase in speech disfluency rates: length of utterance, and lack of structure in the presentation format.A linear model based on utterance length alone was able to predict 77% of all spoken disfluencies in this research.Therefore, design techniques capable of channeling users' speech into briefer sentences potentially could eliminate most spoken disfluencies.Furthermore, changing the structure of the presentation format successfully eliminated 7070 of all disfluent spoken input.The long-term goal of this research is to provide empirical guidance for the design of robust spoken language technology, which eventually may be formulated as a set of user interface guidelines.
Sharon L. Oviatt
CHI1
1994 Integration themes in multimodal human-computer interaction
abstract
This research examines how people integrate spoken and written input during multimodal human-computer interaction. Three studies used a semi-automatic simulation technique to collect data on people's free use of spoken and written input. Within-subject repeated-measures studies were designed, with data analyzed from 44 subjects and 240 tasks. The primary factors that govern people's selection to write versus speak at given points during a human-computer exchange were evaluated. Analyses revealed that people write digits more often than textual content, and proper names more often than other text. A form-based presentation, in comparison with an unconstrained format, also increased the likelihood of writing. However, the most in#uential factor in patterning people 's integrated use of speech and writing is contrastive functionality, or the use of spoken and written input in a contrastiveway to designate a shift in content or functionality, such as original versus correc...
Sharon L. Oviatt, Erik Olsen
ICSLP1
1994 Toward interface design for human language technology: Modality and structure as determinants of linguistic complexity
Sharon L. Oviatt, Phil Cohen 0001, Michelle Wang
Speech Communication1
1992 A rapid semi-automatic simulation technique for investigating interactive speech and handwriting
abstract
This paper describes a new simulation technique designed to support a wide spectrum of empirical studies on the characteristics of spoken, handwritten, and combined pen/voice input to future interactive systems. The simulatidn tech,fique alms: (1) to provide a tool for investigating interactive handwriting and pen systems, on which no simulation research currently is available, (2) to devise a technique appropriate for comparing people's use of speech and writing, such that differences be- tween these communication modalities and their related technologies can be better understood, and (3) to support a very rapid exchange with simulated speech, pen. and pen/voice sys- tems, such th,-tt interactions can be subject-paced. This paper outli,tes the pecifica,ions, general environment. and capabil- ities of a new semi-automatic simulation technique developed to achieve these goals
Sharon L. Oviatt, Phil Cohen 0001, Martin Fong, Michael Frank 0004
ICSLP1
1990 Spoken language in interpreted telephone dialogues
Sharon L. Oviatt, Phil Cohen 0001, Ann Podlozny
ICSLP1
1989 The Effects of Interaction on Spoken Discourse
abstract
Near-term spoken language systems will likely be limited in their interactive capabilities. To design them, we shall need to model how the presence or absence of speaker interaction influences spoken discourse patterns in different types of tasks. In this research, a comprehensive examination is provided of the discourse structure and performance efficiency of both interactive and noninteractive spontaneous speech in a seriated assembly task. More specifically, telephone dialogues and audiotape monologues are compared, which represent opposites in terms of the opportunity for confirmation feedback and clarification subdialogues. Keyboard communication patterns, upon which most natural language heuristics and algorithms have been based, also are contrasted with patterns observed in the two speech modalities. Finally, implications are discussed for the design of near-term limited-interaction spoken language systems.
Sharon L. Oviatt, Phil Cohen 0001
ACL1