Dan Bohus

dblp:00/4118 · DBLP profile ↗
← Back
34ranked-venue papers
14as first author
3since 2021 · last 2024
0000-0002-6283-0590ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 7 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 13 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Human-computer interaction and pervasive computing
4 papers
Human-robot interaction · 51% Human-AI interaction · 49%
Artificial intelligence
3 papers
Planning, search and constraint satisfaction · 33% Question answering and dialogue systems · 22% Video understanding and tracking · 20%

Topics — the 6 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
decision making under uncertainty
0.212013
Look versus Leap: Computing Value of Information with High-Dimensional Streaming Evidence · IJCAI 2013
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue state tracking
0.212013
Discriminative state tracking for spoken dialog systems · ACL (1) 2013
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › decision making under uncertainty
value of information
0.212013
Look versus Leap: Computing Value of Information with High-Dimensional Streaming Evidence · IJCAI 2013
Human-AI interaction › large language model interaction › language-based interaction
natural language interface
0.112019
Demonstrating a Framework for Rapid Development of Physically Situated Interactive Systems · HRI 2019
Machine learning › Probabilistic and Bayesian machine learning
probabilistic inference
0.012013
Look versus Leap: Computing Value of Information with High-Dimensional Streaming Evidence · IJCAI 2013
Natural language and speech › Question answering and dialogue systems
spoken dialogue systems
0.012013
Discriminative state tracking for spoken dialog systems · ACL (1) 2013

Methods — techniques the papers use, named apart from their topics

mistake detection · 1.3intervention type prediction · 1.3field study · 0.2discriminative state tracking · 0.2
YearPublicationVenuePosition
2024 "Uh, This One?": Leveraging Behavioral Signals for Detecting Confusion during Physical Tasks
abstract
A longstanding goal in the AI and HCI research communities is building intelligent assistants to help people with physical tasks. To be effective in this, AI assistants must be aware of not only the physical environment, but also the human user and their cognitive states. In this paper, we specifically consider the detection of confusion, which we operationalize as the moments when a user is “stuck” and needs assistance. We explore how behavioral features such as gaze, head pose, and hand movements differ between periods of confusion vs no-confusion. We present various modeling approaches for detecting confusion that combine behavioral features, length of time, instructional text embeddings, and egocentric video. Although deep networks (e.g., V-Jepa) trained on full video streams perform well in distinguishing confusion from non-confusion, simpler models leveraging lighter weight behavioral features exhibit similarly high performance, even when generalizing to unseen tasks.
Maia Stiber, Dan Bohus, Sean Andrist
ICMI2
2023 HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real World
abstract
Building an interactive AI assistant that can perceive, reason, and collaborate with humans in the real world has been a long-standing pursuit in the AI community. This work is part of a broader research effort to develop intelligent agents that can interactively guide humans through performing tasks in the physical world. As a first step in this direction, we introduce HoloAssist, a large-scale egocentric human interaction dataset, where two people collaboratively complete physical manipulation tasks. The task performer executes the task while wearing a mixed-reality headset that captures seven synchronized data streams. The task instructor watches the performer’s egocentric video in real time and guides them verbally. By augmenting the data with action and conversational annotations and observing the rich behaviors of various participants, we present key insights into how human assistants correct mistakes, intervene in the task completion procedure, and ground their instructions to the environment. HoloAssist spans 166 hours of data captured by 350 unique instructor-performer pairs. Furthermore, we construct and present benchmarks on mistake detection, intervention type prediction, and hand forecasting, along with detailed analysis. We expect HoloAssist will provide an important resource for building AI assistants that can fluidly collaborate with humans in the real world. Data can be downloaded at https://holoassist.github.io/.
Taein Kwon, Mahdi Rad, Bowen Pan, Ishani Chakraborty, Sean Andrist, Dan Bohus, Ashley Feniello, Bugra Tekin, Felipe Vieira Frujeri, Neel Joshi, Marc Pollefeys
ICCV7
2022 Continual Learning about Objects in the Wild: An Interactive Approach
abstract
We introduce a mixed-reality, interactive approach for continually learning to recognize an open-ended set of objects in a user’s surrounding environment. The proposed approach leverages the multimodal sensing, interaction, and rendering affordances of a mixed-reality headset, and enables users to label nearby objects via speech, gaze, and gestures. Image views of each labeled object are automatically captured from varying viewpoints over time, as the user goes about their everyday tasks. The labels provided by the user can be propagated forward and backwards in time and paired with the collected views to update an object recognition model, in order to continually adapt it to the user’s specific objects and environment. We review key challenges for the proposed interactive continual learning approach, present details of an end-to-end system implementation, and report on results and lessons learned from an initial, exploratory case study using the system.
Dan Bohus, Sean Andrist, Ashley Feniello, Nick Saw, Eric Horvitz
ICMI1
2019 Demonstrating a Framework for Rapid Development of Physically Situated Interactive Systems
abstract
We demonstrate an open, extensible framework for enabling faster development and study of physically situated interactive systems. The framework provides a programming model for parallel coordinated computation centered on temporal streams of data, a set of tools for data visualization and processing, and an open ecosystem of components. The demonstration showcases an interaction toolkit of components for systems that interact with people via natural language in the open world.
Sean Andrist, Dan Bohus, Ashley Feniello
HRI2
2017 Rapid development of multimodal interactive systems: a demonstration of platform for situated intelligence
abstract
We demonstrate an open, extensible platform for developing and studying multimodal, integrative-AI systems. The platform provides a time-aware, stream-based programming model for parallel coordinated computation, a set of tools for data visualization, processing, and learning, and an ecosystem of pluggable AI components. The demonstration will showcase three applications built on this platform and highlight how the platform can significantly accelerate development and research in multimodal interactive systems.
Dan Bohus, Sean Andrist, Mihai Jalobeanu
ICMI1
2016 Are You Messing with Me?: Querying about the Sincerity of Interactions in the Open World
abstract
When interacting with robots deployed in the open world, people may often attempt to engage with them in a playful manner or test their competencies. Such engagements are often associated with language and behaviors that fall outside of designed task capabilities and can lead to interaction failures. Detecting when users are driven by play and curiosity can help a robot to understand why some interactions are breaking down, respond more appropriately by conveying its capabilities to its users, and enhance perceptions of its situational awareness and social intelligence. We have been studying the intentions of everyday users in their engagement with a long-lived robot system that provides directions within an office building. We report on a pilot field-study exploring the use of direct queries to elicit the sincerity of user requests, in terms of their actual need for directions. We discuss early results from this initial study and frame research directions and design implications for robots deployed in the wild.
Sean Andrist, Dan Bohus, Eric Horvitz
HRI2
2015 Incremental Coordination: Attention-Centric Speech Production in a Physically Situated Conversational Agent
abstract
Inspired by studies of human-human conversations, we present methods for incrementally coordinating speech production with listeners' visual foci of attention.We introduce a model that considers the demands and availability of listeners' attention at the onset and throughout the production of system utterances, and that incrementally coordinates speech synthesis with the listener's gaze.We present an implementation and deployment of the model in a physically situated dialog system and discuss lessons learned.
Dan Bohus, Eric Horvitz
SIGDIAL Conference2
2014 Managing Human-Robot Engagement with Forecasts and... um... Hesitations
abstract
We explore methods for managing conversational engagement in open-world, physically situated dialog systems. We investigate a self-supervised methodology for constructing forecasting models that aim to anticipate when participants are about to terminate their interactions with a situated system. We study how these models can be leveraged to guide a disengagement policy that uses linguistic hesitation actions, such as filled and non-filled pauses, when uncertainty about the continuation of engagement arises. The hesitations allow for additional time for sensing and inference, and convey the system's uncertainty. We report results from a study of the proposed approach with a directions-giving robot deployed in the wild.
Dan Bohus, Eric Horvitz
ICMI1
2014 UM3I 2014: International Workshop on Understanding and Modeling Multiparty, Multimodal Interactions
abstract
In this paper, we present a brief summary of the international workshop on Modeling Multiparty, Multimodal Interactions. The UM3I 2014 workshop is held in conjunction with the ICMI 2014 conference. The workshop will highlight recent developments and adopted methodologies in the analysis and modeling of multiparty and multimodal interactions, the design and implementation principles of related human-machine interfaces, as well as the identification of potential limitations and ways of overcoming them.
Samer Al Moubayed, Dan Bohus, Anna Esposito, Dirk Heylen, Maria Koutsombogera, Harris Papageorgiou, Gabriel Skantze
ICMI2
2014 Natural Communication about Uncertainties in Situated Interaction
abstract
Physically situated, multimodal interactive systems must often grapple with uncertainties about properties of the world, people, and their intentions and actions. We present methods for estimating and communicating about different uncertainties in situated interaction, leveraging the affordances of an embodied conversational agent. The approach harnesses a representation that captures both the magnitude and the sources of uncertainty, and a set of policies that select and coordinate the production of nonverbal and verbal behaviors to communicate the system's uncertainties to conversational participants. The methods are designed to enlist participants' help in a natural manner to resolve uncertainties arising during interactions. We report on a preliminary implementation of the proposed methods in a deployed system and illustrate the functionality with a trace from a sample interaction.
Tomislav Pejsa, Dan Bohus, Michael F. Cohen, Chit W. Saw, James Mahoney, Eric Horvitz
ICMI2
2014 Crowdsourcing Language Generation Templates for Dialogue Systems
abstract
We explore the use of crowdsourcing to generate natural language in spoken dia-logue systems. We introduce a method-ology to elicit novel templates from the crowd based on a dialogue seed corpus, and investigate the effect that the amount of surrounding dialogue context has on the generation task. Evaluation is performed both with a crowd and with a system de-veloper to assess the naturalness and suit-ability of the elicited phrases. Results indi-cate that the crowd is able to provide rea-sonable and diverse templates within this methodology. More work is necessary be-fore elicited templates can be automati-cally plugged into the system. 1
Margaret Mitchell, Dan Bohus, Ece Kamar
INLG2
2013 Discriminative state tracking for spoken dialog systems
Angeliki Metallinou, Dan Bohus, Jason D. Williams
ACL (1)2
2013 Execution memory for grounding and coordination
Stephanie Rosenthal, Sarjoun Skaff, Manuela M. Veloso, Dan Bohus, Eric Horvitz
HRI4
2013 A framework for multimodal data collection, visualization, annotation and learning
abstract
The development and iterative refinement of inference models for multimodal systems can be challenging and time intensive. We present a framework for multimodal data collection, visualization, annotation, and learning that enables system developers to build models using various machine learning techniques, and quickly iterate through cycles of development, deployment and refinement.
Anne Loomis Thompson, Dan Bohus
ICMI2
2013 Look versus Leap: Computing Value of Information with High-Dimensional Streaming Evidence
Stephanie Rosenthal, Dan Bohus, Ece Kamar, Eric Horvitz
IJCAI2
2012 Learning speaker, addressee and overlap detection models from multimodal streams
abstract
A key challenge in developing conversational systems is fusing streams of information provided by different sensors to make inferences about the behaviors and goals of people. Such systems can leverage visual and audio information collected through cameras and microphone arrays, including the location of various people, their focus of attention, body pose, the sound source direction, prosody, and speech recognition results. In this paper, we explore discriminative learning techniques for making accurate inferences on the problems of speaker, addressee and overlap detection in multiparty human-computer dialog. The focus is on finding ways to leverage within- and across-signal temporal patterns and to automatically construct representations from the raw streams that are informative for the inference problem. We present a novel extension to traditional decision trees which allows them to incorporate and model temporal signals. We contrast these methods with more traditional approaches where a human expert manually engineers relevant temporal features. The proposed approach performs well even with relatively small amounts of training data, which is of practical importance as designing features that are task dependent is time consuming and not always possible.
Oriol Vinyals, Dan Bohus, Rich Caruana
ICMI2
2012 Crowdsourcing the acquisition of natural language corpora: Methods and observations
abstract
We study the opportunity for using crowdsourcing methods to acquire language corpora for use in natural language processing systems. Specifically, we empirically investigate three methods for eliciting natural language sentences that correspond to a given semantic form. The methods convey frame semantics to crowd workers by means of sentences, scenarios, and list-based descriptions. We discuss various performance measures of the crowdsourcing process, and analyze the semantic correctness, naturalness, and biases of the collected language. We highlight research challenges and directions in applying these methods to acquire corpora for natural language processing applications.
William Yang Wang, Dan Bohus, Ece Kamar, Eric Horvitz
SLT2
2011 Decisions about turns in multiparty conversation: from perception to action
abstract
We present a decision-theoretic approach for guiding turn taking in a spoken dialog system operating in multiparty settings. The proposed methodology couples inferences about multiparty conversational dynamics with assessed costs of different outcomes, to guide turn-taking decisions. Beyond considering uncertainties about outcomes arising from evidential reasoning about the state of a conversation, we endow the system with awareness and methods for handling uncertainties stemming from computational delays in its own perception and production. We illustrate via sample cases how the proposed approach makes decisions, and we investigate the behaviors of the proposed methods via a retrospective analysis on logs collected in a multiparty interaction study.
Dan Bohus, Eric Horvitz
ICMI1
2011 Multiparty Turn Taking in Situated Dialog: Study, Lessons, and Directions
Dan Bohus, Eric Horvitz
SIGDIAL Conference1
2009 Leveraging multiple query logs to improve language models for spoken query recognition
abstract
A voice search system requires a speech interface that can correctly recognize spoken queries uttered by users. The recognition performance strongly relies on a robust language model. In this work, we present the use of multiple data sources, with the focus on query logs, in improving ASR language models for a voice search application. Our contributions are three folds: (1) the use of text queries from web search and mobile search in language modeling; (2) the use of web click data to predict query forms from business listing forms; and (3) the use of voice query logs in creating a positive feedback loop. Experiments show that by leveraging these resources, we can achieve recognition performance comparable to, or even better than, that of a previously deploy system where a large amount of spoken query transcripts are used in language modeling.
Xiao Li 0006, Patrick Nguyen, Geoffrey Zweig, Dan Bohus
ICASSP4
2009 Dialog in the open world: platform and applications
abstract
We review key challenges of developing spoken dialog systems that can engage in interactions with one or multiple participants in relatively unconstrained environments. We outline a set of core competencies for open-world dialog, and describe three prototype systems. The systems are built on a common underlying conversational framework which integrates an array of predictive models and component technologies, including speech recognition, head and pose tracking, probabilistic models for scene analysis, multiparty engagement and turn taking, and inferences about user goals and activities. We discuss the current models and showcase their function by means of a sample recorded interaction, and we review results from an observational study of open-world, multiparty dialog in the wild.
Dan Bohus, Eric Horvitz
ICMI1
2009 Models for Multiparty Engagement in Open-World Dialog
Dan Bohus, Eric Horvitz
SIGDIAL Conference1
2009 Learning to Predict Engagement with a Spoken Dialog System in Open-World Settings
Dan Bohus, Eric Horvitz
SIGDIAL Conference1
2009 The RavenClaw dialog management framework: Architecture and systems
Dan Bohus, Alexander I. Rudnicky
Comput. Speech Lang.1
2008 Structured models for joint decoding of repeated utterances
abstract
Due to speech recognition errors, repetition can be a frequent occurrence in voice-search applications. While a proper treatment of this phenomenon requires the joint modeling of two or more utterances simultaneously, currently deployed systems typically treat the utterances independently. In this paper, we analyze the structure of repetitions and find that in at least one commercial directory assistance application, repetitions follow simple structural transformations more than 70 % of the time. We present preliminary results that suggest that significant gains are possible by explicitly modeling this structure in a joint decoding process. Index Terms: speech recognition, minimum bayes risk, joint decoding, repeated utterances
Geoffrey Zweig, Dan Bohus, Xiao Li 0006, Patrick Nguyen
INTERSPEECH2
2008 Joint n-best rescoring for repeated utterances in spoken dialog systems
abstract
Due to speech recognition errors, repetitions are a frequent phenomenon in spoken dialog systems. In previous work (G. Zweig et al., 2008) we have proposed a joint decoding model that can leverage structural relationships between repeated utterances for improving recognition performance. In this paper we extend this work in two directions. First, we propose a direct, classification-based model for the same task. The new model can leverage features that were fundamentally hard to capture in the previous framework (e.g. spellings, false-starts, etc.) and leads to an additional performance improvement. Second, we show how both models can be used to perform a combined rescoring of two n-best lists that are part of a repetition pair.
Dan Bohus, Geoffrey Zweig, Patrick Nguyen, Xiao Li 0006
SLT1
2007 Estimating the Reliability of MDP Policies: a Confidence Interval Approach
Joel R. Tetreault, Dan Bohus, Diane J. Litman
HLT-NAACL2
2006 Doing research on a deployed spoken dialogue system: one year of let's go! experience
abstract
This paper describes our work with Let’s Go, a telephonebased bus schedule information system that has been in use by the Pittsburgh population since March 2005. Results from several studies show that while task success correlates strongly with speech recognition accuracy, other aspects of dialogue such as turn-taking, the set of error recovery strategies, and the initiative style also significantly impact system performance and user behavior. Index Terms: spoken dialogue systems, real-world applications, speech recognition
Antoine Raux, Dan Bohus, Brian Langner, Alan W. Black, Maxine Eskénazi
INTERSPEECH2
2006 Online Supervised Learning of Non-Understanding Recovery Policies
abstract
Spoken dialog systems typically use a limited number of non- understanding recovery strategies and simple heuristic policies to engage them (e.g. first ask user to repeat, then give help, then transfer to an operator). We propose a supervised, online method for learning a non-understanding recovery policy over a large set of recovery strategies. The approach consists of two steps: first, we construct runtime estimates for the likelihood of success of each recovery strategy, and then we use these estimates to construct a policy. An experiment with a publicly available spoken dialog system shows that the learned policy produced a 12.5% relative improvement in the non-understanding recovery rate.
Dan Bohus, Brian Langner, Antoine Raux, Alan W. Black, Maxine Eskénazi, Alexander I. Rudnicky
SLT1
2005 A principled approach for rejection threshold optimization in spoken dialog systems
abstract
A common design pattern in spoken dialog systems is to reject an input when the recognition confidence score falls below a preset rejection threshold. However, this introduces a potentially non-optimal tradeoff between various types of errors such as misunderstandings and false rejections. In this paper, we propose a data-driven method for determining the relative costs of these errors, and then use these costs to optimize state-specific rejection thresholds. We illustrate the use of this approach with data from a spoken dialog system that handles conference room reservations. The results obtained confirm our intuitions about the costs of the errors, and are consistent with anecdotal evidence gathered throughout the use of the system.
Dan Bohus, Alexander I. Rudnicky
INTERSPEECH1
2005 Let's go public! taking a spoken dialog system to the real world
abstract
In this paper, we describe how a research spoken dialog system was made available to the general public. The Let’s Go Public spoken dialog system provides bus schedule information to the Pittsburgh population during off-peak times. This paper describes the changes necessary to make the system usable for the general public and presents analysis of the calls and strategies we have used to ensure high performance. 1.
Antoine Raux, Brian Langner, Dan Bohus, Alan W. Black, Maxine Eskénazi
INTERSPEECH3
2003 Ravenclaw: dialog management using hierarchical task decomposition and an expectation agenda
abstract
We describe RavenClaw, a new dialog management framework developed as a successor to the Agenda [1] architecture used in the CMU Communicator. RavenClaw introduces a clear separation between task and discourse behavior specification, and allows rapid development of dialog management components for spoken dialog systems operating in complex, goal-oriented domains. The system development effort is focused entirely on the specification of the dialog task, while a rich set of domain-independent conversational behaviors are transparently generated by the dialog engine. To date, RavenClaw has been applied to five different domains allowing us to draw some preliminary conclusions as to the generality of the approach. We briefly describe our experience in developing these systems.
Dan Bohus, Alexander I. Rudnicky
INTERSPEECH1
2001 Is this conversation on track?
abstract
Confidence annotation allows a spoken dialog system to accurately assess the likelihood of misunderstanding at the utterance level and to avoid breakdowns in interaction. We describe experiments that assess the utility of features from the decoder, parser and dialog levels of processing. We also investigate the effectiveness of various classifiers, including Bayesian Networks, Neural Networks, SVMs, Decision Trees, AdaBoost and Naive Bayes, to combine this information into an utterancelevel confidence metric. We found that a combination of a subset of the features considered produced promising results with several of the classification algorithms considered, e.g., our Bayesian Network classifier produced a 45.7% relative reduction in confidence assessment error and a 29.6% reduction relative to a handcrafted rule.
Paul Carpenter 0001, Chun Jin, Rong Zhang 0003, Dan Bohus, Alexander I. Rudnicky
INTERSPEECH5
2000 A Web-based Text Corpora Development System
Dan Bohus, Marian Boldea
LREC1