VLDB 2026 Research / reviewers in the wild / expert
Teruhisa Misu
dblp:13/1119
· DBLP profile ↗
70ranked-venue papers
21as first author
21since 2021 · last 2026
0000-0002-6398-9245ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 14 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 14 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 15 · 1 first-author · 7 since 2021Systems, architecture and hardware · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Promoting Prosocial Interactions between Humans with Autonomous AgentsabstractAs robots and autonomous agents integrate into society, understanding their influence on human social dynamics is crucial. We investigate human–robot interactions, focusing on the impact of prosocial behavior by robots on subsequent human interactions and humans’ willingness to exhibit prosocial behavior toward robots. Our study involved a token-collection game in a grid-world environment. Players, human or robot, could become trapped; a prosocial action involved another player freeing the trapped individual. Findings indicate that robots demonstrating prosocial behavior toward humans can inspire prosocial behavior toward others. Humans also show a notable propensity to assist robots. Witnessing robots engage in prosocial behavior may activate social norms related to cooperation, prompting humans to emulate these behaviors. Robots’ actions could improve the saliency of these acts, focusing people’s attention on prosocial behaviors they might not notice otherwise. Overall, the findings suggest that robots can promote prosocial behavior among humans, contributing to a more cooperative social environment. This research has implications for design and implementation of future autonomous systems, emphasizing the importance of social considerations in human-AI interaction studies. Shashank Mehrotra, Teruhisa Misu, Kumar Akash, Mark Steyvers |
ACM Trans. Hum. Robot Interact. | 3 |
| 2025 | Self-Supervised Learning-Based Multimodal Prediction on Prosocial Behavior IntentionsabstractHuman state detection and behavior prediction have seen significant advancements with the rise of machine learning and multimodal sensing technologies. However, predicting prosocial behavior intentions in mobility scenarios, such as helping others on the road, is an underexplored area. Current research faces a major limitation—there are no large, labeled datasets available for prosocial behavior, and small-scale datasets make it difficult to train deep-learning models effectively. To overcome this, we propose a self-supervised learning approach that harnesses multi-modal data from existing physiological and behavioral datasets. By pre-training our model on diverse tasks and fine-tuning it with a smaller, manually labeled prosocial behavior dataset, we significantly enhance its performance. This method addresses the data scarcity issue, providing a more effective benchmark for prosocial behavior prediction, and offering valuable insights for improving intelligent vehicle systems and human-machine interaction. Abinay Reddy Naini, Zhaobo K. Zheng, Teruhisa Misu, Kumar Akash |
ICASSP | 3 |
| 2025 | Toward Informed AV Decision-Making: Computational Model of Well-being and Trust in MobilityabstractFor future human-autonomous vehicle (AV) interactions to be effective and smooth, human-aware systems that analyze and align human needs with automation decisions are essential. Achieving this requires systems that account for human cognitive states. We present a novel computational model in the form of a Dynamic Bayesian Network (DBN) that infers the cognitive states of both AV users and other road users, integrating this information into the AV's decision-making process. Specifically, our model captures the ``well-being'' of both an AV user and an interacting road user as cognitive states alongside trust. Our DBN models infer beliefs over the AV user’s evolving well-being, trust, and intention states, as well as the possible well-being of other road users, based on observed interaction experiences. Using data collected from an interaction study, we refine the model parameters and empirically assess its performance. Finally, we extend our model into a causal inference model (CIM) framework for AV decision-making, enabling the AV to enhance user well-being and trust while balancing these factors with its own operational costs and the well-being of interacting road users. Our evaluation demonstrates the model’s effectiveness in accurately predicting user's states and guiding informed, human-centered AV decisions. Zahra Zahedi, Shashank Mehrotra, Teruhisa Misu, Kumar Akash |
IJCAI | 3 |
| 2024 | Prosociality Matters: How Does Prosocial Behavior in Interdependent Situations Influence the Well-being and Cognition of Road Users?abstractIn hybrid mobility societies, where automated vehicles (AVs) and humans interact in public spaces, the significance of prosocial behaviors intensifies. These behaviors are crucial for the smooth functioning of an interdependent transportation environment, mitigating challenges from the integration of AVs and human-operated systems, and enhancing user well-being by fostering more efficient, less stressful, and inclusive environments. This study explores the impact of receiving prosocial behaviors on cognition, riding behavior, and well-being of micromobility users through interdependent traffic situations within a simulated urban environment. Our mixed design study involved two types of social interactions as between-subject conditions of prosocial and asocial interaction, and three categories of time constraint as within-subject conditions: relaxed, neutral, and pressed. The findings reveal that receiving prosocial and asocial behaviors can affect the state of well-being and trial performance in a mobility environment. Shashank Mehrotra, Kumar Akash, Teruhisa Misu, John D. Lee |
AutomotiveUI | 4 |
| 2024 | Can we enhance prosocial behavior? Using post-ride feedback to improve micromobility interactionsabstractMicromobility devices, such as e-scooters and delivery robots, hold promise for eco-friendly and cost-effective alternatives for future urban transportation. However, their lack of societal acceptance remains a challenge. Therefore, we must consider ways to promote prosocial behavior in micromobility interactions. We investigate how post-ride feedback can encourage the prosocial behavior of e-scooter riders while interacting with sidewalk users, including pedestrians and delivery robots. Using a web-based platform, we measure the prosocial behavior of e-scooter riders. Results found that post-ride feedback can successfully promote prosocial behavior, and objective measures indicated better gap behavior, lower speeds at interaction, and longer stopping time around other sidewalk actors. The findings of this study demonstrate the efficacy of post-ride feedback and provide a step toward designing methodologies to improve the prosocial behavior of mobility users. Sidney T. Scott-Sharoni, Shashank Mehrotra, Kevin Salubre, Miao Song 0007, Teruhisa Misu, Kumar Akash |
AutomotiveUI | 5 |
| 2024 | Prosocial Acts Towards AI Shaped By Reciprocation And Awareness
Kumar Akash, Shashank Mehrotra, Teruhisa Misu, Mark Steyvers |
CogSci | 4 |
| 2024 | Beyond Empirical Windowing: An Attention-Based Approach for Trust Prediction In Autonomous VehiclesabstractHumans’ internal states play a key role in human-machine interaction, leading to the rise of human state estimation as a prominent field. Compared to swift state changes such as surprise and irritation, modeling gradual states like trust and satisfaction are further challenged by label sparsity: long time-series signals are usually associated with a single label, making it difficult to identify the critical span of state shifts. Windowing has been one widely-used technique to enable localized analysis of long time-series data. However, the performance of downstream models can be sensitive to the window size, and determining the optimal window size demands domain expertise and extensive search. To address this challenge, we propose a Selective Windowing Attention Network (SWAN), which employs window prompts and masked attention transformation to enable the selection of attended intervals with flexible lengths. We evaluate SWAN on the task of trust prediction on a new multimodal driving simulation dataset. Experiments show that SWAN significantly outperforms an existing empirical window selection baseline and neural network baselines including CNN-LSTM and Transformer. Furthermore, it shows robustness across a wide span of windowing ranges, compared to the traditional windowing approach. Minxue Niu, Zhaobo K. Zheng, Kumar Akash, Teruhisa Misu |
ICASSP | 4 |
| 2024 | Optimal Driver Warning Generation in Dynamic Driving EnvironmentabstractThe driver warning system that alerts the human driver about potential risks during driving is a key feature of an advanced driver assistance system. Existing driver warning technologies, mainly the forward collision warning and unsafe lane change warning, can reduce the risk of collision caused by human errors. However, the current design methods have several major limitations. Firstly, the warnings are mainly generated in a one-shot manner without modeling the ego driver’s reactions and surrounding objects, which reduces the flexibility and generality of the system over different scenarios. Additionally, the triggering conditions of warning are mostly rule-based threshold-checking given the current state, which lacks the prediction of the potential risk in a sufficiently long future horizon. In this work, we study the problem of optimally generating driver warnings by considering the interactions among the generated warning, the driver behavior, and the states of ego and surrounding vehicles on a long horizon. The warning generation problem is formulated as a partially observed Markov decision process (POMDP). An optimal warning generation framework is proposed as a solution to the proposed POMDP. The simulation experiments demonstrate the superiority of the proposed solution to the existing warning generation methods. Chenran Li, Aolin Xu 0002, Enna Sachdeva, Teruhisa Misu, Behzad Dariush |
ICRA | 4 |
| 2024 | How is the Pilot Doing: VTOL Pilot Workload Estimation by Multimodal Machine Learning on Psycho-physiological SignalsabstractVertical take-off and landing (VTOL) aircraft do not require a prolonged runway, thus allowing them to land almost anywhere. In recent years, their flexibility has made them popular in development, research, and operation. When compared to traditional fixed-wing aircraft and rotorcraft, VTOLs bring unique challenges as they combine many maneuvers from both types of aircraft. Pilot workload is a critical factor for safe and efficient operation of VTOLs. In this work, we conduct a user study to collect multimodal data from 28 pilots while they perform a variety of VTOL flight tasks. We analyze and interpolate behavioral patterns related to their performance and perceived workload. Finally, we build machine learning models to estimate their workload from the collected data. Our results are promising, suggesting that quantitative and accurate VTOL pilot workload monitoring is viable. Such assistive tools would help the research field understand VTOL operations and serve as a stepping stone for the industry to ensure VTOL safe operations and further remote operations. Jong Hoon Park, Lawrence Chen 0004, Ian Higgins, Zhaobo Zheng, Shashank Mehrotra, Kevin Salubre, Mohammadreza Mousaei, Steven Willits, Blaine Levedahl, Timothy Buker, Eliot Xing, Teruhisa Misu, Sebastian A. Scherer, Jean Oh |
RO-MAN | 12 |
| 2023 | Learn-able Evolution Convolutional Siamese Neural Network for Adaptive Driving Style Preference PredictionabstractWe propose a framework for detecting user driving style preference with multimodal signals, to adapt autonomous vehicle driving style to drivers’ preferences in an automatic manner. Mismatch between the automated vehicle driving style and the driver’s preference can lead to more frequent takeovers or even disabling the automation features. We collected multi-modal data from 36 human participants on a driving simulator, including eye gaze, steering grip force, driving maneuvers, brake and throttle pedal inputs as well as foot distance from pedals, pupil diameter, galvanic skin response, heart rate, and situational drive context. Based on the data, we constructed a data-driven framework using convolutional Siamese neural networks (CSNNs) to identify preferred driving styles. The model performance has significant improvement compared to that in the existing literature. In addition, we demonstrated that the proposed framework can improve model performance without network training process using data from target users. This result validates the potential of online model adaption with continued driver-system interaction. We also perform an ablation study on sensing modalities and present the importance of each data channel. Fatemeh Koochaki, Zhaobo K. Zheng, Kumar Akash, Teruhisa Misu |
IV | 4 |
| 2023 | Example-Based Query To Identify Causes of Driving Anomaly with Few Labeled SamplesabstractDriving anomaly detection is important for advanced driver assistance systems (ADAS) to increase driving safety and avoid traffic accidents. However, driving anomaly detection faces many challenges such as numerous and uncertain abnormal patterns observed on the road, sparsity of real anomaly cases documented with accurate labels, and rigid existing systems that rely on manually set thresholds and rules. Previous studies have proposed unsupervised methods for driving anomaly detection in the driver’s behaviors or the road condition by identifying deviations from normal driving conditions. A challenge with unsupervised models is the lack of interpretability, where the cause of the anomaly is not always clear. We address this problem with an example-based query method that combines unsupervised anomaly detection methods with the multi-label k-nearest neighbors (ML-KNN) algorithm to interpret the detected driving anomalies by identifying their possible causes (e.g., surrounding objects or driver’s errors). Our approach relies on a few manually labeled driving segments that are efficiently used as anchors to retrieve the causes of driving anomalies in a given driving segment. These anchors are projected into the embedding created by unsupervised driving anomaly detection systems. The experimental results show that this method can effectively identify the causes of driving anomalies, even for abnormal driving segments triggered by multiple causes. The evaluation shows the flexibility of our proposed solution, where we successfully implement the ML-KNN approach with three alternative feature representations. Yuning Qiu, Teruhisa Misu, Carlos Busso |
IV | 2 |
| 2023 | The Impact of Environmental Features on Drivers' Situation Awareness Using Real-World Driving ScenariosabstractAdvanced driver assistance systems (ADAS) need to account for the driver’s awareness of the environment to be effectively used. This study examines the impact of environmental features (eg, visual complexity, object density, roadway type, lighting) on drivers’ situation awareness (SA). This is achieved using a controlled study with 40 participants. Using a split-plot design, the participants were shown 30 out of 75 real-world driving scenarios displayed in a driving simulator environment. Participants’ responses to Situational Awareness Global Assessment Technique (SAGAT) queries on the type and coordinates of objects in the scene were used to calculate SA scores. A hurdle model was developed to estimate participants’ SA scores. The key findings highlight visual complexity as a significant predictor of SA scores. This predictor was easy to compute and able to capture the complexity of objects that impact road safety as well as the visual clutter in the background. The model showed that drivers were able to identify at least one object of interest in complex environments with high visual complexity and with many objects. A higher proportion of vulnerable road users was associated with a greater likelihood of a non-zero SA score, but the SA score was lower compared to environments with higher proportions of cars. The findings of this study provide insights into the environmental factors to be considered for SA predictive models. Yilun Xing, Sami Park, Kumar Akash, Teruhisa Misu, Linda Ng Boyle |
Int. J. Hum. Comput. Interact. | 4 |
| 2022 | Learning Temporally and Semantically Consistent Unpaired Video-to-Video Translation through Pseudo-Supervision from Synthetic Optical FlowabstractUnpaired video-to-video translation aims to translate videos between a source and a target domain without the need of paired training data, making it more feasible for real applications. Unfortunately, the translated videos generally suffer from temporal and semantic inconsistency. To address this, many existing works adopt spatiotemporal consistency constraints incorporating temporal information based on motion estimation. However, the inaccuracies in the estimation of motion deteriorate the quality of the guidance towards spatiotemporal consistency, which leads to unstable translation. In this work, we propose a novel paradigm that regularizes the spatiotemporal consistency by synthesizing motions in input videos with the generated optical flow instead of estimating them. Therefore, the synthetic motion can be applied in the regularization paradigm to keep motions consistent across domains without the risk of errors in motion estimation. Thereafter, we utilize our unsupervised recycle and unsupervised spatial loss, guided by the pseudo-supervision provided by the synthetic optical flow, to accurately enforce spatiotemporal consistency in both domains. Experiments show that our method is versatile in various scenarios and achieves state-of-the-art performance in generating temporally and semantically consistent videos. Code is available at: https://github.com/wangkaihong/Unsup_Recycle_GAN/. Kaihong Wang, Kumar Akash, Teruhisa Misu |
AAAI | 3 |
| 2022 | Toward Adaptive Driving Styles for Automated Driving with Users' Trust and PreferencesabstractAs autonomous vehicles (AVs) become ubiquitous, users' trust will be critical for the successful adoption of such systems. Prior works have shown that the driving styles of AVs can impact how users trust and rely on such systems. However, users' preferred driving style may vary with changes in trust or road conditions, experience, and personal driving preferences. We explore methods to adapt the driving style of an AV to match the preferred driving style of users to improve their trust in the vehicle. We conducted a pilot study ($n=16$) on a simulated urban environment, where the users experience various static and adaptive driving styles for different pedestrian and traffic-related scenarios. Our results indicate that users best trust AVs that closely match their preferences ($p< 0.05$). We believe that exploring the effects of AV driving style on users' trust and workload will provide necessary steps towards developing human-aware automated systems. Manisha Natarajan, Kumar Akash, Teruhisa Misu |
HRI | 3 |
| 2022 | Incorporating Gaze Behavior Using Joint Embedding With Scene Context for Driver Takeover DetectionabstractDespite the recent advancement in driver assistance systems, most existing solutions and partial automation systems such as SAE Level 2 driving automation systems assume that the driver is in the loop; the human driver must continuously monitor the driving environment. Frequent transition of maneuver control is expected between the driver and the car while using such automation in difficult traffic conditions. In this work, we aim to predict driver takeover timing in order for the system to prepare transition from automation to driver control. While previous studies indicated that eye gaze is an important cue to predict driver takeover, we hypothesize that traffic condition as well as the reliability of the driving automation also have a strong impact. Therefore, we propose an algorithm that jointly consider the driver’s gaze information and contextual driving environment, which is complemented with the vehicle operational and driver physiological signals. Specifically, we consider joint embedding of traffic scene information and gaze behavior using 3DConvolutional Neural Network (3D-CNN). We demonstrate that our algorithm is successfully able to predict driver takeover intent, using user study data from 28 participants collected in simulated driving environments. Yuning Qiu, Carlos Busso, Teruhisa Misu, Kumar Akash |
ICASSP | 3 |
| 2022 | Identification of Adaptive Driving Style Preference through Implicit Inputs in SAE L2 VehiclesabstractA key factor to optimal acceptance and comfort of automated vehicle features is the driving style. Mismatches between the automated and the driver preferred driving styles can make users take over more frequently or even disable the automation features. This work proposes identification of user driving style preference with multimodal signals, so the vehicle could match user preference in a continuous and automatic way. We conducted a driving simulator study with 36 participants and collected extensive multimodal data including behavioral, physiological, and situational data. This includes eye gaze, steering grip force, driving maneuvers, brake and throttle pedal inputs as well as foot distance from pedals, pupil diameter, galvanic skin response, heart rate, and situational drive context. Then, we built machine learning models to identify preferred driving styles, and confirmed that all modalities are important for the identification of user preference. This work paves the road for implicit adaptive driving styles on automated vehicles. Zhaobo Zheng, Kumar Akash, Teruhisa Misu, Vidya Krishnamoorthy, Yuni Lee, Gaojian Huang |
ICMI | 3 |
| 2022 | Driving Anomaly Detection Using Contrastive Multiview Coding to Interpret Cause of AnomalyabstractModern advanced driver assistant systems (ADAS) rely on various types of sensors to monitor the vehicle status, driver's behaviors and road condition. The multimodal systems in the vehicle include sensors, such as accelerometers, pressure sensors, cameras, lidar and radars. When looking at a given scene with multiple modalities, there should be congruent in-formation among different modalities. Exploring the congruent information across modalities can lead to appealing solutions to create robust multimodal representations. This work proposes an unsupervised approach based on contrastive multiview coding (CMC) to capture the correlations in representations extracted from different modalities, learning a more discriminative rep-resentation space for unsupervised anomaly driving detection. We use CMC to train our model to extract view-invariant factors by maximizing the mutual information between mul-tiple representations from a given view, and increasing the distance of views from unrelated segments. We consider the vehicle driving data, driver's physiological data, and external environment data consisting of distances to nearby pedestrians, bicycles, and vehicles. The experimental results on the driving anomaly dataset (DAD) indicate that the CMC representation is effective for driving anomaly detection. The approach is efficient, scalable and interpretable, where the distances in the contrastive embedding for each view can be used to understand potential causes of the detected anomalies. Yuning Qiu, Teruhisa Misu, Carlos Busso |
IROS | 2 |
| 2022 | Domain Knowledge Driven Pseudo Labels for Interpretable Goal-Conditioned Interactive Trajectory PredictionabstractMotion forecasting in highly interactive scenarios is a challenging problem in autonomous driving. In such scenarios, we need to accurately predict the joint behavior of interacting agents to ensure the safe and efficient navigation of autonomous vehicles. Recently, goal-conditioned methods have gained increasing attention due to their advantage in performance and their ability to capture the multimodality in trajec-tory distribution. In this work, we study the joint trajectory prediction problem with the goal-conditioned framework. In particular, we introduce a conditional-variational-autoencoder-based (CVAE) model to explicitly encode different interaction modes into the latent space. However, we discover that the vanilla model suffers from posterior collapse and cannot induce an informative latent space as desired. To address these issues, we propose a novel approach to avoid KL vanishing and induce an interpretable interactive latent space with pseudo labels. The proposed pseudo labels allow us to incorporate domain knowledge on interaction in a flexible manner. We motivate the proposed method using an illustrative toy example. In addition, we validate our framework on the Waymo Open Motion Dataset with both quantitative and qualitative evaluations. Lingfeng Sun, Chen Tang 0001, Yaru Niu, Enna Sachdeva, Chiho Choi, Teruhisa Misu, Masayoshi Tomizuka |
IROS | 6 |
| 2022 | Effects of Augmented-Reality-Based Assisting Interfaces on Drivers' Object-wise Situational Awareness in Highly Autonomous VehiclesabstractAlthough partially autonomous driving (AD) systems are already available in production vehicles, drivers are still required to maintain a sufficient level of situational awareness (SA) during driving. Previous studies have shown that providing information about the AD’s capability using user interfaces can improve the driver’s SA. However, displaying too much information increases the driver’s workload and can distract or overwhelm the driver. Therefore, to design an efficient user interface (UI), it is necessary to understand its effect under different circumstances. In this paper, we focus on a UI based on augmented reality (AR), which can highlight potential hazards on the road. To understand the effect of highlighting on drivers’ SA for objects with different types and locations under various traffic densities, we conducted an in-person experiment with 20 participants on a driving simulator. Our study results show that the effects of highlighting on drivers’ SA varied by traffic densities, object locations and object types. We believe our study can provide guidance in selecting which object to highlight for the AR-based driver-assistance interface to optimize SA for drivers driving and monitoring partially autonomous vehicles. Xiaofeng Gao 0002, Xingwei Wu, Samson Ho, Teruhisa Misu, Kumar Akash |
IV | 4 |
| 2022 | Toward an Adaptive Situational Awareness Support System for Urban DrivingabstractA lack of sufficient situational awareness is a primary cause of traffic crashes due to human error. Redirecting a driver’s attention to critical objects is essential, but alerting driver about all critical objects can lead to distraction. This paper develops and evaluates an adaptive support system that incorporates drivers’ fixations as a proxy for their situational awareness. We implement an experimental system that detects a driver’s gaze on important objects in the traffic scene and adapts a cueing strategy in an augmented reality-based driver awareness assistance interface. We collect and analyze data from 15 participants and show that our adaptive support system strategy is effective without increasing the drivers’ cognitive workload. Finally, we show that our system can increase ratio of drivers’ fixations on critical objects in their view without significantly increasing dwell time per object. Tong Wu 0012, Enna Sachdeva, Kumar Akash, Xingwei Wu, Teruhisa Misu, Jorge Ortiz 0001 |
IV | 5 |
| 2021 | Improving Driver Situation Awareness Prediction using Human Visual Sensory and Memory MechanismabstractSituation awareness (SA) is generally considered as the perception, understanding, and projection of objects’ properties and positions. We believe if the system can sense drivers’ SA, it can appropriately provide warnings for objects that drivers are not aware of. To investigate drivers’ awareness, in this study, a human-subject experiment of driving simulation was conducted for data collection. While a previous predictive model for drivers’ situation awareness utilized drivers’ gaze movement only, this work utilizes object properties, characteristics of human visual sensory and memory mechanism. As a result, the proposed driver SA prediction model achieves over 70% accuracy and outperforms the baselines. Haibei Zhu, Teruhisa Misu, Sujitha Martin, Xingwei Wu, Kumar Akash |
IROS | 2 |
| 2020 | What Driving Says About You: A Small-Sample Exploratory Study Between Personality and Self-Reported Driving Style Among Young Male DriversabstractUnderstanding how personalities relate to driving styles is crucial for improving Advanced Driver Assistance Systems (ADASs) and driver-vehicle interactions. Focusing on the ”high-risk” population of young male drivers, the objective of this study is to investigate the association between personality traits and driving styles. An online survey study was conducted among 46 males aged 21-30 to gauge their personality traits, self-reported driving style, and driving history. Hierarchical Clustering was proposed to identify driving styles and revealed two subgroups of drivers who either had a ”risky” or ”compliant” driving style. Compared to the compliant group, the risky cluster sped more frequently, was easily distracted and affected by negative emotion, and often behaved recklessly. The logit model results showed that the risky driving style was associated with lower Agreeableness and Conscientiousness, but higher driving exposure. An interaction effect was also detected between age and Extraversion to form a risky driving style. Xingwei Wu, Yuki Gorospe, Teruhisa Misu, Y. Huynh, Nimsi Guerrero |
AutomotiveUI | 3 |
| 2020 | Toward Adaptive Trust Calibration for Level 2 Driving AutomationabstractProperly calibrated human trust is essential for successful interaction between humans and automation. However, while human trust calibration can be improved by increased automation transparency, too much transparency can overwhelm human workload. To address this tradeoff, we present a probabilistic framework using a partially observable Markov decision process (POMDP) for modeling the coupled trust-workload dynamics of human behavior in an action-automation context. We specifically consider hands-off Level 2 driving automation in a city environment involving multiple intersections where the human chooses whether or not to rely on the automation. We consider automation reliability, automation transparency, and scene complexity, along with human reliance and eye-gaze behavior, to model the dynamics of human trust and workload. We demonstrate that our model framework can appropriately vary automation transparency based on real-time human trust and workload belief estimates to achieve trust calibration. Kumar Akash, Neera Jain, Teruhisa Misu |
ICMI | 3 |
| 2020 | Toward Real-Time Estimation of Driver Situation Awareness: An Eye-tracking Approach based on Moving Objects of InterestabstractEye-tracking techniques have the potential for estimating driver awareness of road hazards. However, traditional eye-movement measures based on static areas of interest may not capture the unique characteristics of driver eyeglance behavior and challenge the real-time application of the technology on the road. This article proposes a novel method to operationalize driver eye-movement data analysis based on moving objects of interest. A human-subject experiment conducted in a driving simulator demonstrated the potential of the proposed method. Correlation and regression analyses between indirect (i.e., eye-tracking) and direct measures of driver awareness identified some promising variables that feature both spatial and temporal aspects of driver eye-glance behavior relative to objects of interest. Results also suggest that eye-glance behavior might be a promising but insufficient predictor of driver awareness. This work is a preliminary step toward real-time, on-road estimation of driver awareness of road hazards. The proposed method could be further combined with computer-vision techniques such as object recognition to fully automate eye-movement data processing as well as machine learning approaches to improve the accuracy of driver awareness estimation. Hyungil Kim, Sujitha Martin, Ashish Tawari, Teruhisa Misu, Joseph L. Gabbard |
IV | 4 |
| 2020 | Drivers' Attitudes and Perceptions towards A Driving Automation System with Augmented Reality Human-Machine InterfacesabstractInteraction research has been initially focusing on partially and conditionally automated vehicles. Augmented Reality (AR) may provide a promising way to enhance drivers' experience when using autonomous driving (AD) systems. This study sought to gain insights on drivers' subjective assessment of a simulated driving automation system with AR-based support. A driving simulator study was conducted and participants' rating of the AD system in terms of information imparting, nervousness and trust was collected. Cumulative Link Models (CLMs) were developed to investigate the impacts of AR cues, traffic density and intersection complexity on drivers' attitudes towards the presented AD system. Random effects were incorporated in the CLMs to account for the heterogeneity among participants. Results indicated that AR graphic cues could significantly improve drivers' experience by providing advice for their decision-making and mitigating their anxiety and stress. However, the magnitude of AR's effect was impacted by traffic conditions (i.e. diminished at more complex intersections). The study also revealed a strong correlation between self-rated trust and takeover frequency, suggesting takeover and other driving behavior need to be further examined in future studies. Xingwei Wu, Coleman Merenda, Teruhisa Misu, Kyle Tanous, Chihiro Suga, Joseph L. Gabbard |
IV | 3 |
| 2020 | Boosting Standard Classification Architectures Through a Ranking RegularizerabstractWe employ triplet loss as a feature embedding regularizer to boost classification performance. Standard architectures, like ResNet and Inception, are extended to support both losses with minimal hyper-parameter tuning. This promotes generality while fine-tuning pretrained networks. Triplet loss is a powerful surrogate for recently proposed embedding regularizers. Yet, it is avoided due to large batch-size requirement and high computational cost. Through our experiments, we re-assess these assumptions.During inference, our network supports both classification and embedding tasks without any computational overhead. Quantitative evaluation highlights a steady improvement on five fine-grained recognition datasets. Further evaluation on an imbalanced video dataset achieves significant improvement. Triplet loss brings feature embedding capabilities like nearest neighbor to classification models. Code available at http://bit.ly/2LNYEqL. Ahmed Taha 0001, Yi-Ting Chen 0001, Teruhisa Misu, Abhinav Shrivastava, Larry Davis 0001 |
WACV | 3 |
| 2019 | Grounding Human-To-Vehicle Advice for Self-Driving VehiclesabstractRecent success suggests that deep neural control networks are likely to be a key component of self-driving vehicles. These networks are trained on large datasets to imitate human actions, but they lack semantic understanding of image contents. This makes them brittle and potentially unsafe in situations that do not match training data. Here, we propose to address this issue by augmenting training data with natural language advice from a human. Advice includes guidance about what to do and where to attend. We present the first step toward advice giving, where we train an end-to-end vehicle controller that accepts advice. The controller adapts the way it attends to the scene (visual attention) and the control (steering and speed). Attention mechanisms tie controller behavior to salient objects in the advice. We evaluate our model on a novel advisable driving dataset with manually annotated human-to-vehicle advice called Honda Research Institute-Advice Dataset (HAD). We show that taking advice improves the performance of the end-to-end network, while the network cues on a variety of visual features that are provided by advice. The dataset is available at https://usa.honda-ri.com/HAD. Jinkyu Kim 0001, Teruhisa Misu, Yi-Ting Chen 0001, Ashish Tawari, John F. Canny |
CVPR | 2 |
| 2019 | Driving Anomaly Detection with Conditional Generative Adversarial Network using Physiological and CAN-Bus DataabstractNew developments in advanced driver assistance systems (ADAS) can help drivers deal with risky driving maneuvers, preventing potential hazard scenarios. A key challenge in these systems is to determine when to intervene. While there are situations where the needs for intervention or feedback is clear (e.g., lane departure), it is often difficult to determine scenarios that deviate from normal driving conditions. These scenarios can appear due to errors by the drivers, presence of pedestrian or bicycles, or maneuvers from other vehicles. We formulate this problem as a driving anomaly detection, where the goal is to automatically identify cases that require intervention. Towards addressing this challenging but important goal, we propose a multimodal system that considers (1) physiological signals from the driver, and (2) vehicle information obtained from the controller area network (CAN) bus sensor. The system relies on conditional generative adversarial networks (GAN) where the models are constrained by the signals previously observed. The difference of the scores in the discriminator between the predicted and actual signals is used as a metric for detecting driving anomalies. We collected and annotated a novel dataset for driving anomaly detection tasks, which is used to validate our proposed models. We present the analysis of the results, and perceptual evaluations which demonstrate the discriminative power of this unsupervised approach for detecting driving anomalies. Yuning Qiu, Teruhisa Misu, Carlos Busso |
ICMI | 2 |
| 2019 | Deep Multi-Task Learning for Anomalous Driving Detection Using CAN Bus Scalar Sensor DataabstractCorner cases are the main bottlenecks when applying Artificial Intelligence (AI) systems to safety-critical applications. An AI system should be intelligent enough to detect such situations so that system developers can prepare for subsequent planning. In this paper, we propose semi-supervised anomaly detection considering the imbalance of normal situations: In particular, driving data consists of multiple normal situations (e.g., right turn, going straight), some of which (e.g., U-turn) could be as rare as anomalous ones. Existing machine learning based anomaly detection approaches do not fare sufficiently well when applied to such imbalanced data. In this paper, we present a novel multi-task learning (LSTM autoencoder and predictor) based approach that leverages domain-knowledge (maneuver labels) for anomaly detection in driving data. We evaluate the proposed approach both quantitatively and qualitatively on 150 hours of real-world driving data and show improved performance over baseline/existing approaches. Vidyasagar Sadhu, Teruhisa Misu, Dario Pompili |
IROS | 2 |
| 2019 | Effects of "Real-World" Visual Fidelity on AR Interface Assessment: A Case Study Using AR Head-up Display Graphics in DrivingabstractRecent AR research efforts have explored the use of virtual environments to test augmented reality (AR) user interfaces. However, it is yet to be seen what effects the visual fidelity of such virtual environments may have on AR interface assessment, and specifically to what degree assessment results observed in a virtual world would apply to the real world. Automotive AR head-up (HUD) interfaces provide a meaningful application area to examine this problem, especially given that immersive, 3D-graphics-based driving simulators are established tools to examine in-vehicle interfaces safely before testing in real vehicles. In this work, we present an argument that adequately assessing AR interfaces requires a suite of different measures, and that such measures should be considered when debating the appropriateness of virtual environments for AR interface assessment. We present a case study that examines how an AR interface presented via HUD effects driver performance and behavior in different virtual and real environments. Twelve participants completed the study measuring driver task performance, eye gaze behavior and situational awareness during AR guided navigation in low-and high-fidelity virtual simulation, and an on-road environment. Our results suggest that the visual fidelity of the environmental in which an AR interface is assessed, could impact some measures of effectiveness. Discussion is guided by a proposed initial assessment classification for AR user interfaces that may serve to guide future discussions on AR interface evaluation, as well as the suitability of virtual environments for AR assessment. Coleman Merenda, Chihiro Suga, Joseph L. Gabbard, Teruhisa Misu |
ISMAR | 4 |
| 2019 | Effects of Vehicle Simulation Visual Fidelity on Assessing Driver Performance and BehaviorabstractAutomotive manufactures are rapidly developing more advanced in-vehicle systems that seek to provide a driver with more active safety and information in real-time, in particular human machine interfaces (HMIs) using mixed or augmented reality (AR) graphical elements. However, it is difficult to properly test novel AR interfaces in the same way as traditional HMIs via on-road testing. Instead, simulation could likely offer a safer and more financially viable alternative for testing AR HMIs, inconsistent simulation quality may confound HMI research depending on the visual fidelity of each simulation environment. We investigated how visual fidelity in a virtual environment impacts the quality of resulting driver behavior, visual attention, and situational awareness when using the system. We designed two large-scale immersive virtual environments; a “low” graphic fidelity driving simulation representing most current research simulation testbeds and a “high” graphic fidelity environment created in Unreal Engine that represents state of the art graphical presentation. We conducted a user study with 24 participants who navigated a route in a virtual urban environment via direction of AR graphical cues while also monitoring the road scene for pedestrian hazards, and recorded their driving performance, gaze patterns, and subjective feedback via situational awareness survey (SART). Our results show drivers change both their driving and visual behavior depending upon the visual fidelity presented in the virtual scene. We further demonstrate the value of using multi-tiered analysis techniques to more finely examine driver performance and visual attention. Coleman Merenda, Chihiro Suga, Joseph L. Gabbard, Teruhisa Misu |
IV | 4 |
| 2018 | Toward Driving Scene Understanding: A Dataset for Learning Driver Behavior and Causal ReasoningabstractDriving Scene understanding is a key ingredient for intelligent transportation systems. To achieve systems that can operate in a complex physical and social environment, they need to understand and learn how humans drive and interact with traffic scenes. We present the Honda Research Institute Driving Dataset (HDD), a challenging dataset to enable research on learning driver behavior in real-life environments. The dataset includes 104 hours of real human driving in the San Francisco Bay Area collected using an instrumented vehicle equipped with different sensors. We provide a detailed analysis of HDD with a comparison to other driving datasets. A novel annotation methodology is introduced to enable research on driver behavior understanding from untrimmed data sequences. As the first step, baseline algorithms for driver behavior detection are trained and tested to demonstrate the feasibility of the proposed task. Vasili Ramanishka, Yi-Ting Chen 0001, Teruhisa Misu, Kate Saenko |
CVPR | 3 |
| 2018 | Situated reference resolution using visual saliency and crowdsourcing-based priors for a spoken dialog system within vehicles
Teruhisa Misu |
Comput. Speech Lang. | 1 |
| 2018 | Driver Behavior and Performance with Augmented Reality Pedestrian Collision Warning: An Outdoor User StudyabstractThis article investigates the effects of visual warning presentation methods on human performance in augmented reality (AR) driving. An experimental user study was conducted in a parking lot where participants drove a test vehicle while braking for any cross traffic with assistance from AR visual warnings presented on a monoscopic and volumetric head-up display (HUD). Results showed that monoscopic displays can be as effective as volumetric displays for human performance in AR braking tasks. The experiment also demonstrated the benefits of conformal graphics, which are tightly integrated into the real world, such as their ability to guide drivers' attention and their positive consequences on driver behavior and performance. These findings suggest that conformal graphics presented via monoscopic HUDs can enhance driver performance by leveraging the effectiveness of monocular depth cues. The proposed approaches and methods can be used and further developed by future researchers and practitioners to better understand driver performance in AR as well as inform usability evaluation of future automotive AR applications. Hyungil Kim, Joseph L. Gabbard, Alexandre Miranda Añon, Teruhisa Misu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2018 | Augmented Reality Interface Design Approaches for Goal-directed and Stimulus-driven Driving TasksabstractThe automotive industry is rapidly developing new in-vehicle technologies that can provide drivers with information to aid awareness and promote quicker response times. Particularly, vehicles with augmented reality (AR) graphics delivered via head-up displays (HUDs) are nearing mainstream commercial feasibility and will be widely implemented over the next decade. Though AR graphics have been shown to provide tangible benefits to drivers in scenarios like forward collision warnings and navigation, they also create many new perceptual and sensory issues for drivers. For some time now, designers have focused on increasing the realism and quality of virtual graphics delivered via HUDs, and recently have begun testing more advanced 3D HUD systems that deliver volumetric spatial information to drivers. However, the realization of volumetric graphics adds further complexity to the design and delivery of AR cues, and moreover, parameters in this new design space must be clearly and operationally defined and explored. In this work, we present two user studies that examine how driver performance and visual attention are affected when using fixed and animated AR HUD interface design approaches in driving scenarios that require top-down and bottom-up cognitive processing. Results demonstrate that animated design approaches can produce some driving gains (e.g., in goal-directed navigation tasks) but often come at the cost of response time and distance. Our discussion yields AR HUD design recommendations and challenges some of the existing assumptions of world-fixed conformal graphic approaches to design. Coleman Merenda, Hyungil Kim, Kyle Tanous, Joseph L. Gabbard, Blake Feichtl, Teruhisa Misu, Chihiro Suga |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2017 | Understand driver awareness through brake behavior analysis: Reactive versus intended hard brakeabstractDriving is a highly dynamic activity where drivers' awareness to the traffic environment plays the essential role for successful performance. Easy, smooth driving depends on drivers' awareness to develop situation-specific expectations, where infrequent or unexpected situations are not taken into account. Understanding these unexpected situations provides important insight on driver situation awareness and accident prevention. This study takes advantage of recent development of wearable devices and uses driver physiological signals to identify such unexpected situations during driver hard brake. Based on a naturalistic driving dataset, we define two types of hard brake behavior: reactive and intended hard brake. The reactive hard brake relates to drivers reacting to unexpected situations that usually leads to deviated physiological signals due to stress. The intended hard brake relates to planned maneuver implementation that consists of stable physiological signals. By using the human evaluation, we identified the different situations in which these two types of hard brake occurs. Clear difference is observed between these two groups of situations. Our goal is to identify features that are representative of these two type of road environment, especially the situations where unexpected reactive hard brake happens. Following this direction, we extracted features from Lidar depth scanner to represent the road scene, and applied both lasso regression and logistic regression classifier for feature analysis. The regression model achieves high correlation of 0.77 between the prediction and the ground truth while the classification achieves F-score of 0.76. The selected Lidar features can serve as high level road scene representation that facilitate next generation advanced driver assistance systems (ADAS) to prevent accident in unexpected traffic scenarios. Nanxiang Li, Teruhisa Misu, Fei Tao 0003 |
Intelligent Vehicles Symposium | 2 |
| 2016 | Driving maneuver prediction using car sensor and driver physiological signalsabstractThis study presents the preliminary attempt to investigate the usage of driver physiology signals, including electrocardiography (ECG) and respiration wave signals, to predict driving maneuvers. While most studies on driving maneuver prediction uses direct measurements from vehicle or road scene, we believe the mental state changes from the driver when making plans for maneuver can be reflected from the physiological signals. We extract both time and frequency domain features from the physiological signals, and use them as the features to predict the drivers' future maneuver. We formulate the prediction of driver maneuver as a multi-class classification problem by using the features extracted from signal before the driving maneuvers. The multi classes correspond to various types of driving maneuvers including Start, Stop, Lane Switch and Turn. We use the support vector machine (SVM) as the classifier, and compare the performance of using both physiological and car signals (CAN bus) with the baseline classifier that is trained with only car signal. An improved performance is observed when using the physiological features with 0.04 in F-score on average. This improvement is more obvious as the prediction is made earlier. Nanxiang Li, Teruhisa Misu, Ashish Tawari, Alexandre Miranda Añon, Chihiro Suga, Kikuo Fujimura |
ICMI | 2 |
| 2016 | Look at Me: Augmented Reality Pedestrian Warning System Using an In-Vehicle Volumetric Head Up DisplayabstractCurrent pedestrian collision warning systems use either auditory alarms or visual symbols to inform drivers. These traditional approaches cannot tell the driver where the detected pedestrians are located, which is critical for the driver to respond appropriately. To address this problem, we introduce a new driver interface taking advantage of a volumetric head-up display (HUD). In our experimental user study, sixteen participants drove a test vehicle in a parking lot while braking for crossing pedestrians using different interface designs on the HUD. Our results showed that spatial information provided by conformal graphics on the HUD resulted in not only better driver performance but also smoother braking behavior as compared to the baseline. Hyungil Kim, Alexandre Miranda Añon, Teruhisa Misu, Nanxiang Li, Ashish Tawari, Kikuo Fujimura |
IUI | 3 |
| 2015 | Visual Saliency and Crowdsourcing-based Priors for an In-car Situated Dialog SystemabstractThis paper addresses issues in situated language understanding in a moving car. We propose a reference resolution method to identify user queries about specific target objects in their surroundings. We investigate methods of predicting which target object is likely to be queried given a visual scene and what kind of linguistic cues users naturally provide to describe a given target object in a situated environment. We propose methods to incorporate the visual saliency of the visual scene as a prior. Crowdsourced statistics of how people describe an object are also used as a prior. We have collected situated utterances from drivers using our research system, which was embedded in a real vehicle. We demonstrate that the proposed algorithms improve target identification rate by 15.1%. Teruhisa Misu |
ICMI | 1 |
| 2015 | Situated language understanding for a spoken dialog system within vehicles
Teruhisa Misu, Antoine Raux, Rakesh Gupta 0001, Ian Lane |
Comput. Speech Lang. | 1 |
| 2014 | Identification of the Driver's Interest Point using a Head Pose Trajectory for Situated Dialog SystemsabstractThis paper addresses issues existing in situated language understanding in a moving car. Particularly, we propose a method for understanding user queries regarding specific target buildings in their surroundings based on the driver's head pose and speech information. To identify a meaningful head pose motion related to the user query that is among spontaneous motions while driving, we construct a model describing the relationship between sequences of a driver's head pose and the relative direction to an interest point using the Gaussian process regression. We also consider time-varying interest point using kernel density estimation. We collected situated queries from subject drivers by using our research system embedded in a real car. The proposed method achieves an improvement in the target identification rate by 14% in the user-independent training condition and 27% in the user-dependent training condition over the method that uses the head motion at the start-of-speech timing. Young-Ho Kim, Teruhisa Misu |
ICMI | 2 |
| 2014 | Non-monologue HMM-based speech synthesis for service robots: A cloud robotics approachabstractRobot utterances generally sound monotonous, unnatural, and unfriendly because their Text-to-Speech (TTS) systems are not optimized for communication but for text-reading. Here we present a non-monologue speech synthesis for robots. We collected a speech corpus in a non-monologue style in which two professional voice talents read scripted dialogues. Hidden Markov models (HMMs) were then trained with the corpus and used for speech synthesis. We conducted experiments in which the proposed method was evaluated by 24 subjects in three scenarios: text-reading, dialogue, and domestic service robot (DSR) scenarios. In the DSR scenario, we used a physical robot and compared our proposed method with a baseline method using the standard Mean Opinion Score (MOS) criterion. Our experimental results showed that our proposed method's performance was (1) at the same level as the baseline method in the text-reading scenario and (2) exceeded it in the DSR scenario. We deployed our proposed system as a cloud-based speech synthesis service so that it can be used without any cost. Komei Sugiura, Yoshinori Shiga, Hisashi Kawai, Teruhisa Misu, Chiori Hori |
ICRA | 4 |
| 2014 | Crowdsourcing for situated dialog systems in a moving car
Teruhisa Misu |
INTERSPEECH | 1 |
| 2014 | Situated Language Understanding at 25 Miles per HourabstractIn this paper, we address issues in situ-ated language understanding in a rapidly changing environment – a moving car. Specifically, we propose methods for un-derstanding user queries about specific tar-get buildings in their surroundings. Unlike previous studies on physically situated in-teractions such as interaction with mobile robots, the task is very sensitive to tim-ing because the spatial relation between the car and the target is changing while the user is speaking. We collected situated utterances from drivers using our research system, Townsurfer, which is embedded in a real vehicle. Based on this data, we analyze the timing of user queries, spa-tial relationships between the car and tar-gets, head pose of the user, and linguis-tic cues. Optimized on the data, our al-gorithms improved the target identification rate by 24.1 % absolute. 1 Teruhisa Misu, Antoine Raux, Rakesh Gupta 0001, Ian Lane |
SIGDIAL Conference | 1 |
| 2013 | WFST-Based Spoken Dialogue System on Smartphones - Its Development and Implementation for Field UseabstractWe proposed the WFSTDM which is an expandable and adaptable dialogue management platform. The WFSTDM combines various WFSTs and enables us to develop new dialogue management WFSTs necessary for rapid prototyping of spoken dialogue systems. In this paper, we illustrate the outline of the WFSTDM and introduce the WFSTDM builder, a network-based spoken dialogue system development tool. In addition, we go into details about the spoken dialogue system AssisTra for iPhone that we developed as an example to show the implementation of the WFSTDM on smartphones. We also discuss about spoken dialogue systems on smartphones as tools for collecting field data. Etsuo Mizukami, Teruhisa Misu, Chiori Hori |
MDM (2) | 2 |
| 2012 | A bootstrapping approach for SLU portability to a new language by inducting unannotated user queriesabstractThis paper proposes a bootstrapping method of constructing a new spoken language understanding (SLU) system in a target language by utilizing statistical machine translation given an SLU module in some source language. The main challenge in this work is to induct unannotated automatic speech recognition results of user queries in the source language collected through a spoken dialog system, which is under public test. In order to select candidate expressions from among erroneous translation results stemming from problems with speech recognition and machine translation, we use back-translation results to check whether the translation result maintains the semantic meaning of the original sentence. We demonstrate that the proposed scheme can effectively prefer suitable sentences for inclusion in the training data as well as help improve the SLU module for the target language. Teruhisa Misu, Etsuo Mizukami, Hideki Kashioka, Satoshi Nakamura 0001, Haizhou Li 0001 |
ICASSP | 1 |
| 2012 | Reinforcement Learning of Question-Answering Dialogue Policies for Virtual Museum Guides
Teruhisa Misu, Kallirroi Georgila, Anton Leuski, David R. Traum |
SIGDIAL Conference | 1 |
| 2012 | Simultaneous feature selection and parameter optimization for training of dialog policy by reinforcement learningabstractThis paper addresses the problem of feature selection in the reinforcement learning (RL) of the dialog policies of spoken dialog systems. A statistical dialog manager selects the system actions the system should take based on the features derived from the current dialog state and/or the system's belief state. When defining the features used by the system for training the dialog policy, however, finding a set of actually effective features from potentially useful ones is not obvious. In addition, the selection should be done simultaneously with the optimization of the dialog policy. In this paper, we propose an incremental feature selection method for the optimization of a dialog policy by RL, in which improvement of the dialog policy and the feature selection are conducted simultaneously. Experiments in dialog policy optimization by RL with a user simulator demonstrated the following: 1) that the proposed method can find a better dialog policy with fewer policy iterations and 2) the learning speed is comparable with the case where feature selection is conducted in advance. Teruhisa Misu, Hideki Kashioka |
SLT | 1 |
| 2011 | Similarity Based Language Model Construction for Voice Activated Open-Domain Question Answering
István Varga, Kiyonori Ohtake, Kentaro Torisawa, Stijn De Saeger, Teruhisa Misu, Shigeki Matsuda, Jun'ichi Kazama |
IJCNLP | 5 |
| 2011 | User Study of Spoken Decision Support SystemabstractThis paper presents the results of the user evaluation of spo- ken decision support dialogue systems, which help users select from a set of alternatives. Thus far, we have modeled this deci- sion support dialogue as a partially observable Markov decision process (POMDP), and optimized its dialogue strategy to maxi- mize the value of the user’s decision. In this paper, we present a comparative evaluation of the optimized dialogue strategy with several baseline strategies, and demonstrate that the optimized dialogue strategy that was effective in user simulation experi- ments works well in an evaluation by real users. Teruhisa Misu, Kiyonori Ohtake, Chiori Hori, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |
| 2011 | Toward Construction of Spoken Dialogue System that Evokes Users' Spontaneous Backchannels
Teruhisa Misu, Etsuo Mizukami, Yoshinori Shiga, Shinichi Kawamoto, Hisashi Kawai, Satoshi Nakamura 0001 |
SIGDIAL Conference | 1 |
| 2010 | Dialogue Acts Annotation for NICT Kyoto Tour Dialogue Corpus to Construct Statistical Dialogue Systems
Kiyonori Ohtake, Teruhisa Misu, Chiori Hori, Hideki Kashioka, Satoshi Nakamura 0001 |
LREC | 2 |
| 2010 | Modeling Spoken Decision Making Dialogue and Optimization of its Dialogue Strategy
Teruhisa Misu, Komei Sugiura, Kiyonori Ohtake, Chiori Hori, Hideki Kashioka, Hisashi Kawai, Satoshi Nakamura 0001 |
SIGDIAL Conference | 1 |
| 2010 | Dialogue strategy optimization to assist user's decision for spoken consulting dialogue systemsabstractThis paper addresses a user model and dialogue state definition in spoken consulting dialogue systems that help users in making decision. When selecting from a set of alternatives, users have various decision criteria for making decision. Users often do not have a definite goal or criteria for selection, and thus they may find not only what kind of information the system can provide but their own preference or factors that they should emphasize. In this paper, we model such consulting dialogue as partially observable Markov decision process (POMDP). We then present an optimization of dialogue strategy to help users make better decisions. Teruhisa Misu, Komei Sugiura, Kiyonori Ohtake, Chiori Hori, Hideki Kashioka, Hisashi Kawai, Satoshi Nakamura 0001 |
SLT | 1 |
| 2010 | Bayes risk-based dialogue management for document retrieval system with speech interface
Teruhisa Misu, Tatsuya Kawahara |
Speech Commun. | 1 |
| 2009 | Weighted finite state transducer based statistical dialog managementabstractWe proposed a dialog system using a weighted finite-state transducer (WFST) in which user concept and system action tags are input and output of the transducer, respectively. The WFST-based platform for dialog management enables us to combine various statistical models for dialog management (DM), user input understanding and system action generation, and then search the best system action in response to user inputs among multiple hypotheses. To test the potential of the WFST-based DM platform using statistical models, we constructed a dialog system using a human-to-human spoken dialog corpus for hotel reservation, which is annotated with Interchange Format (IF). A scenario WFST and a spoken language understanding (SLU) WFST were obtained from the corpus and then composed together and optimized. We evaluated the detection accuracy of the system next action tags using Mean Reciprocal Ranking (MRR). Finally, we constructed a full WFST-based dialog system by composing SLU, scenario and sentence generation (SG) WFSTs. Humans read the system responses in natural language and judged the quality of the responses. We confirmed that the WFST-based DM platform was capable of handling various spoken language and scenarios when the user concept and system action tags are consistent and distinguishable. Chiori Hori, Kiyonori Ohtake, Teruhisa Misu, Hideki Kashioka, Satoshi Nakamura 0001 |
ASRU | 3 |
| 2009 | Statistical dialog management applied to WFST-based dialog systemsabstractWe have proposed an expandable dialog scenario description and platform to manage dialog systems using a weighted finite-state transducer (WFST) in which user concept and system action tags are input and output of the transducer, respectively. In this paper, we apply this framework to statistical dialog management in which a dialog strategy is acquired from a corpus of human-to-human conversation for hotel reservation. A scenario WFST for dialog management was automatically created from an N-gram model of a tag sequence that was annotated in the corpus with Interchange Format (IF). Additionally, a word-to-concept WFST for spoken language understanding (SLU) was obtained from the same corpus. The acquired scenario WFST and SLU WFST were composed together and then optimized. We evaluated the proposed WFST-based statistic dialog management in terms of correctness to detect the next system actions and have confirmed the automatically acquired dialog scenario from a corpus can manage dialog reasonably on the WFST-based dialog management platform. Chiori Hori, Kiyonori Ohtake, Teruhisa Misu, Hideki Kashioka, Satoshi Nakamura 0001 |
ICASSP | 3 |
| 2009 | Recent advances in WFST-based dialog system
Chiori Hori, Kiyonori Ohtake, Teruhisa Misu, Hideki Kashioka, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2009 | Annotating communicative function and semantic content in dialogue act for construction of consulting dialogue systemsabstractOur goal in this study is to train a dialogue manager that can handle consulting dialogues through spontaneous interactions from a tagged dialogue corpus. We have collected 130 hours of consulting dialogues in sightseeing guidance domain. This paper provides our taxonomy of dialogue act (DA) annotation that can describe two aspects of utterances. One is a communicative function (speech act), and the other is a semantic content of an utterance. We provide an overview of the Kyoto tour guide dialogue corpus and a preliminary analysis using the dialogue act tags. Teruhisa Misu, Kiyonori Ohtake, Chiori Hori, Hideki Kashioka, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |
| 2008 | Dialog management using weighted finite-state transducers
Chiori Hori, Kiyonori Ohtake, Teruhisa Misu, Hideki Kashioka, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2008 | Detection of feeling through back-channels in spoken dialogue
Tatsuya Kawahara, Masayoshi Toyokura, Teruhisa Misu, Chiori Hori |
INTERSPEECH | 3 |
| 2007 | Speech-Based Interactive Information Guidance System using Question-Answering TechniqueabstractThis paper addresses an interactive framework for information navigation based on document knowledge base. In conventional audio guidance systems, such as those deployed in museums, the information flow is one-way and the content is fixed. In order to make an interactive guidance system, we propose the application of question-answering (QA) techniques. Since users tend to use anaphoric expressions in successive questions, we investigate appropriate handling of contextual information based on topic detection, together with the effect of using N-best information in ASR output. Moreover, we apply the QA technique to generation of system-initiative information recommendation. A navigation system on Kyoto city information was implemented. Effectiveness of the proposed techniques was confirmed through a field trial by a number of real novice users. Teruhisa Misu, Tatsuya Kawahara |
ICASSP (4) | 1 |
| 2007 | An Interactive Framework for Document Retrieval and Presentation with Question-Answering Function in Restricted Domain
Teruhisa Misu, Tatsuya Kawahara |
IEA/AIE | 1 |
| 2007 | Bayes risk-based optimization of dialogue management for document retrieval system with speech interfaceabstractAbstract We propose an efficient dialogue management for an informa-tionnavigationsystembasedonadocument knowledge base. Itis expected that incorporation of appropriate N-best candidatesofASRandcontextualinformationwillimprovethesystemper-formance. The system also has several choices in generatingresponses or confirmations. In this paper, this selection is opti-mizedasminimizationofBayesriskbasedonrewardforcorrectinformation presentation and penalty for redundant turns. Wehave evaluated this strategy with our spoken dialogue system“Dialogue Navigator for Kyoto City”, which also has question-answering capability. Effectiveness of the proposed frameworkwas confirmed in the success rate of retrieval and the averagenumber of turns for information access. Index Terms : spoken dialogue system, dialogue management,Bayes risk 1. Introduction The target of spoken dialogue systems is being extended fromsimple databases such as flight information to general docu-ments including manuals and newspaper articles[1]. In suchsystems, the automatic speech recognition (ASR) result of theuser utterance is matched against a set of target documents us-ing the vector space model, and documents with high matchingscores are presented to the user. These kinds of applicationsare expected to be useful especially when retrieving informa-tion with keyboardless devices such as a PDA, a tablet PC anda car navigation system. We have developed “Dialogue Navi-gator for Kyoto City”, which can make interactive guidance byincorporating the question-answering (QA) function as well asgeneral document retrieval.In these systems, making confirmation is needed to elim-inate misunderstandings caused by ASR errors, but so manyredundant confirmations are bothersome to users. There havebeenmanystudiesthatdealwithefficientdialoguemanagementto make confirmation[2, 3, 4]. However, most of them are de-signed for relational database (RDB) query tasks, which havea definite set of keywords, and they are not directly applica-ble to document retrieval tasks, in which every word is used inmatching. In document retrieval tasks, therefore, it is more rea-sonable to make confirmation considering the confidence scoreof retrieval.In addition, our system classifies user utterances to queriesand questions, and generates appropriate responses for respec-tive inputs. Unlike conventional QA tasks, such as TREC QATrack[5], it is not obvious whether the utterance is a query orU1: Please explain the Golden Pavilion.S1: The Golden Pavilion is one of the buildings in theRokuon-ji in Kyoto, and is the main attraction of thetemple sites. The entire pavilion except the basementfloor is covered with pure gold leaf.U2: When was it built?S2: The Golden Pavilion was originally built in 1397to serve as a retirement villa for Shogun AshikagaYoshimitsu.U3: Then, please tell me its history.......Figure 1: Example dialogue of Dialogue Navigator for KyotoCitya question. In addition, an exact answer for a question doesnot necessarily exist in the document sets. Therefore, it is notalways optimal to respond the question with its answer alone.In this paper, we address the extension of conventional op-timization methods of dialogue management, to be applicableto general document retrieval tasks with QA function. Specif-ically, we propose a dialogue management that optimizes thechoicesinresponsegenerationbyminimizingBayesrisk,basedon reward for correct information presentation and penalty forredundant turns, which are defined by the score of documentretrieval and answer extraction. Teruhisa Misu, Tatsuya Kawahara |
INTERSPEECH | 1 |
| 2006 | A bootstrapping approach for developing language model of new spoken dialogue systems by selecting web textsabstractThis paper proposes a bootstrapping method of constructing statistical language models for new spoken dialogue systems by collecting and selecting sentences from the World Wide Web (WWW). To make effective search queries that cover the target domain in full detail, we exploit the document set described about the target domain as seeding data. An important issue is how to filter the retrieved Web pages, since all of the retrieved Web texts are not necessarily suitable as training data. We induct an existing dialogue corpus of different domain to prefer the texts of spoken style. The proposed method was evaluated on two different tasks of software support and sightseeing guidance, and significant reduction of the word error rate was achieved. We show that it is vital to incorporate the dialogue corpus, though not relevant to the target domain, in the text selection phase. Index Terms: speech recognition, language model, spoken dialogue system, web text selection. Teruhisa Misu, Tatsuya Kawahara |
INTERSPEECH | 1 |
| 2006 | Dialogue strategy to clarify user's queries for document retrieval system with speech interface
Teruhisa Misu, Tatsuya Kawahara |
Speech Commun. | 1 |
| 2005 | Dialogue strategy to clarify user's queries for document retrieval system with speech interfaceabstractAbstract This paper proposes a dialogue strategy for clarifying and constraining queries to document retrieval systems with speech input interfaces. It is indispensable for spoken dialogue systems to interpret user’s intention robustly in the presence of speech recognition errors and extraneous expressions characteristic of spontaneous speech. In speech input, moreover, users’ queries tend to be vague, and they may need to be clarified through dialogue in order to extract sufficient information to get meaningful retrieval results. In conventional database query tasks, it is easy to cope with these problems by extracting and confirming keywords based on semantic slots. However, it is not straightforward to apply such a methodology to general document retrieval tasks. In this paper, we first introduce two statistical measures for identifying critical portions to be confirmed. The relevance score (RS) represents the matching degree with the document set. The significance score (SS) detects portions that affect retrieval results. With these measures, the system can generate confirmations to handle speech recognition errors, prior to and after the retrieval, respectively. Then, we propose a dialogue strategy for generating clarifications to narrow down the retrieved items, especially when many documents are matched because of a vague input query. The optimal clarification question is dynamically selected based on information gain (IG) – the reduction in the number of matched items. A set of possible clarification questions is prepared using various knowledge sources. As a bottom-up knowledge source, we extract a list of words that can take a number of objects and potentially causes ambiguity, using a dependency structure analysis of the document texts. This is complemented by top-down knowledge sources of metadata and hand-crafted questions. Our dialogue strategy is implemented and evaluated against a software support knowledge base of 40 K entries. We demonstrate that our strategy significantly improves the success rate of retrieval. Teruhisa Misu, Tatsuya Kawahara |
INTERSPEECH | 1 |
| 2005 | Minimum Bayes-risk decoding considering word significance for information retrieval systemabstractThe paper addresses a new evaluation measure of automatic speech recognition (ASR) and a decoding strategy oriented for speech-based information retrieval (IR). Although word error rate (WER), which treats all words in a uniform manner, has been widely used as an evaluation measure of ASR, significance of words are different in speech understanding or IR. In this paper, we define a new ASR evaluation measure, namely, weighted word error rate (WWER) that gives a weight on errors from a viewpoint of IR. Then, we formulate a decoding method to minimize WWER based on Minimum BayesRisk (MBR) framework, and show that the decoding method improves WWER and IR accuracy. Hiroaki Nanjo, Teruhisa Misu, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2004 | Efficient Confirmation Strategy for Large-scale Text Retrieval Systems with Spoken Dialogue Interface
Kazunori Komatani, Teruhisa Misu, Tatsuya Kawahara, Hiroshi G. Okuno |
COLING | 2 |
| 2004 | Confirmation strategy for document retrieval systems with spoken dialog interfaceabstractAdequate confirmation is indispensable in spoken dialog systems to eliminate misunderstandings caused by speech recognition errors. Spoken language also inherently includes redundant expressions such as disfluency and out-of-domain phrases, which do not contribute to task achievement. It is easy to define a set of keywords to be confirmed for conventional database query tasks, but not straightforward in general document retrieval tasks. In this paper, we propose two statistical measures for identifying portions to be confirmed. A relevance score (RS) represents matching degree with the document set. A significance score (SS) detects portions that consequently affect the retrieval results. With these measures, the system can generate confirmation prior to and posterior to the retrieval, respectively. The strategy is implemented and evaluated with retrieval from software support knowledge base of 40K entries. It is shown that the proposed strategy using the two measures is more efficient than using the conventional confidence measure. Teruhisa Misu, Tatsuya Kawahara, Kazunori Komatani |
INTERSPEECH | 1 |