Na Du

dblp:206/5956 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0002-4383-2451ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 13 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 How Humans Naturally Refer to Targets: Understanding Multimodal Instruction Patterns in Human-Robot Interaction
abstract
Current multimodal instruction-recognition algorithms in human-robot interaction, developed largely from a purely technical perspective, remain rigid and incomplete in their use of human communicative cues. Therefore, a full understanding of how humans naturally refer to targets in interaction is central to enabling robots to interpret and act on user instructions. To investigate this, we collected multimodal behavior data from 30 participants who naturally instructed a robot for household tasks while we systematically varied target distance, direction, and local referent complexity. Our results show that speech instructions were often vague and lacked explicit target-position information. To resolve this ambiguity, multimodal cues are essential: gaze direction provides an order-of-magnitude improvement in target-localization accuracy, while hand pointing, head turns, and speech onset offer reliable temporal anchors for identifying target-directed gaze. We also found that speech patterns varied with distance and local referent complexity, whereas multimodal behaviors shifted with target direction, underscoring the need for context-adaptive recognition and interface design.
Lesong Jia, Makayla Chang, Na Du
CHI4
2026 When to Explain: Modeling User Need for Explanations in Real-World Autonomous Driving
abstract
The integration of artificial intelligence into autonomous vehicles (AVs) raises transparency challenges that can hinder user acceptance and experience. To address when users need AV explanations, we created a large-scale dataset of 3327 diverse driving scenarios, each paired with a user-friendly explanation, and conducted an online study to survey when users need explanations. Using both scenario and user-related factors, our best-performing tree-ensemble models predicted explanation need with great performance (F1=0.72, AUC=0.82). SHAP analyses revealed that while both user- and scenario-related factors matter, factors directly related to driving (AV driving style, AV action, event cause, annual mileage, human driving style) were more contributive than general demographics and environmental factors. Our study delivers a comprehensively annotated dataset that underpins future human-AV interaction research, an explainable model that reliably predicts explanation needs, and valuable insights to inform the design of adaptive AV interfaces for superior user experiences.
Shihong Ling, Yaohan Ding, Yue Wan, Xiaowei Jia, Na Du
CHI6
2026 VRSafe: A Secure Virtual Keyboard to Mitigate Keystroke Inference in Virtual Reality
abstract
Password-based authentication is one of the most commonly used methods for verifying user identities, and its widespread usage continues in virtual reality (VR) applications. As a result, various forms of attacks on password-based authentication in traditional environments such as keystroke inference and shoulder surfing, are still effective in VR applications. While keystroke inference attacks on virtual keyboards have been studied extensively, few efforts have developed an effective and cost-efficient defense strategy to mitigate keystroke inferences in VR. To address this gap, this paper presents a novel QWERTY keyboard called VRSafe that is resilient to keystroke inference attacks. The proposed keyboard carefully introduces false positive keystrokes into the information collected by attackers during the typing process, making the inference of the original password difficult. VRSafe also incorporates a novel malicious login detector that can effectively identify unauthorized login attempts using credentials inferred from keystroke inference attacks with high detection rate and minimal time and memory cost. The proposed design is evaluated through both simulation experiments and a real-world user study, and the results show that VRSafe can significantly reduce the accuracy of keystroke inference attacks while incurring a modest overhead from a usability standpoint.
Na Du, Adam J. Lee, Balaji Palanisamy
CODASPY2
2026 Aligning Task Goals before Execution: Insights from Diverse User Groups into Human-Robot Communication in Domestic Settings
abstract
Integrating domestic robots into everyday life requires not only reliable execution but also prior alignment of task goals between humans and robots. While prior research has examined input interfaces and feedback strategies, it has largely focused on objective performance metrics and often overlooked user variability. To address this gap, we conducted a survey study with 113 participants across four groups: adolescents, younger adults, older adults, and wheelchair users. The survey captured participants’ expectations of future robots (roles, embodiments, and concerns) and their preferences for instruction delivery and robot feedback before execution. Our results reveal both shared and group-specific patterns. Across groups, participants prioritized efficiency in instruction delivery and reliability in robot feedback. Regarding group differences: adolescents emphasized efficiency, wheelchair users valued transparency, and older adults may benefit from additional explanations of novel interaction technologies. Based on these findings, we derive stage-aware, context-sensitive, and group-adaptive design principles and recommendations to guide future robot interfaces.
Lesong Jia, Breelyn Kane Styler, Na Du
HRI4
2026 Modeling Driver Situational Awareness in Takeover Scenarios Using Multimodal Data and Machine Learning
abstract
In conditionally automated driving, drivers out of the control loop may lack situational awareness (SA), leading to inappropriate takeovers. Monitoring a driver’s SA and providing alerts for overlooked objects is critical to enhancing the takeover safety and efficiency. This study aimed to construct predictive models for drivers’ SA of objects during takeover transitions. The model features include drivers’ physiological data before and after takeover requests as well as the environment and object attributes. The ground truth was obtained through a scene reconstruction task, yielding binary SA labels. The Support Vector Machine delivered the best model performance, achieving a macro F1 score of 0.75 and an accuracy of 0.77, when applied with a time window of 2-second pre-takeover request and 4-second post-takeover request. Our model predicts drivers’ SA of specific objects across diverse traffic conditions using short time windows, supporting timely and generalizable driver monitoring and takeover assistance.
Lesong Jia, Na Du
Int. J. Hum. Comput. Interact.2
2025 Explanations Help: Leveraging Human Capabilities to Detect Cyberattacks on Automated Vehicles
Yaohan Ding, Yiheng Feng, Na Du
CHI4
2025 More Than Automation: User Insights into the Functionality and Interface of Wheelchair-Mounted Robotic Arms
abstract
While voice-controlled automated Wheelchair-mounted robotic arms (WMRAs) could potentially offer more natural and simplified interactions than manual control, they introduce new challenges related to human-robot collaboration. To explore user expectations in terms of functionality and interface, we conducted semi-structured interviews with 13 powered wheelchair users who have upper limb impairments. A prototype with a robotic arm, a camera, and a Unity-based simulated interface was developed to help users understand the concept and functionality of the voice-controlled automated WMRA system during interviews. With safety as a priority, we found that users prioritized the WMRA's ability to grasp objects in challenging positions, such as on the ground or at high locations, and emphasized the need for a versatile gripper. Users also expected the WMRA to assist in performing complex daily tasks, with varying expectations for its performance based on task difficulty. Regarding the interface, users sought more information about the system's awareness and task execution, while emphasizing the importance of avoiding information overload. The demand for detailed safety information, such as temperature and gripping force, pointed to the need for enhanced sensor capabilities in the WMRA system. Additionally, concerns about privacy underscored the need for clear communication on privacy policies. Our results provide user-centered insights for automated WMRA, offering design implications and future research directions in areas such as user modeling, hardware, algorithms, and interface development.
Lesong Jia, Breelyn Kane Styler, Na Du
HRI3
2025 DriveBLIP2: Attention-Guided Explanation Generation for Complex Driving Scenarios
abstract
This paper introduces a new framework, Drive-Blip2, built upon the BLIP2-OPT architecture, to generate accurate and contextually relevant explanations for emerging driving scenarios. While existing vision-language models perform well in general tasks, they encounter difficulties in understanding complex, multi-object environments, particularly in real-time applications such as autonomous driving, where the rapid identification of key objects is crucial. To address this limitation, an Attention Map Generator is proposed to highlight significant objects relevant to driving decisions within critical video frames. By directing the model’s focus to these key regions, the generated attention map helps produce clear and relevant explanations, enabling drivers to better understand the vehicle’s decision-making process in critical situations. Evaluations on the DRAMA dataset reveal significant improvements in explanation quality, as indicated by higher BLEU, ROUGE, CIDEr, and SPICE scores compared to baseline models. These findings underscore the potential of targeted attention mechanisms in vision-language models for enhancing explainability in real-time autonomous driving.
Shihong Ling, Yue Wan, Xiaowei Jia, Na Du
IROS4
2025 Watch Out for Explanations: Information Type and Error Type Affect Trust and Situational Awareness in Automated Vehicles
abstract
Trust and situational awareness (SA) are critical for the acceptance and safety of automated vehicles (AVs). While AV explanations with different information types have been studied to enhance drivers' trust and SA, their effectiveness remains unclear when AVs make errors that do not trigger takeover requests. This study investigated the effects of information type, error type, and their interaction on drivers' trust in AVs, SA, and their relationships. We recruited 300 participants in an online video study with a 3 (information type:why,how,why + how) × 3 (error type: false alarm, miss, correct [no error]) mixed design.Howinformation describes the vehicle's action, whilewhyinformation refers to the reason for the vehicle's action. Linear mixed models showed that false alarms and misses were associated with lower SA compared with correct scenarios, but possibly due to different reasons. Compared with correct scenarios, both false alarms and misses were associated with lower trust, with misses even lower than false alarms, possibly due to the varying severity of potential consequences. Compared withwhyandwhy + howinformation,howinformation was generally associated with lower SA and a higher potential of overtrust in false alarms. Trust and SA had a negative linear relationship in misses and false alarms, while no correlations were found in correct scenarios. To mitigate potential overtrust and misinterpretation of situations when AVs make errors, it is crucial to maintain higher SA. We recommend includingwhyinformation in AV explanations and deploying AV decision systems that are less miss-prone.
Yaohan Ding, Lesong Jia, Na Du
IEEE Trans. Hum. Mach. Syst.3
2024 One Size Does Not Fit All: Designing and Evaluating Criticality-Adaptive Displays in Highly Automated Vehicles
abstract
To promote drivers’ overall experiences in highly automated vehicles, we designed three objective criticality-adaptive displays: IO display highlighting Influential Objects, CO display highlighting Critical Objects, and ICO display highlighting Influential and Critical Objects differently. We conducted an online video-based survey study with 295 participants to evaluate them in varying traffic conditions. Results showed that low-trust propensity participants found ICO display more useful while high-trust propensity participants found CO displays more useful. When interacting with vulnerable road users (VRUs), participants had higher situational awareness (SA) but worse non-driving related task (NDRT) performance. Aging and CO displays also led to slower NDRT reactions. Nonetheless, older participants found displays more useful. We recommend providing different criticality-adaptive displays based on drivers’ trust propensity, age, and NDRT choice to enhance driving and NDRT performance and suggest carefully treating objects of different categories in traffic.
Yaohan Ding, Lesong Jia, Na Du
CHI3
2024 Understanding Human-machine Cooperation in Game-theoretical Driving Scenarios amid Mixed Traffic
abstract
Introducing automated vehicles (AVs) on roads may challenge established norms as drivers of human-driven vehicles (HVs) interact with AVs. Our study explored drivers’ decisions in game-theoretical scenarios amid mixed traffic using an online survey study. We manipulated factors including interaction types (HV-HV vs. HV-AV), scenario types (chicken game vs. public goods game), vehicle driving styles (aggressive vs. conservative), and time constraints (high vs. low). The quantitative results showed that human drivers tended to “defect” more, that is, not cooperate, against vehicles with conservative driving styles. The effect of vehicle driving styles was pronounced when interacting with AVs and in chicken game scenarios. Drivers exhibited more “defection” in public goods game scenarios and the effect of scenario types was weakened under high time constraints. Only drivers with moderate driving styles “defected” more in HV-AV interaction. Our qualitative findings provide essential insights into how drivers perceived conditions and formulated strategies for decision-making.
Yutong Zhang 0009, Edmond Awad, Morgan R. Frank, Peng Liu 0030, Na Du
CHI5
2024 Improving Explainable Object-induced Model through Uncertainty for Automated Vehicles
abstract
The rapid evolution of automated vehicles (AVs) has the potential to provide safer, more efficient, and comfortable travel options. However, these systems face challenges regarding reliability in complex driving scenarios. Recent explainable AV architectures neglect crucial information related to inherent uncertainties while providing explanations for actions. To overcome such challenges, our study builds upon the "object-induced" model approach that prioritizes the role of objects in scenes for decision-making and integrates uncertainty assessment into the decision-making process using an evidential deep learning paradigm with a Beta prior. Additionally, we explore several advanced training strategies guided by uncertainty, including uncertainty-guided data reweighting and augmentation. Leveraging the BDD-OIA dataset, our findings underscore that the model, through these enhancements, not only offers a clearer comprehension of AV decisions and their underlying reasoning but also surpasses existing baselines across a broad range of scenarios.
Shihong Ling, Yue Wan, Xiaowei Jia, Na Du
HRI4
2022 Evaluating Effects of Enhanced Autonomy Transparency on Trust, Dependence, and Human-Autonomy Team Performance over Time
abstract
As autonomous systems become more complicated, humans may have difficulty deciphering autonomy-generated solutions and increasingly perceive autonomy as a mysterious black box. The lack of transparency contributes to the lack of trust in autonomy and suboptimal team performance. In response to this concern, researchers have proposed various methods to enhance autonomy transparency and evaluated how enhanced transparency could affect the people’s trust and the human-autonomy team performance. However, the majority of prior studies measured trust at the end of the experiment and averaged behavioral and performance measures across all trials in an experiment, yet overlooked the temporal dynamics of those variables. We have little understanding of how autonomy transparency affects trust, dependence, and performance over time. The present study, therefore, aims to fill the gap and examine such temporal dynamics. We develop a game Treasure Hunter wherein a human uncovers a map for treasures with the help from an intelligent assistant. The intelligent assistant recommends where the human should go next. The rationale behind each recommendation could be conveyed in a display that explicitly lists the option space (i.e., all the possible actions) and the reason why a particular action is the most appropriate in a given context. Results from a human-in-the-loop experiment with 28 participants indicate that by conveying the intelligent assistant’s decision-making rationale via the display, participants’ trust increases significantly and becomes more calibrated over time. Using the display also leads to a higher acceptance of recommendations from the intelligent agent.
Ruikun Luo, Na Du, Xi Jessie Yang
Int. J. Hum. Comput. Interact.2
2022 Predicting Driver Takeover Time in Conditionally Automated Driving
abstract
It is extremely important to ensure a safe takeover transition in conditionally automated driving. One of the critical factors that quantifies the safe takeover transition is takeover time. Previous studies identified the effects of many factors on takeover time, such as takeover lead time, non-driving tasks, modalities of the takeover requests, and scenario urgency. However, there is a lack of research to predict takeover time by considering these factors all at the same time. Toward this end, we used eXtreme Gradient Boosting (XGBoost) to predict the takeover time using a dataset from a meta-analysis study [Zhanget al.(2019)]. In addition, we used SHAP (SHapley Additive exPlanation) to analyze and explain the effects of the predictors on takeover time. We identified seven most critical predictors that resulted in the best prediction performance. Their main effects and interaction effects on takeover time were examined. The results showed that the proposed approach provided both good performance and explainability. Our findings have implications on the design of in-vehicle monitoring and alert systems to facilitate the interaction between the drivers and the automated vehicle.
Jackie Ayoub, Na Du, Xi Jessie Yang, Feng Zhou 0003
IEEE Trans. Intell. Transp. Syst.2
2021 Designing Alert Systems in Takeover Transitions: The Effects of Display Information and Modality
abstract
In conditionally automated driving, in-vehicle alert systems can provide drivers with information to assist their takeovers from automated driving. This study investigated how display modality and information influenced drivers’ acceptance of the in-vehicle alert systems under different event criticality situations. We conducted an online video study with a 3 (information type) × 3 (display modality) × 2 (event criticality) mixed design involving 60 participants. The results showed that considering drivers’ perceived usefulness and ease of use, presenting why only information was not sufficient for takeovers as compared to what will only information and why + what will information. Participants reported higher ease of use in the combination of speech and augmented reality condition when compared to the speech only condition. High event criticality led to drivers’ lower perceived usefulness and more negative opinions of the displays. The findings have implications for the design of in-vehicle alert systems during takeover transitions.
Na Du, Feng Zhou 0003, Dawn M. Tilbury, Lionel P. Robert Jr., Xi Jessie Yang
AutomotiveUI1
2020 Evaluating Effects of Cognitive Load, Takeover Request Lead Time, and Traffic Density on Drivers' Takeover Performance in Conditionally Automated Driving
abstract
In conditionally automated driving, drivers engaged in non-driving related tasks (NDRTs) have difficulty taking over control of the vehicle when requested. This study aimed to examine the relationships between takeover performance and drivers’ cognitive load, takeover request (TOR) lead time, and traffic density. We conducted a driving simulation experiment with 80 participants, where they experienced 8 takeover events. For each takeover event, drivers’ subjective ratings of takeover readiness, objective measures of takeover timing and quality, and NDRT performance were collected. Results showed that drivers had lower takeover readiness and worse performance when they were in high cognitive load, short TOR lead time, and heavy oncoming traffic density conditions. Interestingly, if drivers had low cognitive load, they paid more attention to driving environments and responded more quickly to takeover requests in high oncoming traffic conditions. The results have implications for the design of in-vehicle alert systems to help improve takeover performance.
Na Du, Jinyong Kim, Feng Zhou 0003, Elizabeth Pulver, Dawn M. Tilbury, Lionel P. Robert Jr., Anuj K. Pradhan, Xi Jessie Yang
AutomotiveUI1