VLDB 2026 Research / reviewers in the wild / expert
Jiangtao Gong
dblp:133/8162
· DBLP profile ↗
36ranked-venue papers
5as first author
34since 2021 · last 2026
0000-0002-4310-1894ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 18 · 4 first-author · 17 since 2021Artificial intelligence and machine learning · 16 · 16 since 2021Systems, architecture and hardware · 9 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FreeAskWorld: An Interactive and Closed-Loop Simulator for Human-Centric Embodied AIabstractAs embodied intelligence emerges as a core frontier in artificial intelligence research, simulation platforms must evolve beyond low-level physical interactions to capture complex, human-centered social behaviors. We introduce FreeAskWorld, an interactive simulation framework that integrates large language models (LLMs) for high-level behavior planning and semantically grounded interaction, informed by theories of intention and social cognition. Our framework supports scalable, realistic human-agent simulations and includes a modular data generation pipeline tailored for diverse embodied tasks.To validate the framework, we extend the classic Vision-and-Language Navigation (VLN) task into a semantically enriched Direction Inquiry setting, wherein agents can actively seek and interpret navigational guidance. We present and publicly release FreeAskWorld, a large-scale benchmark dataset comprising reconstructed environments, six diverse task types, 16 core object categories, 63,429 annotated sample frames, and more than 17 hours of interaction data to support training and evaluation of embodied AI systems. We benchmark VLN models, and human participants under both open-loop and closed-loop settings. Experimental results demonstrate that models fine-tuned on FreeAskWorld outperform their original counterparts, achieving enhanced semantic understanding and interaction competency. These findings underscore the efficacy of socially grounded simulation frameworks in advancing embodied AI systems toward sophisticated high-level planning and more naturalistic human-agent interaction. Yuhang Peng, Yizhou Pan, Xinning He, Jihaoyu Yang, Xinyu Yin, Xiaoji Zheng, Jiangtao Gong |
AAAI | 9 |
| 2026 | An LLM-based Simulation Framework for Embodied Conversational Agents in Psychological CounselingabstractDue to privacy concerns, open dialogue datasets for mental health are primarily generated through human or AI synthesis methods. However, the inherent implicit nature of psychological processes, particularly those of clients, poses challenges to the authenticity and diversity of synthetic data. In this paper, we propose ECAs (short for Embodied Conversational Agents), a framework for embodied agent simulation based on Large Language Models (LLMs) that incorporates multiple psychological theoretical principles. Using simulation, we expand real counseling case data into a nuanced embodied cognitive memory space and generate dialogue data based on high-frequency counseling questions. We validated our framework using the D4 dataset. First, we created a public ECAs dataset through batch simulations based on D4. Licensed counselors evaluated our method, demonstrating that it significantly outperforms baselines in simulation authenticity and necessity. Additionally, two LLM-based automated evaluation methods were employed to confirm the higher quality of the generated dialogues compared to the baselines. Lixiu Wu, Yuanrong Tang, Qisen Pan, Xianyang Zhan, Lanxi Xiao, Tianhong Wang 0009, Jiangtao Gong |
AAAI | 9 |
| 2026 | SVATA: A Spatial Visual Attention Tracking and Analysis Platform for Embodied Cognition Research
Xuchao Ren, Jiangtao Gong |
CHI | 5 |
| 2026 | 'I Will Dream Sweet Dreams': Understanding Remote Companionship Volunteer Activities for Factual Orphans in Rural ChinaabstractA significant number of De facto orphans in underdeveloped regions face potential mental health risks, while some novel interventions attempt to involve university student volunteers in providing companionship. However, there is little documented literature on current practices of such emotional support volunteer activities. We conducted a comprehensive investigation of multiple stakeholders involved in an online companionship program facilitated by university student volunteers, designed to provide remote emotional support to these children. Through field observations and interviews, we summarize current practices and identify benefits. We discovered that current remote volunteer initiatives face numerous challenges impeding companionship effectiveness. We summarized four potential technological requirements derived from these challenges. Our study provides the first documented account of remote volunteer companionship activities aimed at improving vulnerable children’s mental health. Our research can inspire future developments of technological solutions for improving emotional companionship in volunteer programs and vulnerable children’s wellbeing. Yuanrong Tang, Yueqing Hu, Tianhong Wang 0009, Hanchao Song, Zhicong Lu, Jiangtao Gong |
CHI | 7 |
| 2025 | Mentigo: An Intelligent Agent for Mentoring Students in the Creative Problem Solving ProcessabstractCreative Problem-Solving (CPS) promotes creative and critical thinking while enhancing real-world problem-solving skills, making it essential for middle school education.However, providing personalized mentorship in CPS projects at scale is challenging due to resource constraints and diverse student needs.To address this, we developed Mentigo, an AI-driven mentor agent designed to guide middle school students through the CPS process.Using a dataset of real classroom interactions, we encoded CPS task stages, adaptive guidance strategies, and personalized feedback mechanisms to inform Mentigo's dynamic mentoring framework powered by large language models (LLMs).A comparative experiment with 12 students and evaluations from five expert educators demonstrated improved student engagement, creativity, and task performance.Our findings highlight design implications for using LLM-based AI mentors to enhance CPS learning in educational environments. Siyu Zha, Yujia Liu 0004, Chengbo Zheng, Fuze Yu, Jiangtao Gong, Ying-Qing Xu |
CHI | 6 |
| 2025 | How Generative Music Affects the ISO Principle-Based Emotion-Focused Therapy: An EEG Study
Jiayu Bao, Yaxing Lyu, Yucheng Jin 0001, Jiangtao Gong |
CogSci | 5 |
| 2025 | A Comprehensive LLM-powered Framework for Driving Intelligence EvaluationabstractEvaluation methods for autonomous driving are crucial for algorithm optimization. However, due to the complexity of driving intelligence, there is currently no comprehensive evaluation method for the level of autonomous driving intelligence. In this paper, we propose an evaluation framework for driving behavior intelligence in complex traffic environments, aiming to fill this gap. We constructed a natural language evaluation dataset of human professional drivers and passengers through naturalistic driving experiments and post-driving behavior evaluation interviews. Based on this dataset, we developed an LLM-powered driving evaluation framework. The effectiveness of this framework was validated through simulated experiments in the CARLA urban traffic simulator and further corroborated by human assessment. Our research provides valuable insights for evaluating and designing more intelligent, human-like autonomous driving agents. The implementation details of the framework11https://github.com/AIR-DISCOVER/Driving-Intellenge-Evaluation-Framework and detailed information about the dataset22https://github.com/AIR-DISCOVER/Driving-Evaluation-Datasetcan be found at the provided links. Shanhe You, Xuewen Luo, Xinhe Liang, Jiashu Yu, Chen Zheng 0005, Jiangtao Gong |
ICRA | 6 |
| 2025 | Designing LLM-simulated Immersive Spaces to Enhance Autistic Children's Social Affordances Understanding in Traffic SettingsabstractOne of the key challenges faced by autistic children is understanding social affordances in complex environments, which further impacts their ability to respond appropriately to social signals. In traffic scenarios, this impairment can even lead to safety concerns. In this paper, we introduce an LLM-simulated immersive projection environment designed to improve this ability in autistic children while ensuring their safety. We first propose 17 design considerations across four major categories, derived from a comprehensive review of previous research. Next, we developed a system called AIroad, which leverages LLMs to simulate drivers with varying social intents, expressed through explicit multimodal social signals. AIroad helps autistic children bridge the gap in recognizing the intentions behind behaviors and learning appropriate responses through various stimuli. A user study involving 14 participants demonstrated that this technology effectively engages autistic children and leads to significant improvements in their comprehension of social affordances in traffic scenarios. Additionally, parents reported high perceived usability of the system. These findings highlight the potential of combining LLM technology with immersive environments for the functional rehabilitation of autistic children in the future. Yancheng Cao, Yangyang He, Shanhe You, Yulin Qiu, Chen Zheng 0005, Xin Tong 0004, Jiangtao Gong |
IUI | 12 |
| 2025 | Embodied Cognition Augmented End2End Autonomous DrivingabstractIn recent years, vision-based end-to-end autonomous driving has emerged as a new paradigm. However, popular end-to-end approaches typically rely on visual feature extraction networks trained under label supervision. This limited supervision framework restricts the generality and applicability of driving models. In this paper, we propose a novel paradigm termed $E^{3}AD$, which advocates for comparative learning between visual feature extraction networks and the general EEG large model, in order to learn latent human driving cognition for enhancing end-to-end planning. In this work, we collected a cognitive dataset for the mentioned contrastive learning process. Subsequently, we investigated the methods and potential mechanisms for enhancing end-to-end planning with human driving cognition, using popular driving models as baselines on publicly available autonomous driving datasets. Both open-loop and closed-loop tests are conducted for a comprehensive evaluation of planning performance. Experimental results demonstrate that the $E^{3}AD$ paradigm significantly enhances the end-to-end planning performance of baseline models. Ablation studies further validate the contribution of driving cognition and the effectiveness of comparative learning process. To the best of our knowledge, this is the first work to integrate human driving cognition for improving end-to-end autonomous driving planning. It represents an initial attempt to incorporate embodied cognitive data into end-to-end autonomous driving, providing valuable insights for future brain-inspired autonomous driving systems. Our code will be made available at https://github.com/AIR-DISCOVER/E-cubed-AD. Ling Niu, Xiaoji Zheng, Ziyuan Yang 0005, Bokui Chen, Jiangtao Gong |
NeurIPS | 7 |
| 2025 | COLP: Scaffolding Children's Online Long-Term Collaborative LearningabstractOnline collaborative learning is increasingly important, yet children still face challenges communicating and working together virtually, limiting their engagement in long-term teamwork. To address this, we designed the Children’s Online Long-term Program (COLP), a 16-week online project-based learning program grounded in multiple learning theories. The program was implemented with 67 upper primary school students (Grades 3–6, ages 8–13) across five provinces in China. Results show that over one-third of participants sustained engagement in online teamwork. Interviews with children and their parents further revealed key communication channels, benefits, and challenges. Notably, parents played multiple roles in supporting their children’s collaboration, especially through modeling and guidance. This study contributes to the design of long-term online collaborative learning interventions for children within computer-supported collaborative learning (CSCL) communities. Siyu Zha, Yuanrong Tang, Jiangtao Gong, Ying-Qing Xu |
Int. J. Hum. Comput. Interact. | 3 |
| 2025 | Designing child-centric AI learning environments: Insights from an LLM-powered creative project-based learning study
Siyu Zha, Yuehan Qiao, Qingyu Hu, Zhongsheng Li, Jiangtao Gong, Ying-Qing Xu |
Int. J. Hum. Comput. Stud. | 5 |
| 2024 | "It Must Be Gesturing Towards Me": Gesture-Based Interaction between Autonomous Vehicles and PedestriansabstractInteracting with pedestrians understandably and efficiently is one of the toughest challenges faced by autonomous vehicles (AVs) due to the limitations of current algorithms and external human-machine interfaces (eHMIs). In this paper, we design eHMIs based on gestures inspired by the most popular method of interaction between pedestrians and human drivers. Eight common gestures were selected to convey AVs’ yielding or non-yielding intentions at uncontrolled crosswalks from previous literature. Through a VR experiment (N1 = 31) and a following online survey (N2 = 394), we discovered significant differences in the usability of gesture-based eHMIs compared to current eHMIs. Good gesture-based eHMIs increase the efficiency of pedestrian-AV interaction while ensuring safety. Poor gestures, however, cause misinterpretation. The underlying reasons were explored: ambiguity regarding the recipient of the signal and whether the gestures are precise, polite, and familiar to pedestrians. Based on this empirical evidence, we discuss potential opportunities and provide valuable insights into developing comprehensible gesture-based eHMIs in the future to support better interaction between AVs and other road users. Xiang Chang, Zihe Chen, Xiaoyan Dong, Tingmin Yan, Haolin Cai, Zherui Zhou, Guyue Zhou, Jiangtao Gong |
CHI | 9 |
| 2024 | Understanding Human-AI Collaboration in Music Therapy Through Co-Design with TherapistsabstractThe rapid development of musical AI technologies has expanded the creative potential of various musical activities, ranging from music style transformation to music generation. However, little research has investigated how musical AIs can support music therapists, who urgently need new technology support. This study used a mixed method, including semi-structured interviews and a participatory design approach. By collaborating with music therapists, we explored design opportunities for musical AIs in music therapy. We presented the co-design outcomes involving the integration of musical AIs into a music therapy process, which was developed from a theoretical framework rooted in emotion-focused therapy. After that, we concluded the benefits and concerns surrounding music AIs from the perspective of music therapists. Based on our findings, we discussed the opportunities and design implications for applying musical AIs to music therapy. Our work offers valuable insights for developing human-AI collaborative music systems in therapy involving complex procedures and specific requirements. Guyue Zhou, Yucheng Jin 0001, Jiangtao Gong |
CHI | 5 |
| 2024 | SurrealDriver: Designing LLM-powered Generative Driver Agent Framework based on Human Drivers' Driving-thinking DataabstractLeveraging advanced reasoning capabilities and extensive world knowledge of large language models (LLMs) to construct generative agents for solving complex real-world problems is a major trend. However, LLMs inherently lack embodiment as humans, resulting in suboptimal performance in many embodied decision-making tasks. In this paper, we introduce a framework for building human-like generative driving agents using post-driving self-report driving-thinking data from human drivers as both demonstration and feedback. To capture high-quality, natural language data from drivers, we conducted urban driving experiments, recording drivers’ verbalized thoughts under various conditions to serve as chain-of-thought prompts and demonstration examples for the LLM-Agent. The framework’s effectiveness was evaluated through simulations and human assessments. Results indicate that incorporating expert demonstration data significantly reduced collision rates by 81.04% and increased human likeness by 50% compared to a baseline LLM-based agent. Our study provides insights into using natural language-based human demonstration data for embodied tasks. The driving-thinking dataset is available at https://github.com/AIR-DISCOVER/Driving-Thinking-Dataset. Zhijie Yi, Xiaoxi Shen, Huiling Peng, Xiaoan Liu, Jingli Qin, Jintao Xie, Peizhong Gao, Guyue Zhou, Jiangtao Gong |
IROS | 12 |
| 2024 | Driving Style Alignment for LLM-powered Driver AgentabstractRecently, LLM-powered driver agents have demonstrated considerable potential in the field of autonomous driving, showcasing human-like reasoning and decision-making abilities. However, current research on aligning driver agent behaviors with human driving styles remains limited, partly due to the scarcity of high-quality natural language data from human driving behaviors. To address this research gap, we propose a multi-alignment framework designed to align driver agents with human driving styles through demonstrations and feedback. Notably, we construct a natural language dataset of human driver behaviors through naturalistic driving experiments and post-driving interviews, offering high-quality human demonstrations for LLM alignment. The framework’s effectiveness is validated through simulation experiments in the CARLA urban traffic simulator and further corroborated by human evaluations. Our research offers valuable insights into designing driving agents with diverse driving styles. The implementation of the framework1and details of the dataset2can be found at the link. Anais Fernandez-Laaksonen, Jiangtao Gong |
IROS | 5 |
| 2024 | Large Language Models Powered Context-aware Motion Prediction in Autonomous DrivingabstractMotion prediction is among the most fundamental tasks in autonomous driving. Traditional methods of motion forecasting primarily encode vector information of maps and historical trajectory data of traffic participants, lacking a comprehensive understanding of overall traffic semantics, which in turn affects the performance of prediction tasks. In this paper, we utilized Large Language Models (LLMs) to enhance the global traffic context understanding for motion prediction tasks. We first conducted systematic prompt engineering, visualizing complex traffic environments and historical trajectory information of traffic participants into image prompts— Transportation Context Map (TC-Map), accompanied by corresponding text prompts. Through this approach, we obtained rich traffic context information from the LLM. By integrating this information into the motion prediction model, we demonstrate that such context can enhance the accuracy of motion predictions. Furthermore, considering the cost associated with LLMs, we propose a cost-effective deployment strategy: enhancing the accuracy of motion prediction tasks at scale with 0.7% LLM-augmented datasets. Our research offers valuable insights into enhancing the understanding of traffic scenes of LLMs and the motion prediction performance of autonomous driving. The source code is available at https://github.com/AIR-DISCOVER/LLM-Augmented-MTR and https://aistudio.baidu.com/projectdetail/7809548. Xiaoji Zheng, Lixiu Wu, Zhijie Yan, Yuanrong Tang, Hao Zhao 0002, Bokui Chen, Jiangtao Gong |
IROS | 8 |
| 2024 | More Than Routing: Joint GPS and Route Modeling for Refine Trajectory Representation LearningabstractTrajectory representation learning plays a pivotal role in supporting various downstream tasks, such as travel time estimation, trajectory classification and Top-k similar trajectory search. Traditional methods in order to filter the noise in GPS trajectories tend to focus on routing-based methods to simplify the trajectories. However, these approaches ignore the motion details contained in the GPS data, limiting the representation capability of trajectory representation learning. To fill this gap, we propose a novel representation learning framework that is Jointly G PS and Route Modeling based on self-supervised technology, namely JGRM. We consider GPS trajectory and route trajectory as the two modals of a single movement observation and fuse information through inter-modal information interaction. Specifically, we develop two encoders, each tailored to capture representations of GPS trajectories and route trajectories respectively. The representations from these two modalities are fed into a shared transformer for inter-modal information interaction. Eventually, we design three self-supervised tasks to train the model. We validate the effectiveness of the proposed method on two real-world datasets through extensive experiments. The experimental results show that JGRM significantly outperforms existing methods in both road segment representation and trajectory representation tasks. Our source code is available at Github https://github.com/mamazi0131/JGRM. Zheyan Tu, Xinhai Chen 0002, Yan Zhang 0122, Deguo Xia, Guyue Zhou, Yu Zheng 0004, Jiangtao Gong |
WWW | 9 |
| 2024 | Beyond digital privacy: Uncovering deeper attitudes toward privacy in cameras among older adults
Ka I Chan, Tongxin Sun, Tongtong Jin, Jihong Jeung, Jiangtao Gong |
Int. J. Hum. Comput. Stud. | 7 |
| 2024 | "I see it as a wellspring for my positive and upward journey in life.": Understanding Current Practices of Assistive Technology's Customized Modification in ChinaabstractDue to the significant differences in physical conditions and living environments of people with disabilities, standardized assistive technologies (ATs) often fail to meet their needs. Modified AT, especially DIY (Do It Yourself) ATs, are a popular solution in many high-income countries, but there is a lack of documentation for low- and middle-income areas, especially in China, where the culture of philanthropy is undeveloped. To understand the current situation in this paper, we conducted semi-structured interviews with 10 individuals with disabilities using modified ATs and 10 individuals involved in providing these including family members, standard assistive device manufacturers, and individuals employed for their modification skills, etc. Based on the results of the thematic analysis, we have summarized the general process of modified ATs for people with disabilities in China and the benefits these devices bring. We found that modified ATs not only make the lives of people with disabilities more comfortable and convenient but also bring them confidence, reduce social pressure, and even help them achieve self-realization. Additionally, we summarized the challenges they encountered before, during, and after the modification, including awareness gaps, family resistance, a lack of a business model, and so on. Specifically, we conducted a special case study about the typical business models and challenges currently faced by AT Modification Organizations in China. Our research provides important design foundations and research insights for the future of universal and personalized production of AT. Haokun Xin, Jiangtao Gong |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2023 | Work with AI and Work for AI: Autonomous Vehicle Safety Drivers' Lived ExperiencesabstractThe development of Autonomous Vehicle (AV) has created a novel job, the safety driver, recruited from experienced drivers to supervise and operate AV in numerous driving missions. Safety drivers usually work with non-perfect AV in high-risk real-world traffic environments for road testing tasks. However, this group of workers is under-explored in the HCI community. To fill this gap, we conducted semi-structured interviews with 26 safety drivers. Our results present how safety drivers cope with defective algorithms and shape and calibrate their perceptions while working with AV. We found that, as front-line workers, safety drivers are forced to take risks accumulated from the AV industry upstream and are also confronting restricted self-development in working for AV development. We contribute the first empirical evidence of the lived experience of safety drivers, the first passengers in the development of AV, and also the grassroots workers for AV, which can shed light on future human-AI interaction research. Mengdi Chu, Keyu Zong, Xin Shu 0009, Jiangtao Gong, Zhicong Lu, Kaimin Guo, Xinyi Dai, Guyue Zhou |
CHI | 4 |
| 2023 | MR.Brick: Designing A Remote Mixed-reality Educational Game System for Promoting Children's Social & Collaborative SkillsabstractChildren are one of the groups most influenced by COVID-19-related social distancing, and a lack of contact with peers can limit their opportunities to develop social and collaborative skills. However, remote socialization and collaboration as an alternative approach is still a great challenge for children. This paper presents MR.Brick, a Mixed Reality (MR) educational game system that helps children adapt to remote collaboration. A controlled experimental study involving 24 children aged six to ten was conducted to compare MR.Brick with the traditional video game by measuring their social and collaborative skills and analyzing their multi-modal playing behaviours. The results showed that MR.Brick was more conducive to children’s remote collaboration experience than the traditional video game. Given the lack of training systems designed for children to collaborate remotely, this study may inspire interaction design and educational research in related fields. Yudan Wu, Shanhe You, Zixuan Guo 0003, Guyue Zhou, Jiangtao Gong |
CHI | 6 |
| 2023 | "I am the follower, also the boss": Exploring Different Levels of Autonomy and Machine Forms of Guiding Robots for the Visually ImpairedabstractGuiding robots, in the form of canes or cars, have recently been explored to assist blind and low vision (BLV) people. Such robots can provide full or partial autonomy when guiding. However, the pros and cons of different forms and autonomy for guiding robots remain unknown. We sought to fill this gap. We designed autonomy-switchable guiding robotic cane and car. We conducted a controlled lab-study (N=12) and a field study (N=9) on BLV. Results showed that full autonomy received better walking performance and subjective ratings in the controlled study, whereas participants used more partial autonomy in the natural environment as demanding more control. Besides, the car robot has demonstrated abilities to provide a higher sense of safety and navigation efficiency compared with the cane robot. Our findings offered empirical evidence about how the BLV community perceived different machine forms and autonomy, which can inform the design of assistive robots. Yan Zhang 0122, Haole Guo, Qihe Chen, Mingming Fan 0001, Guyue Zhou, Jiangtao Gong |
CHI | 9 |
| 2023 | INT2: Interactive Trajectory Prediction at IntersectionsabstractMotion forecasting is an important component in autonomous driving systems. One of the most challenging problems in motion forecasting is interactive trajectory prediction, whose goal is to jointly forecasts the future trajectories of interacting agents. To this end, we present a large-scale interactive trajectory prediction dataset named INT2 for INTeractive trajectory prediction at INTersections. INT2 includes 612,000 scenes, each lasting 1 minute, containing up to 10,200 hours of data. The agent trajectories are auto-labeled by a high-performance offline temporal detection and fusion algorithm, whose quality is further inspected by human judges. Vectorized semantic maps and traffic light information are also included in INT2. Additionally, the dataset poses an interesting domain mismatch challenge. For each intersection, we treat rush-hour and non-rush-hour segments as different domains. We benchmark the best open-sourced interactive trajectory prediction method on INT2 and Waymo Open Motion, under in-domain and cross-domain settings. The dataset, code and models are publicly available at https://github.com/AIRDISCOVER/INT2. Zhijie Yan, Pengfei Li 0007, Zheng Fu, Shaocong Xu, Yongliang Shi, Xiaoxue Chen, Yuhang Zheng 0004, Yang Li 0178, Tianyu Liu 0008, Chuxuan Li, Nairui Luo, Zuoxu Wang, Yifeng Shi, Zhengxiao Han, Jirui Yuan, Jiangtao Gong, Guyue Zhou, Hang Zhao 0021, Hao Zhao 0002 |
ICCV | 19 |
| 2023 | Understanding Embodied Reference with Touch-Line Transformer
Yang Li 0178, Xiaoxue Chen, Hao Zhao 0002, Jiangtao Gong, Guyue Zhou, Federico Rossano, Yixin Zhu 0001 |
ICLR | 4 |
| 2023 | Planning Assembly Sequence with Graph TransformerabstractAssembly Sequence Planning (ASP) is the essential process for modern manufacturing, proven to be NP-complete thus its effective and efficient solution has been a challenge for researchers in the field. In this paper, we present a graph-transformer based framework for the ASP problem which is trained and demonstrated on a self-collected ASP database. The ASP database contains a self-collected set of LEGO models. The LEGO model is abstracted to a heterogeneous graph structure after a thorough analysis of the original structure and feature extraction. The ground truth assembly sequence is first generated by brute-force search and then adjusted manually to be in line with human rational habits. Based on this self-collected ASP dataset, we propose a heterogeneous graph-transformer framework to learn the latent rules for assembly planning. We evaluated the proposed framework in a series of experiments. The results show that the similarity of the predicted and ground truth sequences can reach 0.44, a medium correlation measured by Kendall's τ. Meanwhile, we compared the different effects of node features and edge features and generated a feasible and reasonable assembly sequence as a benchmark for further research. Our dataset and code are available on: htps://github.com/AIR-DISCOVER/ICRA_ASP. Lin Ma 0002, Jiangtao Gong, Hao Chen 0062, Hao Zhao 0002, Wenbing Huang 0001, Guyue Zhou |
ICRA | 2 |
| 2023 | Enable Natural Tactile Interaction for Robot Dog based on Large-format Distributed Flexible Pressure SensorsabstractTouch is an important channel for human-robot interaction, while it is challenging for robots to recognize human touch accurately and make appropriate responses. In this paper, we design and implement a set of large-format distributed flexible pressure sensors on a robot dog to enable natural human-robot tactile interaction. Through a heuristic study, we sorted out 81 tactile gestures commonly used when humans interact with real dogs and 44 dog reactions. A gesture classification algorithm based on ResNet is proposed to recognize these 81 human gestures, and the classification accuracy reaches 98.7%. In addition, an action prediction algorithm based on Transformer is proposed to predict dog actions from human gestures, reaching a 1-gram BLEU score of 0.87. Finally, we compare the tactile interaction with the voice interaction during a freedom human-robot-dog interactive playing study. The results show that tactile interaction plays a more significant role in alleviating user anxiety, stimulating user excitement and improving the acceptability of robot dogs. Lishuang Zhan, Yancheng Cao, Qitai Chen, Haole Guo, Jiasi Gao, Yiyue Luo, Shihui Guo, Guyue Zhou, Jiangtao Gong |
ICRA | 9 |
| 2023 | Annotating Covert Hazardous Driving Scenarios Online: Utilizing Drivers' Electroencephalography (EEG) SignalsabstractAs autonomous driving systems prevail, it is becoming increasingly critical that the systems learn from databases containing fine-grained driving scenarios. Most databases currently available are human-annotated; they are expensive, time-consuming, and subject to behavioral biases. In this paper, we provide initial evidence supporting a novel technique utilizing drivers' electroencephalography (EEG) signals to implicitly label hazardous driving scenarios while passively viewing recordings of real-road driving, thus sparing the need for manual annotation and avoiding human annotators' behavioral biases during explicit report. We conducted an EEG experiment using real-life and animated recordings of driving scenarios and asked participants to report danger explicitly whenever necessary. Behavioral results showed the participants tended to report danger only when overt hazards (e.g., a vehicle or a pedestrian appearing unexpectedly from behind an occlusion) were in view. By contrast, their EEG signals were enhanced at the sight of both an overt hazard and a covert hazard (e.g., an occlusion signalling possible appearance of a vehicle or a pedestrian from behind). Thus, EEG signals were more sensitive to driving hazards than explicit reports. Further, the Time-Series AI (TSAI, [1]) successfully classified EEG signals corresponding to overt and covert hazards. We discuss future steps necessary to materialize the technique in real life. Chen Zheng 0005, Muxiao Zi, Mengdi Chu, Yan Zhang 0122, Jirui Yuan, Guyue Zhou, Jiangtao Gong |
ICRA | 8 |
| 2023 | Can Quadruped Guide Robots be Used as Guide Dogs?abstractQuadruped robots have the potential to guide blind and low vision (BLV) people due to their highly flexible locomotion and emotional value provided by their bionic forms. However, the development of quadruped guide robots rarely involves BLV users' participatory designs and evaluations. In this paper, we conducted two empirical experiments both in indoor controlled and outdoor field scenarios, exploring the benefits and drawbacks of quadruped guide robots. The results show that the nowadays commercial quadruped robots exposed significant disadvantages in usability and trust compared with wheeled robots. It is concluded that the moving gait and walking noise of quadruped robots would limit the guiding effectiveness to a certain extent, and the empathetic effect of its bionic form for BLV users could not be fully reflected. Based on the findings of wheeled robots and quadruped robots' advantages, we discuss the design implications for the future guide robot design for BLV users. This paper reports the first empirical experiment about quadruped guide robots with BLV users and preliminary explores their potential improvement space in substituting guide dogs, which can inspire the further specialized design of quadruped guide robots. Qihe Chen, Yan Zhang 0122, Tingmin Yan, Guyue Zhou, Jiangtao Gong |
IROS | 8 |
| 2023 | Side-by-Side vs Face-to-Face: Evaluating Colocated Collaboration via a Transparent Wall-sized DisplayabstractTraditional wall-sized displays mostly only support side-by-side co-located collaboration, while transparent displays naturally support face-to-face interaction. Many previous works assume transparent displays support collaboration. Yet it is unknown how exactly its afforded face-to-face interaction can support loose or close collaboration, especially compared to the side-by-side configuration offered by traditional large displays. In this paper, we used an established experimental task that operationalizes different collaboration coupling and layout locality, to compare pairs of participants collaborating side-by-side versus face-to-face in each collaborative situation. We compared quantitative measures and collected interview and observation data to further illustrate and explain our observed user behavior patterns. The results showed that the unique face-to-face collaboration brought by transparent display can result in more efficient task performance, different territorial behavior, and both positive and negative collaborative factors. Our findings provided empirical understanding about the collaborative experience supported by wall-sized transparent displays and shed light on its future design. Jiangtao Gong, Mengdi Chu, Minghao Luo, Liuxin Zhang, Yaqiang Wu, Qianying Wang 0002, Can Liu 0003 |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2022 | Remote Co-teaching in Rural Classroom: Current Practices, Impacts, and ChallengesabstractThe shortage of high-quality teachers is one of the biggest educational problems faced by underdeveloped areas. With the development of information and communication technologies (ICTs), China has begun a remote co-teaching intervention program using ICTs for rural classes, forming a unique “co-teaching classroom”. We conducted semi-structured interviews with nine remote urban teachers and twelve local rural teachers. We identified the remote co-teaching classes’ standard practices and co-teachers’ collaborative work process. We also found that remote teachers’ high-quality class directly impacted local teachers and students. Furthermore, interestingly, local teachers were also actively involved in making indirect impacts on their students by deeply coordinating with remote teachers and adapting the resources offered by the remote teachers. We conclude by summarizing and discussing the challenges faced by teachers, lessons learned from the current program, and related design implications to achieve a more adaptive and sustainable ICT4D program design. Siling Guo, Tianchen Sun, Jiangtao Gong, Zhicong Lu, Liuxin Zhang, Qianying Wang 0002 |
CHI | 3 |
| 2022 | Learning with Yourself: a Tangible Twin Robot System to Promote STEM EducationabstractThis paper presents a customized programmable robotic system, TanTwin (Tangible Twin), designed to promote STEM education for K-12 children. Firstly, TanTwin is implemented based on a wheel-robot with standard LEGO bricks. With several deep neural networks, a child can convert a captured portrait of himself/herself into standard LEGO bricks, therefore he/she can build a tangible twin robot of him-selflherself automatically. Besides, to adapt to the customized appearance, the corresponding visual element and content of the robotic system were also changed by a rule-based adaption algorithm. To demonstrate the effectiveness of TanTwin and to investigate whether tangible twin robots could contribute to children's learning, we conducted a controlled experimental study to compare learning with a TanTwin and with a standard robot system through measuring students' cognitive learning outcomes. The pre-/post- knowledge test results indicated that learning with a tangible twin robot leads to significantly better learning outcomes. Given the results, we validate our system and customization technology can promote STEM education. Jiasi Gao, Jiangtao Gong, Guyue Zhou, Haole Guo, Tong Qi |
IROS | 2 |
| 2021 | All in One Group: Current Practices, Lessons and Challenges of Chinese Home-School Communication in IM Group ChatabstractWhen schools and families form a good partnership, children benefit. With the recent flourishing of communication apps, families and schools in China have shifted their primary communication channels to chat groups hosted on popular instant-messenger(IM) tools such as WeChat and QQ. With an interview study consisting of 18 parents and 9 teachers, followed by a survey study with 210 teachers, we found that IM group chat has become the most popular way that the majority of parents and teachers communicate, from among the many different channels available. While there are definite advantages to this kind of group chat, we also found a number of problematic issues, including a lack of privacy and repeated negative feedback shared by both parents and teachers. We discuss our results on how IM-based group chat could affect Chinese teachers’ authoritative figures, affect Chinese teacher’s work-life balance and potentially compromise Chinese students’ privacy. Jiangtao Gong, Zhicong Lu, Qicheng Ding, Yu Zhang 0124, Liuxin Zhang, Qianying Wang 0002 |
CHI | 1 |
| 2021 | HoloBoard: a Large-format Immersive Teaching Board based on pseudo HoloGraphicsabstractIn this paper, we present HoloBoard, an interactive large-format pseduo-holographic display system for lecture based classes. With its unique properties of immersive visual display and transparent screen, we designed and implemented a rich set of novel interaction techniques like immersive presentation, role-play, and lecturing behind the scene that are potentially valuable for lecturing in class. We conducted a controlled experimental study to compare a HoloBoard class with a normal class through measuring students’ learning outcomes and three dimensions of engagement (i.e., behavioral, emotional, and cognitive engagement). We used pre-/post- knowledge tests and multimodal learning analytics to measure students’ learning outcomes and learning experiences. Results indicated that the lecture-based class utilizing HoloBoard lead to slightly better learning outcomes and a significantly higher level of student engagement. Given the results, we discussed the impact of HoloBoard as an immersive media in the classroom setting and suggest several design implications for deploying HoloBoard in immersive teaching practices. Jiangtao Gong, Teng Han, Siling Guo, Jiannan Li, Siyu Zha, Liuxin Zhang, Feng Tian 0001, Qianying Wang 0002, Yong Rui |
UIST | 1 |
| 2021 | Grabbing the Long Tail: A data normalization method for diverse and informative dialogue generation
Zhiqiang Zhan, Yang Zhang 0002, Jiangtao Gong, Qianying Wang 0002, Liuxin Zhang |
Neurocomputing | 4 |
| 2020 | "I can't name it, but I can perceive it" Conceptual and Operational Design of "Tactile Accuracy" Assisting Tactile Image CognitionabstractDesigning a tactile image for blind people is a significant challenge due to the difficulty of recognizing objects on a 2D line drawing image by touch compared to vision. In this paper, we proposed ”tactile accuracy”, a new criterion to evaluate the performance of recognizing 242 raised line images of common objects for 30 subjects (10 blindfolded sighted subjects, 10 congenitally blind subjects, and 10 late blind subjects), instead of the conventional ”naming accuracy” used in the visual image recognition tasks. We used multi-level evaluation criteria including ”tactile accuracy” to systematically analyze the design factors in tactile images. The results showed that using multi-level evaluate criteria could help unveil the tactile cognitive preferences of different types of subjects for personalized learning. Moreover, we reported important design factors that affect tactile image recognition, thus providing guidelines on the design of tactile images. Jiangtao Gong, Wenyuan Yu, Long Ni, Ye Liu 0010, Xiaolan Fu, Ying-Qing Xu |
ASSETS | 1 |
| 2013 | BugMap: a topographic map of bugsabstractA large and complex software system could contain a large number of bugs. It is desirable for developers to understand how these bugs are distributed across the system, so they could have a better overview of software quality. In this paper, we describe BugMap, a tool we developed for visualizing large-scale bug location information. Taken source code and bug data as the input, BugMap can display bug localizations on a topographic map. By examining the topographic map, developers can understand how the components and files are affected by bugs. We apply this tool to visualize the distribution of Eclipse bugs across components/files. The results show that our tool is effective for understanding the overall quality status of a large-scale system and for identifying the problematic areas of the system. Jiangtao Gong, Hongyu Zhang 0002 |
ESEC/SIGSOFT FSE | 1 |