VLDB 2026 Research / reviewers in the wild / expert
Peizhong Gao
dblp:236/9582
· DBLP profile ↗
5ranked-venue papers
3as first author
4since 2021 · last 2024
0000-0002-9965-7795ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SurrealDriver: Designing LLM-powered Generative Driver Agent Framework based on Human Drivers' Driving-thinking DataabstractLeveraging advanced reasoning capabilities and extensive world knowledge of large language models (LLMs) to construct generative agents for solving complex real-world problems is a major trend. However, LLMs inherently lack embodiment as humans, resulting in suboptimal performance in many embodied decision-making tasks. In this paper, we introduce a framework for building human-like generative driving agents using post-driving self-report driving-thinking data from human drivers as both demonstration and feedback. To capture high-quality, natural language data from drivers, we conducted urban driving experiments, recording drivers’ verbalized thoughts under various conditions to serve as chain-of-thought prompts and demonstration examples for the LLM-Agent. The framework’s effectiveness was evaluated through simulations and human assessments. Results indicate that incorporating expert demonstration data significantly reduced collision rates by 81.04% and increased human likeness by 50% compared to a baseline LLM-based agent. Our study provides insights into using natural language-based human demonstration data for embodied tasks. The driving-thinking dataset is available at https://github.com/AIR-DISCOVER/Driving-Thinking-Dataset. Zhijie Yi, Xiaoxi Shen, Huiling Peng, Xiaoan Liu, Jingli Qin, Jintao Xie, Peizhong Gao, Guyue Zhou, Jiangtao Gong |
IROS | 10 |
| 2024 | Mul-O: Encouraging Olfactory Innovation in Various Scenarios Through a Task-Oriented Development PlatformabstractOlfactory interfaces are pivotal in HCI, yet their development is hindered by limited application scenarios, stifling the discovery of new research opportunities. This challenge primarily stems from existing design tools focusing predominantly on odor display devices and the creation of standalone olfactory experiences, rather than enabling rapid adaptation to various contexts and tasks. Addressing this, we introduce Mul-O, a novel task-oriented development platform crafted to aid semi-professionals in navigating the diverse requirements of potential application scenarios and effectively prototyping ideas. Mul-O facilitates the swift association and integration of olfactory experiences into functional designs, system integrations, and concept validations. Comprising a web UI for task-oriented development, an API server for seamless third-party integration, and wireless olfactory display hardware, Mul-O significantly enhances the ideation and prototyping process in multisensory tasks. This was verified by a 15-day workshop attended by 30 participants. The workshop produced seven innovative projects, underscoring Mul-O’s efficacy in fostering olfactory innovation. Peizhong Gao, Fan Liu 0021, Di Wen 0008, Yuze Gao, Linxin Zhang, Chikelei Wang, Yu Zhang 0124, Shao-en Ma, Qi Lu 0001, Haipeng Mi, Ying-Qing Xu |
UIST | 1 |
| 2024 | OdorAgent: Generate Odor Sequences for Movies Based on Large Language ModelabstractNumerous studies have shown that integrating scents into movies enhances viewer engagement and immersion. However, creating such olfactory experiences often requires professional perfumers to match scents, limiting their widespread use. To address this, we propose OdorAgent which combines a LLM with a text-image model to automate video-odor matching. The generation framework is in four dimensions: subject matter, emotion, space, and time. We applied it to a specific movie and conducted user studies to evaluate and compare the effectiveness of different system elements. The results indicate that OdorAgent possesses significant scene adaptability and enables inexperienced individuals to design odor experiences for video and images. Yu Zhang 0124, Peizhong Gao, Fangzhou Kang, Qi Lu 0001, Ying-Qing Xu |
VR | 2 |
| 2023 | Bamboo Agents: Exploring the Potentiality of Digital Craft by Decoding and Recoding ProcessabstractAs an emerging field in HCI, Digital Craft is often involved in debate on its concept that leads to distinctive practices. In this paper, the authors argue that the hybridization of digital power as computing and fabrication and human skills of ideation and hand-making sheds light on an important research direction for future inquiry of digital craft. In particular, the reported project focuses on bamboo craft making to explore the potentiality of digital craft through constructive design research methods. Through the collaborative making with craftspeople, methods of hybridizing digital power and human skills are explored by decoding and recoding the bamboo making process with inventive digital intervention, which enriches bamboo artifacts’ forms and integrates digital fabrication, such as 3D printing. Meanwhile, digital toolkits for bamboo weaving and a digital platform that can make computational design and craft more compatible are created and reported. The authors conclude with reported hybrid craft cases that the hybridization has the potential to stimulate a new generation of artisans. Peizhong Gao, Tanhao Gao, Yanbin Yang, Zhenyuan Liu 0004, Jianyu Shi, Jin Li 0039 |
TEI | 1 |
| 2020 | Visual-Auditory Redirection: Multimodal Integration of Incongruent Visual and Auditory Cues for Redirected WalkingabstractIn this paper, we present a study of redirected walking (RDW) that shifts the positional relationship between visual and auditory cues during curvature manipulation. It has been shown that, when presented with incongruent visual and auditory spatial cues during a localization task, human observers integrate that information based on each cue's relative reliability, which determines their final perception of the target object's location. This multi-modal integration model is known as maximum likelihood estimation (MLE). By altering the visual location of objects that users perceive in virtual reality (VR) through auditory cues during redirection manipulation, we expect fewer users to notice the manipulation, which helps increase the usable curvature gain. Most existing studies on MLE in multi-modal integration have used random-dot stereograms as visual cues under stable motion states. In the present study, we first investigated whether this model holds while walking in VR environment. Our results indicate that in a walking state, users' perceptions of the target object's location shift toward auditory cue as the reliability of vision decreases, in keeping with the trend shown in previous studies on MLE. Based on this result, we then investigated the detection threshold of curvature gains during redirection manipulation under a condition with congruent visual-auditory cues as well as a condition in which users' location perceptions of the target object are considered to be affected by the incongruent auditory cue. We found that the detection threshold of curvature gains was higher with incongruent visual-auditory cues than with congruent cues. These results show that incongruent multimodal cues in VR may have a promising application in the area of redirected walking. Peizhong Gao, Keigo Matsumoto, Takuji Narumi, Michitaka Hirose |
ISMAR | 1 |