VLDB 2026 Research / reviewers in the wild / expert
Chia-Ming Chang 0003
dblp:77/6881-3
· DBLP profile ↗
21ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0003-0390-6361ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 18 · 6 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Don't Worry, Just Follow Me: Prototyping and In-the-Wild Evaluation of Smart Pole Interaction Unit with MobilityabstractPedestrian–automated vehicle (AV) encounters in shared spaces often involve hesitation and ambiguity. Vehicle-mounted external human–machine interfaces (eHMIs) can help, but obscured or poorly timed communications create significant challenges. To address this, we present a mobile smart pole interaction unit (SPIU) with integrated cameras and LED displays, designed as a pedestrian-side system to deliver explicit cues (“WALK,” “STOP”). An in-the-wild evaluation of the SPIU (N = 21) using a four-factor analysis (CarBehavior, Mobility, eHMI, SPIU) showed that the SPIU improved understandability, trust, and perceived safety, and reduced workload compared with the baseline, with a combination (eHMI+SPIU) yielding the strongest results. Beyond these quantitative benefits, participants appreciated the mobility of the SPIU for its “clear” and “easy to decide” mediation. This work contributes to (1) a design and deployment framework for a mobile SPIU and (2) an in-the-wild evaluation protocol for pedestrian–AV interactions in nonsignalized spaces. Our work sparks discussions on real world evaluations involving detailed vehicle kinematics and accessible multimodality (e.g., audio), focusing on the role of personal robots as user-side eHMIs. Vishal Chauhan, Anubhav, Mark Colley, Chia-Ming Chang 0003, Xinyue Gui, Ding Xia, Ehsan Javanmardi, Takeo Igarashi, Kantaro Fujiwara, Manabu Tsukada |
CHI | 4 |
| 2026 | Peeking Ahead of the Field Study: Exploring VLM Personas as Support Tools for Embodied Studies in HCIabstractField studies are irreplaceable but costly, time-consuming, and error-prone, which need careful preparation. Inspired by rapid-prototyping in manufacturing, we propose a fast, low-cost evaluation method using Vision-Language Model (VLM) personas to simulate outcomes comparable to field results. While LLMs show human-like reasoning and language capabilities, autonomous vehicle (AV)-pedestrian interaction requires spatial awareness, emotional empathy, and behavioral generation. This raises our research question: To what extent can VLM personas mimic human responses in field studies? We conducted parallel studies: 1) one real-world study with 20 participants, and 2) one video-study using 20 VLM personas, both on a street-crossing task. We compared their responses and interviewed five HCI researchers on potential applications. Results show that VLM personas mimic human response patterns (e.g., average crossing times of 5.25 s vs. 5.07 s) lack the behavioral variability and depth. They show promise for formative studies, field study preparation, and human data augmentation. Xinyue Gui, Ding Xia, Mark Colley, Vishal Chauhan, Anubhav, Zhongyi Zhou, Ehsan Javanmardi, Stela Hanbyeol Seo, Chia-Ming Chang 0003, Manabu Tsukada, Takeo Igarashi |
CHI | 10 |
| 2025 | A Silent Negotiator? Cross-cultural VR Evaluation of Smart Pole Interaction Units in Dynamic Shared SpacesabstractAs autonomous vehicles (AVs) enter pedestrian-centric environments, existing vehicle-mounted external human–machine interfaces (eHMIs) often fall short in shared spaces due to line-of-sight limitations, inconsistent signaling, and increased decision latency on pedestrians. To address these challenges, we introduce the Smart Pole Interaction Unit (SPIU), an infrastructure-based eHMI that decouples intent signaling from vehicles and provides context-aware, elevated visual cues. We evaluate SPIU using immersive VR-AWSIM simulations in four high-risk urban scenarios: four-way intersections, autonomous mixed traffic, blindspots, and nighttime crosswalks. The experiment was developed in Japan and replicated in Norway, where forty participants engaged in 32 trials each under both SPIU-present and SPIU-absent conditions. Behavioral (response time) and subjective (acceptance scale) data were collected. Results show that SPIU significantly improves pedestrian decision-making, with reductions ranging from 40% to over 80% depending on scenario and cultural context, particularly in complex or low-visibility scenarios. Cross-cultural analyses highlight SPIU’s adaptability across differing urban and social contexts. We release our open-source Smartpole-VR-AWSIM framework to support reproducibility and global advancement of infrastructure-based eHMI research through reproducible and immersive behavioral studies. Vishal Chauhan, Anubhav, Robin Sidhu, Yu Asabe, Kanta Tanaka, Chia-Ming Chang 0003, Xiang Su 0001, Ehsan Javanmardi, Takeo Igarashi, Alex Orsholits, Kantaro Fujiwara, Manabu Tsukada |
VRST | 6 |
| 2025 | Towards the future of pedestrian-AV interaction: Human perception vs. LLM insights on Smart Pole Interaction Unit in shared spaces
Vishal Chauhan, Anubhav, Chia-Ming Chang 0003, Xiang Su 0001, Jin Nakazato, Ehsan Javanmardi, Alex Orsholits, Takeo Igarashi, Kantaro Fujiwara, Manabu Tsukada |
Int. J. Hum. Comput. Stud. | 3 |
| 2025 | The Mokume Dataset and Inverse Modeling of Solid Wood TexturesabstractWe present the Mokume dataset for solid wood texturing consisting of 190 cube-shaped samples of various hard and softwood species documented by high-resolution exterior photographs, annual ring annotations, and volumetric computed tomography (CT) scans. A subset of samples further includes photographs along slanted cuts through the cube for validation purposes. Using this dataset, we propose a three-stage inverse modeling pipeline to infer solid wood textures using only exterior photographs. Our method begins by evaluating a neural model to localize year rings on the cube face photographs. We then extend these exterior 2D observations into a globally consistent 3D representation by optimizing a procedural growth field using a novel iso-contour loss. Finally, we synthesize a detailed volumetric color texture from the growth field. For this last step, we propose two methods with different efficiency and quality characteristics: a fast inverse procedural texture method, and a neural cellular automaton (NCA). We demonstrate the synergy between the Mokume dataset and the proposed algorithms through comprehensive comparisons with unseen captured data. We also present experiments demonstrating the efficiency of our pipeline's components against ablations and baselines. Our code, the dataset, and reconstructions are available via https://mokumeproject.github.io/. Maria Larsson, Hodaka Yamaguchi, Ehsan Pajouheshgar, I-Chao Shen, Kenji Tojo, Chia-Ming Chang 0003, Lars Hansson, Olof Broman, Takashi Ijiri, Ariel Shamir, Wenzel Jakob, Takeo Igarashi |
ACM Trans. Graph. | 6 |
| 2024 | Shrinkable Arm-based eHMI on Autonomous Delivery Vehicle for Effective Communication with Other Road UsersabstractWhen employing autonomous driving technology in logistics, small autonomous delivery vehicles (aka delivery robots) encounter challenges different from passenger vehicles when interacting with other road users. We conducted an online video survey as a pre-study and found that autonomous delivery vehicles need external human-machine interfaces (eHMIs) to ask for help due to their small size and functional limitations. Inspired by everyday human communication, we chose arms as eHMI to show their request through limb motion and gesture. We held an in-house workshop to identify the arm’s requirements for designing a specific arm with shrink-ability (conspicuous when delivering messages but not affect traffic at other times). We prototyped a small delivery robot with a shrinkable arm and filmed the experiment videos. We conducted two studies (a video-based and a 360-degree-photo VR-based) with 18 participants. We demonstrated that arm-on-delivery robots can increase interaction efficiency by drawing more attention and communicating specific information. Xinyue Gui, Mikiya Kusunoki, Bofei Huang, Stela Hanbyeol Seo, Chia-Ming Chang 0003, Haoran Xie 0002, Manabu Tsukada, Takeo Igarashi |
AutomotiveUI | 5 |
| 2024 | Speed Labeling: Non-stop Scrolling for Fast Image LabelingabstractThis study presents “Speed Labeling”, an image-labeling technique to increase the efficiency of easy binary labeling tasks where an annotator can choose a label instantly. We first conduct a formative study to identify the factors affecting the efficiency of easy image labeling: image layout and image transition. Based on these results, we designed a novel labeling technique using non- stop scrolling. In conventional image labeling, the system moves to the next image only after the user assigns a label to the previous image. To maximize efficiency, our technique continuously scrolls images without waiting for the completion of labeling, assuming that the user gives labels at a mostly constant speed. The system dynamically adjusts the scrolling speed based on the labeling speed. Subsequently, we conduct a user study to compare the proposed “non-stop scrolling” technique to the conventional “stop-and-go scrolling” technique in an easy image-labeling task. The results showed that speed labeling requires less time (faster by 7%, 305 more images labeled per man-hour) to complete the labeling task than the conventional technique without a significant increase in errors. In addition, the results showed that speed labeling makes the labeling task more enjoyable for crowd workers and makes them feel more attentive during tasks. Chia-Ming Chang 0003, Xi Yang 0017, Xiang 'Anthony' Chen, Takeo Igarashi |
Graphics Interface | 1 |
| 2024 | "Text + Eye" on Autonomous Taxi to Provide Geospatial Instructions to PassengerabstractWhile text-based external human-machine interface (eHMI) is widely accepted, one limitation is the lack of capability to communicate spatial information such as a different person or location. We built a mixed-eHMI using "eye" as a target-specifier when "text" shows the clear intention to their communication partners. We conducted a pre-experimental observation to develop two testbed scenarios, followed by a video-based user study via life-size projection with a real-car prototype mounted a text display and a set of robotic eyes. The results demonstrated that our proposed "text + eye" combination may represent geospatial information by increasing the success pick-up rate. Xinyue Gui, Ehsan Javanmardi, Stela Hanbyeol Seo, Vishal Chauhan, Chia-Ming Chang 0003, Manabu Tsukada, Takeo Igarashi |
HAI | 5 |
| 2024 | PDFChatAnnotator: A Human-LLM Collaborative Multi-Modal Data Annotation Tool for PDF-Format CatalogsabstractThe document contains substantial unannotated data, necessitating extensive manual labeling efforts. To address this issue, we introduce PDFChatAnnotator, a human-LLM collaborative tool to collect multi-modal data from PDF catalogs. Initially, PDFChatAnnotator automatically employs our proposed multi-modal binding rules to link related data from different modalities and harnesses the information extraction capabilities of large language models (LLMs) to extract specific information from text descriptions. Furthermore, the tool empowers users to guide and refine the LLM’s annotations. During the annotation process, users can influence the LLM through multiple rounds of communication and example establishment via the provided interfaces. To assess the effectiveness of PDFChatAnnotator’s techniques, we conducted a technical evaluation using three catalogs with typical layouts as experimental data. The results showed that all accuracy rates for multi-modal binding exceeded 90%, and both the proposed "example establishment" and "interactive adjustment of requirements" contributed to enhanced accuracy rates. Chia-Ming Chang 0003, Xi Yang 0017 |
IUI | 2 |
| 2024 | SpaceEditing: A Latent Space Editing Interface for Integrating Human Knowledge into Deep Neural NetworksabstractHuman-centered AI aims to bridge the gap between machine decision-making and human understanding. However, even for classification tasks where deep neural networks have achieved superb performance, there are currently few methods that link humans and AI well, especially on domain-specific tasks. In this paper, we propose SpaceEditing, a 2D spatial layout tool that enables human users to interact with the latent space of deep neural networks. During the interaction process, the tool’s algorithm automatically processes user actions, providing feedback to the network and leveraging triplet loss to effectively learn from user-modified information. We evaluate SpaceEditing with three case studies: (1) an archaeology researcher uses a bronze dataset; (2) a deep learning researcher uses a garbage classification dataset; (3) six deep learning beginners use a head pose dataset. The experimental results demonstrate the effectiveness of our tool in integrating human knowledge and improving network performance. Jiafu Wei, Ding Xia, Haoran Xie 0002, Chia-Ming Chang 0003, Xi Yang 0017 |
IUI | 4 |
| 2023 | Efficient Human-in-the-loop System for Guiding DNNs AttentionabstractAttention guidance is used to address dataset bias in deep learning, where the model relies on incorrect features to make decisions. Focusing on image classification tasks, we propose an efficient human-in-the-loop system to interactively direct the attention of classifiers to regions specified by users, thereby reducing the effect of co-occurrence bias and improving the transferability and interpretability of a deep neural network (DNN). Previous approaches for attention guidance require the preparation of pixel-level annotations and are not designed as interactive systems. We herein present a new interactive method that allows users to annotate images via simple clicks. Additionally, we identify a novel active learning strategy that can significantly reduce the number of annotations. We conduct both numerical evaluations and a user study to evaluate the proposed system using multiple datasets. Compared with the existing non-active-learning approach, which typically relies on considerable amounts of polygon-based segmentation masks to fine-tune or train the DNNs, our system can obtain fine-tuned networks on biased datasets in a more time- and cost-efficient manner and offers a more user-friendly experience. Our experimental results show that the proposed system is efficient, reasonable, and reliable. Our code is publicly available at https://github.com/ultratykis/Guiding-DNNs-Attention. Xi Yang 0017, Chia-Ming Chang 0003, Haoran Xie 0002, Takeo Igarashi |
IUI | 3 |
| 2023 | SyncLabeling: A Synchronized Audio Segmentation Interface for Mobile DevicesabstractManual audio segmentation is a time-consuming process, especially when there is more than one sound playing simultaneously that needs to be segmented and annotated (e.g., target and background sounds). In conventional audio annotation interfaces, users need to repeatedly pause and replay the audio to complete an overlap segmentation task, which is very inefficient. In this paper, we propose "SyncLabeling," a synchronized audio segmentation interface for smartphones that allows users to segment and annotate two overlapping sounds in a single audio stream at a time using a game-like labeling interface on mobile devices. We conducted a user study to compare the proposed SyncLabeling interface with a conventional audio annotation interface on four types of audio segmentation tasks. The results showed that the proposed interface is much more efficient than the conventional interface (2.4× faster) under comparable annotation accuracy in most tasks. In addition, more than half of the participants enjoyed using the proposed SyncLabeling interface and showed willingness to use it. Chia-Ming Chang 0003, Xi Yang 0017, Takeo Igarashi |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2022 | NEGraf: A System for Power System Collapse Explanation using Graph Representation and Customized PageRankabstractPower flow simulation produces a colossal amount of complex, integrated, and diverse data. An analyst can get lost in those irregular data when understanding the undergoing of the system. Developing a well-performed analysis method and corresponding visualization techniques is essential. This study proposed NEGraf, a system prototype for transforming data into useful information to assist decision-making in power system monitoring. It can explain the system operating status to support the analyst in recognizing the specific phenomenon in collapse before the blackout, an abnormal period. NEGraf contains a node-edge graph module for data representation, a customized edge-weighted PageRank inference algorithm for data analysis, and a colored graph explanatory interface for information display. We ran a user study for a blackout recognition task. The results show that our interface can better explain intuitively dynamic features in collapse and improve accuracy and recall in predicting blackout than a traditional bar visualization interface. Xinyue Gui, Chia-Ming Chang 0003, Takeo Igarashi |
CW | 2 |
| 2022 | A Drawing Support System for Sketching Aging Anime FacesabstractDrawing the facial features of anime characters at different ages is challenging in the creation process. Since characters’ facial features at different ages have obvious differences, it is difficult, especially for novices to accurately illustrate the age features of anime characters. Conventional data-driven drawing interfaces for anime characters focus on the visual features of hair, emotion, and coloring but fail to provide age-specific drawing guidance. To address this gap, we propose AgeFace, an interactive drawing interface that assists users in creating facial features with age-specific features based on user input strokes. We evaluated the usability of AgeFace by a user experience experiment and a comparison experiment with baseline approaches. The results verified that AgeFace could achieve better performance in usability and better support in the creative process than baseline systems. Sicheng Li 0002, Haoran Xie 0002, Xi Yang 0017, Chia-Ming Chang 0003, Kazunori Miyata |
CW | 4 |
| 2022 | Interactive Drawing Interface for Editing Scene GraphabstractScene graphs have been widely used in various visual applications, such as image retrieval and generation. However, it is time consuming and challenging to draw scene graphs, especially for images with complex scenes. To resolve this issue, we propose an interactive editing interface for scene graph representation in which users can draw scene graphs simply and conveniently. We provide alternatives for frequently used objects, attributes, and relationships in graph drawing. The proposed function design can greatly reduce the drawing time cost and improve drawing quality. We conducted a user study and confirmed that the proposed interface can help users draw the desired scene graph, as well as providing a good user experience. Xusheng Du, Chia-Ming Chang 0003, Xi Yang 0017, Haoran Xie 0002 |
CW | 3 |
| 2022 | An Empirical Study on the Effect of Quick and Careful Labeling Styles in Image Annotation
Chia-Ming Chang 0003, Xi Yang 0017, Takeo Igarashi |
Graphics Interface | 1 |
| 2022 | DualLabel: Secondary Labels for Challenging Image Annotation
Chia-Ming Chang 0003, Xi Yang 0017, Haoran Xie 0002, Takeo Igarashi |
Graphics Interface | 1 |
| 2022 | Fine-tuning Deep Neural Networks by Interactively Refining the 2D Latent Space of Ambiguous ImagesabstractDeep neural networks (DNNs) have achieved excellent results currently in classification, while they may still suffer from ambiguous images which are similar across classes. By contrast, humans have a relatively good ability to distinguish these categories of images. Therefore, we propose a human-in-the-loop solution to assist the network to better classify the images by leveraging human knowledge. To achieve this, we project the high-dimensional latent space trained by the network onto a two-dimensional workspace. The users can interactively modify the projected coordinates of inputs on the workspace using our designed tools, then the modified information will be fed back to the network to fine-tune it, which in turn affects the network's classification results, thereby improving the accuracy of network classification. Jiafu Wei, Haoran Xie 0002, Chia-Ming Chang 0003, Xi Yang 0017 |
IJCAI | 3 |
| 2021 | Spatial Labeling: Leveraging Spatial Layout for Improving Label Quality in Non-Expert Image AnnotationabstractNon-expert annotators (who lack sufficient domain knowledge) are often recruited for manual image labeling tasks owing to the lack of expert annotators. In such a case, label quality may be relatively low. We propose leveraging the spatial layout for improving label quality in non-expert image annotation. In the proposed system, an annotator first spatially lays out the incoming images and labels them on an open space, placing related items together. This serves as a working space (spatial organization) for tentative labeling. During the process, the annotator observes and organizes the similarities and differences between the items. Finally, the annotator provides definitive labels to the images based on the results of the spatial layout. We ran a user study comparing the proposed method and a traditional non-spatial layout in an image labeling task. The results demonstrated that annotators can complete the labeling tasks more accurately using the spatial layout interface than the non-spatial layout interface. Chia-Ming Chang 0003, Chia-Hsien Lee, Takeo Igarashi |
CHI | 1 |
| 2019 | A Hierarchical Task Assignment for Manual Image LabelingabstractManual image labeling (selecting an appropriate “category” for an image) is very tedious and time consuming especially when selecting labels from a large number of categories. In this study, we propose a hierarchical assignment of labeling tasks where the labelers recursively classify images in a category group into sub category groups, working on a single level at a time. This significantly makes each labeler's task easier, reducing the number of choices from 1,000 to 27 on average. In the user study, we compared our hierarchical assignment to a normal (non-hierarchical) assignment for a labeling task. The results show that the hierarchical assignment requires less total time to complete the labeling task. In addition, the learning effect in the labeling process is more profound in the hierarchical assignment. Chia-Ming Chang 0003, Siddharth Deepak Mishra, Takeo Igarashi |
VL/HCC | 1 |
| 2017 | Eyes on a Car: an Interface Design for Communication between an Autonomous Car and a PedestrianabstractSelf-driving technologies have been increasingly developed and tested in recent years (e.g., Volvo's and Google's self-driving cars). However, only a limited number of investigations have so far been conducted into communication between self-driving cars and pedestrians. For example, when a pedestrian is about to cross a street, that pedestrian needs to know the intension of the approaching self-driving car. In the present study, we designed a novel interface known as "Eyes on a Car" to address this problem. We added eyes onto a car so as to establish eye contact communication between that car and pedestrians. The car looks at the pedestrian in order to indicate its intention to stop. This novel interface design was evaluated via a virtual reality (VR) simulated environment featuring a street-crossing scenario. The evaluation results show that pedestrians can make the correct street-crossing decision more quickly if the approaching car has the novel interface "eyes" than in the case of normal cars. In addition, the results show that pedestrians feel safer with regard to crossing a street if the approaching car has eyes and if the eyes look at them. Chia-Ming Chang 0003, Koki Toda, Daisuke Sakamoto, Takeo Igarashi |
AutomotiveUI | 1 |