Huanchen Wang

dblp:42/4352 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0001-9339-1941ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 VisMoDAI: Visual Analytics for Evaluating and Improving Corruption Robustness of Vision-Language Models
abstract
Vision-language (VL) models have shown transformative potential across various critical domains due to their capability to comprehend multi-modal information. However, their performance frequently degrades under distribution shifts, making it crucial to assess and improve robustness against real-world data corruption encountered in practical applications. While advancements in VL benchmark datasets and data augmentation (DA) have contributed to robustness evaluation and improvement, there remain challenges due to a lack of in-depth comprehension of model behavior as well as the need for expertise and iterative efforts to explore data patterns. Given the achievement of visualization in explaining complex models and exploring large-scale data, understanding the impact of various data corruption on VL models aligns naturally with a visual analytics approach. To address these challenges, we introduce VisMoDAI, a visual analytics framework designed to evaluate VL model robustness against various corruption types and identify underperformed samples to guide the development of effective DA strategies. Grounded in the literature review and expert discussions, VisMoDAI supports multi-level analysis, ranging from examining performance under specific corruptions to task-driven inspection of model behavior and corresponding data slice. Unlike conventional works, VisMoDAI enables users to reason about the effects of corruption on VL models, facilitating both model behavior understanding and DA strategy formulation. The utility of our system is demonstrated through case studies and quantitative evaluations focused on corruption robustness in the image captioning task.
Huanchen Wang, Wencheng Zhang, Zhicong Lu, Yuxin Ma 0001
IEEE Trans. Vis. Comput. Graph.1
2025 HarmonyCut: Supporting Creative Chinese Paper-cutting Design with Form and Connotation Harmony
abstract
Chinese paper-cutting, an Intangible Cultural Heritage (ICH), faces challenges from the erosion of traditional culture due to the prevalence of realism alongside limited public access to cultural elements. While generative AI can enhance paper-cutting design with its extensive knowledge base and efficient production capabilities, it often struggles to align content with cultural meaning due to users' and models' lack of comprehensive paper-cutting knowledge. To address these issues, we conducted a formative study (N=7) to identify the workflow and design space, including four core factors (Function, Subject Matter, Style, and Method of Expression) and a key element (Pattern). We then developed HarmonyCut, a generative AI-based tool that translates abstract intentions into creative and structured ideas. This tool facilitates the exploration of suggested related content (knowledge, works, and patterns), enabling users to select, combine, and adjust elements for creative paper-cutting design. A user study (N=16) and an expert evaluation (N=3) demonstrated that HarmonyCut effectively provided relevant knowledge, aiding the ideation of diverse paper-cutting designs and maintaining design quality within the design space to ensure alignment between form and cultural connotation.
Huanchen Wang, Tianrun Qiu, Jiaping Li, Zhicong Lu, Yuxin Ma 0001
CHI1
2025 RAGTrace: Understanding and Refining Retrieval-Generation Dynamics in Retrieval-Augmented Generation
abstract
Retrieval-Augmented Generation (RAG) systems have emerged as a promising solution to enhance large language models (LLMs) by integrating external knowledge retrieval with generative capabilities. While significant advancements have been made in improving retrieval accuracy and response quality, a critical challenge remains that the internal knowledge integration and retrieval-generation interactions in RAG workflows are largely opaque. This paper introduces RAGTrace, an interactive evaluation system designed to analyze retrieval and generation dynamics in RAG-based workflows. Informed by a comprehensive literature review and expert interviews, the system supports a multi-level analysis approach, ranging from high-level performance evaluation to fine-grained examination of retrieval relevance, generation fidelity, and cross-component interactions. Unlike conventional evaluation practices that focus on isolated retrieval or generation quality assessments, RAGTrace enables an integrated exploration of retrieval-generation relationships, allowing users to trace knowledge sources and identify potential failure cases. The system's workflow allows users to build, evaluate, and iterate on retrieval processes tailored to their specific domains of interest. The effectiveness of the system is demonstrated through case studies and expert evaluations on real-world RAG applications.
Sizhe Cheng, Jiaping Li, Huanchen Wang, Yuxin Ma 0001
UIST3
2025 Polymind: Parallel Visual Diagramming with Large Language Models to Support Prewriting Through Microtasks
abstract
Prewriting is the process of generating and organising ideas before a first draft. It consists of a combination of informal, iterative, and semi-structured strategies such as visual diagramming, which poses a challenge for collaborating with large language models (LLMs) in a turn-taking conversational manner. We present Polymind, a visual diagramming tool that leverages multiple LLM-powered agents to support prewriting. The system features a parallel collaboration workflow in place of the turn-taking conversational interactions. It defines multiple ''microtasks'' to simulate group collaboration scenarios such as collaborative writing and group brainstorming. Instead of repetitively prompting a chatbot for various purposes, Polymind enables users to orchestrate multiple microtasks simultaneously. Users can configure and delegate customised microtasks, and manage their microtasks by specifying task requirements and toggling visibility and initiative. Our evaluation revealed that, compared to ChatGPT, users had more customizability over collaboration with Polymind, and were thus able to quickly expand personalised writing ideas during prewriting.
Qian Wan 0004, Jiannan Li, Huanchen Wang, Zhicong Lu
Proc. ACM Hum. Comput. Interact.3
2025 DanModCap: Designing a Danmaku Moderation Tool for Video-Sharing Platforms that Leverages Impact Captions with Large Language Models
abstract
Online video platforms have gained increased popularity due to their ability to support information consumption and sharing and the diverse social interactions they afford. Danmaku, a real-time commentary feature that overlays user comments on a video, has been found to improve user engagement, however, the use of Danmaku can lead to toxic behaviors and inappropriate comments. To address these issues, we propose a proactive moderation approach inspired by Impact Captions, a visual technique used in East Asian variety shows. Impact Captions combine textual content and visual elements to construct emotional and cognitive resonance. Within the context of this work, Impact Captions were used to guide viewers towards positive Danmaku-related activities and elicit more pro-social behaviors. Leveraging Impact Captions, we developed DanModCap, an moderation tool that collected and analyzed Danmaku and used it as input to large generative language models to produce Impact Captions. Our evaluation of DanModCap demonstrated that Impact Captions reduced negative antagonistic emotions, increased users' desire to share positive content, and elicited self-control in Danmaku social action to fostering proactive community maintenance behaviors. Our approach highlights the benefits of using LLM-supported content moderation methods for proactive moderation in a large-scale live content contexts.
Siying Hu, Huanchen Wang, Yu Zhang 0097, Piaohong Wang, Zhicong Lu
Proc. ACM Hum. Comput. Interact.2
2024 Critical Heritage Studies as a Lens to Understand Short Video Sharing of Intangible Cultural Heritage on Douyin
abstract
Intangible Cultural Heritage (ICH) faces numerous threats that can lead to its destruction. While the emergence of short video platforms provides opportunities for fostering innovation and communication among ICH practitioners and viewers, it is still understudied how different stakeholders present, explain, and manage ICH via short videos. To address this, we conduct a mixed-method study of ICH-related videos on Douyin, a popular short video platform in China with an extensive user base and wealth of ICH content. By adopting the Critical Heritage Studies (CHS) framework, we propose a taxonomy of frames that construct the landscape of ICH short videos and then investigate the interactions among different groups regarding power, identity, and knowledge. Additionally, we analyze viewer responses to different frames and groups based on audience metrics (e.g., # of likes and comments) and comments. Our research reveals that government-affiliated and indigenous groups dominate the promotion and presentation of ICH on Douyin. Contrary to previous literature, viewer responses show a preference for videos from external ICH groups and ordinary individuals, suggesting a tendency to counter authority and exclusivity associated with ICH. Moreover, it highlights a lack of sustainable debates and negotiations among different groups involved in ICH discourse. Situated within CHS, we provide design implications for ICH safeguarding and sustainability through short videos and online media.
Huanchen Wang, Minzhu Zhao, Wanyang Hu, Yuxin Ma 0001, Zhicong Lu
CHI1
2023 Understanding Communication Strategies and Viewer Engagement with Science Knowledge Videos on Bilibili
abstract
As a popular form of online media, videos have been widely used to communicate scientific knowledge on video-sharing platforms. These science knowledge videos take advantage of rich and multi-modality information which has the potential to provoke public engagement with science knowledge and promote self-learning. However, how communicators strategically make science knowledge videos to engage viewers, and how specific communication strategies correlate with viewer engagement remain under-explored. In this paper, we first established a taxonomy of communication strategies currently used in science knowledge videos on Bilibili and then examined the correlations between communication strategies and viewers’ behavioral, emotional, and cognitive engagements measured by post-video comments. Our findings revealed the landscape of rich science communication strategies in science knowledge videos and further uncovered the correlations between these strategies and viewer engagements. We situated our results within prior research on science communication and HCI, and provided design implications for video-sharing platforms to support effective science communication.
Yu Zhang 0097, Changyang He, Huanchen Wang, Zhicong Lu
CHI3
2022 Learning Latent Road Correlations from Trajectories
abstract
A core component of the Intelligent Transportation System (ITS) is road network, which forms the most basic transport infrastructure, and becomes widely applied in many traffic applications. In most traffic models, the spatial representation of road network is learned only through static graph connection while dynamic driver preference and traffic conditions in the real world are ignored. Therefore, in this paper, a novel trajectory-based road network representation is proposed. By mining vehicle trajectories, our proposed method can learn dynamic route choice through embeddings of each road in a next-hop prediction model. Then road correlations are calculated by the embeddings to build a latent correlation graph that can be applied in various traffic-related applications. Extensive experiment results prove the effectiveness and rationality of our proposed approach.
Zheng Dong 0006, Quanjun Chen, Renhe Jiang, Huanchen Wang, Xuan Song 0001
IEEE Big Data4
2022 A Geomagnetic Sensor Dataset for Traffic Flow Prediction
abstract
Traffic state prediction is essential in Intelligent Transportation Systems for surveillance, management, and daily commuting. For developing high-accuracy prediction models, real-world traffic state datasets are necessary for training model parameters and evaluating prediction results. However, limited by the existing traffic collection devices, most of the current open datasets for traffic state prediction cannot obtain accurate traffic flow information. In contrast, some datasets directly use detection devices in freeway systems, so they cannot reflect complex urban traffic states. Therefore, a dataset from advanced devices that can record the flow from point to point on an urban road network attracts more attention and drives the progress of research on traffic state prediction models. To deal with the above issues, we introduce a Suburban Traffic Flow dataset using Geomagnetic sensors, or STF-G dataset, constructed for traffic flow prediction. The STF-G dataset consists of 2.5 billion vehicle driving scenarios and 319 corresponding geomagnetic sensors. The data was collected over 20 months and processed with two regional road graphs. We also do the Benchmark experiments in STF-G for analyzing and evaluating the performance of graph neural network models in traffic flow prediction and compare them to the other datasets with the same baseline.
Huanchen Wang, Quanjun Chen, Zheng Dong 0006, Xuan Song 0001, Donglong Yang, Manxia Liu
IEEE Big Data1