Yu Zhang 0043

dblp:50/671-43 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0002-9035-0463ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 DALL: Data Labeling via Data Programming and Active Learning Enhanced by Large Language Models
abstract
Deep learning models for natural language processing rely heavily on high-quality labeled datasets. However, existing labeling approaches often struggle to balance label quality with labeling cost. To address this challenge, we propose DALL, a text labeling framework that integrates data programming, active learning, and large language models. DALL introduces a structured specification that allows users and large language models to define labeling functions via configuration, rather than code. Active learning identifies informative instances for review, and the large language model analyzes these instances to help users correct labels and to refine or suggest labeling functions. We implement DALL as an interactive labeling system for text labeling tasks. Comparative, ablation, and usability studies demonstrate DALL’s efficiency, the effectiveness of its modules, and its usability.
Guozheng Li 0002, Shaoxiang Wang, Yu Zhang 0043, Pengcheng Cao, Chi Harold Liu
CHI4
2026 How Historians Use Visualization: A Corpus-Based Taxonomy and Mixed-Methods Analysis
abstract
Abstract Visualization in historical research is shifting from isolated attempts to systematic practices. However, data‐driven evidence about how historians actually use visualization remains scarce. We present a corpus‐driven, mixed‐methods study that combines analysis of images from 4,142 research articles across history and digital humanities journals with a collaboratively developed visualization taxonomy and a semi‐automatic labeling pipeline. We construct a corpus of 14,021 images, classify 4,831 visualization instances using a hierarchical, domain‐informed taxonomy, and analyze patterns of visualization adoption across venues, history subfields, and time. To interpret these patterns, we conduct interviews with 11 historians and use Hi‐FigAtlas system as a boundary object to support joint inspection of the corpus. We identify distinct roles for visualizations in historical research: primary‐source, evidence‐synthesis, communicative, confirmative, and exploratory. We further find that while historians pursue diverse goals with figures, persistent epistemological and practical barriers, such as uncertainty, provenance, justification burden, and publication constraints, impede the adoption of visualization. This work contributes a grounded account of visualization use in historical scholarship and points to opportunities to better support domain‐specific needs.
Xinyue Chen 0003, Yu Zhang 0043, Weili Zheng, Chiteng Ma, Xiaoru Yuan
Comput. Graph. Forum2
2026 Calli-VA: A Visual Analytics System for Analyzing and Comparing Chinese Calligraphic Styles
abstract
Chinese calligraphy is a quintessential element of Chinese cultural heritage. Analyzing and comparing calligraphic styles not only enhances the appreciation, learning, and advancement of calligraphy but also provides valuable insights into ancient China. However, such analysis remains challenging due to the limited scalability and possible inconsistencies of qualitative methods, as well as usability and misalignment issues in conventional quantitative approaches. We propose Calli-VA, a visual analytics system, to address these challenges. Calli-VA extracts character images and their corresponding strokes from original works and characterizes each character using systematic criteria. During analysis, the system defines the analysis scope by overview and uncovers relationships between characters. Explanation and recommendation mechanisms are integrated to help users understand patterns and guide further exploration. A documentation feature allows users to record and share their findings. We demonstrate the effectiveness of Calli-VA through three case studies and expert feedback.
Jincheng Li 0004, Jinpeng Wu, Shaocong Tan, Lin Du 0011, Yu Zhang 0043, Chaofan Yang, Jiadi Zhang, Rebecca Ruige Xu, Lu Bai 0001, Xiaoru Yuan
IEEE Trans. Vis. Comput. Graph.5
2026 DKMap: Interactive Exploration of Vision-Language Alignment in Multimodal Embeddings via Dynamic Kernel Enhanced Projection
abstract
Examining vision-language alignment in multimodal embeddings is crucial for various tasks, such as evaluating generative models and filtering pretraining data. The intricate nature of high-dimensional features necessitates dimensionality reduction (DR) methods to explore alignment of multimodal embeddings. However, existing DR methods fail to account for cross-modal alignment metrics, resulting in severe occlusion of points with divergent metrics clustered together, inaccurate contour maps from over-aggregation, and insufficient support for multi-scale exploration. To address these problems, this paper introduces DKMap, a novel DR visualization technique for interactive exploration of multimodal embeddings through Dynamic Kernel enhanced projection. First, rather than performing dimensionality reduction and contour estimation sequentially, we introduce a kernel regression supervised t-SNE that directly integrates post-projection contour mapping into the projection learning process, ensuring cross-modal alignment mapping accuracy. Second, to enable multi-scale exploration with dynamic zooming and progressively enhanced local detail, we integrate validation-constrained a refinement of a generalized t-kernel with quad-tree-based multi-resolution technique, ensuring reliable kernel parameter tuning without overfitting. DKMap is implemented as a multi-platform visualization tool, featuring a web-based system for interactive exploration and a Python package for computational notebook analysis. Quantitative comparisons with baseline DR techniques demonstrate DKMap's superiority in accurately mapping cross-modal alignment metrics. We further demonstrate generalizability and scalability of DKMap with three usage scenarios, including visualizing million-scale text-to-image corpus, comparatively evaluating generative models, and exploring a billion-scale pretraining dataset.
Chenxi Ruan, Yu Zhang 0043, Zikun Deng, Wei Zeng 0004
IEEE Trans. Vis. Comput. Graph.3
2026 ProactiveVA: Proactive Visual Analytics with LLM-Based UI Agent
abstract
Visual analytics (VA) is typically applied to complex data, thus requiring complex tools. While visual analytics empowers analysts in data analysis, analysts may get lost in the complexity occasionally. This highlights the need for intelligent assistance mechanisms. However, even the latest LLM-assisted VA systems only provide help when explicitly requested by the user, making them insufficiently intelligent to offer suggestions when analysts need them the most. We propose a ProactiveVA framework in which LLM-powered UI agent monitors user interactions and delivers context-aware assistance proactively. To design effective proactive assistance, we first conducted a formative study analyzing help-seeking behaviors in user interaction logs, identifying when users need proactive help, what assistance they require, and how the agent should intervene. Based on this analysis, we distilled key design requirements in terms of intent recognition, solution generation, interpretability and controllability. Guided by these requirements, we develop a three-stage UI agent pipeline including perception, reasoning, and acting. The agent autonomously perceives users' needs from VA interaction logs, providing tailored suggestions and intuitive guidance through interactive exploration of the system. We implemented the framework in two representative types of VA systems, demonstrating its generalizability, and evaluated the effectiveness through an algorithm evaluation, case and expert study and a user study. We also discuss current design trade-offs of proactive VA and areas for further exploration.
Yuheng Zhao, Xueli Shu, Liwen Fan, Yu Zhang 0043, Siming Chen 0001
IEEE Trans. Vis. Comput. Graph.5
2025 ZuantuSet: A Collection of Historical Chinese Visualizations and Illustrations
abstract
Historical visualizations are a valuable resource for studying the history of visualization and inspecting the cultural context where they were created. When investigating historical visualizations, it is essential to consider contributions from different cultural frameworks to gain a comprehensive understanding. While there is extensive research on historical visualizations within the European cultural framework, this work shifts the focus to ancient China, a cultural context that remains underexplored by visualization researchers. To this aim, we propose a semi-automatic pipeline to collect, extract, and label historical Chinese visualizations. Through the pipeline, we curate ZuantuSet, a dataset with over 71K visualizations and 108K illustrations. We analyze distinctive design patterns of historical Chinese visualizations and their potential causes within the context of Chinese history and culture. We illustrate potential usage scenarios for this dataset, summarize the unique challenges and solutions associated with collecting historical Chinese visualizations, and outline future research directions.
Xiyao Mei, Yu Zhang 0043, Chaofan Yang, Xiaoru Yuan
CHI2
2025 InReAcTable: LLM-powered Interactive Visual Data Story Construction from Tabular Data
abstract
Insights in tabular data capture valuable patterns that help analysts understand critical information.Organizing related insights into visual data stories is crucial for in-depth analysis.However, constructing such stories is challenging because of the complexity of the inherent relations between extracted insights.Users face difficulty sifting through a vast number of discrete insights to integrate specific ones into a unified narrative that meets their analytical goals.Existing methods either heavily rely on user expertise, making the process inefficient, or employ automated approaches that cannot fully capture their evolving goals.In this paper, we introduce InRe-AcTable, a framework that enhances visual data story construction by establishing both structural and semantic connections between data insights.Each user interaction triggers the Acting module, which utilizes an insight graph for structural filtering to narrow the search space, followed by the Reasoning module using the retrievalaugmented generation method based on large language models for semantic filtering, ultimately providing insight recommendations aligned with the user's analytical intent.Based on the InReAcTable framework, we develop an interactive prototype system that guides users to construct visual data stories aligned with their analytical requirements.We conducted a case study and a user experiment to demonstrate the utility and effectiveness of the InReAcTable framework and the prototype system for interactively building visual data stories.
Gerile Aodeng, Guozheng Li 0002, Yunshan Feng, Yu Zhang 0043, Chi Harold Liu
UIST5
2025 BiaSeer: A Visual Analytics System for Identifying and Understanding Media Bias
abstract
Media bias refers to bias in news reporting and coverage that exists pervasively. By identifying media bias, social scientists can understand the different perspectives held by media outlets in news reporting. Existing studies only focus on the analysis of media bias of isolated incidents, but neglect their sustained characteristics. Thus, they cannot provide a comprehensive understanding of specific news topics. We develop BiaSeer, a visual analytics system for identifying and understanding sustained bias of media outlets. BiaSeer employs an overview-to-detail approach for interactive identification of media bias. The overview assists users in determining the analysis scope of media outlets. In addition, it visualizes the variances in coverage patterns between selected media outlets using a matrix visualization to facilitate the identification of biased news articles. BiaSeer visualizes the sustained bias in the context of the evolution of events. It first summarizes news articles into events based on a keyword co-occurrence graph and then connects events into a narrative structure using a path-aware story tree construction method. In addition, BiaSeer integrates a sustained bias computation algorithm and enables analysts to compare the narrative structures of different media outlets using the juxtaposition-based visualization approach. We conducted a user experiment to validate the effectiveness of BiaSeer in helping social scientists understand news topics and the usability of visualization designs. To examine the effectiveness of BiaSeer, we conducted a case study with social scientists on the topics of the Russia-Ukraine conflict. The results demonstrate the utility and usability of BiaSeer in efficiently analyzing media bias and attaining a well-rounded understanding of news topics.
Guozheng Li 0002, Shiyu Han, Jihe Wu, Jiale Hu, Yu Zhang 0043, Chi Harold Liu
Proc. ACM Hum. Comput. Interact.5
2025 Viewpoint Recommendation for Point Cloud Labeling Through Interaction Cost Modeling
abstract
Semantic segmentation of 3D point clouds is important for many applications, such as autonomous driving. To train semantic segmentation models, labeled point cloud segmentation datasets are essential. Meanwhile, point cloud labeling is time-consuming for annotators, which typically involves tuning the camera viewpoint and selecting points by lasso. To reduce the time cost of point cloud labeling, we propose a viewpoint recommendation approach to reduce annotators' labeling time costs. We adapt Fitts' law to model the time cost of lasso selection in point clouds. Using the modeled time cost, the viewpoint that minimizes the lasso selection time cost is recommended to the annotator. We build a data labeling system for semantic segmentation of 3D point clouds that integrates our viewpoint recommendation approach. The system enables users to navigate to recommended viewpoints for efficient annotation. Through an ablation study, we observed that our approach effectively reduced the data labeling time cost. We also qualitatively compare our approach with previous viewpoint selection approaches on different datasets.
Yu Zhang 0043, Chongke Bi, Siming Chen 0001
IEEE Trans. Vis. Comput. Graph.1
2025 VisTaxa: Developing a Taxonomy of Historical Visualizations
abstract
Historical visualizations are a rich resource for visualization research. While taxonomy is commonly used to structure and understand the design space of visualizations, existing taxonomies primarily focus on contemporary visualizations and largely overlook historical visualizations. To address this gap, we describe an empirical method for taxonomy development. We introduce a coding protocol and the VisTaxa system for taxonomy labeling and comparison. We demonstrate using our method to develop a historical visualization taxonomy by coding 400 images of historical visualizations. We analyze the coding result and reflect on the coding process. Our work is an initial step toward a systematic investigation of the design space of historical visualizations.
Yu Zhang 0043, Xinyue Chen 0003, Weili Zheng, Yuhan Guo 0004, Guozheng Li 0002, Siming Chen 0001, Xiaoru Yuan
IEEE Trans. Vis. Comput. Graph.1
2025 LightVA: Lightweight Visual Analytics With LLM Agent-Based Task Planning and Execution
abstract
Visual analytics (VA) requires analysts to iteratively propose analysis tasks based on observations and execute tasks by creating visualizations and interactive exploration to gain insights. This process demands skills in programming, data processing, and visualization tools, highlighting the need for a more intelligent, streamlined VA approach. Large language models (LLMs) have recently been developed as agents to handle various tasks with dynamic planning and tool-using capabilities, offering the potential to enhance the efficiency and versatility of VA. We propose LightVA, a lightweight VA framework that supports task decomposition, data analysis, and interactive exploration through human-agent collaboration. Our method is designed to help users progressively translate high-level analytical goals into low-level tasks, producing visualizations and deriving insights. Specifically, we introduce an LLM agent-based task planning and execution strategy, employing a recursive process involving a planner, executor, and controller. The planner is responsible for recommending and decomposing tasks, the executor handles task execution, including data analysis, visualization generation and multi-view composition, and the controller coordinates the interaction between the planner and executor. Building on the framework, we develop a system with a hybrid user interface that includes a task flow diagram for monitoring and managing the task planning process, a visualization panel for interactive data exploration, and a chat view for guiding the model through natural language instructions. We examine the effectiveness of our method through a usage scenario and an expert study.
Yuheng Zhao, Linbing Xiang, Zifei Guo, Cagatay Turkay, Yu Zhang 0043, Siming Chen 0001
IEEE Trans. Vis. Comput. Graph.7
2025 LEVA: Using Large Language Models to Enhance Visual Analytics
abstract
Visual analytics supports data analysis tasks within complex domain problems. However, due to the richness of data types, visual designs, and interaction designs, users need to recall and process a significant amount of information when they visually analyze data. These challenges emphasize the need for more intelligent visual analytics methods. Large language models have demonstrated the ability to interpret various forms of textual data, offering the potential to facilitate intelligent support for visual analytics. We propose LEVA, a framework that uses large language models to enhance users' VA workflows at multiple stages: onboarding, exploration, and summarization. To support onboarding, we use large language models to interpret visualization designs and view relationships based on system specifications. For exploration, we use large language models to recommend insights based on the analysis of system status and data to facilitate mixed-initiative exploration. For summarization, we present a selective reporting strategy to retrace analysis history through a stream visualization and generate insight reports with the help of large language models. We demonstrate how LEVA can be integrated into existing visual analytics systems. Two usage scenarios and a user study suggest that LEVA effectively aids users in conducting visual analytics.
Yuheng Zhao, Yu Zhang 0043, Zekai Shao 0001, Cagatay Turkay, Siming Chen 0001
IEEE Trans. Vis. Comput. Graph.3
2024 A Spatial Constraint Model for Manipulating Static Visualizations
abstract
We introduce a spatial constraint model to characterize the positioning and interactions in visualizations, thereby facilitating the activation of static visualizations. Our model provides users with the capability to manipulate visualizations through operations such as selection, filtering, navigation, arrangement, and aggregation. Building upon this conceptual framework, we propose a prototype system designed to activate pre-existing visualizations by imbuing them with intelligent interactions. This augmentation is accomplished through the integration of visual objects with forces. The instantiation of our spatial constraint model enables seamless animated transitions between distinct visualization layouts. To demonstrate the efficacy of our approach, we present usage scenarios that involve the activation of visualizations within real-world contexts.
Can Liu 0004, Yu Zhang 0043, Cong Wu 0004, Chen Li 0078, Xiaoru Yuan
ACM Trans. Interact. Intell. Syst.2
2024 Simulation-based Optimization of User Interfaces for Quality-assuring Machine Learning Model Predictions
abstract
Quality-sensitive applications of machine learning (ML) require quality assurance (QA) by humans before the predictions of an ML model can be deployed. QA for ML (QA4ML) interfaces require users to view a large amount of data and perform many interactions to correct errors made by the ML model. An optimized user interface (UI) can significantly reduce interaction costs. While UI optimization can be informed by user studies evaluating design options, this approach is not scalable, because there are typically numerous small variations that can affect the efficiency of a QA4ML interface. Hence, we propose using simulation to evaluate and aid the optimization of QA4ML interfaces. In particular, we focus on simulating the combined effects of human intelligence in initiating appropriate interaction commands and machine intelligence in providing algorithmic assistance for accelerating QA4ML processes. As QA4ML is usually labor-intensive, we use the simulated task completion time as the metric for UI optimization under different interface and algorithm setups. We demonstrate the usage of this UI design method in several QA4ML applications.
Yu Zhang 0043, Martijn Tennekes, Tim J. A. de Jong, R. Lyana Curier, Bob Coecke, Min Chen 0001
ACM Trans. Interact. Intell. Syst.1
2024 CoInsight: Visual Storytelling for Hierarchical Tables With Connected Insights
abstract
Extracting data insights and generating visual data stories from tabular data are critical parts of data analysis. However, most existing studies primarily focus on tabular data stored as flat tables, typically without leveraging the relations between cells in the headers of hierarchical tables. When properly used, rich table headers can enable the extraction of many additional data stories. To assist analysts in visual data storytelling, an approach is needed to organize these data insights efficiently. In this work, we propose CoInsight, a system to facilitate visual storytelling for hierarchical tables by connecting insights. CoInsight extracts data insights from hierarchical tables and builds insight relations according to the structure of table headers. It further visualizes related data insights using a nested graph with edge bundling. We evaluate the CoInsight system through a usage scenario and a user experiment. The results demonstrate the utility and usability of CoInsight for converting data insights in hierarchical tables into visual data stories.
Guozheng Li 0002, Runfei Li, Yunshan Feng, Yu Zhang 0043, Yuyu Luo, Chi Harold Liu
IEEE Trans. Vis. Comput. Graph.4
2024 OldVisOnline: Curating a Dataset of Historical Visualizations
abstract
With the increasing adoption of digitization, more and more historical visualizations created hundreds of years ago are accessible in digital libraries online. It provides a unique opportunity for visualization and history research. Meanwhile, there is no large-scale digital collection dedicated to historical visualizations. The visualizations are scattered in various collections, which hinders retrieval. In this study, we curate the first large-scale dataset dedicated to historical visualizations. Our dataset comprises 13K historical visualization images with corresponding processed metadata from seven digital libraries. In curating the dataset, we propose a workflow to scrape and process heterogeneous metadata. We develop a semi-automatic labeling approach to distinguish visualizations from other artifacts. Our dataset can be accessed with OldVisOnline, a system we have built to browse and label historical visualizations. We discuss our vision of usage scenarios and research opportunities with our dataset, such as textual criticism for historical visualizations. Drawing upon our experience, we summarize recommendations for future efforts to improve our dataset.
Yu Zhang 0043, Ruike Jiang, Liwenhan Xie, Yuheng Zhao, Can Liu 0004, Tianhong Ding, Siming Chen 0001, Xiaoru Yuan
IEEE Trans. Vis. Comput. Graph.1
2022 OneLabeler: A Flexible System for Building Data Labeling Tools
abstract
Labeled datasets are essential for supervised machine learning. Various data labeling tools have been built to collect labels in different usage scenarios. However, developing labeling tools is time-consuming, costly, and expertise-demanding on software development. In this paper, we propose a conceptual framework for data labeling and OneLabeler based on the conceptual framework to support easy building of labeling tools for diverse usage scenarios. The framework consists of common modules and states in labeling tools summarized through coding of existing tools. OneLabeler supports configuration and composition of common software modules through visual programming to build data labeling tools. A module can be a human, machine, or mixed computation procedure in data labeling. We demonstrate the expressiveness and utility of the system through ten example labeling tools built with OneLabeler. A user study with developers provides evidence that OneLabeler supports efficient building of diverse data labeling tools.
Yu Zhang 0043, Yun Wang 0012, Bin B. Zhu, Siming Chen 0001, Dongmei Zhang 0001
CHI1
2021 MI3: Machine-initiated Intelligent Interaction for Interactive Classification and Data Reconstruction
abstract
In many applications, while machine learning (ML) can be used to derive algorithmic models to aid decision processes, it is often difficult to learn a precise model when the number of similar data points is limited. One example of such applications is data reconstruction from historical visualizations, many of which encode precious data, but their numerical records are lost. On the one hand, there is not enough similar data for training an ML model. On the other hand, manual reconstruction of the data is both tedious and arduous. Hence, a desirable approach is to train an ML model dynamically using interactive classification, and hopefully, after some training, the model can complete the data reconstruction tasks with less human interference. For this approach to be effective, the number of annotated data objects used for training the ML model should be as small as possible, while the number of data objects to be reconstructed automatically should be as large as possible. In this article, we present a novel technique for the machine to initiate intelligent interactions to reduce the user’s interaction cost in interactive classification tasks. The technique of machine-initiated intelligent interaction (MI3) builds on a generic framework featuring active sampling and default labeling. To demonstrate the MI3 approach, we use the well-known cholera map visualization by John Snow as an example, as it features three instances of MI3 pipelines. The experiment has confirmed the merits of the MI3 approach.
Yu Zhang 0043, Bob Coecke, Min Chen 0001
ACM Trans. Interact. Intell. Syst.1
2020 BarcodeTree: Scalable Comparison of Multiple Hierarchies
abstract
We propose BarcodeTree (BCT), a novel visualization technique for comparing topological structures and node attribute values of multiple trees. BCT can provide an overview of one hundred shallow and stable trees simultaneously, without aggregating individual nodes. Each BCT is shown within a single row using a style similar to a barcode, allowing trees to be stacked vertically with matching nodes aligned horizontally to ease comparison and maintain space efficiency. We design several visual cues and interactive techniques to help users understand the topological structure and compare trees. In an experiment comparing two variants of BCT with icicle plots, the results suggest that BCTs make it easier to visually compare trees by reducing the vertical distance between different trees. We also present two case studies involving a dataset of hundreds of trees to demonstrate BCT's utility.
Guozheng Li 0002, Yu Zhang 0043, Yu Dong 0001, Christy Jie Liang, Jinson Zhang, Michael J. McGuffin, Xiaoru Yuan
IEEE Trans. Vis. Comput. Graph.2
2017 Interaction+: Interaction enhancement for web-based visualizations
abstract
In this work, we present Interaction+, a tool that enhances the interactive capability of existing web-based visualizations. Different from the toolkits for authoring interactions during the visualization construction, Interaction+ takes existing visualizations as input, analyzes the visual objects, and provides users with a suite of interactions to facilitate the visual exploration, including selection, aggregation, arrangement, comparison, filtering, and annotation. Without accessing the underlying data or process how the visualization is constructed, Interaction+ is application-independent and can be employed in various visualizations on the web. We demonstrate its usage in two scenarios and evaluate its effectiveness with a qualitative user study.
Min Lu 0002, Christy Jie Liang, Yu Zhang 0043, Guozheng Li 0002, Siming Chen 0001, Zongru Li, Xiaoru Yuan
PacificVis3