Xiwen Cai

dblp:207/1244 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
5 papers
Visualization and visual analytics · 98% Multimedia analysis and retrieval · 2%
Human-computer interaction and pervasive computing
3 papers
User interface design and tools · 59% Human-AI interaction · 34% Usability and user experience research · 8%
Databases, data mining, and information retrieval
5 papers
Data integration and cleaning · 73% Query processing and optimization · 14% Knowledge graphs · 12%

Topics — the 10 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning › data quality
data validation
1.012026
RuleScope: Semantic-Aware Authoring of Data Validation Rules · IEEE Trans. Vis. Comput. Graph. 2026
Visualization and visual analytics
interactive data analysis
1.012026
RuleScope: Semantic-Aware Authoring of Data Validation Rules · IEEE Trans. Vis. Comput. Graph. 2026
User interface design and tools › programming environments
computational notebooks
0.912025
Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling Scripts · CHI 2025
Visualization and visual analytics › temporal data visualization
event sequence visualization
0.812024
A Comparative Study on Fixed-Order Event Sequence Visualizations: Gantt, Extended Gantt, and Stringline Charts · IEEE Trans. Vis. Comput. Graph. 2024
Visualization and visual analytics › temporal data visualization
gantt chart
0.812024
A Comparative Study on Fixed-Order Event Sequence Visualizations: Gantt, Extended Gantt, and Stringline Charts · IEEE Trans. Vis. Comput. Graph. 2024
Visualization and visual analytics › data visualization › image visualization
image collection visualization
0.412019
A Semantic-Based Method for Visualizing Large Image Collections · IEEE Trans. Vis. Comput. Graph. 2019
Visualization and visual analytics
visual analytics
0.412019
A Semantic-Based Method for Visualizing Large Image Collections · IEEE Trans. Vis. Comput. Graph. 2019
Data integration and cleaning
data wrangling
0.312025
Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling Scripts · CHI 2025
Usability and user experience research › experimental design
controlled experiment
0.212024
A Comparative Study on Fixed-Order Event Sequence Visualizations: Gantt, Extended Gantt, and Stringline Charts · IEEE Trans. Vis. Comput. Graph. 2024
Multimedia analysis and retrieval › multimedia analysis › multimedia content description
image captioning
0.112019
A Semantic-Based Method for Visualizing Large Image Collections · IEEE Trans. Vis. Comput. Graph. 2019

Methods — techniques the papers use, named apart from their topics

user study · 3.5matrix-based visualization · 3.0LLM · 3.0large language model · 2.0interactive tree view · 2.0mixed-initiative interaction · 1.7knowledge graph construction · 1.7key phrase extraction · 1.7data-aware code suggestion · 1.7code parsing · 1.7task error rate · 1.5in-lab experiment · 1.5completion time · 1.5
YearPublicationVenuePosition
2026 Cerebra: Aligning Implicit Knowledge in Interactive SQL Authoring
abstract
LLM-driven tools have significantly lowered barriers to writing SQL queries. However, user instructions are often underspecified, assuming the model understands implicit knowledge, such as dataset schemas, domain conventions, and task-specific requirements, that isn’t explicitly provided. This results in frequently erroneous scripts that require users to repeatedly clarify their intent. Additionally, users struggle to validate generated scripts because they cannot verify whether the model correctly applied implicit knowledge. We present Cerebra, an interactive NL-to-SQL tool that aligns implicit knowledge between users and LLMs during SQL authoring. Cerebra automatically retrieves implicit knowledge from historical SQL scripts based on user instructions, presents this knowledge in an interactive tree view for code review, and supports iterative refinement to improve generated scripts. To evaluate the effectiveness and usability of Cerebra, we conducted a user study with 16 participants, demonstrating its improved support for customized SQL authoring. The source code of Cerebra is available at https://github.com/zjuidg/CHI26-Cerebra.
Yunfan Zhou, Qiming Shi, Zhongsu Luo, Xiwen Cai, Yanwei Huang, Daehyun Kim 0005, Di Weng, Yingcai Wu
CHI4
2026 Topology optimization of piezoelectric structures based on multi-modal input neural network
Jianhua Xiang, Yudong Yang, Xiwen Cai, Yongfeng Zheng
Eng. Appl. Artif. Intell.3
2026 RuleScope: Semantic-Aware Authoring of Data Validation Rules
abstract
Data validation is a crucial step in data analytics workflows that assesses and ensures the reliability of data flowing into analytical processes. One common approach to data validation involves defining validation rules, which provide explicit constraints and conditions that data must satisfy. However, creating accurate and effective validation rules remains challenging for many practitioners. This challenge stems from the need for practitioners to understand both data structures and their domain-specific semantic relationships. Recent studies have proposed automated approaches to generate validation rules by deriving patterns from data properties. However, these approaches generate rules with limited interpretability and lack support for rule verification and modification, making the rules difficult to understand and adapt. To address these limitations in current validation rule authoring approaches, we present RuleScope, an interactive system for authoring data validation rules through semantic-aware rule generation, visualization, and refinement. RuleScope employs an LLM-based workflow to generate interpretable rules by analyzing data semantics and incorporating domain knowledge. To facilitate rule comprehension, we design a matrix-based visualization that helps users understand rules and analyze validation results. Additionally, RuleScope enables users to interactively refine rules. We evaluate the LLM-based workflow through model evaluation on datasets from different domains and assess RuleScope's usability and effectiveness through two case studies and a user study.
Zhongsu Luo, Di Weng, Xiwen Cai, Xinhuan Shu, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.5
2025 Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling Scripts
abstract
Data analysts frequently employ code completion tools in writing custom scripts to tackle complex tabular data wrangling tasks. However, existing tools do not sufficiently link the data contexts such as schemas and values with the code being edited. This not only leads to poor code suggestions, but also frequent interruptions in coding processes as users need additional code to locate and understand relevant data. We introduce Xavier, a tool designed to enhance data wrangling script authoring in computational notebooks. Xavier maintains users' awareness of data contexts while providing data-aware code suggestions. It automatically highlights the most relevant data based on the user's code, integrates both code and data contexts for more accurate suggestions, and instantly previews data transformation results for easy verification. To evaluate the effectiveness and usability of Xavier, we conducted a user study with 16 data analysts, showing its potential to streamline data wrangling scripts authoring.
Yunfan Zhou, Xiwen Cai, Qiming Shi, Yanwei Huang, Haotian Li 0001, Huamin Qu, Di Weng, Yingcai Wu
CHI2
2025 HYPNOS: Interactive Data Lineage Tracing for Data Transformation Scripts
abstract
In a formal data analysis workflow, data validation is a necessary step that helps data analysts verify the quality of the data and ensure the reliability of the results. Data analysts usually need to validate the result when encountering an unexpected result, such as an abnormal record in a table. In order to understand how a specific record is derived, they would backtrace it in the pipeline step by step via checking the code lines, exposing the intermediate tables, and finding the data records from which it is derived. However, manually reviewing code and backtracing data requires certain expertise, while inspecting the traced records in multiple tables and interpreting their relationships is tedious. In this work, we propose HYPNOS, a visualization system that supports interactive data lineage tracing for data transformation scripts. HYPNOS uses a lineage module for parsing and adapting code to capture both schema-level and instance-level data lineage from data transformation scripts. Then, it provides users with a lineage view for obtaining an overview of the data transformation process and a detail view for tracing instance-level data lineage and inspecting details. HYPNOS reveals different levels of data relationships and helps users with data lineage tracing. We demonstrate the usability and effectiveness of HYPNOS through a use case, interviews of four expert users, and a user study.
Xiwen Cai, Xiaodong Ge, Shuainan Ye, Di Weng, Datong Wei, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.1
2025 Linking Text and Visualizations via Contextual Knowledge Graph
abstract
The integration of visualizations and text is commonly found in data news, analytical reports, and interactive documents. For example, financial articles are presented along with interactive charts to show the changes in stock prices on Yahoo Finance. Visualizations enhance the perception of facts in the text while the text reveals insights of visual representation. However, effectively combining text and visualizations is challenging and tedious, which usually involves advanced programming skills. This paper proposes a semi-automatic pipeline that builds links between text and visualization. To resolve the relationship between text and visualizations, we present a method which structures a visualization and the underlying data as a contextual knowledge graph, based on which key phrases in the text are extracted, grouped, and mapped with visual elements. To support flexible customization of text-visualization links, our pipeline incorporates user knowledge to revise the links in a mixed-initiative manner. To demonstrate the usefulness and the versatility of our method, we replicate prior studies or cases in crafting interactive word-sized visualizations, annotating visualizations, and creating text-chart interactions based on a prototype system. We carry out two preliminary model tests and a user study and the results and user feedbacks suggest our method is effective.
Xiwen Cai, Di Weng, Taotao Fu, Siwei Fu, Yongheng Wang, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.1
2025 CodeLin: An in situ visualization method for understanding data transformation scripts
abstract
Understanding data transformation scripts is an essential task for data analysts who write code to process data. However, this can be challenging, especially when encountering unfamiliar scripts. Comments can help users understand data transformation code, but well-written comments are not always present. Visualization methods have been proposed to help analysts understand data transformations, but they generally require a separate view, which may distract users and entail efforts for connecting visualizations and code. In this work, we explore the use of in situ program visualization to help data analysts understand data transformation scripts. We present CodeLin, a new visualization method that combines word-sized glyphs for presenting transformation semantics and a lineage graph for presenting data lineage in an in situ manner. Through a use case, code pattern demonstrations, and a preliminary user study, we demonstrate the effectiveness and usability of CodeLin. We further discuss how visualization can help users understand data transformation code.
Xiwen Cai, Zhongsu Luo, Di Weng, Shuainan Ye, Yingcai Wu
Vis. Informatics1
2024 A Comparative Study on Fixed-Order Event Sequence Visualizations: Gantt, Extended Gantt, and Stringline Charts
abstract
We conduct two in-lab experiments (N = 93) to evaluate the effectiveness of Gantt charts, extended Gantt charts, and stringline charts for visualizing fixed-order event sequence data. We first formulate five types of event sequences and define three types of sequence elements: point events, interval events, and the temporal gaps between them. Our two experiments focus on event sequences with a pre-defined, fixed order and measure task error rates and completion time. The first experiment shows single sequences and assesses the three charts' performance in comparing event duration or gap. The second experiment shows multiple sequences and evaluates how well the charts reveal temporal patterns. The results suggest that when visualizing single fixed-order event sequences, 1) Gantt and extended Gantt charts lead to comparable error rates in the duration-comparing task; 2) Gantt charts exhibit either shorter or equal completion time than extended Gantt charts; 3) both Gantt and extended Gantt charts demonstrate shorter completion times than stringline charts; 4) however, stringline charts outperform the other two charts with fewer errors in the comparing task when event type counts are high. Additionally, when visualizing multiple point-based fixed-order event sequences, stringline charts require less time than Gantt charts for people to find temporal patterns. Based on these findings, we discuss design opportunities for visualizing fixed-order event sequences and discuss future avenues for optimizing these charts.
Junxiu Tang, Fumeng Yang, Jiang Wu 0012, Yifang Wang 0001, Xiwen Cai, Lingyun Yu 0001, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.6
2023 Complementary surrogate-assisted differential evolution algorithm for expensive multi-objective problems under a limited computational budget
Xiwen Cai, Gan Ruan, Bo Yuan 0006, Liang Gao 0001
Inf. Sci.1
2021 A Surrogate-Assisted Multiswarm Optimization Algorithm for High-Dimensional Computationally Expensive Problems
abstract
This article presents a surrogate-assisted multiswarm optimization (SAMSO) algorithm for high-dimensional computationally expensive problems. The proposed algorithm includes two swarms: the first one uses the learner phase of teaching-learning-based optimization (TLBO) to enhance exploration and the second one uses the particle swarm optimization (PSO) for faster convergence. These two swarms can learn from each other. A dynamic swarm size adjustment scheme is proposed to control the evolutionary progress. Two coordinate systems are used to generate promising positions for the PSO in order to further enhance its search efficiency on different function landscapes. Moreover, a novel prescreening criterion is proposed to select promising individuals for exact function evaluations. Several commonly used benchmark functions with their dimensions varying from 30 to 200 are adopted to evaluate the proposed algorithm. The experimental results demonstrate the superiority of the proposed algorithm over three state-of-the-art algorithms.
Xiwen Cai, Liang Gao 0001, Weiming Shen 0001
IEEE Trans. Cybern.2
2020 A Surrogate-Assisted Offspring Generation Method for Expensive Multi-objective optimization Problems
abstract
Surrogate-assisted multi-objective evolutionary algorithms have been commonly used to solve multi-objective expensive problems. In this paper, we investigate whether the surrogate-assisted offspring generation method can improve the optimization efficiency of multi-objective evolutionary algorithms. We first construct a surrogate model for each objective function. After that, some candidate solutions from the surrogate models are used to produce promising offspring for the multi-objective evolutionary algorithm. In addition, a prescreening criterion based on reference vectors and the nondominated rank is used to select the surviving offspring and exactly evaluated individuals. The pre-screening criterion can ensure the diversity and convergence of the offspring, and reduce function evaluations. Benchmark problems with their dimensions varying from 8 to 30 are used to test the effects of the surrogate-assisted offspring generation method under the framework of using the pre-screening criterion. Experimental results show that using the candidate solutions from surrogate models can enhance the performance of its basic algorithm on most of the problems.
Liang Gao 0001, Weiming Shen 0001, Xiwen Cai
CEC4
2020 Surrogate-assisted classification-collaboration differential evolution for expensive constrained optimization problems
Zan Yang, Haobo Qiu, Liang Gao 0001, Xiwen Cai
Inf. Sci.4
2020 Efficient Generalized Surrogate-Assisted Evolutionary Algorithm for High-Dimensional Expensive Problems
abstract
Engineering optimization problems usually involve computationally expensive simulations and many design variables. Solving such problems in an efficient manner is still a major challenge. In this paper, a generalized surrogate-assisted evolutionary algorithm is proposed to solve such high-dimensional expensive problems. The proposed algorithm is based on the optimization framework of the genetic algorithm (GA). This algorithm proposes to use a surrogate-based trust region local search method, a surrogate-guided GA (SGA) updating mechanism with a neighbor region partition strategy and a prescreening strategy based on the expected improvement infilling criterion of a simplified Kriging in the optimization process. The SGA updating mechanism is a special characteristic of the proposed algorithm. This mechanism makes a fusion between surrogates and the evolutionary algorithm. The neighbor region partition strategy effectively retains the diversity of the population. Moreover, multiple surrogates used in the SGA updating mechanism make the proposed algorithm optimize robustly. The proposed algorithm is validated by testing several high-dimensional numerical benchmark problems with dimensions varying from 30 to 100, and an overall comparison is made between the proposed algorithm and other optimization algorithms. The results show that the proposed algorithm is very efficient and promising for optimizing high-dimensional expensive problems.
Xiwen Cai, Liang Gao 0001, Xinyu Li 0001
IEEE Trans. Evol. Comput.1
2019 Surrogate's Optima Assisted Evolutionary Algorithm for Optimization of Expensive Problems
abstract
In this paper, an efficient surrogate's optima assisted evolutionary optimization algorithm is proposed for the optimization of computationally expensive problems, which sometimes involve costly simulation analysis. The proposed algorithm uses the global optimum and local optima of the surrogates to speed up the evolutionary optimization process. Moreover, the optimization efficiency of the proposed algorithm can be enhanced by using a surrogate prescreening strategy. In order to validate the proposed algorithm, it is tested on several common numerical benchmark problems of 30 dimensions and compared with several optimization algorithms. The results show that the proposed algorithm is very promising for the optimization of the expensive problems.
Xiwen Cai, Liang Gao 0001
CEC1
2019 A Novel Surrogate-assisted Differential Evolution for Expensive Optimization Problems with both Equality and Inequality Constraints
abstract
Surrogates have recently shown excellent abilities in assisting evolutionary algorithms for solving computationally expensive constrained optimization problems (ECOPs). However, the effectiveness of such surrogate-assisted evolutionary algorithms has only been verified on ECOPs with inequality constraints. In this paper, a Novel Surrogate-Assisted Differential Evolution (NSADE) algorithm is proposed for solving ECOPs with equality and inequality constraints, in which a trial vector generation mechanism and two surrogate-assisted local search phases are carried out iteratively. The trial vector generation mechanism based on information exchange between the historical elite solution set and current population is utilized to balance exploiting potential areas and exploring unknown areas. Then the expectation improvement-based local search is used to not only guide the current population to move towards feasible region but also alleviate the inaccuracy of the surrogate on the constraint boundary. Finally, a solution identification-based local search is utilized to further optimize two different types of historical elite solutions. Empirical studies on fifteen widely used benchmark problems demonstrate that the proposed NSADE can effectively obtain high-quality feasible solutions on ECOPs with equality constraints under a limited computational budget.
Zan Yang, Haobo Qiu, Liang Gao 0001, Xiwen Cai
CEC6
2019 An efficient surrogate-assisted particle swarm optimization algorithm for high-dimensional expensive problems
Xiwen Cai, Haobo Qiu, Liang Gao 0001, Xinyu Shao
Knowl. Based Syst.1
2019 A User Study on the Capability of Three Geo-Based Features in Analyzing and Locating Trajectories
abstract
Visual analysis is widely applied to study human mobility due to the ability of integrating contextual information from multiple data sources. Analyzing trajectory data through visualization improves the efficiency and accuracy of the analysis, yet it may induce exposure of the location privacy. To balance the location privacy and analysis effectiveness, this work focuses on the behaviors of different geo-based contexts in the process of trajectory interpretation. Three types of geo-based contexts are identified after surveying 94 related literatures. We further conduct experiments to investigate their capability by evaluating how they benefit the analysis, and whether they lead to location privacy exposure. Finally, we report and discuss interesting findings, and provide guidelines to the design of privacy-preserving analysis approaches for human periodic trajectories.
Xumeng Wang, Tianlong Gu, Xiwen Cai, Tianyi Lao, Yingcai Wu, Wei Chen 0001
IEEE Trans. Intell. Transp. Syst.4
2019 A Semantic-Based Method for Visualizing Large Image Collections
abstract
Interactive visualization of large image collections is important and useful in many applications, such as personal album management and user profiling on images. However, most prior studies focus on using low-level visual features of images, such as texture and color histogram, to create visualizations without considering the more important semantic information embedded in images. This paper proposes a novel visual analytic system to analyze images in a semantic-aware manner. The system mainly comprises two components: a semantic information extractor and a visual layout generator. The semantic information extractor employs an image captioning technique based on convolutional neural network (CNN) to produce descriptive captions for images, which can be transformed into semantic keywords. The layout generator employs a novel co-embedding model to project images and the associated semantic keywords to the same 2D space. Inspired by the galaxy metaphor, we further turn the projected 2D space to a galaxy visualization of images, in which semantic keywords and images are visually encoded as stars and planets. Our system naturally supports multi-scale visualization and navigation, in which users can immediately see a semantic overview of an image collection and drill down for detailed inspection of a certain group of images. Users can iteratively refine the visual layout by integrating their domain knowledge into the co-embedding process. Two task-based evaluations are conducted to demonstrate the effectiveness of our system.
Xiao Xie, Xiwen Cai, Junpei Zhou, Nan Cao 0001, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.2
2017 Ensemble of surrogate models using sign based cross validation error
abstract
Surrogate model methods are usually used as a time-saving approach to reduce the computational burden of expensive computer simulations, while the appropriate surrogate model for an unknown problem is often difficult to choose. In this paper, a SCG (Ensemble of surrogate models using Sign based Cross validation error with Global correction) ensemble modeling method based on pointwise local measures is proposed. The SCG method is constructed from three aspects: (i) sign based pointwise local weight, (ii) global correction parameter, and (iii) distance metric. SCG is used in four benchmark problems to test the effectiveness of this new method. The test results show that the SCG model has higher accuracy and robustness, while spending almost the same modeling time compared to other existing ensemble models.
Haobo Qiu, Xiwen Cai, Liang Gao 0001
CSCWD4