Zhongsu Luo

dblp:249/9854 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0003-0885-2742ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Human-computer interaction and pervasive computing
4 papers
User interface design and tools · 61% Human-AI interaction · 39%
Databases, data mining, and information retrieval
5 papers
Data integration and cleaning · 90% Query processing and optimization · 10%
Computer graphics and multimedia
2 papers
Visualization and visual analytics · 100%
Software engineering, system software, and programming languages
2 papers
Program synthesis and code generation · 77% Program analysis · 23%
Artificial intelligence
2 papers
Language models and text generation · 60% Deep learning architectures and training · 40%

Topics — the 6 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visualization and visual analytics
interactive data analysis
1.922026
RuleScope: Semantic-Aware Authoring of Data Validation Rules · IEEE Trans. Vis. Comput. Graph. 2026
Ferry: Toward Better Understanding of Input/Output Space for Data Wrangling Scripts · IEEE Trans. Vis. Comput. Graph. 2025
User interface design and tools › visualization
program visualization
1.322023
Revealing the Semantics of Data Wrangling Scripts With Comantics · IEEE Trans. Vis. Comput. Graph. 2023
Visualizing the Scripts of Data Wrangling With Somnus · IEEE Trans. Vis. Comput. Graph. 2023
Data integration and cleaning
data wrangling
1.122025
Ferry: Toward Better Understanding of Input/Output Space for Data Wrangling Scripts · IEEE Trans. Vis. Comput. Graph. 2025
Visualizing the Scripts of Data Wrangling With Somnus · IEEE Trans. Vis. Comput. Graph. 2023
Data integration and cleaning › data quality
data validation
1.012026
RuleScope: Semantic-Aware Authoring of Data Validation Rules · IEEE Trans. Vis. Comput. Graph. 2026
Program synthesis and code generation
code generation with language models
0.912025
ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User Prompts · UIST 2025
Machine learning › Deep learning architectures and training
convolutional neural network
0.212023
Revealing the Semantics of Data Wrangling Scripts With Comantics · IEEE Trans. Vis. Comput. Graph. 2023

Methods — techniques the papers use, named apart from their topics

large language model · 3.7matrix-based visualization · 3.0LLM · 3.0visualization · 2.6constraint generation model · 2.6interactive tree view · 2.0slot filling · 2.0siamese convolutional neural network · 2.0provenance graph · 1.3glyph design · 1.3
YearPublicationVenuePosition
2026 Cerebra: Aligning Implicit Knowledge in Interactive SQL Authoring
abstract
LLM-driven tools have significantly lowered barriers to writing SQL queries. However, user instructions are often underspecified, assuming the model understands implicit knowledge, such as dataset schemas, domain conventions, and task-specific requirements, that isn’t explicitly provided. This results in frequently erroneous scripts that require users to repeatedly clarify their intent. Additionally, users struggle to validate generated scripts because they cannot verify whether the model correctly applied implicit knowledge. We present Cerebra, an interactive NL-to-SQL tool that aligns implicit knowledge between users and LLMs during SQL authoring. Cerebra automatically retrieves implicit knowledge from historical SQL scripts based on user instructions, presents this knowledge in an interactive tree view for code review, and supports iterative refinement to improve generated scripts. To evaluate the effectiveness and usability of Cerebra, we conducted a user study with 16 participants, demonstrating its improved support for customized SQL authoring. The source code of Cerebra is available at https://github.com/zjuidg/CHI26-Cerebra.
Yunfan Zhou, Qiming Shi, Zhongsu Luo, Xiwen Cai, Yanwei Huang, Daehyun Kim 0005, Di Weng, Yingcai Wu
CHI3
2026 RuleScope: Semantic-Aware Authoring of Data Validation Rules
abstract
Data validation is a crucial step in data analytics workflows that assesses and ensures the reliability of data flowing into analytical processes. One common approach to data validation involves defining validation rules, which provide explicit constraints and conditions that data must satisfy. However, creating accurate and effective validation rules remains challenging for many practitioners. This challenge stems from the need for practitioners to understand both data structures and their domain-specific semantic relationships. Recent studies have proposed automated approaches to generate validation rules by deriving patterns from data properties. However, these approaches generate rules with limited interpretability and lack support for rule verification and modification, making the rules difficult to understand and adapt. To address these limitations in current validation rule authoring approaches, we present RuleScope, an interactive system for authoring data validation rules through semantic-aware rule generation, visualization, and refinement. RuleScope employs an LLM-based workflow to generate interpretable rules by analyzing data semantics and incorporating domain knowledge. To facilitate rule comprehension, we design a matrix-based visualization that helps users understand rules and analyze validation results. Additionally, RuleScope enables users to interactively refine rules. We evaluate the LLM-based workflow through model evaluation on datasets from different domains and assess RuleScope's usability and effectiveness through two case studies and a user study.
Zhongsu Luo, Di Weng, Xiwen Cai, Xinhuan Shu, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.1
2025 ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User Prompts
Zhongsu Luo, Yunfan Zhou, Xinhuan Shu, Di Weng, Yingcai Wu
UIST3
2025 Ferry: Toward Better Understanding of Input/Output Space for Data Wrangling Scripts
abstract
Understanding the input and output of data wrangling scripts is crucial for various tasks like debugging code and onboarding new data. However, existing research on script understanding primarily focuses on revealing the process of data transformations, lacking the ability to analyze the potential scope, i.e., the space of script inputs and outputs. Meanwhile, constructing input/output space during script analysis is challenging, as the wrangling scripts could be semantically complex and diverse, and the association between different data objects is intricate. To facilitate data workers in understanding the input and output space of wrangling scripts, we summarize ten types of constraints to express table space and build a mapping between data transformations and these constraints to guide the construction of the input/output for individual transformations. Then, we propose a constraint generation model for integrating table constraints across multiple transformations. Based on the model, we develop Ferry, an interactive system that extracts and visualizes the data constraints describing the input and output space of data wrangling scripts, thereby enabling users to grasp the high-level semantics of complex scripts and locate the origins of faulty data transformations. Besides, Ferry provides example input and output data to assist users in interpreting the extracted constraints and checking and resolving the conflicts between these constraints and any uploaded dataset. Ferry's effectiveness and usability are evaluated through two usage scenarios and two case studies, including understanding, debugging, and checking both single and multiple scripts, with and without executable data. Furthermore, an illustrative application is presented to demonstrate Ferry's flexibility.
Zhongsu Luo, Xinhuan Shu, Di Weng, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.1
2025 CodeLin: An in situ visualization method for understanding data transformation scripts
abstract
Understanding data transformation scripts is an essential task for data analysts who write code to process data. However, this can be challenging, especially when encountering unfamiliar scripts. Comments can help users understand data transformation code, but well-written comments are not always present. Visualization methods have been proposed to help analysts understand data transformations, but they generally require a separate view, which may distract users and entail efforts for connecting visualizations and code. In this work, we explore the use of in situ program visualization to help data analysts understand data transformation scripts. We present CodeLin, a new visualization method that combines word-sized glyphs for presenting transformation semantics and a lineage graph for presenting data lineage in an in situ manner. Through a use case, code pattern demonstrations, and a preliminary user study, we demonstrate the effectiveness and usability of CodeLin. We further discuss how visualization can help users understand data transformation code.
Xiwen Cai, Zhongsu Luo, Di Weng, Shuainan Ye, Yingcai Wu
Vis. Informatics3
2023 Visualizing the Scripts of Data Wrangling With Somnus
abstract
Data workers use various scripting languages for data transformation, such as SAS, R, and Python. However, understanding intricate code pieces requires advanced programming skills, which hinders data workers from grasping the idea of data transformation at ease. Program visualization is beneficial for debugging and education and has the potential to illustrate transformations intuitively and interactively. In this article, we explore visualization design for demonstrating the semantics of code pieces in the context of data transformation. First, to depict individual data transformations, we structure a design space by two primary dimensions, i.e., key parameters to encode and possible visual channels to be mapped. Then, we derive a collection of 23 glyphs that visualize the semantics of transformations. Next, we design a pipeline, named Somnus, that provides an overview of the creation and evolution of data tables using a provenance graph. At the same time, it allows detailed investigation of individual transformations. User feedback on Somnus is positive. Our study participants achieved better accuracy with less time using Somnus, and preferred it over carefully-crafted textual description. Further, we provide two example applications to demonstrate the utility and versatility of Somnus.
Siwei Fu, Guoming Ding, Zhongsu Luo, Wei Chen 0001, Hujun Bao, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.4
2023 Revealing the Semantics of Data Wrangling Scripts With Comantics
abstract
Data workers usually seek to understand the semantics of data wrangling scripts in various scenarios, such as code debugging, reusing, and maintaining. However, the understanding is challenging for novice data workers due to the variety of programming languages, functions, and parameters. Based on the observation that differences between input and output tables highly relate to the type of data transformation, we outline a design space including 103 characteristics to describe table differences. Then, we develop Comantics, a three-step pipeline that automatically detects the semantics of data transformation scripts. The first step focuses on the detection of table differences for each line of wrangling code. Second, we incorporate a characteristic-based component and a Siamese convolutional neural network-based component for the detection of transformation types. Third, we derive the parameters of each data transformation by employing a "slot filling" strategy. We design experiments to evaluate the performance of Comantics. Further, we assess its flexibility using three example applications in different domains.
Zhongsu Luo, Siwei Fu, Yongheng Wang, Mingliang Xu 0001, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.2