EDBT 2026 Demo / reviewers in the wild / expert
Di Weng
dblp:190/2257
· DBLP profile ↗
42ranked-venue papers
5as first author
38since 2021 · last 2026
0000-0003-2712-7274ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 2 first-author · 23 since 2021Human-computer interaction and ubiquitous computing · 11 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TSEditor: Interactive Time Series Editing for Privacy Preservation
Kaicheng Shao, Yuanzhe Jin, Xumeng Wang, Zikun Deng, Di Weng, Yingcai Wu |
CHI | 7 |
| 2026 | Cerebra: Aligning Implicit Knowledge in Interactive SQL AuthoringabstractLLM-driven tools have significantly lowered barriers to writing SQL queries. However, user instructions are often underspecified, assuming the model understands implicit knowledge, such as dataset schemas, domain conventions, and task-specific requirements, that isn’t explicitly provided. This results in frequently erroneous scripts that require users to repeatedly clarify their intent. Additionally, users struggle to validate generated scripts because they cannot verify whether the model correctly applied implicit knowledge. We present Cerebra, an interactive NL-to-SQL tool that aligns implicit knowledge between users and LLMs during SQL authoring. Cerebra automatically retrieves implicit knowledge from historical SQL scripts based on user instructions, presents this knowledge in an interactive tree view for code review, and supports iterative refinement to improve generated scripts. To evaluate the effectiveness and usability of Cerebra, we conducted a user study with 16 participants, demonstrating its improved support for customized SQL authoring. The source code of Cerebra is available at https://github.com/zjuidg/CHI26-Cerebra. Yunfan Zhou, Qiming Shi, Zhongsu Luo, Xiwen Cai, Yanwei Huang, Daehyun Kim 0005, Di Weng, Yingcai Wu |
CHI | 7 |
| 2026 | A Declarative Grammar for Interactive Trajectory Visualization: Interaction as First-Class Component
Shifu Chen, Xiaodan Miao, Dazhen Deng, Zikun Deng, Di Weng, Yingcai Wu |
PacificVis | 5 |
| 2026 | TrajectoryCurer: Visual Analysis of Trajectory Data Quality
Xiaodan Miao, Sitong Pan, Shifu Chen, Di Weng, Yingcai Wu |
PacificVis | 4 |
| 2026 | GeoAuthor: Linking Text and Visualization for Geographic Article AuthoringabstractArticles containing geographic information are widely distributed and commonly used in daily life, frequently incorporating geographic visualizations as illustrations. However, the creation of such articles remains cumbersome, necessitating authors to switch between authoring text and illustrations, thereby disrupting immersive writing. Our interviews corroborated this observation and revealed the primary challenge in the traditional process stems from the low synchronization frequency between text and geographic visualizations during creation, coupled with weak visual links, forcing users to mentally maintain this synchronization and thereby increasing their cognitive burden. In response, we developed GeoAuthor, which facilitates the interactive creation of geographic articles by automatically synchronizing text creation with geographic visualizations with rich visual links. This bidirectional approach ensures that the written content and visual representations remain consistent and mutually informative throughout the creation process. Our evaluation demonstrated the efficacy of GeoAuthor, indicating its capacity to streamline the process of creating geographic articles. Zhenning Chen, Hanbei Zhan, Shifu Chen, Zikun Deng, Di Weng, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2026 | KEditVis: A Visual Analytics System for Knowledge Editing of Large Language ModelsabstractLarge Language Models (LLMs) demonstrate exceptional capabilities in factual question answering, yet they sometimes provide incorrect responses. To address this issue, knowledge editing techniques have emerged as effective methods for correcting factual information in LLMs. However, typical knowledge editing workflows struggle with identifying the optimal set of model layers for editing and rely on summary indicators that provide insufficient guidance. This lack of transparency hinders effective comparison and identification of optimal editing strategies. In this paper, we present KEditVis, a novel visual analytics system designed to assist users in gaining a deeper understanding of knowledge editing through interactive visualizations, improving editing outcomes, and discovering valuable insights for the future development of knowledge editing algorithms. With KEditVis, users can select appropriate layers as the editing target, explore the reasons behind ineffective edits, and perform more targeted and effective edits. Our evaluation, including usage scenarios, expert interviews, and a user study, validates the effectiveness and usability of the system. Zhenning Chen, Hanbei Zhan, Yanwei Huang, Xin Wu 0003, Dazhen Deng, Di Weng, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2026 | RCInvestigator: Towards Better Investigation of Anomaly Root Causes in Cloud Computing SystemsabstractRoot cause analysis (RCA) is critical for maintaining the availability and efficiency of cloud computing systems. However, identifying root causes from the large-scale, high-dimensional monitoring data generated by these complex environments is a significant challenge. Current approaches often rely on time-consuming manual analysis to ensure flexibility and reliability, while recent automated methods lack the crucial insights provided by domain experts. To bridge this gap, we propose RCInvestigator, a visual analytics system that facilitates interactive root cause investigation by establishing a tight collaboration between human experts and machine analysis. Our approach addresses three key challenges: a) modeling databases for the root cause investigation, b) inferring root causes from large-scale time series, and c) building comprehensible investigation results. We demonstrate the effectiveness and utility of RCInvestigator through two real-world case studies, which received positive feedback from domain experts. Yunfan Zhou, Shandan Zhou, Weiwei Cui 0001, Qingwei Lin, Thomas Moscibroda, Di Weng, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2026 | RuleScope: Semantic-Aware Authoring of Data Validation RulesabstractData validation is a crucial step in data analytics workflows that assesses and ensures the reliability of data flowing into analytical processes. One common approach to data validation involves defining validation rules, which provide explicit constraints and conditions that data must satisfy. However, creating accurate and effective validation rules remains challenging for many practitioners. This challenge stems from the need for practitioners to understand both data structures and their domain-specific semantic relationships. Recent studies have proposed automated approaches to generate validation rules by deriving patterns from data properties. However, these approaches generate rules with limited interpretability and lack support for rule verification and modification, making the rules difficult to understand and adapt. To address these limitations in current validation rule authoring approaches, we present RuleScope, an interactive system for authoring data validation rules through semantic-aware rule generation, visualization, and refinement. RuleScope employs an LLM-based workflow to generate interpretable rules by analyzing data semantics and incorporating domain knowledge. To facilitate rule comprehension, we design a matrix-based visualization that helps users understand rules and analyze validation results. Additionally, RuleScope enables users to interactively refine rules. We evaluate the LLM-based workflow through model evaluation on datasets from different domains and assess RuleScope's usability and effectiveness through two case studies and a user study. Zhongsu Luo, Di Weng, Xiwen Cai, Xinhuan Shu, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | StructVizor: Interactive Profiling of Semi-Structured Textual DataabstractData profiling plays a critical role in understanding the structure of complex datasets and supporting numerous downstream tasks, such as social media analytics and financial fraud detection. While existing research predominantly focuses on structured data formats, a substantial portion of semi-structured textual data still requires ad-hoc and arduous manual profiling to extract and comprehend its internal structures. In this work, we propose StructVizor, an interactive profiling system that facilitates sensemaking and transformation of semi-structured textual data. Our tool mainly addresses two challenges: a) extracting and visualizing the diverse structural patterns within data, such as how information is organized or related, and b) enabling users to efficiently perform various wrangling operations on textual data. Through automatic data parsing and structure mining, StructVizor enables visual analytics of structural patterns, while incorporating novel interactions to enable profile-based data wrangling. A comparative user study involving 12 participants demonstrates the system's usability and its effectiveness in supporting exploratory data analysis and transformation tasks. Yanwei Huang, Yan Miao, Di Weng, Adam Perer, Yingcai Wu |
CHI | 3 |
| 2025 | RidgeBuilder: Interactive Authoring of Expressive Ridgeline Plots
Yangtian Liu, Junxin Li, Yanwei Huang, Yue Shangguan, Zikun Deng, Di Weng, Yingcai Wu |
CHI | 7 |
| 2025 | Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling ScriptsabstractData analysts frequently employ code completion tools in writing custom scripts to tackle complex tabular data wrangling tasks. However, existing tools do not sufficiently link the data contexts such as schemas and values with the code being edited. This not only leads to poor code suggestions, but also frequent interruptions in coding processes as users need additional code to locate and understand relevant data. We introduce Xavier, a tool designed to enhance data wrangling script authoring in computational notebooks. Xavier maintains users' awareness of data contexts while providing data-aware code suggestions. It automatically highlights the most relevant data based on the user's code, integrates both code and data contexts for more accurate suggestions, and instantly previews data transformation results for easy verification. To evaluate the effectiveness and usability of Xavier, we conducted a user study with 16 data analysts, showing its potential to streamline data wrangling scripts authoring. Yunfan Zhou, Xiwen Cai, Qiming Shi, Yanwei Huang, Haotian Li 0001, Huamin Qu, Di Weng, Yingcai Wu |
CHI | 7 |
| 2025 | ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User Prompts
Zhongsu Luo, Yunfan Zhou, Xinhuan Shu, Di Weng, Yingcai Wu |
UIST | 6 |
| 2025 | HYPNOS: Interactive Data Lineage Tracing for Data Transformation ScriptsabstractIn a formal data analysis workflow, data validation is a necessary step that helps data analysts verify the quality of the data and ensure the reliability of the results. Data analysts usually need to validate the result when encountering an unexpected result, such as an abnormal record in a table. In order to understand how a specific record is derived, they would backtrace it in the pipeline step by step via checking the code lines, exposing the intermediate tables, and finding the data records from which it is derived. However, manually reviewing code and backtracing data requires certain expertise, while inspecting the traced records in multiple tables and interpreting their relationships is tedious. In this work, we propose HYPNOS, a visualization system that supports interactive data lineage tracing for data transformation scripts. HYPNOS uses a lineage module for parsing and adapting code to capture both schema-level and instance-level data lineage from data transformation scripts. Then, it provides users with a lineage view for obtaining an overview of the data transformation process and a detail view for tracing instance-level data lineage and inspecting details. HYPNOS reveals different levels of data relationships and helps users with data lineage tracing. We demonstrate the usability and effectiveness of HYPNOS through a use case, interviews of four expert users, and a user study. Xiwen Cai, Xiaodong Ge, Shuainan Ye, Di Weng, Datong Wei, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | Linking Text and Visualizations via Contextual Knowledge GraphabstractThe integration of visualizations and text is commonly found in data news, analytical reports, and interactive documents. For example, financial articles are presented along with interactive charts to show the changes in stock prices on Yahoo Finance. Visualizations enhance the perception of facts in the text while the text reveals insights of visual representation. However, effectively combining text and visualizations is challenging and tedious, which usually involves advanced programming skills. This paper proposes a semi-automatic pipeline that builds links between text and visualization. To resolve the relationship between text and visualizations, we present a method which structures a visualization and the underlying data as a contextual knowledge graph, based on which key phrases in the text are extracted, grouped, and mapped with visual elements. To support flexible customization of text-visualization links, our pipeline incorporates user knowledge to revise the links in a mixed-initiative manner. To demonstrate the usefulness and the versatility of our method, we replicate prior studies or cases in crafting interactive word-sized visualizations, annotating visualizations, and creating text-chart interactions based on a prototype system. We carry out two preliminary model tests and a user study and the results and user feedbacks suggest our method is effective. Xiwen Cai, Di Weng, Taotao Fu, Siwei Fu, Yongheng Wang, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Relation-Driven Query of Multiple Time SeriesabstractQuerying time series based on their relations is a crucial part of multiple time series analysis. By retrieving and understanding time series relations, analysts can easily detect anomalies and validate hypotheses in complex time series datasets. However, current relation extraction approaches, including knowledge- and data-driven ones, tend to be laborious and do not support heterogeneous relations. By conducting a formative study with 11 experts, we concluded six time series relations, including correlation, causality, similarity, lag, arithmetic, and meta, and summarized three pain points in querying time series involving these relations. We proposed RelaQ, an interactive system that supports the time series query via relation specifications. RelaQ allows users to intuitively specify heterogeneous relations when querying multiple time series, understand the query results based on a scalable, multi-level visualization, and explore possible relations beyond the existing queries. RelaQ is evaluated with two cases and a user study with 12 participants, showing promising effectiveness and usability. Zikun Deng, Weiwei Cui 0001, Di Weng, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | Ferry: Toward Better Understanding of Input/Output Space for Data Wrangling ScriptsabstractUnderstanding the input and output of data wrangling scripts is crucial for various tasks like debugging code and onboarding new data. However, existing research on script understanding primarily focuses on revealing the process of data transformations, lacking the ability to analyze the potential scope, i.e., the space of script inputs and outputs. Meanwhile, constructing input/output space during script analysis is challenging, as the wrangling scripts could be semantically complex and diverse, and the association between different data objects is intricate. To facilitate data workers in understanding the input and output space of wrangling scripts, we summarize ten types of constraints to express table space and build a mapping between data transformations and these constraints to guide the construction of the input/output for individual transformations. Then, we propose a constraint generation model for integrating table constraints across multiple transformations. Based on the model, we develop Ferry, an interactive system that extracts and visualizes the data constraints describing the input and output space of data wrangling scripts, thereby enabling users to grasp the high-level semantics of complex scripts and locate the origins of faulty data transformations. Besides, Ferry provides example input and output data to assist users in interpreting the extracted constraints and checking and resolving the conflicts between these constraints and any uploaded dataset. Ferry's effectiveness and usability are evaluated through two usage scenarios and two case studies, including understanding, debugging, and checking both single and multiple scripts, with and without executable data. Furthermore, an illustrative application is presented to demonstrate Ferry's flexibility. Zhongsu Luo, Xinhuan Shu, Di Weng, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | ChronoDeck: A Visual Analytics Approach for Hierarchical Time Series AnalysisabstractHierarchical time series data comprises a collection of time series aggregated at multiple levels based on categorical, geographical, or physical constraints, the analysis of which aids analysts across various domains like retail, finance, and energy, in gaining valuable insights and making informed decisions. However, existing interactive exploratory analysis approaches for hierarchical time series data fall short in analyzing time series across different aggregation levels and supporting more complex analytical tasks beyond common ones like summarize and compare. These limitations motivate us to develop a new visual analytics approach. We first generalize a taxonomy to delineate various tasks in hierarchical time series analysis, derived from literature survey and expert interviews. Based on this taxonomy, we develop ChronoDeck, an interactive system that incorporates a multi-column hierarchical time series visualization for implementing various analytical tasks and distilling insights from the data. ChronoDeck visualizes each aggregation level of hierarchical time series with a combination of coordinated dimensionality reduction and small multiples visualizations, alongside interactions including highlight, align, filter, and select, assisting users in the visualization, comparison, and transformation of hierarchical time series, as well as identifying the entities of interest. The effectiveness of ChronoDeck is demonstrated by case studies on three real-world datasets and expert interviews. Lingyu Meng, Keyi Yang, Jiabin Xu, Zikun Deng, Di Weng, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | CodeLin: An in situ visualization method for understanding data transformation scriptsabstractUnderstanding data transformation scripts is an essential task for data analysts who write code to process data. However, this can be challenging, especially when encountering unfamiliar scripts. Comments can help users understand data transformation code, but well-written comments are not always present. Visualization methods have been proposed to help analysts understand data transformations, but they generally require a separate view, which may distract users and entail efforts for connecting visualizations and code. In this work, we explore the use of in situ program visualization to help data analysts understand data transformation scripts. We present CodeLin, a new visualization method that combines word-sized glyphs for presenting transformation semantics and a lineage graph for presenting data lineage in an in situ manner. Through a use case, code pattern demonstrations, and a preliminary user study, we demonstrate the effectiveness and usability of CodeLin. We further discuss how visualization can help users understand data transformation code. Xiwen Cai, Zhongsu Luo, Di Weng, Shuainan Ye, Yingcai Wu |
Vis. Informatics | 4 |
| 2025 | PVeSight: Dimensionality reduction-based anomaly detection and visual analysis of photovoltaic stringsabstractEfficient and accurate detection of anomalies in photovoltaic (PV) strings is essential for ensuring the normal operation of PV power stations. Most existing studies focus on developing automated anomaly detection models based on temporal abnormalities in PV strings. However, since analyzing anomalies often requires domain knowledge, existing automated methods have significant limitations in assisting experts to understand the causes and impact of these anomalies. In close collaboration with domain experts, this work has summarized the specific user requirements for PV string anomaly detection and designed PVeSight, an interactive visual analysis system to help experts discover and analyze anomalies in PV strings. We use dimensionality reduction techniques to generate string pattern map. These maps are used for anomaly detection, classifying anomalies, comparative analysis between strings, and hierarchical analysis under inverters and combiner boxes. This helps experts trace the causes of anomalies and acquire valuable insights into anomalous PV strings. Through case studies and expert evaluation, we verified the usability and effectiveness of PVeSight for PV string anomaly detection. Yurun Yang, Xinjing Yi, Yingqiang Jin, Dazhen Deng, Di Weng, Yingcai Wu |
Vis. Informatics | 8 |
| 2024 | Table Illustrator: Puzzle-based interactive authoring of plain tablesabstractPlain tables excel at displaying data details and are widely used in data presentation, often polished to an elaborate appearance for readability in many scenarios. However, existing authoring tools fail to provide both flexible and efficient support for altering the table layout and styles, motivating us to develop an intuitive and swift tool for table prototyping. To this end, we contribute Table Illustrator, a table authoring system taking a novel visual metaphor, puzzle, as the primary interaction unit. Through combinations and configurations on puzzles, the system enables rapid table construction and supports a diverse range of table layouts and styles. The tool design is informed by practical challenges and requirements from interviews with 10 table practitioners and a structured design space based on an analysis of over 2,500 real-world tables. User studies showed that Table Illustrator achieved comparable performance to Microsoft Excel while reducing users’ completion time and perceived workload. Yanwei Huang, Yurun Yang, Xinhuan Shu, Di Weng, Yingcai Wu |
CHI | 5 |
| 2024 | A Deep Spatiotemporal Trajectory Representation Learning Framework for ClusteringabstractLearning trajectory representations is essential in many Location Based Services (LBS) applications. Most traditional methods extract trajectory representations based on manually defined features, while deep learning-based methods can reduce part of the human effort. We propose a Deep Spatiotemporal Trajectory Clustering (DSTC) framework to tackle the Spatiotemporal Trajectory Representation Learning towards the Clustering-friendly space (STRLC) problem. Solving the STRLC problem is not a trivial task because: (1) Defining a uniform token size for datasets with an uneven density of trajectory data is challenging. (2) Measuring the similarity between trajectories spanning time zero in the time dimension is a problem to be solved. (3) It requires first learning a vector that can represent the overall characteristics of spatiotemporal trajectories and then mapping it to a more suitable space for clustering. To tackle these challenges, we first utilize the density-based clustering method to define tokens representing the trajectory points automatically. Then, we use polar coordinates to represent the temporal dimension of trajectories. Additionally, we improve the learned trajectory representations in a clustering-oriented latent space end to end. Experiments conducted on benchmark datasets demonstrate that DSTC achieves better accuracy than existing methods. Moreover, the representations learned from spatiotemporal trajectory data in the real world can be used to identify popular routes during the day. Yongheng Wang, Zhengxuan Lin, Xiongnan Jin, Xing Jin 0002, Di Weng, Yingcai Wu |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | Visualizing Large-Scale Spatial Time Series with GeoChronabstractIn geo-related fields such as urban informatics, atmospheric science, and geography, large-scale spatial time (ST) series (i.e., geo-referred time series) are collected for monitoring and understanding important spatiotemporal phenomena. ST series visualization is an effective means of understanding the data and reviewing spatiotemporal phenomena, which is a prerequisite for in-depth data analysis. However, visualizing these series is challenging due to their large scales, inherent dynamics, and spatiotemporal nature. In this study, we introduce the notion of patterns of evolution in ST series. Each evolution pattern is characterized by 1) a set of ST series that are close in space and 2) a time period when the trends of these ST series are correlated. We then leverage Storyline techniques by considering an analogy between evolution patterns and sessions, and finally design a novel visualization called GeoChron, which is capable of visualizing large-scale ST series in an evolution pattern-aware and narrative-preserving manner. GeoChron includes a mining framework to extract evolution patterns and two-level visualizations to enhance its visual scalability. We evaluate GeoChron with two case studies, an informal user study, an ablation study, parameter analysis, and running time analysis. Zikun Deng, Shifu Chen, Tobias Schreck, Dazhen Deng, Tan Tang, Mingliang Xu 0001, Di Weng, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2024 | Multilevel Visual Analysis of Aggregate Geo-NetworksabstractNumerous patterns found in urban phenomena, such as air pollution and human mobility, can be characterized as many directed geospatial networks (geo-networks) that represent spreading processes in urban space. These geo-networks can be analyzed from multiple levels, ranging from the macro-level of summarizing all geo-networks, meso-level of comparing or summarizing parts of geo-networks, and micro-level of inspecting individual geo-networks. Most of the existing visualizations cannot support multilevel analysis well. These techniques work by: 1) showing geo-networks separately with multiple maps leads to heavy context switching costs between different maps; 2) summarizing all geo-networks into a single network can lead to the loss of individual information; 3) drawing all geo-networks onto one map might suffer from the visual scalability issue in distinguishing individual geo-networks. In this study, we propose GeoNetverse, a novel visualization technique for analyzing aggregate geo-networks from multiple levels. Inspired by metro maps, GeoNetverse balances the overview and details of the geo-networks by placing the edges shared between geo-networks in a stacked manner. To enhance the visual scalability, GeoNetverse incorporates a level-of-detail rendering, a progressive crossing minimization, and a coloring technique. A set of evaluations was conducted to evaluate GeoNetverse from multiple perspectives. Zikun Deng, Shifu Chen, Xiao Xie, Guodao Sun, Mingliang Xu 0001, Di Weng, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | Interactive Table Synthesis With Natural LanguageabstractTables are a ubiquitous data format for insight communication. However, transforming data into consumable tabular views remains a challenging and time-consuming task. To lower the barrier of such a task, research efforts have been devoted to developing interactive approaches for data transformation, but many approaches still presume that their users have considerable knowledge of various data transformation concepts and functions. In this study, we leverage natural language (NL) as the primary interaction modality to improve the accessibility of average users to performing complex data transformation and facilitate intuitive table generation and editing. Designing an NL-driven data transformation approach introduces two challenges: 1) NL-driven synthesis of interpretable pipelines and 2) incremental refinement of synthesized tables. To address these challenges, we present NL2Rigel, an interactive tool that assists users in synthesizing and improving tables from semi-structured text with NL instructions. Based on a large language model and prompting techniques, NL2Rigel can interpret the given NL instructions into a table synthesis pipeline corresponding to Rigel specifications, a declarative language for tabular data transformation. An intuitive interface is designed to visualize the synthesis pipeline and the generated tables, helping users understand the transformation process and refine the results efficiently with targeted NL instructions. The comprehensiveness of NL2Rigel is demonstrated with an example gallery, and we further confirmed NL2Rigel's usability with a comparative user study by showing that the task completion time with NL2Rigel is significantly shorter than that with the original version of Rigel with comparable completion rates. Yanwei Huang, Yunfan Zhou, Changhao Pan, Xinhuan Shu, Di Weng, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | JsonCurer: Data Quality Management for JSON Based on an Aggregated SchemaabstractHigh-quality data is critical to deriving useful and reliable information. However, real-world data often contains quality issues undermining the value of the derived information. Most existing research on data quality management focuses on tabular data, leaving semi-structured data under-exploited. Due to the schema-less and hierarchical features of semi-structured data, discovering and fixing quality issues is challenging and time-consuming. To address the challenge, this paper presents JsonCurer, an interactive visualization system to assist with data quality management in the context of JSON data. To have an overview of quality issues, we first construct a taxonomy based on interviews with data practitioners and a review of 119 real-world JSON files. Then we highlight a schema visualization that presents structural information, statistical features, and quality issues of JSON data. Based on a similarity-based aggregation technique, the visualization depicts the entire JSON data with a concise tree, where summary visualizations are given above each node, and quality issues are illustrated using Bubble Sets across nodes. We evaluate the effectiveness and usability of JsonCurer with two case studies. One is in the domain of data analysis while the other concerns quality assurance in MongoDB documents. Siwei Fu, Di Weng, Yongheng Wang, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | GeoCamera: Telling Stories in Geographic Visualizations with Camera MovementsabstractIn geographic data videos, camera movements are frequently used and combined to present information from multiple perspectives. However, creating and editing camera movements requires significant time and professional skills. This work aims to lower the barrier of crafting diverse camera movements for geographic data videos. First, we analyze a corpus of 66 geographic data videos and derive a design space of camera movements with a dimension for geospatial targets and one for narrative purposes. Based on the design space, we propose a set of adaptive camera shots and further develop an interactive tool called GeoCamera. This interactive tool allows users to flexibly design camera movements for geographic visualizations. We verify the expressiveness of our tool through case studies and evaluate its usability with a user study. The participants find that the tool facilitates the design of camera movements. Wenchao Li 0005, Zhan Wang 0001, Yun Wang 0012, Di Weng, Liwenhan Xie, Siming Chen 0001, Huamin Qu |
CHI | 4 |
| 2023 | Contextual Self-attentive Temporal Point Process for Physical Decommissioning Prediction of Cloud AssetsabstractAs cloud computing continues to expand globally, the need for effective management of decommissioned cloud assets in data centers becomes increasingly important. This work focuses on predicting the physical decommissioning date of cloud assets as a crucial component in reverse cloud supply chain management and data center warehouse operation. The decommissioning process is modeled as a contextual self-attentive temporal point process, which incorporates contextual information to model sequences with parallel events and provides more accurate predictions with more seen historical data. We conducted extensive offline and online experiments in 20 sampled data centers. The results show that the proposed methodology achieves the best performance compared with baselines and improves remarkable 94% prediction accuracy in online experiments. This modeling methodology can be extended to other domains with similar workflow-like processes. Fangkai Yang, Lu Wang 0029, Bo Qiao 0001, Di Weng, Xiaoting Qin, Gregory Weber, Durgesh Nandini Das, Srinivasan Rakhunathan, Ranganathan Srikanth, Qingwei Lin, Dongmei Zhang 0001 |
KDD | 5 |
| 2023 | A survey of urban visual analytics: Advances and future directionsabstractDeveloping effective visual analytics systems demands care in characterization of domain problems and integration of visualization techniques and computational models. Urban visual analytics has already achieved remarkable success in tackling urban problems and providing fundamental services for smart cities. To promote further academic research and assist the development of industrial urban analytics systems, we comprehensively review urban visual analytics studies from four perspectives. In particular, we identify 8 urban domains and 22 types of popular visualization, analyze 7 types of computational method, and categorize existing systems into 4 types based on their integration of visualization techniques and computational models. We conclude with potential research directions and opportunities. Zikun Deng, Di Weng, Mingliang Xu 0001, Yingcai Wu |
Comput. Vis. Media | 2 |
| 2023 | Rigel: Transforming Tabular Data by Declarative MappingabstractWe present Rigel, an interactive system for rapid transformation of tabular data. Rigel implements a new declarative mapping approach that formulates the data transformation procedure as direct mappings from data to the row, column, and cell channels of the target table. To construct such mappings, Rigel allows users to directly drag data attributes from input data to these three channels and indirectly drag or type data values in a spreadsheet, and possible mappings that do not contradict these interactions are recommended to achieve efficient and straightforward data transformation. The recommended mappings are generated by enumerating and composing data variables based on the row, column, and cell channels, thereby revealing the possibility of alternative tabular forms and facilitating open-ended exploration in many data transformation scenarios, such as designing tables for presentation. In contrast to existing systems that transform data by composing operations (like transposing and pivoting), Rigel requires less prior knowledge on these operations, and constructing tables from the channels is more efficient and results in less ambiguity than generating operation sequences as done by the traditional by-example approaches. User study results demonstrated that Rigel is significantly less demanding in terms of time and interactions and suits more scenarios compared to the state-of-the-art by-example approach. A gallery of diverse transformation cases is also presented to show the potential of Rigel's expressiveness. Di Weng, Yanwei Huang, Xinhuan Shu, Guodao Sun, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | ECoalVis: Visual Analysis of Control Strategies in Coal-fired Power PlantsabstractImproving the efficiency of coal-fired power plants has numerous benefits. The control strategy is one of the major factors affecting such efficiency. However, due to the complex and dynamic environment inside the power plants, it is hard to extract and evaluate control strategies and their cascading impact across massive sensors. Existing manual and data-driven approaches cannot well support the analysis of control strategies because these approaches are time-consuming and do not scale with the complexity of the power plant systems. Three challenges were identified: a) interactive extraction of control strategies from large-scale dynamic sensor data, b) intuitive visual representation of cascading impact among the sensors in a complex power plant system, and c) time-lag-aware analysis of the impact of control strategies on electricity generation efficiency. By collaborating with energy domain experts, we addressed these challenges with ECoalVis, a novel interactive system for experts to visually analyze the control strategies of coal-fired power plants extracted from historical sensor data. The effectiveness of the proposed system is evaluated with two usage scenarios on a real-world historical dataset and received positive feedback from experts. Di Weng, Zikun Deng, Haoran Xu 0003, Honglei Yin, Xianyuan Zhan, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | Positive emotions help rank negative reviews for sellers and producers in e-commerceabstractNegative reviews, the poor ratings in postpurchase evaluation, play an indispensable role in e-commerce, especially in shaping future sales and firm equities. However, extant studies seldom examine their potential value for sellers and producers in enhancing capabilities of providing better services and products. For those who exploited the helpfulness of reviews in the view of e-commerce keepers, the ranking approaches were developed for customers instead. To fill this gap, in terms of combining description texts and emotion polarities, the aim of the ranking method in this study is to provide the most helpful negative reviews under a certain product attribute for online sellers and producers. By applying a more reasonable evaluating procedure, experts with related backgrounds are hired to vote for the ranking approaches. Our ranking method turns out to be more reliable for ranking negative reviews for sellers and producers, demonstrating a better performance than the baselines like BM25 with a result of 8% higher. In this paper, we also enrich the previous understandings of emotions in valuing reviews. Specifically, it is surprisingly found that positive emotions are more helpful rather than negative emotions in ranking negative reviews. The unexpected strengthening from positive emotions in ranking suggests that less polarized reviews on negative experience in fact offer more rational feedbacks and thus more helpfulness to the sellers and producers. The presented ranking method could provide e-commerce practitioners with an efficient and effective way to leverage negative reviews from online consumers. Di Weng, Yang Yang 0126, Jichang Zhao |
DSAA | 1 |
| 2022 | Nebula: A Coordinating Grammar of GraphicsabstractIn multiple coordinated views (MCVs), visualizations across views update their content in response to users' interactions in other views. Interactive systems provide direct manipulation to create coordination between views, but are restricted to limited types of predefined templates. By contrast, textual specification languages enable flexible coordination but expose technical burden. To bridge the gap, we contribute Nebula, a grammar based on natural language for coordinating visualizations in MCVs. The grammar design is informed by a novel framework based on a systematic review of 176 coordinations from existing theories and applications, which describes coordination by demonstration, i.e., how coordination is performed by users. With the framework, Nebula specification formalizes coordination as a composition of user- and coordination-triggered interactions in origin and destination views, respectively, along with potential data transformation between the interactions. We evaluate Nebula by demonstrating its expressiveness with a gallery of diverse examples and analyzing its usability on cognitive dimensions. Xinhuan Shu, Di Weng, Junxiu Tang, Siwei Fu, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Visual Cascade Analytics of Large-Scale Spatiotemporal DataabstractMany spatiotemporal events can be viewed as contagions. These events implicitly propagate across space and time by following cascading patterns, expanding their influence, and generating event cascades that involve multiple locations. Analyzing such cascading processes presents valuable implications in various urban applications, such as traffic planning and pollution diagnostics. Motivated by the limited capability of the existing approaches in mining and interpreting cascading patterns, we propose a visual analytics system called VisCas. VisCas combines an inference model with interactive visualizations and empowers analysts to infer and interpret the latent cascading patterns in the spatiotemporal context. To develop VisCas, we address three major challenges 1) generalized pattern inference; 2) implicit influence visualization; and 3) multifaceted cascade analysis. For the first challenge, we adapt the state-of-the-art cascading network inference technique to general urban scenarios, where cascading patterns can be reliably inferred from large-scale spatiotemporal data. For the second and third challenges, we assemble a set of effective visualizations to support location navigation, influence inspection, and cascading exploration, and facilitate the in-depth cascade analysis. We design a novel influence view based on a three-fold optimization strategy for analyzing the implicit influences of the inferred patterns. We demonstrate the capability and effectiveness of VisCas with two case studies conducted on real-world traffic congestion and air pollution datasets with domain experts. Zikun Deng, Di Weng, Yuxuan Liang 0002, Jie Bao 0003, Yu Zheng 0004, Tobias Schreck, Mingliang Xu 0001, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | Compass: Towards Better Causal Analysis of Urban Time SeriesabstractThe spatial time series generated by city sensors allow us to observe urban phenomena like environmental pollution and traffic congestion at an unprecedented scale. However, recovering causal relations from these observations to explain the sources of urban phenomena remains a challenging task because these causal relations tend to be time-varying and demand proper time series partitioning for effective analyses. The prior approaches extract one causal graph given long-time observations, which cannot be directly applied to capturing, interpreting, and validating dynamic urban causality. This paper presents Compass, a novel visual analytics approach for in-depth analyses of the dynamic causality in urban time series. To develop Compass, we identify and address three challenges: detecting urban causality, interpreting dynamic causal relations, and unveiling suspicious causal relations. First, multiple causal graphs over time among urban time series are obtained with a causal detection framework extended from the Granger causality test. Then, a dynamic causal graph visualization is designed to reveal the time-varying causal relations across these causal graphs and facilitate the exploration of the graphs along the time. Finally, a tailored multi-dimensional visualization is developed to support the identification of spurious causal relations, thereby improving the reliability of causal analyses. The effectiveness of Compass is evaluated with two case studies conducted on the real-world urban datasets, including the air pollution and traffic speed datasets, and positive feedback was received from domain experts. Zikun Deng, Di Weng, Xiao Xie, Jie Bao 0003, Yu Zheng 0004, Mingliang Xu 0001, Wei Chen 0001, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | A Visualization Approach for Monitoring Order Processing in E-Commerce WarehouseabstractThe efficiency of warehouses is vital to e-commerce. Fast order processing at the warehouses ensures timely deliveries and improves customer satisfaction. However, monitoring, analyzing, and manipulating order processing in the warehouses in real time are challenging for traditional methods due to the sheer volume of incoming orders, the fuzzy definition of delayed order patterns, and the complex decision-making of order handling priorities. In this paper, we adopt a data-driven approach and propose OrderMonitor, a visual analytics system that assists warehouse managers in analyzing and improving order processing efficiency in real time based on streaming warehouse event data. Specifically, the order processing pipeline is visualized with a novel pipeline design based on the sedimentation metaphor to facilitate real-time order monitoring and suggest potentially abnormal orders. We also design a novel visualization that depicts order timelines based on the Gantt charts and Marey's graphs. Such a visualization helps the managers gain insights into the performance of order processing and find major blockers for delayed orders. Furthermore, an evaluating view is provided to assist users in inspecting order details and assigning priorities to improve the processing performance. The effectiveness of OrderMonitor is evaluated with two case studies on a real-world warehouse dataset. Junxiu Tang, Yuhua Zhou, Tan Tang, Di Weng, Boyang Xie, Lingyun Yu 0001, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2021 | Pareto-Optimal Transit Route Planning With Multi-Objective Monte-Carlo Tree SearchabstractPlanning ideal transit routes in the complex urban environment can improve the performance and efficiency of public transportation systems effectively. However, finding such routes is computationally difficult due to the huge solution space constituted by billions of possible routes. Considering the limited scalability of exact search methods, heuristic search methods were proposed to boost the efficiency and incorporate flexible constraints. Nevertheless, the existing methods conceal multiple criteria in an objective, and thus evaluating the performance of the generated route becomes challenging due to the lack of comparable alternatives. Inspired by the prior study, we formulate the definition of pareto-optimal transit routes based on multiple criteria. However, extracting these routes remains challenging because: A) the sheer volume of possible transit routes; and B) the sparsity of pareto-optimal routes. We address these challenges by developing an efficient search framework: for challenge A, a random search method is developed based on Monte Carlo tree search where the unproductive solution subspaces are pruned progressively to reduce the search cost; and for challenge B, an estimation method is derived to guide the search process by assessing the value for each solution subspace. The superior effectiveness of our approach in approximating the pareto-optimal transit routes was demonstrated by the comprehensive evaluation based on the real-world data. Di Weng, Jie Bao 0003, Yu Zheng 0004, Yingcai Wu |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Towards Better Detection and Analysis of Massive Spatiotemporal Co-Occurrence PatternsabstractWith the rapid development of sensing technologies, massive spatiotemporal data have been acquired from the urban space with respect to different domains, such as transportation and environment. Numerous co-occurrence patterns (e.g., traffic speed <; 10km/h, weather = foggy, and air quality = unhealthy) between the transportation data and other types of data can be obtained with given spatiotemporal constraints (e.g., within 3 kilometers and lasting for 2 hours) from these heterogeneous data sources. Such patterns present valuable implications for many urban applications, such as traffic management, pollution diagnosis, and transportation planning. However, extracting and understanding these patterns is beyond manual capability because of the scale, diversity, and heterogeneity of the data. To address this issue, a novel visual analytics system called CorVizor is proposed to identify and interpret these co-occurrence patterns. CorVizor comprises two major components. The first component is a co-occurrence mining framework involving three steps, namely, spatiotemporal indexing, co-occurring instance generation, and pattern mining. The second component is a visualization technique called CorView that implements a level-of-detail mechanism by integrating tailored visualizations to depict the extracted spatiotemporal co-occurrence patterns. The case studies and expert interviews are conducted to demonstrate the effectiveness of CorVizor. Yingcai Wu, Di Weng, Zikun Deng, Jie Bao 0003, Mingliang Xu 0001, Zhangye Wang, Yu Zheng 0004, Zhiyu Ding, Wei Chen 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Towards Better Bus Networks: A Visual Analytics ApproachabstractBus routes are typically updated every 3-5 years to meet constantly changing travel demands. However, identifying deficient bus routes and finding their optimal replacements remain challenging due to the difficulties in analyzing a complex bus network and the large solution space comprising alternative routes. Most of the automated approaches cannot produce satisfactory results in real-world settings without laborious inspection and evaluation of the candidates. The limitations observed in these approaches motivate us to collaborate with domain experts and propose a visual analytics solution for the performance analysis and incremental planning of bus routes based on an existing bus network. Developing such a solution involves three major challenges, namely, a) the in-depth analysis of complex bus route networks, b) the interactive generation of improved route candidates, and c) the effective evaluation of alternative bus routes. For challenge a, we employ an overview-to-detail approach by dividing the analysis of a complex bus network into three levels to facilitate the efficient identification of deficient routes. For challenge b, we improve a route generation model and interpret the performance of the generation with tailored visualizations. For challenge c, we incorporate a conflict resolution strategy in the progressive decision-making process to assist users in evaluating the alternative routes and finding the most optimal one. The proposed system is evaluated with two usage scenarios based on real-world data and received positive feedback from the experts. Index Terms-Bus route planning, spatial decision-making, urban data visual analytics. Di Weng, Chengbo Zheng, Zikun Deng, Mingze Ma, Jie Bao 0003, Yu Zheng 0004, Mingliang Xu 0001, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2020 | AirVis: Visual Analytics of Air Pollution PropagationabstractAir pollution has become a serious public health problem for many cities around the world. To find the causes of air pollution, the propagation processes of air pollutants must be studied at a large spatial scale. However, the complex and dynamic wind fields lead to highly uncertain pollutant transportation. The state-of-the-art data mining approaches cannot fully support the extensive analysis of such uncertain spatiotemporal propagation processes across multiple districts without the integration of domain knowledge. The limitation of these automated approaches motivates us to design and develop AirVis, a novel visual analytics system that assists domain experts in efficiently capturing and interpreting the uncertain propagation patterns of air pollution based on graph visualizations. Designing such a system poses three challenges: a) the extraction of propagation patterns; b) the scalability of pattern presentations; and c) the analysis of propagation processes. To address these challenges, we develop a novel pattern mining framework to model pollutant transportation and extract frequent propagation patterns efficiently from large-scale atmospheric data. Furthermore, we organize the extracted patterns hierarchically based on the minimum description length (MDL) principle and empower expert users to explore and analyze these patterns effectively on the basis of pattern topologies. We demonstrated the effectiveness of our approach through two case studies conducted with a real-world dataset and positive feedback from domain experts. Zikun Deng, Di Weng, Jie Bao 0003, Yu Zheng 0004, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2019 | SRVis: Towards Better Spatial Integration in Ranking VisualizationabstractInteractive ranking techniques have substantially promoted analysts' ability in making judicious and informed decisions effectively based on multiple criteria. However, the existing techniques cannot satisfactorily support the analysis tasks involved in ranking large-scale spatial alternatives, such as selecting optimal locations for chain stores, where the complex spatial contexts involved are essential to the decision-making process. Limitations observed in the prior attempts of integrating rankings with spatial contexts motivate us to develop a context-integrated visual ranking technique. Based on a set of generic design requirements we summarized by collaborating with domain experts, we propose SRVis, a novel spatial ranking visualization technique that supports efficient spatial multi-criteria decision-making processes by addressing three major challenges in the aforementioned context integration, namely, a) the presentation of spatial rankings and contexts, b) the scalability of rankings' visual representations, and c) the analysis of context-integrated spatial rankings. Specifically, we encode massive rankings and their cause with scalable matrix-based visualizations and stacked bar charts based on a novel two-phase optimization framework that minimizes the information loss, and the flexible spatial filtering and intuitive comparative analysis are adopted to enable the in-depth evaluation of the rankings and assist users in selecting the best spatial alternative. The effectiveness of the proposed technique has been evaluated and demonstrated with an empirical study of optimization methods, two case studies, and expert interviews. Di Weng, Zikun Deng, Feiran Wu, Jingmin Chen, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2018 | HomeFinder Revisited: Finding Ideal Homes with Reachability-Centric Multi-Criteria Decision MakingabstractFinding an ideal home is a difficult and laborious process. One of the most crucial factors in this process is the reachability between the home location and the concerned points of interest, such as places of work and recreational facilities. However, such importance is unrecognized in the extant real estate systems. By characterizing user requirements and analytical tasks in the context of finding ideal homes, we designed ReACH, a novel visual analytics system that assists people in finding, evaluating, and choosing a home based on multiple criteria, including reachability. In addition, we developed an improved data-driven model for approximating reachability with massive taxi trajectories. This model enables users to interactively integrate their knowledge and preferences to make judicious and informed decisions. We show the improvements in our model by comparing the theoretical complexities with the prior study and demonstrate the usability and effectiveness of the proposed system with task-based evaluation. Di Weng, Heming Zhu, Jie Bao 0003, Yu Zheng 0004, Yingcai Wu |
CHI | 1 |
| 2017 | SmartAdP: Visual Analytics of Large-scale Taxi Trajectories for Selecting Billboard LocationsabstractThe problem of formulating solutions immediately and comparing them rapidly for billboard placements has plagued advertising planners for a long time, owing to the lack of efficient tools for in-depth analyses to make informed decisions. In this study, we attempt to employ visual analytics that combines the state-of-the-art mining and visualization techniques to tackle this problem using large-scale GPS trajectory data. In particular, we present SmartAdP, an interactive visual analytics system that deals with the two major challenges including finding good solutions in a huge solution space and comparing the solutions in a visual and intuitive manner. An interactive framework that integrates a novel visualization-driven data mining model enables advertising planners to effectively and efficiently formulate good candidate solutions. In addition, we propose a set of coupled visualizations: a solution view with metaphor-based glyphs to visualize the correlation between different solutions; a location view to display billboard locations in a compact manner; and a ranking view to present multi-typed rankings of the solutions. This system has been demonstrated using case studies with a real-world dataset and domain-expert interviews. Our approach can be adapted for other location selection problems such as selecting locations of retail stores or restaurants using trajectory data. Dongyu Liu, Di Weng, Jie Bao 0003, Yu Zheng 0004, Huamin Qu, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 2 |