VLDB 2026 Research / reviewers in the wild / expert
Yating Wei
dblp:221/4329
· DBLP profile ↗
11ranked-venue papers
2as first author
9since 2021 · last 2024
0000-0003-0743-7558ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A visual analysis approach for data imputation via multi-party tabular data correlation strategiesabstractData imputation is an essential pre-processing task for data governance, aimed at filling in incomplete data. However, conventional data imputation methods can only partly alleviate data incompleteness using isolated tabular data, and they fail to achieve the best balance between accuracy and efficiency. In this paper, we present a novel visual analysis approach for data imputation. We develop a multi-party tabular data association strategy that uses intelligent algorithms to identify similar columns and establish column correlations across multiple tables. Then, we perform the initial imputation of incomplete data using correlated data entries from other tables. Additionally, we develop a visual analysis system to refine data imputation candidates. Our interactive system combines the multi-party data imputation approach with expert knowledge, allowing for a better understanding of the relational structure of the data. This significantly enhances the accuracy and efficiency of data imputation, thereby enhancing the quality of data governance and the intrinsic value of data assets. Experimental validation and user surveys demonstrate that this method supports users in verifying and judging the associated columns and similar rows using their domain knowledge. Dongming Han, Jiacheng Pan, Yating Wei, Yingchaojie Feng, Luoxuan Weng, Ketian Mao, Yuankai Xing, Jianshu Lv, Qiucheng Wan, Wei Chen 0001 |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2024 | Erratum to: A visual analysis approach for data imputation via multi-party tabular data correlation strategies
Dongming Han, Jiacheng Pan, Yating Wei, Yingchaojie Feng, Luoxuan Weng, Ketian Mao, Yuankai Xing, Jianshu Lv, Qiucheng Wan, Wei Chen 0001 |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2024 | A visual analysis approach for data transformation via domain knowledge and intelligent models
Chengcan Chu, Minfeng Zhu 0001, Yating Wei, Jiacheng Pan, Dongming Han, Xuwei Tan, Wei Chen 0001 |
Multim. Syst. | 5 |
| 2024 | Towards Efficient Visual Simplification of Computational Graphs in Deep Neural NetworksabstractA computational graph in a deep neural network (DNN) denotes a specific data flow diagram (DFD) composed of many tensors and operators. Existing toolkits for visualizing computational graphs are not applicable when the structure is highly complicated and large-scale (e.g., BERT (Devlin et al. 2019)). To address this problem, we propose leveraging a suite of visual simplification techniques, including a cycle-removing method, a module-based edge-pruning algorithm, and an isomorphic subgraph stacking strategy. We design and implement an interactive visualization system that is suitable for computational graphs with up to 10 thousand elements. Experimental results and usage scenarios demonstrate that our tool reduces 60% elements on average and hence enhances the performance for recognizing and diagnosing DNN models. Our contributions are integrated into an open-source DNN visualization toolkit, namely, MindInsight [2]. Rusheng Pan, Yating Wei, Han Gao 0016, Gongchang Ou, Caleb Chen Cao, Jingli Xu, Tong Xu 0001, Wei Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Visual Diagnostics of Parallel Performance in Training Large-Scale DNN ModelsabstractDiagnosing the cluster-based performance of large-scale deep neural network (DNN) models during training is essential for improving training efficiency and reducing resource consumption. However, it remains challenging due to the incomprehensibility of the parallelization strategy and the sheer volume of complex data generated in the training processes. Prior works visually analyze performance profiles and timeline traces to identify anomalies from the perspective of individual devices in the cluster, which is not amenable for studying the root cause of anomalies. In this article, we present a visual analytics approach that empowers analysts to visually explore the parallel training process of a DNN model and interactively diagnose the root cause of a performance issue. A set of design requirements is gathered through discussions with domain experts. We propose an enhanced execution flow of model operators for illustrating parallelization strategies within the computational graph layout. We design and implement an enhanced Marey's graph representation, which introduces the concept of time-span and a banded visual metaphor to convey training dynamics and help experts identify inefficient training processes. We also propose a visual aggregation technique to improve visualization efficiency. We evaluate our approach using case studies, a user study and expert interviews on two large-scale models run in a cluster, namely, the PanGu- α 13B model (40 layers), and the Resnet model (50 layers). Yating Wei, Gongchang Ou, Han Gao 0016, Caleb Chen Cao, Luoxuan Weng, Jiaying Lu 0005, Rongchen Zhu, Wei Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | Interactive visual analytics of parallel training strategies for DNN models
Yating Wei, Gongchang Ou, Han Gao 0016, Minfeng Zhu 0001, Wei Chen 0001 |
Comput. Graph. | 2 |
| 2023 | VIS+AI: integrating visualization with artificial intelligence for efficient data analysisabstractAbstract Visualization and artificial intelligence (AI) are well-applied approaches to data analysis. On one hand, visualization can facilitate humans in data understanding through intuitive visual representation and interactive exploration. On the other hand, AI is able to learn from data and implement bulky tasks for humans. In complex data analysis scenarios, like epidemic traceability and city planning, humans need to understand large-scale data and make decisions, which requires complementing the strengths of both visualization and AI. Existing studies have introduced AI-assisted visualization as AI4VIS and visualization-assisted AI as VIS4AI. However, how can AI and visualization complement each other and be integrated into data analysis processes are still missing. In this paper, we define three integration levels of visualization and AI. The highest integration level is described as the framework of VIS+AI, which allows AI to learn human intelligence from interactions and communicate with humans through visual interfaces. We also summarize future directions of VIS+AI to inspire related studies. Xumeng Wang, Ziliang Wu, Wenqi Huang 0002, Yating Wei, Zhaosong Huang, Mingliang Xu 0001, Wei Chen 0001 |
Frontiers Comput. Sci. | 4 |
| 2023 | Federated Visualization: A Privacy-Preserving Strategy for Aggregated Visual QueryabstractWe present a novel privacy preservation strategy for aggregated visual query of decentralized data. The key idea is to imitate the flowchart of the federated learning framework, and reformulate the visualization process within a federated infrastructure. The federation of visualization is fulfilled by leveraging a shared global module that composes the encrypted externalizations of transformed visual features of data pieces in local modules. We design two implementations of federated visualization: a prediction-based scheme, and a query-based scheme. We demonstrate the effectiveness of our approach with a set of visual forms, and verify its robustness with evaluations. We report the value of federated visualization in real scenarios with an expert review. Wei Chen 0001, Yating Wei, Shuyue Zhou, Bingru Lin, Zhiguang Zhou |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | VADAF: Visualization for Abnormal Client Detection and Analysis in Federated LearningabstractFederated Learning (FL) provides a powerful solution to distributed machine learning on a large corpus of decentralized data. It ensures privacy and security by performing computation on devices (which we refer to as clients) based on local data to improve the shared global model. However, the inaccessibility of the data and the invisibility of the computation make it challenging to interpret and analyze the training process, especially to distinguish potential client anomalies. Identifying these anomalies can help experts diagnose and improve FL models. For this reason, we propose a visual analytics system, VADAF, to depict the training dynamics and facilitate analyzing potential client anomalies. Specifically, we design a visualization scheme that supports massive training dynamics in the FL environment. Moreover, we introduce an anomaly detection method to detect potential client anomalies, which are further analyzed based on both the client model’s visual and objective estimation. Three case studies have demonstrated the effectiveness of our system in understanding the FL training process and supporting abnormal client detection and analysis. Linhao Meng, Yating Wei, Rusheng Pan, Shuyue Zhou, Jianwei Zhang 0015, Wei Chen 0001 |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2020 | RSATree: Distribution-Aware Data Representation of Large-Scale Tabular Datasets for Flexible Visual QueryabstractAnalysts commonly investigate the data distributions derived from statistical aggregations of data that are represented by charts, such as histograms and binned scatterplots, to visualize and analyze a large-scale dataset. Aggregate queries are implicitly executed through such a process. Datasets are constantly extremely large; thus, the response time should be accelerated by calculating predefined data cubes. However, the queries are limited to the predefined binning schema of preprocessed data cubes. Such limitation hinders analysts' flexible adjustment of visual specifications to investigate the implicit patterns in the data effectively. Particularly, RSATree enables arbitrary queries and flexible binning strategies by leveraging three schemes, namely, an R-tree-based space partitioning scheme to catch the data distribution, a locality-sensitive hashing technique to achieve locality-preserving random access to data items, and a summed area table scheme to support interactive query of aggregated values with a linear computational complexity. This study presents and implements a web-based visual query system that supports visual specification, query, and exploration of large-scale tabular data with user-adjustable granularities. We demonstrate the efficiency and utility of our approach by performing various experiments on real-world datasets and analyzing time and space complexity. Honghui Mei, Wei Chen 0001, Yating Wei, Shuyue Zhou, Bingru Lin, Ying Zhao 0001, Jiazhi Xia |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2020 | Evaluating Perceptual Bias During Geometric Scaling of ScatterplotsabstractScatterplots are frequently scaled to fit display areas in multi-view and multi-device data analysis environments. A common method used for scaling is to enlarge or shrink the entire scatterplot together with the inside points synchronously and proportionally. This process is called geometric scaling. However, geometric scaling of scatterplots may cause a perceptual bias, that is, the perceived and physical values of visual features may be dissociated with respect to geometric scaling. For example, if a scatterplot is projected from a laptop to a large projector screen, then observers may feel that the scatterplot shown on the projector has fewer points than that viewed on the laptop. This paper presents an evaluation study on the perceptual bias of visual features in scatterplots caused by geometric scaling. The study focuses on three fundamental visual features (i.e., numerosity, correlation, and cluster separation) and three hypotheses that are formulated on the basis of our experience. We carefully design three controlled experiments by using well-prepared synthetic data and recruit participants to complete the experiments on the basis of their subjective experience. With a detailed analysis of the experimental results, we obtain a set of instructive findings. First, geometric scaling causes a bias that has a linear relationship with the scale ratio. Second, no significant difference exists between the biases measured from normally and uniformly distributed scatterplots. Third, changing the point radius can correct the bias to a certain extent. These findings can be used to inspire the design decisions of scatterplots in various scenarios. Yating Wei, Honghui Mei, Ying Zhao 0001, Shuyue Zhou, Bingru Lin, Haojing Jiang, Wei Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |