EDBT 2026 Demo / reviewers in the wild / expert
Zhaokang Wang
dblp:155/5427
· DBLP profile ↗
19ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-8123-9018ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scaling Inter-procedural Dataflow Analysis on the CloudabstractApart from forming the backbone of compiler optimization, static dataflow analysis has been widely applied in a vast variety of applications, such as bug detection, privacy analysis, and program comprehension. Despite its importance, performing inter-procedural dataflow analysis on large-scale programs is well-known to be challenging. In this article, we propose a novel distributed analysis framework supporting the general inter-procedural dataflow analysis. Inspired by large-scale graph processing, we devise dedicated distributed worklist algorithms for both whole-program analysis and incremental analysis. We implement these algorithms and develop a distributed framework called BigDataflow running on a large-scale cluster. The experimental results validate the promising performance of BigDataflow—BigDataflow can finish analyzing the program of million lines of code in minutes. Compared with the state-of-the-art, BigDataflow achieves much more analysis efficiency. Zewen Sun, Duanchen Xu, Yiyu Zhang, Yun Qi, Zhaokang Wang, Yue Li 0006, Xuandong Li, Qingda Lu, Wenwen Peng, Shengjian Guo, Zhiqiang Zuo 0002 |
ACM Trans. Program. Lang. Syst. | 7 |
| 2025 | HSC: Scalable Task Scheduling in Large-Scale Edge Environments
Zhaokang Wang, Yanchao Zhao |
NPC (1) | 2 |
| 2024 | Fluid-Shuttle: Efficient Cloud Data Transmission Based on Serverless Computing CompressionabstractNowadays, there exists a lot of cross-region data transmission demand on the cloud. It is promising to use serverless computing for data compressing to save the total data size. However, it is challenging to estimate the data transmission time and monetary cost with serverless compression. In addition, minimizing the data transmission cost is non-trivial due to the enormous parameter space. This paper focuses on this problem and makes the following contributions: 1) We propose empirical data transmission time and monetary cost models based on serverless compression. It can also predict compression information, e.g., ratio and speed using chunk sampling and machine learning techniques. 2) For single-task cloud data transmission, we propose two efficient parameter search methods based on Sequential Quadratic Programming (SQP) and Eliminate then Divide and Conquer (EDC) with proven error upper bounds. Besides, we propose a parameter fine-tuning strategy to deal with transmission bandwidth variance. 3) Furthermore, for multi-task scenarios, a parameter search method based on dynamic programming and numerical computation is proposed. We have implemented the system called Fluid-Shuttle, which includes straggler optimization, cache optimization, and the autoscaling decompression mechanism. Finally, we evaluate the performance of Fluid-Shuttle with various workloads and applications on the real-world AWS serverless computing platform. Experimental results show that the proposed approach can improve the parameter search efficiency by over$3\times $compared with the state-of-art methods and achieves better parameter quality. In addition, our approach achieves higher time efficiency and lower monetary cost compared with competing cloud data transmission approaches. Rong Gu 0001, Shulin Wang, Haipeng Dai 0001, Zhaokang Wang, Wenjie Bao, Jiaqi Zheng 0001, Yaofeng Tu, Yihua Huang 0001, Lianyong Qi, Xiaolong Xu 0001, Wan-Chun Dou, Guihai Chen |
IEEE/ACM Trans. Netw. | 5 |
| 2023 | Time and Cost-Efficient Cloud Data Transmission based on Serverless Computing Compression
Rong Gu 0001, Haipeng Dai 0001, Shulin Wang, Zhaokang Wang, Yaofeng Tu, Yihua Huang 0001, Guihai Chen |
INFOCOM | 5 |
| 2023 | BigDataflow: A Distributed Interprocedural Dataflow Analysis FrameworkabstractAbstract: Apart from forming the backbone of compiler optimization, static dataflow analysis has been widely applied in a vast variety of applications, such as bug detection, privacy analysis, program comprehension, etc. Despite its importance, performing interprocedural dataflow analysis on large-scale programs is well known to be challenging.In this paper, we propose a novel distributed analysis framework supporting the general interprocedural dataflow analysis.Inspired by large-scale graph processing, we devise a dedicated distributed worklist algorithm tailored for interprocedural dataflow analysis. We implement the algorithm and develop a distributed framework called BigDataflow running on a large-scale cluster.The experimental results validate the promising performance of BigDataflow – it can finish analyzing the program of millions lines of code in minutes. Compared with the state-of-the-art, BigDataflow achieves much more analysis efficiency. Zewen Sun, Duanchen Xu, Yiyu Zhang, Yun Qi, Zhiqiang Zuo 0002, Zhaokang Wang, Yue Li 0006, Xuandong Li, Qingda Lu, Wenwen Peng, Shengjian Guo |
ESEC/SIGSOFT FSE | 7 |
| 2023 | Coral: federated query join order optimization based on deep reinforcement learning
Rong Gu 0001, Liangliang Yin, Lingyi Song, Chunfeng Yuan, Zhaokang Wang, Yihua Huang 0001 |
World Wide Web (WWW) | 7 |
| 2022 | Octopus-DF: Unified DataFrame-based cross-platform data analytic system
Rong Gu 0001, Zhaokang Wang, Yang Che, Yihua Huang 0001 |
Parallel Comput. | 4 |
| 2021 | UniGPS: A Unified Programming Framework for Distributed Graph ProcessingabstractThe industry and academia have proposed many distributed graph processing systems. However, the existing systems are not friendly enough for users like data analysts and algorithm engineers. On the one hand, the programming models and interfaces differ a lot in the existing systems, leading to high learning costs and program migration costs. On the other hand, these graph processing systems are tightly bound to the underlying distributed computing platforms, requiring users to be familiar with distributed computing. To improve the usability of distributed graph processing, we propose a unified distributed graph programming framework UniGPS. Firstly, we propose a unified cross-platform graph programming model VCProg for UniGPS. VCProg hides details of distributed computing from users. It is compatible with the popular graph programming models Pregel, GAS, and Push-Pull. VCProg-based programs can be executed by compatible distributed graph processing systems without modification, reducing the learning overheads of users. Secondly, UniGPS supports Python as the programming language. We propose an interprocess-communication-based execution environment isolation mechanism to enable Java/C++-based graph processing systems to call user-defined methods written in Python. The experimental results show that UniGPS enables users to process big graphs beyond the memory capacity of a single machine without sacrificing usability. UniGPS shows near-linear data scalability and machine scalability. Zhaokang Wang, Yifan Qi, Chunfeng Yuan, Yihua Huang 0001 |
ICPADS | 1 |
| 2021 | Empirical analysis of performance bottlenecks in graph neural network training and inference with GPUs
Zhaokang Wang, Yunpan Wang, Chunfeng Yuan, Rong Gu 0001, Yihua Huang 0001 |
Neurocomputing | 1 |
| 2021 | SparkDQ: Efficient generic big data quality management on distributed data-parallel computation
Rong Gu 0001, Zhaokang Wang, Xiaolong Xu 0001, Chunfeng Yuan, Yihua Huang 0001 |
J. Parallel Distributed Comput. | 4 |
| 2021 | VSIM: Distributed local structural vertex similarity calculation on big graphs
Zhaokang Wang, Chunfeng Yuan, Rong Gu 0001, Yihua Huang 0001 |
J. Parallel Distributed Comput. | 1 |
| 2021 | Alchemy: Distributed financial quantitative analysis system with high-level programming modelabstractAbstract Nowadays, the financial securities investment and transaction are processed digitally. The securities companies have accumulated a large amount of financial data. Given that, quantitative analysis has become a common method for securities investors. However, as the scale of data continues to grow, data storing and processing have increasingly become a big challenge to financial quantitative analysts. First, most existing stand‐alone quantitative analysis systems are hard to accelerate data processing in a distributed way. Second, existing distributed computing frameworks such as Apache Spark demands financial quantitative analysts with expertise knowledge on distributed computing. Third, transition from traditional stand‐alone financial quantitative analysis (FQA) system to such distributed computing system often introduces learning efforts and development overheads. To solve these problems, we propose Alchemy, a distributed processing platform tailored for FQA with high‐level programming interface. Alchemy offers distributed computing service and provides a pipeline‐based programming model which shares many similarities with the traditional stand‐alone systems, reducing learning efforts for financial quantitative analysts. The performance of Alchemy is evaluated in both experimental environment and a real‐world production environment in Huatai Securites which is a leading financial company in China. The results show that Alchemy is able to achieve up to 300 times speedup compared against existing financial quantitative analysis systems. Rong Gu 0001, Zhaokang Wang, Chunfeng Yuan, Yihua Huang 0001 |
Softw. Pract. Exp. | 4 |
| 2021 | Towards Efficient Large-Scale Interprocedural Program Static Analysis on Distributed Data-Parallel ComputationabstractStatic program analysis has been widely applied along the whole process of the program development for bug detection, code optimization, testing, etc. Although researchers have made significant work in static program analysis, it is still challenging to perform sophisticated interprocedural analysis on large-scale modern software. The underlying reason is that interprocedural analysis for large-scale modern software is highly computation- and memory-intensive, leading to poor efficiency and scalability. In this article, we introduce an efficient distributed and scalable solution for sophisticated static analysis. Specifically, we propose a data-parallel algorithm and a join-process-filter computation model for the CFL-reachability-based interprocedural analysis. Based on that, an efficient distributed static analysis engine called BigSpa is developed, which is composed of an offline batch static program analysis system and an online incremental static program analysis system. The BigSpa system has high generality and can support all kinds of static analysis tasks that can be expressed as CFL reachability problems. The performance of BigSpa is evaluated on real-world large-scale software datasets. Our experiments show that the offline batch system can exceed an order of magnitude compared with the most advanced analysis tools available on performance, and for incremental analysis with small batch updates on the same data sets, the online analysis system can achieve near real-time response, which is very fast and flexible. Rong Gu 0001, Zhiqiang Zuo 0002, Han Yin, Zhaokang Wang, Linzhang Wang, Xuandong Li, Yihua Huang 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2021 | Towards Efficient Distributed Subgraph Enumeration Via Backtracking-Based FrameworkabstractFinding or monitoring subgraph instances that are isomorphic to a given pattern graph in a data graph is a fundamental query operation in many graph analytic applications, such as network motif mining and fraud detection. Existing distributed methods are inefficient in communication. They have to shuffle partial matching results during the distributed multiway join. The partial matching results may be much larger than the data graph itself. To overcome the drawback, we develop the Batch-BENU framework for distributed subgraph enumeration on static data graphs. Batch-BENU executes a group of local search tasks in parallel. Each task enumerates subgraphs around a vertex in the data graph, guided by a backtracking-based execution plan. To handle large-scale data graphs that may exceed the memory capacity of a single machine, Batch-BENU stores the data graph in a distributed database. Each task queries adjacency sets of the data graph on demand, shuffling the data graph instead of partial matching results. To support incremental subgraph enumeration on dynamic data graphs, we propose the Streaming-BENU framework. Streaming-BENU turns the problem of enumerating incremental matching results into enumerating all matching results of incremental pattern graphs at each time step. We implement Batch-BENU and Streaming-BENU with the local database cache and the load balance optimization to improve their efficiency. Extensive experiments show that Batch-BENU and Streaming-BENU can scale to big graphs and complex pattern graphs. They outperform the state-of-the-art distributed methods by up to one and two orders of magnitude, respectively. Zhaokang Wang, Guowang Chen, Chunfeng Yuan, Rong Gu 0001, Yihua Huang 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | BENU: Distributed Subgraph Enumeration with Backtracking-Based FrameworkabstractGiven a small pattern graph and a large data graph, the task of subgraph enumeration is to find all the subgraphs of the data graph that are isomorphic to the pattern graph. The state-of-the-art distributed algorithms like SEED and CBF turn subgraph enumeration into a distributed multi-way join problem. They are inefficient in communication as they have to shuffle partial matching results that are much larger than the data graph itself during the join. They also spend non-trivial costs on constructing indexes for data graphs. Different from those join-based algorithms, we develop a new backtracking-based framework BENU for distributed subgraph enumeration. BENU divides a subgraph enumeration task into a group of local search tasks that can be executed in parallel. Each local search task follows a backtracking-based execution plan to enumerate subgraphs. The data graph is stored in a distributed database and is queried as needed. BENU only queries the necessary edges of the data graph and avoids shuffling partial matching results. We also develop an efficient implementation for BENU. We set up an in-memory database cache on each machine. Taking advantage of the inter-task and intra-task locality, the cache significantly reduces the communication cost with controllable memory usage. We conduct extensive experiments to evaluate the performance of BENU. The results show that BENU is scalable and outperforms the state-of-the-art methods by up to an order of magnitude. Zhaokang Wang, Rong Gu 0001, Chunfeng Yuan, Yihua Huang 0001 |
ICDE | 1 |
| 2019 | BigSpa: An Efficient Interprocedural Static Analysis Engine in the CloudabstractStatic program analysis is widely used in various application areas to solve many practical problems. Although researchers have made significant achievements in static analysis, it is still too challenging to perform sophisticated interprocedural analysis on large-scale modern software. The underlying reason is that interprocedural analysis for large-scale modern software is highly computation- and memory-intensive, leading to poor scalability. We aim to tackle the scalability problem by proposing a novel big data solution for sophisticated static analysis. Specifically, we propose a data-parallel algorithm and a join-process-filter computation model for the CFL-reachability based interprocedural analysis and develop an efficient distributed static analysis engine in the cloud, called BigSpa. Our experiments validated that BigSpa running on a cluster scales greatly to perform precise interprocedural analyses on millions of lines of code, and runs an order of magnitude or more faster than the existing state-of-the-art analysis tools. Zhiqiang Zuo 0002, Rong Gu 0001, Zhaokang Wang, Yihua Huang 0001, Linzhang Wang, Xuandong Li |
IPDPS | 4 |
| 2018 | Penguin: Efficient Query-Based Framework for Replaying Large Scale Historical DataabstractIn the big data era, there are many demands for efficient and easy to use data replay services over large scale historical data. For example, stock security trading and on-line e-business services need historical data replay services to conduct system testing or ex-post review and analysis, just like replaying videos for security monitoring. Stream processing systems are designed for processing stream data, thus can not perform complex replay jobs over the static historical data. Database management systems support easy to use complex queries, but lack stream processing abilities. In this paper, we present a data replay model that combines the stream replay and complex query ability together, to allow the applications to replay large scale historical data from various data sources. First, to meet the demands of flexible replay semantics, we designed a set of easy to use replay operators to describe various replay behaviors and semantics. Users can use these operators to build up their complex replay jobs with diversified requirements. Then, we proposed a query mechanism to provide a flexible data loading service. Next, we presented the Penguin framework to support the proposed query-based replay model along with the replay operators. Penguin enables users to develop high-throughput and easy to use replay services over various large scale data sources with tunable replay speeds. To further improve the data replay performance, we proposed four system-level optimizations, including caching loading task results, cascading merging intermediate record queues, producing the replay queue in parallel, and caching remote file streams. Experimental results over replaying millions of records demonstrate that Penguin can achieve up to 4x and 144x speedup in data preparation and up to 16x and 9x speedup in replay speed compared to Apache Phoenix and Apache Hive respectively. As a case study, Penguin has been deployed in production environments of some Securities companies to provide online historical stock data replay services to large number of stock market users. Rong Gu 0001, Zhaokang Wang, Chunfeng Yuan, Yihua Huang 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2015 | Efficient large scale distributed matrix computation with sparkabstractMatrix computation is the core of many massive data-intensive analytical applications such mining social networks, recommendation systems and nature language processing. Due to the importance of matrix computation, it has been widely studied for many years. In the Big Data ear, as the scale of the matrix grows, traditional single-node matrix computation systems can hardly cope with such large data and computation. Existing distributed matrix computation solutions are still not efficient enough, or have poor fault tolerance and usability. In this paper, we propose Marlin, an efficient distributed matrix computation library which is built on top of Spark. Marlin contains several distributed matrix operation algorithms and provides high-level matrix computation primitives for users. In Marlin, we proposed three distributed matrix multiplication algorithms for different situations. Based on this, we designed an adaptive model to choose the best approach for different problems. Moreover, to improve the computation performance, instead of naively using Spark, we put forward some optimizations including taking advantage of the native linear algebra library, reducing shuffle communication and increasing parallelism. Experimental results show that Marlin is over an order of magnitude faster than R (a widely-used statistical computing system) and the existing distributed matrix operation algorithms based on MapReduce. Moreover, Marlin achieves comparable performance to the specialized MPI-based matrix multiplication algorithm SUMMA but uses a general dataflow engine and gains common dataflow features such as scalability and fault tolerance. Rong Gu 0001, Zhaokang Wang, Xusen Yin, Chunfeng Yuan, Yihua Huang 0001 |
IEEE BigData | 3 |
| 2015 | iPLAR: Towards Interactive Programming with Parallel Linear Algebra in R
Zhaokang Wang, Shiqing Fan, Rong Gu 0001, Chunfeng Yuan, Yihua Huang 0001 |
ICA3PP (4) | 1 |