VLDB 2026 Research / reviewers in the wild / expert
Zheheng Liang
dblp:300/5689
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2025
0009-0007-7656-2092ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FreewayML: An Adaptive and Stable Streaming Learning Framework for Dynamic Data StreamsabstractStreaming (machine) learning (SML) can capture dynamic changes in real-time data and perform continuous updates. It has been widely applied in real-world scenarios such as network security, financial regulation, and energy supply. How-ever, due to the sensitivity and lightweight nature of SML models, existing work suffers from low robustness, sudden decline, and catastrophic forgetting when facing unexpected data distribution drifts. Previous studies have attempted to enhance the stability of SML through methods such as data selection, replay, and constraints. However, these methods are typically designed for specific feature spaces and specific ML algorithms. In this paper, we introduce a shift graph based on the distances between data distributions and define three distinct data shift patterns. For these three patterns, we design three adaptive mechanisms, (a) multi-time granularity models, (b) coherent experience clustering, and (c) historical knowledge reuse, that are triggered by a strategy selector, with the goal of enhancing the accuracy and stability of SML. We implement an adaptive and stable SML framework, FreewayML, on top of PyTorch, which is suitable for most SML models. Experimental results show that FreewayML significantly outperforms existing SML systems in both stability and accuracy, with a comparable throughput and latency. Zheheng Liang, Lijie Xu, Wentao Wu 0001, Mingchao Wu, Wuqiang Shen, Wei Wang 0009 |
ICDE | 2 |
| 2025 | Autolink: An Adaptive High-Throughput Streaming Processing System for Distributed EnvironmentabstractStreaming processing is widely applied in fields such as smart grids and sensor detection, where Flink has become the de facto industry standard. Since stream data exhibits dynamic speeds and computing clusters possess uneven computational capabilities, Flink faces challenges in maintaining adequate single-node performance and cluster scalability. Consequently, it struggles to process high-speed real-time data efficiently. To address these challenges, this paper first proposes a fine-grained multi-level watermark mechanism to minimize unnecessary out-of-order processing. Furthermore, dynamic load distribution strategy tailored to diverse distribution characteristics is designed. Additionally, we introduce a demand-resource awareness mechanism that adaptively selects the optimal strategy by monitoring data distribution patterns and cluster resource availability. Based on these designs, we implemented a prototype system called Autolink on the Flink framework. Experimental evaluations using typical datasets demonstrate that Autolink achieves more than 3.2×single-node performance and enhances cluster scalability with a 2.0× speedup ratio. Zheheng Liang, Jijun Zeng, Yingwei Liang |
JCC | 1 |
| 2024 | GraphFlow: A Fast and Accurate Distributed Streaming Graph Computation ModelabstractStreaming graph computation has been widely applied in many fields, e.g., social network analysis and online product recommendation. However, existing streaming graph computation approaches still present limitations on accuracy and efficiency. To improve the accuracy, some distributed systems use the sequential graph update method based on an incremental computation model. However, these systems cannot handle the dynamic graph update concurrently. The speculation-based parallel updating model can parallelize the graph computation, however, it is restricted due to ignoring the original messages when updating a graph. Streaming graph computation usually requires high accuracy and low latency. As such, it is challenging to utilize incremental computation while simultaneously supplying concurrent processing guarantees.To overcome these challenges, in this paper, we first analyze a number of classical graph algorithms and summarize three principles that graph algorithms should satisfy in streaming scenarios. Based on these principles, we propose GraphFlow, a streaming graph computation model. GraphFlow achieves fast and accurate computation by utilizing incremental state update and propagation. To reduce the impact of concurrent update conflicts, GraphFlow provides a fine-grained lock based parallel update strategy. We implement GraphFlow framework and evaluate its performance and concurrent update conflict probability on real-world datasets. Meanwhile, we compare GraphFlow with two existing representative graph processing systems. Experimental results show GraphFlow achieves low latency and outperforms other graph processing systems given large datasets. Zheheng Liang, Yingying Zheng, Chaosheng Yao, Jiayan Wang, Lijie Xu, Shuping Ji, Wei Wang 0049, Shikai Duan |
ICPADS | 1 |
| 2023 | Performance Diagnosis for Microservice-Based Systems via Intra-/Inter-Trace AnalysisabstractDiagnosing performance issues is a slow and labor-intensive process, especially for modern complex microservice systems. In this paper, we propose a novel performance diagnosis framework for microservice application based on intra-/inter-trace analysis. For a slow request to be diagnosed, our approach first determines whether the anomaly is caused by on-path service by aggregating and comparing normal/abnormal traces. Considering that some performance problems may actually be caused by off-path services which do not lie in the path of the abnormal traces, our approach further identifies traces which have temporal relation with the abnormal trace by proposed inter-trace analysis technique, and locates the root cause based on traffic analysis. Zheheng Liang, Guoquan Wu, Zhenyue Long |
APSEC | 1 |
| 2023 | A Reinforcement Learning Approach to Generating Test Cases for Web ApplicationsabstractWeb applications play an important role in modern society. Quality assurance of web applications requires lots of manual efforts. In this paper, we propose WebQT, an automatic test case generator for web applications based on reinforcement learning. Specifically, to increase testing efficiency, we design a new reward model, which encourages the agent to mimic human testers to interact with the web applications. To alleviate the problem of state redundancy, we further propose a novel state abstraction technique, which can identify different web pages with the same functionality as the same state, and yields a simplified state space. We evaluate WebQT on seven open-source web applications. The experimental results show that WebQT achieves 45.4% more code coverage along with higher efficiency than the state-of-the-art technique. In addition, WebQT also reveals 69 exceptions in 11 real-world web applications. Xiaoning Chang, Zheheng Liang, Zhenyue Long, Guoquan Wu, Yu Gao 0002, Wei Chen 0018, Jun Wei 0001, Tao Huang 0001 |
AST | 2 |
| 2023 | Sound Predictive Fuzzing for Multi-threaded ProgramsabstractDeveloping correct multi-threaded programs is challenging and concurrency bugs can be easily introduced. Many of them, known as concurrency vulnerabilities, can be exploited to launch attacks. Fuzzing is shown to be a practical and effective technique to expose vulnerabilities. However, existing works on fuzzing concurrency vulnerabilities almost all follow the framework (like AFL++) designed for fuzzing sequential vulnerabilities. Unlike sequential vulnerabilities, concurrency ones cannot be easily triggered. Concurrency vulnerabilities rely on both inputs and thread interleaving to be exposed while existing fuzzing techniques mainly focus on how to generate effective inputs. We present a new framework based on an existing fuzzing technique, AFL++, to integrate the predictive techniques for effective concurrency vulnerability detection. For every input (the original and the mutated ones), we call a predictive tool such that, even if a concurrency vulnerability is not really triggered, it can be predicted. To overcome heavy efficiency challenges existing in predictive tools, we propose to selectively call a predictive tool based on concurrency coverage criteria. We have selected a sound predictive tool SeqCheck and adapted it to propose our fuzzing framework PredFuzz. We compared our tool with two tools, AFL++ integrated with Google ThreadSanitizer and AFL++ directly integrated with SeqCheck, on six previously studied multi-threaded programs. The experimental results showed that PredFuzz detected significantly more vulnerabilities than AFL++ integrated with ThreadSanitizer and about 70% vulnerabilities detected by AFL++ directly integrated with SeqCheck. Besides, it is extremely efficient without compromising the fuzzing speed of AFL++: it added a smaller slowdown to AFL++ than ThreadSanitizer did and achieved a speedup of more than 1,000x when compared to AFL++ directly integrated with SeqCheck. Yuqi Guo 0002, Zheheng Liang, Jinqiu Wang, Zijiang Yang 0006, Wuqiang Shen, Yan Cai 0001 |
COMPSAC | 2 |
| 2023 | Characterizing Flaky Tests in Node.js ApplicationsabstractRegression testing is an important means of assessing the quality of Node.js applications. However, non-deterministic executions inside Node.js framework could make test cases intermittently pass or fail on the same version of code, which are called flaky tests. Flaky tests can cause unreliable test results, and make developers waste a significant amount of time debugging the bugs that do not belong to the target application. In this paper, we conduct an empirical study on 87 flaky tests from 7 popular Node.js applications, and analyze the non-determinism that causes these flaky tests. Through this study, there is a wide range of non-determinism to cause flaky tests, including non-deterministic event triggering order, non-deterministic function calls, non-deterministic process/thread scheduling order, non-deterministic execution of asynchronous tasks and non-deterministic event triggering data. The result reveals that, existing approaches on event race detection are not sufficient for flaky test detection. In future, researchers can design flaky test detection approaches targeted at different categories of non-determinism. Xiaoning Chang, Zheheng Liang, Guoquan Wu, Yu Gao 0002, Wei Chen 0018, Jun Wei 0001, Zhenyue Long, Tao Huang 0001 |
ASE | 2 |
| 2022 | A multi-source spatio-temporal data cube for large-scale geospatial analysisabstractData management and analysis are challenging with big Earth observation (EO) data. Expanding upon the rising promises of data cubes for analysis-ready big EO data, we propose a new geospatial infrastructure layered over a data cube to facilitate big EO data management and analysis. Compared to previous work on data cubes, the proposed infrastructure, GeoCube, extends the capacity of data cubes to multi-source big vector and raster data. GeoCube is developed in terms of three major efforts: formalize cube dimensions for multi-source geospatial data, process geospatial data query along these dimensions, and organize cube data for high-performance geoprocessing. This strategy improves EO data cube management and keeps connections with the business intelligence cube, which provides supplementary information for EO data cube processing. The paper highlights the major efforts and key research contributions to online analytical processing for dimension formalization, distributed cube objects for tiles, and artificial intelligence enabled prediction of computational intensity for data cube processing. Case studies with data from Landsat, Gaofen, and OpenStreetMap demonstrate the capabilities and applicability of the proposed infrastructure. Peng Yue 0002, Zhipeng Cao 0003, Shuaifeng Zhao, Boyi Shangguan, Liangcun Jiang, Lei Hu 0001, Zhe Fang, Zheheng Liang |
Int. J. Geogr. Inf. Sci. | 9 |