VLDB 2026 Research / reviewers in the wild / expert
Wenzhao Zhang
dblp:148/0855
· DBLP profile ↗
18ranked-venue papers
9as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 5 first-author · 10 since 2021Systems, architecture and hardware · 6 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enabling Realtime Stream Processing with Partial Computation Offloading
Jiamei Lv, Wenzhao Zhang, Yixiao Teng, Yi Gao 0001, Wei Dong 0001 |
ICDCS | 3 |
| 2026 | Satellite-Terrestrial Collaborative Inference for IoRT: Optimizing Latency and Energy Efficiency
Shujun Han, Wenzhao Zhang, Xiaodong Xu 0001, Mengying Sun, Ping Zhang 0003 |
IEEE Internet Things J. | 3 |
| 2025 | Task-Oriented Cloud-Edge-Device Collaborative Semantic Communication: Trade-off Between Privacy-Preserving and QoAISabstractIn this paper, we formulate a Secure Hierarchical Semantic Communication (SH-SC) framework that leverages cloud-edge-device collaboration to enable efficient, robust, and privacy-preserving semantic communications. Firstly, we propose a quantization-aware efficient semantic communication (SemCom) model pre-training scheme running in the cloud. In particular, a semantic quantization method is applied to reduce the data required for transmission, and a quantization-aware multi-splitting points training method is proposed to mitigate the accuracy loss caused by quantization. Secondly, we propose a robust SemCom model deployment strategy in local device and honest but curious edge server for privacy-preserving, where a post-training quantization method on the device is proposed to reduce the computational overhead and enhance privacy preservation. Thirdly, we propose a SemCom model based adaptive device-edge collaborative inferencing mechanism for SemCom quality of AI services (QoAIS), where a Joint Quantization Device-Edge Collaboration Semantic Communication (JQDESC) scheme is formulated. Moreover, we provide a theoretical analysis of the privacy preservation of the proposed quantization scheme against model inversion attack through back-propagation and quantization error accumulation. Experimental results demonstrate that our proposed JQDESC scheme effectively protects privacy under various adversarial capabilities, and has better performance in memory usage and end-to-end latency while maintaining similar accuracy. Guanwu Jiang, Shujun Han, Xiaodong Xu 0001, Wenzhao Zhang, Ping Zhang 0003 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Effective Energy Efficiency Computation Offloading in NOMA-Based MEC Networks with Delay Violation Probability GuaranteeabstractWe investigate the joint communication and computation problem in non-orthogonal multiple access based mobile edge computation networks for massive intelligent machine type communication. We model the whole task offloading process as a double tandem queues model and formulate an optimization problem to maximize effective energy efficiency while guaranteeing the End-to-End delay violation probability. To solve this problem, we propose a Joint Transmission Power allocation and Computation Resources allocation (JTPCR) algorithm. Specifically, we first minimize the task processing delay to obtain the optimal computation resource constrained by total computation resources and maximum tolerable delay. In addition, we exploit the Dinkelbach method to solve the fractional programming problem when optimizing the transmission power. We introduce an auxiliary variable and obtain the lower bounding concave approximation of channel capacity through a path-following method. Finally, we propose the alternating direction method of the multipliers to obtain the optimal transmission power. Simulation results show that the proposed JTPCR algorithm outperforms the comparison schemes under both scenarios: finite transmission blocklength and infinite transmission blocklength. Wenzhao Zhang, Shujun Han, Mengying Sun, Xiaodong Xu 0001 |
WCNC | 1 |
| 2024 | S2E-DECI: Secrecy and Energy-Efficient Dual-Aware Device-Edge Co-Inference for AIoTabstractThis article proposes a secrecy and energy-efficient device-edge co-inference scheme for resource-constrained Artificial Intelligence of Things (AIoT) devices with physical layer security assistance. Our approach leverages split learning, where the AIoT device executes the initial part of the AI model, and the mobile edge computing server (MECs) computes the remainder, reducing energy consumption (EC) and inference delay. We measure secrecy capacity under the finite blocklength regime to address the vulnerability of intermediate feature data (IFD) to eavesdropping over wireless channels and its short block length characteristics. The objective is to minimize the average EC of the device-edge co-inference by jointly optimizing deep neural network (DNN) model partitioning and resource allocation. We formulate a distributed reinforcement learning-based joint DNN model partitioning and resource allocation (DRPA) algorithm, which uses knowledge-based reinforcement learning for optimal DNN partitioning and a convex optimization approach for resource allocation. Simulation results demonstrate that the DRPA algorithm achieves near-optimal performance, closely matching the results of exhaustive search methods. Shujun Han, Wenzhao Zhang, Xiaodong Xu 0001, Bizhu Wang, Mengying Sun, Xiaofeng Tao 0001, Ping Zhang 0003 |
IEEE Internet Things J. | 2 |
| 2024 | Understanding Differencing Algorithms for Mobile Application UpdatesabstractMobile application updates occur frequently, and they continue to add considerable traffic over the Internet. Differencing algorithms, which compute a small delta between the new version and the old version, are often employed to reduce the update overhead. Researchers have proposed many differencing algorithms over the years. Unfortunately, it is currently unknown how these algorithms quantitatively perform for different categories of applications. It is also challenging to know the impacts of different techniques and whether a technique in one algorithm can be integrated into another algorithm for further performance improvement. This paper conducts the first systematic study to understand the performance of four widely used differencing algorithms for mobile application updates, including xdelta3, bsdiff, archive-patcher, and HDiffPatch with respect to five key metrics, including compression ratio, differencing time/memory overhead, and reconstruction time/memory overhead. We perform measurements for 200 mobile applications, and analyze key techniques (such as decompressing-before-differencing, sliding window, and copy instructions merging) that influence the performance of these algorithms. We have provided four important findings which give insights to further optimize for performance improvement. Guided by these insights, we have also proposed a novel algorithm,sdiff, which achieves the smallest compression ratio to state-of-the-art algorithms by combining an appropriately chosen set of key techniques. Tong Sun 0006, Lewei Jin, Wenzhao Zhang, Yi Gao 0001, Wei Dong 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Reducing End-to-End Latency of Trigger-Action IoT Programs on Containerized Edge PlatformsabstractIoT rule engines are important middlewares that allow users to easily create custom trigger-action programs (TAPs) and interact with the physical world. Users expect their TAPs to give a timely response within a certain deadline. Existing works provide this support by boosting the process of trigger event identification. Many IoT rule engines now run in containerized environments, bringing about new challenges and opportunities. Prior solutions can no longer satisfy the need of mitigating the end-to-end latency of containerized TAPs. In this work, we propose EdgeRuler, which couples the IoT rule engine and the container runtime to assure the performance of latency-critical TAPs. To enable such capability, EdgeRuler precisely models the end-to-end latency by exploiting information from both the physical and the cyber world. EdgeRuler then enforces a deadline-aware life-cycle control and resource provision for meeting the TAP constraints in a lightweight and efficient way. We prototype and evaluate EdgeRuler on top of production-ready open-source components, which shows that EdgeRuler reduces the end-to-end latency by 28.6%-96.2% compared to existing scheduling algorithms and 68.4%-89.1% to that of the state-of-the-art IoT rule engines, incurring negligible runtime overhead. Wenzhao Zhang, Yixiao Teng, Yi Gao 0001, Wei Dong 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | LinkLab 2.0: A Multi-tenant Programmable IoT Testbed for Experimentation with Edge-Cloud Integration
Wei Dong 0001, Borui Li 0001, Kaijie Gong, Wenzhao Zhang, Yi Gao 0001 |
NSDI | 6 |
| 2023 | WiEdge: Edge Computing for Audio Sensing Applications With Accurate Wireless Link PredictionabstractAudio sensing applications on embedded and mobile devices have recently enjoyed increasing popularity. Their performance can be significantly improved by edge computing which offloads computation-intensive tasks to edge servers through wireless links. The quality of wireless links is essential to offloading performance. However, existing edge computing solutions can hardly predict the link quality accurately and efficiently in a dynamic wireless environment, resulting in less optimal offloading decisions and unsatisfied user-perceived Quality of Experience (QoE). In this article, we present WiEdge, a distributed edge computing framework for audio sensing applications with accurate wireless link prediction. By combining cross-layer information extracted from recently received WiFi beacons, TCP-level statistics, and the past throughput observations, WiEdge can predict the throughput of wireless links accurately and efficiently in the near future. Based on the prediction, WiEdge makes optimal offloading decisions for QoE maximization. We formulate the offloading decision problem as a stochastic optimal control problem and propose an efficient solution based on model predictive control from the control-theoretic perspective. We implement WiEdge and evaluate its performance extensively in three representative real-world scenarios. Results show that WiEdge achieves high prediction accuracy and improves average normalized QoE by 2%, 11%, and 40% in three different scenarios, compared with state-of-the-art approaches. Chenhong Cao, Wei Dong 0001, Wenzhao Zhang, Yi Gao 0001 |
IEEE Internet Things J. | 3 |
| 2023 | Providing Realtime Support for Containerized Edge ServicesabstractContainers have emerged as a popular technology for edge computing platforms. Although there are varieties of container orchestration frameworks, e.g., Kubernetes to provide high-reliable services for cloud infrastructure, providing real-time support at the containerized edge systems (CESs) remains a challenge. In this paper, we propose EdgeMan , a holistic edge service management framework for CESs, which consists of (1) a model-assisted event-driven lightweight online scheduling algorithm to provide request-level execution plans; (2) a bottleneck-metric-aware progressive resource allocation mechanism to improve resource efficiency. We then build a testbed that installed three containerized services with different latency sensitivities for concrete evaluation. Additionally, we adopt real-world data traces from Alibaba and Twitter for large-scale emulations. Extensive experiments demonstrate that the deadline miss ratio of time-sensitive services run with EdgeMan is reduced by 85.9% on average compared with that of existing methods in both industry and academia. Wenzhao Zhang, Yi Gao 0001, Wei Dong 0001 |
ACM Trans. Internet Techn. | 1 |
| 2023 | A Low-code Development Framework for Cloud-native Edge SystemsabstractCustomizing and deploying an edge system are time-consuming and complex tasks because of hardware heterogeneity, third-party software compatibility, diverse performance requirements, and so on. In this article, we present TinyEdge, a holistic framework for the low-code development of edge systems. The key idea of TinyEdge is to use a top-down approach for designing edge systems. Developers select and configure TinyEdge modules to specify their interaction logic without dealing with the specific hardware or software. Taking the configuration as input, TinyEdge automatically generates the deployment package and estimates the performance with sufficient profiling. TinyEdge provides a unified development toolkit to specify module dependencies, functionalities, interactions, and configurations. We implement TinyEdge and evaluate its performance using real-world edge systems. Results show that: (1) TinyEdge achieves rapid customization of edge systems, reducing 44.15% of development time and 67.79% of lines of code on average compared with the state-of-the-art edge computing platforms; (2) TinyEdge builds compact modules and optimizes the latent circular dependency detection and message routing efficiency; (3) TinyEdge performance estimation has low absolute errors in various settings. Wenzhao Zhang, Yuxuan Zhang 0003, Hongchang Fan, Yi Gao 0001, Wei Dong 0001 |
ACM Trans. Internet Techn. | 1 |
| 2022 | EdgeMan: Ensuring Real-Time Service for Containerized Edge SystemsabstractContainers have emerged as a popular technology for edge computing platforms. Although there are varieties of container orchestration frameworks, e.g., Kubernetes to provide high-reliable services for cloud infrastructure, ensuring realtime service at the containerized edge systems (CESs) remains a challenge. In this paper, we propose Edgeman,a holistic edge service management framework for CESs, which consists of (1) a model-assisted event-driven lightweight online scheduling algorithm to provide request-level execution plans; (2) a bottleneck-metric-aware progressive resource allocation mechanism to improve resource efficiency. We then build a testbed that installed three containerized services with different latency sensitivities for concrete evaluation. Besides, we adopt real-world data traces from Alibaba and Twitter for large-scale emulations. Extensive experiments demonstrate that the deadline miss ratio of Edgemanis reduced 85.9% on average compared with existing methods in both industry and academia. Wenzhao Zhang, Wei Dong 0001, Geng Ren, Yi Gao 0001 |
MSN | 1 |
| 2020 | TinyEdge: Enabling Rapid Edge System Customization for IoT ApplicationsabstractCustomizing and deploying an edge system is a time-consuming and complex task, considering the hardware heterogeneity, third-party software compatibility, diverse performance requirements, etc. In this paper, we present TinyEdge, a holistic system for the rapid customization of edge systems. The key idea of TinyEdge is to use a top-down approach for designing the software and estimating the performance of the customized edge systems under different hardware specifications. Developers select and conFigure modules to specify the critical logic of their interactions, without dealing with the specific hardware or software. Taking the configuration as input, TinyEdge automatically generates the deployment package and estimate the performance after sufficient profiling. TinyEdge provides a unified customization framework for modules to specify their dependencies, functionalities, interactions, and configurations. We implement TinyEdge and evaluate its performance using real-world edge systems. Results show that: 1) TinyEdge achieves rapid customization of edge systems, reducing 44.15% of customization time and 67.79% lines of code on average compared with the state-of-the-art edge platforms; 2) TinyEdge builds compact modules and optimizes the latent circular dependency detection and message queuing efficiency; 3) TinyEdge performance estimation has low average absolute error in various settings. Wenzhao Zhang, Yuxuan Zhang 0003, Hongchang Fan, Yi Gao 0001, Wei Dong 0001 |
SEC | 1 |
| 2016 | Exploring memory hierarchy and network topology for runtime AMR data sharing across scientific applicationsabstractRuntime data sharing across applications is of great importance for avoiding high I/O overhead for scientific data analytics. Sharing data on a staging space running on a set of dedicated compute nodes is faster than writing data to a slow disk-based parallel file system (PFS) and then reading it back for post-processing. Originally, the staging space has been purely based on main memory (DRAM), and thus was several orders of magnitude faster than the PFS approach. However, storing all the data produced by large-scale simulations on DRAM is impractical. Moving data from memory to SSD-based burst buffers is a potential approach to address this issue. However, SSDs are about one order of magnitude slower than DRAM. To optimize data access performance over the staging space, methods such as prefetching data from SSDs according to detected spatial access patterns and distributing data across the network topology have been explored. Although these methods work well for uniform mesh data, which they were designed for, they are not well suited for adaptive mesh refinement (AMR) data. Two major issues must be addressed before constructing such a memory hierarchy and topology-aware runtime AMR data sharing framework: (1) spatial access pattern detection and prefetching for AMR data; (2) AMR data distribution across the network topology at runtime. We propose a framework that addresses these challenges and demonstrate its effectiveness with extensive experiments on AMR data. Our results show the framework's spatial access pattern detection and prefetching methods demonstrate about 26% performance improvement for client analytical processes. Moreover, the framework's topology-aware data placement can improve overall data access performance by up to 18%. Wenzhao Zhang, Houjun Tang, Stephen Ranshous, Surendra Byna, Daniel F. Martin, Kesheng Wu, Bin Dong 0002, Scott Klasky, Nagiza F. Samatova |
IEEE BigData | 1 |
| 2016 | Usage Pattern-Driven Dynamic Data Layout ReorganizationabstractAs scientific simulations and experiments move toward extremely large scales and generate massive amounts of data, the data access performance of analytic applications becomes crucial. A mismatch often happens between write and read patterns of data accesses, typically resulting in poor read performance. Data layout reorganization has been used to improve the locality of data accesses. However, current data reorganizations are static and focus on generating a single (or set of) optimized layouts that rely on prior knowledge of exact future access patterns. We propose a framework that dynamically recognizes the data usage patterns, replicates the data of interest in multiple reorganized layouts that would benefit common read patterns, and makes runtime decisions on selecting a favorable layout for a given read pattern. This framework supports reading individual elements and chunks of a multi-dimensional array of variables. Our pattern-driven layout selection strategy achieves multi-fold speedups compared to reading from the original dataset. Houjun Tang, Surendra Byna, Steve Harenberg, Xiaocheng Zou, Wenzhao Zhang, Kesheng Wu, Bin Dong 0002, Oliver Rübel, Kristofer E. Bouchard, Scott Klasky, Nagiza F. Samatova |
CCGrid | 5 |
| 2016 | AMRZone: A Runtime AMR Data Sharing Framework for Scientific ApplicationsabstractFrameworks that facilitate runtime data sharingacross multiple applications are of great importance for scientificdata analytics. Although existing frameworks work well overuniform mesh data, they can not effectively handle adaptive meshrefinement (AMR) data. Among the challenges to construct anAMR-capable framework include: (1) designing an architecturethat facilitates online AMR data management, (2) achievinga load-balanced AMR data distribution for the data stagingspace at runtime, and (3) building an effective online indexto support the unique spatial data retrieval requirements forAMR data. Towards addressing these challenges to supportruntime AMR data sharing across scientific applications, wepresent the AMRZone framework. Experiments over real-worldAMR datasets demonstrate AMRZone's effectiveness at achievinga balanced workload distribution, reading/writing large-scaledatasets with thousands of parallel processes, and satisfyingqueries with spatial constraints. Moreover, AMRZone's performance and scalability are even comparable with existing state-of-the-art work when tested over uniform mesh data with up to16384 cores, in the best case, our framework achieves a 46% performance improvement. Wenzhao Zhang, Houjun Tang, Steve Harenberg, Surendra Byna, Xiaocheng Zou, Dharshi Devendran, Daniel F. Martin, Kesheng Wu, Bin Dong 0002, Scott Klasky, Nagiza F. Samatova |
CCGrid | 1 |
| 2016 | In Situ Storage Layout Optimization for AMR Spatio-temporal Read AccessesabstractAnalyses of large simulation data often concentrate on regions in space and in time that contain important information. As simulations adopt Adaptive Mesh Refinement (AMR), the data records from a region of interest could be widely scattered on storage devices and accessing interesting regions results in significantly reduced I/O performance. In this work, we study the organization of block-structured AMR data on storage to improve performance of spatio-temporal data accesses. AMR has a complex hierarchical multi-resolution data structure that does not fit easily with the existing approaches that focus on uniform mesh data. To enable efficient AMR read accesses, we develop an in situ data layout optimization framework. Our framework automatically selects from a set of candidate layouts based on a performance model, and reorganizes the data before writing to storage. We evaluate this framework with three AMR datasets and access patterns derived from scientific applications. Our performance model is able to identify the best layout scheme and yields up to a 3X read performance improvement compared to the original layout. Though it is not possible to turn all read accesses into contiguous reads, we are able to achieve 90% of contiguous read throughput with the optimized layouts on average. Houjun Tang, Surendra Byna, Steve Harenberg, Wenzhao Zhang, Xiaocheng Zou, Daniel F. Martin, Bin Dong 0002, Dharshi Devendran, Kesheng Wu, David Trebotich, Scott Klasky, Nagiza F. Samatova |
ICPP | 4 |
| 2015 | Exploring Memory Hierarchy to Improve Scientific Data Read PerformanceabstractImproving read performance is one of the major challenges with speeding up scientific data analytic applications. Utilizing the memory hierarchy is one major line of researches to address the read performance bottleneck. Related methods usually combine solide-state-drives(SSDs) with dynamic random-access memory(DRAM) and/or parallel file system(PFS) to mitigate the speed and space gap between DRAM and PFS. However, these methods are unable to handle key performance issues plaguing SSDs, namely read contention that may cause up to 50% performance reduction. In this paper, we propose a framework that exploits the memory hierarchy resource to address the read contention issues involved with SSDs. The framework employs a general purpose online read algorithm that able to detect and utilize memory hierarchy resource to relieve the problem. To maintain a near optimal operating environment for SSDs, the framework is able to orchastrate data chunks across different memory layers to facilitate the read algorithm. Compared to existing tools, our framework achieves up to 50% read performance improvement when tested on datasets from real-world scientific simulations. Wenzhao Zhang, Houjun Tang, Xiaocheng Zou, Steve Harenberg, Qing Liu 0002, Scott Klasky, Nagiza F. Samatova |
CLUSTER | 1 |