VLDB 2026 Research / reviewers in the wild / expert
Wei Zhang 0173
dblp:10/4661-173
· DBLP profile ↗
27ranked-venue papers
8as first author
22since 2021 · last 2026
0000-0003-2615-2603ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 7 first-author · 14 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Security and privacy · 2 · 2 since 2021Theory of computation · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Path-Sensitive Abstract Interpretation for WCET EstimationabstractWorst-Case Execution Time (WCET) analysis provides an upper bound on a program’s execution time and is fundamental to the design and verification of real-time systems. Accurate modeling of cache behavior is critical for WCET estimation, as cache-miss latency is typically orders of magnitude larger than cache-hit latency. Since cache behavior is path dependent, existing methods commonly use abstract interpretation to estimate cache behaviors without enumerating all paths. However, conventional abstract interpretation is context-agnostic—adopting the most conservative case across paths—and thus may produce an overestimated WCET bound. To bridge the gap between scalability and accuracy, we propose a path-sensitive abstract-interpretation-based cache analysis that maintains a set of cache states drawn from critical execution paths to derive context-aware cache behavior. This path-sensitive cache analysis integrates seamlessly into standard WCET frameworks, resulting in tight yet provably sound WCET bounds. Experiments show that our approach improves WCET accuracy by an average of 24.83% without sacrificing scalability. Shangshang Xiao, Mengxia Sun, Wei Zhang 0173, Naijun Zhan, Lei Ju 0001 |
Proc. ACM Program. Lang. | 3 |
| 2025 | Co-Prime: A Co-design Framework for Privacy Preserving Machine Learning on FPGAabstractIn enormous privacy-sensitive machine learning application domains with collaborative data acquisition from multiple participants, secure multi-party computation (MPC) becomes a promising solution for privacy-preserving machine learning (PPML). Secret sharing protocols is a prevalent MPC strategy, where frequent data distribution and recombination are applied to uphold the confidentiality of participants' data. A key challenge for practical deployment of secret sharing protocols in PPML is the massive and unbalanced computation and communication workloads occurred in various linear and non-linear stages of machine learning. The imbalance could be further amplified when powerful hardware accelerators are designed to reduce the computation latency. In this work, we propose Co-Prime, an FPGA-based 3PC framework for efficient PPML without assistance from a secure third party. Co-Prime integrates protocol and hardware co-optimizations to mitigate the communication bottlenecks in secret sharing schemes. Particularly, Co-Prime proposes a novel protocol conversion technique that seamlessly converts data formats to adaptively adopt preferred protocols in various stages of PPML. Accelerator-friendly MPC primitives and system-level design space exploration schemes are designed to achieve latency hiding through overlapping computation and network communication. Finally, it enables direct interaction with data streams via network communication modules on FPGAs to further reduce the network communication overhead. Experimental results demonstrate significant performance improvements over existing privacy-preserving machine learning frameworks, with 2-18x speedup in inference latency across various LAN/WAN environments and neural network models. Jiming Xu, Lei Ju 0001, Wei Zhang 0173 |
CCS | 6 |
| 2025 | When to Skip: Sparsity-Aware Acceleration for Intermittent Neural Network InferenceabstractDeep neural networks (DNNs) are increasingly being utilized on battery-less IoT devices to facilitate intelligent edge computing. These devices, which lack batteries, often encounter frequent power interruptions and operate under the intermittent computing paradigm. Under this paradigm, system states are periodically backed up before power failures, allowing for the resumption of execution from the latest backup to maintain progress. However, the shift on the applications from traditional control-based embedded tasks to memory and computation-intensive DNN tasks presents new challenges for intermittent computing. Particularly, the significant size of intermediate results during DNN inference results in significant backup and resumption overhead. Our research indicates that exploiting the sparsity of DNN weights and inputs in intermittent systems can effectively mitigate both backup and resumption overhead without compromising inference accuracy. In this study, we introduce a novel sparse-aware acceleration framework for intermittent DNN inference. This framework selectively skips redundant computations during DNN inference, reducing both intermittent system backups and computation overhead. Experimental results on an MSP430 device show that our approach achieves up to$2.98 \times$speedup and an average of$2.56 \times$latency reduction compared to state-of-the-art baselines, while maintaining high inference accuracy. Wei Zhang 0173, Lei Ju 0001 |
HPCC | 3 |
| 2025 | Demo: Real-Time Inference on GPU-Based Heterogeneous SoCs with GPU Cache Locking
Kehao Ma, Wei Zhang 0173, Mengying Zhao, Lei Ju 0001 |
RTCSA | 2 |
| 2025 | Tight Cache Contention Analysis for WCET Estimation on Multicore SystemsabstractWCET (Worst-Case Execution Time) estimation on multicore architecture is particularly challenging mainly due to the complex accesses over cache shared by multiple cores. Existing analysis identifies possible contentions between parallel tasks by leveraging the partial order of the tasks or their program regions. Unfortunately, they overestimate the number of cache misses caused by a remote block access without considering the actual cache state and the number of accesses. This paper reports a new analysis for inter-core cache contention. Based on the order of program regions in a task, we first identify memory references that could be affected if a remote access occurs in a region. Afterwards, a fine-grained contention analysis is constructed that computes the number of cache misses based on the access quantity of local and remote blocks. We demonstrate that the overall inter-core cache interference of a task can be obtained via dynamic programming. Experiments show that compared to existing methods, the proposed analysis reduces inter-core cache interference and WCET estimations by$\mathbf{5 2. 3 1 \%}$and$\mathbf{8. 9 4 \%}$on average, without significantly increasing computation overhead. Shuai Zhao 0004, Jieyu Jiang, Shenlin Cai, Yaowei Liang, Chen Jie, Yinjie Fang, Wei Zhang 0173, Guoquan Zhang, Yaoyao Gu, Ouyang Ouyang, Wanli Chang 0001 |
RTSS | 7 |
| 2025 | DAHE: Parameter-Adaptive and Memory Efficient FPGA Acceleration of Homomorphic EncryptionabstractWhile homomorphic encryption (HE) has been well-recognized as a promising data privacy protection technique, there are many challenges to the real-world deployment of HE applications. In this work, we propose a design flow for parameter-adaptive and memory-efficient FPGA acceleration of homomorphic encryption. In the framework, we explore the correlations between HE parameter selection to meet various design objectives and the huge design space due to underlying FPGA hardware resource allocation. Particularly, we demonstrate that adaptive management of the FPGA memory hierarchy is crucial to supporting diverse cryptosystem parameter selection for application-level security, accuracy, and performance requirements. We propose a resource-efficient and flexible micro-architectural design for HE operations, where data access patterns in various pipeline execution stages are optimized for high memory bandwidth utilization. Furthermore, a memory-aware performance model is built for automatic design space exploration for cryptosystem parameter selection and hardware resource provisioning. Experimental results show 1.50X and 1.16X speedup for the NTT and Rotation operations w.r.t. the state-of-the-art FPGA implementation. Meanwhile, the proposed framework generates flexible and high-performance accelerator code for real HE application kernels with different cryptosystem parameters on a wide range of FPGA devices. Yilan Zhu, Honghui You, Wei Zhang 0173, Jiming Xu, Qian Lou, Shoumeng Yan, Lei Ju 0001 |
IEEE Trans. Computers | 3 |
| 2025 | WCET Estimation for CNN Inference on FPGA SoC With Multi-DPU EnginesabstractThe Deep Learning Processor Unit (DPU) released in the official Xilinx Vitis AI toolchain stands as a commercial off-the-shelf solution tailored for accelerating convolutional neural network (CNN) inference on Xilinx FPGA devices. While most FPGA accelerator focus on high performance and energy-efficiency, analyzing the worst-case execution time (WCET) bound is essential for using CNN accelerations in real-time embedded systems design. In this work, we show that in a multi-DPU environment, the observed worst-case inference time for a CNN inference task could become 3X larger w.r.t. the best case inference time, which prompts the prominent importance of a static timing analysis for FPGA-based CNN inference. We propose, to the best of the authors’ knowledge, the first static timing analysis framework for CNN inference in a multi-DPU environment. The proposed framework introduces a generalized timing behavior model for shared bus arbitration and memory access contention between parallel running DPU engines. Additionally, it incorporates a fine-grained memory access contention analysis that takes into account the characteristics of deep learning applications. For a single-DPU environment, the analysis result is 27% tighter in average compared with the state-of-the-art results. Furthermore, our proposed method produces relatively tight estimated results in the multi-DPU environment. Wei Zhang 0173, Yunlong Yu 0004, Nan Guan, Naijun Zhan, Lei Ju 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2024 | Cache-aware Task Decomposition for Efficient Intermittent Computing SystemsabstractEnergy harvesting offers a scalable and cost-effective power solution for IoT devices, but it introduces the challenge of frequent and unpredictable power failures due to the unstable environment. To address this, intermittent computing has been proposed, which periodically backs up the system state to non-volatile memory (NVM), enabling robust and sustainable computing even in the face of unreliable power supplies. In modern processors, write back cache is extensively utilized to enhance system performance. However, it poses a challenge during backup operations as it buffers updates to memory, potentially leading to inconsistent system states. One solution is to adopt a write-through cache, which avoids the inconsistency issue but incurs increased memory access latency for each write reference. Some existing work enforces a cache flushing before backups to maintain a consistent system state, resulting in significant backup overhead. In this paper, we point out that although cache delays updates to the main memory, it may preserve a recoverable system state in the main memory. Leveraging this characteristic, we propose a cache-aware task decomposition method that divides an application into multiple tasks, ensuring that no dirty cache lines are evicted during their execution. Furthermore, the cache-aware task decomposition maintains an unchanged memory state during the execution of each task, enabling us to parallelize the backup process with task execution and effectively hide the backup latency. Experimental results with different power traces demonstrate the effectiveness of the proposed system. Wei Zhang 0173, Mengying Zhao, Zimeng Zhou, Lei Ju 0001 |
DAC | 2 |
| 2024 | Freshness-aware Data Backup for Batteryless Sensing SystemsabstractBatteryless sensing systems rely on energy harvested from the environment to execute. However, as the harvested energy is generally weak and unstable, the system may experience frequent power failures during processing and sensing. To make forward progress across power outages, the system backs up the system state from static random access memory (SRAM) to non-volatile memory (NVM) before power failures and then restores it upon reboot. Moreover, to avoid losing the collected data, existing approaches save all the collected data from SRAM to NVM before system-off. The data saving and the frequent system reboots consume a lot of energy and time and thus cause a long blocking time. However, the data stored in SRAM can be retained for a short period even after the system is turned off, as the data retention voltage of SRAM is lower than the minimum operating voltage of the microcontroller unit (MCU). In this paper, we leverage the SRAM data retention capability to retain data with a short lifetime on SRAM, while only save data with a long lifetime to NVM. Consequently, the backup overhead is significantly reduced. However, a design challenge is to decide the turn-off voltage to minimize the blocking time. Specifically, turning off at a higher voltage leaves more energy for a longer retention time and results in lower data saving overhead. But this may also cause more on-offs, leading to more system states saving and system states restoring overhead. To address this challenge, the paper proposes a method to adaptively compute the optimal turn-off voltage. Experimental results show that the proposed method can significantly reduce the blocking time caused by data saving and system reboots. The system can collect more data and exhibits improved responsiveness in sensing the environment. Yunlong Yu 0004, Wei Zhang 0173, Songran Liu, Mingsong Lv, Nan Guan, Lei Ju 0001 |
HPCC | 3 |
| 2024 | SoTimer: A Software-based Timekeeper for Energy Harvesting SystemsabstractThe proliferation of IoT devices has made energy harvesting systems an appealing solution for power numerous IoT devices without the limitations imposed by battery life constraints. Despite its potential, energy harvesting systems often encounter frequent power failures leading to interruptions of tasks due to the generally weak energy output they rely on. In modern computer systems, maintaining a consistent power flow is crucial for tasks such as synchronization and system stability. However, in energy harvesting systems, maintaining this consistency is challenging as the system is hard to track time during power failures. Current approaches rely on additional hardware and exploit physical phenomena to estimate power-off duration, but these methods are susceptible to environmental variations like temperature, humidity, and radiation, and are not easily applicable to commercial-off-the-shelf (COTS) micro-controllers. In this study, we introduce SoTimer, a software-based timekeeping solution designed to mitigate the impact of environmental changes without the need for extra hardware. The core concept of SoTimer lies in the stability of charging power between adjacent power-on and power-off periods, as evidenced by extensive measurements under different energy resources. Building on this insight, we propose a machine learning algorithm to predict power-off times based on power conditions observed during power-on phases. Furthermore, we introduce optimization techniques tailored for MCUs with limited computational capabilities to ensure efficient inference. Experimental results demonstrate that SoTimer achieves a high inference accuracy and a negligible runtime overhead. Yunlong Yu 0004, Wei Zhang 0173, Lei Ju 0001 |
HPCC | 3 |
| 2024 | Data-Dependent WAR Analysis for Efficient Task-Based Intermittent Computing
Juxin Niu, Yunlong Yu 0004, Wei Zhang 0173, Nan Guan |
SETTA | 3 |
| 2024 | Cache Behavior Analysis with SP-Relative Addressing for WCET Estimation
Shangshang Xiao, Mengxia Sun, Wei Zhang 0173, Naijun Zhan, Lei Ju 0001 |
SETTA | 3 |
| 2024 | A GPU-Based Privacy-Preserving Machine Learning Acceleration SchemeabstractAs the application of artificial intelligence expands, privacy-preserving machine learning has become a critical research focus. Secret sharing, as a commonly used privacy-preserving technique, has broad application prospects due to its lightweight ciphertext computation characteristics. While secret sharing offers lightweight computation, most schemes are implemented on CPU platforms, leaving room for exploration on GPUs. This paper proposes a GPU-based acceleration scheme for privacy-preserving machine learning, utilizing the ABY3 secret sharing protocol and a 32-bit integer ring, while supporting regularization techniques. Experimental results on standard neural networks, such as VGG-16 and AlexNet, demonstrate a significant performance improvement. The proposed approach achieves a 95× improvement over the CPU-based Falcon scheme and a 3.5× improvement over the GPU-based CryptGPU scheme during privacy training, while reducing communication volume by over 50% in both inference and training phases. Zengrui Huang, Zhiyong Zhang 0006, Wei Zhang 0173, Lei Ju 0001 |
TrustCom | 4 |
| 2023 | Accelerating DNN Inference with Heterogeneous Multi-DPU EnginesabstractThe Deep Learning Processor (DPU) programmable engine released by the official Xilinx Vitis AI toolchain has become one of the commercial off-the-shelf (COTS) solutions for Convolutional Neural Networks (CNNs) inference on Xilinx FPGAs. While modern FPGA devices generally have enough hardware resources to accommodate multi-DPUs simultaneously, the Xilinx toolchain currently only supports the deployment of multiple homogeneous DPUs engines that running independent inference tasks (task-level parallelism). In this work, we demonstrate that deployment of multiple heterogeneous DPU engines makes better resource efficiency for a given FPGA device. Moreover, we show that pipelined execution of a CNN inference task over heterogeneous multi-DPU engines may further improve overall inference throughput with carefully designed CNN layers-to-DPU mapping and scheduling. Finally, for a given CNN model and an FPGA device, we propose a comprehensive framework that automatically determines the optimal heterogeneous DPU deployment, and adaptively chooses the execution scheme between task-level and pipelined parallelism. Compared with the state-of-the-art solution with homogeneous multi-DPU engines and network-level parallelism, the proposed framework shows an average improvement of 13% (up-to 19%) and 6.6% (up-to 10%) on the Xilinx Zynq UltraScale+ MPSoC ZCU104 and ZCU102 platforms, respectively. Zelin Du, Wei Zhang 0173, Zimeng Zhou, Zili Shao, Lei Ju 0001 |
DAC | 2 |
| 2023 | Work or Sleep: Freshness-Aware Energy Scheduling for Wireless Powered Communication Networks with Interference ConsiderationabstractThis paper explores how to schedule energy to optimize the information freshness in wireless powered communication networks (WPCNs) when considering channel interference among adjacent sensor nodes. We introduce Age of Information (AoI) to quantitatively evaluate the information freshness and formulate the AoI optimization problem. Unlike prior works focusing on system optimization for WPCNs while ignoring channel interference in energy transfer, this work reveals situations where channel interference among adjacent sensor nodes cannot be neglected and explores optimizing information freshness with interference consideration. To take the phenomena into account, we propose an energy scheduling solution to detect the channel interference and then judiciously determine the energy and time allocation for individual sensor nodes to improve the AoI performance as well as the system throughput. We implement a multi-node WPCN testbed to validate the functional correctness of the proposed solution, and extensive experiments have demonstrated the effectiveness of the proposed solution. The experimental results show that the proposed solution can reduce the average AoI by 54.7% and the average throughput by 49.8% on average compared to the state-of-the-art solutions. Lei Ju 0001, Chun Jason Xue, Mingliang Zhou 0001, Wei Zhang 0173, Zimeng Zhou |
DAC | 5 |
| 2023 | Light Flash Write for Efficient Firmware Update on Energy-harvesting IoT DevicesabstractFirmware update is an essential service on Internet-of-Things (IoT) devices to fix vulnerabilities and add new functionalities. Firmware update is energy-consuming since it involves intensive flash erase/write operations. Nowadays, IoT devices are increasingly powered by energy harvesting. As the energy output of the harvesters on IoT devices is typically tiny and unstable, a firmware update will likely experience power failures during its progress and fail to complete. This paper presents an approach to increase the success rate of firmware update on energy-harvesting IoT devices. The main idea is to first conduct a lightweight flash write with reduced erase/write time (and thus less energy consumed) to quickly save the new firmware image to flash memory before a power failure occurs. To ensure a long data retention time, a reinforcement step follows to re-write the new firmware image on the flash with default erase/write configuration when the system is not busy and has free energy. Experiments conducted with different energy scenarios show that our approach can significantly increase the success rate and the efficiency of firmware update on energy-harvesting IoT devices. Songran Liu, Mingsong Lv, Wei Zhang 0173, Xu Jiang 0004, Chuancai Gu, Tao Yang 0024, Wang Yi 0001, Nan Guan |
DATE | 3 |
| 2023 | Adaptive Task-Based Intermittent Computing System With Parallel State BackupabstractEnergy harvesting promises to power billions of Internet of Things devices without being restricted by battery life. Since the energy harvester generally outputs weak and unstable energy, the system may suffer frequent and unpredictable power failures, thus falling into cyclically reboots without forward progress. The task-based intermittent computing system which periodically backs up system states into nonvolatile memory (NVM) is proposed to solve the nonprogress problem, with the nontrivial cost of frequent backups. How to reduce the backup overhead becomes a major research problem for intermittent computing. This article, for the first time, proposes to parallelize state backup and program execution with asynchronous direct memory access (DMA) to hide the backup latency into the program’s execution. But, straightforwardly executing the state backup and the program in parallel may cause an inconsistent system state. In specific, the system state may be modified by the program during backup, and therefore may be backed up incorrectly and further cause the system to deliver an incorrect computation result. We make a deep analysis on the system behavior and observe that, although the system state may be backed up incorrectly, the incorrect backup will be covered by the subsequent correct backups soon as the backup operations are performed frequently. In addition, only a small part of variables among all the program states may cause incorrect computation result. So, in this article, we aggressively allow incorrect backups to occur and propose a backup error detection method and a fault-tolerant backup management to guarantee the correctness of the system’s execution. To augment the parallel backup method, an adaptive execution method is further proposed to reduce the number of backups and balance the ratio between task execution time and backup latency. We design a run-time system to implement the proposed approach, and experimental results conducted on an STM32F7-based platform show that the proposed method can achieve a$2.6\times $average speedup. Wei Zhang 0173, Qianling Zhang, Mingsong Lv, Songran Liu, Zimeng Zhou, Qiulin Chen, Nan Guan, Lei Ju 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Optimizing Worst Case Data Freshness in RF-Powered Networked Embedded SystemsabstractMaintaining real-time data freshness plays a critical role in ensuring system correctness and optimizing the system performance in networked embedded systems (NESs). To quantitatively measure the freshness of the collected real-time data, the concept of Age of Information (AoI) has been extensively studied in recent years. This article explores how to minimize the worst case AoI of real-time data in radio-frequency (RF)-powered NESs. In such systems, one hybrid access point (HAP) transfers wireless power to a set of distributed sensor nodes, and in the meantime, receives the information from these sensor nodes. We utilize the metric of AoI to measure the data freshness and present a comprehensive analysis of the worst case AoI of the real-time data in the target system. Based on the analysis, an optimal energy schedule solution is designed to judiciously determine individual sensor nodes’ energy and time allocation to minimize the worst case AoI. Considering the varying importance of different information and sensor nodes in the target system, we further propose the optimal time and energy allocation scheme for minimizing the weighted worst case AoI. A multinode RF-powered NES testbed is implemented to validate the functional correctness of our solutions. The results show that our solutions significantly outperform the state-of-the-art solutions, reducing the worst case AoI and weighted worst case AoI by 69.3% and 75.1% on average, respectively. Zimeng Zhou, Chenchen Fu, Chun Jason Xue, Song Han 0002, Wei Zhang 0173, Lei Ju 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | Precise and scalable shared cache contention analysis for WCET estimationabstractWorst-Case Execution Time (WCET) analysis for real-time tasks must precisely predict cache hit/miss of memory accesses. While bringing great performance benefits, multi-core processors significantly complicate the cache analysis problem due to the shared cache contentions among different cores. Existing methods pessimistically consider that memory references of parallel executing tasks will contend with each other as long as they are mapped to the same cache line. However, in reality, numerous shared cache contentions are mutually exclusive, due to the partial orders among the programs executed in parallel. The presence of shared cache contentions greatly exacerbates the computational complexity of the WCET computation, as finding the longest path needs exploring an exponentially large partial ordering space. In this paper, we propose a quantitative method with O(n2) time complexity to precisely estimate the worst-case extra execution time (WCEET) caused by shared cache contentions. The proposed method can be easily integrated into the abstract-interpretation based WCET estimation framework. Experiments with MRTC benchmarks show that our method can averagely tighten the WCET estimation by 13% without sacrificing the analysis efficiency. Wei Zhang 0173, Mingsong Lv, Wanli Chang 0001, Lei Ju 0001 |
DAC | 1 |
| 2022 | TICK: Tiny Client for BlockchainsabstractIn order to be deployed on storage-limited devices, blockchains generally provide lightweight clients which only store all the block headers rather than all blocks. However, a lightweight client is hard to verify a newly issued transaction, thus making the zero-confirmation transactions between lightweight clients impossible. In particular, transaction verification needs to verify that each referred output of the transaction is not previously spent. The conventional lightweight client design is unscalable as it can only support such an operation in the complexity of$O$($N_{T}$), where$N_{T}$is the total number of transactions in the system. The latest proposals suggest summarizing all the unspent outputs in an ordered Merkle tree. Therefore, a light client can request proof of presence and/or absence of an element in it to prove whether a referred output is previously spent or not, in the complexity of$O$(log($N_{U}$)), where$N_{U}$is the total number of unspent output in the system. However, updating such ordered Merkle tree is slow, thus making the system impractical—by our evaluation, when a new block is generated in Bitcoin, it takes more than one minute to update the ordered Merkle tree. We propose a practical client, TICK, to solve this problem. TICK uses the AVL hash tree to store all the unspent outputs. The AVL hash tree can be updated in the time of$O$($M$*log($N_{U}$)), where$M$is the number of elements that need to be inserted or removed from the AVL hash tree. By evaluation, when a new block is generated, the AVL hash tree can be updated within 1 s. Similarly, the proof can also be generated in the time of$O$(log($N_{U}$)). Therefore,${\textsf {TICK}}$is practical and scalable. Benefited by the AVL hash tree, a storage-limited device can efficiently and cryptographically verify transactions. In addition, rather than requiring new miners to download the entire blockchain before mining, TICK allows new miners to download only a small portion of data to start mining. We implement TICK for Bitcoin and provide an experimental evaluation on its performance by using the current Bitcoin blockchain data. Our result shows that the proof for verifying whether an output of a transaction is spent or not is only several kB. The verification is very fast—generating a proof generally takes less than 1 ms and verifying a proof even takes much less time. In addition, to start mining, new miners only need to download several GB data, rather than downloading over 230-GB data. Wei Zhang 0173, Jiangshan Yu, Qingqiang He, Nan Guan |
IEEE Internet Things J. | 1 |
| 2021 | Intermittent Computing with Efficient State Backup by Asynchronous DMAabstractEnergy harvesting promises to power billions of Internet-of-Things devices without being restricted by battery life. The energy output of harvesters is typically weak and highly unstable, so computing systems must frequently back up program states into non-volatile memory to ensure a program will progress in the presence of frequent power failures. However, state backup is a time-consuming process. In existing solutions for this problem, state backup is conducted sequentially with program execution, which considerably impact system performance. This paper proposes techniques to parallelize state backup and program execution with asynchronous DMA. The challenge is that program states can be incorrectly backed up, which may further cause the program to deliver incorrect computation. Our main idea is to allow errors to occur in parallel state backup and program execution, and detect the errors at the end of the state backup. Moreover, we propose a technique that allows the system to tolerate backup errors during execution without harming logical correctness. We designed a run-time system to implement the proposed approach. Experimental results on an STM32F7-based platform show that execution performance can be considerably improved by parallelizing state backup and program execution. Wei Zhang 0173, Songran Liu, Mingsong Lv, Qiulin Chen, Nan Guan |
DATE | 1 |
| 2021 | Surviving Transient Power Failures with SRAM Data RetentionabstractMany computing systems, such as those powered by energy harvesting or deployed in harsh working environment, may experience unpredictable and frequent transient power failures in their life time. The systems may fail to deliver correct computation results or never progress, as computation is frequently interrupted by the power failures. A possible solution could be frequently saving program states to non-volatile memory (NVM), such as using checkpoints, so that the system can incrementally progress. However, this approach is too costly, since frequent NVM writes is time and energy consuming, and may wear out the NVM device. In this work, we propose an approach to enable a system to use volatile SRAM to correctly progress in the presence of transient power failures, since SRAM is capable of retaining its data for seconds or minutes with the charge remained in the battery/capacitor after the CPU core stops at its brown-out voltage. The main problem is to validate whether the data in SRAM are actually retained during power failures. In our approach, we validate only a subset of the program states with Cyclic Redundancy Check for efficiency. The validation technique requires maintaining a backup version of the program states, which additionally provides the system with the ability to progress incrementally. We implement a run-time system with the proposed approach. Experimental results on an MSP430 platform show that the system can correctly progress on SRAM in the presence of transient power failures with low overhead. Songran Liu, Wei Zhang 0173, Mingsong Lv, Qiulin Chen, Nan Guan |
DATE | 2 |
| 2020 | LATICS: A Low-Overhead Adaptive Task-Based Intermittent Computing SystemabstractEnergy harvesting promises to power billions of Internet-of-Things devices without being restricted by battery life. The energy output of harvesters is typically tiny and highly unstable, so the computing system must store program states into nonvolatile memory frequently to preserve the execution progress in the presence of frequent power failures. Task-based intermittent computing is a promising paradigm to provide such capability, where each task executes atomically and only states across task boundaries need to be saved. This article presents LATICS, a low-overhead adaptive task-based intermittent computing system, which dynamically decides the granularity of atomic execution to avoid unnecessarily frequent state saving when energy supply is sufficient. The novel feature of LATICS is to drastically reduce the amount of states to be saved at task boundaries compared with existing solutions. Notably, we disclose that skipping state saving at some task boundary may cause the system to store more states at other places, and thus leads to higher overall overhead. Therefore, LATICS enforces mandatory state saving at certain task boundaries regardless of the current energy condition to reduce state saving overhead. We implement LATICS on a real energy-harvesting platform based on MSP430 and experimentally compare against the state-of-the-art under different settings. The experimental results show that LATICS significantly reduces state saving overhead and improves execution efficiency compared to existing solutions. Songran Liu, Wei Zhang 0173, Mingsong Lv, Qiulin Chen, Nan Guan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Scope-Aware Useful Cache Block Calculation for Cache-Related Pre-Emption Delay Analysis With Set-Associative Data CachesabstractTiming analysis of real-time systems must consider cache-related pre-emption delay (CRPD) costs when pre-emptive scheduling is used. While most previous work on CRPD analysis only considers instruction caches, the CRPD incurred on data caches is actually more significant. The state-of-the-art CRPD analysis methods are based on useful cache block (UCB) calculation. Unfortunately, as shown in this article, directly extending the existing UCB calculation techniques from instruction caches to data caches will lead to both unsoundness and significant imprecision. To solve these problems, we develop a new UCB calculation technique for data caches, which redefines the analysis unit (to address the unsoundness in the existing method) and precisely captures the dynamic cache access behavior by taking the temporal scopes of memory blocks into consideration. The experimental results show that our new technique yields substantially tighter CRPD estimations comparing with the state-of-the-art. Wei Zhang 0173, Nan Guan, Lei Ju 0001, Yue Tang 0001, Weichen Liu 0001, Zhiping Jia |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | Scope-aware data cache analysis for OpenMP programs on multi-core processors
He Du, Wei Zhang 0173, Nan Guan, Wang Yi 0001 |
J. Syst. Archit. | 2 |
| 2018 | Analyzing Data Cache Related Preemption Delay With Multiple PreemptionsabstractTiming analysis of real-time tasks under preemptive scheduling must take cache-related preemption delay (CRPD) into account. Typically, a task may be preempted more than once during the execution in each period. To bound the total CRPD of${k}$preemptions, existing CRPD analysis techniques estimate the CRPD at each program point, and use the sum of the${k}$-largest CRPD among all program points as the total CRPD upper bound. In this paper, we disclose that the above-mentioned approach, although works well for instruction caches, leads to significant overestimation when dealing with data caches. This is because on data caches, the CRPD of preemptions at different program points may have correlations, and the total CRPD of multiple preemptions is in general smaller than the simple sum of the worst-case CRPD of each preemption. To address this problem, we propose a new technique to efficiently explore the correlation among the CRPD of different preemptions, and thus more precisely calculate the total CRPD. Experiments with benchmark programs show that the proposed technique leads to substantially tighter total CRPD estimation with multiple preemptions comparing with the state-of-the-art. Wei Zhang 0173, Nan Guan, Lei Ju 0001, Weichen Liu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2017 | Scope-Aware Useful Cache Block Analysis for Data Cache Related Preemption DelayabstractStatic timing analysis is crucial for design of realtime systems. While the worst-case execution time of a task is typically computed or measured in a single task environment, the presence of caches imposes additional cache related preemption delay (CRPD) cost to the lower priority tasks in a preemptive multi-tasking system. In this work, we show that existing instruction CRPD analysis techniques cannot be straightforwardly extended for safe and precise data CRPD analysis. In order to capture the dynamic behavior of the data memory references, we introduce the notion of temporal scopes into the abstract cache state (ACS) to capture the data memory blocks that must or may reside in the cache during certain time intervals of program execution. Based on the improved ACS representation, we present a temporal scope aware useful cache block (UCB) calculation for safe and tight estimation of the data CRPD cost. Experimental results show that the proposed technique leads to substantially tighter CRPD estimation, and is applicable to programs with complex data reference patterns. Wei Zhang 0173, Fan Gong, Lei Ju 0001, Nan Guan, Zhiping Jia |
RTAS | 1 |