Zimeng Zhou

dblp:160/1659 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 TensorNTT: Architecture-Aware Optimizations for Number-Theoretic Transform on Tensor Core Unit
Xiangkai Yin, Shuoyu Wang, Zimeng Zhou, Lei Ju 0001, Zhuoran Ji
IEEE Big Data4
2025 FPGA-TrustZone: Security Extension of TrustZone to FPGA for SoC-FPGA Heterogeneous Architecture
abstract
To address the growing security issues faced by ARM-based mobile devices today, TrustZone was adopted to provide a trusted execution environment (TEE) to protect sensitive data. Such TrustZone-based models have been proven to be effective, but they target CPU architectures and do not work for the security of widely used heterogeneous computing platforms such as FPGAs. To solve this issue, we propose a comprehensive SoC-FPGA security framework, FPGA-TrustZone, to support FPGA TEE by extending the security of ARM TrustZone. Experiments on real SoC-FPGA hardware development boards show that FPGA-TrustZone provides high security with low performance overhead.
Xindong Fan, Shuchen Wang, Lei Ju 0001, Zimeng Zhou
DAC6
2025 A Weakly Centralized Hierarchical Sensitive Data Sharing Scheme Based on Edge Computing
Lifeng Ma, Chuanlin Huang, Shaodong Feng, Yongwang Liu, Zimeng Zhou, Kaifa Zheng
ICA3PP (4)8
2025 Adaptive ML-KEM: A Configurable HW-SW Architecture for Post-Quantum Cryptography
abstract
This paper presents a hardware/software (HW/SW) co-designed Module-Lattice-based Key-Encapsulation Mechanism (ML-KEM) accelerator featuring an ARM processor for runtime scheduling and a reconfigurable FPGA for efficient cryptographic kernel execution. Our design supports seamless, bitstream-free switching among ML-KEM's three security levels (1/3/5), enabling real-time key generation, encapsulation, and decapsulation. We propose a Preprocessed Multi-path Delay Commutator NTT (PMDC-NTT) architecture, achieving 15.4% logic reduction and 14.2% latency improvement over conventional MDC-NTT. The proposed HW/SW co-design achieves over$6.7 \times$speedup through pipelined task parallelism, demonstrating significant advantages in both performance and flexibility over prior works.
Lei Ju 0001, Zimeng Zhou
ICCD5
2024 Cache-aware Task Decomposition for Efficient Intermittent Computing Systems
abstract
Energy harvesting offers a scalable and cost-effective power solution for IoT devices, but it introduces the challenge of frequent and unpredictable power failures due to the unstable environment. To address this, intermittent computing has been proposed, which periodically backs up the system state to non-volatile memory (NVM), enabling robust and sustainable computing even in the face of unreliable power supplies. In modern processors, write back cache is extensively utilized to enhance system performance. However, it poses a challenge during backup operations as it buffers updates to memory, potentially leading to inconsistent system states. One solution is to adopt a write-through cache, which avoids the inconsistency issue but incurs increased memory access latency for each write reference. Some existing work enforces a cache flushing before backups to maintain a consistent system state, resulting in significant backup overhead. In this paper, we point out that although cache delays updates to the main memory, it may preserve a recoverable system state in the main memory. Leveraging this characteristic, we propose a cache-aware task decomposition method that divides an application into multiple tasks, ensuring that no dirty cache lines are evicted during their execution. Furthermore, the cache-aware task decomposition maintains an unchanged memory state during the execution of each task, enabling us to parallelize the backup process with task execution and effectively hide the backup latency. Experimental results with different power traces demonstrate the effectiveness of the proposed system.
Wei Zhang 0173, Mengying Zhao, Zimeng Zhou, Lei Ju 0001
DAC4
2024 Optimizing code allocation for hybrid on-chip memory in IoT systems
Zimeng Zhou, Fang-Wei Fu 0001
Integr.2
2024 A new McEliece-type cryptosystem using Gabidulin-Kronecker product codes
Jincheng Zhuang, Zimeng Zhou, Fang-Wei Fu 0001
Theor. Comput. Sci.3
2023 Accelerating DNN Inference with Heterogeneous Multi-DPU Engines
abstract
The Deep Learning Processor (DPU) programmable engine released by the official Xilinx Vitis AI toolchain has become one of the commercial off-the-shelf (COTS) solutions for Convolutional Neural Networks (CNNs) inference on Xilinx FPGAs. While modern FPGA devices generally have enough hardware resources to accommodate multi-DPUs simultaneously, the Xilinx toolchain currently only supports the deployment of multiple homogeneous DPUs engines that running independent inference tasks (task-level parallelism). In this work, we demonstrate that deployment of multiple heterogeneous DPU engines makes better resource efficiency for a given FPGA device. Moreover, we show that pipelined execution of a CNN inference task over heterogeneous multi-DPU engines may further improve overall inference throughput with carefully designed CNN layers-to-DPU mapping and scheduling. Finally, for a given CNN model and an FPGA device, we propose a comprehensive framework that automatically determines the optimal heterogeneous DPU deployment, and adaptively chooses the execution scheme between task-level and pipelined parallelism. Compared with the state-of-the-art solution with homogeneous multi-DPU engines and network-level parallelism, the proposed framework shows an average improvement of 13% (up-to 19%) and 6.6% (up-to 10%) on the Xilinx Zynq UltraScale+ MPSoC ZCU104 and ZCU102 platforms, respectively.
Zelin Du, Wei Zhang 0173, Zimeng Zhou, Zili Shao, Lei Ju 0001
DAC3
2023 Work or Sleep: Freshness-Aware Energy Scheduling for Wireless Powered Communication Networks with Interference Consideration
abstract
This paper explores how to schedule energy to optimize the information freshness in wireless powered communication networks (WPCNs) when considering channel interference among adjacent sensor nodes. We introduce Age of Information (AoI) to quantitatively evaluate the information freshness and formulate the AoI optimization problem. Unlike prior works focusing on system optimization for WPCNs while ignoring channel interference in energy transfer, this work reveals situations where channel interference among adjacent sensor nodes cannot be neglected and explores optimizing information freshness with interference consideration. To take the phenomena into account, we propose an energy scheduling solution to detect the channel interference and then judiciously determine the energy and time allocation for individual sensor nodes to improve the AoI performance as well as the system throughput. We implement a multi-node WPCN testbed to validate the functional correctness of the proposed solution, and extensive experiments have demonstrated the effectiveness of the proposed solution. The experimental results show that the proposed solution can reduce the average AoI by 54.7% and the average throughput by 49.8% on average compared to the state-of-the-art solutions.
Lei Ju 0001, Chun Jason Xue, Mingliang Zhou 0001, Wei Zhang 0173, Zimeng Zhou
DAC6
2023 Adaptive Task-Based Intermittent Computing System With Parallel State Backup
abstract
Energy harvesting promises to power billions of Internet of Things devices without being restricted by battery life. Since the energy harvester generally outputs weak and unstable energy, the system may suffer frequent and unpredictable power failures, thus falling into cyclically reboots without forward progress. The task-based intermittent computing system which periodically backs up system states into nonvolatile memory (NVM) is proposed to solve the nonprogress problem, with the nontrivial cost of frequent backups. How to reduce the backup overhead becomes a major research problem for intermittent computing. This article, for the first time, proposes to parallelize state backup and program execution with asynchronous direct memory access (DMA) to hide the backup latency into the program’s execution. But, straightforwardly executing the state backup and the program in parallel may cause an inconsistent system state. In specific, the system state may be modified by the program during backup, and therefore may be backed up incorrectly and further cause the system to deliver an incorrect computation result. We make a deep analysis on the system behavior and observe that, although the system state may be backed up incorrectly, the incorrect backup will be covered by the subsequent correct backups soon as the backup operations are performed frequently. In addition, only a small part of variables among all the program states may cause incorrect computation result. So, in this article, we aggressively allow incorrect backups to occur and propose a backup error detection method and a fault-tolerant backup management to guarantee the correctness of the system’s execution. To augment the parallel backup method, an adaptive execution method is further proposed to reduce the number of backups and balance the ratio between task execution time and backup latency. We design a run-time system to implement the proposed approach, and experimental results conducted on an STM32F7-based platform show that the proposed method can achieve a$2.6\times $average speedup.
Wei Zhang 0173, Qianling Zhang, Mingsong Lv, Songran Liu, Zimeng Zhou, Qiulin Chen, Nan Guan, Lei Ju 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 Optimizing Worst Case Data Freshness in RF-Powered Networked Embedded Systems
abstract
Maintaining real-time data freshness plays a critical role in ensuring system correctness and optimizing the system performance in networked embedded systems (NESs). To quantitatively measure the freshness of the collected real-time data, the concept of Age of Information (AoI) has been extensively studied in recent years. This article explores how to minimize the worst case AoI of real-time data in radio-frequency (RF)-powered NESs. In such systems, one hybrid access point (HAP) transfers wireless power to a set of distributed sensor nodes, and in the meantime, receives the information from these sensor nodes. We utilize the metric of AoI to measure the data freshness and present a comprehensive analysis of the worst case AoI of the real-time data in the target system. Based on the analysis, an optimal energy schedule solution is designed to judiciously determine individual sensor nodes’ energy and time allocation to minimize the worst case AoI. Considering the varying importance of different information and sensor nodes in the target system, we further propose the optimal time and energy allocation scheme for minimizing the weighted worst case AoI. A multinode RF-powered NES testbed is implemented to validate the functional correctness of our solutions. The results show that our solutions significantly outperform the state-of-the-art solutions, reducing the worst case AoI and weighted worst case AoI by 69.3% and 75.1% on average, respectively.
Zimeng Zhou, Chenchen Fu, Chun Jason Xue, Song Han 0002, Wei Zhang 0173, Lei Ju 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 coxHE: A software-hardware co-design framework for FPGA acceleration of homomorphic computation
abstract
Data privacy becomes a crucial concern in the AI and big data era. Fully homomorphic encryption (FHE) is a promising data privacy protection technique where the entire computation is performed on encrypted data. However, the dramatic increase of the computation workload restrains the usage of FHE for the real-world applications. In this paper, we propose an FPFA accelerator design framework for CKKS-based HE. While the KeySwitch operations are the primary performance bottleneck of FHE computation, we propose a low latency design of KeySwitch module with reduced intra-operation data dependency. Compared with the state-of-the-art FPGA based key-switch implementation that is based on Verilog, the proposed high-level synthesis (HLS) based design reduces the operation latency by 40%. Furthermore, we propose an automated design space exploration framework which generates optimal encryption parameters and accelerators for a given application kernel and the target FPGA device. Experimental results for a set of real HE application kernels on different FPGA devices show that our HLS-based flexible design framework produces substantially better accelerator design compared with a fixed-parameter HE accelerator in terms of security, approximation error, and overall performance.
Mingqin Han, Yilan Zhu, Qian Lou, Zimeng Zhou, Shanqing Guo, Lei Ju 0001
DATE4
2020 FPGA-based Compaction Engine for Accelerating LSM-tree Key-Value Stores
abstract
With the rapid growth of big data, LSM-tree based key-value stores are widely applied due to its high efficiency in write performance. Compaction plays a critical role in LSM-tree, which merges old data and could significantly reduce the overall throughput of the whole system especially for write-intensive workloads. Hardware acceleration for database is a popular trend in recent years. In this paper, we design and implement an FPGA-based compaction engine to accelerate compaction in LSM-tree based key-value stores. To take full advantage of the pipeline mechanism on FPGA, the key-value separation and index-data block separation strategies are proposed. In order to improve the compaction performance, the bandwidth of FPGA-chip is fully utilized. In addition, the proposed acceleration engine is integrated with a classic LSM-tree based store without modifications on the original storage format. The experimental results demonstrate that the proposed FPGA-based compaction engine can achieve up to 92.0x acceleration ratio compared with CPU baseline, and achieve up to 6.4x improvement on the throughput of random writes.
Xuan Sun 0003, Jinghuan Yu, Zimeng Zhou, Chun Jason Xue
ICDE3
2020 Maintaining Real-Time Data Freshness in Wireless Powered Communication Networks
abstract
This paper studies how to maintain real-time data freshness in the emerging wireless powered communication networks (WPCNs). In a WPCN, one or multiple hybrid access points (HAPs) with constant power supply transfer the energy and receive the status update through wireless to and from a set of distributed sensor nodes simultaneously. We utilize the concept of Age of Information (AoI) to quantitatively measure the freshness of the sensor data, and formulate the AoI optimization problem. Based on this problem formulation, we first explore the optimal time allocation method to minimize the average AoI for sensor nodes in single-hop WPCNs. This method is then extended to derive the optimal time allocation for two-hop WPCNs, where some sensor nodes may transmit status update to the HAP through relay nodes. In this more complex scenario, we apply a two-phase method to 1) identify all the two-hop node candidates and their associated relay nodes, and 2) determine the best assignment and the corresponding time allocation to optimize the average AoI. A 5-node WPCN testbed is developed to validate the functional correctness of the proposed methods. Extensive simulations are also conducted for performance evaluation under more comprehensive settings. The experimental results show that the proposed methods can reduce the average AoI by 83.4% on average compared to the state-of-the-art methods.
Zimeng Zhou, Zelin Yun, Chenchen Fu, Chun Jason Xue, Song Han 0002
RTSS1
2020 3D hypothesis clustering for cross-view matching in multi-person motion capture
abstract
We present a multiview method for markerless motion capture of multiple people. The main challenge in this problem is to determine crossview correspondences for the 2D joints in the presence of noise. We propose a 3D hypothesis clustering technique to solve this problem. The core idea is to transform joint matching in 2D space into a clustering problem in a 3D hypothesis space. In this way, evidence from photometric appearance, multiview geometry, and bone length can be integrated to solve the clustering problem efficiently and robustly. Each cluster encodes a set of matched 2D joints for the same person across different views, from which the 3D joints can be effectively inferred. We then assemble the inferred 3D joints to form full-body skeletons for all persons in a bottom-up way. Our experiments demonstrate the robustness of our approach even in challenging cases with heavy occlusion, closely interacting people, and few cameras. We have evaluated our method on many datasets, and our results show that it has significantly lower estimation errors than many state-of-the-art methods.
Miaopeng Li, Zimeng Zhou, Xinguo Liu
Comput. Vis. Media2
2020 Energy-Constrained Data Freshness Optimization in Self-Powered Networked Embedded Systems
abstract
This article explores how to optimize the freshness of real-time data for energy harvesting (EH)-based networked embedded systems (NESs) with energy constraints. We introduce the concept of age of information (AoI) to quantitatively measure the data freshness and present a comprehensive analysis on the average AoI of the real-time data with stochastic update arrival and energy replenishment patterns for single-source EH-based systems. An optimal offline solution and an effective online solution are designed to select a sequence of real-time data updates (while discarding the remaining ones) and determine their corresponding transmission time to minimize the average AoI. We further extend these findings to multisource EH-based NESs, and present an optimal offline solution and an efficient online solution to schedule updates for each data source to optimize the average AoI. The correctness of the analysis and the effectiveness of the proposed solutions have been validated through extensive experiments by comparing to the state-of-the-art methods. According to the experimental results, the proposed solutions reduce the average AoI by 47.2% and 69.1% on average comparing to the state-of-the-art solutions for single-source and multisource EH-based NESs, respectively, with low harvesting rates.
Zimeng Zhou, Chenchen Fu, Chun Jason Xue, Song Han 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2020 Cross Refinement Techniques for Markerless Human Motion Capture
abstract
This article presents a global 3D human pose estimation method for markerless motion capture. Given two calibrated images of a person, it first obtains the 2D joint locations in the images using a pre-trained 2D Pose CNN, then constructs the 3D pose based on stereo triangulation. To improve the accuracy and the stability of the system, we propose two efficient optimization techniques for the joints. The first one, called cross-view refinement, optimizes the joints based on epipolar geometry. The second one, called cross-joint refinement, optimizes the joints using bone-length constraints. Our method automatically detects and corrects the unreliable joint, and consequently is robust against heavy occlusion, symmetry ambiguity, motion blur, and highly distorted poses. We evaluate our method on a number of benchmark datasets covering indoors and outdoors, which showed that our method is better than or on par with the state-of-the-art methods. As an application, we create a 3D human pose dataset using the proposed motion capture system, which contains about 480K images of both indoor and outdoor scenes, and demonstrate the usefulness of the dataset for human pose estimation.
Miaopeng Li, Zimeng Zhou, Xinguo Liu
ACM Trans. Multim. Comput. Commun. Appl.2
2019 Transmit or Discard: Optimizing Data Freshness in Networked Embedded Systems with Energy Harvesting Sources
abstract
This paper explores how to optimize the freshness of real-time data in energy harvesting based networked embedded systems. We introduce the concept of Age of Information (AoI) to quantitatively measure the data freshness and present a comprehensive analysis on the average AoI of the real-time data with stochastic update arrival and energy replenishment rates. Both an optimal offline solution and an effective online solution are designed to judiciously select a subset of the real-time data updates and determine their corresponding transmission times to optimize the average AoI subject to energy constraints. Our extensive experiments have validated the effectiveness of the proposed solutions, and showed that these two methods can significantly improve the average AoI by 47.2% comparing to the state-of-the-art solutions for low energy replenishment rate.
Zimeng Zhou, Chenchen Fu, Chun Jason Xue, Song Han 0002
DAC1
2019 A Hardware-Accelerated Solution for Hierarchical Index-Based Merge-Join(Extended Abstract)
abstract
Hardware acceleration through field programmable gate arrays (FPGAs) has recently become a technique of growing interest for many data-intensive applications. Join query is one of the most fundamental database query types useful in relational database management systems. However, the available solutions so far have been beset by higher costs in comparison with other query types. In this paper, we develop a novel solution to accelerate the processing of sort-merge join queries with low match rates. Specifically, our solution makes use of hierarchical indexes to identify result-yielding regions in the solution space in order to take advantage of result sparseness. Further, in addition to one-dimensional equi-join query processing, our solution supports processing of multidimensional similarity join queries. Experimental results show that our solution is superior to the best existing method in a low match rate setting; the method achieves a speedup factor of 4.8 for join queries with a match rate of 5%.
Zimeng Zhou, Chenyun Yu, Sarana Nutanong, Yufei Cui, Chenchen Fu, Chun Jason Xue
ICDE1
2019 A Hardware-Accelerated Solution for Hierarchical Index-Based Merge-Join
abstract
Hardware acceleration through field programmable gate arrays (FPGAs) has recently become a technique of growing interest for many data-intensive applications. Join query is one of the most fundamental database query types useful in relational database management systems. However, the available solutions so far have been beset by higher costs in comparison to other query types. In this paper, we develop a novel solution to accelerate the processing of sort-merge join queries with low match rates. Specifically, our solution makes use of hierarchical indexes to identify result-yielding regions in the solution space in order to take advantage of result sparseness. Further, in addition to one-dimensional equi-join query processing, our solution supports processing of multidimensional similarity join queries. Experimental results show that our solution is superior to the best existing method in a low match rate setting; the method achieves a speedup factor of 4.8 for join queries with a match rate of 5 percent.
Zimeng Zhou, Chenyun Yu, Sarana Nutanong, Yufei Cui, Chenchen Fu, Chun Jason Xue
IEEE Trans. Knowl. Data Eng.1
2019 Multi-Person Pose Estimation Using Bounding Box Constraint and LSTM
abstract
This paper presents a new method for single-image pose estimation of multiple people combining the traditional bottom-up and the top-down methods. Specifically, we extract features from the input image by a residual network and use a multistage CNN to learn both the confidence maps of joints and the connection relationships, between joints. During testing, we perform the network feedforwarding in a bottom-up manner, and then use the predicted confidence maps, the connection relationships, and the corresponding bounding boxes to parse the poses of all people in a top-down manner. In contrast to the previous top-down methods, our method is robust to bounding box shift and tightness, works well for largely overlapped people, and achieves faster running speed. In contrast to the bottom-up method, our method avoids mistake propagation across different people, and addresses disconnected joints effectively. To estimate human pose from videos, we impose a weight-sharing scheme to the multi-stage CNN, and rewrite it as a recurrent neural network. Thus, we can reuse the prediction results from the previous frames so as to reduce the total stage number, yielding significantly faster speed in invoking the network on videos. And we adopt LSTM units between frames to capture the temporal correlation among video frames. We found that LSTM handles input-quality degradation in videos well and successfully stabilizes the sequential outputs.
Miaopeng Li, Zimeng Zhou, Xinguo Liu
IEEE Trans. Multim.2
2018 Bottom-up Pose Estimation of Multiple Person with Bounding Box Constraint
abstract
In this work, we propose a new method for multi-person pose estimation which combines the traditional bottom-up and the top-down methods. Specifically, we perform the network feed-forwarding in a bottom-up manner, and then parse the poses with bounding box constraints in a top-down manner. In contrast to the previous top-down methods, our method is robust to bounding box shift and tightness. We extract features from an original image by a residual network and train the network to learn both the confidence maps of joints and the connection relationships between joints. During testing, the predicted confidence maps, the connection relationships and the bounding boxes are used to parse the poses of all persons. The experimental results showed that our method learns more accurate human poses especially in challenging situations and gains better time performance, compared with the bottom-up and the top-down methods.
Miaopeng Li, Zimeng Zhou, Xinguo Liu
ICPR2
2015 Managing hybrid on-chip scratchpad and cache memories for multi-tasking embedded systems
abstract
On-chip memory management is essential in design of high performance and energy-efficient embedded systems. While many off-the-shelf embedded processors employ a hybrid on-chip SRAM architecture including both scratchpad memories (SPMs) and caches, many existing work on SPM management ignore the synergy between caches and SPMs. In this work, we propose a static SPM allocation strategy for the hybrid on-chip memory architecture in a multi-tasking environment, which minimizes the overall access latency and energy consumption of the instruction memory subsystem. We capture cache conflict misses via a fine-grained temporal cache behavior model. An integer linear programming (ILP) based formulation is proposed to generate an function-level SPM allocation scheme, where both intra- and inter-task cache interference as well as access frequency are captured for an optimal memory subsystem design. Compared with the state-of-the-art static SPM allocation strategy in a multitasking environment, experimental results show that our SPM management scheme achieves 30.51% further improvement in instruction memory subsystem performance, and up to 34.92% in terms of energy saving.
Zimeng Zhou, Lei Ju 0001, Zhiping Jia, Xin Li 0002
ASP-DAC1