Mingzhe Zhang 0005

dblp:118/5481-5 · DBLP profile ↗
← Back
49ranked-venue papers
4as first author
37since 2021 · last 2026
0000-0002-6440-7550ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 42 · 4 first-author · 31 since 2021Software engineering, systems software and programming languages · 9 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Security and privacy · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Framework for Developing and Optimizing Fully Homomorphic Encryption Programs on GPUs
abstract
In sensitive domains such as healthcare and finance, machine learning increasingly employs Fully Homomorphic Encryption (FHE) to secure both user data and models. Although FHE's intrinsic parallelism naturally aligns with GPU architectures, optimizing GPU kernels alone remains insufficient for efficient end-to-end FHE application development. The inherent complexity of FHE schemes and intricate GPU-specific details impede developers from focusing on high-level program logic. Additionally, FHE's high memory requirements, fine-grained memory operations, and redundant computations introduce further optimization challenges, resulting in inefficiencies even when GPU kernels are individually optimized. This paper introduces EasyFHE, a framework designed to simplify the development and optimization of GPU-accelerated FHE applications. Similar to PyTorch, EasyFHE provides high-level interfaces for defining computational logic while automatically handling low-level tasks, such as implementation selection and memory management. Furthermore, it incorporates an optimization framework that systematically addresses performance bottlenecks by applying tailored optimization passes during the lowering from high-level FHE programs to GPU kernels. Compared to state-of-the-art open-source GPU FHE libraries, EasyFHE uniquely supports FHE programs with memory requirements exceeding typical GPU capacities, achieving an average speedup of 2.88× with a peak of 4.39×.
Jianyu Zhao 0004, Xueyu Wu 0001, Guang Fan 0001, Mingzhe Zhang 0005, Shoumeng Yan, Lei Ju 0001, Zhuoran Ji
ASPLOS (2)4
2026 Falcon: Algorithm-Hardware Co-Design for Efficient Fully Homomorphic Encryption Accelerator
abstract
Fully homomorphic encryption (FHE) enables computation on encrypted data without compromising privacy, positioning it as a promising solution for secure cloud computing. However, its substantial computational overhead impedes practical deployment, prompting the development of dedicated hardware accelerators. In practice, when deploying cryptographic algorithm optimizations on FHE accelerators, hardware constraints typically such as limited memory capacity, often lead to a disparity between theoretical algorithmic advantage and achievable hardware efficiency.
Liang Kong 0005, Xianglong Deng, Guang Fan 0001, Shengyu Fan, Yilan Zhu, Geng Yang 0001, Yisong Chang, Shoumeng Yan, Mingzhe Zhang 0005
ASPLOS (2)10
2026 FHEFusion: Enabling Operator Fusion in FHE Compilers for Depth-Efficient DNN Inference
abstract
Operator fusion is essential for accelerating FHE-based DNN inference because it reduces multiplicative depth and, in turn, lowers the cost of ciphertext operations by keeping them at lower ciphertext levels. Existing approaches either rely on manual optimizations, which miss cross-operator opportunities, or on compiler pattern matching, which lacks generality. Standard DNN graphs omit FHE-specific behaviors, while fully lowering to primitive FHE operations introduces excessive granularity and obstructs effective optimization.We present FHEFusion, a compiler framework for the CKKS scheme that enables fusion through a new IR. This IR preserves high-level DNN semantics while introducing FHE-aware operators—masking and compaction (Strided_Slice)—that are central to CKKS, thereby exposing broader fusion opportunities. Guided by algebraic rules and an FHE-aware cost model, FHEFusion reduces multiplicative depth and identifies profitable fusions. Integrated into ANT-ACE, a state-of-the-art FHE compiler, FHEFusion outperforms NGRAPH, the only framework with graph-level fusion, achieving up to 3.02× (average 1.40×) speedup across seven DNNs (13 variants from different RELU approximations) on CPUs, while maintaining inference accuracy.
Tianxiang Sui, Jianxin Lai, Long Li 0015, Yan Liu 0082, Qing Zhu 0008, Linjie Xiao, Mingzhe Zhang 0005, Jingling Xue
CGO9
2026 Xerxes: Extensive Exploration of Scalable Hardware Systems with CXL-Based Simulation Framework
Yuda An, Shushu Yi, Bo Mao 0003, Qiao Li 0001, Mingzhe Zhang 0005, Diyu Zhou, Ke Zhou 0001, Nong Xiao 0001, Guangyu Sun 0003, Yingwei Luo, Jie Zhang 0048
FAST5
2026 An Efficient and Scalable Hardware Architecture for Number Theoretic Transform on FPGA with Design Automation
abstract
Fully Homomorphic Encryption (FHE) has become a promising approach to protecting data privacy in emerging application scenarios. Unfortunately, FHE suffers from significant processing speed degradation compared to plaintext computation, with one of the primary bottlenecks being the time-consuming Number Theoretic Transform (NTT). Therefore, accelerating NTT to accommodate various FHE parameters is crucial to advancing FHE towards practical use. With highly reconfigurable and performant logical fabrics, Field Programmable Gate Arrays (FPGAs) have exhibited great potential in NTT acceleration. By decomposing large-point NTT with strong data dependency into independent and simple small-point NTTs, the emerging Ten-step NTT (TNTT) algorithms intuitively enable higher parallelism and thereby have the potential to explore better performance compared to traditional algorithms. However, our quantitative analysis reveals that TNTT exhibits significant performance degradation as parallelism increases due to additional varying-size transpositions and Hadamard products. This paper proposes AutoNest, an efficient and scalable hardware architecture, along with an accelerator auto-generation framework for TNTT. The proposed hardware architecture maximizes performance by 1) adopting a 2D block decomposition dataflow to address critical path delays in transpose logic, thereby improving clock frequency. 2) integrating algorithm-level costfree twiddle factor fusion to reduce the number of modular multiplications in Hadamard products, thereby allowing higher parallelism on chip. Moreover, we also deliver an accelerator generation framework conducting automated design space exploration to elaborate a performant TNTT architecture under the target FPGAs' resource budget for user-defined FHE parameters. Experimental results on the AMD-Xilinx U280 FPGA demonstrate that NTT accelerators generated by AutoNest achieve an average speedup of$2.31 \times$compared to prior designs.
Yilan Zhu, Geng Yang 0001, Xingyu Tian, Dilshan Kumarathunga, Liang Kong 0005, Xianglong Deng, Shengyu Fan, Guang Fan 0001, Guiming Shi, Bo Zhang 0098, Yisong Chang, Shoumeng Yan, Zhenman Fang, Mingzhe Zhang 0005
HPCA15
2026 HyperDrive: Hierarchical Exploitation of Memory Efficiency for GPU-Based FHE Acceleration
Guang Fan 0001, Liang Kong 0005, Yilan Zhu, Geng Yang 0001, Shengyu Fan, Xianglong Deng, Fangyu Zheng, Jian Weng, Meng Li 0004, Yisong Chang, Shoumeng Yan, Mingzhe Zhang 0005
ISCA16
2026 MNEMOS: A GPU-Based TFHE Acceleration Framework with Memory Access Optimization
Xianglong Deng, Guang Fan 0001, Shengyu Fan, Mingzhe Zhang 0005
ISCA9
2026 A Bitwidth-Flexible Modular Multiplier with Shift-Free Accumulation for Efficient NTT Acceleration in FHE
Shengyu Fan, Xianglong Deng, Rui Hou 0001, Mingzhe Zhang 0005
ISCAS6
2026 TensorFHE+: Fully Homomorphic Encryption Acceleration Based on Linear Algebra
abstract
Fully Homomorphic Encryption (FHE) enables encrypted data processing on untrusted cloud servers, crucial for privacy-sensitive applications. Despite its potential, performance overheads (about 10, 000× slower) limit adoption. ASIC accelerators outperform GPUs/FPGAs by optimizing specific operations but rely on costly 7nm processes and large on-chip memory, hindering cost-effective deployment. Balancing efficiency with manufacturing constraints remains critical. This paper presents TensorFHE+, a GPU-optimized FHE acceleration framework leveraging Tensor Cores to accelerate Number Theoretic Transform (NTT) operations. Key innovations include: 1) Decomposing CKKS kernels into vector/matrix operations for hardware utilization; 2) Vectorized modulo arithmetic; 3) Data layout optimization for memory efficiency. Evaluated on NVIDIA A100, TensorFHE+ outperforms TensorFHE [1] by 1.44× in average (up to 1.69× on ResNet-20) and surpasses prior GPU implementations [2], [3]. The design also demonstrates compatibility with commercial linear algebra accelerators, enabling efficient FHE deployment.
Yintai Sun, Shengyu Fan, Zhenhua Yin, Xinkai Song, Xing Hu 0001, Zidong Du, Qi Guo 0001, Weizhi Xu 0001, Rui Hou 0001, Dan Meng 0002, Song Bian 0001, Mingzhe Zhang 0005
IEEE Trans. Computers12
2026 Exploration of Karatsuba Algorithm for Efficient Barrett Modular Multiplication
Bo Zhang 0098, Mingzhe Zhang 0005, Shoumeng Yan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 The Future of Fully Homomorphic Encryption System: From a Storage I/O Perspective
Erci Xu, Shengyu Fan, Xianglong Deng, Guiming Shi, Guang Fan 0001, Liang Kong 0005, Yilan Zhu, Shoumeng Yan, Mingzhe Zhang 0005
APPT11
2025 WPC: Weight Plaintext Compression for CNN Inference based on RNS-CKKS
abstract
Convolutional neural network (CNN) inference based on RNS-CKKS enables secure processing on encrypted data but introduces significant weight size overhead. Weight plaintext, weight in RNS-CKKS format, can reach tens to hundreds of gigabytes. Existing compression methods either add high computational cost or yield low compression rates. In this work, we propose WPC, Weight Plaintext Compression, to compress weight plaintext for RNS-CKKS-based CNN inference. We observe that the transformation from the weight in CNN models to the weight plaintext in RNS-CKKS format involves an operation akin to the Discrete Fourier Transform, which shifts data between the time and frequency domains while retaining redundant information from periodic and discrete data. Based on this observation, we first introduce the Periodic Transmit Theorem, which states that periodic patterns can be preserved during the transformation process, thereby enabling compression. We then propose Channel Innermost Packing Scheme and Rotation Padding to rearrange the weight data into periodic patterns for compression. Results show that WPC achieves 1.25 to 2.18 times speedup on an A100 GPU and 46.08 to 139.11 times compression rate.
Guiming Shi, Shengyu Fan, Xianglong Deng, Liang Kong 0005, Jingwei Cai, Shuwen Deng, Mingzhe Zhang 0005, Kaisheng Ma
CCS9
2025 Corrosion Hammer: A Self-Activated Bit-Flip Attack to the Processing-In-Memory Accelerator
abstract
In this paper, taking ReRAM-based PIM accelerators as an example, we present a novel attack framework called Corrosion Hammer, which builds based on the Bit Flip Attack (BFA).Unlike previous BFA methods that require explicit memory fault injection techniques, such as Row Hammer, to modify sensitive bits in the victim Neural Network model, Corrosion Hammer implants the trojan during the hardware-software co-design phase and flips sensitive bits using read disturbance, which is a common noise in ReRAM caused by normal read operations.Furthermore, we explore the impact of inputs on the activation time consumption of the trojan and propose a method to expedite activation using normal input.Our experimental results demonstrate that Corrosion Hammer achieves an extremely covert trojan implantation and activation method, with an adversarial attack success rate of 92.46%.Additionally, using a specially designed method, Trojan activation is 61.98× faster compared to activation in an undisturbed normal operation state.It provides a way to significantly speed up the Trojan activation.
Mengxin Zheng, Shengyu Fan, Qian Lou, Rui Hou 0001, Dan Meng 0002, Mingzhe Zhang 0005
CF7
2025 ALLMod: Exploring Area-Efficiency of LUT-based Large Number Modular Reduction via Hybrid Workloads
abstract
Modular arithmetic, particularly modular reduction, is widely used in cryptographic applications such as homomorphic encryption (HE) and zero-knowledge proofs (ZKP). High-bit-width operations are crucial for enhancing security; however, they are computationally intensive due to the large number of modular operations required. The lookup-table-based (LUT-based) approach, a “space-for-time” technique, reduces computational load by segmenting the input number into smaller bit groups, pre-computing modular reduction results for each segment, and storing these results in LUTs. While effective, this method incurs significant hardware overhead due to extensive LUT usage. In this paper, we introduce ALLMod, a novel approach that improves the area efficiency of LUT-based largenumber modular reduction by employing hybrid workloads. Inspired by the iterative method, ALLMod splits the bit groups into two distinct workloads, achieving lower area costs without compromising throughput. We first develop a template to facilitate workload splitting and ensure balanced distribution. Then, we conduct design space exploration to evaluate the optimal timing for fusing workload results, enabling us to identify the most efficient design under specific constraints. Extensive evaluations show that ALLMod achieves up to $\lt sup\gt1\lt/sup\gt|.65 \times$ and $3 \times$ improvements in area efficiency over conventional LUT-based methods for bit-widths of 128 and 8,192, respectively.
Fangxin Liu, Haomin Li 0002, Zongwu Wang, Bo Zhang 0098, Mingzhe Zhang 0005, Shoumeng Yan, Li Jiang 0002, Haibing Guan
DAC5
2025 WarpDrive: GPU-Based Fully Homomorphic Encryption Acceleration Leveraging Tensor and CUDA Cores
abstract
The application of Fully Homomorphic Encryption (FHE) is rapidly gaining traction as a means to maintain data confidentiality while performing computations on encrypted data. Given the accessibility and computational power, GPUs hold promise for significantly accelerating FHE operations. However, existing GPU-based acceleration solutions face several formidable challenges, notably the extensive occurrence of pipeline stalls induced by memory access and suboptimal harnessing of GPU hardware. This paper presents WarpDrive, a comprehensive framework for GPU-based FHE acceleration. Through sophisticated computation decomposition and fine-grained memory access design, WarpDrive significantly reduces the number of instructions by $\mathbf{7 3 \%}$ and pipeline stalls by $\mathbf{8 6 \%}$ compared to the state-of-the-art solution. Additionally, WarpDrive features a framework that supports the concurrent utilization of CUDA Cores and Tensor Cores within the NTT operation, for the first time, achieving performance that surpasses that of any single type of processing unit. Furthermore, we fully exploit the intra-ciphertext parallelism to elevate both computation and memory utilization, achieving up to $2.12 \times$ improvements without the need for ciphertext batching. Experimental results demonstrate that our optimizations highly enhance the performance of homomorphic operations. On an NVIDIA A100 GPU, WarpDrive achieves a throughput of 1218 KOPS for NTT and 305 KOPS for homomorphic multiplication, outperforming the state-of-the-art GPU solution (TensorFHE) by factors of $13.4 \times$ and $3.5 \times$, respectively. For the specific FHE workload, even under a much smaller batch size, our approach achieves $2.8 \times$ the performance of TensorFHE.
Guang Fan 0001, Mingzhe Zhang 0005, Fangyu Zheng, Shengyu Fan, Xianglong Deng, Wenxu Tang, Liang Kong 0005, Shoumeng Yan
HPCA2
2025 CipherPrune: Efficient and Scalable Private Transformer Inference
abstract
Private Transformer inference using cryptographic protocols offers promising solutions for privacy-preserving machine learning; however, it still faces significant runtime overhead (efficiency issues) and challenges in handling long-token inputs (scalability issues). We observe that the Transformer's operational complexity scales quadratically with the number of input tokens, making it essential to reduce the input token length. Notably, each token varies in importance, and many inputs contain redundant tokens. Additionally, prior private inference methods that rely on high-degree polynomial approximations for non-linear activations are computationally expensive. Therefore, reducing the polynomial degree for less important tokens can significantly accelerate private inference. Building on these observations, we propose \textit{CipherPrune}, an efficient and scalable private inference framework that includes a secure encrypted token pruning protocol, a polynomial reduction protocol, and corresponding Transformer network optimizations. At the protocol level, encrypted token pruning adaptively removes unimportant tokens from encrypted inputs in a progressive, layer-wise manner. Additionally, encrypted polynomial reduction assigns lower-degree polynomials to less important tokens after pruning, enhancing efficiency without decryption. At the network level, we introduce protocol-aware network optimization via a gradient-based search to maximize pruning thresholds and polynomial reduction conditions while maintaining the desired accuracy. Our experiments demonstrate that CipherPrune reduces the execution overhead of private Transformer inference by approximately $6.1\times$ for 128-token inputs and $10.6\times$ for 512-token inputs, compared to previous methods, with only a marginal drop in accuracy. The code is publicly available at https://github.com/UCF-Lou-Lab-PET/cipher-prune-inference.
Yancheng Zhang, Mengxin Zheng, Mimi Xie, Mingzhe Zhang 0005, Lei Jiang 0001, Qian Lou
ICLR5
2025 Analysis of Bit-Flip Attacks on Encrypted Neural Networks
abstract
With the swift progression of artificial intelligence and deep learning, neural networks have achieved remarkable success in domains such as image recognition, natural language processing, and autonomous driving. However, the proliferation of data scales and the extensive deployment of computational resources have engendered significant privacy concerns for users. In scenarios involving personal sensitive data, the safeguarding of privacy is of utmost importance. Homomorphic encryption technology, particularly the CKKS scheme, is capable of performing computations with minimal computational error while preserving data privacy, and it has been extensively utilized in encrypted neural networks. This paper studies Bit-Flip Attacks (BFAs) on encrypted neural networks under the RNS-CKKS scheme. We empirically analyze the effects of bit flips at different memory locations—covering ciphertext data, model weights, and evaluation keys—and report their observable outcomes (silent misclassification, irregular yet decodable outputs, or computation aborts). Under a realistic threat model where the adversary cannot precisely target bytes nor observe model predictions, BFAs can corrupt results but do not leak additional information. Our findings indicate that key corruption often produces conspicuous anomalies that offer detection potential for defenders.
Yilan Zhu, Rui Hou 0001, Dan Meng 0002, Shengyu Fan, Mingzhe Zhang 0005
ICPADS6
2025 FAST: An FHE Accelerator for Scalable-parallelism with Tunable-bit
abstract
Fully Homomorphic Encryption (FHE) enables direct computation on encrypted data, providing substantial security advantages in cloud-based modern society.However, FHE suffers from significant computational overhead compared to plaintext computation, hindering its adoption in real-world applications.While many accelerators have been designed to address performance bottlenecks, most do not fully leverage cryptographic optimization technologies, leaving room for further performance enhancements.In this work, we propose FAST, an FHE accelerator incorporating recent cryptographic optimizations, including hoisting technology and the gadget decomposition key-switching method (named KLSS method).We analyze ciphertext level consumption throughout application execution and observe that workload requirements vary significantly with different ciphertext levels for both hybrid and KLSS key-switching methods.Additionally, we note the differing computational precision requirements for these key-switching methods.Based on these observations, we designed a versatile framework that supports multiple key-switching methods during a single application execution and integrates hoisting technology.
Shengyu Fan, Xianglong Deng, Liang Kong 0005, Guiming Shi, Guang Fan 0001, Dan Meng 0002, Rui Hou 0001, Mingzhe Zhang 0005
ISCA8
2025 Neo: Towards Efficient Fully Homomorphic Encryption Acceleration using Tensor Core
abstract
Fully Homomorphic Encryption (FHE) is an emerging cryptographic technique for privacy-preserving computation, which enables computations on the encrypted data.Nonetheless, the massive computational demands of FHE prevent its further application to real-world workloads.To tackle this problem, several studies focus on the ASIC-based acceleration for FHE.However, the rapid evolution of FHE algorithms poses challenges to the generality of ASIC accelerator design.By contrast, a number of works rely on GPGPUs for FHE accelerations, due to the high parallelism and flexibility provided by GPGPUs.In this work, we propose a GPGPU-based acceleration solution that supports the Cheon-Kim-Kim-Song (CKKS) scheme by further exploiting Tensor Core(TCU) capabilities.In our study, we * Both author contributed equally to this research.
Xianglong Deng, Shengyu Fan, Dan Meng 0002, Rui Hou 0001, Mingzhe Zhang 0005
ISCA8
2025 XHarvest: Rethinking High-Performance and Cost-Efficient SSD Architecture with CXL-Driven Harvesting
abstract
The occasional nature of I/O bursts in production clusters makes the substantial and expensive SSD internal hardware resources (e.g., computation and memory resources) always underutilized, resulting in cost inefficiency.Open-Channel SSD (OCSSD), as a pioneering solution, removes the SSD internal resources but rather leverages the host-side resources to serve I/O requests.Unfortunately, it faces adoption obstacles due to the heavy resource contention with user applications, hampered host-SSD collaboration, and proprietary firmware leakage risks.Tackling these challenges, we propose XHarvest, a new cost-efficient and high-performance SSD architecture, which harnesses compute express link (CXL) and trusted execution environment (TEE) to facilitate dynamic, efficient, and secure host resource harvesting.It reserves moderate SSD internal resources to isolate SSD internal tasks and applications under regular I/O loads while coping with occasional I/O bursts via dynamic host resource harvesting.To this end, XHarvest executes the firmware within the host-side TEE without disclosing sensitive
Shushu Yi, Xianzhang Chen, Chenxi Wang 0005, Shengwen Liang, Zhe Wang 0017, Nong Xiao 0001, Qiao Li 0001, Mingzhe Zhang 0005, Jie Zhang 0048
ISCA10
2025 HAWK: Fully Homomorphic Encryption Accelerator with Fixed-Word Key Decomposition Switching
Liang Kong 0005, Shengyu Fan, Xianglong Deng, Guang Fan 0001, Guiming Shi, Yilan Zhu, Geng Yang 0001, Shoumeng Yan, Mingzhe Zhang 0005
MICRO10
2025 LEGOSim: A Unified Parallel Simulation Framework for Multi-chiplet Heterogeneous Integration
Tiantian Lin, Xiaohang Wang 0001, Ling Wang 0005, Zhulin Zheng, Yingtao Jiang, Amit Kumar Singh 0002, Jieming Yin, Sihai Qiu, Mingzhe Zhang 0005, Kui Ren 0001
MICRO13
2025 LP-HENN: fully homomorphic encryption accelerator with high energy efficiency
abstract
Abstract Fully homomorphic encryption (FHE) enables direct computation on encrypted data without decryption, ensuring data privacy in cloud computing scenarios and preventing the leakage of sensitive information. However, the computational overhead of HE typically exceeds that of plaintext computation by 4 to 5 orders of magnitude, while energy consumption is 5 to 6 orders of magnitude higher. These substantial performance and energy overheads significantly hinder the widespread adoption of FHE. This paper proposed LP-HENN, a novel low-power and energy-efficient FHE accelerator architecture that leverages a RISC-V vector coprocessor and ReRAM crossbar arrays. LP-HENN targets power-constrained application scenarios such as edge devices, aiming to provide highly energy-efficient acceleration support for FHE applications. LP-HENN leverages the collaborative work of the vector processor and ReRAM crossbars, employing optimization strategies to achieve full pipelining and minimize memory access. Furthermore, this paper proposed a parameter selection model for early-stage architecture design, which achieves an optimal balance between performance and energy consumption through the collaborative optimization of multiple parameters. Experimental results show that, for an FHE-based convolutional neural network (HE-CNN) inference application, LP-HENN achieves a 31.82Ã- and 11920.56Ã- improvement in performance and energy efficiency, respectively, compared to CPU. Compared to FxHENN, the state-of-the-art FPGA-based FHE accelerator with high energy efficiency for edge devices, LP-HENN achieves a 2.36Ã- and 10.04Ã- improvement in performance and energy efficiency, respectively. The energy efficiency of LP-HENN is comparable to that of F1, the state-of-the-art ASIC FHE accelerator, while featuring a low power design suitable for edge computing.
Zhuoyu Tian, Shengyu Fan, Xianglong Deng, Rui Hou 0001, Dan Meng 0002, Mingzhe Zhang 0005
Cybersecur.7
2025 Corrosion Hammer: a self-activated bit-flip attack to the processing-in-memory accelerator
abstract
Abstract The Resistive Random-Access-Memory (ReRAM) crossbar-based Processing-In-Memory (PIM) accelerator shows great promise in accelerating neural networks (NNs). This technique boasts low energy consumption and exceptional performance in multiplication and accumulations (MAC) operations, making ReRAM-based PIM accelerators an ideal solution for intelligent computing in wearable and low-power mobile devices. However, security concerns related to PIM have not been adequately addressed. In this paper, we present a new attack framework called SolutionName for ReRAM-based PIM accelerators. SolutionName builds upon the Bit Flip Attack (BFA), a weight modification attack that manipulates the NN function by flipping specific bits in the deployed quantized NN model. Unlike previous BFA methods that require explicit memory fault injection techniques, such as Row Hammer, to modify sensitive bits in the victim NN, SolutionName implants the trojan during the hardware-software co-design phase and flips sensitive bits using read disturbance. Read disturbance is a common noise in ReRAM caused by normal read operations. This approach enables the trojan to be activated quietly during normal use, eliminating the need for explicit attacks. Furthermore, we explore the impact of inputs on the activation time of the trojan and propose a method to expedite activation using normal input. Our experimental results demonstrate that SolutionName achieves an extremely covert trojan implantation and activation method, with an adversarial attack success rate of 94.38%. Additionally, with a specially designed method, the trojan activation can be accelerated on average by 61.98 $$\times $$ × , providing controllable activation.
Mengxin Zheng, Shengyu Fan, Qian Lou, Rui Hou 0001, Dan Meng 0002, Mingzhe Zhang 0005
Cybersecur.7
2025 On Improving the Performance of Intra- and Inter-chiplet Interconnection Networks in Multi-chiplet Systems for Accelerating FHE Encrypted Neural Network Applications
abstract
Fully Homomorphic Encryption (FHE) is regarded as a promising way to protect data privacy with encrypted computation. Due to high computation overhead, hardware based FHE accelerators were proposed to speed up FHE applications. To support complicated FHE-encrypted neural network applications, multi-chiplet based FHE accelerators were further proposed for scaling up system size, whereas one of the challenges is designing efficient intra- and inter-chiplet interconnection networks to accelerate data transfer. Conventional regular topologies like mesh or Kite either lead to high inter-chiplet transmission latency or excessive power consumption as these topologies assume uniform bandwidth or radix for nodes/links, ignoring the highly irregular distribution of inter-chiplet communication volumes. On the other hand, the problem of generating customized intra- and inter-chiplet interconnection networks has high complexity and previous network-on-chip topology generation works cannot efficiently improve the performance of intra- and inter-chiplet interconnection networks. In this article, the intra- and inter-chiplet interconnection optimization problem is defined, aiming to minimize the execution time of FHE applications under cost and power constraints. To efficiently solve this problem, we propose a bilevel optimization algorithm, which decomposes the problem into three sub-problems: (1) FHE parameters selection, (2) task-to-core mapping, and (3) intra-/inter-chiplet interconnection network topology generation. These sub-problems are then solved iteratively. Experimental results demonstrate that our proposed method reduces execution time by 51.66%, 43.16%, 39.44%, 43.34%, and 27.70% compared with REED and four multi-chiplet based FHE accelerators with mesh, Kite, Butterfly, and Florets as inter-chiplet interconnection networks. Therefore, the proposed method can effectively accelerate FHE applications on large-scale multi-chiplet systems.
Zewei Lai, Jinhui Ye, Xiaohang Wang 0001, Zheang Fu, Amit Kumar Singh 0002, Yingtao Jiang, Kui Ren 0001, Mei Yang 0001, Sihai Qiu, Mingzhe Zhang 0005
ACM Trans. Embed. Comput. Syst.13
2025 An Efficient Delta Compression Framework Seamlessly Integrated into Inline Deduplication
abstract
Delta compression can complement data deduplication by further minimizing redundancy through the compression of non-duplicate data chunks. When adding delta compression to deduplication-based backup systems, however, two primary challenges arise that degrade performance of inline deduplication. First, extra I/Os are introduced along the critical paths of backup and restoration for retrieving base chunks, slowing the system. Second, rewriting techniques prohibit specific data chunks from serving as base chunks for delta compression to improve restore performance, resulting in a loss of compression efficiency. In this paper, we introduce LoopDelta, a framework that seamlessly integrates delta compression into inline deduplication for backup storage, addressing the aforementioned challenges by using three techniques: (1) dual-locality-based similarity tracking leverages both logical and physical locality to detect most of the similar chunks, which, due to their locality, can be prefetched by piggybacking on routine operations during deduplication, thereby eliminating extra I/Os during backup; (2) cache-aware filter identifies base chunks requiring extra I/Os during restore and prevents their referencing, thus eliminating extra restore I/Os; and (3) inversed delta compression, which reverses the roles of base and target chunks in the traditional delta compression approach, thereby allowing for the delta compression of data chunks that are otherwise prohibited as base chunks due to rewriting techniques. Experiments show that LoopDelta increases the compression ratio by 1.28 to 11.33 times over basic deduplication, without significantly affecting backup throughput, and enhances restore performance by up to 3.57 times.
Wenbin Zeng, Hong Jiang 0001, Dan Feng 0001, Zichen Xu 0001, Shuibing He, Mingzhe Zhang 0005, Dan Wu 0010
ACM Trans. Storage7
2024 Alchemist: A Unified Accelerator Architecture for Cross-Scheme Fully Homomorphic Encryption
abstract
The use of cross-scheme fully homomorphic encryption (FHE) in privacy-preserving applications present to be a new challenge to hardware accelerator design. Existing accelerator architectures with customized polynomial-level operator abstraction fail to efficiently handle hybrid FHE schemes due to the mismatch between computational demands and available hardware resources under various parameter settings. In this work, we propose a new accelerator architecture that consists of a novel finer-grained low-level operator, i.e., Meta-OP, that not only mathematically supports a diverse range of polynomial operations, but is also hardware-friendly for accelerator design without complex topological logic. We then design a new slot-based data management scheme to efficiently handle the distinct memory access patterns over the Meta-OP. With a slot-based data management approach, Alchemist can accelerate both arithmetic and logic FHE workloads with high hardware utilization rates. In the experiment, we show that Alchemist is up to 24,829X faster than CPU. For arithmetic FHE, compared with the SOTA ASIC accelerators, Alchemist achieves a 29.4X performance per area improvement on average. For logic FHE, compared with the SOTA ASIC accelerators, Alchemist achieves a 7.0X overall speed up on average.
Jianan Mu, Husheng Han, Shangyi Shi, Jing Ye 0001, Zizhen Liu, Shengwen Liang, Meng Li 0004, Mingzhe Zhang 0005, Song Bian 0001, Xing Hu 0001, Huawei Li 0001, Xiaowei Li 0001
DAC8
2024 Flagger: Cooperative Acceleration for Large-Scale Cross-Silo Federated Learning Aggregation
abstract
Cross-silo federated learning (FL) leverages homomorphic encryption (HE) to obscure the model updates from the clients. However, HE poses the challenges of complex cryptographic computations and inflated ciphertext sizes. As cross-silo FL scales to accommodate larger models and more clients, the overheads of HE can overwhelm a CPU-centric aggregator architecture, including excessive network traffic, enormous data volume, intricate computations, and redundant data movements. Tackling these issues, we propose Flagger, an efficient and high-performance FL aggregator. Flagger meticulously integrates the data processing unit (DPU) with computational storage drives (CSD), employing these two distinct near-data processing (NDP) accelerators as a holistic architecture to collaboratively enhance FL aggregation. With the delicate delegation of complex FL aggregation tasks, we build Flagger-DPU and Flagger-CSD to exploit both in-network and in-storage HE acceleration to streamline FL aggregation. We also implement Flagger-Runtime, a dedicated software layer, to coordinate NDP accelerators and enable direct peer-to-peer data exchanges, markedly reducing data migration burdens. Our evaluation results reveal that Flagger expedites the aggregation in FL training iterations by ${436\%}$ on average, compared with traditional CPU-centric aggregators.
Xiurui Pan, Yuda An, Shengwen Liang, Bo Mao 0003, Mingzhe Zhang 0005, Qiao Li 0001, Myoungsoo Jung, Jie Zhang 0048
ISCA5
2024 Trinity: A General Purpose FHE Accelerator
abstract
Fully Homomorphic Encryption (FHE) is crucial for privacy-preserving computing, which allows direct computation on encrypted data. While various FHE schemes have been proposed, none of them efficiently support both arithmetic FHE and logic FHE simultaneously. To address this issue, researchers explore the combination of different FHE schemes within a single application and propose algorithms for the conversion between them. Unfortunately, all prior ASIC-based FHE accelerators are designed to support a single FHE scheme, and none of them supports the acceleration for FHE scheme conversion. This necessitates FHE acceleration systems to integrate multiple accelerators for different schemes, leading to increased system complexity and hindering performance enhancement. In this paper, we present the first multi-modal FHE accelerator based on a unified architecture, which efficiently supports CKKS, TFHE, and their conversion scheme within a single accelerator. To achieve this goal, we first analyze the theoretical foundations of the aforementioned schemes and highlight their composition from a finite number of arithmetic kernels. Then, we investigate the challenges for efficiently supporting these kernels within a unified architecture, which include 1) concurrent support for NTT and FFT, 2) maintaining high hardware utilization across various polynomial lengths, and 3) ensuring consistent performance across diverse arithmetic kernels. To tackle these challenges, we propose a novel FHE accelerator named Trinity, which in-corporates algorithm optimizations, hardware component reuse, and dynamic workload scheduling to enhance the acceleration of CKKS, TFHE, and their conversion scheme. By adaptive select the proper allocation of components for NTT and MAC, Trinity maintains high utilization across NTTs with various polynomial lengths and imbalanced arithmetic workloads. The experiment results show that, for the pure CKKS and TFHE workloads, the performance of our Trinity outperforms the state-of-the- art accelerator for CKKS (SHARP) and TFHE (Morphling) by 1.49 x and 4.23 x, respectively. Moreover, Trinity achieves 919.3 x performance improvement for the FHE-conversion scheme over the CPU-based implementation. Notably, despite the performance improvement, the hardware overhead of Trinity is only 85 % of the summed circuit areas of SHARP and Morphling.
Xianglong Deng, Shengyu Fan, Zhicheng Hu, Zhuoyu Tian, Jiangrui Yu, Dingyuan Cao 0002, Dan Meng 0002, Rui Hou 0001, Meng Li 0004, Qian Lou, Mingzhe Zhang 0005
MICRO12
2024 Skyway: Accelerate Graph Applications with a Dual-Path Architecture and Fine-Grained Data Management
Mo Zou, Mingzhe Zhang 0005, Rujia Wang, Xian-He Sun, Xiaochun Ye, Dongrui Fan
J. Comput. Sci. Technol.2
2023 TensorFHE: Achieving Practical Computation on Encrypted Data Using GPGPU
abstract
In the cloud computing era, privacy protection is becoming pervasive in a broad range of applications (e.g., machine learning, data mining, etc). Fully Homomorphic Encryption (FHE) is considered the perfect solution as it enables privacy-preserved computation on untrusted servers. Unfortunately, the prohibitive performance overhead blocks the wide adoption of FHE (about 10, 000× slower than the normal computation). As heterogeneous architectures have gained remarkable success in several fields, achieving high performance for FHE with specifically designed accelerators seems to be a natural choice. Until now, most FHE accelerators have focused on efficiently implementing one FHE operation at a time based on ASIC and with significantly higher performance than GPU and FPGA. However, recent state-of-the-art FHE accelerators rely on an expensive and large on-chip storage and a high-end manufacturing process (i.e., 7nm), which increase the cost of FHE adoption.In this paper, we propose TensorFHE, an FHE acceleration solution based on GPGPU for real applications on encrypted data. TensorFHE utilizes Tensor Core Units (TCUs) to boost the computation of Number Theoretic Transform (NTT), which is the part of FHE with highest time-cost. Moreover, TensorFHE focuses on performing as many FHE operations as possible in a certain time period rather than reducing the latency of one operation. Based on such an idea, TensorFHE introduces operation-level batching to fully utilize the data parallelism in GPGPU. We experimentally prove that it is possible to achieve comparable performance with GPGPU as with state-of-the-art ASIC accelerators. TensorFHE performs 913 KOPS and 88 KOPS for NTT and HMULT (key FHE kernels) within NVIDIA A100 GPGPU, which is 2.61× faster than state-of-the-art FHE implementation on GPGPU; Moreover, TensorFHE provides comparable performance to the ASIC FHE accelerators, which makes it even 2.9× faster than the F1+ with a specific workload. Such a pure software acceleration based on commercial hardware with high performance can open up usage of state-of-the-art FHE algorithms for a broad set of applications in real systems.
Shengyu Fan, Weizhi Xu 0001, Rui Hou 0001, Dan Meng 0002, Mingzhe Zhang 0005
HPCA6
2023 Poseidon: Practical Homomorphic Encryption Accelerator
abstract
With the development of the important solution for privacy computing, the explosion of data size and computing intensity in Fully Homomorphic Encryption (FHE) has brought enormous challenges to the hardware design. In this paper, we propose a practical FHE accelerator - "Poseidon", which focuses on improving the hardware resource and bandwidth consumption. Poseidon supports complex FHE operations like Bootstrapping, Keyswitch, Rotation and so on, under limited FPGA resources. It refines these operations by abstracting five key operators: Modular Addition (MA), Modular Multiplication (MM), Number Theoretic Transformation (NTT), Automorphsim and Shared Barret Reduction (SBT). These operators are combined and reused to implement higher-level FHE operations. To utilize the FPGA resources more efficiently and improve the parallelism, we adopt the radix-based NTT algorithm and propose HFAuto, an optimized automorphism implementation suitable for FPGA. Then, we design the hardware accelerator based on the optimized key operators and HBM to maximize computational efficiency. We evaluate Poseidon with four domain-specific FHE benchmarks on Xilinx Alveo U280 FPGA. Empirical results show that the efficient reuse of the operator cores and on-chip storage enables superior performance compared with the state-of-the-art GPU, FPGA and accelerator ASICs. We highlight the following results: (1) up to 370× speedup over CPU for the basic operations of FHE; (2) up to 1300×/52× speedup over CPU and the FPGA solution for the key operators; (3) up to 10.6×/8.7× speedup over GPU and the ASIC solution for the FHE benchmark.
Yinghao Yang 0001, Huaizhi Zhang, Shengyu Fan, Mingzhe Zhang 0005, Xiaowei Li 0001
HPCA5
2022 VNet: a versatile network to train real-time semantic segmentation models on a single GPU
Wenxing Li, Ning Lin, Mingzhe Zhang 0005, Xiaoming Chen 0003, Xiaowei Li 0001
Sci. China Inf. Sci.3
2021 Streamline Ring ORAM Accesses through Spatial and Temporal Optimization
abstract
Memory access patterns could leak temporal and spatial information in a sensitive program; therefore, obfuscated memory access patterns are desired from the security perspective. Oblivious RAM (ORAM) has been the favored candidate to eliminate the access pattern leakage through randomly remapping data blocks around the physical memory space. Meanwhile, accessing memory with ORAM protocols results in significant memory bandwidth overhead. For each memory request, after going through the ORAM obfuscation, the main memory needs to service tens of actual memory accesses, and only one real access out of them is useful for the program execution. Besides, to ensure the memory bus access patterns are indistinguishable, extra dummy blocks need to be stored and transmitted, which cause memory space waste and poor performance. In this work, we introduce a new framework, String ORAM, that accelerates the Ring ORAM accesses with Spatial and Temporal optimization schemes. First, we identify that dummy blocks could significantly waste memory space and propose a compact ORAM organization that leverages the real blocks in memory to obfuscate the memory access pattern. Then, we identify the inefficiency of current transaction-based Ring ORAM scheduling on DRAM devices and propose an effective scheduling technique that can overlap the time spent on row buffer misses while ensuring correctness and security. With a minimal modification on the hardware and software, and negligible impact on security, the framework reduces 30.05% execution time and up to 40% memory space overhead compared to the state-of-the-art bandwidth-efficient Ring ORAM.
Dingyuan Cao 0002, Mingzhe Zhang 0005, Xiaochun Ye, Dongrui Fan, Yuezhi Che, Rujia Wang
HPCA2
2021 BitX: Empower Versatile Inference with Hardware Runtime Pruning
abstract
Classic DNN pruning mostly leverages software-based methodologies to tackle the accuracy/speed tradeoff, which involves complicated procedures like critical parameter searching, fine-tuning and sparse training to find the best plan. In this paper, we explore the opportunities of hardware runtime pruning and propose a hardware runtime pruning methodology, termed as “BitX” to empower versatile DNN inference. It targets the abundant useless bits in the parameters, pinpoints and prunes these bits on-the-fly in the proposed BitX accelerator. The versatility of BitX lies in: (1) software effortless; (2) orthogonal to the software-based pruning; and (3) multi-precision support (including both floating point and fixed point). Empirical studies on image classification and object detection models highlight the following results: (1) up to 4.82x speedup over the original non-pruned DNN and 14.76x speedup collaborated with the software-pruned DNN; (2) up to 0.07% and 0.9% higher accuracy for the floating-point and fixed-point DNN, respectively; (3) 2.00x and 3.79x performance improvement over the state-of-the-art accelerators, with 0.039 mm2 and 68.62 mW (floating-point 32), 36.41 mW(16-bit fixed point) power consumption under TSMC 28 nm technology library.
Mingzhe Zhang 0005, Liang Chang 0002, Xiaowei Li 0001
ICPP5
2021 CoPIM: A Concurrency-aware PIM Workload Offloading Architecture for Graph Applications
abstract
Processing-in-Memory (PIM) is considered a promising solution to improve the performance of graph-computing applications by minimizing the data movement between the host and memory. Which workload to offload and how to offload it to PIM logic determine whether the PIM architecture is well utilized. Offloading too much or too little workload from the host processor to the PIM side could hurt overall performance. On the other hand, the offloading granularity needs to be representative without losing generality. In this paper, we present CoPIM, a novel PIM workload offloading architecture that can dynamically determine which portion of the graph workload can benefit more from PIM-side computation. CoPIM focuses on the loop code blocks of graph applications and evaluates the necessity of offloading based on a concurrent memory access model. We also provide detailed architectural designs to support the offloading. In this way, CoPIM reduces the size of offloading instructions and also improves the overall performance with less energy consumption. The experimental results show that compared with other state-of-the-art PIM workload offloading frameworks, CoPIM achieves a speedup by the geometric mean of 19.5% and 11.4% than PEI and GraphPIM, respectively. On the other hand, CoPIM also reduces the un-core energy consumption by 6.8% and 6.5% on average over PEI and GraphPIM, respectively.
Mingzhe Zhang 0005, Rujia Wang, Xiaoming Chen 0003, Xingqi Zou, Xiaoyang Lu, Yinhe Han 0001, Xian-He Sun
ISLPED2
2021 Distilling Bit-level Sparsity Parallelism for General Purpose Deep Learning Acceleration
abstract
Along with the rapid evolution of deep neural networks, the ever-increasing complexity imposes formidable computation intensity to the hardware accelerator. In this paper, we propose a novel computing philosophy called “bit interleaving” and the associate accelerator design called “Bitlet” to maximally exploit the bit-level sparsity. Apart from existing bit-serial/parallel accelerators, Bitlet leverages the abundant “sparsity parallelism” in the parameters to enforce the inference acceleration. Bitlet is versatile by supporting diverse precisions on a single platform, including floating-point 32 and fixed-point from 1b to 24b. The versatility enables Bitlet feasible for both efficient inference and training. Empirical studies on 12 domain-specific deep learning applications highlight the following results: (1) up to 81 × /21 × energy efficiency improvement for training/inference over recent high performance GPUs; (2) up to 15 × /8 × higher speedup/efficiency over state-of-the-art fixed-point accelerators; (3) 1.5mm2 area and scalable power consumption from 570mW (float32) to 432mW (16b) and 365mW (8b) @28nm TSMC; (4) highly configurable justified by ablation and sensitivity studies.
Liang Chang 0002, Zixuan Zhu 0001, Shengjian Lu, Yanhuan Liu, Mingzhe Zhang 0005
MICRO7
2020 Architecting Effectual Computation for Machine Learning Accelerators
abstract
Inference efficiency is the predominant design consideration for modern machine learning accelerators. The ability of executing multiply-and-accumulate (MAC) significantly impacts the throughput and energy consumption during inference. However, MAC operation suffers from significant ineffectual computations that severely undermines the inference efficiency and must be appropriately handled by the accelerator. The ineffectual computations are manifested in two ways: first, zero values as the input operands of the multiplier, waste time and energy but contribute nothing to the model inference; second, zero bits in nonzero values occupy a large portion of multiplication time but are useless to the final result. In this article, we propose an ineffectual-free yet cost-effective computing architecture, called split-and-accumulate (SAC) with two essential bit detection mechanisms to address these intractable problems in tandem. It replaces the conventional MAC operation in the accelerator by only manipulating the essential bits in the parameters (weights) to accomplish the partial sum computation. Besides, it also eliminates multiplications without any accuracy loss, and supports a wide range of precision configurations. Based on SAC, we propose an accelerator family called Tetris and demonstrate its application in accelerating state-of-the-art deep learning models. Tetris includes two implementations designed for either high performance (i.e., cloud applications) or low power consumption (i.e., edge devices), respectively, contingent to its built-in essential bit detection mechanism. We evaluate our design with Vivado HLS platform and achieve up to 6.96× performance enhancement, and up to 55.1× energy efficiency improvement over conventional accelerator designs.
Mingzhe Zhang 0005, Yinhe Han 0001, Qi Wang 0025, Huawei Li 0001, Xiaowei Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2019 FindeR: Accelerating FM-Index-Based Exact Pattern Matching in Genomic Sequences through ReRAM Technology
abstract
Genomics is the critical key to enabling precision medicine, ensuring global food security and enforcing wildlife conservation. The massive genomic data produced by various genome sequencing technologies presents a significant challenge for genome analysis. Because of errors from sequencing machines and genetic variations, approximate pattern matching (APM) is a must for practical genome analysis. Recent work proposes FPGA, ASIC and even process-in-memory-based accelerators to boost the APM throughput by accelerating dynamic-programming-based algorithms (e.g., Smith-Waterman). However, existing accelerators lack the efficient hardware acceleration for the exact pattern matching (EPM) that is an even more critical and essential function widely used in almost every step of genome analysis including assembly, alignment, annotation and compression. State-of-the-art genome analysis adopts the FM-Index that augments the space-efficient BWT with additional data structures permitting fast EPM operations. But the FM-Index is notorious for poor spatial locality and massive random memory accesses. In this paper, we propose a ReRAM-based process-in-memory architecture, FindeR, to enhance the FM-Index EPM search throughput in genomic sequences. We build a reliable and energy-efficient Hamming distance unit to accelerate the computing kernel of FM-Index search using commodity ReRAM chips without introducing extra CMOS logic. We further architect a full-fledged FM-Index search pipeline and improve its search throughput by lightweight scheduling on the NVDIMM. We also create a system library for programmers to invoke FindeR to perform EPMs in genome analysis. Compared to state-of-the-art accelerators, FindeR improves the FM-Index search throughput by 83% ~ 30K× and throughput per Watt by 3.5×~42.5K×.
Farzaneh Zokaee, Mingzhe Zhang 0005, Lei Jiang 0001
PACT2
2019 Magma: A Monolithic 3D Vertical Heterogeneous ReRAM-based Main Memory Architecture
abstract
3D vertical ReRAM (3DV-ReRAM) emerges as one of the most promising alternatives to DRAM due to its good scalability beyond 10nm. Monolithic 3D (M3D) integration enables 3DV-ReRAM to improve its array area efficiency by stacking peripheral circuits underneath an array. A 3DV-ReRAM array has to be large enough to fully cover the peripheral circuits, but such large array size significantly increases its access latency. In this paper, we propose Magma, a M3D stacked heterogeneous ReRAM array architecture, for future main memory systems by stacking a large unipolar 3DV-ReRAM array on the top of a small bipolar 3DV-ReRAM array and peripheral circuits shared by two arrays. We further architect the small bipolar array as a direct-mapped cache for the main memory system. Compared to homogeneous ReRAMs, on average, Magma improves the system performance by 11.4%, reduces the system energy by 24.3% and obtains > 5-year lifetime.
Farzaneh Zokaee, Mingzhe Zhang 0005, Xiaochun Ye, Dongrui Fan, Lei Jiang 0001
DAC2
2019 When Deep Learning Meets the Edge: Auto-Masking Deep Neural Networks for Efficient Machine Learning on Edge Devices
abstract
Deep neural network (DNN) has demonstrated promising performance in various machine learning tasks. Due to the privacy issue and the unpredictable transmission latency, inferring DNN models directly on edge devices trends the development of intelligent systems, like self-driving cars, smart Internet-of-Things (IoTs) and autonomous robotics. The on-device DNN model is obtained by expensive training via vast volumes of high-quality training data in the cloud datacenter, and then deployed into these devices, expecting it to work effectively at the edge. However, edge device always deals with low-quality images caused by compression or environmental noise pollutions. The well-trained model, though could work perfectly on the cloud, cannot adapt to these edge-specific conditions without remarkable accuracy drop. In this paper, we propose an automated strategy, called "AutoMask", to embrace effective machine learning and accelerate DNN inference on edge devices. AutoMask comprises end-to-end trainable software strategies and cost-effective hardware accelerator architecture to improve the adaptability of the device without compromising the constrained computation and storage resources. Extensive experiments, over ImageNet dataset and various state-of-the-art DNNs, show that AutoMask achieves significant inference acceleration and storage reduction while maintains comparable accuracy level on embedded Xilinx Z7020 FPGA, as well as NVIDIA Jetson TX2.
Ning Lin, Xing Hu 0001, Jingliang Gao, Mingzhe Zhang 0005, Xiaowei Li 0001
ICCD5
2019 Balancing Performance and Energy Efficiency of ONoC by Using Adaptive Bandwidth
abstract
On-chip optical links can consume considerable energy. Previous work shows that there exists a trade-off between the optical link bandwidth and energy consumption. In this paper, we analyze the temporal behavior of on-chip communication in typical applications and make the following observation: both ONoC and processor cores are not always busy, and there are significant amounts of time during which these components are relatively idle. Base on such an observation, we present a novel technique called Shifting Link to dynamically changing the bandwidth for optical link according to the workload situation. To evaluate our proposed method, we implement the Shifting Link in FlexiShare ONoC architecture. The experiment result shows that, on geometric mean, Shifting Link reduces the energy consumption by 35.0% with only 5.8% decrease in the performance.
Mingzhe Zhang 0005, Lunkai Zhang, Fred Chong, Zhiyong Liu 0002
ICCD1
2019 Quick-and-Dirty: An Architecture for High-Performance Temporary Short Writes in MLC PCM
abstract
MLC PCM provides high-density data storage and extended data retention; therefore it is a promising alternative for DRAM main memory. However, its low write performance is a major obstacle to commercialization. One opportunity for improving the latency of MLC PCM writes is to use fewer SET iterations in a single write. Unfortunately, this comes with a cost: the data written by these short writes have remarkably shorter retentions and thus need frequent refreshes. As a result, it is impractical to use these short-latency, short-retention writes globally. In this paper, we analyze the temporal behavior of write operations in typical applications and show that the write operations are bursty in nature, that is, during some time intervals the memory is subject to a large number of writes, while during other time intervals there hardly any memory operations take place. Based on this observation, we propose Quick-and-Dirty (QnD), a lightweight scheme to improve the performance of MLC PCM. When the write performance becomes the system bottleneck, QnD performs some write operations using the short-latency, short-retention write mode. Then, when the memory system is relatively quiet, QnD uses idle-memory intervals to refresh the data written by short-latency, short-retention writes in order to mitigate the short retention problem. Our experimental results show that QnD improves performance by 30.9 percent on geometric mean while still providing acceptable memory lifetime (7.58 years on geometric mean). We also provide sensitivity studies of the aggressiveness, memory coverage and granularity of QnD technique.
Mingzhe Zhang 0005, Lunkai Zhang, Lei Jiang 0001, Fred Chong, Zhiyong Liu 0002
IEEE Trans. Computers1
2018 Mmalloc: A Dynamic Memory Management on Many-core Coprocessor for the Acceleration of Storage-intensive Bioinformatics Application
Mingzhe Zhang 0005, Jingrong Zhang, Rui Yan 0009, Zhiyong Liu 0002, Fa Zhang 0001, Xuefeng Cui
BIBM2
2017 Balancing Performance and Lifetime of MLC PCM by Using a Region Retention Monitor
abstract
Multi Level Cell (MLC) Phase Change Memory (PCM) is an enhancement of PCM technology, which provides higher capacity by allowing multiple digital bits to be stored in a single PCM cell. However, the retention time of MLC PCM is limited by the resistance drift problem and refresh operations are required. Previous work shows that there exists a trade-off between write latency and retention-a write scheme with more SET iterations and smaller current provides a longer retention time but at the cost of a longer write latency. Otherwise, a write scheme with fewer SET iterations achieves high performance for writes but requires a greater number of refresh operations due to its significantly reduced retention time, and this hurts the lifetime of MLC PCM. In this paper, we show that only a small part of memory (i.e., hot memory regions) will be frequently accessed in a given period of time. Based on such an observation, we propose Region Retention Monitor (RRM), a novel structure that records and predicts the write frequency of memory regions. For every incoming memory write operation, RRM select a proper write latency for it. Our evaluations show that RRM helps the system improves the balance between system performance and memory lifetime. On the performance side, the system with RRM bridges 77.2% of the performance gap between systems with long writes and systems with short writes. On the lifetime side, a system with RRM achieves a lifetime of 6.4 years, while systems using only long writes and short writes achieve lifetimes of 10.6 and 0.3 years, respectively. Also, we can easily control the aggressiveness of RRM through an attribute called hot threshold. A more aggressively configured RRM can achieve the performance which is only 3.5% inferior than the system using static short writes, while still achieve a lifetime of 5.78 years.
Mingzhe Zhang 0005, Lunkai Zhang, Lei Jiang 0001, Zhiyong Liu 0002, Fred Chong
HPCA1
2017 Quick-and-Dirty: Improving Performance of MLC PCM by Using Temporary Short Writes
abstract
Low write performance is a major obstacle to the commercialization of MLC PCM. One opportunity for improving the latency of MLC PCM writes is to use fewer SET iterations in a single write. Unfortunately, the data written by these short writes have significantly shorter retention time and thus need frequent refreshes. As a result, it is impractical to use these short-latency, short-retention writes globally. In this paper, we analyze the temporal behavior of write operations in typical applications and propose Quick-and-Dirty (QnD), a lightweight scheme to improve the performance of MLC PCM. QnD dynamically performs the short-latency, short-retention write when write operations are bursty, and then uses short-latency, short-retention writes to mitigate the short retention problem when memory system is relatively quiet. Our experimental results show that QnD improves performance by 30.9% on geometric mean while still providing acceptable memory lifetime (7.58 years on geometric mean). We also provide sensitivity studies of the aggressiveness, memory coverage and granularity of QnD technique.
Mingzhe Zhang 0005, Lunkai Zhang, Lei Jiang 0001, Fred Chong, Zhiyong Liu 0002
ICCD1
2015 FreeRider: Non-Local Adaptive Network-on-Chip Routing with Packet-Carried Propagation of Congestion Information
abstract
Non-local adaptive routing techniques, which utilize statuses of both local and distant links to make routing decisions, have recently been shown to be effective solutions for promoting the performance of Network-on-Chip (NoC). The essence of non-local adaptive routing was an additional network dedicated to propagate congestion information of distant links on the NoC. While the dedicated Congestion Propagation Network (CPN) helps routers to make promising routing decisions, it incurs additional wiring and power costs and becomes an unnecessary decoration when the load of NoC is light. Moreover, the CPN has to be extended if one would utilize more sophisticated congestion information to enhance the performance of NoC, bringing in even larger wiring and power costs. This paper proposes an innovative non-local adaptive routing technique called FreeRider, which does not use a dedicated CPN but instead leverages free bits in head flits of existing packets to carry and propagate rich congestion information without introducing additional wires or flits. In order to balance the network load, FreeRider adopts a novel three-stage strategy of output link selection, which adequately utilizes the propagated information to make routing decisions. Experimental results on both synthetic traffic patterns and application traces show that FreeRider achieves better throughput, shorter latency, and smaller power consumption than a state-of-the-art adaptive routing technique with dedicated CPN.
Shaoli Liu, Tianshi Chen 0002, Ling Li 0001, Xi Li 0003, Mingzhe Zhang 0005, Chao Wang 0003, Haibo Meng, Xuehai Zhou, Yunji Chen
IEEE Trans. Parallel Distributed Syst.5
2014 SpongeDirectory: flexible sparse directories utilizing multi-level memristors
abstract
Cache-coherent shared memory is critical for programmability in many-core systems. Several directory-based schemes have been proposed, but dynamic, non-uniform sharing make efficient directory storage challenging, with each giving up storage space, performance or energy.
Lunkai Zhang, Dmitri B. Strukov, Heba Saadeldeen, Dongrui Fan, Mingzhe Zhang 0005, Diana Franklin
PACT5
2013 SimICT: A fast and flexible framework for performance and power evaluation of large-scale architecture
abstract
Simulation is an important method to evaluate future computer systems. However, the increasing complexity of the target systems has made the development of simulators very difficult. Furthermore, detailed simulation of large-scale parallel architecture is so slow that full evaluation of real application becomes a great challenge. This paper presents SimICT, a fast and flexible simulation framework which aims at performance and power evaluation for large-scale architecture. SimICT uses component-based design to improve its flexibility of building target systems. It also introduces an automatic parallel mechanism with relaxed synchronization to speed up the simulation. Finally, it provides a graphic configuration interface to ease the use difficulty. Based on this framework, various existing models, such as performance and power modeling tools, can be integrated to produce a holistic simulation platform.
Xiaochun Ye, Dongrui Fan, Ninghui Sun, Shibin Tang, Mingzhe Zhang 0005, Hao Zhang 0009
ISLPED5