EDBT 2026 Demo / reviewers in the wild / expert
Haozhe Zhu
dblp:208/4162
· DBLP profile ↗
24ranked-venue papers
0as first author
24since 2021 · last 2026
0000-0002-6412-3996ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 21 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GauTracer: Extending Ray Tracing Accelerator for Gaussian-Based Scene Representation
Lizhou Wu, Kunchen Zou, Yuzheng Lin, Chixiao Chen, Xiaoyang Zeng, Haozhe Zhu |
ISCA | 6 |
| 2026 | CIM-Pruner: A Dual-Mode Compute-In-Memory Macro for Efficient VLMs with Intra-Chunk Token Pruning and Merging
Zhuojun Han, Siqi He, Chixiao Chen, Haozhe Zhu |
ISCAS | 4 |
| 2026 | FIT-Print: Toward False-Claim-Resistant Model Ownership Verification via Targeted FingerprintabstractModel fingerprinting has emerged as a crucial mechanism for safeguarding the intellectual property of open-source models, offering a non-intrusive approach that requires no modifications to the protected model. However, our analysis reveals that existing fingerprinting techniques are fundamentally vulnerable to false claim attacks, wherein adversaries can fraudulently assert ownership over independent third-party models. We demonstrate that this vulnerability stems from the untargeted nature of current methods, which evaluate model similarity based on arbitrary sample outputs rather than alignment with a specific, predefined reference. To mitigate this vulnerability, we introduce FIT-Print, a targeted fingerprinting paradigm that actively counters false claim attacks. Specifically, FIT-Print leverages optimization to transform the fingerprint into a verifiable, targeted signature. Building upon this foundation, we propose two black-box fingerprinting methods, the bit-wise FIT-ModelDiff and the list-wise FIT-LIME, which utilize output distances and feature attributions as robust model signatures, respectively. Extensive evaluations across benchmark models and datasets show that our framework perfectly neutralizes false claim attacks (100% defense success rate) and eliminates false alarms on independent models (0.0%), all while maintaining a 100% ownership verification rate against diverse model reuse techniques. Shuo Shao 0002, Haozhe Zhu, Yiming Li 0004, Hongwei Yao, Tianwei Zhang 0004, Zhan Qin |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2026 | COPIC: A Codesigned Accelerator for Efficient Octree-Based Learned Point Cloud Compression at EdgeabstractWith the growing adoption of light detection and ranging (LiDAR) technology and the increasing resolution of captured data, efficient point cloud compression (PCC) has become a critical challenge. While PCC is essential for reducing bandwidth and storage demands, existing octree-based learned PCC methods face two fundamental bottlenecks: 1) octree construction and serialization that involve irregular, pointwise memory accesses that are difficult to parallelize and 2) autoregressive entropy model inference that requires heavy computation. These issues lead to high latency and energy consumption, making real-time deployment on edge devices impractical. This article introduces COPIC, a hardware accelerator designed for real-time PCC on edge devices. COPIC integrates two key components. The first is OctCAM, a content-addressable-memory (CAM)-based module for rapid octree construction and serialization. By enabling hardware-level parallel search, OctCAM significantly reduces the latency and energy consumption caused by irregular memory access patterns in octree processing. The second is PCC-Engine, a dedicated neural network inference module. Combined with a lightweight PCC-TinyAttention network and a window-based key-value caching approach, it further minimizes execution latency and power usage. Experimental results show that COPIC achieves a$25.9\times $speedup compared to GPU implementations, with a total processing latency of 40.1 ms and an average energy consumption of 0.23 J per frame, while delivering a compression rate 12.4% higher than standard proposals. Jiapei Zheng, Yutong Su, Siqi He, Haozhe Zhu, Qi Liu 0010, Chixiao Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | Hydra: Harnessing Expert Popularity for Efficient Mixture-of-Expert Inference on Chiplet SystemabstractThe rapid growth of model sizes in advanced artificial intelligence algorithms, particularly in Transformerbased large language models (LLMs), has led to significant computational overhead. Mixture-of-Expert (MoE) models offer a solution through their sparsely gating mechanism but introduce new challenges of extensive all-to-all communication and model computational inefficiencies. This paper presents Hydra, a software/hardware co-design aimed at accelerating MoE inference on chiplet-based architectures. In software, Hydra employs a popularity-aware expert mapping strategy to optimize interchiplet communication. In hardware, it incorporates Content Addressable Memory (CAM) to eliminate expensive explicit token (un)-permutation based on sparse matrix multiplications and a redundant-calculation-skipping softmax engine to bypass unnecessary division and exponential operations. Evaluated in 22 nm technology, Hydra achieves latency reductions of $14.2 \times$ and $3.5 \times$ and power reductions of $169.1 \times$ and $18.9 \times$ over GPU and state-of-the-art MoE accelerator, respectively, thereby offering a scalable and efficient solution for MoE model deployment. Siqi He, Haozhe Zhu, Jiapei Zheng, Lizhou Wu, Bo Jiao 0003, Qi Liu 0010, Xiaoyang Zeng, Chixiao Chen |
DAC | 2 |
| 2025 | DRAFT: Decoupling Backpropagation from Pre-trained Backbone for Efficient Transformer Fine-Tuning on EdgeabstractTransformers have demonstrated outstanding performance across diverse applications recently, necessitating finetuning to optimize their performance for downstream tasks. However, fine-tuning remains challenging due to the substantial computational costs and storage overhead of backpropagation (BP). The existing fine-tuning techniques require the BP through the massive pre-trained backbone weights for computing the input gradient, resulting in significant computing overhead and memory footprint for resource-constrained edge devices. To address the challenge, this work proposes an algorithm-hardware co-design framework, DRAFT, for efficient Transformer finetuning by decoupling the BP from the backbone weights, thereby efficiently reducing the BP overhead. The framework employs Feedback Decoupling Approximation (FDA), an efficient finetuning algorithm that decouples BP into two low-complexity pathways: trainable adapter pathway and sparse ternary Bypass Network (BPN) pathway. The two pathways work collaboratively to approximate the conventional BP process. Further, a DRAFT accelerator is proposed, featuring a reconfigurable design with lightweight sparse gather networks and dynamic workflows to fully harness the sparsity and data parallelism inherent to the FDA. Experimental results demonstrate that DRAFT achieves a speedup of $4.9 \times$ and an energy efficiency improvement of $4.2 \times$ on average compared to baseline fine-tuning methods across multiple fine-tuning tasks with negligible accuracy loss. Zhirui Huang, Shiwei Liu 0002, Haozhe Zhu, Qi Liu 0010, Chixiao Chen |
DAC | 3 |
| 2025 | PIMoE: Towards Efficient MoE Transformer Deployment on NPU-PIM System through Throttle-Aware Task OffloadingabstractMixture-of-experts (MoE) technique holds significant promise for scaling up Transformer models. However, the data transfer overhead and imbalanced workload hinder efficient deployment. This work presents PIMoE, a heterogeneous system combining processing-in-memory (PIM) and neural-processing-unit (NPU) to facilitate efficient MoE Transformer inference. We propose a throttle-aware task offloading method that addresses workload imbalance between NPU and PIM, achieving optimal task distribution. Furthermore, we design a near-memory-controller data condenser to address the mismatch of sparse data layout between NPU and PIM, enhancing data transfer efficiency. Experimental results demonstrate that PIMoE achieves $4.5 \times$ speedup and $13.7 \times$ greater energy efficiency compared to the A 100, and $1.4 \times$ speedup over a state-of-the-art MoE platform. Lizhou Wu, Haozhe Zhu, Siqi He, Xuanda Lin, Xiaoyang Zeng, Chixiao Chen |
DAC | 2 |
| 2025 | EIGEN: Enabling Efficient 3DIC Interconnect with Heterogeneous Dual-Layer Network-on-Active-InterposerabstractChiplet-based 3DICs have emerged as modular solutions for large-scale, high-performance computing systems. However, unlike monolithic Network-on-Chips (NoCs), 3DICs using active interposers encounter difficulties in managing heterogeneous and high-load network traffic patterns. The emerging field of Networks-on-Active-Interposer (NOAI), independently designed from top dies’ interconnects, is desired to address these challenges by supporting flexible topologies, low-latency memory access, and reduced traffic congestion. To satisfy the requirements, we propose a heterogeneous dual-layer interconnect architecture, EIGEN, for chiplet-interposer systems, along with a reinforcement learning (RL)-based routing framework, which can provide efficient and flexible communication for chiplet-based 3DICs. EIGEN features an application-aware switch-programable interconnection layer (AspLayer) and a dynamic packet-routing interconnection layer (DynLayer) on the interposer, respectively. To optimize inter-chiplet data communication, we also develop an RL-based path routing framework tailored to this dual-layer architecture. The RL framework’s state space includes topology, network, and memory metrics, while the RL reward is set as the product of average packet latency and link utilization. The effectiveness of EIGEN is demonstrated through different applications where CPU, GPU, and AI accelerator chiplets are integrated via NOAI. Simulation results show that the proposed EIGEN achieves up to $\mathbf{6 7. 1 6 \%}$ latency reduction, $\mathbf{5 3. 8 9 \%}$ hops reduction and $\mathbf{1 1. 2 1 \%}$ runtime reduction compared to the state-of-the-art (SOTA) chiplets interconnect architectures. Furthermore, sensitivity analysis of network scaling shows that the EIGEN architecture and framework exhibit strong scalability, with latency reduction of $\mathbf{2 0. 9 6 \%}$ to $\mathbf{5 2. 7 7 \%}$ and hop reduction of $\mathbf{1 9. 7 2 \%}$ to $\mathbf{4 3. 4 1 \%}$ as the system scales from $4 \times 4$ to $32 \times 32$ chiplets. Siyao Jia, Bo Jiao 0003, Haozhe Zhu, Chixiao Chen, Qi Liu 0010, Ming Liu 0022 |
HPCA | 3 |
| 2025 | GauPRE: A Pattern-based Rendering Engine for Gaussian Splatting on Edge Deviceabstract3D Gaussian Splatting (3DGS)–based rendering has gained increasing attention due to its state-of-the-art quality and broad applications in Augmented and Virtual Reality (AR/VR). However, the deployment of 3DGS in edge systems (~10FPS) faces challenge in achieving real-time (≥30FPS) performance due to computational resource constraints. Profiling reveals that the Gauss-Tile rasterization critically impacts pipeline efficiency due to its decoupled nature and computation-accuracy trade-off, which current hardware devices cannot effectively resolve. To this end, we propose a software-hardware co-design that performs rasterization via pattern matching. Our solution adopts flood encoding to represent tile coverage efficiently. It integrates a pattern-aware rasterization unit (PRU) and compresses the codebook using k-means clustering, preserving accuracy with minimal area overhead. At the architectural level, GauPRE adopts early depth-sorting and tile-group-wise rasterization to fuse coverage testing and alpha blending, enabling seamless data flow. We further integrated GauPRE with the GPU to support the end-to-end 3DGS rendering pipeline. Results demonstrate that the GPU+GauPRE delivers an 11.6× end-to-end speedup over the Jetson Orin Nano with only 0.1% area overhead. Meanwhile, the standalone GauPRE rasterization engine achieves 2.09× higher throughput compared to a state-of-the-art 3DGS accelerator. Yuzheng Lin, Lizhou Wu, Chixiao Chen, Xiaoyang Zeng, Haozhe Zhu |
ICCAD | 5 |
| 2025 | A 0.22 pJ/bit Processing-in-Controller GEMV Macro with Weight Prefetch for Efficient Near-Memory Computingabstract3D-stacked DRAM is a key technology enabling the development of large language models (LLMs). However, the intensive computational demands lead to substantial energy consumption resulting from large-scale data movement. Integrating processing-in-memory (PIM) within DRAM has been employed to reduce data movement, but it elevates manufacturing costs and incurs additional area and power consumption. On the other hand, implementing processing-near-memory (PNM) outside the DRAM controller fails to eliminate the high energy consumption associated with interconnects during data transfer. To address these challenges, we propose a processing-in-controller (PIC) architecture aimed at 3D-stacked DRAM for efficient data movement. A size-scalable General Matrix-Vector Multiplication (GEMV) structure is proposed, supporting configurations ranging from 16 × 16 to 128 × 128, thereby enhancing its adaptability for a variety of applications. To address the low throughput bottleneck caused by DDR read latency, a ping-pong buffer with weight prefetching is introduced in the PNM. Based on 28 nm CMOS technology, experimental results demonstrate that the energy consumption of this PIC is 0.22 pJ/bit, while data movement efficiency improves by nearly 92% compared to traditional DRAM operations. Jiapei Zheng, Siqi He, Lizhou Wu, Chen Mu, Haozhe Zhu, Liyu Lin, Qi Liu 0010, Chixiao Chen |
ISCAS | 6 |
| 2025 | Federated-Continual Dynamic Segmentation of Histopathology Guided by Barlow Continuity
Niklas Babendererde, Haozhe Zhu, Moritz Fuchs, Jonathan Stieber, Anirban Mukhopadhyay 0003 |
WACV | 2 |
| 2025 | SSDM: Generated image interaction method based on spatial sparsity for diffusion models
Zhuochao Yang, Jingjing Liu 0004, Haozhe Zhu, Wanquan Liu |
Neurocomputing | 3 |
| 2025 | RT-FLOW: FPGA Implementation of Real-Time Optical-Flow-Based SLAM for High-Speed Tracking and High-Quality MappingabstractSimultaneous Localization and Mapping (SLAM) is pivotal for autonomous robotics, yet feature-based SLAM systems struggle with sparse environmental representations and robustness under dynamic conditions. Optical-flow-based SLAM (OpF-SLAM) addresses these limitations by leveraging pixel-level motion data for dense mapping; however, its computational intensity hinders real-time deployment. This paper presents RT-FLOW, an FPGA-based accelerator for OpF-SLAM that achieves real-time performance through three key innovations: 1) A feature-context encoding engine that exploits inter-frame similarity to resolve data dependency in correlation construction, reducing latency by 77.5%. 2) A heterogeneous mixed-precision flow update engine guided by correlation sparsity, enabling 3.7× faster optical flow computation with negligible accuracy loss. 3) A pivoting-free linear solver using Householder transformations for stable pose optimization. Implemented on Xilinx XCZU7EV FPGA, RT-FLOW processes full-image pixels per frame at 65 fps with an energy efficiency of 0.358 μJ/point, outperforming previous FPGA designs. Evaluated on benchmark datasets, RT-FLOW demonstrates robustness in diverse environments while maintaining sub-110mJ/frame energy consumption. This work bridges the gap between algorithmic potential and hardware feasibility for high-density SLAM, empowering next-generation mobile robots with real-time scene understanding capabilities. Siqi He, Qi Liu 0010, Xiaoyang Zeng, Chixiao Chen, Haozhe Zhu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2024 | ARCTIC: Agile and Robust Compute-In-Memory Compiler with Parameterized INT/FP Precision and Built-In Self TestabstractDigital Compute-in-Memory (DCIM) architectures are playing an increasingly vital role in artificial intelligence (AI) applications due to their significant energy efficiency enhancement. Coupling memory and computing logic in DCIM requires extensive customization of custom cells and layouts, thus increasing design complexity and implementation effort. To adapt to the swiftly evolving AI algorithms, DCIM compiler for agile customization is required. Previous DCIM compilers accelerate the customization process but only focus on integer computation. Moreover, with technology node scaling down, design-for-test circuits are critical for robust chip design, while previous built-in-self-test (BIST) schemes for traditional memory fail to offer support for DCIM. This paper presents ARCTIC, an agile and robust DCIM compiler supporting parameterized integer/floating-point formats with corresponding BIST circuits. To support variable precision formats (including integer and floating-point), ARCTIC applies adaptive topology and layout optimization schemes for optimal performance. The compiler is also equipped with DCIM-friendly MarchCIM BIST circuits for efficient post-silicon tests with negligible area overhead. The energy efficiency of the generated DCIM macros remains competent with the state-of-the-art counterparts. Haozhe Zhu, Siqi He, Chengchen Wang, Xiankui Xiong, Haidong Tian, Xiaoyang Zeng, Chixiao Chen |
DATE | 2 |
| 2024 | A 19.7 TFLOPS/W Multiply-less Logarithmic Floating-Point CIM Architecture with Error-Reduced Compensated Approximate AdderabstractThe growing demand for high-precision neural network training and inference has driven the necessity for floating-point (FP) compute-in-memory (CIM) architectures. However, compared to the extensively studied INT-CIM, the energy efficiency of FP-CIM still requires further optimization and enhancement. This work presents an energy-efficient multiply-less digital SRAM-based FP-CIM architecture. Specifically, to improve the energy efficiency and minimize the area requirement, we propose to employ logarithmic approximate FP multiplication (LAM) within the FP-CIM architecture. The LAM approximates FP multiplication by converting it into a straightforward addition operation, thereby reducing the power consumption and area. Additionally, we propose an approximate adder with error-reduced compensation to address critical path delay issues associated with carry propagation, further minimizing power consumption and area overhead. A 24Kb SRAM CIM macro with the proposed techniques is designed in a 28nm CMOS technology and occupies an area of 0.033 mm2. The simulation results show that our work achieves an energy efficiency of 19.7 TFLOPS/W with bfloat16 representation at 0.9V and 200MHz. Siqi He, Haozhe Zhu, Jinglei Liu, Zhenping Hu, Xiaoyang Zeng, Chixiao Chen |
ISCAS | 4 |
| 2024 | GauSPU: 3D Gaussian Splatting Processor for Real-Time SLAM Systemsabstract3D Gaussian Splatting (3DGS) has recently emerged as a promising technique in the realms of 3D vision and robotics. Its capacity for rapid rendering and high-fidelity reconstruction makes it an attractive candidate for integration into Simultaneous Localization and Mapping (SLAM) systems. However, existing 3DGS-based SLAM systems still suffer from inadequate tracking throughput due to tremendous recursion in volume rendering and irregular memory access for gradient backpropagation. To address these challenges, this paper proposes GauSPU, an algorithm-hardware co-designed accelerator for supporting real-time 3DGS-based SLAM. On the algorithm side, we present a sparse-tile-sampling (STS) method for efficient pose tracking. The STS focuses on informative image regions, discarding the rest to alleviate computational workload while maintaining accuracy. At the hardware level, we make twofold efforts. Firstly, we design a sparsity-adaptive ray recursion unit (SA-RRU) to accelerate volume rendering by leveraging irregular spatial sparsity. The SA-RRU introduces a sub-tile-wise execution pattern and a Morton-based thread allocation scheme to optimize sparsity utilization. Additionally, a sparsity-aware task dispatcher ensures efficient fine-grained task scheduling. Secondly, we propose a memory-access-relaxed backpropagation engine (MAR-BE) for efficient gradient aggregation. It comprises a gradient buffer unit (GBU) for coalescing partial gradients and a pose backward unit (PBU) for pipeline-fused backpropagation, collaboratively eliminating the costly atomic operations. Sufficient experiments demonstrate that, through the integration of GauSPU and GPU, the system achieves a throughput of 33.6 FPS for real-time pose tracking in 3DGS-SLAM, presenting a significant$63.9\times$improvement in energy efficiency compared to the RTX3090 baseline. Lizhou Wu, Haozhe Zhu, Siqi He, Jiapei Zheng, Chixiao Chen, Xiaoyang Zeng |
MICRO | 2 |
| 2024 | FPIA: Communication-Aware Multi-Chiplet Integration With Field-Programmable Interconnect Fabric on Reusable Silicon InterposerabstractSilicon interposer re-usage is drawing attention for cost-effective multi-chiplet integrated systems. To address the communication awareness of inter/off-chiplet interconnect, the paper proposes a field-programmable interconnect fabric and develops its corresponding automatic physical integration tool. The tile-based fabric consists of turnout, cross-over boxes and parallel tracks. It features micro-bump-wise connecting flexibility and hardware efficiency. The automation flow performs chiplet location optimization and efficient bump-to-bump routing, supporting multi-lane bus interconnect and miscellaneous external ports. The methodology is validated by 9 different integration scenarios, where the routability is guaranteed when the local resource utilization ratio approaches 94.5%. The data’s maximum interconnect latency is 2.2 ns and the energy consumption is 1.18 pJ/bit at a bitrate of 1 Gbps. The latency consumes$16.5\times \sim ~53.4\times $fewer clock cycles than the state-of-the-art network-on-package-based reusable interposer architectures. Bo Jiao 0003, Haozhe Zhu, Jundong Zhu, Dexin Wen, Lingli Wang, Jun Tao 0001, Chixiao Chen, Yinhe Han 0001, Qi Liu 0010, Ninghui Sun, Ming Liu 0022 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | Hi-NeRF: A Multicore NeRF Accelerator With Hierarchical Empty Space Skipping for Edge 3-D RenderingabstractNeural radiance field (NeRF) has proved to be promising in augmented/virtual-reality applications. However, the deployment of NeRF on edge devices suffers from inadequate throughput due to redundant ray sampling and congested memory access. To address these challenges, this article proposes Hi-NeRF, a multirendering-core accelerator for efficient edge NeRF rendering. On the architecture level, a hierarchical empty space skipping (HESS) scheme is adopted, which efficiently locates the effective samples with fewer skipping steps and thus accelerates the ray marching process. Furthermore, to alleviate the memory access bottleneck, a vertex-interleaved mapping (VIM) method that eliminates memory bank conflicts is also proposed. On the hardware level, ineffective sample filters (ISFs) and voxel access filters (VCFs) are introduced to further exploit spatial sparsity and data locality at run-time. The experimental results show that our work achieves$2.67\times $rendering throughput and$11.2\times $energy efficiency compared to a SOTA NeRF rendering accelerator. The energy efficiency can be improved by$561\times $compared to a commercial GPU. Lizhou Wu, Haozhe Zhu, Jiapei Zheng, Yinuo Cheng, Qi Liu 0010, Xiaoyang Zeng, Chixiao Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | A $2.53 \mu \mathrm{W}/\text{channel}$ Event-Driven Neural Spike Sorting Processor with Sparsity-Aware Computing-In-Memory MacrosabstractSpike sorting processors with high energy efficiency are widely used in large-scale neural signal processing tasks to monitor the activity of neurons in brains. This paper presents a low-power processor for high-accuracy spike sorting and on-chip incremental learning using an algorithm-hardware co-design approach. The processor introduces an event-driven mechanism with adaptive-threshold detection to conditionally activate the system in order to reduce power consumption. Sparsity-aware computing-in-memory (CIM) macros are also developed in our design to store templates and perform complicated computations efficiently. The prototype is designed using 28nm technology with an area of 0.018 mm2/channel and an overall power efficiency of$\mathbf{2.53} \mu \mathbf{W}/\mathbf{channel}$and 84nW/(channel.cluster) at the voltage of 0.72V. Moreover, the accuracy of the whole design can reach 94.5% in a 32-channel scenario. Hao Jiang 0024, Jiapei Zheng, Yunzhengmao Wang, Jinshan Zhang 0006, Haozhe Zhu, Liangjian Lyu, Yingping Chen, Chixiao Chen, Qi Liu 0010 |
ISCAS | 5 |
| 2023 | A Scalable Die-to-Die Interconnect with Replay and Repair Schemes for 2.5D/3D IntegrationabstractChiplet is a critical technology in the post-Moore era, and the die-to-die (D2D) interconnect is essential for communication between chiplets. Meanwhile, several edge-computing devices based on 2.5D/3D chiplet have recently emerged. However, a lightweight D2D interconnect for 2.5D/3D edge-computing systems is lacking. Given the differences between 2.5D/3D integration, a scalable D2D interconnect with replay and repair schemes is presented in this paper. A credit-based flow control scheme and a custom replay scheme are presented for high efficiency. An effective detection and repair scheme is proposed to enhance fault tolerance for the D2D interconnect. Compared with a previous D2D interconnect design, the proposed D2D interconnect delivers 1.07/1.09Gbps throughput ($\sim 2.4\times/\sim 3.9\times \text{for}\ \text{write}/\text{read}$) and significantly reduced energy/bit with only ∼1.7× increased hardware cost. Additionally, compared with a previous chip-to-chip interconnect design, the proposed D2D interconnect can be configured down to power consumption as low as 0.55pJ/bit and 38.40Gbps throughput, achieving ∼2.5× throughput and significantly reduced latency with a negligible increase in hardware cost. Bo Jiao 0003, Jinshan Zhang 0006, Shiwei Liu 0002, Hao Jiang 0024, Jun Tao 0001, Wenning Jiang, Qi Liu 0010, Lihua Zhang 0002, Haozhe Zhu, Chixiao Chen |
ISCAS | 10 |
| 2022 | A 11.6μ W Computing-on-Memory-Boundary Keyword Spotting Processor with Joint MFCC-CNN Ternary QuantizationabstractThis paper presents an ultra-low-power keyword spotting processor using an algorithm-architecture co-design approach. Joint MFCC-CNN ternary weight quantization is proposed to reduce power consumption. The Mel filter and the DCT module are merged into one matrix multiplication. The merged coefficients and the weights of the rest NN classifier are ternary-quantized, causing less than 3% accuracy loss but 39 × energy efficiency improvement. Moreover, a Computing-on-Memory-Boundary macro is adopted to store the quantized coefficients and weights, and perform matrix multiplications. Compared to the existing computing-in-memory technology, the proposed technique can reduce power consumption due to the higher utilization ratio. To verify the proposed techniques, a keyword spotting processor prototype is designed with 28nm CMOS technology. Simulation results show that the prototype achieves power consumption of 11.6μ W under a power supply of 0.72V and a clock frequency of 250KHz. Xinru Jia, Haozhe Zhu, Yunzheng Wang, Jinshan Zhang 0006, Xiankui Xiong, Dong Xu 0015, Chixiao Chen, Qi Liu 0010 |
ISCAS | 2 |
| 2021 | A 0.57-GOPS/DSP Object Detection PIM Accelerator on FPGAabstractThe paper presents an object detection accelerator featuring a processing-in-memory (PIM) architecture on FPGAs. PIM architectures are well known for their energy efficiency and avoidance of the memory wall. In the accelerator, a PIM unit is developed using BRAM and LUT based counters, which also helps to improve the DSP performance density. The overall architecture consists of 64 PIM units and three memory buffers to store inter-layer results. A shrunk and quantized Tiny-YOLO network is mapped to the PIM accelerator, where DRAM access is fully eliminated during inference. The design achieves a throughput of 201.6 GOPs at 100MHz clock rate and correspondingly, a performance density of 0.57 GOPS/DSP. Bo Jiao 0003, Jinshan Zhang 0006, Yuanyuan Xie, Shunli Wang 0001, Haozhe Zhu, Xiaoyang Kang 0001, Zhiyan Dong, Lihua Zhang 0002, Chixiao Chen |
ASP-DAC | 5 |
| 2021 | Computing Utilization Enhancement for Chiplet-based Homogeneous Processing-in-Memory Deep Learning ProcessorsabstractThis paper presents a design strategy of chiplet-based processing-in-memory systems for deep neural network applications. Monolithic silicon chips are area and power limited, failing to catch the recent rapid growth of deep learning algorithms. The paper first demonstrates a straightforward layer-wise method that partitions the workload of a monolithic accelerator to a multi-chiplet pipeline. A quantitative analysis shows that the straightforward separation degrades the overall utilization of computing resources due to the reduced on-chiplet memory size, thus introducing a higher memory wall. A tile interleaving strategy is proposed to overcome such degradation. This strategy can segment one layer to different chiplets which maximizes the computing utilization. To facilitate the strategy, the modification of the chiplet system hardware is also discussed. To validate the proposed strategy, a nine-chiplet processing-in-memory system is evaluated with a custom-designed object detection network. Each chiplet can achieve a peak performance of 204.8GOPS at a 100-MHz rate. The peak performance of the overall system is 1.711TOPS, where no off-chip memory access is needed. By the tile interleaving strategy, the utilization is improved from 53.9 to 92.8 Bo Jiao 0003, Haozhe Zhu, Jinshan Zhang 0006, Shunli Wang 0001, Xiaoyang Kang 0001, Lihua Zhang 0002, Mingyu Wang 0001, Chixiao Chen |
ACM Great Lakes Symposium on VLSI | 2 |
| 2021 | ALPINE: An Agile Processing-in-Memory Macro Compilation FrameworkabstractProcessing-in-Memory architectures and circuit designs are playing significant roles in the recent energy-efficient machine learning chips. This paper proposes a PIM macro compilation framework called ALPINE to speed up previously tedious and error-prone PIM design flow, paving the way towards open-source and process-portable PIM chips. Relying on an extensible PIM standard cell library, ALPINE can generate the corresponding topology according to the specification, and process placement and routing. The proposed PIM macro is compatible with different storage devices such as SRAM and RRAM, and can support various quantization bit-widths and dataflows. To verify the effectiveness, a 128×128 SRAM-based PIM macro instance is implemented, and the simulation results show that it can achieve an energy efficiency of 19.05TOPS/W under 65nm CMOS technology. The macro performance is not inferior to the state-of-the-art custom PIM designs. Jinshan Zhang 0006, Bo Jiao 0003, Yunzhengmao Wang, Haozhe Zhu, Lihua Zhang 0002, Chixiao Chen |
ACM Great Lakes Symposium on VLSI | 4 |