Yuhao Shu

dblp:291/6338 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0002-0357-4507ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 2 first-author · 16 since 2021
YearPublicationVenuePosition
2026 A Complementary 3T-Based eDRAM Macro for High-Density Dual-Direction CAM and Logic-in-Memory
abstract
Content-addressable memory (CAM) is regarded as an attractive solution for data-intensive applications with high-density search demands. To further improve functional flexibility yet at a low cost, several CAM macros have been developed to support multiple bit-wise logic operations. However, conventional SRAM-based CAM designs are constrained by the large bitcell area, posing significant challenges to achieve higher density. To address this issue, we propose a complementary 3T (C3T) based embedded dynamic random access memory (eDRAM) macro for high-density dual-direction CAM searching and logic-in-memory operations. First, we propose a compact C3T bitcell featuring a pair of complementary decoupled read ports, enabling dual-port read and efficient CAM operations. Second, we present a compact dynamic-circuit-based sense amplifier (DSA) to optimize the area of readout peripheral circuitry while mitigating the read bit line saturation issue. Additionally, we implement dual-direction CAM searching and logic-in-memory operations exploiting the C3T-based eDRAM macro. A 4 Kb C3T-based eDRAM macro has been validated in a commercial 40-nm CMOS process. Post-layout results demonstrate a 53% reduction in the bitcell area and a 58.1% reduction in the macro area compared to the state-of-the-art 6T compute SRAM. Moreover, the proposed design achieves a maximum frequency of 578 MHz for binary CAM (BCAM) searching operations and 694 MHz for logic operations, with energy consumption of 1.12 fJ/bit and 26.8 fJ/bit, respectively.
Lintao Lan, Yuhao Shu, Hui Wang 0036, Yajun Ha
IEEE Trans. Circuits Syst. I Regul. Pap.6
2026 Design of an Aging-Aware Memory With BTI-Mitigated SA and System-Visible Lifetime Management
abstract
The deployment of large language models (LLMs) is critically dependent on the key-value (KV) cache. However, the “write-less, read-many” access pattern of the SRAM tier of KV cache in LLMs poses a reliability challenge. Combined with CMOS scaling, it markedly accelerates aging mechanisms like bias temperature instability (BTI). BTI increases the threshold voltage and slows the devices, posing a severe threat to the reliability of the high-speed KV Cache SRAMs. To address this challenge, this paper proposes an architecture that couples in-situ, path-faithful sensing while exposing interfaces and guidance for system-level reliability management. In this paper,First, we propose an aging-enhanced SRAM cell and sensing amplifier (SA) and their configurable circuit topology to support an aging sensing interface.Second, we present a reconfigurable TDC-based in-situ aging detection circuit. It either detects the aging of cells through the aging sensing interface or replaces the aged cells to maintain their original timing performance.Third, we present a lifetime management strategy that uses the aging detection circuit to monitor the aging status of memory, control its operation modes, and alert its end-of-life alarm. A 28-nm 64-Kbit memory circuit has been constructed to validate the optimizations above. Experiment results show that our proposed aging-aware memory achieves a lifetime of 13.1 years, and after 10 years of operation, the presented aging-aware memory can achieve a maximum operating frequency of 1.47 GHz. Compared with state-of-the-art designs, our proposed design achieves a system lifetime improvement of$\gt 3.9\times $, a maximum SA FoM of 29.5, which is about$2\times $that of SOTA SA designs.
Jianwen Luo 0004, Lintao Lan, Yuhao Shu, Yajun Ha
IEEE Trans. Circuits Syst. I Regul. Pap.7
2026 ACIMC: A 342.7-TOPS/mm2 eDRAM-Based Analog Cryogenic In-Memory Computing Macro
abstract
Cryogenic in-memory computing (IMC) emerges as a promising approach for achieving high computing density and parallel data processing at extremely low temperatures. However, existing IMC macros usually utilize single-bit storage per cell, impeding further improvements in computing density. This article presents a 128-Kb embedded dynamic random access memory (eDRAM)-based analog cryogenic IMC (ACIMC) macro with three key techniques to achieve high computing density. First, we optimize an area-efficient dual three-transistor-zero-capacitor (3T0C) eDRAM bitcell to support 4-bit signed weight storage. Second, we design an area-efficient nonlinear write circuit that ensures a linear mapping between the digital weight and the resulting computing current. Third, we present a fast 4-bit flash analog-to-digital converter (ADC) featuring the column-generated reference voltage scheme and reference-storage sense amplifiers to achieve high-speed cryogenic convolutions. Measurement results from our test chip show that the proposed ACIMC achieves a computing density of 342.7 TOPS/mm2and an energy efficiency of 391.2 TOPS/W for 4b$\times 4$b cryogenic convolutions. Moreover, the retention time of our ACIMC is improved to 82.6 ms at 4.2 K.
Yuhao Shu, Hongtu Zhang, Hao Sun 0035, Weiqiang Liu 0001, Yajun Ha
IEEE Trans. Circuits Syst. I Regul. Pap.1
2026 M3CAM: An MLC RRAM-Based Multi-Bit CAM Design Supporting In-Memory Operation of Multi-State Hamming Distance
abstract
The multi-state Hamming distance (MSHD) is a crucial metric for evaluating the similarity of symbolic sequences in data-intensive search applications, such as genomic analysis. MSHD search is performed by comparing inputs against database entries, which can be efficiently accelerated by content-addressable memories (CAMs), such as multi-level cell (MLC) RRAM CAMs. However, designing MLC RRAM-based MSHD CAM faces critical challenges in area efficiency, primarily due to large CAM cells, complex MSHD computing circuits, and bulky variation-compensating input circuits. To address these issues, we propose three techniques to develop M3CAM, an MLC RRAM-based multi-bit CAM (MCAM) supporting in-memory operation of MSHD. First, we propose a 5T1R MCAM cell with MLC RRAM to support dense symbol matching. Second, we propose an MSHD in-memory operation circuit with only three transistors to support dense MSHD computation. Third, we propose a feedback-driven adaptive input DAC to enable minimal area-overhead compensation for RRAM variation. Compared to the state-of-the-art, the proposed M3CAM reduces MCAM cell area by 74.2%, and reduces MSHD computing circuit transistors by 20%, while expanding the MSHD search range to$16\times $.
Tiankuo Zheng, Chenxin Jiang, Yuhao Shu, Chunmeng Dou, Yajun Ha
IEEE Trans. Circuits Syst. I Regul. Pap.4
2026 DSHD-CAM: High-Throughput RRAM CAM Leveraging Dynamic Shifted Hamming Distance for Genome Analysis
abstract
Genome analysis has been critical in various applications, such as infectious disease control. High throughput is an essential requirement for genome analysis in data-intensive scenarios, which requires acceleration by content-addressable memory (CAM) with parallel comparison capability. However, existing genome analysis CAMs still face inadequate throughput issues due to large cell area, excessive array storage redundancy, and inefficient comparison algorithms. To address these issues, we propose a high-throughput dynamic shifted Hamming distance (SHD) resistive random access memory (RRAM)-based CAM (DSHD-CAM) that leverages the characteristics of genome analysis. First, we propose a compact RRAM-based CAM cell utilizing one-hot encoding and time-domain computation to minimize cell area. Second, we propose a dense CAM array utilizing an efficient storage scheme to reduce array storage redundancy. Third, we propose a dynamic SHD search algorithm filtering out low-match-potential cases to reduce search latency. Compared to the state-of-the-art (SOTA), DSHD-CAM achieves an average of$7.58\times $higher throughput under the same area constraints while maintaining competitive sensitivity and precision.
Chenxin Jiang, Tiankuo Zheng, Rui Li 0095, Yuhao Shu, Yajun Ha
IEEE Trans. Very Large Scale Integr. Syst.5
2025 RSQC: Recursive Sparse QUBO Construction for Quantum Annealing Machines
abstract
Quantum annealing algorithms have shown commercial potential in solving some instances of combinatorial optimization problems. However, existing mapping for general optimization problems into a compatible format for quantum annealing yields dense topology and complicated weighting, which limits the size of solvable problems on practical quantum annealing machines. To address this issue, we propose a novel mapping framework with three new techniques. First, to address the issue from general constraints, we introduce a recursive methodology to map constraints into interconnected Boolean gates and small algebraic cliques, which yields sparse topology and hardware-friendly biases/interactions. Second, to better address frequently-used constraints, we introduce a specialized penalty set based on this methodology with detailed optimizations. Third, to address the issue from the objective, we reformulate the complicated objective into a single multi-bit variable and apply binary search to its range, which turns each search step into a constraint-only problem. Compared with the state-of-the-art, experimental results and analysis over an exhaustive scan for operand bit-widths from 1 to 64 show that: (1) the growth order of the number of physical qubits with regard to operand bit-widths is reduced fromO(w2) toO(w), while the number is reduced by a factor of 10-1 in the best case; (2) the dynamic range of biases/interactions is reduced fromO(22w) to−2in the best case. For the same optimization problem, our framework reduces the requirement of the number of physical qubits and machine precision, and shortens the time from problem to machine.
Jianwen Luo 0004, Yuhao Shu, Yajun Ha
IEEE Trans. Computers2
2025 An Energy-Efficient and Real-Time FPGA-Based Point Cloud Registration Framework with Ultra-Fast and Configurable Multi-Mode Correspondence Search
abstract
Point cloud registration is a fundamental task in LiDAR-based localization and mapping, widely employed in robotics and autonomous vehicles. However, existing registration solutions lose geometric topology continuity and lack scanline-aware, configurable correspondence search, restricting their real-time applicability. To solve these issues, we propose a fundamentally re-architected, energy-efficient FPGA framework for real-time point cloud registration, featuring a configurable, multi-mode correspondence search engine. First, we introduce a scanline-aided range-projection structure (SA-RPS) that reorganizes LiDAR points within configurable segmentation domains into contiguous memory while preserving scanline topology, enabling efficient and flexible multi-mode correspondence search. Second, we develop a deeply pipelined, ultra-fast SA-RPS-based correspondence search (SA-RPS-CS) accelerator that supports dynamic configuration of search mode and parallelism and incorporates a sliding-window cache and scanline-aware K-selection module for high-throughput, multi-mode correspondence extraction. Third, we present a co-designed registration framework that integrates the accelerator with dynamic parameter configuration, enabling adaptive, real-time processing across diverse SLAM scenarios. Experimental results demonstrate that the proposed SA-RPS-CS accelerator delivers \(2.3\times\) – \(32.4\times\) faster search and \(1.8\times\) – \(26.2\times\) higher energy efficiency than previous state-of-the-art FPGA designs, achieving real-time registration for 64-channel LiDAR at 21.5 FPS with negligible loss in accuracy.
Hao Sun 0035, Yuhao Shu, Jianzhong Xiao, Weixiong Jiang, Hui Wang 0036, Yajun Ha
ACM Trans. Reconfigurable Technol. Syst.3
2025 A 5T0C eDRAM-Based Content Addressable Memory for High-Density Searching and Logic-in-Memory
abstract
With the development of big data, there is an increasing demand for high-density searching, where content-addressable memory (CAM) presents an attractive solution for its ability to perform parallel searches. However, this goal is constrained by the difficulty of further reducing the area of SRAM cells, which is commonly used in traditional CAM implementations. To address this issue, we propose a novel CAM with a compact five-transistor-zero-capacitor (5T0C)-embedded dynamic random access memory (eDRAM) for high-density searching and logic-in-memory applications. First, we propose the 5T0C eDRAM gain cell featuring a 3T0C write port and a decoupled read port of 2T to achieve data storage and searching operations. Second, we present a reconfigurable sense amplifier (RSA) design with two different reference voltages to optimize the area overhead of peripheral circuits and support logic operations. Moreover, the 5T0C eDRAM-based CAM can be employed to achieve high-density searching and logic operations. We have validated the eDRAM-based CAM array in the 40-nm CMOS process. The postlayout simulation results show that our design achieves over 15% higher memory density compared to the state-of-the-art 6T SRAM. Additionally, it supports a maximum frequency of 637 and 658 MHz for binary CAM (BCAM) searching and logic operations, while consuming 0.91 and 27.47 fJ/bit at 1.1 V, respectively.
Yuhao Shu, Lintao Lan, Hongtu Zhang, Yajun Ha
IEEE Trans. Very Large Scale Integr. Syst.2
2024 The Optimization of Aging-aware 8T SRAM for FPGA Configuration Memory
abstract
Bias temperature instability (BTI) has posed increasingly long-term reliability issues in modern static random access memory (SRAM) applications, especially for configuration memory in FPGA. In this work, we present an aging-aware 8T SRAM design for the implementation of FPGA configuration memory. First, we adopt a body-source short scheme to bias the p-body at VDD/2, which effectively mitigates the negative BTI for PMOS. Second, we bias the leaking transistors in the supercutoff region with an optimized dynamic leakage-suppression logic, which helps to reduce the leakage power (thus lower temperature) and further mitigate the impact of BTI. Third, we achieve an operating voltage range from 0.9V to 0.6V of the 8T SRAM, which benefits the optimization of lifetimes and leakage power in the idle mode (0.6V). Compared with the state-of-the-art, post-layout results with the TSMC 28nm aging model show that our 8T SRAM-based FPGA achieves the highest figure of merit (FoM) for lifetimes with a 2.33× extension. Moreover, it further achieves a 65.2% reduction in leakage power in the idle mode.
Yuhao Shu, Yajun Ha
ISCAS3
2024 EarFDA: A Lightweight and Energy-Efficient Fall Detection Accelerator for Ear-Worn Devices
abstract
Fall detection systems are crucial in preventing severe injuries among people with limited mobility. However, current systems either employ multiple sensors and redundant features, or adopt computation-intensive detection methods, resulting in a high computational burden and inefficient energy consumption. To address this issue, we propose a lightweight and energy-efficient fall detection accelerator for ear-worn devices, utilizing only one sensor and limited features. First, we propose an accurate and robust pre-processing method that employs a wider overlapping sliding window to extract a superior feature combination. Second, based on the extracted features, we develop a lightweight detection network with low-bit quantization, integrating hybrid convolutions. Third, we design a full-dataflow hardware accelerator for the proposed network to enhance energy efficiency. Experimental results on a public dataset show that the proposed FPGA accelerator achieves superior accuracy compared to the state-of-the-art algorithms running on the CPU, with a speed improvement of at least 31.4 times. Additionally, our accelerator significantly boosts energy efficiency by at least 274 times.
Zhaodong Lv, Hao Sun 0035, Yuhao Shu, Yajun Ha
ISCAS3
2023 CSDB-eDRAM: A 16Kb Energy-Efficient 4T CSDB Gain Cell eDRAM with over 16.6s Retention Time and 49.23uW/Kb at 4.2K for Cryogenic Computing
abstract
Gain-cell based eDRAM is an appealing candidate as the main memory in cryogenic computing for its high density and low power consumption. However, existing eDRAMs fail to achieve higher energy efficiency due to the higher energy consumption in either the retention or dynamic access operations. To solve this issue, we propose three techniques to achieve a 16Kb energy-efficient CSDB-eDRAM for cryogenic memory implementation. First, we propose a 4T CSDB-GC that is able to significantly improve the retention time. Second, we propose a wordline voltage off-chip tuning method to enhance the dual-port read speed and read-disturb free operations. Third, we introduce a bitline split scheme to reduce the dynamic power overhead of each access operation. Measurement results from our fabricated chip show that the dynamic power of our CSDB-eDRAM has been reduced to 49.23 uW/Kb at 1.41 GHz, which outperforms the state-of-the-art by$\mathbf{11.4}\times$. It also achieves the best data retention time of 16.67 s at 4.2 K. Moreover, a negligible retention power of 0.11 pW/Kb can be achieved.
Yuhao Shu, Hongtu Zhang, Hao Sun 0035, Yajun Ha
ISCAS1
2023 RPS-KNN: An Ultra-Fast FPGA Accelerator of Range-Projection-Structure K-Nearest-Neighbor Search for LiDAR Odometry in Smart Vehicles
abstract
KNN (K-Nearest-Neighbors) search has been widely used in LiDAR-related applications. As LiDAR's point clouds become more massive, it is a great challenge to implement a fast and energy-efficient KNN implementation. Previous works consume much time in either building an efficient data structure or searching in the data structure. To solve this issue, we propose a high-locality data structure RPS (range-projection-structure) and an ultra-fast FPGA accelerator of the building and searching process. First, we propose a novel data structure RPS which ensures the points with similar projection locations and range scales are stored in the continuous locations of a memory. Second, we propose a highly-parallel method to build the RPS by projecting the points into a point cloud matrix and parallelly processing the points in a column. Third, based on RPS, we propose a highly-parallel KNN search algorithm, which can quickly narrow the search region and select the KNN from neighboring points in parallel. Experimental results show that our method achieves 13.7 times faster than other FPGA implementations. Moreover, energy efficiency results show that our proposed method is 27.4 times and 50.7 times higher than the state-of-the-art implementations on FPGA and GPU platforms, respectively.
Jianzhong Xiao, Hao Sun 0035, Hongtu Zhang, Chengzhang He, Yuhao Shu, Yajun Ha
ISCAS7
2023 An Energy-Efficient Stream-Based FPGA Implementation of Feature Extraction Algorithm for LiDAR Point Clouds With Effective Local-Search
abstract
Feature extraction is a fundamental and essential step in light detection and ranging (LiDAR) based simultaneously localization and mapping (SLAM) algorithms. Considering the run-time requirement of feature extraction and the stringent battery constraint in smart vehicles, it is a great challenge to develop fast and highly energy-efficient feature extraction implementation for massive point clouds. Unfortunately, existing implementations not only fail to exploit the available parallelism but also fail to make full use of the local information to optimize the computations. To solve the issue, we propose three novel techniques to achieve a fast and energy-efficient FPGA implementation of the feature extraction algorithm with effective local search. First, we propose a low-complexity projection method and a column-scanning scheduler to organize the irregular and sparse point cloud into a well-organized point cloud matrix. Second, based on the point cloud matrix, we exploit its local information and propose a high-parallel method to detect the coarse-grain feature points. Third, we propose a high-parallel conditional priority queue to progressively and evenly select the fine-grain feature points. Experimental results on the KITTI dataset show that our method implemented on the ZCU104 FPGA board achieves the best accuracy and reaches 584 frames per second (FPS) for the feature extraction of a 64-laser LiDAR’s point cloud. Moreover, our proposal achieves the best energy efficiency, which is on average 11.7 times and 9.0 times higher than the state-of-the-art implementations on the GPU and FPGA platforms, respectively.
Hao Sun 0035, Yuhao Shu, Yajun Ha
IEEE Trans. Circuits Syst. I Regul. Pap.4
2023 WDVR-RAM: A 0.25-1.2 V, 2.6-76 POPS/W Charge-Domain In-Memory-Computing Binarized CNN Accelerator for Dynamic AIoT Workloads
abstract
In-memory computing (IMC) is an effective approach to accelerate the interference tasks of binary neural networks (BNN), which has been widely used in artificial intelligence of things (AIoT) applications. However, previous researches only focus on optimizing energy efficiency for a narrow voltage range. This poses a severe limitation for some AIoT applications, because they may have very dynamic workloads that require the energy efficiency optimization of BNN for a wide dynamic voltage range (WDVR). To address this issue, we have developed a novel IMC-based BNN accelerator, supporting energy-efficient operations in a wide dynamic voltage range. First, we analyze different charge-domain architectures in terms of their errors and energy characteristics, and decide the best architecture that is suitable for a wide dynamic voltage range. Second, we introduce a lazy convolution bitline reset (LCBR) scheme to further optimize the energy of multiplication and accumulation (MAC) within the entire voltage range. Third, we design a subthreshold differential batch-normalization amplifier (SDBNA) array to compensate for the inference accuracy loss for the lower end of the voltage range. A 16-Kb WDVR-RAM has been designed and fabricated in a 55-nm CMOS process. Measurement results show that the test chip achieves a peak energy efficiency of 76193-2645 TOPS/W at 0.25-1.2 V for MAC, and a peak energy efficiency of 56141-1973 TOPS/W at 0.25-1.2V for both MAC and batch normalization (BN) layers. For throughput, our work obtains a peak computing density range of 692-144078 GOPS/mm2 with 28.31-0.17 ms/frame (CIFAR-10). Moreover, it also achieves the highest 14.75% and 31.87% recovery for Top-1 accuracy observed for MNIST and CIFAR-10 so far, respectively.
Hongtu Zhang, Yuhao Shu, Hao Sun 0035, Yajun Ha
IEEE Trans. Circuits Syst. I Regul. Pap.2
2022 A Reliable 8T SRAM for High-Speed Searching and Logic-in-Memory Operations
abstract
To efficiently implement searching and logic functions with the SRAM-based in-memory computing (IMC), we need to perform computations on bitlines (BLs) (called compute access) via multiple wordline (WL) activations. However, this may cause prominent read disturbance when the IMC is implemented with the standard 6 T SRAM. To address this reliability issue, existing solutions adopt either auxiliary assistance circuits or alternative bitcell topologies, but they lead to substantial overheads of the access speed or array density. In this article, we propose a novel 8T compute SRAM (CSRAM) for reliable and high-speed in-memory searching and compound logic-in-memory computations. Our 8T CSRAM features a pair of pMOS access transistors and split-WLs dedicated to the compute access. A thorough circuit-level analysis reveals that the pMOS-based compute access port is essential for significantly mitigating the read disturbance. Moreover, we propose an elevated precharge voltage scheme and a low-skewed inverter-based sensing amplifier to improve the sensing speed. We have validated the proposed 8T CSRAM design in a 16 Kb array with a 28-nm CMOS technology. Compared to the state-of-the-art 8 T CSRAM, results show that our design is not only reliable but also 3.1 times faster, with a maximum operating frequency upping to 2.44 GHz.
Yuqi Wang 0004, Yuhao Shu, Weixiong Jiang, Yajun Ha
IEEE Trans. Very Large Scale Integr. Syst.4
2022 FODM: A Framework for Accurate Online Delay Measurement Supporting All Timing Paths in FPGA
abstract
Voltage and frequency scaling (VFS) has been widely used to improve energy efficiency, lifespan, and system reliability by converting conservative timing margins into$V_{\text {dd}}$reduction. Along these lines, to investigate the potential implementation of VFS technique in exploring the timing margins under different voltages and frequencies,in situor online circuit delay measurement is required to monitor all timing paths, which are usually ended with terminal registers. The previously reported online delay measurement approaches require the output of a terminal register to be measurable. However, some FPGA timing paths are ended with embedded hardcores such as DSPs or BRAMs. It is impossible to measure the output of the terminal register inside a hardcore. To address the issue, we propose an online delay monitor (ODM) that can accurately measure the delay of any type of timing path in real-time conditions. The ODM is mainly composed of two shadow registers and a phase-shifted clock. The shadow registers use a phase-shifted clock signal as the input and the output signal of the combinational logic as the clock. In addition, we present an automatic tool and its corresponding design flow (FODM) for inserting an ODM to monitor a path. Compared with the state-of-the-art, our experimental results indicate that the proposed method has the ability to accurately measure the delays online for all the potential timing paths, regardless of their path termination types. Moreover, we demonstrate an average measurement error of only 1.51% using eight floating-point operators at different voltages.
Weixiong Jiang, Heng Yu 0001, Hongtu Zhang, Yuhao Shu, Rui Li 0095, Yajun Ha
IEEE Trans. Very Large Scale Integr. Syst.4