Hongtu Zhang

dblp:139/1148 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ACIMC: A 342.7-TOPS/mm2 eDRAM-Based Analog Cryogenic In-Memory Computing Macro
abstract
Cryogenic in-memory computing (IMC) emerges as a promising approach for achieving high computing density and parallel data processing at extremely low temperatures. However, existing IMC macros usually utilize single-bit storage per cell, impeding further improvements in computing density. This article presents a 128-Kb embedded dynamic random access memory (eDRAM)-based analog cryogenic IMC (ACIMC) macro with three key techniques to achieve high computing density. First, we optimize an area-efficient dual three-transistor-zero-capacitor (3T0C) eDRAM bitcell to support 4-bit signed weight storage. Second, we design an area-efficient nonlinear write circuit that ensures a linear mapping between the digital weight and the resulting computing current. Third, we present a fast 4-bit flash analog-to-digital converter (ADC) featuring the column-generated reference voltage scheme and reference-storage sense amplifiers to achieve high-speed cryogenic convolutions. Measurement results from our test chip show that the proposed ACIMC achieves a computing density of 342.7 TOPS/mm2and an energy efficiency of 391.2 TOPS/W for 4b$\times 4$b cryogenic convolutions. Moreover, the retention time of our ACIMC is improved to 82.6 ms at 4.2 K.
Yuhao Shu, Hongtu Zhang, Hao Sun 0035, Weiqiang Liu 0001, Yajun Ha
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 A 5T0C eDRAM-Based Content Addressable Memory for High-Density Searching and Logic-in-Memory
abstract
With the development of big data, there is an increasing demand for high-density searching, where content-addressable memory (CAM) presents an attractive solution for its ability to perform parallel searches. However, this goal is constrained by the difficulty of further reducing the area of SRAM cells, which is commonly used in traditional CAM implementations. To address this issue, we propose a novel CAM with a compact five-transistor-zero-capacitor (5T0C)-embedded dynamic random access memory (eDRAM) for high-density searching and logic-in-memory applications. First, we propose the 5T0C eDRAM gain cell featuring a 3T0C write port and a decoupled read port of 2T to achieve data storage and searching operations. Second, we present a reconfigurable sense amplifier (RSA) design with two different reference voltages to optimize the area overhead of peripheral circuits and support logic operations. Moreover, the 5T0C eDRAM-based CAM can be employed to achieve high-density searching and logic operations. We have validated the eDRAM-based CAM array in the 40-nm CMOS process. The postlayout simulation results show that our design achieves over 15% higher memory density compared to the state-of-the-art 6T SRAM. Additionally, it supports a maximum frequency of 637 and 658 MHz for binary CAM (BCAM) searching and logic operations, while consuming 0.91 and 27.47 fJ/bit at 1.1 V, respectively.
Yuhao Shu, Lintao Lan, Hongtu Zhang, Yajun Ha
IEEE Trans. Very Large Scale Integr. Syst.7
2024 Precise and Efficient Third-party Java Libraries Identification Tool for Collaborative Software
abstract
Collaborative systems frequently depend on various software components, like third-party libraries (TPLs), to execute their functions and expedite the development of the system. The security of an entire collaboration system can be compromised by a TPL that is vulnerable, particularly in an industrial setting. Unfortunately, current TPL detection tools encounter difficulties in precisely identifying version levels and exhibit inefficiency in detecting TPLs on a large scale.To address these challenges, we recommend JHunter, a precise and efficient tool for detecting TPL version details. Our approach involves introducing a novel concept called the attribute class dependency graph (ACDG) as a feature at the package level for TPLs. We then utilise a graph neural network-based method to compare the similarity of ACDGs and identify a list of candidate TPLs. Later, we use more detailed class-level features, such as Control Flow Graphs (CFGs), and constant features to determine version-specific information. We collected 19,095 different versions of TPLs from Maven to build our feature database. Our analysis demonstrates the effectiveness of JHunter on a real-world dataset, achieving F1 scores of 99.34% and 97.28% at the library and version levels, respectively, surpassing previous state-of-the-art (SOTA) results.
Hongtu Zhang, Jingdong Guo, Laile Xi, Sidy Tambadou, Fang Zuo, Hong Li 0004
CSCWD2
2023 CSDB-eDRAM: A 16Kb Energy-Efficient 4T CSDB Gain Cell eDRAM with over 16.6s Retention Time and 49.23uW/Kb at 4.2K for Cryogenic Computing
abstract
Gain-cell based eDRAM is an appealing candidate as the main memory in cryogenic computing for its high density and low power consumption. However, existing eDRAMs fail to achieve higher energy efficiency due to the higher energy consumption in either the retention or dynamic access operations. To solve this issue, we propose three techniques to achieve a 16Kb energy-efficient CSDB-eDRAM for cryogenic memory implementation. First, we propose a 4T CSDB-GC that is able to significantly improve the retention time. Second, we propose a wordline voltage off-chip tuning method to enhance the dual-port read speed and read-disturb free operations. Third, we introduce a bitline split scheme to reduce the dynamic power overhead of each access operation. Measurement results from our fabricated chip show that the dynamic power of our CSDB-eDRAM has been reduced to 49.23 uW/Kb at 1.41 GHz, which outperforms the state-of-the-art by$\mathbf{11.4}\times$. It also achieves the best data retention time of 16.67 s at 4.2 K. Moreover, a negligible retention power of 0.11 pW/Kb can be achieved.
Yuhao Shu, Hongtu Zhang, Hao Sun 0035, Yajun Ha
ISCAS2
2023 RPS-KNN: An Ultra-Fast FPGA Accelerator of Range-Projection-Structure K-Nearest-Neighbor Search for LiDAR Odometry in Smart Vehicles
abstract
KNN (K-Nearest-Neighbors) search has been widely used in LiDAR-related applications. As LiDAR's point clouds become more massive, it is a great challenge to implement a fast and energy-efficient KNN implementation. Previous works consume much time in either building an efficient data structure or searching in the data structure. To solve this issue, we propose a high-locality data structure RPS (range-projection-structure) and an ultra-fast FPGA accelerator of the building and searching process. First, we propose a novel data structure RPS which ensures the points with similar projection locations and range scales are stored in the continuous locations of a memory. Second, we propose a highly-parallel method to build the RPS by projecting the points into a point cloud matrix and parallelly processing the points in a column. Third, based on RPS, we propose a highly-parallel KNN search algorithm, which can quickly narrow the search region and select the KNN from neighboring points in parallel. Experimental results show that our method achieves 13.7 times faster than other FPGA implementations. Moreover, energy efficiency results show that our proposed method is 27.4 times and 50.7 times higher than the state-of-the-art implementations on FPGA and GPU platforms, respectively.
Jianzhong Xiao, Hao Sun 0035, Hongtu Zhang, Chengzhang He, Yuhao Shu, Yajun Ha
ISCAS5
2023 WDVR-RAM: A 0.25-1.2 V, 2.6-76 POPS/W Charge-Domain In-Memory-Computing Binarized CNN Accelerator for Dynamic AIoT Workloads
abstract
In-memory computing (IMC) is an effective approach to accelerate the interference tasks of binary neural networks (BNN), which has been widely used in artificial intelligence of things (AIoT) applications. However, previous researches only focus on optimizing energy efficiency for a narrow voltage range. This poses a severe limitation for some AIoT applications, because they may have very dynamic workloads that require the energy efficiency optimization of BNN for a wide dynamic voltage range (WDVR). To address this issue, we have developed a novel IMC-based BNN accelerator, supporting energy-efficient operations in a wide dynamic voltage range. First, we analyze different charge-domain architectures in terms of their errors and energy characteristics, and decide the best architecture that is suitable for a wide dynamic voltage range. Second, we introduce a lazy convolution bitline reset (LCBR) scheme to further optimize the energy of multiplication and accumulation (MAC) within the entire voltage range. Third, we design a subthreshold differential batch-normalization amplifier (SDBNA) array to compensate for the inference accuracy loss for the lower end of the voltage range. A 16-Kb WDVR-RAM has been designed and fabricated in a 55-nm CMOS process. Measurement results show that the test chip achieves a peak energy efficiency of 76193-2645 TOPS/W at 0.25-1.2 V for MAC, and a peak energy efficiency of 56141-1973 TOPS/W at 0.25-1.2V for both MAC and batch normalization (BN) layers. For throughput, our work obtains a peak computing density range of 692-144078 GOPS/mm2 with 28.31-0.17 ms/frame (CIFAR-10). Moreover, it also achieves the highest 14.75% and 31.87% recovery for Top-1 accuracy observed for MNIST and CIFAR-10 so far, respectively.
Hongtu Zhang, Yuhao Shu, Hao Sun 0035, Yajun Ha
IEEE Trans. Circuits Syst. I Regul. Pap.1
2022 FODM: A Framework for Accurate Online Delay Measurement Supporting All Timing Paths in FPGA
abstract
Voltage and frequency scaling (VFS) has been widely used to improve energy efficiency, lifespan, and system reliability by converting conservative timing margins into$V_{\text {dd}}$reduction. Along these lines, to investigate the potential implementation of VFS technique in exploring the timing margins under different voltages and frequencies,in situor online circuit delay measurement is required to monitor all timing paths, which are usually ended with terminal registers. The previously reported online delay measurement approaches require the output of a terminal register to be measurable. However, some FPGA timing paths are ended with embedded hardcores such as DSPs or BRAMs. It is impossible to measure the output of the terminal register inside a hardcore. To address the issue, we propose an online delay monitor (ODM) that can accurately measure the delay of any type of timing path in real-time conditions. The ODM is mainly composed of two shadow registers and a phase-shifted clock. The shadow registers use a phase-shifted clock signal as the input and the output signal of the combinational logic as the clock. In addition, we present an automatic tool and its corresponding design flow (FODM) for inserting an ODM to monitor a path. Compared with the state-of-the-art, our experimental results indicate that the proposed method has the ability to accurately measure the delays online for all the potential timing paths, regardless of their path termination types. Moreover, we demonstrate an average measurement error of only 1.51% using eight floating-point operators at different voltages.
Weixiong Jiang, Heng Yu 0001, Hongtu Zhang, Yuhao Shu, Rui Li 0095, Yajun Ha
IEEE Trans. Very Large Scale Integr. Syst.3