EDBT 2026 Demo / reviewers in the wild / expert
Hao Sun 0035
dblp:82/2248-35
· DBLP profile ↗
12ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0003-4518-2533ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 1 first-author · 10 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Event-Driven Asynchronous Graph Neural Network FPGA Accelerator for Real-Time Edge VisionabstractEvent-based asynchronous graph neural networks (GNNs) provide a promising solution for real-time edge vision. By leveraging microsecond-level input latency, asynchronous computation, and sparse storage, they show significant potential for low-latency processing under resource constraints. However, existing FPGA-based accelerators for event-driven asynchronous GNNs cannot meet real-time performance owing to critical bottlenecks in memory utilization, parallelism, and computational redundancy. To address these challenges, we propose a novel FPGA accelerator for event-driven asynchronous GNNs, with three key contributions: 1) Memory-efficient graph feature storage with improved readout parallelism to reduce data access time, 2) Parallelism-enhanced hierarchical graph construction with low dependency to reduce computation time, and 3) Redundancy-free parallel graph convolution with reusable partial computation caching to reduce computation time. The proposed accelerator was deployed on a Xilinx ZCU102 MPSoC platform and evaluated on the N-CARS dataset for car recognition. Compared to the state-of-the-art (SOTA), our proposed design achieves an average$27.59\times $speedup with a latency of$0.58\mu $s while delivering higher accuracy and comparable resource consumption. Tianhang Liu, Guangyao Yan, Runhua Wang, Rui Li 0095, Shijie Meng, Hao Sun 0035, Yajun Ha |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2026 | ACIMC: A 342.7-TOPS/mm2 eDRAM-Based Analog Cryogenic In-Memory Computing MacroabstractCryogenic in-memory computing (IMC) emerges as a promising approach for achieving high computing density and parallel data processing at extremely low temperatures. However, existing IMC macros usually utilize single-bit storage per cell, impeding further improvements in computing density. This article presents a 128-Kb embedded dynamic random access memory (eDRAM)-based analog cryogenic IMC (ACIMC) macro with three key techniques to achieve high computing density. First, we optimize an area-efficient dual three-transistor-zero-capacitor (3T0C) eDRAM bitcell to support 4-bit signed weight storage. Second, we design an area-efficient nonlinear write circuit that ensures a linear mapping between the digital weight and the resulting computing current. Third, we present a fast 4-bit flash analog-to-digital converter (ADC) featuring the column-generated reference voltage scheme and reference-storage sense amplifiers to achieve high-speed cryogenic convolutions. Measurement results from our test chip show that the proposed ACIMC achieves a computing density of 342.7 TOPS/mm2and an energy efficiency of 391.2 TOPS/W for 4b$\times 4$b cryogenic convolutions. Moreover, the retention time of our ACIMC is improved to 82.6 ms at 4.2 K. Yuhao Shu, Hongtu Zhang, Hao Sun 0035, Weiqiang Liu 0001, Yajun Ha |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2026 | RTRA: A Robust and Trustworthy Real-Time V2V Authentication Protocol Based on Geohash and Reputation Mechanisms
Hao Sun 0035, Shuqin Luo, Ziyuan Pu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | GVDF-PIA: Group-Based Vehicular Digital Forensics and Proxy-Assisted Integrity Auditing for Intelligent TransportationabstractVehicular digital forensics is crucial for the security of intelligent transportation systems. However, current vehicular forensics approaches suffer from a single source of evidence, weak verification of evidence authenticity, and inadequate integrity assurance for cloud-stored data during vehicle offline periods. To address these issues, we propose a group-based vehicular digital forensics and proxy-assisted integrity auditing (GVDF-PIA) scheme with the following three techniques. First, we design a multi-perspective digital forensics and efficient storage scheme based on vehicular groups, enabling multi-perspective forensics and improving storage efficiency. Second, we construct a proxy-assisted decentralized integrity auditing mechanism that effectively verifies cloud-stored data even when vehicles are offline. Third, we introduce a time-hash-chain-based method for temporal verification of digital evidence, ensuring the authenticity and chronological integrity of vehicular data, thus making it legally enforceable. Security analysis and experimental results based on general evaluation standards demonstrate that the proposed scheme significantly enhances forensic reliability, while also improving storage efficiency, validating its practical applicability in intelligent transportation systems. Hao Sun 0035, Ziyuan Pu |
IEEE Internet Things J. | 2 |
| 2025 | An FPGA-Based Real-Time Loop Closure Detection Framework With Ultra-Fast Descriptor GeneratorabstractLoop closure detection (LCD) is crucial in LiDAR-based Simultaneous Localization and Mapping (SLAM) for smart vehicles, demanding both real-time performance and high-accuracy. Unfortunately, although the state-of-the-art LCD algorithms offer high-accuracy, the large search space in clustering, the high complexity in descriptor computation, and the slow speed in retrieval prevent them from achieving real-time performance. To address the issue, we propose three key techniques to achieve a real-time FPGA-based LCD framework. First, we reduce the clustering time by designing a Range Image-based clustering accelerator that significantly reduces the search space and achieves high-parallelism. Second, we reduce the descriptor computation time by designing an accelerator that selects only high-quality features and employs simplified operations to achieve low complexity. Third, we reduce the retrieval time by proposing a novel dual-descriptor cross-verification mechanism that uses fewer descriptors while maintaining high-accuracy. Compared to the state-of-the-art, experimental results show that our LCD accelerator demonstrates 128.8x performance improvement across multiple datasets in various scenes and achieves the required real-time performance. Shijie Meng, Weixiong Jiang, Jinjie Huang, Hao Sun 0035, Yajun Ha |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | An Energy-Efficient and Real-Time FPGA-Based Point Cloud Registration Framework with Ultra-Fast and Configurable Multi-Mode Correspondence SearchabstractPoint cloud registration is a fundamental task in LiDAR-based localization and mapping, widely employed in robotics and autonomous vehicles. However, existing registration solutions lose geometric topology continuity and lack scanline-aware, configurable correspondence search, restricting their real-time applicability. To solve these issues, we propose a fundamentally re-architected, energy-efficient FPGA framework for real-time point cloud registration, featuring a configurable, multi-mode correspondence search engine. First, we introduce a scanline-aided range-projection structure (SA-RPS) that reorganizes LiDAR points within configurable segmentation domains into contiguous memory while preserving scanline topology, enabling efficient and flexible multi-mode correspondence search. Second, we develop a deeply pipelined, ultra-fast SA-RPS-based correspondence search (SA-RPS-CS) accelerator that supports dynamic configuration of search mode and parallelism and incorporates a sliding-window cache and scanline-aware K-selection module for high-throughput, multi-mode correspondence extraction. Third, we present a co-designed registration framework that integrates the accelerator with dynamic parameter configuration, enabling adaptive, real-time processing across diverse SLAM scenarios. Experimental results demonstrate that the proposed SA-RPS-CS accelerator delivers \(2.3\times\) – \(32.4\times\) faster search and \(1.8\times\) – \(26.2\times\) higher energy efficiency than previous state-of-the-art FPGA designs, achieving real-time registration for 64-channel LiDAR at 21.5 FPS with negligible loss in accuracy. Hao Sun 0035, Yuhao Shu, Jianzhong Xiao, Weixiong Jiang, Hui Wang 0036, Yajun Ha |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2024 | EarFDA: A Lightweight and Energy-Efficient Fall Detection Accelerator for Ear-Worn DevicesabstractFall detection systems are crucial in preventing severe injuries among people with limited mobility. However, current systems either employ multiple sensors and redundant features, or adopt computation-intensive detection methods, resulting in a high computational burden and inefficient energy consumption. To address this issue, we propose a lightweight and energy-efficient fall detection accelerator for ear-worn devices, utilizing only one sensor and limited features. First, we propose an accurate and robust pre-processing method that employs a wider overlapping sliding window to extract a superior feature combination. Second, based on the extracted features, we develop a lightweight detection network with low-bit quantization, integrating hybrid convolutions. Third, we design a full-dataflow hardware accelerator for the proposed network to enhance energy efficiency. Experimental results on a public dataset show that the proposed FPGA accelerator achieves superior accuracy compared to the state-of-the-art algorithms running on the CPU, with a speed improvement of at least 31.4 times. Additionally, our accelerator significantly boosts energy efficiency by at least 274 times. Zhaodong Lv, Hao Sun 0035, Yuhao Shu, Yajun Ha |
ISCAS | 2 |
| 2023 | CSDB-eDRAM: A 16Kb Energy-Efficient 4T CSDB Gain Cell eDRAM with over 16.6s Retention Time and 49.23uW/Kb at 4.2K for Cryogenic ComputingabstractGain-cell based eDRAM is an appealing candidate as the main memory in cryogenic computing for its high density and low power consumption. However, existing eDRAMs fail to achieve higher energy efficiency due to the higher energy consumption in either the retention or dynamic access operations. To solve this issue, we propose three techniques to achieve a 16Kb energy-efficient CSDB-eDRAM for cryogenic memory implementation. First, we propose a 4T CSDB-GC that is able to significantly improve the retention time. Second, we propose a wordline voltage off-chip tuning method to enhance the dual-port read speed and read-disturb free operations. Third, we introduce a bitline split scheme to reduce the dynamic power overhead of each access operation. Measurement results from our fabricated chip show that the dynamic power of our CSDB-eDRAM has been reduced to 49.23 uW/Kb at 1.41 GHz, which outperforms the state-of-the-art by$\mathbf{11.4}\times$. It also achieves the best data retention time of 16.67 s at 4.2 K. Moreover, a negligible retention power of 0.11 pW/Kb can be achieved. Yuhao Shu, Hongtu Zhang, Hao Sun 0035, Yajun Ha |
ISCAS | 3 |
| 2023 | RPS-KNN: An Ultra-Fast FPGA Accelerator of Range-Projection-Structure K-Nearest-Neighbor Search for LiDAR Odometry in Smart VehiclesabstractKNN (K-Nearest-Neighbors) search has been widely used in LiDAR-related applications. As LiDAR's point clouds become more massive, it is a great challenge to implement a fast and energy-efficient KNN implementation. Previous works consume much time in either building an efficient data structure or searching in the data structure. To solve this issue, we propose a high-locality data structure RPS (range-projection-structure) and an ultra-fast FPGA accelerator of the building and searching process. First, we propose a novel data structure RPS which ensures the points with similar projection locations and range scales are stored in the continuous locations of a memory. Second, we propose a highly-parallel method to build the RPS by projecting the points into a point cloud matrix and parallelly processing the points in a column. Third, based on RPS, we propose a highly-parallel KNN search algorithm, which can quickly narrow the search region and select the KNN from neighboring points in parallel. Experimental results show that our method achieves 13.7 times faster than other FPGA implementations. Moreover, energy efficiency results show that our proposed method is 27.4 times and 50.7 times higher than the state-of-the-art implementations on FPGA and GPU platforms, respectively. Jianzhong Xiao, Hao Sun 0035, Hongtu Zhang, Chengzhang He, Yuhao Shu, Yajun Ha |
ISCAS | 2 |
| 2023 | An Energy-Efficient Stream-Based FPGA Implementation of Feature Extraction Algorithm for LiDAR Point Clouds With Effective Local-SearchabstractFeature extraction is a fundamental and essential step in light detection and ranging (LiDAR) based simultaneously localization and mapping (SLAM) algorithms. Considering the run-time requirement of feature extraction and the stringent battery constraint in smart vehicles, it is a great challenge to develop fast and highly energy-efficient feature extraction implementation for massive point clouds. Unfortunately, existing implementations not only fail to exploit the available parallelism but also fail to make full use of the local information to optimize the computations. To solve the issue, we propose three novel techniques to achieve a fast and energy-efficient FPGA implementation of the feature extraction algorithm with effective local search. First, we propose a low-complexity projection method and a column-scanning scheduler to organize the irregular and sparse point cloud into a well-organized point cloud matrix. Second, based on the point cloud matrix, we exploit its local information and propose a high-parallel method to detect the coarse-grain feature points. Third, we propose a high-parallel conditional priority queue to progressively and evenly select the fine-grain feature points. Experimental results on the KITTI dataset show that our method implemented on the ZCU104 FPGA board achieves the best accuracy and reaches 584 frames per second (FPS) for the feature extraction of a 64-laser LiDAR’s point cloud. Moreover, our proposal achieves the best energy efficiency, which is on average 11.7 times and 9.0 times higher than the state-of-the-art implementations on the GPU and FPGA platforms, respectively. Hao Sun 0035, Yuhao Shu, Yajun Ha |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | WDVR-RAM: A 0.25-1.2 V, 2.6-76 POPS/W Charge-Domain In-Memory-Computing Binarized CNN Accelerator for Dynamic AIoT WorkloadsabstractIn-memory computing (IMC) is an effective approach to accelerate the interference tasks of binary neural networks (BNN), which has been widely used in artificial intelligence of things (AIoT) applications. However, previous researches only focus on optimizing energy efficiency for a narrow voltage range. This poses a severe limitation for some AIoT applications, because they may have very dynamic workloads that require the energy efficiency optimization of BNN for a wide dynamic voltage range (WDVR). To address this issue, we have developed a novel IMC-based BNN accelerator, supporting energy-efficient operations in a wide dynamic voltage range. First, we analyze different charge-domain architectures in terms of their errors and energy characteristics, and decide the best architecture that is suitable for a wide dynamic voltage range. Second, we introduce a lazy convolution bitline reset (LCBR) scheme to further optimize the energy of multiplication and accumulation (MAC) within the entire voltage range. Third, we design a subthreshold differential batch-normalization amplifier (SDBNA) array to compensate for the inference accuracy loss for the lower end of the voltage range. A 16-Kb WDVR-RAM has been designed and fabricated in a 55-nm CMOS process. Measurement results show that the test chip achieves a peak energy efficiency of 76193-2645 TOPS/W at 0.25-1.2 V for MAC, and a peak energy efficiency of 56141-1973 TOPS/W at 0.25-1.2V for both MAC and batch normalization (BN) layers. For throughput, our work obtains a peak computing density range of 692-144078 GOPS/mm2 with 28.31-0.17 ms/frame (CIFAR-10). Moreover, it also achieves the highest 14.75% and 31.87% recovery for Top-1 accuracy observed for MNIST and CIFAR-10 so far, respectively. Hongtu Zhang, Yuhao Shu, Hao Sun 0035, Yajun Ha |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | TAIT: One-Shot Full-Integer Lightweight DNN Quantization via Tunable Activation Imbalance TransferabstractBoth parameter quantization and depthwise convolution are essential measures to provide high-accuracy, lightweight, and resource-friendly solutions when deploying deep neural networks (DNNs) onto edge-AI devices. However, combining the two methodologies may lead to adverse effects: It either suffers from significant accuracy loss or long finetuning time. Besides, contemporary quantization methods are only selectively applied to weight and activation values but not bias and scaling factor values, making them less practical for ASIC/FPGA accelerators. To solve these issues, we propose a novel quantization framework that is effectively optimized for depthwise convolution networks. We discover that the uniformity of the value range within a tensor can serve as a predictor for the tensor’s quantization error. Under the guidance of this predictor, we develop a mechanism called Tunable Activation Imbalance Transfer (TAIT), which tunes the value range uniformity between an activated feature map and its latter weights. Moreover, TAIT fully supports full-integer quantization. We demonstrate TAIT on SkyNet and deploy it on FPGA. Compared to the state-of-the-art, our quantization framework and system design achieve 2.2%+ IoU, $2.4 \times$ speed, and $1.8 \times$ energy efficiency improvements, without any requirement of finetuning. Weixiong Jiang, Heng Yu 0001, Hao Sun 0035, Rui Li 0095, Yajun Ha |
DAC | 4 |