VLDB 2026 Research / reviewers in the wild / expert
Zhanhong Huang
dblp:339/7433
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0003-1103-1104ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HSCIM: A High Security Compute-In-Memory Architecture with PUF based on TST-MRAMabstractWith the rapid development of the Internet of Things (IOT) in the decades, compute-in-memory (CIM) architecture which addresses the Von-Neumann bottleneck are drawing significant attention with a gradually improving demand of high security. Physical unclonable function (PUF) emerges as a promising candidate attributed to its dependence on physical properties of devices instead of traditional key systems. Toggle spin torques magnetoresistive random access memory (TST-MRAM) is considered as satisfying storage technology due to its non-volatility and low power consumption. In this paper, a high security compute-in-memory (HSCIM) architecture based on TST-MRAM is proposed, aiming to generating PUF signals for encryption and guaranteeing security during computation. Simulation results show that with the proposed architecture, a 16-bit multiply-and-accumulate (MAC) operation along with data encryption and decryption can be performed within four clock cycles, achieving an efficient utilization of device resource. Junyi Mai, Feilan Zhao, Zhanhong Huang, Yongkui Yang, Enyi Yao |
ISCAS | 3 |
| 2025 | Parallel Lidar Ground Segmentation: From Mechanical to Solid-State SensorsabstractReal-time and efficient LiDAR ground segmentation is crucial for autonomous perception and edge computing applications. This paper introduces an FPGA-based parallel segmentation framework that significantly reduces computational latency while maintaining high accuracy. Through a stream-pipelined design, the proposed method achieves significant speedup over CPU-based implementations while benefiting from the denser and more evenly distributed point clouds of Solid-State LiDAR (SSL). Extensive evaluations on the SemanticKITTI dataset and a self-collected SSL dataset demonstrate the effectiveness of our approach across both mechanical and SSL systems, highlighting its potential for real-world deployment in autonomous vehicles and edge-based perception platforms. Zhanhong Huang, Antony García, Witek Jachimczyk, Xinming Huang 0001 |
VTC2025-Spring | 2 |
| 2025 | Stream-Based LiDAR Point Cloud Ground Segmentation for Autonomous VehicleabstractThis paper introduces a hardware-friendly refined depth ground segmentation method and its FPGA deployment, designed for real-time applications with stringent power and energy constraints. Our approach outperforms traditional CPU solutions, achieving over 10x lower power consumption and orders of magnitude reduction in energy usage per frame. Leveraging a one-pass stream-based architecture, the method supports various LiDAR configurations with low latency and consistent performance. Experimental results demonstrate superior segmentation accuracy in both point-wise and area-wise evaluations, establishing our approach as a practical and efficient solution for autonomous vehicle applications. Zhanhong Huang, Antony García, Witek Jachimczyk, Xinming Huang 0001 |
VTC2025-Spring | 2 |
| 2025 | An Efficient L-Shape Based Object Detection Using Density Distribution in LiDAR Point Cloud for Autonomous DrivingabstractThis paper presents a novel and efficient L-shape detection framework designed for autonomous driving applications. We propose a fast geometry-based L-shape fitting algorithm that ensures high computational efficiency and consistent execution across diverse data sources, while preserving accurate shape alignment. Furthermore, a cluster-wise, octree-based Bird's Eye View (BEV) classifier is introduced to determine both oriented and axis-aligned bounding box alignments. Extensive experiments validate the proposed method's effectiveness and accuracy in pose estimation, object detection, and speed estimation. The system achieves an object velocity estimation error within 0.5 meters per second. Real-world evaluations confirm the robustness and applicability of the approach in practical autonomous driving scenarios. Zhanhong Huang, Xinming Huang 0001 |
VTC2025-Spring | 2 |
| 2025 | A Spin Scale-Aware Self-Adaptive Ising Annealing Processing Architecture for Combinatorial Optimization ProblemsabstractThe Ising annealing processor has emerged as a promising approach to accelerate the discovery of the optimal solutions for a wide range of combinatorial optimization problems (COPs), by mapping various COPs into a unified Ising model. However, fixed computational strategies and inflexible architectures make previous designs suffer from a low hardware resource utilization rate when the numbers of the total required and real-time flipped spins vary across different COPs and iteration steps. In this paper, a novel spin scale-aware self-adaptive Ising annealing processing architecture (AIAPA) is proposed to address this problem, with an adaptive computational strategy, a custom instruction set, multi-traffic mode routers, and a fully-pipelined computing array. It can dynamically adapt to the varying scenarios during the Ising annealing process to maximize the performance of limited hardware resources. Its prototype, supporting 65k fully-connected spins, is implemented on an FPGA platform, operates at a clock frequency of 188 MHz. The AIAPA achieves up to a 24.22 times faster annealing speed compared to the state-of-the-art FPGA design on the max-cut optimization problem while maintaining a high convergence accuracy. Dong Jiang 0002, Xiangrui Wang, Zhanhong Huang, Longyuan Kang, Simei Yang, Enyi Yao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | A Parallel Tempering Processing Architecture with Multi-Spin Update for Fully-Connected Ising ModelsabstractCombinatorial optimization problems (COPs) are notoriously difficult to solve for classic Von-Neumann computers, which are ubiquitous in various domains. As a state-of-the-art hardware acceleration scheme for COPs, Ising machines are one of the promising research directions for the next generation of computing, but still suffer from the low solution accuracy and speed due to the high complexity of the fully-connected Ising model. In this work, a novel parallel tempering processing architecture (PTPA) is proposed with the modified parallel tempering algorithm, aimed at reducing search time and improving the solution quality. Several techniques are developed to further reduce hardware overhead and enhance parallelism, including the independent pipelined spin update architecture, approximated probability equations, and compact random number generators. Its prototype is implemented on FPGA with eight replicas, each replica containing 1,024 fully-connected spins and at most 64 concurrent update spins. The proposed design achieves an average cut accuracy of 99.43% within 1ms solution time on various G-set problems. Compared with the CPU-based parallel tempering implementation, it enhances the speed of solving the max-cut problems by 5,160 times. Yang Zhang 0120, Xiangrui Wang, Dong Jiang 0002, Zhanhong Huang, Gaopeng Fan, Enyi Yao |
DATE | 4 |
| 2024 | A Novel, Efficient and Accurate Method for Lidar Camera CalibrationabstractAs autonomous systems evolve, the precise calibration of lidar and camera sensors remains a pivotal concern. Among the myriad of available techniques, target-based calibration methods, which employ planar boards with distinct geometry and image patterns, have been a popular choice. These methods simplify the task of extracting corresponding features between the image and lidar point cloud. But many of these approaches also face a significant challenge, which is their sensitivity to lidar resolution and Field of View (FOV), which may degrade the reliability of the calibration results. Therefore, our research introduces a novel calibration method using a uniquely designed acrylic checkerboard which allows the lidar beam to pass through the white grids and reflect back from the black grids. This innovative technique sidesteps the common challenges associated with lidar feature extraction. Our method’s distinct advantage lies in its ability to perform accurate calibrations at close distances, owing to the efficient feature extraction from both lidar and camera sensors. This novel, efficient, and accurate method can provide state-of-the-art results for camera lidar calibration in the field. Please also check our Github repository: https://github.com/WPI-APA-Lab/Acrylic-Board-Lidar-Camera-Calibration Zhanhong Huang, Antony García, Xinming Huang 0001 |
ICRA | 1 |
| 2024 | DCAP: A Scalable Decoupled-Clustering Annealing Processor for Large-Scale Traveling Salesman ProblemsabstractThe Traveling Salesman Problem (TSP) is one of the most well-known NP-hard combinatorial optimization problems (COPs). Many social production problems can be effectively represented as instances of TSPs. However, solving large-scale TSPs remains a significant challenge for conventional Von Neumann computers. Many studies have proposed annealing processors to address large-scale COPs, but most of them focus on unconstrained problems, such as the Maxcut problem. In this paper, a scalable decoupled-clustering annealng processor (DCAP) for efficiently handling large-scale TSPs is presented. A decoupled hierarchical clustering algorithm is proposed for higher convergence speed and improved scalability. Several techniques have been developed in hardware to minimize area overhead and processing time, including a modified spin connection topology for the Ising model, an area-efficient random threshold generator, a one-step spin update scheme and a dynamic prediction method. The DCAP prototype is implemented on FPGA with an operating frequency of 125MHz. We tested our design on various TSP instances from the TSPLIB. Results show that our design outperforms the CPU- and GPU-based Neuro-Ising scheme by achieving maximum speedups of$780\times $and a 42% improvement in accuracy. With multi-chip interconnection, DCAP is able to handle problems of scale up to 85900 cities. Zhanhong Huang, Yang Zhang 0120, Xiangrui Wang, Dong Jiang 0002, Enyi Yao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | An Annealing Processor based on 1k-Spin Fully-Connected Ising Model for Combinatorial Optimization ProblemsabstractCombinatorial optimization problems (COPs) find extensive applications in industrial and social scenarios such as transportation and communication. As the size of NP-hard COPs increases, it becomes impossible to obtain the optimal solution using an enumerative method. Recently, Ising model based annealing processors have received increasing attention due to their potential for rapidly converging to the near-optimal solutions after mapping the problem to them. This paper presents a novel annealing processor (AP) with 1024 fully-connected spins based on a modified Ising model annealing algorithm, which is more suitable for hardware implementation compared to conventional simulated annealing (SA) algorithm. The prototype is implemented using FPGA with the operation frequency up to 100MHz. We tested our design on various G-set problems with an average cut accuracy of 99.19% achieved. The proposed design outperforms the conventional CPU-based method by achieving a max speedup of 2204x for G51. Zhanhong Huang, Xiangrui Wang, Dong Jiang 0002, Yukang Huang, Enyi Yao |
ISCAS | 1 |
| 2023 | A Scalable Annealing Processing Architecture for Fully-Connected Ising ModelsabstractCombinational Optimization Problems (COPs) are prevalent in many different fields. Most of these problems are NP-hard and challenging for computers with conventional Von-Neumann architecture. Ising machines with numerous spins have the potential to solve these problems by emulating the natural annealing process of solid matter. Recent research has explored the hardware implementation of Ising machines to accelerate the convergence process of such problems at room temperature. However, most of them are suffering from low scalability and low parallel processing capability due to the huge hardware cost and high complexity. In this paper, a scalable annealing processing architecture for Ising processor is described to address these issues with a NoC computing paradigm, a distributed storage scheme, and a fully pipelined structure design. The prototype is synthesized using FPGA with the maximum operation frequency of 270MHz, achieving about 32 times faster than conventional simulated annealing method when solving the max-cut problem. Dong Jiang 0002, Xiangrui Wang, Zhanhong Huang, Yukang Huang, Enyi Yao |
ISCAS | 3 |
| 2023 | A Network-on-Chip-Based Annealing Processing Architecture for Large-Scale Fully Connected Ising ModelabstractCombinatorial optimization problems are prevalent in many different fields. Most of these problems are NP-hard and challenging for computers with conventional Von-Neumann architecture. Ising machines with a number of spins have the potential to solve these problems by emulating the natural annealing process of solid matter. Recent research has explored some hardware implementation methods of Ising machines to accelerate the convergence process of such problems at room temperature. However, most of them are suffering from low scalability and low parallel processing capability due to the huge hardware cost and high complexity. In this paper, a novel network-on-chip-based annealing processing architecture (NoCAPA) for a large-scale Ising processor is described to address these issues with a NoC computing paradigm, a distributed storage scheme, and a fully pipelined structure design. Several techniques are developed to further increase convergence speed and reduce hardware resource consumption, including a dynamic multithread parallel update algorithm, a router with merge and deflection abilities, and a unique multiply-accumulate operation. The prototype is implemented in FPGA with the maximum operation frequency of 200MHz, achieving up to$120.5\times $faster than conventional simulated annealing method when solving the max-cut problem while supporting high scalability. Dong Jiang 0002, Xiangrui Wang, Zhanhong Huang, Yongkui Yang, Enyi Yao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |