Hui Wang 0036

dblp:39/721-36 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 7 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 A Complementary 3T-Based eDRAM Macro for High-Density Dual-Direction CAM and Logic-in-Memory
abstract
Content-addressable memory (CAM) is regarded as an attractive solution for data-intensive applications with high-density search demands. To further improve functional flexibility yet at a low cost, several CAM macros have been developed to support multiple bit-wise logic operations. However, conventional SRAM-based CAM designs are constrained by the large bitcell area, posing significant challenges to achieve higher density. To address this issue, we propose a complementary 3T (C3T) based embedded dynamic random access memory (eDRAM) macro for high-density dual-direction CAM searching and logic-in-memory operations. First, we propose a compact C3T bitcell featuring a pair of complementary decoupled read ports, enabling dual-port read and efficient CAM operations. Second, we present a compact dynamic-circuit-based sense amplifier (DSA) to optimize the area of readout peripheral circuitry while mitigating the read bit line saturation issue. Additionally, we implement dual-direction CAM searching and logic-in-memory operations exploiting the C3T-based eDRAM macro. A 4 Kb C3T-based eDRAM macro has been validated in a commercial 40-nm CMOS process. Post-layout results demonstrate a 53% reduction in the bitcell area and a 58.1% reduction in the macro area compared to the state-of-the-art 6T compute SRAM. Moreover, the proposed design achieves a maximum frequency of 578 MHz for binary CAM (BCAM) searching operations and 694 MHz for logic operations, with energy consumption of 1.12 fJ/bit and 26.8 fJ/bit, respectively.
Lintao Lan, Yuhao Shu, Hui Wang 0036, Yajun Ha
IEEE Trans. Circuits Syst. I Regul. Pap.7
2026 A 50 μW/Gbps/Lane Power-Efficient MIPI D-PHY Receiver With Architecture-Level Adaptive and Structural Optimizations for Micro-Displays
abstract
Achieving high power efficiency in Mobile Industry Processor Interface (MIPI) D-PHY receivers is crucial for micro-display chips in AR/VR systems, where stringent power constraints exist. However, existing designs often sacrifice power efficiency for higher data rates due to architectural limitations, neglecting optimization for low-power applications. To address this issue, we propose a receiver architecture that substantially enhances power efficiency through three key techniques. First, we improve the gain-bandwidth product (GBW) by employing an autonomous gain scheduling analog front-end (AFE) that dynamically tunes the gain while reducing drive current. Second, we reduce clocking overhead by introducing a self-monitoring interferometric deserializer that enables clock-free pre-scaling and halves the DDR sampling frequency. Third, we increase transition speed and minimize short-circuit power by utilizing a chaotic topological flow actuator (CTFA) with multi-path current feedthrough. Compared to prior state-of-the-art designs, the proposed receiver achieves a power efficiency of$50~\mu $W/Gbps/lane ($42~\mu $A/Gbps/lane), reducing power and current consumption by 46% and 45%, respectively, using a standard 180-nm process.
Haoran Zeng, Yingqi Feng, Tianai Li, Hang Ye 0007, Zunkai Huang, Hui Wang 0036, Yongxin Zhu 0001, Qiliang Li, Yajun Ha
IEEE Trans. Circuits Syst. I Regul. Pap.6
2025 Accelerating Elliptic Curve Digital Signature Verification on FPGA for Secure Communication
Chujun Feng, Ning Ni 0001, Congyu Lin, Yongxin Zhu 0001, Hui Wang 0036
WASA (1)6
2025 Fast FPGA Accelerator of Graph Cut Algorithm With Threshold Global Relabel and Inertial Push
abstract
Graph cut algorithms are popular in optimization tasks related to min-cut and max-flow problems. However, modern FPGA graph cut algorithm accelerators still need performance and memory resource utilization optimization. On the one hand, they suffer from redundant computations in the heuristic global relabel algorithm and slow convergence speeds during the pushing operation. On the other hand, they can only handle 8-bit 2-D grid graphs with limited size. To address the challenges, first, we propose a novel threshold global relabel algorithm that divides the graph into sleeping and active regions, significantly reducing redundant computations in the sleeping region. Second, we introduce an inertial push technique that imparts flow inertia to break flow barriers and accelerate the algorithm’s convergence. Third, to fully utilize the memory resource in FPGA, we propose an efficient memory layout that divides the memory into read-write and read-only regions. Compared to the state-of-the-art, our FPGA accelerator can efficiently handle 16-bit 2-D grid graphs with 2 million nodes and achieve up to a$2.49\times $improvement in execution time with the same memory usage.
Guangyao Yan, Hui Wang 0036, Yajun Ha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2025 An Energy-Efficient and Real-Time FPGA-Based Point Cloud Registration Framework with Ultra-Fast and Configurable Multi-Mode Correspondence Search
abstract
Point cloud registration is a fundamental task in LiDAR-based localization and mapping, widely employed in robotics and autonomous vehicles. However, existing registration solutions lose geometric topology continuity and lack scanline-aware, configurable correspondence search, restricting their real-time applicability. To solve these issues, we propose a fundamentally re-architected, energy-efficient FPGA framework for real-time point cloud registration, featuring a configurable, multi-mode correspondence search engine. First, we introduce a scanline-aided range-projection structure (SA-RPS) that reorganizes LiDAR points within configurable segmentation domains into contiguous memory while preserving scanline topology, enabling efficient and flexible multi-mode correspondence search. Second, we develop a deeply pipelined, ultra-fast SA-RPS-based correspondence search (SA-RPS-CS) accelerator that supports dynamic configuration of search mode and parallelism and incorporates a sliding-window cache and scanline-aware K-selection module for high-throughput, multi-mode correspondence extraction. Third, we present a co-designed registration framework that integrates the accelerator with dynamic parameter configuration, enabling adaptive, real-time processing across diverse SLAM scenarios. Experimental results demonstrate that the proposed SA-RPS-CS accelerator delivers \(2.3\times\) – \(32.4\times\) faster search and \(1.8\times\) – \(26.2\times\) higher energy efficiency than previous state-of-the-art FPGA designs, achieving real-time registration for 64-channel LiDAR at 21.5 FPS with negligible loss in accuracy.
Hao Sun 0035, Yuhao Shu, Jianzhong Xiao, Weixiong Jiang, Hui Wang 0036, Yajun Ha
ACM Trans. Reconfigurable Technol. Syst.6
2023 Fast FPGA Accelerator of Graph Cut Algorithm with Out-of-order Parallel Execution in Folding Grid Architecture
abstract
Graph cut is a popular approach to solving optimization tasks related to Min-cut/Max-flow problems. However, existing FPGA accelerators of graph cut have difficulty in handling large grid graphs and achieving real-time performance. To address the issue, we propose a novel folding grid architecture that maps an actual one-layered large 2-dimension grid graph into a virtual multi-layered small 2-dimension grid graph. The new architecture not only enables the virtual multi-layered grid graph to execute on a small-size processor array but also adds the potential to concurrently execute grid graph nodes in different layers. In addition, we also propose a novel out-of-order parallel execution technique to fully utilize the architecture parallelism potential. Compared to the state-of-the-art, experimental results show that our design can solve the graph cut problem for grid graphs of 1920 × 1080 nodes in real-time (above 60fps) and achieve a 5.4× improvement in execution time with similar FPGA resources.
Guangyao Yan, Hui Wang 0036, Yajun Ha
DAC3
2022 Ultra-Fast FPGA Implementation of Graph Cut Algorithm With Ripple Push and Early Termination
abstract
Graph cut has been a popular approach widely used to solve the minimum cut problem, which is prevalent in computer vision tasks, although not limited to this field. Push-relabel is considered as one of the promising algorithms of graph cut due to its good potential to be parallelized. However, existing implementations often not only fail to fully exploit the available parallelism but also fail to make full use of the application context to reduce redundant computations. Therefore, they are not competent for application scenarios with high resolution and real-time requirements. To address the issue, we propose three novel techniques to achieve an ultra-fast and efficient FPGA implementation of a push-relabel algorithm. First, we propose a ripple push technique that significantly parallelizes push operations so as to accelerate the push-relabel convergence process. Second, we propose an early-termination technique that effectively removes redundant computations of the push-relabel algorithm. We also theoretically prove the correctness of our early-termination technique. Third, we propose a highly parallelized search technique called flood irrigation search (FIS). It quickly judges early termination conditions based on a pixel parallel architecture. Our implementation focuses on performance-sensitive applications that divide large images into small graph tiles with a specific size. Compared to the state-of-the-art FPGA implementations of push-relabel algorithms, experimental results show that our method can at least achieve$8.93\times $improvement of execution time.
Guangyao Yan, Fupeng Chen, Hui Wang 0036, Yajun Ha
IEEE Trans. Circuits Syst. I Regul. Pap.4
2021 SparkNoC: An energy-efficiency FPGA-based accelerator using optimized lightweight CNN for edge computing
Zunkai Huang, Hui Wang 0036, Victor Chang 0001, Yongxin Zhu 0001, Songlin Feng
J. Syst. Archit.4
2020 Anomaly Detection Based on RBM-LSTM Neural Network for CPS in Advanced Driver Assistance System
abstract
Advanced Driver Assistance System (ADAS) is a typical Cyber Physical System (CPS) application for human–computer interaction. In the process of vehicle driving, we use the information from CPS on ADAS to not only help us understand the driving condition of the car but also help us change the driving strategies to drive in a better and safer way. After getting the information, the driver can evaluate the feedback information of the vehicle, so as to enhance the ability to assist in driving of the ADAS system. This completes a complete human–computer interaction process. However, the data obtained during the interaction usually form a large dimension, and irrelevant features sometimes hide the occurrence of anomalies, which poses a significant challenge to us to better understand the driving states of the car. To solve this problem, we propose an anomaly detection framework based on RBM-LSTM. In this hybrid framework, RBM is trained to extract general underlying features from data collected by CPS, and LSTM is trained from the features learned by RBM. This framework can effectively improve the prediction speed and present a good prediction accuracy to show vehicle driving condition. Besides, drivers are allowed to evaluate the prediction results, so as to improve the accuracy of prediction. Through the experimental results, we can find that the proposed framework not only simplifies the training of the entire neural network and increases the training speed but also greatly improves the accuracy of the interaction-driven data analysis. It is a valid method to analyze the data generated during the human interaction.
Hanlin Zhu, Yongxin Zhu 0001, Victor Chang 0001, Cong He, Ching-Hsien Hsu, Hui Wang 0036, Songlin Feng, Zunkai Huang
ACM Trans. Cyber Phys. Syst.7
2019 Designing efficient accelerator of depthwise separable convolutional neural network on FPGA
Zunkai Huang, Hui Wang 0036, Songlin Feng
J. Syst. Archit.5