EDBT 2026 Demo / reviewers in the wild / expert
Hanwei Fan
dblp:287/4851
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-1177-2108ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FPPS: An FPGA-Based Point Cloud Processing SystemabstractPoint cloud processing is a computational bottleneck in autonomous driving systems, especially for real-time applications, while energy efficiency remains a critical system constraint. This work presents FPPS, an FPGA-accelerated point cloud processing system designed to optimize the iterative closest point (ICP) algorithm, a classic cornerstone of 3D localization and perception pipelines. Evaluated on the widely used KITTI benchmark dataset, the proposed system achieves up to 35×(and an runtime-weighted average of 15.95×) speedup over a state-of-the-art CPU baseline while maintaining equivalent registration accuracy. Notably, the design improves average power efficiency by 8.58×, offering a compelling balance between performance and energy consumption. These results position FPPS as a viable solution for resource-constrained embedded autonomous platforms where both latency and power are key design priorities. Linfeng Du, Hanwei Fan, Wei Zhang 0012 |
ISCAS | 3 |
| 2025 | Invited Paper: CURE-Fuzz: Curiosity-Driven Reinforcement Learning for Agile Hardware TestingabstractModern processors feature complex architectures that necessitate the generation of extensive test programs to ensure functional correctness, making testing the most time-consuming stage of the processor design flow. Existing automated verification frameworks for agile design exhibit significant limitations, such as fixed program structures restricting flexibility, uncontrolled control flows leading to invalid instructions, and low coverage of the vast state space. To address these limitations, we propose CURE-Fuzz, a curiosity-driven reinforcement learning framework designed to enhance agile hardware testing. By integrating a hierarchical test generation model with a curiosity-driven exploration mechanism, CURE-Fuzz enables precise control over test program structure and dependencies while efficiently navigating unexplored processor states. Evaluations on Rocket and Boom core demonstrate that CURE-Fuzz achieves higher coverage and exhibits superior bug detection capabilities compared to state-of-the-art fuzzers. Hanwei Fan, Binguang Zhao, Yangdi Lyu, Jiang Xu 0001, Wei Zhang 0001 |
ICCAD | 1 |
| 2025 | ScanNow: A Scan Window-Based Sparse Matrix Multiplication Accelerator DesignabstractSparse matrix-matrix multiplication (SpMM) is a prevailing kernel in scientific and artificial intelligence applications. However, the irregular memory access behaviors caused by diverse sparse patterns in SpMM lead to a significant performance bottleneck for traditional computing platforms, driving the demand for dedicated hardware accelerators. Unfortunately, the existing hardware accelerators are not flexible enough to handle diverse SpMM workloads due to their rigid task dispatch schemes, limiting the overall performance.To address the challenge, this paper introduces ScanNow, a scan window-based SpMM accelerator design. ScanNow first presents a scan window-based task dispatch scheme designed to minimize the time-consuming irregular memory access behaviors by exploiting potential data reuse opportunities. The proposed task dispatch scheme effectively and dynamically balances input and output data reuse for row-wise dataflow according to real-time sparse patterns. Moreover, a tailored hardware architecture, featuring multiple memory components and customized schedulers, is proposed to accommodate the task dispatch scheme and improve computation parallelism. Experiments demonstrate that ScanNow achieves an average speedup of 1.67x across a wide range of SpMM workloads compared to the state-of-the-art hardware accelerator. Chaofang Ma, Hanwei Fan |
ICCAD | 4 |
| 2025 | FLEX: Leveraging FPGA-CPU Synergy for Mixed-Cell-Height Legalization AccelerationabstractLegalization is a critical yet time-consuming step in very large-scale integration (VLSI) design, tasked with iteratively relocating standard cells to eliminate overlaps while resolving design rule violations. This process is repeatedly invoked during VLSI physical design. However, increasing spatial constraints and complex design rules impose significant challenges on existing CPU- and GPU-based legalizers, including suboptimal task assignment, inefficient algorithm, and long hardware idle time caused by processing tasks with irregular computational patterns in parallel. Linfeng Du, Yipu Zhang 0002, Chaofang Ma, Hanwei Fan, Jiang Xu 0001, Wei Zhang 0012 |
ICPP | 6 |
| 2024 | Explainable Fuzzy Neural Network with Multi-Fidelity Reinforcement Learning for Micro-Architecture Design Space ExplorationabstractWith the continuous advancement of processors, modern micro-architecture designs have become increasingly complex. The vast design space presents significant challenges for human designers, making design space exploration (DSE) algorithms a significant tool for μ-arch design. In recent years, efforts have been made in the development of DSE algorithms, and promising results have been achieved. However, the existing DSE algorithms, e.g., Bayesian Optimization and ensemble learning, suffer from poor interpretability, hindering designers' understanding of the decision-making process. To address this limitation, we propose utilizing Fuzzy Neural Networks to induce and summarize knowledge and insights from the DSE process, enhancing interpretability and controllability. Furthermore, to improve efficiency, we introduce a multi-fidelity reinforcement learning approach, which primarily conducts exploration using cheap but less precise data, thereby substantially diminishing the reliance on costly data. Experimental results show that our method achieves excellent results with a very limited sample budget and successfully surpasses the current state-of-the-art. Our DSE framework is open-sourced and available at https://github.com/fanhanwei/FNN_MFRL_ArchDSE/. Hanwei Fan, Sicheng Li 0001, Tingyuan Liang, Wei Zhang 0012 |
DAC | 1 |
| 2024 | A Modular Branch Predictor Performance Analysis Framework for Fast Design Space ExplorationabstractAs modern processor designs scale up and workloads become more complex, the selection of the branch predictor (BP) and the optimization of its internal parameters are increasingly critical in striking a balance between performance and resource usage. However, current fast performance evaluation models and micro-architectural Design Space Exploration (DSE) frameworks provide limited support for BP components, especially regarding internal parameter adjustments. In this work, we propose a modular BP performance analysis framework that provides fast performance feedback for different BP configurations. Our framework includes a pattern analyzer equipped with more accurate metrics for quantifying an application's predictability of branch behavior, a classification module for selecting the appropriate BP type, and an analytical model set that reflects the impact of internal parameter adjustments of various BPs on both performance and storage resource usage, thereby supporting DSE. Experimental results on three benchmarks confirm the framework's effectiveness, as our proposed model exhibits better correlation while reflecting more parameter changes than previous work. To the best of our knowledge, this is the first analytical model framework that supports comprehensive BP type and parameter adjustments. Hanwei Fan, Sicheng Li 0001, Tingyuan Liang, Wei Zhang 0012 |
DATE | 2 |
| 2022 | Bayesian Optimization with Clustering and Rollback for CNN Auto Pruning
Hanwei Fan, Jiandong Mu, Wei Zhang 0012 |
ECCV (23) | 1 |
| 2022 | A Contrastive-Learning-Based Method for Alert-Scene CategorizationabstractWhether it’s a driver warning or an autonomous driving system, ADAS needs to decide when to alert the driver of danger or take over control. This research formulates the problem as an alert-scene categorization one and proposes a method using contrastive learning. Given a front-view video of a driving scene, a set of anchor points is marked by a human driver, where an anchor point indicates that the semantic attribute of the current scene is different from that of the previous one. The anchor frames are then used to generate contrastive image pairs to train a feature encoder and obtain a scene similarity measure, so as to expand the distance of the scenes of different categories in the feature space. Each scene category is explicitly modeled to capture the meta pattern on the distribution of scene similarity values, which is then used to infer scene categories. Experiments are conducted using front-view videos that were collected during driving at a cluttered dynamic campus. The scenes are categorized into no alert, longitudinal alert, and lateral alert. The results are studied at both feature encoding, category modeling, and reasoning aspects. By comparing precision with two full supervised end-to-end baseline models, the proposed method demonstrates competitive or superior performance. However, it remains still questions: how to generate ground truth data and how to evaluate performance in ambiguous situations, which leads to future works. Shaochi Hu, Hanwei Fan, Biao Gao, Huijing Zhao |
IV | 2 |