EDBT 2026 Demo / reviewers in the wild / expert
Guangyao Yan
dblp:317/6007
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0002-8256-7435ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LibSCAT: Library-Based Formal Verification of Heavily Optimized Multipliers via GNN-Guided Reference SelectionabstractFormal verification of heavily optimized multipliers is a critical yet challenging problem in both industry and academia. Current approaches suffer from fundamental limitations: Symbolic Computer Algebra (SCA) techniques struggle with heavily optimized multipliers, Satisfiability (SAT)-based approaches require structurally similar reference designs, and hybrid methods fail to handle Booth multipliers. On the other hand, industrial design flows possess extensive libraries of verified multipliers for optimization workflows, creating an underutilized opportunity for library-based verification. Yet optimal reference selection becomes challenging due to large-scale libraries and optimization-obscured architectural relationships. To address these challenges, we propose LibSCAT, a verification framework that leverages large-scale reference libraries in a scalable manner. First, we propose a reference library-based methodology that adaptively combines SCA and SAT techniques through intelligent reference selection and predictive method choice. Second, we propose a Siamese Graph Neural Network model that captures multiplier structural relationships in latent space from reverse-engineered graphs, generating robust embeddings for efficient reference selection. Third, we propose a Random Forest-based predictor that leverages learned embeddings for accurate selection of verification strategies. Experimental results show our method achieves 88.2% success on heavily optimized simple partial product multipliers and 94.0% success on heavily optimized Booth multipliers, significantly outperforming state-of-the-art methods. Rui Li 0095, Masahiro Fujita 0004, Heng Yu 0001, Guangyao Yan, Lin Li 0079, Yajun Ha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | Event-Driven Asynchronous Graph Neural Network FPGA Accelerator for Real-Time Edge VisionabstractEvent-based asynchronous graph neural networks (GNNs) provide a promising solution for real-time edge vision. By leveraging microsecond-level input latency, asynchronous computation, and sparse storage, they show significant potential for low-latency processing under resource constraints. However, existing FPGA-based accelerators for event-driven asynchronous GNNs cannot meet real-time performance owing to critical bottlenecks in memory utilization, parallelism, and computational redundancy. To address these challenges, we propose a novel FPGA accelerator for event-driven asynchronous GNNs, with three key contributions: 1) Memory-efficient graph feature storage with improved readout parallelism to reduce data access time, 2) Parallelism-enhanced hierarchical graph construction with low dependency to reduce computation time, and 3) Redundancy-free parallel graph convolution with reusable partial computation caching to reduce computation time. The proposed accelerator was deployed on a Xilinx ZCU102 MPSoC platform and evaluated on the N-CARS dataset for car recognition. Compared to the state-of-the-art (SOTA), our proposed design achieves an average$27.59\times $speedup with a latency of$0.58\mu $s while delivering higher accuracy and comparable resource consumption. Tianhang Liu, Guangyao Yan, Runhua Wang, Rui Li 0095, Shijie Meng, Hao Sun 0035, Yajun Ha |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | Fast FPGA Accelerator of Graph Cut Algorithm With Threshold Global Relabel and Inertial PushabstractGraph cut algorithms are popular in optimization tasks related to min-cut and max-flow problems. However, modern FPGA graph cut algorithm accelerators still need performance and memory resource utilization optimization. On the one hand, they suffer from redundant computations in the heuristic global relabel algorithm and slow convergence speeds during the pushing operation. On the other hand, they can only handle 8-bit 2-D grid graphs with limited size. To address the challenges, first, we propose a novel threshold global relabel algorithm that divides the graph into sleeping and active regions, significantly reducing redundant computations in the sleeping region. Second, we introduce an inertial push technique that imparts flow inertia to break flow barriers and accelerate the algorithm’s convergence. Third, to fully utilize the memory resource in FPGA, we propose an efficient memory layout that divides the memory into read-write and read-only regions. Compared to the state-of-the-art, our FPGA accelerator can efficiently handle 16-bit 2-D grid graphs with 2 million nodes and achieve up to a$2.49\times $improvement in execution time with the same memory usage. Guangyao Yan, Hui Wang 0036, Yajun Ha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | QuantTPM: Efficient Mixed-Precision Quantization Framework for Tractable Probabilistic ModelsabstractTractable probabilistic models (TPMs) can perform reliable probabilistic inference and enhance the reasoning capabilities of edge devices, such as aiding decision-making for autonomous vehicles. To deploy TPMs in edge scenarios with constrained hardware resources and energy, efficient quantization algorithms are necessary. However, the traditional quantization methods for neural networks are not applicable to TPMs due to the irregular model structure and highly varying data distribution. To address the issues, we propose QuantTPM, a mixed-precision quantization framework designed to enhance the energy and resource efficiency of TPM inference. First, we reformulate the irregular model structure into a unified format, as irregular structures are inefficient for hardware implementation. Second, we divide the reformulated model graph into hierarchical levels, so as to assign appropriate quantization bit-widths for different levels with varying precision requirements. Third, we decompose the entire mixed-precision quantization search into several steps with smaller search spaces, so as to reduce the algorithm complexity and save search time. Compared with state-of-the-art works, our mixed-precision quantization framework achieves, on average,$3.7\times $weight compression,$6.0\times $resource efficiency, and$4.8\times $energy consumption, while maintaining competitive accuracy. Guangyao Yan, Weixiong Jiang, Yajun Ha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Fast FPGA Accelerator of Graph Cut Algorithm with Out-of-order Parallel Execution in Folding Grid ArchitectureabstractGraph cut is a popular approach to solving optimization tasks related to Min-cut/Max-flow problems. However, existing FPGA accelerators of graph cut have difficulty in handling large grid graphs and achieving real-time performance. To address the issue, we propose a novel folding grid architecture that maps an actual one-layered large 2-dimension grid graph into a virtual multi-layered small 2-dimension grid graph. The new architecture not only enables the virtual multi-layered grid graph to execute on a small-size processor array but also adds the potential to concurrently execute grid graph nodes in different layers. In addition, we also propose a novel out-of-order parallel execution technique to fully utilize the architecture parallelism potential. Compared to the state-of-the-art, experimental results show that our design can solve the graph cut problem for grid graphs of 1920 × 1080 nodes in real-time (above 60fps) and achieve a 5.4× improvement in execution time with similar FPGA resources. Guangyao Yan, Hui Wang 0036, Yajun Ha |
DAC | 1 |
| 2022 | Ultra-Fast FPGA Implementation of Graph Cut Algorithm With Ripple Push and Early TerminationabstractGraph cut has been a popular approach widely used to solve the minimum cut problem, which is prevalent in computer vision tasks, although not limited to this field. Push-relabel is considered as one of the promising algorithms of graph cut due to its good potential to be parallelized. However, existing implementations often not only fail to fully exploit the available parallelism but also fail to make full use of the application context to reduce redundant computations. Therefore, they are not competent for application scenarios with high resolution and real-time requirements. To address the issue, we propose three novel techniques to achieve an ultra-fast and efficient FPGA implementation of a push-relabel algorithm. First, we propose a ripple push technique that significantly parallelizes push operations so as to accelerate the push-relabel convergence process. Second, we propose an early-termination technique that effectively removes redundant computations of the push-relabel algorithm. We also theoretically prove the correctness of our early-termination technique. Third, we propose a highly parallelized search technique called flood irrigation search (FIS). It quickly judges early termination conditions based on a pixel parallel architecture. Our implementation focuses on performance-sensitive applications that divide large images into small graph tiles with a specific size. Compared to the state-of-the-art FPGA implementations of push-relabel algorithms, experimental results show that our method can at least achieve$8.93\times $improvement of execution time. Guangyao Yan, Fupeng Chen, Hui Wang 0036, Yajun Ha |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |