Rui Li 0095

dblp:96/4282-95 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-2953-9742ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 6 first-author · 11 since 2021
YearPublicationVenuePosition
2026 LibSCAT: Library-Based Formal Verification of Heavily Optimized Multipliers via GNN-Guided Reference Selection
abstract
Formal verification of heavily optimized multipliers is a critical yet challenging problem in both industry and academia. Current approaches suffer from fundamental limitations: Symbolic Computer Algebra (SCA) techniques struggle with heavily optimized multipliers, Satisfiability (SAT)-based approaches require structurally similar reference designs, and hybrid methods fail to handle Booth multipliers. On the other hand, industrial design flows possess extensive libraries of verified multipliers for optimization workflows, creating an underutilized opportunity for library-based verification. Yet optimal reference selection becomes challenging due to large-scale libraries and optimization-obscured architectural relationships. To address these challenges, we propose LibSCAT, a verification framework that leverages large-scale reference libraries in a scalable manner. First, we propose a reference library-based methodology that adaptively combines SCA and SAT techniques through intelligent reference selection and predictive method choice. Second, we propose a Siamese Graph Neural Network model that captures multiplier structural relationships in latent space from reverse-engineered graphs, generating robust embeddings for efficient reference selection. Third, we propose a Random Forest-based predictor that leverages learned embeddings for accurate selection of verification strategies. Experimental results show our method achieves 88.2% success on heavily optimized simple partial product multipliers and 94.0% success on heavily optimized Booth multipliers, significantly outperforming state-of-the-art methods.
Rui Li 0095, Masahiro Fujita 0004, Heng Yu 0001, Guangyao Yan, Lin Li 0079, Yajun Ha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2026 ESACO: Fast E-Graph Extraction via Orchestrated Simulated Annealing-Based Local Search and Ant Colony Optimization-Based Global Search
abstract
Equality graphs (E-graphs) offer a compact representation for vast sets of equivalent implementations, proving invaluable in hardware synthesis and program optimization. Nevertheless, extracting the optimal implementation from an e-graph constitutes an NP-hard challenge. Current extraction methods face critical limitations: heuristic-based approaches fail to produce high-quality solutions, GPU-accelerated techniques lack determinism and demand excessive memory, exact ILP methods struggle with scalability, and specialized solvers only function for particular e-graph types. To address this, we present ESACO, a novel deterministic framework that rapidly and consistently converges to high-quality solutions across diverse benchmarks by effectively combining Simulated Annealing (SA) for local refinement with Ant Colony Optimization (ACO) for global search. First, we develop a synergistic hybrid-heuristic framework that orchestrates complementary search paradigms, harmonizing ACO’s global exploration capabilities with SA’s targeted local exploitation mechanisms. Second, we introduce an SA-based local search method that employs novel rip-up and repair moves for efficiently refining promising solutions. Third, we propose an ACO-based global search algorithm incorporating strategic restart mechanisms to effectively explore the complex solution space while escaping local optima. Experimental results demonstrate that ESACO achieves up to 42× speedup using a single thread compared to state-of-the-art GPU-accelerated methods while maintaining or improving solution quality.
Rui Li 0095, Lin Li 0079, Heng Yu 0001, Yajun Ha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2026 Event-Driven Asynchronous Graph Neural Network FPGA Accelerator for Real-Time Edge Vision
abstract
Event-based asynchronous graph neural networks (GNNs) provide a promising solution for real-time edge vision. By leveraging microsecond-level input latency, asynchronous computation, and sparse storage, they show significant potential for low-latency processing under resource constraints. However, existing FPGA-based accelerators for event-driven asynchronous GNNs cannot meet real-time performance owing to critical bottlenecks in memory utilization, parallelism, and computational redundancy. To address these challenges, we propose a novel FPGA accelerator for event-driven asynchronous GNNs, with three key contributions: 1) Memory-efficient graph feature storage with improved readout parallelism to reduce data access time, 2) Parallelism-enhanced hierarchical graph construction with low dependency to reduce computation time, and 3) Redundancy-free parallel graph convolution with reusable partial computation caching to reduce computation time. The proposed accelerator was deployed on a Xilinx ZCU102 MPSoC platform and evaluated on the N-CARS dataset for car recognition. Compared to the state-of-the-art (SOTA), our proposed design achieves an average$27.59\times $speedup with a latency of$0.58\mu $s while delivering higher accuracy and comparable resource consumption.
Tianhang Liu, Guangyao Yan, Runhua Wang, Rui Li 0095, Shijie Meng, Hao Sun 0035, Yajun Ha
IEEE Trans. Circuits Syst. I Regul. Pap.5
2026 FPGA Routing Congestion Prediction via Graph Learning-Aided Conditional GAN
abstract
Routing congestion prediction expedites the closure of FPGA placement and routing (PnR). Current prediction methods employ convolutional models, taking advantage of their capacity of dealing with image-style inputs. However, these methods neglect the direct representation of circuit netlist and its information fusion with placement scheme. Moreover, the limited size of the convolutional kernel struggles to capture circuit connectivity in distant geometric regions. To address these issues, this article presents a graph-based routing congestion prediction framework that fuses the information contained in the circuit’s topological netlist and geometric placement scheme, and leverages a conditional generative adversarial network (cGAN) model to achieve optimized prediction performance compared to contemporary approaches. Our framework encompasses three key components: (1) the HeteroGraph, a heterogeneous graph that integrates a netlist subgraph and a layout subgraph by space mapping edges; (2) the HeteroGNN, a heterogeneous graph neural network that learns the latent features of both the circuit netlist and placement scheme through dual-space message-passing; and (3) the HeteroGNN-embedded cGAN, a model that combines the HeteroGNN with a cGAN for accurate FPGA routing congestion prediction. Compared to state-of-the-art approaches, our method reduces the routing congestion prediction’s root-mean-square error by 18.2% on the VTR7 benchmarks and by 15.0% on the large-scale Titan23 benchmarks. The code associated with this article can be found at https://github.com/AIPnR/FPGA_Hetero_Congestion_Prediction .
Qingyu Yang 0004, Jingjin Li, Rui Li 0095, Yuting He 0002, Yajun Ha, LinLin Shen, Ruibin Bai, Heng Yu 0001
ACM Trans. Design Autom. Electr. Syst.3
2026 DSHD-CAM: High-Throughput RRAM CAM Leveraging Dynamic Shifted Hamming Distance for Genome Analysis
abstract
Genome analysis has been critical in various applications, such as infectious disease control. High throughput is an essential requirement for genome analysis in data-intensive scenarios, which requires acceleration by content-addressable memory (CAM) with parallel comparison capability. However, existing genome analysis CAMs still face inadequate throughput issues due to large cell area, excessive array storage redundancy, and inefficient comparison algorithms. To address these issues, we propose a high-throughput dynamic shifted Hamming distance (SHD) resistive random access memory (RRAM)-based CAM (DSHD-CAM) that leverages the characteristics of genome analysis. First, we propose a compact RRAM-based CAM cell utilizing one-hot encoding and time-domain computation to minimize cell area. Second, we propose a dense CAM array utilizing an efficient storage scheme to reduce array storage redundancy. Third, we propose a dynamic SHD search algorithm filtering out low-match-potential cases to reduce search latency. Compared to the state-of-the-art (SOTA), DSHD-CAM achieves an average of$7.58\times $higher throughput under the same area constraints while maintaining competitive sensitivity and precision.
Chenxin Jiang, Tiankuo Zheng, Rui Li 0095, Yuhao Shu, Yajun Ha
IEEE Trans. Very Large Scale Integr. Syst.4
2025 RefSCAT: Formal Verification of Logic-Optimized Multipliers via Automated Reference Multiplier Generation and SCA-SAT Synergy
abstract
Formally verifying logic-optimized integer multipliers remains a crucial yet insufficiently addressed problem in both industry and academia, presenting significant verification challenges, particularly when verifying the large-scale logic-optimized multipliers with diverse architectures. Satisfiability (SAT)-based methods require structurally similar and known correct reference multipliers, which may not always be readily accessible. Symbolic computer algebra (SCA) techniques can verify multipliers without references but encounter difficulties with optimized multipliers due to unclear adder boundaries. To enable effective formal verification of the optimized multipliers, we propose the RefSCAT framework, which contains a reference multiplier generator that produces references structurally similar to the optimized multiplier with clear adder boundaries, enabling a synergistic SCA-SAT verification flow. First, we propose a reverse engineering algorithm that extracts the essential adder tree from the optimized multiplier, ensuring similarity. Second, since only a partial netlist is extractable after optimization, we propose a constraint satisfaction algorithm to complete the generation using only adders while following the extracted netlist, ensuring both similarity and clear adder boundaries. Third, leveraging the generated reference, we propose a synergized SCA-SAT verification flow that verifies the generated reference using SCA and then uses it as a correct reference for the SAT-based verification. The experiments demonstrate that RefSCAT can successfully verify logic-optimized multipliers with diverse partial-product-based architectures up to 128 bits, outperforming the state-of-the-art methods by verifying at least 29% more benchmarks.
Rui Li 0095, Lin Li 0079, Heng Yu 0001, Masahiro Fujita 0004, Weixiong Jiang, Yajun Ha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 RefSCAT-2.0: Formal Verification of Large-Scale Optimized Multipliers via Quantum-Inspired Ant Colony Optimization-Based Reference Generation
abstract
Formal verification of large-scale optimized integer multipliers remains a critical yet insufficiently addressed challenge in industry and academia. Current methods employ reference multiplier generators to automatically construct structurally similar reference multipliers, which are then used by Satisfiability (SAT)-based techniques to verify equivalence with optimized multipliers. However, these approaches face limitations when generating references for large-scale optimized multipliers within acceptable timeframes. To address these limitations, we introduce the RefSCAT-2.0 framework, designed to rapidly produce high-quality large-scale reference multipliers. Firstly, we generate the macro-architecture to determine the number of adders required for constructing the reference multiplier. We propose a novel Integer Linear Programming (ILP)-based macro-architecture generation algorithm that minimizes the number of allocated adders, thereby reducing the overall problem complexity. Secondly, we organize the allocated adders into groups to simplify the subsequent generation process. We present a multi-level scheduler that automatically decomposes adders into groups with minimized interdependencies, ensuring both the quality of generation and a reduction in overall generation complexity. Thirdly, we generate the micro-architecture for each scheduled group, wherein we finalize the connections between adders. We present a graph-based design space representation coupled with a quantum-inspired ant colony optimization (QACO)-based generation algorithm that can efficiently explores the micro-architectures of each scheduled group. Experimental results show that RefSCAT-2.0 successfully verifies all 124 cases in a 256-bit optimized multiplier benchmark suite, outperforming SCA-based tcad22revsca and hybrid RefSCATTCAD24 methods which solve only 24 cases each.
Rui Li 0095, Lin Li 0079, Heng Yu 0001, Masahiro Fujita 0004, Weixiong Jiang, Yajun Ha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 A Recursion and Lock Free GPU-Based Logic Rewriting Framework Exploiting Both Intranode and Internode Parallelism
abstract
Logic rewriting is an effective but time-consuming technique to optimize the multilevel logic network by rewriting subnetworks of the input network with other logic equivalent structures. However, contemporary multithread rewriting algorithms either fail to parallelize the subprocedures of rewriting for individual nodes (intranode parallelism) or require locks to ensure the mutual exclusive among the scheduled nodes that are rewritten concurrently (internode parallelism), hence inevitably decreasing the degrees of parallelism and the scalability. This article proposes a novel GPU-based logic rewriting acceleration framework to address the mentioned issues in two phases. First, to exploit the intranode parallelism, we propose recursion-free algorithms that parallelize subprocedures of rewriting, which was hard to achieve due to the highly recursive nature of original rewriting algorithms. Second, to exploit the internode parallelism, we propose a work scheduler that can schedule mutually exclusive nodes and a GPU-friendly data structure that can support efficient concurrent operations. The new work scheduler and data structure allow simultaneously processing plenty of nodes without using locks. Experimental results show that our method can achieve on average$3.81\times $speedup, compared to the state-of-the-art GPU-parallel method with the same quality of results.
Lin Li 0079, Rui Li 0095, Yajun Ha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 Criticality-Aware Negotiation-Driven Scrubbing Scheduling for Reliability Maximization in SRAM-Based FPGAs
abstract
Memory scrubbing is a resource-efficient technique to ensure the high reliability of SRAM-based FPGAs by refreshing the configuration memory just before its execution. To maximize reliability, a scrubbing scheduling algorithm is expected to scrub as many tasks as possible. Unfortunately, contemporary scheduling algorithms either suboptimally handle scrubbing conflicts under bursty requests from multiple user tasks or discriminate against low-criticality tasks by giving them very low scrubbing opportunities. Besides, exploring the architectural support for scrubbing problems may bring considerable potential for reliability improvements. However, this direction of scheduling-architecture co-optimization has not been well studied so far. In this article, we propose a negotiation-based dynamic scrubbing framework, which addresses the above-mentioned issues in three phases: 1) we propose a negotiation-driven scrubbing scheduling algorithm, which temporarily allows and iteratively reduces the conflicts of scrubbing tasks in order to accommodate more scrubbing tasks to be scheduled; 2) we develop a logistic probability model to prevent scheduling starvation of a set of mix-criticality tasks by dynamically legalizing conflicting ones, considering both the criticality and schedulability of each task; and 3) we develop a dynamic voltage/frequency scaling-based multi-ICAPs allocation algorithm to co-optimize with FPGA architectural features for reliability maximization. Compared to the state-of-the-art, experimental results show that our work achieves up to 31.46% improvement in terms of reliability for contemporary SRAM-based FPGAs.
Rui Li 0095, Heng Yu 0001, Lin Li 0079, Yajun Ha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 FODM: A Framework for Accurate Online Delay Measurement Supporting All Timing Paths in FPGA
abstract
Voltage and frequency scaling (VFS) has been widely used to improve energy efficiency, lifespan, and system reliability by converting conservative timing margins into$V_{\text {dd}}$reduction. Along these lines, to investigate the potential implementation of VFS technique in exploring the timing margins under different voltages and frequencies,in situor online circuit delay measurement is required to monitor all timing paths, which are usually ended with terminal registers. The previously reported online delay measurement approaches require the output of a terminal register to be measurable. However, some FPGA timing paths are ended with embedded hardcores such as DSPs or BRAMs. It is impossible to measure the output of the terminal register inside a hardcore. To address the issue, we propose an online delay monitor (ODM) that can accurately measure the delay of any type of timing path in real-time conditions. The ODM is mainly composed of two shadow registers and a phase-shifted clock. The shadow registers use a phase-shifted clock signal as the input and the output signal of the combinational logic as the clock. In addition, we present an automatic tool and its corresponding design flow (FODM) for inserting an ODM to monitor a path. Compared with the state-of-the-art, our experimental results indicate that the proposed method has the ability to accurately measure the delays online for all the potential timing paths, regardless of their path termination types. Moreover, we demonstrate an average measurement error of only 1.51% using eight floating-point operators at different voltages.
Weixiong Jiang, Heng Yu 0001, Hongtu Zhang, Yuhao Shu, Rui Li 0095, Yajun Ha
IEEE Trans. Very Large Scale Integr. Syst.5
2021 TAIT: One-Shot Full-Integer Lightweight DNN Quantization via Tunable Activation Imbalance Transfer
abstract
Both parameter quantization and depthwise convolution are essential measures to provide high-accuracy, lightweight, and resource-friendly solutions when deploying deep neural networks (DNNs) onto edge-AI devices. However, combining the two methodologies may lead to adverse effects: It either suffers from significant accuracy loss or long finetuning time. Besides, contemporary quantization methods are only selectively applied to weight and activation values but not bias and scaling factor values, making them less practical for ASIC/FPGA accelerators. To solve these issues, we propose a novel quantization framework that is effectively optimized for depthwise convolution networks. We discover that the uniformity of the value range within a tensor can serve as a predictor for the tensor’s quantization error. Under the guidance of this predictor, we develop a mechanism called Tunable Activation Imbalance Transfer (TAIT), which tunes the value range uniformity between an activated feature map and its latter weights. Moreover, TAIT fully supports full-integer quantization. We demonstrate TAIT on SkyNet and deploy it on FPGA. Compared to the state-of-the-art, our quantization framework and system design achieve 2.2%+ IoU, $2.4 \times$ speed, and $1.8 \times$ energy efficiency improvements, without any requirement of finetuning.
Weixiong Jiang, Heng Yu 0001, Hao Sun 0035, Rui Li 0095, Yajun Ha
DAC5
2020 DVFS-Based Scrubbing Scheduling for Reliability Maximization on Parallel Tasks in SRAM-based FPGAs
abstract
To obtain high reliability but avoiding the huge area overhead of traditional triple modular redundancy (TMR) methods in SRAM-based FPGAs, scrubbing based methods reconfigure the configuration memory of each task just before its execution. However, due to the limitation of the FPGA reconfiguration module that can only scrub one task at a time, parallel tasks may leave stringent timing requirements to schedule their scrubbing processes. Thus the scrubbing requests may be either delayed or omitted, leading to a less reliable system. To address this issue, we propose a novel optimal DVFS-based scrubbing algorithm to adjust the execution time of user tasks, thus significantly enhance the chance to schedule scrubbing successfully for parallel tasks. Besides, we develop an approximation algorithm to speed up its optimal version and develop a novel K-Means based method to reduce the memory usage of the algorithm. Compared to the state-of-the-art, experimental results show that our work achieves up to 36.11% improvement on system reliability with comparable algorithm execution time and memory consumption.
Rui Li 0095, Heng Yu 0001, Weixiong Jiang, Yajun Ha
DAC1
2020 An Accurate FPGA Online Delay Monitor Supporting All Timing Paths
abstract
Accurate circuit delay measurement is essential for various purposes such as aging detection, health monitoring, and dynamic voltage and frequency scaling. State-of-the-art measurement techniques exhibit several limitations. For example, they are insufficiently informative by only returning binary results on the status of the circuit being normal or abnormal. More importantly, current approaches are not applicable for measuring the delay of timing paths that end with DSPs and BRAMs. To address the issues, we propose a novel online delay monitor (ODM) for modern FPGA platforms that (1) accurately returns the numerical delay values, (2) and is compatible with all types of timing paths in FPGAs. Our proposed ODM is achieved by employing a shadow register triggered by the output signal of a combinational circuit to sample a phase shifting clock. Besides, our design is capable of conveniently measuring the clock jitters, so we are able to propose an associated jitter management scheme to ensure correct ODM sampling. Experimental results show that our ODM achieves an error within 2% with respect to the ground truth.
Weixiong Jiang, Rui Li 0095, Heng Yu 0001, Yajun Ha
ISCAS2