Chunyan Pei

dblp:346/6385 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2026
0009-0001-6898-7927ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Timing-Aware Optimization of Die-Level Routing and TDM Assignment for Multi-FPGA Systems
abstract
The escalating scale and complexity of modern circuits demand multi-FPGA emulation platforms that incorporate multi-die architectures. However, most existing routers remain FPGA-level, optimizing wire-length or total Time-Division Multiplexing (TDM) ratios while disregarding die-level load imbalance and path-level slack. They result in suboptimal performance and timing violations. In this paper, we propose a timing-aware co-optimization framework for die-level routing and TDM assignment, explicitly linking physical constraints to critical path timing slack. The proposed flow features a timing-aware load-balanced die-level router with timing path compression and a timing graph-based TDM assignment. Experiments on industrial designs show that the proposed method improves the worst-path slack by 98% over the existing methods.
Haoyuan Li 0004, Chunyan Pei, Jianwang Zhai, Wenjian Yu
ASP-DAC3
2026 Automated Parameter Tuning for Multi-FPGA Partitioning: A Preference-Guided Approach
abstract
Parameter tuning for multi-FPGA partitioning algorithms represents a bottleneck in modern chip emulation and verification workflows. Current multilevel partitioning tools require manual configuration of various parameters, where each evaluation can take tens of seconds to minutes, making exhaustive search impractical and expert-driven tuning both time-consuming and suboptimal. To automate this process, we propose a preference-guided Bayesian optimization framework specifically designed for industrial FPGA partitioning parameter tuning under limited evaluation budgets. Our approach maximizes the minimum timing slack by incorporating domain-specific insights: we exploit the strong correlation between cutsize and timing performance through a priority-based ranking scheme that guides a pairwise Gaussian process to learn configuration preferences. Additionally, we introduce a kernel input transformation that properly handles the mixed discrete-continuous parameter space typical in EDA tools. Our method converges faster with fewer evaluations and achieves the best timing slack in 60–70% of cases on industrial circuit benchmarks compared to existing methods including standard Bayesian optimization, quasi-random sampling, and state-of-the-art preference learning techniques. The proposed framework reduces parameter tuning from days of manual effort to hours of automated optimization, offering practitioners a deployment-ready solution that improves both design quality and engineering productivity.
Yutao Dai, Shengbo Tong, Chunyan Pei, Zhuohua Liu, Yi Liu 0013, Rui Wang 0014, Wenjian Yu
ASP-DAC3
2026 HGNN-Part: A High-Quality Hypergraph Partitioner Based on Hypergraph Generative Model
abstract
Hypergraph partitioning is a fundamental combinatorial optimization problem with critical applications in VLSI design. While recent deep learning based approaches have shown promise for this problem, they rely on graph neural networks (GNNs) that require transforming hypergraphs into normal graphs, thereby losing the high-order relationships in hypergraph structures. In this work, we propose a novel framework that directly utilizes hypergraph neural networks (HGNNs) to exploit the high-order interactions in hypergraphs. We develop an efficient normalized cut loss computation algorithm optimized for GPU training and apply randomized matrix decomposition techniques to significantly accelerate the eigenvector computation required for node feature extraction without sacrificing quality. To address the scarcity of open-source hypergraph data, we release a comprehensive dataset with 164 VLSI hypergraphs collected from various EDA contests and benchmarks. Extensive experiments on the ISPD98 and ISPD05 benchmarks demonstrate that our method achieves superior partitioning quality compared to state-of-the-art approaches, including multilevel methods (hMETIS), spectral methods (SpecPart, K-SpecPart), and recent deep learning based approaches (MedPart, GenPart). Furthermore, training on our expanded dataset yields additional performance gains, validating the framework’s ability to leverage larger training data effectively.
Shengbo Tong, Rufan Zhou, Chunyan Pei, Wenjian Yu
DATE3
2025 Efficient Hypergraph Modeling of VLSI Circuits for the MFS-Based Emulation and Simulation Acceleration
abstract
As the scale of integrated circuit (IC) design continues to expand, the multi-FPGA system (MFS) is widely employed for logic emulation and simulation acceleration which ensures the functional correctness of logic circuits. During this process, circuit partitioning becomes a dispensable step. In this work, we address the hypergraph modeling techniques for the MFS-orientated circuit partitioning. Firstly, an efficient adaptive flattening algorithm considering multi-dimensional resource constraints and based on dynamic programming (DP) is proposed. Then, a parallel algorithm for clock modeling is proposed. With them, an efficient tool of hypergraph modeling is developed. Experiments on industrial benchmarks with up to sixty million cells have validated the efficiency and correctness of the proposed techniques. The results also demonstrate the benefit of the adaptive flattening to the subsequent hypergraph partitioning, and the significant acceleration effects of the proposed DP-based adaptive flattening and the parallel clock modeling algorithms.
Chunyan Pei, Shengbo Tong, Wenjian Yu
ASP-DAC2
2025 Deep Learning Inspired Capacitance Extraction Techniques
abstract
With the advancement of integrated circuit (IC), the process technology becomes more complicated and the design margin shrinks. Thus, the parasitic extraction is more demanded during IC design. In this invited paper, we survey the research progress on IC capacitance extraction, especially the usage of deep-learning technologies in relevant problems. Firstly, a method based on graph neural network (GNN) for predicting the parasitic capacitances in the pre-layout design stage is presented. It exhibits potential benefit for the optimization of SRAM design. Then, the deep-learning-inspired methods for post-layout capacitance extraction are presented, including CNN-Cap, NAS-Cap and GNN-Cap, etc. They can revamp the accuracy drawback of layout parasitic extraction (LPE) method and the efficiency drawback of 3-D capacitance field solver. Lastly, we briefly review the deep-learning technique for improving the accuracy of the random walk based 3-D capacitance solver for the structures under the advanced process technology.
Wenjian Yu, Shan Shen, Dingcheng Yang, Haoyuan Li 0004, Jiechen Huang, Chunyan Pei
ASP-DAC6
2025 BlasPart: A Deterministic Parallel Partitioner for Balanced Large-Scale Hypergraph Partitioning
abstract
Balanced hypergraph partitioning is a fundamental problem in applications like VLSI design, high-performance computing, etc. Nowadays, large-scale hypergraphs become more common due to the increasing complexity of modern systems. Thus, fast and high-quality deterministic partitioning algorithms are largely in demand. Regarding the quality of partitioning, balance is a critical concern when the number of partitions increases. In this paper, we propose BlasPart, a deterministic parallel algorithm for balanced large-scale hypergraph partitioning. BlasPart leverages a recursive multilevel bisection framework to achieve high-quality partitions while ensuring deterministic outcomes. A level-dependent balance constraint is also proposed to further improve the efficiency and effectiveness of the proposed partitioner. Extensive experiments, with comparisons to the state-of-the-art partitioners (hMETIS, BiPart, and Mt-KaHyParSDet), demonstrate that BlasPart achieves better balance and scalability while maintaining competitive partitioning quality and efficiency. BlasPart runs $3.33 \times$ faster than Mt-KaHyPar-SDet on average for a 4096-way partitioning task on six benchmarks.
Shengbo Tong, Chunyan Pei, Wenjian Yu
DAC2
2024 Deep-Learning-Based Pre-Layout Parasitic Capacitance Prediction on SRAM Designs
abstract
To achieve higher system energy efficiency, SRAM in SoCs is often customized. The parasitic effects cause notable discrepancies between pre-layout and post-layout circuit simulations, leading to difficulty in converging design parameters and excessive design iterations. Is it possible to well predict the parasitics based on the pre-layout circuit, so as to perform parasitic-aware pre-layout simulation? In this work, we propose a deep-learning-based 2-stage model to accurately predict these parasitics in pre-layout stages. The model combines a Graph Neural Network (GNN) classifier and Multi-Layer Perceptron (MLP) regressors, effectively managing class imbalance of the net parasitics in SRAM circuits. We also employ Focal Loss to mitigate the impact of abundant internal net samples and integrate subcircuit information into the graph to abstract the hierarchical structure of schematics. Experiments on 4 real SRAM designs show that our approach not only surpasses the state-of-the-art model in parasitic prediction by a maximum of 19X reduction of error but also significantly boosts the simulation process by up to 598X speedup.
Shan Shen, Dingcheng Yang, Chunyan Pei, Bei Yu 0001, Wenjian Yu
ACM Great Lakes Symposium on VLSI4
2024 EasyPart: An Effective and Comprehensive Hypergraph Partitioner for FPGA-based Emulation
abstract
Logic verification becomes more and more important for the design of large-scale digital integrated circuits (ICs). This makes FPGA-based hardware emulation an imperative step in the design flow, and how to effectively partition and map the circuit netlist into the multi-FPGA system (MFS) for emulation is of concern. In this paper, we present EasyPart, an effective and comprehensive hypergraph partitioner for the FPGA-based hardware emulation. EasyPart can handle the practical constraints in the MFS for logic emulation and includes novel techniques for pursuing minimum hop during topology-driven partitioning and treating the interconnection constraints. We have evaluated EasyPart against state-of-the-art partitioners on public benchmarks. The results show that EasyPart can reduce the cutsize with a comparable or shorter runtime. EasyPart is capable of finding non-hop solutions with better robustness and performance compared to previous work. It also achieves significant improvements in terms of time division multiplexing (TDM) ratio and maximum hop when tested on industrial cases.
Shengbo Tong, Haoyuan Li 0004, Chunyan Pei, Wenjian Yu, Shengjun Liu 0001
ICCAD4