Peiyu Liao

dblp:282/5353 · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0003-1220-1363ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 5 first-author · 19 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Submodular Maximization-inspired Adaptive Routing Bend Space Planning
abstract
Routing can greatly impact tape-out chip performance by determining the physical layout of metal wire segments. As designs grow in complexity and size, modern routing frameworks struggle to manage limited routing resources among numerous nets efficiently. In this paper, we introduce a novel adaptive routing bend space planning framework, ARSP, that adaptively adjusts the routing bend space for each net based on the availability of routing resources throughout the routing flow. ARSP is built on a well-defined submodular maximization problem and uses an efficient approximation algorithm to ensure sub-optimal performance. Integrating ARSP with state-of-the-art routing flows shows an average improvement of 6.34% and 5.11% in reducing shorts and spacing violations, respectively. Additionally, our adaptive planning framework outperforms all static routing space planning strategies in both effectiveness and efficiency, showing the necessity of adaptive planning.
Siting Liu 0002, Peng Xu 0052, Peiyu Liao, Keren Zhu 0001, Yibo Lin, Bei Yu 0001
DATE3
2026 RegPlace: Regularity-Aware Placement for Full-System DNN Accelerator Designs
abstract
The rise of deep neural network accelerators demands physical design tools that recognize spatial regularity patterns. Traditional placers, unaware of the regularity of spatial arrays, produce suboptimal solutions. This work proposes RegPlace, a regularity-aware placement algorithm for full-system DNN accelerators that automatically identifies processing elements using graph convolutional networks and employs a variance-based soft regularity loss to guide optimization. Compared to state-of-the-art methods, our approach achieves up to 6% wirelength reduction while maintaining comparable runtime, with post-placement metrics further confirming its effectiveness.
Jiaxi Jiang, Yuan Pu 0001, Yuxuan Zhao 0001, Peiyu Liao, Zuodong Zhang, Yibo Lin, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 3D-Flow: Flow-based Standard Cell Legalization for 3D ICs
abstract
The standard-cell placement legalization is a critical step in the physical design. The emerging 3D ICs have brought challenges to traditional legalizers on efficiency and effectiveness. In this work, we present a fast flow-based legalization algorithm, 3D-Flow, that minimizes cell displacement in a 3D solution space. Our legalizer resolves overflowed bins by finding the shortest augmenting path on a 3D grid graph, utilizing an effective branch-and-bound algorithm. Moreover, a post-optimization with a cycle-canceling algorithm is proposed to minimize the maximum displacement. Our approach leverages the global perspective inherent in network flow methods, considering multiple dies in 3D ICs to minimize cell displacement. Experimental results on ICCAD 2022 and 2023 contest benchmarks demonstrate our proposed algorithm achieves up to 13% and 43% less average and maximum cell displacement compared to state-of-the-art legalizers in a similar runtime.
Yuxuan Zhao 0001, Peiyu Liao, Bei Yu 0001
DAC2
2025 Ultrafast Density Gradient Accumulation in 3D Analytical Placement with Divergence Theorem
abstract
Density gradient accumulation plays a pivotal role in 3D analytical placement. Analytical placers rely on this fundamental operation during the backward step of each iteration to compute the gradient of the density penalty for every node. This primitive operation thus constitutes a significant runtime bottleneck, especially for mixed-size designs with large macros. Furthermore, this bottleneck becomes increasingly critical as the grid size in 3D placement is considerably larger than that in conventional 2D placement. In this paper, we propose an algorithm inspired by the divergence theorem to reduce the time complexity of density gradient accumulation. We also present our implementations of this algorithm for both CPU and GPU versions. Experimental results demonstrate that our method achieves more than 3× end-to-end runtime speedup on CPU and GPU compared to the SOTA analytical 3D placer.
Peiyu Liao, Yuxuan Zhao 0001, Siting Liu 0002, Bei Yu 0001
ICCAD1
2025 H3D: Heterogeneous Resources Aware Global Router for Face-to-Face Bonded 3D ICs
abstract
The emerging 3D ICs have brought challenges to traditional routers in deciding the intra-die and inter-die interconnects. Existing pseudo-3D flows rely on 2D IC routing engines, combined with a 3D via legalization step to complete routing. The separation of intra-die and inter-die routing significantly degrades solution quality. To address this issue, we propose H3D, the first native 3D global router designed for face-to-face bonded 3D ICs. H3D constructs a heterogeneous routing grid to represent routing and hybrid bonding terminal (HBT) resources. We develop dedicated dynamic programming-based algorithms to optimize the HBT number and locations for 3D Steiner trees on the heterogeneous grid. Specifically, H3D minimizes the number of HBTs by traversing the Steiner tree in reverse depth-first order and relocates HBTs to legal locations with minimal wirelength in reverse breadth-first order, leveraging the convexity of L1 distance for efficient optimization. Experimental results on various real-world designs demonstrate that H3D achieves 12% shorter wirelength, 25% fewer HBTs, and 1.9× speedup compared to state-of-the-art approaches.
Yuxuan Zhao 0001, Siting Liu 0002, Peiyu Liao, Bei Yu 0001
ICCAD4
2025 Layout Decomposition via Boolean Satisfiability
abstract
Multiple patterning lithography (MPL) has been introduced in the integrated circuits manufacturing industry to enhance feature density as the technology node advances. A crucial step of MPL is assigning layout features to different masks, namely layout decomposition. Exact algorithms like integer linear programming (ILP) can solve layout decomposition to optimality but lack scalability for dense patterns. Relaxation algorithms (e.g., linear programming and semi-definite programming) and heuristics (e.g., exact cover) are capable of handling large cases at the cost of inferior solution quality. These methods rely on different mathematical solvers and expert-designed heuristics to offer a balance between solution quality and computational efficiency. In this article, we propose a unified layout decomposition framework comprising three algorithms: 1) satisfiability (SAT)-exact; 2) SAT-bilevel; and 3) SAT-fast, all leveraging the capabilities of Boolean SAT solvers. The SAT-exact ensures optimality, but with faster convergence than ILP, SAT-bilevel addresses the decomposition as a bilevel optimization problem for rapid near-optimal solutions, and SAT-fast handles very large layouts in an incremental manner. Experimental results demonstrate our framework’s superiority over existing state-of-the-art methods in terms of solution quality and runtime.
Hongduo Liu, Peiyu Liao, Mengchuan Zou, Xijun Li, Mingxuan Yuan, Tsung-Yi Ho, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 Analytical Heterogeneous Die-to-Die 3-D Placement With Macros
abstract
This article presents an innovative approach to 3-D mixed-size placement in heterogeneous face-to-face (F2F) bonded 3-D ICs. We propose an analytical framework that utilizes a dedicated density model and a bistratal wirelength model, effectively handling macros and standard cells in a 3-D solution space. A novel 3-D preconditioner is developed to resolve the topological and physical gap between macros and standard cells. Additionally, we propose a mixed-integer linear programming (MILP) formulation for macro rotation to optimize wirelength. Our framework is implemented with full-scale GPU acceleration, leveraging an adaptive 3-D density accumulation algorithm and an incremental wirelength gradient algorithm. Experimental results on ICCAD 2023 contest benchmarks demonstrate that our framework can achieve 5.9% quality score improvement compared to the first-place winner with 4.0$\times $runtime speedup. Additional experiments on modern RISC-V designs further validate the generalizability and superiority of our framework.
Yuxuan Zhao 0001, Peiyu Liao, Siting Liu 0002, Jiaxi Jiang, Yibo Lin, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 p-Laplacian Adaptation for Generative Pre-trained Vision-Language Models
abstract
Vision-Language models (VLMs) pre-trained on large corpora have demonstrated notable success across a range of downstream tasks. In light of the rapidly increasing size of pre-trained VLMs, parameter-efficient transfer learning (PETL) has garnered attention as a viable alternative to full fine-tuning. One such approach is the adapter, which introduces a few trainable parameters into the pre-trained models while preserving the original parameters during adaptation. In this paper, we present a novel modeling framework that recasts adapter tuning after attention as a graph message passing process on attention graphs, where the projected query and value features and attention matrix constitute the node features and the graph adjacency matrix, respectively. Within this framework, tuning adapters in VLMs necessitates handling heterophilic graphs, owing to the disparity between the projected query and value space. To address this challenge, we propose a new adapter architecture, p-adapter, which employs p-Laplacian message passing in Graph Neural Networks (GNNs). Specifically, the attention weights are re-normalized based on the features, and the features are then aggregated using the calibrated attention matrix, enabling the dynamic exploitation of information with varying frequencies in the heterophilic attention graphs. We conduct extensive experiments on different pre-trained VLMs and multi-modal tasks, including visual question answering, visual entailment, and image captioning. The experimental results validate our method's significant superiority over other PETL methods. Our code is available at https://github.com/wuhy68/p-Adapter/.
Haoyuan Wu, Xinyun Zhang 0001, Peng Xu 0052, Peiyu Liao, Xufeng Yao, Bei Yu 0001
AAAI4
2024 Parallel Gröbner Basis Rewriting and Memory Optimization for Efficient Multiplier Verification
abstract
Formal verification of integer multipliers is a significant but time-consuming problem. This paper introduces a novel approach that emphasizes the acceleration of symbolic computer algebra (SCA)-based verification systems from the perspective of efficient implementation instead of traditional algorithm enhancement. Our first strategy involves leveraging parallel computing to accelerate the rewriting process of the Gröbner basis. Confronting the issue of frequent memory operations during the Gröbner basis reduction phase, we propose a double buffering scheme coupled with an operator scheduler to minimize memory allocation and deallocation. These unique contributions are integrated into a state-of-the-art verification tool and result in substantial improvements in verification speed, demonstrating more than 15× speedup for a 1024×1024 multiplier.
Hongduo Liu, Peiyu Liao, Junhua Huang, Hui-Ling Zhen, Mingxuan Yuan, Tsung-Yi Ho, Bei Yu 0001
DATE2
2024 Multi-Electrostatics Based Placement for Non-Integer Multiple-Height Cells
abstract
A circuit design incorporating non-integer multi-height (NIMH) cells, such as a combination of 8-track and 12-track cells, offers increased flexibility in optimizing area, timing, and power simultaneously. The conventional approach for placing NIMH cells involves using commercial tools to generate an initial global placement, followed by a legalization process that divides the block area into row regions with specific heights and relocates cells to rows of matching height. However, such placement flow often causes significant disruptions in the initial placement results, resulting in inferior wirelength. To address this issue, we propose a novel multi-electrostatics-based global placement algorithm that utilizes the NIMH-aware clustering method to dynamically generate rows. This algorithm directly tackles the global placement problem with NIMH cells. Specifically, we utilize an augmented Lagrangian formulation along with a preconditioning technique to achieve high-quality solutions with fast and robust numerical convergence. Experimental results on the OpenCores benchmarks demonstrate that our algorithm achieves about 12% improvements on HPWL with 23.5X speed up on average, outperforming state-of-the-art approaches. Furthermore, our placement solutions demonstrate a substantial improvement in WNS and TNS by 22% and 49% respectively. These results affirm the efficiency and effectiveness of our proposed algorithm in solving row-based placement problems for NIMH cells.
Yu Zhang 0189, Yuan Pu 0001, Fangzhou Liu 0005, Peiyu Liao, Kai-Yuan Chao, Keren Zhu 0001, Yibo Lin, Bei Yu 0001
ISPD4
2024 Analytical Die-to-Die 3-D Placement With Bistratal Wirelength Model and GPU Acceleration
abstract
In this paper, we present a new analytical 3D placement framework with a bistratal wirelength model for F2Fbonded 3D ICs with heterogeneous technology nodes based on the electrostatic-based density model. The proposed framework, enabled GPU-acceleration, is capable of efficiently determining node partitioning and locations simultaneously, leveraging the dedicated 3D wirelength model and density model. The experimental results on ICCAD 2022 contest benchmarks demonstrate that our proposed 3D placement framework can achieve up to 6.1% wirelength improvement and 4.1% on average compared to the first-place winner with much fewer vertical interconnections and up to 9.8× runtime speedup. Notably, the proposed framework also outperforms the state-of-the-art 3D analytical placer by up to 3.3% wirelength improvement and 2.1% on average with up to 8.8× acceleration on large cases using GPUs.
Peiyu Liao, Yuxuan Zhao 0001, Dawei Guo, Yibo Lin, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 On a Moreau Envelope Wirelength Model for Analytical Global Placement
abstract
Analytical placement is proven to be effective in global placement. The differentiability of wirelength models is very critical to gradient-based numerical optimization. Most previous works approximate the non-smooth half-perimeter wirelength (HPWL) model with various differentiable functions. In this paper, we propose a new differentiable wirelength model using the Moreau envelope to approximate HPWL. By combining the state-of-the-art electrostatic-based placement algorithm, the experimental results demonstrate that our proposed algorithm can achieve up to 5.4% HPWL improvement and more than 1% on average compared to the most widely-used nonlinear wirelength model.
Peiyu Liao, Hongduo Liu, Yibo Lin, Bei Yu 0001, Martin D. F. Wong
DAC1
2023 Layout Decomposition via Boolean Satisfiability
abstract
Multiple patterning lithography (MPL) has been introduced in the integrated circuits manufacturing industry to enhance feature density as the technology node advances. A crucial step of MPL is assigning layout features to different masks, namely layout decomposition. Exact algorithms like integer linear programming (ILP) can solve layout decomposition to optimality but lacks scalability for very dense patterns. Approximation algorithms (e.g., linear programming, semi-definite programming) and heuristics (e.g., Exact-Cover) are capable of handling large cases but can only get inferior solutions. In this paper, we propose a new exact algorithm that tackles layout decomposition by solving a series of boolean satisfiability instances. Our algorithm can preserve optimality and achieve more than 4× speedup compared to ILP. In addition, we provide an approximation algorithm by reformulating the layout decomposition to a bilevel optimization problem. Experiments show that our approximation algorithm can attain higher solution quality compared to SDP and heuristics within faster convergence.
Hongduo Liu, Peiyu Liao, Mengchuan Zou, Xijun Li, Mingxuan Yuan, Tsung-Yi Ho, Bei Yu 0001
DAC2
2023 FastGR: Global Routing on CPU-GPU with Heterogeneous Task Graph Scheduler (Extended Abstract)
abstract
Running time is a key metric across the standard physical design flow stages. However, with the rapid growth in design sizes, routing runtime has become the runtime bottleneck in the physical design flow. To improve the effectiveness of the modern global router, we propose a global routing framework with GPU-accelerated routing algorithms and a heterogeneous task graph scheduler, called FastGR. Its runtime-oriented version FastGRL achieves 2.489× speedup compared with the state-of-the-art global router. Furthermore, the GPU-accelerated L-shape pattern routing used in FastGRL can contribute to 9.324× speedup over the sequential algorithm on CPU. Its quality-oriented version FastGRH offers further quality improvement over FastGRL with similar acceleration.
Siting Liu 0002, Yuan Pu 0001, Peiyu Liao, Hongzhong Wu, Rui Zhang 0040, Zhitang Chen, Wenlong Lv, Yibo Lin, Bei Yu 0001
IJCAI3
2023 DREAMPlace 4.0: Timing-Driven Placement With Momentum-Based Net Weighting and Lagrangian-Based Refinement
abstract
Optimizing timing is critical to the design closure of integrated circuits (ICs). However, most existing algorithms for circuit placement focus on the optimization of wirelength instead of timing metrics. This article presents a timing-driven placement framework. It consists of a global placement stage based on net weighting with momentum, and a detailed placement stage based on the Lagrangian multipliers. By improving the preconditioners and timing engines to facilitate net weighting and discrete local search, we have achieved superior timing improvement on benchmarks from ICCAD 2015 contest, including worst negative slack (WNS) and total negative slack (TNS).
Peiyu Liao, Dawei Guo, Zizheng Guo 0001, Siting Liu 0002, Yibo Lin, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 FastGR: Global Routing on CPU-GPU With Heterogeneous Task Graph Scheduler
abstract
Running time is a key metric across the standard physical design flow stages. However, with the rapid growth in design sizes, routing runtime has become the runtime bottleneck in the physical design flow. As a result, speeding routing becomes a critical and pressing task for IC design automation. Aside from the running time, we need to evaluate the quality of the global routing solution since a poor global routing engine degrades the solution performance after the entire routing stage. This work takes both of them into consideration. We propose a global routing framework with GPU-accelerated routing algorithms and a heterogeneous task graph scheduler, called FastGR, to accelerate the procedure of the modern global router and improve its effectiveness. Its runtime-oriented version$\text {FastGR}^{\text {L}}$achieves$2.489\times $speedup compared with the state-of-the-art global router. Furthermore, the GPU-accelerated L-shape pattern routing algorithm used in$\text {FastGR}^{\text {L}}$can contribute to$9.324\times $speedup over the sequential algorithm on CPU. Its quality-oriented version$\text {FastGR}^{\text {H}}$offers a 27.855% improvement of the number of shorts over the runtime-oriented version and still gets$1.970\times $faster than the most advanced global router.
Siting Liu 0002, Yuan Pu 0001, Peiyu Liao, Hongzhong Wu, Rui Zhang 0040, Zhitang Chen, Wenlong Lv, Yibo Lin, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 CTM-SRAF: Continuous Transmission Mask-Based Constraint-Aware Subresolution Assist Feature Generation
abstract
In the lithography process, subresolution assist features (SRAFs), as an essential resolution enhancement technique (RET), is applied to improve the pattern fidelity and enlarge the process window. In this article, we propose a robust constraint-aware SRAF generation method based on continuous transmission mask (CTM). The intensity distribution on the CTM is extracted to guide the SRAF generation. The SRAF insertion also honors the design rules, which is formulated as integer programming with quadratic constraints and solved by a fast yet efficient algorithm. A fast probe-based SRAF evolution method is proposed to determine the shapes of SRAFs. The effectiveness and efficiency are demonstrated based on the experimental results.
Ziyang Yu 0001, Peiyu Liao, Yuzhe Ma, Bei Yu 0001, Martin D. F. Wong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 DREAMPlace 4.0: Timing-driven Global Placement with Momentum-based Net Weighting
abstract
Timing optimization is critical to integrated circuit (IC) design closure. Existing global placement algorithms mostly focus on wirelength optimization without considering timing. In this paper, we propose a timing-driven global placement algorithm leveraging a momentum-based net weighting strategy. Besides, we improve the preconditioner to incorporate our net weighting scheme. Experimental results on ICCAD 2015 contest benchmarks demonstrate that our algorithm can significantly improve total negative slack (TNS) and meanwhile be beneficial to worse negative slack (WNS).
Peiyu Liao, Siting Liu 0002, Zhitang Chen, Wenlong Lv, Yibo Lin, Bei Yu 0001
DATE1
2022 FastGR: Global Routing on CPU-GPU with Heterogeneous Task Graph Scheduler
abstract
Routing is an essential step to integrated circuits (IC) design closure. With the rapid increase of design scales, routing has become the runtime bottleneck in the physical design flow. Thus, accelerating routing becomes a vital and urgent task for IC design automation. This paper proposes a global routing framework running on hybrid CPU-GPU platforms with a heterogeneous task scheduler and a GPU-accelerated pattern routing algorithm. We demonstrate that the task scheduler can lead to 2.307 × speedup compared with the widely-adopted batch-based parallelization strategy on CPU and the GPU-accelerated pattern routing algorithm can contribute to 10.877 × speedup over the sequential algorithm on CPU. Finally, the combined techniques can achieve 2.426 × speedup without quality degradation compared with the state-of-the-art global router.
Siting Liu 0002, Peiyu Liao, Rui Zhang 0040, Zhitang Chen, Wenlong Lv, Yibo Lin, Bei Yu 0001
DATE2
2021 Physical Synthesis for Advanced Neural Network Processors
abstract
The remarkable breakthroughs in deep learning have led to a dramatic thirst for computational resources to tackle interesting real-world problems. Various neural network processors have been proposed for the purpose, yet, far fewer discussions have been made on the physical synthesis for such specialized processors, especially in advanced technology nodes. In this paper, we review several physical synthesis techniques for advanced neural network processors. We especially argue that datapath design is an essential methodology in the above procedures due to the organized computational graph of neural networks. As a case study, we investigate a wafer-scale deep learning accelerator placement problem in detail.
Zhuolun He, Peiyu Liao, Siting Liu 0002, Yuzhe Ma, Yibo Lin, Bei Yu 0001
ASP-DAC2
2021 Global Placement with Deep Learning-Enabled Explicit Routability Optimization
abstract
Placement and routing (PnR) is the most time-consuming part of the physical design flow. Recognizing the routing performance ahead of time can assist designers and design tools to optimize placement results in advance. In this paper, we propose a fully convolutional network model to predict congestion hotspots and then incorporate this prediction model into a placement engine, DREAMPlace, to get a more route-friendly result. The experimental results on ISPD2015 benchmarks show that with the superior accuracy of the prediction model, our proposed approach can achieve up to 9.05% reduction in congestion rate and 5.30% reduction in routed wirelength compared with the state-of-the-art.
Siting Liu 0002, Qi Sun 0002, Peiyu Liao, Yibo Lin, Bei Yu 0001
DATE3
2020 Learn to Floorplan through Acquisition of Effective Local Search Heuristics
abstract
Automatic heuristic design through reinforcement learning opens a promising direction for solving computationally difficult problems. Unlike most previous works that aimed at solution construction, we explore the possibility of acquiring local search heuristics through massive search experiments. To illustrate the applicability, an agent is trained to perform a walk in the search space by selecting a candidate neighbor solution at each step. Specifically, we target the floorplanning problem, where a neighbor solution is generated through perturbing the sequence pair encoding of a floorplan. Experimental results demonstrate the efficacy of the acquired heuristics as well as the potential of automatic heuristic design.
Zhuolun He, Yuzhe Ma, Peiyu Liao, Ngai Wong 0001, Bei Yu 0001, Martin D. F. Wong
ICCD4