Yunqi Shi

dblp:285/4409 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Timing-driven Detailed Placement via TimingMask-guided Path-level Optimization
abstract
Timing-driven detailed placement is a critical stage in very large scale integrated (VLSI) design, aiming to locally adjust cell positions to further improve circuit timing performance. Existing methods commonly adopt proxy metrics as optimization objectives, such as weighted wirelength and approximate delay. However, these surrogate metrics are not fully aligned with the final timing metrics obtained through static timing analysis (STA), often leading to suboptimal timing results. Besides, methods based directly on STA tools suffer from very low search efficiency, making the cost of timing optimization prohibitive. To address these issues, we propose an effective timing-driven detailed placement method via TimingMask-guided path-level optimization. One core of our method is the TimingMask guidance mechanism, which integrates both arc delay and path slack information based on the RC timing model, thereby providing more targeted and effective guidance for refinement of critical cells. Meanwhile, our method adopts a path-level timing evaluation strategy with incremental updates, accelerating the optimization process while preserving timing accuracy. Experimental results on the ICCAD 2015 contest benchmarks demonstrate that our method significantly outperforms state-of-the-art detailed placement methods such as DREAMPlace4.0 DP, achieving an average improvement of 25.3% in total negative slack (TNS) and 21.7% in worst negative slack (WNS).
Ruo-Tong Chen, Chengrui Gao, Ke Xue 0001, Yunqi Shi, Xi Lin 0001, Mingxuan Yuan, Chao Qian 0001, Zhi-Hua Zhou
DATE5
2026 Dynamic Algorithm Configuration for Global Placement
abstract
Placement is a vital step in the physical design flow of very large-scale integration (VLSI) circuits. GPU-accelerated analytical placement algorithms, such as DREAMPlace, have achieved high-quality performance with dramatic speedup. The algorithm configurations of the analytical placer have a significant impact on its convergence and final performance. However, its tuning process is difficult and time-consuming. Recently, AutoDMP tries to search for optimal static algorithm configurations using Bayesian optimization, but the performance is still limited due to its static strategy, which cannot leverage information during algorithm execution. In this paper, we propose the dynamic algorithm configuration framework for DREAMPlace (DACDMP), using reinforcement learning (RL) to learn the dynamic control policy of the most critical hyperparameter, i.e., the learning rate. Moreover, to address the insufficiency of optimization, we increase the number of optimization steps in each Lagrangian relaxation problem, thereby improving the solution’s optimality. DACDMP outperforms the current leading methods, i.e., DREAMPlace 4.0, AutoDMP, and Xplace. For example, compared to DREAMPlace 4.0, it achieves an average improvement of 2.75% in wirelength, 18.74% in worst negative slack (WNS), 44.60% in total negative slack (TNS), and 29.39% in the number of violation points on the ICCAD 2015 benchmark.
Ke Xue 0001, Ruo-Tong Chen, Yunqi Shi, Mingxuan Yuan, Chao Qian 0001, Zhi-Hua Zhou
DATE4
2026 Reinforcement Learning for Hybrid Bonding Terminal Legalization in 3D ICs
abstract
Hybrid bonding (HB) in 3D ICs enables scaling but introduces overlap challenges from large pitch requirements. Existing legalization methods use exhaustive sliding-window scanning, resulting in significant computational inefficiency. To address this, we propose a reinforcement learning (RL) approach that adaptively selects subregions for targeted displacement optimization. The learned policy generalizes to unseen designs without fine-tuning. Experimental results on open-source and industrial benchmarks show our method fully eliminates overlaps with minimal displacement and reduced runtime compared with baselines.
Wanqi Ren, Chengrui Gao, Yunqi Shi, Mingzhou Fan, Ke Xue 0001, Chenjian Ding, Mingxuan Yuan, Chao Qian 0001
DATE3
2025 ReMaP: Macro Placement by Recursively Prototyping and Periphery-Guided Relocating
abstract
We introduce the ReMaP framework, which generates expert-quality macro placements through recursively prototyping and periphery-guided relocating. A key innovation is ABPlace, an angle-based analytical method that arranges macros along an ellipse to facilitate a rough distribution near the periphery, while optimizing dataflow, minimizing overlap, and ensuring convergence. Based on the results of ABPlace, an efficient heuristic is proposed to position macros along the chip’s periphery, mirroring practices often employed by experts. Our framework outperforms three leading macro placers in both WNS and TNS across eight test cases, achieving improvements up to 34.15% in WNS and 65.39% in TNS, as tested on the popular OpenROAD-flow-scripts infrastructure. Additionally, our parameter autotuning method further improves timing by 8.75%.
Yunqi Shi, Xi Lin 0001, Shixiong Kai, Ke Xue 0001, Mingxuan Yuan, Chao Qian 0001, Zhi-Hua Zhou
DAC1
2025 Timing-Driven Global Placement by Efficient Critical Path Extraction
abstract
Timing optimization during the global placement of integrated circuits has been a significant focus for decades, yet it remains a complex, unresolved issue. Recent analytical methods typically use pin-level timing information to adjust net weights, which is fast and simple but neglects the path-based nature of the timing graph. The existing path-based methods, however, cannot balance the accuracy and efficiency due to the exponential growth of number of critical paths. In this work, we propose a GPU-accelerated timing-driven global placement framework, integrating accurate path-level information into the efficient DREAMPlace infrastructure. It optimizes the fine-grained pin-to-pin attraction objective and is facilitated by efficient critical path extraction. We also design a quadratic distance loss function specifically to align with the RC timing model. Experimental results demonstrate that our method significantly outperforms the current leading timing-driven placers, achieving an average improvement of 40.5% in total negative slack (TNS) and 8.3% in worst negative slack (WNS), as well as an improvement in half-perimeter wirelength (HPWL).
Yunqi Shi, Shixiong Kai, Xi Lin 0001, Ke Xue 0001, Mingxuan Yuan, Chao Qian 0001
DATE1
2024 Reinforcement Learning Policy as Macro Regulator Rather than Macro Placer
abstract
In modern chip design, placement aims at placing millions of circuit modules, which is an essential step that significantly influences power, performance, and area (PPA) metrics. Recently, reinforcement learning (RL) has emerged as a promising technique for improving placement quality, especially macro placement. However, current RL-based placement methods suffer from long training times, low generalization ability, and inability to guarantee PPA results. A key issue lies in the problem formulation, i.e., using RL to place from scratch, which results in limits useful information and inaccurate rewards during the training process. In this work, we propose an approach that utilizes RL for the refinement stage, which allows the RL policy to learn how to adjust existing placement layouts, thereby receiving sufficient information for the policy to act and obtain relatively dense and precise rewards. Additionally, we introduce the concept of regularity during training, which is considered an important metric in the chip design industry but is often overlooked in current RL placement methods. We evaluate our approach on the ISPD 2005 and ICCAD 2015 benchmark, comparing the global half-perimeter wirelength and regularity of our proposed method against several competitive approaches. Besides, we test the PPA performance using commercial software, showing that RL as a regulator can achieve significant PPA improvements. Our RL regulator can fine-tune placements from any method and enhance their quality. Our work opens up new possibilities for the application of RL in placement, providing a more effective and efficient approach to optimizing chip design. Our code is available at \url{https://github.com/lamda-bbo/macro-regulator}.
Ke Xue 0001, Ruo-Tong Chen, Xi Lin 0001, Yunqi Shi, Shixiong Kai, Chao Qian 0001
NeurIPS4
2023 Macro Placement by Wire-Mask-Guided Black-Box Optimization
abstract
The development of very large-scale integration (VLSI) technology has posed new challenges for electronic design automation (EDA) techniques in chip floorplanning. During this process, macro placement is an important subproblem, which tries to determine the positions of all macros with the aim of minimizing half-perimeter wirelength (HPWL) and avoiding overlapping. Previous methods include packing-based, analytical and reinforcement learning methods. In this paper, we propose a new black-box optimization (BBO) framework (called WireMask-BBO) for macro placement, by using a wire-mask-guided greedy procedure for objective evaluation. Equipped with different BBO algorithms, WireMask-BBO empirically achieves significant improvements over previous methods, i.e., achieves significantly shorter HPWL by using much less time. Furthermore, it can fine-tune existing placements by treating them as initial solutions, which can bring up to 50% improvement in HPWL. WireMask-BBO has the potential to significantly improve the quality and efficiency of chip floorplanning, which makes it appealing to researchers and practitioners in EDA and will also promote the application of BBO. Our code is available at https://github.com/lamda-bbo/WireMask-BBO.
Yunqi Shi, Ke Xue 0001, Song Lei, Chao Qian 0001
NeurIPS1