VLDB 2026 Research / reviewers in the wild / expert
Zhaoqi Fu
dblp:321/5669
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2025
0009-0009-6307-7324ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An Effective and Efficient Cross-Link Insertion for Non-Tree Clock Network SynthesisabstractClock skew introduces significant challenge to the overall system performance. Existing non-tree solutions like cross-link insertion often come with limitations, such as the over-consumption of resource and power. In this work, we propose a cross-link insertion algorithm that effectively reduces the clock skew with minimal power overhead, and prioritize delay optimization on the paths with high sensitivity to the skew. The experimental results from the ISPD 2010 benchmarks show a 17% reduction in the mean of clock skew, a 45% decrease in the standard deviation of clock skew, and a 13% lower power consumption versus the advanced non-tree solutions in literature. Jinghao Ding, Jiazhi Wen, Zhaoqi Fu, Mengshi Gong, Yuanrui Qi, Wenxin Yu 0001, Jinjia Zhou |
DATE | 4 |
| 2025 | Timing-Driven Global Placement With Hybrid Heuristics and Nadam-Based Net WeightingabstractTiming optimization is critical to the entire design flow of the very-large-scale integrated (VLSI) circuit, and Global Placement is pivotal in achieving timing closure within the design flow of very-large-scale integration circuits. However, most global placement algorithms focus on optimizing wirelength rather than timing. Therefore, we propose a novel timing-driven global placement algorithm to address this gap. This paper proposes a timing-driven global placement algorithm utilizing a Nadam-based net-weighting strategy. Additionally, we employ a hybrid heuristic approach for adaptive dynamic adjustment of net weights. The experimental results on the ICCAD 2015 contest benchmarks show that compared to the RePlAce, our algorithm significantly improves WNS and TNS by 40.7% and 56.5%, respectively. Linhao Lu, Wenxin Yu 0001, Hongwei Tian, Chengjin Li, Xinmiao Li, Zhaoqi Fu, Zhengjie Zhao, Jingwei Lu |
DATE | 6 |
| 2025 | An Effective Macro Placement Framework with Reinforcement Learning and Monte Carlo Tree Search
Jinghao Ding, Wenxin Yu 0001, Yuanrui Qi, Zhaoqi Fu, Mengshi Gong, I-Chyn Wey, Jinjia Zhou |
ACM Great Lakes Symposium on VLSI | 4 |
| 2024 | A P&R Co- Optimization Engine for Reducing CongestionabstractPlacement and routing (P&R) are two crucial stages in the physical design process to optimize different objectives. For instance, placement often focuses on optimizing the half-perimeter wirelength (HPWL) and estimated congestion while routing attempts to minimize the wirelength and the number of overflows. The misalignment of objectives inevitably leads to a significant decline in solution quality. Therefore, this paper is an efficient Formula R co-optimization engine that bridges the gap between placement and routing disparities. To progressively alleviate routing congestion issues, we perform cell movements on cells causing overflows and then re-run the routing process. Comparing our experimental results with CUGR [5], our method reduced 17.1 in overflow, with slight decreases in wirelength and vias. Dongliang Xia, Wenxin Yu 0001, Zhaoqi Fu, Zejun Gan, Chengjin Li |
ACM Great Lakes Symposium on VLSI | 3 |
| 2024 | Track Assignment Using Gradient Indication and Simulated AnnealingabstractTrack assignment has become a crucial step within the physical design flow. In this work, we proposed a novel track assignment approach using gradients estimation and simulated annealing. Specifically, we formulate the overlap cost function and derive the gradients to estimate the quality of candidate positions for iroutes movement. A simulated annealer is devised to consume the gradients indication and perturb the track assignment solution from the global perspective. Experimental results show 5.15% lower overlap cost in average of all 10 benchmarks from DAC 2012 Routability-Driven Placement Contest, compared to the state-of-the-art approach NTA [1]. Yuanrui Qi, Zejun Gan, Jinghao Ding, Zhaoqi Fu, Mengshi Gong, Wenxin Yu 0001 |
ISCAS | 4 |
| 2023 | Clock Aware Low Power PlacementabstractIn modern VLSI design, more than 30% of power consumption is caused by clock networks due to their large capacitance demand and high switching frequency. Prior research on clock power optimization primarily focuses on improving the routing and synthesis processes, where the planning flexibility is restricted by the placed registers. In this paper, we develop a novel co-optimization framework to conduct analytic placement and clock tree synthesis simultaneously. A numerical engine is proposed to balance the power demand between clock and regular signal networks by effectively and efficiently updating the clock tree topology, generating the network synthesis solution, and aligning it with the placement objective using the ePlace infrastructure. The experimental results validate the high performance of our proposed algorithm on all eight CLKISPD05 benchmarks, achieving a 45.1% reduction in clock-net wirelength and a 12.7% reduction in total switching power compared to RePlAce. Moreover, our algorithm outperforms the state-of-the-art clock aware placement algorithm SimPL+Lopper, achieving a 25% reduction in clock-net wirelength and a 10.5% reduction in total switching power. Jinghao Ding, Linhao Lu, Zhaoqi Fu, Mengshi Gong, Yuanrui Qi, Wenxin Yu 0001 |
ICCAD | 3 |
| 2022 | An Efficient Maze Routing Algorithm for Fast Global RoutingabstractMaze routing remains the most time-consuming step for modern global routers. Previous works accelerate the maze routing by routing multiple regions or nets simultaneously. This paper presents a novel parallel maze router with bidirectional path search and dynamic routing scheduling, which exhibits higher efficiency than all the previous routers. On the ISPD 2008 benchmark suite, our router outperforms the fastest global routers SPRoute and FastRoute 4.1 by an average speedup of 1.95x and 10.03x, while the difference on the total overflow and wirelength is negligible. Zhaoqi Fu, Wenxin Yu 0001, Xin Cheng 0004 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2022 | Optimal Region-based Mixed-Cell-Height Detailed Placement Considering Complex Minimum-Implant-Area ConstraintsabstractWe propose a minimum-implant-area (MIA) aware detailed placement algorithm for multi-row-height standard cells. Specifically, (1) we calculate the optimal regions for all the cells, and (2) we cluster each group of cells with the same threshold voltage and width less than the minimum implant width and then reshape each cluster. (4) We develop an enhanced legalization algorithm to minimize the total wirelength. (5) We solve the remaining inter-row violations by greedily shifting the concerning cells with minimum displacement. Compared with the state-of-the-art work [6], the experimental results show that on average of all the ISPD 2014 benchmarks A [1] our algorithm reduces the wirelength by 6% and runs 6.02x faster with all the MIA violations resolved. Wenxin Yu 0001, Zhaoqi Fu, Xin Cheng 0004 |
ACM Great Lakes Symposium on VLSI | 3 |