Mengshi Gong

dblp:348/4524 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0005-8984-0923ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Timing-Driven Detailed Placement with Collaborative Topology Reconstruction
abstract
Placement is a critical step in the physical design, as it largely determines the potential for subsequent optimization. In this work, we propose a timing-driven detailed placement framework: first, a simplified RC-tree model is employed for flip-flop–buffer compensation; then, a gradient-augmented global heuristic algorithm is incorporated; and finally, timing improvement is achieved through local collaborative optimization. A comprehensive evaluation on eight ICCAD 2015 benchmarks demonstrates the effectiveness of our approach. Compared to DREAMPlace4.0-DP, a state-of-the-art timing-driven placer, our framework achieves an average improvement of 19.60% in WNS and 55.74% in TNS, while introducing less disturbance to the global placement. Moreover, it delivers a 0.80% reduction in HPWL and reduces runtime by 20.24%.
Zhengjie Zhao, Wenxin Yu 0001, Mengshi Gong, Youzhi Zheng, Xinmiao Li, Wenyu Liu 0018, Jingwei Lu
DATE4
2025 An Effective and Efficient Cross-Link Insertion for Non-Tree Clock Network Synthesis
abstract
Clock skew introduces significant challenge to the overall system performance. Existing non-tree solutions like cross-link insertion often come with limitations, such as the over-consumption of resource and power. In this work, we propose a cross-link insertion algorithm that effectively reduces the clock skew with minimal power overhead, and prioritize delay optimization on the paths with high sensitivity to the skew. The experimental results from the ISPD 2010 benchmarks show a 17% reduction in the mean of clock skew, a 45% decrease in the standard deviation of clock skew, and a 13% lower power consumption versus the advanced non-tree solutions in literature.
Jinghao Ding, Jiazhi Wen, Zhaoqi Fu, Mengshi Gong, Yuanrui Qi, Wenxin Yu 0001, Jinjia Zhou
DATE5
2025 An Effective Macro Placement Framework with Reinforcement Learning and Monte Carlo Tree Search
Jinghao Ding, Wenxin Yu 0001, Yuanrui Qi, Zhaoqi Fu, Mengshi Gong, I-Chyn Wey, Jinjia Zhou
ACM Great Lakes Symposium on VLSI5
2024 Track Assignment Using Gradient Indication and Simulated Annealing
abstract
Track assignment has become a crucial step within the physical design flow. In this work, we proposed a novel track assignment approach using gradients estimation and simulated annealing. Specifically, we formulate the overlap cost function and derive the gradients to estimate the quality of candidate positions for iroutes movement. A simulated annealer is devised to consume the gradients indication and perturb the track assignment solution from the global perspective. Experimental results show 5.15% lower overlap cost in average of all 10 benchmarks from DAC 2012 Routability-Driven Placement Contest, compared to the state-of-the-art approach NTA [1].
Yuanrui Qi, Zejun Gan, Jinghao Ding, Zhaoqi Fu, Mengshi Gong, Wenxin Yu 0001
ISCAS5
2023 An Efficient and Robust Algorithm for Common Path Pessimism Removal In Static Timing Analysis
abstract
Common path pessimism removal (CPPR) is an essential step in modern static timing analysis (STA) to avoid unnecessary circuit overdesign caused by extra pessimism. However, current CPPR approaches exhibit good performance yet poor scalability. This work proposes an efficient algorithm based on timing graph pruning to address the CPPR scalability issue. On average, on industry benchmarks from TAU 2014 CAD contest, our algorithm runs 23.1% faster versus OpenTimer, the state-of-the-art open-source STA engine in literature.
Mengshi Gong, Wenxin Yu 0001
ACM Great Lakes Symposium on VLSI1
2023 Clock Aware Low Power Placement
abstract
In modern VLSI design, more than 30% of power consumption is caused by clock networks due to their large capacitance demand and high switching frequency. Prior research on clock power optimization primarily focuses on improving the routing and synthesis processes, where the planning flexibility is restricted by the placed registers. In this paper, we develop a novel co-optimization framework to conduct analytic placement and clock tree synthesis simultaneously. A numerical engine is proposed to balance the power demand between clock and regular signal networks by effectively and efficiently updating the clock tree topology, generating the network synthesis solution, and aligning it with the placement objective using the ePlace infrastructure. The experimental results validate the high performance of our proposed algorithm on all eight CLKISPD05 benchmarks, achieving a 45.1% reduction in clock-net wirelength and a 12.7% reduction in total switching power compared to RePlAce. Moreover, our algorithm outperforms the state-of-the-art clock aware placement algorithm SimPL+Lopper, achieving a 25% reduction in clock-net wirelength and a 10.5% reduction in total switching power.
Jinghao Ding, Linhao Lu, Zhaoqi Fu, Mengshi Gong, Yuanrui Qi, Wenxin Yu 0001
ICCAD5