Zhipeng Huang 0009

dblp:229/4173 · DBLP profile ↗
← Back
17ranked-venue papers
1as first author
14since 2021 · last 2026
0009-0004-1762-479XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 1 first-author · 14 since 2021
YearPublicationVenuePosition
2026 AiTPO: KAN-UNet Heterogeneous Network for Timing Prediction and Optimization at Global Routing
abstract
Routing is a critical stage in achieving timing closure in integrated circuit design. Due to the time-consuming flow of detailed routing (DR), the lack of accurate routing information, and the impact of congestion during global routing (GR), rapidly obtaining precise timing information at the global routing stage to guide subsequent timing optimization is a significant challenge. These challenges lead to substantial discrepancies between the estimated timing at GR stage and the actual results after post-DR, resulting in inaccurate evaluations of chip performance. To address this issue, we propose an effective timing prediction and optimization framework, AiTPO. The innovative KAN-UNet heterogeneous timing prediction model effectively combines UNet and KAN networks. By fusing spatial features extracted by UNet with numerical data, the model gains the capability to learn complex relationships across multi-modal data, thereby enhancing robustness and accuracy. Additionally, with the accurate timing evaluation, we introduce two timing optimization strategies during global routing to enhance timing performance. The first strategy involves net ordering based on predicted significant delay nets, prioritizing the routing of more timing-critical nets to reduce detours caused by congestion. The second strategy employs timing estimation to select the most optimal topology from multiple candidates generated by the enhanced A* algorithm, where congestion is considered as a cost factor. Which contributes to optimizing Worst Negative Slack (WNS) and Total Negative Slack (TNS). Experimental results on the real circuits under 28nm process node show that the wire delay prediction accuracy with the proposed KAN-UNet model improves by 34.6% and 25.4% in terms of Mean Absolute Error (MAE) and Max Absolute Error (MaxAE), respectively, compared to GR-based estimations and demonstrate the effectiveness of our timing optimization strategies, which lead to a 2.0% and 4.2% improvement in TNS and WNS, respectively.
Zhisheng Zeng, Simin Tao, Zhipeng Huang 0009, Biwei Xie, Wei Gao 0003
ACM Trans. Design Autom. Electr. Syst.4
2025 Toward Advancing 3D-ICs Physical Design: Challenges and Opportunities
abstract
As the demand for higher integration density and performance efficiency continues to grow, 3D stacking has emerged as a promising solution. In 3D ICs, the complexity of physical design and the optimization space is significantly increasing. Therefore, researching high-quality 3D native instead of pesudo 3D physical design has become even more important. This paper reviews recent advancements and persistent challenges in 3D physical design, focusing on F2F bonding technologies. Then, this paper discusses several issues that still require further research and some overlooked problems, with the hope of helping researchers develop higher-quality 3D native physical design tools in the future.
Xueyan Zhao, Zhisheng Zeng, Zhipeng Huang 0009, Biwei Xie, Yungang Bao
ASP-DAC4
2025 iCTS: Iterative and Hierarchical Clock Tree Synthesis With Skew-Latency-Load Tree
abstract
The advancement of modern clock tree synthesis (CTS) encounters a bottleneck, primarily due to the difficulty in achieving multiobjective co-optimization among complex design processes. To concurrently optimize skew, latency, and load capacitance, we propose an iterative and hierarchical CTS framework, which is composed of clustering, topology generation and routing, buffering, and optimization. First, we introduce a capacitance-based metric to achieve adaptive balanced clustering and optimize the cluster results through simulated annealing. Second, to construct a clock tree with lower latency, load capacitance, and skew, we introduce the skew-latency-load tree (SLLT), which combines the advantages of bound skew tree and Steiner shallow-light tree, and we propose an effective SLLT construction algorithm. Third, to further optimize CTS result by buffering, we introduce the critical wirelength evaluation (CWE) to evaluate the capability of each buffer, and propose the insertion delay estimation (IDE) to reduce the evaluation bias during buffering, then design the iterative skew convergence algorithm (ISCA) to achieve complete convergence of skew. We validate our solution using 28 nm process technology. Compared to our method, the commercial tool increases skew, latency, and clock capacitance by 39.5%, 13.0%, and 18.5%, respectively, while the OpenROAD by 101.6%, 50.7%, and 25.5%, respectively.
Zhipeng Huang 0009, Bei Yu 0001, Wenxing Zhu, Jian Chen 0011, Zhixue He
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 iEDA: An Open-source infrastructure of EDA
abstract
By leveraging the power of open-source software, the EDA tool offers a cost-effective and flexible solution for designers, researchers, and hobbyists alike. Open-source EDA promotes collaboration, innovation, and knowledge sharing within the EDA community. It emphasizes the role of the toolchain in accelerating the development of electronic systems, reducing design costs, and improving design quality. This paper presents an open-source EDA project, iEDA, aiming to build a basic infrastructure for EDA technology evolution and closing the industrial-academic gap in the EDA area. As the foundation for developing EDA tools and researching EDA algorithms and technologies, iEDA is mainly composed of file system, database, manager, operator and interface. To demonstrate the effectiveness of iEDA, we implement and tape out four chips of different scales (from 700k to 500M gates) on different process nodes (110nm and 28nm) with iEDA. iEDA is publicly available on the project home page https://github.com/OSCC-Project/iEDA.
Zengrong Huang, Simin Tao, Zhipeng Huang 0009, Chunan Zhuang, Yihang Qiu, Guojie Luo, Huawei Li 0001, Haihua Shen, Mingyu Chen 0001, Dongbo Bu, Wenxing Zhu, Ye Cai 0001, Xiaoming Xiong, Yi Heng, Peng Zhang 0007, Bei Yu 0001, Biwei Xie, Yungang Bao
ASPDAC4
2024 iPD: An Open-source intelligent Physical Design Toolchain
abstract
Open-source electronic design automation (EDA) shows promising potential in unleashing EDA innovation and lowering the cost of chip design. The open-source EDA toolchain is a comprehensive set of software tools designed to facilitate the design, analysis, and verification of electronic circuits and systems. We developed a physical design EDA toolchain (named iPD) from netlist to GDS-II, including design, analysis, and verification. iPD now covers the whole flow of physical design (including floorplan, placement, clock tree synthesis, routing, timing optimization etc.), part of the analysis tools (timing analysis and power analysis), and part of the verification tools (design rule check). For more friendly support EDA research and development and chip design, we design a reliability, extendibility, ease-of-use, and feature richness physical design toolchain. This paper introduces the software structure, functions, and metrics of the iPD toolchain.
Simin Tao, Shijian Chen, Zhisheng Zeng, Zhipeng Huang 0009, Hongxi Wu, Zengrong Huang, Liwei Ni, Xueyan Zhao, Shuaiying Long, Xiaoze Lin, Fuxing Huang, Yihang Qiu, Zheqing Shao, Jikang Liu, Yuyao Liang, Biwei Xie, Yungang Bao, Bei Yu 0001
ASPDAC5
2024 Toward Controllable Hierarchical Clock Tree Synthesis with Skew-Latency-Load Tree
abstract
Clock tree synthesis (CTS) constructs an efficient clock tree, meeting design constraints and minimizing resource usage. It serves as a bridge between placement and routing, facilitating concurrent optimization of multiple design objectives. To construct a clock tree with lower latency and load capacitance while maintaining a specified skew constraint, we introduce skew-latency-load tree (SLLT) which combines the merits of bound skew tree and Steiner shallow-light tree, along with an analysis and demonstration of the boundaries of these two tree types. We propose a method for constructing SLLT, which significantly reduces both the maximum latency and load capacitance compared to previous methods while ensuring skew control. Combining this routing topology generation method, we introduce a hierarchical CTS framework, and it is constructed by integrating partition schemes and buffering optimization techniques. We validate our solution at 28nm process technology, demonstrating superior performance compared to the solutions of OpenROAD and advanced commercial tool. Our approach outperforms in all metrics (max latency, skew, buffer number, clock capacitance), achieving a significant reduction in latency of 29.45% compared to OpenROAD and 6.75% compared to the commercial tool.
Zhipeng Huang 0009, Bei Yu 0001, Wenxing Zhu
DAC2
2024 Net Resource Allocation: A Desirable Initial Routing Step
abstract
In modern IC design, routing significantly impacts chip performance, power, area, and design iteration count. Critical challenges in routing include generating a rectilinear Steiner minimum tree (RSMT) for each net and handling routing resources among nets. Due to limited resources and net order, congestion is inevitable in VLSI circuit routing. Most competitive routers address congestion after routing without prior net guidance, leading to difficulty in managing resources among nets. We suggest introducing a net resource allocation step to tackle routing and congestion as a potentially desirable initial routing stage. Firstly, we introduce the net region probability density (NRPD) concept to achieve suitable net resource allocation. Using a prior NRPD, we model the resource allocation problem as linear programming (LP). We solve the LP problem and obtain a posterior NRPD for each net on each grid. Based on the posterior NRPD and congestion map, we introduce a cost scheme to guide net routing. This cost scheme supports a weighted RSMT construction technique for better topological solutions. We propose an iterative method for global routing and track assignment, improving detailed routing quality and optimizing design rule violations. Experimental results show the effectiveness of net resource allocation and demonstrate the superior performance of our router over OpenROAD's router across multiple metrics.
Zhisheng Zeng, Jikang Liu, Zhipeng Huang 0009, Ye Cai 0001, Biwei Xie, Yungang Bao
DAC3
2024 Simultaneous Conjugate Gradient and iAFF-UNet for Accurate IR Drop Calculation
abstract
IR drop analysis has become a computationally challenging problem with the shrinking of advanced process nodes. Solving the IR drop problem is time-consuming and an accurate and fast IR drop calculator is crucial for shortening the design cycle. In this work, we introduce an innovative IR drop calculation framework based on the conjugate gradient method and iAFFUNet network. iAFFUNet incorporates the UNet structure with the iterative attention feature fusion (iAFF) blocks to refine conventional approaches of feature concatenation and fusion. iAFF blocks employ multi-scale channel attention modules to enhance feature representation. Furthermore, we leverage intermediate results from the conjugate gradient method as augmented features and utilize graph attention networks for initial value calculation, thereby expediting the iteration process. Alternatively, the matrix operation process can be further accelerated using GPU optimization. During the training phase, we adopt a transfer learning strategy by fine-tuning limited real circuit datasets based on a pre-trained model obtained from training with a substantial amount of synthetic circuit datasets. Experimental results on the ICCAD 2023 contest real hidden testcases under the Nangate 45nm process node show that our model achieves an average improvement of 48.7% and 53.9 % in MAE compared to the contest's champion and the second place, respectively. Additionally, our model achieves a 39.8% reduction in CPU runtime compared to the champion of the contest.
Yipei Xu, Simin Tao, Zhipeng Huang 0009, Biwei Xie, Wei Gao 0003
ICCD4
2024 Instance-level Timing Learning and Prediction at Placement using Res-UNet Network
abstract
Instance level post-routing timing analysis at the placement stage is of great importance for timing optimization such as gate sizing and cell movement etc. Determining the timing bottlenecks accurately and in a fast way has become significantly meaningful for accelerating the timing closure since the time-consuming iteration cycle. In this work, we propose an instance-level timing prediction framework to identify the critical cells of post-routing at the placement stage, which constructs a pixel level image-to-image timing hotspot map translation task using an encoder-decoder based Res-UNet. The network framework combines the strengths of residual learning and basic U-Net, helping in collecting the local and global features of the entire layout of the circuit over different spatial scales. Experimental results on ISCAS’89 benchmark circuits under the 28nm process node demonstrated that with the proposed model, the average prediction accuracy of the critical cells classification achieves 90.7% for unseen designs in terms of the value of the F1-score. Moreover, the framework has achieved a speedup of three orders of magnitude compared with the conventional design flow.
Simin Tao, Zhipeng Huang 0009, Biwei Xie, Ge Li 0002
ISCAS3
2024 AiTO: Simultaneous gate sizing and buffer insertion for timing optimization with GNNs and RL
Hongxi Wu, Zhipeng Huang 0009, Wenxing Zhu
Integr.2
2023 iPL-3D: A Novel Bilevel Programming Model for Die-to-Die Placement
abstract
Die-to-die (D2D) placement is a more challenging stage in achieving higher performance with complex constraints, critically impacting timing, power, yield, cost, etc. Existing placers often rely on indirect objectives (e.g., considering cut sizes in tier assignment), which can lead to a loss of the overall solution space utilization and may even deviate from the actual objective. To address this issue, this paper leverages the natural dominance relationship between decision variables to transform the original problem into a bilevel programming problem equivalently. Additionally, an alternating optimization framework is introduced to enhance the exploration of the overall solution space. On the one hand, we propose two tier optimization operators for simultaneous optimization of wirelength and #terminal in global and detailed perspectives; On the other hand, we present a near-optimal terminal legalization algorithm following an efficient multi-tier co-placement. Compared with the top three winners of the ICCAD'22 contest, our placer achieves 4.33%, 4.42%, and 5.88% smaller wire-length, 79.61 %, 16.74%, and 15.76% fewer #terminal and competitive runtime. Moreover, our placer always uses the fewest #terminal and achieves amazing wirelength reduction when the terminal size changes.
Xueyan Zhao, Shijian Chen, Yihang Qiu, Jiangkao Li, Zhipeng Huang 0009, Biwei Xie, Yungang Bao
ICCAD5
2022 A Robust Global Routing Engine with High-Accuracy Cell Movement under Advanced Constraints
abstract
Placement and routing are typically defined as two separate problems to reduce the design complexity. However, such a divide-and-conquer approach inevitably incurs the degradation of solution quality due to the correlation/objectives of placement and routing are not entirely consistent. Besides, with various constraints (e.g., timing, R/C characteristic, voltage area, etc.) imposed by advanced circuit designs, bridging the gap between placement and routing while satisfying the advanced constraints has become more challenging. In this paper, we develop a robust global routing engine with high-accuracy cell movement under advanced constraints to narrow the gap and improve the routing solution. We first present a routing refinement technique to obtain the convergent routing result based on fixed placement, which provides more accurate information for subsequent cell movement. To achieve fast and high-accuracy position prediction for cell movement, we construct a lookup table (LUT) considering complex constraints/objectives (e.g., routing direction and layer-based power consumption), and generate a timing-driven gain map for each cell based on the LUT. Finally, based on the prediction, we propose an alternating cell movement and cluster movement scheme followed by partial rip-up and reroute to optimize the routing solution. Experimental results on the ICCAD 2020 contest benchmarks show that our algorithm achieves the best total scores among all published works. Compared with the champion of the ICCAD 2021 contest, experimental results on the ICCAD 2021 contest benchmarks show that our algorithm achieves better solution quality in shorter runtime.
Ziran Zhu, Fuheng Shen, Yangjie Mei, Zhipeng Huang 0009, Jianli Chen
ICCAD4
2022 Novel Proximal Group ADMM for Placement Considering Fogging and Proximity Effects
abstract
Fogging and proximity effects (FPEs) are two major factors that cause inaccurate exposure and layout pattern distortions in e-beam lithography. In this article, we propose an analytical placement algorithm that considers both FPEs. We formulate the global placement problem as a separable minimization problem with linear constraints, where different objectives can be tackled one by one in an alternating fashion. Then, we propose a novel proximal group alternating direction method of multipliers (ADMM) to solve the separable minimization problem with two subproblems, where the first subproblem (associated with wirelength and density) is solved by the steepest descent method without line search, and the second one (associated with the FPEs) is handled by an analytical scheme. We prove the property of global convergence of the proximal group ADMM method. Finally, the FPEs-aware legalization and detailed placement are employed to legalize and improve the placement result. The experimental results show that our algorithm is effective and efficient for the addressed problem. Our algorithm achieved 5.7% smaller fogging variation, 6.8% lower proximity variation, and 5.4% lower runtime with a minor wirelength overhead compared with the state-of-the-art work.
Jianli Chen, Zhipeng Huang 0009, Ziran Zhu, Zheng Peng 0002, Wenxing Zhu, Yao-Wen Chang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Late Breaking Results: An Effective Legalization Algorithm for Heterogeneous FPGAs with Complex Constraints
abstract
The modern FPGA placement problem has become much more challenging than ever with various emerging design constraints, such as the location (including relative location (RLOC) and location range) and chain constraints, which have not been considered in the literature. In this paper, we propose a combinatorial algorithm for FPGA legalization with location and chain constraints. We first identify a virtual range to cluster an instance with the RLOC constraint and formulate minimum cost integer linear programming. Besides, we use an adaptive algorithm to deal with the chain-aware legalization problem for better quality and runtime trade-offs. Finally, a legalization algorithm based on minimum cost maximum flow (MCMF) is used to improve the solution quality further. Compared with the state-of-the-art work, experimental results show that our proposed algorithm can achieve respectively 4.4% and 4.9% smaller average and maximum movements, 1.7% smaller routed wirelength, and 7.6% shorter routing runtime
Zhipeng Huang 0009, Ziran Zhu, Jun Yu 0010, Jianli Chen
DAC1
2020 An Efficient EPIST Algorithm for Global Placement with Non-Integer Multiple-Height Cells *
abstract
With the increasing design requirements of modern circuits, a standard-cell library often contains cells of different row heights to address various trade-offs among performance, power, and area. However, maintaining all standard cells with integer multiples of a single-row height could cause some area overheads and increase power consumption. In this paper, we present an analytical placer to directly consider a circuit design with non-integer multiple-height standard cells and additional layout constraints. The region of different cell heights is adaptively generated by the global placement result. In particular, an exact penalty iterative shrinkage and thresholding (EPIST) algorithm is employed to efficiently optimize the global placement problem. The convergence of the algorithm is proved, and the acceleration strategy is proposed to improve the performance of our algorithm. Compared with the state-of-the-art works, experimental results based on the 2017 CAD Contest at ICCAD benchmarks show that our algorithm achieves the best wirelength and area for every benchmark. In particular, our proposed EPIST algorithm provides a new direction for effectively solving large-scale nonlinear optimization problems with non-smooth terms, which are often seen in real-world applications.
Jianli Chen, Zhipeng Huang 0009, Wenxing Zhu, Jun Yu 0010, Yao-Wen Chang
DAC2
2020 Mixed-cell-height legalization considering complex minimum width constraints and half-row fragmentation effect
Ziran Zhu, Zhipeng Huang 0009, Wenxing Zhu, Jianli Chen, Hanbin Zhou, Senhua Dong
Integr.2
2018 Analytical solution of Poisson's equation and its application to VLSI global placement
abstract
Poisson's equation has been used in VLSI global placement for describing the potential field induced by a given charge density distribution. Unlike previous global placement methods that solve Poisson's equation numerically, in this paper, we provide an analytical solution of the equation to calculate the potential energy of an electrostatic system. The analytical solution is derived based on the separation of variables method and an exact density function to model the block distribution in a placement region, which is an infinite series and converges absolutely. Using the analytical solution, we give a fast computation scheme of Poisson's equation and develop an effective and efficient global placement algorithm called Pplace. Experimental results show that our Pplace achieves smaller placement wirelength than ePlace and NTUplace3, two leading wirelength-driven placers. With the pervasive applications of Poisson's equation in scientific fields, in particular, our effective, efficient, and robust computation scheme for its analytical solution can provide substantial impacts to these fields.
Wenxing Zhu, Zhipeng Huang 0009, Jianli Chen, Yao-Wen Chang
ICCAD2