Zhen Zhuang

dblp:230/7755 · DBLP profile ↗
← Back
28ranked-venue papers
9as first author
24since 2021 · last 2026
0000-0002-2972-8770ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 8 first-author · 21 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Partitioning-free 3D-IC Floorplanning
abstract
3D integration with fine-pitch hybrid bonding offers a promising path to alleviate interconnect bottlenecks in conventional two-dimensional (2D) ICs, yet efficient 3D floorplanning remains challenging due to the enlarged solution space and non-uniform inter-die communication latency. Existing methods either extend 2D representations into 3D, leading to combinatorial complexity, or adopt partitioning-first pipelines that fix block-to-die assignments early and hinder joint optimization of floorplan, die assignment, and vertical connectivity. In this work, we present \textsc{Great3D}, a partitioning-free 3D floorplanning framework that directly optimizes a native 3D floorplan. \textsc{Great3D} formulates a unified objective that couples interconnect cost with a cycles-per-instruction (CPI)-derived latency term to capture the system-level impact of face-to-face (F2F) bonding. Algorithmically, it combines an SDP-based 3D global embedding with a dynamic-programming refinement for die assignment, followed by 2D continuous refinement with practical design constraints. \textcolor{blue}{Experiments on the GSRC and ATPlace benchmark suites show that \textsc{Great3D} consistently achieves strong wirelength and CPI quality against state-of-the-art 3D floorplanners. On GSRC, it reduces total wirelength by up to about $70\%$ (and by $2.40$--$2.74\times$ on average) over competing 3D-native floorplanners, and its dynamic-programming die-assignment stage further improves CPI by $9.5$--$17.8\%$, while maintaining competitive runtime on instances of up to a few hundred blocks.}
Shuo Ren 0001, Zhen Zhuang, Rongliang Fu, Leilei Jin, Libo Shen, Bei Yu 0001, Tsung-Yi Ho
ASP-DAC2
2026 Graph Attention-Based Current Crowding Analysis at TSV Interfaces in 3D Power Delivery Networks
Zhen Zhuang, Tsung-Yi Ho, Sung Kyu Lim
ASP-DAC2
2026 DPO-3D: Differentiable Power Delivery Network Optimization via Flexible Modeling for Routability and IR-Drop Tradeoff in Face-to-Face 3D ICs
Zhen Zhuang, Yuxuan Zhao 0001, Bei Yu 0001, Sung Kyu Lim, Tsung-Yi Ho
ASP-DAC1
2026 FastRW: An Efficient Random Walk Method for Steady-State Thermal Analysis
abstract
Thermal simulation is increasingly critical in modern IC design and manufacturing. Random walk methods based on the Feynman-Kac formula enable efficient local temperature estimation without computing the full temperature field. However, in practical scenarios without Dirichlet boundary conditions, these methods often require excessively long paths and heuristic truncation rules. In this work, we revisit Feynman-Kac sampling and derive an exact characterization of the truncation error: the expected residual contribution is a simple scalar multiple of the temperature at the truncation point. This insight leads to FastRW, a random-walk framework that safely applies aggressive truncation. FastRW uses a cheap, noisy prior temperature field to approximate the residual term and shorten individual paths, and further exploits cross-relations among query points through a Bayesian posterior update to reduce the number of required walks. Experiments on 3DIC steady-state thermal benchmarks show that FastRW achieves over 6× speedup over prior Feynman-Kac-based methods with better accuracy.
Zixiao Wang 0001, Tianshu Hou, Zhen Zhuang, Tsung-Yi Ho, Farzan Farnia, Bei Yu 0001
DATE4
2026 ETLA-3D: Equivalent Thin Layer Aggregation based Thermal FEM for Hybrid Bonding F2F 3D ICs
abstract
In 3D face-to-face (F2F) hybrid bonding ICs, sub-micrometer thin layers lead to an extreme aspect ratio between the lateral dimensions and the vertical thickness. This poses major challenges for finite element method (FEM) thermal simulation. To address this, we introduce ETLA-3D, a thermal FEM methodology based on equivalent thin-layer aggregation, designed specifically for hybrid bonding F2F 3D ICs. The method consolidates the physical properties of thin layers into their neighboring layers by introducing new integral terms into the FEM weak form, greatly reducing the complexity of meshing, the simulation degrees of freedom (DoFs) and the computational cost, while preserving accuracy. Experimental results show that ETLA-3D achieves up to 695.8 × faster runtime compared to the commercial FEM tool (COMSOL Multiphysics), with a maximum absolute error of less than 1.1°C. By combining high accuracy with exceptional efficiency, ETLA-3D establishes a reliable and efficient FEM framework to model the thermal behavior of F2F 3D ICs.
Zhen Zhuang, Darong Huang 0003, Luis Costero, Rongmei Chen, David Atienza 0001, Tsung-Yi Ho
DATE2
2026 Parallel Delay-Driven Layer Assignment Leveraging Hierarchical Task Graph Modeling for Advanced Technology Nodes
abstract
Very large scale integration (VLSI) circuits typically consist of millions of nets, posing significant challenges for efficient physical design. Interconnect delay has become a critical factor for timing performance in technology nodes at 5nm and beyond. Additionally, the coupling effect among the wires increases the complexity of delay optimization. Moreover, tapering constraints are essential in advanced technology nodes to ensure manufacturability. Furthermore, the ever-increasing scale of modern designs necessitates a high-performance computing (HPC) framework to accelerate delay-driven layer assignment in advanced technology nodes. To address these challenges, we propose ParDelay, a parallel delay-driven layer assignment leveraging hierarchical task graph modeling while considering tapering constraints for advanced technology nodes, which includes the following five key techniques: 1) A general deterministic parallel framework is proposed for delay-driven layer assignment, leveraging a hierarchical task graph to enable both internet and inter-node parallelism. 2) A delay- and overflow-driven tapering repairing strategy is proposed to eliminate tapering violations while further optimizing net delay. 3) A local delay-critical net filtering method is proposed to analyze local delay criticality to guide layer assignment, thereby minimizing delay while eliminating overflow. 4) To mitigate the coupling effect, we propose a net shielding algorithm that reduces wire density for maximum delay candidate nets to optimize maximum delay. 5) A delay-aware refinement strategy is proposed to classify nets by their delay rank and assign distinct non-default-rule (NDR) wire permissions and refinement objectives, thereby reducing delay. Experimental results demonstrate that, compared to existing layer assignment algorithms and parallel routing frameworks, our approach effectively reduces delay, via count, and runtime under the tapering constraints.
Zhen Zhuang, Genggeng Liu, Wen-Hao Liu 0001, Tsung-Yi Ho, Ting-Chi Wang
IEEE Trans. Computers2
2026 HiePlace: Efficient Hierarchical PCB Placement
abstract
Due to the rapid expansion of printed circuit board (PCB) designs, accompanied by diverse design rules and specific constraints, there has been a substantial increase in manual design engineering efforts. To address this challenge, industries are seeking productivity improvements through automated placement techniques. However, existing placers primarily target VLSI placement and do not align well with PCBs’ unique characteristics. This mismatch arises from both the customization of PCBs and the complexity of the problem, which involves considering various constraints such as priorities, irregularities, and alignment. This paper introduces HiePlace, an efficient mathematical programming (MP)-based placement framework designed explicitly for PCBs. It aims to address the diverse constraints and achieve better performance. To address the issue of time-consuming computation in the direct MP-based algorithm, we present two innovative acceleration techniques: (1) In the initial stage, we introduce a dynamic programming approach to prioritize the placement of core components. This technique effectively reduces the solution space and enhances the overall placement quality. (2) Additionally, we propose a relaxation algorithm to minimize the number of boolean variables and further narrow down the solution space. This approach enables more efficient placement results by considering the problems specific constraints. Experimental results show that the proposed framework produces 7.7× speed up and 66% cost reduction.
Shanyi Li, Zhen Zhuang, Weihua Sheng, Bei Yu 0001, Tsung-Yi Ho
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2026 Multilayer Package Power/Ground Planes Synthesis With Balanced DC IR Drops: A Game-Theoretic Optimization Approach
abstract
Recently, the challenge of integrating an increasing number of transistors on a single die to adhere to Moores Law has spurred the need for innovative packaging solutions. Power/ground planes are integral to packages, and designers typically strive to maximize their size. This provides shielding and maintains constant impedance for adjacent high-speed signal wires, benefiting signal integrity. Additionally, large power/-ground planes help reduce DC IR drops, enhancing power integrity. However, the necessity for multiple power/ground nets, each requiring independent power/ground planes within a package, makes the optimal allocation of limited free space a complex task. This paper introduces a game-theoretic optimization method aimed at evenly mitigating DC IR drops across the multi-layer package power/ground planes. In the formulated game of achieving the ideal power/ground plane design, we can enhance the use of package space and realize a design with evenly distributed DC IR drops across all power/ground planes. This is accomplished by adjusting strategies and reaching a state of Nash equilibrium in the allocation of free space. Additionally, we propose a rapid multi-layer power/ground plane DC IR drop evaluation and a power/ground plane legalization method to bolster our optimization method.
Siyuan Liang 0002, Zhen Zhuang, Kai-Yuan Chao, Bei Yu 0001, Tsung-Yi Ho
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2026 Cracking the Code of Backdoor Attacks With Confidence Consistency
abstract
It is widely believed that increasing the amount of training data enhances the intelligence of deep learning models, which in turn heightens dependence on external datasets. However, these datasets are susceptible to adversarial poisoning, allowing attackers to insert backdoors that trigger misclassifications. Although various defense strategies have been suggested, training-phase defenses (TPDs) appear most promising, as they can significantly lower the attack success rate (ASR) without greatly impacting model performance. Nevertheless, creating TPDs that achieve both high accuracy and low ASR is challenging due to two main issues: 1) Many solutions require additional clean samples that match the distribution of the poisoned dataset, which is not always practical in real-world scenarios; 2) most existing solutions have high computational costs, sometimes requiring five to ten times the expense of standard training, which severely limits their practical use. To tackle these challenges, we introduce Confidence Consistency Detection (CCD), an efficient and lightweight training-phase backdoor detection method. CCD is particularly advantageous in situations where clean data is scarce or unavailable, as it completely eliminates the need for external clean samples. Moreover, CCD significantly reduces computational costs to just 25% to 50% of existing solutions (1.7 times the standard training time), providing a notable improvement over current TPD methods. The core innovation of CCD lies in its ability to utilize the high confidence shown by backdoor samples during the early stages of model training for precise detection. Specifically, we initially train a model on a poisoned dataset for a few epochs, followed by intra-class loss fine-tuning to increase sensitivity to poisoned samples. We then create preliminary sets of poisoned and clean samples by assessing the consistency of confidence variations before and after fine-tuning. These sets guide the model training, enabling the detection of high-confidence poisoned samples. Extensive experiments demonstrate that CCD effectively reduces the attack success rate (ASR) to 1.43%, while having a negligible impact on the model’s clean accuracy. In detecting poisoned samples, CCD achieves a 99% true positive rate (TPR) and a 0.033% false positive rate (FPR), setting a new benchmark in the field.
Zhen Zhuang, Yijian Ding, Jun Shao 0001, Donghai Zhu, Huiyan Wang 0002
IEEE Trans. Inf. Forensics Secur.1
2026 Adaptive Redistribution Layer Routing for Chiplet-Package Co-Design in 2.5D System
abstract
2.5D packaging has become a popular alternative to integrate advanced logic and memory chiplets for high-performance computing and artificial intelligence systems. In the conventional design flow, chiplets and packages are independently designed and then integrated at the assembly stage. To bridge the gap between chiplet designs and package designs, existing chiplet-package co-design methods iteratively optimize chiplet layouts to improve the performance of the entire system. However, Redistribution Layer (RDL) routing, which finishes the interconnections between chiplets at the package level and significantly affects the system performance, is neglected in the existing co-design flows. Therefore, this article proposes an effective chiplet-package co-design flow focusing on the RDL routing to optimize the package system performance dynamically. The proposed co-design flow can fill in the missing link, package-level co-optimization, of previous design flows. In the proposed co-design flow, we propose an efficient RDL routing algorithm to iteratively optimize the substrate layout based on the cross-boundary timing context extracted from both chiplets and the package. The proposed RDL routing algorithm has two critical techniques, including (1) a Maximal Independent Set-based (MIS-based) pin assignment method to dynamically optimize the pin positions of nets and (2) a network-flow-based router to generate routing layouts. Experimental results show that the proposed design flow can gradually improve the maximum frequency of a real design to the target performance, 400 MHz.
Zhen Zhuang, Weishiun Hung, M. D. Arafat Kabir, Yarui Peng, Tsung-Yi Ho
ACM Trans. Design Autom. Electr. Syst.1
2025 Fast Routing Algorithm for Mask Stitching Region of Ultra Large Wafer Scale Integration
abstract
Interposer-based packaging has gained tremendous popularity in integrating advanced logic and memory chiplets for artificial intelligence and high-performance computing systems. The size of the silicon interposer is the critical bottleneck in improving the performance of integrated systems by mounting more and more advanced chiplets, such as high bandwidth memory (HBM). Nowadays, ultra large wafer scale integration is a popular alternative to integrate large amounts of advanced chiplets on a big wafer scale silicon interposer. However, wafer scale silicon interposers cannot be manufactured by one mask due to the reticle limitation. Therefore, the mask stitching technique is used to manufacture ultra large systems by applying multiple masks for different sub-regions of an ultra large silicon interposer. To achieve the alignment of two adjacent sub-regions manufactured by different masks, the two sub-regions have an overlapped stitching region. Previous algorithms cannot handle the special design rules of mask stitching regions and are not efficient enough to generate high-quality routing solutions. In this work, we propose a fast routing algorithm for mask stitching regions to efficiently solve the special design rules. The time complexity of the proposed algorithm is O(n log n), where n is the number of nets. Compared with state-of-the-art work, our algorithm can achieve 100% routability with an effective reduction of wirelength. Furthermore, the proposed algorithm can achieve a speedup of thousands of times.
Zhen Zhuang, Quan Chen 0007, Hao Yu 0001, Tsung-Yi Ho
ASP-DAC1
2025 Accelerating k-means ++ Algorithm
Jiehao Liang, Somdeb Sarkhel, Zhao Song 0002, Chenbo Yin 0002, Zhen Zhuang, Danyang Zhuo
IEEE Big Data5
2025 GNN-MLS: Signal Routing in Mixed-Node 3D ICs through GNN-Assisted Metal Layer Sharing
abstract
Native 3D Integrated Circuit (3D IC) design offers enhanced performance and density but faces challenges in signal routing due to limited true 3D EDA tool support. Pseudo-3D flows bridge this gap but lack cross-tier optimization, critical for both mixed-node and homogeneous designs. Metal Layer Sharing (MLS) addresses this by enabling cross-tier routing co-optimization but risks timing degradation if not applied strategically. Additionally, MLS creates open connections in hybrid-bonded 3D ICs, making chips untestable. We propose GNN-MLS, a Graph Neural Network-based framework for precise MLS net selection, combined with a tailored DFT solution for robust testability. Experiments show GNN-MLS reduces timing violations by 79% and improves WNS and TNS by 81% and 94%, and moves designs closer to true 3D ICs.
Pruek Vanna-Iampikul, Zhen Zhuang, Tsung-Yi Ho, Sung Kyu Lim
DAC3
2025 ChronoTE: Crosstalk-Aware Timing Estimation for Routing Optimization via Edge-Enhanced GNNs
abstract
Accurate timing estimation during the routing stage is critical for modern VLSI design closure, especially under increasing crosstalk effects in advanced technology nodes. During the routing process, the crosstalk effect is usually modeled by predicting coupling capacitance with congestion information. However, such estimations are often overly pessimistic, as crosstalk-induced delay is influenced not only by coupling capacitance but also by the relative arrival times of signals. In this work, we propose ChronoTE, a novel edge-enhanced graph neural network (GNN) framework that performs crosstalk-aware net delay estimation by jointly modeling physical topology and timing characteristics. By embedding timing-window-aware features into edge representations, ChronoTE enables accurate delay prediction without requiring full routing or parasitic extraction. Experimental results on industrial-scale open-source designs demonstrate that ChronoTE, by delivering sign-off quality delay estimation in the early global routing stage, significantly accelerates design closure and contributes to area reduction.
Leilei Jin, Rongliang Fu, Zhen Zhuang, Liang Xiao 0001, Fangzhou Liu 0005, Bei Yu 0001, Tsung-Yi Ho
ICCAD3
2025 MMPack: Multi-Mask Co-Design for Ultra-Large Wafer-Scale Package Integration
abstract
Interposer-based packaging has emerged as a pivotal technology for integrating advanced logic and memory chiplets in artificial intelligence (AI) and high-performance computing (HPC) systems. To accommodate growing system complexity, ultra-large wafer-scale integration employs expanded silicon interposers to support more chiplets. However, manufacturing such interposers exceeds the limits of single-mask lithography, requiring mask stitching, a technique that introduces unique physical design constraints and structural discontinuities. Additionally, thermo-mechanical stress, particularly near through-silicon vias (TSVs) and stitching regions, poses critical reliability challenges that conventional floorplanning methods fail to address. This paper presents MMPack, a hierarchical analytical framework for multi-mask chiplet-package co-design. Our approach integrates three key innovations: (1) a performance-driven partitioning algorithm that minimizes inter-chiplet and inter-mask communication overhead; (2) a stitching-aware hierarchical floorplanning strategy based on alternating optimization to address mask boundary constraints; and (3) a stress-aware post-processing step that employs an analytical model to reduce critical stress concentrations while preserving floorplanning quality. Experimental results demonstrate that MMPack significantly enhances both architectural performance and mechanical reliability while maintaining efficient layout and runtime scalability. These results highlight the practicality of our framework for enabling robust, high-performance designs in next-generation wafer-scale integration systems.
Shanyi Li, Zhen Zhuang, Siyuan Liang 0002, Bei Yu 0001, Tsung-Yi Ho
ICCAD2
2025 ML-Based Fine-Grained Modeling of DC Current Crowding in Power Delivery TSVs for Face-to-Face 3D ICs
Zhen Zhuang, Bei Yu 0001, Tsung-Yi Ho, Martin D. F. Wong, Sung Kyu Lim
ISPD2
2025 Hierarchical Partitioning-Based Interchip Redistribution Layer Routing for Fan-Out Wafer-Level Packaging
Haoyang Xu, Xing Huang 0001, Zhen Zhuang, Zhiwen Yu 0001, Bin Guo 0001, Kai-Yuan Chao, Bei Yu 0001, Tsung-Yi Ho, Martin D. F. Wong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 Floorplet: Performance-Aware Floorplan Framework for Chiplet Integration
abstract
A chiplet is an integrated circuit (IC) that encompasses a well-defined subset of an overall systems functionality. In contrast to traditional monolithic system-on-chips (SoCs), chipletbased architecture can reduce costs and increase reusability, representing a promising avenue for continuing Moore’s Law. Despite the advantages of multi-chiplet architectures, floorplan design in a chiplet-based architecture has received limited attention. Conflicts between cost and performance necessitate a trade-off in chiplet floorplan design since additional latency introduced by advanced packaging can decrease performance. Consequently, balancing performance, cost, area, and reliability is of paramount importance. To address this challenge, we propose Floorplet (Floorplan chiplet), a framework comprising simulation tools for performance reporting and comprehensive models for cost and reliability optimization. Our framework employs the open-source Gem5 simulator to establish the relationship between performance and floorplan for the first time, guiding the floorplan optimization of multi-chiplet architecture. The experimental results show that our method decreases inter-chiplet communication costs by 24.81%.
Shixin Chen, Shanyi Li, Zhen Zhuang, Su Zheng, Zheng Liang 0003, Tsung-Yi Ho, Bei Yu 0001, Alberto L. Sangiovanni-Vincentelli
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 A Robust Multilayer X-Architecture Global Routing System Based on Particle Swarm Optimization
abstract
Global routing is an extremely important stage of very large scale integration (VLSI) physical design. With the rise of nano-scale integrated circuit design, the multilayer global routing problem has attracted considerable research interest during the past few years. In this article, a multilayer X-architecture global routing (ML-XGR) system based on particle swarm optimization (PSO), called FZU-Router, is proposed to solve the ML-XGR problem for the first time. FZU-Router contains a multilayer X-architecture integer linear programming (MX-ILP) model and a multilayer X-architecture PSO (MX-PSO) algorithm, which are presented to formulate and solve the ML-XGR problem, respectively. Moreover, four effective strategies are designed to enhance the efficiency of FZU-Router: 1) a strategy for generating new routing modes is proposed to strengthen the robustness of encoding strategy of MX-PSO; 2) a strategy for combining MX-PSO with maze routing is proposed to improve the routability; 3) a strategy for reducing the channel capacity is proposed to make better use of optimization ability of MX-PSO; and 4) a strategy for dynamic resource assignment is proposed to make better use of routing resources and shorten the running time. Experimental results on multiple benchmarks confirm that the proposed FZU-Router leads to fewer total overflow and shorter total wirelength compared with the state-of-the-art routers.
Genggeng Liu, Zhen Zhuang, Zhenyu Pei, Min Gan, Xing Huang 0001, Wenzhong Guo
IEEE Trans. Syst. Man Cybern. Syst.3
2023 Multi-Product Optimization for 3D Heterogeneous Integration with D2W Bonding
abstract
3D heterogeneous integration enables the integration of multiple heterogeneous chiplets into the same package with the effective reduction of package size and interconnection latency. According to the market requirement, chiplets with robust re-usability and effective cost reduction can be selected from a library to form different package products for enlarging total profit. Since die-to-wafer (D2W) bonding enables the chiplets with different sizes to be bonded in a package, it is a more flexible option for 3D heterogeneous integration compared with the conventional wafer-to-wafer (W2W) bonding. However, this promising technique creates new issues, including 1) flexible chiplet bonding enabling more than one chiplet to be bonded with a base chiplet to construct multiple products and 2) degraded bonding leading to the degradation of performance. In this work, a distributed integer-linear-programming-based (ILP-based) method is proposed to efficiently maximize the profits of multiple package products considering the issues of cost-addition 3D heterogeneous integration with D2W bonding. Compared with the baseline, the distributed ILP-based method can achieve the best profits while achieving a 5.96X speedup. To the best of our knowledge, this is the first work to solve the multi-product optimization problem for 3D heterogeneous integration with D2W bonding.
Zhen Zhuang, Kai-Yuan Chao, Bei Yu 0001, Tsung-Yi Ho, Martin D. F. Wong
ICCAD1
2022 TRADER: A Practical Track-Assignment-Based Detailed Router
abstract
As the last stage of VLSI routing, detailed routing should consider complicated design rules in order to meet the manufacturability of chips. With the continuous development of VLSI technology node, the design rules are changing and increasing which makes detailed routing a hard task. In this paper, we present a practical track-assignment-based detailed router to deal with the most representative design rules in modern designs. The proposed router consists of four major stages: (1) a graph-based track assignment algorithm is proposed to optimize the design rule violations of an entire die area; (2) an effective rip-up and reroute method is used to reduce the design rule violations in local regions; (3) a segment migration algorithm is proposed to reduce short violations; and (4) a stack via optimization technique is proposed to reduce minimum area violations. Practical benchmarks from 2019 ISPD contest are used to evaluate the proposed router. Compared with the state-of-the-art detailed router, Dr. CU 2.0, the number of violations can be reduced by up to 35.11 % with an average reduction rate of 10.08 %. The area of short can be reduced by up to 61.49 % with an average reduction rate of 44.80 %.
Zhen Zhuang, Genggeng Liu, Tsung-Yi Ho, Bei Yu 0001, Wenzhong Guo
DATE1
2022 Multi-Package Co-Design for Chiplet Integration
abstract
Due to the cost and design complexity associated with advanced technology nodes, it is difficult for traditional monolithic System-on-Chip to follow the Moore's Law, which means the economic benefits have been weakened. Semiconductor industries are looking for advanced packages to improve the economic advantages. Since the multi-chiplet architecture supporting heterogeneous integration has the robust re-usability and effective cost reduction, chiplet integration has become the mainstream of advanced packages. Nowadays, the number of mounted chiplets in a package is continuously increasing with the requirement of high system performance. However, the large area caused by the increasing of chiplets leads to the serious reliability issues, including warpage and bump stress, which worsens the yield and cost. The multi-package architecture, which can distribute chiplets to multiple packages and use less area of each package, is a popular alternative to enhance the reliability and reduce the cost in advanced packages. However, the primary challenge of the multi-package architecture lies in the tradeoff between the inter-package costs, i.e., the interconnection among packages, and the intra-package costs, i.e., the reliability caused by warpage and bump stress. Therefore, a co-design methodology is indispensable to optimize multiple packages simultaneously to improve the quality of the whole system. To tackle this challenge, we adopt mathematical programming methods in the multi-package co-design problem regarding the nature of the synergistic optimization of multiple packages. To the best of our knowledge, this is the first work to solve the multi-package co-design problem.
Zhen Zhuang, Bei Yu 0001, Kai-Yuan Chao, Tsung-Yi Ho
ICCAD1
2022 SPTA: A Scalable Parallel ILP-Based Track Assignment Algorithm with Two-Stage Partition
abstract
Routability has always been a very challenging issue in Very Large Scale Integrated (VLSI) circuit design. The routability is considered in track assignment so that the global routing results can better match the requirements of detailed routing. However, existing heuristic track assignment algorithms are prone to local optimality, which cannot provide the accurate routability estimation. To overcome this limitation, we propose a scalable parallel Integer Linear Programming (ILP)-based track assignment algorithm, called SPTA, which employs a two-stage partition strategy. First, by taking into account both the global and local nets, all wires are assigned to tracks, making full use of the information from the global routing results. Second, an efficient ILP model for track assignment is proposed to minimize the overlap between iroutes1, thus significantly improving routability. Third, a two-stage partition strategy is designed to reduce the runtime. Finally, a panel-subpanel-level parallelism is proposed to further speed up the algorithm without sacrificing the quality of the solutions. Experimental results show that SPTA has a better routability estimation compared with the existing algorithms.
Yidan Jing, Liliang Yang, Zhen Zhuang, Genggeng Liu, Xing Huang 0001, Wen-Hao Liu 0001, Ting-Chi Wang
VLSI-SoC3
2021 ALIFRouter: A Practical Architecture-Level Inter-FPGA Router for Logic Verification
abstract
As the scale of VLSI circuits increases rapidly, multi-FPGA prototyping systems have been widely used for logic verification. Due to the limited number of connections between FPGAs, however, the routability of prototyping systems is a bottleneck. As a consequence, timing division multiplexing (TDM) technique has been proposed to improve the usability of prototyping systems, but it causes a dramatic increase in system delay. In this paper, we propose ALIFRouter, a practical architecture-level inter-FPGA router, to improve the chip performance by reducing the corresponding system delay. ALIFRouter consists of three major stages, including i) routing topology generation, ii) TDM ratio assignment, and iii) system delay optimization. Additionally, a multi-thread parallelization method is integrated into the three stages to improve the efficiency of ALIFRouter. With the proposed algorithm, major performance indicators of multi-FPGA systems such as signal multiplexing ratio can be improved significantly.
Zhen Zhuang, Xing Huang 0001, Genggeng Liu, Wenzhong Guo, Weikang Qian, Wen-Hao Liu 0001
DATE1
2020 MiniDelay: Multi-Strategy Timing-Aware Layer Assignment for Advanced Technology Nodes
abstract
Layer assignment, a major step in global routing of integrated circuits, is usually performed to assign segments of nets to multiple layers. Besides the traditional optimization goals such as overflow and via count, interconnect delay plays an important role in determining chip performance and has been attracting much attention in recent years. Accordingly, in this paper, we propose MiniDelay, a timing-aware layer assignment algorithm to minimize delay for advanced technology nodes, taking both wire congestion and coupling effect into account. MiniDelay consists of the following three key techniques: 1) a non-default-rule routing technique is adopted to reduce the delay of timing critical nets, 2) an effective congestion assessment method is proposed to optimize delay of nets and via count simultaneously, and 3) a net scalpel technique is proposed to further reduce the maximum delay of nets, so that the chip performance can be improved in a global manner. Experimental results on multiple benchmarks confirm that the proposed algorithm leads to lower delay and few vias, while achieving the best solution quality among the existing algorithms with the shortest runtime.
Xinghai Zhang, Zhen Zhuang, Genggeng Liu, Xing Huang 0001, Wen-Hao Liu 0001, Wenzhong Guo, Ting-Chi Wang
DATE2
2020 MSFRoute: Multi-Stage FPGA Routing for Timing Division Multiplexing Technique
abstract
As the scale of VLSI circuits and fabrication costs increase rapidly, multi-FPGA prototyping systems are widely adopted in industry to make logic verification faster and cheaper. Since routing signals can usually exceed the number of I/O pins in an FPGA, timing division multiplexing (TDM) technique is required to solve this problem. FPGA routing for developing a prototyping system is a big challenge due to the signal delay of TDM. This paper presents MSFRoute, a multi-stage FPGA routing framework for timing division multiplexing technique, to optimize the signal delay and the routability for prototyping systems. In this work, a TDM ratios assignment algorithm with an efficient parallelization method is proposed to optimize inter-FPGA signal delay. Meanwhile, we propose a practical system clock period optimization method to solve critical signal delay problem. Experimental results show that our routing framework reduces TDM ratios by up to 88.3% with an average reduction rate of 41.8%. With the proposed parallelization method, total flow of MSFRoute can get up to 4.38X speedup with a 2.77X speedup on average.
Zhen Zhuang, Genggeng Liu, Xing Huang 0001, Xiaotao Jia, Wen-Hao Liu 0001, Wenzhong Guo
ACM Great Lakes Symposium on VLSI1
2020 A unified algorithm based on HTS and self-adapting PSO for the construction of octagonal and rectilinear SMT
Genggeng Liu, Zhisheng Chen 0002, Zhen Zhuang, Wenzhong Guo
Soft Comput.3
2019 RDTA: An Efficient Routability-Driven Track Assignment Algorithm
abstract
This paper presents a routability-driven track assignment algorithm (RDTA) to efficiently estimate routability. Routability has become a very challenging issue in modern IC design and it can be effectively estimated by routing congestion. Track assignment is a stage towards bridging the gap between global routing and detailed routing. And track assignment can also analyze routing congestion more accurate than other routing stages like global routing. Some works have used track assignment to estimate routability. In this work, wire segments extracted from a global routing solution are assigned to proper tracks considering local nets, via locations and pin access. The overlap between wire segments extracted from a global routing solution can be effectively reduced by the algorithm so that the solution of RDTA can effectively estimate the routability.
Genggeng Liu, Zhen Zhuang, Wenzhong Guo, Ting-Chi Wang
ACM Great Lakes Symposium on VLSI2