Ziran Zhu

dblp:193/3138 · DBLP profile ↗
← Back
40ranked-venue papers
9as first author
32since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 37 · 8 first-author · 29 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Comprehensive Delay-Aware Net Weighting Framework for Timing-Driven Global Placement
Lixin Chen, Keyu Peng, Jinghui Zhou, Shuting Cai, Ziran Zhu
ASP-DAC7
2026 A Timing-Driven Hierarchical Macro Placement Framework for Large-Scale Complex IP Blocks
Lixin Chen, Jinghui Zhou, Ziran Zhu
ASP-DAC4
2026 GPU-Accelerated Global Routing with Balanced Timing and Congestion Optimization
abstract
As integrated circuit (IC) designs continue to scale in complexity, global routing faces increasing challenges in managing timing and congestion simultaneously. This paper proposes a GPU-accelerated global routing framework that effectively balances timing optimization and congestion mitigation. The proposed framework begins with a preprocessing stage that partitions ultralarge nets, followed by timing path construction, and decomposes nets based on estimated pin slack to enhance scalability and timing sensitivity. For critical nets, we propose a timing and congestion driven GPU-accelerated hybrid 3D pattern routing method. Specifically, an Elmore-based timing weight calculation method is proposed to efficiently capture the timing criticality of routing paths, and the resulting weights are ordered for more targeted and effective timing optimization. Then, a well-designed cost scheme is proposed to better balance timing and congestion. Finally, we develop a GPU-accelerated hybrid 3D pattern routing strategy that combines L -shape and sparse Z -shape patterns to improve routing efficiency. After routing critical nets, the remaining non-critical nets are routed using a congestion-driven GPU-accelerated routing engine that supports flexible detours to alleviate congestion and utilize residual routing resources. Compared with the champion of the ISPD 2025 contest, experimental results on the ISPD 2025 contest benchmarks show that our algorithm achieves 19.4% better weighted scores and 1 6. 2 % faster runtime.
Jinghui Zhou, Fuxing Huang, Lixin Chen, Xinglin Zheng, Ziran Zhu
ASP-DAC5
2026 Late Breaking Results: RL-Based Macro Placement with Cell Clustering and Rudy Modeling for Routability Optimization
Youwen Wang, Xinglin Zheng, Keyu Peng, Ziran Zhu
DATE5
2025 Late Breaking Results: Customized Diffusion Model Empowered by Heterogeneous Graph Network for Effective Floorplanning
abstract
Floorplanning is a critical phase in VLSI physical design, focusing on determining block positions while optimizing wirelength under specified area constraints. However, classical analytical-based floorplanners are highly sensitive to the quality of initial solutions and existing learningbased methods often suffer from high computational inefficiency and complexity. In this paper, we propose a customized diffusion model to directly generate high-quality initial floorplans. By leveraging a classical analytical-based floorplanner on top of this initial floorplan, the final floorplanning results are significantly improved. To enhance feature extraction, a heterogeneous graph neural network (HGNN) is developed to explicitly incorporate block-to-block and pin-to-block relationships from the netlist during the diffusion process. Additionally, a novel guidance sampling function is introduced to optimize both wirelength and overlap, effectively reducing the required sampling steps while maintaining competitive initial solutions. Experimental results demonstrate that integrating our proposed diffusion model with an advanced analytical-based floorplanner achieves at least 4.8% reduction in runtime and 3.0% reduction in HPWL compared to the original floorplanner and other diffusion-based methods.
Xinglin Zheng, Keyu Peng, Youwen Wang, Wenxing Zhu, Ziran Zhu
DAC6
2025 Comprehensive Placement and Routing Framework with Guaranteed In-Cell Routability for Synthesizing Complementary-FET Cells
abstract
As the technology node advances beyond 5 nm, the conventional FinFET architecture encounters substantial scaling issues. ComplementaryFET (CFET) technology, characterized by the vertical stacking of P-FET over N-FET or vice versa, has emerged as a promising solution. However, the inherent characteristics of CFET architecture, particularly the scarcity of routing resources, pose significant obstacles to in-cell routability and layout generation. In this paper, we develop a comprehensive placement and routing framework for synthesizing CFET cells. We first present a partitioning technique followed by a heuristic quality maintenance strategy for large-scale cells to ensure scalability and efficiency. Then, we propose a novel satisfiability modulo theories (SMT)-based placement method that incorporates partial routing to achieve minimum-width placement while ensuring in-cell routability. Particularly, the placement method also determines the pin positions for each net, which simplifies subsequent routing complexity. Finally, we propose a progressive metal routing method to address the challenges of routing resource scarcity and unidirectional routing in CFET technology, which includes a manual-inspired M0 routing followed by an integral linear programming (ILP)-based M1 and M2 routing. Compared with the state-of-the-art CFET cell generators, experimental results show that our algorithm achieves the smallest cell width for all tested cells, with 7 out of 30 cells exhibiting smaller widths. For the remaining 23 cells, which have the same cell width as those in other generators, our algorithm achieves the smallest M2 usage and total metal length.
Zhengzhe Zheng, Yinuo Wu, Keyu Peng, Ziran Zhu
DAC5
2025 Multiscale Feature Attention and Transformer Based Congestion Prediction for Routability-Driven FPGA Macro Placement
abstract
As routability has emerged as a critical task in modern field-programmable gate array (FPGA) physical design, it is desirable to develop an effective congestion prediction model during the placement stage. Given that the interconnection congestion level is a critical metric for measuring the routability of FPGA placement, we utilize that level as the model training label. In this paper, we propose a multiscale feature attention (MFA) and transformer based congestion prediction model to extract placement features and strengthen their association with congested areas for effective FPGA macro placement. A convolutional neural network (CNN) component is first designed to extract multiscale features from grid-based placement. Then, a well-designed MFA block is proposed that utilizes the dual attention mechanism on both spatial and channel dimensions to enhance the representation of each multiscale feature. By incorporating MFA blocks and CNN's output at each skip connection layer, our model substantially enhances its capability to learn features and recover more precise congestion level maps. Furthermore, multiple transformer layers that employ dynamic attention mechanisms are utilized to extract global information, which can significantly improve the difference between various congestion levels and enhance the ability to identify these levels. Based on the ten most congested and challenging benchmarks from the MLCAD 2023 FPGA macro placement contest, experimental results show that our model outperforms existing congestion prediction models. Furthermore, our model can achieve the best routability and score among the contest winners when integrated into the macro placer based on DREAMPlaceFPGA.
Xinglin Zheng, Youwen Wang, Keyu Peng, Ziran Zhu
DATE5
2025 DiSPlace: Diffusion-Sharing-Driven Transistor-Level Placement Beyond Standard-Cell Boundaries for DTCO
abstract
As the increasing demands of design technology co-optimization (DTCO) in advanced nodes, the rigid configurations of standard cells impose significant limitations on wirelength and area optimization. A more flexible alternative is to place transistors directly on the design canvas, allowing for precise transistor-level adjustments that reduce wirelength and minimize design area. In this paper, we propose DiSPlace, a novel diffusion-sharing-driven transistor-level placement algorithm beyond standard-cell boundaries to fully leverage DTCO. We first present an in-cell placement based transistor pairing method to pair PMOS and NMOS transistors with the same gate net, followed by incorporating Gaussian perturbations to generate an initial placement. Then, we propose the first diffusion-sharing-driven global placement framework. It begins with the construction of diffusion sharing nets to guide transistor placement, followed by an analytical model for simultaneously optimizing diffusion sharing, wirelength, and density. Besides, a nonlinear optimization with adaptive penalty adjustment is presented to solve the analytical model effectively and efficiently. Finally, we develop a satisfiability modulo theories (SMT)-based detailed placement method to optimize design area and wirelength while ensuring legal placement. A diffusion-sharing-aware partitioning technique is also developed to enhance the scalability and efficiency of the SMT-based method. Compared to a standard-cell-based placer and the state-of-the-art transistor-level placer, our algorithm achieves significant improvements, reducing wirelength by 18% and 11%, and design area by 24% and 4%, respectively. These results highlight the effectiveness of DiSPlace in achieving high-quality placements for transistor-level designs.
Keyu Peng, Yinuo Wu, Zhengzhe Zheng, Ziran Zhu, Chao Wang 0068, Jun Yang 0006
ICCAD5
2025 NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
abstract
Compositional training has been the de-facto paradigm in existing Multimodal Large Language Models (MLLMs), where pre-trained vision encoders are connected with pre-trained LLMs through continuous multimodal pre-training. However, the multimodal scaling property of this paradigm remains difficult to explore due to the separated training. In this paper, we focus on the native training of MLLMs in an end-to-end manner and systematically study its design space and scaling property under a practical setting, i.e., data constraint. Through careful study of various choices in MLLM, we obtain the optimal meta-architecture that best balances performance and training cost. After that, we further explore the scaling properties of the native MLLM and indicate the positively correlated scaling relationship between visual encoders and LLMs. Based on these findings, we propose a native MLLM called NaViL, combined with a simple and cost-effective recipe. Experimental results on 14 multimodal benchmarks confirm the competitive performance of NaViL against existing MLLMs. Besides that, our findings and results provide in-depth insights for the future study of native MLLMs.
Changyao Tian, Hao Li 0069, Gen Luo, Xizhou Zhu, Weijie Su 0002, Hanming Deng, Jinguo Zhu, Ziran Zhu, Lewei Lu, Wenhai Wang, Hongsheng Li 0001, Jifeng Dai
NeurIPS9
2025 Two stage Ordered Escape Routing combined with LP and heuristic algorithm for large scaled PCB
Disi Lin, Chuandong Chen, Rongshan Wei, Qinghai Liu, Ziran Zhu, Zhifeng Lin, Jianli Chen
Integr.6
2025 Routability-Driven Macro Placement Engine for Modern FPGAs With Complex Cascade Shape and Region Constraints
abstract
Field-programmable gate array (FPGA) macro placement holds a crucial role within the FPGA physical design flow since it substantially influences the subsequent stages of cell placement and routing. With the increasing number of macros and the complex cascade shape and region constraints imposed by modern FPGAs, the routability and macro placement have become much more challenging. In this paper, we propose an effective and efficient routability-driven macro placement algorithm for modern FPGAs with cascade shape and region constraints. To reserve adequate space for cell placement and guarantee routability, we first develop a routability-driven mixed-size analytical global placement that evenly distributes both macros and cells while considering cascade shape and region constraints. Particularly, the proposed global placement engine integrates a well-trained congestion prediction model, targeting benchmarks with high routing congestion to enhance overall routability. Then, we propose an integer linear programming (ILP)-based cascade shape legalization followed by matching-based macro legalization to remove macro overlaps while satisfying the region constraints. Finally, a routability-driven detailed macro placement is proposed to refine the solution. Compared with the winners of the MLCAD 2023 FPGA macro placement contest and state-of-the-art works, experimental results show that our algorithm achieves the best overall score and routability.
Keyu Peng, Jianli Chen, Jun Yang 0006, Ziran Zhu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2025 Dual Multimodal Fusions With Convolution and Transformer Layers for VLSI Congestion Prediction
abstract
In very large scale integration (VLSI) circuit physical design, precise congestion prediction during placement is crucial for enhancing routability and accelerating design processes. Existing congestion prediction models often encounter challenges in handling multimodal information and lack effective fusion of placement and netlist features, limiting their prediction accuracy. In this article, we present a novel congestion prediction model that leverages dual multimodal fusions with convolution and transformer layers to effectively capture the multiscale placement information and enhance congestion prediction accuracy. We first adopt convolutional neural networks (CNNs) to extract grid-based placement features and heterogeneous graph convolutional networks (HGCNs) to extract netlist information. To help the model understand the correlation between different modalities, we then propose an early feature fusion (EFF) to integrate netlist knowledge into multiscale placement features at multimodal interaction subspace. Besides, a deep feature fusion (DFF) method is proposed to further fuse multimodal features, which has multiple vision transformer layers based on adaptive attention enhancement technology. These layers include self-attention (SA) to boost intramodal features and cross-attention (CA) to perform cross-modal feature fusion on netlist and grid-based placement features. Finally, the output features of DFF are sent into the cascaded decoder to recover the congestion map by exploiting several upsampling layers and merging with EFF features. Compared with the existing state-of-the-art congestion prediction models, experimental results demonstrate that our model not only outperforms them in prediction accuracy, but also excels in reducing routing congestion when integrated into the placer DREAMPlace.
Youwen Wang, Xinglin Zheng, Keyu Peng, Ziran Zhu, Jianli Chen, Jun Yang 0006
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 Variational Label-Correlation Enhancement for Congestion Prediction
abstract
As the complexity of Integrated Circuits (ICs) rises, accurate routing and congestion prediction, crucial for identifying early design flaws, become essential to expedite circuit design and conserve resources in the lengthy physical design process. Despite the advancements in current congestion prediction methodologies, an essential aspect that has been largely overlooked is the spatial label-correlation between different grids in congestion prediction. The spatial label-correlation is a fundamental characteristic of circuit design, where the congestion status of a grid is not isolated but inherently influenced by the conditions of its neighboring grids. In order to fully exploit the inherent spatial label-correlation between neighboring grids, we propose a novel approach, VALCE, i.e., VAriational Label-Correlation Enhancement for Congestion Prediction, which considers the local label-correlation in the congestion map, associating the estimated congestion value of each grid with a local label-correlation weight influenced by its surrounding grids. VALCE leverages variational inference techniques to estimate this weight, thereby enhancing the regression model’s performance by incorporating spatial dependencies. Experiment results validate the superior effectiveness of VALCE on the public available ISPD2011 and DAC2012 benchmarks using the superblue circuit line.
Congyu Qiao, Ning Xu 0009, Xin Geng 0001, Ziran Zhu, Jun Yang 0006
ASPDAC5
2024 Late Breaking Results: Routability-Driven FPGA Macro Placement Considering Complex Cascade Shape and Region Constraints
abstract
Field-programmable gate array (FPGA) macro placement holds a crucial role within the FPGA physical design flow since it substantially influences the subsequent stages of cell placement and routing. In this paper, we propose an effective and efficient routability-driven macro placement algorithm for modern FPGAs with cascade shape and region constraints. To reserve adequate space for cell placement and guarantee routability, we first develop a routability-driven mixed-size analytical global placement (GP) that evenly distributes both macros and cells while considering cascade shape and region constraints. Then, we propose an integer linear programming (ILP)-based cascade shape legalization (LG) followed by matching-based macro legalization to remove macro overlaps while satisfying the region constraints. Finally, a routability-driven detailed macro placement is proposed to refine the solution. Compared with the top contestants of the MLCAD 2023 contest, experimental results show that our algorithm achieves the best overall score and routability.
Keyu Peng, Jun Yang 0006, Ziran Zhu
DAC5
2024 Idempotence and Perceptual Image Compression
abstract
Idempotence is the stability of image codec to re-compression. At the first glance, it is unrelated to perceptual image compression. However, we find that theoretically: 1) Conditional generative model-based perceptual codec satisfies idempotence; 2) Unconditional generative model with idempotence constraint is equivalent to conditional generative codec. Based on this newfound equivalence, we propose a new paradigm of perceptual image codec by inverting unconditional generative model with idempotence constraints. Our codec is theoretically equivalent to conditional generative codec, and it does not require training new models. Instead, it only requires a pre-trained mean-square-error codec and unconditional generative model. Empirically, we show that our proposed approach outperforms state-of-the-art methods such as HiFiC and ILLM, in terms of Fréchet Inception Distance (FID). The source code is provided in https://github.com/tongdaxu/Idempotence-and-Perceptual-Image-Compression.
Tongda Xu, Ziran Zhu, Dailan He, Yanghao Li, Zhe Wang 0070, Hongwei Qin, Yan Wang 0105, Ya-Qin Zhang
ICLR2
2024 Noise Dimension of GAN: An Image Compression Perspective
abstract
Generative adversial network (GAN) is a type of generative model that maps a high-dimensional noise to samples in target distribution. However, the dimension of noise required in GAN is not well understood. Previous approaches view GAN as a mapping from a continuous distribution to another continous distribution. In this paper, we propose to view GAN as a discrete sampler instead. From this perspective, we build a connection between the minimum noise required and the bits to losslessly compress the images. Furthermore, to understand the behaviour of GAN when noise dimension is limited, we propose divergence-entropy trade-off. This trade-off depicts the best divergence we can achieve when noise is limited. And as rate distortion trade-off, it can be numerically solved when source distribution is known. Finally, we verifies our theory with experiments on image generation.
Ziran Zhu, Tongda Xu, Yan Wang 0105
ICME1
2024 A fast and high-performance global router with enhanced congestion control
Xiqiong Bai, Yilu Chen, Zhifeng Lin, Zhijie Cai, Ziran Zhu, Jianli Chen
Integr.6
2024 An effective routability-driven packing algorithm for large-scale heterogeneous FPGAs
Zijun Li 0005, Ziran Zhu, Jianli Chen
Integr.2
2024 High-Performance Placement Engine for Modern Large-Scale FPGAs With Heterogeneity and Clock Constraints
abstract
As field-programmable gate array (FPGA) architectures continue to evolve and become more complex, the heterogeneity and clock constraints imposed by modern FPGAs have posed significant challenges to FPGA placement. This article proposes a high-performance placement engine for modern large-scale FPGAs with heterogeneity and clock constraints. To improve efficiency and scalability, we develop a clustering method considering both internal/external connectivity and the balance of block types to build the hierarchy. In each hierarchy level, we propose a hybrid penalty and augmented Lagrangian method (HPALM) to convert the FPGA global placement with heterogeneity and clock constraints into a series of unconstrained optimization subproblems, then use the Adam method to solve each subproblem. In particular, we prove that the HPALM is globally convergent for global placement. Besides, a matching-based IP block legalization is developed to legalize the DSPs and RAMs, and a multistage packing is presented to cluster LUTs and FFs into HCLBs. Finally, we propose a history-based legalization to legalize CLBs in an FPGA, and a simulated-annealing-based detailed placement is presented to reduce the wirelength while maintaining legality. Compared with the state-of-the-art works, experimental results based on the ISPD 2017 contest benchmarks show that the proposed algorithm can achieve the shortest routed wirelength in a reasonable runtime.
Ziran Zhu, Yangjie Mei, Kangkang Deng, Jianli Chen, Jun Yang 0006, Yao-Wen Chang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2024 Subgraph matching-based reference placement for printed circuit board designs
Ziran Zhu, Miaodi Su, Haiyuan Su, Yifeng Xiao, Jianli Chen, Yao-Wen Chang
J. Supercomput.1
2023 Efficient Global Optimization for Large Scaled Ordered Escape Routing
abstract
Ordered Escape Routing (OER) problem, which is an NP-hard problem, is critical in PCB design. Primary methods based on integer linear programming (ILP) or heuristic algorithms work well on small-scale PCBs with fewer pins. However, when dealing with large-scale instances, the performance of ILP strategies suffers dramatically as the number of variables increases due to time-consuming preprocessing. As for heuristic algorithms, ripping-up and rerouting is adopted to increase resource utilization, which frequently causes time violation. In this paper, we propose an efficient ILP-based routing engine for dense PCB to simultaneously minimize wiring length and runtime, considering the specific routing constraints. By weighting the length, we first model the OER problem as a special network flow problem. Then we separate the non-crossing constraint from typical ILP modeling to reduce the number of integral variables greatly. In addition, considering the congestion of routing resources, the ILP method is proposed to detect congestion. Finally, unlike the traditional schemes that deal with negotiated congestion, our approach works by reducing the local area capacity and then allowing the global automatic optimization of congestion. Compared with the state-of-the-art work, experimental results show that our algorithm can solve cases in larger scale in high routing quality of less length and reduce routing time by 76%.
Chuandong Chen, Dishi Lin, Rongshan Wei, Qinghai Liu, Ziran Zhu, Jianli Chen
ASP-DAC5
2023 Disjoint-Path and Golden-Pin Based Irregular PCB Routing with Complex Constraints
abstract
PCB routing becomes time-consuming as the complexity of PCB design increases. Unlike traditional schemes that treat the two essential PCB routing processes separately, namely, escape and bus routing, we consider the continuity between them and present a golden-pin-based routing scheme to find the desired solution with angle and topology constraints. Further, conventional rip-up and reroute methods are often ineffective and inefficient for congestion alleviation and routability optimization. We construct a component graph by modeling components as vertices and applying the minimum weight vertex covering method to improve the routability. A self-adaptable ordering method is presented for escape routing to arrange the pin order on the component boundary, guaranteeing successful bus routing. In addition, escape routing is performed based on a disjoint path method. We construct a dynamic Hanan grid in bus routing and utilize a novel congestion adjustment technique to improve solution quality. Compared with FreeRouting and Allegro, the experiment results show that our algorithm achieves high routability and a significant 90% runtime reduction.
Qinghai Liu, Qinfei Tang, Jiarui Chen, Chuandong Chen, Ziran Zhu, Jianli Chen, Yao-Wen Chang
DAC5
2022 Voronoi Diagram Based Heterogeneous Circuit Layout Centerline Extraction for Mask Verification
abstract
Modern circuit layout centerline extraction is an essential step in estimating the parasitic inductance and verifying the layout performance in mask verification. As the continued feature size shrinking and the complexity of modern circuit design keeps growing, heterogeneous layout centerline extraction has become even more challenging. In this paper, we first formulate a Voronoi diagram-based problem transformation to collect all centerline points. Then, a graph-based initial centerline generation algorithm is presented to handle all invalid centerline points effectively. Finally, a heterogeneity-aware centerline optimization method is proposed to generate optimized design-violation-free centerline results for irregular structures. Compared with the state-of-the-art commercial 3D-RC parasitic parameter extraction tool RCExplorer and the 1st place in the 2019 EDA Elite Challenge Contest, experimental results show that our algorithm achieves the best average precision ratio of 99.7% on centerline extraction while satisfying all design constraints.
Xiqiong Bai, Ziran Zhu, Jianli Chen, Jun Yu 0010, Yao-Wen Chang
ASP-DAC2
2022 Subgraph matching based reference placement for PCB designs: late breaking results
abstract
Reference placement is promising to handle the increasing complexity in PCB design. We model the netlist into a graph and use a subgraph matching algorithm to find the isomorphism of the placed template in component combination to reuse the placement. The state-of-the-art VF3 algorithm can achieve high matching accuracy while suffering from high computation time in large-scale instances. Thus, we propose the D2BS algorithm to guarantee matching quality and efficiency. We build and filter the candidate set (CS) according to designed features to construct the CS structure. In the CS optimization, a graph diversity tolerance strategy is adopted to achieve inexact matching. Then, hierarchical match is developed to search the template embeddings in the CS structure guided by branch backtracking and matched nodes snatching. Experimental results show that D2BS outperforms VF3 in accuracy and runtime, achieving 100% accuracy on PCB instances.
Miaodi Su, Yifeng Xiao, Haiyuan Su, Ziran Zhu, Jianli Chen, Yao-Wen Chang
DAC7
2022 High-performance placement for large-scale heterogeneous FPGAs with clock constraints
abstract
With the increasing complexity of the field-programmable gate array (FPGA) architecture, heterogeneity and clock constraints have greatly challenged FPGA placement. In this paper, we present a high-performance placement algorithm for large-scale heterogeneous FPGAs with clock constraints. We first propose a connectivity-aware and type-balanced clustering method to construct the hierarchy and improve the scalability. In each hierarchy level, we develop a novel hybrid penalty and augmented Lagrangian method to formulate the heterogeneous and clock-aware placement as a sequence of unconstrained optimization subproblems and adopt the Adam method to solve each unconstrained optimization subproblem. Then, we present a matching-based IP blocks legalization to legalize the RAMs and DSPs, and a multi-stage packing technique is proposed to cluster FFs and LUTs into HCLBs. Finally, history-based legalization is developed to legalize CLBs in an FPGA. Based on the ISPD 2017 clock-aware FPGA placement contest benchmarks, experimental results show that our algorithm achieves the smallest routed wirelength for all the benchmarks among all published works in a reasonable runtime.
Ziran Zhu, Yangjie Mei, Zijun Li 0005, Jingwen Lin, Jianli Chen, Jun Yang 0006, Yao-Wen Chang
DAC1
2022 A Robust Global Routing Engine with High-Accuracy Cell Movement under Advanced Constraints
abstract
Placement and routing are typically defined as two separate problems to reduce the design complexity. However, such a divide-and-conquer approach inevitably incurs the degradation of solution quality due to the correlation/objectives of placement and routing are not entirely consistent. Besides, with various constraints (e.g., timing, R/C characteristic, voltage area, etc.) imposed by advanced circuit designs, bridging the gap between placement and routing while satisfying the advanced constraints has become more challenging. In this paper, we develop a robust global routing engine with high-accuracy cell movement under advanced constraints to narrow the gap and improve the routing solution. We first present a routing refinement technique to obtain the convergent routing result based on fixed placement, which provides more accurate information for subsequent cell movement. To achieve fast and high-accuracy position prediction for cell movement, we construct a lookup table (LUT) considering complex constraints/objectives (e.g., routing direction and layer-based power consumption), and generate a timing-driven gain map for each cell based on the LUT. Finally, based on the prediction, we propose an alternating cell movement and cluster movement scheme followed by partial rip-up and reroute to optimize the routing solution. Experimental results on the ICCAD 2020 contest benchmarks show that our algorithm achieves the best total scores among all published works. Compared with the champion of the ICCAD 2021 contest, experimental results on the ICCAD 2021 contest benchmarks show that our algorithm achieves better solution quality in shorter runtime.
Ziran Zhu, Fuheng Shen, Yangjie Mei, Zhipeng Huang 0009, Jianli Chen
ICCAD1
2022 Timing-Aware Fill Insertions With Design-Rule and Density Constraints
abstract
Metal fill insertion has become an essential step in reducing dielectric thickness variation and improving pattern uniformity, which is important in mitigating process variations, thereby achieving better manufacturing yield. However, metal fills could induce coupling capacitance, which is not often considered in existing works that typically focus more on pattern density uniformity, incurring significant problems in timing closure. However, it is a great challenge to consider three types of capacitances (i.e., area, fringe, and lateral capacitances) with design rules and density constraints at the fill insertion stage simultaneously. This article presents an efficient timing-aware fill insertion algorithm for minimizing the total capacitance and fill amount, considering the density constraints. First, we present an initial metal fill insertion and design-rule-aware legalization to obtain an initial fill insertion solution quickly. Second, from critical conductors to powers/grounds in a circuit, we divide conductors into different equivalent paths and then construct a capacitance graph to reduce the capacitance of each equivalent path globally. Third, we propose a density-aware coupling capacitance optimization method and a fast Monte Carlo-based fill selection to further reduce the coupling capacitance between any pair of conductors. Finally, we present a density-aware fill deletion method to reduce the fill amount. We evaluate the performance of our algorithm on the benchmarks of the 2018 CAD Contest at ICCAD and its official contest evaluator. Compared with the first-place team of the contest and the state-of-the-artwork, experimental results show that our algorithm achieves the lowest total capacitance and the least fill amount in a comparable runtime.
Xiqiong Bai, Ziran Zhu, Jianli Chen, Tingshen Lan, Jun Yu 0010, Wenxing Zhu, Yao-Wen Chang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 Novel Proximal Group ADMM for Placement Considering Fogging and Proximity Effects
abstract
Fogging and proximity effects (FPEs) are two major factors that cause inaccurate exposure and layout pattern distortions in e-beam lithography. In this article, we propose an analytical placement algorithm that considers both FPEs. We formulate the global placement problem as a separable minimization problem with linear constraints, where different objectives can be tackled one by one in an alternating fashion. Then, we propose a novel proximal group alternating direction method of multipliers (ADMM) to solve the separable minimization problem with two subproblems, where the first subproblem (associated with wirelength and density) is solved by the steepest descent method without line search, and the second one (associated with the FPEs) is handled by an analytical scheme. We prove the property of global convergence of the proximal group ADMM method. Finally, the FPEs-aware legalization and detailed placement are employed to legalize and improve the placement result. The experimental results show that our algorithm is effective and efficient for the addressed problem. Our algorithm achieved 5.7% smaller fogging variation, 6.8% lower proximity variation, and 5.4% lower runtime with a minor wirelength overhead compared with the state-of-the-art work.
Jianli Chen, Zhipeng Huang 0009, Ziran Zhu, Zheng Peng 0002, Wenxing Zhu, Yao-Wen Chang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 Mixed-Cell-Height Placement With Drain-to-Drain Abutment and Region Constraints
abstract
Along with device scaling, the drain-to-drain abutment (DDA) and fence region constraints arise as emerging challenges in modern circuit designs, incurring additional difficulties, especially for designs with mixed-cell-height standard cells which have prevailed in advanced technology. This article presents the first work to address the mixed-cell-height placement problem considering the DDA and fence region constraints from post-global placement throughout the detailed placement. Our algorithm consists of three major stages: 1) preprocessing; 2) legalization; and 3) detailed placement. At the preprocessing stage, we align cells to the desired rows that meet the region constraint, considering the total cell displacement and the distribution ratio of source nodes to drain nodes simultaneously. After deciding the cell ordering of every row, we first propose an interval concept to handle fixed macros and fence regions and then apply the robust modulus-based matrix splitting iteration method to remove all cell overlaps with minimized total displacement at the legalization stage. For detailed placement, unlike the existing works that can handle the DDA constraint only for single rows, we propose a satisfiability-based approach that considers the whole layout to fix the DDA violations more effectively. Besides, we further present an integer linear program (ILP)-based method to optimize the cell displacement without increasing the DDA violations. Compared with a shortest-path method, experimental results show that our proposed algorithm can significantly reduce cell violations, average cell displacement, and maximum cell displacement, in a comparable runtime.
Jianli Chen, Ziran Zhu, Longkun Guo, Yu-Wei Tseng, Yao-Wen Chang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Late Breaking Results: An Effective Legalization Algorithm for Heterogeneous FPGAs with Complex Constraints
abstract
The modern FPGA placement problem has become much more challenging than ever with various emerging design constraints, such as the location (including relative location (RLOC) and location range) and chain constraints, which have not been considered in the literature. In this paper, we propose a combinatorial algorithm for FPGA legalization with location and chain constraints. We first identify a virtual range to cluster an instance with the RLOC constraint and formulate minimum cost integer linear programming. Besides, we use an adaptive algorithm to deal with the chain-aware legalization problem for better quality and runtime trade-offs. Finally, a legalization algorithm based on minimum cost maximum flow (MCMF) is used to improve the solution quality further. Compared with the state-of-the-art work, experimental results show that our proposed algorithm can achieve respectively 4.4% and 4.9% smaller average and maximum movements, 1.7% smaller routed wirelength, and 7.6% shorter routing runtime
Zhipeng Huang 0009, Ziran Zhu, Jun Yu 0010, Jianli Chen
DAC4
2021 Late Breaking Results: Heterogeneous Circuit Layout Centerline Extraction for Mask Verification
abstract
With the continued feature-size shrinking in modern circuit designs, the layout performance estimation and parasitic import calculation based on the extracted centerline result play an important role in mask verification. Most previous works on layout centerline extraction focus on identifying the connectivity among the devices in a mask layout, with few ones collecting accurate centerline information for mask verification while considering design constraints. In this paper, we first formulate the centerline extraction problem as a Voronoi diagram to collect centerline points. Then, we present a graph-based invalid centerline removal algorithm to generate an initial centerline result. Finally, a complexity-driven centerline optimization method is proposed to further optimize the centerline while considering design constraints. Compared with the commercial 3D-RC parasitic parameter extraction tool RCExplorer and the 1st place in the 2019 EDA Elite Challenge Contest, experimental results show that our algorithm achieves the highest average precision ratio of 99.8% on centerline extraction while satisfying all design constraints in the shortest runtime.
Xiqiong Bai, Ziran Zhu, Lichong Sun, Jianli Chen
DAC2
2021 A Robust Modulus-Based Matrix Splitting Iteration Method for Mixed-Cell-Height Circuit Legalization
abstract
Modern circuits often contain standard cells of different row heights to meet various design requirements. Taller cells give larger drive strengths and higher speed at the cost of larger areas and power. Multi-row height standard cells incur challenging issues for layout designs, especially the mixed-cell-height legalization problem with heterogeneous cell structures. Honoring the good cell positions from global placement, we present in this article a robust modulus-based matrix splitting iteration method (RMMSIM) to solve the mixed-cell-height legalization problem. Fixing the cell ordering from global placement and relaxing the right-boundary constraints, our proposed method first converts the problem into an equivalent linear complementarity problem (LCP), and then properly splits the matrices in the LCP so that the RMMSIM can solve the LCP optimally. The RMMSIM effectively explores the sparse characteristic of a circuit, and takes only linear time per iteration; as a result, it can solve the QP very efficiently. Finally, an allocation scheme for illegal cells is used to align such cells to placement sites on rows and fix the placement of out-of-right-boundary cells, if any. Experimental results show the effectiveness and efficiency of our proposed algorithm. In addition, the RMMSIM convergence and optimality are theoretically proved and empirically validated. In particular, this article provides a new RMMSIM formulation for various optimization problems that require solving large-scale convex quadratic programming problems efficiently.
Jianli Chen, Ziran Zhu, Wenxing Zhu, Yao-Wen Chang
ACM Trans. Design Autom. Electr. Syst.2
2020 Hamiltonian Path Based Mixed-Cell-Height Legalization for Neighbor Diffusion Effect Mitigation
abstract
In modern circuit designs, standard cells are designed with different heights based on the power, area, and other characteristics to address various design requirements. For those cells with different heights, in particular, there are inter-cell diffusion steps if the diffusion heights of neighboring cells are different, called the neighbor diffusion effect (NDE) which has become critical in advanced technology nodes. In this paper, we present a Hamiltonian-path-based mixed-cell-height legalization algorithm for NDE mitigation. We first present a row assignment method considering both cell displacements and diffusion steps to assign cells to their desired rows that meet the power-rail alignment constraints. Then, we propose a Hamiltonian-path-based diffusion-step reduction method to effectively reduce the NDE violations while preserving the global placement solution. Particularly, we develop a 2-approximation algorithm to find a minimum weight Hamiltonian path connecting two vertices, and a 1.5-approximation algorithm to find a minimum weight Hamiltonian path with a specified end vertex. Finally, we present an NDE-aware legalization method with design compaction to resolve overlaps and NDE violations. Experimental results show that our algorithm can resolve all NDE violations without any area overhead in reasonable runtime.
Jianli Chen, Ziran Zhu, Qinghai Liu, Wenxing Zhu, Yao-Wen Chang
DAC2
2020 Mixed-cell-height legalization considering complex minimum width constraints and half-row fragmentation effect
Ziran Zhu, Zhipeng Huang 0009, Wenxing Zhu, Jianli Chen, Hanbin Zhou, Senhua Dong
Integr.1
2020 Mixed-Cell-Height Legalization Considering Technology and Region Constraints
abstract
Mixed-cell-height circuits have become popular in advanced technologies for better power, area, routability, and performance tradeoffs. With technology and region constraints imposed by modern circuit designs, the mixed-cell-height legalization problem has become even more challenging. Additionally, an ideal legalization method should minimize both the average and maximum cell movements to preserve the quality of a given placement as much as possible. In this article, we present an effective and efficient mixed-cell-height legalization algorithm to consider technology and region constraints while minimizing the average and maximum cell movements. We first present a fence region handling technique to unify the fence regions and the default region. To obtain a desired cell assignment, we then propose a movement-aware cell reassignment method by iteratively reassigning cells in locally dense areas to their desired rows. After cell reassignment, a technology-aware legalization is presented to remove cell overlaps while satisfying the technology constraints. Finally, we propose a technology-aware refinement to further reduce the average and maximum cell movements without increasing the technology constraints violations. Compared with the champion of the 2017 CAD Contest at ICCAD and the state-of-the-art work, experimental results based on the 2017 CAD Contest at ICCAD benchmarks show that our algorithm achieves the best average and maximum cell movements and significantly fewer technology constraint violations, in a comparable runtime. The experimental results based on the modified 2015 ISPD Contest benchmarks also demonstrate the effectiveness of our algorithm in minimizing the average and maximum cell movements, compared with state-of-the-art mixed-cell-height legalizers.
Ziran Zhu, Jianli Chen, Wenxing Zhu, Yao-Wen Chang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2018 Generalized augmented lagrangian and its applications to VLSI global placement
abstract
Global placement dominates the circuit placement process in its solution quality and efficiency. With increasing design complexity and various design constraints, it is desirable to develop an efficient, high-quality global placement algorithm for modern large-scale circuit designs. In this paper, we first analyze the properties of four nonlinear optimization methods (the quadratic penalty method, the Lagrange multiplier method, and two augmented Lagrangian methods) for global placement, and then develop a generalized augmented Lagrangian method to solve this problem. Our proposed method preserves the advantages of the quadratic penalty method and the augmented Lagrangian method, and provides a smooth progress from the quadratic penalty method to the augmented Lagrangian method. We prove that the proposed generalized augmented Lagrangian method is globally convergent for the original global placement problem, even with different constraints. Compared with the other four popular optimization methods, experimental results show that our method achieves the best quality and is robust for handling different objectives. In particular, our generalized augmented Lagrangian formulation is theoretically sound and can solve generic large-scale constrained nonlinear optimization problems, which are widely used in many fields.
Ziran Zhu, Jianli Chen, Zheng Peng 0002, Wenxing Zhu, Yao-Wen Chang
DAC1
2018 Mixed-cell-height legalization considering technology and region constraints
abstract
Mixed-cell-height circuits have become popular in advanced technologies for better power, area, routability, and performance tradeoffs. With the technology and region constraints imposed by modern circuit designs, the mixed-cell-height legalization problem has become more challenging. In this paper, we present an effective and efficient legalization algorithm for mixed-cell-height circuit designs with technology and region constraints. We first present a fence region handling technique to unify the fence regions and the default ones. To obtain a desired cell assignment, we then propose a movement-aware cell reassignment method by iteratively reassigning cells in locally dense areas to their desired rows. After cell reassignment, a technology-aware legalization is presented to remove cell overlaps while satisfying the technology constraints. Finally, we propose a technology-aware refinement to further reduce the average and maximum cell movements without increasing the technology constraints violations. Compared with the champion of the 2017 ICCAD CAD Contest and the state-of-the-art work, experimental results show that our algorithm achieves the best average and maximum cell movements and significantly fewer technology constraint violations, in a comparable runtime.
Ziran Zhu, Jianli Chen, Wenxing Zhu, Yao-Wen Chang
ICCAD1
2017 Toward Optimal Legalization for Mixed-Cell-Height Circuit Designs
abstract
Modern circuits often contain standard cells of different row heights to meet various design requirements. Higher cells give larger drive strengths at the costs of larger areas and power. Multi-row-height standard cells incur challenging issues to layout designs, especially the mixed-cell-height legalization problem due to the heterogeneous cell structures. Honoring the good cell positions from global placement, we present in this paper a fast and near-optimal algorithm to solve the legalization problem. Fixing the cell ordering from global placement and relaxing the right boundary constraints, we first convert the problem into a linear complementarity problem (LCP). With the converted LCP, we split its matrices to meet the convergence requirement of a modulus-based matrix splitting iteration method (MMSIM), and then apply the MMSIM to solve the LCP. This MMSIM method guarantees the optimality if no cells are placed beyond the right boundary of a chip. Finally, a Tetris-like allocation approach is used to align cells to placement sites on rows and fix the placement of out-of-right-boundary cells, if any. Experimental results show that our proposed algorithm can achieve the best cell displacement and wirelength among all published methods in reasonable runtimes. The MMSIM optimality is theoretically proven and empirically validated. In particular, our formulation provides new generic solutions and research directions for various optimization problems that require solving large-scale quadratic programs efficiently.
Jianli Chen, Ziran Zhu, Wenxing Zhu, Yao-Wen Chang
DAC2
2017 An adaptive hybrid memetic algorithm for thermal-aware non-slicing VLSI floorplanning
Jianli Chen, Ziran Zhu, Wenxing Zhu
Integr.3
2017 Discrete Relaxation Method for Triple Patterning Lithography Layout Decomposition
abstract
In this paper, we consider the triple patterning lithography layout decomposition problem. To address the problem, a discrete relaxation theory is built. For designing a discrete relaxation based decomposition framework, we propose a surface projection method for identifying native conflicts in a layout, and then constructing the conflict graph. Guided by the theory, the conflict graph is reduced to small size subgraphs by vertex removals, which is a discrete relaxation. Furthermore, by ignoring stitch insertions and assigning weights to features, the layout decomposition problem on the small subgraphs is further relaxed to a 0-1 program, which is solved by the Branch-and-Bound method. To obtain a feasible solution of the original problem, legalization methods are introduced to legalize a relaxation solution. At the legalization stage, we prior utilize one-stitch insertion to eliminate conflicts, and use a backtrack coloring algorithm to obtain a better solution. We test our decomposition approach on the ISCAS-85 & 89 benchmarks. Comparisons of experimental results show that our approach finds solutions of some benchmarks better than those by the state-of-the-art decomposers. Especially, according to our discrete relaxation theory, some optimal decompositions are obtained.
Ziran Zhu, Wenxing Zhu
IEEE Trans. Computers2