Yuxuan Zhao 0001

dblp:173/0246-1 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0001-5995-4763ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 4 first-author · 15 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DPO-3D: Differentiable Power Delivery Network Optimization via Flexible Modeling for Routability and IR-Drop Tradeoff in Face-to-Face 3D ICs
Zhen Zhuang, Yuxuan Zhao 0001, Bei Yu 0001, Sung Kyu Lim, Tsung-Yi Ho
ASP-DAC3
2026 HPPlacer: A High-Precision Slack-Aware Global Placement Engine
abstract
Timing-driven global placement plays a decisive role in the final performance of very large-scale integration (VLSI) circuits, but is consistently challenged by the trade-off between design accuracy and efficiency. Most existing methods rely on coarse-grained net-weighting strategies. While these approaches are straightforward to implement, they cannot precisely identify and optimize complex timing paths, such as paths with sharing effects or large slack deviations. To overcome this bottleneck, we propose a high-precision slack-aware global placement engine called HPPlacer, which includes the following three key techniques: 1) a local clock buffer-to-flip-flop connection optimization method, 2) a path-level differentiable timing optimization model, and 3) a dynamic adjustment mechanism-based pin-pair weighting strategy. With the proposed method, efficient chip placement with excellent timing behaviors can be generated automatically within a short period of time. The experimental results on multiple benchmark circuits confirm that HPPlacer leads to significant improvements in both timing performance and wirelength compared to state-of-the-art placement tools.
Qinggong Shen, Haoyang Xu, Zhiwen Yu 0001, Bin Guo 0001, Yuxuan Zhao 0001, Bei Yu 0001, Tsung-Yi Ho, Xing Huang 0001
DATE6
2026 IncreMacro-3D: Incremental Macro Placement for Face-to-Face Stacked Memory-on-Logic 3D ICs
abstract
Face-to-face stacked 3D ICs, such as memory-on-logic (MoL) architectures, have emerged as a promising solution to overcome the limitations of traditional 2D integration by offering enhanced performance, power efficiency, and density. Given the increasing design complexity of modern system-on-chips (SoCs), achieving high-quality macro placement is critical, as it plays a decisive role in determining the final performance, power, and area (PPA) metrics. However, existing RTL-to-GDS 3D physical design flows for MoL 3D ICs rely heavily on manual macro placement, which becomes increasingly challenging and time-consuming for modern SoCs with a vast number of macros. In this paper, we introduce an innovative macro placement algorithm, IncreMacro-3D, which employs graph neural network-based macro repartitioning and 3D macro position refinement, thereby facilitating subsequent steps in 3D physical design flow. The experimental results on several benchmark circuits demonstrate that the proposed approach can reduce the routed wirelength, worst negative slack (WNS), total negative slack (TNS), and total power consumption by 6.1%, 44.2%, 62.8%, and 0.6% compared to state-of-the-art analytical placer for MoL 3D ICs.
Lancheng Zou, Sing Sen Ye, Yuan Pu 0001, Jiaxi Jiang, Siting Liu 0002, Yuxuan Zhao 0001, Bei Yu 0001
DATE7
2026 RegPlace: Regularity-Aware Placement for Full-System DNN Accelerator Designs
abstract
The rise of deep neural network accelerators demands physical design tools that recognize spatial regularity patterns. Traditional placers, unaware of the regularity of spatial arrays, produce suboptimal solutions. This work proposes RegPlace, a regularity-aware placement algorithm for full-system DNN accelerators that automatically identifies processing elements using graph convolutional networks and employs a variance-based soft regularity loss to guide optimization. Compared to state-of-the-art methods, our approach achieves up to 6% wirelength reduction while maintaining comparable runtime, with post-placement metrics further confirming its effectiveness.
Jiaxi Jiang, Yuan Pu 0001, Yuxuan Zhao 0001, Peiyu Liao, Zuodong Zhang, Yibo Lin, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2026 DeepVerifier: Learning to Update Test Sequences for Coverage-Guided Verification
abstract
Verification is critical in ensuring the reliable operation of modern, complex computing systems. However, as processor designs become increasingly sophisticated, conventional static verification techniques struggle to generate high-quality test sequences that achieve comprehensive coverage. Dynamic simulation-based approaches, which leverage coverage-driven objectives, can increase confidence in correct processor functionality but often suffer from low verification efficiency due to the generation of redundant test sequences and significant computational overhead. To address these challenges, this paper presents DeepVerifier, a novel coverage-guided test generation framework that leverages data-driven learning of existing test sequences and their associated coverage feedback. DeepVerifier uses a language model to learn the semantic representations of test sequences, ensure adherence to syntax constraints, and estimate the relationship between test sequences and coverage scores. By updating test sequences with higher coverage, DeepVerifier can significantly improve the efficiency and effectiveness of the verification process. Experimental results of verifying an out-of-order RISC-V microprocessor demonstrate that the framework accurately estimates the coverage scores of test sequences and updates high-quality sequences that contribute to higher coverage. This coverage-guided test generation technique holds promise for enhancing the reliability of modern processor designs.
Yuntao Lu, Yuxuan Zhao 0001, Ziyue Zheng, Yangdi Lyu, Bei Yu 0001
ACM Trans. Design Autom. Electr. Syst.3
2026 PPD: A Portable and Highly Parallel Dispatching System for Deep Learning
Wendong Xu, Yuhao Ji, Yueting Li 0001, Yuxuan Zhao 0001, Zhengwu Liu, Bei Yu 0001, Ngai Wong 0001
ACM Trans. Design Autom. Electr. Syst.5
2025 A Systematic Approach for Multi-objective Double-side Clock Tree Synthesis
abstract
As the scaling of semiconductor devices nears its limits, utilizing the back-side space of silicon has emerged as a new trend for future integrated circuits. With intense interest, several works have hacked existing backend tools to explore the potential of synthesizing double-side clock trees via nano Through-Silicon-Vias (nTSVs). However, these works lack a systematic perspective on design resource allocation and multi-objective optimization. We propose a systematic approach to design clock trees with double-side metal layers, including hierarchical clock routing, concurrent buffers and nTSVs insertion, and skew refinement. Compared with the state-of-the-art (SOTA) methods, the widely-used open-source tool, our algorithm outperforms them in latency, skew, wirelength, and the number of buffers and nTSVs.
Xun Jiang 0002, Yuxuan Zhao 0001, Zizheng Guo 0001, Heng Wu 0007, Bei Yu 0001, Sung Kyu Lim, Runsheng Wang, Ru Huang 0001, Yibo Lin
DAC3
2025 3D-Flow: Flow-based Standard Cell Legalization for 3D ICs
abstract
The standard-cell placement legalization is a critical step in the physical design. The emerging 3D ICs have brought challenges to traditional legalizers on efficiency and effectiveness. In this work, we present a fast flow-based legalization algorithm, 3D-Flow, that minimizes cell displacement in a 3D solution space. Our legalizer resolves overflowed bins by finding the shortest augmenting path on a 3D grid graph, utilizing an effective branch-and-bound algorithm. Moreover, a post-optimization with a cycle-canceling algorithm is proposed to minimize the maximum displacement. Our approach leverages the global perspective inherent in network flow methods, considering multiple dies in 3D ICs to minimize cell displacement. Experimental results on ICCAD 2022 and 2023 contest benchmarks demonstrate our proposed algorithm achieves up to 13% and 43% less average and maximum cell displacement compared to state-of-the-art legalizers in a similar runtime.
Yuxuan Zhao 0001, Peiyu Liao, Bei Yu 0001
DAC1
2025 Ultrafast Density Gradient Accumulation in 3D Analytical Placement with Divergence Theorem
abstract
Density gradient accumulation plays a pivotal role in 3D analytical placement. Analytical placers rely on this fundamental operation during the backward step of each iteration to compute the gradient of the density penalty for every node. This primitive operation thus constitutes a significant runtime bottleneck, especially for mixed-size designs with large macros. Furthermore, this bottleneck becomes increasingly critical as the grid size in 3D placement is considerably larger than that in conventional 2D placement. In this paper, we propose an algorithm inspired by the divergence theorem to reduce the time complexity of density gradient accumulation. We also present our implementations of this algorithm for both CPU and GPU versions. Experimental results demonstrate that our method achieves more than 3× end-to-end runtime speedup on CPU and GPU compared to the SOTA analytical 3D placer.
Peiyu Liao, Yuxuan Zhao 0001, Siting Liu 0002, Bei Yu 0001
ICCAD2
2025 H3D: Heterogeneous Resources Aware Global Router for Face-to-Face Bonded 3D ICs
abstract
The emerging 3D ICs have brought challenges to traditional routers in deciding the intra-die and inter-die interconnects. Existing pseudo-3D flows rely on 2D IC routing engines, combined with a 3D via legalization step to complete routing. The separation of intra-die and inter-die routing significantly degrades solution quality. To address this issue, we propose H3D, the first native 3D global router designed for face-to-face bonded 3D ICs. H3D constructs a heterogeneous routing grid to represent routing and hybrid bonding terminal (HBT) resources. We develop dedicated dynamic programming-based algorithms to optimize the HBT number and locations for 3D Steiner trees on the heterogeneous grid. Specifically, H3D minimizes the number of HBTs by traversing the Steiner tree in reverse depth-first order and relocates HBTs to legal locations with minimal wirelength in reverse breadth-first order, leveraging the convexity of L1 distance for efficient optimization. Experimental results on various real-world designs demonstrate that H3D achieves 12% shorter wirelength, 25% fewer HBTs, and 1.9× speedup compared to state-of-the-art approaches.
Yuxuan Zhao 0001, Siting Liu 0002, Peiyu Liao, Bei Yu 0001
ICCAD1
2025 Invited: Physical Design for Advanced 3D ICs: Challenges and Solutions
abstract
As technology scaling predicted by Moore's law slows down, 3D integrated circuits (3D ICs) have emerged as a promising alternative to enhance performance while maintaining cost-effectiveness. With the advancement of fabrication and bonding technologies, wafer-level 3D integration enables fine-grain 3D interconnects that maximize the benefits in power, performance, and area (PPA). However, a multitude of challenges have obstructed traditional electronic design automation (EDA) methodologies for 3D IC implementations. This paper surveys the major challenges in the physical design of advanced 3D ICs. We provide a comprehensive review of existing solutions, analyzing their advantages and disadvantages in depth. Finally, we discuss open problems and research opportunities in the development of native 3D EDA tools.
Yuxuan Zhao 0001, Lancheng Zou, Bei Yu 0001
ISPD1
2025 Analytical Heterogeneous Die-to-Die 3-D Placement With Macros
abstract
This article presents an innovative approach to 3-D mixed-size placement in heterogeneous face-to-face (F2F) bonded 3-D ICs. We propose an analytical framework that utilizes a dedicated density model and a bistratal wirelength model, effectively handling macros and standard cells in a 3-D solution space. A novel 3-D preconditioner is developed to resolve the topological and physical gap between macros and standard cells. Additionally, we propose a mixed-integer linear programming (MILP) formulation for macro rotation to optimize wirelength. Our framework is implemented with full-scale GPU acceleration, leveraging an adaptive 3-D density accumulation algorithm and an incremental wirelength gradient algorithm. Experimental results on ICCAD 2023 contest benchmarks demonstrate that our framework can achieve 5.9% quality score improvement compared to the first-place winner with 4.0$\times $runtime speedup. Additional experiments on modern RISC-V designs further validate the generalizability and superiority of our framework.
Yuxuan Zhao 0001, Peiyu Liao, Siting Liu 0002, Jiaxi Jiang, Yibo Lin, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2024 Analytical Die-to-Die 3-D Placement With Bistratal Wirelength Model and GPU Acceleration
abstract
In this paper, we present a new analytical 3D placement framework with a bistratal wirelength model for F2Fbonded 3D ICs with heterogeneous technology nodes based on the electrostatic-based density model. The proposed framework, enabled GPU-acceleration, is capable of efficiently determining node partitioning and locations simultaneously, leveraging the dedicated 3D wirelength model and density model. The experimental results on ICCAD 2022 contest benchmarks demonstrate that our proposed 3D placement framework can achieve up to 6.1% wirelength improvement and 4.1% on average compared to the first-place winner with much fewer vertical interconnections and up to 9.8× runtime speedup. Notably, the proposed framework also outperforms the state-of-the-art 3D analytical placer by up to 3.3% wirelength improvement and 2.1% on average with up to 8.8× acceleration on large cases using GPUs.
Peiyu Liao, Yuxuan Zhao 0001, Dawei Guo, Yibo Lin, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 AutoGraph: Optimizing DNN Computation Graph for Parallel GPU Kernel Execution
abstract
Deep learning frameworks optimize the computation graphs and intra-operator computations to boost the inference performance on GPUs, while inter-operator parallelism is usually ignored. In this paper, a unified framework, AutoGraph, is proposed to obtain highly optimized computation graphs in favor of parallel executions of GPU kernels. A novel dynamic programming algorithm, combined with backtracking search, is adopted to explore the optimal graph optimization solution, with the fast performance estimation from the mixed critical path cost. Accurate runtime information based on GPU Multi-Stream launched with CUDA Graph is utilized to determine the convergence of the optimization. Experimental results demonstrate that our method achieves up to 3.47x speedup over existing graph optimization methods. Moreover, AutoGraph outperforms state-of-the-art parallel kernel launch frameworks by up to 1.26x.
Yuxuan Zhao 0001, Qi Sun 0002, Zhuolun He, Bei Yu 0001
AAAI1
2022 GTuner: tuning DNN computations on GPU via graph attention network
abstract
It is an open problem to compile DNN models on GPU and improve the performance. A novel framework, GTuner, is proposed to jointly learn from the structures of computational graphs and the statistical features of codes to find the optimal code implementations. A Graph ATtention network (GAT) is designed as the performance estimator in GTuner. In GAT, graph neural layers are used to propagate the information in the graph and a multi-head self-attention module is designed to learn the complicated relationships between the features. Under the guidance of GAT, the GPU codes are generated through auto-tuning. Experimental results demonstrate that our method outperforms the previous arts remarkably.
Qi Sun 0002, Xinyun Zhang 0001, Hao Geng, Yuxuan Zhao 0001, Haisheng Zheng, Bei Yu 0001
DAC4
2022 Mixed-Cell-Height Legalization on CPU-GPU Heterogeneous Systems
abstract
Legalization conducts refinements on post-globalplacement cell location to compromise design constraints and parameters. These include placement fence regions, power/ground rail alignments, timing, wire length and etc. In advanced technology nodes, designs can easily contain millions of mutiple-row standard cells, which challenges the scalability of modern legalization algorithms. In this paper, for the first time, we investigate dedicated legalization algorithms on heterogeneous platforms, which promises intelligent usage of CPU and GPU resources and hence provides new algorithm design methodologies for large scale physical design problems. Experimental results on IC/CAD 2017 and ISPD 2015 contest benchmarks demonstrate the effectiveness and the efficiency of the proposed algorithm, compared to the state-of-the-art legalization solution for mixedcell-height designs.
Kit Fung, Yuxuan Zhao 0001, Yibo Lin, Bei Yu 0001
DATE3