Hao-Hsiang Hsiao

dblp:343/0877 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0003-4865-0654ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 6 first-author · 11 since 2021
YearPublicationVenuePosition
2026 C3PO: Commercial-Quality Global Placement via Coherent, Concurrent Timing, Routability, and Wirelength Optimization
abstract
Despite achieving orders-of-magnitude runtime speedup, GPU-accelerated placers (GPU-Placers) still have extremely limited industrial adoption, largely due to the wide gaps in Power, Performance, and Area (PPA) metrics compared to those well-established CPU-centric commercial Physical Design (PD) tools. To overcome this issue, we introduce C3PO, the first commercial-quality, differentiable, multi-objective global placer that performs concurrent timing, routability, and wirelength optimization in a coherent manner with custom CUDA kernels. Particularly, we propose a convex-based framework that dynamically computes objective weights at each placement iteration by solving a quadratic problem, eliminating the need of manual parameter tuning. In the experiments, we rigorously validate C3PO with an industry-leading commercial PD tool and demonstrate that on 8 designs from TILOS [1] and IWLS [2] in ASAP 7nm [3], C3PO consistently outperforms the commercial tool by up to 16.7% in routed wirelength and 19.6% in switching power with complete full-flow validation.
Yi-Chen Lu, Hao-Hsiang Hsiao, Rongjian Liang, Haoxing Ren
ASP-DAC2
2026 Differentiable Tier Assignment for Timing and Congestion-Aware Routing in 3D ICs
abstract
State-of-the-art (SOTA) 3D physical design (PD) flows extend commercial 2D place-and-route (P&R) tools to enable signoff-quality 3D IC implementation through double metal stacking and inter-die metal layer sharing. While metal layer sharing introduces additional routing resources, the substantially higher manufacturing cost of face-to-face (F2F) inter-die vias compared to intra-die vias necessitates 3D-aware routing strategies to manage routability-cost trade-offs. To address this, we propose differentiable routing guidance for 3D ICs (DRG-3D), a GPUaccelerated differentiable optimization framework that provides routing guidance for 3D ICs. DRG-3D formulates a fully differentiable objective that simultaneously optimizes key 3D design metrics: routing congestion, wirelength, via cost, and F 2 F -via cost, which enables efficient and scalable gradient-based optimization over large-scale netlists. Experimental results show that DRG-3D outperforms the SOTA Pin-3D flow, achieving up to 8.37% reduction in routing overflow, 23.99% reduction in total negative slack (TNS), and 18.05% reduction in post-route timing violations.
Yuan-Hsiang Lu, Hao-Hsiang Hsiao, Yi-Chen Lu, Haoxing Ren, Sung Kyu Lim
ASP-DAC2
2026 A Hybrid Reinforcement Learning Framework for Efficient Physical Design Parameter Tuning
abstract
Traditional Design Space Exploration (DSE) methods in Physical Design (PD), such as Bayesian Optimization (BO) and Ant Colony Optimization (ACO), as well as state-of-the-art commercial tools like Synopsys DSO.ai, typically treat the design flow as a black box, lacking insight into the underlying designs. This hinders their ability to generalize across unseen designs. In this article, we introduce FastTuner, an innovative Reinforcement Learning (RL) agent that leverages Graph Neural Networks (GNNs) and Transformers to understand the underlying designs and enable rapid DSE on unseen designs across various PD stages. Our approach incorporates an attention-based framework for autoregressive and conditional parameter tuning and introduces a power, performance and area (PPA) estimator to predict end-of-flow PPA metrics, significantly accelerating RL reward computation. Extensive evaluations on seven industrial designs using the TSMC 28nm technology node demonstrate that FastTuner significantly outperforms existing state-of-the-art DSE techniques in both optimization quality and runtime, achieving improvements of up to 79.38% in Total Negative Slack (TNS), 12.22% in total power, and more than 50x reduction in runtime.
Hao-Hsiang Hsiao, Yi-Chen Lu, Pruek Vanna-Iampikul, Sung Kyu Lim
ACM Trans. Design Autom. Electr. Syst.1
2025 InsightAlign: A Transferable Physical Design Recipe Recommender Based on Design Insights
abstract
Physical design tools have complex workflows with many different ways of optimizing power, performance, and area (PPA) out of a large number of options and hyperparameters in different engines and functionalities. Black-box optimization techniques are widely adopted to automate quality-of-result (QoR) exploration. Such exploration often proves impractical in real-world customer environments due to high computational demands, lengthy exploration cycles, and the need for large parallel jobs. To reduce the exploration space for viable compute resource requirements, we propose a novel design methodology to enable transferable learning by incorporating design insights crafted on top of physical design experts’ experience and streamlining QoR exploration as a sequence generation task for best recipe selection. We apply language model-inspired alignment techniques to learn the ranking of different recipe sets, enabling our model to generalize beyond known-good manually tuned expert design recipes. Extensive evaluations demonstrate our method’s superior QoRs and runtime performance on unseen industrial designs and rigorous benchmarks.
Hao-Hsiang Hsiao, Sudipto Kundu, Wei-Ting Chan, Deyuan Guo, Sung Kyu Lim
DAC1
2025 DCO-3D: Differentiable Congestion Optimization in 3D ICs
abstract
State-of-the-art 3D IC flows fail to consider 3D congestion during earlier stages, leading to excessive use of end-of-flow ECO resources for routability correction that severely degrades full-chip Power, Performance, and Area metrics. We present DCO-3D, a Machine Learning-based routability-aware 3D PD flow that performs early post-route congestion prediction using Siamese Networks and resolves the predicted hotspots using a fully differentiable 3D cell spreading with Graph Neural Network. On 6 industrial designs in a commercial 3nm node, DCO-3D improves Pin-3D, the known best Pin-3D flow, by up to 47.2% in overflow, 86.2% in TNS and 5.1% in power at signoff.
Hao-Hsiang Hsiao, Yi-Chen Lu, Pruek Vanna-Iampikul, Anthony Agnesina, Rongjian Liang, Yuan-Hsiang Lu, Haoxing Ren, Sung Kyu Lim
DAC1
2025 BUFFALO: PPA-Configurable, LLM-based Buffer Tree Generation via Group Relative Policy Optimization
abstract
Buffer insertion is a critical netlist optimization technique in Physical Design (PD) that balances trade-offs between Power, Performance, and Area (PPA) metrics. Traditional buffering methods rely heavily on local heuristics, which do not scale and often result in globally sub-optimal solutions. Prior Machine Learning (ML) techniques such as BufFormer attempted to alleviate this limitation but remain prohibitively time-consuming (and sub-optimal) due to their incremental nature. In this paper, we introduce BUFFALO, a generative buffer insertion framework that, for the first time in PD, formulates buffer tree generation as a sequence-to-sequence task solved by Large Language Models (LLMs). Particularly, given a design, BUFFALO performs single-shot generation of buffer trees for all fanout-violating nets and INSTA-selected timing critical nets. Furthermore, Group Relative Policy Optimization (GRPO), a Reinforcement Learning (RL) technique, is employed to refine predicted solutions in a PPA-configurable manner. Experimental results on 9 full-chip designs in a 7nm node demonstrate that BUFFALO outperforms an industry-leading commercial PD tool by 71% in Total Negative Slack (TNS), 67.69% in Worst Negative Slack (WNS), and 83x in runtime without incurring additional power consumption.
Hao-Hsiang Hsiao, Yi-Chen Lu, Sung Kyu Lim, Haoxing Ren
ICCAD1
2025 Invited Paper: LLM-Enhanced GPU-Optimized Physical Design at Scale
abstract
Modern Physical Design (PD) flows face a dual challenge: proprietary, heterogeneous design data and the rapid evolution of process nodes, both of which block models from transferring to new chips. To overcome these hurdles, we demonstrate a unified, data-driven framework that distills critical netlist optimization moves, including gate sizing, buffer insertion, and cell relocation, into "optimization primitives" learned by Large Language Models (LLMs). Particularly, we develop a high-quality, synthetic optimization data generation pipeline with commercial tools at scale, while using a GPU-accelerated differentiable Static Timing Analysis (STA) engine to create fast feedback loop, enabling end-to-end gradient propagation to guide model learning. By training on both real and synthetic data across multiple technology generations, our approach captures fundamental PD optimization patterns that transfer seamlessly to unseen designs, overcoming the constraints of fragmented design representations and proprietary data in industrial PD flows.
Yi-Chen Lu, Hao-Hsiang Hsiao, Haoxing Ren
ICCAD2
2024 ML-based Physical Design Parameter Optimization for 3D ICs: From Parameter Selection to Optimization
abstract
While various studies have shown effective parameter optimizations for specific designs, there is limited exploration of parameter optimization within the domain of 3D Integrated Circuits. We present the first comprehensive study, both qualitatively and quantitatively, comparing five state-of-the-art (SOTA) techniques for parameter optimization applied to 3D ICs. Additionally, we introduce an end-to-end machine learning-based framework, encompassing important parameter selection through optimization, all without human intervention. Extensive studies across six industrial designs under the TSMC 28nm technology node reveal that our proposed framework outperforms SOTA techniques in three different optimization objectives in both optimization quality and runtime.
Hao-Hsiang Hsiao, Pruek Vanna-Iampikul, Yi-Chen Lu, Sung Kyu Lim
DAC1
2024 FastTuner: Transferable Physical Design Parameter Optimization using Fast Reinforcement Learning
abstract
Current state-of-the-art Design Space Exploration (DSE) methods in Physical Design (PD), including Bayesian optimization (BO) and Ant Colony Optimization (ACO), mainly rely on black-boxed rather than parametric (e.g., neural networks) approaches to improve end-of-flow Power, Performance, and Area (PPA) metrics, which often fail to generalize across unseen designs as netlist features are not properly leveraged. To overcome this issue, in this paper, we develop a Reinforcement Learning (RL) agent that leverages Graph Neural Networks (GNNs) and Transformers to perform "fast" DSE on unseen designs by sequentially encoding netlist features across different PD stages. Particularly, an attention-based encoder-decoder framework is devised for "conditional" parameter tuning, and a PPA estimator is introduced to predict end-of-flow PPA metrics for RL reward estimation. Extensive studies across 7 industrial designs under the TSMC 28nm technology node demonstrate that the proposed framework FastTuner, significantly outperforms existing state-of-the-art DSE techniques in both optimization quality and runtime. where we observe improvements up to 79.38% in Total Negative Slack (TNS), 12.22% in total power, and 50x in runtime.
Hao-Hsiang Hsiao, Yi-Chen Lu, Pruek Vanna-Iampikul, Sung Kyu Lim
ISPD1
2024 GAN-Place: Advancing Open Source Placers to Commercial-quality Using Generative Adversarial Networks and Transfer Learning
abstract
Recently, GPU-accelerated placers such as DREAMPlace and Xplace have demonstrated their superiority over traditional CPU-reliant placers by achieving orders of magnitude speed up in placement runtime. However, due to their limited focus in placement objectives (e.g., wirelength and density), the placement quality achieved by DREAMPlace or Xplace is not comparable to that of commercial tools. In this article, to bridge the gap between open source and commercial placers, we present a novel placement optimization framework named GAN-Place that employs generative adversarial learning to transfer the placement quality of the industry-leading commercial placer, Synopsys ICC2, to existing open source GPU-accelerated placers (DREAMPlace and Xplace). Without the knowledge of the underlying proprietary algorithms or constraints used by the commercial tools, our framework facilitates transfer learning to directly enhance the open source placers by optimizing the proposed differentiable loss that denotes the “similarity” between DREAMPlace- or Xplace-generated placements and those in commercial databases. Experimental results on seven industrial designs not only show that our GAN-Place immediately improves the Power, Performance, and Area metrics at the placement stage but also demonstrates that these improvements last firmly to the post-route stage, where we observe improvements by up to 8.3% in wirelength, 7.4% in power, and 37.6% in Total Negative Slack on a commercial CPU benchmark.
Yi-Chen Lu, Haoxing Ren, Hao-Hsiang Hsiao, Sung Kyu Lim
ACM Trans. Design Autom. Electr. Syst.3
2023 DREAM-GAN: Advancing DREAMPlace towards Commercial-Quality using Generative Adversarial Learning
abstract
DREAMPlace is a renowned open-source placer that provides GPU-acceleratable infrastructure for placements of Very-Large-Scale-Integration (VLSI) circuits. However, due to its limited focus on wirelength and density, existing placement solutions of DREAMPlace are not applicable to industrial design flows. To improve DREAMPlace towards commercial-quality without knowing the black-boxed algorithms of the tools, in this paper, we present DREAM-GAN, a placement optimization framework that advances DREAMPlace using generative adversarial learning. At each placement iteration, aside from optimizing the wirelength and density objectives of the vanilla DREAMPlace, DREAM-GAN computes and optimizes a differentiable loss that denotes the similarity score between the underlying placement and the tool-generated placements in commercial databases. Experimental results on 5 commercial and OpenCore designs using an industrial design flow implemented by Synopsys ICC2 not only demonstrate that DREAM-GAN significantly improves the vanilla DREAMPlace at the placement stage across each benchmark, but also show that the improvements last firmly to the post-route stage, where we observe improvements by up to 8.3% in wirelength and 7.4% in total power.
Yi-Chen Lu, Haoxing Ren, Hao-Hsiang Hsiao, Sung Kyu Lim
ISPD3