VLDB 2026 Research / reviewers in the wild / expert
Pruek Vanna-Iampikul
dblp:279/8587
· DBLP profile ↗
20ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0002-2897-7142ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 6 first-author · 19 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Boosting Scalability and Performance: Macro Placement for Flexible 3D-Stacked ML AcceleratorsabstractMachine learning (ML) accelerators increasingly demand architectural flexibility to support diverse workloads. While reconfigurable designs like Self-Adaptive Reconfigurable Arrays (SARA) improve adaptability, they suffer from physical design bottlenecks-particularly in memory-intensive settings where dense memory integration inflates area, wirelength, and energy consumption. To address these challenges, we propose a scalable multi-tier memory-on-logic (MoL) integration framework that co-optimizes macro placement and standard-cell layout for flexible ML accelerators. Our methodology combines architectural insights with subgradient-based convex optimization for global placement, followed by simulated annealing for overlap resolution and multi-tier partitioning. The proposed flow significantly improves area efficiency, routing congestion, and timing closure. Experimental results on ML workloads demonstrate up to 3.4 × runtime speedup, 4.4 × energy-delay product (EDP) improvement, and $2.3 \times$ chip area reduction compared to conventional two-tier MoL baselines, enabling practical deployment of high-performance, adaptive 3D ML accelerators. Canlin Zhang, Pruek Vanna-Iampikul, Tushar Krishna, Sung Kyu Lim |
ASP-DAC | 3 |
| 2026 | A Hybrid Reinforcement Learning Framework for Efficient Physical Design Parameter TuningabstractTraditional Design Space Exploration (DSE) methods in Physical Design (PD), such as Bayesian Optimization (BO) and Ant Colony Optimization (ACO), as well as state-of-the-art commercial tools like Synopsys DSO.ai, typically treat the design flow as a black box, lacking insight into the underlying designs. This hinders their ability to generalize across unseen designs. In this article, we introduce FastTuner, an innovative Reinforcement Learning (RL) agent that leverages Graph Neural Networks (GNNs) and Transformers to understand the underlying designs and enable rapid DSE on unseen designs across various PD stages. Our approach incorporates an attention-based framework for autoregressive and conditional parameter tuning and introduces a power, performance and area (PPA) estimator to predict end-of-flow PPA metrics, significantly accelerating RL reward computation. Extensive evaluations on seven industrial designs using the TSMC 28nm technology node demonstrate that FastTuner significantly outperforms existing state-of-the-art DSE techniques in both optimization quality and runtime, achieving improvements of up to 79.38% in Total Negative Slack (TNS), 12.22% in total power, and more than 50x reduction in runtime. Hao-Hsiang Hsiao, Yi-Chen Lu, Pruek Vanna-Iampikul, Sung Kyu Lim |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2025 | DCO-3D: Differentiable Congestion Optimization in 3D ICsabstractState-of-the-art 3D IC flows fail to consider 3D congestion during earlier stages, leading to excessive use of end-of-flow ECO resources for routability correction that severely degrades full-chip Power, Performance, and Area metrics. We present DCO-3D, a Machine Learning-based routability-aware 3D PD flow that performs early post-route congestion prediction using Siamese Networks and resolves the predicted hotspots using a fully differentiable 3D cell spreading with Graph Neural Network. On 6 industrial designs in a commercial 3nm node, DCO-3D improves Pin-3D, the known best Pin-3D flow, by up to 47.2% in overflow, 86.2% in TNS and 5.1% in power at signoff. Hao-Hsiang Hsiao, Yi-Chen Lu, Pruek Vanna-Iampikul, Anthony Agnesina, Rongjian Liang, Yuan-Hsiang Lu, Haoxing Ren, Sung Kyu Lim |
DAC | 3 |
| 2025 | GNN-MLS: Signal Routing in Mixed-Node 3D ICs through GNN-Assisted Metal Layer SharingabstractNative 3D Integrated Circuit (3D IC) design offers enhanced performance and density but faces challenges in signal routing due to limited true 3D EDA tool support. Pseudo-3D flows bridge this gap but lack cross-tier optimization, critical for both mixed-node and homogeneous designs. Metal Layer Sharing (MLS) addresses this by enabling cross-tier routing co-optimization but risks timing degradation if not applied strategically. Additionally, MLS creates open connections in hybrid-bonded 3D ICs, making chips untestable. We propose GNN-MLS, a Graph Neural Network-based framework for precise MLS net selection, combined with a tailored DFT solution for robust testability. Experiments show GNN-MLS reduces timing violations by 79% and improves WNS and TNS by 81% and 94%, and moves designs closer to true 3D ICs. Pruek Vanna-Iampikul, Zhen Zhuang, Tsung-Yi Ho, Sung Kyu Lim |
DAC | 2 |
| 2025 | Closing the Gap: Advantages of Block-Level over Gate-Level in 3D IC Design for Advanced NodesabstractGate-level 3D ICs have demonstrated substantial power and performance improvements over 2D ICs. However, at advanced technology nodes, achieving sufficient hybrid bond density within a reduced chip footprint requires a 200nm bond pitch—beyond the limits of current manufacturing capabilities. To address this issue, we focus on block-level 3D IC design which has fewer top-level connections. Previous methodologies for block-level 3D IC design suffer from several limitations: (1) reliance on slow simulated annealing-based floorplanning, (2) suboptimal 3D design flows lacking post-route optimization, (3) a lack of attention to the critical initial step of soft block design, and (4) excessive hybrid bond usage that necessitates a 500nm pitch. To overcome these challenges, we propose a comprehensive methodology that includes: (1) a fast, gradient-based analytical solver, (2) Fence-3D flow, delivering up to 15% improvement in power-delay product, (3) ML-based congestion-aware soft block sizing, delivering up to 6.6% improvement, and (4) a partial-MLS method that selectively applies Metal Layer Sharing, reducing hybrid bond count by up to 71%. Collectively, these techniques enable the use of a 1um hybrid bond pitch in 3nm block-level 3D ICs. Min Gyu Park, Pruek Vanna-Iampikul, Sung Kyu Lim |
ICCAD | 2 |
| 2025 | 3D Acceleration for Mixture-of-Experts and Multi-Head Attention Spiking Transformers with Dynamic Head PruningabstractSpiking Neural Networks (SNNs) provide a brain-inspired and event-driven mechanism that is believed to be critical to unlock energy-efficient deep learning. On the other hand, mixture-of-experts (MoE) models mirror the parallel distributed processing of the nervous system, and expand model capacity without scaling up the number of computational operations. However, there is currently a lack of hardware support for highly parallel distributed processing in spiking based MoE models. This paper introduces the first 3D hardware architecture and design methodology for Mixture-of-Experts and Multi-Head Attention spiking transformers. By leveraging 3D integration with memory-on-logic and logic-on-logic stacking and exploring energy-efficient dynamic head pruning, we explore such brain-inspired accelerators with spatially stackable circuitry, demonstrating significant improvements of energy efficiency and latency over conventional 2D CMOS integration. Boxun Xu, Junyoung Hwang, Pruek Vanna-Iampikul, Yuxuan Yin, Sung Kyu Lim, Peng Li 0001 |
ICCAD | 3 |
| 2025 | Placement-Aware 3D Net-to-Pad Assignment for Array-Style Hybrid Bonding 3D ICsabstractHybrid bonding is emerging as a key technology for 3D integration, offering finer bonding pitches that address the high interconnect density requirements of modern VLSI applications. In advanced node technologies, where metal pitches are significantly smaller than bonding pitches, 3D net assignment becomes critical for achieving optimal design performance. Existing approaches primarily focus on either ensuring the legality of the assignment or optimizing the flexibility of 3D net locations for timing purposes in isolation. This limitation restricts the performance improvements of 3D designs over traditional 2D counterparts. To overcome these challenges, we introduce AnchorGrid, a novel 3D net assignment framework designed to concurrently assign 3D nets to legal locations while supporting their movement to enhance timing optimization. By modeling 3D nets as pairs of specialized ''anchor'' cells, accompanied by relative placement constraints, precise movement and alignment are achieved during the pre-route optimization phase, before final placement onto grid-based locations. Experimental results on advanced node commercial designs demonstrate that AnchorGrid achieves up to a 24.35% improvement in power, performance, and area (PPA) metrics, while reducing design rule check (DRC) violations by 90%, outperforming state-of-the-art methods. Pruek Vanna-Iampikul, Jun-Sik Yoon, Chaeryung Park, Gary Yeap, Sung Kyu Lim |
ISPD | 1 |
| 2025 | Glass Interposer Integration of Logic and Memory Chiplets: PPA and Power/Signal Integrity BenefitsabstractGlass interposers have become a compelling option for 2.5-D heterogeneous integration compared to silicon. It allows 3-D stacking configuration between the embedded dies and the conventional flip-chip dies mounted directly on top at low cost. Furthermore, the interconnect pitch and through-glass-via (TGV) diameter in glass are becoming comparable to their counterparts in silicon. In this study, we investigate the power, performance, area (PPA), signal integrity (SI) and power integrity (PI) advantages of 3-D stacking afforded by glass interposers over silicon interposers. Our research employs a chiplet/package co-design approach, progressing from an register-transfer-level description of RISC-V chiplets to final graphic data system (GDS) layouts, utilizing TSMC 28 nm for chiplets and Georgia Tech’s 3-D glass packaging for the interposer. Compared to silicon, glass interposers offer a$2.6\times $reduction in area, a$21\times $reduction in wire length, a 17.72% reduction in full-chip power consumption, a 64.7% increase in SI and a$10\times $improvement in PI, with a 35% increase in thermal. Furthermore, we provide a detailed comparative analysis with 3-D Silicon technologies. It not only highlights the competitive advantages of glass interposers, but also provides critical insights into each design’s potential limitations and optimization opportunities. Pruek Vanna-Iampikul, Seungmin Woo, Serhat Erdogan, Lingjun Zhu, Mohanalingam Kathaperumal, Ravi Agarwal, Ram Gupta, Kevin Rinebold, Madhavan Swaminathan, Sung Kyu Lim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | GNN-assisted Back-side Clock Routing Methodology for Advance TechnologiesabstractThe back-side metal layers exhibit lower parasitics compared to the front-side layers in advanced technologies, making them suitable for clock-net distribution. In this study, we explore the advantages of using back-side metal layers for clock routing, which is shared with a power delivery network. Our Graph Neural Network (GNN) based framework, effectively distributes the clock-tree between the front and back sides. We address the back-side clock nets' creation by incorporating back-side buffers. Our results demonstrate better clock and full-chip metrics represented by an increase of up to 13% in the effective frequency with equivalent power consumption, using 3 nm technology. Nesara Eranna Bethur, Pruek Vanna-Iampikul, Odysseas Zografos, Lingjun Zhu, Giuliano Sisto, Dragomir Milojevic, Alberto García Ortiz, Geert Hellings, Julien Ryckaert, Francky Catthoor, Sung Kyu Lim |
DAC | 2 |
| 2024 | ML-based Physical Design Parameter Optimization for 3D ICs: From Parameter Selection to OptimizationabstractWhile various studies have shown effective parameter optimizations for specific designs, there is limited exploration of parameter optimization within the domain of 3D Integrated Circuits. We present the first comprehensive study, both qualitatively and quantitatively, comparing five state-of-the-art (SOTA) techniques for parameter optimization applied to 3D ICs. Additionally, we introduce an end-to-end machine learning-based framework, encompassing important parameter selection through optimization, all without human intervention. Extensive studies across six industrial designs under the TSMC 28nm technology node reveal that our proposed framework outperforms SOTA techniques in three different optimization objectives in both optimization quality and runtime. Hao-Hsiang Hsiao, Pruek Vanna-Iampikul, Yi-Chen Lu, Sung Kyu Lim |
DAC | 2 |
| 2024 | AI-Driven Evaluation and Optimization of Bump Pitch Effects on Chiplet and Interposer Design Qualityabstract2.5D integration is gaining popularity primarily due to its ability to facilitate intellectual property (IP) reuse. Unlike conventional 2D and 3D approaches, 2.5D integration requires a more complex design and analysis process and is highly sensitive to changes in design parameters. However, research on the sensitivity of 2.5D design parameters is notably scarce, with most studies still concentrating on 2D and 3D. In this paper, we propose an AI-driven model for predicting sensitivity and an optimization methodology for 2.5D parameters, with a particular focus on bump pitch. Our approach employs advanced machine learning models to accurately predict how variations in bump pitch impact the power, performance, and area of chiplets, as well as the footprint, signal integrity, power integrity, and thermal integrity of the interposer layout. We also utilize Bayesian optimization to identify the optimal bump pitch for specific design objectives. Experimental validation of our model demonstrates high accuracy, with average relative errors of 2.69% for interpolation and 2.7% for extrapolation. Furthermore, optimization results, tailored by adjusting weights for various potential design goals, show an average improvement of 11% in area, wire length, and signal integrity-driven optimization, and 9% in power and thermal integrity-driven optimization. Seungmin Woo, Pruek Vanna-Iampikul, Sung Kyu Lim |
ICCAD | 2 |
| 2024 | Spiking Transformer Hardware Accelerators in 3D IntegrationabstractSpiking neural networks (SNNs) are powerful models of spatiotemporal computation and are well suited for deployment on resource-constrained edge devices and neuromorphic hardware due to their low power consumption. Leveraging attention mechanisms similar to those found in their artificial neural network counterparts, recently emerged spiking transformers have showcased promising performance and efficiency by capitalizing on the binary nature of spiking operations. Recognizing the current lack of dedicated hardware support for spiking transformers, this paper presents the first work on 3D spiking transformer hardware architecture and design methodology. We present an architecture and physical design co-optimization approach tailored specifically for spiking transformers. Through memory-on-logic and logic-on-logic stacking enabled by 3D integration, we demonstrate significant energy and delay improvements compared to conventional 2D CMOS integration. Boxun Xu, Junyoung Hwang, Pruek Vanna-Iampikul, Sung Kyu Lim, Peng Li 0001 |
ICCAD | 3 |
| 2024 | FastTuner: Transferable Physical Design Parameter Optimization using Fast Reinforcement LearningabstractCurrent state-of-the-art Design Space Exploration (DSE) methods in Physical Design (PD), including Bayesian optimization (BO) and Ant Colony Optimization (ACO), mainly rely on black-boxed rather than parametric (e.g., neural networks) approaches to improve end-of-flow Power, Performance, and Area (PPA) metrics, which often fail to generalize across unseen designs as netlist features are not properly leveraged. To overcome this issue, in this paper, we develop a Reinforcement Learning (RL) agent that leverages Graph Neural Networks (GNNs) and Transformers to perform "fast" DSE on unseen designs by sequentially encoding netlist features across different PD stages. Particularly, an attention-based encoder-decoder framework is devised for "conditional" parameter tuning, and a PPA estimator is introduced to predict end-of-flow PPA metrics for RL reward estimation. Extensive studies across 7 industrial designs under the TSMC 28nm technology node demonstrate that the proposed framework FastTuner, significantly outperforms existing state-of-the-art DSE techniques in both optimization quality and runtime. where we observe improvements up to 79.38% in Total Negative Slack (TNS), 12.22% in total power, and 50x in runtime. Hao-Hsiang Hsiao, Yi-Chen Lu, Pruek Vanna-Iampikul, Sung Kyu Lim |
ISPD | 3 |
| 2023 | Glass Interposer Integration of Logic and Memory Chiplets: PPA and Power/Signal Integrity BenefitsabstractGlass interposers enable 3D stacking between the chiplets embedded into the substrate and the ones stacked directly on top, which is not possible in silicon. In this work, we demonstrate the benefits of such stacking in glass interposers over silicon in terms of key system-level metrics including area, wirelength, signal, power, and thermal integrity. We achieve this goal with GDS layouts of both chiplets and interposers and sign-off simulations. Our experiments show that glass offers 2.6X area, 21X wirelength, 17.72% full-chip power, 64.7% signal integrity, and 10X power integrity improvement over silicon at the cost of 15% increase in temperature. Pruek Vanna-Iampikul, Lingjun Zhu, Serhat Erdogan, Mohanalingam Kathaperumal, Ravi Agarwal, Ram Gupta, Kevin Rinebold, Sung Kyu Lim |
DAC | 1 |
| 2023 | A 3D Implementation of Convolutional Neural Network for Fast InferenceabstractLow latency inference has many applications in edge machine learning. In this paper, we present a run-time configurable convolutional neural network (CNN) inference ASIC design for low-latency edge machine learning. By implementing a 5-stage pipelined CNN inference model in a 3D ASIC technology, we demonstrate that the model distributed on two dies utilizing face-to-face (F2F) 3D integration achieves superior performance. Our experimental results show that the design based on 3D integration achieves 43% better energy-delay product when compared to the traditional 2D technology. Narasinga Rao Miniskar, Pruek Vanna-Iampikul, Aaron R. Young, Sung Kyu Lim, Frank Liu 0001, Jieun Yoo, Corrinne Mills, Farah Fahim, Jeffrey S. Vetter |
ISCAS | 2 |
| 2023 | Snap-3D: A Constrained Placement-Driven Physical Design Methodology for High Performance 3-D ICsabstract3-D integration technology is one of the leading options to advance Moore’s Law beyond conventional scaling. One of the 3-D integration choice is the heterogeneous integration with the benefits of power saving over the homogeneous integration. With the lack of commercial 3-D tools, existing 3-D physical design flows utilize 2-D commercial tools to perform 3-D integrated circuit (3-D IC) physical synthesis. Specifically, these flows build 2-D designs first and then convert them into 3-D designs. However, several works demonstrate that design qualities degrade during this 2-D–3-D transformation and some of the flows do not support heterogeneous integration. In this article, we propose Snap-3D, a constraint-driven placement approach to build commercial-quality 3-D ICs, which supports both homogeneous and heterogeneous 3-D ICs. Our key idea is based on the observation that if the standard cell height is contracted and partitioned into multiple tiers, any commercial 2-D placer can place them onto the row structure and naturally achieve high-quality 3-D placement. This methodology is shown to optimize power, performance, and area (PPA) metrics across different tiers simultaneously and minimize the aforementioned design quality loss. Experimental results on seven industrial designs demonstrate that Snap-3D achieves up to 10.9% wirelength, 9% power, and 25% performance improvements compared with state-of-the-art 3-D design flows. Pruek Vanna-Iampikul, Chengjia Shao, Yi-Chen Lu, Sai Pentapati, Yun Heo, Jae-Seung Choi, Sung Kyu Lim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | GNN-based Multi-bit Flip-flop Clustering and Post-clustering Design Optimization for Energy-efficient 3D ICsabstractIn high-performance three-dimensional Integrated Circuits (3D ICs), clock networks consume a large portion of the full-chip power. However, no previous 3D IC work has ever optimized 3D clock networks for both power and performance simultaneously, which results in sub-optimal 3D designs. To overcome this issue, in this article, we propose a GNN-based flip-flop clustering algorithm that merges single-bit flip-flops into multi-bit flip-flops in an unsupervised manner, which jointly optimizes the power and performance metrics of clock networks. Moreover, we integrate our algorithm into the state-of-the-art 3D physical design flow and verify the integration, which leads to a better 3D full-chip design. Experimental results on eight industrial benchmarks demonstrate that the algorithm achieves improvements up to 18% in total power and 8.2% in performance over the state-of-the-art 3D flow. Pruek Vanna-Iampikul, Yi-Chen Lu, Da Eun Shim, Sung Kyu Lim |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2022 | Design Automation and Test Solutions for Monolithic 3D ICsabstractMonolithic 3D (M3D) is an emerging heterogeneous integration technology that overcomes the limitations of the conventional through-silicon-via (TSV) and provides significant performance uplift and power reduction. However, the ultra-dense 3D interconnects impose significant challenges during physical design on how to best utilize them. Besides, the unique low-temperature fabrication process of M3D requires dedicated design-for-test mechanisms to verify the reliability of the chip. In this article, we provide an in-depth analysis on these design and test challenges in M3D. We also provide a comprehensive survey of the state-of-the-art solutions presented in the literature. This article encompasses all key steps on M3D physical design, including partitioning, placement, clock routing, and thermal analysis and optimization. In addition, we provide an in-depth analysis of various fault mechanisms, including M3D manufacturing defects, delay faults, and MIV (monolithic inter-tier via) faults. Our design-for-test solutions include test pattern generation for pre/post-bond testing, built-in-self-test, and test access architectures targeting M3D. Lingjun Zhu, Arjun Chaudhuri, Sanmitra Banerjee, Gauthaman Murali, Pruek Vanna-Iampikul, Krishnendu Chakrabarty, Sung Kyu Lim |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2021 | Snap-3D: A Constrained Placement-Driven Physical Design Methodology for Face-to-Face-Bonded 3D ICsabstract3D integration technology is one of the leading options that can advance Moore's Law beyond conventional scaling. Due to the absence of commercial 3D placers and routers, existing 3D physical design flows rely heavily on 2D commercial tools to handle 3D IC physical synthesis. Specifically, these flows build 2D designs first and then convert them into 3D designs. However, several works demonstrate that design qualities degrade during this 2D-3D transformation. In this paper, we overcome this issue with our Snap-3D, a constraint-driven placement approach to build commercial-quality 3D ICs. Our key idea is based on the observation that if the standard cell height is contracted by one half and partitioned into multiple tiers, any commercial 2D placer can place them onto the row structure and naturally achieve high-quality 3D placement. This methodology is shown to optimize power, performance, and area (PPA) metrics across different tiers simultaneously and minimize the aforementioned design quality loss. Experimental results on 7 industrial designs demonstrate that Snap-3D achieves up to 5.4% wirelength, 10.1% power, and 92.3% total negative slack improvements compared with state-of-the-art 3D design flows. Pruek Vanna-Iampikul, Chengjia Shao, Yi-Chen Lu, Sai Pentapati, Sung Kyu Lim |
ISPD | 1 |
| 2020 | RTL-to-GDS Design Tools for Monolithic 3D ICsabstractIn this paper, we propose RTL-to-GDS design flow for monolithic 3D ICs (M3D) built with carbon nanotube field-effect transistors and resistive memory. Our tool flow is based on commercial 2D tools and smart ways to extend them to conduct M3D design and simulation. We provide a post-route optimization flow, which exploits the full potential of the underlying M3D process design kit (PDK) for power, performance and area (PPA) optimization. We also conduct IR-drop and thermal analysis on M3D designs to improve the reliability. To enhance the testability of our M3D designs, we develop design-for-test (DFT) methodologies and integrate a low-overhead built-in self-test module into our design for testing inter-layer vias (ILVs) as well as logic circuitries in the individual tiers. Our benchmark design is RISC-V Rocketcore, which is an open source processor. Our experiments show 8.1% of power, 19.6% of wirelength and 55.7% of area savings with M3D designs at iso-performance compared to its 2D counterpart. In addition, our IR-drop and thermal analyses indicate acceptable power and thermal integrity in our M3D design. Gauthaman Murali, Pruek Vanna-Iampikul, Dae Hyun Kim 0004, Arjun Chaudhuri, Sanmitra Banerjee, Krishnendu Chakrabarty, Saibal Mukhopadhyay, Sung Kyu Lim |
ICCAD | 3 |