Heechun Park

dblp:139/2586 · DBLP profile ↗
← Back
34ranked-venue papers
8as first author
25since 2021 · last 2026
0000-0003-2796-518XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 34 · 8 first-author · 25 since 2021Software engineering, systems software and programming languages · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021
YearPublicationVenuePosition
2026 Standard Cell Layout Generation: Methodological Evolution and Architectural Impacts
abstract
Over the past decade and beyond, standard cell layout generation has been a cornerstone of digital integrated circuit design, evolving from manual design practices to advanced AI-assisted methodologies. This survey systematically reviews the evolution of standard cell layout generation across two major axes: methodological advances and transistor and cell architecture changes. First, we trace the methodological progression from manual design practices and early design-rule-based approaches to heuristic algorithms, exact optimization methods, and the most recent AI-driven paradigms. Then, we examine the evolution of transistor and cell structures: The transistor from planar CMOS to the emerging vertical devices (CFETs and Flip-FETs); and the cell architecture from the multi-row structures to the integration of buried power rails (BPR) and backside metal layers. Finally, we highlight future research directions including transistor–cell–chip co-design methodologies, optimization techniques that leverage unique electrical characteristics of emerging CMOS devices. This comprehensive survey aims to provide insights into current state-of-the-art techniques while highlighting promising avenues for future innovation in standard cell layout generation for next-generation technologies.
Junghyun Yoon, Ikkyum Kim, Gyumin Kim, Sojung Park, Heechun Park
ASP-DAC5
2026 SMT-Based Optimal Transistor Folding and Placement for Standard Cell Layout Generation
abstract
As semiconductor technology continues to scale, standard cell layout generation becomes increasingly challenging, yet remains critical for Design Technology Co-Optimization (DTCO). This paper presents a satisfiability modulo theories (SMT)-based methodology that tightly integrates transistor folding and placement to optimize standard cell layouts. Our approach offers two key contributions: (1) a unified SMT-based framework for simultaneous transistor folding and placement, which explores folding configurations beyond mandatory constraints to achieve globally optimal placement with improved area and routability; and (2) a mono-dummy insertion strategy based on the longest common subsequence (LCS) algorithm, which aligns PMOS and NMOS transistor chains to enhance diffusion sharing and further reduce layout area. Experimental results using the ASAP7 7 nm PDK show that our SMT-based transistor placement framework produces fully routable layouts with $4.12 \%$ smaller average cell area compared to manually crafted counterparts. Moreover, when combined with heuristic in-cell routing, our method successfully generates layouts for 172 standard cells, including complex cells such as high fan-in gates and flip-flops with asynchronous reset, with transistor placement completing in 836.7 seconds. This work demonstrates that our simultaneous optimization of transistor folding and placement enables both higher layout quality and practical automation for advanced standard cell design.
Junghyun Yoon, Heechun Park
ASP-DAC2
2026 Critical-Path-Centric 3D IC Placement for Timing Optimization
abstract
The vertical integration of 3D ICs introduces physical design complexities beyond those of conventional 2D ICs, especially in the multi-tier placement stage. Existing pseudo-3D design flows typically follow a sequential process of 2D placement followed by tier partitioning, which focuses on balancing cell distribution across tiers but largely overlooks timing-related factors. In this paper, we propose a timing-driven 3D IC placement framework that explicitly prioritizes critical paths throughout the placement process. Starting from an initial 2D placement, we apply an ILP-based tier assignment that encourages critical-path cells to reside on the same tier, reducing unnecessary vertical transitions that degrade timing. We then perform critical-path-aware planar refinement, where the locations of critical-path cells are iteratively adjusted toward path-centric targets, with displacement magnitudes determined by timing criticality. Finally, tier partitioning for non-critical cells is conducted using an enhanced cost function that minimizes interference with previously optimized critical paths, thereby preventing relocation during subsequent legalization. Experimental results demonstrate 13.0% improvement in worst negative slack (WNS) and 7.85% reduction in total negative slack (TNS), validating the effectiveness of our critical path-driven placement strategy.
Sojung Park, Heechun Park
DATE2
2026 Breaking Standard Cell Margin Constraints for Area-Efficient VLSI Design
abstract
In standard-cell-based VLSI design, fixed margins at cell boundaries are necessary to prevent short violations between adjacent transistors carrying different signals. However, these margins are redundant for most abutted cell pairs and incur non-negligible area overhead when accumulated across the chip. In this paper, we present a novel VLSI design optimization framework that eliminates redundant margins by strategically merging adjacent cells into margin-free cells (MF-cells), which preserve the same functionality with reduced area due to the removal of inter-cell margins. Precisely, we identify optimal cell pairs for merging from an initial standard-cell-based placement using a maximum weighted matching (MWM) algorithm. Each identified pair is replaced with an MF-cell and placed at an optimal position using a placement algorithm that minimizes wirelength and routing congestion. Compared to the conventional standard-cell-based design, we achieve on average 3.9% reduction in total cell area and 4.7% reduction in full chip area, leading to 2.7% reduction in total wire length and 2.1% improvement in timing performance while maintaining comparable power consumption. Our framework is a practical approach to achieve meaningful area and timing improvements, which is fully compatible with commercial standard-cell-based VLSI design flow.
Junghyun Yoon, Jooyeon Jeong, Heechun Park
DATE3
2026 Congestion-Aware Integrated Optimization of Multi-Bit Flip-Flop Clustering and Placement
abstract
Multi-bit flip-flops (MBFFs) play an increasingly important role in modern VLSI design due to their ability to significantly reduce clock power by sharing clock drivers across multiple flip-flops (FFs). Conventional MBFF utilization techniques primarily focus on aggressively merging single-bit flip-flops (SBFFs) into MBFFs while preserving timing, aiming to maximize clock power reduction. However, such approaches often overlook the routing congestion that arises as merged FFs become concentrated in a single region. In this paper, we propose a congestion-aware MBFF clustering and placement methodology that simultaneously considers routing congestion as well as conventional power-performance-area (PPA). We first generate an initial solution through aggressive clustering and placement based on FF proximity to maximize the utilization of MBFFs. Then the solution is refined by applying debanking operations and a minimum-cost maximum-flow (MCMF)-based merging algorithm under a timing- and congestion-aware cost model. Finally, the generated MBFFs are placed within timing-preserving regions in a congestion-critical order, effectively mitigating routing congestion and placement degradation induced by aggressive merging. Experimental results demonstrate that the proposed method significantly increases MBFF utilization and reduces clock power and total power by 43.15% and 16.67%, respectively, compared to conventional MBFF optimization approaches. Furthermore, routing congestion overflow and design rule violations (DRVs) are reduced by 4.83% and 3.98%, respectively, while maintaining timing comparable to the original SBFF-based designs.
Jiyun Ham, Heechun Park
ISLPED2
2026 BS-CSO: Backside-Assisted Clock-Signal Co-Optimization Framework
Ikkyum Kim, Geonhyeong Park, Heechun Park
ISLPED3
2026 Post-Placement Optimization with Pin-Level Graph Isomorphism Network and Recurrent Q-Learning
abstract
Post-placement optimization has been actively studied using techniques such as gate sizing, threshold-voltage (Vth) swapping, and buffer insertion. Recently, many approaches have adopted machine learning, particularly reinforcement learning (RL), to automate these optimizations. However, existing RL-based methods fall short of comprehensive modeling of global circuit characteristics, failing to jointly optimize timing and power while covering only narrow combinations of optimization actions. In this work, we propose a deep recurrent Q-network (DRQN)-based framework for post-placement circuit optimization that jointly leverages all three operations. The circuit is sequentially clustered based on timing criticality, and each cluster is modeled using a pin-level Graph Isomorphism Network with Edge features (GINE) combined with a long short-term memory (LSTM) to capture sequential dependencies across gates. Based on this representation, the RL agent simultaneously determines gate sizing and Vth selection, and buffer insertion for each gate and its driving net. To enable integrated optimization of timing and power, we design a reward function that applies symmetric logarithmic compression to each metric at a fixed scale, allowing controllable trade-offs across the design. Experimental results demonstrate that the proposed framework reduces TNS by 96% and total power by 22% on average, compared to 69% and 5% by a timing-only RL baseline, and ablation studies validate the contributions of the GINE model and recurrent architecture.
Kijung Kong, Heechun Park
ISLPED2
2025 PPA-Aware Tier Partitioning for 3D IC Placement with ILP Formulation
abstract
3D ICs are renowned for their potential to enable high-performance and low-power designs by utilizing denser and shorter inter-tier connections. In the physical design flow of 3D ICs, the placement stage includes a differentiated design step to assign instances to different tiers, i.e., top or bottom, called tier partitioning. Despite its importance to overall circuit performance, previous tier partitioning approaches have not taken power-performance-area (PPA) optimization into account, leading to degradation in timing and increased power consumption. In this paper, we propose a novel tier partitioning method in 3D IC placement that concurrently optimizes all PPA-relevant aspects, i.e., intertier cuts, overlapping areas, tier transitions along timing-critical paths, and local/global area balance. We first reduce the problem complexity with netlist clustering based on logical and physical relations, and then formulate an integer-linear programming (ILP) model for each cluster to find an optimal solution. Experiments on various benchmarks demonstrate that our method achieves significant improvements over previous tier partitioning results in terms of all PPA metrics, including 1.23% reduction in power consumption, 24.08% reduction in total negative slack (TNS), and 3.44% reduction in wirelength on average.
Eunsol Jeong, Taewhan Kim 0001, Heechun Park
ASP-DAC3
2025 Clustering-Driven Bonding Terminal Legalization with Reinforcement Learning for F2F 3D ICs
abstract
3D ICs have garnered significant attention as a means to overcome the scaling limits of Moore's Law. Specifically, Face-to-Face (F2F) 3D ICs with hybrid bonding have become a promising solution as they enable heterogeneous stacking with different technology nodes. However, the state-of-the-art F2F 3D IC design flow faces a challenge in locating bonding terminals, since their large bonding pitches compared to the adjacent metal layers cause the existing routing engine to generate a huge amount of design rule violations (DRVs) from the overlapped pitches. In this paper, we introduce our clustering-driven bonding terminal legalization method that efficiently honors bonding pitches. Starting from the optimal 3D via locations with lots of bonding pitch overlaps, our method legalizes all bonding terminals by leveraging both aspects of minimizing displacements (i.e., maintaining routing quality) and grid-based assignments (i.e., overlap elimination) through applying grid-based legalization on clustered bonding terminals. Moreover, we apply reinforcement learning for clustering to obtain the optimal bonding terminal clusters that utilize both advantages. Experiments demonstrate that we successfully eliminate all bonding terminal pitch overlaps with minimal displacement of 3D vias from their optimal locations, thereby reducing the overhead of timing and power degradation than the previous approach.
Gyumin Kim, Heechun Park
ASP-DAC2
2025 Mitigating Routability Problems in Complementary-FET-based VLSI Designs
abstract
As semiconductor technology scales beyond 5 nm, complementary FET (CFET) that stacks P-FET and N-FET enables extreme cell area scaling. However, reduced standard cell height with CFET presents routability challenges due to limited back-end-of-line (BEOL) routing resources. In this paper, we address two key routability problems, i.e., pin accessibility and routing congestion, by employing various pin-extended cells on demand. Moreover, we introduce an end-to-end VLSI design framework that further alleviates routing congestion using partial placement blockages. Experimental results demonstrate that our framework eliminates most design rule violations (DRVs) related to routability while maintaining the area advantages of CFET technology.
Junghyun Yoon, Heechun Park
DAC2
2025 Timing-Driven Detailed Placement with Unsupervised Graph Learning
abstract
Detailed placement is a crucial stage in VLSI design that starts from the global placement result to determine the final legal locations of each cell through fine-grained optimization. Traditional detailed placement methods focus on minimizing the half-perimeter wire length (HPWL) as in global placement. However, incorporating timing-driven placement becomes essential with the increasing complexity of VLSI designs and tighter performance constraints. In this paper, we propose a timing-driven detailed placement framework that leverages unsupervised graph learning techniques. Specifically, we integrate timing-related metrics into the objective function for detailed placement and formulate it into the loss function of a graph neural network (GNN) model. The loss function includes overlap, legality, and timing-related arc lengths, with appropriate weights using Bayesian optimization. Experimental results show that our framework achieves comparable or improved HPWL while significantly reducing total negative slack (TNS) by 5.5%, compared to existing methods.
Dhoui Lim, Heechun Park
DATE2
2025 A Unified Design Flow for Homogeneous and Heterogeneous 3D Integration with Fine-Pitch Hybrid Bonding
abstract
Fine-pitch hybrid bonding offers significant power, performance, and area (PPA) benefits for 3D ICs. However, existing electronic design automation (EDA) tools lack native support for multi-tier integration, prompting the use of pseudo-3D design flows based on conventional 2D tools. Still, these flows often overlook key 3D-specific constraints, such as via count limitations and overlap avoidance, resulting in 3D via overflow and design rule violations of 3D via pitch overlaps. Furthermore, in the context of heterogeneous 3D ICs with different technology nodes, prior works fail to incorporate critical design considerations, including timing-aware cell assignment, congestion management, and wirelength control. In this paper, we propose a unified 3D IC design framework addressing these limitations through four key techniques: (1) Generalized footprint contraction to support both heterogeneous and homogeneous 3D integrations; (2) A constraint-driven tier partitioning for pitch-based 3D via budgeting that reduces 3D via count by up to 62% and maintaining 3D via utilization between 45–89%; (3) A timing-aware gain to steer performance-critical cells to the faster die, improving final WNS and TNS by up to 62% and 74%, respectively. (4) A post-route 3D via legalization stage eliminates all overlapping 3D vias utilizing reinforcement learning (RL) that minimizes design quality degradation. Evaluations on various homogeneous and heterogeneous 3D integrations on representative benchmarks show that our framework consistently improves timing, wirelength, and power, while removing all 3D via violations compared to conventional pseudo-3D flows.
Gyumin Kim, Heechun Park
ICCAD2
2025 Timing-Driven Macro Placement with Connectivity-Aware Clustering
abstract
Macro placement in the floorplanning stage is crucial as it marks the beginning of the entire VLSI physical design flow. Despite its significant impact on the final design quality, identifying the optimal macro placement for sign-off power-performance-area (PPA) remains challenging and often relies on empirical methods. Although recent algorithms for automatic macro placement leverage the relationships among macros and standard cells, they failed to fully address timing-related objectives. In this paper, we propose a timing-driven macro placement framework that utilizes connectivity and clocked element (e.g., flip-flop) information among macros for clustering and integrates dedicated timing costs throughout the entire process. We first utilize the Louvain algorithm to iteratively conduct multi-level clustering on a directed graph derived from the design RTL, with edge weights assigned based on the ratio of the number of logic gates and clocked elements between macros. At each clustering level, we utilize simulated annealing (SA) to determine the clusters' shape and placement. After the last iteration, we conduct macro placement per cluster using SA with our timing-related cost function. Finally, we refine the macro placement result to escape from the bottom-left arrangement, again guided by the timing-related cost function. Experimental results demonstrate that our macro placement framework successfully improves timing metrics, WNS and TNS, by 16.87% and 18.49% over state-of-the-art open-source macro placement, and by 6.59% and 15.54% over a commercial macro placement engine, respectively.
Gangmin Jeon, Heechun Park
ISLPED2
2024 Memristive Logic-in-Memory Implementation with Area Efficiency and Parallelism
abstract
In-memory computing that performs logic operations inside the memory unit has emerged as an alternative to the traditional von Neumann architecture. In particular, logic-in-memory (LiM) architecture with memristor crossbars shows promising results in terms of power and performance, which operates the Boolean functions within the crossbar in parallel by utilizing memristor-aided logic (MAGIC) operations. In this work, we propose an efficient LUT mapping framework of Boolean logic functions into the memristor crossbar, which leverages the parallelism of logic computation to maximize performance and enhances area efficiency by optimizing crossbar utility. In detail, we map the LUTs of NoN (NOR-of-NOR) into the crossbar in a topological order in a way that maximizes the horizontal and vertical NOR operations. Moreover, we add an algorithmic rearranging stage for the intermediate results to more utilize the crossbar space during the next LUT placement. Our experiments prove that the proposed LUT mapping framework incorporates parallel operations with higher area efficiency, thereby achieving a 13.81 % speedup on average with a smaller crossbar dimension compared to the previous approach. We also successfully map the benchmark circuits with a large number of LUTs into the small-sized crossbar, in which the previous method has failed.
Ikkyum Kim, Heechun Park
ICCD2
2024 DTOC-P: Deep-Learning-Driven Timing Optimization Using Commercial EDA Tool With Practicality Enhancement
abstract
Deep learning (DL) models have recently paid considerable attention to timing prediction in the place-and-route (P&R) flow. As yet, the DL-based prior works are confined to timing prediction at the time-consuming routing stage, and very few have addressed the timing prediction problem at the placement, i.e., at the pre-route stage. Moreover, no work has addressed a seamless link of timing prediction at the pre-route stage to the final timing optimization through commercial P&R tools. In this work, we introduce a novel framework called DTOC-P that seamlessly integrates deep-learning-driven timing optimization into cutting-edge commercial P&R tools. Our framework is composed of two phases: (1) the pre-route timing prediction phase that performs DL-driven arc delay and arc output slew prediction with an elaborated hierarchical model; (2) the timing optimization phase which incorporates commercial P&R tools with DL-driven prediction outcomes to perform timing optimization. In addition, DTOC-P framework achieves enhanced practicality with the application of continual learning in the timing prediction phase, and the concept of anomaly detection in the timing optimization phase. Experimental results show that our DTOC-P framework improves pre-route prediction accuracy by up to 55% and 47% on arc delay and arc output, which are further enhanced to encompass a broader range of designs by continual learning supported in DTOC-P, practically using a tenfold reduced training time compared to re-training all datasets from scratch. In terms of timing optimization, our experiments reveal that DTOC-P framework improves WNS, TNS, and the number of timing violation paths by up to 12%, 41%, and 34%, respectively, which is a remarkable progress compared to its predecessor through the integration of anomaly detection that excludes potential outliers to effectively protects against erroneous timing updates during the timing optimization phase.
Jaehoon Ahn, Kyungjoon Chang, Kyumyung Choi, Taewhan Kim 0001, Heechun Park
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 Comprehensive Physical Design Flow Incorporating 3-D Connections for Monolithic 3-D ICs
abstract
In this paper, we propose a comprehensive physical design flow specifically tailored for Monolithic 3D (M3D) integration, a transformative technology for high-density and highperformance IC design in the post-Moore era. Unlike conventional RTL-to-GDS flows that heavily focus on utilizing commercial 2D design tools, our design flow delves deep into the suboptimal issues inherent in implementing cross-tier connections, which are not adequately addressed by 2D tools. Our proposed flow provides seamless optimization for such connections through three key design stages following pseudo-3D placement: (1) 3D routing-aware tier partitioning that induces subtle imbalances in cell area distribution between tiers to maximize the utilization of monolithic inter-tier vias (MIVs); (2) MIV-guided detailed placement that optimizes the placement by strategically utilizing reserved whitespace for enhanced 3D connections; and (3) MIV-aware 3D routing that takes full advantage of the finetuned placement result. Experiment results using open-source benchmark circuits in advanced 7nm technology nodes show that our proposed M3D design flow achieves up to 9.92% wirelength reduction per 3D net, resulting in 76.70% improvement in worst negative slack, and an equivalently improved 60.28% energy-delay-product over the state-of-the-art M3D design flow on average even with challenging design conditions. We provide valuable insights into various factors for efficient and high-quality M3D IC design with effective solutions.
Suwan Kim, Heechun Park
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 DTOC: integrating Deep-learning driven Timing Optimization into the state-of-the-art Commercial EDA tool
abstract
Recently, deep-learning (DL) models have paid a considerable attention to timing prediction in the placement and routing (P&R) flow. As yet, the DL-based prior works are confined to timing prediction at the time-consuming global routing stage, and very few have addressed the timing prediction problem at the placement, i.e., at the pre-route stage. This is because it is not easy to “accurately” predict various timing parameters at the pre-route stage. Moreover, no work has addressed a seamless link of timing prediction at the pre-route stage to the final timing optimization through making use of commercial P&R tools. In this work, we propose a framework called DTOC, to be used at the pre-route stage for this end. Precisely, the framework is composed of two models: (1) a DL-driven arc delay and arc output slew prediction model, performing in two levels: (level-1) predicting net resistance (R), net capacitance (C), and arc length (Len), followed by (level-2) predicting arc delay and arc output slew from the R/C/Len prediction obtained in (level-1); (2) a timing optimization model, which uses the inference outcomes in our DL-driven prediction model to enable the commercial P&R tools to calculate the full path delays, setting update timing margins on paths, so that the P&R tools should use more accurate margins on timing optimization. Experimental results show that, by using our DTOC framework during timing optimization in P&R, we improve the pre-route prediction accuracy on arc delay and arc output slew by 20~26% on average, and improve the WNS, TNS, and the number of timing violation paths by 50~63 % on average.
Kyungjoon Chang, Jaehoon Ahn, Heechun Park, Kyu-Myung Choi, Taewhan Kim 0001
DATE3
2023 Eliminating Minimum Implant Area Violations With Design Quality Preservation
abstract
Minimum implant area (MIA) violation has emerged in the sub-micrometer technology which requires a certain amount of threshold voltage ($V_{\text {t}}$) area for the fabrication. Elimination of MIA violations in the sign-off layout thus becomes an inevitable task for a high-performance multiple-$V_{\text {t}}$design. Conventional approaches as well as the previous efforts to remove MIA violations bring severe defects to the final design in that locally moving cells or reassigning$V_{\text {t}}\text{s}$make the timing constraints unsatisfied or power consumption to be exploded. In this article, we propose a comprehensive MIA violation removal algorithm that fully and systematically controls the timing budget and power overhead with three sequential steps: 1) removing intra-row MIA violations by$V_{\text {t}}$reassignment under timing preservation and minimal power increments; 2) removing inter-row MIA violations with a theoretically optimal$V_{\text {t}}$reassignment while satisfying timing constraints; and 3) refining$V_{\text {t}}$reassignment to recover the power loss without violating both MIA constraints and timing closure. Moreover, we introduce a preprocessing algorithm at the preroute stage to remove a huge amount of MIA violations in advance for an additional runtime reduction without design quality degradation. Experiments through benchmark circuits show that our proposed approach completely resolve MIA violations while ensuring no timing violation and using 34.6% less power overhead on average than the conventional approaches and previous works. In addition, our preprocessing step reduces 45%–88% of MIA violations before the routing stage, which incurs 41% faster MIA removal on average in the final stage with similar design quality.
Eunsol Jeong, Taewhan Kim 0001, Heechun Park
IEEE Trans. Very Large Scale Integr. Syst.3
2022 A Systematic Removal of Minimum Implant Area Violations under Timing Constraint
abstract
Fixing minimum implant area (MIA) violations in the post-route layout is an essential and inevitable task for the high-performance designs employing multiple threshold voltages. Unlike the conventional approaches, which have tried to locally move cells or reassign$V_{t}$(threshold voltage) of some cells in a way to resolve the MIA violations with little or no consideration of timing constraint, our proposed approach fully and systematically controls the timing budget during the removal of MIA violations. Precisely, our solution consists of three sequential steps: (1) performing critical path aware cell selection for$V_{t}$reassignment to fix the intra-row MIA violations while considering timing constraint and minimal power increments; (2) performing a theoretically optimal$V_{t}$reassignment to fix the inter-row MIA violations while satisfying both of the intra-row MIA and timing constraints; (3) refining$V_{t}$reassignment to further reduce the power consumption while meeting intra- and inter-row MIA constraints as well as timing constraints. Experiments through benchmark circuits show that our proposed approach is able to completely resolve MIA violations while ensuring no timing violation and achieving much less power increments over that by the conventional approaches.
Eunsol Jeong, Heechun Park, Taewhan Kim 0001
DATE2
2022 Tightly Linking 3D Via Allocation Towards Routing Optimization for Monolithic 3D ICs
abstract
Monolithic 3D (M3D) is a revolutionary technology for high-density and high-performance chip design in the post-Moore era. However, it suffers from considerable thermal confinement due to the transistor stacking and insulating materials between the layers. As a way of reducing power, thereby mitigating the thermal problem, we propose a comprehensive physical design methodology that incorporates two new important items, one is blockage aware MIV (monolithic inter-tier via) placement and the other is 3D net ordering for routing, intending to optimize wire length. Precisely, we propose a three-step approach: (1) retrieving the MIV region candidates for each 3D net, (2) fine-tuning placement to secure MIV spots in the presence of blockages, and (3) performing M3D routing with net ordering to consider the fine-tuned placement result. We implement the proposed M3D design flow by utilizing commercial 2D IC EDA tools while providing seamless optimization for cross-tier connections. In the meantime, our experiments confirm that proposed M3D design flow saves wire length per cross-tier net by up to 41.42%, which corresponds to 7.68% less total net switching power, equivalently 36.79% lower energy-delay-product over the conventional state-of-the-art M3D design flow.
Suwan Kim, Sehyeon Chung, Taewhan Kim 0001, Heechun Park
ISLPED4
2022 Speeding-up neuromorphic computation for neural networks: Structure optimization approach
Heechun Park, Taewhan Kim 0001
Integr.1
2022 Built-in Self-Test and Fault Localization for Inter-Layer Vias in Monolithic 3D ICs
abstract
Monolithic 3D (M3D) integration provides massive vertical integration through the use of nanoscale inter-layer vias (ILVs). However, high integration density and aggressive scaling of the inter-layer dielectric make ILVs especially prone to defects. We present a low-cost built-in self-test (BIST) method that requires only two test patterns to detect opens, stuck-at faults, and bridging faults (shorts) in ILVs. We also propose an extended BIST architecture for fault detection, called Dual-BIST, to guarantee zero ILV fault masking due to single BIST faults and negligible ILV fault masking due to multiple BIST faults. We analyze the impact of coupling between adjacent ILVs arranged in a 1D array in block-level partitioned designs. Based on this analysis, we present a novel test architecture called Shared-BIST with the added functionality of localizing single and multiple faults, including coupling-induced faults. We introduce a systematic clustering-based method for designing and integrating a delay bank with the Shared-BIST architecture for testing small-delay defects in ILVs with minimal yield loss. Simulation results for four two-tier M3D benchmark designs highlight the effectiveness of the proposed BIST framework.
Arjun Chaudhuri, Sanmitra Banerjee, Heechun Park, Bon Woong Ku, Sukeshwar Kannan, Krishnendu Chakrabarty, Sung Kyu Lim
ACM J. Emerg. Technol. Comput. Syst.4
2021 Pseudo-3D Physical Design Flow for Monolithic 3D ICs: Comparisons and Enhancements
abstract
Studies have shown that monolithic 3D ( M3D ) ICs outperform the existing through-silicon-via ( TSV ) -based 3D ICs in terms of power, performance, and area ( PPA ) metrics, primarily due to the orders of magnitude denser vertical interconnections offered by the nano-scale monolithic inter-tier vias. In order to facilitate faster industry adoption of the M3D technologies, physical design tools and methodologies are essential. Recent academic efforts in developing an EDA algorithm for 3D ICs, mainly targeting placement using TSVs, are inadequate to provide commercial-quality GDS layouts. Lately, pseudo-3D approaches have been devised, which utilize commercial 2D IC EDA engines with tricks that help them operate as an efficient 3D IC CAD tool. In this article, we provide thorough discussions and fair comparisons (both qualitative and quantitative) of the state-of-the-art pseudo-3D design flows, with analysis of limitations in each design flow and solutions to improve their PPA metrics. Moreover, we suggest a hybrid pseudo-3D design flow that achieves both benefits. Our enhancements and the inter-mixed design flow, provide up to an additional 26% wirelength, 10% power consumption, and 23% of power-delay-product improvements.
Heechun Park, Bon Woong Ku, Kyungwook Chang, Da Eun Shim, Sung Kyu Lim
ACM Trans. Design Autom. Electr. Syst.1
2021 Allocation of Always-On State Retention Storage for Power Gated Circuits - Steady-State- Driven Approach
abstract
It is generally known that a considerable portion of flip-flops in circuits is occupied by the ones with mux-feedback loop (called self-loop), which is the critical (inherently unavoidable) bottleneck in minimizing total (always-on) storage size for the allocation of nonuniform multibits for retaining flip-flop states in power gated circuits. This is because it is necessary to replace every self-loop flip-flop with a distinct retention flip-flop with at least one-bit storage for retaining its state since there is no clue where the flip-flop state, when waking up, comes from, i.e., from the mux-feedback loop or from the driving flip-flops other than itself. This work breaks this bottleneck by safely treating a large portion of the self-loop flip-flops as if they were the same as the flip-flops with no self-loop. Specifically, we design a novel mechanism of steady-state monitoring, operating for a few cycles just before sleeping, on a partial set of self-loop flip-flops, by which the expensive state retention storage is never be needed for the monitored flip-flops, contributing to a significant saving on the total size of the always-on state retention storage for power gating. Through experiments with benchmark circuits, it is shown that our proposed method is able to reduce the total number of retention bits by 27.12% on average when at most 2-bit retention flip-flop is used, saving standby power by 19.41% compared with the state-of-the-art conventional method.
Taehwan Kim 0007, Heechun Park, Taewhan Kim 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2021 Clock Delivery Network Design and Analysis for Interposer-Based 2.5-D Heterogeneous Systems
abstract
The 2-D CMOS process technology scaling may have reached its pinnacle, yet it is not feasible to manufacture all computing elements at lower technological nodes. This has opened a new branch of chip designing that allows chiplets on different technological nodes to be integrated into a single package using interposers, the passive interconnection mediums. However, establishing a high-frequency communication over an entirely passive layer is one of the significant design challenges of 2.5-D systems. In this article, we present a robust clocking architecture for a 2.5-D system consisting of 64 processor cores. This clocking scheme consists of two major components, namely, interposer clocking and on-chiplet clocking. The interposer clocking consists of clocks used to achieve global synchronicity and clocks for interchiplet communication established using the AIB protocol. We synthesized these clocking components using commercial EDA tools and analyzed them using standard tools, on-chip, and package models. We also compare these results against a 2-D design of the same benchmark and another 2.5-D clocking architecture. Our experiments show that the absolute clock power is up to 16% less, and the ratio of clock power to system power is up to 4% less in the 2.5-D design than its 2-D counterpart.
Gauthaman Murali, Heechun Park, Eric Qin 0001, Hakki Mert Torun, Majid Ahadi Dolatsara, Madhavan Swaminathan, Tushar Krishna, Sung Kyu Lim
IEEE Trans. Very Large Scale Integr. Syst.2
2020 Pseudo-3D Approaches for Commercial-Grade RTL-to-GDS Tool Flow Targeting Monolithic 3D ICs
abstract
Despite the recent academic efforts to develop Electronic Design Automation (EDA) algorithms for 3D ICs, the current market does not have commercial 3D computer-aided design (CAD) tools. Insteadpseudo-3D alternative design flows have been devised which utilize commercial 2D CAD engines with tricks that help them operate as a fairly-efficient 3D CAD tool. In this paper we provide detailed discussions and fair power-performance-area (PPA) comparisons of state-of-the-art pseudo-3D design flows. We also analyze the limitations of each design flow and provide solutions with better PPA and various design options. Our experiments using commercial PDK, GDS layouts, and sign-off simulations demonstrate that we achieve up to 26% wirelength and 10% power consumption reduction for pseudo-3D design flows. We also provide a partitioning-first scheme to partitioning-last design flow which increases design freedom with tolerable PPA degradation.
Heechun Park, Bon Woong Ku, Kyungwook Chang, Da Eun Shim, Sung Kyu Lim
ISPD1
2020 Architecture, Chip, and Package Codesign Flow for Interposer-Based 2.5-D Chiplet Integration Enabling Heterogeneous IP Reuse
abstract
A new trend in system-on-chip (SoC) design is chiplet-based IP reuse using 2.5-D integration. Complete electronic systems can be created through the integration of chiplets on an interposer, rather than through a monolithic flow. This approach expands access to a large catalog of off-the-shelf intellectual properties (IPs), allows reuse of them, and enables heterogeneous integration of blocks in different technologies. In this article, we present a highly integrated design flow that encompasses architecture, circuit, and package to build and simulate heterogeneous 2.5-D designs. Our target design is 64core architecture based on Reduced Instruction Set Computer (RISC)-V processor. We first chipletize each IP by adding logical protocol translators and physical interface modules. We convert a given register transfer level (RTL) for 64-core processor into chiplets, which are enhanced with our centralized network-onchip. Next, we use our tool to obtain physical layouts, which is subsequently used to synthesize chip-to-chip I/O drivers and these chiplets are placed/routed on a silicon interposer. Our package models are used to calculate power, performance, and area (PPA) and reliability of 2.5-D design. Our design space exploration (DSE) study shows that 2.5-D integration incurs 1.29× power and 2.19× area overheads compared with 2-D counterpart. Moreover, we perform DSE studies for power delivery scheme and interposer technology to investigate the tradeoffs in 2.5-D integrated chip (IC) designs.
Gauthaman Murali, Heechun Park, Eric Qin 0001, Hyoukjun Kwon, Venakata Chaitanya Krishna Chekuri, Nael Mizanur Rahman, Nihar Dasari, Minah Lee, Hakki Mert Torun, Kallol Roy, Madhavan Swaminathan, Saibal Mukhopadhyay, Tushar Krishna, Sung Kyu Lim
IEEE Trans. Very Large Scale Integr. Syst.3
2019 Architecture, Chip, and Package Co-design Flow for 2.5D IC Design Enabling Heterogeneous IP Reuse
abstract
A new trend in complex SoC design is chiplet-based IP reuse using 2.5D integration. In this paper we present a highly-integrated design flow that encompasses architecture, circuit, and package to build and simulate heterogeneous 2.5D designs. We chipletize each IP by adding logical protocol translators and physical interface modules. These chiplets are placed/routed on a silicon interposer next. Our package models are then used to calculate PPA and signal/power integrity of the overall system. Our design space exploration study using our tool flow shows that 2.5D integration incurs 2.1x PPA overhead compared with 2D SoC counterpart.
Gauthaman Murali, Heechun Park, Eric Qin 0001, Hyoukjun Kwon, Venakata Chaitanya Krishna Chekuri, Nihar Dasari, Minah Lee, Hakki Mert Torun, Kallol Roy, Madhavan Swaminathan, Saibal Mukhopadhyay, Tushar Krishna, Sung Kyu Lim
DAC3
2019 RTL-to-GDS Tool Flow and Design-for-Test Solutions for Monolithic 3D ICs
abstract
Monolithic 3D IC overcomes the limitation of the existing through-silicon-via (TSV) based 3D IC by providing denser vertical connections with nano-scale inter-layer vias (ILVs). In this paper, we demonstrate a thorough RTL-to-GDS design flow for monolithic 3D IC, which is based on commercial 2D place-and-route (P&R) tools and clever ways to extend them to handle 3D IC designs and simulations. We also provide a low-cost built-in-self-test (BIST) method to detect various faults that can occur on ILVs. Lastly, we present a resistive random access memory (ReRAM) compiler that generates memory modules that are to be integrated in monolithic 3D ICs.
Heechun Park, Kyungwook Chang, Bon Woong Ku, Daehyun Kim 0002, Arjun Chaudhuri, Sanmitra Banerjee, Saibal Mukhopadhyay, Krishnendu Chakrabarty, Sung Kyu Lim
DAC1
2019 Built-in Self-Test for Inter-Layer Vias in Monolithic 3D ICs
abstract
Monolithic 3D integration provides massive vertical integration through the use of nanoscale inter-layer vias (ILVs). However, high integration density and aggressive scaling of the inter-layer dielectric make ILVs especially prone to defects. We present a low-cost built-in self-test (BIST) method to detect opens, stuck-at faults (SAFs), and bridging faults (shorts) in ILVs. Two test patterns-all-1s and all-0s-are applied to the input side of a set of ILVs (e.g., making up a bus between two tiers). On the adjacent tier (the output side of the ILVs), the test responses are compacted to a 2-bit signature through space compaction. We prove that this compaction solution does not introduce any fault aliasing. Simulations results using HSPICE and M3D benchmark designs show that the proposed BIST method requires low area overhead and test time, but provides effective fault localization and the detectability of a wide range of resistive faults.
Arjun Chaudhuri, Sanmitra Banerjee, Heechun Park, Bon Woong Ku, Krishnendu Chakrabarty, Sung Kyu Lim
ETS3
2019 Hybrid asynchronous circuit generation amenable to conventional EDA flow
Heechun Park, Taewhan Kim 0001
Integr.1
2018 Structure optimizations of neuromorphic computing architectures for deep neural network
abstract
This work addresses a new structure optimization of neuromorphic computing architectures. This enables to speed up the DNN (deep neural network) computation twice as fast as, theoretically, that of the existing architectures. Precisely, we propose a new structural technique of mixing both of the dendritic and axonal based neuromorphic cores in a way to totally eliminate the inherent non-zero waiting time between cores in the DNN implementation. In addition, in conjunction with the new architecture we propose a technique of maximally utilizing computation units so that the resource overhead of total computation units can be minimized. We have provided a set of experimental data to demonstrate the effectiveness (i.e., speed and area) of our proposed architectural optimizations: ~2× speedup with no accuracy penalty on the neuromorphic computation or improved accuracy with no additional computation time.
Heechun Park, Taewhan Kim 0001
DATE1
2015 Synthesis of TSV Fault-Tolerant 3-D Clock Trees
abstract
In through-silicon-via (TSV) based 3-D integrated chips (ICs), synthesizing 3-D clock tree is one of the most challenging tasks. Since the clock signal is delivered to clock sinks (e.g., latches, flip-flops) through TSVs, any fault on a TSV in the clock tree may cause a chip failure. Therefore, ensuring the reliability of clock TSVs in 3-D ICs is highly important. To cope with clock TSV reliability problem effectively, we propose a new circuit cell called slew-controlled TSV fault-tolerant unit (SC-TFU) which overcomes the limited capability of the conventional TFUs and propose a full solution to the problem of designing and synthesizing 3-D TSV fault-tolerant clock tree based on SC-TFUs. Precisely, for a presynthesized 3-D clock tree, we solve the problem in three steps: 1) performing a comprehensive TSV pairing algorithm to maximally allocate SC-TFUs; 2) replacing TSV pairs obtained in step 1 with SC-TFUs followed by TSV tripling to maximize TSV fault-tolerance under wire and time constraints; and 3) performing a global clock skew tuning process on the SC-TFU embedded 3-D clock tree produced in step 2. Through out experiments, two outstanding benefits are confirmed: 1) our synthesis using SC-TFUs enables a large number of clock TSVs to be paired or tripled to ensure a very high degree of TSV fault-tolerance and 2) our synthesis flow effectively performs tuning of global clock skew whose variation is caused by the inclusion of TSV fault-tolerant cells into 3-D clock trees.
Heechun Park, Taewhan Kim 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2013 Comprehensive technique for designing and synthesizing TSV fault-tolerant 3D clock trees
abstract
Recently, to cope with clock TSV (Through-Silicon-Via) reliability problem efficiently, a new circuit structure called TSV Fault-tolerant Unit (TFU) and the allocation method of TFUs have been proposed. However, the existing design methods partially or never addressed following key issues: (1) the feasibility of TSV pairing for TFU allocation, (2) maximizing TSV pairing, (3) supporting the slew and delay control capability in TFU for the cases of pre-bond testing as well as post-bond stage, and (4) minimizing the impact of TFU insertion on the clock skew of the whole 3D clock tree. In this work, we propose a full solution to the problem of designing and synthesizing TSV fault-tolerant clock tree from a 3D clock tree, which effectively addresses above key issues.
Heechun Park, Taewhan Kim 0001
ICCAD1