EDBT 2026 Demo / reviewers in the wild / expert
Dae Hyun Kim 0004
dblp:85/5149-4 · also Dae-Hyun Kim 0004, Daehyun Kim 0004
· DBLP profile ↗
31ranked-venue papers
7as first author
7since 2021 · last 2024
0000-0001-8275-5949ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 7 first-author · 4 since 2021Theory of computation · 3 · 2 since 2021Software engineering, systems software and programming languages · 2Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Construction of All Multilayer Monolithic RSMTs and Its Application to Monolithic 3D IC RoutingabstractMonolithic three-dimensional (M3D) integration allows ultra-thin silicon tier stacking in a single package. The high-density stacking is acquiring interest and is becoming more popular for smaller footprint areas, shorter wirelength, higher performance, and lower power consumption than the conventional planar fabrication technologies. The physical design of M3D integrated circuits requires several design steps, such as three-dimensional (3D) placement, 3D clock-tree synthesis, 3D routing, and 3D optimization. Among these, 3D routing is significantly time consuming due to countless routing blockages. Therefore, 3D routers proposed in the literature insert monolithic interlayer vias (MIVs) and perform tier-by-tier routing in two substeps. In this article, we propose an algorithm to build a routing topology database (DB) used to construct all multilayer monolithic rectilinear Steiner minimum trees on the 3D Hanan grid. To demonstrate the effectiveness of the DB in various applications, we use the DB to construct timing-driven 3D routing topologies and perform congestion-aware global routing on 3D designs. We anticipate that the algorithm and the DB will help 3D routers reduce the runtime of the MIV insertion step and improve the quality of the 3D routing. Monzurul Islam Dewan, Sheng-En David Lin, Dae Hyun Kim 0004 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2023 | Dual-Purpose Hardware Algorithms and Architectures - Part 2: Integer DivisionabstractInteger division is different from floating-point division in that (1) the execution time of an integer division is highly dependent on the leading 1 locations of operands, (2) x and −x have different magnitude parts if x is a two’s complement integer, and (3) rounding is not necessary. In this paper, we apply the interval-analysis-based division algorithm proposed in "Dual-Purpose Hardware Algorithms and Architecture – Part 1: Floating-Point Division" [1] to offline and online integer division. We implement four online integer dividers using the algorithm, compare them with other dividers, and present detailed simulation results with in-depth analysis of the dividers. We find that the online dividers outperform the offline dividers when the waiting time for online operands goes up and the number of quotient bits to obtain goes down. Jihee Seo, Dae Hyun Kim 0004 |
ARITH | 2 |
| 2023 | Dual-Purpose Hardware Algorithms and Architectures - Part 1: Floating-Point DivisionabstractDivision is a time-consuming, but frequently-used arithmetic operation, so an enormous amount of effort has been made to improve the performance of dividers. Most of the division algorithms in the literature are offline algorithms that minimize the execution time of a single division, whereas some others are online algorithms that maximize the throughput (# divisions executed per time). In this paper, we propose an interval-analysis-based normal-binary floating-point division algorithm that can be used for both offline and online division. We implement two offline and four online dividers using the algorithm and compare them with recently-proposed offline and online dividers. The simulation results show that the offline versions are the best for a Binary64 offline division, whereas the online versions are the best for a Binary64 online division. Jihee Seo, Dae Hyun Kim 0004 |
ARITH | 2 |
| 2022 | Bayesian Optimization over Permutation SpacesabstractOptimizing expensive to evaluate black-box functions over an input space consisting of all permutations of d objects is an important problem with many real-world applications. For example, placement of functional blocks in hardware design to optimize performance via simulations. The overall goal is to minimize the number of function evaluations to find high-performing permutations. The key challenge in solving this problem using the Bayesian optimization (BO) framework is to trade-off the complexity of statistical model and tractability of acquisition function optimization. In this paper, we propose and evaluate two algorithms for BO over Permutation Spaces (BOPS). First, BOPS-T employs Gaussian process (GP) surrogate model with Kendall kernels and a Tractable acquisition function optimization approach to select the sequence of permutations for evaluation. Second, BOPS-H employs GP surrogate model with Mallow kernels and a Heuristic search approach to optimize the acquisition function. We theoretically analyze the performance of BOPS-T to show that their regret grows sub-linearly. Our experiments on multiple synthetic and real-world benchmarks show that both BOPS-T and BOPS-H perform better than the state-of-the-art BO algorithm for combinatorial spaces. To drive future research on this important problem, we make new resources and real-world benchmarks available to the community. Aryan Deshwal, Syrine Belakaria, Janardhan Rao Doppa, Dae Hyun Kim 0004 |
AAAI | 4 |
| 2022 | A ReRAM Memory Compiler for Monolithic 3D Integrated Circuits in a Carbon Nanotube ProcessabstractWe present a ReRAM memory compiler for monolithic 3D (M3D) integrated circuits (IC). We develop ReRAM architectures for M3D ICs using 1T-1R bit cells and single and multiple tiers of transistors for access and peripheral circuits. The compiler includes an automated flow for generation of subarrays of different dimensions and larger arrays of a target capacity by integrating multiple subarrays. The compiler is demonstrated using an M3D process design kit (PDK) based on a Carbon Nanotube Transistor technology. The PDK includes multiple layers of transistors and back-end-of-the-line integrated ReRAM. Simulations show the compiled ReRAM macros with multiple tiers of transistors reduces footprint and improves performance over the macros with single-tier transistors. The compiler creates layout views that are exported into library exchange format or graphic data system for full-array assembly and schematic/symbol views to extract per-bit read/write energy and read latency. Comparison of the proposed M3D subarray architectures with baseline 2D subarrays, generated with a custom-designed set of bit cells and peripherals, demonstrate up to 48% area reduction and 13% latency improvement. Dae Hyun Kim 0004, Sung Kyu Lim, Saibal Mukhopadhyay |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2022 | Design Automation Algorithms for the NP-Separate VLSI Design MethodologyabstractThe NP-Separate design methodology for very-large-scale integration (VLSI) design fine-controls the sizes of transistors, thereby achieving significant power, performance, and area improvement compared to the conventional standard-cell-based design methodology. NP-Separate uses NP cells formed by merging and routing N and P cells having only NFETs and PFETs, respectively. The NP cell formation, however, should be automated to design large circuits using the NP-Separate design methodology. In this paper, we propose design automation algorithms to create NP cells automatically. Simulation results show that the automated NP-Separate reduces the design time significantly, decreases the coupling capacitance by 13%, the critical path delay by 6%, and the power consumption by 10% on average compared to the manual NP-Separate designs. We also propose a detailed placement algorithm to generate more compact VLSI layouts with a little wirelength overhead. The combined effect reduces the coupling capacitance by 10%, the critical path delay by 5%, and the power consumption by 10% on average compared to the manual NP-Separate designs. Monzurul Islam Dewan, Dae Hyun Kim 0004 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2021 | HeM3D: Heterogeneous Manycore Architecture Based on Monolithic 3D Vertical IntegrationabstractHeterogeneous manycore architectures are the key to efficiently execute compute- and data-intensive applications. Through-silicon-via (TSV)-based 3D manycore system is a promising solution in this direction as it enables the integration of disparate computing cores on a single system. Recent industry trends show the viability of 3D integration in real products (e.g., Intel Lakefield SoC Architecture, the AMD Radeon R9 Fury X graphics card, and Xilinx Virtex-7 2000T/H580T, etc.). However, the achievable performance of conventional TSV-based 3D systems is ultimately bottlenecked by the horizontal wires (wires in each planar die). Moreover, current TSV 3D architectures suffer from thermal limitations. Hence, TSV-based architectures do not realize the full potential of 3D integration. Monolithic 3D (M3D) integration, a breakthrough technology to achieve “More Moore and More Than Moore,” opens up the possibility of designing cores and associated network routers using multiple layers by utilizing monolithic inter-tier vias (MIVs) and hence, reducing the effective wire length. Compared to TSV-based 3D integrated circuits (ICs), M3D offers the “true” benefits of vertical dimension for system integration: the size of an MIV used in M3D is over 100 × smaller than a TSV. This dramatic reduction in via size and the resulting increase in density opens up numerous opportunities for design optimizations in 3D manycore systems: designers can use up to millions of MIVs for ultra-fine-grained 3D optimization, where individual cores and routers can be spread across multiple tiers for extreme power and performance optimization. In this work, we demonstrate how M3D-enabled vertical core and uncore elements offer significant performance and thermal improvements in manycore heterogeneous architectures compared to its TSV-based counterpart. To overcome the difficult optimization challenges due to the large design space and complex interactions among the heterogeneous components (CPU, GPU, Last Level Cache, etc.) in a M3D-based manycore chip, we leverage novel design-space exploration algorithms to trade off different objectives. The proposed M3D-enabled heterogeneous architecture, called HeM3D , outperforms its state-of-the-art TSV-equivalent counterpart by up to 18.3% in execution time while being up to 19°C cooler. Aqeeb Iqbal Arka, Biresh Kumar Joardar, Ryan Gary Kim, Dae Hyun Kim 0004, Janardhan Rao Doppa, Partha Pratim Pande |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2020 | Power, Performance, and Thermal Trade-offs in M3D-enabled Manycore ChipsabstractMonolithic 3D (M3D) technology enables unprecedented degrees of integration on a single chip. The miniscule monolithic inter-tier vias (MIVs) in M3D are the key behind higher transistor density and more flexibility in designing circuits compared to conventional through silicon via (TSV)-based architectures. This results in significant performance and energyefficiency improvements in M3D-based systems. Moreover, the thin inter-layer dielectric (ILD) used in M3D provides better thermal conductivity compared to TSV-based solutions and eliminates the possibility of thermal hotspots. However, the fabrication of M3D circuits still suffers from several non-ideal effects. The thin ILD layer may cause electrostatic coupling between tiers. Furthermore, the low-temperature annealing degrades the top-tier transistors and bottom-tier interconnects. An NoC-based manycore design needs to consider all these M3D- process related non-idealities. In this paper, we discuss various design challenges for an M3D-enabled manycore chip. We present the power-performance-thermal trade-offs associated with these emerging manycore architectures. Shouvik Musavvir, Anwesha Chatterjee, Ryan Gary Kim, Dae Hyun Kim 0004, Janardhan Rao Doppa, Partha Pratim Pande |
DATE | 4 |
| 2020 | RTL-to-GDS Design Tools for Monolithic 3D ICsabstractIn this paper, we propose RTL-to-GDS design flow for monolithic 3D ICs (M3D) built with carbon nanotube field-effect transistors and resistive memory. Our tool flow is based on commercial 2D tools and smart ways to extend them to conduct M3D design and simulation. We provide a post-route optimization flow, which exploits the full potential of the underlying M3D process design kit (PDK) for power, performance and area (PPA) optimization. We also conduct IR-drop and thermal analysis on M3D designs to improve the reliability. To enhance the testability of our M3D designs, we develop design-for-test (DFT) methodologies and integrate a low-overhead built-in self-test module into our design for testing inter-layer vias (ILVs) as well as logic circuitries in the individual tiers. Our benchmark design is RISC-V Rocketcore, which is an open source processor. Our experiments show 8.1% of power, 19.6% of wirelength and 55.7% of area savings with M3D designs at iso-performance compared to its 2D counterpart. In addition, our IR-drop and thermal analyses indicate acceptable power and thermal integrity in our M3D design. Gauthaman Murali, Pruek Vanna-Iampikul, Dae Hyun Kim 0004, Arjun Chaudhuri, Sanmitra Banerjee, Krishnendu Chakrabarty, Saibal Mukhopadhyay, Sung Kyu Lim |
ICCAD | 5 |
| 2020 | NP-Separate: A New VLSI Design Methodology for Area, Power, and Performance OptimizationabstractUse of standard cells in the very-large-scale integration (VLSI) design enables very short time to market even for complex microprocessors. Thus, most VLSI layouts are designed using standard cells. In this article, we propose a new design methodology, namely, NP-Separate, to reduce the power consumption and area and increase the performance of a VLSI layout more effectively than the standard-cell-based design methodology. NP-Separate uses N cells and P cells composed of NFETs and PFETs only, respectively, thereby providing a higher degree of flexibility than using standard cells. Our simulation results for several benchmark circuits show that NP-Separate reduces the layout area by 9%, power consumption by 10%, power-delay product by 18%, and energy-delay product by 26% on average while satisfying given timing constraints compared to standard-cell-based designs. Monzurul Islam Dewan, Dae Hyun Kim 0004 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Construction of All Rectilinear Steiner Minimum Trees on the Hanan Grid and Its Applications to VLSI DesignabstractA rectilinear Steiner minimum tree (RSMT) is a rectilinear Steiner tree connecting a given set of pins with the shortest wirelength. RSMT construction is one of the most frequently used algorithms in the physical design automation, including floorplanning, placement, routing, and interconnect estimation and optimization. Thus, efficient algorithms to construct RSMTs have been developed for many years in academia and industry. Unfortunately, RSMT construction is an NP-hard problem, so even a fast RSMT construction algorithm, such as GeoSteiner is too slow to use in physical design automation tools. FLUTE, a fast lookup-table-based RSMT construction algorithm, builds and uses a routing topology database to quickly construct RSMTs. In this paper, we present an algorithm to build a database (ARSMT DB) to construct all RSMTs on the Hanan grid for a given set of pins. ARSMT DB constructs all RSMTs in almost no time, so numerous applications could use it for various purposes. We apply the ARSMT DB to two applications, timing-driven RSMT construction and congestion-aware global routing, and show that the ARSMT DB can help reduce source-to-critical-sink lengths, source-to-critical-sink delays, and routing congestion significantly. Since the size of the original ARSMT DB is too large, we present techniques to reduce the database size. Sheng-En David Lin, Dae Hyun Kim 0004 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Inter-Tier Process-Variation-Aware Monolithic 3-D NoC Design Space ExplorationabstractMonolithic 3-D (M3D) technology enables high density integration, performance, and energy efficiency by sequentially stacking tiers on top of each other. M3D-based network-on-chip (NoC) architectures can exploit these benefits by adopting tier partitioning for intra-router stages. However, conventional fabrication methods are infeasible for M3D-enabled designs due to temperature-related issues. This has necessitated lower temperature and temperature-resilient techniques for M3D fabrication, leading to inferior performance of transistors in the top tier and interconnects in the bottom tier. The resulting inter-tier process variation leads to the performance degradation of M3D-enabled NoCs. In this article, we demonstrate that without considering inter-tier process variation, an M3D-enabled NoC architecture overestimates the energy-delay-product (EDP) on average by 50.8% for a set of SPLASH-2 and PARSEC benchmarks. As a countermeasure, we adopt a process variation-aware design approach. The proposed design and optimization method distributes the intra-router stages and inter-router links among the tiers to mitigate the adverse effects of process variation. Experimental results show that the NoC architecture under consideration improves the EDP by 27.4% on average across all benchmarks compared to the process-oblivious design. Shouvik Musavvir, Anwesha Chatterjee, Ryan Gary Kim, Dae Hyun Kim 0004, Partha Pratim Pande |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2019 | High-Throughput Multiplier Architectures Enabled by Intra-Unit Fast ForwardingabstractIn this paper, we propose a pipelined multiplier architecture that can resolve data dependencies. The proposed architecture generates partial results in the pipeline stages of the multiplier and forwards the partial results back to the pipeline stages through so-called fast-forwarding paths, thereby enabling an execution of dependent multiplications with a minimum delay penalty. We apply the architecture to a normal binary multiplier (NBBE-2) and two redundant binary multipliers (RBBE-4 and CRBBE-4) and compare the execution time, clock period, area, and power consumption of the multipliers. The simulation results show that the proposed architecture achieves up to 30% execution time reduction. Jihee Seo, Dae Hyun Kim 0004 |
ARITH | 2 |
| 2019 | Dependency-Resolving Intra-Unit Pipeline Architecture for High-Throughput MultipliersabstractIn this paper, we propose two dependency-resolving intra-unit pipeline architectures to design high-throughput multipliers. Simulation results show that the proposed multipliers achieve approximately 2.4× to 4.3× execution time reduction at a cost of 4.4% area and 3.7% power overheads for highly-dependent multiplications. Jihee Seo, Dae Hyun Kim 0004 |
DATE | 2 |
| 2019 | Construction of All Multilayer Monolithic Rectilinear Steiner Minimum Trees on the 3D Hanan Grid for Monolithic 3D IC RoutingabstractMonolithic three-dimensional~(3D) integration enables stacking multiple ultra-thin silicon tiers in a single package, thereby providing smaller footprint area, shorter wirelength, higher performance, and lower power consumption than conventional planar fabrication technologies. Physical design of monolithic 3D integrated circuits~(ICs) requires several design steps such as 3D placement, 3D clock-tree synthesis, 3D routing, and 3D optimization. Among the steps, 3D routing is very time-consuming due to numerous routing blockages. Thus, 3D routing is typically performed in two sub-steps, monolithic inter-layer via~(MIV) insertion and tier-by-tier routing. In this paper, we propose an algorithm to build a routing topology database that can be used to construct all multilayer monolithic rectilinear Steiner minimum trees on the 3D Hanan grid. The database will help 3D routers reduce the runtime of the MIV insertion step and improve the quality of the 3D routing. Sheng-En David Lin, Dae Hyun Kim 0004 |
ISPD | 2 |
| 2018 | Construction of All Rectilinear Steiner Minimum Trees on the Hanan GridabstractGiven a set of pins, a Rectilinear Steiner Minimum Tree (RSMT) connects the pins using only rectilinear edges with the minimum wirelength. RSMT construction is heavily used at various design steps such as floorplanning, placement, routing, and interconnect estimation and optimization, so fast algorithms to construct RSMTs have been developed for many years. However, RSMT construction is an NP-hard problem, so even a fast RSMT construction algorithm such as GeoSteiner [7] is too slow to use in electronic design automation (EDA) tools. FLUTE, a lookup-table-based RSMT construction algorithm, builds and uses a routing topology database to quickly construct RSMTs[5]. However, FLUTE outputs only one RSMT for a given set of pin locations. In this paper, we develop an algorithm to build a database of all RSMTs on the Hanan grid for up to nine pins. The database will be able to help minimize routing congestion and maximize the routability in the design of modern very-large-scale integration layouts. Sheng-En David Lin, Dae Hyun Kim 0004 |
ISPD | 2 |
| 2018 | Design Space Exploration of 3D Network-on-Chip: A Sensitivity-based Optimization ApproachabstractHigh-performance and energy-efficient Network-on-Chip (NoC) architecture is one of the crucial components of the manycore processing platforms. A very promising NoC architecture recently proposed in the literature is the three-dimensional small-world NoC (3D SWNoC). Due to short vertical links in 3D integration and the robustness of small-world networks, the 3D SWNoC architecture outperforms its other 3D counterparts. However, the performance of 3D SWNoC is highly dependent on the placement of the links and associated routers. In this article, we propose a sensitivity-based link placement algorithm (SEN) to optimize the performance of 3D SWNoC. The sensitivity of a link in a NoC measures the importance of the link. The SEN algorithm optimizes the performance of 3D SWNoC by calculating the sensitivities of all the links in the NoC and removing the least important link repeatedly. We compare the performance of SEN algorithm with simulated annealing- (SA) and recently proposed machine-learning-based (ML) optimization algorithm. The optimized 3D SWNoC obtained by the proposed SEN algorithm achieves, on average, 11.5% and 13.6% lower latency and 18.4% and 21.7% lower energy-delay product than those optimized by the SA and ML algorithms respectively. In addition, the SEN algorithm is 26 to 33 times faster than the SA algorithm for the optimization of 64-, 128-, and 256-core 3D SWNoC designs. The performance gain provided by the SEN-, SA-, and ML-based methods also depend on the characteristics of the benchmarks under consideration. If the traffic pattern generated by a benchmark does not have enough variation, then the ML-based method does not have adequate opportunity to optimize the network. However, we find that ML-based methodology has faster convergence time than SEN and SA for bigger systems. The ML-based optimization algorithm is almost 4 and 97 times faster than the SEN- and SA-based algorithm for a system with 256 cores. Sourav Das 0002, Dae Hyun Kim 0004, Janardhan Rao Doppa, Partha Pratim Pande |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2018 | Analysis of Performance Benefits of Multitier Gate-Level Monolithic 3-D Integrated CircuitsabstractVertical interconnects used in monolithic 3-D integrated circuits (3-D ICs), so-called monolithic interlayer vias (MIVs), are as small as local vias. Thus, redesigning an existing 2-D IC layout in a monolithic 3-D IC generally results in shorter wire length than the 2-D IC layout. In addition, MIVs have almost negligible resistance and capacitance, so their impact on signal delay is very small. Thus, redesigning a 2-D IC layout in a monolithic 3-D IC is expected to improve its performance significantly. Some researchers designed several monolithic 3-D IC layouts and showed their timing benefits in the literature. In this paper, we present analytical models for performance (timing) benefits of multitier gate-level monolithic 3-D ICs. The analytical models we develop in this paper can be used to quickly estimate the performance benefits multitier gate-level monolithic 3-D integration provides without physically redesigning 2-D IC layouts in 3-D. Inki Hong, Dae Hyun Kim 0004 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | Detailed-Placement-Enabled Dynamic Power Optimization of Multitier Gate-Level Monolithic 3-D ICsabstractMonolithic 3-D integration is expected to provide significantly higher degree of device density than through-silicon-via-based 3-D integration due mainly to its nano-scale intertier connections. By stacking more than two device layers (multitier) within a 3-D chip, further wirelength reduction could be achieved, which can lead to additional performance and power benefits. In this paper, we propose a detailed placement algorithm called nonuniform-scaling-based placement to optimize the dynamic power consumption of multitier gate-level monolithic 3-D ICs. We also introduce delay- and length-based timing constraints to prevent potential degradation of the performance metric during placement. Under the same timing constraints, our algorithm reduces dynamic power consumption more effectively than the uniform-scaling-based placement algorithm by 2% to 14%. Sheng-En David Lin, Dae Hyun Kim 0004 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2015 | Small-World Network Enabled Energy Efficient and Robust 3D NoC ArchitecturesabstractThree dimensional (3D) Network-on-Chip (NoC) architectures enable design of low power and high performance communication fabrics for multicore chips. In spite of achievable performance benefits, 3D NoCs are still bottlenecked by the planar interconnects. To exploit the benefits introduced by the vertical dimension, it is imperative to explore novel 3D NoC architectures. In this paper, we propose design of a small-world (SW) network based 3D NoCs. We demonstrate that the proposed 3D SW NoC outperforms its conventional 3D mesh-based counterparts. On average, it provides ~25% reduction in the energy delay product (EDP) compared to 3D MESH without introducing any additional link overhead in presence of conventional SPLASH-2 and PARSEC benchmarks. The proposed 3D SW NoC is more robust in presence of TSV failures and performs better than fault-free 3D MESH even in the presence of 25% TSVs failure. Sourav Das 0002, Dae Hyun Kim 0004, Partha Pratim Pande |
ACM Great Lakes Symposium on VLSI | 3 |
| 2015 | Optimizing 3D NoC Design for Energy Efficiency: A Machine Learning ApproachabstractThree-dimensional (3D) Network-on-Chip (NoC) is an emerging technology that has the potential to achieve high performance with low power consumption for multicore chips. However, to fully realize their potential, we need to consider novel 3D NoC architectures. In this paper, inspired by the inherent advantages of small-world (SW) 2D NoCs, we explore the design space of SW network-based 3D NoC architectures. We leverage machine learning to intelligently explore the design space to optimize the placement of both planar and vertical communication links for energy efficiency. We demonstrate that the optimized 3D SW NoC designs perform significantly better than their 3D MESH counterparts. On an average, the 3D SW NoC shows 35% energy-delay-product (EDP) improvement over 3D MESH for the nine PARSEC and SPLASH2 benchmarks considered in this work. The highest performance improvement of 43% was achieved for RADIX. Interestingly, even after reducing the number of vertical links by 50%, the optimized 3D SW NoC performs 25% better than the fully connected 3D MESH, which is a strong indication of the effectiveness of our optimization methodology. Sourav Das 0002, Janardhan Rao Doppa, Dae Hyun Kim 0004, Partha Pratim Pande, Krishnendu Chakrabarty |
ICCAD | 3 |
| 2015 | Design and Analysis of 3D-MAPS (3D Massively Parallel Processor with Stacked Memory)abstractThis paper describes the architecture, design, analysis, and simulation and measurement results of the 3D-MAPS (3D massively parallel processor with stacked memory) chip built with a 1.5 V, 130 nm process technology and a two-tier 3D stacking technology using 1.2$\micro\hbox{m}$-diameter, 6$\micro \hbox{m}$-height through-silicon vias (TSVs) and$3.4\nbsp\micro\hbox{m}$-diameter face-to-face bond pads. 3D-MAPS consists of a core tier containing 64 cores and a memory tier containing 64 memory blocks. Each core communicates with its dedicated 4KB SRAM block using face-to-face bond pads, which provide negligible data transfer delay between the core and the memory tiers. The maximum operating frequency is 277 MHz and the maximum memory bandwidth is 70.9 GB/s at 277 MHz. The peak measured memory bandwidth usage is 63.8 GB/s and the peak measured power is approximately 4 W based on eight parallel benchmarks. Dae Hyun Kim 0004, Krit Athikulwongse, Michael B. Healy, Mohammad M. Hossain, Moongon Jung, Ilya Khorosh, Gokul Kumar, Young-Joon Lee, Dean L. Lewis, Tzu-Wei Lin, Chang Liu 0034, Shreepad Panth, Mohit Pathak, Minzhen Ren, Guanhao Shen, Taigon Song, Dong Hyuk Woo, Xin Zhao 0001, Joungho Kim, Ho Choi, Gabriel H. Loh, Hsien-Hsin S. Lee, Sung Kyu Lim |
IEEE Trans. Computers | 1 |
| 2014 | TSV-Aware Interconnect Distribution Models for Prediction of Delay and Power Consumption of 3-D Stacked ICsabstract3-D integrated circuits (3-D ICs) are expected to have shorter wirelength, better performance, and less power consumption than 2-D ICs. These benefits come from die stacking and use of through-silicon vias (TSVs) fabricated for interconnections across dies. However, the use of TSVs has several negative impacts such as area and capacitance overhead. To predict the quality of 3-D ICs more accurately, TSV-aware 3-D wirelength distribution models considering the negative impacts were developed. In this paper, we apply an optimal buffer insertion algorithm to the TSV-aware 3-D wirelength distribution models and present various prediction results on wirelength, delay, and power consumption of 3-D ICs. We also apply the framework to 2-D and 3-D ICs built with various combinations of process and TSV technologies and predict the quality of today and future 3-D ICs. Dae Hyun Kim 0004, Saibal Mukhopadhyay, Sung Kyu Lim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2014 | Backend Dielectric Reliability Full Chip SimulatorabstractBackend dielectric breakdown degrades the reliability of circuits. A methodology to estimate chip lifetime because of backend dielectric breakdown is presented. It incorporates failures because of parallel tracks, the width effect, and field enhancement due to line ends. It also includes the operating temperature and activity. Muhammad Bashir, Chang-Chih Chen, Linda S. Milor, Dae Hyun Kim 0004, Sung Kyu Lim |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2013 | Block-level designs of die-to-wafer bonded 3D ICs and their design quality tradeoffsabstractIn 3D ICs, block-level designs provide various advantages over designs done at other granularity such as gate-level because they promote the reuse of IP blocks. In this paper, we study block-level 3D-IC designs, where the footprint of the dies in the stack are different. This happens in case of die-to-wafer bonding, which is more popular choice for near-term low-cost 3D designs. We study design quality tradeoffs among three different ways to place through-silicon vias (TSVs): TSV-farm, TSV-distributed, and TSV-whitespace. In our holistic approach, we use wirelength, power, performance, temperature, and mechanical stress metrics to conduct comprehensive comparative studies on the three design styles. In addition, we provide analysis on the impact of TSV size and pitch on the design quality of these three styles. Krit Athikulwongse, Dae Hyun Kim 0004, Moongon Jung, Sung Kyu Lim |
ASP-DAC | 2 |
| 2013 | Study of Through-Silicon-Via Impact on the 3-D Stacked IC LayoutabstractThe technology of through-silicon vias (TSVs) enables fine-grained integration of multiple dies into a single 3-D stack. TSVs occupy significant silicon area due to their sheer size, which has a great effect on the quality of 3-D integrated chips (ICs). Whereas well-managed TSVs alleviate routing congestion and reduce wirelength, excessive or ill-managed TSVs increase the die area and wirelength. In this paper, we investigate the impact of the TSV on the quality of 3-D IC layouts. Two design schemes, namely TSV co-placement (irregular TSV placement) and TSV site (regular TSV placement), and accompanying algorithms to find and optimize locations of gates and TSVs are proposed for the design of 3-D ICs. Two TSV assignment algorithms are also proposed to enable the regular TSV placement. Simulation results show that the wirelength of 3-D ICs is shorter than that of 2-D ICs by up to 25%. Dae Hyun Kim 0004, Krit Athikulwongse, Sung Kyu Lim |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2012 | Block-level 3D IC design with through-silicon-via planningabstractSince re-designing and re-optimizing existing logic, memory, and IP blocks in a 3D fashion significantly increases design cost, near-term three-dimensional integrated circuit (3D IC) design will focus on reusing existing 2D blocks. One way to reuse 2D blocks in the 3D IC design is to first perform 3D floorplanning, insert signal through-silicon vias (TSVs) for 3D inter-block connections, and then route the blocks. In this paper, we propose algorithms (finding signal TSV locations, assigning TSVs to whitespace blocks, and manipulating whitespace blocks) for post-floorplanning signal TSV planning in the block-level 3D IC design. Experimental results show that our signal TSV planner outperforms the state-of-the-art TSV-aware 3D floorplanner by 7% to 38% with respect to wirelength. In addition, our multiple TSV insertion algorithm outperforms a single TSV insertion algorithm by 27% to 37%. Dae Hyun Kim 0004, Rasit Onur Topaloglu, Sung Kyu Lim |
ASP-DAC | 1 |
| 2010 | Design method and test structure to characterize and repair TSV defect induced signal degradation in 3D systemabstractIn this paper we present a test structure and design methodology for testing, characterization, and self-repair of TSVs in 3D ICs. The proposed structure can detect the signal degradation through TSVs due to resistive shorts and variations in TSV. For TSVs with moderate signal degradations, the proposed structure reconfigures itself as signal recovery circuit to improve signal fidelity. The paper presents the design of the test/recovery structure, the test methodologies, and demonstrates its effectiveness through stand alone simulations as well as in a full-chip physical design of a 3D IC. Minki Cho, Chang Liu 0034, Dae Hyun Kim 0004, Sung Kyu Lim, Saibal Mukhopadhyay |
ICCAD | 3 |
| 2009 | A study of Through-Silicon-Via impact on the 3D stacked IC layoutabstractThrough-Silicon-Via (TSV) is the enabling technology for the fine-grained 3D integration of multiple dies into a single stack. These TSVs occupy non-negligible silicon area because of their sheer size. This significant silicon area occupied by the TSVs and the interconnections made to the TSVs greatly affect area, power, performance, and reliability of 3D IC layouts. Well-managed TSVs alleviate congestion, reduce wirelength, and improve performance, whereas excessive TSVs not only increase the die area, but also have negative impact on many design objectives. In this paper, we study the impact of TSV on various aspects of 3D layouts. We use GDSII layouts of 2D and 3D designs, and thoroughly compare the pros and cons of TSV usage. We propose a new force-directed 3D gate-level placement that efficiently handles TSVs. In addition, we present an algorithm that assigns TSVs to nets to complete routing that involves TSVs. This algorithm, together with our 3D placer, is integrated into a commercial P&R tool to generate fully validated GDSII layouts. Our experiments based on synthesized benchmarks indicate that our algorithms help generate GDSII layouts of 3D designs that are optimized in terms of area, wirelength, and metal layer count. Dae Hyun Kim 0004, Krit Athikulwongse, Sung Kyu Lim |
ICCAD | 1 |
| 2008 | Bus-aware microarchitectural floorplanningabstractIn this paper we present the first bus-aware microarchitectural floorplanning. Our goal is to study the impact of bus routability on other important floorplanning objectives including area, performance, power, and thermal. We developed a fast performance-aware bus routing algorithm, which is integrated into the floorplanning engine to ensure routability while optimizing other conflicting objectives. Our related experiments performed on high performance processors show that we obtain 100% routability at the cost of minimal increase on area, performance, and power objectives under thermal constraint. Dae Hyun Kim 0004, Sung Kyu Lim |
ASP-DAC | 1 |
| 2008 | Global bus route optimization with application to microarchitectural design explorationabstractCircuit and processor designs will continue to increase in complexity for the foreseeable future. With these increasing sizes comes the use of wide buses to move large amounts of data from one place to another. Bus routing has therefore become increasingly important. In this paper, we present a new bus routing algorithm that globally optimizes both the floorplan and the bus routes themselves. Our algorithm is based on creating a range of feasible bus positions and then using Linear Programming to optimally solve for bus locations. We present this algorithm for use in microarchitectures and explore several different optimization objectives, including performance, floorplan area, and power consumption. Our results demonstrate that this algorithm is effective for efficiently generating feasible routes for complex modern designs and provides better results than previous approaches. Dae Hyun Kim 0004, Sung Kyu Lim |
ICCD | 1 |