Kyu-Myung Choi

dblp:94/1330 · DBLP profile ↗
← Back
23ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0001-8153-8344ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 23 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 8 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 ML-driven Design Technology Co-Optimization Framework for Advanced Technology Nodes
abstract
The goal of design and technology co-optimization (DTCO) is to find a combination of parameter options (i.e., parameter setting values) of target process technology that enables to produce a target design implementation of optimal PPA (performance, power, area). Since the number of parameters sharply increases as the technology scales, recently, lots of attention has been paid to automating this DTCO process in both semiconductor foundry and academic research community. This paper addresses the problem of a full DTCO automation that deals with analyzing the numerous parameter options at advanced technology nodes. Precisely, we develop a machine learning (ML) based DTCO automation framework, which supports three key features: (1) an effective analysis on the changes of DTCO parameter options within an acceptable runtime; (2) a full exploration of chip/block-level PPA metrics through automatic standard cell (SC) library generation, for which we develop a new technique that enables to accelerate the iterative physical design process; (3) supporting both of Complementary FET (CFET) based SCs and multi-row-height SCs to account for future generation technology. Through experiments with benchmark circuits, it is shown that our DTCO automation framework is able to accurately predict the direction and magnitude of PPA changes of target designs with 5x sampling efficiency. In addition, it is shown that our SC layout generator supporting CFET and multi-rowheight SCs provides a timely DTCO process relevant to ongoing technology advancements.
Hyunbae Seo, Handong Cho, Sehyeon Chung, Kyu-Myung Choi, Taewhan Kim 0001
ASP-DAC4
2024 Standard Cell Layout Generator Amenable to Design Technology Co-Optimization in Advanced Process Nodes
abstract
To generate standard cell (SC) layouts of competitive quality, pin accessibility and in-cell routing congestion should be thoroughly taken into account. In this work, we develop a new tool to address this issue. Precisely, we (1) develop a technology compilation module that can convert diverse cell architectures and design rules into grid based design parameters and layer configuration, (2) generate optimal FET placement using metrics that can accurately and efficiently predict intra-cell pin accessibility and in-cell routing congestion, and (3) introduce the concept of ghost-via and ghost-metal, and formulate in-cell routing using satisfiability modulo theory for pin separation and extension. Experimental results show that our system is able to synthesize SC layouts with a routing completion rate of 95~98 %, which is far better than the previous SC layout generator, and produce layouts comparable to the ARM's hand-crafted layouts. In addition, the design implementations produced by using our 2-layer ID SC library exhibit on average 76.6% fewer design rule violations (DRVs) with similar or better quality of timing and area, while in comparison with that produced by using the library of hand-crafted ARM SCs, the implementations produced by using our L-layer 2D SC library exhibit on average 11.7% smaller area with comparable timing and DRV count.
Handong Cho, Hyunbae Seo, Sehyeon Chung, Kyu-Myung Choi, Taewhan Kim 0001
DATE4
2024 BOXGB: Design Parameter Optimization with Systematic Integration of Bayesian Optimization and XGBoost
abstract
Finding design flow parameters that ensure a high quality of final chip is a very important task, but requires an excessive amount of effort and time. In this work, we automate this task by proposing a machine learning (ML)-based design space optimization (DSO) framework. Rather than simply applying one ML model exclusively or multiple ones in a naive manner, we develop a comprehensive chain of ML engines which is able to explore the design parameter space more economically but effectively to make a fast convergence on finding the best parameter set. Specifically, we solve the DSO problem in three steps: (1) random sampling of parameter sets and then performing design evaluation to produce an initial ML training dataset; (2) iteratively, downsizing parameter dimension through Principal Component Analysis (PCA) followed by sampling through an exploration-centric mechanism which is internally driven by Bayesian Optimization (BO) and then evaluating the sample; (3) iteratively, sampling through an exploitation-centric mechanism driven by XGBoost regression and then checking anomaly by using XGBoost classification followed by evaluating the sample if it's not anomaly. From our experiments with benchmark designs, it is shown that our approach is able to find design parameter sets which are far better than that found by the prior state-of-the-art ML-based approaches, even with fewer number of design evaluations (i.e., EDA tool runs). In addition, in comparison with the designs produced by using the default parameter setting, our DSO framework is able to improve the design PPA metrics by$5\sim 30{\%}$I.
Chanhee Jeon, Doyeon Won, Jaewan Yang, Kyu-Myung Choi, Taewhan Kim 0001
DATE4
2023 DTOC: integrating Deep-learning driven Timing Optimization into the state-of-the-art Commercial EDA tool
abstract
Recently, deep-learning (DL) models have paid a considerable attention to timing prediction in the placement and routing (P&R) flow. As yet, the DL-based prior works are confined to timing prediction at the time-consuming global routing stage, and very few have addressed the timing prediction problem at the placement, i.e., at the pre-route stage. This is because it is not easy to “accurately” predict various timing parameters at the pre-route stage. Moreover, no work has addressed a seamless link of timing prediction at the pre-route stage to the final timing optimization through making use of commercial P&R tools. In this work, we propose a framework called DTOC, to be used at the pre-route stage for this end. Precisely, the framework is composed of two models: (1) a DL-driven arc delay and arc output slew prediction model, performing in two levels: (level-1) predicting net resistance (R), net capacitance (C), and arc length (Len), followed by (level-2) predicting arc delay and arc output slew from the R/C/Len prediction obtained in (level-1); (2) a timing optimization model, which uses the inference outcomes in our DL-driven prediction model to enable the commercial P&R tools to calculate the full path delays, setting update timing margins on paths, so that the P&R tools should use more accurate margins on timing optimization. Experimental results show that, by using our DTOC framework during timing optimization in P&R, we improve the pre-route prediction accuracy on arc delay and arc output slew by 20~26% on average, and improve the WNS, TNS, and the number of timing violation paths by 50~63 % on average.
Kyungjoon Chang, Jaehoon Ahn, Heechun Park, Kyu-Myung Choi, Taewhan Kim 0001
DATE4
2020 Synthesis of Hardware Performance Monitoring and Prediction Flow Adapting to Near-Threshold Computing and Advanced Process Nodes
abstract
An elaborate hardware performance monitor (HPM) has become increasingly important for handling huge performance variation of near-threshold computing and recent process technologies. In this paper, we propose a new approach to the problem of predicting critical path delays (CPDs) using HPM. Precisely, for a target circuit or system, we formulate the problem of finding an efficient combination of ring oscillators (ROs) for accurate prediction of CPDs on the circuit as a mixed integer second-order cone programming and propose a method of minimizing the total number of ROs for a given pessimism level of prediction. Then, we propose a prediction flow of CPDs through statistical estimation of process parameters from measurements of the customized HPM and machine learning based delay mapping from the estimation. For a set of benchmark circuits tested using 28nm PDK and 0.6V operation, it is shown that our approach is very effective, reducing the pessimism of CPDs and minimum supply voltages by 6.7~52.9% and 20.6~50.8% over those of conventional approaches, respectively.
Jeongwoo Heo, Kwangok Jeong, Taewhan Kim 0001, Kyu-Myung Choi
ASP-DAC4
2020 SRAM on-chip monitoring methodology for high yield and energy efficient memory operation at near threshold voltage
Taehwan Kim 0007, Kwangok Jeong, Jungyun Choi, Taewhan Kim 0001, Kyu-Myung Choi
Integr.5
2019 Design Rule Evaluation Framework Using Automatic Cell Layout Generator for Design Technology Co-Optimization
abstract
This paper proposes a complete and full automation framework of evaluating design rules (DRs) to facilitate the process of design technology co-optimization (DTCO), which is highly demanded in 14-nm and beyond technologies. Our proposed framework explores the changes of DRs and evaluates the impacts on the number and types of DR violations as well as the resulting cell/chip layout area. Precisely, the core engine of our DR evaluation framework for DTCO, the automatic cell layout generator, consists of key enabling techniques for standard cell layout optimization. They are integrated coherently to seamlessly support the advanced process technologies using FinFET transistors, complex DRs, and double patterning (DP) lithography. Also, the tight integration of our automatic cell layout generation into the DR evaluation framework with diverse analysis features enables the DTCO process to be much faster and more efficient. We provide a set of experimental data not only to show how much our proposed enabling techniques are effective in optimizing layouts but also to show how effectively our framework explores and analyzes the DTCO parameters (e.g., ground DRs and DP DRs).
Kyeongrok Jo, Seyong Ahn, Jungho Do, Taejoong Song, Taewhan Kim 0001, Kyu-Myung Choi
IEEE Trans. Very Large Scale Integr. Syst.6
2018 Cohesive techniques for cell layout optimization supporting 2D metal-1 routing completion
abstract
This work addresses the problem of automatically synthesizing compact standard cell layouts with 2D metal-1 routing under design rule constraints. Precisely, we propose a set of new highly impacting techniques dedicated solely to the generation of cell layouts with 2D metal-1 routing completion. Those are (1) netlist decomposition (2) transistor chaining combined with transistor folding, (3) gate poly ordering combined with fast routing congestion estimation, and (4) 2D single-layer routing with minimal resource. It is shown from experiments that our proposed layout generator is able to produce layouts of quality comparable to the expert's manual ones, but spending just one hour for 56 representative cells generation.
Kyeongrok Jo, Seyong Ahn, Taewhan Kim 0001, Kyu-Myung Choi
ASP-DAC4
2010 An industrial perspective of 3D IC integration technology: from the viewpoint of design technology
abstract
3D IC integration is very important to overcome the technology scaling barriers and to satisfy mobile devices' demand. In this paper, we describe the challenges we are facing in developing 3D IC design methodology, especially in the case of TSV-SiP (Logic-Memory die stacking). Also, appropriate development approaches are proposed. The EDA tools for TSV-SiP, which are initially provided by extending current conventional tools, will be gradually enhanced to better support 3D IC designs.
Kyu-Myung Choi
ASP-DAC1
2010 Power gating: Circuits, design methodologies, and best practice for standard-cell VLSI designs
abstract
Power Gating has become one of the most widely used circuit design techniques for reducing leakage current. Its concept is very simple, but its application to standard-cell VLSI designs involves many careful considerations. The great complexity of designing a power-gated circuit originates from the side effects of inserting current switches, which have to be resolved by a combination of extra circuitry and customized tools and methodologies. In this tutorial we survey these design considerations and look at the best practice within industry and academia. Topics include output isolation and data retention, current switch design and sizing, and physical design issues such as power networks, increases in area and wirelength, and power grid analysis. Designers can benefit from this tutorial by obtaining a better understanding of implications of power gating during an early stage of VLSI designs. We also review the ways in which power gating has been improved. These include reducing the sizes of switches, cutting transition delays, applying power gating to smaller blocks of circuitry, and reducing the energy dissipated in mode transitions. Power Gating has also been combined with other circuit techniques, and these hybrids are also reviewed. Important open problems are identified as a stimulus to research.
Youngsoo Shin, Jun Seomun, Kyu-Myung Choi, Takayasu Sakurai
ACM Trans. Design Autom. Electr. Syst.3
2010 Supply Switching With Ground Collapse for Low-Leakage Register Files in 65-nm CMOS
abstract
Power-gating has been widely used to reduce subthreshold leakage current. However, the extent of leakage saving through power-gating diminishes with technology scaling due to gate leakage of data-retention circuit elements. Furthermore, power-gating involves substantial increase of area and wirelength. A circuit technique called supply switching with ground collapse (SSGC) has recently been proposed to overcome the limitation of power-gating. The circuit technique is successfully applied to the register file of ARM9 microprocessor in a 1.2 V, 65-nm CMOS process, and the measured result is reported for the first time. The leakage current is reduced by a factor of 960 on average of 83 dies at 25°C , and by a factor of 150 at 85°C. Compared to a register file implemented in conventional power-gating, leakage current is cut by a factor of 2.2, demonstrating that SSGC can be a substitute for power-gating in nanometer CMOS.
Hyung-Ock Kim, Bong Hyun Lee, Jong-Tae Kim, Jung Yun Choi, Kyu-Myung Choi, Youngsoo Shin
IEEE Trans. Very Large Scale Integr. Syst.5
2009 45nm Low-power Embedded Pseudo-SRAM with ECC-based Auto-adjusted Self-refresh Scheme
abstract
In this paper, a low-power embedded pseudo-SRAM adopting novel auto-adjusted self-refresh control scheme has been designed. The proposed self-refresh control scheme automatically extends the self-refresh period by monitoring the number of failed cells using error correction code (ECC). The scheme can provide a substantial reduction of data-retention power consumption by choosing an optimal self-refresh period regardless of process, voltage, and temperature (PVT) variations. A 4-Mb embedded pseudo-SRAM designed in a 45-nm embedded DRAM technology providing 1.1-V 166-MHz random cycle operation achieves 57-uW data retention power consumption at room temperature.
Suk-Soo Pyo, Cheol-Ha Lee, Gyun-Hong Kim, Kyu-Myung Choi, Young-Hyun Jun, Bai-Sun Kong
ISCAS4
2008 An industrial perspective of power-aware reliable SoC design
abstract
Reliable SoC design is becoming one of important real design problems since the fast pace of semiconductor scaling and the introduction of new device structures and materials incur more reliability problems than can be solved in the given time frame (2 years/technology node). Reliable design mostly requires resource overhead (additional power consumption, silicon area, and execution time) to recover from errors. Minimizing the overhead in the reliable SoC design will give SoC industries a competitive edge. Especially, in the case of mobile SoC, mastering the overhead of power consumption is absolutely imperative. In this paper, we investigate reliable SoC design in terms of reducing the overhead of power consumption. First, we review the current practice of reliable SoC design and assess its impact on power consumption. Then, we present our perspective on new design methodology towards power-aware reliable SoC design.
Soo-Kwan Eo, Sungjoo Yoo, Kyu-Myung Choi
ASP-DAC3
2008 A practical approach of memory access parallelization to exploit multiple off-chip DDR memories
abstract
3D stacked memory enables more off-chip DDR memories. Redesigning existing IPs to exploit the increased memory parallelism will be prohibitively costly. In our work, we propose a practical approach to exploit the increased bandwidth and reduced latency of multiple off-chip DDR memories while reusing existing IPs without modification. The proposed approach is based on two new concepts: transaction id renaming and distributed soft arbitration. We present two on-chip network components, request parallelizer and read data serializer, to realize the concepts. Experiments with synthetic test cases and an industrial strength DTV SoC design show that the proposed approach gives significant improvements in total execution cycle (21.6%) and average memory access latency (31.6%) in the DTV case with a small area overhead (30.1% in the on-chip network, and less than 1.4% in the entire chip).
Woo-Cheol Kwon, Sungjoo Yoo, Sung-Min Hong, Byeong Min, Kyu-Myung Choi, Soo-Kwan Eo
DAC5
2008 Dynamic Voltage Scaling of Supply and Body Bias Exploiting Software Runtime Distribution
abstract
This paper presents a method of dynamic voltage scaling (DVS) that tackles both switching and leakage power with combined Vdd/Vbsscaling and gives minimum average energy consumption exploiting the runtime distribution of software execution. We present a mathematical formulation of the DVS problem and an efficient numerical solution. Experimental results show that the presented method shows up to 44% further reduction in energy consumption compared with existing methods. Especially, when the leakage power consumption is significant, i.e. when temperature is high, the presented method is proven to be the most effective.
Sungpack Hong, Sungjoo Yoo, Byeong Bin, Kyu-Myung Choi, Soo-Kwan Eo, Taehwan Kim 0007
DATE4
2008 An Open-Loop Flow Control Scheme Based on the Accurate Global Information of On-Chip Communication
abstract
3D stacked memory is being adopted as a promising solution to offer high bandwidth and low latency in memory access. Compared with the on-chip network design with conventional off chip memory, it gives a new problem of minimizing communication conflicts since multiple concurrent high bandwidth data transfers will flow through the on-chip network. In order to tackle this problem, we propose applying an open-loop flow control scheme based on the accurate global information (destination and status) of on-chip communication. The proposed open-loop flow control scheme exploits the information and selectively buffers and arbitrates data transfers to remove conflicts at destinations in a preventive manner. As an implementation of the presented scheme, we present on-chip buffers called Buf3D's that share the global information with each other to perform the selective buffering and arbitration of data transfers. Experiments with synthetic test cases and an industrial strength DTV design show that the proposed method improves aggregate memory bandwidth significantly (average 19.0 %~25.8 % in the synthetic cases and up to 18.4 % in the DTV case) with a small area overhead (15.2 % in the DTV case) of on-chip network.
Woo-Cheol Kwon, Sung-Min Hong, Sungjoo Yoo, Byeong Min, Kyu-Myung Choi, Soo-Kwan Eo
DATE5
2006 PowerViP: Soc power estimation framework at transaction level
abstract
In this work, we propose a SoC power estimation framework built on our system-level simulation environment. Our framework provides designers with the system-level power profile in a cycle-accurate manner. We target the framework to run fast and accurately, which is enabled by adopting different modeling techniques depending on the power characteristics of various IP blocks. The framework can be applied to any target SoC design
Ikhwan Lee, Sungjoo Yoo, Eui-Young Chung, Kyu-Myung Choi, Jeong-Taek Kong, Soo-Kwan Eo
ASP-DAC6
2006 Worst case execution time analysis for synthesized hardware
abstract
We propose a hardware performance estimation flow for fast design space exploration, based on worst-case execution time analysis algorithms for software analysis. Test cases on some real-world applications show that our flow provides a fight upper bound of the execution time, and many useful hints to the designer.
Jun-hee Yoo, Xingguang Feng, Kiyoung Choi, Eui-Young Chung, Kyu-Myung Choi
ASP-DAC5
2006 A systematic IP and bus subsystem modeling for platform-based system design
abstract
The topic on platform-based system modeling has received a great deal of attention today. One of the important tasks that significantly affect the effectiveness and efficiency of the system modeling is the modeling of IP components and communication between IPs. To be effective, it is generally accepted that the system modeling should be performed in two steps; In the first step, a fast but some inaccurate system modeling is considered to facilitate the simultaneous development of software and hardware. The second step then refines the models of the software and hardware blocks (i.e., IPs) to increase the simulation accuracy for the system performance analysis. Here, one critical factor required for a successful system modeling is a systematic modeling of the IP blocks and bus subsystem connecting the IPs. In this respect, this work addresses the problem of systematic modeling of the IPs and bus subsystem in different levels of refinements. In the experiments, we found that by applying our proposed IP and bus modeling methods to the MPEG-4 application, we are able to achieve 4/spl times/ performance improvement and at the same time, reduce the software development time by 35%, compared to that by conventional modeling methods.
Junhyung Um, Woo-Cheol Kwon, Sungpack Hong, Young-Taek Kim, Kyu-Myung Choi, Jeong-Taek Kong, Soo-Kwan Eo, Taewhan Kim 0001
DATE5
2006 Runtime distribution-aware dynamic voltage scaling
abstract
We propose a new intra-task dynamic voltage scaling (DVS) method to capture an important fact of ‘software runtime distribution ’ and integrate it into DVS effectively. Specifically, the proposed method finds performance levels, for a given software runtime distribution, i.e. statistical profiling of execution cycles (neither the execution cycle of worst-case execution path nor the worst-case execution cycles of basic blocks), which leads to a minimal energy consumption while satisfying the given deadline constraints. Experimental results report that the proposed method gives 19.2%~33.3 % further energy reduction compared with the best-known methods for two industrial multimedia software programs, H.264 decoder and MPEG4 decoder. 1.
Sungpack Hong, Sungjoo Yoo, HoonSang Jin, Kyu-Myung Choi, Jeong-Taek Kong, Soo-Kwan Eo
ICCAD4
2005 Fast and Accurate Transaction Level Modeling of an Extended AMBA2.0 Bus Architecture
abstract
A transaction level modeling (TLM) approach is used to meet the simulation speed as well as cycle accuracy for large scale SoC performance analysis. We implemented the transaction-level model of a proprietary bus called AHB+ which supports an extended AMBA2.0 protocol. The AHB+ transaction-level model is shown to be 353 times faster than the pin-accurate RTL model, while maintaining 97% accuracy on average. We also present the TLM development procedure of a bus architecture.
Young-Taek Kim, Youngduk Kim, Chulho Shin, Eui-Young Chung, Kyu-Myung Choi, Jeong-Taek Kong, Soo-Kwan Eo
DATE6
2004 Fast Exploration of Parameterized Bus Architecture for Communication-Centric SoC Design
abstract
For successful SoC design, efficient and scalable communication architecture is crucial. Some bus interconnects now provide configurable structures to meet this requirement of an SoC design. Furthermore, bus IP vendors provide software tools that automatically generate RTL codes of a bus once its designer configures it. Configurability, however, imposes more challenges upon designers because complexity involved in optimization increases exponentially as the number of parameters grows. In this paper, we present a novel approach with which effort requirement can be dramatically reduced. An automated optimization tool we developed is used and it exploits a genetic algorithm for fast design exploration. This paper shows that the time for the optimizing task can be reduced by more than 90% when the tool is used and, more significantly the task can be done without an expert's hand while ending up with a better solution.
Chulho Shin, Young-Taek Kim, Eui-Young Chung, Kyu-Myung Choi, Jeong-Taek Kong, Soo-Kwan Eo
DATE4
2003 An MTCMOS design methodology and its application to mobile computing
abstract
The Multi-Threshold CMOS (MTCMOS) technology provides a solution to the high performance and low power design requirements of modern designs. While the low Vth transistors are used to implement the desired function, the high Vth transistors are used to cut off the leakage current. In this paper, we (i) examine the effectiveness of the MTCMOS technology for the Samsung's 0.18?m process, (ii) propose a new special flip-flop which keeps a valid data during the sleep mode, and (iii) develop a methodology which takes into account the new design issues related to the MTCMOS technology. Towards validating the proposed technique, a Personal Digital Assistant (PDA) processor has been implemented using the MTCMOS design methodology, and the 0.18?m process. The fabricated PDA processor operates at 333MHz, and consumes about 2?W of leakage power. Whereas the performance of the MTCMOS implementation is the same as that of the generic CMOS implementation, three orders of reduction in the leakage power has been achieved.
Hyo-Sig Won, Kyo-Sun Kim, Kwang-Ok Jeong, Ki-Tae Park, Kyu-Myung Choi, Jeong-Taek Kong
ISLPED5