EDBT 2026 Demo / reviewers in the wild / expert
Youngsoo Shin
dblp:81/750
· DBLP profile ↗
126ranked-venue papers
22as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 118 · 18 first-author · 18 since 2021Software engineering, systems software and programming languages · 8 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-authorHuman-computer interaction and ubiquitous computing · 5 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Integrated Re-Fragmentation and Curve Correction for Curvilinear Optical Proximity CorrectionabstractCurvilinear optical proximity correction (OPC) treats segments as curves rather than lines, and offers improved correction accuracy. Once a set of segments is identified through fragmentation, it remains the same throughout OPC iterations, limiting OPC performance in runtime and accuracy. We propose additional step of re-fragmentation that can be integrated with curve correction inside the OPC iteration loop. Two machine learning (ML) models are applied for quick re-fragmentation: (1) U-Net is used to detect the critical segments, and (2) MLP identifies the point along the critical curve where the curve should be divided. The reference samples to train the MLP are generated through Bayesian optimization. Compared to standard OPC, the proposed method yields a substantial reduction in OPC iterations (from 20 to 15) and runtime (from 188s to 152s), on average of test clips, when target vertex placement error (VPE) is given. If the number of OPC iterations is fixed (to 25), the proposed method yields average VPE of 1.12 nm, much smaller than 1.65 nm achievable through standard OPC. Junha Jang, Youngsoo Shin |
ASP-DAC | 4 |
| 2026 | Smart Parenting, Smarter Planet: Designing Human-Centered IoT Solutions for Eco-Friendly MotherhoodabstractThis study explores a new approach to designing computing systems that support eco-friendly parenting behaviors. We address a gap in research on sustainable practices in everyday interactions with Internet-of-Things (IoT) technologies. Focusing on the unique challenges of newborn parents managing energy consumption, the research applies a Human-Centered Design (HCD) approach to develop tailored solutions. Using digital ethnography and Focus Group Interview (FGI), the study identifies key pain points and opportunities for embedding sustainability into daily parenting routines. Through iterative prototyping and user evaluations, the study creates the Eco-parenting Mode, an IoT interface featuring personalized automation, real-time energy monitoring, and hands-free control. These features enhance both usability and sustainability. The findings advance Human-Computer Interaction (HCI) and Interaction Design (IxD) by showing how IoT systems can support dynamic caregiving while preserving user agency. Practical contributions include strategies for creating context-aware, inclusive IoT platforms that promote sustainable behaviors in family environments. Youngsoo Shin, Chanhee Shin, Hyunsoo Jang |
Int. J. Hum. Comput. Interact. | 1 |
| 2026 | Empathy in action: An empirical exploration of user perspectives on conversational agent empathyabstract• A holistic framework of conversational agent (CA) empathy is developed and validated, integrating user insights with established empathy theories. • Open-ended surveys reveal core empathic behaviors in CAs, leading to the identification of four key agent traits: Active Listening, Personalization, Emotional Expressivity, and Persona Attractiveness. • The Interpersonal Reactivity Index (IRI) is applied to examine the influence of these traits on user perceptions of cognitive and affective empathy components. • Structural equation modeling confirms the impact of each trait on user evaluations of perspective-taking, fantasy, empathic concerns, and personal distress, providing guidance for designing empathetic CAs. As Conversational Agents (CAs) increasingly interact with users on social and emotional levels, understanding how these agents convey empathy has become a critical challenge. This paper reports on an exploratory mixed-methods study that propose and empirically explores an initial framework for CA empathic responsiveness. The research proceeds in two sequential studies. Study 1 first conducted a qualitative analysis of open-ended surveys (N=166, U.S. and South Korea) to identify key user-defined empathic behaviors. Through thematic analysis, these insights were integrated with existing empathy theories to derive a framework of four core agent characteristics: Active Listening (AL), Personalization (PE), Emotional Expressivity (EE), and Persona Attractiveness (PA). Study 2 then conducted a quantitative investigation in South Korea (N=200) using Structural Equation Modeling (SEM) to test this framework. Empathic responsiveness was operationalized adopting Agent Empathic Reactivity Index (AERI), a validated CA-specific adaptation of the Interpersonal Reactivity Index (IRI), assessing perspective-taking, fantasy, empathic concerns, and personal distress. SEM results confirmed all 12 hypothesized paths. AL and PE strongly enhanced perspective-taking and empathic concerns, while PE also significantly reduced personal distress. Notably, EE and PA had dual effects: they improved positive dimensions, such as fantasy, but also significantly increased users’ perception of the agent’s personal distress. These findings highlight the delicate balance required in designing emotionally resonant CAs. This work advances the theoretical understanding of multidimensional agent empathy and provides actionable, nuanced guidance for designers aiming to build trust and foster long-term user relationships. Bumho Lee, Youngsoo Shin, Byounghyun Yoo |
Int. J. Hum. Comput. Stud. | 2 |
| 2026 | On-Chip Warpage Extraction Through Ring Oscillator Delay TestingabstractWafer warpage is important and critical issue in 3D-IC as well as 2D-IC. Accurate warpage extraction is necessary to screen out wafers with excessive bending before they proceed to subsequent fabrication steps. The conventional method relies on optical measurement, which can be applied only to a few sample wafers and is not practical for full-wafer inspection in high-volume manufacturing. We propose an on-chip warpage extraction technique that leverages delay variation in ring oscillators (ROs) between prebond and postbond test. A key is to derive the mobility change of nMOS and pMOS transistors from RO delay change. The mobility change is subsequently converted into mechanical stress, which becomes a basis of constructing wafer surface map. Experimental results demonstrate that the proposed method extracts wafer warpage with an accuracy of 94% compared to actual measurement. Youngsoo Shin |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | An Island Style Multi-Objective Evolutionary Framework for Synthesis of Memristor-Aided LogicabstractThe optimal in-memory mapping onto memristor crossbars involves competing design goals: minimizing crossbar utilization, reducing delay, and achieving an even layout. Existing heuristic algorithms struggle to address these objectives simultaneously, often yielding suboptimal solutions. This paper introduces an automatic design solution to optimize multiple objectives concurrently. Specifically, it proposes an island-style evolutionary algorithm for multi-objective optimization of in-memory mapping. This algorithm produces a set of solutions, corresponding to Pareto points. Each point can be stored in a library of mapping solutions, which can be chosen when corresponding design is re-used as a macro. Experimental evaluation on IWLS benchmarks demonstrates the effectiveness of this approach in addressing multiple design objectives efficiently. Umar Afzaal, Seunggyu Lee, Youngsoo Shin |
ASP-DAC | 3 |
| 2025 | ML for Computational Lithography: Practical RecipesabstractOptical lithography, also called photolithography, consists of two processes as shown in Figure 1: mask pattern is exposed to a light source to form a photoresist (PR) pattern during optical process; and PR pattern goes through chemical (etch) process to form the final pattern on the wafer. Computational lithography comprises mathematical and algorithmic approaches to improve the resolution attainable in these processes. Its key components are Youngsoo Shin |
ASP-DAC | 1 |
| 2025 | Fast and Accurate Analysis of Power Distribution Network Impedance for DRAM DesignabstractIn DRAM circuits, package and on-chip power distribution network (PDN) should be analyzed as a whole, necessitating impedance analysis. Circuit-level simulation at a number of sample frequency points is a standard practice but is very time consuming. We propose a multi-layer perceptron (MLP) machine learning model for quick prediction of the resonant frequency of each VDD-VSS node pair. Such resonant frequencies are clustered; frequency points for impedance analysis are sampled, densely around clustered resonant frequencies and sparsely outside of clusters. Experimental results show that our method reduces the number of frequency points by 91.6% and the runtime of impedance analysis by 90.4% compared to standard practice, with less than 3% error in predicted maximum impedance. Minseung Shin, Youngsoo Shin |
ISCAS | 3 |
| 2025 | Toward personalized AI-powered recommender systems to support users' daily music choice experiences
Youngsoo Shin, Krithik Ranjan, Michael C. Kowalski |
Int. J. Hum. Comput. Stud. | 1 |
| 2024 | Fast IR-Drop Prediction of Analog Circuits Using Recurrent Synchronized GCN and Y-Net ModelabstractIR-drop analysis of analog circuits is a challenge because the current waveforms of target transistors, with connection to VDD or VSS, are extracted through transistor-level simulation, and the analysis itself, in particular dynamic one, is computationally expensive. We introduce two ML models for high-speed analysis. (1) Recurrent synchronized graph convolutional network (RS-GCN) is used for quick prediction of current waveforms. Each subcircuit is modeled with recurrent-GCN, in which recurrent connection is for the analysis in discrete time series. Recurrent-GCNs are synchronized to take account of common connections including VDD, VSS, and the inputs and outputs of subcircuits. Experiments show that RS-GCN takes only 0.85% of SPICE runtime, while prediction error is 14% on average. (2) Y-Net is applied for actual IR-drop analysis of small layout partition, one by one. Pad location and PDN resistance are provided as one 2D input of Y-Net; they are encoded and go through GCNs to account for neighbor layout partitions. Current map, derived from RS-GCN, becomes the second input. Final IR-drop map is extracted from the decoder. Experiments demonstrate that Y-Net, in conjunction with RS-GCN for current extraction, takes 2.5% of runtime from popular commercial solution with 15% prediction inaccuracy. Seunggyu Lee, Daijoon Hyun, Younggwang Jung, Gangmin Cho, Youngsoo Shin |
DATE | 5 |
| 2024 | Integrated Netlist Synthesis and In-Memory Mapping for Memristor-Aided LogicabstractMemristive memory (memristor) enables logic operations within the memory array, where memristors in the same row or column serve as a logic gate. Logic functions are implemented in the memory through netlist synthesis and in-memory mapping, which assigns each gate operation to specific memristors. The goal is to minimize latency, which represents the number of clock cycles required to complete the operations. While multiple gate operations can be executed in the same clock cycle, additional cycles may be needed for copy operations to align the gate operations. Therefore, assigning each operation to a clock cycle is a challenge. Furthermore, the results of in-memory mapping vary depending on the input netlist. To further reduce latency, an integrated approach is necessary to provide an optimal netlist. We propose two approaches: (1) graph coloring-based in-memory mapping, where the gates are colored to assign sets of gates that operate simultaneously, and (2) integration with mapping-aware netlist synthesis, which iteratively revises the input netlist based on latency evaluation; an incremental method is employed to accelerate the process. Experiments demonstrate that the coloring-based in-memory mapping reduces latency by 17% compared to the state-of-the-art method. The integrated approach achieves an additional 15% reduction in latency. Seunggyu Lee, Youngsoo Shin |
ACM Great Lakes Symposium on VLSI | 3 |
| 2024 | Accurate Interpolation of Library Timing Parameters Through Recurrent Convolutional Neural NetworkabstractInterpolation is used to approximate the timing parameters of logic cells not specified in timing tables. Bilinear interpolation has been taken for granted in the industry, but the error increases as the nonlinearity of the timing parameters increases. In this article, we propose machine learning (ML)-based interpolation to obtain more accurate timing parameters. Recurrent convolutional neural network (R-CNN) is employed and various ranges of table entries form a sequence of input data, in which the recurrent network allows them to influence the interpolation. In addition, variational autoencoder (VAE) is used to capture the distribution feature of the table. ML interpolation is parallelized in GPU to minimize the runtime overhead from numerous arithmetic operations. Experimental results demonstrate that ML interpolation reduces timing parameter error by 19.7% and path delay error by 3.4% compared to bilinear interpolation at the cost of 13% runtime overhead. Daijoon Hyun, Younggwang Jung, Youngsoo Shin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | Decap Insertion With Local Cell Relocation Minimizing IR-Drop Violations and Routing DRVsabstractDecoupling capacitor (decap) cells are inserted near function cells of high switching activities so that their IR-drop can be suppressed. Decaps become more complex these days while a number of metal layers are used for internal connection, thereby starting to manifest themselves as routing blockage. Postplacement decap insertion with both IR-drop violations and routing design rule violations (DRVs) being taken into account is addressed for the first time. Local cell relocation is performed to reduce the number of decaps in the actual decap insertion step. U-Net integrated with a graph convolutional network (GCN) is introduced to predict the DRV probability, which drives decap insertion. The problem of decap insertion is then formulated as mixed integer quadratically constrained programming (MIQCP) and a heuristic algorithm is presented for practical application. Experiments with a few test circuits demonstrate that the increase in routing DRV is reduced by 26% on average with no IR-drop violations, compared to conventional methods that do not explicitly consider DRVs. This brings a 60% reduction in routing runtime and a 33% improvement in total negative slack (TNS). Daijoon Hyun, Younggwang Jung, Youngsoo Shin |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | Decoupling Capacitor Insertion Minimizing IR-Drop Violations and Routing DRVsabstractDecoupling capacitor (decap) cells are inserted near function cells of high switching activities so that their IR-drop can be suppressed. Their design becomes more complex and uses higher metal layers, thereby starting to manifest themselves as routing blockage. Post-placement decap insertion, with a goal of minimizing both IR-drop violations and routing design rule violations (DRVs), is addressed for the first time. U-Net with graph convolutional network is introduced to predict routing DRV penalty. The decap insertion problem is formulated and a heuristic algorithm is presented. Experiments with a few test circuits demonstrate that DRVs are reduced by 16% on average with no IR-drop violations, compared to a conventional method which does not explicitly consider DRVs. This results in 48% reduction in routing runtime and 23% improvement in total negative slack. Daijoon Hyun, Younggwang Jung, Insu Cho, Youngsoo Shin |
ASP-DAC | 4 |
| 2023 | Lightning Talk 21: EDA with ML, Rule-Based, or Both?abstractMachine learning (ML) has been effectively applied to many applications these days mainly because huge amount of data is available through internet. This is not the case in semiconductor industry, where data is not shared between companies or even inside a single organization. ML model, in particular complex one with many parameters, is in danger of overfit when data volume is small and becomes inefficient.Rule-based (or expert) system has been considered a part of AI (artificial intelligence), which also includes ML, and has been popular in many EDA applications. This paper tries to compare the two methods when data volume is high and low. The expectation is that ML is more efficient with high data volume and rule-based is less sensitive to the amount of data and so can be a better choice in low data volume. In addition, combining both methods in a way that rules are revised with some guidance from ML model is investigated so that rule-based method can be a good option in low data volume.Two example applications are considered: OPC refragmentation, which has been addressed through random forest classifier (RFC) ML model [1], and placement utilization in low aspect ratio design, where CNN has been applied [3] to identify utilization distribution over layout sub-regions. Youngsoo Shin |
DAC | 1 |
| 2023 | Power Distribution Network Optimization Using HLA-GCN for Routability EnhancementabstractPower distribution network (PDN) consumes many routing resources to satisfy IR-drop constraints. With the increasing IR drop and the decreasing metal tracks in recent technology, the design of PDN becomes very important for circuit routing. In this paper, post-placement PDN optimization is proposed for routability enhancement. For a given regular PDN, we iteratively remove partial straps that have a small impact on IR-drop while improving routing overflow. Hierarchical layout-aware graph convolutional network (HLA-GCN) is introduced to find the candidate areas for strap removal, and one area is selected based on scoring. This process is applied twice to reduce the candidates for strap removal, and one strap is finally chosen after identifying the actual impact on IR-drop and routing congestion. This method is enabled by fast incremental IR-drop analysis using PDN-GCN, which classifies nodes with voltage change to update only those nodes in the modified nodal analysis. Experimental results address that the proposed method reduces routing overflow by 16% in an acceptable time, where IR-drop values are updated quickly with high accuracy of less than 2% error. Younggwang Jung, Daijoon Hyun, Soyoon Choi, Youngsoo Shin |
ICCAD | 4 |
| 2023 | Multisource Clock Tree Synthesis Through Sink Clustering and Fast Clock Latency PredictionabstractMultisource clock tree consists of a number of local clock trees rooted at respective tap drivers, which are then connected to a clock source through H-tree. We address two key problems for the synthesis of multisource clock tree: clock sink clustering for constructing local clock trees, and the decision of the number of trees. Weight-balanced k-means clustering is applied for the first problem; sinks of the same cluster are localized and the load capacitances of tap drivers are balanced as much as possible. The number of trees can be searched in exhaustive fashion, while clock latency of local trees is estimated with fast CNN-based model. Experiments with a few test circuits demonstrate that clock latency is reduced by 11.8 % on average, while synthesis runtime is reduced by 64% thanks to CNN model. Byungho Choi, Yonghwi Kwon 0002, Umar Afzaal, Youngsoo Shin |
ISCAS | 4 |
| 2023 | Block-Level Power Net Routing of Analog Circuit Using Reinforcement LearningabstractA mixed-signal IC consists of a number of blocks driven by one or more supply voltages. Power net routing determines power wire width and routing path in such a way that routing area is minimized and IR-drop constraints are all satisfied. We propose two-stage power net routing, consisting of trial- and main-routing. In trial routing, reinforcement learning (RL) is applied to find the routing path and wire width assuming larger routing grids. In main routing, we take each net one by one and apply integer linear programming (ILP) to determine the exact path along much smaller grids, while the result from trial routing is respected. Experiments with a few test circuits indicate that the proposed method yields on average of 11% smaller routing area compared to the state-of-the-art method; IR-drop constraints are all satisfied while only 87% are satisfied with the state-of-the-art. Gangmin Cho, Youngsoo Shin |
ISCAS | 3 |
| 2023 | Airgap Insertion and Layer Reassignment Under Setup and Hold Timing ConstraintsabstractAirgap formed in intermetal dielectric (IMD) reduces coupling capacitance, and thus can be utilized for timing optimization. Metal layers with airgap are limited due to high cost of airgap formation. Layer reassignment is to relocate some timing critical wires in nonairgap layers to airgap layers while noncritical wires in airgap layers are reassigned to nonairgap layers. Airgap insertion is to determine the amount of airgaps that are inserted for each critical wires in airgap layers. The two problems are solved in unified fashion with a goal of maximizing setup total negative slack (TNS) while satisfying hold constraints and design rules. They can be formulated as mixed-integer quadratically constrained programming (MIQCP). So, for practical application, a heuristic algorithm is presented and is experimentally compared to MIQCP with small examples. The experiments demonstrate that setup TNS and setup worst negative slack (WNS) are improved by 37% and 8%, respectively; they are improved by 26% and 5% with a simple-minded approach. The algorithm is also parallelized for application to larger circuits; runtime is decreased by 69% with eight threads. Daijoon Hyun, Younggwang Jung, Youngsoo Shin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Routability Optimization of Extreme Aspect Ratio Design through Non-uniform Placement Utilization and Selective Flip-flop StackingabstractCircuits that are placed with very low (or high) aspect ratio are susceptible to routing overflows. Such designs are difficult to close and usually end up with larger area with low area utilization. In this article, we propose two routability optimization methods to implement designs even with very low (or high) aspect ratio and high area utilization. First, we find the best assignment of non-uniform placement utilization through convolutional neural network model, and cell placement is performed while respecting the placement utilization. This allows many cells to be spread out over the entire design rather than being centered. The experiments show that most overflows of 16.5% occurring in cell placement are removed with 23.1% reduction in wire length; this is the result of further improving overflow of 9.8% compared to a conventional method. In the second, some flip-flops are selectively stacked to reduce the routing resources used for clock routing. U-Net model is built with graph attention network to predict the congestion after clock-tree synthesis, and the flip-flops in highly congested areas are selected for stacking. The proposed method improves the overflows, which occurs after clock-tree synthesis, by 22.1%. Daijoon Hyun, Sunwha Koh, Younggwang Jung, Youngsoo Shin |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2021 | Routability Optimization for Extreme Aspect Ratio Design Using Convolutional Neural NetworkabstractCircuits that are placed with very low (or high) aspect ratio are susceptible to routing overflows. Such designs are difficult to close and usually end up with larger area with low area utilization. We observe that non-uniform setting of utilization target greatly helps in these designs, specifically low utilization in the center and gradually higher utilization toward the ends. We introduce a convolutional neural network (CNN) model to predict the setting of utilization target values. Experiments indicate that routing congestion overflows are reduced by 29% on average of test designs with 40% reduction in wirelength. Sunwha Koh, Younggwang Jung, Daijoon Hyun, Youngsoo Shin |
ISCAS | 4 |
| 2021 | Dynamic IR Drop Prediction Using Image-to-Image Translation Neural NetworkabstractDynamic IR drop analaysis is very time consuming, so it is only applied in signoff stage before tapeout. U-net model, which is an image-to-image translation neural network, is employed for quick analysis of dynamic IR drop. A number of feature maps are used for u-net input: a map of effective PDN resistance seen from each gate, a map of current consumption of each gate (in particular time instance), and a map of relative distance to nearest power supply pad. A layout is partitioned into a grid of regions and IR drop is predicted region-by-region. For fast prediction, (1) analysis is performed only in time windows which are estimated to cause high IR drop, and (2) effective PDN resistance is approximated through a proposed simplification method. Experiments with a few test circuits demonstrate that dynamic IR drop is predicted 20 times faster than commercial analysis package with 15% error. Yonghwi Kwon 0002, Giyoon Jung, Daijoon Hyun, Youngsoo Shin |
ISCAS | 4 |
| 2020 | Integrated Airgap Insertion and Layer Reassignment for Circuit Timing optimizationabstractAirgap is an intentional void formed in inter-metal dielectric (IMD). It brings about reduced coupling capacitance, and so can be used to improve circuit timing. Airgap can be utilized in a limited number of metal layers due to its high process cost. For given airgap layers, two problems should be addressed to insert airgap: relocate some metal segments in non-airgap layers into airgap layers (called layer reassignment) and determine the amount of airgap for each metal segment in airgap layers (airgap insertion). Two problems are solved together in this paper with a goal of maximizing setup total negative slack (TNS) while assuring no hold violations. It is formulated as mixed integer quadratically constrained programming (MIQCP); heuristic algorithm is proposed for practical application and its performance against MIQCP is experimentally assessed using small test circuits. Experiments demonstrate that TNS and WNS are improved by 35% and 10%, respectively, while simple minded approach achieves 6% and 4% less improvements compared to the proposed method. Younggwang Jung, Daijoon Hyun, Youngsoo Shin |
ASP-DAC | 3 |
| 2020 | Fast ECO Leakage Optimization Using Graph Convolutional NetworkabstractAt the very late design stage, engineering change order (ECO) leakage optimization is often performed to swap some cells for the ones with lower leakage, e.g. the cells with higher threshold voltage (Vth) or with longer gate length. It is very effective but time consuming due to iterative nature of swap and timing check with correction. We introduce a graph convolutional network (GCN) for quick ECO leakage optimization. GCN receives a number of input parameters that model the current timing information of a netlist as well as the connectivity of the cells in a form of a weighted connectivity matrix. Once it is trained, GCN predicts exact Vth (with Vth given by commercial ECO leakage optimization as a reference) of 83% of cells, on average of test circuits. The remaining 17% of cells are responsible for some negative timing slack. To correct such timing as well as to remove any minimum implant width (MIW) violations, we propose a heuristic Vth reassignment. The combined GCN and heuristic achieve 52% reduction of leakage, which can be compared to 61% reduction from commercial ECO, but with less than half of runtime. Yonghwi Kwon 0002, Youngsoo Shin |
ACM Great Lakes Symposium on VLSI | 3 |
| 2020 | Pre-layout clock tree estimation and optimization using artificial neural networkabstractClock tree synthesis (CTS) takes place in a very late design stage, so most of the time, power consumption is analyzed while a circuit does not contain a clock tree. We build an artificial neural network (ANN) to estimate the number of clock buffers and apply to each clock gater as well as clock source in ideal clock network. Clock structure is then constructed using such estimated clock buffers. Experiments with a few test circuits demonstrate very high accuracy of this method, average clock power estimation error less than 5%. The proposed method also allows us to find the possible minimum number of clock buffers with optimized clock parameters (e.g. target skew, clock transition time). The possible minimum number of buffers can be found by binary search algorithm and on each step of the algorithm, trained ANN is used to find such clock parameters for the target number of buffers. Using proposed clock parameter optimization, we found that the number of buffers in clock network can be reduced by 31%, on average. Sunwha Koh, Yonghwi Kwon 0002, Youngsoo Shin |
ISLPED | 3 |
| 2019 | Accurate Wirelength Prediction for Placement-Aware Synthesis through Machine LearningabstractPlacement-aware synthesis, which combines logic synthesis with virtual placement and routing (P&R) to better take account of wiring, has been popular for timing closure. The wirelength after virtual placement is correlated to actual wirelength, but correlation is not strong enough for some chosen paths. An algorithm to predict the actual wirelength from placement-aware synthesis is presented. It extracts a number of parameters from a given virtual path. A handful of synthetic parameters are compiled through linear discriminant analysis (LDA), and they are submitted to a few machine learning models. The final prediction of actual wirelength is given by the weighted sum of prediction from such machine learning models, in which weight is determined by the population of neighbors in parameter space. Experiments indicate that the predicted wirelength is 93% accurate compared to actual wirelength; this can be compared to conventional virtual placement, in which wirelength is predicted with only 79% accuracy. Daijoon Hyun, Yuepeng Fan, Youngsoo Shin |
DATE | 3 |
| 2019 | Clock Gating Synthesis of Netlist with Cyclic Logic PathsabstractGate-level clock gating is to synthesize clock gating structure (grouping of registers and extracting gating function of each group) from a netlist. We note that a simpler gating function can be derived from a cyclic logic path that connects the input and output of the same register. Another benefit comes from the fact that simplifying the cyclic paths using the derived gating function as don't-care is straightforward. A key problem in this approach is to extract a set of cyclic paths of each register, such that power consumption is minimized and circuit timing is left intact. Experiments demonstrate that power consumption is reduced by 49% on average of test circuits (with initial ungated netlist as a reference), while a sample previous gate-level clock gating achieves 34% of power saving. Yonghwi Kwon 0002, Inhak Han, Youngsoo Shin |
ICCAD | 3 |
| 2019 | Endurance Enhancement of Multi-Level Cell Phase Change MemoryabstractPhase change memory (PCM) is a promising device for its good scalability and negligible standby power consumption. Multi-level cell (MLC) PCM allows higher memory density, but it suffers from reduced endurance due to frequent RESET operations during writing. Inter-state direct write (ISDW) method is proposed, in which intermediate states `01' and `10' are reached without RESET initialization. A new MLC PCM model is presented, which takes account of phase configuration of each MLC PCM state; the feasibility of ISDW is assessed using the model. Compression-based RESET removal encoding (CRE) is also proposed to further reduce the number of RESET operations. Experiments demonstrate that the proposed methods achieve 38.4× enhancement of cell endurance; the writing energy dissipation is reduced to 31% on average of test cases. Cheongwon Lee, Youngsoo Song, Youngsoo Shin |
ICCAD | 3 |
| 2019 | Standard Cell Layout Design and Placement Optimization for TFET-Based CircuitsabstractTunneling Field-Effect Transistors (TFETs) have a potential to decrease supply voltage of integrated circuits thanks to the superior subthreshold swing. However, the source and drain of TFETs are doped in different types (one in n+and the other in p+), which raises challenges in fabrication in sub-10nm processes. We propose a method to optimize standard cell layouts for TFETs, in which consistent doping profile is maintained in the vertical direction so that design rule violations due to small spacing between implantation masks are resolved. We also notice that the footprints of some standard cells turn out to be rectilinear. A post-placement optimization method to join the cell layouts is also addressed. We finally propose a TFET fabrication process using self-aligned quadruple patterning (SAQP), which can enable TFET fabrication in sub-10nm processes. Our proposed methods bring about 4.5% area reduction, based on experiments with a set of test circuits. Youngsoo Song, Jinwook Jung, Youngsoo Shin |
ISCAS | 3 |
| 2019 | Neural Network Classifier-Based OPC With Imbalanced Training DataabstractMachine learning-guided optical proximity correction, called ML-OPC in this paper, has recently been proposed to alleviate long runtime of model-based OPC. ML-OPC using regression methods has been presented but with limited prediction accuracy. We propose neural network classifier-based OPC (NNC-OPC), in which a neural network classifier serves as a mask bias model. A few techniques are applied to enhance basic NNC-OPC: parameterization of layout segments using polar Fourier transform signals, dimensionality reduction through weighted principal component analysis, and sampling of training layout segments. Training segments are typically imbalanced over the range of mask biases, which may cause large prediction error for segments that appear less frequently. This is resolved by three techniques: 1) synthetic data generation; 2) class reorganization; and 3) an adaptive learning rate. Experiments with NNC-OPC with all techniques applied indicate that prediction error of mask bias and training time are reduced by 29% and 80%, respectively, compared to state-of-the-art ML-OPC with regression methods. Suhyeong Choi, Seongbo Shim, Youngsoo Shin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Integrated Approach of Airgap Insertion for Circuit Timing OptimizationabstractAirgap technology enables air to be introduced in inter-metal dielectric (IMD). Airgap between certain wires reduces coupling capacitance due to the reduced permittivity; this can be utilized to decrease circuit delay. We propose an integrated approach of airgap insertion with the goal of circuit timing optimization. It consists of three sub-problems. We first select the layers that employ airgap, called airgap layers, that maximize total negative slack (TNS) improvement; this yields TNS improvement of 7% to 15% and worst negative slack (WNS) improvement of 2% to 8%, compared to a simple assumption of airgap layers. Second, we reassign the layers of wires such that more wires on critical paths can be placed in airgap layers. This is formulated as integer linear programming (ILP), and a more practical heuristic algorithm is also proposed. It provides an additional 17% TNS improvement and 6% WNS improvement. Finally, we perform airgap insertion through ILP formulation, where a number of design rules are modeled with linear constraints. To reduce the heavy runtime of ILP, a layout partitioning technique is also applied. It implements a feasible airgap mask in a manageable time where the amount of inserted airgap is close to the optimal solution. Daijoon Hyun, Youngsoo Shin |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2019 | Integrated Latch Placement and Cloning for Timing OptimizationabstractThis article presents an algorithm for integrated timing-driven latch placement and cloning. Given a circuit placement, the proposed algorithm relocates some latches while circuit timing is improved. Some latches are replicated to further improve the timing; the number of replicated latches along with their locations are automatically determined. After latch cloning, each of the replicated latches is set to drive a subset of the fanouts that have been driven by the original single latch. The proposed algorithm is then extended such that relocation and cloning are applied to some latches together with their neighbor logic gates. Experimental results demonstrate that the worst negative slack and the total negative slack are improved by 24% and 59%, respectively, on average of test circuits. The negative impacts on circuit area and power consumption are both marginal, at 0.7% and 1.9% respectively. Jinwook Jung, Gi-Joon Nam, Woohyun Chung, Youngsoo Shin |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2019 | Cut Optimization for Redundant Via Insertion in Self-Aligned Double PatterningabstractRedundant via (RV) insertion helps prevent via defects and hence leads to yield enhancement. However, RV insertion in self-aligned double patterning (SADP) processes is challenging since cut optimization has to be considered together. In SADP, parallel one-dimensional metal lines are divided into signal wires and dummy wires by line-end cuts. If an RV is inserted, signal wires need to be extended to connect to the RV. To this end, an additional cut, which we call RV cut, is introduced to make a space for the extension. Since RV cuts and line-end cuts are manufactured with the same mask set, design rules between those cuts have to be honored, which incurs proper distribution and mask assignment to individual cuts. In this article, we address a problem of integrated RV insertion and cut optimization. We show that the problem can be formulated as an integer linear programming (ILP). We also propose a heuristic algorithm is presented for practical application, in which potential locations of RVs are first identified and used to properly insert as many RVs as possible while minimizing the conflict between RV cuts. Our experimental results demonstrate that 75% of vias receive RVs with 8% increase in total wire length, which is only slightly worse than the optimal result obtained by ILP. Youngsoo Song, Daijoon Hyun, Jingon Lee, Jinwook Jung, Youngsoo Shin |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2018 | Automatic insertion of airgap with design rule constraintsabstractAirgap is a technology that enables air to be used as IMD (inter metal dielectric). It brings about reduced coupling capacitance, which helps reduce circuit delay and power consumption. Airgap is constrained by a number of design rules. Manual insertion of airgap while design rules are all respected is inconvenient and time consuming. We address automatic airgap insertion in this paper, in which the goal is to insert maximum amount of airgap in selected paths (e.g. timing critical paths) while related design rules are all honored. Our approach consists of three steps: (1) layout is decomposed into a set of sublayouts, such that airgap can be inserted in each sublayout independently, (2) each of large sublayouts is further partitioned in heuristic fashion, and (3) airgap insertion in each sublayout (or each partition of sublayout) is performed through ILP (integer linear programming). Experiments indicate that runtime is manageable; the impact of airgap on circuit delay is also demonstrated, e.g. 5.6% improvement of worst slack on average of test circuits. Daijoon Hyun, Youngsoo Shin |
ASP-DAC | 2 |
| 2018 | Fast Timing Analysis of Non-Tree Clock Network with Shorted WiresabstractA non-tree clock network, such as crosslink and mesh, includes some shorted wires to reduce clock skew. A short-circuit current that flows through the shorted wires makes conventional static timing analysis (STA) inapplicable. Transistor-level simulation may be applied but takes long time. We address a fast timing analysis of non-tree clock network. A partial circuit made of drivers, shorted wires, and receivers is extracted and represented as voltagedependent current sources with π-model of RC load. Given voltage waveforms at driver inputs, we calculate the waveform at each shorted node by repeating nodal analysis for each time step; the waveform is represented as piecewise linear function. As the waveform propagates to receiver input via RC tree, the responses for all linear segments are obtained and merged into a full waveform. The waveform at receiver input then passes through receiver to produce a linear waveform at receiver output. Finally, timing parameters from the waveform at receiver output are transferred to STA, such that it utilizes the parameters to analyze the remaining circuit from receiver outputs to clock sinks. Experiments with a few test circuits demonstrate that analysis time is reduced by 10× with only 1% error on average (both in delay and transition time) compared to SPICE. Kiwon Yoon, Daijoon Hyun, Youngsoo Shin |
ACM Great Lakes Symposium on VLSI | 3 |
| 2018 | Library Optimization for Near-Threshold Voltage DesignabstractA circuit operating at near-threshold voltage (NTV) dissipates much less energy, but it suffers from significant increase in cell delay as well as delay variation. In this paper, we address two library optimization methods for NTV design: (1) transistor lengths are increased to benefit from reverse short channel effect (RSCE), and (2) each flip-flop is optimized into a few versions with different timing parameters by redistributing clock signals to clocked transistors. Flip-flops are remapped to optimized ones via integer linear programming (ILP); a goal is to minimize total negative slack (TNS) under hold time constraints, which is a critical concern in NTV design. Experiments demonstrate that our proposed method achieves 22%, 56%, and 13% reductions in clock period, energy dissipation, and circuit area, respectively, on average of a few test circuits in 55-nm technology. Daijoon Hyun, Jaewoo Seo, Youngsoo Shin |
ISCAS | 3 |
| 2018 | Transient Clock Power Estimation of Pre-CTS NetlistabstractClock tree synthesis (CTS) is performed in a very late stage of design. Power estimation, therefore, can only be done without clock network in most design stages, which is not desirable given that clock network is usually the biggest power consumer. One may adopt an estimate of clock power, but its dynamic nature arising from clock gating brings a challenge in the estimation of clock power in a pre-CTS design. In this paper, we (1) estimate the clock tree components (clock gating cells (CGCs) and buffers as well as their wireloads) by using artificial neural networks (ANNs) and (2) use them while gating or ungating of each CGC is identified from a netlist cycle-by-cycle to estimate transient clock power consumption. Experiments with a few test circuits indicate that (1) the estimation of clock tree components causes the error of 13% on average, and (2) the estimated clock power waveform is very close to the actual waveform with average error of only 2%. Yonghwi Kwon 0002, Jinwook Jung, Inhak Han, Youngsoo Shin |
ISCAS | 4 |
| 2018 | Fast Timing Analysis of Transistor-Level Full Custom Digital CircuitsabstractThis paper presents a fast timing analysis methodology that can be applied to full-custom digital circuits. Given a transistor-level circuit netlist, we build a timing graph that consists of gates and wires. Each gate is modeled as a set of equivalent RC networks representing a specific input pattern while taking into account of stacked transistor effect and Miller effect. Wires are also modeled into RC trees using a Rectilinear Steiner minimal tree. An improved RC delay model that takes input transition time into account is used for computing the propagation delays of the RC networks, which is also proposed in this paper. Gates and wires are then modeled into a hardware description language (HDL) so that the timing analysis is performed using an off-the-shelf function simulator. Experimental results on a few test circuits indicate that up to 1800× faster timing analysis can be realized compared to SPICE; the average error of the proposed delay model is 11.2%. Jingon Lee, Jinwook Jung, Youngsoo Shin |
ISCAS | 3 |
| 2018 | Module grouping to reduce the area of test wrappers in SoCs
Youngsoo Shin |
Integr. | 2 |
| 2018 | Memory-Efficient Parametric Semiglobal MatchingabstractAccurate stereo matching for depth extraction requires a large memory space, which restricts its use in resource-limited systems. The problem is aggravated by the recent trend of applications requiring significantly high pixel resolution and disparity levels. To alleviate the high memory requirement, we propose to represent the aggregation costs as a Gaussian mixture model (GMM) function. Only a set of GMM parameters is stored and used instead of all the costs for each pixel. We also propose GMM parameter update-based aggregation along multiple paths. To preserve the accuracy of the disparity map, we employ a depth confidence measure and propose an update rule for the slanted surface of an object. Experimental results over the KITTI dataset show that the proposed method reduces the memory requirement to less than 5% of that of semiglobal matching, while the accuracy is maintained at the level of state-of-the-art semiglobal and local methods. Yeongmin Lee, Min-Gyu Park, Youngbae Hwang, Youngsoo Shin, Chong-Min Kyung |
IEEE Signal Process. Lett. | 4 |
| 2018 | OWARU: Free Space-Aware Timing-Driven Incremental Placement With Critical Path SmoothingabstractThis paper presents an incremental timing-driven placement tool, named OWARU. It optimizes timing critical paths through a free space-aware path smoothing: the gates on such paths are relocated to free spaces around the smoothed paths, while incremental static timing analysis is involved to accurately assess timing changes due to the relocation. OWARU is extended to accommodate gate sizing and layer assignment to demonstrate the effectiveness of unified physical synthesis optimizations and incremental placement. The goal is to show that OWARU is an ideal platform for timing closure at later stages of a physical design flow. OWARU is applied on a set of test circuits from 14-nm high-performance commercial microprocessors, which originally failed in timing closure. On average, the worst slack is improved by 63.6%, which corresponds to 5.0% of the clock period; total negative slack is improved by 69.1%. Jinwook Jung, Gi-Joon Nam, Lakshmi N. Reddy, Iris Hui-Ru Jiang, Youngsoo Shin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2018 | Folded Circuit Synthesis: Min-Area Logic Synthesis Using Dual-Edge-Triggered Flip-FlopsabstractThe area required by combinational logic of a sequential circuit based on standard flip-flops can be reduced by identifying subcircuits that are identical. Pairs of matching subcircuits can then be replaced by circuits in which dual-edge-triggered flip-flops operate on multiplexed data at the rising and falling edges of the clock signal. We show how to modify the Boolean network describing a combinational logic to increase the opportunities for folding, without affecting its function. Experiments with benchmark circuits achieved an average reduction in circuit area of 18%. Inhak Han, Youngsoo Shin |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2017 | Pin Accessibility-Driven Cell Layout Redesign and Placement OptimizationabstractThe layout of standard cells is very dense these days, so some pins are hard to get access to. This is in particular true in complex cells with many pins (e.g. AOI) and in the layout where many of those cells are densely packed without much whitespace. We redesign those complex cells, so a library now contains both original cell and its new version with easier pin access; a systematic method is proposed to pick candidate cells for redesign and to dictate how redesign should be performed. We also introduce a measure of inaccessibility of pins in a cell, named IOC. Placement optimization is performed, which uses IOC to determine which cells should be replaced by its redesigned version and how whitespace should be redistributed. Experiments with 12 test circuits indicate that the number of routing errors (after the initial placement) is reduced by 82% on average, and the subsequent detailed routing takes 72% less runtime. Jaewoo Seo, Jinwook Jung, Youngsoo Shin |
DAC | 4 |
| 2017 | Cut mask optimization for multi-patterning directed self-assembly lithographyabstractLine-end cut process has been used to create very fine metal wires in sub-14nm technology. Cut patterns split regular line patterns into a number of wire segments with some segments being used as actual routing wires. In sub-7nm technology, cuts are smaller than optical resolution limit, and a directed self-assembly lithography with multiple patterning (MP-DSAL) is considered as a patterning solution. We address cut mask optimization problem for MP-DSAL, in which cut locations are determined in such a way that cuts are grouped into manufacturable clusters and assigned to one of masks without MP coloring conflicts; minimizing wire extensions is also pursued in the process. Only a restricted version of this problem has been addressed before while we do not assume any such restrictions. The problem is formulated as ILP first, and a fast heuristic algorithm is also proposed for application to larger circuits. Experimental results indicate that the ILP can remove all coloring conflicts, and reduce total wire extensions by 93% on average compared to those obtained by the restricted approach. Heuristic achieves a similar result with less than 1% of coloring conflicts and 91% reduction in total wire extensions. Wachirawit Ponghiran, Seongbo Shim, Youngsoo Shin |
DATE | 3 |
| 2017 | Timing-aware wire width optimization for SADP processabstractWire width optimization for SADP process is addressed, which involves a decision of how cut- and block-masks should form; a goal is to reduce wire delay in timing critical paths. The problem is formulated using a graph: a vertex corresponds to wire segment with its maximum length for widening as a vertex weight; an edge represents a potential conflict between two candidate wire segments that we wish to widen. A maximum weight independent set corresponds to an ideal solution. For a few circuits that we test, wire resistance of timing critical nets is reduced by 18.5% on average, which leads to 9.9% reduction in clock period. Youngsoo Song, Youngsoo Shin |
DATE | 3 |
| 2017 | Redundant Via Insertion with Cut Optimization for Self-Aligned Double PatterningabstractLine-end cuts are employed to enable 1D gridded designs in self-aligned double patterning (SADP) process. Due to the minimum spacing constraints between adjacent cuts, cut optimization is important component. However, it brings a new challenge to redundant via (RV) insertion. As the cuts for RVs are not taken into account during line-end cut optimization, inserting some RVs may cause coloring conflicts or design rule violations. Youngsoo Song, Jinwook Jung, Youngsoo Shin |
ACM Great Lakes Symposium on VLSI | 3 |
| 2017 | Redundant Via insertion in SADP process with cut merging and optimizationabstractLine-end cuts in self-aligned double patterning (SADP) process are employed for printing lD-gridded patterns. Redundant via (RV) requires another cut, named RV cut, to be introduced, which may cause coloring conflicts or design rule violations with adjacent line-end cuts. RV insertion should be coordinated together with cut optimization so that maximum number of RVs are inserted while incurring no coloring conflicts among cuts. A technique named cut merging is addressed to remove cut conflicts and thereby increase the number of RV candidates. Cut redistribution and color assignment (for both line-end and RV cuts) are also taken into account to further increase RV candidates. Youngsoo Song, Jinwook Jung, Youngsoo Shin |
VLSI-SoC | 3 |
| 2017 | Design for experience innovation: understanding user experience in new product developmentabstractIn providing a better experience to users in terms of product usage, we focus on the important concept of a user-centred design (UCD), and explore a new approach to user experience (UX), with the effort to understand experience-driven innovation. Based on the conceptual framework of experiential network and the results of multiple case studies covering 643 successfully designed products or services providing an optimised UX, we categorise the UX context into the following four representative types: individualisation, combination, integration, and ecosystem. Furthermore, we identify the essential UCD concepts that reflect the core needs and expectations of users in each of the designed contexts, that is, specialty, usefulness, usability, and fluency. Finally, we discuss the dynamic concepts that help achieve a successful experience innovation. We expect these findings to play a crucial role in the development of novel design concepts or strategies, not only to better understand the needs of contemporary users, but also to better understand the dynamics of innovation. Youngsoo Shin, Chaerin Im, Hyosun Oh, Jinwoo Kim 0001 |
Behav. Inf. Technol. | 1 |
| 2017 | Fast Verification of Guide-Patterns for Directed Self-Assembly LithographyabstractGuide-patterns (GPs) are critical to the construction of contacts and vias in directed self-assembly (DSA) lithography. Simulations can be used to verify GPs, but runtime is excessive. Instead, we categorize the shapes of GPs using a small number of geometric parameters. Then a verification function is built to predict whether a GP will produce the required contacts, as follows: a vector in parameter space is constructed to represent each GP in a test set; the acceptability of each GP is then assessed by DSA simulation, and each vector is tagged “good” or “bad” accordingly; next, the parameter space is deformed to convert a radial distribution into one in which the good and bad vectors can be separated by a hyper-plane, which finally becomes the verification function. We also show how to reduce the dimensionality of the parameter space by principal component analysis, and how to generalize the geometric description of GPs to allow different types of GP to be verified in a uniform fashion. The proposed GP verification is demonstrated in 10 nm technology. Seongbo Shim, Youngsoo Shin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | Lithography Defect Probability and Its Application to Physical Design OptimizationabstractModern standard cells contain intercell margins at the left and right ends for better lithography. We introduce defect probability, which is the probability that a lithography defect occurs if the margins between two adjacent cells are missing. Computing the defect probability of all cell pairs is impractical due to lengthy lithography simulations and huge number of cell pair combinations. Two approximate methods are employed to make this computation possible: reducing the range of optical proximity correction and grouping cell pairs of similar geometry at the cell boundary. We also present how the cell layout can be modified for a lower defect probability with no impact on the cell electrical parameters. Defect probability is applied to two physical design optimization problems. In the automatic placement, we consider that all cells are initially without margins. We want to locate two cells adjacent if their defect probability is zero (or negligibly small) or insert margins in between; this is achieved using the average defect probability as one of the cost terms of the placement. Experiments in 28-nm commercial library demonstrate an 8% reduction in the area with a 4% shorter wirelength. In the second application, we assume that the standard placement using cells with margins have been performed. We want to identify redundant margins that can be removed while the defect probability is kept zero. We take a step forward and shuffle the location of a few consecutive cells in the same row so that more redundant margins are identified. Once all the redundant margins are removed, newly created whitespace is distributed to reduce routing congestion in highly congested areas. Experiments indicate a 48% reduction in the number of overflow routing grids. Seongbo Shim, Woohyun Chung, Youngsoo Shin |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Buffer insertion to remove hold violations at multiple process cornersabstractBuffer insertion to remove hold violations at multiple process corners is addressed for the first time. The problem is formulated as integer linear programming (ILP); it is combined with circuit partitioning heuristic so that larger circuits can also be handled. A heuristic buffer insertion algorithm is then proposed and compared to ILP, which demonstrates only a slight increase of the number of buffers (2.4% on average). Two additional intuitive methods are implemented to demonstrate why new heuristic algorithm is needed: conventional buffer insertion at each process corner one by one and conventional buffer insertion at all process corners simultaneously followed by combining insertion results. Inhak Han, Daijoon Hyun, Youngsoo Shin |
ASP-DAC | 3 |
| 2016 | Mask optimization for directed self-assembly lithography: Inverse DSA and inverse lithographyabstractIn directed self-assembly lithography (DSAL), a mask contains the images of guide patterns (GPs), which are patterned on a wafer through optical lithography; the wafer then goes through DSA process to pattern contacts. Mask design for DSAL, which is the opposite of the above processes, consists of two key steps, inverse DSA and inverse lithography, which we address in this paper. In inverse DSA, we progressively refine GPs until they produce target contacts as closely as possible. GP is defined as a function of a few geometry parameters, and how sensitive the contacts are to the parameters are calculated which then guides how much the GP should be refined. In inverse lithography, mask is progressively refined so that target GPs are produced. Mask is defined by pixel values and their gradient guides the direction that the mask should be refined. There are too many pixels for gradient calculation; the method to approximate calculation is proposed. Inverse DSA and inverse lithography are extended to handle process variations. We modify basic inverse lithography so that the resulting mask becomes less sensitive to lithography variations; basic inverse DSA is modified so that it provides the way this sensitivity can be checked. Seongbo Shim, Youngsoo Shin |
ASP-DAC | 2 |
| 2016 | Redundant via insertion for multiple-patterning directed-self-assembly lithographyabstractIn sub-7nm technology, the size and pitch of vias are much smaller than optical resolution limit, and directed self-assembly lithography with multiple patterning technology (MP-DSAL) has been proposed as a solution. In MP-DSAL, vias that are close are clustered and patterned together via DSAL process, and via clusters that are close are printed using different masks via MP. Redundant vias, which are typically used for better via manufacturability, should be inserted very carefully in MP-DSAL because some redundant vias may cause large and complex via clusters, which are undesirable in DSAL; some other redundant vias may cause mask assignment of via clusters impossible, often called MP coloring conflict. Seongbo Shim, Woohyun Chung, Youngsoo Shin |
DAC | 3 |
| 2016 | Redundant via insertion in directed self-assembly lithography
Woohyun Chung, Seongbo Shim, Youngsoo Shin |
DATE | 3 |
| 2016 | OWARU: free space-aware timing-driven incremental placementabstractThis paper proposes a powerful new technique called “OWARU”1 that re-places and re-sizes multiple gates simultaneously to improve the most critical paths of a design. In essence, it is an incremental timing-driven placement technique integrated with gate sizing optimization that runs in conjunction with static timing analysis to guarantee a WYSIWYG 2 property. The OWARU technique offers several key advantages over previous techniques such as geometrical path straightening via the Bézier-curve algorithm, free space awareness to guarantee a legal placement solution, and an accurate true timing mode. The Bézier-curve geometric smoothing algorithm is extended with new anchor placement techniques to further improve the path placement. Free space aware placement algorithm is further enhanced with multiple gate optimization. The preliminary results are promising. We applied the OWARU technique at the end of industrial strength physical synthesis optimization on high performance microprocessor designs. The technique was extremely effective in improving the most critical path of the tested designs. On timing critical paths that were not fully closed from the previous physical synthesis optimization, the WS (worst slack) is improved by 5.3% of the total clock period and the TNS (total negative slack) improved by 91.3% on average. Jinwook Jung, Gi-Joon Nam, Lakshmi N. Reddy, Iris Hui-Ru Jiang, Youngsoo Shin |
ICCAD | 5 |
| 2016 | Crosslink insertion for minimizing OCV clock skewabstractCrosslinks may be inserted in a few clock tree nodes to reduce on-chip variation induced clock skew, simply called OCV skew. A change in clock transition and clock latency should be accurately estimated and be reflected in crosslink insertion algorithm, which we study. Fast estimation of OCV skew is important, which we also address. Crosslink insertion problem is modeled into a graph, and is solved through integer linear programming (ILP) as well as a fast heuristic. Experiments in 28-nm technology indicate that maximum OCV skew is reduced by 44% and 35%, on average, by ILP and heuristic algorithm, respectively. Kiwon Yoon, Seongbo Shim, Youngsoo Shin |
ISCAS | 3 |
| 2016 | Wakeup scheduling and its buffered tree synthesis for power gating circuits
Seungwhun Paik, Seokhyeong Kang, Youngsoo Shin |
Integr. | 4 |
| 2016 | Synthesis of Dual-Mode Circuits Through Library Design, Gate Sizing, and Clock-Tree OptimizationabstractA dual-mode circuit is a circuit that has two operating modes: a default high-performance mode at nominal voltage and a secondary low-performance near-threshold voltage (NTV) mode. A key problem that we address is to maximize NTV mode clock frequency. Some cells that are particularly slow in NTV mode are optimized through transistor sizing and stack removal; static noise margin of each gate is extracted and appended in a library so that function failures can be checked and removed during synthesis. A new gate-sizing algorithm is proposed that takes account of timing slacks at both modes. A new sensitivity measure is introduced for this purpose; binary search is then applied to find the maximum NTV mode frequency. Clock-tree synthesis is reformulated to minimize clock skew at both modes. This is motivated by the fact that the proportion of load-dependent delay along clock paths, as well as clock-path delays themselves, should be made equal. Experiments on some test circuits indicate that NTV mode clock period is reduced by 24%, on average; clock skew at NTV decreases by 13%, on average; and NTV mode energy-delay product is reduced by 20%, on average. Seokhyeong Kang, Youngsoo Shin |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2016 | One-Cycle Correction of Timing Errors in Pipelines With Standard Clocked ElementsabstractOne of the most aggressive uses of dynamic voltage scaling is timing speculation, which in turn requires fast correction of timing errors. The fastest existing error correction technique imposes a one-cycle time penalty only, but it is restricted to two-phase transparent latch-based pipelines. We perform one-cycle error correction by gating only the main latch in each stage of the pipeline that precedes a failed stage. This new method is applicable to widely used clocking elements, such as flip-flops and pulsed latches. Because it prevents inputs arriving at a stage, which is stalled, it can also be used in pipelines with multiple fan-in, fan-out, and looping. Simulations show an energy saving of 8%-12% with a target throughput of 0.9 instructions per cycle, and 15%-18% when the target is 0.8. Insup Shin, Jae-Joon Kim, Yu-Shiang Lin, Youngsoo Shin |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2015 | Identifying redundant inter-cell margins and its application to reducing routing congestion
Woohyun Chung, Seongbo Shim, Youngsoo Shin |
DATE | 3 |
| 2015 | Defect Probability of Directed Self-Assembly Lithography: Fast Identification and Post-Placement OptimizationabstractIn directed self-assembly lithography (DSAL), an inter-cell cluster of contacts, which crosses the boundary of cells, is more likely to cause patterning failure because corresponding guide pattern (GP) has not been verified beforehand. All forms of inter-cell clusters can systematically be identified and grouped, which allows us to define DSA defect probability when two arbitrary cells are placed side by side. We then address post-placement optimization, in which some cells are flipped and some cells are swapped with their adjacent cells so that the number of whitespaces inserted in between cell pairs of high defect probability is minimized. Experiments with a few test circuits demonstrate 11% increase of placement density, on average, with no expected DSA defects. Seongbo Shim, Woohyun Chung, Youngsoo Shin |
ICCAD | 3 |
| 2015 | Physical synthesis of DNA circuits with spatially localized gatesabstractWith the current DNA nanotechnology, we are now able to arrange DNA molecules on a DNA origami to compose a logic gate. This in turn realizes a spatially localized DNA circuit, on which the logic gates are placed on the specific locations as in electronic circuits. In this paper, we address three key problems in designing large-scale spatially localized DNA circuits. An AND gate, made of four hairpins, functions in stochastic manner and sometimes outputs a wrong result. Given tolerable error probability at each circuit output, we address how the probability that each AND gate functions correctly can be determined, which in turn determines the location of constituent hairpins. In the second problem, we study how hairpins are arranged on a DNA origami to minimize the area of a whole circuit, which determines the area of the origami board. The third problem regards the DNA domain assignment so that connected gates can communicate without interference. Jinwook Jung, Daijoon Hyun, Youngsoo Shin |
ICCD | 3 |
| 2015 | Physical design and mask optimization for directed self-assembly lithography (DSAL)abstractIn DSAL, contact holes are indirectly formed through guide patterns (GPs). Thus the integrity of GPs is very important, in particular when GP shape is large and complex. The limitation in GP shape calls for careful consideration in physical design stage. In mask optimization, synthesizing ideal GP shape and verifying whether synthesized GP is correct are important but difficult problem. Some challenges in physical design and mask optimization are reviewed in this paper with possible solutions. Seongbo Shim, Youngsoo Shin |
VLSI-SoC | 2 |
| 2015 | Message from the technical program chairsabstractOn behalf of the Technical Program Committee of the 23rd IFIP/IEEE International Conference on Very Large Scale Integration (VLSI-SoC 2015), we welcome you to Daejeon, Korea and thank you for joining us at this important event. Youngsoo Shin, Chi-Ying Tsui |
VLSI-SoC | 1 |
| 2015 | Prosocial Activists in SNS: The Impact of Isomorphism and Social Presence on Prosocial BehaviorsabstractThe advent of information and communication technology has made people practice prosocial behavior in social networking services (SNSs) more easily. For this reason, the aim of the study was to identify the social and individual factors that induce prosociality in SNS. The concept of isomorphism for categorizing the characteristics of each social networks was adopted. The study also considered the concept of social presence for representing each individual. The experiment manipulated types of isomorphism (Mimetic, Normative, and Coercive) and degrees of social presence in an experimental SNS context. The study also measured individuals’ intention and activity of prosocial behavior. The experiment results indicate that mimetic and normative isomorphic conditions induce higher levels of prosocial intention and activity than coercive isomorphic condition. Also, a higher degree of social presence induces a higher level of prosocial intention. More interesting, the impact of mimetic condition is stronger when the social presence is higher. Youngsoo Shin, Bumho Lee, Jinwoo Kim 0001 |
Int. J. Hum. Comput. Interact. | 1 |
| 2014 | Lithographic defect aware placement using compact standard Cells without inter-cell marginabstractConventional standard cells contain extra space, called inter-cell margin, to prevent potential defects caused by lithography process. Margin is indeed necessary between some cell pairs, but there are also lots of cell pairs that do not yield any defects (or have very low probability of defects) when they are placed without margin. We address a new placement problem using standard cells without inter-cell margin. Placement should be done such that defect probability is made as small as possible while standard objectives such as wirelength is also pursued. The key in this approach is efficient computation of defect probabilities of all cell pairs and arranging them as a table that is referred to by a placer. We study how the cell pairs can be grouped by examining similar patterns along cell boundary, which greatly reduces the number of defect probability computation. The proposed placement method was evaluated on a few test circuits using 28-nm technology. Chip area was reduced by 10.8% on average with average and maximum defect probability kept below 0.4% and 4.1%, respectively. Seongbo Shim, Yoojong Lee, Youngsoo Shin |
ASP-DAC | 3 |
| 2014 | Power minimization of pipeline architecture through 1-cycle error correction and voltage scalingabstractWe present a new 1-cycle timing error correction method, which enables aggressive voltage scaling in a pipelined architecture. The proposed method differs from the state-of-the-art in that the pipeline stage where the timing error occurs can continue to receive input data without halting to avoid data collision. The feature allows the pipeline to avoid recurring clock gating when timing errors happen at multiple stages or timing errors continue to occur at a certain stage. Compared to a state-of-the-art method, the proposed method shows 2-6% energy reduction for a 5-stage pipeline and 7-11% reduction for a 10-stage pipeline. In addition, the proposed logic to propagate clock gating signal is much simpler than that of the previous method [1] by eliminating reverse propagation path of clock gating signal. Insup Shin, Jae-Joon Kim, Youngsoo Shin |
ASP-DAC | 3 |
| 2014 | Simplifying Clock Gating Logic by Matching Factored FormsabstractGate-level clock gating starts with a netlist, with partial or no gating applied; some flip-flops are then selected for further gating to reduce the circuit's power consumption, and a gating logic of the smallest possible size must then be synthesized. We show how to do this by factored form matching, in which gating functions in factored forms are matched, as far as possible, with factored forms of the Boolean functions of existing combinational nodes in the circuit; additional gates are then introduced, but only for the portion of gating functions that are not matched. Strong matching identifies matches that are explicitly present in the factored forms, and weak matching seeks matches that are implicit in the logic and thus are more difficult to discover. Factored form matching reduces gating logic by an average of 24%, over a few test circuits, for which Boolean division only achieves an average reduction of 8%. Inhak Han, Youngsoo Shin |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Pulsed-latch ASIC synthesis in industrial design flowabstractFlip-flop has long been used as a sequencing element of choice in ASIC design; commercial synthesis tools have also been developed in this context. This work has been motivated by a question of whether existing CAD tools can be employed from RTL to layout while pulsed latch replaces flip-flop as a sequencing element. Two important problems have been identified and their solutions are proposed: placement of pulse generators and latches for integrity of pulse shape, and design of special scan latches and their selective use to reduce hold violations. A reference design flow has also been set up using published documents, in order to assess the proposed one. In 40-nm technology, the proposed flow achieves 20% reduction in circuit area and 30% reduction in power consumption, on average of 12 test circuits. Duckhwan Kim 0001, Youngsoo Shin |
ASP-DAC | 3 |
| 2013 | Analysis and minimization of short-circuit current in mesh clock networkabstractMesh clock network is very effective at reducing clock skew. But mesh causes a large increase of power consumption, in particular due to shorted buffers. We first analyze the short-circuit power consumption of the mesh clock network. It is observed that skew distribution of premesh tree is important in determining the amount of short-circuit power. We then propose a new clock buffer, which practically eliminates short-circuit current in a mesh network. Experiments on a few test circuits using 40-nm technology indicate that clock power consumption is reduced by 13.0% on average with 4.8% of area increase; this can be compared to buffer sizing, which only achieves 5.6% saving of power. Seongbo Shim, Minyoung Mo, Youngsoo Shin |
ICCD | 4 |
| 2013 | A pipeline architecture with 1-cycle timing error correction for low voltage operationsabstractWe present a new timing error correction scheme which allows each pipeline stage to halt for one cycle only. The small timing penalty for the error correction operation in the proposed scheme makes it possible to eliminate the extra timing guardband that was needed to accommodate timing uncertainty due to process variations. As a result, lower supply voltage can be used with the proposed scheme for low power operations. Compared to the previous 1-cycle error correction scheme which uses two-phase transparent latch based pipeline [1], the proposed scheme can be applied to the pipeline based on more popular clocking elements such as flip-flop or pulsed latch. Insup Shin, Jae-Joon Kim, Yu-Shiang Lin, Youngsoo Shin |
ISLPED | 4 |
| 2012 | Introducing irregularity to routing architecture of structured ASIC for better routabilityabstractImposing regularity presents a fundamental limitation to any structured ASIC, or more generally any programmable logic device. It has been recently shown that irregularity can be introduced in structured ASIC, in particular in programmable logic elements of structured ASIC, through a special photolithography process, and the degree of irregularity can be customized for each particular design by manufacturing a few extra masks. We experiment how irregularity can be introduced to routing architecture of structured ASIC. When a whole routing area is made of an array of two routing architectures, the area is reduced by 8% to 16% (compared to standard structured ASIC) due to less white space, which is made possible by improved routability; the total wirelength is reduced by 6% to 14%. The new routing architectures and routing algorithm specific to the architectures are presented. Insup Shin, Donkyu Baek, Youngsoo Shin |
FPT | 3 |
| 2012 | Clock Gating Synthesis of Pulsed-Latch CircuitsabstractPulsed-latch circuits, in which latches are triggered by a short pulse, can reduce power consumption as well as increasing performance; and they can largely be designed using conventional computer-aided design tools. We explore the automatic synthesis of clock-gating logic for pulsed-latch circuits in which gating is implemented by enabling and disabling several pulse generators. The key problem is to arrange that each group of latches contains physically close latches, so that a short pulse from a pulse generator is delivered safely, and to ensure that the latches in a group have similar Boolean gating conditions because their clock is gated and ungated together. The resulting gating conditions should be implemented using as little extra logic as possible; for this purpose we rely on Boolean division, with an internal node of existing logic being used as the divisor. The proposed clock gating synthesis is assessed in 45-nm technology. Seungwhun Paik, Inhak Han, Youngsoo Shin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2012 | Synthesis of Active-Mode Power-Gating CircuitsabstractActive leakage is transient, which can be suppressed by design techniques such as dual-Vt. Active-mode power-gating (AMPG) can further reduce active leakage by power-gating groups of gates that perform computations with results that are not loaded due to clock-gating. AMPG involves several challenges; the grouping of gates must take circuit timing into account, and current switches need to be sized to preserve power network integrity as well as circuit timing. We propose solutions to these problems in the content of the entire process of synthesizing AMPG circuits. The physical design of AMPG circuits is also difficult due to the large number of virtual ground rails that must be mutually isolated. We address these issues by integrating placement with power network synthesis. Experiments on several test circuits implemented in 45-nm technology demonstrate the effectiveness of AMPG in the circuits that we synthesized, in terms of power consumption, area, wirelength, and timing. Jun Seomun, Insup Shin, Youngsoo Shin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2012 | Maximizing Frequency and Yield of Power-Constrained Designs Using Programmable Power-GatingabstractA large spread of leakage power due to process variations impacts the total power consumption of integrated circuits (ICs) substantially. This in turn may reduce frequency and/or yield of power-constrained designs. Facing such challenges, we propose two methods using power-gating (PG) devices whose effective width can be adjusted during a post-silicon tuning process. In the first method, we consider processors exhibiting substantial core-to-core frequency and leakage power variations while only a global voltage/frequency domain is supported. Since each core in a processor often has its own PG device, the total width each PG device and the global voltage are tuned jointly to maximize the global frequency for a given power constraint. Our experiment demonstrates that the maximum frequency of 2-, 4-, 8-, and 16-core processors is improved by 5%-21%. In the second method, we take rejected dies due to excessive leakage power. We adjust the width of PG devices such that the dies satisfy their given power constraint. Our experiment shows that 88%-98% of discarded dies violating their power constraint are recovered. Nam Sung Kim, Abhishek A. Sinkar, Jun Seomun, Youngsoo Shin |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2012 | HLS-dv: A High-Level Synthesis Framework for Dual-Vdd ArchitecturesabstractDual supply voltage design is widely accepted as an effective way to reduce the power consumption of CMOS circuits. In this paper, we propose a comprehensive design framework that includes dual- scheduling, dual- allocation, controller synthesis as well as layout generation. In particular, we address a problem of high-level synthesis with objective of minimizing power consumption of storage units and multiplexers using dual- ; this is made possible by utilizing timing slack that is left in the data-path after operation scheduling. We use integer linear programming (ILP) and also provide heuristic algorithms to solve the dual- register and connection allocation. The physical layout of dual-circuits has to separate power rails of and cells from each other. We propose a voltage island based placement algorithm to relieve this restriction and allow more flexibility of placement. In experiments on benchmark designs implemented in 1.08 V (with Vddlof 0.8 V) 65-nm CMOS technology, both switching and leakage power are reduced by 20% on average, respectively, compared to data-path with dual-Vddapplied to functional units alone. Detailed analysis of area and wirelength is performed to assess feasibility of the proposed method. Insup Shin, Seungwhun Paik, Dongwan Shin, Youngsoo Shin |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2011 | Selectively patterned masks: Structured ASIC with asymptotically ASIC performanceabstractStructured ASIC, which consists of a homogeneous array of tiles, suffers from large delay and area due to its inherent regularity. A new lithography method called selectively patterned masks (SPM) is proposed. It exploits special masks called masking masks and double exposure technique to allow more than one types of tiles to be patterned on the same wafer. The result is a heterogeneous array of tiles, which relaxes regularity in structured ASIC. A new structured ASIC based on SPM is proposed; tile and routing architectures, design flow, and tile packing and routing algorithm are all addressed. Experiments in 45-nm technology show that, compared to ASIC, the proposed structured ASIC exhibits 2.0 times of area when circuits are optimized for area and 1.2 times of delay when they are optimized for delay. Both figures represent substantial improvement over conventional structured ASIC. Donkyu Baek, Insup Shin, Seungwhun Paik, Youngsoo Shin |
ASP-DAC | 4 |
| 2011 | Pulser gating: A clock gating of pulsed-latch circuitsabstractA pulsed-latch is an ideal sequencing element for low-power ASIC designs due to its smaller capacitance and simple timing model. Clock gating of pulsed-latch circuits can be realized by gating a pulse generator (or pulser), which we call pulser gating. The problem of pulser gating synthesis is formulated for the first time. Given a gate-level netlist with location of latches, we first extract the gating function of each latch; the gating functions are merged to reduce the amount of extra logic while gating probability is not sacrificed too much. We also have to take account of proximity of latches, because a pulser, which is gated by merged gating function, and its latches have to be physically close for safe delivery of pulse. The heuristic algorithm that considers all three factors (similarity of gating functions, literal count to implement gating functions, and proximity of latches) is proposed and assessed in terms of power saving and area using 45-nm technology. Inhak Han, Seungwhun Paik, Youngsoo Shin |
ASP-DAC | 4 |
| 2011 | Thermal signature: a simple yet accurate thermal index for floorplan optimizationabstractA floorplanning has a potential to reduce chip temperature due to the conductive nature of heat. If floorplan optimization, which is usually based on simulated annealing, is employed to reduce temperature, its evaluation should be done extremely fast with high accuracy. A new thermal index, named thermal signature, is proposed. It approximates the temperature calculation, which is done by taking the product of Green's function and power density integrated over space. The correlation coefficient between thermal signature and temperature is shown to be quite high, more than 0.7 in many examples. A floorplanner that uses thermal signature is constructed and assessed using real design examples in 32-nm technology. It produces a floorplan whose maximum temperature is 11.4°C smaller than that of standard floorplan, on average, in reasonable amount of runtime. Inhak Han, Sachin S. Sapatnekar, Youngsoo Shin |
DAC | 4 |
| 2011 | Implementation of pulsed-latch and pulsed-register circuits to minimize clocking powerabstractA pulsed-latch can be modeled as a fast flip-flop. This allows conventional flip-flop designs to be migrated to pulsed-latch versions by simple replacement to reduce the clocking power. A key step in the migration process is to insert pulsers, which generate clock pulse to drive local latches; the number of pulsers as well as the wirelength of clock routing must be minimized to reduce the clocking power. We formulate a pulser insertion problem to find a set of latch groups where each group shares a pulser and its load constraint is satisfied; both an ILP formulation and a heuristic algorithm are presented to solve the problem. Experimental results of circuits implemented with 32-nm CMOS technology show that the clocking power of pulsed-latch designs obtained by our approach is 5.9% less than that of greedy approach; this is 44.7% less than that of flip-flop designs. We also consider the problem of pulsed-register where a pulser is integrated with multiple latches. A concept of logical distance is explored during our clustering algorithm to minimize the overhead of signal wirelength when converting flip-flops to pulsed-registers. Compared with flip-flop circuits, signal wirelength is increased by 6.3%, which is 1.4% smaller than without considering logical distance, while reducing the clocking power by 24%. Seungwhun Paik, Gi-Joon Nam, Youngsoo Shin |
ICCAD | 3 |
| 2011 | Pulsed-Latch Aware Placement for Timing-Integrity OptimizationabstractUtilizing pulsed-latches in circuit designs is one emerging solution to timing improvements. Pulsed-latches, driven by a brief clock signal generated from pulse generators, possess superior design parameters over flip-flops. If the pulse generator and pulsed-latches are not placed properly, however, pulse-width degradations at pulsed-latches and thus timing violations might occur. In this paper, we present a unified placement framework for pulsed-latches to maintain the timing integrity. Our new placer has the following distinguished features: 1) a multilevel analytical placement framework to effectively prevent the potential pulse-width distortion problem; 2) a physical-location aware pulse-generator insertion algorithm to identify each desired group of a pulse generator and latches; and 3) a new optimization gradient for global placement to consider the impact of load capacitance of generators. Experimental results show that our placement flow can effectively consider pulse-width integrity and thus achieve much smaller total/worst negative slacks with marginal wirelength overheads, compared to a leading commercial and an academic placement flows. Yi-Lin Chuang, Youngsoo Shin, Yao-Wen Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2011 | Retiming Pulsed-Latch Circuits With Regulating Pulse WidthabstractA pulsed-latch is an ideal sequencing element for high-performance application-specific integrated circuit designs due to its simple timing model and reduced sequencing overhead. The possibility of time-borrowing while a latch is transparent is deliberately ignored in pulsed-latch circuits to simplify the timing model. However, using more than one pulse width allows another form of time-borrowing, which preserves the simple timing model. The associated problem of allocating pulse widths is called pulse width allocation (PWA); and we combine it with retiming to achieve a shorter clock period in pulsed-latch circuits than can be obtained by retiming or PWA alone, with less requirement for extra latches than standard retiming. An exact solution can be obtained by an integer linear programming, but this is restricted to small circuits. We therefore introduce a practical heuristic, which performs clock skew scheduling to find the minimum clock period and then brings the clock skew as nearly back to zero as possible by converting it to combined retiming and PWA. Experiments with 45 nm technology circuits suggest that the heuristic algorithm achieves a clock period that is close to minimum in most circuits, with an average 23% reduction compared to initial circuit, at an average cost of a 16% increase in the number of latches. On the same benchmarks, standard retiming achieves a 20% reduction in clock period with a 29% increase in the number of latches. The cost in extra area and energy is reasonable. Seungwhun Paik, Seonggwan Lee, Youngsoo Shin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2011 | Design and Optimization of Power-Gated Circuits With Autonomous Data RetentionabstractPower gating has been widely employed to reduce subthreshold leakage. Data retention elements (flip-flops and isolation circuits) are used to preserve circuit states during standby mode, if the states are needed again after wake-up. These elements must be controlled by an external power management unit, causing a network of control signals implemented with extra wires and buffers. A power-gated circuit with autonomous data retention (APG) is proposed to remove the overhead involved in control signals. Retention elements in APG derive their control by detecting rising potential of virtual ground rails when power gating starts, i.e., they control themselves without explicit control signals. Design of retention elements for APG is addressed to facilitate safe capturing of circuit states. Experiments with 65-nm technology demonstrate that, compared to standard power gating, total wirelength, and average wiring congestion are reduced by 8.6% and 4.1% on average, respectively, at a cost of 6.8% area increase. In order to fast charge virtual ground rails, a pMOS switch driven by a short pulse is employed to directly provide charges to virtual ground. This helps retention elements avoid short-circuit current while making transition to standby mode. The optimization procedure for sizing pMOS switch and deciding pulse width is addressed, and assessed with 65-nm technology. Experiments show that, compared to standard power gating, APG reduces the delay to enter and exit the standby mode by 65.6% and 28.9%, respectively, with corresponding energy dissipation during the period cut by 46.1% and 36.5%. Standby mode leakage power consumption is also reduced by 15.8% on average. Jun Seomun, Youngsoo Shin |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | Statistical time borrowing for pulsed-latch circuit designsabstractPulsed-latch inherits the advantage of latch in less sequencing overhead while taking the advantage of flip-flop in its convenience during timing analysis. Even though this advantage comes from the fact that pulsed-latch uses a short pulse, it is still capable of a small amount of time borrowing. A problem of allocating pulse width (out of a few predefined widths), where each width is modeled by a random variable, is formulated for minimizing the clock period of pulsed-latch circuits; this is equivalent to assigning a random variable that represents the amount of time borrowed by the combinational block between each latch pair. A statistical approach is important in this problem because assuming +3¿ of all pulse widths does not represent the worst case. An allocation algorithm called SPWA as well as an algorithm to compute timing yield is proposed. In experiments with 45-nm technology, compared to the case of no time borrowing, the clock period was reduced by 12.2% and 11.7% on average when the yield constraint Ycis 0.85 and 0.95, respectively; this is compared to the deterministic counterpart called DPWA, which reduced the clock period by 7.6% and 7.3%. More importantly, DPWA failed to satisfy the yield constraints in four (out of eleven) circuits while the yield constraints were always satisfied in SPWA. Seungwhun Paik, Lee-eun Yu, Youngsoo Shin |
ASP-DAC | 3 |
| 2010 | Bounded potential slack: enabling time budgeting for dual-Vt allocation of hierarchical designabstractTime budgeting, which assigns timing assertion at block boundary, is a crucial step in hierarchical design. The proportion of high- and low-Vtgates of each block, which determines overall leakage power consumption, is dictated by timing assertion, yet dual-Vtallocation is not taken into account during conventional time budgeting. Bounded potential slack is introduced as a measure of dual-Vtallocation, and is experimentally shown to be strongly correlated with the percentage of high-Vtgates. A new time budgeting is proposed with objective of achieving bounded potential slack, which is formulated as a linear programming problem. In experiments with example hierarchical designs implemented in 45-nm commercial technology, the proposed time budgeting reduced leakage power by 32% on average compared to conventional time budgeting, when both are followed by the same dual-Vtallocation. The time budgeting is also applied to voltage island design, where each block can have its own Vddwith mix of high- and low-Vtgates. Jun Seomun, Seungwhun Paik, Youngsoo Shin |
ASP-DAC | 3 |
| 2010 | Pulsed-latch aware placement for timing-integrity optimizationabstractUtilizing pulsed latches in a circuit is one emerging solution to timing improvements. Pulsed latches, driven by a brief clock signal generated from pulse generators, possess superior design parameters over flip-flops. If pulse generators and pulsed latches are not placed properly, however, pulse-width degradations at pulsed latches and thus timing violations might occur. In this paper, we introduce the pulsed-latch aware placement problem for timing integrity and present a unified placement framework to tackle this problem. Our new placer has the following distinguished features: (1) a multilevel pulsed-latch aware analytical placement framework to effectively prevent the potential pulse-width distortion problem, (2) a physical-information aware latch grouping algorithm to identify each desired group of a pulse generator and pulsed latches, and (3) a new optimization gradient for global placement to consider the impact of load capacitance of generators. Experimental results show that our placement flow can effectively consider pulse-width integrity and thus achieve much smaller total/worst negative slacks with marginal wirelength overheads, compared to a leading commercial and an academic placement flows. Yi-Lin Chuang, Youngsoo Shin, Yao-Wen Chang |
DAC | 3 |
| 2010 | Synthesis and implementation of active mode power gating circuitsabstractActive leakage current is much larger (~ 10x) than standby leakage current, and takes a large proportion (30% to 40%) of active power consumption. Active mode power gating (AMPG) has been proposed to extend the application of basic power gating to reducing active leakage; it relies on clock-gating signals to cut the power off a part of combinational gates. The problem to select those gates while integrity of circuit behavior remains intact has not been solved yet. We identify three constraints to solve this problem, namely functional, timing, and current constraints. The problem of synthesizing AMPG circuits is then laid out, and synthesis algorithm is proposed; a group of gates that can be power-gated by each clock-gating signal and the size of footer that is attached to the group constitute a synthesis output. The layout methodology for standard cell designs is proposed to assess AMPG circuits in area and wirelength. Experiments in 1.1 V, 45-nm technology demonstrate that active leakage is reduced by 16% on average compared to clock-gated circuits. Jun Seomun, Insup Shin, Youngsoo Shin |
DAC | 3 |
| 2010 | Wakeup synthesis and its buffered tree construction for power gating circuit designsabstractA power gating circuit suffers from large amount of rush current when it wakes up, especially when all switch cells are turned on at the same time. If each switch cell is turned on in different instant of time, the rush current can be reduced. It is shown in this paper that the rush current can be reduced even more if signal transition time (or signal slew) to each switch cell is adjusted. The wakeup synthesis that we define is to determine the turn-on time and signal slew of each switch cell; the goal is to minimize wakeup delay while rush current is kept below the maximum value that is allowed. The corresponding synthesis algorithm is proposed. The determined turn-on time and signal slew are implemented using a buffered tree, where a source is a wakeup signal and sinks are multiple switch cells; the synthesis algorithm to generate the tree is proposed. The wakeup synthesis and buffered tree construction are integrated into a design flow that receives a netlist of power gating circuit as an input and produces a layout of netlist with wakeup network embedded. Experiments in an industrial 1.1 V, 45-nm technology demonstrate that the wakeup delay is reduced by 43% on average of example circuits compared with 2-pass turn-on, which is widely used. Seungwhun Paik, Youngsoo Shin |
ISLPED | 3 |
| 2010 | Pulse Width Allocation and Clock Skew Scheduling: Optimizing Sequential Circuits Based on Pulsed LatchesabstractPulsed latches, latches driven by a brief clock pulse, offer the same convenience of timing verification and optimization as flip-flop-based circuits, while retaining the advantages of latches over flip-flops. But a pulsed latch that uses a single pulse width has a lower bound on its clock period, limiting its capacity to deal with higher frequencies or operate at lowerVdd. The limitation still exists even when clock skew scheduling is employed, since the amount of skew that can be assigned and realized is practically limited due to process variation. For the first time, we formulate the problem of allocating pulse widths, out of a small discrete number of predefined widths, and scheduling clock skews, within a predefined upper bound on skew, for optimizing pulsed latch-based sequential circuits. We then present an algorithm calledPWCS_Optimize(pulse width allocation and clock skew scheduling, PWCS) to solve the problem. The allocated skews are realized through synthesis of local clock trees between pulse generators and latches, and a global clock tree between a clock source and pulse generators. Experiments with 65-nm technology demonstrate that combining a small number of different pulse widths with clock skews of up to 10% of the clock period yield the minimum achievable clock period for many benchmark circuits. The results have an average figure of merit of 0.86, where 1.0 indicates a minimum clock period, and the average reduction in area by 11%. The design flow includingPWCS_Optimize, placement and routing, and synthesis of local and global clock trees is presented and assessed with example circuits. Hyein Lee 0003, Seungwhun Paik, Youngsoo Shin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2010 | HLS-l: A High-Level Synthesis Framework for Latch-Based ArchitecturesabstractLevel-sensitive latches are widely used in high-performance custom designs while edge-triggered flip-flops are predominantly used in application-specific integrated circuits. We consider a latch as a basis for storage and address each step of high-level synthesis (HLS), including scheduling, allocation, and control synthesis. While the use of latches provides an opportunity to reduce the latency during the scheduling, the register allocation has to take extra conflicts caused by latch into account, and the control synthesis has to be tailored to support the latch-based data-path. Optimization potentials specific to this HLS are identified and solutions are proposed. Specifically, the register allocation can be improved by refining the operation schedule in a way to reduce the number of edges in a register conflict graph; the latency can be reduced by adjusting the clock duty cycle in a way to generate a tighter schedule. All the steps of HLS and optimization procedures were integrated into a framework called HLS-l. It was tested on benchmark designs implemented in 1.1-V, 45 nm complementary metal-oxide-semiconductor technology. Compared to the conventional HLS, HLS-l was able to reduce the latency by 18.2% on average with 9.2% less area and 16.0% less power consumption. The application of HLS-l to an industrial example is demonstrated through the design of a module extracted from H.264/advanced video coding. Seungwhun Paik, Insup Shin, Youngsoo Shin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2010 | Power gating: Circuits, design methodologies, and best practice for standard-cell VLSI designsabstractPower Gating has become one of the most widely used circuit design techniques for reducing leakage current. Its concept is very simple, but its application to standard-cell VLSI designs involves many careful considerations. The great complexity of designing a power-gated circuit originates from the side effects of inserting current switches, which have to be resolved by a combination of extra circuitry and customized tools and methodologies. In this tutorial we survey these design considerations and look at the best practice within industry and academia. Topics include output isolation and data retention, current switch design and sizing, and physical design issues such as power networks, increases in area and wirelength, and power grid analysis. Designers can benefit from this tutorial by obtaining a better understanding of implications of power gating during an early stage of VLSI designs. We also review the ways in which power gating has been improved. These include reducing the sizes of switches, cutting transition delays, applying power gating to smaller blocks of circuitry, and reducing the energy dissipated in mode transitions. Power Gating has also been combined with other circuit techniques, and these hybrids are also reviewed. Important open problems are identified as a stimulus to research. Youngsoo Shin, Jun Seomun, Kyu-Myung Choi, Takayasu Sakurai |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2010 | Supply Switching With Ground Collapse for Low-Leakage Register Files in 65-nm CMOSabstractPower-gating has been widely used to reduce subthreshold leakage current. However, the extent of leakage saving through power-gating diminishes with technology scaling due to gate leakage of data-retention circuit elements. Furthermore, power-gating involves substantial increase of area and wirelength. A circuit technique called supply switching with ground collapse (SSGC) has recently been proposed to overcome the limitation of power-gating. The circuit technique is successfully applied to the register file of ARM9 microprocessor in a 1.2 V, 65-nm CMOS process, and the measured result is reported for the first time. The leakage current is reduced by a factor of 960 on average of 83 dies at 25°C , and by a factor of 150 at 85°C. Compared to a register file implemented in conventional power-gating, leakage current is cut by a factor of 2.2, demonstrating that SSGC can be a substitute for power-gating in nanometer CMOS. Hyung-Ock Kim, Bong Hyun Lee, Jong-Tae Kim, Jung Yun Choi, Kyu-Myung Choi, Youngsoo Shin |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2009 | Register allocation for high-level synthesis using dual supply voltagesabstractReducing the power consumption of memory elements is known to be the most influential in minimizing total power consumption, since designs tend to use more memories these days. In this paper, we address a problem of high-level synthesis with the objective of minimizing power consumption of storage using dual-Vdd. Specifically, we propose a complete design framework that starts from dual-Vdd scheduling, dual-Vdd allocation, and controller synthesis down to the final layout. Its main feature is dual-Vdd register allocation, which exploits timing slacks left in the data-path after operation scheduling. In experiments on benchmark designs implemented in 1.08 V (with Vddl of 0.8 V), 65-nm CMOS technology, both switching and leakage power were reduced by 20% on average, respectively, compared to data-path with dual-Vdd applied to functional units alone. Detailed analysis of slack histogram, area, wirelength, and congestion were performed to assess feasibility of the design framework. Insup Shin, Seungwhun Paik, Youngsoo Shin |
DAC | 3 |
| 2009 | HLS-l: High-level synthesis of high performance latch-based circuitsabstractAn inherent performance gap between custom designs and ASICs is one of the reasons why many designers still start their designs from register transfer level (RTL) description rather than from behavioral description, which can be synthesized to RTL via high-level synthesis (HLS). Sequencing overhead is one of the factors for this performance gap; the choice between latch and flip-flop is not typically taken into account during HLS, even though it affects all the steps of HLS. HLS-l is a new design framework that employs high-performance latches during scheduling, allocation, and controller synthesis. Its main feature is a new scheduler that is based on a concept of phase step (as opposed to conventional control step), which allows us scheduling in finer granularity, register allocation that resolves the conflict of latch being read and written at the same time, and controller synthesis that exploits dual-edge triggered storage elements to support phase step based scheduling. In experiments on benchmark designs implemented in 1.2 V, 65 -nm CMOS technology, HLS-l reduced latency by 16.6% on average, with 9.5% less circuit area, compared to the designs produced by conventional HLS. Seungwhun Paik, Insup Shin, Youngsoo Shin |
DATE | 3 |
| 2009 | Retiming and time borrowing: Optimizing high-performance pulsed-latch-based circuitsabstractPulsed-latches take advantage of both latches in their high performance and flip-flops in their convenience of timing analysis. To minimize the clock period of pulsed-latch-based circuits for a higher performance, a problem of combined retiming and time borrowing is formulated, where the latter is enabled by using a handful of different pulse widths. The problem is first approached by formulating it as an integer linear programming to lay a theoretical foundation. A heuristic approach is proposed, which solves the problem by performing clock skew scheduling for the minimum clock period and gradually converting skew into a combination of retiming and time borrowing. Experiments with 45-nm technology demonstrate that the clock period close to the minimum can be achieved for all benchmark circuits with an average of 1.03x with less use of extra latches compared to the conventional retiming. Seonggwan Lee, Seungwhun Paik, Youngsoo Shin |
ICCAD | 3 |
| 2009 | Frequency and yield optimization using power gates in power-constrained designsabstractManufactured dies exhibit a large spread of maximum frequency and leakage power due to process variations, which have been increasing with technology scaling. Reducing the spread is very important for maximizing the frequency and the yield of power-constrained designs, because otherwise many dies that do not satisfy frequency or power constraints would be discarded. In this paper, we propose two optimization methods to improve the maximum operating frequency and the yield using power gates that already exist in many power-constrained designs. In the first method, we consider the designs of multiple cores, where each of them can be independently power-gated. When each core shows different frequencies due to within-die variations, the strength of a power gate in each core is adjusted to make their maximum operating frequencies even. This allows faster cores to consume less active leakage power, reducing the total power consumption well below a power constraint in a globally-clocked design. We subsequently increase global supply voltage for higher overall frequency until the power constraint is satisfied. In our experiments assuming multicore processors with 2--16 cores, the maximum operating frequency was improved by 4-23%. In the second method, we take leaky-but-fast dies (which otherwise would be discarded) and adjust the strength of the power gates such that they can operate in an acceptable power and frequency region. The problem is extended to designs employing a frequency binning strategy, where we have an additional objective of maximizing the number of dies for higher frequency bins. In our experiments with ISCAS benchmark circuits, most discarded fast-but leaky dies were recovered using the second method. Nam Sung Kim, Jun Seomun, Abhishek A. Sinkar, Jungseob Lee, Tae Hee Han, Ken Choi, Youngsoo Shin |
ISLPED | 7 |
| 2009 | ssr HLShbox-ssr pg: High-Level Synthesis of Power-Gated CircuitsabstractA problem inherent in power-gated circuits is the overhead of state-retention storage required to preserve the circuit state in standby mode.HLS-pgis a new design framework that takes power gating into account, from scheduling, allocation, and controller synthesis to the final circuit layout. Its main feature is a new scheduler that minimizes the number of retention registers required at the power-gating control step. In experiments on benchmark designs implemented in 0.9-V 65-nm technology,HLS-pgreduced leakage current by 20.7% on average, with 5.0% less area and 4.1% less wirelength, compared to the power-gated circuits produced by conventional high-level synthesis. Eunjoo Choi, Changsik Shin, Youngsoo Shin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2009 | Semicustom Design of Zigzag Power-Gated Circuits in Standard Cell ElementsabstractZigzag power gating (ZPG) can overcome the long wake-up delay of standard power gating, but its requirement for both nMOS and pMOS current switches, in a zigzag pattern, requires complicated power networks, limiting application to custom designs. We propose a design framework for cell-based semicustom design of ZPG circuits, using a new power network architecture that allows the unmodified conventional logic cells to be combined with custom circuitry such as ZPG flip-flops, input forcing circuits, and current switches. The design flow, from register transfer level description to layout, is described and applied to a 32-b microprocessor design using a 1.2-V 65-nm triple-well bulk CMOS process. The use of a sleep vector in ZPG requires additional switching power when entering standby mode and returning to active mode. The switching power should be minimized so that is does not outweigh the leakage saved by employing ZPG scheme. We formulate the selection of a sleep vector as a multiobjective optimization problem, minimizing both the transition energy and the total wirelength of a design. We solve the problem by employing multiobjective genetic-based algorithm. Experimental results show an average saving of 39% in transition energy and 8% in total wirelength for several benchmark circuits in 65-nm technology. Youngsoo Shin, Seungwhun Paik, Hyung-Ock Kim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2009 | Minimizing leakage power of sequential circuits through mixed-Vt flip-flops and multi-Vt combinational gatesabstractThe current use of multi- V t to control leakage power targets combinational gates, even though sequential elements such as flip-flops and latches also contribute appreciable leakage. We can, nevertheless, apply multi- V t to flip-flops, but few can take advantage of high- V t , which causes abrupt changes in timing. We combine low- and high- V t at the transistor level to design mixed- V t flip-flops with reduced leakage, an unchanged footprint, and a small increase in either setup time or clock-to-Q delay, but not both. An allocation algorithm for two V t s determines the V t (mixed, high, or low) of each flip-flop and the V t of each combinational gate (high or low) in a sequential circuit. Experiments with 65-nm technology show an average leakage saving of 42% compared to conventional multi- V t approaches; the leakage of flip-flops alone is cut by 78%. This saving is largely unaffected by die-to-die or within-die process variations, which we show through simulations. Standard deviation of leakage caused by process variation is also reduced due to less use of low- V t devices. We also extend our approach to three V t s, and obtain a further 14% reduction in leakage. Chungki Oh, Youngsoo Shin |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2008 | Statistical mixed Vt allocation of body-biased circuits for reduced leakage variationabstractLeakage current is susceptible to variation of transistor parameters and environment such as temperature, which results in wide spread in leakage distribution. The spread can be reduced by employing body biasing: reverse body bias for too leaky dies and forward body bias for too slow dies. We investigate body biasing of mixed Vtcircuits. It is shown that the conventional body biasing has limitation in reducing leakage variation of mixed Vtcircuits. This is because low- and high-Vtdevices do not track each other and their body biasing sensitivities are different. We present alternative body biasing scheme that targets compensating die-to-die variation of low Vt. Under this body biasing scheme, within-die profiles of low- and high-Vt, which we need for statistical allocation of mixed Vt, get wider thus become different from the original ones. We present an analytical procedure to derive new within-die profiles. Experiments with 45-nm predictive model show that the spread in leakage distribution (ratio of maximum and minimum leakage) can be reduced to 4.5 as opposed to 9.4 from conventional body biasing on mixed Vtcircuits. Jinseob Jeong, Seungwhun Paik, Youngsoo Shin |
ASP-DAC | 3 |
| 2008 | Multiobjective optimization of sleep vector for zigzag power-gated circuits in standard cell elementsabstractZigzag power gating (ZPG) has been proposed to alleviate the drawback of power gating in its long wake-up delay, thereby broadening the application of power gating to suppressing active- as well as standby-leakage. However, complicated power network due to the use of nMOS and pMOS switches in zigzag fashion has limited its application to custom circuits. Heterogeneous use of power rails inevitably incurs overhead of area and wirelength during physical design. Furthermore, the use of sleep vector causes additional switching power when entering standby mode and returning back to active mode. The switching power should be minimized not to outweigh the leakage saving by employing ZPG scheme. In this paper, we propose a complete power network architecture, which allows us to use unmodified standard cell elements for implementing ZPG circuits. We formulate selecting sleep vector as a multi-objective optimization problem, minimizing transition energy and total wirelength. We solve the problem by employing multiobjective genetic-based algorithm. Experimental results show the saving of 39% in transition energy and 8% in wirelength, on average, for several benchmark circuits in 65-nm technology. The complete design flow starting from RTL description down to layout is proposed, and assessed with 65-nm technology. Seungwhun Paik, Youngsoo Shin |
DAC | 2 |
| 2008 | Pulse width allocation with clock skew scheduling for optimizing pulsed latch-based sequential circuitsabstractPulsed latches, latches driven by a brief clock pulse, offer the convenience of flip-flop-like timing verification and optimization, while retaining superior design parameters of latches over flip-flops. But, pulsed latch-based design using a single pulse width has a limitation in reducing clock period. The limitation still exists even if clock skew scheduling is employed, since the amount of skew that can be assigned is practically limited due to process variations. The problem of allocating pulse width (out of discrete number of predefined widths) and scheduling clock skew (within prescribed upper bound) is formulated, for the first time, for optimizing pulsed latch-based sequential circuits. An allocation algorithm called PWCS_Optimize is proposed to solve the problem. Experiments with 65-nm technology demonstrate that small number of variety of pulse widths (up to 5) combined with clock skews (up to 10% of clock period) yield minimum clock period for many benchmark circuits. The design flow including PWCS_Optimize, placement and routing, and synthesis of local and global clock trees is presented and assessed with example circuits. Hyein Lee 0003, Seungwhun Paik, Youngsoo Shin |
ICCAD | 3 |
| 2008 | 3-D thermal simulation with dynamic power profilesabstractOn-chip temperature and temperature gradient have been emerging as important design criteria as technology is scaled down to nano-meter regime. There have been several approaches to analyze or simulate the thermal behavior of chips, but all the approaches assume constant average power consumption of each block, which is reasonable when the change in power is localized and transient. However, as the aggressive power management techniques are employed in block level of granularity, power consumption of blocks become fluctuating a lot, which yields a large error with the conventional thermal analysis. A 3-D thermal simulation, with time-varying power consumption of blocks, is proposed in this paper. The partial differential heat conduction equation is solved with finite difference method, and we also employ alternating direction implicit method to decrease the computational complexity. The prototype simulator was designed and tested on several examples. Eunjoo Choi, Youngsoo Shin |
ISCAS | 2 |
| 2008 | Power-gating-aware high-level synthesisabstractA problem inherent in designing power-gated circuits is the overhead of the state-retention storage required to preserve the circuit state in standby mode. Reducing the amount of retention storage is known to be the most influential factor in minimizing the loss of the benefit (i.e. power saving) by power-gating. In this paper, we address a new problem of high-level synthesis with the objective of minimizing the size of retention storage to be used in the power-gated circuits. Specifically, we propose a complete design framework, called HLS-pg, that starts from the power-gating-aware scheduling, allocation, and controller synthesis down to the final circuit layout. The key contribution of the work is to solve the power-gating-aware scheduling problem, namely, scheduling operations that minimizes the number of retention registers required at the power-gating control step, while satisfying resource and latency constraints. In experiments on benchmark designs implemented in 65-nm CMOS technology, HLS-pg generates circuits with 27% less leakage current, with 6% less circuit area and wirelength, compared to the power-gated circuits produced by conventional highlevel synthesis. Eunjoo Choi, Changsik Shin, Youngsoo Shin |
ISLPED | 4 |
| 2008 | Skewed Flip-Flop and Mixed-Vt Gates for Minimizing Leakage in Sequential CircuitsabstractMixedVthas been widely used to control leakage without affecting circuit performance. However, existing approaches only target combinational circuits, even though sequential elements such as flip-flops contribute an appreciable proportion of the total leakage. Applying highVtto ordinary flip-flops would reduce the number of combinational gates that can be assigned to highVt, because any timing slacks would be absorbed by the increased setup guard time and propagation delay of the high-Vtflip-flops. A skewed flip-flop (SFF) can be constructed by replacing a subset of transistors in a conventional flip-flop with low-leakage devices, such as large-Lgatetransistors. In terms of leakage and delay, SFFs exhibit very skewed characteristic, which depends on the transistors that are replaced. Our algorithm selectively substitutes SFFs for conventional flip-flops in sequential circuits so as to reduce the leakage while continuing to satisfy the timing constraint. When combined with the mixed-Vtcombinational circuits, this achieves an average leakage saving of 15% compared to mixedVtalone. The leakage of the flip-flops themselves is cut by 25% on average. Jun Seomun, Youngsoo Shin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2007 | Simultaneous Control of Subthreshold and Gate Leakage Current in Nanometer-Scale CMOS CircuitsabstractPower gating has been widely used to reduce subthreshold leakage. However, its efficiency degrades very fast with technology scaling due to the gate leakage of circuits specific to power gating, such as storage elements and output interface circuits with a data-retention capability. A new scheme called supply switching with ground collapse is proposed to control both gate and subthreshold leakage in nanometer-scale CMOS circuits. Compared to power gating, the leakage is cut by a factor of 6.3 with 65nm and 8.6 with 45nm technology. Various issues in implementing the proposed scheme using standard-cell elements are addressed, from RTL to layout. The proposed design flow is demonstrated on a commercial design with 90nm technology, and the leakage saving by a factor of 32 is observed with 3% and 6% of increase in area and wirelength, respectively. Youngsoo Shin, Sewan Heo, Hyung-Ock Kim, Jung Yun Choi |
ASP-DAC | 1 |
| 2007 | Skewed Flip-Flop Transformation for Minimizing Leakage in Sequential CircuitsabstractMixed Vt has been widely used to control leakage without affecting circuit performance. However, current approaches target the combinational circuits even though sequential elements, such as flip-flops, contribute an appreciable proportion of the total leakage. A skewed flip-flop (SFF) is obtained by slightly increasing the gate length of a subset of the transistors in a conventional flip-flop. The resulting SFF will exhibit very skewed characteristics in terms of leakage and delay, which depend on the transistors that are replaced. We present an algorithm that selectively substitutes SFFs for conventional flip-flops in sequential circuits, such that the timing constraint is still satisfied while the leakage from the flip-flops is reduced. When combined with the mixed Vt technique, an average leakage saving of 16% is achieved, compared to the use of mixed Vt alone. Jun Seomun, Youngsoo Shin |
DAC | 3 |
| 2007 | Minimizing leakage power in sequential circuits by using mixed Vt flip-flopsabstractDual Vthas been widely used to control leakage, while, at the same time, satisfying circuit performance. However, current approaches target the combinational circuits even though sequential elements, such as flip-flops and latches, contribute an appreciable proportion of the total leakage. The use of dual Vtflip-flops is limited to circuits of large timing slack, because introducing high Vtflip-flops in place of low Vtones yields abrupt change in timing. We propose mixed Vtflip-flops, which are designed by using both low and high Vt, but in different transistors. Compared to low Vtflip-flop, the mixed Vtflip-flops exhibit increased delay, but either on setup time or on clock-to- Q delay but not on both, while their leakage is greatly reduced. We extend the conventional sensitivity-based dual Vt allocation algorithm to incorporate mixed Vtflip-flops together with dual Vtcombinational gates. Experimental results show that an average leakage saving of 31% is achieved, compared to the use of dual Vton combinational subcircuits alone. The leakage of the flip-flops themselves is cut by 57% on average. Youngsoo Shin |
ICCAD | 2 |
| 2007 | Supply Switching With Ground Collapse: Simultaneous Control of Subthreshold and Gate Leakage Current in Nanometer-Scale CMOS CircuitsabstractPower gating has been widely used to reduce subthreshold leakage. However, the efficiency of power gating degrades very fast with technology scaling, which we demonstrate by experiment. This is due to the gate leakage of circuits specific to power gating, such as storage elements and output interface circuits with a data-retention capability. A new scheme called supply switching with ground collapse is proposed to control both gate and subthreshold leakage in nanometer-scale CMOS circuits. Compared to power gating, the leakage is cut by a factor of 6.3 with 65-nm and 8.6 with 45-nm technology. Various issues in implementing the proposed scheme using standard-cell elements are addressed, from register transfer level to layout. These include the choice of standby supply voltage with circuits that support it, a power network architecture for designs based on standard-cell elements, a current switch design methodology, several circuit elements specific to the proposed scheme, and the design flow that encompasses all the components. The proposed design flow is demonstrated on a commercial design with 90-nm technology, and the leakage saving by a factor of 32 is observed with 3% and 6% of increase in area and wirelength, respectively. Youngsoo Shin, Sewan Heo, Hyung-Ock Kim, Jung Yun Choi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2006 | Analysis and optimization of gate leakage current of power gating circuitsabstractPower gating is widely accepted as an efficient way to suppress subthreshold leakage current. Yet, it suffers from gate leakage current, which grows very fast with scaling down of gate oxide. We try to understand the sources of leakage current in power gating circuits and show that input MOSFETs plays a crucial role in determining total gate leakage current. It is also shown that the choice of a current switch in terms of polarity, threshold voltage, and size has a significant impact on total leakage current. From the observation of the importance of input MOSFETs, we propose the power optimization of power gating circuits through input control. Hyung-Ock Kim, Youngsoo Shin |
ASP-DAC | 2 |
| 2006 | Physical design methodology of power gating circuits for standard-cell-based designabstractThe application of power gating circuits to semicustom design based on standard-cell elements is limited due to the requirement of customizing cells that are tailored for power gating or the requirement of customizing physical design methodologies for placement and power network. We propose a new power network architecture that enables use of conventional standard-cell elements. A few custom library elements are developed wherever needed, including output interface circuits and data retention storage elements. A novel method of current switch design is also described. The proposed methodology is applied to ISCAS benchmark circuits, and also to a commercial Viterbi decoder with 0.18/spl mu/m CMOS technology. Hyung-Ock Kim, Youngsoo Shin, Hyuk Kim, Iksoo Eo |
DAC | 2 |
| 2005 | μITRON-LP: power-conscious real-time OS based on cooperative voltage scaling for multimedia applicationsabstractThis work presents a cooperative dynamic power management method and its implementation. The implementation consists of design of a real-time OS, applications including MPEG-4, and development of a supporting hardware platform with an off-the-shelf processor. We describe several factors that are important in the implementation and discuss its efficiency through experiment. The experimental results with the prototype system shows that 74% power saving is possible in multi-task multimedia environment. Hiroshi Kawaguchi 0001, Youngsoo Shin, Takayasu Sakurai |
IEEE Trans. Multim. | 2 |
| 2004 | Architecting voltage islands in core-based system-on-a-chip designsabstractVoltage islands enable core-level power optimization for System-on-Chip (SoC) designs by utilizing a unique supply voltage for each core. Architecting voltage islands involves island partition creation, voltage level assignment and floorplanning. The task of island partition creation and level assignment have to be done simultaneously in a floorplanning context due to the physical constraints involved in the design process. This leads to a floorplanning problem formulation that is very different from the traditional floorplanning for ASIC-style design.In this paper, we define the problem of architecting voltage islands in core-based designs and present a new algorithm for simultaneous voltage island partitioning, voltage level assignment and physical-level floorplanning. Application of the proposed algorithm to a few benchmark and industrial examples is demonstrated using a prototype tool. Results show power savings of 14%--28%, depending on the constraints imposed on the number of voltage islands and other physical-level parameters. Jingcao Hu, Youngsoo Shin, Nagu R. Dhanwada, Radu Marculescu |
ISLPED | 2 |
| 2002 | Power distribution analysis of VLSI interconnects using model orderreductionabstractThe analysis and simulation of effects induced by interconnects become increasingly important as the scale of process technologies steadily shrinks. While most analyses focus on the timing aspects of interconnects, power consumption is also important. In this paper, the power distribution analysis of interconnects is studied using a reduced-order model. The relation between power consumption and the poles and residues of a transfer function (either exact or approximated) is derived, and a simple yet accurate driver model is developed, allowing power consumption to be computed efficiently. Application of the proposed method to RC networks is demonstrated using a prototype tool. Youngsoo Shin, Takayasu Sakurai |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2001 | Coupling-Driven Bus Design for Low-Power Application-Specific SystemsabstractIn modern embedded systems including communication and multimedia applications, large fraction of power is consumed during memory access and data transfer. Thus, buses should be designed and optimized to consume reasonable power while delivering sufficient performance. In this paper, we address a bus ordering problem for low-power application-specific systems. A heuristic algorithm is proposed to determine the order in a way that effective lateral component of capacitance is reduced, thereby reducing the power consumed by buses. Experimental results for various examples indicate that the average power saving from 30% to 46.7% depending on capacitance components can be obtained without any circuit overhead. Youngsoo Shin, Takayasu Sakurai |
DAC | 1 |
| 2001 | Estimation of power distribution in VLSI interconnectsabstractThe analysis and simulation of effects induced by VLSI interconnects become increasingly important as the scale of process technologies steadily shrinks. While most analyses focus on the timing aspects of interconnects, power consumption is also important. In this paper, the power distribution estimation of interconnects is studied using a reduced-order model. The relation between power consumption and the poles and residues of a transfer function is derived, and an appropriate driver model is developed, allowing power consumption to be computed efficiently. Application of the proposed method to RC networks is demonstrated using a prototype tool. 1. Youngsoo Shin, Takayasu Sakurai |
ISLPED | 1 |
| 2001 | Partial bus-invert coding for power optimization of application-specific systemsabstractThis paper presents two bus coding schemes for power optimization of application-specific systems: partial pus-invert coding and its extension to multiway partial bus-invert coding. In the first scheme, only a selected subgroup of bus lines is encoded to avoid unnecessary inversion of relatively inactive and/or uncorrelated bus lines which are not included in the subgroup. In the extended scheme, we partition a bus into multiple subbuses by clustering highly correlated bus lines and then encode each subbus independently. We describe a heuristic algorithm of partitioning a bus into subbuses for each encoding scheme. Experimental results for various examples indicate that both encoding schemes are highly efficient for application-specific systems. Youngsoo Shin, Soo-Ik Chae, Kiyoung Choi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2001 | Narrow bus encoding for low-power DSP systemsabstractHigh levels of integration in integrated circuits often lead to the problem of running out of pins. Narrow data buses can be used to alleviate this problem provided that the degraded performance due to wait cycles can be tolerated. We address bus coding methods for low-power core-based systems incorporating narrow buses. We show that transition signaling combined with bus-invert coding, which we call BITS coding, is particularly suitable for the data patterns of typical DSP applications on narrow data buses. The application of BITS coding to real circuit design is limited by the extra bus line introduced, which changes the pinout of the chip. We propose a new coding method, which does not require the extra bus line but retains the advantage of BITS. Youngsoo Shin, Kiyoung Choi, Young-Hoon Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2000 | Narrow bus encoding for low power systemsabstractHigh integration in integrated circuits often leads to the problem of running out of pins.Narrow data buses can be used to alleviate this problem at the cost of performance degradation due to wait cycles.In this paper, we address bus coding methods for low power core-based systems incorporating narrow buses.Although the conventional Bus-Invert code performs well for completely random patterns, we show that transition signaling combined with Bus-Invert, which we call BITS coding, can achieve much more power saving for data patterns of typical DSP applications.The application of BITS coding to a real circuit design is limited by the overhead of the encoder and decoder circuits and the extra bus line introduced.We propose an approximate version of BITS coding, which do not require the extra bus line while retaining the advantage of BITS coding. Youngsoo Shin, Kiyoung Choi |
ASP-DAC | 1 |
| 2000 | Schedulability-driven performance analysis of multiple mode embedded real-time systemsabstractProviding multiple modes to support dynamically changing environments, standards, and new services is prevalent in embedded systems, especially in mobile radio systems. Because such a system frequently contains time-constrained tasks, it is important to analyze the temporal requirements as well as the functional correctness. This paper presents a method to analyze temporal requirements imposed on an embedded real-time system supporting multiple modes. While most performance analysis methods focus only on testing the feasibility of a task or a system, our method goes further by addressing the problem of locating hot spots of a system thereby helping the designer to choose among alternative designs or architectures. We formally define the analysis problem and show that it is very unlikely to be solved efficiently. We present a heuristic algorithm, which is accurate and fast enough to be used in iterative processes in system-level analysis and design. The analysis problem is extended to accommodate probabilistic behavior exhibited by soft real-time tasks. Youngsoo Shin, Daehong Kim, Kiyoung Choi |
DAC | 1 |
| 2000 | Power Optimization of Real-Time Embedded Systems on Variable Speed ProcessorsabstractPower efficient design of real-time embedded systems based on programmable processors becomes more important as system functionality is increasingly realized through software. This paper presents a power optimization method for real-time embedded applications on a variable speed processor. The method combines off-line and on-line components. The off-line component determines the lowest possible maximum processor speed while guaranteeing deadlines of all tasks. The on-line component dynamically varies the processor speed or brings a processor into a power-down mode according to the status of task set in order to exploit execution time variations and idle intervals. Experimental results show that the proposed method obtains a significant power reduction across several kinds of applications. Youngsoo Shin, Kiyoung Choi, Takayasu Sakurai |
ICCAD | 1 |
| 1999 | Power Conscious Fixed Priority Scheduling for Hard Real-Time SystemsabstractPower efficient design of real-time systems based on programmable processors becomes more important as system functionality is increasingly realized through software. This paper presents a powerefficient version of a widely used fixed priority scheduling method. The method yields a power reduction by exploiting slack times, both those inherent in the system schedule and those arising from variations of execution times. The proposed run-time mechanism is simple enough to be implemented in most kernels. Experimental results show that the proposed scheduling method obtains a significant power reduction across several kinds of applications. 1 Introduction Recently, power consumption has been a critical design constraint in the design of digital systems due to widely used portable systems such as cellular phones and PDAs, which require low power consumption with high speed and complex functionality. The design of such systems often involves reprogrammable processors such as microprocessors... Youngsoo Shin, Kiyoung Choi |
DAC | 1 |
| 1998 | Partial bus-invert coding for power optimization of system level busabstractWe presen t a partial bus-in vertcoding scheme for po wer optim ization of system level bus. In the proposed sch eme, we select a su b-group of bus lines involved in b us encoding to a void unnecessary inversion of b us lines not in the sub-group thereby redu cing th e total number of bus transitions. We propose a heuristic algorithm that selects the sub-grou p of bus lines for b us encoding. Ex periments on benchmark examples in dicate that the partial bus-in vert coding reduces the tot al bus tran sitions b y 62.6% on the av erage, compared to that of the unencoded patterns. Youngsoo Shin, Soo-Ik Chae, Kiyoung Choi |
ISLPED | 1 |
| 1996 | Software synthesis through task decomposition by dependency analysisabstractLatency tolerance is one of main problems of software synthesis in the design of hardware-software mixed systems. This paper presents a methodology for speeding up systems through latency tolerance which is obtained by decomposition of tasks and generation of an efficient scheduler. The task decomposition process focuses on the dependency analysis of system i/o operations. Scheduling of the decomposed tasks is performed in a mixed static and dynamic fashion. Experimental results show the significance of our approach. Youngsoo Shin, Kiyoung Choi |
ICCAD | 1 |
| 1996 | An integrated hardware-software cosimulation environment with automated interface generationabstractWe present a hardware-software cosimulation environment for heterogeneous systems. To be an efficient system verification environment for the rapid prototyping of heterogeneous systems, the environment provides following features: interface transparency, smooth transition to cosynthesis, simulation acceleration, and integrated user interface and internal representation. Among them, the first two are more important than the others. To support these two features, we have developed automatic interface generation schemes. As demonstrating experiments, two heterogeneous systems performing same function with different target architectures were cosimulated and prototyped successfully in our environment. The experimental results show that our environment can be a useful heterogeneous system specification/verification environment for rapid prototyping. Kyuseok Kim, Youngsoo Shin, Kiyoung Choi |
RSP | 3 |
| 1995 | An integrated hardware-software cosimulation environment for heterogeneous systems prototypingabstractNo abstract available. Kyuseok Kim, Youngsoo Shin, Taekyoon Ahn, Wonyong Sung, Kiyoung Choi, Soonhoi Ha |
ASP-DAC | 3 |
| 1995 | Efficient Prototyping System Based on Incremental Design and Module-by-Module VerificationabstractThis paper presents an efficient hardware prototyping methodologies of digital systems. A is based on a low-cost and flexible prototyping system which consists of a general-purpose CPU and a FPGA-based custom board. Using our prototyping methodologies such as incremental system design and module-by-module verification, we can map partial system specification into hardware prototype, which is implemented by programming FPGAs on custom board. This allows flexible and efficient system verification as well as reduction in cost and time of prototype building. Youngsoo Shin, Kyuseok Kim, Jae-Hee Won, Kiyoung Choi |
ISCAS | 2 |