VLDB 2026 Research / reviewers in the wild / expert
Seokhyeong Kang
dblp:12/8003
· DBLP profile ↗
96ranked-venue papers
1as first author
63since 2021 · last 2026
0000-0003-3015-1806ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 94 · 1 first-author · 61 since 2021Software engineering, systems software and programming languages · 23 · 18 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Heterogeneous Graph-based Gate Sizer Integrating Graph Attention Network and TransformerabstractGate sizing is a critical step in achieving the target power, performance, and area (PPA) in chip design. In recent years, machine learning (ML) methods have recently emerged as a new paradigm for gate sizing. Their promising results have gained attention; however, the practical applicability and performance of existing works is limited by at least one of the following factors: (1) long runtime due to test-time optimization or autoregressive prediction; (2) limited exploration of architectural choices; (3) an overly simplified data representation, known as a homogeneous graph, which merges pins and cells into a single node. To improve both practicality and performance, we introduce a novel MLbased gate sizer, dubbed DPH-Sizer, which directly predicts the appropriate gate sizes using a heterogeneous graph that separates cells and their pins into distinct nodes. This heterogeneous graph explicitly captures the relationships between different circuit elements, leading to enhanced performance. Lastly, we propose InterCell and Intra-Cell GAT blocks to explicitly capture both intracell and inter-cell information. These are followed by transformer blocks, which are placed at the end of the GAT stack to capture global path-level features. In our experiments, we validate each of the proposed components and demonstrate that DPH-Sizer maintains power consumption within 2.0% on average while achieving improvements of 54.3% in timing (WNS) and 1.3% in area metrics. Jinmo Ahn, Jinoh Cho, Jaemin Seo, Jakang Lee, Seokhyeong Kang |
ASP-DAC | 6 |
| 2026 | Au-MEDAL: Adaptable Grid Router with Metal Edge Detection And Layer Integration
Andrew B. Kahng, Seokhyeong Kang, Jakang Lee, Dooseok Yoon |
ASP-DAC | 2 |
| 2026 | REvolution: An Evolutionary Framework for RTL Generation driven by Large Language ModelsabstractLarge Language Models (LLMs) are used for Register-Transfer Level (RTL) code generation, but they face two main challenges: functional correctness and Power, Performance, and Area (PPA) optimization. Iterative, feedbackbased methods partially address these, but they are limited to local search, hindering the discovery of a global optimum. This paper introduces REvolution, a framework that combines Evolutionary Computation (EC) with LLMs for automatic RTL generation and optimization. REvolution evolves a population of candidates in parallel, each defined by a design strategy, RTL implementation, and evaluation feedback. The framework includes a dual-population algorithm that divides candidates into Fail and Success groups for bug fixing and PPA optimization, respectively. An adaptive mechanism further improves search efficiency by dynamically adjusting the selection probability of each prompt strategy according to its success rate. Experiments on the VerilogEval and RTLLM benchmarks show that REvolution increased the initial pass rate of various LLMs by up to 24.0 percentage points. The DeepSeek-V3 model achieved a final pass rate of 95.5%, comparable to state-of-the-art results, without the need for separate training or domain-specific tools. Additionally, the generated RTL designs showed significant PPA improvements over reference designs. This work introduces a new RTL design approach by combining LLMs’ generative capabilities with EC’s broad search power, overcoming the local-search limitations of previous methods. Kyungjun Min, Kyumin Cho, Junhwan Jang, Seokhyeong Kang |
ASP-DAC | 4 |
| 2026 | TANGRAM: A Novel ILP-based On-Track Bus Routing via Placement and Compression of PolygonsabstractBus routing is an advanced topic of signal routing. Unlikely to the classical routing, the bus routing problem has complex constraints such as topology consistency and channel compactness. Existing bus routing algorithms mostly rely on the iterative maze routing, which is heavily time-consuming and sensitive to net ordering, thereby easily succumbing to suboptimality. To overcome this limitation, we propose a novel bus routing algorithm using placement and compression of routing pattern polygons. Critically, our method does not rely on maze routing, thus highly fast and effective. Experimental results show that the proposed method achieves an average of 1.8% quality improvement over the best known results of ICCAD 2018 contest benchmarks. Jaekyung Im, Seokhyeong Kang |
DATE | 2 |
| 2026 | Efficient Down-sampling in Hybrid Neural Networks using Adversarial AutoencodersabstractEarly convolutional layers in hybrid neural networks enable efficient down-sampling but pose a significant burden on inference latency and energy consumption. We propose a method to replace the conventional down-sampling block with lightweight autoencoders to enhance applicability in edge devices. We introduce an adversarial training strategy to align the autoencoder’s latent features with the original stem, ensuring compatibility with succeeding layers. Applying our method to MobileViTV2-050 yields a 1.23x speedup and a 47% reduction in Energy-Delay Product (EDP) with only a 1.0% accuracy drop on ImageNet-1K. Jonghyeon Nam, JoonSeok Kim, Eunji Kwon, Seokhyeong Kang |
DATE | 4 |
| 2026 | Timing- and Power-Aware Differentiable Repair of Minimum Implant Area Violations
Jinoh Cho, Minhyuk Kweon, Jinmo Ahn, Jeyeong Park, Seokhyeong Kang |
ISLPED | 5 |
| 2026 | Fusing the Attention Training Dataflow with Local Safe Softmax and Model-Independent Tiling
JoonSeok Kim, Daeheon Lee, DongHwan Yoon, Jonghyeon Nam, Seokhyeong Kang |
ISLPED | 5 |
| 2026 | Invited: Post-Placement Buffering and Sizing ContestabstractThe ISPD 2026 Contest [22] challenges participants to develop post-detailed placement buffering and sizing tools that optimize timing and fix electrical rule check (ERC) violations under real-world constraints. Unlike prior contests, this contest emphasizes practical physical design challenges including fixed macros and I/Os, power delivery network (PDN) blockages, soft placement blockages, and fixed routing resources. The contest provides eight public benchmarks and four hidden benchmarks, with a range from 15K to 1.4M instances, in the ASAP7 7nm technology node [4] with multi-threshold voltage cell libraries. Evaluation is performed using the open-source OpenROAD infrastructure, with scoring based on timing (total negative slack), power (dynamic and leakage) and penalties for ERC violations, displacement, routing congestion and runtime. This paper describes the contest problem formulation, benchmarks, evaluation methodology, a review of related contests and a two-year roadmap for continuation in the ISPD 2027 Contest. Andrew B. Kahng, Seokhyeong Kang, Sayak Kundu, Yiting Liu 0002, Davit Markarian, Seonghyeon Park, Zhiang Wang |
ISPD | 2 |
| 2026 | CTRL-B: Back-End-of-Line Configuration Optimization Using Cross-Domain Transferable Reinforcement Learning
Sung-Yun Lee, Jinoh Cho, Daijoon Hyun, Seokhyeong Kang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | Leveraging Machine Learning Techniques for Traditional EDA Workflow EnhancementabstractAs technology nodes advance and feature sizes shrink, the increasing complexity of design rules and routing congestion has resulted in greater design challenges and rising costs. Machine learning (ML) models offer significant potential to enhance design quality by enabling early prediction and optimization during the design flow. However, only a few works have validated the effectiveness of ML model when integrated to the traditional design flow. This paper will cover the effectiveness of ML-enhanced design workflow with some practical applications. Additionally, we will address which problems should be solved to achieve successful ML integration. Jinoh Cho, Jaekyung Im, Kyungjun Min, Seonghyeon Park, Jaemin Seo, Jongho Yoon 0001, Seokhyeong Kang |
ASP-DAC | 8 |
| 2025 | PC-Opt: Partition and Conquest-based Optimizer using Multi-Agents for Complex Analog CircuitsabstractRecent research in electronic design automation (EDA) tools has focused on utilizing artificial intelligence (AI) for sizing analog circuit designs. Still, there has been a lack of focus on optimizing complex analog circuits. To optimize complex analog circuits within a few circuit simulations, we propose a partition-and-conquest-based optimizer (PC-Opt). PC-Opt assigns distinct actor-critic roles within a multi-agent system, facilitating the partitioning of complex analog circuits and conquering their optimization challenges. Partial differential training is developed for the proper prediction of each actor, which merges each other and then predicts the optimized entire circuit. To generate a compact and non-biased dataset for network training, a concentrated sampling method is devised. Experimental results on three circuits demonstrate the effectiveness of PC-Opt. Youngchang Choi, Sejin Park 0001, Ho-Jin Lee, Kyongsu Lee, Jae-Yoon Sim, Seokhyeong Kang |
ASP-DAC | 6 |
| 2025 | LIBMixer: An all-MLP Architecture for Cell Library Characterization towards Design Space OptimizationabstractCell library characterization is a fundamental stage of electronic design automation (EDA), as it provides essential electrical models for circuit simulation and design quality assessment. However, the development of advanced nodes demands increasing computational resources and engineering efforts for characterization. We introduce LIBMixer, a machine learning-based framework for fast and accurate library characterization designed to enhance design space optimization. Leveraging multi-layer perceptron architectures, LIBMixer efficiently manages complex relationships between technology and electrical characteristics. It achieves a 31.5× faster runtime than conventional EDA tools while enhancing alignment. Compared to state-of-the-art methods, LIBMixer targets 6.4× more standard cells for both power and timing information. This scalability improvement enables the practical synthesis of IP cores, demonstrating high correlation across of power, performance, and area results. Additionally, Pareto fronts of synthesis design results with LIBMixer-inferred libraries closely match those from foundry files. Experimental results highlight the effectiveness of LIBMixer as a fast and reliable alternative for PVT analysis. Sunggyu Jang, Jakang Lee, Seokhyeong Kang |
ASP-DAC | 4 |
| 2025 | ParaFormer: A Hybrid Graph Neural Network and Transformer Approach for Pre-Routing Parasitic RC PredictionabstractPredicting the quality of post-route design at an early stage can reduce overall design time. To achieve this, we propose ParaFormer, a pre-routing parasitic RC prediction framework. This framework integrates a heterogeneous graph neural network (HGNN) and a graph transformer to capture the topological and geometric information of circuit data. The HGNN model represents circuit data as heterogeneous graphs to learn complex topological relationships, while the graph transformer calculates attention between each net to learn geometric relationships. Our framework predicts parasitic RC, enabling RC tree modeling and SPEF file generation. This allows the predicted results to be utilized in timing and power analysis using commercial tools. Additionally, we incorporate gradient normalization to reduce the imbalance between different objectives in multi-task learning, improving overall model performance. Experimental results show that ParaFormer achieves R2 scores of 0.9901 and 0.9630 for resistance and capacitance, respectively. In timing analysis, it achieves R2 scores of 0.9749 for wire delay and 0.9876 for cell delay, with a MAPE of 1.45% in power analysis. These results indicate that our method is highly effective for timing and power prediction in the early design stage. Jongho Yoon 0001, Jakang Lee, Junseok Hur, Seokhyeong Kang |
ASP-DAC | 5 |
| 2025 | FedEDA: Federated Learning Framework for Privacy-Preserving Machine Learning in EDAabstractIn advanced nodes, the optimization of Power, Performance, and Area (PPA) is becoming increasingly complex, requiring significant resources and time for circuit design and optimization using Electronic Design Automation (EDA). As a key approach to overcome these challenges, Machine Learning (ML) techniques have been widely studied in the field of EDA. However, security concerns around Intellectual Property (IP) limit access to real-world circuit data, making it difficult to gather sufficient data for training ML models. This lack of available circuit benchmarks restricts progress in ML research. In this study, we propose FedEDA, which, to the best of our knowledge, is the first Federated Learning (FL) aggregation algorithm specifically designed for EDA. FedEDA addresses concerns about IP security by exchanging model weights among FL participants instead of sharing raw data. Furthermore, FedEDA leverages Rent’s Rule and circuit size to capture the hierarchical structure of circuits, mitigating issues related to data imbalance among participants and improving the quality of weight aggregation on EDA data. We demonstrate the applicability of FedEDA across various EDA tasks, including routability, parasitic RC, and wirelength prediction. FedEDA outperforms existing FL algorithms in EDA tasks, demonstrating superior performance. Seonghyeon Park, Seokhyeong Kang |
DAC | 4 |
| 2025 | Late Breaking Results: Fine-Tuning LLMs for Test Stimuli GenerationabstractThe understanding and reasoning capabilities of large language models (LLMs) with text data have made them widely used for test stimuli generation. Existing studies have primarily focused on methods such as prompt engineering or providing feedback to the LLMs’ generated outputs to improve test stimuli generation. However, these approaches have not been successful in enhancing the LLMs’ domain-specific performance in generating test stimuli. In this paper, we introduce a framework for finetuning LLMs for test stimuli generation through dataset generation and reinforcement learning (RL). Our dataset generation approach creates a table-shaped test stimuli dataset, which helps ensure that the LLM produces consistent outputs. Additionally, our two-stage fine-tuning process involves training the LLMs on domain-specific data and using RL to provide feedback on the generated outputs, further enhancing the LLMs’ performance in test stimuli generation. Experimental results confirm that our framework improves syntax correctness and code coverage of test stimuli, outperforming commercial models. Hyeonwoo Park, Seonghyeon Park, Seokhyeong Kang |
DAC | 3 |
| 2025 | Late Breaking Results: A Geometric Diffusion Model for Macro Placement GenerationabstractMacro placement is crucial in VLSI design, directly impacting circuit performance. We introduce MacroDiff, a diffusion-based macro placement generative model that captures wirelength relationships instead of directly predicting macro coordinates. By leveraging wirelength as an intermediate representation, MacroDiff naturally preserves circuit connectivity, reduces placement constraints, and enhances solution flexibility while inherently handling rotational and translational invariances. Experiments on ISPD2005 benchmarks show that MacroDiff reduces macro overlap by 91.6%, lowers macro legalization displacement by 74.4%, and improves half-perimeter wirelength (HPWL) by 7.0%. While maintaining the efficiency of generative approaches, MacroDiff generates high-quality placements more reliably, narrowing the gap with state-of-the-art methods. The source code for this work is available at https://github.com/jhy00n/MacroDiff. Jongho Yoon 0001, Jinsung Jeon, Seokhyeong Kang |
DAC | 3 |
| 2025 | Neural Circuit Parameter Prediction for Efficient Quantum Data LoadingabstractQuantum machine learning (QML) has demonstrated the potential to outperform classical machine learning algorithms in various fields. However, encoding classical data into quantum states, known as quantum data loading, remains a challenge. Existing methods achieve high accuracy in loading single data, but lack efficiency for large-scale data loading tasks. In this work, we propose Neural Circuit Parameter Prediction, a novel method that leverages classical deep neural networks to predict the parameters of parameterized quantum circuits directly from the input data. This approach benefits from the batch inference capability of neural networks and improves the accuracy of quantum data loading. We introduce real-valued parameterization of quantum circuits and a three-phase training strategy to further enhance training efficiency and accuracy. Experimental results on MNIST dataset show that our method achieves a 17.31 % improvement in infidelity score and 108 times faster runtime compared to existing methods. Our approach provides an efficient solution for quantum data loading, enabling the practical deployment of QML algorithms on large-scale datasets. Sunghye Park, Seokhyeong Kang |
DATE | 3 |
| 2025 | Improving LLM-Based Verilog Code Generation with Data Augmentation and RLabstractLarge language models (LLMs) have recently attracted significant attention for their potential in Verilog code generation. However, existing LLM-based methods face several challenges, including data scarcity and the high computational cost of generating prompts for fine-tuning. Motivated by these challenges, we explore methods to augment training datasets, develop more efficient and effective prompts for fine-tuning, and implement training methods incorporating electronic design automation (EDA) tools. Our proposed framework for fine-tuning LLMs for Verilog code generation includes (1) abstract syntax tree (AST)-based data augmentation, (2) output-relevant code masking, a prompt generation method based on the logical structure of Verilog code, and (3) reinforcement learning with tool feedback (RLTF), a fine-tuning method using EDA tool results. Experimental studies confirm that our framework significantly improves syntax and functional correctness, outperforming commercial and non-commercial models on open-source benchmarks. Kyungjun Min, Seonghyeon Park, Hyeonwoo Park, Jinoh Cho, Seokhyeong Kang |
DATE | 5 |
| 2025 | SO3-Cell: Standard Cell Layout Automation Framework for Simultaneous Optimization of Topology, Placement, and RoutingabstractWe propose SO3-Cell, the first automatic standard cell layout generation framework that optimizes three key steps simultaneously using Mixed-Integer Linear Programming (MILP). SO3-Cell simultaneously performs circuit topology optimization, transistor placement, and internal cell routing to achieve an optimized layout solution. Our optimization objective is to minimize metal usage while enhancing cell layout flexibility within a given area.We introduce design space pruning techniques to mitigate the complexity of larger designs, such as a full adder, a reset flip-flop (FF), and a 2-bit FF. We successfully generate a layout for a 44-transistor 2-bit FF within 25,862 seconds, demonstrating the scalability and robustness of the SO3-Cell framework. We evaluate the block-level PPA impact of the proposed cell-layout improvements, demonstrating a 35.0% reduction in power, a 2.2% increase in frequency, and a 31.1% reduction in area. Chung-Kuan Cheng, Andrew B. Kahng, Byeonggon Kang, Seokhyeong Kang, Jakang Lee, Bill Lin 0001 |
ICCAD | 4 |
| 2025 | A Parallel Analytical Legalization Algorithm via Alternating Direction Method of MultipliersabstractLegalization tries to resolve the cell overlaps and align every cells to the placement sites while honoring the global placement results. The existing legalization works are mostly relying on heuristical cell-by-cell search within a window and this makes large suboptimality in their algorithm. Only a few works have attempted to solve the legalization problem through analytical method, but they also suffered huge runtime overhead and suboptimality due to the discrete nature of legalization. In this paper, we revisit the classical legalization problem and propose a new parallel analytical legalization method based on alternating direction method of multipliers (ADMM) with heterogeneous CPU-GPU parallelism. Experimental results show that our method significantly improves the solution quality compared to existing open-source legalizers. Jaekyung Im, Seokhyeong Kang |
ICCAD | 2 |
| 2025 | Enhancing Timing Closure via Spatially Embedded Graph Transformer with Low Power/Area OverheadabstractAs technology scales and operating frequencies increase, achieving timing closure in digital circuits has become significantly challenging. Traditional post-placement optimization often fails to account for routing variations, leading to persistent timing violations and lengthy design iterations. To address this issue, we introduce a Spatially Embedded Graph Transformer (SEGT) for accurate timing prediction at the placement stage, considering the effects of subsequent processes. By incorporating distance-aware self-attention and a positional encoding scheme tailored to the geometric structure of circuit designs, SEGT effectively captures complex interactions between non-adjacency nodes and the positional context of circuits. Compared to previous timing prediction networks, SEGT improves the R2score from 0.915 to 0.959 in post-routing stage delay prediction and from 0.857 to 0.938 in post-routing path slack prediction. Leveraging accurate timing predictions of SEGT, a preemptive timing closure framework is developed to handle routing-induced violations proactively before post-placement optimization. Our framework identifies post-routing timing violation paths with slacks and adjusts the required arrival time by precisely the amount needed to clear each violation, thereby reducing additional design overhead. Experimental results of the proposed timing closure enhancement framework with SEGT predictions showed a substantial reduction in total negative slack across test designs while also lowering power and area overhead compared to previous works, demonstrating significant timing improvement and superior design efficiency. By enhancing the accuracy of timing prediction, this work enables a proactive and efficient timing closure approach, bridging the gap between placement and routing stages in physical design. Joonyoung Seo, Jonghyeon Nam, Howoo Jang, Yoonseok Jung, Seokhyeong Kang |
ICCAD | 5 |
| 2025 | Diffusion-Enhanced Graph Transformer with Reinforcement Learning for Transferable Analog Circuit OptimizerabstractWe propose a Diffusion-Enhanced Graph Transformer (DEGT) for analog circuit optimization that overcomes the limitations of traditional vector- and graph-based approaches. Conventional methods struggle to capture the complex connectivity of analog circuits and often require expert-imposed heuristic constraints on the sizing of some transistors to greatly reduce the searching space. In contrast, our method introduces three key innovations. First, an enhanced graph representation combined with a transformer architecture conveys circuit information to the machine learning network without any loss, enabling effective incremental knowledge transfer across various circuit designs. Second, the proposed DEGT quantifies the influence of each device by considering connection distances and path configurations, thereby providing a comprehensive, topology-aware representation of device interactions. Third, a violation handling method autonomously trains non-functional regions in the design space, eliminating the need for expert-imposed constraints or circuit classifications. Experimental evaluations demonstrate that the proposed optimizer consistently improves the figure of merit for a circuit with each round of incremental knowledge transfer using data from different circuits. These results highlight the potential of our approach to advance autonomous analog circuit design by reducing the reliance on expert intervention and improving overall optimization performance. Ho-Jin Lee, Kyeong-Jun Lee, Jae-Hoon Lee, Kyu-Jin Choi, Geunyong Choi, Youngchang Choi, Kyongsu Lee, Seokhyeong Kang, Jae-Yoon Sim |
ISLPED | 8 |
| 2025 | Invited: Artificial Netlist Generation for Enhanced Circuit Data AugmentationabstractOptimizing power, performance, and area (PPA) at advanced nodes has become an increasingly challenging and complex task. To address these challenges, approaches such as machine learning (ML) and design-technology co-optimization (DTCO) have emerged as promising solutions. However, their effectiveness is limited by the lack of diverse training data and prolonged turnaround times (TAT). Artificial data has been widely used in various fields to address the limitations of real-world data. By augmenting datasets, artificial data improve the robustness of ML models against input perturbations, leading to improved performance. Similarly, in the physical design flow, artificial data has great potential for overcoming the scarcity of real-world circuit data [1], [2], [3]. Artificial circuits proposed in previous studies are typically designed for specific applications. By developing a method to generate artificial circuit which resemble real circuits, we can address data scarcity and TAT challenges in physical design. In this talk, we will discuss how leveraging artificial circuits to explore a wide range of circuit characteristics can enhance ML model performance for unseen real-world circuits and accelerate the PPA exploration flow. Seokhyeong Kang |
ISPD | 1 |
| 2024 | SkyPlace: A New Mixed-size Placement Framework using Modularity-based Clustering and SDP RelaxationabstractElectrostatics-based placement has made a great success and inspired many placement algorithms. However, the recent direction of improvement is missing two important problems for mixed-size placement - 1) how to initialize placement and 2) how to handle large macros in the analytical placement. In this paper, we propose our new mixed-size placer, SkyPlace which is enhanced by novel placement initialization using macro-aware clustering and semidefinite programming. Experimental results show that SkyPlace clearly outperforms the leading-edge placer on academic benchmarks. Jaekyung Im, Seokhyeong Kang |
DAC | 2 |
| 2024 | PPA-Relevant Clustering-Driven Placement for Large-Scale VLSI DesignsabstractToday's place-and-route (P&R) flows are increasingly challenged by complexity and scale of modern designs. Often, heuristics must trade off between turnaround time and quality of PPA outcomes. This paper presents a clustered placement methodology that improves both turnaround time and final-routed solution quality. Our PPA-aware clustering considers timing, power and logical hierarchy during netlist clustering, effectively reducing problem size and accelerating global placement runtime while improving post-route PPA metrics. Additionally, our machine learning (ML)-accelerated virtualized P&R methodology predicts the best cluster shapes (i.e., aspect ratios and utilizations) to use in P&R of the clustered netlist. With the open-source OpenROAD tool, our methods achieve up to 47% (average: 36%) global placement runtime improvement with similar half-perimeter wirelength (HPWL) and 90% (29%) improvement in post-route total negative slack (TNS). With the commercial Cadence Innovus tool, our methods achieve up to 3.92% (1%) improvement in power and 99% (49%) improvement in TNS. Andrew B. Kahng, Seokhyeong Kang, Sayak Kundu, Kyungjun Min, Seonghyeon Park, Bodhisatta Pramanik |
DAC | 2 |
| 2024 | RL-PTQ: RL-based Mixed Precision Quantization for Hybrid Vision TransformersabstractExisting quantization approaches incur significant accuracy loss when compressing hybrid convolution and transformer models with low bit-width. This paper presents RL-PTQ, a novel post-training quantization (PTQ) framework utilizing reinforcement learning (RL). Our focus is on determining the most effective bit-width and observer for quantization configurations tailored for mixed precision by grouping layers and addressing the challenges of quantization of hybrid transformers. We achieved the highest quantized accuracy for MobileViTs compared to the previous PTQ methods [5--7]. Furthermore, our quantized model on Processing In Memory (PIM) architecture exhibited an energy efficiency enhancement of 10.1× and 22.6× compared to the baseline model, on the state-of-the-art PIM accelerator [15] and GPU, respectively. Eunji Kwon, Minxuan Zhou, Tajana Rosing, Seokhyeong Kang |
DAC | 5 |
| 2024 | HiLight: A Comprehensive Framework for High-Performance and Lightweight Scalability in Surface Code CommunicationabstractIn pursuing fault-tolerant quantum computing (FTQC), the surface code (SC) serves as a key quantum error correction protocol. The double-defect mode of the SC enables long-range two-qubit communication via braiding. However, intersecting braiding paths create communication bottlenecks, leading to increased circuit latency. Sunghye Park, Seokhyeong Kang |
DAC | 3 |
| 2024 | Improvement of Mixed Track - Height Standard-Cell PlacementabstractIn sub-Snm nodes, track-height of standard cells must be aggressively scaled down while preserving design PPA. This requirement brings the challenge of placing a set of cells that have mixed track-heights, subject to the constraint that cells with the same height must be placed together in an “island” of cell rows. We apply integer linear programming (ILP) to solve the row assignment problem and improve the runtime of ILP with clustering, using a cost function that combines half-perimeter wirelength and displacement from a starting unconstrained placement. Considering the row assignment solution, we define fence-regions, which enable an existing place-and-route (P&R) tool to place the cells while considering the row-island constraints. Experimental results show that our proposed method can on average reduce final-routed wirelength by 8.5% and total power by 3.3 %, with worst negative slack and total negative slack reductions of 24.0% and 13.0%, compared with the previous state-of-the-art method [10]. Andrew B. Kahng, Seokhyeong Kang, Minhyuk Kweon |
DATE | 2 |
| 2024 | ViT- ToGo: Vision Transformer Accelerator with Grouped Token PruningabstractVision Transformer ($V$iT) has gained prominence for its performance in various vision tasks but comes with considerable computational and memory demands, posing a challenge when deploying it on resource-constrained edge devices. To address this limitation, various token pruning methods have been proposed to reduce the computation. However, the majority of token pruning techniques do not account for practical use in actual embedded devices, which demand a significant reduction in computational load. In this paper, we introduce ViT-ToGo, a$V$iT accelerator with grouped token pruning. This enables the parallel execution of the$V$iT models and the token pruning process. We implement grouped token pruning with a head-wise importance estimator which simplifies the process need for token pruning, including sorting and reordering. Our proposed method achieves up to 66 % reduction in the number of tokens, resulting in up to 36% reduction in GFLOPs, with only a minimal accuracy drop of around 1 %. Furthermore, the hardware implementation incurs a marginal resource overhead of 1.13% in average. Seungju Lee, Kyumin Cho, Eunji Kwon, Sejin Park 0001, Seojeong Kim, Seokhyeong Kang |
DATE | 6 |
| 2024 | Trans-Net: Knowledge-Transferring Analog Circuit Optimizer with a Netlist-Based Circuit RepresentationabstractFinding an optimal point in the design space of analog circuits requires a substantial time-consuming effort even for skillful circuit designers. There have been extensive studies on automated sizing of transistors in analog circuits based on machine learning (ML) algorithms. However, the previous approaches suffer from lack of expandability and necessitate an inevitable retraining process of the given model to apply for optimization of different circuits. The graph-based representation of a circuit with reinforcement learning (RL) achieved a knowledge transfer when optimizing the same circuit with different process technologies. However, it can be hardly applied to different circuit topologies due to the failure of generalizing the training of RL agent. This paper introduces Trans-Net, an analog circuit optimizer that is capable of supporting the knowledge transfer across different circuits as well as different process technologies with a circuit representation that defines the circuit topology by one-to-one mapping from SPICE netlist. The proposed analog circuit optimizer successfully supports multiple circuits within a single ML model, showcasing its effectiveness on five different circuit topologies across three different process technologies. Ho-Jin Lee, Kyeong-Jun Lee, Youngchang Choi, Kyongsu Lee, Seokhyeong Kang, Jae-Yoon Sim |
DATE | 5 |
| 2024 | CTRL-B: Back-End-Of-Line Configuration Pathfinding Using Cross-Technology Transferable Reinforcement LearningabstractIn advanced technology nodes, the impact of the back-end-of-line (BEOL) on chip performance and power consumption becomes progressively significant. This paper presents a BEOL configuration pathfinding framework using the proposed cross-technology transferable reinforcement learning (CTRL) model. We optimize the BEOL technology parameters, including the metal stack, geometry, and design rules, to enhance the power efficiency and performance resulting from the physical design process. First, we extract various design and technology features and embed them into metal-type-wise deep neural networks. Our multi-policy model selects a BEOL parameter configuration that is estimated to improve chip power efficiency and performance. We employ a policy gradient algorithm complemented by various training strategies, such as data normalization, action-reward buffer, and network optimization, to expedite training convergence. Furthermore, we transfer the pathfinding knowledge from the trained model in the old node to the new model for efficient BEOL configuration pathfinding in the advanced node, referred to as cross-technology transfer learning. The proposed framework achieved 19 % reduced total power consumption, 68 % improved worst negative slack, and 87 % improved total negative slack with a high reward efficiency, on average, in several designs. Further, we demonstrate that the reward efficiency of our proposed CTRL model outperforms that of the conventional fine-tuning method in transferring knowledge by 24 %. Sung-Yun Lee, Kyungjun Min, Seokhyeong Kang |
DATE | 3 |
| 2024 | Unveiling the Black-Box: Leveraging Explainable AI for FPGA Design Space OptimizationabstractWith significant advancements in various design methodologies, modern integrated circuits have experienced noteworthy improvements in power, performance, and area. Among various methodologies, design space optimization (DSO), which automatically explores electronic design automation (EDA) tool parameters for a given design, has been extensively studied in recent years. In this study, we propose an approach to fine-tuning an effective FPGA design space to suit a specific design. By utilizing our ML-based prediction and explainable artificial intelligence (XAI) approach, we quantify parameter contribution scores, which reveal the correlation between each parameter and the final timing results. Using the valuable insights from the parameter contribution scores, we can refine the design space only with effective parameters for subsequent timing optimization. During the optimization, our framework improved the maximum operating frequency by 26% on average in six test designs. To accomplish this, our framework required even 47% fewer FPGA compilations than the baseline, demonstrating its superior capacity for achieving fast convergence. Jaemin Seo, Sejin Park 0001, Seokhyeong Kang |
DATE | 3 |
| 2024 | RL-Fill: Timing-Aware Fill Insertion using Reinforcement LearningabstractWe introduce RL-Fill, a novel reinforcement learning framework for timing-aware fill insertion. RL-Fill first generates a large number of fills in the empty spaces and then removes the timing-critical fills as determined by the policy network. Towards faster convergence and stability, our framework employs a two-phase training process. In the first phase, we train the policy with offline expert data using an imitation learning scheme. In the second phase, we further optimize the policy with online data using reinforcement learning. Moreover, we propose a new data augmentation method, LayoutMix, to ensure data-efficient training despite limited number of expert data. Our results demonstrate that RL-Fill is competitive to the commercial tool and outperforms the previous machine learning-based method in timing metrics while adhering density constraints. Jinoh Cho, Seonghyeon Park, Jakang Lee, Sung-Yun Lee, Jinmo Ahn, Seokhyeong Kang |
ICCAD | 6 |
| 2024 | Improving Timing & Power Trade-off in Post-place Optimization Using Multi-agent Reinforcement LearningabstractIn recent years, post-place optimization has emerged as a critical stage in physical design, aiming to improve power, performance, and area (PPA). Among various optimization techniques, buffer insertion, gate sizing, and Vth assignment have become leading optimization techniques for decades. However, these techniques are traditionally applied sequentially, leading to a critical suboptimality problem. Each technique obtains its own iterations performed step by step, preventing the best-optimal optimization for a target instance. To address this limitation, we propose a novel reinforcement learning (RL) based post-place optimization framework that performs various optimization techniques simultaneously. Moreover, to overcome the persistent headache of chip design, timing and power trade-off, we employ multiple agents that target each specific objective. By leveraging deep RL and graph neural network (GNN), our framework models optimal policies, dynamically selecting the most effective action given a target instance. Consequently, our framework obtained an improved Pareto-frontier set compared to comparison baselines while exhibiting 80% and 43% improvements in total negative slack and leakage power, respectively. The results demonstrate that our method outperforms weighted-sum-based co-optimization methods in optimizing timing and power. Jaemin Seo, Sejin Park 0001, Seokhyeong Kang |
ICCAD | 3 |
| 2024 | Mobile Transformer Accelerator Exploiting Various Line Sparsity and Tile-Based Dynamic QuantizationabstractTransformer models are difficult to employ in mobile devices due to their memory-and computation-intensive properties. Accordingly, there is ongoing research on various methods for compressing transformer models, such as pruning and quantization. However, general computing platforms such as central processing units (CPUs) and graphics processing units (GPUs) are not energy-efficient to accelerate the pruned model because the unstructured sparsity they exhibit causes degradation of parallelism. In this paper, we propose a low-power accelerator for transformers that can handle various levels of structured sparsity induced by line pruning with different granularity. Our approach accelerates pruned transformers in a head-wise and line-wise manner. We present a head reorganization and shuffling method that supports head-wise skip operations and resolves the load imbalance problem among processing engines (PEs) caused by the varying number of operations in each head. Furthermore, we implemented a sparse quantized general matrix-to-matrix multiplication (SQ-GEMM) module that supports line-wise skipping and on-the-fly tile-based dynamic quantization of activations. As a result, compared to mobile GPU and CPU, the proposed accelerator improved the energy efficiency by 2.9× and 12.3× for the detection transformer (DETR), and 3.0× and 12.4× for the vision transformer (ViT) models, respectively. In addition, our proposed mobile accelerator achieved the highest energy efficiency among the current state-of-the-art FPGA-based transformer accelerators. Eunji Kwon, Jongho Yoon 0001, Seokhyeong Kang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | MA-Opt: Reinforcement Learning-Based Analog Circuit Optimization Using Multi-ActorsabstractThere is a need for electronic design automation (EDA) tools for analog circuit design since analog circuit design requires substantial human effort and expertise. Using reinforcement learning (RL)-inspired methodologies, this study presents MA-Opt, an analog circuit optimizer. We propose MA-Opt to provide multiple predictions of optimized circuit designs through the use of multiple actors. Multiple actors can be exploited effectively by sharing a memory that affects the loss function of network training, resulting in an accelerated optimization of circuits. Furthermore, we introduce a cooperative near-sampling method deploying a synergistic effect and then optimizing the design. The efficiency of MA-Opt was demonstrated by simulating three analog circuits and comparing the results to other methods. In the experiment, the use of multiple actors with a shared elite solution set and the cooperative near-sampling method proved to be effective. MA-Opt achieved minimum target metrics up to 34$\%$better than DNN-Opt within the same number of simulations while satisfying all given constraints. Moreover, at identical runtime, MA-Opt exhibited better Figure of Merits (FoMs) in comparison to DNN-Opt. Youngchang Choi, Sejin Park 0001, Minjeong Choi, Kyongsu Lee, Seokhyeong Kang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | Graph Partitioning Approach for Fast Quantum Circuit SimulationabstractOwing to the exponential increase in computational complexity, the fast simulation of the large quantum circuit has become very difficult. This is an important challenge for the utilization of quantum computers because it is closely related to the verification of quantum computation by classical machines. The Hybrid Schrödinger-Feynman simulation seems to be a promising solution, but its application is very limited. To solve this drawback, we propose an improved simulation method based on graph partitioning. Experimental results show that our approach significantly reduces the simulation time of the Hybrid Schrödinger-Feynman simulation. Jaekyung Im, Seokhyeong Kang |
ASP-DAC | 2 |
| 2023 | Reinforcement Learning-based Analog Circuit Optimizer using gm/ID for SizingabstractDesigning analog circuits incurs high time costs because designers must consider numerous design variables or trade-off relationships of circuit performance based on a lot of knowledge and experience. To reduce design time, various machine learning methods have been used to optimize analog circuits by learning the correlation between the device size and the circuit performance. However, it is difficult to train the correlation because of its high non-linearity and wide design space. In this paper, this study proposes a new framework to optimize analog circuit designs by combining reinforcement learning (RL) and the sensitivity analysis with gm/IDsizing, which is more intuitive for interpreting circuit performance. Furthermore, the universal value function approximator (UVFA), previously proposed in RL, is modified more simply to make it easier to find the target design. Additionally, the dataset is rearranged and sampled by the criteria that are established based on the principle of circuit operation, which helps to orient the agent to learn the circuit operation. Using the proposed methods, we optimize three types of differential amplifiers with common mode feedback circuits and obtain the best circuit design. Compared to baseline, we find the optimal point using modified UVFA, and moreover, reduce the number of iterations by 42.2%, 39.5%, and 37.5%, respectively, for the three test cases. Minjeong Choi, Youngchang Choi, Kyongsu Lee, Seokhyeong Kang |
DAC | 4 |
| 2023 | MA-Opt: Reinforcement Learning-based Analog Circuit Optimization using Multi-ActorsabstractAnalog circuit design requires significant human efforts and expertise; therefore, electronic design automation (EDA) tools for analog design are needed. This study presents MA-Opt that is an analog circuit optimizer using reinforcement learning (RL)-inspired framework. MA-Opt using multiple actors is proposed to provide various predictions of optimized circuit designs in parallel. Sharing a specific memory that affects the loss function of network training is proposed to exploit multiple actors effectively, accelerating circuit optimization. Moreover, we devise a novel method to tune the most optimized design in previous simulations into a more optimized design. To demonstrate the efficiency of the proposed framework, MA-Opt was simulated for three analog circuits and the results were compared with those of other methods. The experimental results indicated the strength of using multiple actors with a shared elite solution set and the near-sampling method. Within the same number of simulations, while satisfying all given constraints, MA-Opt obtained minimum target metrics up to 24% better than DNN-Opt. Furthermore, MA-Opt obtained better Figure of Merits (FoMs) than DNN-Opt at the same runtime. Youngchang Choi, Minjeong Choi, Kyongsu Lee, Seokhyeong Kang |
DATE | 4 |
| 2023 | Routability Prediction using Deep Hierarchical Classification and RegressionabstractRoutability prediction can forecast the locations where design rule violations occur without routing and thus can speed up the design iterations by skipping the time-consuming routing tasks. This paper investigated (i) how to predict the routability on a continuous value and (ii) how to improve the prediction accuracy for the minority samples. We propose a deep hierarchical classification and regression (HCR) model that can detect hotspots with the number of violations. The hierarchical inference flow can prevent the model from overfitting to the majority samples in imbalanced data. In addition, we introduce a training method for the proposed HCR model that uses Bayesian optimization to find the ideal modeling parameters quickly and incorporates transfer learning for the regression model. We achieved an R2 score of 0.71 for the regression and increased the Fl score in the binary classification by 94% compared to previous work [6]. Jakang Lee, Seokhyeong Kang |
DATE | 3 |
| 2023 | Mobile Accelerator Exploiting Sparsity of Multi-Heads, Lines, and Blocks in Transformers in Computer VisionabstractIt is difficult to employ transformer models for computer vision in mobile devices due to their memory- and computation-intensive properties. Accordingly, there is ongoing research on various methods for compressing transformer models, such as pruning. However, general computing platforms such as central processing units (CPUs) and graphics processing units (GPUs) are not energy-efficient to accelerate the pruned model due to their structured sparsity. This paper proposes a low-power accelerator for transformers with various sizes of structured sparsity induced by pruning with different granularity. In this study, we can accelerate a transformer that has been pruned in a head-wise, line-wise, or block-wise manner. We developed a head scheduling algorithm to support head-wise skip operations and resolve the processing engine (PE) load imbalance problem caused by different number of operations in one head. Moreover, we implemented a sparse general matrix-to-matrix multiplication (sparse GEMM) module that supports line-wise and block-wise skipping. As a result, when compared with a mobile GPU and mobile CPU respectively, our proposed accelerator achieved$6.1\times$and$13.6\times$improvements in energy efficiency for the detection transformer (DETR) model and achieved approximately$2.6\times$and$7.9\times$improvements in the energy efficiency on average for the vision transformer (ViT) models. Eunji Kwon, Haena Song, Seokhyeong Kang |
DATE | 4 |
| 2023 | RL-Legalizer: Reinforcement Learning-based Cell Priority Optimization in Mixed-Height Standard Cell LegalizationabstractCell legalization order has a substantial effect on the quality of modern VLSI designs, which use mixed-height standard cells. In this paper, we propose a deep reinforcement learning framework to optimize cell priority in the legalization phase of various designs. We extract the selected features of movable cells and their surroundings, then embed them into cell-wise deep neural networks. We then determine cell priority and legalize them in order using a pixel-wise search algorithm. The proposed framework uses a policy gradient algorithm and several training techniques, including grid-cell subepisode, data normalization, reduced-dimensional state, and network optimization. We aim to resolve the suboptimality of existing sequential legalization algorithms with respect to displacement and wirelength. On average, our proposed framework achieved 34% lower legalization costs in various benchmarks compared to that of the state-of-the-art legalization algorithm. Sung-Yun Lee, Seonghyeon Park, Minjae Kim 0005, Le Pham Tuyen 0001, Seokhyeong Kang |
DATE | 6 |
| 2023 | FPGA-Based Accelerator for Rank-Enhanced and Highly-Pruned Block-Circulant Neural NetworksabstractNumerous network compression methods have been proposed to deploy deep neural networks in a resource-constrained embedded system. Among them, block-circulant matrix (BCM) compression is one of the promising hardware-friendly methods for both acceleration and compression. However, it has several limitations; (i) limited representation due to the structural characteristic of circulant matrix, (ii) limitation of the compression parameter, (iii) need to specialize the dataflow for BCM-compressed network accelerators. In this paper, rank-enhanced and highly-pruned block-circulant matrices compression (RP-BCM) framework is proposed to overcome these limitations. RP-BCM comprises two stages: Hadamard-BCM and BCM-wise pruning. Moreover, a dedicated skip scheme is introduced to processing element design for exploiting high-parallelism with BCM-wise sparsity. Furthermore, we propose specialized dataflow for a BCM-compressed network on a resource-constrained FPGA. As a result, the proposed method achieves parameter reduction and FLOPs reduction for ResNet-50 in ImageNet by 92.4% and 77.3%, respectively. Moreover, the proposed hardware design achieves$3.1\times$improvement in energy efficiency on the Xilinx PYNQ-Z2 FPGA board for ResNet-18 on ImageNet compared to the GPU. Haena Song, Jongho Yoon 0001, Eunji Kwon, Tae-Hyun Oh, Seokhyeong Kang |
DATE | 6 |
| 2023 | ClusterNet: Routing Congestion Prediction and Optimization Using Netlist Clustering and Graph Neural NetworksabstractAccurately predicting routing congestion caused by netlist topology is essential as circuit designs become increasingly complex. To correctly predict routing congestion, the use of graph neural networks (GNNs) has gained great attention. However, existing GNN-based methods have limitations in capturing crucial netlist information and effectively representing complex topologies. In this work, we propose a novel approach, ClusterNet, to predict routing congestion caused by netlist topology. Our approach leverages netlist clustering to overcome these limitations. We first divide the netlist into highly connected clusters using the Leiden algorithm, enabling an analysis of the local netlist topology. We then predict routing congestion by exploiting GNNs to generate cluster embeddings that capture the detailed netlist topology. In addition, we introduce a cluster padding method that utilizes the trained model to mitigate routing congestion. By applying the proposed ClusterNet, we can accurately predict and optimize routing congestion from specific cluster topologies. Our experimental results demonstrated improved prediction performance, with a mean absolute error of 0.056 and an R2 score of 0.669. Furthermore, routing congestion optimization significantly improved the total negative slack and reduced the number of failing endpoints by 14.5% and 9.9%, respectively. Kyungjun Min, Seongbin Kwon, Sung-Yun Lee, Sunghye Park, Seokhyeong Kang |
ICCAD | 6 |
| 2023 | Routability Prediction and Optimization Using Explainable AIabstractMachine learning (ML) techniques have been widely studied to predict routability in early-stage. To reduce the design turn-around time during the placement and routing iterations, it is crucial to predict the design rule violation (DRV) hotspots precisely before actual detailed routing. However, complex network architectures of ML make it challenging for humans to understand how ML generates predictions and to identify the factors that significantly influence the predictions. This black-box nature of ML limits the efficient integration of the prediction techniques into an optimization process. Explainable artificial intelligence enables the interpretation of decision rationales in the ML model and brings us the reasons underlying the prediction of the model. In this paper, we propose a routability optimization framework that analyzes the input features relevant to the predicted DRV hotspots using an explainable model and selects the most suitable optimization methods. The proposed framework comprises three steps - (1) predicting DRV hotspots in the early-global routing stage, (2) calculating how much each input feature contributes to the predictions and (3) applying a proper optimization method to improve the routability. We reduced the number of DRVs by 78% on average in 16 design layouts without degrading the design Quality. Seonghyeon Park, Seongbin Kwon, Seokhyeong Kang |
ICCAD | 4 |
| 2023 | Multi-Source Transfer Learning for Design Technology Co-OptimizationabstractIn advanced technology nodes, pitch scaling have not kept up with the Moore's Law. To continue progression, the design technology co-optimization (DTCO) has been proposed. However, implementing DTCO requires significant time cost and resources due to iterative trials. In addition, optimal design and technology option depend on each design, thus it should start from scratch whenever the target design changes. We present a DTCO framework based on Bayesian optimization that efficiently explores design feedback for optimization. In addition, our framework incorporates a multi-source transfer Gaussian process (MTGP) that ensures robust optimization even for unseen designs. MTGP significantly improves prediction and generalization performance by integrating multiple single source transfer Gaussian processes. Our framework, on average, reduced the mean absolute error of power and area by 47.3% and 24.1%, respectively, and power and area by 37.3% and 19.9%, respectively, compared to the reference, in 7nm technology nodes. Jakang Lee, Seonghyeon Park, Seokhyeong Kang |
ISLPED | 4 |
| 2023 | Construction of Realistic Place-and-Route Benchmarks for Machine Learning ApplicationsabstractMany design optimization methods using machine learning (ML) techniques have been investigated to reduce the number of design iterations in the physical design flow. The demand for big data to support ML research has been increasing, but the lack of place-and-route (P&R) benchmarks is one of the major problems. We propose a framework to construct realistic P&R benchmarks for use in training ML applications. The framework can organize the P&R database using an artificial netlist generator, which can create any gate-level netlist from user-specified input parameters that represent the topological characteristics of the circuit. We show that a training dataset that contains many artificial gate-level netlists can improve the generalizability of the model to predict the routability for unseen real circuits without using expensive real-world data. Compared to the model that had been trained with real-world circuits, we improved the F1 score in predicting the timing and routing failure by 26.4% and 54.5%, respectively. Sung-Yun Lee, Kyungjun Min, Seokhyeong Kang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Adaptive FSP: Adaptive Architecture Search with Filter Shape Pruning
Aeri Kim, Seungju Lee, Eunji Kwon, Seokhyeong Kang |
ACCV (1) | 4 |
| 2022 | Signal-Integrity-Aware Interposer Bus Routing in 2.5D Heterogeneous IntegrationabstractWe propose a fast interposer bus router that observes the complex design rules of silicon interposer layers and optimizes the signal integrity. By escaping highly integrated physical layers (PHYs) of chiplets and sharing the same bus topology, our router compactly interconnects thousands of bump I/Os within a short timeframe. In addition, we secure the maximum wire pitch and guard the signal wires to optimize the signal integrity in high bandwidth. Compared with the results of a commercial EDA tool, our router is about five times faster and the results are verified to transmit signal in a target data rate with 30% improved eye width and 35% improved eye height for industrial designs. Our router can provide practical routing results for the upcoming 2.5D ICs that have more chiplets and require higher bandwidth than the existing chips. Sung-Yun Lee, Kyungjun Min, Seokhyeong Kang |
ASP-DAC | 4 |
| 2022 | A fast and scalable qubit-mapping method for noisy intermediate-scale quantum computersabstractThis paper presents an efficient qubit-mapping method that redesigns a quantum circuit to overcome the limitations of qubit connectivity. We propose a recursive graph-isomorphism search to generate the scalable initial mapping. In the main mapping, we use an adaptive look-ahead window search to resolve the connectivity constraint within a short runtime. Compared with the state-of-the-art method [15], our proposed method reduced the number of additional gates by 23% on average and the runtime by 68% for the three largest benchmark circuits. Furthermore, our method improved circuit stability by reducing the circuit depth and thus can be a step forward towards fault tolerance. Sunghye Park, Minhyuk Kweon, Jae-Yoon Sim, Seokhyeong Kang |
DAC | 5 |
| 2022 | Design and Evaluation Frameworks for Advanced RISC-based Ternary ProcessorabstractIn this paper, we introduce the design and veri-fication frameworks for developing a fully-functional emerging ternary processor. Based on the existing compiling environments for binary processors, for the given ternary instructions, the software-level framework provides an efficient way to convert the given programs to the ternary assembly codes. We also present a hardware-level framework to rapidly evaluate the performance of a ternary processor implemented in arbitrary design technology. As a case study, the fully-functional 9-trit advanced RISC-based ternary (ART-9) core is newly developed by using the proposed frameworks. Utilizing 24 custom ternary instructions, the 5-stage ART-9 prototype architecture is successfully verified by a number of test programs including dhrystone benchmark in a ternary domain, achieving the processing efficiency of 57.8 DMIPS/W and$3.06\times 10^{6}$DMIPS/W in the FPGA-level ternary-logic emulations and the emerging CNTFET ternary gates, respectively. Dongyun Kam, Jung Gyu Min, Jongho Yoon 0001, Sunmean Kim, Seokhyeong Kang, Youngjoo Lee 0002 |
DATE | 5 |
| 2022 | GAN-Dummy Fill: Timing-aware Dummy Fill Method using GANabstractThe chemical mechanical polishing (CMP) dummy fill method is commonly used for the planarization of the CMP process, resulting in the development of many automated methods. We propose a dummy fill method using a generative adversarial network (GAN) that improves the existing dummy fill methods in terms of the uniformity of metal density and timing of critical nets. The dummy patterns created were similar to those of existing methods. However, the GAN dummy fill method applies additional optimizations to make the CMP dummy fill pattern efficient. The method learns by adding density and parasitic capacitance to the loss function of the GAN. Compared to dummy patterns generated from commercial tools, dummy patterns generated from GAN-dummy fill reduced the negative timing slack due to parasitic capacitance by up to 45%. Myong Kong, Minhyuk Kweon, Seokhyeong Kang |
ACM Great Lakes Symposium on VLSI | 4 |
| 2022 | MCQA: Multi-Constraint Qubit Allocation for Near-FTQC DeviceabstractIn response to the rapid development of quantum processors, quantum software must be advanced by considering the actual hardware limitations. Among the various design automation problems in quantum computing, qubit allocation modifies the input circuit to match the hardware topology constraints. In this work, we present an effective heuristic approach for qubit allocation that considers not only the hardware topology but also other constraints for near-fault-tolerant quantum computing (near-FTQC). We propose a practical methodology to find an effective initial mapping to reduce both the number of gates and circuit latency. We then perform dynamic scheduling to maximize the number of gates executed in parallel in the main mapping phase. Our experimental results with a Surface-17 processor confirmed a substantial reduction in the number of gates, latency, and runtime by 58%, 28%, and 99%, respectively, compared with the previous method [18]. Moreover, our mapping method is scalable and has a linear time complexity with respect to the number of gates. Sunghye Park, Jae-Yoon Sim, Seokhyeong Kang |
ICCAD | 4 |
| 2022 | CPR: Crossbar-grain Pruning for an RRAM-based Accelerator with Coordinate-based Weight MappingabstractResistive random access memory (RRAM)-based crossbar arrays with the process-in-memory (PIM) approach are emerging as a promising technique for accelerating deep neural networks (DNNs) with their high-speed and multi-level programming characteristics. The computation performance of crossbars can be improved by pruning techniques, however, tightly coupled bitlines and wordlines in PIM architectures make the exploitation of data sparsity rather difficult. Therefore, the hardware dependency has to be carefully considered for the application of pruning techniques. In this work, we develop an RRAM-based DNN accelerator, CPR, with (i) a coordinate-based weight mapping method and (ii) a crossbar-grain pruning algorithm. This mapping method locates weights to different crossbar arrays according to their spatial location, thereby increasing input data reuse and reducing waste of energy and latency. We minimize the unused space in crossbars to increase the area efficiency by combining multiple crossbar subarrays. Moreover, by retaining only the desired number of crossbar subarrays in the processing-element (PE) and pruning the rest, we preserve the high computing parallelism without violating the hardware dependency. The overhead of additional circuits is not significant compared to the overall chip design. The experimental results show that our weight mapping method outperforms the conventional method by 1.6 × in TOPs/W with a smaller area. The proposed pruning algorithm reduces the latency by 71%, energy by 70%, and area by 84% from baseline implementation. Seokhyeong Kang |
ICCD | 2 |
| 2022 | Lightweight Speaker Recognition in Poincaré SpacesabstractThis letter proposes a lightweight model for speaker recognition by leveraging a hyperbolic space. The speaker recognition performance heavily depends on the distinctiveness of speaker embeddings induced by metric learning. However, most state-of-the-art embedding methods are typically based on the Euclidean metric space, which does not account for inherent hierarchical structures of speech voice characteristics. The recent development of the neural hyperbolic geometry has demonstrated its effectiveness to model continuous hierarchical structures, which have been typically cumbersome to model by standard deep neural networks. This facet provides an additional by-product of a compact representation. Inspired by the favorable geometry of the hyperbolic geometry, we developed a hyperbolic ResNet for speaker recognition. We found that in smaller dimension regimes than typical cases, the learned speaker embeddings are more discriminative; in other words, more compact at the same level of performance. Our experiments on the large-scale VoxCeleb datasets show that, given the limited channel dimensions of neural networks, our method consistently has favorable performance against the standard ResNet for both speaker recognition and verification tasks. Sung-Bin Kim, Seokhyeong Kang, Tae-Hyun Oh |
IEEE Signal Process. Lett. | 3 |
| 2022 | CHAMP: Channel Merging Process for Cost-Efficient Highly-Pruned CNN AccelerationabstractThis paper presents an advanced offline scheduling scheme to improve the accelerator efficiency, especially for the highly-pruned convolutional neural networks (HP-CNNs). Based on the existing outlier-aware accelerator design, we demonstrate the efficiency drop of HP-CNN processing for the first time, and element-wise channel merging is proposed to make a dense processing sequence even for the highly-pruned model. The dedicated hardware architecture is also presented to process the merged channels with the minimum hardware-level overheads, improving the energy efficiency for handling HP-CNNs by preserving the hardware utilization. We further investigate the optimal size of accumulator and multiplexer, in addition to the number of merged channels, exploiting the attractive energy-performance trade-offs. As a result, unlike the practical HP-CNNs for the on-device solutions, the proposed method enhances the overall efficiency by up to 33% compared to the state-of-the-art schemes. Hyeokjun Kwon, Younghoon Byun, Seokhyeong Kang, Youngjoo Lee 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | Approach to Improve the Performance Using Bit-level Sparsity in Neural NetworksabstractThis paper presents a convolutional neural network (CNN) accelerator that can skip zero weights and handle outliers, which are few but have a significant impact on the accuracy of CNNs, to achieve speedup and increase the energy efficiency of CNN. We propose an offline weight-scheduling algorithm which can skip zero weights and combine two non-outlier weights simultaneously using bit-level sparsity of CNNs. We use a reconfigurable multiplier-and-accumulator (MAC) unit for two purposes; usually used to compute combined two non-outliers and sometimes to compute outliers. We further improve the speedup of our accelerator by clipping some of the outliers with negligible accuracy loss. Compared to DaDianNao [7] and Bit-Tactical [16] architectures, our CNN accelerator can improve the speed by 3.34 and 2.31 times higher and reduce energy consumption by 29.3% and 30.2%, respectively. Yesung Kang, Eunji Kwon, Seunggyu Lee, Younghoon Byun, Youngjoo Lee 0002, Seokhyeong Kang |
DATE | 6 |
| 2021 | MDARTS: Multi-objective Differentiable Neural Architecture SearchabstractIn this work, we present a differentiable neural architecture search (NAS) method that takes into account two competing objectives, quality of result (QoR) and quality of service (QoS) with hardware design constraints. NAS research has recently received a lot of attention due to its ability to automatically find architecture candidates that can outperform handcrafted ones. However, the NAS approach which complies with actual HW design constraints has been under-explored. A naive NAS approach for this would be to optimize a combination of two criteria of QoR and QoS, but the simple extension of the prior art often yields degenerated architectures, and suffers from a sensitive hyperparameter tuning. In this work, we propose a multi-objective differential neural architecture search, called MDARTS. MDARTS has an affordable search time and can find Pareto frontier of QoR versus QoS. We also identify the problematic gap between all the existing differentiable NAS results and those final post-processed architectures, where soft connections are binarized. This gap leads to performance degradation when the model is deployed. To mitigate this gap, we propose a separation loss that discourages indefinite connections of components by implicitly minimizing entropy. Hyun-jeong Kwon, Eunji Kwon, Youngchang Choi, Tae-Hyun Oh, Seokhyeong Kang |
DATE | 6 |
| 2021 | Machine Learning Framework for Early Routability Prediction with Artificial Netlist GeneratorabstractRecent routability research has exploited a machine learning (ML)-based modeling methodologies to consider various routability factors that are derived from placement solution. These factors are very related to the circuit characteristics (e.g., pin density, routing congestion, demand of routing resources, etc), and lack of circuit benchmarks in training can lead to poor predictability for ‘unseen’ circuit designs. In this paper, we propose a machine learning (ML) framework for early routability prediction modeling. The method includes a new artificial netlist generator (ANG) that generates an artificial gate-level netlist from the user-specified topology characteristics of synthetic circuit, even with real world circuit-like. In this framework, we exploit that ANG that supports obtaining ground truths for use in training ML-based model, the training dataset that have a wide range of topological characteristics provides strong ability to inference noisy, previous-unseen data. Compared to a design-specific training dataset [4] that is used for routability prediction modeling, we increase the test accuracy of binary classification (‘pass' or ‘fail’) on timing, DRC and routability by 6.3%, 8.6% and 6.6%, and reduce the generalization error [12] by as much as 87% compared to design-specific training dataset [4]. Hyun-jeong Kwon, Sung-Yun Lee, Seungwon Kim, Mingyu Woo, Seokhyeong Kang |
DATE | 6 |
| 2021 | Design and Analysis of a Low-Power Ternary SRAMabstractThis paper proposes the design of a ternary inverter that uses low current as input voltage is VDD/2. When the supply voltage is set to 1 V, current supplied by a voltage source as an input voltage VDD/2 is reduced by 22.75% from 1.89μA to 1.46μA. By connecting ternary inverters back-to-back, a trit-storage element is implemented as a ternary SRAM cell. This paper also presents the first verification of read/write schemes that consider noise margins. Youngchang Choi, Sunmean Kim, Kyongsu Lee, Seokhyeong Kang |
ISCAS | 4 |
| 2021 | Reinforcement Learning-Based Power Management Policy for Mobile Device SystemsabstractThis paper presents a power management policy that utilizes reinforcement learning to increase the power efficiency of mobile device systems based on a multiprocessor system-on-a-chip (MPSoC). The proposed policy predicts a system’s characteristics and learns power management controls to adapt to the variations in the system. We consider the behavioral characteristics of systems that run on mobile devices under diverse scenarios. Therefore, the policy can flexibly manage the system power regardless of the application scenario and achieve lower energy consumption without compromising the user satisfaction. The average energy per unit quality of service (QoS) of the proposed policy is lower than that of the previous six dynamic voltage/frequency scaling governors by 31.66%. Furthermore, we reduce the runtime overhead by implementing the proposed policy as hardware. We implemented the policy on the field programmable gate array (FPGA) and construct a communication interface between the central processing units (CPUs) and the hardware of the proposed policy. Decision-making by the hardware-implemented policy is 3.92 times faster than by the software-implemented policy. Eunji Kwon, Sodam Han, Yoonho Park, Jongho Yoon 0001, Seokhyeong Kang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2021 | Variation-Aware SRAM Cell Optimization Using Deep Neural Network-Based Sensitivity AnalysisabstractUnder process, voltage, and temperature variations, SRAM cell stability largely fluctuates from the nominal value. In the design step, SRAM cell optimization while ignoring the fluctuation induces the yield loss for the stability. Variation-aware optimization of an SRAM cell can prevent the yield loss problem by considering the mean and variance of SRAM cell stability when finding optimal design parameters. This paper proposes a novel SRAM optimization method that uses a deep neural network (DNN). Multiple DNNs from ensemble techniques represent the mean and variance of SRAM cell stability for the nominal design parameters. Subsequent sensitivity analysis of DNN extracts the K design parameters that have the most dominant effects on the mean and variance of SRAM cell stability. Then multidimensional optimization is used to find the optimal values of these K parameters to maximize the mean stability while minimizing its variance. The proposed method achieved an average of 2% error compared to MC simulation. The proposed optimization method takes only 561 s to provide the most optimal design parameter values of an SRAM cell. Hyun-jeong Kwon, Young Hwan Kim, Seokhyeong Kang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | Enhanced Power Delivery Pathfinding for Emerging 3-D Integration TechnologyabstractIn advanced technology nodes, emerging 3-D integration technology is a promising “More Than Moore” lever for continued scaling of system capability and value. In the 3-D integrated circuit (3-D IC) implementation, the power delivery network (PDN) is crucial to meeting design specifications. However, determining the optimal PDN design is nontrivial. On the one hand, to meet the voltage (IR) drop requirement, a denser power mesh is desired. On the other hand, to meet the timing requirement, more routing resource is needed for signal routing. Moreover, additional competition between signal routing and power routing is caused by intertier vertical interconnects in 3-D IC. In this article, we propose a power delivery pathfinding methodology for emerging 3-D integration, which seeks to identify a “near-optimal” (or, very high quality) PDN for a given BEOL stack, vertical interconnection, and PDN specification. Compared with previous works, our methodology can explore richer solution spaces as it supports different PDN layer combinations and PDN layer configurations. We develop models for routability and worst IR drop to help reduce iterations between PDN design and circuit design in 3-D IC implementation. We present validations and demonstrate improvement in IR drop and routability with real design blocks in 28- and 14-nm foundry technology nodes. Andrew B. Kahng, Seokhyeong Kang, Seungwon Kim, Bangqi Xu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | Late Breaking Results: Reinforcement Learning-based Power Management Policy for Mobile Device SystemsabstractThis paper presents a power management policy that exploits reinforcement learning to increase power efficiency of mobile device systems. Our Q-learning-based policy predicts a system’s characteristics and learns power management controls to adapt to the system’s variations. Therefore, we can flexibly manage the system power regardless of the application scenario and can achieve lower energy per QoS compared to previous dynamic voltage/frequency scaling governors. To minimize the process overhead, we implemented our power management policy as hardware; the hardware-implemented policy reduced the average latency up to 40× compared to the software-implemented policy. Eunji Kwon, Sodam Han, Yoonho Park, Young Hwan Kim, Seokhyeong Kang |
DAC | 5 |
| 2020 | Analysis and Solution of CNN Accuracy Reduction over Channel Loop TilingabstractOwing to the growth of the size of convolutional neural networks (CNNs), quantization and loop tiling (also called loop breaking) are mandatory to implement CNN on an embedded system. However, channel loop tiling of quantized CNNs induces unexpected errors. We explain why channel loop tiling of quantized CNNs induces the unexpected errors, and how the errors affect the accuracy of state-of-the-art CNNs. We also propose a method to recover accuracy under channel tiling by compressing and decompressing the most-significant bits of partial sums. Using the proposed method, we can recover accuracy by 12.3% with only 1% circuit area overhead and an additional 2% of power consumption. Yesung Kang, Yoonho Park, Eunji Kwon, Taeho Lim, Sangyun Oh, Mingyu Woo, Seokhyeong Kang |
DATE | 8 |
| 2020 | GRLC: grid-based run-length compression for energy-efficient CNN acceleratorabstractConvolutional neural networks (CNNs) require a huge amount of off-chip DRAM access, which accounts for most of its energy consumption. Compression of feature maps can reduce the energy consumption of DRAM access. However, previous compression methods show poor compression ratio if the feature maps are either extremely sparse or dense. To improve the compression ratio efficiently, we have exploited the spatial correlation and the distribution of non-zero activations in output feature maps. In this work, we propose a grid-based run-length compression (GRLC) and have implemented a hardware for the GRLC. Compared with a previous compression method [1], GRLC reduces 11% of the DRAM access and 5% of the energy consumption on average in VGG-16, ExtractionNet and ResNet-18. Yoonho Park, Yesung Kang, Eunji Kwon, Seokhyeong Kang |
ISLPED | 5 |
| 2020 | Compact Topology-Aware Bus Routing for Design RegularityabstractIn bus routing, if signal bits in a bus structure share a common routing topology, routability is increased by avoiding twisted patterns and variation immunity. The bus routing problem has become significantly important because of increasing complexity of bus structures for multichip-module, I/O pins, or on-chip memories in advanced technology. We present and evaluate a compact topology-aware bus routing method that can both compactly synthesize the routing topology of the bus and minimize design rule violations even in designs with high bus density and high track utilization. Our proposed method completed the bus routing in the runtime limit of the ICCAD-2018 contest and achieved 66% reduction in total cost compared with the winner of that contest. SangGi Do, Sung-Yun Lee, Seokhyeong Kang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Additive Statistical Leakage Analysis Using Exponential Mixture ModelabstractVariation-aware leakage analysis becomes an essential design process as the technology node continuously shrinks. This article proposes a novel additive statistical leakage analysis method that uses exponential mixture model (EMM) to estimate the leakage distribution. Using a few leakage data for sub-blocks of an input circuit, we estimate any shape of leakage distribution regardless of new process nodes or operating conditions. Leakage distribution of an input circuit can be obtained by adding the leakage distributions of the sub-blocks. The proposed addition step sequentially adds the leakage distributions of sub-blocks that are expressed as EMMs. Before the addition step, we improve the accuracy by handling linear dependence among leakage simulation data of sub-blocks. In addition, we propose a method to reduce the number of components of an EMM to prevent exponential increase in runtime and memory during the addition process. The proposed method achieved 43.6 times improvement in goodness-of-fit of the estimated cumulative density functions compared to the best results of other analytic model-based methods. Hyun-jeong Kwon, Sung-Yun Lee, Young Hwan Kim, Seokhyeong Kang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | Power Delivery Pathfinding for Emerging Die-to-Wafer Integration TechnologyabstractIn advanced technology nodes, emerging die-to-wafer (D2W) integration technology is a promising "More Than Moore" lever for continued scaling of system capability and value. In D2W 3D IC implementation, the power delivery network (PDN) is crucial to meeting design specifications. However, determining the optimal PDN design is nontrivial. On the one hand, to meet the IR drop requirement, denser power mesh is desired. On the other hand, to meet the timing requirement for a high-utilization design, more routing resource should be available for signal routing. Moreover, additional competition between signal routing and power routing is caused by inter-tier vertical interconnects in 3D IC. In this paper, we propose a power delivery pathfinding methodology for emerging die-to-wafer integration, which seeks to identify an optimal or near-optimal PDN for a given design and PDN specification. Our pathfinding methodology exploits models for routability and worst IR drop, which helps reduce iterations between PDN design and circuit design in 3D IC implementation. We present validations with real design examples and a 28nm foundry technology. Andrew B. Kahng, Seokhyeong Kang, Seungwon Kim, Kambiz Samadi, Bangqi Xu |
DATE | 2 |
| 2019 | Fence-Region-Aware Mixed-Height Standard Cell LegalizationabstractWe propose a fence-region-aware mixed-height standard cell legalization that can optimize the placement of standard cells that have more than a two row height in various shapes of the fence region. The algorithm consists of pre-legalization and mixed-height standard cell legalization steps to prioritize cell legalization; then a quality refinement step that uses simulated annealing reduces the displacement. Our proposed method achieved 63% improvement in the average quality score and 72% improvement in average runtime, compared to the winners of the ICCAD-2017 contest. SangGi Do, Mingyu Woo, Seokhyeong Kang |
ACM Great Lakes Symposium on VLSI | 3 |
| 2018 | An optimal gate design for the synthesis of ternary logic circuitsabstractOver the last few decades, CMOS-based digital circuits have been steadily developed. However, because of the power density limits, device scaling may soon come to an end, and new approaches for circuit designs are required. Multi-valued logic (MVL) is one of the new approaches, which increases the radix for computation to lower the complexity of the circuit. For the MVL implementation, ternary logic circuit designs have been proposed previously, though they could not show advantages over binary logic, because of unoptimized synthesis techniques. In this paper, we propose a methodology to design ternary gates by modeling pull-up and pull-down operations of the gates. Our proposed methodology makes it possible to synthesize ternary gates with a minimum number of transistors. From HSPICE simulation results, our ternary designs show significant power-delay product reductions; 49 % in the ternary full adder and 62 % in the ternary multiplier compared to the existing methodology. We have also compared the number of transistors in CMOS-based binary logic circuits and ternary device-based logic circuits. Sunmean Kim, Taeho Lim, Seokhyeong Kang |
ASP-DAC | 3 |
| 2018 | Fast chip-package-PCB coanalysis methodology for power integrity of multi-domain high-speed memory: A case studyabstractThe power integrity of high-speed interfaces is an increasingly important issue in mobile memory systems. However, because of complicated design variations such as adjacent VDD domain coupling, conventional case-specific modeling is limited in analyzing trends in results from parametric variations. Moreover, conventional industrial methods can be simulated only after the design layout is completed and it requires a lot of back-annotation processes, which result in delayed delays time to market. In this paper, we propose a chip-package-PCB coanalysis methodology applied to our multi-domain high-speed memory system model with a current generation method. Our proposed parametric simulation model can analyze the tendency of power integrity results from variable sweeps and Monte Carlo simulations, and it shows a significantly reduced runtime compared to the conventional EDA methodology under JEDEC LPPDR4 environment. Seungwon Kim, Ki Jin Han, Seokhyeong Kang |
DATE | 4 |
| 2017 | Fast Predictive Useful Skew Methodology for Timing-Driven Placement OptimizationabstractIncremental timing-driven placement (TDP) is one of the most crucial steps for timing closure in a physical design. The need for high-performance incremental TDP continues to grow, but prior studies have focused on optimizing only setup timing slacks, which can be easily stuck in local optima. In this paper, we present a useful skew methodology based on a maximum mean weight cycle (MMWC) approach in the incremental TDP. The proposed useful skew methodology finds an optimal clock latency for each flip-flop, and the clock latency is implemented by moving the flip-flops and/ or reassigning them to local clock buffers. With the proposed TDP method, we effectively reduce the early slack of ICCAD 2015 contest benchmarks, and achieve 124(%) and 78(%) of total quality score improvement compared to the 2015 contest winner, and early slack histogram compression (EHC) method, respectively. Moreover, with fewer iterations in the optimization, the runtime of our predictive useful skew method is an average of 7.4 times faster than an EHC method. Seungwon Kim, SangGi Do, Seokhyeong Kang |
DAC | 3 |
| 2017 | GRASP based metaheuristics for layout pattern classificationabstractLayout pattern classification has been recently utilized in IC design. It clusters hotspot patterns for design-space analysis or yield optimization. In pattern classification, an optimal clustering is essential, as well as its runtime and accuracy. Within the research-oriented infrastructure used in the ICCAD 2016 contest, we have developed a fast metaheuristic for the pattern classification that utilizes the Greedy Randomized Adaptive Search Procedure (GRASP). Our proposed metaheuristic outperforms the best-reported results on all of the ICCAD 2016 benchmarks. In addition, we achieve up to a 50% cluster count reduction, and improve a runtime significantly compared to a commercial EDA tool provided in the ICCAD 2016 contest [1]. Mingyu Woo, Seungwon Kim, Seokhyeong Kang |
ICCAD | 3 |
| 2016 | Novel approximate synthesis flow for energy-efficient FIR filterabstractThe portability of emerging computing systems demands further reduction in the power consumption of their components. Approximate computing can reduce power consumption by using a simplified or an inaccurate circuit. In this paper, the energy efficiency of a finite impulse response (FIR) filter is improved through approximate computing. We propose an approximate synthesis technique for an energy-efficient FIR filter with an acceptable level of accuracy. We employ the common subexpression elimination (CSE) algorithm to implement the FIR filter and replace conventional adder/subtractors with approximate ones. While yielding acceptable rates of accuracy, the proposed flow can attain a maximum energy saving of 50.7% in comparison with conventional FIR filter designs. Yesung Kang, Seokhyeong Kang |
ICCD | 3 |
| 2016 | Wakeup scheduling and its buffered tree synthesis for power gating circuits
Seungwhun Paik, Seokhyeong Kang, Youngsoo Shin |
Integr. | 3 |
| 2016 | Novel Adaptive Power-Gating Strategy and Tapered TSV Structure in Multilayer 3D ICabstractAmong power dissipation components, leakage power has become more dominant with each successive technology node. Power-gating techniques have been widely used to reduce the standby leakage energy. In this work, we investigate a power-gating strategy for through-silicon via (TSV)-based 3D IC stacking structures. Power-gating control is becoming more complicated as more dies are stacked. We combine the on-chip PDN and TSV in a multilayered 3D IC to perform power-gating analysis of the static and dynamic voltage drops and in-rush current. Then, we propose a novel power-gating strategy that optimizes the in-rush current profile, subject to the voltage-drop constraints. Our power-gating strategy provides a minimal wake-up latency such that the voltage noise safety margins are not violated. In addition, the layer dependency of the 3D IC on the power gating is analyzed in terms of the wake-up time reduction. We achieve an average wake-up time reduction of 43% for all cases with our adaptive power-gating method that exploits location (or layer) information regarding the aggressors in a 3D IC. A tapered TSV architecture based on the layer dependency has been analyzed; it exhibits up to 18% wake-up time reduction compared to that of circuits with uniform TSVs. Seungwon Kim, Seokhyeong Kang, Ki Jin Han |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2016 | Synthesis of Dual-Mode Circuits Through Library Design, Gate Sizing, and Clock-Tree OptimizationabstractA dual-mode circuit is a circuit that has two operating modes: a default high-performance mode at nominal voltage and a secondary low-performance near-threshold voltage (NTV) mode. A key problem that we address is to maximize NTV mode clock frequency. Some cells that are particularly slow in NTV mode are optimized through transistor sizing and stack removal; static noise margin of each gate is extracted and appended in a library so that function failures can be checked and removed during synthesis. A new gate-sizing algorithm is proposed that takes account of timing slacks at both modes. A new sensitivity measure is introduced for this purpose; binary search is then applied to find the maximum NTV mode frequency. Clock-tree synthesis is reformulated to minimize clock skew at both modes. This is motivated by the fact that the proportion of load-dependent delay along clock paths, as well as clock-path delays themselves, should be made equal. Experiments on some test circuits indicate that NTV mode clock period is reduced by 24%, on average; clock skew at NTV decreases by 13%, on average; and NTV mode energy-delay product is reduced by 20%, on average. Seokhyeong Kang, Youngsoo Shin |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2015 | An optimal operating point by using error monitoring circuits with an error-resilient techniqueabstractFor applications related to human, such as Internet of Things (loT) and wearable devices, near threshold voltage (NTV) technology has been proposed for the trade-off between performance and energy consumption. However, errorresilient techniques are required in the circuits to improve reliability of the NTV operation. In this paper, we propose a low-overhead error-resilient system and a design flow for NTV operations. We use a new monitoring circuit, which can detect timing errors and find an optimal operation point of the system. Also, we propose two different methodologies, which are slack-based methodology and sensitivity-based methodology. From the proposed monitoring system and the sensitivitybased sorting algorithm, benchmark results show that the optimal designs provide up to 46% monitoring area reduction maintaining similar error detection ability of the conventional error-resilient design. Seungwon Kim, Seokhyeong Kang |
VLSI-SoC | 4 |
| 2015 | An Improved Methodology for Resilient Design ImplementationabstractResilient design techniques are used to (i) ensure correct operation under dynamic variations and to (ii) improve design performance (e.g., timing speculation). However, significant overheads (e.g., 16% and 14% energy penalties due to throughput degradation and additional circuits) are incurred by existing resilient design techniques. For instance, resilient designs require additional circuits to detect and correct timing errors. Further, when there is an error, the additional cycles needed to restore a previous correct state degrade throughput, which diminishes the performance benefit of using resilient designs. In this work, we describe an improved methodology for resilient design implementation to minimize the costs of resilience in terms of power, area, and throughput degradation. Our methodology uses two levers: selective-endpoint optimization (i.e., sensitivity-based margin insertion) and clock skew optimization. We integrate the two optimization techniques in an iterative optimization flow which comprehends toggle rate information and the trade-off between cost of resilience and margin on combinational paths. Since the error-detection network can result in up to 9% additional wirelength cost, we also propose a matching-based algorithm for construction of the error-detection network to minimize this resilience overhead. Further, our implementations comprehend the impacts of signoff corners (in particular, hold constraints, and use of typical vs. slow libraries) and process variation, which are typically omitted in previous studies of resilience trade-offs. Our proposed flow achieves energy reductions of up to 21% and 10% compared to a conventional (with only margin used to attain robustness) design and a brute-force implementation (i.e., a typical resilient design, where resilient endpoints are (greedily) instantiated at timing-critical endpoints), respectively. We show that these benefits increase in the context of an adaptive voltage scaling strategy. Andrew B. Kahng, Seokhyeong Kang, Jiajia Li 0002, José Pineda de Gyvez |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2014 | A new methodology for reduced cost of resilienceabstractResilient design techniques are used to (i) ensure correct operation under dynamic variations; and (ii) improve design performance (e.g., through timing speculation). However, significant overheads (e.g., 17% and 15% energy penalties due to throughput degradation and additional circuits) are incurred by existing resilient design techniques. For instance, resilient designs require additional circuits to detect and correct timing errors. Further, when there is an error, the additional cycles needed to restore a previous correct state degrade throughput, which diminishes the performance benefit of using resilient designs. In this work, we propose a methodology for resilient design implementation to minimize the costs of resilience in terms of power, area and throughput degradation. Our methodology uses two levers: selective-endpoint optimization (i.e., sensitivity-based margin insertion) and clock skew optimization. We integrate the two optimization techniques in an iterative optimization flow which comprehends toggle rate information and the tradeoff between cost of resilience and margin on combinational paths. Our proposed flow achieves energy reductions of up to 19% and 21% compared to a conventional design (with only margin used to attain robustness) and a brute-force implementation, respectively. These benefits increase in the context of an adaptive voltage scaling strategy. Andrew B. Kahng, Seokhyeong Kang, Jiajia Li 0002 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2013 | Smart non-default routing for clock power reductionabstractAt advanced process nodes, non-default routing rules (NDRs) are integral to clock network synthesis methodologies. NDRs apply wider wire widths and spacings to address electromigration constraints, and to reduce parasitic and delay variations. However, wider wires result in larger driven capacitance and dynamic power. In this work, we quantify the potential for capacitance and power reduction through the application of "smart" NDR (SNDR) that substitute narrower-width NDRs on selected clock network segments, while maintaining skew, slew, delay and EM reliability criteria. We propose a practical methodology to apply smart NDRs in standard clock tree synthesis flows. Our studies with a 32/28nm library and open-source benchmarks confirm substantial (average of 9.2%) clock wire capacitance reduction and an average of 4.9% clock switching power savings over the current fixed-NDR methodology, without loss of QoR in the clock distribution. Andrew B. Kahng, Seokhyeong Kang, Hyein Lee 0001 |
DAC | 2 |
| 2013 | Active-mode leakage reduction with data-retained power gatingabstractPower gating is one of the most effective solutions available to reduce leakage power. However, power gating is not practically usable in an active mode due to the overheads of inrush current and data retention. In this work, we propose a data-retained power gating (DRPG) technique which enables power gating of flip-flops during active mode. More precisely, we combine clock gating and power gating techniques, with the flip-flops being power-gated during clock masked periods. We introduce a retention switch which retains data during the power gating. With the retention switch, correct logic states and functionalities are guaranteed without additional control circuitry. The proposed technique can achieve significant active-mode leakage reduction over conventional designs with small area and performance overheads. In studies with a 65nm foundry library and open-source benchmarks, DRPG achieves up to 25.7% active-mode leakage savings (11.8% savings on average) over conventional designs. Andrew B. Kahng, Seokhyeong Kang, Bongil Park |
DATE | 2 |
| 2013 | High-performance gate sizing with a signoff timerabstractProcess and device scaling in late-CMOS technologies highlight leakage power as a critical challenge for the semiconductor industry. Careful gate sizing and Vth-swapping can reduce leakage, but prior optimizations based on convex or dynamic programming (i) are often based on unrealistic assumptions about circuit delay and slew propagation, (ii) fail to handle practical design rules such as transition time or load upper bounds, and (iii) do not scale well to input complexities when full extracted parasitics are available. Seeing substantial opportunities for improvement, we present a multithreaded, stochastic optimization (Trident2.0) for gate sizing and Vthassignment to minimize leakage power subject to capacitance, slew and timing constraints. Scalability and high performance of Trident2.0 are validated on ISPD-2013 Gate Sizing Contest benchmarks. Andrew B. Kahng, Seokhyeong Kang, Hyein Lee 0001, Igor L. Markov, Pankit Thapar |
ICCAD | 2 |
| 2013 | Statistical analysis and modeling for error composition in approximate computation circuitsabstractAggressive requirements for low power and high performance in VLSI designs have led to increased interest in approximate computation. Approximate hardware modules can achieve improved energy efficiency compared to accurate hardware modules. While a number of previous works have proposed hardware modules for approximate arithmetic, these works focus on solitary approximate arithmetic operations. To utilize the benefit of approximate hardware modules, CAD tools should be able to quickly and accurately estimate the output quality of composed approximate designs. A previous work [10] proposes an interval-based approach for evaluating the output quality of certain approximate arithmetic designs. However, their approach uses sampled error distributions to store the characterization data of hardware, and its accuracy is limited by the number of intervals used during characterization. In this work, we propose an approach for output quality estimation of approximate designs that is based on a lookup table technique that characterizes the statistical properties of approximate hardwares and a regression-based technique for composing statistics to formulate output quality. These two techniques improve the speed and accuracy for several error metrics over a set of multiply-accumulator testcases. Compared to the interval-based modeling approach of [10], our approach for estimating output quality of approximate designs is 3.75ß more accurate for comparable runtime on the testcases and achieves 8.4ß runtime reduction for the error composition flow. We also demonstrate that our approach is applicable to general testcases. Wei-Ting Jonas Chan, Andrew B. Kahng, Seokhyeong Kang, Rakesh Kumar 0002, John Sartori |
ICCD | 3 |
| 2013 | Many-Core Token-Based Adaptive Power GatingabstractAmong power dissipation components, leakage power has become more dominant with each successive technology node. Leakage energy waste can be reduced by power gating. In this paper, we extend token-based adaptive power gating (TAP), a technique to power gate an actively executing core during memory accesses, to many-core Chip Multi-Processors (CMPs). TAP works by tracking every system memory request and its estimated time of arrival so that a core may power gate itself without performance or energy loss. Previous work on TAPshows several benefits compared to earlier state-of-the-art techniques, including zero performance hit and 2.58 times average energy savings for out-of-order cores. We show that TAP can adapt to increasing memory contention by increasing power-gated time by 3.69 times compared to a low memory-pressure case. We also scale TAP to many-core architectures with a distributed wake-up controller that is capable of supporting staggered wake-ups and able to power gate each core for 99.07% of the time, achieved by a non-scalable centralized scheme. Andrew B. Kahng, Seokhyeong Kang, Tajana Rosing, Richard D. Strong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2013 | Enhancing the Efficiency of Energy-Constrained DVFS DesignsabstractThe proliferation of embedded systems and mobile devices has created an increasing demand for low-energy hardware. Dynamic voltage and frequency scaling (DVFS) is a popular energy reduction technique that allows a hardware design to reduce average power consumption while still enabling the design to meet a high-performance target when necessary. To conserve energy, many DVFS-based embedded and mobile devices often spend a large fraction of their lifetimes in a low-power mode. However, DVFS designs produced by conventional multimode CAD flows tend to have significant energy overheads when operating outside of the peak performance mode, even when they are operating in a low-power mode. A dedicated core can be added for low-energy operation, but has a high cost in terms of area and leakage. In this paper, we explore the DVFS design space to identify the factors that affect DVFS efficiency. Based on our insights, we propose two design-level techniques to enhance the energy efficiency of DVFS for energy constrained systems. First, we present a context-aware DVFS design flow that considers the intrinsic characteristics of the hardware design, as well as the operating scenario—including the relative amounts of time spent in different modes, the range of performance scalability, and the target efficiency metric—to optimize the design for maximum energy efficiency. We also present a selective replication-based DVFS design methodology that identifies hardware modules for which context-aware multimode design may be inefficient and creates dedicated module replicas for different operating modes for such modules. We show that context-aware design can reduce average power by up to 20% over a conventional multimode design flow. Selective replication can reduce average power by an additional 4%. We also use the generated insights to identify microarchitectural decisions that impact DVFS efficiency. We show that the benefits from the proposed design-level techniques increase when microarchitectural transformations are allowed. Andrew B. Kahng, Seokhyeong Kang, Rakesh Kumar 0002, John Sartori |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | Accuracy-configurable adder for approximate arithmetic designsabstractApproximation can increase performance or reduce power consumption with a simplified or inaccurate circuit in application contexts where strict requirements are relaxed. For applications related to human senses, approximate arithmetic can be used to generate sufficient results rather than absolutely accurate results. Approximate design exploits a tradeoff of accuracy in computation versus performance and power. However, required accuracy varies according to applications, and 100% accurate results are still required in some situations. In this paper, we propose an accuracy-configurable approximate (ACA) adder for which the accuracy of results is configurable during runtime. Because of its configurability, the ACA adder can adaptively operate in both approximate (inaccurate) mode and accurate mode. The proposed adder can achieve significant throughput improvement and total power reduction over conventional adder designs. It can be used in accuracy-configurable applications, and improves the achievable tradeoff between performance/power and quality. The ACA adder achieves approximately 30% power reduction versus the conventional pipelined adder at the relaxed accuracy requirement. Andrew B. Kahng, Seokhyeong Kang |
DAC | 2 |
| 2012 | MAPG: Memory access power gatingabstractIn mobile systems, the problems of short battery life and increased temperature are exacerbated by wasted leakage power. Leakage power waste can be reduced by power-gating a core while it is stalled waiting for a resource. In this work, we propose and model memory access power gating (MAPG), a low-overhead technique to enable power gating of an active core when it stalls during a long memory access. We describe a programmable two-stage power gating switch design that can vary a core's wake-up delay while maintaining voltage noise limits and leakage power savings. We also model the processor power distribution network and the effect of memory access power gating on neighboring cores. Last, we apply our power gating technique to actual benchmarks, and examine energy savings and overheads from power gating stalled cores during long memory accesses. Our analyses show the potential for over 38% energy savings given “perfect” power gating on memory accesses; we achieve energy savings exceeding 20% for a practical, counter-based implementation. Kwangok Jeong, Andrew B. Kahng, Seokhyeong Kang, Tajana Rosing, Richard D. Strong |
DATE | 3 |
| 2012 | Sensitivity-guided metaheuristics for accurate discrete gate sizingabstractThe well-studied gate-sizing optimization is a major contributor to IC power-performance tradeoffs. Viable optimizers must accurately model circuit timing, satisfy a variety of constraints, scale to large circuits, and effectively utilize a large (but finite) number of possible gate configurations, including Vt and Lg. Within the research-oriented infrastructure used in the ISPD 2012 Gate Sizing Contest, we develop a metaheuristic approach to gate sizing that integrates timing and power optimization, and handles several types of constraints. Our solutions are evaluated using a rigorous protocol that computes circuit delay with Synopsys PrimeTime. Our implementation Trident outperforms the best-reported results on all but one of the ISPD 2012 benchmarks. Compared to the 2012 contest winner, we further reduce leakage power by an average of 43%. Andrew B. Kahng, Seokhyeong Kang, Myung-Chul Kim, Igor L. Markov |
ICCAD | 3 |
| 2012 | TAP: token-based adaptive power gatingabstractWe propose a low-overhead technique, Token-Based Adaptive Power Gating (TAP), to power gate an actively executing out-of-order core during memory accesses. TAP tracks every system memory request, providing a lower-bound estimate for the response time. TAP also tracks the state of every power-gateable core in the system, to provide minimal latency wake-up modes to cores such that voltage noise safety margins are not violated. A power-gating switch that utilizes TAP can deterministically power gate its core with energy savings up to 22.39% and no performance hit. Andrew B. Kahng, Seokhyeong Kang, Tajana Rosing, Richard D. Strong |
ISLPED | 2 |
| 2012 | Construction of realistic gate sizing benchmarks with known optimal solutionsabstractGate sizing in VLSI design is a widely-used method for power or area recovery subject to timing constraints. Several previous works have proposed gate sizing heuristics for power and area optimization. However, finding the optimal gate sizing solution is NP-hard, and the suboptimality of sizing solutions has not been sufficiently quantified for each heuristic. Thus, the need for further research has been unclear. Andrew B. Kahng, Seokhyeong Kang |
ISPD | 2 |
| 2012 | Recovery-Driven Design: Exploiting Error Resilience in Design of Energy-Efficient ProcessorsabstractConventional computer-aided design (CAD) methodologies optimize a processor module for correct operation and prohibit timing violations during nominal operation. We propose recovery-driven design, a design approach that optimizes a processor module for a target timing error rate (ER) instead of correct operation. The target ER is chosen based on how many errors can be gainfully tolerated by a hardware or software error resilience mechanism. We show that significant power benefits are possible from a recovery-driven design approach that deliberately allows errors caused by voltage overscaling to occur during nominal operation, while relying on an error resilience technique to tolerate these errors. We present a detailed evaluation and analysis of such a CAD methodology that minimizes the power of a processor module for a target ER. We show how this design-level methodology can be extended to design recovery-driven processors—processors that are optimized to take advantage of hardware or software error resilience. We also discuss a gradual slack recovery-driven design approach that optimizes for a range of ERs to create soft processors—processors that have graceful failure characteristics and the ability to trade throughput or output quality for additional energy savings over a range of ERs. We demonstrate significant power benefits over conventional design—11.8% on average over all modules and ER targets, and up to 29.1% for individual modules. Processor-level benefits were 19.0%, on average. Benefits increase when recovery-driven design is coupled with an error resilience mechanism or when the number of available voltage domains increases. Andrew B. Kahng, Seokhyeong Kang, Rakesh Kumar 0002, John Sartori |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2010 | Slack redistribution for graceful degradation under voltage overscalingabstractModern digital IC designs have a critical operating point, or “wall of slack”, that limits voltage scaling. Even with an error-tolerance mechanism, scaling voltage below a critical voltage - so-called overscaling - results in more timing errors than can be effectively detected or corrected. This limits the effectiveness of voltage scaling in trading off system reliability and power. We propose a design-level approach to trading off reliability and voltage (power) in, e.g., microprocessor designs. We increase the range of voltage values at which the (timing) error rate is acceptable; we achieve this through techniques for power-aware slack redistribution that shift the timing slack of frequently-exercised, near-critical timing paths in a power- and area-efficient manner. The resulting designs heuristically minimize the voltage at which the maximum allowable error rate is encountered, thus minimizing power consumption for a prescribed maximum error rate and allowing the design to fail more gracefully. Compared with baseline designs, we achieve a maximum of 32.8% and an average of 12.5% power reduction at an error rate of 2%. The area overhead of our techniques, as evaluated through physical implementation (synthesis, placement and routing), is no more than 2.7%. Andrew B. Kahng, Seokhyeong Kang, Rakesh Kumar 0002, John Sartori |
ASP-DAC | 2 |
| 2010 | Recovery-driven design: a power minimization methodology for error-tolerant processor modulesabstractConventional CAD methodologies optimize a processor module for correct operation, and prohibit timing violations during nominal operation. In this paper, we propose recovery-driven design, a design approach that optimizes a processor module for a target timing error rate instead of correct operation. We show that significant power benefits are possible from a recovery-driven design flow that deliberately allows errors caused by voltage overscaling ([10],[3]) to occur during nominal operation, while relying on an error recovery technique to tolerate these errors. We present a detailed evaluation and analysis of such a CAD methodology that minimizes the power of a processor module for a target error rate. We demonstrate power benefits of up to 25%, 19%, 22%, 24%, 20%, 28%, and 20% versus traditional P&R at error rates of 0.125%, 0.25%, 0.5%, 1%, 2%, 4%, and 8%, respectively. Coupling recovery-driven design with an error recovery technique enables increased efficiency and additional power savings. Andrew B. Kahng, Seokhyeong Kang, Rakesh Kumar 0002, John Sartori |
DAC | 2 |
| 2010 | Designing a processor from the ground up to allow voltage/reliability tradeoffsabstractCurrent processor designs have a critical operating point that sets a hard limit on voltage scaling. Any scaling beyond the critical voltage results in exceeding the maximum allowable error rate, i.e., there are more timing errors than can be effectively and gainfully detected or corrected by an error-tolerance mechanism. This limits the effectiveness of voltage scaling as a knob for reliability/power tradeoffs. In this paper, we present power-aware slack redistribution, a novel design-level approach to allow voltage/reliability tradeoffs in processors. Techniques based on power-aware slack redistribution reapportion timing slack of the frequently-occurring, near-critical timing paths of a processor in a power- and area-efficient manner, such that we increase the range of voltages over which the incidence of operational (timing) errors is acceptable. This results in soft architectures — designs that fail gracefully, allowing us to perform reliability/power tradeoffs by reducing voltage up to the point that produces maximum allowable errors for our application. The goal of our optimization is to minimize the voltage at which a soft architecture encounters the maximum allowable error rate, thus maximizing the range over which voltage scaling is possible and minimizing power consumption for a given error rate. Our experiments demonstrate 23% power savings over the baseline design at an error rate of 1%. Observed power reductions are 29%, 29%, 19%, and 20% for error rates of 2%, 4%, 8%, and 16% respectively. Benefits are higher in the face of error recovery using Razor. Area overhead of our techniques is up to 2.7%. Andrew B. Kahng, Seokhyeong Kang, Rakesh Kumar 0002, John Sartori |
HPCA | 2 |