Jongho Yoon 0001

dblp:211/5332-1 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-9295-4850ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Leveraging Machine Learning Techniques for Traditional EDA Workflow Enhancement
abstract
As technology nodes advance and feature sizes shrink, the increasing complexity of design rules and routing congestion has resulted in greater design challenges and rising costs. Machine learning (ML) models offer significant potential to enhance design quality by enabling early prediction and optimization during the design flow. However, only a few works have validated the effectiveness of ML model when integrated to the traditional design flow. This paper will cover the effectiveness of ML-enhanced design workflow with some practical applications. Additionally, we will address which problems should be solved to achieve successful ML integration.
Jinoh Cho, Jaekyung Im, Kyungjun Min, Seonghyeon Park, Jaemin Seo, Jongho Yoon 0001, Seokhyeong Kang
ASP-DAC7
2025 ParaFormer: A Hybrid Graph Neural Network and Transformer Approach for Pre-Routing Parasitic RC Prediction
abstract
Predicting the quality of post-route design at an early stage can reduce overall design time. To achieve this, we propose ParaFormer, a pre-routing parasitic RC prediction framework. This framework integrates a heterogeneous graph neural network (HGNN) and a graph transformer to capture the topological and geometric information of circuit data. The HGNN model represents circuit data as heterogeneous graphs to learn complex topological relationships, while the graph transformer calculates attention between each net to learn geometric relationships. Our framework predicts parasitic RC, enabling RC tree modeling and SPEF file generation. This allows the predicted results to be utilized in timing and power analysis using commercial tools. Additionally, we incorporate gradient normalization to reduce the imbalance between different objectives in multi-task learning, improving overall model performance. Experimental results show that ParaFormer achieves R2 scores of 0.9901 and 0.9630 for resistance and capacitance, respectively. In timing analysis, it achieves R2 scores of 0.9749 for wire delay and 0.9876 for cell delay, with a MAPE of 1.45% in power analysis. These results indicate that our method is highly effective for timing and power prediction in the early design stage.
Jongho Yoon 0001, Jakang Lee, Junseok Hur, Seokhyeong Kang
ASP-DAC1
2025 Late Breaking Results: A Geometric Diffusion Model for Macro Placement Generation
abstract
Macro placement is crucial in VLSI design, directly impacting circuit performance. We introduce MacroDiff, a diffusion-based macro placement generative model that captures wirelength relationships instead of directly predicting macro coordinates. By leveraging wirelength as an intermediate representation, MacroDiff naturally preserves circuit connectivity, reduces placement constraints, and enhances solution flexibility while inherently handling rotational and translational invariances. Experiments on ISPD2005 benchmarks show that MacroDiff reduces macro overlap by 91.6%, lowers macro legalization displacement by 74.4%, and improves half-perimeter wirelength (HPWL) by 7.0%. While maintaining the efficiency of generative approaches, MacroDiff generates high-quality placements more reliably, narrowing the gap with state-of-the-art methods. The source code for this work is available at https://github.com/jhy00n/MacroDiff.
Jongho Yoon 0001, Jinsung Jeon, Seokhyeong Kang
DAC1
2024 Mobile Transformer Accelerator Exploiting Various Line Sparsity and Tile-Based Dynamic Quantization
abstract
Transformer models are difficult to employ in mobile devices due to their memory-and computation-intensive properties. Accordingly, there is ongoing research on various methods for compressing transformer models, such as pruning and quantization. However, general computing platforms such as central processing units (CPUs) and graphics processing units (GPUs) are not energy-efficient to accelerate the pruned model because the unstructured sparsity they exhibit causes degradation of parallelism. In this paper, we propose a low-power accelerator for transformers that can handle various levels of structured sparsity induced by line pruning with different granularity. Our approach accelerates pruned transformers in a head-wise and line-wise manner. We present a head reorganization and shuffling method that supports head-wise skip operations and resolves the load imbalance problem among processing engines (PEs) caused by the varying number of operations in each head. Furthermore, we implemented a sparse quantized general matrix-to-matrix multiplication (SQ-GEMM) module that supports line-wise skipping and on-the-fly tile-based dynamic quantization of activations. As a result, compared to mobile GPU and CPU, the proposed accelerator improved the energy efficiency by 2.9× and 12.3× for the detection transformer (DETR), and 3.0× and 12.4× for the vision transformer (ViT) models, respectively. In addition, our proposed mobile accelerator achieved the highest energy efficiency among the current state-of-the-art FPGA-based transformer accelerators.
Eunji Kwon, Jongho Yoon 0001, Seokhyeong Kang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 FPGA-Based Accelerator for Rank-Enhanced and Highly-Pruned Block-Circulant Neural Networks
abstract
Numerous network compression methods have been proposed to deploy deep neural networks in a resource-constrained embedded system. Among them, block-circulant matrix (BCM) compression is one of the promising hardware-friendly methods for both acceleration and compression. However, it has several limitations; (i) limited representation due to the structural characteristic of circulant matrix, (ii) limitation of the compression parameter, (iii) need to specialize the dataflow for BCM-compressed network accelerators. In this paper, rank-enhanced and highly-pruned block-circulant matrices compression (RP-BCM) framework is proposed to overcome these limitations. RP-BCM comprises two stages: Hadamard-BCM and BCM-wise pruning. Moreover, a dedicated skip scheme is introduced to processing element design for exploiting high-parallelism with BCM-wise sparsity. Furthermore, we propose specialized dataflow for a BCM-compressed network on a resource-constrained FPGA. As a result, the proposed method achieves parameter reduction and FLOPs reduction for ResNet-50 in ImageNet by 92.4% and 77.3%, respectively. Moreover, the proposed hardware design achieves$3.1\times$improvement in energy efficiency on the Xilinx PYNQ-Z2 FPGA board for ResNet-18 on ImageNet compared to the GPU.
Haena Song, Jongho Yoon 0001, Eunji Kwon, Tae-Hyun Oh, Seokhyeong Kang
DATE2
2022 Design and Evaluation Frameworks for Advanced RISC-based Ternary Processor
abstract
In this paper, we introduce the design and veri-fication frameworks for developing a fully-functional emerging ternary processor. Based on the existing compiling environments for binary processors, for the given ternary instructions, the software-level framework provides an efficient way to convert the given programs to the ternary assembly codes. We also present a hardware-level framework to rapidly evaluate the performance of a ternary processor implemented in arbitrary design technology. As a case study, the fully-functional 9-trit advanced RISC-based ternary (ART-9) core is newly developed by using the proposed frameworks. Utilizing 24 custom ternary instructions, the 5-stage ART-9 prototype architecture is successfully verified by a number of test programs including dhrystone benchmark in a ternary domain, achieving the processing efficiency of 57.8 DMIPS/W and$3.06\times 10^{6}$DMIPS/W in the FPGA-level ternary-logic emulations and the emerging CNTFET ternary gates, respectively.
Dongyun Kam, Jung Gyu Min, Jongho Yoon 0001, Sunmean Kim, Seokhyeong Kang, Youngjoo Lee 0002
DATE3
2021 Reinforcement Learning-Based Power Management Policy for Mobile Device Systems
abstract
This paper presents a power management policy that utilizes reinforcement learning to increase the power efficiency of mobile device systems based on a multiprocessor system-on-a-chip (MPSoC). The proposed policy predicts a system’s characteristics and learns power management controls to adapt to the variations in the system. We consider the behavioral characteristics of systems that run on mobile devices under diverse scenarios. Therefore, the policy can flexibly manage the system power regardless of the application scenario and achieve lower energy consumption without compromising the user satisfaction. The average energy per unit quality of service (QoS) of the proposed policy is lower than that of the previous six dynamic voltage/frequency scaling governors by 31.66%. Furthermore, we reduce the runtime overhead by implementing the proposed policy as hardware. We implemented the policy on the field programmable gate array (FPGA) and construct a communication interface between the central processing units (CPUs) and the hardware of the proposed policy. Decision-making by the hardware-implemented policy is 3.92 times faster than by the software-implemented policy.
Eunji Kwon, Sodam Han, Yoonho Park, Jongho Yoon 0001, Seokhyeong Kang
IEEE Trans. Circuits Syst. I Regul. Pap.4