Seonghyeon Park

dblp:258/6927 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EpiCaR: Knowing What You Don't Know Matters for Better Reasoning in LLMs
abstract
Improving the reasoning abilities of large language models (LLMs) has largely relied on iterative self-training with model-generated data.While effective at boosting accuracy, existing approaches primarily reinforce successful reasoning paths, incurring a substantial calibration cost: models become overconfident and lose the ability to represent uncertainty.This failure has been characterized as a form of model collapse in alignment, where predictive distributions degenerate toward low-variance point estimates.We address this issue by reframing open-ended reasoning training as an epistemic learning problem, in which models must learn not only how to reason, but also when their reasoning should be trusted.We propose epistemically-calibrated reasoning (EPICAR) as a training objective that jointly optimizes reasoning performance and calibration, and instantiate it within an iterative supervised fine-tuning framework using explicitly extracted meta-cognitive self-evaluation signals.Experiments on Llama-3 and Qwen-3 families demonstrate that our approach achieves Pareto-superiority over standard baselines in both accuracy and calibration, particularly in models with sufficient reasoning capacity (e.g., 3B+).This framework generalizes effectively to OOD mathematical reasoning (GSM8K) and code generation (MBPP).Ultimately, our approach enables a 3× reduction in the overall inference compute budget, matching the K = 30 majority-vote performance of STaR with only K = 10 confidence-weighted samples, entirely without the multi-model overhead of external verifiers.
Je Won Yeom, Jaewon Sok, Seonghyeon Park, Jeongjae Park, Taesup Kim
ACL (1)3
2026 Invited: Post-Placement Buffering and Sizing Contest
abstract
The ISPD 2026 Contest [22] challenges participants to develop post-detailed placement buffering and sizing tools that optimize timing and fix electrical rule check (ERC) violations under real-world constraints. Unlike prior contests, this contest emphasizes practical physical design challenges including fixed macros and I/Os, power delivery network (PDN) blockages, soft placement blockages, and fixed routing resources. The contest provides eight public benchmarks and four hidden benchmarks, with a range from 15K to 1.4M instances, in the ASAP7 7nm technology node [4] with multi-threshold voltage cell libraries. Evaluation is performed using the open-source OpenROAD infrastructure, with scoring based on timing (total negative slack), power (dynamic and leakage) and penalties for ERC violations, displacement, routing congestion and runtime. This paper describes the contest problem formulation, benchmarks, evaluation methodology, a review of related contests and a two-year roadmap for continuation in the ISPD 2027 Contest.
Andrew B. Kahng, Seokhyeong Kang, Sayak Kundu, Yiting Liu 0002, Davit Markarian, Seonghyeon Park, Zhiang Wang
ISPD6
2025 Leveraging Machine Learning Techniques for Traditional EDA Workflow Enhancement
abstract
As technology nodes advance and feature sizes shrink, the increasing complexity of design rules and routing congestion has resulted in greater design challenges and rising costs. Machine learning (ML) models offer significant potential to enhance design quality by enabling early prediction and optimization during the design flow. However, only a few works have validated the effectiveness of ML model when integrated to the traditional design flow. This paper will cover the effectiveness of ML-enhanced design workflow with some practical applications. Additionally, we will address which problems should be solved to achieve successful ML integration.
Jinoh Cho, Jaekyung Im, Kyungjun Min, Seonghyeon Park, Jaemin Seo, Jongho Yoon 0001, Seokhyeong Kang
ASP-DAC5
2025 FedEDA: Federated Learning Framework for Privacy-Preserving Machine Learning in EDA
abstract
In advanced nodes, the optimization of Power, Performance, and Area (PPA) is becoming increasingly complex, requiring significant resources and time for circuit design and optimization using Electronic Design Automation (EDA). As a key approach to overcome these challenges, Machine Learning (ML) techniques have been widely studied in the field of EDA. However, security concerns around Intellectual Property (IP) limit access to real-world circuit data, making it difficult to gather sufficient data for training ML models. This lack of available circuit benchmarks restricts progress in ML research. In this study, we propose FedEDA, which, to the best of our knowledge, is the first Federated Learning (FL) aggregation algorithm specifically designed for EDA. FedEDA addresses concerns about IP security by exchanging model weights among FL participants instead of sharing raw data. Furthermore, FedEDA leverages Rent’s Rule and circuit size to capture the hierarchical structure of circuits, mitigating issues related to data imbalance among participants and improving the quality of weight aggregation on EDA data. We demonstrate the applicability of FedEDA across various EDA tasks, including routability, parasitic RC, and wirelength prediction. FedEDA outperforms existing FL algorithms in EDA tasks, demonstrating superior performance.
Seonghyeon Park, Seokhyeong Kang
DAC3
2025 Late Breaking Results: Fine-Tuning LLMs for Test Stimuli Generation
abstract
The understanding and reasoning capabilities of large language models (LLMs) with text data have made them widely used for test stimuli generation. Existing studies have primarily focused on methods such as prompt engineering or providing feedback to the LLMs’ generated outputs to improve test stimuli generation. However, these approaches have not been successful in enhancing the LLMs’ domain-specific performance in generating test stimuli. In this paper, we introduce a framework for finetuning LLMs for test stimuli generation through dataset generation and reinforcement learning (RL). Our dataset generation approach creates a table-shaped test stimuli dataset, which helps ensure that the LLM produces consistent outputs. Additionally, our two-stage fine-tuning process involves training the LLMs on domain-specific data and using RL to provide feedback on the generated outputs, further enhancing the LLMs’ performance in test stimuli generation. Experimental results confirm that our framework improves syntax correctness and code coverage of test stimuli, outperforming commercial models.
Hyeonwoo Park, Seonghyeon Park, Seokhyeong Kang
DAC2
2025 Improving LLM-Based Verilog Code Generation with Data Augmentation and RL
abstract
Large language models (LLMs) have recently attracted significant attention for their potential in Verilog code generation. However, existing LLM-based methods face several challenges, including data scarcity and the high computational cost of generating prompts for fine-tuning. Motivated by these challenges, we explore methods to augment training datasets, develop more efficient and effective prompts for fine-tuning, and implement training methods incorporating electronic design automation (EDA) tools. Our proposed framework for fine-tuning LLMs for Verilog code generation includes (1) abstract syntax tree (AST)-based data augmentation, (2) output-relevant code masking, a prompt generation method based on the logical structure of Verilog code, and (3) reinforcement learning with tool feedback (RLTF), a fine-tuning method using EDA tool results. Experimental studies confirm that our framework significantly improves syntax and functional correctness, outperforming commercial and non-commercial models on open-source benchmarks.
Kyungjun Min, Seonghyeon Park, Hyeonwoo Park, Jinoh Cho, Seokhyeong Kang
DATE2
2024 PPA-Relevant Clustering-Driven Placement for Large-Scale VLSI Designs
abstract
Today's place-and-route (P&R) flows are increasingly challenged by complexity and scale of modern designs. Often, heuristics must trade off between turnaround time and quality of PPA outcomes. This paper presents a clustered placement methodology that improves both turnaround time and final-routed solution quality. Our PPA-aware clustering considers timing, power and logical hierarchy during netlist clustering, effectively reducing problem size and accelerating global placement runtime while improving post-route PPA metrics. Additionally, our machine learning (ML)-accelerated virtualized P&R methodology predicts the best cluster shapes (i.e., aspect ratios and utilizations) to use in P&R of the clustered netlist. With the open-source OpenROAD tool, our methods achieve up to 47% (average: 36%) global placement runtime improvement with similar half-perimeter wirelength (HPWL) and 90% (29%) improvement in post-route total negative slack (TNS). With the commercial Cadence Innovus tool, our methods achieve up to 3.92% (1%) improvement in power and 99% (49%) improvement in TNS.
Andrew B. Kahng, Seokhyeong Kang, Sayak Kundu, Kyungjun Min, Seonghyeon Park, Bodhisatta Pramanik
DAC5
2024 AI-Based Mental Health Assessment for Adolescents Using Their Daily Digital Activities
abstract
Adolescents and their parents hesitate to acknowledge mental health issues until symptoms severely worsen, making timely treatment challenging. Moreover, infrequent psychiatric consultations often fail to adjust treatments to the dynamic nature of mental health states. To address these issues, our paper proposes an AI-based mental health assessment framework for adolescent mental health through non-invasively collected data from daily digital activities on their mobile devices, including tablets and smartphones. For this, we collect fifteen different types of passive sensor data across three primary categories of activities: studying, smartphone using, and metaverse gaming. Additionally, each adolescent completes self-survey reports on eight different disorders which are used as labels. Then, feature extraction is conducted based on this dataset, which yields 1,523 features that could function as potential digital biomarkers of mental health conditions in adolescents. Utilizing these features, our algorithm named CAMP: Customizable Automated Machine learning Process incorporates simulated annealing for feature selection. This approach enables the construction of AI models for mental health assessment that are finely tuned to domain specific strategies. Our experiments show that our proposed framework can significantly improve models' performance.
Joonsung Lee, Taehwi Lee, Soeun Baek, Seonghyun Jin, Haeun Yoo, Youngeun Cho, Seonghyeon Park, Kwangsu Cho, Chang-Gun Lee
DSAA8
2024 RL-Fill: Timing-Aware Fill Insertion using Reinforcement Learning
abstract
We introduce RL-Fill, a novel reinforcement learning framework for timing-aware fill insertion. RL-Fill first generates a large number of fills in the empty spaces and then removes the timing-critical fills as determined by the policy network. Towards faster convergence and stability, our framework employs a two-phase training process. In the first phase, we train the policy with offline expert data using an imitation learning scheme. In the second phase, we further optimize the policy with online data using reinforcement learning. Moreover, we propose a new data augmentation method, LayoutMix, to ensure data-efficient training despite limited number of expert data. Our results demonstrate that RL-Fill is competitive to the commercial tool and outperforms the previous machine learning-based method in timing metrics while adhering density constraints.
Jinoh Cho, Seonghyeon Park, Jakang Lee, Sung-Yun Lee, Jinmo Ahn, Seokhyeong Kang
ICCAD2
2023 RL-Legalizer: Reinforcement Learning-based Cell Priority Optimization in Mixed-Height Standard Cell Legalization
abstract
Cell legalization order has a substantial effect on the quality of modern VLSI designs, which use mixed-height standard cells. In this paper, we propose a deep reinforcement learning framework to optimize cell priority in the legalization phase of various designs. We extract the selected features of movable cells and their surroundings, then embed them into cell-wise deep neural networks. We then determine cell priority and legalize them in order using a pixel-wise search algorithm. The proposed framework uses a policy gradient algorithm and several training techniques, including grid-cell subepisode, data normalization, reduced-dimensional state, and network optimization. We aim to resolve the suboptimality of existing sequential legalization algorithms with respect to displacement and wirelength. On average, our proposed framework achieved 34% lower legalization costs in various benchmarks compared to that of the state-of-the-art legalization algorithm.
Sung-Yun Lee, Seonghyeon Park, Minjae Kim 0005, Le Pham Tuyen 0001, Seokhyeong Kang
DATE2
2023 Routability Prediction and Optimization Using Explainable AI
abstract
Machine learning (ML) techniques have been widely studied to predict routability in early-stage. To reduce the design turn-around time during the placement and routing iterations, it is crucial to predict the design rule violation (DRV) hotspots precisely before actual detailed routing. However, complex network architectures of ML make it challenging for humans to understand how ML generates predictions and to identify the factors that significantly influence the predictions. This black-box nature of ML limits the efficient integration of the prediction techniques into an optimization process. Explainable artificial intelligence enables the interpretation of decision rationales in the ML model and brings us the reasons underlying the prediction of the model. In this paper, we propose a routability optimization framework that analyzes the input features relevant to the predicted DRV hotspots using an explainable model and selects the most suitable optimization methods. The proposed framework comprises three steps - (1) predicting DRV hotspots in the early-global routing stage, (2) calculating how much each input feature contributes to the predictions and (3) applying a proper optimization method to improve the routability. We reduced the number of DRVs by 78% on average in 16 design layouts without degrading the design Quality.
Seonghyeon Park, Seongbin Kwon, Seokhyeong Kang
ICCAD1
2023 Multi-Source Transfer Learning for Design Technology Co-Optimization
abstract
In advanced technology nodes, pitch scaling have not kept up with the Moore's Law. To continue progression, the design technology co-optimization (DTCO) has been proposed. However, implementing DTCO requires significant time cost and resources due to iterative trials. In addition, optimal design and technology option depend on each design, thus it should start from scratch whenever the target design changes. We present a DTCO framework based on Bayesian optimization that efficiently explores design feedback for optimization. In addition, our framework incorporates a multi-source transfer Gaussian process (MTGP) that ensures robust optimization even for unseen designs. MTGP significantly improves prediction and generalization performance by integrating multiple single source transfer Gaussian processes. Our framework, on average, reduced the mean absolute error of power and area by 47.3% and 24.1%, respectively, and power and area by 37.3% and 19.9%, respectively, compared to the reference, in 7nm technology nodes.
Jakang Lee, Seonghyeon Park, Seokhyeong Kang
ISLPED3
2022 Guaranteeing Safety Despite Physical Errors in Cyber-Physical Systems
abstract
This paper considers a cyber-physical system with a so-called “self-looping” node that repeats the inner-loop for physical situation awareness, i.e., more loops for more harsh physical situations. Regarding such a self-looping node, we observe the existence of physical errors that make the looping useless and eventually cause a critical failure. To prevent such a critical failure despite a physical error, this paper proposes a novel mechanism by introducing “time wall” and “safety backup”. The time wall limits the time budget for the self-looping node so as to switch to the safety backup while still meeting the deadline to prevent critical failure despite physical errors. Our experiments through both simulation and actual implementation show that the proposed mechanism gives a comparable accuracy with the existing methods in normal cases while completely preventing the critical failure in physical error cases.
Jongwoo Han, Seonghyeon Park, Haejoo Jeon, Chang-Gun Lee
RTAS2