Subeen Park

dblp:314/6049 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 67% Optimization for machine learning · 17% Question answering and dialogue systems · 17%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › robustness
distributionally robust optimization
0.912025
Sufficient Invariant Learning for Distribution Shift · CVPR 2025
Machine learning › Trustworthy machine learning › robustness
distribution shift
0.912025
Sufficient Invariant Learning for Distribution Shift · CVPR 2025
Natural language and speech › Question answering and dialogue systems
domain-specific question answering
0.912025
TIDES: Technical Information Discovery and Extraction System · EMNLP 2025
Machine learning › Trustworthy machine learning › out-of-distribution generalization
invariant learning
0.912025
Sufficient Invariant Learning for Distribution Shift · CVPR 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
Sufficient Invariant Learning for Distribution Shift · CVPR 2025
Machine learning › Optimization for machine learning › gradient-based optimization
sharpness-aware minimization
0.912025
Sufficient Invariant Learning for Distribution Shift · CVPR 2025
Storage systems
crash consistency
0.912025
Analyzing and Enhancing ArckFS: An Anecdotal Example of Benefits of Artifact Evaluation · SOSP 2025
Storage systems
file systems
0.912025
Analyzing and Enhancing ArckFS: An Anecdotal Example of Benefits of Artifact Evaluation · SOSP 2025
Storage systems › file systems › file system design
persistent memory file system
0.912025
Analyzing and Enhancing ArckFS: An Anecdotal Example of Benefits of Artifact Evaluation · SOSP 2025

Methods — techniques the papers use, named apart from their topics

prompt-based large language model · 1.7TF-IDF · 1.7sharpness-aware minimization · 0.9group distributionally robust optimization · 0.9artifact evaluation · 0.9
YearPublicationVenuePosition
2026 Enhancing LLMs for Manufacturing Information Extraction
Subeen Park, Hakyung Lee, Ryunyi Lee, Hyo-won Suh, Kyungwoo Song
PAKDD (4)1
2025 Sufficient Invariant Learning for Distribution Shift
abstract
Learning robust models under distribution shifts between training and test datasets is a fundamental challenge in machine learning. While learning invariant features across environments is a popular approach, it often assumes that these features are fully observed in both training and test sets—a condition frequently violated in practice. When models rely on invariant features absent in the test set, their robustness in new environments can deteriorate. To tackle this problem, we introduce a novel learning principle called the Sufficient Invariant Learning (SIL) framework, which focuses on learning a sufficient subset of invariant features rather than relying on a single feature. After demonstrating the limitation of existing invariant learning methods, we propose a new algorithm, Adaptive Sharpness-aware Group Distributionally Robust Optimization (ASGDRO), to learn diverse invariant features by seeking common flat minima across the environments. We theoretically demonstrate that finding a common flat minima enables robust predictions based on diverse invariant features. Empirical evaluations on multiple datasets, including our new benchmark, confirm ASGDRO’s robustness against distribution shifts, highlighting the limitations of existing methods. Code: https://github.com/MLAI-Yonsei/SIL-ASGDRO.
Taero Kim, Subeen Park, Sungjun Lim 0002, Yonghan Jung, Krikamol Muandet, Kyungwoo Song
CVPR2
2025 TIDES: Technical Information Discovery and Extraction System
abstract
Addressing the challenges in QA for specific technical domains requires identifying relevant portions of extensive documents and generating answers based on this focused content.Traditional pre-trained LLMs often struggle with domain-specific terminology, while fine-tuned LLMs demand substantial computational resources.To overcome these limitations, we propose TIDES, Technical Information Distillation and Extraction System.TIDES is a trainingfree approach that combines traditional TF-IDF techniques with prompt-based LLMs in a hybrid process, effectively addressing complex technical questions.It uses TF-IDF to identify and prioritize domain-specific words that are rare in other documents and LLMs to refine the candidate pool by focusing on the most relevant segments in documents through multiple stages.Our approach improves the precision and efficiency of QA systems in technical contexts without LLM retraining.
Subeen Park, Hakyung Lee, YongTaek Lim, Hyo-won Suh, Kyungwoo Song
EMNLP2
2025 Analyzing and Enhancing ArckFS: An Anecdotal Example of Benefits of Artifact Evaluation
abstract
We analyze and enhance Trio and ArckFS by Zhou et al. (SOSP 2023), high-performance NVM file system architecture and file system. A group of authors from KAIST initiated this study through a careful review of the paper and the released artifact, seeking to enhance the Trio work. Their analysis identifies (1) insufficient clarity in the paper on the handling of multi-inode operations, and (2) several implementation bugs in ArckFS that cause occasional operation failures or potential crash inconsistencies during inode creation.
Jonguk Jeon, Subeen Park, Sanidhya Kashyap, Sudarsun Kannan, Diyu Zhou, Jeehoon Kang
SOSP2