Sungrae Park

dblp:151/3635 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0001-5338-0113ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 ZERA: Zero-init Instruction Evolving Refinement Agent - From Zero Instructions to Structured Prompts via Principle-based Optimization
abstract
Automatic Prompt Optimization (APO) improves large language model (LLM) performance by refining prompts for specific tasks.However, prior APO methods typically focus only on user prompts, rely on unstructured feedback, and require large sample sizes and long iteration cycles-making them costly and brittle.We propose ZERA (Zero-init instruction Evolving Refinement Agent), a novel framework that jointly optimizes both system and user prompts through principled, low-overhead refinement.ZERA scores prompts using eight evaluation principles with automatically inferred weights, and revises prompts based on these structured critiques.This enables fast convergence to high-quality prompts using minimal examples and short iteration cycles.We evaluate ZERA across five LLMs and nine diverse datasets spanning reasoning, summarization, and code generation tasks.Experimental results demonstrate consistent improvements over strong baselines.Further ablation studies highlight the contribution of each component to more effective prompt construction.
Seungyoun Yi, Minsoo Khang, Sungrae Park
EMNLP3
2025 KIEval: Evaluation Metric for Document Key Information Extraction
Minsoo Khang, Sang Chul Jung, Sungrae Park, Teakgyu Hong
ICDAR (3)3
2024 SRFormer: Text Detection Transformer with Incorporated Segmentation and Regression
abstract
Existing techniques for text detection can be broadly classified into two primary groups: segmentation-based and regression-based methods. Segmentation models offer enhanced robustness to font variations but require intricate post-processing, leading to high computational overhead. Regression-based methods undertake instance-aware prediction but face limitations in robustness and data efficiency due to their reliance on high-level representations. In our academic pursuit, we propose SRFormer, a unified DETR-based model with amalgamated Segmentation and Regression, aiming at the synergistic harnessing of the inherent robustness in segmentation representations, along with the straightforward post-processing of instance-level regression. Our empirical analysis indicates that favorable segmentation predictions can be obtained at the initial decoder layers. In light of this, we constrain the incorporation of segmentation branches to the first few decoder layers and employ progressive regression refinement in subsequent layers, achieving performance gains while minimizing computational load from the mask. Furthermore, we propose a Mask-informed Query Enhancement module. We take the segmentation result as a natural soft-ROI to pool and extract robust pixel representations, which are then employed to enhance and diversify instance queries. Extensive experimentation across multiple benchmarks has yielded compelling findings, highlighting our method's exceptional robustness, superior training and data efficiency, as well as its state-of-the-art performance. Our code is available at https://github.com/retsuh-bqw/SRFormer-Text-Det.
Qingwen Bu, Sungrae Park, Minsoo Khang, Yichuan Cheng
AAAI2
2022 BROS: A Pre-trained Language Model Focusing on Text and Layout for Better Key Information Extraction from Documents
abstract
Key information extraction (KIE) from document images requires understanding the contextual and spatial semantics of texts in two-dimensional (2D) space. Many recent studies try to solve the task by developing pre-trained language models focusing on combining visual features from document images with texts and their layout. On the other hand, this paper tackles the problem by going back to the basic: effective combination of text and layout. Specifically, we propose a pre-trained language model, named BROS (BERT Relying On Spatiality), that encodes relative positions of texts in 2D space and learns from unlabeled documents with area-masking strategy. With this optimized training scheme for understanding texts in 2D space, BROS shows comparable or better performance compared to previous methods on four KIE benchmarks (FUNSD, SROIE*, CORD, and SciTSR) without relying on visual features. This paper also reveals two real-world challenges in KIE tasks--(1) minimizing the error from incorrect text ordering and (2) efficient learning from fewer downstream examples--and demonstrates the superiority of BROS over previous methods.
Teakgyu Hong, Donghyun Kim 0012, Mingi Ji, Wonseok Hwang, Daehyun Nam, Sungrae Park
AAAI6
2022 Domain Generalization by Mutual-Information Regularization with Pre-trained Models
Junbum Cha, Kyungjae Lee 0001, Sungrae Park, Sanghyuk Chun
ECCV (23)3
2022 Multi-modal Text Recognition Networks: Interactive Enhancements Between Visual and Semantic Features
Byeonghu Na, Yoonsik Kim, Sungrae Park
ECCV (28)3
2022 Contrastive Learning for Knowledge Tracing
abstract
Knowledge tracing is the task of understanding student’s knowledge acquisition processes by estimating whether to solve the next question correctly or not. Most deep learning-based methods tackle this problem by identifying hidden representations of knowledge states from learning histories. However, due to the sparse interactions between students and questions, the hidden representations can be easily over-fitted and often fail to capture student’s knowledge states accurately. This paper introduces a contrastive learning framework for knowledge tracing that reveals semantically similar or dissimilar examples of a learning history and stimulates to learn their relationships. To deal with the complexity of knowledge acquisition during learning, we carefully design the components of contrastive learning, such as architectures, data augmentation methods, and hard negatives, taking into account pedagogical rationales. Our extensive experiments on six benchmarks show statistically significant improvements from the previous methods. Further analysis shows how our methods contribute to improving knowledge tracing performances.
Wonsung Lee, Jaeyoon Chun, Youngmin Lee, Kyoungsoo Park, Sungrae Park
WWW5
2021 Show, Attend and Distill: Knowledge Distillation via Attention-based Feature Matching
abstract
Knowledge distillation extracts general knowledge from a pretrained teacher network and provides guidance to a target student network. Most studies manually tie intermediate features of the teacher and student, and transfer knowledge through predefined links. However, manual selection often constructs ineffective links that limit the improvement from the distillation. There has been an attempt to address the problem, but it is still challenging to identify effective links under practical scenarios. In this paper, we introduce an effective and efficient feature distillation method utilizing all the feature levels of the teacher without manually selecting the links. Specifically, our method utilizes an attention based meta network that learns relative similarities between features, and applies identified similarities to control distillation intensities of all possible pairs. As a result, our method determines competent links more efficiently than the previous approach and provides better performance on model compression and transfer learning tasks. Further qualitative analyses and ablative studies describe how our method contributes to better distillation.
Mingi Ji, Byeongho Heo, Sungrae Park
AAAI3
2021 SynthTIGER: Synthetic Text Image GEneratoR Towards Better Text Recognition Models
Moonbin Yim, Yoonsik Kim, Hancheol Cho, Sungrae Park
ICDAR (4)4
2021 SWAD: Domain Generalization by Seeking Flat Minima
abstract
Domain generalization (DG) methods aim to achieve generalizability to an unseen target domain by using only training data from the source domains. Although a variety of DG methods have been proposed, a recent study shows that under a fair evaluation protocol, called DomainBed, the simple empirical risk minimization (ERM) approach works comparable to or even outperforms previous methods. Unfortunately, simply solving ERM on a complex, non-convex loss function can easily lead to sub-optimal generalizability by seeking sharp minima. In this paper, we theoretically show that finding flat minima results in a smaller domain generalization gap. We also propose a simple yet effective method, named Stochastic Weight Averaging Densely (SWAD), to find flat minima. SWAD finds flatter minima and suffers less from overfitting than does the vanilla SWA by a dense and overfit-aware stochastic weight sampling strategy. SWAD shows state-of-the-art performances on five DG benchmarks, namely PACS, VLCS, OfficeHome, TerraIncognita, and DomainNet, with consistent and large margins of +1.6% averagely on out-of-domain accuracy. We also compare SWAD with conventional generalization methods, such as data augmentation and consistency regularization methods, to verify that the remarkable performance improvements are originated from by seeking flat minima, not from better in-domain generalizability. Last but not least, SWAD is readily adaptable to existing DG methods without modification; the combination of SWAD and an existing DG method further improves DG performances. Source code is available at https://github.com/khanrc/swad.
Junbum Cha, Sanghyuk Chun, Hancheol Cho, Seunghyun Park 0001, Yunsung Lee, Sungrae Park
NeurIPS7
2020 Scale down Transformer by Grouping Features for a Lightweight Character-level Language Model
abstract
This paper introduces a method that efficiently reduces the computational cost and parameter size of Transformer.The proposed model, refer to as Group-Transformer, splits feature space into multiple groups, factorizes the calculation paths, and reduces computations for the group interaction.Extensive experiments on two benchmark tasks, enwik8 and text8, prove our model's effectiveness and efficiency in small-scale Transformers.To the best of our knowledge, Group-Transformer is the first attempt to design Transformer with the group strategy, widely used for efficient CNN architectures.
Sungrae Park, Geewook Kim, Junyeop Lee, Junbum Cha, Jihoon Kim 0002, Hwalsuk Lee
COLING1
2020 Character Region Attention for Text Spotting
Youngmin Baek, Seung Shin, Jeonghun Baek, Sungrae Park, Junyeop Lee, Daehyun Nam, Hwalsuk Lee
ECCV (29)4
2020 Dirichlet Variational Autoencoder
Weonyoung Joo, Wonsung Lee, Sungrae Park, Il-Chul Moon
Pattern Recognit.3
2019 Adversarial Dropout for Recurrent Neural Networks
Sungrae Park, Kyungwoo Song, Mingi Ji, Wonsung Lee, Il-Chul Moon
AAAI1
2019 Hierarchical Context Enabled Recurrent Neural Network for Recommendation
abstract
A long user history inevitably reflects the transitions of personal interests over time. The analyses on the user history require the robust sequential model to anticipate the transitions and the decays of user interests. The user history is often modeled by various RNN structures, but the RNN structures in the recommendation system still suffer from the long-term dependency and the interest drifts. To resolve these challenges, we suggest HCRNN with three hierarchical contexts of the global, the local, and the temporary interests. This structure is designed to withhold the global long-term interest of users, to reflect the local sub-sequence interests, and to attend the temporary interests of each transition. Besides, we propose a hierarchical context-based gate structure to incorporate our interest drift assumption. As we suggest a new RNN structure, we support HCRNN with a complementary bi-channel attention structure to utilize hierarchical context. We experimented the suggested structure on the sequential recommendation tasks with CiteULike, MovieLens, and LastFM, and our model showed the best performances in the sequential recommendations.
Kyungwoo Song, Mingi Ji, Sungrae Park, Il-Chul Moon
AAAI3
2019 What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model Analysis
abstract
Many new proposals for scene text recognition (STR) models have been introduced in recent years. While each claim to have pushed the boundary of the technology, a holistic and fair comparison has been largely missing in the field due to the inconsistent choices of training and evaluation datasets. This paper addresses this difficulty with three major contributions. First, we examine the inconsistencies of training and evaluation datasets, and the performance gap results from inconsistencies. Second, we introduce a unified four-stage STR framework that most existing STR models fit into. Using this framework allows for the extensive evaluation of previously proposed STR modules and the discovery of previously unexplored module combinations. Third, we analyze the module-wise contributions to performance in terms of accuracy, speed, and memory demand, under one consistent set of training and evaluation datasets. Such analyses clean up the hindrance on the current comparisons to understand the performance gain of the existing modules. Our code is publicly available.
Jeonghun Baek, Geewook Kim, Junyeop Lee, Sungrae Park, Dongyoon Han, Sangdoo Yun, Seong Joon Oh, Hwalsuk Lee
ICCV4
2018 Adversarial Dropout for Supervised and Semi-Supervised Learning
abstract
Recently, training with adversarial examples, which are generated by adding a small but worst-case perturbation on input examples, has improved the generalization performance of neural networks. In contrast to the biased individual inputs to enhance the generality, this paper introduces adversarial dropout, which is a minimal set of dropouts that maximize the divergence between 1) the training supervision and 2) the outputs from the network with the dropouts. The identified adversarial dropouts are used to automatically reconfigure the neural network in the training process, and we demonstrated that the simultaneous training on the original and the reconfigured network improves the generalization performance of supervised and semi-supervised learning tasks on MNIST, SVHN, and CIFAR-10. We analyzed the trained model to find the performance improvement reasons. We found that adversarial dropout increases the sparsity of neural networks more than the standard dropout. Finally, we also proved that adversarial dropout is a regularization term with a rank-valued hyper-parameter that is different from a continuous-valued parameter to specify the strength of the regularization.
Sungrae Park, Jun-Keon Park, Su-Jin Shin, Il-Chul Moon
AAAI1
2018 Diagnosis Prediction via Medical Context Attention Networks Using Deep Generative Modeling
abstract
Predicting the clinical outcome of patients from the historical electronic health records (EHRs) is a fundamental research area in medical informatics. Although EHRs contain various records associated with each patient, the existing work mainly dealt with the diagnosis codes by employing recurrent neural networks (RNNs) with a simple attention mechanism. This type of sequence modeling often ignores the heterogeneity of EHRs. In other words, it only considers historical diagnoses and does not incorporate patient demographics, which correspond to clinically essential context, into the sequence modeling. To address the issue, we aim at investigating the use of an attention mechanism that is tailored to medical context to predict a future diagnosis. We propose a medical context attention (MCA)-based RNN that is composed of an attention-based RNN and a conditional deep generative model. The novel attention mechanism utilizes the derived individual patient information from conditional variational autoencoders (CVAEs). The CVAE models a conditional distribution of patient embeddings and his/her demographics to provide the measurement of patient's phenotypic difference due to illness. Experimental results showed the effectiveness of the proposed model.
Wonsung Lee, Sungrae Park, Weonyoung Joo, Il-Chul Moon
ICDM2
2018 Hierarchical prescription pattern analysis with symptom labels
Su-Jin Shin, Je-Yong Oh, Sungrae Park, Il-Chul Moon
Pattern Recognit. Lett.3
2017 Identifying prescription patterns with a topic model of diseases and medications
Sungrae Park, Doosup Choi, Won Chul Cha, Chu Hyun Kim, Il-Chul Moon
J. Biomed. Informatics1
2015 Associative topic models with numerical time series
Sungrae Park, Wonsung Lee, Il-Chul Moon
Inf. Process. Manag.1
2015 Efficient extraction of domain specific sentiment lexicon with active learning
Sungrae Park, Wonsung Lee, Il-Chul Moon
Pattern Recognit. Lett.1
2014 Disease-medicine topic model for prescription record mining
abstract
Analyzing patient records is important for improving the quality of medical services and for understanding each patient's historical diseases. However, the huge size of the data requires statistical analysis procedures. In this paper, we proposed a probabilistic model-the disease-medicine topic model (DMTM)-to explore connected knowledge about diseases and medicines. In the model, diseases and medicines are modeled using generative process. We used the latent Dirichlet allocation, which is one of the most popular topic models, as the baseline model. Then, we compared the qualities of topic representations quantitatively and qualitatively. The comparison results showed that the topics derived from the DMTM are clearer to identify and that specific patterns were found in the diseases and medicines. In the case of topic network analysis, these specific patterns were proved using centrality measurements.
Sungrae Park, Doosup Choi, Wonsung Lee, Dain Jung, Il-Chul Moon
SMC1