EDBT 2026 Demo / reviewers in the wild / expert
Sungwon Han 0001
dblp:72/5688-1
· DBLP profile ↗
24ranked-venue papers
10as first author
19since 2021 · last 2026
0000-0002-1129-760XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 9 first-author · 13 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Denoising attribution maps through gradient analysis of critical parametersabstractMany post-hoc explainable feature attribution techniques analyze gradient propagation and understand decisions of deep learning models. However, conventional gradient analysis used to generate such maps can be noisy, potentially compromising the reliability of explanations. In this work, we introduce a robust method for computing feature attributions by identifying critical parameters and refining gradient propagation through these parameters. This method reduces the impact of non-critical parameters, mitigating the effect of feature leakage and randomized initialization, which introduce noise in attribution maps. We implemented this concept as an add-on module, called CriGrad , and evaluated its efficacy using three benchmarks and seven explainable models. Our results show that focusing on critical parameters improves explanability in 93% of the cases, demonstrating its effectiveness and improved reliability. SeungEon Lee 0001, Heejin Bin, Sungwon Han 0001, Meeyoung Cha |
Pattern Recognit. | 3 |
| 2025 | Generalizable Disaster Damage Assessment via Change Detection with Vision Foundation ModelabstractThe increasing frequency and intensity of natural disasters call for rapid and accurate damage assessment. In response, disaster benchmark datasets from high-resolution satellite imagery have been constructed to develop methods for detecting damaged areas. However, these methods face significant challenges when applied to previously unseen regions due to the limited geographical and disaster-type diversity in the existing datasets. We introduce DAVI (Disaster Assessment with VIsion foundation model), a novel approach that addresses domain disparities and detects structural damage at the building level without requiring ground-truth labels for target regions. DAVI combines task-specific knowledge from a model trained on source regions with task-agnostic knowledge from an image segmentation model to generate pseudo labels indicating potential damage in target regions. It then utilizes a two-stage refinement process, which operate at both pixel and image levels, to accurately identify changes in disaster-affected areas. Our evaluation, including a case study on the 2023 Türkiye earthquake, demonstrates that our model achieves exceptional performance across diverse terrains (e.g., North America, Asia, and the Middle East) and disaster types (e.g., wildfires, hurricanes, and tsunamis). This confirms its robustness in disaster assessment without dependence on ground-truth labels and highlights its practical applicability. Kyeongjin Ahn, Sungwon Han 0001, Sungwon Park 0001, Meeyoung Cha |
AAAI | 2 |
| 2025 | Measuring Fine-Grained Urban Air Temperature with Satellite ImageryabstractRecent studies on the urban heat island phenomenon reveal how rapid urbanization intensifies temperature disparities in urban cores, highlighting the need for sustainable urban planning solutions. Analyzing the problems caused by these effects requires high-resolution climate data; however, physical weather stations often lack sufficient regional coverage and resolution. Proposals for alternative methods have attempted to bridge this gap, but they fall short in capturing regional characteristics adequately or necessitate obtaining difficult-to-get input data. This research proposes to use satellite data, where the visual spectrum provides rich information about the degree of human development and is easy to obtain, to measure urban air temperature. Our model, UrbanHeat, uses multi-resolution satellite imagery and employs land surface temperature and global climate data as proxy labels to predict air temperature at a granular scale. The results show that the model provides predictions at a much finer scale while showing superior performance in measuring ordinal relationships between points by capturing both local and broad land cover details of the region. Our case studies demonstrate how predictions at high resolution can help protect vulnerable populations from extreme heat (e.g., elders or developing countries) and contribute to sustainable urban development worldwide. Minhyuk Song, Sungwon Han 0001, SeungEon Lee 0001, Donghyun Ahn, Meeyoung Cha |
AAAI | 2 |
| 2025 | Retrieval Augmented Time Series ForecastingabstractTime series forecasting uses historical data to predict future trends, leveraging the relationships between past observations and available features. In this paper, we propose RAFT, a retrieval-augmented time series forecasting method to provide sufficient inductive biases and complement the model’s learning capacity. When forecasting the subsequent time frames, we directly retrieve historical data candidates from the training dataset with patterns most similar to the input, and utilize the future values of these candidates alongside the inputs to obtain predictions. This simple approach augments the model’s capacity by externally providing information about past patterns via retrieval modules. Our empirical evaluations on ten benchmark datasets show that RAFT consistently outperforms contemporary baselines with an average win ratio of 86%. Sungwon Han 0001, SeungEon Lee 0001, Meeyoung Cha, Sercan Ö. Arik, Jinsung Yoon |
ICML | 1 |
| 2025 | Adversarial Style Augmentation via Large Language Model for Robust Fake News DetectionabstractThe spread of fake news harms individuals and presents a critical social challenge that must be addressed. Although numerous algorithmic and insightful features have been developed to detect fake news, many of these features can be manipulated with style-conversion attacks, especially with the emergence of advanced language models, making it more difficult to differentiate from genuine news. This study proposes adversarial style augmentation, AdStyle, designed to train a fake news detector that remains robust against various style-conversion attacks. The primary mechanism involves the strategic use of LLMs to automatically generate a diverse and coherent array of style-conversion attack prompts, enhancing the generation of particularly challenging prompts for the detector. Experiments indicate that our augmentation strategy significantly improves robustness and detection performance when evaluated on fake news benchmark datasets. Sungwon Park 0001, Sungwon Han 0001, Xing Xie 0001, Jae-Gil Lee 0001, Meeyoung Cha |
WWW | 2 |
| 2025 | Enhancing Domain Generalization for Robust Machine-Generated Text DetectionabstractLarge language models have revolutionized text generation, offering significant benefits while also posing threats to society, such as copyright infringement and misinformation. To prevent harmful use, the task of detecting machine-generated content has become an important research topic, though it remains particularly challenging across diverse content domains. This paper presents DGRM, an innovative add-on module designed to improve the domain generalization capability of existing machine-generated text detectors. Our model consists of two training components. (1) Feature disentanglement separates a text's embedding into target-specific and common attributes, thereby enhancing semantic domain generalization across different content domains. (2) Feature regularization applies constraints to these attributes to extract additional target-relevant information and ensure detection consistency under syntactic perturbations—thus achieving syntactic domain generalization. Evaluation over multiple datasets demonstrates that incorporating our module substantially improves the detection of machine-generated text across semantically and syntactically diverse domains. We hope our work contributes to mitigating the harmful use of language models. Sungwon Park 0001, Sungwon Han 0001, Meeyoung Cha |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | Large Language Models Can Automatically Engineer Features for Few-Shot Tabular LearningabstractLarge Language Models (LLMs), with their remarkable ability to tackle challenging and unseen reasoning problems, hold immense potential for tabular learning, that is vital for many real-world applications. In this paper, we propose a novel in-context learning framework, FeatLLM, which employs LLMs as feature engineers to produce an input data set that is optimally suited for tabular predictions. The generated features are used to infer class likelihood with a simple downstream machine learning model, such as linear regression and yields high performance few-shot learning. The proposed FeatLLM framework only uses this simple predictive model with the discovered features at inference time. Compared to existing LLM-based approaches, FeatLLM eliminates the need to send queries to the LLM for each sample at inference time. Moreover, it merely requires API-level access to LLMs, and overcomes prompt size limitations. As demonstrated across numerous tabular datasets from a wide range of domains, FeatLLM generates high-quality rules, significantly (10% on average) outperforming alternatives such as TabLLM and STUNT. Sungwon Han 0001, Jinsung Yoon, Sercan Ö. Arik, Tomas Pfister |
ICML | 1 |
| 2024 | Uncertainty-Aware Face Embedding With Contrastive Learning for Open-Set EvaluationabstractWhile advances in deep learning have enabled novel applications in various fields, face recognition in open-set scenarios remains a complex task, owing to the challenges posed by the extensive volume of low-quality face images. We introduce a new approach for recognizing faces in unconstrained open-set settings by leveraging uncertainty-aware embeddings through contrastive learning. Our model, called UCFace, effectively regulates the contribution of each face image based on the face uncertainty derived from image quality as an inverse proxy. Face embeddings are reinterpreted as a probabilistic distribution within the embedding space, where the degree of sharpness (i.e., distribution concentration) reflects the underlying uncertainty and probability density is used as a similarity metric to facilitate contrastive learning. Experiments on a wide range of face datasets, including those with high, mixed, and real-world low-resolution face images, demonstrate that UCFace enhances open-set face recognition performance by integrating the aspect of uncertainty. Kyeongjin Ahn, SeungEon Lee 0001, Sungwon Han 0001, Cheng-Yaw Low, Meeyoung Cha |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | Unified Neural Topic Model via Contrastive Learning and Term WeightingabstractTwo types of topic modeling predominate: generative methods that employ probabilistic latent models and clustering methods that identify semantically coherent groups.This paper newly presents UTopic (Unified neural Topic model via contrastive learning and term weighting) that combines the advantages of these two types.UTopic uses contrastive learning and term weighting to learn knowledge from a pretrained language model and discover influential terms from semantically coherent clusters.Experiments show that the generated topics have a high-quality topic-word distribution in terms of topic coherence, outperforming existing baselines across multiple topic coherence measures.We demonstrate how our model can be used as an add-on to existing topic models and improve their performance. Sungwon Han 0001, Mingi Shin, Sungkyu Park, Changwook Jung, Meeyoung Cha |
EACL | 1 |
| 2023 | Towards Attack-tolerant Federated Learning via Critical Parameter AnalysisabstractFederated learning is used to train a shared model in a decentralized way without clients sharing private data with each other. Federated learning systems are susceptible to poisoning attacks when malicious clients send false updates to the central server. Existing defense strategies are ineffective under non-IID data settings. This paper proposes a new defense strategy, FedCPA (Federated learning with Critical Parameter Analysis). Our attack-tolerant aggregation method is based on the observation that benign local models have similar sets of top-k and bottom-k critical parameters, whereas poisoned local models do not. Experiments with different attack scenarios on multiple datasets demonstrate that our model outperforms existing defense strategies in defending against poisoning attacks. Sungwon Han 0001, Sungwon Park 0001, Fangzhao Wu, Sundong Kim, Bin B. Zhu, Xing Xie 0001, Meeyoung Cha |
ICCV | 1 |
| 2023 | FedDefender: Client-Side Attack-Tolerant Federated LearningabstractFederated learning enables learning from decentralized data sources without compromising privacy, which makes it a crucial technique. However, it is vulnerable to model poisoning attacks, where malicious clients interfere with the training process. Previous defense mechanisms have focused on the server-side by using careful model aggregation, but this may not be effective when the data is not identically distributed or when attackers can access the information of benign clients. In this paper, we propose a new defense mechanism that focuses on the client-side, called FedDefender, to help benign clients train robust local models and avoid the adverse impact of malicious model updates from attackers, even when a server-side defense cannot identify or remove adversaries. Our method consists of two main components: (1) attack-tolerant local meta update and (2) attack-tolerant global knowledge distillation. These components are used to find noise-resilient model parameters while accurately extracting knowledge from a potentially corrupted global model. Our client-side defense strategy has a flexible structure and can work in conjunction with any existing server-side strategies. Evaluations of real-world scenarios across multiple datasets show that the proposed method enhances the robustness of federated learning against model poisoning attacks. Sungwon Park 0001, Sungwon Han 0001, Fangzhao Wu, Sundong Kim, Bin B. Zhu, Xing Xie 0001, Meeyoung Cha |
KDD | 2 |
| 2023 | DualFair: Fair Representation Learning at Both Group and Individual Levels via Contrastive Self-supervisionabstractAlgorithmic fairness has become an important machine learning problem, especially for mission-critical Web applications. This work presents a self-supervised model, called DualFair, that can debias sensitive attributes like gender and race from learned representations. Unlike existing models that target a single type of fairness, our model jointly optimizes for two fairness criteria—group fairness and counterfactual fairness—and hence makes fairer predictions at both the group and individual levels. Our model uses contrastive loss to generate embeddings that are indistinguishable for each protected group, while forcing the embeddings of counterfactual pairs to be similar. It then uses a self-knowledge distillation method to maintain the quality of representation for the downstream tasks. Extensive analysis over multiple datasets confirms the model’s validity and further shows the synergy of jointly addressing two fairness criteria, suggesting the model’s potential value in fair intelligent Web applications. Sungwon Han 0001, SeungEon Lee 0001, Fangzhao Wu, Sundong Kim, Chuhan Wu, Xiting Wang, Xing Xie 0001, Meeyoung Cha |
WWW | 1 |
| 2023 | Active Learning for Human-in-the-Loop Customs InspectionabstractWe study the human-in-the-loop customs inspection scenario, where an AI-assisted algorithm supports customs officers by recommending a set of imported goods to be inspected. If the inspected items are fraudulent, the officers can levy extra duties. These logs are then used as additional training data for the next iterations. Choosing to inspect suspicious items first leads to an immediate gain in customs revenue, yet such inspections may not bring new insights for learning dynamic traffic patterns. On the other hand, inspecting uncertain items can help acquire new knowledge, which will be used as a supplementary training resource to update the selection systems. Based on multiyear customs datasets from three countries, we demonstrate that some degree of exploration is necessary to cope with domain shifts in the trade data. The results show that a hybrid strategy of selecting likely fraudulent and uncertain items will eventually outperform the exploitation-only strategy. Sundong Kim, Tung-Duong Mai, Sungwon Han 0001, Sungwon Park 0001, Thi Nguyen Duc Khanh, Jaechan So, Karandeep Singh, Meeyoung Cha |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Learning Economic Indicators by Aggregating Multi-Level Geospatial InformationabstractHigh-resolution daytime satellite imagery has become a promising source to study economic activities. These images display detailed terrain over large areas and allow zooming into smaller neighborhoods. Existing methods, however, have utilized images only in a single-level geographical unit. This research presents a deep learning model to predict economic indicators via aggregating traits observed from multiple levels of geographical units. The model first measures hyperlocal economy over small communities via ordinal regression. The next step extracts district-level features by summarizing interconnection among hyperlocal economies. In the final step, the model estimates economic indicators of districts via aggregating the hyperlocal and district information. Our new multi-level learning model substantially outperforms strong baselines in predicting key indicators such as population, purchasing power, and energy consumption. The model is also robust against data shortage; the trained features from one country can generalize to other countries when evaluated with data gathered from Malaysia, the Philippines, Thailand, and Vietnam. We discuss the multi-level model's implications for measuring inequality, which is the essential first step in policy and social science research on inequality and poverty. Sungwon Park 0001, Sungwon Han 0001, Donghyun Ahn, Jaeyeon Kim, Jeasurk Yang, Susang Lee, Seunghoon Hong, Hyunjoo Yang, Meeyoung Cha |
AAAI | 2 |
| 2022 | FedX: Unsupervised Federated Learning with Cross Knowledge Distillation
Sungwon Han 0001, Sungwon Park 0001, Fangzhao Wu, Sundong Kim, Chuhan Wu, Xing Xie 0001, Meeyoung Cha |
ECCV (30) | 1 |
| 2022 | Self-explaining deep models with logic rule reasoningabstractWe present SELOR, a framework for integrating self-explaining capabilities into a given deep model to achieve both high prediction performance and human precision. By “human precision”, we refer to the degree to which humans agree with the reasons models provide for their predictions. Human precision affects user trust and allows users to collaborate closely with the model. We demonstrate that logic rule explanations naturally satisfy them with the expressive power required for good predictive performance. We then illustrate how to enable a deep model to predict and explain with logic rules. Our method does not require predefined logic rule sets or human annotations and can be learned efficiently and easily with widely-used deep learning modules in a differentiable way. Extensive experiments show that our method gives explanations closer to human decision logic than other methods while maintaining the performance of the deep learning model. SeungEon Lee 0001, Xiting Wang, Sungwon Han 0001, Xiaoyuan Yi, Xing Xie 0001, Meeyoung Cha |
NeurIPS | 3 |
| 2022 | Using Web Data to Reveal 22-Year History of Sneaker DesignsabstractWeb data and computational models can play important roles in analyzing cultural trends. The current study presents an analysis of 23,492 sneaker images and metadata collected from a global reselling shop, StockX.com. Based on data encompassing 22 years from 1999 to 2020, we propose a sneaker design index that helps track changes in the design characteristics of sneakers using a contrastive learning method. Our data suggest that sneaker designs have been employing brighter colors and lower hue and saturation values over time. We also observe how popular brands have continued to build their unique identities in shape-related design space. The embedding analysis also predicts which sneakers will likely see a high premium in the reselling market, suggesting viable algorithm-driven investment and design strategies. The current work is one of the first publicly available studies to analyze product design evolution over a long historical period and has implications for the novel use of Web data to understand cultural patterns that are otherwise difficult to assess. Sungkyu Park, Hyeonho Song, Sungwon Han 0001, Berhane Weldegebriel, Lev Manovich, Emanuele Arielli, Meeyoung Cha |
WWW | 3 |
| 2021 | Elsa: Energy-based Learning for Semi-supervised Anomaly Detection
Sungwon Han 0001, Hyeonho Song, SeungEon Lee 0001, Sungwon Park 0001, Meeyoung Cha |
BMVC | 1 |
| 2021 | Improving Unsupervised Image Clustering With Robust LearningabstractUnsupervised image clustering methods often introduce alternative objectives to indirectly train the model and are subject to faulty predictions and overconfident results. To overcome these challenges, the current research proposes an innovative model RUC that is inspired by robust learning. RUC’s novelty is at utilizing pseudo-labels of existing image clustering models as a noisy dataset that may include misclassified samples. Its retraining process can revise misaligned knowledge and alleviate the overconfidence problem in predictions. The model’s flexible structure makes it possible to be used as an add-on module to other clustering methods and helps them achieve better performance on multiple datasets. Extensive experiments show that the proposed model can adjust the model confidence with better calibration and gain additional robustness against adversarial noise. Sungwon Park 0001, Sungwon Han 0001, Sundong Kim, Danu Kim, Sungkyu Park, Seunghoon Hong, Meeyoung Cha |
CVPR | 2 |
| 2020 | Lightweight and Robust Representation of Economic Scales from Satellite ImageryabstractSatellite imagery has long been an attractive data source providing a wealth of information regarding human-inhabited areas. While high-resolution satellite images are rapidly becoming available, limited studies have focused on how to extract meaningful information regarding human habitation patterns and economic scales from such data. We present READ, a new approach for obtaining essential spatial representation for any given district from high-resolution satellite imagery based on deep neural networks. Our method combines transfer learning and embedded statistics to efficiently learn the critical spatial characteristics of arbitrary size areas and represent such characteristics in a fixed-length vector with minimal information loss. Even with a small set of labels, READ can distinguish subtle differences between rural and urban areas and infer the degree of urbanization. An extensive evaluation demonstrates that the model outperforms state-of-the-art models in predicting economic scales, such as the population density in South Korea (R2=0.9617), and shows a high use potential in developing countries where district-level economic scales are unknown. Sungwon Han 0001, Donghyun Ahn, Hyunji Cha, Jeasurk Yang, Sungwon Park 0001, Meeyoung Cha |
AAAI | 1 |
| 2020 | A Comprehensive and Adversarial Approach to Self-Supervised Representation LearningabstractSelf-supervised representation learning aims to generate effective representations for data instances without the need for manual labels, also known as unsupervised embedding learning, which has been a critical challenge in many existing semi-supervised and supervised learning tasks. This paper proposes a new self-supervised learning approach, called Super-AND, which extends the memory-based pretraining method AND model [13]. Super-AND has its unique set of losses that combines data augmentation in neighborhood discovery for more accurate anchor selection in embedding learning and further presents an adversarial training manner to learn more confident embeddings under the unsupervised setting. Experimental results exhibit that Super-AND outperforms all existing state-of-the-art self-supervised representation learning approaches and achieves an accuracy of 89.2% on the image classification task for CIFAR-10. Yizhan Xu, Sungwon Han 0001, Sungwon Park 0001, Meeyoung Cha, Cheng-Te Li |
IEEE BigData | 2 |
| 2020 | Mitigating Embedding and Class Assignment Mismatch in Unsupervised Image Classification
Sungwon Han 0001, Sungwon Park 0001, Sungkyu Park, Sundong Kim, Meeyoung Cha |
ECCV (24) | 1 |
| 2020 | Learning to Score Economic Development from Satellite ImageryabstractReliable and timely measurements of economic activities are fundamental for understanding economic development and designing government policies. However, many developing countries still lack reliable data. In this paper, we introduce a novel approach for measuring economic development from high-resolution satellite images in the absence of ground truth statistics. Our method consists of three steps. First, we run a clustering algorithm on satellite images that distinguishes artifacts from nature (siCluster). Second, we generate a partial order graph of the identified clusters based on the level of economic development, either by human guidance or by low-resolution statistics (siPog). Third, we use a CNN-based sorter that assigns differentiable scores to each satellite grid based on the relative ranks of clusters (siScore). The novelty of our method is that we break down a computationally hard problem into sub-tasks, which involves a human-in-the-loop solution. With the combination of unsupervised learning and the partial orders of dozens of urban vs. rural clusters, our method can estimate the economic development scores of over 10,000 satellite grids consistently with other baseline development proxies (Spearman correlation of 0.851). This efficient method is interpretable and robust; we demonstrate how to apply our method to both developed (e.g., South Korea) and developing economies (e.g., Vietnam and Malawi). Sungwon Han 0001, Donghyun Ahn, Sungwon Park 0001, Jeasurk Yang, Susang Lee, Hyunjoo Yang, Meeyoung Cha |
KDD | 1 |
| 2019 | Learning Sleep Quality from Daily LogsabstractPrecision psychiatry is a new research field that uses advanced data mining over a wide range of neural, behavioral, psychological, and physiological data sources for classification of mental health conditions. This study presents a computational framework for predicting sleep efficiency of insomnia sufferers. A smart band experiment is conducted to collect heterogeneous data, including sleep records, daily activities, and demographics, whose missing values are imputed via Improved Generative Adversarial Imputation Networks (Imp-GAIN). Equipped with the imputed data, we predict sleep efficiency of individual users with a proposed interpretable LSTM-Attention (LA Block) neural network model. We also propose a model, Pairwise Learning-based Ranking Generation (PLRG), to rank users with high insomnia potential in the next day. We discuss implications of our findings from the perspective of a psychiatric practitioner. Our computational framework can be used for other applications that analyze and handle noisy and incomplete time-series human activity data in the domain of precision psychiatry. Sungkyu Park, Cheng-Te Li, Sungwon Han 0001, Cheng Hsu, Sang Won Lee 0004, Meeyoung Cha |
KDD | 3 |