EDBT 2026 Demo / reviewers in the wild / expert
Ho Bae
dblp:199/1782
· DBLP profile ↗
14ranked-venue papers
2as first author
12since 2021 · last 2025
0000-0002-5238-3547ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 since 2021Security and privacy · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dependable Code Repair with LLMs: AI-Driven Vulnerability Detection and Automated PatchingabstractThe rapid proliferation of software vulnerabilities has created an urgent need for intelligent, automated methods to detect and mitigate security flaws at scale. Traditional vulnerability analysis depends heavily on manual inspection and domain-specific expertise, which are increasingly inadequate in the era of generative AI-driven code development. This research proposes an AI-based automated vulnerability detection and secure code generation framework that leverages multi-modal datasets, including source code and binaries, to achieve end-to- end automation across the vulnerability lifecycle: detection, patch generation, and validation. The system integrates explainable AI (XAI)-based vulnerability cause analysis, generative patch synthesis, system-level defensive code generation, Rust-based memory safety transformation, and differential privacy mechanisms for model confidentiality. Developed through a Korea- U.S. joint research initiative, this project aims to establish an internationally deployable platform for trustworthy and privacy- preserving AI -driven software security. The proposed research contributes both foundational methods and operational tools toward self-healing, explainable, and secure-by-design software ecosystems. Sungmin Han, Hyoungshick Kim, Hojoon Lee 0001, Hyungon Moon, Yuseok Jeon, Ho Bae, Donghyun Yeo, Gail-Joon Ahn, Sangkyun Lee 0002 |
PRDC | 6 |
| 2025 | Regularizing Hard Examples Improves Adversarial RobustnessabstractRecent studies have validated that pruning hard-to-learn examples from training improves the generalization performance of neural networks (NNs). In this study, we investigate this intriguing phenomenon---the negative effect of hard examples on generalization---in adversarial training. Particularly, we theoretically demonstrate that the increase in the difficulty of hard examples in adversarial training is significantly greater than the increase in the difficulty of easy examples. Furthermore, we verify that hard examples are only fitted through memorization of the label in adversarial training. We conduct both theoretical and empirical analyses of this memorization phenomenon, showing that pruning hard examples in adversarial training can enhance the model's robustness. However, the challenge remains in finding the optimal threshold for removing hard examples that degrade robustness performance. Based upon these observations, we propose a new approach, difficulty proportional label smoothing (DPLS), to adaptively mitigate the negative effect of hard examples, thereby improving the adversarial robustness of NNs. Notably, our experimental result indicates that our method can successfully leverage hard examples while circumventing the negative effect. Hyungyu Lee, Saehyung Lee, Ho Bae, Sungroh Yoon |
J. Mach. Learn. Res. | 3 |
| 2024 | VFLIP: A Backdoor Defense for Vertical Federated Learning via Identification and Purification
Yungi Cho, Woorim Han, Miseon Yu, Younghan Lee 0001, Ho Bae, Yunheung Paek |
ESORICS (4) | 5 |
| 2024 | DAFA: Distance-Aware Fair Adversarial TrainingabstractThe disparity in accuracy between classes in standard training is amplified during adversarial training, a phenomenon termed the robust fairness problem. Existing methodologies aimed to enhance robust fairness by sacrificing the model's performance on easier classes in order to improve its performance on harder ones. However, we observe that under adversarial attacks, the majority of the model's predictions for samples from the worst class are biased towards classes similar to the worst class, rather than towards the easy classes. Through theoretical and empirical analysis, we demonstrate that robust fairness deteriorates as the distance between classes decreases. Motivated by these insights, we introduce the Distance-Aware Fair Adversarial Training (DAFA) methodology, which addresses robust fairness by taking into account the similarities between classes. Specifically, our method assigns distinct adversarial margins and loss weights to each class and adjusts them to encourage a trade-off in robustness among similar classes. Experimental results across various datasets demonstrate that our method not only maintains average robust accuracy but also significantly improves the worst robust accuracy, indicating a marked improvement in robust fairness compared to existing methods. Hyungyu Lee, Saehyung Lee, Hyemi Jang, Junsung Park 0001, Ho Bae, Sungroh Yoon |
ICLR | 5 |
| 2023 | Privacy-Preserving Publishing of Individual-Level Medical Data for Cloud ServicesabstractDeep learning (DL) has been extensively adopted in many applications, including disease prediction. Most DL-based applications are executed on a cloud server because the DL models are too large and complicated to be executed on the client-side. De facto cloud-hosted inferences lead to privacy concerns regarding services that operate on personal medical data. Nevertheless, given the recent development of DL-based applications for health-diagnosis services, these applications have become a dominant means of healthcare support in our daily lives. To prevent the misuse of personal medical data, several techniques have been developed to preserve sensitive information, with a trade-off between privacy and utility. A simple method that offers privacy preservation and good prediction performance involves the deployment of a diagnostic method to the client side. However, doing so makes DL models more vulnerable to adversaries. To this end, we propose a deep private generative framework that guarantees user-data privacy while maintaining the original class information and protecting the models from reverse engineering. Experimentation with practical deep neural networks on benchmark disease datasets demonstrates that the proposed method decreases the mutual information between the original data and synthetic data by nearly 80% while preserving a prediction accuracy of nearly 95% of the original prediction accuracy. Ho Bae, Heonseok Ha, Siwon Kim |
BIBM | 1 |
| 2023 | FLGuard: Byzantine-Robust Federated Learning via Ensemble of Contrastive Models
Younghan Lee 0001, Yungi Cho, Woorim Han, Ho Bae, Yunheung Paek |
ESORICS (4) | 4 |
| 2023 | New Insights for the Stability-Plasticity Dilemma in Online Continual Learning
Dahuin Jung, Sunwon Hong, Hyemi Jang, Ho Bae, Sungroh Yoon |
ICLR | 5 |
| 2023 | PUCA: Patch-Unshuffle and Channel Attention for Enhanced Self-Supervised Image DenoisingabstractAlthough supervised image denoising networks have shown remarkable performance on synthesized noisy images, they often fail in practice due to the difference between real and synthesized noise. Since clean-noisy image pairs from the real world are extremely costly to gather, self-supervised learning, which utilizes noisy input itself as a target, has been studied. To prevent a self-supervised denoising model from learning identical mapping, each output pixel should not be influenced by its corresponding input pixel; This requirement is known as J-invariance. Blind-spot networks (BSNs) have been a prevalent choice to ensure J-invariance in self-supervised image denoising. However, constructing variations of BSNs by injecting additional operations such as downsampling can expose blinded information, thereby violating J-invariance. Consequently, convolutions designed specifically for BSNs have been allowed only, limiting architectural flexibility. To overcome this limitation, we propose PUCA, a novel J-invariant U-Net architecture, for self-supervised denoising. PUCA leverages patch-unshuffle/shuffle to dramatically expand receptive fields while maintaining J-invariance and dilated attention blocks (DABs) for global context incorporation. Experimental results demonstrate that PUCA achieves state-of-the-art performance, outperforming existing methods in self-supervised image denoising. Hyemi Jang, Junsung Park 0001, Dahuin Jung, Jaihyun Lew, Ho Bae, Sungroh Yoon |
NeurIPS | 5 |
| 2023 | Exploring Clustered Federated Learning's Vulnerability against Property Inference AttackabstractClustered federated learning (CFL) is an advanced technique in the field of federated learning (FL) that addresses the issue of catastrophic forgetting caused by non-independent and identically distributed (non-IID) datasets. CFL achieves this by clustering clients based on the similarity of their datasets and training a global model for each cluster. Despite the effectiveness of CFL in mitigating performance degradation resulting from non-IID datasets, the potential risk of privacy leakages in CFL has not been thoroughly studied. Previous work evaluated the risk of privacy leakages in FL using the property inference attack (PIA), which extracts information about unintended properties (i.e., attributes that differ from the target attribute of the global model’s main task). In this paper, we explore the potential risk of unintended property leakage in CFL by subjecting it to both passive and active PIAs. Our empirical analysis shows that the passive PIA performance on CFL is substantially better than that on FL in terms of the attack AUC score. Moreover, we propose an enhanced active PIA method tailored for CFL to improve the attack performance. Our method introduces a scale-up parameter that amplifies the impact of malicious local updates, resulting in better performance than the previous technique. Furthermore, we demonstrate that the vulnerability of CFL can be alleviated by applying differential privacy (DP) mechanisms at the client-level. Unlike previous works, which have shown that applying DP to FL can induce a high utility loss, our empirical results indicate that DP can be used as a defense mechanism in CFL, leading to a better trade-off between privacy and utility. Yungi Cho, Younghan Lee 0001, Ho Bae, Yunheung Paek |
RAID | 4 |
| 2023 | PixelSteganalysis: Pixel-Wise Hidden Information Removal With Low Visual DegradationabstractRecently, the field of steganography has experienced rapid developments based on deep learning (DL). DL based steganography distributes secret information over all the available bits of the cover image, thereby posing difficulties in using conventional steganalysis methods to detect, extract or remove hidden secret images. However, our proposed framework is the first to effectively disable covert communications and transactions that use DL based steganography. We propose a DL based steganalysis technique that effectively removes secret images by restoring the distribution of the original images. We formulate a problem and address it by exploiting sophisticated pixel distributions and an edge distribution of images by using a deep neural network. Based on the given information, we remove the hidden secret information at the pixel level. We evaluate our technique by comparing it with conventional steganalysis methods using three public benchmarks. As the decoding method of DL based steganography is approximate (lossy) and is different from the decoding method of conventional steganography, we also introduce a new quantitative metric called the destruction rate (DT). The experimental results demonstrate performance improvements of 10–20$\%$in both the decoded rate and the DT. Dahuin Jung, Ho Bae, Hyun-Soo Choi, Sungroh Yoon |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2022 | Membership Privacy-Preserving GAN
Heonseok Ha, Uiwon Hwang, Jaehee Jang, Ho Bae, Sungroh Yoon |
BMVC | 4 |
| 2022 | DNA Privacy: Analyzing Malicious DNA Sequences Using Deep Neural NetworksabstractRecent advances in next-generation sequencing technologies have led to the successful insertion of video information into DNA using synthesized oligonucleotides. Several attempts have been made to embed larger data into living organisms. This process of embedding messages is called steganography and it is used for hiding and watermarking data to protect intellectual property. In contrast, steganalysis is a group of algorithms that serves to detect hidden information from covert media. Various methods have been developed to detect messages embedded in conventional covert channels. However, conventional steganalysis algorithms are mostly limited to common covert media. Most common detection approaches, such as frequency analysis-based methods, often overlook important signals when directly applied to DNA steganography and are easily bypassed by recently developed steganography techniques. To address the limitations of conventional approaches, a sequence-learning-based malicious DNA sequence analysis method based on neural networks has been proposed. The proposed method learns intrinsic distributions and identifies distribution variations using a classification score to predict whether a sequence is to be a coding or non-coding sequence. Based on our experiments and results, we have developed a framework to safeguard security against DNA steganography. Ho Bae, Seonwoo Min, Hyun-Soo Choi, Sungroh Yoon |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2019 | Comprehensive ensemble in QSAR prediction for drug discoveryabstractBACKGROUND: Quantitative structure-activity relationship (QSAR) is a computational modeling method for revealing relationships between structural properties of chemical compounds and biological activities. QSAR modeling is essential for drug discovery, but it has many constraints. Ensemble-based machine learning approaches have been used to overcome constraints and obtain reliable predictions. Ensemble learning builds a set of diversified models and combines them. However, the most prevalent approach random forest and other ensemble approaches in QSAR prediction limit their model diversity to a single subject. RESULTS: The proposed ensemble method consistently outperformed thirteen individual models on 19 bioassay datasets and demonstrated superiority over other ensemble approaches that are limited to a single subject. The comprehensive ensemble method is publicly available at http://data.snu.ac.kr/QSAR/ . CONCLUSIONS: We propose a comprehensive ensemble method that builds multi-subject diversified models and combines them through second-level meta-learning. In addition, we propose an end-to-end neural network-based individual classifier that can automatically extract sequential features from a simplified molecular-input line-entry system (SMILES). The proposed individual models did not show impressive results as a single model, but it was considered the most important predictor when combined, according to the interpretation of the meta-learning. Sunyoung Kwon, Ho Bae, Jeonghee Jo, Sungroh Yoon |
BMC Bioinform. | 2 |
| 2018 | Quantized Memory-Augmented Neural NetworksabstractMemory-augmented neural networks (MANNs) refer to a class of neural network models equipped with external memory (such as neural Turing machines and memory networks). These neural networks outperform conventional recurrent neural networks (RNNs) in terms of learning long-term dependency, allowing them to solve intriguing AI tasks that would otherwise be hard to address. This paper concerns the problem of quantizing MANNs. Quantization is known to be effective when we deploy deep models on embedded systems with limited resources. Furthermore, quantization can substantially reduce the energy consumption of the inference procedure. These benefits justify recent developments of quantized multi layer perceptrons, convolutional networks, and RNNs. However, no prior work has reported the successful quantization of MANNs. The in-depth analysis presented here reveals various challenges that do not appear in the quantization of the other networks. Without addressing them properly, quantized MANNs would normally suffer from excessive quantization error which leads to degraded performance. In this paper, we identify memory addressing (specifically, content-based addressing) as the main reason for the performance degradation and propose a robust quantization method for MANNs to address the challenge. In our experiments, we achieved a computation-energy gain of 22× with 8-bit fixed-point and binary quantization compared to the floating-point implementation. Measured on the bAbI dataset, the resulting model, named the quantized MANN (Q-MANN), improved the error rate by 46% and 30% with 8-bit fixed-point and binary quantization, respectively, compared to the MANN quantized using conventional techniques. Sei Joon Kim, Seil Lee, Ho Bae, Sungroh Yoon |
AAAI | 4 |