Saerom Park

dblp:209/8156 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-2687-7105ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 MIG-COW: Transferable Adversarial Attacks on Deepfake Detectors via Gradient Decomposition
Wonjune Seo, Joonhyuk Baek, Yeseong Jung, Saerom Park
ACM Multimedia4
2024 Fair Sampling in Diffusion Models through Switching Mechanism
abstract
Diffusion models have shown their effectiveness in generation tasks by well-approximating the underlying probability distribution. However, diffusion models are known to suffer from an amplified inherent bias from the training data in terms of fairness. While the sampling process of diffusion models can be controlled by conditional guidance, previous works have attempted to find empirical guidance to achieve quantitative fairness. To address this limitation, we propose a fairness-aware sampling method called \textit{attribute switching} mechanism for diffusion models. Without additional training, the proposed sampling can obfuscate sensitive attributes in generated data without relying on classifiers. We mathematically prove and experimentally demonstrate the effectiveness of the proposed method on two key aspects: (i) the generation of fair data and (ii) the preservation of the utility of the generated data.
Jinseong Park 0001, Hoki Kim, Jaewook Lee 0001, Saerom Park
AAAI5
2024 Privacy-Preserving Embedding via Look-up Table Evaluation with Fully Homomorphic Encryption
abstract
In privacy-preserving machine learning (PPML), homomorphic encryption (HE) has emerged as a significant primitive, allowing the use of machine learning (ML) models while protecting the confidentiality of input data. Although extensive research has been conducted on implementing PPML with HE by developing the efficient construction of private counterparts to ML models, the efficient HE implementation of embedding layers for token inputs such as words remains inadequately addressed. Thus, our study proposes an efficient algorithm for privacy-preserving embedding via look-up table evaluation with HE(HELUT) by developing an encrypted indicator function (EIF) that assures high precision with the use of the approximate HE scheme(CKKS). Based on the proposed EIF, we propose the CodedHELUT algorithm to facilitate an encrypted embedding layer for the first time. CodedHELUT leverages coded inputs to improve overall efficiency and optimize memory usage. Our comprehensive empirical analysis encompasses both synthetic tables and real-world largescale word embedding models. CodedHELUT algorithm achieves amortized evaluation time of 0.018-0.242s for GloVe6B50d, 0.104-01.298s for GloVe42300d, 0.262-3.283s for GPT-2 and BERT embedding layers while maintaining high precision (16 bits)
Jaeyun Kim, Saerom Park, Joohee Lee, Jung Hee Cheon
ICML2
2024 Privacy-preserving inference resistant to model extraction attacks
Junyoung Byun, Jaewook Lee 0001, Saerom Park
Expert Syst. Appl.4
2023 Efficient homomorphic encryption framework for privacy-preserving regression
Junyoung Byun, Saerom Park, Jaewook Lee 0001
Appl. Intell.2
2023 Efficient differentially private kernel support vector classifier for multi-class classification
Jinseong Park 0001, Junyoung Byun, Jaewook Lee 0001, Saerom Park
Inf. Sci.5
2022 Privacy-Preserving Fair Learning of Support Vector Machine with Homomorphic Encryption
abstract
Fair learning has received a lot of attention in recent years since machine learning models can be unfair in automated decision-making systems with respect to sensitive attributes such as gender, race, etc. However, to mitigate the discrimination on the sensitive attributes and train a fair model, most fair learning methods have required to get access to the sensitive attributes in training or validation phases. In this study, we propose a privacy-preserving training algorithm for a fair support vector machine classifier based on Homomorphic Encryption (HE), where the privacy of both sensitive information and model secrecy can be preserved. The expensive computational costs of HE can be significantly improved by protecting only the sensitive information, introducing refined formulation and low-rank approximation using shared eigenvectors. Through experiments on the synthetic and real-world data, we demonstrate the effectiveness of our algorithm in terms of accuracy and fairness and show that our method significantly outperforms other privacy-preserving solutions in terms of better trade-offs between accuracy and fairness. To the best of our knowledge, our algorithm is the first privacy-preserving fair learning algorithm using HE.
Saerom Park, Junyoung Byun, Joohee Lee
WWW1
2022 Fairness Audit of Machine Learning Models with Confidential Computing
abstract
Algorithmic discrimination is one of the significant concerns in applying machine learning models to a real-world system. Many researchers have focused on developing fair machine learning algorithms without discrimination based on legally protected attributes. However, the existing research has barely explored various security issues that can occur while evaluating model fairness and verifying fair models. In this study, we propose a fairness audit framework that assesses the fairness of ML algorithms while addressing potential security issues such as data privacy, model secrecy, and trustworthiness. To this end, our proposed framework utilizes confidential computing and builds a chain of trust through enclave attestation primitives combined with public scrutiny and state-of-the-art software-based security techniques, enabling fair ML models to be securely certified and clients to verify a certified one. Our micro-benchmarks on various ML models and real-world datasets show the feasibility of the fairness certification implemented with Intel SGX in practice. In addition, we analyze the impact of data poisoning, which is an additional threat during data collection for fairness auditing. Based on the analysis, we illustrate the theoretical curves of fairness gap and minimal group size and the empirical results of fairness certification on poisoned datasets.
Saerom Park, Yeon-sup Lim
WWW1
2021 Stability Analysis of Denoising Autoencoders Based on Dynamical Projection System
abstract
In this study, we give a stability analysis of denoising autoencoder(DAE) from the novel perspective of dynamical systems when the input density is defined as a distribution on a manifold. We demonstrate the connection between the corrupted distribution and the learned reconstruction function of a nonlinear DAE, which motivates the use of a dynamic projection system (DPS) associated with the learned reconstruction function. Utilizing the constructed DPS, we prove that the high-density region of the corrupted data distribution asymptotically converges to the data manifold. Then, we show that the region is the attracting stable equilibrium manifold of the DPS which is completely stable. These results serve a theoretical basis of the DAE in recognizing the high-density region of the highly corrupted data with large deviations through the DPS. The effectiveness of this analysis is verified by conducting experiments on several toy examples and real image datasets with various types of noise.
Saerom Park, Jaewook Lee 0001
IEEE Trans. Knowl. Data Eng.1
2020 Lipschitz-Certifiable Training with a Tight Outer Bound
abstract
Verifiable training is a promising research direction for training a robust network. However, most verifiable training methods are slow or lack scalability. In this study, we propose a fast and scalable certifiable training algorithm based on Lipschitz analysis and interval arithmetic. Our certifiable training algorithm provides a tight propagated outer bound by introducing the box constraint propagation (BCP), and it efficiently computes the worst logit over the outer bound. In the experiments, we show that BCP achieves a tighter outer bound than the global Lipschitz-based outer bound. Moreover, our certifiable training algorithm is over 12 times faster than the state-of-the-art dual relaxation-based method; however, it achieves comparable or better verification performance, improving natural accuracy. Our fast certifiable training algorithm with the tight outer bound can scale to Tiny ImageNet with verification accuracy of 20.1\% ($\ell_2$-perturbation of $\epsilon=36/255$). Our code is available at \url{https://github.com/sungyoon-lee/bcp}.
Sungyoon Lee, Jaewook Lee 0001, Saerom Park
NeurIPS3
2019 Learning of indiscriminate distributions of document embeddings for domain adaptation
abstract
Natural language processing (NLP) is an important application area in domain adaptation because properties of texts depend on their corpus. However, a textual input is not fundamentally represented as the numerical vector. Many domain adaptation methods for NLP have been developed on the basis of n umerical representations of texts instead of textual inputs. Thus, we develop a distributed representation learning method of words and documents for domain adaptation. The developed method addresses the domain separation problem of document embeddings from different domains, that is, the supports of the embeddings are separable across domains and the distributions of the embeddings are discriminated. We propose a new method based on negative sampling. The proposed method learns document embeddings by assuming that a noise distribution is dependent on a domain. The proposed method moves a document embedding close to the embeddings of the important words in the document and keeps the embedding away from the word embeddings that occur frequently in both domains. For Amazon reviews, we verified that the proposed method outperformed other representation methods in terms of indiscriminability of the distributions of the document embeddings through experiments such as visualizing them and calculating a proxy A-distance measure. We also performed sentiment classification tasks to validate the effectiveness of document embeddings. The proposed method achieved consistently better results than other methods. In addition, we applied the learned document embeddings to the domain adversarial neural network method, which is a popular deep learning-based domain adaptation model. The proposed method obtained not only better performance on most datasets but also more stable convergences for all datasets than the other methods. Therefore, the proposed method are applicable to other domain adaptation methods for NLP using numerical representations of documents or words.
Saerom Park, Jaewook Lee 0001
Intell. Data Anal.1
2019 Semi-supervised distributed representations of documents for sentiment analysis
Saerom Park, Jaewook Lee 0001, Kyoungok Kim
Neural Networks1
2018 Learning representative exemplars using one-class Gaussian process regression
Youngdoo Son, Sujee Lee, Saerom Park, Jaewook Lee 0001
Pattern Recognit.3