EDBT 2026 Demo / reviewers in the wild / expert
Junghye Lee
dblp:234/8224
· DBLP profile ↗
8ranked-venue papers in the field
0as first author
8since 2021 · last 2027
—ORCID · conflict
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 4Data Mining & Knowledge Discovery · 2Information Retrieval & Web Search · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Who invests matters: A multimodal contrastive learning approach to successful startup-investor matching
Yeokyung Hwang, Eunbi Jeon, Ana Theodora Balaci, Junseok Hwang, Junghye Lee |
Inf. Process. Manag. | 5 |
| 2025 | Federated Gradient Boosting for Financial Fraud Detection: An Empirical Study in the Banking SectorabstractThe development of effective fraud detection systems (FDS) is hindered by strict data privacy regulations that prevent centralized data sharing. Federated learning (FL) has emerged as a promising alternative, enabling collaborative model training without exposing sensitive data. While FL has been explored in the healthcare domain, research on its application to financial fraud detection remains relatively limited. Specifically, FL research on real-world banking fraud types-with detailed customer, account, and transaction data-remains underexplored. We present the first empirical study of federated gradient boosting models for financial fraud detection in the banking sector, motivated by their superior performance over deep learning models on tabular fraud data. We evaluate and compare four representative federated gradient boosting models using both a private multi-fraud banking dataset from the Financial Security Institute (FSI) and a publicly available banking dataset, under various scenarios. Key findings include the consistent superiority of FedXGBBagging (a federated gradient boosting model), general vulnerability to data quantity skew, performance instability under bank join/dropout, and limitations in detecting localized banking fraud types such as ATM skimming. The findings from our empirical study highlight challenges and design considerations for deploying FL-based FDSs in the banking sector. In-Young Ko, Taek-Ho Lee, Junghye Lee |
CIKM | 4 |
| 2024 | CAFO: Feature-Centric Explanation on Time Series ClassificationabstractIn multivariate time series (MTS) classification, finding the important features (e.g., sensors) for model performance is crucial yet challenging due to the complex, high-dimensional nature of MTS data, intricate temporal dynamics, and the necessity for domain-specific interpretations. Current explanation methods for MTS mostly focus on time-centric explanations, apt for pinpointing important time periods but less effective in identifying key features. This limitation underscores the pressing need for a feature-centric approach, a vital yet often overlooked perspective that complements time-centric analysis. To bridge this gap, our study introduces a novel feature-centric explanation and evaluation framework for MTS, named CAFO (Channel Attention and Feature Orthgonalization). CAFO employs a convolution-based approach with channel attention mechanisms, incorporating a depth-wise separable channel attention module (DepCA) and a QR decomposition-based loss for promoting feature-wise orthogonality. We demonstrate that this orthogonalization enhances the separability of attention distributions, thereby refining and stabilizing the ranking of feature importance. This improvement in feature-wise ranking enhances our understanding of feature explainability in MTS. Furthermore, we develop metrics to evaluate global and class-specific feature importance. Our framework's efficacy is validated through extensive empirical analyses on two major public benchmarks and real-world datasets, both synthetic and self-collected, specifically designed to highlight class-wise discriminative features. The results confirm CAFO's robustness and informative capacity in assessing feature importance in MTS classification tasks. This study not only advances the understanding of feature-centric explanations in MTS but also sets a foundation for future explorations in feature-centric explanations. The codes are available at https://github.com/eai-lab/CAFO. Seok-Ju Hahn, Yoontae Hwang, Junghye Lee, Seulki Lee 0002 |
KDD | 4 |
| 2024 | HarmoSATE: Harmonized embedding-based self-attentive encoder to improve accuracy of privacy-preserving federated predictive analysisabstractAccurate privacy-preserving prediction using electronic health record (EHR) data distributed in multiple hospitals is essential to enable stakeholders related to healthcare services to obtain useful information without privacy leakage. In this paper, we propose harmonized embedding-based self-attentive encoder (HarmoSATE), which is a new method for privacy-preserving federated predictive analysis. We extract contextual embeddings of local institutions using Word2Vec, and then harmonize locally-trained embeddings using a neural network-based harmonization technique. The proposed method uses a deep representative encoder based on self-attention to learn complex and dynamic patterns inherent to harmonized embeddings of medical concepts. To evaluate our method, we implemented experiments using sequential medical codes collected from the Medical Information Mart for Intensive Care-III dataset in a distributed setting. It achieved a significant increase in average AUC, ranging from 3% to 8% depending on the experiments compared to baseline models, demonstrating superior prediction accuracy of a patient's diagnosis in the next admission. HarmoSATE can be a useful alternative to obtain accurate and practical results for various predictive tasks that use sensitive and distributed EHR data while preserving patients' privacy. Taek-Ho Lee, Suhyeon Kim, Junghye Lee, Chi-Hyuck Jun |
Inf. Sci. | 3 |
| 2023 | Word2Vec-based efficient privacy-preserving shared representation learning for federated recommendation system in a cross-device settingabstractRecommendation systems have required centralized storage of user data, but due to privacy concerns, recent studies adopted federated learning (FL) that discloses intermediate statistics instead of raw data to build privacy-preserving federated recommendation systems. However, they suffer from inefficiencies in privacy-preserving mechanisms and inaccuracies in simple algorithms that ignore sequential information. This study proposes an extension of Word2Vec for a privacy-preserving federated sequential recommendation system (PPFSRS). This method exploits sequential information to generate contextual item representations for accurate recommendations while concealing privacy-sensitive features efficiently. Specifically, we mixed updates from negative samples to inhibit the direct leakage of purchased items from model updates. In addition, our method computes approximate model updates that can occur when sensitive features only belong to negative samples to prevent inference attacks. In experiments, we used benchmark datasets for recommendation and simulated highly distributed data such that each user stores historical data locally. While preserving privacy with reasonable complexity, the proposed method showed little degradation in recommendation performance compared to FL-based Word2Vec without privacy-preserving mechanisms. Utilizing contextual item representations trained by our method from highly distributed data will be a practical starting point for PPFSRS in a cross-device setting. Taek-Ho Lee, Suhyeon Kim, Junghye Lee, Chi-Hyuck Jun |
Inf. Sci. | 3 |
| 2022 | Connecting Low-Loss Subspace for Personalized Federated LearningabstractDue to the curse of statistical heterogeneity across clients, adopting a personalized federated learning method has become an essential choice for the successful deployment of federated learning-based services. Among diverse branches of personalization techniques, a model mixture-based personalization method is preferred as each client has their own personalized model as a result of federated learning. It usually requires a local model and a federated model, but this approach is either limited to partial parameter exchange or requires additional local updates, each of which is helpless to novel clients and burdensome to the client's computational capacity. As the existence of a connected subspace containing diverse low-loss solutions between two or more independent deep networks has been discovered, we combined this interesting property with the model mixture-based personalized federated learning method for improved performance of personalization. We proposed SuPerFed, a personalized federated learning method that induces an explicit connection between the optima of the local and the federated model in weight space for boosting each other. Through extensive experiments on several benchmark datasets, we demonstrated that our method achieves consistent gains in both personalization performance and robustness to problematic scenarios possible in realistic services. Seok-Ju Hahn, Minwoo Jeong, Junghye Lee |
KDD | 3 |
| 2022 | Risk score-embedded deep learning for biological age estimation: Development and validationabstractThe health index measures a person’s overall health status which provides useful information for people to manage their health, so developing a precise and relevant health index is urgent. Currently, many researchers have studied the biological age (BA) estimation, one of the beneficial health indices, by applying machine learning and deep learning techniques to health data. However, most of them have focused on the chronological age prediction or basic latent feature extraction methods. In this paper, we present a new algorithm to estimate BA, called Risk Score-Embedded Autoencoder-based BA (RSAE-BA). RSAE-BA can provide an accurate health index by using deep representation learning with an individual’s health risk. We first proposed a notion of risk score (RS) calculation to monitor a person’s health risk. Then we extracted latent features by using an autoencoder embedding the RS, and used them to generate BA. To evaluate RSAE-BA, we presented a new BA validation method using the RS, which is applicable to both unlabeled and labeled data. We compared the results of RSAE-BA with existing methods, and demonstrated the accuracy of RSAE-BA and its applicability to predict disease incidence. We believe that RSAE-BA will be a useful alternative method to measure a person’s health. Suhyeon Kim, Eun-Sol Lee, Chiehyeon Lim, Junghye Lee |
Inf. Sci. | 5 |
| 2021 | Bilingual autoencoder-based efficient harmonization of multi-source private data for accurate predictive modelingabstractSharing electronic health record data is essential for advanced analysis, but may put sensitive information at risk. Several studies have attempted to address this risk using contextual embedding, but with many hospitals involved, they are often inefficient and inflexible. Thus, we propose a bilingual autoencoder-based model to harmonize local embeddings in different spaces. Cross-hospital reconstruction of embeddings makes encoders map embeddings from hospitals to a shared space and align them spontaneously. We also suggest two-phase training to prevent distortion of embeddings during harmonization with hospitals that have biased information. In experiments, we used medical event sequences from the Medical Information Mart for Intensive Care-III dataset and simulated the situation of multiple hospitals. For evaluation, we measured the alignment of events from different hospitals and the prediction accuracy of a patient’s diagnosis in the next admission in three scenarios in which local embeddings do not work. The proposed method efficiently harmonizes embeddings in different spaces, increases prediction accuracy, and gives flexibility to include new hospitals, so is superior to previous methods in most cases. It will be useful in predictive tasks to utilize distributed data while preserving private information. Taek-Ho Lee, Junghye Lee, Chi-Hyuck Jun |
Inf. Sci. | 2 |