Junghye Lee

dblp:234/8224 · DBLP profile ↗
← Back
21ranked-venue papers
2as first author
18since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 8 · 8 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2027 Who invests matters: A multimodal contrastive learning approach to successful startup-investor matching
Yeokyung Hwang, Eunbi Jeon, Ana Theodora Balaci, Junseok Hwang, Junghye Lee
Inf. Process. Manag.5
2026 Multi-modal deep learning-based fashion recommendation with styles
Wonho Sohn, Suhyeon Kim, Dongcheol Lim, Hyerim Park, Wonji Lee, Hwansung Yu, Junghye Lee
Knowl. Based Syst.8
2025 Federated Gradient Boosting for Financial Fraud Detection: An Empirical Study in the Banking Sector
abstract
The development of effective fraud detection systems (FDS) is hindered by strict data privacy regulations that prevent centralized data sharing. Federated learning (FL) has emerged as a promising alternative, enabling collaborative model training without exposing sensitive data. While FL has been explored in the healthcare domain, research on its application to financial fraud detection remains relatively limited. Specifically, FL research on real-world banking fraud types-with detailed customer, account, and transaction data-remains underexplored. We present the first empirical study of federated gradient boosting models for financial fraud detection in the banking sector, motivated by their superior performance over deep learning models on tabular fraud data. We evaluate and compare four representative federated gradient boosting models using both a private multi-fraud banking dataset from the Financial Security Institute (FSI) and a publicly available banking dataset, under various scenarios. Key findings include the consistent superiority of FedXGBBagging (a federated gradient boosting model), general vulnerability to data quantity skew, performance instability under bank join/dropout, and limitations in detecting localized banking fraud types such as ATM skimming. The findings from our empirical study highlight challenges and design considerations for deploying FL-based FDSs in the banking sector.
In-Young Ko, Taek-Ho Lee, Junghye Lee
CIKM4
2025 Diffusion Federated Dataset
abstract
Diffusion models have demonstrated decent generation quality, yet their deployment in federated learning scenarios remains challenging. Due to data heterogeneity and a large number of parameters, conventional parameter averaging schemes often fail to achieve stable collaborative training of diffusion models. We reframe collaborative synthetic data generation as a cooperative sampling procedure from a mixture of decentralized distributions, each encoded by a pre-trained local diffusion model. This leverages the connection between diffusion and energy-based models, which readily supports compositional generation thereof. Consequently, we can directly obtain refined synthetic dataset, optionally with differential privacy guarantee, even without exchanging diffusion model parameters. Our framework reduces communication overhead while maintaining the generation quality, realized through an unadjusted Langevin algorithm with a convergence guarantee.
Seok-Ju Hahn, Junghye Lee
NeurIPS2
2025 A graph convolutional network for time series classification using recurrence plots
abstract
Abstract Time series classification (TSC) is a crucial task across various domains, and its performance heavily depends on the quality of input representations. Among various representations, the recurrence plot (RP) effectively captures topological recurrence, the unique property of time series data. However, conventional convolutional neural networks (CNNs) cannot fully exploit this property since they treat the RP as grid-like data. In this study, we propose RP-GCN, a novel approach that uses a graph convolutional network (GCN) to exploit topological recurrence inherent in the RP, thereby improving TSC performance. Our method transforms a multivariate time series into graphs where state matrices act as node feature matrices and RPs serve as adjacency matrices, enabling graph convolution to utilize recurrence relationships. We evaluated RP-GCN on 35 benchmark multivariate time series classification datasets and demonstrated superior accuracy and efficient inference time compared to existing methods.
Hyewon Kang, Taek-Ho Lee, Junghye Lee
Appl. Intell.3
2025 Inter-country trade similarity graph-based long short-term memory for port throughput prediction
Wonho Sohn, Dongcheol Lim, Suhyeon Kim, Junghye Lee
Eng. Appl. Artif. Intell.4
2025 Machine learning for disease-specific prediction of high-cost patients
Inwoo Tae, Hyeongwoo Kong, Junghye Lee
Eng. Appl. Artif. Intell.3
2025 TMF-GNN: Temporal matrix factorization-based graph neural network for multivariate time series forecasting with missing values
abstract
Missing data in multivariate time series (MTS) is a very common issue, often caused by unreliable sensors and data storage or transmission problems. Particularly, such missing data can cause some errors and biases in the MTS forecasting tasks of real-world applications, implying that proper handling of the missing data is essential. Therefore, in this study, we propose a new method for MTS forecasting with missing values, called a temporal matrix factorization-based graph neural network (TMF-GNN), to improve predictive performance outcomes. TMF-GNN basically uses the concept of TMF, which reconstructs partially observed MTS data. We newly present a data-adaptive regularization method for TMF based on graph-based and sequential deep learning algorithms to capture both the variable-wise and time-wise information of MTS data affected by missingness. We demonstrate the feasibility of the proposed method by conducting various experiments on three MTS datasets and show how it outperforms baseline methods. We believe that our study will have an impact on several MTS-related tasks and that it can be a useful alternative for handling missing values in MTS data. • Missing data is a pervasive problem in multivariate time series forecasting. • We proposed a method to optimize time series forecasting with missing data. • We proposed an approach using graph-based and temporal deep learning models. • The proposed method outperformed existing methods in forecasting accuracy. • Our approach can mine temporal patterns and inter-correlations in time series data.
Suhyeon Kim, Taek-Ho Lee, Junghye Lee
Expert Syst. Appl.3
2024 Pursuing Overall Welfare in Federated Learning through Sequential Decision Making
abstract
In traditional federated learning, a single global model cannot perform equally well for all clients. Therefore, the need to achieve the *client-level fairness* in federated system has been emphasized, which can be realized by modifying the static aggregation scheme for updating the global model to an adaptive one, in response to the local signals of the participating clients. Our work reveals that existing fairness-aware aggregation strategies can be unified into an online convex optimization framework, in other words, a central server's *sequential decision making* process. To enhance the decision making capability, we propose simple and intuitive improvements for suboptimal designs within existing methods, presenting $\texttt{AAggFF}$. Considering practical requirements, we further subdivide our method tailored for the *cross-device* and the *cross-silo* settings, respectively. Theoretical analyses guarantee sublinear regret upper bounds for both settings: $\mathcal{O}(\sqrt{T \log{K}})$ for the cross-device setting, and $\mathcal{O}(K \log{T})$ for the cross-silo setting, with $K$ clients and $T$ federation rounds. Extensive experiments demonstrate that the federated system equipped with $\texttt{AAggFF}$ achieves better degree of client-level fairness than existing methods in both practical settings.
Seok-Ju Hahn, Gi-Soo Kim, Junghye Lee
ICML3
2024 CAFO: Feature-Centric Explanation on Time Series Classification
abstract
In multivariate time series (MTS) classification, finding the important features (e.g., sensors) for model performance is crucial yet challenging due to the complex, high-dimensional nature of MTS data, intricate temporal dynamics, and the necessity for domain-specific interpretations. Current explanation methods for MTS mostly focus on time-centric explanations, apt for pinpointing important time periods but less effective in identifying key features. This limitation underscores the pressing need for a feature-centric approach, a vital yet often overlooked perspective that complements time-centric analysis. To bridge this gap, our study introduces a novel feature-centric explanation and evaluation framework for MTS, named CAFO (Channel Attention and Feature Orthgonalization). CAFO employs a convolution-based approach with channel attention mechanisms, incorporating a depth-wise separable channel attention module (DepCA) and a QR decomposition-based loss for promoting feature-wise orthogonality. We demonstrate that this orthogonalization enhances the separability of attention distributions, thereby refining and stabilizing the ranking of feature importance. This improvement in feature-wise ranking enhances our understanding of feature explainability in MTS. Furthermore, we develop metrics to evaluate global and class-specific feature importance. Our framework's efficacy is validated through extensive empirical analyses on two major public benchmarks and real-world datasets, both synthetic and self-collected, specifically designed to highlight class-wise discriminative features. The results confirm CAFO's robustness and informative capacity in assessing feature importance in MTS classification tasks. This study not only advances the understanding of feature-centric explanations in MTS but also sets a foundation for future explorations in feature-centric explanations. The codes are available at https://github.com/eai-lab/CAFO.
Seok-Ju Hahn, Yoontae Hwang, Junghye Lee, Seulki Lee 0002
KDD4
2024 HarmoSATE: Harmonized embedding-based self-attentive encoder to improve accuracy of privacy-preserving federated predictive analysis
abstract
Accurate privacy-preserving prediction using electronic health record (EHR) data distributed in multiple hospitals is essential to enable stakeholders related to healthcare services to obtain useful information without privacy leakage. In this paper, we propose harmonized embedding-based self-attentive encoder (HarmoSATE), which is a new method for privacy-preserving federated predictive analysis. We extract contextual embeddings of local institutions using Word2Vec, and then harmonize locally-trained embeddings using a neural network-based harmonization technique. The proposed method uses a deep representative encoder based on self-attention to learn complex and dynamic patterns inherent to harmonized embeddings of medical concepts. To evaluate our method, we implemented experiments using sequential medical codes collected from the Medical Information Mart for Intensive Care-III dataset in a distributed setting. It achieved a significant increase in average AUC, ranging from 3% to 8% depending on the experiments compared to baseline models, demonstrating superior prediction accuracy of a patient's diagnosis in the next admission. HarmoSATE can be a useful alternative to obtain accurate and practical results for various predictive tasks that use sensitive and distributed EHR data while preserving patients' privacy.
Taek-Ho Lee, Suhyeon Kim, Junghye Lee, Chi-Hyuck Jun
Inf. Sci.3
2023 Multitask Deep Learning for Human Activity, Speed, and Body Weight Estimation Using Commercial Smart Insoles
abstract
Healthcare professionals and individual users use wearable devices equipped with various sensors for healthcare management. Recently, the joint usage of artificial intelligence and these wearable sensors has played an essential role in healthcare management by providing a wide range of applications such as fitness tracking, gym activity monitoring, patient rehabilitation monitoring, and disease detection. These tasks eventually aim to enhance personal well-being and better manage the user’s physical health by monitoring different activity types and body weight changes. Here, we present an efficient multi-task learning framework based on commercial smart insoles that can solve three tasks related to physical health management: activity classification, speed estimation, and body weight estimation. Our multi-task framework converts the sensor data from the smart insole to a recurrence plot, which shows significant performance improvement compared to processing the raw time series data. In addition, we utilized a modified MobileNetV2 as our backbone network, which has a total parameter of less than 100K and a computational budget of 0.34G of multiply-accumulate operations. Furthermore, we collected a vast dataset from 72 users carrying out 16 experiments, which contains the largest number of people for multi-task learning purposes using smart insoles. Extensive experiments show that the proposed multi-task learning framework is extremely efficient while outperforming or leading to comparable performance against single-task models.
Hyewon Kang, Jaewan Yang, Haneul Jung, Seulki Lee 0002, Junghye Lee
IEEE Internet Things J.6
2023 Word2Vec-based efficient privacy-preserving shared representation learning for federated recommendation system in a cross-device setting
abstract
Recommendation systems have required centralized storage of user data, but due to privacy concerns, recent studies adopted federated learning (FL) that discloses intermediate statistics instead of raw data to build privacy-preserving federated recommendation systems. However, they suffer from inefficiencies in privacy-preserving mechanisms and inaccuracies in simple algorithms that ignore sequential information. This study proposes an extension of Word2Vec for a privacy-preserving federated sequential recommendation system (PPFSRS). This method exploits sequential information to generate contextual item representations for accurate recommendations while concealing privacy-sensitive features efficiently. Specifically, we mixed updates from negative samples to inhibit the direct leakage of purchased items from model updates. In addition, our method computes approximate model updates that can occur when sensitive features only belong to negative samples to prevent inference attacks. In experiments, we used benchmark datasets for recommendation and simulated highly distributed data such that each user stores historical data locally. While preserving privacy with reasonable complexity, the proposed method showed little degradation in recommendation performance compared to FL-based Word2Vec without privacy-preserving mechanisms. Utilizing contextual item representations trained by our method from highly distributed data will be a practical starting point for PPFSRS in a cross-device setting.
Taek-Ho Lee, Suhyeon Kim, Junghye Lee, Chi-Hyuck Jun
Inf. Sci.3
2022 Connecting Low-Loss Subspace for Personalized Federated Learning
abstract
Due to the curse of statistical heterogeneity across clients, adopting a personalized federated learning method has become an essential choice for the successful deployment of federated learning-based services. Among diverse branches of personalization techniques, a model mixture-based personalization method is preferred as each client has their own personalized model as a result of federated learning. It usually requires a local model and a federated model, but this approach is either limited to partial parameter exchange or requires additional local updates, each of which is helpless to novel clients and burdensome to the client's computational capacity. As the existence of a connected subspace containing diverse low-loss solutions between two or more independent deep networks has been discovered, we combined this interesting property with the model mixture-based personalized federated learning method for improved performance of personalization. We proposed SuPerFed, a personalized federated learning method that induces an explicit connection between the optima of the local and the federated model in weight space for boosting each other. Through extensive experiments on several benchmark datasets, we demonstrated that our method achieves consistent gains in both personalization performance and robustness to problematic scenarios possible in realistic services.
Seok-Ju Hahn, Minwoo Jeong, Junghye Lee
KDD3
2022 Risk score-embedded deep learning for biological age estimation: Development and validation
abstract
The health index measures a person’s overall health status which provides useful information for people to manage their health, so developing a precise and relevant health index is urgent. Currently, many researchers have studied the biological age (BA) estimation, one of the beneficial health indices, by applying machine learning and deep learning techniques to health data. However, most of them have focused on the chronological age prediction or basic latent feature extraction methods. In this paper, we present a new algorithm to estimate BA, called Risk Score-Embedded Autoencoder-based BA (RSAE-BA). RSAE-BA can provide an accurate health index by using deep representation learning with an individual’s health risk. We first proposed a notion of risk score (RS) calculation to monitor a person’s health risk. Then we extracted latent features by using an autoencoder embedding the RS, and used them to generate BA. To evaluate RSAE-BA, we presented a new BA validation method using the RS, which is applicable to both unlabeled and labeled data. We compared the results of RSAE-BA with existing methods, and demonstrated the accuracy of RSAE-BA and its applicability to predict disease incidence. We believe that RSAE-BA will be a useful alternative method to measure a person’s health.
Suhyeon Kim, Eun-Sol Lee, Chiehyeon Lim, Junghye Lee
Inf. Sci.5
2021 A multi-stage data mining approach for liquid bulk cargo volume analysis based on bill of lading data
abstract
Liquid bulk cargo (LBC) volume analysis has received considerably great attention recently since LBC is a valuable and high-demand cargo. Thus, it is important to establish an analysis system for LBC volume, as it can help inform strategies for port planning and management. Nevertheless, LBC volume analysis is a challenging task for researchers because trends in LBC volume are highly volatile and non-stationary. In this paper, a new framework for enabling informative LBC volume analysis based on bill of lading (BL) data is proposed, which consists of three parts: item segmentation, exploratory volume analysis, and volume prediction. Firstly, an innovative item segmentation system using item texts of BL data was developed, which can generate subcategory as well as category information of LBC items that existing system cannot provide. Next, exploratory volume analysis was performed to understand the volume characteristics of each categorized and subcategorized item in terms of geography and timeline. Lastly, manifold learning- and deep learning-based time series techniques were proposed to increase LBC volume prediction accuracy compared with existing statistical models. The experimental results for volume prediction show the accuracy increased by 34% and 18% in average at category and subcategory levels over baseline models. It is believed that our proposed method will be helpful for stakeholders in maritime logistics, giving them the insights that they need to make better decisions.
Suhyeon Kim, Wonho Sohn, Dongcheol Lim, Junghye Lee
Expert Syst. Appl.4
2021 An efficient multivariate feature ranking method for gene selection in high-dimensional microarray data
abstract
Classification of microarray data plays a significant role in the diagnosis and prediction of cancer. However, its high-dimensionality (>tens of thousands) compared to the number of observations (
Junghye Lee, In Young Choi, Chi-Hyuck Jun
Expert Syst. Appl.1
2021 Bilingual autoencoder-based efficient harmonization of multi-source private data for accurate predictive modeling
abstract
Sharing electronic health record data is essential for advanced analysis, but may put sensitive information at risk. Several studies have attempted to address this risk using contextual embedding, but with many hospitals involved, they are often inefficient and inflexible. Thus, we propose a bilingual autoencoder-based model to harmonize local embeddings in different spaces. Cross-hospital reconstruction of embeddings makes encoders map embeddings from hospitals to a shared space and align them spontaneously. We also suggest two-phase training to prevent distortion of embeddings during harmonization with hospitals that have biased information. In experiments, we used medical event sequences from the Medical Information Mart for Intensive Care-III dataset and simulated the situation of multiple hospitals. For evaluation, we measured the alignment of events from different hospitals and the prediction accuracy of a patient’s diagnosis in the next admission in three scenarios in which local embeddings do not work. The proposed method efficiently harmonizes embeddings in different spaces, increases prediction accuracy, and gives flexibility to include new hospitals, so is superior to previous methods in most cases. It will be useful in predictive tasks to utilize distributed data while preserving private information.
Taek-Ho Lee, Junghye Lee, Chi-Hyuck Jun
Inf. Sci.2
2020 Word2vec-based latent semantic analysis (W2V-LSA) for topic modeling: A study on blockchain technology trend analysis
abstract
Blockchain has become one of the core technologies in Industry 4.0. To help decision-makers establish action plans based on blockchain, it is an urgent task to analyze trends in blockchain technology. However, most of existing studies on blockchain trend analysis are based on effort demanding full-text investigation or traditional bibliometric methods whose study scope is limited to a frequency-based statistical analysis. Therefore, in this paper, we propose a new topic modeling method called Word2vec-based Latent Semantic Analysis (W2V-LSA), which is based on Word2vec and Spherical k-means clustering to better capture and represent the context of a corpus. We then used W2V-LSA to perform an annual trend analysis of blockchain research by country and time for 231 abstracts of blockchain-related papers published over the past five years. The performance of the proposed algorithm was compared to Probabilistic LSA, one of the common topic modeling techniques. The experimental results confirmed the usefulness of W2V-LSA in terms of the accuracy and diversity of topics by quantitative and qualitative evaluation. The proposed method can be a competitive alternative for better topic modeling to provide direction for future research in technology trend analysis and it is applicable to various expert systems related to text mining.
Suhyeon Kim, Haecheong Park, Junghye Lee
Expert Syst. Appl.3
2020 Markov blanket-based universal feature selection for classification and regression of mixed-type data
Junghye Lee, Jun-Yong Jeong, Chi-Hyuck Jun
Expert Syst. Appl.1
2020 Secure and Differentially Private Logistic Regression for Horizontally Distributed Data
abstract
Scientific collaborations benefit from sharing information and data from distributed sources, but protecting privacy is a major concern. Researchers, funders, and the public in general are getting increasingly worried about the potential leakage of private data. Advanced security methods have been developed to protect the storage and computation of sensitive data in a distributed setting. However, they do not protect against information leakage from the outcomes of data analyses. To address this aspect, studies on differential privacy (a state-of-the-art privacy protection framework) demonstrated encouraging results, but most of them do not apply to distributed scenarios. Combining security and privacy methodologies is a natural way to tackle the problem, but naive solutions may lead to poor analytical performance. In this paper, we introduce a novel strategy that combines differential privacy methods and homomorphic encryption techniques to achieve the best of both worlds. Using logistic regression (a popular model in biomedicine), we demonstrated the practicability of building secure and privacy-preserving models with high efficiency (less than 3 min) and good accuracy [<;1% of difference in the area under the receiver operating characteristic curve (AUC) against the global model] using a few real-world datasets.
Miran Kim, Junghye Lee, Lucila Ohno-Machado, Xiaoqian Jiang
IEEE Trans. Inf. Forensics Secur.2