EDBT 2026 Demo / reviewers in the wild / expert
Chi-Hyuck Jun
dblp:37/5898
· DBLP profile ↗
7ranked-venue papers in the field
0as first author
3since 2021 · last 2024
0000-0003-0911-7347ORCID · reported
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 6Data Mining & Knowledge Discovery · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | HarmoSATE: Harmonized embedding-based self-attentive encoder to improve accuracy of privacy-preserving federated predictive analysisabstractAccurate privacy-preserving prediction using electronic health record (EHR) data distributed in multiple hospitals is essential to enable stakeholders related to healthcare services to obtain useful information without privacy leakage. In this paper, we propose harmonized embedding-based self-attentive encoder (HarmoSATE), which is a new method for privacy-preserving federated predictive analysis. We extract contextual embeddings of local institutions using Word2Vec, and then harmonize locally-trained embeddings using a neural network-based harmonization technique. The proposed method uses a deep representative encoder based on self-attention to learn complex and dynamic patterns inherent to harmonized embeddings of medical concepts. To evaluate our method, we implemented experiments using sequential medical codes collected from the Medical Information Mart for Intensive Care-III dataset in a distributed setting. It achieved a significant increase in average AUC, ranging from 3% to 8% depending on the experiments compared to baseline models, demonstrating superior prediction accuracy of a patient's diagnosis in the next admission. HarmoSATE can be a useful alternative to obtain accurate and practical results for various predictive tasks that use sensitive and distributed EHR data while preserving patients' privacy. Taek-Ho Lee, Suhyeon Kim, Junghye Lee, Chi-Hyuck Jun |
Inf. Sci. | 4 |
| 2023 | Word2Vec-based efficient privacy-preserving shared representation learning for federated recommendation system in a cross-device settingabstractRecommendation systems have required centralized storage of user data, but due to privacy concerns, recent studies adopted federated learning (FL) that discloses intermediate statistics instead of raw data to build privacy-preserving federated recommendation systems. However, they suffer from inefficiencies in privacy-preserving mechanisms and inaccuracies in simple algorithms that ignore sequential information. This study proposes an extension of Word2Vec for a privacy-preserving federated sequential recommendation system (PPFSRS). This method exploits sequential information to generate contextual item representations for accurate recommendations while concealing privacy-sensitive features efficiently. Specifically, we mixed updates from negative samples to inhibit the direct leakage of purchased items from model updates. In addition, our method computes approximate model updates that can occur when sensitive features only belong to negative samples to prevent inference attacks. In experiments, we used benchmark datasets for recommendation and simulated highly distributed data such that each user stores historical data locally. While preserving privacy with reasonable complexity, the proposed method showed little degradation in recommendation performance compared to FL-based Word2Vec without privacy-preserving mechanisms. Utilizing contextual item representations trained by our method from highly distributed data will be a practical starting point for PPFSRS in a cross-device setting. Taek-Ho Lee, Suhyeon Kim, Junghye Lee, Chi-Hyuck Jun |
Inf. Sci. | 4 |
| 2021 | Bilingual autoencoder-based efficient harmonization of multi-source private data for accurate predictive modelingabstractSharing electronic health record data is essential for advanced analysis, but may put sensitive information at risk. Several studies have attempted to address this risk using contextual embedding, but with many hospitals involved, they are often inefficient and inflexible. Thus, we propose a bilingual autoencoder-based model to harmonize local embeddings in different spaces. Cross-hospital reconstruction of embeddings makes encoders map embeddings from hospitals to a shared space and align them spontaneously. We also suggest two-phase training to prevent distortion of embeddings during harmonization with hospitals that have biased information. In experiments, we used medical event sequences from the Medical Information Mart for Intensive Care-III dataset and simulated the situation of multiple hospitals. For evaluation, we measured the alignment of events from different hospitals and the prediction accuracy of a patient’s diagnosis in the next admission in three scenarios in which local embeddings do not work. The proposed method efficiently harmonizes embeddings in different spaces, increases prediction accuracy, and gives flexibility to include new hospitals, so is superior to previous methods in most cases. It will be useful in predictive tasks to utilize distributed data while preserving private information. Taek-Ho Lee, Junghye Lee, Chi-Hyuck Jun |
Inf. Sci. | 3 |
| 2020 | Regularization-based model tree for multi-output regression
Jun-Yong Jeong, Ju-Seok Kang, Chi-Hyuck Jun |
Inf. Sci. | 3 |
| 2018 | Variable Selection and Task Grouping for Multi-Task LearningabstractWe consider multi-task learning, which simultaneously learns related prediction tasks, to improve generalization performance. We factorize a coefficient matrix as the product of two matrices based on a low-rank assumption. These matrices have sparsities to simultaneously perform variable selection and learn and overlapping group structure among the tasks. The resulting bi-convex objective function is minimized by alternating optimization, where sub-problems are solved using alternating direction method of multipliers and accelerated proximal gradient descent. Moreover, we provide the performance bound of the proposed method. The effectiveness of the proposed method is validated for both synthetic and real-world datasets. Jun-Yong Jeong, Chi-Hyuck Jun |
KDD | 2 |
| 2017 | Instance categorization by support vector machines to adjust weights in AdaBoost for imbalanced data classification
Wonji Lee, Chi-Hyuck Jun, Jong-Seok Lee |
Inf. Sci. | 2 |
| 2014 | Designing of a new monitoring t-chart using repetitive sampling
Muhammad Aslam 0002, Nasrullah Khan, Muhammad Azam 0001, Chi-Hyuck Jun |
Inf. Sci. | 4 |