Chi-Hyuck Jun

dblp:37/5898 · DBLP profile ↗
← Back
30ranked-venue papers
2as first author
7since 2021 · last 2024
0000-0003-0911-7347ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 4 since 2021Databases, data management, data science and information retrieval · 7 · 3 since 2021Computer networks · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2024 HarmoSATE: Harmonized embedding-based self-attentive encoder to improve accuracy of privacy-preserving federated predictive analysis
abstract
Accurate privacy-preserving prediction using electronic health record (EHR) data distributed in multiple hospitals is essential to enable stakeholders related to healthcare services to obtain useful information without privacy leakage. In this paper, we propose harmonized embedding-based self-attentive encoder (HarmoSATE), which is a new method for privacy-preserving federated predictive analysis. We extract contextual embeddings of local institutions using Word2Vec, and then harmonize locally-trained embeddings using a neural network-based harmonization technique. The proposed method uses a deep representative encoder based on self-attention to learn complex and dynamic patterns inherent to harmonized embeddings of medical concepts. To evaluate our method, we implemented experiments using sequential medical codes collected from the Medical Information Mart for Intensive Care-III dataset in a distributed setting. It achieved a significant increase in average AUC, ranging from 3% to 8% depending on the experiments compared to baseline models, demonstrating superior prediction accuracy of a patient's diagnosis in the next admission. HarmoSATE can be a useful alternative to obtain accurate and practical results for various predictive tasks that use sensitive and distributed EHR data while preserving patients' privacy.
Taek-Ho Lee, Suhyeon Kim, Junghye Lee, Chi-Hyuck Jun
Inf. Sci.4
2023 Dynamic mutual information-based feature selection for multi-label learning
abstract
In classification problems, feature selection is used to identify important input features to reduce the dimensionality of the input space while improving or maintaining classification performance. Traditional feature selection algorithms are designed to handle single-label learning, but classification problems have recently emerged in multi-label domain. In this study, we propose a novel feature selection algorithm for classifying multi-label data. This proposed method is based on dynamic mutual information, which can handle redundancy among features controlling the input space. We compare the proposed method with some existing problem transformation and algorithm adaptation methods applied to real multi-label datasets using the metrics of multi-label accuracy and hamming loss. The results show that the proposed method demonstrates more stable and better performance for nearly all multi-label datasets.
Kyung-jun Kim, Chi-Hyuck Jun
Intell. Data Anal.2
2023 Word2Vec-based efficient privacy-preserving shared representation learning for federated recommendation system in a cross-device setting
abstract
Recommendation systems have required centralized storage of user data, but due to privacy concerns, recent studies adopted federated learning (FL) that discloses intermediate statistics instead of raw data to build privacy-preserving federated recommendation systems. However, they suffer from inefficiencies in privacy-preserving mechanisms and inaccuracies in simple algorithms that ignore sequential information. This study proposes an extension of Word2Vec for a privacy-preserving federated sequential recommendation system (PPFSRS). This method exploits sequential information to generate contextual item representations for accurate recommendations while concealing privacy-sensitive features efficiently. Specifically, we mixed updates from negative samples to inhibit the direct leakage of purchased items from model updates. In addition, our method computes approximate model updates that can occur when sensitive features only belong to negative samples to prevent inference attacks. In experiments, we used benchmark datasets for recommendation and simulated highly distributed data such that each user stores historical data locally. While preserving privacy with reasonable complexity, the proposed method showed little degradation in recommendation performance compared to FL-based Word2Vec without privacy-preserving mechanisms. Utilizing contextual item representations trained by our method from highly distributed data will be a practical starting point for PPFSRS in a cross-device setting.
Taek-Ho Lee, Suhyeon Kim, Junghye Lee, Chi-Hyuck Jun
Inf. Sci.4
2021 An Improved Time-Series Forecasting Model Using Time Series Decomposition and GRU Architecture
Hyun Jae Jo, Won Joong Kim, Hyo Kyung Goh, Chi-Hyuck Jun
ICONIP (6)4
2021 An efficient multivariate feature ranking method for gene selection in high-dimensional microarray data
abstract
Classification of microarray data plays a significant role in the diagnosis and prediction of cancer. However, its high-dimensionality (>tens of thousands) compared to the number of observations (
Junghye Lee, In Young Choi, Chi-Hyuck Jun
Expert Syst. Appl.3
2021 Mutual information-based multi-output tree learning algorithm
abstract
A tree model with low time complexity can support the application of artificial intelligence to industrial systems. Variable selection based tree learning algorithms are more time efficient than existing Classification and Regression Tree (CART) algorithms. To our best knowledge, there is no attempt to deal with categorical input variable in variable selection based multi-output tree learning. Also, in the case of multi-output regression tree, a conventional variable selection based algorithm is not suitable to large datasets. We propose a mutual information-based multi-output tree learning algorithm that consists of variable selection and split optimization. The proposed method discretizes each variable based on k-means into 2–4 clusters and selects the variable for splitting based on the discretized variables using mutual information. This variable selection component has relatively low time complexity and can be applied regardless of output dimension and types. The proposed split optimization component is more efficient than an exhaustive search. The performance of the proposed tree learning algorithm is similar to or better than that of a multi-output version of CART algorithm on a specific dataset. In addition, with a large dataset, the time complexity of the proposed algorithm is significantly reduced compared to a CART algorithm.
Hyun-Seok Kang, Chi-Hyuck Jun
Intell. Data Anal.2
2021 Bilingual autoencoder-based efficient harmonization of multi-source private data for accurate predictive modeling
abstract
Sharing electronic health record data is essential for advanced analysis, but may put sensitive information at risk. Several studies have attempted to address this risk using contextual embedding, but with many hospitals involved, they are often inefficient and inflexible. Thus, we propose a bilingual autoencoder-based model to harmonize local embeddings in different spaces. Cross-hospital reconstruction of embeddings makes encoders map embeddings from hospitals to a shared space and align them spontaneously. We also suggest two-phase training to prevent distortion of embeddings during harmonization with hospitals that have biased information. In experiments, we used medical event sequences from the Medical Information Mart for Intensive Care-III dataset and simulated the situation of multiple hospitals. For evaluation, we measured the alignment of events from different hospitals and the prediction accuracy of a patient’s diagnosis in the next admission in three scenarios in which local embeddings do not work. The proposed method efficiently harmonizes embeddings in different spaces, increases prediction accuracy, and gives flexibility to include new hospitals, so is superior to previous methods in most cases. It will be useful in predictive tasks to utilize distributed data while preserving private information.
Taek-Ho Lee, Junghye Lee, Chi-Hyuck Jun
Inf. Sci.3
2020 Markov blanket-based universal feature selection for classification and regression of mixed-type data
Junghye Lee, Jun-Yong Jeong, Chi-Hyuck Jun
Expert Syst. Appl.3
2020 Regularization-based model tree for multi-output regression
Jun-Yong Jeong, Ju-Seok Kang, Chi-Hyuck Jun
Inf. Sci.3
2019 Machine learning models based on the dimensionality reduction of standard automated perimetry data for glaucoma diagnosis
Sudong Lee, Ji-Hyung Lee, Young-Geun Choi, Hee-Cheon You, Ja-Heon Kang, Chi-Hyuck Jun
Artif. Intell. Medicine6
2019 Hybrid data stream clustering by controlling decision error
abstract
Data stream clustering is an unsupervised learning method for sequential data. Data stream clustering has some challenging issues, such as handling limited memory, dealing with evolving clusters, and detecting noise data. We propose a hybrid data stream clustering method that combines model-based c lustering and density-based clustering. The proposed method finds evolving clusters quickly and obtains cluster information easily. We use multiple hypothesis testing to handle noise data by controlling a decision error. In this testing method, we employ the positive false discovery rate as the decision error. We use a density-based algorithm to discover cluster evolution from newly arrived data. Then, we estimate a Gaussian mixture model and update the clustering results by combining past cluster information and the cluster information for newly arrived data. We applied the proposed method to several synthetic and real datasets. The experimental results demonstrate that the proposed method works effectively for a data stream that includes noise data. In addition, the proposed method yields robust results relative to input parameters compared to an existing density-based data stream clustering method.
Jeonghwa Lee, Taek-Ho Lee, Chi-Hyuck Jun
Intell. Data Anal.3
2018 Variable Selection and Task Grouping for Multi-Task Learning
abstract
We consider multi-task learning, which simultaneously learns related prediction tasks, to improve generalization performance. We factorize a coefficient matrix as the product of two matrices based on a low-rank assumption. These matrices have sparsities to simultaneously perform variable selection and learn and overlapping group structure among the tasks. The resulting bi-convex objective function is minimized by alternating optimization, where sub-problems are solved using alternating direction method of multipliers and accelerated proximal gradient descent. Moreover, we provide the performance bound of the proposed method. The effectiveness of the proposed method is validated for both synthetic and real-world datasets.
Jun-Yong Jeong, Chi-Hyuck Jun
KDD2
2018 Rough set model based feature selection for mixed-type data with feature space decomposition
Kyung-jun Kim, Chi-Hyuck Jun
Expert Syst. Appl.2
2018 Fast incremental learning of logistic model tree using least angle regression
Sudong Lee, Chi-Hyuck Jun
Expert Syst. Appl.2
2017 Instance categorization by support vector machines to adjust weights in AdaBoost for imbalanced data classification
Wonji Lee, Chi-Hyuck Jun, Jong-Seok Lee
Inf. Sci.2
2014 Improved churn prediction in telecommunication industry by analyzing a large network
Kyoungok Kim, Chi-Hyuck Jun, Jaewook Lee 0001
Expert Syst. Appl.2
2014 Designing of a new monitoring t-chart using repetitive sampling
Muhammad Aslam 0002, Nasrullah Khan, Muhammad Azam 0001, Chi-Hyuck Jun
Inf. Sci.4
2013 Ranking evaluation of institutions based on a Bayesian network having a latent variable
Chi-Hyuck Jun
Knowl. Based Syst.2
2013 PCA-based high-dimensional noisy data clustering via control of decision errors
Jeonghwa Lee, Chi-Hyuck Jun
Knowl. Based Syst.2
2012 Learning Bayesian network structure using Markov blanket decomposition
Anh Tuan Bui, Chi-Hyuck Jun
Pattern Recognit. Lett.2
2011 Discriminant analysis of binary data following multivariate Bernoulli distribution
Chi-Hyuck Jun
Expert Syst. Appl.2
2011 Stability-based validation of bicluster solutions
Youngrok Lee, Jeonghwa Lee, Chi-Hyuck Jun
Pattern Recognit.3
2010 A causal discovery algorithm using multiple regressions
Young-Hun Choi, Chi-Hyuck Jun
Pattern Recognit. Lett.2
2009 A simple and fast algorithm for K-medoids clustering
Hae-Sang Park, Chi-Hyuck Jun
Expert Syst. Appl.2
2009 Classifying genes according to predefined patterns by controlling false discovery rate
Hae-Sang Park, Chi-Hyuck Jun, Joo-Yeon Yoo
Expert Syst. Appl.2
2008 Flexible patient rule induction method for optimizing process variables in discrete type
Il-Gyo Chong, Chi-Hyuck Jun
Expert Syst. Appl.2
2006 Variables sampling plans for Weibull distributed lifetimes under sudden death testing
abstract
Sudden death testing can be utilized for deciding upon the lot acceptance of manufactured parts. Variables single, and double sampling plans are proposed for the lot acceptance of parts whose life follows a Weibull distribution with known shape parameter. The proposed plans are different from the existing ones in that the lot acceptance criteria do not depend on the estimated scale parameter. Design parameters of both sampling plans are determined by using the usual two-point approach. The number of groups is determined independently of the group size, and even independently of the shape parameter. Also, the double sampling plan can reduce the average number of groups required. The effects of mis-specification of the shape parameter on the probability of accepting the lots under the single sampling plan are analyzed & discussed.
Chi-Hyuck Jun, Saminathan Balamurali
IEEE Trans. Reliab.1
2006 Frequency Insertion Strategy for Channel Assignment Problem
Won-young Shin, Soo Chang, Jaewook Lee 0001, Chi-Hyuck Jun
Wirel. Networks4
2005 Classification-based collaborative filtering using market basket data
Jong-Seok Lee, Chi-Hyuck Jun, Jaewook Lee 0001
Expert Syst. Appl.2
2000 Teletraffic - Theory and Applications: H. Akimaru, K. Kawashima
Chi-Hyuck Jun
Comput. Commun.1