Hyunjung Shin

dblp:s/HyunjungShin · also Hyunjung (Helen) Shin · DBLP profile ↗
← Back
61ranked-venue papers
14as first author
16since 2021 · last 2025
0000-0001-8347-8277ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 8 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 29 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Security and privacy · 2Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 BIGPN: Biologically informed graph propagational network for plasma proteomic profiling of neurodegenerative biomarkers
Sunghong Park, Dong-Gi Lee, Masaud Shah, Hyunjung Shin, Hyun Goo Woo
Artif. Intell. Medicine5
2025 NeuroFANN: identification of neuropathological subtypes in dementia with plasma proteins by using functionally annotated neural network
abstract
Dementia diagnosis relies on identifying neuropathological features, such as beta-amyloid (Aβ) deposition, medial temporal lobe atrophy (MTA), and white matter hyperintensity (WMH). Recently, plasma protein biomarkers have emerged as a cost-effective and less invasive tool for identifying neuropathological features, enhanced by machine learning (ML) for precise diagnosis. However, most ML studies fail to account for protein-protein interactions (PPIs) and synergetic effects between proteins, overlooking their collective contributions to disease mechanisms. Additionally, the lack of consideration for functional properties may result in the redundant and imbalanced representation of proteins and their functions, potentially limiting the effectiveness of dementia diagnosis. In this study, we propose NeuroFANN, a method designed to classify three neuropathological subtypes in dementia-positivity for Aβ, MTA, and WMH-using plasma protein biomarkers. A key feature of NeuroFANN is the combination of the PPI network-based synergetic effects with the functional annotation-based protein biomarker clustering. NeuroFANN extracts synergetic effects by propagating independent effects of proteins across the PPI network, which are then aggregated in functional protein clusters, thereby enabling global PPI awareness and capturing the biological properties of protein biomarkers. From a South Korean cohort, 54 proteins were identified as plasma protein biomarkers for dementia subtypes and grouped into 16 clusters. NeuroFANN outperformed comparison methods in classifying dementia subtypes, with its core components validated as key contributors to superior performance. Additionally, the risk scores predicted by NeuroFANN showed a strong association with longitudinal cognitive decline, demonstrating its potential as a valuable diagnostic tool in clinical settings.
Sunghong Park, Doyoon Kim, Ji-Hye Choi, Changhyung Hong, Sangjoon Son, Hyunwoong Roh, Hyunjung Shin, Hyun Goo Woo
Briefings Bioinform.7
2025 PPIxGPN: plasma proteomic profiling of neurodegenerative biomarkers with protein-protein interaction-based eXplainable graph propagational network
abstract
Neurodegenerative diseases involve progressive neuronal dysfunction, requiring the identification of specific pathological features for accurate diagnosis. While cerebrospinal fluid analysis and neuroimaging are commonly used, their invasive nature and high costs limit clinical applicability. Recently advances in plasma proteomics offer a less invasive and cost-effective alternative, further enhanced by machine learning (ML). However, most ML-based studies overlook synergetic effects from protein-protein interactions (PPIs), which play a key role in disease mechanisms. Although graph convolutional network and its extensions can utilize PPIs, they rely on locality-based feature aggregation, overlooking essential components and emphasizing noisy interactions. Moreover, expanding those methods to cover broader PPIs results in complex model architectures that reduce explainability, which is crucial in medical ML models for clinical decision-making. To address these challenges, we propose Protein-Protein Interaction-based eXplainable Graph Propagational Network (PPIxGPN), a novel ML model designed for plasma proteomic profiling of neurodegenerative biomarkers. PPIxGPN captures synergetic effects between proteins by integrating PPIs with independent effects of proteins, leveraging globality-based feature aggregation to represent comprehensive PPI properties. This process is implemented using a single graph propagational layer, enabling PPIxGPN to be configured by shallow architecture, thereby PPIxGPN ensures high model explainability, enhancing clinical applicability by providing interpretable outputs. Experimental validation on the UK Biobank dataset demonstrated the superior performance of PPIxGPN in neurodegenerative risk prediction, outperforming comparison methods. Furthermore, the explainability of PPIxGPN facilitated detailed analyses of the discriminative significance of synergistic effects, the predictive importance of proteins, and the longitudinal changes in biomarker profiles, highlighting its clinical relevance.
Sunghong Park, Dong-Gi Lee, Seung Ho Kim, Hyeon Jin Hwang, Hyunjung Shin, Hyun Goo Woo
Briefings Bioinform.6
2025 Prospective domain adaptation for longitudinal data
Sunghong Park, Jeongheun Yeon, Dong-Gi Lee, Sangjoon Son, Hyun Goo Woo, Hyunjung Shin
Knowl. Based Syst.6
2025 Explainable multiplex graph propagational network with multimodal neuroimage integration for dementia subtype diagnosis
Sunghong Park, Dong-gi Lee, Do Kyoon Kim, Yonghyun Nam, Bumhee Park, Narae Kim, Seulgi Lee, Changhyung Hong, Sangjoon Son, Hyunwoong Roh, Hyun Goo Woo, Hyunjung Shin
Neural Networks13
2025 Extremely fast graph integration for semi-supervised learning via Gaussian fields with Neumann approximation
Taehwan Yun, Myung Jun Kim, Hyunjung Shin
Pattern Recognit.3
2024 Identification of molecular subtypes of dementia by using blood-proteins interaction-aware graph propagational network
abstract
Plasma protein biomarkers have been considered promising tools for diagnosing dementia subtypes due to their low variability, cost-effectiveness, and minimal invasiveness in diagnostic procedures. Machine learning (ML) methods have been applied to enhance accuracy of the biomarker discovery. However, previous ML-based studies often overlook interactions between proteins, which are crucial in complex disorders like dementia. While protein-protein interactions (PPIs) have been used in network models, these models often fail to fully capture the diverse properties of PPIs due to their local awareness. This drawback increases the chance of neglecting critical components and magnifying the impact of noisy interactions. In this study, we propose a novel graph-based ML model for dementia subtype diagnosis, the graph propagational network (GPN). By propagating the independent effect of plasma proteins on PPI network, the GPN extracts the globally interactive effects between proteins. Experimental results showed that the interactive effect between proteins yielded to further clarify the differences between dementia subtype groups and contributed to the performance improvement where the GPN outperformed existing methods by 10.4% on average.
Sunghong Park, Changhyung Hong, Sangjoon Son, Hyunwoong Roh, Doyoon Kim, Hyunjung Shin, Hyun Goo Woo
Briefings Bioinform.6
2024 Network based Enterprise Profiling with Semi-Supervised Learning
Sunghong Park, Kanghee Park, Hyunjung Shin
Expert Syst. Appl.3
2024 In-house data adaptation to public data: Multisite MRI harmonization to predict Alzheimer's disease conversion
Sunghong Park, Sangjoon Son, Kanghee Park, Yonghyun Nam, Hyunjung Shin
Expert Syst. Appl.5
2024 Mutual Domain Adaptation
Sunghong Park, Myung Jun Kim, Kanghee Park, Hyunjung Shin
Pattern Recognit.4
2023 Discovering comorbid diseases using an inter-disease interactivity network based on biobank-scale PheWAS data
abstract
MOTIVATION: Understanding comorbidity is essential for disease prevention, treatment and prognosis. In particular, insight into which pairs of diseases are likely or unlikely to co-occur may help elucidate the potential relationships between complex diseases. Here, we introduce the use of an inter-disease interactivity network to discover/prioritize comorbidities. Specifically, we determine disease associations by accounting for the direction of effects of genetic components shared between diseases, and categorize those associations as synergistic or antagonistic. We further develop a comorbidity scoring algorithm to predict whether diseases are more or less likely to co-occur in the presence of a given index disease. This algorithm can handle networks that incorporate relationships with opposite signs. RESULTS: We finally investigate inter-disease associations among 427 phenotypes in UK Biobank PheWAS data and predict the priority of comorbid diseases. The predicted comorbidities were verified using the UK Biobank inpatient electronic health records. Our findings demonstrate that considering the interaction of phenotype associations might be helpful in better predicting comorbidity. AVAILABILITY AND IMPLEMENTATION: The source code and data of this study are available at https://github.com/dokyoonkimlab/DiseaseInteractiveNetwork. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yonghyun Nam, Sang-Hyuk Jung, Jae-Seung Yun, Vivek Sriram, Pankhuri Singhal, Marta Byrska-Bishop, Anurag Verma, Hyunjung Shin, Woong-Yang Park, Hong-Hee Won, Do Kyoon Kim
Bioinform.8
2023 Fast Prediction for Criminal Suspects through Neighbor Mutual Information-Based Latent Network
abstract
One of the interesting characteristics of crime data is that criminal cases are often interrelated. Criminal acts may be similar, and similar incidents may occur consecutively by the same offender or by the same criminal group. Among many machine learning algorithms, network‐based approaches are well‐suited to reflect these associative characteristics. Applying machine learning to criminal networks composed of cases and their associates can predict potential suspects. This narrows the scope of an investigation, saving time and cost. However, inference from criminal networks is not straightforward as it requires being able to process complex information entangled with case‐to‐case, person‐to‐person, and case‐to‐person connections. Besides, being useful at a crime scene requires urgency. However, predictions from network‐based machine learning algorithms are generally slow when the data is large and complex in structure. These limitations are an immediate barrier to any practical use of the criminal network geared by machine learning. In this study, we propose a criminal network‐based suspect prediction framework. The network we designed has a unique structure, such as a sandwich panel, in which one side is a network of crime cases and the other side is a network of people such as victims, criminals, and witnesses. The two networks are connected by relationships between the case and the persons involved in the case. The proposed method is then further developed into a fast inference algorithm for large‐scale criminal networks. Experiments on benchmark data showed that the fast inference algorithm significantly reduced execution time while still being competitive in performance comparisons of the original algorithm and other existing approaches. Based on actual crime data provided by the Korean National Police, several examples of how the proposed method is applied are shown.
Jong Ho Jhee, Myung Jun Kim, Myeonggeon Park, Jeongheun Yeon, Hyunjung Shin
Int. J. Intell. Syst.5
2023 Prospective classification of Alzheimer's disease conversion from mild cognitive impairment
Sunghong Park, Changhyung Hong, Dong-Gi Lee, Kanghee Park, Hyunjung Shin
Neural Networks5
2022 Clinical Decision Support to Reduce Nephrotoxic Medication-Associated Acute Kidney Injury in Non-Critically Ill Hospitalized Children
Bayley Bennett, Hyunjung Shin, Swaminathan Kandaswamy, Evan Orenstein, Edwin Ray
AMIA2
2021 Polypharmacy side-effect prediction with enhanced interpretability based on graph feature attention network
abstract
MOTIVATION: Polypharmacy side effects should be carefully considered for new drug development. However, considering all the complex drug-drug interactions that cause polypharmacy side effects is challenging. Recently, graph neural network (GNN) models have handled these complex interactions successfully and shown great predictive performance. Nevertheless, the GNN models have difficulty providing intelligible factors of the prediction for biomedical and pharmaceutical domain experts. METHOD: A novel approach, graph feature attention network (GFAN), is presented for interpretable prediction of polypharmacy side effects by emphasizing target genes differently. To artificially simulate polypharmacy situations, where two different drugs are taken together, we formulated a node classification problem by using the concept of line graph in graph theory. RESULTS: Experiments with benchmark datasets validated interpretability of the GFAN and demonstrated competitive performance with the graph attention network in a previous work. And the specific cases in the polypharmacy side-effect prediction experiments showed that the GFAN model is capable of very sensitively extracting the target genes for each side-effect prediction. AVAILABILITY AND IMPLEMENTATION: https://github.com/SunjooBang/Polypharmacy-side-effect-prediction.
Sunjoo Bang, Jong Ho Jhee, Hyunjung Shin
Bioinform.3
2021 Customer sentiment analysis with more sensibility
Sunghong Park, Kanghee Park, Hyunjung Shin
Eng. Appl. Artif. Intell.4
2020 Dementia key gene identification with multi-layered SNP-gene-disease network
abstract
MOTIVATION: Recently, various approaches for diagnosing and treating dementia have received significant attention, especially in identifying key genes that are crucial for dementia. If the mutations of such key genes could be tracked, it would be possible to predict the time of onset of dementia and significantly aid in developing drugs to treat dementia. However, gene finding involves tremendous cost, time and effort. To alleviate these problems, research on utilizing computational biology to decrease the search space of candidate genes is actively conducted. In this study, we propose a framework in which diseases, genes and single-nucleotide polymorphisms are represented by a layered network, and key genes are predicted by a machine learning algorithm. The algorithm utilizes a network-based semi-supervised learning model that can be applied to layered data structures. RESULTS: The proposed method was applied to a dataset extracted from public databases related to diseases and genes with data collected from 186 patients. A portion of key genes obtained using the proposed method was verified in silico through PubMed literature, and the remaining genes were left as possible candidate genes. AVAILABILITY AND IMPLEMENTATION: The code for the framework will be available at http://www.alphaminers.net/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Dong-Gi Lee, Myungjun Kim, Sangjoon Son, Changhyung Hong, Hyunjung Shin
Bioinform.5
2020 Inference on historical factions based on multi-layered network of historical figures
Myungjun Kim, Dong-Gi Lee, Sangkuk Lee, Hyunjung Shin
Expert Syst. Appl.5
2020 Corrections to "Comorbidity Scoring With Causal Disease Networks"
abstract
Presents corrections to author information in the above named paper.
Jong Ho Jhee, Sunjoo Bang, Dong-Gi Lee, Hyunjung Shin
IEEE ACM Trans. Comput. Biol. Bioinform.4
2019 Disease gene identification based on generic and disease-specific genome networks
abstract
SUMMARY: Immune diseases have a strong genetic component with Mendelian patterns of inheritance. While the tight association has been a major understanding in the underlying pathophysiology for the category of immune diseases, the common features of these diseases remain unclear. Based on the potential commonality among immune genes, we design Gene Ranker for key gene identification. Gene Ranker is a network-based gene scoring algorithm that initially constructs a backbone network based on protein interactions. Patient gene expression networks are added into the network. An add-on process screens the networks of weighted gene co-expression network analysis (WGCNA) on the samples of immune patients. Gene Ranker is disease-specific; however, any WGCNA network that passes the screening procedure can be added on. With the constructed network, it employs the semi-supervised learning for gene scoring. RESULTS: The proposed method was applied to immune diseases. Based on the resulting scores, Gene Ranker identified potential key genes in immune diseases. In scoring validation, an average area under the receiver operating characteristic curve of 0.82 was achieved, which is a significant increase from the reference average of 0.76. Highly ranked genes were verified through retrieval and review of 27 million PubMed literatures. As a typical case, 20 potential key genes in rheumatoid arthritis were identified: 10 were de facto genes and the remaining were novel. AVAILABILITY AND IMPLEMENTATION: Gene Ranker is available at http://www.alphaminers.net/GeneRanker/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yonghyun Nam, Jong Ho Jhee, Ji-Hyun Lee, Hyunjung Shin
Bioinform.5
2019 Disease Pathway Cut for Multi-Target drugs
abstract
BACKGROUND: Biomarker discovery studies have been moving the focus from a single target gene to a set of target genes. However, the number of target genes in a drug should be minimum to avoid drug side-effect or toxicity. But still, the set of target genes should effectively block all possible paths of disease progression. METHODS: In this article, we propose a network based computational analysis for target gene identification for multi-target drugs. The min-cut algorithm is employed to cut all the paths from onset genes to apoptotic genes on a disease pathway. If the pathway network is completely disconnected, development of disease will not further go on. The genes corresponding to the end points of the cutting edges are identified as candidate target genes for a multi-target drug. RESULTS AND CONCLUSIONS: The proposed method was applied to 10 disease pathways. In total, thirty candidate genes were suggested. The result was validated with gene set enrichment analysis software, PubMed literature review and de facto drug targets.
Sunjoo Bang, Sangjoon Son, Hyunjung Shin
BMC Bioinform.4
2019 Drug repurposing with network reinforcement
abstract
BACKGROUND: Drug repurposing has been motivated to ameliorate low probability of success in drug discovery. For the recent decade, many in silico attempts have received primary attention as a first step to alleviate the high cost and longevity. Such study has taken benefits of abundance, variety, and easy accessibility of pharmaceutical and biomedical data. Utilizing the research friendly environment, in this study, we propose a network-based machine learning algorithm for drug repurposing. Particularly, we show a framework on how to construct a drug network, and how to strengthen the network by employing multiple/heterogeneous types of data. RESULTS: The proposed method consists of three steps. First, we construct a drug network from drug-target protein information. Then, the drug network is reinforced by utilizing drug-drug interaction knowledge on bioactivity and/or medication from literature databases. Through the enhancement, the number of connected nodes and the number of edges between them become more abundant and informative, which can lead to a higher probability of success of in silico drug repurposing. The enhanced network recommends candidate drugs for repurposing through drug scoring. The scoring process utilizes graph-based semi-supervised learning to determine the priority of recommendations. CONCLUSIONS: The drug network is reinforced in terms of the coverage and connections of drugs: the drug coverage increases from 4738 to 5442, and the drug-drug associations as well from 808,752 to 982,361. Along with the network enhancement, drug recommendation becomes more reliable: AUC of 0.89 was achieved lifted from 0.79. For typical cases, 11 recommended drugs were shown for vascular dementia: amantadine, conotoxin GV, tenocyclidine, cycloeucine, etc.
Yonghyun Nam, Myungjun Kim, Hang-Seok Chang, Hyunjung Shin
BMC Bioinform.4
2019 The translational network for metabolic disease - from protein interaction to disease co-occurrence
abstract
BACKGROUND: The recent advances in human disease network have provided insights into establishing the relationships between the genotypes and phenotypes of diseases. In spite of the great progress, it yet remains as only a map of topologies between diseases, but not being able to be a pragmatic diagnostic/prognostic tool in medicine. It can further evolve from a map to a translational tool if it equips with a function of scoring that measures the likelihoods of the association between diseases. Then, a physician, when practicing on a patient, can suggest several diseases that are highly likely to co-occur with a primary disease according to the scores. In this study, we propose a method of implementing 'n-of-1 utility' (n potential diseases of one patient) to human disease network-the translational disease network. RESULTS: We first construct a disease network by introducing the notion of walk in graph theory to protein-protein interaction network, and then provide a scoring algorithm quantifying the likelihoods of disease co-occurrence given a primary disease. Metabolic diseases, that are highly prevalent but have found only a few associations in previous studies, are chosen as entries of the network. CONCLUSIONS: The proposed method substantially increased connectivity between metabolic diseases and provided scores of co-occurring diseases. The increase in connectivity turned the disease network info-richer. The result lifted the AUC of random guessing up to 0.72 and appeared to be concordant with the existing literatures on disease comorbidity.
Yonghyun Nam, Dong-Gi Lee, Sunjoo Bang, Ju Han Kim, Jae-Hoon Kim 0004, Hyunjung Shin
BMC Bioinform.6
2019 Semi-supervised learning for hierarchically structured networks
Myungjun Kim, Dong-Gi Lee, Hyunjung Shin
Pattern Recognit.3
2019 Comorbidity Scoring with Causal Disease Networks
abstract
In recent years, there has been numerous studies constructing a disease network with diverse sources of data. Many researchers attempted to extend the usage of the disease network by employing machine learning algorithms on various problems such as prediction of comorbidity. The relations between diseases can further be specified into causal relations. When causality is laid on the edges in the network, prediction for comorbid diseases can be more improved. However, not many machine learning algorithms have been developed to concern causality. In this study, we exploit a network based machine learning algorithm that generates comorbidity scores from a causal disease network. In order to find comorbid diseases, semi-supervised scoring for causal networks is proposed. It computes scores of entire nodes in the network when a specific node is labeled. Each score is calculated one at a time and affects to the others along causal edges. The algorithm iterates until it converges. We compared the scoring results of the causal disease network and those of simple association network. As a gold standard, we referenced the values of relative risk from prevalence database, HuDiNe. Scoring by the proposed method provides clearer distinguishability between the top-ranked diseases in the comorbidity list. This is a benefit because it allows the choosing of the most significant ones on an easier fashion. To present typical use of the resulting list, comorbid diseases of Huntington disease and pnuemonia are validated via PubMed literature, respectively.
Jong Ho Jhee, Sunjoo Bang, Dong-Gi Lee, Hyunjung Shin
IEEE ACM Trans. Comput. Biol. Bioinform.4
2018 Historical inference based on semi-supervised learning
abstract
In the past, most historical research has been manually carried out by exploring historical facts reading between the lines of documents. Nowadays, historical big data has become electronically available and advances in machine learning techniques allow us to analyze the vast amount of historical data. From a historical perspective, making inferences about political stances of historical figures is important for grasping historical rivalries and power structures of an era. Thus, in this paper, we propose an approach to the systematic inference of power mechanisms based on a human network constructed from historical data. In this network, humans are linked according to the degree of kinship using genealogy records, and identified by political stances on agendas recorded in the annals of a dynasty as a political force. And then, a machine learning algorithm, semi-supervised learning, classifies humans who cannot identify political stances as political forces that reflect the links of the networks. The data consist of the genealogy of the Andong Gwon clan, a record of family relations of 10,243 people from the 10th to 15th century Korea, and the Annals of the Joseon Dynasty, a historical volume that describes historical facts of the Joseon Dynasty for 472 years and is composed of 1894 fascicles and 888 books. From the data, we construct a human network based on a historically meaningful period (1443–1488), and classify people into two political forces using the proposed method. We suggest that this machine learning approach to historical study could be utilized as a potent reference tool devoid of the subjectivism of human experts in the field of history.
Dong-Gi Lee, Sangkuk Lee, Myungjun Kim, Hyunjung Shin
Expert Syst. Appl.4
2017 Visual analytics for biomedical cluster subdivision: a design study with psychiatrists
abstract
In the last few years, Electronic Health Records (EHRs) have collected large-sized medical data. While EHR allows doctors to approach medication data easily, they suffer difficulties to analyze multi-dimensional medical data. We present multi-dimensional visual analytics tools to support analyzing multi-dimensional dataset by the combination of 3D RadVis and parallel coordinate. Also, we propose user-driven research design process to prospect for visualization development. This study is now in progress to interview domain experts to analyze the usability of the tools.
Hyoji Ha, Hyunwoo Han, Sungyun Bae, Sangjoon Son, Changhyung Hong, Hyunjung Shin, Kyungwon Lee
CGI7
2017 Data-driven dementia diagnosis record visualization system
abstract
In this study, we propose 'Dementia Tracker' which is a visualization system with 21,094 dementia records for 8 years. This system makes it easy to understand complex dementia record data. In addition, the patient's own record is not only well-read, but also can be compared with people who have similar degree of dementia through the group filter function. Therefore, the current dementia situation can be understood more easily, and the future dementia can be predicted and prevented in advance.
Wenjun Guo, Sangjun Son, Changhyung Hong, Hyunjung Shin, Kyungwon Lee
VINCI5
2016 Causality modeling for directed disease network
abstract
MOTIVATION: Causality between two diseases is valuable information as subsidiary information for medicine which is intended for prevention, diagnostics and treatment. Conventional cohort-centric researches are able to obtain very objective results, however, they demands costly experimental expense and long period of time. Recently, data source to clarify causality has been diversified: available information includes gene, protein, metabolic pathway and clinical information. By taking full advantage of those pieces of diverse information, we may extract causalities between diseases, alternatively to cohort-centric researches. METHOD: In this article, we propose a new approach to define causality between diseases. In order to find causality, three different networks were constructed step by step. Each step has different data sources and different analytical methods, and the prior step sifts causality information to the next step. In the first step, a network defines association between diseases by utilizing disease-gene relations. And then, potential causalities of disease pairs are defined as a network by using prevalence and comorbidity information from clinical results. Finally, disease causalities are confirmed by a network defined from metabolic pathways. RESULTS: The proposed method is applied to data which is collected from database such as MeSH, OMIM, HuDiNe, KEGG and PubMed. The experimental results indicated that disease causality that we found is 19 times higher than that of random guessing. The resulting pairs of causal-effected diseases are validated on medical literatures. AVAILABILITY AND IMPLEMENTATION: http://www.alphaminers.net CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sunjoo Bang, Jae-Hoon Kim 0004, Hyunjung Shin
Bioinform.3
2015 Knowledge boosting: a graph-based integration approach with multi-omics data and genomic knowledge for cancer clinical outcome prediction
abstract
OBJECTIVE: Cancer can involve gene dysregulation via multiple mechanisms, so no single level of genomic data fully elucidates tumor behavior due to the presence of numerous genomic variations within or between levels in a biological system. We have previously proposed a graph-based integration approach that combines multi-omics data including copy number alteration, methylation, miRNA, and gene expression data for predicting clinical outcome in cancer. However, genomic features likely interact with other genomic features in complex signaling or regulatory networks, since cancer is caused by alterations in pathways or complete processes. METHODS: Here we propose a new graph-based framework for integrating multi-omics data and genomic knowledge to improve power in predicting clinical outcomes and elucidate interplay between different levels. To highlight the validity of our proposed framework, we used an ovarian cancer dataset from The Cancer Genome Atlas for predicting stage, grade, and survival outcomes. RESULTS: Integrating multi-omics data with genomic knowledge to construct pre-defined features resulted in higher performance in clinical outcome prediction and higher stability. For the grade outcome, the model with gene expression data produced an area under the receiver operating characteristic curve (AUC) of 0.7866. However, models of the integration with pathway, Gene Ontology, chromosomal gene set, and motif gene set consistently outperformed the model with genomic data only, attaining AUCs of 0.7873, 0.8433, 0.8254, and 0.8179, respectively. CONCLUSIONS: Integrating multi-omics data and genomic knowledge to improve understanding of molecular pathogenesis and underlying biology in cancer should improve diagnostic and prognostic indicators and the effectiveness of therapies.
Do Kyoon Kim, Je-Gun Joung, Kyung-Ah Sohn 0001, Hyunjung Shin, Yu Rang Park, Marylyn D. Ritchie, Ju Han Kim
J. Am. Medical Informatics Assoc.4
2013 A Hybrid Cancer Prognosis System Based on Semi-Supervised Learning and Decision Trees
Yonghyun Nam, Hyunjung Shin
ICONIP (2)2
2013 Stock Price Prediction Based on Hierarchical Structure of Financial Networks
Kanghee Park, Hyunjung Shin
ICONIP (2)2
2013 Prediction of movement direction in crude oil prices based on semi-supervised learning
Hyunjung Shin, Tianya Hou, Kanghee Park, Chan-Kyoo Park, Sunghee Choi
Decis. Support Syst.1
2013 Robust predictive model for evaluating breast cancer survivability
Kanghee Park, Amna Ali, Do Kyoon Kim, Yeolwoo An, Minkoo Kim, Hyunjung Shin
Eng. Appl. Artif. Intell.6
2013 Stock price prediction based on a complex interrelation network of economic factors
Kanghee Park, Hyunjung Shin
Eng. Appl. Artif. Intell.2
2013 Sharpened graph ensemble for semi-supervised learning
abstract
The generalization ability of a machine learning algorithm varies on the specified values to the model parameters and the degree of noise in the learning dataset. If the dataset has an enough amount of labeled data points, the optimal value for the model parameter can be found via validation by usi ng a subset of the given dataset. However, for semi-supervised learning – one of the most recent learning algorithms, this is not as available as in conventional supervised learning. In semi-supervised learning, it is assumed that the dataset is given with only a few labeled data points. Therefore, holding out some of labeled data points for validation is not easy. The lack of labeled data points, furthermore, makes it difficult to estimate the degree of noise in the dataset. To circumvent the addressed difficulties, we propose to employ ensemble learning and graph sharpening. The former replaces the model parameter selection procedure to an ensemble network of the committee members trained with various values of model parameter. The latter, on the other hand, improves the performance of algorithms by removing unhelpful information caused by noise. The experimental results demonstrate the applicability of the proposed method for many real-world problems with no concern for the technical difficulties, by selecting the best parameter values and mitigating the influence of noise.
Inae Choi, Kanghee Park, Hyunjung Shin
Intell. Data Anal.3
2013 Research and applications: Breast cancer survivability prediction using labeled, unlabeled, and pseudo-labeled patient data
abstract
BACKGROUND: Prognostic studies of breast cancer survivability have been aided by machine learning algorithms, which can predict the survival of a particular patient based on historical patient data. However, it is not easy to collect labeled patient records. It takes at least 5 years to label a patient record as 'survived' or 'not survived'. Unguided trials of numerous types of oncology therapies are also very expensive. Confidentiality agreements with doctors and patients are also required to obtain labeled patient records. PROPOSED METHOD: These difficulties in the collection of labeled patient data have led researchers to consider semi-supervised learning (SSL), a recent machine learning algorithm, because it is also capable of utilizing unlabeled patient data, which is relatively easier to collect. Therefore, it is regarded as an algorithm that could circumvent the known difficulties. However, the fact is yet valid even on SSL that more labeled data lead to better prediction. To compensate for the lack of labeled patient data, we may consider the concept of tagging virtual labels to unlabeled patient data, that is, 'pseudo-labels,' and treating them as if they were labeled. RESULTS: Our proposed algorithm, 'SSL Co-training', implements this concept based on SSL. SSL Co-training was tested using the surveillance, epidemiology, and end results database for breast cancer and it delivered a mean accuracy of 76% and a mean area under the curve of 0.81.
Hyunjung Shin
J. Am. Medical Informatics Assoc.2
2012 The Effect of Priming of Individualism/Collectivism on the Müller-Lyer Illusion
Yiwon Hyun, Hyunjung Shin, Myeong-Ho Sohn
CogSci3
2012 Differential Effects of the Cultural Orientation Dimensions on Global Precedence
Mijung Joo, Hyunmin Kang, Hyunjung Shin, Jaesik Lee
CogSci3
2012 Cultural priming and the scene perception
Bia Kim, Yoonkyoung Lee, Goeun Lee, Hyunjung Shin
CogSci5
2012 The influence of cultural dispositions on the scene perception
Yoonkyoung Lee, Bia Kim, Yiwon Hyun, Cheon Woo Shin, Jaesik Lee, Hyunjung Shin
CogSci6
2012 A scoring model to detect abusive billing patterns in health insurance claims
Hyunjung Shin, Hayoung Park, Junwoo Lee, Won Chul Jhee
Expert Syst. Appl.1
2012 Synergistic effect of different levels of genomic data for cancer clinical outcome prediction
Do Kyoon Kim, Hyunjung Shin, Young Soo Song, Ju Han Kim
J. Biomed. Informatics2
2011 A hybrid cognitive architecture based on the relationship between consciousness, memory and attention
Eunsook Kim, Hyunjung Shin
CogSci2
2011 A hybrid cognitive architecture based on the relationship between consciousness, memory and attention
Eunsook Kim, Hyunjung Shin
CogSci2
2011 Fast Content-Based File Type Identification
Irfan Ahmed 0001, Kyung-suk Lhee, Hyunjung Shin
IFIP Int. Conf. Digital Forensics3
2011 Semantically enabled and statistically supported biological hypothesis testing with tissue microarray databases
abstract
BACKGROUND: Although many biological databases are applying semantic web technologies, meaningful biological hypothesis testing cannot be easily achieved. Database-driven high throughput genomic hypothesis testing requires both of the capabilities of obtaining semantically relevant experimental data and of performing relevant statistical testing for the retrieved data. Tissue Microarray (TMA) data are semantically rich and contains many biologically important hypotheses waiting for high throughput conclusions. METHODS: An application-specific ontology was developed for managing TMA and DNA microarray databases by semantic web technologies. Data were represented as Resource Description Framework (RDF) according to the framework of the ontology. Applications for hypothesis testing (Xperanto-RDF) for TMA data were designed and implemented by (1) formulating the syntactic and semantic structures of the hypotheses derived from TMA experiments, (2) formulating SPARQLs to reflect the semantic structures of the hypotheses, and (3) performing statistical test with the result sets returned by the SPARQLs. RESULTS: When a user designs a hypothesis in Xperanto-RDF and submits it, the hypothesis can be tested against TMA experimental data stored in Xperanto-RDF. When we evaluated four previously validated hypotheses as an illustration, all the hypotheses were supported by Xperanto-RDF. CONCLUSIONS: We demonstrated the utility of high throughput biological hypothesis testing. We believe that preliminary investigation before performing highly controlled experiment can be benefited.
Young Soo Song, Chan Hee Park, Hee-Joon Chung, Hyunjung Shin, Jihun Kim 0001, Ju Han Kim
BMC Bioinform.4
2010 Graph sharpening
Hyunjung Shin, N. Jeremy Hill, Andreas Martin Lisewski, Joon-Sang Park
Expert Syst. Appl.1
2009 On Improving the Accuracy and Performance of Content-Based File Type Identification
Irfan Ahmed 0001, Kyung-suk Lhee, Hyunjung Shin
ACISP3
2009 Protein functional class prediction with a combined graph
Hyunjung Shin, Koji Tsuda, Bernhard Schölkopf
Expert Syst. Appl.1
2008 Semi-supervised Learning with Ensemble Learning and Graph Sharpening
Inae Choi, Hyunjung Shin
IDEAL2
2007 Graph sharpening plus graph integration: a synergy that improves protein functional classification
abstract
MOTIVATION: Predicting protein function is a central problem in bioinformatics, and many approaches use partially or fully automated methods based on various combination of sequence, structure and other information on proteins or genes. Such information establishes relationships between proteins that can be modelled most naturally as edges in graphs. A priori, however, it is often unclear which edges from which graph may contribute most to accurate predictions. For that reason, one established strategy is to integrate all available sources, or graphs as in graph integration, in the hope that the positive signals will add to each other. However, in the problem of functional prediction, noise, i.e. the presence of inaccurate or false edges, can still be large enough that integration alone has little effect on prediction accuracy. In order to reduce noise levels and to improve integration efficiency, we present here a recent method in graph-based learning, graph sharpening, which provides a theoretically firm yet intuitive and practical approach for disconnecting undesirable edges from protein similarity graphs. This approach has several attractive features: it is quick, scalable in the number of proteins, robust with respect to errors and tolerant of very diverse types of protein similarity measures. RESULTS: We tested the classification accuracy in a test set of 599 proteins with remote sequence homology spread over 20 Gene Ontology (GO) functional classes. When compared to integration alone, graph sharpening plus integration of four vastly different molecular similarity measures improved the overall classification by nearly 30% [0.17 average increase in the area under the ROC curve (AUC)]. Moreover, and partially through the increased sparsity of the graphs induced by sharpening, this gain in accuracy came at negligible computational cost: sharpening and integration took on average 4.66 (+/-4.44) CPU seconds. AVAILABILITY: Software and Supplementary data will be available on http://mammoth.bcm.tmc.edu/
Hyunjung Shin, Andreas Martin Lisewski, Olivier Lichtarge
Bioinform.1
2007 Neighborhood Property-Based Pattern Selection for Support Vector Machines
abstract
The support vector machine (SVM) has been spotlighted in the machine learning community because of its theoretical soundness and practical performance. When applied to a large data set, however, it requires a large memory and a long time for training. To cope with the practical difficulty, we propose a pattern selection algorithm based on neighborhood properties. The idea is to select only the patterns that are likely to be located near the decision boundary. Those patterns are expected to be more informative than the randomly selected patterns. The experimental results provide promising evidence that it is possible to successfully employ the proposed algorithm ahead of SVM training.
Hyunjung Shin, Sungzoon Cho
Neural Comput.1
2006 Graph Based Semi-supervised Learning with Sharper Edges
Hyunjung Shin, N. Jeremy Hill, Gunnar Rätsch
ECML1
2006 Response modeling with support vector machines
Hyunjung Shin, Sungzoon Cho
Expert Syst. Appl.1
2005 Invariance of neighborhood relation under input space to feature space mapping
Hyunjung Shin, Sungzoon Cho
Pattern Recognit. Lett.1
2003 Fast Pattern Selection Algorithm for Support Vector Classifiers: Time Complexity Analysis
Hyunjung Shin, Sungzoon Cho
IDEAL1
2003 How many neighbors to consider in pattern pre-selection for support vector classifiers?
abstract
Training support vector classifiers (SVC) requires large memory and long cpu time when the pattern set is large. To alleviate the computational burden in SVC training, we previously proposed a preprocessing algorithm which selects only the patterns in the overlap region around the decision boundary, based on neighborhood properties. The k-nearest neighbors' class label entropy for each pattern was used to estimate the pattern's proximity to the decision boundary. The value of parameter k is critical, yet has been determined by a rather ad-hoc fashion. We propose in this paper a systematic procedure to determine k and show its effectiveness through experiments.
Hyunjung Shin, Sungzoon Cho
IJCNN1
2003 Fast Pattern Selection for Support Vector Classifiers
Hyunjung Shin, Sungzoon Cho
PAKDD1
2002 Pattern Selection for Support Vector Classifiers
Hyunjung Shin, Sungzoon Cho
IDEAL1
2000 Observational Learning with Modular Networks
Hyunjung Shin, Hyoungjoo Lee, Sungzoon Cho
IDEAL1