Hyun Goo Woo

dblp:71/3270 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0002-0916-893XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2025 BIGPN: Biologically informed graph propagational network for plasma proteomic profiling of neurodegenerative biomarkers
Sunghong Park, Dong-Gi Lee, Masaud Shah, Hyunjung Shin, Hyun Goo Woo
Artif. Intell. Medicine6
2025 NeuroFANN: identification of neuropathological subtypes in dementia with plasma proteins by using functionally annotated neural network
abstract
Dementia diagnosis relies on identifying neuropathological features, such as beta-amyloid (Aβ) deposition, medial temporal lobe atrophy (MTA), and white matter hyperintensity (WMH). Recently, plasma protein biomarkers have emerged as a cost-effective and less invasive tool for identifying neuropathological features, enhanced by machine learning (ML) for precise diagnosis. However, most ML studies fail to account for protein-protein interactions (PPIs) and synergetic effects between proteins, overlooking their collective contributions to disease mechanisms. Additionally, the lack of consideration for functional properties may result in the redundant and imbalanced representation of proteins and their functions, potentially limiting the effectiveness of dementia diagnosis. In this study, we propose NeuroFANN, a method designed to classify three neuropathological subtypes in dementia-positivity for Aβ, MTA, and WMH-using plasma protein biomarkers. A key feature of NeuroFANN is the combination of the PPI network-based synergetic effects with the functional annotation-based protein biomarker clustering. NeuroFANN extracts synergetic effects by propagating independent effects of proteins across the PPI network, which are then aggregated in functional protein clusters, thereby enabling global PPI awareness and capturing the biological properties of protein biomarkers. From a South Korean cohort, 54 proteins were identified as plasma protein biomarkers for dementia subtypes and grouped into 16 clusters. NeuroFANN outperformed comparison methods in classifying dementia subtypes, with its core components validated as key contributors to superior performance. Additionally, the risk scores predicted by NeuroFANN showed a strong association with longitudinal cognitive decline, demonstrating its potential as a valuable diagnostic tool in clinical settings.
Sunghong Park, Doyoon Kim, Ji-Hye Choi, Changhyung Hong, Sangjoon Son, Hyunwoong Roh, Hyunjung Shin, Hyun Goo Woo
Briefings Bioinform.8
2025 PPIxGPN: plasma proteomic profiling of neurodegenerative biomarkers with protein-protein interaction-based eXplainable graph propagational network
abstract
Neurodegenerative diseases involve progressive neuronal dysfunction, requiring the identification of specific pathological features for accurate diagnosis. While cerebrospinal fluid analysis and neuroimaging are commonly used, their invasive nature and high costs limit clinical applicability. Recently advances in plasma proteomics offer a less invasive and cost-effective alternative, further enhanced by machine learning (ML). However, most ML-based studies overlook synergetic effects from protein-protein interactions (PPIs), which play a key role in disease mechanisms. Although graph convolutional network and its extensions can utilize PPIs, they rely on locality-based feature aggregation, overlooking essential components and emphasizing noisy interactions. Moreover, expanding those methods to cover broader PPIs results in complex model architectures that reduce explainability, which is crucial in medical ML models for clinical decision-making. To address these challenges, we propose Protein-Protein Interaction-based eXplainable Graph Propagational Network (PPIxGPN), a novel ML model designed for plasma proteomic profiling of neurodegenerative biomarkers. PPIxGPN captures synergetic effects between proteins by integrating PPIs with independent effects of proteins, leveraging globality-based feature aggregation to represent comprehensive PPI properties. This process is implemented using a single graph propagational layer, enabling PPIxGPN to be configured by shallow architecture, thereby PPIxGPN ensures high model explainability, enhancing clinical applicability by providing interpretable outputs. Experimental validation on the UK Biobank dataset demonstrated the superior performance of PPIxGPN in neurodegenerative risk prediction, outperforming comparison methods. Furthermore, the explainability of PPIxGPN facilitated detailed analyses of the discriminative significance of synergistic effects, the predictive importance of proteins, and the longitudinal changes in biomarker profiles, highlighting its clinical relevance.
Sunghong Park, Dong-Gi Lee, Seung Ho Kim, Hyeon Jin Hwang, Hyunjung Shin, Hyun Goo Woo
Briefings Bioinform.7
2025 KF-NIPT: K-mer and fetal fraction-based estimation of chromosomal anomaly from NIPT data
abstract
BACKGROUND: Non-Invasive Prenatal Testing (NIPT) is a technique that allows pregnant women to screen for chromosomal abnormalities in their developing fetus without the need for invasive procedures like amniocentesis or chorionic villus sampling. However, current methods to detect anomaly from maternal cell-free DNAs (cfDNAs) that are based on the sequence read counts calculating z-scores face challenges with false positives and negatives. To address these challenges, we aimed to develop a novel NIPT algorithm named KF-NIPT, which is derived from the initials of k-mer and fetal fraction used in its development with the goal of significantly improving accuracy. RESULTS: We developed a KF-NIPT, a new algorithm that estimate chromosomal anomaly by calculating K-mer-based sequence depth and fetal fraction from the whole genome sequencing (WGS) data. Moreover, we implemented a modified preprocessing pipeline for the WGS data, correcting the biases of the genomic mapping quality and the GC contents. The performance of our method was evaluated using publicly available NIPT data. We could demonstrate that our method has better accuracy and sensitivity compared to those of the previous methods. CONCLUSIONS: We found that using k-mer and fetal fraction reduces errors in NIPT and have integrated this into a pipeline, showing that the traditional read count-based z-score method can be improved. KF-NIPT is implemented in the R and Python environment. The source code is available at https://github.com/eastbrain/KF-NIPT . KF-NIPT has been tested on Ubuntu Linux-64 server and Linux-64 on Windows using a WSL (Windows Subsystem for Linux).
Dongin Kim, Ji Yeon Sohn, Jin Hee Cho, Ji-Hye Choi, Gwi-young Oh, Hyun Goo Woo
BMC Bioinform.6
2025 Prospective domain adaptation for longitudinal data
Sunghong Park, Jeongheun Yeon, Dong-Gi Lee, Sangjoon Son, Hyun Goo Woo, Hyunjung Shin
Knowl. Based Syst.5
2025 Explainable multiplex graph propagational network with multimodal neuroimage integration for dementia subtype diagnosis
Sunghong Park, Dong-gi Lee, Do Kyoon Kim, Yonghyun Nam, Bumhee Park, Narae Kim, Seulgi Lee, Changhyung Hong, Sangjoon Son, Hyunwoong Roh, Hyun Goo Woo, Hyunjung Shin
Neural Networks12
2024 Identification of molecular subtypes of dementia by using blood-proteins interaction-aware graph propagational network
abstract
Plasma protein biomarkers have been considered promising tools for diagnosing dementia subtypes due to their low variability, cost-effectiveness, and minimal invasiveness in diagnostic procedures. Machine learning (ML) methods have been applied to enhance accuracy of the biomarker discovery. However, previous ML-based studies often overlook interactions between proteins, which are crucial in complex disorders like dementia. While protein-protein interactions (PPIs) have been used in network models, these models often fail to fully capture the diverse properties of PPIs due to their local awareness. This drawback increases the chance of neglecting critical components and magnifying the impact of noisy interactions. In this study, we propose a novel graph-based ML model for dementia subtype diagnosis, the graph propagational network (GPN). By propagating the independent effect of plasma proteins on PPI network, the GPN extracts the globally interactive effects between proteins. Experimental results showed that the interactive effect between proteins yielded to further clarify the differences between dementia subtype groups and contributed to the performance improvement where the GPN outperformed existing methods by 10.4% on average.
Sunghong Park, Changhyung Hong, Sangjoon Son, Hyunwoong Roh, Doyoon Kim, Hyunjung Shin, Hyun Goo Woo
Briefings Bioinform.7
2022 CGV: Cancer Genome Viewer, a web service for integrative cancer genome and pharmacogenomic data analysis
abstract
MOTIVATION: Multi-omic profiling data, such as The Cancer Genome Atlas and pharmacogenomic data, facilitate research into cancer mechanisms and drug development. However, it is not easy for researchers to connect, integrate and analyze huge and heterogeneous data, which is a major obstacle to the utilization of cancer genomic data. RESULTS: We developed Cancer Genome Viewer (CGV), a user-friendly web service that provides functions to integrate and visualize cancer genome data and pharmacogenomic data. Users can easily select and customize the samples to be analyzed with the pre-defined selection options for patients' clinic-pathological features from multiple datasets. Using the customized dataset, users can perform subsequent data analyses comprehensively, including gene set analysis, clustering or survival analysis. CGV also provides pre-calculated drug response scores from pharmacogenomic data, which may facilitate the discovery of new cancer targets and therapeutics. AVAILABILITY AND IMPLEMENTATION: CGV web service is implemented with the R Shiny application at http://cgv.sysmed.kr and the source code is freely available at https://git.sysmed.kr/sysmed_public/cgv. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ji-Hye Choi, Hui-Seon Choi, Seong-Ho Cho, Ji-Hye Lee, Hyun Goo Woo
Bioinform.5
2020 scTyper: a comprehensive pipeline for the cell typing analysis of single-cell RNA-seq data
abstract
BACKGROUND: Recent advances in single-cell RNA sequencing (scRNA-seq) technology have enabled the identification of individual cell types, such as epithelial cells, immune cells, and fibroblasts, in tissue samples containing complex cell populations. Cell typing is one of the key challenges in scRNA-seq data analysis that is usually achieved by estimating the expression of cell marker genes. However, there is no standard practice for cell typing, often resulting in variable and inaccurate outcomes. RESULTS: We have developed a comprehensive and user-friendly R-based scRNA-seq analysis and cell typing package, scTyper. scTyper also provides a database of cell type markers, scTyper.db, which contains 213 cell marker sets collected from literature. These marker sets include but are not limited to markers for malignant cells, cancer-associated fibroblasts, and tumor-infiltrating T cells. Additionally, scTyper provides three customized methods for estimating cell-type marker expression, including nearest template prediction (NTP), gene set enrichment analysis (GSEA), and average expression values. DNA copy number inference method (inferCNV) has been implemented with an improved modification that can be used for malignant cell typing. The package also supports the data preprocessing pipelines by Cell Ranger from 10X Genomics and the Seurat package. A summary reporting system is also implemented, which may facilitate users to perform reproducible analyses. CONCLUSIONS: scTyper provides a comprehensive and user-friendly analysis pipeline for cell typing of scRNA-seq data with a curated cell marker database, scTyper.db.
Ji-Hye Choi, Hye In Kim, Hyun Goo Woo
BMC Bioinform.3
2019 SEQprocess: a modularized and customizable pipeline framework for NGS processing in R package
abstract
BACKGROUNDS: Next-Generation Sequencing (NGS) is now widely used in biomedical research for various applications. Processing of NGS data requires multiple programs and customization of the processing pipelines according to the data platforms. However, rapid progress of the NGS applications and processing methods urgently require prompt update of the pipelines. Recent clinical applications of NGS technology such as cell-free DNA, cancer panel, or exosomal RNA sequencing data also require appropriate customization of the processing pipelines. Here, we developed SEQprocess, a highly extendable framework that can provide standard as well as customized pipelines for NGS data processing. RESULTS: SEQprocess was implemented in an R package with fully modularized steps for data processing that can be easily customized. Currently, six pre-customized pipelines are provided that can be easily executed by non-experts such as biomedical scientists, including the National Cancer Institute's (NCI) Genomic Data Commons (GDC) pipelines as well as the popularly used pipelines for variant calling (e.g., GATK) and estimation of allele frequency, RNA abundance (e.g., TopHat2/Cufflink), or DNA copy numbers (e.g., Sequenza). In addition, optimized pipelines for the clinical sequencing from cell-free DNA or miR-Seq are also provided. The processed data were transformed into R package-compatible data type 'ExpressionSet' or 'SummarizedExperiment', which could facilitate subsequent data analysis within R environment. Finally, an automated report summarizing the processing steps are also provided to ensure reproducibility of the NGS data analysis. CONCLUSION: SEQprocess provides a highly extendable and R compatible framework that can manage customized and reproducible pipelines for handling multiple legacy NGS processing tools.
Taewoon Joo, Ji-Hye Choi, Ji-Hye Lee, So Eun Park, Youngsic Jeon, Sae Hoon Jung, Hyun Goo Woo
BMC Bioinform.7
2017 Pan-cancer analysis of systematic batch effects on somatic sequence variations
abstract
BACKGROUND: The Cancer Genome Atlas (TCGA) is a comprehensive database that includes multi-layered cancer genome profiles. Large-scale collection of data inevitably generates batch effects introduced by differences in processing at various stages from sample collection to data generation. However, batch effects on the sequence variation and its characteristics have not been studied extensively. RESULTS: We systematically evaluated batch effects on somatic sequence variations in pan-cancer TCGA data, revealing 999 somatic variants that were batch-biased with statistical significance (P < 0.00001, Fisher's exact test, false discovery rate ≤ 0.0027). Most of the batch-biased variants were associated with specific sample plates. The batch-biased variants, which had a unique mutational spectrum with frequent indel-type mutations, preferentially occurred at sites prone to sequencing errors, e.g., in long homopolymer runs. Non-indel type batch-biased variants were frequent at splicing sites with the unique consensus motif sequence 'TTDTTTAGTT'. Furthermore, some batch-biased variants occur in known cancer genes, potentially causing misinterpretation of mutation profiles. CONCLUSIONS: Our strategy for identifying batch-biased variants and characterising sequence patterns might be useful in eliminating false variants and facilitating correct interpretation of sequence profiles.
Ji-Hye Choi, Seong-Eui Hong, Hyun Goo Woo
BMC Bioinform.3
2007 GAzer: gene set analyzer
abstract
UNLABELLED: Gene Set Analyzer (GAzer) is a web-based integrated gene set analysis tool covering previously reported parametric and non-parametric models. Based on a simulation test for the reported algorithms, we classified and implemented three main statistical methods consisting of the z-statistic, gene permutation and sample permutation for ten gene set categories including Gene Ontology (GO) for human, mouse, rat and yeast. This tool identifies significantly altered gene sets scored by z-statistics and P-values from the z-test or permutation test and provides q-values and Bonferroni P-values to correct multiple hypothesis testing. GAzer allows users to observe changes in expression of each gene in a gene set or to see the significance of the gene sets containing a gene(s) of interest, thus allowing interactive data analysis both at the gene and gene set level. Moreover, GAzer offers extensive annotation for each gene. AVAILABILITY: The GAzer gene set analyzer is freely available at http://integromics.kobic.re.kr/GAzer/. SUPPLEMENTARY INFORMATION: This can be found on the web page (http://integromics.kobic.re.kr/GAzer/supplement.jsp).
Sang-Bae Kim, Sungjin Yang, Seon-Kyu Kim, Sang Cheol Kim, Hyun Goo Woo, David J. Volsky, Seon-Young Kim, In-Sun Chu
Bioinform.5