Heewon Park

dblp:159/4758 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MetaCCI: meta cell-cell interaction inference and its application to CCIs characteristics of MDS
abstract
MOTIVATION: Cell-cell interactions (CCIs) are fundamental to multicellular organisms and play crucial roles in diverse biological processes and disease mechanisms. Understanding CCIs is vital for deciphering disease pathogenesis and developing therapeutic strategies. Although numerous computational methods have been developed to infer CCIs from complex biological data, most existing approaches rely primarily on single-gene expression levels and ligand-receptor databases, often failing to capture the nuanced network-wide changes characteristic of disease states. RESULT: We propose MetaCCI, a novel computational strategy that integrates meta-information into CCI inference by extending the traditional gene expression-based analysis to a gene regulatory network framework. MetaCCI meticulously combines established ligand-receptor pairs with quantitative insights into gene behavior within complex gene networks, enabling the precise extraction of relevant targets for CCI inference. Subsequently, CCI inference was performed using an eigen cell co-expression network, providing a more holistic view of cell-cell communication. Monte Carlo simulations demonstrated that MetaCCI consistently outperforms existing methods in CCI inference. We applied MetaCCI to characterize cell-cell communication in Myelodysplastic Syndromes (MDS). Our results identified distinct interaction patterns in MDS compared with normal cell populations, specifically highlighting the loss of CCIs between "Dendritic cells and Hematopoietic precursor cells" and between "Dendritic cells and Hematopoietic multipotent progenitor cells" as characteristic features of MDS. Furthermore, FABP5, CD63, and HMGB1 were identified as MDS-specific markers. These findings suggest that diminished CCIs involving dendritic cells, hematopoietic precursor cells, and multipotent progenitor cells are pivotal to MDS pathogenesis. AVAILABILITY AND IMPLEMENTATION: The MetaCCI software is freely available at https://github.com/HeewonGitHub/MetaCCI. An archived version of the software and example datasets used in this study is available at Zenodo: https://doi.org/10.5281/zenodo.20101527.
Heewon Park, Seiya Imoto, Satoru Miyano
Bioinform.1
2025 ASAP: Unsupervised Post-training with Label Distribution Shift Adaptive Learning Rate
abstract
In real-world applications, machine learning models face online label shift, where label distributions change over time. Effective adaptation requires careful learning rate selection: too low slows adaptation and too high causes instability. We propose ASAP (Adaptive Shift Aware Post-training), which dynamically adjusts the learning rate by computing the cosine distance between current and previous unlabeled outputs and mapping it within a bounded range. ASAP requires no labels, model ensembles, or past inputs, using only the previous softmax output for fast, lightweight adaptation. Experiments across multiple datasets and shift scenarios show ASAP consistently improves accuracy and efficiency, making it practical for unsupervised model adaptation.
Heewon Park, Mugon Joe, Miru Kim, Minhae Kwon
CIKM1
2025 CluVar: clustering of variants using autoencoder for inferring cancer subclones from single cell RNA sequencing data
abstract
Tumor tissues are composed of malignant subclones with diverse genetic profiles. Reconstructing the evolutionary trajectory of these subclones is crucial for understanding how tumors acquire malignant traits. However, current approaches to subclonal tree reconstruction are limited either by their reliance on single-cell DNA sequencing (scDNA-seq) that involve a small number of cells and thus yield low-resolution results, or using single-cell RNA sequencing (scRNA-seq) data, which despite including larger cell populations, remain susceptible to bias from high dropout rates and technical noise. Here, we introduce CluVar, an autoencoder-based framework for inferring the phylogeny of cancer subclones from scRNA-seq data using mutation profile analysis. To address the extensive missing variant information inherent in scRNA-seq datasets, CluVar incorporates a customized loss function and multiple hidden layers optimized for clustering. CluVar demonstrated superior performance in reconstructing phylogenetic trees of cancer subclones under a range of erroneous conditions. When applied to cancer scRNA-seq data, the phylogenetic tree predicted using CluVar aligned well with the transcriptomic profiles. These findings highlight its utility for tracing evolutionary trajectories and identifying novel variants associated with cancer progression.
Chae Won Kim, Heewon Park, Yuchang Seong, Minhae Kwon, Junil Kim
Briefings Bioinform.2
2025 Powerful gene network enrichment analysis and its application to severe COVID-19 gene network
abstract
Understanding complex disease mechanisms requires research methods beyond individual gene analysis to capture the coordinated behavior of genes within regulatory networks. Traditional gene set enrichment approaches such as over-representation analysis and gene set enrichment analysis focus primarily on gene lists and often overlook the intricate network structures that control cellular processes. Although a gene network enrichment analysis strategy (GbNEA) has been proposed, this method assesses enrichment significance via phenotype permutation and the Kolmogorov-Smirnov test, which lowers statistical power and increases the computational burden due to repeated gene network re-estimation. To overcome these limitations, we developed a novel approach, powerful gene network enrichment analysis (PGNEA), which characterizes gene networks by integrating gene expression, regulatory effects, and hubness. PGNEA evaluates the enrichment of phenotype-specific gene networks by quantifying differences in gene activity patterns and assesses statistical significance by evaluating permutation of gene activity rather than phenotype permutation. This approach exhibits significantly enhanced computational efficiency and statistical sensitivity. We demonstrated the advantages of PGNEA through Monte Carlo simulations and applied it to whole-blood RNA-seq data obtained from the Japan COVID-19 Task Force. PGNEA successfully identified viral infection-related pathways enriched in severe COVID-19 gene networks, including those linked to "COVID-19," "HIV-1 infection," "Hepatitis B," "Influenza A," "Measles," and "Kaposi sarcoma-associated herpesvirus infection." Notably, key molecular markers such as PIK3, NF-B family members, FOXA, JUN, and CXCL8 were identified, with strong and consistent molecular interplays between CXCL8 and NFKBIA. These findings underscore the potential of PGNEA as an efficient tool for identifying biologically meaningful pathways and network-level mechanisms associated with various phenotypes, including severe viral infections.
Heewon Park, Seiya Imoto, Satoru Miyano
Briefings Bioinform.1
2025 Gene behaviors-based network enrichment analysis and its application to reveal immune disease pathways enriched with COVID-19 severity-specific gene networks
abstract
MOTIVATION: Gene network analysis is essential for understanding the complex mechanisms underlying diseases, which often involve disruptions in molecular networks rather than individual genes. Despite the availability of large-scale omics datasets and computational tools for gene network analysis, interpretation of the biological relevance of these extensive networks remains challenging. RESULTS: We propose a novel computational strategy, gene behaviors-based network enrichment analysis, which systematically identifies functional pathways enriched in phenotype-specific gene networks. Our novel method incorporates comprehensive network characteristics, i.e. gene expression levels, edge strengths, and structural patterns of edges, to rank genes based on activity and assess pathway enrichment, effectively identifying functional pathways enriched within these networks. Through simulation studies, our strategy demonstrated superior performance compared with that of existing methods in identifying enriched pathways. We applied this strategy to whole-blood RNA-seq data from 1102 COVID-19 samples provided by the Japan COVID-19 Task Force. The analysis revealed immune disease pathways enriched with COVID-19 severity-specific gene networks, including "Systemic lupus erythematosus" in asymptomatic and severe samples and "Inflammatory bowel disease," "Primary immunodeficiency," and "Rheumatoid arthritis" in mild samples. Key biomarkers of COVID-19, such as CXCL8, S100A9, and HLA class I genes, have been identified as critical hub genes and the main players within these networks. AVAILABILITY AND IMPLEMENTATION: Code is available in Figshare (https://doi.org/10.6084/m9.figshare.29093648.v3).
Heewon Park, Seiya Imoto, Satory Miyano
Bioinform.1
2025 Personalized Split Federated Learning With Early Exit: Pretraining and Online Learning Against Label Shifts
abstract
Advancements in artificial intelligence (AI) have enabled Internet of Things (IoT) devices to offer intelligent services, improving system adaptability and scalability. Split federated learning (SFL) has emerged as a promising approach for privacy-sensitive and resource-constrained IoT devices, addressing computational and privacy challenges. In the SFL framework, IoT devices serve as clients and do not share raw client data with the server. Instead, they offload computationally intensive tasks to the server. However, an SFL-based IoT system faces three key challenges. First, it struggles to personalize client models when client data distributions are heterogeneous. Second, it encounters a trade-off between communication overhead and data privacy. Third, it suffers from severe performance degradation when label distributions shift after deploying the pre-trained model. To address these issues, we propose a novel early-exit SFL framework consisting of a pre-training phase and an online learning phase. In the pre-training phase, we introduce a personalized SFL training method to tailor each client model to its data distribution and a surrogate target generation method to train the server’s large model. In the online learning phase, client models are updated to handle label distribution shifts and maintain performance by leveraging the server’s large model as a teacher through knowledge distillation. For the real-time inference, early-exit is available by passing through only the client’s model. Extensive simulations demonstrate that the proposed framework achieves superior accuracy and lower communication costs compared to state-of-the-art methods. Notably, our method outperforms existing methods with average improvements of 18.49% in pre-training settings and 30.96% under diverse label distribution shift scenarios.
Miru Kim, Heewon Park, Minhae Kwon
IEEE Internet Things J.2
2024 Visual Localization in Repetitive and Symmetric Indoor Parking Lots using 3D Key Text Graph
abstract
Indoor parking lots are the GPS-denied spaces to which vision-based localization approaches have usually been applied to solve localization problems. However, due to the repetitiveness and symmetry of the spaces, visual localization methods commonly confront difficulties in estimating precise 3D poses. In this study, we propose four novel modules that improve localization precision by imposing the existing methods with the spatial discerning ability. The first module constructs a key text graph that represents the topology of key texts in the space and becomes the basis for discerning repetitiveness and symmetry. Next, the orientation filtering module estimates the unknown 3D orientation of the query image and resolves spatial symmetric ambiguity. The similarity scoring module sorts out the top-scored database images, discerning the spatial repetitiveness based on detected key text bounding boxes. Our pose verification module evaluates the pose confidence of top-scored candidates and determines the most reliable pose. Our method has been validated in two real indoor parking lots, achieving new state-of-the-art performance levels.
Joohyung Kim, Gunhee Koo, Heewon Park, Nakju Lett Doh
ICRA3
2023 AutoMetric: Towards Measuring Open-Source Software Quality Metrics Automatically
abstract
In modern software development, open-source software (OSS) plays a crucial role. Although some methods exist to verify the safety of OSS, the current automation technologies fall short. To address this problem, we propose AutoMetric, an automatic technique for measuring security metrics for OSS in repository level. Using AutoMetric which only collects repository addresses of the projects, it is possible to inspect many projects simultaneously regardless of its size and scope. AutoMetric contains five metrics: Mean Time to Update (MU), Mean Time to Commit (MC), Number of Contributors (NC), Inactive Period (IP), and Branch Protection (BP). These metrics can be calculated quickly even if the source code changes. By comparing metrics in AutoMetric with 2,675 reported vulnerabilities in GitHub Advisory Database (GAD), the result shows that the more frequent updates and commits and the shorter the inactivity period, the more vulnerabilities were found.
Taejun Lee, Heewon Park, Heejo Lee
AST2
2022 PredictiveNetwork: predictive gene network estimation with application to gastric cancer drug response-predictive network analysis
abstract
BACKGROUND: Gene regulatory networks have garnered a large amount of attention to understand disease mechanisms caused by complex molecular network interactions. These networks have been applied to predict specific clinical characteristics, e.g., cancer, pathogenicity, and anti-cancer drug sensitivity. However, in most previous studies using network-based prediction, the gene networks were estimated first, and predicted clinical characteristics based on pre-estimated networks. Thus, the estimated networks cannot describe clinical characteristic-specific gene regulatory systems. Furthermore, existing computational methods were developed from algorithmic and mathematics viewpoints, without considering network biology. RESULTS: To effectively predict clinical characteristics and estimate gene networks that provide critical insights into understanding the biological mechanisms involved in a clinical characteristic, we propose a novel strategy for predictive gene network estimation. The proposed strategy simultaneously performs gene network estimation and prediction of the clinical characteristic. In this strategy, the gene network is estimated with minimal network estimation and prediction errors. We incorporate network biology by assuming that neighboring genes in a network have similar biological functions, while hub genes play key roles in biological processes. Thus, the proposed method provides interpretable prediction results and enables us to uncover biologically reliable marker identification. Monte Carlo simulations shows the effectiveness of our method for feature selection in gene estimation and prediction with excellent prediction accuracy. We applied the proposed strategy to construct gastric cancer drug-responsive networks. CONCLUSION: We identified gastric drug response predictive markers and drug sensitivity/resistance-specific markers, AKR1B10, AKR1C3, ANXA10, and ZNF165, based on GDSC data analysis. Our results for identifying drug sensitive and resistant specific molecular interplay are strongly supported by previous studies. We expect that the proposed strategy will be a useful tool for uncovering crucial molecular interactions involved a specific biological mechanism, such as cancer progression or acquired drug resistance.
Heewon Park, Seiya Imoto, Satoru Miyano
BMC Bioinform.1
2017 A Novel Adaptive Penalized Logistic Regression for Uncovering Biomarker Associated with Anti-Cancer Drug Sensitivity
abstract
We propose a novel adaptive penalized logistic regression modeling strategy based on Wilcoxon rank sum test (WRST) to effectively uncover driver genes in classification. In order to incorporate significance of gene in classification, we first measure significance of each gene by gene ranking method based on WRST, and then the adaptive L1-type penalty is discriminately imposed on each gene depending on the measured importance degree of gene. The incorporating significance of genes into adaptive logistic regression enables us to impose a large amount of penalty on low ranking genes, and thus noise genes are easily deleted from the model and we can effectively identify driver genes. Monte Carlo experiments and real world example are conducted to investigate effectiveness of the proposed approach. In Sanger data analysis, we introduce a strategy to identify expression modules indicating gene regulatory mechanisms via the principal component analysis (PCA), and perform logistic regression modeling based on not a single gene but gene expression modules. We can see through Monte Carlo experiments and real world example that the proposed adaptive penalized logistic regression outperforms feature selection and classification compared with existing L1-type regularization. The discriminately imposed penalty based on WRST effectively performs crucial gene selection, and thus our method can improve classification accuracy without interruption of noise genes. Furthermore, it can be seen through Sanger data analysis that the method for gene expression modules based on principal components and their loading scores provides interpretable results in biological viewpoints.
Heewon Park, Yuichi Shiraishi, Seiya Imoto, Satoru Miyano
IEEE ACM Trans. Comput. Biol. Bioinform.1