Seung-Hyun Hong

dblp:69/3325 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
6since 2021 · last 2024
0000-0003-2848-6666ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 6 since 2021
YearPublicationVenuePosition
2024 FuncNet: A Machine Learning Framework Capable of Enhancing Neural Cell Type Classification via Functional Subtype Clustering
abstract
Cell type identification in single cell RNA sequencing (scRNA-seq) remains an essential yet challenging task. Existing algorithms assign cell type nomenclature typically by (i) subgrouping cells by computing proximity of cellular gene expression patterns at the single cell level, (ii) identifying marker genes in each cluster, and (iii) performing a lookup of these genes in the literature and databases to assign cell subtype labels. Unfortunately, this approach falls short when classifying cells based on functions, because biological function(s) of a group of cells cannot be determined based on a limited set of marker genes. Here we propose a new modeling formalism, namely FuncNet, which is a form of neural network-based method capable of discerning the functional profiles of individual cells. The key idea of FuncNet is to incorporate "a secondary learning arm" into the otherwise classical multi-layer neural network in order to bring in "prior knowledge" into the overall learning process. The prior knowledge that is brought into the FuncNet learning is Gene Ontology Biological Processes that should have been captured in the cells being studied. We applied FuncNet to three publicly available single cell transcriptomics datasets from human and mouse brains. The results demonstrate that FuncNet can successfully reproduce the functional nomenclatures associated with brain cells in the literature. FuncNet suggests a new way of subtyping cells, specifically for the functional subtyping of single cells—an area which has not been well established in the current single cell data analysis practices.
Chenyu Zhang 0007, Honglin Wang, Seung-Hyun Hong, Riqiang Yan, Dong-Guk Shin
BIBM3
2024 vSPACE: exploring virtual spatial representation of articular chondrocytes at the single-cell level
abstract
SUMMARY: vSPACE is a web-based application presenting a spatial representation of scRNAseq data obtained from human articular cartilage by emulating the concept of spatial transcriptomics technology, but virtually. This virtual 2D plot presentation of human articular cartage cells generates several zonal distribution patterns, for one or multiple genes at a time, revealing patterns that scientists can appreciate as imputed spatial distribution patterns along the zonal axis. AVAILABILITY AND IMPLEMENTATION: vSPACE is implemented in Python Dash as a web-based toolbox designed for data visualization of zonal gene expression patterns in articular cartilage chondrocytes. This tool is freely accessible at: https://vspace.cse.uconn.edu/The source code and extra materials for this service can be downloaded from: https://github.com/zhacheny/vSPACE.
Chenyu Zhang 0007, Honglin Wang, Yeonsoo Chung, Seung-Hyun Hong, Merissa Olmer, Hannah Swahn, Martin Lotz, Peter F. Maye, David W. Rowe, Dong-Guk Shin
Bioinform.4
2023 Using Biological Processes as Prior Knowledge Identifies New Microglial Immune Signatures at Single Cell Level in Alzheimer's Disease
abstract
The brain’s resident immune cells, microglia, play important roles in the pathological process of Alzheimer's disease (AD). Upon encountering amyloid-β (Aβ) plaque accumulation, microglia change its state to produce proinflammatory cytokines to uptake and clear Aβ but during the chronic state of neuroinflammation they become responsible for neurodegeneration. In this paper we propose a new method of analyzing microglia undergoing multi-stage state changes from homeostasis (HOM) to disease associated microglia (DAM). Our single cell gene expression data analysis method differs from the conventional marker based subtyping methods in the sense that it uses prior knowledge, namely, Gene Ontology’s Biological Processes known for immune functions and other related ones. In this regard, our method can be thought of supervised clustering rather than UMAP/tSNE style unbiased subgrouping of cells. The strengths of this "prior knowledge" using method are multi-faceted: (i) one can assign meaningful functional nomenclatures to identified subtypes of cells, (ii) one can refine GO Biological Process genes to be context-specific (e.g., phagocytosis responsible by "microglia" as opposed to "general immune cells"), and (iii) additional context-specific biological processes can be identified in addition to the ones used as "prior knowledge". We illustrate the advantages of our semi-supervised cell clustering method using a set of publicly available human AD and mouse AD model gene expression datasets.
Chenyu Zhang 0007, Honglin Wang, Seung-Hyun Hong, Riqiang Yan, Dong-Guk Shin
BIBM3
2022 Pola Viz Reveals Microglia Polarization at Single Cell Level in Alzheimer's Disease
abstract
Microglia are macrophages residing in the brain and spinal cord and responsible for neurological immune defense. A growing number of studies point out the importance of microglia’s role in Alzheimer’s disease (AD) pathology. Like macrophages in other tissues, microglia are polarized to M1/M2 states, Ml inhibiting cell inflammatory and causing tissue damage and M2 resolving inflammation and repairing damaged tissue. Thanks to scRNA-seq technology, single cell gene expression data sets for microglia have become available. However, examining the polarization of microglia at single cell level to study their association with AD is difficult for reasons such as complexity (too many genes, too many cells), expression dropouts, etc. We propose a scRNA-seq data visualization framework, namely, Pola Viz that can associate the polarization states of microglia with gene regulatory pathways that are potentially responsible for defining M1/M2 states or transition between them. For its two-dimensional placement of single cells, the x-axis depicts each cell’s M1/M2 or its transition score and the y-axis denotes what we call, the pathway “route” scores designed to assign the degree of contribution of the concerned pathway routes to each cell’s M1/M2 state. We conducted two case studies involving publicly available AD single cell mRNA datasets, one from human brain and one from mouse brain. Our method identified pathway routes that are highly correlated with known functional roles of microglia. It also discovers similarities between mice and human AD microglia activation such as T cell receptor signaling pathway, cell cycle, NF-kappa B signaling pathways, all highly disturbed at late stage of AD. Our method provides a new way of examining the M1/M2 polarization of microglia at single cell level
Chenyu Zhang 0007, Pujan Joshi, Honglin Wang, Seung-Hyun Hong, Riqiang Yan, Dong-Guk Shin
BIBM4
2021 Identification of Crosstalk between Biological Pathway Routes in Cancer Cohorts
abstract
Signal transduction pathways can affect each other through crosstalk between one or more shared components. In this paper, we present a framework that uses a route-based approach to identify crosstalk between signaling pathway routes in different cancer cohorts. The framework identifies all possible routes originating from a ligand/receptor (LR) of one pathway to a transcription factor (TF) of another pathway. Then, the downstream crosstalk routes, which originate from the TF of one pathway to ligand of another pathway, are identified. For each crosstalk route, activity scores and p-values are computed. Overall route activity in each cancer cohort was assessed in terms of two summary metrics, “Proportion of Significance” (PS) and “Average Route Score” (ARS). Case studies of five human cancer cohorts from The Cancer Genome Atlas (TCGA) repository are presented to demonstrate the discovery value of this approach.
Pujan Joshi, Honglin Wang, Salvatore Jaramillo, Seung-Hyun Hong, Charles Giardina, Dong-Guk Shin
BIBM4
2021 ctBuilder: A framework for building pathway crosstalks by combining single cell data with bulk cell data
abstract
Biomedical research scientists routinely use publicly available gene regulatory pathway databases to extract interpretation from high-throughput multi-omics studies. However, deficiencies of these existing pathway resources include: (i) the databases store “individualized” pathways as a collection of networks thus making discovery of crosstalk across the boundaries of curated pathways difficult, and (ii) curated pathways themselves are incomplete, particularly from the disease perspective. We propose a computational pathway extension framework, called ctBuilder, which aims to tackle the deficiencies. It uses single cell gene expression data to identify potential candidates to tune the pathway extension “context-specific” and uses bulk cell gene expression data sets to further refine and improve likelihoods of the regulatory relationships estimated to interlink constituents involved between pathways. For demonstration, we applied ctBuilder to 5 non-alcoholic steatohepatis related studies from GEO. We show how a pair of known pathways, TNF signaling vs. NF-kB signaling, can be extended to include crosstalk subnetworks specifically related to the liver disease. Discovered relationships among them merit wetlab experiments for validation.
Honglin Wang, Pujan Joshi, Seung-Hyun Hong, Dong-Ju Shin, Dong-Guk Shin
BIBM3
2020 Identification of Key Biological Pathway Routes in Cancer Cohorts
abstract
Over the last two decades, various pathway analysis methods have been proposed to investigate complex biological interactions with omics data. Topology based (TB) pathway analysis techniques are generally considered to have better performance than non-topology-based methods. However, these methods score an entire pathway as a unit, where the relevance of individual routes within the pathway is lost. In this paper, a novel route-based pathway analysis framework is discussed that can effectively process entire cohorts of gene expression data sets and identify significant pathway routes in the given cohort. The framework begins with identifying all possible transcription factor (TF) centric routes from KEGG signaling pathways. For each route in a pathway, activity scores and p-values are calculated for samples in the given cohort. Overall route activity in a cohort is assessed in terms of two summary metrics, “Proportion of Significance” (PS) and “Average Route Score” (ARS). Case studies of two human cancer cohorts from The Cancer Genome Atlas (TCGA) repository are presented.
Pujan Joshi, Brent Basso, Honglin Wang, Seung-Hyun Hong, Charles Giardina, Dong-Guk Shin
BIBM4
2020 cTAP: A Machine Learning Framework for Predicting Target Genes of a Transcription Factor using a Cohort of Gene Expression Data Sets
abstract
Identifying target genes of a transcription factor is crucial in biomedical research. Thanks to ChIP-seq technology, scientists can estimate potential genome-wide target genes of a transcription factor. However, finding the consistently behaving Up/Down targets of a transcription factor in a given biological context is difficult because it requires analysis of a large number of studies under the same or comparable context. We present a transcription target prediction method, called Cohort-based TF target prediction system (cTAP). This method assumes that the pathway involving the transcription factor of interest is featured with multiple functional groups of marker genes pertaining to the concerned biological process. It uses the notion of gene-presence and gene-absence in addition to log2 ratios of gene expression values for the prediction. Target prediction is made by applying multiple machine-learning models that learn the patterns of genepresence and gene-absence from log2 ratio and four types of Z scores from the normalized cohort's gene expression data. The learned patterns are then associated with the putative targets of the concerned transcription factor to elicit genes exhibiting Up/Down gene regulation patterns “consistently” within the cohort. Totally 11 publicly available GEO data sets related to osteoclastogenesis are used in our experiment. The learned models using gene-presence and gene-absence produce target genes different from using only log2 ratios such as CASP1, BID, and IRF5. Our literature survey reveals that all these predicted targets have known roles in bone remodeling, specifically related to immune and osteoclasts, suggesting confidence in our method and potential merit for a wet-lab experiment for validation.
Honglin Wang, Pujan Joshi, Seung-Hyun Hong, Peter F. Maye, David W. Rowe, Dong-Guk Shin
BIBM3
2016 Deep pathway analysis incorporating mutation information and gene expression data
abstract
We propose a new way of analyzing biological pathways in which the analysis combines both transcriptome data and mutation information and uses the outcome to identify routes of aberrant pathways potentially responsible for the etiology of disease. Each pathway route is encoded as a Bayesian Network which is initialized with a sequence of conditional probabilities which are designed to encode directionality of regulatory relationships encoded in the pathways, i.e. activation and inhibition relationships. First, we demonstrate the effectiveness of our model through simulation in which the model is able to discern patients in Test Group from ones in Control Group. Second, we apply our model to analyze the Breast Cancer data set, available from TCGA, against some pathways available from KEGG. Our experiment with this published data reaffirms the claims reported from the original breast cancer PAM50 subtype study. Our model can further analyze the patients of each subtype based on the identified route of aberration. For example, our analysis shows that complex biological process patterns are presented for HER2+ patients potentially suggesting our method's use for producing refined subtyping. We manage to find commonly perturbed pathway routes for HER2+ patients. We claim such “deep” pathway analysis could be very useful in designing a personalized therapy.
Tham H. Hoang, Pujan Joshi, Seung-Hyun Hong, Dong-Guk Shin
BIBM4
2013 A software framework integrating gene expression patterns, binding site analysis and gene ontology to hypothesize gene regulation relationships
abstract
One known challenge in analyzing gene expression data is to combine analysis outcomes obtained disparately by applying multiple, independent meta-analysis methods. Here we present an integrative computational system that narrows down biological hypotheses by integrating gene expression patterns, transcription factor (TF) binding site analysis outcomes, and Gene Ontology (GO) enrichment analysis outcomes. This system identifies regulated genes from microarray experiments through statistical processes, categorizes similarly behaving groups of genes and then carries out binding site analysis and gene function enrichment analysis based on some significant clusters. The output is an ordered set of "putative" pair-wise relationships between TFs and their potential target genes. The relationships are ranked based on their closeness to the experimental context. We demonstrate the effectiveness of our framework using two independent microarray data sets.
Pujan Joshi, Baikang Pei, Seung-Hyun Hong, Ivo Kalajzic, Dong-Ju Shin, David W. Rowe, Dong-Guk Shin
BIBM3