EDBT 2026 Demo / reviewers in the wild / expert
Dong-Guk Shin
dblp:12/405
· DBLP profile ↗
40ranked-venue papers
9as first author
9since 2021 · last 2024
0009-0004-4649-7363ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 30 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Resolving multiple structural variation callers and platforms by matrix transformation and imputationabstractAdvances in genome sequencing technologies have increased the ability to detect structural variations but this process remains imperfect due to inference limitations for paired-end and long read platforms. In addition to the higher sequencing cost of longer reads, it has been shown that long read platforms have lower inference performance for certain structural variants than the lower cost paired reads. Therefore, it is important to establish an automated information-driven methodology for combining the best parts of any given platform and structural variation calling algorithm to produce a high-quality SV estimate of a new unstudied individual. We detail a novel method for clustering similar variants from an arbitrary number of callers and platforms by transforming them into a standardized matrix with imputed missing values. Especially useful in the new formulation is the ability to represent new unseen variants in a shared space (e.g., a cloud or local DB) with the previously studied examples. This allows our method to be extended for online or semi-supervised learning where gold standard data sets derived from the 1000 Genomes Phase2, Phase3, and Human Genome SV project Phase 1 and Phase 2 can be used to eliminate background noise. We compare our novel ensemble method with leading individual callers and other ensemble methods and show an increase in performance. We showcase the importance of offering the analysis transparency in identifying disease-specific SVs (i.e., orofacial cleft lip and palate) using selected samples from the Gabrielle Miller Kids First Asian Orofacial Cleft cohort. Our work is written in python3 and is fully open source along with several preprocessed supervised learning datasets. Timothy James Becker, Michelle Le, Deem Quinn, Yeonsoo Chung, Dong-Guk Shin |
BIBM | 5 |
| 2024 | FuncNet: A Machine Learning Framework Capable of Enhancing Neural Cell Type Classification via Functional Subtype ClusteringabstractCell type identification in single cell RNA sequencing (scRNA-seq) remains an essential yet challenging task. Existing algorithms assign cell type nomenclature typically by (i) subgrouping cells by computing proximity of cellular gene expression patterns at the single cell level, (ii) identifying marker genes in each cluster, and (iii) performing a lookup of these genes in the literature and databases to assign cell subtype labels. Unfortunately, this approach falls short when classifying cells based on functions, because biological function(s) of a group of cells cannot be determined based on a limited set of marker genes. Here we propose a new modeling formalism, namely FuncNet, which is a form of neural network-based method capable of discerning the functional profiles of individual cells. The key idea of FuncNet is to incorporate "a secondary learning arm" into the otherwise classical multi-layer neural network in order to bring in "prior knowledge" into the overall learning process. The prior knowledge that is brought into the FuncNet learning is Gene Ontology Biological Processes that should have been captured in the cells being studied. We applied FuncNet to three publicly available single cell transcriptomics datasets from human and mouse brains. The results demonstrate that FuncNet can successfully reproduce the functional nomenclatures associated with brain cells in the literature. FuncNet suggests a new way of subtyping cells, specifically for the functional subtyping of single cells—an area which has not been well established in the current single cell data analysis practices. Chenyu Zhang 0007, Honglin Wang, Seung-Hyun Hong, Riqiang Yan, Dong-Guk Shin |
BIBM | 5 |
| 2024 | vSPACE: exploring virtual spatial representation of articular chondrocytes at the single-cell levelabstractSUMMARY: vSPACE is a web-based application presenting a spatial representation of scRNAseq data obtained from human articular cartilage by emulating the concept of spatial transcriptomics technology, but virtually. This virtual 2D plot presentation of human articular cartage cells generates several zonal distribution patterns, for one or multiple genes at a time, revealing patterns that scientists can appreciate as imputed spatial distribution patterns along the zonal axis. AVAILABILITY AND IMPLEMENTATION: vSPACE is implemented in Python Dash as a web-based toolbox designed for data visualization of zonal gene expression patterns in articular cartilage chondrocytes. This tool is freely accessible at: https://vspace.cse.uconn.edu/The source code and extra materials for this service can be downloaded from: https://github.com/zhacheny/vSPACE. Chenyu Zhang 0007, Honglin Wang, Yeonsoo Chung, Seung-Hyun Hong, Merissa Olmer, Hannah Swahn, Martin Lotz, Peter F. Maye, David W. Rowe, Dong-Guk Shin |
Bioinform. | 10 |
| 2023 | Using Biological Processes as Prior Knowledge Identifies New Microglial Immune Signatures at Single Cell Level in Alzheimer's DiseaseabstractThe brain’s resident immune cells, microglia, play important roles in the pathological process of Alzheimer's disease (AD). Upon encountering amyloid-β (Aβ) plaque accumulation, microglia change its state to produce proinflammatory cytokines to uptake and clear Aβ but during the chronic state of neuroinflammation they become responsible for neurodegeneration. In this paper we propose a new method of analyzing microglia undergoing multi-stage state changes from homeostasis (HOM) to disease associated microglia (DAM). Our single cell gene expression data analysis method differs from the conventional marker based subtyping methods in the sense that it uses prior knowledge, namely, Gene Ontology’s Biological Processes known for immune functions and other related ones. In this regard, our method can be thought of supervised clustering rather than UMAP/tSNE style unbiased subgrouping of cells. The strengths of this "prior knowledge" using method are multi-faceted: (i) one can assign meaningful functional nomenclatures to identified subtypes of cells, (ii) one can refine GO Biological Process genes to be context-specific (e.g., phagocytosis responsible by "microglia" as opposed to "general immune cells"), and (iii) additional context-specific biological processes can be identified in addition to the ones used as "prior knowledge". We illustrate the advantages of our semi-supervised cell clustering method using a set of publicly available human AD and mouse AD model gene expression datasets. Chenyu Zhang 0007, Honglin Wang, Seung-Hyun Hong, Riqiang Yan, Dong-Guk Shin |
BIBM | 5 |
| 2022 | A framework for associating structural variants with cell-specific transcription factors and histone modifications in defect phenotypesabstractStructural Variations (SVs) naturally occur in healthy human populations and account for more genomic difference than Single Nucleotide Variations (SNVs). However, because of the low prevalence in gene coding regions and the ubiquity in the intergenic regions, determining SV association for defect phenotypes remains challenging. We developed a framework that classifies five categories of SVs using defect associated genes along with cell-specific histone modifications (HM) concurrently with transcription factor binding sites (TFBS). We completed a comprehensive structural variation analysis of 17 family trios consisting of healthy paternal and maternal genomes along with the non-syndromic Orofacial cleft (OFC) child genomes which are publicly available from the Kids First Data Resource Portal. For the integrative analysis, we used ChIP-seq data from Neural Crest and Mesenchymal Stem Cell types from ENCODE along with TFBS taken from JASPAR. We found that the OFC children had elevated regulatory SVs when classified with our framework, aligning with prior studies regarding the complex non-syndromic phenotype. This result supports the use of a defect-specific gene paralog lists integrated with HMs and TFBS to identify plausible SV regions connected to a defect phenotype. Timothy James Becker, Dashzeveg Bayarsaihan, Dong-Guk Shin |
BIBM | 3 |
| 2022 | Pola Viz Reveals Microglia Polarization at Single Cell Level in Alzheimer's DiseaseabstractMicroglia are macrophages residing in the brain and spinal cord and responsible for neurological immune defense. A growing number of studies point out the importance of microglia’s role in Alzheimer’s disease (AD) pathology. Like macrophages in other tissues, microglia are polarized to M1/M2 states, Ml inhibiting cell inflammatory and causing tissue damage and M2 resolving inflammation and repairing damaged tissue. Thanks to scRNA-seq technology, single cell gene expression data sets for microglia have become available. However, examining the polarization of microglia at single cell level to study their association with AD is difficult for reasons such as complexity (too many genes, too many cells), expression dropouts, etc. We propose a scRNA-seq data visualization framework, namely, Pola Viz that can associate the polarization states of microglia with gene regulatory pathways that are potentially responsible for defining M1/M2 states or transition between them. For its two-dimensional placement of single cells, the x-axis depicts each cell’s M1/M2 or its transition score and the y-axis denotes what we call, the pathway “route” scores designed to assign the degree of contribution of the concerned pathway routes to each cell’s M1/M2 state. We conducted two case studies involving publicly available AD single cell mRNA datasets, one from human brain and one from mouse brain. Our method identified pathway routes that are highly correlated with known functional roles of microglia. It also discovers similarities between mice and human AD microglia activation such as T cell receptor signaling pathway, cell cycle, NF-kappa B signaling pathways, all highly disturbed at late stage of AD. Our method provides a new way of examining the M1/M2 polarization of microglia at single cell level Chenyu Zhang 0007, Pujan Joshi, Honglin Wang, Seung-Hyun Hong, Riqiang Yan, Dong-Guk Shin |
BIBM | 6 |
| 2021 | Identification of Crosstalk between Biological Pathway Routes in Cancer CohortsabstractSignal transduction pathways can affect each other through crosstalk between one or more shared components. In this paper, we present a framework that uses a route-based approach to identify crosstalk between signaling pathway routes in different cancer cohorts. The framework identifies all possible routes originating from a ligand/receptor (LR) of one pathway to a transcription factor (TF) of another pathway. Then, the downstream crosstalk routes, which originate from the TF of one pathway to ligand of another pathway, are identified. For each crosstalk route, activity scores and p-values are computed. Overall route activity in each cancer cohort was assessed in terms of two summary metrics, “Proportion of Significance” (PS) and “Average Route Score” (ARS). Case studies of five human cancer cohorts from The Cancer Genome Atlas (TCGA) repository are presented to demonstrate the discovery value of this approach. Pujan Joshi, Honglin Wang, Salvatore Jaramillo, Seung-Hyun Hong, Charles Giardina, Dong-Guk Shin |
BIBM | 6 |
| 2021 | ctBuilder: A framework for building pathway crosstalks by combining single cell data with bulk cell dataabstractBiomedical research scientists routinely use publicly available gene regulatory pathway databases to extract interpretation from high-throughput multi-omics studies. However, deficiencies of these existing pathway resources include: (i) the databases store “individualized” pathways as a collection of networks thus making discovery of crosstalk across the boundaries of curated pathways difficult, and (ii) curated pathways themselves are incomplete, particularly from the disease perspective. We propose a computational pathway extension framework, called ctBuilder, which aims to tackle the deficiencies. It uses single cell gene expression data to identify potential candidates to tune the pathway extension “context-specific” and uses bulk cell gene expression data sets to further refine and improve likelihoods of the regulatory relationships estimated to interlink constituents involved between pathways. For demonstration, we applied ctBuilder to 5 non-alcoholic steatohepatis related studies from GEO. We show how a pair of known pathways, TNF signaling vs. NF-kB signaling, can be extended to include crosstalk subnetworks specifically related to the liver disease. Discovered relationships among them merit wetlab experiments for validation. Honglin Wang, Pujan Joshi, Seung-Hyun Hong, Dong-Ju Shin, Dong-Guk Shin |
BIBM | 5 |
| 2021 | Deep Pathway Analysis V2.0: A Pathway Analysis Framework Incorporating Multi-Dimensional Omics DataabstractPathway analysis is essential in cancer research particularly when scientists attempt to derive interpretation from genome-wide high-throughput experimental data. If pathway information is organized into a network topology, its use in interpreting omics data can become very powerful. In this paper, we propose a topology-based pathway analysis method, called DPA V2.0, which can combine multiple heterogeneous omics data types in its analysis. In this method, each pathway route is encoded as a Bayesian network which is initialized with a sequence of conditional probabilities specifically designed to encode directionality of regulatory relationships defined in the pathway. Unlike other topology-based pathway tools, DPA is capable of identifying pathway routes as representatives of perturbed regulatory signals. We demonstrate the effectiveness of our model by applying it to two well-established TCGA data sets, namely, breast cancer study (BRCA) and ovarian cancer study (OV). The analysis combines mRNA-seq, mutation, copy number variation, and phosphorylation data publicly available for both TCGA data sets. We performed survival analysis and patient subtype analysis and the analysis outcomes revealed the anticipated strengths of our model. We hope that the availability of our model encourages wet lab scientists to generate extra data sets to reap the benefits of using multiple data types in pathway analysis. The majority of pathways distinguished can be confirmed by biological literature. Moreover, the proportion of correctly indentified pathways is 10 percent higher than previous work where only mRNA-seq and mutation data is incorporated for breast cancer patients. Consequently, such an in-depth pathway analysis incorporating more diverse data can give rise to the accuracy of perturbed pathway detection. Dong-Guk Shin |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2020 | TensorSV: structural variation inference using tensors and variable topology neural networksabstractGenomic variation inference is an important sub problem in whole genome assembly. For the well-established human genome, additional information is available from the reference that allows data representations to depart from assembly graphs. One recent method using a convolutional neural network (CNN) and image-based representation, outperformed other algorithms in resolving single nucleotide variations (SNVs). This same representation was used to call deletion structural variations (SVs) but was unable to be applied to other SV types like inversions. We present a variable topology method along with a novel data representation suitable for a wide range of genomic variation inference. We demonstrate the effectiveness of this representation for SVs by training CNN ensembles with tensors derived from the 1000 genomes phase 3 high coverage dataset for detecting deletion and inversion types and comparing the results to well established methods. Our pure ML method facilitates a new direction in SV inference technique, where feature selection and region filtering are no longer needed to maintain low false positive rates. Timothy James Becker, Dong-Guk Shin |
BIBM | 2 |
| 2020 | Identification of Key Biological Pathway Routes in Cancer CohortsabstractOver the last two decades, various pathway analysis methods have been proposed to investigate complex biological interactions with omics data. Topology based (TB) pathway analysis techniques are generally considered to have better performance than non-topology-based methods. However, these methods score an entire pathway as a unit, where the relevance of individual routes within the pathway is lost. In this paper, a novel route-based pathway analysis framework is discussed that can effectively process entire cohorts of gene expression data sets and identify significant pathway routes in the given cohort. The framework begins with identifying all possible transcription factor (TF) centric routes from KEGG signaling pathways. For each route in a pathway, activity scores and p-values are calculated for samples in the given cohort. Overall route activity in a cohort is assessed in terms of two summary metrics, “Proportion of Significance” (PS) and “Average Route Score” (ARS). Case studies of two human cancer cohorts from The Cancer Genome Atlas (TCGA) repository are presented. Pujan Joshi, Brent Basso, Honglin Wang, Seung-Hyun Hong, Charles Giardina, Dong-Guk Shin |
BIBM | 6 |
| 2020 | cTAP: A Machine Learning Framework for Predicting Target Genes of a Transcription Factor using a Cohort of Gene Expression Data SetsabstractIdentifying target genes of a transcription factor is crucial in biomedical research. Thanks to ChIP-seq technology, scientists can estimate potential genome-wide target genes of a transcription factor. However, finding the consistently behaving Up/Down targets of a transcription factor in a given biological context is difficult because it requires analysis of a large number of studies under the same or comparable context. We present a transcription target prediction method, called Cohort-based TF target prediction system (cTAP). This method assumes that the pathway involving the transcription factor of interest is featured with multiple functional groups of marker genes pertaining to the concerned biological process. It uses the notion of gene-presence and gene-absence in addition to log2 ratios of gene expression values for the prediction. Target prediction is made by applying multiple machine-learning models that learn the patterns of genepresence and gene-absence from log2 ratio and four types of Z scores from the normalized cohort's gene expression data. The learned patterns are then associated with the putative targets of the concerned transcription factor to elicit genes exhibiting Up/Down gene regulation patterns “consistently” within the cohort. Totally 11 publicly available GEO data sets related to osteoclastogenesis are used in our experiment. The learned models using gene-presence and gene-absence produce target genes different from using only log2 ratios such as CASP1, BID, and IRF5. Our literature survey reveals that all these predicted targets have known roles in bone remodeling, specifically related to immune and osteoclasts, suggesting confidence in our method and potential merit for a wet-lab experiment for validation. Honglin Wang, Pujan Joshi, Seung-Hyun Hong, Peter F. Maye, David W. Rowe, Dong-Guk Shin |
BIBM | 6 |
| 2019 | HFM: Hierarchical Feature Moment Extraction for Multi-Omic Data VisualizationabstractSequencing the DNA of the estimated 7.5 billion living humans would generate 1.4 zettabytes of data. However, given current per-read rendering techniques, just one DNA alignment file which is around 200 gigabytes can be resource intensive to visualize at arbitrary scale. Going from human DNA and RNA sequencing data to biological insight is a process that requires domain knowledge in addition to computational methods that are bound by time and space. We address these limitations by integrating a parallel out-of-core feature extraction algorithm with a disk-based hierarchical data store that provides several orders of magnitude speed-up for common analysis and visualization tasks. To demonstrate the effectiveness of our strategy, we have developed a web-based REST service that serves translated data to a real-time genomic viewer, which in turn renders standardized moments as stacked-area graphs of features in milliseconds for multiple samples using a familiar genome browser interface. Unlike per-read techniques which read a variable number of rows from the sequence alignment file depending on the region of interest, our data structure returns a controllable data size of that region, making the technique ideally suited for visualization and macro-level insight of large cohorts. The strategy works well for high-coverage single coordinate-based visualization but could be extended for use in other long-range visualization techniques. We detail our open-source Cython/Python based implementation as well as our prototype web-based visualization tool and then compare the resulting performance and against established visualization tools. Timothy James Becker, Dong-Guk Shin |
BIBM | 2 |
| 2019 | Recurrent Neural Network for Gene Regulation Network Construction on Time Series Expression DataabstractWe propose a new way of exploring potential transcription factor targets in which the Recurrent Neural Network (RNN) is used to model time series gene expression data. Once the training of the RNN is completed, inference is performed through feeding the RNN artificially constructed signals. These artificial signals emulate the original gene expression data and the transcriptional factor of interest is set to be zero constantly to model the knockout state of the transcription factor. The predicted expression patterns of the other genes from the RNN are then used to measure the likelihood that the gene is regulated by the knocked out transcriptional factor. After repeating the same process for each gene as Transcription Factor in the dataset, we construct a gene regulation network with edge weights assigned. We demonstrate the effectiveness of our model by comparing our method with existing popular approaches. The result shows that our RNN method can identify transcription factor targets with higher accuracies than most of existing approaches. Overall, our RNN model trained on time series gene expression data can be useful for discovering transcription factor targets as well as building a gene regulation network. Pujan Joshi, Dong-Guk Shin |
BIBM | 3 |
| 2019 | TriPOINT: a software tool to prioritize important genes in pathways and their non-coding regulatorsabstractSUMMARY: Current approaches for pathway analyses focus on representing gene expression levels on graph representations of pathways and conducting pathway enrichment among differentially expressed genes. However, gene expression levels by themselves do not reflect the overall picture as non-coding factors play an important role to regulate gene expression. To incorporate these non-coding factors into pathway analyses and to systematically prioritize genes in a pathway we introduce a new software: Triangulation of Perturbation Origins and Identification of Non-Coding Targets. Triangulation of Perturbation Origins and Identification of Non-Coding Targets is a pathway analysis tool, implemented in Java that identifies the significance of a gene under a condition (e.g. a disease phenotype) by studying graph representations of pathways, analyzing upstream and downstream gene interactions and integrating non-coding regions that may be regulating gene expression levels. AVAILABILITY AND IMPLEMENTATION: The TriPOINT open source software is freely available at https://github.uconn.edu/ajt06004/TriPOINT under the GPL v3.0 license. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Asa Thibodeau, Dong-Guk Shin |
Bioinform. | 2 |
| 2016 | Deep pathway analysis incorporating mutation information and gene expression dataabstractWe propose a new way of analyzing biological pathways in which the analysis combines both transcriptome data and mutation information and uses the outcome to identify routes of aberrant pathways potentially responsible for the etiology of disease. Each pathway route is encoded as a Bayesian Network which is initialized with a sequence of conditional probabilities which are designed to encode directionality of regulatory relationships encoded in the pathways, i.e. activation and inhibition relationships. First, we demonstrate the effectiveness of our model through simulation in which the model is able to discern patients in Test Group from ones in Control Group. Second, we apply our model to analyze the Breast Cancer data set, available from TCGA, against some pathways available from KEGG. Our experiment with this published data reaffirms the claims reported from the original breast cancer PAM50 subtype study. Our model can further analyze the patients of each subtype based on the identified route of aberration. For example, our analysis shows that complex biological process patterns are presented for HER2+ patients potentially suggesting our method's use for producing refined subtyping. We manage to find commonly perturbed pathway routes for HER2+ patients. We claim such “deep” pathway analysis could be very useful in designing a personalized therapy. Tham H. Hoang, Pujan Joshi, Seung-Hyun Hong, Dong-Guk Shin |
BIBM | 5 |
| 2016 | QuIN: A Web Server for Querying and Visualizing Chromatin Interaction NetworksabstractUNLABELLED: Recent studies of the human genome have indicated that regulatory elements (e.g. promoters and enhancers) at distal genomic locations can interact with each other via chromatin folding and affect gene expression levels. Genomic technologies for mapping interactions between DNA regions, e.g., ChIA-PET and HiC, can generate genome-wide maps of interactions between regulatory elements. These interaction datasets are important resources to infer distal gene targets of non-coding regulatory elements and to facilitate prioritization of critical loci for important cellular functions. With the increasing diversity and complexity of genomic information and public ontologies, making sense of these datasets demands integrative and easy-to-use software tools. Moreover, network representation of chromatin interaction maps enables effective data visualization, integration, and mining. Currently, there is no software that can take full advantage of network theory approaches for the analysis of chromatin interaction datasets. To fill this gap, we developed a web-based application, QuIN, which enables: 1) building and visualizing chromatin interaction networks, 2) annotating networks with user-provided private and publicly available functional genomics and interaction datasets, 3) querying network components based on gene name or chromosome location, and 4) utilizing network based measures to identify and prioritize critical regulatory targets and their direct and indirect interactions. AVAILABILITY: QuIN's web server is available at http://quin.jax.org QuIN is developed in Java and JavaScript, utilizing an Apache Tomcat web server and MySQL database and the source code is available under the GPLV3 license available on GitHub: https://github.com/UcarLab/QuIN/. Asa Thibodeau, Eladio J. Márquez, Oscar Luo, Yijun Ruan, Francesca Menghi, Dong-Guk Shin, Michael L. Stitzel, Paola Vera-Licona, Duygu Ucar |
PLoS Comput. Biol. | 6 |
| 2013 | A software framework integrating gene expression patterns, binding site analysis and gene ontology to hypothesize gene regulation relationshipsabstractOne known challenge in analyzing gene expression data is to combine analysis outcomes obtained disparately by applying multiple, independent meta-analysis methods. Here we present an integrative computational system that narrows down biological hypotheses by integrating gene expression patterns, transcription factor (TF) binding site analysis outcomes, and Gene Ontology (GO) enrichment analysis outcomes. This system identifies regulated genes from microarray experiments through statistical processes, categorizes similarly behaving groups of genes and then carries out binding site analysis and gene function enrichment analysis based on some significant clusters. The output is an ordered set of "putative" pair-wise relationships between TFs and their potential target genes. The relationships are ranked based on their closeness to the experimental context. We demonstrate the effectiveness of our framework using two independent microarray data sets. Pujan Joshi, Baikang Pei, Seung-Hyun Hong, Ivo Kalajzic, Dong-Ju Shin, David W. Rowe, Dong-Guk Shin |
BIBM | 7 |
| 2009 | Computing Consistency Between Microarray Data and Known Gene Regulation RelationshipsabstractMicroarray experiments produce expression patterns for thousands of genes at once. On the other hand, biomedical literature contains large amounts of gene regulation relationship information accumulated over the years. One obvious requirement is an automated way of comparing microarray data with the collection of known gene regulation relationships. Such an automated comparison is imperative because it can help biologists rapidly understand the context of a given microarray experiment. In addition, the consistency measure can be used to either validate or refute the hypothesis being tested using the microarray experiment. In this paper we present a systematic way of examining the consistency between a given set of microarray data and known gene regulation relationships. We first introduce a simple gene regulation network model with two separate algorithms designed to isolate a maximally consistent network. Subsequently, we extend the model to take into account multiple regulating factors for a single gene while highlighting both consistencies and inconsistencies. We illustrate the effectiveness of our approach with two practical examples, one that picks the peroxisome proliferator-activated receptor (PPAR) pathway as highly consistent from multiple pathways of Kyoto encyclopedia of genes and genomes (KEGG), and another that isolates key regulatory relationships involving nfkb1 and others known for macrophage's counter response to inflammation. Dong-Guk Shin, Saira Ali Kazmi, Baikang Pei, Yoo-Ah Kim, Jeffrey Maddox, Ravi Nori, Winfried Krueger, David W. Rowe |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2008 | Meta Analysis of Microarray Data Using Gene Regulation PathwaysabstractUsing microarray technology for genetic analysis in biological experiments requires computationally intensive tools to interpret results. The main objective here is to develop a ldquometa-analysisrdquo tool that enables researchers to ldquosprayrdquo microarray data over a network of relevant gene regulation relationships, extracted from a database of published gene regulatory pathway models. The consistency of the data from a microarray experiment is evaluated to determine if it agrees or contradicts with previous findings. The database is limited to ldquoactivaterdquo and ldquoinhibitrdquo gene regulatory relationships at this point and a heuristic graph based approach is developed for consistency checking. Predictions are made for the regulation of genes that were not a part of the microarray experiment, but are related to the experiment through regulatory relationships. This meta-analysis will not only highlight consistent findings but also pinpoint genes that were missed in earlier experiments and should be considered in subsequent analysis. Saira Ali Kazmi, Yoo-Ah Kim, Baikang Pei, Ravi Nori, David W. Rowe, Hsin-Wei Wang, Dong-Guk Shin |
BIBM | 8 |
| 2008 | Reverse Engineering of Gene Regulatory Network by Integration of Prior Global Gene Regulatory InformationabstractA Bayesian network is a model to study the structures of gene regulatory networks. It has the ability to integrate information from both prior knowledge and experimental data. Some previous works have explored the advantage of using prior knowledge. Unfortunately, most of the existing works only utilize biological knowledge about local structures of each gene in the network. In this study, we propose an approach to efficiently integrate global ordering information into model learning, where the ordering information specifies the indirect relationships among genes. We study the model behaviors with synthetic data. We demonstrate that, compared with a traditional Bayesian network model that uses only local prior knowledge, utilizing additional global ordering knowledge can significantly improve the modelpsilas performance. The magnitude of this improvement depends on how much global ordering information is integrated and how much noise the data includes. Baikang Pei, David W. Rowe, Dong-Guk Shin |
BIBM | 3 |
| 2006 | A Computational Inference Framework for analyzing Gene Regulation Pathway using Microarray DataabstractMicroarray experiments produce gene expression data at such a high speed and volume that it is imperative to use highly specialized computational tools for their analyses. One group of such computational tools deals with, namely, "meta-analysis" of microarray data. This step attempts to extract biological interpretations from the identified gene expression pattern. One particular aspect of meta-analysis is sorting out which gene regulation pathways are active and/or inhibited. The focus on this paper is to propose a computational framework with which scientists can compare microarray data with known gene regulation networks that are formed by two known binary gene regulation relationships, activate and inhibit. Using this framework scientists can conduct numerous analysis tasks including (i) identify active or inhibited sub-networks out of massively interconnected gene regulation pathways, (ii) find key genes, namely hubs, that are inferred to be widely involved in multiple aspects of gene regulation, (iii) identify regions of the network that contradict the known regulation, and (iv) estimate the direction of expression of genes that were not included in the microarray experiment. One known utility of this meta-analysis is to help scientists to identify a group of genes that they have missed in earlier experiments and should include in their subsequent experiments. Another utility is to enable them to isolate the group of genes that should be followed more closely using higher accuracy gene expression assays. We introduce three separate but inter-related meta-analysis methodologies, namely, FCFS, Hub majority, and SDP-based. We illustrate our proposed framework using microarray data derived from noggin treated osteoblast cells. The example clearly shows finding sub-networks related to B-catenin which should be expected and thus demonstrates the effectiveness of our proposed framework Dong-Guk Shin, John Bluis, Yoo-Ah Kim, Winfried Krueger, Jeffrey Maddox, Ravi Nori, Nathan Viniconis, Hsin-Wei Wang, David W. Rowe |
BIBE | 1 |
| 2003 | Nodal Distance Algorithm: Calculating a Phylogenetic Tree Comparison MetriabstractMaintaining a phylogenetic relationship repository requires the development of tools that are useful for mining the data stored in the repository. One way to query a database of phylogenetic information would be to compare phylogenetic trees. Because the only existing tree comparison methods are computationally intensive, this is not a reasonable task. Presented here is the nodal distance algorithm which has significantly less computation time than the most widely used comparison method, the partition metric. When the metric is calculated for trees where one species has been repositioned to a distant part of the tree no further computation is required as is needed for the partition metric. The nodal distance algorithm provides a method for comparing large sets of phylogenetic trees in a reasonable amount of time. John Bluis, Dong-Guk Shin |
BIBE | 2 |
| 2001 | Comparing Trees in a Phylogenetic Relationship RepositoryabstractScientists are able to use many different phylogenetic analysis tools to assist them in their research. The data collection and data analysis processes can take days or weeks to complete. Having this data available in a repository would reduce this process time and allow researchers to spend more time analyzing data instead of collecting it. We collected data for 21 completely sequenced genomes and created an intuitive interface for browsing the genomes. The interface allows users to search this data including the pre-calculated phylogenetic trees that are stored in the database. We have also developed a new method for comparing a large set of phylogenetic trees so users can search the database based on a user given hypothesis tree. John Bluis, Ravi Nori, Hsin-Wei Wang, Pinglei Zhou, J. Peter Gogarten, Dong-Guk Shin |
BIBE | 6 |
| 2000 | CBM: A Visual Query Interface Model based on Annotated Cartoon DiagramsabstractThis paper describes a novel visual query interface model, called Cartoon-Based Model (CBM), which aims at assisting novice database users to understand the content of a database with the views that they are mostly familiar with.CBM is different from previously proposed graphical query interfaces, many of which are based on semantic data models.CBM suggests a new dictum for building a next generation user-friendly database query interface: Visual query interfaces should be built based upon the "perception" (e.g., cartoon diagrams in our case) that the user would possess about the data content, not the artificially modeled view that a database designer would possess.A prototype has been constructed to demonstrate the soundness and feasibility of the proposed model. Dong-Guk Shin, Ravi Shankar Nori |
Advanced Visual Interfaces | 1 |
| 2000 | A Graph-based Meta-data Framework for Interoperation between Genome DatabasesabstractThe proliferation, diversity and complexity of genome databases pose a significant challenge to the multidatabase research community. We propose a meta-data framework based on a graph modeling technique and show how this framework can be applied to making two genome databases interoperable. Our technique maps individual database schemas expressed in heterogeneous data models into a common graph representation. This graph-based framework is designed to express both regular data and meta-data. Modeling both types of data is necessary for establishing database correspondences. We show how inter-database relationships can be expressed in terms of correspondences amongst basic graph units, including nodes, edges and the "value path". We apply these basic graph units to perform query interoperation between two databases, namely DB/12 (Database of Human Chromosome 12) and GDB (Genome Database). Kei-Hoi Cheung, Dong-Guk Shin |
BIBE | 2 |
| 1999 | Xyntagma: a Graphical Query Interface for the ACeDB Genome DatabasesabstractAs the number of non-expert database users who wish to access genome databases directly grows, it becomes increasingly imperative to develop an easy-to-use query interface for such databases. Relying on a graphical solution is one way of addressing the problem. A graphical query interface provides a built-in schema browsing mechanism, which enables users to express queries more easily than with a textual query interface. We summarize our experience of building a graphical query interface for ACeDB genome databases. ACeDB is a pseudo-object oriented database system which was primarily developed to store biological data and to be used by biologists. Our interface, entitled Xyntagma, allows the user to browse an object-oriented schema and enables him to graphically express ad hoc queries that conform to the ACeDB query syntax. Wally Grajewski, Dong-Guk Shin |
CBMS | 2 |
| 1999 | A Multistrategy Approach to Classification Learning in Databases
Changhwan Lee, Dong-Guk Shin |
Data Knowl. Eng. | 2 |
| 1999 | Using Hellinger distance in a nearest neighbour classifier for relational databases
Changhwan Lee, Dong-Guk Shin |
Knowl. Based Syst. | 2 |
| 1998 | Automatic query mapping among genomic databases: a pilot exploration
Kei-Hoi Cheung, Prakash M. Nadkarni, Perry L. Miller, Dong-Guk Shin |
AMIA | 4 |
| 1998 | An epistemological display query interfaceabstractDatabase users without sufficient understanding of database schemas are normally daunted by the task of query formulation. To alleviate these concerns, we have come up with a visual query system that allows users to express queries at the conceptual level. This interface works by representing a targeted database schema at the conceptual level and allows the user to enter restriction, projection and join conditions in an interactive fashion. The mapping between the conceptual level and the relational schema of a targeted database is accomplished via the usage of the knowledge representation language Lk. We have built a prototype that works with human genome databases. Dong-Guk Shin, Wally Grajewski, Lung-Yung Chu |
AVI | 1 |
| 1998 | A metadata approach to query interoperation between molecular biology databasesabstractMOTIVATION: Molecular biology databases have been proliferating rapidly. Their heterogeneity and complexity pose a great challenge to efforts in database interoperation. To minimize the efforts of interoperating heterogeneous databases, it is useful to develop a system that lets a user of a particular genomic database access another related database as if the latter is structurally similar to the former. RESULTS: We extend a structurally simple model-the entity-attribute-value (EAV) model-to describe uniformly metadata relating to individual databases. Such metadata, which are necessary for performing database comparisons, include descriptions of primitive database objects (including entities, attributes, domain values and entity relationships) and specification of correspondences among the database objects. We show how to decompose SQL queries and map them from one database to another based on the EAV representation of the basic database objects. A prototype system is implemented to demonstrate query interoperation between two chromosome map databases. AVAILABILITY: Freely available (Cold Fusion source code and an Access database containing the mapping knowledge) upon request from the author. CONTACT: [email protected] Kei-Hoi Cheung, Prakash M. Nadkarni, Dong-Guk Shin |
Bioinform. | 3 |
| 1998 | A Methodology of Constructing Canonical Form Database Schemas in a Multiple Heterogeneous Database EnvironmentabstractThe Query Clearing House (QCH) model aims at achieving an ideal heterogeneous multiple database environment in which users can submit queries without concern for the location of the data or the specifics of the relevant database schemas. One important prerequisite for building such an environment is to organize local database schemas in a cohesive way so that a systematic method of determining data relevancy can be developed. In this paper, we propose a method for converting multiple local relational schemas into canonical form expressions. The resulting canonical form schema expressions include much more descriptive data semantics than what is offered by the original database schemas. These expressions also include information for mapping between the terms used in the canonical form expressions and their counterparts implemented in local databases. We include a number of examples illustrating how the canonical form schema expressions are produced.Request access from your librarian to read this article's full text. Jeong Seok Lim, Dong-Guk Shin |
J. Database Manag. | 2 |
| 1994 | A Context-Sensitive Discretization of Numeric Attributes for Classification Learning
Changhwan Lee, Dong-Guk Shin |
ECAI | 2 |
| 1994 | EEL: An Instance-Based Learning Method for Databases
Changhwan Lee, Dong-Guk Shin |
IEA/AIE | 2 |
| 1994 | An Expection-Driven Response Understanding ParadigmabstractThis paper describes a model that can account for ad hoc user-responses to argument interrogative type of system-initiated questions. Successful implementation of the model can provide an alternative solution that is more effective than the menu-driven approach that has been proposed as a meager solution to enable the system to ask a question to the user. The proposed model assumes that when the system asks a question, it maintains an expectation of the potential answers. The system then uses the expectation as the focus to perform the most likely interpretation of the user's response. Without using such a focus the interpretation process could be unbounded. The interpretation process is mapped into a heuristic search problem. The interpretation process results in identifying a particular expectation-response relationship type, which the system can use to tailor its response strategy with respect to the given user-response. A prototype has been constructed to demonstrate the soundness of the proposed model.> Dong-Guk Shin |
IEEE Trans. Knowl. Data Eng. | 1 |
| 1992 | Redesigning, implementing and integrating Escherichia coli genome software tools with an object-oriented database systemabstractThis paper reports our exploratory work to redesign, implement and integrate a collection of genome software tools with an object-oriented database system. Our software tools deal with genome data from Escherichia coli K-12, a bacterium that has been studied intensively and provides richer data sets than any other living organism. The object-oriented DBMS used for the integration is ONTOS, a commercial object-oriented system from Ontologic Inc. This redesign and implementation task was performed in two steps. First, C programs were converted into C++, and then the C++ version programs were modified and integrated with an object-oriented modeling of the data to form an ONTOS database application. The first step helps us develop a conceptual view for a DBMS-independent object-oriented construct. The second step elucidates what additional DBMS-dependent modification steps are needed to provide persistency to the objects. Examples are included to illustrate steps of the redesign and implementation. Overall, the outcome of this project demonstrates that programs and data can be successfully integrated with an object-oriented database, while providing the objects with persistency and shareability. This paper includes discussions using concrete examples on what advantage the object-oriented database approach provides over the relational database approach. Dong-Guk Shin, Changhwan Lee, Kenneth E. Rudd, Claire M. Berg |
Comput. Appl. Biosci. | 1 |
| 1991 | Fragmenting Relations Horizontally Using a Knowledge-Based ApproachabstractIn distributed DBMSs, one major issue in developing a horizontal fragmentation technique is what criteria to use to guide the fragmentation. The authors propose to use, in addition to typical user queries, particular knowledge about the data itself. Use of this knowledge allows revision of typical user queries into more precise forms. The revised query expressions produce better estimations of user reference clusters to the database than the original query expressions. The estimated user reference clusters form a basis to partition relations horizontally. In the proposed approach, an ordinary many-sorted language is extended to represent the queries and knowledge compatibly. This knowledge is identified in terms of five axiom schemata. An inference procedure is developed to apply the knowledge to the queries deductively.> Dong-Guk Shin, Keki B. Irani |
IEEE Trans. Software Eng. | 1 |
| 1990 | Am/AG Model: a Hierarchical Social System Metaphor for Distributed Problem SolvingabstractThis work explores a distributed problem solving (DPS) approach, namely the AM/AG (Amplification/Aggregation) model. The AM/AG model is a hierarchic social system metaphor for DPS based on Mintzberg’s model of organizations. At the core of the model are information flow mechanisms, namely, amplification and aggregation. Amplification is a process of decomposing a given task, called an agenda, into a set of subtasks with magnified degree of specificity and distributing them to multiple processing units downward in the hierarchy. Aggregation is a process of combining the results reported from multiple processing units into a unified view, called a resolution, and promoting the conclusion upward in the hierarchy. Amplification is discussed in detail. A set of generative rules is introduced. Each rule specifies a set of actions for transforming an input agenda into other forms with higher specificity. The proposed model can be used to account for the memory recall process which makes associations between vast amounts of related concepts, sorts out the combined results, and promotes the most plausible ones. An example of memory recall is used to illustrate the model. Dong-Guk Shin, Joseph Leone |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1988 | Knowledge-aided engineering environment for design and manufacturingabstractThe development of a knowledge-aided engineering environment, an emerging concept that aims to exploit the commonality and the adaptability of the knowledge underlying design and manufacturing, is presented. The following objectives are critical for achieving this end: (1) the definition of an integrated set of conceptual and physical data models that can support the unique requirements of the engineering and manufacturing environments and also promote the sharing of common knowledge; (2) the production of a mechanism for the automatic generation of such models and their accompanying software for specific application domains; (3) the creation of a strategy for knowledge-driven dialogue management; and (4) the specification of a rule base that supports the integration of the dialogue manager and the data model generator into a functional prototype. An example illustrating the KAD/KAM (knowledge-aided design/knowledge-aided manufacturing) environment is presented.> Dong-Guk Shin, Fred J. Maryanski |
COMPSAC | 1 |