EDBT 2026 Demo / reviewers in the wild / expert
Do Kyoon Kim
dblp:13/8425 · also Dokyoon Kim
· DBLP profile ↗
20ranked-venue papers
4as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GeOKG: geometry-aware knowledge graph embedding for Gene Ontology and genesabstractMOTIVATION: Leveraging deep learning for the representation learning of Gene Ontology (GO) and Gene Ontology Annotation (GOA) holds significant promise for enhancing downstream biological tasks such as protein-protein interaction prediction. Prior approaches have predominantly used text- and graph-based methods, embedding GO and GOA in a single geometric space (e.g. Euclidean or hyperbolic). However, since the GO graph exhibits a complex and nonmonotonic hierarchy, single-space embeddings are insufficient to fully capture its structural nuances. RESULTS: In this study, we address this limitation by exploiting geometric interaction to better reflect the intricate hierarchical structure of GO. Our proposed method, Geometry-Aware Knowledge Graph Embeddings for GO and Genes (GeOKG), leverages interactions among various geometric representations during training, thereby modeling the complex hierarchy of GO more effectively. Experiments at the GO level demonstrate the benefits of incorporating these geometric interactions, while gene-level tests reveal that GeOKG outperforms existing methods in protein-protein interaction prediction. These findings highlight the potential of using geometric interaction for embedding heterogeneous biomedical networks. AVAILABILITY AND IMPLEMENTATION: https://github.com/ukjung21/GeOKG. Chang-Uk Jeong, Jaesik Kim, Do Kyoon Kim, Kyung-Ah Sohn 0001 |
Bioinform. | 3 |
| 2025 | Explainable multiplex graph propagational network with multimodal neuroimage integration for dementia subtype diagnosis
Sunghong Park, Dong-gi Lee, Do Kyoon Kim, Yonghyun Nam, Bumhee Park, Narae Kim, Seulgi Lee, Changhyung Hong, Sangjoon Son, Hyunwoong Roh, Hyun Goo Woo, Hyunjung Shin |
Neural Networks | 4 |
| 2024 | Uncovering genetic associations in the human diseasome using an endophenotype-augmented disease networkabstractMOTIVATION: Many diseases, particularly cardiometabolic disorders, exhibit complex multimorbidities with one another. An intuitive way to model the connections between phenotypes is with a disease-disease network (DDN), where nodes represent diseases and edges represent associations, such as shared single-nucleotide polymorphisms (SNPs), between pairs of diseases. To gain further genetic understanding of molecular contributors to disease associations, we propose a novel version of the shared-SNP DDN (ssDDN), denoted as ssDDN+, which includes connections between diseases derived from genetic correlations with intermediate endophenotypes. We hypothesize that a ssDDN+ can provide complementary information to the disease connections in a ssDDN, yielding insight into the role of clinical laboratory measurements in disease interactions. RESULTS: Using PheWAS summary statistics from the UK Biobank, we constructed a ssDDN+ revealing hundreds of genetic correlations between diseases and quantitative traits. Our augmented network uncovers genetic associations across different disease categories, connects relevant cardiometabolic diseases, and highlights specific biomarkers that are associated with cross-phenotype associations. Out of the 31 clinical measurements under consideration, HDL-C connects the greatest number of diseases and is strongly associated with both type 2 diabetes and heart failure. Triglycerides, another blood lipid with known genetic causes in non-mendelian diseases, also adds a substantial number of edges to the ssDDN. This work demonstrates how association with clinical biomarkers can better explain the shared genetics between cardiometabolic disorders. Our study can facilitate future network-based investigations of cross-phenotype associations involving pleiotropy and genetic heterogeneity, potentially uncovering sources of missing heritability in multimorbidities. AVAILABILITY AND IMPLEMENTATION: The generated ssDDN+ can be explored at https://hdpm.biomedinfolab.com/ddn/biomarkerDDN. Jakob Woerner, Vivek Sriram, Yonghyun Nam, Anurag Verma, Do Kyoon Kim |
Bioinform. | 5 |
| 2024 | Non-invasive prediction of massive transfusion during surgery using intraoperative hemodynamic monitoring data
Doyun Kwon, Young Mi Jung, Hyung-Chul Lee, Tae Kyong Kim, Garam Lee, Do Kyoon Kim, Seung-Bo Lee, Seung Mi Lee |
J. Biomed. Informatics | 7 |
| 2024 | Frequency Domain Deep Learning With Non-Invasive Features for Intraoperative Hypotension PredictionabstractBACKGROUND: Intraoperative hypotension can lead to postoperative organ dysfunction. Previous studies primarily used invasive arterial pressure as the key biosignal for the detection of hypotension. However, these studies had limitations in incorporating different biosignal modalities and utilizing the periodic nature of biosignals. To address these limitations, we utilized frequency-domain information, which provides key insights that time-domain analysis cannot provide, as revealed by recent advances in deep learning. With the frequency-domain information, we propose a deep-learning approach that integrates multiple biosignal modalities. METHODS: We used the discrete Fourier transform technique, to extract frequency information from biosignal data, which we then combined with the original time-domain data as input for our deep learning model. To improve the interpretability of our results, we incorporated recent interpretable modules for deep-learning models into our analysis. RESULTS: We constructed 75 994 segments from the data of 3226 patients to predict hypotension during surgery. Our proposed frequency-domain deep-learning model outperformed conventional approaches that rely solely on time-domain information. Notably, our model achieved a greater increase in AUROC performance than the time-domain deep learning models when trained on non-invasive biosignal data only (AUROC 0.898 [95% CI: 0.885-0.91] vs. 0.853 [95% CI: 0.839-0.867]). Further analysis revealed that the 1.5-3.0 Hz frequency band played an important role in predicting hypotension events. CONCLUSION: Utilizing the frequency domain not only demonstrated high performance on invasive data but also showed significant performance improvement when applied to non-invasive data alone. Our proposed framework offers clinicians a novel perspective for predicting intraoperative hypotension. Jeong-Hyeon Moon, Garam Lee, Seung Mi Lee, Jiho Ryu, Do Kyoon Kim, Kyung-Ah Sohn 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2023 | Discovering comorbid diseases using an inter-disease interactivity network based on biobank-scale PheWAS dataabstractMOTIVATION: Understanding comorbidity is essential for disease prevention, treatment and prognosis. In particular, insight into which pairs of diseases are likely or unlikely to co-occur may help elucidate the potential relationships between complex diseases. Here, we introduce the use of an inter-disease interactivity network to discover/prioritize comorbidities. Specifically, we determine disease associations by accounting for the direction of effects of genetic components shared between diseases, and categorize those associations as synergistic or antagonistic. We further develop a comorbidity scoring algorithm to predict whether diseases are more or less likely to co-occur in the presence of a given index disease. This algorithm can handle networks that incorporate relationships with opposite signs. RESULTS: We finally investigate inter-disease associations among 427 phenotypes in UK Biobank PheWAS data and predict the priority of comorbid diseases. The predicted comorbidities were verified using the UK Biobank inpatient electronic health records. Our findings demonstrate that considering the interaction of phenotype associations might be helpful in better predicting comorbidity. AVAILABILITY AND IMPLEMENTATION: The source code and data of this study are available at https://github.com/dokyoonkimlab/DiseaseInteractiveNetwork. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yonghyun Nam, Sang-Hyuk Jung, Jae-Seung Yun, Vivek Sriram, Pankhuri Singhal, Marta Byrska-Bishop, Anurag Verma, Hyunjung Shin, Woong-Yang Park, Hong-Hee Won, Do Kyoon Kim |
Bioinform. | 11 |
| 2022 | Mediation Analysis and Mixed-Effects Models for the Identification of Stage-specific Imaging Genetics Patterns in Alzheimer's DiseaseabstractAlzheimer's disease (AD) is one of the most common and severe forms of Senile Dementia. Genome-wide association studies (GWAS) have identified dozens of AD susceptible loci. To better understand potential mechanism-of-action for AD, quantitative brain imaging features have been studied as mediators linking genetic variants to AD outcomes. In this study, Mediation analysis, Chow test and Mixed-effects Models are used to investigate the biological pathways by which genetic variants affect both brain structures/functions and disease diagnosis. We analyzed the imaging and genetics data collected from the Alzheimer's Disease Neuroimaging Initiative (ADNI) project, including a Polygenic Hazard Score (PHS) and 13 imaging quantitative traits (QTs) extracted from the AV45 PET scans quantifying the amyloid deposition in different brain regions of subjects from four separate diagnostic groups. Mediation analysis assessed the mediating effects of image QTs between PHS and diagnosis, whereas Chow test and Linear Mixed-Effects models were used to characterize intra-group differences in the associations between genetic scores and imaging QTs for different disease stages. Results show that promising stage-specific imaging QTs that mediate the genetic effect of the studied PHS on disease status have been identified, providing novel insights into the predictive power of the PHS and the mediating power of amyloid imaging QTs with respect to multiple stages over the AD progression. Daniele Pala, Xia Ning, Do Kyoon Kim, Li Shen 0001 |
BIBM | 4 |
| 2022 | Identifying genes associated with brain volumetric differences through tissue specific transcriptomic inference from GWAS summary dataabstractBACKGROUND: Brain volume has been widely studied in the neuroimaging field, since it is an important and heritable trait associated with brain development, aging and various neurological and psychiatric disorders. Genome-wide association studies (GWAS) have successfully identified numerous associations between genetic variants such as single nucleotide polymorphisms and complex traits like brain volume. However, it is unclear how these genetic variations influence regional gene expression levels, which may subsequently lead to phenotypic changes. S-PrediXcan is a tissue-specific transcriptomic data analysis method that can be applied to bridge this gap. In this work, we perform an S-PrediXcan analysis on GWAS summary data from two large imaging genetics initiatives, the UK Biobank and Enhancing Neuroimaging Genetics through Meta Analysis, to identify tissue-specific transcriptomic effects on two closely related brain volume measures: total brain volume (TBV) and intracranial volume (ICV). RESULTS: As a result of the analysis, we identified 10 genes that are highly associated with both TBV and ICV. Nine out of 10 genes were found to be associated with TBV in another study using a different gene-based association analysis. Moreover, most of our discovered genes were also found to be correlated with multiple cognitive and behavioral traits. Further analyses revealed the protein-protein interactions, associated molecular pathways and biological functions that offer insight into how these genes function and interact with others. CONCLUSIONS: These results confirm that S-PrediXcan can identify genes with tissue-specific transcriptomic effects on complex traits. The analysis also suggested novel genes whose expression levels are related to brain volumetric traits. This provides important insights into the genetic mechanisms of the human brain. Hung Mai, Jingxuan Bao, Paul M. Thompson, Do Kyoon Kim, Li Shen 0001 |
BMC Bioinform. | 4 |
| 2021 | Interpretable temporal graph neural network for prognostic prediction of Alzheimer's disease using longitudinal neuroimaging dataabstractAlzheimer's disease (AD) is a progressive neurodegenerative brain disorder characterized by memory loss and cognitive decline. Early detection and accurate prognosis of AD is an important research topic, and numerous machine learning methods have been proposed to solve this problem. However, traditional machine learning models are facing challenges in effectively integrating longitudinal neuroimaging data and biologically meaningful structure and knowledge to build accurate and interpretable prognostic predictors. To bridge this gap, we propose an interpretable graph neural network (GNN) model for AD prognostic prediction based on longitudinal neuroimaging data while embracing the valuable knowledge of structural brain connectivity. In our empirical study, we demonstrate that 1) the proposed model outperforms several competing models (i.e., DNN, SVM) in terms of prognostic prediction accuracy, and 2) our model can capture neuroanatomical contribution to the prognostic predictor and yield biologically meaningful interpretation to facilitate better mechanistic understanding of the Alzheimer's disease. Source code is available at https://github.com/JaesikKim/temporal-GNN. Mansu Kim, Jaesik Kim, Jeffrey Qu, Heng Huang 0001, Qi Long, Kyung-Ah Sohn 0001, Do Kyoon Kim, Li Shen 0001 |
BIBM | 7 |
| 2021 | Multi-layered network-based pathway activity inference using directed random walks: application to predicting clinical outcomes in urologic cancerabstractMOTIVATION: To better understand the molecular features of cancers, a comprehensive analysis using multi-omics data has been conducted. In addition, a pathway activity inference method has been developed to facilitate the integrative effects of multiple genes. In this respect, we have recently proposed a novel integrative pathway activity inference approach, iDRW and demonstrated the effectiveness of the method with respect to dichotomizing two survival groups. However, there were several limitations, such as a lack of generality. In this study, we designed a directed gene-gene graph using pathway information by assigning interactions between genes in multiple layers of networks. RESULTS: As a proof-of-concept study, it was evaluated using three genomic profiles of urologic cancer patients. The proposed integrative approach achieved improved outcome prediction performances compared with a single genomic profile alone and other existing pathway activity inference methods. The integrative approach also identified common/cancer-specific candidate driver pathways as predictive prognostic features in urologic cancers. Furthermore, it provides better biological insights into the prioritized pathways and genes in an integrated view using a multi-layered gene-gene network. Our framework is not specifically designed for urologic cancers and can be generally applicable for various datasets. AVAILABILITY AND IMPLEMENTATION: iDRW is implemented as the R software package. The source codes are available at https://github.com/sykim122/iDRW. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. So Yeon Kim, Eun Kyung Choe, Manu K. Shivakumar, Do Kyoon Kim, Kyung-Ah Sohn 0001 |
Bioinform. | 4 |
| 2021 | HiG2Vec: hierarchical representations of Gene Ontology and genes in the Poincaré ballabstractMOTIVATION: Knowledge manipulation of Gene Ontology (GO) and Gene Ontology Annotation (GOA) can be done primarily by using vector representation of GO terms and genes. Previous studies have represented GO terms and genes or gene products in Euclidean space to measure their semantic similarity using an embedding method such as the Word2Vec-based method to represent entities as numeric vectors. However, this method has the limitation that embedding large graph-structured data in the Euclidean space cannot prevent a loss of information of latent hierarchies, thus precluding the semantics of GO and GOA from being captured optimally. On the other hand, hyperbolic spaces such as the Poincaré balls are more suitable for modeling hierarchies, as they have a geometric property in which the distance increases exponentially as it nears the boundary because of negative curvature. RESULTS: In this article, we propose hierarchical representations of GO and genes (HiG2Vec) by applying Poincaré embedding specialized in the representation of hierarchy through a two-step procedure: GO embedding and gene embedding. Through experiments, we show that our model represents the hierarchical structure better than other approaches and predicts the interaction of genes or gene products similar to or better than previous studies. The results indicate that HiG2Vec is superior to other methods in capturing the GO and gene semantics and in data utilization as well. It can be robustly applied to manipulate various biological knowledge. AVAILABILITYAND IMPLEMENTATION: https://github.com/JaesikKim/HiG2Vec. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jaesik Kim, Do Kyoon Kim, Kyung-Ah Sohn 0001 |
Bioinform. | 2 |
| 2020 | Liver imaging features by convolutional neural network to predict the metachronous liver metastasis in stage I-III colorectal cancer patients based on preoperative abdominal CT scanabstractAbstract Background Introducing deep learning approach to medical images has rendered a large amount of un-decoded information into usage in clinical research. But mostly, it has been focusing on the performance of the prediction modeling for disease-related entity, but not on the clinical implication of the feature itself. Here we analyzed liver imaging features of abdominal CT images collected from 2019 patients with stage I – III colorectal cancer (CRC) using convolutional neural network (CNN) to elucidate its clinical implication in oncological perspectives. Results CNN generated imaging features from the liver parenchyma. Dimension reduction was done for the features by principal component analysis. We designed multiple prediction models for 5-year metachronous liver metastasis (5YLM) using combinations of clinical variables (age, sex, T stage, N stage) and top principal components (PCs), with logistic regression classification. The model using “1 st PC (PC1) + clinical information” had the highest performance (mean AUC = 0.747) to predict 5YLM, compared to the model with clinical features alone (mean AUC = 0.709). The PC1 was independently associated with 5YLM in multivariate analysis (beta = − 3.831, P < 0.001). For the 5-year mortality rate, PC1 did not contribute to an improvement to the model with clinical features alone. For the PC1, Kaplan-Meier plots showed a significant difference between PC1 low vs. high group. The 5YLM-free survival of low PC1 was 89.6% and the high PC1 was 95.9%. In addition, PC1 had a significant correlation with sex, body mass index, alcohol consumption, and fatty liver status. Conclusion The imaging features combined with clinical information improved the performance compared to the standardized prediction model using only clinical information. The liver imaging features generated by CNN may have the potential to predict liver metastasis. These results suggest that even though there were no liver metastasis during the primary colectomy, the features of liver imaging can impose characteristics that could be predictive for metachronous liver metastasis. Eun Kyung Choe, So Yeon Kim, Hua Sun Kim, Kyu Joo Park, Do Kyoon Kim |
BMC Bioinform. | 6 |
| 2017 | Using knowledge-driven genomic interactions for multi-omics data analysis: metadimensional models for predicting clinical outcomes in ovarian carcinomaabstractIt is common that cancer patients have different molecular signatures even though they have similar clinical features, such as histology, due to the heterogeneity of tumors. To overcome this variability, we previously developed a new approach incorporating prior biological knowledge that identifies knowledge-driven genomic interactions associated with outcomes of interest. However, no systematic approach has been proposed to identify interaction models between pathways based on multi-omics data. Here we have proposed such a novel methodological framework, called metadimensional knowledge-driven genomic interactions (MKGIs). To test the utility of the proposed framework, we applied it to an ovarian cancer dataset including multi-omics profiles from The Cancer Genome Atlas to predict grade, stage, and survival outcome. We found that each knowledge-driven genomic interaction model, based on different genomic datasets, contains different sets of pathway features, which suggests that each genomic data type may contribute to outcomes in ovarian cancer via a different pathway. In addition, MKGI models significantly outperformed the single knowledge-driven genomic interaction model. From the MKGI models, many interactions between pathways associated with outcomes were found, including the mitogen-activated protein kinase (MAPK) signaling pathway and the gonadotropin-releasing hormone (GnRH) signaling pathway, which are known to play important roles in cancer pathogenesis. The beauty of incorporating biological knowledge into the model based on multi-omics data is the ability to improve diagnosis and prognosis and provide better interpretability. Thus, determining variability in molecular signatures based on these interactions between pathways may lead to better diagnostic/treatment strategies for better precision medicine. Do Kyoon Kim, Ruowang Li, Anastasia Lucas, Shefali S. Verma, Scott M. Dudek, Marylyn D. Ritchie |
J. Am. Medical Informatics Assoc. | 1 |
| 2015 | Knowledge boosting: a graph-based integration approach with multi-omics data and genomic knowledge for cancer clinical outcome predictionabstractOBJECTIVE: Cancer can involve gene dysregulation via multiple mechanisms, so no single level of genomic data fully elucidates tumor behavior due to the presence of numerous genomic variations within or between levels in a biological system. We have previously proposed a graph-based integration approach that combines multi-omics data including copy number alteration, methylation, miRNA, and gene expression data for predicting clinical outcome in cancer. However, genomic features likely interact with other genomic features in complex signaling or regulatory networks, since cancer is caused by alterations in pathways or complete processes. METHODS: Here we propose a new graph-based framework for integrating multi-omics data and genomic knowledge to improve power in predicting clinical outcomes and elucidate interplay between different levels. To highlight the validity of our proposed framework, we used an ovarian cancer dataset from The Cancer Genome Atlas for predicting stage, grade, and survival outcomes. RESULTS: Integrating multi-omics data with genomic knowledge to construct pre-defined features resulted in higher performance in clinical outcome prediction and higher stability. For the grade outcome, the model with gene expression data produced an area under the receiver operating characteristic curve (AUC) of 0.7866. However, models of the integration with pathway, Gene Ontology, chromosomal gene set, and motif gene set consistently outperformed the model with genomic data only, attaining AUCs of 0.7873, 0.8433, 0.8254, and 0.8179, respectively. CONCLUSIONS: Integrating multi-omics data and genomic knowledge to improve understanding of molecular pathogenesis and underlying biology in cancer should improve diagnostic and prognostic indicators and the effectiveness of therapies. Do Kyoon Kim, Je-Gun Joung, Kyung-Ah Sohn 0001, Hyunjung Shin, Yu Rang Park, Marylyn D. Ritchie, Ju Han Kim |
J. Am. Medical Informatics Assoc. | 1 |
| 2015 | Predicting censored survival data based on the interactions between meta-dimensional omics data in breast cancerabstractEvaluation of survival models to predict cancer patient prognosis is one of the most important areas of emphasis in cancer research. A binary classification approach has difficulty directly predicting survival due to the characteristics of censored observations and the fact that the predictive power depends on the threshold used to set two classes. In contrast, the traditional Cox regression approach has some drawbacks in the sense that it does not allow for the identification of interactions between genomic features, which could have key roles associated with cancer prognosis. In addition, data integration is regarded as one of the important issues in improving the predictive power of survival models since cancer could be caused by multiple alterations through meta-dimensional genomic data including genome, epigenome, transcriptome, and proteome. Here we have proposed a new integrative framework designed to perform these three functions simultaneously: (1) predicting censored survival data; (2) integrating meta-dimensional omics data; (3) identifying interactions within/between meta-dimensional genomic features associated with survival. In order to predict censored survival time, martingale residuals were calculated as a new continuous outcome and a new fitness function used by the grammatical evolution neural network (GENN) based on mean absolute difference of martingale residuals was implemented. To test the utility of the proposed framework, a simulation study was conducted, followed by an analysis of meta-dimensional omics data including copy number, gene expression, DNA methylation, and protein expression data in breast cancer retrieved from The Cancer Genome Atlas (TCGA). On the basis of the results from breast cancer dataset, we were able to identify interactions not only within a single dimension of genomic data but also between meta-dimensional omics data that are associated with survival. Notably, the predictive power of our best meta-dimensional model was 73% which outperformed all of the other models conducted based on a single dimension of genomic data. Breast cancer is an extremely heterogeneous disease and the high levels of genomic diversity within/between breast tumors could affect the risk of therapeutic responses and disease progression. Thus, identifying interactions within/between meta-dimensional omics data associated with survival in breast cancer is expected to deliver direction for improved meta-dimensional prognostic biomarkers and therapeutic targets. Do Kyoon Kim, Ruowang Li, Scott M. Dudek, Marylyn D. Ritchie |
J. Biomed. Informatics | 1 |
| 2014 | An integrated analysis of genome-wide DNA methylation and genetic variants underlying etoposide-induced cytotoxicity in European and African populations
Ruowang Li, Do Kyoon Kim, Scott M. Dudek, Marylyn D. Ritchie |
EvoApplications | 2 |
| 2013 | Robust predictive model for evaluating breast cancer survivability
Kanghee Park, Amna Ali, Do Kyoon Kim, Yeolwoo An, Minkoo Kim, Hyunjung Shin |
Eng. Appl. Artif. Intell. | 3 |
| 2013 | Research and applications: Extracting coordinated patterns of DNA methylation and gene expression in ovarian cancerabstractOBJECTIVE: DNA methylation, a regulator of gene expression, plays an important role in diverse biological processes including developmental process, carcinogenesis and aging. In particular, aberrant DNA methylation has been largely observed in several types of cancers. Currently, it is important to extract disease-specific gene sets associated with the regulation of DNA methylation. MATERIALS AND METHODS: Here we propose a novel approach to find the minimum regulatory units of genes, co-methylated and co-expressed gene pairs (MEGP) that are highly correlated gene pairs between DNA methylation and gene expression showing the co-regulatory relationship. To evaluate whether our method is applicable to extract disease-associated genes, we applied our method to a large-scale dataset from the Cancer Genome Atlas extracting significantly associated MEGP and analyzed their functional correlation. RESULTS: We observed that many MEGP physically interacted with each other and showed high semantic similarity with gene ontology terms. Furthermore, we performed gene set enrichment tests to identify how they are correlated in a complex biological process. Our MEGP were highly enriched in the biological pathway associated with ovarian cancers. CONCLUSIONS: Our approach is useful for discovering coordinated epigenetic markers associated with specific diseases. Je-Gun Joung, Do Kyoon Kim, Kyung-Hwa Kim, Ju Han Kim |
J. Am. Medical Informatics Assoc. | 2 |
| 2012 | Synergistic effect of different levels of genomic data for cancer clinical outcome prediction
Do Kyoon Kim, Hyunjung Shin, Young Soo Song, Ju Han Kim |
J. Biomed. Informatics | 1 |
| 2010 | TMA-TAB: A spreadsheet-based document for exchange of tissue microarray data based on the tissue microarray-object modelabstractThe importance of tissue microarrays (TMA) as clinical validation tools for cDNA microarray results is increasing, whereas researchers are still suffering from TMA data management issues. After we developed a comprehensive data model for TMA data storage, exchange and analysis, TMA-OM, we focused our attention on the development of a user-friendly exchange format with high expressivity in order to promote data communication of TMA results and TMA-OM supportive database applications. We developed TMA-TAB, a spreadsheet-based data format for TMA data submission to the TMA-OM supportive TMA database system. TMA-TAB was developed by simplifying, modifying and reorganizing classes, attributes and templates of TMA-OM into five entities: experiment, block, slide, core_in_block, and core_in_slide. Five tab-delimited formats (investigation design format, block description format, slide description format, core clinicohistopathological data format, and core result data format) were made, each representing the entities of experiment, block, slide, core_in_block, and core_in_slide. We implemented TMA-TAB import and export modules on Xperanto-TMA, a TMA-OM supportive database application, to facilitate data submission. Development and implementation of TMA-TAB and TMA-OM provide a strong infrastructure for powerful and user-friendly TMA data management. Young Soo Song, Hye Won Lee, Yu Rang Park, Do Kyoon Kim, Jaehyun Sim, Hyunseok Peter Kang, Ju Han Kim |
J. Biomed. Informatics | 4 |