Jihoon Kim 0001

dblp:51/4060-1 · DBLP profile ↗
← Back
31ranked-venue papers
6as first author
4since 2021 · last 2023
0000-0002-5351-238XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 30 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2023 Blockchain-enabled immutable, distributed, and highly available clinical research activity logging system for federated COVID-19 data analysis from multiple institutions
abstract
OBJECTIVE: We aimed to develop a distributed, immutable, and highly available cross-cloud blockchain system to facilitate federated data analysis activities among multiple institutions. MATERIALS AND METHODS: We preprocessed 9166 COVID-19 Structured Query Language (SQL) code, summary statistics, and user activity logs, from the GitHub repository of the Reliable Response Data Discovery for COVID-19 (R2D2) Consortium. The repository collected local summary statistics from participating institutions and aggregated the global result to a COVID-19-related clinical query, previously posted by clinicians on a website. We developed both on-chain and off-chain components to store/query these activity logs and their associated queries/results on a blockchain for immutability, transparency, and high availability of research communication. We measured run-time efficiency of contract deployment, network transactions, and confirmed the accuracy of recorded logs compared to a centralized baseline solution. RESULTS: The smart contract deployment took 4.5 s on an average. The time to record an activity log on blockchain was slightly over 2 s, versus 5-9 s for baseline. For querying, each query took on an average less than 0.4 s on blockchain, versus around 2.1 s for baseline. DISCUSSION: The low deployment, recording, and querying times confirm the feasibility of our cross-cloud, blockchain-based federated data analysis system. We have yet to evaluate the system on a larger network with multiple nodes per cloud, to consider how to accommodate a surge in activities, and to investigate methods to lower querying time as the blockchain grows. CONCLUSION: Blockchain technology can be used to support federated data analysis among multiple institutions.
Tsung-Ting Kuo, Anh Pham, Maxim E. Edelson, Jihoon Kim 0001, Yash Gupta, Lucila Ohno-Machado, David M. Anderson, Chandrasekar Balacha, Tyler Bath, Sally L. Baxter, Andrea Becker-Pennrich, Douglas S. Bell, Elmer V. Bernstam, Ngan Chau, Michele E. Day, Jason N. Doctor, Scott L. DuVall, Robert El-Kareh, Renato Florian, Robert W. Follett, Benjamin P. Geisler, Alessandro Ghigi, Assaf Gottlieb, Christian Hinske, Zhaoxian Hu, Diana Ir, Xiaoqian Jiang, Katherine K. Kim, Tara K. Knight, Jejo Koola, Ulrich Mansmann, Michael E. Matheny, Daniella Meeker, Zongyang Mou, Larissa Neumann, Nghia H. Nguyen, Nicholas R. Anderson 0001, Eunice Park, Paulina Paul, Mark J. Pletcher, Kai W. Post, Clemens Rieder, Clemens Scherer, Lisa M. Schilling, Andrey Soares, Spencer L. SooHoo, Ekin Soysal, Steven Covington, Brian Tep, Brian Toy, Baocheng Wang, Zhen R. Wu, Hua Xu 0001, Yong K. Choi, Kai Zheng 0002, Yujia Zhou 0003, Rachel A Zucker
J. Am. Medical Informatics Assoc.4
2023 WICOX: Weight-Based Integrated Cox Model for Time-to-Event Data in Distributed Databases Without Data-Sharing
abstract
To exploit large-scale biomedical data, the application of common data models and the establishment of data networks are being actively carried out worldwide. However, due to the privacy issues, it is difficult to share data distributed among institutions. In this study, we developed and evaluated weight-based integrated Cox model (WICOX) as a privacy-protecting method without sharing patient-level information across institutions. WICOX generates a weight for each institutional model and builds an integrated model of multi-institutional data based on these weights. WICOX does not require iterative communication until the centralized parameter converges. We performed experiments to show the weight characteristic of our algorithm based on 10 hospitals (2910 intensive care unit (ICU) stays in total) from the electronic intensive care unit Collaborative Research Database to predict time to ICU mortality with eight risk factors. Compared with the centralized Cox model, WICOX showed biases from 0 to 0.68E-2, from 0.00E-2 to 4.98E-2, and from 0.74E-2 to 1.7E-2 for time-dependent AUC, log hazard ratio, and survival rate, respectively. In addition, through simulation results using real 10 hospitals, WICOX showed robust results in accuracy under any composition of hospitals. The results of the experiments highlight that WICOX has robust characteristics and provides predictive performance and statistical inference results nearly the same as those of the centralized model. WICOX is a non-iterative method using the weight of institutional model for implementing the Cox model across multiple institutions in a privacy-preserving manner.
Ji A. Park, Tae H. Kim, Jihoon Kim 0001, Yu Rang Park
IEEE J. Biomed. Health Informatics3
2022 The evolving privacy and security concerns for genomic data analysis and sharing as observed from the iDASH competition
abstract
Concerns regarding inappropriate leakage of sensitive personal information as well as unauthorized data use are increasing with the growth of genomic data repositories. Therefore, privacy and security of genomic data have become increasingly important and need to be studied. With many proposed protection techniques, their applicability in support of biomedical research should be well understood. For this purpose, we have organized a community effort in the past 8 years through the integrating data for analysis, anonymization and sharing consortium to address this practical challenge. In this article, we summarize our experience from these competitions, report lessons learned from the events in 2020/2021 as examples, and discuss potential future research directions in this emerging field.
Tsung-Ting Kuo, Xiaoqian Jiang, Haixu Tang, XiaoFeng Wang 0001, Arif Ozgun Harmanci, Miran Kim, Kai W. Post, Diyue Bu, Tyler Bath, Jihoon Kim 0001, Weijie Liu 0004, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.10
2021 Privacy-protecting, reliable response data discovery using COVID-19 patient observations
abstract
OBJECTIVE: To utilize, in an individual and institutional privacy-preserving manner, electronic health record (EHR) data from 202 hospitals by analyzing answers to COVID-19-related questions and posting these answers online. MATERIALS AND METHODS: We developed a distributed, federated network of 12 health systems that harmonized their EHRs and submitted aggregate answers to consortia questions posted at https://www.covid19questions.org. Our consortium developed processes and implemented distributed algorithms to produce answers to a variety of questions. We were able to generate counts, descriptive statistics, and build a multivariate, iterative regression model without centralizing individual-level data. RESULTS: Our public website contains answers to various clinical questions, a web form for users to ask questions in natural language, and a list of items that are currently pending responses. The results show, for example, that patients who were taking angiotensin-converting enzyme inhibitors and angiotensin II receptor blockers, within the year before admission, had lower unadjusted in-hospital mortality rates. We also showed that, when adjusted for, age, sex, and ethnicity were not significantly associated with mortality. We demonstrated that it is possible to answer questions about COVID-19 using EHR data from systems that have different policies and must follow various regulations, without moving data out of their health systems. DISCUSSION AND CONCLUSIONS: We present an alternative or a complement to centralized COVID-19 registries of EHR data. We can use multivariate distributed logistic regression on observations recorded in the process of care to generate results without transferring individual-level data outside the health systems.
Jihoon Kim 0001, Larissa Neumann, Paulina Paul, Michele E. Day, Michael Aratow, Douglas S. Bell, Jason N. Doctor, Christian Hinske, Xiaoqian Jiang, Katherine K. Kim, Michael E. Matheny, Daniella Meeker, Mark J. Pletcher, Lisa M. Schilling, Spencer L. SooHoo, Hua Xu 0001, Kai Zheng 0002, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.1
2020 Longitudinal Cohort Data Transformation Based on a Common Data Model and Metadata Standards: Examples from the Strong Heart Study
Jihoon Kim 0001, Paulina Paul, Pravina Kota, Yu R. Park, Julie A. Stoner, Elisa Lee, Lucila Ohno-Machado
AMIA1
2020 Privacy-preserving model learning on a blockchain network-of-networks
abstract
OBJECTIVE: To facilitate clinical/genomic/biomedical research, constructing generalizable predictive models using cross-institutional methods while protecting privacy is imperative. However, state-of-the-art methods assume a "flattened" topology, while real-world research networks may consist of "network-of-networks" which can imply practical issues including training on small data for rare diseases/conditions, prioritizing locally trained models, and maintaining models for each level of the hierarchy. In this study, we focus on developing a hierarchical approach to inherit the benefits of the privacy-preserving methods, retain the advantages of adopting blockchain, and address practical concerns on a research network-of-networks. MATERIALS AND METHODS: We propose a framework to combine level-wise model learning, blockchain-based model dissemination, and a novel hierarchical consensus algorithm for model ensemble. We developed an example implementation HierarchicalChain (hierarchical privacy-preserving modeling on blockchain), evaluated it on 3 healthcare/genomic datasets, as well as compared its predictive correctness, learning iteration, and execution time with a state-of-the-art method designed for flattened network topology. RESULTS: HierarchicalChain improves the predictive correctness for small training datasets and provides comparable correctness results with the competing method with higher learning iteration and similar per-iteration execution time, inherits the benefits of the privacy-preserving learning and advantages of blockchain technology, and immutable records models for each level. DISCUSSION: HierarchicalChain is independent of the core privacy-preserving learning method, as well as of the underlying blockchain platform. Further studies are warranted for various types of network topology, complex data, and privacy concerns. CONCLUSION: We demonstrated the potential of utilizing the information from the hierarchical network-of-networks topology to improve prediction.
Tsung-Ting Kuo, Jihoon Kim 0001, Rodney A. Gabriel
J. Am. Medical Informatics Assoc.2
2019 Evaluating and sharing global genetic ancestry in biomedical datasets
abstract
Genetic ancestry is a critical co-factor to study phenotype-genotype associations using cohorts of human subjects. Most publicly available molecular datasets are, however, missing this information or only share self-reported race and ethnicity, representing a limitation to identify and repurpose datasets to investigate the contribution of ancestry to diseases and traits. We propose an analytical framework to enrich the metadata from publicly available cohorts with genetic ancestry information and a resulting diversity score at continental resolution, calculated directly from the data. We illustrate this framework using The Cancer Genome Atlas datasets searched through the DataMed Data Discovery Index. Data repositories and contributors can use this framework to provide genetic diversity measurements for controlled access datasets, minimizing the work involved in requesting a dataset that may ultimately prove inadequate for a researcher's purpose. With the increasing global scale of human genetics research, studies on disease risk and susceptibility would benefit greatly from the adequate estimation and sharing of genetic diversity in publicly available datasets following a framework such as the one presented.
Olivier Harismendy, Jihoon Kim 0001, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.2
2017 Project A-G-T-C: Adding Genetic Test data in Clinical data warehouse
Jonathan Wickes, Hyeon-Eui Kim, Hyun-Dae Kim, Jihoon Kim 0001, Lucila Ohno-Machado, Olivier Harismendy
AMIA4
2017 PRINCESS: Privacy-protecting Rare disease International Network Collaboration via Encryption through Software guard extensionS
abstract
Motivation: We introduce PRINCESS, a privacy-preserving international collaboration framework for analyzing rare disease genetic data that are distributed across different continents. PRINCESS leverages Software Guard Extensions (SGX) and hardware for trustworthy computation. Unlike a traditional international collaboration model, where individual-level patient DNA are physically centralized at a single site, PRINCESS performs a secure and distributed computation over encrypted data, fulfilling institutional policies and regulations for protected health information. Results: To demonstrate PRINCESS' performance and feasibility, we conducted a family-based allelic association study for Kawasaki Disease, with data hosted in three different continents. The experimental results show that PRINCESS provides secure and accurate analyses much faster than alternative solutions, such as homomorphic encryption and garbled circuits (over 40 000× faster). Availability and Implementation: https://github.com/achenfengb/PRINCESS_opensource. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Feng Chen 0016, Shuang Wang 0002, Xiaoqian Jiang, Sijie Ding, Yao Lu 0006, Jihoon Kim 0001, Süleyman Cenk Sahinalp, Chisato Shimizu, Jane C. Burns, Victoria J. Wright, Eileen Png, Martin L. Hibberd, David D. Lloyd, Amalio Telenti, Cinnamon S. Bloss, Dov Fox, Kristin E. Lauter, Lucila Ohno-Machado
Bioinform.6
2017 Mechanisms to protect the privacy of families when using the transmission disequilibrium test in genome-wide association studies
abstract
MOTIVATION: Inappropriate disclosure of human genomes may put the privacy of study subjects and of their family members at risk. Existing privacy-preserving mechanisms for Genome-Wide Association Studies (GWAS) mainly focus on protecting individual information in case-control studies. Protecting privacy in family-based studies is more difficult. The transmission disequilibrium test (TDT) is a powerful family-based association test employed in many rare disease studies. It gathers information about families (most frequently involving parents, affected children and their siblings). It is important to develop privacy-preserving approaches to disclose TDT statistics with a guarantee that the risk of family 're-identification' stays below a pre-specified risk threshold. 'Re-identification' in this context means that an attacker can infer that the presence of a family in a study. METHODS: In the context of protecting family-level privacy, we developed and evaluated a suite of differentially private (DP) mechanisms for TDT. They include Laplace mechanisms based on the TDT test statistic, P-values, projected P-values and exponential mechanisms based on the TDT test statistic and the shortest Hamming distance (SHD) score. RESULTS: Using simulation studies with a small cohort and a large one, we showed that that the exponential mechanism based on the SHD score preserves the highest utility and privacy among all proposed DP methods. We provide a guideline on applying our DP TDT in a real dataset in analyzing Kawasaki disease with 187 families and 906 SNPs. There are some limitations, including: (1) the performance of our implementation is slow for real-time results generation and (2) handling missing data is still challenging. AVAILABILITY AND IMPLEMENTATION: The software dpTDT is available in https://github.com/mwgrassgreen/dpTDT. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Meng Wang 0007, Zhanglong Ji, Shuang Wang 0002, Jihoon Kim 0001, Xiaoqian Jiang, Lucila Ohno-Machado
Bioinform.4
2017 iCONCUR: informed consent for clinical data and bio-sample use for research
abstract
Background: Implementation of patient preferences for use of electronic health records for research has been traditionally limited to identifiable data. Tiered e-consent for use of de-identified data has traditionally been deemed unnecessary or impractical for implementation in clinical settings. Methods: We developed a web-based tiered informed consent tool called informed consent for clinical data and bio-sample use for research (iCONCUR) that honors granular patient preferences for use of electronic health record data in research. We piloted this tool in 4 outpatient clinics of an academic medical center. Results: Of patients offered access to iCONCUR, 394 agreed to participate in this study, among whom 126 patients accessed the website to modify their records according to data category and data recipient. The majority consented to share most of their data and specimens with researchers. Willingness to share was greater among participants from an Human Immunodeficiency Virus (HIV) clinic than those from internal medicine clinics. The number of items declined was higher for for-profit institution recipients. Overall, participants were most willing to share demographics and body measurements and least willing to share family history and financial data. Participants indicated that having granular choices for data sharing was appropriate, and that they liked being informed about who was using their data for what purposes, as well as about outcomes of the research. Conclusion: This study suggests that a tiered electronic informed consent system is a workable solution that respects patient preferences, increases satisfaction, and does not significantly affect participation in research.
Hyeon-Eui Kim, Elizabeth A. Bell, Jihoon Kim 0001, Amy M. Sitapati, Joe Ramsdell, Claudiu Farcas, Dexter Friedman, Stephanie Feudjio Feupe, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.3
2016 Developing a predictive model for discharge delay in the Post-Anesthesia Care Unit
Rodney A. Gabriel, Jihoon Kim 0001, Lucila Ohno-Machado
AMIA2
2016 Does Health Status Affect Patient Preferences for Sharing Clinical Data for Research?
Imho Jang, Diana Guijarro, Jimmy Quach, Jihoon Kim 0001, Hyeon-Eui Kim, Elizabeth A. Bell, Robert El-Kareh, Lucila Ohno-Machado
AMIA4
2016 Personas in online health communities
Jina Huh, Bum Chul Kwon, Sung-Hee Kim, Sukwon Lee, Jaegul Choo, Jihoon Kim 0001, Minje Choi, Ji Soo Yi
J. Biomed. Informatics6
2014 MAGI: a Node.js web service for fast microRNA-Seq analysis in a GPU infrastructure
abstract
SUMMARY: MAGI is a web service for fast MicroRNA-Seq data analysis in a graphics processing unit (GPU) infrastructure. Using just a browser, users have access to results as web reports in just a few hours->600% end-to-end performance improvement over state of the art. MAGI's salient features are (i) transfer of large input files in native FASTA with Qualities (FASTQ) format through drag-and-drop operations, (ii) rapid prediction of microRNA target genes leveraging parallel computing with GPU devices, (iii) all-in-one analytics with novel feature extraction, statistical test for differential expression and diagnostic plot generation for quality control and (iv) interactive visualization and exploration of results in web reports that are readily available for publication. AVAILABILITY AND IMPLEMENTATION: MAGI relies on the Node.js JavaScript framework, along with NVIDIA CUDA C, PHP: Hypertext Preprocessor (PHP), Perl and R. It is freely available at http://magi.ucsd.edu.
Jihoon Kim 0001, Eric Levy, Alex Ferbrache, Petra Stepanowsky, Claudiu Farcas, Shuang Wang 0002, Stefan Brunner, Tyler Bath, Yuan Wu 0003, Lucila Ohno-Machado
Bioinform.1
2014 HUGO: Hierarchical mUlti-reference Genome cOmpression for aligned reads
abstract
BACKGROUND AND OBJECTIVE: Short-read sequencing is becoming the standard of practice for the study of structural variants associated with disease. However, with the growth of sequence data largely surpassing reasonable storage capability, the biomedical community is challenged with the management, transfer, archiving, and storage of sequence data. METHODS: We developed Hierarchical mUlti-reference Genome cOmpression (HUGO), a novel compression algorithm for aligned reads in the sorted Sequence Alignment/Map (SAM) format. We first aligned short reads against a reference genome and stored exactly mapped reads for compression. For the inexact mapped or unmapped reads, we realigned them against different reference genomes using an adaptive scheme by gradually shortening the read length. Regarding the base quality value, we offer lossy and lossless compression mechanisms. The lossy compression mechanism for the base quality values uses k-means clustering, where a user can adjust the balance between decompression quality and compression rate. The lossless compression can be produced by setting k (the number of clusters) to the number of different quality values. RESULTS: The proposed method produced a compression ratio in the range 0.5-0.65, which corresponds to 35-50% storage savings based on experimental datasets. The proposed approach achieved 15% more storage savings over CRAM and comparable compression ratio with Samcomp (CRAM and Samcomp are two of the state-of-the-art genome compression algorithms). The software is freely available at https://sourceforge.net/projects/hierachicaldnac/with a General Public License (GPL) license. LIMITATION: Our method requires having different reference genomes and prolongs the execution time for additional alignments. CONCLUSIONS: The proposed multi-reference-based compression algorithm for aligned reads outperforms existing single-reference based algorithms.
Pinghao Li, Xiaoqian Jiang, Shuang Wang 0002, Jihoon Kim 0001, Hongkai Xiong, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.4
2014 Detecting inappropriate access to electronic health records using collaborative filtering
Aditya Krishna Menon, Xiaoqian Jiang, Jihoon Kim 0001, Jaideep Vaidya, Lucila Ohno-Machado
Mach. Learn.3
2012 Selecting Cases for Whom Additional Tests Can Improve Prognostication
Xiaoqian Jiang, Jihoon Kim 0001, Yuan Wu 0003, Lucila Ohno-Machado
AMIA2
2012 A patient-driven adaptive prediction technique to improve personalized risk estimation for clinical decision support
abstract
OBJECTIVE: Competing tools are available online to assess the risk of developing certain conditions of interest, such as cardiovascular disease. While predictive models have been developed and validated on data from cohort studies, little attention has been paid to ensure the reliability of such predictions for individuals, which is critical for care decisions. The goal was to develop a patient-driven adaptive prediction technique to improve personalized risk estimation for clinical decision support. MATERIAL AND METHODS: A data-driven approach was proposed that utilizes individualized confidence intervals (CIs) to select the most 'appropriate' model from a pool of candidates to assess the individual patient's clinical condition. The method does not require access to the training dataset. This approach was compared with other strategies: the BEST model (the ideal model, which can only be achieved by access to data or knowledge of which population is most similar to the individual), CROSS model, and RANDOM model selection. RESULTS: When evaluated on clinical datasets, the approach significantly outperformed the CROSS model selection strategy in terms of discrimination (p<1e-14) and calibration (p<0.006). The method outperformed the RANDOM model selection strategy in terms of discrimination (p<1e-12), but the improvement did not achieve significance for calibration (p=0.1375). LIMITATIONS: The CI may not always offer enough information to rank the reliability of predictions, and this evaluation was done using aggregation. If a particular individual is very different from those represented in a training set of existing models, the CI may be somewhat misleading. CONCLUSION: This approach has the potential to offer more reliable predictions than those offered by other heuristics for disease risk estimation of individual patients.
Xiaoqian Jiang, Aziz A. Boxwala, Robert El-Kareh, Jihoon Kim 0001, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.4
2012 Calibrating predictive model estimates to support personalized medicine
abstract
OBJECTIVE: Predictive models that generate individualized estimates for medically relevant outcomes are playing increasing roles in clinical care and translational research. However, current methods for calibrating these estimates lose valuable information. Our goal is to develop a new calibration method to conserve as much information as possible, and would compare favorably to existing methods in terms of important performance measures: discrimination and calibration. MATERIAL AND METHODS: We propose an adaptive technique that utilizes individualized confidence intervals (CIs) to calibrate predictions. We evaluate this new method, adaptive calibration of predictions (ACP), in artificial and real-world medical classification problems, in terms of areas under the ROC curves, the Hosmer-Lemeshow goodness-of-fit test, mean squared error, and computational complexity. RESULTS: ACP compared favorably to other calibration methods such as binning, Platt scaling, and isotonic regression. In several experiments, binning, isotonic regression, and Platt scaling failed to improve the calibration of a logistic regression model, whereas ACP consistently improved the calibration while maintaining the same discrimination or even improving it in some experiments. In addition, the ACP algorithm is not computationally expensive. LIMITATIONS: The calculation of CIs for individual predictions may be cumbersome for certain predictive models. ACP is not completely parameter-free: the length of the CI employed may affect its results. CONCLUSIONS: ACP can generate estimates that may be more suitable for individualized predictions than estimates that are calibrated using existing methods. Further studies are necessary to explore the limitations of ACP.
Xiaoqian Jiang, Melanie Osl, Jihoon Kim 0001, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.3
2012 iDASH: integrating data for analysis, anonymization, and sharing
abstract
iDASH (integrating data for analysis, anonymization, and sharing) is the newest National Center for Biomedical Computing funded by the NIH. It focuses on algorithms and tools for sharing data in a privacy-preserving manner. Foundational privacy technology research performed within iDASH is coupled with innovative engineering for collaborative tool development and data-sharing capabilities in a private Health Insurance Portability and Accountability Act (HIPAA)-certified cloud. Driving Biological Projects, which span different biological levels (from molecules to individuals to populations) and focus on various health conditions, help guide research and development within this Center. Furthermore, training and dissemination efforts connect the Center with its stakeholders and educate data owners and data consumers on how to share and use clinical and biological data. Through these various mechanisms, iDASH implements its goal of providing biomedical and behavioral researchers with access to data, software, and a high-performance computing environment, thus enabling them to generate and test new hypotheses.
Lucila Ohno-Machado, Vineet Bafna, Aziz A. Boxwala, Brian E. Chapman, Wendy W. Chapman, Kamalika Chaudhuri, Michele E. Day, Claudiu Farcas, Nathaniel D. Heintzman, Xiaoqian Jiang, Hyeon-Eui Kim, Jihoon Kim 0001, Michael E. Matheny, Frederic S. Resnic, Staal Amund Vinterbo
J. Am. Medical Informatics Assoc.12
2012 Grid Binary LOgistic REgression (GLORE): building shared models without sharing data
abstract
OBJECTIVE: The classification of complex or rare patterns in clinical and genomic data requires the availability of a large, labeled patient set. While methods that operate on large, centralized data sources have been extensively used, little attention has been paid to understanding whether models such as binary logistic regression (LR) can be developed in a distributed manner, allowing researchers to share models without necessarily sharing patient data. MATERIAL AND METHODS: Instead of bringing data to a central repository for computation, we bring computation to the data. The Grid Binary LOgistic REgression (GLORE) model integrates decomposable partial elements or non-privacy sensitive prediction values to obtain model coefficients, the variance-covariance matrix, the goodness-of-fit test statistic, and the area under the receiver operating characteristic (ROC) curve. RESULTS: We conducted experiments on both simulated and clinically relevant data, and compared the computational costs of GLORE with those of a traditional LR model estimated using the combined data. We showed that our results are the same as those of LR to a 10(-15) precision. In addition, GLORE is computationally efficient. LIMITATION: In GLORE, the calculation of coefficient gradients must be synchronized at different sites, which involves some effort to ensure the integrity of communication. Ensuring that the predictors have the same format and meaning across the data sets is necessary. CONCLUSION: The results suggest that GLORE performs as well as LR and allows data to remain protected at their original sites.
Yuan Wu 0003, Xiaoqian Jiang, Jihoon Kim 0001, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.3
2011 AnyExpress: Integrated toolkit for analysis of cross-platform gene expression data using a fast interval matching algorithm
abstract
BACKGROUND: Cross-platform analysis of gene express data requires multiple, intricate processes at different layers with various platforms. However, existing tools handle only a single platform and are not flexible enough to support custom changes, which arise from the new statistical methods, updated versions of reference data, and better platforms released every month or year. Current tools are so tightly coupled with reference information, such as reference genome, transcriptome database, and SNP, which are often erroneous or outdated, that the output results are incorrect and misleading. RESULTS: We developed AnyExpress, a software package that combines cross-platform gene expression data using a fast interval-matching algorithm. Supported platforms include next-generation-sequencing technology, microarray, SAGE, MPSS, and more. Users can define custom target transcriptome database references for probe/read mapping in any species, as well as criteria to remove undesirable probes/reads. AnyExpress offers scalable processing features such as binding, normalization, and summarization that are not present in existing software tools. As a case study, we applied AnyExpress to published Affymetrix microarray and Illumina NGS RNA-Seq data from human kidney and liver. The mean of within-platform correlation coefficient was 0.98 for within-platform samples in kidney and liver, respectively. The mean of cross-platform correlation coefficients was 0.73. These results confirmed those of the original and secondary studies. Applying filtering produced higher agreement between microarray and NGS, according to an agreement index calculated from differentially expressed genes. CONCLUSION: AnyExpress can combine cross-platform gene expression data, process data from both open- and closed-platforms, select a custom target reference, filter out undesirable probes or reads based on custom-defined biological features, and perform quantile-normalization with a large number of microarray samples. AnyExpress is fast, comprehensive, flexible, and freely available at http://anyexpress.sourceforge.net.
Jihoon Kim 0001, Kiltesh Patel, Hyunchul Jung, Winston Patrick Kuo, Lucila Ohno-Machado
BMC Bioinform.1
2011 Full impact of laboratory information system requires direct use by clinical staff: cluster randomized controlled trial
abstract
OBJECTIVE: To evaluate the time to communicate laboratory results to health centers (HCs) between the e-Chasqui web-based information system and the pre-existing paper-based system. METHODS: Cluster randomized controlled trial in 78 HCs in Peru. In the intervention group, 12 HCs had web access to results via e-Chasqui (point-of-care HCs) and forwarded results to 17 peripheral HCs. In the control group, 22 point-of-care HCs received paper results directly and forwarded them to 27 peripheral HCs. Baseline data were collected for 15 months. Post-randomization data were collected for at least 2 years. Comparisons were made between intervention and control groups, stratified by point-of-care versus peripheral HCs. RESULTS: For point-of-care HCs, the intervention group took less time to receive drug susceptibility tests (DSTs) (median 9 vs 16 days, p<0.001) and culture results (4 vs 8 days, p<0.001) and had a lower proportion of 'late' DSTs taking >60 days to arrive (p<0.001) than the control. For peripheral HCs, the intervention group had similar communication times for DST (median 22 vs 19 days, p=0.30) and culture (10 vs 9 days, p=0.10) results, as well as proportion of 'late' DSTs (p=0.57) compared with the control. CONCLUSIONS: Only point-of-care HCs with direct access to the e-Chasqui information system had reduced communication times and fewer results with delays of >2 months. Peripheral HCs had no benefits from the system. This suggests that health establishments should have point-of-care access to reap the benefits of electronic laboratory reporting.
Joaquin A. Blaya, Sonya S. Shin, Carmen Contreras, Gloria Yale, Carmen Suares, Luis Asencios, Jihoon Kim 0001, Pablo Rodriguez 0003, Peter Cegielski, Hamish S. F. Fraser
J. Am. Medical Informatics Assoc.7
2011 Using statistical and machine learning to help institutions detect suspicious access to electronic health records
abstract
OBJECTIVE: To determine whether statistical and machine-learning methods, when applied to electronic health record (EHR) access data, could help identify suspicious (ie, potentially inappropriate) access to EHRs. METHODS: From EHR access logs and other organizational data collected over a 2-month period, the authors extracted 26 features likely to be useful in detecting suspicious accesses. Selected events were marked as either suspicious or appropriate by privacy officers, and served as the gold standard set for model evaluation. The authors trained logistic regression (LR) and support vector machine (SVM) models on 10-fold cross-validation sets of 1291 labeled events. The authors evaluated the sensitivity of final models on an external set of 58 events that were identified as truly inappropriate and investigated independently from this study using standard operating procedures. RESULTS: The area under the receiver operating characteristic curve of the models on the whole data set of 1291 events was 0.91 for LR, and 0.95 for SVM. The sensitivity of the baseline model on this set was 0.8. When the final models were evaluated on the set of 58 investigated events, all of which were determined as truly inappropriate, the sensitivity was 0 for the baseline method, 0.76 for LR, and 0.79 for SVM. LIMITATIONS: The LR and SVM models may not generalize because of interinstitutional differences in organizational structures, applications, and workflows. Nevertheless, our approach for constructing the models using statistical and machine-learning techniques can be generalized. An important limitation is the relatively small sample used for the training set due to the effort required for its construction. CONCLUSION: The results suggest that statistical and machine-learning methods can play an important role in helping privacy officers detect suspicious accesses to EHRs.
Aziz A. Boxwala, Jihoon Kim 0001, Janice M. Grillo, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.2
2011 A multi-layered framework for disseminating knowledge for computer-based decision support
abstract
BACKGROUND: There are several challenges in encoding guideline knowledge in a form that is portable to different clinical sites, including the heterogeneity of clinical decision support (CDS) tools, of patient data representations, and of workflows. METHODS: We have developed a multi-layered knowledge representation framework for structuring guideline recommendations for implementation in a variety of CDS contexts. In this framework, guideline recommendations are increasingly structured through four layers, successively transforming a narrative text recommendation into input for a CDS system. We have used this framework to implement rules for a CDS service based on three guidelines. We also conducted a preliminary evaluation, where we asked CDS experts at four institutions to rate the implementability of six recommendations from the three guidelines. CONCLUSION: The experience in using the framework and the preliminary evaluation indicate that this approach has promise in creating structured knowledge, to implement in CDS systems, that is usable across organizations.
Aziz A. Boxwala, Beatriz H. S. C. Rocha, Saverio M. Maviglia, Vipul Kashyap, Seth Meltzer, Jihoon Kim 0001, Ruslana Tsurikova, Adam Wright, Marilyn D. Paterno, Amanda Fairbanks, Blackford Middleton
J. Am. Medical Informatics Assoc.6
2011 Trends in biomedical informatics: most cited topics from recent years
abstract
Biomedical informatics is a young, highly interdisciplinary field that is evolving quickly. It is important to know which published topics in generalist biomedical informatics journals elicit the most interest from the scientific community, and whether this interest changes over time, so that journals can better serve their readers. It is also important to understand whether free access to biomedical informatics articles impacts their citation rates in a significant way, so authors can make informed decisions about unlock fees, and journal owners and publishers understand the implications of open access. The topics and JAMIA articles from years 2009 and 2010 that have been most cited according to the Web of Science are described. To better understand the effects of free access in article dissemination, the number of citations per month after publication for articles published in 2009 versus 2010 was compared, since there was a significant change in free access to JAMIA articles between those years. Results suggest that there is a positive association between free access and citation rate for JAMIA articles.
Hyeon-Eui Kim, Xiaoqian Jiang, Jihoon Kim 0001, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.3
2010 DSGeo: Software tools for cross-platform analysis of gene expression data in GEO
Ronilda C. Lacson, Erik Pitzer, Jihoon Kim 0001, Pedro A. F. Galante, Christian Hinske, Lucila Ohno-Machado
J. Biomed. Informatics3
2009 Towards large-scale sample annotation in gene expression repositories
abstract
BACKGROUND: Large repositories of biomedical research data are most useful to translational researchers if their data can be aggregated for efficient queries and analyses. However, inconsistent or non-existent annotations describing important sample details such as name of tissue or cell line, histopathological type, and subject characteristics like demographics, treatment, and survival are seldom present in data repositories, making it difficult to aggregate data. RESULTS: We created a flexible software tool that allows efficient annotation of samples using a controlled vocabulary, and report on its use for the annotation of over 12,500 samples. CONCLUSION: While the amount of data is very large and seemingly poorly annotated, a lot of information is still within reach. Consistent tool-based re-annotation enables many new possibilities for large scale interpretation and analyses that would otherwise be impossible.
Erik Pitzer, Ronilda C. Lacson, Christian Hinske, Jihoon Kim 0001, Pedro A. F. Galante, Lucila Ohno-Machado
BMC Bioinform.4
2007 Difference-based clustering of short time-course microarray data with replicates
abstract
BACKGROUND: There are some limitations associated with conventional clustering methods for short time-course gene expression data. The current algorithms require prior domain knowledge and do not incorporate information from replicates. Moreover, the results are not always easy to interpret biologically. RESULTS: We propose a novel algorithm for identifying a subset of genes sharing a significant temporal expression pattern when replicates are used. Our algorithm requires no prior knowledge, instead relying on an observed statistic which is based on the first and second order differences between adjacent time-points. Here, a pattern is predefined as the sequence of symbols indicating direction and the rate of change between time-points, and each gene is assigned to a cluster whose members share a similar pattern. We evaluated the performance of our algorithm to those of K-means, Self-Organizing Map and the Short Time-series Expression Miner methods. CONCLUSIONS: Assessments using simulated and real data show that our method outperformed aforementioned algorithms. Our approach is an appropriate solution for clustering short time-course microarray data with replicates.
Jihoon Kim 0001, Ju Han Kim
BMC Bioinform.1
2004 ChromoViz: multimodal visualization of gene expression data onto chromosomes using scalable vector graphics
abstract
SUMMARY: ChromoViz is an R package for the visualization of microarray gene expression data, cross-species and cross-platform comparisons, as well as non-expression genomic data obtained from public databases onto chromosomes. Chromosomal visualization format is proposed for the clear decoupling of the data layer from the procedure layer and the combined visualization of genomic data from heterogeneous data sources. Visualization with Javascript-enabled scalable vector graphics enables interactive visualization and navigation of data objects on the Web. AVAILABILITY: http://www.snubi.org/software/ChromoViz/
Jihoon Kim 0001, Hee-Joon Chung, Chan Hee Park, Woong-Yang Park, Ju Han Kim
Bioinform.1