EDBT 2026 Demo / reviewers in the wild / expert
Nansu Zong
dblp:44/7805
· DBLP profile ↗
20ranked-venue papers
5as first author
14since 2021 · last 2024
0000-0003-0066-9524ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Stratifying heart failure patients with graph neural network and transformer using Electronic Health Records to optimize drug response predictionabstractOBJECTIVES: Heart failure (HF) impacts millions of patients worldwide, yet the variability in treatment responses remains a major challenge for healthcare professionals. The current treatment strategies, largely derived from population based evidence, often fail to consider the unique characteristics of individual patients, resulting in suboptimal outcomes. This study aims to develop computational models that are patient-specific in predicting treatment outcomes, by utilizing a large Electronic Health Records (EHR) database. The goal is to improve drug response predictions by identifying specific HF patient subgroups that are likely to benefit from existing HF medications. MATERIALS AND METHODS: A novel, graph-based model capable of predicting treatment responses, combining Graph Neural Network and Transformer was developed. This method differs from conventional approaches by transforming a patient's EHR data into a graph structure. By defining patient subgroups based on this representation via K-Means Clustering, we were able to enhance the performance of drug response predictions. RESULTS: Leveraging EHR data from 11 627 Mayo Clinic HF patients, our model significantly outperformed traditional models in predicting drug response using NT-proBNP as a HF biomarker across five medication categories (best RMSE of 0.0043). Four distinct patient subgroups were identified with differential characteristics and outcomes, demonstrating superior predictive capabilities over existing HF subtypes (best mean RMSE of 0.0032). DISCUSSION: These results highlight the power of graph-based modeling of EHR in improving HF treatment strategies. The stratification of patients sheds light on particular patient segments that could benefit more significantly from tailored response predictions. CONCLUSIONS: Longitudinal EHR data have the potential to enhance personalized prognostic predictions through the application of graph-based AI techniques. Shaika Chowdhury, Yongbin Chen, Pengyang Li, Sivaraman Rajaganapathy, Andrew Wen, Xiao Ma 0019, Qiying Dai, Yue Yu 0012, Sunyang Fu, Xiaoqian Jiang, Zhe He 0001, Sunghwan Sohn, Xiaoke Liu, Suzette J. Bielinski, Alanna M. Chamberlain, James R. Cerhan, Nansu Zong |
J. Am. Medical Informatics Assoc. | 17 |
| 2024 | A taxonomy for advancing systematic error analysis in multi-site electronic health record-based clinical concept extractionabstractBACKGROUND: Error analysis plays a crucial role in clinical concept extraction, a fundamental subtask within clinical natural language processing (NLP). The process typically involves a manual review of error types, such as contextual and linguistic factors contributing to their occurrence, and the identification of underlying causes to refine the NLP model and improve its performance. Conducting error analysis can be complex, requiring a combination of NLP expertise and domain-specific knowledge. Due to the high heterogeneity of electronic health record (EHR) settings across different institutions, challenges may arise when attempting to standardize and reproduce the error analysis process. OBJECTIVES: This study aims to facilitate a collaborative effort to establish common definitions and taxonomies for capturing diverse error types, fostering community consensus on error analysis for clinical concept extraction tasks. MATERIALS AND METHODS: We iteratively developed and evaluated an error taxonomy based on existing literature, standards, real-world data, multisite case evaluations, and community feedback. The finalized taxonomy was released in both .dtd and .owl formats at the Open Health Natural Language Processing Consortium. The taxonomy is compatible with several different open-source annotation tools, including MAE, Brat, and MedTator. RESULTS: The resulting error taxonomy comprises 43 distinct error classes, organized into 6 error dimensions and 4 properties, including model type (symbolic and statistical machine learning), evaluation subject (model and human), evaluation level (patient, document, sentence, and concept), and annotation examples. Internal and external evaluations revealed strong variations in error types across methodological approaches, tasks, and EHR settings. Key points emerged from community feedback, including the need to enhancing clarity, generalizability, and usability of the taxonomy, along with dissemination strategies. CONCLUSION: The proposed taxonomy can facilitate the acceleration and standardization of the error analysis process in multi-site settings, thus improving the provenance, interpretability, and portability of NLP models. Future researchers could explore the potential direction of developing automated or semi-automated methods to assist in the classification and standardization of error analysis. Sunyang Fu, Liwei Wang 0010, Andrew Wen, Nansu Zong, Anamika Kumari, Rui Zhang 0028, Yanshan Wang, Jennifer L. St. Sauver, Sunghwan Sohn |
J. Am. Medical Informatics Assoc. | 5 |
| 2023 | Using artificial intelligence to learn optimal regimen plan for Alzheimer's diseaseabstractBACKGROUND: Alzheimer's disease (AD) is a progressive neurological disorder with no specific curative medications. Sophisticated clinical skills are crucial to optimize treatment regimens given the multiple coexisting comorbidities in the patient population. OBJECTIVE: Here, we propose a study to leverage reinforcement learning (RL) to learn the clinicians' decisions for AD patients based on the longitude data from electronic health records. METHODS: In this study, we selected 1736 patients from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database. We focused on the two most frequent concomitant diseases-depression, and hypertension, thus creating 5 data cohorts (ie, Whole Data, AD, AD-Hypertension, AD-Depression, and AD-Depression-Hypertension). We modeled the treatment learning into an RL problem by defining states, actions, and rewards. We built a regression model and decision tree to generate multiple states, used six combinations of medications (ie, cholinesterase inhibitors, memantine, memantine-cholinesterase inhibitors, hypertension drugs, supplements, or no drugs) as actions, and Mini-Mental State Exam (MMSE) scores as rewards. RESULTS: Given the proper dataset, the RL model can generate an optimal policy (regimen plan) that outperforms the clinician's treatment regimen. Optimal policies (ie, policy iteration and Q-learning) had lower rewards than the clinician's policy (mean -3.03 and -2.93 vs. -2.93, respectively) for smaller datasets but had higher rewards for larger datasets (mean -4.68 and -2.82 vs. -4.57, respectively). CONCLUSIONS: Our results highlight the potential of using RL to generate the optimal treatment based on the patients' longitude records. Our work can lead the path towards developing RL-based decision support systems that could help manage AD with comorbidities. Kritib Bhattarai, Sivaraman Rajaganapathy, Trisha Das, Yejin Kim 0001, Yongbin Chen, Qiying Dai, Xiaoqian Jiang, Nansu Zong |
J. Am. Medical Informatics Assoc. | 9 |
| 2022 | Modeling a Cancer Symptom Control Domain Using HL7 FHIR: Applicability of the Minimal Common Oncology Data Elements (mCODE)
Nan Huo, Yue Yu 0012, Nansu Zong, Andrea Cheville, Claude J. Nanjo, Eric Prud'hommeaux, Deirdre Pachman, Guohui Xiao 0001, Emily R. Pfaff, Christopher G. Chute, Guoqian Jiang, Kathryn J. Ruddy |
AMIA | 3 |
| 2022 | A Comparative Study on the Capability of Real-World Antineoplastic Drug Data Collection by CanMED, ATC and HemOnc
Yue Yu 0012, Kathryn J. Ruddy, Nan Huo, Nansu Zong, Deirdre Pachman, Christopher G. Chute, Emily R. Pfaff, Andrea Cheville, Guoqian Jiang |
AMIA | 4 |
| 2022 | BETA: a comprehensive benchmark for computational drug-target predictionabstractInternal validation is the most popular evaluation strategy used for drug-target predictive models. The simple random shuffling in the cross-validation, however, is not always ideal to handle large, diverse and copious datasets as it could potentially introduce bias. Hence, these predictive models cannot be comprehensively evaluated to provide insight into their general performance on a variety of use-cases (e.g. permutations of different levels of connectiveness and categories in drug and target space, as well as validations based on different data sources). In this work, we introduce a benchmark, BETA, that aims to address this gap by (i) providing an extensive multipartite network consisting of 0.97 million biomedical concepts and 8.5 million associations, in addition to 62 million drug-drug and protein-protein similarities and (ii) presenting evaluation strategies that reflect seven cases (i.e. general, screening with different connectivity, target and drug screening based on categories, searching for specific drugs and targets and drug repurposing for specific diseases), a total of seven Tests (consisting of 344 Tasks in total) across multiple sampling and validation strategies. Six state-of-the-art methods covering two broad input data types (chemical structure- and gene sequence-based and network-based) were tested across all the developed Tasks. The best-worst performing cases have been analyzed to demonstrate the ability of the proposed benchmark to identify limitations of the tested methods for running over the benchmark tasks. The results highlight BETA as a benchmark in the selection of computational strategies for drug repurposing and target discovery. Nansu Zong, Ning Li 0045, Andrew Wen, Victoria Ngo, Yue Yu 0012, Ming Huang 0006, Shaika Chowdhury, Chao Jiang 0002, Sunyang Fu, Richard Weinshilboum, Guoqian Jiang, Lawrence Hunter |
Briefings Bioinform. | 1 |
| 2022 | Correction to: Characterizing phenotypic abnormalities associated with high-risk individuals developing lung cancer using electronic health records from the All of Us researcher workbenchabstractJournal of the American Medical Informatics Association, Volume 28, Issue 11, November 2021, Pages 2313–2324, https://doi.org/10.1093/jamia/ocab174 The revised manuscript entitled “Characterizing phenotypic abnormalities associated with high-risk individuals developing lung cancer using electronic health records from the All of Us researcher workbench,” by Jiang, et al. has been submitted by the authors as a replacement for the originally published version of the manuscript. Following publication, the authors were alerted to the article’s noncompliance with the All of Us Research Program’s Data Access and Use Policies. The policies violated are that authorized data users will not: After an extensive effort to obscure values that correspond to fewer than 20 participants to align with the All of Us Research Program’s policies, the authors have submitted a corrected manuscript. This version of the manuscript has been reviewed by the editorial team and is being republished after retraction of the original version of the... Jie Na, Nansu Zong, David E. Midthun, Guoqian Jiang |
J. Am. Medical Informatics Assoc. | 2 |
| 2022 | FHIR-Ontop-OMOP: Building clinical knowledge graphs in FHIR RDF with the OMOP Common data ModelabstractBACKGROUND: Knowledge graphs (KGs) play a key role to enable explainable artificial intelligence (AI) applications in healthcare. Constructing clinical knowledge graphs (CKGs) against heterogeneous electronic health records (EHRs) has been desired by the research and healthcare AI communities. From the standardization perspective, community-based standards such as the Fast Healthcare Interoperability Resources (FHIR) and the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) are increasingly used to represent and standardize EHR data for clinical data analytics, however, the potential of such a standard on building CKG has not been well investigated. OBJECTIVE: To develop and evaluate methods and tools that expose the OMOP CDM-based clinical data repositories into virtual clinical KGs that are compliant with FHIR Resource Description Framework (RDF) specification. METHODS: We developed a system called FHIR-Ontop-OMOP to generate virtual clinical KGs from the OMOP relational databases. We leveraged an OMOP CDM-based Medical Information Mart for Intensive Care (MIMIC-III) data repository to evaluate the FHIR-Ontop-OMOP system in terms of the faithfulness of data transformation and the conformance of the generated CKGs to the FHIR RDF specification. RESULTS: A beta version of the system has been released. A total of more than 100 data element mappings from 11 OMOP CDM clinical data, health system and vocabulary tables were implemented in the system, covering 11 FHIR resources. The generated virtual CKG from MIMIC-III contains 46,520 instances of FHIR Patient, 716,595 instances of Condition, 1,063,525 instances of Procedure, 24,934,751 instances of MedicationStatement, 365,181,104 instances of Observations, and 4,779,672 instances of CodeableConcept. Patient counts identified by five pairs of SQL (over the MIMIC database) and SPARQL (over the virtual CKG) queries were identical, ensuring the faithfulness of the data transformation. Generated CKG in RDF triples for 100 patients were fully conformant with the FHIR RDF specification. CONCLUSION: The FHIR-Ontop-OMOP system can expose OMOP database as a FHIR-compliant RDF graph. It provides a meaningful use case demonstrating the potentials that can be enabled by the interoperability between FHIR and OMOP CDM. Generated clinical KGs in FHIR RDF provide a semantic foundation to enable explainable AI applications in healthcare. Guohui Xiao 0001, Emily R. Pfaff, Eric Prud'hommeaux, David Booth, Deepak K. Sharma, Nan Huo, Yue Yu 0012, Nansu Zong, Kathryn J. Ruddy, Christopher G. Chute, Guoqian Jiang |
J. Biomed. Informatics | 8 |
| 2022 | Developing an ETL tool for converting the PCORnet CDM into the OMOP CDM to facilitate the COVID-19 data integration
Yue Yu 0012, Nansu Zong, Andrew Wen, Sijia Liu 0002, Daniel J. Stone, David Knaack, Alanna M. Chamberlain, Emily R. Pfaff, Davera Gabriel, Christopher G. Chute, Nilay Shah, Guoqian Jiang |
J. Biomed. Informatics | 2 |
| 2021 | Automatic Construction of Biomedical Knowledge Graph from Covid-19 Literature
Chao Jiang 0002, Victoria Ngo, Richard Chapman 0001, Guoqian Jiang, Nansu Zong |
AMIA | 6 |
| 2021 | Disparity analysis of patient portal messaging use for COVID-19 in urban versus rural locality
Ming Huang 0006, Andrew Wen, Liwei Wang 0010, Sijia Liu 0002, Yanshan Wang, Nansu Zong, Yue Yu 0012, Julie E. Prigge, Brian Costello, Nilay D. Shah, Henry Ting, Chyke Doubeni, Jungwei Fan 0001, Christi A. Patten |
AMIA | 7 |
| 2021 | FHIRTime: Standardizing Temporal Patterns Identified from Clinical Narratives Using HL7 FHIR
Daniel J. Stone, Sijia Liu 0002, Yuan Luo 0001, Andrew Wen, Nansu Zong, Luke V. Rasmussen, Prakash Adekkanattu, Pascal S. Brandt, Jennifer A. Pacheco, Fei Wang 0001, Cui Tao, Jyotishman Pathak, Guoqian Jiang |
AMIA | 5 |
| 2021 | Drug-target prediction utilizing heterogeneous bio-linked network embeddingsabstractTo enable modularization for network-based prediction, we conducted a review of known methods conducting the various subtasks corresponding to the creation of a drug-target prediction framework and associated benchmarking to determine the highest-performing approaches. Accordingly, our contributions are as follows: (i) from a network perspective, we benchmarked the association-mining performance of 32 distinct subnetwork permutations, arranging based on a comprehensive heterogeneous biomedical network derived from 12 repositories; (ii) from a methodological perspective, we identified the best prediction strategy based on a review of combinations of the components with off-the-shelf classification, inference methods and graph embedding methods. Our benchmarking strategy consisted of two series of experiments, totaling six distinct tasks from the two perspectives, to determine the best prediction. We demonstrated that the proposed method outperformed the existing network-based methods as well as how combinatorial networks and methodologies can influence the prediction. In addition, we conducted disease-specific prediction tasks for 20 distinct diseases and showed the reliability of the strategy in predicting 75 novel drug-target associations as shown by a validation utilizing DrugBank 5.1.0. In particular, we revealed a connection of the network topology with the biological explanations for predicting the diseases, 'Asthma' 'Hypertension', and 'Dementia'. The results of our benchmarking produced knowledge on a network-based prediction framework with the modularization of the feature selection and association prediction, which can be easily adapted and extended to other feature sources or machine learning algorithms as well as a performed baseline to comprehensively evaluate the utility of incorporating varying data sources. Nansu Zong, Rachael Sze Nga Wong, Yue Yu 0012, Andrew Wen, Ming Huang 0006, Ning Li 0045 |
Briefings Bioinform. | 1 |
| 2021 | Characterizing phenotypic abnormalities associated with high-risk individuals developing lung cancer using electronic health records from the All of Us researcher workbenchabstractOBJECTIVE: The study sought to test the feasibility of conducting a phenome-wide association study to characterize phenotypic abnormalities associated with individuals at high risk for lung cancer using electronic health records. MATERIALS AND METHODS: We used the beta release of the All of Us Researcher Workbench with clinical and survey data from a population of 225 000 subjects. We identified 3 cohorts of individuals at high risk to develop lung cancer based on (1) the 2013 U.S. Preventive Services Task Force criteria, (2) the long-term quitters of cigarette smoking criteria, and (3) the younger age of onset criteria. We applied the logistic regression analysis to identify the significant associations between individuals' phenotypes and their risk categories. We validated our findings against a lung cancer cohort from the same population and conducted an expert review to understand whether these associations are known or potentially novel. RESULTS: We found a total of 214 statistically significant associations (P < .05 with a Bonferroni correction and odds ratio > 1.5) enriched in the high-risk individuals from 3 cohorts, and 15 enriched in the low-risk individuals. Forty significant associations enriched in the high-risk individuals and 13 enriched in the low-risk individuals were validated in the cancer cohort. Expert review identified 15 potentially new associations enriched in the high-risk individuals. CONCLUSIONS: It is feasible to conduct a phenome-wide association study to characterize phenotypic abnormalities associated in high-risk individuals developing lung cancer using electronic health records. The All of Us Research Workbench is a promising resource for the research studies to evaluate and optimize lung cancer screening criteria. Jie Na, Nansu Zong, David E. Midthun, Yuan Luo 0001, Guoqian Jiang |
J. Am. Medical Informatics Assoc. | 2 |
| 2018 | DataMed - an open source discovery index for finding biomedical datasetsabstractOBJECTIVE: Finding relevant datasets is important for promoting data reuse in the biomedical domain, but it is challenging given the volume and complexity of biomedical data. Here we describe the development of an open source biomedical data discovery system called DataMed, with the goal of promoting the building of additional data indexes in the biomedical domain. MATERIALS AND METHODS: DataMed, which can efficiently index and search diverse types of biomedical datasets across repositories, is developed through the National Institutes of Health-funded biomedical and healthCAre Data Discovery Index Ecosystem (bioCADDIE) consortium. It consists of 2 main components: (1) a data ingestion pipeline that collects and transforms original metadata information to a unified metadata model, called DatA Tag Suite (DATS), and (2) a search engine that finds relevant datasets based on user-entered queries. In addition to describing its architecture and techniques, we evaluated individual components within DataMed, including the accuracy of the ingestion pipeline, the prevalence of the DATS model across repositories, and the overall performance of the dataset retrieval engine. RESULTS AND CONCLUSION: Our manual review shows that the ingestion pipeline could achieve an accuracy of 90% and core elements of DATS had varied frequency across repositories. On a manually curated benchmark dataset, the DataMed search engine achieved an inferred average precision of 0.2033 and a precision at 10 (P@10, the number of relevant results in the top 10 search results) of 0.6022, by implementing advanced natural language processing and terminology services. Currently, we have made the DataMed system publically available as an open source package for the biomedical community. Anupama E. Gururaj, Ibrahim Burak Özyurt, Ruiling Liu, Ergin Soysal, Trevor Cohen, Firat Tiryaki, Yueling Li, Nansu Zong, Min Jiang 0007, Deevakar Rogith, Mandana Salimi, Hyeon-Eui Kim, Philippe Rocca-Serra, Alejandra N. González-Beltrán, Claudiu Farcas, Todd Johnson, Ronald Margolis, George Alter, Susanna-Assunta Sansone, Ian Fore, Lucila Ohno-Machado, Jeffrey S. Grethe, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 9 |
| 2017 | Explorative Analyses on Indexing OMOP based Clinical Datasets with DATS
Ethan Park, Diana Guijarro, Imho Jang, Stephen Trac, Nansu Zong, Hyeon-Eui Kim |
AMIA | 5 |
| 2017 | Deep mining heterogeneous networks of biomedical linked data to predict novel drug-target associationsabstractMOTIVATION: A heterogeneous network topology possessing abundant interactions between biomedical entities has yet to be utilized in similarity-based methods for predicting drug-target associations based on the array of varying features of drugs and their targets. Deep learning reveals features of vertices of a large network that can be adapted in accommodating the similarity-based solutions to provide a flexible method of drug-target prediction. RESULTS: We propose a similarity-based drug-target prediction method that enhances existing association discovery methods by using a topology-based similarity measure. DeepWalk, a deep learning method, is adopted in this study to calculate the similarities within Linked Tripartite Network (LTN), a heterogeneous network generated from biomedical linked datasets. This proposed method shows promising results for drug-target association prediction: 98.96% AUC ROC score with a 10-fold cross-validation and 99.25% AUC ROC score with a Monte Carlo cross-validation with LTN. By utilizing DeepWalk, we demonstrate that: (i) this method outperforms other existing topology-based similarity computation methods, (ii) the performance is better for tripartite than with bipartite networks and (iii) the measure of similarity using network topology outperforms the ones derived from chemical structure (drugs) or genomic sequence (targets). Our proposed methodology proves to be capable of providing a promising solution for drug-target prediction based on topological similarity with a heterogeneous network, and may be readily re-purposed and adapted in the existing of similarity-based methodologies. AVAILABILITY AND IMPLEMENTATION: The proposed method has been developed in JAVA and it is available, along with the data at the following URL: https://github.com/zongnansu1982/drug-target-prediction . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Nansu Zong, Hyeon-Eui Kim, Victoria Ngo, Olivier Harismendy |
Bioinform. | 1 |
| 2017 | Constructing faceted taxonomy for heterogeneous entities based on object properties in linked data
Nansu Zong, Hong-Gee Kim, Sejin Nam |
Data Knowl. Eng. | 1 |
| 2017 | xStore: Federated temporal query processing for large scale RDF triples on a cloud environment
Jae-Hong Eom, Sejin Nam, Nansu Zong, Dong-Hyuk Im, Hong-Gee Kim |
Neurocomputing | 4 |
| 2015 | Aligning ontologies with subsumption and equivalence relations in Linked Data
Nansu Zong, Sejin Nam, Jae-Hong Eom, Hyunwhan Joe, Hong-Gee Kim |
Knowl. Based Syst. | 1 |