Juan Zhao 0003

dblp:32/757-3 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
7since 2021 · last 2023
0000-0003-1429-0662ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 7 first-author · 6 since 2021Security and privacy · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2023 Evaluating resources composing the PheMAP knowledge base to enhance high-throughput phenotyping
abstract
OBJECTIVE: A previous study, PheMAP, combined independent, online resources to enable high-throughput phenotyping (HTP) using electronic health records (EHRs). However, online resources offer distinct quality descriptions of diseases which may affect phenotyping performance. We aimed to evaluate the phenotyping performance of single resource-based PheMAPs and investigate an optimized strategy for HTP. MATERIALS AND METHODS: We compared how each resource produced top-ranked concept unique identifiers (CUIs) by term frequency-inverse document frequency with Jaccard matrices comparing single resources and the original PheMAP. We correlated top-ranked concepts from each resource to features used in established Phenotype KnowledgeBase (PheKB) algorithms for hypothyroidism, type II diabetes mellitus (T2DM), and dementias. Using resources separately, we calculated multiple phenotype risk scores for individuals from Vanderbilt University Medical Center's BioVU DNA Biobank and compared phenotyping performance against rule-based eMERGE algorithms. Lastly, we implemented an ensemble strategy which classified patient case/control status based upon PheMAP resource agreement. RESULTS: Jaccard similarity matrices indicate that the similarity of CUIs comprising single resource-based PheMAPs varies. Single resource-based PheMAPs generated from MedlinePlus and MedicineNet outperformed others but only encompass 81.6% of overall disease phenotypes. We propose the PheMAP-Ensemble which provides higher average accuracy and precision than the combined average accuracy and precision of single resource-based PheMAPs. While offering complete phenotype coverage, PheMAP-Ensemble significantly increases phenotyping recall compared to the original iteration. CONCLUSIONS: Resources comprising the PheMAP produce different phenotyping performance when implemented individually. The ensemble method significantly improves the quality of PheMAP by fully utilizing dissimilar resources to capture accurate phenotyping data from EHRs.
Nicholas C. Wan, Ali A Yaqoob, Henry H. Ong, Juan Zhao 0003, Wei-Qi Wei
J. Am. Medical Informatics Assoc.4
2023 Evaluating and mitigating bias in machine learning models for cardiovascular disease prediction
Fuchen Li, Patrick Wu, Henry H. Ong, Josh F. Peterson, Wei-Qi Wei, Juan Zhao 0003
J. Biomed. Informatics6
2022 Evaluating Phenotype Classification Using Synthesized Online Content
Wei-Qi Wei, Juan Zhao 0003, Henry H. Ong
AMIA3
2022 Bassa-ML - A Blockchain and Model Card Integrated Federated Learning Provenance Platform
abstract
Federated learning is a collaborative/distributed machine learning system which is designed to address the privacy issues in centralized machine learning systems. The transparency and provenance of a machine learning model are important aspects of federated learning systems since they impact peoples’ lives in various domains (e.g., from healthcare to personal finance to employment). However, most of the existing federated learning systems deal with centralized coordinators which are vulnerable to attacks and privacy breaches. Also, they do not provide any standard transparency and provenance mechanisms for the resulting models. In this paper, we propose a blockchain and Model Card-based integrated federated learning system "Bassa-ML" providing enhanced transparency and trust for the models. Model parameter sharing, local model generation, model averaging, and model sharing functions are implemented using smart contracts. The generated models, model training information, and model reports are stored in the blockchain ledger as Model Card Objects. This results in enhanced transparency and auditability to the federated learning process.
Eranga Bandara, Sachin Shetty, Abdul Rahman, Ravi Mukkamala, Juan Zhao 0003, Xueping Liang
CCNC5
2021 Early Detect COVID-19 Presenting Symptoms and Characteristics Using Natural Language Processing on Electronic Health Records
Juan Zhao 0003, Monika E. Grabowska, Wei-Qi Wei
AMIA1
2021 DDIWAS: High-throughput electronic health record-based screening of drug-drug interactions
abstract
OBJECTIVE: We developed and evaluated Drug-Drug Interaction Wide Association Study (DDIWAS). This novel method detects potential drug-drug interactions (DDIs) by leveraging data from the electronic health record (EHR) allergy list. MATERIALS AND METHODS: To identify potential DDIs, DDIWAS scans for drug pairs that are frequently documented together on the allergy list. Using deidentified medical records, we tested 616 drugs for potential DDIs with simvastatin (a common lipid-lowering drug) and amlodipine (a common blood-pressure lowering drug). We evaluated the performance to rediscover known DDIs using existing knowledge bases and domain expert review. To validate potential novel DDIs, we manually reviewed patient charts and searched the literature. RESULTS: DDIWAS replicated 34 known DDIs. The positive predictive value to detect known DDIs was 0.85 and 0.86 for simvastatin and amlodipine, respectively. DDIWAS also discovered potential novel interactions between simvastatin-hydrochlorothiazide, amlodipine-omeprazole, and amlodipine-valacyclovir. A software package to conduct DDIWAS is publicly available. CONCLUSIONS: In this proof-of-concept study, we demonstrate the value of incorporating information mined from existing allergy lists to detect DDIs in a real-world clinical setting. Since allergy lists are routinely collected in EHRs, DDIWAS has the potential to detect and validate DDI signals across institutions.
Patrick Wu, Scott D. Nelson, Juan Zhao 0003, Cosby A. Stone Jr., QiPing Feng, Qingxia Chen, Eric A. Larson, Bingshan Li, Nancy J. Cox, C. Michael Stein, Elizabeth Phillips, Dan M. Roden, Joshua C. Denny, Wei-Qi Wei
J. Am. Medical Informatics Assoc.3
2021 ConceptWAS: A high-throughput method for early identification of COVID-19 presenting symptoms and characteristics from clinical notes
Juan Zhao 0003, Monika E. Grabowska, Vern Eric Kerchberger, Joshua C. Smith, H. Nur Eken, QiPing Feng, Josh F. Peterson, S. Trent Rosenbloom, Kevin B. Johnson, Wei-Qi Wei
J. Biomed. Informatics1
2020 Detecting National Institutes of Health's funding interests and trends using Machine Learning
Juan Zhao 0003, QiPing Feng, Wei-Qi Wei
AMIA1
2020 PheMap: a multi-resource knowledge base for high-throughput phenotyping within electronic health records
abstract
OBJECTIVE: Developing algorithms to extract phenotypes from electronic health records (EHRs) can be challenging and time-consuming. We developed PheMap, a high-throughput phenotyping approach that leverages multiple independent, online resources to streamline the phenotyping process within EHRs. MATERIALS AND METHODS: PheMap is a knowledge base of medical concepts with quantified relationships to phenotypes that have been extracted by natural language processing from publicly available resources. PheMap searches EHRs for each phenotype's quantified concepts and uses them to calculate an individual's probability of having this phenotype. We compared PheMap to clinician-validated phenotyping algorithms from the Electronic Medical Records and Genomics (eMERGE) network for type 2 diabetes mellitus (T2DM), dementia, and hypothyroidism using 84 821 individuals from Vanderbilt Univeresity Medical Center's BioVU DNA Biobank. We implemented PheMap-based phenotypes for genome-wide association studies (GWAS) for T2DM, dementia, and hypothyroidism, and phenome-wide association studies (PheWAS) for variants in FTO, HLA-DRB1, and TCF7L2. RESULTS: In this initial iteration, the PheMap knowledge base contains quantified concepts for 841 disease phenotypes. For T2DM, dementia, and hypothyroidism, the accuracy of the PheMap phenotypes were >97% using a 50% threshold and eMERGE case-control status as a reference standard. In the GWAS analyses, PheMap-derived phenotype probabilities replicated 43 of 51 previously reported disease-associated variants for the 3 phenotypes. For 9 of the 11 top associations, PheMap provided an equivalent or more significant P value than eMERGE-based phenotypes. The PheMap-based PheWAS showed comparable or better performance to a traditional phecode-based PheWAS. PheMap is publicly available online. CONCLUSIONS: PheMap significantly streamlines the process of extracting research-quality phenotype information from EHRs, with comparable or better performance to current phenotyping approaches.
Neil S. Zheng, QiPing Feng, Vern Eric Kerchberger, Juan Zhao 0003, Todd L. Edwards, Nancy J. Cox, C. Michael Stein, Dan M. Roden, Joshua C. Denny, Wei-Qi Wei
J. Am. Medical Informatics Assoc.4
2019 Deep Learning Using Electronic Health Records and Genetic Data to Predict Cardiovascular Diseases
Juan Zhao 0003, QiPing Feng, Patrick Wu, Joshua C. Denny, Wei-Qi Wei
AMIA1
2019 Transfer learning for detecting unknown network attacks
abstract
Network attacks are serious concerns in today’s increasingly interconnected society. Recent studies have applied conventional machine learning to network attack detection by learning the patterns of the network behaviors and training a classification model. These models usually require large labeled datasets; however, the rapid pace and unpredictability of cyber attacks make this labeling impossible in real time. To address these problems, we proposed utilizing transfer learning for detecting new and unseen attacks by transferring the knowledge of the known attacks. In our previous work, we have proposed a transfer learning-enabled framework and approach, called HeTL, which can find the common latent subspace of two different attacks and learn an optimized representation, which was invariant to attack behaviors’ changes. However, HeTL relied on manual pre-settings of hyper-parameters such as relativeness between the source and target attacks. In this paper, we extended this study by proposing a clustering-enhanced transfer learning approach, called CeHTL, which can automatically find the relation between the new attack and known attack. We evaluated these approaches by stimulating scenarios where the testing dataset contains different attack types or subtypes from the training set. We chose several conventional classification models such as decision trees, random forests, KNN, and other novel transfer learning approaches as strong baselines. Results showed that proposed HeTL and CeHTL improved the performance remarkably. CeHTL performed best, demonstrating the effectiveness of transfer learning in detecting new network attacks.
Juan Zhao 0003, Sachin Shetty, Jan Wei Pan, Charles A. Kamhoua, Kevin A. Kwiat
EURASIP J. Inf. Secur.1
2019 Detecting time-evolving phenotypic topics via tensor factorization on electronic health records: Cardiovascular disease case study
Juan Zhao 0003, David J. Schlueter, Patrick Wu, Vern Eric Kerchberger, S. Trent Rosenbloom, Quinn Stanton Wells, QiPing Feng, Joshua C. Denny, Wei-Qi Wei
J. Biomed. Informatics1
2018 Using Topic Modeling to Identify Relationship between LPA Variant and Disease Phenotypes
Juan Zhao 0003, QiPing Feng, Patrick Wu, Joshua C. Denny, Wei-Qi Wei
AMIA1
2018 A Reliable Data Provenance and Privacy Preservation Architecture for Business-Driven Cyber-Physical Systems Using Blockchain
abstract
Cyber-physical systems (CPS) including power systems, transportation, industrial control systems, etc. support both advanced control and communications among system components. Frequent data operations could introduce random failures and malicious attacks or even bring down the whole system. The dependency on a central authority increases the risk of single point of failure. To establish an immutable data provenance scheme for CPS, the authors adopt blockchain and propose a decentralized architecture to assure data integrity. In business-driven CPS, end users are required to share their personal information with multiple third parties. To prevent data leakage and preserve user privacy, the authors isolate and feed different information retrieval requests using tokens specifically generated for each type of request. Providing both traceability of data operations, and unlinkability of end user activities, a robust blockchain-based CPS is prototyped. Evaluation indicates the architecture is capable of assured data provenance validation and user privacy preservation at a low overhead.
Xueping Liang, Sachin Shetty, Deepak K. Tosh, Juan Zhao 0003, Danyi Li, Jihong Liu
Int. J. Inf. Secur. Priv.4
2017 Towards Decentralized Accountability and Self-sovereignty in Healthcare Systems
Xueping Liang, Sachin Shetty, Juan Zhao 0003, Daniel Bowden, Danyi Li, Jihong Liu
ICICS3
2017 Integrating blockchain for data sharing and collaboration in mobile healthcare applications
abstract
Enabled by mobile and wearable technology, personal health data delivers immense and increasing value for healthcare, benefiting both care providers and medical research. The secure and convenient sharing of personal health data is crucial to the improvement of the interaction and collaboration of the healthcare industry. Faced with the potential privacy issues and vulnerabilities existing in current personal health data storage and sharing systems, as well as the concept of self-sovereign data ownership, we propose an innovative user-centric health data sharing solution by utilizing a decentralized and permissioned blockchain to protect privacy using channel formation scheme and enhance the identity management using the membership service supported by the blockchain. A mobile application is deployed to collect health data from personal wearable devices, manual input, and medical devices, and synchronize data to the cloud for data sharing with healthcare providers and health insurance companies. To preserve the integrity of health data, within each record, a proof of integrity and validation is permanently retrievable from cloud database and is anchored to the blockchain network. Moreover, for scalable and performance considerations, we adopt a tree-based data processing and batching method to handle large data sets of personal health data collected and uploaded by the mobile platform.
Xueping Liang, Juan Zhao 0003, Sachin Shetty, Jihong Liu, Danyi Li
PIMRC2
2015 An Influence Field Perspective on Predicting User's Retweeting Behavior
Jianjun Yu, Kejun Dong, Juan Zhao 0003, Kai Nan
WAIM4
2014 Recommending funding collaborators with scholar social networks
abstract
Applying for research funding projects is becoming one of the most important ways for scientists to carry on the research. How to find an appropriate collaborator/applicant is a major concern for scientists. Social networks provide one means of visualizing existing and potential collaborations. In this paper, we study the funding collaborators recommendation problems. We solve the problem by starting with analyzing the researchers' motivations for finding collaboration, which are (i) to form a competitive team (ii) to expand cooperation circle, which little work noticed. We model the funding relation as a complex network called co-applicant network. Based on that, we propose a utility function to take all the aspects of recommendation into account. And we propose a novel recommendation algorithm by modeling the utility function based on the group relations in the co-applicant network. We experiment our approaches on National Science Foundation of China (NSFC) funding projects and achieve effective results.
Juan Zhao 0003, Kejun Dong, Jianjun Yu
DSAA1
2013 DSN: A Knowledge-Based Scholar Networking Practice Towards Research Community
abstract
In this paper, we carry out a knowledge-based scholar network practice towards Research community, named Research Social Networking, shortly DSN, by setting up a large knowledge base of scientists. We discuss key technologies in the paper, including scholar disambiguation and relationship extraction with the better performance evaluation than traditional methods. The DSN system has been implemented and integrated with Duckling cloud service, known as Research Online, with more than 60 thousand scientists and 100 thousand papers.
Juan Zhao 0003, Kejun Dong, Jianjun Yu
e-Science1