EDBT 2026 Demo / reviewers in the wild / expert
Victor M. Castro
dblp:148/5518
· DBLP profile ↗
22ranked-venue papers
5as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 5 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Comparing patient-reported symptoms and structured clinician documentation in electronic health recordsabstractOBJECTIVES: Real-world data (RWD) analyses primarily rely on structured clinical documentation collected through routine clinical care or driven by medical billing requirements. Patient-reported outcome measures (PROMs), integrated into electronic health records (EHRs), are an additional data source that could offer valuable insights into a patient's perspective and contribute to a more comprehensive understanding of health outcomes in RWD studies. This study aims to characterize agreement between PROMs symptoms and structured clinical documentation of these symptoms by clinicians in EHRs. MATERIALS AND METHODS: A cross-sectional study of 913 244 adult primary care annual physical visits between January 1, 2019 and December 31, 2023. We compared differences in prevalence and agreement of patient-reported symptoms (PRS) and structured clinician documentation (CD) across 15 respiratory, gastrointestinal, cardiometabolic, and neuropsychiatric symptoms. RESULTS: Patient-reported symptom prevalence were significantly higher compared to CD across most symptoms including joint pain (33% PRS vs 12%), headaches (17% PRS vs 8.8% CD), and sleep disturbance (24% PRS vs 10% CD). Clinicians documented anxiety (11% PRS vs 23% CD) and depression (6.6% PRS vs 15.4% CD) symptoms using structured code at higher rates than patients reported them. Agreement between symptom self-report and clinician-documented structured codes was low to moderate (κ: 0.06-0.39). DISCUSSION: Primary care patients self-report symptoms up to ten times more frequently than clinicians document them with structured codes in the EHR. CONCLUSION: This work demonstrates the value and feasibility of incorporating PRSs in RWD studies to reduce misclassification and more holistically capture a patient's health. Victor M. Castro, Vivian S. Gainer, Danielle M. Crookes, Shawn N. Murphy, Justin Manjourides |
J. Am. Medical Informatics Assoc. | 1 |
| 2024 | Semi-supervised Double Deep Learning Temporal Risk Prediction (SeDDLeR) with Electronic Health Records
Isabelle-Emmanuella Nogues, Jun Wen 0001, Yihan Zhao, Clara-Lea Bonzel, Victor M. Castro, Yucong Lin, Shike Xu, Jue Hou 0001, Tianxi Cai |
J. Biomed. Informatics | 5 |
| 2022 | The Mass General Brigham Biobank Portal: an i2b2-based data repository linking disparate and high-dimensional patient data to support multimodal analyticsabstractOBJECTIVE: Integrating and harmonizing disparate patient data sources into one consolidated data portal enables researchers to conduct analysis efficiently and effectively. MATERIALS AND METHODS: We describe an implementation of Informatics for Integrating Biology and the Bedside (i2b2) to create the Mass General Brigham (MGB) Biobank Portal data repository. The repository integrates data from primary and curated data sources and is updated weekly. The data are made readily available to investigators in a data portal where they can easily construct and export customized datasets for analysis. RESULTS: As of July 2021, there are 125 645 consented patients enrolled in the MGB Biobank. 88 527 (70.5%) have a biospecimen, 55 121 (43.9%) have completed the health information survey, 43 552 (34.7%) have genomic data and 124 760 (99.3%) have EHR data. Twenty machine learning computed phenotypes are calculated on a weekly basis. There are currently 1220 active investigators who have run 58 793 patient queries and exported 10 257 analysis files. DISCUSSION: The Biobank Portal allows noninformatics researchers to conduct study feasibility by querying across many data sources and then extract data that are most useful to them for clinical studies. While institutions require substantial informatics resources to establish and maintain integrated data repositories, they yield significant research value to a wide range of investigators. CONCLUSION: The Biobank Portal and other patient data portals that integrate complex and simple datasets enable diverse research use cases. i2b2 tools to implement these registries and make the data interoperable are open source and freely available. Victor M. Castro, Vivian S. Gainer, Nich Wattanasin, Barbara Benoit, Andrew Cagan, Bhaswati Ghosh, Sergey Goryachev, Reeta Metta, Heekyong Park, Taowei David Wang, Michael Mendis, Martin Rees, Christopher Herrick, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 1 |
| 2022 | Multiview Incomplete Knowledge Graph Integration with application to cross-institutional EHR data harmonizationabstractOBJECTIVE: The growing availability of electronic health records (EHR) data opens opportunities for integrative analysis of multi-institutional EHR to produce generalizable knowledge. A key barrier to such integrative analyses is the lack of semantic interoperability across different institutions due to coding differences. We propose a Multiview Incomplete Knowledge Graph Integration (MIKGI) algorithm to integrate information from multiple sources with partially overlapping EHR concept codes to enable translations between healthcare systems. METHODS: The MIKGI algorithm combines knowledge graph information from (i) embeddings trained from the co-occurrence patterns of medical codes within each EHR system and (ii) semantic embeddings of the textual strings of all medical codes obtained from the Self-Aligning Pretrained BERT (SAPBERT) algorithm. Due to the heterogeneity in the coding across healthcare systems, each EHR source provides partial coverage of the available codes. MIKGI synthesizes the incomplete knowledge graphs derived from these multi-source embeddings by minimizing a spherical loss function that combines the pairwise directional similarities of embeddings computed from all available sources. MIKGI outputs harmonized semantic embedding vectors for all EHR codes, which improves the quality of the embeddings and enables direct assessment of both similarity and relatedness between any pair of codes from multiple healthcare systems. RESULTS: With EHR co-occurrence data from Veteran Affairs (VA) healthcare and Mass General Brigham (MGB), MIKGI algorithm produces high quality embeddings for a variety of downstream tasks including detecting known similar or related entity pairs and mapping VA local codes to the relevant EHR codes used at MGB. Based on the cosine similarity of the MIKGI trained embeddings, the AUC was 0.918 for detecting similar entity pairs and 0.809 for detecting related pairs. For cross-institutional medical code mapping, the top 1 and top 5 accuracy were 91.0% and 97.5% when mapping medication codes at VA to RxNorm medication codes at MGB; 59.1% and 75.8% when mapping VA local laboratory codes to LOINC hierarchy. When trained with 500 labels, the lab code mapping attained top 1 and 5 accuracy at 77.7% and 87.9%. MIKGI also attained best performance in selecting VA local lab codes for desired laboratory tests and COVID-19 related features for COVID EHR studies. Compared to existing methods, MIKGI attained the most robust performance with accuracy the highest or near the highest across all tasks. CONCLUSIONS: The proposed MIKGI algorithm can effectively integrate incomplete summary data from biomedical text and EHR data to generate harmonized embeddings for EHR codes for knowledge graph modeling and cross-institutional translation of EHR codes. Doudou Zhou, Ziming Gan, Alina Patwari, Everett Neil Rush, Clara-Lea Bonzel, Vidul Ayakulangara Panickan, Chuan Hong, Yuk-Lam Ho, Tianrun A. Cai, Lauren Costa, Victor M. Castro, Shawn N. Murphy, Gabriel A. Brat, Griffin M. Weber, Paul Avillach, John Michael Gaziano, Kelly Cho, Katherine P. Liao, Tianxi Cai |
J. Biomed. Informatics | 13 |
| 2021 | Evaluation of the Portability of Natural Language Processing-based Computable Phenotypes in the eMERGE Network
Jennifer A. Pacheco, Luke V. Rasmussen, Ken Wiley, Thomas N. Person, David J. Cronkite, Sunghwan Sohn, Shawn N. Murphy, Justin H. Gundelach, Vivian S. Gainer, Victor M. Castro, Cong Liu 0020, Todd Lingren, Frank D. Mentch, Agnes S. Sundaresan, Garrett Eickelberg, Valerie Willis, Al'ona Furmanchuk, Roshan Patel, David Carrell, Marc S. Williams, Elizabeth W. Karlson, Jodell E. Linder, Yuan Luo 0001, Chunhua Weng, Wei-Qi Wei |
AMIA | 10 |
| 2021 | A Novel Data Portal to Enable COVID-19 Data Integration and Analysis
Nich Wattanasin, Victor M. Castro, Vivian S. Gainer, Barbara Benoit, Andrew Cagan, Reeta Metta, Shawn N. Murphy |
AMIA | 2 |
| 2021 | Temporally informed random forests for suicide risk predictionabstractOBJECTIVE: Suicide is one of the leading causes of death worldwide, yet clinicians find it difficult to reliably identify individuals at high risk for suicide. Algorithmic approaches for suicide risk detection have been developed in recent years, mostly based on data from electronic health records (EHRs). Significant room for improvement remains in the way these models take advantage of temporal information to improve predictions. MATERIALS AND METHODS: We propose a temporally enhanced variant of the random forest (RF) model-Omni-Temporal Balanced Random Forests (OT-BRFs)-that incorporates temporal information in every tree within the forest. We develop and validate this model using longitudinal EHRs and clinician notes from the Mass General Brigham Health System recorded between 1998 and 2018, and compare its performance to a baseline Naive Bayes Classifier and 2 standard versions of balanced RFs. RESULTS: Temporal variables were found to be associated with suicide risk: Elevated suicide risk was observed in individuals with a higher total number of visits as well as those with a low rate of visits over time, while lower suicide risk was observed in individuals with a longer period of EHR coverage. RF models were more accurate than Naive Bayesian classifiers at predicting suicide risk in advance (area under the receiver operating curve = 0.824 vs. 0.754, respectively). The proposed OT-BRF model performed best among all RF approaches, yielding a sensitivity of 0.339 at 95% specificity, compared to 0.290 and 0.286 for the other 2 RF models. Temporal variables were assigned high importance by the models that incorporated them. DISCUSSION: We demonstrate that temporal variables have an important role to play in suicide risk detection and that requiring their inclusion in all RF trees leads to increased predictive performance. Integrating temporal information into risk prediction models helps the models interpret patient data in temporal context, improving predictive performance. Ilkin Bayramli, Victor M. Castro, Yuval Barak-Corren, Emily M. Madsen, Matthew K. Nock, Jordan W. Smoller, Ben Y. Reis |
J. Am. Medical Informatics Assoc. | 2 |
| 2021 | Validation of an internationally derived patient severity phenotype to support COVID-19 analytics from electronic health record dataabstractOBJECTIVE: The Consortium for Clinical Characterization of COVID-19 by EHR (4CE) is an international collaboration addressing coronavirus disease 2019 (COVID-19) with federated analyses of electronic health record (EHR) data. We sought to develop and validate a computable phenotype for COVID-19 severity. MATERIALS AND METHODS: Twelve 4CE sites participated. First, we developed an EHR-based severity phenotype consisting of 6 code classes, and we validated it on patient hospitalization data from the 12 4CE clinical sites against the outcomes of intensive care unit (ICU) admission and/or death. We also piloted an alternative machine learning approach and compared selected predictors of severity with the 4CE phenotype at 1 site. RESULTS: The full 4CE severity phenotype had pooled sensitivity of 0.73 and specificity 0.83 for the combined outcome of ICU admission and/or death. The sensitivity of individual code categories for acuity had high variability-up to 0.65 across sites. At one pilot site, the expert-derived phenotype had mean area under the curve of 0.903 (95% confidence interval, 0.886-0.921), compared with an area under the curve of 0.956 (95% confidence interval, 0.952-0.959) for the machine learning approach. Billing codes were poor proxies of ICU admission, with as low as 49% precision and recall compared with chart review. DISCUSSION: We developed a severity phenotype using 6 code classes that proved resilient to coding variability across international institutions. In contrast, machine learning approaches may overfit hospital-specific orders. Manual chart review revealed discrepancies even in the gold-standard outcomes, possibly owing to heterogeneous pandemic conditions. CONCLUSIONS: We developed an EHR-based severity phenotype for COVID-19 in hospitalized patients and validated it at 12 international sites. Jeffrey G. Klann, Hossein Estiri, Griffin M. Weber, Bertrand Moal, Paul Avillach, Chuan Hong, Amelia L. M. Tan, Brett K. Beaulieu-Jones, Victor M. Castro, Thomas Maulhardt, Alon Geva, Alberto Malovini, Andrew M. South, Shyam Visweswaran, Michele Morris, Malarkodi J. Samayamuthu, Gilbert S. Omenn, Kee Yuan Ngiam, Kenneth D. Mandl, Martin Boeker, Karen L. Olson, Danielle L. Mowery, Robert W. Follett, David A. Hanauer, Riccardo Bellazzi, Jason H. Moore, Ne-Hooi Will Loh, Douglas S. Bell, Kavishwar B. Wagholikar, Luca Chiovato, Valentina Tibollo, Siegbert Rieg, Anthony L. L. J. Li, Vianney Jouhet, Emily Schriver, Zongqi Xia, Meghan Hutch, Yuan Luo 0001, Isaac S. Kohane, Gabriel A. Brat, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 9 |
| 2020 | Generative Transfer Learning for Measuring Plausibility of EHR Diagnosis Records Over Time
Hossein Estiri, Sebastien Vasey, Jeffrey G. Klann, Victor M. Castro, Shawn N. Murphy |
AMIA | 4 |
| 2020 | High-throughput Phenotyping with EHR Sequences
Shawn N. Murphy, Hossein Estiri, Zachary H. Strasser, Kavishwar B. Wagholikar, Victor M. Castro |
AMIA | 5 |
| 2020 | HistoriView: An Interactive and Scalable Visual Exploratory Plugin for Longitudinal Patient Data Review
Heekyong Park, Taowei David Wang, Vivian S. Gainer, Victor M. Castro, Nich Wattanasin, Shawn N. Murphy |
AMIA | 4 |
| 2020 | sureLDA: A multidisease automated phenotyping method for the electronic health recordabstractOBJECTIVE: A major bottleneck hindering utilization of electronic health record data for translational research is the lack of precise phenotype labels. Chart review as well as rule-based and supervised phenotyping approaches require laborious expert input, hampering applicability to studies that require many phenotypes to be defined and labeled de novo. Though International Classification of Diseases codes are often used as surrogates for true labels in this setting, these sometimes suffer from poor specificity. We propose a fully automated topic modeling algorithm to simultaneously annotate multiple phenotypes. MATERIALS AND METHODS: Surrogate-guided ensemble latent Dirichlet allocation (sureLDA) is a label-free multidimensional phenotyping method. It first uses the PheNorm algorithm to initialize probabilities based on 2 surrogate features for each target phenotype, and then leverages these probabilities to constrain the LDA topic model to generate phenotype-specific topics. Finally, it combines phenotype-feature counts with surrogates via clustering ensemble to yield final phenotype probabilities. RESULTS: sureLDA achieves reliably high accuracy and precision across a range of simulated and real-world phenotypes. Its performance is robust to phenotype prevalence and relative informativeness of surogate vs nonsurrogate features. It also exhibits powerful feature selection properties. DISCUSSION: sureLDA combines attractive properties of PheNorm and LDA to achieve high accuracy and precision robust to diverse phenotype characteristics. It offers particular improvement for phenotypes insufficiently captured by a few surrogate features. Moreover, sureLDA's feature selection ability enables it to handle high feature dimensions and produce interpretable computational phenotypes. CONCLUSIONS: sureLDA is well suited toward large-scale electronic health record phenotyping for highly multiphenotype applications such as phenome-wide association studies . Yuri Ahuja, Doudou Zhou, Zeling He, Jiehuan Sun, Victor M. Castro, Vivian S. Gainer, Shawn N. Murphy, Chuan Hong, Tianxi Cai |
J. Am. Medical Informatics Assoc. | 5 |
| 2019 | sureLDA: A Novel Multi-Disease Automated Phenotyping Method for the Electronic Health Record
Yuri Ahuja, Doudou Zhou, Zeling He, Jiehuan Sun, Victor M. Castro, Vivian S. Gainer, Shawn N. Murphy, Chuan Hong, Tianxi Cai |
AMIA | 5 |
| 2019 | High-throughput multimodal automated phenotyping (MAP) with application to PheWASabstractOBJECTIVE: Electronic health records linked with biorepositories are a powerful platform for translational studies. A major bottleneck exists in the ability to phenotype patients accurately and efficiently. The objective of this study was to develop an automated high-throughput phenotyping method integrating International Classification of Diseases (ICD) codes and narrative data extracted using natural language processing (NLP). MATERIALS AND METHODS: We developed a mapping method for automatically identifying relevant ICD and NLP concepts for a specific phenotype leveraging the Unified Medical Language System. Along with health care utilization, aggregated ICD and NLP counts were jointly analyzed by fitting an ensemble of latent mixture models. The multimodal automated phenotyping (MAP) algorithm yields a predicted probability of phenotype for each patient and a threshold for classifying participants with phenotype yes/no. The algorithm was validated using labeled data for 16 phenotypes from a biorepository and further tested in an independent cohort phenome-wide association studies (PheWAS) for 2 single nucleotide polymorphisms with known associations. RESULTS: The MAP algorithm achieved higher or similar AUC and F-scores compared to the ICD code across all 16 phenotypes. The features assembled via the automated approach had comparable accuracy to those assembled via manual curation (AUCMAP 0.943, AUCmanual 0.941). The PheWAS results suggest that the MAP approach detected previously validated associations with higher power when compared to the standard PheWAS method based on ICD codes. CONCLUSION: The MAP approach increased the accuracy of phenotype definition while maintaining scalability, thereby facilitating use in studies requiring large-scale phenotyping, such as PheWAS. Katherine P. Liao, Jiehuan Sun, Tianrun A. Cai, Nicholas B. Link, Chuan Hong, Jie Huang 0030, Jennifer E. Huffman, Jessica L. Gronsbell, Yuk-Lam Ho, Victor M. Castro, Vivian S. Gainer, Shawn N. Murphy, Christopher J. O'Donnell, John Michael Gaziano, Kelly Cho, Peter Szolovits, Isaac S. Kohane, Sheng Yu 0002 |
J. Am. Medical Informatics Assoc. | 11 |
| 2018 | High-Throughput Multimodal Automated Phenotyping (MAP) Incorporating Natural Language Processing with Application to PheWAS
Katherine P. Liao, Jiehuan Sun, Tianrun A. Cai, Nicholas B. Link, Chuan Hong, Jie Huang 0030, Jennifer E. Huffman, Jessica L. Gronsbell, Lauren Costa, Victor M. Castro, Vivian S. Gainer, Shawn N. Murphy, John Michael Gaziano, Kelly Cho, Peter Szolovits, Isaac S. Kohane, Sheng Yu 0002, Tianxi Cai |
AMIA | 10 |
| 2017 | Building Better Timeline Interactions for Patient Chart Reviews
Heekyong Park, Taowei David Wang, Vivian S. Gainer, Victor M. Castro, Shawn N. Murphy |
AMIA | 4 |
| 2015 | Stratification of Risk for Fall Resulting in Hospital Readmission through Medication Side Effects Profiles
Thomas H. McCoy Jr., Victor M. Castro, Roy H. Perlis |
AMIA | 2 |
| 2015 | Computable Phenotypes enabled by the i2b2 Validation Platform
Shawn N. Murphy, Vivian S. Gainer, Victor M. Castro, Alyssa P. Goodson, Lori C. Phillips, Sheng Yu 0002, Tianxi Cai |
AMIA | 3 |
| 2014 | Integrating Information from Unstructured Text with Structured Clinical Data from an Electronic Medical Record to Improve Patient Cohort Identification
Victor M. Castro, Sergey Goryachev, Christopher Herrick, Vivian S. Gainer, Martin Rees, Shawn N. Murphy |
AMIA | 1 |
| 2014 | Evaluation of matched control algorithms in EHR-based phenotyping studies: A case study of inflammatory bowel disease comorbidities
Victor M. Castro, W. Kay Apperson, Vivian S. Gainer, Ashwin N. Ananthakrishnan, Alyssa P. Goodson, Taowei David Wang, Christopher Herrick, Shawn N. Murphy |
J. Biomed. Informatics | 1 |
| 2013 | Amassing Pediatric Brain MRI's to Understand "Normal" using Mi2b2
Shawn N. Murphy, Christopher Herrick, Victor M. Castro, Randy L. Gollub, Nathaniel Reynolds, Patricia Ellen Grant |
AMIA | 3 |
| 2012 | Implementing a pharmacovigilance framework using data from electronic medical records
Victor M. Castro, Vivian S. Gainer, Christopher Herrick, Shawn N. Murphy, Wannapa Kay Mahamaneerat |
AMIA | 1 |