VLDB 2026 Research / reviewers in the wild / expert
Johanna Loomba
dblp:318/5993
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0003-3673-5423ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | National COVID Cohort Collaborative data enhancements: a path for expanding common data modelsabstractOBJECTIVE: To support long COVID research in National COVID Cohort Collaborative (N3C), the N3C Phenotype and Data Acquisition team created data designs to aid contributing sites in enhancing their data. Enhancements include long COVID specialty clinic indicator; Admission, Discharge, and Transfer transactions; patient-level social determinants of health; and in-hospital use of oxygen supplementation. MATERIALS AND METHODS: For each enhancement, we defined the scope and wrote guidance on how to prepare and populate the data in a standardized way. RESULTS: As of June 2024, 29 sites have added at least one data enhancement to their N3C pipeline. DISCUSSION: The use of common data models is critical to the success of N3C; however, these data models cannot account for all needs. Project-driven data enhancement is required. This should be done in a standardized way in alignment with common data model specifications. Our approach offers a useful pathway for enhancing data to improve fit for purpose. CONCLUSION: In this initiative, we rapidly produced project-specific data modeling guidance and documentation in support of long COVID research while maintaining a commitment to terminology standards and harmonized data. Kellie M. Walters, Marshall Clark, Sofia Dard, Stephanie S. Hong, Elizabeth Kelly, Kristin Kostka, Adam M. Lee, Robert T. Miller, Michele Morris, Matvey Palchuk, Emily R. Pfaff, Adam B. Wilcox, Alexis Graves, Alfred Anzalone, Amin Manna, Amit Saha, Amy Olex, Andrea Zhou, Andrew E. Williams, Andrew Southerland, Andrew T. Girvin, Anita Walden, Anjali A Sharathkumar, Benjamin R. C. Amor, Benjamin Bates, Brian Hendricks, Caleb Alexander, Carolyn T. Bramante, Cavin Ward-Caviness, Charisse R. Madlock-Brown, Christine Suver, Christopher G. Chute, Christopher Dillon, Chunlei Wu, Clare Schmitt, Cliff Takemoto, Dan Housman, Davera Gabriel, David Eichmann, Diego Mazzotti, Don Brown, Eilis A. Boudreau, Elaine L. Hill, Elizabeth Zampino, Emily Carlson Marti, Evan French, Farrukh M. Koraishy, Federico Mariona, Fred W. Prior, George Sokos, Greg Martin, Harold P. Lehmann, Heidi Spratt, Hemalkumar Mehta, Hythem Sidky, J. W. Awori Hayanga, Jami Pincavitch, Jaylyn Clark, Jeremy Richard Harper, Jessica Islam, Jin Ge, Joel Gagnier, Joel H. Saltz, Johanna Loomba, John Buse, Jomol P. Mathew, Joni L. Rutter, Julie A. McMurry, Justin Guinney, Justin Starren, Karen Crowley, Katie Rebecca Bradwell, Ken Wilkins, Kenneth R. Gersing, Kenrick Dwain Cato, Kimberly Murray, Lavance Northington, Lee Allan Pyles, Leonie Misquitta, Lesley Cottrell, Lili M. Portilla, Mariam Deacy, Mark M. Bissell, Mary Emmett, Mary Morrison Saltz, Melissa A. Haendel, Meredith C. B. Adams, Meredith Temple-O'Connor, Michael G. Kurilla, Nabeel Qureshi, Nasia Safdar, Nicole Garbarini, Noha Sharafeldin, Ofer Sadan, Patricia A. Francis, Penny Wung Burgoon, Peter N. Robinson, Philip R. O. Payne, Rafael Fuentes, Randeep Jawa, Rebecca Erwin-Cohen, Rena Patel, Richard A. Moffitt, Richard L. Zhu, Rishi Kamaleswaran, Robert Hurley, Saiju Pyarajan, Samuel G. Michael, Samuel Bozzette, Sandeep Mallipattu, Satyanarayana Vedula, Scott Chapman, Shawn T. O'Neil, Soko Setoguchi, Tellen D. Bennett, Tiffany Callahan, Umit Topaloglu, Usman Sheikh, Valery Gordon, Vignesh Subbian, Warren A. Kibbe, Wenndy Hernandez, Will Beasley, Will Cooper, William Hillegass, Xiaohan Tanner Zhang |
J. Am. Medical Informatics Assoc. | 66 |
| 2023 | A Bayesian Hierarchical Analysis on the Disparity of Emergency Department Visits for COVID-19: A Cohort Study Using National COVID Cohort Collaborative (N3C) DataabstractSince the COVID-19 pandemic in 2020, there are numerous studies and researches on the long term effect of COVID-19 on both patient level and social level with disparities noted in infection rates and outcomes. However, differences in the COVID related healthcare decisions after patients present for care (e.g. hospitalization after emergency department visit) has not been studied much at a national level. The National COVID Cohort Collaborative (N3C) provides researchers with abundant data collected from different clinical sites, making it suitable for Bayesian hierarchical modeling while analyzing the disparity in hospitalization after emergency department visit, where prior information or belief could be easily included in the modeling process by adjusting the prior distribution of parameters. In this analysis, we select demographic information (age, sex, race and ethnicity) and the Charlson Comorbidity Index (CCI) as features and study the relationship between these features and whether a patient would be hospitalized after having a COVID related visit to an emergency department (ED). Johanna Loomba, Andrea Zhou, Suchetha Sharma, Saurav Sengupta, Donald E. Brown |
ICMLA | 2 |
| 2023 | Determining Risk Factors for Long COVID Using Positive Unlabeled Learning on Electronic Health Records Data from NIH N3CabstractPost-acute sequelae of SARS-Co V-2 infection (PASC), also known as Long COVID, is an emerging medical condition in the aftermath of the COVID-19 pandemic. Research on this disease is limited by its newness and the lack of reliable controls, which can hinder model development. The National COVID Cohort Collaborative (N3C)11https://ncats.nih.gov/n3c contains Electronic Health Record (EHR) data for 7 million COVID positive patients from 76 sites across the United States, of which there are fifty thousand Long COVID patients. For this study, we model our risk factor analysis as Positive Unlabeled (PU) problem, where we treat Long COVID patients as the positive sample and rest of the COVID positive patients as unlabeled data. We first curate reliable controls using a PU modeling technique called bagging. We then use this cohort of positive and the curated negative samples to model risk factors for Long COVID. We utilize an attention-based deep learning approach using Long Short Term Memory (LSTM) networks on historical diagnosis data prior to COVID-19 infection, to first predict for Long COVID and then extract the model attention values to score input diagnoses for each patient. Using this process, we achieve an Area Under the Receiver Operating Characteristic (AUROC) of 0.93 (0.88 F1 Score) for the prediction task, significantly outperforming the same model trained on randomly selected controls. We then use a scoring process to rank different input diagnoses for each correctly classified patient with attention values extracted from the trained model and find the temporal distribution of top diagnosis codes which, when represented graphically, becomes a helpful tool to for physicians to investigate diagnosis patterns that effect Long COVID and also evaluate model trustworthiness. Saurav Sengupta, Johanna Loomba, Suchetha Sharma, Scott A. Chapman, Donald E. Brown |
ICMLA | 2 |
| 2022 | Vital Measurements of Hospitalized COVID-19 Patients as a Predictor of Long COVID: An EHR-based Cohort Study from the RECOVER Program in N3CabstractIt is shown that various symptoms could remain in the stage of post-acute sequelae of SARS-CoV-2 infection (PASC), otherwise known as Long COVID. A number of COVID patients suffer from heterogeneous symptoms, which severely impact recovery from the pandemic. While scientists are trying to give an unambiguous definition of Long COVID, efforts in prediction of Long COVID could play an important role in understanding the characteristic of this new disease. Vital measurements (e.g. oxygen saturation, heart rate, blood pressure) could reflect body's most basic functions and are measured regularly during hospitalization, so among patients diagnosed COVID positive and hospitalized, we analyze the vital measurements of first 7 days since the hospitalization start date to study the pattern of the vital measurements and predict Long COVID with the information from vital measurements. Johanna Loomba, Suchetha Sharma, Donald E. Brown |
BIBM | 2 |
| 2022 | Analyzing historical diagnosis code data from NIH N3C and RECOVER Programs using deep learning to determine risk factors for Long CovidabstractPost-acute sequelae of SARS-CoV-2 infection (PASC) or Long COVID is an emerging medical condition that has been observed in several patients with a positive diagnosis for COVID-19. Historical Electronic Health Records (EHR) like diagnosis codes, lab results and clinical notes have been analyzed using deep learning and have been used to predict future clinical events. In this paper, we propose an interpretable deep learning approach to analyze historical diagnosis code data from the National COVID Cohort Collective (N3C)1to find the risk factors contributing to developing Long COVID. Using our deep learning approach, we are able to predict if a patient is suffering from Long COVID from a temporally ordered list of diagnosis codes up to 45 days post the first COVID positive test or diagnosis for each patient, with an accuracy of 70.48%. We are then able to examine the trained model using Gradient-weighted Class Activation Mapping (GradCAM) to give each input diagnoses a score. The highest scored diagnosis were deemed to be the most important for making the correct prediction for a patient. We also propose a way to summarize these top diagnoses for each patient in our cohort and look at their temporal trends to determine which codes contribute towards a positive Long COVID diagnosis. Saurav Sengupta, Johanna Loomba, Suchetha Sharma, Donald E. Brown, Lorna E. Thorpe, Melissa A. Haendel, Christopher G. Chute, Stephanie S. Hong |
BIBM | 2 |
| 2022 | The iTHRIV Commons: a cross-institution information and health research data sharing architecture and web applicationabstractOBJECTIVE: The integrated Translational Health Research Institute of Virginia (iTHRIV) aims to develop an information architecture to support data workflows throughout the research lifecycle for cross-state teams of translational researchers. MATERIALS AND METHODS: The iTHRIV Commons is a cross-state harmonized infrastructure supporting resource discovery, targeted consultations, and research data workflows. As the front end to the iTHRIV Commons, the iTHRIV Research Concierge Portal supports federated login, personalized views, and secure interactions with objects in the ITHRIV Commons federation. The canonical use-case for the iTHRIV Commons involves an authenticated user connected to their respective high-security institutional network, accessing the iTHRIV Research Concierge Portal web application on their browser, and interfacing with multi-component iTHRIV Commons Landing Services installed behind the firewall at each participating institution. RESULTS: The iTHRIV Commons provides a technical framework, including both hardware and software resources located in the cloud and across partner institutions, that establishes standard representation of research objects, and applies local data governance rules to enable access to resources from a variety of stakeholders, both contributing and consuming. DISCUSSION: The launch of the Commons API service at partner sites and the addition of a public view of nonrestricted objects will remove barriers to data access for cross-state research teams while supporting compliance and the secure use of data. CONCLUSIONS: The secure architecture, distributed APIs, and harmonized metadata of the iTHRIV Commons provide a methodology for compliant information and data sharing that can advance research productivity at Hub sites across the CTSA network. Johanna Loomba, Glenn S. Wasson, Ravi Kiran Reddy Chamakuri, Pabitra Kumar Dash, Stephen G. Patterson, Mary M. A. Potter, Jason Edward Krisch, Martha M. Tenzer, Karen C. Johnston, Donald E. Brown |
J. Am. Medical Informatics Assoc. | 1 |
| 2022 | Synergies between centralized and federated approaches to data quality: a report from the national COVID cohort collaborativeabstractOBJECTIVE: In response to COVID-19, the informatics community united to aggregate as much clinical data as possible to characterize this new disease and reduce its impact through collaborative analytics. The National COVID Cohort Collaborative (N3C) is now the largest publicly available HIPAA limited dataset in US history with over 6.4 million patients and is a testament to a partnership of over 100 organizations. MATERIALS AND METHODS: We developed a pipeline for ingesting, harmonizing, and centralizing data from 56 contributing data partners using 4 federated Common Data Models. N3C data quality (DQ) review involves both automated and manual procedures. In the process, several DQ heuristics were discovered in our centralized context, both within the pipeline and during downstream project-based analysis. Feedback to the sites led to many local and centralized DQ improvements. RESULTS: Beyond well-recognized DQ findings, we discovered 15 heuristics relating to source Common Data Model conformance, demographics, COVID tests, conditions, encounters, measurements, observations, coding completeness, and fitness for use. Of 56 sites, 37 sites (66%) demonstrated issues through these heuristics. These 37 sites demonstrated improvement after receiving feedback. DISCUSSION: We encountered site-to-site differences in DQ which would have been challenging to discover using federated checks alone. We have demonstrated that centralized DQ benchmarking reveals unique opportunities for DQ improvement that will support improved research analytics locally and in aggregate. CONCLUSION: By combining rapid, continual assessment of DQ with a large volume of multisite data, it is possible to support more nuanced scientific questions with the scale and rigor that they require. Emily R. Pfaff, Andrew T. Girvin, Davera Gabriel, Kristin Kostka, Michele Morris, Matvey Palchuk, Harold P. Lehmann, Benjamin R. C. Amor, Mark Bissell, Katie R. Bradwell, Sigfried Gold, Stephanie S. Hong, Johanna Loomba, Amin Manna, Julie A. McMurry, Emily Niehaus, Nabeel Qureshi, Anita Walden, Xiaohan Tanner Zhang, Richard L. Zhu, Richard A. Moffitt, Christopher G. Chute, William G. Adams, Shaymaa Al-Shukri, Alfred Anzalone, Ahmad Baghal, Tellen D. Bennett, Elmer V. Bernstam, Mark M. Bissell, Brian Bush, Thomas R. Campion Jr., Victor Castro, Jack Chang, Deepa D. Chaudhari, Wenjin Chen, San Chu, James J. Cimino, Keith A. Crandall, Mark Crooks, Sara J. Deakyne Davies, John Dipalazzo, David A. Dorr, Daniel Eckrich, Sarah E. Eltinge, Daniel G. Fort, Georgiy Golovko, Snehil Gupta, Melissa A. Haendel, Janos G. Hajagos, David A. Hanauer, Brett M. Harnett, Ronald Horswell, Nancy Huang, Steven G. Johnson, Michael Kahn, Kamil Khanipov, Curtis Kieler, Katherine Ruiz De Luzuriaga, Sarah E. Maidlow, Ashley Martinez, Jomol Mathew, James C. McClay, Gabriel McMahan, Brian Melancon, Stéphane M. Meystre, Lucio Miele, Hiroki Morizono, Ray Pablo, Lav P. Patel, Jimmy Phuong, Daniel J. Popham, Claudia P. Pulgarin, Indra Neil Sarkar, Nancy Sazo, Soko Setoguchi, Selvin Soby, Sirisha Surampalli, Christine Suver, Uma Maheswara Reddy Vangala, Shyam Visweswaran, James von Oehsen, Kellie M. Walters, Laura K. Wiley, David A. Williams, Adrian H. Zai |
J. Am. Medical Informatics Assoc. | 13 |