VLDB 2026 Research / reviewers in the wild / expert
Shyam Visweswaran
dblp:54/110
· DBLP profile ↗
70ranked-venue papers
9as first author
23since 2021 · last 2025
0000-0002-2079-8684ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 65 · 7 first-author · 22 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | In defense of empathic informaticsabstractOBJECTIVES: To explore the potential effects of recent restrictions on discussions regarding diversity, equity, and inclusion (DEI) in the field of biomedical informatics. MATERIALS AND METHODS: Executive orders issued by the U.S. federal government regarding diversity and gender issues are discussed in the context of implications for biomedical informatics research. RESULTS: Restrictions on specific terminology can hinder research into critical topics such as bias and fairness in clinical artificial intelligence and machine learning algorithms. Additionally, these limitations may narrow the scope of questions that informatics research can address and obstruct efforts to enhance the diversity of perspectives within the field. DISCUSSION: Responding to these threats requires a community response. The American Medical Informatics Association (AMIA) can help the informatics community present a united front in support of DEI research in multiple ways. CONCLUSION: The informatics community should take a strong and unambiguous response to support diversity, equity, and inclusion of underrepresented perspectives in the field. Harry Hochheiser, Shyam Visweswaran |
J. Am. Medical Informatics Assoc. | 2 |
| 2025 | A resource for Logical Observation Identifiers Names and Codes terms that may be associated with identifying informationabstractOBJECTIVES: The primary objective was to compile a comprehensive list of Logical Observation Identifiers Names and Codes (LOINC) terms that may be associated with patient, healthcare provider, and healthcare facility identifying information. MATERIALS AND METHODS: We developed a 2-step procedure for identifying LOINC terms, which consists of a keyword search of Long Common Names and filtering on selected property values, followed by expert physician review to confirm and categorize the terms. RESULTS: The final list comprises 1309 LOINC terms potentially associated with identifying information of patients, providers, and facilities. This list is publicly available on GitHub. DISCUSSION: Compared with electronic health record data coded with other terminologies, LOINC-coded data present unique challenges for deidentification, and a resource of LOINC terms that may be associated with identifying information will be helpful for this purpose. CONCLUSION: This resource is valuable for deidentifying LOINC-coded data, ensuring compliance with the Health Insurance Portability and Accountability Act (HIPAA), and preserving the privacy of patients, providers, and facilities. Mehdi Nourelahi, Eugene Sadhu, Malarkodi J. Samayamuthu, Shyam Visweswaran |
J. Am. Medical Informatics Assoc. | 4 |
| 2024 | Mammo-CLIP: A Vision Language Foundation Model to Enhance Data Efficiency and Robustness in Mammography
Shantanu Ghosh, Clare B. Poynton, Shyam Visweswaran, Kayhan Batmanghelich |
MICCAI (12) | 3 |
| 2024 | Towards cross-application model-agnostic federated cohort discoveryabstractOBJECTIVES: To demonstrate that 2 popular cohort discovery tools, Leaf and the Shared Health Research Information Network (SHRINE), are readily interoperable. Specifically, we adapted Leaf to interoperate and function as a node in a federated data network that uses SHRINE and dynamically generate queries for heterogeneous data models. MATERIALS AND METHODS: SHRINE queries are designed to run on the Informatics for Integrating Biology & the Bedside (i2b2) data model. We created functionality in Leaf to interoperate with a SHRINE data network and dynamically translate SHRINE queries to other data models. We randomly selected 500 past queries from the SHRINE-based national Evolve to Next-Gen Accrual to Clinical Trials (ENACT) network for evaluation, and an additional 100 queries to refine and debug Leaf's translation functionality. We created a script for Leaf to convert the terms in the SHRINE queries into equivalent structured query language (SQL) concepts, which were then executed on 2 other data models. RESULTS AND DISCUSSION: 91.1% of the generated queries for non-i2b2 models returned counts within 5% (or ±5 patients for counts under 100) of i2b2, with 91.3% recall. Of the 8.9% of queries that exceeded the 5% margin, 77 of 89 (86.5%) were due to errors introduced by the Python script or the extract-transform-load process, which are easily fixed in a production deployment. The remaining errors were due to Leaf's translation function, which was later fixed. CONCLUSION: Our results support that cohort discovery applications such as Leaf and SHRINE can interoperate in federated data networks with heterogeneous data models. Nicholas J. Dobbins, Michele Morris, Eugene Sadhu, Douglas MacFadden, Marc-Danie Nazaire, William Simons, Griffin M. Weber, Shawn N. Murphy, Shyam Visweswaran |
J. Am. Medical Informatics Assoc. | 9 |
| 2024 | Extraction of sleep information from clinical notes of Alzheimer's disease patients using natural language processingabstractOBJECTIVES: Alzheimer's disease (AD) is the most common form of dementia in the United States. Sleep is one of the lifestyle-related factors that has been shown critical for optimal cognitive function in old age. However, there is a lack of research studying the association between sleep and AD incidence. A major bottleneck for conducting such research is that the traditional way to acquire sleep information is time-consuming, inefficient, non-scalable, and limited to patients' subjective experience. We aim to automate the extraction of specific sleep-related patterns, such as snoring, napping, poor sleep quality, daytime sleepiness, night wakings, other sleep problems, and sleep duration, from clinical notes of AD patients. These sleep patterns are hypothesized to play a role in the incidence of AD, providing insight into the relationship between sleep and AD onset and progression. MATERIALS AND METHODS: A gold standard dataset is created from manual annotation of 570 randomly sampled clinical note documents from the adSLEEP, a corpus of 192 000 de-identified clinical notes of 7266 AD patients retrieved from the University of Pittsburgh Medical Center (UPMC). We developed a rule-based natural language processing (NLP) algorithm, machine learning models, and large language model (LLM)-based NLP algorithms to automate the extraction of sleep-related concepts, including snoring, napping, sleep problem, bad sleep quality, daytime sleepiness, night wakings, and sleep duration, from the gold standard dataset. RESULTS: The annotated dataset of 482 patients comprised a predominantly White (89.2%), older adult population with an average age of 84.7 years, where females represented 64.1%, and a vast majority were non-Hispanic or Latino (94.6%). Rule-based NLP algorithm achieved the best performance of F1 across all sleep-related concepts. In terms of positive predictive value (PPV), the rule-based NLP algorithm achieved the highest PPV scores for daytime sleepiness (1.00) and sleep duration (1.00), while the machine learning models had the highest PPV for napping (0.95) and bad sleep quality (0.86), and LLAMA2 with finetuning had the highest PPV for night wakings (0.93) and sleep problem (0.89). DISCUSSION: Although sleep information is infrequently documented in the clinical notes, the proposed rule-based NLP algorithm and LLM-based NLP algorithms still achieved promising results. In comparison, the machine learning-based approaches did not achieve good results, which is due to the small size of sleep information in the training data. CONCLUSION: The results show that the rule-based NLP algorithm consistently achieved the best performance for all sleep concepts. This study focused on the clinical notes of patients with AD but could be extended to general sleep information extraction for other diseases. Sonish Sivarajkumar, Thomas Yu Chow Tam, Haneef Ahamed Mohammad, Samuel Viggiano, David Oniani, Shyam Visweswaran, Yanshan Wang |
J. Am. Medical Informatics Assoc. | 6 |
| 2024 | Fairness and inclusion methods for biomedical informatics research
Shyam Visweswaran, Yuan Luo 0001, Mor Peleg |
J. Biomed. Informatics | 1 |
| 2024 | MedSyn: Text-Guided Anatomy-Aware Synthesis of High-Fidelity 3-D CT ImagesabstractThis paper introduces an innovative methodology for producing high-quality 3D lung CT images guided by textual information. While diffusion-based generative models are increasingly used in medical imaging, current state-of-the-art approaches are limited to low-resolution outputs and underutilize radiology reports' abundant information. The radiology reports can enhance the generation process by providing additional guidance and offering fine-grained control over the synthesis of images. Nevertheless, expanding text-guided generation to high-resolution 3D images poses significant memory and anatomical detail-preserving challenges. Addressing the memory issue, we introduce a hierarchical scheme that uses a modified UNet architecture. We start by synthesizing low-resolution images conditioned on the text, serving as a foundation for subsequent generators for complete volumetric data. To ensure the anatomical plausibility of the generated samples, we provide further guidance by generating vascular, airway, and lobular segmentation masks in conjunction with the CT images. The model demonstrates the capability to use textual input and segmentation tasks to generate synthesized images. Algorithmic comparative assessments and blind evaluations conducted by 10 board-certified radiologists indicate that our approach exhibits superior performance compared to the most advanced models based on GAN and diffusion techniques, especially in accurately retaining crucial anatomical features such as fissure lines and airways. This innovation introduces novel possibilities. This study focuses on two main objectives: (1) the development of a method for creating images based on textual prompts and anatomical components, and (2) the capability to generate new images conditioning on anatomical elements. The advancements in image generation can be applied to enhance numerous downstream tasks. Yanwu Xu 0003, Li Sun 0010, Wei Peng 0009, Shuyue Jia, Katelyn Morrison, Adam Perer, Afrooz Zandifar, Shyam Visweswaran, Motahhare Eslami, Kayhan Batmanghelich |
IEEE Trans. Medical Imaging | 8 |
| 2023 | A new method for estimating the probability of causal relationships from observational data: Application to the study of the short-term effects of air pollution on cardiovascular and respiratory disease
Bryan Andrews, Chirayu Wongchokprasitti, Shyam Visweswaran, Chirag M. Lakhani, Chirag J. Patel, Gregory F. Cooper |
Artif. Intell. Medicine | 3 |
| 2023 | A broadly applicable approach to enrich electronic-health-record cohorts by identifying patients with complete data: a multisite evaluationabstractOBJECTIVE: Patients who receive most care within a single healthcare system (colloquially called a "loyalty cohort" since they typically return to the same providers) have mostly complete data within that organization's electronic health record (EHR). Loyalty cohorts have low data missingness, which can unintentionally bias research results. Using proxies of routine care and healthcare utilization metrics, we compute a per-patient score that identifies a loyalty cohort. MATERIALS AND METHODS: We implemented a computable program for the widely adopted i2b2 platform that identifies loyalty cohorts in EHRs based on a machine-learning model, which was previously validated using linked claims data. We developed a novel validation approach, which tests, using only EHR data, whether patients returned to the same healthcare system after the training period. We evaluated these tools at 3 institutions using data from 2017 to 2019. RESULTS: Loyalty cohort calculations to identify patients who returned during a 1-year follow-up yielded a mean area under the receiver operating characteristic curve of 0.77 using the original model and 0.80 after calibrating the model at individual sites. Factors such as multiple medications or visits contributed significantly at all sites. Screening tests' contributions (eg, colonoscopy) varied across sites, likely due to coding and population differences. DISCUSSION: This open-source implementation of a "loyalty score" algorithm had good predictive power. Enriching research cohorts by utilizing these low-missingness patients is a way to obtain the data completeness necessary for accurate causal analysis. CONCLUSION: i2b2 sites can use this approach to select cohorts with mostly complete EHR data. Jeffrey G. Klann, Darren W. Henderson, Michele Morris, Hossein Estiri, Griffin M. Weber, Shyam Visweswaran, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 6 |
| 2023 | Informative missingness: What can we learn from patterns in missing laboratory data in the electronic health record?
Amelia L. M. Tan, Emily J. Getzen, Meghan Hutch, Zachary H. Strasser, Alba Gutiérrez-Sacristán, Trang T. Le, Arianna Dagliati, Michele Morris, David A. Hanauer, Bertrand Moal, Clara-Lea Bonzel, William Yuan, Lorenzo Chiudinelli, Priyam Das, Harrison G. Zhang, Bruce J. Aronow, Paul Avillach, Gabriel A. Brat, Tianxi Cai, Chuan Hong, William G. La Cava, He Hooi Will Loh, Yuan Luo 0001, Shawn N. Murphy, Kee Yuan Hgiam, Gilbert S. Omenn, Lav P. Patel, Malarkodi J. Samayamuthu, Emily R. Shriver, Zahra Shakeri Hossein Abad, Byorn W. L. Tan, Shyam Visweswaran, Griffin M. Weber, Zongqi Xia, Bertrand Verdy, Qi Long, Danielle L. Mowery, John H. Holmes |
J. Biomed. Informatics | 32 |
| 2023 | Special issue on fairness and inclusion in biomedical informatics research: technical and social perspectives
Shyam Visweswaran, Yuan Luo 0001, Mor Peleg |
J. Biomed. Informatics | 1 |
| 2022 | Delivering Real World Patient Data for Clinical and Translational Research: Approaches from Four Institutions
Christopher A. Harle, Daniella Meeker, Shyam Visweswaran, Thomas R. Campion Jr., Boyd M. Knosp |
AMIA | 3 |
| 2022 | Developing and Evaluation of Computational Phenotypes of Metastatic Breast Cancer Using All of Us Data
Israel Dilan-Pantojas, Shyam Visweswaran, Michael J. Becich, Xia Jiang, Richard D. Boyce |
AMIA | 3 |
| 2022 | Research data warehouse best practices: catalyzing national data sharing through informatics innovationabstractResearch Patient Data Repositories (RPDRs) have become essential infrastructure for traditional Clinical and Translational Science Award (CTSA) programs and increasingly for a wide range of research consortia and learning health system networks.1–5 Almost every institution with a CTSA or Clinical Translational Research (CTR) program (found in states with lower amounts of National Institutes of Health funding) hosts an RPDR for the benefit of affiliated researchers. These repositories aim to enable healthcare research based upon the patient populations they serve. Within the institution, RPDRs are valuable for a range of research activities. They are used to identify patients for clinical trial recruitment using privacy-preserving methods to search and extract specific cohorts of trial-eligible patients.6 They aid in developing and validating computable phenotypes that are increasingly important for accurately identifying patient cohorts in a reproducible fashion.7 RPDRs provide de-identified patient data for population health research and support a growing body of artificial intelligence to predict patient outcomes.8 Further, clinical studies can often be simulated using data from an RPDR.9 Beyond the institution, aggregates of de-identified datasets from multiple institutions linked with privacy-preserving hash codes provide an unprecedented opportunity to conduct population health research, perform comparative effectiveness analyses and apply artificial intelligence methods over large and diverse populations.10 The data contained within the RPDR vary across institutions, based on institutional strengths and weaknesses; the papers published in this issue reflect that variability (see Table 1). Data are commonly acquired from local electronic health records (EHRs) and other clinical information systems that capture information during clinical care. Data consist of diagnoses, problem lists, procedures, prescribed medications, laboratory exams, and many types of free-text reports. Overall, the benefits of the RPDR for accelerating translational research can be significant. For example, at Harvard, in 2006, between $94 and $136 million in annual research funding was linked to the use of data from the RPDR.11 Shawn N. Murphy, Shyam Visweswaran, Michael J. Becich, Thomas R. Campion Jr., Boyd M. Knosp, Genevieve B. Melton, Leslie Lenert |
J. Am. Medical Informatics Assoc. | 2 |
| 2022 | Synergies between centralized and federated approaches to data quality: a report from the national COVID cohort collaborativeabstractOBJECTIVE: In response to COVID-19, the informatics community united to aggregate as much clinical data as possible to characterize this new disease and reduce its impact through collaborative analytics. The National COVID Cohort Collaborative (N3C) is now the largest publicly available HIPAA limited dataset in US history with over 6.4 million patients and is a testament to a partnership of over 100 organizations. MATERIALS AND METHODS: We developed a pipeline for ingesting, harmonizing, and centralizing data from 56 contributing data partners using 4 federated Common Data Models. N3C data quality (DQ) review involves both automated and manual procedures. In the process, several DQ heuristics were discovered in our centralized context, both within the pipeline and during downstream project-based analysis. Feedback to the sites led to many local and centralized DQ improvements. RESULTS: Beyond well-recognized DQ findings, we discovered 15 heuristics relating to source Common Data Model conformance, demographics, COVID tests, conditions, encounters, measurements, observations, coding completeness, and fitness for use. Of 56 sites, 37 sites (66%) demonstrated issues through these heuristics. These 37 sites demonstrated improvement after receiving feedback. DISCUSSION: We encountered site-to-site differences in DQ which would have been challenging to discover using federated checks alone. We have demonstrated that centralized DQ benchmarking reveals unique opportunities for DQ improvement that will support improved research analytics locally and in aggregate. CONCLUSION: By combining rapid, continual assessment of DQ with a large volume of multisite data, it is possible to support more nuanced scientific questions with the scale and rigor that they require. Emily R. Pfaff, Andrew T. Girvin, Davera Gabriel, Kristin Kostka, Michele Morris, Matvey Palchuk, Harold P. Lehmann, Benjamin R. C. Amor, Mark Bissell, Katie R. Bradwell, Sigfried Gold, Stephanie S. Hong, Johanna Loomba, Amin Manna, Julie A. McMurry, Emily Niehaus, Nabeel Qureshi, Anita Walden, Xiaohan Tanner Zhang, Richard L. Zhu, Richard A. Moffitt, Christopher G. Chute, William G. Adams, Shaymaa Al-Shukri, Alfred Anzalone, Ahmad Baghal, Tellen D. Bennett, Elmer V. Bernstam, Mark M. Bissell, Brian Bush, Thomas R. Campion Jr., Victor Castro, Jack Chang, Deepa D. Chaudhari, Wenjin Chen, San Chu, James J. Cimino, Keith A. Crandall, Mark Crooks, Sara J. Deakyne Davies, John Dipalazzo, David A. Dorr, Daniel Eckrich, Sarah E. Eltinge, Daniel G. Fort, Georgiy Golovko, Snehil Gupta, Melissa A. Haendel, Janos G. Hajagos, David A. Hanauer, Brett M. Harnett, Ronald Horswell, Nancy Huang, Steven G. Johnson, Michael Kahn, Kamil Khanipov, Curtis Kieler, Katherine Ruiz De Luzuriaga, Sarah E. Maidlow, Ashley Martinez, Jomol Mathew, James C. McClay, Gabriel McMahan, Brian Melancon, Stéphane M. Meystre, Lucio Miele, Hiroki Morizono, Ray Pablo, Lav P. Patel, Jimmy Phuong, Daniel J. Popham, Claudia P. Pulgarin, Indra Neil Sarkar, Nancy Sazo, Soko Setoguchi, Selvin Soby, Sirisha Surampalli, Christine Suver, Uma Maheswara Reddy Vangala, Shyam Visweswaran, James von Oehsen, Kellie M. Walters, Laura K. Wiley, David A. Williams, Adrian H. Zai |
J. Am. Medical Informatics Assoc. | 81 |
| 2022 | An atomic approach to the design and implementation of a research data warehouseabstractOBJECTIVE: As a long-standing Clinical and Translational Science Awards (CTSA) Program hub, the University of Pittsburgh and the University of Pittsburgh Medical Center (UPMC) developed and implemented a modern research data warehouse (RDW) to efficiently provision electronic patient data for clinical and translational research. MATERIALS AND METHODS: We designed and implemented an RDW named Neptune to serve the specific needs of our CTSA. Neptune uses an atomic design where data are stored at a high level of granularity as represented in source systems. Neptune contains robust patient identity management tailored for research; integrates patient data from multiple sources, including electronic health records (EHRs), health plans, and research studies; and includes knowledge for mapping to standard terminologies. RESULTS: Neptune contains data for more than 5 million patients longitudinally organized as Health Insurance Portability and Accountability Act (HIPAA) Limited Data with dates and includes structured EHR data, clinical documents, health insurance claims, and research data. Neptune is used as a source for patient data for hundreds of institutional review board-approved research projects by local investigators and for national projects. DISCUSSION: The design of Neptune was heavily influenced by the large size of UPMC, the varied data sources, and the rich partnership between the University and the healthcare system. It includes several unique aspects, including the physical warehouse straddling the University and UPMC networks and management under an HIPAA Business Associates Agreement. CONCLUSION: We describe the design and implementation of an RDW at a large academic healthcare system that uses a distinctive atomic design where data are stored at a high level of granularity. Shyam Visweswaran, Brian McLay, Nickie Cappella, Michele Morris, John T. Milnes, Steven E. Reis, Jonathan C. Silverstein, Michael J. Becich |
J. Am. Medical Informatics Assoc. | 1 |
| 2022 | SurvMaximin: Robust federated approach to transporting survival risk prediction models
Harrison G. Zhang, Xin Xiong 0006, Chuan Hong, Griffin M. Weber, Gabriel A. Brat, Clara-Lea Bonzel, Yuan Luo 0001, Rui Duan 0004, Nathan P. Palmer, Meghan Hutch, Alba Gutiérrez-Sacristán, Riccardo Bellazzi, Luca Chiovato, Kelly Cho, Arianna Dagliati, Hossein Estiri, Noelia García-Barrio, Romain Griffier, David A. Hanauer, Yuk-Lam Ho, John H. Holmes, Mark S. Keller, Jeffrey G. Klann, Sehi L'Yi, Sara Lozano-Zahonero, Sarah E. Maidlow, Adeline Makoudjou, Alberto Malovini, Bertrand Moal, Jason H. Moore, Michele Morris, Danielle L. Mowery, Shawn N. Murphy, Antoine Neuraz, Kee Yuan Ngiam, Gilbert S. Omenn, Lav P. Patel, Miguel Pedrera-Jiménez, Andrea Prunotto, Malarkodi J. Samayamuthu, Fernando J. Sanz Vidorreta, Emily Schriver, Petra Schubert, Pablo Serrano-Balazote, Andrew M. South, Amelia L. M. Tan, Byorn W. L. Tan, Valentina Tibollo, Patric Tippmann, Shyam Visweswaran, Zongqi Xia, William Yuan, Daniela Zöller, Isaac S. Kohane, Paul Avillach, Zijian Guo 0003, Tianxi Cai |
J. Biomed. Informatics | 51 |
| 2021 | Using Distribution Divergence to Predict Changes in the Performance of Clinical Predictive Models
Mohammadamin Tajgardoon, Shyam Visweswaran |
AIME | 2 |
| 2021 | Comparison of Population-wide Explanations for Predicting the Outcomes of Patients with Community-Acquired Pneumonia
Eddie Pérez Claudio, Shyam Visweswaran, Harry Hochheiser |
AMIA | 2 |
| 2021 | Patient-Specific Modeling with Lazy Random Forest (LazyRF)
Adriana Johnson, Gregory F. Cooper, Shyam Visweswaran |
AMIA | 3 |
| 2021 | The National COVID Cohort Collaborative (N3C): Rationale, design, infrastructure, and deploymentabstractOBJECTIVE: Coronavirus disease 2019 (COVID-19) poses societal challenges that require expeditious data and knowledge sharing. Though organizational clinical data are abundant, these are largely inaccessible to outside researchers. Statistical, machine learning, and causal analyses are most successful with large-scale data beyond what is available in any given organization. Here, we introduce the National COVID Cohort Collaborative (N3C), an open science community focused on analyzing patient-level data from many centers. MATERIALS AND METHODS: The Clinical and Translational Science Award Program and scientific community created N3C to overcome technical, regulatory, policy, and governance barriers to sharing and harmonizing individual-level clinical data. We developed solutions to extract, aggregate, and harmonize data across organizations and data models, and created a secure data enclave to enable efficient, transparent, and reproducible collaborative analytics. RESULTS: Organized in inclusive workstreams, we created legal agreements and governance for organizations and researchers; data extraction scripts to identify and ingest positive, negative, and possible COVID-19 cases; a data quality assurance and harmonization pipeline to create a single harmonized dataset; population of the secure data enclave with data, machine learning, and statistical analytics tools; dissemination mechanisms; and a synthetic data pilot to democratize data access. CONCLUSIONS: The N3C has demonstrated that a multisite collaborative learning health network can overcome barriers to rapidly build a scalable infrastructure incorporating multiorganizational clinical data for COVID-19 analytics. We expect this effort to save lives by enabling rapid collaboration among clinicians, researchers, and data scientists to identify treatments and specialized care and thereby reduce the immediate and long-term impacts of COVID-19. Melissa A. Haendel, Christopher G. Chute, Tellen D. Bennett, David Eichmann, Justin Guinney, Warren A. Kibbe, Philip R. O. Payne, Emily R. Pfaff, Peter N. Robinson, Joel H. Saltz, Heidi Spratt, Christine Suver, John Wilbanks, Adam B. Wilcox, Andrew E. Williams, Chunlei Wu, Clair Blacketer, Robert L. Bradford, James J. Cimino, Marshall Clark, Evan W. Colmenares, Patricia A. Francis, Davera Gabriel, Alexis Graves, Raju Hemadri, Stephanie S. Hong, George Hripcsak, Dazhi Jiao, Jeffrey G. Klann, Kristin Kostka, Adam M. Lee, Harold P. Lehmann, Lora Lingrey, Robert T. Miller, Michele Morris, Shawn N. Murphy, Karthik Natarajan, Matvey Palchuk, Usman Sheikh, Harold R. Solbrig, Shyam Visweswaran, Anita Walden, Kellie M. Walters, Griffin M. Weber, Xiaohan Tanner Zhang, Richard L. Zhu, Benjamin R. C. Amor, Andrew T. Girvin, Amin Manna, Nabeel Qureshi, Michael G. Kurilla, Samuel G. Michael, Lili M. Portilla, Joni L. Rutter, Christopher P. Austin, Kenneth R. Gersing |
J. Am. Medical Informatics Assoc. | 41 |
| 2021 | Validation of an internationally derived patient severity phenotype to support COVID-19 analytics from electronic health record dataabstractOBJECTIVE: The Consortium for Clinical Characterization of COVID-19 by EHR (4CE) is an international collaboration addressing coronavirus disease 2019 (COVID-19) with federated analyses of electronic health record (EHR) data. We sought to develop and validate a computable phenotype for COVID-19 severity. MATERIALS AND METHODS: Twelve 4CE sites participated. First, we developed an EHR-based severity phenotype consisting of 6 code classes, and we validated it on patient hospitalization data from the 12 4CE clinical sites against the outcomes of intensive care unit (ICU) admission and/or death. We also piloted an alternative machine learning approach and compared selected predictors of severity with the 4CE phenotype at 1 site. RESULTS: The full 4CE severity phenotype had pooled sensitivity of 0.73 and specificity 0.83 for the combined outcome of ICU admission and/or death. The sensitivity of individual code categories for acuity had high variability-up to 0.65 across sites. At one pilot site, the expert-derived phenotype had mean area under the curve of 0.903 (95% confidence interval, 0.886-0.921), compared with an area under the curve of 0.956 (95% confidence interval, 0.952-0.959) for the machine learning approach. Billing codes were poor proxies of ICU admission, with as low as 49% precision and recall compared with chart review. DISCUSSION: We developed a severity phenotype using 6 code classes that proved resilient to coding variability across international institutions. In contrast, machine learning approaches may overfit hospital-specific orders. Manual chart review revealed discrepancies even in the gold-standard outcomes, possibly owing to heterogeneous pandemic conditions. CONCLUSIONS: We developed an EHR-based severity phenotype for COVID-19 in hospitalized patients and validated it at 12 international sites. Jeffrey G. Klann, Hossein Estiri, Griffin M. Weber, Bertrand Moal, Paul Avillach, Chuan Hong, Amelia L. M. Tan, Brett K. Beaulieu-Jones, Victor M. Castro, Thomas Maulhardt, Alon Geva, Alberto Malovini, Andrew M. South, Shyam Visweswaran, Michele Morris, Malarkodi J. Samayamuthu, Gilbert S. Omenn, Kee Yuan Ngiam, Kenneth D. Mandl, Martin Boeker, Karen L. Olson, Danielle L. Mowery, Robert W. Follett, David A. Hanauer, Riccardo Bellazzi, Jason H. Moore, Ne-Hooi Will Loh, Douglas S. Bell, Kavishwar B. Wagholikar, Luca Chiovato, Valentina Tibollo, Siegbert Rieg, Anthony L. L. J. Li, Vianney Jouhet, Emily Schriver, Zongqi Xia, Meghan Hutch, Yuan Luo 0001, Isaac S. Kohane, Gabriel A. Brat, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 14 |
| 2021 | Predicting outcomes in central venous catheter salvage in pediatric central line-associated bloodstream infectionabstractOBJECTIVE: Central line-associated bloodstream infections (CLABSIs) are a common, costly, and hazardous healthcare-associated infection in children. In children in whom continued access is critical, salvage of infected central venous catheters (CVCs) with antimicrobial lock therapy is an alternative to removal and replacement of the CVC. However, the success of CVC salvage is uncertain, and when it fails the catheter has to be removed and replaced. We describe a machine learning approach to predict individual outcomes in CVC salvage that can aid the clinician in the decision to attempt salvage. MATERIALS AND METHODS: Over a 14-year period, 969 pediatric CLABSIs were identified in electronic health records. We used 164 potential predictors to derive 4 types of machine learning models to predict 2 failed salvage outcomes, infection recurrence and CVC removal, at 10 time points between 7 days and 1 year from infection onset. RESULTS: The area under the receiver-operating characteristic curve varied from 0.56 to 0.83, and key predictors varied over time. The infection recurrence model performed better than the CVC removal model did. CONCLUSIONS: Machine learning-based outcome prediction can inform clinical decision making for children. We developed and evaluated several models to predict clinically relevant outcomes in the context of CVC salvage in pediatric CLABSI and illustrate the variability of predictors over time. Lorne W. Walker, Andrew J. Nowalk, Shyam Visweswaran |
J. Am. Medical Informatics Assoc. | 3 |
| 2020 | Patient-Specific Modeling with Personalized Decision Paths
Adriana Johnson, Gregory F. Cooper, Shyam Visweswaran |
AMIA | 3 |
| 2019 | Team-Centered Informatics: A Necessary Adaptation to Translational and Implementation Science?
Suresh K. Bhavnani, Shyam Visweswaran, Erich Kummerfeld, Carlos Clark, Rebekah Penton |
AMIA | 2 |
| 2019 | Insights from a dissertation on the development of a Learning Electronic Medical Record System: data-driven, context-aware learning
Andrew J. King 0002, Shyam Visweswaran, Harry Hochheiser, Gilles Clermont, Gregory F. Cooper |
AMIA | 2 |
| 2019 | Curating EHR data in the All of Us Research Program
Karthik Natarajan, Robert J. Carroll, Thomas R. Campion Jr., Joan Grand, Shyam Visweswaran |
AMIA | 5 |
| 2019 | An Empirical Investigation of Instance-Specific Causal Bayesian Network LearningabstractSignificant progress has been made in developing algorithms for learning graphical causal models from data. Most of these algorithms learn a causal structure that is shared by all the instances (e.g., patients) in the training dataset. However, different instances may not all share the same causal structure. We introduced an instance-specific method called IGES [15] that learns a causal model for each instance T by using the features of T and the instances in the training dataset. In the current paper, we study the empirical performance of the IGES method on several biomedical datasets. The results provide support that instance-specific structure exists and is important to model in these real domains. Fattaneh Jabbari, Shyam Visweswaran, Gregory F. Cooper |
BIBM | 2 |
| 2019 | Translational bioinformatics in mental health: open access data sources and computational biomarker discoveryabstractMental illness is increasingly recognized as both a significant cost to society and a significant area of opportunity for biological breakthrough. As -omics and imaging technologies enable researchers to probe molecular and physiological underpinnings of multiple diseases, opportunities arise to explore the biological basis for behavioral health and disease. From individual investigators to large international consortia, researchers have generated rich data sets in the area of mental health, including genomic, transcriptomic, metabolomic, proteomic, clinical and imaging resources. General data repositories such as the Gene Expression Omnibus (GEO) and Database of Genotypes and Phenotypes (dbGaP) and mental health (MH)-specific initiatives, such as the Psychiatric Genomics Consortium, MH Research Network and PsychENCODE represent a wealth of information yet to be gleaned. At the same time, novel approaches to integrate and analyze data sets are enabling important discoveries in the area of mental and behavioral health. This review will discuss and catalog into an organizing framework the increasingly diverse set of MH data resources available, using schizophrenia as a focus area, and will describe novel and integrative approaches to molecular biomarker discovery that make use of mental health data. Jessica D. Tenenbaum, Krithika Bhuvaneshwar, Jane P. Gagliardi, Kate Fultz Hollis, Peilin Jia, Radhakrishnan Nagarajan, Gopalkumar Rakesh, Vignesh Subbian, Shyam Visweswaran, Zhongming Zhao, Leon Rozenblit |
Briefings Bioinform. | 10 |
| 2019 | Using machine learning to selectively highlight patient information
Andrew J. King 0002, Gregory F. Cooper, Gilles Clermont, Harry Hochheiser, Milos Hauskrecht, Dean F. Sittig, Shyam Visweswaran |
J. Biomed. Informatics | 7 |
| 2019 | Estimating and Controlling the False Discovery Rate of the PC Algorithm Using Edge-specific P-ValuesabstractMany causal discovery algorithms infer graphical structure from observational data. The PC algorithm in particular estimates a completed partially directed acyclic graph (CPDAG), or an acyclic graph containing directed edges identifiable with conditional independence testing. However, few groups have investigated strategies for estimating and controlling the false discovery rate (FDR) of the edges in the CPDAG. In this article, we introduce PC with p-values (PC-p), a fast algorithm that robustly computes edge-specific p-values and then estimates and controls the FDR across the edges. PC-p specifically uses the p-values returned by many conditional independence (CI) tests to upper bound the p-values of more complex edge-specific hypothesis tests. The algorithm then estimates and controls the FDR using the bounded p-values and the Benjamini-Yekutieli FDR procedure. Modifications to the original PC algorithm also help PC-p accurately compute the upper bounds despite non-zero Type II error rates. Experiments show that PC-p yields more accurate FDR estimation and control across the edges in a variety of CPDAGs compared to alternative methods. Eric V. Strobl, Peter Spirtes, Shyam Visweswaran |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2018 | Utility of Visual Analytics for Identifying Patient Subgroups in EMRs: Insights for Accelerating Precision Medicine
Suresh K. Bhavnani, Revathi Sellappan, Jonathan Starkey, Winston Chan, Tianlong Chen 0004, Shyam Visweswaran |
AMIA | 7 |
| 2018 | Software Package to Load Data from REDCap to PCORnet CDM 4.0
Charles D. Borromeo, William Shirey, Nickie Cappella, Shyam Visweswaran, Jonathan C. Silverstein, Michael J. Becich |
AMIA | 4 |
| 2018 | Workflow for Developing i2b2 Ontologies from Source Terminologies in ACT
Charles D. Borromeo, William Shirey, Michele Morris, Malarkodi J. Samayamuthu, Shyam Visweswaran |
AMIA | 6 |
| 2018 | Design of a Learning Electronic Medical Record: A Qualitative Study of ICU Clinicians' Information Needs and Practices
Luca Calzoni, Gilles Clermont, Gregory F. Cooper, Shyam Visweswaran, Harry Hochheiser |
AMIA | 4 |
| 2018 | Using Machine Learning to Predict the Information Seeking Behavior of Clinicians Using an Electronic Medical Record System
Andrew J. King 0002, Gregory F. Cooper, Harry Hochheiser, Gilles Clermont, Milos Hauskrecht, Shyam Visweswaran |
AMIA | 6 |
| 2018 | Social Context Sentence Classification from Psychiatric Reports using Positive and Unlabeled Learning
José D. Posada, Lingyun Shi, Sergio M. Castro, Shyam Visweswaran, Neal D. Ryan, Henk Harkema, Rich Tsui |
AMIA | 4 |
| 2018 | A Computable Phenotype Library Plugin for i2b2
Michele Morris, Shyam Visweswaran |
AMIA | 3 |
| 2017 | Vicinity Exploration: Enabling User-Driven Visual Search of Multiple Machine Learning Models for Precision Medicine
Suresh K. Bhavnani, Archana Ayyaswami, Tianlong Chen 0004, Shyam Visweswaran, Gowtham Bellala, Kevin E. Bassler |
AMIA | 4 |
| 2017 | Exploring Novel Graphical Representations of Clinical Data in a Learning EMR
Luca Calzoni, Gilles Clermont, Gregory F. Cooper, Shyam Visweswaran, Harry Hochheiser |
AMIA | 4 |
| 2017 | Automated annotation and classification of BI-RADS assessment from radiology reportsabstractThe Breast Imaging Reporting and Data System (BI-RADS) was developed to reduce variation in the descriptions of findings. Manual analysis of breast radiology report data is challenging but is necessary for clinical and healthcare quality assurance activities. The objective of this study is to develop a natural language processing (NLP) system for automated BI-RADS categories extraction from breast radiology reports. We evaluated an existing rule-based NLP algorithm, and then we developed and evaluated our own method using a supervised machine learning approach. We divided the BI-RADS category extraction task into two specific tasks: (1) annotation of all BI-RADS category values within a report, (2) classification of the laterality of each BI-RADS category value. We used one algorithm for task 1 and evaluated three algorithms for task 2. Across all evaluations and model training, we used a total of 2159 radiology reports from 18 hospitals, from 2003 to 2015. Performance with the existing rule-based algorithm was not satisfactory. Conditional random fields showed a high performance for task 1 with an F-1 measure of 0.95. Rules from partial decision trees (PART) algorithm showed the best performance across classes for task 2 with a weighted F-1 measure of 0.91 for BIRADS 0-6, and 0.93 for BIRADS 3-5. Classification performance by class showed that performance improved for all classes from Naïve Bayes to Support Vector Machine (SVM), and also from SVM to PART. Our system is able to annotate and classify all BI-RADS mentions present in a single radiology report and can serve as the foundation for future studies that will leverage automated BI-RADS annotation, to provide feedback to radiologists as part of a learning health system loop. Sergio M. Castro, Eugene Tseytlin, Olga Medvedeva, Kevin J. Mitchell, Shyam Visweswaran, Tanja Bekhuis, Rebecca S. Jacobson |
J. Biomed. Informatics | 5 |
| 2016 | An informatics research agenda to support precision medicine: seven key areasabstractThe recent announcement of the Precision Medicine Initiative by President Obama has brought precision medicine (PM) to the forefront for healthcare providers, researchers, regulators, innovators, and funders alike. As technologies continue to evolve and datasets grow in magnitude, a strong computational infrastructure will be essential to realize PM's vision of improved healthcare derived from personal data. In addition, informatics research and innovation affords a tremendous opportunity to drive the science underlying PM. The informatics community must lead the development of technologies and methodologies that will increase the discovery and application of biomedical knowledge through close collaboration between researchers, clinicians, and patients. This perspective highlights seven key areas that are in need of further informatics research and innovation to support the realization of PM. Jessica D. Tenenbaum, Paul Avillach, Marge M. Benham-Hutchins, Matthew K. Breitenstein, Erin L. Crowgey, Mark A. Hoffman, Xia Jiang, Subha Madhavan, John E. Mattison, Radhakrishnan Nagarajan, Bisakha Ray, Dmitriy Shin, Shyam Visweswaran, Zhongming Zhao, Robert R. Freimuth |
J. Am. Medical Informatics Assoc. | 13 |
| 2016 | Outlier-based detection of unusual patient-management actions: An ICU study
Milos Hauskrecht, Iyad Batal, Charmgil Hong, Gregory F. Cooper, Shyam Visweswaran, Gilles Clermont |
J. Biomed. Informatics | 6 |
| 2015 | Development and Preliminary Evaluation of a Prototype of a Learning Electronic Medical Record System
Andrew J. King 0002, Gregory F. Cooper, Harry Hochheiser, Gilles Clermont, Shyam Visweswaran |
AMIA | 5 |
| 2015 | Knowledge transfer via classification rules using functional mapping for integrative modeling of gene expression dataabstractBACKGROUND: Most 'transcriptomic' data from microarrays are generated from small sample sizes compared to the large number of measured biomarkers, making it very difficult to build accurate and generalizable disease state classification models. Integrating information from different, but related, 'transcriptomic' data may help build better classification models. However, most proposed methods for integrative analysis of 'transcriptomic' data cannot incorporate domain knowledge, which can improve model performance. To this end, we have developed a methodology that leverages transfer rule learning and functional modules, which we call TRL-FM, to capture and abstract domain knowledge in the form of classification rules to facilitate integrative modeling of multiple gene expression data. TRL-FM is an extension of the transfer rule learner (TRL) that we developed previously. The goal of this study was to test our hypothesis that "an integrative model obtained via the TRL-FM approach outperforms traditional models based on single gene expression data sources". RESULTS: To evaluate the feasibility of the TRL-FM framework, we compared the area under the ROC curve (AUC) of models developed with TRL-FM and other traditional methods, using 21 microarray datasets generated from three studies on brain cancer, prostate cancer, and lung disease, respectively. The results show that TRL-FM statistically significantly outperforms TRL as well as traditional models based on single source data. In addition, TRL-FM performed better than other integrative models driven by meta-analysis and cross-platform data merging. CONCLUSIONS: The capability of utilizing transferred abstract knowledge derived from source data using feature mapping enables the TRL-FM framework to mimic the human process of learning and adaptation when performing related tasks. The novel TRL-FM methodology for integrative modeling for multiple 'transcriptomic' datasets is able to intelligently incorporate domain knowledge that traditional methods might disregard, to boost predictive power and generalization performance. In this study, TRL-FM's abstraction of knowledge is achieved in the form of functional modules, but the overall framework is generalizable in that different approaches of acquiring abstract knowledge can be integrated into this framework. Henry Ogoe, Shyam Visweswaran, Xinghua Lu 0001, Vanathi Gopalakrishnan |
BMC Bioinform. | 2 |
| 2015 | Comparison of machine learning classifiers for influenza detection from emergency department free-text reports
Arturo L. Pineda, Ye Ye 0002, Shyam Visweswaran, Gregory F. Cooper, Michael M. Wagner 0001, Fu-Chiang Tsui |
J. Biomed. Informatics | 3 |
| 2013 | Decision Path Models for Patient-Specific Modeling of Patient Outcomes
Antonio Luiz S. Ferreira, Gregory F. Cooper, Shyam Visweswaran |
AMIA | 3 |
| 2013 | Data-driven identification of unusual clinical actions in the ICU
Milos Hauskrecht, Shyam Visweswaran, Gregory F. Cooper, Gilles Clermont |
AMIA | 2 |
| 2013 | Deep Multiple Kernel LearningabstractDeep learning methods have predominantly been applied to large artificial neural networks. Despite their state-of-the-art performance, these large networks typically do not generalize well to datasets with limited sample sizes. In this paper, we take a different approach by learning multiple layers of kernels. We combine kernels at each layer and then optimize over an estimate of the support vector machine leave-one-out error rather than the dual objective function. Our experiments on a variety of datasets show that each layer successively increases performance with only a few base kernels. Eric V. Strobl, Shyam Visweswaran |
ICMLA (1) | 2 |
| 2013 | Outlier detection for patient monitoring and alerting
Milos Hauskrecht, Iyad Batal, Michal Valko, Shyam Visweswaran, Gregory F. Cooper, Gilles Clermont |
J. Biomed. Informatics | 4 |
| 2012 | Network Label Propagation Yields Reproducible Biomarker SNPs in Alzheimer's Datasets
Matthew E. Stokes, Shyam Visweswaran |
AMIA | 2 |
| 2012 | The role of complementary bipartite visual analytical representations in the analysis of SNPs: a case study in ancestral informative markersabstractOBJECTIVE: Several studies have shown how sets of single-nucleotide polymorphisms (SNPs) can help to classify subjects on the basis of their continental origins, with applications to case-control studies and population genetics. However, most of these studies use dimensionality-reduction methods, such as principal component analysis, or clustering methods that result in unipartite (either subjects or SNPs) representations of the data. Such analyses conceal important bipartite relationships, such as how subject and SNP clusters relate to each other, and the genotypes that determine their cluster memberships. METHODS: To overcome the limitations of current methods of analyzing SNP data, the authors used three bipartite analytical representations (bipartite network, heat map with dendrograms, and Circos ideogram) that enable the simultaneous visualization and analysis of subjects, SNPs, and subject attributes. RESULTS: The results demonstrate (1) novel insights into SNP data that are difficult to derive from purely unipartite views of the data, (2) the strengths and limitations of each method, revealing the role that each play in revealing novel insights, and (3) implications for how the methods can be used for the analysis of SNPs in genomic studies associated with disease. CONCLUSION: The results suggest that bipartite representations can reveal new patterns in SNP data compared with existing unipartite representations. However, the novel insights require multiple representations to discover, verify, and comprehend the complex relationships. The results therefore motivate the need for a complementary visual analytical framework that guides the use of multiple bipartite representations to analyze complex relationships in SNP data. Suresh K. Bhavnani, Gowtham Bellala, Sundar Victor, Kevin E. Bassler, Shyam Visweswaran |
J. Am. Medical Informatics Assoc. | 5 |
| 2012 | Building an automated SOAP classifier for emergency department reports
Danielle L. Mowery, Janyce Wiebe, Shyam Visweswaran, Henk Harkema, Wendy W. Chapman |
J. Biomed. Informatics | 3 |
| 2011 | Learning genetic epistasis using Bayesian network scoring criteriaabstractBACKGROUND: Gene-gene epistatic interactions likely play an important role in the genetic basis of many common diseases. Recently, machine-learning and data mining methods have been developed for learning epistatic relationships from data. A well-known combinatorial method that has been successfully applied for detecting epistasis is Multifactor Dimensionality Reduction (MDR). Jiang et al. created a combinatorial epistasis learning method called BNMBL to learn Bayesian network (BN) epistatic models. They compared BNMBL to MDR using simulated data sets. Each of these data sets was generated from a model that associates two SNPs with a disease and includes 18 unrelated SNPs. For each data set, BNMBL and MDR were used to score all 2-SNP models, and BNMBL learned significantly more correct models. In real data sets, we ordinarily do not know the number of SNPs that influence phenotype. BNMBL may not perform as well if we also scored models containing more than two SNPs. Furthermore, a number of other BN scoring criteria have been developed. They may detect epistatic interactions even better than BNMBL.Although BNs are a promising tool for learning epistatic relationships from data, we cannot confidently use them in this domain until we determine which scoring criteria work best or even well when we try learning the correct model without knowledge of the number of SNPs in that model. RESULTS: We evaluated the performance of 22 BN scoring criteria using 28,000 simulated data sets and a real Alzheimer's GWAS data set. Our results were surprising in that the Bayesian scoring criterion with large values of a hyperparameter called α performed best. This score performed better than other BN scoring criteria and MDR at recall using simulated data sets, at detecting the hardest-to-detect models using simulated data sets, and at substantiating previous results using the real Alzheimer's data set. CONCLUSIONS: We conclude that representing epistatic interactions using BN models and scoring them using a BN scoring criterion holds promise for identifying epistatic genetic variants in data. In particular, the Bayesian scoring criterion with large values of a hyperparameter α appears more promising than a number of alternatives. Xia Jiang, Richard E. Neapolitan, M. Michael Barmada, Shyam Visweswaran |
BMC Bioinform. | 4 |
| 2011 | Application of an efficient Bayesian discretization method to biomedical dataabstractBACKGROUND: Several data mining methods require data that are discrete, and other methods often perform better with discrete data. We introduce an efficient Bayesian discretization (EBD) method for optimal discretization of variables that runs efficiently on high-dimensional biomedical datasets. The EBD method consists of two components, namely, a Bayesian score to evaluate discretizations and a dynamic programming search procedure to efficiently search the space of possible discretizations. We compared the performance of EBD to Fayyad and Irani's (FI) discretization method, which is commonly used for discretization. RESULTS: On 24 biomedical datasets obtained from high-throughput transcriptomic and proteomic studies, the classification performances of the C4.5 classifier and the naïve Bayes classifier were statistically significantly better when the predictor variables were discretized using EBD over FI. EBD was statistically significantly more stable to the variability of the datasets than FI. However, EBD was less robust, though not statistically significantly so, than FI and produced slightly more complex discretizations than FI. CONCLUSIONS: On a range of biomedical datasets, a Bayesian discretization method (EBD) yielded better classification performance and stability but was less robust than the widely used FI discretization method. The EBD discretization method is easy to implement, permits the incorporation of prior knowledge and belief, and is sufficiently fast for application to high-dimensional data. Jonathan L. Lustgarten, Shyam Visweswaran, Vanathi Gopalakrishnan, Gregory F. Cooper |
BMC Bioinform. | 2 |
| 2011 | The application of naive Bayes model averaging to predict Alzheimer's disease from genome-wide dataabstractOBJECTIVE: Predicting patient outcomes from genome-wide measurements holds significant promise for improving clinical care. The large number of measurements (eg, single nucleotide polymorphisms (SNPs)), however, makes this task computationally challenging. This paper evaluates the performance of an algorithm that predicts patient outcomes from genome-wide data by efficiently model averaging over an exponential number of naive Bayes (NB) models. DESIGN: This model-averaged naive Bayes (MANB) method was applied to predict late onset Alzheimer's disease in 1411 individuals who each had 312,318 SNP measurements available as genome-wide predictive features. Its performance was compared to that of a naive Bayes algorithm without feature selection (NB) and with feature selection (FSNB). MEASUREMENT: Performance of each algorithm was measured in terms of area under the ROC curve (AUC), calibration, and run time. RESULTS: The training time of MANB (16.1 s) was fast like NB (15.6 s), while FSNB (1684.2 s) was considerably slower. Each of the three algorithms required less than 0.1 s to predict the outcome of a test case. MANB had an AUC of 0.72, which is significantly better than the AUC of 0.59 by NB (p<0.00001), but not significantly different from the AUC of 0.71 by FSNB. MANB was better calibrated than NB, and FSNB was even better in calibration. A limitation was that only one dataset and two comparison algorithms were included in this study. CONCLUSION: MANB performed comparatively well in predicting a clinical outcome from a high-dimensional genome-wide dataset. These results provide support for including MANB in the methods used to predict outcomes from large, genome-wide datasets. Shyam Visweswaran, Gregory F. Cooper |
J. Am. Medical Informatics Assoc. | 2 |
| 2010 | Bayesian rule learning for biomedical data miningabstractMOTIVATION: Disease state prediction from biomarker profiling studies is an important problem because more accurate classification models will potentially lead to the discovery of better, more discriminative markers. Data mining methods are routinely applied to such analyses of biomedical datasets generated from high-throughput 'omic' technologies applied to clinical samples from tissues or bodily fluids. Past work has demonstrated that rule models can be successfully applied to this problem, since they can produce understandable models that facilitate review of discriminative biomarkers by biomedical scientists. While many rule-based methods produce rules that make predictions under uncertainty, they typically do not quantify the uncertainty in the validity of the rule itself. This article describes an approach that uses a Bayesian score to evaluate rule models. RESULTS: We have combined the expressiveness of rules with the mathematical rigor of Bayesian networks (BNs) to develop and evaluate a Bayesian rule learning (BRL) system. This system utilizes a novel variant of the K2 algorithm for building BNs from the training data to provide probabilistic scores for IF-antecedent-THEN-consequent rules using heuristic best-first search. We then apply rule-based inference to evaluate the learned models during 10-fold cross-validation performed two times. The BRL system is evaluated on 24 published 'omic' datasets, and on average it performs on par or better than other readily available rule learning methods. Moreover, BRL produces models that contain on average 70% fewer variables, which means that the biomarker panels for disease prediction contain fewer markers for further verification and validation by bench scientists. Vanathi Gopalakrishnan, Jonathan L. Lustgarten, Shyam Visweswaran, Gregory F. Cooper |
Bioinform. | 3 |
| 2010 | Learning patient-specific predictive models from clinical data
Shyam Visweswaran, Derek C. Angus, Margaret Hsieh, Lisa A. Weissfeld, Donald Yealy, Gregory F. Cooper |
J. Biomed. Informatics | 1 |
| 2010 | Learning Instance-Specific Predictive Models
Shyam Visweswaran, Gregory F. Cooper |
J. Mach. Learn. Res. | 1 |
| 2009 | Measuring Stability of Feature Selection in Biomedical Datasets
Jonathan L. Lustgarten, Vanathi Gopalakrishnan, Shyam Visweswaran |
AMIA | 3 |
| 2009 | A Bayesian Method for Identifying Genetic Interactions
Shyam Visweswaran, An-Kwok Ian Wong, M. Michael Barmada |
AMIA | 1 |
| 2009 | Knowledge-based variable selection for learning rules from proteomic dataabstractBACKGROUND: The incorporation of biological knowledge can enhance the analysis of biomedical data. We present a novel method that uses a proteomic knowledge base to enhance the performance of a rule-learning algorithm in identifying putative biomarkers of disease from high-dimensional proteomic mass spectral data. In particular, we use the Empirical Proteomics Ontology Knowledge Base (EPO-KB) that contains previously identified and validated proteomic biomarkers to select m/zs in a proteomic dataset prior to analysis to increase performance. RESULTS: We show that using EPO-KB as a pre-processing method, specifically selecting all biomarkers found only in the biofluid of the proteomic dataset, reduces the dimensionality by 95% and provides a statistically significantly greater increase in performance over no variable selection and random variable selection. CONCLUSION: Knowledge-based variable selection even with a sparsely-populated resource such as the EPO-KB increases overall performance of rule-learning for disease classification from high-dimensional proteomic mass spectra. Jonathan L. Lustgarten, Shyam Visweswaran, Robert P. Bowser, William R. Hogan, Vanathi Gopalakrishnan |
BMC Bioinform. | 2 |
| 2008 | Assessing the Performance Characteristics of Signals Used by a Clinical Event Monitor to Detect Adverse Drug Reactions in the Nursing Home
Steven M. Handler, Joseph T. Hanlon, Subashan Perera, Melissa I. Saul, Douglas B. Fridsma, Shyam Visweswaran, Stephanie A. Studenski, Yazan F. Roumani, Nicholas G. Castle, David A. Nace, Michael J. Becich |
AMIA | 6 |
| 2008 | Improving Classification Performance with Discretization on Biomedical Datasets
Jonathan L. Lustgarten, Vanathi Gopalakrishnan, Himanshu Grover, Shyam Visweswaran |
AMIA | 4 |
| 2008 | Analysis of a Failed Clinical Decision Support System for Management of Congestive Heart Failure
Rajiv Wadhwa, Douglas B. Fridsma, Melissa I. Saul, Louis E. Penrod, Shyam Visweswaran, Gregory F. Cooper, Wendy W. Chapman |
AMIA | 5 |
| 2007 | Evidence-based Anomaly Detection in Clinical Domains
Milos Hauskrecht, Michal Valko, Branislav Kveton, Shyam Visweswaran, Gregory F. Cooper |
AMIA | 4 |
| 2005 | Deriving the Expected Utility of a Predictive Model When the Utilities Are Uncertain
Gregory F. Cooper, Shyam Visweswaran |
AMIA | 2 |
| 2005 | Patient-Specific Models for Predicting the Outcomes of Patients with Community Acquired Pneumonia
Shyam Visweswaran, Gregory F. Cooper |
AMIA | 1 |
| 2004 | Instance-Specific Bayesian Model Averaging for ClassificationabstractClassification algorithms typically induce population-wide models that are trained to perform well on average on expected future instances. We introduce a Bayesian framework for learning instance-specific models from data that are optimized to predict well for a particular instance. Based on this framework, we present a that performs selective model averaging over a restricted class of Bayesian networks. On experimental evaluation, this algorithm shows superior performance over model selection. We intend to apply such instance-specific algorithms to improve the performance of patient-specific predictive models induced from medical data. instance-specific algorithm called ISA Shyam Visweswaran, Gregory F. Cooper |
NIPS | 1 |
| 2003 | Detecting Adverse Drug Events in Discharge Summaries Using Variations on the Simple Bayes Model
Shyam Visweswaran, Paul Hanbury, Melissa I. Saul, Gregory F. Cooper |
AMIA | 1 |