VLDB 2026 Research / reviewers in the wild / expert
Michael J. Becich
dblp:76/5057
· DBLP profile ↗
22ranked-venue papers
1as first author
5since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | REDCap and the National Mesothelioma Virtual Bank - a scalable and sustainable model for rare disease biorepositoriesabstractOBJECTIVE: Rare disease research requires data sharing networks to power translational studies. We describe novel use of Research Electronic Data Capture (REDCap), a web application for managing clinical data, by the National Mesothelioma Virtual Bank, a federated biospecimen, and data sharing network. MATERIALS AND METHODS: National Mesothelioma Virtual Bank (NMVB) uses REDCap to integrate honest broker activities, enabling biospecimen and associated clinical data provisioning to investigators. A Web Portal Query tool was developed to source and visualize REDCap data in interactive, faceted search, enabling cohort discovery by public users. An AWS Lambda function behind an API calculates the counts visually presented, while protecting record level data. The user-friendly interface, quick responsiveness, automatic generation from REDCap, and flexibility to new data, was engineered to sustain the NMVB research community. RESULTS: NMVB implementations enabled a network of 8 research institutions with over 2000 mesothelioma cases, including clinical annotations and biospecimens, and public users' cohort discovery and summary statistics. NMVB usage and impact is demonstrated by high website visits (>150 unique queries per month), resource use requests (>50 letter of interests), and citations (>900) to papers published using NMVB resources. DISCUSSION: NMVB's REDCap implementation and query tool is a framework for implementing federated and integrated rare disease biobanks and registries. Advantages of this framework include being low-cost, modular, scalable, and efficient. Future advances to NVMB's implementations will include incorporation of -omics data and development of downstream analysis tools to advance mesothelioma and rare disease research. CONCLUSION: NVMB presents a framework for integrating biobanks and patient registries to enable translational research for rare diseases. Rumana Rashid, Susan Copelli, Jonathan C. Silverstein, Michael J. Becich |
J. Am. Medical Informatics Assoc. | 4 |
| 2022 | Developing and Evaluation of Computational Phenotypes of Metastatic Breast Cancer Using All of Us Data
Israel Dilan-Pantojas, Shyam Visweswaran, Michael J. Becich, Xia Jiang, Richard D. Boyce |
AMIA | 4 |
| 2022 | Research Patient Data Repositories: Perspectives from JAMIA Special Issue Editors on the Next Generation of Multi-Institutional Data Sharing
Genevieve B. Melton, Leslie Lenert, Michael J. Becich, Shawn N. Murphy, Thomas R. Campion Jr. |
AMIA | 3 |
| 2022 | Research data warehouse best practices: catalyzing national data sharing through informatics innovationabstractResearch Patient Data Repositories (RPDRs) have become essential infrastructure for traditional Clinical and Translational Science Award (CTSA) programs and increasingly for a wide range of research consortia and learning health system networks.1–5 Almost every institution with a CTSA or Clinical Translational Research (CTR) program (found in states with lower amounts of National Institutes of Health funding) hosts an RPDR for the benefit of affiliated researchers. These repositories aim to enable healthcare research based upon the patient populations they serve. Within the institution, RPDRs are valuable for a range of research activities. They are used to identify patients for clinical trial recruitment using privacy-preserving methods to search and extract specific cohorts of trial-eligible patients.6 They aid in developing and validating computable phenotypes that are increasingly important for accurately identifying patient cohorts in a reproducible fashion.7 RPDRs provide de-identified patient data for population health research and support a growing body of artificial intelligence to predict patient outcomes.8 Further, clinical studies can often be simulated using data from an RPDR.9 Beyond the institution, aggregates of de-identified datasets from multiple institutions linked with privacy-preserving hash codes provide an unprecedented opportunity to conduct population health research, perform comparative effectiveness analyses and apply artificial intelligence methods over large and diverse populations.10 The data contained within the RPDR vary across institutions, based on institutional strengths and weaknesses; the papers published in this issue reflect that variability (see Table 1). Data are commonly acquired from local electronic health records (EHRs) and other clinical information systems that capture information during clinical care. Data consist of diagnoses, problem lists, procedures, prescribed medications, laboratory exams, and many types of free-text reports. Overall, the benefits of the RPDR for accelerating translational research can be significant. For example, at Harvard, in 2006, between $94 and $136 million in annual research funding was linked to the use of data from the RPDR.11 Shawn N. Murphy, Shyam Visweswaran, Michael J. Becich, Thomas R. Campion Jr., Boyd M. Knosp, Genevieve B. Melton, Leslie Lenert |
J. Am. Medical Informatics Assoc. | 3 |
| 2022 | An atomic approach to the design and implementation of a research data warehouseabstractOBJECTIVE: As a long-standing Clinical and Translational Science Awards (CTSA) Program hub, the University of Pittsburgh and the University of Pittsburgh Medical Center (UPMC) developed and implemented a modern research data warehouse (RDW) to efficiently provision electronic patient data for clinical and translational research. MATERIALS AND METHODS: We designed and implemented an RDW named Neptune to serve the specific needs of our CTSA. Neptune uses an atomic design where data are stored at a high level of granularity as represented in source systems. Neptune contains robust patient identity management tailored for research; integrates patient data from multiple sources, including electronic health records (EHRs), health plans, and research studies; and includes knowledge for mapping to standard terminologies. RESULTS: Neptune contains data for more than 5 million patients longitudinally organized as Health Insurance Portability and Accountability Act (HIPAA) Limited Data with dates and includes structured EHR data, clinical documents, health insurance claims, and research data. Neptune is used as a source for patient data for hundreds of institutional review board-approved research projects by local investigators and for national projects. DISCUSSION: The design of Neptune was heavily influenced by the large size of UPMC, the varied data sources, and the rich partnership between the University and the healthcare system. It includes several unique aspects, including the physical warehouse straddling the University and UPMC networks and management under an HIPAA Business Associates Agreement. CONCLUSION: We describe the design and implementation of an RDW at a large academic healthcare system that uses a distinctive atomic design where data are stored at a high level of granularity. Shyam Visweswaran, Brian McLay, Nickie Cappella, Michele Morris, John T. Milnes, Steven E. Reis, Jonathan C. Silverstein, Michael J. Becich |
J. Am. Medical Informatics Assoc. | 8 |
| 2018 | Software Package to Load Data from REDCap to PCORnet CDM 4.0
Charles D. Borromeo, William Shirey, Nickie Cappella, Shyam Visweswaran, Jonathan C. Silverstein, Michael J. Becich |
AMIA | 6 |
| 2016 | The Human Microbiome: Informatics Challenges and Opportunities
Alexander V. Alekseyenko, Michael J. Becich, Todd Z. DeSantis, Jack A. Gilbert, Georg K. Gerber |
AMIA | 2 |
| 2015 | The center for causal discovery of biomedical knowledge from big dataabstractThe Big Data to Knowledge (BD2K) Center for Causal Discovery is developing and disseminating an integrated set of open source tools that support causal modeling and discovery of biomedical knowledge from large and complex biomedical datasets. The Center integrates teams of biomedical and data scientists focused on the refinement of existing and the development of new constraint-based and Bayesian algorithms based on causal Bayesian networks, the optimization of software for efficient operation in a supercomputing environment, and the testing of algorithms and software developed using real data from 3 representative driving biomedical projects: cancer driver mutations, lung disease, and the functional connectome of the human brain. Associated training activities provide both biomedical and data scientists with the knowledge and skills needed to apply and extend these tools. Collaborative activities with the BD2K Consortium further advance causal discovery tools and integrate tools and resources developed by other centers. Gregory F. Cooper, Ivet Bahar, Michael J. Becich, Panayiotis V. Benos, Jeremy M. Berg, Jeremy U. Espino, Clark Glymour, Rebecca S. Jacobson, Michelle Kienholz, Adrian V. Lee, Xinghua Lu 0001, Richard Scheines |
J. Am. Medical Informatics Assoc. | 3 |
| 2014 | Brief communication: PaTH: towards a learning health system in the Mid-Atlantic regionabstractThe PaTH (University of Pittsburgh/UPMC, Penn State College of Medicine, Temple University Hospital, and Johns Hopkins University) clinical data research network initiative is a collaborative effort among four academic health centers in the Mid-Atlantic region. PaTH will provide robust infrastructure to conduct research, explore clinical outcomes, link with biospecimens, and improve methods for sharing and analyzing data across our diverse populations. Our disease foci are idiopathic pulmonary fibrosis, atrial fibrillation, and obesity. The four network sites have extensive experience in using data from electronic health records and have devised robust methods for patient outreach and recruitment. The network will adopt best practices by using the open-source data-sharing tool, Informatics for Integrating Biology and the Bedside (i2b2), at each site to enhance data sharing using centrally defined common data elements, and will use the Shared Health Research Information Network (SHRINE) for distributed queries across the network. Waqas Amin, Fu-Chiang Tsui, Charles D. Borromeo, Cynthia H. Chuang, Jeremy U. Espino, Daniel Ford, Wenke Hwang, Wishwa Kapoor, Harold P. Lehmann, G. Daniel Martich, Sally C. Morton, Anuradha Paranjape, William Shirey, Aaron A. Sorensen, Michael J. Becich, Rachel Hess |
J. Am. Medical Informatics Assoc. | 15 |
| 2012 | Finding Collaborators: Towards Interactive Tools for Research Network Systems
Charles D. Borromeo, Titus Schleyer, Michael J. Becich, Harry Hochheiser |
AMIA | 3 |
| 2011 | The Biomedical Resource Ontology (BRO) to enable resource discovery in clinical and translational researchabstractThe biomedical research community relies on a diverse set of resources, both within their own institutions and at other research centers. In addition, an increasing number of shared electronic resources have been developed. Without effective means to locate and query these resources, it is challenging, if not impossible, for investigators to be aware of the myriad resources available, or to effectively perform resource discovery when the need arises. In this paper, we describe the development and use of the Biomedical Resource Ontology (BRO) to enable semantic annotation and discovery of biomedical resources. We also describe the Resource Discovery System (RDS) which is a federated, inter-institutional pilot project that uses the BRO to facilitate resource discovery on the Internet. Through the RDS framework and its associated Biositemaps infrastructure, the BRO facilitates semantic search and discovery of biomedical resources, breaking down barriers and streamlining scientific research that will improve human health. Jessica D. Tenenbaum, Patricia L. Whetzel, Kent Anderson, Charles D. Borromeo, Ivo D. Dinov, Davera Gabriel, Beth A. Kirschner, Barbara Mirel, Timothy D. Morris, Natasha F. Noy, Csongor Nyulas, David Rubenson, Paul R. Saxman, Nancy Whelan, Zachary C. Wright, Brian D. Athey, Michael J. Becich, Geoffrey S. Ginsburg, Mark A. Musen, Kevin A. Smith 0001, Alice F. Tarantal, Daniel L. Rubin, Peter Lyster |
J. Biomed. Informatics | 18 |
| 2010 | Unintended consequences of health information technology: A need for biomedical informatics
Elmer V. Bernstam, William R. Hersh, Ida Sim, David Eichmann, Jonathan C. Silverstein, Jack W. Smith, Michael J. Becich |
J. Biomed. Informatics | 7 |
| 2008 | Assessing the Performance Characteristics of Signals Used by a Clinical Event Monitor to Detect Adverse Drug Reactions in the Nursing Home
Steven M. Handler, Joseph T. Hanlon, Subashan Perera, Melissa I. Saul, Douglas B. Fridsma, Shyam Visweswaran, Stephanie A. Studenski, Yazan F. Roumani, Nicholas G. Castle, David A. Nace, Michael J. Becich |
AMIA | 11 |
| 2008 | Development of an instrument for measuring clinicians' power perceptions in the workplace
Christa E. Bartos, Douglas B. Fridsma, Brian S. Butler, Louis E. Penrod, Michael J. Becich, Rebecca S. Jacobson |
J. Biomed. Informatics | 5 |
| 2007 | Editorial: Lessons Learned from the Shared Pathology Informatics Network (SPIN): A Scalable Network for Translational Research and Public HealthabstractThe article by McMurry et al. in the current issue of JAMIA describes an innovative architecture to support National Health Information Networks (NHIN) that comprises a “… distributed approach to data storage in order to protect privacy and enable strong institutional autonomy to engender participation. The architecture provides oversight and transparency to ensure patient trust and allows variable levels of access according to investigator needs and institutional policies, defining a self-scaling architecture that encourages voluntary regional collaborations that coalesce to form a nationwide network …”.1 This work moves informatics a critical step forward in providing an open architecture that can support translational research and interface with appropriate depth to systems for public heath and clinical care. This linkage is crucially important for the sharing of biospecimens, and a valuable resource in this “–omics” era. Modern molecular medicine drives the demand for extensively annotated tissue specimens for basic science and translational research. Availability of such specimens can support and facilitate clinical trials, biomarker development, and discovery of new targets for novel treatments. The test bed used for McMurry's architecture has been the Shared Pathology Informatics Network (SPIN).2–4 The NIH-funded SPIN project inspired development of a number of very interesting open source technologies. These include tools for: de-identification,5, 6 autocoding,7, 8 and structuring of clinical data for use in tissue banking informatics.9, 10 This demonstrates the potential of interoperable architectures to serve the needs of care providers, investigators, and public health authorities. A key aspect of the work by McMurry is that their project can successfully “… protect patient privacy, grant institutional autonomy, and exploit legacy systems and data sharing agreements …”.1 This work includes successful development, integration, and support for de-identification systems and creation of honest broker mechanisms for the effective delivery of information to users. This addresses a frequently ignored component for the delivery of translational research by providing an effective bridge between the clinical and research communities. Public health needs are frequently ignored in the development of such systems. There is critical need for such services in biosurveillance and public health informatics. The Harvard group's work ensures “… both national anonymized coverage for routine analysis and provider authorized re-identification during emergency investigations …”.1 Hence their scalable architecture provides a multi-faceted and effective solution long sought after in biomedical informatics. It interconnects clinical, translational research, and public health informatics stakeholders. The recent NIH Roadmap initiative specifically recognizes innovative architectures such as the one proposed by McMurry et al.1 as vital components in advancing the understanding of disease and in improving health.11 Indeed, a trans-NIH Informatics Committee (TNIC) has been established to coordinate all informatics activity under the Roadmap. As one specific example of an NIH Roadmap activity, the National Electronics Clinical Trials and Research (NECTAR) network exists to enhance the efficiency of clinical research networks through informatics and other technologies.12 As a result, investigators will more easily broaden the scope of their research.12 A second example is the recently funded cancer Biomedical Informatics Grid (caBIG) project13–16 coordinated by the U.S. National Cancer Institute. The nationwide caBIG investigators group collectively develops and evaluates informatics tools and networks to support cancer research, and specifically translational research. A third Roadmap example is the NIH's seven recently funded National Centers for Biomedical Computing (NCBC). Two of the NCBC sites focus on translational informatics and a third focuses on biomedical ontology critical to the integration of clinical and research informatics efforts.17 The NIH and other agencies of the Department of Health and Human Services place increasing emphasis on, and ascribe greater significance to informatics research and development activities. The first U.S. National Health Information Technology Coordinator was established in 2004.18, 19 The President charged the Coordinator with facilitating widespread deployment of health information technology within ten years to realize substantial improvements in health-care safety, quality and efficiency.20, 21 Translational research will benefit immensely from the national focus on deployment of interoperable systems for health care. The current era has been heralded as the “Decade of Informatics,” as evidenced by the aforementioned national efforts, including adoption of key standards in health care (CDA, HL7, CCOW, LOINC, CDISC, BRIDG and others).22 The standard-related efforts, as well as corresponding terminologies and ontologies (e.g., UMLS and SNOMED) must be diffusely embedded into software development projects to ensure interoperability.23 The architecture proposed by Mcmurry et al.1 provides significant new opportunities for wide-scale adoption of informatics and research tools. Michael J. Becich |
J. Am. Medical Informatics Assoc. | 1 |
| 2004 | The tissue microarray data exchange specification: implementation by the Cooperative Prostate Cancer Tissue ResourceabstractBACKGROUND: Tissue Microarrays (TMAs) have emerged as a powerful tool for examining the distribution of marker molecules in hundreds of different tissues displayed on a single slide. TMAs have been used successfully to validate candidate molecules discovered in gene array experiments. Like gene expression studies, TMA experiments are data intensive, requiring substantial information to interpret, replicate or validate. Recently, an open access Tissue Microarray Data Exchange Specification has been released that allows TMA data to be organized in a self-describing XML document annotated with well-defined common data elements. While this specification provides sufficient information for the reproduction of the experiment by outside research groups, its initial description did not contain instructions or examples of actual implementations, and no implementation studies have been published. The purpose of this paper is to demonstrate how the TMA Data Exchange Specification is implemented in a prostate cancer TMA. RESULTS: The Cooperative Prostate Cancer Tissue Resource (CPCTR) is funded by the National Cancer Institute to provide researchers with samples of prostate cancer annotated with demographic and clinical data. The CPCTR now offers prostate cancer TMAs and has implemented a TMA database conforming to the new open access Tissue Microarray Data Exchange Specification. The bulk of the TMA database consists of clinical and demographic data elements for 299 patient samples. These data elements were extracted from an Excel database using a transformative Perl script. The Perl script and the TMA database are open access documents distributed with this manuscript. CONCLUSIONS: TMA databases conforming to the Tissue Microarray Data Exchange Specification can be merged with other TMA files, expanded through the addition of data elements, or linked to data contained in external biological databases. This article describes an open access implementation of the TMA Data Exchange Specification and provides detailed guidance to researchers who wish to use the Specification. Jules J. Berman, Milton W. Datta, André Alexander Kajdacsy-Balla, Jonathan Melamed, Jan Orenstein, Kevin Dobbin, Ashok Patel, Rajiv Dhir, Michael J. Becich |
BMC Bioinform. | 9 |
| 2004 | Tests for finding complex patterns of differential expression in cancers: towards individualized medicineabstractBACKGROUND: Microarray studies in cancer compare expression levels between two or more sample groups on thousands of genes. Data analysis follows a population-level approach (e.g., comparison of sample means) to identify differentially expressed genes. This leads to the discovery of 'population-level' markers, i.e., genes with the expression patterns A > B and B > A. We introduce the PPST test that identifies genes where a significantly large subset of cases exhibit expression values beyond upper and lower thresholds observed in the control samples. RESULTS: Interestingly, the test identifies A > B and B < A pattern genes that are missed by population-level approaches, such as the t-test, and many genes that exhibit both significant overexpression and significant underexpression in statistically significantly large subsets of cancer patients (ABA pattern genes). These patterns tend to show distributions that are unique to individual genes, and are aptly visualized in a 'gene expression pattern grid'. The low degree of among-gene correlations in these genes suggests unique underlying genomic pathologies and high degree of unique tumor-specific differential expression. We compare the PPST and the ABA test to the parametric and non-parametric t-test by analyzing two independently published data sets from studies of progression in astrocytoma. CONCLUSIONS: The PPST test resulted findings similar to the nonparametric t-test with higher self-consistency. These tests and the gene expression pattern grid may be useful for the identification of therapeutic targets and diagnostic or prognostic markers that are present only in subsets of cancer patients, and provide a more complete portrait of differential expression in cancer. James Lyons-Weiler, Satish Patel, Michael J. Becich, Tony E. Godfrey |
BMC Bioinform. | 3 |
| 2003 | Design and analysis of a content-based pathology image retrieval systemabstractA prototype, content-based image retrieval system has been built employing a client/server architecture to access supercomputing power from the physician's desktop. The system retrieves images and their associated annotations from a networked microscopic pathology image database based on content similarity to user supplied query images. Similarity is evaluated based on four image feature types: color histogram, image texture, Fourier coefficients, and wavelet coefficients, using the vector dot product as a distance metric. Current retrieval accuracy varies across pathological categories depending on the number of available training samples and the effectiveness of the feature set. The distance measure of the search algorithm was validated by agglomerative cluster analysis in light of the medical domain knowledge. Results show a correlation between pathological significance and the image document distance value generated by the computer algorithm. This correlation agrees with observed visual similarity. This validation method has an advantage over traditional statistical evaluation methods when sample size is small and where domain knowledge is important. A multi-dimensional scaling analysis shows a low dimensionality nature of the embedded space for the current test set. Arthur W. Wetzel, John R. Gilbertson, Michael J. Becich |
IEEE Trans. Inf. Technol. Biomed. | 4 |
| 2000 | Defining the role of anatomic pathology images in the multimedia electronic medical record-a preliminary report
Rebecca S. Jacobson, Cynthia S. Gadd, Gregory J. Naus, Michael J. Becich, Henry J. Lowe |
AMIA | 4 |
| 2000 | Prototype Web-based continuing medical education using FlashPix images
Adam B. Landman, Yukako Yagi, John R. Gilbertson, Robert Dawson, Alberto M. Marchevsky, Michael J. Becich |
AMIA | 6 |
| 1998 | A Graphical User Interface for Content-Based Image Retrieval Engine that Allows Remote Server Access Through the Internet
Arthur W. Wetzel, Yukako Yagi, Michael J. Becich |
AMIA | 4 |
| 1997 | Application of UMLS indexing systems to a WWW-based tool for indexing of digital images
Charlie Hatton, James W. Woods, Rajiv Dhir, Sheldon Bastacky, Jonathan Epstein, Gary Miller, Joel Greenson, Kirk Wojno, Michael J. Becich |
AMIA | 9 |