VLDB 2026 Research / reviewers in the wild / expert
Warren A. Kibbe
dblp:65/7007
· DBLP profile ↗
13ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0001-5622-7659ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | National COVID Cohort Collaborative data enhancements: a path for expanding common data modelsabstractOBJECTIVE: To support long COVID research in National COVID Cohort Collaborative (N3C), the N3C Phenotype and Data Acquisition team created data designs to aid contributing sites in enhancing their data. Enhancements include long COVID specialty clinic indicator; Admission, Discharge, and Transfer transactions; patient-level social determinants of health; and in-hospital use of oxygen supplementation. MATERIALS AND METHODS: For each enhancement, we defined the scope and wrote guidance on how to prepare and populate the data in a standardized way. RESULTS: As of June 2024, 29 sites have added at least one data enhancement to their N3C pipeline. DISCUSSION: The use of common data models is critical to the success of N3C; however, these data models cannot account for all needs. Project-driven data enhancement is required. This should be done in a standardized way in alignment with common data model specifications. Our approach offers a useful pathway for enhancing data to improve fit for purpose. CONCLUSION: In this initiative, we rapidly produced project-specific data modeling guidance and documentation in support of long COVID research while maintaining a commitment to terminology standards and harmonized data. Kellie M. Walters, Marshall Clark, Sofia Dard, Stephanie S. Hong, Elizabeth Kelly, Kristin Kostka, Adam M. Lee, Robert T. Miller, Michele Morris, Matvey Palchuk, Emily R. Pfaff, Adam B. Wilcox, Alexis Graves, Alfred Anzalone, Amin Manna, Amit Saha, Amy Olex, Andrea Zhou, Andrew E. Williams, Andrew Southerland, Andrew T. Girvin, Anita Walden, Anjali A Sharathkumar, Benjamin R. C. Amor, Benjamin Bates, Brian Hendricks, Caleb Alexander, Carolyn T. Bramante, Cavin Ward-Caviness, Charisse R. Madlock-Brown, Christine Suver, Christopher G. Chute, Christopher Dillon, Chunlei Wu, Clare Schmitt, Cliff Takemoto, Dan Housman, Davera Gabriel, David Eichmann, Diego Mazzotti, Don Brown, Eilis A. Boudreau, Elaine L. Hill, Elizabeth Zampino, Emily Carlson Marti, Evan French, Farrukh M. Koraishy, Federico Mariona, Fred W. Prior, George Sokos, Greg Martin, Harold P. Lehmann, Heidi Spratt, Hemalkumar Mehta, Hythem Sidky, J. W. Awori Hayanga, Jami Pincavitch, Jaylyn Clark, Jeremy Richard Harper, Jessica Islam, Jin Ge, Joel Gagnier, Joel H. Saltz, Johanna Loomba, John Buse, Jomol P. Mathew, Joni L. Rutter, Julie A. McMurry, Justin Guinney, Justin Starren, Karen Crowley, Katie Rebecca Bradwell, Ken Wilkins, Kenneth R. Gersing, Kenrick Dwain Cato, Kimberly Murray, Lavance Northington, Lee Allan Pyles, Leonie Misquitta, Lesley Cottrell, Lili M. Portilla, Mariam Deacy, Mark M. Bissell, Mary Emmett, Mary Morrison Saltz, Melissa A. Haendel, Meredith C. B. Adams, Meredith Temple-O'Connor, Michael G. Kurilla, Nabeel Qureshi, Nasia Safdar, Nicole Garbarini, Noha Sharafeldin, Ofer Sadan, Patricia A. Francis, Penny Wung Burgoon, Peter N. Robinson, Philip R. O. Payne, Rafael Fuentes, Randeep Jawa, Rebecca Erwin-Cohen, Rena Patel, Richard A. Moffitt, Richard L. Zhu, Rishi Kamaleswaran, Robert Hurley, Saiju Pyarajan, Samuel G. Michael, Samuel Bozzette, Sandeep Mallipattu, Satyanarayana Vedula, Scott Chapman, Shawn T. O'Neil, Soko Setoguchi, Tellen D. Bennett, Tiffany Callahan, Umit Topaloglu, Usman Sheikh, Valery Gordon, Vignesh Subbian, Warren A. Kibbe, Wenndy Hernandez, Will Beasley, Will Cooper, William Hillegass, Xiaohan Tanner Zhang |
J. Am. Medical Informatics Assoc. | 124 |
| 2022 | Standardizing, harmonizing, and protecting data collection to broaden the impact of COVID-19 research: the rapid acceleration of diagnostics-underserved populations (RADx-UP) initiativeabstractOBJECTIVE: The Rapid Acceleration of Diagnostics-Underserved Populations (RADx-UP) program is a consortium of community-engaged research projects with the goal of increasing access to Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) tests in underserved populations. To accelerate clinical research, common data elements (CDEs) were selected and refined to standardize data collection and enhance cross-consortium analysis. MATERIALS AND METHODS: The RADx-UP consortium began with more than 700 CDEs from the National Institutes of Health (NIH) CDE Repository, Disaster Research Response (DR2) guidelines, and the PHENotypes and eXposures (PhenX) Toolkit. Following a review of initial CDEs, we made selections and further refinements through an iterative process that included live forums, consultations, and surveys completed by the first 69 RADx-UP projects. RESULTS: Following a multistep CDE development process, we decreased the number of CDEs, modified the question types, and changed the CDE wording. Most research projects were willing to collect and share demographic NIH Tier 1 CDEs, with the top exception reason being a lack of CDE applicability to the project. The NIH RADx-UP Tier 1 CDE with the lowest frequency of collection and sharing was sexual orientation. DISCUSSION: We engaged a wide range of projects and solicited bidirectional input to create CDEs. These RADx-UP CDEs could serve as the foundation for a patient-centered informatics architecture allowing the integration of disease-specific databases to support hypothesis-driven clinical research in underserved populations. CONCLUSION: A community-engaged approach using bidirectional feedback can lead to the better development and implementation of CDEs in underserved populations during public health emergencies. Gabriel A. Carrillo, Michael Cohen-Wolkowiez, Emily M. D'agostino, Keith Marsolo, Lisa M. Wruck, Laura Johnson, James Topping, Al Richmond, Giselle Corbie, Warren A. Kibbe |
J. Am. Medical Informatics Assoc. | 10 |
| 2022 | Demonstrating an approach for evaluating synthetic geospatial and temporal epidemiologic data utility: results from analyzing >1.8 million SARS-CoV-2 tests in the United States National COVID Cohort Collaborative (N3C)abstractOBJECTIVE: This study sought to evaluate whether synthetic data derived from a national coronavirus disease 2019 (COVID-19) dataset could be used for geospatial and temporal epidemic analyses. MATERIALS AND METHODS: Using an original dataset (n = 1 854 968 severe acute respiratory syndrome coronavirus 2 tests) and its synthetic derivative, we compared key indicators of COVID-19 community spread through analysis of aggregate and zip code-level epidemic curves, patient characteristics and outcomes, distribution of tests by zip code, and indicator counts stratified by month and zip code. Similarity between the data was statistically and qualitatively evaluated. RESULTS: In general, synthetic data closely matched original data for epidemic curves, patient characteristics, and outcomes. Synthetic data suppressed labels of zip codes with few total tests (mean = 2.9 ± 2.4; max = 16 tests; 66% reduction of unique zip codes). Epidemic curves and monthly indicator counts were similar between synthetic and original data in a random sample of the most tested (top 1%; n = 171) and for all unsuppressed zip codes (n = 5819), respectively. In small sample sizes, synthetic data utility was notably decreased. DISCUSSION: Analyses on the population-level and of densely tested zip codes (which contained most of the data) were similar between original and synthetically derived datasets. Analyses of sparsely tested populations were less similar and had more data suppression. CONCLUSION: In general, synthetic data were successfully used to analyze geospatial and temporal trends. Analyses using small sample sizes or populations were limited, in part due to purposeful data label suppression-an attribute disclosure countermeasure. Users should consider data fitness for use in these cases. Jason A. Thomas, Randi E. Foraker, Noa Zamstein, Jon D. Morrow, Philip R. O. Payne, Adam B. Wilcox, Melissa A. Haendel, Christopher G. Chute, Kenneth R. Gersing, Anita Walden, Tellen D. Bennett, David Eichmann, Justin Guinney, Warren A. Kibbe, Emily R. Pfaff, Peter N. Robinson, Joel H. Saltz, Heidi Spratt, Justin Starren, Christine Suver, Chunlei Wu, Davera Gabriel, Stephanie S. Hong, Kristin Kostka, Harold P. Lehmann, Richard A. Moffitt, Michele Morris, Matvey Palchuk, Xiaohan Tanner Zhang, Richard L. Zhu, Benjamin R. C. Amor, Mark M. Bissell, Marshall Clark, Andrew T. Girvin, Adam M. Lee, Robert T. Miller, Kellie M. Walters, Yooree Chae, Connor Cook, Alexandra Dest, Racquel R. Dietz, Thomas Dillon, Patricia A. Francis, Rafael Fuentes, Alexis Graves, Andrew J. Neumann, Shawn T. O'Neil, Usman Sheikh, Andréa M. Volz, Elizabeth Zampino, Christopher P. Austin, Samuel Bozzette, Mariam Deacy, Nicole Garbarini, Michael G. Kurilla, Samuel G. Michael, Joni L. Rutter, Meredith Temple-O'Connor, Katie Rebecca Bradwell, Amin Manna, Nabeel Qureshi, Mary Morrison Saltz, Julie A. McMurry, Carolyn T. Bramante, Jeremy Richard Harper, Wenndy Hernandez, Farrukh M. Koraishy, Federico Mariona, Saidulu Mattapally, Amit Saha, Satyanarayana Vedula, Yujuan Fu, Nisha Mathews, Ofer Mendelevitch |
J. Am. Medical Informatics Assoc. | 14 |
| 2021 | The National COVID Cohort Collaborative (N3C): Rationale, design, infrastructure, and deploymentabstractOBJECTIVE: Coronavirus disease 2019 (COVID-19) poses societal challenges that require expeditious data and knowledge sharing. Though organizational clinical data are abundant, these are largely inaccessible to outside researchers. Statistical, machine learning, and causal analyses are most successful with large-scale data beyond what is available in any given organization. Here, we introduce the National COVID Cohort Collaborative (N3C), an open science community focused on analyzing patient-level data from many centers. MATERIALS AND METHODS: The Clinical and Translational Science Award Program and scientific community created N3C to overcome technical, regulatory, policy, and governance barriers to sharing and harmonizing individual-level clinical data. We developed solutions to extract, aggregate, and harmonize data across organizations and data models, and created a secure data enclave to enable efficient, transparent, and reproducible collaborative analytics. RESULTS: Organized in inclusive workstreams, we created legal agreements and governance for organizations and researchers; data extraction scripts to identify and ingest positive, negative, and possible COVID-19 cases; a data quality assurance and harmonization pipeline to create a single harmonized dataset; population of the secure data enclave with data, machine learning, and statistical analytics tools; dissemination mechanisms; and a synthetic data pilot to democratize data access. CONCLUSIONS: The N3C has demonstrated that a multisite collaborative learning health network can overcome barriers to rapidly build a scalable infrastructure incorporating multiorganizational clinical data for COVID-19 analytics. We expect this effort to save lives by enabling rapid collaboration among clinicians, researchers, and data scientists to identify treatments and specialized care and thereby reduce the immediate and long-term impacts of COVID-19. Melissa A. Haendel, Christopher G. Chute, Tellen D. Bennett, David Eichmann, Justin Guinney, Warren A. Kibbe, Philip R. O. Payne, Emily R. Pfaff, Peter N. Robinson, Joel H. Saltz, Heidi Spratt, Christine Suver, John Wilbanks, Adam B. Wilcox, Andrew E. Williams, Chunlei Wu, Clair Blacketer, Robert L. Bradford, James J. Cimino, Marshall Clark, Evan W. Colmenares, Patricia A. Francis, Davera Gabriel, Alexis Graves, Raju Hemadri, Stephanie S. Hong, George Hripcsak, Dazhi Jiao, Jeffrey G. Klann, Kristin Kostka, Adam M. Lee, Harold P. Lehmann, Lora Lingrey, Robert T. Miller, Michele Morris, Shawn N. Murphy, Karthik Natarajan, Matvey Palchuk, Usman Sheikh, Harold R. Solbrig, Shyam Visweswaran, Anita Walden, Kellie M. Walters, Griffin M. Weber, Xiaohan Tanner Zhang, Richard L. Zhu, Benjamin R. C. Amor, Andrew T. Girvin, Amin Manna, Nabeel Qureshi, Michael G. Kurilla, Samuel G. Michael, Lili M. Portilla, Joni L. Rutter, Christopher P. Austin, Kenneth R. Gersing |
J. Am. Medical Informatics Assoc. | 6 |
| 2012 | GI Diaries©: A model iOS application for the collection of patient reported outcomes in clinical care
Andrew Gawron, Mark Wimbiscus Yoon, Laura Wimbiscus Yoon, Kristina Verkaik, Meghan Thompson, Abel N. Kho, Warren A. Kibbe, John E. Pandolfino |
AMIA | 7 |
| 2012 | Integrating Research Recruitment into a Clinical Patient Portal
Luke V. Rasmussen, David Were, Jeff Lunt, Steve Lee, Carl Christensen, Warren A. Kibbe, Justin Starren |
AMIA | 6 |
| 2011 | Direct2Experts: a pilot national network to demonstrate interoperability among research-networking platformsabstractResearch-networking tools use data-mining and social networking to enable expertise discovery, matchmaking and collaboration, which are important facets of team science and translational research. Several commercial and academic platforms have been built, and many institutions have deployed these products to help their investigators find local collaborators. Recent studies, though, have shown the growing importance of multiuniversity teams in science. Unfortunately, the lack of a standard data-exchange model and resistance of universities to share information about their faculty have presented barriers to forming an institutionally supported national network. This case report describes an initiative, which, in only 6 months, achieved interoperability among seven major research-networking products at 28 universities by taking an approach that focused on addressing institutional concerns and encouraging their participation. With this necessary groundwork in place, the second phase of this effort can begin, which will expand the network's functionality and focus on the end users. Griffin M. Weber, William K. Barnett, Mike Conlon, David Eichmann, Warren A. Kibbe, Holly J. Falk-Krzesinski, Michael Halaas, Layne Johnson, Eric Meeks, Donald Mitchell, Titus Schleyer, Sarah C. Stallings, Michael Warden, Maninder Kahlon |
J. Am. Medical Informatics Assoc. | 5 |
| 2010 | Comparison of Beta-value and M-value methods for quantifying methylation levels by microarray analysisabstractBACKGROUND: High-throughput profiling of DNA methylation status of CpG islands is crucial to understand the epigenetic regulation of genes. The microarray-based Infinium methylation assay by Illumina is one platform for low-cost high-throughput methylation profiling. Both Beta-value and M-value statistics have been used as metrics to measure methylation levels. However, there are no detailed studies of their relations and their strengths and limitations. RESULTS: We demonstrate that the relationship between the Beta-value and M-value methods is a Logit transformation, and show that the Beta-value method has severe heteroscedasticity for highly methylated or unmethylated CpG sites. In order to evaluate the performance of the Beta-value and M-value methods for identifying differentially methylated CpG sites, we designed a methylation titration experiment. The evaluation results show that the M-value method provides much better performance in terms of Detection Rate (DR) and True Positive Rate (TPR) for both highly methylated and unmethylated CpG sites. Imposing a minimum threshold of difference can improve the performance of the M-value method but not the Beta-value method. We also provide guidance for how to select the threshold of methylation differences. CONCLUSIONS: The Beta-value has a more intuitive biological interpretation, but the M-value is more statistically valid for the differential analysis of methylation levels. Therefore, we recommend using the M-value method for conducting differential methylation analysis and including the Beta-value statistics when reporting the results to investigators. Pan Du 0004, Chiang-Ching Huang, Nadereh Jafari, Warren A. Kibbe, Lifang Hou, Simon M. Lin |
BMC Bioinform. | 5 |
| 2009 | From disease ontology to disease-ontology lite: statistical methods to adapt a general-purpose ontology for the test of gene-ontology associationsabstractSubjective methods have been reported to adapt a general-purpose ontology for a specific application. For example, Gene Ontology (GO) Slim was created from GO to generate a highly aggregated report of the human-genome annotation. We propose statistical methods to adapt the general purpose, OBO Foundry Disease Ontology (DO) for the identification of gene-disease associations. Thus, we need a simplified definition of disease categories derived from implicated genes. On the basis of the assumption that the DO terms having similar associated genes are closely related, we group the DO terms based on the similarity of gene-to-DO mapping profiles. Two types of binary distance metrics are defined to measure the overall and subset similarity between DO terms. A compactness-scalable fuzzy clustering method is then applied to group similar DO terms. To reduce false clustering, the semantic similarities between DO terms are also used to constrain clustering results. As such, the DO terms are aggregated and the redundant DO terms are largely removed. Using these methods, we constructed a simplified vocabulary list from the DO called Disease Ontology Lite (DOLite). We demonstrated that DOLite results in more interpretable results than DO for gene-disease association tests. The resultant DOLite has been used in the Functional Disease Ontology (FunDO) Web application at http://www.projects.bioinformatics.northwestern.edu/fundo. Pan Du 0004, Jared Flatow, Michelle Holko, Warren A. Kibbe, Simon M. Lin |
Bioinform. | 6 |
| 2008 | lumi: a pipeline for processing Illumina microarrayabstractUNLABELLED: Illumina microarray is becoming a popular microarray platform. The BeadArray technology from Illumina makes its preprocessing and quality control different from other microarray technologies. Unfortunately, most other analyses have not taken advantage of the unique properties of the BeadArray system, and have just incorporated preprocessing methods originally designed for Affymetrix microarrays. lumi is a Bioconductor package especially designed to process the Illumina microarray data. It includes data input, quality control, variance stabilization, normalization and gene annotation portions. In specific, the lumi package includes a variance-stabilizing transformation (VST) algorithm that takes advantage of the technical replicates available on every Illumina microarray. Different normalization method options and multiple quality control plots are provided in the package. To better annotate the Illumina data, a vendor independent nucleotide universal identifier (nuID) was devised to identify the probes of Illumina microarray. The nuID annotation packages and output of lumi processed results can be easily integrated with other Bioconductor packages to construct a statistical data analysis pipeline for Illumina data. AVAILABILITY: The lumi Bioconductor package, www.bioconductor.org Pan Du 0004, Warren A. Kibbe, Simon M. Lin |
Bioinform. | 2 |
| 2007 | Application of Wavelet Transform to the MS-based Proteomics Data PreprocessingabstractMass spectrometry (MS) has become one of the major detection technologies for high-throughput proteomics. The preprocessing of mass spectra is crucial for its subsequent analysis like biomarker discovery or protein identification. Wavelet transform is gradually becoming an important methodology in the MS data preprocessing. This paper reviews the application of wavelet transforms in quality control, smoothing and peak detection of MS data preprocessing. It also proposes an improved Discrete Wavelet Transform (DWT) smoothing algorithm, which utilizes the cross-level DWT coefficients information during smoothing. Most of the algorithms described in this paper are included or will be included in the BioconductorMassSpecWaveletpackage. Pan Du 0004, Simon M. Lin, Warren A. Kibbe, Haihui Wang |
BIBE | 3 |
| 2006 | Improved peak detection in mass spectrum by incorporating continuous wavelet transform-based pattern matchingabstractMOTIVATION: A major problem for current peak detection algorithms is that noise in mass spectrometry (MS) spectra gives rise to a high rate of false positives. The false positive rate is especially problematic in detecting peaks with low amplitudes. Usually, various baseline correction algorithms and smoothing methods are applied before attempting peak detection. This approach is very sensitive to the amount of smoothing and aggressiveness of the baseline correction, which contribute to making peak detection results inconsistent between runs, instrumentation and analysis methods. RESULTS: Most peak detection algorithms simply identify peaks based on amplitude, ignoring the additional information present in the shape of the peaks in a spectrum. In our experience, 'true' peaks have characteristic shapes, and providing a shape-matching function that provides a 'goodness of fit' coefficient should provide a more robust peak identification method. Based on these observations, a continuous wavelet transform (CWT)-based peak detection algorithm has been devised that identifies peaks with different scales and amplitudes. By transforming the spectrum into wavelet space, the pattern-matching problem is simplified and in addition provides a powerful technique for identifying and separating the signal from the spike noise and colored noise. This transformation, with the additional information provided by the 2D CWT coefficients can greatly enhance the effective signal-to-noise ratio. Furthermore, with this technique no baseline removal or peak smoothing preprocessing steps are required before peak detection, and this improves the robustness of peak detection under a variety of conditions. The algorithm was evaluated with SELDI-TOF spectra with known polypeptide positions. Comparisons with two other popular algorithms were performed. The results show the CWT-based algorithm can identify both strong and weak peaks while keeping false positive rate low. AVAILABILITY: The algorithm is implemented in R and will be included as an open source module in the Bioconductor project. Pan Du 0004, Warren A. Kibbe, Simon M. Lin |
Bioinform. | 2 |
| 2001 | An assessment of cancer clinical trials vocabulary and IT infrastructure in the U.S
K. Gwaltney, Christopher G. Chute, Douglas Hageman, Warren A. Kibbe, Kathleen A. McCormick, Dianne M. Reeves, Lawrence W. Wright |
AMIA | 4 |