VLDB 2026 Research / reviewers in the wild / expert
Matthew Scotch
dblp:01/6138
· DBLP profile ↗
22ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0001-5100-9724ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 21 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EpiScale: Large-Scale Simulation of Infectious Disease Based on Human MobilityabstractWe present a demonstration of a highly scalable, spatially explicit infectious disease simulation that models the spread of disease across all 220,000+ census block groups in the United States using a compartmental Susceptible-Infectious-Recovered (SIR) epidemiological framework. To achieve this unprecedented scale and resolution, our system leverages efficient sparse matrix and vector operations alongside statistical approximations of large numbers of independent random events via Poisson and Normal distributions. The resulting simulation produces realistic spatiotemporal dynamics that align with empirical patterns observed in major epidemics, including the COVID-19 outbreak. Our live demonstration at the conference will highlight the simulation's computational efficiency and interactive capabilities. Starting from the conference venue in Minneapolis, participants will be able to configure disease parameters and observe the geographic spread of infection in real time, offering both an educational and analytical perspective on pandemic modeling. Ruochen Kong 0001, Taylor Anderson 0001, David J. Heslop, Matthew Scotch, Flora D. Salim, C. Raina MacIntyre, Andreas Züfle |
SIGSPATIAL/GIS | 4 |
| 2025 | A Probabilistic Framework for Imputing Genetic Distances in Spatiotemporal Pathogen ModelsabstractPathogen genome data offers valuable structure for spatial models, but its utility is limited by incomplete sequencing coverage. We propose a probabilistic framework for inferring genetic distances between unsequenced cases and known sequences within defined transmission chains, using time-aware evolutionary distance modeling. The method estimates pairwise divergence from collection dates and observed genetic distances, enabling biologically plausible imputation grounded in observed divergence patterns, without requiring sequence alignment or known transmission chains. Applied to highly pathogenic avian influenza A/H5 cases in wild birds in the United States, this approach supports scalable, uncertainty-aware augmentation of genomic datasets and enhances the integration of evolutionary information into spatiotemporal modeling workflows. Haley Stone, Jing Du 0003, Hao Xue 0001, Matthew Scotch, David J. Heslop, Andreas Züfle, C. Raina MacIntyre, Flora D. Salim |
SIGSPATIAL/GIS | 4 |
| 2025 | Simulated Infectious Diseases Datasets with Controlled Data BiasabstractMassive datasets related to infectious diseases became available after the COVID-19 pandemic, supporting data-driven approaches in modeling and forecasting infectious diseases. However, these approaches are known to exacerbate data biases present in the training data such as having certain demographic groups being over or underrepresented in the data. Such data collection biases may propagate through the modeling and prediction pipelines to decision-making, and the consequences are relatively unknown. Therefore, efforts are needed to understand how data collection bias affects data-driven infectious disease models. This datasets and benchmarks paper provides a suite of datasets, each corresponding to a simulated disease spread among a population of 5000 simulated agents over 90 days in Atlanta and San Francisco. For each dataset, we provide not only the full (simulated ground truth) of the disease spread in terms of when, where, and by whom the disease spreads, but also information on which cases are observed when different types and degrees of data collection bias are applied. The agents' characteristics, check-ins, and social network data are also available to support downstream tasks. Additionally, we also describe how to use the simulation to re-generate the data and to generate new datasets in different regions and with different parameters. With the provided datasets and the simulation tools, researchers studying the spread of infectious diseases may better understand, account for, and correct the systematic bias caused by the inherent real-world data bias, and hence improve the prediction of infectious diseases. Ruochen Kong 0001, Taylor Anderson 0001, Matthew Scotch, David J. Heslop, Yonchanok Khaokaew, Hao Xue 0001, Li Xiong 0001, C. Raina MacIntyre, Flora D. Salim, Andreas Züfle |
KDD (2) | 3 |
| 2023 | Foundational domains and competencies for baccalaureate health informatics educationabstractBACKGROUND: Foundational domains are the building blocks of educational programs. The lack of foundational domains in undergraduate health informatics (HI) education can adversely affect the development of rigorous curricula and may impede the attainment of CAHIIM accreditation of academic programs. OBJECTIVE: This White Paper presents foundational domains developed by AMIA's Academic Forum Baccalaureate Education Committee (BEC) which include corresponding competencies (knowledge, skills, and attitudes) that are intended for curriculum development and CAHIIM accreditation quality assessment for undergraduate education in applied health informatics. METHODS: The AMIA BEC used the previously published master's foundational domains as a guide to creating a set of competencies for health informatics at the undergraduate level to assess graduates from undergraduate health informatics programs for competence at graduation. A consensus method was used to adapt the domains for undergraduate level course work and harmonize the foundational domains with the currently adapted domains for HI master's education. RESULTS: Ten foundational domains were developed to support the development and evaluation of baccalaureate health informatics education. DISCUSSION: This article will inform future work towards building CAHIIM accreditation standards to ensure that higher education institutions meet acceptable levels of quality for undergraduate health informatics education. Saif S. Khairat, Sue S. Feldman, Arif Rana, Mohammad Faysel, Saptarshi Purkayastha, Matthew Scotch, Christina Eldredge |
J. Am. Medical Informatics Assoc. | 6 |
| 2022 | Towards Standards for Undergraduate Health Informatics Education
Christina Eldredge, Saif Khairat, Sue S. Feldman, Matthew Scotch |
AMIA | 4 |
| 2021 | High-throughput sequencing of SARS-CoV-2 in wastewater provides insights into circulating variants
Rafaela S. Fontenele, Simona Kraberger, James Hadfield, Erin M. Driver, Devin Bowes, LaRinda A. Holland, Temitope O. Faleye, Sangeet Adhikari, Rosa Inchausti, Wydale K. Holmes, Stephanie Deitrick, Darrell Duty, Aruni Bhatnagar, Ray A. Yeager, Rochelle H. Holm, Kevin Dixon, Tim Constantine, Melissa A. Wilson, Efrem S. Lim, Xiaofang Jiang, Rolf U. Halden, Matthew Scotch, Arvind Varsani |
AMIA | 24 |
| 2020 | GeoBoost2: a natural languageprocessing pipeline for GenBank metadata enrichment for virus phylogeographyabstractSUMMARY: We present GeoBoost2, a natural language-processing pipeline for extracting the location of infected hosts for enriching metadata in nucleotide sequences repositories like National Center of Biotechnology Information's GenBank for downstream analysis including phylogeography and genomic epidemiology. The increasing number of pathogen sequences requires complementary information extraction methods for focused research, including surveillance within countries and between borders. In this article, we describe the enhancements from our earlier release including improvement in end-to-end extraction performance and speed, availability of a fully functional web-interface and state-of-the-art methods for location extraction using deep learning. AVAILABILITY AND IMPLEMENTATION: Application is freely available on the web at https://zodo.asu.edu/geoboost2. Source code, usage examples and annotated data for GeoBoost2 is freely available at https://github.com/ZooPhy/geoboost2. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Arjun Magge, Davy Weissenbacher, Karen O'Connor, Tasnia Tahsin, Graciela Gonzalez-Hernandez, Matthew Scotch |
Bioinform. | 6 |
| 2018 | Deep neural networks and distant supervision for geographic location mention extractionabstractMotivation: Virus phylogeographers rely on DNA sequences of viruses and the locations of the infected hosts found in public sequence databases like GenBank for modeling virus spread. However, the locations in GenBank records are often only at the country or state level, and may require phylogeographers to scan the journal articles associated with the records to identify more localized geographic areas. To automate this process, we present a named entity recognizer (NER) for detecting locations in biomedical literature. We built the NER using a deep feedforward neural network to determine whether a given token is a toponym or not. To overcome the limited human annotated data available for training, we use distant supervision techniques to generate additional samples to train our NER. Results: Our NER achieves an F1-score of 0.910 and significantly outperforms the previous state-of-the-art system. Using the additional data generated through distant supervision further boosts the performance of the NER achieving an F1-score of 0.927. The NER presented in this research improves over previous systems significantly. Our experiments also demonstrate the NER's capability to embed external features to further boost the system's performance. We believe that the same methodology can be applied for recognizing similar biomedical entities in scientific literature. Arjun Magge, Davy Weissenbacher, Abeed Sarker, Matthew Scotch, Graciela Gonzalez-Hernandez |
Bioinform. | 4 |
| 2018 | GeoBoost: accelerating research involving the geospatial metadata of virus GenBank recordsabstractSummary: GeoBoost is a command-line software package developed to address sparse or incomplete metadata in GenBank sequence records that relate to the location of the infected host (LOIH) of viruses. Given a set of GenBank accession numbers corresponding to virus GenBank records, GeoBoost extracts, integrates and normalizes geographic information reflecting the LOIH of the viruses using integrated information from GenBank metadata and related full-text publications. In addition, to facilitate probabilistic geospatial modeling, GeoBoost assigns probability scores for each possible LOIH. Availability and implementation: Binaries and resources required for running GeoBoost are packed into a single zipped file and freely available for download at https://tinyurl.com/geoboost. A video tutorial is included to help users quickly and easily install and run the software. The software is implemented in Java 1.8, and supported on MS Windows and Linux platforms. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Tasnia Tahsin, Davy Weissenbacher, Karen O'Connor, Arjun Magge, Matthew Scotch, Graciela Gonzalez-Hernandez |
Bioinform. | 5 |
| 2017 | Bayesian phylogeography of influenza A/H3N2 for the 2014-15 season in the United States using three frameworks of ancestral state reconstructionabstractAncestral state reconstructions in Bayesian phylogeography of virus pandemics have been improved by utilizing a Bayesian stochastic search variable selection (BSSVS) framework. Recently, this framework has been extended to model the transition rate matrix between discrete states as a generalized linear model (GLM) of genetic, geographic, demographic, and environmental predictors of interest to the virus and incorporating BSSVS to estimate the posterior inclusion probabilities of each predictor. Although the latter appears to enhance the biological validity of ancestral state reconstruction, there has yet to be a comparison of phylogenies created by the two methods. In this paper, we compare these two methods, while also using a primitive method without BSSVS, and highlight the differences in phylogenies created by each. We test six coalescent priors and six random sequence samples of H3N2 influenza during the 2014-15 flu season in the U.S. We show that the GLMs yield significantly greater root state posterior probabilities than the two alternative methods under five of the six priors, and significantly greater Kullback-Leibler divergence values than the two alternative methods under all priors. Furthermore, the GLMs strongly implicate temperature and precipitation as driving forces of this flu season and nearly unanimously identified a single root state, which exhibits the most tropical climate during a typical flu season in the U.S. The GLM, however, appears to be highly susceptible to sampling bias compared with the other methods, which casts doubt on whether its reconstructions should be favored over those created by alternate methods. We report that a BSSVS approach with a Poisson prior demonstrates less bias toward sample size under certain conditions than the GLMs or primitive models, and believe that the connection between reconstruction method and sampling bias warrants further investigation. Daniel Magee, Marc A. Suchard, Matthew Scotch |
PLoS Comput. Biol. | 3 |
| 2016 | A high-precision rule-based extraction system for expanding geospatial metadata in GenBank recordsabstractOBJECTIVE: The metadata reflecting the location of the infected host (LOIH) of virus sequences in GenBank often lacks specificity. This work seeks to enhance this metadata by extracting more specific geographic information from related full-text articles and mapping them to their latitude/longitudes using knowledge derived from external geographical databases. MATERIALS AND METHODS: We developed a rule-based information extraction framework for linking GenBank records to the latitude/longitudes of the LOIH. Our system first extracts existing geospatial metadata from GenBank records and attempts to improve it by seeking additional, relevant geographic information from text and tables in related full-text PubMed Central articles. The final extracted locations of the records, based on data assimilated from these sources, are then disambiguated and mapped to their respective geo-coordinates. We evaluated our approach on a manually annotated dataset comprising of 5728 GenBank records for the influenza A virus. RESULTS: We found the precision, recall, and f-measure of our system for linking GenBank records to the latitude/longitudes of their LOIH to be 0.832, 0.967, and 0.894, respectively. DISCUSSION: Our system had a high level of accuracy for linking GenBank records to the geo-coordinates of the LOIH. However, it can be further improved by expanding our database of geospatial data, incorporating spell correction, and enhancing the rules used for extraction. CONCLUSION: Our system performs reasonably well for linking GenBank records for the influenza A virus to the geo-coordinates of their LOIH based on record metadata and information extracted from related full-text articles. Tasnia Tahsin, Davy Weissenbacher, Robert Rivera, Rachel Beard, Mari Firago, Garrick L. Wallstrom, Matthew Scotch, Graciela Gonzalez-Hernandez |
J. Am. Medical Informatics Assoc. | 7 |
| 2015 | Analyses of Merging Clinical and Viral Genetic Data for Influenza Surveillance
Daniel Magee, Rachel Beard, Matthew Scotch |
AMIA | 3 |
| 2015 | Knowledge-driven geospatial location resolution for phylogeographic models of virus migrationabstractUNLABELLED: Diseases caused by zoonotic viruses (viruses transmittable between humans and animals) are a major threat to public health throughout the world. By studying virus migration and mutation patterns, the field of phylogeography provides a valuable tool for improving their surveillance. A key component in phylogeographic analysis of zoonotic viruses involves identifying the specific locations of relevant viral sequences. This is usually accomplished by querying public databases such as GenBank and examining the geospatial metadata in the record. When sufficient detail is not available, a logical next step is for the researcher to conduct a manual survey of the corresponding published articles. MOTIVATION: In this article, we present a system for detection and disambiguation of locations (toponym resolution) in full-text articles to automate the retrieval of sufficient metadata. Our system has been tested on a manually annotated corpus of journal articles related to phylogeography using integrated heuristics for location disambiguation including a distance heuristic, a population heuristic and a novel heuristic utilizing knowledge obtained from GenBank metadata (i.e. a 'metadata heuristic'). RESULTS: For detecting and disambiguating locations, our system performed best using the metadata heuristic (0.54 Precision, 0.89 Recall and 0.68 F-score). Precision reaches 0.88 when examining only the disambiguation of location names. Our error analysis showed that a noticeable increase in the accuracy of toponym resolution is possible by improving the geospatial location detection. By improving these fundamental automated tasks, our system can be a useful resource to phylogeographers that rely on geospatial metadata of GenBank sequences. . Davy Weissenbacher, Tasnia Tahsin, Rachel Beard, Mari Figaro, Robert Rivera, Matthew Scotch, Graciela Gonzalez-Hernandez |
Bioinform. | 6 |
| 2014 | Comparison of ARIMA and Random Forest time series models for prediction of avian influenza H5N1 outbreaksabstractBACKGROUND: Time series models can play an important role in disease prediction. Incidence data can be used to predict the future occurrence of disease events. Developments in modeling approaches provide an opportunity to compare different time series models for predictive power. RESULTS: We applied ARIMA and Random Forest time series models to incidence data of outbreaks of highly pathogenic avian influenza (H5N1) in Egypt, available through the online EMPRES-I system. We found that the Random Forest model outperformed the ARIMA model in predictive ability. Furthermore, we found that the Random Forest model is effective for predicting outbreaks of H5N1 in Egypt. CONCLUSIONS: Random Forest time series modeling provides enhanced predictive ability over existing time series models for the prediction of infectious disease outbreaks. This result, along with those showing the concordance between bird and human outbreaks (Rabinowitz et al. 2012), provides a new approach to predicting these dangerous outbreaks in bird populations based on existing, freely available data. Our analysis uncovers the time-series structure of outbreak severity for highly pathogenic avain influenza (H5N1) in Egypt. Michael J. Kane, Natalie Price, Matthew Scotch, Peter Rabinowitz |
BMC Bioinform. | 3 |
| 2013 | Creating a MRSA Ontology to Support Categorization of MRSA Infections
Susana B. Martins, Samson W. Tu, Richard Martinello, Michael Rubin, Philip Foulis, Stephen Luther, Tyler Forbush, Matthew Scotch, Brad Doebbelling, Mary K. Goldstein |
AMIA | 8 |
| 2011 | The Yale cTAKES extensions for document classification: architecture and applicationabstractBACKGROUND: Open-source clinical natural-language-processing (NLP) systems have lowered the barrier to the development of effective clinical document classification systems. Clinical natural-language-processing systems annotate the syntax and semantics of clinical text; however, feature extraction and representation for document classification pose technical challenges. METHODS: The authors developed extensions to the clinical Text Analysis and Knowledge Extraction System (cTAKES) that simplify feature extraction, experimentation with various feature representations, and the development of both rule and machine-learning based document classifiers. The authors describe and evaluate their system, the Yale cTAKES Extensions (YTEX), on the classification of radiology reports that contain findings suggestive of hepatic decompensation. RESULTS AND DISCUSSION: The F(1)-Score of the system for the retrieval of abdominal radiology reports was 96%, and was 79%, 91%, and 95% for the presence of liver masses, ascites, and varices, respectively. The authors released YTEX as open source, available at http://code.google.com/p/ytex. Vijay Garla, Vincent Lo Re III, Zachariah Dorey-Stein, Farah Kidwai-Khan, Matthew Scotch, Julie A. Womack, Amy Justice, Cynthia Brandt |
J. Am. Medical Informatics Assoc. | 5 |
| 2011 | Enhancing phylogeography by improving geographical information from GenBankabstractPhylogeography is a field that focuses on the geographical lineages of species such as vertebrates or viruses. Here, geographical data, such as location of a species or viral host is as important as the sequence information extracted from the species. Together, this information can help illustrate the migration of the species over time within a geographical area, the impact of geography over the evolutionary history, or the expected population of the species within the area. Molecular sequence data from NCBI, specifically GenBank, provide an abundance of available sequence data for phylogeography. However, geographical data is inconsistently represented and sparse across GenBank entries. This can impede analysis and in situations where the geographical information is inferred, and potentially lead to erroneous results. In this paper, we describe the current state of geographical data in GenBank, and illustrate how automated processing techniques such as named entity recognition, can enhance the geographical data available for phylogeographic studies. Matthew Scotch, Indra Neil Sarkar, Changjiang Mei, Robert Leaman, Kei-Hoi Cheung, Pierina Ortiz, Ashutosh Singraur, Graciela Gonzalez-Hernandez |
J. Biomed. Informatics | 1 |
| 2010 | Brief review: Use of statistical analysis in the biomedical informatics literatureabstractStatistics is an essential aspect of biomedical informatics. To examine the use of statistics in informatics research, a literature review of recent articles in two high-impact factor biomedical informatics journals, the Journal of American Medical Informatics Association (JAMIA) and the International Journal of Medical Informatics was conducted. The use of statistical methods in each paper was examined. Articles of original investigations from 2000 to 2007 were reviewed. For each journal, the results by statistical methods were analyzed as: descriptive, elementary, multivariable, other regression, machine learning, and other statistics. For both journals, descriptive statistics were most often used. Elementary statistics such as t tests, chi(2), and Wilcoxon tests were much more frequent in JAMIA, while machine learning approaches such as decision trees and support vector machines were similar in occurrence across the journals. Also, the use of diagnostic statistics such as sensitivity, specificity, precision, and recall, was more frequent in JAMIA. These results highlight the use of statistics in informatics and the need for biomedical informatics scientists to have, as a minimum, proficiency in descriptive and elementary statistics. Matthew Scotch, Mona Duggal, Cynthia Brandt, Zhenqui Lin, Richard N. Shiffman |
J. Am. Medical Informatics Assoc. | 1 |
| 2008 | Case Report: Development of Grid-like Applications for Public Health Using Web 2.0 Mashup TechniquesabstractDevelopment of public health informatics applications often requires the integration of multiple data sources. This process can be challenging due to issues such as different file formats, schemas, naming systems, and having to scrape the content of web pages. A potential solution to these system development challenges is the use of Web 2.0 technologies. In general, Web 2.0 technologies are new internet services that encourage and value information sharing and collaboration among individuals. In this case report, we describe the development and use of Web 2.0 technologies including Yahoo! Pipes within a public health application that integrates animal, human, and temperature data to assess the risk of West Nile Virus (WNV) outbreaks. The results of development and testing suggest that while Web 2.0 applications are reasonable environments for rapid prototyping, they are not mature enough for large-scale public health data applications. The application, in fact a "systems of systems," often failed due to varied timeouts for application response across web sites and services, internal caching errors, and software added to web sites by administrators to manage the load on their servers. In spite of these concerns, the results of this study demonstrate the potential value of grid computing and Web 2.0 approaches in public health informatics. Matthew Scotch, Kevin Y. Yip, Kei-Hoi Cheung |
J. Am. Medical Informatics Assoc. | 1 |
| 2008 | HCLS 2.0/3.0: Health care and life sciences data mashup using Web 2.0/3.0abstractWe describe the potential of current Web 2.0 technologies to achieve data mashup in the health care and life sciences (HCLS) domains, and compare that potential to the nascent trend of performing semantic mashup. After providing an overview of Web 2.0, we demonstrate two scenarios of data mashup, facilitated by the following Web 2.0 tools and sites: Yahoo! Pipes, Dapper, Google Maps and GeoCommons. In the first scenario, we exploited Dapper and Yahoo! Pipes to implement a challenging data integration task in the context of DNA microarray research. In the second scenario, we exploited Yahoo! Pipes, Google Maps, and GeoCommons to create a geographic information system (GIS) interface that allows visualization and integration of diverse categories of public health data, including cancer incidence and pollution prevalence data. Based on these two scenarios, we discuss the strengths and weaknesses of these Web 2.0 mashup technologies. We then describe Semantic Web, the mainstream Web 3.0 technology that enables more powerful data integration over the Web. We discuss the areas of intersection of Web 2.0 and Semantic Web, and describe the potential benefits that can be brought to HCLS research by combining these two sets of technologies. Kei-Hoi Cheung, Kevin Y. Yip, Jeffrey P. Townsend, Matthew Scotch |
J. Biomed. Informatics | 4 |
| 2002 | The Companion Project - A Wearable Computing Device for the Patient for Reducing Medical Errors and Improving Patient Care
Matthew Scotch, George Hripcsak |
AMIA | 1 |
| 2002 | The sublanguage of cross-coverage
Peter D. Stetson, Stephen B. Johnson, Matthew Scotch, George Hripcsak |
AMIA | 3 |