Dustin Machi

dblp:141/5357 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
6since 2021 · last 2024
0000-0002-9081-6526ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 4 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2024 Novel multi-cluster workflow system to support real-time HPC-enabled epidemic science: Investigating the impact of vaccine acceptance on COVID-19 spread
Parantapa Bhattacharya, Dustin Machi, Jiangzhuo Chen, Stefan Hoops, Bryan L. Lewis, Henning S. Mortveit, Srinivasan Venkatramanan, Mandy L. Wilson, Achla Marathe, Przemyslaw J. Porebski, Brian Klahn, Joseph Outten, Anil Vullikanti, Dawen Xie, Abhijin Adiga, Shawn Brown, Christopher L. Barrett, Madhav V. Marathe
J. Parallel Distributed Comput.2
2023 A Network Synthesis and Analytics Pipeline with Applications to Sustainable Energy in Smart Grid
abstract
Transitioning to clean and low-carbon energy is becoming a crucial goal for many entities in the energy systems sector such as governments, power utilities, and policymakers. This shift to clean energy is supported by a diverse portfolio of data products such as satellite data, smart meter data, power networks, green energy datasets (e.g., solar installations & electric vehicles), microgrid networks, and building stock data. Among these, network datasets are becoming increasingly common in addressing a wide array of issues in residential energy, especially in applications that focus on social good. Thus, streamlining the process of generating different types of networks will be helpful. In this work, we propose a versatile network synthesis and analytics pipeline developed using software design principles that make it modular, scalable, and extensible. Three case studies are presented to illustrate the significance of network data in sustainable energy applications.
Swapna Thorve, Aparna Kishore, Dustin Machi, S. S. Ravi, Madhav V. Marathe
e-Science3
2022 A Web-Based System for Contagion Simulations on Networked Populations
abstract
Motivated by a wide range of applications, research on agent-based models of contagion propagation over networks has attracted a lot of attention in the literature. Many of the available software systems for simulating such agent-based models require users to download software, build the executable, and set up execution environments. Further, running the resulting executable may require access to high performance computing clusters. Our work describes an open access software system (NetSimS) that works under the “Modeling and Simulation as a Service” (MSaaS) paradigm. It enables users to run simulations by selecting models and parameter values, initial conditions, and networks through a web interface. The system supports a variety of models and networks with millions of nodes and edges. In addition to the simulator, the system includes components that enable users to choose initial conditions for simulations in a variety of ways, to analyze the data generated through simulations, and to produce plots from the data. We describe the components of NetSimS and carry out a performance evaluation of the system. We also discuss two case studies carried out on large networks using the system. NetSimS is a major component within net.science, a cyberinfrastructure for network science.
Tanvir Ferdousi, Aparna Kishore, Lucas Machi, Dustin Machi, Chris J. Kuhlman, S. S. Ravi
e-Science4
2021 AI-Driven Agent-Based Models to Study the Role of Vaccine Acceptance in Controlling COVID-19 Spread in the US
abstract
We study the role of vaccine acceptance in controlling the spread of COVID-19 in the US using AI-driven agent-based models. Our study uses a 288 million node social contact network spanning all 50 US states plus Washington DC, comprised of 3300 counties, with 12.59 billion daily interactions. The highly-resolved agent-based models use realistic information about disease progression, vaccine uptake, production schedules, acceptance trends, prevalence, and social distancing guidelines. Developing a national model at this resolution that is driven by realistic data requires a complex scalable workflow, model calibration, simulation, and analytics components. Our workflow optimizes the total execution time and helps in improving overall human productivity.This work develops a pipeline that can execute US-scale models and associated workflows that typically present significant big data challenges. Our results show that, when compared to faster and accelerating vaccinations, slower vaccination rates due to vaccine hesitancy cause averted infections to drop from 6.7M to 4.5M, and averted total deaths to drop from 39.4K to 28.2K nationwide. This occurs despite the fact that the final vaccine coverage is the same in both scenarios. Improving vaccine acceptance by 10% in all states increases averted infections from 4.5M to 4.7M (a 4.4% improvement) and total deaths from 28.2K to 29.9K (a 6% increase) nationwide. The analysis also reveals interesting spatio-temporal differences in COVID-19 dynamics as a result of vaccine acceptance. To our knowledge, this is the first national-scale analysis of the effect of vaccine acceptance on the spread of COVID-19, using detailed and realistic agent-based models.
Parantapa Bhattacharya, Dustin Machi, Jiangzhuo Chen, Stefan Hoops, Bryan L. Lewis, Henning S. Mortveit, Srinivasan Venkatramanan, Mandy L. Wilson, Achla Marathe, Przemyslaw J. Porebski, Brian Klahn, Joseph Outten, Anil Vullikanti, Dawen Xie, Abhijin Adiga, Shawn Brown, Christopher L. Barrett, Madhav V. Marathe
IEEE BigData2
2021 Scalable Epidemiological Workflows to Support COVID-19 Planning and Response
abstract
The COVID-19 global outbreak represents the most significant epidemic event since the 1918 influenza pandemic. Simulations have played a crucial role in supporting COVID-19 planning and response efforts. Developing scalable workflows to provide policymakers quick responses to important questions pertaining to logistics, resource allocation, epidemic forecasts and intervention analysis remains a challenging computational problem. In this work, we present scalable high performance computing-enabled workflows for COVID-19 pandemic planning and response. The scalability of our methodology allows us to run fine-grained simulations daily, and to generate county-level forecasts and other counterfactual analysis for each of the 50 states (and DC), 3140 counties across the USA. Our workflows use a hybrid cloud/cluster system utilizing a combination of local and remote cluster computing facilities, and using over 20,000 CPU cores running for 6-9 hours every day to meet this objective. Our state (Virginia), state hospital network, our university, the DOD and the CDC use our models to guide their COVID-19 planning and response efforts. We began executing these pipelines March 25, 2020, and have delivered and briefed weekly updates to these stakeholders for over 30 weeks without interruption.
Dustin Machi, Parantapa Bhattacharya, Stefan Hoops, Jiangzhuo Chen, Henning S. Mortveit, Srinivasan Venkatramanan, Bryan L. Lewis, Mandy L. Wilson, Arindam Fadikar, Tom Maiden, Christopher L. Barrett, Madhav V. Marathe
IPDPS1
2021 A genomic data resource for predicting antimicrobial resistance from laboratory-derived antimicrobial susceptibility phenotypes
abstract
Antimicrobial resistance (AMR) is a major global health threat that affects millions of people each year. Funding agencies worldwide and the global research community have expended considerable capital and effort tracking the evolution and spread of AMR by isolating and sequencing bacterial strains and performing antimicrobial susceptibility testing (AST). For the last several years, we have been capturing these efforts by curating data from the literature and data resources and building a set of assembled bacterial genome sequences that are paired with laboratory-derived AST data. This collection currently contains AST data for over 67 000 genomes encompassing approximately 40 genera and over 100 species. In this paper, we describe the characteristics of this collection, highlighting areas where sampling is comparatively deep or shallow, and showing areas where attention is needed from the research community to improve sampling and tracking efforts. In addition to using the data to track the evolution and spread of AMR, it also serves as a useful starting point for building machine learning models for predicting AMR phenotypes. We demonstrate this by describing two machine learning models that are built from the entire dataset to show where the predictive power is comparatively high or low. This AMR metadata collection is freely available and maintained on the Bacterial and Viral Bioinformatics Center (BV-BRC) FTP site ftp://ftp.bvbrc.org/RELEASE_NOTES/PATRIC_genomes_AMR.txt.
Margo VanOeffelen, Marcus Nguyen, Derya Aytan-Aktug, Thomas S. Brettin, Emily M. Dietrich, Ron Kenyon, Dustin Machi, Chunhong Mao, Robert Olson, Gordon D. Pusch, Maulik Shukla, Rick L. Stevens, Veronika Vonstein, Andrew S. Warren, Alice R. Wattam, Hyun Seung Yoo, James J. Davis 0002
Briefings Bioinform.7
2020 From 5Vs to 6Cs: Operationalizing Epidemic Data Management with COVID-19 Surveillance
abstract
The COVID-19 pandemic brought to the forefront an unprecedented need for experts, as well as citizens, to visualize spatio-temporal disease surveillance data. Web application dashboards were quickly developed to fill t his g ap, b ut a ll of these dashboards supported a particular niche view of the pandemic (ie, current status or specific r egions). I n t his paper, we describe our work developing our COVID-19 Surveillance Dashboard, which offers a unique view of the pandemic while also allowing users to focus on the details that interest them. From the beginning, our goal was to provide a simple visual tool for comparing, organizing, and tracking near-real-time surveillance data as the pandemic progresses. In developing this dashboard, we also identified 6 key metrics which we propose as a standard for the design and evaluation of real-time epidemic science dashboards. Our dashboard was one of the first released to the public, and continues to be actively visited. Our own group uses it to support federal, state and local public health authorities, and it is used by individuals worldwide to track the evolution of the COVID-19 pandemic, build their own dashboards, and support their organizations as they plan their responses to the pandemic.
Akhil Sai Peddireddy, Dawen Xie, Pramod Patil, Mandy L. Wilson, Dustin Machi, Srinivasan Venkatramanan, Brian Klahn, Przemyslaw J. Porebski, Parantapa Bhattacharya, Shirish Dumbre, Erin Raymond, Madhav V. Marathe
IEEE BigData5
2019 Mechanistic and data-driven agent-based models to explain human behavior in online networked group anagram games
abstract
In anagram games, players are provided with letters for forming as many words as possible over a specified time duration. Anagram games have been used in controlled experiments to study problems such as collective identity, effects of goal-setting, internal-external attributions, test anxiety, and others. The majority of work on anagram games involves individual players. Recently, work has expanded to group anagram games where players cooperate by sharing letters. In this work, we analyze experimental data from online social networked experiments of group anagram games. We develop mechanistic and data-driven models of human decision-making to predict detailed game player actions (e.g., what word to form next). With these results, we develop a composite agent-based modeling and simulation platform that incorporates the models from data analysis. We compare model predictions against experimental data, which enables us to provide explanations of human decision-making and behavior. Finally, we provide illustrative case studies using agent-based simulations to demonstrate the efficacy of models to provide insights that are beyond those from experiments alone.
Vanessa Cedeno-Mieles, Xinwei Deng, Yihui Ren 0001, Abhijin Adiga, Christopher L. Barrett, Saliya Ekanayake, Gizem Korkmaz, Chris J. Kuhlman, Dustin Machi, Madhav V. Marathe, S. S. Ravi, Brian J. Goode, Naren Ramakrishnan, Parang Saraf, Nathan Self, Noshir S. Contractor, Joshua M. Epstein, Michael W. Macy
ASONAM10
2019 PATRIC as a unique resource for studying antimicrobial resistance
abstract
The Pathosystems Resource Integration Center (PATRIC, www.patricbrc.org) is designed to provide researchers with the tools and services that they need to perform genomic and other 'omic' data analyses. In response to mounting concern over antimicrobial resistance (AMR), the PATRIC team has been developing new tools that help researchers understand AMR and its genetic determinants. To support comparative analyses, we have added AMR phenotype data to over 15 000 genomes in the PATRIC database, often assembling genomes from reads in public archives and collecting their associated AMR panel data from the literature to augment the collection. We have also been using this collection of AMR metadata to build machine learning-based classifiers that can predict the AMR phenotypes and the genomic regions associated with resistance for genomes being submitted to the annotation service. Likewise, we have undertaken a large AMR protein annotation effort by manually curating data from the literature and public repositories. This collection of 7370 AMR reference proteins, which contains many protein annotations (functional roles) that are unique to PATRIC and RAST, has been manually curated so that it projects stably across genomes. The collection currently projects to 1 610 744 proteins in the PATRIC database. Finally, the PATRIC Web site has been expanded to enable AMR-based custom page views so that researchers can easily explore AMR data and design experiments based on whole genomes or individual genes.
Dionysios A. Antonopoulos, Rida Assaf, Ramy K. Aziz, Thomas S. Brettin, Christopher Bun, Neal Conrad, James J. Davis 0002, Emily M. Dietrich, Terry Disz, Svetlana Gerdes, Ron Kenyon, Dustin Machi, Chunhong Mao, Daniel E. Murphy-Olson, Eric K. Nordberg, Gary J. Olsen, Robert Olson, Ross A. Overbeek, Bruce D. Parrello, Gordon D. Pusch, John Santerre, Maulik Shukla, Rick L. Stevens, Margo VanOeffelen, Veronika Vonstein, Andrew S. Warren, Alice R. Wattam, Fangfang Xia, Hyun Seung Yoo
Briefings Bioinform.12
2018 Generative Modeling of Human Behavior and Social Interactions Using Abductive Analysis
abstract
Abduction is an inference approach that uses data and observations to identify plausible (and preferably, best) explanations for phenomena. Applications of abduction (e.g., robotics, genetics, image understanding) have largely been devoid of human behavior. Here, we devise and execute an iterative abductive analysis process that is driven by the social sciences: behaviors and interactions among groups of human subjects. One goal is to understand intra-group cooperation and its effect on fostering collective identity. We build an online game platform; perform and analyze controlled laboratory experiments; form hypotheses; build, exercise, and evaluate network-based agent-based models; and evaluate the hypotheses in multiple abductive iterations, improving our understanding as the process unfolds. While the experimental results are of interest, the paper's thrust is methodological, and indeed establishes the potential of iterative abductive looping for the (computational) social sciences.
Yihui Ren 0001, Vanessa Cedeno-Mieles, Xinwei Deng, Abhijin Adiga, Christopher L. Barrett, Saliya Ekanayake, Brian J. Goode, Gizem Korkmaz, Chris J. Kuhlman, Dustin Machi, Madhav V. Marathe, Naren Ramakrishnan, S. S. Ravi, Parang Saraf, Nathan Self, Noshir S. Contractor, Joshua M. Epstein, Michael W. Macy
ASONAM11
2015 RNA-Rocket: an RNA-Seq analysis resource for infectious disease research
abstract
MOTIVATION: RNA-Seq is a method for profiling transcription using high-throughput sequencing and is an important component of many research projects that wish to study transcript isoforms, condition specific expression and transcriptional structure. The methods, tools and technologies used to perform RNA-Seq analysis continue to change, creating a bioinformatics challenge for researchers who wish to exploit these data. Resources that bring together genomic data, analysis tools, educational material and computational infrastructure can minimize the overhead required of life science researchers. RESULTS: RNA-Rocket is a free service that provides access to RNA-Seq and ChIP-Seq analysis tools for studying infectious diseases. The site makes available thousands of pre-indexed genomes, their annotations and the ability to stream results to the bioinformatics resources VectorBase, EuPathDB and PATRIC. The site also provides a combination of experimental data and metadata, examples of pre-computed analysis, step-by-step guides and a user interface designed to enable both novice and experienced users of RNA-Seq data. AVAILABILITY AND IMPLEMENTATION: RNA-Rocket is available at rnaseq.pathogenportal.org. Source code for this project can be found at github.com/cidvbi/PathogenPortal. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary materials are available at Bioinformatics online.
Andrew S. Warren, Cristina Aurrecoechea, Brian P. Brunk, Prerak Desai, Scott J. Emrich, Gloria I. Giraldo-Calderón, Omar S. Harb, Deborah Hix, Daniel Lawson, Dustin Machi, Chunhong Mao, Michael McClelland, Eric K. Nordberg, Maulik Shukla, Leslie B. Vosshall, Alice R. Wattam, Rebecca Will, Hyun Seung Yoo, Bruno W. S. Sobral
Bioinform.10