VLDB 2026 Research / reviewers in the wild / expert
Shweta Purawat
dblp:165/9171
· DBLP profile ↗
7ranked-venue papers
1as first author
5since 2021 · last 2024
0000-0002-5183-2750ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Streamlined Edge Computing for Fire Science and Management using WIFIRE EdgeabstractIn recent years, frequent and highly destructive megafires become one of the biggest climate-induced disasters. Fire behavior models using data from many emerging sources can inform decision support tools to respond to and mitigate such megafires. Emerging edge sensing and computing technologies within the fire environment can enhance the speed, reliability, and efficiency of wildland fire management, leading to better prevention, faster response times, and more effective mitigation of fire-related disasters. However, a unified system that streamlines the integration of edge technology advances within fire science and management workflows is needed. This paper presents the design and demonstrated case studies of the WIFIRE Edge Platform that facilitates the integration of sensing and AI capabilities at the edge. The initial attack and prescribed burn concept scenarios are described, highlighting the sensor deployment and utilization at the fire front. Ilkay Altintas, Shweta Purawat, Ismael Pérez, Jenny Lee, Melissa Floca, Jessica Block, Josh Breslow, Daniel Crawl |
e-Science | 2 |
| 2023 | Preserving File Provenance Using Principles of Blockchain to Ensure Scientific ReproducibilityabstractReproducibility plays an essential role in scientific research to ensure accuracy and serves as a foundation for future advancements. Scientific reproducibility becomes particularly challenging when dealing with vast amounts of input files that change hands or move across different laboratories or organizations. Preserving the provenance of data files ensures critical information about the originality of data files is captured to support the reproducibility of scientific research. The paper focuses on capturing and verifying input and output data file provenance using the principles of blockchain. The technique stores the hashes of data files in a database along with user and workflow information. It allows the workflow to verify the data against the hashes at any point. The method is demonstrated using Parflow, a Hydrologic model, as a proof-of-concept. Rizbanul Hasan, Shweta Purawat, Catherine Mills Olschanowsky, Ilkay Altintas |
e-Science | 2 |
| 2023 | Automating the Evaluation of Datasets for FAIR and CARE PrinciplesabstractThe FAIR and CARE data principles are critical to ensuring widespread and equitable access to open data. They provide guidelines for what should be contained in metadata and how certain types of data should be handled. This study examines how large, collective data hubs such as the WIFIRE Data Commons can implement the FAIR and CARE data principles through evaluating the datasets hosted on the platform. An automation pipeline was developed to check for specified criteria in these principles, allowing fast integration of the principles on a large scale data hub. This pipeline can be expanded to check for all FAIR and CARE criteria, and similar pipelines can be created for a variety of other data hubs. Automating for the FAIR and CARE principles will help simplify organization of open data, allowing for a greater expansion of open science. Rujula Yete, Shweta Purawat, Ismael Pérez, Daniel Crawl, Ilkay Altintas |
e-Science | 2 |
| 2021 | TemPredict: A Big Data Analytical Platform for Scalable Exploration and Monitoring of Personalized Multimodal Data for COVID-19abstractA key takeaway from the COVID-19 crisis is the need for scalable methods and systems for ingestion of big data related to the disease, such as models of the virus, health surveys, and social data, and the ability to integrate and analyze the ingested data rapidly. One specific example is the use of the Internet of Things and wearables (i.e., the Oura ring) to collect large-scale individualized data (e.g., temperature and heart rate) continuously and to create personalized baselines for detection of disease symptoms. Individualized data, when collected, has great potential to be linked with other datasets making it possible to combine individual and societal scale models for further understanding the disease. However, the volume and variability of such data require novel big data approaches to be developed as infrastructure for scalable use. This paper presents the data pipeline and big data infrastructure for the TemPredict project, which, to the best of our knowledge, is the largest public effort to gather continuous physiological data for time-series analysis. This effort unifies data ingestion with the development of a novel end-to-end cyberinfrastructure to enable the curation, cleaning, alignment, sketching, and passing of the data, in a secure manner, by the researchers making use of the ingested data for their COVID-19 detection algorithm development efforts. We present the challenges, the closed-loop data pipelines, and the secure infrastructure to support the development of time-sensitive algorithms for alerting individuals based on physiological predictors illness, enabling early intervention. Shweta Purawat, Subhasis Dasgupta, Jining Song, Shakti Davis, Kajal T. Claypool, Sandeep Chandra, Ashley E. Mason, Varun K. Viswanath, Amit Klein 0002, Patrick Kasl, YingJing Wen, Benjamin L. Smarr, Amarnath Gupta, Ilkay Altintas |
IEEE BigData | 1 |
| 2021 | Modular performance prediction for scientific workflows using Machine Learning
Alok Singh 0004, Shweta Purawat, Arvind Rao, Ilkay Altintas |
Future Gener. Comput. Syst. | 2 |
| 2019 | A demonstration of modularity, reuse, reproducibility, portability and scalability for modeling and simulation of cardiac electrophysiology using Kepler WorkflowsabstractMulti-scale computational modeling is a major branch of computational biology as evidenced by the US federal interagency Multi-Scale Modeling Consortium and major international projects. It invariably involves specific and detailed sequences of data analysis and simulation, often with multiple tools and datasets, and the community recognizes improved modularity, reuse, reproducibility, portability and scalability as critical unmet needs in this area. Scientific workflows are a well-recognized strategy for addressing these needs in scientific computing. While there are good examples if the use of scientific workflows in bioinformatics, medical informatics, biomedical imaging and data analysis, there are fewer examples in multi-scale computational modeling in general and cardiac electrophysiology in particular. Cardiac electrophysiology simulation is a mature area of multi-scale computational biology that serves as an excellent use case for developing and testing new scientific workflows. In this article, we develop, describe and test a computational workflow that serves as a proof of concept of a platform for the robust integration and implementation of a reusable and reproducible multi-scale cardiac cell and tissue model that is expandable, modular and portable. The workflow described leverages Python and Kepler-Python actor for plotting and pre/post-processing. During all stages of the workflow design, we rely on freely available open-source tools, to make our workflow freely usable by scientists. Pei-Chi Yang, Shweta Purawat, Pek U. Ieong, Mao-Tsuen Jeng, Kevin R. DeMarco, Igor Vorobyov, Andrew D. McCulloch, Ilkay Altintas, Rommie E. Amaro, Colleen E. Clancy |
PLoS Comput. Biol. | 2 |
| 2015 | Big data provenance: Challenges, state of the art and opportunitiesabstractAbility to track provenance is a key feature of scientific workflows to support data lineage and reproducibility. The challenges that are introduced by the volume, variety and velocity of Big Data, also pose related challenges for provenance and quality of Big Data, defined as veracity. The increasing size and variety of distributed Big Data provenance information bring new technical challenges and opportunities throughout the provenance lifecycle including recording, querying, sharing and utilization. This paper discusses the challenges and opportunities of Big Data provenance related to the veracity of the datasets themselves and the provenance of the analytical processes that analyze these datasets. It also explains our current efforts towards tracking and utilizing Big Data provenance using workflows as a programming model to analyze Big Data. Jianwu Wang 0001, Daniel Crawl, Shweta Purawat, Mai H. Nguyen, Ilkay Altintas |
IEEE BigData | 3 |