VLDB 2026 Research / reviewers in the wild / expert
Daniel Crawl
dblp:41/3549 · also Dan Crawl
· DBLP profile ↗
23ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0003-1013-8241ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 8 since 2021Software engineering, systems software and programming languages · 9 · 6 since 2021Databases, data management, data science and information retrieval · 8 · 2 since 2021Artificial intelligence and machine learning · 7 · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An Agentic Approach to Generate Conversational Narratives on the Immersive ForestabstractWildfires are becoming increasingly destructive and costly each year, affecting lives, damaging infrastructure, and degrading ecosystems. To address this growing threat, the fire and land management community need smarter, data-driven tools to understand the landscape and plan their essential prescribed burns that reduce hazardous fuels. With the ultimate goal to minimize the devastation of wildfires by enabling proactive and data-driven fuel management at landscape scale, this paper presents an approach that builds heterogeneous remote sensing data into a temporal-spatial knowledge graph, then queries it with a Large Language Model (LLM) based agent, providing a natural language interface to an extensive system of granular landscape knowledge and metrics. We demonstrate how we build knowledge graphs from LiDAR-derived vegetation metrics using GraphDB enabling precise location and time-based insights. We present a user-facing system intended to respond to queries about the effects of prescribed burns over time. Built around an LLM Agent (e.g., OpenAI, LLaMA) orchestrated with LangChain and LangGraph, the system allows users to interact with complex fire and fuel data through a natural language chat interface. It also includes a web search tool for retrieving external fire-related content to enrich responses. While this work operates as a standalone knowledge system, it was first envisioned as a method of generating conversational narratives while navigating a virtual 3D forest environment in our prototype immersive visualization application called Immersive Forest. Isaac Nealey, Nicholas Scherer, Blake Crowther, Jennifer Du, Pitchayarasm Kunghae, Wesley Schiller, Mai Nguyen, Daniel Crawl, Ilkay Altintas |
eScience | 8 |
| 2025 | Towards a Federated Approach to Complex Digital TwinsabstractIn recent years, digital twins have emerged as an advanced data-driven technology for modeling and monitoring complex physical systems to inform actionable decision-making. One area where digital twins can be especially impactful is urgent computing applications, such as natural disasters and hazards, where timely and reliable decision support is needed. However, it is challenging to streamline dynamic data with modeling and simulation methods for urgent computing applications due to the use of heterogeneous multimodal data sources, Internet of Things (IoT) networks, and high-performance computing (HPC) systems to power edge-to-cloud capabilities. In this work, we describe a reference architecture that uses federation and dynamic composability to power digital twins for complex physical systems. We demonstrate the usefulness of our approach with the Firemap architecture on WIFIRE, an end-to-end cyberinfrastructure for real-time and data-driven simulation, prediction, and visualization of wildfire behavior. Hena Ahmed, Daniel Crawl, Ilkay Altintas |
HPDC | 2 |
| 2024 | Near Real-Time Wildfire Damage Assessment using Aerial Thermal Imagery and Machine LearningabstractThis project aims at developing an AI system to provide a reliable assessment of the structural damage caused by wildfires in the first burn period. Our approach uses multimodal data, including multispectral aerial images, historical post-fire damage assessment data, and building footprints, to create an association between damage data and structure footprints. We use these associations to generate features and use machine learning methods to assess the level of damage to structures. The resulting AI-driven system can be used to provide wildfire-induced structural damage assessments in near-real-time using only aerial images for future fires. We provide damage assessment results on several megafires in California to demonstrate the applicability of our approach to real wildfire scenarios. Saqib Azim, Mai H. Nguyen, Daniel Crawl, Jessica Block, Rawaf Al Rawaf, Francesca Hart, Robert Scott, Ilkay Altintas |
IEEE Big Data | 3 |
| 2024 | Streamlined Edge Computing for Fire Science and Management using WIFIRE EdgeabstractIn recent years, frequent and highly destructive megafires become one of the biggest climate-induced disasters. Fire behavior models using data from many emerging sources can inform decision support tools to respond to and mitigate such megafires. Emerging edge sensing and computing technologies within the fire environment can enhance the speed, reliability, and efficiency of wildland fire management, leading to better prevention, faster response times, and more effective mitigation of fire-related disasters. However, a unified system that streamlines the integration of edge technology advances within fire science and management workflows is needed. This paper presents the design and demonstrated case studies of the WIFIRE Edge Platform that facilitates the integration of sensing and AI capabilities at the edge. The initial attack and prescribed burn concept scenarios are described, highlighting the sensor deployment and utilization at the fire front. Ilkay Altintas, Shweta Purawat, Ismael Pérez, Jenny Lee, Melissa Floca, Jessica Block, Josh Breslow, Daniel Crawl |
e-Science | 8 |
| 2023 | Visualization and Labeling of Terrestrial LiDAR Data for Three-Dimensional Fuel ClassificationabstractWildland fire modeling tools can ingest high resolution 3D vegetation models as inputs. However, data used to build the surface fuels in these models is often at a 30-meter resolution, which does not necessarily provide sufficient detail for accurate modeling of fires. Terrestrial laser scans are increasingly being used to collect detailed vegetation data that could be integrated with new approaches to fuel and fire modeling, but manual segmentation of scans is not scalable beyond a small number of scans. There is a need to automatically segment these high resolution point clouds as they are collected in the field, such that they may be leveraged by fuel and fire models for wildland fire response and mitigation and other applied climate science. This paper summarizes our early work on a labeling, visualization and machine learning pipeline for detailed segmentation of fuels. Specific contributions are: (1) a labeling approach involving 3 dimensional segmentation of point clouds using a point cloud processing engine; (2) a visualization approach using a computer graphics engine; and (3) early results from a deep learning modeling approach for fuel segmentation by category (live and dead) and size class (1, 10, 100 and 1000 hour fuels). Ivannia Gomez Moreno, Isaac Nealey, Daniel Roten, Mai H. Nguyen, Daniel Crawl, Kate O'Laughlin, Melissa Floca, Scott Pokswinski, Ilkay Altintas |
e-Science | 5 |
| 2023 | TrueTrees: A Scalable Workflow for the Integration of Airborne LiDAR Scanning Data into Fuel Models for Prescribed Fire SimulationsabstractLong-standing fire suppression policies, global warming and human influence at the urban-wildland interface are fueling a global wildfire crisis. Prescribed burns are increasingly being recognized as an essential procedure to reduce fuel (biomass) overgrowth and mitigate the size and severity of uncontrolled wildfires. Next-generation three-dimensional (3D) fire behavior simulations, which can help land managers to identify risks and improve planning for successful prescribed burns, depend on accurate 3D fuel structure models. We introduce TrueTrees, a workflow that integrates tree-level observations from airborne lidar surveys into FastFuels 3D fuel models. The workflow is optimized and distributed to allow processing of point cloud data for typical burn units within minutes. TrueTrees is implemented into a prototype of BurnPro3D, a user-friendly science-driven decision support platform for prescribed burn planners. Daniel Roten, Lucas A. Wells, Daniel Crawl, Russell Parsons, Anthony Marcozzi, Rodman R. Linn, John Kevin Hiers, Ilkay Altintas |
e-Science | 3 |
| 2023 | Automating the Evaluation of Datasets for FAIR and CARE PrinciplesabstractThe FAIR and CARE data principles are critical to ensuring widespread and equitable access to open data. They provide guidelines for what should be contained in metadata and how certain types of data should be handled. This study examines how large, collective data hubs such as the WIFIRE Data Commons can implement the FAIR and CARE data principles through evaluating the datasets hosted on the platform. An automation pipeline was developed to check for specified criteria in these principles, allowing fast integration of the principles on a large scale data hub. This pipeline can be expanded to check for all FAIR and CARE criteria, and similar pipelines can be created for a variety of other data hubs. Automating for the FAIR and CARE principles will help simplify organization of open data, allowing for a greater expansion of open science. Rujula Yete, Shweta Purawat, Ismael Pérez, Daniel Crawl, Ilkay Altintas |
e-Science | 4 |
| 2022 | Machine Learning for Improved Post-fire Debris Flow Likelihood PredictionabstractTimely prediction of debris flow probabilities in areas impacted by wildfires is crucial to mitigate public exposure to this hazard during post-fire rainstorms. This paper presents a machine learning approach to amend an existing dataset of post-fire debris flow events with additional features reflecting existing vegetation type and geology, and train traditional and deep learning methods on a randomly selected subset of the data. The developed methods achieve AUC (area under the receiver operational characteristic curve) values of 0.93 (random forest) and 0.92 (neural network) on the test set, representing a significant improvement over a logistic regression model currently used (AUC 0.79). The paper also overviews a distributed, Kubernetesbased big data processing pipeline to efficiently retrieve features in areas impacted by new fires, and deploy the methods for real-time prediction of debris flow hazards. Daniel Roten, Jessica Block, Daniel Crawl, Jenny Lee, Ilkay Altintas |
IEEE Big Data | 3 |
| 2022 | A Science-Enabled Virtual Reality Demonstration to Increase Social Acceptance of Prescribed BurnsabstractIncreasing social acceptance of prescribed burns is an important element of ramping up these controlled burns to the scale required to effectively mitigate destructive wildfires through reduction of excessive fire fuel loads. As part of a Design Challenge, students created concept designs for physical or virtual installations that would increase public understanding and acceptance of prescribed burns as an important tool for ending devastating megafires. The proposals defined how the public would interact with the installation and the learning goals for participants. This poster provides an overview of the virtual reality (VR) pipeline created to develop working prototypes of the immersive experiences and VR games that were proposed by the finalists in the design challenge. Isaac Nealey, Daniela Encinas Pacheco, Ivannia Gomez Moreno, Melissa Floca, Daniel Crawl, Ilkay Altintas |
e-Science | 5 |
| 2019 | Scaling Deep Learning-Based Analysis of High-Resolution Satellite Imagery with Distributed ProcessingabstractHigh-resolution satellite imagery is a rich source of data applicable to a variety of domains, ranging from demo-graphics and land use to agriculture and hazard assessment. We have developed an end-to-end analysis pipeline that uses deep learning and unsupervised learning to process high-resolution satellite imagery and have applied it to various applications in previous work. As high-resolution satellite imagery is large-volume data, scalability is important to be able to analyze data from large geographical areas. To add scalability to our process, we converted our original pipeline, implemented using the Caffe deep learning library and the Python machine learning library Scikit-Learn, to other platforms that make use of distributed computation. Specifically, to add scalability, we use Keras for deep learning, and evaluate two different distributed platforms, Spark and Dask, for unsupervised learning. We report on results in scaling up our satellite analysis pipeline. Mai H. Nguyen, Daniel Crawl, Jessica Block, Ilkay Altintas |
IEEE BigData | 3 |
| 2019 | Understanding a Rapidly Expanding Refugee Camp Using Convolutional Neural Networks and Satellite ImageryabstractIn summer 2017, close to one million Rohingya, an ethnic minority group in Myanmar, have fled to Bangladesh due to the persecution of Muslims. This large influx of refugees has resided around existing refugee camps. Because of this dramatic expansion, the newly established Kutupalong-Balukhali expansion site lacked basic infrastructure and public service. While Non-Governmental Organizations (NGOs) such as Refugee Relief and Repatriation Commissioner (RRCC) conducted a series of counting exercises to understand the demographics of refugees, our understanding of camp formation is still limited. Since the household type survey is time-consuming and does not entail geo-information, we propose to use a combination of high-resolution satellite imagery and machine learning (ML) techniques to assess the spatiotemporal dynamics of the refugee camp. Four Very-High Resolution (VHR) images (i.e., World View-2) are analyze to compare the camp pre-and post-influx. Using deep learning and unsupervised learning, we organized the satellite image tiles of a given region into geographically relevant categories. Specifically, we used a pre-trained convolutional neural network (CNN) to extract features from the image tiles, followed by cluster analysis to segment the extracted features into similar groups. Our results show that the size of the built-up area increased significantly from 0.4 km² in January 2016 and 1.5 km² in May 2017 to 8.9 km² in December 2017 and 9.5 km² in February 2018. Through the benefits of unsupervised machine learning, we further detected the densification of the refugee camp over time and were able to display its heterogeneous structure. The developed method is scalable and applicable to rapidly expanding settlements across various regions. And thus a useful tool to enhance our understanding of the structure of refugee camps, which enables us to allocate resources for humanitarian needs to the most vulnerable populations. Susanne Benz, Hogeun Park, Daniel Crawl, Jessica Block, Mai H. Nguyen, Ilkay Altintas |
eScience | 4 |
| 2018 | Land Cover Classification at the Wildland Urban Interface using High-Resolution Satellite Imagery and Deep LearningabstractLand cover classification analysis from satellite imagery is important for monitoring change in ecosystems and urban growth over time. However, the land cover classifications that are widely available in the United States are generated at a low spatial and temporal resolution, so that the spatial distribution between vegetation and urban areas in the wildland urban interface is difficult to measure. High spatial and temporal resolution analysis is essential for understanding and managing changing environments in these regions. This paper describes an end to end satellite data ingestion and analysis pipeline using deep learning on high resolution satellite imagery for generating pixel-based land cover classification. Mai H. Nguyen, Jessica Block, Daniel Crawl, Vincent Siu, Akshit Bhatnagar, Federico Rodríguez, Alison Kwan, Namrita Baru, Ilkay Altintas |
IEEE BigData | 3 |
| 2017 | Automated scalable detection of location-specific Santa Ana conditions from weather data using unsupervised learningabstractSouthern California's dry climate and fire-prone vegetation make the area vulnerable to extreme wildfire conditions. These conditions are exacerbated by Santa Ana weather patterns, which are characterized by very low humidity and gusty winds blowing in from the deserts. We present an approach using unsupervised learning to model and detect Santa Ana conditions based on sensor measurements from weather stations. Our approach uses cluster analysis to capture weather patterns specific to the region surrounding each weather station. A method is provided to automatically determine the Santa Ana cluster for each cluster model using dynamic, data-driven criteria. The resulting cluster models are applied to real-time sensor measurements to provide location-specific and time-specific detection of Santa Ana conditions. The Spark distributed platform is leveraged to scale the system to large datasets from multiple weather stations, and the Kepler workflow system is used to provide a GUI-based, easy-to-use interface to the underlying system. Results of testing our approach on an existing network of weather stations are presented. Our scalability experiment shows that the approach can process up to one million live sensor measurements in less than one minute on one machine. The proposed system can be used to aid in wildfire management and prevention by focusing firefighting efforts on regions with increased wildfire risks. Mai H. Nguyen, Daniel Crawl, Dylan Uys, Ilkay Altintas |
IEEE BigData | 2 |
| 2017 | An Unsupervised Deep Learning Approach for Satellite Image Analysis with Applications in Demographic AnalysisabstractHigh resolution satellite imagery is a growing source of data with potential applications in many diverse domains. Efficient large scale analysis of this rich data can lead to unprecedented discoveries with societal impact. We present a new framework for organizing collections of satellite images into demographically relevant categories using unsupervised learning techniques. Our framework first extracts features using pre-trained Convolutional Neural Networks from tiles of high resolution satellite images of a city. The k-means algorithm is then applied to these features to organize images into visually similar groups. The resulting clustered images are validated using demographic data. The cluster model is then applied to six different cities around the world to test the transferability of our methods. Finally, the discovered image clusters are visualized in our customized web interface to enable demographers, social scientists, and economists to understand the organization of a city. Jessica Block, Mehrdad Yazdani, Mai H. Nguyen, Daniel Crawl, Marta Jankowska, John J. Graham, Thomas A. DeFanti, Ilkay Altintas |
eScience | 4 |
| 2017 | Research Objects for Interworkability among Global Environmental and Geophysical DataabstractThe provisioning and exploitation at a global scale of environmental and geophysical data requires advanced automation and governance mechanisms that enable (meta)data interoperability but also the exchange of formalized scientific concepts and methods. In this paper we introduce recent efforts in such direction, based on scientific workflows and research objects as enablers of such vision. The former enable the integration of different web services in a higher-level data processing artifact while the latter enhances governance around data product validation and consistency, result reproducibility and credit to the principal investigators and data providers. This paper provides a concise overview of our project, current status and next steps. José Manuél Gómez-Pérez, Chuck Meertens, Fran Boler, Henry Loescher, Christine Laney, Daniel Crawl, Ilkay Altintas |
eScience | 6 |
| 2017 | Data Hub Architecture for Smart CitiesabstractToday large amount of data is generated by cities. Many of the datasets are openly available and are contributed by different sectors, government bodies and institutions. The new data can affect our understanding of the issues faced by cities and can support evidence based policies. However usage of data is limited due to difficulty in assimilating data from different sources. Open datasets often lack uniform structure which limits its analysis using traditional database systems. In this paper we present Citadel, a data hub for cities. Citadel's goal is to support end to end knowledge discovery cyber-infrastructure for effective analysis and policy support. Citadel is designed to ingest large amount of heterogeneous data and supports multiple use cases by encouraging data sharing in cities. Our poster presents the proposed features, architecture, implementation details and initial results. Jason Koh, Sandeep Singh Sandha, Bharathan Balaji, Daniel Crawl, Ilkay Altintas, Rajesh K. Gupta 0001, Mani Srivastava 0001 |
SenSys | 4 |
| 2016 | A scalable approach for location-specific detection of Santa Ana conditionsabstractSanta Ana conditions are hot, dry, windy weather conditions that can greatly increase the dangers of wildfires in southern California. We present a machine learning approach to detect Santa Ana conditions based on sensor measurements from weather stations. Cluster analysis is performed on historical weather data to build models to identify Santa Ana patterns. A separate model is built using data from each weather station to capture the patterns specific to the microclimate of each region. Real-time sensor data from a weather station can then be processed to determine if the region surrounding that station is experiencing Santa Ana conditions. Results can be used as a warning system to focus firefighting efforts on regions with increased wildfire risks. Through the use of the Kepler workflow system and distributed computing with Spark, data from several weather stations can be processed in parallel using a scalable clustering algorithm, allowing our approach to scale to large datasets from multiple weather stations. Mai H. Nguyen, Dylan Uys, Daniel Crawl, Charles Cowart, Ilkay Altintas |
IEEE BigData | 3 |
| 2015 | Big data provenance: Challenges, state of the art and opportunitiesabstractAbility to track provenance is a key feature of scientific workflows to support data lineage and reproducibility. The challenges that are introduced by the volume, variety and velocity of Big Data, also pose related challenges for provenance and quality of Big Data, defined as veracity. The increasing size and variety of distributed Big Data provenance information bring new technical challenges and opportunities throughout the provenance lifecycle including recording, querying, sharing and utilization. This paper discusses the challenges and opportunities of Big Data provenance related to the veracity of the datasets themselves and the provenance of the analytical processes that analyze these datasets. It also explains our current efforts towards tracking and utilizing Big Data provenance using workflows as a programming model to analyze Big Data. Jianwu Wang 0001, Daniel Crawl, Shweta Purawat, Mai H. Nguyen, Ilkay Altintas |
IEEE BigData | 2 |
| 2013 | Approaches to Distributed Execution of Scientific Workflows in KeplerabstractThe Kepler scientific workflow system enables creation, execution and sharing of workflows across a broad range of scientific and engineering disciplines while also facilitating remote and distributed execution of workflows. In this paper, we present Marcin Plóciennik, Tomasz Zok, Ilkay Altintas, Jianwu Wang 0001, Daniel Crawl, David Abramson 0001, Frederic Imbeaux, Bernard Guillerminet, Marcos López-Caniego, Isabel Campos Plasencia, Wojciech Pych, Pawel Ciecielag, Bartek Palak, Michal Owsiak, Yann Frauel |
Fundam. Informaticae | 5 |
| 2010 | Monitoring data quality in KeplerabstractData quality is an important component of modern scientific discovery. Many scientific discovery processes consume data from a diverse array of resources such as streaming sensor networks, web services, and databases. The validity of a scientific computation's results is highly dependent on the quality of these input data. Scientific workflow systems are being increasingly used to automate scientific computations by facilitating experiment design, data capture, integration, processing, and analysis. These workflows may execute in grid or cloud environments, and if the data produced during workflow execution is deemed unusable or low in quality, execution should stop to prevent wasting these valuable resources. We propose an approach in the Kepler scientific workflow system for monitoring data quality and demonstrate its use for oceanography and bioinformatics domains. Aisa Na'im, Daniel Crawl, Maria Indrawan, Ilkay Altintas, Shulei Sun |
HPDC | 2 |
| 2010 | Fault-Tolerance in Dataflow-Based Scientific Workflow ManagementabstractThis paper addresses the challenges of providing fault-tolerance in scientific workflow management. The specification and handling of faults in scientific workflows should be defined precisely in order to ensure the consistent execution against the process-specific requirements. We identified a number of typical failure patterns that occur in real-life scientific workflow executions. Following the intuitive recovery strategies that correspond to the identified patterns, we developed the methodologies that integrate recovery fragments into fault-prone scientific workflow models. Compared to the existing fault-tolerance mechanisms, the propositions reduce the effort of workflow designers by defining recovery fragments automatically. Furthermore, the developed framework implements the necessary mechanisms to capture the faults from the different layers of a scientific workflow management architecture. Experience indicates that the framework can be employed effectively to model, capture and tolerate the typical failure patterns that we identified. Ustun Yildiz, Pierre Mouallem, Mladen A. Vouk, Daniel Crawl, Ilkay Altintas |
SERVICES | 4 |
| 2010 | A Fault-Tolerance Architecture for Kepler-Based Distributed Scientific Workflows
Pierre Mouallem, Daniel Crawl, Ilkay Altintas, Mladen A. Vouk, Ustun Yildiz |
SSDBM | 2 |
| 2006 | Using Location Dependence to Manage Mobile DataabstractFragmentation, (portions of data located on many separate devices), and versioning, (different versions of the same datum located on different devices), of user data are increasingly prevalent and important problems as users work on a more diverse array of mobile computing devices. A common solution to these problems is to cache data on the current device. This paper describes and evaluates a mechanism for addressing fragmentation and versioning using affinity relationships. Affinity relationships that represent the current data interests of users provide hints as to what data to cache on mobile computing devices. Trace data collected over a 15 month period has been used to inform the design of the affinity mechanism. Trace-based simulation demonstrates the effectiveness of an affinity-directed approach to reduce fragmentation and versioning with low overhead Daniel Crawl, Joseph Dunn, John K. Bennett, Avneesh Bhatnagar, William Evan Speight |
MobiQuitous | 1 |