VLDB 2026 Research / reviewers in the wild / expert
Mai H. Nguyen
dblp:54/3961
· DBLP profile ↗
16ranked-venue papers
6as first author
5since 2021 · last 2024
0000-0002-4945-1334ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 2 since 2021Databases, data management, data science and information retrieval · 9 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Near Real-Time Wildfire Damage Assessment using Aerial Thermal Imagery and Machine LearningabstractThis project aims at developing an AI system to provide a reliable assessment of the structural damage caused by wildfires in the first burn period. Our approach uses multimodal data, including multispectral aerial images, historical post-fire damage assessment data, and building footprints, to create an association between damage data and structure footprints. We use these associations to generate features and use machine learning methods to assess the level of damage to structures. The resulting AI-driven system can be used to provide wildfire-induced structural damage assessments in near-real-time using only aerial images for future fires. We provide damage assessment results on several megafires in California to demonstrate the applicability of our approach to real wildfire scenarios. Saqib Azim, Mai H. Nguyen, Daniel Crawl, Jessica Block, Rawaf Al Rawaf, Francesca Hart, Robert Scott, Ilkay Altintas |
IEEE Big Data | 2 |
| 2024 | Prescribed Fire Modeling using Knowledge-Guided Machine Learning for Land ManagementabstractIn recent years, the increasing threat of devastating wildfires has underscored the need for effective prescribed fire management. Process-based computer simulations have traditionally been employed to plan prescribed fires for wildfire prevention. However, even simplified process models are too compute-intensive to be used for real-time decision-making. Traditional ML methods used for fire modeling offer computational speedup but struggle with physically inconsistent predictions, biased predictions due to class imbalance, biased estimates for fire spread metrics (e.g., burned area, rate of spread), and limited generalizability in out-of-distribution wind conditions. This paper introduces a novel machine learning (ML) framework that enables rapid emulation of prescribed fires while addressing these concerns. To overcome these challenges, the framework incorporates domain knowledge in the form of physical constraints, a hierarchical modeling structure to capture the interdependence among variables of interest, and also leverages pre-existing source domain data to augment training data and learn the spread of fire more effectively. Notably, improvement in fire metric (e.g., burned area) estimates offered by our framework makes it useful for fire managers, who often rely on these estimates to make decisions about prescribed burn management. Furthermore, our framework exhibits better generalization capabilities than the other ML-based fire modeling methods across diverse wind conditions and ignition patterns. Somya Sharma Chatterjee, Kelly Lindsay, Neel Chatterjee, Rohan Patil, Ilkay Altintas De Callafon, Michael S. Steinbach, Daniel Giron, Mai H. Nguyen, Vipin Kumar 0001 |
SDM | 8 |
| 2023 | Visualization and Labeling of Terrestrial LiDAR Data for Three-Dimensional Fuel ClassificationabstractWildland fire modeling tools can ingest high resolution 3D vegetation models as inputs. However, data used to build the surface fuels in these models is often at a 30-meter resolution, which does not necessarily provide sufficient detail for accurate modeling of fires. Terrestrial laser scans are increasingly being used to collect detailed vegetation data that could be integrated with new approaches to fuel and fire modeling, but manual segmentation of scans is not scalable beyond a small number of scans. There is a need to automatically segment these high resolution point clouds as they are collected in the field, such that they may be leveraged by fuel and fire models for wildland fire response and mitigation and other applied climate science. This paper summarizes our early work on a labeling, visualization and machine learning pipeline for detailed segmentation of fuels. Specific contributions are: (1) a labeling approach involving 3 dimensional segmentation of point clouds using a point cloud processing engine; (2) a visualization approach using a computer graphics engine; and (3) early results from a deep learning modeling approach for fuel segmentation by category (live and dead) and size class (1, 10, 100 and 1000 hour fuels). Ivannia Gomez Moreno, Isaac Nealey, Daniel Roten, Mai H. Nguyen, Daniel Crawl, Kate O'Laughlin, Melissa Floca, Scott Pokswinski, Ilkay Altintas |
e-Science | 4 |
| 2022 | Effects of data and entity ablation on multitask learning models for biomedical entity recognition
Nicholas E. Rodriguez, Mai H. Nguyen, Bridget T. McInnes |
J. Biomed. Informatics | 2 |
| 2021 | Generalization in Cardiac Image SegmentationabstractDeep learning methods have achieved great success in medical imaging applications. Although data is very crucial for deep learning models, the medical imaging domain is restricted by the limited size of datasets and differences between them. This makes it difficult for models to generalize across datasets and achieve robust performance in practical settings. Therefore, knowing how to combine different datasets together to take advantage of new data while retaining performance on previous data becomes an important problem. In this paper, we focus on the task of cardiac image semantic segmentation and synthesize five different real-world scenarios to find the optimal training approach for deep learning models to achieve good generalization across different datasets. Zhengjie Xu, Garrison W. Cottrell, Mai H. Nguyen |
IEEE BigData | 4 |
| 2019 | Scaling Deep Learning-Based Analysis of High-Resolution Satellite Imagery with Distributed ProcessingabstractHigh-resolution satellite imagery is a rich source of data applicable to a variety of domains, ranging from demo-graphics and land use to agriculture and hazard assessment. We have developed an end-to-end analysis pipeline that uses deep learning and unsupervised learning to process high-resolution satellite imagery and have applied it to various applications in previous work. As high-resolution satellite imagery is large-volume data, scalability is important to be able to analyze data from large geographical areas. To add scalability to our process, we converted our original pipeline, implemented using the Caffe deep learning library and the Python machine learning library Scikit-Learn, to other platforms that make use of distributed computation. Specifically, to add scalability, we use Keras for deep learning, and evaluate two different distributed platforms, Spark and Dask, for unsupervised learning. We report on results in scaling up our satellite analysis pipeline. Mai H. Nguyen, Daniel Crawl, Jessica Block, Ilkay Altintas |
IEEE BigData | 1 |
| 2019 | Understanding a Rapidly Expanding Refugee Camp Using Convolutional Neural Networks and Satellite ImageryabstractIn summer 2017, close to one million Rohingya, an ethnic minority group in Myanmar, have fled to Bangladesh due to the persecution of Muslims. This large influx of refugees has resided around existing refugee camps. Because of this dramatic expansion, the newly established Kutupalong-Balukhali expansion site lacked basic infrastructure and public service. While Non-Governmental Organizations (NGOs) such as Refugee Relief and Repatriation Commissioner (RRCC) conducted a series of counting exercises to understand the demographics of refugees, our understanding of camp formation is still limited. Since the household type survey is time-consuming and does not entail geo-information, we propose to use a combination of high-resolution satellite imagery and machine learning (ML) techniques to assess the spatiotemporal dynamics of the refugee camp. Four Very-High Resolution (VHR) images (i.e., World View-2) are analyze to compare the camp pre-and post-influx. Using deep learning and unsupervised learning, we organized the satellite image tiles of a given region into geographically relevant categories. Specifically, we used a pre-trained convolutional neural network (CNN) to extract features from the image tiles, followed by cluster analysis to segment the extracted features into similar groups. Our results show that the size of the built-up area increased significantly from 0.4 km² in January 2016 and 1.5 km² in May 2017 to 8.9 km² in December 2017 and 9.5 km² in February 2018. Through the benefits of unsupervised machine learning, we further detected the densification of the refugee camp over time and were able to display its heterogeneous structure. The developed method is scalable and applicable to rapidly expanding settlements across various regions. And thus a useful tool to enhance our understanding of the structure of refugee camps, which enables us to allocate resources for humanitarian needs to the most vulnerable populations. Susanne Benz, Hogeun Park, Daniel Crawl, Jessica Block, Mai H. Nguyen, Ilkay Altintas |
eScience | 6 |
| 2019 | Ten simple rules for writing and sharing computational analyses in Jupyter NotebooksabstractAuthor(s): Rule, Adam; Birmingham, Amanda; Zuniga, Cristal; Altintas, Ilkay; Huang, Shih-Cheng; Knight, Rob; Moshiri, Niema; Nguyen, Mai H; Rosenthal, Sara Brin; Pérez, Fernando; Rose, Peter W | Editor(s): Lewitter, Fran Adam Rule, Amanda Birmingham, Cristal Zuñiga, Ilkay Altintas, Shih-Cheng Huang, Rob Knight 0001, Niema Moshiri, Mai H. Nguyen, Sara Brin Rosenthal, Peter W. Rose |
PLoS Comput. Biol. | 8 |
| 2018 | Land Cover Classification at the Wildland Urban Interface using High-Resolution Satellite Imagery and Deep LearningabstractLand cover classification analysis from satellite imagery is important for monitoring change in ecosystems and urban growth over time. However, the land cover classifications that are widely available in the United States are generated at a low spatial and temporal resolution, so that the spatial distribution between vegetation and urban areas in the wildland urban interface is difficult to measure. High spatial and temporal resolution analysis is essential for understanding and managing changing environments in these regions. This paper describes an end to end satellite data ingestion and analysis pipeline using deep learning on high resolution satellite imagery for generating pixel-based land cover classification. Mai H. Nguyen, Jessica Block, Daniel Crawl, Vincent Siu, Akshit Bhatnagar, Federico Rodríguez, Alison Kwan, Namrita Baru, Ilkay Altintas |
IEEE BigData | 1 |
| 2018 | Analytics Pipeline for Left Ventricle Segmentation and Volume Estimation on Cardiac MRI Using Deep LearningabstractThe left ventricle (LV) is the largest chamber in the heart and plays a critical role in cardiac function. Noninvasive cardiac imaging modalities (e.g., cardiac magnetic resonance (CMR), transesophageal echocardiography (TEE), and computed tomography (CT)) are commonly used to study LV size and function in addition to other cardiac structural aspects such as valvular disease, and are invaluable tools for the diagnosis and management of heart disease. However, the process of analyzing cardiac images is time-consuming and labor-intensive. Automatic LV segmentation and volume estimation from cardiac images are thus essential in providing efficient and consistent analysis. We discuss findings from our investigation into different techniques for processing and analyzing CMR images and present the methods giving best performance in an end -to -end analytics pipeline for LV segmentation and volume estimation. This pipeline can serve as an initial step towards analyzing CMR at scale to aid in non-invasive cardiac disease detection. Mai H. Nguyen, Ehab Abdelmaguid, Jolene Huang, Sanjay Kenchareddy, Disha Singla, Laura Wilke, Marcus Bobar, Eric D. Carruth, Dylan Uys, Ilkay Altintas, Evan D. Muse, Giorgio Quer, Steven R. Steinhubl |
eScience | 1 |
| 2017 | Automated scalable detection of location-specific Santa Ana conditions from weather data using unsupervised learningabstractSouthern California's dry climate and fire-prone vegetation make the area vulnerable to extreme wildfire conditions. These conditions are exacerbated by Santa Ana weather patterns, which are characterized by very low humidity and gusty winds blowing in from the deserts. We present an approach using unsupervised learning to model and detect Santa Ana conditions based on sensor measurements from weather stations. Our approach uses cluster analysis to capture weather patterns specific to the region surrounding each weather station. A method is provided to automatically determine the Santa Ana cluster for each cluster model using dynamic, data-driven criteria. The resulting cluster models are applied to real-time sensor measurements to provide location-specific and time-specific detection of Santa Ana conditions. The Spark distributed platform is leveraged to scale the system to large datasets from multiple weather stations, and the Kepler workflow system is used to provide a GUI-based, easy-to-use interface to the underlying system. Results of testing our approach on an existing network of weather stations are presented. Our scalability experiment shows that the approach can process up to one million live sensor measurements in less than one minute on one machine. The proposed system can be used to aid in wildfire management and prevention by focusing firefighting efforts on regions with increased wildfire risks. Mai H. Nguyen, Daniel Crawl, Dylan Uys, Ilkay Altintas |
IEEE BigData | 1 |
| 2017 | An Unsupervised Deep Learning Approach for Satellite Image Analysis with Applications in Demographic AnalysisabstractHigh resolution satellite imagery is a growing source of data with potential applications in many diverse domains. Efficient large scale analysis of this rich data can lead to unprecedented discoveries with societal impact. We present a new framework for organizing collections of satellite images into demographically relevant categories using unsupervised learning techniques. Our framework first extracts features using pre-trained Convolutional Neural Networks from tiles of high resolution satellite images of a city. The k-means algorithm is then applied to these features to organize images into visually similar groups. The resulting clustered images are validated using demographic data. The cluster model is then applied to six different cities around the world to test the transferability of our methods. Finally, the discovered image clusters are visualized in our customized web interface to enable demographers, social scientists, and economists to understand the organization of a city. Jessica Block, Mehrdad Yazdani, Mai H. Nguyen, Daniel Crawl, Marta Jankowska, John J. Graham, Thomas A. DeFanti, Ilkay Altintas |
eScience | 3 |
| 2016 | Determining feature extractors for unsupervised learning on satellite imagesabstractAdvances in satellite imagery presents unprecedented opportunities for understanding natural and social phenomena at global and regional scales. Although the field of satellite remote sensing has evaluated imperative questions to human and environmental sustainability, scaling those techniques to very high spatial resolutions at regional scales remains a challenge. Satellite imagery is now more accessible with greater spatial, spectral and temporal resolution creating a data bottleneck in identifying the content of images. Because satellite images are unlabeled, unsupervised methods allow us to organize images into coherent groups or clusters. However, the performance of unsupervised methods, like all other machine learning methods, depends on features. Recent studies using features from pre-trained networks have shown promise for learning in new datasets. This suggests that features from pre-trained networks can be used for learning in temporally and spatially dynamic data sources such as satellite imagery. It is not clear, however, which features from which layer and network architecture should be used for learning new tasks. In this paper, we present an approach to evaluate the transferability of features from pre-trained Deep Convolutional Neural Networks for satellite imagery. We explore and evaluate different features and feature combinations extracted from various deep network architectures, and systematically evaluate over 2,000 network-layer combinations. In addition, we test the transferability of our engineered features and learned features from an unlabeled dataset to a different labeled dataset. Our feature engineering and learning are done on the unlabeled Draper Satellite Chronology dataset, and we test on the labeled UC Merced Land dataset to achieve near state-of-the-art classification results. These results suggest that even without any or minimal training, these networks can generalize well to other datasets. This method could be useful in the task of clustering unlabeled images and other unsupervised machine learning tasks. Behnam Hedayatnia, Mehrdad Yazdani, Mai H. Nguyen, Jessica Block, Ilkay Altintas |
IEEE BigData | 3 |
| 2016 | A scalable approach for location-specific detection of Santa Ana conditionsabstractSanta Ana conditions are hot, dry, windy weather conditions that can greatly increase the dangers of wildfires in southern California. We present a machine learning approach to detect Santa Ana conditions based on sensor measurements from weather stations. Cluster analysis is performed on historical weather data to build models to identify Santa Ana patterns. A separate model is built using data from each weather station to capture the patterns specific to the microclimate of each region. Real-time sensor data from a weather station can then be processed to determine if the region surrounding that station is experiencing Santa Ana conditions. Results can be used as a warning system to focus firefighting efforts on regions with increased wildfire risks. Through the use of the Kepler workflow system and distributed computing with Spark, data from several weather stations can be processed in parallel using a scalable clustering algorithm, allowing our approach to scale to large datasets from multiple weather stations. Mai H. Nguyen, Dylan Uys, Daniel Crawl, Charles Cowart, Ilkay Altintas |
IEEE BigData | 1 |
| 2015 | Big data provenance: Challenges, state of the art and opportunitiesabstractAbility to track provenance is a key feature of scientific workflows to support data lineage and reproducibility. The challenges that are introduced by the volume, variety and velocity of Big Data, also pose related challenges for provenance and quality of Big Data, defined as veracity. The increasing size and variety of distributed Big Data provenance information bring new technical challenges and opportunities throughout the provenance lifecycle including recording, querying, sharing and utilization. This paper discusses the challenges and opportunities of Big Data provenance related to the veracity of the datasets themselves and the provenance of the analytical processes that analyze these datasets. It also explains our current efforts towards tracking and utilizing Big Data provenance using workflows as a programming model to analyze Big Data. Jianwu Wang 0001, Daniel Crawl, Shweta Purawat, Mai H. Nguyen, Ilkay Altintas |
IEEE BigData | 4 |
| 1997 | Tau Net A neural network for modeling temporal variability
Mai H. Nguyen, Garrison W. Cottrell |
Neurocomputing | 1 |