VLDB 2026 Research / reviewers in the wild / expert
Austin Coursey
dblp:301/9228
· DBLP profile ↗
12ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0003-1774-6442ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Novel Approach to Evaluating the Effectiveness of Large Language Models for Multimodal Analysis of Embodied Learning in ClassroomsabstractThis paper presents an approach that uses Large Language Models (LLMs) as late-fusion interpreters to synthesize multimodal signals from embodied classroom activities and infer students’ metacognitive behaviors. Our multimodal pipeline analyzes students’ movements, gaze, gestures, and speech within a mixed-reality simulation displayed on a classroom screen to support enactment and learning. Vision- and speech-derived features are fused at the interpretive layer via zero-shot prompting, self-consistency reasoning, and targeted prompt engineering to derive planning, enacting, monitoring, reflecting, and interacting behaviors. We investigate whether LLMs can reliably integrate modality-specific analytics to produce accurate behavioral labeling and whether an LLM-as-a-Judge can validate them at scale. To address scalability and reduce human burden, we introduce an automated evaluation protocol employing LLM-as-a-Judge to assess classification quality, enabling rapid, iterative benchmarking of model variants and prompt strategies. Using a balanced corpus of human-validated segments and perturbed controls, we compare text-only language models (e.g., GPT-5) with visual–language models (e.g., Qwen2.5-VL) that incorporate direct visual processing. Results indicate late-fusion, text-based LLMs can outperform VLMs on behavior judgment without raw video, and precision- or recall-oriented prompts adjust decision boundaries for subtle or brief segments. These findings position LLMs as effective late-fusion mechanisms for multimodal learning analytics and demonstrate the viability of LLM-as-a-Judge for scalable, human-in-the-loop evaluation. Joyce Horn Fonteles, Nithin Sivakumaran, Clayton Cohn, Austin Coursey, Shoubin Yu, Elias Stengel-Eskin, T. S. Ashwin, Mohit Bansal, Gautam Biswas |
LAK | 4 |
| 2025 | Data-Driven Fault Detection and Isolation Enhanced with System Structural Relationships (DX Competition)abstractFault detection and isolation are becoming increasingly important as modern systems become more complex. To encourage the development of new fault detection solutions that can operate with limited noisy data and an incomplete mathematical model, the DX 2025 LiU-ICE competition for diagnosis of the air path of an internal combustion engine was introduced. In this paper, we present our winning solution to this competition. Our fault detection architecture starts with a semi-supervised Transformer Autoencoder trained to reconstruct nominal data. Detected faults are then passed through a rule-based fault persistence filter that aims to suppress false positives. Once a fault is detected, we use four neural networks trained to estimate features determined from structural analysis of a partial system model. The residuals of these networks are fed to a supervised fault classification network that estimates the fault probabilities. With this architecture, we achieved an 87% detection rate with a 0% false alarm rate on the provided competition data. Additionally, our isolation architecture assigned the correct fault 73.8% probabilty on average. On unseen competition data from a new driving cycle, we achieved a 100% detection rate and assigned the correct fault 66.2% probability on average. On the other hand, the Transformer Autoencoder failed to transfer to the new driving conditions, causing many false alarms. We discuss ways future work can reduce this. Austin Coursey, Abel Díaz-González, Marcos Quiñones-Grueiro, Gautam Biswas |
DX | 1 |
| 2025 | A Data-Driven Particle Filter Approach for System-Level Prediction of Remaining Useful LifeabstractAccurate estimation of the remaining useful life (RUL) of industrial systems is a critical component of predictive maintenance strategies. This work presents a data-driven method for RUL prediction that also quantifies uncertainty, drawing inspiration from model-based particle filtering techniques. Instead of simulating system state transitions, we model degradation as a stochastic process governed by performance metrics and use a Bayesian particle filtering framework to infer its underlying parameters. Our approach bypasses traditional state-space modeling by directly estimating the end-of-life distribution from observed performance data. Key characteristics of the filter, such as propagation noise and observation correction strength, are adapted over time based on current observations and past predictive performance, enabling better capture of future uncertainty. We evaluate the proposed method using an unmanned aerial vehicle simulation dataset developed for system-level prognostics research, which includes high-fidelity degradation signals and ground-truth system performance metrics for validating predictive accuracy. Abel Díaz-González, Austin Coursey, Marcos Quiñones-Grueiro, Gautam Biswas |
DX | 2 |
| 2025 | Safe to Fly? Real-Time Flight Mission Feasibility Assessment for Drone Package Delivery OperationsabstractEnsuring flight safety for small unmanned aerial systems (sUAS) requires continuous in-flight monitoring and decision-making, as unexpected events can alter power consumption and deplete battery energy faster than anticipated. Such events may result in insufficient battery capacity to complete a mission, thereby compromising flight safety. In this paper, we present an online feasibility assessment and contingency management framework that continuously monitors the aircraft’s battery state and the energy required to complete the flight in real-time, which enables informed decision-making to enhance flight safety. The framework consists of two main components: power consumption prediction and battery voltage trajectory prediction. The power consumption prediction is conducted using a model that is based on momentum theory, while the voltage trajectory prediction is performed using a Neural Ordinary Differential Equation (Neural ODE)-based data-driven model. By integrating these two components, the framework evaluates the feasibility of a flight mission in real time and determines whether to proceed with the mission or initiate rerouting. We evaluate the framework’s performance in a drone delivery scenario in the Dallas–Fort Worth (DFW) area, where the aircraft encounters an unexpected energy depletion event mid-flight. The proposed framework is tasked with assessing the feasibility of completing the mission and, if necessary, rerouting the aircraft for an emergency landing. The results demonstrate that the framework accurately and efficiently detects energy insufficiencies in real-time and re-routes the aircraft to a [3] predefined emergency landing site. Abenezer Taye, Austin Coursey, Marcos Quiñones-Grueiro, Gautam Biswas |
DX | 2 |
| 2024 | Quantifying the Sim-To-Real Gap in UAV Disturbance RejectionabstractDue to the safety risks and training sample inefficiency, it is often preferred to develop controllers in simulation. However, minor differences between the simulation and the real world can cause a significant sim-to-real gap. This gap can reduce the effectiveness of the developed controller. In this paper, we examine a case study of transferring an octorotor reinforcement learning controller from simulation to the real world. First, we quantify the effectiveness of the real-world transfer by examining safety metrics. We find that although there is a noticeable (around 100%) increase in deviation in real flights, this deviation may not be considered unsafe, as it will be within > 2m safety corridors. Then, we estimate the densities of the measurement distributions and compare the Jensen-Shannon divergences of simulated and real measurements. From this, we show that the vehicle’s orientation is significantly different between simulated and real flights. We attribute this to a different flight mode in real flights where the vehicle turns to face the next waypoint. We also find that the reinforcement learning controller actions appear to correctly counteract disturbance forces. Then, we analyze the errors of a measurement autoencoder and state transition model neural network applied to real data. We find that these models further reinforce the difference between the simulated and real attitude control, showing the errors directly on the flight paths. Finally, we discuss important lessons learned in the sim-to-real transfer of our controller. Austin Coursey, Marcos Quiñones-Grueiro, Gautam Biswas |
DX | 1 |
| 2024 | Data-Driven RUL Prediction Using Performance Metrics (Short Paper)
Abel Díaz-González, Austin Coursey, Marcos Quiñones-Grueiro, Chetan S. Kulkarni, Gautam Biswas |
DX | 2 |
| 2024 | FT-AED: Benchmark Dataset for Early Freeway Traffic Anomalous Event DetectionabstractEarly and accurate detection of anomalous events on the freeway, such as accidents, can improve emergency response and clearance. However, existing delays and mistakes from manual crash reporting records make it a difficult problem to solve. Current large-scale freeway traffic datasets are not designed for anomaly detection and ignore these challenges. In this paper, we introduce the first large-scale lane-level freeway traffic dataset for anomaly detection. Our dataset consists of a month of weekday radar detection sensor data collected in 4 lanes along an 18-mile stretch of Interstate 24 heading toward Nashville, TN, comprising over 3.7 million sensor measurements. We also collect official crash reports from the Tennessee Department of Transportation Traffic Management Center and manually label all other potential anomalies in the dataset. To show the potential for our dataset to be used in future machine learning and traffic research, we benchmark numerous deep learning anomaly detection models on our dataset. We find that unsupervised graph neural network autoencoders are a promising solution for this problem and that ignoring spatial relationships leads to decreased performance. We demonstrate that our methods can reduce reporting delays by over 10 minutes on average while detecting 75% of crashes. Our dataset and all preprocessing code needed to get started are publicly released at https://vu.edu/ft-aed/ to facilitate future research. Austin Coursey, Junyi Ji, Marcos Quiñones-Grueiro, William Barbour, Yuhang Zhang 0009, Tyler Derr, Gautam Biswas, Daniel B. Work |
NeurIPS | 1 |
| 2023 | Aspirations and Practice of ML Model Documentation: Moving the Needle with Nudging and TraceabilityabstractThe documentation practice for machine-learned (ML) models often falls short of established practices for traditional software, which impedes model accountability and inadvertently abets inappropriate or misuse of models. Recently, model cards, a proposal for model documentation, have attracted notable attention, but their impact on the actual practice is unclear. In this work, we systematically study the model documentation in the field and investigate how to encourage more responsible and accountable documentation practice. Our analysis of publicly available model cards reveals a substantial gap between the proposal and the practice. We then design a tool named DocML aiming to (1) nudge the data scientists to comply with the model cards proposal during the model development, especially the sections related to ethics, and (2) assess and manage the documentation quality. A lab study reveals the benefit of our tool towards long-term documentation quality and accountability. Avinash Bhat, Austin Coursey, Grace Hu, Sixian Li, Nadia Nahar, Shurui Zhou, Christian Kästner, Jin L. C. Guo |
CHI | 2 |
| 2023 | Anomaly Detection for Multi-Zone Buildings Using Cluster-Trained LSTM AutoencodersabstractThe optimal energy performance of building operations is affected by component faults, which may go unnoticed for long periods of time. Significant energy savings can be achieved if faulty behaviors are detected and rectified in a timely manner. In this work, we adopt an unsupervised approach for anomaly detection that combines automatic data engineering using clustering methods with Long-Short Term Memory (LSTM)-based Autoencoders. First, data engineering is used to extract multiple operating modes from nominal data of building operations. Then, an LSTM-based Autoencoder is trained to capture the characteristics of non-linear and temporal dynamics for each operating mode. Finally, the ensemble of models can be used for anomaly detection after training has been completed. We benchmark variants of our approach against state-of-the-art Autoencoders for anomaly detection by using a recently developed experimental dataset provided by the ASHRAE Research Project RP-1312. Unsupervised anomaly detection is a challenging task due to the lack of faulty labels and the need to identify faults while avoiding false alarms. Our novel approach improves the average true positive rate for fault detection by 11.4% against a state-of-the-art plain LSTM Autoencoder while keeping the false alarm rate around 5 % without having to use labeled fault data. Austin Coursey, Marcos Quiñones-Grueiro, Gautam Biswas, Timothy Darrah |
CoDIT | 1 |
| 2023 | On Learning Data-Driven Models For In-Flight Drone Battery Discharge Estimation From Real DataabstractAccurate estimation of the battery state of charge (SOC) for unmanned aerial vehicles (UAV) in-flight monitoring is essential for the safety and survivability of the system. Successful physics-based models of the battery have been developed in the past, however, these models do not take into account the effects of mission profile and environmental conditions during flight on the battery power consumption. Recently, data-driven methods have become popular given their ease of use and scalability. Yet, most benchmarking experiments have been conducted on simulated battery datasets. In this work, we compare different data-driven models for battery SOC estimation of a hexacopter UAV system using real flight data. We analyze the importance of a number of flight variables under different environmental conditions to determine the factors that affect battery SOC over the course of the flight. Our experiments demonstrate that additional flight variables are necessary to create an accurate SOC estimation model through data-driven methods. Austin Coursey, Marcos Quiñones-Grueiro, Gautam Biswas |
SMARTCOMP | 1 |
| 2023 | Large-scale End-of-Life Prediction of Hard Disks in Distributed DatacentersabstractOn a daily basis, data centers process huge volumes of data backed by the proliferation of inexpensive hard disks. Data stored in these disks serve a range of critical functional needs from financial, and healthcare to aerospace. As such, premature disk failure and consequent loss of data can be catastrophic. To mitigate the risk of failures, cloud storage providers perform condition-based monitoring and replace hard disks before they fail. By estimating the remaining useful life of hard disk drives, one can predict the time-to-failure of a particular device and replace it at the right time, ensuring maximum utilization whilst reducing operational costs. In this work, large-scale predictive analyses are performed using severely skewed health statistics data by incorporating customized feature engineering and a suite of sequence learners. Past work suggests using LSTMs as an excellent approach to predicting remaining useful life. To this end, we present an encoder-decoder LSTM model where the context gained from understanding health statistics sequences aid in predicting an output sequence of the number of days remaining before a disk potentially fails. The models developed in this work are trained and tested across an exhaustive set of all of the 10 years of S.M.A.R.T. health data in circulation from Backblaze and on a wide variety of disk instances. It closes the knowledge gap on what full-scale training achieves on thousands of devices and advances the state-of-the-art by providing tangible metrics for evaluation and generalization for practitioners looking to extend their workflow to all years of health data in circulation across disk manufacturers. The encoder-decoder LSTM posted an RMSE of 0.83 during training and 0.86 during testing over the exhaustive 10-year data while being able to generalize competitively over other drives from the Seagate family. Rohan Mohapatra, Austin Coursey, Saptarshi Sengupta |
SMARTCOMP | 2 |
| 2021 | Remaining Useful Life Estimation of Hard Disk Drives using Bidirectional LSTM NetworksabstractPhysical and cloud storage services are well-served by functioning and reliable high-volume storage systems. Recent observations point to hard disk reliability as one of the most pressing reliability issues in data centers containing massive volumes of storage devices such as HDDs. In this regard, early detection of impending failure at the disk level aids in reducing system downtime and reduces operational loss making proactive health monitoring a priority for AIOps in such settings. In this work, we introduce methods of extracting meaningful attributes associated with operational failure and of pre-processing the highly imbalanced health statistics data for subsequent prediction tasks using data-driven approaches. We use a Bidirectional LSTM with a multi-day look back period to learn the temporal progression of health indicators and baseline them against vanilla LSTM and Random Forest models to come up with several key metrics that establish the usefulness of and superiority of our model under some tightly defined operational constraints. For example, using a 15 day look back period, our approach can predict the occurrence of disk failure with an accuracy of 96.4% considering test data 60 days before failure. This helps to alert operations maintenance well in-advance about potential mitigation needs. In addition, our model reports a mean absolute error of 0.12 for predicting failure up to 60 days in advance, placing it among the state-of-the-art in recent literature. Austin Coursey, Gopal Nath, Srikanth Prabhu, Saptarshi Sengupta |
IEEE BigData | 1 |