VLDB 2026 Research / reviewers in the wild / expert
Akul Goyal
dblp:334/9010
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0003-2484-6185ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | What We Talk About When We Talk About Logs: Understanding the Effects of Dataset Quality on Endpoint Threat Detection ResearchabstractEndpoint threat detection research hinges on the availability of worthwhile evaluation benchmarks, but experimenters' understanding of the contents of benchmark datasets is often limited. Typically, attention is only paid to the realism of attack behaviors, which comprises only a small percentage of the audit logs in the dataset, while other characteristics of the data are inscrutable and unknown. We propose a new set of questions for what to talk about when we talk about logs (i.e., datasets): What activities are in the dataset? We introduce a novel visualization that succinctly represents the totality of 100+ GB datasets by plotting the occurrence of provenance graph neighborhoods in a time series. How synthetic is the background activity? We perform autocorrelation analysis of provenance neighborhoods in the training split to identify process behaviors that occur at predictable intervals in the test split. Finally, How conspicuous is the malicious activity? We quantify the proportion of attack behaviors that are observed as benign neighborhoods in the training split as compared to previously-unseen attack neighborhoods. We then validate these questions by profiling the classification performance of state-of-the-art intrusion detection systems (R-CAID, FLASH, KAIROS, GNN) against a battery of public benchmark datasets (DARPA Transparent Computing and OpTC, ATLAS, ATLASv2). We demonstrate that synthetic background activities dramatically inflate True Negative Rates, while conspicuous malicious activities artificially boost True Positive Rates. Further, by explicitly controlling for these factors, we provide a more holistic picture of classifier performance. This work will elevate the dialogue surrounding threat detection datasets and will increase the rigor of threat detection experiments. Jason Liu 0002, Muhammad Adil Inam, Akul Goyal, Andy Riddle, Kim Westfall, Adam Bates 0001 |
SP | 3 |
| 2025 | A Software-Based Approach for Detecting Tire Blowouts Through Inertial Measurement Anomalies in Autonomous Vehicles with Machine LearningabstractA prominent issue among ground vehicles is tire blowout or other catastrophic tire failure, which occurs when at least one of the vehicle's tires rapidly deflates or detaches completely. This can cause a driver to lose control of the vehicle's steering which can lead to dangerous, potentially fatal accidents. These incidents can become even more perilous when tire failures occur on vehicles without a human driver behind the wheel, specifically in autonomous vehicles. Therefore, we propose a method for detecting tire failures by analyzing data from the inertial measurement unit (IMU) for anomalies in accelerometer and gyroscope values, which signify such incidents. To test our method's accuracy, we simulate tire failure in highway conditions by physically detaching a wheel on our RACECAR, a 1/14 scale model autonomous car, while it performs line following on a curved path. We develop and compare a threshold-based algorithm, a Z-score-based algorithm, a convolutional neural network (CNN), and a recurrent neural network (RNN) to predict when blowouts occur during our trials, the latter three of which validate the use of IMUs in the detection of tire failures. Viraaj Seth, Raj Petkar, Saketh Ayyagari, Stephen Cai, Akul Goyal, Sven Kirchner, Ross Greer |
VTC2025-Spring | 5 |
| 2024 | R-CAID: Embedding Root Cause Analysis within Provenance-based Intrusion DetectionabstractIn modern enterprise security, endpoint detection products fire an alert when process activity matches known attack behavior patterns. Human analysts then perform Root Cause Analysis (RCA) over event logs to determine if the alert is indicative of an actual attack. Data Provenance can help to automate RCA by representing event logs as a causal dependency graphs; in fact, researchers are now examining whether provenance-based anomaly detection should replace pattern-based detection altogether. Unfortunately, we observe that current approaches leverage off-the-shelf graph embedding techniques that are unable to associate events with their root causes. This shortcoming not only fails to capitalize on the RCA capabilities of provenance, but also leaves provenance-based IDS vulnerable to mimicry and evasion attacks.This work presents the design and implementation of R-CAID, a novel approach to incorporate RCA into provenance-based IDS. R-CAID precomputes each node’s root causes during graph construction, then directly links those nodes to their root causes during embedding. Further, R-CAID’s classification model is node/process-level, rather than graph/system-level, bringing it more in line with the precision of commercial systems. Under a passive adversary model, we find that R-CAID consistently outperforms baseline graph neural networks, sequence-based log IDS, and even a commercial endpoint detection system. Under a white-box active adversary model, R-CAID maintains a high level of performance (e.g., for DARPA Theia, 0.94 AUC adversarial down from 0.99 passive). R-CAID achieves this by associating each system entity with its immutable and unforgeable root causes, preventing adversaries from being able to masquerade as legitimate processes. This work is thus the first to demonstrate the promise of provenance-based IDS in a manner that avoids the pitfalls of mimicry and evasion. Akul Goyal, Gang Wang 0011, Adam Bates 0001 |
SP | 1 |
| 2023 | Sometimes, You Aren't What You Do: Mimicry Attacks against Provenance Graph Host Intrusion Detection Systems
Akul Goyal, Xueyuan Han, Gang Wang 0011, Adam Bates 0001 |
NDSS | 1 |
| 2023 | SoK: History is a Vast Early Warning System: Auditing the Provenance of System IntrusionsabstractAuditing, a central pillar of operating system security, has only recently come into its own as an active area of public research. This resurgent interest is due in large part to the notion of data provenance, a technique that iteratively parses audit log entries into a dependency graph that explains the history of system execution. Provenance facilitates precise threat detection and investigation through causal analysis of sophisticated intrusion behaviors. However, the absence of a foundational audit literature, combined with the rapid publication of recent findings, makes it difficult to gain a holistic picture of advancements and open challenges in the area.In this work, we survey and categorize the provenance-based system auditing literature, distilling contributions into a layered taxonomy based on the audit log capture and analysis pipeline. Recognizing that the Reduction Layer remains a key obstacle to the further proliferation of causal analysis technologies, we delve further on this issue by conducting an ambitious independent evaluation of 8 exemplar reduction techniques against the recently-released DARPA Transparent Computing datasets. Our experiments uncover that past approaches frequently prune an overlapping set of activities from audit logs, reducing the synergistic benefits from applying them in tandem; further, we observe an inverse relation between storage efficiency and anomaly detection performance. However, we also observe that log reduction techniques are able to synergize effectively with data compression, potentially reducing log retention costs by multiple orders of magnitude. We conclude by discussing promising future directions for the field. Muhammad Adil Inam, Yinfang Chen, Akul Goyal, Jason Liu 0002, Jaron Mink, Noor Michael, Sneha Gaur, Adam Bates 0001, Wajih Ul Hassan |
SP | 3 |
| 2022 | FAuST: Striking a Bargain between Forensic Auditing's Security and ThroughputabstractSystem logs are invaluable to forensic audits, but grow so large that in practice fine-grained logs are quickly discarded – if captured at all – preventing the real-world use of the provenance-based investigation techniques that have gained popularity in the literature. Encouragingly, forensically-informed methods for reducing the size of system logs are a subject of frequent study. Unfortunately, many of these techniques are designed for offline reduction in a central server, meaning that the up-front cost of log capture, storage, and transmission must still be paid at the endpoints. Moreover, to date these techniques exist as isolated (and, often, closed-source) implementations; there does not exist a comprehensive framework through which the combined benefits of multiple log reduction techniques can be enjoyed. Muhammad Adil Inam, Akul Goyal, Jason Liu 0002, Jaron Mink, Noor Michael, Sneha Gaur, Adam Bates 0001, Wajih Ul Hassan |
ACSAC | 2 |