VLDB 2026 Research / reviewers in the wild / expert
Thomas Moyer
dblp:13/4686 · also Thomas M. Moyer
· DBLP profile ↗
20ranked-venue papers
3as first author
6since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 7 · 1 first-authorArtificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 3Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Detecting VoIP Data Streams: Approaches Using Hidden Representation LearningabstractThe use of voice-over-IP technology has rapidly expanded over the past several years, and has thus become a significant portion of traffic in the real, complex network environment. Deep packet inspection and middlebox technologies need to analyze call flows in order to perform network management, load-balancing, content monitoring, forensic analysis, and intelligence gathering. Because the session setup and management data can be sent on different ports or out of sync with VoIP call data over the Real-time Transport Protocol (RTP) with low latency, inspection software may miss calls or parts of calls. To solve this problem, we engineered two different deep learning models based on hidden representation learning. MAPLE, a matrix-based encoder which transforms packets into an image representation, uses convolutional neural networks to determine RTP packets from data flow. DATE is a density-analysis based tensor encoder which transforms packet data into a three-dimensional point cloud representation. We then perform density-based clustering over the point clouds as latent representations of the data, and classify packets as RTP or non-RTP based on their statistical clustering features. In this research, we show that these tools may allow a data collection and analysis pipeline to begin detecting and buffering RTP streams for later session association, solving the initial drop problem. MAPLE achieves over ninety-nine percent accuracy in RTP/non-RTP detection. The results of our experiments show that both models can not only classify RTP versus non-RTP packet streams, but could extend to other network traffic classification problems in real deployments of network analysis pipelines. Maya Kapoor, Michael Napolitano, Jonathan Quance, Thomas Moyer, Siddharth Krishnan |
AAAI | 4 |
| 2023 | Defeasible-PROV: Conflict Resolution in Smart Building DevicesabstractProgrammable Logic Controllers (PLCs) are an integral component for managing automation processes of smart buildings. PLCs use protocols which make these control systems vulnerable to many common attacks due to which it is possible to create conflicts on certain devices of smart buildings thereby disrupting functionality. In this paper, we propose DEFEASIBLE-PROV, a system for resolving conflicts in the system by detecting the conflict creating sensors and conflict impacted actuators. Our tool is capable of blocking conflict creating rules in the system. Our evaluation results show that our proposed methodology contributes significantly to conflict resolution in the system. Abdullah Al Farooq, Zac Taylor, Kyle Ruona, Thomas Moyer |
COMPSAC | 4 |
| 2022 | Flurry: A Fast Framework for Provenance Graph Generation for Representation LearningabstractRepresentation learning and deep learning on data provenance graphs has yielded insightful new angles for intrusion detection in cybersecurity systems and is rapidly expanding as a research topic in the scientific community. In order to train the learning models, system execution data must be created and captured which represent realistic cyberattacks of a wide variety. Furthermore, these graphs must contain equally authentic benign workflows relative to the host and its multi-faceted processes. In order to support dynamic generation of provenance graphs to rapidly train new provenance graph learning models, we present Flurry, an end-to-end framework which simulates attack and benign activity and generates provenance graphs in multiple formats exportable to graph learning systems. In this demonstration, we showcase Flurry's ability to simulate both pre-configured and user-defined cyberattacks as well as benign behavior, and convert these captures into provenance graphs. We investigate the spectral properties of these graphs and perform attack classification experiments comparing three graph neural network models. Our results are comparable to those in the original learning models' research papers, showing that the Flurry graphs provide an ideal baseline and an extensible framework for graph representation learning on provenance graphs. In our demo, we show that Flurry will bring brand-new expandability and proficiency to the provenance graph learning community's available data. Maya Kapoor, Joshua Melton, Michael Ridenhour, Thomas Moyer, Siddharth Krishnan |
CIKM | 4 |
| 2022 | Deep Packet Inspection at Scale: Search Optimization Through Locality-Sensitive HashingabstractDeep packet inspection is a primary tool for security specialists, surveillance analysts, and network engineers to lawfully intercept and analyze network traffic. In order to process this data or select streams of interest from the large amount of data flowing in today’s internet, solutions must be capable of identifying network traffic as quickly and accurately as possible. The ever-increasing diversity of data as well as sheer size has rendered the current regular expression matching and filtering solutions ineffective. We propose locality-sensitive hash embedding techniques Alpine and Palm for packet analysis. The fixed size of hashes as well as the adaptability of distance measures is proven to address the network traffic classification problem in our experiments and improves scalability over current state-of-the-art, automata-based search engines. In this paper, we analyze the system’s ability to classify network traffic by many data layer protocols and traffic types with over 99% accuracy. The model is also proven effective in areas where the regular expressions are inapplicable, such as traffic profiling. Finally, we provide real benchmarks of the system’s ability to scale to large signature and hash sets with much improved performance, demonstrating real-world applicability and generalizability of locality-sensitive hashing to deep packet inspection technology. Maya Kapoor, Siddharth Krishnan, Thomas Moyer |
NCA | 3 |
| 2021 | OPD: Network Packet Distribution after Achieving Equilibrium to Mitigate DDOS AttackabstractCrossfire Denial of Service (DDoS) is a new and organized type of attack to make the services of a target organization unavailable. The adversaries, in such attack, adveraremploy BOTs and decoy servers to flood critical links that are used to communicate with the target. In this paper, we propose an optimization framework, OPD (Optimized Packet Distributor), for distributing the packets in the most balanced way after a crossfire DDoS attack gets detected in a network. We formulate the network packet distribution problem as a non-linear link weight function and solved the optimization problem with a very efficient algorithm, Frank Wolf (FW). FW is the most efficient algorithm for solving convex set optimization problem. It achieves network equilibrium (or converges) with very few iterations. OPD will help the commonly used routing algorithm of Software Defined Network (SDN) by providing an optimized number of packets for each link. When a router sends a packet, it will follow the packet distribution proposed by OPD. We simulated a crossfire attack in NS2 simulator. The number of packets traveling in each link is collected for a given time frame from NS2. Based on this packet distribution, OPD proposes desired packet flows for the network. After evaluating the result of OPD, we have found that it can distribute the packets to the links that were less congested during the time of crossfire DDoS attack. Abdullah Al Farooq, Thomas Moyer, Dewan Tanvir Ahmed |
COMPSAC | 2 |
| 2021 | PROV-GEM: Automated Provenance Analysis Framework using Graph EmbeddingsabstractData provenance graphs, detailed traces of system behavior, are a popular construct to analyze and forecast malicious cyber activity like advanced persistent threats (APT). A critical limitation of existing analysis techniques is the lack of an automated analytic framework to predict APTs. In this work, we address that limitation by augmenting efficient capture and storage mechanisms to include automated analysis. Specifically, we propose PROV-GEM, a deep graph learning framework to identify malicious anomalous behavior from provenance data. Since data provenance graphs are complex datasets often expressed as heterogeneous attributed multiplex networks, we use a unified relation-aware embedding framework to capture the necessary contexts and associated interactions between the various entities manifest in the data. Furthermore, provenance graphs by nature are rich detailed structures that are heavily attributed compared to other complex systems that have been used traditionally in graph machine learning applications. Towards that end, our framework uniquely captures “multi-embeddings” that can represent varied contexts of nodes and their multi-faceted nature. We demonstrate the efficacy of our embeddings by applying PROV-GEM to two publicly available APT provenance graph datasets from StreamSpot and Unicorn. PROV-GEM achieves strong performance on both datasets with a 99% accuracy and 97% F1-score on the StreamSpot dataset, and a 97% accuracy and 89% F1-score on the Unicorn dataset, equaling or outperforming comparable state-of-the-art APT threat detection models. Unlike other frameworks, PROV-GEM utilizes an efficient graph convolutional approach coupled with relational self-attention to generate rich graph embeddings that capture the complex topology of data provenance graphs, providing an effective automated analytic framework for APT detection. Maya Kapoor, Joshua Melton, Michael Ridenhour, Siddharth Krishnan, Thomas Moyer |
ICMLA | 5 |
| 2019 | IoTC2: A Formal Method Approach for Detecting Conflicts in Large Scale IoT Systems
Abdullah Al Farooq, Ehab Al-Shaer, Thomas Moyer, Krishna Kant 0001 |
IM | 3 |
| 2018 | Runtime Analysis of Whole-System ProvenanceabstractIdentifying the root cause and impact of a system intrusion remains a foundational challenge in computer security. Digital provenance provides a detailed history of the flow of information within a computing system, connecting suspicious events to their root causes. Although existing provenance-based auditing techniques provide value in forensic analysis, they assume that such analysis takes place only retrospectively. Such post-hoc analysis is insufficient for realtime security applications; moreover, even for forensic tasks, prior provenance collection systems exhibited poor performance and scalability, jeopardizing the timeliness of query responses. We present CamQuery, which provides inline, realtime provenance analysis, making it suitable for implementing security applications. CamQuery is a Linux Security Module that offers support for both userspace and in-kernel execution of analysis applications. We demonstrate the applicability of CamQuery to a variety of runtime security applications including data loss prevention, intrusion detection, and regulatory compliance. In evaluation, we demonstrate that CamQuery reduces the latency of realtime query mechanisms, while imposing minimal overheads on system execution. CamQuery thus enables the further deployment of provenance-based technologies to address central challenges in computer security. Thomas Pasquier, Xueyuan Han, Thomas Moyer, Adam Bates 0001, Olivier Hermant, David M. Eyers, Jean Bacon, Margo I. Seltzer |
CCS | 3 |
| 2018 | Towards Scalable Cluster Auditing through Grammatical Inference over Provenance Graphs
Wajih Ul Hassan, Mark Lemay, Nuraini Aguse, Adam Bates 0001, Thomas Moyer |
NDSS | 5 |
| 2017 | Practical whole-system provenance captureabstractData provenance describes how data came to be in its present form. It includes data sources and the transformations that have been applied to them. Data provenance has many uses, from forensics and security to aiding the reproducibility of scientific experiments. We present CamFlow, a whole-system provenance capture mechanism that integrates easily into a PaaS offering. While there have been several prior whole-system provenance systems that captured a comprehensive, systemic and ubiquitous record of a system's behavior, none have been widely adopted. They either A) impose too much overhead, B) are designed for long-outdated kernel releases and are hard to port to current systems, C) generate too much data, or D) are designed for a single system. CamFlow addresses these shortcoming by: 1) leveraging the latest kernel design advances to achieve efficiency; 2) using a self-contained, easily maintainable implementation relying on a Linux Security Module, NetFilter, and other existing kernel facilities; 3) providing a mechanism to tailor the captured provenance data to the needs of the application; and 4) making it easy to integrate provenance across distributed systems. The provenance we capture is streamed and consumed by tenant-built auditor applications. We illustrate the usability of our implementation by describing three such applications: demonstrating compliance with data regulations; performing fault/intrusion detection; and implementing data loss prevention. We also show how CamFlow can be leveraged to capture meaningful provenance without modifying existing applications. Thomas Pasquier, Xueyuan Han, Mark Goldstein, Thomas Moyer, David M. Eyers, Margo I. Seltzer, Jean Bacon |
SoCC | 4 |
| 2017 | Transparent Web Service Auditing via Network Provenance FunctionsabstractDetecting and explaining the nature of attacks in distributed web services is often difficult -- determining the nature of suspicious activity requires following the trail of an attacker through a chain of heterogeneous software components including load balancers, proxies, worker nodes, and storage services. Unfortunately, existing forensic solutions cannot provide the necessary context to link events across complex workflows, particularly in instances where application layer semantics (e.g., SQL queries, RPCs) are needed to understand the attack. In this work, we present a transparent provenance-based approach for auditing web services through the introduction of Network Provenance Functions (NPFs). NPFs are a distributed architecture for capturing detailed data provenance for web service components, leveraging the key insight that mediation of an application's protocols can be used to infer its activities without requiring invasive instrumentation or developer cooperation. We design and implement NPF with consideration for the complexity of modern cloud-based web services, and evaluate our architecture against a variety of applications including DVDStore, RUBiS, and WikiBench to show that our system imposes as little as 9.3% average end-to-end overhead on connections for realistic workloads. Finally, we consider several scenarios in which our system can be used to concisely explain attacks. NPF thus enables the hassle-free deployment of semantically rich provenance-based auditing for complex applications workflows in the Cloud. Adam Bates 0001, Wajih Ul Hassan, Kevin R. B. Butler, Alin Dobra, Bradley Reaves, Patrick T. Cable II, Thomas Moyer, Nabil Schear |
WWW | 7 |
| 2017 | Taming the Costs of Trustworthy Provenance through Policy ReductionabstractProvenance is an increasingly important tool for understanding and even actively preventing system intrusion, but the excessive storage burden imposed by automatic provenance collection threatens to undermine its value in practice. This situation is made worse by the fact that the majority of this metadata is unlikely to be of interest to an administrator, instead describing system noise or other background activities that are not germane to the forensic investigation. To date, storing data provenance in perpetuity was a necessary concession in even the most advanced provenance tracking systems in order to ensure the completeness of the provenance record for future analyses. In this work, we overcome this obstacle by proposing a policy-based approach to provenance filtering , leveraging the confinement properties provided by Mandatory Access Control (MAC) systems in order to identify and isolate subdomains of system activity for which to collect provenance. We introduce the notion of minimal completeness for provenance graphs, and design and implement a system that provides this property by exclusively collecting provenance for the trusted computing base of a target application. In evaluation, we discover that, while the efficacy of our approach is domain dependent, storage costs can be reduced by as much as 89% in critical scenarios such as provenance tracking in cloud computing data centers. To the best of our knowledge, this is the first policy-based provenance monitor to appear in the literature. Adam Bates 0001, Jing (Dave) Tian, Grant Hernandez, Thomas Moyer, Kevin R. B. Butler, Trent Jaeger |
ACM Trans. Internet Techn. | 4 |
| 2016 | Bootstrapping and maintaining trust in the cloud
Nabil Schear, Patrick T. Cable II, Thomas Moyer, Bryan Richard, Robert Rudd |
ACSAC | 3 |
| 2015 | Trustworthy Whole-System Provenance for the Linux Kernel
Adam Bates 0001, Jing (Dave) Tian, Kevin R. B. Butler, Thomas Moyer |
USENIX Security Symposium | 4 |
| 2012 | Scalable Integrity-Guaranteed AJAX
Thomas Moyer, Trent Jaeger, Patrick D. McDaniel |
APWeb | 1 |
| 2012 | Scalable Web Content AttestationabstractThe web is a primary means of information sharing for most organizations and people. Currently, a recipient of web content knows nothing about the environment in which that information was generated other than the specific server from whence it came (and even that information can be unreliable). In this paper, we develop and evaluate the Spork system that uses the Trusted Platform Module (TPM) to tie the web server integrity state to the web content delivered to browsers, thus allowing a client to verify that the origin of the content was functioning properly when the received content was generated and/or delivered. We discuss the design and implementation of the Spork service and its browser-side Firefox validation extension. In particular, we explore the challenges and solutions of scaling the delivery of mixed static and dynamic content to a large number of clients using exceptionally slow TPM hardware. We perform an in-depth empirical analysis of the Spork system within Apache web servers. This analysis shows Spork can deliver nearly 8,000 static or over 6,500 dynamic integrity-measured web objects per second. More broadly, we identify how TPM-based content web services can scale to large client loads with manageable overheads and deliver integrity-measured content with manageable overhead. Thomas Moyer, Kevin R. B. Butler, Joshua Schiffman, Patrick D. McDaniel, Trent Jaeger |
IEEE Trans. Computers | 1 |
| 2010 | An architecture for enforcing end-to-end access control over web applicationsabstractThe web is now being used as a general platform for hosting distributed applications like wikis, bulletin board messaging systems and collaborative editing environments. Data from multiple applications originating at multiple sources all intermix in a single web browser, making sensitive data stored in the browser subject to a broad milieu of attacks (cross-site scripting, cross-site request forgery and others). The fundamental problem is that existing web infrastructure provides no means for enforcing end-to-end security on data. To solve this we design an architecture using mandatory access control (MAC) enforcement. We overcome the limitations of traditional MAC systems, implemented solely at the operating system layer, by unifying MAC enforcement across virtual machine, operating system, networking and application layers. We implement our architecture using Xen virtual machine management, SELinux at the operating system layer, labeled IPsec for networking and our own label-enforcing web browser, called FlowwolF. We tested our implementation and find that it performs well, supporting data intermixing while still providing end-to-end security guarantees. Boniface Hicks, Sandra Julieta Rueda, Dave King 0002, Thomas Moyer, Joshua Schiffman, Yogesh Sreenivasan, Patrick D. McDaniel, Trent Jaeger |
SACMAT | 4 |
| 2009 | Scalable Web Content AttestationabstractThe Web is a primary means of information sharing for most organizations and people. Currently, a recipient of Web content knows nothing about the environment in which that information was generated other than the specific server from whence it came (and even that information can be unreliable). In this paper, we develop and evaluate the Spork system that uses the trusted platform module (TPM) to tie the Web server integrity state to the Web content delivered to browsers, thus allowing a client to verify that the origin of the content was functioning properly when the received content was generated and/or delivered. We discuss the design and implementation of the Spork service and its browser-side Firefox validation extension. In particular, we explore the challenges and solutions of scaling the delivery of mixed static and dynamic content using exceptionally slow TPM hardware. We perform an in-depth empirical analysis of the Spork system within Apache Web servers. This analysis shows Spork can deliver nearly 8,000 static or over 7,000 dynamic integrity-measured Web objects per-second. More broadly, we identify how TPM-based content Web services can scale with manageable overheads and deliver integrity-measured content with manageable overhead. Thomas Moyer, Kevin R. B. Butler, Joshua Schiffman, Patrick D. McDaniel, Trent Jaeger |
ACSAC | 1 |
| 2009 | Justifying Integrity Using a Virtual Machine VerifierabstractEmerging distributed computing architectures, such as grid and cloud computing, depend on the high integrity execution of each system in the computation. While integrity measurement enables systems to generate proofs of their integrity to remote parties, we find that current integrity measurement approaches are insufficient to prove runtime integrity for systems in these architectures. Integrity measurement approaches that are flexible enough have an incomplete view of runtime integrity, possibly leading to false integrity claims, and approaches that provide comprehensive integrity do so only for computing environments that are too restrictive. In this paper, we propose an architecture for building comprehensive runtime integrity proofs for general purpose systems in distributed computing architectures. In this architecture, we strive for classical integrity, using an approximation of the Clark-Wilson integrity model as our target. Key to building such integrity proofs is a carefully crafted host system whose long-term integrity can be justified easily using current techniques and a new component, called a VM verifier, which comprehensively enforces our integrity target on VMs. We have built a prototype based on the Xen virtual machine system for SELinux VMs, and find that distributed compilation can be implemented, providing accurate proofs of our integrity target with less than 4% overhead. Joshua Schiffman, Thomas Moyer, Christopher Shal, Trent Jaeger, Patrick D. McDaniel |
ACSAC | 2 |
| 2009 | Configuration management at massive scale: system design and experienceabstractThe development and maintenance of network device configurations is one of the central challenges faced by large network providers. Current network management systems fail to meet this challenge primarily because of their inability to adapt to rapidly evolving customer and provider-network needs, and because of mismatches between the conceptual models of the tools and the services they must support. In this paper, we present the Presto configuration management system that attempts to address these failings in a comprehensive and flexible way. Developed for and used during the last 5 years within a large ISP network, Presto constructs device-native configurations based on the composition of configlets representing different services or service options. Configlets are compiled by extracting and manipulating data from external systems as directed by the Presto configuration scripting and template language. We outline the configuration management needs of large-scale network providers, introduce the PRESTO system and configuration language, and reflect upon our experiences developing PRESTO configured VPN and VoIP services. In doing so, we describe how PRESTO promotes healthy configuration management practices. William Enck, Thomas Moyer, Patrick D. McDaniel, Subhabrata Sen, Panagiotis Sebos, Sylke Spoerel, Albert G. Greenberg, Yu-Wei Eric Sung, Sanjay G. Rao, William Aiello |
IEEE J. Sel. Areas Commun. | 2 |