EDBT 2026 Demo / reviewers in the wild / expert
Ashish Gehani
dblp:21/2267
· DBLP profile ↗
43ranked-venue papers
11as first author
15since 2021 · last 2025
0000-0002-3940-2467ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 19 · 7 first-author · 5 since 2021Software engineering, systems software and programming languages · 9 · 1 first-author · 4 since 2021Systems, architecture and hardware · 6 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SCALPEL: Structured Content Access Logging and Pruning for Efficient LayoutsabstractContainerizing scientific workflows helps ensure their reproducibility. Including all data required for deterministic re-execution aids the process. In data-intensive climate science and other high-performance computing domains, workflows routinely process large-scale data archives through parallel frameworks such as Dask. Packaging these archives inflates container images, with vast swaths of data that are never accessed by the application driving up transfer costs and hindering deployment. We present SCALPEL, a framework for semantic carving - selectively retaining only the data an application actually consumes during analysis, while excluding data that is accessed but not used.We provide two complementary carving modes, each operating at a different level of observability: (1) value-level carving, which traces data access at the resolution of individual rows (in tabular datasets) or specific multidimensional elements (in array datasets), enabling deterministic re-execution with precisely the same inputs; and (2) partition/chunk-level carving, a configurable alternative that tracks data access at the coarser granularity of table partitions or array chunks, facilitating flexible re-execution that can accommodate additional inputs within these broader segments. By interposing on HDF5 I/O operations and Dask’s execution layer, both widely adopted technologies for large-scale data storage and parallel computation, we capture application-specific data access patterns, achieving reductions in container size by several orders of magnitude. We demonstrate these reductions in data-intensive scientific workflows, including precipitation-driven climate modeling and other geophysical workloads. Raffay Atiq, Ashish Gehani, Tanu Malik, Fareed Zaffar |
eScience | 2 |
| 2024 | Access-Based Carving of Data for Efficient Reproducibility of ContainersabstractScientific applications often depend on data produced from computational models. Model-generated data can be prohibitively large. Current mechanisms for sharing and distributing applications, such as containers, assume all model data is saved and included with a program or is downloaded during build time to support its successful re-execution. However, including model data increases the sizes of containers. This increases the cost and time required for deployment and further reuse. We present ABCD (Access-Based Carving of Data), a framework for specializing I/O libraries which, given an application, automates the process of identifying and including only a subset of the data accessed by the program. To do this we show how such specialization can be achieved at two levels of granularity: at a library level and at a system call level. The different levels help to include data for a single parameter run or over several parameter runs. We show several orders of magnitude reduction in data size via the specialization of HDF5 I/O libraries associated with model-based data-intensive applications, such as those operating on precipitation and geophysical data. Rohan Tikmany, Aniket Modi, Raffay Atiq, Moaz Reyad, Ashish Gehani, Tanu Malik |
CCGrid | 5 |
| 2024 | 6G-XSec: Explainable Edge Security for Emerging OpenRAN ArchitecturesabstractThe evolution from 5G to 6G cellular networks signifies a crucial advancement towards enhanced robustness and automation driven by the promise of ubiquitous Artificial Intelligence (AI) to overhaul network operations, commonly referred to as AIOps. However, 6G network operators also need to deal with evolving threats at the edge to ensure data integrity and availability. We introduce 6G-XSEC, the first framework that seeks to automatically monitor, analyze, and explain anomalies and threats at the cellular network edge. Our framework enhances the emerging Open Radio Access Network (O-RAN) control plane with run-time analytic capabilities and explainability. A distinguishing aspect of our framework is the use of expert referencing, a coupling of lightweight unsupervised deep learning-based anomaly detection with large language models (LLMs) to first detect, analyze, and subsequently explain complicated real-world cellular threats and anomalies at run-time, based on enhanced security telemetry from the O-RAN data plane. We build a prototype 6G-XSEC framework and evaluate it against 5 end-to-end cellular attacks from the literature, achieving 100% detection rate with our best model. We also propose effective LLM prompt templates for attack analysis and present qualitative results from 5 popular LLMs. Haohuang Wen, Prakhar Sharma, Vinod Yegneswaran, Phillip A. Porras, Ashish Gehani, Zhiqiang Lin 0001 |
HotNets | 5 |
| 2024 | Kondo: Efficient Provenance-Driven Data DebloatingabstractIsolation increases upfront costs of provisioning containers. This is due to unnecessary software and data in container images. While several static and dynamic analysis methods for pruning unnecessary software are known, less attention has been paid to pruning unnecessary data. In this paper, we address the problem of determining and reducing unused data within a containerized application. Current data lineage methods can be used to detect data files that are never accessed in any of the observed runs, but this leads to a pessimistic amount of debloating. It is our observation that while an application may access a data file, it often accesses only a small portion of it over all its runs. Based on this observation, we present an approach and a tool Kondo, which aims to identify the set of all possible offsets that could be accessed within the data files over all executions of the application. Kondo works by fuzzing the parameter inputs to the application, and running it on the fuzzed inputs, with vastly fewer runs than brute force execution over all possible parameter valuations. Our evaluation on realistic benchmarks shows that Kondo is able to achieve 63% reduction in data file sizes and 98% recall against the set of all required offsets, on average. Aniket Modi, Rohan Tikmany, Tanu Malik, Raghavan Komondoor, Ashish Gehani, Deepak D'Souza |
ICDE | 5 |
| 2024 | 5G-Spector: An O-RAN Compliant Layer-3 Cellular Attack Detection Service
Haohuang Wen, Phillip A. Porras, Vinod Yegneswaran, Ashish Gehani, Zhiqiang Lin 0001 |
NDSS | 4 |
| 2023 | SoK: A Tale of Reduction, Security, and Correctness - Evaluating Program Debloating Paradigms and Their Compositions
Muaz Ali, M. Faraz Karim, Ayesha Naeem, Rukhshan Haroon, Huzaifah Nadeem, Waseem Sabir, Fahad Shaon, Fareed Zaffar, Vinod Yegneswaran, Ashish Gehani, Sazzadur Rahaman |
ESORICS (4) | 12 |
| 2023 | autoMPI: Automated Multiple Perspective Attack Investigation With Semantics Aware Execution PartitioningabstractMultiple Perspective attack Investigation (MPI) is a technique to partition application dependencies based on high-level semantics. It facilitates provenance analysis by generating succinct causal graphs. It involves an annotation process that identifies variables and data structures corresponding to the partitions and the communication channels between them. Though the amount of annotation is small, this process requires a detailed understanding of the source code. In this work, autoMPI, we extend the capability ofMPIby automating the identifying annotation requirements. We leverage a hybrid analysis approach, performing a differential analysis based on crafted inputs. Static analysis is conducted to identify the annotation sites within the application code afterward automatically. Our evaluation shows the proposed approach can significantly facilitate the annotation process. It correctly identifies all required annotation sites within an average 16 seconds analysis time for the majority of analyzed programs with average precision and recall 72.5% and 100%, respectively. Mohannad Alhanahnah, Shiqing Ma, Ashish Gehani, Gabriela F. Ciocarlie, Vinod Yegneswaran, Somesh Jha, Xiangyu Zhang 0001 |
IEEE Trans. Software Eng. | 3 |
| 2022 | Provenance-based Workflow Diagnostics Using Program SpecificationabstractWorkflow management systems (WMS) help automate and coordinate scientific modules and monitor their execution. WMSes are also used to repeat a workflow application with different inputs to test sensitivity and reproducibility of runs. However, when differences arise in outputs across runs, current WMSes do not audit sufficient provenance metadata to determine where the execution first differed. This increases diagnostic time and leads to poor quality diagnostic results. In this paper, we use program specification to precisely determine locations where workflow execution differs. We use existing provenance audited to isolate modules where execution differs. We show that using program specification comes at some increased storage overhead due to mapping of provenance data flows onto program specification, but leads to better quality diagnostics in terms of the number of differences found and their location relative to comparing provenance metadata audited within current WMSes. Tanu Malik, Iyad Kanj, Ashish Gehani |
HIPC | 4 |
| 2022 | PACED: Provenance-based Automated Container Escape DetectionabstractThe security of container-based microservices relies heavily on the isolation of operating system resources that is provided by namespaces. However, vulnerabilities exist in the isolation of containers that may be exploited by attackers to gain access to the host. These are commonly referred to as container escape attacks. While prior work has identified vulnerabilities in namespace isolation, no general container escape detection and warning system has been presented. We present Paced, a novel, realtime system to detect container-escape attacks. We define what constitutes a cross-namespace event and how such events can be used to detect a container escape attack. We develop a provenance-based approach to isolate cross-namespace events and propose a rule—privileged_flow—to detect attacks on Docker and Kubernetes environments. We evaluate our detection method on a suite of contemporary CVEs with container escape exploits, bad container configurations, and benchmarks. Paced achieves near-perfect accuracy with no false negatives. We release our implementation and datasets as free, open-source software. Mashal Abbas, Shahpar Khan, Abdul Monum, Fareed Zaffar, Rashid Tahir, David M. Eyers, Hassaan Irshad, Ashish Gehani, Vinod Yegneswaran, Thomas Pasquier |
IC2E | 8 |
| 2022 | Trimmer: Context-Specific Code ReductionabstractWe present Trimmer, a state-of-the-art tool for reducing code size. Trimmer reduces code sizes by specializing programs with respect to constant inputs provided by developers. The static data can be provided as command-line options or through configuration files. The constants define the features that must be retained, which in turn determine the features that are unused in a specific deployment (and can therefore be removed). Trimmer includes sophisticated compiler transformations for input specialization, supports precise yet efficient context-sensitive inter-procedural constant propagation, and introduces a custom loop unroller. Trimmer is easy-to-use and extensively parameterized. We discuss how Trimmer can be configured by developers to explicitly trade analysis precision and specialization time. We also provide a high-level description of Trimmer’s static analysis passes. The source code is publicly available at: https://github.com/ashish-gehani/Trimmer. A video demonstration can be found here: https://youtu.be/6pAuJ68INnI. Aatira Anum Ahmad, Mubashir Anwar, Hashim Sharif, Ashish Gehani, Fareed Zaffar |
ASE | 4 |
| 2022 | CHEX: Multiversion Replay with Ordered CheckpointsabstractIn scientific computing and data science disciplines, it is often necessary to share application workflows and repeat results. Current tools containerize application workflows, and share the resulting container for repeating results. These tools, due to containerization, do improve sharing of results. However, they do not improve the efficiency of replay. In this paper, we present the multiversion replay problem, which arises when multiple versions of an application are containerized, and each version must be replayed to repeat results. To avoid executing each version separately, we develop CHEX , which checkpoints program state and determines when it is permissible to reuse program state across versions. It does so using system call-based execution lineage. Our capability to identify common computations across versions enables us to consider optimizing replay using an in-memory cache, based on a checkpoint-restore-switch system. We show the multiversion replay problem is NP-hard, and propose efficient heuristics for it. CHEX reduces overall replay time by sharing common computations but avoids storing a large number of checkpoints. We demonstrate that CHEX maintains lightweight package sharing, and improves the total time of multiversion replay by 50% on average. Naga Nithin Manne, Shilvi Satpati, Tanu Malik, Amitabha Bagchi, Ashish Gehani, Amitabh Chaudhary |
Proc. VLDB Endow. | 5 |
| 2022 | Trimmer: An Automated System for Configuration-Based Software DebloatingabstractSoftware bloat has negative implications for security, reliability, and performance. To counter bloat, we proposeTrimmer, a static analysis-based system for pruning unused functionality.Trimmerremoves code that is unused with respect to user-provided command-line arguments and application-specific configuration files.Trimmeruses concrete memory tracking and a custom inter-procedural constant propagation analysis that facilitates dead code elimination. Our system supports both context-sensitive and context-insensitive constant propagation. We show that context-sensitive constant propagation is important for effective software pruning in most applications. We introducesparse constant propagationthat performs constant propagation only for configuration-hosting variables and show that it performs better (higher code size reductions) compared to constant propagation for all program variables. Overall, our results show thatTrimmerreduces binary sizes for real-world programs with reasonable analysis times. Across 20 evaluated programs, we observe a mean binary size reduction of 22.7 percent and a maximum reduction of 62.7 percent. For 5 programs, we observe performance speedups ranging from 5 to 53 percent. Moreover, we show that winnowing software applications can reduce the program attack surface by removing code that contains exploitable vulnerabilities. We find that debloating usingTrimmerremoves CVEs in 4 applications. Aatira Anum Ahmad, Abdul Rafae Noor, Hashim Sharif, Usama Hameed, Shoaib Asif, Mubashir Anwar, Ashish Gehani, Fareed Zaffar, Junaid Haroon Siddiqui |
IEEE Trans. Software Eng. | 7 |
| 2021 | ALchemist: Fusing Application and Audit Logs for Precise Attack Provenance without Instrumentation
Shiqing Ma, Zhuo Zhang 0002, Guanhong Tao 0001, Xiangyu Zhang 0001, Dongyan Xu, Vincent Urias, Han Wei Lin, Gabriela F. Ciocarlie, Vinod Yegneswaran, Ashish Gehani |
NDSS | 11 |
| 2021 | CLARION: Sound and Clear Provenance Tracking for Microservice Deployments
Xutong Chen, Hassaan Irshad, Yan Chen 0004, Ashish Gehani, Vinod Yegneswaran |
USENIX Security Symposium | 4 |
| 2021 | TRACE: Enterprise-Wide Provenance Tracking for Real-Time APT DetectionabstractWe present TRACE, a comprehensive provenance tracking system for scalable, real-time, enterprise-wide APT detection. TRACE uses static analysis to identify program unit structures and inter-unit dependences, such that the provenance of an output event includes the input events within the same unit. Provenance collected from individual hosts are integrated to facilitate construction of a distributed enterprise-wide causal graph. We describe the evolution of TRACE over a four-year period, during which our improvements to the system focused on performance, scalability, and fidelity. In this time span, the system call coverage increased (from 47 to 66) while the time and space overhead reduced by over one and two orders of magnitude, respectively. We also provide results from five adversarial engagements where an independent team of system evaluators conducted APT attacks and assessed system performance. The input from our system was used by three other teams to implement real-time APT detection logic. Retrospective analysis revealed that TRACE provided sufficient evidence to detect over 80% of the attack stages across all evaluations. By the last engagement, temporal and spatial overhead had been reduced significantly to 18% and 10%, respectively. Hassaan Irshad, Gabriela F. Ciocarlie, Ashish Gehani, Vinod Yegneswaran, Kyu Hyung Lee, Jignesh M. Patel, Somesh Jha, Yonghwi Kwon 0001, Dongyan Xu, Xiangyu Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2019 | Longitudinal Analysis of Misuse of Bitcoin
Karim M. El Defrawy, Ashish Gehani, Alexandre Matton |
ACNS | 2 |
| 2019 | ProvMark: A Provenance Expressiveness Benchmarking SystemabstractSystem level provenance is of widespread interest for applications such as security enforcement and information protection. However, testing the correctness or completeness of provenance capture tools is challenging and currently done manually. In some cases there is not even a clear consensus about what behavior is correct. We present an automated tool, ProvMark, that uses an existing provenance system as a black box and reliably identifies the provenance graph structure recorded for a given activity, by a reduction to subgraph isomorphism problems handled by an external solver. ProvMark is a beginning step in the much needed area of testing and comparing the expressiveness of provenance systems. We demonstrate ProvMark's usefuless in comparing three capture systems with different architectures and distinct design philosophies. Sheung Chi Chan, James Cheney, Pramod Bhatotia, Thomas Pasquier, Ashish Gehani, Hassaan Irshad, Lucian Carata, Margo I. Seltzer |
Middleware | 5 |
| 2018 | Wholly!: A Build System For The Modern Software Stack
Loic Gelle, Hassen Saïdi, Ashish Gehani |
FMICS | 3 |
| 2018 | TRIMMER: application specialization for code debloatingabstractWith the proliferation of new hardware architectures and ever-evolving user requirements, the software stack is becoming increasingly bloated. In practice, only a limited subset of the supported functionality is utilized in a particular usage context, thereby presenting an opportunity to eliminate unused features. In the past, program specialization has been proposed as a mechanism for enabling automatic software debloating. In this work, we show how existing program specialization techniques lack the analyses required for providing code simplification for real-world programs. We present an approach that uses stronger analysis techniques to take advantage of constant configuration data, thereby enabling more effective debloating. We developed Trimmer, an application specialization tool that leverages user-provided configuration data to specialize an application to its deployment context. The specialization process attempts to eliminate the application functionality that is unused in the user-defined context. Our evaluation demonstrates Trimmer can effectively reduce code bloat. For 13 applications spanning various domains, we observe a mean binary size reduction of 21% and a maximum reduction of 75%. We also show specialization reduces the surface for code-reuse attacks by reducing the number of exploitable gadgets. For the evaluated programs, we observe a 20% mean reduction in the total gadget count and a maximum reduction of 87%. Hashim Sharif, Muhammad Abubakar, Ashish Gehani, Fareed Zaffar |
ASE | 3 |
| 2018 | MCI : Modeling-based Causality Inference in Audit Logging for Attack Investigation
Yonghwi Kwon 0001, Fei Wang 0001, Weihang Wang 0001, Kyu Hyung Lee, Wen-Chuan Lee, Shiqing Ma, Xiangyu Zhang 0001, Dongyan Xu, Somesh Jha, Gabriela F. Ciocarlie, Ashish Gehani, Vinod Yegneswaran |
NDSS | 11 |
| 2018 | Kernel-Supported Cost-Effective Audit Logging for Causality Tracking
Shiqing Ma, Juan Zhai, Yonghwi Kwon 0001, Kyu Hyung Lee, Xiangyu Zhang 0001, Gabriela F. Ciocarlie, Ashish Gehani, Vinod Yegneswaran, Dongyan Xu, Somesh Jha |
USENIX ATC | 7 |
| 2017 | Low-Leakage Secure Search for Boolean Expressions
Fernando Krell, Gabriela F. Ciocarlie, Ashish Gehani, Mariana Raykova 0001 |
CT-RSA | 3 |
| 2017 | Automated Categorization of Onion Sites for Analyzing the Darkweb EcosystemabstractOnion sites on the darkweb operate using the Tor Hidden Service (HS) protocol to shield their locations on the Internet, which (among other features) enables these sites to host malicious and illegal content while being resistant to legal action and seizure. Identifying and monitoring such illicit sites in the darkweb is of high relevance to the Computer Security and Law Enforcement communities. We have developed an automated infrastructure that crawls and indexes content from onion sites into a large-scale data repository, called LIGHTS, with over 100M pages. In this paper we describe Automated Tool for Onion Labeling (ATOL), a novel scalable analysis service developed to conduct a thematic assessment of the content of onion sites in the LIGHTS repository. ATOL has three core components -- (a) a novel keyword discovery mechanism (ATOLKeyword) which extends analyst-provided keywords for different categories by suggesting new descriptive and discriminative keywords that are relevant for the categories; (b) a classification framework (ATOLClassify) that uses the discovered keywords to map onion site content to a set of categories when sufficient labeled data is available; (c) a clustering framework (ATOLCluster) that can leverage information from multiple external heterogeneous knowledge sources, ranging from domain expertise to Bitcoin transaction data, to categorize onion content in the absence of sufficient supervised data. The paper presents empirical results of ATOL on onion datasets derived from the LIGHTS repository, and additionally benchmarks ATOL's algorithms on the publicly available 20 Newsgroups dataset to demonstrate the reproducibility of its results. On the LIGHTS dataset, ATOLClassify gives a 12% performance gain over an analyst-provided baseline, while ATOLCluster gives a 7% improvement over state-of-the-art semi-supervised clustering algorithms. We also discuss how ATOL has been deployed and externally evaluated, as part of the LIGHTS system. Shalini Ghosh, Ariyam Das, Phillip A. Porras, Vinod Yegneswaran, Ashish Gehani |
KDD | 5 |
| 2016 | To route or to secure: Tradeoffs in ICNs over MANETsabstractInformation-Centric Networks (ICNs) operating over Mobile Ad hoc Networks (MANETs) are challenged by the node churn, evolving topologies, and limited resources of the underlying network. The complex interplay of publishers, subscribers, and brokers brings with it a corresponding set of security concerns, where precisely-defined trust boundaries are needed to guarantee the confidentiality and integrity of all data objects in the ecosystem. Building a practical framework that can service users efficiently requires understanding the motivations and actions of the participants. We explore several tradeoffs between efficiency and the security of data objects in such environments, using ICEMAN - a real-wold implementation of an ICN that operates on MANETs. Since our findings are based on an actual system, they have significant implications for building efficient ICNs that have security designed in at the outset (rather than added later when options may be limited). We empirically establish that there is a strong interplay between the need to have more specific information for efficient routing and the need to ensure trust and confidentiality in such a decentralized system. Hasanat Kazmi, Hasnain Lakhani, Ashish Gehani, Rashid Tahir, Fareed Zaffar |
NCA | 3 |
| 2015 | Decentralized Authorization and Privacy-Enhanced Routing for Information-Centric NetworksabstractAs information-centric networks are deployed in increasingly diverse settings, there is a growing need to protect the privacy of participants. We describe the design, implementation, and evaluation of a security framework that achieves this. It ensures the integrity and confidentiality of published content, the associated descriptive metadata, and the interests of subscribers. Mariana Raykova 0001, Hasnain Lakhani, Hasanat Kazmi, Ashish Gehani |
ACSAC | 4 |
| 2015 | Using Provenance Patterns to Vet Sensitive Behaviors in Android Apps
Chao Yang 0022, Guangliang Yang 0001, Ashish Gehani, Vinod Yegneswaran, Dawood Tariq, Guofei Gu |
SecureComm | 3 |
| 2014 | PIDGIN: privacy-preserving interest and content sharing in opportunistic networksabstractOpportunistic networks have recently received considerable attention from both industry and researchers. These networks can be used for many applications without the need for a dedicated IT infrastructure. In the context of opportunistic networks, content sharing in particular has attracted significant attention. To support content sharing, opportunistic networks often implement a publish-subscribe system in which users may publish their own content and indicate interest in other content through subscriptions. Using a smartphone, any user can act as a broker by opportunistically forwarding both published content and interests within the network. Unfortunately, opportunistic networks are faced with serious privacy and security issues. Untrusted brokers can not only compromise the privacy of subscribers by learning their interests but also can gain unauthorised access to the disseminated content. This paper addresses the research challenges inherent to the exchange of content and interests without: (i) compromising the privacy of subscribers, and (ii) providing unauthorised access to untrusted brokers. Specifically, this paper presents an interest and content sharing solution that addresses these security challenges and preserves privacy in opportunistic networks. We demonstrate the feasibility and efficiency of the solution by implementing a prototype and analysing its performance on smart phones. Muhammad Rizwan Asghar, Ashish Gehani, Bruno Crispo, Giovanni Russello |
AsiaCCS | 2 |
| 2013 | Analytical models for risk-based intrusion response
Bugra Çaskurlu, Ashish Gehani, Cemal Çagatay Bilgin, K. Subramani 0001 |
Comput. Networks | 2 |
| 2012 | SPADE: Support for Provenance Auditing in Distributed Environments
Ashish Gehani, Dawood Tariq |
Middleware | 1 |
| 2011 | Ensuring Security and Availability through Model-Based Cross-Layer Adaptation
Minyoung Kim 0002, Mark-Oliver Stehr, Ashish Gehani, Carolyn L. Talcott |
UIC | 3 |
| 2010 | Tracking and Sketching Distributed Data ProvenanceabstractCurrent provenance collection systems typically gather metadata on remote hosts and submit it to a central server. In contrast, several data-intensive scientific applications require a decentralized architecture in which each host maintains an authoritative local repository of the provenance metadata gathered on that host. The latter approach allows the system to handle the large amounts of metadata generated when auditing occurs at fine granularity, and allows users to retain control over their provenance records. The decentralized architecture, however, increases the complexity of auditing, tracking, and querying distributed provenance. We describe a system for capturing data provenance in distributed applications, and the use of provenance sketches to optimize subsequent data provenance queries. Experiments with data gathered from distributed workflow applications demonstrate the feasibility of a decentralized provenance management system and improvements in the efficiency of provenance queries. Tanu Malik, Ligia Nistor, Ashish Gehani |
eScience | 3 |
| 2010 | Mendel: efficiently verifying the lineage of data modified in multiple trust domainsabstractData is routinely created, disseminated, and processed in distributed systems that span multiple administrative domains. To maintain accountability while the data is transformed by multiple parties, a consumer must be able to check the lineage of the data and deem it trustworthy. If integrity is not ensured, the consequences can be significant, particularly when the data cannot easily be reproduced. Verifying the provenance of a piece of data generated using inputs from multiple administrative domains is likely to require the use of numerous public keys that originate at external institutions. Current methods for verifying the integrity of such data from other users will not scale for provenance metadata since scores of verifications may be needed to validate a single file's lineage graph. We describe Mendel, a protocol with a three-pronged strategy that combines eager signature verification, lazy trust establishment, and cryptographic ordering witnesses to yield fast lineage verification in distributed multi-domain environments. Further, we show how decisional lineage queries, that is whether one file is the ancestor of the other, can be answered with high probability in constant time. Ashish Gehani, Minyoung Kim 0002 |
HPDC | 1 |
| 2010 | Efficient querying of distributed provenance storesabstractCurrent projects that automate the collection of provenance information use a centralized architecture for managing the resulting metadata - that is, provenance is gathered at remote hosts and submitted to a central provenance management service. In contrast, we are developing a completely decentralized system with each computer maintaining the authoritative repository of the provenance gathered on it. Our model has several advantages, such as scaling to large amounts of metadata generation, providing low-latency access to provenance metadata about local data, avoiding the need for synchronization with a central service after operating while disconnected from the network, and letting users retain control over their data provenance records. We describe the SPADE project's support for tracking data provenance in distributed environments, including how queries can be optimized with provenance sketches, pre-caching, and caching. Ashish Gehani, Minyoung Kim 0002, Tanu Malik |
HPDC | 1 |
| 2009 | System Support for Forensic Inference
Ashish Gehani, Florent Kirchner, Natarajan Shankar |
IFIP Int. Conf. Digital Forensics | 1 |
| 2008 | Parameterized access control: from design to prototypeabstractPeer-to-peer overlays provide a substrate well suited to building distributed storage systems. Applications that use the infrastructure need the ability to control access to their data. However, traditional authorization services were not designed to operate in the face of network partitions, malicious nodes, and on an Internet-wide scale. Ashish Gehani, Surendar Chandra |
SecureComm | 1 |
| 2007 | Bonsai: Balanced Lineage AuthenticationabstractThe provenance of a piece of data is of utility to a wide range of applications. Its availability can be drastically increased by automatically collecting lineage information during filesystem operations. However, when data is processed by multiple users in independent administrative domains, the resulting filesystem metadata can be trusted only if it has been cryptographically certified. This has three ramifications: it slows down filesystem operations, it requires more storage for metadata, and verification depends on attestations from remote nodes. We show that current schemes do not scale in a distributed environment. In particular, as data is processed, the latency of filesystem operations will degrade exponentially. Further, the amount of storage needed for the lineage metadata will grow at a similar rate. Next, we examine a completely decentralized scheme that has fast filesystem operations with minimal storage overhead. We demonstrate that its verification operation will fail with an exponentially increasing likelihood as more nodes are unreachable (because of being powered off or disconnected from the network). Finally, we present a new scheme, Bonsai, where the verification failure is significantly reduced by tolerating a small increase in filesystem latency and storage overhead for certification compared to file systems without lineage certification. Ashish Gehani, Ulf Lindqvist |
ACSAC | 1 |
| 2007 | Automated Storage Reclamation Using Temporal Importance AnnotationsabstractThis work focuses on scenarios that require the storage of large amounts of data. Such systems require the ability to either continuously increase the storage space or reclaim space by deleting contents. Traditionally, storage systems relegated object reclamation to applications. In this work, content creators explicitly annotate the object using a temporal importance function. The storage system uses this information to evict less important objects. The challenge is to design importance functions that are simple and expressive. We describe a two step temporal importance function. We introduce the notion of storage importance density to quantify the importance levels for which the storage is full. Using extensive simulations and observations of a university wide lecture video capture and storage application, we show that our abstraction allows the users to express the amount of persistence for each individual object. Surendar Chandra, Ashish Gehani, Xuwen Yu |
ICDCS | 2 |
| 2007 | Super-Resolution Video Analysis for Forensic Investigations
Ashish Gehani, John H. Reif |
IFIP Int. Conf. Digital Forensics | 1 |
| 2007 | VEIL: A System for Certifying Video ProvenanceabstractTraditionally, a consumer decided how much to trust a piece of data based on its source. As digital video cameras and editors become ubiquitous, an arbitrary video object is increasingly likely to be produced using a range of operations that combine clips from a multitude of sources. A consumer can determine the assurance level of the data by knowing its lineage. We describe a system to embed the provenance of the video into the data itself. As long as the video contains a predefined threshold of data (from the spatial and temporal domains), the entire lineage can be ascertained. We embed the metadata using subpixel linear interpolation between similar blocks in proximal frames. It can then be extracted in real time using a novel method for computing the embedded interpolation. We implemented the process in C and report the performance overhead it introduces for playing video files. We also characterize the tradeoff between the auxiliary channel's capacity (which limits the amount of provenance metadata that can be embedded) and the extent to which the video can be edited (in the spatial or temporal domains) while retaining complete lineage. Ashish Gehani, Ulf Lindqvist |
ISM | 1 |
| 2006 | PAST: Probabilistic Authentication of Sensor TimestampsabstractSensor networks are deployed to monitor the physical environment in public and vulnerable locations. It is not economically viable to house sensors in tamper-resilient enclosures as they are deployed in large numbers. As a result, an adversary can subvert the integrity of the data being produced by gaining physical access to a sensor and altering its code. If the sensor output is timestamped, then tainted data can be distinguished once the time of attack is determined. To prevent the adversary from generating fraudulent timestamps, the data must be authenticated using a forward-secure protocol. Previous work requires the computation of n hashes to verify the (n + 1)threading. This paper describes PAST, a protocol that allows timestamps to be authenticated with high probability using a small constant number of readings. In particular, PAST is parameterized so that the metadata overhead (and associated power consumption) can be reduced at the cost of lower confidence in the authentication guarantee. Our protocol allows arbitrary levels of assurance for the integrity of timestamps (with logarithmically increasing storage costs) while tolerating any predefined fraction of compromised base stations. Unlike prior schemes, PAST does not depend on synchronized clocks Ashish Gehani, Surendar Chandra |
ACSAC | 1 |
| 2006 | Augmenting storage with an intrusion response primitive to ensure the security of critical dataabstractHosts connected to the Internet continue to suffer attacks with high frequency. The use of an intrusion detector allows potential threats to be flagged. When an alarm is raised, preventive action can be taken. A primary goal of such action is to assure the security of the data stored in the system. If this operation is effected manually, the delay between the alarm and the response may be enough for an intruder to cause significant damage.The alternative proposed in this paper is to provide a response primitive for intrusion detectors to utilize in automating the response. We describe RICE, a modification to the Java file subsystem that provides such functionality for data that is deemed to be threatened by an attack. If it is activated when an intrusion appears likely to succeed, it guarantees the confidentiality, integrity and availability of the protected data even after a system is compromised.In particular, RICE allows cryptographic encapsulation of data to be reduced to simple key deletion so that it can be effected rapidly. Further, it uses digitally signed hashes of file deltas to allow untained data to be distinguished from the rest. Finally, file deltas are replicated at a remote node to ensure that changes made by an attacker can be undone using the remote replicas. Ashish Gehani, Surendar Chandra, Gershon Kedem |
AsiaCCS | 1 |
| 2005 | Paranoid: A Global Secure File Access Control SystemabstractThe Paranoid file system is an encrypted, secure, global file system with user managed access control. The system provides efficient peer-to-peer application transparent file sharing. This paper presents the design, implementation and evaluation of the Paranoid file system and its access-control architecture. The system lets users grant safe, selective, UNIX-like, file access to peer groups across administrative boundaries. Files are kept encrypted and access control translates into key management. The system uses a novel transformation key scheme to effect access revocation. The file system works seamlessly with existing applications through the use of interposition agents. The interposition agents provide a layer of indirection making it possible to implement transparent remote file access and data encryption/decryption without any kernel modifications. System performance evaluations show that encryption and remote file-access overheads are small, demonstrating that the Paranoid system is practical Fareed Zaffar, Gershon Kedem, Ashish Gehani |
ACSAC | 3 |
| 2004 | RheoStat: Real-Time Risk Management
Ashish Gehani, Gershon Kedem |
RAID | 1 |