VLDB 2026 Research / reviewers in the wild / expert
Erez Zadok
dblp:71/1341
· DBLP profile ↗
96ranked-venue papers
8as first author
17since 2021 · last 2025
0000-0001-5248-9184ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 77 · 8 first-author · 14 since 2021Databases, data management, data science and information retrieval · 17 · 2 since 2021Software engineering, systems software and programming languages · 11Security and privacy · 4Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhanced File System Testing through Input and Output CoverageabstractEffective file system testing relies on coverage to detect bugs and enhance reliability. We analyzed real file system bugs and found a weak correlation between code coverage, the most commonly used metric, and test effectiveness; many bugs were in covered code but remained undetected. Our study also showed that covering diverse file system inputs and outputs---system call arguments and return values---can be key to detecting the majority of observed bugs. Geoffrey H. Kuenning, Kamal Parvez, Scott A. Smolka, Erez Zadok |
SYSTOR | 5 |
| 2025 | The Past, Present, and Future of Storage Technologies (Part 1 of 2)
Geoffrey H. Kuenning, Youjip Won, Ming Zhao 0002, Erez Zadok |
ACM Trans. Storage | 4 |
| 2025 | The Past, Present, and Future of Storage Technologies (part 2 of 2)abstractThe Past, Present, and Future of Storage Technologies (part 2 of 2)Any good research project begins with a "literature review"-a process of looking for papers relevant to a topic of interest, reviewing them to identify those more useful while discarding the rest, then looking for more papers, and repeating this process over and over until you feel you have reached some saturation point.That is the point when you're coming up against the same papers you have seen already (which we like to call the "transitive closure" point).This review stage is fairly time-consuming but also critical; in fact, many research papers get rejected for neglecting to cite important related work.Similarly, practitioners may run into roadblocks that they could have avoided if they had been aware of all the existing literature.So, are you a new graduate student or an employee at a company who is interested in innovating in a given storage technology?Do you wish you could find a single publication that would summarize (almost) everything there is to know about a specific technology?If so, this Special Issue (both parts) is hopefully for you because, unlike conference papers, our journal articles have no page limits and thus allow authors to discuss any technology with as much detail as needed and include a comprehensive bibliography for those interested in more.Incidentally, we found out that survey papers tend to get well cited and received.We believe this is because they become a "go to" source on a topic, thus saving the readers from having to read and cite many other papers.Authors of survey papers may see their articles cited over and over.In 2023, TOS's Editor-in-Chief ( EiC ) reached out to a few senior people to brainstorm ideas for special issues.Because putting together a special issue is a huge task, he recruited several Associate Editors (AEs) to help.We settled on an idea particularly suitable for journals: survey papers.And we decided to focus on the bottom of the storage stack: storage technologies and media.We also debated whether we should focus on futuristic technologies, current, or past ones.In the end, we opted to include everything: all storage technologies, regardless of their age or maturity, can teach us something useful.Normally, TOS authors submit their full manuscript for review.However, because this special issue's survey nature would likely mean longer papers, we wanted to provide better direction to authors.So, we posted a CFP asking the prospective authors to submit a short one-page abstract.We provided guidance to prospective authors as to what makes a good survey paper.Specifically, authors would have to survey many related papers in their chosen area, so as to make their survey the most comprehensive paper on the topic to date.Secondly, it was not enough to just summarize past papers; authors also needed to provide insight into why and how a given technology evolved over the years, and where it might go in the future. Geoffrey H. Kuenning, Youjip Won, Ming Zhao 0002, Erez Zadok |
ACM Trans. Storage | 4 |
| 2025 | Into the Void: Mapping the Unseen Gaps in High Dimensional DataabstractWe present a comprehensive pipeline, integrated with a visual analytics system called GapMiner, capable of exploring and exploiting untapped opportunities within the empty regions of high-dimensional datasets. Our approach utilizes a novel Empty-Space Search Algorithm (ESA) to identify the center points of these uncharted voids, which represent reservoirs for potentially valuable new configurations. Initially, this process is guided by user interactions through GapMiner, which visualizes Empty-Space Configurations (ESCs) within the context of the dataset and allows domain experts to explore and refine ESCs for subsequent validation in domain experiments or simulations. These activities iteratively enhance the dataset and contribute to training a connected deep neural network (DNN). As training progresses, the DNN gradually assumes the role of identifying and validating high-potential ESCs, reducing the need for direct user involvement. Once the DNN achieves sufficient accuracy, it autonomously guides the exploration of optimal configurations by predicting performance and refining configurations through a combination of gradient ascent and improved empty-space searches. Domain experts were actively involved throughout the system's development. Our findings demonstrate that this methodology consistently generates superior novel configurations compared to conventional randomization-based approaches. We illustrate its effectiveness in multiple case studies with diverse objectives. Tyler Estro, Geoffrey H. Kuenning, Erez Zadok, Klaus Mueller 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Metis: File System Model Checking via Versatile Input and State Exploration
Manish Adkar, Gerard J. Holzmann, Geoffrey H. Kuenning, Scott A. Smolka, Erez Zadok |
FAST | 8 |
| 2024 | Secure Archival is Hard... Really HardabstractArchival systems are often tasked with storing highly valuable data that may be targeted by malicious actors. When the lifetime of the secret data is on the order of decades to centuries, the threat of improved cryptanalysis casts doubt on the long-term security of cryptographic techniques, which rely on hardness assumptions that are hard to prove over archival time scales. This threat makes the design of secure archival systems exceptionally difficult. Some archival systems turn a blind eye to this issue, hoping that current cryptographic techniques will not be broken; others often use techniques---such as secret sharing---that are impractical at scale. This position paper sheds light on the core challenges behind building practically viable secure long-term archives; we identify promising research avenues towards this goal. Maliha Tabassum, Soumya Chowdary Daruru, Gaurav Kulhare, Arvin Wang, Ethan L. Miller, Erez Zadok |
HotStorage | 7 |
| 2024 | Accelerating multi-tier storage cache simulations using knee detection
Tyler Estro, Mário Antunes 0001, Pranav Bhandari, Anshul Gandhi, Geoffrey H. Kuenning, Carl A. Waldspurger, Avani Wildani, Erez Zadok |
Perform. Evaluation | 9 |
| 2023 | Input and Output Coverage Needed in File System TestingabstractFile systems need testing to discover bugs and to help ensure reliability. Many file system testing tools are evaluated based on their code coverage. We analyzed recently reported bugs in Ext4 and BtrFS and found a weak correlation between code coverage and test effectiveness: many bugs are missed because they depend on specific inputs, even though the code was covered by a test suite. Our position is that coverage of system call inputs and outputs is critically important for testing file systems. We thus suggest input and output coverage as criteria for file system testing, and show how they can improve the effectiveness of testing. We built a prototype called IOCov to evaluate the input and output coverage of file system testing tools. IOCov identified many untested cases (specific inputs and outputs or ranges thereof) for both CrashMonkey and xfstests. Additionally, we discuss a method and associated metrics to identify over- and under-testing using IOCov. Gautam Ahuja, Geoffrey H. Kuenning, Scott A. Smolka, Erez Zadok |
HotStorage | 5 |
| 2023 | Guiding Simulations of Multi-Tier Storage Caches Using Knee DetectionabstractSimulating storage cache hierarchies enables efficient exploration of their configuration space, including diverse topologies, parameters and policies, and devices with varied performance characteristics, while avoiding expensive physical experiments. Miss Ratio Curves (MRCs) efficiently characterize the performance of a cache over a range of cache sizes. These useful tools reveal “key points” for cache simulation, such as knees in the curve that immediately follow sharp cliffs. Unfortunately, there are no automated techniques for efficiently finding key points in MRCs, and the cross-application of existing knee-detection algorithms yields inaccurate results. We present a multi-stage framework that identifies key points in any MRC, for both stack-based (e.g., LRU) and more sophis-ticated eviction algorithms (e.g., ARC). Our approach quickly locates candidates using efficient hash-based sampling, curve simplification, knee detection, and novel post-processing filters. We introduce Z-Method, a new multi-knee detection algorithm that employs statistical outlier detection to choose promising points robustly and efficiently. We evaluate our framework against seven other knee-detection algorithms, using both ARC and LRU MRCs from 106 diverse real-world workloads, and apply it to identify key points in multi-tier MRCs. Compared to naive approaches, our framework reduces the total number of points needed to accurately identify the best two-tier cache hierarchies by an average factor of approximately$5.5\times$for ARC and$7.7\times$for LRU. Tyler Estro, Mário Antunes 0001, Pranav Bhandari, Anshul Gandhi, Geoffrey H. Kuenning, Carl A. Waldspurger, Avani Wildani, Erez Zadok |
MASCOTS | 9 |
| 2023 | F3: Serving Files Efficiently in Serverless ComputingabstractServerless platforms offer on-demand computation and represent a significant shift from previous platforms that typically required resources to be pre-allocated (e.g., virtual machines). As serverless platforms have evolved, they have become suitable for a much wider range of applications than their original use cases. However, storage access remains a pain point that holds serverless back from becoming a completely generic computation platform. Alex Merenstein, Vasily Tarasov, Ali Anwar 0001, Scott Guthridge, Erez Zadok |
SYSTOR | 5 |
| 2023 | Improving Storage Systems Using Machine LearningabstractOperating systems include many heuristic algorithms designed to improve overall storage performance and throughput. Because such heuristics cannot work well for all conditions and workloads, system designers resorted to exposing numerous tunable parameters to users—thus burdening users with continually optimizing their own storage systems and applications. Storage systems are usually responsible for most latency in I/O-heavy applications, so even a small latency improvement can be significant. Machine learning (ML) techniques promise to learn patterns, generalize from them, and enable optimal solutions that adapt to changing workloads. We propose that ML solutions become a first-class component in OSs and replace manual heuristics to optimize storage systems dynamically. In this article, we describe our proposed ML architecture, called KML. We developed a prototype KML architecture and applied it to two case studies: optimizing readahead and NFS read-size values. Our experiments show that KML consumes less than 4 KB of dynamic kernel memory, has a CPU overhead smaller than 0.2%, and yet can learn patterns and improve I/O throughput by as much as 2.3× and 15× for two case studies—even for complex, never-seen-before, concurrently running mixed workloads on different storage devices. Ibrahim Umit Akgun, Ali Selman Aydin, Andrew Burford, Michael McNeill, Michael Arkhangelskiy, Erez Zadok |
ACM Trans. Storage | 6 |
| 2023 | PC-Expo: A Metrics-Based Interactive Axes Reordering Method for Parallel Coordinate DisplaysabstractParallel coordinate plots (PCPs) have been widely used for high-dimensional (HD) data storytelling because they allow for presenting a large number of dimensions without distortions. The axes ordering in PCP presents a particular story from the data based on the user perception of PCP polylines. Existing works focus on directly optimizing for PCP axes ordering based on some common analysis tasks like clustering, neighborhood, and correlation. However, direct optimization for PCP axes based on these common properties is restrictive because it does not account for multiple properties occurring between the axes, and for local properties that occur in small regions in the data. Also, many of these techniques do not support the human-in-the-loop (HIL) paradigm, which is crucial (i) for explainability and (ii) in cases where no single reordering scheme fits the users' goals. To alleviate these problems, we present PC-Expo, a real-time visual analytics framework for all-in-one PCP line pattern detection and axes reordering. We studied the connection of line patterns in PCPs with different data analysis tasks and datasets. PC-Expo expands prior work on PCP axes reordering by developing real-time, local detection schemes for the 12 most common analysis tasks (properties). Users can choose the story they want to present with PCPs by optimizing directly over their choice of properties. These properties can be ranked, or combined using individual weights, creating a custom optimization scheme for axes reordering. Users can control the granularity at which they want to work with their detection scheme in the data, allowing exploration of local regions. PC-Expo also supports HIL axes reordering via local-property visualization, which shows the regions of granular activity for every axis pair. Local-property visualization is helpful for PCP axes reordering based on multiple properties, when no single reordering scheme fits the user goals. A comprehensive evaluation was done with real users and diverse datasets confirm the efficacy of PC-Expo in data storytelling with PCPs. Anjul Kumar Tyagi, Tyler Estro, Geoffrey H. Kuenning, Erez Zadok, Klaus Mueller 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | SpecNFS: A Challenge Dataset Towards Extracting Formal Models from Natural Language SpecificationsabstractCan NLP assist in building formal models for verifying complex systems? We study this challenge in the context of parsing Network File System (NFS) specifications. We define a semantic-dependency problem over SpecIR, a representation language we introduce to model sentences appearing in NFS specification documents (RFCs) as IF-THEN statements, and present an annotated dataset of 1,198 sentences. We develop and evaluate semantic-dependency parsing systems for this problem. Evaluations show that even when using a state-of-the-art language model, there is significant room for improvement, with the best models achieving an F1 score of only 60.5 and 33.3 in the named-entity-recognition and dependency-link-prediction sub-tasks, respectively. We also release additional unlabeled data and other domain-related texts. Experiments show that these additional resources increase the F1 measure when used for simple domain-adaption and transfer-learning-based approaches, suggesting fruitful directions for further research Sayontan Ghosh, Amanpreet Singh, Alex Merenstein, Scott A. Smolka, Erez Zadok, Niranjan Balasubramanian |
LREC | 6 |
| 2021 | CNSBench: A Cloud Native Storage Benchmark
Alex Merenstein, Vasily Tarasov, Ali Anwar 0001, Deepavali Bhagwat, Julie Lee, Lukas Rupprecht, Dimitrios Skourtis, Erez Zadok |
FAST | 9 |
| 2021 | A Machine Learning Framework to Improve Storage System PerformanceabstractStorage systems and their OS components are designed to accommodate a wide variety of applications and dynamic workloads. Storage components inside the OS contain various heuristic algorithms to provide high performance and adaptability for different workloads. These heuristics may be tunable via parameters, and some system calls allow users to optimize their system performance. These parameters are often predetermined based on experiments with limited applications and hardware. Thus, storage systems often run with these predetermined and possibly suboptimal values. Tuning these parameters manually is impractical: one needs an adaptive, intelligent system to handle dynamic and complex workloads. Machine learning (ML) techniques are capable of recognizing patterns, abstracting them, and making predictions on new data. ML can be a key component to optimize and adapt storage systems. In this position paper, we propose KML, an ML framework for storage systems. We implemented a prototype and demonstrated its capabilities on the well-known problem of tuning optimal readahead values. Our results show that KML has a small memory footprint, introduces negligible overhead, and yet enhances throughput by as much as 2.3x. Ibrahim Umit Akgun, Ali Selman Aydin, Aadil Shaikh, Lukas Velikov, Erez Zadok |
HotStorage | 5 |
| 2021 | Model-Checking Support for File System DevelopmentabstractDeveloping and maintaining a file system is time-consuming, typically requiring years of effort. Developers often test compliance with APIs such as POSIX with hand-written regression suites that, alas, examine only a fraction of a file system's state space. Conversely, formal model checking can explore vast state spaces efficiently, increasing confidence in the file system's implementation. Yet model checking is not currently part of file system development. Our position is that file systems should be designed a priori to facilitate model checking. To this end, we introduce MCFS, an architecture for efficient and comprehensive file-system model checking. MCFS relies on two new APIs that save and restore a file system's in-memory and on-disk state. We describe our earlier attempts at model-checking file systems, including unsuccessful or inefficient ones. Those attempts led us to develop VeriFS, which implements the new APIs. We illustrate MCFS's model-checking principles with VeriFS, a FUSE-based file system we were able to quickly develop with MCFS's help. Gomathi Ganesan, Gerard J. Holzmann, Scott A. Smolka, Erez Zadok, Geoffrey H. Kuenning |
HotStorage | 6 |
| 2021 | Introduction to the Special Issue on USENIX ATC 2020abstractNo abstract available. Ada Gavrilovska, Erez Zadok |
ACM Trans. Storage | 2 |
| 2020 | Carver: Finding Important Parameters for Storage System Tuning
Geoffrey H. Kuenning, Erez Zadok |
FAST | 3 |
| 2020 | Desperately Seeking ... Optimal Multi-Tier Cache Configurations
Tyler Estro, Pranav Bhandari, Avani Wildani, Erez Zadok |
HotStorage | 4 |
| 2020 | The Case for Benchmarking Control Operations in Cloud Native Storage
Alex Merenstein, Vasily Tarasov, Ali Anwar 0001, Deepavali Bhagwat, Lukas Rupprecht, Dimitrios Skourtis, Erez Zadok |
HotStorage | 7 |
| 2020 | Re-Animator: Versatile High-Fidelity Storage-System Tracing and ReplayingabstractModern applications use storage systems in complex and often surprising ways. Tracing system calls is a common approach to understanding applications' behavior, allowing offline analysis and enabling replay in other environments. But current system-call tracing tools have drawbacks: (1) they often omit some information---such as raw data buffers---needed for full analysis; (2) they have high overheads; (3) they often use non-portable trace formats; and (4) they may not offer useful and scalable analysis and replay tools. Ibrahim Umit Akgun, Geoffrey H. Kuenning, Erez Zadok |
SYSTOR | 3 |
| 2020 | Supporting Transactions for Bulk NFSv4 CompoundsabstractMore applications nowadays use network and cloud storage; and modern network file system protocols support compounding operations---packing more operations in one request (e.g., NFSv4, SMB). This is known to improve overall throughput and latency by reducing the number of network round trips. It has been reported that by utilizing compounds, NFSv4 performance, especially in high-latency networks, can be improved by orders of magnitude. Alas, with more operations packed into a single message, partial failures become more likely---some server-side operations succeed while others fail to execute. This places a greater challenge on client-side applications to recover from such failures. To solve this and simplify application development, we designed and built TC-NFS, an NFSv4-based network file system with transactional compound execution. We evaluated TC-NFS with different workloads, compounding degrees, and network latencies. Compared to an already existing NFSv4 system that fully utilizes compounds, our end-to-end transactional support adds as little as ~1.1% overhead but as much as ~25× overhead for some intense micro- and macro-workloads. Akshay Aurora, Ming Chen 0013, Erez Zadok |
SYSTOR | 4 |
| 2020 | Analyzing the distribution fit for storage workload and Internet traffic traces
Muhammad Wajahat, Aditya Yele, Tyler Estro, Anshul Gandhi, Erez Zadok |
Perform. Evaluation | 5 |
| 2019 | Graphs Are Not Enough: Using Interactive Visual Analytics in Storage Research
Geoffrey H. Kuenning, Klaus Mueller 0001, Anjul Kumar Tyagi, Erez Zadok |
HotStorage | 5 |
| 2019 | Distribution Fitting and Performance Modeling for Storage TracesabstractUnderstanding I/O workloads and modeling their performance is important for optimizing storage systems. A useful first step towards understanding the characteristics of storage workloads is to analyze their inter-arrival times and service requirements. If these characteristics are found to follow certain probability distributions, then corresponding stochastic models can be employed to efficiently estimate the performance of storage workloads. Such approaches have been explored in other domains using an assortment of distributions, including the Normal, Weibull, and Exponential. However, our analysis and others' past attempts revealed that none of those distributions provided a good fit for storage workloads. We analyzed over 200 traces across 4 different workload families using 20 widely used distributions, including ones seldom used for storage modeling. We found that the Hyper-exponential distribution with just two phases H_2 was superior in modeling the storage traces compared to other distributions under five diverse metrics of accuracy, including metrics that assess the risk of over-fitting. Based on these results, we developed a Markov-chain-based stochastic model that accurately estimates the storage system performance across several workload traces. To highlight the applicability of our model, we conducted what-if analyses to investigate the performance impact of workload variability and garbage collection under various scenarios. Muhammad Wajahat, Aditya Yele, Tyler Estro, Anshul Gandhi, Erez Zadok |
MASCOTS | 5 |
| 2019 | Kurma: secure geo-distributed multi-cloud storage gatewaysabstractCloud storage is highly available, scalable, and cost-efficient. Yet, many cannot store data in cloud due to security concerns and legacy infrastructure such as network-attached storage (NAS). We describe Kurma, a cloud storage gateway system that allows NAS-based programs to seamlessly and securely access cloud storage. To share files among distant clients, Kurma maintains a unified file-system namespace by replicating metadata across geo-distributed gateways. Kurma stores only encrypted data blocks in clouds, keeps file-system and security metadata on-premises, and can verify data integrity and freshness without any trusted third party. Kurma uses multiple clouds to prevent cloud outage and vendor lock-in. Kurma's performance is 52--91% that of a local NFS server while providing geo-replication, confidentiality, integrity, and high availability. Ming Chen 0013, Erez Zadok |
SYSTOR | 2 |
| 2019 | Performance and Resource Utilization of FUSE User-Space File SystemsabstractTraditionally, file systems were implemented as part of operating systems kernels, which provide a limited set of tools and facilities to a programmer. As the complexity of file systems grew, many new file systems began being developed in user space. Low performance is considered the main disadvantage of user-space file systems but the extent of this problem has never been explored systematically. As a result, the topic of user-space file systems remains rather controversial: while some consider user-space file systems a “toy” not to be used in production, others develop full-fledged production file systems in user space. In this article, we analyze the design and implementation of a well-known user-space file system framework, FUSE, for Linux. We characterize its performance and resource utilization for a wide range of workloads. We present FUSE performance and also resource utilization with various mount and configuration options, using 45 different workloads that were generated using Filebench on two different hardware configurations. We instrumented FUSE to extract useful statistics and traces, which helped us analyze its performance bottlenecks and present our analysis results. Our experiments indicate that depending on the workload and hardware used, performance degradation (throughput) caused by FUSE can be completely imperceptible or as high as −83%, even when optimized; and latencies of FUSE file system operations can be increased from none to 4× when compared to Ext4. On the resource utilization side, FUSE can increase relative CPU utilization by up to 31% and underutilize disk bandwidth by as much as −80% compared to Ext4, though for many data-intensive workloads the impact was statistically indistinguishable. Our conclusion is that user-space file systems can indeed be used in production (non-“toy”) settings, but their applicability depends on the expected workloads. Bharath Kumar Reddy Vangoor, Prafful Agarwal, Manu Mathew, Arun Ramachandran, Swaminathan Sivaraman, Vasily Tarasov, Erez Zadok |
ACM Trans. Storage | 7 |
| 2018 | Towards Better Understanding of Black-box Auto-Tuning: A Comparative Analysis for Storage Systems
Vasily Tarasov, Sachin Tiwari, Erez Zadok |
USENIX ATC | 4 |
| 2018 | Cluster and Single-Node Analysis of Long-Term Deduplication PatternsabstractDeduplication has become essential in disk-based backup systems, but there have been few long-term studies of backup workloads. Most past studies either were of a small static snapshot or covered only a short period that was not representative of how a backup system evolves over time. For this article, we first collected 21 months of data from a shared user file system; 33 users and over 4,000 snapshots are covered. We then analyzed the dataset, examining a variety of essential characteristics across two dimensions: single-node deduplication and cluster deduplication. For single-node deduplication analysis, our primary focus was individual-user data. Despite apparently similar roles and behavior among all of our users, we found significant differences in their deduplication ratios. Moreover, the data that some users share with others had a much higher deduplication ratio than average. For cluster deduplication analysis, we implemented seven published data-routing algorithms and created a detailed comparison of their performance with respect to deduplication ratio, load distribution, and communication overhead. We found that per-file routing achieves a higher deduplication ratio than routing by super-chunk (multiple consecutive chunks), but it also leads to high data skew (imbalance of space usage across nodes). We also found that large chunking sizes are better for cluster deduplication, as they significantly reduce data-routing overhead, while their negative impact on deduplication ratios is small and acceptable. We draw interesting conclusions from both single-node and cluster deduplication analysis and make recommendations for future deduplication systems design. Zhen Jason Sun, Geoffrey H. Kuenning, Sonam Mandal, Philip Shilane, Vasily Tarasov, Nong Xiao 0001, Erez Zadok |
ACM Trans. Storage | 7 |
| 2017 | On the Performance Variation in Modern Storage Stacks
Vasily Tarasov, Hari Prasath Raman, Dean Hildebrand, Erez Zadok |
FAST | 5 |
| 2017 | vNFS: Maximizing NFS Performance with Compounds and Vectorized I/O
Ming Chen 0013, Dean Hildebrand, Henry Nelson, Jasmit Saluja, Ashok Sankar Harihara Subramony, Erez Zadok |
FAST | 6 |
| 2017 | To FUSE or Not to FUSE: Performance of User-Space File Systems
Bharath Kumar Reddy Vangoor, Vasily Tarasov, Erez Zadok |
FAST | 3 |
| 2017 | POSIX is Dead! Long Live... errr... What Exactly?
Erez Zadok, Dean Hildebrand, Geoffrey H. Kuenning, Keith A. Smith |
HotStorage | 1 |
| 2017 | vNFS: Maximizing NFS Performance with Compounds and Vectorized I/OabstractModern systems use networks extensively, accessing both services and storage across local and remote networks. Latency is a key performance challenge, and packing multiple small operations into fewer large ones is an effective way to amortize that cost, especially after years of significant improvement in bandwidth but not latency. To this end, the NFSv4 protocol supports a compounding feature to combine multiple operations. Yet compounding has been underused since its conception because the synchronous POSIX file-system API issues only one (small) request at a time. We propose vNFS , an NFSv4.1-compliant client that exposes a vectorized high-level API and leverages NFS compound procedures to maximize performance. We designed and implemented vNFS as a user-space RPC library that supports an assortment of bulk operations on multiple files and directories. We found it easy to modify several UNIX utilities, an HTTP/2 server, and Filebench to use vNFS. We evaluated vNFS under a wide range of workloads and network latency conditions, showing that vNFS improves performance even for low-latency networks. On high-latency networks, vNFS can improve performance by as much as two orders of magnitude. Ming Chen 0013, Geetika Babu Bangera, Dean Hildebrand, Farhaan Jalia, Geoffrey H. Kuenning, Henry Nelson, Erez Zadok |
ACM Trans. Storage | 7 |
| 2016 | Using Hints to Improve Inline Block-layer Deduplication
Sonam Mandal, Geoffrey H. Kuenning, Dongju Ok, Varun Shastry, Philip Shilane, Sun Zhen, Vasily Tarasov, Erez Zadok |
FAST | 8 |
| 2016 | A long-term user-centric analysis of deduplication patternsabstractDeduplication has become essential in disk-based backup systems, but there have been few long-term studies of backup workloads. Most past studies either were of a small static snapshot or covered only a short period that was not representative of how a backup system evolves over time. For this paper, we collected 21 months of data from a shared user file system; 33 users and over 4,000 snapshots are covered. We analyzed the data set for a variety of essential characteristics. However, our primary focus was individual user data. Despite apparently similar roles and behavior in all of our users, we found significant differences in their deduplication ratios. Moreover, the data that some users share with others had a much higher deduplication ratio than average. We analyze this behavior and make recommendations for future deduplication systems design. Geoffrey H. Kuenning, Sonam Mandal, Philip Shilane, Vasily Tarasov, Nong Xiao 0001, Erez Zadok |
MSST | 7 |
| 2016 | SeMiNAS: A Secure Middleware for Wide-Area Network-Attached StorageabstractUtility computing is being gradually realized as exemplified by cloud computing. Outsourcing computing and storage to global-scale cloud providers benefits from high accessibility, flexibility, scalability, and cost-effectiveness. However, users are uneasy outsourcing the storage of sensitive data due to security concerns. We address this problem by presenting SeMiNAS---an efficient middleware system that allows files to be securely outsourced to providers and shared among geo-distributed offices. SeMiNAS achieves end-to-end data integrity and confidentiality with a highly efficient authenticated-encryption scheme. SeMiNAS leverages advanced NFSv4 features, including compound procedures and data-integrity extensions, to minimize extra network round trips caused by security meta-data. SeMiNAS also caches remote files locally to reduce accesses to providers over WANs. We designed, implemented, and evaluated SeMiNAS, which demonstrates a small performance penalty of less than 26% and an occasional performance boost of up to 19% for Filebench workloads. Ming Chen 0013, Erez Zadok, Arun O. Vasudevan, Kelong Wang |
SYSTOR | 2 |
| 2015 | Terra Incognita: On the Practicality of User-Space File Systems
Vasily Tarasov, Kumar Sourav, Sagar Trehan, Erez Zadok |
HotStorage | 5 |
| 2015 | Parametric Optimization of Storage Systems
Erez Zadok, Aashray Arora, Akhilesh Chaganti, Arvind Chaudhary, Sonam Mandal |
HotStorage | 1 |
| 2015 | Newer Is Sometimes Better: An Evaluation of NFSv4.1abstractThe popular Network File System (NFS) protocol is 30 years old. The latest version, NFSv4, is more than ten years old but has only recently gained stability and acceptance. NFSv4 is vastly different from its predecessors: it offers a stateful server, strong security, scalability/WAN features, and callbacks, among other things. Yet NFSv4's efficacy and ability to meet its stated design goals had not been thoroughly studied until now. This paper compares NFSv4.1's performance with NFSv3 using a wide range of micro- and macro-benchmarks on a testbed configured to exercise the core protocol features. We (1) tested NFSv4's unique features, such as delegations and statefulness; (2) evaluated performance comprehensively with different numbers of threads and clients, and different network latencies and TCP/IP features; (3) found, fixed, and reported several problems in Linux's NFSv4.1 implementation, which helped improve performance by up to 11X; and (4) discovered, analyzed, and explained several counter-intuitive results. Depending on the workload, NFSv4.1 was up to 67\% slower than NFSv3 in a low-latency network, but exceeded NFSv3's performance by up to 2.9X in a high-latency environment. Moreover, NFSv4.1 outperformed NFSv3 by up to 172X when delegations were used. Ming Chen 0013, Dean Hildebrand, Geoffrey H. Kuenning, Soujanya Shankaranarayana, Erez Zadok |
SIGMETRICS | 6 |
| 2015 | On the Trade-Offs among Performance, Energy, and Endurance in a Versatile Hybrid DriveabstractThere are trade-offs among performance, energy, and device endurance for storage systems. Designs optimized for one dimension or workload often suffer in another. Therefore, it is important to study the trade-offs to enable adaptation to workloads and dimensions. As Flash SSD has emerged, hybrid drives have been studied more closely. However, hybrids are mainly designed for high throughput, efficient energy consumption, or improving endurance—leaving quantitative study on the trade-offs unexplored. Past endurance studies also lack a concrete model to help study the trade-offs. Last, previous designs are often based on inflexible policies that cannot adapt easily to changing conditions. We designed and developed GreenDM , a versatile hybrid drive that combines Flash-based SSDs with traditional HDDs. The SSD can be used as cache or as primary storage for hot data. We present our endurance model together with GreenDM to study these trade-offs. GreenDM presents a block interface and requires no modifications to existing software. GreenDM offers tunable parameters to enable the system to adapt to many workloads. We have designed, developed, and carefully evaluated GreenDM with a variety of workloads using commodity SSD and HDD drives. We demonstrate the importance of versatility to enable adaptation to various workloads and dimensions. Ming Chen 0013, Amanpreet Mukker, Erez Zadok |
ACM Trans. Storage | 4 |
| 2015 | Introduction to the Special Issue on USENIX FAST 2015abstractNo abstract available. Jiri Schindler, Erez Zadok |
ACM Trans. Storage | 2 |
| 2015 | Visual Correlation Analysis of Numerical and Categorical Data on the Correlation MapabstractCorrelation analysis can reveal the complex relationships that often exist among the variables in multivariate data. However, as the number of variables grows, it can be difficult to gain a good understanding of the correlation landscape and important intricate relationships might be missed. We previously introduced a technique that arranged the variables into a 2D layout, encoding their pairwise correlations. We then used this layout as a network for the interactive ordering of axes in parallel coordinate displays. Our current work expresses the layout as a correlation map and employs it for visual correlation analysis. In contrast to matrix displays where correlations are indicated at intersections of rows and columns, our map conveys correlations by spatial proximity which is more direct and more focused on the variables in play. We make the following new contributions, some unique to our map: (1) we devise mechanisms that handle both categorical and numerical variables within a unified framework, (2) we achieve scalability for large numbers of variables via a multi-scale semantic zooming approach, (3) we provide interactive techniques for exploring the impact of value bracketing on correlations, and (4) we visualize data relations within the sub-spaces spanned by correlated variables by projecting the data into a corresponding tessellation of the map. Zhiyuan Zhang 0006, Kevin T. McDonnell, Erez Zadok, Klaus Mueller 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2014 | On the Importance of Evaluating Storage Systems' $Costs
Amanpreet Mukker, Erez Zadok |
HotStorage | 3 |
| 2014 | Linux NFSv4.1 Performance Under a Microscope
Ming Chen 0013, Dean Hildebrand, Geoffrey H. Kuenning, Soujanya Shankaranarayana, Vasily Tarasov, Arun O. Vasudevan, Erez Zadok, Ksenia Zakirova |
LISA | 7 |
| 2013 | Building workload-independent storage with VT-trees
Pradeep Shetty, Richard P. Spillane, Ravikant Malpani, Binesh Andrews, Justin Seyster, Erez Zadok |
FAST | 6 |
| 2013 | Virtual machine workloads: the case for new benchmarks for NAS
Vasily Tarasov, Dean Hildebrand, Geoffrey H. Kuenning, Erez Zadok |
FAST | 4 |
| 2013 | Improving I/O Performance Using Virtual Disk Introspection
Vasily Tarasov, Dean Hildebrand, Renu Tewari, Geoffrey H. Kuenning, Erez Zadok |
HotStorage | 6 |
| 2012 | Power consumption in enterprise-scale backup storage systems
Kevin M. Greenan, Andrew W. Leung, Erez Zadok |
FAST | 4 |
| 2012 | Extracting flexible, replayable models from large block traces
Vasily Tarasov, Santhosh Kumar, Jack Ma, Dean Hildebrand, Anna Povzner, Geoffrey H. Kuenning, Erez Zadok |
FAST | 7 |
| 2012 | Adaptive Runtime Verification
Ezio Bartocci, Radu Grosu, Atul Karmarkar, Scott A. Smolka, Scott D. Stoller, Erez Zadok, Justin Seyster |
RV | 6 |
| 2012 | Generating Realistic Datasets for Deduplication Analysis
Vasily Tarasov, Amar Mudrankit, Will Buik, Philip Shilane, Geoffrey H. Kuenning, Erez Zadok |
USENIX ATC | 6 |
| 2012 | InterAspect: aspect-oriented instrumentation with GCC
Justin Seyster, Ketan Dixit, Xiaowan Huang, Radu Grosu, Klaus Havelund, Scott A. Smolka, Scott D. Stoller, Erez Zadok |
Formal Methods Syst. Des. | 8 |
| 2012 | Don't Thrash: How to Cache Your Hash on FlashabstractThis paper presents new alternatives to the well-known Bloom filter data structure. The Bloom filter, a compact data structure supporting set insertion and membership queries, has found wide application in databases, storage systems, and networks. Because the Bloom filter performs frequent random reads and writes, it is used almost exclusively in RAM, limiting the size of the sets it can represent. This paper first describes the quotient filter, which supports the basic operations of the Bloom filter, achieving roughly comparable performance in terms of space and time, but with better data locality. Operations on the quotient filter require only a small number of contiguous accesses. The quotient filter has other advantages over the Bloom filter: it supports deletions, it can be dynamically resized, and two quotient filters can be efficiently merged. The paper then gives two data structures, the buffered quotient filter and the cascade filter, which exploit the quotient filter advantages and thus serve as SSD-optimized alternatives to the Bloom filter. The cascade filter has better asymptotic I/O performance than the buffered quotient filter, but the buffered quotient filter outperforms the cascade filter on small to medium data sets. Both data structures significantly outperform recently-proposed SSD-optimized Bloom filter variants, such as the elevator Bloom filter, buffered Bloom filter, and forest-structured Bloom filter. In experiments, the cascade filter and buffered quotient filter performed insertions 8.6--11 times faster than the fastest Bloom filter variant and performed lookups 0.94--2.56 times faster. Michael A. Bender, Martin Farach-Colton, Rob Johnson 0001, Russell Kraner, Bradley C. Kuszmaul, Dzejla Medjedovic, Pablo Montes, Pradeep Shetty, Richard P. Spillane, Erez Zadok |
Proc. VLDB Endow. | 10 |
| 2012 | Software monitoring with controllable overhead
Xiaowan Huang, Justin Seyster, Sean Callanan, Ketan Dixit, Radu Grosu, Scott A. Smolka, Scott D. Stoller, Erez Zadok |
Int. J. Softw. Tools Technol. Transf. | 8 |
| 2011 | An efficient multi-tier tablet server storage architectureabstractDistributed, structured data stores such as Big Table, HBase, and Cassandra use a cluster of machines, each running a database-like software system called the Tablet Server Storage Layer or TSSL. A TSSL's performance on each node directly impacts the performance of the entire cluster. In this paper we introduce an efficient, scalable, multi-tier storage architecture for tablet servers. Our system can use any layered mix of storage devices such as Flash SSDs and magnetic disks. Our experiments show that by using a mix of technologies, performance for certain workloads can be improved beyond configurations using strictly two-tier approaches with one type of storage technology. We utilized, adapted, and integrated cache-oblivious algorithms and data structures, as well as Bloom filters, to improve scalability significantly. We also support versatile, efficient transactional semantics. We analyzed and evaluated our system against the storage layers of Cassandra and Hadoop HBase. We used wide range of workloads and configurations from read- to write-optimized, as well as different input sizes. We found that our system is 3--10× faster than existing systems; that using proper data structures, algorithms, and techniques is critical for scalability, especially on modern Flash SSDs; and that one can fully support versatile transactions without sacrificing performance. Richard P. Spillane, Pradeep Shetty, Erez Zadok, Sagar Dixit, Shrikar Archak |
SoCC | 3 |
| 2011 | Benchmarking File System Benchmarking: It *IS* Rocket Science
Vasily Tarasov, Saumitra Bhanage, Erez Zadok, Margo I. Seltzer |
HotOS | 3 |
| 2011 | Don't Thrash: How to Cache Your Hash on Flash
Michael A. Bender, Martin Farach-Colton, Rob Johnson 0001, Bradley C. Kuszmaul, Dzejla Medjedovic, Pablo Montes, Pradeep Shetty, Richard P. Spillane, Erez Zadok |
HotStorage | 9 |
| 2011 | Redflag: A Framework for Analysis of Kernel-Level Concurrency
Justin Seyster, Prabakar Radhakrishnan, Samriti Katoch, Abhinav Duggal, Scott D. Stoller, Erez Zadok |
ICA3PP (1) | 6 |
| 2011 | Runtime Verification with State Estimation
Scott D. Stoller, Ezio Bartocci, Justin Seyster, Radu Grosu, Klaus Havelund, Scott A. Smolka, Erez Zadok |
RV | 7 |
| 2011 | On the energy consumption and performance of systems softwareabstractModels of energy consumption and performance are necessary to understand and identify system behavior, prior to designing advanced controls that can balance out performance and energy use. This paper considers the energy consumption and performance of servers running a relatively simple file-compression workload. We found that standard techniques for system identification do not produce acceptable models of energy consumption and performance, due to the intricate interplay between the discrete nature of software and the continuous nature of energy and performance. This motivated us to perform a detailed empirical study of the energy consumption and performance of this system with varying compression algorithms and compression levels, file types, persistent storage media, CPU DVFS levels, and disk I/O schedulers. Our results identify and illustrate factors that complicate the system's energy consumption and performance, including nonlinearity, instability, and multi-dimensionality. Our results provide a basis for future work on modeling energy consumption and performance to support principled design of controllable energy-aware systems. Radu Grosu, Priya Sehgal, Scott A. Smolka, Scott D. Stoller, Erez Zadok |
SYSTOR | 6 |
| 2010 | Evaluating Performance and Energy in File System Server Workloads
Priya Sehgal, Vasily Tarasov, Erez Zadok |
FAST | 3 |
| 2010 | Exporting kernel page caching for efficient user-level I/OabstractThe modern file system is still implemented in the kernel, and is statically linked with other kernel components. This architecture has brought performance and efficient integration with memory management. However kernel development is slow and modern storage systems must support an array of features, including distribution across a network, tagging, searching, deduplication, checksumming, snap-shotting, file pre-allocation, real time I/O guarantees for media, and more. To move complex components into user-level however will require an efficient mechanism for handling page faulting and zero-copy caching, write ordering, synchronous flushes, interaction with the kernel page write-back thread, and secure shared memory. We implement such a system, and experiment with a user-level object store built on top. Our object store is a complete re-design of the traditional storage stack and demonstrates the efficiency of our technique, and the flexibility it grants to user-level storage systems. Our current prototype file system incurs between a 1% and 6% overhead on the default native file system EXT3 for in-cache system workloads. Where the native kernel file system design has traditionally found its primary motivation. For update and insert intensive metadata workloads that are out-of-cache, we perform 39 times better than the native EXT3 file system, while still performing only 2 times worse on out-of-cache random lookups. Richard P. Spillane, Sagar Dixit, Shrikar Archak, Saumitra Bhanage, Erez Zadok |
MSST | 5 |
| 2010 | Aspect-Oriented Instrumentation with GCC
Justin Seyster, Ketan Dixit, Xiaowan Huang, Radu Grosu, Klaus Havelund, Scott A. Smolka, Scott D. Stoller, Erez Zadok |
RV | 8 |
| 2010 | Optimizing energy and performance for server-class file system workloadsabstractRecently, power has emerged as a critical factor in designing components of storage systems, especially for power-hungry data centers. While there is some research into power-aware storage stack components, there are no systematic studies evaluating each component's impact separately. Various factors like workloads, hardware configurations, and software configurations impact the performance and energy efficiency of the system. This article evaluates the file system's impact on energy consumption and performance. We studied several popular Linux file systems, with various mount and format options, using the FileBench workload generator to emulate four server workloads: Web, database, mail, and fileserver, on two different hardware configurations. The file system design, implementation, and available features have a significant effect on CPU/disk utilization, and hence on performance and power. We discovered that default file system options are often suboptimal, and even poor. In this article we show that a careful matching of expected workloads and hardware configuration to a single software configuration—the file system—can improve power-performance efficiency by a factor ranging from 1.05 to 9.4 times. Priya Sehgal, Vasily Tarasov, Erez Zadok |
ACM Trans. Storage | 3 |
| 2009 | Enabling Transactional File Access via Lightweight Kernel Extensions
Richard P. Spillane, Sachin Gaikwad, Manjunath Chinni, Erez Zadok, Charles P. Wright |
FAST | 4 |
| 2009 | Energy and performance evaluation of lossless file data compression on server systemsabstractData compression has been claimed to be an attractive solution to save energy consumption in high-end servers and data centers. However, there has not been a study to explore this. In this paper, we present a comprehensive evaluation of energy consumption for various file compression techniques implemented in software. We apply various compression tools available on Linux to a variety of data files, and we try them on server class and workstation class systems. We compare their energy and performance results against raw reads and writes. Our results reveal that software based data compression cannot be considered as a universal solution to reduce energy consumption. Various factors like the type of the data file, the compression tool being used, the read-to-write ratio of the workload, and the hardware configuration of the system impact the efficacy of this technique. In some cases, however, we found compression to save substantial energy and improve performance. Rachita Kothiyal, Vasily Tarasov, Priya Sehgal, Erez Zadok |
SYSTOR | 4 |
| 2009 | DHIS: discriminating hierarchical storageabstractA typical storage hierarchy comprises of components with varying performance and cost characteristics, providing multiple options for data placement. We propose and evaluate a hierarchical storage system, DHIS, that uses application-level hints to discriminate between data with different access characteristics, and then customizes its placement and caching policies to each type. The data placement decisions in DHIS are made in an online fashion, during data creation. Most existing solutions that attempt to customize data layout require moving data around, based on access characteristics. DHIS uses two kinds of information to make its decisions. First, it uses knowledge about higher-level pointers between blocks (for example, file system pointers) to understand the relationship between blocks and consequently, their importance. Second, DHIS defines a set of generic attributes that the higher layers can use to annotate data, conveying various properties such as importance, access pattern, etc. Based on these attributes, DHIS dynamically decides to place the data in the hierarchy best suited for its requirements. By doing so, DHIS solves a critical problem faced by storage vendors and developers of higher level storage software, in terms of choosing the most efficient policy among many alternatives. Through several benchmarks, we show that DHIS's data placement decisions improve performance significantly. Chaitanya Yalamanchili, Kiron Vijayasankar, Erez Zadok, Gopalan Sivathanu |
SYSTOR | 3 |
| 2008 | Software monitoring with bounded overheadabstractIn this paper, we introduce the new technique of high-confidence software monitoring (HCSM), which allows one to perform software monitoring with bounded overhead and concomitantly achieve high confidence in the observed error rates. HCSM is formally grounded in the theory of supervisory control of finite-state automata: overhead is controlled, while maximizing confidence, by disabling interrupts generated by the events being monitored - and hence avoiding the overhead associated with processing these interrupts - for as short a time as possible under the constraint of a user-supplied target overhead Otarget. HCSM is a general technique for software monitoring in that HCSM-based instrumentation can be attached at any system interface or API. A generic controller implements the optimal control strategy described above. As a proof of concept, and as a practical framework for software monitoring, we have implemented HCSM-based monitoring for both bounds checking and memory leak detection. We have further conducted an extensive evaluation of HCSM's performance on several real-world applications, including the Lighttpd Web server, and a number of special-purpose micro-benchmarks. Our results demonstrate how confidence grows in a monotonically increasing fashion with the target overhead, and that tight confidence intervals can be obtained for each target-overhead level. Sean Callanan, David J. Dean, Michael Gorbovitski, Radu Grosu, Justin Seyster, Scott A. Smolka, Scott D. Stoller, Erez Zadok |
IPDPS | 8 |
| 2008 | DARC: dynamic analysis of root causes of latency distributionsabstractOSprof is a versatile, portable, and efficient profiling methodology based on the analysis of latency distributions. Although OSprof has offers several unique benefits and has been used to uncover several interesting performance problems, the latency distributions that it provides must be analyzed manually. These latency distributions are presented as histograms and contain distinct groups of data, called peaks, that characterize the overall behavior of the running code. By automating the analysis process, we make it easier to take advantage of OSprof's unique features. Avishay Traeger, Ivan Deras, Erez Zadok |
SIGMETRICS | 3 |
| 2008 | Selective Versioning in a Secure Disk System
Swaminathan Sundararaman, Gopalan Sivathanu, Erez Zadok |
USENIX Security Symposium | 3 |
| 2008 | A nine year study of file system and storage benchmarkingabstractBenchmarking is critical when evaluating performance, but is especially difficult for file and storage systems. Complex interactions between I/O devices, caches, kernel daemons, and other OS components result in behavior that is rather difficult to analyze. Moreover, systems have different features and optimizations, so no single benchmark is always suitable. The large variety of workloads that these systems experience in the real world also adds to this difficulty. In this article we survey 415 file system and storage benchmarks from 106 recent papers. We found that most popular benchmarks are flawed and many research papers do not provide a clear indication of true performance. We provide guidelines that we hope will improve future performance evaluations. To show how some widely used benchmarks can conceal or overemphasize overheads, we conducted a set of experiments. As a specific example, slowing down read operations on ext2 by a factor of 32 resulted in only a 2--5% wall-clock slowdown in a popular compile benchmark. Finally, we discuss future work to improve file system and storage benchmarking. Avishay Traeger, Erez Zadok, Nikolai Joukov, Charles P. Wright |
ACM Trans. Storage | 2 |
| 2007 | Model Predictive Control for Memory ProfilingabstractWe make two contributions in the area of memory profiling. The first is a real-time, memory-profiling toolkit we call Memcov that provides both allocation/deallocation and access profiles of a running program. Memcov requires no recompilation or relinking and significantly reduces the barrier to entry for new applications of memory profiling by providing a clean, non-invasive way to perform two major functions: processing of the stream of memory-allocation events in real time and monitoring of regions in order to receive notification the next time they are hit. Our second contribution is an adaptive memory profiler and leak detector called MemcovMPC. Built on top of Memcov, MemcovMPCuses model predictive control to derive an optimal control strategy for leak detection that maximizes the number of areas monitored for leaks, while minimizing the associated runtime overhead. When it observes that an area has not been accessed for a user-definable period of time, it reports it as a potential leak. Our approach requires neither mark-and-sweep leak detection nor static analysis, and reports a superset of the memory leaks actually occurring as the program runs. The set of leaks reported by MemcovMPCcan be made to approximate the actual set more closely by lengthening the threshold period. Sean Callanan, Radu Grosu, Justin Seyster, Scott A. Smolka, Erez Zadok |
IPDPS | 5 |
| 2007 | RAIF: Redundant Array of Independent Filesystems
Nikolai Joukov, Arun M. Krishnakumar, Chaitanya Patti, Abhishek Rai, Sunil Satnur, Avishay Traeger, Erez Zadok |
MSST | 7 |
| 2007 | Extending ACID semantics to the file systemabstractAn organization's data is often its most valuable asset, but today's file systems provide few facilities to ensure its safety. Databases, on the other hand, have long provided transactions. Transactions are useful because they provide atomicity, consistency, isolation, and durability (ACID). Many applications could make use of these semantics, but databases have a wide variety of nonstandard interfaces. For example, applications like mail servers currently perform elaborate error handling to ensure atomicity and consistency, because it is easier than using a DBMS. A transaction-oriented programming model eliminates complex error-handling code because failed operations can simply be aborted without side effects. We have designed a file system that exports ACID transactions to user-level applications, while preserving the ubiquitous and convenient POSIX interface. In our prototype ACID file system, called Amino, updated applications can protect arbitrary sequences of system calls within a transaction. Unmodified applications operate without any changes, but each system call is transaction protected. We also built a recoverable memory library with support for nested transactions to allow applications to keep their in-memory data structures consistent with the file system. Our performance evaluation shows that ACID semantics can be added to applications with acceptable overheads. When Amino adds atomicity, consistency, and isolation functionality to an application, it performs close to Ext3. Amino achieves durability up to 46% faster than Ext3, thanks to improved locality. Charles P. Wright, Richard P. Spillane, Gopalan Sivathanu, Erez Zadok |
ACM Trans. Storage | 4 |
| 2006 | Compiler-assisted software verification using plug-insabstractWe present Protagoras, a new plug-in architecture for the GNU compiler collection that allows one to modify GCC's internal representation of the program under compilation. We illustrate the utility of Protagoras by presenting plug-ins for both compile-time and runtime software verification and monitoring. In the compile-time case, we have developed plug-ins that interpret the GIMPLE intermediate representation to verify properties statically. In the runtime case, we have developed plug-ins for GCC to perform memory leak detection, array bounds checking, and reference-count access monitoring. Sean Callanan, Radu Grosu, Xiaowan Huang, Scott A. Smolka, Erez Zadok |
IPDPS | 5 |
| 2006 | Operating System Profiling via Latency Analysis
Nikolai Joukov, Avishay Traeger, Rakesh Iyer, Charles P. Wright, Erez Zadok |
OSDI | 5 |
| 2006 | Type-Safe Disks
Gopalan Sivathanu, Swaminathan Sundararaman, Erez Zadok |
OSDI | 3 |
| 2006 | Versatility and Unix semantics in namespace unificationabstractAdministrators often prefer to keep related sets of files in different locations or media, as it is easier to maintain them separately. Users, however, prefer to see all files in one location for convenience. One solution that accommodates both needs is virtual namespace unification---providing a merged view of several directories without physically merging them. For example, namespace unification can merge the contents of several CD-ROM images without unpacking them, merge binary directories from different packages, merge views from several file servers, and more. Namespace unification can also enable snapshotting by marking some data sources read-only and then utilizing copy-on-write for the read-only sources. For example, an OS image may be contained on a read-only CD-ROM image---and the user's configuration, data, and programs could be stored in a separate read-write directory. With copy-on-write unification, the user need not be concerned about the two disparate file systems.It is difficult to maintain Unix semantics while offering a versatile namespace unification system. Past efforts to provide such unification often compromised on the set of features provided or Unix compatibility---resulting in an incomplete solution that users could not use.We designed and implemented a versatile namespace unification system called Unionfs . Unionfs maintains Unix semantics while offering advanced namespace unification features: dynamic insertion and removal of namespaces at any point in the merged view, mixing read-only and read-write components, efficient in-kernel duplicate elimination, NFS interoperability, and more. Since releasing our Linux implementation, it has been used by thousands of users and over a dozen Linux distributions, which helped us discover and solve many practical problems. Charles P. Wright, Jay Dave, Puja Gupta, Harikesavan Krishnan, David P. Quigley, Erez Zadok, Mohammad Nayyer Zubair |
ACM Trans. Storage | 6 |
| 2006 | On incremental file system developmentabstractDeveloping file systems from scratch is difficult and error prone. Using layered, or stackable, file systems is a powerful technique to incrementally extend the functionality of existing file systems on commodity OSes at runtime. In this article, we analyze the evolution of layering from historical models to what is found in four different present day commodity OSes: Solaris, FreeBSD, Linux, and Microsoft Windows. We classify layered file systems into five types based on their functionality and identify the requirements that each class imposes on the OS. We then present five major design issues that we encountered during our experience of developing over twenty layered file systems on four OSes. We discuss how we have addressed each of these issues on current OSes, and present insights into useful OS and VFS features that would provide future developers more versatile solutions for incremental file system development. Erez Zadok, Rakesh Iyer, Nikolai Joukov, Gopalan Sivathanu, Charles P. Wright |
ACM Trans. Storage | 1 |
| 2005 | Increasing distributed storage survivability with a stackable RAID-like file systemabstractWe have designed a stackable file system called Redundant Array of Independent Filesystems (RAIF). It combines the data survivability properties and performance benefits of traditional RAIDs with the unprecedented flexibility of composition, improved security, and ease of development of stackable file systems. RAIF can be mounted on top of any combination of other file systems including network, distributed, disk-based, and memory-based file systems. Existing encryption, compression, antivirus, and consistency checking stackable file systems can be mounted above and below RAIF, to efficiently cope up with slow or unsecure branches. Individual files can be distributed across branches, replicated, stored with parity, or stored with erasure correction coding to recover from failures on multiple branches. Per-file incremental recovery, storage type migration, and load-balancing are especially well suited for grid storages. In this paper, we describe the current RAIF design, provide preliminary performance results and discuss current status and future directions. Nikolai Joukov, Abhishek Rai, Erez Zadok |
CCGRID | 3 |
| 2005 | Accurate and Efficient Replaying of File System Traces
Nikolai Joukov, Timothy Wong, Erez Zadok |
FAST | 3 |
| 2005 | Versatile, portable, and efficient OS profiling via latency analysisabstractOperating systems are complex and their behavior depends on many factors. Source code, if available, does not directly help understand the OS's behavior, as the behavior depends on actual workloads and external inputs. Runtime profiling is a key technique for understanding the behavior and mutual-influence of modern OS components. Such profiling is useful to prove new concepts, debug problems, and optimize the performance of existing OSs. Unfortunately, existing profiling methods lack in important areas: they do not provide much of the necessary information about the OS's behavior; they require OS modification and therefore are not portable; or they exact high overheads thus perturbing the profiled OS. Nikolai Joukov, Rakesh Iyer, Avishay Traeger, Charles P. Wright, Erez Zadok |
SOSP | 5 |
| 2004 | Tracefs: A File System to Trace Them All
Akshat Aranya, Charles P. Wright, Erez Zadok |
FAST | 3 |
| 2004 | A Versatile and User-Oriented Versioning File System
Kiran-Kumar Muniswamy-Reddy, Charles P. Wright, Andrew Himmer, Erez Zadok |
FAST | 4 |
| 2004 | I3FS: An In-Kernel Integrity Checker and Intrusion Detection File System
Anand Kashyap, Gopalan Sivathanu, Erez Zadok |
LISA | 4 |
| 2004 | Reducing Storage Management Costs via Informed User-Based Policies
Erez Zadok, Jeffrey Osborn, Ariye Shater, Charles P. Wright, Kiran-Kumar Muniswamy-Reddy, Jason Nieh |
MSST | 1 |
| 2004 | Avfs: An On-Access Anti-Virus File System
Yevgeniy Miretskiy, Abhijith Das, Charles P. Wright, Erez Zadok |
USENIX Security Symposium | 4 |
| 2003 | Cosy: Develop in User-Land, Run in Kernel-Mode
Amit Purohit, Charles P. Wright, Joseph Spadavecchia, Erez Zadok |
HotOS | 4 |
| 2003 | NCryptfs: A Secure and Convenient Cryptographic File System
Charles P. Wright, Michael C. Martino, Erez Zadok |
USENIX ATC, General Track | 3 |
| 2002 | Toward Cost-Sensitive Modeling for Intrusion Detection and ResponseabstractIntrusion detection systems (IDSs) must maximize the realization of security goals while minimizing costs. In this paper, we study the problem of building cost-sensitive intrusion detection models. We examine the major cost factors associated with an IDS, which include development cost, operational cost, damage cost due to successful intrusions, and the cost of manual and automated response to intrusions. These cost factors can be qualified according to a defined attack taxonomy and site-specific security policies and priorities. We define cost models to formulate the total expected cost of an IDS, and present cost-sensitive machine learning techniques that can produce detection models that are optimized for user-defined cost metrics. Empirical experiments show that our cost-sensitive modeling and deployment techniques are effective in reducing the overall cost of intrusion detection. Wenke Lee, Wei Fan 0001, Salvatore J. Stolfo, Erez Zadok |
J. Comput. Secur. | 5 |
| 2001 | Data Mining Methods for Detection of New Malicious ExecutablesabstractA serious security threat today is malicious executables, especially new, unseen malicious executables often arriving as email attachments. These new malicious executables are created at the rate of thousands every year and pose a serious security threat. Current anti-virus systems attempt to detect these new malicious programs with heuristics generated by hand. This approach is costly and oftentimes ineffective. We present a data mining framework that detects new, previously unseen malicious executables accurately and automatically. The data mining framework automatically found patterns in our data set and used these patterns to detect a set of new malicious binaries. Comparing our detection methods with a traditional signature-based method, our method more than doubles the current detection rates for new malicious executables. Matthew G. Schultz, Eleazar Eskin, Erez Zadok, Salvatore J. Stolfo |
S&P | 3 |
| 2001 | Fast Indexing: Support for Size-Changing Algorithms in Stackable File Systems
Erez Zadok, Johan M. Andersen, Ion Badulescu, Jason Nieh |
USENIX ATC, General Track | 1 |
| 2000 | FiST: A Language for Stackable File Systems
Erez Zadok, Jason Nieh |
USENIX ATC, General Track | 1 |
| 1999 | Extending File Systems Using Stackable Templates
Erez Zadok, Ion Badulescu, Alex Shender |
USENIX ATC, General Track | 1 |
| 1993 | HLFSD: Delivering Email to Your $HOME
Erez Zadok, Alexander Dupuy |
LISA | 1 |