EDBT 2026 Demo / reviewers in the wild / expert
Ethan L. Miller
dblp:m/EthanLMiller
· DBLP profile ↗
94ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0003-2994-9060ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 81 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-authorSoftware engineering, systems software and programming languages · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3Security and privacy · 2Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Valet: Efficient Data Placement on Modern SSDsabstractThe increasing demand for ssds coupled with scaling difficulties has left manufacturers scrambling for newer ssd interfaces which promise better performance and durability. While these interfaces reduce the rigidity of traditional abstractions, they require application or system-level changes that can impact the stability, security, and portability of systems. To make matters worse, such changes are rendered futile with the introduction of next-generation interfaces. It is therefore no surprise that such interfaces have seen limited adoption, leaving behind a graveyard of experimental interfaces ranging from open-channel ssds to stream ssds. Devashish R. Purandare, Peter Alvaro, Avani Wildani, Darrell D. E. Long, Ethan L. Miller |
SoCC | 5 |
| 2024 | Secure Archival is Hard... Really HardabstractArchival systems are often tasked with storing highly valuable data that may be targeted by malicious actors. When the lifetime of the secret data is on the order of decades to centuries, the threat of improved cryptanalysis casts doubt on the long-term security of cryptographic techniques, which rely on hardness assumptions that are hard to prove over archival time scales. This threat makes the design of secure archival systems exceptionally difficult. Some archival systems turn a blind eye to this issue, hoping that current cryptographic techniques will not be broken; others often use techniques---such as secret sharing---that are impractical at scale. This position paper sheds light on the core challenges behind building practically viable secure long-term archives; we identify promising research avenues towards this goal. Maliha Tabassum, Soumya Chowdary Daruru, Gaurav Kulhare, Arvin Wang, Ethan L. Miller, Erez Zadok |
HotStorage | 6 |
| 2023 | TMC: Near-Optimal Resource Allocation for Tiered-Memory SystemsabstractMain memory dominates data center server cost, and hence data center operators are exploring alternative technologies such as CXL-attached and persistent memory to improve cost without jeopardizing performance. Introducing multiple tiers of memory introduces new challenges, such as selecting the appropriate memory configuration for a given workload mix. In particular, we observe that inefficient configurations increase cost by up to 2.6× for clients, and resource stranding increases cost by 2.2× for cloud operators. To address this challenge, we introduce TMC, a system for recommending cloud configurations according to workload characteristics and the dynamic resource utilization of a cluster. Whereas prior work utilized extensive simulation or costly machine learning techniques, incurring significant search costs, our approach profiles applications to reveal internal properties that lead to fast and accurate performance estimations. Our novel configuration-selection algorithm incorporates a new heuristic, packing penalty, to ensure that recommended configurations will also achieve good resource efficiency. Our experiments demonstrate that TMC reduces the search cost by up to 4× over the state-of-the-art, while improving resource utilization by up to 17% as compared to a naive policy that requests optimal tiered memory allocations in isolation. Yuanjiang Ni, Pankaj Mehra, Ethan L. Miller, Heiner Litz |
SoCC | 3 |
| 2023 | Persimmon: an append-only ZNS-first filesystemabstractWhile NAND flash has become the centerpiece of modern data center storage, legacy interfaces impact its performance and lifetime. Emulating in-place updates on SSDs results in frequent garbage collection, causing slowdowns, wear, and write amplification. Changing how we utilize modern SSDs to unlock their full potential is necessary. Even though Zoned Namespace SSDs provide an efficient append-only interface, filesystems on such drives still depend upon in-place updates and fixed metadata addresses, making their use with zoned storage complex and inefficient.We present Persimmon, a fork of the f2fs filesystem built with append-only metadata structures and tuned for zoned names-paces. Persimmon updates f2fs with new in-memory structures for better management in a zoned context, improved checkpoint logic, and append-only metadata management. Persimmon reduces tail latency, background garbage collection, and write amplification. Persimmon adapts its layout to zoned device constraints to unlock greater utilization and better drive cleanup. Devashish R. Purandare, Sam Schmidt, Ethan L. Miller |
ICCD | 3 |
| 2021 | Don't Let RPCs Constrain Your APIabstractAs data becomes increasingly distributed, traditional RPC and data serialization limits performance, result in rigidity, and hamper expressivity. We believe that technology trends including high-density persistent memory, high-speed networks, and programmable switches make this the right time to revisit prior research on distributed shared memory, global addressing, and content-based networking. Our vision combines the code mobility of RPC with first-class data references in a global address space by co-designing the OS and the network around pervasive data identity. We have initial results showing the promise of the proposed co-design. Daniel Bittman, Robert Soulé, Ethan L. Miller, Vishal Shrivastav, Pankaj Mehra, Matthew Boisvert, Avi Silberschatz, Peter Alvaro |
HotNets | 3 |
| 2021 | Twizzler: A Data-centric OS for Non-volatile MemoryabstractByte-addressable, non-volatile memory (NVM) presents an opportunity to rethink the entire system stack. We present Twizzler, an operating system redesign for this near-future. Twizzler removes the kernel from the I/O path, provides programs with memory-style access to persistent data using small (64 bit), object-relative cross-object pointers, and enables simple and efficient long-term sharing of data both between applications and between runs of an application. Twizzler provides a clean-slate programming model for persistent data, realizing the vision of Unix in a world of persistent RAM. We show that Twizzler is simpler, more extensible, and more secure than existing I/O models and implementations by building software for Twizzler and evaluating it on NVM DIMMs. Most persistent pointer operations in Twizzler impose less than 0.5 ns added latency. Twizzler operations are up to faster than Unix , and SQLite queries are up to faster than on PMDK. YCSB workloads ran 1.1– faster on Twizzler than on native and NVM-optimized SQLite backends. Daniel Bittman, Peter Alvaro, Pankaj Mehra, Darrell D. E. Long, Ethan L. Miller |
ACM Trans. Storage | 5 |
| 2020 | Geomancy: Automated Performance Enhancement through Data Layout OptimizationabstractThe size and complexity of large storage systems, such as high-performance computing (HPC) systems, inhibit rapid effective restructuring of data layouts to maintain performance as workloads shift. To address this issue, we have developed Geomancy, a tool that models the placement of data within a distributed storage system and reacts to drops in performance. Our approach to optimizing throughput offers benefits for storage systems such as avoiding potential bottlenecks and increasing overall I/O throughput from 11% to 30%. Oceane Bel, Kenneth Chang, Nathan R. Tallent, Dirk Düllmann, Ethan L. Miller, Faisal Nawab, Darrell D. E. Long |
ISPASS | 5 |
| 2020 | Twizzler: a Data-Centric OS for Non-Volatile Memory
Daniel Bittman, Peter Alvaro, Pankaj Mehra, Darrell D. E. Long, Ethan L. Miller |
USENIX ATC | 5 |
| 2020 | The Future of the Past: Challenges in Archival Storage
Ethan L. Miller |
USENIX ATC | 1 |
| 2019 | Optimizing Systems for Byte-Addressable NVM by Reducing Bit Flipping
Daniel Bittman, Darrell D. E. Long, Peter Alvaro, Ethan L. Miller |
FAST | 4 |
| 2019 | A Tale of Two Abstractions: The Case for Object Space
Daniel Bittman, Peter Alvaro, Darrell D. E. Long, Ethan L. Miller |
HotStorage | 4 |
| 2019 | String Figure: A Scalable and Elastic Memory Network ArchitectureabstractDemand for server memory capacity and performance is rapidly increasing due to expanding working set sizes of modern applications, such as big data analytics, inmemory computing, deep learning, and server virtualization. One promising techniques to tackle this requirements is memory networking, whereby a server memory system consists of multiple 3D die-stacked memory nodes interconnected by a high-speed network. However, current memory network designs face substantial scalability and flexibility challenges. This includes (1) maintaining high throughput and low latency in large-scale memory networks at low hardware cost, (2) efficiently interconnecting an arbitrary number of memory nodes, and (3) supporting flexible memory network scale expansion and reduction without major modification of the memory network design or physical implementation. To address the challenges, we propose String Figure1, a highthroughput, elastic, and scalable memory network architecture. String Figure consists of (1) an algorithm to generate random topologies that achieve high network throughput and nearoptimal path lengths in large-scale memory networks, (2) a hybrid routing protocol that employs a mix of computation and look up tables to reduce the overhead of both in routing, (3) a set of network reconfiguration mechanisms that allow both static and dynamic network expansion and reduction. Our experiments using RTL simulation demonstrate that String Figure can interconnect over one thousand memory nodes with a shortest path length within five hops across various traffic patterns and real workloads. Matheus Ogleari, Ye Yu 0001, Chen Qian 0001, Ethan L. Miller, Jishen Zhao |
HPCA | 4 |
| 2019 | SSP: Eliminating Redundant Writes in Failure-Atomic NVRAMs via Shadow Sub-PagingabstractNon-Volatile Random Access Memory (NVRAM) technologies are closing the performance gap between traditional storage and memory. However, the integrity of persistent data structures after an unclean shutdown remains a major concern. Logging is commonly used to ensure consistency of NVRAM systems, but it imposes significant performance overhead and causes additional wear out by writing extra data into NVRAM. Our goal is to eliminate the extra writes that are needed to achieve consistency. SSP (i) exploits a novel cache-line-level remapping mechanism to eliminate redundant data copies in NVRAM, (ii) minimizes the storage overheads using page consolidation and (iii) removes failure-atomicity overheads from the critical path, significantly improving the performance of NVRAM systems. Our evaluation results demonstrate that SSP reduces overall write traffic by up to 1.8×, reduces extra NVRAM writes in the critical path by up to 10× and improves transaction throughput by up to 1.6×, compared to a state-of-the-art logging design. Yuanjiang Ni, Jishen Zhao, Heiner Litz, Daniel Bittman, Ethan L. Miller |
MICRO | 5 |
| 2019 | A Persistent Problem: Managing Pointers in NVMabstractByte-addressable non-volatile memory (NVM) placed alongside DRAM promises a fundamental shift in software abstractions, yet many approaches to using NVM promise merely incremental improvement by relying on old interfaces and archaic abstractions. We assert that redesigning the core programming model presented by the operating system is vital to best exploiting this technology. We are developing Twizzler, an OS that presents an effective programming model for NVM sufficient to construct persistent data structures that can be easily and globally shared without serialization costs. We consider and evolve a key-value store that runs on Twizzler, and demonstrate how our programming model improves programmability with early experiments indicating performance need not be lost and may be improved. Daniel Bittman, Peter Alvaro, Ethan L. Miller |
PLOS@SOSP | 3 |
| 2018 | Alpha Entanglement Codes: Practical Erasure Codes to Archive Data in Unreliable EnvironmentsabstractData centres that use consumer-grade disks drives and distributed peer-to-peer systems are unreliable environments to archive data without enough redundancy. Most redundancy schemes are not completely effective for providing high availability, durability and integrity in the long-term. We propose alpha entanglement codes, a mechanism that creates a virtual layer of highly interconnected storage devices to propagate redundant information across a large scale storage system. Our motivation is to design flexible and practical erasure codes with high fault-tolerance to improve data durability and availability even in catastrophic scenarios. By "flexible and practical", we mean code settings that can be adapted to future requirements and practical implementations with reasonable trade-offs between security, resource usage and performance. The codes have three parameters. Alpha increases storage overhead linearly but increases the possible paths to recover data exponentially. Two other parameters increase fault-tolerance even further without the need of additional storage. As a result, an entangled storage system can provide high availability, durability and offer additional integrity: it is more difficult to modify data undetectably. We evaluate how several redundancy schemes perform in unreliable environments and show that alpha entanglement codes are flexible and practical codes. Remarkably, they excel at code locality, hence, they reduce repair costs and become less dependent on storage locations with poor availability. Our solution outperforms Reed-Solomon codes in many disaster recovery scenarios. Vero Estrada-Galiñanes, Ethan L. Miller, Pascal Felber, Jehan-François Pâris |
DSN | 2 |
| 2018 | Reducing NVM Writes with Optimized Shadow Paging
Yuanjiang Ni, Jishen Zhao, Daniel Bittman, Ethan L. Miller |
HotStorage | 4 |
| 2018 | Steal but No Force: Efficient Hardware Undo+Redo Logging for Persistent Memory SystemsabstractPersistent memory is a new tier of memory that functions as a hybrid of traditional storage systems and main memory. It combines the benefits of both: the data persistence of storage with the fast load/store interface of memory. Most previous persistent memory designs place careful control over the order of writes arriving at persistent memory. This can prevent caches and memory controllers from optimizing system performance through write coalescing and reordering. We identify that such write-order control can be relaxed by employing undo+redo logging for data in persistent memory systems. However, traditional software logging mechanisms are expensive to adopt in persistent memory due to performance and energy overheads. Previously proposed hardware logging schemes are inefficient and do not fully address the issues in software. To address these challenges, we propose a hardware undo+redo logging scheme which maintains data persistence by leveraging the write-back, write-allocate policies used in commodity caches. Furthermore, we develop a cache force-write-back mechanism in hardware to significantly reduce the performance and energy overheads from forcing data into persistent memory. Our evaluation across persistent memory microbenchmarks and real workloads demonstrates that our design significantly improves system throughput and reduces both dynamic energy and memory traffic. It also provides strong consistency guarantees compared to software approaches. Matheus Ogleari, Ethan L. Miller, Jishen Zhao |
HPCA | 2 |
| 2018 | Using Simulation to Design Scalable and Cost-Efficient Archival Storage SystemsabstractThe need for reliable and cost-effective data storage grows as digital information becomes increasingly ubiquitous. Archival systems must store valuable data for years while adapting to changing user needs, capacity, and performance requirements. Storage devices differ in terms of performance, capacity, reliability, acquisition cost, power consumption, and the rates at which their features change over time. As a result, choosing the best storage technology to use for an archive has become increasingly challenging with the proliferation of new technologies alongside existing ones. We have designed a simulator that models the capacity, performance, acquisition cost, and power cost of an archival system using the characteristics of the drives and media that comprise it. We simulate and compare four storage technologies that exhibit different cost and performance characteristics: tape, optical disc, hard disk, and NAND flash SSD. We evaluate the total cost of ownership for each storage technology within an archival system, and we explore the effect that prospective technological advancements and growth rates over time may have on the relative cost and viability of each storage technology for archival systems. We show that the lifecycle and upgrade cost of drives are significant cost factors for removable media archives. We observe that increasing performance requires adding more drives to an archival system, and the cost of each drive dominates the cost to increase performance. We compare trends in storage technologies to suggest developments that could minimize the long-term total cost of ownership for archival systems. We show that hard disks and flash could become cost-competitive with tape-based archives by adopting new designs to minimize infrastructure and electricity costs. James Byron, Darrell D. E. Long, Ethan L. Miller |
MASCOTS | 3 |
| 2018 | Efficient Reconstruction Techniques for Disaster Recovery in Secret-Split DatastoresabstractIncreasingly, archival systems are relying on authentication-based techniques that leverage secret-splitting rather than encryption to secure data for long-term storage. Secret-splitting data across multiple independent repositories reduces complexities in key management, eliminates the need for updates due to encryption algorithm deprecation over time, and reduces the risk of insider compromise. While reconstruction of stored data objects is straightforward if a user-maintained index is available, the system must also support disaster recovery incase the index is unavailable. Designing a mechanism for efficient index-free reconstruction, that does not increase the risk of attacker compromise, is a challenge. Reconstruction requires the association of chunks that make up an object, which is the kind of information attackers can use to identify chunks they must steal to illicitly obtain data. We propose two new techniques, the set-subset reconstruction and secret-split secure hash (S3H) reconstruction, which allow chunks of data to be correlated and quickly reconstructed without providing useful information to an attacker. Both techniques operate on the entire collections of secret-split chunks in the archive. While they can efficiently rebuild an entire archive, they are inefficient and impractical for rebuilding single objects, making them useless for attackers that do not have access to all of the data. These techniques can each be tuned to trade-off between reconstruction performance and security, reducing overall runtime from O(NK) (for N objects requiring K recombined chunks each to return the original object) to between O(N) and O(N2). These runtimes are practical for archives containing as many as 107objects for the secret-split secure hash method and 109objects for the set-subset method. Larger archives can run these techniques with manageable runtimes by grouping data into separate smaller collections and running the algorithms on each collection in parallel. Sinjoni Mukhopadhyay, Joel Cameron Frank, Justin King, Daniel Bittman, Darrell D. E. Long, Ethan L. Miller |
MASCOTS | 6 |
| 2018 | Inkpack: A Secure, Data-Exposure Resistant Storage SystemabstractRemoving hard drives from a data center may expose sensitive data, such as encryption keys or passwords. To prevent exposure, data centers have security policies in place to physically secure drives in the system, and securely delete data from drives that are removed. Despite advances in security technology and best practices, implementation of these security measures is often done incorrectly. We anticipate that physical security will fail, and fixing the issue after the failure is costly and ineffective. Oceane Bel, Kenneth Chang, Daniel Bittman, Darrell D. E. Long, Hiroshi Isozaki, Ethan L. Miller |
SYSTOR | 6 |
| 2017 | CAPES: unsupervised storage performance tuning using neural network-based deep reinforcement learningabstractParameter tuning is an important task of storage performance optimization. Current practice usually involves numerous tweak-benchmark cycles that are slow and costly. To address this issue, we developed CAPES, a model-less deep reinforcement learning-based unsupervised parameter tuning system driven by a deep neural network (DNN). It is designed to find the optimal values of tunable parameters in computer systems, from a simple client-server system to a large data center, where human tuning can be costly and often cannot achieve optimal performance. CAPES takes periodic measurements of a target computer system's state, and trains a DNN which uses Q-learning to suggest changes to the system's current parameter values. CAPES is minimally intrusive, and can be deployed into a production system to collect training data and suggest tuning actions during the system's daily operation. Evaluation of a prototype on a Lustre file system demonstrates an increase in I/O throughput up to 45% at saturation point. Yan Li 0006, Kenneth Chang, Oceane Bel, Ethan L. Miller, Darrell D. E. Long |
SC | 4 |
| 2016 | Pilot: A Framework that Understands How to Do Performance Benchmarks the Right WayabstractCarrying out even the simplest performance benchmark requires considerable knowledge of statistics and computer systems, and painstakingly following many error-prone steps, which are distinct skill sets yet essential for getting statistically valid results. As a result, many performance measurements in peer-reviewed publications are flawed. Among many problems, they fall short in one or more of the following requirements: accuracy, precision, comparability, repeatability, and control of overhead. This is a serious problem because poor performance measurements misguide system design and optimization. We propose a collection of algorithms and heuristics to automate these steps. They cover the collection, storing, analysis, and comparison of performance measurements. We implement these methods as a readily-usable open source software framework called Pilot, which can help to reduce human error and shorten benchmark time. Evaluation of Pilot on various benchmarks show that it can reduce the cost and complexity of running benchmarks, and can produce better measurement results. Yan Li 0006, Yash Gupta, Ethan L. Miller, Darrell D. E. Long |
MASCOTS | 3 |
| 2016 | RESAR: Reliable Storage at Exabyte ScaleabstractStored data needs to be protected against device failure and irrecoverable sector read errors, yet doing so at exabyte scale can be challenging given the large number of failures that must be handled. We have developed RESAR (Robust, Efficient, Scalable, Autonomous, Reliable) storage, an approach to storage system redundancy that only uses XOR-based parity and employs a graph to lay out data and parity. The RESAR layout offers greater robustness and higher flexibility for repair at the same overhead as a declustered version of RAID 6. For instance, a RESAR-based layout with 16 data disklets per stripe has about 50 times lower probability of suffering data loss in the presence of a fixed number of failures than a corresponding RAID 6 organization. RESAR uses a layer of virtual storage elements to achieve better manageability, a broader potential for energy savings, as well as easier adoption of heterogeneous storage devices. Thomas J. E. Schwarz, Ahmed Amer, Tom M. Kroeger, Ethan L. Miller, Darrell D. E. Long, Jehan-François Pâris |
MASCOTS | 4 |
| 2016 | Effects of prolonged media usage and long-term planning on archival systemsabstractIn archival systems, storage media are often replaced much earlier than their expected service life in exchange for other benefits of new media, such as higher capacity, bandwidth, and I/O operations per second, or lower costs. In an era of decreasing media density growth rates, retiring media early by considering only short-term benefits while discarding potential long-term cost benefits could have a negative long-term impact on an archival system's economics. To extend an archival system's life, at low cost, while limiting performance degradation, we suggest extending media lifetime past manufacturer recommendations as well as increasing the horizon for planning and provisioning future media purchases. We present a cost-benefit analysis of the impact of prolonged media usage and long-term planning. Through Monte Carlo simulation, we simulate the behavior of an archival system using tapes, hard disk drives (HDDs), solid state devices (SSDs), and Blu-ray discs. We show that leaving older media in the archival system makes economic sense for SSDs without significantly affecting reliability; we show cost improvements of approximately 10% for SSDs for a low annual media density growth rate, such as 5%, which would have been a loss of 35%, for a high annual media density rate, such as 20%. We show that, for SSDs and hard disks, the optimal planning time of an archival system is at least as long as the media service life. Combining prolonged media usage with an extended planning horizon reduced costs by 15% for a system using SSDs. Avani Wildani, Ethan L. Miller, David S. H. Rosenthal, Darrell D. E. Long |
MSST | 3 |
| 2016 | Classifying Data to Reduce Long-Term Data Movement in Shingled Write DisksabstractShingled magnetic recording (SMR) is a means of increasing the density of hard drives that brings a new set of challenges. Due to the nature of SMR disks, updating in place is not an option. Holes left by invalidated data can only be filled if the entire band is reclaimed, and a poor band compaction algorithm could result in spending a lot of time moving blocks over the lifetime of the device. We propose using write frequency to separate blocks to reduce data movement and develop a band compaction algorithm that implements this heuristic. We demonstrate how our algorithm results in improved data management, resulting in an up to 45% reduction in required data movements when compared to naive approaches to band management. Stephanie N. Jones, Ahmed Amer, Ethan L. Miller, Darrell D. E. Long, Rekha Pitchumani, Christina R. Strong |
ACM Trans. Storage | 3 |
| 2016 | Can We Group Storage? Statistical Techniques to Identify Predictive Groupings in Storage System AccessesabstractStoring large amounts of data for different users has become the new normal in a modern distributed cloud storage environment. Storing data successfully requires a balance of availability, reliability, cost, and performance. Typically, systems design for this balance with minimal information about the data that will pass through them. We propose a series of methods to derive groupings from data that have predictive value, informing layout decisions for data on disk. Unlike previous grouping work, we focus on dynamically identifying groupings in data that can be gathered from active systems in real time with minimal impact using spatiotemporal locality. We outline several techniques we have developed and discuss how we select particular techniques for particular workloads and application domains. Our statistical and machine-learning-based grouping algorithms answer questions such as “What can a grouping be based on?” and “Is a given grouping meaningful for a given application?” We design our models to be flexible and require minimal domain information so that our results are as broadly applicable as possible. We intend for this work to provide a launchpad for future specialized system design using groupings in combination with caching policies and architectural distinctions such as tiered storage to create the next generation of scalable storage systems. Avani Wildani, Ethan L. Miller |
ACM Trans. Storage | 2 |
| 2015 | ASCAR: Automating contention management for high-performance storage systemsabstractHigh-performance parallel storage systems, such as those used by supercomputers and data centers, can suffer from performance degradation when a large number of clients are contending for limited resources, like bandwidth. These contentions lower the efficiency of the system and cause unwanted speed variances. We present the Automatic Storage Contention Alleviation and Reduction system (ASCAR), a storage traffic management system for improving the bandwidth utilization and fairness of resource allocation. ASCAR regulates I/O traffic from the clients using a rule based algorithm that controls the congestion window and rate limit. The rule-based client controllers are fast responding to burst I/O because no runtime coordination between clients or with a central coordinator is needed; they are also autonomous so the system has no scale-out bottleneck. Finding optimal rules can be a challenging task that requires expertise and numerous experiments. ASCAR includes a SHAred-nothing Rule Producer (SHARP) that produces rules in an unsupervised manner by systematically exploring the solution space of possible rule designs and evaluating the target workload under the candidate rule sets. Evaluation shows that our ASCAR prototype can improve the throughput of all tested workloads - some by as much as 35%. ASCAR improves the throughput of a NASA NPB BTIO checkpoint workload by 33.5% and reduces its speed variance by 55.4% at the same time. The optimization time and controller overhead are unrelated to the scale of the system; thus, it has the potential to support future large-scale systems that can have millions of clients and thousands of servers. As a pure client-side solution, ASCAR needs no change to either the hardware or server software. Yan Li 0006, Xiaoyuan Lu, Ethan L. Miller, Darrell D. E. Long |
MSST | 3 |
| 2015 | Percival: A searchable secret-split datastoreabstractMaintaining information privacy is challenging when sharing data across a distributed long-term datastore. In such applications, secret splitting the data across independent sites has been shown to be a superior alternative to fixed-key encryption; it improves reliability, reduces the risk of insider threat, and removes the issues surrounding key management. However, the inherent security of such a datastore normally precludes it from being directly searched without reassembling the data; this, however, is neither computationally feasible nor without risk since reassembly introduces a single point of compromise. As a result, the secret-split data must be pre-indexed in some way in order to facilitate searching. Previously, fixed-key encryption has also been used to securely pre-index the data, but in addition to key management issues, it is not well suited for long term applications. To meet these needs, we have developed Percival: a novel system that enables searching a secret-split datastore while maintaining information privacy. We leverage salted hashing, performed within hardware security modules, to access prerecorded queries that have been secret split and stored in a distributed environment; this keeps the bulk of the work on each client, and the data custodians blinded to both the contents of a query as well as its results. Furthermore, Percival does not rely on the datastore's exact implementation. The result is a flexible design that can be applied to both new and existing secret-split datastores. When testing Percival on a corpus of approximately one million files, it was found that the average search operation completed in less than one second. Joel Cameron Frank, Shayna M. Frank, Lincoln Thurlow, Tom M. Kroeger, Ethan L. Miller, Darrell D. E. Long |
MSST | 5 |
| 2015 | Classifying data to reduce long term data movement in shingled write disksabstractShingled Magnetic Recording (SMR) is a means of increasing the density of hard drives that brings a new set of challenges. Due to the nature of SMR disks, updating in place is not an option. Holes left by invalidated data can only be filled if the entire band is reclaimed, and a poor band compaction algorithm could result in spending a lot of time moving blocks over the lifetime of the device. We propose using write frequency to separate blocks to reduce data movement and develop a band compaction algorithm that implements this heuristic. We demonstrate how our algorithm results in improved data management, resulting in an up to 47% reduction in required data movements when compared to naive approaches to band management. Stephanie N. Jones, Ahmed Amer, Ethan L. Miller, Darrell D. E. Long, Rekha Pitchumani, Christina R. Strong |
MSST | 3 |
| 2015 | Realistic request arrival generation in storage benchmarksabstractBenchmarks are widely used to perform apples-to-apples comparison in a controlled and reliable fashion. Benchmarks must model real world workload behavior. In recent years, to meet web scale demands, Key-Value (KV) stores have emerged as a vital component of cloud serving systems. The Yahoo! Cloud Serving Benchmark (YCSB) has emerged as the standard benchmark for evaluating key-value systems, and has been preferred by both the industry and academia. Though YCSB provides a variety of options to generate realistic workloads, like most benchmarks it has ignored the temporal characteristics of generated workloads. YCSB's constant-rate request arrival process is unrealistic and fails to capture the real world arrival patterns. Existing workload studies on disk, filesystem, key-value system, network, and web traffic all show that they all exhibit some common temporal properties such as burstiness, self similarity, long range dependence, and diurnal activity. In this work, we show that the commonly observed traffic patterns can be modeled using the three categories of arrival processes: a)Poisson, b)Self similar, and c)Envelope-guided process. The three categories presented are a necessary and sufficient set of request arrival models that all storage benchmarks should provide. To demonstrate the ease of incorporating the models in benchmarks, we have modified YCSB to generate workloads based on all three models, and show the effect of realistic request arrivals through an example database evaluation. Rekha Pitchumani, Shayna M. Frank, Ethan L. Miller |
MSST | 3 |
| 2015 | Purity: Building Fast, Highly-Available Enterprise Flash Storage from Commodity ComponentsabstractAlthough flash storage has largely replaced hard disks in consumer class devices, enterprise workloads pose unique challenges that have slowed adoption of flash in ``performance tier'' storage appliances. In this paper, we describe Purity, the foundation of Pure Storage's Flash Arrays, the first all-flash enterprise storage system to support compression, deduplication, and high-availability. John Colgrove, John D. Davis, John Hayes, Ethan L. Miller, Cary Sandvig, Russell Sears, Ari Tamches, Neil Vachharajani, Feng Wang 0003 |
SIGMOD Conference | 4 |
| 2015 | SMRDB: key-value data store for shingled magnetic recording disksabstractShingled Magnetic Recording (SMR) disks employ a shingled write process that overlaps the data tracks on the disk surface like the shingles on a roof, thereby increasing disk areal density with minimal manufacturing changes. While these disks have the same read behavior as current disks, random writes and in-place data updates are no longer possible, since a write to a track must overwrite and destroy data on all tracks that it overlaps. Rekha Pitchumani, James P. Hughes 0001, Ethan L. Miller |
SYSTOR | 3 |
| 2014 | An Economic Perspective of Disk vs. Flash Media in Archival StorageabstractFor three decades, Kryder's law correctly predicted an exponential increase in bit density on disk platters, leading to an exponential drop in cost per gigabyte, and thus to an entrenched expectation that if data could be stored for a few years the incremental cost of storing it forever would be minimal. However, disk now is over 7 times as expensive as Kryder's law would have predicted, and industry projections suggest that in 2020 the gap will reach 200 times, disrupting this expectation. Our model shows that archives based upon alternative media are surprisingly cost competitive with archives based upon traditional disk media over the long-term. We propose using Archival Flash for long-term data preservation, with the trade off between longer data retention period and lower write cycles. Avani Wildani, Ethan L. Miller, Daniel C. Rosenthal, Ian F. Adams, Christina E. Strong, Andy Hospodor |
MASCOTS | 3 |
| 2014 | PERSES: Data Layout for Low Impact FailuresabstractGrowth in disk capacity continues to outpace advances in read speed and device reliability. This has led to storage systems spending increasing amounts of time in a degraded state while failed disks reconstruct. Users and applications that do not use the data on the failed or degraded drives are negligibly impacted by the failure, increasing the perceived performance of the system. We leverage this observation with PERSES, a statistical data allocation scheme to reduce the performance impact of reconstruction after disk failure. PERSES reduces degradation from the perspective of the user by clustering data on disks such that data with high probability of co-access is placed on the same device as often as possible. Trace-driven simulations show that, by laying out data with PERSES, we can reduce the perceived time lost due to failure over three years by up to 80% compared to arbitrary allocation. Avani Wildani, Ethan L. Miller, Ian F. Adams, Darrell D. E. Long |
MASCOTS | 2 |
| 2014 | Muninn: a Versioning Flash Key-Value Store Using an Object-based Storage ModelabstractWhile non-volatile memory (NVRAM) devices have the potential to alleviate the trade-off between performance, scalability, and energy in storage and memory subsystems, a block interface and storage subsystems designed for slow I/O devices make it difficult to efficiently exploit NVRAMs in a portable and extensible way. Yangwook Kang, Rekha Pitchumani, Thomas Marlette, Ethan L. Miller |
SYSTOR | 4 |
| 2014 | A File By Any Other Name: Managing File Names with MetadataabstractFile names are one of the earliest computing abstractions, a string of characters to uniquely identify a file for the system, and to help users remember the contents when they look for it later. They are also a rich source of semantic metadata about files. However, this metadata is unstructured and opaque to the rest of the system. As a result, metadata in file names is often error-prone, and hard to search for. File names can and should be more meaningful and reliable, while simplifying application design and encouraging users and applications to provide more metadata for search. Aleatha Parker-Wood, Darrell D. E. Long, Ethan L. Miller, Philippe Rigaux, Andy Isaacson |
SYSTOR | 3 |
| 2014 | Random Slicing: Efficient and Scalable Data Placement for Large-Scale Storage SystemsabstractThe ever-growing amount of data requires highly scalable storage solutions. The most flexible approach is to use storage pools that can be expanded and scaled down by adding or removing storage devices. To make this approach usable, it is necessary to provide a solution to locate data items in such a dynamic environment. This article presents and evaluates the Random Slicing strategy, which incorporates lessons learned from table-based, rule-based, and pseudo-randomized hashing strategies and is able to provide a simple and efficient strategy that scales up to handle exascale data. Random Slicing keeps a small table with information about previous storage system insert and remove operations, drastically reducing the required amount of randomness while delivering a perfect load distribution. Alberto Miranda, Sascha Effert, Yangwook Kang, Ethan L. Miller, Ivan Popov, André Brinkmann, Tom Friedetzky, Toni Cortes |
ACM Trans. Storage | 4 |
| 2013 | Horus: fine-grained encryption-based security for large-scale storage
Yan Li 0006, Nakul Sanjay Dhotre, Yasuhiro Ohara, Tom M. Kroeger, Ethan L. Miller, Darrell D. E. Long |
FAST | 5 |
| 2013 | Screaming fast Galois field arithmetic using intel SIMD instructions
James S. Plank, Kevin M. Greenan, Ethan L. Miller |
FAST | 3 |
| 2013 | HANDS: A heuristically arranged non-backup in-line deduplication systemabstractDeduplicating in-line data on primary storage is hampered by the disk bottleneck problem, an issue which results from the need to keep an index mapping portions of data to hash values in memory in order to detect duplicate data without paying the performance penalty of disk paging. The index size is proportional to the volume of unique data, so placing the entire index into RAM is not cost effective with a deduplication ratio below 45%. HANDS reduces the amount of in-memory index storage required by up to 99% while still achieving between 30% and 90% of the deduplication a full memory-resident index provides, making primary deduplication cost effective in workloads with deduplication rates as low as 8%. HANDS is a framework that dynamically pre-fetches fingerprints from disk into memory cache according to working sets statistically derived from access patterns. We use a simple neighborhood grouping as our statistical technique to demonstrate the effectiveness of our approach. HANDS is modular and requires only spatio-temporal data, making it suitable for a wide range of storage systems without the need to modify host file systems. Avani Wildani, Ethan L. Miller, Ohad Rodeh |
ICDE | 2 |
| 2013 | Validating Storage System InstrumentationabstractThere is a large body of work-such as system administration and intrusion detection-that relies upon storage system logs and snapshots. These solutions rely on accurate system records, however, little effort has been made to verify the correctness of logging instrumentation and log reliability. We present a solution, called ExDiff, that uses expectation differencing to validate storage system logs. Our solution can identify development errors such as the omission of a logging point and runtime errors such as log crashes. ExDiff uses metadata snapshots and activity logs to predict the expected state of the system and compares that with the system's actual state. Mismatches between the expected and actual metadata states can then be used to highlight gaps in log coverage, as well as aid in identifying specific types of missing entries. We show that ExDiff provides valuable insight to system designers, administrators and researchers by accurately identifying gaps in log coverage, providing clues useful in isolating specific types of missing log entries, and highlighting potential misunderstandings in logged action. Ian F. Adams, Mark W. Storer, Avani Wildani, Ethan L. Miller, Brian A. Madden |
MASCOTS | 4 |
| 2013 | Single-Snapshot File System AnalysisabstractMetadata snapshots are a common method for gaining insight into file systems due to their small size and relative ease of acquisition. Since they are static, most researchers have used them for relatively simple analyses such as file size distributions and age of files. We hypothesize that it is possible to gain much richer insights into file system and user behavior by clustering features in metadata snapshots and comparing the entropy within clusters to the entropy within natural partitions such as directory hierarchies. We discuss several different methods for gaining deeper insights into metadata snapshots, and show a small proof of concept using data from Los Alamos National Laboratories. In our initial work, we see evidence that it is possible to identify user locality information, traditionally the purview of dynamic traces, using a single static snapshot. Avani Wildani, Ian F. Adams, Ethan L. Miller |
MASCOTS | 3 |
| 2013 | Enabling cost-effective data processing with smart SSDabstractThis paper explores the benefits and limitations of in-storage processing on current Solid-State Disk (SSD) architectures. While disk-based in-storage processing has not been widely adopted, due to the characteristics of hard disks, modern SSDs provide high performance on concurrent random writes, and have powerful processors, memory, and multiple I/O channels to flash memory, enabling in-storage processing with almost no hardware changes. In addition, offloading I/O tasks allows a host system to fully utilize devices' internal parallelism without knowing the details of their hardware configurations. To leverage the enhanced data processing capabilities of modern SSDs, we introduce the Smart SSD model, which pairs in-device processing with a powerful host system capable of handling data-oriented tasks without modifying operating system code. By isolating the data traffic within the device, this model promises low energy consumption, high parallelism, low host memory footprint and better performance. To demonstrate these capabilities, we constructed a prototype implementing this model on a real SATA-based SSD. Our system uses an object-based protocol for low-level communication with the host, and extends the Hadoop MapReduce framework to support a Smart SSD. Our experiments show that total energy consumption is reduced by 50% due to the low-power processing inside a Smart SSD. Moreover, a system with a Smart SSD can outperform host-side processing by a factor of two or three by efficiently utilizing internal parallelism when applications have light trafic to the device DRAM under the current architecture. Yangwook Kang, Yang-Suk Kee, Ethan L. Miller, Chanik Park |
MSST | 3 |
| 2012 | Evolutionary Trends in a Supercomputing Tertiary Storage EnvironmentabstractTracking archival usage and data migration in a long term supercomputing system is critical to understanding not only how users' needs and habits have changed over time, but also how the archive itself evolves in response to these external factors. Yet this type of study has not previously been performed. To address this need, we conducted an in-depth comparison of user initiated file activity on the mass storage system (MSS) at the National Center for Atmospheric Research (NCAR) during two periods, one in the early 1990s, and another nearly twenty years later. In addition to confirming earlier findings, our analysis turned up three surprising results. First, the read: write ratio went from 2:1 in the earlier trace to 1:2 in the later trace, a reduction of a factor of four in reads relative to writes. Second, only 30% of the current archive was accessed during the three year period of the study, in stark contrast to the 80% seen in the 1992 trace analysis. Third, access latency to the first byte of data actually got slower despite much faster computers and storage devices. These findings indicate that archival behavior has shifted towards a write-heavy workload, and that future archives can be more optimized for write activity than previously believed. Furthermore it may be worth considering the value of data being archived when it is stored, since later retrieval is increasingly less likely. Joel Cameron Frank, Ethan L. Miller, Ian F. Adams, Daniel C. Rosenthal |
MASCOTS | 2 |
| 2012 | Emulating a Shingled Write DiskabstractShingled Magnetic Recording technology is expected to play a major role in the next generation of hard disk drives. But it introduces some unique challenges to system software researchers and prototype hardware is not readily available for the broader research community. It is crucial to work on system software in parallel to hardware manufacturing, to ensure successful and effective adoption of this technology. In this work, we present a novel Shingled Write Disk (SWD) emulator that uses a hard disk utilizing traditional Perpendicular Magnetic Recording (PMR) and emulates a Shingled Write Disk on top of it. We implemented the emulator as a pseudo block device driver and evaluated the performance overhead incurred by employing the emulator. The emulator has a slight overhead which is only measurable during pure sequential reads and writes. The moment disk head movement comes into picture, due to any random access, the emulator overhead becomes so insignificant as to become immeasurable. Rekha Pitchumani, Andy Hospodor, Ahmed Amer, Yangwook Kang, Ethan L. Miller, Darrell D. E. Long |
MASCOTS | 5 |
| 2012 | Usage behavior of a large-scale scientific archiveabstractArchival storage systems for scientific data have been growing in both size and relevance over the past two decades, yet researchers and system designers alike must rely on limited and obsolete knowledge to guide archival management and design. To address this issue, we analyzed three years of filelevel activities from the NCAR mass storage system, providing valuable insight into a large-scale scientific archive with over 1600 users, tens of millions of files, and petabytes of data. Our examination of system usage showed that, while a subset of users were responsible for most of the activity, this activity was widely distributed at the file level. We also show that the physical grouping of files and directories on media can improve archival storage system performance. Based on our observations, we provide suggestions and guidance for both future scientific archival system designs as well as improved tracing of archival activity. Ian F. Adams, Brian A. Madden, Joel Cameron Frank, Mark W. Storer, Ethan L. Miller, Gene Harano |
SC | 5 |
| 2012 | Understanding data survivability in archival storage systemsabstractPreserving data for a long period of time in the face of faults, large and small, is crucial for designing reliable archival storage systems. However, the survivability of data is different from the reliability of storage because typically, data are stored in more than one storage at a given moment. Previous studies of reliability ignore the former. We present a framework for relating data survivability and storage reliability, and use the framework to gauge the impact of rare but large-scale events on data survivability. We also present a method to track all copies of data and the condition of all the online and offline media, devices and systems on which they are stored uninterruptedly over the whole lifetime of the data. With this method, the survivability of the data can be closely monitored, and potential dangers can be handled in a timely manner. A better understanding of data survivability can be used in reducing unnecessary data replicas, thus reducing the cost. Yan Li 0006, Ethan L. Miller, Darrell D. E. Long |
SYSTOR | 2 |
| 2012 | Analysis of Workload Behavior in Scientific and Historical Long-Term Data RepositoriesabstractThe scope of archival systems is expanding beyond cheap tertiary storage: scientific and medical data is increasingly digital, and the public has a growing desire to digitally record their personal histories. Driven by the increase in cost efficiency of hard drives, and the rise of the Internet, content archives have become a means of providing the public with fast, cheap access to long-term data. Unfortunately, designers of purpose-built archival systems are either forced to rely on workload behavior obtained from a narrow, anachronistic view of archives as simply cheap tertiary storage, or extrapolate from marginally related enterprise workload data and traditional library access patterns. To close this knowledge gap and provide relevant input for the design of effective long-term data storage systems, we studied the workload behavior of several systems within this expanded archival storage space. Our study examined several scientific and historical archives, covering a mixture of purposes, media types, and access models---that is, public versus private. Our findings show that, for more traditional private scientific archival storage, files have become larger, but update rates have remained largely unchanged. However, in the public content archives we observed, we saw behavior that diverges from the traditional “write-once, read-maybe” behavior of tertiary storage. Our study shows that the majority of such data is modified---sometimes unnecessarily---relatively frequently, and that indexing services such as Google and internal data management processes may routinely access large portions of an archive, accounting for most of the accesses. Based on these observations, we identify areas for improving the efficiency and performance of archival storage systems. Ian F. Adams, Mark W. Storer, Ethan L. Miller |
ACM Trans. Storage | 3 |
| 2011 | Reliable and randomized data distribution strategies for large scale storage systemsabstractThe ever-growing amount of data requires highly scalable storage solutions. The most flexible approach is to use storage pools that can be expanded and scaled down by adding or removing storage devices. To make this approach usable, it is necessary to provide a solution to locate data items in such a dynamic environment. This paper presents and evaluates the Random Slicing strategy, which incorporates lessons learned from table-based, rule-based, and pseudo-randomized hashing strategies and is able to provide a simple and efficient strategy that scales up to handle exascale data. Random Slicing keeps a small table with information about previous storage system insert and remove operations, drastically reducing the required amount of randomness while delivering a perfect load distribution. Alberto Miranda, Sascha Effert, Yangwook Kang, Ethan L. Miller, André Brinkmann, Toni Cortes |
HiPC | 4 |
| 2011 | Object-based SCM: An efficient interface for Storage Class MemoriesabstractStorage Class Memory (SCM) has become increasingly popular in enterprise systems as well as embedded and mobile systems. However, replacing hard drives with SCMs in current storage systems often forces either major changes in file systems or suboptimal performance, because the current block-based interface does not deliver enough information to the device to allow it to optimize data management for specific device characteristics such as the out-of-place update. To alleviate this problem and fully utilize different characteristics of SCMs, we propose the use of an object-based model that provides the hardware and firmware the ability to optimize performance for the underlying implementation, and allows drop-in replacement for devices based on new types of SCM. We discuss the design of object-based SCMs and implement an object-based flash memory prototype. By analyzing different design choices for several subsystems, such as data placement policies and index structures, we show that our object-based model provides comparable performance to other flash file systems while enabling advanced features such as object-level reliability. Yangwook Kang, Jingpei Yang, Ethan L. Miller |
MSST | 3 |
| 2011 | Efficiently identifying working sets in block I/O streamsabstractIdentifying groups of blocks that tend to be read or written together in a given environment is the first step towards powerful techniques for device failure isolation and power management. For example, identified groups can be placed together on a single disk, avoiding excess drive activity across an exascale storage system. Unlike previous grouping work, we focus on identifying groupings in data that can be gathered from real, running systems with minimal impact. Using temporal, spatial, and access ordering information from an enterprise data set, we identified a set of groupings that consistently appear, indicating that these are working sets that are likely to be accessed together. We present several techniques to obtain groupings along with a discussion of what techniques best apply to particular types of real systems. We intend to use these preliminary results to inform our search for new types of workloads with a goal of identifying properties of easily separable workloads across different systems and dynamically moving groups in these workloads to reduce disk activity in large storage systems. Avani Wildani, Ethan L. Miller, Lee Ward |
SYSTOR | 2 |
| 2010 | Examining Energy Use in Heterogeneous Archival Storage SystemsabstractControlling energy usage in data centers, and storage in particular, continues to rise in importance. Many systems and models have examined energy efficiency through intelligent spin-down of disks and novel data layouts, yet little work has been done to examine how power usage over the course of months to years is impacted by the characteristics of the storage devices chosen for use. Long-term power usage is particularly important for archival storage systems, since it is a large contributor to overall system cost. In this work, we begin exploring the impact that broad policies (e.g. utilize high-bandwidth devices first) have upon the power efficiency of a disk based archival storage system of heterogeneous devices over the course of a year. Using a discrete event simulator, we found that even simple heuristic policies for allocating space can have significant impact on the power usage of a system. We show that our system growth policies can cause power usage to vary from 10% higher to 18% lower than a naive random data allocation scheme. We also found that under low read rates power is dominated by that used in standby modes. Most interestingly, we found cases where concentrating data on fewer devices yielded increased power usage. Ian F. Adams, Ethan L. Miller, Mark W. Storer |
MASCOTS | 2 |
| 2010 | Efficient Storage Management for Object-based Flash MemoryabstractFlash memory has become increasingly popular in today's storage systems. However, replacing hard drives with flash memory in current systems often either requires major file system changes or causes performance degradation due to the limitations of block-based interface and out-of-place updates required by flash. To alleviate this problem, we propose an object-based model for flash memory that gives the hardware and firmware the ability to optimize performance for the underlying implementation. Based on this model, we propose two new data placement policies that exploit richer information from an object-based interface. Using simulation, we show that cleaning overhead can be reduced by up to 9% by separating data and metadata. Segregating the access time from metadata can further reduce the cleaning overhead by up to 23%. Yangwook Kang, Jingpei Yang, Ethan L. Miller |
MASCOTS | 3 |
| 2010 | Design issues for a shingled write disk systemabstractIf the data density of magnetic disks is to continue its current 30-50% annual growth, new recording techniques are required. Among the actively considered options, shingled writing is currently the most attractive one because it is the easiest to implement at the device level. Shingled write recording trades the inconvenience of the inability to update in-place for a much higher data density by a using a different write technique that overlaps the currently written track with the previous track. Random reads are still possible on such devices, but writes must be done largely sequentially. In this paper, we discuss possible changes to disk-based data structures that the adoption of shingled writing will require. We first explore disk structures that are optimized for large sequential writes with little or no sequential writing, even of metadata structures, while providing acceptable read performance. We also examine the usefulness of non-volatile RAM and the benefits of object-based interfaces in the context of shingled disks. Finally, through the analysis of recent device traces, we demonstrate the surprising stability of written device blocks, with general purpose workloads showing that more than 93% of device blocks remain unchanged over a day, and that for more specialized workloads less than 0.5% of a shingled-write disk's capacity would be needed to hold randomly updated blocks. Ahmed Amer, Darrell D. E. Long, Ethan L. Miller, Jehan-François Pâris, Thomas J. E. Schwarz |
MSST | 3 |
| 2010 | Security Aware Partitioning for efficient file system searchabstractIndex partitioning techniques-where indexes are broken into multiple distinct sub-indexes-are a proven way to improve metadata search speeds and scalability for large file systems, permitting early triage of the file system. A partitioned metadata index can rule out irrelevant files and quickly focus on files that are more likely to match the search criteria. Also, in a large file system that contains many users, a user's search should not include confidential files the user doesn't have permission to view. To meet these two parallel goals, we propose a new partitioning algorithm, Security Aware Partitioning, that integrates security with the partitioning method to enable efficient and secure file system search. In order to evaluate our claim of improved efficiency, we compare the results of Security Aware Partitioning to six other partitioning methods, including implementations of the metadata partitioning algorithms of SmartStore and Spyglass, two recent systems doing partitioned search in similar environments. We propose a general set of criteria for comparing partitioning algorithms, and use them to evaluate the partitioning algorithms. Our results show that Security Aware Partitioning can provide excellent search performance at a low computational cost to build indexes, O(n). Based on metrics such as information gain, we also conclude that expensive clustering algorithms do not offer enough benefit to make them worth the additional cost in time and memory. Aleatha Parker-Wood, Christina E. Strong, Ethan L. Miller, Darrell D. E. Long |
MSST | 3 |
| 2009 | Adding aggressive error correction to a high-performance compressing flash file systemabstractWhile NAND flash memories have rapidly increased in both capacity and performance and are increasingly used as a storage device in many embedded systems, their reliability has decreased both because of increased density and the use of multi-level cells (MLC). Current MLC technology only specifies the minimum requirement for an error correcting code (ECC), but provides no additional protection in hardware. However, existing flash file systems such as YAFFS and JFFS2 rely upon ECC to survive small numbers of bit errors, but cannot survive the larger numbers of bit errors or page failures that are becoming increasingly common as flash file systems scale to multiple gigabytes. Yangwook Kang, Ethan L. Miller |
EMSOFT | 2 |
| 2009 | Spyglass: Fast, Scalable Metadata Search for Large-Scale Storage Systems
Andrew W. Leung, Minglong Shao, Timothy Bisson, Shankar Pasupathy, Ethan L. Miller |
FAST | 5 |
| 2009 | Protecting against rare event failures in archival systemsabstractDigital archives are growing rapidly, necessitating stronger reliability measures than RAID to avoid data loss from device failure. Mirroring, a popular solution, is too expensive over time. We present a compromise solution that uses multi-level redundancy coding to reduce the probability of data loss from multiple simultaneous device failures. This approach handles small-scale failures of one or two devices efficiently while still allowing the system to survive rare-event, larger-scale failures of four or more devices. In our approach, each disk is split into a set of fixed size disklets which are used to construct reliability stripes. To protect against rare event failures, reliability stripes are grouped into larger super-groups, each of which has a corresponding super-parity; super-parity is only used to recover data when disk failures overwhelm the redundancy in a single reliability stripe. Super-parity can be stored on a variety of devices such as NV-RAM and always-on disks to offset write bottlenecks while still keeping the number of active devices low. Our calculations of failure probabilities show that adding super-parity allows our system to absorb many more disk failures without data loss. Through discrete event simulation, we found that adding super-groups has a significant impact on mean time to data loss and that rebuilds are slow but not unmanageable. Finally, we showed that robustness against rare events can be achieved for a fraction of total system cost. Avani Wildani, Thomas J. E. Schwarz, Ethan L. Miller, Darrell D. E. Long |
MASCOTS | 3 |
| 2009 | The effectiveness of deduplication on virtual machine disk imagesabstractVirtualization is becoming widely deployed in servers to efficiently provide many logically separate execution environments while reducing the need for physical servers. While this approach saves physical CPU resources, it still consumes large amounts of storage because each virtual machine (VM) instance requires its own multi-gigabyte disk image. Moreover, existing systems do not support ad hoc block sharing between disk images, instead relying on techniques such as overlays to build multiple VMs from a single "base" image. Keren Jin, Ethan L. Miller |
SYSTOR | 2 |
| 2009 | POTSHARDS - a secure, recoverable, long-term archival storage systemabstractUsers are storing ever-increasing amounts of information digitally, driven by many factors including government regulations and the public's desire to digitally record their personal histories. Unfortunately, many of the security mechanisms that modern systems rely upon, such as encryption, are poorly suited for storing data for indefinitely long periods of time; it is very difficult to manage keys and update cryptosystems to provide secrecy through encryption over periods of decades. Worse, an adversary who can compromise an archive need only wait for cryptanalysis techniques to catch up to the encryption algorithm used at the time of the compromise in order to obtain “secure” data. To address these concerns, we have developed POTSHARDS, an archival storage system that provides long-term security for data with very long lifetimes without using encryption. Secrecy is achieved by using unconditionally secure secret splitting and spreading the resulting shares across separately managed archives. Providing availability and data recovery in such a system can be difficult; thus, we use a new technique, approximate pointers, in conjunction with secure distributed RAID techniques to provide availability and reliability across independent archives. To validate our design, we developed a prototype POTSHARDS implementation. In addition to providing us with an experimental testbed, this prototype helped us to understand the design issues that must be addressed in order to maximize security. Mark W. Storer, Kevin M. Greenan, Ethan L. Miller, Kaladhar Voruganti |
ACM Trans. Storage | 3 |
| 2008 | Reliability of flat XOR-based erasure codes on heterogeneous devicesabstractXOR-based erasure codes are a computationally-efficient means of generating redundancy in storage systems. Some such erasure codes provide irregular fault tolerance: some subsets of failed storage devices of a given size lead to data loss, whereas other subsets of failed storage devices of the same size are tolerated. Many storage systems are composed of heterogeneous devices that exhibit different failure and recovery rates, in which different placements- mappings of erasure-coded symbols to storage devices-of a flat XOR-based erasure code lead to different reliabilities. We have developed redundancy placement algorithms that utilize the structure of flat XOR-based erasure codes and a simple analytic model to determine placements that maximize reliability. Simulation studies validate the utility of the simple analytic reliability model and the efficacy of the redundancy placement algorithms. Kevin M. Greenan, Ethan L. Miller, Jay J. Wylie |
DSN | 2 |
| 2008 | Workload-based configuration of MEMS-based storage devices for mobile systemsabstractBecause of its small form factor, high capacity, and expected low cost, MEMS-based storage is a suitable storage technology for mobile systems. However, flash memory may outperform MEMS-based storage in terms of performance, and energy-efficiency. The problem is that MEMS-based storage devices have a large number (i.e., thousands) of heads, and to deliver peak performance, all heads must be deployed simultaneously to access each single sector. Since these devices are mechanical and thus some housekeeping information is needed for each head, this results in a huge capacity loss and increases the energy consumption of MEMS-based storage with respect to flash. Mohammed G. Khatib, Ethan L. Miller, Pieter H. Hartel |
EMSOFT | 2 |
| 2008 | Pergamum: Replacing Tape with Energy Efficient, Reliable, Disk-Based Archival Storage
Mark W. Storer, Kevin M. Greenan, Ethan L. Miller, Kaladhar Voruganti |
FAST | 3 |
| 2008 | Optimizing Galois Field Arithmetic for Diverse Processor Architectures and Applications
Kevin M. Greenan, Ethan L. Miller, Thomas J. E. Schwarz |
MASCOTS | 2 |
| 2008 | Measurement and Analysis of Large-Scale Network File System Workloads
Andrew W. Leung, Shankar Pasupathy, Garth R. Goodson, Ethan L. Miller |
USENIX ATC | 4 |
| 2007 | Scalable security for petascale parallel file systemsabstractPetascale, high-performance file systems often hold sensitive data and thus require security, but authentication and authorization can dramatically reduce performance. Existing security solutions perform poorly in these environments because they cannot scale with the number of nodes, highly distributed data, and demanding workloads. To address these issues, we developed Maat, a security protocol designed to provide strong, scalable security to these systems. Maat introduces three new techniques. Extended capabilities limit the number of capabilities needed by allowing a capability to authorize I/O for any number of client-file pairs. Automatic Revocation uses short capability lifetimes to allow capability expiration to act as global revocation, while supporting non-revoked capability renewal. Secure Delegation allows clients to securely act on behalf of a group to open files and distribute access, facilitating secure joint computations. Experiments on the Maat prototype in the Ceph petascale file system show an overhead as little as 6--7%. Andrew W. Leung, Ethan L. Miller, Stephanie N. Jones |
SC | 2 |
| 2007 | POTSHARDS: Secure Long-Term Storage Without Encryption
Mark W. Storer, Kevin M. Greenan, Ethan L. Miller, Kaladhar Voruganti |
USENIX ATC | 3 |
| 2006 | Reliability mechanisms for file systems using non-volatile memory as a metadata storeabstractPortable systems such as cell phones and portable media players commonly use non-volatile RAM (NVRAM) to hold all of their data and metadata, and larger systems can store metadata in NVRAM to increase file system performance by reducing synchronization and transfer overhead between disk and memory data structures. Unfortunately, wayward writes from buggy software and random bit flips may result in an unreliable persistent store. We introduce two orthogonal and complementary approaches to reliably storing file system structures in NVRAM. First, we reinforce hardware and operating system memory consistency by employing page-level write protection and error correcting codes. Second, we perform on-line consistency checking of the filesystem structures by replaying logged file system transactions on copied data structures; a structure is consistent if the replayed copy matches its live counterpart. Our experiments show that the protection mechanisms can increase fault tolerance by six orders of magnitude while incurring an acceptable amount of overhead on writes to NVRAM. Since NVRAM is much faster and consumes far less power than disk-based storage, the added overhead of error checking leaves an NVRAM-based system both faster and more reliable than a disk-based system. Additionally, our techniques can be implemented on systems lacking hardware support for memory management, allowing them to be used on lowend and embedded systems without an MMU. Kevin M. Greenan, Ethan L. Miller |
EMSOFT | 2 |
| 2006 | Store, Forget, and Check: Using Algebraic Signatures to Check Remotely Administered StorageabstractThe emerging use of the Internet for remote storage and backup has led to the problem of verifying that storage sites in a distributed system indeed store the data; this must often be done in the absence of knowledge of what the data should be. We use m/n erasure-correcting coding to safeguard the stored data and use algebraic signatures hash functions with algebraic properties for verification. Our scheme primarily utilizes one such algebraic property: taking a signature of parity gives the same result as taking the parity of the signatures. To make our scheme collusionresistant, we blind data and parity by XORing them with a pseudo-random stream. Our scheme has three advantages over existing techniques. First, it uses only small messages for verification, an attractive property in a P2P setting where the storing peers often only have a small upstream pipe. Second, it allows verification of challenges across random data without the need for the challenger to compare against the original data. Third, it is highly resistant to coordinated attempts to undetectably modify data. These signature techniques are very fast, running at tens to hundreds of megabytes per second. Because of these properties, the use of algebraic signatures will permit the construction of large-scale distributed storage systems in which large amounts of storage can be verified with minimal network bandwidth. Thomas J. E. Schwarz, Ethan L. Miller |
ICDCS | 2 |
| 2006 | Providing High Reliability in a Minimum Redundancy Archival Storage SystemabstractInter-file compression techniques store files as sets of references to data objects or chunks that can be shared among many files. While these techniques can achieve much better compression ratios than conventional intra-file compression methods such as Lempel-Ziv compression, they also reduce the reliability of the storage system because the loss of a few critical chunks can lead to the loss of many files. We show how to eliminate this problem by choosing for each chunk a replication level that is a function of the amount of data that would be lost if that chunk were lost. Experiments using actual archival data show that our technique can achieve significantly higher robustness than a conventional approach combining data mirroring and intra-file compression while requiring about half the storage space. Deepavali Bhagwat, Kristal T. Pollack, Darrell D. E. Long, Thomas J. E. Schwarz, Ethan L. Miller, Jehan-François Pâris |
MASCOTS | 5 |
| 2006 | Ceph: A Scalable, High-Performance Distributed File System
Sage A. Weil, Scott A. Brandt, Ethan L. Miller, Darrell D. E. Long, Carlos Maltzahn |
OSDI | 3 |
| 2006 | Grid resource management - CRUSH: controlled, scalable, decentralized placement of replicated dataabstractEmerging large-scale distributed storage systems are faced with the task of distributing petabytes of data among tens or hundreds of thousands of storage devices. Such systems must evenly distribute data and workload to efficiently utilize available resources and maximize system performance, while facilitating system growth and managing hardware failures. We have developed CRUSH, a scalable pseudorandom data distribution function designed for distributed object-based storage systems that efficiently maps data objects to storage devices without relying on a central directory. Because large systems are inherently dynamic, CRUSH is designed to facilitate the addition and removal of storage while minimizing unnecessary data movement. The algorithm accommodates a wide variety of data replication and reliability mechanisms and distributes data in terms of user-defined policies that enforce separation of replicas across failure domains. Sage A. Weil, Scott A. Brandt, Ethan L. Miller, Carlos Maltzahn |
SC | 3 |
| 2006 | Using MEMS-based storage in computer systems - device modeling and managementabstractMEMS-based storage is an emerging nonvolatile secondary storage technology. It promises high performance, high storage density, and low power consumption. With fundamentally different architectural designs from magnetic disk, MEMS-based storage exhibits unique two-dimensional positioning behaviors and efficient power state transitions. We model these low-level, device-specific properties of MEMS-based storage and present request scheduling algorithms and power management strategies that exploit the full potential of these devices. Our simulations show that MEMS-specific device management policies can significantly improve system performance and reduce power consumption. Scott A. Brandt, Darrell D. E. Long, Ethan L. Miller |
ACM Trans. Storage | 4 |
| 2005 | Disk Infant Mortality in Large Storage SystemsabstractAs disk drives have dropped in price relative to tape, the desire for the convenience and speed of online access to large data repositories has, led to the deployment of petabyte-scale disk farms with thousands of disks. Unfortunately, the very large size of these repositories renders them vulnerable to previously rare failure modes such as multiple, unrelated disk failures leading to data loss. While some business models, such as free email servers, may be able to tolerate some occurrence of data loss, others, including premium online services and storage of simulation results at a national laboratory, cannot. This paper describes the effect of infant mortality on long-term failure rates of systems that must preserve their data for decades. Our failure models incorporate the well-known "bathtub curve," which reflects the higher failure rates of new disk drives, a lower, constant failure rate during the remainder of the design life span, and increased failure rates as components wear out. Large systems are vulnerable to the "cohort effect" that occurs when many disks are simultaneously replaced by new disks. Our more accurate disk models and simulations have yielded predictions of system lifetimes that are more pessimistic than existing models that assume a constant disk failure rate. Thus, larger system scale requires designers to take disk infant mortality into account. Qin Xin 0005, Thomas J. E. Schwarz, Ethan L. Miller |
MASCOTS | 3 |
| 2005 | Richer File System Metadata Using Links and AttributesabstractTraditional file systems provide a weak and inadequate structure for meaningful representations of file interrelationships and other context-providing metadata. Existing designs, which store additional file-oriented metadata either in a database, on disk, or both are limited by the technologies upon which they depend. Moreover, they do not provide for user-defined relationships among files. To address these issues, we created the linking file system (LiFS), a file system design in which files may have both arbitrary user- or application-specified attributes, and attributed links between files. In order to assure performance when accessing links and attributes, the system is designed to store metadata in non-volatile memory. This paper discusses several use cases that take advantage of this approach and describes the user-space prototype we developed to test the concepts presented. Alexander Ames, Carlos Maltzahn, Nikhil Bobb, Ethan L. Miller, Scott A. Brandt, Alisa Neeman, Adam Hiatt, Deepa Tuteja |
MSST | 4 |
| 2005 | Impact of Failure on Interconnection Networks for Large Storage SystemsabstractRecent advances in large-capacity, low-cost storage devices have led to active research in design of large-scale storage systems built from commodity devices for supercomputing applications. Such storage systems, composed of thousands of storage devices, are required to provide high system bandwidth and petabyte-scale data storage. A robust network interconnection is essential to achieve high bandwidth, low latency, and reliable delivery during data transfers. However, failures, such as temporary link outages and node crashes, are inevitable. We discuss the impact of potential failures on network interconnections in very large-scale storage systems and analyze the trade-offs among several storage network topologies by simulations. Our results suggest that a good interconnect topology be essential to fault-tolerance of a petabyte-scale storage system. Qin Xin 0005, Ethan L. Miller, Thomas J. E. Schwarz, Darrell D. E. Long |
MSST | 2 |
| 2004 | Evaluation of Distributed Recovery in Large-Scale Storage Systems
Qin Xin 0005, Ethan L. Miller, Thomas J. E. Schwarz |
HPDC | 2 |
| 2004 | Caching Support for Push-Pull Data Dissemination using Data-Snooping Routers
Ismail Ari, Ethan L. Miller |
ICPADS | 2 |
| 2004 | Replication Under Scalable Hashing: A Family of Algorithms for Scalable Decentralized Data DistributionabstractSummary form only given. Typical algorithms for decentralized data distribution work best in a system that is fully built before it first used; adding or removing components results in either extensive reorganization of data or load imbalance in the system. We have developed a family of decentralized algorithms, RUSH (replication under scalable hashing), that maps replicated objects to a scalable collection of storage servers or disks. RUSH algorithms distribute objects to servers according to user-specified server weighting. While all RUSH variants support addition of servers to the system, different variants have different characteristics with respect to lookup time in petabyte-scale systems, performance with mirroring (as opposed to redundancy codes), and storage server removal. All RUSH variants redistribute as few objects as possible when new servers are added or existing servers are removed, and all variants guarantee that no two replicas of a particular object are ever placed on the same server. Because there is no central directory, clients can compute data locations in parallel, allowing thousands of clients to access objects on thousands of servers simultaneously. R. J. Honicky, Ethan L. Miller |
IPDPS | 2 |
| 2004 | Interconnection Architectures for High-Performance Object-Based File Systems
Andy Hospodor, Ethan L. Miller |
MSST | 2 |
| 2004 | OBFS: A File System for Object-Based Storage Devices
Feng Wang 0003, Scott A. Brandt, Ethan L. Miller, Darrell D. E. Long |
MSST | 3 |
| 2004 | File System Workload Analysis For Large Scientific Computing Applications
Feng Wang 0003, Qin Xin 0005, Scott A. Brandt, Ethan L. Miller, Darrell D. E. Long, Tyce T. McLarty |
MSST | 5 |
| 2004 | Dynamic Metadata Management for Petabyte-Scale File SystemsabstractIn petabyte-scale distributed file systems that decouple read and write from metadata operations, behavior of the metadata server cluster will be critical to overall system performance and scalability. We present a dynamic subtree partitioning and adaptive metadata management system designed to efficiently manage hierarchical metadata workloads that evolve over time. We examine the relative merits of our approach in the context of traditional workload partitioning strategies, and demonstrate the performance, scalability and adaptability advantages in a simulation environment. Sage A. Weil, Kristal T. Pollack, Scott A. Brandt, Ethan L. Miller |
SC | 4 |
| 2002 | Strong Security for Network-Attached Storage
Ethan L. Miller, Darrell D. E. Long, William E. Freeman, Benjamin C. Reed |
FAST | 1 |
| 2001 | HeRMES: High-Performance Reliable MRAM-Enabled StorageabstractMagnetic RAM (MRAM) is a new memory technology with access and cost characteristics comparable to those of conventional dynamic RAM (DRAM) and the non-volatility of magnetic media such as disk. Simply replacing DRAM with MRAM will make main memory non-volatile, but it will not improve file system performance. However, effective use of MRAM in a file system has the potential to significantly improve performance over existing file systems. The HeRMES file system will use MRAM to dramatically improve file system performance by using it as a permanent store for both file system data and metadata. In particular, metadata operations, which make up over 50% of all file system requests [14], are nearly free in HeRMES because they do not require any disk accesses. Data requests will also be faster, both because of increased metadata request speed and because using MRAM as a non-volatile cache will allow HeRMES to better optimize data placement on disk. Though MRAM capacity is too small to replace disk entirely, HeRMES will use MRAM to provide high-speed access to relatively small units of data and metadata, leaving most file data stored on disk. Ethan L. Miller, Scott A. Brandt, Darrell D. E. Long |
HotOS | 1 |
| 1999 | An Experimental Analysis of Cryptographic Overhead in Performance-Critical SystemsabstractThis paper studies the performance implications of using cryptographic controls in performance-critical systems. Full cryptographic controls beyond basic authentication are considered and experimentally validated in the concept of network file systems. This paper demonstrates that processor speeds have become fast enough to support cryptographic controls in many performance-critical systems. Integrity and authentication using keyed-hash and RSA as well as confidentiality using RC5 are tested. This analysis demonstrates that full cryptographic controls are feasible in a distributed network file system, by showing the performance overhead for including signature, hash and encryption algorithms on various embedded and workstation computers. The results from these experiments are used to predict the performance impact using three proposed network disk security schemes. William E. Freeman, Ethan L. Miller |
MASCOTS | 2 |
| 1998 | Self-Similarity in File SystemsabstractWe demonstrate that high-level file system events exhibit self-similar behaviour, but only for short-term time scales of approximately under a day. We do so through the analysis of four sets of traces that span time scales of milliseconds through months, and that differ in the trace collection method, the filesystems being traced, and the chronological times of the tracing. Two sets of detailed, short-term file system trace data are analyzed; both are shown to have self-similar like behaviour, with consistent Hurst parameters (a measure of self-similarity) for all file system traffic as well as individual classes of file system events. Long-term file system trace data is then analyzed, and we discover that the traces' high variability and self-similar behaviour does not persist across time scales of days, weeks, and months. Using the short-term trace data, we show that sources of file system traffic exhibit ON/OFF source behaviour, which is characterized by highly variably lengthed bursts of activity, followed by similarly variably lengthed periods of inactivity. This ON/OFF behaviour is used to motivate a simple technique for synthesizing a stream of events that exhibit the same self-similar short-term behaviour as was observed in the file system traces. Steve D. Gribble, Gurmeet Singh Manku, Drew S. Roselli, Eric A. Brewer, Timothy J. Gibson, Ethan L. Miller |
SIGMETRICS | 6 |
| 1998 | Performance Measurements of Tertiary Storage Devices
Theodore Johnson, Ethan L. Miller |
VLDB | 2 |
| 1997 | Interactive Volumetric Information Visualization
David S. Ebert, Chris Shaw 0002, Amen Zwa, Ethan L. Miller, D. Aaron Roberts |
Graphics Interface | 4 |
| 1997 | RAMA: An Easy-to-Use, High-Performance Parallel File System
Ethan L. Miller, Randy H. Katz |
Parallel Comput. | 1 |
| 1995 | RAMA: Easy Access to a High-Bandwidth Massively Parallel File System
Ethan L. Miller, Randy H. Katz |
USENIX | 1 |
| 1994 | RAID-II: A High-Bandwidth Network File ServerabstractIn 1989, the RAID (Redundant Arrays of Inexpensive Disks) group at U.C. Berkeley built a prototype disk array called RAID-I. The bandwidth delivered to clients by RAID-I was severely limited by the memory system bandwidth of the disk array's host workstation. They designed their second prototype, RAID-II, to deliver more of the disk array bandwidth to file server clients. A custom-built crossbar memory system called the XBUS board connects the disks directly to the high-speed network, allowing data for large requests to bypass the server workstation. RAID-II runs Log-Structured File System (LFS) software to optimize performance for bandwidth-intensive applications. The RAID-II hardware with a single XBUS controller board delivers 20 megabytes/second for large, random read operations and up to 31 megabytes/second for sequential read operations. A preliminary implementation of LFS on RAID-II delivers 21 megabytes/second on large read requests and 15 megabytes/second on large write operations.> Ann L. Drapeau, Ken Shirriff, John H. Hartman, Ethan L. Miller, Srinivasan Seshan, Randy H. Katz, Ken Lutz, David A. Patterson 0001, Edward K. Lee 0001, Peter M. Chen, Garth A. Gibson |
ISCA | 4 |
| 1994 | Performance and Design Evaluation of the RAID-II Storage Server
Peter M. Chen, Edward K. Lee 0001, Ann L. Drapeau, Ken Lutz, Ethan L. Miller, Srinivasan Seshan, Ken Shirriff, David A. Patterson 0001, Randy H. Katz |
Distributed Parallel Databases | 5 |
| 1991 | Input/output behavior of supercomputing applicationsabstractThe collection and analysis of supercomputer I/O traces and their use in a collection of buffering and caching simulations are described. This serves two purposes. First, it gives a model of how individual applications running on supercomputers request file system I/O, allowing system designer to optimize I/O hardware and file system algorithms to that model. Second, the buffering simulations show what resources are needed to maximize the CPU utilization of a supercomputer given a very bursty I/O request rate. By using read-ahead and write-behind in a large solid stated disk, one or two applications were sufficient to fully utilize a Cray Y-MP CPU. Ethan L. Miller, Randy H. Katz |
SC | 1 |