EDBT 2026 Demo / reviewers in the wild / expert
David Pease
dblp:85/6115
· DBLP profile ↗
18ranked-venue papers
1as first author
1since 2021 · last 2025
0009-0005-8716-3770ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Storage systems · 90% Performance modeling and evaluation · 7% Processor architecture and microarchitecture · 2% |
Topics — the 12 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › magnetic storage
tape storage |
1.2 | 3 | 2025 | Magnetic Tape Storage Technology · ACM Trans. Storage 2025 Seamlessly integrating disk and tape in a multi-tiered distributed file system · ICDE 2015 File-based media workflows using ltfs tapes · ACM Multimedia 2010 |
Storage systems › archival storage
tape libraries |
0.9 | 1 | 2025 | Magnetic Tape Storage Technology · ACM Trans. Storage 2025 |
Storage systems
distributed storage |
0.2 | 1 | 2015 | A Tale of Two Erasure Codes in HDFS · FAST 2015 |
Storage systems › storage reliability
erasure coding |
0.2 | 1 | 2015 | A Tale of Two Erasure Codes in HDFS · FAST 2015 |
Storage systems
file systems |
0.2 | 1 | 2015 | Seamlessly integrating disk and tape in a multi-tiered distributed file system · ICDE 2015 |
Storage systems › file systems › distributed file system
HDFS |
0.2 | 1 | 2015 | A Tale of Two Erasure Codes in HDFS · FAST 2015 |
Storage systems
storage reliability |
0.2 | 1 | 2015 | A Tale of Two Erasure Codes in HDFS · FAST 2015 |
Hardware accelerators and domain-specific architectures
bioinformatics accelerator |
0.1 | 1 | 2005 | The UCSC Kestrel Parallel Processor · IEEE Trans. Parallel Distributed Syst. 2005 |
Processor architecture and microarchitecture › SIMD
SIMD processor |
0.1 | 1 | 2005 | The UCSC Kestrel Parallel Processor · IEEE Trans. Parallel Distributed Syst. 2005 |
Storage systems › storage performance
storage quality of service |
0.0 | 1 | 2004 | Polus: Growing Storage QoS Management Beyond a "4-Year Old Kid" · FAST 2004 |
Processor architecture and microarchitecture
SIMD |
0.0 | 1 | 2005 | The UCSC Kestrel Parallel Processor · IEEE Trans. Parallel Distributed Syst. 2005 |
Parallel and multicore computing › data parallelism
SIMD vectorization |
0.0 | 1 | 2005 | The UCSC Kestrel Parallel Processor · IEEE Trans. Parallel Distributed Syst. 2005 |
Methods — techniques the papers use, named apart from their topics
timing-based servo · 0.9magnetic recording · 0.9POSIX file system interface · 0.2LTFS · 0.2performance analysis · 0.1architectural design · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Magnetic Tape Storage TechnologyabstractMagnetic tape provides a cost-effective way to retain the exponentially increasing volumes of data being created in recent years. The low cost per terabyte combined with tape’s low energy consumption make it an appealing option for storing infrequently accessed data and has resulted in a resurgence in use of the technology. Magnetic tape as a digital data storage technology was first commercialized in the early 1950’s and has evolved continuously since then. Despite its long history, tape has significant potential for continued capacity and data rate scaling. This article strives to provide an overview of linear magnetic tape technology, usage, history, and future outlook. After a short introduction, the article delves into the details of how modern tape drives and media operate, including the basic mechanism and physics of magnetic recording, current tape media technology, state-of-the-art tape head technology, tape layout and encoding, data retrieval, timing-based servo and mechatronics of a tape drive, and the capabilities of current drives. This is followed by a discussion of tape libraries, an overview of tape library performance modeling research, operating system-level and application-level tape support, tape use cases, and the future scaling potential and outlook of tape. The article concludes with a history of tape hardware, media, usage and software. Mark A. Lantz, Simeon Furrer, Martin Petermann, Hugo E. Rothuizen, Stella Brach, Luzius Kronig, Ilias Iliadis, Beat Weiss, Edwin R. Childers, David Pease |
ACM Trans. Storage | 10 |
| 2019 | Agni: An Efficient Dual-access File System over Object StorageabstractObject storage is a low-cost, scalable component of cloud ecosystems. However, interface incompatibilities and performance limitations inhibit its adoption for emerging cloud-based workloads. Users are compelled to either run their applications over expensive block storage-based file systems or use inefficient file connectors over object stores. Dual access, the ability to read and write the same data through file systems interfaces and object storage APIs, has promise to improve performance and eliminate storage sprawl. Kunal Lillaney, Vasily Tarasov, David Pease, Randal C. Burns |
SoCC | 3 |
| 2019 | The Case for Dual-access File Systems over Object Storage
Kunal Lillaney, Vasily Tarasov, David Pease, Randal C. Burns |
HotStorage | 3 |
| 2015 | A Tale of Two Erasure Codes in HDFS
Mingyuan Xia 0001, Mohit Saxena, Mario Blaum, David Pease |
FAST | 4 |
| 2015 | Seamlessly integrating disk and tape in a multi-tiered distributed file systemabstractThe explosion of data volumes in enterprise environments and limited budgets have triggered the need for multi-tiered storage systems. With the bulk of the data being extremely infrequently accessed, tape is a natural fit for storing such data. In this paper we present our approach to a file storage system that seamlessly integrates disk and tape, enabling a bottomless and cost-effective storage architecture that can scale to accommodate Big Data requirements. The proposed system offers access to data through a POSIX filesystem interface under a single global namespace, optimizing the placement of data across disk and tape tiers. Using a self-contained, standardized and open filesystem format on the removable tape media, the proposed system avoids dependence on proprietary software and external metadata servers to access the data stored on tape. By internally managing the tape tier resources, such as tape drives and cartridges, the system relieves the user from the burden of dealing with the complexities of tape storage. Our implementation, which is based on the GPFS and LTFS filesystems, demonstrates the applicability of the proposed architecture in real-world environments. Our experimental evaluation has shown that this is a very promising approach in terms scalability, performance and manageability. The proposed system has been productized by IBM as LTFS Enterprise Edition. Ioannis Koltsidas, Slavisa Sarafijanovic, Martin Petermann, Nils Haustein, Harald Seipp, Robert Haas 0001, Jens Jelitto, Thomas Weigold, Edwin R. Childers, David Pease, Evangelos Eleftheriou |
ICDE | 10 |
| 2015 | An I/O scheduler for dual-partitioned tapesabstractFor a long time, tapes have had a single logical partition that was efficiently operated by dedicated software in batch mode. Nowadays, after the introduction of the LTO 5 standard, tapes support logical partitioning and, with the arrival of the Linear Tape File System (LTFS), their contents are exposed through the file system interface as regular files and directories. As a consequence, tape medium can now be accessed by concurrent processes that may create files in parallel. This creates two potential problems: interleaved data blocks and expensive partition switches. This paper presents the design and implementation of Unified, an I/O scheduler for LTFS that addresses these problems. We observe the effectiveness of delayed writes on decreased file fragmentation and the reduction of partition switches with the use of buffering and of redundant file copies. We also find out that software-based read prefetching, commonly used to manage disk devices, does not improve read times on tapes, but rather introduces potential overhead. With the techniques described in this paper, the Unified scheduler allows tape operations to be performed close to the raw hardware speed. Lucas Correia Villa Real, Michael Richmond, Brian Biskeborn, David Pease |
NAS | 4 |
| 2014 | DedupT: Deduplication for tape systemsabstractDeduplication is a commonly-used technique on disk-based storage pools. However, deduplication has not been used for tape-based pools: tape characteristics, such as high mount and seek times combined with data fragmentation resulting from deduplication create a toxic combination that leads to unacceptably high retrieval times. This work proposes DedupT, a system that efficiently supports deduplication on tape pools. This paper (i) details the main challenges to enable efficient deduplication on tape libraries, (ii) presents a class of solutions based on graph-modeling of similarity between data items that enables efficient placement on tapes; and (iii) presents the design and evaluation of novel cross-tape and on-tape chunk placement algorithms that alleviate tape mount time overhead and reduce on-tape data fragmentation. Using 4.5 TB of real-world workloads, we show that DedupT retains at least 95% of the deduplication efficiency. We show that DedupT mitigates major retrieval time overheads, and, due to reading less data, is able to offer better restore performance compared to the case of restoring non-deduplicated data. Abdullah Gharaibeh, Cornel Constantinescu, Maohua Lu, Ramani Routray, Prasenjit Sarkar, David Pease, Matei Ripeanu |
MSST | 7 |
| 2014 | Taming IO Spikes in Enterprise and Campus VM DeploymentabstractEnterprises and campuses have widely employed virtualization to improve resource utilization and save capital costs. The virtualized storage infrastructure usually comprises a pool of storage nodes (disks, RAID groups and storage servers), each consolidating a number of virtual machines (VMs). The IO workload on each storage node is highly bursty where a few overloaded periods, namely IO spikes, incur lags and greatly affect VM performance. In this paper, we design VMpart, a system that automatically reconfigures VM deployment to remove IO spikes from a deployed VM environment. Firstly, VMpart collects IO parameters and identifies IO spikes for deployed VMs during production periods. Secondly, during the maintenance phase, VMpart divides every VM based on its disk partitions and reconfigures the system with a fine-grained optimized partition deployment for better load balancing. We set up a representative VM environment with desktop and server VMs used in enterprises and campuses to evaluate VMpart. Our experiments show that the optimized deployment reduces the average booting time of desktop VMs during boot storm from 9% to 28%, improves server VM throughput by more than 60% and reduces VM storage migration time by about 57%. Mingyuan Xia 0001, Pin Zhou, David Pease, Xue (Steve) Liu |
SYSTOR | 3 |
| 2013 | PDS Cloud: Long Term Digital Preservation in the CloudabstractThe emergence of the cloud and advanced object-based storage services provides opportunities to support novel models for long term preservation of digital assets. Among the benefits of this approach is leveraging the cloud's inherent scalability and redundancy to dynamically adapt to evolving needs of digital preservation. PDS Cloud is an OAIS-based preservation-aware storage service employing multiple heterogeneous cloud providers. It materializes the logical concept of a preservation information-object into physical cloud storage objects. Preserved information can be interpreted by deploying virtual appliances in the compute cloud provisioned with cloud storage data objects together with their designated rendering software. PDS Cloud has a hierarchical data model supporting independent tenants whose assets are organized in multiple aggregations based on content and value. Continuous changes to data objects, life-cycle activities, virtual appliances and cloud providers are applied in a manner transparent to the client. PDS Cloud is being developed as an infrastructure component of the European Union ENSURE project, where it is used for preservation of medical and financial data. Simona Rabinovici-Cohen, John M. Marberg, Kenneth Nagin, David Pease |
IC2E | 4 |
| 2012 | CloudDT: Efficient tape resource management using deduplication in cloud backup and archival services
Abdullah Gharaibeh, Cornel Constantinescu, Maohua Lu, Ramani Routray, Prasenjit Sarkar, David Pease, Matei Ripeanu |
CNSM | 7 |
| 2011 | Guest EditorialabstractNo abstract available. André Brinkmann, David Pease |
ACM Trans. Storage | 2 |
| 2010 | File-based media workflows using ltfs tapesabstractWhile digital video cameras have existed for over two decades digital video cassettes are still the primary storage medium in professional video archives. One of the major inhibitors in the transition to file-based workflows and media archives is the lack of an affordable, portable and archive compatible storage medium for the vast amounts of content produced. Arnon Amir, David Pease, Rainer Richter, Brian Biskeborn, Michael Richmond, Lucas Correia Villa Real |
ACM Multimedia | 2 |
| 2010 | The Linear Tape File SystemabstractWhile there are many financial and practical reasons to prefer tape storage over disk for various applications, the difficultly of using tape in a general way is a major inhibitor to its wider usage. We present a file system that takes advantage of a new generation of tape hardware to provide efficient access to tape using standard, familiar system tools and interfaces. The Linear Tape File System (LTFS) makes using tape as easy, flexible, portable, and intuitive as using other removable and sharable media, such as a USB drive. David Pease, Arnon Amir, Lucas Correia Villa Real, Brian Biskeborn, Michael Richmond, Atsushi Abe |
MSST | 1 |
| 2005 | Security vs Performance: Tradeoffs using a Trust FrameworkabstractWe present an architecture of a trust framework that can be used to intelligently tradeoff between security and performance in a SAN file system. The primary idea is to differentiate between various clients in the system based on their trustworthiness and provide them with differing levels of security and performance. Client trustworthiness reflects its expected behavior and is evaluated in an online fashion using a customizable trust model. We also describe the interface of the trust framework with an example block level security solution for an out-of-band virtualization based SAN file system (SAN FS). The proposed framework can be easily extended to provide differential treatment based on data sensitivity, using a configurable parameter of the trust model. This allows associating stringent security requirements for more sensitive data, while trading off security for better performance for less critical data, a situation regularly desired in an enterprise. Aameek Singh, Sandeep Gopisetty, Linda Duyanovich, Kaladhar Voruganti, David Pease, Ling Liu 0001 |
MSST | 5 |
| 2005 | A Hybrid Access Model for Storage Area NetworksabstractWe present HSAN & a hybrid storage area network, which uses both in-band (like NFS (R. Sandberg et al., 1985)) and out-of-band visualization (like SAN FS (J. Menon et al., 2003)) access models. HSAN uses hybrid servers that can serve as both metadata and NAS servers to intelligently decide the access model per each request, based on the characteristics of requested data. This is in contrast to existing efforts that merely provide concurrent support for both models and do not exploit model appropriateness for requested data. The HSAN hybrid model is implemented using low overhead cache-admission and cache-replacement schemes and aims to improve overall response times for a wide variety of workloads. Preliminary analysis of the hybrid model indicates performance improvements over both models. Aameek Singh, Sandeep Gopisetty, Kaladhar Voruganti, David Pease, Ling Liu 0001 |
MSST | 4 |
| 2005 | An Architecture for Lifecycle Management in Very Large File SystemsabstractWe present a policy-based architecture STEPS for lifecycle management (LCM) in a mass scale distributed file system. The STEPS architecture is designed in the context of IBM's SAN file system (SFS) and leverages the parallelism and scalability offered by SFS, while providing a centralized point of control for policy-based management. The architecture uses novel concepts like policy cache and rate-controlled migration for efficient and non-intrusive execution of the LCM functions, while ensuring that the architecture scales with very large number of files. The architecture has been implemented and used for lifecycle management in a distributed deployment of SFS with heterogeneous data. We conduct experiments on the implementation to study the performance of the architecture. We observed that STEPS is highly scalable with increase in the number as well as the size of the file objects hosted by SFS. The performance study also demonstrated that most of the efficiency of policy execution is derived from policy cache. Further, a rate-control mechanism is necessary to ensure that users are isolated from LCM operations. Akshat Verma, David Pease, Upendra Sharma, Marc A. Kaplan, Jim Rubas, Rohit Jain, Murthy V. Devarakonda, Mandis Beigi |
MSST | 2 |
| 2005 | The UCSC Kestrel Parallel ProcessorabstractThe architectural landscape of high-performance computing stretches from superscalar uniprocessor to explicitly parallel systems, to dedicated hardware implementations of algorithms. Single-purpose hardware can achieve the highest performance and uniprocessors can be the most programmable. Between these extremes, programmable and reconfigurable architectures provide a wide range of choice in flexibility, programmability, computational density, and performance. The UCSC Kestrel parallel processor strives to attain single-purpose performance while maintaining user programmability. Kestrel is a single-instruction stream, multiple-data stream (SIMD) parallel processor with a 512-element linear array of 8-bit processing elements. The system design focuses on efficient high-throughput DNA and protein sequence analysis, but its programmability enables high performance on computational chemistry, image processing, machine learning, and other applications. The Kestrel system has had unexpected longevity in its utility due to a careful design and analysis process. Experience with the system leads to the conclusion that programmable SIMD architectures can excel in both programmability and performance. This work presents the architecture, implementation, applications, and observations of the Kestrel project at the University of California at Santa Cruz. Andrea Di Blas, David M. Dahle, Mark Diekhans, Leslie Grate, Jeffrey D. Hirschberg, Kevin Karplus, Hansjörg Keller, Mark Kendrick, Francisco J. Mesa-Martinez, David Pease, Eric Rice, Angela Schultz, Don Speck, Richard Hughey |
IEEE Trans. Parallel Distributed Syst. | 10 |
| 2004 | Polus: Growing Storage QoS Management Beyond a "4-Year Old Kid"
Sandeep Uttamchandani, Kaladhar Voruganti, Sudarshan M. Srinivasan, John Palmer, David Pease |
FAST | 5 |