David D. Chambliss

dblp:89/2753 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4Systems, architecture and hardware · 3Graphics, computer vision, multimedia, augmented reality and games · 3Security and privacy · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Storage systems · 53% Performance modeling and evaluation · 47%
Databases, data mining, and information retrieval
1 paper
Distributed and cloud data management · 77% Query processing and optimization · 23%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
workload characterization
0.212015
Visualizing Block IO Workloads · ACM Trans. Storage 2015
Distributed and cloud data management
query offloading
0.112008
Integration of Server, Storage and Database Stack: Moving Processing Towards Data · ICDE 2008
Query processing and optimization › OLAP
star query
0.012008
Integration of Server, Storage and Database Stack: Moving Processing Towards Data · ICDE 2008
Storage systems › computational storage
in-storage computing
0.012008
Integration of Server, Storage and Database Stack: Moving Processing Towards Data · ICDE 2008

Methods — techniques the papers use, named apart from their topics

trace visualization · 0.2value proposition analysis · 0.2
YearPublicationVenuePosition
2016 Quick Access to Compressed Data in Storage Systems
abstract
Summary form only given. Primary storage systems that compress data in real time, use some form of on disk metadata to perform the virtualization needed in storing compressed data. Usually this metadata is in the form of B-trees (eventually compressed) and stored on disk. For random accesses to compressed data, where the metadata is not in cache, this additional layer significantly slows down random reads and writes. Our solution is to use much less metadata that only provides an approximation of the location of compressed data on disk and can be easily stored in the memory of the storage system. Read operations are extended to compensate for the imprecise position information in the metadata, and index marks embedded in the data are used to locate the required data within the expanded read. The data placement of written data is constrained to be described by the reduced metadata. The placement uses a piecewise linear scheme based on the locality in compressibility of data and we support this assumption with experiments.
Cornel Constantinescu, David D. Chambliss
DCC2
2015 Visualizing Block IO Workloads
abstract
Massive block IO systems are the workhorses powering many of today’s largest applications. Databases, health care systems, and virtual machine images are examples for block storage applications. The massive scale of these workloads, and the complexity of the underlying storage systems, makes it difficult to pinpoint problems when they occur. This work attempts to shed light on workload patterns through visualization, aiding our intuition. We describe our experience in the last 3 years of analyzing and visualizing customer traces from XIV, an IBM enterprise block storage system. We also present results from applying the same visualization technology to Linux filesystems. We show how visualization aids our understanding of workloads and how it assists in resolving customer performance problems.
Ohad Rodeh, Haim Helman, David D. Chambliss
ACM Trans. Storage3
2013 Random Extraction from Compressed Data - A Practical Study
abstract
Modern primary storage systems support or intend to add support for real time compression usually based on some flavor of the LZ77 and/or Huffman algorithm. There is a fundamental tradeoff in adding real time (adaptive) compression to such a system: to get good compression the amount of compressed data (the independently compressed block) should be large, to be able to read quickly from random places the blocks should be small. One idea is to let the independently compressed blocks be large but to be able to start decompressing the needed part of the block from a random location inside the compressed block. We explore this idea and compare it with a few alternatives, experimenting with the \zlib\ code base.
Cornel Constantinescu, Joseph S. Glider, Dilip Simha, David D. Chambliss
DCC4
2012 Insights for data reduction in primary storage: a practical analysis
abstract
There has been increasing interest in deploying data reduction techniques in primary storage systems. This paper analyzes large datasets in four typical enterprise data environments to find patterns that can suggest good design choices for such systems. The overall data reduction opportunity is evaluated for deduplication and compression, separately and combined, then in-depth analysis is presented focusing on frequency, clustering and other patterns in the collected data. The results suggest ways to enhance performance and reduce resource requirements and system cost while maintaining data reduction effectiveness. These techniques include deciding which files to compress based on file type and size, using duplication affinity to guide deployment decisions, and optimizing the detection and mapping of duplicate content adaptively when large segments account for most of the opportunity.
Maohua Lu, David D. Chambliss, Joseph S. Glider, Cornel Constantinescu
SYSTOR2
2011 IO Tetris: Deep Storage Consolidation for the Cloud via Fine-Grained Workload Analysis
abstract
Intelligent workload consolidation in storage systems leads to better Return On Investment (ROI), in terms of more efficient use of data center resources, better Quality of Service (QoS), and lower power consumption. This is particularly significant yet challenging in a cloud environment, in which a large set of different workloads multiplex on a shared, heterogeneous infrastructure. However, the increasing availability of fine grained workload logging facilities allows better insights to be gained from workload profiles. As a consequence, consolidation can be done more deeply, according to a detailed understanding of how well given workloads mix. We describe IO Tetris, which takes a first look at fine-grained consolidation in large-scale storage systems by leveraging temporal patterns found in real-world I/O traces gathered from enterprise storage environments. The core functionality of IO Tetris consists of two stages. A grouping stage performs hierarchical grouping of storage workloads to find complementary groupings that consolidate well together over time and conflicting ones that do not. After that, a migration stage examines the discovered groupings to determine how to maximize resource utilization efficiency while minimizing migration costs. Experiments based on customer I/O traces from a high-end enterprise class IBM storage controller show that a non-trivial number of IO Tetris groupings exist in real-world storage workloads, and that these groupings can be leveraged to achieve better storage consolidation in a cloud setting.
Ramani Routray, David M. Eyers, David D. Chambliss, Prasenjit Sarkar, Douglas Willcocks, Peter R. Pietzuch
IEEE CLOUD4
2011 Mixing Deduplication and Compression on Active Data Sets
abstract
Many new storage systems provide some form of data reduction. We examine data reduction methods that might be suitable for \emph{primary} storage systems serving active data (as contrasted with backup and archive systems), by analysis of file sets found in different active data environments. We address questions of: how effective are compression and variations of deduplication, both separately and in combination, when deduplication and compression are combined, which should be applied first, what will the tradeoff be between the different methods in their use of MIPS relative to the data reduction achieved, and what degree of data reduction should be expected for different data types.
Cornel Constantinescu, Joseph S. Glider, David D. Chambliss
DCC3
2010 Effective Quality of Service Differentiation for Real-world Storage Systems
abstract
Data storage is an integral part of IT infrastructures, where Quality of Service (QoS) differentiation amongst customers and their applications is essential for many. Achieving this objective in a production environment is nontrivial, because these environments are complex and dynamic. Numerous practical and engineering constraints render the task even more challenging. This paper presents SLED-2, a QoS differentiation solution that meets these challenges in offering effective protection to the performance of important workloads at the expense of less important workloads when needed. SLED-2 uses a customized feedback heuristic that rate-limits selected I/O streams. This approach is unique in that it accounts for a number of important practical considerations, including fine-grained controls, errors in storage systems models, and inexpensive and safe QoS management. SLED-2 has been implemented for the IBM DS8000 series storage servers and shown to be highly effective in a set of hostile and practical scenarios using test facilities for IBM storage products.
David D. Chambliss, Prashant Pandey 0005, William Shearman, Joseph Hyde
MASCOTS2
2008 Integration of Server, Storage and Database Stack: Moving Processing Towards Data
abstract
Storage architecture includes more and more processing power for increasing requirement of reliability, managibility and scalability. For example, an IBM storage server is equipped with 4 or 8 state-of-the-art processors and gigabytes of memories. This trend enables analyzing data locally inside a storage server. Processing data locally is appealing under the following circumstances: (1) huge reduction of data flowing to the host, (2) reduction of CPU consumption on host. Accordingly, the benefits are (1) less data traffic through IO channel to the host, (2) better utilization of host bufferpool, and (3) enabling more workload on the host. One crucial task is to understand how DBMS can benefit from such hardware. That is to identify which database operations are beneficial to be offloaded given a query workload in a particular setting. For certain operations, we establish value proposition via various approaches and show the analytical and experimental results. In particular, starjoin queries are commonly used in business warehouses. We propose to offload a portion of a starjoin query from host to the POWER5 P processors on a storage server, which dramatically reduces the amount of channel IO and host CPU consumption. Moreover, the query elapsed time is improved via the exploitation of the state-of-the-art P processors on a storage server.
Lin Qiao 0001, Vijayshankar Raman, Inderpal Narang, Prashant Pandey 0005, David D. Chambliss, Gene Fuh, James A. Ruddy, Ying-Lin Chen, Kou-Horng Yang, Fen-Ling Ling
ICDE5
2003 Performance Virtualization for Large-Scale Storage Systems
abstract
Current data centers require storage capacities of hundreds of terabytes to petabytes. Time-critical applications such as online transactions processing depend on getting adequate performance from the storage subsystem: otherwise, they fail. It is difficult to provide predictable quality of service at this level of complexity, because I/O workloads are extremely variable and device behavior is poorly understood. Ensuring that unrelated but competing workloads do not affect each other's performance is still more difficult, and equally necessary. We present SLEDS (Service Level Enforcement Discipline for Storage), a distributed controller that provides statistical performance guarantees on a storage system built from commodity components. SLEDS can adaptively handle unpredictable workload variations so that each client continues to get the performance it needs even in the presence of misbehaving, competing peers. After evaluating the SLEDS on a heterogeneous mid-range storage system, we found that it is vastly superior to the raw system in its ability to provide performance guarantees, while only introducing a negligible overhead.
David D. Chambliss, Guillermo A. Alvarez, Prashant Pandey 0005, Divyesh Jadav, Ram Menon, Tzongyu P. Lee
SRDS1