Philip Shilane

dblp:06/4348 · DBLP profile ↗
← Back
12ranked-venue papers in the field
1as first author
5since 2021 · last 2024
0000-0003-1235-0502ORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 10 (1 first)Database Systems & Data Management · 2
YearPublicationVenuePosition
2024 Physical vs. Logical Indexing with IDEA: Inverted Deduplication-Aware Index
Asaf Levi, Philip Shilane, Sarai Sheinvald, Gala Yadgar
FAST2
2023 Dataset Similarity Detection for Global Deduplication in the DD File System
abstract
Deduplication has become a widely used technique to reduce space requirements for storage systems by replacing redundant chunks of data with references. While storage systems continue to grow in size, there remain practical limits to the size of any deduplication node, and enterprise businesses may have dozens to hundreds of nodes. It is important to place datasets on nodes in a multi-node environment to take advantage of deduplication savings globally. For customers of the DD File System (DDFS)1, we provide the Global Deduplication Service that advises customers on data placement to maximize deduplication-related space savings. This paper describes our currently shipping approach that uses a Fingerprint Dictionary to intelligently cluster customer data and generate a plan to relocate datasets to improve global deduplication. We report results from thousands of deployed systems at customer sites. We have also developed a further improvement using MinHashes that lowers resource requirements, and we provide proofs of the similarity estimates. Our results on a real-world dataset show that MinHashes improve the clustering speed up to 400X relative to our previous method and reduce memory consumption up to 260X.
Tony Wong, Smriti Thakkar, Kao-Feng Hsieh, Zachary Tom, Hetaben Saraiya, Philip Shilane
ICDE6
2022 DedupSearch: Two-Phase Deduplication Aware Keyword Search
Nadav Elias, Philip Shilane, Sarai Sheinvald, Gala Yadgar
FAST2
2021 The Dilemma between Deduplication and Locality: Can Both be Achieved?
Xiangyu Zou, Jingsong Yuan, Philip Shilane, Wen Xia, Haijun Zhang 0002, Xuan Wang 0002
FAST3
2021 Odess: Speeding up Resemblance Detection for Redundancy Elimination by Fast Content-Defined Sampling
abstract
Multiple data reduction techniques have been investigated to lower storage costs for a wide variety of customers. In this work, we focus on similarity-based delta compression, which calculates and stores the difference of very similar, but non-duplicate, chunks in storage systems. Delta compression is often implemented along with deduplication and has been shown to achieve a much higher compression ratio. Currently, the N-Transform method is the most popular and widely-used approach to generate features for data content (e.g. chunks) to detect similar candidates (and then apply delta compression). For delta compression systems, though, the throughput of N-Transform is often the bottleneck. Finesse is a high throughput variant of N-Transform, but it suffers from lower detection accuracy and compression ratio. The computation overhead of N-Transform consists of two parts: calculating the rolling hash across data and applying time-consuming transforms on each hash. In this work, we propose Odess, a fast resemblance detection approach, that uses a novel Content-Defined Sampling method to generate a much smaller proxy hash set and then applies transforms on this small hash set. This reduces the calculations in the transform step from being the bottleneck. Meanwhile, Odess also leverages the faster Gear hash to generate rolling hashes. Thus, Odess greatly reduces the computational overhead for resemblance detection while achieving high detection accuracy and high compression ratio. Our evaluation results show that Odess is ~ 5.4× (Finesse) and ~ 26.9× (N-Transform) faster (on average) at generating features for resemblance detection. When considering an end-to-end data reduction storage system, Odess increases throughput by ~ 1.36× (Finesse) and ~ 2.76× (N-Transform) while maintaining the compression ratio of N-Transform and increasing the compression ratio ~ 1.22× over Finesse.
Xiangyu Zou, Wen Xia, Philip Shilane, Haoliang Tan, Haijun Zhang 0002, Xuan Wang 0002
ICDE4
2017 The Logic of Physical Garbage Collection in Deduplicating Storage
Fred Douglis, Abhinav Duggal, Philip Shilane, Tony Wong, Shiqin Yan, Fabiano C. Botelho
FAST3
2016 Using Hints to Improve Inline Block-layer Deduplication
Sonam Mandal, Geoffrey H. Kuenning, Dongju Ok, Varun Shastry, Philip Shilane, Sun Zhen, Vasily Tarasov, Erez Zadok
FAST5
2014 Migratory compression: coarse-grained data reordering to improve compressibility
Guanlin Lu, Fred Douglis, Philip Shilane, Grant Wallace
FAST4
2013 Memory efficient sanitization of a deduplicated storage system
Fabiano C. Botelho, Philip Shilane, Windsor W. Hsu
FAST2
2012 WAN optimized replication of backup datasets using stream-informed delta compression
Philip Shilane, Mark Huang, Grant Wallace, Windsor W. Hsu
FAST1
2012 Characteristics of backup workloads in production systems
Grant Wallace, Fred Douglis, Hangwei Qian, Philip Shilane, Stephen Smaldone, Mark Chamness, Windsor W. Hsu
FAST4
2011 Tradeoffs in Scalable Data Routing for Deduplication Clusters
Wei Dong 0003, Fred Douglis, Kai Li 0001, R. Hugo Patterson, Sazzala Reddy, Philip Shilane
FAST6