EDBT 2026 Demo / reviewers in the wild / expert
Haoliang Tan
dblp:233/3132
· DBLP profile ↗
10ranked-venue papers
7as first author
8since 2021 · last 2026
0009-0003-9919-1926ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 4 first-author · 4 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Once Rolling Hashing is Enough: Exploiting Rolling Hash Reuse in Delta CompressionabstractIn backup storage, delta compression successfully achieves a much higher data reduction ratio than chunk-level deduplication by applying a finer granularity in redundancy detection and elimination. However, it introduces additional, intensive computation overhead and results in a 50%–70% worsening backup throughput. Haoliang Tan, Wenhao Ou, Xiangyu Zou, Yanqi Pan, Zhaoquan Gu, Wen Xia |
EuroSys | 1 |
| 2026 | The Enduring Potential of Locality: Efficient Routing for Fine-Grained Cluster Deduplication
Xiangyu Zou, Philip Shilane, Haoliang Tan, Wen Xia |
IWQoS | 4 |
| 2026 | Improving the Restore Performance of Fine-Grained Deduplication on Docker Container Image StorageabstractThe rapid growth of container images leads to heavy storage pressures on image registries. Fine-grained deduplication, such as at the file and chunk levels, is a promising technique for reducing the storage space in image registries, compared to Docker's native coarse-grained image layer deduplication. However, fine-grained deduplication often incurs significant image restoration latency due to fragmented I/Os, resulting in up to 8× restoration I/O slowdowns. Moreover, existing restore-optimized deduplication techniques are not tailored to the characteristics of the container image, resulting in low deduplication ratios of images. Consequently, restoring performance remains a critical barrier to the practical adoption of fine-grained deduplication in container image registries. To address this challenge, we propose MiDedup, a restore-friendly fine-grained deduplication approach tailored for container image registries. MiDedup is built on an underexplored observation: the redundancy across layers of images follows a non-uniform distribution. Based on this insight, MiDedup introduces three core techniques: ① Across-layer-aware reorganization, which reorganizes deduplicated data into a compact, sequential layout, significantly reduces fragmented I/Os during image restore. ② Popularity-aware rewriting, which selectively rewrites hot image layers to further improve the balance between deduplication ratio and restore performance. ③ Hybrid-granularity deduplication, which combines file-level and chunk-level deduplication to reduce metadata and fragmented data. Experiments on the real-world Docker image dataset show that MiDedup reduces the 50–80% restore I/O overhead of fine-grained deduplication, which is comparable to the restore performance of native layer-level deduplication, while achieving 1.2–7.3× higher deduplication ratios. Haoliang Tan, Wenhao Ou, Xiangyu Zou, Lisha Qin, Zhaoquan Gu, Wen Xia |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2025 | A Comprehensive Study of Data Reduction Methods on Docker Container ImagesabstractThe rapid growth of Docker images in cloud infrastructures has intensified storage and network demands, posing challenges to QoS for registries and deployments. This paper systematically evaluates four data reduction methods—file-level and block-level deduplication, delta compression, and local compression—on 18 representative images. We quantify their tradeoffs in compression, computation, I/O overhead, and restore performance, revealing that (1) filtering large files ($>64 \text{KB}$) preserves 80% of redundancy elimination at half the cost, (2) category-based reorganization reduces restore latency by 83%, and (3) fixed-size chunking with optimized blocks balances memory and compression. Based on these insights, we propose adaptive strategies for efficient container storage. Haoliang Tan, Lisha Qin, Xiangyu Zou, Zhaoquan Gu, Wen Xia |
IWQoS | 1 |
| 2024 | SuperDelta: Multiple Referenced Base Chunks Scheme for Fine-grained Deduplication Backup Storage SystemabstractDeduplication-based techniques are popular in backup storage systems for reducing data volume. To maximize data reduction, existing fine-grained deduplication approaches not only eliminate duplicate chunks but also delta-compress non-duplicate chunks as delta relative to their similar (base) chunks. However, each chunk may have multiple similar chunks, and delta compression usually only selects one of them as the base chunk, i.e., a one-to-one scheme. This scheme benefits the restore performance because it needs to read only one (instead of multiple) base chunk in decompressing delta chunks, while it also wastes the potential compressibility among other similar chunks.In this paper, we propose SuperDelta to further exploit compressibility across multiple similar chunks and to preserve the restore performance advantage of the one-to-one scheme as much as possible. It is based on three techniques. (1) To further eliminate redundancy among similar chunks, SuperDelta applies a "Multiple Referenced Base Chunks" (MRBC) scheme instead of the one-to-one scheme. It combines several similar pairs of chunks in delta encoding to recover possibly lost compressibility in "boundary shift" problems. (2) To avoid the negative side effects of MRBC on restore performance, SuperDelta introduces a rebase scheme to rebuild simple reference paths among duplicate and similar chunks. It significantly simplifies the restore workflow, but also costs slightly more storage space because of impacting the workflow of redundancy detection. (3) To compensate for the additional storage cost, SuperDelta applies a space-recycle scheme to remove derived data when they become old while ensuring the optimized restore performance of the latest backups.Experiments on four real-world backup datasets show that SuperDelta increases the overall compression ratio by 1.05~2.40 times than the traditional one-to-one fine-grained deduplication without significantly affecting the backup and restore throughput. Haoliang Tan, Xiangyu Zou, Binzhaoshuo Wan, Zhaoquan Gu, Wen Xia |
DCC | 1 |
| 2024 | MiDedup: A Restore-Friendly Deduplication Method on Docker Image Storage Systems
Lisha Qin, Haoliang Tan, Xiangyu Zou, Wenhao Ou, Rubing Huang, Wen Xia |
NPC (1) | 2 |
| 2024 | The Design of Fast Delta Encoding for Delta Compression Based Storage SystemsabstractDelta encoding is a data reduction technique capable of calculating the differences (i.e., delta) among very similar files and chunks. It is widely used for various applications, such as synchronization replication, backup/archival storage, cache compression, and so on. However, delta encoding is computationally costly due to its time-consuming word-matching operations for delta calculation. Existing delta encoding approaches either run at a slow encoding speed, such as Xdelta and Zdelta, or at a low compression ratio, such as Ddelta and Edelta. In this article, we propose Gdelta, a fast delta encoding approach with a high compression ratio. The key idea behind Gdelta is the combined use of five techniques: (1) employing an improved Gear-based rolling hash to replace Adler32 hash for fast scanning overlapping words of similar chunks, (2) adopting a quick array-based indexing for word-matching, (3) applying a sampling indexing scheme to reduce the cost of traditional building full indexes for base chunks’ words, (4) skipping unmatched words to accelerate delta encoding through non-redundant areas, and (5) last but not least, after word-matching, further batch compressing the remainder to improve the compression ratio. Our evaluation results driven by seven real-world datasets suggest that Gdelta achieves encoding/decoding speedups of 3.5X∼25X over the classic Xdelta and Zdelta approaches while increasing the compression ratio by about 10%∼240%. Haoliang Tan, Wen Xia, Xiangyu Zou, Qing Liao 0001, Zhaoquan Gu |
ACM Trans. Storage | 1 |
| 2021 | Odess: Speeding up Resemblance Detection for Redundancy Elimination by Fast Content-Defined SamplingabstractMultiple data reduction techniques have been investigated to lower storage costs for a wide variety of customers. In this work, we focus on similarity-based delta compression, which calculates and stores the difference of very similar, but non-duplicate, chunks in storage systems. Delta compression is often implemented along with deduplication and has been shown to achieve a much higher compression ratio. Currently, the N-Transform method is the most popular and widely-used approach to generate features for data content (e.g. chunks) to detect similar candidates (and then apply delta compression). For delta compression systems, though, the throughput of N-Transform is often the bottleneck. Finesse is a high throughput variant of N-Transform, but it suffers from lower detection accuracy and compression ratio. The computation overhead of N-Transform consists of two parts: calculating the rolling hash across data and applying time-consuming transforms on each hash. In this work, we propose Odess, a fast resemblance detection approach, that uses a novel Content-Defined Sampling method to generate a much smaller proxy hash set and then applies transforms on this small hash set. This reduces the calculations in the transform step from being the bottleneck. Meanwhile, Odess also leverages the faster Gear hash to generate rolling hashes. Thus, Odess greatly reduces the computational overhead for resemblance detection while achieving high detection accuracy and high compression ratio. Our evaluation results show that Odess is ~ 5.4× (Finesse) and ~ 26.9× (N-Transform) faster (on average) at generating features for resemblance detection. When considering an end-to-end data reduction storage system, Odess increases throughput by ~ 1.36× (Finesse) and ~ 2.76× (N-Transform) while maintaining the compression ratio of N-Transform and increasing the compression ratio ~ 1.22× over Finesse. Xiangyu Zou, Wen Xia, Philip Shilane, Haoliang Tan, Haijun Zhang 0002, Xuan Wang 0002 |
ICDE | 5 |
| 2020 | Exploring the Potential of Fast Delta Encoding: Marching to a Higher Compression RatioabstractDelta compression (or called delta encoding) is a data reduction technique capable of calculating the differences (i.e., delta) among the very similar files and chunks, and is thus widely used for optimizing synchronization replication, backup/archival storage, cache compression, etc. However, delta compression is costly because of its time-consuming word-matching operations for delta calculation. Existing delta encoding approaches, are either at a slow encoding speed, such as Xdelta and Zdelta, or at a low compression ratio, such as Ddelta and Edelta. In this paper, we propose Gdelta, a fast delta encoding approach with a high compression ratio, that improves the delta encoding speed by employing an improved fast Gear-based rolling hash for scanning fine-grained words, and a quick array-based indexing scheme for word-matching, and then, after word-matching, further batch compressing the rest to improve the compression ratio. Our evaluation results driven by six real-world datasets suggest that Gdelta achieves encoding/decoding speedups of 2X~4X over the classic Xdelta and Zdelta approaches while increasing the compression ratio by about 10%~120%. Haoliang Tan, Xiangyu Zou, Qing Liao 0001, Wen Xia |
CLUSTER | 1 |
| 2019 | Object Affordances Graph Network for Action Recognition
Haoliang Tan, Le Wang 0003, Qilin Zhang 0004, Zhanning Gao, Nanning Zheng 0001, Gang Hua 0001 |
BMVC | 1 |