EDBT 2026 Demo / reviewers in the wild / expert
Lisha Qin
dblp:402/1790
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving the Restore Performance of Fine-Grained Deduplication on Docker Container Image StorageabstractThe rapid growth of container images leads to heavy storage pressures on image registries. Fine-grained deduplication, such as at the file and chunk levels, is a promising technique for reducing the storage space in image registries, compared to Docker's native coarse-grained image layer deduplication. However, fine-grained deduplication often incurs significant image restoration latency due to fragmented I/Os, resulting in up to 8× restoration I/O slowdowns. Moreover, existing restore-optimized deduplication techniques are not tailored to the characteristics of the container image, resulting in low deduplication ratios of images. Consequently, restoring performance remains a critical barrier to the practical adoption of fine-grained deduplication in container image registries. To address this challenge, we propose MiDedup, a restore-friendly fine-grained deduplication approach tailored for container image registries. MiDedup is built on an underexplored observation: the redundancy across layers of images follows a non-uniform distribution. Based on this insight, MiDedup introduces three core techniques: ① Across-layer-aware reorganization, which reorganizes deduplicated data into a compact, sequential layout, significantly reduces fragmented I/Os during image restore. ② Popularity-aware rewriting, which selectively rewrites hot image layers to further improve the balance between deduplication ratio and restore performance. ③ Hybrid-granularity deduplication, which combines file-level and chunk-level deduplication to reduce metadata and fragmented data. Experiments on the real-world Docker image dataset show that MiDedup reduces the 50–80% restore I/O overhead of fine-grained deduplication, which is comparable to the restore performance of native layer-level deduplication, while achieving 1.2–7.3× higher deduplication ratios. Haoliang Tan, Wenhao Ou, Xiangyu Zou, Lisha Qin, Zhaoquan Gu, Wen Xia |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | A Cost-Effective and Decompression-Transparent Compressor for OLTP-Oriented DatabasesabstractThe row-oriented store model is the cornerstone component of modern online transaction processing (OLTP) database systems. In response to the massive increase in data within database systems, compression techniques are employed to enhance storage efficiency. Regrettably, current compression methods suffer from either the amplification issue due to coarse compression granularity or inefficient decompression operations, thus usually decreasing the speed of query processing. To this end, we present DPTC, a cost-effective and decompression-transparent approach designed to compress data pages, the basic storage unit of OLTP database systems. Specifically, (1) DPTC applies a row-wise decompression-oriented structure to track the first occurrence of redundant data in compressed data, which effectively supports the decompression of individual records from pages, thereby avoiding unwarranted decompression in record access. Moreover, (2) DPTC employs an in-page dynamic packing strategy, which determines the compression units based on the impact of each data reduction operation on the compression gains and eliminates gains-inefficient data reductions. Furthermore, (3) DPTC utilizes a SIMD-based mechanism that leverages the characteristics of operations within the decompression process to improve the decompression speed. Our evaluation results confirm that DPTC is efficient in terms of decompression speed and compression ratio. Within an OLTP database system, DPTC yields throughput improvements of up to 4.28 x in TPC-C and reduces latency by up to 33.3% for data point queries in a row-oriented storage engine. Hao Hu 0015, Qiyang Zheng, Xiangyu Zou, Lisha Qin, Wanchuan Zhang, Zhaoheng Jiang, Dingwen Tao, Hongpeng Wang 0002, Wen Xia |
ICDE | 4 |
| 2025 | An Effective Uncorrectable Memory Error Prediction Framework by Exploiting UPH Indicators in Production EnvironmentsabstractUCEs (Uncorrectable memory errors) pose significant challenges to cloud computing systems, often resulting in catastrophic failures and crashes. Researchers have explored prediction approaches to address this issue. Previous studies have provided insights into memory error prediction, focusing on memory module part numbers and relationships between error code data. However, these efforts face challenges due to insufficient data features and suboptimal optimization, especially in production environments where hardware/software sparing techniques are widely deployed, the UCE ratio is low, and long lead time is required. To address these issues, our study first collect a large amount of memory data from different vendors in Huawei's production environment, which has deployed hardware/software sparing techniques, to provide more general data. Second, we exploit new indicators termed UPH (Unique, Pinx, and History) from this data, which play a crucial role in predicting UCEs. UPH offers a more profound understanding of the factors contributing to UCEs and demonstrates higher precision and recall. Then, we integrate existing indicators and UPH into our prediction framework and demonstrate the significance of UPH through indicator importance assessments. We also optimize the framework by determining an optimal sampling window. In production environments with long lead time and low UCE ratio, we improve the framework by implementing noise reduction, self-history learning, and a new scenario-based model selection approach. Experimental results demonstrate 19 % - 27 % increase in UCE prediction recall with 4 %-11 % increase in precision under different scenarios, outperforming state-of-the-art methods in production environments. Xiaobo Zheng, Lisha Qin, Wen Xia, Chentao Wu, Yunfei Gu, Qicong Lin, Huifang Jiao, Rubing Huang |
IPDPS | 2 |
| 2025 | A Comprehensive Study of Data Reduction Methods on Docker Container ImagesabstractThe rapid growth of Docker images in cloud infrastructures has intensified storage and network demands, posing challenges to QoS for registries and deployments. This paper systematically evaluates four data reduction methods—file-level and block-level deduplication, delta compression, and local compression—on 18 representative images. We quantify their tradeoffs in compression, computation, I/O overhead, and restore performance, revealing that (1) filtering large files ($>64 \text{KB}$) preserves 80% of redundancy elimination at half the cost, (2) category-based reorganization reduces restore latency by 83%, and (3) fixed-size chunking with optimized blocks balances memory and compression. Based on these insights, we propose adaptive strategies for efficient container storage. Haoliang Tan, Lisha Qin, Xiangyu Zou, Zhaoquan Gu, Wen Xia |
IWQoS | 2 |
| 2024 | Leveraging Partitioning to Mitigate Concurrent Conflicts in Disaggregated Memory Key-Value StoresabstractThe adoption of disaggregated memory (DM) in key-value (KV) storage systems is considered a cost-effective and efficient solution for addressing the significant performance challenges encountered by conventional KV storage systems. However, these systems must handle substantial concurrent requests, making it essential to detect and resolve conflicts to ensure the data correctness. Existing approaches guarantee the correctness of concurrent operations by Compare And Swap (CAS) but consume more network round-trip times (RTTs) to degrade performance. In addition, previous methods incur additional overhead when DM nodes fail.To address the above issues, this paper introduces AKV, a high-performance Agent-Based Key-Value Store on disaggregated memory. AKV partitions keys according to specified strategies, where a single partition’s keys are managed by the same agent to handle read and write requests from multiple clients. This design mitigates the likelihood of concurrency conflicts by enforcing fine-grained serialization of requests within each partition. Specifi-cally, to partition keys, AKV proposes load-aware and affinity-aware strategies. To handle concurrent requests in a fine-grained serialized manner, AKV introduces a partition-level concurrency control scheme without RDMA_CAS. To detect the agent failure and recovery for high availability, AKV proposes a decentralized approach without additional management servers. We evaluate AKV with micro and real-world benchmarks. Experimental results show that AKV outperforms the state-of-the-art KV stores on DM by up to 1.8 × in throughput. Lisha Qin, Hao Hu 0015, Wen Xia |
HPCC | 2 |
| 2024 | MiDedup: A Restore-Friendly Deduplication Method on Docker Image Storage Systems
Lisha Qin, Haoliang Tan, Xiangyu Zou, Wenhao Ou, Rubing Huang, Wen Xia |
NPC (1) | 1 |