VLDB 2026 Research / reviewers in the wild / expert
Hao Hu 0015
dblp:67/6924-15
· DBLP profile ↗
12ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0003-0726-9953ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving Compression Ratio of Lossy Compression on HPC Datasets via Modeling-Based Arithmetic CodingabstractHPC applications generate massive amounts of data that impose significant burdens on both storage and I/O systems. Although lossy compressors have been widely adopted in this scenario to reduce data volume, the SOTA approach fails to fully exploit redundancy because it relies on separate techniques that operate at incompatible granularities. Their suboptimal compression ratios leave I/O as the dominant bottleneck in data dumps/loads. Therefore, we propose MAC, a compression framework built upon existing SZ compressor. It leverages the alignment between HPC system characteristics and modeling-based arithmetic coding to balance the compression ratio improvement and time cost. Instead of applying Huffman coding and dictionary-based compressors like zstd or gzip sequentially on quantization factors as SZ does, MAC replaces them with an adaptive arithmetic encoder. Specifically, MAC first employs bit-packing on incoming quantization factors to reduce overhead, as these factors are typically small enough that standard 4-byte storage would impede processing efficiency. The system then constructs context with the knowledge of the length of each quantization factor, utilizing hash tables to store and retrieve historical occurrences. By leveraging two models with distinct prefix-matching strategies and integrating them via a logistic mixer, MAC yields substantial compression gains. This architecture ensures that compression and decompression latencies remain low enough to accelerate overall dump and load operations. Experiments show that MACSZ achieves a compression ratio improvement of over 25%, which translates directly into enhanced throughput on HPC cluster architectures as Fig 1 and 2 demonstrate. Zhichao Yang 0019, Xiangyu Zou, Hao Hu 0015, Wen Xia |
DCC | 4 |
| 2026 | CuCM: A GPU-Powered Context-Mixing Compressor for Archival StorageabstractThe explosive growth of global data has created an increasing demand for archival storage, where efficient compression is crucial to reduce capacity cost. However, existing archival compressors face a fundamental tradeoff: mainstream methods (e.g., ZSTD with level$21 / 22$) offer limited compression ratios, while context-mixing compressors (e.g., LPAQ) achieve higher ratios but are often too slow for practical use. Therefore, we present CuCM, a GPU-powered context-mixing compressor to overcome this tradeoff. By introducing pre-learning and batch update mechanisms, CuCM resolves the data dependencies inherent in the autoregressive modeling of contextmixing compressors. During compression, CuCM processes each predefined vector as a single unit. It utilizes the current model to predict the probability distribution for the entire vector, deferring model updates until the vector is fully processed. During decompression, CuCM employs an aggressive look-ahead strategy, preassuming bit values for context construction. It then retains only the outcomes of hypotheses that remain consistent with the actual decoded data. Experiments like figure 1 show that CuCM achieves up to$12.6 \times$higher throughput than LPAQ while maintaining comparable compression ratios across both general-purpose and archival datasets. Zhichao Yang 0019, Xiangyu Zou, Hao Hu 0015, Wen Xia |
DCC | 4 |
| 2026 | A High-Performance Persistent Transactional Memory System via Cooperative Concurrency Control
Hao Hu 0015, Xinrui Zheng, Yizou Chen, Xiangyu Zou, Erci Xu, Hongpeng Wang 0002, Wen Xia |
HPDC | 1 |
| 2025 | Apic: A Precomputation-Based Integer Compressor for OLTP DatabasesabstractCurrent compressors for OLTP databases perform well on text but face challenges with integers, although integers are a critical component of the workload. Most existing integer compressors are ineffective as a complementary solution, since they compress integers together and cannot decompress a certain integer individually, making them incompatible with the data access requirement of OLTP databases. To this end, we propose Apic, a precomputation-based arithmetic coding to efficiently compress each integers (a very tiny unit), though small data are always hard to compress, and ensure compatibility with OLTP datasets. Specifically, Apic presents Bitwidth-aware Precomputed Frequency and Prefixaware Precomputed Decoding to tackle challenges of applying arithmetic coding in this scenario, such as the substantial space costs of symbol frequencies and decompression complexity. Evaluations on real-world and desensitized commercial datasets suggest that Apic improves the compression ratio by up to 80% on integers over VByte, while preserving comparable decompression speed and thus query performance. Xiangyu Zou, Kaiwen Deng, Hao Hu 0015, Wen Xia |
DCC | 4 |
| 2025 | The Logic of Fingerprint Upgrade in Deduplicated StorageabstractCryptographically strong hashes are an integral part of deduplicated storage to identify unique chunks, so they are called fingerprints. Many systems were designed with MD5 or SHA-1 hashes, which seemed sufficiently strong at the time. With the ever-increasing computational power and ongoing security attacks, the security and reliability of the original fingerprints in deduplicated storage may be compromised, making them vulnerable to attacks, which can result in data loss or leakage. This may be a concern for deduplicated storage customers that have already purchased systems, who now are required to switch to a stronger hash function such as SHA-256, possibly due to future legal requirements. This reveals the need for upgrading original fingerprints to stronger hashes on existing duplicated storage systems, but there is currently a lack of research on how to perform such an upgrade effectively. Fingerprint upgrade entails not only adding a new fingerprint calculation function but also replacing original fingerprints in existing, large data structures with new fingerprints. One direct method is to first reconstruct all logical files and then deduplicate them again using new fingerprint calculations. However, such a method will repeatedly process duplicate chunks, resulting in significant space and time overhead. In this paper, we propose an upgrade method that decouples computationally and spatially expensive operations from repetitive operations, so that we can release the space of duplicate chunks and avoid excessive space consumption. We also propose leveraging the characteristics of deduplicated data and file similarities to optimize the cache line and data access pattern, respectively, which further reduces the time overhead. Experimental results demonstrate that our upgrade method is practical and efficient, achieving a throughput of$\sim 1.7 \text{GB} / \mathrm{s}$. Boju Chen, Philip Shilane, Xiangyu Zou, Wen Xia, Hao Hu 0015 |
ICCD | 6 |
| 2025 | ALPHA: A Scalable Lock-Free Partitioned Hash Index for Persistent Memory on NUMA ArchitecturesabstractHash indexes are widely used in modern dataintensive applications to support efficient query performance. While persistent memory (PM) provides larger capacity and byteaddressability, we identify that PM-based hash indexes suffer from scalability bottlenecks under high concurrency. This stems from NUMA access imposing significant costs, exacerbating the inherently request-agnostic I/O issues in concurrent read/write operations. To this end, we present ALPHA, a highly scalable hash index designed for persistent memory. The key idea with ALPHA lies in exploiting the natural hash-based data access features of hash indexes to reduce I/O irrelevant to data requests. We propose a NUMA-friendly framework to reduce the negative impacts of remote access through the cross-thread delegation mechanism. To further improve throughput, we use two techniques to address read and write issues of the PM-oriented hash index. First, we propose a novel two-layer hash structure to minimize PM access by filtering unnecessary key retrievals during data reads. Second, we devise a partition-based finegrained data access mechanism that enables lock-free concurrent writes without excessive PM write overheads. Our comprehensive experimental evaluation shows that ALPHA outperforms state-of-the-art PM-based hash indexes by up to$22.17 \times$in throughput while reducing tail latency by up to 97.1 %. Qiyang Zheng, Hao Hu 0015, Yanqi Pan, Wen Xia |
ICCD | 2 |
| 2025 | A Cost-Effective and Decompression-Transparent Compressor for OLTP-Oriented DatabasesabstractThe row-oriented store model is the cornerstone component of modern online transaction processing (OLTP) database systems. In response to the massive increase in data within database systems, compression techniques are employed to enhance storage efficiency. Regrettably, current compression methods suffer from either the amplification issue due to coarse compression granularity or inefficient decompression operations, thus usually decreasing the speed of query processing. To this end, we present DPTC, a cost-effective and decompression-transparent approach designed to compress data pages, the basic storage unit of OLTP database systems. Specifically, (1) DPTC applies a row-wise decompression-oriented structure to track the first occurrence of redundant data in compressed data, which effectively supports the decompression of individual records from pages, thereby avoiding unwarranted decompression in record access. Moreover, (2) DPTC employs an in-page dynamic packing strategy, which determines the compression units based on the impact of each data reduction operation on the compression gains and eliminates gains-inefficient data reductions. Furthermore, (3) DPTC utilizes a SIMD-based mechanism that leverages the characteristics of operations within the decompression process to improve the decompression speed. Our evaluation results confirm that DPTC is efficient in terms of decompression speed and compression ratio. Within an OLTP database system, DPTC yields throughput improvements of up to 4.28 x in TPC-C and reduces latency by up to 33.3% for data point queries in a row-oriented storage engine. Hao Hu 0015, Qiyang Zheng, Xiangyu Zou, Lisha Qin, Wanchuan Zhang, Zhaoheng Jiang, Dingwen Tao, Hongpeng Wang 0002, Wen Xia |
ICDE | 1 |
| 2025 | PDPH: A Performant Dynamic Perfect Hash Index with Rehashing OptimizationsabstractA dynamic perfect hashing index offers collision-free assignment for entries, ensuring excellent query performance, an essential metric for the Quality of Service (QoS) in database and big data systems. Its core mechanism is rehashing, which resolves collisions during insertions by reassigning entries to distinct slots. However, rehashing is costly, as hash collisions increase linearly with the number of entries in the hash bucket, and the exponential complexity of finding a feasible reassignment. Therefore, rehashing significantly limits the insertion performance of dynamic perfect hashing indexes. To this end, we propose PDPH, a hybrid approach that enhances insertion performance by combining static and dynamic perfect hashing techniques. Specifically, (1) PDPH leverages the complementary strengths of both techniques to reduce the rehashing complexity to a linear level, rather than the exponential overhead of previous approaches. (2) It further reduces the frequency of rehashing by enforcing multiple insertions to share a single rehashing operation. (3) It also limits the scale of rehashing by allowing partial rehashing operations. Evaluations on realworld datasets suggest that, compared to the state-of-the-art approach (MapEmbed), PDPH enhances insertion throughput by up to$2.32 \times$, while maintaining query performance. Xiangyu Zou, Wen Xia, Hao Hu 0015 |
IWQoS | 4 |
| 2025 | MANS: Efficient and Portable ANS Encoding for Multi-Byte Integer Data on CPUs and GPUsabstractLossless compression is a classic technique for reducing data storage and transmission requirements. Asymmetric Numeral Systems (ANS) is a high-throughput, high-ratio lossless compression algorithm, but it lacks effective support for multi-byte data and cross-platform compatibility. To address this issue, we propose an Adaptive Data Mapping (ADM) scheme, which maps multi-byte integer data into single-byte space based on the data’s characteristics, improving the compression ratio of ANS while maintaining low encoding redundancy. We also optimize the ADM algorithm and the ANS encoder for GPU and CPU architectures, respectively, and combine them to create an efficient and portable ANS encoding method for multi-byte integer data, called MANS. Experimental results show that MANS improves compression ratios by an average of 1.24 ×, achieves 870.27MB/s throughput on CPUs, and delivers up to 288.45 × and 135.86 × speedups on an NVIDIA A100 and an AMD MI210 GPU compared to the CPU version—demonstrating its efficiency and portability across platforms. Wenjing Huang 0002, Jinwu Yang, Shengquan Yin, Haoxu Li, Yida Gu, Xing Jing, Shiyuan Fu, Hao Hu 0015, Guangming Tan, Dingwen Tao |
SC | 10 |
| 2024 | Leveraging Partitioning to Mitigate Concurrent Conflicts in Disaggregated Memory Key-Value StoresabstractThe adoption of disaggregated memory (DM) in key-value (KV) storage systems is considered a cost-effective and efficient solution for addressing the significant performance challenges encountered by conventional KV storage systems. However, these systems must handle substantial concurrent requests, making it essential to detect and resolve conflicts to ensure the data correctness. Existing approaches guarantee the correctness of concurrent operations by Compare And Swap (CAS) but consume more network round-trip times (RTTs) to degrade performance. In addition, previous methods incur additional overhead when DM nodes fail.To address the above issues, this paper introduces AKV, a high-performance Agent-Based Key-Value Store on disaggregated memory. AKV partitions keys according to specified strategies, where a single partition’s keys are managed by the same agent to handle read and write requests from multiple clients. This design mitigates the likelihood of concurrency conflicts by enforcing fine-grained serialization of requests within each partition. Specifi-cally, to partition keys, AKV proposes load-aware and affinity-aware strategies. To handle concurrent requests in a fine-grained serialized manner, AKV introduces a partition-level concurrency control scheme without RDMA_CAS. To detect the agent failure and recovery for high availability, AKV proposes a decentralized approach without additional management servers. We evaluate AKV with micro and real-world benchmarks. Experimental results show that AKV outperforms the state-of-the-art KV stores on DM by up to 1.8 × in throughput. Lisha Qin, Hao Hu 0015, Wen Xia |
HPCC | 4 |
| 2023 | DRPTM: A Decoupled Read-efficient High-scalable Persistent Transactional MemoryabstractPersistent transactional memory (PTM) exploits transactions to provide an easy crash-consistent interface for persistent memory (PM). However, because of the substantial reader-side overhead brought on by the low bandwidth and long persistence latency of PM, present PTM research cannot scale effectively. This paper proposes a highly scalable PTM system, DRPTM, which allows nearly non-overhead reads without lowering the isolation level. DRPTM decouples persistence latency from concurrency control and traces the read-only copy maintained in logs as a lightweight read set. The evaluation shows that DRPTM significantly outperforms the state-of-the-art PTM systems for various workloads and achieves near-linear scalability. Wenkai Liang, Hao Hu 0015, Xiangyu Zou, Wen Xia, Yanqi Pan |
DAC | 2 |
| 2023 | EEPH: An Efficient Extendible Perfect Hashing for Hybrid PMem-DRAMabstractIn recent years, the performance of hash indexes has been significantly improved by exploiting emerging persistent memory (PMem). However, the performance improvement of hash indexes mainly comes from exploiting the hardware features of PMem. Only a few studies optimize the hash index itself to fully exploit the potential of PMem. Interestingly, many of these studies improve the performance of write, but disregard the performance of read, of hash indexes on PMem. With extensive experimental evaluation, we find the major reason for inefficient read in the hash index on PMem is that the overhead of hash collision processing is expensive.To address that, we propose a novel Efficient Extendible Perfect Hashing (EEPH) on PMem-DRAM hybrid data layout to improve read performance of hash indexes. Specifically, we reduce the overhead of dynamic perfect hashing extension on PMem by combing extendible hashing. We then design a hybrid data layout to unlock the inherent read strengths of perfect hashing (i.e., zero collision). Last, we devise a complement move algorithm to efficiently guarantee the zero collision of perfect hashing when data move is conducted on PMem. We compare EEPH with the state-of-the-art hash indexes on PMem by conducting comprehensive experiments on several real-world read-intensive and read-skew workloads. The experimental results confirm the superiority of our EEPH as it achieves up to 2.21× higher throughput and about 1/3 of the 99th percentile latency than state-of-the-art hash indexes. Hao Hu 0015, Dingbang Liu, Bo Tang 0016, Wen Xia |
ICDE | 2 |