VLDB 2026 Research / reviewers in the wild / expert
Wenjing Huang 0002
dblp:07/3842-2
· DBLP profile ↗
9ranked-venue papers
1as first author
9since 2021 · last 2026
0009-0009-5541-3519ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM TrainingabstractHandling communication overhead in large-scale tensor-parallel training remains a critical challenge due to the dense, near-zero distributions of intermediate tensors, which exacerbate errors under frequent communication and introduce significant computational overhead during compression. To this end, we propose TACO (Tensor-parallel Adaptive COmmunication compression), a robust FP8-based framework for compressing TP intermediate tensors. First, we employ a data-driven reshaping strategy combined with an Adaptive Scale–Hadamard Transform to enable high-fidelity FP8 quantization, while its Dual-Scale Quantization mechanism ensures numerical stability throughout training. Second, we design a highly fused compression operator to reduce memory traffic and kernel launch overhead, allowing efficient overlap with communication. Finally, we integrate TACO with existing state-of-the-art methods for Data and Pipeline Parallelism to develop a compression-enabled 3D-parallel training framework. Detailed experiments on GPT models and Qwen model demonstrate up to 1.87 × end-to-end throughput improvement while maintaining near-lossless accuracy, validating the effectiveness and efficiency of TACO in large-scale training. Xingjian Tian, Bing Lu 0001, Shengkai Lyu, Shengquan Yin, Wenjing Huang 0002, Hairui Zhao 0002, Guangming Tan, Dingwen Tao |
HPDC | 7 |
| 2026 | A Fully GPU-Accelerated Framework for High-Performance Configuration Interaction Selection with Neural Network Quantum StatesabstractAI-driven methods have demonstrated considerable success in tackling the central challenge of accurately solving the Schrödinger equation for complex many-body systems. Among neural network quantum state (NNQS) approaches, the NNQS-SCI (Selected Configuration Interaction) method stands out as a state-of-the-art technique, recognized for its high accuracy and scalability. However, its application to larger systems is severely constrained by a hybrid CPU-GPU architecture. Specifically, centralized CPU-based global de-duplication creates a severe scalability barrier due to communication bottlenecks, while host-resident coupled-configuration generation induces prohibitive computational overheads. We introduce QiankunNet-cuSCI, a fully GPU-accelerated SCI framework designed to overcome these bottlenecks. It first integrates a distributed, load-balanced global de-duplication algorithm to minimize redundancy and communication overhead at scale. To address compute limitations, it employs specialized, fine-grained CUDA kernels for exact coupled configuration generation. Finally, to break the single-GPU memory barrier exposed by this full acceleration, it incorporates a GPU memory-centric runtime featuring GPU-side pooling, streaming mini-batches, and overlapped offloading. This design enables much larger configuration spaces and shifts the bottleneck from host-side limitations back to on-device inference. Our evaluation demonstrates that our work fundamentally expands the scale of solvable problems. On an NVIDIA A100 cluster with 64 GPUs, our work achieves up to 2.32 × end-to-end speedup over the highly-optimized NNQS-SCI baseline while preserving the same chemical accuracy. Furthermore, it demonstrates excellent distributed performance, maintaining over 90% parallel efficiency in strong scaling tests. Daran Sun, Bowen Kan, Haoquan Long, Hairui Zhao 0002, Haoxu Li, Ankang Feng, Wenjing Huang 0002, Yida Gu, Honghui Shang, Yunquan Zhang, Dingwen Tao, Ninghui Sun, Guangming Tan |
HPDC | 9 |
| 2026 | ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs
Jinwu Yang, Jiaan Wu, Xinyang Ma, Hairui Zhao 0002, Yida Gu, Yuanhong Huang, Wenjing Huang 0002, Yili Ma, Zhongzhe Hu, Shaoteng Liu, Jiaxun Lu, Guangming Tan, Dingwen Tao |
ISCA | 9 |
| 2026 | PRISM: An Efficient GPU-Based Lossy Compression Framework for Progressive Data Retrieval with Multi-Level InterpolationabstractWith the exponential growth of computing power, large-scale scientific simulations are producing massive volumes of data, leading to critical storage and I/O challenges. Error-bounded lossy compression has become one of the most effective solutions for reducing data size while preserving accuracy. Meanwhile, to achieve high-performance compression on such large datasets, leveraging GPUs has become increasingly essential. GPU-based lossy compressors deliver strong performance, but typically support only single-precision decompression, limiting their ability to meet the diverse accuracy requirements of scientific workflows. Progressive compressors can address this limitation by enabling on-demand precision retrieval. However, existing progressive lossy compressors on GPU still suffer from low throughput. To overcome these challenges, we present PRISM, a GPU-based progressive lossy compressor that achieves both high throughput and multi-precision retrieval, which introduces a high performance progressive framework that integrates the multiple interpolation predictors, efficient bitplane extraction, and an enhanced lossless compression that combines sign-absolute coding with zero-aware parallel algorithms. Evaluations on representative real-world datasets from five scientific domains show that PRISM significantly outperforms state-of-the-art progressive compressors on GPU, reducing retrieval data volume by over 15.6× and achieving up to 20.1× higher throughput on the NVIDIA H100 GPU under the same error bounds. Bing Lu 0001, Hairui Zhao 0002, Dejun Luo, Wenjing Huang 0002, Yida Gu, Jinyang Liu 0003, Guangming Tan, Dingwen Tao |
PPoPP | 5 |
| 2026 | CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model TrainingabstractAs training scales grow, collective communication libraries (CCL) increasingly face anomalies arising from complex interactions among hardware, software, and environmental factors. These anomalies typically manifest as slow/hang communication, the most frequent and time-consuming category to diagnose. However, traditional diagnostic methods remain inaccurate and inefficient, frequently requiring hours or even days for root cause analysis. To address this, we propose CCL-D, a high-precision diagnostic system designed to detect and locate slow/hang anomalies in large-scale distributed training. CCL-D integrates a rank-level real-time probe with an intelligent decision analyzer. The probe measures cross-layer anomaly metrics using a lightweight distributed tracing framework to monitor communication traffic. The analyzer performs automated anomaly detection and root-cause location, precisely identifying the faulty GPU rank. Deployed on a 4,000-GPU cluster over one year, CCL-D achieved near-complete coverage of known slow/hang anomalies and pinpointed affected ranks within 6 minutes—substantially outperforming existing solutions. Yida Gu, Fakang Wang, Jianhao Fu, Zhenhang Sun, Qianyu Zhang 0001, Hairui Zhao 0002, Wenjing Huang 0002, Jinwu Yang, Yueyuan Zhou, Qian Zhao 0021, Haoxu Li, Zhan Wang 0003, Guangming Tan, Dingwen Tao |
PPoPP | 9 |
| 2026 | KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
Xinyang Ma, Dejun Luo, Hairui Zhao 0002, Bing Lu 0001, Wenjing Huang 0002, Yida Gu, Jinyang Liu 0003, Dingwen Tao, Guangming Tan |
SIGCOMM | 6 |
| 2026 | TSUE+: An Efficient Update Framework With Swift Recycling Mechanism for Erasure-Coded Cluster File SystemsabstractCompared to replication-based storage systems, erasure-coded storage incurs significantly higher overhead during data updates. To address this issue, various parity logging methods have been proposed. Nevertheless, due to the long update path and substantial amount of random I/O involved in erasure code update processes, the resulting long latency and low through put often fail to meet the requirements of high performance applications. To address this challenge, we propose TSUE+, an efficient update framework with a swift recycling mechanism. TSUE+ divides the update process into two distinct stages: in the synchronous stage, data updates are stored in the format of replica data logs, eliminating random I/O by trading space for time; in the asynchronous stage, the recorded update logs are recycled and merged into original data and parity blocks, thereby reclaiming the storage overhead incurred in the synchronization phase. By converting random I/O operations into sequential ones based on data logs, TSUE+ effectively reduces update latency; furthermore, it significantly minimizes recycling overhead using a three-layer log structure and by leveraging the spatio-temporal locality of access patterns. We evaluated TSUE+ and other state of-the-art (SOTA) update mechanisms under diverse encoding schemes, using heterogeneous storage devices—including HDDs, SATA SSDs, NVMe SSDs, and PMEM—and multiple real-world and synthetic workloads: the MSR Cambridge trace, the Alibaba Cloud trace, the Tencent Cloud trace, and multiple synthetic worst-case workloads. Among all the platforms, TSUE+ has achieved significant performance improvements compared to other update methods, it also indicates that TSUE+ can be applied to various storage devices. Additionally, we provided percentile-based tail latency tests and update tests under the worst-case environment, which further demonstrated the ro bustness of TSUE+. Moreover, by enabling prompt log recycle and avoiding unnecessary overwrites and improving update granularity through locality-aware recycling, TSUE+ not only improves update performance but also mitigates write wear on SSD devices, thereby extending their operational lifespan. Yida Gu, Wenjing Huang 0002, Yili Ma, Dong Dai 0001, Guangming Tan, Dingwen Tao |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | TSUE: A Two-Stage Data Update Method for an Erasure Coded Cluster File SystemabstractCompared to replication-based storage systems, erasure-coded storage incurs significantly higher overhead during data updates. To address this issue, various parity logging methods have been proposed. Nevertheless, due to the long update path and substantial amount of random I/O involved in erasure code update processes, the resulting long latency and low throughput often fail to meet the requirements of high performance applications. To this end, we propose a two-stage data update method called TSUE. TSUE divides the update process into a synchronous stage that records updates in a data log, and an asynchronous stage that recycles the log in real-time. TSUE effectively reduces update latency by transforming random I/O into sequential I/O, and it significantly reduces recycle overhead by utilizing a three-layer log and the spatio-temporal locality of access patterns. In SSDs cluster, TSUE significantly improves update performance, achieving improvements of 7.6× under Ali-Cloud trace, 5× under Ten-Cloud trace, while it also extends the SSD's lifespan by up to 13× through reducing the frequencies of reads/writes and of erase operations. Yida Gu, Wenjing Huang 0002, Dong Dai 0001, Guangming Tan, Dingwen Tao |
HPDC | 4 |
| 2025 | MANS: Efficient and Portable ANS Encoding for Multi-Byte Integer Data on CPUs and GPUsabstractLossless compression is a classic technique for reducing data storage and transmission requirements. Asymmetric Numeral Systems (ANS) is a high-throughput, high-ratio lossless compression algorithm, but it lacks effective support for multi-byte data and cross-platform compatibility. To address this issue, we propose an Adaptive Data Mapping (ADM) scheme, which maps multi-byte integer data into single-byte space based on the data’s characteristics, improving the compression ratio of ANS while maintaining low encoding redundancy. We also optimize the ADM algorithm and the ANS encoder for GPU and CPU architectures, respectively, and combine them to create an efficient and portable ANS encoding method for multi-byte integer data, called MANS. Experimental results show that MANS improves compression ratios by an average of 1.24 ×, achieves 870.27MB/s throughput on CPUs, and delivers up to 288.45 × and 135.86 × speedups on an NVIDIA A100 and an AMD MI210 GPU compared to the CPU version—demonstrating its efficiency and portability across platforms. Wenjing Huang 0002, Jinwu Yang, Shengquan Yin, Haoxu Li, Yida Gu, Xing Jing, Shiyuan Fu, Hao Hu 0015, Guangming Tan, Dingwen Tao |
SC | 1 |