EDBT 2026 Demo / reviewers in the wild / expert
Xinjue Zheng
dblp:384/3993
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0007-7321-9264ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SSFusion: Tensor Fusion with Selective Sparsification for Efficient Distributed DNN Training
Zhangqiang Ming, Yuchong Hu, Yuanhao Shu, Wenxiang Zhou, Xinjue Zheng, Dan Feng 0001 |
ICDE | 6 |
| 2025 | Saving Memory via Residual Reduction for DNN Training with Compressed Communication
Xinjue Zheng, Zhangqiang Ming, Yuchong Hu, Chenxuan Yao, Wenxiang Zhou, Dan Feng 0001 |
Euro-Par (2) | 1 |
| 2025 | SAFusion: Efficient Tensor Fusion with Sparsification Ahead for High-Performance Distributed DNN TrainingabstractDistributed deep neural networks (DNN) training systems deployed across workers have been widely used in various domains, while the communication overhead among workers for synchronizing gradient tensors often becomes the performance bottleneck. To optimize communication efficiency, state-of-the-art studies often apply both two techniques: i) gradient sparsification compression, which truncates the gradient to its largest elements to reduce the communication traffic, and ii) tensor fusion, which merges multiple gradient tensors within a fusion buffer to transmit them together to reduce the communication startup overhead. However, we find that existing studies often apply gradient sparsification after tensor fusion (we call sparsification-behind tensor fusion), which leads to a fact that a lot of fused gradient tensors are missed after the sparsification, thus impairing the convergence performance. Zhangqiang Ming, Yuchong Hu, Xinjue Zheng, Wenxiang Zhou, Dan Feng 0001 |
HPDC | 3 |
| 2025 | LowDiff: Efficient Frequent Checkpointing via Low-Cost Differential for High-Performance Distributed Training SystemsabstractDistributed training of large deep-learning models often leads to failures, so checkpointing is commonly employed for recovery. State-of-the-art studies focus on frequent checkpointing for fast recovery from failures. However, it generates numerous checkpoints, incurring substantial costs and thus degrading training performance. Recently, differential checkpointing has been proposed to reduce costs, but it is limited to recommendation systems, so its application to general distributed training systems remains unexplored. Chenxuan Yao, Yuchong Hu, Xinjue Zheng, Wenxiang Zhou |
SC | 5 |
| 2024 | ADTopk: All-Dimension Top-k Compression for High-Performance Data-Parallel DNN TrainingabstractData-parallel deep neural networks (DNN) training systems deployed across nodes have been widely used in various domains, while the system performance is often bottlenecked by the communication overhead among workers for synchronizing gradients. Top-k sparsification compression is the de facto approach to alleviate the communication bottleneck, which truncates the gradient to its largest k elements before sending it to other nodes. Zhangqiang Ming, Yuchong Hu, Wenxiang Zhou, Xinjue Zheng, Chenxuan Yao, Dan Feng 0001 |
HPDC | 4 |