VLDB 2026 Research / reviewers in the wild / expert
Tongrui Liu
dblp:392/8896
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UPServe: Backend Agnostic Proxy for Black-box Heterogeneous LLM Scheduling
Haorui Wan, Chenyang Hei, Fuliang Li, Chengxi Gao, Yuhan Jia, Tongrui Liu, Xingwei Wang 0001 |
IWQoS | 6 |
| 2026 | HeteCCL: Synthesizing Near-Optimal Collective Communication Schedules for Heterogeneous GPU Clusters
Chenyang Hei, Jiamin Cao, Chengxi Gao, Xiuzhu Sha, Tongrui Liu, Dengke Zhang, Ennan Zhai, Xingwei Wang 0001 |
NSDI | 6 |
| 2026 | SeaAnchor-GS: Identity-Anchored Gaussian Splatting for High-Fidelity Dynamic Underwater Scene ReconstructionabstractDynamic underwater novel-view synthesis remains challenging because refraction, scattering, and particle interference undermine correspondence reliability, geometric consistency, and appearance stability. We propose SeaAnchor-GS, a robust dynamic 3D Gaussian splatting framework for underwater scene reconstruction. Each Gaussian is associated with a persistent identity embedding and deformed via identity–time conditioning, which improves deformation estimation under unstable underwater observations. A dual-branch residual dynamics module captures both dominant motion and fine-scale variations, while confidence-guided sampling, progressive deformation activation, and neighborhood-consistency regularization enhance optimization robustness and model compactness. Extensive experiments on dynamic underwater benchmarks show consistent improvements in reconstruction fidelity and perceptual quality, with a favorable balance between quality, efficiency, and representation compactness. Additional results on a static underwater benchmark suggest that the proposed representation remains competitive beyond the dynamic setting. Yaoming Zhuang, Tongrui Liu, Yifan Chao, Hao Wu 0064, Chengdong Wu 0001, Zhanlin Liu |
IEEE Signal Process. Lett. | 2 |
| 2025 | Canvas: Scalable Collective Communication Scheduling for Large-Scale GPU ClustersabstractState-of-the-art deep learning models rely on large GPU clusters and various parallelism strategies, which in turn depend on collective communication (CC) operators to synchronize data. While vendor libraries (e.g., NCCL, RCCL) provide standard CC algorithms, they often suffer from bandwidth bottlenecks in imbalanced topologies. Recent synthesis-based methods improve performance but face three key limitations: poor scalability due to the combinatorial explosion of scheduling space, lack of support for multistage execution, and suboptimal communication throughput. We propose Canvas, a scalable and near-optimal CC scheduling framework that addresses these challenges. Canvas introduces: (1) Hierarchical synthesis to decompose the global scheduling problem into tractable subproblems for scalability. (2) Collective decomposition to enable structured, multi-stage algorithm generation. (3) Cross-micro-batch pipeline scheduling to parallelize communication across micro-batches and maximize link utilization. Evaluations show that Canvas achieves up to 1.98× bandwidth speedup over TACCL and 3.56× over TE-CCL, and synthesizes algorithms for 512-GPU topologies within 1.77 hours, whereas TACCL fails to produce results within 24 hours. Chenyang Hei, Fuliang Li, Chengxi Gao, Tongrui Liu, Xiuzhu Sha, Xingwei Wang 0001 |
ICNP | 5 |
| 2025 | TuCCL: Tailored and Unified Configuration Optimizations for High-Performance Collective Communication LibraryabstractModern distributed training systems face escalating communication bottlenecks as GPU clusters scale to accommodate large models. While collective communication libraries and automated synthesizers address algorithmic efficiency, they suffer from three critical limitations including labor-intensive manual intervention requirement, overreliance on predefined input optimization, and suboptimal isolated configuration optimization. To solve these problems, we present TuCCL, a systematic framework that co-optimizes communication algorithms and runtime parameters through three innovations: Topology-Aware Sketch Generation that automatically produces high-performance primitives, Hierarchical Configuration Optimization modeling nonlinear parameter-performance relationships, and Multi-phase Resource-Aware Configuration Optimization enabling joint configuration tuning with adaptive search space pruning. Evaluations demonstrate TuCCL’s superiority over state-of-the-art systems with 1.75x–11.49x bandwidth improvements for AllGather/AllReduce on NVIDIA V100/A100 clusters, 90.2% faster configuration search than grid methods, and 1.22x-2.52x end-to-end training speedups across diverse model scales. Chenyang Hei, Fuliang Li, Tongrui Liu, Chengxi Gao, Xiuzhu Sha, Xingwei Wang 0001 |
ICNP | 4 |
| 2025 | ResCCL: Resource-Efficient Scheduling for Collective CommunicationabstractAs distributed deep learning training (DLT) systems scale, collective communication has become a significant performance bottleneck. While current approaches optimize bandwidth utilization and task completion time, existing communication libraries (CCLs) backends fail to efficiently manage GPU resources during algorithm execution, limiting the performance of advanced algorithms. This paper proposes ResCCL, a novel CCL backend designed for Resource-Efficient Scheduling to address key limitations in current systems. ResCCL enhances execution efficiency by optimizing scheduling at the primitive level (e.g., send and recvReduceCopy), enabling flexible thread block (TB) allocation, and generating lightweight communication kernels to minimize runtime overhead. Our approach tackles the global scheduling problem, reduces idle TB resources, and enhances communication bandwidth. Evaluation results demonstrate that ResCCL achieves up to 2.5× improvement in bandwidth performance compared to both NCCL and MSCCL. It reduces SM resource overhead by 77.8% and increases TB utilization by 41.6% while running the same algorithms. In end-to-end DLT, ResCCL boosts Megatron's throughput by up to 39%. Tongrui Liu, Chenyang Hei, Fuliang Li, Chengxi Gao, Jiamin Cao, Ennan Zhai, Xingwei Wang 0001 |
SIGCOMM | 1 |
| 2024 | Financial Fraud Defense Strategy based on Gradient Compensated Asynchronous Federated LearningabstractAsynchronous federated learning (AFL) allows participants to immediately submit trained models without waiting for other participants, building upon the federated learning (FL). Due to privacy concerns, FL is more susceptible to financial fraud, compounded by the gradient delay issues introduced by asynchronous submissions, making defense against financial fraud more challenging. Motivated by the above finding, we propose a secure and privacy-preserving AFL defense method for image datasets with an implanted backdoor via filtering redundant neurons (BDAFL), enhancing its resilience against such attacks without compromising privacy. We utilize gradient compensation to mitigate the impact of delays introduced by asynchrony. To counter financial fraud, we employ an anomaly detection algorithm based on neurons’ weights and ensemble distillation to eliminate the affected neurons implanted with the backdoor, rendering the attack ineffective. Extensive experiments demonstrate the effectiveness and superiority of our approach. Tongrui Liu, Yizhi Zhou, Zhipeng Song, Xibei Jia, Heng Qi |
ICPADS | 1 |