Xiandong Lu

dblp:440/1423 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
0009-0007-0105-8561ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 61% Distributed systems · 39%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems › fault tolerance
checkpointing
1.012026
Frequent Checkpointing through Mergeable Delta Compression and Semi-Reliable Transmission · INFOCOM 2026
Storage systems
data compression
1.012026
Frequent Checkpointing through Mergeable Delta Compression and Semi-Reliable Transmission · INFOCOM 2026
Storage systems › data compression
delta compression
1.012026
Frequent Checkpointing through Mergeable Delta Compression and Semi-Reliable Transmission · INFOCOM 2026
Distributed systems
fault tolerance
0.312026
Frequent Checkpointing through Mergeable Delta Compression and Semi-Reliable Transmission · INFOCOM 2026

Methods — techniques the papers use, named apart from their topics

semi-reliable transmission · 1.0mergeable delta compression · 1.0
YearPublicationVenuePosition
2026 Frequent Checkpointing through Mergeable Delta Compression and Semi-Reliable Transmission
Jie Gui, Xiandong Lu, Qun Huang 0001
INFOCOM2
2026 Theseus: Runtime-Adaptive GPU Collective Communication with Hot-Swappable Schedules
abstract
Current GPU Collective Communication Libraries (CCLs) employ predefined schedules optimized for stable environments. Their supported schedules and selection logic are fixed at communicator initialization, which fails to account for evolving runtime conditions, such as workload characteristics and hardware health status. Consequently, long-running GPU jobs experience suboptimal performance after hours or days of execution, which translates into longer job completion times and wasted GPU cluster resources. To address this problem, we present Theseus, a novel CCL backend that provides schedule-level runtime adaptivity. It admits user-defined schedules and selection policies. As runtime conditions change, Theseus selects suitable schedules using cluster-wide runtime attributes beyond CCL-internal metrics. Moreover, it hot-swaps from the previous schedule consistently across GPUs with low overhead. Theseus acts as a drop-in replacement to facilitate integration. We evaluate Theseus extensively on various GPU workloads with intuitive policies. Compared with NCCL, Theseus achieves up to 1.61X speedup of communication time in stable environments and 2.46X in dynamic environments. It improves end-to-end job completion time by up to 1.84X while incurring comparable or lower overhead.
Rui Ding 0014, Xiandong Lu, Xunpeng Liu, Xuran Hao, Houyuan Zhu, Anyi Xu, Sinuo Cao, Haifeng Sun 0004, Qun Huang 0001, Jiamin Cao
SIGCOMM2