Dong Xu 0024

dblp:09/3493-24 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0002-4093-628XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2026 CXL-CCL: Inter-Node Collective GPU-Communication Using a CXL Shared Memory Pool
abstract
Large language models (LLMs) training or inference across multiple nodes introduces significant pressure on GPU memory and interconnect bandwidth. The Compute Express Link (CXL) shared memory pool offers a scalable solution by enabling memory sharing across nodes, reducing over-provisioning and improving resource utilization. We propose CXL-CCL, a collective communication library, leveraging the CXL shared memory pool to support cross-node GPU operations without relying on traditional RDMA-based networking. Our design addresses the challenges in synchronization, data interleaving, and communication parallelization faced by using the CXL shared memory pool for collective communications. Evaluating on multiple nodes with a TITAN-II CXL switch and six Micron CZ120 memory cards, we show that CXL-CCL achieves highly efficient collective operations across hosts, demonstrating CXL’s potential for scalable, memory-centric GPU communication. Our evaluation demonstrates that CXL-CCL achieves average performance improvements of 1.34 × for AllGather, 1.84 × for Broadcast, 1.94 × for Gather, and 1.07 × for Scatter, compared to the original RDMA-based implementation over 200 Gbps InfiniBand. In addition, an LLM training case study shows 1.11 × speedup compared with the InfiniBand while reducing interconnect hardware cost by 2.75 ×.
Dong Xu 0024, Han Meng, Dengcheng Zhu, Liguang Xie, Wu Xiang, Henry Hu, Hui Zhang 0033, Dong Li 0001
ICS1
2024 MTM: Rethinking Memory Profiling and Migration for Multi-Tiered Large Memory
abstract
Multi-terabyte large memory systems are often characterized by more than two memory tiers with different latency and bandwidth. Multi-tiered large memory systems call for rethinking of memory profiling and migration because of the unique problems unseen in the traditional memory systems with smaller capacity and fewer tiers. We develop MTM, an application-transparent Multi-Tiered Memory management framework, based on three principles: (1) connecting the control of profiling overhead with the profiling mechanism for high-quality profiling; (2) building a universal page migration policy on the complex multi-tiered memory for high performance; and (3) introducing huge page awareness. We evaluate MTM using common big-data applications with realistic working sets (hundreds of GB to 1 TB). MTM outperforms seven solutions by up to 42% (17% on average).
Jie Ren 0015, Dong Xu 0024, Junhee Ryu, Kwangsik Shin, Daewoo Kim, Dong Li 0001
EuroSys2
2024 Enabling Large Dynamic Neural Network Training with Learning-based Memory Management
abstract
Dynamic neural network (DyNN) enables high computational efficiency and strong representation capability. However, training DyNN can face a memory capacity problem because of increasing model size or limited GPU memory capacity. Managing tensors to save GPU memory is challenging, because of the dynamic structure of DyNN. We present DyNN-Offload, a memory management system to train DyNN. DyNN-Offload uses a learned approach (using a neural network called the pilot model) to increase predictability of tensor accesses to facilitate memory management. The key of DyNN-Offload is to enable fast inference of the pilot model in order to reduce its performance overhead, while providing high inference (or prediction) accuracy. DyNNOffload reduces input feature space and model complexity of the pilot model based on a new representation of DyNN; DyNNOffload converts the hard problem of making prediction for individual operators into a simpler problem of making prediction for a group of operators in DyNN. DyNN-Offload enables 8 × larger DyNN training on a single GPU compared with using PyTorch alone (unprecedented with any existing solution). Evaluating with AlphaFold (a production-level, large-scale DyNN), we show that DyNN-Offload outperforms unified virtual memory (UVM) and dynamic tensor rematerialization (DTR), the most advanced solutions to save GPU memory for DyNN, by 3 × and 2.1 × respectively in terms of maximum batch size.
Jie Ren 0015, Dong Xu 0024, Shuangyan Yang, Christian Navasca, Chenxi Wang 0005, Guoqing Harry Xu, Dong Li 0001
HPCA2
2024 Efficient Tensor Offloading for Large Deep-Learning Model Training based on Compute Express Link
abstract
The deep learning models (DL) are becoming bigger, easily beyond the memory capacity of a single accelerator. The recent progress in large DL training utilizes CPU memory as an extension of accelerator memory and offloads tensors to CPU memory to save accelerator memory. This solution transfers tensors between the two memories, creating a major performance bottleneck. We identify two problems during tensor transfers: (1) the coarse-grained tensor transfer creating difficulty in hiding transfer overhead, and (2) the redundant transfer that unnecessarily migrates value-unchanged bytes from CPU to accelerator. We introduce a cache coherence interconnect based on Compute Express Link (CXL) to build a cache coherence domain between CPU memory and accelerator memory. By slightly extending CXL to support an update cache-coherence protocol and avoiding unnecessary data transfers, we reduce training time by $33.7 \%$ (up to $55.4 \%$) without changing model convergence and accuracy, compared with the state-of-the-art work in DeepSpeed [62].
Dong Xu 0024, Kwangsik Shin, Daewoo Kim, Hyeran Jeon, Dong Li 0001
SC1
2024 FlexMem: Adaptive Page Profiling and Migration for Tiered Memory
Dong Xu 0024, Junhee Ryu, Kwangsik Shin, Dong Li 0001
USENIX ATC1