Zhirun Yue

dblp:430/6856 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
0009-0005-6357-7995ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 46% Memory systems · 30% High-performance computing · 23%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
large-scale training
1.012026
Efficient Tensor Offloading Based on CXL Memory Pool for Extreme Scale Deep Learning · IEEE Trans. Computers 2026
Hardware accelerators and domain-specific architectures
memory-constrained training
1.012026
Efficient Tensor Offloading Based on CXL Memory Pool for Extreme Scale Deep Learning · IEEE Trans. Computers 2026
Memory systems
memory disaggregation
1.012026
Efficient Tensor Offloading Based on CXL Memory Pool for Extreme Scale Deep Learning · IEEE Trans. Computers 2026
Hardware accelerators and domain-specific architectures › accelerator offloading
tensor offloading
1.012026
Efficient Tensor Offloading Based on CXL Memory Pool for Extreme Scale Deep Learning · IEEE Trans. Computers 2026
Memory systems
hybrid memory
0.312026
Efficient Tensor Offloading Based on CXL Memory Pool for Extreme Scale Deep Learning · IEEE Trans. Computers 2026

Methods — techniques the papers use, named apart from their topics

hybrid object-based memory management · 1.0dynamic adaptive pipelining · 1.0NUMA-aware memory allocation · 1.0
YearPublicationVenuePosition
2026 CxlRatioOpt : A Predictive Optimizer for Performance Trade-offs in CXL-Tiered Memory
Bowen Wang 0013, Haiyuan Wan, Zhirun Yue, Fangming Liu
ISCAS6
2026 Efficient Tensor Offloading Based on CXL Memory Pool for Extreme Scale Deep Learning
abstract
The exponential growth of deep learning models imposes severe memory constraints on GPUs, significantly increasing training costs. The prevailing solution for addressing memory constraints is tensor offloading, exemplified by ZeRO Infinity, which leverages GPUs, CPUs, and NVMe SSDs to enable large-scale model training. However, ZeRO-Infinity faces significant performance bottlenecks due to memory access imbalance, coarse-grained tensor transfers, NVMe bandwidth and latency limits, and software complexity. Compute Express Link (CXL) emerges as a promising technology for building disaggregated memory pools, yet its integration into large-scale training remains underexplored.This paper introduces an efficient CXL memory pool into the system for tensor offloading and leveraging CXL protocol features for hardware acceleration. The proposed design incorporates NUMA-aware memory allocation and a Dynamic Adaptive Pipelining (DAP) strategy to enhance communication–computation overlap, along with a hybrid object-based memory management scheme to mitigate fragmentation and improve allocation efficiency. Experimental evaluation on an 8- GPU system with CXL Type-3 expansion cards demonstrates up to 72.7% throughput improvement, 62.9% latency reduction, and support for training 1.14× larger models compared with SSD based ZeRO-Infinity. This is the first work to integrate CXL with ZeRO-Infinity for large-scale training, offering practical insights for future CXL-based heterogeneous systems.
Dongwei Xu, Fangming Liu, Bowen Wang 0013, Haiyuan Wan, Zhirun Yue
IEEE Trans. Computers10