Yueran Tang

dblp:413/4442 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0007-8049-9233ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
memory-efficient training
1.012026
STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal Planning · EuroSys 2026
Memory systems › memory management
memory allocation
1.012026
STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal Planning · EuroSys 2026
Memory systems
memory management
1.012026
STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal Planning · EuroSys 2026
Machine learning › Efficient and distributed learning
distributed training
0.312026
STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal Planning · EuroSys 2026

Methods — techniques the papers use, named apart from their topics

spatio-temporal planning · 2.0
YearPublicationVenuePosition
2026 STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal Planning
abstract
The rapid scaling of large language models (LLMs) has significantly increased GPU memory pressure, which is further aggravated by training optimization techniques such as virtual pipeline and recomputation that disrupt tensor lifespans and introduce considerable memory fragmentation. Such fragmentation stems from the use of online GPU memory allocators in popular deep learning frameworks like PyTorch, which disregard tensor lifespans. As a result, this inefficiency can waste as much as 43% of memory and trigger out-of-memory errors, undermining the effectiveness of optimization methods.
Zixiao Huang 0001, Hao Lin 0005, Chunyang Zhu, Yueran Tang, Quanlu Zhang, Zhenhua Li 0001, Shengen Yan, Zhenhua Zhu 0002, Guohao Dai 0001, Yu Wang 0002
EuroSys5