EDBT 2026 Demo / reviewers in the wild / expert
Wei Cao 0010
dblp:54/6265-10
· DBLP profile ↗
3ranked-venue papers
0as first author
2since 2021 · last 2022
0000-0003-1620-2419ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Memory systems · 75% Hardware accelerators and domain-specific architectures · 25% | |
| Software engineering, system software, and programming languages
1 paper |
Runtime systems and virtual machines · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator |
0.6 | 1 | 2022 | An Application-oblivious Memory Scheduling System for DNN Accelerators · ACM Trans. Archit. Code Optim. 2022 |
Memory systems › memory management
DNN training memory management |
0.6 | 1 | 2022 | An Application-oblivious Memory Scheduling System for DNN Accelerators · ACM Trans. Archit. Code Optim. 2022 |
Memory systems › memory controller
memory scheduling |
0.6 | 1 | 2022 | An Application-oblivious Memory Scheduling System for DNN Accelerators · ACM Trans. Archit. Code Optim. 2022 |
Memory systems › virtual memory management
memory swapping |
0.6 | 1 | 2022 | An Application-oblivious Memory Scheduling System for DNN Accelerators · ACM Trans. Archit. Code Optim. 2022 |
Runtime systems and virtual machines
runtime memory management |
0.2 | 1 | 2022 | An Application-oblivious Memory Scheduling System for DNN Accelerators · ACM Trans. Archit. Code Optim. 2022 |
Methods — techniques the papers use, named apart from their topics
memory behavior tracing · 1.1application-oblivious scheduling · 1.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | An Application-oblivious Memory Scheduling System for DNN AcceleratorsabstractDeep Neural Networks (DNNs) tend to go deeper and wider, which poses a significant challenge to the training of DNNs, due to the limited memory capacity of DNN accelerators. Existing solutions for memory-efficient DNN training are densely coupled with the application features of DNN workloads, e.g., layer structures or computational graphs of DNNs are necessary for these solutions. This would result in weak versatility for DNNs with sophisticated layer structures or complicated computation graphs. These schemes usually need to be re-implemented or re-adapted due to the new layer structures or the unusual operators in the computational graphs introduced by these DNNs. In this article, we review the memory pressure issues of DNN training from the perspective of runtime systems and model the memory access behaviors of DNN workloads. We identify the iterative, regularity , and extremalization properties of memory access patterns for DNN workloads. Based on these observations, we propose AppObMem, an application-oblivious memory scheduling system. AppObMem automatically traces the memory behaviors of DNN workloads and schedules the memory swapping to reduce the memory pressure of the device accelerators without the perception of high-level information of layer structures or computation graphs. Evaluations on a variety of DNN models show that, AppObMem obtains 40–60% memory savings with acceptable performance loss. AppObMem is also competitive with other open sourced SOTA schemes. Jiansong Li, Xueying Wang 0003, Xiaobing Chen, Guangli Li, Peng Zhao 0008, Xianzhi Yu, Yongxin Yang, Wei Cao 0010, Lei Liu 0030, Xiaobing Feng 0002 |
ACM Trans. Archit. Code Optim. | 9 |
| 2021 | Pinpointing the Memory Behaviors of DNN TrainingabstractThe training of deep neural networks (DNNs) is usually memory-hungry due to the limited device memory capacity of DNN accelerators. Characterizing the memory behaviors of DNN training is critical to optimize the device memory pressures. In this work, we pinpoint the memory behaviors of each device memory block of GPU during training by instrumenting the memory allocators of the runtime system. Our results show that the memory access patterns of device memory blocks are stable and follow an iterative fashion. These observations are useful for the future optimization of memory-efficient training from the perspective of raw memory access patterns. Jiansong Li, Guangli Li, Peng Zhao 0008, Xueying Wang 0003, Xiaobing Chen, Xianzhi Yu, Yongxin Yang, Zihan Jiang 0006, Wei Cao 0010, Lei Liu 0030, Xiaobing Feng 0002 |
ISPASS | 10 |
| 2020 | Compiler-Assisted Operator Template Library for DNN Accelerators
Jiansong Li, Wei Cao 0010, Guangli Li, Xueying Wang 0003, Lei Liu 0030, Xiaobing Feng 0002 |
NPC | 2 |