EDBT 2026 Demo / reviewers in the wild / expert
Yuanwei Wang
dblp:180/5118
· DBLP profile ↗
9ranked-venue papers
1as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MEPipe: Democratizing LLM Training with Memory-Efficient Slice-Level Pipeline Scheduling on Cost-Effective AcceleratorsabstractThe training of large language models (LLMs) typically needs costly GPUs, such as NVIDIA A100 or H100. They possess substantial high-bandwidth on-chip memory and rapid interconnects like NVLinks. The exorbitant expenses associated with LLM training pose not just an economic challenge but also a societal one, as it restricts the ability to train LLMs from scratch to a selected few organizations. Zhenbo Sun, Shengqi Chen 0001, Yuanwei Wang, Jian Sha, Guanyu Feng |
EuroSys | 3 |
| 2024 | AdaPipe: Optimizing Pipeline Parallelism with Adaptive Recomputation and PartitioningabstractLarge language models (LLMs) have demonstrated powerful capabilities, requiring huge memory with their increasing sizes and sequence lengths, thus demanding larger parallel systems. The broadly adopted pipeline parallelism introduces even heavier and unbalanced memory consumption. Recomputation is a widely employed technique to mitigate the problem but introduces extra computation overhead. Zhenbo Sun, Huanqi Cao, Yuanwei Wang, Guanyu Feng, Shengqi Chen 0001, Haojie Wang 0004 |
ASPLOS (3) | 3 |
| 2023 | TriCache: A User-Transparent Block Cache Enabling High-Performance Out-of-Core Processing with In-Memory ProgramsabstractOut-of-core systems rely on high-performance cache sub-systems to reduce the number of I/O operations. Although the page cache in modern operating systems enables transparent access to memory and storage devices, it suffers from efficiency and scalability issues on cache misses, forcing out-of-core systems to design and implement their own cache components, which is a non-trivial task. This study proposes TriCache, a cache mechanism that enables in-memory programs to efficiently process out-of-core datasets without requiring any code rewrite. It provides a virtual memory interface on top of the conventional block interface to simultaneously achieve user transparency and sufficient out-of-core performance. A multi-level block cache design is proposed to address the challenge of per-access address translations required by a memory interface. It can exploit spatial and temporal localities in memory or storage accesses to render storage-to-memory address translation and page-level concurrency control adequately efficient for the virtual memory interface. Our evaluation shows that in-memory systems operating on top of TriCache can outperform Linux OS page cache by more than one order of magnitude, and can deliver performance comparable to or even better than that of corresponding counterparts designed specifically for out-of-core scenarios. Guanyu Feng, Huanqi Cao, Xiaowei Zhu 0001, Bowen Yu 0003, Yuanwei Wang, Zixuan Ma, Shengqi Chen 0001 |
ACM Trans. Storage | 5 |
| 2022 | TriCache: A User-Transparent Block Cache Enabling High-Performance Out-of-Core Processing with In-Memory Programs
Guanyu Feng, Huanqi Cao, Xiaowei Zhu 0001, Bowen Yu 0003, Yuanwei Wang, Zixuan Ma, Shengqi Chen 0001 |
OSDI | 5 |
| 2022 | Scaling graph traversal to 281 trillion edges with 40 million coresabstractGraph processing, especially high-performance graph traversal, plays a more and more important role in data analytics. The successor of Sunway TaihuLight, New Sunway, is equipped with nearly 10 PB memory and over 40 million cores, which brings the opportunity to process hundreds of trillions of edges graphs. However, the graph with an unprecedented scale also brings severe performance challenges, including load imbalance, poor locality, and irregular access of graph traversal workload. Huanqi Cao, Yuanwei Wang, Haojie Wang 0004, Heng Lin, Zixuan Ma, Wanwang Yin |
PPoPP | 2 |
| 2022 | BaGuaLu: targeting brain scale pretrained models with over 37 million coresabstractLarge-scale pretrained AI models have shown state-of-the-art accuracy in a series of important applications. As the size of pretrained AI models grows dramatically each year in an effort to achieve higher accuracy, training such models requires massive computing and memory capabilities, which accelerates the convergence of AI and HPC. However, there are still gaps in deploying AI applications on HPC systems, which need application and system co-design based on specific hardware features. Zixuan Ma, Jiaao He, Jiezhong Qiu, Huanqi Cao, Yuanwei Wang, Zhenbo Sun, Liyan Zheng 0001, Haojie Wang 0004, Shizhi Tang, Tianyu Zheng, Junyang Lin, Guanyu Feng, Zeqiang Huang, Aohan Zeng, Jianwei Zhang 0012, Runxin Zhong, Tianhui Shi, Jie Tang 0001, Hongxia Yang, Xin Liu 0086, Jidong Zhai |
PPoPP | 5 |
| 2022 | Scaling Graph 500 SSSP to 140 Trillion Edges with over 40 Million CoresabstractThe SSSP kernel was first introduced into the Graph 500 benchmark in 2017. However, there has been no result from a full-scale world-top supercomputer. The primary reason is the poor work-inefficiency of existing algorithms at large scales. In this paper, we propose an SSSP implementation for The Newest Generation Sunway Supercomputer,including an SSSP algorithm to achieve work-efficiency, along with an adaptive dense/sparse-mode selection approach to achieve communication-efficiency. Our implementation reaches 7638 GTEPS, with 103158 processors (over 40 million cores), and achieves 3.7× in performance and 512× in graph size compared with the current top one on the Graph 500 SSSP list. Based on our experience of running extreme-scale SSSP, we uncover the root cause of its poor scalability: the weight distribution allows edges with weights close to zero, making the SSSP tree deeper on larger graphs. We further explore a scalability-friendly weight distribution by setting a non-zero lower bound to the edge weights. Yuanwei Wang, Huanqi Cao, Zixuan Ma, Wanwang Yin |
SC | 1 |
| 2021 | Encouraging Compiler Optimization Practice for Undergraduate Students through CompetitionabstractAI and other emerging applications demand domain-specific architectures which require compiler techniques such as back-end generation for different architectures and optimizations. However, traditional undergraduate compiler courses emphasize the front-end, while code generation and optimization are rarely involved. To motivate universities to have more industry-friendly compiler courses, we have designed a national compiler design competition for undergraduates to include compiler techniques beyond parsing. Moreover, we provided reliable, continuous cloud storage and an online evaluation platform for distributed competitors. In the 9-week Competition in 2020, each team (up to 4 students) was required to implement a compiler for a given SysY language and given target hardware (Raspberry Pi 4B). The performance was evaluated by executing code generated by the compiler on real hardware. Finally,21 of 72 teams successfully passed all functional test cases; 12 of 21 teams implemented optimizations showing significant speedup over gcc -O0; furthermore, compilers of the top 3 teams performed better than gcc -O2 on the given 10 performance test cases. Some advanced optimization techniques, such as multithreading and SIMD, were used by some teams. This paper summarizes the competition and further thoughts on compiler courses. Yu Zhang 0086, Chunming Hu, Mingliang Zeng, Yitong Huang, Yuanwei Wang |
ITiCSE (1) | 6 |
| 2016 | Nonlinear 2D shape registration via thin-plate spline and Lie group representation
Shihui Ying, Yuanwei Wang, Zhijie Wen, Yuping Lin |
Neurocomputing | 2 |