EDBT 2026 Demo / reviewers in the wild / expert
Zhiyuan Ai
dblp:203/1934
· DBLP profile ↗
3ranked-venue papers
2as first author
1since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Storage systems · 52% Hardware accelerators and domain-specific architectures · 37% GPUs and heterogeneous computing · 11% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 50% Deep learning architectures and training · 50% | |
| Databases, data mining, and information retrieval
1 paper |
Graph data management · 100% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training › mixture of experts
mixture-of-experts inference |
0.9 | 1 | 2025 | KTransformers: Unleashing the Full Potential of CPU/GPU Hybrid Inference for MoE Models · SOSP 2025 |
Machine learning › Efficient and distributed learning
model inference |
0.9 | 1 | 2025 | KTransformers: Unleashing the Full Potential of CPU/GPU Hybrid Inference for MoE Models · SOSP 2025 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.9 | 1 | 2025 | KTransformers: Unleashing the Full Potential of CPU/GPU Hybrid Inference for MoE Models · SOSP 2025 |
Storage systems › out-of-core computation
out-of-core graph processing |
0.7 | 2 | 2019 | Clip: A Disk I/O Focused Parallel Out-of-Core Graph Processing System · IEEE Trans. Parallel Distributed Syst. 2019 Squeezing out All the Value of Loaded Data: An Out-of-core Graph Processing System with Reduced Disk I/O · USENIX ATC 2017 |
Graph data management › graph processing
out-of-core graph processing |
0.4 | 1 | 2019 | Clip: A Disk I/O Focused Parallel Out-of-Core Graph Processing System · IEEE Trans. Parallel Distributed Syst. 2019 |
Graph data management › graph algorithms
parallel graph algorithms |
0.4 | 1 | 2019 | Clip: A Disk I/O Focused Parallel Out-of-Core Graph Processing System · IEEE Trans. Parallel Distributed Syst. 2019 |
Storage systems › i/o optimization
disk i/o optimization |
0.4 | 1 | 2019 | Clip: A Disk I/O Focused Parallel Out-of-Core Graph Processing System · IEEE Trans. Parallel Distributed Syst. 2019 |
GPUs and heterogeneous computing
CPU-GPU heterogeneous computing |
0.3 | 1 | 2025 | KTransformers: Unleashing the Full Potential of CPU/GPU Hybrid Inference for MoE Models · SOSP 2025 |
Storage systems › flash and SSD › solid-state drive
NVMe SSD |
0.1 | 1 | 2019 | Clip: A Disk I/O Focused Parallel Out-of-Core Graph Processing System · IEEE Trans. Parallel Distributed Syst. 2019 |
Storage systems › i/o optimization
disk i/o reduction |
0.1 | 1 | 2017 | Squeezing out All the Value of Loaded Data: An Out-of-core Graph Processing System with Reduced Disk I/O · USENIX ATC 2017 |
Methods — techniques the papers use, named apart from their topics
CPU/GPU hybrid inference · 1.7semi-external mode · 0.8disk i/o locality optimization · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | KTransformers: Unleashing the Full Potential of CPU/GPU Hybrid Inference for MoE ModelsabstractDue to the sparse nature of Mixture-of-Experts (MoE) models, they are particularly suitable for hybrid CPU/GPU inference, especially in low-concurrency scenarios. This hybrid approach leverages both the large, cost-effective memory capacity of CPU/DRAM and the high bandwidth of GPU/VRAM. However, existing hybrid solutions remain bottlenecked by CPU computation limits and CPU-GPU synchronization overheads, severely restricting their ability to efficiently run state-of-the-art large MoE models, such as the 671B DeepSeek-V3/R1. Hongtao Chen, Weiyu Xie, Boxin Zhang, Jingqi Tang, Shaoyuan Chen, Ziwei Yuan, Chengyu Qiu, Yuening Zhu, Qingliang Ou, Jiaqi Liao, Xianglin Chen, Zhiyuan Ai, Yongwei Wu 0001 |
SOSP | 15 |
| 2019 | Clip: A Disk I/O Focused Parallel Out-of-Core Graph Processing SystemabstractExisting parallel out-of-core graph processing systems focus on improving disk I/O locality, which leads to restrictions on their programming models. Although improving the locality, these constraints also restrict the expressiveness and hence only sub-optimal algorithms are supported. These sub-optimal algorithms typically incur sequential, but much larger, amount of disk I/O. In this paper, we explore a fundamentally different tradeoff: less total amount of I/O rather than better locality. We show that out-of-core graph processing systems uniquely provide the opportunities to lift the restrictions of the programming model in a feasible manner. To demonstrate the ideas, we build Clip, which enables more efficient algorithms that require much less amount of total disk I/O. Our experiments show that the algorithms that can be only implemented in Clip are much faster than the original disk-locality-optimized algorithms. We also further extend our technique's scope of application by providing a semi-external mode. Our analysis and evaluation demonstrate that semi-external is not only feasible for many cases, but also be able to deliver a significant speedup for important graph applications. Moreover, we further improve the performance of originally supported applications by designing more optimizations and evaluate our system on NVMe SSD. Zhiyuan Ai, Yongwei Wu 0001, Xuehai Qian, Kang Chen 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | Squeezing out All the Value of Loaded Data: An Out-of-core Graph Processing System with Reduced Disk I/O
Zhiyuan Ai, Yongwei Wu 0001, Xuehai Qian, Kang Chen 0001 |
USENIX ATC | 1 |