Zhiyuan Ai

dblp:203/1934 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
1since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Storage systems · 52% Hardware accelerators and domain-specific architectures · 37% GPUs and heterogeneous computing · 11%
Artificial intelligence
1 paper
Efficient and distributed learning · 50% Deep learning architectures and training · 50%
Databases, data mining, and information retrieval
1 paper
Graph data management · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › mixture of experts
mixture-of-experts inference
0.912025
KTransformers: Unleashing the Full Potential of CPU/GPU Hybrid Inference for MoE Models · SOSP 2025
Machine learning › Efficient and distributed learning
model inference
0.912025
KTransformers: Unleashing the Full Potential of CPU/GPU Hybrid Inference for MoE Models · SOSP 2025
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.912025
KTransformers: Unleashing the Full Potential of CPU/GPU Hybrid Inference for MoE Models · SOSP 2025
Storage systems › out-of-core computation
out-of-core graph processing
0.722019
Clip: A Disk I/O Focused Parallel Out-of-Core Graph Processing System · IEEE Trans. Parallel Distributed Syst. 2019
Squeezing out All the Value of Loaded Data: An Out-of-core Graph Processing System with Reduced Disk I/O · USENIX ATC 2017
Graph data management › graph processing
out-of-core graph processing
0.412019
Clip: A Disk I/O Focused Parallel Out-of-Core Graph Processing System · IEEE Trans. Parallel Distributed Syst. 2019
Graph data management › graph algorithms
parallel graph algorithms
0.412019
Clip: A Disk I/O Focused Parallel Out-of-Core Graph Processing System · IEEE Trans. Parallel Distributed Syst. 2019
Storage systems › i/o optimization
disk i/o optimization
0.412019
Clip: A Disk I/O Focused Parallel Out-of-Core Graph Processing System · IEEE Trans. Parallel Distributed Syst. 2019
GPUs and heterogeneous computing
CPU-GPU heterogeneous computing
0.312025
KTransformers: Unleashing the Full Potential of CPU/GPU Hybrid Inference for MoE Models · SOSP 2025
Storage systems › flash and SSD › solid-state drive
NVMe SSD
0.112019
Clip: A Disk I/O Focused Parallel Out-of-Core Graph Processing System · IEEE Trans. Parallel Distributed Syst. 2019
Storage systems › i/o optimization
disk i/o reduction
0.112017
Squeezing out All the Value of Loaded Data: An Out-of-core Graph Processing System with Reduced Disk I/O · USENIX ATC 2017

Methods — techniques the papers use, named apart from their topics

CPU/GPU hybrid inference · 1.7semi-external mode · 0.8disk i/o locality optimization · 0.8
YearPublicationVenuePosition
2025 KTransformers: Unleashing the Full Potential of CPU/GPU Hybrid Inference for MoE Models
abstract
Due to the sparse nature of Mixture-of-Experts (MoE) models, they are particularly suitable for hybrid CPU/GPU inference, especially in low-concurrency scenarios. This hybrid approach leverages both the large, cost-effective memory capacity of CPU/DRAM and the high bandwidth of GPU/VRAM. However, existing hybrid solutions remain bottlenecked by CPU computation limits and CPU-GPU synchronization overheads, severely restricting their ability to efficiently run state-of-the-art large MoE models, such as the 671B DeepSeek-V3/R1.
Hongtao Chen, Weiyu Xie, Boxin Zhang, Jingqi Tang, Shaoyuan Chen, Ziwei Yuan, Chengyu Qiu, Yuening Zhu, Qingliang Ou, Jiaqi Liao, Xianglin Chen, Zhiyuan Ai, Yongwei Wu 0001
SOSP15
2019 Clip: A Disk I/O Focused Parallel Out-of-Core Graph Processing System
abstract
Existing parallel out-of-core graph processing systems focus on improving disk I/O locality, which leads to restrictions on their programming models. Although improving the locality, these constraints also restrict the expressiveness and hence only sub-optimal algorithms are supported. These sub-optimal algorithms typically incur sequential, but much larger, amount of disk I/O. In this paper, we explore a fundamentally different tradeoff: less total amount of I/O rather than better locality. We show that out-of-core graph processing systems uniquely provide the opportunities to lift the restrictions of the programming model in a feasible manner. To demonstrate the ideas, we build Clip, which enables more efficient algorithms that require much less amount of total disk I/O. Our experiments show that the algorithms that can be only implemented in Clip are much faster than the original disk-locality-optimized algorithms. We also further extend our technique's scope of application by providing a semi-external mode. Our analysis and evaluation demonstrate that semi-external is not only feasible for many cases, but also be able to deliver a significant speedup for important graph applications. Moreover, we further improve the performance of originally supported applications by designing more optimizations and evaluate our system on NVMe SSD.
Zhiyuan Ai, Yongwei Wu 0001, Xuehai Qian, Kang Chen 0001
IEEE Trans. Parallel Distributed Syst.1
2017 Squeezing out All the Value of Loaded Data: An Out-of-core Graph Processing System with Reduced Disk I/O
Zhiyuan Ai, Yongwei Wu 0001, Xuehai Qian, Kang Chen 0001
USENIX ATC1