VLDB 2026 Research / reviewers in the wild / expert
Yuxuan Liu 0019
dblp:42/7844-19
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0008-2770-7926ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI ArchitecturesabstractThe rapid scaling of large language models (LLMs) has unveiled critical limitations in current hardware architectures, including constraints in memory capacity, computational efficiency, and interconnection bandwidth.DeepSeek-V3, trained on 2,048 NVIDIA H800 GPUs, demonstrates how hardware-aware model co-design can effectively address these challenges, enabling cost-efficient training and inference at scale.This paper presents an in-depth analysis of the DeepSeek-V3/R1 model architecture and its AI infrastructure, highlighting key innovations such as Multi-head Latent Attention (MLA) for enhanced memory efficiency, Mixture of Experts (MoE) architectures for optimized computation-communication trade-offs, FP8 mixed-precision training to unlock the full potential of hardware capabilities, and a Multi-Plane Network Topology to minimize * Yuqing Wang and Liyue Zhang are the corresponding authors of this paper. Chenggang Zhao, Chengqi Deng, Chong Ruan, Damai Dai, Huazuo Gao, Jiashi Li, Panpan Huang, Shangyan Zhou, Shirong Ma, Wenfeng Liang, Ying He 0018, Yuxuan Liu 0019, Y. X. Wei |
ISCA | 14 |
| 2024 | Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep LearningabstractThe rapid progress in Deep Learning (DL) and Large Language Models (LLMs) has exponentially increased demands of computational power and bandwidth. This, combined with the high costs of faster computing chips and interconnects, has significantly inflated High Performance Computing (HPC) construction costs. To address these challenges, we introduce the Fire-Flyer AI-HPC architecture, a synergistic hardware-software co-design framework and its best practices. For DL training, we deployed the Fire-Flyer 2 with 10,000 PCIe A100 GPUs, achieved performance approximating the DGX-A100 while reducing costs by half and energy consumption by $40 \%$. We specifically engineered HFReduce to accelerate allreduce communication and implemented numerous measures to keep our Computation-Storage Integrated Network congestion-free. Through our software stack, including HaiScale, 3FS, and HAI-Platform, we achieved substantial scalability by overlapping computation and communication. Our system-oriented experience from DL training provides valuable insights to drive future advancements in AI-HPC. Xiao Bi, Guanting Chen 0002, Shanhuang Chen, Chengqi Deng, Honghui Ding, Kai Dong 0003, Qiushi Du, Kang Guan, Jianzhong Guo, Yongqiang Guo, Zhe Fu 0009, Ying He 0018, Panpan Huang, Jiashi Li, Wenfeng Liang, Xiaodong Liu 0021, Xin Liu 0126, Yiyuan Liu, Yuxuan Liu 0019, Shanghao Lu, Xiaotao Nie, Tian Pei, Junjie Qiu, Zehui Ren, Zhangli Sha, Xuecheng Su, Xiaowen Sun, Yixuan Tan, Minghui Tang, Ziwei Xie, Yiliang Xiong, Shengfeng Ye, Shuiping Yu, Yukun Zha, Mingchuan Zhang, Yichao Zhang 0004, Chenggang Zhao, Yao Zhao 0005, Shangyan Zhou, Shunfeng Zhou, Yuheng Zou |
SC | 21 |
| 2023 | CPS: A Cooperative Para-virtualized Scheduling Framework for Manycore MachinesabstractToday's cloud platforms offer large virtual machine (VM) instances with multiple virtual CPUs (vCPU) on manycore machines. These machines typically have a deep memory hierarchy to enhance communication between cores. Although previous researches have primarily focused on addressing the performance scalability issues caused by the double scheduling problem in virtualized environments, they mainly concentrated on solving the preemption problem of synchronization primitives and the traditional NUMA architecture. This paper specifically targets a new aspect of scalability issues caused by the absence of runtime hypervisor-internal states (RHS). We demonstrate two typical RHS problems, namely the invisible pCPU (physical CPU) load and dynamic cache group mapping. These RHS problems result in a collapse in VM performance and low CPU utilization because the guest VM lacks visibility into the latest runtime internal states maintained by the hypervisor, such as pCPU load and vCPU-pCPU mappings. Consequently, the guest VM makes inefficient scheduling decisions. Yuxuan Liu 0019, Tianqiang Xu, Zeyu Mi, Zhichao Hua 0001, Binyu Zang, Haibo Chen 0001 |
ASPLOS (4) | 1 |
| 2023 | Security and Performance in the Delegated User-level Virtualization
Dingji Li, Zeyu Mi, Yuxuan Liu 0019, Binyu Zang, Haibing Guan, Haibo Chen 0001 |
OSDI | 4 |