VLDB 2026 Research / reviewers in the wild / expert
Lixiang Luo
dblp:31/2051
· DBLP profile ↗
3ranked-venue papers
1as first author
2since 2021 · last 2025
0000-0001-9850-3201ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Vela: A Virtualized LLM Training System with GPU Direct RoCEabstractVela is a cloud-native system designed for LLM training workloads built using off-the-shelf hardware, Linux KVM-based virtualization, and a virtualized RDMA over Converged Ethernet (RoCE) network. Vela virtual machines (VMs) support peer-to-peer DMA between the GPUs and SRIOV-based network interface. In this paper, we share Vela's key architectural aspects with details from an NVIDIA A100 GPU-based deployment in one of the IBM Cloud data centers. Throughout the paper, we share insights and experiences from designing, building, and operating the system over a ~2.5 year timeframe to highlight the capabilities of readily available software and hardware technologies and the improvement opportunities for future AI systems, thereby making AI infrastructure more accessible to a broader community. As we evaluated the system for performance at ~1500 GPU scale, we achieved ~80% of the ideal throughput while training a 50 billion parameter decoder model using model parallelism, and ~70% per GPU FLOPS compared to a single VM with the High-Performance Linpack benchmark. Apoorve Mohan, Robert Walkup, Bengi Karaçali, Ming-Hung Chen, Abdullah Kayi, Liran Schour, Shweta Salaria, Sophia Wen, I-Hsin Chung, Abdul Alim, Constantinos Evangelinos, Lixiang Luo, Marc Dombrowa, Laurent Schares, Ali Sydney, Pavlos Maniotis, Sandhya Koteshwara, Brent Tang, Joel Belog, Rei Odaira, Vasily Tarasov, Eran Gampel, Drew Thorstensen, Talia Gershon, Seetharami Seelam |
ASPLOS (2) | 12 |
| 2022 | NVMe Virtualization for Cloud Virtual MachinesabstractPublic clouds are rapidly moving to support Non-Volatile Memory Express (NVMe) based storage to meet the ever-increasing I/O throughput and latency demands of modern workloads. They provide NVMe storage through virtual machines (VMs) where multiple VMs running on a host may share a physical NVMe device. The virtualization method used to share the NVMe capability has important performance, usability and security implications. In this paper, we propose three NVMe storage virtualization methods: PCI device passthrough, virtual block device method, and Storage Performance Development Kit (SPDK) virtual host target method. We evaluate these virtualization methods in terms of performance, scalability, CPU overhead, technology maturity, security, and availability to use one or more of these methods in IBM public cloud. Lixiang Luo, I-Hsin Chung, Seetharami R. Seelam, Ming-Hung Chen, Yun Joon Soh |
ICPE | 1 |
| 2020 | Parameter Optimization of Reduced Fluid Model via Sparse Point MeasurementsabstractModel order reduction is a modeling method that can facilitate the analysis and the control synthesis of distributed parameter systems, which are governed by partial differential equations. Many techniques have been proposed to further improve the property of the reduced order model based on full state information, which is usually not available in practical applications. In this paper, a reduced-order modeling method based on sparse point measurements is proposed for the benchmark system of flow past a cylinder. This method involves solving the dynamic optimization problems online, whose objective functions measure the differences between the predicted and observed variables on each time horizon. In this method, the sensor placement is also discussed, and we determine the sensor locations depending on a model-free optimization technique. Several numerical examples of different scenarios are simulated not only to illustrate the effectiveness of the proposed method but also to verify the discussion on the sensor placement strategy. Chao Xu 0001, Lixiang Luo |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |