Luo Mai

dblp:81/10470 · DBLP profile ↗
← Back
24ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-3594-1092ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 4 since 2021Computer networks · 4 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Systems, architecture and hardware · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
abstract
The rise of long-context Large Language Models (LLMs) amplifies memory and bandwidth demands during autoregressive decoding, as the Key-Value (KV) cache grows with each generated token. Low-bit KV-cache quantization (e.g., 4-bit or 2-bit) can reduce memory footprint while preserving accuracy, but existing systems suffer from slow decoding due to their exclusive reliance on CUDA cores, neglecting Tensor Cores—the primary source of compute on modern GPUs. We present BitDecoding, a new long-context LLMs inference system with low-bit KV cache. BitDecoding enables efficient low-bit KV cache decoding by cooperatively leveraging CUDA Cores and Tensor Cores. It introduces methods for automatically inducing optimized layouts to exploit Tensor Cores, along with novel warp-level parallelization strategies for dequantization. For unified system support, BitDecoding includes a query transformation module supporting diverse attention variants, a quantization kernel to support both tensor-wise and channelwise scaling used in various quantization algorithms with high performance, and a dequantization kernel with a softwaredefined pipeline to coordinate CUDA and Tensor Cores execution for mix-precision operations. In addition, architecture-specific optimizations leverage Hopper's warpgroup tensor instructions and Blackwell's native low-precision tensor formats to maximize decoding throughput on the latest GPU generations. Evaluated on Blackwell, Hopper, Ada, and Ampere architectures, BitDecoding attains on average a 7.5× decoding speedup over FP16 FlashDecoding-v2, and further reaches up to 8.6× with native MXFP4 formats on Blackwell, while surpassing the state-of-the-art low-bit system QServe by up to 4.3×. On LLaMA-3.1-8B with a 128K context, BitDecoding reduces singlebatch decoding latency by 3×, demonstrating substantial improvements for long-context generation, and is open sourced at https://github.com/OpenBitSys/BitDecoding.
Dayou Du, Shijie Cao, Jianyi Cheng, Luo Mai, Ting Cao 0003, Mao Yang 0004
HPCA4
2025 MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
abstract
The sparse Mixture-of-Experts (MoE) architecture is increasingly favored for scaling Large Language Models (LLMs) efficiently, but it depends on heterogeneous compute and memory resources. These factors jointly affect system Cost, Accuracy, and Performance (CAP), making trade-offs inevitable. Existing benchmarks often fail to capture these trade-offs accurately, complicating practical deployment decisions. To address this, we introduce MoE-CAP, a benchmark specifically designed for MoE systems. Our analysis reveals that achieving an optimal balance across CAP is difficult with current hardware; MoE systems typically optimize two of the three dimensions at the expense of the third—a dynamic we term the MoE-CAP trade-off. To visualize this, we propose the CAP Radar Diagram. We further introduce sparsity-aware performance metrics—Sparse Memory Bandwidth Utilization (S-MBU) and Sparse Model FLOPS Utilization (S-MFU)—to enable accurate performance benchmarking of MoE systems across diverse hardware platforms and deployment scenarios. This benchmark is available on Github: https://github.com/sparse-generative-ai/MoE-CAP.
Yinsicheng Jiang, Yao Fu 0013, Yeqi Huang, Ping Nie, Zhan Lu, Leyang Xue, Congjie He, Man-Kit Sit, Jilong Xue, Ziming Miao, Dayou Du, Tairan Xu, Edoardo Maria Ponti, Luo Mai
NeurIPS16
2025 WaferLLM: Large Language Model Inference at Wafer Scale
Congjie He, Yeqi Huang, Pei Mu 0003, Ziming Miao, Jilong Xue, Lingxiao Ma, Fan Yang 0024, Luo Mai
OSDI8
2024 Learning High-Frequency Functions Made Easy with Sinusoidal Positional Encoding
abstract
Fourier features based positional encoding (PE) is commonly used in machine learning tasks that involve learning high-frequency features from low-dimensional inputs, such as 3D view synthesis and time series regression with neural tangent kernels. Despite their effectiveness, existing PEs require manual, empirical adjustment of crucial hyperparameters, specifically the Fourier features, tailored to each unique task. Further, PEs face challenges in efficiently learning high-frequency functions, particularly in tasks with limited data. In this paper, we introduce sinusoidal PE (SPE), designed to efficiently learn adaptive frequency features closely aligned with the true underlying function. Our experiments demonstrate that SPE, without hyperparameter tuning, consistently achieves enhanced fidelity and faster training across various tasks, including 3D view synthesis, Text-to-Speech generation, and 1D regression. SPE is implemented as a direct replacement for existing PEs. Its plug-and-play nature lets numerous tasks easily adopt and benefit from SPE.
Chuanhao Sun, Zhihang Yuan, Kai Xu 0014, Luo Mai, N. Siddharth 0001, Mahesh K. Marina
ICML4
2024 ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
Yao Fu 0013, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete, Dmitrii Ustiugov, Yuvraj Patel, Luo Mai
OSDI7
2024 Tenplex: Dynamic Parallelism for Deep Learning using Parallelizable Tensor Collections
abstract
Deep learning (DL) jobs use multi-dimensional parallelism, i.e., combining data, model, and pipeline parallelism, to use large GPU clusters efficiently. Long-running jobs may experience changes to their GPU allocation: (i) resource elasticity during training adds or removes GPUs; (ii) hardware maintenance may require redeployment on different GPUs; and (iii) GPU failures force jobs to run with fewer devices. Current DL frameworks tie jobs to a set of GPUs and thus lack support for these scenarios. In particular, they cannot change the multi-dimensional parallelism of an already-running job in an efficient and model-independent way.
Marcel Wagenländer, Bo Zhao 0019, Luo Mai, Peter R. Pietzuch
SOSP4
2023 GEAR: A GPU-Centric Experience Replay System for Large Reinforcement Learning Models
abstract
This paper introduces a distributed, GPU-centric experience replay system, GEAR, designed to perform scalable reinforcement learning (RL) with large sequence models (such as transformers). With such models, existing systems such as Reverb face considerable bottlenecks in memory, computation, and communication. GEAR, however, optimizes memory efficiency by enabling the memory resources on GPU servers (including host memory and device memory) to manage trajectory data. Furthermore, it facilitates decentralized GPU devices to expedite various trajectory selection strategies, circumventing computational bottlenecks. GEAR is equipped with GPU kernels capable of collecting trajectories using zero-copy access to host memory, along with remote-directed-memory access over InfiniBand, improving communication efficiency. Cluster experiments have shown that GEAR can achieve performance levels up to 6× greater than Reverb when training state-of-the-art large RL models. GEAR is open-sourced at https:// github.com/bigrl-team/gear.
Hanjing Wang, Man-Kit Sit, Congjie He, Ying Wen 0001, Weinan Zhang 0001, Jun Wang 0012, Yaodong Yang 0001, Luo Mai
ICML8
2023 Large sequence models for sequential decision-making: a survey
Muning Wen, Runji Lin, Hanjing Wang, Yaodong Yang 0001, Ying Wen 0001, Luo Mai, Jun Wang 0012, Haifeng Zhang 0002, Weinan Zhang 0001
Frontiers Comput. Sci.6
2023 TorchOpt: An Efficient Library for Differentiable Optimization
abstract
Differentiable optimization algorithms often involve expensive computations of various meta-gradients. To address this, we design and implement TorchOpt, a new PyTorch-based differentiable optimization library. TorchOpt provides an expressive and unified programming interface that simplifies the implementation of explicit, implicit, and zero-order gradients. Moreover, TorchOpt has a distributed execution runtime capable of parallelizing diverse operations linked to differentiable optimization tasks across CPU and GPU devices. Experimental results demonstrate that TorchOpt achieves a 5.2× training time speedup in a cluster. TorchOpt is open-sourced at https://github.com/metaopt/torchopt and has become a PyTorch Ecosystem project.
Xidong Feng, Bo Liu 0039, Xuehai Pan, Yao Fu 0013, Luo Mai, Yaodong Yang 0001
J. Mach. Learn. Res.6
2022 MegBA: A GPU-Based Distributed Library for Large-Scale Bundle Adjustment
Wenteng Liang, Ran Yan 0003, Luo Mai, Shiwen Liu
ECCV (37)4
2022 A Theoretical Understanding of Gradient Bias in Meta-Reinforcement Learning
abstract
Gradient-based Meta-RL (GMRL) refers to methods that maintain two-level optimisation procedures wherein the outer-loop meta-learner guides the inner-loop gradient-based reinforcement learner to achieve fast adaptations. In this paper, we develop a unified framework that describes variations of GMRL algorithms and points out that existing stochastic meta-gradient estimators adopted by GMRL are actually \textbf{biased}. Such meta-gradient bias comes from two sources: 1) the compositional bias incurred by the two-level problem structure, which has an upper bound of $\mathcal{O}\big(K\alpha^{K}\hat{\sigma}_{\text{In}}|\tau|^{-0.5}\big)$ \emph{w.r.t.} inner-loop update step $K$, learning rate $\alpha$, estimate variance $\hat{\sigma}^{2}_{\text{In}}$ and sample size $|\tau|$, and 2) the multi-step Hessian estimation bias $\hat{\Delta}_{H}$ due to the use of autodiff, which has a polynomial impact $\mathcal{O}\big((K-1)(\hat{\Delta}_{H})^{K-1}\big)$ on the meta-gradient bias. We study tabular MDPs empirically and offer quantitative evidence that testifies our theoretical findings on existing stochastic meta-gradient estimators. Furthermore, we conduct experiments on Iterated Prisoner's Dilemma and Atari games to show how other methods such as off-policy learning and low-bias estimator can help fix the gradient bias for GMRL algorithms in general.
Bo Liu 0039, Xidong Feng, Luo Mai, Haifeng Zhang 0002, Jun Wang 0012, Yaodong Yang 0001
NeurIPS4
2022 Ekko: A Large-Scale Deep Learning Recommender System with Low-Latency Model Update
Chijun Sima, Yao Fu 0013, Man-Kit Sit, Liyi Guo, Xuri Gong, Junyu Wu, Haidong Rong, Pierre-Louis Aublin, Luo Mai
OSDI11
2021 Efficient Reinforcement Learning Development with RLzoo
abstract
Many multimedia developers are exploring for adopting Deep Reinforcement Learning (DRL) techniques in their applications. They however often find such an adoption challenging. Existing DRL libraries provide poor support for prototyping DRL agents (i.e., models), customising the agents, and comparing the performance of DRL agents. As a result, the developers often report low efficiency in developing DRL agents. In this paper, we introduce RLzoo, a new DRL library that aims to make the development of DRL agents efficient. RLzoo provides developers with (i) high-level yet flexible APIs for prototyping DRL agents, and further customising the agents for best performance, (ii) a model zoo where users can import a wide range of DRL agents and easily compare their performance, and (iii) an algorithm that can automatically construct DRL agents with custom components (which are critical to improve agent's performance in custom applications). Evaluation results show that RLzoo can effectively reduce the development cost of DRL agents, while achieving comparable performance with existing DRL libraries.
Tianyang Yu, Hongming Zhang 0003, Yanhua Huang, Quancheng Guo, Luo Mai, Hao Dong 0003
ACM Multimedia7
2021 Fast and Flexible Human Pose Estimation with HyperPose
abstract
Estimating human pose is an important yet challenging task in multimedia applications. Existing pose estimation libraries target reproducing standard pose estimation algorithms. When it comes to customising these algorithms for real-world applications, none of the existing libraries can offer both the flexibility of developing custom pose estimation algorithms and the high-performance of executing these algorithms on commodity devices. In this paper, we introduce Hyperpose, a novel flexible and high-performance pose estimation library. Hyperpose provides expressive Python APIs that enable developers to easily customise pose estimation algorithms for their applications. It further provides a model inference engine highly optimised for real-time pose estimation. This engine can dynamically dispatch carefully designed pose estimation tasks to CPUs and GPUs, thus automatically achieving high utilisation of hardware resources irrespective of deployment environments. Extensive evaluation results show that Hyperpose can achieve up to 3.1x~7.3x higher pose estimation throughput compared to state-of-the-art pose estimation libraries without compromising estimation accuracy. By 2021, Hyperpose has received over 1000 stars on GitHub and attracted users from both industry and academy.
Yixiao Guo, Luo Mai, Hao Dong 0003
ACM Multimedia4
2021 Move Fast and Meet Deadlines: Fine-grained Real-time Stream Processing with Cameo
Shivaram Venkataraman, Indranil Gupta, Luo Mai, Rahul Potharaju
NSDI4
2020 KungFu: Making Training in Distributed Machine Learning Adaptive
Luo Mai, Marcel Wagenländer, Konstantinos Fertakis, Andrei-Octavian Brabete, Peter R. Pietzuch
OSDI1
2019 Crossbow: Scaling Deep Learning with Small Batch Sizes on Multi-GPU Servers
abstract
Deep learning models are trained on servers with many GPUs, and training must scale with the number of GPUs. Systems such as TensorFlow and Caffe2 train models with parallel synchronous stochastic gradient descent: they process a batch of training data at a time, partitioned across GPUs, and average the resulting partial gradients to obtain an updated global model. To fully utilise all GPUs, systems must increase the batch size, which hinders statistical efficiency. Users tune hyper-parameters such as the learning rate to compensate for this, which is complex and model-specific. We describe Crossbow, a new single-server multi-GPU system for training deep learning models that enables users to freely choose their preferred batch size---however small---while scaling to multiple GPUs. Crossbow uses many parallel model replicas and avoids reduced statistical efficiency through a new synchronous training method. We introduce SMA, a synchronous variant of model averaging in which replicas independently explore the solution space with gradient descent, but adjust their search synchronously based on the trajectory of a globally-consistent average model. Crossbow achieves high hardware efficiency with small batch sizes by potentially training multiple model replicas per GPU, automatically tuning the number of replicas to maximise throughput. our experiments show that Crossbow improves the training time of deep learning models on an 8-GPU server by 1.3--4X compared to TensorFlow.
Alexandros Koliousis, Pijika Watcharapichat, Matthias Weidlich 0001, Luo Mai, Paolo Costa, Peter R. Pietzuch
Proc. VLDB Endow.4
2018 Chi: A Scalable and Programmable Control Plane for Distributed Stream Processing Systems
abstract
Stream-processing workloads and modern shared cluster environments exhibit high variability and unpredictability. Combined with the large parameter space and the diverse set of user SLOs, this makes modern streaming systems very challenging to statically configure and tune. To address these issues, in this paper we investigate a novel control-plane design, Chi, which supports continuous monitoring and feedback, and enables dynamic re-configuration. Chi leverages the key insight of embedding control-plane messages in the data-plane channels to achieve a low-latency and flexible control plane for stream-processing systems. Chi introduces a new reactive programming model and design mechanisms to asynchronously execute control policies, thus avoiding global synchronization. We show how this allows us to easily implement a wide spectrum of control policies targeting different use cases observed in production. Large-scale experiments using production workloads from a popular cloud provider demonstrate the flexibility and efficiency of our approach.
Luo Mai, Kai Zeng 0002, Rahul Potharaju, Steve Suh, Shivaram Venkataraman, Paolo Costa, Terry Kim, Saravanam Muthukrishnan, Vamsi Kuppa, Sudheer Dhulipalla, Sriram Rao
Proc. VLDB Endow.1
2017 TensorLayer: A Versatile Library for Efficient Deep Learning Development
abstract
Recently we have observed emerging uses of deep learning techniques in multimedia systems. Developing a practical deep learning system is arduous and complex. It involves labor-intensive tasks for constructing sophisticated neural networks, coordinating multiple network models, and managing a large amount of training-related data. To facilitate such a development process, we propose TensorLayer which is a Python-based versatile deep learning library. TensorLayer provides high-level modules that abstract sophisticated operations towards neuron layers, network models, training data and dependent training jobs. In spite of offering simplicity, it has transparent module interfaces that allows developers to flexibly embed low-level controls within a backend engine, with the aim of supporting fine-grain tuning towards training. Real-world cluster experiment results show that TensorLayeris able to achieve competitive performance and scalability in critical deep learning tasks. TensorLayer was released in September 2016 on GitHub. Since after, it soon become one of the most popular open-sourced deep learning library used by researchers and practitioners.
Hao Dong 0003, Akara Supratak, Luo Mai, Fangde Liu, Axel Oehmichen, Simiao Yu, Yike Guo
ACM Multimedia3
2017 Emu: Rapid Prototyping of Networking Services
Nik Sultana, Salvator Galea, David Greaves, Marcin Wójcik, Jonny Shipton, Richard G. Clegg, Luo Mai, Pietro Bressana, Robert Soulé, Richard Mortier, Paolo Costa, Peter R. Pietzuch, Jon Crowcroft, Andrew W. Moore 0002, Noa Zilberman
USENIX ATC7
2016 FLICK: Developing and Running Application-Specific Network Services
Abdul Alim, Richard G. Clegg, Luo Mai, Lukas Rupprecht, Eric Seckler, Paolo Costa, Peter R. Pietzuch, Alexander L. Wolf, Nik Sultana, Jon Crowcroft, Anil Madhavapeddy, Andrew W. Moore 0002, Richard Mortier, Masoud Koleini, Luis Oviedo, Matteo Migliavacca, Derek McAuley
USENIX ATC3
2014 NetAgg: Using Middleboxes for Application-specific On-path Aggregation in Data Centres
abstract
Data centre applications for batch processing (e.g. map/reduce frameworks) and online services (e.g. search engines) scale by distributing data and computation across many servers. They typically follow a partition/aggregation pattern: tasks are first partitioned across servers that process data locally, and then those partial results are aggregated. This data aggregation step, however, shifts the performance bottleneck to the network, which typically struggles to support many-to-few, high-bandwidth traffic between servers.
Luo Mai, Lukas Rupprecht, Abdul Alim, Paolo Costa, Matteo Migliavacca, Peter R. Pietzuch, Alexander L. Wolf
CoNEXT1
2013 Supporting application-specific in-network processing in data centres
abstract
No abstract available.
Luo Mai, Lukas Rupprecht, Paolo Costa, Matteo Migliavacca, Peter R. Pietzuch, Alexander L. Wolf
SIGCOMM1
2011 Load Balanced Rendezvous Data Collection in Wireless Sensor Networks
abstract
We study the rendezvous data collection problem for the mobile sink in wireless sensor networks. We introduce to jointly optimize trajectory planning for the mobile sink and workload balancing for the network. By doing so, the mobile sink is able to efficiently collect network-wide data within a given delay bound and the network can eliminate the energy bottleneck to dramatically prolong its lifetime. Such a joint optimization problem is shown to be NP-hard and we propose an approximation algorithm, named RPS-LB, to approach the optimal solution. In RPS-LB, according to observed properties of the median reference structure in the network, a series of Rendezvous Points (RPs) are selected to construct the trajectory for the mobile sink and the derived approximation ratio of RPSLB guarantees that the formed trajectory is comparable with the optimal solution. The workload allocated to each RP is proven to be balanced mathematically. We then relax the assumption that mobile sink knows the location of each sensor node and present a localized, fully distributed version, RPS-LB-D, which largely improves the system applicability in practice. We verify the effectiveness of our proposals via extensive experiments.
Luo Mai, Longfei Shangguan, Chao Lang, Junzhao Du, Hui Liu 0006, Zhenjiang Li 0001, Mo Li 0001
MASS1