VLDB 2026 Research / reviewers in the wild / expert
Lingxiao Ma
dblp:57/3203
· DBLP profile ↗
28ranked-venue papers
3as first author
22since 2021 · last 2026
0009-0009-9524-5476ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 13 · 1 first-author · 12 since 2021Systems, architecture and hardware · 9 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FastTTS: Accelerating Test-Time Scaling for Edge LLM Reasoning
Hao Mark Chen, Zhiwen Mo, Guanxi Lu, Shuang Liang 0012, Lingxiao Ma, Wayne Luk, Hongxiang Fan |
ASPLOS (2) | 5 |
| 2026 | MetaAttention: A Unified and Performant Attention Framework across Hardware BackendsabstractComputing attention is the backbone of transformer-based models like large language models. However, the increasing diversity of attention algorithms presents significant challenges for unleashing hardware performance. State-of-the-art variants like FlashAttention target a specific attention algorithm or hardware platform, which fail to generalize to other algorithms and platforms. Yu Cheng 0030, Lei Wang 0222, Yuqing Xia, Ziming Miao, Lingxiao Ma, Fan Yang 0024, Jilong Xue, Zhi Yang 0001, Mao Yang 0004, Xingda Wei, Haibo Chen 0001 |
PPoPP | 6 |
| 2025 | T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on EdgeabstractThe deployment of Large Language Models (LLMs) on edge devices is increasingly important to enhance on-device intelligence. Weight quantization is crucial for reducing the memory footprint of LLMs on devices. However, low-bit LLMs necessitate mixed precision matrix multiplication (mpGEMM) of low precision weights and high precision activations during inference. Existing systems, lacking native support for mpGEMM, resort to dequantize weights for high precision computation. Such an indirect way can lead to a significant inference overhead. Jianyu Wei, Shijie Cao, Ting Cao 0003, Lingxiao Ma, Lei Wang 0222, Yanyong Zhang, Mao Yang 0004 |
EuroSys | 4 |
| 2025 | NeuStream: Bridging Deep Learning Serving and Stream ProcessingabstractModern Deep Neural Network (DNN) exhibits a pattern where multiple sub-models are executed, guided by control flows such as loops and switch/merge operations. This dynamic nature introduces complexities in batching the requests of such DNNs for efficient execution on GPUs. In this paper, we present NeuStream, a programming model and runtime system for serving deep learning workloads using stream processing. NeuStream decomposes the inference workflow into modules and forms them into a streaming processing system where a request flows through. Based on such abstraction, NeuStream is able to batch requests at fine-grained module granularity. To maximize serving goodput, NeuStream exploits a two-level scheduling approach to decide the best batching requests and resource allocation for each module while satisfying service level objectives (SLOs). Our evaluation of NeuStream on a set of modern DNNs like Large Language Models (LLM) and diffusion models, etc., shows that NeuStream significantly improves goodput compared to state-of-the-art DNN serving systems. Yu Cheng 0030, Ziming Miao, Lingxiao Ma, Jilong Xue, Zhi Yang 0001 |
EuroSys | 6 |
| 2025 | LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM InferenceabstractLarge Language Model (LLM) inference becomes resource-intensive, prompting a shift toward low-bit model weights to reduce the memory footprint and improve efficiency.Such low-bit LLMs necessitate the mixed-precision matrix multiplication (mpGEMM), an important yet underexplored operation involving the multiplication of lower-precision weights with higher-precision activations.Off-theshelf hardware does not support this operation natively, leading to indirect, thus inefficient, dequantization-based implementations.In this paper, we study the lookup table (LUT)-based approach for mpGEMM and find that a conventional LUT implementation fails to achieve the promised gains.To unlock the full potential of LUT-based mpGEMM, we propose LUT Tensor Core, a softwarehardware co-design for low-bit LLM inference.LUT Tensor Core differentiates itself from conventional LUT designs through: 1) * Work is done during internship at Microsoft Research. Zhiwen Mo, Lei Wang 0222, Jianyu Wei, Zhichen Zeng 0002, Shijie Cao, Lingxiao Ma, Naifeng Jing, Ting Cao 0003, Jilong Xue, Fan Yang 0024, Mao Yang 0004 |
ISCA | 6 |
| 2025 | PipeThreader: Software-Defined Pipelining for Efficient DNN Execution
Yu Cheng 0030, Lei Wang 0222, Yining Shi 0001, Yuqing Xia, Lingxiao Ma, Jilong Xue, Yang Wang 0053, Zhiwen Mo, Fan Yang 0024, Mao Yang 0004, Zhi Yang 0001 |
OSDI | 5 |
| 2025 | WaferLLM: Large Language Model Inference at Wafer Scale
Congjie He, Yeqi Huang, Pei Mu 0003, Ziming Miao, Jilong Xue, Lingxiao Ma, Fan Yang 0024, Luo Mai |
OSDI | 6 |
| 2025 | BitNet: 1-bit Pre-training for Large Language ModelsabstractThe increasing size of large language models (LLMs) has posed challenges for deployment and raised concerns about environmental impact due to high energy consumption. Previous research typically applies quantization after pre-training. While these methods avoid the need for model retraining, they often cause notable accuracy loss at extremely low bit-widths. In this work, we explore the feasibility and scalability of 1-bit pre-training. We introduce BitNet b1 and BitNet b1.58, the scalable and stable 1-bit Transformer architecture designed for LLMs. Specifically, we introduce BitLinear as a drop-in replacement of the nn.Linear layer in order to train 1-bit weights from scratch. Experimental results show that BitNet b1 achieves competitive performance, compared to state-of-the-art 8-bit quantization methods and FP16 Transformer baselines. With the ternary weight, BitNet b1.58 matches the half-precision Transformer LLM with the same model size and training tokens in terms of both perplexity and end-task performance, while being significantly more cost-effective in terms of latency, memory, throughput, and energy consumption. More profoundly, BitNet defines a new scaling law and recipe for training new generations of LLMs that are both high-performance and cost-effective. It enables a new computation paradigm and opens the door for designing specific hardware optimized for 1-bit LLMs. Hongyu Wang 0009, Shuming Ma, Lingxiao Ma, Lei Wang 0222, Wenhui Wang 0003, Li Dong 0004, Shaohan Huang, Huaijie Wang, Jilong Xue, Yi Wu 0013, Furu Wei |
J. Mach. Learn. Res. | 3 |
| 2025 | Squeezer: Efficient Multi-DNN Inference for Edge Video Analytics via Cross-Model SchedulingabstractVideo analytics at the edge is becoming increasingly prevalent in many scenarios, such as smart campuses and intelligent factories. These applications often consist of multiple subtasks, which necessitates the optimization for multi-DNN (Deep Neural Network) inference. Due to limited consideration over cross-model scheduling, current practices cannot fully leverage available computing resources, leading to suboptimal performance. To address this, we propose Squeezer, a multiDNN serving framework that holistically schedules multiple DNN models on an edge server with a single GPU. Squeezer decouples the cross-model scheduling into a two-layered approach, which involves (1) balanced operator grouping which partitions operators of multiple DNN models into groups, significantly reducing the scheduling complexity and (2) kernel scheduler which orchestrates parallel execution within each group by considering the interplay among kernels running in parallel, thereby enabling cross-model optimizations in multi-DNN inference. Performance evaluation results demonstrate that Squeezer outperforms state-of-the-art baselines, achieving up to 1.91× improvement in system throughput. Lingxiao Ma, Ziyan Fu 0001, Yuanchun Li 0003, Ju Ren 0001, Yaoxue Zhang, Yunxin Liu 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Ladder: Enabling Efficient Low-Precision Deep Learning Computing through Hardware-aware Tensor Transformation
Lei Wang 0222, Lingxiao Ma, Shijie Cao, Quanlu Zhang, Jilong Xue, Yining Shi 0001, Ningxin Zheng, Ziming Miao, Fan Yang 0024, Ting Cao 0003, Yuqing Yang 0001, Mao Yang 0004 |
OSDI | 2 |
| 2024 | ConvStencil: Transform Stencil Computation to Matrix Multiplication on Tensor CoresabstractTensor Core Unit (TCU) is increasingly integrated into modern high-performance processors to enhance matrix multiplication performance. However, constrained to its over-specification, its potential for improving other critical scientific operations like stencil computations remains untapped. Yuetao Chen, Kun Li 0016, Donglin Bai, Lei Wang 0222, Lingxiao Ma, Yunquan Zhang, Ting Cao 0003, Mao Yang 0004 |
PPoPP | 6 |
| 2024 | Scaling Deep Learning Computation over the Inter-Core Connected Intelligence Processor with T10abstractAs AI chips incorporate numerous parallelized cores to scale deep learning (DL) computing, inter-core communication is enabled recently by employing high-bandwidth and low-latency interconnect links on the chip (e.g., Graphcore IPU). It allows each core to directly access the fast scratchpad memory in other cores, which enables new parallel computing paradigms. However, without proper support for the scalable inter-core connections in current DL compilers, it is hard for developers to exploit the benefits of this new architecture. Yuqi Xue, Yu Cheng 0030, Lingxiao Ma, Ziming Miao, Jilong Xue, Jian Huang 0006 |
SOSP | 4 |
| 2023 | Welder: Scheduling Deep Learning Memory Access via Tile-graph
Yining Shi 0001, Zhi Yang 0001, Jilong Xue, Lingxiao Ma, Yuqing Xia, Ziming Miao, Yuxiao Guo 0001, Fan Yang 0024, Lidong Zhou |
OSDI | 4 |
| 2023 | Optimizing Dynamic Neural Networks with Brainstorm
Weihao Cui, Zhenhua Han, Lingji Ouyang, Yichuan Wang 0002, Ningxin Zheng, Lingxiao Ma, Yuqing Yang 0001, Fan Yang 0024, Jilong Xue, Lili Qiu, Lidong Zhou, Quan Chen 0002, Haisheng Tan, Minyi Guo |
OSDI | 6 |
| 2023 | Cocktailer: Analyzing and Optimizing Dynamic Control Flow in Deep Learning
Chen Zhang 0001, Lingxiao Ma, Jilong Xue, Yining Shi 0001, Ziming Miao, Fan Yang 0024, Jidong Zhai, Zhi Yang 0001, Mao Yang 0004 |
OSDI | 2 |
| 2023 | PIT: Optimization of Dynamic Sparse Deep Learning Models via Permutation Invariant TransformationabstractDynamic sparsity, where the sparsity patterns are unknown until runtime, poses a significant challenge to deep learning. The state-of-the-art sparsity-aware deep learning solutions are restricted to pre-defined, static sparsity patterns due to significant overheads associated with preprocessing. Efficient execution of dynamic sparse computation often faces the misalignment between the GPU-friendly tile configuration for efficient execution and the sparsity-aware tile shape that minimizes coverage wastes (non-zero values in tensor). Ningxin Zheng, Huiqiang Jiang, Quanlu Zhang, Zhenhua Han, Lingxiao Ma, Yuqing Yang 0001, Fan Yang 0024, Chengruidong Zhang, Lili Qiu, Mao Yang 0004, Lidong Zhou |
SOSP | 5 |
| 2023 | FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device PlacementabstractWith the increasing data volume, there is a trend of using large-scale pre-trained models to store the knowledge into an enormous number of model parameters. The training of these models is composed of lots of dense algebras, requiring a huge amount of hardware resources. Recently, sparsely-gated Mixture-of-Experts (MoEs) are becoming more popular and have demonstrated impressive pretraining scalability in various downstream tasks. However, such a sparse conditional computation may not be effective as expected in practical systems due to the routing imbalance and fluctuation problems. Generally, MoEs are becoming a new data analytics paradigm in the data life cycle and suffering from unique challenges at scales, complexities, and granularities never before possible. In this paper, we propose a novel DNN training framework, FlexMoE, which systematically and transparently address the inefficiency caused by dynamic dataflow. We first present an empirical analysis on the problems and opportunities of training MoE models, which motivates us to overcome the routing imbalance and fluctuation problems by a dynamic expert management and device placement mechanism. Then we introduce a novel scheduling module over the existing DNN runtime to monitor the data flow, make the scheduling plans, and dynamically adjust the model-to-hardware mapping guided by the real-time data traffic. A simple but efficient heuristic algorithm is exploited to dynamically optimize the device placement during training. We have conducted experiments on both NLP models (e.g., BERT and GPT) and vision models (e.g., Swin). And results show FlexMoE can achieve superior performance compared with existing systems on real-world workloads --- FlexMoE outperforms DeepSpeed by 1.70x on average and up to 2.10x, and outperforms FasterMoE by 1.30x on average and up to 1.45x. Xiaonan Nie, Xupeng Miao, Zilong Wang 0033, Jilong Xue, Lingxiao Ma, Gang Cao 0003, Bin Cui 0001 |
Proc. ACM Manag. Data | 6 |
| 2022 | SparTA: Deep-Learning Model Sparsity via Tensor-with-Sparsity-Attribute
Ningxin Zheng, Quanlu Zhang, Lingxiao Ma, Yuqing Yang 0001, Fan Yang 0024, Yang Wang 0053, Mao Yang 0004, Lidong Zhou |
OSDI | 4 |
| 2022 | ROLLER: Fast and Efficient Tensor Compilation for Deep Learning
Hongyu Zhu 0003, Yijia Diao, Shanbin Ke, Chen Zhang 0001, Jilong Xue, Lingxiao Ma, Yuqing Xia, Fan Yang 0024, Mao Yang 0004, Lidong Zhou, Asaf Cidon, Gennady Pekhimenko |
OSDI | 8 |
| 2022 | CuWide: Towards Efficient Flow-Based Training for Sparse Wide Models on GPUsabstractWide models such as generalized linear models and factorization-based models have been extensively used in various predictive applications, e.g., recommendation, CTR prediction, and image recognition. Due to the memory bounded property of the models, the performance improvement on CPU is reaching the limitation. GPU is known to have many computation units and high memory bandwidth, and becomes a promising platform for training machine learning models. However, the GPU training for the wide models is far from optimal due to the sparsity and irregularity in wide models. The existing GPU-based wide models are even slower than the ones using CPU. The classical training schema of the wide models does not optimized for the GPU architecture, which suffers from large amount of random memory accesses and redundant read/write of intermediate values. In this paper, we propose an efficient GPU-training framework for the large-scale wide models, named cuWide. To fully benefit from the memory hierarchy of GPU, cuWide applies a new flow-based schema for training, which leverages the spatial and temporal locality of wide models to drastically reduce the amount of communication with GPU global memory. To do so, we adopt a bigraph computation model to efficiently realize the flow-based schema and exploit three flexible interfaces for programming. Further, we use the 2D partition of mini-batch (in sample and feature dimensions) with proposed graph abstraction to optimize GPU memory access for sparse data, and apply several spatial-temporal caching mechanisms (importance-based model caching and cross-stage accumulation caching mechanisms) to achieve a high performance kernel. To efficiently implement cuWide, we also propose several GPU-oriented optimizations, including feature-oriented data layout to enhance the data locality, replication mechanism to reduce update conflicts in shared memory, and multi-stream scheduling to overlap data transferring and kernel computing. We show that cuWide can be up to more than 20× faster than the state-of-the-art GPU solutions and multi-core CPU solutions. Xupeng Miao, Lingxiao Ma, Zhi Yang 0001, Yingxia Shao, Bin Cui 0001, Lele Yu, Jiawei Jiang 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | CuWide: Towards Efficient Flow-based Training for Sparse Wide Models on GPUs (Extended Abstract)abstractIn this paper, we propose an efficient GPU-training framework for the large-scale wide models, named cuWide. To fully benefit from the memory hierarchy of GPU, cuWide applies a new flow-based schema for training, which leverages the spatial and temporal locality of wide models to drastically reduce the amount of communication with GPU global memory. Comprehensive experiments show that cuWide can be up to more than 20× faster than the state-of-the-art GPU solutions and multi-core CPU solutions. Xupeng Miao, Lingxiao Ma, Zhi Yang 0001, Yingxia Shao, Bin Cui 0001, Lele Yu, Jiawei Jiang 0001 |
ICDE | 2 |
| 2021 | Heterogeneity-Aware Distributed Machine Learning Training via Partial ReduceabstractAll-reduce is the key communication primitive used in distributed data-parallel training due to the high performance in the homogeneous environment. However, All-reduce is sensitive to stragglers and communication delays as deep learning has been increasingly deployed on the heterogeneous environment like cloud. In this paper, we propose and analyze a novel variant of all-reduce, called partial-reduce, which provides high heterogeneity tolerance and performance by decomposing the synchronous all-reduce primitive into parallel-asynchronous partial-reduce operations. We provide theoretical guarantees, proving that partial-reduce converges to a stationary point at the similar sub-linear rate as distributed SGD. To enforce the convergence of the partial-reduce primitive, we further propose a dynamic staleness-aware distributed averaging algorithm and implement a novel group generation mechanism to prevent possible update isolation in heterogeneous environments. We build a prototype system in the real production cluster and validate its performance under different workloads. The experiments show that it is 1.21x-2x faster than other state-of-the-art baselines. Xupeng Miao, Xiaonan Nie, Yingxia Shao, Zhi Yang 0001, Jiawei Jiang 0001, Lingxiao Ma, Bin Cui 0001 |
SIGMOD Conference | 6 |
| 2020 | PCGCN: Partition-Centric Processing for Accelerating Graph Convolutional NetworkabstractInspired by the successes of convolutional neural networks (CNN) in computer vision, the convolutional operation has been moved beyond low-dimension grids (e.g., images) to high-dimensional graph-structured data (e.g., web graphs, social networks), leading to graph convolutional network (GCN). And GCN has been gaining popularity due to its success in real-world applications such as recommendation, natural language processing, etc. Because neural network and graph propagation have high computation complexity, GPUs have been introduced to both neural network training and graph processing. However, it is notoriously difficult to perform efficient GCN computing on data parallel hardware like GPU due to the sparsity and irregularity in graphs. In this paper, we present PCGCN, a novel and general method to accelerate GCN computing by taking advantage of the locality in graphs. We experimentally demonstrate that real-world graphs usually have the clustering property that can be used to enhance the data locality in GCN computing. Then, PCGCN proposes to partition the whole graph into chunks according to locality and process subgraphs with a dual-mode computing strategy which includes a selective and a full processing methods for sparse and dense subgraphs, respectively. Compared to existing state-of-the-art implementations of GCN on real-world and synthetic datasets, our implementation on top of TensorFlow achieves up to 8.8× speedup over the fastest one of the baselines. Chao Tian 0001, Lingxiao Ma, Zhi Yang 0001, Yafei Dai |
IPDPS | 2 |
| 2020 | Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks
Lingxiao Ma, Zhi Yang 0001, Jilong Xue, Youshan Miao, Wenxiang Hu, Fan Yang 0024, Lidong Zhou |
OSDI | 1 |
| 2019 | SeerNet: Predicting Convolutional Neural Network Feature-Map Sparsity Through Low-Bit QuantizationabstractIn this paper we present a novel and general method to accelerate convolutional neural network (CNN) inference by taking advantage of feature map sparsity. We experimentally demonstrate that a highly quantized version of the original network is sufficient in predicting the output sparsity accurately, and verify that leveraging such sparsity in inference incurs negligible accuracy drop compared with the original network. To accelerate inference, for each convolution layer our approach first obtains a binary sparsity mask of the output feature maps by running inference on a quantized version of the original network layer, and then conducts a full-precision sparse convolution to find out the precise values of the non-zero outputs. Compared with existing work, our approach avoids the overhead of training additional auxiliary networks, while is still applicable to general CNN networks without being limited to certain application domains. Shijie Cao, Lingxiao Ma, Wencong Xiao, Chen Zhang 0001, Yunxin Liu 0001, Lanshun Nie, Zhi Yang 0001 |
CVPR | 2 |
| 2019 | NeuGraph: Parallel Deep Neural Network Computation on Large Graphs
Lingxiao Ma, Zhi Yang 0001, Youshan Miao, Jilong Xue, Ming Wu 0007, Lidong Zhou, Yafei Dai |
USENIX ATC | 1 |
| 2017 | Garaph: Efficient GPU-accelerated Graph Processing on a Single Machine with Balanced Replication
Lingxiao Ma, Zhi Yang 0001, Jilong Xue, Yafei Dai |
USENIX ATC | 1 |
| 2005 | CoopStreaming: A Novel Peer-to-Peer System for Fast Live Media Streaming
Jianwei Yin, Weipeng Yao, Lingxiao Ma, Jinxiang Dong |
WAIM | 3 |