EDBT 2026 Demo / reviewers in the wild / expert
Mengdi Wu
dblp:283/8285
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0001-6672-1252ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Cloud and datacenter computing · 28% Distributed systems · 25% Parallel and multicore computing · 25% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% | |
| Artificial intelligence
2 papers |
Efficient and distributed learning · 100% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing › inference serving
LLM serving |
1.0 | 1 | 2026 | FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees · NSDI 2026 |
Compilers and program optimization › compiler optimization
superoptimization |
0.9 | 1 | 2025 | Mirage: A Multi-Level Superoptimizer for Tensor Programs · OSDI 2025 |
Compilers and program optimization › deep learning compiler
tensor program optimization |
0.9 | 1 | 2025 | Mirage: A Multi-Level Superoptimizer for Tensor Programs · OSDI 2025 |
Distributed systems
distributed machine learning |
0.9 | 1 | 2025 | GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism · ASPLOS (1) 2025 |
Parallel and multicore computing
pipeline parallelism |
0.9 | 1 | 2025 | GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism · ASPLOS (1) 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.6 | 1 | 2022 | Finding the Task-Optimal Low-Bit Sub-Distribution in Deep Neural Networks · ICML 2022 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.6 | 1 | 2022 | Finding the Task-Optimal Low-Bit Sub-Distribution in Deep Neural Networks · ICML 2022 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.5 | 2 | 2025 | Mirage: A Multi-Level Superoptimizer for Tensor Programs · OSDI 2025 GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism · ASPLOS (1) 2025 |
Machine learning › Efficient and distributed learning
inference serving |
0.3 | 1 | 2026 | FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees · NSDI 2026 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN training |
0.3 | 1 | 2025 | GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism · ASPLOS (1) 2025 |
Methods — techniques the papers use, named apart from their topics
token-level scheduling · 2.0model partitioning · 0.9graph pipeline parallelism · 0.9task-objective optimization · 0.6gaussian mixture · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Identity Testing for Circuits with Exponentiation GatesabstractMotivated by practical applications in the design of optimization compilers for neural networks, we initiated the study of identity testing problems for arithmetic circuits augmented with exponentiation gates that compute the real function x↦ e^x. These circuits compute real functions of form P(→x)/P'(→x), where both P(→x) and P'(→x) are exponential polynomials ∑_{i = 1}^k f_i(→x)⋅ exp((g_i(→x))/(h_i(→x))), for polynomials f_i(→x),g_i(→x), and h_i(→x). We formalize a black-box query model over finite fields for this class of circuits, which is mathematical simple and reflects constraints faced by real-world neural network compilers. We proved that a simple and efficient randomized identity testing algorithm achieves perfect completeness and non-trivial soundness. Concurrent with our work, the algorithm has been implemented in the optimization compiler Mirage by Wu et al. (OSDI 2025), demonstrating promising empirical performance in both efficiency and soundness error. Finally, we propose a number-theoretic conjecture under which our algorithm is sound with high probability. Jiatu Li, Mengdi Wu |
ITCS | 2 |
| 2026 | FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
Gabriele Oliaro, Xupeng Miao, Xinhao Cheng, Vineeth Kada, Mengdi Wu, Ruohan Gao, Yingyi Huang, Remi Delacourt, April Yang, Yingcheng Wang, Colin Unger |
NSDI | 5 |
| 2025 | GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline ParallelismabstractDeep neural networks (DNNs) continue to grow rapidly in size, making them infeasible to train on a single device (e.g. GPU). Pipeline parallelism is commonly used in existing DNN systems to support large-scale DNN training by partitioning a DNN into multiple stages, which concurrently perform DNN computation for different micro-batches of training samples in a pipeline fashion. However, existing pipeline-parallel approaches only consider sequential pipeline stages and thus ignore the topology of a DNN, resulting in missed model-parallel opportunities. Byungsoo Jeon, Mengdi Wu, Shiyi Cao, Sunghyun Park 0004, Neeraj Aggarwal, Colin Unger, Daiyaan Arfeen, Peiyuan Liao, Xupeng Miao, Mohammad Alizadeh, Gregory R. Ganger, Tianqi Chen 0001 |
ASPLOS (1) | 2 |
| 2025 | Mirage: A Multi-Level Superoptimizer for Tensor Programs
Mengdi Wu, Xinhao Cheng, Chunan Shi, Jianan Ji, Man Kit Ao, Praveen Velliengiri, Xupeng Miao, Oded Padon |
OSDI | 1 |
| 2024 | LACO: A Latency-Constraint Offline Neural Network Scheduler towards Reliable Self-Driving PerceptionabstractPerception is a crucial and compute-intensive self-driving stage where multiple DNN models periodically process image and point-cloud inputs with constant intervals as latency constraints. This scenario is similar to the multistream scenario in the MLPerf benchmark but with more than one query sequence of diverse intervals from different sensors, and each query also evokes more than one model. In this paper, we call this scenario multi-multistream (M2-stream). Because the periodic inputs arrive continuously, the neural network processor should process them in time to avoid information dropping and maintain self-driving safety. Kaisheng Ma, Zhanhong Tan, Mengdi Wu |
ICCAD | 4 |
| 2022 | Finding the Task-Optimal Low-Bit Sub-Distribution in Deep Neural NetworksabstractQuantized neural networks typically require smaller memory footprints and lower computation complexity, which is crucial for efficient deployment. However, quantization inevitably leads to a distribution divergence from the original network, which generally degrades the performance. To tackle this issue, massive efforts have been made, but most existing approaches lack statistical considerations and depend on several manual configurations. In this paper, we present an adaptive-mapping quantization method to learn an optimal latent sub-distribution that is inherent within models and smoothly approximated with a concrete Gaussian Mixture (GM). In particular, the network weights are projected in compliance with the GM-approximated sub-distribution. This sub-distribution evolves along with the weight update in a co-tuning schema guided by the direct task-objective optimization. Sufficient experiments on image classification and object detection over various modern architectures demonstrate the effectiveness, generalization property, and transferability of the proposed method. Besides, an efficient deployment flow for the mobile CPU is developed, achieving up to 7.46$\times$ inference acceleration on an octa-core ARM CPU. Our codes have been publicly released at https://github.com/RunpeiDong/DGMS. Runpei Dong, Zhanhong Tan, Mengdi Wu, Linfeng Zhang 0001, Kaisheng Ma |
ICML | 3 |