VLDB 2026 Research / reviewers in the wild / expert
Xinhao Cheng
dblp:158/3448
· DBLP profile ↗
6ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative DecodingabstractModern large language model (LLM) applications exhibit diverse service-level objectives (SLOs), from low-latency requirements in interactive coding assistants to more relaxed constraints in data wrangling tasks. Existing LLM serving systems, which rely on uniform batching and scheduling strategies, often fail to meet these heterogeneous SLOs concurrently. We present AdaServe, the first LLM serving system designed to support efficient multi-SLO serving through SLO-customized speculative decoding. AdaServe formulates multi-SLO serving as a constrained optimization problem and introduces a hardware-aware algorithm that constructs a speculation tree tailored to each request's latency target. It features a speculate-select-verify pipeline that enables fine-grained control over decoding speed while maximizing system throughput. AdaServe further adapts to workload variation by dynamically adjusting speculation parameters. Evaluations across diverse workloads show that AdaServe reduces SLO violations by up to 4.3X and improves goodput by up to 1.9X compared to the best-performing baselines, highlighting its effectiveness in multi-SLO serving. Zikun Li, Zhuofu Chen, Remi Delacourt, Gabriele Oliaro, Qinghan Chen, Shuhuai Lin, April Yang, Zhihao Zhang 0001, Zhuoming Chen, Yi-Hsiang Lai, Xinhao Cheng, Xupeng Miao |
EuroSys | 12 |
| 2026 | FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
Gabriele Oliaro, Xupeng Miao, Xinhao Cheng, Vineeth Kada, Mengdi Wu, Ruohan Gao, Yingyi Huang, Remi Delacourt, April Yang, Yingcheng Wang, Colin Unger |
NSDI | 3 |
| 2026 | Multi-State Reliability Modeling and Analysis for More Electric Aircraft Electrical Power System Considering State Transition UncertaintyabstractElectrical power system (EPS) of more electric aircraft (MEA) is a complex multi-state system, of which three issues of multi-state modeling, state transition uncertainty, and computational complexity pose challenges to its reliability modeling and analysis. This paper proposes a multi-state reliability modeling and analysis method for MEA EPS considering state transition uncertainty based on Markov process, fuzzy theory, and universal generating function (UGF). First, an equivalent structural model is established according to system architecture. For each component, a fuzzy state transition model is built based on triangular fuzzy numbers and Markov process. Then, the UGF method is integrated to express and infer the state probability and performance level distribution of system paths and branches, based on which three multi-state reliability indexes are calculated, respectively reliability, average performance, and performance shortfall. The propagation of component state transition uncertainty results in system reliability uncertainty. Their uncertain degree is quantified by a fuzzy entropy-based index. Finally, a case of MEA EPS is studied, and the comparative results show the effectiveness and superiority of the proposed method, supporting the reliability design and maintenance of MEA EPS. Xinhao Cheng, Jiayu Chen 0002, Hongjuan Ge, Peng Li 0038, Yinxiao Hu |
IEEE Trans. Reliab. | 1 |
| 2025 | Mirage: A Multi-Level Superoptimizer for Tensor Programs
Mengdi Wu, Xinhao Cheng, Chunan Shi, Jianan Ji, Man Kit Ao, Praveen Velliengiri, Xupeng Miao, Oded Padon |
OSDI | 2 |
| 2024 | SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and VerificationabstractThis paper introduces SpecInfer, a system that accelerates generative large language model (LLM) serving with tree-based speculative inference and verification. The key idea behind SpecInfer is leveraging small speculative models to predict the LLM's outputs; the predictions are organized as a token tree, whose nodes each represent a candidate token sequence. The correctness of all candidate token sequences represented by a token tree is verified against the LLM in parallel using a novel tree-based parallel decoding mechanism. SpecInfer uses an LLM as a token tree verifier instead of an incremental decoder, which significantly reduces the end-to-end latency and computational requirement for serving generative LLMs while provably preserving model quality. Our evaluation shows that SpecInfer outperforms existing LLM serving systems by 1.5-2.8× for distributed LLM inference and by 2.6-3.5× for offloading-based LLM inference, while preserving the same generative performance. SpecInfer is publicly available at https://github.com/flexflow/FlexFlow/ Xupeng Miao, Gabriele Oliaro, Zhihao Zhang 0001, Xinhao Cheng, Rae Ying Yee Wong, Alan Zhu 0001, Lijie Yang 0003, Xiaoxiang Shi, Chunan Shi, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar |
ASPLOS (3) | 4 |
| 2014 | Green Traffic Compression in Wireless Sensor NetworksabstractEmerging multi-hop machine-to-machine (M2M) communications that likely support a large number of wireless devices create new challenges for spectrum scarcity and energy efficiency. In parallel to pursuing physical layer transmission efficiency, traffic compression to reduce required wireless transmissions suggests a new paradigm of wireless networks. Utilizing the natures of broadcasting and information collection in wireless sensor or machine networks, cognitive traffic compression can be facilitated by our proposed optimal fusion rules and topology compression algorithm. Therefore, only the necessary and connected sensors/machines in M2M networks are required to transmit, to achieve the desirable distortion of information collection (i.e. detection/estimation error). In other words, given the desirable distortion, the number of sensors to transmit, or equivalently the total energy consumption, serves our purpose of energy efficiency for end-to-end networking. Numerical results show successful compression of total network traffic to significantly enhance networking energy efficiency. Kang-Hao Peng, Kwang-Cheng Chen, Shao-Lun Huang, Shao-Chou Hung, Xinhao Cheng |
VTC Spring | 5 |