EDBT 2026 Demo / reviewers in the wild / expert
Xiaoxiang Shi
dblp:25/3181
· DBLP profile ↗
5ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0000-6840-4691ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Agentix: An Efficient Serving Engine for LLM Agents as General Programs
Michael Luo, Xiaoxiang Shi, Colin Cai, Tianjun Zhang, Justin Wong, Yanping Huang, Joseph Gonzalez 0001, Ion Stoica |
NSDI | 2 |
| 2024 | SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and VerificationabstractThis paper introduces SpecInfer, a system that accelerates generative large language model (LLM) serving with tree-based speculative inference and verification. The key idea behind SpecInfer is leveraging small speculative models to predict the LLM's outputs; the predictions are organized as a token tree, whose nodes each represent a candidate token sequence. The correctness of all candidate token sequences represented by a token tree is verified against the LLM in parallel using a novel tree-based parallel decoding mechanism. SpecInfer uses an LLM as a token tree verifier instead of an incremental decoder, which significantly reduces the end-to-end latency and computational requirement for serving generative LLMs while provably preserving model quality. Our evaluation shows that SpecInfer outperforms existing LLM serving systems by 1.5-2.8× for distributed LLM inference and by 2.6-3.5× for offloading-based LLM inference, while preserving the same generative performance. SpecInfer is publicly available at https://github.com/flexflow/FlexFlow/ Xupeng Miao, Gabriele Oliaro, Zhihao Zhang 0001, Xinhao Cheng, Rae Ying Yee Wong, Alan Zhu 0001, Lijie Yang 0003, Xiaoxiang Shi, Chunan Shi, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar |
ASPLOS (3) | 10 |
| 2024 | Improving the Efficiency of Serverless Computing via Core-Level Power ManagementabstractServerless computing has recently become a significant application paradigm in data centers. However, existing power management methods focus on optimizations at the coarse-grained server level, making them unable to handle the characteristics of these short-lived, dynamic serverless functions. In this context, the unawareness of function-level characteristics by the existing power management systems can severely degrade the energy efficiency of the data centers. To address this challenge, we design a function-level power management system. Instead of relying on server-level schedulers, we propose a novel core-level scheduling policy for serverless functions that can efficiently allocate functions to the most suitable CPU core. Additionally, we propose a power management mechanism for serverless computing that can reduce system power consumption with functions’ QoS guaranteed. Our evaluation shows that our system achieves a maximum power saving of 8.5% and an average power saving of 8% across the majority of loads without incurring any loss in tail latency, as compared to the conventional server-level scheduling system. Du Liu, Jing Wang 0055, Xinkai Wang 0003, Chao Li 0009, Lu Zhang 0049, Xiaofeng Hou, Xiaoxiang Shi, Minyi Guo |
CCGrid | 7 |
| 2024 | Aceso: Efficient Parallel DNN Training through Iterative Bottleneck AlleviationabstractMany parallel mechanisms, including data parallelism, tensor parallelism, and pipeline parallelism, have been proposed and combined together to support training increasingly large deep neural networks (DNN) on massive GPU devices. Given a DNN model and GPU cluster, finding the optimal configuration by combining these parallelism mechanisms is an NP-hard problem. Widely adopted mathematical programming approaches have been proposed to search in a configuration subspace, but they are still too costly when scaling to large models over numerous devices. Youshan Miao, Xiaoxiang Shi, Saeed Maleki, Fan Yang 0024, Yungang Bao, Sa Wang |
EuroSys | 4 |
| 2008 | Reconfigurable acceleration of microphone array algorithms for speech enhancementabstractMicrophone arrays play an important role in noise reduction and speech enhancement. Their algorithms are based on beamforming, which reduces the level of localized and ambient noise signals while minimizing distortion to speech from the desired direction via spatial filtering. This paper describes a class of subband beamforming algorithms. The similarity between different algorithms is discussed. To enhance computational efficiency, the algorithms are implemented in frequency domain. A hardware architecture, with bitwidth optimization, is proposed to support the algorithms. An implementation with 7 instances on a Xilinx XC4VSX55 FPGA at 175MHz can run 41.7 times faster than the corresponding pure software implementation on a 3.2GHz Pentium 4 PC. Ka Fai Cedric Yiu, Chun Hok Ho, Nedelko Grbic, Xiaoxiang Shi, Wayne Luk |
ASAP | 5 |