EDBT 2026 Demo / reviewers in the wild / expert
Fahao Chen
dblp:305/3588
· DBLP profile ↗
18ranked-venue papers
11as first author
18since 2021 · last 2026
0000-0002-4345-1296ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 4 first-author · 8 since 2021Systems, architecture and hardware · 7 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Federated Fine-Tuning of Sparsely-Activated Large Language Models on Resource-Constrained DevicesabstractFederated fine-tuning of Mixture-of-Experts (MoE)-based large language models (LLMs) is challenging due to their massive computational requirements and the resource constraints of participants. Existing works attempt to fill this gap through model quantization, computation offloading, or expert pruning. However, they cannot achieve desired performance due to impractical system assumptions and a lack of consideration for MoE-specific characteristics. In this paper, we propose Flux, a system designed to enable federated fine-tuning of MoE-based LLMs across participants with constrained computing resources (e.g., consumer-grade GPUs), aiming to minimize time-to-accuracy. Flux introduces three key innovations: (1) quantization-based local profiling to estimate expert activation with minimal overhead, (2) adaptive layer-aware expert merging to reduce resource consumption while preserving accuracy, and (3) dynamic expert role assignment using an exploration-exploitation strategy to balance tuning and non-tuning experts. Extensive experiments on LLaMA-MoE and DeepSeek-MoE with multiple benchmark datasets demonstrate that Flux significantly outperforms existing methods, achieving up to 4.75× speedup in time-to-accuracy. Fahao Chen, Peng Li 0017, Zhou Su 0001, Dongxiao Yu |
EuroSys | 1 |
| 2026 | Director: Accelerating Distributed MoE Serving via Online Proactive Expert PlacementabstractExpert parallelism has become the prevailing paradigm to serve Mixture-of-Experts (MoE) models. Its efficiency depends on the communication and computation latencies of the GPUs, which are linked to the placement of experts in the GPUs. Existing works for optimizing expert placement focus on leveraging past requests' expert activation patterns. However, they demonstrate deficiencies facing diverse and rapidly changing request patterns, calling for an online, proactive approach. Implementing such an approach requires addressing several challenges: the uncertainty associated with incoming requests' expert activation, the cost of expert migration, and the NP-hard complexity in optimization. Therefore, we present Director, a new distributed MoE serving system that minimizes end-to-end latency via prediction-driven, online expert placement. Director uses either a lightweight cascaded predictor or a low-bit quantized replica for expert activation patterns of incoming requests. An online migration module then enacts the changes with near-zero downtime by executing migrations in compute-bound phases, keeping disruption bounded. At its core, a relaxation-based expert placement optimizer operates under capacity constraints, runs in polynomial time, and achieves a (1+ ϵ) approximation ratio. Finally, we implement a prototype and demonstrate, through extensive experiments, a reduction in end-to-end latency of 11 ~ 55% for popular MoE models (e.g., Mistral, DeepSeek and Qwen) compared to existing work. Qianli Liu, Kaibin Guo, Zicong Hong, Peng Li 0017, Fahao Chen, Song Guo 0001 |
INFOCOM | 5 |
| 2026 | Efficient Mixture-of-Experts Model Inference at the Edge via Adaptive Expert Merging
Ruirui Zhang 0003, Yifei Zou, Peng Li 0017, Fahao Chen, Yupeng Li 0001, Xiuzhen Cheng, Falko Dressler, Dongxiao Yu |
IEEE Trans. Netw. | 4 |
| 2025 | Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference LatencyabstractSpeculative decoding accelerates Large Language Model (LLM) inference by employing a small speculative model (SSM) to generate multiple candidate tokens and verify them using the LLM in parallel. This technique has been widely integrated into LLM inference serving systems. However, inference requests typically exhibit uncertain execution time, which poses a significant challenge of efficiently scheduling requests in these systems. Existing work estimates execution time based solely on predicted output length, which could be inaccurate because execution time depends on both output length and token acceptance rate of verification by the LLM. In this paper, we propose a semi-clairvoyant request scheduling algorithm called Least-Attained/Perceived-Service for Speculative Decoding (LAPS-SD). Given a number of inference requests, LAPS-SD can effectively minimize average inference latency by adaptively scheduling requests according to their features during decoding. When token acceptance rate is dynamic and execution time is difficult to estimate, LAPS-SD maintains multiple priority queues and allows request execution preemption across different queues. Once the token acceptance rate becomes stable, LAPS-SD can accurately estimate the execution time and schedule requests accordingly. Extensive experiments show that LAPS-SD reduces inference latency by approximately 39% compared to state-of-the-art scheduling methods. Ruixiao Li, Fahao Chen |
IJCAI | 2 |
| 2025 | SPIN: Accelerating Large Language Model Inference with Heterogeneous Speculative Models
Fahao Chen, Peng Li 0017, Tom H. Luan, Zhou Su 0001, Jing Deng 0001 |
INFOCOM | 1 |
| 2025 | Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
Qianli Liu, Zicong Hong, Peng Li 0017, Fahao Chen, Song Guo 0001 |
INFOCOM | 4 |
| 2025 | Corrections to "Giant Could Be Tiny: Efficient Inference of Giant Models on Resource-Constrained UAVs"abstractPresents corrections to the paper, (Corrections to “Giant Could Be Tiny: Efficient Inference of Giant Models on Resource-Constrained UAVs”). Fahao Chen, Peng Li 0017, Shengli Pan 0001, Jing Deng 0001 |
IEEE Internet Things J. | 1 |
| 2025 | Efficient multi-job federated learning scheduling with fault tolerance
Boqian Fu, Fahao Chen, Shengli Pan 0001, Peng Li 0017, Zhou Su 0001 |
Peer Peer Netw. Appl. | 2 |
| 2025 | Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token CondensationabstractMixture-of-Experts (MoE) is an emerging technique for scaling large models with sparse activation. MoE models are typically trained in a distributed manner with anexpert parallelismscheme, where experts in each MoE layer are distributed across multiple GPUs. However, the default expert parallelism suffers from the heavy network burden due to the all-to-all intermediate data exchange among GPUs before and after the expert run. Some existing works have proposed to reduce intermediate data exchanges by transferring experts to reduce the network loads, however, which would decrease parallelism level of expert execution and make computation inefficient. The weaknesses of existing works motivate us to explore whether it is possible to reduce inter-GPU traffic while maintaining a high degree of expert parallelism. This paper gives a positive response by presentingLuffy, a communication-efficient distributed MoE training system with two new techniques. First,Luffymigrates sequences among GPUs to hide heavy token pulling paths within GPUs and avoid copying experts over GPUs. Second, we propose token condensation that identifies similar tokens and then eliminates redundant transmissions. We implementLuffybased on PyTorch and evaluate its performance on a testbed of 16 V100 GPUs.Luffysystem can achieve a speedup of up to$2.73\times $compared to state-of-the-art MoE training systems. Fahao Chen, Peng Li 0017, Zicong Hong, Zhou Su 0001, Song Guo 0001 |
IEEE Trans. Netw. | 1 |
| 2025 | Serving Transformer Models via Joint Requst Scheduling and Batching in the Network EdgeabstractTransformers have dominated the field of natural language processing, attributed to their capability to handle sequential input data. There is a surge of work on computational and networking optimizations, aimed at improving the training efficiency of Transformers. However, transformer inference, a cornerstone of myriad AI services, remains relatively underexplored. With the challenge of variable-length inputs, conventional methods adopt padding schemes, resulting in computational waste. Moreover, works on transformer inference often overlook the integration between request scheduling and batching, which play pivotal roles in inference systems. To address these challenges, we introduce TCB, a comprehensiveTransformer inference system that integrates aConcatBatching scheme to reduce computational redundancy by concatenating requests. In addition, we present an online request batching algorithm, designed to augment the throughput of scheduled requests. Consider a muiti-server case, we further introduce a joint request assignment and batching scheduling policy to fully utilize resources on servers while ensuring quality-of-service of inference. Extensive experiments demonstrate that our proposed methods can significantly outperform existing works. Boqian Fu, Fahao Chen, Peng Li 0017, Deze Zeng |
IEEE Trans. Sustain. Comput. | 2 |
| 2024 | Giant Could Be Tiny: Efficient Inference of Giant Models on Resource-Constrained UAVsabstractGiant models, characterized by their billions or even trillions of parameters, has demonstrated unprecedented capabilities in handling complex tasks on Artificial intelligence (AI)-driven UAVs, such as disaster relief, aerial navigation, and manipulation. However, there is an open challenge about the mismatching between the massive computation and memory requirements of giant models and the limited resources on UAVs. Existing works either pose privacy concerns with offloading methods or compromise model accuracy with various model compression techniques. In this paper, we fill the gap by exploiting the Mixture-of-Expert (MoE) model architecture that decouples giant models into multiple tiny experts, so that UAVs can dynamically load a few experts that best match their current input. We consider a general scenario of several edge servers feeding experts to multiple UVAs and formulate a core problem of expert selection and UAV-edge association. Due to the high complexity of this problem, we propose a solution, termed GESolver, based on graph learning, which automatically solves the problem by learning the complicated interaction between edge servers, UAVs, as well as their required experts. We evaluate our proposed method with three popular MoE-based models under various problem settings. The experiments demonstrate that our proposed method can significantly outperform other baselines. Fahao Chen, Peng Li 0017, Shengli Pan 0001, Jing Deng 0001 |
IEEE Internet Things J. | 1 |
| 2024 | Non-Clairvoyant Scheduling of Distributed Machine Learning With Inter-Job and Intra-Job Parallelism on Heterogeneous GPUsabstractDistributed machine learning (DML) has shown great promise in accelerating model training on multiple GPUs. To increase GPU utilization, a common practice is to let multiple learning jobs share GPU clusters, where the most fundamental and critical challenge is how to efficiently schedule these jobs on GPUs. However, existing works about DML job scheduling are constrained to settings with homogeneous GPUs. GPU heterogeneity is common in practice, but its influence on multiple DML job scheduling has been seldom studied. Moreover, DML jobs have internal structures that contain great parallelism potentials, which have not yet been fully exploited in the heterogeneous computing environment. In this paper, we proposeHare, a DML job scheduler that exploits both inter-job and intra-job parallelism in a heterogeneous GPU cluster.Hareadopts a relaxed fixed-scale synchronization scheme that allows independent tasks to be flexibly scheduled within a training round. Given full knowledge of job arrival time and sizes, we propose a fast heuristic algorithm to minimize the average job completion time and derive its theoretical bound is derived. Without prior knowledge of jobs, we propose an online algorithm based on the Heterogeneity-aware Least-Attained Service (HLAS) policy. We evaluateHareusing a small-scale testbed and a trace-driven simulator. The results show that it can outperform the state-of-the-art, achieving a performance improvement of about 2.94×. Fahao Chen, Peng Li 0017, Celimuge Wu, Song Guo 0001 |
IEEE Trans. Cloud Comput. | 1 |
| 2023 | Low-Latency Perception Sharing Services for Connected Autonomous VehiclesabstractConnected autonomous vehicles (CAVs) are promising to improve road safety, thanks to various on-board sensors, such as LiDAR, radars, and stereo cameras. However, perception view could be significantly limited due to occlusions, extreme weather, and far objects. To address these challenges, in this paper, we propose an efficient edge-assisted perception sharing scheme, which enables vehicles to exchange the information about their sensed environment to improve road safety. We formulate perception sharing as an online optimization problem, with the objective of maximizing the total weighted utility, where utility indicates the quality of collected sensor data while weight means the intensity of the vehicle's demand for information in a certain area. To solve this problem, we propose an efficient online heuristic algorithm, which decouples the original problem into multiple sub-problems and solves them alternatively to find the optimal solution. Extensive simulations demonstrate that our proposed method can significantly improve the perception sharing performance. Fahao Chen, Peng Li 0017, Dongxiao Yu, Xiuzhen Cheng |
VTC Fall | 1 |
| 2023 | DGC: Training Dynamic Graphs with Spatio-Temporal Non-Uniformity using Graph Partitioning by ChunksabstractDynamic Graph Neural Network (DGNN) has shown a strong capability of learning dynamic graphs by exploiting both spatial and temporal features. Although DGNN has recently received considerable attention by AI community and various DGNN models have been proposed, building a distributed system for efficient DGNN training is still challenging. It has been well recognized that how to partition the dynamic graph and assign workloads to multiple GPUs plays a critical role in training acceleration. Existing works partition a dynamic graph into snapshots or temporal sequences, which only work well when the graph has uniform spatio-temporal structures. However, dynamic graphs in practice are not uniformly structured, with some snapshots being very dense while others are sparse. To address this issue, we propose DGC, a distributed DGNN training system that achieves a 1.25× - 7.52× speedup over the state-of-the-art in our testbed. DGC's success stems from a new graph partitioning method that partitions dynamic graphs into chunks, which are essentially subgraphs with modest training workloads and few inter connections. This partitioning algorithm is based on graph coarsening, which can run very fast on large graphs. In addition, DGC has a highly efficient run-time, powered by the proposed chunk fusion and adaptive stale aggregation techniques. Extensive experimental results on 3 typical DGNN models and 4 popular dynamic graph datasets are presented to show the effectiveness of DGC. Fahao Chen, Peng Li 0017, Celimuge Wu |
Proc. ACM Manag. Data | 1 |
| 2023 | Edge-Assisted Short Video Sharing With Guaranteed Quality-of-ExperienceabstractAs a rising star of social apps, short video apps, e.g., TikTok, have attracted a large number of mobile users by providing fresh and short video contents that highly match their watching preferences. Meanwhile, the booming growth of short video apps imposes new technical challenges on the existing computation and communication infrastructure. Traditional solutions maintain all videos on the cloud and stream them to users via contend delivery networks or the Internet. However, they incur huge network traffic and long delay that seriously affects users’ watching experiences. In this article, we propose an edge-assisted short video sharing framework to address these challenges by caching some highly preferred videos at edge servers that can be accessed by users via high-speed network connections. Since edge servers have limited computation and storage resources, we design an online algorithm with provable approximation ratio to decide which videos should be cached at edge servers, without the knowledge of future network quality and watching preferences changes. Furthermore, we improve the performance by jointly considering video fetching and user-edge association. Extensive simulations are conducted to evaluate the proposed algorithms under various system settings, and the results show that our proposals outperform existing schemes. Fahao Chen, Peng Li 0017, Deze Zeng, Song Guo 0001 |
IEEE Trans. Cloud Comput. | 1 |
| 2022 | Hare: Exploiting Inter-job and Intra-job Parallelism of Distributed Machine Learning on Heterogeneous GPUsabstractDistributed machine learning (DML) has shown great promise in accelerating model training on multiple GPUs. To increase GPU utilization, a common practice is to let multiple learning jobs share GPU clusters, where the most fundamental and critical challenge is how to efficiently schedule these jobs on GPUs. However, existing works about DML job scheduling are constrained to settings with homogeneous GPUs. GPU heterogeneity is common in practice, but its influence on multiple DML job scheduling has been seldom studied. Moreover, DML jobs have internal structures that contain great parallelism potentials, which have not yet been fully exploited in the heterogeneous computing environment. In this paper, we propose Hare, a DML job scheduler that exploits both inter-job and intra-job parallelism in a heterogeneous GPU cluster. Hare has three novel designs. First, Hare optimizes GPU execution environment to reduce task switching overhead by exploiting unique features of DML scheduling. Second, Hare adopts a relaxed fixed-scale synchronization scheme that allows independent tasks to be flexibly scheduled within a training round. Finally, we propose a fast heuristic algorithm to minimize the total weighted job completion time by jointly considering job features and hardware heterogeneity. Its theoretical bound is derived. We evaluate Hare using a small-scale testbed and a trace-driven simulator. The results show that it can outperform the state-of-the-art by about 2x. Fahao Chen, Peng Li 0017, Celimuge Wu, Song Guo 0001 |
HPDC | 1 |
| 2022 | TCB: Accelerating Transformer Inference Services with Request ConcatenationabstractTransformer has dominated the field of natural language processing because of its strong capability in learning from sequential input data. In recent years, various computing and networking optimizations have been proposed for improving transformer training efficiency. However, transformer inference, as the core of many AI services, has been seldom studied. A key challenge of transformer inference is variable-length input. In order to align these input, existing work has proposed batching schemes by padding zeros, which unfortunately introduces significant computational redundancy. Moreover, existing transformer inference studies are separated from the whole serving system, where both request batching and request scheduling are critical and they have complex interaction. To fill the research gap, we propose TCB, a Transformer inference system with a novel ConcatBatching scheme as well as a jointly designed online scheduling algorithm. ConcatBatching minimizes computational redundancy by concatenating multiple requests, so that batch rows can be aligned with reduced padded zeros. Moreover, we conduct a systemic study by designing an online request scheduling algorithm aware of ConcatBatching. This scheduling algorithm needs no future request information and has provable theoretical guarantee. Experimental results show that TCB can significantly outperform state-of-the-art. Boqian Fu, Fahao Chen, Peng Li 0017, Deze Zeng |
ICPP | 2 |
| 2022 | FedGraph: Federated Graph Learning With Intelligent SamplingabstractFederated learning has attracted much research attention due to its privacy protection in distributed machine learning. However, existing work of federated learning mainly focuses on Convolutional Neural Network (CNN), which cannot efficiently handle graph data that are popular in many applications. Graph Convolutional Network (GCN) has been proposed as one of the most promising techniques for graph learning, but its federated setting has been seldom explored. In this article, we propose FedGraph for federated graph learning among multiple computing clients, each of which holds a subgraph. FedGraph provides strong graph learning capability across clients by addressing two unique challenges. First, traditional GCN training needs feature data sharing among clients, leading to risk of privacy leakage. FedGraph solves this issue using a novel cross-client convolution operation. The second challenge is high GCN training overhead incurred by large graph size. We propose an intelligent graph sampling algorithm based on deep reinforcement learning, which can automatically converge to the optimal sampling policies that balance training speed and accuracy. We implement FedGraph based on PyTorch and deploy it on a testbed for performance evaluation. The experimental results of four popular datasets demonstrate that FedGraph significantly outperforms existing work by enabling faster convergence to higher accuracy. Fahao Chen, Peng Li 0017, Toshiaki Miyazaki, Celimuge Wu |
IEEE Trans. Parallel Distributed Syst. | 1 |