EDBT 2026 Demo / reviewers in the wild / expert
Boyi Liu 0002
dblp:179/0897-2
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0004-5894-7878ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Harnessing Asynchrony to Balance Modalities in Multi-modal Federated Learning
Yiming Ma 0005, Boyi Liu 0002, Zimu Zhou, Yongxin Tong |
DASFAA (3) | 2 |
| 2026 | Towards Asynchronous Client Collaboration in Personalized Federated Learning
Boyi Liu 0002, Zimu Zhou, Yongxin Tong |
INFOCOM | 1 |
| 2026 | FedMosaic: Federated Retrieval-Augmented Generation via Parametric AdaptersabstractRetrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by grounding generation in external knowledge to improve factuality and reduce hallucinations. Yet most deployments assume a centralized corpus, which is infeasible in privacy-aware domains where knowledge remains siloed. This motivates federated RAG (FedRAG), where a central LLM server collaborates with distributed silos without sharing raw documents. In-context RAG violates this requirement by transmitting verbatim documents, whereas parametric RAG encodes documents into light weight adapters that merge with a frozen LLM at inference, avoiding raw-text exchange. We adopt the parametric approach but face two unique challenges induced by FedRAG: high storage and communication from per-document adapters, and destructive aggregation caused by indiscriminately merging multiple adapters. We present FedMosaic, the first federated RAG framework built on parametric adapters. FedMosaic clusters semantically related documents into multi-document adapters with document-specific masks to reduce overhead while preserving specificity, and performs selective adapter aggregation to combine only relevance-aligned, non-conflicting adapters. Experiments show that FedMosaic achieves an average 10.9% higher accuracy than state-of-the-art methods in four categories, while lowering storage costs by 78.8% to 86.3% and communication costs by 91.4%, and never sharing raw documents. Zhilin Liang, Yuxiang Wang 0014, Zimu Zhou, Hainan Zhang 0001, Boyi Liu 0002, Yongxin Tong |
SIGIR | 5 |
| 2025 | DarkDistill: Difficulty-Aligned Federated Early-Exit Network Training on Heterogeneous DevicesabstractEarly-exit networks (EENs), which adapt their computational depths based on input samples, are widely adopted to accelerate inference in edge computing applications. The effectiveness of EENs relies on difficulty-aware training, which tailors shallow exits for simple samples and deep exits for complex ones. However, existing difficulty-aware training schemes assume centralized environments with sufficient data, which become invalid with real-world edge devices. In this paper, we explore difficulty-aware training in a federated manner, where EENs are collaboratively trained on heterogeneous devices. We observe the cross-model exit unalignment phenomenon, a unique problem when aggregating local EENs into a cohesive global model. To address this problem, we design a novel Difficulty-Aligned Reverse Knowledge Distillation scheme named DarkDistill that preserves the difficulty-specific specialization for aggregating heterogeneous local models. Instead of direct parameter averaging, it trains difficulty-conditional data generators, and selectively transfers generated knowledge of specific difficulty among matched exits of heterogeneous EENs. Evaluations show that DarkDistill outperforms the state-of-the-arts in both full-parameter and parameter-efficient fine-tuning of EENs. Lehao Qu, Shuyuan Li, Zimu Zhou, Boyi Liu 0002, Yi Xu 0013, Yongxin Tong |
KDD (2) | 4 |
| 2025 | Poster: Asynchronous Federated Learning Library and Benchmark with AFL-LibabstractAsynchronous Federated Learning (AFL) emerges as a practical paradigm for collaborative model training across IoT devices with heterogeneous compute capabilities, network bandwidth, and online availability. Yet AFL research still lacks a unified, easy-to-use platform for reproducible experimentation under such system configurations. We present AFL-Lib, the first open-source library and benchmark de-signed for AFL. It allows researchers to flexibly configure device types, network conditions, and availability patterns to emulate the system heterogeneity that gives rise to model staleness, a key factor affecting AFL algorithm design. Beyond system-level settings, AFL-Lib supports plug-in modules for personalized and multi-modal federated training, enabling exploration of the interplay between system and data heterogeneity. AFL-Lib implements 10 state-of-the-art AFL algorithms and 4 synchronous baselines, and integrates 12 datasets spanning image, text, and sensor. We will continue to expand AFL-Lib with new algorithms, datasets, and features to support ongoing AFL research. All code and data are publicly available at https://github.com/boyi-liu/AFL-Lib. Boyi Liu 0002, Shuyuan Li, Zimu Zhou, Yiming Ma 0005, Yongxin Tong |
MobiCom | 1 |
| 2024 | CASA: Clustered Federated Learning with Asynchronous ClientsabstractClustered Federated Learning (CFL) is an emerging paradigm to extract insights from data on IoT devices. Through iterative client clustering and model aggregation, CFL adeptly manages data heterogeneity, ensures privacy, and delivers personalized models to heterogeneous devices. Traditional CFL approaches, which operate synchronously, suffer from prolonged latency for waiting slow devices during clustering and aggregation. This paper advocates a shift to asynchronous CFL, allowing the server to process client updates as they arrive. This shift enhances training efficiency yet introduces complexities to the iterative training cycle. To this end, we present CASA, a novel CFL scheme for Clustering-Aggregation Synergy under Asynchrony. Built upon a holistic theoretical understanding of asynchrony's impact on CFL, CASA adopts a bi-level asynchronous aggregation method and a buffer-aided dynamic clustering strategy to harmonize between clustering and aggregation. Extensive evaluations on standard benchmarks show that CASA outperforms representative baselines in model accuracy and achieves 2.28-6.49× higher convergence speed. Boyi Liu 0002, Yiming Ma 0005, Zimu Zhou, Yexuan Shi, Shuyuan Li, Yongxin Tong |
KDD | 1 |
| 2022 | Dynamic Graph Learning Based on Hierarchical Memory for Origin-Destination Demand PredictionabstractRecent years have witnessed a rapid growth of applying deep spatiotemporal methods in traffic forecasting. However, the prediction of origin-destination (OD) demands is still a challenging problem since the number of OD pairs is usually quadratic to the number of stations. In this case, most of the existing spatiotemporal methods fail to handle spatial relations on such a large scale. To address this problem, this paper provides a dynamic graph representation learning framework for OD demands prediction. In particular, a hierarchical memory updater is first proposed to maintain a time-aware representation for each node, and the representations are updated according to the most recently observed OD trips in continuous-time and multiple discrete-time ways. Second, a spatiotemporal propagation mechanism is provided to aggregate representations of neighbor nodes along a random spatiotemporal route which treats origin and destination as two different semantic entities. Last, an objective function is designed to derive the future OD demands according to the most recent node representations, and also to tackle the data sparsity problem in OD prediction. Extensive experiments have been conducted on two real-world datasets, and the experimental results demonstrate the superiority of the proposed method. The code and data are available at https://github.com/Rising0321/HMOD. Ruixing Zhang, Liangzhe Han, Boyi Liu 0002, Jiayuan Zeng, Leilei Sun |
IJCAI | 3 |