EDBT 2026 Demo / reviewers in the wild / expert
Boan Liu
dblp:14/9228
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0002-5877-9999ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OWL: Worker-assisted server bandwidth optimization for efficient communication federated learningabstractEdge computing in federated learning based on centralized architecture often faces communication constraints in large clusters. Although there have been some efforts like computation-communication overlapping and fine-granularity flow scheduling towards how to reduce the communication cost, this is still a matter of ongoing research. Motivated by the underutilization of bandwidth among workers (edge devices) and the replication of deep neural network (DNN) model distributions in data-parallel federated learning, we propose OWL, a novel worker-assisted server bandwidth optimization method. OWL partitions numerous computation branches into groups based on the model's network topology , allowing for overlapping model distribution and computation among workers, thereby leveraging idle communication resources on the workers to compensate for server bandwidth. To address the issue of model distribution congestion on the server, we formulate group partition as an optimization problem , which proves to be NP-hard. We tackle this problem through a divide-and-conquer approach employing an approximation grouping algorithm and a deploying algorithm. Finally, we evaluate the performance of OWL through simulations and a comprehensive real-world case study involving model training on OWL and deployment on edge systems. Experimental results demonstrate that OWL reduces overall training time by up to 20%-69% and improves scalability by over 9.5% compared to state-of-the-art overlapping approaches. Boan Liu, Chuang Hu, Dazhao Cheng |
J. Parallel Distributed Comput. | 2 |
| 2025 | Spread+: Scalable Model Aggregation in Federated Learning With Non-IID DataabstractFederated learning (FL) addresses privacy concerns by training models without sharing raw data, overcoming the limitations of traditional machine learning paradigms. However, the rise of smart applications has accentuated the heterogeneity in data and devices, which presents significant challenges for FL. In particular, data skewness among participants can compromise model accuracy, while diverse device capabilities lead to aggregation bottlenecks, causing severe model congestion. In this article, we introduce Spread+, a hierarchical system that enhances FL by organizing clients into clusters and delegating model aggregation to edge devices, thus mitigating these challenges. Spread+ leverages hedonic coalition formation game to optimize customer organization and adaptive algorithms to regulate aggregation intervals within and across clusters. Moreover, it refines the aggregation algorithm to boost model accuracy. Our experiments demonstrate that Spread+ significantly alleviates the central aggregation bottleneck and surpasses mainstream benchmarks, achieving performance improvements of 49.58% over FAVG and 22.78% over Ring-allreduce. Huanghuang Liang, Boan Liu, Chuang Hu, Dan Wang 0002, Xiaobo Zhou 0002, Dazhao Cheng |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2024 | Diversifying the Mixture-of-Experts Representation for Language Models with Orthogonal OptimizerabstractThe Mixture of Experts (MoE) has emerged as a highly successful technique in deep learning, based on the principle of divide-and-conquer to maximize model capacity without significant additional computational cost. Even in the era of large-scale language models (LLMs), MoE continues to play a crucial role, as some researchers have indicated that GPT-4 adopts the MoE structure to ensure diverse inference results. However, MoE is susceptible to performance degeneracy, particularly evident in the issues of imbalance and homogeneous representation among experts. While previous studies have extensively addressed the problem of imbalance, the challenge of homogeneous representation remains unresolved. In this study, we shed light on the homogeneous representation problem, wherein experts in the MoE fail to specialize and lack diversity, leading to frustratingly high similarities in their representations (up to 99% in a well-performed MoE model). This problem restricts the expressive power of the MoE and, we argue, contradicts its original intention. To tackle this issue, we propose a straightforward yet highly effective solution: OMoE, an orthogonal expert optimizer. Additionally, we introduce an alternating training strategy that encourages each expert to update in a direction orthogonal to the subspace spanned by other experts. Our algorithm facilitates MoE training in two key ways: firstly, it explicitly enhances representation diversity, and secondly, it implicitly fosters interaction between experts during orthogonal weights computation. Through extensive experiments, we demonstrate that our proposed optimization algorithm significantly improves the performance of fine-tuning the MoE model on the GLUE benchmark, SuperGLUE benchmark, question-answering task, and name entity recognition tasks. Boan Liu, Liang Ding 0006, Li Shen 0008, Keqin Peng, Yu Cao 0014, Dazhao Cheng, Dacheng Tao |
ECAI | 1 |
| 2024 | Diffindo: Accelerating Distributed GANs with Auxiliary Generators and DiscriminatorsabstractIntegrating edge computing with Generative Adversarial Networks (GANs) leads to significant communication overhead in centralized federated learning systems due to frequent synchronization between generators and discriminators, as well as the transmission of large generated samples. In addition, this close coupling can cause excessive GPU memory consumption, especially with coarse-grained deployment strategies, resulting in memory thrashing and reduced training speeds. To tackle these issues, we present Diffindo, a novel distributed GAN training system that reduces training time and optimizes GPU memory allocation while maintaining accuracy. By utilizing fine-grained deployment strategies and developing innovative task scheduling algorithms for servers and workers, we enhance training efficiency. Additionally, we implement a computation-communication overlapping strategy to improve resource utilization. Experimental results show that Diffindo outperforms state-of-the-art GAN training systems, achieving 13% higher accuracy and 32% faster training speeds. Boan Liu, Dazhao Cheng |
HPCC | 2 |
| 2024 | JediGAN: A Fully Decentralized Training of GAN with Adaptive Discriminator Averaging and Generator Selection
Boan Liu, Dazhao Cheng |
NPC (1) | 2 |
| 2023 | PAD-Net: An Efficient Framework for Dynamic NetworksabstractDynamic networks, e.g., Dynamic Convolution (DY-Conv) and the Mixture of Experts (MoE), have been extensively explored as they can considerably improve the model's representation power with acceptable computational cost.The common practice in implementing dynamic networks is to convert the given static layers into fully dynamic ones where all parameters are dynamic (at least within a single layer) and vary with the input.However, such a fully dynamic setting may cause redundant parameters and high deployment costs, limiting the applicability of dynamic networks to a broader range of tasks and models.The main contributions of our work are challenging the basic commonsense in dynamic networks and proposing a partially dynamic network, namely PAD-Net, to transform the redundant dynamic parameters into static ones.Also, we further design Iterative Mode Partition to partition dynamic and static parameters efficiently.Our method is comprehensively supported by large-scale experiments with two typical advanced dynamic architectures, i.e., DY-Conv and MoE, on both image classification and GLUE benchmarks.Encouragingly, we surpass the fully dynamic networks by +0.7% top-1 acc with only 30% dynamic parameters for ResNet-50 and +1.9% average score in language understanding with only 50% dynamic parameters for BERT.Code will be released Shwai He, Liang Ding 0006, Daize Dong, Boan Liu, Fuqiang Yu, Dacheng Tao |
ACL (1) | 4 |
| 2022 | Spread: Decentralized Model Aggregation for Scalable Federated LearningabstractFederated learning (FL) is a new distributed machine learning paradigm that enables machine learning on edge devices. One unique feature of FL is that edge devices belong to individuals; and since they are not “owned” by the FL coordinator, but can be “federated” instead, there can potentially be a huge number of edge devices. In the current distributed ML architecture, the parameter server (PS) architecture, model aggregation is centralized. When facing a large number of edge devices, the centralized model aggregation becomes the bottleneck and fundamentally restricts system scalability. Chuang Hu, Huanghuang Liang, Boan Liu, Dazhao Cheng, Dan Wang 0002 |
ICPP | 4 |