EDBT 2026 Demo / reviewers in the wild / expert
Bingshuai Li
dblp:269/4522
· DBLP profile ↗
15ranked-venue papers
0as first author
14since 2021 · last 2025
0000-0001-7655-6919ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 6 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MIDDLE: A Mobility-Driven Device-Edge-Cloud Federated Learning FrameworkabstractFederated learning (FL) can be implemented in large-scale wireless networks in a hierarchical way, introducing edge servers as relays between the cloud server and devices. These devices are dispersed within multiple clusters coordinated by edges. However, the devices are typically mobile users with unpredictable trajectories, and the impact of their mobility on the model training process is not well-studied. In this work, we propose a newMobIlity-Driven feDeratedLEarning framework, namely MIDDLE. MIDDLE addresses unbalanced model updates by capitalizing on model aggregation opportunities on mobile devices due to their mobility across edges. It consists of two components: on-device model aggregation, which aggregates models from different edges carried by mobile devices as they move across edges, and in-edge device selection, adjusting the current edge optimization direction through careful device selection. Theoretical analysis emphasizes that on-device model aggregation can reduce bias in model updating on edges and the cloud, thereby accelerating the FL model convergence. Building on this analysis, we introduce on-device global control averaging, modifying the training process on mobile devices and extending MIDDLE into$\text{MIDDLE}^{+}$. Extensive experimental results validate that MIDDLE and$\text{MIDDLE}^{+}$can reduce the time steps to reach the target accuracy by 19.44% and 20.37% at least, respectively. Songli Zhang, Zhenzhe Zheng 0001, Fan Wu 0006, Bingshuai Li, Yunfeng Shao 0001, Guihai Chen |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Ents: An Efficient Three-party Training Framework for Decision Trees by Communication OptimizationabstractMulti-party training frameworks for decision trees based on secure multi-party computation enable multiple parties to train high-performance models on distributed private data with privacy preservation. The training process essentially involves frequent dataset splitting according to the splitting criterion (e.g. Gini impurity). However, existing multi-party training frameworks for decision trees demonstrate communication inefficiency due to the following issues: (1) They suffer from huge communication overhead in securely splitting a dataset with continuous attributes. (2) They suffer from huge communication overhead due to performing almost all the computations on a large ring to accommodate the secure computations for the splitting criterion. Guopeng Lin, Weili Han, Wenqiang Ruan, Ruisheng Zhou, Lushan Song, Bingshuai Li, Yunfeng Shao 0001 |
CCS | 6 |
| 2024 | Nebula: An Edge-Cloud Collaborative Learning Framework for Dynamic Edge EnvironmentsabstractTo bring the great power of modern DNNs into mobile computing and distributed systems, current practices primarily employ one of the two learning paradigms: cloud-based learning or on-device learning. Despite their distinct advantages, neither of these two paradigms could effectively deal with highly dynamic edge environments reflected in quick data distribution shifts and on-device resource fluctuations. In this work, we propose Nebula, an edge-cloud collaborative learning framework to enable rapid model adaptation for changing edge environments. To achieve this, we first propose a new block-level model decomposition scheme to decompose the large cloud model into multiple combinable modules. With this design, we can agilely derive personalized sub-models with compact sizes for edge devices, and quickly aggregate the updated sub-models to integrate new knowledge learned on the edge into the cloud model. We further propose an end-to-end learning framework that incorporates the modular model design into an efficient model adaptation pipeline, including an offline on-cloud model prototyping and training stage, and an online edge-cloud collaborative adaptation stage. Extensive experiments demonstrate that Nebula improves model performance (e.g., 18.89% accuracy increase) and resource efficiency (e.g., 7.12 × communication cost reduction) in adapting models to dynamic edge environments. Yan Zhuang 0004, Zhenzhe Zheng 0001, Yunfeng Shao 0001, Bingshuai Li, Fan Wu 0006, Guihai Chen |
ICPP | 4 |
| 2024 | MAP: Model Aggregation and Personalization in Federated Learning With Incomplete ClassesabstractIn some real-world applications, data samples are usually distributed on local devices, where federated learning (FL) techniques are proposed to coordinate decentralized clients without directly sharing users’ private data. FL commonly follows the parameter server architecture and contains multiple personalization and aggregation procedures. The natural data heterogeneity across clients, i.e., Non-I.I.D. data, challenges both the aggregation and personalization goals in FL. In this paper, we focus on a special kind of Non-I.I.D. scene where clients own incomplete classes, i.e., each client can only access a partial set of the whole class set. The server aims to aggregate a complete classification model that could generalize to all classes, while the clients are inclined to improve the performance of distinguishing their observed classes. For better model aggregation, we point out that the standard softmax will encounter several problems caused by missing classes and propose “restricted softmax” as an alternative. For better model personalization, we point out that the hard-won personalized models are not well exploited and propose “inherited private model” to store the personalization experience. Our proposed algorithm named MAP could simultaneously achieve the aggregation and personalization goals in FL. Abundant experimental studies verify the superiorities of our algorithm. Xin-Chun Li, Shaoming Song, Yinchuan Li, Bingshuai Li, Yunfeng Shao 0001, Yang Yang 0074, De-Chuan Zhan |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | ODE: An Online Data Selection Framework for Federated Learning With Limited StorageabstractMachine learning (ML) models have been deployed in mobile networks to deal with massive data from different layers to enable automated network management. To overcome high communication cost and severe privacy concerns of centralized ML, federated learning (FL) has been proposed to achieve distributed ML among numerous networked devices. While the computation and communication limitation has been widely studied, the impact of limited storage of mobile devices on the performance of FL is still not explored. Without an effective data selection policy to filter the massive streaming networked data on devices, classical FL can suffer from much longer model training time ($4\times$) and dramatic inference accuracy reduction ($7\%$), observed in our experiments. In this work, we take the first step to consider the online data selection for FL with limited on-device storage. We first define a new data valuation metric for data selection in FL with theoretical guarantee for simultaneously accelerating model convergence and enhancing final accuracy. We further design ODE, an Online Data sElection framework for FL, to coordinate networked devices to store valuable data samples collaboratively. Experimental results on one industrial dataset and three public datasets show the remarkable advantages of ODE over the state-of-the-art approaches. Particularly, on the industrial dataset, ODE achieves as high as$2.5\times$speedup of training time and$6\%$increase in final accuracy, and is robust to various factors in practical environments. Chen Gong 0006, Zhenzhe Zheng 0001, Yunfeng Shao 0001, Bingshuai Li, Fan Wu 0006, Guihai Chen |
IEEE/ACM Trans. Netw. | 4 |
| 2023 | Learning From Your Neighbours: Mobility-Driven Device-Edge-Cloud Federated LearningabstractFederated learning (FL) in large-scale wireless networks is implemented in a hierarchical way by introducing edge servers as relays between the cloud server and devices, where devices are dispersed within multiple clusters coordinated by edges. However, the devices are usually mobile users with unpredictable mobile trajectories, whose effects on the model training process are still less studied. In this work, we propose a new MobIlity-Driven feDerated LEarning framework, namely MIDDLE in wireless networks, which can relieve unbalanced and biased model updates by leveraging the new model aggregation opportunities on mobile devices due to their mobility across edges. Specifically, mobile devices can have different models while traversing across edges, and adequately aggregate these models on the device. By theoretical analysis, we can show that this on-device model aggregation can reduce the bias of model updating on edges and cloud, and then accelerate the convergence of model training in FL. Then, we define a model similarity utility to measure the difference in gradient updates among various models, which guides the adaptive on-device model aggregation and in-edge device selection to facilitate the comprehensive information sharing between edges. Extensive experiment results validate that MIDDLE can achieve 1.51 × −6.85 × speedup on the model training, compared with the state-of-the-art model training approaches in hierarchical FL. Songli Zhang, Zhenzhe Zheng 0001, Fan Wu 0006, Bingshuai Li, Yunfeng Shao 0001, Guihai Chen |
ICPP | 4 |
| 2023 | To Store or Not? Online Data Selection for Federated Learning with Limited StorageabstractMachine learning models have been deployed in mobile networks to deal with massive data from different layers to enable automated network management and intelligence on devices. To overcome high communication cost and severe privacy concerns of centralized machine learning, federated learning (FL) has been proposed to achieve distributed machine learning among networked devices. While the computation and communication limitation has been widely studied, the impact of on-device storage on the performance of FL is still not explored. Without an effective data selection policy to filter the massive streaming data on devices, classical FL can suffer from much longer model training time (4 ×) and significant inference accuracy reduction (7%), observed in our experiments. In this work, we take the first step to consider the online data selection for FL with limited on-device storage. We first define a new data valuation metric for data evaluation and selection in FL with theoretical guarantees for speeding up model convergence and enhancing final model accuracy, simultaneously. We further design ODE, a framework of Online Data sElection for FL, to coordinate networked devices to store valuable data samples. Experimental results on one industrial dataset and three public datasets show the remarkable advantages of ODE over the state-of-the-art approaches. Particularly, on the industrial dataset, ODE achieves as high as 2.5 × speedup of training time and 6% increase in inference accuracy, and is robust to various factors in practical environments. Chen Gong 0006, Zhenzhe Zheng 0001, Fan Wu 0006, Yunfeng Shao 0001, Bingshuai Li, Guihai Chen |
WWW | 5 |
| 2023 | Source-free and black-box domain adaptation via distributionally adversarial training
Kunhong Wu, Yahong Han, Yunfeng Shao 0001, Bingshuai Li, Fei Wu 0001 |
Pattern Recognit. | 5 |
| 2023 | Active and Compact Entropy Search for High-Dimensional Bayesian OptimizationabstractEntropy search and its derivative methods are one class of Bayesian Optimization methods that achieve active exploration of black-box functions. They maximize the information gain about the position in the input space where the black-box function gets the global optimum. However, existing entropy search methods suffer from harassment caused by high dimensional optimization problems. On the one hand, the computation for estimating entropies increases exponentially as dimensions increase, which limits the applicability of entropy search to high dimensional problems. On the other hand, many high-dimensional problems have the property that a large number of dimensions have little influence on the objective function, but currently there is no compress mechanism to exclude these redundant dimensions. In this work, we propose Active Compact Entropy Search (AcCES) to fix these defects. Under the guidance of historical evaluation, AcCES actively explores the prevalent inter-dimensional correlations by maximizing the linear or non-linear relationships that may exist between dimensions in the acquisition function, which is ignored by existing Bayesian Optimization methods. In order to build a more compact input space, redundant dimensions are compressed by exploiting inter-dimensional correlations. Experiments demonstrate that AcCES achieves higher query efficiency and optimal results than existing entropy search methods. Run Li, Yahong Han, Yunfeng Shao 0001, Meiyu Qi, Bingshuai Li |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2022 | Federated Learning with Position-Aware NeuronsabstractFederated Learning (FL) fuses collaborative models from local nodes without centralizing users' data. The permutation invariance property of neural networks and the non-i.i.d. data across clients make the locally updated parameters imprecisely aligned, disabling the coordinate-based parameter averaging. Traditional neurons do not explicitly consider position information. Hence, we propose Position-Aware Neurons (PANs) as an alternative, fusing position-related values (i.e., position encodings) into neuron outputs. PANs couple themselves to their positions and minimize the possibility of dislocation, even updating on heterogeneous data. We turn on/off PANs to disable/enable the permutation invariance property of neural networks. PANs are tightly coupled with positions when applied to FL, making parameters across clients pre-aligned and facilitating coordinate-based parameter averaging. PANs are algorithm-agnostic and could universally improve existing FL algorithms. Furthermore, “FL with PANs” is simple to implement and computationally friendly. Xin-Chun Li, Yichu Xu, Shaoming Song, Bingshuai Li, Yinchuan Li, Yunfeng Shao 0001, De-Chuan Zhan |
CVPR | 4 |
| 2022 | Avoid Overfitting User Specific Information in Federated Keyword SpottingabstractKeyword spotting (KWS) aims to discriminate a specific wakeup word from other signals precisely and efficiently for different users.Recent works utilize various deep networks to train KWS models with all users' speech data centralized without considering data privacy.Federated KWS (FedKWS) could serve as a solution without directly sharing users' data.However, the small amount of data, different user habits, and various accents could lead to fatal problems, e.g., overfitting or weight divergence.Hence, we propose several strategies to encourage the model not to overfit user-specific information in FedKWS.Specifically, we first propose an adversarial learning strategy, which updates the downloaded global model against an overfitted local model and explicitly encourages the global model to capture user-invariant information.Furthermore, we propose an adaptive local training strategy, letting clients with more training data and more uniform class distributions undertake more local update steps.Equivalently, this strategy could weaken the negative impacts of those users whose data is less qualified.Our proposed FedKWS-UI could explicitly and implicitly learn user-invariant information in FedKWS.Abundant experimental results on federated Google Speech Commands verify the effectiveness of FedKWS-UI. Xin-Chun Li, Jin-Lin Tang, Shaoming Song, Bingshuai Li, Yinchuan Li, Yunfeng Shao 0001, Le Gan, De-Chuan Zhan |
INTERSPEECH | 4 |
| 2022 | Asymmetric Temperature Scaling Makes Larger Networks Teach Well AgainabstractKnowledge Distillation (KD) aims at transferring the knowledge of a well-performed neural network (the {\it teacher}) to a weaker one (the {\it student}). A peculiar phenomenon is that a more accurate model doesn't necessarily teach better, and temperature adjustment can neither alleviate the mismatched capacity. To explain this, we decompose the efficacy of KD into three parts: {\it correct guidance}, {\it smooth regularization}, and {\it class discriminability}. The last term describes the distinctness of {\it wrong class probabilities} that the teacher provides in KD. Complex teachers tend to be over-confident and traditional temperature scaling limits the efficacy of {\it class discriminability}, resulting in less discriminative wrong class probabilities. Therefore, we propose {\it Asymmetric Temperature Scaling (ATS)}, which separately applies a higher/lower temperature to the correct/wrong class. ATS enlarges the variance of wrong class probabilities in the teacher's label and makes the students grasp the absolute affinities of wrong classes to the target class as discriminative as possible. Both theoretical analysis and extensive experimental results demonstrate the effectiveness of ATS. The demo developed in Mindspore is available at \url{https://gitee.com/lxcnju/ats-mindspore} and will be available at \url{https://gitee.com/mindspore/models/tree/master/research/cv/ats}. Xin-Chun Li, Wen-Shu Fan, Shaoming Song, Yinchuan Li, Bingshuai Li, Yunfeng Shao 0001, De-Chuan Zhan |
NeurIPS | 5 |
| 2022 | Exploring uncertainty in regression neural networks for construction of prediction intervals
Yuandu Lai, Yahong Han, Yunfeng Shao 0001, Meiyu Qi, Bingshuai Li |
Neurocomputing | 6 |
| 2021 | FedPHP: Federated Personalization with Inherited Private Models
Xin-Chun Li, De-Chuan Zhan, Yunfeng Shao 0001, Bingshuai Li, Shaoming Song |
ECML/PKDD (1) | 4 |
| 2020 | Bidirectional Adversarial Training for Semi-Supervised Domain AdaptationabstractSemi-supervised domain adaptation (SSDA) is a novel branch of machine learning that scarce labeled target examples are available, compared with unsupervised domain adaptation. To make effective use of these additional data so as to bridge the domain gap, one possible way is to generate adversarial examples, which are images with additional perturbations, between the two domains and fill the domain gap. Adversarial training has been proven to be a powerful method for this purpose. However, the traditional adversarial training adds noises in arbitrary directions, which is inefficient to migrate between domains, or generate directional noises from the source to target domain and reverse. In this work, we devise a general bidirectional adversarial training method and employ gradient to guide adversarial examples across the domain gap, i.e., the Adaptive Adversarial Training (AAT) for source to target domain and Entropy-penalized Virtual Adversarial Training (E-VAT) for target to source domain. Particularly, we devise a Bidirectional Adversarial Training (BiAT) network to perform diverse adversarial trainings jointly. We evaluate the effectiveness of BiAT on three benchmark datasets and experimental results demonstrate the proposed method achieves the state-of-the-art. Pin Jiang, Aming Wu, Yahong Han, Yunfeng Shao 0001, Meiyu Qi, Bingshuai Li |
IJCAI | 6 |