Danyang Xiao

dblp:185/7116 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0001-6798-9683ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Robust prediction of massive short cloud workloads using online meta learning
Xuan Mo, Jialun Li, Shunjue Chen, Danyang Xiao, Weigang Wu
Inf. Sci.4
2026 Decentralized Federated Distillation With Protected Pruned Models in Few Global Epochs
abstract
Decentralized Federated Distillation (DFD) has emerged as a significant research direction since it not only supports heterogeneous model training but also naturally avoids privacy risks and communication bottlenecks stemming from the central server. DFD shows great potential in various training scenarios, especially in cross-silo federated environments. Existing DFD algorithms have limitations, including reliance on public datasets for distillation, insufficient privacy protection, and high communication overhead. This paper introduces Ring-Distill, a novel DFD algorithm designed specifically for cross-silo federated environments, effectively addressing the aforementioned limitations. Ring-Distill allows clients to complete federated training within a few global epochs (i.e., communication rounds) without sharing public datasets, e.g.,$N$global epochs for a system involving$N$clients. To protect privacy and further reduce communication overhead, Ring-Distill contains a privacy-oriented automatic model pruning mechanism (PAMP) which can automatically create the privacy-preserving compressed model called proxy model for each client. These proxy models are employed for client-to-client distillation during each global epoch. Furthermore, Ring-Distill contains a historical model-based distillation (HMD) mechanism that allows local models to transfer knowledge from multiple historical proxy model replicas stored locally. Due to historical proxy model replicas, the HMD mechanism not only improves the distillation performance but also effectively mitigates client dropout issues. Theoretical analysis and comprehensive experiments show that Ring-Distill has significant advantages in terms of accuracy, privacy, and communication cost.
Danyang Xiao, Jialun Li, Xuan Mo, Weigang Wu, Jiannong Cao 0001
IEEE Trans. Dependable Secur. Comput.1
2025 Integrated and Fungible Scheduling of Deep Learning Workloads Using Multi-Agent Reinforcement Learning
abstract
GPU clusters have been widely used to co-locate various deep learning (DL) workloads in a multi-tenant way. Although such resource sharing can significantly reduce training cost, resource contention and interference among co-located workloads make task scheduling very complex and challenging. To simplify the scheduling problem, existing algorithms usually divide the procedure of scheduling into two sub-tasks, i.e., task placement and resource allocation, and allocate resources according to pre-defined and fixed resource demands. However, such a paradigm significantly constrains the selection of potential scheduling solutions. In this article, we present MAIFS, a novel multi-agent reinforcement learning based scheduling algorithm that handles task placement and resource allocation integratedly, and allows fungible resource allocation based on resource sensitivity of DL workloads. The core of MAIFS lies in two mechanisms. The multi-agent attention mechanism is designed to learn and share inter-related resource state features observed from different agents, which enables agents to explore fungible resource allocation solutions. The dynamic coordination graph mechanism is designed for coordinating interactive task placement decisions of agents during integrated scheduling, so as to mitigate potential task conflicts. Simulated experiments using two large scale production DL workload traces and physical deployment experiments based on a Kubernetes based GPU cluster show that MAIFS can outperform state-of-the-art scheduling algorithms by up to 44% in terms of makespan and 46% in terms of job completion time (JCT).
Jialun Li, Danyang Xiao, Diying Yang, Xuan Mo, Weigang Wu
IEEE Trans. Parallel Distributed Syst.2
2024 Privacy Leakage from Logits Attack and its Defense in Federated Distillation
abstract
Federated Distillation (FD), a popular variant of Federated Learning (FL), has attracted researchers' attention due to its ability to support heterogeneous model training. Generally, FD allows clients to upload logits associated with public datasets for knowledge transfer, yet logits may pose privacy risks. In this study, we provide the first demonstration of the impact of privacy risks caused by logits. Specifically, we design a data reconstruction attack against logits named L-Attack which can reveal sensitive information about the target client without access to the target model. Via the zeroth-order optimization technique, L-Attack involves training a server-side generator that unveils certain features of private data owned by the target client. To defend against L-Attack, we propose a label aggregation-based FD algorithm called LabelAvg which allows clients to upload predicted hard labels for knowledge transfer instead of logits. Due to the insufficient information in labels for distillation, LabelAvg provides a voting-based label smoothing mechanism that enables the server to construct smooth labels from received labels. The generated smooth labels which stand for the consensus among all clients, indicate the approximate probability distribution. Thus, these smoothed labels bear a striking similarity to logits and can be used for distillation. Analysis and experimental results prove LabelAvg is superior to baselines in terms of accuracy, privacy, and communication data volume.
Danyang Xiao, Diying Yang, Jialun Li, Xu Chen 0004, Weigang Wu
DSN1
2024 FedLOC: A Layer Output Based Compression Algorithm for Federated Learning
abstract
Communication cost is a main challenge in Federated Learning (FL). Gradient sparsification is one of the effective ways to reduce communication data volumes by allowing clients to send only a small portion of gradient elements to the server. Existing gradient sparsification methods, such as Top-K, analyze and decide the contribution of gradient elements for uploading by comparing their values. However, gradient elements with large values may not necessarily contain more information. To improve the accuracy while effectively reducing communication costs, in this paper, we propose a novel gradient sparsification-based FL algorithm called FedLOC, which sparsifies gradients not only based on the value of the gradient elements but also on the layer output of the local model. In particular, we design a two-phase compression operation to reduce the communication data volumes. First, FedLOC sparsifies gradients based on their values to construct initial compressed gradients. Subsequently, since the performance of the local model is related to the output of each layer, FedLOC analyzes the contribution of initial compressed gradient elements and constructs masks for the second sparsification operation based on each layer’s extend output derived from the original layer output. Convergence analysis and experimental results show that FedLOC can achieve superior performance in terms of accuracy and convergence when effectively reducing data communication volumes.
Danyang Xiao, Jieying Zhou, Weigang Wu
HPCC2
2024 EvoGWP: Predicting Long-Term Changes in Cloud Workloads Using Deep Graph-Evolution Learning
abstract
Workload prediction plays a crucial role in resource management of large scale cloud datacenters. Although quite a number of methods/algorithms have been proposed, long-term changes have not been explicitly identified and considered. Due to shifty user demands, workload re-locations, or other reasons, the “resource usage pattern” of a workload, which is usually quite stable in a short-term view, may change dynamically in a long-term range. Such long-term dynamic changes may cause significant accuracy degradation for prediction algorithms. How to handle such long-term dynamic changes is an open and challenging issue. In this article, we propose Evolution Graph for Workload Prediction (EvoGWP), a novel method that can predict long-term dynamic changes using a delicately designed graph-based evolution learning algorithm. EvoGWP automatically extracts shapelets to explicitly identify resource usage patterns of workloads in a fine-grained level, and predicts workload changes by considering factors in both temporal and spatial dimensions. We design a two-level importance based shapelet extraction mechanism to mine new usage pattern changes in temporal dimension, and design a novel evolution graph model to fuse the interference among resource usage patterns of different workloads in spatial dimension. By combining temporal extraction of shapelets from each single workload and spatial interference of shapelets among different workloads, we then design a spatio-temporal GNN-based encoder-decoder model to predict the long-term dynamic changes of workloads. Experiments using real trace data from Alibaba, Tencent and Google show that EvoGWP improves the prediction accuracy by up to 58.6% over the state-of-the-art prediction methods. Moreover, EvoGWP can outperform the state-of-the-art prediction methods in terms of model convergence. To the best of our knowledge, this is the first work that explicitly identifies fine-grained workload resource usage patterns to accurately predict long-term dynamic changes of workloads.
Jialun Li, Jieqian Yao, Danyang Xiao, Diying Yang, Weigang Wu
IEEE Trans. Parallel Distributed Syst.3
2023 Bi-level Sampling: Improved Clients Selection in Heterogeneous Settings for Federated Learning
abstract
Client selection (a.k.a., client sampling) is one of the hot topics in Federated Learning (FL). In each communication round, selecting some clients to participate in aggregation can effectively reduce the communication overhead caused by exchanging model parameters. However, due to statistical heterogeneity in FL, selecting clients randomly may affect the performance of aggregated global models. existing approaches regarding client selection firstly cluster clients and then sample(select) some representative clients from each cluster. However, these clustering-based approaches may be either time-intensive or high complexity. To address these issues, In this paper, we introduce Bi-level Sampling, a clustering-based approach for client selection. After multinomial distribution sampling, Bi-level Sampling clusters clients based on weighted per-label mean class scores and then selects participating clients for federated learning in each round. Bi-level Sampling can lead to better client representativity and the reduced variance of the client’s stochastic aggregation weights in FL. Our approach can be integrated into typical FL frameworks. Experimental results show that, compared with state-of-the-art approaches, our approach demonstrates significantly more stable and accurate convergence behavior-getting higher test accuracy and less training time, especially in highly Non-IID settings.
Danyang Xiao, Congcong Zhan, Jialun Li, Weigang Wu
IPCCC1
2023 FedCME: Client Matching and Classifier Exchanging to Handle Data Heterogeneity in Federated Learning
abstract
Data heterogeneity across clients is one of the key challenges in Federated Learning (FL), which may slow down the global model convergence and even weaken global model performance. Most existing approaches tackle the heterogeneity by constraining local model updates through reference to global information provided by the server. This can alleviate the performance degradation on the aggregated global model. Different from existing methods, we focus the information exchange between clients, which could also enhance the effectiveness of local training and lead to generate a high-performance global model. Concretely, we propose a novel FL framework named FedCME by client matching and classifier exchanging. In FedCME, clients with large differences in data distribution will be matched in pairs, and then the corresponding pair of clients will exchange their classifiers at the stage of local training in an intermediate moment. Since the local data determines the local model training direction, our method can correct update direction of classifiers and effectively alleviate local update divergence. Besides, we propose feature alignment to enhance the training of the feature extractor. Experimental results demonstrate that FedCME performs better than FedAvg, FedProx, MOON and FedRS on popular federated learning benchmarks including FMNIST and CIFAR10, in the case where data are heterogeneous.
Jun Nie, Danyang Xiao, Lei Yang 0024, Weigang Wu
MSN2
2023 Learning Scheduling Policies for Co-Located Workloads in Cloud Datacenters
abstract
Co-location, which deploys long running applications and batch-processing applications in the same computing cluster, has become a promising way to improve resource utility for large cloud datacenters. However, co-location brings huge challenges to task scheduling because different types of workloads may affect each other. Existing works on task scheduling rarely focus on the scenario of co-location. This article presents Co-ScheRRL, a scheduling algorithm delicately designed for co-located workloads. Co-ScheRRL consists of two major mechanisms: i) a self-attention encoding mechanism which encodes and represents states of the computing cluster as a set of embedding feature vectors; ii) a deep reinforcement learning (DRL) relational reasoning mechanism which calculates and compares different scheduling actions under different co-located workloads pattern via DRL feedback reward signals based on these feature vectors. Our two mechanisms can tackle complicatedly and dynamically varying behaviors of co-located workloads. With the help of these two mechanisms, Co-ScheRRL is able to construct high-quality scheduling policies. Trace-driven simulation demonstrates that Co-ScheRRL outperforms existing scheduling algorithms in terms of makespan by more than 38.4% and throughput by more than 166.7%.
Jialun Li, Danyang Xiao, Jieqian Yao, Yujie Long, Weigang Wu
IEEE Trans. Cloud Comput.2
2022 EFL-WP: Federated Learning-Based Workload Prediction in Inter-Cloud Environments
abstract
Resource allocation has been always a major concern of cloud providers. Workload prediction can effectively improve resource utility by providing information about resources available in the future. Machine learning-based workload prediction has been widely studied and deployed in large-scale clouds. However, workload prediction in Inter-Cloud environments has not been considered. Since more and more cloud services are orchestrated by containers (or virtual machines) distributed in multiple clouds, intelligent models for workload prediction should be trained based on trace data across clouds. Due to concerns raised by data privacy, the trace data of clouds may not be shared with each other, and general distributed training is not applicable. In this paper, we propose a general and flexible framework, namely EFL-WP, for training workload prediction models in the Inter-Cloud environment. EFL-WP adopts the federated learning approach and allows cloud providers to collaborate to train prediction models without sharing traces. EFL-WP considers the difference among workloads in different cloud (i.e., the Non-IID characteristic of traces) by using two novel techniques: participant selection mechanism and multi-global models aggregation. The former can prevent some local models trained on non-IID traces from participating in global aggregation, which is beneficial to the global model. The latter allows the coordinator to aggregate several global models according to the difference among local models. To further improve accuracy, EFL-WP adopts an ensemble inference strategy. Experimental results show that the proposed framework can be superior to baselines on both Alibaba dataset and Tencent Games Traces.
Danyang Xiao, Bokai Cao, Weigang Wu
IJCNN1
2022 RTGA: Robust ternary gradients aggregation for federated learning
Chengang Yang, Danyang Xiao, Bokai Cao, Weigang Wu
Inf. Sci.2
2022 Iteration number-based hierarchical gradient aggregation for distributed deep learning
Danyang Xiao, Jieying Zhou, Yunfei Du 0001, Weigang Wu
J. Supercomput.1
2022 Mixing Activations and Labels in Distributed Training for Split Learning
abstract
Split Learning (SL) is a distributed machine learning setting that allows several nodes to train neural networks based on model parallelism. Since SL avoids sharing raw data among training nodes, it can protect data privacy by nature. However, recent studies show that, raw data may be reconstructed from activations in training, which may cause data privacy leakage. Besides raw data, label sharing in SL may also cause privacy problems. In order to address these issues, we propose a novel mechanism called multiple activations and labels mix (MALM). By taking advantage of the diversity of sample categories, MALM generates mixed activations that preserve a low distance correlation with the raw data so as to reduce the risk of reconstruction attacks. To protect label information, MALM creates obfuscated labels associated with the raw data so as to prevent adversaries from inferring ground-truth labels. Since clients with few sample categories may not effectively generate mixed activations and obfuscated labels, we propose a bipartite graph based assistant client match technique for MALM, which lets clients with a large number of categories provide mixed activations and obfuscated labels for clients with few categories. Those clients with few categories can mix the obtained mixed activations and obfuscated labels with their own activations and labels. Experimental results show that, compared with baselines, MALM can reduce the risk of raw data and label information leakage with lower cost, while achieving comparable even better model performance.
Danyang Xiao, Chengang Yang, Weigang Wu
IEEE Trans. Parallel Distributed Syst.1
2021 FedVF: Personalized Federated Learning Based on Layer-wise Parameter Updates with Variable Frequency
abstract
Federated learning is a new and increasingly popular distributed machine learning paradigm, which can employ multiple clients such as mobile phones to collaboratively train a deep neural network model under the coordination of the central server. The raw data is always kept locally on the clients, and only the model parameters are communicated between clients and the server, so that data privacy can be largely preserved. To cope with the effect of not independent and identically distributed (Non-IID) data, personalized federated learning that can provide personalized and customized models for clients participating in federated training has been proposed and widely studied. In this paper, we propose a personalized federated learning algorithm that can provide a personalized local model for each client while storing the latest global model on the central server through only one federated training. The deep neural network model is divided into two parts of global layers and personalized layers, and the whole training process is partitioned into earlier and later stages. Based on the division of layer-wise parameters, the cumulative learning strategy of updating parameters on different layers with variable frequency is adopted to better personalize local models in the later stage based on the global features learned in the earlier stage. Compared with existing personalized federated learning algorithms, our proposed algorithm can achieve a good balance between the personalized model and the global model in federated learning, and performs better in communication efficiency and model accuracy.
Yuan Mei 0005, Binbin Guo, Danyang Xiao, Weigang Wu
IPCCC3
2021 EGC: Entropy-based gradient compression for distributed deep learning
Danyang Xiao, Yuan Mei 0005, Di Kuang, Mengqiang Chen, Binbin Guo, Weigang Wu
Inf. Sci.1
2020 Dual-Way Gradient Sparsification for Asynchronous Distributed Deep Learning
abstract
Distributed parallel training using computing clusters is desirable for large scale deep neural networks. One of the key challenges in distributed training is the communication cost for exchanging information, such as stochastic gradients, among training nodes. Recently, gradient sparsification techniques have been proposed to reduce the amount of data exchanged and thus alleviate the network overhead. However, most existing gradient sparsification approaches consider only synchronous parallelism and cannot be applied in asynchronous distributed training.
Zijie Yan, Danyang Xiao, Mengqiang Chen, Jieying Zhou, Weigang Wu
ICPP2
2018 A Semi-Supervised Network Embedding Model for Protein Complexes Detection
abstract
Protein complex is a group of associated polypeptide chains which plays essential roles in biological process. Given a graph representing protein-protein interactions (PPI) network, it is critical but non-trivial to detect protein complexes.In this paper, we propose a semi-supervised network embedding model by adopting graph convolutional networks to effectively detect densely connected subgraphs. We conduct extensive experiment on two popular PPI networks with various data sizes and densities. The experimental results show our approach achieves state-of-the-art performance.
Wei Zhao 0033, Jia Zhu 0003, Min Yang 0007, Danyang Xiao, Gabriel Pui Cheong Fung, Xiaojun Chen 0006
AAAI4
2018 GA-ADE: a novel approach based on graph algorithm to improves the detection of adverse drug events
Xingcheng Wu, Jia Zhu 0003, Danyang Xiao, Xueqin Lin, Rui Ding 0007
Multim. Tools Appl.3
2016 PCMiner: An Extensible System for Analysing and Detecting Protein Complexes
Danyang Xiao, Jia Zhu 0003, Yong Tang 0001, Lingxiao Chen, Jingmin Wei
APWeb (2)1
2016 A novel feature selection strategy for friends recommendation
abstract
With the social network being widely used, people would like to use the friends recommendation provided by a social websites. There are lots of methods to make the recommendation results more accurately and efficiently. By considering the feature selection strategy in the stage of data preprocessing, we propose a novel friend recommendation system using a classification model, which formulates the recommendation problem. We compare the performance of four classifiers, and draw a conclusion that our proposed method can get higher accuracy.
Rui Ding 0007, Jia Zhu 0003, Yong Tang 0001, Xueqin Lin, Danyang Xiao, Haoye Dong
CSCWD5