Shiyao Ma

dblp:183/1810 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 12 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Clustering-Based User Selection in Federated Learning: Metadata Exploitation for 3GPP Networks
Shiyao Ma, Ke Zhang 0008, Chen Sun 0006, Wenqi Zhang 0002
WCNC2
2025 Privacy-Preserving Multilayer Community Detection via Federated Learning
abstract
Existing frameworks of privacy-preserving multilayer community detection have room for improving detection performance and reducing communication overhead. To address these issues, we propose a novel privacy-preserving multilayer community detection framework based on federated learning which is called federated multilayer community detection (FMCD). First, we propose a novel aggregation strategy by utilizing the network average degree of local networks to aggregate the parameters uploaded by clients in the step of aggregation, which can improve the performance of community detection. Second, we design a training procedure to complete multilayer community detection in multiorganizations, which can reduce communication overhead by transmitting merged community information instead of the global parameter. Finally, experiment results on synthetic and real networks with different criteria illustrate that FMCD can achieve significant performance gains, compared with state-of-the-art algorithms.
Shiyao Ma
IEEE Trans. Comput. Soc. Syst.1
2024 Federated Learning with CSMA Based User Selection for IoT Applications
abstract
User selection has became crucial for improving energy efficiency in communication of federated learning (FL) over wireless networks. However, centralized user selection causes additional system complexity. This study proposes a network intrinsic approach of distributed user selection that leverages the radio resource competition mechanism in random access. Taking the carrier sensing multiple access (CSMA) mechanism as an example of random access, we manipulate the contention window (CW) size to prioritize certain users for obtaining radio resources in each round of training. Training data bias is used as a target scenario for FL with user selection. Prioritization is based on the distance between the newly trained local model and the global model of the previous round. To avoid “excessive contribution” by certain users, a counting mechanism is used to ensure fairness. Simulations with various datasets demonstrate that the proposed method can rapidly achieve convergence similar to that of the centralized user selection approach.
Chen Sun 0006, Shiyao Ma, Songtao Wu, Qiang Tong 0002, Wenqi Zhang 0002
ICC2
2023 A General Solution for Straggler Effect and Unreliable Communication in Federated Learning
abstract
The straggler effect is the main bottleneck for Federated Learning (FL), where the performance of training is degraded by the slowest member. Another significant problem is unreliable communication, which somehow has been neglected in previous studies. That is, the transmission of local models is not successful every time. In this paper, we find that the problems of straggler effect and unreliable communication are implicitly caused by time divergence of User Equipments (UEs) in each training round. Based on this, we propose our solutions for these two problems and show that our solutions can be merged into a general one: the problem of the straggler effect and unreliable communication can be solved with a simple UE selection method. This method consists of two steps: First, we cluster UEs into several groups based on UEs' physical parameters or performance metrics; Second, in each training round, only UEs from the same group are chosen for FL operation. Full explanations are given why the time divergence is statistically reduced, and therefore it can mitigate the aforementioned two problems. Our solutions are further illustrated with some examples and validated by simulations.
Tianming Zang, Shiyao Ma, Chen Sun 0006, Wei Chen 0002
ICC3
2023 Robust Semisupervised Federated Learning for Images Automatic Recognition in Internet of Drones
abstract
Air access networks have been recognized as a significant driver of various Internet of Things (IoT) services and applications. In particular, the aerial computing network infrastructure centered on the Internet of Drones has set off a new revolution in automatic image recognition. This emerging technology relies on sharing ground-truth-labeled data between unmanned aerial vehicle (UAV) swarms to train a high-quality automatic image recognition model. However, such an approach will bring data privacy and data availability challenges. To address these issues, we first present a semisupervised federated learning (SSFL) framework for privacy-preserving UAV image recognition. Specifically, we propose a model parameter mixing strategy to improve the naive combination of federated learning and semisupervised learning methods under two realistic scenarios (labels-at-client and labels-at-server), which is referred to as federated mixing (FedMix). Furthermore, there are significant differences in the number, features, and distribution of local data collected by UAVs using different camera modules in different environments, i.e., statistical heterogeneity. To alleviate the statistical heterogeneity problem, we propose an aggregation rule based on the frequency of the client’s participation in training, namely, the FedFreq aggregation rule, which can adjust the weight of the corresponding local model according to its frequency. Numerical results demonstrate that the performance of our proposed method is significantly better than those of the current baseline and is robust to different non-independent and identically distributed(IID) levels of client data.
Zhe Zhang 0043, Shiyao Ma, Zhaohui Yang 0001, Zehui Xiong, Jiawen Kang 0001, Yi Wu 0021, Kejia Zhang 0002, Dusit Niyato
IEEE Internet Things J.2
2023 Personalized Saliency in Task-Oriented Semantic Communications: Image Transmission and Performance Analysis
abstract
Semantic communication, as a promising technology, has emerged to break through the Shannon limit, which is envisioned as the key enabler and fundamental paradigm for future 6G networks and applications, e.g., smart healthcare. In this paper, we focus on UAV image-sensing-driven task-oriented semantic communications scenarios. The majority of existing work has focused on designing advanced algorithms for high-performance semantic communication. However, the challenges, such as energy-hungry and efficiency-limited image retrieval manner, and semantic encoding without considering user personality, have not been explored yet. These challenges have hindered the widespread adoption of semantic communication. To address the above challenges, at the semantic level, we first design an energy-efficient task-oriented semantic communication framework with a triple-based scene graph for image information. We then design a new personalized semantic encoder based on user interests to meet the requirements of personalized saliency. Moreover, at the communication level, we study the effects of dynamic wireless fading channel on semantic transmission mathematically and thus design an optimal multi-user resource allocation scheme by using game theory. Numerical results based on real-world datasets clearly indicate that the proposed framework and schemes significantly enhance the personalization and anti-interference performance of semantic communication, and are also efficient to improve the communication quality of semantic communication services.
Jiawen Kang 0001, Hongyang Du 0001, Zonghang Li, Zehui Xiong, Shiyao Ma, Dusit Niyato
IEEE J. Sel. Areas Commun.5
2022 Privacy-Preserving Anomaly Detection in Cloud Manufacturing Via Federated Transformer
abstract
With the rapid development of cloud manufacturing, industrial production with edge computing as the core architecture has been greatly developed. However, edge devices often suffer from abnormalities and failures in industrial production. Therefore, detecting these abnormal situations timely and accurately is crucial for cloud manufacturing. As such, a straightforward solution is that the edge device uploads the data to the cloud for anomaly detection. However, Industry 4.0 puts forward higher requirements for data privacy and security so that it is unrealistic to upload data from edge devices directly to the cloud. Considering the abovementioned severe challenges, this article customizes a weakly supervised edge computing anomaly detection framework, i.e., federated learning-based transformer framework (FedAnomaly), to deal with the anomaly detection problem in cloud manufacturing. Specifically, we introduce federated learning (FL) framework that allows edge devices to train an anomaly detection model in collaboration with the cloud without compromising privacy. To boost the privacy performance of the framework, we add differential privacy noise to the uploaded features. To further improve the ability of edge devices to extract abnormal features, we use the transformer to extract the feature representation of abnormal data. In this context, we design a novel collaborative learning protocol to promote efficient collaboration between FL and transformer. Furthermore, extensive case studies on four benchmark datasets verify the effectiveness of the proposed framework. To the best of our knowledge, this is the first time integrating FL and transformer to deal with anomaly detection problems in cloud manufacturing.
Shiyao Ma, Jiangtian Nie, Jiawen Kang 0001, Lingjuan Lyu, Ryan Wen Liu, Ruihui Zhao, Ziyao Liu, Dusit Niyato
IEEE Trans. Ind. Informatics1
2019 Adia: Achieving High Link Utilization with Coflow-Aware Scheduling in Data Center Networks
abstract
Link utilization has received extensive attention since data centers become the most pervasive platform for data-parallel applications. A specific job of such applications involves communication among multiple machines. The recently proposed coflow abstraction depicts such communication through a group of parallel flows, and captures application performance through corresponding communication requirements. Existing techniques to improve link utilization, however, either restrict themselves to achieving work conservation, or merely focus on flow-level metrics and ignore coflow-level performance. In this paper, we address the coflow-aware scheduling problem with the objective of maximizing link utilization. Through theoretic analyses, we formulate the coflow-aware scheduling problem as a NP-hard open shop scheduling problem with heterogeneous concurrency. We design Adia, a hierarchical scheduling framework to conduct both inter- and intra- link scheduling. The design of Adia leverages priority-based scheduling while guarantees work-conserving and starvation-free bandwidth allocation at the same time. We also prove Adia's algorithm is two-approximate in terms of link utilization. Extensive simulation results on ns3 further show that Adia outperforms both per-flow mechanisms coflow schemes in terms of link utilization, and achieves similar coflow performance in comparison with the state-of-art coflow scheduling schemes.
Jingjie Jiang, Shiyao Ma, Bo Li 0001, Baochun Li
IEEE Trans. Cloud Comput.2
2018 Unraveling the RTT-fairness Problem for BBR: A Queueing Model
abstract
BBR is a congestion-based congestion control algorithm recently proposed by Google. It proactively measures the bottleneck bandwidth and round trip times (RTTs) of a connection pipe, based on which it governs its sending behaviors. Despite the significant throughput gains and latency reduction, some experimental studies reveal that BBR may result in a salient RTT-fairness problem, in that short-RTT flows can be starved of bandwidth allocation when comnetina with lons-R'I'T flows. In this paper, we study BBR's RTT-fairness problem from a theoretic perspective. We present a closed-form solution that characterizes the intrinsic dynamics of BBR flows and their interactions. Specifically, we model BBR's sending behaviors and bandwidth dynamics, based on which we establish an exponential relationship between the flows' bandwidth shares and their RTTs. We show that the degree of unfairness is dictated by the RTT ratio between two flows, irrespective of the other network parameters, such as the initial sending rates or link capacity. In particular, when the RTT ratio of the two flows is greater than 2, the short-RTT flow is starved of bandwidth allocation ( ≤ 0.1%), Our theoretical results are corroborated by simulations in a wide range of settings.
Yuechen Tao, Jingjie Jiang, Shiyao Ma, Wei Wang 0030, Bo Li 0001
GLOBECOM3
2017 Maximizing link utilization with coflow-aware scheduling in datacenter networks
abstract
Link utilization has received extensive attention since datacenters become the most prevalent platform for data-parallel computing applications. A specific job of such applications involves communication among multiple machines. The coflow abstraction depicts such communication and captures application performance through corresponding network requirements. Existing techniques to improve link utilization, however, either restrict themselves to work conservation, or merely focus on flow-level metrics and ignore coflow-level performance. In this paper, we address the coflow-aware scheduling problem with the objective of maximizing link utilization. Through theoretic analyses, we formulate the coflow-aware scheduling problem as a NP-hard open shop scheduling problem with heterogeneous concurrency. Despite the hardness of this problem, we design Maluca, a hierarchical scheduling framework to conduct both inter- and intra-link scheduling. Maluca's algorithm is not only starvation-free and work-conserving, but also 2-approximate in terms of link utilization. Extensive simulation results demonstrate that Maluca outperforms both per-flow and coflow schemes in terms of link utilization, and achieves similar coflow performance in comparison with the state-of-art coflow scheduling schemes.
Jingjie Jiang, Shiyao Ma, Bo Li 0001, Baochun Li, Jiangchuan Liu
ICC2
2017 Coflex: Navigating the fairness-efficiency tradeoff for coflow scheduling
abstract
Fair and efficient coflow scheduling improves application-level networking performance in today's datacenters. Ideally, a coflow scheduler should provide isolation guarantees on the minimum coflow progress to achieve predictable networking performance. Network operators, on the other hand, strive to decrease the average coflow completion time (CCT). Unfortunately, optimal isolation guarantees and minimum average CCT are conflicting objectives and cannot be achieved at the same time. Existing coflow schedulers either optimize isolation guarantees at the expense of long CCTs (e.g., HUG [1]), or decrease the average CCT without performance isolation (e.g., Varys and Aalo [2], [3]). The lack of a smooth tradeoff in between poses a dilemma between low efficiency and no performance isolation. To bridge this gap, we develop a new coflow scheduler, Coflex, to navigate this tradeoff. Coflex allows network operators to specify the desired level of isolation guarantee using a tunable fairness knob, while at the same time decreasing the average CCT. Both our real-world deployments and trace-driven simulations have shown that Coflex offers a smooth tradeoff between fairness and efficiency. At an appropriate tradeoff level, Coflex outperforms fair schedulers by 2 × in minimizing the average CCT.
Wei Wang 0030, Shiyao Ma, Bo Li 0001, Baochun Li
INFOCOM2
2016 Custody: Towards Data-Aware Resource Sharing in Cloud-Based Big Data Processing
abstract
With the advent of big data processing frameworks, the performance of data-parallel applications is heavily affected by the time it takes to read input data, making it important to improve data locality. Existing methods in achieving data locality have primarily focused on selecting machines to place tasks of applications. Nevertheless, the set of machines that an application can choose from is determined by a cluster manager, which is oblivious to the location of data in existing resource sharing frameworks. In this paper, we design, implement and evaluate Custody, a new cluster management framework that helps to maximize data locality by allocating the executor processes with local access to data to those applications in need. Custody achieves this objective by dynamically collecting runtime information of an application's input data and by effectively allocating executors among and within applications through theoretic analyses of the data-aware resource sharing problem. With significantly better data locality, Custody avoids unnecessary network transfers and thus expedites job completion times. Our experimental results on a 100-node cluster demonstrate that Custody can improve the data locality for input tasks by 36.9% in comparison with Spark's default cluster manager. Meanwhile, it reduces the job completion times by 14.9% due to fewer network transfers.
Shiyao Ma, Jingjie Jiang, Bo Li 0001, Baochun Li
CLUSTER1
2016 Chronos: Meeting coflow deadlines in data center networks
abstract
Guaranteed performance for data-parallel applications is important for both service providers and cloud data centers that host such services. A job of data-parallel applications involves communication among multiple machines to transmit intermediate results. Such communication comprises a collection of parallel flows, which is abstracted as a coflow in recent proposals. In this paper, we study the problem of meeting deadlines for coflows in data center networks. Existing flow-level scheduling schemes are insufficient to guarantee the coflow-level performance, since a coflow can meet its deadline only when all its constituent flows finish on time. Due to the scarce bandwidth on the network bottleneck, it is vital to coordinate concurrent coflows to meet as many deadlines as possible. We present Chronos, a scheduling framework that captures the correlation of flows belonging to a coflow, and handles the resource allocation among multiple concurrent coflows. Chronos is work-conserving and starvation-free without integrating complicated admission control mechanisms. We show via extensive simulations on ns3 that Chronos can make 1.6× more coflows meet their deadlines compared to flow-level schemes.
Shiyao Ma, Jingjie Jiang, Bo Li 0001, Baochun Li
ICC1
2016 Tailor: Trimming Coflow Completion Times in Datacenter Networks
abstract
Tasks in a data-parallel job communicate with each other through a number of concurrent flows, which is described as a coflow. These flows are correlated in the sense that the performance of a coflow is dictated by the flow that takes the longest time to complete. Minimizing coflow completion times, however, turns out to be a challenge, given the correlation across flows and how they are routed collectively through a datacenter network. In this paper, we propose Tailor, a simple yet effective mechanism with the objective of trimming the coflow completion times in a datacenter network. To achieve our objective, Tailor takes advantage of OpenFlow in a software-defined datacenter network. By monitoring and rerouting live flows to links with lighter loads, Tailor guarantees that the coflow completion time is minimized dynamically and converges to its lower bound. Our experimental results in both Mininet and large-scale simulations have shown that Tailor is much more effective than flow-level schemes when it comes to reducing coflow completion times. It also outperforms existing scheduling-only coflow mechanisms and achieves similar performance with the state-of-the-art hybrid mechanism, yet with much lower complexity.
Jingjie Jiang, Shiyao Ma, Bo Li 0001, Baochun Li
ICCCN2
2016 Maximizing container-based network isolation in parallel computing clusters
abstract
Data-parallel applications, especially those associated with user-facing web services, have struggled to enhance their worst case performance. It is therefore important to improve the minimum amount of resources guaranteed for applications in a cluster. Existing cluster management frameworks, however, provide isolation for computation resources (such as CPU) only, and are oblivious to network isolation guarantees. In this paper, we design, implement and evaluate Libra, a new cluster management framework that helps to maximize the isolation guarantee for the bandwidth requirements from applications. We start with a theoretical analysis of the network sharing problem, which contains two key steps: container placement and bandwidth allocation. By collecting the status of access links and the bandwidth demand of applications, we coordinate the placement of containers to minimize the system bottleneck such that the bandwidth guarantee for applications can be optimized. We further embrace host-based rate limiting to ensure such maximized bandwidth guarantee can be reached without hurting network utilization. Both our testbed-based experiments and large-scale simulations demonstrate that Libra significantly improves the network isolation guarantee: in comparison with existing cluster managers and network schedulers, the performance gain is more than 105.59%. Meanwhile, it improves application performance by 57.71% and maintains high network utilization.
Shiyao Ma, Jingjie Jiang, Bo Li 0001, Baochun Li
ICNP1
2016 Symbiosis: Network-aware task scheduling in data-parallel frameworks
abstract
Even with the recent proliferation of in-memory computation in data-parallel frameworks (such as Spark), transfers over the network are still time-consuming. Similar to computation, network transfers serve as main roadblocks as we try to minimize job completion times. Existing schedulers were designed as isolated solutions that focused on computation or network performance only. Without any coordination, the utilization of computation and network resources may become unbalanced, leading to a reduced level of overall resource utilization. In this paper, we design, implement, and evaluate Symbiosis, a network-aware task scheduler designed to coordinate computation-bound and network-bound tasks in a large cluster, so that resources are utilized in a more balanced fashion. Symbiosis is an online scheduler that predicts resource imbalance before launching tasks, and correct such imbalance by co-locating computation-bound and network-bound tasks in the same executor process. As a guiding principle, it is engineered to be practically implemented within and to complement existing data-parallel frameworks. We have implemented Symbiosis within Spark, and carried out our experiments on a 100-node cluster. We show convincing evidence that Symbiosis reduces job completion times by 11.9% in comparison to Spark's current scheduler with little overhead.
Jingjie Jiang, Shiyao Ma, Bo Li 0001, Baochun Li
INFOCOM2