EDBT 2026 Demo / reviewers in the wild / expert
Zhihao Qu
dblp:173/0285
· DBLP profile ↗
77ranked-venue papers
5as first author
68since 2021 · last 2026
0000-0001-7538-1985ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 32 · 3 first-author · 27 since 2021Systems, architecture and hardware · 29 · 2 first-author · 25 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Quantization-Aware Incentive Mechanism for Communication-Efficient Federated Learning
Hengrui Cui, Zhihao Qu, Bin Tang 0002 |
ICDCS | 2 |
| 2026 | Incentivizing and Orchestrating Cloud-Edge LLM Speculative Decoding via Auctions
Mingtao Ji, Lei Jiao 0002, Bin Tang 0002, Zhihao Qu |
ICDCS | 4 |
| 2026 | Trilogy: Tag Information Collection in Multi-Category Commodity RFID Systems
Zhihao Qu, Jia Liu 0008, Yingchi Mao, Bin Tang 0002 |
ICDCS | 2 |
| 2026 | MOM-VI: Mobility-aware joint offloading and migration with traffic flow prediction for vehicle-infrastructure collaboration
Shihong Hu, Kaiyue Li, Zhihao Qu, Bin Tang 0002 |
Future Gener. Comput. Syst. | 3 |
| 2026 | Battery Lifetime Extension in Heterogeneous Satellite Edge Computing: A Lyapunov-DRL ApproachabstractLow Earth Orbit (LEO) satellite mobile edge computing (SMEC) has emerged as a pivotal technology for delivering low-delay communication and computational services to under-served regions. However, ensuring sustainable operation of SMEC systems remains challenging due to limited onboard energy and battery degradation, which is critically influenced by the Depth of Discharge (DoD) of battery. This paper investigates the collaborative DoD optimization problem in heterogeneous SMEC, formulating it as a long-term stochastic optimization aimed at minimizing DoD while maintaining system stability under dynamic energy supply and stochastic task arrivals. To address this problem, we propose a Lyapunov-guided optimization framework that integrates Lyapunov optimization with a multi-agent deep reinforcement learning algorithm. Specifically, we employ Lyapunov optimization to transform the long-term objective into a sequence of per-time-slot subproblems. Subsequently, given the high complexity of minimizing the derived Lyapunov drift function directly, we utilize a Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm to solve these subproblems. This approach leverages MADDPG’s proven capability in handling complex sequential decision-making processes while enabling real-time adaptive optimization in dynamic environments. Extensive simulations demonstrate that the proposed framework achieves significant reduction in DoD, bounded delay, and stable queue dynamics in diverse scenarios. Liang Zhong 0002, Shen Tian, Deze Zeng, Zhihao Qu, Chengyu Hu 0002 |
IEEE Internet Things J. | 4 |
| 2026 | Analytic personalized federated meta-learning
Shunxian Gu, Chaoqun You, Deke Guo, Zhihao Qu, Bangbang Ren, Zaipeng Xie, Lailong Luo |
Pattern Recognit. | 4 |
| 2026 | Enhance and reuse: A dual-mechanism approach to boost deep forest for label distribution learning
Jia-Le Xu, Shen-Huan Lyu, Yu-Nian Wang, Zhihao Qu, Bin Tang 0002 |
Pattern Recognit. | 5 |
| 2026 | Compressing model with few class-imbalance samples: An out-of-distribution expedition
Tian-Shuang Wu, Shen-Huan Lyu, Yanyan Wang 0001, Zhihao Qu |
Pattern Recognit. Lett. | 5 |
| 2026 | Reinforcement Learning-Based Equipment Combination Selection Optimization for Multi-Layer Kill WebsabstractModern network-centric operations increasingly rely on multi-layer Kill Webs (KWs), enabling redundant and non-linear sensing-to-strike pathways while introducing a combinatorial equipment selection problem under uncertainty and resource constraints. This paper formulates a multi-layer KW equipment combination selection as a sequential decision-making problem by explicitly modeling heterogeneous equipment capabilities, resource constraints, and the network topology. To address this problem, we developed an RL learning-based optimization framework, where a multi-objective reward function integrates normalized relevance, operational risk, and timeliness, with a penalty mechanism for infeasible or incomplete kill-chain closure. Based on the jointly captured state information (e.g., network structure, equipment attributes, target characteristics, and resource availability), an Actor-Critic (AC) algorithm is developed to learn adaptive equipment combination selection across different operational stages using temporal-difference advantage estimation and entropy regularization. Simulation results under diverse battlefield scenarios demonstrate that the proposed framework consistently outperforms Deep Q-Network (DQN), Proximal Policy Optimization (PPO), and Particle Swarm Optimization (PSO), achieving at least a 19.6% improvement in overall operational effectiveness while maintaining low decision latency. Chao Fang 0001, Waris Ali, Yingshan Li, Zhihao Qu, Deze Zeng |
IEEE Trans. Cloud Comput. | 6 |
| 2026 | Lightweight Adaptive Quantization Algorithms for Federated Learning With Heterogeneous Clients
Hengrui Cui, Zhihao Qu, Bin Tang 0002, Yue Zeng 0002 |
IEEE Trans. Mob. Comput. | 2 |
| 2026 | Time-Efficient Identifying Key Tag Distribution in Large-Scale RFID SystemsabstractWith the proliferation of RFID-enabled applications, large-scale RFID systems often require multiple readers to ensure full coverage of numerous tags. In such systems, we sometimes pay more attention to a subset of tags instead of all, which are called key tags. This paper studies an under-investigated problemkey tag distribution identification, which aims to identify which key tags are beneath which readers. This is crucial for efficiently managing specific items of interest, which can quickly pinpoint key tags and help RFID readers covering these tags collaborate to improve the tag inventory efficiency. We propose a protocol called Kadept that identifies the key tag distribution by designing a sophisticated Cuckoo filter that teases out key tags as well as assigns each of them a singleton slot for response. With this design, a great number of trivial (non-key) tags will keep silent and free up bandwidth resources for key tags, and each key tag is sorted in a collision-free way and can be identified with only 1-bit response, which significantly improves the time efficiency. To enhance the scalability and efficiency of Kadept for high key tag proportions, we propose E-Kadept protocol, which accelerates the identification process by designing an incremental Cuckoo filter that reduces false positives and improves space efficiency. We theoretically analyze how to optimize protocol parameters of Kadept and E-Kadept, and conduct extensive simulations under different tag distribution scenarios. Compared with the state-of-the-art, E-Kadept can improve the time efficiency by a factor of 1.75×, when the ratio of key tags to all tags is 0.3. Yanyan Wang 0001, Jia Liu 0008, Zhihao Qu, Shen-Huan Lyu, Bin Tang 0002 |
IEEE Trans. Mob. Comput. | 3 |
| 2026 | AFedLF: Adaptive Layer Freezing of Foundation Models in Heterogeneous Federated LearningabstractThe rise of pre-trained foundation models (FMs) has popularized the trend of fine-tuning FMs to fit downstream tasks, while Federated Learning (FL) has become the de-facto approach for training distributed data with privacy-preservation. However, fine-tuning FMs in FL faces overwhelming overheads due to its bulky nature. While freezing parameters in FM have the potential to accelerate FL training, existing freezing strategies statically freeze parameters on specified or already converged layers, incur severe accuracy degradation, and resource-inefficiency in heterogeneous environments. In this paper, we propose AFedLF, an adaptive freezing framework for FM in FL, to accelerate its wall-clock time for convergence without losing its final accuracy. However, this poses great challenges, as different freezing strategies lead to different accuracy gains and time overheads, while unfreezing more layers may bring marginal accuracy gains but significant time overheads. To address this challenge, AFedLF mathematically establishes a correlation between the freezing strategy and the accuracy gain and time overhead, and allocates adaptive freezing strategies to clients, based on our insight that unfreezing more layers on devices with strong computation and communication capabilities helps improve resource efficiency. Besides, AFedLF incorporates our well-designed intermediate result caching scheme with constant approximation ratios utilizing the limited storage capacity on mobile devices to cache intermediate results to skip forward propagation, further saving wall-clock time. Finally, we implemented AFedLF using an open-source FL benchmark, and extensive trace-driven experimental results showed that AFedLF accelerates wall-clock time by up to 6.1× compared to state-of-the-art solutions, without sacrificing accuracy. Yue Zeng 0002, Jie Zhang 0076, Song Guo 0001, Zhihao Qu, Zicong Hong, Bin Tang 0002, Junlong Zhou, Jiaying Yu |
IEEE Trans. Mob. Comput. | 5 |
| 2026 | A Unified Simulation Platform and Computation-Reuse Algorithm for Task Scheduling in Vehicle-Infrastructure CollaborationabstractVehicle-Infrastructure Collaboration (VIC) integrates vehicles and roadside infrastructure using advanced communication technologies, forming a crucial component of intelligent transportation systems (ITS). In a VIC system, tasks generated by vehicles can either be processed locally or offloaded to nearby edge servers. Current research often focuses on optimizing task scheduling but overlooks the inherent spatiotemporal correlations among tasks, which can lead to redundant computations due to similar tasks producing identical results. Additionally, the diversity in research scenarios and model constructions has resulted in the absence of a unified simulation verification platform, making it difficult to compare and validate various scheduling algorithms. To address these challenges, we have developed a comprehensive VIC simulation platform (CVSP). This platform not only features vehicle simulation capabilities like those of SUMO for modeling vehicle movement, but it also allows for the customization of driving scenarios and configurations, including edge resource settings, and incorporates a unified algorithm execution module for evaluating the performance of scheduling algorithms. Using CVSP can provide a clearer understanding of the strengths and weaknesses of scheduling algorithms, which in turn benefits the development of VIC systems. To tackle the spatiotemporal correlations observed in vehicular tasks, we propose a branch and bound algorithm based on computation-reuse (BB-CR). This algorithm integrates a computational reuse model derived from fused vehicle and road data. We simulate both non-congested and congested scenarios of vehicles traversing intersections to validate the performance of the baseline and BB-CR on the CVSP. The results highlight the versatility of the CVSP, showing that the BB-CR algorithm reduces system costs and minimizes the probability of task loss compared to the baseline. The code is publicly available at https://github.com/lzp991105/comprehensive-VIC-simulation-platform.git. Shihong Hu, Zhihao Qu, Bin Tang 0002, Xiongxiong Xu |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2025 | Efficient Target Tag Information Collection in Commodity RFID SystemsabstractWith the proliferation of RFID-enabled applications, large-scale RFID systems containing numerous tags are becoming increasingly common. Efficient management of these systems requires the ability to quickly collect information from specific subsets of tags, known as target tags. However, existing works often rely on hardware modifications, limiting their applicability to commercial RFID systems. To address this, this paper proposes the Target Tag Information Collection (TTIC) protocol and its enhanced version E-TTIC, both designed for efficient target tag collection using off-the-shelf RFID devices. TTIC leverages the Cuckoo filter to eliminate non-target tags and accurately collect target tags. It first inserts all target tags into the Cuckoo filter and then designs select commands compatible with commercial RFID readers to identify target tags while silencing non-target tags. On top of TTIC, E-TTIC further reduces the collection time by designing a novel Cuckoo filter that requires fewer select commands. We implement our protocols in a commodity RFID system. Extensive experiments show that E-TTIC can improve time efficiency by up to 80 % compared to the baseline. Zhenni Cao, Yanyan Wang 0001, Zhihao Qu, Bin Tang 0002 |
ICC | 3 |
| 2025 | Multi-cluster Layer-Sharing Container Scheduling in Cloud-Edge Collaboration
Yaoting Cao, Shihong Hu, Zhihao Qu, Lingling Hao |
ICIC (15) | 3 |
| 2025 | LCO-AGQ: A Lightweight Client-Oriented Adaptive Gradient Quantization Algorithm for Federated Learning
Hengrui Cui, Zhihao Qu, Bin Tang 0002 |
INFOCOM | 2 |
| 2025 | Multi-Range Query in Commodity RFID SystemsabstractRange Query (RQ) is to check whether there are any RFID tags with data beyond a given range. With about 46 billion RFID tags sold worldwide in 2023, time-efficient RQ becomes increasingly important for practical use, which can help users quickly pinpoint the target tags (if any) and give an early warning (e.g., fire alarm) to them for taking urgent actions and reducing the potential risk. However, existing work can deal with only a single range rather than multiple ranges that are very common in real-world applications. For example, foods in the refrigerator and the freezer have different temperature ranges for safe storing; treating them as one would probably give rise to query errors. In this paper, we study an under-investigated problem called multi-range query, which aims to achieve RQ in an RFID system with multiple query ranges. We propose a tailored protocol called anomalous tag identification (ATI) that quickly separates target tags from others and avoids querying all tags for saving communication overhead. In ATI, we design a fixedlength encoding vector together with standards-compliant select commands to deal with different ranges individually, without the need for any hardware modification. We implement the proposed protocols in commodity RFID systems. Experimental results show that ATI is superior to the baseline under different parameters, in terms of the time efficiency and space efficiency. Yanyan Wang 0001, Jia Liu 0008, Zhihao Qu, Shen-Huan Lyu, Bin Tang 0002 |
IWQoS | 3 |
| 2025 | Federated Learning via TEE-Based Dual-Branch Architecture and Interaction-Aware Pruning
Mingyang Xie, Zhihao Qu |
NPC (2) | 4 |
| 2025 | Layer-wise Adaptive Compression Method under Non-IID Settings for Federated LearningabstractFederated learning (FL) enables collaborative model training while preserving data privacy through decentralized data storage. However, the frequent transmission of high-dimensional model updates between FL clients and the central server incurs substantial communication overhead. Although prior studies compress model updates to reduce transmission overhead, fixed-rate schemes retain two major limitations: insensitivity to client-level Non-IID and uniform layer-wise compression, resulting in undercompression or over-compression of different layers. To overcome these issues, we propose a Layer-wise Adaptive Compression in Non-IID Situation (LWACN) algorithm, which applies global-local parameter similarity and client label entropy to measure the degree of client non-IID in compression. Moreover, we introduce window loss fluctuation and layer importance to mitigate the mismatching problem caused by the constant compression rate. Extensive experiments demonstrate that LWACN exhibits a better convergence rate and generalization ability than fixed compression. Specifically, compared to the state-of-the-art method, LWACN reduces transmission cost by up to 19.4%, and improves the final model accuracy by 5.5%. Ziyuan Feng, Zhihao Qu |
SMC | 4 |
| 2025 | Reliability-aware hybrid SFC backup and deployment in edge computing
Yue Zeng 0002, Shanshan Lin, Bin Tang 0002, Xiaoliang Wang 0001, Zhihao Qu, Song Guo 0001, Junlong Zhou |
Comput. Networks | 6 |
| 2025 | Multilevel Spatial-Temporal Joint Large Language Model for Traffic Prediction in Symbiotic IoTabstractThe Symbiotic Internet of Things (IoT) represents a collaborative framework wherein edge devices and AI models coordinate resource utilization and interact via 6G connectivity to optimize operational performance and system efficiency. Given the interdependent nature of this symbiotic relationship—where the performance and efficiency of each participant are significantly enhanced by the others—accurate traffic prediction becomes crucial. Despite the ability of existing deep learning models to model spatial-temporal dependencies, they face challenges in feature engineering and natural language feature fusion, particularly in few-shot learning scenarios. The recent advancement of pre-trained LLMs has demonstrated superior performance in time series analysis through their language comprehension and generalization capabilities. This work presents an innovative Multi-level Spatial-Temporal Joint Large Language Model (MSTJLLM) designed for traffic forecasting. The model incorporates multi-level embeddings of information flow, integrating historical features, dynamic spatial-temporal features, and prompt text features to capture complex dependencies. This approach aids LLMs in better understanding information flow through prompt text. Fine-tuning strategies enable the LLM to maintain language comprehension while enhancing spatial-temporal prediction accuracy. Tests on three real-world information flowsets demonstrate MSTJLLM’s robust prediction capability in both data-rich and data-scarce scenarios. Zhengwei Xu 0001, Shaopeng Xu, Zhihao Qu |
IEEE Internet Things J. | 3 |
| 2025 | Enhance learning efficiency of oblique decision tree via feature concatenation
Shen-Huan Lyu, Yi-Xiao He, Yanyan Wang 0001, Zhihao Qu, Bin Tang 0002 |
Inf. Sci. | 4 |
| 2025 | Scaling Persistent In-Memory Key-Value Stores Over Modern Tiered, Heterogeneous Memory HierarchiesabstractRecent advances in ultra-fast non-volatile memories (e.g., 3D XPoint) and high-speed interconnect fabrics (e.g., RDMA) enable a high-performance tiered, heterogeneous memory system, effectively overcoming the cost, scaling, and capacity limitations in DRAM-based key-value stores. To fully unleash the performance potential of such memory systems, this paper presents BonsaiKV+, a key-value store that makes the best use of different components in a modern RDMA-enabled heterogeneous memory system. The core of BonsaiKV+ is a tri-layer architecture that achieves efficient, elastic scaling up/out using a set of novel mechanisms and techniques—pipelined tiered indexing, NVM congestion control mechanisms, fine-grained data striping, and NUMA-aware data management—to leverage hardware strengths and tackle device deficiencies. We compare BonsaiKV+ with state-of-the-art key-value stores using a variety of YCSB workloads. Evaluation results demonstrate that BonsaiKV+ outperforms others by up to 7.30$\times$, 18.89$\times$, and 13.67$\times$in read-, write-, and scan-intensive scenarios, respectively. Miao Cai 0001, Junru Shen, Zhihao Qu |
IEEE Trans. Computers | 4 |
| 2025 | FedQClip: Accelerating Federated Learning via Quantized Clipped SGDabstractFederated Learning (FL) has emerged as a promising technique for collaboratively training machine learning models among multiple participants while preserving privacy-sensitive data. However, the conventional parameter server architecture presents challenges in terms of communication overhead when employing iterative optimization methods such as Stochastic Gradient Descent (SGD). Although communication compression techniques can reduce the traffic cost of FL during each training round, they often lead to degraded convergence rates, mainly due to compression errors and data heterogeneity. To address these issues, this paper presents FedQClip, an innovative approach that combines quantization and Clipped SGD. FedQClip leverages an adaptive step size inversely proportional to the$\ell_{2}$norm of the gradient, effectively mitigating the negative impacts of quantized errors. Additionally, clipped operations can be applied locally and globally to further expedite training. Theoretical analyses provide evidence that, even under the settings of Non-IID (non-independent and identically distributed) data, FedQClip achieves a convergence rate of$\mathcal{O}(\frac{1}{\sqrt{T}})$, effectively addressing the convergence degradation caused by compression errors. Furthermore, our theoretical analysis highlights the importance of selecting an appropriate number of local updates to enhance the convergence of FL training. Through extensive experiments, we demonstrate that FedQClip outperforms state-of-the-art methods in terms of communication efficiency and convergence rate. Zhihao Qu, Ninghui Jia, Shihong Hu, Song Guo 0001 |
IEEE Trans. Computers | 1 |
| 2025 | ExpertDRL: Request Dispatching and Instance Configuration for Serverless Edge Inference With Foundation ModelsabstractThe prevalence of the pre-training & fine-tuning paradigm enables machine learning models to quickly adapt to various downstream tasks by fine-tuning pre-trained foundation models (FMs), greatly facilitating various IoT applications that rely on model inference in dynamic edge serverless environments. Efficiently dispatching inference requests and configuring instances to batch inference requests can significantly enhance resource efficiency. However, existing serverless inference solutions are tailored for traditional models, make coarse-grained request dispatching and instance configuration decisions, fail to exploit the shared model backbone characteristics of the FM and capture delayed rewards in dynamic environments, and ignore communication latency between edge sites, resulting in high costs and constraint violations. In this paper, we leverage our insight that fine-grained batch inference requests can effectively exploit the shared model backbone feature of FM to save monetary costs. We propose an algorithm that incorporates deep reinforcement learning (DRL) and expert intervention for fine-grained request dispatching and instance configuration, where the DRL component outputs fractional solutions as guidance, while the expert intervention module integrates our insights—batching reduces monetary costs at the expense of increased inference latency, whereas higher configurations shorten inference latency. This module rounds fractional solutions and adjusts instance configurations to search for optimal solutions while satisfying constraints, with theoretical guarantees rigorously proved. Finally, we conducted our experiments on an OpenFaas-based platform and simulator, and extensive trace-driven evaluation results show that ExpertDRL can save costs by up to 85.14% and improve request acceptance ratio by up to 26.93%, compared to the state-of-the-art solution. Yue Zeng 0002, Junlong Zhou, Zhihao Qu, Song Guo 0001, Tianjian Gong |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | SPAVM: A SFC Placement and VNF Migration Framework for VNF Instance Reuse in Vehicle-Infrastructure CollaborationabstractIn the context of Vehicle-Infrastructure Collaboration (VIC), this study addresses the critical challenge of optimizing Service Function Chains (SFCs) deployment within resource-constrained and delay-sensitive vehicular edge networks. By leveraging Virtual Network Function (VNF) technology, which shifts network services from traditional hardware to a more agile, container-based edge computing architecture, we aim to enhance the Quality of Service (QoS) for vehicular users. SFCs, composed of multiple VNFs arranged in a specific sequence, are pivotal for delivering a range of functional services essential for QoS enhancement. Our objective is to optimize the placement of SFCs by reusing VNF instances to minimize both delay and operational costs. The reuse of VNF instances, however, introduces two significant challenges: the effective placement of SFCs within vehicular edge networks and the optimization of VNF instance containers positioning. To address these challenges, we propose a robust framework called SPAVM (Service Placement and VNF Migration), which consists of two primary components: one for SFC placement and the other for VNF migration. For SFC placement, we introduce the Service Placement based on VNF Instance Reuse algorithm (SPVIR), which maximizes the utilization of existing VNF container resources. For VNF migration, we propose the VNF Instance Migration algorithm (VIMA), which considers edge server connectivity to determine optimal migration targets for VNF instance containers. Extensive experiments validate the proposed algorithms' performance, demonstrating their effectiveness in reducing delay and costs, thereby enhancing the overall efficiency of vehicular edge networks. Shihong Hu, Zhihao Qu |
IEEE Trans. Serv. Comput. | 3 |
| 2024 | Mask-Encoded Sparsification: Mitigating Biased Gradients in Communication-Efficient Split LearningabstractThis paper introduces a novel framework designed to achieve a high compression ratio in Split Learning (SL) scenarios where resource-constrained devices are involved in large-scale model training. Our investigations demonstrate that compressing feature maps within SL leads to biased gradients that can negatively impact the convergence rates and diminish the generalization capabilities of the resulting models. Our theoretical analysis provides insights into how compression errors critically hinder SL performance, which previous methodologies underestimate. To address these challenges, we employ a narrow bit-width encoded mask to compensate for the sparsification error without increasing the order of time complexity. Supported by rigorous theoretical analysis, our framework significantly reduces compression errors and accelerates the convergence. Extensive experiments also verify that our method outperforms existing solutions regarding training efficiency and communication complexity. Our code can be found at https://github.com/BinaryMus/MaskSparsification. Zhihao Qu, Shen-Huan Lyu, Miao Cai 0001 |
ECAI | 2 |
| 2024 | On Efficient Zygote Container Planning and Task Scheduling for Edge Native Application AccelerationabstractEdge native applications usually consist of several dependent tasks encapsulated in containers and started on-demand in the edge cloud. Unfortunately, the application performance is deeply affected by the notorious cold startup problem of containers. Pre-warming Zygote container pre-imported certain common packages has been proven as an effective startup acceleration solution. Since a Zygote can be shared among colocated tasks that require identical common packages, not only the Zygote planning but also the task scheduling decisions shall be carefully made to maximize the benefit of the Zygotes pre-warmed in limited memory. Additionally, task dependency necessitates co-locating highly dependent tasks on the same server, naturally raising a dilemma in task scheduling. To this end, in this paper, we investigate the problem of how to plan Zygote and schedule tasks for application completion time minimization, which is proved to be NP-hard. We further propose a Priority and Popularity (P&P) based edge native application acceleration algorithm. Both theoretical analysis and extensive experiments demonstrate the effectiveness of our proposed algorithm. The experiment results show that P&P can reduce the application completion time by 11.7%. Yuepeng Li, Lin Gu 0002, Zhihao Qu, Lifeng Tian, Deze Zeng |
INFOCOM | 3 |
| 2024 | Container-Aware Service Function Chains Placement and Optimization in Vehicular Edge ComputingabstractIn the domain of Vehicular Edge Computing (VEC), this paper addresses the complex problem of Service Function Chain (SFC) placement, which is crucial for the efficient deployment of cloud applications in vehicular environments. Our approach leverages Virtual Network Function (VNF) technology, transitioning network services from traditional hardware to more flexible, container-based edge computing frameworks. The primary objective is to organize these VNFs into functional SFCs, thereby minimizing service delays in VEC systems. SFC placement faces two significant challenges. The first is the sequential dependency among VNFs within an SFC, which adds substantial complexity to container deployment. The second challenge is the cold start delay of VNF containers, a critical issue in scenarios requiring swift response times, which negatively impacts service quality in VEC applications. To address these challenges, we propose a comprehensive model for SFC placement that considers the deployment status of VNF containers, acknowledging the complexity of this NP-hard problem. The core of our contribution is the development of the Single SFC Placement Algorithm (SSPA), a sophisticated, greedy-based approach designed for the effective placement of individual SFCs. We enhance this algorithm by incorporating a Particle Swarm Optimization (PSO) technique, making it capable of efficiently handling the placement of multiple SFCs. Our solution aims to improve edge resource utilization and mitigate startup delays associated with the initial activation of containers, thereby reducing service delays. Extensive experimental evaluations demonstrate that our algorithm achieves a 34.3% average reduction in service delay compared to four baselines and a 5.8% average reduction compared to other container-aware algorithms. Shihong Hu, Zhihao Qu |
ISPA | 3 |
| 2024 | Identifying Key Tag Distribution in Large-Scale RFID SystemsabstractWith the proliferation of RFID-enabled applications, multiple readers are required for the complete coverage of numerous tags in a large-scale RFID system. In this scenario, we sometimes pay more attention to a subset of tags instead of all, which are referred to as key tags. In this paper, we study an under-investigated problem key tag distribution identification, which aims to identify which key tags are beneath which readers. This is crucial for efficiently managing specific items of interest, which can quickly pinpoint key tags and help RFID readers covering these tags collaborate to improve the tag inventory efficiency. Since key tags typically make up a small part of all tags, it is time consuming to deal with all tags in the traditional way. We propose a protocol called Kadept that identifies the key tag distribution by using a sophisticatedly designed filter that teases out key tags as well as assigns each of them a singleton slot for response. With this design, a great number of trivial (non-key) tags will keep silent and free up bandwidth resources for key tags, and each key tag is sorted in a collision-free way and can be identified with only 1-bit response, which significantly improves the time efficiency. We theoretically analyze how can we optimize protocol parameters of Kadept and conduct extensive simulations under different tag distribution scenarios. Compared with the state-of-the-art, Kadept can improve the time efficiency by a factor of 3.7×, when the ratio of key tags to all tags is 0.1. Yanyan Wang 0001, Jia Liu 0008, Shen-Huan Lyu, Zhihao Qu, Bin Tang 0002 |
IWQoS | 4 |
| 2024 | Gradient-Aware Incremental Network Quantization
Jiao Meng, Zhihao Qu, Shihong Hu |
NPC (2) | 2 |
| 2024 | Personalized Federated Learning with Feature Alignment via Knowledge Distillation
Guangfei Qi, Zhihao Qu, Shen-Huan Lyu, Ninghui Jia |
PRICAI (2) | 2 |
| 2024 | Boosting MLPs on Graphs via Distillation in Resource-Constrained EnvironmentabstractGraph Neural Networks (GNNs) have emerged as a powerful technique across various applications, due to their effective message-passing mechanism. However, their deployment is constrained by limited computational resources, energy concerns, and low-latency processing requirements. While existing works employ logit-based knowledge distillation from GNNs to guide Multilayer Perceptrons (MLPs) training, these methods may lead to reduced accuracy and compromised robustness. These drawbacks arise from two primary factors: the insufficient exploitation of the rich information embedded within the graph structures and the inherent susceptibility of MLPs to noisy data. To tackle these issues, we propose a Mixed Multi-order Knowledge Distillation (MMKD) method, which combines the GNN's logits with hidden layer information through the multi-order distillation to improve the accuracy of the MLP. Moreover, we employ both raw data and perturbed data as input, enhancing the density of knowledge extraction as well as the MLPs' generalization. Extensive experiments across seven benchmark datasets verify the superior performance of our approach in terms of effectiveness and robustness. In comparison with the baseline, our approach achieves an accuracy improvement of up to 8.68% in typical GNN tasks. Zhihao Qu, Ninghui Jia, Shihong Hu, Deze Zeng |
SMC | 2 |
| 2024 | Enhancing on-Device Inference Security Through TEE-Integrated Dual-Network ArchitectureabstractTrusted Execution Environment(TEE) offers a secure data processing zone for model inference. Due to its limited resources, existing solutions like MirrorNet usually deploy a lightweight model within the TEE for sensitive data, and a backbone model outside for the rest. However, these approaches do not inherently limit the learning ability of the backbone model, which could acquire inference capabilities similar to the lightweight model, inevitably weakening the security. To counter this, we propose FakeNN, an innovative mechanism with a dual-network architecture, that intentionally guides the backbone model towards low predictive performance, thereby reducing its ability to infer sensitive information. We further improve the accuracy of the entire model by integrating a channel attention mechanism which reduces the transmission of redundant information. We conduct extensive experiments and the results demonstrate that FakeNN substantially expands the performance gap between the non-secure and secure TEE models, with improvements ranging from 3.16% to 66.42% compared to MirrorNet. This enhancement strengthens the system's security without negatively impacting the accuracy of the model. Zhihao Qu, Ninghui Jia |
SMC | 2 |
| 2024 | Optimized Power Control for Privacy-Preserving Over-the-Air Federated Edge Learning With Device SamplingabstractOver-the-air federated edge learning (Air-FEEL) shows promise as a distributed machine learning paradigm for edge devices. By leveraging the superposition property of a multiple access channel (MAC), Air-FEEL can achieve low communication latency during training while enhancing the data privacy of edge devices, though at the expense of compromised learning performance. Recent studies suggest that optimizing the convergence speed of Air-FEEL can be accomplished by regulating the transmission power of edge devices while ensuring their differential privacy (DP). In this paper, we advance by incorporating device sampling in Air-FEEL (Air-FEEL-DS) to improve privacy and reduce device energy consumption, where each edge device decides randomly and independently whether to participate in each training round. Firstly, we theoretically characterize both the DP guarantee and convergence performance of Air-FEEL-DS. Then, we formulate a power control optimization problem to optimize the convergence speed while ensuring a specified DP guarantee. Despite the non-convex nature of this problem, we propose an efficient algorithm by linking it to a variant, transforming the variant into a convex problem, and demonstrating that the convex problem accommodates an efficient waterfilling-like algorithm. Finally, simulation results show that our proposed power control scheme achieves much faster convergence for Air-FEEL-DS than the channel inversion method, and has close convergence performance with significantly lower energy consumption compared to Air-FEEL with optimized power control but without device sampling. Bin Tang 0002, Bei Hu, Zhihao Qu |
IEEE Internet Things J. | 3 |
| 2024 | A Protocol Stack for Large-Scale RFID Systems: Mitigating Reader and Tag CollisionsabstractWith the proliferation of RFID-enabled applications, multiple RFID readers (or reader antennas) must be used to provide full coverage to any deployment area beyond the communication range of a single reader. However, reader collision together with tag-to-tag collision seriously degrades the system performance or even blocks out some tags from being read. To address this problem, this paper proposes a time-efficient protocol stack that is tailored to the tag identification in a multi-reader RFID system, which consists of two protocols: one is for eliminating reader collision (RCE) and the other is for avoiding tag-to-tag collision (TCE). In RCE, we enable multiple readers to work in parallel for maximizing the number of tags to be read per unit of time. Where RCE shines is that it well addresses the problem of unbalanced load by each reader due to uneven tag distributions. Namely, in existing work, the readers with fewer tags covered will finish the tag identification earlier than other readers (with more tags). After that, these readers have to wait for all readers until they are done. This is a waste of time. The solution of RCE is to take the reader with the minimum number of tags as the reference and iteratively to update the set of concurrent readers. Besides, in TCE, we use bit-level tag response to resolve tag-to-tag collision and increase the number of useful slots, so does the read throughput. Theoretical analysis and simulation results show that the above two protocols can jointly improve the inventory efficiency and reduce the identification time by more than 50%, in comparison to the state-of-the-art. We validate TIMR’s performance by comparing its results on a real-world library data set with simulated data, demonstrating consistency across both settings. Yanyan Wang 0001, Jia Liu 0008, Zhihao Qu, Wei Xiang 0001 |
IEEE Internet Things J. | 4 |
| 2024 | SafeDRL: Dynamic Microservice Provisioning With Reliability and Latency Guarantees in Edge EnvironmentsabstractAs a key technology of 5G, network function virtualization enables each monolithic service to be divided into microservices, facilitating their deployment and management in edge environments. One of the most critical issues in 5G is how to support dynamically arriving mission-critical services with low-latency and high-reliability requirements in distributed edge environments. However, most existing works focus on how to provide reliable services without considering latency, and their heuristics struggle to cope with high-dimensional constraints and complex environments with heterogeneous infrastructure and services. In this paper, we propose a SafeDRL algorithm to resource-efficiently support these dynamically arriving services while meeting their reliability and latency requirements. Specifically, we first formulate the problem as an integer nonlinear programming and prove its NP-hardness. To tackle this problem, our SafeDRL algorithm captures delayed rewards in dynamic environments by reinforcement learning, and corrects constraint violations with high-quality feasible solutions based on expert intervention, and prunes unnecessary backup instances for optimality. The algorithm is proved to have a bounded approximation ratio in general cases. Extensive trace-driven simulations show that, compared with the state-of-the-art solution, SafeDRL can save resource costs by up to 49.32% and improve the service acceptance ratio by up to 55% with acceptable execution time. Yue Zeng 0002, Zhihao Qu, Song Guo 0001, Jie Zhang 0076, Jing Li 0093, Bin Tang 0002 |
IEEE Trans. Computers | 2 |
| 2024 | Joint Service Request Scheduling and Container Retention in Serverless Edge Computing for Vehicle-Infrastructure CollaborationabstractLightweight and layered structure containers in serverless edge computing (SEC) provide flexible service configurations and computing for vehicles with diverse service requests in the Vehicle-Infrastructure Collaboration (VIC) environment. Despite progress in service request scheduling for the VIC system, the effect of layer sharing between different service images on request scheduling has not been fully explored. Additionally, the cold-start latency of service containers in SEC can significantly degrade the responsiveness of vehicle services, and container retention is proposed to minimize its impact and improve overall system performance. However, the existing research neglects the complex coupling relationship between request scheduling and container retention decisions, while focusing on the single decision optimization problem. Consequently, minimizing system costs by single decision optimization may not achieve the effect of joint decision optimization. To bridge this gap, we study the joint service request scheduling and container retention problem based on layer sharing and container caching. First, we model the joint decision problem with specific constraints and aim to minimize the long-term system cost while considering vehicle mobility. Second, an online co-decision scheme called Onco is proposed to solve the problem, which incorporates request scheduling and container retention for multiple vehicle services. Finally, both synthetic and real trace-driven simulation experiments have been conducted to evaluate the performance of Onco. The experimental results show that Onco outperforms state-of-the-art baselines in terms of system cost reduction and response time improvement. Shihong Hu, Zhihao Qu, Bin Tang 0002, Guanghui Li 0001, Weisong Shi |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | Joint Controller Placement and Flow Assignment in Software-Defined Edge Networks
Shunpeng Hua, Yue Zeng 0002, Zhihao Qu, Bin Tang 0002 |
ICA3PP (5) | 4 |
| 2023 | An Efficient Fault Tolerance Strategy for Multi-task MapReduce Models Using Coded Distributed Computing
Zaipeng Xie, Chenghong Xu, Zhihao Qu, Wen-Zhan Song 0001 |
ICA3PP (7) | 6 |
| 2023 | Anchor Sampling for Federated Learning with Partial Client ParticipationabstractCompared with full client participation, partial client participation is a more practical scenario in federated learning, but it may amplify some challenges in federated learning, such as data heterogeneity. The lack of inactive clients’ updates in partial client participation makes it more likely for the model aggregation to deviate from the aggregation based on full client participation. Training with large batches on individual clients is proposed to address data heterogeneity in general, but their effectiveness under partial client participation is not clear. Motivated by these challenges, we propose to develop a novel federated learning framework, referred to as FedAMD, for partial client participation. The core idea is anchor sampling, which separates partial participants into anchor and miner groups. Each client in the anchor group aims at the local bullseye with the gradient computation using a large batch. Guided by the bullseyes, clients in the miner group steer multiple near-optimal local updates using small batches and update the global model. By integrating the results of the two groups, FedAMD is able to accelerate the training process and improve the model performance. Measured by $\epsilon$-approximation and compared to the state-of-the-art methods, FedAMD achieves the convergence by up to $O(1/\epsilon)$ fewer communication rounds under non-convex objectives. Empirical studies on real-world datasets validate the effectiveness of FedAMD and demonstrate the superiority of the proposed algorithm: Not only does it considerably save computation and communication costs, but also the test accuracy significantly improves. Feijie Wu, Song Guo 0001, Zhihao Qu, Shiqi He, Jing Gao 0004 |
ICML | 3 |
| 2023 | Joint Service Placement and Container Retention for Serverless-Based Vehicular Edge ComputingabstractLightweight containers in serverless-based vehicular edge computing (SVEC) offer flexible service provisioning for edge service providers (ESPs), enabling quick response to various service requests from mobile vehicles and improving the quality of service delivery. However, the unpredictability of vehicle mobility and the variability of service requests pose significant challenges to service placement for ESPs. Moreover, the cold-start latency experienced by service containers in SVEC can greatly hinder the responsiveness of vehicle services. To mitigate this impact and enhance the overall system performance, container retention is introduced as a solution. In this paper, we study the joint service placement and container retention problem in the dynamic SVEC system. First, we model the joint decision problem with specific constraints and aim to maximize the profit of each edge considering vehicle mobility. Second, an online co-decision scheme called CMU-O is proposed to solve the problem, which adopts an improved upper confidence bound method based on dynamic osmotic pressure (UCB-OP). Finally, the experimental results demonstrate that the service request success rate of the CMU-O is on average 7% higher than the baselines. Furthermore, the profits obtained under the CMU-O are also on average 23% higher than the baselines. Shihong Hu, Zhihao Qu, Bin Tang 0002 |
ICPADS | 2 |
| 2023 | Handover Analysis with Spatially Correlated Blockage ModelabstractIn vehicular networks, handover can occur due to frequent changes in communication link status and transmission distance caused by the mobility of vehicles. Although handover is necessary for maintaining stable communication performance, it may cause interruptions in computing services within the distributed vehicular system. Recent literature has analysed the impact of blockage and mobility on handover performance. However, these works often assume that blockage probability during vehicle movement follows an independent distribution for ease of derivation. This assumption ignores some obstacles that can consecutively affect the blockage probability of the communication link during movement, leading to an inaccurate evaluation of the handover performance. To address this issue, this paper constructs a vehicular network scenario based on a spatially correlated blockage model where roadside obstacles and vehicle movement trajectories can jointly affect the communication link status between vehicles and roadside units. Using this model, we characterize the correlation between blockage probabilities during movement and theoretically derive a closedform expression for handover probability. Simulation results verify the accuracy of the analytical results while also revealing how factors such as vehicle speed, roadside unit density, blockage length affect handover performance. Bin Tang 0002, Zhihao Qu |
MSN | 4 |
| 2023 | Federated Learning with Common Representation Learning Criterion and Personalized PredictorabstractFederated learning (FL) enables model training on decentralized devices while preserving data privacy. However, data heterogeneity poses a significant challenge to FL, and various approaches have been proposed to address it. Existing research has mainly focused on either enhancing global models or customizing personalized models for clients. This paper proposes a novel approach, FedCRC, that decouples the machine learning model into a representation extractor and predictor. This enables us to enhance both generalization and personalization, thereby addressing the challenge of data heterogeneity in FL. The approach employs a stable global predictor to unify the representation learning criterion during the training of the representation extractor. Additionally, a personalized predictor is trained for each client to achieve a personalized model tailored to the local data distribution. Our FedCRC algorithm was evaluated on multiple benchmark datasets with varying distributions, covering diverse settings. Extensive experimental results demonstrate the effectiveness of our method. Wenzhong Wang, Zaipeng Xie, Bingzhe Yu, Zhihao Qu, Hongli Cao |
SMC | 4 |
| 2023 | BonsaiKV: Towards Fast, Scalable, and Persistent Key-Value Stores with Tiered, Heterogeneous Memory SystemabstractEmerging NUMA/CXL-based tiered memory systems with heterogeneous memory devices such as DRAM and NVMM deliver ultrafast speed, large capacity, and data persistence all at once, offering great promise to high-performance in-memory key-value stores. To fully unleash the performance potential of such memory systems, this paper presents BonsaiKV, a key-value store that makes the best use of different components in a tiered memory system. The core of BonsaiKV is a tri-layer hierarchical storage architecture that separates data indexing, persistence, and scalability from each other and realizes each of them within a specialized software-hardware layer. We design BonsaiKV with a set of novel techniques, including collaborative tiered indexing, NVMM congestion control mechanisms, fine-grained data striping, and NUMA-aware data management, to leverage hardware strengths and tackle device deficiencies. We compare BonsaiKV with state-of-the-art NVMM-optimized key-value stores and persistent index structures using a variety of YCSB workloads. Evaluation results demonstrate that BonsaiKV outperforms others by up to 7.69×, 19.59×, and 12.86× in read-, write- and scan-intensive scenarios, respectively. Miao Cai 0001, Junru Shen, Zhihao Qu |
Proc. VLDB Endow. | 4 |
| 2023 | Mobility-Aware Proactive Flow Setup in Software-Defined Mobile Edge NetworksabstractThe software-defined network (SDN) enabled mobile edge network greatly facilitates network resource management and promotes many emerging applications. However, user mobility may cause the SDN controller to set flow rules frequently, introduce additional flow setup latency, cause delay jitter, and undermine latency-sensitive services. Proactive flow setup is an effective way to eliminate flow setup latency, but existing work fails to maximize the flow setup hit ratio, a metric for evaluating the quality of proactive flow setup decisions, which is critical for latency-sensitive services. In this paper, we study how to proactively set flow rules to maximize the flow setup hit ratio under limited available network resources to eliminate the flow setup latency as much as possible. Then, we formalize the proactive flow setup problem as two integer linear programming problems under two typical routing strategies, default routing and dynamic routing. Both problems are proved to be NP-hard. To tackle these two problems, we propose a linear programming-based polynomial-time approximation algorithm for the default routing case and a greedy-based heuristic algorithm for the dynamic routing case. Extensive trace-driven experimental and simulation results verify that our algorithms can improve the flow setup hit ratio by up to 30.99% compared to existing solutions. Yue Zeng 0002, Bin Tang 0002, Sanglu Lu, Feng Xu 0008, Song Guo 0001, Zhihao Qu |
IEEE Trans. Commun. | 7 |
| 2023 | From Deterioration to Acceleration: A Calibration Approach to Rehabilitating Step Asynchronism in Federated OptimizationabstractIn the setting of federated optimization, where a global model is aggregated periodically, step asynchronism occurs when participants conduct model training by efficiently utilizing their computational resources. It is well acknowledged that step asynchronism leads to objective inconsistency under non-i.i.d. data, which degrades the model’s accuracy. To address this issue, we propose a new algorithmFedaGrac, which calibrates the local direction to a predictive global orientation. Taking advantage of the estimated orientation, we guarantee that the aggregated model does not excessively deviate from the global optimum while fully utilizing the local updates of faster nodes. We theoretically prove thatFedaGracholds an improved order of convergence rate than the state-of-the-art approaches and eliminates the negative effect of step asynchronism. Empirical results show that our algorithm accelerates the training and enhances the final accuracy. Feijie Wu, Song Guo 0001, Haozhao Wang, Haobo Zhang 0002, Zhihao Qu, Jie Zhang 0076 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2023 | RuleDRL: Reliability-Aware SFC Provisioning With Bounded Approximations in Dynamic EnvironmentsabstractAs a key enabling technology for 5G, network function virtualization abstracts services into software-based service function chains (SFCs), facilitating mission-critical services with high-reliability requirements. However, it is challenging to cost-effectively provide reliable SFCs in dynamic environments due to delayed rewards caused by future SFC requests, limited infrastructure resources, and heterogeneity in hardware and software reliability. Although deep reinforcement learning (DRL) can effectively capture delayed rewards in dynamic environments, its trial-and-error exploration in a vast solution space with massive infeasible solutions may lead to frequent constraint violations and traps in poor local optima. To address these challenges, we propose a RuleDRL algorithm that combines the capability of DRL to capture delayed rewards and the strength of rule-based schemes to explore high-quality solutions without violating constraints. Specifically, we first formulate the reliable SFC provision problem as an integer nonlinear programming problem, which is proven to be NP-hard. Then, we jointly design DRL and rule-based schemes that are coupled to make the final decision and establish a bounded approximation ratio in general cases. Extensive trace-driven simulations show that RuleDRL can save the total cost by up to 65.67% and improve the SFC acceptance ratio by up to 82%, compared to the state-of-the-art solution. Yue Zeng 0002, Zhihao Qu, Song Guo 0001, Bin Tang 0002, Jing Li 0093, Jie Zhang 0076 |
IEEE Trans. Serv. Comput. | 2 |
| 2022 | Sign bit is enough: a learning synchronization framework for multi-hop all-reduce with ultimate compressionabstractTraditional one-bit compressed stochastic gradient descent can not be directly employed in multi-hop all-reduce, a widely adopted distributed training paradigm in network-intensive high-performance computing systems such as public clouds. According to our theoretical findings, due to the cascading compression, the training process has considerable deterioration on the convergence performance. To overcome this limitation, we implement a sign-bit compression-based learning synchronization framework, Marsit. It prevents cascading compression via an elaborate bit-wise operation for unbiased sign aggregation and its specific global compensation mechanism for mitigating compression deviation. The proposed framework retains the same theoretical convergence rate as non-compression mechanisms. Experimental results demonstrate that Marsit reduces up to 35% training time while preserving the same accuracy as training without compression. Feijie Wu, Shiqi He, Song Guo 0001, Zhihao Qu, Haozhao Wang, Weihua Zhuang, Jie Zhang 0076 |
DAC | 4 |
| 2022 | Toward Performance Efficient UAV Task Scheduling in Cloud Native EdgeabstractUnmanned Aerial Vehicle (UAV) has been widely applied in many domains. But the computation and energy resource limitation severely hinders its development and application. Mobile Edge Computing (MEC) emerges as a promising platform to process the tasks offloaded from the UAVs to effectively improve the Quality-of-Service (QoS). To this vision, it is first required that the edge servers must be deployed with the needed service to handle the offloaded task. Fortunately, by exploring cloud native computing technology, it is possible to deploy container-based microservice to MEC in a prompt way. In this case, it raises the task scheduling problem on whether to deploy a new service or to utilize an existing service to balance the overhead between data transmission and the microservice deployment (i.e., container image pulling) for overall task completion time minimization. In this paper, the problem is first formulated in Integer Linear Programming (ILP) form, and proved to be NP-hard. We further propose an incentive-based request scheduling algorithm. Experiments based on track-driven simulations show that the total completion time for all tasks is reduced by 21.37% compared to the state-of-the-art solution. Deze Zeng, Zhihao Qu |
GLOBECOM | 3 |
| 2022 | Advanced Persistent Threat Detection in Smart Grid Clouds Using Spatiotemporal Context-Aware Graph EmbeddingabstractAdvanced persistent threat (APT) attacks have caused severe damage to many core information infrastructures. To tackle this issue, the graph-based methods have been proposed due to their ability for learning complex interaction patterns of network entities with discrete graph snapshots. However, such methods are challenged by the computer networking model characterized by a natural continuous-time dynamic heterogeneous graph. In this paper, we propose a heterogeneous graph neural network based APT detection method in smart grid clouds. Our model is an encoder-decoder structure. The encoder uses heterogeneous temporal memory and attention embedding modules to capture contextual information of interactions of network entities from the time and spatial dimensions respectively. We implement a prototype and conduct extensive experiments on real-world cyber-security datasets with more than 10 million records. Experimental results show that our method can achieve superior detection performance than state-of-the-art methods. Weiyong Yang, Xingshen Wei, Zhihao Qu |
GLOBECOM | 6 |
| 2022 | FedDGIC: Reliable and Efficient Asynchronous Federated Learning with Gradient CompensationabstractAsynchronous federated learning is a distributed machine learning paradigm that may alleviate the impact of straggler nodes and improve the efficiency of federated training. However, some nodes can become sluggish, and node dropout may frequently happen for various reasons, such as network connection constraints, energy deficits, and system faults. Consequently, the global model may deviate from the desired convergence direction and lead to suboptimal results. This work proposes an asynchronous federated learning framework, FedDGIC, to mitigate the impact of the node dropout problem. The proposed framework can improve training efficiency by utilizing a dynamic grouping algorithm with gradient compensation. Experiments are performed in a real federated learning environment using two datasets, i.e., MNIST and CIFAR-10. Compared with three state-of-the-art methods, the proposed FedDGIC can significantly improve training efficiency and provide reliable asynchronous federated learning. Zaipeng Xie, Junchen Jiang, Zhihao Qu, Hanxiang Liu |
ICPADS | 4 |
| 2022 | Hierarchical Channel-spatial Encoding for Communication-efficient Collaborative LearningabstractIt witnesses that the collaborative learning (CL) systems often face the performance bottleneck of limited bandwidth, where multiple low-end devices continuously generate data and transmit intermediate features to the cloud for incremental training. To this end, improving the communication efficiency by reducing traffic size is one of the most crucial issues for realistic deployment. Existing systems mostly compress features at pixel level and ignore the characteristics of feature structure, which could be further exploited for more efficient compression. In this paper, we take new insights into implementing scalable CL systems through a hierarchical compression on features, termed Stripe-wise Group Quantization (SGQ). Different from previous unstructured quantization methods, SGQ captures both channel and spatial similarity in pixels, and simultaneously encodes features in these two levels to gain a much higher compression ratio. In particular, we refactor feature structure based on inter-channel similarity and bound the gradient deviation caused by quantization, in forward and backward passes, respectively. Such a double-stage pipeline makes SGQ hold a sublinear convergence order as the vanilla SGD-based optimization. Extensive experiments show that SGQ achieves a higher traffic reduction ratio by up to 15.97 times and provides 9.22 times image processing speedup over the uniform quantized training, while preserving adequate model accuracy as FP32 does, even using 4-bit quantization. This verifies that SGQ can be applied to a wide spectrum of edge intelligence applications. Qihua Zhou, Song Guo 0001, Yi Liu 0057, Jie Zhang 0076, Jiewei Zhang, Tao Guo 0004, Zhenda Xu, Zhihao Qu |
NeurIPS | 9 |
| 2022 | FedALP: An Adaptive Layer-Based Approach for Improved Personalized Federated Learning
Zaipeng Xie, Zhihao Qu, Bin Tang 0002, Weiyi Zhao |
WASA (2) | 3 |
| 2022 | A Comprehensive Survey on Training Acceleration for Large Machine Learning Models in IoTabstractThe ever-growing artificial intelligence (AI) applications have greatly reshaped our world in many areas, e.g., smart home, computer vision, natural language processing, etc. Behind these applications are usually machine learning (ML) models with extremely large size, which require huge data sets for accurate training to mine the value contained in the big data. Large ML models, however, can consume tremendous computing resources to achieve decent performance and thus, it is difficult to train them in resource-constrained Internet of Things (IoT) environments, which would prevent further development and application of AI techniques in the future. To deal with such challenges, there are many efforts on accelerating the training process for large ML models in IoT. In this article, we provide a comprehensive review on the recent advances toward reducing the computing cost during the training stage while maintaining comparable model accuracy. Specifically, the optimization algorithms that aim to improve the convergence rate are emphasized over various distributed learning architectures that exploit ubiquitous computing resources. Then, the article elaborates the computation hardware acceleration and communication optimization for collaborative training among multiple learning entities. Finally, the remaining challenges, future opportunities, and possible directions are discussed. Haozhao Wang, Zhihao Qu, Qihua Zhou, Haobo Zhang 0002, Boyuan Luo, Wenchao Xu 0001, Song Guo 0001, Ruixuan Li 0001 |
IEEE Internet Things J. | 2 |
| 2022 | FastCache: A write-optimized edge storage system via concurrent merging cache for IoT applications
Lin Qian, Zhihao Qu, Miao Cai 0001, Xiaoliang Wang 0001, Weiguo Duan |
J. Syst. Archit. | 2 |
| 2022 | Heterogeneity-Aware Gradient Coding for Tolerating and Leveraging StragglersabstractDistributed gradient descent has been widely adopted in the machine learning field because considerable computing resources are available when facing the huge volume of data. Specifically, the gradient over the whole data is cooperatively computed by multiple workers. However, its performance can be severely affected by slow workers, namely stragglers. Recently, coding-based approaches have been introduced to mitigate the straggler problem, but they could hardly deal with the heterogeneity among workers. Besides, they always discard the results of stragglers causing huge resource waste. In this article, we first investigate how to tolerate stragglers by discarding their results and then seek to leverage the stragglers. For tolerating stragglers, we propose a heterogeneity-aware coding scheme that encodes gradients adaptive to the computing capability of workers. Theoretically, this scheme is optimal for stragglers tolerance. Relying on the scheme, we further propose an algorithm called DHeter-aware to exploit the gradients of stragglers which we called delayed gradients. Moreover, theoretical results characterized for DHeter-aware exhibits the same convergence rate as the gradient descent without delayed gradients. Experiments on various tasks and clusters demonstrate that our coding scheme outperforms all the state-of-the-art methods and the DHeter-aware further accelerates the coding scheme by achieving 25 percent time savings. Haozhao Wang, Song Guo 0001, Bin Tang 0002, Ruixuan Li 0001, Yutong Yang, Zhihao Qu, Yi Wang 0004 |
IEEE Trans. Computers | 6 |
| 2022 | Adaptive Federated Learning on Non-IID Data With Resource ConstraintabstractFederated learning (FL) has been widely recognized as a promising approach by enabling individual end-devices to cooperatively train a global model without exposing their own data. One of the key challenges in FL is the non-independent and identically distributed (Non-IID) data across the clients, which decreases the efficiency of stochastic gradient descent (SGD) based training process. Moreover, clients with different data distributions may cause bias to the global model update, resulting in a degraded model accuracy. To tackle the Non-IID problem in FL, we aim to optimize the local training process and global aggregation simultaneously. For local training, we analyze the effect of hyperparameters (e.g., the batch size, the number of local updates) on the training performance of FL. Guided by the toy example and theoretical analysis, we are motivated to mitigate the negative impacts incurred by Non-IID data via selecting a subset of participants and adaptively adjust their batch size. A deep reinforcement learning based approach has been proposed to adaptively control the training of local models and the phase of global aggregation. Extensive experiments on different datasets show that our method can improve the model accuracy by up to 30 percent, as compared to the state-of-the-art approaches. Jie Zhang 0076, Song Guo 0001, Zhihao Qu, Deze Zeng, Yufeng Zhan, Rajendra Akerkar |
IEEE Trans. Computers | 3 |
| 2022 | Partial Synchronization to Accelerate Federated Learning Over Relay-Assisted Edge NetworksabstractFederated Learning (FL) is a promising machine learning paradigm to cooperatively train a global model with highly distributed data located on mobile devices. Aiming to optimize the communication efficiency for gradient aggregation and model synchronization among large-scale devices, we propose a relay-assisted FL framework. By breaking the traditional transmission-order constraint and exploiting the broadcast characteristic of relay nodes, we design a novel synchronization scheme named Partial Synchronization Parallel (PSP), in which models and gradients are transmitted simultaneously and aggregated at relay nodes, resulting in traffic reduction. We prove that PSP has the same convergence rate as the sequential synchronization approaches via rigorous analysis. To further accelerate the training process, we integrate PSP with any unbiased and error-bounded compression technologies and prove that the convergence properties of the resulting scheme still hold. Extensive experiments are conducted in a distributed cluster environment with real-world datasets and the results demonstrate that our proposed approach reduces the training time up to 37 percent compared to state-of-the-art methods. Zhihao Qu, Song Guo 0001, Haozhao Wang, Yi Wang 0004, Albert Y. Zomaya, Bin Tang 0002 |
IEEE Trans. Mob. Comput. | 1 |
| 2022 | Error-Compensated Sparsification for Communication-Efficient Decentralized Training in Edge EnvironmentabstractCommunication has been considered as a major bottleneck in large-scale decentralized training systems since participating nodes iteratively exchange large amounts of intermediate data with their neighbors. Although compression techniques like sparsification can significantly reduce the communication overhead in each iteration, errors caused by compression will be accumulated, resulting in a severely degraded convergence rate. Recently, the error compensation method for sparsification has been proposed in centralized training to tolerate the accumulated compression errors. However, the analog technique and the corresponding theory about its convergence in decentralized training are still unknown. To fill in the gap, we design a method named ECSD-SGD that significantly accelerates decentralized training via error-compensated sparsification. The novelty lies in that we identify the component of the exchanging information in each iteration (i.e., the sparsified model update) and make targeted error compensation over the component. Our thorough theoretical analysis shows that ECSD-SGD supports arbitrary sparsification ratio and achieves the same convergence rate as the non-sparsified decentralized training methods. We also conduct extensive experiments on multiple deep learning models to validate our theoretical findings. Results show that ECSD-SGD outperforms all the start-of-the-art sparsified methods in terms of both the convergence speed and the final generalization accuracy. Haozhao Wang, Song Guo 0001, Zhihao Qu, Ruixuan Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | Adaptive Vertical Federated Learning on Unbalanced FeaturesabstractMost of the existing FL systems focus on a data-parallel architecture where training data are partitioned by samples among several parties. In some real-life applications, however, partitioning by features is also of practical relevance and the number of features is usually unbalanced among parties. The corresponding learning framework is referred to as Vertical Federated Learning (VFL). Though some pioneering work focused on VFL, the convergence properties of VFL on unbalanced features, especially when parties conduct different numbers of local updates concerning heterogeneous computational capabilities are still unknown. In this article, we propose a new learning framework to improve the training efficiency of VFL on unbalanced features. Given the number of features and the computational capability owned by each party, our thorough theoretical analysis exhibits that the number of local updates conducted by each party has a great effect on the convergence rate and the computational complexity, both of which jointly determine the overall training efficiency in an interrelated and sophisticated way. Based on our theoretical findings, we formulate an optimization problem and derive the optimal solution by selecting an adaptive number of local training rounds for each party. Extensive experiments on various datasets and models demonstrate that our approach significantly improves the training efficiency of VFL. Jie Zhang 0076, Song Guo 0001, Zhihao Qu, Deze Zeng, Haozhao Wang, Albert Y. Zomaya |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | FastCache: A Client-Side Cache with Variable-Position Merging Schema in Network Storage System
Lin Qian, Xiaoliang Wang 0001, Zhihao Qu, Weiguo Duan |
ICA3PP (2) | 4 |
| 2021 | Octo: INT8 Training with Loss-aware Compensation and Backward Quantization for Tiny On-device Learning
Qihua Zhou, Song Guo 0001, Zhihao Qu, Jingcai Guo, Zhenda Xu, Jiewei Zhang, Tao Guo 0004, Boyuan Luo, Jingren Zhou 0001 |
USENIX ATC | 3 |
| 2021 | Communication-efficient Federated Learning via Quantized Clipped SGD
Ninghui Jia, Zhihao Qu |
WASA (1) | 2 |
| 2021 | Scheduling coflows of multi-stage jobs under network resource constraints
Yue Zeng 0002, Bin Tang 0002, Songtao Guo, Zhihao Qu |
Comput. Networks | 5 |
| 2021 | On-Device Learning Systems for Edge Intelligence: A Software and Hardware Synergy PerspectiveabstractModern machine learning (ML) applications are often deployed in the cloud environment to exploit the computational power of clusters. However, this in-cloud computing scheme cannot satisfy the demands of emerging edge intelligence scenarios, including providing personalized models, protecting user privacy, adapting to real-time tasks, and saving resource cost. In order to conquer the limitations of conventional in-cloud computing, there comes the rise of on-device learning, which makes the end-to-end ML procedure totally on user devices, without unnecessary involvement of the cloud. In spite of the promising advantages of on-device learning, implementing a high-performance on-device learning system still faces with many severe challenges, such as insufficient user training data, backward propagation (BP) blocking, and limited peak processing speed. Observing the substantial improvement space in the implementation and acceleration of on-device learning systems, we intend to present a comprehensive analysis of the latest research progress and point out potential optimization directions from the system perspective. This survey presents a software and hardware synergy of on-device learning techniques, covering the scope of model-level neural network design, algorithm-level training optimization, and hardware-level instruction acceleration. We hope this survey could bring fruitful discussions and inspire the researchers to further promote the field of edge intelligence. Qihua Zhou, Zhihao Qu, Song Guo 0001, Boyuan Luo, Jingcai Guo, Zhenda Xu, Rajendra Akerkar |
IEEE Internet Things J. | 2 |
| 2021 | LOSP: Overlap Synchronization Parallel With Local Compensation for Fast Distributed TrainingabstractWhen running in Parameter Server (PS), the Distributed Stochastic Gradient Descent (D-SGD) incurs significant communication delays and huge communication overhead due to the model synchronization. Moreover, considering the heterogeneity of computational capability among workers, traditional synchronization modes incur under-utilization of computational resources because fast workers have to wait for slow ones finishing the computation. Although our previous work OSP can effectively solve these problems by overlapping the computation and communication procedures and allowing adaptive multiple local updates in distributed training, it causes the staleness problem brought by the overlap, yielding a performance degradation. In this paper, we propose a new method named LOSP by introducing local compensation to our previous synchronization mechanism, which mitigates adverse effects caused by the overlapping synchronization. We theoretically prove that LOSP (1) preserves the same convergence rate as the sequential SGD for non-convex problems, and (2) exhibits good scalability due to the linear speedup property with respect to both the number of workers and the average number of local updates. Evaluations show that LOSP significantly improves performance over the state-of-the-art ones in terms of both convergence accuracy and communication cost. Haozhao Wang, Zhihao Qu, Song Guo 0001, Ningqi Wang, Ruixuan Li 0001, Weihua Zhuang |
IEEE J. Sel. Areas Commun. | 2 |
| 2021 | Petrel: Heterogeneity-Aware Distributed Deep Learning Via Hybrid SynchronizationabstractThe parameter server (PS) paradigm has achieved great success in deploying large-scale distributed Deep Learning (DL) systems. However, these systems implicitly assume that the cluster is homogeneous and this assumption does not hold in many realworld cases. Although the previous efforts are paid to address heterogeneity, they mainly prioritize the contribution of fast workers and reduce the involvement of slow workers, resulting in the limitations of workload imbalance and computation inefficiency. We reveal that grouping workers into communities, an abstraction proposed by us, and handling parameter synchronization at the community level can conquer these limitations and accelerate the training convergence progress. The inspiration of community comes from our exploration of prior knowledge about the similarity between workers, which is often neglected by previous work. These observations motivate us to propose a new synchronization mechanism named Community-aware Synchronous Parallel (CASP), which uses the Asynchronous Advantage Actor-Critic (A3C)-based algorithm to intelligently determine community configuration and fully improve the synchronization performance. The whole idea has been implemented in a prototype system called Petrel that achieves a good balance between convergence efficiency and communication overhead. The evaluation under various benchmarks with multiple metrics and baseline comparison demonstrates the effectiveness of Petrel. Specifically, Petrel accelerates the training convergence speed by up to 1.87 x faster and reduces communication traffic by up to 26.85 percent, on average, over the non-community synchronization mechanisms. Qihua Zhou, Song Guo 0001, Zhihao Qu, Peng Li 0017, Li Li 0012, Minyi Guo, Kun Wang 0005 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | Intermediate Value Size Aware Coded MapReduceabstractMapReduce is a commonly used framework for parallel processing of data-intensive tasks, but its performance usually suffers from heavy communication load incurred by the shuffling of intermediate values (IVs) among computing servers. Recently, the Coded MapReduce framework is proposed which uses a coding scheme named coded distributed computing (CDC) to trade the communication load with extra computation in MapReduce. CDC can achieve the optimal computation-communication tradeoff when all the IVs have the same size. However, in many practical applications, the sizes of IVs can vary over a large range, leading to inferior performance. In this paper, we introduce a generalized CDC scheme which takes the sizes of IVs into account and then propose a combinatorial optimization problem aiming to minimize the communication load when the computation load is fixed. We show that the problem is NP-hard, and further propose a very efficient algorithm which achieves an approximation ratio of 2. Experiments conducted on Alibaba Cloud show that, compared to the original CDC scheme, our proposed IV size aware approach can significantly reduce the communication load and achieve a lower total execution time. Yamei Dong, Bin Tang 0002, Zhihao Qu, Sanglu Lu |
ICPADS | 4 |
| 2020 | Joint Service Placement and Computation Offloading in Mobile Edge Computing: An Auction-based ApproachabstractThe emerging applications, e.g., virtual reality, online games, and Internet of Vehicles, have computation-intensive and latency-sensitive requirements. Mobile edge computing (MEC) is a powerful paradigm that significantly improves the quality of service (QoS) of these applications by offloading computation and deploying services at the network edge. Existing works on service placement in MEC usually ignore the impact of the different requirements of QoS among service providers (SPs), which is common in many applications such that online game requires extremely low latency and online video requires extremely large bandwidth. Considering the competitive relationship among SPs, we propose an auction-based resource allocation mechanism. We formulate the problem as a social welfare maximization problem to maximize effectiveness of allocated resources while maintaining economic robustness. According to our theoretical analysis, this problem is NP-hard, and thus it is practically impossible to derive the optimal solution. To tackle this, we design multiple rounds of iterative auctions mechanism (MRIAM), which divides resources into blocks and allocates them through multiple rounds of auctions. Finally, we conduct extensive experiments and demonstrate that our auction-based mechanism is effective in resource allocation and robust in economics. Zhihao Qu, Bin Tang 0002 |
ICPADS | 2 |
| 2020 | Physical-Layer Arithmetic for Federated Learning in Uplink MU-MIMO Enabled Wireless NetworksabstractFederated learning is a very promising machine learning paradigm where a large number of clients cooperatively train a global model using their respective local data. In this paper, we consider the application of federated learning in wireless networks featuring uplink multiuser multiple-input and multiple-output (MU-MIMO), and aim at optimizing the communication efficiency during the aggregation of client-side updates by exploiting the inherent superposition of radio frequency (RF) signals. We propose a novel approach named Physical-Layer Arithmetic (PhyArith), where the clients encode their local updates into aligned digital sequences which are converted into RF signals for sending to the server simultaneously, and the server directly recovers the exact summation of these updates as required from the superimposed RF signal by employing a customized sum-product algorithm. PhyArith is compatible with commodity devices due to the use of full digital operation in both the client-side encoding and the server-side decoding processes, and can also be integrated with other updates compression based acceleration techniques. Simulation results show that PhyArith further improves the communication efficiency by 1.5 to 3 times for training LeNet-5, compared with solutions only applying updates compression. Tao Huang 0007, Zhihao Qu, Bin Tang 0002, Lei Xie 0004, Sanglu Lu |
INFOCOM | 3 |
| 2020 | Rateless802.11: Extending WiFi applicability in extremely poor channels
Tao Huang 0007, Bin Tang 0002, Zhihao Qu, Sanglu Lu |
Comput. Networks | 4 |
| 2020 | A Learning-Based Incentive Mechanism for Federated LearningabstractInternet of Things (IoT) generates large amounts of data at the network edge. Machine learning models are often built on these data, to enable the detection, classification, and prediction of the future events. Due to network bandwidth, storage, and especially privacy concerns, it is often impossible to send all the IoT data to the data center for centralized model training. To address these issues, federated learning has been proposed to let nodes use the local data to train models, which are then aggregated to synthesize a global model. Most of the existing work has focused on designing learning algorithms with provable convergence time, but other issues, such as incentive mechanism, are unexplored. Although incentive mechanisms have been extensively studied in network and computation resource allocation, yet they cannot be applied to federated learning directly due to the unique challenges of information unsharing and difficulties of contribution evaluation. In this article, we study the incentive mechanism for federated learning to motivate edge nodes to contribute model training. Specifically, a deep reinforcement learning-based (DRL) incentive mechanism has been designed to determine the optimal pricing strategy for the parameter server and the optimal training strategies for edge nodes. Finally, numerical experiments have been implemented to evaluate the efficiency of the proposed DRL-based incentive mechanism. Yufeng Zhan, Peng Li 0017, Zhihao Qu, Deze Zeng, Song Guo 0001 |
IEEE Internet Things J. | 3 |
| 2020 | Cooperative Caching for Multiple Bitrate Videos in Small Cell EdgesabstractCaching popular videos at mobile edge servers (MESs) has been confirmed as a promising method to improve mobile users (MUs) perceived quality of experience (QoE) and to alleviate the server load. However, with the multiple bitrate encoding techniques prevalently employed in modern streaming services, caching deployment is challenging for the following three facts: (1) cooperative caching should be explored for MUs located at overlapped coverage areas of MESs; (2) there exists tradeoff consideration for caching either high bitrate videos or high diversity videos; and (3) the relationship between MU perceived QoE and MU received bitrate, known as QoE function, varies in different services. Aiming to maximize the MU perceived QoE, we formulate the multiple bitrate video caching problem, and prove this problem is NP-hard for any given positive and strictly increasing QoE function. We then propose a polynomial complexity algorithm based on a general QoE function, which can achieve an approximate ratio arbitrarily close to 1/2. Specifically, for a linear QoE function, we explore useful property of optimal solutions, based on which more efficient algorithms are proposed. We demonstrate the effectiveness of our solutions via both theoretical analysis and extensive simulations. Zhihao Qu, Bin Tang 0002, Song Guo 0001, Sanglu Lu, Weihua Zhuang |
IEEE Trans. Mob. Comput. | 1 |
| 2019 | Multi-Path Routing Oriented Flow Statistics Collection in Software Defined NetworksabstractIn Software Defined Networks (SDNs), one of the key tasks in control plane is the monitoring and measurement of the whole network. A typical SDN consists of a set of switches and a logically centralized controller responsible for network state monitoring and flow scheduling, to which low cost and efficient flow statistics collection plays an important role. However, existing flow statistics collection methods mainly focus on single-path routing and cannot accurately capture the flow statistics in the case of multi-path routing (MPR), which is widely used in modern networks. In this paper, we are motivated to propose a Multi-Path oriented Flow Statistics Collection (MFSC) strategy to minimize the total communication cost for flow statistics collection. The problem is first formulated into an integer linear programming (ILP) form. After analyzing the complexity of this problem, we present a relaxation-based algorithm with an approximation factor p, where p is the maximum number of switches passed by each flow. The experiment results demonstrate that our proposed algorithm can reduce the total communication cost by over 14% compared with the traditional solutions. Jie Zhang 0076, Song Guo 0001, Deze Zeng, Zhihao Qu |
ICPADS | 4 |
| 2015 | Fast Cooperative Content Distribution over Hybrid Wireless NetworksabstractRecently, device-to-device (D2D) communications have been leveraged to offload the traffic on cellular networks. In this paper, we focus on the content distribution problem over a hybrid cellular and local D2D network where mobile devices cooperatively download a same content. We assume that the transmissions over D2D communications are scheduled in a centralized fashion so as to achieve a high transmission efficiency. Aiming at minimizing the content download time, we formulate the minimum content download time (MinCD) problem as an integer linear programming problem, and show its hardness and inapproximability. For the case that only a single channel is available for D2D communications, we propose an asymptotically optimal algorithm. For the multi-channel case, we further propose a heuristic algorithm with low time complexity and demonstrate its efficiency via extensive simulations. Zhihao Qu, Bin Tang 0002, Sanglu Lu |
GLOBECOM | 1 |
| 2015 | Energy-Aware Cost-Effective Cooperative Mobile Streaming on Smartphones over Hybrid Wireless NetworksabstractThe ever-increasing demands on mobile streaming over smartphones make the cellular networks always occupied by heavy load under traditional base-station-to-device (B2D) based streaming architecture, and even degrade the quality of service (QoS) seriously. To offload the traffic of cellular networks and provide scalable mobile streaming services with guaranteed QoS, in this paper we propose a device-to-device (D2D) communication motivated cooperative streaming framework by exploiting the capacity of both WiFi interface and cellular interface equipped with smartphones. Specifically, under the energy constraint of individual smartphone, we develop technique to minimize the over traffic of the cellular network by efficiently disseminating video over the D2D network with multi-hop routing supported. We formulate such an energy-aware cost-effective video dissemination problem as an integer linear programming problem, and show it to be NP-hard and even hard to approximate. We further present an energy allocation based algorithm and a simulated annealing heuristic algorithm which provide a trade-off between the performance and complexity to support the dissemination scheduling of cooperative mobile streaming. We evaluate the performance effectiveness of our proposal via both theoretical analysis and extensive simulation. Zhihao Qu, Bin Tang 0002, Sanglu Lu, Song Guo 0001 |
ICPP | 1 |