VLDB 2026 Research / reviewers in the wild / expert
Zhiqing Tang
dblp:231/6044
· DBLP profile ↗
53ranked-venue papers
8as first author
48since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 32 · 5 first-author · 29 since 2021Software engineering, systems software and programming languages · 10 · 1 first-author · 9 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RAP: Resource-Adaptive Planning for Efficient LLM Tool Calling in Edge Computing
Zhiqing Tang, Jianxiong Guo, Jiong Lou, Tian Wang 0001, Weijia Jia 0001 |
INFOCOM | 2 |
| 2026 | Adaptive Two-timescale Joint Service Placement and Request Scheduling for Efficient Edge AIGC
Changfu Xu, Xiao Mao, Zhiqing Tang, Haodong Zou, Yuzhu Liang |
INFOCOM | 4 |
| 2026 | Accelerating Diffusion Model Inference via Semantics-Aware Trajectory Reuse and Adaptive Scheduling
Hanshuai Cui, Zhiqing Tang, Zhi Yao, Weijia Jia 0001 |
NOSSDAV | 2 |
| 2026 | JUST: Cost-efficient joint cluster upgrade sequencing and task scheduling for containerized edge computing
Zhiqing Tang, Wenmian Yang, Jianxiong Guo, Tian Wang 0001, Weijia Jia 0001 |
Comput. Networks | 3 |
| 2026 | Physical Layer Security of Coupled Phase Shifts STAR-RIS-Aided NOMA System Under Hybrid Far- and Near-Field ScenariosabstractNear-field (NF) communications have attracted considerable interest, particularly with the implementation of extremely large-scale antenna arrays (ELAA). Additionally, the increase in communication frequencies and the expansion of reconfigurable intelligent surface (RIS) apertures contribute to this growing field. This paper investigates the synergy of simultaneously transmitting and reflecting (STAR)-RIS and non-orthogonal multiple access (NOMA) for secure transmission under hybrid far-field (FF) and NF scenarios. The secrecy sum rate (SSR) maximization problem is formulated by joint optimization of the power allocation, the beamforming at access point (AP), and the transmission/reflection coefficients (TRCs). Specifically, we consider the transmit power budget, unit-norm conditions, coupled phase shifts (CPS), quality of service requirements, and decoding order. To tackle this extremely challenging problem, we combine the successive convex approximation (SCA), Riemannian exact penalty method via smoothing, and penalty dual decomposition (PDD) and successfully develop an efficient iterative algorithm. Simulation results reveal that the proposed design exhibits superior effectiveness when compared to other traditional benchmarks. Lei Shi 0001, Zhiqing Tang, Lingfeng Shen, Wanming Hao, Jie Li 0002 |
IEEE Internet Things J. | 3 |
| 2026 | Stackelberg Game with Zero-Determinant Strategy for Incentive Mechanism Design in Socially Aware Mobile CrowdsensingabstractIn Mobile Crowdsensing (MCS), incentive mechanisms are crucial for encouraging mobile users to join tasks while users selfishly pursue personal benefit maximization. While most existing studies focus on the interaction between the requester and users, the internal value of socially aware user relationships remains underexplored. Users naturally form social connections, assisting or collaborating on tasks, but current mechanisms often neglect asymmetric social effects, which can lead to unequal willingness to cooperate and eventual breakdowns in collaboration (e.g., less profitable users refusing to cooperate). To end this, we propose an integrated incentive mechanism that models the interaction between the requester and users as a two-stage Stackelberg Game (SG) while accounting for pairwise asymmetric social effects. Pairwise cooperation is governed by the Iterated Prisoner’s Dilemma (IPD), with users employing Zero-Determinant (ZD) strategies to ensure cooperation despite unequal payoffs. Additionally, a plug-and-play sub-algorithm is introduced to filter low-quality or malicious users simultaneously and evaluate task redundancy, enhancing system robustness. We rigorously prove the existence of the Nash equilibrium, design an efficient iterative algorithm for our proposed mechanism, and validate its effectiveness through extensive experiments on real-world social datasets, which demonstrate that our method significantly improves system utility and cooperation stability while ensuring quality of service requirements. Gailun Zeng, Jianxiong Guo, Chuanwen Luo, Zhiqing Tang, Tian Wang 0001, Weijia Jia 0001 |
ACM Trans. Knowl. Discov. Data | 4 |
| 2026 | Efficient Layer-Granularity Unloading for LLMs in Edge ComputingabstractAdvancements in edge computing and container technology have made it increasingly popular and convenient to deploy Large Language Models (LLMs) through containers at the edge. However, the limited GPU resources of edge servers make it impractical to retain the model in GPU memory for long periods due to the high memory cost, especially when they remain idle without user requests. Existing work unloads the entire idle models to reduce memory costs on edge servers, but reloading them introduces significant loading delays that affect task Quality of Service (QoS). Therefore, efficient management of idle models is a critical issue that has been largely neglected in existing research and requires urgent attention. To address this gap, this paper studies the problem of idle model management from the perspective of the trade-off between memory cost and loading delay under the QoS constraint. A novel layer-granularity model unloading method is proposed, which leverages the layered characteristics of the model. We formulate an online joint optimization problem to determine which layers to unload and when, and present a layer-granularity unloading strategy inspired by the ski rental problem to solve it. We implement a real system with layer-granularity unloading for LLMs on NVIDIA GPUs and validate the effectiveness of the proposed method. Experimental results show it effectively trades off memory cost and loading delay, improving overall performance by up to 39.6%. Zhenzheng Li, Zhiqing Tang, Jianxiong Guo, Weijia Jia 0001, Wei Zhao 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2026 | Hybrid Learning for Cold-Start-Aware Microservice Scheduling in Dynamic Edge EnvironmentsabstractWith the rapid growth of IoT devices and their diverse workloads, container-based microservices deployed at edge nodes have emerged as a lightweight, scalable solution. However, existing microservice scheduling algorithms often assume static resource availability, which is unrealistic when multiple containers are assigned to an edge node. Besides, containers suffer from cold-start inefficiencies during early-stage training in currently popular reinforcement learning (RL) algorithms. In this paper, we propose a hybrid learning framework that combines offline imitation learning (IL) with online Soft Actor-Critic (SAC) optimization to enable cold-start-aware microservice scheduling with dynamic resource allocation. We first formulate a delay-and-energy-aware scheduling problem and construct a rule-based expert to generate demonstration data for behavior cloning. Then, a GRU-enhanced policy network is designed within the policy network to extract correlations among multiple decisions by separately encoding slow-evolving node states and fast-changing microservice features, and an action selection mechanism is provided to speed up convergence. Extensive experiments show that our method significantly accelerates convergence and achieves superior final performance. Compared with baselines, our algorithm improves the total objective by 50% and convergence speed by 70%, and demonstrates the highest stability and robustness across various edge configurations. Jingxi Lu, Jianxiong Guo, Xingjian Ding, Zhiqing Tang, Tian Wang 0001, Weijia Jia 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2026 | EAT: QoS-Aware Edge-Collaborative AIGC Task Scheduling via Attention-Guided Diffusion Reinforcement Learning
Zhiqing Tang, Jiong Lou, Zhi Yao, Tian Wang 0001, Yinglong Wang 0001, Weijia Jia 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2026 | Smart Server Selection: Enhancing QoE Through a Budget-Aware Bandit in Meta Computing
Yandi Li, Jianxiong Guo, Yupeng Li 0001, Zhiqing Tang, Xingjian Ding, Tian Wang 0001, Weijia Jia 0001 |
IEEE Trans. Netw. | 4 |
| 2026 | Adaptive Request Scheduling and Load Balancing for Edge Deployed Large Language ModelsabstractThe inference services of Large Language Models (LLMs) play a crucial role in many fields. Deploying LLMs at the network edge can effectively reduce response latency and enhance domain-specific knowledge. However, due to the dynamic and unpredictable nature of edge user requests and the limited resources of edge servers, improper request scheduling can lead to server load imbalance and increased inference latency. To address this challenge, we model edge LLM inference under dynamic workloads and constrained edge resources by fully considering mutual influences and trade-offs between inference performance, inference latency, and edge resource utilization. We propose a Workload Prediction and Dynamic Request Scheduling (WPDS) algorithm for edge computing environments. The WPDS algorithm first evaluates and prioritizes heterogeneous requests by extracting features and importance scores from user requests. A Transformer-based encoder is used to extract load characteristics of edge servers over different time periods to predict future GPU resource demands. Then, a soft actor-critic reinforcement learning model dynamically schedules requests to appropriate LLM instances, capturing the dynamic relationship between request processing and edge resource utilization from complex state spaces. We evaluate our algorithm in a real Kubernetes-based edge inference prototype system. Experimental results indicate that our approach reduces average inference time by approximately 16.6% compared to the best-performing heuristic baseline algorithm, while also achieving more balanced GPU utilization across edge servers. Fangyi Mou, Zhiqing Tang, Weijia Jia 0001, Wei Zhao 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2026 | Adapting Multi-Model Inference Pipelines With Diffusion-Based Reinforcement Learning in Edge Computing
Jinhao Sheng, Zhiqing Tang, Jianxiong Guo, Kun Yue, Tian Wang 0001, Weijia Jia 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2026 | Reverse-Offloading Incentive Mechanism With Long-Term Optimization in Social-Aware Cloud-Edge Systems
Gailun Zeng, Jianxiong Guo, Xingjian Ding, Zhiqing Tang, Tian Wang 0001, Weijia Jia 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2026 | Latency Minimization for IRS-Enhanced Wideband MIMO-OFDM MEC Networks With Practical Reflection ModelabstractIntelligent reflecting surface (IRS) has been considered a promising technology to be applied to mobile edge computing (MEC) systems, especially when offloading links are blocked or weak. However, most existing works are restricted to narrow-band channel and ideal IRS reflection model, which is not practical and may lead to significant performance degradation. Thus, we consider an IRS-enhanced wideband MEC system with practical IRS reflection model. Our objective is to minimize the weighted latency of all devices by jointly optimizing the offloading data volume, edge computing resources, BS receiving vector, and IRS basic phase shift (BPS). Since the formulated problem is non-convex, we employ the block coordinate descent (BCD) technique to decouple it into two subproblems to alternatively optimize computing and communication resources. In particular, the computing resource optimization subproblem is solved based on Karush-Kuhn-Tucker (KKT) conditions and bisection search method. While the communication resource optimization subproblem is first transformed into a weighted sum-rate maximization problem based on LDR technique and KKT conditions. Then leveraging the equivalence between sum-rate maximization and MSE minimization, it is converted into a multi-variable problem that can be effectively solved using BCD technique. Simulation results show that the proposed schemes can reduce latency by 16% compared to baseline schemes when the number of IRS elements is 100, confirming the effectiveness of considering practical IRS reflection model for wideband MEC systems. Nana Li 0001, Wanming Hao, Xingwang Li 0001, Zhengyu Zhu 0001, Zhiqing Tang, Shouyi Yang |
IEEE Trans. Wirel. Commun. | 5 |
| 2025 | POSFed: Tackling Non-IID Challenges in One-Shot Federated Learning via PersonalizationabstractFederated Learning (FL) enables collaborative model training across distributed clients without requiring the exchange of raw data. However, existing One-Shot FL (OSFL) methods, designed for communication efficiency by reducing fed-erated rounds to one, suffer substantial performance degradation when faced with highly non-IID data across clients, primarily due to critical distribution shifts: label shift, feature shift, and concept shift. In this paper, we introduce POSFed, a new personalized three-stage approach, to systematically address these fundamen-tal limitations: (1) Each client locally generates robust and label-agnostic synthetic datasets via self-supervising learning, ensuring essential knowledge is captured despite local distribution shifts; (2) The server aggregates all synthetic datasets to train a global feature extractor, capturing generalizable and transferable rep-resentations across heterogeneous client data; and (3) Each client efficiently adapts the feature extractor by learning a personalized classification head on its own data, enabling effective local customization and mitigating both feature and concept shifts. Extensive experiments across multiple benchmarks demonstrate that POSFed significantly outperforms state-of-the-art methods, achieving performance comparable to multi-round personalized approaches while using only one communication round. By ensuring both superior personalization and practical communication efficiency, POSFed establishes a feasible paradigm for FL under more realistic and heterogeneous conditions. Code is available at https://github.com/I643204431IPOSFed. Xuanzhe Xiao, Jianxiong Guo, Zhiqing Tang, Qiufen Ni, Weili Wu 0001 |
ICDM | 4 |
| 2025 | Enhancing Federated Learning in IoT Through Dynamic Spectral Clustering
Changyu Zheng, Junli Gong, Jianxiong Guo, Zhiqing Tang, Tian Wang 0001 |
WASA (2) | 4 |
| 2025 | LR2Scheduler: layer-aware, resource-balanced, and request-adaptive container scheduling for edge computing
Wentao Peng, Zhiqing Tang, Jianxiong Guo, Jiong Lou, Tian Wang 0001, Weijia Jia 0001 |
CCF Trans. Pervasive Comput. Interact. | 2 |
| 2025 | Transmit Antenna Selection and Power Allocation Optimization for Non-Orthogonal Multiple Access Systems with Statistical Channel State InformationabstractABSTRACT This paper considers a downlink multiple input single output (MISO) non‐orthogonal multiple access (NOMA) system over Nakagami‐m fading channels, where a multi‐antenna base station (BS) serves several single‐antenna users with the statistical channel state information (CSI) of each user. We propose a novel low‐complexity transmit antenna selection by head user (TAS‐head) strategy for the first time to exploit the spatial diversity of multiple antennas. Based on our proposed TAS‐head strategy, we derive a closed‐form expression of the exact outage probability (OP). We further analyse the asymptotic OP and diversity order in high signal‐to‐noise ratio (SNR) regime. Finally, we formulate a power allocation optimization problem to maximize sum throughput under outage constraints. We also design an Adam algorithm in combination with numerical differentiation method to obtain a suboptimal solution. Monte Carlo (MC) simulations verify the accuracy of our derived exact OP. Results show that our proposed TAS‐head strategy is more effective than its benchmarks (TAS‐near/far and TAS‐maj). Furthermore, we prove that PA‐TDR criterion achieves better performance than PA‐ACG in scenarios where the descending order of target data rate is the same with that of channel condition. Our designed Adam algorithm turns out to be more effective in comparison with genetic algorithm (GA) in multi‐user case. Results indicate that our proposed TAS‐head strategy is an efficient method to meet users' QoS requirements, especially in low SNR (or transmit power) regime. Zhuo Han, Wanming Hao, Shouyi Yang, Zhiqing Tang |
IET Commun. | 4 |
| 2025 | On container vulnerabilities in edge computing: A fix-on-deployment approach
Tianhui Meng, Jianxiong Guo, Zhiqing Tang, Weijia Jia 0001 |
J. Syst. Archit. | 4 |
| 2025 | Look Closer to Your Enemy: Learning to Attack via Teacher-Student MimickingabstractDeep neural networks have significantly advanced person re-identification (ReID) applications in the realm of the industrial internet, yet they remain vulnerable. Thus, it is crucial to study the robustness of ReID systems, as there are risks of adversaries using these vulnerabilities to compromise industrial surveillance systems. Current adversarial methods focus on generating attack samples using misclassification feedback from victim models (VMs), neglecting VM's cognitive processes. We seek to address this by producing authentic ReID attack instances through VM cognition decryption. This approach boasts advantages like better transferability to open-set ReID tests, easier VM misdirection, and enhanced creation of realistic and undetectable assault images. However, the task of deciphering the cognitive mechanism in VM is widely considered to be a formidable challenge. In this paper, we propose a novel inconspicuous and controllable ReID attack baseline, LCYE (LookCloser toYourEnemy), to generate adversarial query images. Specifically, LCYE first distills VM's knowledge via teacher-student memory mimicking the proxy task. This knowledge prior serves as an unambiguous cryptographic token, encapsulating elements deemed indispensable and plausible by the VM, with the intent of facilitating precise adversarial misdirection. Further, benefiting from the multiple opposing task framework of LCYE, we investigate the interpretability and generalization of ReID models from the view of the adversarial attack, including cross-domain adaption, cross-model consensus, and online learning process. Extensive experiments on four ReID benchmarks show that our method outperforms other state-of-the-art attackers with a large margin in white-box, black-box, and target attacks. The source code can be found athttps://github.com/MingjieWang0606/LCYE-attack_reid. Mingjie Wang 0001, Jianxiong Guo, Dingwen Xiao, Zhiqing Tang |
IEEE Trans. Big Data | 5 |
| 2025 | Layer-Aware Cost-Effective Container Updates With Edge-Cloud Collaboration in Edge ComputingabstractContainers have become popular for deploying applications in Edge Computing (EC) for their seamless integration and easy deployment. Frequent container updates are essential to enhance performance and introduce new challenges for cutting-edge applications such as large language models and digital twins. However, traditional container update methods result in substantial download costs and task interruptions, which are unacceptable for latency-sensitive tasks in resource-constrained EC. Existing work has largely overlooked the layered structure of container images. By leveraging this layered structure, duplicate downloads can be reduced, and various layers can be transferred from other edges, reducing burden on the remote cloud. In this paper, we model the layer-aware container update problem with edge-cloud collaboration to minimize update and scheduling costs. We present the Layer-aware Edge-cloud collaborative Container Update (LECU) algorithm based on reinforcement learning to make container update decisions. Moreover, a task scheduling algorithm is devised to schedule tasks affected by container updates to other edges, minimizing the impact of task interruptions. We implement our LECU algorithm on an edge system with real-world data traces to demonstrate its effectiveness and conduct larger-scale simulations to evaluate its scalability. Results demonstrate that our algorithms reduce container update and task scheduling costs by 14% and 19%, respectively, compared to baselines. Hanshuai Cui, Zhiqing Tang, Yuan Wu 0001, Weijia Jia 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | Two-Stage Deep Energy Optimization in IRS-Assisted UAV-Based Edge Computing SystemsabstractIntegrating wireless-powered Mobile Edge Computing (MEC) with Unmanned Aerial Vehicles (UAVs) leverages computation offloading services for mobile devices, significantly enhancing the mobility and control of MEC networks. However, current research has not focused on customizing system designs for Terahertz (THz) communication networks. When dealing with THz communication, one must account for blockage vulnerability due to severe THz wave propagation attenuation and insufficient diffraction. The Intelligent Reflecting Surface (IRS) can effectively address these limitations in the model, enhancing spectrum efficiency and coverage capabilities while reducing blockage vulnerability in THz networks. In this paper, we introduce an upgraded MEC system that integrates IRS and UAVs into THz communication networks, focusing on a binary offloading policy for studying the computation offloading problem. Our primary objective is to optimize the energy consumption of both UAVs and User Electronic Devices, alongside refining the phase shift of the IRS reflector. The problem is a Mixed Integer Non-Linear Programming problem known as NP-hard. To tackle this challenge, we propose a two-stage deep learning-based optimization framework named Iterative Order-Preserving Policy Optimization (IOPO). Unlike exhaustive search methods, IOPO continually updates offloading decisions through an order-preserving quantization method, thereby accelerating convergence and reducing computational complexity, especially when handling complex problems with extensive solution spaces. The numerical results demonstrate that the proposed algorithm significantly improves energy efficiency and achieves near-optimal performance compared to benchmark methods. Jianqiu Wu, Zhongyi Yu, Jianxiong Guo, Zhiqing Tang, Tian Wang 0001, Weijia Jia 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Blockchain-Enabled Multiple Sensitive Task-Offloading Mechanism for MEC ApplicationsabstractAs mobile devices proliferate and mobile applications diversify, Mobile Edge Computing (MEC) has become widely adopted to efficiently allocate computing resources at the network edge and alleviate network congestion. In the MEC initial phase, the absence of vital information presents challenges in devising task-offloading policies, and identifying malicious devices responsible for providing inaccurate feedback is complex. To fill in such gaps, we introduce a consortium blockchain-enabledCommitteeVoting basedTaskOffloadingModel (CVTOM) to collaboratively formulate resource allocation policies and establish deterrence against malicious servers producing erroneous results intentionally. Different voting principle mechanisms of each committee member are first designed in a Blockchain-enabled system which helps to represent the system's resource status. Additionally, we propose a Multi-armed Bandits relatedThompsonSampling basedAdaptivePreferenceOptimization (TSAPO) algorithm for task-offloading policy, enhancing the timely identification of potent edge servers to improve computing resource utilization which first considers dynamic edge server space and parallel computing scenarios. The solid proof process greatly contributes to the theoretical analysis of the TSAPO. The simulation experiments demonstrate the delay and budget can be reduced by around 25% and 10% respectively, showcasing the superior performance of our approach. Yang Xu 0013, Hangfan Li, Cheng Zhang 0035, Zhiqing Tang, Xiaoxiong Zhong, Ju Ren 0001, Hongbo Jiang 0001, Yaoxue Zhang |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Online Worker Scheduling for Maximizing Long-Term Utility in Crowdsourcing with Unknown QualityabstractSpatiotemporal Mobile CrowdSourcing (MCS) is a new intelligent sensing paradigm for large-scale data acquisition where requesters can recruit a crowd of workers to perform data collection tasks. How to recruit suitable workers in a dynamic environment to maximize platform utility is a key issue and has become a research hotspot. Many past studies have made great efforts in this regard, but most of them either assume that the worker quality is known in advance or ignore the limitations of workers’ short-term ability to provide resources. In this article, we consider a platform-centered online spatiotemporal MCS system where mobile workers have both long-term and short-term constraints for providing resources, and their quality is unknown to the platform, while the platform has a long-term budget constraint for recruiting workers. We aim to find an online worker scheduling scheme to maximize the platform’s long-term utility without violating the constraints of both workers and the platform. To address this problem, we first transform the long-term utility maximization problem into a real-time utility maximization problem by leveraging the Lyapunov optimization, then design algorithms based on the Upper Confidence Bound (UCB) and Markov approximation to solve each real-time utility maximization problem with unknown worker quality. We demonstrate that our UCB-based algorithm has a sublinear regret and prove that our proposed framework has a performance guarantee for the addressed problem. Finally, we evaluate our design through numerical simulation experiments, and the results demonstrate the effectiveness of our algorithm. Pengfei Lin 0001, Xingjian Ding, Jianxiong Guo, Zhiqing Tang, Deying Li 0001, Weili Wu 0001 |
ACM Trans. Internet Techn. | 5 |
| 2025 | Optimizing Communication Efficiency through Training Potential in Multi-Modal Federated LearningabstractMulti-modal Federated Learning (FL) is a type of FL that considers utilizing multiple modalities of data to improve overall performance. While multi-modal data brings richer information, it also introduces more significant communication overhead. Reducing this overhead hinges on two key strategies: increasing the convergence speed of the training or reducing the communication overhead in each communication round. However, few studies have considered these two strategies simultaneously and formed a unified optimization framework. Thus, we propose a joint client and modality selection framework to reduce communication overhead. Modality selection executed on each client assigns weights to modalities based on their contribution to training potential, aiming at accelerating the convergence. Client selection executed on the server assigns weights to clients by considering different metrics, especially total training potential after the modality selection. We validate our proposed method on the five widely used open-source datasets, achieving satisfactory accuracy while reducing the total communication overhead to 2.43%–14.24% compared to without selection on different datasets, significantly outperforming existing state-of-the-art (SOTA) methods. Code is available at https://github.com/1643204431/OCETPMMFL . Jianxiong Guo, Xingjian Ding, Zhiqing Tang, Tian Wang 0001, Weili Wu 0001, Weijia Jia 0001 |
ACM Trans. Internet Techn. | 4 |
| 2025 | Cloud-Edge System for Scheduling Unpredictable LLM Requests With Combinatorial BanditabstractThe rapid growth in demand for large language models (LLMs) has strained cloud-edge infrastructure. While edges offer low latency and clouds provide vast resources, scheduling LLM requests efficiently remains a major challenge due to their unpredictable processing times, which leads to Headof-Line (HOL) blocking that degrades system throughput and responsiveness. To address this, we introduce the Online CloudEdge Collaborative Request Scheduling (OCE-CRS) framework. OCE-CRS models the proactive scheduling of LLM requests as a contextual combinatorial bandit problem. At its core is our novel Combinatorial Neural Delayed Upper Confidence Bound (CN DUCB) algorithm, which learns to predict request processing times from the semantic content of the request prompt alone. This enables an inspired policy based on Shortest Job First (SJF) that prioritizes shorter jobs for edge execution, simultaneously maximizing throughput and mitigating HOL blocking. To prevent time-consuming neural network training from blocking scheduling decisions, we employ an asynchronous mechanism. This decouples model updates from the real-time scheduling loop, effectively handling the resultant delayed feedback where observations from past rounds are used in later training steps. We provide a theoretical sublinear regret bound for our algorithm. Extensive experiments validate that OCE-CRS significantly improves throughput, Job Completion Time (JCT), and queueing delay, demonstrating superior performance and robustness in both static and continuous batching environments. Yandi Li, Jianxiong Guo, Zhiqing Tang, Xingjian Ding, Juncheng Wang 0001, Tian Wang 0001, Weijia Jia 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2025 | Online Layer-Aware Joint Request Scheduling, Container Placement, and Resource Provision in Edge ComputingabstractContainers have emerged as a pivotal tool for service deployment in edge computing. Before running the container, an image composed of several layers must exist locally. Recent strategies have utilized layer-sharing in images to reduce deployment delays. However, existing research only focuses on a single aspect of container orchestration, like container placement, neglecting the joint optimization of the entire orchestration process. To fill in such gaps, this article introduces an online strategy that considers layer-aware container orchestration, encompassing request scheduling, container placement, and resource provision. The goal is to reduce costs, adapt to evolving user demands, and adhere to system constraints. We present an online optimization problem that accounts for various real-world factors in orchestration, including container and server expenses. An online algorithm is proposed, integrating a regularization-based approach and stepwise rounding to address this optimization problem efficiently. The regularization approach separates time-dependent container placement and server wake-up costs, requiring only current information and past decisions. The stepwise rounding process generates feasible solutions that meet system constraints, reducing computational costs. Additionally, a competitive ratio proof is provided for the proposed algorithm. Extensive evaluations demonstrate that our approach achieves about 20% performance enhancement compared to baseline algorithms. Zhenzheng Li, Jiong Lou, Zhiqing Tang, Jianxiong Guo, Tian Wang 0001, Weijia Jia 0001, Wei Zhao 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2025 | Enhancing LLM QoS Through Cloud-Edge Collaboration: A Diffusion-Based Multi-Agent Reinforcement Learning ApproachabstractLarge Language Models (LLMs) are widely used across various domains, but deploying them in cloud data centers often leads to significant response delays and high costs, undermining Quality of Service (QoS) at the network edge. Although caching LLM request results at the edge using vector databases can greatly reduce response times and costs for similar requests, this approach has been overlooked in prior research. To address this, we propose a novelVector database-assisted cloud-Edge collaborativeLLM QoSOptimization (VELO) framework that caches LLM request results at the edge using vector databases, thereby reducing response times for subsequent similar requests. Unlike methods that modify LLMs directly, VELO leaves the LLM's internal structure intact and is applicable to various LLMs. Building on VELO, we formulate the QoS optimization problem as a Markov Decision Process (MDP) and design an algorithm based on Multi-Agent Reinforcement Learning (MARL). Our algorithm employs a diffusion-based policy network to extract the LLM request features, determining whether to request the LLM in the cloud or retrieve results from the edge's vector database. Implemented in a real edge system, our experimental results demonstrate that VELO significantly enhances user satisfaction by simultaneously reducing delays and resource consumption for edge users of LLMs. Our DLRS algorithm improves performance by 15.0% on average for similar requests and by 14.6% for new requests compared to the baselines. Zhi Yao, Zhiqing Tang, Wenmian Yang, Weijia Jia 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2024 | Adaptive Configuration Selection for Multi-Model Inference Pipelines in Edge ComputingabstractThe growing demand for real-time processing tasks is driving the need for multi-model inference pipelines on edge devices. However, cost-effectively deploying these pipelines while optimizing Quality of Service (QoS) and costs poses significant challenges. Existing solutions often neglect device resource constraints, focusing mainly on inference accuracy and cost efficiency. To address this, we develop a framework for configuring multi-model inference pipelines. Specifically: 1) We model the decision-making problem by considering the pipeline’s QoS, costs, and device resource limitations. 2) We create a feature extraction module using residual networks and a load prediction model based on Long Short-Term Memory (LSTM) to gather comprehensive node and pipeline status information. Then, we implement a Reinforcement Learning (RL) algorithm based on policy gradients for online configuration decisions. 3) Experiments conducted in a real Kubernetes cluster show that our approach significantly improve QoS while reducing costs and shorten decision-making time for complex pipelines compared to baseline algorithms. Jinhao Sheng, Zhiqing Tang, Jianxiong Guo, Tian Wang 0001 |
HPCC | 2 |
| 2024 | Dynamic Offloading Control for Waste Sorting Based on Deep Q-Network
Jianxiong Guo, Zhiqing Tang, Xingjian Ding, Tian Wang 0001 |
ICA3PP (2) | 4 |
| 2024 | Efficient Serverless Function Scheduling in Edge ComputingabstractServerless computing is a promising approach for edge computing since its inherent features, e.g., lightweight virtualization, rapid scalability, and economic efficiency. However, there are two challenges existing in serverless edge computing: significant cold start latency and request blocking. Previous studies have not successfully resolved these challenges, which affect the Quality of Experience. In this paper, we formulate the Serverless Function Scheduling (SFS) problem in resource-limited edge computing, aiming to minimize the average response time. To solve this intractable scheduling problem, we first consider a simplified offline form of the SFS problem and design a polynomial-time optimal scheduling algorithm. Inspired by this optimal algorithm, we propose an Enhanced Shortest Function First (ESFF) algorithm, including function creation and function replacement. To avoid frequent cold starts, ESFF selectively decides the initialization of new function instances when receiving requests. To deal with request blocking, ESFF judiciously replaces serverless functions based on the function weight at the completion time of requests. Extensive simulations based on real-world serverless request traces are conducted, and the results show that ESFF consistently and substantially outperforms existing baselines under different settings. Jiong Lou, Zhiqing Tang, Shijing Yuan, Jie Li 0002, Weijia Jia 0001, Chentao Wu |
ICC | 2 |
| 2024 | VELO: A Vector Database-Assisted Cloud-Edge Collaborative LLM QoS Optimization FrameworkabstractThe Large Language Model (LLM) has gained significant popularity and is extensively utilized across various domains. Most LLM deployments occur within cloud data centers, where they encounter substantial response delays and incur high costs, thereby impacting the Quality of Services (QoS) at the network edge. Leveraging vector database caching to store LLM request results at the edge can substantially mitigate response delays and cost associated with similar requests, which has been overlooked by previous research. Addressing these gaps, this paper introduces a novel Vector database-assisted cloud-Edge collaborative LLM QoS Optimization (VELO) framework. Firstly, we propose the VELO framework, which ingeniously employs vector database to cache the results of some LLM requests at the edge to reduce the response time of subsequent similar requests. Diverging from direct optimization of the LLM, our VELO framework does not necessitate altering the internal structure of LLM and is broadly applicable to diverse LLMs. Subsequently, building upon the VELO framework, we formulate the QoS optimization problem as a Markov Decision Process (MDP) and devise an algorithm grounded in Multi-Agent Reinforcement Learning (MARL) to decide whether to request the LLM in the cloud or directly return the results from the vector database at the edge. Moreover, to enhance request feature extraction and expedite training, we refine the policy network of MARL and integrate expert demonstrations. Finally, we implement the proposed algorithm within a real edge system. Experimental findings confirm that our VELO framework substantially enhances user satisfaction by concurrently diminishing delay and resource consumption for edge users utilizing LLMs. Zhi Yao, Zhiqing Tang, Jiong Lou, Ping Shen, Weijia Jia 0001 |
ICWS | 2 |
| 2024 | Container Scheduling with Dynamic Computing Resource for Microservice Deployment in Edge ComputingabstractWith the massive increase of Internet of Things devices and their data, executing applications by using micro-service architecture has emerged as the predominant trend. As the container technology emerges, microservices can be lightweightly deployed in resource-constrained edge nodes. However, existing container scheduling algorithms often overlook the allocation of computing resources on edge servers. When multiple containers are assigned to an edge node, it is usually assumed that they share a CPU frequency, which is obviously unrealistic. In this paper, we first formulate an online container-based microservice scheduling problem with dynamic computing power to minimize the total delay and energy consumption, where we need to determine the assignment between microservices and edge nodes and the allocation of computing power to each microservice in an edge node. Then, we propose a Soft Actor-Critic (SAC) based reinforcement learning algorithm to address this problem, where a GRU unit is designed in the policy network to extract the correlation among multiple decisions, and an action selection mechanism is given to speed up the convergence. Finally, a simulated scheduling system is implemented to validate our algorithm, which demonstrates that our algorithm outperforms the commonly used baselines by up to 65% in terms of the total objective on average. Jingxi Lu, Jianxiong Guo, Xingjian Ding, Zhiqing Tang, Tian Wang 0001 |
MSN | 5 |
| 2024 | LRScheduler: A Layer-aware and Resource-adaptive Container Scheduler in Edge ComputingabstractLightweight containers provide an efficient approach for deploying computation-intensive applications in net-work edge. The layered storage structure of container images can further reduce the deployment cost and container startup time. Existing researches discuss layer sharing scheduling theoretically but with little attention paid to the practical implementation. To fill in this gap, we propose and implement a Layer-aware and Resource-adaptive container Scheduler (LRScheduler) in edge computing. Specifically, we first utilize container image layer information to design and implement a node scoring and container scheduling mechanism. This mechanism can effectively reduce the download cost when deploying containers, which is very important in edge computing with limited bandwidth. Then, we design a dynamically weighted and resource-adaptive mechanism to enhance load balancing in edge clusters, increasing layer sharing scores when resource load is low to use idle resources effectively. Our scheduler is built on the scheduling framework of Kubernetes, enabling full process automation from task information acquisition to container deployment. Testing on a real system has shown that our design can effectively reduce the container deployment cost as compared with the default scheduler. Zhiqing Tang, Wentao Peng, Jianxiong Guo, Jiong Lou, Hanshuai Cui, Tian Wang 0001, Yuan Wu 0001, Weijia Jia 0001 |
MSN | 1 |
| 2024 | QoS-Aware Energy-Efficient Multi-UAV Offloading Ratio and Trajectory Control Algorithm in Mobile-Edge ComputingabstractMultiple unmanned aerial vehicle (UAV)-assisted mobile-edge computing (MEC) leverages UAVs equipped with computational resources as mobile-edge servers, providing flexibility and low-latency connections, especially beneficial in smart cities and the Internet of Things (IoT). Maximizing Quality of Services (QoS) while minimizing energy consumption necessitates developing a suitable offloading ratio and trajectory control algorithm for UAVs. However, existing research on UAV control algorithms overlooks significant challenges like the heterogeneity of user equipments (UEs) and offloading failures. Furthermore, there is a dearth of experimental validation in large-scale UAV-assisted MEC scenarios. To bridge these gaps, we introduce a QoS-aware energy-efficient multi-UAV offloading ratio and trajectory control algorithm (QEMUOT). Specifically, 1) a composite UE mobility model is proposed to enhance system heterogeneous modeling, encompassing models for high-speed, low-speed, and fixed UEs; 2) QEMUOT is devised using multiagent reinforcement learning algorithms to determine offloading ratio and trajectory control decisions. To tackle sparse reward space and offloading failures, we employ expert demonstrations for pretraining and enhance reward mechanisms; and 3) experimental simulations illustrate that our algorithm outperforms baseline algorithms in user QoS with reduced energy consumption and demonstrates superior scalability in scenarios with numerous UAVs and UEs. Jiajie Yin, Zhiqing Tang, Jiong Lou, Jianxiong Guo, Tian Wang 0001, Weijia Jia 0001 |
IEEE Internet Things J. | 2 |
| 2024 | Online Container Scheduling With Fast Function Startup and Low Memory Cost in Edge ComputingabstractExtending serverless computing to the edge has emerged as a promising approach to support service, but startup containerized serverless functions lead to the cold-start delay. Recent research has introduced container caching methods to alleviate the cold-start delay, including cache as the entire container or the Zygote container. However, container caching incurs memory costs. The system must ensure fast function startup and low memory cost of edge servers, which has been overlooked in the literature. This paper aims to jointly optimize startup delay and memory cost. We formulate an online joint optimization problem that encompasses container scheduling decisions, including invocation distribution, container startup, and container caching. To solve the problem, we propose an online algorithm with a competitive ratio and low computational complexity. The proposed algorithm decomposes the problem into two subproblems and solves them sequentially. Each container is assigned a randomized strategy, and these container-level decisions are merged to constitute overall container caching decisions. Furthermore, a greedy-based subroutine is designed to solve the subproblem associated with invocation distribution and container startup decisions. Experiments on the real-world dataset indicate that the algorithm can reduce average startup delay by up to 23% and lower memory costs by up to 15%. Zhenzheng Li, Jiong Lou, Jianfei Wu, Jianxiong Guo, Zhiqing Tang, Ping Shen, Weijia Jia 0001, Wei Zhao 0001 |
IEEE Trans. Computers | 5 |
| 2024 | Startup-Aware Dependent Task Scheduling With Bandwidth Constraints in Edge ComputingabstractIn edge computing, applications can be scheduled in the granularity of inter-dependent tasks to proximate edge servers to achieve high performance. Before execution, the edge server must initialize the corresponding runtime environment, named task startup. However, existing studies on dependent task scheduling severely ignore bandwidth constraints during task startups, which is impractical and incurs a long startup latency. To fill in this gap, we first model the task startup process with bandwidth constraints on edge servers. Then, we formulate the dependent task scheduling problem with startup latency in heterogeneous edge computing. To efficiently generate schedules and satisfy the real-time requirements in edge computing, a novel low-complexity list scheduling algorithm integrated with cloud clone, Startup-aware Dependent Task Scheduling (SDTS), is proposed. Constrained by bandwidth and computation resources, SDTS first coordinates task startup, dependent data transmission, and task execution to optimize each task’s finish time. Then, a cloud clone for each task is deployed to utilize scalable resources and initialized runtime environments. Furthermore, task scheduling refinement is designed to release the bandwidth and computation resources consumed by redundant tasks and improve the schedule. Extensive simulations based on real-world datasets show that SDTS substantially reduces 30%-60% makespan compared with existing baselines. Jiong Lou, Zhiqing Tang, Weijia Jia 0001, Wei Zhao 0001, Jie Li 0002 |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Joint Resource Overbooking and Container Scheduling in Edge ComputingabstractContainers have gained popularity in Edge Computing (EC) networks due to their lightweight and flexible deployment advantage. In resource-constrained EC environments, overbooking container resources can substantially improve resource utilization. However, existing work overlooks the complex interplay between resource provisioning and container scheduling, which may result in performance degradation or inefficient resource utilization due to highly dynamic resource heterogeneity in EC. To address this issue, this paper presents a novel joint Resource Overbooking and Container Scheduling (ROCS) algorithm. Our approach accounts for resource heterogeneity and the geographical distribution of edge nodes, and we formulate the ROCS problem to consolidate various costs and revenues into a single profit metric for service providers. To enhance resource utilization and maximize the profit of the service providers, we develop an efficient algorithm that operates within a hybrid action space scheme by leveraging soft actor-critic reinforcement learning. Furthermore, we introduce a risk assessment mechanism to mitigate overbooking risks. Large-scale simulations with real-world data traces demonstrate the efficacy of our proposed ROCS algorithm, validating its advantage of improving resource utilization within EC networks. Zhiqing Tang, Fangyi Mou, Jiong Lou, Weijia Jia 0001, Yuan Wu 0001, Wei Zhao 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | Multi-User Layer-Aware Online Container Migration in Edge-Assisted Vehicular NetworksabstractIn edge-assisted vehicular networks, containers are very suitable for deploying applications and providing services due to their lightweight and rapid deployment. To provide high-quality services, many existing studies show that the containers need to be migrated to follow the vehicles’ trajectory. However, it has been conspicuously neglected by existing work that making full use of the complex layer-sharing information of containers among multiple users can significantly reduce migration latency. In this paper, we propose a novel online container migration algorithm to reduce the overall task latency. Specifically: 1) we model the multi-user layer-aware online container migration problem in edge-assisted vehicular networks, comprehensively considering the initialization latency, computation latency, and migration latency. 2) A feature extraction method based on attention and long short-term memory is proposed to fully extract the multi-user layer-sharing information. Then, a policy gradient-based reinforcement learning algorithm is proposed to make the online migration decisions. 3) The experiments are conducted with real-world data traces. Compared with the baselines, our algorithms effectively reduce the total latency by 8% to 30% on average. Zhiqing Tang, Fangyi Mou, Jiong Lou, Weijia Jia 0001, Yuan Wu 0001, Wei Zhao 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | Latency-Aware Container Scheduling in Edge Cluster Upgrades: A Deep Reinforcement Learning ApproachabstractIn Mobile Edge Computing (MEC), Internet of Things (IoT) devices offload computationally-intensive tasks to edge nodes, where they are executed within containers, reducing the reliance on centralized cloud infrastructure. Cluster software upgrades are essential to maintain the efficient and secure operation of edge clusters. However, traditional cloud cluster upgrade strategies are ill-suited for edge clusters due to their geographically distributed nature and resource limitations. Therefore, it is crucial to properly schedule containers during edge cluster upgrades to minimize the impact on running tasks. This article proposes a latency-aware container scheduling algorithm for efficient edge cluster upgrading. Specifically: 1) We formulate the online container scheduling problem for edge cluster upgrade to minimize the total task latency. 2) We propose a policy gradient-based reinforcement learning algorithm that addresses this problem by considering the characteristics of MEC, including heterogeneous resources, image distribution, and low-latency requirements. Subsequently, a location feature extraction method based on self-attention is designed to fully extract and utilize edge node distribution. 3) Experiments based on simulated and real-world data traces demonstrate that our algorithm reduces total task latency by approximately 30% compared to baseline algorithms. Hanshuai Cui, Zhiqing Tang, Jiong Lou, Weijia Jia 0001, Wei Zhao 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2023 | Joint Task Scheduling and Container Image Caching in Edge ComputingabstractIn Edge Computing (EC), containers have been increasingly used to deploy applications to provide mobile users services. Each container must run based on a container image file that exists locally. However, it has been conspicuously neglected by existing work that effective task scheduling combined with dynamic container image caching is a promising way to reduce the container image download time with the limited bandwidth resource of edge nodes. To fill in such gaps, in this paper, we propose novel joint Task Scheduling and Image Caching (TSIC) algorithms, specifically: 1) We consider the joint task scheduling and image caching problem and formulate it as a Markov Decision Process (MDP), taking the communication delay, waiting delay, and computation delay into consideration; 2) To solve the MDP problem, a TSIC algorithm based on deep reinforcement learning is proposed with the customized state and action spaces and combined with an adaptive caching update algorithm. 3) A real container system is implemented to validate our algorithms. The experiments show that our strategy outperforms the existing baseline approaches by 23% and 35% on average in terms of total delay and waiting delay, respectively. Fangyi Mou, Zhiqing Tang, Jiong Lou, Jianxiong Guo, Wenhua Wang 0003, Tian Wang 0001 |
MSN | 2 |
| 2023 | Pricing Model for Dynamic Resource Overbooking in Edge ComputingabstractEdge Computing (EC) with cloud-like Quality of Service (QoS) can find its wide applications in various resource-constrained smart cities where the resource requirements can be different during peak and off-peak periods. During off-peak periods, there are often many resources that have been requested but not used, which can be reused to obtain higher profit. However, to the best of our knowledge, there is no effective pricing model or overbooking mechanism in EC. To fill in this gap, a novel pricing model for dynamic resource overbooking is proposed in this paper, specifically: 1) To meet the needs of different users in EC, methods of on-demand, daily, auction, and the new spot billing are designed, in which resources can be overbooked. 2) An auction approach with pricing rule and winner determination rule is designed for auction billing, which is proved to guarantee individual rationality, computational efficiency, and truthfulness. 3) To make more use of the auction approach to utilize idle resources, a dynamic resource overbooking mechanism is introduced, including a cancellation policy and a resource prediction method. The mechanism is validated with real-world data-trace. Experimental results show that the dynamic resource overbooking mechanism maximizes the profit of edge nodes with a high QoS Satisfaction ratio of on-demand and daily billing. Zhiqing Tang, Fuming Zhang, Weijia Jia 0001, Wei Zhao 0001 |
IEEE Trans. Cloud Comput. | 1 |
| 2023 | Cost-Effective Scheduling for Dependent Tasks With Tight Deadline Constraints in Mobile Edge ComputingabstractIn Mobile Edge Computing (MEC), latency-sensitive mobile applications comprising dependent tasks can be scheduled to edge or cloud servers to reduce latency and execution costs. However, existing algorithms based on deadline distribution can hardly satisfy tight application deadlines in heterogeneous MEC due to lacking a global view of the future impacts on descendant tasks. To fill in this gap, we formulate the deadline-constrained cost optimization problem for dependent task scheduling in MEC and propose a low-complexity scheduling algorithm that considers a single task's future impacts in two stages. Specifically: (1) In the edge scheduling stage, each task is scheduled according to its successors’ latest start times instead of its sub-deadline to alleviate the lateness of its successors. An edge-only schedule plan is generated by scheduling tasks only on edge servers to save execution costs. (2) In the cloud offloading stage, in order to utilize the powerful cloud resources to satisfy the deadline, the edge-only schedule plan missing the deadline is efficiently modified by properly offloading multiple successive tasks to the cloud. Simulation results show the substantial advantage of the proposed algorithm over baselines in both online and offline scenarios. Jiong Lou, Zhiqing Tang, Songli Zhang, Weijia Jia 0001, Wei Zhao 0001, Jie Li 0002 |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | Layer Dependency-Aware Learning Scheduling Algorithms for Containers in Mobile Edge ComputingabstractDue to the features of lightweight and easy deployment, the use of containers has emerged as a promising approach for Mobile Edge Computing (MEC). Before running the container, an image composed of several layers must exist locally. However, it has been conspicuously neglected by existing work that task scheduling at the granularity of the layer instead of the image can significantly reduce the task completion time to further meet the real-time requirement and resource efficiency in resource-limited MEC. To bridge the gap, considering the complex dependency between layers and images, a novel layer dependency-aware container scheduling algorithm is proposed to reduce the total task completion time. Specifically: 1) We model the online layer dependency-aware scheduling problem for containers in a heterogeneous MEC, considering the layer download time and task computation time. 2) A policy gradient algorithm is proposed to solve this problem, and the high-dimensional and low-dimensional relations for layer dependencies are extracted with improved action selection. 3) Experiments based on the real-world data trace show that the proposed algorithm outperforms the image-based and layer-based baseline algorithms by 54% and 19% on average, respectively. Zhiqing Tang, Jiong Lou, Weijia Jia 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | Energy-Efficient Joint Task Assignment and Migration in Data Centers: A Deep Reinforcement Learning ApproachabstractEnergy-efficient task scheduling in data centers is a critical issue and has drawn wide attention. However, the task execution times are mixed and hard to estimate in a real-world data center. It has been conspicuously neglected by existing work that scheduling decisions made at tasks’ arrival times are likely to cause energy waste or idle resources over time. To fill in such gaps, in this paper, we jointly consider assignment and migration for mixed duration tasks and devise a novel energy-efficient task scheduling algorithm. Task assignment can improve resource utilization, and migration is required when long-running tasks run in low-load servers. Specifically: 1) We formulate mixed duration task scheduling as a large-scale Markov Decision Process (MDP) problem; 2) To solve such a large-scale MDP problem, we design an efficient Deep Reinforcement Learning (DRL) algorithm to make assignment and migration decisions. To make the DRL algorithm more practical in real scenarios, multiple optimizations are proposed to achieve online training; 3) Experiments with real-world data have shown that our algorithm outperforms the existing baselines 14% on average in terms of energy consumption while keeping the same level of Quality of Service (QoS). Jiong Lou, Zhiqing Tang, Weijia Jia 0001 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2023 | Efficient Container Assignment and Layer Sequencing in Edge ComputingabstractContainers are becoming a popular way of running applications in edge computing. Before running the application, the edge node must download the application’s container image consisting of multiple layers. However, given the limited bandwidth in edge computing, the container startup latency due to long image download time seriously affects the real-time performance. In this article, we jointly determine the container assignment and the layer download sequence to reduce the total startup latency. We formulate the Container Assignment and Layer Sequencing (CALS) problem and prove its NP-hardness. A Layer-Aware Scheduling Algorithm (LASA) is proposed, fully considering layer sharing among images. First, layers shared by the same set of images are grouped to reduce CALS’s problem scale without affecting the optimal result. Second, considering both layer sharing and existing layer size on edge nodes, a layer-aware algorithm is designed to assign containers to appropriate edge nodes. Finally, to determine the layer download sequence on each edge node, an approximation algorithm is proposed. We further analyze the approximation ratio of LASA in the case of identical edge nodes with sufficient capacity. Extensive experiments based on real-world data show the effectiveness of LASA, which reduces the total startup latency by 40% to 60%. Jiong Lou, Hao Luo 0012, Zhiqing Tang, Weijia Jia 0001, Wei Zhao 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2022 | Efficient instance reuse approach for service function chain placement in mobile edge computing
Songli Zhang, Weijia Jia 0001, Zhiqing Tang, Jiong Lou, Wei Zhao 0001 |
Comput. Networks | 3 |
| 2022 | Representation and Reinforcement Learning for Task Scheduling in Edge ComputingabstractRecently, many deep reinforcement learning (DRL)-based task scheduling algorithms have been widely used in edge computing (EC) to reduce energy consumption. Unlike the existing algorithms considering fixed and fewer edge nodes (servers) and tasks, in this article, a representation model with a DRL based algorithm is proposed to adapt the dynamic change of nodes and tasks and solve the dimensional disaster in DRL caused by a massive scale. Specifically, 1) we apply the representation learning models to describe the different nodes and tasks in EC, i.e., nodes and tasks are mapped to corresponding vector sub-spaces to reduce the dimensions and store the vector space efficiently. 2) With the space after dimensionality reduction, a DRL-based algorithm is employed to learn the vector representations of nodes and tasks and make scheduling decisions. 3) The experiments are conducted with the real-world data set, and the results show that the proposed representation model with DRL-based algorithm outperforms the baselines 18.04 and 9.94 percent on average regarding energy consumption and service level agreement violation (SLAV), respectively. Zhiqing Tang, Wei Jia 0001, Wenmian Yang, Yongjian You |
IEEE Trans. Big Data | 1 |
| 2020 | Dependent Task Offloading for Multiple Jobs in Edge ComputingabstractThe dependent task offloading problem for one single job in edge computing (EC) has drawn attention widely. Unlike most existing approaches that only focus on a single job, we aim to solve the dependent task offloading problem for multiple jobs, which is more general in the real world. To solve this problem, we propose a deep reinforcement learning (DRL) based multi-job dependent task offloading algorithm. Specifically, 1) we model edge nodes, jobs, and tasks in a resource-limited EC scenario, where the dependent tasks of multiple jobs are offloaded to the nodes to be processed. Then we model the task offloading decision as a Markov decision process (MDP) problem to minimize the transmission cost and computation cost. 2) To represent the state space of MDP and to accelerate decision-making in EC, we propose a DRL-based algorithm with the aid of graph convolutional network (GCN) to extract the dependency information of different tasks and then improve the action selection process. 3) We conduct experiments with real-world trace, demonstrating our algorithm outperforms the baseline algorithms 13.78% on average in regarding to offloading cost. Zhiqing Tang, Jiong Lou, Fuming Zhang, Weijia Jia 0001 |
ICCCN | 1 |
| 2020 | WiFind: Driver Fatigue Detection with Fine-Grained Wi-Fi Signal FeaturesabstractDriver fatigue is a leading factor in road accidents that can cause severe fatalities. Existing fatigue detection works focus on vision and electroencephalography (EEG) based means of detection. However, vision-based approaches suffer from view-blocking or vision distortion problems and EEG-based systems are intrusive, and the drivers have to use/wear the devices with inconvenience or additional costs. In our work, we propose a novel Wi-Fi signals based fatigue detection approach, called WiFind to overcome the drawbacks as associated with the current works. WiFind is simple and (wearable) device-free. It can detect the fatigue symptoms in the vehicle without relying on any visual image or video. By applying self-adaptive method, it can recognize the body features of drivers in multiple modes. It applies Hilbert-Huang transform (HHT) based pattern extract method results in accuracy increase in motion detection mode. WiFind can be easily deployed in a commodity Wi-Fi infrastructure, and we have evaluated its performance in real driving environments. The experimental results have shown that WiFind can achieve the recognition accuracy of 89.6 percent in a single driver scenario. Weijia Jia 0001, Hongjian Peng, Na Ruan, Zhiqing Tang, Wei Zhao 0001 |
IEEE Trans. Big Data | 4 |
| 2019 | Online Joint Scheduling of Delay-Sensitive and Computation-Oriented Tasks in Edge ComputingabstractIn the context of Edge Computing (EC) and Internet of Things (IoT), numerous tasks are offloaded from mobile users and sensor devices to edge nodes for further processing to reduce delay and solve the problem of insufficient local computation resources. These tasks can be mainly divided into delay-sensitive and computation-oriented tasks. The former tasks depend on the service provided by the container, while the latter tasks are submitted as a batch with task dependencies. Considering the heterogeneity of edge nodes, joint task scheduling can effectively improve resource utilization. However, relatively few researches consider the different characteristics of tasks like container constraints and task dependencies in joint task scheduling in EC. In order to fill in this gap, we propose a deep deterministic policy gradient (DDPG) based online joint task scheduling (OJTS) algorithm. Specifically, 1) We first model the problem of joint scheduling of delay-sensitive and computation-oriented tasks in resource-constrained EC scenario with the goals of maximizing system utility and minimizing system cost (weighted sum of the number and duration of unfinished tasks). 2) Then, we propose a deep reinforcement learning (DRL) algorithm to solve the above problem and make appropriate adjustments to the original network structure according to the scheduling decision. 3) Through validation on real-world trace, OJTS can improve the system utility by 26.0% and overall reward by 51.2% compared with baselines and meet real-time decision-making requirements. Fuming Zhang, Zhiqing Tang, Jiong Lou, Weijia Jia 0001 |
MSN | 2 |
| 2019 | Migration Modeling and Learning Algorithms for Containers in Fog ComputingabstractFog Computing (FC) is a flexible architecture to support distributed domain-specific applications with cloud-like quality of service. However, current FC still lacks the mobility support mechanism when facing many mobile users with diversified application quality requirements. Such mobility support mechanism can be critical such as in the industrial internet where human, products, and devices are moveable. To fill in such gaps, in this paper we propose novel container migration algorithms and architecture to support mobility tasks with various application requirements. Our algorithms are realized from three aspects: 1) We consider mobile application tasks can be hosted in a container of a corresponding fog node that can be migrated, taking the communication delay and computational power consumption into consideration; 2) We further model such container migration strategy as multiple dimensional Markov Decision Process (MDP) spaces. To effectively reduce the large MDP spaces, efficient deep reinforcement learning algorithms are devised to achieve fast decision-making and 3) We implement the model and algorithms as a container migration prototype system and test its feasibility and performance. Extensive experiments show that our strategy outperforms the existing baseline approaches 2.9, 48.5 and 58.4 percent on average in terms of delay, power consumption, and migration cost, respectively. Zhiqing Tang, Fuming Zhang, Weijia Jia 0001, Wei Zhao 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2018 | A Dynamic Resource Overbooking Mechanism in Fog ComputingabstractFog Computing (FC - similarly edge computing) as new computing paradigm can support distributed domain-specific or area-specific applications with cloud-like quality of service (QoS). This promising paradigm thus can find its wide applications in various industrial scenarios and smart cities in which the resource requirements will be divided into peak-hour or non-peak-hour. To deal with such features of applications, a flexible resource allocation approach based on pricing model can be critical for the success of such paradigm. To the best of our knowledge, we have not seen such pricing based resource allocation approach ever been reported for FC scenarios. In this paper, we propose a novel pricing based dynamic resource allocation model through overbooking mechanism, and it is realized through three steps: 1) According to different QoS requirements of user tasks, methods of on-demand billing, daily billing, and auction billing are designed, in which we allow the resource to be overbooked; 2) For auction billing, we design an auction approach including pricing rule and winner determination rule. We prove that our auction approach guarantees individual rationality, computational efficiency, and truthfulness. 3) To overbook as much resource as possible with a high degree of QoS satisfaction of on-demand and daily billing, we overbook the resource based on a resource utilization prediction using neural network and service level agreement violation feedback. In the end, we validate the mechanism with real-world data trace. Experimental results show that our auction approach achieves desirable properties, and our dynamic resource overbooking mechanism maximizes the profit of nodes with a high degree of QoS satisfaction of on-demand and daily billing and a high resource utilization prediction accuracy rate. Fuming Zhang, Zhiqing Tang, Mingcheng Chen, Weijia Jia 0001 |
MASS | 2 |