VLDB 2026 Research / reviewers in the wild / expert
Yibo Jin 0001
dblp:184/7591-1
· DBLP profile ↗
31ranked-venue papers
8as first author
28since 2021 · last 2025
0000-0001-5560-1338ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 22 · 5 first-author · 21 since 2021Systems, architecture and hardware · 8 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Orchestrating In-Network Aggregation for Distributed Machine Learning via In-Band Network Telemetry
Mingtao Ji, Yibo Jin 0001, Zhuzhong Qian, Tuo Cao |
J. Comput. Sci. Technol. | 2 |
| 2024 | INTaaS: Provisioning In-band Network Telemetry as a service via online learning
Mingtao Ji, Chenwei Su, Yitao Fan, Yibo Jin 0001, Zhuzhong Qian, Yu Chen 0038, Tuo Cao, Sheng Zhang 0001 |
Comput. Networks | 4 |
| 2024 | Crowdsourcing Upon Learning: Energy-Aware Dispatch With Guarantee for Video AnalyticsabstractOver the last decade, the mobile crowdsourcing has become a paradigm to conduct the manual annotation and further analytics by recruited workers, with their rewards depending on the result quality. Existing dispatchers cannot precisely capture the resource-quality trade-off for video analytics, because the configurations supported by recruited workers are limited, and workers’ availability changes over time. To determine the most suitable configurations as well as workers for video analytics, we formulate a non-linear mixed program in long term, maximizing the crowdsourcing profit. Based on previous results under various configurations and workers, we design an algorithm via a series of subproblems to decide the configurations adaptively upon the prediction of workers’ feedbacks. Such prediction is based on volatile multi-armed bandit to capture workers’ availability and stochastic changes on resource uses. Furthermore, we extend the proposed algorithms to the multi-worker selection scenario where the platform needs to determine a candidate worker set instead of a single worker for video analytics. Via rigorous proof, the regret is ensured upon the Lyapunov optimization and the bandit, measuring the gap between the online decisions and the offline optimum. Extensive trace-driven experiments show that our proposed algorithm improves the profit by 37% compared with other algorithms. Yu Chen 0038, Sheng Zhang 0001, Yibo Jin 0001, Zhuzhong Qian, Mingjun Xiao, Yu Liang 0001, Sanglu Lu |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | Spatial and Temporal Detection With Attention for Real-Time Video Analytics at EdgesabstractThe detection of objects via neural networks plays a key role in various video analytics, but consumes huge resources. Due to the limited computing capability at edges, such real-time detections should be precisely used for the objects that need the most attention. Unfortunately, as the target objects keep moving, existing systems fail to consider both the distribution of objects and object movements over regions, and existing tracking mechanisms are easily affected by background content. Therefore, we propose spatial and temporal detection with attention for analytics, to increase the quality of detections for those targets. However, the attention shift over regions, the uncertainty of detections, and the constrained edge resources essentially hamper us from efficient analytics. We propose an adaptive partition planner to divide the frame into regions to achieve spatial attention. Afterwards, we design a detection planner to orchestrate the detection model temporally for each region by an online mechanism, via a queue-based adaptation. The spatial and temporal attention are integrated to maximize the accumulative detection accuracy. Via rigorous proof, both dynamic regret regarding detection accuracy and the real-time requirement for the video analytics are ensured. The testbed experiments confirm the superiority of our approach over multiple state-of-the-art algorithms. Sheng Zhang 0001, Yibo Jin 0001, Fangwen Cheng, Zhuzhong Qian, Sanglu Lu |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | ViChaser: Chase Your Viewpoint for Live Video Streaming With Block-Oriented Super-ResolutionabstractThe usage of live streaming services has led to a substantial increase in live video traffic. However, the perceived quality of experience of users is frequently limited by variations in the upstream bandwidth of streamers. To address this issue, several adaptive bitrate (ABR) algorithms have been developed to mitigate bandwidth variations. Nevertheless, the ability of users to enjoy high-quality live streams remains limited. While neural-enhanced approaches, such as super-resolution, offer significant quality improvements, frame-oriented super-resolution leads to excessive inference delay that violates the real-time feature of live streaming. In response, we propose ViChaser, which examines block-oriented super-resolution for live streaming. ViChaser performs neural super-resolution on potential blocks of interest in the media server, corresponding to the user’s viewpoint, and uses online learning to adapt to the dynamic content of the video. Additionally, ViChaser utilizes the Lyapunov framework to efficiently allocate uplink bandwidth for original low-quality live video and high-quality labels. The experimental results demonstrate that ViChaser achieves 1.2–1.5 dB higher video quality in Peak-Signal-to-Noise-Ratio than WebRTC and increases processing speed by 11–16 fps relative to LiveNAS. Ning Chen 0010, Sheng Zhang 0001, Zhi Ma 0002, Yu Chen 0038, Yibo Jin 0001, Jie Wu 0001, Zhuzhong Qian, Yu Liang 0001, Sanglu Lu |
IEEE/ACM Trans. Netw. | 5 |
| 2023 | INTView: Adaptive Planner for In-Band Network Telemetry without DetoursabstractNetwork visualization is essential for network operators to diagnose ongoing network failures and understand the quality of the network. In-Band Network Telemetry (INT) supports network visualization by inserting P4 switch state information (e.g., queue length, hop latency, and link utilization) into the specific INT packets. In order to achieve network-wide coverage, paths of INT packets need to be delicately designed to ensure non-overlapping, high performance, and low overhead. However, existing INT path planning solutions ignore capturing the dynamic network status and fail to obtain the optimum. In this paper, we model the INT path planning based on the directed Edge Cover problem upon the dynamic network status, with the objective of minimizing the INT path latency. Although it is actually a non-linear integer program problem, by adopting delicate transformations, we design approximation algorithms with performance guarantees. We implement our system prototype INTView based on INTCollector upon real devices, i.e., Barefoot Wedge100BF and Inspur Rack. Extensive evaluation upon realistic settings shows that our proposed algorithm achieves 2x performance improvement regarding the completion time, compared with state-of-the-art schemas. Mingtao Ji, Chenwei Su, Zhuzhong Qian, Yu Chen 0038, Yibo Jin 0001, Sheng Zhang 0001 |
ICC | 6 |
| 2023 | Crowd2: Multi-agent Bandit-based Dispatch for Video Analytics upon CrowdsourcingabstractMany crowdsourcing platforms are emerging, leveraging the resources of recruited workers to execute various outsourcing tasks, mainly for those computing-intensive video analytics with high quality requirements. Although the profit of each platform is strongly related to the quality of analytics feedback, due to the uncertainty on diverse performance of workers and the conflicts of interest over platforms, it is non-trivial to determine the dispatch of tasks with maximum benefits. In this paper, we design a decentralized mechanism for a Crowd of Crowdsourcing platforms, denoted as Crowd2, optimizing the worker selection to maximize the social welfare of these platforms in a long-term scope, under the consideration of both proportional fairness and dynamic flexibility. Concretely, we propose a video analytics dispatch algorithm based on multi-agent bandit, for which the more accurate profit estimates are attained via the decoupling of multi-knapsack based mapping problem. Via rigorous proofs, a sub-linear regret bound for social welfare of crowdsourcing profits is achieved while both fairness and flexibility are ensured. Extensive trace-driven experiments demonstrate that Crowd2improves the social welfare by 36.8%, compared with other alternatives. Yu Chen 0038, Sheng Zhang 0001, Yibo Jin 0001, Ning Chen 0010, Mingtao Ji, Mingjun Xiao |
INFOCOM | 4 |
| 2023 | Orchestrating Blockchain with Decentralized Federated Learning in Edge NetworksabstractDecentralized federated learning across edge networks can leverage blockchain with consensus mechanisms for training information exchange among participants over costly and distrustful wide-area networks. However, it is non-trivial to optimally operate the blockchain to support decentralized federated learning due to the complex cost structure of blockchain operations, the balance between blockchain overhead and model convergence, and the dynamics and uncertainties of edge network environments. To overcome these challenges, we formulate a non-linear time-varying integer program that jointly places blockchain nodes and determines the number of training iterations to minimize the long-term blockchain computation and communication cost. We then design an online polynomial-time approximation algorithm that decomposes the problem and solves the subproblems alternately on the fly using only estimated inputs. We rigorously prove the sublinear regret of our approach. We further implement our approach with a prototype system, and conduct extensive trace-driven experiments to validate the superiority of our approach over other alternatives. Yibo Jin 0001, Lei Jiao 0002, Zhuzhong Qian, Ruiting Zhou, Lingjun Pu |
SECON | 1 |
| 2023 | Scheduling In-Band Network Telemetry With Convergence-Preserving Federated LearningabstractConducting federated learning across distributed sites with In-Band Network Telemetry (INT) based data collection faces critical challenges, including control decisions of different frequencies, convergence of the models being trained, and resource provisioning coupled over time. To study this problem, we formulate a non-linear mixed-integer program to optimize the long-term INT overhead, resource cost, and federated learning cost. We then design polynomial-time online algorithms to solve this problem with only observable inputs on the fly, featuring laziness-aware resource adaption, online-learning-based INT flow selection and model aggregation control, as well as expectation-preserving randomized dependent rounding. We rigorously prove the parameterized-constant competitive ratio of our approach against the offline optimum, and the time-averaged constraint violation that vanishes in the long run. With extensive trace-driven evaluations, we confirm the superiority of our approach over other alternative approaches for reducing total cost and the efficacy of our trained models for solving real machine learning problems, reducing the real-time cost by 34% on average. Yibo Jin 0001, Lei Jiao 0002, Mingtao Ji, Zhuzhong Qian, Sheng Zhang 0001, Ning Chen 0010, Sanglu Lu |
IEEE/ACM Trans. Netw. | 1 |
| 2022 | Multi-server Multi-user Game at Edges for Heterogeneous Video AnalyticsabstractIn past years, artificial intelligence related services and applications have boomed, which require high computation, high bandwidth and low latency. Edge computing is regarded as an appropriate solution for them, especially video analytics. In this paper, we study the multi-server multi-user heterogeneous video analytics offloading problem, where users select appropriate edge servers and then offload their raw video data to the servers for essential analytics. To deal with the cooperation and conflicts among users and get a stable situation where each user has no incentive to change the offloading decision unilaterally, we formulate the video analytics offloading problem as a multiplayer game. Based on the goal of minimizing the overall delay, we design the potential optimal server selection strategy and then propose a game theory-based algorithm, through which the Nash equilibrium can be reached. Furthermore, we analyze its near-optimal performance via rigorous proof. Finally, extensive trace-driven experiments show that our method improves the overall delay by 48% on average, compared with other algorithms. Yu Chen 0038, Sheng Zhang 0001, Yibo Jin 0001, Zhuzhong Qian, Sanglu Lu |
ICC | 3 |
| 2022 | Learning for Crowdsourcing: Online Dispatch for Video Analytics with GuaranteeabstractCrowdsourcing enables a paradigm to conduct the manual annotation and the analytics by those recruited workers, with their rewards relevant to the quality of the results. Existing dispatchers fail to capture the resource-quality trade-off for video analytics, since the configurations supported by various workers are different, and the workers’ availability is essentially dynamic. To determine the most suitable configurations as well as workers for video analytics, we formulate a non-linear mixed program in a long-term scope, maximizing the profit for the crowdsourcing platform. Based on previous results under various configurations and workers, we design an algorithm via a series of subproblems to decide the configurations adaptively upon the prediction of the worker rewards. Such prediction is based on volatile multi-armed bandit to capture the workers’ availability and stochastic changes on resource uses. Via rigorous proof, the regret is ensured upon the Lyapunov optimization and the bandit, measuring the gap between the online decisions and the offline optimum. Extensive trace-driven experiments show that our algorithm improves the platform profit by 37%, compared with other algorithms. Yu Chen 0038, Sheng Zhang 0001, Yibo Jin 0001, Zhuzhong Qian, Mingjun Xiao, Ning Chen 0010, Zhi Ma 0002 |
INFOCOM | 3 |
| 2022 | RCM: Residue-aware Consolidation for Heterogeneous MLaaS ClusterabstractWith the rapid development of Machine Learning (ML), Machine-Learning-as-a-Service (MLaaS) clusters appear in large numbers to support cloud platforms services, which adopt virtual machine (VM) to improve the availability, resilience and security. However, low energy efficiency is a major problem in such clusters. Previous work focused on reducing the number of physical machines by centralizing resources migration. Nevertheless, for ML tasks with frequent memory switching, blind migration is not worth the cost because the remaining time is less than the migration time, since the migration time can not be ignore due to the memory intensive of ML tasks. Therefore, this paper explores how the remaining time and memory replacement states in ML tasks, which we summarize as residue, affect migration, and proposes an online residue-aware migration algorithm based on Lyapunov optimization. Through rigorous proof, the gap between the algorithm and the optimal solution is ensured. Extensive simulations show that the proposed algorithm is better than the previous migration. Kefeng Wu, Chunlei Xu, Xiongfeng Hu, Yibo Jin 0001, Zhuzhong Qian |
IPCCC | 5 |
| 2022 | User-Perceived QoE Adaptation for Accelerated Playback in Mobile Video StreamingabstractUser-perceived quality of experience (QoE) is critical as mobile video streaming experiences a substantial growth. User's demands are becoming diversified where accelerated play-back is the preference of a considerable part of users. However, the limited and fluctuate mobile bandwidth is often not capable of satisfying user's demand of watching video at 2x or higher speed because of consequential frequent rebuffering. Previous adaptive bitrate (ABR) algorithms hardly consider the variety of user playback rates. In this work, we fully exploit the relation between user-perceived, i.e., subjective video quality and the characteristic of video content. The result of our motivational experiments shows that viewers are less sensitive to the bitrate variation and playback rate alternation if there is higher degree of motion in the video. With above guidelines, we adaptively adjust the quality configuration and playback rate to significantly reduce the rebuffering while achieving similar or even higher subjective quality. Then we formulate subjective quality and playback rate adaption as a QoE maximization problem and propose the content based subjective quality and playback rate adaptation algorithm (CSP) utilizing Lyapunov optimization technique. Via rigorous proof, the time-average QoE achieved by CSP is in$O(1/V)$gap compared to optimal value, where$V$is the control parameter. Extensive evaluations confirm the superiority of our proposed algorithm over other state-of-the-art algorithms under both normal and accelerated playback rate. Xiongfeng Hu, Yibo Jin 0001, Kefeng Wu, Zhuzhong Qian, Sanglu Lu |
MSN | 2 |
| 2022 | Energy-efficient Federated Learning via Stabilization-aware On-device Update ScalingabstractFederated learning is emerging as a major learning paradigm, which enables multiple devices to train a model col-laboratively and to keep the privacy of data. However, substantial computation-intensive iterations are performed on devices before the training completion, which incurs heavy consumption of the energy. Along with the stabilization of those model parameters being trained, such on-device training iterations are redundant gradually over time. Thus, we propose to scale the update results obtained from reduced iterations as the substitute for on-device training, based on current model status and device heterogeneity. We thus formulate a time-varying integer program, to minimize cumulative energy consumption over devices, subject to a long-term constraint regarding the model convergence. We then design a polynomial-time online algorithm upon system dynamics, which essentially balances the energy consumption and the model quality being trained. Via rigorous proofs, our approach only incurs sub linear regret, compared with its optimum, and ensures related model convergence. Extensive testbed experiments for real training confirm the superiority of our approach, over multiple alternatives, under various scenarios, decreasing at least 30.2% energy consumption, while preserving the accuracy of the model. Suwei Xu, Yibo Jin 0001, Zhuzhong Qian, Sheng Zhang 0001, Zhenjie Lin |
SECON | 2 |
| 2022 | Focus! Provisioning Attention-aware Detection for Real-time On-device Video AnalyticsabstractThe detection of objects via neural networks plays a key role in various video analytics, but consumes huge resources. Due to the limited on-device computing capability, such real-time detections should be precisely used for the objects that need the most attention. Unfortunately, as the target objects keep moving, existing systems fail to conduct adaptive detections over multiple regions in a video, and existing tracking mechanisms are easily affected by background contents. Therefore, we propose to design attention-aware on-device detection for analytics, to increase the quality of detections for those targets. However, the uncertainty of detections, the attention shift over regions, and the provisioning of on-device resources essentially hamper us from efficient analytics. We formulate such a scenario as a non-linear integer program in long-term scope, to maximize the detection accuracy. Afterwards, we design an online mechanism to orchestrate the detection model for each region in the video to cope with the moves of the targets, via a queue-based adaptation and the randomized rounding. Via rigorous proof, both dynamic regret regarding detection accuracy and the real-time requirement for the video analytics are ensured. The testbed experiments confirm the superiority of our approach over multiple state-of-the-art algorithms. Yibo Jin 0001, Sheng Zhang 0001, Fangwen Cheng, Zhuzhong Qian, Sanglu Lu |
SECON | 2 |
| 2022 | Adaptive provisioning for mobile cloud gaming at edges
Tuo Cao, Yibo Jin 0001, Xiongfeng Hu, Sheng Zhang 0001, Zhuzhong Qian, Sanglu Lu |
Comput. Networks | 2 |
| 2022 | Inference replication at edges via combinatorial multi-armed bandit
Hesheng Sun, Yibo Jin 0001, Yanfang Zhu, Zhuzhong Qian, Sheng Zhang 0001, Sanglu Lu |
J. Syst. Archit. | 3 |
| 2022 | Adaptive Configuration Selection and Bandwidth Allocation for Edge-Based Video AnalyticsabstractMajor cities worldwide have millions of cameras deployed for surveillance, business intelligence, traffic control, crime prevention, etc. Real-time analytics on video data demands intensive computation resources and high energy consumption. Traditional cloud-based video analytics relies on large centralized clusters to ingest video streams. With edge computing, we can offload compute-intensive analysis tasks to nearby servers, thus mitigating long latency incurred by data transmission via wide area networks. When offloading video frames from the front-end device to an edge server, the application configuration (i.e., frame sampling rate and frame resolution) will impact several metrics, such as energy consumption, analytics accuracy and user-perceived latency. In this paper, we study the configuration selection and bandwidth allocation for multiple video streams, which are connected to the same edge node sharing an upload link. We propose an efficient online algorithm, called JCAB, which jointly optimizes configuration adaption and bandwidth allocation to address a number of key challenges in edge-based video analytics systems, including edge capacity limitation, unknown network variation, intrusive dynamics of video contents. Our algorithm is developed based on Lyapunov optimization and Markov approximation, works online without requiring future information, and achieves a provable performance bound. We also extend the proposed algorithms to the multi-edge scenario in which each user or video stream has an additional choice about which edge server to connect. Extensive evaluation results show that the proposed solutions can effectively balance the analytics accuracy and energy consumption while keeping low system latency in a variety of settings. Sheng Zhang 0001, Yibo Jin 0001, Jie Wu 0001, Zhuzhong Qian, Mingjun Xiao, Sanglu Lu |
IEEE/ACM Trans. Netw. | 3 |
| 2022 | LOCUS: User-Perceived Delay-Aware Service Placement and User Allocation in MEC EnvironmentabstractIn the multi-access edge computing environment, app vendors deploy their services and applications at the network edges, and edge users offload their computation tasks to edge servers. We study the user-perceived delay-aware service placement and user-allocation problem in edge environment. We model the MEC-enabled network, where the user-perceived delay consists of computing delay and transmission delay. The total cost in the offloading system is defined as the sum of service placement, edge server usage and energy consumption cost, and we need to minimize the total cost by determining the overall service-placing decision and user-allocation decision, while guaranteeing that the user-perceived delay requirement of each user is fulfilled. Our considered problem is formulated as a Mixed Integer Linear Programming problem, and we prove its NP-hardness. Due to the intractability of the considered problem, we propose a LOCal-search based algorithm for USer-perceived delay-aware service placement and user-allocation in edge environment, named LOCUS, which starts with a feasible solution and then repeatedly reduces the total cost by performing local-search steps. After that, we analyze the time complexity of LOCUS and prove that it achieves provable guaranteed performance. Finally, we compare LOCUS with other existing methods and show its good performance through experiments. Yu Chen 0038, Sheng Zhang 0001, Yibo Jin 0001, Zhuzhong Qian, Mingjun Xiao, Jidong Ge, Sanglu Lu |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | $run$ runData: Re-Distributing Data via Piggybacking for Geo-Distributed Data Analytics Over EdgesabstractEfficiently analyzing geo-distributed datasets is emerging as a major demand in a cloud-edge system. Since the datasets are often generated in closer proximity to end users, traditional works mainly focus on offloading proper tasks from those hotspot edges to the datacenter to decrease the overall completion time of submitted jobs in a one-shot manner. However, optimizing the completion time of current job alone is insufficient in a long-term scope since some datasets would be used multiple times. Instead, optimizing the data distribution is much more efficient and could directly benefit forthcoming jobs, although it may postpone the execution of current one. Unfortunately, due to the throwaway feature of data fetcher, existing data analytics systems fail to re-distribute corresponding data out of hotspot edges after the execution of data analytics. In order to minimize the overall completion time for a sequence of jobs as well as to guarantee the performance of current one, we propose to re-distribute the data along with task offloading, and formulate corresponding ε-bounded data-driven task scheduling problem over wide area network under the consideration of edge heterogeneity. We design an online schemarunData, which offloads proper tasks and related data via piggybacking to the datacenter based on delicately calculated probabilities. Through rigorous theoretical analysis,runData is proved concentrated on its optimum with high probability. We implementrunData based on Spark and HDFS. Both testbed results and trace-driven simulations show that run Data re-distributes proper data via piggybacking and achieves up to 37 percent reduction on average response time compared with state-of-the-art schemas. Yibo Jin 0001, Zhuzhong Qian, Song Guo 0001, Sheng Zhang 0001, Lei Jiao 0002, Sanglu Lu |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | TRAN: Task Replication with Guarantee via Multi-armed BanditabstractWith the rapid development of edge computing, edge clusters need to deal with a tremendous amount of tasks, making some edge clusters overloaded, which further translates into task completion lag. Previous works usually copy the tasks from overloaded edges to idle edges so as to reduce the task queuing and computing delay. However, the completion delay of tasks copied to different edges cannot be predicted before the replication decision is made, which affects the overall task replication performance. In this paper, we propose an online task replication algorithm based on the predictions derived from multi-armed bandit. Via rigorous proof, the regret is ensured to be sub-linear upon the bandit, measuring the gap between the online decisions and the offline optimum. Extensive simulations are conducted to confirm the superiority of the proposed algorithm over state-of-the-art replication strategies. Bowen Peng, Jingmian Wang, Weiwei Miao, Zeng Zeng, Yibo Jin 0001, Sheng Zhang 0001, Zhuzhong Qian |
ICPADS | 6 |
| 2021 | Learning for Learning: Predictive Online Control of Federated Learning with Edge ProvisioningabstractOperating federated learning optimally over distributed cloud-edge networks is a non-trivial task, which requires to manage data transference from user devices to edges, resource provisioning at edges, and federated learning between edges and the cloud. We formulate a non-linear mixed integer program, minimizing the long-term cumulative cost of such a federated learning system while guaranteeing the desired convergence of the machine learning models being trained. We then design a set of novel polynomial-time online algorithms to make adaptive decisions by solving continuous solutions and converting them to integers to control the system on the fly, based only on the predicted inputs about the dynamic and uncertain cloud-edge environments via online learning. We rigorously prove the competitive ratio, capturing the multiplicative gap between our approach using predicted inputs and the offline optimum using actual inputs. Extensive evaluations with real-world training datasets and system parameters confirm the empirical superiority of our approach over multiple state-of-the-art algorithms. Yibo Jin 0001, Lei Jiao 0002, Zhuzhong Qian, Sheng Zhang 0001, Sanglu Lu |
INFOCOM | 1 |
| 2021 | Edge-assisted Online On-device Object Detection for Real-time Video AnalyticsabstractReal-time on-device object detection for video analytics fails to meet the accuracy requirement due to limited resources of mobile devices while offloading object detection inference to edges is time-consuming due to the transference of video data over edge networks. Based on the system with both on-device object tracking and edge-assisted analysis, we formulate a non-linear time-coupled program over time, maximizing the overall accuracy of object detection by deciding the frequency of edge-assisted inference, under the consideration of both dynamic edge networks and the constrained detection latency. We then design a learning-based online algorithm to adjust the threshold for triggering edge-assisted inference on the fly in terms of the object tracking results, which essentially controls the deviation of on-device tracking between two consecutive frames in the video, by only taking previously observable inputs. We rigorously prove that our approach only incurs sub-linear dynamic regret for the optimality objective. At last, we implement our proposed online schema, and extensive testbed results with real-world traces confirm the empirical superiority over alternative algorithms, in terms of up to 36% improvement on detection accuracy with ensured detection latency. Mengxi Hanyao, Yibo Jin 0001, Zhuzhong Qian, Sheng Zhang 0001, Sanglu Lu |
INFOCOM | 2 |
| 2021 | Soudain: Online Adaptive Profile Configuration for Real-time Video AnalyticsabstractSince the real-time video analytics with high accuracy requirement is resource-consuming, the profiles regarding such resource-accuracy trade-off are needed before the analytics for better resource allocation at resource-constrained edges. With the inner changes of the video contents, outdated profiles fail to capture the trade-off dynamically over time, which requires the profiles to be updated periodically and incurs an overwhelming resource overhead. Thus, we present Soudain, which dynamically adjusts the configurations in profiles and corresponding profiling intervals to capture the inner changes of multiple video streams at edges. Upon the fine-grained decisions for profiles, we propose an integer program to maximize the accuracy of video analytics in a long-term scope with resource constraint, and then design an algorithm to adjust the profiles in an online manner. We implement Soudain upon the server with GPU. Our testbed evaluations confirm that, by using the live video streams derived from real-world traffic cameras, Soudain ensures the real-time requirement and achieves up to 25% improvement on the detection accuracy, compared with multiple state-of-the-art alternatives. Yibo Jin 0001, Weiwei Miao, Zeng Zeng, Zhuzhong Qian, Jingmian Wang, Mingxian Zhou, Tuo Cao |
IWQoS | 2 |
| 2021 | Service Placement and Bandwidth Allocation for MEC-enabled Mobile Cloud GamingabstractMobile cloud gaming (MCG), which is potential to deliver high-quality gaming experience to users anywhere and anytime, suffers from tremendous wide-area traffic and long network delays. Mobile edge computing (MEC), where cloud computing capabilities are pushed to the network edge, can help by providing gaming services in the proximity to users. However, since the quality of experience (QoE), i.e., the gaming experience, is easily impaired by long network delays and low frame rates, the performance of MEC-enabled MCG highly depends on the placement of gaming services and the allocation of related bandwidth. Furthermore, due to the erratic mobility of users, migrating services to follow such mobility decreases the impairment but incurs extra system cost, leading to the performance-cost tradeoff. To address these challenges, in this paper, we jointly investigate service placement and bandwidth allocation for MEC-enabled MCG. Considering the system dynamics, we propose to minimize the QoE impairment in a long time scope under a cost constraint for long-term migrations. To solve the problem, we develop an online two-layer iterative algorithm OnTrial. Rigorous theoretical analyses demonstrate that OnTrial achieves a near-optimal performance and bounds the potential violation of the migration cost constraint. Simulation results show that OnTrial outperforms other algorithms by at least 25% on the long-term QoE impairment. Tuo Cao, Zhuzhong Qian, Mingxian Zhou, Yibo Jin 0001 |
WOWMOM | 5 |
| 2021 | Budget-Aware Online Control of Edge Federated Learning on Streaming Data With Stochastic InputsabstractPerforming federated learning continuously in edge networks while training data are dynamically and unpredictably streamed to the devices faces critical challenges, including the global model convergence, the long-term resource budget, and the uncertain stochastic network and execution environment. We formulate an integer program to capture all these challenges, which minimizes the cumulative total latency of stream learning on device and federated learning between devices and the edge server. We then decouple the problem, design an online learning algorithm for controlling the number of local model updates via a convex-concave reformulation and rectified gradient-descent steps, and design a bandit learning algorithm for selecting the edge server for global model aggregations by incorporating the budget information to strike the exploit-explore balance. We rigorously prove the sub-linear regret regarding the optimization objective and the sub-linear constraint violation regarding the maximal on-device load, while guaranteeing the convergence of the global model trained. Extensive evaluations with real-world training data and input traces confirm the empirical superiority of our approach over multiple state-of-the-art algorithms. Yibo Jin 0001, Lei Jiao 0002, Zhuzhong Qian, Sheng Zhang 0001, Sanglu Lu |
IEEE J. Sel. Areas Commun. | 1 |
| 2021 | Cuttlefish: Neural Configuration Adaptation for Video Analysis in Live Augmented RealityabstractInstead of relying on remote clouds, today's Augmented Reality (AR) applications usually send videos to nearby edge servers for analysis (such as objection detection) so as to optimize the user's quality of experience (QoE), which is often determined by not only detection latency but also detection accuracy, playback fluency, etc. Therefore, many studies have been conducted to help adaptively choose best video configuration, e.g., resolution and frame per second (fps), based on network bandwidth to further improve QoE. However, we notice that the video content itself has significant impacts on the configuration selection, e.g., the videos with high-speed objects must be encoded with a high fps to meet the user's fluency requirement. In this article, we aim to adaptively select configurations that match the time-varying network condition as well as the video content. We design Cuttlefish, a system that generates video configuration decisions using reinforcement learning (RL). Cuttlefish trains a neural network model that picks a configuration for the next encoding slot based on observations collected by AR devices. Cuttlefish does not rely on any pre-programmed models or specific assumptions on the environments. Instead, it learns to make configuration decisions solely through observations of the resulting performance of historical decisions. Cuttlefish automatically learns the adaptive configuration policy for diverse AR video streams and obtains a gratifying QoE. We compared Cuttlefish to several state-of-the-art bandwidth-based and velocity-based methods using trace-driven and real world experiments. The results show that Cuttlefish achieves a 18.4-25.8 percent higher QoE than the others. Ning Chen 0010, Siyi Quan, Sheng Zhang 0001, Zhuzhong Qian, Yibo Jin 0001, Jie Wu 0001, Sanglu Lu |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2021 | DeepSlicing: Collaborative and Adaptive CNN Inference With Low LatencyabstractThe booming of Convolutional Neural Networks (CNNs) has empowered lots of computer-vision applications. Due to its stringent requirement for computing resources, substantial research has been conducted on how to optimize its deployment and execution on resource-constrained devices. However, previous works have several weaknesses, including limited support for various CNN structures, fixed scheduling strategies, overlapped computations, high synchronization overheads, etc. In this article, we present DeepSlicing, a collaborative and adaptive inference system that adapts to various CNNs and supports customized flexible fine-grained scheduling. As a built-in functionality, DeepSlicing has supported typical CNNs including GoogLeNet, ResNet, etc. By partitioning both model and data, we also design an efficient scheduler, Proportional Synchronized Scheduler (PSS), which achieves the trade-off between computation and synchronization. Based on PyTorch, we have implemented DeepSlicing on the testbed with real-world edge settings that consists of 8 heterogeneous Raspberry Pi's. The results indicate that DeepSlicing with PSS outperforms the existing systems dramatically, e.g., the inference latency and memory footprint are reduced up to 5.79× and 14.72×, respectively. Shuai Zhang 0058, Sheng Zhang 0001, Zhuzhong Qian, Jie Wu 0001, Yibo Jin 0001, Sanglu Lu |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2020 | Resource-Efficient and Convergence-Preserving Online Participant Selection in Federated LearningabstractFederated learning achieves the privacy-preserving training of models on mobile devices by iteratively aggregating model updates instead of raw training data to the server. Since excessive training iterations and model transferences incur heavy usage of computation and communication resources, selecting appropriate devices and excluding unnecessary model updates can help save the resource usage. We formulate an online time-varying non-linear integer program to minimize the cumulative resource usage over time while achieving the desired long-term convergence of the model being trained. We design an online learning algorithm to make fractional control decisions based on both previous system dynamics and previous training results, and also design an online randomized rounding algorithm to convert the fractional decisions into integers without violating any constraints. We rigorously prove that our online approach only incurs sub-linear dynamic regret for the optimality loss and sub-linear dynamic fit for the long-term convergence violation. We conduct extensive trace-driven evaluations and confirm the empirical superiority of our approach over alternative algorithms in terms of up to 27% reduction on the resource usage while sacrificing only 4% reduction on accuracy. Yibo Jin 0001, Lei Jiao 0002, Zhuzhong Qian, Sheng Zhang 0001, Sanglu Lu, Xiaoliang Wang 0001 |
ICDCS | 1 |
| 2020 | Provisioning Edge Inference as a Service via Online LearningabstractProvisioning machine learning inference as a service at the mobile network edge for distributed users in an online setting faces multiple challenges, including the accuracy-resource trade-off for model selection, the time-coupled decision for model distribution, and the unpredictable user inference workload. To overcome such challenges, we firstly model an online time-varying non-linear integer program of maximizing the overall service's inference accuracy through dynamic model instance selection, delivery and workload distribution. Afterwards, we design an online learning algorithm to make fractional control decisions, which alternates between minimizing an outer problem and maximizing an inner problem of an equivalent convex-concave formulation by only taking previously observable inputs. We further design a randomized rounding algorithm to convert the fractional decisions into integers. We rigorously prove that our approach only incurs sub-linear dynamic regret for the optimality loss and sub-linear dynamic fit for the long-term constraints violation. Finally, we conduct extensive evaluations with real- world data and confirm the empirical superiority of our approach over state-of-the-art algorithms in terms of up to 30% reduction on accuracy loss and 34% reduction on constraints violation. Yibo Jin 0001, Lei Jiao 0002, Zhuzhong Qian, Sheng Zhang 0001, Ning Chen 0010, Sanglu Lu, Xiaoliang Wang 0001 |
SECON | 1 |
| 2018 | ran-GJS: Orchestrating Data Analytics for Heterogeneous Geo-distributed EdgesabstractMany organizations and companies have deployed not only datacenters but also large number of geo-distributed heterogeneous edges to provide fast data analytics services. Since large volume of data transmission across WAN can be costly, existing works mainly focus on pre-processing data in-place to avoid transmission. However, the heterogeneity of edges on either local computing capacity or network bandwidth limits the efficient use on scarce resource, which may result in long task completion time. To cope with dynamic demands on scarce resource, we take the heterogeneity of both computing capacity and network bandwidth of geo-distributed edges into consideration when assigning data analytical tasks and their associated data between the central datacenter and edges such that the overall latency can be reduced. We formulate the geo-distributed data-task joint scheduling problem (GJS), show its NP-hardness, and propose a near-optimal randomized scheduling algorithm (ran-GJS). ran-GJS can be proved concentrated around its optimum value with high probability, i.e., 1--O(e--t2) where t is the concentration bound by using Martingale Analysis. The experimental results obtained form both extensive simulations and Yarn-based prototype show that ran-GJS significantly speeds up the geo-distributed analytics with a gain on average completion time of at least 28% over state-of-the-art baseline algorithms. Yibo Jin 0001, Zhuzhong Qian, Song Guo 0001, Sheng Zhang 0001, Xiaoliang Wang 0001, Sanglu Lu |
ICPP | 1 |