EDBT 2026 Demo / reviewers in the wild / expert
Fang Dong 0001
dblp:75/2871-1
· DBLP profile ↗
116ranked-venue papers
18as first author
65since 2021 · last 2026
0000-0001-6770-326XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 10 first-author · 15 since 2021Computer networks · 31 · 2 first-author · 27 since 2021Human-computer interaction and ubiquitous computing · 27 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 7 · 4 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Software engineering, systems software and programming languages · 6 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Otter: Mitigating Background Distractions of Wide-Angle Few-Shot Action Recognition with Enhanced RWKV
Wenbo Huang 0001, Jinghui Zhang 0001, Guang Li 0008, Lei Zhang 0130, Fang Dong 0001, Takahiro Ogawa 0001, Miki Haseyama |
AAAI | 7 |
| 2026 | PipeNN: A Predictor-Free Pipeline for Energy-Latency Co-Optimization of Heterogeneous Mobile DAG-DNN Inference
Yukun Tian, Tianwei Jiang, Ruiting Zhou, Fang Dong 0001, Mengyang Liu |
ICDCS | 7 |
| 2026 | Decentralizing Compressed Sensing for Federated Learning with Hardware-Software Codesign
Fan Sun 0005, Fang Dong 0001, Dian Shen |
INFOCOM | 2 |
| 2026 | Adaptive region encoding for efficient video object detection in edge computing
Lisha Gao, Zhenxuan Xu, Zhaowu Huang, Fang Dong 0001 |
Knowl. Based Syst. | 6 |
| 2026 | Corrigendum: DESIGN: Online Device Selection and Edge Association for Federated Synergy Learning-enabled AIoTabstractThis is a corrigendum for the article “DESIGN: Online Device Selection and Edge Association for Federated Synergy Learning-enabled AIoT” published in ACM Trans. Intell. Syst. Technol. 15, 5, Article 104 (November 2024), 28 pages. Shucun Fu, Fang Dong 0001, Dian Shen, Runze Chen 0001, Jiangshan Hao |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2026 | Popularity-Aware Layer-Wise Caching and Function Scheduling for Dynamic Workflow at the EdgeabstractServerless Edge Computing (SEC) has emerged as a promising paradigm for delivering low-latency, resource-efficient services for edge-native applications, which are implemented as dependent functions, forming Directed Acyclic Graph (DAG) workflows. Unfortunately, the application's performance is hindered by the notorious issue of cold startup, especially in resource-constrained SEC environments. Layer- wise container caching has been proven to be an effective startup acceleration solution in SEC, due to its fine granularity and flexibility. However, due to the dynamic nature of call graphs and the skewness in function popularity in DAG workflows, as well as the heterogeneity of container layer cold start time and edge computing environments, the performance of existing layer- wise caching mechanisms degrades significantly. To solve this problem, we propose an efficient DAG workflow deployment method in SEC to minimize the application completion time (ACT) in the long term. We model the problem as a joint optimization of container layer- wise caching and function scheduling, which is a Time-coupled Integer Nonlinear Programming (TINLP) problem. To solve it, we first convert it to an Integer Linear Programming (ILP) problem and propose an online algorithm with theoretical performance guarantees. Extensive experiments demonstrate that our method achieves up to$2.92\times$speedup in ACT reduction. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang |
IEEE Trans. Mob. Comput. | 2 |
| 2026 | Balloon: An Adaptive Indexing System for Distributed Collaborative Edge Data QueryabstractEdge computing facilitates the development and implementation of new mechanisms for data storage on edge servers located near users, thereby enabling low-latency edge data retrieval services. However, individual edge servers have limited storage capacity, which restricts their ability to meet edge users' diverse data demands. They often need to cooperate with neighbors. Edge data indexing systems play a crucial role in helping edge servers locate the requested data. Existing edge data indexing systems typically rely on static indexing structures, which may lead to significantly reduced query accuracy or excessive system memory overhead in dynamic edge data query (EDQ) scenarios, making it difficult to meet the high scalability requirements of edge computing systems. In this paper, we address these challenges and propose a novel adaptive edge data indexing system called Balloon, which dynamically scales memory usage to balance query accuracy and memory overhead as data volume varies at runtime. Specifically, each edge server maintains a resizable Adaptive Bloom Filter (ABF) for its local data, and builds a Balloon tree that aggregates the ABFs of neighboring servers within a latency constraint. The size of the Balloon tree dynamically scales with data volume to balance query accuracy and memory overhead. To validate Balloon, we implement and conduct comprehensive experiments using an edge storage system with 50 edge servers. The results indicate that, compared to the state-of-the-art edge indexing system, Balloon improves query accuracy by 26.74% and reduces query time by 18.16%. In the meantime, it reduces up to 35.66% memory overhead. Siyu Tan, Fang Dong 0001, Qiang He 0001, Shuting Qiu, Yun Yang 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2026 | Joint Optimization of DNN Model Caching and Request Routing in Mobile Edge ComputingabstractMobile edge computing (MEC) can pre-cache deep neural networks (DNNs) near end-users, providing low-latency services and improving users’ quality of experience (QoE). However, caching all DNN models at capacity-limited edge servers is difficult, and the impact of model loading time on QoE remains underexplored. We explore dynamic DNNs by disassembling a complete DNN model into interrelated submodels to enable fine-grained joint optimization of submodel caching and request routing to balance inference precision and loading latency. In this paper, we study the joint dynamic model caching and request routing problem in MEC networks, aiming to maximize user request inference precision under constraints of server resources, latency, and model loading time. We propose CoCaR, an offline algorithm based on linear programming and random rounding that optimizes joint decisions with a provable performance bound. Furthermore, we develop an online extension, CoCaROL, to adapt to dynamic and unpredictable request patterns. The simulation results demonstrate that CoCaR improves the average inference precision for user requests by 40.1% over state-of-the-art baselines. In addition, CoCaR-OL achieves an improvement of at least 32.3% in users’ QoE over competitive baselines. Shuting Qiu, Fang Dong 0001, Siyu Tan, Ruiting Zhou, Dian Shen, Patrick P. C. Lee, Qilin Fan |
IEEE Trans. Netw. | 2 |
| 2026 | Enabling Efficient Synergistic Multi-view Inference Across Heterogeneous Edge DevicesabstractMulti-view inference (MVI), which accepts images from multiple viewpoints as input of deep neural networks, is proposed to improve the inference accuracy of conventional single-view models. However, existing mechanisms face challenges in feature fusion and computation efficiency: (1) features from inter-view and intra-view contribute differently to inference, and uniform feature fusion limits MVI accuracy; (2) the sophisticated process and tremendous computational workload of MVI cause a considerable increase in inference latency. This article addresses the above challenges and enables high-accuracy and low-latency MVI for edge intelligence by proposing an end-to-edge synergistic multi-view inference (SMVI) framework. SMVI integrates the f eature f u sion module based on pairwise m utual- a ttention (FUMA), which incorporates the differences between features, enhancing MVI accuracy. To optimize the computation of FUMA-based SMVI, we present a joint optimization algorithm of r esource a llocation and m odel p artition (RAMP) to reduce MVI latency, considering device heterogeneity, dynamic network connection, and resource limitation in heterogeneous edge environments. We developed an SMVI prototype system with heterogeneous embedded GPUs and evaluated its performance in real-world MVI scenarios. Extensive experiments demonstrate that the proposed mechanism achieves a notable MVI accuracy improvement of approximately 4% and accelerates the process by 4.08 × compared to state-of-the-art approaches. Fang Dong 0001, Runze Chen 0001, Shucun Fu, Wangbing Cheng, Ruiting Zhou |
ACM Trans. Sens. Networks | 1 |
| 2025 | Manta: Enhancing Mamba for Few-Shot Action Recognition of Long Sub-SequenceabstractIn few-shot action recognition (FSAR), long sub-sequences of video naturally express entire actions more effectively. However, the high computational complexity of mainstream Transformer-based methods limits their application. Recent Mamba demonstrates efficiency in modeling long sequences, but directly applying Mamba to FSAR overlooks the importance of local feature modeling and alignment. Moreover, long sub-sequences within the same class accumulate intra-class variance, which adversely impacts FSAR performance. To solve these challenges, we propose a Matryoshka MAmba and CoNtrasTive LeArning framework (Manta). Firstly, the Matryoshka Mamba introduces multiple Inner Modules to enhance local feature representation, rather than directly modeling global features. An Outer Module captures dependencies of timeline between these local features for implicit temporal alignment. Secondly, a hybrid contrastive learning paradigm, combining both supervised and unsupervised methods, is designed to mitigate the negative effects of intra-class variance accumulation. The Matryoshka Mamba and the hybrid contrastive learning paradigm operate in two parallel branches within Manta, enhancing Mamba for FSAR of long sub-sequence. Manta achieves new state-of-the-art performance on prominent benchmarks, including SSv2, Kinetics, UCF101, and HMDB51. Extensive empirical studies prove that Manta significantly improves FSAR of long sub-sequence from multiple perspectives. Wenbo Huang 0001, Jinghui Zhang 0001, Guang Li 0008, Lei Zhang 0130, Shuoyuan Wang, Fang Dong 0001, Jiahui Jin 0001, Takahiro Ogawa 0001, Miki Haseyama |
AAAI | 6 |
| 2025 | Bilateral Virtual Companions: The Impact of Virtual Humans' Movement and Voice Realism on User Perception and Experience in Multi-user VR CinemasabstractDespite the increasing prevalence of online social interaction, challenges such as insufficient immersion and lack of interactivity still persist. To overcome these limitations, this study developed a multi-user virtual reality (VR) cinema system with motion capture (Mocap) and multi-user VR technology. The system demonstrated strengths in overcoming physical space restrictions and saving travel costs, providing a more enriched interactive experience and fulfilling social needs under special circumstances, and facilitating metaverse applications. Furthermore, to give insight into the further design of multi-user VR cinemas, this study investigated the impact of virtual humans' (VHs') characteristics on user perception and experience of bilateral virtual companions, by evaluating the impact of movement and voice realism on immersion, social presence and intimacy. The results show that while neither movement nor voice realism significantly influences immersion, voice realism rather than movement realism significantly affects social presence and intimacy. Jingfeng Hu, Ding Ding 0002, Xiangyu Xu 0001, Jinghui Zhang 0001, Jiahui Jin 0001, Fang Dong 0001 |
CSCWD | 6 |
| 2025 | HiSpTRSV: Exploring Tile-Level Parallelism for SpTRSV Acceleration on FPGAsabstractSparse Triangular Solve (SpTRSV) is a critical level2 kernel in sparse Basic Linear Algebra Subprograms (BLAS). While Field-Programmable Gate Array (FPGA) accelerators for SpTRSV focus on optimizing individual tiles, they overlook intertile parallelism. Designing an inter-tile parallelism accelerator poses challenges, including constructing fine-grained dependency graph, handling communication overhead, and balancing workloads. HiSpTRSV addresses these challenges through dependency graph parsing, tile-based highly parallel algorithm, filtering mechanisms, and bidirectional matching with modular indexing. Experiments show that HiSpTRSV outperforms the state-of-the-art SpTRSV accelerator in terms of a 34.3% performance improvement. HiSpTRSV achieves a $3.58 \times$ speedup and $9.59 \times$ higher energy efficiency compared to GPUs. Fan Sun 0005, Fang Dong 0001, Dian Shen |
DAC | 2 |
| 2025 | FaSei: Fast Serverless Edge Inference with Synergistic Lazy Loading and Layer-wise Caching
Zhaowu Huang, Fang Dong 0001, Xiaolin Guo, Daheng Yin |
INFOCOM | 2 |
| 2025 | CoCaR: Enabling Efficient Dynamic DNN-Based Model Caching and Request Routing in MEC
Shuting Qiu, Fang Dong 0001, Siyu Tan, Dian Shen, Ruiting Zhou, Qilin Fan |
INFOCOM | 2 |
| 2025 | InfiniCL: Elastic Continual Learning for Resource-Constrained Edge DevicesabstractOn-device continual learning (CL) enables lifelong and privacy-preserving learning for various edge intelligent applications. Increasing the number of model parameters as new learning tasks emerge is effective in ensuring learning quality but inefficient in memory cost, especially for resource-constrained devices. In this paper, we introduce InfiniCL, the first ondevice CL system that dynamically balances memory cost and learning quality. A key idea behind InfiniCL is elastic continual learning: selectively freezing layers in the expanding model and periodically distilling the model, preventing unbounded memory growth while preserving learning quality for new tasks. This novel CL paradigm opens a new challenging problem: how to decide the memory allocation of the model and data to achieve better learning Quality of Service (QoS) under the limited memory budget? To alleviate this challenge, we further propose a Bayesian Optimization-driven algorithm to jointly optimize layer freezing selection and data-model memory allocation. Evaluations show that InfiniCL outperforms state-of-the-art methods on diverse memory constraints, achieving 5.34-7.15% and 2.72-9.36% higher accuracy on CIFAR-100 and ImageNet-100, respectively. Chenyu Lu, Mengyang Liu, Fang Dong 0001, Borui Li 0001, Ruiting Zhou, Shiyao Ji |
IWQoS | 3 |
| 2025 | Online Deployment of Dynamic Edge DAG Serverless Functions Toward Fast StartupabstractServerless computing converts services into multiple containerized functions, greatly enhancing the flexibility of edge applications. To guarantee Service Level Objectives (SLOs), it is common to pre-warm containers for all functions. However, we have observed that requests with a Directed Acyclic Graph (DAG) structure commonly involve only a subset of functions. Pre-warming all containers may lead to over-provisioning, and the fixed optimization strategy for pre-warming becomes ineffective in scenarios where the calling probabilities of functions dynamically change. To overcome the above challenges, we constructed a dynamic DAG model with the consideration of function calling probability, aiming to minimize the request execution time. Based on the model, we proposed an Expectation-based Rounding Optimization (EBRO) algorithm to progressively find the offline optimal container perwarm and deployment strategy with theoretical performance guarantee. Then, we implemented an Online Pre-warming and Tasks Scheduling (OPTS) algorithm to adjust pre-warming and deployment locations based on real-time edge resources and container states. Finally, extensive experiments on a edge cluster show that, compared with baselines, EBRO has the lowest average request execution time and cold start time, with reductions reaching 78.4% and 88.5% respectively. OPTS can potentially reduce the request execution time by up to 25.2%, and reduce the cold start time by a maximum of 80.1 %. Ruiting Zhou, Haodong Tian, Fang Dong 0001 |
IWQoS | 5 |
| 2025 | ADPTD: Adaptive Data Partition With Unbiased Task Dispatching for Video Analytics at the EdgeabstractRecently, edge-assisted methods have been proposed as a promising technique to deliver fast and accurate on-device video analytics by partitioning frame data and dispatching them to edge servers for parallel execution. However, the data partition (DP) reduces the detection latency but decreases accuracy since objects may cross the boundaries of adjacent blocks. The effect of DP on the accuracy and latency depends on multiple vital parameters (e.g., target size, density, network, and computing resources) in an unknown and time-varying fashion. Moreover, these parameters are determined by the application scenarios and edge environment, which are uncertain and heterogeneous at the edge. Hence, how to partition frames to strike a balance between accuracy and latency is a nontrivial and intractable problem. To this end, we propose an online learning-based device-edge–cloud collaboration framework, ADPTD, to guide DP at the edge. We propose an optimal task dispatching algorithm (OTD) to minimize detection latency. Then, we propose a multiarmed bandit-based algorithm to pick a DP strategy and invoke OTD to dispatch tasks in each time slot. Theoretical analysis reveals that ADPTD achieves sublinear regret. Extensive experimental results show that ADPTD outperforms the state-of-the-art methods, achieving a latency reduction of up to$2.53\times $and improving accuracy by up to 49.4%. Zhaowu Huang, Fang Dong 0001, Haopeng Zhu, Mengyang Liu, Dian Shen, Ruiting Zhou, Xiaolin Guo, Baijun Chen |
IEEE Internet Things J. | 2 |
| 2025 | Multi-Dimensional Training Optimization for Efficient Federated Synergy LearningabstractEdge learning (EL) is an end-to-edge collaborative learning paradigm enabling devices to participate in model training and data analysis, opening countless opportunities for edge intelligence. As a promising EL framework, federated synergy learning (FSyL) mitigates the computation and communication overhead on resource-constrained devices by offloading partial model layers to the edge server for synergistic training. Nevertheless, due to the system and statistical heterogeneity, naively using existing FSyL methods is significantly time-consuming and causes accuracy degradation. Motivated by this issue, this paper introduces a novel FSyL framework that integrates multi-dimensional training optimization and formulates the edge learning cost minimization (ELCM) problem. To tackle the ELCM efficiently, we designOL-MG, anOnLineModel Splitting and Resource ProvisioningGame. Specifically, we first reformulate and decompose the original ELCM based on data quality evaluation. Then, given a model splitting decision, we determine the optimal resource provisioning in Sub-problem1, based on which optimal model splitting in Sub-problem2 is modeled as a potential game. Subsequently, we introduce a decentralized algorithm to find a Nash equilibrium (NE) solution. Furthermore, we further extendOL-MGto support a budget-aware multi-edge scenario. Extensive experiments demonstrate that the proposed mechanism significantly outperforms state-of-the-art methods in cost-saving and accuracy improvement. Shucun Fu, Fang Dong 0001, Runze Chen 0001, Dian Shen, Jinghui Zhang 0001, Qiang He 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | Resource-Efficient DNN Inference With Early Exiting in Serverless Edge ComputingabstractServerless Edge Computing (SEC) has gained widespread adoption in improving resource utilization due to its triggered event-driven model. However, deploying deep neural network (DNN) inference services directly in SEC leads to resource inefficiencies, which stem from two key factors. First, existing methods adopt model-wise function encapsulation, which requires the entire DNN model to occupy memory throughout its execution lifecycle. This increases both memory footprint and occupancy time. Second, uniform DNN inference for diversity input leads to redundant computations and additional inference time. To this end, we propose REDI, a novel framework that leverages fine-grained block-wise function encapsulation and progressive inference to provide resource-efficient DNN inference while ensuring latency requirements. REDI enables the release of memory from already inferred shallow networks and allows each request to exit early based on input data complexity, eliminating redundant computations. To fully unleash the potential, REDI jointly considers resource heterogeneity, data diversity, and environment dynamics to investigate the block-wise function placement problem. We introduce an uncertainty-aware online learning-driven algorithm with bounded regret. Finally, we conduct extensive trace-driven experiments to evaluate our methods, demonstrating that REDI achieves a significant speedup of up to$6.52\times$in terms of resource usage cost compared to state-of-the-art methods. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang, Jinghui Zhang 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | DiG-In-GNN: Discriminative Feature Guided GNN-Based Fraud Detector against Inconsistencies in Multi-Relation Fraud GraphabstractFraud detection on multi-relation graphs aims to identify fraudsters in graphs. Graph Neural Network (GNN) models leverage graph structures to pass messages from neighbors to the target nodes, thereby enriching the representations of those target nodes. However, feature and structural inconsistency in the graph, owing to fraudsters' camouflage behaviors, diminish the suspiciousness of fraud nodes which hinders the effectiveness of GNN-based models. In this work, we propose DiG-In-GNN, Discriminative Feature Guided GNN against Inconsistency, to dig into graphs for fraudsters. Specifically, we use multi-scale contrastive learning from the perspective of the neighborhood subgraph where the target node is located to generate guidance nodes to cope with the feature inconsistency. Then, guided by the guidance nodes, we conduct fine-grained neighbor selection through reinforcement learning for each neighbor node to precisely filter nodes that can enhance the message passing and therefore alleviate structural inconsistency. Finally, the two modules are integrated together to obtain discriminable representations of the nodes. Experiments on three fraud detection datasets demonstrate the superiority of the proposed method DiG-In-GNN, which obtains up to 20.73% improvement over previous state-of-the-art methods. Our code can be found at https://github.com/GraphBerry/DiG-In-GNN. Jinghui Zhang 0001, Zhengjia Xu, Dingyang Lyu 0001, Dian Shen, Jiahui Jin 0001, Fang Dong 0001 |
AAAI | 7 |
| 2024 | Rendering Super Resolution Video Streaming Efficiently with in-Network ComputingabstractEmerging live video streaming applications, e.g., Ultra High Definition videos and interactive video streaming, have put forward new demands for ultra-high bandwidth and reduced delay to match the desired quality of experience. Since current on-device Super-Resolution (SR) approaches are hindered by the limited end-device capabilities, we are motivated to take advantage of Mobile Edge Computing (MEC) and Computing in the Network technologies, such that SR videos can be processed on a more powerful infrastructure by integrating the resources from end-devices through the edge and all the way to the cloud. However, the integration of SR and MEC is non-trivial due to the challenges introduced by the features of SR tasks, and the heterogeneous nature of MEC resources. In this paper, we endeavor to explore and solve these challenges by presenting AVSA, which renders SR live video streaming efficiently with in-network computing. AVSA can adaptively allocate SR work-loads under heterogeneous resources and yield a cost-effective workload allocation with theoretical performance guarantees. Simulation results show that, compared with the state-of-the-art methods, our method achieves up to 13.06× speedup in terms of end-to-end latency. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang, Baijun Chen, Daheng Yin |
HPCC | 2 |
| 2024 | SkewCache: Skewed Layer-wise Caching for Function Chains in Serverless Edge ComputingabstractIn serverless edge computing (SEC), traditional monolithic applications are encapsulated in multiple dependent functions, in which event-driven service provisioning introduces the cold start problem. Existing works adopt uniform container caching for each function to mitigate the cold start problem, which remains with relatively low efficiency due to ignoring the skewed invocation frequency and the heterogeneity of cold-start behaviors of functions. In this paper, we propose SkewCache, an efficient layer-wise container caching framework for frequency-skewed function chains in SEC. Based on the container’s layer structure, SkewCache enables fine-grained and balanced container caching on multiple edge servers, which takes into account both the invocation frequency skewness and the cold start latency of each function. The problem is modeled as an integer nonlinear programming (INLP) to minimize the application completion time (ACT). To solve the INLP, we first convert it to an equivalent integer linear programming (ILP) form. Then, we propose an approximation algorithm to solve the ILP with a guaranteed approximation ratio. To evaluate the performance of the proposed algorithm, we conduct intensive simulations and the results show that our algorithms outperform baselines, achieving up to 1.94× speedup in terms of ACT reduction. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang, Haodong Tian |
HPCC | 2 |
| 2024 | Joint Layer-wise Caching and Request Routing for Serverless Inference Acceleration at the EdgeabstractIn recent years, the serverless paradigm has been introduced into edge computing, enhancing resource utilization. However, it has also introduced the issue of cold start, which hinders low-latency responses. This cold start problem is significantly exacerbated when it comes to Artificial Intelligence (AI) applications, due to the necessity of loading bulky deep learning (DL) model files. Existing methods typically address the cold start issue by caching complete containers, which is inefficient in resource-constrained edge servers. This is because massive DL models may not fit in the limited cache space or lead to fragmentation of the cache memory. To solve the above problem, we seek a layer-wise model caching method that regards the DL model as a chain structure of multiple layers and caches a part of the model layers. However, heterogeneity in cold starts and edge computing’s complex environments pose challenges in selecting model layers and servers. In this paper, we investigate the joint layer-wise caching and request routing problem to minimize application completion time (ACT). We propose an approximation algorithm with a provable performance guarantee. Finally, intensive trace-driven simulations show that our algorithm achieves 3.03× speedup in terms of ACT reduction. Zhaowu Huang, Fang Dong 0001, Xiaolin Guo, Haodong Tian |
HPCC | 2 |
| 2024 | HASFL: Harnessing Heterogeneous Models Across Diverse Devices for Enhanced Federated LearningabstractRecent advancements in federated learning have shown promising results in resource-constrained edge environments. However, with mobile devices becoming more capable of collecting data, individual client models are unable to utilize the available data due to their devices’ limited support for complex model training. Conversely, non-portable devices possess substantial computational resources, but the data they autonomously collect is insufficient to support the training of complex models. In this paper, we introduce HASFL, a novel split federated learning (SFL) framework that supports model structure heterogeneity across devices and decouples computation from the model. Through circular group training, HASFL enables mobile devices to utilize complex models to train their own data while ensuring that non-portable devices harness the data collected by mobile users. HASFL effectively addresses the challenges of applying advanced machine learning models in resource-constrained environments, leveraging the collective power of distributed devices without compromising data security. We implemented a circular group allocation method using the online algorithm to ensure cooperative training among heterogeneous models within each group while minimizing training time. In addition, we have conducted experiments to evaluate the performance of HASFL on various datasets and model architectures and analyzed the communication overhead of HASFL. The experimental results demonstrate that HASFL supports the training of heterogeneous models and significantly enhances the model’s accuracy with a relatively small increase in communication overhead. Jiangshan Hao, Fang Dong 0001, Bingheng Cen, Shucun Fu, Ruiting Zhou, Ding Ding 0002 |
ICPP | 2 |
| 2024 | FSVFG: Towards Immersive Full-Scene Volumetric Video Streaming with Adaptive Feature GridabstractGiven the truly immersive viewing experiences, full-scene volumetric videos have received increasing attention from both academia and industry. Their vast data volumes, however, present significant challenges for real-time streaming over today's bandwidth-limited Internet. Considering the vast amount of full-scene volumetric data to be streamed and the limited bandwidth on the Internet, achieving adaptive full-scene volumetric video streaming over the Internet presents a significant challenge. Inspired by the advantages offered by neural fields, especially the feature grid method, we propose FSVFG, a novel full-scene volumetric video streaming system integrated feature grids as the representation of volumetric content. FSVFG employs an incremental training approach for feature grids and stores the features and residuals between adjacent grids as frames. To support adaptive streaming, we delve into the data structure and rendering processes of feature grids and propose bandwidth adaptation mechanisms. The mechanisms involve a coarse ray-marching for the selection of features and residuals to be sent, and achieve variable bitrate streaming by Level-of-Detail (LoD) and residual filtering. Based on these mechanisms, FSVFG achieves adaptive streaming by adaptively balancing the transmission of feature and residual according to the available bandwidth. Our preliminary results demonstrate the effectiveness of FSVFG, demonstrating its ability to improve visual quality and reduce bandwidth requirements of full-scene volumetric video streaming. Daheng Yin, Jianxin Shi 0005, Miao Zhang 0003, Zhaowu Huang, Jiangchuan Liu, Fang Dong 0001 |
ACM Multimedia | 6 |
| 2024 | B2-Bandit: Budgeted Pricing With Blocking Constraints for Metaverse Crowdsensing Under UncertaintyabstractMetaverse has been viewed as the next generation of human-computer interaction, which requires collecting information from both the physical and virtual world. One potential way is to employ virtual service providers (VSPs) to finish collection tasks by designing posted-pricing mechanisms via the crowdsensing platform. As VSPs’ costs and values are usually unknown, learning the optimal posted-pricing policy under uncertainty is undoubtedly critical to utilize the budget efficiently. However, existing posted-pricing learning algorithms assume that agents provide services without blocking and agents’ attributes follow an independent identical distribution, both of which are unrealistic in Metaverse, e.g., VSPs should continuously sense the physical world to make provided services realistic, which makes the long working VSP unavailable/blocked for a certain period of time. In this paper, we address the budgeted pricing problem under uncertainty by considering blocking constraints and unknown non-identical VSPs’ attributes. The problem is modeled as a Budgeted-pricing Blocking Bandit (B2-bandit) problem, which remains unaddressed even for the oracle case with known VSPs’ information. We thus first propose a pricing policy for the oracle case with an instance-dependent approximation ratio to the global optimum. For the general B2-bandit problem with unknown information, we propose an online learning algorithm satisfying blocking constraints and incurring an accumulated regret up to$O(MK\log B)$as compared to the oracle approximation algorithm, where$M,K,B$are the number of VSPs, candidate prices and the budget, respectively. Experiments on real datasets validate that the proposed algorithm improves more than 172% accumulated value compared to baseline pricing algorithms. Xiang Liu 0014, Weiwei Wu 0001, Chenchen Fu, Fang Dong 0001, Junzhou Luo |
IEEE J. Sel. Areas Commun. | 6 |
| 2024 | Optimal Harvest-Then-Transmit Scheduling for Throughput Maximization in Time-Varying RF Powered SystemsabstractEnergy harvesting is a promising technique to address the energy hunger problem for thousands of wireless devices. In Radio Frequency (RF) energy harvesting systems, a wireless device first harvests energy and then transmits data with this energy, hence the ‘harvest-then-transmit’ (HTT) principle is widely adopted. We must carefully design the HTT schedule, i.e., schedule the timing between harvesting and transmission, and decide the data transmission power such that the throughput can be maximized with the limited harvested energy. Distinct from existing work, we assume energy harvested from RF sources is time-varying, which is more practical but more difficult to handle. We first discover a surprising result that the optimal transmission power is independent of the transmission time, but solely depends on the RF harvesting power, for a simple case when the energy harvesting is stable. We then obtain an optimal offline HTT-scheduling for the general case that allows the RF harvesting power to vary with time. To the best of our knowledge, it is the first optimal HTT-scheduling algorithm that achieves maximum data throughput for time-varying RF powered systems. Finally, an efficient online heuristic algorithm is designed based on the offline optimality properties. Simulations show that the proposed online algorithm has superior performance, which achieves more than 90% of the offline maximum throughput in most cases. Feng Shan, Junzhou Luo, Qiao Jin 0003, Liwen Cao, Weiwei Wu 0001, Zhen Ling 0001, Fang Dong 0001 |
IEEE J. Sel. Areas Commun. | 7 |
| 2024 | Privacy-preserving model splitting and quality-aware device association for federated edge learningabstractAbstract Federated edge learning (FEEL) provides a promising device‐edge collaborative learning paradigm, which enables edge devices to parallel participate in model co‐creation while preserving user privacy, opening countless opportunities to enable edge intelligence. With the growing demand for intelligent services, extensive FEEL deployment is inevitable. Nevertheless, existing FL schemes neglect two unique features (i.e., resource heterogeneity and data heterogeneity) in real‐world edge learning and thus may negatively affect the training efficiency and accuracy. Specifically, (1) heterogeneous and limited device resources cause massive laggards, which bring intolerable training delay; (2) heterogeneous data distribution causes device quality divergence, bringing severe training accuracy degradation. This article proposes a split‐based FEEL framework and an adaptive model splitting and quality‐aware device association scheme (MSDA) to tackle the aforementioned challenges. MSDA contains two levels: at the model splitting level, according to device capability and model structure, an adaptive splitting mechanism is proposed to provide a low‐latency and privacy‐preserving model splitting strategy for each device and guide subsequent device association. At the device association level, each device is simulated as a player with a quality weight in the potential game. Then a quality‐aware decentralized device association mechanism is designed to ensure that more high‐quality devices upload local updates before the deadline with the help of the edge server. Finally, experimental results demonstrate that MSDA yields significant improvements, achieving up to 3.1 training speedup and 39% accuracy improvement compared to state‐of‐the‐art methods. Shucun Fu, Fang Dong 0001, Dian Shen, Tianyang Lu |
Softw. Pract. Exp. | 2 |
| 2024 | DESIGN: Online Device Selection and Edge Association for Federated Synergy Learning-enabled AIoTabstractThe artificial intelligence of things (AIoT) is an emerging technology that enables numerous AIoT devices to participate in big data analytics and machine learning (ML) model training, providing various customized intelligent services for industry manufacturing. Federated learning (FL) empowers AIoT applications with privacy-preserving distributed model training without sharing raw data. However, due to IoT devices’ limited computing and memory resources, existing FL approaches for AIoT applications cannot support efficient large-scale model training. Federated synergy learning (FSyL) is a promising collaborative paradigm that alleviates the computation and communication overhead on resource-constrained AIoT devices via offloading part of the ML model to the edge server for end-to-edge collaborative training. Existing FSyL works neither efficiently address the inter-round device selection to improve model diversity nor determine the intra-round edge association to reduce the training cost, which hinders the applications of FSyL-enable AIoT. Motivated by this issue, this article first investigates the bottlenecks of executing FSyL in AIoT. It builds an optimization model of joint inter-round device selection and intra-round edge association for balancing model diversity and training cost. To tackle the intractable coupling problem, we present a framework named Online DEvice SelectIon and EdGe AssociatioN for Cost-Diversity Tradeoffs FSyL (DESIGN). First, the edge association subproblem is extracted from the original problem, and game theory determines the optimal association decision for an arbitrary device selection. Then, based on the optimal association decision, device selection is modeled as a combinatorial multi-armed bandit (CMAB) problem. Finally, we propose an online mechanism to obtain joint DESIGN decisions. The performance of DESIGN is theoretically analyzed and experimentally evaluated on real-world datasets. The results show that DESIGN can achieve up to \(84.3\%\) in cost-saving with an accuracy improvement of \(23.6\%\) compared with the state-of-the-art. Shucun Fu, Fang Dong 0001, Dian Shen, Runze Chen 0001, Jiangshan Hao |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2024 | eMPTCP: A Framework to Fully Extend Multipath TCPabstractMPTCP provides the basic multipath support for network applications to deliver high throughput and robust communication. However, the original MPTCP is designed with limited extensibility. Various research works have tried to extend MPTCP to attain better performance or richer functionalities. These existing approaches either modify the kernel implementation of MPTCP, which involve considerable engineering efforts and may accidentally introduce safety issues, or control MPTCP via userspace tools, which suffer from restricted functionality support. To address this issue, we propose eMPTCP, an easy-to-use framework to fully extend MPTCP without safety risks. Internally, eMPTCP has a modular and pluggable model which allows operators to specify a comprehensive MPTCP extension as a chain of sub-policies. eMPTCP further enforces the policies through packet header manipulations. To ensure safety, eMPTCP is implemented using eBPF. Despite the stringent constraints of eBPF, we show that it is possible to implement an elaborated framework for a fully extensible MPTCP. Through verifying MPTCP in a number of real-world cases and extensive experiments, we show that eMPTCP is able to support a wide range of MPTCP extensions, while the overhead of eMPTCP operations in the kernel is in the scale of nanosecond, and the extra processing time accounts for only about 0.63% of flows’ transmission time. Dian Shen, Bin Yang 0027, Junxue Zhang 0001, Fang Dong 0001, John C. S. Lui |
IEEE/ACM Trans. Netw. | 4 |
| 2024 | Addressing Heterogeneity in Federated Learning with Client Selection via Submodular OptimizationabstractFederated learning (FL) has been proposed as a privacy-preserving distributed learning paradigm, which differs from traditional distributed learning in two main aspects: the systems heterogeneity, meaning that clients participating in training have significant differences in systems performance including CPU frequency, dataset size, and transmission power, and the statistical heterogeneity, indicating that the data distribution among clients exhibits Non-Independent Identical Distribution. Therefore, the random selection of clients will significantly reduce the training efficiency of FL. In this article, we propose a client selection mechanism considering both systems and statistical heterogeneity, which aims to improve the time-to-accuracy performance by trading off the impact of systems performance differences and data distribution differences among the clients on training efficiency. First, client selection is formulated as a combinatorial optimization problem that jointly optimizes systems and statistical performance. Then, we generalize it to a submodular maximization problem with knapsack constraint, and propose the Iterative Greedy with Partial Enumeration (IGPE) algorithm to greedily select the suitable clients. Then, the approximation ratio of IGPE is analyzed theoretically. Extensive experiments verify that the time-to-accuracy performance of the IGPE algorithm outperforms other compared algorithms in a variety of heterogeneous environments. Jinghui Zhang 0001, Fa Xin, Fang Dong 0001, Junzhou Luo |
ACM Trans. Sens. Networks | 5 |
| 2024 | Joint Optimization of Device Selection and Resource Allocation for Multiple Federations in Federated Edge LearningabstractFederated edge learning (FEEL) is a promising collaborative paradigm, which employs edge devices (EDs) to train machine learning models for a federation. It opens countless opportunities to enable edge intelligence. The increasingly diversified demands for intelligent services are driving the deployment of various federations at the edge. Existing works on FEEL focus on a single federation and ignore inter-federation device competition and intra-device resource allocation, which hinders the applications of FEEL. To address this issue, this article first investigates the bottlenecks of executing multiple federations and builds a joint optimization model as a two-stage Stackelberg game involving device selection and resource allocation. To tackle the problem efficiently, we present a game-theoretical approach namedDeviceSelection andResourceAllocation forMultipleFederationsGame (DSRAMF-G). First, following the arbitrary device selection of leaders (i.e., federations), the time cost minimization of followers (i.e., EDs) is modeled as a convex problem to obtain the optimal resource allocation. Then, based on followers’ optimal responses, device selection is modeled as a congestion game. We prove the existence of the Nash equilibrium and propose a decentralized mechanism. Finally, extensive experiments show that DSRAMF-G significantly outperforms the state-of-the-art methods, achieving up to 5.9x training speedup and 2.8x resource-savings. Shucun Fu, Fang Dong 0001, Dian Shen, Jinghui Zhang 0001, Zhaowu Huang, Qiang He 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2023 | HyperRTV: Neural-Enhanced Adaptive Real-Time Video Streaming Based on Terminal-Edge CollaborationabstractVideo has become the primary source of Internet traffic due to the advance of streaming media technology and the surge in user demand for real-time video streaming applications. In this case, neural-enhanced video streaming and mobile edge computing are proposed to improve video content quality on nearby MEC (Mobile Edge Computing) devices and to maintain low interaction delay under limited bandwidth. In order to maximize the architectural advantages brought by MEC, we designed HyperRTV, a terminal-edge collaborative real-time video transmission system with neural-enhanced streaming. HyperRTV consists of three key components: 1) a video super-resolution structure that uses hardware-accelerated DNNs to satisfy latency limits; 2) an adaptive bitrate controller that can dynamically adjust the bitrate under various network conditions; 3) a heuristic task offloading strategy based on the ant colony algorithm to further meet the heterogeneity and dynamics of the MEC environment. Experiments show that HyperRTV decreases bandwidth consumption by 46% on average and achieves a 48%-65% reduction in latency, consistently maintaining better visual quality and higher stability in unpredictable networks. Our task offloading strategy can reduce the execution and transfer time by 16.1%-19.5% in a MEC environment with multi-task requests. Baijun Chen, Daheng Yin, Lifei Teng, Fang Dong 0001 |
CSCWD | 4 |
| 2023 | Accelerate Multi-view Inference with End-edge Collaborative ComputingabstractMulti-view inference can utilize visual information from several views like a human being and significantly improve accuracy in some scenes, but it inevitably incurs more computing overhead than traditional DNN inference. To meet the requirement of low latency in typical scenarios, we consider utilizing model partition technique of edge computing to speed up multi-view inference, and design a multi-view end-edge co-inference execution framework (MV-IEF) which can make use of both end and edge resources for multi-view inference tasks. However, when employing the framework simply, the efficiency of multi-view inference will be constrained by network dynamics and heterogeneity of devices corresponding to multiple views. To break this constraint, we establish an optimization model based on the framework to minimize the multi-view inference time and solve it on the basis of game theory. And meanwhile, we propose a joint optimization algorithm for multi-view resource allocation and model partition (MV-JRAMP), which can make remarkable decisions of resource allocation and model partiton according to network status and computing capabilities of devices. Finally, we build a prototype and evaluate the performance of MV-JRAMP. Experiments show that MV-JRAMP can accelerate multi-view inference by up to 3.71×. Wangbing Cheng, MinFeng Zhang, Fang Dong 0001, Shucun Fu |
CSCWD | 3 |
| 2023 | MIA-FedDL: A Membership Inference Attack against Federated Distillation LearningabstractFederated learning, which gathers the model parameters of numerous IoT devices without accessing user data, is used to train high-quality models and enhance the quantity and quality of available data. Federated distillation learning strategies are suggested as a means of addressing the difficulties associated with federated learning in Non-IID environment. However, recent research has demonstrated that federated learning falls short of providing complete privacy protection and is still open to malevolent attackers’ inference attacks. In this paper, we investigate a malicious attacker’s membership inference attack in a federated distillation learning context. We initially concentrate on federated distillation learning in the Non-IID environment, and then we propose a membership inference attack technique started by a malicious client that can successfully infer the privacy of other clients without getting more information. Comprehensive experimental results show the effectiveness of our MIA-FedDL and quantify user privacy leakage in federated distillation learning. Fang Dong 0001 |
CSCWD | 2 |
| 2023 | FreezePipe: An Efficient Dynamic Pipeline Parallel Approach Based on Freezing Mechanism for Distributed DNN TrainingabstractDeep Neural Network (DNN) training on a large scale is extremely time-consuming and computationally intensive, which is accelerated by distributed training. In recent years, pipeline parallelism has been developed, which enables partitioning the model across several devices, e.g. GPU, and training efficiency is improved by dividing data batches into micro-batches, with each of them processed by a different stage of the model. Currently, parallel training assumes pipeline placement and partitioning are static, with parameters updating each iteration, without accounting for freezing. This results in computational resources not being fully utilized. In this paper, we propose FreezePipe, a novel method for optimizing deep learning training that combines the freezing mechanism with pipeline parallel training. In FreezePipe, a lightweight method for determining the freezing strategy based on gradient changes is employed. Considering that resources need to be released based on the frozen layer, a lightweight model partitioning algorithm was designed to determine the optimal strategy for pipeline partitioning. Experimental results show that FreezePipe can reduce the training time by 64.5% compared to Torchgpipe on CIFAR-10 dataset without compromising any model performance. Caishan Weng, Zhiyang Shu, Zhengjia Xu, Jinghui Zhang 0001, Junzhou Luo, Fang Dong 0001, Zhengang Wang |
CSCWD | 6 |
| 2023 | Joint Optimization of Task Offloading and Resource Allocation for Edge Video AnalyticsabstractWith the development of artificial intelligence technology and intelligent devices, people show great interest in intelligent applications and services, but it is impossible to complete these compute-intensive AI tasks locally, especially video analysis tasks. Edge computing is regarded as an appropriate solution to these problems. In this paper, we study the multi-user multi-server edge-end collaboration video analytics task offloading problem aiming at minimizing the overall delay for each device to finish its task. Each device chooses whether to execute the task locally or to offload the task to an edge server, and which edge server to select. At the theoretical level, we model the joint problem of task offloading and resource allocation as a mixed integer programming problem. We first determine the optimal resource allocation policy with a given task offloading decision profile. Then, task offloading problem is modeled as a congestion game and propose a decentralized mechanism to achieve a Nash equilibrium. Moreover, experimental results demonstrate that the proposed method is efficient and can significantly and steadily improve the system performance, reducing the overall delay by 33.96% on average, compared with other algorithms. Zhenxuan Xu, Yunzhou Xie, Fang Dong 0001, Shucun Fu, Jiangshan Hao |
CSCWD | 3 |
| 2023 | CoCV: Heterogeneous Processors Collaboration Mechanism for End-to-End Execution of Intelligent Computer Vision Tasks on Mobile DevicesabstractObject detection, image classification, and various other computer vision tasks have become prevalent on mobile devices. These computer vision tasks are typically executed with three stages: pre-processing, inference, and post-processing. Mobile SoC, serving as the computing unit on mobile devices, typically consists of heterogeneous processors like CPU, GPU, and NPU. However, during the execution of a computer vision task, current available frameworks only achieve the parallelism of CPU and GPU in the inference stage. While during pre- and post-processing, only CPU is used, leaving GPU and NPU on the SoC to be idle. For mobile applications that require low latency, the overhead of pre-processing and post-processing stages often account for more than 50% of the total latency, which becoming a performance bottleneck of the entire task. To reduce latency, it is imperative to fully utilize the idle heterogeneous processors (GPU, NPU) on the SoC and achieve heterogeneous processors parallelism in all three stages during execution. In this paper, we propose CoCV, a heterogeneous processor parallel computing system for computer vision tasks on mobile devices. In CoCV, we are the first to build an image processing operator library for heterogeneous parallel computing on mobile devices. Besides, we design a task allocation scheduling algorithm to guide the partitioning of processing tasks during execution, which ensuring a relatively balanced workload between different processors. A cross-stage operator chaining technique is also proposed to reduce the data sharing overhead among different processors during execution. We build a prototype system and evaluate it with different computer vision tasks. The results show up to 33% latency reduction for end-to-end tasks and 2.32× speedup compared with the current best solution. Ye Wan, Mengyang Liu, Guangtong Li, Fang Dong 0001 |
ICPADS | 4 |
| 2023 | Adaptive Overlap Padding and Resolution Selection for Frame Split-based Edge Video AnalyticsabstractFor providing accurate and fast on-device high-resolution video analytics, edge-assisted methods are widely proposed by using lower-resolution frames and splitting them with overlap padding. The use of low-resolution frames can significantly reduce the computational workload of video analytics. Dividing a frame into multiple overlapping parts simultaneously improves parallelism and maintains processing accuracy. However, accuracy and latency serve as a pair of tradeoff metrics, and prioritizing one to optimize overlap padding or resolution selection will compromise the other metric. Fortunately, we have discovered that in real-world video analytics scenarios with varying object sizes, it is not necessary to simultaneously achieve a high overlap padding size and high resolution. Hence, how to set appropriate overlap padding size and resolution to strike a balance between accuracy and latency in practical scenarios is a nontrivial and intractable problem. To this end, we propose an online learning-based method to achieve adaptive overlap padding and resolution selection, called APR. We model the problem as an integer programming and propose a Muli-armed bandit (MAB) theory-based algorithm to solve it. We discretize the continuum overlap padding size into a finite set to narrow explore space and set the frame split strategy as context information to achieve fast convergence. Theoretical analysis reveals APR achieves sub-linear regret. Extensive experimental results show APR outperforms the benchmark methods, achieving up to 2.06 × speedup in terms of latency and 0.18× increase in accuracy. Haopeng Zhu, Zhaowu Huang, Xiaolin Guo, Mengyang Liu, Baijun Chen, Fang Dong 0001 |
ICPADS | 6 |
| 2023 | ASFL: Adaptive Semi-asynchronous Federated Learning for Balancing Model Accuracy and Total Latency in Mobile Edge NetworksabstractFederated learning (FL) is a new paradigm for privacy-preserving learning. This is particularly appealing in the mobile edge network (MEN), in which devices collectively train a global model with their own set of data. It is, however, routinely difficult for FL algorithms to satisfy different training task preferences in terms of the total latency and model accuracy due to a number of factors including the straggler effect, data heterogeneity, communication bottleneck and device mobility. To this end, we propose an Adaptive Semi-asynchronous Federated Learning (ASFL) framework, which adaptively balances the total latency and model accuracy according to the task preferences in MEN. Specifically, ASFL conducts a two-stage operation: i) Device selection stage. Each global round selects a set of devices that can maximize the model accuracy to eliminate data heterogeneity and communication bottlenecks; ii) Training stage. We first define a latency-accuracy objective value to model the balance between the latency and accuracy. Then in each global round, we use a deep reinforcement learning (DRL) algorithm based on soft actor-critic with discrete actions to intelligently derive the number of picked devices (i.e., participants in the current global aggregation) and the lag tolerance at each global round to maximize the latency-accuracy objective value. Extensive experiments show that ASFL can improve the latency-accuracy objective value by up to 94% compared with three state-of-the-art FL frameworks. Jieling Yu, Ruiting Zhou, Chen Chen 0067, Bo Li 0001, Fang Dong 0001 |
ICPP | 5 |
| 2023 | WAEVSR: Enabling Collaborative Live Video Super-Resolution in Wide-Area MEC EnvironmentabstractLive video streaming is increasingly popular for its rich content and real-time interactions, but its demand for bandwidth has put a heavy burden on backbone networks. To save bandwidth, recent studies have proposed neural-enhanced live video streaming that deploys deep neural networks (DNNs) for video super-resolution (VSR) on end devices or nearby edge devices to enhance video quality by taking low-resolution frames as input and producing high-resolution output frames. In this solution, the high computational demands of high-quality VSR DNNs make them difficult to support on single end or edge device, necessitating the use of distributed resources in edge facilities. However, the distributed deployment of high-quality VSR DNNs for low-latency inference remains challenging due to the inherent data dependencies of VSR DNNs and the heterogeneity and dynamics of edge facilities. In this paper, we present WAEVSR, a novel collaborative neural-enhanced live video super-resolution system that enables effective leverage of distributed resources to maximize the latency-bounded quality in wide-area MEC environments. WAEVSR consists of two key components: 1) It deploys a parallel-friendly video super-resolution DNN among edge devices, 2) with an inference controller based on the variable-size sliding window to balance the latency and quality of distributed inference in the heterogeneous and dynamics MEC environment. Prototype-based evaluation shows that WAEVSR can achieve 2.5 × lower end-to-end latency than traditional super-resolution serving with a 0.01 drop in SSIM score. The case study also demonstrates its higher stability on latency than vanilla distributed MEC deployment. Daheng Yin, Fang Dong 0001, Baijun Chen, Dian Shen, Ruiting Zhou, Xiaolin Guo, Zhaowu Huang |
IWQoS | 2 |
| 2023 | ROIAdaptor: Adaptive Task Offloading of ROI-Encoded Videos for Edge Video AnalyticsabstractReal-time analytics on video data demands intensive computation resources and high bandwidth consumption. Edge computing enables us to offload resource-intensive analytics tasks to nearby edge servers, effectively reducing the extended latency. Numerous studies have applied ROI encoding technology to decrease video data size, thereby decreasing transmission latency. However, ROI encoding will reduce inference accuracy and indirectly affect inference latency. Existing works have ignored these impacts, leading to significant performance degradation. In this paper, we first identified the necessity of considering video QP settings when offloading, and investigated the impacts of changing QP settings on accuracy and latency to demonstrate that adaptive offloading based on QP settings can improve the efficiency of video analytics. We formulate the problem of minimizing average latency to meet real-time requirements under accuracy constraints, and propose an online offloading algorithm called ROIAdaptor, based on a contextual multi-armed bandit method. Our algorithm is developed based on LinUCB, a contextual multi-armed bandit method, and operates online with historical information, achieving a provable performance bound. Simulation results show that ROIAdaptor can reduce the overall latency by an average of 35.7% and meet the requirements for real-time video analysis, with virtually no loss in accuracy. Zhenxuan Xu, Xiaolin Guo, Zhaowu Huang, Shucun Fu, Fang Dong 0001 |
MSN | 5 |
| 2023 | Joint Quality Evaluation, Model Splitting and Resource Provisioning for Split Edge LearningabstractEdge learning (EL) is an end-edge collaborative learning paradigm that enables numerous edge devices to participate in model training and data analysis, opening countless opportunities to enable edge intelligence. As is a promising EL approach, split edge learning (SPEL) alleviates the computation and communication overhead on resource-constrained devices via offloading part of the machine learning (ML) model to the edge server for cooperative training. Nevertheless, due to the system and statistical heterogeneity of the edge environment, naively using existing SPEL methods brings significantly time-consuming and accuracy degradation. Specifically, system heterogeneity causes intolerable time costs in each training round, while statistical heterogeneity further results in weight divergence and more training rounds to achieve global convergence. Motivated by this issue, this paper designs an efficient SPEL scheme to minimize the total time cost of participating devices. Specifically, we propose a novel SPEL framework and formulate the edge learning cost minimization (ELCM) problem that involves jointly optimizing model splitting and resource provisioning. We design OL-MG, i.e., OnLine Model Splitting and Resource Provisioning Game scheme, to solve the ELCM problem. In OL-MG, we first transform and decompose the original ELCM into two subproblems based on data quality evaluation. Second, we determine the optimal resource provisioning of Sub-problem1 with a given model splitting decision, based on which optimal model splitting of Sub-problem2 is modeled as a potential game. Then, we propose a decentralized algorithm to find a Nash equilibrium (NE) solution for the ELCM problem. Experimental results from both hardware prototype and simulation demonstrate that OL-MG outperforms the state-of-the-art methods, achieving up to 3.1x training cost savings and 40% accuracy improvement. Shucun Fu, Fang Dong 0001, Dian Shen, Qiang He 0001 |
SECON | 2 |
| 2023 | Label Information Enhanced Fraud Detection against Low Homophily in GraphsabstractNode classification is a substantial problem in graph-based fraud detection. Many existing works adopt Graph Neural Networks (GNNs) to enhance fraud detectors. While promising, currently most GNN-based fraud detectors fail to generalize to the low homophily setting. Besides, label utilization has been proved to be significant factor for node classification problem. But we find they are less effective in fraud detection tasks due to the low homophily in graphs. In this work, we propose GAGA, a novel Group AGgregation enhanced TrAnsformer, to tackle the above challenges. Specifically, the group aggregation provides a portable method to cope with the low homophily issue. Such an aggregation explicitly integrates the label information to generate distinguishable neighborhood information. Along with group aggregation, an attempt towards end-to-end trainable group encoding is proposed which augments the original feature space with the class labels. Meanwhile, we devise two additional learnable encodings to recognize the structural and relational context. Then, we combine the group aggregation and the learnable encodings into a Transformer encoder to capture the semantic information. Experimental results clearly show that GAGA outperforms other competitive graph-based fraud detectors by up to 24.39% on two trending public datasets and a real-world industrial dataset from Baidu. Even more, the group aggregation is demonstrated to outperform other label utilization methods (e.g., C&S, BoT/UniMP) in the low homophily setting. Jinghui Zhang 0001, Zhengjie Huang, Weibin Li 0004, Shikun Feng, Ziheng Ma, Yu Sun 0029, Dianhai Yu, Fang Dong 0001, Jiahui Jin 0001, Beilun Wang, Junzhou Luo |
WWW | 9 |
| 2023 | DeepMetricCorr: Fast flow correlation for data center networks with deep metric learning
Zunyi Liu, Dian Shen, Jiaang Bao, Fang Dong 0001, Jiong You |
Comput. Networks | 4 |
| 2023 | PipePar: Enabling fast DNN pipeline parallel training in heterogeneous GPU clusters
Jinghui Zhang 0001, Geng Niu, Qiangsheng Dai, Fang Dong 0001, Zhiang Wu 0001 |
Neurocomputing | 6 |
| 2023 | Deep-Reinforcement-Learning-Based Production Scheduling in Industrial Internet of ThingsabstractThe unprecedented prosperity of the Industrial Internet of Things (IIoT) promotes the traditional industry transforming into intelligent manufacturing so that the whole production process can be comprehensively controlled to achieve flexible production. Intelligent scheduling, as one of the key enabling techniques, is desired to allocate the production of several machines by an efficient solution with minimum makespan. Existing approaches adopt a fixed search paradigm based on expert knowledge to seek satisfactory solutions. However, considering the varying data distribution and large sized of the practical problems, these methods fail to guarantee the quality of the obtained solution under the real-time requirement. To address this challenge, we formulate the production scheduling problem as a Markov decision process (MDP) and specifically design a job scheduling model made up of a job batching module for the hybrid flow-shop scheduling problem on batch processing machines (HFSP-BPM). Our proposed model consists of an actor network that learns the action under different conditions and a critic network that evaluates the action of the actor. We analyze the convergence of the model under different parameter settings to determine the optimal parameter. Extensive numerical experiments on both publicly available data set and real steel plant production data set demonstrate that the proposed deep reinforcement learning (DRL) approach compared with other baselines, more than 6% average improvements can be observed in many instances. Zihui Luo, Chengling Jiang, Liang Liu 0001, Xiaolong Zheng 0002, Huadong Ma, Fang Dong 0001, Fucun Li |
IEEE Internet Things J. | 6 |
| 2023 | PADP-FedMeta: A personalized and adaptive differentially private federated meta learning mechanism for AIoT
Fang Dong 0001, Xinghua Ge, Qinya Li, Jinghui Zhang 0001, Dian Shen, Xiao Liu 0004, Gang Li 0009, Fan Wu 0006, Junzhou Luo |
J. Syst. Archit. | 1 |
| 2023 | Noise-aware Local Model Training Mechanism for Federated LearningabstractAs a new paradigm in training intelligent models, federated learning is widely used to train a global model without requiring local data to be uploaded from end devices. However, there are often mislabeled samples (i.e., noisy samples) in the dataset, which will cause the model update to deviate from the correct direction during the training process, thus reducing the convergence accuracy of the global model. Existing works employ noisy label correction techniques to reduce the impact of noisy samples on model updates by correcting labels; however, such methods necessitate the use of prior knowledge and additional communication costs, which cannot be directly applied to federated learning due to data privacy concerns and limited communication resources. Therefore, this paper proposes a noise-aware local model training method that corrects the noisy labels directly at the end device under the constraints of federated learning. By constructing a label correction model, a joint optimization problem is formally defined for optimizing both the label correction model and the client-side local training model (e.g., classification model). As a solution to this optimization problem, we propose a robustness training algorithm using label correction, along with a cross-validation data sampling algorithm that updates both models simultaneously. It is verified through experiments that the mechanism can effectively improve the model convergence accuracy on noisy datasets in federated learning scenarios. Jinghui Zhang 0001, Dingyang Lyu 0001, Qiangsheng Dai, Fa Xin, Fang Dong 0001 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2023 | OL-EUA: Online User Allocation for NOMA-Based Mobile Edge ComputingabstractMobile edge computing (MEC) raises a variety of new challenges for app vendors, including the Edge User Allocation (EUA) problem. EUA aims to allocate as many app users as possible in an MEC system to minimum edge servers in the system. In non-orthogonal multiple access (NOMA)-based MEC system, multiple app users can be allocated to the same subchannel on an edge server through transmit power allocation based on their intra-cell and inter-cell interference. However, allocating excessive app users to the same subchannel may result in severe interference and consequently impact app users’ data rates. In addition, in an MEC system, app users join and depart randomly, and thus need to be allocated in an online manner. Existing EUA approaches suffer from poor performance in dynamic real-world NOMA-based MEC systems because they allocate app users in an offline manner and do not consider the complication caused by NOMA. In this paper, we propose OL-EUA, an OnLine approach for solving dynamic EUA problems in NOMA-based MEC systems. Its performance is theoretically analyzed and experimentally evaluated on a public dataset. Guangming Cui, Qiang He 0001, Xiaoyu Xia 0001, Feifei Chen 0001, Fang Dong 0001, Hai Jin 0001, Yun Yang 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2023 | Multi-Exit DNN Inference Acceleration Based on Multi-Dimensional Optimization for Edge IntelligenceabstractEdge intelligence, as a prospective paradigm for accelerating DNN inference, is mostly implemented by model partitioning which inevitably incurs the large transmission overhead of DNN's intermediate data. A popular solution introduces multi-exit DNNs to reduce latency by enabling early exits. However, existing work ignores the correlation between exit settings and synergistic inference, causing incoordination of device-to-edge. To address this issue, this paper first investigates the bottlenecks of executing multi-exit DNNs in edge computing and builds a novel model for inference acceleration with exit selection, model partition, and resource allocation. To tackle the intractable coupling subproblems, we propose a Multi-exit DNN inference Acceleration framework based on Multi-dimensional Optimization (MAMO). In MAMO, the exit selection subproblem is first extracted from the original problem. Then, bidirectional dynamic programming is employed to determine the optimal exit setting for an arbitrary multi-exit DNN. Finally, based on the optimal exit setting, a DRL-based policy is developed to learn joint decisions of model partition and resource allocation. We deploy MAMO on a real-world testbed and evaluate its performance in various scenarios. Extensive experiments show that it can adapt to heterogeneous tasks and dynamic networks, and accelerate DNN inference by up to 13.7x compared with the state-of-the-art. Fang Dong 0001, Huitian Wang, Dian Shen, Zhaowu Huang, Qiang He 0001, Jinghui Zhang 0001, Liangsheng Wen |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | Energy-Efficient General PoI-Visiting by UAV With a Practical Flight Energy ModelabstractUnmanned aerial vehicles (UAVs) are being widely exploited for various applications,e.g., traversing to collect data from ground sensors, patrolling to monitor key facilities, moving to aid mobile edge computing. We summarize these UAV applications and formulate a problem, namely thegeneral waypoint-based PoI-visiting problem. Since energy is critical due to the limited onboard storage capacity, we aim at minimizing flight energy consumption. In our problem, we pay special attention to the energy consumption for turning and switching operations on flight planning, which are usually ignored in the literature but play an important role in practical UAV flights according to our real-world measurement experiments. We propose specially designed graph parts to model the turning and switching cost and thus transfer the problem into a classic graph problem,i.e., general traveling salesman problem, which can be efficiently solved. Theoretical analysis shows that such problem transformation has the graph redefinition approximation ratio upper bound,$max\lbrace \Theta /\delta ,2\rbrace$, where$\Theta$is related to the designed graph parts and$\delta$is a constant. Finally, we evaluate our proposed algorithm by simulations. The results show that it costs less than 107% of the optimal minimum energy consumption for small scale problems and costs only 50% as much energy as a naive algorithm for large scale problems. Feng Shan, Runqun Xiong, Fang Dong 0001, Junzhou Luo, Suyang Wang |
IEEE Trans. Mob. Comput. | 4 |
| 2023 | Enabling Distributed and Optimal RDMA Resource Sharing in Large-Scale Data Center Networks: Modeling, Analysis, and ImplementationabstractRemote Direct Memory Access (RDMA) suffers from unfairness issues and performance degradation when multiple applications share RDMA network resources. Hence, an efficient resource scheduling mechanism is urged to optimally allocates RDMA resources among applications. However, traditional Network Utility Maximization (NUM) based solutions are inadequate for RDMA due to three challenges: 1) The standard NUM-oriented algorithm cannot deal with coupling variables introduced by multiple dependent RDMA operations; 2) The stringent constraint of RDMA on-board resources complicates the standard NUM by bringing extra optimization dimensions; 3) Naively applying traditional algorithms for NUM suffers from scalability issues in solving a large-scale RDMA resource scheduling problem. In this paper, we present how to optimally share the RDMA resources in large-scale data center networks with a distributed manner. First, we propose Distributed RDMA NUM (DRUM) to model the RDMA resource scheduling problem as a new variation of the NUM problem. Second, we present distributed algorithms to efficiently solve the large-scale, interdependent RDMA resource sharing problem for different RDMA use cases. Through theoretical analysis, the convergence and parallelism of proposed algorithms are guaranteed. Finally, we implement the algorithms as a kernel-level indirection module in the real-world RDMA environment, so as to provide end-to-end resource sharing and performance guarantee. Through extensive evaluations by large-scale simulations and testbed experiments, we show that our method significantly improves applications’ performance under resource contention, achieving$1.7-3.1\times $higher throughput, and in a dynamic context, the largest performance improvement reaches 98.1% and 64.1% in terms of latency and throughput, respectively. Dian Shen, Junzhou Luo, Fang Dong 0001, Xiaolin Guo, Ciyuan Chen, John C. S. Lui |
IEEE/ACM Trans. Netw. | 3 |
| 2023 | Editorial for Resource Management at the Edge for Future web, Mobile, and IoT Applications
Qiang He 0001, Fang Dong 0001, Chenshu Wu, Yun Yang 0001 |
World Wide Web (WWW) | 2 |
| 2022 | Last-mile Matters: Mitigating the Tail Latency of Virtualized Networks with Multipath Data PlaneabstractVirtualized network has become the cornerstone of today's large-scale cloud data centers. In particular, the data plane of virtualized network, consisting of virtual switch, virtual router and other software network functionalities, performs all network packets processing of virtual machines (VMs). However, current virtualized data plane solutions incur drastic performance interference with co-resident VMs, and thus suffer from unpredictable network performance, especially in terms of tail latency. In this work, we show that the performance issue stems from the fact that CPU plays a dual role of both communication and computation in virtualized networks. A number of virtual network components and their complex packets processing create an undue burden on the hosts' CPUs and in turn cause the mutual performance interference among VMs and networks. To address this issue, we present a multipath data plane solution, where the traffic of VMs can be adaptively and seamlessly offloaded to the adjacent hosts. At the core of this design is to optimize the VM traffic allocation among multiple paths. We formulate the VM multipath traffic allocation problem with coupled variables of computing and network resources, which were only considered as mutually independent in prior researches. Then we present a distributed algorithm to efficiently solve the large-scale, interdependent global optimization problem, with convergence and optimality guarantees. Through extensive simulations and real-world testbed experiments, we show that our solution delivers consistent performance improvement (up to$6.7\times$improvement in aggregate throughput and$21.4\times$reduction in tail latency, respectively) in the dynamic cloud system. Dian Shen, Yi Zhai 0004, Fang Dong 0001, Junzhou Luo |
CLUSTER | 3 |
| 2022 | Federated Learning Client Selection Mechanism Under System and Data HeterogeneityabstractFederated learning (FL) has been proposed to train a global model by distributed architecture, while keeping the training data local. Owing to the large scale of clients in FL, all clients to participate in training is not feasible. The heterogeneity of clients, including system and data heterogeneity, also poses huge challenge to the client selection problem. Traditional client selection mechanisms can’t handle these heterogeneities effectively, which lead to poor training efficiency. Hence, this paper comprehensively considers system and data heterogeneity to select clients and dynamically adjust the number of selected clients. For system heterogeneity, we build the latency model to predict the training time for selecting clients with best performance including CPU frequency, the size of dataset and transmission power. Besides, for data heterogeneity, the cluster model is established to cluster clients for alleviating the accuracy jitter owing to the non independent and identically distributed (Non-IID) dataset. We formulate the client selection problem aiming to minimize the overall training time on the premise of accuracy, and design the Federated Client Cluster and latency-Prediction Selection (FCCPS) algorithm to solve this problem. With extensive simulations, we show that the FCCPS algorithm can reduce the training time by up to 21% on Cifar-10 dataset and 13% on FashionMNIST dataset, as compared to FedAvg. Fan Xin, Jinghui Zhang 0001, Junzhou Luo, Fang Dong 0001 |
CSCWD | 4 |
| 2022 | Towards the Full Extensibility of Multipath TCP with eMPTCPabstractMPTCP provides the basic multipath support for network applications to deliver high throughput and robust communication. However, the original MPTCP is designed with limited extensibility. Various research works have tried to extend MPTCP to attain better performance or richer functionalities. These existing approaches either modify the kernel implementation of MPTCP, which involve considerable engineering efforts and may accidentally introduce security issues, or control MPTCP via user-space tools, which suffer from restricted functionality support. To address this issue, we propose eMPTCP, an easy-to-use framework to fully extend MPTCP without security risks. Internally, eMPTCP has a modular and pluggable model which allows operators to specify a comprehensive MPTCP extension as a chain of sub-policies. eMPTCP further enforces the policies through packet header manipulations. To ensure safety, eMPTCP is implemented using eBPF. Despite the stringent constraints of eBPF, we show that it is possible to implement an elaborated framework for a fully extensible MPTCP. Through verifying MPTCP in a number of real-world cases and extensive experiments, we show that eMPTCP is able to support a wide range of MPTCP extensions, while the overhead of eMPTCP operations in the kernel is in the scale of nanosecond, and the extra processing time accounts for only about 0.63% of flows' transmission time. Bin Yang 0027, Dian Shen, Junxue Zhang 0001, Fang Dong 0001, Junzhou Luo, John C. S. Lui |
ICNP | 4 |
| 2022 | Enabling Latency-Sensitive DNN Inference via Joint Optimization of Model Surgery and Resource Allocation in Heterogeneous EdgeabstractNowadays, edge computing is widely adopted to resolve the emerging deep neural networks (DNNs)-driven intelligence scenarios with the requirement of low-latency and high-accuracy, which includes heterogeneous end devices and DNNs. In such scenarios, the influx of data and computation into a shared edge server incurs prohibitive latency. Thus, we exploit the advantage of Multi-exit DNNs (ME-DNNs) that tasks can exit early at appropriate depths to save inference time. However, naively using ME-DNNs in the heterogeneous edge still fails to deliver fast inference due to improper model surgery and resource allocation. Zhaowu Huang, Fang Dong 0001, Dian Shen, Huitian Wang, Xiaolin Guo, Shucun Fu |
ICPP | 2 |
| 2022 | Formulating Interference-aware Data Delivery Strategies in Edge Storage SystemsabstractNetworked edge servers constitute an edge storage system in edge computing (EC). Upon users’ requests, data must be delivered from edge servers in the system or from the cloud to users. Existing studies of edge storage systems have unfortunately neglected the fact that an excessive number of users accessing the same edge server for data may impact users’ data rates seriously due to the wireless interference. Thus, users must first be allocated to edge servers properly for ensuring their data rates. After that, requested data can be delivered to users to minimize their average data delivery latency. In this paper, we formulate this Interference-aware Data Delivery at the network Edge (IDDE) problem, and demonstrate its NP-hardness. To tackle it effectively and efficiently, we propose IDDE-G, a novel approach that first finds a Nash equilibrium as the strategy for allocating users. Then, it finds an approximate strategy for delivering requested data to allocated users. We analyze the performance of IDDE-G theoretically and evaluate its performance experimentally to demonstrate the effectiveness and efficiency of IDDE-G on solving the IDDE problem. Xiaoyu Xia 0001, Feifei Chen 0001, Qiang He 0001, Guangming Cui, John C. Grundy, Mohamed Almorsy, Fang Dong 0001 |
ICPP | 7 |
| 2022 | Exploiting the Computational Path Diversity with In-network Computing for MECabstractWith Computing in the Network technologies, Mobile Edge Computing (MEC) has expanded the resource distribution and tightly integrated computing-network capabilities from the end-devices, through the edge, to the cloud infrastructure, including at points in between. Thus, edge computing is able to deliver a more collaborative processing, better service responding to the increasing application needs in low latency processing. In the presence of integrated computing-network resources and their increased capacity, current proximity-to-data methods in edge computing lead to sub-optimal performance in terms of processing latency. Addressing this issue, this paper presents a Low-latency Adaptive Workload Allocation framework (LAWA) to harness the growing in-network computing resources to deliver low latency processing capabilities for emerging latency-constrained applications. LAWA defines an application by its computational source and destination. Considering the diversity of computing and network resources, we try to find an optimal computational path and its workload allocation. We model the problem as a mixed integer programming problem. To solve this problem, we propose the computational pathfinding and workload allocation algorithms with optimality guarantees. Experimental results show that, comparing with the state-of-the-art methods, our method achieves up to 8.04× speedup, in terms of end-to-end latency. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang, Zhenyang Ni, Yulong Jiang, Daheng Yin |
SECON | 2 |
| 2021 | PipePar: A Pipelined Hybrid Parallel Approach for Accelerating Distributed DNN TrainingabstractLarge scale DNN training tasks are exceedingly compute-intensive and time-consuming, which are usually executed on highly-parallel platforms. Data and model parallelization is a common way to speed up the training progress across devices. However, they tend to achieve sub-optimal performance due to the communication overheads and unbalanced load among servers. Recent emerging pipelining solutions mitigate the above issues, incorporating the advantages of data and model parallelism. In this paper, we make a step further towards optimizing the execution of pipelining. We introduce PipePar, a pipeline-parallel DNN training method that provides optimized execution strategies of layer-stacked DNNs. PipePar considers the entire tensor partition space of pipelining and explores potential hybrid parallel configurations of each stage in the pipeline. Additionally, we notice the network heterogeneity between different GPU servers and it is inevitable to transfer tensors with different bandwidths and latency. So, taking into account both computation and communication capacity of different GPU servers, PipePar is intended to find a elastic load distribution strategy at different levels. We evaluate PipePar with a set of real-world DNNs on 4 GPU servers. Our experimental results show that PipePar is able to find an efficient strategy that are up to 2.16× faster than state-of-the-art hybrid parallelization approaches. Jiange Li, Jinghui Zhang 0001, Jiahui Jin 0001, Fang Dong 0001 |
CSCWD | 5 |
| 2021 | Towards Tunable RDMA Parameter Selection at Runtime for Datacenter ApplicationsabstractBecause of the low-latency and high-throughput benefits of RDMA, an increasing number of collaborative applications in datacenters are re-designed with RDMA to boost the performance. Among various low-level hardware primitives provided by RDMA, exposed as parameters of APIs, the application designers select and hardcode them to exploit all the performance benefits of RDMA. However, with the dynamic nature of datacenter application, the hardcoded and fixed parameter selection fails to take full advantages of RDMA capabilities, which can cause up to 35% throughput performance loss. To address this issue, we present a tunable RDMA parameter selection framework, which allows parameter tuning at runtime, adaptive to the dynamic application and server status. To attain the native RDMA performance, we use a lightweight decision tree to reduce the overhead of RDMA parameter selection. Finally, we implement the tunable RDMA parameter selection framework with native RDMA API to provide a more abstract API. To demonstrate the effectiveness of our method, we implement a key-value service based on the abstract API. Experiment results show that our implementation has only a very small overhead compared with the native RDMA, while the optimized key-value service achieves 112% more throughput than Pilaf and 66% more throughput than FaRM. Fang Dong 0001, Dian Shen, Chengtian Zhang, Jinghui Zhang 0001, Junzhou Luo |
CSCWD | 2 |
| 2021 | A CPU Load-awared Virtual Router Placement Strategy in Cloud NetworkabstractWith the increment of the scale of users and networks, network virtualization technology has been widely used by service providers in cloud networks to address elastic network demands. As a fundamental network virtualization component realizing cross-tenant traffic routing, the placement of virtual routers has become a considerable factor influencing the network performance. And through in-depth experiments, we also found that CPU load augments may well incur a significant degradation in network throughput performance, revealing the problems of the placement of virtual routers in existing cloud network modes: 1) Ignore the bandwidth loss caused by the CPU load variation. 2) Lack of theoretical support for the optimal scheme. Based on these above, we propose a CPU load-aware virtual router placement strategy, which balances the computing load situation of each virtual router, and adopts branch and bound algorithm and convex optimization to achieve the approximate optimal placement within 0.1% error. We have evaluated our strategy in our cloud testbed, and find a 20% improvement in terms of cross-tenant throughput compared with the worst case of existing strategies, Fang Dong 0001, Dian Shen, Yi Zhai 0004, Ciyuan Chen |
CSCWD | 2 |
| 2021 | Enabling Low Latency Edge Intelligence based on Multi-exit DNNs in the WildabstractIn recent years, deep neural networks (DNNs) have witnessed a booming of artificial intelligence Internet of Things applications with stringent demands across high accuracy and low latency. A widely adopted solution is to process such computation-intensive DNNs inference tasks with edge computing. Nevertheless, existing edge-based DNN processing methods still cannot achieve acceptable performance due to the intensive transmission data and unnecessary computation. To address the above limitations, we take the advantage of Multi-exit DNNs (ME-DNNs) that allows the tasks to exit early at different depths of the DNN during inference, based on the input complexity. However, naively deploying ME-DNNs in edge still fails to deliver fast and consistent inference in the wild environment. Specifically, 1) at the model-level, unsuitable exit settings will increase additional computational overhead and will lead to excessive queuing delay; 2) at the computation-level, it is hard to sustain high performance consistently in the dynamic edge computing environment. In this paper, we present a Low Latency Edge Intelligence Scheme based on Multi-Exit DNNs (LEIME) to tackle the aforementioned problem. At the model-level, we propose an exit setting algorithm to automatically build optimal ME-DNNs with lower time complexity; At the computation-level, we present a distributed offloading mechanism to fine-tune the task dispatching at runtime to sustain high performance in the dynamic environment, which has the property of close-to-optimal performance guarantee. Finally, we implement a prototype system and extensively evaluate it through testbed and large-scale simulation experiments. Experimental results demonstrate that LEIME significantly improves applications' performance, achieving 1.1–18.7 × speedup in different situations. Zhaowu Huang, Fang Dong 0001, Dian Shen, Junxue Zhang 0001, Huitian Wang, Guangxing Cai, Qiang He 0001 |
ICDCS | 2 |
| 2021 | Dynamic Path Based DNN Synergistic Inference Acceleration in Edge Computing EnvironmentabstractDeep Neural Networks (DNNs) have achieved excellent performance in intelligent applications. Nevertheless, it is elusive for devices with limited resources to support computationally intensive DNNs, while employing the cloud may lead to prohibitive latency. Better solutions are exploiting edge computing and reducing unnecessary computation. Multi-exit DNN based on the early exit mechanism has an impressive effect in the latter, and in edge computing paradigm, model partition on multi-exit chain DNNs is proved to accelerate inference effectively. However, despite reducing computations to some extent, multiple exits may lead to instability of performance due to variable sample quality, performance inferior to the original model especially in the worst case. Furthermore, nowadays DNNs are universally characterized by a directed acyclic graph (DAG), complicating the partition of multi-exit DNN exceedingly. To solve the issues, in this paper, considering online exit prediction and model execution optimization for multi-exit DNN, we propose a Dynamic Path based DNN Synergistic inference acceleration framework (DPDS), where exit designators are designed to avoid iterative entry for exits; to further promote computational synergy in the edge, the multi-exit DNN is dynamically partitioned according to network environment to achieve fine-grained computing offloading. Experimental results show that DPDS can significantly accelerate DNN inference by 1.87× to 6.78×. Huitian Wang, Fang Dong 0001, Wei Zhao 0023 |
ICPADS | 4 |
| 2020 | Distributed and Optimal RDMA Resource Scheduling in Shared Data Center NetworksabstractRemote Direct Memory Access (RDMA) suffers from unfairness issues and performance degradation when multiple applications share RDMA network resources. Hence, an efficient resource scheduling mechanism is urged to optimally allocates RDMA resources among applications. However, traditional Network Utility Maximization (NUM) based solutions are inadequate for RDMA due to three challenges: 1) The standard NUM-oriented algorithm cannot deal with coupling variables introduced by multiple dependent RDMA operations; 2) The stringent constraint of RDMA on-board resources complicates the standard NUM by bringing extra optimization dimensions; 3) Naively applying traditional algorithms for NUM suffers from scalability and convergence issues in solving a large-scale RDMA resource scheduling problem. Dian Shen, Junzhou Luo, Fang Dong 0001, Xiaolin Guo, John C. S. Lui |
INFOCOM | 3 |
| 2020 | S-MAC: Achieving High Scalability via Adaptive Scheduling in LPWANabstractLow Power Wide Area Networks (LPWAN) are an emerging well-adopted platform to connect the Internet-of-Things. With the growing demands for LPWAN in IoT, the number of supported end-devices cannot meet the IoT deployment requirements. The core problem is the transmission collisions when large-scale end-devices transmit concurrently. The previous research mainly includes transmission scheduling strategies, collision detection and avoidance mechanism. The use of these existing approaches to address the above limitations in LPWAN may introduce excessive communication overhead, end-devices cost, power consumption, or hardware complexity. In this paper, we present S-MAC, an adaptive MAC-layer scheduler for LPWAN. The key innovation of S-MAC is to take advantage of the periodic transmission characteristics of LPWAN applications and also the collision behaviour features of LoRa PHY-layer to enhance the scalability. Technically, S-MAC is capable of adaptively perceiving clock drift of end-devices, adaptively identifying the join and exit of end-devices, and adaptively performing the scheduling strategy dynamically. Meanwhile, it is compatible with native LoRaWAN, and adaptable to existing Class A, B and C devices. Extensive implementations and evaluations on commodity devices show that S-MAC increases the number of connected end-devices by 4.06× and improves network throughput by 4.01× with PRR requirement of > 95%. Zhuqing Xu, Junzhou Luo, Zhimeng Yin 0001, Tian He 0001, Fang Dong 0001 |
INFOCOM | 5 |
| 2020 | Emerging intelligent big data analytics for cloud and edge computingabstractIntelligent big data analytics is an emerging paradigm in the age of big data, analytics, and artificial intelligence, and it exploits how to use artificial intelligence to enhance big data analytics for various applications.1 As cloud computing cannot meet the strict computing time requirement in latency-critical big data analysis applications, edge computing has emerged as a solution to address the drawbacks of cloud-based solutions by moving computation physically closer to the network edge where data are generated. However, edge computing does not have sufficient resources for complex intelligent big data analytics tasks. Consequently, this special issue is focused on exploiting key techniques of intelligent big data analytics by involving cloud and edge computing. Presented with an avalanche of biological interactions data, computational biology is now facing greater challenges on big data analysis and requires more studies to mine and integrate cloud-based multiomics data, especially when the data are related to infectious diseases. Meanwhile, machine learning techniques have recently succeeded in different computational biology tasks. For this reason, Chen et al2 proposed APEX2S, a novel two-layer machine learning model, for discovery of the protein-protein interactions data. APEX2S calibrated the focus for host-pathogen protein-protein interactions study, aiming to apply machine learning techniques for learning the interactions data and making predictions. To date, there are a wide variety of applications of human action recognition, such as surveillance, robotics, health care, video searching, and human-computer interaction. However, there are many challenges involved in human action recognition in videos, such as cluttered backgrounds, occlusions, viewpoint variation, execution rate, and camera motion. To solve this, Zhao et al3 proposed a novel action recognition method to improve the recognition accuracy by adopting the key frame extraction and multi-feature fusion techniques. A key frame extraction method based on node contribution weighting is proposed to extract video key frames, and different convolutional neural networks are used to obtain corresponding classification results and merge, so as to better complement the information in different flows. Many applications are now deployed on Virtual Machines (VMs) or even Spot VMs elastically rented from public Clouds. To save costs, interval-priced VMs are not released until the ends of rented intervals. Such delays of control effects make existing methods rent or release excess VMs leading to over controls. Fluctuating prices make Spot VMs unreliable due to unexpected termination which makes fault-tolerant strategies crucial. In order to decrease the VM rental cost while guaranteeing the SLA and robustness, Cai et al4 proposed a hybrid control method UCM which takes advantage of queuing-model-based loosely coupled controllers, unequal-interval-based collaborating method, and an existing group-based fault tolerant strategy. Lidar-based city objects detection is an interesting topic along with the development of Laser scan equipment which has been widely applied in various applications such as 3D building reconstruction, navigation, and so on. Superpixel segmentations are widely applied to image processing or computer vision tasks. Many experiments have proven that superpixels generated from atomic meaningful pixel regions, can improve the processing efficiency while losing little information of the original image. Therefore, Mao et al5 describes a city object detection algorithm for airborne Lidar images using superpixel segmentation and DenseNet classification. A three-block DenseNet is applied to classify the superpixels into four main types of city objects (Building, road, field, and railway). In addition, a graph based neighborhood adjustment algorithm is designed to further improve the classification results. Virtual network embedding (VNE) aims to solve how to efficiently allocate physical resources to a virtual network. However, this issue has been proved to be an NP-hard problem. To address the challenge, Wang et al6 formalize the problem as a mixed integer programming problem and propose a novel VNE method based on reinforcement learning. Then to solve this problem, Wang et al6 introduce a pointer network to generate virtual node mapping strategies through an attention mechanism, and design a reward function related to link resource consumption to build the connection between node mapping and link mapping stages of VNE. Dual-hop 60 GHz wireless networks which support relay-assisted dual-hop transmission have been widely adopted in recent years, aiming to prolong communication distance and bypass obstacles in 60 GHz band. However, it is very challenging to perform link scheduling in such dual-hop architecture while considering several factors, that is, reducing network power consumption, avoiding overloaded APs/relays and adapting to network dynamics. To this end, Wu et al7 investigate the problem of energy efficient link scheduling with load constraints (ELL), and propose solutions to deal with network dynamics by presenting a fine-grained energy model for dual-hop 60 GHz networks and proposing a polynomial-time global scheduling algorithm. Job-pool based workload estimation has attract a lot of attention recently, which analyzes the characteristics of existing tasks' workloads to estimate the currently running tasks' workload. However, the workload patterns of some tasks do have seasonality and trend, and conventional per-job based regression methods may yield better workload prediction results. Also, in some cases, some new tasks may not follow the workload patterns of existing tasks in the pool. Thus, Yu et al8 develop an integrated scheme which combines clustering and regression for workload prediction. Exorbitant resources are required to train a deep neural network (DNN). Often researchers deploy an approach that uses distributed parallel training to acquire larger models faster on GPUs. This approach has its detriments, though; on one hand, a GPU's expanded capacity to compute also produces bigger bottlenecks in inter-GPU's communications during model training, and multi-GPU systems lead to complex connectivity. Workload schedulers then end up having to consider hardware topology and requirements for workload communication, in hopes of allocating GPU resources to optimize execution time and improve usage in a heterogeneous environment. On the other hand, the high memory requirements to train a DNN model make running the training processes on GPUs onerous. To contend with this, Zhang et al9 introduce two execution optimization methods based on pipeline-hybrid parallelism in a GPU cluster with heterogeneous networking. Sketch is a compact data structure used to summarize data streams. It is widely used in the measurement of network traffic, and its accuracy is higher than traditional methods. Currently, there are some typical sketches: Count-Min Sketch, CU Sketch, and Count Sketch. According to the characteristics of network traffic, Zhu et al10 propose a new sketch framework called Self-Adaption Sketch, which combines Sketch and Bloom Filter. In the framework, the sketch is created dynamically and the memory space is adjusted timely according to the network traffic by using the concept carrying. We thank the authors for their contributions, including those whose papers are not included in this special issue. We also would like to acknowledge thoughtful work from many reviewers who provided valuable evaluations and recommendations. Fang Dong 0001, Jianming Yong |
Concurr. Comput. Pract. Exp. | 1 |
| 2020 | Towards tenant demand-aware bandwidth allocation strategy in cloud datacenter
Jiuxin Cao, Zhuo Ma 0002, Jue Xie, Xiangying Zhu, Fang Dong 0001, Bo Liu 0004 |
Future Gener. Comput. Syst. | 5 |
| 2020 | Accelerating Skycube Computation with Partial and Parallel Processing for Service SelectionabstractRecently researchers use skyline techniques to optimize service selection procedure, where they can filter those low-quality web services from the large amount of candidates and return a much smaller high-quality service set. The skycube concept is adopted for quickly responding to the skyline queries with different combinations of Quality of Web Service (QoWS) parameters. As the skycube computation is quite time-consuming, it is a compelling challenge to accelerate this procedure. However, the current solutions usually have a number of redundant computations which will significantly affect the efficiency. To address such drawbacks, after an in-depth analysis of skycube computation procedure, we introduce a partial skycube, which only consists of the skylines with frequently used combinations of QoWS. Then the computational relationships between the skyline on one subspace and its parent-space are studied. Based on the relationships, we develop ParCube algorithm to speedup partial skycube computation by reusing the intermediate comparison results. Meanwhile, at the execution phase, ParCube can be further optimized with parallel execution mode and optimized scheduling strategy. Finally, we evaluate the efficiency and scalability of ParCube on both single machine and cluster environment. The results show that ParCube can efficiently compute partial skycube and scale well in cluster environment. Fang Dong 0001, Junzhou Luo, Jiahui Jin 0001, Jiyuan Shi, Jun Shen 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2020 | Facilitating Application-Aware Bandwidth Allocation in the Cloud with One-Step-Ahead Traffic InformationabstractBandwidth allocation to virtual machines (VMs) has a significant impact on the performance of communication-intensive big data applications hosted in VMs. It is crucial to accurately determine how much bandwidth to be reserved for VMs and when to adjust it. Past approaches typically resort to predicting the long-term network demands of applications for bandwidth allocation. However, lacking of prediction accuracy, these methods lead to the unpredictable application performance. Recently, it is conceded that the network demands of applications can only be accurately derived right before each of their execution phases. Hence, it is challenging to timely allocate the bandwidth to VMs with limited information. In this paper, we design and implement AppBag, an Application-aware Bandwidth guarantee framework, which allocates the accurate bandwidth to VMs with one-step-ahead traffic information. We propose an algorithm to allocate the bandwidth to VMs and map them onto feasible hosts. To reduce the overhead when adjusting the allocation, an efficient Lazy Migration (LM) algorithm is proposed with bounded performance. We conduct extensive evaluations using real-world applications, showing that AppBag can handle the bandwidth requests at run-time, while reducing the execution time of applications by 47.3 percent and the global traffic by 36.7 percent, compared to the state-of-the-art methods. Dian Shen, Junzhou Luo, Fang Dong 0001, Jiahui Jin 0001, Junxue Zhang 0001, Jun Shen 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2019 | ParaNF: Enabling Delay-Balanced Network Function Parallelism in NFVabstractIn Network Function Virtualization (NFV), multiple network functions cooperate to provide various network services. To reduce the end-to-end latency through a chain of network functions, research hotspots have turned to complete NF parallelism frameworks. However, several issues remain in them such as the manual dependency analysis on NFs and the excessive parallelism for NFs. Therefore, in this paper, we present ParaNF, an effective delay-balanced NF parallelism framework. ParaNF mainly consists of two logical components. First, the ParaNF orchestrator conducts a dynamic dependency analysis to find out which NFs can be parallelized and then conducts a delay-balanced NF parallelism optimization strategy. Second, the ParaNF infrastructure performs light-weight, dynamic packet copying and merging guided by an efficient label mechanism to support high-performance NF parallelism. We implement a ParaNF prototype with DPDK. Our evaluations show that ParaNF not only realizes the line-speed packet processing, but also achieves significant reduction in latency by up to 47% than the traditional SFC and 35% than OpenBox. Junzhou Luo, Fang Dong 0001, Dian Shen |
CSCWD | 3 |
| 2019 | ADDA: Adaptive Distributed DNN Inference Acceleration in Edge Computing EnvironmentabstractImplementing intelligent mobile applications on IoT devices with DNN technology has become an inevitable trend. Due to the limitations of the size of DNN model deployed onto end devices and the instability of wide-area network transmission, either End-only mode or Cloud-only mode cannot guarantee the reasonable latency and recognition accuracy simultaneously. A better solution is to exploit the edge computing, where the existing edge computing execution framework and offloading mechanism for DNN inference suffer unnecessary computational overheads and underutilized computing capacity of end and edge. To address these shortcomings, an adaptive distributed DNN inference acceleration framework for edge computing environment is proposed in this paper, where DNN computation path optimization and DNN computation partition optimization are taken into consideration. The evaluations demonstrate that our method can effectively accelerate the DNN inference compared to the state-of-the-art methods. Huitian Wang, Guangxing Cai, Zhaowu Huang, Fang Dong 0001 |
ICPADS | 4 |
| 2019 | Cloud computing-based big data processing and intelligent analyticsabstractCloud computing-based big data processing and intelligent analyticsCloud and big data have become the big things today in many systems especially regarding of information processing and intelligent analytics.Big data analytics is the use of advanced analytic techniques against very large, diverse data sets, and it allows analysts, researchers to make better decisions using data that was previously inaccessible or unusable.Due to the urgent demand on high capacity of computation and storage resources, cloud computing has been acknowledged as the primary computing paradigm for massive data storage, processing under various circumstances and different requirement.1 Moreover, edge computing pushes the cloud frontier to the edge of the network and extends cloud computing to be able to address more application scenarios.This special track plans to solicit novel and original manuscripts in the above topics with an emphasize on ''Cloud Computing-based Big Data Processing and Intelligent Analytics.''From those submitted papers for the 6th International Conference on Advanced Cloud and Big Data (CBD 2018) held in Lanzhou, China on August 12 to August 14, 2018, nine papers are selected that target the following research issues in cloud computing and big data:• Service deployment and task scheduling in cloud computing and edge computing.• Cloud storage system design and optimization.• Case studies of big data in cloud-based system.• Network performance optimization in data centers.• Approximate big data analysis and processing.• Security threats and solutions in cloud computing and big data processing.Recently, mobile edge computing (MEC) has become a fascinating technology trend for future computing paradigm, while at the same time, it is also facing a great challenge that is how to make full use of edge resources to provide a seamless support for compute-intensive latency-sensitive applications.Most of the existing works assume that tasks can be executed upon every edge server, but the assumption does not hold in practical scenarios because a specific application task often corresponds to a certain service that provides the corresponding running environment.How to decide service deployment of so many types of services among multiple edge servers is also a big challenge.To address the challenge, Zhou et al 2 study dynamic service deployment for latency-sensitive applications and first model the long-term budget-constrained latency minimization problem as a multi-slot latency minimization problem based on the Lyapunov framework.Furthermore, the task scheduling optimization is also considered here, which makes every edge server be fully utilized in an even more efficient collaborative manner.How to repair data blocks in the erasure coding storage system is of great challenge as considering the incast problem introduced at the new node.Current solutions mainly rely on path planning and resource allocation and waste a large amount of storage and bandwidth resources unavoidably.Xia et al 3 propose the incast problem to be resolved economically via the in-network aggregation, where a set of in-network methods to repair a failed data block in the erasure coding storage systems.Compared with the existing methods, the in-network methods are capable of reducing the bandwidth consumption as well as achieving higher repair speed.With the continuous development of intelligent transportation systems (ITSs), various sensor data can be used to detect traffic information, but there is a lack of practical value to guide public traveling.Xu et al 4 propose a real-time traffic index model of expressways by using a traffic index to evaluate the actual conditions of expressways.The model considers the actual situation of floating and non-floating vehicles on expressways.Included is the realization of the complete calculation model of real-time traffic index estimation, including highway section division, spatial topology map matching, driving route calculation, and road congestion status judgment.For roads without floating car coverage, the weighted-moving-average time-series prediction method is used to predict the traffic index, so that the running condition of all roads in the network can be analyzed completely.A spotlight has shined on scientific workflows in recent years, as a result of their enormous impact on big data-related scientific areas.Large-scale scientific workflow scheduling across global data centers requires the scheduling framework to optimize data movement cost by leveraging these distributed data centers.However, challenges regarding of data-intensive workflow execution in multiple geo-distributed data centers still exist, such as data dependency, intermediate data placement, etc. Scientific workflow's data and task co-scheduling aim to solve the above problem, which is known to be NP-hard.Zhang et al 5 propose a novel approach based on the multilevel graph coarsening and un-coarsening framework, together with a specialized hybrid genetic algorithm having distinctive graph partition driven features of repair and local improvement, for scheduling data-intensive scientific workflows in geo-distributed data centers and optimizing the cross-data center data transfer volume. Fang Dong 0001, Chenshu Wu, Shangce Gao |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | GStar: an efficient framework for answering top-k star queries on billion-node knowledge graphs
Jiahui Jin 0001, Junzhou Luo, Samamon Khemmarat, Fang Dong 0001, Lixin Gao 0001 |
World Wide Web | 4 |
| 2018 | (WIP) Evaluation of a Cloud-Based System for Delivering Adaptive Micro Open Education Resource to Fresh LearnersabstractIn this paper, we present an online computation approach implemented in a cloud-based system to assist open education resource (OER) providers and instructors dealing with the sparsity of data in micro OER recommendation. An algorithmic framework is provided to realize the novel micro OER recommendation system based on heuristic rules. These rules can also optimize the approaches to blending new-coming micro OERs into established learning paths. Comparing with different widely used recommender systems, our evaluation shows the proposed heuristic algorithms for online computation performs satisfactorily in terms of precision and recall values. Geng Sun 0002, Tingru Cui, Fang Dong 0001, Jun Shen 0001, Shiping Chen 0001, Jiayin Lin |
IEEE CLOUD | 3 |
| 2018 | Ensemble Machine Learning Systems for the Estimation of Steel Quality ControlabstractRecent advances in the steel industry have encountered challenges in soliciting decision making solutions for quality control of products based on data mining techniques. In this paper, we present a steel quality control prediction system encompassing with real-world data as well as comprehensive data analysis results. The core process is cautiously designed as a regression problem, which is then best handled by grouping various learning algorithms with their massive resource of historical production datasets. The characteristics of the currently most popular learning models used in regression problem analysis are as well investigated and compared. The performance indicates our steel quality control prediction system based on ensemble machine learning model can offer promising result whilst delivering high usability for local manufacturers to address the production problem by aid of development of machine learning techniques. Furthermore, real-world deployment of this system is demonstrated and discussed. Finally, future directions and the performance expectation are pointed out. Fucun Li, Jianqing Wu 0002, Fang Dong 0001, Jiayin Lin, Geng Sun 0002, Huaming Chen, Jun Shen 0001 |
IEEE BigData | 3 |
| 2018 | An Effective Model for Edge-Side Collaborative Storage in Data-Intensive Edge ComputingabstractEdge Computing is a new computing paradigm that performs data processing at the edge of the network (i.e., edge servers) to lower data processing latency. Existing research works have paid lots of attention to how to offload computation tasks from terminals to edge servers, but most of them ignored how to store tasks' necessary data like pretrained models or databases on edge servers. Recently, the data-intensive tasks like deep learning and augmented reality are becoming common, which need large data storages and powerful computation resources. This leads to a cumbersome challenge, since many lightweight edge servers have limited resources. If an edge server does not have a task's necessary data, it needs to offload the task to cloud data centers or download the necessary data from the cloud. Both cases could increase the data processing latency. To address this problem, this paper proposes an edge-side collaborative storage framework (ECS). In ECS, the edge servers collaboratively store and process data-intensive tasks' necessary data. Particularly, if an edge server does not have the necessary data, it will forward the task to the nearest servers that contain the data. An effective iterative data placement algorithm is also proposed to improve ECS's performance. The experimental results show that ECS is 2× better than the traditional non-shared storage framework in terms of the cache hit rate. Junzhou Luo, Jiahui Jin 0001, Runqun Xiong, Fang Dong 0001 |
CSCWD | 5 |
| 2018 | Advances in cloud computing and big data analyticsabstractAdvances in cloud computing and Fang Dong 0001, Jun Shen 0001, Qiang He 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2018 | Throughput Maximization for the Wireless Powered Communication in Green CitiesabstractWireless power transfer (WPT) is a recently developed technique to perfectly address the energy problem for smart city sensors that do not have a readily wired power supply. In a radio-frequency-powered wireless communication system, sensors first harvest energy via the radio-frequency WPT, and then, transmit sensed data to the receiver. The “harvest-then-transmit” protocol is used to coordinate the two operations, in which we wish to optimally decide when to harvest energy, when to transmit data, and what transmission rate should be used such that the data are maximally transmitted. Unlike existing works, we assumed that the wireless transferred power is dynamically changing, instead of being constant, which is more realistic in green smart city applications, e.g., Industry 4.0 workshop, smart transportation, and smart buildings, where environments are continuously changing, so the wireless transferred power is affected dynamically. In this paper, we present an optimal scheduling algorithm for the offline case where the varying WPT is known in advance. Based on the optimal principles learned from the offline case, we have designed an efficient online algorithm. Finally, we report our simulation results that demonstrate that our online scheduling algorithm can adaptively and efficiently achieve high data throughput. Feng Shan, Junzhou Luo, Weiwei Wu 0001, Fang Dong 0001, Xiaojun Shen 0002 |
IEEE Trans. Ind. Informatics | 4 |
| 2017 | Virtual network fault diagnosis mechanism based on fault injectionabstractDiagnosing faults in virtual networks is always a popular research area. Existing researches primarily focus on diagnosing faults in physical networks, while they could not identify the faults introduced by virtual networks. Besides, the high complexity of algorithms and the requirement for modifying hardware may limit their scope of use. To address these drawbacks, in this paper, we propose a novel approach to diagnose faults in virtual networks. The rational of our approach is that the faults can be identified when located in the packet traces, with the knowledge that the possible known faults that can happen in that location. To achieve this goal, we apply packet marking, fault injection and machine learning techniques to provide precise fault diagnosis. Experimental results show that our approach can efficiently identify 73% of the faults while for virtual network-specific faults, our approach can diagnose 86% of them. Our system can also support real-time or near real-time fault analysis. Fang Dong 0001, Dian Shen, Runqun Xiong, Jiahui Jin 0001 |
CSCWD | 2 |
| 2017 | GScheduler: Optimizing resource provision by using GPU usage pattern extraction in cloud environmentsabstractGPU-based clusters are widely chosen for accelerating a variety of scientific applications in high-end cloud environments. With their growing popularity, there is a necessity for improving the system throughput and decreasing the turnaround time for co-executing applications on the same GPU device. However, resource contention among multiple applications on a multi-tasked GPU leads to the performance degradation of applications. Previous works are not accurate enough to learn the characteristics of GPU application before execution, or cannot get such information timely, which may lead to misleading scheduling decisions. In this paper, we present GScheduler, a framework to detect and reduce interference for co-executing applications on the GPU-based cloud. The most important feature of GScheduler is to utilize GPU usage pattern extractor for detecting interference between applications. It is composed of key function-call graph extractor and key GPU resource usage vector extractor, the former is used to detect the similarity of GPU usage mode between applications, while the latter is used to calculate the similarity of GPU resource requirements in-between. In addition, an interference aware scheduler is proposed to minimize the interference. We evaluated our framework with 26 diverse, real-world CUDA applications. When compared with state-of the-art interference-oblivious schedulers, our framework improves system throughput by 36% on average, and achieves a 30.5% reduction of turnaround time on average. Zhuqing Xu, Fang Dong 0001, Jiahui Jin 0001, Junzhou Luo, Jun Shen 0001 |
SMC | 2 |
| 2017 | Recent advances in big data analysis and applicationabstractRecent advances in big data analysis and applicationWith the rapid development of information and network technology in recent years, massive data from many different kinds of applications, such as social network, e-business, intelligent transportation and medical diagnosis [1], can be generated, collected, and aggregated under various circumstances and scenarios like WAN, LAN, mobile Internet, Internet of things, and so on [2].The potential value of big data can only be exploited and unleashed by means of efficient big data analysis and application.Meanwhile, big data-related security has also drawn much attention from both the academy and industry with its increasing importance.For these reasons, new methodologies and technologies need to be proposed and developed for advancing both big data analytics and applications.This special issue focuses on a new strategic research area that addresses "Recent Advances in Big Data Analysis and Application."From those submitted papers for the 3rd International Conference on Advanced Cloud and Big Data (CBD 2015) held in Yangzhou, Jiangsu, China, on October 30 to November 1, 2015, 9 papers are selected that target the following research issues in big data:• Big data processing and optimization mechanism, • big data security and privacy protection, and • big data analysis application. Fang Dong 0001, Junzhou Luo |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | New trends and innovative methods in cloud computing and big dataabstractWith the rapid development of network technology in recent years, more and more innovative computing models and related applications appear, such as mobile device based cloud computing, social network, intelligent manufacturing and e-business, et al. Under these situations, massive data will be generated, stored and processed under various circumstances and different requirement. order to effectively store and analyze the big data based on cloud computing architecture to realize intelligent network-based applications, the new methodologies and technologies for both cloud computing and big data need to be proposed and developed. This special track focuses on a new strategic research area that addresses "New Trends and Innovative Methods in Cloud Computing and Big Data." From those submitted papers for the 4th International Conference on Advanced Cloud and Big Data (CBD 2016) held in Chengdu, Sichuan, China on August 13th -August 16th, 2016, five papers are selected that target the following research issues in big data: Fang Dong 0001, Jun Shen 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | Enabling application-aware flexible graph partition mechanism for parallel graph processing systemsabstractSummary With the emerging of the large‐scale graph data,Pregel‐like graph parallel processing systems have been an essential tool to efficiently process the graph data. The first step to use thePregel‐like systems is to partition the graph into multiple blocks and distribute them on multiple machines. The partition strategy plays a significant role in determining the performance because a good partition could both ensure load balance and optimize network communication overhead, and vice versa. However, existing partition strategies fail to meet the requirements because they suffer from the following drawbacks: (1) they ignore the application features and (2) they ignore the multi‐application feature in productive environment. To overcome those drawbacks, we proposed thesuperblockpartition strategy, which utilizes theatomic blocksgenerated by pre‐processing of the original graph and could be constructed and re‐constructed dynamically according to the submitted applications in real time. The hash‐based and clustering‐based pre‐partition methods are covered in details. The application feature extraction method and heuristicsuperblockpartition algorithm are proposed to construct the superblocks. Experimental results show that thesuperblockpartition strategy could boost the graph processing performance and its partition efficiency also outperforms the hash‐based and topology optimal partition strategy. Copyright © 2016 John Wiley & Sons, Ltd. Fang Dong 0001, Junxue Zhang 0001, Junzhou Luo, Dian Shen, Jiahui Jin 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | Towards a fast and secure design for enterprise-oriented cloud storage systemsabstractSummary With the rapid development of information technology, enormous volumes of data are being generated by enterprises at all times. The management and storage of these large‐scale data have always been challenging enterprises. As these data are usually shared among users in a collaborative manner, secure data access and access performance are 2 key concerns for data storage of enterprises. However, current solutions fail to meet the requirements of enterprises since they suffer from the following drawbacks: (1) they do not support fine‐grained access control and cannot meet the strict secure data access requirements of enterprises, and (2) they suffer from the unpredictable access latency. Thus in this paper, we propose Frostor, an enterprise‐oriented cloud storage system, which addresses the secure data access issue through a user account and IP‐based fine‐grained access control mechanism, and guarantees the access performance via a two‐level performance optimization mechanism. We further implement Frostor and deploy it on the testbed environment in a real data center. Extensive evaluations have shown that Frostor implements fine‐grained access control, while achieving a significant reduction (≥60%) on access latency. Fang Dong 0001, Dian Shen, Zhuqing Xu, Junzhou Luo |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | Resource provisioning optimization for service hosting on cloud platformabstractWith the popularity of cloud computing technology, service hosting is used as a typical model to deploy different kinds of services on cloud platform. In recent years, how to effectively provide resources for service hosting has attracted more and more attention. However, most of the existing works only focused on how to effectively provide virtual machines for service hosting. They ignored how to efficiently place these virtual machines into physical servers, when considering multidimensional resource requirements. This may result in unreasonable virtual machine placement in servers, thereby causing the underutilization of resource. To address this problem, we propose a novel resource provisioning method including virtual machine provisioning for hosting service and virtual machine placement in servers. The proposed method decides how many virtual machines should be provided for each service by utilizing queuing theory. Then based on the virtual machines to be provided, the proposed method models the virtual machine placement problem as a variant of cutting stock problem, and decides how many servers should be provided by solving this problem. The proposed method is evaluated by simulations. Experimental results show the proposed method achieves a better performance than these baseline methods. Jiyuan Shi, Fang Dong 0001, Jinghui Zhang 0001, Jiahui Jin 0001, Junzhou Luo |
CSCWD | 2 |
| 2016 | AppBag: Application-Aware Bandwidth Allocation for Virtual Machines in Cloud EnvironmentabstractIt is challenging to allocate the network bandwidth to virtual machines(VMs) hosting communication-intensive applications. Due to the temporal and spatial variability of the hosted applications, it is crucial how much bandwidth to be reserved for each VM and when to adjust it. Prior approaches typically resort to predicting the applications' network demands, according to which the VMs are placed once for all or periodically migrated. However, recent works conceded that the network demands of applications can only be accurately derived right before each execution phase. In this paper, we propose AppBag, an Application-aware Bandwidth guarantee framework which allocates the bandwidth to VMs using only one-stepahead information. An efficient VM migration algorithm is then proposed to adjust the bandwidth allocation and corresponding VM placement, subjected to the network demands variation in future execution phases. We further implement AppBag with OpenStack and deploy it on the testbed environment in our data center. Extensive evaluations using popular applications show that AppBag can handle the bandwidth requests at run-time while improving applications' performance and reducing the global traffic in the data center fabric. Dian Shen, Junzhou Luo, Fang Dong 0001, Junxue Zhang 0001 |
ICPP | 3 |
| 2016 | A client-side directory prefetching mechanism for GlusterFSabstractDistributed file system has the characteristics of large capacity, good scalability and high reliability, which make it widely used in many areas involving large-scale data storage. It offers simplified, highly-available services for users to access data. However, due to the non-metadata design, the performance of traversal operation on large directories in those non-metadata distributed file systems is poor. With the increasing amount of files, it severely affects the user experience. In this paper, we present a directory prefetching mechanism on the client side to reduce directory traversal operation latency in non-metadata distributed file system. The mechanism, combined with the client's cache, adopts the directory access history to predict future access pattern and fetches the content of the directory without user intervention. Our goal is to reduce the overall access latency in the non-metadata distributed file system in order to better satisfy the user experience. Fang Dong 0001, Junxue Zhang 0001, Zhuqing Xu, Junzhou Luo |
SMC | 2 |
| 2015 | Computing service Skyeube for web service selectionabstractResearchers in service computing area introduce the service Skyline to optimize web service selection. It can eliminate those low-quality web services for large amounts of candidates and return a much smaller and high-quality set to the user. But there is one obvious limitation for these work that they can only compute Skyline on one combination of QoWS parameters. However, in practical, different users may be interested in different combinations of QoWS parameters, and existing work cannot afford such requirement for different QoWS preference. In this paper, we introduce the service Skyeube which consists of Skyline on all possible combinations of QoWS parameters. As it is computed previously in off-line manner, using Skyeube can speed up the response time in real-time web service selection. Unfortunately, the current Skycbue computation solutions suffer from the issue of dimension scalability. To overcome this problem, in this paper, the computational relationship between Skyline computation on one subspace and its super-space are studied. Then a novel computational model, which can compute Skyline on related subspaces by reusing the duplicate comparison results, is developed. Based on this model, a Column-sorting based Skyeube computation algorithm, called CSBSC, is proposed to compute Skyeube much more efficiently. The simulations demonstrate the efficiency and scalability of our CSBSC. Fang Dong 0001, Junzhou Luo |
CSCWD | 2 |
| 2015 | Two-Phase Online Virtual Machine Placement in Heterogeneous Cloud Data CenterabstractWith the rapid development and popularity of cloud computing technology, more and more Collaborative Virtual Environment (CVE) systems are migrated to cloud computing environment to improve the effectiveness of resource usage. Virtual Machine (VM) placement in cloud data center is a key issue of providing high-efficient cloud platform for CVE system. However, most existing VM placement algorithms ignore the following characteristics of actual cloud environment: (1) VMs deployment requests arrive and leave dynamically, (2) Cloud data center usually consists of many heterogeneous Physical Machines (PMs). Ignoring these two characteristics result in an inefficient and unbalanced use of multiple resources of PMs. Thus using these algorithms directly will lead to a poor resource utilization. In this article, we propose a two-phase online VM placement algorithm, which helps the cloud data center to minimize different resource usages and aims at a more efficient use of multiple resources. Our algorithm selects the most suitable PM type for VM based on Cosine Similarity, and adaptively maps VMs to PMs by using an approximation algorithm. The proposed algorithm is evaluated by simulations. Experimental results show our proposed algorithm ensures a more efficient use of multiple resources over the existing approaches. Jiyuan Shi, Fang Dong 0001, Jinghui Zhang 0001, Junzhou Luo, Ding Ding 0002 |
SMC | 2 |
| 2015 | Entropy-based denial-of-service attack detection in cloud data centerabstractSummary Cloud data centers today usually lack network resource isolation. Meanwhile, it is easy to deploy and terminate large number of malicious virtual machines in a few seconds, while the administrator is probably difficult to identify these malicious virtual machines immediately. These features open doors for attackers to launch denial‐of‐service (DoS) attacks that target at degrading the quality of cloud service. This paper studies an attack scenario that malicious tenants use cloud resources to launch DoS attack targeting at data center subnets. Unlike traditional data flow‐based detections, which heavily depend on the pattern of data flows, we propose an approach that takes advantage of virtual machine status including CPU usage and network usage to identify the attack. We notice that malicious virtual machines exhibit similar status patterns when attack is launched. Based on this observation, information entropy is applied in monitoring the status of virtual machines to identify the attack behaviors. We conduct our experiments in the campus‐wide data center, and the results show our detection system can promptly and accurately response to DoS attacks. Copyright © 2015 John Wiley & Sons, Ltd. Jiuxin Cao, Fang Dong 0001, Xiangying Zhu |
Concurr. Comput. Pract. Exp. | 3 |
| 2015 | Towards optimized scheduling for data-intensive scientific workflow in multiple datacenter environmentabstractSummary In the big data era, scientific workflow exhibits the characteristics of data intensity and becomes increasingly popular in scientific domains. Efficient scheduling of data‐intensive scientific workflow in a multiple datacenter (DC) environment has been a long‐standing challenge. Most of previous work on data‐intensive scientific workflow scheduling primarily focused on the optimization of reducing the volumes of data transfer between workflow tasks. In this paper, novel scheduling strategies for the execution of data‐intensive scientific workflow in multi‐DC environment are proposed aiming at the optimization of the overall data transfer time. A novel DC selection approach is proposed to minimize the number of DCs having enough storage capacity for the execution of scientific workflow as well as optimized inter‐DC network bandwidth for efficient data transfer between workflow tasks. A k‐means clustering‐based data placement strategy is adopted to intelligently place the initial data of scientific workflow thereby reducing the volume of initial data transfer between different DCs. A multilevel task replication scheduling strategy is invented to reduce the volumes of intermediate data transfer between DCs during the runtime of the scientific workflow. Simulations spanning a broad range of scientific workflow and multi‐DC settings are performed in order to verify the proposed approaches. The numerical results show that our combined scheduling strategy significantly reduces the overall data transfer time and data transfer volume when scientific workflow is scheduled in multi‐DC environment. Copyright © 2015 John Wiley & Sons, Ltd. Jinghui Zhang 0001, Junzhou Luo, Fang Dong 0001, Junxue Zhang 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2015 | Stochastic modeling of dynamic right-sizing for energy-efficiency in cloud data centers
Dian Shen, Junzhou Luo, Fang Dong 0001, Wei Wang 0089, Guoqing Jin, Weidong Li 0001 |
Future Gener. Comput. Syst. | 3 |
| 2014 | A budget and deadline aware scientific workflow resource provisioning and scheduling mechanism for cloudabstractCurrently in large-scale scientific experiments, scientists often submit scientific workflow jobs at different time. From the view of system, the entire workload is a stream of jobs submitted at an unpredictable time and different job has different priority and deadline. Moreover the cost of performing these jobs cannot exceed a certain budget constraint. Therefore how to perform scientific workflow applications efficiently in cloud has become the urgent problem. However most of existing work didn't consider unpredictable submission time of jobs, as well as budget and deadline constrains. In this paper, we design an elastic resource provisioning and task scheduling mechanism to perform scientific workflows in cloud. Our goal is to complete as many high-priority workflows as possible under budget and deadline constrains. This mechanism consists of three phases: workflow preprocessing, elastic resource provisioning and task scheduling. We perform evaluation with real AMS experiment scientific computing data under different budget constraints. We also consider inaccurate task execution time, VM provisioning delays and task failures in evaluation. The results show that our mechanism achieves a better performance than these reference mechanisms. In addition, the inaccurate task execution time, VM provisioning delays, and task failures do not bring significant impact to mechanism's performance. Jiyuan Shi, Junzhou Luo, Fang Dong 0001, Jinghui Zhang 0001 |
CSCWD | 3 |
| 2014 | A Sampling-Based Hybrid Approximate Query Processing System in the CloudabstractSampling-based approximate query processing method provides the way, in which the users can save their time and resources for 'Big Data' analytical applications, if the estimated results can satisfy the accuracy expectation earlier before a long wait for the final accurate results. Online aggregation (OLA) is such an attractive technology to respond aggregation queries by calculating approximate results with the confidence interval getting tighter over time. It has been built into the MapReuduce-based cloud system for big data analytics, which allows users to monitor the query progress and save money by killing the computation earlier once sufficient accuracy has been obtained. Unfortunately, there exists a major obstacle that is the estimation failure of OLA affects the OLA performance, which is resulted from the biased sample set that violates the unbiased assumption of OLA sampling. To handle this problem, we first propose a hybrid approximate query processing model to improve the overall OLA performance, where a dynamic scheme switching mechanism is deliberately designed to switch unpromising OLA queries into the bootstrap scheme for further processing, avoiding the whole dataset scanning resulted from the OLA estimation failure. In addition, we also present a progressive estimation method to reduce the false positive ratio of our dynamic scheme switching mechanism. Moreover, we have implemented our hybrid approximate query processing system in Hadoop, and conducted extensive experiments on the TPC-H benchmark for skewed data distribution. Our results demonstrate that our hybrid system can produce acceptable approximate results within a time period one order of magnitude shorter compared to the original OLA over Hadoop. Yuxiang Wang 0001, Junzhou Luo, Aibo Song, Fang Dong 0001 |
ICPP | 4 |
| 2014 | Game theory based dynamic resource allocation for hybrid environment with cloud and big data applicationabstractVirtualization based cloud and big data applications have been widely adopted in various fields. Because deploying the big data applications on the cloud will cause obvious performance degradation, the cloud and big data applications are provided with fixed resource separately. However, the traditional fixed resource allocation mechanism has two drawbacks: (1) low resource utility and (2) unresponsiveness to the performance degradation. To address these drawbacks, the cloud and big data hybrid environment is designed, where fair resource allocation is used to ensure fairness between cloud and big data applications while virtual machine migration is used to make each virtual machine in cloud application reach its own satisfactory. Herein, game theory is used to model the conflict and negotiation between cloud and big data applications. Firstly, the Nash Equilibrium is used to discover the best strategy for both applications. Secondly, as for virtual machine migration, we use Nash Bargaining game to present the situation where virtual machines compete for more resources allocation while their minimal demand is ensured. Finally, experiments are carried out to prove that the hybrid environment outperforms the traditional method both in resource utility and application performance. Junxue Zhang 0001, Fang Dong 0001, Dian Shen, Junzhou Luo |
SMC | 2 |
| 2014 | OATS: online aggregation with two-level sharing strategy in cloud
Yuxiang Wang 0001, Junzhou Luo, Aibo Song, Fang Dong 0001 |
Distributed Parallel Databases | 4 |
| 2013 | Scheduling Parallel Task Graphs on non-dedicated heterogeneous multicluster platform with Moldable Task DuplicationabstractWorkflow applications structured as Parallel Task Graphs (PTG) exhibit both data and task parallelism and arise in scientific and industrial domains. Most of previous works regarding PTG scheduling only target dedicated multicluster platform. In this paper we develop a scheduling algorithm, MTD (Moldable Task Duplication with forward migration of duplicated predecessors), which applies to non-dedicated heterogeneous multicluster platforms. Our novel contribution is that in MTD, dynamic critical task determination accounts for the heterogeneity and fluctuations of multicluster platform within the hypothetical deadline, and the strategy of moldable task duplication with forward migrations of duplicated predecessors is invented to fully exploit the flexibility of data-parallel tasks. Simulations show that our approach can achieve better average PTG makespan than its competitors. Jinghui Zhang 0001, Junzhou Luo, Fang Dong 0001 |
CSCWD | 3 |
| 2013 | Dynamic Resource Management in a HPC and Cloud Hybrid Environment
Fang Dong 0001, Junzhou Luo |
ICA3PP (1) | 2 |
| 2013 | A Personalized Hybrid Recommendation System Oriented to E-Commerce Mass Data in the CloudabstractPersonalized recommendation technology in E-commerce is widespread to solve the problem of product information overload. However, with the further growth of the number of E-commerce users and products, the original recommendation algorithms and systems will face several new challenges: (1) to model user's interests more accurately, (2) to provide more diverse recommendation modes, and (3) to support large-scale expansion. To address these challenges, from the actual demands of E-commerce applications (as Made-in-China website), a personalized hybrid recommendation system, which can support massive data set, is designed and implemented in this paper by using Cloud technology. Hereinto, the recommendation algorithms are designed based on a novel user interesting model for different scenarios, and the massive data parallel processing techniques in Cloud computing is utilized to realize the effective execution of recommendation algorithms. Finally, several experiments are presented to highlight the system performance. Fang Dong 0001, Junzhou Luo, Yuxiang Wang 0001, Jun Shen 0001 |
SMC | 1 |
| 2013 | Partition-Based Online Aggregation with Shared Sampling in the Cloud
Yuxiang Wang 0001, Junzhou Luo, Aibo Song, Fang Dong 0001 |
J. Comput. Sci. Technol. | 4 |
| 2013 | Scheduling of scientific workflow in non-dedicated heterogeneous multicluster platform
Jinghui Zhang 0001, Junzhou Luo, Fang Dong 0001 |
J. Syst. Softw. | 3 |
| 2012 | Improving Online Aggregation Performance for Skewed Data Distribution
Yuxiang Wang 0001, Junzhou Luo, Aibo Song, Jiahui Jin 0001, Fang Dong 0001 |
DASFAA (1) | 5 |
| 2012 | TASS: Transaction Assurance in Service SelectionabstractAs there are various risks of failure when Web Services are deployed in unreliable environment, the execution of a composite service requires the assurance of the transaction mechanism. However, existing QoS-aware composition approaches do not consider the transactional constraints during service selection. We address this issue by considering the combination of transactional and QoS requirements. Firstly, the novel construction and processing rules are proposed to guarantee the atomic consistency of the composite service and the correctness of these rules is proved subsequently. Then, on basis of these rules, an Ant Colony System based service selection algorithm is presented to guarantee the end-to-end QoS constraints on the premise of ensuring the atomic consistency during service selection. Therein, an optimization strategy is suggested to shrink the searching space of the algorithm tremendously. Finally, experimental results show the efficiency and effectiveness of the algorithm and demonstrate further the correctness of the construction and processing rules through simulations. Jiuxin Cao, Gongrui Zhu, Bo Liu 0004, Fang Dong 0001 |
ICWS | 5 |
| 2012 | Performance evaluation and analysis of SEU Cloud Computing Platform - Using general benchmarks and real world AMS applicationabstractCloud computing, as a popular technique to support and achieve CSCW, is gaining increasing importance in recent years, where the virtualization becomes the key technique. However, although utilizing virtualization can implement more efficient and flexible resource allocation, it may also come at the cost of increased system complexity and dynamics. In order to effectively adapt to performance fluctuations for ensuring high-performance, a generic approach to predict the performance influences of cloud platforms is highly desirable. To address this request, in this paper, the major factors that affect the performance of cloud and the relevant variation discipline are evaluated and analyzed thoroughly using a series of benchmarks in SEU (Southeast University) Cloud Computing Platform, where not only a general methodology on quantifying the performance influence but also the most important impact factors are proposed. Moreover, we use a real world application, as AMS experiment, to further evaluate the relevant performance. Fang Dong 0001, Junzhou Luo, Jiahui Jin 0001 |
SMC | 1 |
| 2012 | An effective data aggregation based adaptive long term CPU load prediction mechanism on computational grid
Fang Dong 0001, Junzhou Luo, Aibo Song, Jiuxin Cao, Jun Shen 0001 |
Future Gener. Comput. Syst. | 1 |
| 2011 | BAR: An Efficient Data Locality Driven Task Scheduling Algorithm for Cloud ComputingabstractLarge scale data processing is increasingly common in cloud computing systems like MapReduce, Hadoop, and Dryad in recent years. In these systems, files are split into many small blocks and all blocks are replicated over several servers. To process files efficiently, each job is divided into many tasks and each task is allocated to a server to deals with a file block. Because network bandwidth is a scarce resource in these systems, enhancing task data locality(placing tasks on servers that contain their input blocks) is crucial for the job completion time. Although there have been many approaches on improving data locality, most of them either are greedy and ignore global optimization, or suffer from high computation complexity. To address these problems, we propose a heuristic task scheduling algorithm called Balance-Reduce(BAR), in which an initial task allocation will be produced at first, then the job completion time can be reduced gradually by tuning the initial task allocation. By taking a global view, BAR can adjust data locality dynamically according to network state and cluster workload. The simulation results show that BAR is able to deal with large problem instances in a few seconds and outperforms previous related algorithms in term of the job completion time. Jiahui Jin 0001, Junzhou Luo, Aibo Song, Fang Dong 0001, Runqun Xiong |
CCGRID | 4 |
| 2011 | Load-aware based adaptive rescheduling mechanism for workflow applicationabstractIn order to integrate the massive distributed resources to accomplish the complex engineering applications cooperatively, workflow scheduling is an important aspect. However, as the available computing power of Grid resources is changing dynamically, static scheduling scheme will lead to low performance in real Grid environment. Therefore, the rescheduling mechanism should be taken into consideration. Although a few relevant mechanisms have been proposed in recent year, as they do not consider the essence of dynamic feature and the relevant algorithms are too simple, they can not obtain a good enough result yet. To address these problems, a load-aware based adaptive rescheduling mechanism for DAG application called LAR is proposed. Therein, during application running, the load exception will be detected and the execution state of application will be judged to decide whether the rescheduling process should be triggered. And in rescheduling stage, an effective rescheduling algorithm which utilizes the latest prediction information is present. The simulation results show that our mechanism can outperform the relevant algorithms in NRSL, and can effectively solve the performance decreasing problem in real Grid environment. Fang Dong 0001, Junzhou Luo, Aibo Song, Jiuxin Cao |
CSCWD | 1 |
| 2011 | QoS Preference-Aware Replica Selection Strategy Using MapReduce-Based PGA in Data GridsabstractData replication is an important technique to reduce access latency and bandwidth consumption in Grid environment. As one of the major functions of data replication, replica selection determines the best replica according to some specific criteria in Data Grid environment, where the data resources are limited and Grid users compete for these resources. In this paper, we focus mainly on a novel QoS preference-aware replica selection strategy which will meet individual QoS sensitivity (IQS) constraints for different users/applications. We first present a framework that characterize QoS properties of replica services and establish its mathematical model by introducing quantification methods. In order to deal with the IQS constraints and to perceive Grid users' QoS preferences accurately, we propose a QoS preference acquisition algorithm based on Analytic Hierarchy Process (AHP). We then design and implement a novel effective and efficient parallel genetic algorithm (PGA) based on Map Reduce paradigm for optimizing the objective function which corresponds to the optimal replica. Simulation results show that our strategy has a better performance in validity as well as scalability, and the optimal replica can always be obtained for Grid users with different IQS constraints under Data Grid environments that vary in system loads, scheduling strategies and user types. Runqun Xiong, Junzhou Luo, Aibo Song, Bo Liu 0004, Fang Dong 0001 |
ICPP | 5 |
| 2010 | Data Aggregation based Adaptive Long term load Prediction mechanism in Grid environmentabstractIn recent years, as a popular technique to support CSCW, Grid computing is becoming more and more attractive. Hereinto, as the CPU load information can guide task scheduling process greatly, the long-term CPU load prediction becomes a very hot research field and has been widely studied. However, as the prediction errors will be accumulated gradually and meanwhile the relevant parameters' optimal values may change dynamically with the variance of load series, the previous prediction algorithms usually can not obtain good prediction accuracy when the length of prediction interval is quite large. To address these feature, a Data Aggregation based Adaptive Long term load Prediction mechanism called DA2LP is proposed in this paper. Therein, in order to reduce the number of prediction step and increase the amount of useful input load information, the data aggregation concept is introduced to integrate with AR model. Meanwhile, with the observation and analysis of the relevant parameters' impact on prediction accuracy in our prediction model, an adaptive parameter selection mechanism is proposed, where the optimal relevant parameters can be adapted automatically to enhance prediction accuracy during the prediction process. The experiments show that our proposed mechanism can outperform significantly the previous prediction methods in mean square error (MSE) for long term load prediction. Fang Dong 0001, Junzhou Luo, Aibo Song, Jiuxin Cao |
CSCWD | 1 |
| 2010 | SLA-Based Resource Co-Allocation in Multi-Cluster GridabstractResource co-allocation is a crucial but challenging problem for Grid Computing. With the emergence of WS-Resource Framework and Open Grid Services Architecture, resource co-allocation is commonly associated with a service level agreement (SLA) to determine the achieved QoS level. In this paper, we present an approach for resource co-allocation in Multi-cluster Grid that maximizes the user satisfaction degree while satisfying the QoS requirements defined in a SLA. The QoS metrics considered in this paper include deadline, cluster availability, service reliability and budget. Experimental results are presented to show the effectiveness of our approach. Wei Wang 0089, Junzhou Luo, Aibo Song, Fang Dong 0001 |
GLOBECOM | 4 |
| 2010 | Resource Load Based Stochastic DAGs Scheduling Mechanism for Grid EnvironmentabstractThe dynamic feature is one of the most important differences between Grid and traditional heterogeneous distributed systems, thus the most significant challenge for task scheduling in Grid environment is how to relieve the resource performance dynamism effectively. However, the existing schedule algorithms usually suppose that computation or communication times are deterministic and static, thus they will lead to bad performance in the practical Grid environment. To address this problem, a mechanism which is used to estimate the probability distribution of task execution time based on resource load is proposed. And then a Resource Load based Stochastic DAGs Scheduling algorithm for Grid environments is introduced. The simulation results show that our mechanism can achieve a significant improvement in several metrics (such as normalized real schedule length) and can relieve the influence brought by the dynamic nature of Grid effectively. Fang Dong 0001, Junzhou Luo, Aibo Song, Jiahui Jin 0001 |
HPCC | 1 |
| 2010 | A novel task scheduling algorithm based on dynamic critical path and effective duplication for pervasive computing environmentabstractAbstract In order to effectively utilize massive heterogeneous resources and provide transparent computing capability to upper applications, task scheduling as the key issue of pervasive computing system becomes significantly important. Previous proposed priority and duplication based task scheduling algorithms, which can be applied in pervasive computing environment, usually have following limitations: critical path cannot be calculated accurately while neglecting the effect of resource availability in scheduling; in duplication based resource allocation stage, duplications without restriction would lead to some negative effects on final schedule length (SL). For the purpose of solving these problems, a novel task scheduling algorithm based on dynamic critical path (DCP) and effective duplication, called DCPED, is presented in this paper. In DCPED, a more accurate DCP calculation method which takes resource availability into account is introduced. Meanwhile an effective task duplication strategy is proposed to eliminate ineffective duplications and make an optimized schedule result by using space compression technique and dynamic critical path length (DCPL) based evaluation technique respectively. Finally, simulation results show that DCPED can outperform previous algorithms significantly in NSL and speedup rate metrics. Especially, it is very effective for utilizing computing resources and scheduling the fine‐grain and large‐scale workflow applications in pervasive computing system. Copyright © 2008 John Wiley & Sons, Ltd. Junzhou Luo, Fang Dong 0001, Jiuxin Cao, Aibo Song |
Wirel. Commun. Mob. Comput. | 2 |
| 2007 | QoS Matching Offset Oriented Resource Clustering Scheduling Algorithm in Grid EnvironmentabstractWith more and more research carried on in the QoS of Grid, QoS-based Grid task scheduling algorithm has become a hot research aspect. In this paper, various existing QoS-based Grid scheduling algorithms are analyzed firstly. And by introducing the conception of QoS matching offset between tasks and resources in Grid scheduling, resources and tasks can be clustering upon their offset in order that the resources in scheduling are able to be allocated on demand. Meanwhile, we take into consideration some parameters in the scheduling like QoS benefit value obtained by the task and some restricted condition: deadline of the task. The simulation results show that the algorithm's performance is better than most of the proposed algorithms in the aspects of effective resource utility, resource load balance, task acceptance rate and average QoS benefit value. Fang Dong 0001, Junzhou Luo |
CSCWD | 1 |
| 2007 | QoS Deviation Distance Based Negotiation Algorithm in Grid Resource Advance ReservationabstractHow to guarantee user's QoS (Quality of Service) demands becomes increasingly important in service-oriented grid environment. Current research on grid resource advance reservation, a well-known and effective mechanism to guarantee QoS, forces on proposing theoretic architecture to support advance reservation. Detailed research on advance reservation is scarce. For that, SNAP (Service Negotiation and Acquisition Protocol) is extended to support QoS description in fine granularity and new calculation method of QoS deviation distance is proposed. Then, advance reservation state switching is analyzed and QoS deviation distance based negotiation algorithm is addressed. Preliminary results show that proposed negotiation algorithm that considers QoS deviation distance can produce remarkable improvement in user satisfaction and resource utilization. Zhiang Wu 0001, Junzhou Luo, Fang Dong 0001, Xudong Ni |
CSCWD | 3 |