VLDB 2026 Research / reviewers in the wild / expert
Dong Yuan 0001
dblp:15/2771-1
· DBLP profile ↗
123ranked-venue papers
11as first author
59since 2021 · last 2026
0000-0003-1130-0888ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 37 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 20 · 18 since 2021Software engineering, systems software and programming languages · 20 · 1 first-author · 3 since 2021Computer networks · 12 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 11 since 2021Security and privacy · 11 · 7 since 2021Databases, data management, data science and information retrieval · 9 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scaling Law Analysis in Federated Learning: How to Select the Optimal Model Size?abstractThe recent success of large language models (LLMs) has sparked a growing interest in training large-scale models. As the model size continues to scale, concerns are growing about the depletion of high-quality, well-curated training data. This has led practitioners to explore training approaches like Federated Learning (FL), which can leverage the abundant data on edge devices while maintaining privacy. However, the decentralization of training datasets in FL introduces challenges to scaling large models, a topic that remains under-explored. This paper fills this gap and provides qualitative insights on generalizing the previous model scaling experience to federated learning scenarios. Specifically, we derive a PAC-Bayes (Probably Approximately Correct Bayesian) upper bound for the generalization error of models trained with stochastic algorithms in federated settings and quantify the impact of distributed training data on the optimal model size by finding the analytic solution of model size that minimizes this bound. Our theoretical results demonstrate that the optimal model size has a negative power law relationship with the number of clients if the total training compute is unchanged. Besides, we also find that switching to FL with the same training compute will inevitably reduce the upper bound of generalization performance that the model can achieve through training, and that estimating the optimal model size in federated scenarios should depend on the average training compute across clients. Furthermore, we also empirically validate the correctness of our results with extensive training runs on different models, network settings, and datasets. Xuanyu Chen, Dong Yuan 0001 |
AAAI | 4 |
| 2026 | Optimal Look-back Horizon for Time Series Forecasting in Federated LearningabstractSelecting an appropriate look-back horizon remains a fundamental challenge in time series forecasting (TSF), particularly in federated learning scenarios where data is decentralized, heterogeneous, and often non-independent. While recent work has explored horizon selection by preserving forecasting-relevant information in an intrinsic space, these approaches are primarily restricted to centralized and independently distributed settings. This paper presents a principled framework for adaptive horizon selection in federated time series forecasting through an intrinsic space formulation. We introduce a synthetic data generator that captures essential temporal structures in client data, including autoregressive dependencies, seasonality, and trend, while incorporating client-specific heterogeneity. Building on this model, we define a transformation that maps time series windows into an intrinsic representation space with well-defined geometric and statistical properties. We then derive a decomposition of the forecasting loss into a Bayesian term, which reflects irreducible uncertainty, and an approximation term, which accounts for finite-sample effects and limited model capacity. Our analysis shows that while increasing the look-back horizon improves the identifiability of deterministic patterns, it also increases approximation error due to higher model complexity and reduced sample efficiency. We prove that the total forecasting loss is minimized at the smallest horizon where the irreducible loss starts to saturate, while the approximation loss continues to rise. This work provides a rigorous theoretical foundation for adaptive horizon selection for time series forecasting in federated learning. Dahao Tang, Zhibo Jin, Dong Yuan 0001 |
AAAI | 6 |
| 2026 | MosaicUFL: A Fine-grained Unlabeled Client Selection Framework for Efficient Federated Learning
Zikang Wen, Huaming Chen, Dong Yuan 0001 |
IPDPS | 6 |
| 2026 | K&L: Penetrating Backdoor Defense with Key and Locks
Xinyi Wang 0005, Jiayu Zhang 0001, Zhibo Jin, Dong Yuan 0001, Huaming Chen |
WWW | 5 |
| 2026 | Optimal schedule for periodic jobs with discretely controllable processing times on two machinesabstractIn many real-world situations, the processing time of computational jobs can be shortened by lowering the processing quality. This is referred to as discretely controllable processing time, where the original processing time can be shortened to a number of levels with lower processing qualities. In this paper, we study the scheduling problem of periodic jobs with discretely controllable processing times on two machines. The problem is NP-hard, as directly solving it through dynamic programming leads to exponential computational complexity. This is because we need to memorise a set of processed jobs to avoid reprocessing. In order to address this issue, we prove the Ordered Scheduling Structure (OSS) Property and the Consecutive Decision Making (CDM) Property. The OSS Property allows us to search for an optimal solution in which jobs on the same machine are orderly started. The CDM Property allows us to memorise only two jobs to completely avoid the job reprocessing. These two properties greatly decrease the searching space, and the resultant dynamic programming solution to find an optimal solution is with pseudo-polynomial computational complexity. Wei Bao 0001, Dong Yuan 0001, Albert Y. Zomaya |
J. Parallel Distributed Comput. | 4 |
| 2026 | A Hybrid Online and Offline Requests Inference Serving System for LLM in Private Computer EnvironmentabstractWhile advancements in Large Language Models (LLMs) have broadened their applications, performing multitask LLM inference on a single GPU remains challenging due to insufficient GPU memory to load all model parameters. Existing methods that offload unneeded parameters to main memory and prefetch them back introduce high latency due to data transfer overhead. We propose FixGen, a single GPU LLM online-offline mixed inference serving system that supports multitask inference. First, we develop FixPool to optimize memory management by centralizing storage in a fixed memory space and optimizing parameter storage procedures, thus reducing PCIe resource consumption and improving efficiency. Secondly, we use a router to select appropriate operators to compute requests in the prefill and decode stages within the same batch to reduce the response latency for sudden online requests. FixGen reduces the online request-response latency between 1.06 and 1.84 times compared to the conventional method while maintaining the throughput to offline requests on a single GPU from OPT-6.7b to OPT-30b. Dong Yuan 0001 |
IEEE Trans. Computers | 3 |
| 2026 | FedSCOPE: A Comprehensive Evaluation Framework for Federated Learning in Human-Centered Social Computing
Dong Yuan 0001, Weiping Ding 0001 |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2026 | Joint Optimization of Resource Allocation and Request Batching for Multi-Tenant Inference Serving on GPUabstractThe GPU technology has significantly aided Deep Learning (DL), especially in enhancing the performance of inference services. Tenants deploy inference models on the GPU, which are then uniformly scheduled and executed by an inference serving system. In resource-constrained environments, a single GPU needs to handle requests from multiple tenants. The diver sity of inference tasks, varying request frequencies, and different model architectures make designing an efficient inference serving system a significant challenge. Most current research discusses resource allocation and request batching separately, overlooking the critical connection between them. In such complex inference environments, this connection is particularly crucial. To rapidly process requests from various tenants in such a dynamic environment, we leverage the connection between resource allocation and request batching to design DRS: Deep Reinforcement Scheduler. In DRS, we use the Deep Deterministic Policy Gradient (DDPG) as our scheduling algorithm and NVIDIA Multi-Process Service (MPS) for spatial parallelism in sharing a single GPU among multiple tenants. By observing environmental information, we can rapidly adjust the GPU allocation for different tenants and f ind the proper request batch size, thereby maintaining high efficiency. In experiments, DRS achieves a speedup of up to 2.23× and 24 × compared to the baselines with the Makespan and Job Completion Time (JCT) metrics. Dong Yuan 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | PUMPS: Skeleton-Agnostic Point-Based Universal Motion Pre-Training for Synthesis in Human Motion Tasks
Clinton Mo, Kun Hu 0008, Chengjiang Long, Dong Yuan 0001, Wan-Chi Siu, Zhiyong Wang 0001 |
ICCV | 4 |
| 2025 | CrossFL: A Cross-Round Federated Learning Framework with Asynchronous Client Selection
Zikang Wen, Zihao Yao, Huaming Chen, Dong Yuan 0001 |
PAKDD (2) | 7 |
| 2025 | DBSSL: A Scheme to Detect Backdoor Attacks in Self-Supervised Learning ModelsabstractRecently, self-supervised learning has garnered significant attention for its ability to extract high-quality features from unlabeled data. However, existing research indicates that backdoor attacks can pose significant threats to self-supervised learning. Additionally, due to the substantial training overhead, traditional backdoor detection methods in supervised learning, such as meta-learning, are not well-suited for self-supervised learning. Current backdoor detection methods for self-supervised learning can only defend against backdoor attacks triggered by patch. To address this challenge, we proposes DBSSL, a scheme that can efficiently detect backdoor attacks in self-supervised learning models. In DBSSL, we connect backdoor attacks with adversarial attacks and use adversarial perturbations for detection. Specifically, we first conduct targeted adversarial attacks on the samples, classifying them into different classes. Then, we employ the Z-Score method to analyze these adversarial perturbations. If outliers are present, we can infer that the self-supervised learning model has been backdoored. Unlike previous works, our scheme is better suited for downstream users with constrained computation resources and exhibits excellent detection capabilities for backdoor attacks caused by patch-type or global triggers. Abundant theoretical analysis has demonstrated the feasibility of our scheme. Moreover, extensive experiments show that our method performs well in detecting prevalent backdoor attacks in self-supervised learning, achieving detection accuracy exceeding 95%. Yuxian Huang, Geng Yang 0002, Dong Yuan 0001, Shui Yu 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | On Security Weaknesses and Vulnerabilities in Deep Learning SystemsabstractThe security guarantee of AI-enabled software systems (particularly using deep learning techniques as a functional core) is pivotal against the adversarial attacks exploiting software vulnerabilities. However, little attention has been paid to a systematic investigation of vulnerabilities in such systems. A common situation learned from the open source software community is that deep learning engineers frequently integrate off-the-shelf or open-source learning frameworks into their ecosystems. In this work, we specifically look into deep learning (DL) framework and perform the firstsystematicstudy of vulnerabilities in DL systems through a comprehensive analysis of identified vulnerabilities from Common Vulnerabilities and Exposures (CVE) and open-source DL tools, including TensorFlow, Caffe, OpenCV, Keras, and PyTorch. We propose a two-stream data analysis framework to explore vulnerability patterns from various databases. We investigate the unique DL frameworks and libraries development ecosystems that appear to be decentralized and fragmented. By revisiting the Common Weakness Enumeration (CWE) List, which provides the traditional software vulnerability related practices, we observed that it is more challenging to detect and fix the vulnerabilities throughout the DL systems lifecycle. Moreover, we conducted a large-scale empirical study of 3,049 DL vulnerabilities to better understand the patterns of vulnerability and the challenges in fixing them. Zhongzheng Lai, Huaming Chen, Ruoxi Sun 0001, Yu Zhang 0177, Minhui Xue 0001, Dong Yuan 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2025 | Threats and Defenses in the Federated Learning Life Cycle: A Comprehensive Survey and ChallengesabstractFederated learning (FL) offers innovative solutions for privacy-preserving distributed machine learning (ML). Different from centralized data collection algorithms, FL enables participants to locally train their model and only share the model updates for aggregation. Since private data never leaves the end node, FL effectively mitigates privacy leakage during collaborative training. Despite its promising potential, FL is vulnerable to various attacks due to its distributed nature, affecting the entire life cycle of FL services. These threats can harm the model's utility or compromise participants' privacy, either directly or indirectly. In response, numerous defense frameworks have been proposed, demonstrating effectiveness in specific settings and scenarios. To provide a clear understanding of the current research landscape, this article reviews the most representative and state-of-the-art threats and defense frameworks throughout the FL service life cycle. We start by identifying FL threats that harm utility and privacy, including those with potential or direct impacts. Then, we dive into the defense frameworks, analyze the relationship between threats and defenses, and compare the trade-offs among different defense strategies. We subsequently revisit these studies to evaluate their practicality in real-world scenarios and conclude by summarizing existing research bottlenecks and outlining future directions. We hope this survey sheds light on trustworthy FL research and contributes to the FL community. Zhongliang Guo 0001, Huaming Chen, Dong Yuan 0001, Weiping Ding 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Adaptformer: An Adaptive Multimodal Deep Decomposition Approach for Power Consumption Forecasting
Dahao Tang, Dong Yuan 0001 |
ADMA (5) | 6 |
| 2024 | Holistic Evaluation Metrics for Federated LearningabstractFederated Learning (FL) is a promising technology that enables multiple participants collaboratively train a joint model without sharing their local data. Owing the privacy protection nature, FL has attracted interests from industry, leading to its deployment across diverse domains such as smartphones, institutions, and Internet of Things (IoTs). Despite the development of various FL algorithms aimed at enhancing FL performance from multiple perspectives, their evaluation typically hinges on a single metric like accuracy, failing to account for the unique demands of different use cases. Thus, how to comprehensively evaluate an FL algorithm and determine the most suitable candidate for a designated use case remains an open question. To address this research gap, we introduce the Holistic Evaluation Metrics (HEM) for federated learning. In this work, we consider the application scenarios of IoT, Smartphones, and Institutions as the principal FL use cases. Subsequently, we determine the components of the evaluation metric and the corresponding importance vector for each use case. The HEM index is then generated by integrating these metric components with their importance vectors. We evaluate various FL algorithms in different use cases through HEM, our experimental results demonstrate that HEM can effectively evaluate and select the appropriate FL algorithms in different use cases. Jehad Ibrahim, Huaming Chen, Dong Yuan 0001 |
CSCWD | 4 |
| 2024 | Motion Keyframe Interpolation for Any Human Skeleton via Temporally Consistent Point Cloud Sampling and Reconstruction
Clinton Mo, Kun Hu 0008, Chengjiang Long, Dong Yuan 0001, Zhiyong Wang 0001 |
ECCV (82) | 4 |
| 2024 | CAKD: A Correlation-Aware Knowledge Distillation Framework Based on Decoupling Kullback-Leibler DivergenceabstractIn knowledge distillation, a primary focus has been on transforming and balancing multiple distillation components. In this work, we emphasize the importance of thoroughly examining each distillation component, as we observe that not all elements are equally crucial. From this perspective, we decouple the Kullback-Leibler (KL) divergence into three unique elements: Binary Classification Divergence (BCD), Strong Correlation Divergence (SCD), and Weak Correlation Divergence (WCD). Each of these elements presents varying degrees of influence. Leveraging these insights, we present the Correlation-Aware Knowledge Distillation (CAKD) framework. CAKD is designed to prioritize the facets of the distillation components that have the most substantial influence on predictions, thereby optimizing knowledge transfer from teacher to student models. Our experiments demonstrate that adjusting the effect of each element enhances the effectiveness of knowledge transformation. Furthermore, evidence shows that our novel CAKD framework consistently outperforms the baseline across diverse models and datasets. Our work further highlights the importance and effectiveness of closely examining the impact of different parts of distillation process. Zao Zhang, Huaming Chen, Pei Ning, Dong Yuan 0001 |
ICDM | 5 |
| 2024 | Radio Frequency Signal based Human Silhouette Segmentation: A Sequential Diffusion ApproachabstractRadio frequency (RF) signals have been proved to be flexible for human silhouette segmentation (HSS) under complex environments. Existing studies are mainly based on a one-shot approach, which lacks a coherent projection ability from the RF domain. Additionally, the spatio-temporal patterns have not been fully explored for human motion dynamics in HSS. Therefore, we propose a two-stage Sequential Diffusion Model (SDM) to progressively synthesize high-quality segmentation jointly with the considerations on motion dynamics. Cross-view transformation blocks are devised to guide the diffusion model in a multi-scale manner for comprehensively characterizing human related patterns in an individual frame such as directional projection from signal planes. Moreover, spatio-temporal blocks are devised to fine-tune the frame-level model to incorporate spatio-temporal contexts and motion dynamics, enhancing the consistency of the segmentation maps. Comprehensive experiments on a public benchmark - HIBER demonstrate the state-of-the-art performance of our method with an IoU 0.732. Our code is available at https://github.com/ph-w2000/SDM. Penghui Wen, Kun Hu 0008, Dong Yuan 0001, ChangYang Li, Zhiyong Wang 0001 |
ICME | 3 |
| 2024 | Edge-LLM: A Collaborative Framework for Large Language Model Serving in Edge ComputingabstractThe rapid advancement and extensive implementation of Large Language Models (LLMs) are milestones in the realm of artificial intelligence. Although Parameter-Efficient Transfer Learning (PETL), a.k.a. Adapter, methods have reduced the barrier for fine-tuning and inference on LLMs, it becomes a challenge to efficiently deploy and fine-tuning different adapter models needed for massive AI applications. With the popularity of SoC chips, the computing power of edge devices has improved significantly. To meet the computational resources required by LLM applications and improve quality of service (QoS), we propose Edge-LLM, a server-node collaboration framework for large-scale language model serving, to efficiently utilize edge resources to accelerate LLM fine-tuning and inference in resource-constrained scenarios. In the framework, we implement an adaptive quantization strategy, FM cache mechanism, and value density first (VDF) scheduling algorithm to reduce GPU overhead and accelerate LLM computation. The experimental results demonstrate that Edge-LLM can significantly improve overall computational speed by a factor of 17, decrease the number of tasks experiencing timeouts by 63%, and reduce GPU overhead by up to 43%. Fenglong Cai, Dong Yuan 0001, Li-Zhen Cui 0001 |
ICWS | 2 |
| 2024 | GE-AdvGAN: Improving the transferability of adversarial samples by gradient editing-based adversarial generative modelabstractAdversarial generative models, such as Generative Adversarial Networks (GANs), are widely applied for generating various types of data, i.e., images, text, and audio. Accordingly, its promising performance has led to the GAN-based adversarial attack methods in the white-box and black-box attack scenarios. The importance of transferable black-box attacks lies in their ability to be effective across different models and settings, more closely aligning with real-world applications. However, it remains challenging to retain the performance in terms of transferable adversarial examples for such methods. Meanwhile, we observe that some enhanced gradient-based transferable adversarial attack algorithms require prolonged time for adversarial sample generation. Thus, in this work, we propose a novel algorithm named GE-AdvGAN to enhance the transferability of adversarial samples whilst improving the algorithm's efficiency. The main approach is via optimising the training process of the generator parameters. With the functional and characteristic similarity analysis, we introduce a novel gradient editing (GE) mechanism and verify its feasibility in generating transferable samples on various models. Moreover, by exploring the frequency domain information to determine the gradient editing direction, GE-AdvGAN can generate highly transferable adversarial samples while minimizing the execution time in comparison to the state-of-the-art transferable adversarial attack algorithms. The performance of GE-AdvGAN is comprehensively evaluated by large-scale experiments on different datasets, which results demonstrate the superiority of our algorithm. The code for our algorithm is available at: https://github.com/LMBTough/GE-advGAN. Huaming Chen, Xinyi Wang 0005, Jiayu Zhang 0001, Zhibo Jin, Kim-Kwang Raymond Choo, Jun Shen 0001, Dong Yuan 0001 |
SDM | 8 |
| 2024 | VPPFL: A verifiable privacy-preserving federated learning scheme against poisoning attacks
Yuxian Huang, Geng Yang 0002, Hao Zhou 0034, Hua Dai 0003, Dong Yuan 0001, Shui Yu 0001 |
Comput. Secur. | 5 |
| 2024 | Coarse-to-Fine: A hierarchical DNN inference framework for edge computingabstractDeep neural networks (DNNs) have been increasingly used in recent years to achieve higher inference accuracy; however, implementing deeper networks in edge-computing environments can be challenging. Current methods for accelerating CNN inference focus on finding a trade-off between accuracy and latency under an assumed uniform distribution, ignoring the impact of real-world data distributions. To address this, we propose the Coarse-to-Fine (C2F) framework, which includes a C2F model and a corresponding C2F inference architecture to better exploit distributional differences in the edge environment. The C2F model is derived from various adaptations of Convolutional Neural Networks (CNNs). By deconstructing the original CNNs into multiple smaller models, the C2F model increases memory consumption within an acceptable range to improve inference speed without sacrificing accuracy. The C2F architecture deploys C2F models more logically in complex edge environments, reducing inference costs and memory consumption. We conduct experiments on the CIFAR dataset with different backbone networks and show that our C2F framework can simultaneously reduce latency and improve accuracy in complex edge environments. Zao Zhang, Wei Bao 0001, ChangYang Li, Dong Yuan 0001 |
Future Gener. Comput. Syst. | 5 |
| 2024 | Joint Video Denoising and Super-Resolution Network for IoT CamerasabstractIoT (Internet of Things) cameras have widely been deployed over the last few years. These cameras are often with limited hardware so that they can only capture noisy videos in low resolution. In this work, we propose the joint video denoising and super-resolution network for IoT cameras, which consists of the noise-robust moving-attention (NRMA) module and the noise-eliminated upsampling (NEU) module. In NRMA, we adopt a coarse-to-fine approach by first extracting the coarse flow and then refining through bi-directional feature propagation among adjacent frames. In NEU, we further utilize inner-frame features for noise-elimination and upsampling. Through this approach, we avoid the negative effects brought by applying denoising and super-resolution in tandem, and enhance the reconstruction of moving objects by the embedded attention layers in NRMA. We conduct our experiments on both synthetic datasets, which utilize existing data with additive white Gaussian noise (AWGN), and a realistic dataset captured using a pair of IoT and professional cameras. Our extensive experimental results demonstrate that our proposed method significantly reduces noise and enhances detail in both types of datasets. Notably, our approach outperforms the state-of-the-art benchmark (RealBasicVSR) by an average of 5.24 dB on the existing datasets (with noise level σ = 20) and by 0.95 dB on the realistic dataset in terms of PSNR. Liming Ge, Wei Bao 0001, Xinyi Sheng, Dong Yuan 0001, Bing Bing Zhou, Zhiyong Wang 0001 |
IEEE Internet Things J. | 4 |
| 2024 | Contribution-wise Byzantine-robust aggregation for Class-Balanced Federated Learning
Weiping Ding 0001, Huaming Chen, Wei Bao 0001, Dong Yuan 0001 |
Inf. Sci. | 5 |
| 2024 | FastPTM: Fast weights loading of pre-trained models for parallel inference service provisioning
Fenglong Cai, Dong Yuan 0001, Wei He 0020, Wei Guo 0017, Li-Zhen Cui 0001 |
Parallel Comput. | 2 |
| 2024 | Disentangled Orchestration on Cyber RangesabstractCyber ranges require networked applications to test cyberspace events effectively. As testing becomes more advanced, it involves multiple real-world applications with flexible execution orders. However, it is increasingly challenging to orchestrate large-scale, chained, and heterogeneous Internet applications. State-of-the-art orchestration techniques face scalability issues due to inefficient representation models and entangled scheduling of events and applications. To address these issues, we present Wukong, a disentangled orchestration system in cyber ranges that disaggregates the scheduling and execution of workflows and their applications in a decentralized coordination approach. First, we overcome the heterogeneity of events with a workflow model that encodes event chains with compositional Directed Acyclic Graphs (DAGs) and unified event triggers. Second, Wukong disaggregates the execution of DAGs and applications with push-pull decentralized coordination over distributed agents. Our evaluation of Wukong on a real-world cyber range demonstrates its expressive, scalable, and efficient abilities for automatically emulating diverse event chains. The storage footprint of compositional modeling is up to 57 times smaller than that of baseline models. Wukong's response delay is 1.52 to 2.74 times shorter than state-of-the-art orchestration engines, and the scheduling delay is up to 2.16 times smaller than the baseline approach. Yongquan Fu, Weihong Han, Dong Yuan 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2023 | Real-EVE: Real-Time Edge-Assist Video Enhancement for Joint Denoising and Super-Resolution
Liming Ge, Wei Bao 0001, Dong Yuan 0001, Bing Bing Zhou |
ICA3PP (1) | 3 |
| 2023 | Hierarchical Federated Learning with Adaptive Momentum in Multi-Tier NetworksabstractIn this paper, we propose and analyze HierAdMo, a three-tier adaptive momentum accelerated client-edge-cloud Federated Learning (FL) algorithm. HierAdMo combines the momentum acceleration on both worker and edge levels. However, simply combining these two levels of momenta may lead to disagreement between them, negatively influencing convergence performance. To this end, we embed an online adaptive method that scales down the momentum when disagreement occurs. We provide mathematical proof for the convergence of HierAdMo for non-i.i.d. data and the tighter convergence upper bound compared with a version of HierAdMo without adaptation (HierAdMo-R). Finally, extensive experiments based on real-world datasets are conducted, verifying that HierAdMo outperforms existing mainstream benchmarks and achieves the optimal or near-optimal convergence performance compared with HierAdMo-R under a wide range of settings. Zhengjie Yang, Sen Fu, Wei Bao 0001, Dong Yuan 0001, Bing Bing Zhou |
ICDCS | 4 |
| 2023 | Trajectory Alignment based Multi-Scaled Temporal Attention for Efficient Video TransformerabstractAlthough the video transformer gets remarkable accuracy on video recognition tasks, it is hard to be deployed in resource-constrained scenarios due to the high computational cost. A method that dynamically modifies and trains the transformer model, ensuring that the computational cost matches the deployment scenario requirement, would be an effective solution to this challenge. In this paper, we propose a method for modifying large-scale video transformers with trajectory alignment based multi-scaled temporal attention (TAMS) schemes to reduce the computational cost significantly while losing accuracy slightly. In the temporal dimension, we adopt multi-scaled sparsity patterns in hierarchical transformer blocks. In the spatial dimension, we use region selection to force the transformer to focus on high-importance regions while not corrupting the spatial context. Our method reduces up to 40% computational cost of state-of-the-art large-scale video transformers with a slight accuracy drop (~ 7%) on the video recognition task. Zao Zhang, Dong Yuan 0001, Yu Zhang 0177, Wei Bao 0001 |
ICME | 2 |
| 2023 | ITIF: Integrated Transformers Inference Framework for Multiple Tenants on GPUabstractTransformer models, which have gained prominence in recent years, serve as the backbone for a wide range of deep learning applications, from natural language processing to computer vision. These models have become indispensable for various inference tasks, but their implementation often faces challenges related to GPU utilization and system throughput. Current GPU-based inference frameworks typically treat each model individually, leading to suboptimal resource management and reduced performance. In response to these limitations, we introduce ITIF: Integrated Transformers Inference Framework for multi-tenants with a shared backbone. By allowing multiple tenants to share a single backbone Transformer model on a single GPU, ITIF consolidates operators from diverse multi-tenant inference models, which in turn optimizes GPU utilization and system throughput. Our proposed framework, ITIF, represents a significant step towards improving the efficiency of deep learning, particularly for large-scale cloud providers hosting numerous models with a shared backbone. In our experiments, ITIF improved performance by up to 2.40 × compared with the baselines. Zao Zhang, Wei Bao 0001, Dong Yuan 0001 |
ICPP | 4 |
| 2023 | ParaTra: A Parallel Transformer Inference Framework for Concurrent Service Provision in Edge ComputingabstractEdge computing has been widely used to deploy and service deep learning applications. Equipped with GPUs, edge nodes can process concurrent incoming inference requests of the deep learning model. However, existing methods for inference tasks do not allow efficient parallel handling of user requests. This paper investigates the popular Transformer deep learning model and develops ParaTra, a parallel transformer inference framework for providing parallel inference services to users. In the framework, the Transformer model is partitioned and deployed in users’ devices and the edge node to efficiently utilize their processing power. The concurrent inference tasks with different sizes are dynamically packaged in a scheduling queue and sent in batch to an encoder-decoder pipeline for processing. ParaTra can significantly reduce the overheads of parallel processing and the usage of GPU memory. Experiment results show that ParaTra can save up to 37.1% of GPU memory usage and improve 8.4 times of processing speed. Fenglong Cai, Dong Yuan 0001, Mengwei Xie, Wei He 0020, Lanju Kong, Wei Guo 0017, Yali Jiang 0004, Li-Zhen Cui 0001 |
ICWS | 2 |
| 2023 | Real-time Night Surveillance Video Retrieval through Calibrated Denoising and Super-resolutionabstractReal-time video surveillance cameras have been widely deployed over the last few years. In case of incidents such as natural disasters, it provides vital guidance in real time to aid the rescue operations. However, the quality of the captured video is far from satisfactory due to the limited camera hardware and low network bandwidth. Noise is often observed especially at night and the resolution is low. To this end, we are motivated to retrieve the nighttime surveillance video through calibrated denoising and super-resolution. We only use the preceding and current frames, but not the future frames. Thus, we avoid the additional delay of waiting for future frames, which is not suitable for real-time video applications. We design the pipeline semantically beneficial for both denoising and super-resolution, and achieve high rendering quality, especially for real-world noise. Moreover, we propose a novel calibration method for collecting paired noisy and clean observations in the real world, which provides more effective training data. We conduct experiments using the real-world dataset collected under low-light conditions, and benchmark with state-of-the-art video denoising and super-resolution methods. Results show that our method achieves significant performance gain while introducing small delay compared with the benchmarks, suitable for real-time videos. Liming Ge, Wei Bao 0001, Xinyi Sheng, Dong Yuan 0001, Bing Bing Zhou |
IJCNN | 4 |
| 2023 | Enhancing federated learning robustness in adversarial environment through clustering Non-IID featuresabstractFederated Learning (FL) enables many clients to train a joint model without sharing the raw data. While many byzantine-robust FL methods have been proposed, FL remains vulnerable to security attacks such as poisoning attacks and evasion attacks due to its distributed adversarial environment. Additionally, real-world training data used in FL are usually Non-Independent and Identically Distributed (Non-IID), which further weakens the robustness of the existing FL methods (such as Krum, Median, Trimmed-Mean, etc.), thereby making it possible for a global model in FL to be broken in extreme Non-IID scenarios. In this work, we mitigate the aforementioned weaknesses of existing FL methods in Non-IID and adversarial scenarios by proposing a new FL framework called Mini-Federated Learning (Mini-FL). Mini-FL follows the general FL approach but considers the Non-IID sources of FL and aggregates the gradients by groups. Specifically, Mini-FL first performs unsupervised learning for the gradients received to define the grouping policy. Then, the server divides the gradients received into different groups according to the grouping policy defined and performs byzantine-robust aggregation. Finally, the server calculates the weighted mean of gradients from each group to update the global model. Owning the strong generality, Mini-FL can utilize the most existing byzantine-robust method. We demonstrate that Mini-FL effectively enhances FL robustness and achieves greater global accuracy than existing FL methods when against security attacks and in Non-IID settings. Dong Yuan 0001, Abubakar Sadiq Sani, Wei Bao 0001 |
Comput. Secur. | 2 |
| 2023 | A Unifying Probabilistic Framework for Partially Labeled Data LearningabstractPartially labeled data learning (PLDL), including partial label learning (PLL) and partial multi-label learning (PML), has been widely used in nowadays data science. Researchers attempt to construct different specific models to deal with the different classification tasks for PLL and PML scenarios respectively. The main challenge in training classifiers for PLL and PML is how to deal with ambiguities caused by the noisy false-positive labels in the candidate label set. The state-of-the-art strategy for both scenarios is to perform disambiguation by identifying the ground-truth label(s) directly from the candidate label set, which can be summarized into two categories: 'the identifying method' and 'the embedding method'. However, both kinds of methods are constructed by hand-designed heuristic modeling under considerations like feature/label correlations with no theoretical interpretation. Instead of adopting heuristic or specific modeling, we propose a novel unifying framework called A Unifying Probabilistic Framework for Partially Labeled Data Learning (UPF-PLDL), which is derived from a clear probabilistic formulation, and brings existing research on PLL and PML under one theoretical interpretation with respect to information theory. Furthermore, the proposed UPF-PLDL also unifies 'the identifying method' and 'the embedding method' into one integrated framework, which naturally incorporates the feature and label correlation considerations. Comprehensive experiments on synthetic and real-world datasets for both PLL and PML scenarios clearly demonstrate the superiorities of the derived framework. Xiuwen Gong, Dong Yuan 0001, Wei Bao 0001, Fulin Luo |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Geo-Ellipse-Indistinguishability: Community-Aware Location Privacy Protection for Directional DistributionabstractDirectional distribution analysis has long served as a fundamental functionality in abstracting dispersion and orientation of spatial datasets. Spatial datasets that describe sensitive information of individuals such as health status and home addresses must be used and shared cautiously to protect individuals' privacy. There is an inherent tension between the need of accurate directional distribution result and the requirement of individuals' location privacy. Plenty of excellent location privacy protection approaches such as geo-indistinguishability can provide strong protection for locations but considerably at the expense of statistical quality of subsequent directional distribution analysis. In this paper, to protect individual location data for directional distribution, we define the geographic feature of community with covariance matrix and then propose ageo-ellipse-indistinguishabilityprivacy notion incorporating this covariance matrix. As an instantiation of metric differential privacy,geo-ellipse-indistinguishabilityguarantees pairwise inputs cannot be distinguishable with the level proportional to privacy budget and Mahalanobis distance between them, given a randomized output. We also present elliptical privacy mechanisms to achieve this privacy definition on the basis of gamma distribution and multivariate normal distribution. We finally evaluate the empirical utility of the proposed mechanism in New York home addresses database. Our experiments demonstrate that under the same privacy level, our proposed elliptical approach can achieve significantly higher directional distribution utility than circular noise function based method. Ying Zhao 0012, Dong Yuan 0001, Jia Tina Du, Jinjun Chen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Discriminative Metric Learning for Partial Label LearningabstractOne simple strategy to deal with ambiguity in partial label learning (PLL) is to regard all candidate labels equally as the ground-truth label, and then solve the PLL problem using existing multiclass classification algorithms. However, due to the noisy false-positive labels in the candidate set, these approaches are readily mislead and do not generalize well in testing. Consequently, the method of identifying the ground-truth label straight from the candidate label set has grown popular and effective. When the labeling information in PLL is ambiguous, we ought to take advantage of the data's underlying structure, such as label and feature interdependencies, to conduct disambiguation. Furthermore, while metric learning is an excellent method for supervised learning classification that takes feature and label interdependencies into account, it cannot be used to solve the weekly supervised learning PLL problem directly due to the ambiguity of labeling information in the candidate label set. In this article, we propose an effective PLL paradigm called discriminative metric learning for partial label learning (DML-PLL), which aims to learn a Mahanalobis distance metric discriminatively while identifying the ground-truth label iteratively for PLL. We also design an efficient algorithm to alternatively optimize the metric parameter and the latent ground-truth label in an iterative way. Besides, we prove the convergence of the designed algorithms by two proposed lemmas. We additionally study the computational complexity of the proposed DML-PLL in terms of training and testing time for each iteration. Extensive experiments on both controlled UCI datasets and real-world PLL datasets from diverse domains demonstrate that the proposed DML-PLL regularly outperforms the compared approaches in terms of prediction accuracy. Xiuwen Gong, Dong Yuan 0001, Wei Bao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Hierarchical Federated Learning With Momentum Acceleration in Multi-Tier NetworksabstractIn this article, we propose Hierarchical Federated Learning with Momentum Acceleration (HierMo), a three-tier worker-edge-cloud federated learning algorithm that applies momentum for training acceleration. Momentum is calculated and aggregated in the three tiers. We provide convergence analysis for HierMo, showing a convergence rate of$\mathcal {O}(\frac{1}{T})$. In the analysis, we develop a new approach to characterize model aggregation, momentum aggregation, and their interactions. Based on this result, we prove that HierMo achieves a tighter convergence upper bound compared with HierFAVG without momentum. We also propose HierOPT, which optimizes the aggregation periods (worker-edge and edge-cloud aggregation periods) to minimize the loss given a limited training time. By conducting the experiment, we verify that HierMo outperforms existing mainstream benchmarks under a wide range of settings. In addition, HierOPT can achieve a near-optimal performance when we test HierMo under different aggregation periods. Zhengjie Yang, Sen Fu, Wei Bao 0001, Dong Yuan 0001, Albert Y. Zomaya |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2023 | A Novel Graph-Based Computation Offloading Strategy for Workflow Applications in Mobile Edge ComputingabstractWith the fast development of mobile edge computing (MEC), there is an increasing demand for running complex applications on the edge. These complex applications can be represented as workflows where task dependencies are explicitly specified. To achieve better Quality of Service (QoS), computation offloading is widely used in the MEC environment. However, many existing computation offloading strategies only focus on independent computation tasks but overlook the task dependencies. Meanwhile, most of these strategies are based on search algorithms which are often time-consuming and hence not suitable for many delay-sensitive complex applications in MEC. Therefore, a highly efficient graph-based strategy was proposed in our recent work but it can only deal with simple workflow applications with linear (namely sequential) structure. For solving these problems, a novel graph-based strategy is proposed for workflow applications in MEC. Specifically, this strategy can deal with complex workflow applications with nonlinear (viz. parallel, selective and iterative) structures. Meanwhile, the offloading decision plan with the lowest energy consumption of the end-device under deadline constraint can be found by using the graph-based partition technique. We have comprehensively evaluated our strategy on FogWorkflowSim platform for complex workflow applications. Extensive numerical results demonstrate that the end device's energy consumption can be effectively reduced by 7.81% and 9.51% compared with PSO and GA by the proposed strategy. Meanwhile, the strategy running time is 1% and 0.2% of PSO and GA, respectively. Xuejun Li 0001, Dong Yuan 0001, Jia Xu 0010, Xiao Liu 0004 |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | Dynamic path learning in decision trees using contextual banditsabstractAbstract We present a novel online decision-making solution, where the optimal path of a given decision tree is dynamically found based on the contextual bandits analysis. At each round, the learner finds a path in the decision tree by making a sequence of decisions following the tree structure and receives an outcome when a terminal node is reached. At each decision node, the environment information is observed to hint on which child node to visit, resulting in a better outcome. The objective is to learn the context-specific optimal decision for each decision node to maximize the accumulated outcome. In this paper, we propose Dynamic Path Identifier (DPI), a learning algorithm where the contextual bandit is applied to every decision node, and the observed outcome is used as the reward of the previous decisions of the same round. The technical difficulty of DPI is the high exploration challenge caused by the width (i.e., the number of paths) of the tree as well as the large context space. We mathematically prove that DPI’s regret per round approached zero as the number of the rounds approaches infinity. We also prove that the regret is not a function of the number of paths in the tree. Numerical evaluations are provided to complement the theoretical analysis. Weiyu Ju, Dong Yuan 0001, Wei Bao 0001, Liming Ge, Bing Bing Zhou |
World Wide Web (WWW) | 2 |
| 2022 | Partial Multi-Label Learning via Large Margin Nearest Neighbour EmbeddingsabstractTo deal with ambiguities in partial multi-label learning (PML), existing popular PML research attempts to perform disambiguation by direct ground-truth label identification. However, these approaches can be easily misled by noisy false-positive labels in the iteration of updating the model parameter and the latent ground-truth label variables. When labeling information is ambiguous, we should depend more on underlying structure of data, such as label and feature correlations, to perform disambiguation for partially labeled data. Moreover, large margin nearest neighbour (LMNN) is a popular strategy that considers data structure in classification. However, due to the ambiguity of labeling information in PML, traditional LMNN cannot be used to solve the PML problem directly. In addition, embedding is an effective technology to decrease the noise information of data. Inspried by LMNN and embedding technology, we propose a novel PML paradigm called Partial Multi-label Learning via Large Margin Nearest Neighbour Embeddings (PML-LMNNE), which aims to conduct disambiguation by projecting labels and features into a lower-dimension embedding space and reorganize the underlying structure by LMNN in the embedding space simultaneously. An efficient algorithm is designed to implement the proposed method and the convergence rate of the algorithm is analyzed. Moreover, we present a theoretical analysis of the generalization error bound for the proposed PML-LMNNE, which shows that the generalization error converges to the sum of two times the Bayes error over the labels when the number of instances goes to infinity. Comprehensive experiments on artificial and real-world datasets demonstrate the superiorities of the proposed PML-LMNNE. Xiuwen Gong, Dong Yuan 0001, Wei Bao 0001 |
AAAI | 2 |
| 2022 | Partial Label Learning via Label Influence FunctionabstractTo deal with ambiguities in partial label learning (PLL), state-of-the-art strategies implement disambiguations by identifying the ground-truth label directly from the candidate label set. However, these approaches usually take the label that incurs a minimal loss as the ground-truth label or use the weight to represent which label has a high likelihood to be the ground-truth label. Little work has been done to investigate from the perspective of how a candidate label changing a predictive model. In this paper, inspired by influence function, we develop a novel PLL framework called Partial Label Learning via Label Influence Function (PLL-IF). Moreover, we implement the framework with two specific representative models, an SVM model and a neural network model, which are called PLL-IF+SVM and PLL-IF+NN method respectively. Extensive experiments conducted on various datasets demonstrate the superiorities of the proposed methods in terms of prediction accuracy, which in turn validates the effectiveness of the proposed PLL-IF framework. Xiuwen Gong, Dong Yuan 0001, Wei Bao 0001 |
ICML | 2 |
| 2022 | Semi-Online Multi-Machine with Restart Scheduling for Integrated Edge and Cloud Computing SystemsabstractWe study the multi-machine task scheduling problem in an integrated serverless edge and cloud computing system, where tasks can be scheduled locally on edge processors or offloaded to cloud servers, with the objective of minimizing the makespan, i.e., the total time to finish all tasks. The system is semi-online, where the edge processing delays of the tasks are known as priori, but the cloud processing delays remain unknown due to the uncertainty introduced by uploading and loading delay (loading the software environment). The problem is NP-hard in nature, and therefore we resort to approximation schemes and propose a novel algorithm named multi-machine with restart scheduling (MRS). MRS utilizes task restart, where a task that is cancelled will be restarted later when its processing time exceeds the threshold, and the threshold can be adaptively adjusted. We derive an competitive ratio for MRS so that its worst-case gap from the optimal solution is bounded. We also implement the MRS scheduler in a real-world system, which schedules a diverse set of Deep Neural Network (DNN) inference tasks. It shows that MRS achieves significant reduction in makespan compared to existing benchmark schemes. Liming Ge, Wei Bao 0001, Dong Yuan 0001, Nguyen Hoang Tran, Bing Bing Zhou, Albert Y. Zomaya |
ICPP | 4 |
| 2022 | Competitive Analysis for Multi-Commodity Ski-Rental ProblemabstractWe investigate an extended version of the classical ski-rental problem with multiple commodities. A customer uses a set of commodities altogether, and he/she needs to choose payment options to cover the usage of each commodity without the knowledge of the future. The payment options of each commodity include (1) renting: to pay for an on-demand usage and (2) buying: to pay for the lifetime usage. It is a novel extension of the classical ski-rental problem which deals with only one commodity. To address this problem, we propose a new online algorithm called the Multi-Object Break-Even (MOBE) algorithm and conduct competitive analysis. We show that the tight lower and upper bounds of MOBE algorithm's competitive ratio are e/e-1 and 2 respectively against adaptive adversary under arbitrary renting and buying prices. We further prove that MOBE algorithm is an optimal online algorithm if commodities have the same rent-to-buy ratio. Numerical results verify our theoretical conclusion and demonstrate the advantages of MOBE in a real-world scenario. Binghan Wu, Wei Bao 0001, Dong Yuan 0001 |
IJCAI | 3 |
| 2022 | Orchestrating Heterogeneous Cyber-range Event Chains With Serverless-container WorkflowabstractCyber ranges need to run versatile network applications to increase the fidelity of the tests. With the growing complexity of cyberspace events that involve tens to hundreds of diverse applications and flexible execution orders of applications, it is increasingly challenging to orchestrate large-scale, complicated chains of heterogeneous Internet applications. State-of-the-art orchestration techniques do not scale out well due to the inefficient representation model and scheduling of network-centric and correlated Internet application activities. We present a serverless-container workflow orchestration scheme called Wukong. First, we overcome the heterogeneity of events with a workflow model that encodes event chains with compositional DAGs and unified serverless-container event triggers. Second, Wukong scales the scheduling of serverless-container workflows by automatically decomposing DAGs and push-pull coordinated event executions over distributed serverless-container runtime agents. Our evaluation on a real-world cyber range shows that Wukong is expressive, scalable and efficient for automatically emulating diverse event chains, in that the compositional modeling reduces the storage footprint over 57 to 58 times compared to baseline models, the response delay of Wukong is 1.52 to 2.74 times shorter than state-of-the-art orchestration engines, and the scheduling delay is 1.14 to 2.16 times smaller than those of the baseline approach. Yongquan Fu, Weihong Han, Dong Yuan 0001 |
MASCOTS | 3 |
| 2022 | Edge-assisted deep video denoising and super-resolution for real-time surveillance at nightabstractVideo surveillance cameras have been extensively deployed over the last few years. In case of incidents such as natural disaster rescue, it provides vital guidance in real-time. However, due to the limited camera hardware and network bandwidth, noise are observed especially at night and the resolution is low. To tackle these two issues, we design and implement EADV, an Edge-Assisted Deep Video denoising and super-resolution system for real-time surveillance at night. We demonstrate the video quality enhancement using a camera, a displayer, and an edge server. The low-quality video captured by the camera is enhanced by the server and shown on the displayer. The enhanced real-time video is smooth and the performance uplift is observable. Liming Ge, Wei Bao 0001, Dong Yuan 0001, Bing Bing Zhou |
MobiCom | 3 |
| 2022 | A Novel Graph-based Computation Offloading Strategy for Workflow Applications in Mobile Edge Computingabstract[J1C2 Presentation Abstract at IEEE SERVICES 2022 for IEEE Transactions on Services Computing DOI 10.1109/TSC.2022.3180067] Xuejun Li 0001, Dong Yuan 0001, Jia Xu 0010, Xiao Liu 0004 |
SERVICES | 3 |
| 2022 | SPrivAD: A secure and privacy-preserving mutually dependent authentication and data access scheme for smart communities
Abubakar Sadiq Sani, Elisa Bertino, Dong Yuan 0001, Ke Meng 0001, Zhao Yang Dong |
Comput. Secur. | 3 |
| 2022 | Generalized Large Margin $k$NN for Partial Label LearningabstractTo deal with noises in partial label learning (PLL), existing approaches try to perform disambiguation either by identifying the ground-truth label or by averaging the candidate labels. However, these methods can be easily misled by the false-positive noisy labels in the candidate set, and fail to generalize well in testing. When labeling information is ambiguous, learning paradigms should depend more on underlying data structure. Large margin nearest neighbour (LMNN) is a popular strategy to consider instance and class correlations in supervised learning, but can not be directly used in weakly-supervised PLL due to the ambiguity of labeling information. In this paper, we first define similarly and differently labeled pairs as well as the similarity weight to evaluate the similarties between any two instances. We then propose a novel PLL method called Generalized Large Margin$k$NN for Partial Label Learning (GLMNN-PLL), which adapts the framework of LMNN to PLL by modifying the constraint from ‘the same class’ to ‘similarly-labeled’. GLMNN-PLL aims to learn a new metric and perform disambiguation by reorganizing the underlying data structure, that is, making similarly labeled instances closer to each other while making differently labeled instances seperated by a large margin. As two close instances with shared labels do not necessarily belong to the same class, we put a weight on each instance pair. An efficient algorithm is designed to optimize the proposed method and the convergence is analyzed in this paper. Moreover, we present a theoretical analysis of the generalization error bound for GLMNN-PLL. Comprehensive experiments on controlled UCI datasets as well as real-world partial label datasets from various domains demonstrate the superiorities of the proposed method. Xiuwen Gong, Dong Yuan 0001, Wei Bao 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Top- Partial Label MachineabstractTo deal with ambiguities in partial label learning (PLL), the existing PLL methods implement disambiguations, by either identifying the ground-truth label or averaging the candidate labels. However, these methods can be easily misled by the false-positive labels in the candidate label set. We find that these ambiguities often originate from the noise caused by highly correlated or overlapping candidate labels, which leads to the difficulty in identifying the ground-truth label on the first attempt. To give the trained models more tolerance, we first propose the top-k partial loss and convex top-k partial hinge loss. Based on the losses, we present a novel top-k partial label machine (TPLM) for partial label classification. An efficient optimization algorithm is proposed based on accelerated proximal stochastic dual coordinate ascent (Prox-SDCA) and linear programming (LP). Moreover, we present a theoretical analysis of the generalization error for TPLM. Comprehensive experiments on both controlled UCI datasets and real-world partial label datasets demonstrate that the proposed method is superior to the state-of-the-art approaches. Xiuwen Gong, Dong Yuan 0001, Wei Bao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Federated Learning With Nesterov Accelerated GradientabstractFederated learning (FL) is a fast-developing technique that allows multiple workers to train a global model based on a distributed dataset. Conventional FL (FedAvg) employs gradient descent algorithm, which may not be efficient enough. Momentum is able to improve the situation by adding an additional momentum step to accelerate the convergence and has demonstrated its benefits in both centralized and FL environments. It is well-known that Nesterov Accelerated Gradient (NAG) is a more advantageous form of momentum, but it is not clear how to quantify the benefits of NAG in FL so far. This motives us to propose FedNAG, which employs NAG in each worker as well as NAG momentum and model aggregation in the aggregator. We provide a detailed convergence analysis of FedNAG and compare it with FedAvg. Extensive experiments based on real-world datasets and trace-driven simulation are conducted, demonstrating that FedNAG increases the learning accuracy by 3–24% and decreases the total training time by 11–70% compared with the benchmarks under a wide range of settings. Zhengjie Yang, Wei Bao 0001, Dong Yuan 0001, Nguyen Hoang Tran, Albert Y. Zomaya |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Competitive Analysis for Two-Level Ski-Rental ProblemabstractIn this paper, we study a two-level ski-rental problem. There are multiple commodities, each one can be “rented” (paying for on-demand usage) or “purchased” (paying for life-time usage). There is also a combo purchase available so that all commodities can be purchased as a combo. Since the usages of the commodities in future are not known in advance, to minimize the overall cost, we design an online algorithm to decide if we rent a commodity, purchase a commodity, or make a combo purchase. We first propose a deterministic online algorithm. It can achieve 3 competitive ratio, which is optimal and tight. Next, we further propose a randomized online algorithm, leading to a e^σ/(e^σ-1) competitive ratio, where σ is the ratio between the price of a single commodity and the price of combo purchase. Finally, we apply simulation to verify the theoretical competitive ratios and evaluate the actual performance against benchmarks. Binghan Wu, Wei Bao 0001, Dong Yuan 0001 |
AAAI | 3 |
| 2021 | Crypto-Chain: A Relay Resilience Framework for Smart VehiclesabstractRecent findings show that smart vehicles can be exposed to relay attacks resulting from weaknesses in cryptographic operations, such as authentication and key derivation, or poor implementation of these operations. Relay attacks refer to attacks in which authentication is evaded without needing to attack a smart vehicle itself. They are a recurrent problem in practice. In this paper, we formulate the necessary relay resilience settings for strengthening authentication and key derivation and achieving the secure design and efficient implementation of cryptographic protocols based on universal composability, which allows the modular design and analysis of cryptographic protocols. We introduce Crypto-Chain, a relay resilience framework that extends Kusters’s universal composition theorem on a fixed number of protocol systems to prevent bypass of cryptographic operations and avoid implementation errors. Our framework provides an ideal crypto-chain functionality that supports several cryptographic primitives. Furthermore, we provide an ideal functionality for mutual authentication and key derivation in Crypto-Chain by which cryptographic protocols can use cryptographic operations, knowledge about the computation time of the operations, and cryptographic timestamps to ensure relay resilience. As a proof of concept, we first propose and implement a mutual authentication and key derivation protocol (MKD) that confirms the efficiency and relay resilience capabilities of Crypto-Chain and then apply Crypto-Chain to fix two protocols used in smart vehicles, namely Megamos Crypto and Hitag-AES/Pro. Abubakar Sadiq Sani, Dong Yuan 0001, Elisa Bertino, Zhao Yang Dong |
ACSAC | 2 |
| 2021 | Learning Early Exit for Deep Neural Network Inference on Mobile Devices through Multi-Armed BanditsabstractWe present a novel learning framework that utilizes the early exit of Deep Neural Network (DNN), a device-only solution that reduces the latency of inference by sacrificing a reasonable degree of accuracy. Choosing the optimal exit point is challenging as the delay and the accuracy of each exit point are random and cannot be known in advance. The problem is further complicated as the overall duration of the processing is also unknown. To this end, we propose Learning Early Exit (LEE), an online learning scheme based on multi-armed bandits analysis. LEE efficiently learns the optimal exit point for mobile-based DNN inference while simultaneously balancing the exploration-exploitation trade-off. LEE differs from the standard bandit analyses in two ways: the reward of choosing each exit point addresses the confidence-latency trade-off, and the time duration between each action is random (i.e., the latency of each action is random). LEE addresses the aforementioned challenges and it achieves asymptotically optimal performance. We implement a real-world system with a real-time testbed that can be deployed in a driving system. DNN models with multiple exit points are trained and deployed in the testbed so that the performance of LEE and benchmark schemes can be tested and compared. The result denotes that LEE substantially outperforms the benchmark schemes. Weiyu Ju, Wei Bao 0001, Dong Yuan 0001, Liming Ge, Bing Bing Zhou |
CCGRID | 3 |
| 2021 | Dynamic Early Exit Scheduling for Deep Neural Network Inference through Contextual BanditsabstractRecent advances in Deep Neural Networks (DNNs) have dramatically improved the accuracy of DNN inference, but also introduce larger latency. In this paper, we investigate how to utilize early exit, a novel method that allows inference to exit at earlier exit points at the cost of an acceptable amount of accuracy. Scheduling the optimal exit point on a per-instance basis is challenging because the realized performance (i.e., confidence and latency) of each exit point is random and the statistics vary in different scenarios. Moreover, the performance has dependencies among the exit points, further complicating the problem. Therefore, the optimal exit scheduling decision cannot be known in advance but should be learned in an online fashion. To this end, we propose Dynamic Early Exit (DEE), a real-time online learning algorithm based on contextual bandit analysis. DEE observes the performance at each exit point as context and decides whether to exit or keep processing. Unlike standard contextual bandit analyses, the rewards of the decisions in our problem are temporally dependent. Furthermore, the performances of the earlier exit points are inevitably explored more compared to the later ones, which poses an unbalance exploration-exploitation trade-off. DEE addresses the aforementioned challenges, where its regret per inference asymptotically approaches zero. We compare DEE with four benchmark schemes in the real-world experiment. The experiment result shows that DEE can improve the overall performance by up to 98.1% compared to the best benchmark scheme. Weiyu Ju, Wei Bao 0001, Liming Ge, Dong Yuan 0001 |
CIKM | 4 |
| 2021 | Fast Multi-label LearningabstractEmbedding approaches have become one of the most pervasive techniques for multi-label classification. However, the training process of embedding methods usually involves a complex quadratic or semidefinite programming problem, or the model may even involve an NP-hard problem. Thus, such methods are prohibitive on large-scale applications. More importantly, much of the literature has already shown that the binary relevance (BR) method is usually good enough for some applications. Unfortunately, BR runs slowly due to its linear dependence on the size of the input data. The goal of this paper is to provide a simple method, yet with provable guarantees, which can achieve competitive performance without a complex training process. To achieve our goal, we provide a simple stochastic sketch strategy for multi-label classification and present theoretical results from both algorithmic and statistical learning perspectives. Our comprehensive empirical studies corroborate our theoretical findings and demonstrate the superiority of the proposed methods. Xiuwen Gong, Dong Yuan 0001, Wei Bao 0001 |
IJCAI | 2 |
| 2021 | Understanding Partial Multi-Label Learning via Mutual InformationabstractTo deal with ambiguities in partial multilabel learning (PML), state-of-the-art methods perform disambiguation by identifying ground-truth labels directly. However, there is an essential question:“Can the ground-truth labels be identified precisely?". If yes, “How can the ground-truth labels be found?". This paper provides affirmative answers to these questions. Instead of adopting hand-made heuristic strategy, we propose a novel Mutual Information Label Identification for Partial Multilabel Learning (MILI-PML), which is derived from a clear probabilistic formulation and could be easily interpreted theoretically from the mutual information perspective, as well as naturally incorporates the feature/label relevancy considerations. Extensive experiments on synthetic and real-world datasets clearly demonstrate the superiorities of the proposed MILI-PML. Xiuwen Gong, Dong Yuan 0001, Wei Bao 0001 |
NeurIPS | 2 |
| 2021 | A fast and scalable authentication scheme in IOT for smart living
Jianhua Li 0002, Jiong Jin, Lingjuan Lyu, Dong Yuan 0001, Longxiang Gao, Chao Shen 0001 |
Future Gener. Comput. Syst. | 4 |
| 2021 | A Universally Composable Key Exchange Protocol for Advanced Metering Infrastructure in the Energy InternetabstractThe increasing adoption of multiway communications in the advanced metering infrastructure (AMI) of the energy Internet, which is known as the Internet-based smart grid, raises a new question about the security of customers' sensitive data and how the data can be protected from growing cyber attacks such as side-channel and false data injection attacks. The dynamic nature of remote connect/disconnect of components in the AMI also brings new types of security threats. To achieve secure multiway communications and remote connect/disconnect of components, the AMI requires a key exchange protocol (KEP) that meets a number of its security requirements such as confidentiality, integrity, availability, identification, authentication, and access control. In this context, in this article we present a KEP that uses an ideal crypto functionality and an ideal AMI key exchange functionality based on universal composability, which allows modular design and analysis of cryptographic protocols. The former functionality enables AMI components or users to perform authenticated cryptographic operations, while the later functionality enables the users to meet the AMI security requirements before generating a shared secret session key, which can be used in an ideal manner. We carry out experiments to validate the performance of our protocol, and the results show that our protocol offers better performance benefits compared to the existing related protocols and is suitable for the Energy Internet. We further demonstrate the usefulness of our ideal functionalities as a security reinforcement for a widely used KEP, namely the Elliptic Curve Diffie-Hellman. Abubakar Sadiq Sani, Dong Yuan 0001, Wei Bao 0001, Zhao Yang Dong |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | eDeepSave: Saving DNN Inference using Early Exit During Handovers in Mobile Edge EnvironmentabstractRecent advances in deep neural networks (DNNs) have substantially improved the accuracy of intelligent applications. One effective scheme known as DNN partition further improves the speed of the inference by partitioning the DNN to a mobile device and its connected edge server to jointly process the inference. However, one of the challenges is how to maintain the service during handovers to avoid interruptions. Inspired by the recently developed early exit technique, where the DNN inference can be accelerated by leaving at an earlier exit point, we propose eDeepSave, a promising solution to save a large portion of video frames that cannot be handled during handovers. eDeepSave comprises three subschemes: (1) save the partially completed frames that are affected when the handover begins. (2) determine which frames we should save during a handover to maximize the number of saved frames. (3) repartition the last arriving frame before the end of the handover with a provable performance bound so that the frames after the handover can be processed without experiencing congestion. We build up a real-world prototype for the field experiments and extensive simulations, showing that eDeepSave can save up to 100% of the affected frames during handover. Weiyu Ju, Dong Yuan 0001, Wei Bao 0001, Liming Ge, Bing Bing Zhou |
ACM Trans. Sens. Networks | 2 |
| 2020 | Online Metric Learning for Multi-Label ClassificationabstractExisting research into online multi-label classification, such as online sequential multi-label extreme learning machine (OSML-ELM) and stochastic gradient descent (SGD), has achieved promising performance. However, these works lack an analysis of loss function and do not consider label dependency. Accordingly, to fill the current research gap, we propose a novel online metric learning paradigm for multi-label classification. More specifically, we first project instances and labels into a lower dimension for comparison, then leverage the large margin principle to learn a metric with an efficient optimization algorithm. Moreover, we provide theoretical analysis on the upper bound of the cumulative loss for our method. Comprehensive experiments on a number of benchmark multi-label datasets validate our theoretical approach and illustrate that our proposed online metric learning (OML) algorithm outperforms state-of-the-art methods. Xiuwen Gong, Dong Yuan 0001, Wei Bao 0001 |
AAAI | 2 |
| 2020 | Accelerating on-device DNN inference during service outage through scheduling early exit
Wei Bao 0001, Dong Yuan 0001, Liming Ge, Nguyen Hoang Tran, Albert Y. Zomaya |
Comput. Commun. | 3 |
| 2020 | Cost effective dynamic data placement for efficient access of social networks
Hourieh Khalajzadeh, Dong Yuan 0001, Bing Bing Zhou, John C. Grundy, Yun Yang 0001 |
J. Parallel Distributed Comput. | 2 |
| 2020 | VR-Rides: An object-oriented application framework for immersive virtual reality exergamesabstractSUMMARY Exercise can improve health and well‐being. With this in mind, immersive virtual reality (VR) games are being developed to promote physical activity, and are generally evaluated through user studies. However, building such applications is time consuming and expensive. This paper introduces VR‐Rides, an object‐oriented application framework focused on the development of experiment‐oriented VR exergames. Following the modular programming pattern, this framework facilitates the integration of different hardware (such as VR devices, sensors, and physical activity devices) within immersive VR experiences that overlay game narratives on Google Street View panoramas. Combining software engineering and interaction patterns, modules of VR‐Rides can be easily added and managed in the Unity game engine. We evaluate the code efficiency and development effort across our VR exergames developed using VR‐Rides. The reliability, maintainability, and usability of our framework are also demonstrated via code metrics analysis and user studies. The results show that investing in a systematic approach to reusing code and design can be a worthwhile effort for researchers beyond software engineering. Yifan Wang 0031, Kiran Ijaz, Dong Yuan 0001, Rafael A. Calvo |
Softw. Pract. Exp. | 3 |
| 2020 | Prune and Plant: Efficient Placement and Parallelism of Virtual Network FunctionsabstractNetwork function virtualization (NFV) is a promising solution to realize a variety of network services. By definition, virtual network functions (VNFs) are chained together to realize different services. However, chaining is not an ideal solution as service latency grows linearly with respect to the length of the chain. Motivated by the fact that many VNFs can be parallelized, we investigate parallelism of VNFs for acceleration. The dependency of the VNFs is characterized by a directed acyclic graph (DAG). We aim to deploy the VNFs in the right place and process them in parallel without violating the DAG, to minimize the overall delay. However, directly solving the delay minimization problem is NP-hard, and it may also introduce a large number of duplicated packets to burden the system. To deal with these issues, we propose the Prune and Plant (P&P) scheme with polynomial computational complexity, to reduce the overall delay while limiting the number of duplicated packets. P&P comprises two stages: in the Prune stage, we prune the original DAG into a series-parallel graph (SP-graph), which eliminates NP-hardness while maintaining parallelism of VNFs. In the Plant stage, we find the optimal placement for the VNFs with respect to the SP-graph. By both simulation and prototyping, we demonstrate that P&P significantly outperforms benchmark schemes. Wei Bao 0001, Dong Yuan 0001, Bing Bing Zhou, Albert Y. Zomaya |
IEEE Trans. Computers | 2 |
| 2019 | Xyreum: A High-Performance and Scalable Blockchain for IIoT Security and PrivacyabstractAs cyber attacks to Industrial Internet of Things (IIoT) remain a major challenge, blockchain has emerged as a promising technology for IIoT security due to its decentralization and immutability characteristics. Existing blockchain designs, however, introduce high computational complexity and latency challenges which are unsuitable for IIoT. This paper proposes Xyreum, a new high-performance and scalable blockchain for enhanced IIoT security and privacy. Xyreum uses a Time-based Zero-Knowledge Proof of Knowledge (T-ZKPK) with authenticated encryption to perform Mutual Multi-Factor Authentication (MMFA). T-ZKPK properties are also used to support Key Establishment (KE) for securing transactions. Our approach for reaching consensus, which is a blockchain group decision-making process, is based on lightweight cryptographic algorithms. We evaluate our scheme with respect to security, privacy, and performance, and the results show that, compared with existing relevant blockchain solutions, our scheme is secure, privacy-preserving, and achieves a significant decrease in computation complexity and latency performance with high scalability. Furthermore, we explain how to use our scheme to strengthen the security of the REMME protocol, a blockchain-based security protocol deployed in several application domains. Abubakar Sadiq Sani, Dong Yuan 0001, Wei Bao 0001, Phee Lep Yeoh, Zhao Yang Dong, Branka Vucetic, Elisa Bertino |
ICDCS | 2 |
| 2019 | SEE: Scheduling Early Exit for Mobile DNN Inference during Service OutageabstractIn recent years, the rapid development of edge computing enables us to process a wide variety of intelligent applications at the edge, such as real-time video analytics. However, edge computing could suffer from service outage caused by the fluctuated wireless connection or congested computing resource. During the service outage, the only choice is to process the deep neural network (DNN) inference at the local mobile devices. The obstacle is that due to the limited resource, it may not be possible to complete inference tasks on time. Inspired by the recently developedearly exit of DNNs, where we can exit DNN at earlier layers to shorten the inference delay by sacrificing an acceptable level of accuracy, we propose to adopt such mechanism to process inference tasks during the service outage. The challenge is how to obtain the optimal schedule with diverse early exit choices. To this end, we formulate an optimal scheduling problem with the objective to maximize a general overall utility. However, the problem is in the form of integer programming, which cannot be solved by a standard approach. We therefore prove the Ordered Scheduling structure, indicating that a frame arrived earlier must be scheduled earlier. Such structure greatly decreases the searching space for an optimal solution. Then, we propose the Scheduling Early Exit (SEE) algorithm based on dynamic programming, to solve the problem optimally with polynomial computational complexity. Finally, we conduct trace-driven simulations and compare SEE with two benchmarks. The result shows that SEE can outperform the benchmarks by 50.9%. Wei Bao 0001, Dong Yuan 0001, Liming Ge, Nguyen Hoang Tran, Albert Y. Zomaya |
MSWiM | 3 |
| 2019 | New Parallel Algorithms for All Pairwise Computation on Large HPC ClustersabstractAll pairwise computation is defined as performing computation between every pair of the elements in a given dataset. It is often a necessary first step in a number of bioinformatics applications. Many of such applications require multiple terabytes of main memory and take multiple peta floating point operations to complete the computation. Therefore, large HPC clusters are needed to tackle these large-scale computational problems. Conventionally designed parallel algorithms using data partitioning may have a scalability issue, i.e., for a given problem of fixed size the efficiency may decrease if the number of compute nodes is increased (Amdahl's law). In this paper we introduce a new method for parallel algorithm design. Using this method we first design an efficient one-dimensional (1D) ring algorithm and then a two-dimensional (2D) algorithm based on the 1D ring for all pairwise computation. When increasing the compute nodes, instead of reducing the block size, we make multiple copies of the original data blocks in the 1D ring and distribute them across the added compute nodes in the other dimension. By properly organizing the compute nodes the communication overhead can be reduced to a minimum in this two-dimensional setting. Experiments on a Cray XC40 HPC supercomputer show that our new algorithms are very efficient and scalable for large-scale all pairwise computation on large HPC clusters. Wei Bao 0001, Pengyi Yang, Dong Yuan 0001, Bing Bing Zhou |
PDCAT | 5 |
| 2019 | Cyber security framework for Internet of Things-based Energy Internet
Abubakar Sadiq Sani, Dong Yuan 0001, Jiong Jin, Longxiang Gao, Shui Yu 0001, Zhao Yang Dong |
Future Gener. Comput. Syst. | 2 |
| 2019 | A novel directional and non-local-convergent particle swarm optimization based workflow scheduling in cloud-edge environment
Ying Xie 0002, Yuanwei Zhu, Yeguo Wang, Yongliang Cheng, Rongbin Xu, Abubakar Sadiq Sani, Dong Yuan 0001, Yun Yang 0001 |
Future Gener. Comput. Syst. | 7 |
| 2019 | A highly efficient algorithm towards optimal data storage and regeneration cost in multiple clouds
Dong Yuan 0001, Li-Zhen Cui 0001, Bing Bing Zhou |
Future Gener. Comput. Syst. | 2 |
| 2019 | Universally Composable Key Bootstrapping and Secure Communication Protocols for the Energy InternetabstractThe Energy Internet is an advanced smart grid solution to increase energy efficiency by jointly operating multiple energy resources via the Internet. However, such an increasing integration of energy resources requires secure and efficient communication in the Energy Internet. To address such a requirement, we propose a new secure key bootstrapping protocol to support the integration and operation of energy resources. By using a universal composability model that provides a strong security notion for designing and analyzing cryptographic protocols, we define an ideal functionality that supports several cryptographic primitives used in this paper. Furthermore, we provide an ideal functionality for key bootstrapping and secure communication, which allows exchanged session keys to be used for secure communication in an ideal manner. We propose the first secure key bootstrapping protocol that enables a user to verify the identities of other users before key bootstrapping. We also present a secure communication protocol for unicast and multicast communications. The ideal functionalities help in the design and analysis of the proposed protocols. We perform some experiments to validate the performance of our protocols, and the results show that our protocols are superior to the existing related protocols and are suitable for the Energy Internet. As a proof of concept, we apply our functionalities to a practical key bootstrapping protocol, namely generic bootstrapping architecture. Abubakar Sadiq Sani, Dong Yuan 0001, Wei Bao 0001, Zhao Yang Dong, Branka Vucetic, Elisa Bertino |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2018 | A Lightweight Security and Privacy-Enhancing Key Establishment for Internet of Things ApplicationsabstractRecent findings show that many mission critical Internet of Things (IoT) applications are exposed to increasing security risks. The complex and dynamic nature of the IoT and its applications also bring new types of security threats. To achieve end-to-end secure communication, IoT applications need key establishment schemes with integrated security fundamentals such as identification and authentication of IoT components, as well as integrity, confidentiality, availability and authenticity of data, to prevent security attacks from weakening and disrupting the communication. In this context, we present a new lightweight key establishment scheme that comprises a novel Identity-Based Credentials (IBC) mechanism and key establishment protocol. The IBC mechanism enables an IoT component to securely disclose a single identity for security support and privacy enhancement for key establishment. In this paper, we model IoT application attributes to develop our lightweight security and privacy enhancing key establishment scheme. The formal verification and analysis show that, compared to the existing schemes, our proposed scheme is resilient against more types of security attacks, and incurs lower computational and communication costs in IoT applications. Abubakar Sadiq Sani, Dong Yuan 0001, Phee Lep Yeoh, Wei Bao 0001, Shiping Chen 0001, Branka Vucetic |
ICC | 2 |
| 2018 | A Truthful Mechanism for Optimally Purchasing IaaS Instances and Scheduling Parallel Jobs in Service Clouds
Bingbing Zheng, Li Pan 0001, Dong Yuan 0001, Shijun Liu, Yuliang Shi, Lu Wang 0007 |
ICSOC | 3 |
| 2018 | A Truthful Mechanism for Scheduling and Pricing Pleasingly Parallel Jobs in a Service CloudabstractAs more and more users outsource their job executions to service clouds, effective job scheduling and pricing models are needed to solve resource and service competitions between users. Considering the particularity of scheduling and pricing problems in a service cloud whose goal is generally social welfare maximization, current commonly used models, such as fixed-pricing schemes, have obvious shortcomings and thus are unfeasible. Therefore, in this paper, we propose a randomized mechanism to schedule and charge job executions in service clouds. Our proposed mechanism can schedule jobs in a flexible way to achieve approximate social welfare maximization while guaranteeing non-preemption. Flexibility means the number of instances which are allocated to a job can be changed over time. The mechanism is truthful in expectation, computationally efficient and individually rational. The theoretical analysis shows that our mechanism can achieve an expected social welfare approximation ratio α, which can be 2 in some situations. Extensive simulations show that our proposed mechanism can efficiently solve the job scheduling problem in service clouds. Bingbing Zheng, Li Pan 0001, Dong Yuan 0001, Shijun Liu |
ICWS | 3 |
| 2018 | Ins and Outs: Optimal Caching and Re-Caching Policies in Mobile NetworksabstractCaching is essential for data-intensive mobile applications to reduce duplicated data transmission. In this paper, we study the optimal probabilistic caching and re-caching policies in mobile networks as the file popularity may change over time. We propose a Probabilistic File Re-caching (PFR) policy to match the updated popularity. Through PFR, files cached (resp. not cached) are probabilistically opted out (resp. opted in). PFR is with substantial advantages: (1) PFR can handle a huge combinatorial number of all possible situations. (2) The expected number of replaced files is minimized. (3) The computational complexity of PFR is low. Second, we study a utility optimization problem in the mobile network, in order to optimally decide the probability that each file is cached and whether a file should be downloaded from a peer device or directly from the server. Even though the optimization problem is non-convex programming in nature, we devise a computationally efficient Optimal Probabilistic Caching and Requesting (OPCR) policy, through decoupling the decision variables, to derive a globally optimal solution. Finally, we develop a real-world prototype and conduct trace-driven simulations to validate and evaluate our proposed PFR and OPCR policies. Wei Bao 0001, Dong Yuan 0001, Keqi Shi, Weiyu Ju, Albert Y. Zomaya |
MobiHoc | 2 |
| 2018 | sFog: Seamless Fog Computing Environment for Mobile IoT ApplicationsabstractFog computing is a promising solution to provide low-latency and ubiquitously available computation offloading services to widely distributed Internet of Things (IoT) devices with limited computing capabilities. One obstacle, however, is how to seamlessly hand over mobile IoT devices among different fog nodes to avoid service interruption. In this paper, we propose seamless fog (sFog), a new framework supporting efficient congestion control and seamless handover schemes. Intrinsically, sFog improves system performance during handovers (achieved by the handover scheme), and guarantees the performance does not degrade when handovers do not occur (achieved by the congestion control scheme). Through the congestion control scheme, jobs are efficiently offloaded without causing unnecessary system idling; through the handover scheme, jobs are pre-migrated to the target fog node when a handover is about to occur, in order to reduce migration delay. In order to evaluate the performance of sFog, we propose a theoretical framework and establish a real-world prototype. Both the theoretical and experimental results show that sFog achieves substantial delay reductions compared with traditional benchmark handover schemes. Wei Bao 0001, Dong Yuan 0001, Zhengjie Yang, Bing Bing Zhou, Stewart Adams, Albert Y. Zomaya |
MSWiM | 2 |
| 2018 | DCDedupe: Selective Deduplication and Delta Compression with Effective Routing for Distributed Storage
Binqi Zhang, Chen Wang 0008, Bing Bing Zhou, Dong Yuan 0001, Albert Y. Zomaya |
J. Grid Comput. | 4 |
| 2018 | Virtual Fog: A Virtualization Enabled Fog Computing Framework for Internet of ThingsabstractThe prosperity of Internet of Things (IoT) and the success of rich Cloud services have expedited the emergence of a new computing paradigm called Fog computing, which promotes the processing of data at the proximity of their sources. Complementary to the Cloud, Fog promises to offer many appealing features, such as low latency, low cost, high multitenancy, high scalability, and to consolidate the IoT ecosystem. Although the Fog concept has been widely adopted in many areas, a comprehensive realization has yet been adequately researched. To address all these issues, in this paper, object virtualization is investigated to overcome obstacles resulting from resource constraints on sensory-level nodes while service virtualization is explored to easily create tailored applications for end users. Moreover, network function virtualization is studied to perform the flexibility of network service provisioning. Grounded on object virtualization, network function virtualization and service virtualization, a layered framework that encompasses smart objects, Fog and Cloud is presented to illustrate the realization of virtual Fog along IoT continuum. This proposed virtual Fog framework is applied to a smart living case for verification, then quantitative analysis is conducted to demonstrate the low latency, low operating expense, high multitenancy and scalability, followed by an experimental evaluation to further confirm that delay and jitter can be decreased through virtualization. Jianhua Li 0002, Jiong Jin, Dong Yuan 0001, Hongke Zhang |
IEEE Internet Things J. | 3 |
| 2018 | An Algorithm for Finding the Minimum Cost of Storing and Regenerating Datasets in Multiple CloudsabstractThe proliferation of cloud computing allows users to flexibly store, re-compute or transfer large generated datasets with multiple cloud service providers. However, due to the pay-as-you-go model, the total cost of using cloud services depends on the consumption of storage, computation and bandwidth resources which are three key factors for the cost of IaaS-based cloud resources. In order to reduce the total cost for data, given cloud service providers with different pricing models on their resources, users can flexibly choose a cloud service to store a generated dataset, or delete it and choose a cloud service to regenerate it whenever reused. However, finding the minimum cost is a complicated yet unsolved problem. In this paper, we propose a novel algorithm that can calculate the minimum cost for storing and regenerating datasets in clouds, i.e., whether datasets should be stored or deleted, and furthermore where to store or to regenerate whenever they are reused. This minimum cost also achieves the best trade-off among computation, storage and bandwidth costs in multiple clouds. Comprehensive analysis and rigid theorems guarantee the theoretical soundness of the paper, and general (random) simulations conducted with popular cloud service providers' pricing models demonstrate the excellent performance of our approach. Dong Yuan 0001, Li-Zhen Cui 0001, Wenhao Li 0006, Xiao Liu 0004, Yun Yang 0001 |
IEEE Trans. Cloud Comput. | 1 |
| 2018 | A Genetic Algorithm Based Data Replica Placement Strategy for Scientific Applications in CloudsabstractCloud computing is a promising distributed computing platform for big data applications, e.g., scientific applications, since excessive resources can be obtained from cloud services for processing and storing both existing and generated application datasets. However, when tasks process big data stored in distributed data centers, the inevitable data movements will cause huge bandwidth cost and execution delay. In this paper, we construct a tripartite graph based model to formulate the data replica placement problem and propose a genetic algorithm based data replica placement strategy for scientific applications to reduce data transmissions in cloud. Our approach can reduce 1) the size of moved data, 2) the time of data movement and 3) the number of movements. We conduct experiments to compare the proposed strategy with the random placement strategy used in Hadoop Distributed Files System (HDFS), which demonstrates that our strategy has better performance for scientific applications in clouds. Li-Zhen Cui 0001, Lingxi Yue, Yuliang Shi, Hui Li 0048, Dong Yuan 0001 |
IEEE Trans. Serv. Comput. | 6 |
| 2017 | Crowd-enabled Pareto-Optimal Objects Finding Employing Multi-Pairwise-Comparison QuestionsabstractToday, Pareto-optimal objects finding has been applied in various fields, such as group decision making and opinion collection. Many of the existing solutions to this problem require explicit attributes for objects. However, these attributes cannot be obtained sometimes. To address this issue, we propose an algorithm, which uses preference relations given by crowdsourcing, to find Pareto-optimal objects with shorter latency and lower monetary costs. It employs two multi-pairwise-comparison question models: BEST-form and BETTER-form questions. Multiple BEST (or BETTER) questions can be sent to crowds concurrently. Extensive experimental results show that the number of questions reduces greatly. In addition, the numerical results show that the latency is significantly shortened at a reasonable monetary cost, compared with the existing methods. Chang Liu 0040, Yinan Zhang 0002, Lei Liu 0003, Li-Zhen Cui 0001, Dong Yuan 0001, Chunyan Miao |
CIKM | 5 |
| 2017 | A novel dynamic checkpoint selection strategy for time-constrained massive cloud business workflowsabstractThe utilization of cloud computing resources contributes to a fast growth of e-business and e-government where businesses and government agencies often require processing a large volume of service requests in a constrained period of time. In order to achieve on-time completion of a large volume of time-constrained parallel processes, this paper proposes a novel dynamic checkpoint selection strategy (CSSDM) for monitoring cloud business workflows. CSSDMmonitors the execution time points rather than activities along the business workflows. The paper also provides some parameters that can track the execution process of each virtual machine. This mechanism significantly increases the efficiency in monitoring large volume processes so as to improve the quality of service in both e-business and e-government sectors. Simulation results show that CSSDMhas excellent performance in different scenarios even when certain amounts of noises are added. Rongbin Xu, Yeguo Wang, Ying Xie 0002, Dong Yuan 0001 |
CSCWD | 5 |
| 2017 | An optimized data storage strategy by computational performance and monetary cost with data importance in the cloudabstractThe rapid development of cloud computing service allows scientific research data to be easily moved into cloud. Users pay for computational and storing resources to deal with their data in the cloud therefore how to manage those resources has become an urgent issue to be solved. Large application data need to be handled every day, thus a good storage strategy can save monetary cost as well as improve computational performance. In order to satisfy different demands of users, various storage strategies have been proposed. Aiming at proposing an optimized strategy, the paper constructs data importance and time importance models to evaluate data sets, which makes strategy by choosing appropriate data sets to store based on outcomes. Combining these two models, our strategy performances well both in monetary cost and computational performance. Various experiments show excellent results that the strategy performs better than normal storage strategies in both computational performance and monetary cost. The time complexity of our strategy is O(n2). Kangkang Zhao, Dong Yuan 0001, Ying Xie 0002, Rongbin Xu |
CSCWD | 2 |
| 2017 | A Novel Data Set Importance Based Cost-Effective and Computation-Efficient Storage Strategy in the CloudabstractThe rapid development of cloud computing service allows data and computation intensive applications to be easily moved into cloud. Users pay for computing and storing resources to deal with their data in the cloud, therefore how to manage those resources has become an important issue. In order to satisfy different demands of users, various storage strategies have been proposed. Aiming at presenting a cost-effective and computation-efficient strategy, the paper considers both the monetary cost and response time to decide whether a generated data set should be stored or not based on the proposed data importance model. The main focus of this model is to compare the attributes and relationships between data sets. To better illustrate the effectiveness and efficiency of our strategy, both theoretical analysis and simulations are conducted in this paper. Various experiments show excellent results that our strategy outperforms its counterparts in both cost and computation. Rongbin Xu, Kangkang Zhao, Dong Yuan 0001, Ying Xie 0002, Yun Yang 0001 |
ICWS | 4 |
| 2017 | Selling Reserved Instances through Pay-as-You-Go Model in Cloud ComputingabstractCurrent Infrastructure-as-a-Service (IaaS) clouds offer both on-demand and reservation instance purchasing options. Users can combine these two options dynamically to serve time-varying demands while minimizing their instance acquisition costs. However, when future demands are unknown, it is far from trivial for cloud users to make optimal instance purchasing decisions. To deal with this problem, a carefully designed online algorithm can be employed to guide users in acquiring instances without any prior knowledge of future demands while guaranteeing a competitive ratio. In this paper, we propose an instance reselling model, in which a cloud user can temporarily rent out its idle reserved instances to other users through pay-as-you-go model. We also design online instance acquisition strategies which achieve a better competitive ratio than previous methods. Through extensive simulations based on both synthetic data and real-world traces, we show that our online algorithm under the proposed reselling model can outperform previous models and achieve significant cost savings. Dong Yuan 0001, Li Pan 0001, Shijun Liu, Xiangxu Meng |
ICWS | 2 |
| 2017 | Towards an Analysis of Traffic Shaping and Policing in Fog Networks Using Stochastic Fluid ModelsabstractThis paper gives models and analytic techniques for studying shaping and policing data traffic in fog networks. The traffic in these networks is expected to be highly diverse and bursty, and regulation will be required as an integral part of congestion control. We generalize the Leaky Bucket model to shape and police traffic source for rate-based congestion control in high-speed fog networks. In particular, the Markov modulated fluid sources reflect the bursty characteristics of data traffic. To measure the performance of the model in shaping and policing traffic, we derive four performance metrics. The experimental results show that with proper design the Leaky Bucket model effectively controls a 4-way trade-off between throughput, loss probability, delay and burstiness of data traffic. Numerical results also reveal that the model performance is sensitive to certain traffic source characteristics. Jiaojiao Jiang 0001, Longxiang Gao, Jiong Jin, Tom H. Luan, Shui Yu 0001, Dong Yuan 0001, Yong Xiang 0001, Dongfeng Yuan |
MobiQuitous | 6 |
| 2017 | Cost-Effective Processing in Fog-Integrated Internet of Things EcosystemsabstractThe emerging Internet of Things (IoT) paradigm creates a growing need to analyze a significant amount of data produced by the interconnected IoT devices. Since IoT devices have limited computation capabilities, Fog Computing is a natural complement, to provide distributed, location-aware, and easy-to-access computation resources. In this work, we address the problem of application processing and data offloading in a Fog-integrated IoT ecosystem. By leveraging the Lyapunov optimization technique, we design an online and distributed system control policy called the Distributed Weighted Backpressure (DWB) policy that asymptotically minimizes the cost of IoT devices. A three-way tradeoff among queue backlogs, communication cost, and computation cost is then investigated. Finally, simulation study has been conducted to validate the correctness and usefulness of the proposed DWB policy. Wei Bao 0001, Wei Li 0058, Flávia Coimbra Delicato, Paulo F. Pires, Dong Yuan 0001, Bing Bing Zhou, Albert Y. Zomaya |
MSWiM | 5 |
| 2017 | Towards secure energy internet communication scheme: An identity-based key bootstrapping protocol supporting unicast and multicastabstractIt is expected that there are a variety of energy resources to be jointly operated by the power system in the future. Through jointly operating all types of energy resources via the Internet, Energy Internet is a promising solution to increase energy efficiency. However, the increasing integration of energy resources inevitably imposes challenges in secure communication in the Energy Internet. Resourceful and reliable communication with high Quality of Security Service (QoSS) are crucial to the success of information exchange in Energy Internet. In this light, we advocate a new identity-based key bootstrapping protocol to support energy resources integration and operation, and satisfy the increasing requirements in QoSS. The proposed scheme enables a component to verify the identities of other components and ensure the authenticity and integrity of messages for key bootstrapping, which supports an identity-based communication paradigm for unicast and multicast communication. A system model is defined, and the identity-based communication paradigm is applied to address communication security concerns in real-time. The security and performance analyses show that the proposed scheme is superior to existing schemes. Abubakar Sadiq Sani, Dong Yuan 0001, Wei Bao 0001, Zhao Yang Dong |
NCA | 2 |
| 2017 | Near-optimal dynamic priority scheduling strategy for instance-intensive business workflows in cloud computingabstractSummary Utilization of cloud computing resources has made a fast growth in e‐business. Business and government agencies often need to handle large volume of service requests, the so‐called instance‐intensive business processes in a constrained period. On‐time completion for instance‐intensive business processes within the constrained time is a very important issue. In the past few years, traditional optimal task scheduling has been well researched and proven to be a nondeterministic polynomial (NP) time–complete problem. So many heuristic and metaheuristic algorithms are put forward to solve the issue with near‐optimal solutions. However, most of them just treat a single workflow instance as a multistep task without considering that steps within a task can be different types of activities. To explain multistep features of business workflows, a typical motivating instance‐intensive business example of security exchange and a multistep scheduling model for business workflows are introduced in this paper. Then our near‐optimal dynamic priority scheduling (DPS) strategy is proposed on the basis of the idea of Min‐Min heuristic algorithm and greedy philosophy. Compared to the first come first served and constrained Min‐Min by makespan and standard deviation, DPS can make a more optimized choice in each round of scheduling towards overall outcome. To show the effectiveness of DPS, theoretical minimum execution time (METtheory) is used as a benchmark for evaluation based on simulation. The results show that the ratios between METtheory and DPS are more than 98.5% by scheduling different orders of magnitude tasks from 1000 to 1 000 000. In particular, the ratio between METtheory and DPS is nearly 99.9% with 1 000 000 tasks, which means that our DPS can get the near‐optimal result when scheduling large number of tasks. Rongbin Xu, Yeguo Wang, Dong Yuan 0001, Ying Xie 0002, Yun Yang 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2017 | An Energy-Efficient Storage Strategy for Cloud Datacenters Based on Variable K-Coverage of a HypergraphabstractDistributed storage systems, e.g., Hadoop Distributed File System (HDFS), have been widely used in datacenters for handling large amounts of data due to their excellent performance in terms of fault tolerance, reliability and scalability. However, these storage systems usually adopt the same replication and storage strategy to guarantee data availability, i.e., creating the same number of replicas for all data sets and randomly storing them across data nodes. Such strategies do not fully consider the difference requirements of data availability on different data sets. More servers than necessary should thus be used to store replicas of rarely-used data, which will lead to increased energy consumption. To address this issue, we propose an energy-efficient storage strategy for cloud datacenters based on a novel hypergraph coverage model. According to users' requirements of data availability in different applications, our proposed algorithm can selectively determine the corresponding minimum hyperedge coverage, which represents the minimum set of data nodes required in the datacenter. Hence, some other data nodes can be turned off for the purpose of energy saving. We have also implemented our proposed algorithm as a dynamic runtime strategy in a HDFS based prototype datacenter for performance evaluation. Experimental results show that the variable hypergraph coverage based strategy can not only reduce energy consumption, but can also improve the network performance in the datacenter. Ting Yang 0002, Haibo Pen, Wei Li 0058, Dong Yuan 0001, Albert Y. Zomaya |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | Improving Cloud-Based Online Social Network Data Placement and ReplicationabstractOnline social networks make it more convenient for people to find and communicate with other people based on shared interests, ideas, association with different groups, etc. Common social networks such as Facebook and Twitter have hundreds of millions or even billions of users scattered all around the world sharing interconnected data. Users demand low latency access to not only their own data but also their friends' data, often very large, e.g. videos, pictures etc. However, social network service providers have a limited monetary capital to store every piece of data everywhere to minimise users' data access latency. Geo-distributed cloud services with virtually unlimited capabilities are suitable for large scale social networks data storage in different geographical locations. Key problems including how to optimally store and replicate these huge datasets and how to distribute the requests to different datacenters are addressed in this paper. A novel genetic algorithm-based approach is used to find a near-optimal number of replicas for every user's data and a near-optimal placement of replicas to minimise monetary cost while satisfying latency requirements for all users. Experiments on a Facebook dataset demonstrate our technique's effectiveness in outperforming other representative placement and replication strategies. Hourieh Khalajzadeh, Dong Yuan 0001, John C. Grundy, Yun Yang 0001 |
CLOUD | 2 |
| 2016 | A cost-optimal service selection approach for collaborative workflow execution in cloudsabstractToday, there has been a strong demand of distributed collaboration in design and manufacturing, due to the acceleration of economic globalization and the popularity of virtual enterprises (VE) model. Because of the characteristics of cloud computing, such as elasticity and on-demand computing, it is promising to deploy and execute collaborative workflows that contain multiple tasks and services such as Computer-Aided Design (CAD) software components on cloud resources for supporting collaboration across enterprises. Specifically, how to cost-effectively select appropriate services to execute workflows within deadlines while without violating multiple constraints becomes an important issue. In this paper, through investigating the practical requirements of collaborative design workflow, we first formulate the issue of the cost-optimal cloud service selection for collaborative workflow executions as a multi-dimensional optimization problem with multiple constraints. Then we propose an effective approach based on genetic algorithms to address this problem for obtaining near-optimal solutions. Based on workload data derived from real-world systems, we conduct experiments which show that our approach outperforms traditional greedy algorithms in finding better solutions and it also provides real-time performance guarantees in real-world cloud computing environments. Li Pan 0001, Dong Yuan 0001, Shijun Liu, Lei Wu 0002, Xiangxu Meng |
CSCWD | 3 |
| 2016 | A Cloud Robotics Framework of Optimal Task Offloading for Smart City ApplicationsabstractCloud robotics is an emerging paradigm that enables autonomous robotic agents to communicate and collaborate with cloud computing infrastructures. It further complements Internet of Things (IoT) to improve the performance of smart city applications. By offloading heavy data- intensive computation to the ubiquitous cloud, quality of service (QoS) guarantee can be ensured. Unlike their mobile counterpart, the robots have unique characteristics of mobility, skill- learning, data collection and decision-making capabilities, which makes offloading decisions significantly complex. This paper proposes a generic cloud robotics framework to realize smart city vision while taking into consideration its various complexities. Specifically, task offloading is formulated as a constrained optimization problem capable of handling Direct Acyclic Graph (DAG) known as task flow. Given the constraints, a genetic algorithm (GA) based scheme is further developed to solve the problem. The performance of the algorithm is verified by evaluating the results via three benchmarks. To the best of our knowledge, this is one of the first attempts of task offloading approach for smart city applications of cloud robotics. Akhlaqur Rahman, Jiong Jin, Antonio L. Cricenti, Ashfaqur Rahman, Dong Yuan 0001 |
GLOBECOM | 5 |
| 2016 | An Optimal and Iterative Pricing Model for Multiclass IaaS Cloud Services
Li Pan 0001, Shijun Liu, Lei Wu 0002, Li-Zhen Cui 0001, Dong Yuan 0001 |
ICSOC | 6 |
| 2016 | Workflow temporal verification for monitoring parallel business processesabstractAbstract Workflow temporal verification is conducted to guarantee on‐time completion, which is one of the most important QoS (Quality of Service) dimensions for business processes running in the cloud. However, as today's business systems often need to handle a large number of concurrent customer requests, conventional response‐time based process monitoring strategies conducted in a one‐by‐one fashion cannot be applied efficiently to a large batch of parallel processes because of significant time overhead. Similar situations may also exist in software companies where multiple software projects are carried out at the same time by software developers. To address such a problem, based on a novel runtime throughput consistency model, this paper proposes a QoS‐aware throughput based checkpoint selection strategy, which can dynamically select a small number of checkpoints along the system timeline to facilitate the temporal verification of throughput constraints and achieve the target on‐time completion rate. Experimental results demonstrate that our strategy can achieve the best efficiency and effectiveness compared with the state‐of‐the‐art as and other representative response‐time based checkpoint selection strategies. Copyright © 2016 John Wiley & Sons, Ltd. Xiao Liu 0004, Dingxian Wang, Dong Yuan 0001, Futian Wang, Yun Yang 0001 |
J. Softw. Evol. Process. | 3 |
| 2016 | Ensuring Cloud Data Reliability with Minimum Replication by Proactive Replica CheckingabstractData reliability and storage costs are two primary concerns for current Cloud storage systems. To ensure data reliability, the widely used multi-replica (typically three) replication strategy in current Clouds incurs a huge extra storage consumption, resulting in a huge storage cost for data-intensive applications in the Cloud in particular. In order to reduce the Cloud storage consumption while meeting the data reliability requirement, in this paper we present a cost-effective data reliability management mechanism named PRCR based on a generalized data reliability model. By using a proactive replica checking approach, while the running overhead for PRCR is negligible, PRCR ensures reliability of the massive Cloud data with the minimum replication, which can also serve as a cost effectiveness benchmark for replication based approaches. Our simulation indicates that, compared with the conventional three-replica strategy, PRCR can reduce from one-third to two-thirds of the Cloud storage space consumption, hence significantly lowering the storage cost in a Cloud. Wenhao Li 0006, Yun Yang 0001, Dong Yuan 0001 |
IEEE Trans. Computers | 3 |
| 2015 | A framework for convergence of cloud services and Internet of thingsabstractToday, Cloud Computing and the Internet of things are two “major forces” that drive the development of new Information Technology (IT) solutions. Many Internet of things (IoT) based large-scale applications rely on a cloud platform for data processing and storage. However, big data generated or collected by large-scale geo-distributed devices needs to be transferred to the cloud, often becoming a bottleneck for the system. In this paper, we propose a framework that integrates popular cloud services with a network of IoT devices. In the framework, novel methods have been designed for reliable and efficient data transportation. This framework provides a convergence of cloud services and devices that will ease the development of IoT based, cloud-enabled applications. We have implemented a prototype of the framework to demonstrate the convergence of popular cloud services and IoT technologies. Dong Yuan 0001, Jiong Jin, John C. Grundy, Yun Yang 0001 |
CSCWD | 1 |
| 2015 | Sliding mode-like congestion control for communication networks with heterogeneous applicationsabstractThis paper develops a fair and efficient congestion control framework using robust sliding mode control. It not only considers the limitations of current optimal congestion control approach, but also guarantees the performance of heterogeneous applications with different Quality of Service (QoS) requirements. By further proposing enhanced sliding mode-like congestion control algorithms, the paper addresses two critical issues raised from previous work [1], namely, the rigorous stability of the system and the sensitivity of design parameter. The thorough treatment makes the framework practically applicable, as well as retains its appealing properties. Moreover, the paper summarizes the applications of sliding mode approach in communication networks and highlights its great potential ahead. Jiong Jin, Dong Yuan 0001, Jinchuan Zheng |
ICC | 2 |
| 2015 | Dynamic On-the-Fly Minimum Cost Benchmarking for Storing Generated Scientific Datasets in the CloudabstractMassive computation power and storage capacity of cloud computing systems enable users to either store large generated scientific datasets in the cloud or delete and then regenerate them whenever reused. Due to the pay-as-you-go model, the more datasets we store, the more storage cost we need to pay, alternatively, we can delete some generated datasets to save the storage cost but more computation cost is incurred for regeneration whenever the datasets are reused. Hence, there should exist a trade-off between computation and storage in the cloud, where different storage strategies lead to different total costs. The minimum cost, which reflects the best trade-off, is an important benchmark for evaluating the cost-effectiveness of different storage strategies. However, the current benchmarking approach is neither efficient nor practical to be applied on the fly at runtime. In this paper, we propose a novel Partitioned Solution Space based approach with efficient algorithms for dynamic yet practical on-the-fly minimum cost benchmarking of storing generated datasets in the cloud. In this approach, we pre-calculate all the possible minimum cost storage strategies and save them in different partitioned solution spaces. The minimum cost storage strategy represents the minimum cost benchmark, and whenever the datasets storage cost changes at runtime in the cloud (e.g. new datasets are generated and/or existing datasets' usage frequencies are changed), our algorithms can efficiently retrieve the current minimum cost storage strategy from the partitioned solution space and update the benchmark. By dynamically keeping the benchmark updated, our approach can be practically utilised on the fly at runtime in the cloud, based on which the minimum cost benchmark can be either proactively reported or instantly responded upon request. Case studies and experimental results based on Amazon cloud show the efficiency, scalability and practicality of our approach. Dong Yuan 0001, Xiao Liu 0004, Yun Yang 0001 |
IEEE Trans. Computers | 1 |
| 2014 | Throughput based temporal verification for monitoring large batch of parallel processesabstractOn-time completion is one of the most important QoS (Quality of Service) dimensions for business processes running in the cloud. While today’s business systems often need to handle thousands of concurrent user requests, process monitoring is basically conducted in a one by one fashion. It is possible to repeat the strategies for monitoring a single process a thousand times to monitor a thousand parallel processes. However, the time overhead will be a thousand-fold increase as well, which brings a big challenge for process monitoring. In this paper, based on a novel runtime throughput consistency model, we propose a QoS-aware throughput based checkpoint selection strategy which can dynamically select a small number of checkpoints along the system timeline to facilitate the temporal verification of throughput constraints and achieve the target on-time completion rate. The experimental results demonstrate that our strategy can achieve the best efficiency and effectiveness compared with the state-of-the-art as well as other representative response-time based checkpoint selection strategies. Xiao Liu 0004, Dingxian Wang, Dong Yuan 0001, Futian Wang, Yun Yang 0001 |
ICSSP | 3 |
| 2014 | A gaussian fields based mining method for semi-automating staff assignment in workflow applicationabstractStaff assignment is a very important task in the research of workflow resource management. Currently, many well-known workflow applications still rely on human resource assigners such as process initiator or process monitor to perform staff assignment task. In this paper, we propose a semi-automatic workflow staff assignment method which can decrease the workload of staff assigner based on a novel semi-supervised machine learning framework. Our method can be applied to learn all kinds of activities that each actor is capable of based on the workflow event log. After we have learned all labeled data, we can suggest a suitable actor to undertake the specified activities when a new process is assigned. With the proposed method, we can get an average prediction accuracy of 97% and 91% on the data sets of two manufacturing enterprise applications respectively. Rongbin Xu, Xiao Liu 0004, Ying Xie 0002, Dong Yuan 0001, Yun Yang 0001 |
ICSSP | 4 |
| 2014 | Do we need to handle every temporal violation in scientific workflow systems?abstractScientific processes are usually time constrained with overall deadlines and local milestones. In scientific workflow systems, due to the dynamic nature of the underlying computing infrastructures such as grid and cloud, execution delays often take place and result in a large number of temporal violations. Since temporal violation handling is expensive in terms of both monetary costs and time overheads, an essential question aroused is “do we need to handle every temporal violation in scientific workflow systems?” The answer would be “true” according to existing works on workflow temporal management which adopt the philosophy similar to the handling of functional exceptions, that is, every temporal violation should be handled whenever it is detected. However, based on our observation, the phenomenon of self-recovery where execution delays can be automatically compensated for by the saved execution time of subsequent workflow activities has been entirely overlooked. Therefore, considering the nonfunctional nature of temporal violations, our answer is “not necessarily true.” To take advantage of self-recovery, this article proposes a novel adaptive temporal violation handling point selection strategy where this phenomenon is effectively utilised to avoid unnecessary temporal violation handling. Based on simulations of both real-world scientific workflows and randomly generated test cases, the experimental results demonstrate that our strategy can significantly reduce the cost on temporal violation handling by over 96% while maintaining extreme low violation rate under normal circumstances. Xiao Liu 0004, Yun Yang 0001, Dong Yuan 0001, Jinjun Chen |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2013 | An Algorithm for Cost-Effectively Storing Scientific Datasets with Multiple Service Providers in the CloudabstractThe proliferation of cloud computing allows scientists to deploy computation and data intensive applications without infrastructure investment, where large generated datasets can be flexibly stored with multiple cloud service providers. Due to the pay-as-you-go model, the total application cost largely depends on the usage of computation, storage and bandwidth resources, and cutting the cost of cloud-based data storage becomes a big concern for deploying scientific applications in the cloud. In this paper, we propose a novel algorithm that can automatically decide whether a generated dataset should be 1) stored in the current cloud, 2) deleted and re-generated whenever reused or 3) transferred to cheaper cloud service for storage. The algorithm finds the trade-off among computation, storage and bandwidth costs in the cloud, which are three key factors for the cost of storing generated application datasets with multiple cloud service providers. Simulations conducted with popular cloud service providers' pricing models show that the proposed algorithm is highly cost-effective to be utilised in the cloud. Dong Yuan 0001, Xiao Liu 0004, Wenhao Li 0006, Dahai Cao, Yun Yang 0001 |
e-Science | 1 |
| 2013 | Selecting checkpoints along the time line: a novel temporal checkpoint selection strategy for monitoring a batch of parallel business processesabstractNowadays, most business processes are running in a parallel, distributed and time-constrained manner. How to guarantee their on-time completion is a challenging issue. In the past few years, temporal checkpoint selection which selects a subset of workflow activities for verification of temporal consistency has been proved to be very successful in monitoring single, complex and large size scientific workflows. An intuitive approach is to apply those strategies to individual business processes. However, in such a case, the total number of checkpoints will be enormous, namely the cost for system monitoring and exception handling could be excessive. To address such an issue, we propose a brand new idea which selects time points along the workflow execution time line as checkpoints to monitor a batch of parallel business processes simultaneously instead of individually. Based on such an idea, a set of new definitions as well as a time-point based checkpoint selection strategy are presented in this paper. Our preliminary results demonstrate that it can achieve an order of magnitude reduction in the number of checkpoints while maintaining satisfactory on-time completion rates compared with the state-of-the-art activity-point based checkpoint selection strategy. Xiao Liu 0004, Yun Yang 0001, Dahai Cao, Dong Yuan 0001 |
ICSE | 4 |
| 2013 | A clustering based coscheduling strategy for efficient scientific workflow execution in cloud computingabstractSUMMARY Due to its advantages of cost‐effectiveness, on‐demand provisioning and easy for sharing, cloud computing has grown in popularity with the research community for deploying scientific applications such as workflows. Although such interests continue growing and scientific workflows are widely deployed in collaborative cloud environments that consist of a number of data centers, there is an urgent need for exploiting strategies which can place application datasets across globally distributed data centers and schedule tasks according to the data layout to reduce both latency and makespan for workflow execution. In this paper, by utilizing dependencies among datasets and tasks, we propose an efficient data and task coscheduling strategy that can place input datasets in a load balance way and meanwhile, group the mostly related datasets and tasks together. Moreover, data staging is used to overlap task execution with data transmission in order to shorten the start time of tasks. We build a simulation environment on Tianhe supercomputer for evaluating the proposed strategy and run simulations by random and realistic workflows. The results demonstrate that the proposed strategy can effectively improve scheduling performance while reducing the total volume of data transfer across data centers. Concurrency and Computation: Practice and Experience, 2013.© 2013 Wiley Periodicals, Inc. Kefeng Deng, Kaijun Ren, Junqiang Song, Dong Yuan 0001, Yang Xiang 0001, Jinjun Chen |
Concurr. Comput. Pract. Exp. | 4 |
| 2013 | A market-oriented hierarchical scheduling strategy in cloud workflow systems
Zhangjun Wu, Xiao Liu 0004, Zhiwei Ni, Dong Yuan 0001, Yun Yang 0001 |
J. Supercomput. | 4 |
| 2013 | A Highly Practical Approach toward Achieving Minimum Data Sets Storage Cost in the CloudabstractMassive computation power and storage capacity of cloud computing systems allow scientists to deploy computation and data intensive applications without infrastructure investment, where large application data sets can be stored in the cloud. Based on the pay-as-you-go model, storage strategies and benchmarking approaches have been developed for cost-effectively storing large volume of generated application data sets in the cloud. However, they are either insufficiently cost-effective for the storage or impractical to be used at runtime. In this paper, toward achieving the minimum cost benchmark, we propose a novel highly cost-effective and practical storage strategy that can automatically decide whether a generated data set should be stored or not at runtime in the cloud. The main focus of this strategy is the local-optimization for the tradeoff between computation and storage, while secondarily also taking users' (optional) preferences on storage into consideration. Both theoretical analysis and simulations conducted on general (random) data sets as well as specific real world applications with Amazon's cost model show that the cost-effectiveness of our strategy is close to or even the same as the minimum cost benchmark, and the efficiency is very high for practical runtime utilization in the cloud. Dong Yuan 0001, Yun Yang 0001, Xiao Liu 0004, Wenhao Li 0006, Li-Zhen Cui 0001, Meng Xu 0012, Jinjun Chen |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2012 | A Cost-Effective Mechanism for Cloud Data Reliability Management Based on Proactive Replica CheckingabstractIn current Cloud computing environments, management of data reliability has become a challenge. For data-intensive scientific applications, storing data in the Cloud with the typical 3-replica replication strategy for managing the data reliability would incur huge storage cost. To address this issue, in this paper we present a novel cost-effective data reliability management mechanism named PRCR, which proactively checks the availability of replicas for maintaining data reliability. Our simulation indicates that, comparing with the typical 3-replica replication strategy, PRCR can reduce the storage space consumption by one-third to two-thirds, hence reduce the storage cost significantly in the Cloud. Wenhao Li 0006, Yun Yang 0001, Jinjun Chen, Dong Yuan 0001 |
CCGRID | 4 |
| 2012 | A data dependency based strategy for intermediate data storage in scientific cloud workflow systemsabstractSUMMARY Many scientific workflows are data intensive where large volumes of intermediate data are generated during their execution. Some valuable intermediate data need to be stored for sharing or reuse. Traditionally, they are selectively stored according to the system storage capacity, determined manually. As doing science in the cloud has become popular nowadays, more intermediate data can be stored in scientific cloud workflows based on a pay‐for‐use model. In this paper, we build an intermediate data dependency graph (IDG) from the data provenance in scientific workflows. With the IDG, deleted intermediate data can be regenerated, and as such we develop a novel intermediate data storage strategy that can reduce the cost of scientific cloud workflow systems by automatically storing appropriate intermediate data sets with one cloud service provider. The strategy has significant research merits, i.e. it achieves a cost‐effective trade‐off of computation cost and storage cost and is not strongly impacted by the forecasting inaccuracy of data sets' usages. Meanwhile, the strategy also takes the users' tolerance of data accessing delay into consideration. We utilize Amazon's cost model and apply the strategy to general random as well as specific astrophysics pulsar searching scientific workflows for evaluation. The results show that our strategy can reduce the overall cost of scientific cloud workflow execution significantly. Copyright © 2010 John Wiley & Sons, Ltd. Dong Yuan 0001, Yun Yang 0001, Xiao Liu 0004, Gaofeng Zhang, Jinjun Chen |
Concurr. Comput. Pract. Exp. | 1 |
| 2012 | A trust-based noise injection strategy for privacy protection in cloudabstractSUMMARY Cloud promises users that they can present and deploy IT services in a pay‐as‐you‐go fashion in an open and virtualized cloud environment while saving huge capital investment in their own IT infrastructure. In this sense, protection of users' privacy is critical and has become one of the most concerned issues as otherwise users may eventually lose the confidence and passion of deploying cloud in practice. Under some special cloud circumstances, some users' privacy, such as plans or habits, could be induced from their service requests by service providers without permissions from users. In this regard, obfuscation strategy can protect this kind of privacy by injecting ‘noise’ service requests to confuse potential ‘immoral’ service providers. However, existing noise obfuscation strategies focus on single noise injection whereas investigation of noise injection architecture has been neglected. Especially, a common service pattern in inter‐clouds environment, the cooperative service process including different service providers, makes the risk of privacy serious and uncontrollable by the spread of users' privacy. To address this, we present a novel trust‐based noise injection strategy for privacy protection in cloud. To support the strategy, we describe our noise injection architecture in cloud which specializes in the relations between various service roles in inter‐clouds based on our trust model. The simulation can demonstrate that our noise injection strategy could significantly improve the effectiveness of privacy protection. Copyright © 2011 John Wiley & Sons, Ltd. Gaofeng Zhang, Yun Yang 0001, Dong Yuan 0001, Jinjun Chen |
Softw. Pract. Exp. | 3 |
| 2011 | A Local-Optimisation Based Strategy for Cost-Effective Datasets Storage of Scientific Applications in the CloudabstractMassive computation power and storage capacity of cloud computing systems allow scientists to deploy computation and data intensive applications without infrastructure investment, where large application datasets can be stored in the cloud. However, due to the pay-as-you-go model, the datasets should be strategically stored in order to reduce the overall application cost. In this paper, by utilising Data Dependency Graph (DDG) from data provenances in scientific applications, deleted datasets can be regenerated, and as such we develop a novel cost-effective datasets storage strategy that can automatically store appropriate datasets in the cloud. This strategy achieves a localised optimal trade-off between computation and storage, meanwhile also taking users' tolerance of data accessing delay into consideration. Simulations conducted on general (random) datasets and a specific astrophysics pulsar searching application with Amazon's cost model show that our strategy can reduce the application cost significantly. Dong Yuan 0001, Yun Yang 0001, Xiao Liu 0004, Jinjun Chen |
IEEE CLOUD | 1 |
| 2011 | A Weighted K-Means Clustering Based Co-scheduling Strategy towards Efficient Execution of Scientific Workflows in Collaborative Cloud EnvironmentsabstractDue to the advantages of cost-effectiveness, on-demand resource provision and easy for sharing, cloud computing has grown in popularity with research community for deploying scientific applications such as workflows. When such interest continues growing and workflows are widely performed in collaborative cloud environments that consist of a number of data centers, there is an urgent need for exploiting strategies which can place the application data across globally distributed data centers and schedule tasks according to the data layout to reduce both the latency and make span for workflow execution. In this paper, by utilising dependencies among datasets and tasks, we propose an efficient data and task co scheduling strategy that can place input datasets in a load balance way and meanwhile group the mostly related datasets and tasks together. We build a simulation environment on Tianhe supercomputer to evaluate the proposed strategy and run simulations by random and realistic workflows. The results demonstrate that the proposed strategy can effectively improve workflows performance while reducing the total volume of data transfer across data centers. Kefeng Deng, Lingmei Kong, Junqiang Song, Kaijun Ren, Dong Yuan 0001 |
DASC | 5 |
| 2011 | A Novel Cost-Effective Dynamic Data Replication Strategy for Reliability in Cloud Data CentresabstractNowadays, large-scale Cloud-based applications have put forward higher demand for storage ability of data centres. Data in the Cloud need to be stored with high efficiency and cost effectiveness while meeting the requirement of reliability. While in current Cloud systems, typical 3-replicas data replication strategies are applied for data reliability, in this paper we propose a novel cost-effective dynamic data replication strategy which facilitates an incremental replication method to reduce the storage cost and meet the data reliability requirement at the same time. This replication strategy works very well especially for data which are only used temporarily and/or have a relatively low reliability requirement. The simulation shows that our replication strategy for reliability can reduce the data storage cost in data centres substantially. Wenhao Li 0006, Yun Yang 0001, Dong Yuan 0001 |
DASC | 3 |
| 2011 | A Generic QoS Framework for Cloud Workflow SystemsabstractDue to the dynamic nature of cloud computing, how to achieve satisfactory QoS (Quality of Service) in cloud workflow systems becomes a challenge. Meanwhile, since QoS requirements have many dimensions, a unified system design for different QoS management components is required to reduce the system complexity and software development cost. Therefore, this paper proposes a generic QoS framework for cloud workflow systems. Covering the major stages of a workflow lifecycle, the framework consists of four components, viz. QoS requirement specification, QoS-aware service selection, QoS consistency monitoring and QoS violation handling. While there are many QoS dimensions, this paper illustrates a concrete performance framework as a case study and briefly touches others. We also demonstrate the system implementation and evaluate the effectiveness of the performance framework in our cloud workflow system. Xiao Liu 0004, Yun Yang 0001, Dong Yuan 0001, Gaofeng Zhang, Wenhao Li 0006, Dahai Cao |
DASC | 3 |
| 2011 | On-demand minimum cost benchmarking for intermediate dataset storage in scientific cloud workflow systems
Dong Yuan 0001, Yun Yang 0001, Xiao Liu 0004, Jinjun Chen |
J. Parallel Distributed Comput. | 1 |
| 2011 | A novel general framework for automatic and cost-effective handling of recoverable temporal violations in scientific workflow systems
Xiao Liu 0004, Zhiwei Ni, Zhangjun Wu, Dong Yuan 0001, Jinjun Chen, Yun Yang 0001 |
J. Syst. Softw. | 4 |
| 2011 | A novel statistical time-series pattern based interval forecasting strategy for activity durations in workflow systems
Xiao Liu 0004, Zhiwei Ni, Dong Yuan 0001, Yuan-Chun Jiang, Zhangjun Wu, Jinjun Chen, Yun Yang 0001 |
J. Syst. Softw. | 3 |
| 2010 | Handling Recoverable Temporal Violations in Scientific Workflow Systems: A Workflow Rescheduling Based StrategyabstractDue to the complex nature of scientific workflow systems, the violations of temporal QoS constraints often take place and may severely affect the usefulness of the execution's results. Therefore, to deliver satisfactory QoS, temporal violations need to be recovered effectively. However, there are two fundamental issues which have so far not been well addressed: how to define recoverable temporal violations, how to design corresponding exception handling strategies. In this paper, we first propose a probability based temporal consistency model to define the temporal violations which are statistically recoverable by light-weight exception handling strategies. Afterwards, a novel Ant Colony Optimisation based two-stage workflow local rescheduling strategy (ACOWR) is proposed to handle detected recoverable temporal violations in an automatic and cost-effective fashion. The simulation experiments conducted in our scientific workflow system demonstrate the excellent performance of our handling strategy in reducing both local and global temporal violation rates. Xiao Liu 0004, Jinjun Chen, Zhangjun Wu, Zhiwei Ni, Dong Yuan 0001, Yun Yang 0001 |
CCGRID | 5 |
| 2010 | An Effective Framework of Light-Weight Handling for Three-Level Fine-Grained Recoverable Temporal Violations in Scientific WorkflowsabstractTemporal violations may often take place and deteriorate the overall QoS of scientific workflows. To handle temporal violations in an automatic and cost-effective fashion, we need to resolve the following issues: 1) how to define fine-grained recoverable temporal violations, 2) which light-weight effective exception handling strategies to be facilitated. This paper proposes an effective exception handling framework. Based on a probability based temporal consistency model, the probability range for recoverable temporal violations is divided into three levels of fine-grained temporal violations. Afterwards, three corresponding light-weight exception handling strategies including TDA (Time Deficit Allocation), ACOWR (Ant Colony Optimisation based two-stage Workflow local Rescheduling) and TDA+ACOWR (the combined strategy of TDA and ACOWR) are presented. The experimental results demonstrate the excellent performance of our framework in reducing both local and global temporal violations. Xiao Liu 0004, Zhiwei Ni, Zhangjun Wu, Dong Yuan 0001, Jinjun Chen, Yun Yang 0001 |
ICPADS | 4 |
| 2010 | A cost-effective strategy for intermediate data storage in scientific cloud workflow systemsabstractMany scientific workflows are data intensive where a large volume of intermediate data is generated during their execution. Some valuable intermediate data need to be stored for sharing or reuse. Traditionally, they are selectively stored according to the system storage capacity, determined manually. As doing science on cloud has become popular nowadays, more intermediate data can be stored in scientific cloud workflows based on a pay-for-use model. In this paper, we build an Intermediate data Dependency Graph (IDG) from the data provenances in scientific workflows. Based on the IDG, we develop a novel intermediate data storage strategy that can reduce the cost of the scientific cloud workflow system by automatically storing the most appropriate intermediate datasets in the cloud storage. We utilise Amazon's cost model and apply the strategy to an astrophysics pulsar searching scientific workflow for evaluation. The results show that our strategy can reduce the overall cost of scientific cloud workflow execution significantly. Dong Yuan 0001, Yun Yang 0001, Xiao Liu 0004, Jinjun Chen |
IPDPS | 1 |
| 2010 | A data placement strategy in scientific cloud workflows
Dong Yuan 0001, Yun Yang 0001, Xiao Liu 0004, Jinjun Chen |
Future Gener. Comput. Syst. | 1 |
| 2008 | An Algorithm in SwinDeW-C for Scheduling Transaction-Intensive Cost-Constrained Cloud WorkflowsabstractThe concept of cloud computing has been wide spreading very recently. Cloud computing has many unique advantages which can be utilised to facilitate (cloud) workflow execution. Transaction-intensive cost-constrained cloud workflows are workflows with a large number of workflow instances (i.e. transaction intensive) bounded by a certain budget for execution (i.e. cost constrained) in a cloud computing environment (i.e. cloud workflows). However, there are not any specific scheduling algorithms so far for transaction-intensive cost-constrained cloud workflows. This paper presents a novel scheduling algorithm which considers the characteristics of cloud computing to accommodate transaction-intensive cost-constrained workflows by compromising execution time and cost with user input enabled on the fly. The simulation performed demonstrates that the algorithm can reduce the mean execution cost while meeting the user-requested deadline. Yun Yang 0001, Jinjun Chen, Xiao Liu 0004, Dong Yuan 0001, Hai Jin 0001 |
eScience | 5 |
| 2007 | A Process-semantic Repository for Automatic Workflow Modeling in Web Service EnvironmentabstractWith the fast development of Web service technologies, we need workflow management system to assure a mass of different Web services to work cooperatively. In this paper, we propose a cooperative process-semantic repository based on probability, through which we could achieve automatic modeling of workflow according to custom's requirements in the Web service environment. Based on transfer probability, the components of the process-semantic repository and its self-adapting mechanism are also described. A prototype of e-travel system was then conceived, and applied in tourism. This e-travel system could automatically generate travel processes for the tourists. The efficiency analysis of the process-semantic repository is also done through some experiment data. Dong Yuan 0001 |
CSCWD | 1 |