VLDB 2026 Research / reviewers in the wild / expert
Rui Han 0001
dblp:87/7513-1
· DBLP profile ↗
50ranked-venue papers
21as first author
29since 2021 · last 2026
0000-0001-6894-1921ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 11 first-author · 7 since 2021Computer networks · 10 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Differential Privacy Noise Injection for Decentralized Federated Learning of Visual Recognition Tasks
Junyan Ouyang, Siqi Du, Rui Han 0001, Chi Harold Liu, Jianxin Zhao 0001, Xiaoning Wu, Lydia Y. Chen |
Int. J. Comput. Vis. | 3 |
| 2026 | EdgeTail: Mitigating Long-Tail Visual Problems in Continual Learning at EdgeabstractLarge vision and language models deployed at edge encounter continuously evolving input distributions, including not only new tasks but also highly unbalanced long-tail classes. For example, smart-surveillance cameras frequently capture common objects such as pedestrians and cars, while only occasionally observing tail classes like horseback riders or stroller pushers. However, most existing long-tail mitigation techniques are designed for fixed pretraining data. This often leads to poor accuracy of tail classes on evolving data and incurs high computational costs on edge devices. In this article, we propose EdgeTail, a lightweight long-tail mitigation method for edge-side continual learning. EdgeTail’s key design features are: (i) optimal long-tailed mitigation solution search, which adaptively selects the best long-tail learning method for the current distribution/task; and (ii) graph attention classifier and multi-branch adapter, which improves the quality and stability of tail class representations with small overheads. We implement EdgeTail in PyTorch and extensively evaluate it against state-of-the-art methods. The results show that EdgeTail improves the average accuracy by 36.09% under fixed training windows and by 31.12% under different training window sizes. Yuzhong Ouyang, Xiaoning Wu, Rui Han 0001, Anjie Luo, Chi Harold Liu, Jing Chen 0030, Ying Guo 0030 |
ACM Trans. Internet Things | 4 |
| 2026 | ConvertNet: Training-Time Model Scaling for In-Vehicle Continuous LearningabstractVehicles are transforming from transportation means to mobile intelligent terminals, powered by artificial intelligence (AI) technologies such as large language (LLM) and vision large models (VLM). To maintain high learning accuracy in all situations, in-vehicle AI models encounter the acute need to continuously learn and dynamically retrain AI models on the fly. The volatile resource demands of such stochastic retraining jobs and inference jobs of higher priority however can not be accommodated simultaneously by existing training systems, which rely on pre-generated compressed models of fixed architecture. In this paper, we proposeConvertNet, a novel continuous learning system that effectively scales up/down a compressed model through training-time neuron memorizer (TNM) - a max heap of neurons in tree structure. Based on TNM,ConvertNetestimates the model's resource-accuracy trade-off per neuron, thus efficiently utilizing the limited available resources to scale up and (re)train the model to improve its accuracy. Evaluated on six representative in-vehicle scenarios, comparative experiments against eleven state-of-the-art techniques show thatConvertNetachieves as much as 22.33% improvement in learning accuracy, and reduces energy consumption by 3.65x. Rui Han 0001, Chi Harold Liu, Guoren Wang, Lydia Y. Chen |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | ElasticFed: Collaborative Large-Small Transformer Training for Federated Continual Learning at EdgeabstractExecuting transformer-based applications on edge devices encounters challenging scenarios of continually learning new tasks. Federated continual learning (FCL) is a prevalent framework that supports model training using local data across edge devices. However, state-of-the-art FCL techniques either cause high computation and communication costs on large transformer models, or train small/compressed models, limiting both the learning capacity/accuracy on clients' local data and the global knowledge exchange among them. In this paper, we propose ElasticFed, a neuron-grained scaling approach for large-small transformer collaborative training in edge-based FCL. ElasticFed's key design features are (i) a proxy mechanism in local training, which constructs a compact model consisting of the original transformer's most important neurons to the current task, thus collaboratively training both models with small overheads; and (ii) a neuron-grained global aggregator, which se-lectively aggregates different clients' knowledge belonging to the most important neurons, thus maximizing the positive knowledge transfer with small communication costs. The comparison results against state-of-the-art techniques show that ElasticFed improves accuracy by 59.69% under the same training time. Compared to the techniques on original transformers, ElasticFed reduces training time and communication costs by 1.5x and 2.9x with small accuracy losses of 0.88% Yunlai Cheng, Rui Han 0001, Chi Harold Liu |
INFOCOM | 2 |
| 2025 | Adaptive ensemble optimization for memory-related hyperparameters in retraining DNN at edge
Yidong Xu, Rui Han 0001, Xiaojiang Zuo, Junyan Ouyang, Chi Harold Liu, Lydia Y. Chen |
Future Gener. Comput. Syst. | 2 |
| 2025 | Accuracy-aware differential privacy in federated learning of large transformer models
Junyan Ouyang, Rui Han 0001, Xiaojiang Zuo, Yunlai Cheng, Chi Harold Liu |
J. Inf. Secur. Appl. | 2 |
| 2025 | MIFS: An adaptive multipath information fused self-supervised framework for drug discovery
Qun Liu 0005, Rui Han 0001, Yike Guo, Guoyin Wang 0001 |
Neural Networks | 3 |
| 2025 | EdgeTA: Neuron-Grained Scaling of Foundation Models in Edge-Side RetrainingabstractFoundation models (FMs) such as large language models are becoming the backbone technology for artificial intelligence systems. It is particularly challenging to deploy multiple FMs on edge devices, which not only have limited computational resources, but also encounter unseen input data from evolving domains or learning tasks. When new data arrives, existing prior art of FM mainly focuses on retraining compressed models of predetermined network architectures, limiting the feasibility of edge devices to efficiently achieve high accuracy for FMs. In this paper, we propose EdgeTA, a neuron-grained FM scaling system to maximize the overall accuracy of FMs promptly in response to their data dynamics. EdgeTA's key design features in scaling are (i) proxy mechanism, which adaptively transforms a FM into a compact architecture retaining the most important neurons to the input data, and (ii) neuron-grained scheduler, which jointly optimizes model sizes and resource allocation for all FMs on edge devices. Under tight retraining window and limited device resources, the design of EdgeTA can achieve most of the original FM's accuracy with much smaller retraining costs. We implement EdgeTA on FMs of natural language processing, computer vision and multimodal applications. Comparison results against state-of-the-art techniques show that our approach improves accuracy by 21.88% and reduces memory footprint and energy consumptions by 27.14% and 65.65%, while further achieving 15.96% overall accuracy improvement via neuron-grained scheduling. Rui Han 0001, Chi Harold Liu, Guoren Wang, Song Guo 0001, Lydia Y. Chen |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | Loci: Federated Continual Learning of Heterogeneous Tasks at EdgeabstractFederated continual learning (FCL) has attracted growing attention in achieving collaborative model training among edge clients, each of which learns its local model for a sequence of tasks. Most existing FCL approaches aggregate clients’ latest local models to exchange knowledge. This unfortunately deviates from real-world scenarios where each model is optimized independently using the client’s own dynamic data and different clients have heterogeneous tasks. These tasks not only have distinct class labels (e.g., animals or vehicles) but also differ in input feature distributions. The aggregated model thus often shifts to a higher loss value and incurs accuracy degradation. In this article, we depart from the model-grained view of aggregation and transform it into multiple task-grained aggregations. Each aggregation allows a client to learn from other clients to improve its model accuracy on one task. To this end, we propose Loci to provide abstractions for clients’ past and peer task knowledge using compact model weights, and develop a communication-efficient approach to train each client’s local model by exchanging its tasks’ knowledge with the most accuracy relevant one from other clients. Through its general-purpose API, Loci can be used to provide efficient on-device training for existing deep learning applications of graph, image, nature language processing, and multimodal data. Using extensive comparative evaluations, we show Loci improves the model accuracy by 32.48% without increasing training time, reduces communication cost by 83.6%, and achieves more improvements when scale (task/client number) increases. Yaxin Luopan, Rui Han 0001, Xiaojiang Zuo, Chi Harold Liu, Guoren Wang, Lydia Y. Chen |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | APK-MRL: An Adaptive Pre-training Framework with Knowledge-enhanced for Molecular Representation LearningabstractAs a dominant pre-training paradigm, molecular contrastive learning (MCL) has been proven effective in learning molecular representations with unlabeled data. However, high data dependency and scarce domain knowledge caused by data augmentation in MCL limit the model’s generalization and performance. To address these issues, we propose an adaptive pre-training framework with knowledge-enhanced for molecular representation learning, named APK-MRL. It seamlessly integrates diverse prior information on hierarchical skeletons and the chemical semantics of molecules, aiming to obtain stronger stability and generalization. Extensive computational experiments demonstrate that APK-MRL can achieve competitive performances over state-of-the-art baselines on both drug-target interaction and molecular properties prediction tasks. All code is released at https://github.com/lukcats/APK-MRL. Qun Liu 0005, Rui Han 0001, Li Liu 0030, Yike Guo, Guoyin Wang 0001 |
BIBM | 3 |
| 2024 | FedViT: Federated continual learning of vision transformer at edge
Xiaojiang Zuo, Yaxin Luopan, Rui Han 0001, Chi Harold Liu, Guoyin Wang 0001, Lydia Y. Chen |
Future Gener. Comput. Syst. | 3 |
| 2024 | ElasticDNN: On-Device Neural Network Remodeling for Adapting Evolving Vision Domains at EdgeabstractExecuting deep neural networks (DNN) based vision tasks on edge devices encounters challenging scenarios of significant and continually evolving data domains (e.g. background or subpopulation shift). With limited resources, the state-of-the-art domain adaptation (DA) methods either cause high training overheads on large DNN models, or incur significant accuracy losses when adapting small/compressed models in an online fashion. The inefficient resource scheduling among multiple applications further degrades their overall model accuracy. In this paper, we present ElasticDNN, a framework that enables online DNN remodeling for applications encountering evolving domain drifts at edge. Its first key component is the master-surrogate DNN models, which can dynamically generate a small surrogate DNN by retaining and training the large master DNN's most relevant regions pertinent to the new domain. The second novelty of ElasticDNN is the filter-grained resource scheduling, which allocates GPU resources based on online accuracy estimation and DNN remodeling of co-running applications. We fully implement ElasticDNN and demonstrate its effectiveness through extensive experiments. The results show that, compared to existing online DA methods using the same model sizes, ElasticDNN improves accuracy by 23.31% and reduces adaption time by 35.67x. In the more challenging multi-application scenario, ElasticDNN improves accuracy by an average of 25.91%. Rui Han 0001, Chi Harold Liu, Guoren Wang, Lydia Y. Chen |
IEEE Trans. Computers | 2 |
| 2023 | FedKNOW: Federated Continual Learning with Signature Task Knowledge Integration at EdgeabstractDeep Neural Networks (DNNs) have been ubiquitously adopted in internet of things and are becoming an integral of our daily life. When tackling the evolving learning tasks in real world, such as classifying different types of objects, DNNs face the challenge to continually retrain themselves according to the tasks on different edge devices. Federated continual learning is a promising technique that offers partial solutions but yet to overcome the following difficulties: the significant accuracy loss due to the limited on-device processing, the negative knowledge transfer caused by the limited communication of non-IID data, and the limited scalability on the tasks and edge devices. In this paper, we propose FedKNOW, an accurate and scalable federated continual learning framework, via a novel concept of signature task knowledge. FedKNOW is a client side solution that continuously extracts and integrates the knowledge of signature tasks which are highly influenced by the current task. Each client of FedKNOW is composed of a knowledge extractor, a gradient restorer and, most importantly, a gradient integrator. Upon training for a new task, the gradient integrator ensures the prevention of catastrophic forgetting and mitigation of negative knowledge transfer by effectively combining signature tasks identified from the past local tasks and other clients’ current tasks through the global model. We implement FedKNOW in PyTorch and extensively evaluate it against state-of-the-art techniques using popular federated continual learning benchmarks. Extensive evaluation results on heterogeneous edge devices show that FedKNOW improves model accuracy by 63.24% without increasing model training time, reduces communication cost by 34.28%, and achieves more improvements under difficult scenarios such as large numbers of tasks or clients, and training different complex networks. Yaxin Luopan, Rui Han 0001, Chi Harold Liu, Guoren Wang, Lydia Y. Chen |
ICDE | 2 |
| 2023 | EdgeVisionBench: A Benchmark of Evolving Input Domains for Vision Applications at EdgeabstractVision applications powered by deep neural networks (DNNs) are widely deployed on edge devices and solve the learning tasks of incoming data streams whose class label and input feature continuously evolve, known as domain shift. Despite its prominent presence in real-world edge scenarios, existing benchmarks used by domain adaptation methods overlook evolving domains and under represent their shifts in label and feature distributions. To address this gap, we present EdgeVisionBench, a benchmark seeking to generate evolving domains of various types and reflect their realistic label and feature shifts encountered by edge-based vision applications. To facilitate evaluating domain adaptation methods on edge devices, we provide an open-source package that automates workload generation, contains popular DNN models and compression techniques, and standardizes evaluations with interactive interfaces. Code and datasets are available at https://github.com/LINC-BIT/EdgeVisionBench. Rui Han 0001, Chi Harold Liu, Guoren Wang, Lydia Y. Chen |
ICDE | 2 |
| 2023 | Evaluating Differential Privacy in Federated Continual LearningabstractIn recent years, the privacy-protecting framework Differential Privacy (DP) has achieved remarkable success and has been widely studied. However, there is a lack of work on DP in the area of Federated Continual Learning (FCL), which is a combination of Federated Learning (FL) and Continual Learning (CL). This paper presents a formal definition of DP-FCL and evaluates several DP-FCL methods based on Gaussian DP (GDP) and Individual DP (IDP). The experimental results indicate that gradient modification based CL strategies are not practical in DP-FCL. To the best of our knowledge, this is the first work to experimentally study DP-FCL, which can provide a reference for future research in this area. Junyan Ouyang, Rui Han 0001, Chi Harold Liu |
VTC Fall | 2 |
| 2023 | Delay-Sensitive Energy-Efficient UAV Crowdsensing by Deep Reinforcement LearningabstractMobile crowdsensing (MCS) by unmanned aerial vehicles (UAVs) servicing delay-sensitive applications becomes popular by navigating a group of UAVs to take advantage of their equipped high-precision sensors and durability for data collection in harsh environments. In this paper, we aim to simultaneously maximize collected data amount, geographical fairness, and minimize the energy consumption of all UAVs, as well as to guarantee the data freshness by setting a deadline in each timeslot. Specifically, we propose a centralized control, distributed execution framework by decentralized deep reinforcement learning (DRL) for delay-sensitive and energy-efficient UAV crowdsensing, called “DRL-eFresh”. It includes a synchronous computational architecture with GRU sequential modeling to generate multi-UAV navigation decisions. Also, we derive an optimal time allocation solution for data collection while considering all UAV efforts and avoiding much data dropout due to limited data upload time and wireless data rate. Simulation results show that DRL-eFresh significantly improves the energy efficiency, as compared to the best baseline DPPO, by 14% and 22% on average when varying different sensing ranges and number of PoIs, respectively. Zipeng Dai, Chi Harold Liu, Rui Han 0001, Guoren Wang, Kin K. Leung, Jian Tang 0008 |
IEEE Trans. Mob. Comput. | 3 |
| 2022 | Human-Drone Collaborative Spatial Crowdsourcing by Memory-Augmented and Distributed Multi-Agent Deep Reinforcement LearningabstractSpatial crowdsourcing (SC) has been proved quite successful by employing human participants to achieve certain tasks like Uber and Gigwalk. Meanwhile, with the fast devel-opment of unmanned aerial vehicles (e.g., drones), they have become a new source of data collectors equipped with a variety of different sensors. In this paper, we propose a novel SC scenario, enabling human participants to work collaboratively with drones in the presence of multiple charging stations to achieve certain data collection tasks, like videography and surveillance. We propose a novel deep reinforcement learning (D RL) framework called “FD- MAPPO (Cubic Map)”, which consists of a fully de-centralized multi-agent DRL (MADRL) algorithm called “Fully Decentralized Multi-Agent Proximal Policy Optimization (FD-MAPPO)”, and a spatiotemporal memory augmented neural network with novel cubic writing and spatially contextual reading mechanisms called “Cubic Map”. Cubic Map extracts long-term spatiotemporal features, navigates drones to accurately locate the position of the target, i.e., charging stations or sensors. Extensive results on two real datasets of KAIST and NCSU campuses show that FD- MAPPO (Cubic Map) consistently outperforms six other baselines in terms of efficiency. Yu Wang 0115, Chi Harold Liu, Chengzhe Piao, Ye Yuan 0001, Rui Han 0001, Guoren Wang, Jian Tang 0008 |
ICDE | 5 |
| 2022 | AoI-minimal UAV Crowdsensing by Model-based Graph Convolutional Reinforcement LearningabstractMobile Crowdsensing (MCS) with smart devices has become an appealing paradigm for urban sensing. With the development of 5G-and-beyond technologies, unmanned aerial vehicles (UAVs) become possible for real-time applications, including wireless coverage, search and even disaster response. In this paper, we consider to use a group of UAVs as aerial base stations (BSs) to move around and collect data from multiple MCS users, forming a UAV crowdsensing campaign (UCS). Our goal is to maximize the collected data, geographical coverage whiling minimizing the age-of-information (AoI) of all mobile users simultaneously, with efficient use of constrained energy reserve. We propose a model-based deep reinforcement learning (DRL) framework called "GCRL-min(AoI)", which mainly consists of a novel model-based Monte Carlo tree search (MCTS) structure based on state-of-the-art approach MCTS (AlphaZero). We further improve it by adding a spatial UAV-user correlation extraction mechanism by a relational graph convolutional network (RGCN), and a next state prediction module to reduce the dependance of experience data. Extensive results and trajectory visualization on three real human mobility datasets in Purdue University, KAIST and NCSU show that GCRL-min(AoI) consistently outperforms five baselines, when varying different number of UAVs and maximum coupling loss in terms of four metrics. Zipeng Dai, Chi Harold Liu, Yuxiao Ye, Rui Han 0001, Ye Yuan 0001, Guoren Wang, Jian Tang 0008 |
INFOCOM | 4 |
| 2022 | EdgeTuner: Fast Scheduling Algorithm Tuning for Dynamic Edge-Cloud Workloads and ResourcesabstractEdge-cloud jobs are rapidly prevailing in many application domains, posing the challenge of using both resource-strenuous edge devices and elastic cloud resources. Efficient resource allocation on such jobs via scheduling algorithms is essential to guarantee their performance, e.g. latency. Deep reinforcement learning (DRL) is increasingly adopted to make scheduling decisions but faces the conundrum of achieving high rewards at a low training overhead. It is unknown if such a DRL can be applied to timely tune the scheduling algorithms that are adopted in response to fast changing workloads and resources. In this paper, we propose EdgeTuner to effectively leverage DRL to select scheduling algorithms online for edge-cloud jobs. The enabling features of EdgeTuner are sophisticated DRL model that captures complex dynamics of Edge-Cloud jobs/tasks and an effective simulator to emulate the response times of short-running jobs in accordance to dynamically changing scheduling algorithms. EdgeTuner trains DRL agents offline by directly interacting with the simulator. We implement EdgeTuner on Kubernetes scheduler and extensively evaluate it on Kubernetes cluster testbed driven by the production traces. Our results show that EdgeTuner outperforms prevailing scheduling algorithms by achieving significant lower job response time while accelerating DRL training speed by more than 180x. Rui Han 0001, Shilin Wen, Chi Harold Liu, Ye Yuan 0001, Guoren Wang, Lydia Y. Chen |
INFOCOM | 1 |
| 2022 | Lightweight and Accurate DNN-Based Anomaly Detection at EdgeabstractDeep neural networks (DNNs) have been showing significant success in various anomaly detection applications such as smart surveillance and industrial quality control. It is increasingly important to detect anomalies directly on edge devices, because of high responsiveness requirements and tight latency constraints. The accuracy of DNN-based solutions rely on large model capacity and thus long training and inference time, making them inapplicable on resource strenuous edge devices. It is hence imperative to scale DNN model sizes in correspondence to the run-time system requirements, i.e. meeting deadlines with minimal accuracy losses, which are highly dependent on the platforms and real-time system status. Existing scaling techniques either take long training time to pre-generate scaling options or disturb the unsteady training process of anomaly detection DNNs, lacking the adaptability to heterogeneous edge systems and incurring low inference accuracies. In this paper, we present LightDNN to scale DNN models for anomaly detection applications at edge, featuring high detection accuracies with lightweight training and inference time. To this end, LightDNN quickly extracts and compresses blocks in a DNN, and provides large scaling space (e.g. 1 million options) by dynamically combining these compressed blocks online. At run-time, LightDNN predicts the DNN's inference latency according to the monitored system status, and optimizes the combination of blocks to maximize its accuracy under deadline constraints. We implement and extensively evaluate LightDNN on both CPU and GPU edge platforms using 8 popular anomaly detection workloads. Comparative experiments with state-of-the-art methods show that our approach provides 145.8 to 0.56 trillion times more scaling options without increasing training and inference overheads, thus achieving as much as 15.05% increase in accuracy under the same deadlines. Rui Han 0001, Gaofeng Xin, Chi Harold Liu, Guoren Wang, Lydia Y. Chen |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Federated Learning With Heterogeneity-Aware Probabilistic Synchronous Parallel on EdgeabstractWith the massive amount of data generated from mobile devices and the increase of computing power of edge devices, the paradigm of Federated Learning has attracted great momentum. In federated learning, distributed and heterogeneous nodes collaborate to learn model parameters. However, while providing benefits such as privacy by design and reduced latency, the heterogeneous network present challenges to the synchronisation methods, or barrier control methods, used in training, regarding system progress and model convergence etc. The design of these barrier mechanisms is critical for the performance and scalability of federated learning systems. We propose a new barrier control technique called Probabilistic Synchronous Parallel (PSP). In contrast to existing mechanisms, it introduces a sampling primitive that composes with existing barrier control mechanisms to produce a family of mechanisms with improved convergence speed and scalability. Our proposal is supported with a convergence analysis of PSP-based SGD algorithm. In practice, we also propose heuristic techniques that further improve the efficiency of PSP. We evaluate the performance of proposed methods using the federated learning specific FEMNSIT dataset. The evaluation results show that PSP can effectively achieve good balance between system efficiency and model accuracy, mitigating the challenge of heterogeneity in federated learning. Jianxin Zhao 0001, Rui Han 0001, Yongkai Yang, Benjamin Catterall, Chi Harold Liu, Lydia Y. Chen, Richard Mortier, Jon Crowcroft, Liang Wang 0009 |
IEEE Trans. Serv. Comput. | 2 |
| 2022 | MespaConfig: Memory-Sparing Configuration Auto-Tuning for Co-Located In-Memory Cluster Computing JobsabstractDistributed in-memory computing frameworks usually have lots of parameters (e.g., the buffer size of shuffle) to form a configuration for each execution. A well-tuned configuration can bring large improvements of performance. However, to improve resource utilization, jobs are often share the same cluster, which causes dynamic cluster load conditions. According to our observation, the variation of cluster load reduces effectiveness of configuration tuning. Besides, as a common problem of cluster computing jobs, overestimation of resources also occurs during configuration tuning. It is challenging to efficiently find the optimal configuration in a shared cluster with the consideration of memory-sparing. In this article, we introduce MespaConfig, a job-level configuration optimizer for distributed in-memory computing jobs. Advancements of MespaConfig over previous work are features including memory-sparing and load-sensitive. We evaluate MespaConfig by 6 typical Spark programs under different load conditions. The evaluation results show that MespaConfig improves the performance of six typical programs by up to 12× compared with default configurations. MespaConfig also achieves at most 41 percent reduction of configuration memory usage and reduces the optimization time overhead by 10.8× compared with the state-of-the-art approach. Zan Zong, Lijie Wen 0001, Xuming Hu, Rui Han 0001, Chen Qian 0003, Li Lin 0011 |
IEEE Trans. Serv. Comput. | 4 |
| 2021 | LABELNET: Recovering Noisy Labels
Amirmasoud Ghiassi, Robert Birke, Rui Han 0001, Lydia Y. Chen |
IJCNN | 3 |
| 2021 | Mobile Crowdsensing for Data Freshness: A Deep Reinforcement Learning ApproachabstractData collection by mobile crowdsensing (MCS) is emerging as data sources for smart city applications, however how to ensure data freshness has sparse research exposure but quite important in practice. In this paper, we consider to use a group of mobile agents (MAs) like UAVs and driverless cars which are equipped with multiple antennas to move around in the task area to collect data from deployed sensor nodes (SNs). Our goal is to minimize the age of information (AoI) of all SNs and energy consumption of MAs during movement and data upload. To this end, we propose a centralized deep reinforcement learning (DRL)-based solution called "DRL-freshMCS" for controlling MA trajectory planning and SN scheduling. We further utilize implicit quantile networks to maintain the accurate value estimation and steady policies for MAs. Then, we design an exploration and exploitation mechanism by dynamic distributed prioritized experience replay. We also derive the theoretical lower bound for episodic AoI. Extensive simulation results show that DRL-freshMCS significantly reduces the episodic AoI per remaining energy, compared to five baselines when varying different number of antennas and data upload thresholds, and number of SNs. We also visualize their trajectories and AoI update process for clear illustrations. Zipeng Dai, Hao Wang 0193, Chi Harold Liu, Rui Han 0001, Jian Tang 0008, Guoren Wang |
INFOCOM | 4 |
| 2021 | LegoDNN: block-grained scaling of deep neural networks for mobile visionabstractDeep neural networks (DNNs) have become ubiquitous techniques in mobile and embedded systems for applications such as image/object recognition and classification. The trend of executing multiple DNNs simultaneously exacerbate the existing limitations of meeting stringent latency/accuracy requirements on resource constrained mobile devices. The prior art sheds light on exploring the accuracy-resource tradeoff by scaling the model sizes in accordance to resource dynamics. However, such model scaling approaches face to imminent challenges: (i) large space exploration of model sizes, and (ii) prohibitively long training time for different model combinations. In this paper, we present LegoDNN, a lightweight, block-grained scaling solution for running multi-DNN workloads in mobile vision systems. LegoDNN guarantees short model training times by only extracting and training a small number of common blocks (e.g. 5 in VGG and 8 in ResNet) in a DNN. At run-time, LegoDNN optimally combines the descendant models of these blocks to maximize accuracy under specific resources and latency constraints, while reducing switching overhead via smart block-level scaling of the DNN. We implement LegoDNN in TensorFlow Lite and extensively evaluate it against state-of-the-art techniques (FLOP scaling, knowledge distillation and model compression) using a set of 12 popular DNN models. Evaluation results show that LegoDNN provides 1,296x to 279,936x more options in model sizes without increasing training time, thus achieving as much as 31.74% improvement in inference accuracy and 71.07% reduction in scaling energy consumptions. Rui Han 0001, Chi Harold Liu, Guoren Wang, Jian Tang 0008, Lydia Y. Chen |
MobiCom | 1 |
| 2021 | Enhancing Robustness of On-Line Learning Models on Highly Noisy DataabstractClassification algorithms have been widely adopted to detect anomalies for various systems, e.g., IoT, cloud and face recognition, under the common assumption that the data source is clean, i.e., features and labels are correctly set. However, data collected from the wild can be unreliable due to careless annotations or malicious data transformation for incorrect anomaly detection. In this article, we extend a two-layer on-line data selection framework: Robust Anomaly Detector (RAD) with a newly designed ensemble prediction where both layers contribute to the final anomaly detection decision. To adapt to the on-line nature of anomaly detection, we consider additional features of conflicting opinions of classifiers, repetitive cleaning, and oracle knowledge. We on-line learn from incoming data streams and continuously cleanse the data, so as to adapt to the increasing learning capacity from the larger accumulated data set. Moreover, we explore the concept of oracle learning that provides additional information of true labels for difficult data points. We specifically focus on three use cases, (i) detecting 10 classes of IoT attacks, (ii) predicting 4 classes of task failures of big data jobs, and (iii) recognising 100 celebrities faces. Our evaluation results show that RAD can robustly improve the accuracy of anomaly detection, to reach up to 98.95 percent for IoT device attacks (i.e., +7%), up to 85.03 percent for cloud task failures (i.e., +14%) under 40 percent label noise, and for its extension, it can reach up to 77.51 percent for face recognition (i.e., +39%) under 30 percent label noise. The proposed RAD and its extensions are general and can be applied to different anomaly detection algorithms. Zilong Zhao 0001, Robert Birke, Rui Han 0001, Bogdan Robu, Sara Bouchenak, Sonia Ben Mokhtar, Lydia Y. Chen |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2021 | SlimML: Removing Non-Critical Input Data in Large-Scale Iterative Machine LearningabstractThe core of many large-scale machine learning (ML) applications, such as neural networks (NN), support vector machine (SVM), and convolutional neural network (CNN), is the training algorithm that iteratively updates model parameters by processing massive datasets. From a plethora of studies aiming at accelerating ML, being data parallelization and parameter server, the prevalent assumption is that all data points are equivalently relevant to model parameter updating. In this article, we challenge this assumption by proposing a criterion to measure a data point's effect on model parameter updating, and experimentally demonstrate that the majority of data points are non-critical in the training process. We develop a slim learning framework, termed SlimML, which trains the ML models only on the critical data and thus significantly improves training performance. To such an end, SlimML efficiently leverages a small number of aggregated data points per iteration to approximate the criticalness of original input data instances. The proposed approach can be used by changing a few lines of code in a standard stochastic gradient descent (SGD) procedure, and we demonstrate experimentally, on NN regression, SVM classification, and CNN training, that for large datasets, it accelerates model training process by an average of 3.61 times while only incurring accuracy losses of 0.37 percent. Rui Han 0001, Chi Harold Liu, Shilin Li, Lydia Y. Chen, Guoren Wang, Jian Tang 0008, Jieping Ye |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2021 | Accurate Differentially Private Deep Learning on the EdgeabstractDeep learning (DL) models are increasingly built on federated edge participants holding local data. To enable insight extractions without the risk of information leakage, DL training is usually combined with differential privacy (DP). The core theme is to tradeoff learning accuracy by adding statistically calibrated noises, particularly to local gradients of edge learners, during model training. However, this privacy guarantee unfortunately degrades model accuracy due to edge learners' local noises, and the global noise aggregated at the central server. Existing DP frameworks for edge focus on local noise calibration via gradient clipping techniques, overlooking the heterogeneity and dynamic changes of local gradients, and their aggregated impact on accuracy. In this article, we present a systematical analysis that unveils the influential factors capable of mitigating local and aggregated noises, and design PrivateDL to leverage these factors in noise calibration so as to improve model accuracy while fulfilling privacy guarantee. PrivateDL features on: (i) sampling-based sensitivity estimation for local noise calibration and (ii) combining large batch sizes and critical data identification in global training. We implement PrivateDL on the popular Laplace/Gaussian DP mechanisms and demonstrate its effectiveness using Intel BigDL workloads, i.e., considerably improving model accuracy by up to 5X when comparing against existing DP frameworks. Rui Han 0001, Junyan Ouyang, Chi Harold Liu, Guoren Wang, Dapeng Oliver Wu, Lydia Y. Chen |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | Accelerating Gossip-Based Deep Learning in Heterogeneous Edge Computing PlatformsabstractWith the exponential growth of data created at the network edge, decentralized and Gossip-based training of deep learning (DL) models on edge computing (EC) gains tremendous research momentum, owing to its capability to learn from resource-strenuous edge nodes with limited network connectivity. Today's edge devices are extremely heterogeneous, e.g., hardware and software stacks, and result in high performance variation of training time and inducing extra delay to synchronize and converge. The large body of prior art accelerates DL, being data or model parallelization, via a centralized server, e.g., parameter server scheme, which may easily turn into the system bottleneck or single point of failure. In this artice, we propose EdgeGossip, a framework specifically designed to accelerate the training process of decentralized and Gossip-based DL training for heterogeneous EC platforms. EdgeGossip features on: (i) low performance variation among multiple EC platforms during iterative training, and (ii) accuracy-aware training to fastly obtain best possible model accuracy. We implement EdgeGossip based on popular Gossip algorithms and demonstrate its effectiveness using real-world DL workloads, i.e., considerably reducing model training time by an average of 2.70 times while only incurring accuracy losses of 0.78 percent. Rui Han 0001, Shilin Li, Chi Harold Liu, Gaofeng Xin, Lydia Y. Chen |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | Accelerating Deep Learning Systems via Critical Set Identification and Model CompressionabstractModern distributed engines are increasingly deployed to accelerate large-scaled deep learning (DL) training jobs. While the parallelism of distributed workers/nodes promises the scalability, the computation and communication overheads of the underlying iterative solving algorithms, e.g., stochastic gradient decent, unfortunately become the bottleneck for distributed DL training jobs. Existing approaches address such limitations by designing more efficient synchronization algorithms and model compressing techniques, but do not adequately address issues relating to processing massive datasets. In this article, we propose ClipDL, which accelerates the deep learning systems by simultaneously decreasing the number of model parameters as well as reducing the computations on critical data only. The core component of ClipDL is the estimation of critical set based on the observation that large proportions of input data have little influence on model parameter updating in many prevalent DL algorithms. We implemented ClipDL on Spark (a popular distributed engine for big data) and BigDL (based on de-factor distributed DL training architecture, parameter server), and integrated it with representative model compression techniques. The exhaustive experiments on real DL applications and datasets show ClipDL accelerates model training process by an average of 2.32 times while only incurring accuracy losses of 1.86 percent. Rui Han 0001, Chi Harold Liu, Shilin Li, Shilin Wen, Xue (Steve) Liu |
IEEE Trans. Computers | 1 |
| 2020 | Facial expression recognition with convolutional neural networks via a new face cropping and rotation strategy
Kuan Li, Yi Jin 0002, Muhammad Waqar Akram, Rui Han 0001, Jiongwei Chen |
Vis. Comput. | 4 |
| 2019 | Workload-Adaptive Configuration Tuning for Hierarchical Cloud SchedulersabstractCluster schedulers provide flexible resource sharing mechanism for best-effort cloud jobs, which occupy a majority in modern datacenters. Properly tuning a scheduler's configurations is the key to these jobs' performance because it decides how to allocate resources among them. Today's cloud scheduling systems usually rely on cluster operators to set the configuration and thus overlook the potential performance improvement through optimally configuring the scheduler according to the heterogeneous and dynamic cloud workloads. In this paper, we introduce AdaptiveConfig, a run-time configurator for cluster schedulers that automatically adapts to the changing workload and resource status in two steps. First, a comparison approach estimates jobs' performances under different configurations and diverse scheduling scenarios. The key idea here is to transform a scheduler's resource allocation mechanism and their variable influence factors (configurations, scheduling constraints, available resources, and workload status) into business rules and facts in a rule engine, thereby reasoning about these correlated factors in job performance comparison. Second, a workload-adaptive optimizer transforms the cluster-level searching of huge configuration space into an equivalent dynamic programming problem that can be efficiently solved at scale. We implement AdaptiveConfig on the popular YARN Capacity and Fair schedulers and demonstrate its effectiveness using real-world Facebook and Google workloads, i.e., successfully finding best configurations for most of scheduling scenarios and considerably reducing latencies by a factor of two with low optimization time. Rui Han 0001, Chi Harold Liu, Zan Zong, Lydia Y. Chen, Wending Liu, Jianfeng Zhan |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | AdaptiveConfig: Run-Time Configuration of Cluster Schedulers for Cloud Short-Running JobsabstractCluster schedulers provide flexible resource sharing mechanism for short-running jobs, which occupy a majority of cloud jobs. A scheduler's configuration decides how to allocate resources among jobs and hence it is crucial to their performances. Today's cloud platforms usually rely on cluster administrators to set this configuration, thus it is difficult to optimally configure the scheduler so as to minimize the latencies of heterogeneous and dynamically changing jobs in the cloud. In this paper, we introduce AdaptiveConfig, a run-time configurator for cluster schedulers that automatically adapts to the changing workload and resource status. This includes: (1) an estimator to calculate jobs' performances under different configurations and various scheduling scenarios. The key idea here is to transform a scheduler's resource allocation mechanisms and their variable influence factors (configuration parameters, scheduling constraints, available resources, and workload status) into business rules and facts in a rule engine, thereby reasoning about these correlated factors in job performance estimation. (2) A run-time optimizer that efficiently searches the configuration space to find the optimal configuration for the current workload. We implemented AdaptiveConfig on the popular YARN Capacity and Fair schedulers and demonstrate its effectiveness using workloads of Facebook jobs, i.e. considerably reducing latencies by 2.22 times (and up to 4.50 times) with low optimization overheads. Rui Han 0001, Zan Zong, Lydia Y. Chen, Jianfeng Zhan |
ICDCS | 1 |
| 2018 | Benchmarking Big Data Systems: A ReviewabstractWith the fast development of big data systems in recent years, a variety of open-source benchmarks have been built to evaluate and compare the workloads on these systems, and to promote their technology improvement. However, to date no comprehensive survey has been written on this topic. This paper attempts to fill the void by presenting a review of the state-of-the-art big data benchmarking efforts. The paper first gives an overview of popular open-source benchmarks from the point of view of big data systems. It then reviews the three important aspects of benchmarking - workload generation techniques, workload input data generation techniques, and metrics used to assess systems. For each aspect, the paper divides the surveyed benchmarks into different categories and describes some representative benchmarks, rather than all benchmarks listed, in each category, following the discussion of potential research directions to motivate future work in this area. Rui Han 0001, Lizy Kurian John, Jianfeng Zhan |
IEEE Trans. Serv. Comput. | 1 |
| 2017 | CloudMix: Generating Diverse and Reducible Workloads for Cloud SystemsabstractThe prosperity of cloud computing offers common infrastructures to a wide range of applications. Understanding these applications' workload behaviors is the premise of designing, managing, and optimizing cloud systems. Considering the heterogeneity and diversity of cloud workloads, for the sake of fairness, cloud benchmarks must be able to accurately replicate their behaviors in cloud systems, including both the usages of cloud resources and the micro-architectural behaviors beyond the virtualization layer. Furthermore, workloads spanning long durations are usually required to achieve representativeness in evaluation. Hence the more challenging issue is to significantly reduce the evaluation duration while still preserving their workload characteristics. This paper presents our efforts towards generating cloud workloads of diverse behaviors and reducible durations. Our benchmark tool, CloudMix, employs a repository of reducible workload blocks (RWBs) as the high level abstraction of workload behaviors, including usages of the two most important cloud resources (CPU and memory) and their pairing micro-architectural operations. CloudMix further introduces an efficient methodology to combine RWBs to synthesize and replicate diverse cloud workloads in real-world traces. The effectiveness of CloudMix is demonstrated by generating a variety of reducible workloads according to a Google cluster trace and by applying these workloads in job scheduling optimization on Hadoop YARN. The evaluation results show: (i) when the workload durations are reduced by 100 times, the replication errors of workload behaviors are smaller than 2.08%; (ii) when providing fast evaluations (workload durations are reduced by 10 to 100 times) to recommend the optimal setting in YARN job scheduling, the performance degradation in the recommended setting is just 0.69% compared to that of the actual optimal setting. CloudMix is publicly available from the project home page http://prof.ict.ac.cn/BigDataBench/multi tenancy/. Rui Han 0001, Zan Zong, Fan Zhang 0047, José Luis Vázquez-Poletti, Zhen Jia 0001, Lei Wang 0004 |
CLOUD | 1 |
| 2017 | AccurateML: Information-aggregation-based approximate processing for fast and accurate machine learning on MapReduceabstractThe growing demands of processing massive datasets have promoted irresistible trends of running machine learning applications on MapReduce. When processing large input data, it is often of greater values to produce fast and accurate enough approximate results than slow exact results. Existing techniques produce approximate results by processing parts of the input data, thus incurring large accuracy losses when using short job execution times, because all the skipped input data potentially contributes to result accuracy. We address this limitation by proposing AccurateML that aggregates information of input data in each map task to create small aggregated data points. These aggregated points enable all map tasks producing initial outputs quickly to save computation times and decrease the outputs' size to reduce communication times. Our approach further identifies the parts of input data most related to result accuracy, thus first using these parts to improve the produced outputs to minimize accuracy losses. We evaluated AccurateML using real machine learning applications and datasets. The results show: (i) it reduces execution times by 30 times with small accuracy losses compared to exact results; (ii) when using the same execution times, it achieves 2.71 times reductions in accuracy losses compared to existing approximate processing techniques. Rui Han 0001, Fan Zhang 0047 |
INFOCOM | 1 |
| 2017 | Work-in-Progress: Maximizing Model Accuracy in Real-time and Iterative Machine LearningabstractAs iterative machine learning (ML) (e.g. neural network based supervised learning and k-means clustering) becomes more ubiquitous in our daily life, it is becoming increasingly important to complete model training quickly to support real-time decision making, while still achieving high model accuracy (e.g. low prediction errors) that is critical for profits of ML tasks. Motivated by the observation that the small proportions of accuracy-critical input data can contribute to large parts of model accuracy in many iterative ML applications, this paper introduces a system middleware to maximize model accuracy by spending the limited time budget on the most accuracy-related input data. To achieve this, our approach employs a fast method to divide the input data into multiple parts of similar points and represents each part with an aggregated data point. Using these points, it quickly estimates the correlations between different parts and model accuracy, thus allowing ML tasks to process the most accuracy-related parts first. We incorporate our approach with two popular supervised and unsupervised ML algorithms on Spark and demonstrate its benefits in providing high model accuracy under short deadlines. Rui Han 0001, Fan Zhang 0047, Lydia Y. Chen, Jianfeng Zhan |
RTSS | 1 |
| 2017 | CLAP: Component-Level Approximate Processing for Low Tail Latency and High Result Accuracy in Cloud Online ServicesabstractModern latency-critical online services such as search engines often process requests by consulting large input data spanning massive parallel components. Hence the tail latency of these components determines the service latency. To trade off result accuracy for tail latency reduction, existing techniques use the components responding before a specified deadline to produce approximate results. However, they skip a large proportion of components when load gets heavier, thus incurring large accuracy losses. In this paper, we propose CLAP to enable component-level approximate processing of requests for low tail latency and small accuracy losses. CLAP aggregates information of input data to create small aggregated data points. Using these points, CLAP reduces latency variance of parallel components and allows them to produce initial results quickly; CLAP also identifies the parts of input data most related to requests' result accuracies, thus first using these parts to improve the produced results to minimize accuracy losses. We evaluated CLAP using real services and datasets. The results show: (i) CLAP reduces tail latency by 6.46 times with accuracy losses of 2.2 percent compared to existing exact processing techniques; (ii) when using the same latency, CLAP reduces accuracy losses by 31.58 times compared to existing approximate processing techniques. Rui Han 0001, Siguang Huang, Jianfeng Zhan |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | Understanding Big Data Analytics Workloads on Modern ProcessorsabstractBig data analytics workloads are very significant ones in modern data centers, and it is more and more important to characterize their representative workloads and understand their behaviors so as to improve the performance of data center computer systems. In this paper, we embark on a comprehensive study to understand the impacts and performance implications of the big data analytics workloads on the systems equipped with modern superscalar out-of-order processors. After investigating three most important application domains in Internet services in terms of page views and daily visitors, we choose 11 representative data analytics workloads and characterize their micro-architectural behaviors by using hardware performance counters. Our study reveals that the big data analytics workloads share many inherent characteristics, which place them in a different class from the traditional workloads and the scale-out services. To further understand the characteristics of big data analytics workloads, we perform correlation analysis to identify the most key factors that affect cycles per instruction (CPI). Also, we reveal that the increasing complexity of the big data software stacks will put higher pressures on the modern processor pipelines. Zhen Jia 0001, Jianfeng Zhan, Lei Wang 0004, Chunjie Luo, Wanling Gao, Rui Han 0001, Lixin Zhang 0002 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2016 | AccuracyTrader: Accuracy-Aware Approximate Processing for Low Tail Latency and High Result Accuracy in Cloud Online ServicesabstractModern latency-critical online services such as search engines often process requests by consulting large input data spanning massive parallel components. Hence the tail latency of these components determines the service latency. To trade off result accuracy for tail latency reduction, existing techniques use the components responding before a specified deadline to produce approximate results. However, they may skip a large proportion of components when load gets heavier, thus incurring large accuracy losses. This paper presents AccuracyTrader that produces approximate results with small accuracy losses while maintaining low tail latency. AccuracyTrader aggregates information of input data on each component to create a small synopsis, thus enabling all components producing initial results quickly using their synopses. AccuracyTrader also uses synopses to identify the parts of input data most related to arbitrary requests' result accuracy, thus first using these parts to improve the produced results in order to minimize accuracy losses. We evaluated AccuracyTrader using workloads in real services. The results show: (i) AccuracyTrader reduces tail latency by over 40 times with accuracy losses of less than 7% compared to existing exact processing techniques, (ii) when using the same latency, AccuracyTrader reduces accuracy losses by over 13 times comparing to existing approximate processing techniques. Rui Han 0001, Siguang Huang, Fei Tang 0003, Fu-Gui Chang, Jianfeng Zhan |
ICPP | 1 |
| 2015 | Interference-Aware Component Scheduling for Reducing Tail Latency in Cloud Interactive ServicesabstractLarge-scale interactive services usually divide requests into multiple sub-requests and distribute them to a large number of server components for parallel execution. Hence the tail latency (i.e. The slowest component's latency) of these components determines the overall service latency. On a cloud platform, each component shares and competes node resources such as caches and I/O bandwidths with its co-located jobs, hence inevitably suffering from their performance interference. In this paper, we study the short-running jobs in a 12k-node Google cluster to illustrate the dynamic resource demands of these jobs, resulting in both individual components' latency variability over time and across different nodes and hence posing a major challenge to maintain low tail latency. Given this motivation, this paper introduces a dynamic and interference-aware scheduler for large-scale, parallel cloud services. At each scheduling interval, it collects workload and resource contention information of a running service, and predicts both the component latency on different nodes and the overall service performance. Based on the predicted performance, the scheduler identifies straggling components and conducts near-optimal component-node allocations to adapt to the changing workloads and performance interferences. We demonstrate that, using realistic workloads, the proposed approach achieves significant reductions in tail latency compared to the basic approach without scheduling. Rui Han 0001, Siguang Huang, Chenrong Shao, Shulin Zhan, Jianfeng Zhan, José Luis Vázquez-Poletti |
ICDCS | 1 |
| 2015 | PCS: Predictive Component-Level Scheduling for Reducing Tail Latency in Cloud Online ServicesabstractModern latency-critical online services often rely on composing results from a large number of server components. Hence the tail latency (e.g. The 99th percentile of response time), rather than the average, of these components determines the overall service performance. When hosted on a cloud environment, the components of a service typically co-locate with short batch jobs to increase machine utilizations, and share and contend resources such as caches and I/O bandwidths with them. The highly dynamic nature of batch jobs in terms of their workload types and input sizes causes continuously changing performance interference to individual components, hence leading to their latency variability and high tail latency. However, existing techniques either ignore such fine-grained component latency variability when managing service performance, or rely on executing redundant requests to reduce the tail latency, which adversely deteriorate the service performance when load gets heavier. In this paper, we propose PCS, a predictive and component-level scheduling framework to reduce tail latency for large-scale, parallel online services. It uses an analytical performance model to simultaneously predict the component latency and the overall service performance on different nodes. Based on the predicted performance, the scheduler identifies straggling components and conducts near-optimal component-node allocations to adapt to the changing performance interferences from batch jobs. We demonstrate that, using realistic workloads, the proposed scheduler reduces the component tail latency by an average of 67.05% and the average overall service latency by 64.16% compared with the state-of-the-art techniques on reducing tail latency. Rui Han 0001, Siguang Huang, Chenrong Shao, Shulin Zhan, Jianfeng Zhan, José Luis Vázquez-Poletti |
ICPP | 1 |
| 2014 | Enabling cost-aware and adaptive elasticity of multi-tier cloud applications
Rui Han 0001, Moustafa Ghanem, Li Guo 0002, Yike Guo, Michelle Osmond |
Future Gener. Comput. Syst. | 1 |
| 2013 | Elastic algorithms for guaranteeing quality monotonicity in big data miningabstractWhen mining large data volumes in big data applications users are typically willing to use algorithms that produce acceptable approximate results satisfying the given resource and time constraints. Two key challenges arise when designing such algorithms. The first relates to reasoning about tradeoffs between the quality of data mining output, e.g. prediction accuracy for classification tasks and available resource and time budgets. The second is organizing the computation of the algorithm to guarantee producing better quality of results as more budget is used. Little work has addressed these two challenges together in a generic way. In this paper, we propose a novel framework for developing elastic big data mining algorithms. Based on Shannon's entropy, an information-theoretic approach is introduced to reason about how result quality is affected by the allocated budget. This is then used to guide the development of algorithms that adapt to the available time budgets while guaranteeing producing better quality results as more budgets are used. We demonstrate the application of the framework by developing elastic k-Nearest Neighbour (kNN) classification and collaborative filtering (CF) recommendation algorithms as two examples. The core of both elastic algorithms is to use a naïve kNN classification or CF algorithm over R-tree data structures that successively approximate the entire datasets. Experimental evaluation was performed using prediction accuracy as quality metric on real datasets. The results show that elastic mining algorithms indeed produce results with consistent increase in observable qualities, i.e., prediction accuracy, in practice. Rui Han 0001, Lei Nie 0008, Moustafa Ghanem, Yike Guo |
IEEE BigData | 1 |
| 2012 | Elastic Application Container: A Lightweight Approach for Cloud Resource ProvisioningabstractVirtual machine (VM) based virtual infrastructure has been adopted widely in cloud computing environment for elastic resource provisioning. Performing resource management using VMs, however, is a heavyweight task. In practice, we have identified two scenarios where VM based resource management is less feasible and less resource-efficient. In this paper, we propose a lightweight resource management model that is called Elastic Application Container (EAC). EAC is a virtual resource unit for delivering better resource efficiency and more scalable cloud applications. We describe the EAC system architecture and components, and also present an algorithm for EAC resource provisioning. We also describe an implementation of the EAC-oriented platform to support multi-tenant cloud use. To evaluate our approach and implementation, we conducted experiments and collected performance data by comparing VM-based and EAC-based resource management with regards to their feasibility and resource-efficiency. The experiment results show that our proposed EAC-based resource management approach outperforms the VM-based approach in terms of feasibility and resource-efficiency. Sijin He, Li Guo 0002, Yike Guo, Chao Wu 0001, Moustafa Ghanem, Rui Han 0001 |
AINA | 6 |
| 2012 | Lightweight Resource Scaling for Cloud ApplicationsabstractElastic resource provisioning is a key feature of cloud computing, allowing users to scale up or down resource allocation for their applications at run-time. To date, most practical approaches to managing elasticity are based on allocation/de-allocation of the virtual machine (VM) instances to the application. This VM-level elasticity typically incurs both considerable overhead and extra costs, especially for applications with rapidly fluctuating demands. In this paper, we propose a lightweight approach to enable cost-effective elasticity for cloud applications. Our approach operates fine-grained scaling at the resource level itself (CPUs, memory, I/O, etc) in addition to VM-level scaling. We also present the design and implementation of an intelligent platform for light-weight resource management of cloud applications. We describe our algorithms for light-weight scaling and VM-level scaling and show their interaction. We then use an industry standard benchmark to evaluate the effectiveness of our approach and compare its performance against traditional approaches. Rui Han 0001, Li Guo 0002, Moustafa Ghanem, Yike Guo |
CCGRID | 1 |
| 2012 | Does the Cloud need new algorithms? An introduction to elastic algorithmsabstractCloud computing has emerged as a cost-effective way to deliver metered computing resources. Within a Cloud, elasticity of resource usage is typically realized through the “on-demand” provision principle supported by the “Pay-as-You-Go” business model. However, little, or no work, has investigated elasticity of algorithms for Cloud computing. In this paper, we introduce novel research on elastic algorithms (EA) where the computation itself is organized in a “Pay-as-You-Go” fashion. In contrast to conventional algorithms, where computation is a deterministic process that only produces an “ali-or-nothing” result, an EA generates a sequence of approximate results corresponding to its resource consumption. As more resources are consumed, better results will be derived. In this sense, the quality of the algorithm is elastic to its resource consumption. In the paper, we formalize the proeprties of elasticity and also formalize desirable properties for elastic algorithms themselves. We illustrate the design of an EA for kNN classification in the context of machine learning and discuss its properties. Finally we provide an ambitious agenda for future research in this area. Yike Guo, Moustafa Ghanem, Rui Han 0001 |
CloudCom | 3 |
| 2012 | Modelling and performance analysis of clinical pathways using the stochastic process algebra PEPAabstractBACKGROUND: Hospitals nowadays have to serve numerous patients with limited medical staff and equipment while maintaining healthcare quality. Clinical pathway informatics is regarded as an efficient way to solve a series of hospital challenges. To date, conventional research lacks a mathematical model to describe clinical pathways. Existing vague descriptions cannot fully capture the complexities accurately in clinical pathways and hinders the effective management and further optimization of clinical pathways. METHOD: Given this motivation, this paper presents a clinical pathway management platform, the Imperial Clinical Pathway Analyzer (ICPA). By extending the stochastic model performance evaluation process algebra (PEPA), ICPA introduces a clinical-pathway-specific model: clinical pathway PEPA (CPP). ICPA can simulate stochastic behaviours of a clinical pathway by extracting information from public clinical databases and other related documents using CPP. Thus, the performance of this clinical pathway, including its throughput, resource utilisation and passage time can be quantitatively analysed. RESULTS: A typical clinical pathway on stroke extracted from a UK hospital is used to illustrate the effectiveness of ICPA. Three application scenarios are tested using ICPA: 1) redundant resources are identified and removed, thus the number of patients being served is maintained with less cost; 2) the patient passage time is estimated, providing the likelihood that patients can leave hospital within a specific period; 3) the maximum number of input patients are found, helping hospitals to decide whether they can serve more patients with the existing resource allocation. CONCLUSIONS: ICPA is an effective platform for clinical pathway management: 1) ICPA can describe a variety of components (state, activity, resource and constraints) in a clinical pathway, thus facilitating the proper understanding of complexities involved in it; 2) ICPA supports the performance analysis of clinical pathway, thereby assisting hospitals to effectively manage time and resources in clinical pathway. Xian Yang 0001, Rui Han 0001, Yike Guo, Jeremy T. Bradley, Benita Cox, Robert Dickinson, Richard Kitney |
BMC Bioinform. | 2 |
| 2011 | A Deployment Platform for Dynamically Scaling Applications in the CloudabstractSimplifying the process of deploying applications is almost essential in the cloud. However, existing techniques can automate applications' initial deployment but have not yet adequately addressed their scaling problem. In this paper, a deployment platform to enable a novel dynamic scaling technique is introduced. This platform employs: (i) an extensible specification that describes all aspects of applications, (ii) a flexible analytical model that determines how many servers to be deployed for an application in each scaling. The platform's ability to handle dynamic workloads and to scale applications quickly enough to maintain the response time target is demonstrated. Rui Han 0001, Li Guo 0002, Yike Guo, Sijin He |
CloudCom | 1 |
| 2010 | Dynamically Analyzing Time Constraints in Workflow Systems with Fixed-Date ConstraintabstractIn workflow management systems (WFMSs), time management plays an essential role in controlling the lifecycle of business processes. Especially, run-time analysis of time constraints is necessary to help process manger proactively detect possible deadline violations and appropriately handle these violations. Traditional time constraint analyses either present deterministic results which are too restrictive in highly uncertain workflow processes, or only consider static analysis at workflow build-time. For such an issue, this paper proposes a dynamic approach for analyzing time constraints during process execution. To be specific, based on a Petri-net-extended stochastic model, this approach first analyzes activity instances’ continuous probabilities of satisfying time constraints when a process instance is initiated. Afterwards, during the execution of this process instance, the approach dynamically updates these probabilities whenever an activity instance is completed. Moreover, an example process instance in real-world WFMSs shows the practicality of our approach. Rui Han 0001, Lijie Wen 0001, Jianmin Wang 0001 |
APWeb | 1 |