Young Choon Lee

dblp:52/6328 · DBLP profile ↗
← Back
92ranked-venue papers
21as first author
11since 2021 · last 2026
0000-0001-8560-6199ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 48 · 17 first-author · 2 since 2021Computer networks · 7 · 4 since 2021Software engineering, systems software and programming languages · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorSecurity and privacy · 3 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Multi-Provider Caching in Multi-Tier Fog Networks
Ferdous Sharifi, Shaahin Hessabi, Young Choon Lee
ICFEC3
2025 Agnos-L2: Adaptive Layer 2 Framework for Heterogeneous Blockchain Ecosystems
Zhongli Dong, Young Choon Lee, Albert Y. Zomaya
PDCAT2
2025 Demo: P4 Based In-network ML with Federated Learning to Secure and Slice IoT Networks
abstract
Recent cyberattacks have increasingly targeted distributed networking environments like IoT networks. To detect these attacks, hidden under network traffic encryption, many centralized Machine Learning (ML) based solutions have been introduced, which are not well suited for IoT networks. This work proposes PIFL a practical approach to secure IoT networks by combining federated learning, in-network ML using P4-enabled devices, software-defined networks, and binarized neural networks. PIFL detects compromised edge devices and isolates them into separate network slices based on trust parameters derived from their behavior. We demonstrate the feasibility of PIFL using an experimental testbed with three intelligent network devices and seven IoT devices implemented on Raspberry Pi devices.
Chamara Manoj Madarasingha Kattadige, Thilini Dahanayaka, Kanchana Thilakarathna, Suranga Seneviratne, Young Choon Lee, Salil S. Kanhere, Albert Y. Zomaya, Aruna Seneviratne, Phil Ridley
WoWMoM5
2025 Poster: Unified Fog Node Utilization for Multiple Content Providers through Cluster-Based Cooperative Caching
abstract
Efficient content caching is crucial for video streaming services, enhancing user experience and conserving network bandwidth. While traditional Content Delivery Networks (CDNs) address this need to a certain extent, fog computing is emerging as a complementary solution. In particular, this new computing paradigm utilizes nodes positioned between users and the cloud continuum (fog nodes). However, the limited capacity of fog nodes poses a challenge. Previous studies have addressed this challenge by focusing on cooperative caching, considering factors like popularity and user location, yet often overlooked the shared use of fog node storage by multiple content providers (MCPs). This paper introduces CCoFog Caching (CCo-Fog), a cluster-based cooperative content caching strategy that allocates fog node storage capacity among content providers using user clustering and the multi-tier feature of fog networks. It also proposes a content placement algorithm that considers popularity to determine the number of content copies in the network. Evaluation with real-world data shows that CCo-Fog significantly improves latency by 54% and hit ratio by1|6% compared to existing strategies.
Ferdous Sharifi, Young Choon Lee, Shaahin Hessabi
WoWMoM2
2023 Mobility-Aware Fog Offloading
Ferdous Sharifi, Ali Rasaii, Melika Honarmand, Shaahin Hessabi, Young Choon Lee
APNOMS5
2023 Smart Contract Data Monitoring and Visualization
abstract
Blockchain technology has attracted significant industry, academic, and governmental attention since its emerged in 2008. Blockchain use cases are now being explored by traditional, transaction-oriented businesses in the finance, insurance, logistics and healthcare sectors to name a few. This has expanded further with the widespread use of Internet of Things (IoT) devices. Massive amounts of data are generated by IoT devices and are recorded in the blockchain. While blockchain provides many advantages, such as immutability and transparency, its serialized nature makes impossible to read in a single query. Multiple requests are required even for simple tasks, such as displaying an account's transaction history. This further leads to the difficulty in understanding the data in the blockchain. In this paper, we address the problem of smart contract visualization in a real-time manner. To this end, we design a visualization dashboard for smart contracts. A visual aid for massive amounts of data helps users understand the blockchain's overall activities, uncover operational risks and provide critical intelligence by visualising unusual activities and connections. Such insights may enable the user to investigate and predict any anomalies or reveal any network vulnerabilities. Cattle farm selected as a use case because the voluminous data can be acquired from IoT sensors on the farm cattle. Our dashboard has been proven to help visualize the life cycle of animals, the distribution of activities and time factor analysis. This visualization can give a user a better perspective of the token functions and results as well as animal management issues.
Seng Kuang Yap, Zhongli Dong, Mark Toohey, Young Choon Lee, Albert Y. Zomaya
ICBC4
2023 Catch the Intruder: Collaborative and Personalized Malware Detection By On-Device Application Fingerprinting
abstract
The vulnerability of smartphones to cyber attacks has been a serious concern to users arising from the integrity of installed applications (mobile apps). These apps are to provide legitimate and diversified on-the-go services. However, some have uncovered ways to penetrate smartphones for malicious behaviors. While some development and distribution regulations, such as Google Play Protect are in place, their effectiveness is often limited due primarily to falling behind the emergence of new malware. This paper presents an Analytic-based deep neural network Android Malware detection (ADAM) to detect potentially dangerous apps based on a set of features and patterns, i.e., application fingerprints, extracted from mobile apps. In particular, ADAM uses these features and patterns to train feature-specific DNNs to have consensus on the application labels when their ground truth is unknown. In addition, ADAM leverages the transfer learning technique to obtain its adjustability to new applications across smartphones. This is done by reusing the pre-trained model(s) and making them more adaptable by model personalization and federated learning (FL) techniques. This adjustability facilitates collaborative detection and independent labeling across smartphones, assisted by FL guards, which protect ADAM against poisoning attacks through model analysis. ADAM relies on a diverse dataset containing more than 153000 applications with over 41000 extracted features for DNNs training. ADAM’s feature-specific DNNs, on average, achieved more than 98% accuracy compared to Play Protect and antivirus software, resulting in an outstanding performance against data manipulation attacks.
Amirmohammad Pasdar, Young Choon Lee, Seok-Hee Hong 0001
ICWS2
2022 Train Me to Fight: Machine-Learning Based On-Device Malware Detection for Mobile Devices
abstract
Mobile applications (apps) on smartphones have become a primary means to bring a wide variety of services on the go. These apps are provided by third-party developers and service providers. These apps are increasingly diverse, so as are malware. As a result, current signature-based protection approaches are ineffective against new malware. This poses privacy and security risks, increasing smartphones' vulnerability to cyber attacks. In this paper, we present a novel Deep neural network-based On-device Malware Detection (DOM) that employs model personalization and transfer learning for enhancing real-time ondevice detection performance. DOM consists of two on-device machine learning models referred to as generic and personalized models and dynamically analyzes applications to extract a comprehensive set of features. The generic model is a fine-tuned deep neural network (DNN) for labeling applications whose ground truth is not available. In contrast, the personalized model is a lightweight trainable model created by retaining the majority of the generic DNN layers and trainable parameters and adding a new lightweight neural network. The personalized model is further improved with the help of federated learning, which aggregates the personalized model parameters. We have used over 32000 real-world applications from different repositories to train and evaluate DOM. Experiments show that the generic DNN model achieves 98.41% accuracy, and the personalized model has also demonstrated outstanding performance detection with an accuracy of 87%. DOM is very lightweight and uses less than 4% memory consumption.
Amirmohammad Pasdar, Young Choon Lee, Tongliang Liu, Seok-Hee Hong 0001
CCGRID2
2022 Power Control Framework for Green Data Centers
abstract
In recent years, renewable energy, such as wind and photovoltaic electric power has been increasingly integrated into data center power provisioning systems to address high energy consumption of data centers. However, in reality, the intermittency and randomness of renewable energy (power supply fluctuation) is detrimental to the reliable operation of sophisticated IT equipment in those so-called green data centers. In this article, we address the problem of data center power regulation explicitly taking into account the unreliability and instability of renewable energy sources. To this extent, we design a novel data center power control framework that smoothens the power fluctuation and instability of renewable energy sources. The core of our framework is two power regulation optimization algorithms. In particular, a server workload scheduling algorithm deals with high frequency fluctuations while an Uninterruptable Power Supply (UPS) power regulation algorithm handles low frequency and large extent power fluctuations. These algorithms are also designed to satisfy service level agreement (SLA) and standby power supply capacity. We have conducted an extensive evaluation study using trace data of a real data center of 30000-node cluster with 50 x 250 UPS battery groups and 24-hour power generation data from real wind farm and photovoltaic power station. The experimental results show our framework effectively smoothens fluctuations of data center power supply, more effective use of renewable energy, and extend the UPS batteries’ lives to reduce the skyrocketed data center operating expenses.
Ting Yang 0002, Yucheng Hou, Young Choon Lee, Albert Y. Zomaya
IEEE Trans. Cloud Comput.3
2021 iEdge: An IoT-assisted Edge Computing Framework
abstract
Edge computing has emerged as a viable solution to bridge the gap between distributed Internet of Things (IoT) devices and centralized distant clouds. In particular, small-scale servers are deployed at the edge of network (i.e., edge servers) to `help' cloud servers process data IoT devices constantly generate. However, these edge servers often struggle to deal with emerging applications that require real-time data processing in situ, such as real-time facial recognition. In this paper, we present iEdge as an IoT-assisted edge computing framework that enables the seamless execution of applications across an edge server and nearby IoT devices. The seamless execution in essence has been realized by transforming platform-dependent monolithic applications to cross-platform composite applications and offloading some tasks/functions of these composite applications to IoT devices considering device context. We have evaluated iEdge using a prototype implementation with a real-time facial recognition application. Experimental results show that iEdge effectively harnesses smart IoT devices as a consolidated edge computing execution environment and enables such an application to process more video streams than typical `edge-only' computing.
Hochul Lee, Seyul Lee, Young Choon Lee, Hyuck Han, Sooyong Kang
PerCom3
2021 Virtual reality: A survey of enabling technologies and its applications in IoT
Miao Hu 0001, Xianzhuo Luo, Young Choon Lee, Yipeng Zhou, Di Wu 0001
J. Netw. Comput. Appl.4
2020 The Power of ARM64 in Public Clouds
abstract
ARM processors, with their low power consumption and heat dissipation, have been highly successful in embedded systems. In the recent past, there have been attempts to adopt these energy-efficient processors for servers in data centers. However, a fundamental question remains open with ARM-based systems on server side is whether they are capable of handling compute-intensive workloads at scale. This paper gives our answer to this question with an empirical approach. We study the performance characteristics of the Amazon Graviton Processor - an ARM64 processor with the Cortex-A72 micro-architecture - using the A1 (Graviton) product family on AWS EC2, with comparisons to the I3 and M5 product families based on Intel Xeon processors. We use a combination of micro benchmark and performance counters to identify the lack of L3 cache and the slower memory access speed limit Graviton's capability in achieving higher performance. We confirm Graviton's capability in handling various large-scale horizontally scalable compute-intensive workloads, including multi-tier web service, video transcoding and terabyte scale sorting. In our large-scale evaluations, the test worker fleet has up to 1600 vCPU cores, which is by far the largest ARM64 cluster that has been reported. We observe that the A1 product family achieves the same price-performance in multi-tier web service, up to 37% cost saving in video transcoding, and up to 65% cost saving in terabyte scale sorting, as compared with the I3 and M5 product families.
Qingye Jiang, Young Choon Lee, Albert Y. Zomaya
CCGRID2
2020 EdgeSum: Edge-Based Video Summarization with Dash Cams
abstract
The following topics are dealt with: cloud computing; Internet of Things; learning (artificial intelligence); virtual machines; mobile computing; security of data; resource allocation; scheduling; software engineering; regression analysis.
Jayden King, Lily Huang, Di Wu 0001, Yipeng Zhou, Young Choon Lee
IC2E5
2020 ANN-Assisted Multi-cloud Scheduling Recommender
Amirmohammad Pasdar, Tahereh Hassanzadeh, Young Choon Lee, Bernard Mans
ICONIP (4)3
2020 Robust Scheduling for Large-Scale Distributed Systems
abstract
In large-scale distributed systems, such as clouds, failures are rather the norm than the exception. These failures include job failures, server failures, network outage and power failure. Among them, server failures are most common. With the wide adoption of cloud computing, the impact of server failures in clouds is far greater than that in traditional computer clusters as jobs of different tenants are often co-located (multi-tenancy). In this paper, we address the problem of robust scheduling, with realistic failure modeling, to minimize such impact on the execution of (co-located) jobs. To this end, we develop four online failure-aware (FA) scheduling algorithms, FAFF-WJ, FAFF-FC, FABF-WJ and FABF-FC, considering the availability and reliability of servers. In particular, FF (First-Fit) and BF (Best-Fit) indicate how the availability of servers is checked while WJ (Waiting Job) and FC (Failure Count) differ primarily in whether the reliability is measured from job's perspective or server's perspective. All four algorithms are designed essentially by combining these availability and reliability check methods. We evaluate our scheduling algorithms with failures generated based on our failure modeling of six real-world server failure traces. Our evaluation results show the effectiveness of our scheduling algorithms in robust job execution, with respect to both performance and cost.
Young Choon Lee, Jayden King, Young Ki Kim, Seok-Hee Hong 0001
TrustCom1
2020 Hybrid scheduling for scientific workflows on hybrid clouds
Amirmohammad Pasdar, Young Choon Lee, Khaled Almiani
Comput. Networks2
2020 Lightweight Power Monitoring Framework for Virtualized Computing Environments
abstract
The pervasive use of virtualization techniques in today's datacenters poses challenges in power monitoring since it is not possible to directly measure the power consumption of a virtual entity such as a virtual machine (VM) and a container. In this paper, we present cWatts++, a lightweight virtual power meter that enables accurate power usage measurement in virtualized computing environments such as VMs and containers of Cloud data centers. At the core of cWatts++ is its application-agnostic power model. To this end, we devise two power models (eventModel and raplModel) that are driven by CPU event counters and the Running Average Power Limit (RAPL) feature of modern Intel CPUs, respectively. While eventModel is more generic and, thus, applicable to a wide range of workloads, raplModel is particularly good for CPU-bound workloads. We have evaluated cWatts++ with its two power models in a real system using the PARSEC benchmark suite and our in-house benchmarks. Our evaluation study demonstrates that these power models have an average error of 4.55 and 1.25 percent, respectively, compared with actual power usage measurements of a real power meter, Cabac Power-Mate.
James Phung, Young Choon Lee, Albert Y. Zomaya
IEEE Trans. Computers2
2020 Automated Fine-Grained CPU Cap Control in Serverless Computing Platform
abstract
Serverless computing has emerged as a new cloud computing execution model that liberates users and application developers from explicitly managing `physical' resources, leaving such a resource management burden to service providers. In this article, we study the problem of resource allocation for multi-tenant serverless computing platforms explicitly taking into account workload fluctuations including sudden surges. In particular, we investigate different root causes of performance degradation in these platforms where tenants (their applications) have different workload characteristics. To this end, we develop a fine-grained CPU cap control solution as a resource manager that dynamically adjusts CPU usage limit (or CPU cap) concerning applications with same/similar performance requirements, i.e., application groups. The adjustment of CPU caps applies primarily to co-located worker processes of serverless computing platforms to minimize resource contention, which is the major source of performance degradation. The actual adjustment decisions are made based on performance metrics (e.g., throttled time and queue length) using a group-aware scheduling algorithm. The extensive experimental results performed in our local cluster confirm that the proposed resource manager can effectively eliminate the burden of explicit reservation of computing capacity, even when fluctuations and sudden surges in the incoming workload exist. We measure the robustness of the proposed resource manager by comparing it with several heuristics which extensively used in practice, including the enhanced version of round robin and the least length queue scheduling policies, under various workload intensities driven by real-world scenarios. Notably, our resource manager outperforms other heuristics by decreasing skewness and average response time up to 44 and 94 percent, respectively, while it does not over-use the CPU resources.
Young Ki Kim, M. Reza HoseinyFarahabady, Young Choon Lee, Albert Y. Zomaya
IEEE Trans. Parallel Distributed Syst.3
2019 DAGBENCH: A Performance Evaluation Framework for DAG Distributed Ledgers
abstract
Directed Acyclic Graph (DAG) has been emerging as the so-called Blockchain 3.0 after Bitcoin (Blockchain 1.0) and Ethereum (Blockchain 2.0). This new distributed ledger technology is getting significant attention for its high performance and low transaction fee. There have already been several notable implementations, such as IOTA [1], Nano [2] and Byteball [3]. In this paper, we present DAGBENCH as a performance evaluation framework for DAG implementations. DAGBENCH provides a number of sample workloads and adaptors that make effective and easy evaluation of different DAG implementations. It allows any DAG implementation to be evaluated by adding an adaptor. DAGBENCH allows to measure the performance of DAG implementation in terms of throughput, latency, scalability, success indicator, resource consumption, transaction data size and transaction fee. We demonstrate the efficacy of DAGBENCH with different DAG implementations. In particular, we have conducted experiments, on Amazon EC2, with three popular DAG implementations: IOTA, Nano and Byteball. Our experimental results provide the performance comparison between these implementations that helps developers/users effectively evaluate different performance characteristics; and, this enables them to identify bottlenecks and accordingly to improve performance.
Zhongli Dong, Emma Zheng, Young Choon Lee, Albert Y. Zomaya
CLOUD3
2019 Scalable Video Transcoding in Public Clouds
abstract
In this paper, we present the challenges involved in large-scale video transcoding application in public clouds. We introduce the architecture of an existing video transcoding system which is tightly coupled with an existing video sharing service. We examine the horizontal scalability of the video transcoding system on AWS EC2. With an online transaction processing (OLTP) model, the system achieves linear horizontal scalability up to 1,000 vCPU cores, but starts to experience performance degradation beyond that. We analyze the resource consumption pattern of the existing system, then introduce an improved architecture by adding a message queue layer. This effectively decouples the video transcoding system from the video sharing service and converts the OLTP model into a batch processing model. Large-scale evaluations on AWS EC2 indicate that the improved design maintains linear horizontal scalability at 10,100 vCPU cores. The hybrid design of the system allows it to be easily adapted for other batch processing use cases without the need to modify or recompile the application.
Qingye Jiang, Young Choon Lee, Albert Y. Zomaya
CCGRID2
2019 On the Trade-Off Between Performance and Storage Efficiency of Replication-Based Object Storage
abstract
The object storage systems are used to store and manage unstructured data. Most object storage systems provide the replication policy (REP) or erasure code policy (EC) to ensure the reliability and availability of data. In this paper, we study the trade-off between performance and storage efficiency of these policies with respect to different data sizes of user requests. To this end, we present a hybrid policy management system that takes advantage of both policies by automatically changing policy based on data size. We have implemented the hybrid system in OpenStack Swift. Our evaluation results show the throughput of GET request increases up to 36% while improving storage efficiency by up to 53% compared to that using the REP policy.
Hanbeom Jo, Hochul Lee, Young Choon Lee, Hyuck Han, Sooyong Kang
CloudCom4
2019 Visualisation of Distributed Systems Simulation Made Simple
abstract
Distributed (computing) systems come in various sizes and scale. They range from a single workstation computer with several processors, a cluster of compute nodes (servers) to a federation of geographically distributed data centres with millions of servers. Job scheduling is a fundamental aspect for data centre efficiency. In this paper, we present ds-viz as a visualisation aid for ds-sim, a recently developed distributed systems simulator. In particular, ds-viz significantly helps leverage the evaluation and analysis of scheduling algorithms that ds-sim facilitates to design. We show the effectiveness of these tools with some examples.
Jayden King, Young Ki Kim, Young Choon Lee, Seok-Hee Hong 0001
CloudCom3
2019 Toward Cost Efficient Cloud Bursting
Amirmohammad Pasdar, Young Choon Lee, Khaled Almiani
ICSOC2
2019 Holistic Approach for Studying Resource Failures at Scale
abstract
In large-scale distributed systems, such as data centers resource failures are the norm rather than an exception. In this paper, we propose a holistic approach to study resource failures from resource failure modelling to distributed system simulation to failure-aware scheduling algorithm design. In particular, we present (1) a simple and yet practical way to model resource failures using real-world failure traces, (2) a new distributed systems simulator and (3) two failure-aware scheduling algorithms. These scheduling algorithms are designed primarily to validate (1) and (2). Our evaluation results demonstrate the feasibility and effectiveness of our holistic approach.
Young Choon Lee, Jayden King, Seok-Hee Hong 0001
NCA1
2019 Mobile collaborative computing on the fly
Hochul Lee, Young Choon Lee, Hyuck Han, Sooyong Kang
Pervasive Mob. Comput.3
2019 CollaboRoid: Mobile platform support for collaborative applications
Hochul Lee, Young Choon Lee, Sooyong Kang
Pervasive Mob. Comput.3
2018 Decentralized Admission Control for High-Throughput Key-Value Data Stores
abstract
Workload surges are a serious hindrance to per-formance of even high-throughput key-value data stores, such as Cassandra, MongoDB, and more recently Aerospike. In this paper, we present a decentralized admission controller for high-throughput key-value data stores. The proposed controller dynamically regulates the release time of incoming requests explicitly taking into account different Quality of Service (QoS) classes. In particular, an instance of such controller is assigned to each client for its autonomous admission control specific to the client's QoS requirements. These controllers operate in a decentralized manner with only local performance metrics, response time and queue waiting time. Despite the use of such "minimal" run-time state information, our decentralized admission controller is capable of coping with workload surges respecting QoS requirements. The performance evaluation is carried out by comparing the proposed admission controller with the default scheduling policy of Aerospike, in a testbed cluster under various workload intensity rates. Experimental results confirm that the proposed controller improves QoS satisfaction in terms of end-to-end response time by nearly 12 times, on average, compared with that of Aerospike's, in high-rate workload. Results also show decreases of the average and standard deviation of latency up to 31% and 50%, respectively, during workload surges (peak load) in high-rate workload.
Young Ki Kim, M. Reza HoseinyFarahabady, Young Choon Lee, Albert Y. Zomaya
CCGrid3
2018 Janus: A Generic QoS Framework for Software-as-a-Service Applications
abstract
The move from the traditional Software-as-a-Product (SaaP) model to the Software-as-a-Service (SaaS) model is apparent with the wide adoption of cloud computing. Unlike the SaaP model, the SaaS model delivers a diverse set of software features directly from public clouds to a large number of arbitrary users with varying quality of service (QoS) requirements. QoS is typically assured by admission control. However, there are two outstanding issues with traditional QoS systems: (1) they are usually designed and developed with a special purpose, making them difficult to be reused for other use cases; and (2) they have limited scalability (i.e., vertical scalability) due to the write-intensive nature of admission control workload. In this paper, we present Janus - a QoS framework that is generic and scalable for SaaS applications taking full advantage of cloud's inherent horizontal scalability (scaling-out). Janus uses a multi-layer architecture to eliminate the communication between nodes (being scaled out) in the same layer achieving horizontal scalability without sacrificing vertical scalability. Janus ensures accurate admission control (QoS decisions) using a distributed set of leaky buckets with a refill mechanism. Janus also adopts a key-value request-response mechanism for easy integration with the actual application. We extensively evaluate Janus on AWS cloud with both Apache HTTP server benchmarking tool and a photo sharing web application. Our experimental results demonstrate that (a) Janus achieves linear scalability both vertically and horizontally, and (b) Janus can be integrated with existing applications with a minimum amount of code change. In particular, Janus achieves more than 100,000 requests per second with only 10 nodes (4 vCPU cores on each node) in the QoS server layer and 90% of the admission control decisions were made in 3 milliseconds.
Qingye Jiang, Young Choon Lee, Albert Y. Zomaya
CLUSTER2
2018 Dynamic Control of CPU Usage in a Lambda Platform
abstract
Lambda platform is a new concept based on an event-driven server-less computation that empowers application developers to build scalable enterprise software in a virtualized environment without provisioning or managing any physical servers (a server-less solution). In reality, however, devising an effective consolidation method to host multiple Lambda functions into a single machine is challenging. The existing simple resource allocation algorithms, such as the round-robin policy used in many commercial server-less systems, suffer from lack of responsiveness to a sudden surge in the incoming workload. This will result in an unsatisfactory performance degradation that is directly experienced by the end-user of a Lambda application. In this paper, we address the problem of CPU cap management in a Lambda platform for ensuring different QoS enforcement levels in a platform with shared resources, in case of fluctuations and sudden surges in the incoming workload requests. To this end, we present a closed-loop (feedback-based) CPU cap controller, which fulfills the QoS levels enforced by the application owners. The controller adjusts the number of working threads per QoS class and dispatches the outstanding Lambda functions along with the associated events to the most appropriate working thread. The proposed solution reduces the QoS violations by an average of 6.36 times compared to the round-robin policy. It can also maintain the end-to-end response time of applications belonging to the highest priority QoS class close to the target set-point while decreasing the overall response time by up to 52%.
Young Ki Kim, M. Reza HoseinyFarahabady, Young Choon Lee, Albert Y. Zomaya, Raja Jurdak
CLUSTER3
2018 Hierarchical Recursive Resource Sharing for Containerized Applications
Youngjin Kim 0011, Young Choon Lee, Hyuck Han, Sooyong Kang
ICSOC2
2018 Modeling System-Level Power Consumption Profiles Using RAPL
abstract
Efficient power management is essential in ensuring the economic viability of large-scale distributed systems. There is growing interest in applying the Running Average Power Limit (RAPL) feature that is commonly found in today's Intel CPUs to energy monitoring and efficiency applications. We investigate the power consumption characteristics of two different CPUs using the RAPL feature. We present a prototype lightweight software-based virtual power meter that exploits this functionality. Utilizing a simple but very effective application-agnostic power model, it offers comparable or superior performance to existing power models that are more complex. It is portable across a variety of systems. It can be used in containerized or virtualized environments. We demonstrate that our power model has an average error of 1.63 %. This result compares favorably with existing state-of-the-art power models and is achieved using a simple power model. Consequently, our power meter is viable for use in real-world applications such as power estimation for energy-aware scheduling.
James Phung, Young Choon Lee, Albert Y. Zomaya
NCA2
2018 SAMD: Fine-Grained Application Sharing for Mobile Collaboration
abstract
The collective use of ever connected and pervasive mobile devices has been increasingly sought for in mobile collaboration, such as multiplayer mobile gaming and distributed processing. The current model of mobile collaboration requires each device to install a particular, `full' mobile app for a respective collaboration. Besides, collaboration functionalities are typically implemented at application level. In this paper, we present Single Application Multiple Device (SAMD) as a platform-level mobile collaboration framework. A mobile app developed using SAMD is capable of fine-grained application sharing. In particular, SAMD enables devices, agreed to participate in collaboration, to get portions of the app on-the-fly and run them without the prior installation. To achieve this, we have developed three solutions as core functionalities of SAMD: 1) Controller packaging, 2) lookahead transfer and 3) code adaptation. We have implemented SAMD on Android as a proof-of-concept prototype. Our experimental results demonstrate SAMD can provide fine-grained sharing of latency-insensitive applications.
Hochul Lee, Byoungjun Seo, Young Choon Lee, Hyuck Han, Sooyong Kang
PerCom4
2018 On efficient resource use for scientific workflows in clouds
Khaled Almiani, Young Choon Lee, Bernard Mans
Comput. Networks2
2018 Collective Energy-Efficiency Approach to Data Center Networks Planning
abstract
Energy efficiency of data centers (DCs) has become a major concern as DCs continue to grow large often-Energy efficiency of data centers (DCs) has become a major concern as DCs continue to grow large often hosting tens of thousands of servers or even hundreds of thousands of them. Clearly, such a volume of DCs implies scale of data center network (DCN) with a huge number of network nodes and links. The energy consumption of this communication network has skyrocketed and become the same league as computing servers' costs. With the ever-increasing amount of data that need to be stored and processed in DCs, DCN traffic continues to soar drawing increasingly more power. In particular, more than one-third of the total energy in DCs is consumed by communication links, switching and aggregation elements. In this paper, we concern the energy efficiency of data center explicitly taking into account both servers and DCN. To this end, we present VPTCA, as a collective energy-efficiency approach to data center network planning, which deals with virtual machine (VM) placement and communication traffic configuration. VPTCA aims particularly to reduce the energy consumption of DCN by assigning interrelated VMs into the same server or pod, which effectively helps reduce the amount of transmission load. In the layer of traffic message, VPTCA optimally uses switch ports and link bandwidth to balance the load and avoid congestions, enabling DCN to increase its transmission capacity, and saving a significant amount of network energy. In our evaluation via NS-2 simulations, the performance of VPTCA is measured and compared with two well-known DCN management algorithms, Global First Fit and ElasticTree. Based on our experimental results, VPTCA outperforms existing algorithms in providing DCN more transmission capacity with less energy consumption.
Ting Yang 0002, Young Choon Lee, Albert Y. Zomaya
IEEE Trans. Cloud Comput.2
2017 Cloud Bursting Scheduler for Cost Efficiency
abstract
Clouds have been increasingly adopted due to primarily their elasticity and pay-as-you-go (PAYG) pricing. While many organizations outsource the entire ICT solution to public clouds like Amazon Web Services Elastic Compute Cloud (EC2), others consider occasional workload offloading (cloud bursting) due to various reasons including governance and security. In this paper, we present Cloud Bursting Scheduler (CBS), a new cloud bursting algorithm. CBS explicitly takes into account cost factors of private in-house system (or private cloud) and public cloud. In particular, CBS attempts to optimize the cost to performance ratio by offloading jobs to public cloud explicitly taking into account time-varying electricity rates with private clouds and the timeinvariant rental rate of many public clouds. Based on simulation results obtained using real workload traces, CBS saves costs of running workloads by 55% and 12% compared with costs of cloud sourcing and private cloud, respectively. It also improves resource utilization (to 89%) by judiciously (de)activating inhouse resources and dynamically provisioning cloud resources.
Young Choon Lee, Bing Lian
CLOUD1
2017 Platform Support for Mobile Edge Computing
abstract
Computing resources including mobile devices at the edge of a network are increasingly connected and capable of collaboratively processing what's believed to be too complex to them. Collaboration possibilities with today's feature-rich mobile devices go far beyond simple media content sharing, traditional video conferencing and cloud-based software as a services. The realization of these possibilities for mobile edge computing (MEC) requires non-trivial amounts of efforts in enabling multi-device resource sharing. The current practice of mobile collaborative application development remains largely at the application level. In this paper, we present CollaboRoid, a platform-level solution that provides a set of system services for mobile collaboration. CollaboRoid's platform-level design significantly eases the development of mobile collaborative applications promoting MEC. In particular, it abstracts the sharing of not only hardware resources, but also software resources and multimedia contents between multiple heterogeneous mobile devices. We implement CollaboRoid in the application framework layer of the Android stack and evaluate it with several collaboration scenarios on Nexus 5 and 7 devices. Our experimental results show the feasibility of the platform-level collaboration using CollaboRoid in terms of the latency and energy consumption.
Hochul Lee, Young Choon Lee, Hyuck Han, Sooyong Kang
CLOUD3
2017 CtrlCloud: Performance-Aware Adaptive Control for Shared Resources in Clouds
abstract
Consolidating applications of conflicting service level objectives (SLOs) to share virtualized resources in cloud datacenters requires efficient resource management to ensure overall high Quality-of-Service (QoS). Applications of different performance targets often exhibit different resource demands. Thus, it is not trivial to translate individual application SLOs to corresponding resource shares in a shared virtualized environment to meet performance targets. In this paper, we present CtrlCloud, a performance-aware resource controlling system, that adaptively allocates resources, with a resource-share controller and an allocation optimization model. The controller automatically adapts resource demands based on performance deviations, while the optimization model resolves conflicts in resource demands from multiple co-located applications based on their ongoing performance achieved. We implement a proof-of-concept prototype of CtrlCloud in Python on top of Xen hypervisor. Our experimental results indicate that CtrlCloud can optimize allocations of CPU resources across multiple applications to maintain the 95th percentile latency within predefined SLO targets. CtrlCloud also provides QoS differentiation and yet fulfilling of CPU share demands from applications is maximized given resource availability. We further compare CtrlCloud against two other resource allocation methods commonly used in current clouds. CtrlCloud improves resource utilization by allocating resource shares optimal to 'actual needs' as it employs share-performance online modeling.
Omer Y. Adam, Young Choon Lee, Albert Y. Zomaya
CCGrid2
2017 Application-Agnostic Power Monitoring in Virtualized Environments
abstract
Many servers use technologies such as virtualization or containerization to improve server utilization. These technologies pose challenges for power monitoring since it is not possible to directly measure the power use of an abstraction such as a virtual machine. Much work has been done in modeling the power use of CPUs, virtual machines and entire servers, however, there is a scarcity of work in building lightweight power monitoring middleware that can be deployed across a range of systems. In this paper, we present cWatts+ as a prototype lightweight software-based virtual power meter. Utilizing a simple but powerful application-agnostic power model, it offers comparable performance to existing "more complex and heavier-weight" power models. It uses a small number of widely available CPU event counters and the Performance Application Programming Interface Library to estimate power usage on a per-thread basis. It has minimal overhead and is portable across a variety of systems. It can be used in containerized or virtualized environments. We evaluate the estimation performance of cWatts+ for a variety of real-world benchmarks that are relevant to large distributed systems. Also, we examine the importance of including CPU core temperature data in the power model. We demonstrate that our power model has an average error of less than 5%. This result compares favorably with existing state-of-the-art power models and is achieved using a relatively simple power model that exhibits minimal power consumption (overhead). Consequently, our power monitoring middleware is viable for use in real-world applications such as power estimation for energy-aware scheduling.
James Phung, Young Choon Lee, Albert Y. Zomaya
CCGrid2
2017 A QoS-Aware Resource Allocation Controller for Function as a Service (FaaS) Platform
M. Reza HoseinyFarahabady, Young Choon Lee, Albert Y. Zomaya, Zahir Tari
ICSOC2
2017 Serverless Execution of Scientific Workflows
Qingye Jiang, Young Choon Lee, Albert Y. Zomaya
ICSOC2
2017 Resource demand aware scheduling for workflows in clouds
abstract
A major challenge of running applications in clouds is to determine the right number of resources (virtual machines or VMs) to rent in terms of both performance and cost. Such a challenge becomes greater if the application requires to run across multiple resources. In this paper, we address the problem of scheduling scientific workflow applications. The structure of workflows, dictated by precedence/data dependencies, and the diversity of resources in clouds both at large scale make the resource provisioning and task scheduling very complex. To this end, we design the Resource Demand Aware Scheduling (RDAS) algorithm that schedules workflows based on their resource demands and priorities considering workflow structure. RDAS partitions workflows and allocates resources of possibly different capacities/types to the partitions in a “fair” manner such that their execution times do not vary significantly. RDAS turns resource and application heterogeneity (a major hindering factor in clouds) into an opportunity for optimizing resource provisioning for scientific workflows. Based on our experimental results, RDAS demonstrates its capacity of minimizing the overall workflow completion time (makespan) and in turn minimizing costs of the execution. In particular, RDAS outperforms three existing algorithms by 22%, 13% and 33%, on average, in terms of makespan, cost and the number of resources used, respectively.
Khaled Almiani, Young Choon Lee, Bernard Mans
NCA2
2017 adCFS: Adaptive completely fair scheduling policy for containerised workflows systems
abstract
Scientific workflows are increasingly containerised, which requires rethinking central processing unit (CPU) sharing policies to accommodate different workload types. However, container engines running scientific workflows struggle to share the CPU fairly, as workload characteristics are not taken into account. This paper proposes a sharing policy called the Adaptive Completely Fair Scheduling policy (adCFS), which considers the future state of CPU usage and proactively shares CPU cycles between various containers based on their corresponding workload metrics (e.g., CPU usage, task runtime, #tasks). adCFS estimates the weight of workload characteristics and redistributes the CPU based on the corresponding weights. The Markov chain model is used to predict CPU state use, and the adCFS policy is triggered to dynamically allocate containers to the proper CPU portions. Experimental results show enhanced container CPU response time for those containers that run heavy and large jobs: these display 12% faster response time compared with the default CFS (Completely Fair Scheduler). adCFS therefore enhances CFS by considering workload metrics, which leads to the CPU being shared fairly when it is fully used.
Eidah J. Alzahrani, Zahir Tari, Young Choon Lee, Deafallah Alsadie, Albert Y. Zomaya
NCA3
2017 A resource allocation controller for key-value data stores
abstract
Recent distributed key-value data stores, such as Aerospike are getting the momentum with ever-increasing need for large-scale real-time data processing. While these data stores can provide significantly improved performance, they still struggle to meet Quality of Service (QoS) during workload surges. In this paper, we address the problem of QoS-aware resource allocation for burst workloads in key-value data stores. To this end, we design a resource allocation controller, which enables each application to independently regulate the releases of its requests taking into account QoS. In particular, the proposed controller monitors the actual performance metrics of the target system and dynamically releases requests from a buffer owned by each application accordingly. We have implemented the proposed controller in an Aerospike cluster for our performance evaluation. Experiments have been conducted with various workload intensities (with up to 36,180 write operations per second) in comparison with the default Aerospike policy. Experimental results confirm that the proposed controller decreases the overall average latency up to 41% on high-rate workload while maintaining the QoS of high priority applications.
Young Ki Kim, M. Reza HoseinyFarahabady, Young Choon Lee, Albert Y. Zomaya
NCA3
2017 Stochastic Resource Provisioning for Containerized Multi-Tier Web Services in Clouds
abstract
Under today's bursty web traffic, the fine-grained per-container control promises more efficient resource provisioning for web services and better resource utilization in cloud datacenters. In this paper, we present Two-stage Stochastic Programming Resource Allocator (2SPRA). It optimizes resource provisioning for containerized n-tier web services in accordance with fluctuations of incoming workload to accommodate predefined SLOs on response latency. In particular, 2SPRA is capable of minimizing resource over-provisioning by addressing dynamics of web traffic as workload uncertainty in a native stochastic optimization model. Using special-purpose OpenOpt optimization framework, we fully implement 2SPRA in Python and evaluate it against three other existing allocation schemes, in a Docker-based CoreOS Linux VMs on Amazon EC2. We generate workloads based on four real-world web traces of various traffic variations: AOL, WorldCup98, ClarkNet, and NASA. Our experimental results demonstrate that 2SPRA achieves the minimum resource over-provisioning outperforming other schemes. In particular, 2SPRA allocates only 6.16 percent more than application's actual demand on average and at most 7.75 percent in the worst case. It achieves 3x further reduction in total resources provisioned compared to other schemes delivering overall cost-savings of 53.6 percent on average and up to 66.8 percent. Furthermore, 2SPRA demonstrates consistency in its provisioning decisions and robust responsiveness against workload fluctuations.
Omer Y. Adam, Young Choon Lee, Albert Y. Zomaya
IEEE Trans. Parallel Distributed Syst.2
2016 Partitioning-Based Workflow Scheduling in Clouds
abstract
Many applications in science and engineering become increasingly complex and large scale. These applications often consist of a large number of precedence-constrained tasks forming workflows represented by directed acyclic graph (DAG). In recent years, cloud computing has greatly leveraged the elastic and cost-efficient deployment of these applications. However, their effective deployment is largely dependent on the scheduling algorithm adopted. Most existing workflow scheduling algorithms are designed to optimize deadline or budget/cost, i.e., one being the objective and the other being constraint. In this paper, we present the Partitioning-Based Workflow Scheduling (PBWS) algorithm, which liberates the user from explicitly setting the upper bound of deadline and cost. Instead, PBWS adopts a slack parameter that controls the tradeoff point between deadline and cost. In particular, PBWS partitions a workflow into a number of small task graphs (or simply partitions) for which the granularity of such partitions is determined by the slack parameter. Each of these partitions is then matched with the best performing cloud resource in terms of both the overall execution time (makespan) and cost. The size of partitions may change by rearranging tasks between different partitions for the optimization of resource assignment. Our experimental results show that our PBWSalgorithm outperforms two existing algorithms in terms of cost by a large margin with little overhead on makespan.
Khaled Almiani, Young Choon Lee
AINA2
2016 A Model Predictive Controller for Contention-Aware Resource Allocation in Virtualized Data Centers
abstract
Data center efficiency is primarily sought by sharing physical resources, such as processors, memory, and disks in the form of virtual machines or containers among multiple users, i.e., workload consolidation. However, the reality is co-located applications in these virtual platforms compete for resources and interfere with each others' performance, resulting in performance variability/degradation. In this paper, we present the contentionaware resource allocation (CARA) solution, which optimizes data center efficiency. It is essentially devised based on a model predictive control that enables to make judicious consolidation decisions with future system states. CARA consolidates workloads explicitly taking into account the correlation between shared and isolated resource usage patterns. Based on our experimental results, CARA improves the overall resource utilization by 32%, without a significant impact on the quality-of-service (QoS) enforcement level. Such improvement results in a fewer number of active servers and in turn contributes to an overall energy saving by 33%.
M. Reza HoseinyFarahabady, Young Choon Lee, Albert Y. Zomaya, Zahir Tari, Andy Song
MASCOTS2
2016 Constructing Performance-Predictable Clusters with Performance-Varying Resources of Clouds
abstract
Although most current cloud providers, such as Amazon Web Services (AWS) and Microsoft Azure offer different types of computing instances with different capacities, cloud users tend to hire a cluster of instances of particular type to ensure performance predictability for their applications. Nowadays, many large-scale applications including big data analytics applications feature workload patterns that have heterogeneous resource demands, for which, accounting for heterogeneity of virtual cloud instances to allocate would be highly advantageous to the application performance. However, performance predictability has been always an issue in such clusters of heterogeneous resources. In particular, to precisely decide on what instances from which types to enclose in a cluster, such that the desired performance is attained, remains an open question. To this end, we devise a resource allocation mechanism by formulating it as a Mixed-Integer programming model representing an optimization problem. Our resource allocation mechanism incorporates predictable average performance as a unified performance metric, which concerns two key performance-related issues: (a) Performance variation within same-type instances, and (b) Correlations of performance variabilities across different types. Our experimental results demonstrate that target performance is predictable and attainable for clusters of heterogeneous resources. Our mechanism constructs clusters whose performance is within 95 percent of the performance of optimal ones, hence deadlines are always met. By reoptimisation, our mechanism can react to performance mispredictions and support autoscaling for varying workloads. We experimentally verify our findings using clusters on Amazon EC2 with MapReduce workloads, and on a private cloud as well. We conduct comparison experiments with an existing recent resource allocation approach in literature.
Omer Y. Adam, Young Choon Lee, Albert Y. Zomaya
IEEE Trans. Computers2
2015 Effective resource multiplexing for scientific workflows
abstract
Scientific workflows feature complex precedence constraints that are mostly dictated by data dependencies between tasks. The inter-task communication (data staging) in these complex workflow applications incurs significant overheads resulting in a major hindering factor of high performance and effective resource utilization. As the scale of these applications becomes increasingly large due primarily to the recent explosive growth of data, addressing this hindrance is of great practical importance. In this paper, we present a resource multiplexing (RM) technique, which leverages data staging aiming to minimize idle times between execution of tasks due to inter-task communication overheads. In particular, we incorporate RM into our DEWE framework1with making a set of extensions to the framework. The rationale behind RM is each slot or core pairs up the actual workflow task and the RM-enabled file loading DEWE extension (File Client) in the way that their resource usage is complementary. We demonstrate the efficacy of our multiplexing technique in a data-intensive computing environment using an astronomy application. Our results from experiments conducted in Amazon EC2 demonstrate that our multiplexing technique is effective with the reduction in resource idle time between jobs by 57% on average and up to 91%.
Thomas Ryan, Young Choon Lee
APNOMS2
2015 Crowdware: A Framework for GPU-Based Public-Resource Computing with Energy-Aware Incentive Mechanism
abstract
The power of the crowd, more precisely crowdsourced resources, is in its ubiquity. Accounting for traditional desktop/laptop computers and recent mobile computing devices including tablets and smart phones far surpasses the number of servers in cloud data centers. Besides, the capacity and capability of these resources owned by the crowd (crowd-sourced resources) has increased dramatically with GPUs in particular. Although a myriad of public-resource (or volunteer) computing projects, including SETI@home and Milkyway@home, have attracted the participation of crowd-sourced resources at very large scale, the sustainability of such participation is in doubt due primarily to ever-increasing energy costs. In this paper, we present Crowdware, a framework for enabling sustainable GPU-based public-resource computing with a realistic financial incentive mechanism. To this end, we design an auction-based resource allocation algorithm and a profit-based resource participation algorithm, explicitly considering the electricity cost of participating resources. Our results show that Crowdware greatly promotes profitability and cost efficiency for resource providers and resource consumers, respectively. Specifically, Crowdware has enabled the execution of MD5 password recovery jobs, in our testbed, with only 2.2% of the cost of using Amazon EC2 GPU instances while the participation of crowd-sourced resources is profitable with an average profit rate of 9.2%. Crowdware also shows great scalability with its fat-client and thin-server design. Together, Crowdware significantly improves the sustainability of public-resource computing.
Zhongli Dong, Young Choon Lee, Albert Y. Zomaya
CloudCom2
2015 Executing Large Scale Scientific Workflow Ensembles in Public Clouds
abstract
Scientists in different fields, such as high energy physics, earth science, and astronomy are developing large-scale workflow applications. In many use cases, scientists need to run a set of interrelated but independent workflows (i.e., Workflow ensembles) for the entire scientific analysis. As a workflow ensemble usually contains many sub-workflows in each of which hundreds or thousands of jobs exist with precedence constraints, the execution of such a workflow ensemble makes a great concern with cost even using elastic and pay-as-you-go cloud resources. In this paper, we address two main challenges in executing large-scale workflow ensembles in public clouds with both cost and deadline constraints: (1) execution coordination, and (2) resource provisioning. To this end, we develop a new pulling based workflow execution system with a profiling-based resource provisioning strategy. The idea is homogeneity in both scientific workflows and cloud resources can be exploited to remove scheduling overhead (in execution coordination) and to minimize cost meeting deadline. Our results show that our solution system can achieve 80% speed-up, by removing scheduling overhead, compared to the well-known Pegasus workflow management system when running scientific workflow ensembles. Besides, our evaluation using Montage (an astronomical image mosaic engine) workflow ensembles on around 1000-core Amazon EC2 clusters has demonstrated the efficacy of our resource provisioning strategy in terms of cost effectiveness within deadline.
Qingye Jiang, Young Choon Lee, Albert Y. Zomaya
ICPP2
2015 Resource-efficient workflow scheduling in clouds
Young Choon Lee, Hyuck Han, Albert Y. Zomaya, Mazin Yousif
Knowl. Based Syst.1
2015 Adaptive multiple-workflow scheduling with task rearrangement
Young Choon Lee, Alan D. Fekete, Albert Y. Zomaya
J. Supercomput.2
2015 RAMP: reliability-aware elastic instance provisioning for profit maximization
Luke M. Leslie, Young Choon Lee, Albert Y. Zomaya
J. Supercomput.2
2014 Local Resource Shaper for MapReduce
abstract
Resource capacity is often over provisioned to primarily deal with short periods of peak load. Shaping these peaks by shifting them to low utilization periods (valleys) is referred to as "resource consumption shaping". While originally aimed at the data center level, the resource consumption shaping we consider focuses on local resources, like CPU or I/O as we have identified that individual jobs also incur load peaks and valleys on these resources. In this paper, we present Local Resource Shaper (LRS), which limits fairness in resource sharing between co-located MapReduce tasks. LRS enables Hadoop to maximize resource utilization and minimize resource contention independently of job type. Co-located MapReduce tasks are often prone to resource contention (i.e., Load peak) due to similar resource usage patterns particularly with traditional fair resource sharing. In essence, LRS differentiates co-located tasks through active and passive slots that serve as containers for interchangeable map or reduce tasks. LRS lets an active slot consume as much resources as possible, and a passive slot make use of any unused resources. LRS leverages such slot differentiation with its new scheduler, Interleave. Our results show that LRS always outperforms the best static slot configuration with three Hadoop schedulers in terms of both resource utilization and performance.
Peng Lu 0004, Young Choon Lee, Vincent Gramoli, Luke M. Leslie, Albert Y. Zomaya
CloudCom2
2014 Energy-Efficient Data Center Networks Planning with Virtual Machine Placement and Traffic Configuration
abstract
Data Center (DC), the underlying infrastructure of cloud computing, becomes startling large with more powerful computing and communication capability to satisfy the wide spectrum of composite applications. In a large scale DC, a great number of switches connect servers into one complex network. The energy consumption of this communication network has skyrocketed and become the same league as the computing servers' costs. More than one-third of the total energy in DCs is consumed by communication links, switching and aggregation elements. Saving Data Center Network (DCN) energy to improve data center efficiency (power usage effectiveness or PUE) become the key technique in green computing. In this paper, we present VPTCA as an energy-efficient data center network planning solution that collectively deals with virtual machine placement and communication traffic configuration. VPTCA aims to reduce the DCN's energy consumption. In particular, interrelated VMs are assigned into the same server or pod, which effectively helps to reduce the amount of transmission load. In the layer of traffic message, VPTCA optimally uses switch ports and link bandwidth to balance the load and avoid congestions, enabling DCN to increase its transmission capacity, and saving a significant amount of network energy. In our evaluation via NS-2 simulations, the performance of VPTCA is measured and compared with two well-known DCN management algorithms, Global First Fit and Elastic Tree. Based on our experimental results, VPTCA outperforms existing algorithms in providing DCN more transmission capacity with less energy consumption.
Ting Yang 0002, Young Choon Lee, Albert Y. Zomaya
CloudCom2
2014 Running Data-Intensive Scientific Workflows in the Cloud
abstract
The scale of scientific applications becomes increasingly large not only in computation, but also in data. Many of these applications also concern inter-related tasks with data dependencies, hence, they are scientific workflows. The efficient coordination of executing/running scientific workflows is of great practical importance. The core of such coordination is scheduling and resource allocation. In this paper, we present three scheduling heuristics for running large-scale, data-intensive scientific workflows in clouds. In particular, the three heuristic algorithms are designed to leverage slot queue threshold, data locality and data prefetching, respectively. We also demonstrate how these heuristics can be collectively used to tackle different issues in running "data-intensive" workflows in clouds although each of these heuristics can be used independently. The practicality of our algorithms has been realized by actually implementing and incorporating them into our workflow execution system (DEWE). Using Montage, an astronomical image mosaic engine, as an example workflow, and Amazon EC2 as the cloud environment, we evaluate the performance of our heuristics in terms primarily of completion time (make span). We also scrutinize workflow execution showing different execution phases to identify their impact on performance. Our algorithms scale well and reduce make span by up to 27%.
Chiaki Sato, Luke M. Leslie, Young Choon Lee, Albert Y. Zomaya, Rajiv Ranjan 0001
PDCAT3
2014 Efficient allocation of resources in multiple heterogeneous Wireless Sensor Networks
Wei Li 0058, Flávia Coimbra Delicato, Paulo F. Pires, Young Choon Lee, Albert Y. Zomaya, Claudio Miceli, Luci Pirmez
J. Parallel Distributed Comput.4
2014 Randomized approximation scheme for resource allocation in hybrid-cloud environment
M. Reza HoseinyFarahabady, Young Choon Lee, Albert Y. Zomaya
J. Supercomput.2
2014 Pareto-Optimal Cloud Bursting
abstract
Large-scale Bag-of-Tasks (BoT) applications are characterized by their massively parallel, yet independent operations. The use of resources in public clouds to dynamically expand the capacity of a private computer system might be an appealing alternative to cope with such massive parallelism. To fully realize the benefit of this ‘cloud bursting’, the performance to cost ratio (or cost efficiency) must be thoroughly studied and incorporated into scheduling and resource allocation strategies. In this paper, we present PANDA, a framework for static scheduling BoT applications across resources in both private and public clouds. The framework at the core incorporates a fully polynomial-time approximation scheme (FPTAS) as a novel scheduling algorithm, which generates schedules with the best trade-off point between cost and performance; hence Pareto-optimality. We have theoretically discussed the complexity and correctness of our algorithms, and experimentally verified their efficacy and practicality using ISOMAP—a widely-used nonlinear manifold method as a real-world BoT application. Our evaluation conducted in a 'multi-cloud' environment of our 40-core private system and Amazon EC2 public cloud demonstrates the scheduling quality of PANDA is guaranteed to be within a measurable distance from the optimal solution. Results obtained from our experiments show such quality is 8 percent or less from the optimum. We also show the sensitivity and robustness of our scheduling solutions against performance errors in both resources and applications.
M. Reza HoseinyFarahabady, Young Choon Lee, Albert Y. Zomaya
IEEE Trans. Parallel Distributed Syst.2
2013 Exploiting Performance and Cost Diversity in the Cloud
abstract
Infrastructure-as-a-Service (IaaS) platforms, such as Amazon EC2, allow clients access to massive computational power in the form of virtual machines (VMs) known as instances. Amazon hosts three different instance purchasing options, each with its own service level agreement covering availability and pricing. In addition, Amazon offers access to a number of geographical regions, zones, and instance types from which to select. In this paper, we present a resource allocation and job scheduling framework (RAMC-DC), which utilizes Amazon's rich selection of service offerings---particularly within Spot and On-Demand instance purchasing options---aiming to cost efficiently execute deadline-constrained jobs. The framework is capable of ensuring quality of service in terms of cost, deadline compliance and service reliability. Such capacities are realized incorporating a set of novel strategies including execution time and cost approximation, bidding and resource allocation strategies. To the best of our knowledge, RAMC-DC most extensively exploits the service diversity of Amazon EC2, and offers a comprehensive cost efficiency solution that is able to deliver both the performance and reliability of On-Demand instances and the low costs of Spot instances. Experimental results obtained from extensive simulations using Amazon's Spot price traces show that our approach keeps deadline breaches and early-termination rates as low as 0.47% and 0.18%, respectively. This reliable performance is achieved with total costs between 13% and 20% of an equivalent approach using only On-Demand instances.
Luke M. Leslie, Young Choon Lee, Peng Lu 0004, Albert Y. Zomaya
IEEE CLOUD2
2013 Stretch Out and Compact: Workflow Scheduling with Resource Abundance
abstract
Resource abundance is apparent in today's multi-core era. Workflow applications common in science and engineering can take great advantage of such ample resource capacity. Yet, existing workflow scheduling algorithms have been mostly designed on the traditional and contrasting premise of scarce resource capacity relative to what applications require. In this paper, we address the problem of workflow scheduling exploiting resource abundance not only for performance, but also for resource efficiency. We first present the critical-path-first scheduling algorithm, which efficiently stretches out the schedule to proactively preserve critical path length, the shortest possible time of completion. We then develop an algorithm to compact the schedule for resource efficiency. The schedule is compacted by rearranging tasks making use of idle/inefficiency slots present in the schedule due to precedence constraints (synchronization). We have compared the performance of our scheduling algorithm with three previous algorithms and evaluated the efficacy of our schedule compaction algorithm by applying it to those four scheduling algorithms including our own. Experiments were conducted in a simulated environment with various real-world scientific workflows. Results show that our scheduling algorithm achieves a great performance to resource usage ratio. Further, our schedule compaction algorithm reduces resource usage by 33% on average ranging from 11.3% to 90.7% with no make span increases.
Young Choon Lee, Albert Y. Zomaya
CCGRID1
2013 Non-intrusive Slot Layering in Hadoop
abstract
Hadoop, an open source implementation of MapReduce, uses slots to represent resource sharing. The number of slots in a Hadoop cluster node specifies the concurrency of task execution. Thus, the slot configuration has a significant impact on performance. The number of slots is by default hand-configured (static) and slots share resources "fairly". As resource capacity (e.g., #cores) continues to increase and application dynamics becomes increasingly diverse, the current practices of static slot configuration and fair resource sharing may not efficiently utilize resources. Besides, such fair sharing is against priority-based scheduling when high priority jobs are sharing resource with lower priority jobs. In this paper we study the optimization of resource utilization in Hadoop focusing on those two issues of current practices and present a non-intrusive slot layering solution. Our solution approach in essence uses two tiers of slot (Active and Passive) to increase the degree of concurrency with minimal performance interference between them. Tasks in the Passive slots proceed their execution when tasks in the Active slots are not fully using (CPU) resource, and tasks/slots in these tiers are dynamically and adaptively managed. To leverage the effectiveness of slot layering, we develop a layering-aware task scheduler. Our non-intrusive slot layering approach is unique in that (1) it is a generic way to manage resource sharing for parallel and distributed computing models (e.g., MPI and cloud computing) and (2) both overall throughput and high-priority job performance are improved. Our experimental results with 6 representative jobs show 3%-34% improvement in overall throughput and 13%-48% decrease in the executing time of high-priority jobs compared with static configurations.
Peng Lu 0004, Young Choon Lee, Albert Y. Zomaya
CCGRID2
2013 Handling Uncertainty: Pareto-Efficient BoT Scheduling on Hybrid Clouds
abstract
Coping with uncertainty is a challenging and complex problem particularly in hybrid cloud environments-private cloud plus public cloud. Conflicting goals of minimizing the cost and performance, unknown prior knowledge about task running times, and a lack of estimation tools are just a few of the challenges that resource management systems in those environments encounter. The aim in this paper is to find Pareto-optimal schedules for large-scale Bag-of-Tasks (BoT) applications that meet user defined constraints, such as deadline or budget or some tradeoff between them. BoT applications are common in science and engineering and consist of many independent tasks. To achieve the user's chosen Pareto-optimal schedule, we develop a dynamic resource allocation process for hybrid clouds. We also present a hybrid approach to estimating task running times that incorporates several estimators with a feedback control system to cope with the inherent uncertainty in such estimation. Through extensive experiments on a test bed hybrid cloud, using Amazon EC2 as a public cloud, we show that the proposed approach can achieve near optimality with little overhead, and consistently achieves a solution within 2% of the user's chosen Pareto-optimal schedule. Further, we demonstrate that our approach performs better than an extended List scheduling approach by reducing both the total cost and time needed to run the application by almost 20% and 5% on average, respectively.
M. Reza HoseinyFarahabady, Hamid R. Dehghani Samani, Luke M. Leslie, Young Choon Lee, Albert Y. Zomaya
ICPP4
2012 Workload Characteristic Oriented Scheduler for MapReduce
abstract
Applications in many areas are increasingly developed and ported using the Map Reduce framework (more specifically, Hadoop) to exploit (data) parallelism. The application scope of Map Reduce has been extended beyond the original design goal which was large-scale data processing. This extension inherently makes a need for scheduler to explicitly take into account characteristics of job for two main goals of efficient resource use and performance improvement. In this paper, we study Map Reduce scheduling strategies to effectively deal with different workload characteristics CPU intensive and I/O intensive. We present the Workload Characteristic Oriented Scheduler (WCO), which strives for co-locating tasks of possibly different Map Reduce jobs with complementing resource usage characteristics. WCO is characterized by its essentially dynamic and adaptive scheduling decisions using information obtained from its characteristic estimator. Workload characteristics of tasks are primarily estimated by sampling with the help of some static task selection strategies, e.g., Java byte code analysis. Results obtained from extensive experiments using 11 benchmarks in a 4-node local cluster and a 51-node Amazon EC2 cluster show 17% performance improvement on average in terms of throughput in the situation of co-existing diverse workloads.
Peng Lu 0004, Young Choon Lee, Chen Wang 0008, Bing Bing Zhou, Junliang Chen 0002, Albert Y. Zomaya
ICPADS2
2012 Non-clairvoyant Assignment of Bag-of-Tasks Applications Across Multiple Clouds
abstract
Bag-of-Tasks applications are often composed of a large number of independent tasks, hence, they can easily scale out. With public clouds, the (dynamic) expansion of resource capacity in private clouds is much facilitated. Clearly, cost efficiently running BoT applications in a multi-cloud environment is of great practical importance. In this paper, we investigate how efficiently multiple clouds can be exploited for running BoT applications and present a fully polynomial time randomized approximation scheme (FPRAS) as a novel task assignment algorithm for BoT applications. The resulting task assignment can be optimized in terms of cost, make span or the tradeoff between them. The objective function incorporated into our algorithm is devised in the way the optimization objective is tunable based on user preference. Our task assignment decisions are made without any prior knowledge of the processing time of tasks, i.e., non-clairvoyant task assignment. We adopt a Monte Carlo sampling method to estimate unknown task running time. The experimental results shows our algorithm approximates the optimal solution with little overhead.
M. Reza HoseinyFarahabady, Young Choon Lee, Albert Y. Zomaya
PDCAT2
2012 Profit-driven scheduling for cloud services with data access awareness
Young Choon Lee, Chen Wang 0008, Albert Y. Zomaya, Bing Bing Zhou
J. Parallel Distributed Comput.1
2012 Energy efficient utilization of resources in cloud computing systems
Young Choon Lee, Albert Y. Zomaya
J. Supercomput.1
2012 Cashing in on the Cache in the Cloud
abstract
Over the past decades, caching has become the key technology used for bridging the performance gap across memory hierarchies via temporal or spatial localities; in particular, the effect is prominent in disk storage systems. Applications that involve heavy I/O activities, which are common in the cloud, probably benefit the most from caching. The use of local volatile memory as cache might be a natural alternative, but many well-known restrictions, such as capacity and the utilization of host machines, hinder its effective use. In addition to technical challenges, providing cache services in clouds encounters a major practical issue (quality of service or service level agreement issue) of pricing. Currently, (public) cloud users are limited to a small set of uniform and coarse-grained service offerings, such as High-Memory and High-CPU in Amazon EC2. In this paper, we present the cache as a service (CaaS) model as an optional service to typical infrastructure service offerings. Specifically, the cloud provider sets aside a large pool of memory that can be dynamically partitioned and allocated to standard infrastructure services as disk cache. We first investigate the feasibility of providing CaaS with the proof-of-concept elastic cache system (using dedicated remote memory servers) built and validated on the actual system, and practical benefits of CaaS for both users and providers (i.e., performance and profit, respectively) are thoroughly studied with a novel pricing scheme. Our CaaS model helps to leverage the cloud economy greatly in that 1) the extra user cost for I/O performance gain is minimal if ever exists, and 2) the provider's profit increases due to improvements in server consolidation resulting from that performance gain. Through extensive experiments with eight resource allocation strategies, we demonstrate that our CaaS model can be a promising cost-efficient solution for both users and providers.
Hyuck Han, Young Choon Lee, Woong Shin, Hyungsoo Jung 0001, Heon Young Yeom, Albert Y. Zomaya
IEEE Trans. Parallel Distributed Syst.2
2011 Profiling Applications for Virtual Machine Placement in Clouds
abstract
Application profiling is an important technique for efficient resource management. The decision making of scheduling and resource allocation typically takes great advantage of such a technique primarily for improving resource utilization. With the advent of cloud computing as a multitenant virtualized platform, diverse applications are increasingly deployed onto the cloud and they more than often share physical resources. The background load (other applications running on the same physical machine) is therefore an important factor for profiling an application in this cloud computing scenario. In this paper, we present a novel application profiling technique using the canonical correlation analysis (CCA) method, which identifies the relationship between application performance and resource usage. We further devise a performance prediction model based on application profiles generated using CCA. Clearly, our profiling technique with this prediction model has a lot of potentials particularly in virtual machine (VM) placement with performance awareness. Our experimental results demonstrate the capability of our profiling technique and the accuracy of our prediction model.
Anh Vu Do, Junliang Chen 0002, Chen Wang 0008, Young Choon Lee, Albert Y. Zomaya, Bing Bing Zhou
IEEE CLOUD4
2011 Priority-Based Scheduling for Large-Scale Distribute Systems with Energy Awareness
abstract
Large-scale distributed computing systems (LDSs), such as grids and clouds are primarily designed to provide massive computing capacity. These systems dissipate often excessive energy to both power and cool them. Concerns over greening these systems have prompted a call for scheduling policies with energy awareness (e.g., energy proportionality). The dynamic and heterogeneous nature of resources and tasks in LDSs is a major hurdle to be overcome for energy efficiency when designing scheduling policies. In this paper, we address the problem of scheduling tasks with different priorities (deadlines) for energy efficiency exploiting resource heterogeneity. Specifically, our investigation for energy efficiency focuses on two issues: (1) balancing the workload in the way utilization is maximized and (2) power management by controlling execution of tasks on processor for ensuring the energy is optimally consumed. We form a hierarchical scheduler that exploits the multi-core architecture for effective scheduling. Our scheduling approach exploits the diversity of task priority for proper load balancing across heterogeneous processors while observing energy consumption in the system. Simulation experiments prove the efficacy of our approach, and the comparison results indicate our scheduling policy helps improve energy efficiency of the system.
Masnida Hussin, Young Choon Lee, Albert Y. Zomaya
DASC2
2011 Tradeoffs Between Profit and Customer Satisfaction for Service Provisioning in the Cloud
abstract
The recent cloud computing paradigm represents a trend of moving business applications to platforms run by parties located in different administrative domains. A cloud platform is often highly scalable and cost-effective through its pay-as-you-go pricing model. However, being shared by a large number of users, the running of applications in the platform faces higher performance uncertainty compared to a dedicated platform. Existing Service Level Agreements (SLAs) cannot sufficiently address the performance variation issue. In this paper, we use utility theory leveraged from economics and develop a new utility model for measuring customer satisfaction in the cloud. Based on the utility model, we design a mechanism to support utility-based SLAs in order to balance the performance of applications and the cost of running them. We consider an infrastructure-as-a-service type cloud platform (e.g., Amazon EC2), where a business service provider leases virtual machine (VM) instances with spot prices from the cloud and gains revenue by serving its customers. Particularly, we investigate the interaction of service profit and customer satisfaction. In addition, we present two scheduling algorithms that can effectively bid for different types of VM instances to make tradeoffs between profit and customer satisfaction. We conduct extensive simulations based on the performance data of different types of Amazon EC2 instances and their price history. Our experimental results demonstrate that the algorithms perform well across the metrics of profit, customer satisfaction and instance utilization.
Junliang Chen 0006, Chen Wang 0008, Bing Bing Zhou, Young Choon Lee, Albert Y. Zomaya
HPDC5
2011 Reputation-Based Resource Allocation in Market-Oriented Distributed Systems
Masnida Hussin, Young Choon Lee, Albert Y. Zomaya
ICA3PP (1)2
2011 Efficient Energy Management Using Adaptive Reinforcement Learning-Based Scheduling in Large-Scale Distributed Systems
abstract
Energy consumption in large-scale distributed systems, such as computational grids and clouds gains a lot of attention recently due to its significant performance, environmental and economic implications. These systems consume a massive amount of energy not only for powering them, but also cooling them. More importantly, the explosive increase in energy consumption is not linear to resource utilization as only a marginal percentage of energy is consumed for actual computational works. This energy problem becomes more challenging with uncertainty and variability of workloads and heterogeneous resources in those systems. This paper presents a dynamic scheduling algorithm incorporating reinforcement learning for good performance and energy efficiency. This incorporation helps the scheduler observe and adapt to various processing requirements (tasks) and different processing capacities (resources). The learning process of our scheduling algorithm develops an association between the best action (schedule) and the current state of the environment (parallel system). We have also devised a task-grouping technique to help the decision-making process of our algorithm. The grouping technique is adaptive in nature since it incorporates current workload and energy consumption for the best action. Results from our extensive simulations with varying processing capacities and a diverse set of tasks demonstrate the effectiveness of this learning approach.
Masnida Hussin, Young Choon Lee, Albert Y. Zomaya
ICPP2
2011 A parallel bi-objective hybrid metaheuristic for energy-aware scheduling for cloud computing systems
Mohand-Said Mezmaz, Nouredine Melab, Yacine Kessaci, Young Choon Lee, El-Ghazali Talbi, Albert Y. Zomaya, Daniel Tuyttens
J. Parallel Distributed Comput.4
2011 Energy Conscious Scheduling for Distributed Computing Systems under Different Operating Conditions
abstract
Traditionally, the primary performance goal of computer systems has focused on reducing the execution time of applications while increasing throughput. This performance goal has been mostly achieved by the development of high-density computer systems. As witnessed recently, these systems provide very powerful processing capability and capacity. They often consist of tens or hundreds of thousands of processors and other resource-hungry devices. The energy consumption of these systems has become a major concern. In this paper, we address the problem of scheduling precedence-constrained parallel applications on multiprocessor computer systems and present two energy-conscious scheduling algorithms using dynamic voltage scaling (DVS). A number of recent commodity processors are capable of DVS, which enables processors to operate at different voltage supply levels at the expense of sacrificing clock frequencies. In the context of scheduling, this multiple voltage facility implies that there is a trade-off between the quality of schedules and energy consumption. To effectively balance these two performance goals, we have devised a novel objective function and a variant from that. The main difference between the two algorithms is in their measurement of energy consumption. The extensive comparative evaluations conducted as part of this work show that the performance of our algorithms is very compelling in terms of both application completion time and energy consumption.
Young Choon Lee, Albert Y. Zomaya
IEEE Trans. Parallel Distributed Syst.1
2010 Dynamic Job-Clustering with Different Computing Priorities for Computational Resource Allocation
abstract
The diversity of job characteristics such as unstructured/unorganized arrival of jobs and priorities, could lead to inefficient resource allocation. Therefore, the characterization of jobs is an important aspect worthy of investigation, which enables judicious resource allocation decisions achieving two goals (performance and utilization) and improves resource availability.
Masnida Hussin, Young Choon Lee, Albert Y. Zomaya
CCGRID2
2010 Profit-Driven Service Request Scheduling in Clouds
abstract
A primary driving force of the recent cloud computing paradigm is its inherent cost effectiveness. As in many basic utilities, such as electricity and water, consumers/clients in cloud computing environments are charged based on their service usage, hence the term `pay-per-use'. While this pricing model is very appealing for both service providers and consumers, fluctuating service request volume and conflicting objectives (e.g., profit vs. response time) between providers and consumers hinder its effective application to cloud computing environments. In this paper, we address the problem of service request scheduling in cloud computing systems. We consider a three-tier cloud structure, which consists of infrastructure vendors, service providers and consumers, the latter two parties are particular interest to us. Clearly, scheduling strategies in this scenario should satisfy the objectives of both parties. Our contributions include the development of a pricing model-using processor-sharing-for clouds, the application of this pricing model to composite services with dependency consideration (to the best of our knowledge, the work in this study is the first attempt), and the development of two sets of profit-driven scheduling algorithms.
Young Choon Lee, Chen Wang 0008, Albert Y. Zomaya, Bing Bing Zhou
CCGRID1
2010 Linear Combinations of DVFS-Enabled Processor Frequencies to Modify the Energy-Aware Scheduling Algorithms
abstract
The energy consumption issue in distributed computing systems has become quite critical due to environmental concerns. In response to this, many energy-aware scheduling algorithms have been developed primarily by using the dynamic voltage-frequency scaling (DVFS) capability incorporated in recent commodity processors. The majority of these algorithms involve two passes: schedule generation and slack reclamation. The latter is typically achieved by lowering processor frequency for tasks with slacks. In this paper, we revisit this energy reduction technique from a different perspective and propose a new slack reclamation algorithm which uses a linear combination of the maximum and minimum processor frequencies to decrease energy consumption. This algorithm has been evaluated based on results obtained from experiments with three different sets of task graphs: 1,500 randomly generated task graphs, and 300 task graphs of each of two real-world applications (Gauss-Jordan and LU decomposition). The results show that the amount of energy saved in the proposed algorithm is 13.5%, 25.5% and 0.11% for random, LU decomposition and Gauss-Jordan task graphs, respectively, these percentages for the reference DVFSbased algorithm are 12.4%, 24.6% and 0.1%, respectively.
Nikzad Babaii Rizvandi, Javid Taheri, Albert Y. Zomaya, Young Choon Lee
CCGRID4
2010 A bi-objective hybrid genetic algorithm to minimize energy consumption and makespan for precedence-constrained applications using dynamic voltage scaling
abstract
Precedence-constrained parallel applications are one of the most typical application model used in scientific and engineering fields. Almost all efforts, on this kind of applications, have focused on the minimization of makespan (completion time). It is only recently that much attention has been paid to energy consumption. In this paper, we address the precedence-constrained parallel applications on heterogeneous computing systems (HCSs). We propose a new bi-objective hybrid genetic algorithm that takes into account, not only makespan, but also energy consumption. This metaheuristic adopts dynamic voltage scaling (DVS) to minimize energy consumption. Our study provides promising results showing the significance and potential of DVS. The experimental results from our comparative evaluation study confirm the superior performance of our approach over the other known heuristics on the two criteria energy saving and completion time.
Mohand-Said Mezmaz, Young Choon Lee, Nouredine Melab, El-Ghazali Talbi, Albert Y. Zomaya
IEEE Congress on Evolutionary Computation2
2010 On the Effect of Using Third-Party Clouds for Maximizing Profit
Young Choon Lee, Chen Wang 0008, Javid Taheri, Albert Y. Zomaya, Bing Bing Zhou
ICA3PP (1)1
2010 ADREA: A Framework for Adaptive Resource Allocation in Distributed Computing Systems
abstract
Large-scale distributed computing systems (LDCSs) can be best characterized by their dynamic nature particularly in terms of availability and performance. Typically, these systems deal with various types of jobs in many aspects, such as resource requirements, quality of service (QoS) and other temporal constraints. These diverse characteristics in both resources and jobs impose a great burden on scheduling and resource allocation. That is, inefficient resource allocation brings about poor resource utilization issues and often unreliable job execution. We present the Adaptive Reliable Allocation (ADREA) scheme, which attempts to ensure reliable job execution effectively exploiting heterogeneity in both resources and jobs using a novel clustering technique and a dynamic job migration policy. Specifically, ADREA intends to pave the way in producing better performance (e.g., response time, resource utilization) with reliable computation. Extensive simulations with varying processing capacities and different job arrival rates have been carried out to evaluate our scheme. The results demonstrate that the proposed scheme provides better performance over other algorithms as it significantly improves both job completion time and resource utilization.
Masnida Hussin, Young Choon Lee, Albert Y. Zomaya
PDCAT2
2010 Rescheduling for reliable job completion with the support of clouds
Young Choon Lee, Albert Y. Zomaya
Future Gener. Comput. Syst.1
2010 Robust task scheduling for volunteer computing systems
Young Choon Lee, Albert Y. Zomaya, Howard Jay Siegel
J. Supercomput.1
2009 Minimizing Energy Consumption for Precedence-Constrained Applications Using Dynamic Voltage Scaling
abstract
Jobs on high-performance computing systems are deployed mostly with the sole goal of minimizing completion times. This performance demand has been satisfied without paying much attention to power/energy consumption. Consequently, that has become a major concern in high-performance computing systems. In this paper, we address the problem of scheduling precedence-constrained parallel applications on such systems-specifically with heterogeneous resources-accounting for both application completion time and energy consumption. Our scheduling algorithm adopts dynamic voltage scaling (DVS) to minimize energy consumption. DVS can be used with a number of recent commodity processors that are enabled to operate in different voltage supply levels at the expense of sacrificing clock frequencies. In the context of scheduling, this multiple voltage facility implies that there is a trade-off between the quality of schedules and energy consumption. Our algorithm effectively balances these two performance goals using a novel objective function, which takes into account both goals; this claim is verified by the results obtained from our extensive comparative evaluation study.
Young Choon Lee, Albert Y. Zomaya
CCGRID1
2009 Interweaving heterogeneous metaheuristics using harmony search
abstract
In this paper, we present a novel parallel-metaheuristic framework, which enables a set of heterogeneous metaheuristics to be effectively interwoven and coordinated. The key player of this framework is a harmony-search-based coordinator devised using a recent breed of soft computing paradigm called harmony search that mimics the improvisation process of musicians. For the applicability validation and the performance evaluation, we have implemented a parallel hybrid metaheuristic using the framework for the task scheduling problem on multiprocessor computing systems. Experimental results verify that the proposed framework is a compelling approach to parallelize heterogeneous metaheuristics.
Young Choon Lee, Albert Y. Zomaya
IPDPS1
2009 On the Performance of a Dual-Objective Optimization Model for Workflow Applications on Grid Platforms
abstract
In attempts to exploit a diverse set of resources in grids efficiently, numerous assays in resource management, particularly scheduling, have been made. The primary objective of these efforts is the minimization of application completion time; however, they tend to achieve this objective at the expense of redundant resource usage. This paper investigates the problem of scheduling workflow applications on grids and presents a novel scheduling algorithm for the solution of this problem. Our algorithm performs the scheduling by accounting for both completion time and resource usage-dual objectives. Since the performance of grid resources changes dynamically and the accurate estimation of their performance is very difficult, our algorithm incorporates rescheduling to deal with unforeseen performance fluctuations effectively. The paper provides a comparative evaluation study conducted by using an extensive set of experiments. The study demonstrates that the proposed algorithm delivers promising performance in three respects: completion time, resource utilization, and robustness to resource-performance fluctuations.
Young Choon Lee, Riky Subrata, Albert Y. Zomaya
IEEE Trans. Parallel Distributed Syst.1
2008 Resource-centric task allocation in grids with artificial danger model support
abstract
This paper addresses the problem of scheduling bag-of-tasks (BoT) applications in grids and presents a novel heuristic, called the most suitable match with danger model support algorithm (MSMD) for these applications. Unlike previous approaches, MSMD is capable of efficiently dealing with BoT applications regardless of whether they are computationally or data intensive, or a mixture of both; this strength of MSMD is achieved by making scheduling decisions based on the suitability of resource-task matches, instead of completion time. MSMD incorporates an artificial danger model - based on the danger model in immunology - which selectively responds to unexpected behaviors of resources and applications, in order to increase fault-tolerance. The results from our thorough and extensive evaluation study confirm the superior performance of MSMD, and its generic applicability compared with previous approaches that only consider one or the other of the task requirements.
Young Choon Lee, Albert Y. Zomaya
IPDPS1
2008 A Novel State Transition Method for Metaheuristic-Based Scheduling in Heterogeneous Computing Systems
abstract
Much of the recent literature shows a prevalance in the use of metaheuristics in solving a variety of problems in parallel and distributed computing. This is especially ture for problems that have a combinatorial nature, such as scheduling and load balancing. Despite numerous efforts, task scheduling remains one of the most challenging problems in heterogeneous computing environments. In this paper, we propose a new state transitionscheme , called the Duplication-based State Transition (DST) method specially designed for metaheuristics that can be used for the task scheduling problem in heterogeneous computing environments. State transition in metaheuristics is a key component that takes charge of generating variants of a given state. The DST method produces a new state by first overlapping randomly generated states with the current state and then the resultant state is refined by removing ineffectual tasks. The proposed method is incorporated into three different metaheuristics: genetic algorithms (GAs), simulated annealing (SA), and artificial immune system (AISs). They are experimentally evaluated and are also compared with existing algorithms. The experimental results confirm DST's promising impact on the performance of metaheuristics.
Young Choon Lee, Albert Y. Zomaya
IEEE Trans. Parallel Distributed Syst.1
2007 An Artificial Immune System for Heterogeneous Multiprocessor Scheduling with Task Duplication
abstract
In this study, we investigate the task scheduling problem in heterogeneous computing environments and propose a novel scheduling algorithm, called the artificial immune system with duplication (AISD) algorithm that efficiently tackles the problem. The AISD algorithm incorporates the clonal selection principle in the immune system and task duplication into the scheduling process. Based on the performance results obtained from extensive experiments conducted with a comprehensive set of both randomly generated and well-known application task graphs and various system configurations, AISD consistently outperformed the two existing algorithms by a noticeable margin, especially when scheduling communication intensive task graphs.
Young Choon Lee, Albert Y. Zomaya
IPDPS1
2007 Practical Scheduling of Bag-of-Tasks Applications on Grids with Dynamic Resilience
abstract
Over the past decade, the grid has emerged as an attractive platform to tackle various large-scale problems, especially in science and engineering. One primary issue associated with the efficient and effective utilization of heterogeneous resources in a grid is scheduling. Grid scheduling involves a number of challenging issues, mainly due to the dynamic nature of the grid. There are only a handful of scheduling schemes for grid environments that realistically deal with this dynamic nature that have been proposed in the literature. In this paper, two novel scheduling algorithms, called the shared-input-data-based listing (SIL) algorithm and the multiple queues with duplication (MQD) algorithm for bag-of-tasks (BoT) applications in grid environments are proposed. The SIL algorithm targets scheduling data-intensive BoT (DBoT) applications, whereas the MQD algorithm deals with scheduling computationally intensive BoT (CBoT) applications. Their common and primary forte is that they make scheduling decisions without fully accurate performance prediction information. Another point to note is that both scheduling algorithms adopt task duplication as an attempt to reduce serious schedule increases. Our evaluation study employs a number of experiments with various simulation settings. The results show the practicability and competitiveness of our algorithms when compared to existing methods
Young Choon Lee, Albert Y. Zomaya
IEEE Trans. Computers1
2006 Data Sharing Pattern Aware Scheduling on Grids
abstract
These days an increasing number of applications, especially in science and engineering, are dealing with a massive amount of data; hence they are data-intensive. Bioinformatics, data-mining and image processing are some typical areas of data-intensive applications. Such applications tend to be deployed on grids that provide powerful processing capabilities at reasonable cost. One fundamental scheduling issue, that arises when exploiting grids with these types of applications, is the minimization of data transfer. Therefore, the use of an efficient scheduling scheme that takes into account data transfers is rather essential in order to achieve both a shorter application completion time and efficient system utilization. In this paper, a novel scheduling algorithm, called the shared input data based listing (SIL) algorithm for data-intensive bag-of-tasks (DBoT) applications in grid environments is proposed. The algorithm uses a set of task lists that are constructed taking the data sharing pattern into account and that are reorganized dynamically, based on performance of resources, during the execution of the application. The primary goal of this dynamic listing is to minimize data transfer, thus leading to shortening the overall completion time of DBoT applications. SIL further attempts to reduce serious schedule increases by adopting task duplication. In our evaluation study extensive simulation tests with three different types of the DBoT application model have been conducted. Based on the experimental results, SIL noticeably outperforms two previously proposed algorithms in schedule length
Young Choon Lee, Albert Y. Zomaya
ICPP1
2005 A Productive Duplication-Based Scheduling Algorithm for Heterogeneous Computing Systems
Young Choon Lee, Albert Y. Zomaya
HPCC1