EDBT 2026 Demo / reviewers in the wild / expert
M. Reza HoseinyFarahabady
dblp:h/MRHoseinyFarahabady · also M. Reza Hoseinyfarahabady, Mohammad Reza Hoseiny Farahabady, MohammadReza HoseinyFarahabady
· DBLP profile ↗
46ranked-venue papers
35as first author
16since 2021 · last 2026
0000-0002-7851-9377ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 16 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2Theory of computation · 2Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GPU-Accelerated Approximate Nearest Neighbor Search via PCA-Augmented Graph Indexing for Vector Databases
Yuanpeng Wang, M. Reza HoseinyFarahabady, Albert Y. Zomaya |
IPDPS | 2 |
| 2025 | Real-Time Interference-Aware CPU and I/O Capping Mechanism for Multi-Tenant ContainersabstractPerformance interference in multi-tenant container-ized environments-such as those utilizing Linux Containers (LX C)-poses a critical challenge to maintaining Quality of Service (QoS) and adhering to Service Level Agreements (SLAs), especially under conditions of high resource contention. Con-ventional static resource allocation strategies often fail to adapt to dynamic workload behaviors and lack cross-resource coor-dination, leading to inefficiencies and degraded performance. While machine learning-based approaches, such as Long Short-Term Memory (LSTM) predictors, have demonstrated improved forecasting capabilities, their computational complexity and training requirements introduce latency and overhead, rendering them impractical for real-time control in resource-constrained deployments. In this paper, we introduce a lightweight, real-time interference-aware resource management solution that integrates predictive modeling with fine-grained CPU and I/O capping mechanisms using Linux cgroup subsystems. Our solution lever-ages continuous profiling of key performance metrics, including QoS violation frequency, CPU throttling rates, and I/O contention signals, to identify emerging interference patterns across co-located containers. The solution dynamically adjusts resource quotas and scheduling parameters in response to runtime observations, enabling adaptive capacity provisioning with minimal system overhead. We implement and evaluate our solution on a heterogeneous LXC-based container cluster with up to 32 concurrently running containers. Experimental results show that our proposed framework achieves an average speedup of 71.4 % for latency-sensitive high-priority workloads compared to the default LX C scheduler, while significantly reducing interference-induced performance degradation across mixed-priority services. M. Reza HoseinyFarahabady, Albert Y. Zomaya |
CLOUD | 1 |
| 2025 | Accelerating Key-Value Data Structures Using AVX-512 SIMD ExtensionsabstractAdvanced Vector Extensions 512 (AVX-512), a modern SIMD instruction set for x86 architectures, enables data-level parallelism through 512-bit wide ZMM registers capable of processing multiple data elements concurrently within a single instruction cycle. In this study, we present a high-throughput, lock-free, in-memory architecture for key-value data-stores that exploits AVX-512 vector operations to accelerate fundamental operations such as insertion and lookup. Our design introduces an optimized memory layout that partitions the key space into two disjoint regions (primary and secondary) and employs three independent hash functions to identify candidate slots. This asymmetric layout improves key distribution, reduces collision probability, and enhances overall lookup efficiency. Experimental evaluation shows that this strategy yields the lowest insertion failure rate among tested memory partitioning schemes. By leveraging AVX-512 instructions in combination with most optimized memory layout, our implementation achieves insertion throughput within 6% of Intel TBB's highly optimized multithreaded hash map, despite avoiding explicit synchronization or thread-level parallelism. Under workloads with 550 million entries and a 90% miss rate, our approach delivers 4.0-5.1x speedup over standard STL, Boost, Robin-Hood, and Abseil hash maps, and up to$2.5 x$improvement relative to TBB and Abseil. These gains are consistently observed for both 32-bit and 64-bit floating-point key types. The results confirm the viability of AVX-512-centric designs as a cost-effective alternative to thread-level parallelism, particularly in environments where minimizing synchronization overhead and ensuring deterministic execution are critical. Our findings suggest for a paradigm shift in CPU and system architecture, emphasizing wider vector units and improved memory bandwidth utilization as primary levers for scalable high-performance computing. These findings suggest that future extensions of AVX-512 capabilities, such as non-blocking memory loads, expanded vector registers, and asynchronous prefetching, could enhance the efficiency of data-intensive workloads. M. Reza HoseinyFarahabady, Javid Taheri, Albert Y. Zomaya |
CLUSTER | 1 |
| 2025 | Scalable Approximate Nearest Neighbor Search with PCA-Augmented HNSW in Vector Databases
Yuanpeng Wang, M. Reza HoseinyFarahabady, Albert Y. Zomaya |
PDCAT | 2 |
| 2024 | Geo-Distributed Analytical Streaming Architecture for IoT PlatformsabstractThe surge in real-time IoT data introduces scalability and computational challenges, necessitating advanced architectural and technological solutions. Streamed data processing is increasingly adopted across industries to enhance operational efficiency by extracting insights from vast, unstructured datasets. However, complex analytical tasks, such as multi-join queries, often require stateful iterative calculations on high-volume, high-velocity data, which challenges conventional programming models like MapReduce. This paper introduces an architectural model enabling application developers to create intricate streaming computational logic within an IoT platform. Our architecture supports scalable applications across distributed edge-tier nodes, particularly for iterative analytical operations on streamed data. We discuss core concepts and a timestamp model (borrowed from the timely data-flow concept) attached to data items circulating between computational blocks, which can execute concurrently on different edge-tier nodes. Additionally, we detail a buffer management mechanism that dynamically adjusts memory size in each computational block on nodes with limited capacity. This mechanism considers application performance requirements and runtime conditions to optimize buffer sizes. Performance evaluation against cloud-tier alternatives confirms the effectiveness of our solution. Experimental results show a significant reduction in p-99 delay compared to cloud-tier deployment with a database engine for analytical applications involving multi-join operations. M. Reza HoseinyFarahabady, Albert Y. Zomaya |
CLUSTER | 1 |
| 2024 | Controlling Performance Interference in Multi-Tenant Containerized EnvironmentsabstractMulti-tenant containerized environments offer numerous benefits to virtualized computing platforms, providing a lightweight and consistent environment for application development. However, the shared nature of resources among coresident containers introduces interference, potentially leading to performance degradation and violating service level agreements (SLA) set by end-users. Addressing the inherent interference in multi-tenant containerized environments is a challenging yet promising endeavor, as it can significantly impact performance and violate SLAs. This paper presents a lightweight system designed to diagnose and control interference in a multitenant containerized environment. We have implemented the proposed solution on top of the Linux container environment and conducted experiments with diverse CPU- and I/O-intensive workloads, including real-world applications like distributed data processing and event stream processing pipelines. The results highlight that our solution achieves an average prediction error below 24% in CPU-bound workloads, with none surpassing 38% across diverse workloads. M. Reza HoseinyFarahabady, Albert Y. Zomaya |
NCA | 1 |
| 2024 | Containerized Data-Flow Processing for Scalable Real-Time Analytics on Edge Devices
M. Reza HoseinyFarahabady, Albert Y. Zomaya |
PDCAT | 1 |
| 2024 | Out-of-Memory GPU Sorting Using Asynchronous CUDA Streams
M. Reza HoseinyFarahabady, Albert Y. Zomaya |
PDCAT | 1 |
| 2024 | I/O Latency Management in Private Cloud Infrastructures
M. Reza HoseinyFarahabady, Albert Y. Zomaya |
PDCAT | 1 |
| 2023 | Energy efficient resource controller for Apache StormabstractSummary Apache Storm is a distributed processing engine that can reliably process unbounded streams of data for real‐time applications. While recent research activities mostly focused on devising a resource allocation and task scheduling algorithm to satisfy high performance or low latency requirements of Storm applications across a distributed and multi‐core system, finding a solution that can optimize the energy consumption of running applications remains an important research question to be further explored. In this article, we present a controlling strategy for CPU throttling that continuously optimize the level of consumed energy of a Storm platform by adjusting the voltage and frequency of the CPU cores while running the assigned tasks under latency constraints defined by the end‐users. The experimental results running over a Storm cluster with 4 physical nodes (total 24 cores) validates the effectiveness of proposed solution when running multiple compute‐intensive operations. In particular, the proposed controller can keep the latency of analytic tasks, in terms of 99th latency percentile, within the quality of service requirement specified by the end‐user while reducing the total energy consumption by 18% on average across the entire Storm platform. M. Reza HoseinyFarahabady, Javid Taheri, Albert Y. Zomaya, Zahir Tari |
Concurr. Comput. Pract. Exp. | 1 |
| 2022 | Enhancing disk input output performance in consolidated virtualized cloud platforms using a randomized approximation schemeabstractAbstract In a virtualized computer system with shared resources, consolidated virtual services (VSs) fiercely compete with each other to obtain the required capacity of resources, and this causes significant system's performance degradation. The performance of input output (I/O)‐bound applications running inside their own VS is mainly determined by the total time required to schedule every read/write request, plus the actual time needed by the device driver to complete the request. To achieve a right performance isolation of shared resources (e.g., the last level cache, memory bandwidth, and the disk buffer), it is essential to limit the performance degradation level among collocated applications, as simultaneously several I/O operations are requested by VSs, perhaps with different priorities. This article proposes a resource allocation controller that uses a fully polynomial‐time randomized approximation scheme to enable performance isolation of concurrent I/O requests in a shared system with multiple consolidated VSs. This controller uses a Monte Carlo sampling approach to measure and estimate the unknown attributes of operational requests originating from each VS. This is formalized as an optimization problem with the aim to minimize the degree of total quality of service (QoS) violation incidents in the entire platform. We associated a reward function to every working machine that represents the fulfillment degree of quality of service metric among all running VSs. The conducted comprehensive set of experiments showed that the proposed algorithm can reduce the QoS violation incidents by 32%, compared with the result which is obtained by employing the default resource allocation policy embedded in the existing Linux container layer. M. Reza HoseinyFarahabady, Javid Taheri, Albert Y. Zomaya, Zahir Tari, Wei Bao 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2022 | MultiScaler: A Multi-Loop Auto-Scaling Approach for Cloud-Based ApplicationsabstractCloud computing offers a wide range of services through a pool of heterogeneous Physical Machines (PMs) hosted on cloud data centers, where each PM can host several Virtual Machines (VMs). Resource sharing among VMs comes with major benefits, but it can create technical challenges that have a detrimental effect on the performance. To ensure a specific service level requested by the cloud-based applications, there is a need for an approach to assign adequate resources to each VM. To this end, we present our novel Multi-Loop Control approach, calledMultiScaler, to allocate resources to VMs based on the Service Level Agreement (SLA) requirements and the run-time conditions.MultiScaleris mainly composed of three different levels working closely with each other to achieve an optimal resource allocation. We propose a set of tailor-made controllers to monitor VMs and take actions accordingly to regulate contention among collocated VMs, to reallocate resources if required, and to migrate VMs from one PM to another. The evaluation in a VMware cluster have shown that theMultiScalerapproach can meet applications performance goals and guarantee the SLA by assigning the exact resources that the applications require. Compared with sophisticated baselines,MultiScalerproduces significantly better reaction to changes in workloads even under the presence of noisy neighbors. Auday Aldulaimy, Javid Taheri, Andreas Kassler, M. Reza HoseinyFarahabady, Shuiguang Deng, Albert Y. Zomaya |
IEEE Trans. Cloud Comput. | 4 |
| 2021 | Data-Intensive Workload Consolidation in Serverless (Lambda/FaaS) PlatformsabstractA significant amount of research studies in the past years has been devoted on developing efficient mechanisms to control the level of degradation among consolidate workloads in a shared platform. Workload consolidation is a promising feature that is employed by most service providers to reduce the total operating costs in traditional computing systems [1]–[3]. Serverless paradigm - also known as Function as a Service, FaaS, and Lambda - recently emerged as a new virtualization run-time model that disentangles the traditional state of applications' users from the burden of provisioning physical computing resources, leaving the difficulty of providing the adequate resource capacity on the service provider's side. This paper focuses on a number of challenges associated with workload consolidation when a serverless platform is expected to execute several data-intensive functional units. Each functional unit is considered to be the atomic component that reacts to a stream of input data. A serverless application in the proposed model is composed of a series of functional units. Through a systematic approach, we highlight the main challenges for devising an efficient workload consolidation process in a data-intensive serverless platform. To this end, we first study the performance interference among multiple workloads to obtain the capacity of last level cache (LLC). We show how such contention among workloads can lead to a significant throughput degradation on a single physical server. We expand our investigation into a general case with the aim to prevent the total throughput never falling below a predefined utilization level. Based on the empirical results, we develop a consolidation model and then design a computationally efficient controller to optimize the throughput degradation among a platform consists fs multiple machines. The performance evaluation is conducted using modern workloads inspired by data management services, and data analytic benchmark tools in our in-house four node platform showing the efficiency of the proposed solution to mitigate the QoS violation rate for high priority applications by 90% while can enhance the normalized throughput usage of disk devices by 39 %. M. Reza HoseinyFarahabady, Javid Taheri, Albert Y. Zomaya, Zahir Tari |
NCA | 1 |
| 2021 | QSpark: Distributed Execution of Batch & Streaming Analytics in Spark PlatformabstractA significant portion of research work in the past decade has been devoted on developing resource allocation and task scheduling solutions for large-scale data processing platforms. Such algorithms are designed to facilitate deployment of data analytic applications across either conventional cluster computing systems or modern virtualized data-centers. The main reason for such a huge research effort stems from the fact that even a slight improvement in the performance of such platforms can bring a considerable monetary savings for vendors, especially for modern data processing engines that are designed solely to perform high throughput or/and low-latency computations over massive-scale batch or streaming data. A challenging question to be yet answered in such a context is to design an effective resource allocation solution that can prevent low resource utilization while meeting the enforced performance level (such as 99-th latency percentile) in circumstances where contention among applications to obtain the capacity of shared resources is a non negligible performance-limiting parameter. This paper proposes a resource controller system, called QSpark, to cope with the problem of (i) low performance (i.e., resource utilization in the batch mode and p-99 response time in the streaming mode), and (ii) the shared resource interference among collocated applications in a multi-tenancy modern Spark platform. The proposed solution leverages a set of controlling mechanisms for dynamic partitioning of the allocation of computing resources, in a way that it can fulfill the QoS re-quirements of latency-critical data processing applications, while enhancing the throughput for all working nodes without reaching their saturation points. Through extensive experiments in our in-house Spark cluster, we compared the achieved performance of proposed solution against the default Spark resource allocation policy for a variety of Machine Learning (ML), Artificial Intelligence (AI), and Deep Learning (DL) applications. Experimental results show the effectiveness of the proposed solution by reducing the p-99 latency of high priority applications by 32 % during the burst traffic periods (for both batch and stream modes), while it can enhance the QoS satisfaction level by 65 % for applications with the highest priority (compared with the results of default Spark resource allocation strategy). M. Reza HoseinyFarahabady, Javid Taheri, Albert Y. Zomaya, Zahir Tari |
NCA | 1 |
| 2021 | A Learning-Based Scheduler for High Volume Processing in Data Warehouse Using Graph Neural Networks
Vivek Bengre, M. Reza HoseinyFarahabady, Mohammad Pivezhandi, Albert Y. Zomaya, Ali Jannesari |
PDCAT | 2 |
| 2021 | Low Latency Execution Guarantee Under Uncertainty in Serverless Platforms
M. Reza HoseinyFarahabady, Javid Taheri, Albert Y. Zomaya, Zahir Tari |
PDCAT | 1 |
| 2020 | Spark-Tuner: An Elastic Auto-Tuner for Apache Spark StreamingabstractSpark has emerged as one of the most widely and successfully used data analytical engine for large-scale enterprise, mainly due to its unique characteristics that facilitate computations to be scaled out in a distributed environment. This paper deals with the performance degradation due to resource contention among collocated analytical applications with different priority and dissimilar intrinsic characteristics in a shared Spark platform. We propose an auto-tuning strategy of computing resources in a distributed Spark platform for handling scenarios in which submitted analytical applications have different quality of service (QoS) requirements (e.g., latency constraints), while the interference among computing resources is considered as a key performance-limiting parameter. We compared Spark-Tuner to two widely used resource allocation heuristics in a large scale Spark cluster through extensive experimental settings across several traffic patterns with uncertain rate and application types. Experimental results show that with Spark-Tuner, the Spark engine can decrease the p-99 latency of high priority applications by 43% during the high-rate traffic periods, while maintaining the same level of CPU throughput across a cluster. M. Reza HoseinyFarahabady, Javid Taheri, Albert Y. Zomaya, Zahir Tari |
CLOUD | 1 |
| 2020 | Q-Flink: A QoS-Aware Controller for Apache FlinkabstractModern stream-data processing platforms are required to execute processing pipelines over high-volume, yet high-velocity, datasets under tight latency constraints. Apache Flink has emerged as an important new technology of large-scale platform that can distribute processing over a large number of computing nodes in a cluster (i.e., scale-out processing). Flink allows application developers to design and execute queries over continuous raw-inputs to analyze a large amount of streaming data in a parallel and distributed fashion. To increase the throughput of computing resources in stream processing platforms, a service provider might be tempted to use a consolidation strategy to pack as many processing applications as possible on the working nodes, with the hope of increasing the total revenue by improving the overall resource utilization. However, there is a hidden trap for achieving such a higher throughput solely by relying on an interference-oblivious consolidation strategy. In practice, collocated applications in a shared platform can fiercely compete with each others for obtaining the capacity of shared resources (e.g., cache and memory bandwidth) which in turn can lead to a severe performance degradation for all consolidated workloads.This paper addresses the shared resource contention problem associated with the auto-resource controlling mechanism of Apache Flink engine running across a distributed cluster. A controlling strategy is proposed to handle scenarios in which stream processing applications may have different quality of service (QoS) requirements while the resource interference is considered as the key performance-limiting parameter. The performance evaluation is carried out by comparing the proposed controller with the default Flink resource allocation strategy in a testbed cluster with total 32 Intel Xeon cores under different workload traffic with up to 4000 streaming applications chosen from various benchmarking tools. Experimental results demonstrate that the proposed controller can successfully decrease the average latency of high priority applications by 223% during the burst traffic while maintaining the requested QoS enforcement levels. M. Reza HoseinyFarahabady, Ali Jannesari, Javid Taheri, Wei Bao 0001, Albert Y. Zomaya, Zahir Tari |
CCGRID | 1 |
| 2020 | Auto-tuning of large-scale iterative operations on modern streaming platformsabstractAs more analytical applications today require real-time processing over high volume data streams, finding an optimal implementation of traditional algorithms which possess iterative computations are gaining popularity and become crucial in most commercial contexts, particularly in edge processing and cloud applications. In this work, we propose an auto-tuning mechanism for enhancing the run-time performance of real-world iterative and cyclic stream processing applications (Multi-Join Operation as the study case) to correctly adjust the right performance bounds for workloads with different characteristics and data-sizes running on modern streaming data processing platform. M. Reza HoseinyFarahabady, Javid Taheri, Albert Y. Zomaya, Zahir Tari |
CoNEXT | 1 |
| 2020 | A Dynamic Resource Controller for Resolving Quality of Service Issues in Modern Streaming Processing EnginesabstractDevising an elastic resource allocation controller of data analytical applications in virtualized data-center has received a great attention recently, mainly due to the fact that even a slight performance improvement can translate to huge monetary savings in practical large-scale execution. Apache Flink is among modern streamed data processing run-times that can provide both low latency and high throughput computation in to execute processing pipelines over high-volume and high-velocity data-items under tight latency constraints. However, a yet to be answered challenge in a large-scale platform with tens of worker nodes is how to resolve the run-time violation in the quality of service (QoS) level in a multi-tenant data streaming platforms, particularly when the amount of workload generated by different users fluctuates. Studies showed that a static resource allocation algorithm (round-robin), which is used by default in Apache Flink, suffer from lack of responsiveness to sudden traffic surges happening unpredictably during the run-time. In this paper, we address the problem of resource management in a Flink platform for ensuring different QoS enforcement levels in a platform with shared computing resources. The proposed solution applies theoretical principals borrowed from close-loop control theory to design a CPU and memory adjustment mechanism with the primary goal to fulfill the different QoS levels requested by submitted applications while the resource interference is considered as the critical performance-limiting factor. The performance evaluation is carried out by comparing the proposed resource allocation mechanism with two static heuristics (round robin and class-based weighted fair queuing) in a 80-core cluster under multiple traffic patterns resembling sudden changes in the incoming workloads of low-priory streaming applications. The experimental results confirm the stability of the proposed controller to regulate the underlying platform resources to smoothly follow the target values (QoS violation rates). Particularly, the proposed solution can achieve higher efficiency compared to the other heuristics by reducing the response-time of high priority applications by 53% while maintaining the enforced QoS levels during the burst traffic periods. M. Reza HoseinyFarahabady, Javid Taheri, Albert Y. Zomaya, Zahir Tari |
NCA | 1 |
| 2020 | Graceful Performance Degradation in Apache Storm
M. Reza HoseinyFarahabady, Javid Taheri, Albert Y. Zomaya, Zahir Tari |
PDCAT | 1 |
| 2020 | Automated Fine-Grained CPU Cap Control in Serverless Computing PlatformabstractServerless computing has emerged as a new cloud computing execution model that liberates users and application developers from explicitly managing `physical' resources, leaving such a resource management burden to service providers. In this article, we study the problem of resource allocation for multi-tenant serverless computing platforms explicitly taking into account workload fluctuations including sudden surges. In particular, we investigate different root causes of performance degradation in these platforms where tenants (their applications) have different workload characteristics. To this end, we develop a fine-grained CPU cap control solution as a resource manager that dynamically adjusts CPU usage limit (or CPU cap) concerning applications with same/similar performance requirements, i.e., application groups. The adjustment of CPU caps applies primarily to co-located worker processes of serverless computing platforms to minimize resource contention, which is the major source of performance degradation. The actual adjustment decisions are made based on performance metrics (e.g., throttled time and queue length) using a group-aware scheduling algorithm. The extensive experimental results performed in our local cluster confirm that the proposed resource manager can effectively eliminate the burden of explicit reservation of computing capacity, even when fluctuations and sudden surges in the incoming workload exist. We measure the robustness of the proposed resource manager by comparing it with several heuristics which extensively used in practice, including the enhanced version of round robin and the least length queue scheduling policies, under various workload intensities driven by real-world scenarios. Notably, our resource manager outperforms other heuristics by decreasing skewness and average response time up to 44 and 94 percent, respectively, while it does not over-use the CPU resources. Young Ki Kim, M. Reza HoseinyFarahabady, Young Choon Lee, Albert Y. Zomaya |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | Dynamic Control of CPU Cap Allocations in Stream Processing and Data-Flow PlatformsabstractThis paper focuses on Timely dataflow programming model for processing streams of data. We propose a technique to define CPU resource allocation (i.e., CPU capping) with the goal to improve response time latency in such type of applications with different quality of service (QoS) level, as they are concurrently running in a shared multi-core computing system with unknown and volatile demand. The proposed solution predicts the expected performance of the underlying platform using an online approach based on queuing theory and adjusts the corrections required in CPU allocation to achieve the most optimized performance. The experimental results confirms that measured performance of the proposed model is highly accurate while it takes into account the percentiles on the QoS metrics. The theoretical model used for elastic allocation of CPU share in the target platform takes advantage of design principals in model predictive control theory and dynamic programming to solve an optimization problem. While the prediction module in the proposed algorithm tries to predict the temporal changes in the arrival rate of each data flow, the optimization module uses a system model to estimate the interference among collocated applications by continuously monitoring the available CPU utilization in individual nodes along with the number of outstanding messages in every intermediate buffer of all TDF applications. The optimization module eventually performs a cost-benefit analysis to mitigate the total amount of QoS violation incidents by assigning the limited CPU shares among collocated applications. The proposed algorithm is robust (i.e., its worst-case output is guaranteed for arbitrarily volatile incoming demand coming from different data streams), and if the demand volatility is not large, the output is optimal, too. Its implementation is done using the TDF framework in Rust for distributed and shared memory architectures. The experimental results show that the proposed algorithm reduces the average and p99 latency of delay-sensitive applications by 21% and 31.8%, respectively, while can reduce the amount of QoS violation incidents by 98% on average. M. Reza HoseinyFarahabady, Ali Jannesari, Zahir Tari, Javid Taheri, Albert Y. Zomaya |
NCA | 1 |
| 2019 | Real-Time Stream Data Processing at ScaleabstractA typical scenario in a stream data-flow processing engine is that users submit continues queries in order to receive the computational result once a new stream of data arrives. The focus of the paper is to design a dynamic CPU cap controller for stream data-flow applications with real-time constraints, in which the result of computations must be available within a short time period, specified by the user, once a recent update in the input data occurs. It is common that the stream data-flow processing engine is deployed over a cluster of dedicated or virtualized server nodes, e.g., Cloud or Edge platform, to achieve a faster data processing. However, the attributes of incoming stream data-flow might fluctuate in an irregular way. To effectively cope with such unpredictable conditions, the underlying resource manager needs to be equipped with a dynamic resource provisioning mechanism to ensure the real-time requirements of different applications. The proposed solution uses control theory principals to achieve a good utilization of computing resources and a reduced average response time. The proposed algorithm dynamically adjusts the required quality of service (QoS) in an environment when multiple stream & data-flow processing applications concurrently run with unknown and volatile workloads. Our study confirms that such a unpredictable demand can negatively degrade the system performance, mainly due to adverse interference in the utilization of shared resources. Unlike prior research studies which assumes a static or zero correlation among the performance variability among consolidated applications, we presume the prevalence of shared-resource interference among collocated applications as a key performance-limiting parameter and confront it in scenarios where several applications have different QoS requirements with unpredictable workload demands. We design a low-overhead controller to achieve two natural optimization objectives of minimizing QoS violation amount and maximizing the average CPU utilization. The algorithm takes advantage of design principals in model predictive control theory for elastic allocation of CPU share. The experimental results confirm that there is a strong correlation in performance degradation among consolidation strategies and the system utilization for obtaining the capacity of shared resources in a non-cooperative manner. The results confirm that the proposed solution can reduce the average latency of delay-sensitive applications by 17% comparing to the results of a well established heuristic called Class-Based Weighted Fair Queuing (CFWFQ). At the same time, the proposed solution can prevent the QoS violation incidents by 62%. M. Reza HoseinyFarahabady, Ali Jannesari, Wei Bao 0001, Zahir Tari, Albert Y. Zomaya |
PDCAT | 1 |
| 2019 | Disk Throughput Controller for Cloud Data-CentersabstractWith the increasing popularity of virtual machine monitoring (VMM) technologies, performance variability among collocated virtual machines (VMs) can easily become a severe scalability issue. Particularly, it becomes a necessary for administrative team to control the performance degradation level in a shared environment when multiple I/O-intensive applications simultaneously request their I/O operations [1]. Nevertheless, adding several logical layers between the running applications and the physical storage system, as seen in contemporary virtualized storage devices, makes it considerably difficult to build a low overhead controlling mechanism for such systems (while each VM may running a separate operating system instance) [2]. In this paper, we propose a strategy based on control theory for managing the performance of several I/O requests, such as mean response times and read/write throughput in a consolidated environment where multiple virtual services can share access to a storage system. This scheme uses an approach for measuring the characterization of read/write performance attributes of each virtual services and also takes into account the run-time quality of service enforcement levels requested by them. This is formulated as an optimization problem where a reward function is defined to reduce the overall QoS violation incidents among all consolidated virtual services. Performance evaluation is carried out by comparing the proposed solution with the default embedded Linux controller across a range of emulated application workloads in scenarios with multiple consolidated virtual containers. The results confirm that the proposed solution can reduce the overall QoS violation incident rates in scenarios in which the platform operates at a significant traffic load comparing to the default policy in LXC engine. M. Reza HoseinyFarahabady, Zahir Tari, Albert Y. Zomaya |
PDCAT | 1 |
| 2018 | Decentralized Admission Control for High-Throughput Key-Value Data StoresabstractWorkload surges are a serious hindrance to per-formance of even high-throughput key-value data stores, such as Cassandra, MongoDB, and more recently Aerospike. In this paper, we present a decentralized admission controller for high-throughput key-value data stores. The proposed controller dynamically regulates the release time of incoming requests explicitly taking into account different Quality of Service (QoS) classes. In particular, an instance of such controller is assigned to each client for its autonomous admission control specific to the client's QoS requirements. These controllers operate in a decentralized manner with only local performance metrics, response time and queue waiting time. Despite the use of such "minimal" run-time state information, our decentralized admission controller is capable of coping with workload surges respecting QoS requirements. The performance evaluation is carried out by comparing the proposed admission controller with the default scheduling policy of Aerospike, in a testbed cluster under various workload intensity rates. Experimental results confirm that the proposed controller improves QoS satisfaction in terms of end-to-end response time by nearly 12 times, on average, compared with that of Aerospike's, in high-rate workload. Results also show decreases of the average and standard deviation of latency up to 31% and 50%, respectively, during workload surges (peak load) in high-rate workload. Young Ki Kim, M. Reza HoseinyFarahabady, Young Choon Lee, Albert Y. Zomaya |
CCGrid | 2 |
| 2018 | Dynamic Control of CPU Usage in a Lambda PlatformabstractLambda platform is a new concept based on an event-driven server-less computation that empowers application developers to build scalable enterprise software in a virtualized environment without provisioning or managing any physical servers (a server-less solution). In reality, however, devising an effective consolidation method to host multiple Lambda functions into a single machine is challenging. The existing simple resource allocation algorithms, such as the round-robin policy used in many commercial server-less systems, suffer from lack of responsiveness to a sudden surge in the incoming workload. This will result in an unsatisfactory performance degradation that is directly experienced by the end-user of a Lambda application. In this paper, we address the problem of CPU cap management in a Lambda platform for ensuring different QoS enforcement levels in a platform with shared resources, in case of fluctuations and sudden surges in the incoming workload requests. To this end, we present a closed-loop (feedback-based) CPU cap controller, which fulfills the QoS levels enforced by the application owners. The controller adjusts the number of working threads per QoS class and dispatches the outstanding Lambda functions along with the associated events to the most appropriate working thread. The proposed solution reduces the QoS violations by an average of 6.36 times compared to the round-robin policy. It can also maintain the end-to-end response time of applications belonging to the highest priority QoS class close to the target set-point while decreasing the overall response time by up to 52%. Young Ki Kim, M. Reza HoseinyFarahabady, Young Choon Lee, Albert Y. Zomaya, Raja Jurdak |
CLUSTER | 2 |
| 2018 | A Model Predictive Controller for Managing QoS Enforcements and Microarchitecture-Level Interferences in a Lambda PlatformabstractLambda paradigm, also known as Function as a Service (FaaS), is a novel event-driven concept that allows companies to build scalable and reliable enterprise applications in an off-premise computing data-center as a serverless solution. In practice, however, an important goal for the service provider of a Lambda platform is to devise an efficient way to consolidate multiple Lambda functions in a single host. While the majority of existing resource management solutions use only operating-system level metrics (e.g., average utilization of computing and I/O resources) to allocate the available resources among the submitted workloads in a balanced way, a resource allocation schema that is oblivious to the issue of shared-resource contention can result in a significant performance variability and degradation within the entire platform. This paper proposes a predictive controller scheme that dynamically allocates resources in a Lambda platform. This scheme uses a prediction tool to estimate the future rate of every event stream and takes into account the quality of service enforcements requested by the owner of each Lambda function. This is formulated as an optimization problem where a set of cost functions are introduced (i) to reduce the total QoS violation incidents; (ii) to keep the CPU utilization level within an accepted range; and (iii) to avoid the fierce contention among collocated applications for obtaining shared resources. Performance evaluation is carried out by comparing the proposed solution with an enhanced interference-aware version of three well-known heuristics, namely spread, binpack (the two native clustering solutions employed by Docker Swarm) and best-effort resource allocation schema. Experimental results show that the proposed controller improves the overall performance (in terms of reducing the end-to-end response time) by 14.9 percent on average compared to the best result of the other heuristics. The proposed solution also increases the overall CPU utilization by 18 percent on average (for lightweight workloads), while achieves an average 87 percent (maximum 146 percent) improvement in preventing QoS violation incidents. M. Reza HoseinyFarahabady, Albert Y. Zomaya, Zahir Tari |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | QoS- and Contention- Aware Resource Provisioning in a Stream Processing EngineabstractThis paper addresses the shared resource contention problem associated with the auto-parallelization of running queries in distributed stream processing engines. In such platforms, analyzing a large amount of data often requires to execute user-defined queries over continues raw-inputs in a parallel fashion at each single host. However, previous studies showed that the collocated applications can fiercely compete for shared resources, resulting in a severe performance degradation among applications. This paper presents an advanced resource allocation strategy for handling scenarios in which the target applications have different quality of service (QoS) requirements while shared-resource interference is considered as a key performance-limiting parameter. To properly allocate the best possible resource to each query, the proposed controller predicts the performance degradation of the running pane-level as well as the window-level queries when co-running with other queries. This is addressed as an optimization problem where a set of cost functions is defined to achieve the following goals: a) reduce the sum of QoS violation incidents over all machines; b) keep the CPU utilization level within an accepted range; and c) avoid fierce shared resource interference among collocated applications. Particle swarm optimization is used to find an acceptable solution at each round of the controlling period. The performance of the proposed solution is benchmarked with Round-Robin and best-effort strategies, and the experimental results clearly demonstrate that the proposed controller has the following advantages over its opponents: it increases the overall resource utilization by 15% on average while can reduce the average tuple latencies by 14%. It also achieves an average 123% improvement in preventing QoS violation incidents. M. Reza HoseinyFarahabady, Albert Y. Zomaya, Zahir Tari |
CLUSTER | 1 |
| 2017 | A Dynamic Resource Controller for a Lambda ArchitectureabstractLambda architecture is a novel event-driven serverless paradigm that allows companies to build scalable and reliable enterprise applications. As an attractive alternative to traditional service oriented architecture (SOA), Lambda architecture can be used in many use cases including BI tools, in-memory graph databases, OLAP, and streaming data processing. In practice, an important aim of Lambda's service providers is devising an efficient way to co-locate multiple Lambda functions with different attributes into a set of available computing resources. However, previous studies showed that consolidated workloads can compete fiercely for shared resources, resulting in severe performance variability/degradation. This paper proposes a resource allocation mechanism for a Lambda platform based on the model predictive control framework. Performance evaluation is carried out by comparing the proposed solution with multiple resource allocation heuristics, namely enhanced versions of spread and binpack, and best-effort approaches. Results confirm that the proposed controller increases the overall resource utilization by 37% on average and achieves a significant improvement in preventing QoS violation incidents compared to others. M. Reza HoseinyFarahabady, Javid Taheri, Zahir Tari, Albert Y. Zomaya |
ICPP | 1 |
| 2017 | A QoS-Aware Resource Allocation Controller for Function as a Service (FaaS) Platform
M. Reza HoseinyFarahabady, Young Choon Lee, Albert Y. Zomaya, Zahir Tari |
ICSOC | 1 |
| 2017 | A resource allocation controller for key-value data storesabstractRecent distributed key-value data stores, such as Aerospike are getting the momentum with ever-increasing need for large-scale real-time data processing. While these data stores can provide significantly improved performance, they still struggle to meet Quality of Service (QoS) during workload surges. In this paper, we address the problem of QoS-aware resource allocation for burst workloads in key-value data stores. To this end, we design a resource allocation controller, which enables each application to independently regulate the releases of its requests taking into account QoS. In particular, the proposed controller monitors the actual performance metrics of the target system and dynamically releases requests from a buffer owned by each application accordingly. We have implemented the proposed controller in an Aerospike cluster for our performance evaluation. Experiments have been conducted with various workload intensities (with up to 36,180 write operations per second) in comparison with the default Aerospike policy. Experimental results confirm that the proposed controller decreases the overall average latency up to 41% on high-rate workload while maintaining the QoS of high priority applications. Young Ki Kim, M. Reza HoseinyFarahabady, Young Choon Lee, Albert Y. Zomaya |
NCA | 2 |
| 2017 | QoS-aware resource allocation for stream processing engines using priority channelsabstractThis paper addresses the challenging problem of guaranteeing quality-of-service (QoS) requirements associated with parallel running queries in distributed stream processing engines. In such platforms, the real-time processing of streaming data often requires executing a set of user-defined queries over continues data flows. However, previous studies showed that guaranteeing QoS enforcement (such as end-to-end response time) for a collection of applications is a complex problem. This paper presents an advanced resource allocation strategy to tackle such a problem by considering the traffic pattern of individual data streams. To properly allocate resource for streaming queries execution, we define a certain number of priority channels to categorize the streaming data across the system. The resource allocation is addressed as an optimization problem where a set of cost functions is defined to achieve the following goals: a) reduce the sum of QoS violation incidents across all applications; b) increase the CPU utilization level, and (c) avoid the additional costs caused by frequent reconfigurations. The proposed solution does not depend on any assumption about the incoming data rate or the query processing time. The performance of the proposed solution is benchmarked, and the experimental results reveal that the proposed scheme increases the overall resource utilization by 23% on average and reduces the QoS violations by 29% against round-robin strategy. It could also prevent QoS violation incidents at different levels by tuning the cost function. Zahir Tari, M. Reza HoseinyFarahabady, Albert Y. Zomaya |
NCA | 3 |
| 2016 | A Model Predictive Controller for Contention-Aware Resource Allocation in Virtualized Data CentersabstractData center efficiency is primarily sought by sharing physical resources, such as processors, memory, and disks in the form of virtual machines or containers among multiple users, i.e., workload consolidation. However, the reality is co-located applications in these virtual platforms compete for resources and interfere with each others' performance, resulting in performance variability/degradation. In this paper, we present the contentionaware resource allocation (CARA) solution, which optimizes data center efficiency. It is essentially devised based on a model predictive control that enables to make judicious consolidation decisions with future system states. CARA consolidates workloads explicitly taking into account the correlation between shared and isolated resource usage patterns. Based on our experimental results, CARA improves the overall resource utilization by 32%, without a significant impact on the quality-of-service (QoS) enforcement level. Such improvement results in a fewer number of active servers and in turn contributes to an overall energy saving by 33%. M. Reza HoseinyFarahabady, Young Choon Lee, Albert Y. Zomaya, Zahir Tari, Andy Song |
MASCOTS | 1 |
| 2016 | A QoS-aware controller for Apache StormabstractApache Storm has recently emerged as an attractive fault-tolerant open-source distributed data processing platform that has been chosen by many industry leaders to develop real-time applications for processing a huge amount of data in a scalable manner. A key aspect to achieve the best performance in this system lies on the design of an efficient scheduler for component execution, called topology, on the available computing resources. In response to workload fluctuations, we propose an advanced scheduler for Apache Storm that provides improved performance with highly dynamic behavior. While enforcing the required Quality-of-Service (QoS) of individual data streams, the controller allocates computing resources based on decisions that consider the future states of non-controllable disturbance parameters, e.g. arriving rate of tuples or resource utilization in each worker node. The performance evaluation is carried out by comparing the proposed solution with two well-known alternatives, namely the Storm's default scheduler and the best-effort approach (i.e. the heuristic that is based on the first-fit decreasing approximation algorithm). Experimental results clearly show that the proposed controller increases the overall resource utilization by 31% on average compared to the two others solutions, without significant negative impact on the QoS enforcement level. M. Reza HoseinyFarahabady, Hamid R. Dehghani Samani, Albert Y. Zomaya, Zahir Tari |
NCA | 1 |
| 2014 | Randomized approximation scheme for resource allocation in hybrid-cloud environment
M. Reza HoseinyFarahabady, Young Choon Lee, Albert Y. Zomaya |
J. Supercomput. | 1 |
| 2014 | Pareto-Optimal Cloud BurstingabstractLarge-scale Bag-of-Tasks (BoT) applications are characterized by their massively parallel, yet independent operations. The use of resources in public clouds to dynamically expand the capacity of a private computer system might be an appealing alternative to cope with such massive parallelism. To fully realize the benefit of this ‘cloud bursting’, the performance to cost ratio (or cost efficiency) must be thoroughly studied and incorporated into scheduling and resource allocation strategies. In this paper, we present PANDA, a framework for static scheduling BoT applications across resources in both private and public clouds. The framework at the core incorporates a fully polynomial-time approximation scheme (FPTAS) as a novel scheduling algorithm, which generates schedules with the best trade-off point between cost and performance; hence Pareto-optimality. We have theoretically discussed the complexity and correctness of our algorithms, and experimentally verified their efficacy and practicality using ISOMAP—a widely-used nonlinear manifold method as a real-world BoT application. Our evaluation conducted in a 'multi-cloud' environment of our 40-core private system and Amazon EC2 public cloud demonstrates the scheduling quality of PANDA is guaranteed to be within a measurable distance from the optimal solution. Results obtained from our experiments show such quality is 8 percent or less from the optimum. We also show the sensitivity and robustness of our scheduling solutions against performance errors in both resources and applications. M. Reza HoseinyFarahabady, Young Choon Lee, Albert Y. Zomaya |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | Handling Uncertainty: Pareto-Efficient BoT Scheduling on Hybrid CloudsabstractCoping with uncertainty is a challenging and complex problem particularly in hybrid cloud environments-private cloud plus public cloud. Conflicting goals of minimizing the cost and performance, unknown prior knowledge about task running times, and a lack of estimation tools are just a few of the challenges that resource management systems in those environments encounter. The aim in this paper is to find Pareto-optimal schedules for large-scale Bag-of-Tasks (BoT) applications that meet user defined constraints, such as deadline or budget or some tradeoff between them. BoT applications are common in science and engineering and consist of many independent tasks. To achieve the user's chosen Pareto-optimal schedule, we develop a dynamic resource allocation process for hybrid clouds. We also present a hybrid approach to estimating task running times that incorporates several estimators with a feedback control system to cope with the inherent uncertainty in such estimation. Through extensive experiments on a test bed hybrid cloud, using Amazon EC2 as a public cloud, we show that the proposed approach can achieve near optimality with little overhead, and consistently achieves a solution within 2% of the user's chosen Pareto-optimal schedule. Further, we demonstrate that our approach performs better than an extended List scheduling approach by reducing both the total cost and time needed to run the application by almost 20% and 5% on average, respectively. M. Reza HoseinyFarahabady, Hamid R. Dehghani Samani, Luke M. Leslie, Young Choon Lee, Albert Y. Zomaya |
ICPP | 1 |
| 2012 | Non-clairvoyant Assignment of Bag-of-Tasks Applications Across Multiple CloudsabstractBag-of-Tasks applications are often composed of a large number of independent tasks, hence, they can easily scale out. With public clouds, the (dynamic) expansion of resource capacity in private clouds is much facilitated. Clearly, cost efficiently running BoT applications in a multi-cloud environment is of great practical importance. In this paper, we investigate how efficiently multiple clouds can be exploited for running BoT applications and present a fully polynomial time randomized approximation scheme (FPRAS) as a novel task assignment algorithm for BoT applications. The resulting task assignment can be optimized in terms of cost, make span or the tradeoff between them. The objective function incorporated into our algorithm is devised in the way the optimization objective is tunable based on user preference. Our task assignment decisions are made without any prior knowledge of the processing time of tasks, i.e., non-clairvoyant task assignment. We adopt a Monte Carlo sampling method to estimate unknown task running time. The experimental results shows our algorithm approximates the optimal solution with little overhead. M. Reza HoseinyFarahabady, Young Choon Lee, Albert Y. Zomaya |
PDCAT | 1 |
| 2011 | Pancyclicity of OTIS (swapped) networks based on properties of the factor graph
Marzieh Malekimajd, M. Reza HoseinyFarahabady, Ali Movaghar-Rahimabadi, Hamid Sarbazi-Azad |
Inf. Process. Lett. | 2 |
| 2011 | On pancyclicity properties of OTIS-mesh
T. Shafiei, M. Reza HoseinyFarahabady, Ali Movaghar-Rahimabadi, Hamid Sarbazi-Azad |
Inf. Process. Lett. | 2 |
| 2008 | Some topological and combinatorial properties of WK-recursive mesh and WK-pyramid interconnection networks
M. Reza HoseinyFarahabady, Navid Imani, Hamid Sarbazi-Azad |
J. Syst. Archit. | 1 |
| 2007 | On Pancyclicity Properties of OTIS Networks
M. Reza HoseinyFarahabady, Hamid Sarbazi-Azad |
HPCC | 1 |
| 2006 | On the Fault Patterns Properties in the Torus Networks
M. Reza HoseinyFarahabady, Farshad Safaei, Ahmad Khonsari, Mahmood Fathy |
AICCSA | 1 |
| 2006 | Characterization of spatial fault patterns in interconnection networks
M. Reza HoseinyFarahabady, Farshad Safaei, Ahmad Khonsari, Mahmood Fathy |
Parallel Comput. | 1 |
| 2006 | The Grid-Pyramid: A Generalized Pyramid Network
M. Reza HoseinyFarahabady, Hamid Sarbazi-Azad |
J. Supercomput. | 1 |