Farshad Khunjush

dblp:91/6142 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
6since 2021 · last 2024
0000-0002-3339-6051ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 2 first-author · 4 since 2021Computer networks · 3 · 2 since 2021Artificial intelligence and machine learning · 2
YearPublicationVenuePosition
2024 Proactive auto-scaling technique for web applications in container-based edge computing using federated learning model
Javad Dogani, Farshad Khunjush
J. Parallel Distributed Comput.2
2023 Host load prediction in cloud computing with Discrete Wavelet Transformation (DWT) and Bidirectional Gated Recurrent Unit (BiGRU) network
Javad Dogani, Farshad Khunjush, Mehdi Seydali
Comput. Commun.2
2023 Auto-scaling techniques in container-based cloud and edge/fog computing: Taxonomy and survey
Javad Dogani, Reza Namvar, Farshad Khunjush
Comput. Commun.3
2023 Multivariate workload and resource prediction in cloud computing using CNN and GRU by attention mechanism
Javad Dogani, Farshad Khunjush, Mohammad Reza Mahmoudi, Mehdi Seydali
J. Supercomput.2
2022 K-AGRUED: A Container Autoscaling Technique for Cloud-based Web Applications in Kubernetes Using Attention-based GRU Encoder-Decoder
Javad Dogani, Farshad Khunjush, Mehdi Seydali
J. Grid Comput.2
2021 Ignite-GPU: a GPU-enabled in-memory computing architecture on clusters
Amir Hossein Sojoodi, Majid Salimi Beni, Farshad Khunjush
J. Supercomput.3
2020 Fully distributed sleeping compressive data gathering in wireless sensor networks
abstract
One of the most essential challenges of wireless sensor networks (WSNs) is the limited energy of their nodes. Compressive data gathering (CDG) and sleep scheduling are two practical solutions to extend the lifetime of such networks. However, applying these techniques incorrectly may lead to having significant reconstruction errors. This study proposes a fully distributed method to combine these two techniques in WSNs. An essential condition for CDG is the participation of all nodes in each round. The simulations along with the mathematical proof demonstrate that sleep scheduling can only be used during CDG when this condition holds. The proposed method considers this condition and therefore leads to an acceptable reconstruction error. Moreover, the proposed method is fully distributed in which message passing among the nodes for sleep scheduling is not required. The simulation results show that the proposed method improves the lifetime of the WSNs up to 2.2 times compared to non‐sleeping CDG with an equivalent reconstruction error.
Saeed Mehrjoo, Farshad Khunjush, Amir Ghaedi
IET Commun.2
2018 Optimal data aggregation tree in wireless sensor networks based on improved river formation dynamics
abstract
Abstract The restricted energy of nodes is one of the most important challenges in wireless sensor networks. Since data transmissions among nodes consume most of the nodes' energy, thus, minimizing the unnecessary transmissions reduces the consumed energy. One of the sources of this problem is the redundancy of raw data that can be eliminated at the aggregation points. As a result, data aggregation can be considered as an effective strategy to tackle the mentioned issue and to optimize the communication energy consumption. In this paper, the sensor nodes are organized in a tree structure, and the data aggregation are done in intermediate nodes at the junction of tree branches. One of the main characteristics of tree protocols is reduction of energy consumption through optimizing the structure of a data aggregation tree. For this, this paper proposes to apply a swarm intelligent algorithm named river formation dynamics. The simulation results show that the proposed algorithm outperforms in comparison to the famous ant colony optimization algorithm in terms of network lifetime. Simulations show that the proposed algorithm makes nearly 4% and 50% improvement in lifetime of wireless sensor networks than ant colony optimization and shortest path routing, respectively.
Saeed Mehrjoo, Farshad Khunjush
Comput. Intell.2
2016 Adaptive sparse matrix representation for efficient matrix-vector multiplication
Pantea Zardoshti, Farshad Khunjush, Hamid Sarbazi-Azad
J. Supercomput.2
2015 A two-tier design space exploration algorithm to construct GPU performance model
Sayyed Ali Mirsoleimani, Farshad Khunjush, Ali Karami
J. Syst. Archit.2
2015 SEATS: smart energy-aware task scheduling in real-time cloud computing
Seyedmehdi Hosseinimotlagh, Farshad Khunjush, Rashidaldin Samadzadeh
J. Supercomput.2
2015 A statistical performance analyzer framework for OpenCL kernels on Nvidia GPUs
Ali Karami, Farshad Khunjush, Sayyed Ali Mirsoleimani
J. Supercomput.2
2014 A Cooperative Two-Tier Energy-Aware Scheduling for Real-Time Tasks in Computing Clouds
abstract
Customers in a cloud would like to receive the results of their task as soon as possible while paying less. On the other hand, cloud providers aim to mitigate the operational cost of cloud environments. In other words, a limited budget makes providers create efficient cloud systems that utilize the computational powers of the clouds while minimizing their energy consumptions and environmental footprints. One of the prevalent techniques in mitigating the total energy consumptions of data-centers is through using consolidation of virtual machines (VMs). However, it incurs significant overheads on both computing resources and network infrastructure of a cloud. Furthermore, it causes tasks to be accomplished later or even it might lead to System Level Agreement (SLA) violations. To address the aforementioned challenges, we propose a cooperative two-tier task scheduling approach to benefit both cloud providers and their customers. It regulates the execution speeds of real-time tasks in a way that a host reaches the optimum level of utilization instead of migrating its tasks to other hosts. We also propose several predictive global task scheduling policies to map arrived tasks to feasible VMs. The simulation results show that the proposed task scheduling approach not only reduces the total energy consumption of a cloud by 41%, but also has profound impacts on turnaround times of real-time tasks by 85%.
Seyedmehdi Hosseinimotlagh, Farshad Khunjush, Seyedmahyar Hosseinimotlagh
PDP2
2013 A parallel memetic algorithm on GPU to solve the task scheduling problem in heterogeneous environments
abstract
Hybrid metaheuristics have shown their capabilities to solve NP-hard problems. However, they exhibit significantly higher execution times in comparison to deterministic approaches. Parallel techniques are usually leveraged to overcome the execution time bottleneck for various metaheuristics. Recently, GPUs have emerged as general purpose parallel processors and have been harnessed to reduce the execution time of these algorithms. In this work, we propose a novel parallel memetic algorithm which is fully offloaded onto GPUs. In addition, we propose an adaptive sorting strategy in order to achieve maximum possible speedups for discrete optimization problems on GPUs. In order to show the efficacy of our algorithm, a task scheduling problem for heterogeneous environments is chosen as a case study. The output of this problem can have a tangible impact on overall performance of parallel heterogeneous platforms. The achieved results of our approach are promising and show up to 696x speedup in comparison to the sequential approach for various versions of this problem. Moreover, the effects of key parameters of memetic algorithms in terms of execution time and solution quality are investigated.
Sayyed Ali Mirsoleimani, Ali Karami, Farshad Khunjush
GECCO3
2008 Extended characterization of DMA transfers on the Cell BE processor
abstract
The main contributors to message delivery latency in message passing environments are the copying operations needed to transfer and bind a received message to the consuming process/thread. A significant portion of the software communication overhead is attributed to message copying. Recently, a set of factors has been leading high- performance processor architectures toward designs that feature multiple processing cores on a single chip (a.k.a. CMP). The Cell Broadband Engine (BE) shows potential to provide high-performance to parallel applications (e.g., MPI applications). The Cell's non-homogeneous architecture along with small local storage in SPEs impose restrictions and challenges for parallel applications. In this work, we first characterize various data delivery mechanisms in the Cell BE processor; then, we propose techniques to facilitate the delivery of a message in MPI environments implemented in the Cell BE processor. We envision a cluster system comprising several cell processors each supporting several computation threads.
Farshad Khunjush, Nikitas J. Dimopoulos
IPDPS1
2007 Comparing Direct-to-Cache Transfer Policies to TCP/IP and M-VIA During Receive Operations in MPI Environments
Farshad Khunjush, Nikitas J. Dimopoulos
ISPA1