Md Rajib Hossen

dblp:294/6958 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2024
0000-0002-0882-2434ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 54% Energy-efficient computing · 28% Distributed systems · 14%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems › peer-to-peer systems
incentive mechanisms
0.712023
Market Mechanism-Based User-in-the-Loop Scalable Power Oversubscription for HPC Systems · HPCA 2023
Energy-efficient computing
power management
0.712023
Market Mechanism-Based User-in-the-Loop Scalable Power Oversubscription for HPC Systems · HPCA 2023
Energy-efficient computing › datacenter power management
power oversubscription
0.712023
Market Mechanism-Based User-in-the-Loop Scalable Power Oversubscription for HPC Systems · HPCA 2023
Cloud and datacenter computing
resource allocation
0.712023
Market Mechanism-Based User-in-the-Loop Scalable Power Oversubscription for HPC Systems · HPCA 2023
Cloud and datacenter computing
cluster resource management and scheduling
0.612022
Practical Efficient Microservice Autoscaling with QoS Assurance · HPDC 2022
Cloud and datacenter computing › autoscaling
microservice autoscaling
0.612022
Practical Efficient Microservice Autoscaling with QoS Assurance · HPDC 2022
Cloud and datacenter computing › microservices
microservice resource management
0.612022
Practical Efficient Microservice Autoscaling with QoS Assurance · HPDC 2022
Cloud and datacenter computing › quality of service
quality-of-service assurance
0.212022
Practical Efficient Microservice Autoscaling with QoS Assurance · HPDC 2022

Methods — techniques the papers use, named apart from their topics

trace-based simulation · 0.7market mechanism · 0.7rule-based autoscaling · 0.6opportunistic resource reduction · 0.6
YearPublicationVenuePosition
2024 Enabling Workload-Driven Elasticity in MPI-based Ensembles
abstract
Interdisciplinary workflows are evolving to accom-modate the growing resource diversity and parallelism in modern computing systems. The necessity of integrating various components, including multi-scale simulations and Artificial Intelligence and Machine Learning (AI/ML) with ensemble methods, has made the workflows increasingly complex and challenging to manage using traditional High performance computing (HPC) infrastructure. Cloud computing provides capabilities such as container orchestration, automation, and elasticity to manage the growing heterogeneity, scale, and complexity of HPC systems and workflows. Converged computing, a growing movement that integrates HPC and cloud technologies into a seamless environ-ment, can provide a means to bridge the gap between needs and capabilities. In particular, ensemble-based HPC workflows can benefit from the potential efficiency improvements afforded by these capabilities. While MPI - based (Message Passing Interface) workflows have been demonstrated to scale using cloud-native orchestration in Kubernetes, there is little work on understanding the combined impact of autoscaling and elasticity on MPI-based workflows. In this study, we explore the cost and performance of elasticity applied to ensembles of MPI-based HPC simulations using the Flux Operator. We propose a workload-driven autoscaling algorithm that outperforms CPU utilization-based autoscaling for MPI-based ensembles. The efficiency gains afforded by elastic, autoscaled approaches for MPI - based ensembles are described. We demonstrate that the workload-driven algorithm can reduce ensemble completion time by up to$4.7\times$in comparison with CPU utilization-based autoscaling.
Md Rajib Hossen, Vanessa Sochat, Abhik Sarkar, Mohammad A. Islam 0001, Daniel Milroy
CLUSTER1
2023 Market Mechanism-Based User-in-the-Loop Scalable Power Oversubscription for HPC Systems
abstract
Significant power consumption is one of the major challenges for current and future high-performance computing (HPC) systems. All the while, HPC systems generally remain power underutilized, making them a great candidate for applying power oversubscription to reclaim unused capacity. However, an oversubscribed HPC system may occasionally get overloaded. In this paper, we propose MPR (Market-based Power Reduction), a scalable market-based approach where users actively participate in reducing the HPC system’s power consumption to mitigate overloads. In MPR, HPC users bid to supply, in exchange for incentives, the resource reduction required for handling the overloads. Using several real-world trace-based simulations, we extensively evaluate MPR and show that, by participating in MPR, users always receive more rewards than the cost of performance loss. At the same time, the HPC manager enjoys orders of magnitude more resource gain than her incentive payoff to the users. We also demonstrate the real-world effectiveness of MPR on a prototype system.
Md Rajib Hossen, Kishwar Ahmed, Mohammad A. Islam 0001
HPCA1
2022 Practical Efficient Microservice Autoscaling with QoS Assurance
abstract
Cloud applications are increasingly moving away from monolithic services to agile microservices-based deployments. However, efficient resource management for microservices poses a significant hurdle due to the sheer number of loosely coupled and interacting components. The interdependencies between various microservices make existing cloud resource autoscaling techniques ineffective. Meanwhile, machine learning (ML) based approaches that try to capture the complex relationships in microservices require extensive training data and cause intentional SLO violations. Moreover, these ML-heavy approaches are slow in adapting to dynamically changing microservice operating environments. In this paper, we propose PEMA (Practical Efficient Microservice Autoscaling), a lightweight microservice resource manager that finds efficient resource allocation through opportunistic resource reduction. PEMA's lightweight design enables novel workload-aware and adaptive resource management. Using three prototype microservice implementations, we show that PEMA can find efficient resource allocation and save up to 33% resources compared to the commercial rule-based resource allocations.
Md Rajib Hossen, Mohammad A. Islam 0001, Kishwar Ahmed
HPDC1