EDBT 2026 Demo / reviewers in the wild / expert
Subhendu Behera
dblp:268/1478
· DBLP profile ↗
3ranked-venue papers
3as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Distributed systems · 83% Memory systems · 8% Storage systems · 8% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems › fault tolerance
checkpointing |
0.4 | 1 | 2020 | Orchestrating Fault Prediction with Live Migration and Checkpointing · HPDC 2020 |
Distributed systems › fault tolerance › proactive fault tolerance
failure prediction |
0.4 | 1 | 2020 | Orchestrating Fault Prediction with Live Migration and Checkpointing · HPDC 2020 |
Distributed systems
fault tolerance |
0.4 | 1 | 2020 | Orchestrating Fault Prediction with Live Migration and Checkpointing · HPDC 2020 |
Memory systems › cache management › storage caching
burst buffer |
0.1 | 1 | 2020 | Orchestrating Fault Prediction with Live Migration and Checkpointing · HPDC 2020 |
Storage systems
checkpoint storage |
0.1 | 1 | 2020 | Orchestrating Fault Prediction with Live Migration and Checkpointing · HPDC 2020 |
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Predictive Execution of Workflows in a HPC+Cloud EnvironmentabstractMeeting deadlines for data-intensive workflows on HPC systems is challenging as jobs experience varying wait times before resources become available. This impact is significant in hybrid HPC+Cloud scheduling, which can lead to resource idleness, deadline violations, and higher costs. To address these issues, we propose scheduling data-intensive workflows over a combined HPC+Cloud hybrid environment in a deterministic manner by scavenging unused HPC resources. We predict resource availability (RA) of HPC systems, and exploit this prediction to dynamically split resource allocation between HPC's unused and Cloud's on-demand resources to complete a workflow by a given deadline. The deterministic resource allocation allows for preloading input data for workflow tasks, avoiding execution delays. Further, we develop an adaptive scaling algorithm that effectively backs up the targeted HPC allocation on Cloud facilities to avoid workflow execution delays in the event of incorrect RA estimation. Experiments show that our scheduling technique imposes minimal impact on HPC production jobs, saves cost for$>75 \%$workflow runs, suggests accurate budgets with a mean 7.11 % to 14.75 % cost estimation error, and finishes a mean 98 % to 99.4 % of tasks before deadlines. Subhendu Behera, Jae-Seung Yeom, Daniel Milroy, Marc Niethammer, Frank Mueller 0001 |
HiPC | 1 |
| 2022 | P-ckpt: Coordinated Prioritized CheckpointingabstractGood prediction accuracy and adequate lead time to failure are key to the success of failure-aware Check-point/Restart (C/R) models on current and future large-scale High-Performance Computing (HPC) systems. This paper develops a novel checkpointing technique, called p-ckpt, that aims to maintain the performance efficiency of failure-aware C/R models even when failures are predicted with a small lead time. The p-ckpt technique is developed for HPC systems with multi-level memory systems to prioritize checkpoints from vulnerable nodes (nodes with predicted failure) in the event of failure prediction. It applies coordination among the nodes within an application so that vulnerable nodes' checkpoint data is stored to the Parallel File System (PFS) first by assigning priorities based on the lead time to failure. Vulnerable nodes thus have low-latency access on the critical path to the PFS before any failure happens. Further, we create the hybrid p-ckpt model by integrating Live Migration (LM) because of its cost-effectiveness and to reduce checkpoint frequency. Our hybrid p-ckpt C/R model considers prediction lead time and checkpoint latency to the PFS to decide on a feasible proactive action such as p-ckpt and LM. Simulations of six real-world applications for the Summit supercomputer show a ≈53-65% reduction in overhead due to the hybrid p-ckpt model compared to a ≈31-61% reduction in a state-of-the-art solution. We assess our C/R models against multiple failure distributions and consider lead time variability and failure prediction accuracy. Based on this evaluation and assessment, we discuss the trade-offs of using these models and their impact on application overhead. Subhendu Behera, Lipeng Wan 0001, Frank Mueller 0001, Matthew Wolf, Scott Klasky |
IPDPS | 1 |
| 2020 | Orchestrating Fault Prediction with Live Migration and CheckpointingabstractCheckpoint/Restart (C/R) is widely used to provide fault tolerance on High-Performance Computing (HPC) systems. However, Parallel File System (PFS) overhead and failure uncertainty cause significant application overhead. This paper develops an adaptive multi-level C/R model that incorporates a failure prediction and analysis model, which orchestrates failure prediction, checkpointing, checkpoint frequency, and proactive live migration along with the additional benefit of Burst Buffers (BB). It effectively reduces the overheads due to failures, checkpointing, and recovery. Simulation results for the Summit supercomputer yield a reduction of ~20%-86% in application overhead due to BBs, orchestrated failure prediction, and migration. We also observe a ~29% decrease in checkpoint writes to BBs, which can increase the longevity of the BB storage devices. Subhendu Behera, Lipeng Wan 0001, Frank Mueller 0001, Matthew Wolf, Scott Klasky |
HPDC | 1 |