Michael A. Heffner

dblp:66/2523 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
0since 2021 · last 2007
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 89% Parallel and multicore computing · 11%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
distributed coordination
0.112006
Poster reception - DejaVu: transparent user-level checkpointing, migration and recovery for distributed systems · SC 2006
Distributed systems
fault tolerance
0.112006
Poster reception - DejaVu: transparent user-level checkpointing, migration and recovery for distributed systems · SC 2006
Distributed systems › distributed system architecture › distributed operating systems
process migration
0.112006
Poster reception - DejaVu: transparent user-level checkpointing, migration and recovery for distributed systems · SC 2006
Distributed systems › fault tolerance
rollback recovery
0.112006
Poster reception - DejaVu: transparent user-level checkpointing, migration and recovery for distributed systems · SC 2006
Distributed systems › fault tolerance › checkpointing
transparent checkpointing
0.112006
Poster reception - DejaVu: transparent user-level checkpointing, migration and recovery for distributed systems · SC 2006
Parallel and multicore computing
MPI
0.012006
Poster reception - DejaVu: transparent user-level checkpointing, migration and recovery for distributed systems · SC 2006
Parallel and multicore computing
parallel programming models and runtimes
0.012006
Poster reception - DejaVu: transparent user-level checkpointing, migration and recovery for distributed systems · SC 2006

Methods — techniques the papers use, named apart from their topics

state capture instrumentation · 0.1incremental checkpointing · 0.1
YearPublicationVenuePosition
2007 DejaVu: Transparent User-Level Checkpointing, Migration, and Recovery for Distributed Systems
abstract
In this paper, we present a new fault tolerance system called DejaVu for transparent and automatic checkpointing, migration, and recovery of parallel and distributed applications. DejaVu provides a transparent parallel checkpointing and recovery mechanism that recovers from any combination of systems failures without any modification to parallel applications or the OS. It uses a new runtime mechanism for transparent incremental checkpointing that captures the least amount of state needed to maintain global consistency and provides a novel communication architecture that enables transparent migration of existing MPI codes, without source-code modifications. Performance results from the production-ready implementation show less than 5% overhead in real-world parallel applications with large memory footprints.
Joseph F. Ruscio, Michael A. Heffner, Srinidhi Varadarajan
IPDPS2
2006 Poster reception - DejaVu: transparent user-level checkpointing, migration and recovery for distributed systems
abstract
We present a new fault tolerance system, DejaVu, for transparent and automatic checkpointing, migration and recovery of parallel and distributed applications. DejaVu has several novel features. First, it provides a transparent parallel checkpointing and recovery mechanism that recovers from any combination of systems failures without modification to parallel applications or the underlying operating system. Second, it uses a novel instrumentation and state capture mechanism that transparently captures application state. Third, it uses a new runtime mechanism for transparent incremental checkpointing, capturing the least amount of state needed to maintain global consistency. Finally, it provides a novel communication architecture that enables transparent migration of existing MPI codes, without source-code modifications. DejaVu has been implemented for 32 bit and 64 bit Linux platforms on x86 processors interconnected over Infiniband or Gigabit Ethernet networks. Performance results from the production-ready implementation shows less than 5% overhead with real-world parallel applications with large memory footprints.
Joseph F. Ruscio, Michael A. Heffner, Srinidhi Varadarajan
SC2