Daniel T. Graves

dblp:50/9990 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 62% Distributed systems · 19% Hardware reliability and fault tolerance · 19%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › scientific computing systems
adaptive mesh refinement
0.212016
Granularity and the cost of error recovery in resilient AMR scientific applications · SC 2016
Hardware reliability and fault tolerance
error recovery
0.212016
Granularity and the cost of error recovery in resilient AMR scientific applications · SC 2016
Distributed systems
fault tolerance
0.212016
Granularity and the cost of error recovery in resilient AMR scientific applications · SC 2016
High-performance computing › fault tolerance at scale
local recovery
0.212016
Granularity and the cost of error recovery in resilient AMR scientific applications · SC 2016
High-performance computing
scientific computing systems
0.212016
Granularity and the cost of error recovery in resilient AMR scientific applications · SC 2016

Methods — techniques the papers use, named apart from their topics

parameterization · 0.2cost modeling · 0.2
YearPublicationVenuePosition
2016 Granularity and the cost of error recovery in resilient AMR scientific applications
abstract
Supercomputing platforms are expected to have larger failure rates in the future because of scaling and power concerns. The memory and performance impact may vary with error types and failure modes. Therefore, localized recovery schemes will be important for scientific computations, including failure modes where application intervention is suitable for recovery. We present a resiliency methodology for applications using structured adaptive mesh refinement, where failure modes map to granularities within the application for detection and correction. This approach also enables parameterization of cost for differentiated recovery. The cost model is built with tuning parameters that can be used to customize the strategy for different failure rates in different computing environments. We also show that this approach can make recovery cost proportional to the failure rate.
Anshu Dubey, Hajime Fujita 0002, Daniel T. Graves, Andrew A. Chien, Devesh Tiwari
SC3
2014 A survey of high level frameworks in block-structured adaptive mesh refinement packages
Anshu Dubey, Ann S. Almgren, John B. Bell, Martin Berzins, Steven R. Brandt, Greg Bryan, Phillip Colella, Daniel T. Graves, Michael Lijewski, Frank Löffler 0001, Brian W. O'Shea, Erik Schnetter, Brian van Straalen, Klaus Weide
J. Parallel Distributed Comput.8
2011 Petascale Block-Structured AMR Applications without Distributed Meta-data
Brian van Straalen, Phillip Colella, Daniel T. Graves, Noel Keen
Euro-Par (2)3