Anshu Dubey

dblp:79/6344 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
4since 2021 · last 2025
0000-0003-3299-7426ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 7 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 RAPTOR: Practical Numerical Profiling of Scientific Applications
abstract
The proliferation of low-precision units in modern high-performance architectures increasingly burdens domain scientists. Historically, the choice in HPC was easy: can we get away with 32 bit floating-point operations and lower bandwidth requirements, or is FP64 necessary? Driven by Artificial Intelligence, vendors introduce novel low-precision units for vector and tensor operations, and FP64 capabilities stagnate or are reduced. This forces scientists to re-evaluate their codes, but a trivial search-and-replace approach to go from FP64 to FP16 will not suffice.
Faveo Hoerold, Ivan R. Ivanov, Akash Dhruv, William S. Moses, Anshu Dubey, Mohamed Wahib, Jens Domke
SC5
2025 A tool and a methodology to use macros for abstracting variations in code for different computational demands
Anshu Dubey, Youngjun Lee, Tom Klosterman, Emil Vatai
Future Gener. Comput. Syst.1
2025 CG-Kit: Code Generation Toolkit for performant and maintainable variants of source code applied to Flash-X hydrodynamics simulations
Johann Rudi, Youngjun Lee, Aidan H. Chadha, Mohamed Wahib, Klaus Weide, Jared O'Neal, Anshu Dubey
Future Gener. Comput. Syst.7
2021 Towards performance portability in the Spark astrophysical magnetohydrodynamics solver in the Flash-X simulation framework
Sean M. Couch, Jared Carlson, Michael Pajkos, Brian W. O'Shea, Anshu Dubey, Tom Klosterman
Parallel Comput.5
2018 Experience report: refactoring the mesh interface in FLASH, a multiphysics software
abstract
FLASH is a highly-configurable multiphysics software designed for solving a large class of problems that involve fluid flows and need adaptive mesh refinement (AMR). FLASH has been in existence for two decades and has undergone four major revisions. It is now undergoing its fifth major revision to deal with increasingly heterogeneous platforms. The architecture of previous versions of the code and the AMR package at its core, Paramesh, are inadequate to meet the challenges posed by heterogeneity. In this paper we describe our experience with refactoring the mesh interface of the code to work with a more modern AMR library, AMReX. The focus of the paper is the refactoring methodology and the attendant software process that we have found useful to ensure that code quality is maintained during the transition.
Jared O'Neal, Klaus Weide, Anshu Dubey
e-Science3
2017 Trends in Data Locality Abstractions for HPC Systems
abstract
The cost of data movement has always been an important concern in high performance computing (HPC) systems. It has now become the dominant factor in terms of both energy consumption and performance. Support for expression of data locality has been explored in the past, but those efforts have had only modest success in being adopted in HPC applications for various reasons. them However, with the increasing complexity of the memory hierarchy and higher parallelism in emerging HPC systems, locality management has acquired a new urgency. Developers can no longer limit themselves to low-level solutions and ignore the potential for productivity and performance portability obtained by using locality abstractions. Fortunately, the trend emerging in recent literature on the topic alleviates many of the concerns that got in the way of their adoption by application developers. Data locality abstractions are available in the forms of libraries, data structures, languages and runtime systems; a common theme is increasing productivity without sacrificing performance. This paper examines these trends and identifies commonalities that can combine various locality concepts to develop a comprehensive approach to expressing and managing data locality on future large-scale high-performance computing systems.
Didem Unat, Anshu Dubey, Torsten Hoefler, John Shalf, Mark James Abraham, Mauro Bianco, Bradford L. Chamberlain, Romain Cledat, H. Carter Edwards, Hal Finkel, Karl Fürlinger, Frank Hannig, Emmanuel Jeannot, Amir Kamil, Jeff Keasler, Paul H. J. Kelly, Vitus J. Leung, Hatem Ltaief, Naoya Maruyama, Chris J. Newburn, Miquel Pericàs
IEEE Trans. Parallel Distributed Syst.2
2016 Granularity and the cost of error recovery in resilient AMR scientific applications
abstract
Supercomputing platforms are expected to have larger failure rates in the future because of scaling and power concerns. The memory and performance impact may vary with error types and failure modes. Therefore, localized recovery schemes will be important for scientific computations, including failure modes where application intervention is suitable for recovery. We present a resiliency methodology for applications using structured adaptive mesh refinement, where failure modes map to granularities within the application for detection and correction. This approach also enables parameterization of cost for differentiated recovery. The cost model is built with tuning parameters that can be used to customize the strategy for different failure rates in different computing environments. We also show that this approach can make recovery cost proportional to the failure rate.
Anshu Dubey, Hajime Fujita 0002, Daniel T. Graves, Andrew A. Chien, Devesh Tiwari
SC1
2015 Ongoing verification of a multiphysics community code: FLASH
abstract
SUMMARY When developing a complex, multi‐authored code, daily testing on multiple platforms and under a variety of conditions is essential. It is therefore necessary to have a regression test suite that is easily administered and configured, as well as a way to easily view and interpret the test suite results. We describe the methodology for verification of FLASH, a highly capable multiphysics scientific application code with a wide user base. The methodology uses a combination of unit and regression tests and an in‐house testing software that is optimized for operation under limited resources. Although our practical implementations do not always comply with theoretical regression‐testing research, our methodology provides a comprehensive verification of a large scientific code under resource constraints.Copyright © 2013 John Wiley & Sons, Ltd.
Anshu Dubey, Klaus Weide, Dongwook Lee 0004, John Bachan, Christopher S. Daley, Samuel Olofin, Noel T. Taylor, Paul M. Rich, Lynn B. Reid
Softw. Pract. Exp.1
2014 A survey of high level frameworks in block-structured adaptive mesh refinement packages
Anshu Dubey, Ann S. Almgren, John B. Bell, Martin Berzins, Steven R. Brandt, Greg Bryan, Phillip Colella, Daniel T. Graves, Michael Lijewski, Frank Löffler 0001, Brian W. O'Shea, Erik Schnetter, Brian van Straalen, Klaus Weide
J. Parallel Distributed Comput.1
2013 Parallel Algorithms for Using Lagrangian Markers in Immersed Boundary Method with Adaptive Mesh Refinement in FLASH
abstract
Computational fluid dynamics (CFD) are at the forefront of computational mechanics in requiring large-scale computational resources associated with high performance computing (HPC). Many flows of practical interest also include moving and deforming boundaries. High fidelity computations of fluid-structure interactions (FSI) are amongst the most challenging problems in computational mechanics. Additionally, many FSI applications have different resolution requirements in different parts of the domain and therefore requirement adaptive mesh refinement (AMR) for computational efficiency. FLASH is a well established AMR code with an existing Lagrangian framework which could be augmented and exploited to implement an immersed boundary method for simulating fluid-structure interactions atop an existing infrastructure. This paper describes the augmentations to the Lagrangian framework, and the new parallel algorithms added to the FLASH infrastructure that enabled the implementation of immersed boundary method in FLASH. The paper also presents scaling behavior and performance analysis of the implementations.
Prateeti Mohapatra, Anshu Dubey, Christopher S. Daley, Marcos Vanella, Elias Balaras
SBAC-PAD2
2012 Optimization of multigrid based elliptic solver for large scale simulations in the FLASH code
abstract
SUMMARY FLASH is a multiphysics multiscale adaptive mesh refinement (AMR) code originally designed for simulation of reactive flows often found in Astrophysics. With its wide user base and flexible applications configuration capability, FLASH has a dual task of maintaining scalability and portability in all its solvers. The scalability of fully explicit solvers in the code is tied very closely to that of the underlying mesh. Others such as the Poisson solver based on a multigrid method have more complex scaling behavior. Multigrid methods suffer from processor starvation and dominating communication costs at coarser grids with increase in the number of processors. In this paper, we propose a combination of uniform grid mesh with AMR mesh, and the merger of two different sets of solvers to overcome the scalability limitation of the Poisson solver in FLASH. The principal challenge in the proposed merger is the efficiency of the communication algorithm to map the mesh back and forth between uniform grid and AMR. We present two different parallel mapping algorithms and also discuss results from performance studies of the two implementations. Copyright © 2012 John Wiley & Sons, Ltd.
Christopher S. Daley, Marcos Vanella, Anshu Dubey, Klaus Weide, Elias Balaras
Concurr. Comput. Pract. Exp.3
2011 Parallel algorithms for moving Lagrangian data on block structured Eulerian meshes
Anshu Dubey, Katie Antypas, Christopher S. Daley
Parallel Comput.1
2009 Extensible component-based architecture for FLASH, a massively parallel, multiphysics simulation code
Anshu Dubey, Katie Antypas, Murali K. Ganapathy, Lynn B. Reid, Katherine Riley, Daniel J. Sheeler, Andrew R. Siegel, Klaus Weide
Parallel Comput.1
2001 Redistribution strategies for portable parallel FFT: a case study
abstract
Abstract The best approach to parallelize multidimensional FFT algorithms has long been under debate. Distributed transposes are widely used, but they also vary in communication policies and hence performance. In this work we analyze the impact of different redistribution strategies on the performance of parallel FFT, on various machine architectures. We found that some redistribution strategies were consistently superior, while some others were unexpectedly inferior. An in‐depth investigation into the reasons for this behavior is included in this work. Copyright © 2001 John Wiley & Sons, Ltd.
Anshu Dubey, Daniele Tessera
Concurr. Comput. Pract. Exp.1
1994 A General Purpose Subroutine for Fast Fourier Transform on a Distributed Memory Parallel Machine
Anshu Dubey, Mohammad Zubair, Chester E. Grosch
Parallel Comput.1