Dinesh K. Kaushik

dblp:63/3892 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Parallel and multicore computing · 46% High-performance computing · 45% Memory systems · 5%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 87% Computational science and engineering · 13%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › sequence analysis
genomic sequence analysis
0.112011
Highly scalable ab initio genomic motif identification · SC 2011
Bioinformatics and computational biology › sequence analysis
motif discovery
0.112011
Highly scalable ab initio genomic motif identification · SC 2011
Parallel and multicore computing › parallel programming models › hybrid programming models
hybrid MPI/OpenMP
0.112011
Highly scalable ab initio genomic motif identification · SC 2011
Parallel and multicore computing › parallel computing
scalable parallel computing
0.112011
Highly scalable ab initio genomic motif identification · SC 2011
High-performance computing
performance optimization at scale
0.122009
Enabling high-fidelity neutron transport simulations on petascale architectures · SC 2009
Achieving High Sustained Performance in an Unstructured Mesh CFD Application · SC 1999
High-performance computing
scientific computing systems
0.122009
Enabling high-fidelity neutron transport simulations on petascale architectures · SC 2009
Achieving High Sustained Performance in an Unstructured Mesh CFD Application · SC 1999
Parallel and multicore computing
parallel programming models
0.132011
Highly scalable ab initio genomic motif identification · SC 2011
Performance Modeling and Tuning of an Unstructured Mesh CFD Application · SC 2000
Achieving High Sustained Performance in an Unstructured Mesh CFD Application · SC 1999
High-performance computing
performance optimization
0.012000
Performance Modeling and Tuning of an Unstructured Mesh CFD Application · SC 2000
Performance modeling and evaluation
performance tuning
0.012000
Performance Modeling and Tuning of an Unstructured Mesh CFD Application · SC 2000
High-performance computing › scientific computing systems
computational fluid dynamics
0.011999
Achieving High Sustained Performance in an Unstructured Mesh CFD Application · SC 1999
Memory systems › memory hierarchy
memory hierarchy optimization
0.011999
Achieving High Sustained Performance in an Unstructured Mesh CFD Application · SC 1999
Computational science and engineering
computational fluid dynamics
0.012000
Performance Modeling and Tuning of an Unstructured Mesh CFD Application · SC 2000
Memory systems
data layout optimization
0.012000
Performance Modeling and Tuning of an Unstructured Mesh CFD Application · SC 2000
High-performance computing
unstructured mesh computation
0.011999
Achieving High Sustained Performance in an Unstructured Mesh CFD Application · SC 1999

Methods — techniques the papers use, named apart from their topics

master-slave work assignment · 0.2OpenMP · 0.2MPI · 0.2weighted partitioning · 0.2spatial multigrid preconditioner · 0.2matrix-tensor operation optimization · 0.2sparse matrix-vector product analysis · 0.1performance modeling · 0.1implicit unstructured grid simulation · 0.0data reuse optimization · 0.0
YearPublicationVenuePosition
2016 Unstructured computational aerodynamics on many integrated core architecture
Mohammed A. Al Farhan, Dinesh K. Kaushik, David E. Keyes
Parallel Comput.2
2015 Exploring Shared-Memory Optimizations for an Unstructured Mesh CFD Application on Modern Parallel Systems
abstract
In this work, we revisit the 1999 Gordon Bell Prize winning PETSc-FUN3D aerodynamics code, extending it with highly-tuned shared-memory parallelization and detailed performance analysis on modern highly parallel architectures. An unstructured-grid implicit flow solver, which forms the backbone of computational aerodynamics, poses particular challenges due to its large irregular working sets, unstructured memory accesses, and variable/limited amount of parallelism. This code, based on a domain decomposition approach, exposes tradeoffs between the number of threads assigned to each MPI-rank sub domain, and the total number of domains. By applying several algorithm- and architecture-aware optimization techniques for unstructured grids, we show a 6.9X speed-up in performance on a single-node Intel® XeonTM1 E5 2690 v2 processor relative to the out-of-the-box compilation. Our scaling studies on TACC Stampede supercomputer show that our optimizations continue to provide performance benefits over baseline implementation as we scale up to 256 nodes.
Dheevatsa Mudigere, Srinivas Sridharan 0002, Anand M. Deshpande, Jongsoo Park, Alexander Heinecke, Mikhail Smelyanskiy, Bharat Kaul, Pradeep Dubey, Dinesh K. Kaushik, David E. Keyes
IPDPS9
2011 Highly scalable ab initio genomic motif identification
abstract
We present results of scaling an ab initio motif family identification system, Dragon Motif Finder (DMF), to 65,536 processor cores of IBM Blue Gene/P. DMF seeks groups of mutually similar polynucleotide patterns within a set of genomic sequences and builds various motif families from them. Such information is of relevance to many problems in life sciences. Prior attempts to scale such ab initio motif-finding algorithms achieved limited success. We solve the scalability issues using a combination of mixed-mode MPI-OpenMP parallel programming, master-slave work assignment, multi-level workload distribution, multi-level MPI collectives, and serial optimizations. While the scalability of our algorithm was excellent (94% parallel efficiency on 65,536 cores relative to 256 cores on a modest-size problem), the final speedup with respect to the original serial code exceeded 250,000 when serial optimizations are included. This enabled us to carry out many large-scale ab initio motif-finding simulations in a few hours while the original serial code would have needed decades of execution time.
Benoit Marchand, Vladimir B. Bajic, Dinesh K. Kaushik
SC3
2009 Enabling high-fidelity neutron transport simulations on petascale architectures
abstract
The UNIC code is being developed as part of the DOE's Nuclear Energy Advanced Modeling and Simulation (NEAMS) program. UNIC is an unstructured, deterministic neutron transport code that allows a highly detailed description of a nuclear reactor. The primary goal of our simulation efforts is to reduce the uncertainties and biases in reactor design calculations by progressively replacing existing multilevel averaging (homogenization) techniques with more direct solution methods based on first principles. Since the neutron transport equation is seven dimensional (three in space, two in angle, one in energy, and one in time), these simulations are among the most memory and computationally intensive in all of computational science. In order to model the complex physics of a reactor core, billions of spatial elements, hundreds of angles, and thousands of energy groups are necessary, leading to problem sizes with petascale degrees of freedom. Therefore, these calculations exhaust memory resources on current and even next-generation architectures. In this paper, we present UNIC simulation results for two important representative problems in reactor design and analysis---PHENIX and ZPR-6. In each case, UNIC shows good weak scalability on up to 163,840 cores of Blue Gene/P (Argonne) and 122,800 cores of XT5 (Oak Ridge). While our current per processor performance is less than ideal, we demonstrate a clear ability to effectively utilize the leadership computing platforms. Over the coming months, we aim to improve the per processor performance while maintaining the high parallel efficiency by employing better algorithms such as spatial p- and h-multigrid preconditioners, optimized matrix-tensor operations, and weighted partitioning for better load balancing. Combining these additional algorithmic improvements with the availability of larger parallel machines should allow us to realize our long-term goal of explicit geometry coupled multiphysics reactor simulations. In the long run, these high-fidelity simulations will be able to replace expensive mockup experiments and reduce the uncertainty in crucial reactor design and operational parameters.
Dinesh K. Kaushik, Micheal Smith, Allan B. Wollaber, Barry Smith 0002, Andrew R. Siegel, Won Sik Yang
SC1
2008 Improving the Performance of Tensor Matrix Vector Multiplication in Cumulative Reaction Probability Based Quantum Chemistry Codes
Dinesh K. Kaushik, William Gropp, Michael Minkoff, Barry Smith 0002
HiPC1
2001 A Scientific Data Management System for Irregular Applications
abstract
Many scientific applications are I/O intensive and generate large data sets, spanning hundreds or thousands of "files." Management, storage, efficient access, and analysis of this data present an extremely challenging task. We have developed a software system, called Scientific Data Manager (SDM), that uses a combination of parallel file I/O and database support for high-performance scientific data management. SDM provides a high-level API to the user and, internally, uses a parallel file system to store real data and a database to store application-related metadata. In this paper, we describe how we designed and implemented SDM to support irregular applications. SDM can efficiently handle the reading and writing of data in an irregular mesh, as well as the distribution of index values. We describe the SDM user interface and how we have implemented it to achieve high performance. SDM makes extensive use of MPI-IO's noncontiguous collective I/O functions. SDM also uses the concept of a history file to optimize the cost of the index distribution using the metadata stored in database. We present performance results with two irregular applications, a CFD code called FUN3D and a Rayleigh-Taylor instability code, on the SGI Origin2000 at Argonne National Laboratory.
Jaechun No, Rajeev Thakur, Dinesh K. Kaushik, Lori A. Diachin, Alok N. Choudhary
IPDPS3
2001 High-performance parallel implicit CFD
William Gropp, Dinesh K. Kaushik, David E. Keyes, Barry Smith 0002
Parallel Comput.2
2000 Analyzing the Parallel Scalability of an Implicit Unstructured Mesh CFD Code
William Gropp, Dinesh K. Kaushik, Barry Smith 0002, David E. Keyes
HiPC2
2000 Performance Modeling and Tuning of an Unstructured Mesh CFD Application
abstract
This paper describes performance tuning experiences with a three-dimensional unstructured grid Euler flow code from NASA, which we have reimplemented in the PETSc framework and ported to several large-scale machines, including the ASCI Red and Blue Pacific machines, the SGI Origin, the Cray T3E, and Beowulf clusters. The code achieves a respectable level of performance for sparse problems, typical of scientific and engineering codes based on partial differential equations, and scales well up to thousands of processors. Since the gap between CPU speed and memory access rate is widening, the code is analyzed from a memory-centric perspective (in contrast to traditional flop-orientation) to understand its sequential and parallel performance. Performance tuning is approached on three fronts: data layouts to enhance locality of reference, algorithmic parameters, and parallel programming model. This effort was guided partly by some simple performance models developed for the sparse matrix-vector product operation.
William Gropp, Dinesh K. Kaushik, David E. Keyes, Barry Smith 0002
SC2
1999 Achieving High Sustained Performance in an Unstructured Mesh CFD Application
abstract
This paper highlights a three-year project by an interdisciplinary team on a legacy F77 computational fluid dynamics code, with the aim of demonstrating that implicit unstructured grid simulations can execute at rates not far from those of explicit structured grid codes, provided attention is paid to data motion complexity and the reuse of data positioned at the levels of the memory hierarchy closest to the processor, in addition to traditional operation count complexity. The demonstration code is from NASA and the enabling parallel hardware and (freely available) software toolkit are from DOE, but the resulting methodology should be broadly applicable, and the hardware limitations exposed should allow programmers and vendors of parallel platforms to focus with greater encouragement on sparse codes with indirect addressing. This snapshot of ongoing work shows a performance of 15 microseconds per degree of freedom to steady-state convergence of Euler flow on a mesh with 2.8 million vertices using 3072 dual-processor nodes of ASCI Red, corresponding to a sustained floating-point rate of 0.227 Tflop/s.
W. K. Anderson, William Gropp, Dinesh K. Kaushik, David E. Keyes, Barry Smith 0002
SC3