Kapil K. Mathur

dblp:07/6620 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
0since 2021 · last 1994
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 68% High-performance computing · 32%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel programming models
data parallelization
0.011989
Element order and convergence rate of the conjugate gradient method for data parallel stress analysis · SC 1989
Parallel and multicore computing
data-parallel programming
0.011989
Matrix multiplication on the connection machine · SC 1989
High-performance computing › numerical linear algebra
matrix multiplication
0.011989
Matrix multiplication on the connection machine · SC 1989
Parallel and multicore computing
parallel computing
0.011989
Matrix multiplication on the connection machine · SC 1989
Parallel and multicore computing › parallel algorithms › parallel algorithm design
systolic algorithms
0.011989
Matrix multiplication on the connection machine · SC 1989
Parallel and multicore computing › array processor
connection machine
0.011989
Matrix multiplication on the connection machine · SC 1989
Mathematical optimization › iterative methods
conjugate gradient method
0.011989
Element order and convergence rate of the conjugate gradient method for data parallel stress analysis · SC 1989

Methods — techniques the papers use, named apart from their topics

finite element method · 0.0diagonal preconditioning · 0.0conjugate gradient · 0.0systolic algorithm · 0.0matrix-vector multiplication primitive · 0.0
YearPublicationVenuePosition
1994 Multiplication of Matrices of Arbitrary Shape on a Data Parallel Computer
Kapil K. Mathur, S. Lennart Johnsson
Parallel Comput.1
1989 Matrix multiplication on the connection machine
abstract
A data parallel implementation of the multiplication of matrices of arbitrary shapes and sizes is presented. A systolic algorithm based on a rectangular processor layout is used by the implementation. All processors contain submatrices of the same size for a given operand. Matrix-vector multiplication is used as a primitive for local matrix-matrix multiplication in the Connection Machine system CM-2 implementation. The peak performance of the local matrix-matrix multiplication is in excess of 20 Gflops s-1. The overall algorithm including all required data motion has a peak performance of 5.8 Gflops s-1.
S. Lennart Johnsson, Tim Harris 0002, Kapil K. Mathur
SC3
1989 Element order and convergence rate of the conjugate gradient method for data parallel stress analysis
abstract
A data parallel formulation of the finite element method is described. The data structures and the algorithms for stiffness matrix generation and the solution of the equilibrium equations are presented briefly. The generation of the elemental stiffness matrices requires no communication, even though each finite element is distributed over several processors. The conjugate gradient method with a diagonal preconditioner has been used for the solution of the resulting sparse linear system. This formulation has been implemented on the Connection Machine® model CM-2. The simulations reported in this article investigate the influence of the mesh discretization and the interpolation order on the convergence behavior of the conjugate gradient method. A linear dependence of the convergence behavior on the mesh discretization parameter is observed. In addition, the convergence rate depends on the interpolation order p as Ο(p1.6). The peak floating point rate (single-precision) for the evaluation of the stiffness matrix is approximately 2.4 Gflops s-1. The iterative solver peaks at nearly 850 Mflops s-1.
Kapil K. Mathur, S. Lennart Johnsson
SC1