Dhananjay Brahme

dblp:03/3267 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Electronic design automation · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation › hardware verification and test › design for testability
built-in self-test
0.011984
Functional Testing of Microprocessors · IEEE Trans. Computers 1984
Electronic design automation › hardware verification and test › test generation
functional test generation
0.011984
Functional Testing of Microprocessors · IEEE Trans. Computers 1984
Electronic design automation
hardware verification and test
0.011984
Functional Testing of Microprocessors · IEEE Trans. Computers 1984
Electronic design automation › hardware verification and test › VLSI testing
microprocessor testing
0.011984
Functional Testing of Microprocessors · IEEE Trans. Computers 1984

Methods — techniques the papers use, named apart from their topics

reduced graph model · 0.0fault classification · 0.0
YearPublicationVenuePosition
2013 SymSig: A low latency interconnection topology for HPC clusters
abstract
This paper presents the underlying theory and the performance of a cluster using a new 2-hop network topology. This topology is constructed using a symmetric equation and Singer Difference Sets and is called SymSig. The degree of connections at each node with SymSig is about half compared to previous methods using Singer Difference Sets. A comparison with a cluster of Clos topology shows significant advantages. The worst case congestion in SymSig topology for unicast permutation is 2, where as in Clos it is proportional to the radix of the building block switches used. The number of switches required is smaller by about 25%, the size of the cluster is larger by about 15% and the worst bandwidth is better by about 50% for SymSig. These advantages are retained for peta and exascale systems. Its performance on a set of collectives like exchange-all, shift-all, broadcast-all and all-to-all send/receive shows improvements ranging from 39% to 83%. Its performance on a molecular dynamics application GROMMACS shows improvement of upto 33%. This network is particularly suitable for applications that require global all to all communications. The low latency of this network makes it scaleable and an attractive alternative for building peta and exascale systems.
Dhananjay Brahme, Onkar Bhardwaj, Vipin Chaudhary
HiPC1
2010 Parallel Sparse Matrix Vector Multiplication using greedy extraction of boxes
abstract
Parallel Sparse Matrix Vector Multiplication (PSpMV) is a compute intensive kernel used in iterative solvers like Conjugate Gradient, GMRES and Lanzcos. Numerous attempts at optimizing this function have been made that require fine tuning of many hardware and software parameters to achieve optimal performance. We attempt to offer a simple framework that involves (i) Employing a greedy algorithm to extract variable-sized dense sub matrices without zeroes filled in, (ii) Partitioning the sparse matrix in a load balanced manner and maintaining partial information at each node, and (iii) Overlapping communication with computation. Using the aforementioned, we reduce memory traffic and hide communication latencies, and hope to inherently achieve improved cache and register utilization. This paper reports the performance improvements of PSpMV as such and when used in Preconditioned Conjugate Gradient (PCG).
Dhananjay Brahme, Binit Ranjan Mishra, Anup Barve
HiPC1
1984 Functional Testing of Microprocessors
abstract
This paper presents a new and systematic method to generate tests for microprocessors. A functional level model for the microprocessor is used and it is represented by a reduced graph. A new and comprehensive model of the instruction execution process is developed. Various types of faults are analyzed and it is shown that with the use of appropriate codewords all faults can be classified into three types. This gives rise to a systematic procedure to generate tests which is independent of the microprocessor implementation details. Tests are given to detect faults in any microprocessor, first for the READ register instructions, and then for the remaining instructions. These tests can be executed by the microprocessor in a self-test mode, thus dispensing with the need for an external tester.
Dhananjay Brahme, Jacob A. Abraham
IEEE Trans. Computers1