Rashid Kaleem

dblp:38/9698 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-authorSoftware engineering, systems software and programming languages · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 58% High-performance computing · 42%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%
Artificial intelligence
1 paper
Graph learning · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
collective communication
0.312018
Framework for scalable intra-node collective operations using shared memory · SC 2018
Parallel and multicore computing › parallel computing › parallel communication
shared-memory communication
0.312018
Framework for scalable intra-node collective operations using shared memory · SC 2018
Parallel and multicore computing
parallel programming models
0.112011
The tao of parallelism in algorithms · PLDI 2011
Machine learning › Graph learning
graph algorithms
0.012011
The tao of parallelism in algorithms · PLDI 2011

Methods — techniques the papers use, named apart from their topics

dependence graph analysis · 0.4
YearPublicationVenuePosition
2018 Framework for scalable intra-node collective operations using shared memory
Surabhi Jain, Rashid Kaleem, Marc Gamell, Akhil Langer, Dmitry Durnov, Alexander Sannikov, María Jesús Garzarán
SC2
2016 Synchronization Trade-Offs in GPU Implementations of Graph Algorithms
abstract
Although there is an extensive literature on GPU implementations of graph algorithms, we do not yet have a clear understanding of how implementation choices impact performance. As a step towards this goal, we studied how the choice of synchronization mechanism affects the end-to-end performance of complex graph algorithms, using stochastic gradient descent (SGD) as an exemplar. We implemented seven synchronization strategies for this application and evaluated them on two GPU platforms, using both road networks and social network graphs as inputs. Our experiments showed that although none of the seven strategies dominates the rest, it is possible to use properties of the platform and input graph to predict the best strategy.
Rashid Kaleem, Anand Venkat, Sreepathi Pai, Mary W. Hall, Keshav Pingali
IPDPS1
2014 Adaptive heterogeneous scheduling for integrated GPUs
abstract
Many processors today integrate a CPU and GPU on the same die, which allows them to share resources like physical memory and lowers the cost of CPU-GPU communication. As a consequence, programmers can effectively utilize both the CPU and GPU to execute a single application. This paper presents novel adaptive scheduling techniques for integrated CPU-GPU processors. We present two online profiling-based scheduling algorithms: naïve and asymmetric. Our asymmetric scheduling algorithm uses low-overhead online profiling to automatically partition the work of data-parallel kernels between the CPU and GPU without input from application developers. It does profiling on the CPU and GPU in a way that it doesn't penalize GPU-centric workloads that run significantly faster on the GPU. It adapts to application characteristics by addressing: 1) load imbalance via irregularity caused by, e.g., data-dependent control flow, 2) different amounts of work on each kernel call, and 3) multiple kernels with different characteristics. Unlike many existing approaches primarily targeting NVIDIA discrete GPUs, our scheduling algorithm does not require offline processing.
Rashid Kaleem, Rajkishore Barik, Tatiana Shpeisman, Brian T. Lewis, Chunling Hu, Keshav Pingali
PACT1
2014 Efficient Mapping of Irregular C++ Applications to Integrated GPUs
Rajkishore Barik, Rashid Kaleem, Deepak Majeti, Brian T. Lewis, Tatiana Shpeisman, Chunling Hu, Ali-Reza Adl-Tabatabai
CGO2
2011 The tao of parallelism in algorithms
abstract
For more than thirty years, the parallel programming community has used the dependence graph as the main abstraction for reasoning about and exploiting parallelism in "regular" algorithms that use dense arrays, such as finite-differences and FFTs. In this paper, we argue that the dependence graph is not a suitable abstraction for algorithms in new application areas like machine learning and network analysis in which the key data structures are "irregular" data structures like graphs, trees, and sets.
Keshav Pingali, Donald Nguyen, Milind Kulkarni 0001, Martin Burtscher, Muhammad Amber Hassaan, Rashid Kaleem, Tsung-Hsien Lee, Andrew Lenharth, Roman Manevich, Mario Méndez-Lojo, Dimitrios Prountzos
PLDI6