VLDB 2026 Research / reviewers in the wild / expert
Jayvant Anantpur
dblp:39/9698
· DBLP profile ↗
8ranked-venue papers
4as first author
0since 2021 · last 2018
0000-0003-3353-0625ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-authorSoftware engineering, systems software and programming languages · 4 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
GPUs and heterogeneous computing · 76% Distributed systems · 12% Processor architecture and microarchitecture · 12% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 100% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization › accelerator compilation
GPU compiler optimization |
0.3 | 1 | 2017 | Scratchpad Sharing in GPUs · ACM Trans. Archit. Code Optim. 2017 |
GPUs and heterogeneous computing
GPU memory management |
0.3 | 1 | 2017 | Scratchpad Sharing in GPUs · ACM Trans. Archit. Code Optim. 2017 |
GPUs and heterogeneous computing › GPU scheduling
GPU thread scheduling |
0.3 | 1 | 2017 | Scratchpad Sharing in GPUs · ACM Trans. Archit. Code Optim. 2017 |
GPUs and heterogeneous computing › GPU scheduling
warp scheduling |
0.3 | 1 | 2017 | Scratchpad Sharing in GPUs · ACM Trans. Archit. Code Optim. 2017 |
GPUs and heterogeneous computing
GPU performance optimization |
0.2 | 1 | 2016 | Improving GPU Performance Through Resource Sharing · HPDC 2016 |
Distributed systems
resource sharing |
0.2 | 1 | 2016 | Improving GPU Performance Through Resource Sharing · HPDC 2016 |
Processor architecture and microarchitecture › many-core architecture
streaming multiprocessor |
0.2 | 1 | 2016 | Improving GPU Performance Through Resource Sharing · HPDC 2016 |
GPUs and heterogeneous computing › GPU scheduling
thread block scheduling |
0.2 | 1 | 2016 | Improving GPU Performance Through Resource Sharing · HPDC 2016 |
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
CPU-GPU co-processing |
0.1 | 1 | 2011 | Automatic compilation of MATLAB programs for synergistic execution on heterogeneous processors · PLDI 2011 |
GPUs and heterogeneous computing › GPU performance analysis
GPU utilization |
0.1 | 1 | 2016 | Improving GPU Performance Through Resource Sharing · HPDC 2016 |
Compilers and program optimization › parallelization
data parallelism |
0.0 | 1 | 2011 | Automatic compilation of MATLAB programs for synergistic execution on heterogeneous processors · PLDI 2011 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.6shared memory allocation · 0.2register allocation · 0.2data-parallel mapping · 0.2automatic compilation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Reducing GPU Register File Energy
Vishwesh Jatala, Jayvant Anantpur, Amey Karkare |
Euro-Par | 2 |
| 2017 | Taming warp divergence
Jayvant Anantpur, R. Govindarajan |
CGO | 1 |
| 2017 | Scratchpad Sharing in GPUsabstractGeneral-Purpose Graphics Processing Unit (GPGPU) applications exploit on-chip scratchpad memory available in the Graphics Processing Units (GPUs) to improve performance. The amount of thread level parallelism (TLP) present in the GPU is limited by the number of resident threads, which in turn depends on the availability of scratchpad memory in its streaming multiprocessor (SM). Since the scratchpad memory is allocated at thread block granularity, part of the memory may remain unutilized. In this article, we propose architectural and compiler optimizations to improve the scratchpad memory utilization. Our approach, called Scratchpad Sharing , addresses scratchpad under-utilization by launching additional thread blocks in each SM. These thread blocks use unutilized scratchpad memory and also share scratchpad memory with other resident blocks. To improve the performance of scratchpad sharing, we propose Owner Warp First (OWF) scheduling that schedules warps from the additional thread blocks effectively. The performance of this approach, however, is limited by the availability of the part of scratchpad memory that is shared among thread blocks. We propose compiler optimizations to improve the availability of shared scratchpad memory. We describe an allocation scheme that helps in allocating scratchpad variables such that shared scratchpad is accessed for short duration. We introduce a new hardware instruction, relssp , that when executed releases the shared scratchpad memory. Finally, we describe an analysis for optimal placement of relssp instructions, such that shared scratchpad memory is released as early as possible, but only after its last use, along every execution path. We implemented the hardware changes required for scratchpad sharing and the relssp instruction using the GPGPU-Sim simulator and implemented the compiler optimizations in Ocelot framework. We evaluated the effectiveness of our approach on 19 kernels from 3 benchmarks suites: CUDA-SDK, GPGPU-Sim, and Rodinia. The kernels that under-utilize scratchpad memory show an average improvement of 19% and maximum improvement of 92.17% in terms of the number of instruction executed per cycle when compared to the baseline approach, without affecting the performance of the kernels that are not limited by scratchpad memory. Vishwesh Jatala, Jayvant Anantpur, Amey Karkare |
ACM Trans. Archit. Code Optim. | 2 |
| 2016 | Improving GPU Performance Through Resource SharingabstractGraphics Processing Units (GPUs) consisting of Streaming Multiprocessors (SMs) achieve high throughput by running a large number of threads and context switching among them to hide execution latencies. The number of thread blocks, and hence the number of threads that can be launched on an SM, depends on the resource usage--e.g. number of registers, amount of shared memory--of the thread blocks. Since the allocation of threads to an SM is at the thread block granularity, some of the resources may not be used up completely and hence will be wasted. Vishwesh Jatala, Jayvant Anantpur, Amey Karkare |
HPDC | 2 |
| 2015 | PRO: Progress Aware GPU Warp Scheduling AlgorithmabstractGraphics Processing Units (GPUs) contain multiple SIMD cores and each core can run a large number of threads concurrently. Threads in a core are scheduled and executed infixed sized groups, called warps. Each core contains one or more warp schedulers that select and execute warps from a pool of ready warps. In spite of having a large number of concurrent warps - 48 on NVIDIA Fermi architecture GPU - on manyGPGPU applications, current warp scheduling algorithms can not effectively utilize the hardware resources, resulting in stall cycles and loss in performance. The main reason for this is current warp scheduling algorithms mostly focus on long latency operations, especially global memory accesses, and do not take into account factors such as the progress of each thread block and the number of ready warps. In this paper, we propose, PRO, a progress warp scheduling algorithm that not only focuses on finishing individual thread blocks faster but also on reducing the overall execution time. These goals are achieved by dynamically prioritizing thread blocks and warps, based on their progress. We implemented our proposed algorithm in the GPGPU-SIM simulator and evaluated on various applications from GPGPU-SIM, Rosina and CUDASDK benchmark suites. We achieved an average speedup of 1.12x and a maximum speedup of 1.94x over the commonly used LooseRound Robin warp scheduling algorithm. Over the Two Level warp scheduler, our algorithm showed an average speedup of1.13x and a maximum speedup of 1.6x. Our proposed solution requires only a very small increase in the GPU hardware. Jayvant Anantpur, R. Govindarajan |
IPDPS | 1 |
| 2014 | Taming Control Divergence in GPUs through Control Flow LinearizationabstractBranch divergence is a very commonly occurring performance problem in GPGPU in which the execution of diverging branches is serialized to execute only one control flow path at a time. Existing hardware mechanism to reconverge threads using a stack causes duplicate execution of code for unstructured control flow graphs. Also the stack mechanism cannot effectively utilize the available parallelism among diverging branches. Further, the amount of nested divergence allowed is also limited by depth of the branch divergence stack. In this paper we propose a simple and elegant transformation to handle all of the above mentioned problems. The transformation converts an unstructured CFG to a structured CFG without duplicating user code. It incurs only a linear increase in the number of basic blocks and also the number of instructions. Our solution linearizes the CFG using a predicate variable. This mechanism reconverges the divergent threads as early as possible. It also reduces the depth of the reconvergence stack. The available parallelism in nested branches can be effectively extracted by scheduling the basic blocks to reduce the effect of stalls due to memory accesses. It can also increase execution efficiency of nested loops with different trip counts for different threads. We implemented the proposed transformation at PTX level using the Ocelot compiler infrastructure. We evaluated the technique using various benchmarks to show that it can be effective in handling the performance problem due to divergence in unstructured CFGs. Jayvant Anantpur, R. Govindarajan |
CC | 1 |
| 2013 | Runtime dependence computation and execution of loops on heterogeneous systemsabstractGPUs have been used for parallel execution of DOALL loops. However, loops with indirect array references can potentially cause cross iteration dependences which are hard to detect using existing compilation techniques. Applications with such loops cannot easily use the GPU and hence do not benefit from the tremendous compute capabilities of GPUs. In this paper, we present an algorithm to compute at runtime the cross iteration dependences in such loops. The algorithm uses both the CPU and the GPU to compute the dependences. Specifically, it effectively uses the compute capabilities of the GPU to quickly collect the memory accesses performed by the iterations by executing the slice functions generated for the indirect array accesses. Using the dependence information, the loop iterations are levelized such that each level contains independent iterations which can be executed in parallel. Another interesting aspect of the proposed solution is that it pipelines the dependence computation of the future level with the actual computation of the current level to effectively utilize the resources available in the GPU. We use NVIDIA Tesla C2070 to evaluate our implementation using benchmarks from Polybench suite and some synthetic benchmarks. Our experiments show that the proposed technique can achieve an average speedup of 6.4x on loops with a reasonable number of cross iteration dependences. Jayvant Anantpur, R. Govindarajan |
CGO | 1 |
| 2011 | Automatic compilation of MATLAB programs for synergistic execution on heterogeneous processorsabstractMATLAB is an array language, initially popular for rapid prototyping, but is now being increasingly used to develop production code for numerical and scientific applications. Typical MATLAB programs have abundant data parallelism. These programs also have control flow dominated scalar regions that have an impact on the program's execution time. Today's computer systems have tremendous computing power in the form of traditional CPU cores and throughput oriented accelerators such as graphics processing units(GPUs). Thus, an approach that maps the control flow dominated regions to the CPU and the data parallel regions to the GPU can significantly improve program performance. Ashwin Prasad, Jayvant Anantpur, R. Govindarajan |
PLDI | 2 |