VLDB 2026 Research / reviewers in the wild / expert
Abhinav Jangda
dblp:168/8310
· DBLP profile ↗
15ranked-venue papers
11as first author
8since 2021 · last 2026
0000-0002-4849-6776ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 7 first-author · 5 since 2021Software engineering, systems software and programming languages · 9 · 5 first-author · 6 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MSCCL++: Rethinking GPU Communication Abstractions for AI InferenceabstractAI applications increasingly run on fast-evolving, heterogeneous hardware to maximize performance, but general-purpose libraries lag in supporting these features. Performance-minded programmers often build custom communication stacks that are fast but error-prone and non-portable. This paper introduces MSCCL++, a design methodology for developing high-performance, portable communication kernels. It provides (1) a low-level, performance-preserving primitive interface that exposes minimal hardware abstractions while hiding the complexities of synchronization and consistency, (2) a higher-level DSL for application developers to implement workload-specific communication algorithms, and (3) a library of efficient algorithms implementing the standard collective API, enabling adoption by users with minimal expertise. Compared to state-of-the-art baselines, MSCCL++ achieves geomean speedups of 1.7× (up to 5.4×) for collective communication and 1.2× (up to 1.38×) for AI inference workloads. MSCCL++ is in production of multiple AI services provided by Microsoft Azure, and has also been adopted by RCCL, the GPU collective communication library maintained by AMD. MSCCL++ is open source and available at https://github.com/microsoft/mscclpp. Our two years of experience with MSCCL++ suggests that its abstractions are robust, enabling support for new hardware features, such as multimem, within weeks of development. Changho Hwang, Peng Cheng 0005, Roshan Dathathri, Abhinav Jangda, Saeed Maleki, Madan Musuvathi, Olli Saarikivi, Aashaka Shah, Ziyue Yang 0002, Binyang Li, Caio Rocha, Mahdieh Ghazimirsaeed, Sreevatsa Anantharamu |
ASPLOS (2) | 4 |
| 2026 | Compiling Strassen-like Matrix Multiplication Algorithms to Fast CUDA KernelsabstractMatrix multiplication is a key operation in scientific computing and machine learning, with GPU libraries like NVIDIA Cutlass and cuBLAS providing optimized implementations of the three nested loop cubic algorithm. While sub-cubic algorithms, like the Strassen algorithm and its variants, are theoretically faster, their recursive structure makes it challenging to implement efficient GPU kernels. This is why existing approaches either do excessive memory accesses or do not effectively overlap memory accesses and computations, leading to sub-optimal performance compared to theoretical expectations. This paper presents SubCuber, a domain-specific compiler that generates efficient CUDA kernels for Strassen-like matrix multiplication algorithms. SubCuber contains two novel CUDA kernels that are designed to minimize memory input loads and effectively overlap computation with memory loads. To generate efficient code, SubCuber constructs the dependency graph of a Strassen-like algorithm, selects efficient kernel schedules, and applies fusion strategies tailored to a recursion level, matrix sizes, and GPU. Our evaluation on NVIDIA A100 and H200 GPUs shows that for both single- and half-precision floating point matrix multiplications, SubCuber’s generated code outperforms state-of-the-art CUDA implementations for matrix multiplication and the Strassen algorithm. SubCuber is up to 12% faster for one recursion level and 22% for two recursion levels over Cutlass and cuBLAS, while existing approaches are only up to 8% faster for one-level and 16% faster for two-levels. Furthermore, SubCuber makes matrix multiplication in language models like Phi-4 14B, Qwen-3 32B, and LLaMA-3 405b, up to 16% faster for inference scenarios. Abhinav Jangda |
Proc. ACM Program. Lang. | 1 |
| 2024 | A Framework for Fine-Grained Synchronization of Dependent GPU KernelsabstractMachine Learning (ML) models execute several parallel computations including Generalized Matrix Multiplication, Convolution, Dropout, etc. These computations are commonly executed on Graphics Processing Units (GPUs), by dividing the computation into independent processing blocks, known as tiles. Since the number of tiles are usually higher than the execution units of a GPU, tiles are executed on all execution units in one or more waves. However, the number of tiles is not always a multiple of the number of execution units. Thus, tiles executed in the final wave can under-utilize the GPU. To address this issue, we present cuSync, a framework for synchronizing dependent kernels using a user-defined fine-grained synchronization policy to improve the GPU utilization. cuSync synchronizes tiles instead of kernels, which allows executing independent tiles of dependent kernels concurrently. We also present a compiler to generate diverse fine-grained synchronization policies based on dependencies between kernels. Our experiments found that synchronizing CUDA kernels using cuSync reduces the inference times of four popular ML models: MegatronLM GPT-3 by up to 15%, LLaMA by up to 14%, ResNet-38 by up to 22%, and VGG-19 by up to 16% over several batch sizes. Abhinav Jangda, Saeed Maleki, Maryam Mehri Dehnavi, Madan Musuvathi, Olli Saarikivi |
CGO | 1 |
| 2024 | Fast Kronecker Matrix-Matrix Multiplication on GPUsabstractKronecker Matrix-Matrix Multiplication (Kron-Matmul) is the multiplication of a matrix with the Kronecker Product of several smaller matrices. Kron-Matmul is a core operation for many scientific and machine learning computations. State-of-the-art Kron-Matmul implementations utilize existing tensor algebra operations, such as matrix multiplication, transpose, and tensor matrix multiplication. However, this design choice prevents several Kron-Matmul specific optimizations, thus, leaving significant performance on the table. Abhinav Jangda |
PPoPP | 1 |
| 2024 | Knowledge Transfer from High-Resource to Low-Resource Programming Languages for Code LLMsabstractOver the past few years, Large Language Models of Code (Code LLMs) have started to have a significant impact on programming practice. Code LLMs are also emerging as building blocks for research in programming languages and software engineering. However, the quality of code produced by a Code LLM varies significantly by programming language. Code LLMs produce impressive results on high-resource programming languages that are well represented in their training data (e.g., Java, Python, or JavaScript), but struggle with low-resource languages that have limited training data available (e.g., OCaml, Racket, and several others). This paper presents an effective approach for boosting the performance of Code LLMs on low-resource languages using semi-synthetic data. Our approach, called M ulti PL-T, generates high-quality datasets for low-resource languages, which can then be used to fine-tune any pretrained Code LLM. M ulti PL-T translates training data from high-resource languages into training data for low-resource languages in the following way. 1) We use a Code LLM to synthesize unit tests for commented code from a high-resource source language, filtering out faulty tests and code with low test coverage. 2) We use a Code LLM to translate the code from the high-resource source language to a target low-resource language. This gives us a corpus of candidate training data in the target language, but many of these translations are wrong. 3) We use a lightweight compiler to compile the test cases generated in (1) from the source language to the target language, which allows us to filter our obviously wrong translations. The result is a training corpus in the target low-resource language where all items have been validated with test cases. We apply this approach to generate tens of thousands of new, validated training items for five low-resource languages: Julia, Lua, OCaml, R, and Racket, using Python as the source high-resource language. Furthermore, we use an open Code LLM (StarCoderBase) with open training data (The Stack), which allows us to decontaminate benchmarks, train models without violating licenses, and run experiments that could not otherwise be done. Using datasets generated with M ulti PL-T, we present fine-tuned versions of StarCoderBase and Code Llama for Julia, Lua, OCaml, R, and Racket that outperform other fine-tunes of these base models on the natural language to code task. We also present Racket fine-tunes for two very recent models, DeepSeek Coder and StarCoder2, to show that M ulti PL-T continues to outperform other fine-tuning approaches for low-resource languages. The M ulti PL-T approach is easy to apply to new languages, and is significantly more efficient and effective than alternatives such as training longer. Federico Cassano, John Gouwar, Francesca Lucchetti, Claire Schlesinger, Anders Freeman, Carolyn Jane Anderson, Molly Q. Feldman, Michael Greenberg 0002, Abhinav Jangda, Arjun Guha |
Proc. ACM Program. Lang. | 9 |
| 2023 | MultiPL-E: A Scalable and Polyglot Approach to Benchmarking Neural Code GenerationabstractLarge language models have demonstrated the ability to generate both natural language and programming language text. Although contemporary code generation models are trained on corpora with several programming languages, they are tested using benchmarks that are typically monolingual. The most widely used code generation benchmarks only target Python, so there is little quantitative evidence of how code generation models perform on other programming languages. We propose MultiPL-E, a system for translating unit test-driven code generation benchmarks to new languages. We create the first massively multilingual code generation benchmark by using MultiPL-E to translate two popular Python code generation benchmarks to 18 additional programming languages. We use MultiPL-E to extend the HumanEval benchmark [1] and MBPP benchmark [2] to 18 languages that encompass a range of programming paradigms and popularity. Using these new parallel benchmarks, we evaluate the multi-language performance of three state-of-the-art code generation models: Codex [1], CodeGen [3]and InCoder [4]. We find that Codex matches or even exceeds its performance on Python for several other languages. The range of programming languages represented in MultiPL-E allow us to explore the impact of language frequency and language features on model performance. Finally, the MultiPL-E approach of compiling code generation benchmarks to new programming languages is both scalable and extensible, making it straightforward to evaluate new models, benchmarks, and languages. Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen, Luna Phipps-Costin, Donald Pinckney, Ming-Ho Yee, Yangtian Zi, Carolyn Jane Anderson, Molly Q. Feldman, Arjun Guha, Michael Greenberg 0002, Abhinav Jangda |
IEEE Trans. Software Eng. | 13 |
| 2022 | Breaking the computation and communication abstraction barrier in distributed machine learning workloadsabstractRecent trends towards large machine learning models require both training and inference tasks to be distributed. Considering the huge cost of training these models, it is imperative to unlock optimizations in computation and communication to obtain best performance. However, the current logical separation between computation and communication kernels in machine learning frameworks misses optimization opportunities across this barrier. Breaking this abstraction can provide many optimizations to improve the performance of distributed workloads. However, manually applying these optimizations requires modifying the underlying computation and communication libraries for each scenario, which is both time consuming and error-prone. Abhinav Jangda, Amir Hossein Nodehi Sabet, Saeed Maleki, Youshan Miao, Madan Musuvathi, Todd Mytkowicz, Olli Saarikivi |
ASPLOS | 1 |
| 2021 | Accelerating graph sampling for graph machine learning using GPUsabstractRepresentation learning algorithms automatically learn the features of data. Several representation learning algorithms for graph data, such as DeepWalk, node2vec, and Graph-SAGE, sample the graph to produce mini-batches that are suitable for training a DNN. However, sampling time can be a significant fraction of training time, and existing systems do not efficiently parallelize sampling. Abhinav Jangda, Sandeep Polisetty, Arjun Guha, Marco Serafini |
EuroSys | 1 |
| 2020 | Model-Based Warp Overlapped Tiling for Image Processing Programs on GPUsabstractDomain-specific languages that execute image processing pipelines on GPUs, such as Halide and Forma, operate by 1)~dividing the image into overlapped tiles, and 2)~fusing loops to improve memory locality. However, current approaches have limitations: 1)~they require intra thread block synchronization, which has a nontrivial cost, 2)~they must choose between small tiles that require more overlapped computations or large tiles that increase shared memory access (and lowers occupancy), and 3) their autoscheduling algorithms use simplified GPU models that can result in inefficient global memory accesses. Abhinav Jangda, Arjun Guha |
PACT | 1 |
| 2020 | An Effective Fusion and Tile Size Model for PolyMageabstractEffective models for fusion of loop nests continue to remain a challenge in both general-purpose and domain-specific language (DSL) compilers. The difficulty often arises from the combinatorial explosion of grouping choices and their interaction with parallelism and locality. This article presents a new fusion algorithm for high-performance domain-specific compilers for image processing pipelines. The fusion algorithm is driven by dynamic programming and explores spaces of fusion possibilities not covered by previous approaches, and it is also driven by a cost function more concrete and precise in capturing optimization criteria than prior approaches. The fusion model is particularly tailored to the transformation and optimization sequence applied by PolyMage and Halide, two recent DSLs for image processing pipelines. Our model-driven technique when implemented in PolyMage provides significant improvements (up to 4.32×) over PolyMage’s approach (which uses auto-tuning to aid its model) and over Halide’s automatic approach (by up to 2.46×) on two state-of-the-art shared-memory multicore architectures. Abhinav Jangda, Uday Bondhugula |
ACM Trans. Program. Lang. Syst. | 1 |
| 2019 | Swizzle Inventor: Data Movement Synthesis for GPU KernelsabstractUtilizing memory and register bandwidth in modern architectures may require swizzles --- non-trivial mappings of data and computations onto hardware resources --- such as shuffles. We develop Swizzle Inventor to help programmers implement swizzle programs, by writing program sketches that omit swizzles and delegating their creation to an automatic synthesizer. Our synthesis algorithm scales to real-world programs, allowing us to invent new GPU kernels for stencil computations, matrix transposition, and a finite field multiplication algorithm (used in cryptographic applications). The synthesized 2D convolution and finite field multiplication kernels are on average 1.5--3.2x and 1.1--1.7x faster, respectively, than expert-optimized CUDA kernels. Phitchaya Mangpo Phothilimthana, Archibald Samuel Elliott, An Wang 0003, Abhinav Jangda, Bastian Hagedorn, Henrik Barthels, Samuel J. Kaufman, Vinod Grover, Emina Torlak, Rastislav Bodík |
ASPLOS | 4 |
| 2019 | Not So Fast: Analyzing the Performance of WebAssembly vs. Native Code
Abhinav Jangda, Bobby Powers, Emery D. Berger, Arjun Guha |
USENIX ATC | 1 |
| 2019 | Formal foundations of serverless computingabstractServerless computing (also known as functions as a service) is a new cloud computing abstraction that makes it easier to write robust, large-scale web services. In serverless computing, programmers write what are called serverless functions, which are programs that respond to external events. When demand for the serverless function spikes, the platform automatically allocates additional hardware and manages load-balancing; when demand falls, the platform silently deallocates idle resources; and when the platform detects a failure, it transparently retries affected requests. In 2014, Amazon Web Services introduced the first serverless platform, AWS Lambda, and similar abstractions are now available on all major cloud computing platforms. Unfortunately, the serverless computing abstraction exposes several low-level operational details that make it hard for programmers to write and reason about their code. This paper sheds light on this problem by presenting λ λ , an operational semantics of the essence of serverless computing. Despite being a small (half a page) core calculus, λ λ models all the low-level details that serverless functions can observe. To show that λ λ is useful, we present three applications. First, to ease reasoning about code, we present a simplified naive semantics of serverless execution and precisely characterize when the naive semantics and λ λ coincide. Second, we augment λ λ with a key-value store to allow reasoning about stateful serverless functions. Third, since a handful of serverless platforms support serverless function composition, we show how to extend λ λ with a composition language and show that our implementation can outperform prior work. Abhinav Jangda, Donald Pinckney, Yuriy Brun, Arjun Guha |
Proc. ACM Program. Lang. | 1 |
| 2018 | An effective fusion and tile size model for optimizing image processing pipelinesabstractEffective models for fusion of loop nests continue to remain a challenge in both general-purpose and domain-specific language (DSL) compilers. The difficulty often arises from the combinatorial explosion of grouping choices and their interaction with parallelism and locality. This paper presents a new fusion algorithm for high-performance domain-specific compilers for image processing pipelines. The fusion algorithm is driven by dynamic programming and explores spaces of fusion possibilities not covered by previous approaches, and is driven by a cost function more concrete and precise in capturing optimization criteria than prior approaches. The fusion model is particularly tailored to the transformation and optimization sequence applied by PolyMage and Halide, two recent DSLs for image processing pipelines. Our model-driven technique when implemented in PolyMage provides significant improvements (up to 4.32X) over PolyMage's approach (which uses auto-tuning to aid its model), and over Halide's automatic approach (by up to 2.46X) on two state-of-the-art shared-memory multicore architectures. Abhinav Jangda, Uday Bondhugula |
PPoPP | 1 |
| 2017 | RandHeap: Heap Randomization for Mitigating Heap Spray Attacks in Virtual MachinesabstractVirtual machines are an integral component of our present software systems infrastructure, including the web, and are here to stay. Web browsers like Google Chrome and Mozilla Firefox uses virtual machines to execute JavaScript code. Java Virtual Machines (JVMs) use just-in-time compilers to compile Java byte code to machine code. However, with the increasing use of virtual machines, they are also susceptible to security attacks. One such class of attack is the heap spray attack, wherein the attacker populates the heap with malicious code and exploits a vulnerability to jump to the populated malicious code in the heap, thereby enabling arbitrary code execution. In this paper, we present RandHeap, a technique to randomize the heap layout to detect and prevent heap spray attacks. RandHeap randomizes the heap in three different ways: (i) by randomizing object layout, (ii) by randomizing array layout, and (iii) by encrypting data stored on the heap. Using RandHeap, we were able to detect and prevent several heap spray attacks. For the evaluation of RandHeap, we implemented the concept of RandHeap in Google V8 and JikesRVM. We executed Octane 2.0 Benchmarks on Google V8 and Dacapo 9.12 Benchmarks on JikesRVM. Observations show that heap randomization using RandHeap is accompanied with low overhead and modest memory requirement. We implemented heap spraying attacks in Google V8 and JikesRVM and found that RandHeap was able to detect and prevent the attacks successfully. Abhinav Jangda, Mohit Mishra |
PST | 1 |