Karthik Srinivasa Murthy

dblp:304/5215 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0002-6063-9653ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 44% Program verification · 44% Programming languages and type systems · 13%
Artificial intelligence
1 paper
Efficient and distributed learning · 87% Language models and text generation · 13%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 50% Distributed systems · 50%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program verification
automated verification
0.912025
TensorRight: Automated Verification of Tensor Graph Rewrites · Proc. ACM Program. Lang. 2025
Machine learning › Efficient and distributed learning
distributed training
0.712023
Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models · ASPLOS (1) 2023
Machine learning › Efficient and distributed learning › distributed training
model parallelism
0.712023
Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models · ASPLOS (1) 2023
Distributed systems › communication optimization
communication-computation overlap
0.712023
Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models · ASPLOS (1) 2023
Parallel and multicore computing
parallel programming models
0.712023
Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models · ASPLOS (1) 2023
Programming languages and type systems › language semantics › formal semantics
denotational semantics
0.312025
TensorRight: Automated Verification of Tensor Graph Rewrites · Proc. ACM Program. Lang. 2025
Natural language and speech › Language models and text generation
large language model
0.212023
Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models · ASPLOS (1) 2023

Methods — techniques the papers use, named apart from their topics

intra-layer model parallelism · 1.3computation decomposition · 1.3symbolic execution · 0.9bounded rank analysis · 0.9SMT solving · 0.9
YearPublicationVenuePosition
2025 TensorRight: Automated Verification of Tensor Graph Rewrites
abstract
Tensor compilers, essential for generating efficient code for deep learning models across various applications, employ tensor graph rewrites as one of the key optimizations. These rewrites optimize tensor computational graphs with the expectation of preserving semantics for tensors of arbitrary rank and size. Despite this expectation, to the best of our knowledge, there does not exist a fully automated verification system to prove the soundness of these rewrites for tensors of arbitrary rank and size. Previous works, while successful in verifying rewrites with tensors of concrete rank, do not provide guarantees in the unbounded setting. To fill this gap, we introduce T ensor R ight , the first automatic verification system that can verify tensor graph rewrites for input tensors of arbitrary rank and size. We introduce a core language, T ensor R ight DSL, to represent rewrite rules using a novel axis definition, called aggregated-axis , which allows us to reason about an unbounded number of axes. We achieve unbounded verification by proving that there exists a bound on tensor ranks, under which bounded verification of all instances implies the correctness of the rewrite rule in the unbounded setting. We derive an algorithm to compute this rank using the denotational semantics of T ensor R ight DSL. T ensor R ight employs this algorithm to generate a finite number of bounded-verification proof obligations, which are then dispatched to an SMT solver using symbolic execution to automatically verify the correctness of the rewrite rules. We evaluate T ensor R ight ’s verification capabilities by implementing rewrite rules present in XLA ’s algebraic simplifier. The results demonstrate that T ensor R ight can prove the correctness of 115 out of 175 rules in their full generality, while the closest automatic, bounded -verification system can express only 18 of these rules.
Jai Arora, Sirui Lu, Devansh Jain 0001, Tianfan Xu, Farzin Houshmand, Phitchaya Mangpo Phothilimthana, Mohsen Lesani, Praveen Narayanan, Karthik Srinivasa Murthy, Rastislav Bodík, Amit Sabne, Charith Mendis
Proc. ACM Program. Lang.9
2023 Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models
abstract
Large deep learning models have shown great potential with state-of-the-art results in many tasks. However, running these large models is quite challenging on an accelerator (GPU or TPU) because the on-device memory is too limited for the size of these models. Intra-layer model parallelism is an approach to address the issues by partitioning individual layers or operators across multiple devices in a distributed accelerator cluster. But, the data communications generated by intra-layer model parallelism can contribute to a significant proportion of the overall execution time and severely hurt the computational efficiency.
Jinliang Wei, Amit Sabne, Andy Davis, Berkin Ilbeyi, Blake Hechtman, Dehao Chen, Karthik Srinivasa Murthy, Marcello Maggioni, Tongfei Guo, Yuanzhong Xu, Zongwei Zhou
ASPLOS (1)8
2021 A Flexible Approach to Autotuning Multi-Pass Machine Learning Compilers
abstract
Search-based techniques have been demonstrated effective in solving complex optimization problems that arise in domain-specific compilers for machine learning (ML). Unfortunately, deploying such techniques in production compilers is impeded by two limitations. First, prior works require factorization of a computation graph into smaller subgraphs over which search is applied. This decomposition is not only non-trivial but also significantly limits the scope of optimization. Second, prior works require search to be applied in a single stage in the compilation flow, which does not fit with the multi-stage layered architecture of most production ML compilers. This paper presents Xtat, an autotuner for production ML compilers that can tune both graph-level and subgraph-level optimizations across multiple compilation stages. Xtat applies Xtat-M, a flexible search methodology that defines a search formulation for joint optimizations by accurately modeling the interactions between different compiler passes. Xtat tunes tensor layouts, operator fusion decisions, tile sizes, and code generation parameters in XLA, a production ML compiler, using various search strategies. In an evaluation across 150 ML training and inference models on Tensor Processing Units (TPUs) at Google, Xtat offers up to 2.4x and an average 5% execution time speedup over the heavily-optimized XLA compiler.
Phitchaya Mangpo Phothilimthana, Amit Sabne, Nikhil Sarda, Karthik Srinivasa Murthy, Yanqi Zhou, Christof Angermueller, Michael Burrows, Sudip Roy 0002, Ketan Mandke, Rezsa Farahani, Yu Emma Wang, Berkin Ilbeyi, Blake A. Hechtman, Bjarke Roune, Yuanzhong Xu, Samuel J. Kaufman
PACT4