Sampath Rajendra

dblp:332/1579 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0004-2868-0260ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Hardware accelerators and domain-specific architectures · 52% GPUs and heterogeneous computing · 48%
Software engineering, system software, and programming languages
3 papers
Compilers and program optimization · 100%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 100%
Artificial intelligence
1 paper
Kernel, tree and ensemble methods · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
analytical query processing
1.012026
Scalable GPU Acceleration of Scalar Functions in Analytical Databases: Compilation, Benchmarking, and Optimization · Proc. VLDB Endow. 2026
GPUs and heterogeneous computing
GPU computing
1.012026
Scalable GPU Acceleration of Scalar Functions in Analytical Databases: Compilation, Benchmarking, and Optimization · Proc. VLDB Endow. 2026
Compilers and program optimization › compiler construction
retargetable compilation
0.812024
SilvanForge: A Schedule-Guided Retargetable Compiler for Decision Tree Inference · SOSP 2024
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.812024
SilvanForge: A Schedule-Guided Retargetable Compiler for Decision Tree Inference · SOSP 2024
Compilers and program optimization
machine learning compiler
0.612022
Treebeard: An Optimizing Compiler for Decision Tree Based ML Inference · MICRO 2022
Hardware accelerators and domain-specific architectures
domain-specific compilers
0.612022
Treebeard: An Optimizing Compiler for Decision Tree Based ML Inference · MICRO 2022
Compilers and program optimization › domain-specific compilation
query compilation
0.312026
Scalable GPU Acceleration of Scalar Functions in Analytical Databases: Compilation, Benchmarking, and Optimization · Proc. VLDB Endow. 2026
Machine learning › Kernel, tree and ensemble methods › ensemble learning
tree ensembles
0.212022
Treebeard: An Optimizing Compiler for Decision Tree Based ML Inference · MICRO 2022

Methods — techniques the papers use, named apart from their topics

compilation · 3.0benchmarking · 3.0progressive lowering · 1.7SIMD vectorization · 1.7MLIR · 1.7decision tree inference · 1.5
YearPublicationVenuePosition
2026 Scalable GPU Acceleration of Scalar Functions in Analytical Databases: Compilation, Benchmarking, and Optimization
Kaushik Rajan, Sampath Rajendra, Momin Al-Ghosien, Nicolas Bruno, Carlo Curino, Matteo Interlandi, Yinan Li 0009, Lukas M. Maas, Craig Peeper, Surajit Chaudhuri, Johannes Gehrke
Proc. VLDB Endow.2
2024 Welding Natural Language Queries to Analytics IRs with LLMs
Kaushik Rajan, Aseem Rastogi, Akash Lal, Sampath Rajendra, Krithika Subramanian, Krut Patel
CIDR4
2024 SilvanForge: A Schedule-Guided Retargetable Compiler for Decision Tree Inference
abstract
The proliferation of machine learning together with the rapid evolution of the hardware ecosystem has led to a surge in the demand for model inference on a variety of hardware. Decision tree based models are the most popular models on tabular data. This paper is motivated by the problems encountered when targeting inference of these models to run at peak performance on CPU and GPU targets. Existing solutions are neither portable nor achieve the best possible performance for the specific hardware they target. This is because they do not explore and customize optimization configurations to the target processor and the model being used.
Ashwin Prasad, Sampath Rajendra, Kaushik Rajan, R. Govindarajan, Uday Bondhugula
SOSP2
2022 Treebeard: An Optimizing Compiler for Decision Tree Based ML Inference
abstract
Decision tree ensembles are among the most commonly used machine learning models. These models are used in a wide range of applications and are deployed at scale. Decision tree ensemble inference is usually performed with libraries such as XGBoost, LightGBM, and Sklearn. These libraries incorporate a fixed set of optimizations for the hardware targets they support. However, maintaining these optimizations is prohibitively expensive with the evolution of hardware. Further, they do not specialize the inference code to the model being used, leaving significant performance on the table. This paper presents TREEBEARD, an optimizing compiler that progressively lowers the inference computation to optimized CPU code through multiple intermediate abstractions. By applying model-specific optimizations at the higher levels, tree walk optimizations at the middle level, and machine-specific optimizations lower down, TREEBEARD can specialize inference code for each model on each supported CPU target. TREEBEARD combines several novel optimizations at various abstraction levels to mitigate architectural bottlenecks and enable SIMD vectorization of tree walks. We implement TREEBEARD using the MLIR compiler infrastructure and demonstrate its utility by evaluating it on a diverse set of benchmarks. TREEBEARD is significantly faster than state-of-the-art systems, XGBoost, Treelite and Hummingbird, by 2.6×, 4.7× and 5.4× respectively in a single-core execution setting, and by 2.3×, 2.7× and 14× respectively in multi-core settings.
Ashwin Prasad, Sampath Rajendra, Kaushik Rajan, R. Govindarajan, Uday Bondhugula
MICRO2