Harsha Santhanam

dblp:407/8479 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0009-0008-1636-5430ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 67% Storage systems · 33%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Artificial intelligence
1 paper
Knowledge representation and reasoning · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › computational storage
in-storage computing
0.912025
In-Storage Acceleration of Retrieval Augmented Generation as a Service · ISCA 2025
Hardware accelerators and domain-specific architectures › accelerator integration
near-storage accelerator
0.912025
In-Storage Acceleration of Retrieval Augmented Generation as a Service · ISCA 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation
query embedding
0.312025
In-Storage Acceleration of Retrieval Augmented Generation as a Service · ISCA 2025
Information retrieval
retrieval-augmented generation
0.312025
In-Storage Acceleration of Retrieval Augmented Generation as a Service · ISCA 2025
Information retrieval
similarity search
0.312025
In-Storage Acceleration of Retrieval Augmented Generation as a Service · ISCA 2025

Methods — techniques the papers use, named apart from their topics

metamorphic architecture · 2.6in-storage processing · 2.6
YearPublicationVenuePosition
2026 IroKnight: Ownership-Preserving Neural Acceleration for Inference Serving
Harsha Santhanam, Ashwin Rohit Alagiri Rajan, Hadi Esmaeilzadeh
ISCA1
2025 In-Storage Acceleration of Retrieval Augmented Generation as a Service
abstract
Retrieval-augmented generation (RAG) services are rapidly gaining adoption in enterprise settings as they combine information retrieval systems (e.g., databases) with large language models (LLMs) to enhance response generation and reduce hallucinations.By augmenting an LLM's fixed pre-trained knowledge with real-time information retrieval, RAG enables models to effectively extend their context to large knowledge bases by selectively retrieving only the most relevant information.As a result, RAG provides the effect of dynamic updates to the LLM's knowledge without requiring expensive and time-consuming retraining.While some deployments keep the entire database in memory, RAG services are increasingly shifting toward persistent storage to accommodate ever-growing knowledge bases, enhance utility, and improve cost-efficiency.However, this transition fundamentally reshapes the system's performance profile: empirical analysis reveals that the Search & Retrieval phase emerges as the dominant contributor to end-to-end latency.This phase typically involves (1) running a smaller language model to generate query embeddings, (2) executing similarity and relevance checks over varying data structures, and (3) performing frequent, long-latency accesses to persistent storage.To address this triad of challenges, we propose a metamorphic in-storage accelerator architecture that provides the necessary programmability to support diverse RAG algorithms, dynamic data structures, and varying computational patterns.The architecture also supports in-storage execution of smaller language models for query embedding generation while final LLM generation is executed on DGX A100 systems.Experimental results show up to 4.3× and 1.5× improvement in end-to-end throughput compared to conventional retrieval pipelines using Xeon CPUs with NVMe storage and A100 GPUs with DRAM, respectively.
Rohan Mahapatra, Harsha Santhanam, Christopher Priebe, Hanyang Xu 0002, Hadi Esmaeilzadeh
ISCA2