VLDB 2026 Research / reviewers in the wild / expert
Ashwin Krishnan
dblp:317/1410
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0002-8592-3132ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM4PART: From CallGraphs to Bottleneck Partitions with LLM-Enhanced InsightsabstractModern Python applications often use heterogeneous workloads (mixed compute, data, and I/O) and run on CPUs, GPUs, and accelerators across domains such as language processing, finance, and recommendation systems [1], [2]. Understanding their performance requires analyzing how functions interact and contribute to overall execution. Profiling tools such as cProfile, SnakeViz, and pyinstrument expose function-level bottlenecks, but in practice developers rely on manual inspection or limited profiling, which does not scale well and may not reflect realistic workloads. Moreover, profiling typically reports timings at the function level, making it difficult to understand higher-level execution regions composed of multiple interacting functions. Prior work on automated application decomposition includes tools such as Mono2Micro [3] and ServiceCutter [4], which use runtime traces for modularization, while MLIR and HPVM [1] support backend optimizations after architectural decisions but provide limited support for early performance analysis. In contrast, our approach uses lightweight profiling with cprofile and static call graphs integrated with LLM-based analysis to identify and interpret bottleneck partitions before full-scale deployment. Vyuhita Bonthu, Venkatesh Pasumarti, Ashwin Krishnan, Manoj Nambiar 0001 |
ISPASS | 3 |
| 2024 | CAR-LLM: Cloud Accelerator Recommender for Large Language ModelsabstractTransformer-based Large Language Models (LLMs) have garnered significant attention due to their multi-modality and exceptional performance across diverse applications. This surge in popularity has spurred the development of numerous new LLMs and corresponding hardware solutions for efficient deployment. However, deploying LLMs on different accelerators for inference poses a significant challenge due to the vast search space involved. This encompasses considerations such as the number of accelerator chips, instances of accelerators, and the choice of inference framework to meet stringent workload and latency constraints. In this paper, we introduce the Cloud Accelerator Recommender for Large Language Models (CAR-LLM), a framework designed to optimize the deployment of LLMs on available accelerators and hardware across various cloud vendors. CAR-LLM aims to achieve maximum performance with minimal cost by recommending optimal deployment strategies. We outline a cost-effective experimental strategy and investigate key parameters affecting the latency of LLMs on specific hardware. Additionally, we develop a performance model to predict latency and throughput, enhancing deployment efficiency and decision-making for LLM applications. Ashwin Krishnan, Venkatesh Pasumarti, Samarth Inamdar, Arghyajoy Mondal, Manoj Nambiar 0001, Rekha Singhal |
HiPC | 1 |
| 2022 | CMOS Ring-Oscillator-Based Electrochemical Capacitance Imager with Frequency-Division-Multiplexed ReadoutabstractWe present a CMOS electrochemical capacitance imager designed to overcome Debye-length-screening effects in integrated electrochemical assays and featuring frequency-division-multiplexed (FDM) column readout. Our imager contains eight pixel arrays where each pixel contains a 5-MHz-180-MHz ring-oscillator-based capacitance-to-frequency converter (CFC) to detect interfacial capacitance changes at an in-pixel working electrode. We implement FDM readout, which reduces column-readout time compared to a time-division-multiplexed approach, by operating each CFC in a column at a unique nominal frequency while summing the pixel output currents. We report experimental results from electronic performance characterization of a $3.0\times 2.5-\mathrm{mm}^{2}$ 420-pixel electrochemical capacitance imager, fabricated in a 1.8-V 0.18-$\mu{\mathrm{m}}$ CMOS process. Ashwin Krishnan, Peter M. Levine |
ISCAS | 1 |
| 2022 | Performance Model and Profile Guided Design of a High-Performance Session Based Recommendation EngineabstractSession-based recommendation (SBR) systems are widely used in transactional systems to make personalized recommendations to the end-user. In online retail systems, recommendations-based decisions need to be made at a very high rate especially during peak hours. The required computational workload is very high especially when there is a larger number of products involved. Session Based Recommendation (SBR) models incorporate the learning-based product buying pattern from various user interaction sessions and try to recommend the top-K products, the user is likely to purchase. These models comprise several functional layers that widely vary in their compute and data access patterns. To support high recommendation rates, all these layers need a performance optimal implementation, which can be a challenge given the diverse nature of the computations involved. For this reason, one compute platform - whether it is CPU, GPU, or a Field Programmable Gate Array (FPGA) may not be able to provide an optimal implementation for all the layers. In this paper, we describe performance modeling and profile-based design approach to arrive at an optimal implementation, comprising of the hybrid CPU, GPU, and FPGA platforms for NISER - a session-based recommendation model that avoids popularity bias in recommendations. In addition, the design for the CPU-FPGA hybrid platform is implemented for NISER and we observed that experimental results closely follow the results predicted by the performance model for the implemented deployment option. Ashwin Krishnan, Manoj Nambiar 0001, Nupur Sumeet, Sana Iqbal |
ICPE | 1 |