EDBT 2026 Demo / reviewers in the wild / expert
Krishna Teja Chitty-Venkata
dblp:248/5314
· DBLP profile ↗
10ranked-venue papers
9as first author
7since 2021 · last 2026
0000-0002-3027-1915ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 6 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PKAS: Predictive KVCache-Aware Scheduling for Faster LLM and Transformer InferencesabstractWith rising popularity of LLMs, the performance, scalability, and resource-efficiency of inferences become a crucial challenge. The core part of the inference process is the KV cache, which avoids recomputing intermediate attention states, and the batching strategy that batches multiple requests per forward pass to leverage GPU parallelism. KV cache memory grows linearly with sequence length and batch sizes, easily exceeding the limited GPU memory capacity. State-of-the-art inference runtimes use continuous batching to maximize GPU utilization by interleaving the processing of new requests (i.e., prefill requests) with ongoing generation requests (i.e., decode requests). However, existing schedulers greedily admit prefill requests without considering the future KV cache memory required to successfully run the decode phases. This shortsighted approach causes frequent KV cache overflows, which in turn trigger preemption and recomputation of requests, severely degrading both throughput and latency. We propose PKAS, a Predictive KV Cache-Aware Scheduling algorithm to mitigate this inefficiency by reducing preemptions. PKAS uses a low-overhead technique to simulate future KV cache utilization and guide the admissibility for new request candidates. Combined with lightweight output-length predictions, PKAS can make better batching decisions, preventing KV cache overflows and drastically reducing preemptions. Evaluations on diverse models and workloads show that PKAS achieves up to 7.34x higher throughput and 8x lower latency compared to state-of-the-art scheduling, with the largest gains on long-context workloads where KV cache pressure is high. Avinash Maurya, Krishna Teja Chitty-Venkata, Bogdan Nicolae, Antonios Kougkas, Xian-He Sun |
HPDC | 3 |
| 2025 | Langvision-Lora-Nas: Neural Architecture Search for Variable Lora Rank In Vision Language ModelsabstractVision Language Models (VLMs) integrate visual and text modalities to enable multimodal understanding and generation. These models typically combine a Vision Transformer (ViT) as an image encoder and a Large Language Model (LLM) for text generation. LoRA (Low-Rank Adaptation) is an efficient fine-tuning method to adapt pre-trained models to new tasks by introducing low-rank updates to their weights. While LoRA has emerged as a powerful technique for fine-tuning large models by introducing low-rank updates, current implementations assume a fixed rank, potentially limiting flexibility and efficiency across diverse tasks. This paper introduces LangVision-LoRA-NAS, a novel framework that integrates Neural Architecture Search (NAS) with LoRA to optimize VLMs for variable-rank adaptation. Our approach leverages NAS to dynamically search for the optimal LoRA rank configuration tailored to specific multimodal tasks, balancing performance and computational efficiency. Through extensive experiments using the LLaMA-3.2-11B model on several datasets, LangVision-LoRA-NAS demonstrates notable improvement in model performance while reducing fine-tuning costs. Our Base and searched fine-tuned models on LLaMA-3.2-11B-Vision-Instruct can be found here and the code for LangVision-LoRA-NAS can be found here. Krishna Teja Chitty-Venkata, Murali Emani, Venkatram Vishwanath |
ICIP | 1 |
| 2024 | WActiGrad: Structured Pruning for Efficient Finetuning and Inference of Large Language Models on AI Accelerators
Krishna Teja Chitty-Venkata, Varuni Katti Sastry, Murali Emani, Venkatram Vishwanath, Sanjif Shanmugavelu, Sylvia Howland |
Euro-Par (2) | 1 |
| 2023 | A survey of techniques for optimizing transformer inference
Krishna Teja Chitty-Venkata, Sparsh Mittal, Murali Emani, Venkatram Vishwanath, Arun K. Somani |
J. Syst. Archit. | 1 |
| 2022 | Efficient Design Space Exploration for Sparse Mixed Precision Neural ArchitecturesabstractPruning and Quantization are two effective Deep Neural Network (DNN) compression methods for efficient inference on various hardware platforms. Pruning refers to removing unimportant weights or nodes, whereas Quantization converts the floating-point parameters to low-bit fixed integer representation. The pruned and low precision models result in smaller and faster inference models on hardware platforms with almost the same accuracy as the unoptimized network. Tensor Cores in Nvidia Ampere 100 (A100) GPU supports (1) 2:4 fine-grained sparse pruning where 2 out of every 4 elements are pruned, and (2) traditional dense multiplication to achieve a good accuracy and performance trade-off. The A100 Tensor Core also takes advantage of 1-bit, 4-bit, and 8-bit multiplication to speed up the inference of a model. Hence, finding the right matrix type (dense or 2:4 sparse) along with the precision for each layer becomes a combinatorial problem. Neural Architecture Search (NAS) can alleviate such problems by automating the architecture design process instead of a brute-force search. In this paper, we propose (i) Mixed Sparse and Precision Search (MSPS), a NAS framework to search for efficient sparse and mixed-precision quantized model within the predefined search space and fixed backbone neural network (Eg. ResNet50), and (ii) Architecture, Sparse and Precision Search (ASPS) to jointly search for kernel size and number of filters, and sparse-precision combination of each layer. We illustrate the effectiveness of our methods targeting A100 Tensor Core on Nvidia GPUs by searching efficient sparse-mixed precision networks on ResNet50 and achieving better accuracy-latency trade-off models compared to the manually designed Uniform Sparse Int8 networks. Krishna Teja Chitty-Venkata, Murali Emani, Venkatram Vishwanath, Arun K. Somani |
HPDC | 1 |
| 2021 | Array-Aware Neural Architecture SearchabstractConvolutional Neural Networks (CNNs) have exceeded human accuracy in many Computer Vision tasks, such as Image Classification, Object Detection, Image Segmentation, etc. This advancement is due to the efficient manual design of CNNs in initially, followed by automated design through Neural Architecture Search (NAS). In parallel to neural network design, advances in Accelerator hardware design, such as Google’s Tensor Processing Unit (TPU), Eyeriss, etc., also occurred for efficient processing of CNN forward propagation. The heart of these accelerators is an array processor (Systolic Array) of a fixed dimension, that limits the amount of CNN computation that can be carried out in a single clock cycle. While NAS is able to produce efficient neural architectures, the networks need to be co-designed with respect to the underlying array dimensions to obtain the best performance. In this paper, we introduce "Array Aware Neural Architecture Search" to automatically design efficient CNNs for a fixed array-based neural network accelerator. Previous Hardware Aware NAS methods consider a fixed search space for different hardware platforms and search within its predefined space. We explore the search space based on the underlying hardware array dimensions to design a more efficient CNN architectures for optimal performance. We observe that our proposed NAS methods on the CIFAR-10 dataset produce similar accuracy as the baseline network while saving a substantial number of cycles on the Array. Krishna Teja Chitty-Venkata, Arun K. Somani |
ASAP | 1 |
| 2021 | Searching Architecture and Precision for U-net based Image Restoration TasksabstractManually architecting Deep Neural Networks (DNNs) has led to the success of Deep Learning in many domains. However, recent DNNs designed using Neural Architecture Search (NAS) have exceeded manually designed architectures and have significantly reduced the human effort to develop complex networks. Current works use NAS to identify a cell architecture constrained by a fixed order of operations that is then replicated throughout the network. The constraints potentially limit the effectiveness of NAS in converging on a more efficient DNN architecture. In the first part of our paper, we propose “Operation Search,” a search on an enlarged topological space for U-net and its variants that retain efficiency. The idea is to allow for custom cells (operations and their sequence) at various levels of the network to maximize image quality while being sensitive to computation cost. In the second part of our paper, we propose custom quantization at various levels resulting in a mixed-precision network. Additionally, we increase the search efficiency by constraining the search space to use the same precision for both weights and activations at any level. This does not result in computational inefficiency because it matches the operand precisions supported by Tensor Core enabled GPUs. Krishna Teja Chitty-Venkata, Arun K. Somani, Sreenivas Kothandaraman |
ICIP | 1 |
| 2020 | Array Aware Training/Pruning: Methods for Efficient Forward Propagation on Array-based Neural Network AcceleratorsabstractDue to the increase in the use of large-sized Deep Neural Networks (DNNs) over the years, specialized hardware accelerators such as Tensor Processing Unit and Eyeriss have been developed to accelerate the forward pass of the network. The essential component of these devices is an array processor which is composed of multiple individual compute units for efficiently executing Multiplication and Accumulation (MAC) operation. As the size of this array limits the amount of DNN processing of a single layer, the computation is performed in several batches serially leading to extra compute cycles along both the axes. In practice, due to the mismatch between matrix and array sizes, the computation does not map on the array exactly. In this work, we address the issue of minimizing processing cycles on the array by adjusting the DNN model parameters by using a structured hardware array dependent optimization. We introduce two techniques in this paper: Array Aware Training (AAT) for efficient training and Array Aware Pruning (AAP) for efficient inference. Weight pruning is an approach to remove redundant parameters in the network to decrease the size of the network. The key idea behind pruning in this paper is to adjust the model parameters (the weight matrix) so that the array is fully utilized in each computation batch. Our goal is to compress the model based on the size of the array so as to reduce the number of computation cycles. We observe that both the proposed techniques results into similar accuracy as the original network while saving a significant number of processing cycles (75%). Krishna Teja Chitty-Venkata, Arun K. Somani |
ASAP | 1 |
| 2020 | Model Compression on Faulty Array-based Neural Network AcceleratorabstractDue to an increase in the size of Deep Neural Networks (DNNs), special purpose hardware has gained prominence to accelerate the forward pass of the network like Google's Tensor Processing Unit (TPU) and Eyeriss. The heart of these accelerators is a Matrix Multiplication unit, which is based on systolic array architecture. This array processor has a grid-like structure, made of individual Processing Elements (PEs) that can be extended along row and column. A lot of work has been done on the computing array implementation and its reliability concerns in the past. However, their fault tolerance perspective with respect to DNNs is not yet fully understood with a fault model. We, in this paper, first present a fault model i.e., different sequences in which faults can occur on the array. We classify the fault modes into random, row, and column, and study their impact on the accuracy of DNNs followed by overheads of the mitigation strategies. Pruning is a process of removing redundant parameters in the model to decrease the network size for efficient performance. Although several pruning techniques have been developed to reduce the inference time on general-purpose and special-purpose systems, model compression (pruning) under faulty scenario has not yet been explored. In the second part of our work, we co-design a Fault based and Array size based Pruning (FPAP) algorithm with an intent of bypassing the faults and removing the internal redundancy at the same time for efficient inference. We compare our method with recent pruning methods under different fault scenarios and array sizes. We achieve a mean speedup of 4.2x where the baselines achieved 1.6x on ConvNet, NiN, AlexNet, VGG16 over Eyeriss in the case of random faults. Krishna Teja Chitty-Venkata, Arun K. Somani |
PRDC | 1 |
| 2019 | Impact of Structural Faults on Neural Network PerformanceabstractDeep Learning (DL), a subset of Artificial Intelligence (AI), is growing rapidly with possible applications in different domains such as speech recognition, computer vision etc. Deep Neural Network (DNN), the backbone of DL algorithms is a directed graph containing multiple layers with different number of neurons residing in each layer. The use of these networks has been increased in the last few years due to availability of large data sets and huge computation power. As the size of DNN is growing over the years, researchers have developed specialized hardware accelerators to reduce the inference compute time. An example of such domain specific architecture designed for Neural Network acceleration is Tensor Processing Unit (TPU) which outperforms GPU in the inference stage of DNN execution. The heart of this inference engine is a Matrix Multiplication unit which is based on systolic array architecture. The TPU's systolic array is a grid-like structure made of individual processing elements that can be extended along rows and columns. Due to external environmental factors or internal scaling of semiconductor, these systems are often prone to faults which leads to improper calculations and thereby resulting in inaccurate decisions by the DNN. Although a lot of work has been done in the past on the computing array implementation and it's reliability concerns, their fault tolerance behavior for DNN application is not very well understood. It is not even clear what would be the impact of various different faults on the accuracy. We in this work, first study possible mapping strategies to implement a convolution and dense layer weights on TPU systolic array. Next we consider various faults scenarios that may occur in the array. We divide these fault scenarios into low, high row and column faults (Fig. 1(a) pictorially represents column faults) modes with respect to the multiplication unit. Next, we study the impact of these fault models on the overall accuracy of the DNN performance on a faculty TPU unit. The goal is to study the resiliency and overcome the limitations of earlier work. The previous work was very effective in masking the random faults which used pruning of weights (removing weights or connections in the DNN) plus retraining to mask the faults on the array. However, it failed in the case of column faults which is clearly shown in Fig. 1(b). We also propose techniques to mitigate or bypass the row and column faults. Our mapping strategy follows physical_x(i) = i%N and physical_y(j) = j%N where (i,j) represents the index of dense (FC) weight matrix and (physical x(i), physical y(j)) indicates the actual physical location on the array of size N. The convolution filters are linearized with respect to every channel so as to convert them into proper weight matrix and mapped according to the previous mentioned policy. It was shown that DNNs can up to certain faults in the array while retaining the original accuracy (low row faults). The accuracy of the network decreases even with one column faults if it (column) is in the use. As per the results, it is proved that for the same number of row and column faults, the latter has most impact on the network accuracy because pruning input neuron has very little effect than pruning an output neuron. We experimented with three different networks and found the influence of these different faults to be the same. These faults can be mitigated using techniques like Matrix Transpose and Array Reduction which does not require retraining of weights. For low row faults, the original mapping policy can be retained such that weights can be mapped at their exact locations which does not affect the accuracy. Low column faults can be converted into low row faults by transposing the matrix. In the case of high row (column) faults, the entire row (column) has to be avoided to completely bypass the faulty locations. Static mapping of weights along with retraining the network on the array can be effective in the case of random faults. Adapting to change in the case of structured faults can reduce the burden of retraining which happens outside the TPU. Krishna Teja Chitty-Venkata, Arun K. Somani |
ASAP | 1 |