Vikram Nelvoy Rajendiran

dblp:291/5580 · also Vikram N. R · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0001-7477-2124ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Mobile-friendly Image de-noising: Hardware Conscious Optimization for Edge Application
abstract
Image enhancement is a critical task in computer vision and photography that is often entangled with noise. This renders the traditional Image Signal Processing (ISP) ineffective compared to the advances in deep learning. However, the success of such methods is increasingly associated with the ease of their deployment on edge devices, such as smartphones. This work presents a novel mobile-friendly network for image de-noising obtained with Entropy-Regularized differentiable Neural Architecture Search (NAS) on a hardware-aware search space for a U-Net architecture, which is first-of-its-kind. The designed model has 12% less parameters, with ~2-fold improvement in on- device latency and 1.5-fold improvement in the memory footprint for a 0.7% drop in PSNR, when deployed and profiled on Samsung Galaxy S24 Ultra. Compared to the SOTA Swin-Transformer for Image Restoration, the proposed network had competitive accuracy with ~18-fold reduction in GMACs. Further, the network was tested successfully for Gaussian de-noising with 3 intensities on 4 benchmarks and real-world de-noising on 1 benchmark demonstrating its generalization ability.
Srinivas Soumitri Miriyala, Sowmya Vajrala, Hitesh Kumar, Sravanth Kodavanti, Vikram Nelvoy Rajendiran
ICASSP5
2024 Mixed Precision Neural Quantization with Multi-Objective Bayesian Optimization for on-Device Deployment
abstract
Mixed-precision quantization has emerged as a solution in recent times for accurate inference of Deep Neural Networks on edge. However the prior-art is far from being deployable on embedded devices due to various practical limitations. In this work, a pipeline is designed that a) performs layer-wise assignment of bit-precisions, b) builds the quantized graph, c) deploys the graph on the embedded device (Galaxy S23), and d) measures the accuracy and on-device performance. This pipeline is optimized using multi-objective Bayesian Optimization for simultaneously maximizing the accuracy and minimizing the on-device inference time, resulting in a Pareto list. The best configuration among them resulted in 3.16, 2.8, and 2.57 times model compression, 31%, 26% and 18% latency improvement and -0.01, 0.27, and 0.08 accuracy drop for ResNet18, MobileNetV2, and InceptionV3, respectively, on ImageNet, establishing a new benchmark in mixed-precision quantization.
Srinivas Soumitri Miriyala, P. K. Suhas, Utsav Tiwari, Vikram Nelvoy Rajendiran
ICASSP4
2024 Learning Representations from Explainable and Connectionist Approaches for Visual Question Answering
abstract
Reasoning conditioned on visual and linguistic information has gained immense importance in recent times. The prior art in Visual Question Answering (VQA) has been predominantly connectionist in nature. To resolve the issues of connectionist AI models, Symbolic models were proposed that allowed for explainable visual reasoning. In addition to semantic parsing, such models worked towards visual parsing resulting in scene graphs that provided scope for accurate reasoning conditioned on the explainable scene graphs. However, the real scenarios of VQA cannot always be segregated exclusively into connectionist (neural networks) and conceptual modalities. Rather, they are always dependent on the relationships and interactions between the two modalities. In this work, the authors proposed a question-guided attention mechanism that combines the approach of explainable visual reasoning through scene graphs with a cross-modality-based multi-head attention mechanism. The contributions of con-nectionist and conceptual modalities are learned through the semantic parsing of questions in each VQA task. The novel method is tested with the VQA2.0 and GQA and it resulted in 65.31% and 63.06% accuracy, respectively, which is better than the state-of-the-art in explainable AI.
Aakansha Mishra, Srinivas Soumitri Miriyala, Vikram Nelvoy Rajendiran
ICASSP3
2024 Edge Deployable Distributed Evolutionary Optimization based Calibration method for Neural Quantization
abstract
Accuracy drop in neural quantization is addressed in prior-art through Post Training Quantization (PTQ) schemes such as Percentile and Range-based calibration that remain sensitive to the data distribution. On the other hand, the sophisticated methods that efficiently handle the variability in data require their deployment also on the embedded devices significantly increasing the memory and latency. We solve this issue by translating PTQ as a non-linear programming problem, which is then efficiently solved block-wise in distributed manner using an evolutionary algorithm. The quantized models are also deployed on the Galaxy S23 smartphone to measure the on-device performance. MobileNetV2 and ResNet18 in Int8 precision resulted in 0.33 and 0.03 accuracy drop, respectively, which is best by the standards of PTQ. Our approach is the first-of-its-kind hardware-agnostic high-accuracy PTQ method that allows the seamless deployment of quantized networks on embedded devices.
Utsav Tiwari, Srinivas Soumitri Miriyala, Vikram Nelvoy Rajendiran
ICASSP3
2024 Efficient Visual Question Answering on Embedded Devices: Cross-Modality Attention With Evolutionary Quantization
abstract
Visual Question Answering (VQA) lies at the intersection of vision and language domains necessitating learning representations from multiple modalities. While the model development for VQA has witnessed tremendous growth, the efforts for its deployment on embedded devices have been lagging limiting its true potential. In this work, the authors address this challenge by designing a novel hardware-friendly architecture for VQA based on the transformer model with cross-modality attention. The memory footprint of the VQA model is optimized for on-device deployment using a distributed framework for Post Training Quantization (PTQ) formulated as a Non-Linear Programming (NLP) problem. The NLP problem is solved using an Evolutionary algorithm to determine the low-bit representation of the VQA model with minimal accuracy drop compared to the full precision model. The quantized model for VQA with a marginal accuracy drop of less than 2%, resulted in 4 times memory improvement, and over 2 times latency improvement, enabling its successful deployment on the Samsung Galaxy S23 device. The comprehensive study explores the potential of the proposed generic end-to-end pipeline from VQA model development to its deployment.
Aakansha Mishra, Aditya Agarwala, Utsav Tiwari, Vikram Nelvoy Rajendiran, Srinivas Soumitri Miriyala
ICIP4
2024 Efficient Adapter on Pre-trained Visual Feature Reliance in Medical Visual Question Answering
Aakansha Mishra, Prateek Keserwani, Vikram Nelvoy Rajendiran, Ashok K. Senapati
ICPR (28)3
2023 Receptive Field Reliant Zero-Cost Proxies for Neural Architecture Search
abstract
Neural Architecture Search (NAS) is a fast growing technology for automatic design of deep-learning architectures. NAS includes three stages: search space design, search strategy, and evaluation criterion. Among these, the evaluation of various architectures is very cost-intensive task. In this work, we have proposed a set of receptive field reliant zero-cost proxies which need only one iteration of training and thereby reduce the computational time associated with evaluation criterion during the NAS. The proposed zero-cost proxies are based on layer-wise binding of the prune-at-initialization score with its receptive field for more effective measure as compared to the vanilla counterparts to achieve generalizability. The proposed zero-cost proxies are validated on the set of PyTorchCV models, and NAS-Bench-201 benchmarking datasets. The proposed zero-cost proxies have performed better for set of PyTorchCV models and competitively with vanilla counterparts for NAS-Bench-201. The efficiency of the proposed method is also demonstrated in NAS on NAS-Bench-201 using Aging Evolution as controller.
Prateek Keserwani, Srinivas Soumitri Miriyala, Vikram Nelvoy Rajendiran, Pradeep N. Shivamurthappa
ICASSP3
2023 A Generalized Zero-Shot Quantization of Deep Convolutional Neural Networks Via Learned Weights Statistics
abstract
Quantizing the floating-point weights and activations of deep convolutional neural networks to fixed-point representation yields reduced memory footprints and inference time. Recently, efforts have been afoot towards zero-shot quantization that does not require original unlabelled training samples of a given task. These best-published works heavily rely on the learned batch normalization (BN) parameters to infer the range of the activations for quantization. In particular, these methods are built upon either empirical estimation framework or the data distillation approach, for computing the range of the activations. However, the performance of such schemes severely degrades when presented with a network that does not accommodate BN layers. In this line of thought, we propose ageneralized zero-shot quantization(GZSQ) framework that neither requires original data nor relies on BN layer statistics. We have utilized the data distillation approach and leveraged only the pre-trained weights of the model to estimate enriched data for range calibration of the activations. To the best of our knowledge, this is the first work that utilizes the distribution of the pre-trained weights to assist the process of zero-shot quantization. The proposed scheme has significantly outperformed the existing zero-shot works,e.g., an improvement of$\sim$33% in classification accuracy for MobileNetV2 and several other models that are w & w/o BN layers, for a variety of tasks. We have also demonstrated the efficacy of the proposed work across multiple open-source quantization frameworks. Importantly, our work is the first attempt towards the post-training zero-shot quantization of futuristic unnormalized deep neural networks.
Prasen Kumar Sharma, Arun Abraham, Vikram Nelvoy Rajendiran
IEEE Trans. Multim.3
2020 Processor Pipelining Method for Efficient Deep Neural Network Inference on Embedded Devices
abstract
Myriad applications of Deep Neural Networks (DNN) and the race for better accuracy have paved the way for the development of more computationally intensive network architectures. Execution of these heavy networks on embedded devices needs highly efficient real-time DNN inference frameworks. But the sequential architecture of popular DNNs makes it difficult to parallelize its operations among different processors. We propose a novel pipelining method pluggable on top of conventional inference frameworks and capable of parallelizing DNN inference on heterogeneous processors without impacting the accuracy. We partition the network into subnets, by estimating the optimal split points, and pipeline these subnets across multiple processors. The results shows that the proposed method achieves up to 68% improvement in the frames per second (FPS) rate of popular network architectures like VGG19, DenseNet-121 and ResNet-152. Moreover, we show that our method can be used to extract even more performance out of high performance chipsets, by better utilizing the capabilities of its AI processor ecosystem. We also showcase that our method can be easily extended to other low performance chipsets, where this additional performance gain is crucial to deploy real-time AI applications. Our results show performance improvement of up to 47% in the FPS rate on these chipsets without the need of specialized AI hardware.
Akshay Parashar, Arun Abraham, Deepak Chaudhary, Vikram Nelvoy Rajendiran
HiPC4