Nilesh A. Ahuja

dblp:66/1132 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
13since 2021 · last 2025
0000-0001-9467-4628ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 INRet: A General Framework for Accurate Retrieval of INRs for Shapes
abstract
Implicit neural representations (INRs) have become an important method for encoding various data types, such as 3D objects or scenes, images, and videos. They have proven to be particularly effective at representing 3D content, e.g., 3D scene reconstruction from 2D images, novel 3D content creation, as well as the representation, interpolation and completion of 3D shapes. With the widespread generation of 3D data in an INR format, there is a need to support effective organization and retrieval of INRs saved in a data store. A key aspect of retrieval and clustering of INRs in a data store is the formulation of similarity between INRs that would, for example, enable retrieval of similar INRs using a query INR. In this work, we propose INRet (INR Retrieve), a method for determining similarity between INRs that represent shapes, thus enabling accurate retrieval of similar shape INRs from an INR data store. INRet flexibly supports different INR architectures such as INRs with octree grids, triplanes, and hash grids, as well as different implicit functions including signed/unsigned distance function and occupancy field. We demonstrate that our method is more general and accurate than the existing INR retrieval method, which only supports simple MLP INRs and requires the same architecture between the query and stored INRs. Furthermore, compared to converting INRs to other representations (e.g., point clouds or multi-view images) for 3D shape retrieval, INRet achieves higher accuracy while avoiding the conversion overhead.
Yushi Guan, Daniel Kwan, Ruofan Liang, Selvakumar Panneer, Nilesh Jain, Nilesh A. Ahuja, Nandita Vijaykumar
3DV6
2025 ContraGS: Codebook-Condensed and Trainable Gaussian Splatting for Fast, Memory-Efficient Reconstruction
Sankeerth Durvasula, Sharanshangar Muhunthan, Zain Moustafa, Ruofan Liang, Yushi Guan, Nilesh A. Ahuja, Nilesh Jain, Selvakumar Panneer, Nandita Vijaykumar
ICCV7
2025 Retri3D: 3D Neural Graphics Representation Retrieval
abstract
Learnable 3D Neural Graphics Representations (3DNGR) have emerged as promising 3D representations for reconstructing 3D scenes from 2D images. Numerous works, including Neural Radiance Fields (NeRF), 3D Gaussian Splatting (3DGS), and their variants, have significantly enhanced the quality of these representations. The ease of construction from 2D images, suitability for online viewing/sharing, and applications in game/art design downstream tasks make it a vital 3D representation, with potential creation of large numbers of such 3D models. This necessitates large data stores, local or online, to save 3D visual data in these formats. However, no existing framework enables accurate retrieval of stored 3DNGRs. In this work, we propose, Retri3D, a framework that enables accurate and efficient retrieval of 3D scenes represented as NGRs from large data stores using text queries. We introduce a novel Neural Field Artifact Analysis technique, combined with a Smart Camera Movement Module, to select clean views and navigate pre-trained 3DNGRs. These techniques enable accurate retrieval by selecting the best viewing directions in the 3D scene for high-quality visual feature embeddings. We demonstrate that Retri3D is compatible with any NGR representation. On the LERF and ScanNet++ datasets, we show significant improvement in retrieval accuracy compared to existing techniques, while being orders of magnitude faster and storage efficient.
Yushi Guan, Daniel Kwan, Jean Sebastien Dandurand, Ruofan Liang, Nilesh Jain, Nilesh A. Ahuja, Selvakumar Panneer, Nandita Vijaykumar
ICLR8
2025 Rate-Distortion Theory in Coding for Machines and Its Applications
abstract
Recent years have seen a tremendous growth in both the capability and popularity of automatic machine analysis of media, especially images and video. As a result, a growing need for efficient compression methods optimised for machine vision, rather than human vision, has emerged. To meet this growing demand, significant developments have been made in image and video coding for machines. Unfortunately, while there is a substantial body of knowledge regarding rate-distortion theory for human vision, the same cannot be said of machine analysis. In this paper, we greatly extend the current rate-distortion theory for machines, providing insight into important design considerations of machine-vision codecs. We then utilise this newfound understanding to improve several methods for learned image coding for machines. Our proposed methods achieve state-of-the-art rate-distortion performance on several computer vision tasks - classification, instance and semantic segmentation, and object detection.
Alon Harell, Yalda Foroutan, Nilesh A. Ahuja, Parual Datta, Bhavya Kanzariya, V. Srinivasa Somayazulu, Omesh Tickoo, Anderson de Andrade, Ivan V. Bajic
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Split-DNN Computing for Video Analytics
Nagabhushan Eswara, Jaroslaw J. Sydir, V. Srinivasa Somayazulu, Parual Datta, Nilesh A. Ahuja, Omesh Tickoo
ICPR (3)5
2024 A Robust Framework for Evaluation of Unsupervised Time-Series Anomaly Detection
Onat Güngör, Amanda Rios, Priyanka Mudgal, Nilesh A. Ahuja, Tajana Rosing
ICPR (26)4
2023 FRE: A Fast Method For Anomaly Detection And Segmentation
Ibrahima J. Ndiour, Nilesh A. Ahuja, Ergin Utku Genc, Omesh Tickoo
BMVC2
2023 Neural Rate Estimator and Unsupervised Learning for Efficient Distributed Image Analytics in Split-DNN models
abstract
Thanks to advances in computer vision and AI, there has been a large growth in the demand for cloud-based visual analytics in which images captured by a low-powered edge device are transmitted to the cloud for analytics. Use of conventional codecs (JPEG, MPEG, HEVC, etc.) for compressing such data introduces artifacts that can seriously degrade the performance of the downstream analytic tasks. Split-DNN computing has emerged as a paradigm to address such usages, in which a DNN is partitioned into a client-side portion and a server side portion. Low-complexity neural networks called ‘bottleneck units' are introduced at the split point to transform the intermediate layer features into a lower-dimensional representation better suited for compression and transmission. Optimizing the pipeline for both compression and task-performance requires high-quality estimates of the information-theoretic rate of the intermediate features. Most works on compression for image analytics use heuristic approaches to estimate the rate, leading to suboptimal performance. We propose a high-quality ‘neural rateestimator’ to address this gap. We interpret the lower-dimensional bottleneck output as a latent representation of the intermediate feature and cast the rate-distortion optimization problem as one of training an equivalent variational auto-encoder with an appropriate loss function. We show that this leads to improved rate-distortion outcomes. We further show that replacing supervised loss terms (such as cross-entropy loss) by distillation-based losses in a teacher-student framework allows for unsupervised training of bottleneck units without the need for explicit training labels. This makes our method very attractive for real world deployments where access to labeled training data is difficult or expensive. We demonstrate that our method outperforms several state-of-the-art methods by obtaining improved task accuracy at lower bi-trates on image classification and semantic segmentation tasks.
Nilesh A. Ahuja, Parual Datta, Bhavya Kanzariya, V. Srinivasa Somayazulu, Omesh Tickoo
CVPR1
2022 incDFM: Incremental Deep Feature Modeling for Continual Novelty Detection
Amanda Rios, Nilesh A. Ahuja, Ibrahima J. Ndiour, Ergin Utku Genc, Laurent Itti, Omesh Tickoo
ECCV (25)2
2022 Anomalib: A Deep Learning Library for Anomaly Detection
abstract
This paper introduces anomalib1, a novel library for unsupervised anomaly detection and localization. With reproducibility and modularity in mind, this open-source library provides algorithms from the literature and a set of tools to design custom anomaly detection algorithms via a plug-and-play approach. Anomalib comprises state-of-the-art anomaly detection algorithms that achieve top performance on the benchmarks and that can be used off-the-shelf. In addition, the library provides components to design custom algorithms that could be tailored towards specific needs. Additional tools, including experiment trackers, visualizers, and hyper-parameter optimizers, make it simple to design and implement anomaly detection models. The library also supports OpenVINO model-optimization and quantization for real-time deployment. Overall, anomalib is an extensive library for the design, implementation, and deployment of unsupervised anomaly detection models from data to the edge.
Samet Akcay, Diederik J. D. Ameln, Ashwin Vaidya, Barath Lakshmanan, Nilesh A. Ahuja, Ergin Utku Genc
ICIP5
2022 Subspace Modeling for Fast Out-Of-Distribution and Anomaly Detection
abstract
This paper presents a fast, principled approach for detecting anomalous and out-of-distribution (OOD) samples in deep neural networks (DNN). We propose the application of linear statistical dimensionality reduction techniques on the semantic features produced by a DNN, in order to capture the low-dimensional subspace truly spanned by said features. We show that the feature reconstruction error (FRE), which is the ℓ2-norm of the difference between the original feature in the high-dimensional space and the pre-image of its low-dimensional reduced embedding, is highly effective for OOD and anomaly detection. To generalize to intermediate features produced at any given layer, we extend the methodology by applying nonlinear kernel-based methods. Experiments using standard image datasets and DNN architectures demonstrate that our method meets or exceeds best-in-class quality performance, but at a fraction of the computational and memory cost required by the state of the art. It can be trained and run very efficiently, even on a traditional CPU.
Ibrahima J. Ndiour, Nilesh A. Ahuja, Omesh Tickoo
ICIP2
2022 A Low-Complexity Approach to Rate-Distortion Optimized Variable Bit-Rate Compression for Split DNN Computing
abstract
Split computing has emerged as a recent paradigm for implementation of DNN-based AI workloads, wherein a DNN model is split into two parts, one of which is executed on a mobile/client device and the other on an edge-server (or cloud). Data compression is applied to the intermediate tensor from the DNN that needs to be transmitted, addressing the challenge of optimizing the rate-accuracy-complexity trade-off. Existing split-computing approaches adopt ML-based data compression, but require that the parameters of either the entire DNN model, or a significant portion of it, be retrained for different compression levels. This incurs a high computational and storage burden: training a full DNN model from scratch is computationally demanding, maintaining multiple copies of the DNN parameters increases storage requirements, and switching the full set of weights during inference increases memory bandwidth. In this paper, we present an approach that addresses all these challenges. It involves the systematic design and training of bottleneck units - simple, low-cost neural networks - that can be inserted at the point of split. Our approach is remarkably lightweight, both during training and inference, highly effective and achieves excellent rate-distortion performance at a small fraction of the compute and storage overhead compared to existing methods.
Parual Datta, Nilesh A. Ahuja, V. Srinivasa Somayazulu, Omesh Tickoo
ICPR2
2021 E2E Visual Analytics: Achieving >10X Edge/Cloud Optimizations
abstract
As visual analytics continues to rapidly grow, there is a critical need to improve the end-to-end efficiency of visual processing in edge/cloud systems. In this paper, we cover algorithms, systems and optimizations in three major areas for edge/cloud visual processing: (1) addressing storage and retrieval efficiency of visual data and meta-data by employing and optimizing visual data management systems, (2) addressing compute efficiency of visual analytics by taking advantage of co-optimization between the compression and analytics domains and (3) addressing networking (bandwidth) efficiency of visual data compression by tailoring it based on analytics tasks. We describe techniques in each of the above areas and measure its efficacy on state-of-the-art platforms (Intel Xeon), workloads and datasets. Our results show that we can achieve >10X improvements in each area based on novel algorithms, systems, and co-design optimizations. We also outline future research directions based on our findings which outline areas of further performance and efficiency advantages in end-to-end visual analytics.
Chaunte W. Lacewell, Nilesh A. Ahuja, Juan Pablo Muñoz, Parual Datta, Ragaad AlTarawneh, Vui Seng Chua, Nilesh Jain, Omesh Tickoo, Ravi R. Iyer 0001
NAS2
2020 Semantic-Preserving Image Compression
abstract
Video traffic comprises a large majority of the total traffic on the internet today. Uncompressed visual data requires a very large data rate; lossy compression techniques are employed in order to keep the data-rate manageable. Increasingly, a significant amount of visual data being generated is consumed by analytics (such as classification, detection, etc.) residing in the cloud. Image and video compression can produce visual artifacts, especially at lower data-rates, which can result in a significant drop in performance on such analytic tasks. Moreover, standard image and video compression techniques aim to optimize perceptual quality for human consumption by allocating more bits to perceptually significant features of the scene. However, these features may not necessarily be the most suitable ones for semantic tasks. We present here an approach to compress visual data in order to maximize performance on a given analytic task. We train a deep auto-encoder using a multi-task loss to learn the relevant embeddings. An approximate differentiable model of the quantizer is used during training which helps boost the accuracy during inference. We apply our approach on an image classification problem and show that for a given level of compression, it achieves higher classification accuracy than that obtained by performing classification on images compressed using JPEG. Our approach also outperforms the relevant state-of-the-art approach by a significant margin.
Neel Patwa, Nilesh A. Ahuja, V. Srinivasa Somayazulu, Omesh Tickoo, Srenivas Varadarajan, Shashidhar G. Koolagudi
ICIP2
2006 Multidimensional Generalized Sampling Theorem for wavelet Based Image Superresolution
abstract
The multidimensional generalized sampling theorem (GST) developed here provides a theoretical framework for wavelet based image superresolution, a topic of interest to the signal and image processing community during the last few years.
Nilesh A. Ahuja, Nirmal K. Bose
ICIP1
2006 Superresolution and noise filtering using moving least squares
abstract
An irregularly spaced sampling raster formed from a sequence of low-resolution frames is the input to an image sequence superresolution algorithm whose output is the set of image intensity values at the desired high-resolution image grid. The method of moving least squares (MLS) in polynomial space has proved to be useful in filtering the noise and approximating scattered data by minimizing a weighted mean-square error norm, but introducing blur in the process. Starting with the continuous version of the MLS, an explicit expression for the filter bandwidth is obtained as a function of the polynomial order of approximation and the standard deviation (scale) of the Gaussian weight function. A discrete implementation of the MLS is performed on images and the effect of choice of the two dependent parameters, scale and order, on noise filtering and reduction of blur introduced during the MLS process is studied.
Nirmal K. Bose, Nilesh A. Ahuja
IEEE Trans. Image Process.2
2005 Spatiotemporal-bandwidth product of m-dimensional signals
abstract
A system-theoretic proof for the spatiotemporal-bandwidth product (STBP) of multidimensional signals is given. Explicit formulae for the STBP of multivariate Gaussian signals and their derivatives are also supplied.
Nilesh A. Ahuja, Nirmal K. Bose
IEEE Signal Process. Lett.1