Nilesh Jain

dblp:134/6343 · also Nilesh K. Jain · DBLP profile ↗
← Back
19ranked-venue papers
0as first author
18since 2021 · last 2026
0000-0002-9685-2763ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LiveAvatar: Real-time 3D Gaussian Head Avatars from Single Images
abstract
We present a real-time method for generating 3D Gaussian head avatars directly from monocular 2D video streams, without per-subject training or 3D morphable model fitting. We train a single model on a large collection of diverse video clips of human heads to generalize across identities, expressions, and poses. Input images are encoded with a pre-trained vision transformer into feature maps, from which both the static and dynamic components are inferred without relying on facial keypoints or parametric models. Dynamic information is injected into the avatar generation process via cross-attention, enabling both self- and cross-identity reenactment. Experiments on the VFHQ dataset show accurate and robust reconstruction performance and improved view point stability compared to the state-of-the-art. We propose a full end-to-end pipeline which is able to create, animate, and render avatars at over 20 fps, enabling new applications in live telepresence and interactive virtual characters.
Björn Browatzki, Niall Murray, Nilesh Jain
IMX3
2026 ARGO+: Achieving Multi-Level Scalable GNN Training on Distributed Multi-Core Platform
abstract
Graph Neural Networks (GNNs) have emerged as powerful tools for learning from graph-structured data and are widely used in various applications such as traffic prediction, and Electronic Design Automation, among others. However, training GNNs on large-scale graphs with hundreds of millions of nodes and billions of edges is time-consuming, often requiring days or weeks on a single machine. While distributed GNN training across multiple machines has been explored to leverage greater computation and memory resources, existing approaches primarily focus on inter-machine scalability while overlooking intra-machine scalability across multiple cores. As a result, state-of-the-art distributed GNN training frameworks lead to severe resource underutilization and limited training performance in terms of epoch time. This work introduces ARGO+, a novel GNN system designed to achieve multi-level scalability for distributed GNN training by efficiently scaling across both inter- and intra-machine levels. ARGO+ features a two-level graph partitioning strategy, exploiting parallelisms across both graph topology and feature dimensions while minimizing communication overhead and ensuring balanced workloads. In addition, ARGO+ integrates NUMA-aware optimizations to enhance intra-node scalability, addressing inefficiencies in remote-socket data accessing. During runtime, ARGO+ adopts a two-stage parallel training scheme to further reduce communication overheads and instantiates multiple training processes to exploit computation-communication overlapping. We evaluate ARGO+ on a distributed multi-core CPU cluster consisting of 8 machines, each with a dual-socket 80-core Intel Xeon processor. Our results demonstrate that ARGO+ achieves up to 1.89× speedup compared with DistDGL, and 1.53-1.56× speedup compared with state-of-the-art distributed GNN training systems. In addition, ARGO+ is compatible with the Deep Graph Library (DGL), a widely used GNN framework, enabling seamless integration with existing GNN programs. Finally, while ARGO+ adopts various optimizations to improve scalability and training performance, these optimizations do not alter the semantics of the training algorithm; thus, the model accuracy and convergence remain consistent with the original implementation of the algorithm.
Yi-Chien Lin, Sameh Gobriel, Nilesh Jain, Viktor Prasanna 0001
IEEE Trans. Parallel Distributed Syst.4
2025 INRet: A General Framework for Accurate Retrieval of INRs for Shapes
abstract
Implicit neural representations (INRs) have become an important method for encoding various data types, such as 3D objects or scenes, images, and videos. They have proven to be particularly effective at representing 3D content, e.g., 3D scene reconstruction from 2D images, novel 3D content creation, as well as the representation, interpolation and completion of 3D shapes. With the widespread generation of 3D data in an INR format, there is a need to support effective organization and retrieval of INRs saved in a data store. A key aspect of retrieval and clustering of INRs in a data store is the formulation of similarity between INRs that would, for example, enable retrieval of similar INRs using a query INR. In this work, we propose INRet (INR Retrieve), a method for determining similarity between INRs that represent shapes, thus enabling accurate retrieval of similar shape INRs from an INR data store. INRet flexibly supports different INR architectures such as INRs with octree grids, triplanes, and hash grids, as well as different implicit functions including signed/unsigned distance function and occupancy field. We demonstrate that our method is more general and accurate than the existing INR retrieval method, which only supports simple MLP INRs and requires the same architecture between the query and stored INRs. Furthermore, compared to converting INRs to other representations (e.g., point clouds or multi-view images) for 3D shape retrieval, INRet achieves higher accuracy while avoiding the conversion overhead.
Yushi Guan, Daniel Kwan, Ruofan Liang, Selvakumar Panneer, Nilesh Jain, Nilesh A. Ahuja, Nandita Vijaykumar
3DV5
2025 ContraGS: Codebook-Condensed and Trainable Gaussian Splatting for Fast, Memory-Efficient Reconstruction
Sankeerth Durvasula, Sharanshangar Muhunthan, Zain Moustafa, Ruofan Liang, Yushi Guan, Nilesh A. Ahuja, Nilesh Jain, Selvakumar Panneer, Nandita Vijaykumar
ICCV8
2025 Retri3D: 3D Neural Graphics Representation Retrieval
abstract
Learnable 3D Neural Graphics Representations (3DNGR) have emerged as promising 3D representations for reconstructing 3D scenes from 2D images. Numerous works, including Neural Radiance Fields (NeRF), 3D Gaussian Splatting (3DGS), and their variants, have significantly enhanced the quality of these representations. The ease of construction from 2D images, suitability for online viewing/sharing, and applications in game/art design downstream tasks make it a vital 3D representation, with potential creation of large numbers of such 3D models. This necessitates large data stores, local or online, to save 3D visual data in these formats. However, no existing framework enables accurate retrieval of stored 3DNGRs. In this work, we propose, Retri3D, a framework that enables accurate and efficient retrieval of 3D scenes represented as NGRs from large data stores using text queries. We introduce a novel Neural Field Artifact Analysis technique, combined with a Smart Camera Movement Module, to select clean views and navigate pre-trained 3DNGRs. These techniques enable accurate retrieval by selecting the best viewing directions in the 3D scene for high-quality visual feature embeddings. We demonstrate that Retri3D is compatible with any NGR representation. On the LERF and ScanNet++ datasets, we show significant improvement in retrieval accuracy compared to existing techniques, while being orders of magnitude faster and storage efficient.
Yushi Guan, Daniel Kwan, Jean Sebastien Dandurand, Ruofan Liang, Nilesh Jain, Nilesh A. Ahuja, Selvakumar Panneer, Nandita Vijaykumar
ICLR7
2025 Mamba-Shedder: Post-Transformer Compression for Efficient Selective Structured State Space Models
abstract
Juan Pablo Munoz, Jinjie Yuan, Nilesh Jain. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Juan Pablo Muñoz, Jinjie Yuan, Nilesh Jain
NAACL (Long Papers)3
2025 Accelerating GNN Inference via Automated Parallel Execution on Edge Heterogeneous Platforms
abstract
Recently, Graph Neural Networks (GNN) have been integrated into various local applications, such as local community detection and local code assistant, making edge inference increasingly important. To support diverse workloads, state-of-the-art edge devices have evolved into heterogeneous platforms, integrating components like CPU, GPU, and NPU. To this end, we propose GNX, a novel GNN system that accelerates GNN inference on edge heterogeneous platforms by leveraging all the heterogeneous processing units. Given a GNN model and a heterogeneous platform, GNX automatically generates parallel execution plans, consisting of both data and pipeline parallelism. To reduce the complexity of the design space, GNX converts GNN models into coarse-grained blocks and performs the search at the block level. By leveraging the APIs provided by state-of-the-art heterogeneous frameworks, GNX can flexibly schedule various parallel execution plans and seamlessly adjust the workload across the heterogeneous processing units for load-balanced execution. Our study shows that GNX effectively accelerates three widely-used GNN models on two state-of-the-art edge heterogeneous platforms. Compared with the baseline approach that uses only a single processing unit, GNX achieves up to a 2.57× speedup. Compared with adopting data parallelism and a state-of-the-art scheduler, GNX achieves up to 1.90× and 1.79× speedup, respectively. We also discuss the applicability of and extensions to GNX to support other GNN models.
Yi-Chien Lin, Haoyang Fan, Sameh Gobriel, Nilesh Jain, Viktor Prasanna 0001
SBAC-PAD4
2024 LoNAS: Elastic Low-Rank Adapters for Efficient Large Language Models
abstract
Large Language Models (LLMs) continue to grow, reaching hundreds of billions of parameters and making it challenging for Deep Learning practitioners with resource-constrained systems to use them, e.g., fine-tuning these models for a downstream task of their interest. Adapters, such as low-rank adapters (LoRA), have been proposed to reduce the number of trainable parameters in a model, reducing memory requirements and enabling smaller systems to fine-tune these models. Orthogonal to this work, Neural Architecture Search (NAS) has been used to discover compressed and more efficient architectures without sacrificing performance compared to similar base models. This paper introduces a novel approach, LoNAS, to use NAS on language models by exploring a search space of elastic low-rank adapters while reducing memory and compute requirements of full-scale NAS, resulting in high-performing compressed models obtained from weight-sharing super-networks. Compared to models fine-tuned with LoRA, these models contain fewer total parameters, reducing the inference time with only minor decreases in accuracy and, in some cases, even improving accuracy. We discuss the limitations of LoNAS and share observations for the research community regarding its generalization capabilities, which have motivated our follow-up work.
Juan Pablo Muñoz, Jinjie Yuan, Nilesh Jain
LREC/COLING4
2024 EFTNAS: Searching for Efficient Language Models in First-Order Weight-Reordered Super-Networks
abstract
Transformer-based models have demonstrated outstanding performance in natural language processing (NLP) tasks and many other domains, e.g., computer vision. Depending on the size of these models, which have grown exponentially in the past few years, machine learning practitioners might be restricted from deploying them in resource-constrained environments. This paper discusses the compression of transformer-based models for multiple resource budgets. Integrating neural architecture search (NAS) and network pruning techniques, we effectively generate and train weight-sharing super-networks that contain efficient, high-performing, and compressed transformer-based models. A common challenge in NAS is the design of the search space, for which we propose a method to automatically obtain the boundaries of the search space and then derive the rest of the intermediate possible architectures using a first-order weight importance technique. The proposed end-to-end NAS solution, EFTNAS, discovers efficient subnetworks that have been compressed and fine-tuned for downstream NLP tasks. We demonstrate EFTNAS on the General Language Understanding Evaluation (GLUE) benchmark and the Stanford Question Answering Dataset (SQuAD), obtaining high-performing smaller models with a reduction of more than 5x in size without or with little degradation in performance.
Juan Pablo Muñoz, Nilesh Jain
LREC/COLING3
2024 Textual-Visual Logic Challenge: Understanding and Reasoning in Text-to-Image Generation
Peixi Xiong, Michael A. Kozuch, Nilesh Jain
ECCV (5)3
2024 ARGO: An Auto-Tuning Runtime System for Scalable GNN Training on Multi-Core Processor
abstract
As Graph Neural Networks (GNNs) become popular, libraries like PyTorch-Geometric (PyG) and Deep Graph Library (DGL) are proposed; these libraries have emerged as the de facto standard for implementing GNNs because they provide graph-oriented APIs and are purposefully designed to manage the inherent sparsity and irregularity in graph structures. However, these libraries show poor scalability on multi-core processors, which under-utilizes the available platform resources and limits the performance. This is because GNN training is a resource-intensive workload with high volume of irregular data accessing, and existing libraries fail to utilize the memory bandwidth efficiently. To address this challenge, we propose ARGO, a novel runtime system for GNN training that offers scalable performance. ARGO exploits multi-processing and core-binding techniques to improve platform resource utilization. We further develop an auto-tuner that searches for the optimal configuration for multi-processing and core-binding. The auto-tuner works automatically, making it completely transparent from the user. Furthermore, the auto-tuner allows ARGO to adapt to various platforms, GNN models, datasets, etc. We evaluate ARGO on two representative GNN models and four widely-used datasets on two platforms. With the proposed autotuner, ARGO is able to select a near-optimal configuration by exploring only 5% of the design space. ARGO speeds up state-of-the-art GNN libraries by up to 5.06× and 4.54× on a four-socket Ice Lake machine with 112 cores and a two-socket Sapphire Rapids machine with 64 cores, respectively. Finally, ARGO can seamlessly integrate into widely-used GNN libraries (e.g., DGL, PyG) with few lines of code and speed up GNN training.
Yi-Chien Lin, Sameh Gobriel, Nilesh Jain, Gopi Krishna Jha, Viktor Prasanna 0001
IPDPS4
2024 Distributed Training of Neural Radiance Fields: A Performance Characterization
abstract
Implicit neural representation is an emerging method that leverages deep neural networks and learned parameters to represent 3D scenes efficiently and accurately. Neural radiance field (NeRF) is a state-of-art implicit representation that achieves photorealistic 3D reconstruction with compact neural network models. However, as the complexity and scale of the scene increase, training NeRF models with a single GPU proves insufficient for achieving fast training and high-quality reconstruction. To address this challenge, prior works proposed distributed NeRF training methods. This is the first work to conduct a detailed evaluation of two major distributed NeRF training methods and their tradeoffs: distributed data parallel (DDP) and spatial segmentation (SS). We find that DDP training requires cross-device synchronization during training, while SS training incurs additional fusion overhead during inference. Our analysis also reveals that sampling input images is a common key bottleneck in distributed NeRF training. At the beginning of each training iteration, the CPU generates input batches for all GPUs in the cluster by sampling all images in the dataset, causing significant stalls that constitute up to 43.3% of the total training time. To alleviate this bottleneck, we propose a pipelined input sampling strategy that precomputes input samples on the CPU concurrently with model training on the GPUs. Our evaluation demonstrates an average speedup in training time by$1.95\times($up to$2.24\times)$.
Adrian Zhao, Louis Zhang, Sankeerth Durvasula, Nilesh Jain, Selvakumar Panneer, Nandita Vijaykumar
ISPASS5
2023 Mem-Rec: Memory Efficient Recommendation System using Alternative Representation
Gopi Krishna Jha, Anthony Thomas, Nilesh Jain, Sameh Gobriel, Tajana Rosing, Ravi R. Iyer 0001
ACML3
2022 Neuroevolution-enhanced multi-objective optimization for mixed-precision quantization
abstract
Mixed-precision quantization is a powerful tool to enable memory and compute savings of neural network workloads by deploying different sets of bit-width precisions on separate compute operations. In this work, we present a flexible and scalable framework for automated mixed-precision quantization that concurrently optimizes task performance, memory compression, and compute savings through multi-objective evolutionary computing. Our framework centers on Neuroevolution-Enhanced Multi-Objective Optimization (NEMO), a novel search method, which combines established search methods with the representational power of neural networks. Within NEMO, the population is divided into structurally distinct sub-populations, or species, which jointly create the Pareto frontier of solutions for the multi-objective problem. At each generation, species perform separate mutation and crossover operations, and are re-sized in proportion to the goodness of their contribution to the Pareto frontier. In our experiments, we define a graph-based representation to describe the underlying workload, enabling us to deploy graph neural networks trained by NEMO via neuroevolution, to find Pareto optimal configurations for MobileNet-V2, ResNet50 and ResNeXt-101-32×8d. Compared to the state-of-the-art, we achieve competitive results on memory compression and superior results for compute compression. Further analysis reveals that the graph representation and the species-based approach employed by NEMO are critical to finding optimal solutions.
Santiago Miret, Vui Seng Chua, Mattias Marder, Mariano Phielipp, Nilesh Jain, Somdeb Majumdar
GECCO5
2022 EZNAS: Evolving Zero-Cost Proxies For Neural Architecture Scoring
abstract
Neural Architecture Search (NAS) has significantly improved productivity in the design and deployment of neural networks (NN). As NAS typically evaluates multiple models by training them partially or completely, the improved productivity comes at the cost of significant carbon footprint. To alleviate this expensive training routine, zero-shot/cost proxies analyze an NN at initialization to generate a score, which correlates highly with its true accuracy. Zero-cost proxies are currently designed by experts conducting multiple cycles of empirical testing on possible algorithms, datasets, and neural architecture design spaces. This experimentation lowers productivity and is an unsustainable approach towards zero-cost proxy design as deep learning use-cases diversify in nature. Additionally, existing zero-cost proxies fail to generalize across neural architecture design spaces. In this paper, we propose a genetic programming framework to automate the discovery of zero-cost proxies for neural architecture scoring. Our methodology efficiently discovers an interpretable and generalizable zero-cost proxy that gives state of the art score-accuracy correlation on all datasets and search spaces of NASBench-201 and Network Design Spaces (NDS). We believe that this research indicates a promising direction towards automatically discovering zero-cost proxies that can work across network architecture design spaces, datasets, and tasks.
Yash Akhauri, Juan Pablo Muñoz, Nilesh Jain, Ravi R. Iyer 0001
NeurIPS3
2021 A 93 TOPS/Watt Near-Memory Reconfigurable SAD Accelerator for HEVC/AV1/JEM Encoding
abstract
Motion Estimation (ME) is a major bottleneck of a Video encoding pipeline. This paper presents a low power near memory Sum of Absolute Difference (SAD) accelerator for ME. The accelerator is composed of 64 modular SAD Processing Elements (PEs) on a Reconfigurable fabric, offering maximal parallelism to support traditional and futuristic Rate-Distortion-Optimization (RDO) schemes consistent with HEVC/AV1/JEM. The accelerator offers up-to 55% speedup over State-of-art accelerators and a 7x speedup when compared to a 12 core Intel Xeon E5 processor. Our solution achieves 93 TOPS/Watt running at 500MHz frequency, capable of processing real-time 4K 30fps video. Synthesized in 22nm process, the accelerator occupies 0.08mm2 and consumes 5.46mW dynamic power.
Jainaveen Sundaram, Srivatsa Rangachar Srinivasa, Dileep Kurian, Indranil Chakraborty, Sirisha Rani Kale, Nilesh Jain, Tanay Karnik, Ravi R. Iyer 0001, Anuradha Srinivasan
DATE6
2021 E2E Visual Analytics: Achieving >10X Edge/Cloud Optimizations
abstract
As visual analytics continues to rapidly grow, there is a critical need to improve the end-to-end efficiency of visual processing in edge/cloud systems. In this paper, we cover algorithms, systems and optimizations in three major areas for edge/cloud visual processing: (1) addressing storage and retrieval efficiency of visual data and meta-data by employing and optimizing visual data management systems, (2) addressing compute efficiency of visual analytics by taking advantage of co-optimization between the compression and analytics domains and (3) addressing networking (bandwidth) efficiency of visual data compression by tailoring it based on analytics tasks. We describe techniques in each of the above areas and measure its efficacy on state-of-the-art platforms (Intel Xeon), workloads and datasets. Our results show that we can achieve >10X improvements in each area based on novel algorithms, systems, and co-design optimizations. We also outline future research directions based on our findings which outline areas of further performance and efficiency advantages in end-to-end visual analytics.
Chaunte W. Lacewell, Nilesh A. Ahuja, Juan Pablo Muñoz, Parual Datta, Ragaad AlTarawneh, Vui Seng Chua, Nilesh Jain, Omesh Tickoo, Ravi R. Iyer 0001
NAS7
2021 Declarative Data Serving: The Future of Machine Learning Inference on the Edge
abstract
Recent advances in computer architecture and networking have ushered in a new age of edge computing, where computation is placed close to the point of data collection to facilitate low-latency decision making. As the complexity of such deployments grow into networks of interconnected edge devices, getting the necessary data to be in "the right place at the right time" can become a challenge. We envision a future of edge analytics where data flows between edge nodes are declaratively configured through high-level constraints. Using machine learning model-serving as a prototypical task, we illustrate how the heterogeneity and specialization of edge devices can lead to complex, task-specific communication patterns even in relatively simple situations. Without a declarative framework, managing this complexity will be challenging for developers and will lead to brittle systems. We conclude with a research vision for database community that brings our perspective to the emergent area of edge computing.
Ted Shaowang, Nilesh Jain, Dennis Matthews, Sanjay Krishnan
Proc. VLDB Endow.2
2016 Generalized activity recognition using accelerometer in wearable devices for IoT applications
abstract
The proliferation of low power and low cost continuous sensing has generated an immense interest in the area of activity recognition. However, the real time detection is still a challenge for several reasons: requirement from the user to specify the type of activity, complex algorithms, and collection of data from multiple devices. In this paper, we describe a generalized activity recognition system, its applications, and the challenges involved in implementing the algorithm in resource-constrained devices. The distinctive aspects of our study include: 1) automatic detection and recognition of different activities (running, walking, crawling, climbing, and pronating), 2) using just one axis from an accelerometer sensor, and 3) simple features and pattern matching algorithm leading to computationally inexpensive and memory efficient system suitable for resource-constrained wearable devices. The activity recognition model was trained using data collected from 52 unique subjects. The model was mapped onto Intel® Quark™ SE Pattern Matching Engine, and field-tested using eight additional subjects achieving performance up to 91%.
Ebrahim Al Safadi, Fahim Mohammad, Darshan Iyer, Benjamin J. Smiley, Nilesh Jain
AVSS5