Nikhil Parthasarathy

dblp:209/4951 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Efficient and distributed learning · 37% Representation and self-supervised learning · 29% Language models and text generation · 12%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 16 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
data-efficient learning
1.622025
Active Data Curation Effectively Distills Large-Scale Multimodal Models · CVPR 2025
Data curation via joint example selection further accelerates multimodal learning · NeurIPS 2024
Machine learning › Representation and self-supervised learning
pre-training
0.922024
Towards In-context Scene Understanding · NeurIPS 2023
Data curation via joint example selection further accelerates multimodal learning · NeurIPS 2024
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.912025
Active Data Curation Effectively Distills Large-Scale Multimodal Models · CVPR 2025
Machine learning › Efficient and distributed learning › model compression › knowledge distillation › cross-modal distillation
multimodal distillation
0.912025
Active Data Curation Effectively Distills Large-Scale Multimodal Models · CVPR 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.812024
Data curation via joint example selection further accelerates multimodal learning · NeurIPS 2024
Machine learning › Efficient and distributed learning
data curation
0.812024
Data curation via joint example selection further accelerates multimodal learning · NeurIPS 2024
Machine learning › Representation and self-supervised learning › contrastive learning
multimodal contrastive learning
0.812024
Data curation via joint example selection further accelerates multimodal learning · NeurIPS 2024
Natural language and speech › Language models and text generation
in-context learning
0.712023
Towards In-context Scene Understanding · NeurIPS 2023
Natural language and speech › Language models and text generation › large language model training › language model pretraining
in-context pretraining
0.712023
Towards In-context Scene Understanding · NeurIPS 2023
Computer vision › Segmentation and scene understanding
scene understanding
0.712023
Towards In-context Scene Understanding · NeurIPS 2023
Machine learning › Representation and self-supervised learning › pre-training › spatio-temporal pre-training
video pretraining
0.712023
Self-supervised video pretraining yields robust and more human-aligned visual representations · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
approximate bayesian inference
0.312017
Neural Networks for Efficient Bayesian Decoding of Natural Images from Retinal Neurons · NIPS 2017
Computer vision › 3D vision
brain decoding
0.312017
Neural Networks for Efficient Bayesian Decoding of Natural Images from Retinal Neurons · NIPS 2017
Computer vision › Vision and language
image-text retrieval
0.312025
Active Data Curation Effectively Distills Large-Scale Multimodal Models · CVPR 2025
Computer vision › 3D vision
depth estimation
0.212023
Towards In-context Scene Understanding · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning
0.212023
Self-supervised video pretraining yields robust and more human-aligned visual representations · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

contrastive learning · 1.5progressive freezing · 0.9knowledge distillation · 0.9model approximation · 0.8joint example selection · 0.8nearest neighbor retrieval · 0.7attention · 0.7convolutional autoencoder · 0.6artificial neural network · 0.6
YearPublicationVenuePosition
2025 Active Data Curation Effectively Distills Large-Scale Multimodal Models
abstract
Knowledge distillation (KD) is the de facto standard for compressing large-scale multimodal models into smaller ones. Prior works have explored ever more complex KD strategies involving different objectives, teacher-ensembles, and weight inheritance. In this work, we explore an alternative, yet simple approach—active data curation as effective distillation for contrastive multimodal pretraining. Our simple online batch selection method, ACID, outperforms strong KD baselines across various model-,data-and compute-configurations. Further, we find such an active curation strategy to in fact be complementary to standard KD, and can be effectively combined to train highly performant inference-efficient models. Our simple and scalable pretraining framework, ACED, achieves state-of-the-art results across 27 zero-shot classification and image-text retrieval tasks with upto 11% less inference FLOPs. We further demonstrate that ACED yields strong vision-encoders for training generative multimodal models, outperforming larger vision encoders on image-captioning and visual question-answering tasks.
Vishaal Udandarao, Nikhil Parthasarathy, Muhammad Ferjad Naeem, Talfan Evans, Samuel Albanie, Federico Tombari, Yongqin Xian, Alessio Tonioni, Olivier J. Hénaff
CVPR2
2025 LayerLock: Non-Collapsing Representation Learning with Progressive Freezing
Goker Erdogan, Nikhil Parthasarathy, Catalin Ionescu, Drew A. Hudson, Alexander Lerchner, Andrew Zisserman, Mehdi S. M. Sajjadi, João Carreira 0001
ICCV2
2024 Data curation via joint example selection further accelerates multimodal learning
abstract
Data curation is an essential component of large-scale pretraining. In this work, we demonstrate that jointly prioritizing batches of data is more effective for learning than selecting examples independently. Multimodal contrastive objectives expose the dependencies between data and thus naturally yield criteria for measuring the joint learnability of a batch. We derive a simple and tractable algorithm for selecting such batches, which significantly accelerate training beyond individually-prioritized data points. As performance improves by selecting from large super-batches, we also leverage recent advances in model approximation to reduce the computational overhead of scoring. As a result, our approach—multimodal contrastive learning with joint example selection (JEST)—surpasses state-of-the-art pretraining methods with up to 13× fewer iterations and 10× less computation. Essential to the performance of JEST is the ability to steer the data selection process towards the distribution of smaller, well-curated datasets via pretrained reference models, exposing data curation as a new dimension for neural scaling laws.
Talfan Evans, Nikhil Parthasarathy, Hamza Merzic, Olivier J. Hénaff
NeurIPS2
2023 Towards In-context Scene Understanding
abstract
In-context learning––the ability to configure a model's behavior with different prompts––has revolutionized the field of natural language processing, alleviating the need for task-specific models and paving the way for generalist models capable of assisting with any query. Computer vision, in contrast, has largely stayed in the former regime: specialized decoders and finetuning protocols are generally required to perform dense tasks such as semantic segmentation and depth estimation. In this work we explore a simple mechanism for in-context learning of such scene understanding tasks: nearest neighbor retrieval from a prompt of annotated features. We propose a new pretraining protocol––leveraging attention within and across images––which yields representations particularly useful in this regime. The resulting Hummingbird model, suitably prompted, performs various scene understanding tasks without modification while approaching the performance of specialists that have been finetuned for each task. Moreover, Hummingbird can be configured to perform new tasks much more efficiently than finetuned models, raising the possibility of scene understanding in the interactive assistant regime.
Ivana Balazevic, David Steiner 0004, Nikhil Parthasarathy, Relja Arandjelovic, Olivier J. Hénaff
NeurIPS3
2023 Self-supervised video pretraining yields robust and more human-aligned visual representations
abstract
Humans learn powerful representations of objects and scenes by observing how they evolve over time. Yet, outside of specific tasks that require explicit temporal understanding, static image pretraining remains the dominant paradigm for learning visual foundation models. We question this mismatch, and ask whether video pretraining can yield visual representations that bear the hallmarks of human perception: generalisation across tasks, robustness to perturbations, and consistency with human judgements. To that end we propose a novel procedure for curating videos, and develop a contrastive framework which learns from the complex transformations therein. This simple paradigm for distilling knowledge from videos, called VITO, yields general representations that far outperform prior video pretraining methods on image understanding tasks, and image pretraining methods on video understanding tasks. Moreover, VITO representations are significantly more robust to natural and synthetic deformations than image-, video-, and adversarially-trained ones. Finally, VITO’s predictions are strongly aligned with human judgements, surpassing models that were specifically trained for that purpose. Together, these results suggest that video pretraining could be a simple way of learning unified, robust, and human-aligned representations of the visual world.
Nikhil Parthasarathy, S. M. Ali Eslami, João Carreira 0001, Olivier J. Hénaff
NeurIPS1
2017 Neural Networks for Efficient Bayesian Decoding of Natural Images from Retinal Neurons
abstract
Decoding sensory stimuli from neural signals can be used to reveal how we sense our physical environment, and is valuable for the design of brain-machine interfaces. However, existing linear techniques for neural decoding may not fully reveal or exploit the fidelity of the neural signal. Here we develop a new approximate Bayesian method for decoding natural images from the spiking activity of populations of retinal ganglion cells (RGCs). We sidestep known computational challenges with Bayesian inference by exploiting artificial neural networks developed for computer vision, enabling fast nonlinear decoding that incorporates natural scene statistics implicitly. We use a decoder architecture that first linearly reconstructs an image from RGC spikes, then applies a convolutional autoencoder to enhance the image. The resulting decoder, trained on natural images and simulated neural responses, significantly outperforms linear decoding, as well as simple point-wise nonlinear decoding. These results provide a tool for the assessment and optimization of retinal prosthesis technologies, and reveal that the retina may provide a more accurate representation of the visual scene than previously appreciated.
Nikhil Parthasarathy, Eleanor Batty, William Falcon, Thomas Rutten, Mohit Rajpal, E. J. Chichilnisky, Liam Paninski
NIPS1