VLDB 2026 Research / reviewers in the wild / expert
Nikhil Parthasarathy
dblp:209/4951
· DBLP profile ↗
6ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Efficient and distributed learning · 37% Representation and self-supervised learning · 29% Language models and text generation · 12% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 16 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
data-efficient learning |
1.6 | 2 | 2025 | Active Data Curation Effectively Distills Large-Scale Multimodal Models · CVPR 2025 Data curation via joint example selection further accelerates multimodal learning · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning
pre-training |
0.9 | 2 | 2024 | Towards In-context Scene Understanding · NeurIPS 2023 Data curation via joint example selection further accelerates multimodal learning · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.9 | 1 | 2025 | Active Data Curation Effectively Distills Large-Scale Multimodal Models · CVPR 2025 |
Machine learning › Efficient and distributed learning › model compression › knowledge distillation › cross-modal distillation
multimodal distillation |
0.9 | 1 | 2025 | Active Data Curation Effectively Distills Large-Scale Multimodal Models · CVPR 2025 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.8 | 1 | 2024 | Data curation via joint example selection further accelerates multimodal learning · NeurIPS 2024 |
Machine learning › Efficient and distributed learning
data curation |
0.8 | 1 | 2024 | Data curation via joint example selection further accelerates multimodal learning · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning › contrastive learning
multimodal contrastive learning |
0.8 | 1 | 2024 | Data curation via joint example selection further accelerates multimodal learning · NeurIPS 2024 |
Natural language and speech › Language models and text generation
in-context learning |
0.7 | 1 | 2023 | Towards In-context Scene Understanding · NeurIPS 2023 |
Natural language and speech › Language models and text generation › large language model training › language model pretraining
in-context pretraining |
0.7 | 1 | 2023 | Towards In-context Scene Understanding · NeurIPS 2023 |
Computer vision › Segmentation and scene understanding
scene understanding |
0.7 | 1 | 2023 | Towards In-context Scene Understanding · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning › pre-training › spatio-temporal pre-training
video pretraining |
0.7 | 1 | 2023 | Self-supervised video pretraining yields robust and more human-aligned visual representations · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
approximate bayesian inference |
0.3 | 1 | 2017 | Neural Networks for Efficient Bayesian Decoding of Natural Images from Retinal Neurons · NIPS 2017 |
Computer vision › 3D vision
brain decoding |
0.3 | 1 | 2017 | Neural Networks for Efficient Bayesian Decoding of Natural Images from Retinal Neurons · NIPS 2017 |
Computer vision › Vision and language
image-text retrieval |
0.3 | 1 | 2025 | Active Data Curation Effectively Distills Large-Scale Multimodal Models · CVPR 2025 |
Computer vision › 3D vision
depth estimation |
0.2 | 1 | 2023 | Towards In-context Scene Understanding · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning |
0.2 | 1 | 2023 | Self-supervised video pretraining yields robust and more human-aligned visual representations · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 1.5progressive freezing · 0.9knowledge distillation · 0.9model approximation · 0.8joint example selection · 0.8nearest neighbor retrieval · 0.7attention · 0.7convolutional autoencoder · 0.6artificial neural network · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Active Data Curation Effectively Distills Large-Scale Multimodal ModelsabstractKnowledge distillation (KD) is the de facto standard for compressing large-scale multimodal models into smaller ones. Prior works have explored ever more complex KD strategies involving different objectives, teacher-ensembles, and weight inheritance. In this work, we explore an alternative, yet simple approach—active data curation as effective distillation for contrastive multimodal pretraining. Our simple online batch selection method, ACID, outperforms strong KD baselines across various model-,data-and compute-configurations. Further, we find such an active curation strategy to in fact be complementary to standard KD, and can be effectively combined to train highly performant inference-efficient models. Our simple and scalable pretraining framework, ACED, achieves state-of-the-art results across 27 zero-shot classification and image-text retrieval tasks with upto 11% less inference FLOPs. We further demonstrate that ACED yields strong vision-encoders for training generative multimodal models, outperforming larger vision encoders on image-captioning and visual question-answering tasks. Vishaal Udandarao, Nikhil Parthasarathy, Muhammad Ferjad Naeem, Talfan Evans, Samuel Albanie, Federico Tombari, Yongqin Xian, Alessio Tonioni, Olivier J. Hénaff |
CVPR | 2 |
| 2025 | LayerLock: Non-Collapsing Representation Learning with Progressive Freezing
Goker Erdogan, Nikhil Parthasarathy, Catalin Ionescu, Drew A. Hudson, Alexander Lerchner, Andrew Zisserman, Mehdi S. M. Sajjadi, João Carreira 0001 |
ICCV | 2 |
| 2024 | Data curation via joint example selection further accelerates multimodal learningabstractData curation is an essential component of large-scale pretraining. In this work, we demonstrate that jointly prioritizing batches of data is more effective for learning than selecting examples independently. Multimodal contrastive objectives expose the dependencies between data and thus naturally yield criteria for measuring the joint learnability of a batch. We derive a simple and tractable algorithm for selecting such batches, which significantly accelerate training beyond individually-prioritized data points. As performance improves by selecting from large super-batches, we also leverage recent advances in model approximation to reduce the computational overhead of scoring. As a result, our approach—multimodal contrastive learning with joint example selection (JEST)—surpasses state-of-the-art pretraining methods with up to 13× fewer iterations and 10× less computation. Essential to the performance of JEST is the ability to steer the data selection process towards the distribution of smaller, well-curated datasets via pretrained reference models, exposing data curation as a new dimension for neural scaling laws. Talfan Evans, Nikhil Parthasarathy, Hamza Merzic, Olivier J. Hénaff |
NeurIPS | 2 |
| 2023 | Towards In-context Scene UnderstandingabstractIn-context learning––the ability to configure a model's behavior with different prompts––has revolutionized the field of natural language processing, alleviating the need for task-specific models and paving the way for generalist models capable of assisting with any query. Computer vision, in contrast, has largely stayed in the former regime: specialized decoders and finetuning protocols are generally required to perform dense tasks such as semantic segmentation and depth estimation. In this work we explore a simple mechanism for in-context learning of such scene understanding tasks: nearest neighbor retrieval from a prompt of annotated features. We propose a new pretraining protocol––leveraging attention within and across images––which yields representations particularly useful in this regime. The resulting Hummingbird model, suitably prompted, performs various scene understanding tasks without modification while approaching the performance of specialists that have been finetuned for each task. Moreover, Hummingbird can be configured to perform new tasks much more efficiently than finetuned models, raising the possibility of scene understanding in the interactive assistant regime. Ivana Balazevic, David Steiner 0004, Nikhil Parthasarathy, Relja Arandjelovic, Olivier J. Hénaff |
NeurIPS | 3 |
| 2023 | Self-supervised video pretraining yields robust and more human-aligned visual representationsabstractHumans learn powerful representations of objects and scenes by observing how they evolve over time. Yet, outside of specific tasks that require explicit temporal understanding, static image pretraining remains the dominant paradigm for learning visual foundation models. We question this mismatch, and ask whether video pretraining can yield visual representations that bear the hallmarks of human perception: generalisation across tasks, robustness to perturbations, and consistency with human judgements. To that end we propose a novel procedure for curating videos, and develop a contrastive framework which learns from the complex transformations therein. This simple paradigm for distilling knowledge from videos, called VITO, yields general representations that far outperform prior video pretraining methods on image understanding tasks, and image pretraining methods on video understanding tasks. Moreover, VITO representations are significantly more robust to natural and synthetic deformations than image-, video-, and adversarially-trained
ones. Finally, VITO’s predictions are strongly aligned with human judgements, surpassing models that were specifically trained for that purpose. Together, these results suggest that video pretraining could be a simple way of learning unified, robust, and human-aligned representations of the visual world. Nikhil Parthasarathy, S. M. Ali Eslami, João Carreira 0001, Olivier J. Hénaff |
NeurIPS | 1 |
| 2017 | Neural Networks for Efficient Bayesian Decoding of Natural Images from Retinal NeuronsabstractDecoding sensory stimuli from neural signals can be used to reveal how we sense our physical environment, and is valuable for the design of brain-machine interfaces. However, existing linear techniques for neural decoding may not fully reveal or exploit the fidelity of the neural signal. Here we develop a new approximate Bayesian method for decoding natural images from the spiking activity of populations of retinal ganglion cells (RGCs). We sidestep known computational challenges with Bayesian inference by exploiting artificial neural networks developed for computer vision, enabling fast nonlinear decoding that incorporates natural scene statistics implicitly. We use a decoder architecture that first linearly reconstructs an image from RGC spikes, then applies a convolutional autoencoder to enhance the image. The resulting decoder, trained on natural images and simulated neural responses, significantly outperforms linear decoding, as well as simple point-wise nonlinear decoding. These results provide a tool for the assessment and optimization of retinal prosthesis technologies, and reveal that the retina may provide a more accurate representation of the visual scene than previously appreciated. Nikhil Parthasarathy, Eleanor Batty, William Falcon, Thomas Rutten, Mohit Rajpal, E. J. Chichilnisky, Liam Paninski |
NIPS | 1 |