Neehar Kondapaneni

dblp:280/3261 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Image recognition and object detection · 18% Trustworthy machine learning · 16% Representation and self-supervised learning · 16%
Human-computer interaction and pervasive computing
1 paper
Learning and educational technologies · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
1.722025
Representational Difference Explanations · NeurIPS 2025
Representational Similarity via Interpretable Visual Concepts · ICLR 2025
Machine learning › Representation and self-supervised learning › representation analysis
representation similarity
0.912025
Representational Similarity via Interpretable Visual Concepts · ICLR 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.912025
A Closer Look at Benchmarking Self-supervised Pre-training with Image Classification · Int. J. Comput. Vis. 2025
Machine learning › Transfer learning and domain adaptation › transferability estimation
transfer evaluation
0.912025
A Closer Look at Benchmarking Self-supervised Pre-training with Image Classification · Int. J. Comput. Vis. 2025
Computer vision › Image recognition and object detection › visual concept learning
visual concept discovery
0.912025
Representational Similarity via Interpretable Visual Concepts · ICLR 2025
Computer vision › 3D vision
depth estimation
0.812024
Text-Image Alignment for Diffusion-Based Perception · CVPR 2024
Computer vision › Segmentation and scene understanding › semantic segmentation
diffusion-based segmentation
0.812024
Text-Image Alignment for Diffusion-Based Perception · CVPR 2024
Computer vision › Vision and language › cross-modal alignment
image-text alignment
0.812024
Text-Image Alignment for Diffusion-Based Perception · CVPR 2024
Computer vision › 3D vision › depth estimation
monocular depth estimation
0.812024
Text-Image Alignment for Diffusion-Based Perception · CVPR 2024
Computer vision › Image recognition and object detection › object detection
open-vocabulary object detection
0.812024
Text-Image Alignment for Diffusion-Based Perception · CVPR 2024
Computer vision › Segmentation and scene understanding
semantic segmentation
0.812024
Text-Image Alignment for Diffusion-Based Perception · CVPR 2024
Machine learning › Time series and sequential data › temporal sequence modeling
knowledge tracing
0.612022
Visual Knowledge Tracing · ECCV (25) 2022
Computer vision › Image recognition and object detection
image classification
0.312025
Representational Difference Explanations · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

visual knowledge tracing · 1.1representational similarity analysis · 0.9representation analysis · 0.9post hoc XAI · 0.9linear probing · 0.9k-nearest neighbors · 0.9fine-tuning · 0.9model personalization · 0.8diffusion model · 0.8caption generation · 0.8
YearPublicationVenuePosition
2025 Representational Similarity via Interpretable Visual Concepts
abstract
How do two deep neural networks differ in how they arrive at a decision? Measuring the similarity of deep networks has been a long-standing open question. Most existing methods provide a single number to measure the similarity of two networks at a given layer, but give no insight into what makes them similar or dissimilar. We introduce an interpretable representational similarity method (RSVC) to compare two networks. We use RSVC to discover shared and unique visual concepts between two models. We show that some aspects of model differences can be attributed to unique concepts discovered by one model that are not well represented in the other. Finally, we conduct extensive evaluation across different vision model architectures and training protocols to demonstrate its effectiveness.
Neehar Kondapaneni, Oisin Mac Aodha, Pietro Perona
ICLR1
2025 Representational Difference Explanations
abstract
We propose a method for discovering and visualizing the differences between two learned representations, enabling more direct and interpretable model comparisons. We validate our method, which we call Representational Differences Explanations (RDX), by using it to compare models with known conceptual differences and demonstrate that it recovers meaningful distinctions where existing explainable AI (XAI) techniques fail. Applied to state-of-the-art models on challenging subsets of the ImageNet and iNaturalist datasets, RDX reveals both insightful representational differences and subtle patterns in the data. Although comparison is a cornerstone of scientific analysis, current tools in machine learning, namely post hoc XAI methods, struggle to support model comparison effectively. Our work addresses this gap by introducing an effective and explainable tool for contrasting model representations.
Neehar Kondapaneni, Oisin Mac Aodha, Pietro Perona
NeurIPS1
2025 A Closer Look at Benchmarking Self-supervised Pre-training with Image Classification
abstract
Self-supervised learning (SSL) is a machine learning approach where the data itself provides supervision, eliminating the need for external labels. The model is forced to learn about the data's inherent structure or context by solving a pretext task. With SSL, models can learn from abundant and cheap unlabeled data, significantly reducing the cost of training models where labels are expensive or inaccessible. In Computer Vision, SSL is widely used as pre-training followed by a downstream task, such as supervised transfer, few-shot learning on smaller labeled data sets, and/or unsupervised clustering. Unfortunately, it is infeasible to evaluate SSL methods on all possible downstream tasks and objectively measure the quality of the learned representation. Instead, SSL methods are evaluated using in-domain evaluation protocols, such as fine-tuning, linear probing, and k-nearest neighbors (kNN). However, it is not well understood how well these evaluation protocols estimate the representation quality of a pre-trained model for different downstream tasks under different conditions, such as dataset, metric, and model architecture. In this work, we study how classification-based evaluation protocols for SSL correlate and how well they predict downstream performance on different dataset types. Our study includes eleven common image datasets and 26 models that were pre-trained with different SSL methods or have different model backbones. We find that in-domain linear/kNN probing protocols are, on average, the best general predictors for out-of-domain performance. We further investigate the importance of batch normalization for the various protocols and evaluate how robust correlations are for different kinds of dataset domain shifts. In addition, we challenge assumptions about the relationship between discriminative and generative self-supervised methods, finding that most of their performance differences can be explained by changes to model backbones. Supplementary Information: The online version contains supplementary material available at 10.1007/s11263-025-02402-w.
Markus Marks, Manuel Knott 0001, Neehar Kondapaneni, Elijah Cole, Thijs Defraeye, Fernando Pérez-Cruz, Pietro Perona
Int. J. Comput. Vis.3
2024 Text-Image Alignment for Diffusion-Based Perception
abstract
Diffusion models are generative models with impressive text-to-image synthesis capabilities and have spurred a new wave of creative methods for classical machine learning tasks. However, the best way to harness the perceptual knowledge of these generative models for visual tasks is still an open question. Specifically, it is unclear how to use the prompting interface when applying diffusion backbones to vision tasks. We find that automatically generated captions can improve text-image alignment and significantly enhance a model's cross-attention maps, leading to better perceptual performance. Our approach improves upon the current state-of-the-art (SOTA) in diffusionbased semantic segmentation on ADE20K and the current overall SOTA for depth estimation on NYUv2. Furthermore, our method generalizes to the cross-domain setting. We use model personalization and caption modifications to align our model to the target domain and find improvements over unaligned baselines. Our crossdomain object detection model, trained on Pascal VOC, achieves SOTA results on Watercolor2K. Our cross-domain segmentation method, trained on Cityscapes, achieves SOTA results on Dark Zurich-val and Nighttime Driving. Project page: vision.caltech.edu/TADP/ Code page: github.com/damaggu/TADP
Neehar Kondapaneni, Markus Marks, Manuel Knott 0001, Rogério Guimarães, Pietro Perona
CVPR1
2022 Visual Knowledge Tracing
Neehar Kondapaneni, Pietro Perona, Oisin Mac Aodha
ECCV (25)1