Arpita Chowdhury

dblp:228/4019 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Trustworthy machine learning · 54% Image recognition and object detection · 14% Vision and language · 11%

Topics — the 15 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
2.532025
Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation · CVPR 2025
Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained Analysis · CVPR 2025
A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis · ICLR 2024
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
1.932025
Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained Analysis · CVPR 2025
A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis · ICLR 2024
Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation · CVPR 2025
Machine learning › Trustworthy machine learning › interpretability › visual explanation
class activation map
1.722025
Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation · CVPR 2025
Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained Analysis · CVPR 2025
Machine learning › Trustworthy machine learning › interpretability › neural network interpretation
vision transformer explainability
0.912025
Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained Analysis · CVPR 2025
Machine learning › Trustworthy machine learning › interpretability › attention analysis
attention-based explanation
0.812024
A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis · ICLR 2024
Machine learning › Trustworthy machine learning
calibration
0.812024
Fine-Tuning is Fine, if Calibrated · NeurIPS 2024
Machine learning › Learning paradigms › continual learning
catastrophic forgetting
0.812024
Fine-Tuning is Fine, if Calibrated · NeurIPS 2024
Natural language and speech › Language models and text generation › large language model reasoning
comparative reasoning
0.812024
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs · NeurIPS 2024
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.812024
Fine-Tuning is Fine, if Calibrated · NeurIPS 2024
Machine learning › Trustworthy machine learning › calibration
logit calibration
0.812024
Fine-Tuning is Fine, if Calibrated · NeurIPS 2024
Computer vision › Vision and language
multimodal evaluation
0.812024
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs · NeurIPS 2024
Computer vision › Vision and language › vision-language model
multimodal large language model
0.812024
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs · NeurIPS 2024
Machine learning › Deep learning architectures and training
foundation model
0.212024
Fine-Tuning is Fine, if Calibrated · NeurIPS 2024
Machine learning › Transfer learning and domain adaptation
pre-trained models
0.212024
Fine-Tuning is Fine, if Calibrated · NeurIPS 2024
Machine learning › Deep learning architectures and training
transformer
0.212024
A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis · ICLR 2024

Methods — techniques the papers use, named apart from their topics

visual prompt tuning · 0.9multi-head attention · 0.9discriminative region localization · 0.9class-specific prompt learning · 0.9class activation mapping · 0.9multimodal large language model · 0.8empirical study · 0.8cross-attention · 0.8class-specific queries · 0.8CLIP · 0.8
YearPublicationVenuePosition
2025 Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained Analysis
abstract
We present a simple approach to make pre-trained Vision Transformers (ViTs) interpretable for fine-grained analysis, aiming to identify and localize the traits that distinguish visually similar categories, such as bird species. Pretrained ViTs, such as DINO, have demonstrated remarkable capabilities in extracting localized, discriminative features. However, saliency maps like Grad-CAM often fail to identify these traits, producing blurred, coarse heatmaps that highlight entire objects instead. We propose a novel approach, Prompt Class Attention Map (Prompt-CAM), to address this limitation. Prompt-CAM learns class-specific prompts for a pre-trained ViT and uses the corresponding outputs for classification. To correctly classify an image, the true-class prompt must attend to unique image patches not present in other classes’ images (i.e., traits). As a result, the true class’s multi-head attention maps reveal traits and their locations. Implementation-wise, Prompt-CAM is almost a "free lunch," requiring only a modification to the prediction head of Visual Prompt Tuning (VPT). This makes Prompt-CAM easy to train and apply, in stark contrast to other interpretable methods that require designing specific models and training processes. Extensive empirical studies on a dozen datasets from various domains (e.g., birds, fishes, insects, fungi, flowers, food, and cars) validate the superior interpretation capability of Prompt-CAM. The source code and demo are available at https://github.com/Imageomics/Prompt_CAM.
Arpita Chowdhury, Dipanjyoti Paul, Zheda Mai, Jianyang Gu, Kazi Sajeed Mehrab, Elizabeth G. Campolongo, Daniel I. Rubenstein, Charles V. Stewart, Anuj Karpatne, Tanya Y. Berger-Wolf, Yu Su 0001, Wei-Lun Chao
CVPR1
2025 Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation
abstract
Class activation map (CAM) has been widely used to highlight image regions that contribute to class predictions. Despite its simplicity and computational efficiency, CAM often struggles to identify discriminative regions that distinguish visually similar fine-grained classes. Prior efforts address this limitation by introducing more sophisticated explanation processes, but at the cost of extra complexity. In this paper, we propose Finer-CAM, a method that retains CAM’s efficiency while achieving precise localization of discriminative regions. Our key insight is that the deficiency of CAM lies not in "how" it explains, but in "what" it explains. Specifically, previous methods attempt to identify all cues contributing to the target class’s logit value, which inadvertently also activates regions predictive of visually similar classes. By explicitly comparing the target class with similar classes and spotting their differences, Finer-CAM suppresses features shared with other classes and emphasizes the unique, discriminative details of the target class. Finer-CAM is easy to implement, compatible with various CAM methods, and can be extended to multi-modal models for accurate localization of specific concepts. Additionally, Finer-CAM allows adjustable comparison strength, enabling users to selectively highlight coarse object contours or fine discriminative details. Quantitatively, we show that masking out the top 5% of activated pixels by Finer-CAM results in a larger relative confidence drop compared to baselines. The source code and demo are available at https://github.com/Imageomics/Finer-CAM.
Jianyang Gu, Arpita Chowdhury, Zheda Mai, David Carlyn, Tanya Y. Berger-Wolf, Yu Su 0001, Wei-Lun Chao
CVPR3
2024 A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis
abstract
We present a novel usage of Transformers to make image classification interpretable. Unlike mainstream classifiers that wait until the last fully connected layer to incorporate class information to make predictions, we investigate a proactive approach, asking each class to search for itself in an image. We realize this idea via a Transformer encoder-decoder inspired by DEtection TRansformer (DETR). We learn ''class-specific'' queries (one for each class) as input to the decoder, enabling each class to localize its patterns in an image via cross-attention. We name our approach INterpretable TRansformer (INTR), which is fairly easy to implement and exhibits several compelling properties. We show that INTR intrinsically encourages each class to attend distinctively; the cross-attention weights thus provide a faithful interpretation of the prediction. Interestingly, via ''multi-head'' cross-attention, INTR could identify different ''attributes'' of a class, making it particularly suitable for fine-grained classification and analysis, which we demonstrate on eight datasets. Our code and pre-trained models are publicly accessible at the Imageomics Institute GitHub site: https://github.com/Imageomics/INTR.
Dipanjyoti Paul, Arpita Chowdhury, Xinqi Xiong, Feng-Ju Chang, David Carlyn, Samuel Stevens 0001, Kaiya Provost, Anuj Karpatne, Bryan Carstens, Daniel I. Rubenstein, Charles V. Stewart, Tanya Y. Berger-Wolf, Yu Su 0001, Wei-Lun Chao
ICLR2
2024 MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
abstract
The ability to compare objects, scenes, or situations is crucial for effective decision-making and problem-solving in everyday life. For instance, comparing the freshness of apples enables better choices during grocery shopping, while comparing sofa designs helps optimize the aesthetics of our living space. Despite its significance, the comparative capability is largely unexplored in artificial general intelligence (AGI). In this paper, we introduce MLLM-CompBench, a benchmark designed to evaluate the comparative reasoning capability of multimodal large language models (MLLMs). MLLM-CompBench mines and pairs images through visually oriented questions covering eight dimensions of relative comparison: visual attribute, existence, state, emotion, temporality, spatiality, quantity, and quality. We curate a collection of around 40K image pairs using metadata from diverse vision datasets and CLIP similarity scores. These image pairs span a broad array of visual domains, including animals, fashion, sports, and both outdoor and indoor scenes. The questions are carefully crafted to discern relative characteristics between two images and are labeled by human annotators for accuracy and relevance. We use MLLM-CompBench to evaluate recent MLLMs, including GPT-4V(ision), Gemini-Pro, and LLaVA-1.6. Our results reveal notable shortcomings in their comparative abilities. We believe MLLM-CompBench not only sheds light on these limitations but also establishes a solid foundation for future enhancements in the comparative capability of MLLMs.
Jihyung Kil, Zheda Mai, Justin Lee, Arpita Chowdhury, Kerrie Cheng, Lemeng Wang, Wei-Lun Chao
NeurIPS4
2024 Fine-Tuning is Fine, if Calibrated
abstract
Fine-tuning is arguably the most straightforward way to tailor a pre-trained model (e.g., a foundation model) to downstream applications, but it also comes with the risk of losing valuable knowledge the model had learned in pre-training. For example, fine-tuning a pre-trained classifier capable of recognizing a large number of classes to master a subset of classes at hand is shown to drastically degrade the model's accuracy in the other classes it had previously learned. As such, it is hard to further use the fine-tuned model when it encounters classes beyond the fine-tuning data. In this paper, we systematically dissect the issue, aiming to answer the fundamental question, "What has been damaged in the fine-tuned model?" To our surprise, we find that the fine-tuned model neither forgets the relationship among the other classes nor degrades the features to recognize these classes. Instead, the fine-tuned model often produces more discriminative features for these other classes, even if they were missing during fine-tuning! What really hurts the accuracy is the discrepant logit scales between the fine-tuning classes and the other classes, implying that a simple post-processing calibration would bring back the pre-trained model's capability and at the same time unveil the feature improvement over all classes. We conduct an extensive empirical study to demonstrate the robustness of our findings and provide preliminary explanations underlying them, suggesting new directions for future theoretical analysis.
Zheda Mai, Arpita Chowdhury, Ping Zhang 0016, Cheng-Hao Tu 0001, Hong-You Chen, Vardaan Pahuja, Tanya Y. Berger-Wolf, Song Gao 0001, Charles V. Stewart, Yu Su 0001, Wei-Lun Chao
NeurIPS2