Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Aditya Chinchure

dblp:331/8100 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-7814-8238ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Knowledge representation and reasoning · 34% Vision and language · 24% Video understanding and tracking · 11%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
vision-language model
1.622025
Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events · CVPR 2025
From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models · EMNLP 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning
abductive reasoning
0.912025
Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events · CVPR 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.912025
Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events · CVPR 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › nonmonotonic reasoning
defeasible reasoning
0.912025
Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events · CVPR 2025
Computer vision › Video understanding and tracking › deep video understanding
video reasoning
0.912025
Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events · CVPR 2025
Machine learning › Trustworthy machine learning › fairness › fairness in generative models
bias in generative models
0.812024
TIBET: Identifying and Evaluating Biases in Text-to-Image Generative Models · ECCV (79) 2024
Natural language and speech › Language models and text generation
cross-cultural understanding
0.812024
From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models · EMNLP 2024
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.812024
TIBET: Identifying and Evaluating Biases in Text-to-Image Generative Models · ECCV (79) 2024
Computer vision › Vision and language
visual grounding
0.212024
From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

benchmark evaluation · 1.5multiple-choice evaluation · 0.9benchmark construction · 0.9
YearPublicationVenuePosition
2025 Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events
abstract
The commonsense reasoning capabilities of vision-language models (VLMs), especially in abductive reasoning and defeasible reasoning, remain poorly understood. Most benchmarks focus on typical visual scenarios [1], [23], [42], making it difficult to discern whether model performance stems from keen perception and reasoning skills, or reliance on pure statistical recall. We argue that by focusing on atypical events in videos, clearer insights can be gained on the core capabilities of VLMs. Explaining and understanding such out-of-distribution events requires models to extend beyond basic pattern recognition and regurgitation of their prior knowledge. To this end, we introduce Black-SwanSuite, a benchmark for evaluating VLMs’ ability to reason about unexpected events through abductive and defeasible tasks. Our tasks artificially limit the amount of visual information provided to models while questioning them about hidden unexpected events, or provide new visual information that could change an existing hypothesis about the event. We curate a comprehensive benchmark suite comprising over 3,800 MCQ, 4,900 generative and 6,700 yes/no questions, spanning 1,655 videos. After extensively evaluating various state-of-the-art VLMs, including GPT-4o and Gemini 1.5 Pro, as well as open-source VLMs such as LLaVA-Video, we find significant performance gaps of up to 32% from humans on these tasks. Our findings reveal key limitations in current VLMs, emphasizing the need for enhanced model architectures and training strategies. Our data and leaderboard is available at https://blackswan.cs.ubc.ca.
Aditya Chinchure, Sahithya Ravi, Raymond T. Ng, Vered Shwartz, Boyang Li 0001, Leonid Sigal
CVPR1
2024 TIBET: Identifying and Evaluating Biases in Text-to-Image Generative Models
Aditya Chinchure, Pushkar Shukla, Gaurav Bhatt, Kiri Salij, Kartik Hosanagar, Leonid Sigal, Matthew Turk 0001
ECCV (79)1
2024 From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models
abstract
Despite recent advancements in visionlanguage models, their performance remains suboptimal on images from non-western cultures, due to underrepresentation in training datasets.Various benchmarks have been proposed to test models' cultural inclusivity, but they have limited coverage of cultures and do not adequately assess cultural diversity across universal as well as culture-specific local concepts.To address these limitations, we introduce the GLOBALRG benchmark, comprising two challenging tasks: retrieval across universals and cultural visual grounding.The former task entails retrieving culturally-diverse images for universal concepts from 50 countries, while the latter aims at grounding culture-specific concepts within images from 15 countries.Our evaluation across a wide range of models reveals that the performance varies significantly across cultures -underscoring the necessity for enhancing multicultural understanding in vision-language models.Our
Mehar Bhatia, Sahithya Ravi, Aditya Chinchure, Eunjeong Hwang, Vered Shwartz
EMNLP3
2023 VLC-BERT: Visual Question Answering with Contextualized Commonsense Knowledge
abstract
There has been a growing interest in solving Visual Question Answering (VQA) tasks that require the model to reason beyond the content present in the image. In this work, we focus on questions that require commonsense reasoning. In contrast to previous methods which inject knowledge from static knowledge bases, we investigate the incorporation of contextualized knowledge using Commonsense Transformer (COMET), an existing knowledge model trained on human-curated knowledge bases. We propose a method to generate, select, and encode external commonsense knowledge alongside visual and textual cues in a new pre-trained Vision-Language-Commonsense transformer model, VLC-BERT. Through our evaluation on the knowledge-intensive OK-VQA and A-OKVQA datasets, we show that VLC-BERT is capable of outperforming existing models that utilize static knowledge bases. Furthermore, through a detailed analysis, we explain which questions benefit, and which don’t, from contextualized commonsense knowledge from COMET. Code: https://github.com/aditya10/VLC-BERT
Sahithya Ravi, Aditya Chinchure, Leonid Sigal, Renjie Liao 0001, Vered Shwartz
WACV2