Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Amita Kamath

dblp:267/9823 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0002-5005-415XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Vision and language · 36% Efficient and distributed learning · 21% Language models and text generation · 6%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
compositionality
1.422024
The Hard Positive Truth About Vision-Language Compositionality · ECCV (14) 2024
Text encoders bottleneck compositionality in contrastive vision-language models · EMNLP 2023
Machine learning › Efficient and distributed learning
model compression
0.812024
Matryoshka Query Transformer for Large Vision-Language Models · NeurIPS 2024
Computer vision › Vision and language › vision-language model
multimodal large language model
0.812024
Matryoshka Query Transformer for Large Vision-Language Models · NeurIPS 2024
Machine learning › Efficient and distributed learning › model compression › token compression
visual token compression
0.812024
Matryoshka Query Transformer for Large Vision-Language Models · NeurIPS 2024
Machine learning › Efficient and distributed learning › model compression › token compression
visual token reduction
0.812024
Matryoshka Query Transformer for Large Vision-Language Models · NeurIPS 2024
Computer vision › Vision and language › vision-language model
contrastive vision-language model
0.712023
Text encoders bottleneck compositionality in contrastive vision-language models · EMNLP 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning
spatial reasoning
0.712023
What's "up" with vision-language models? Investigating their struggle with spatial reasoning · EMNLP 2023
Computer vision › 3D vision › 3d scene understanding
spatial relation understanding
0.712023
What's "up" with vision-language models? Investigating their struggle with spatial reasoning · EMNLP 2023
Natural language and speech › Language models and text generation
text representation
0.712023
Text encoders bottleneck compositionality in contrastive vision-language models · EMNLP 2023
Natural language and speech › Information extraction and text analysis
concept expansion
0.612022
Webly Supervised Concept Expansion for General Purpose Vision Models · ECCV (36) 2022
Computer vision › Image recognition and object detection
object detection
0.612022
Towards General Purpose Vision Systems: An End-to-End Task-Agnostic Vision-Language Architecture · CVPR 2022
Computer vision › Vision and language
visual question answering
0.612022
Towards General Purpose Vision Systems: An End-to-End Task-Agnostic Vision-Language Architecture · CVPR 2022
Machine learning › Learning paradigms › weakly supervised learning
webly supervised learning
0.612022
Webly Supervised Concept Expansion for General Purpose Vision Models · ECCV (36) 2022
Machine learning › Transfer learning and domain adaptation
domain shift
0.412020
Selective Question Answering under Domain Shift · ACL 2020
Machine learning › Trustworthy machine learning › uncertainty estimation
selective classification
0.412020
Selective Question Answering under Domain Shift · ACL 2020
Machine learning › Trustworthy machine learning
calibration
0.112020
Selective Question Answering under Domain Shift · ACL 2020

Methods — techniques the papers use, named apart from their topics

query transformer · 0.8matryoshka representation learning · 0.8vision-language pretraining · 0.7recovery probe · 0.7fine-tuning · 0.7contrastive learning · 0.7calibrator training · 0.4
YearPublicationVenuePosition
2026 Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
abstract
Abstract The lack of reasoning capabilities in Vision-Language Models (VLMs) has remained at the forefront of research discourse. We posit that this behavior stems from a reporting bias in their training data. That is, how people communicate about visual content by default omits tacit information needed to supervise some types of reasoning; e.g., “at the game today!” is a more likely caption than “a photo of 37 people standing behind a field”. We investigate the data underlying the popular VLMs OpenCLIP, LLaVA-1.5 and Molmo through the lens of theories from pragmatics, and find that reporting bias results in insufficient representation of four reasoning skills (spatial, temporal, negation, and counting), despite the corpora being of web-scale, and/or synthetically generated. With a set of curated benchmarks, we demonstrate that: (i) VLMs perform poorly on the aforementioned types of reasoning suppressed in the training data by reporting bias; (ii) contrary to popular belief, scaling data size, model size, and to multiple languages does not result in emergence of these skills by default; but, promisingly, (iii) incorporating annotations specifically collected to obtain tacit information is effective. Our findings highlight the need for more intentional training data curation methods, rather than counting on scale for emergence of reasoning capabilities.
Amita Kamath, Jack Hessel, Khyathi Raghavi Chandu, Jena D. Hwang, Kai-Wei Chang 0001, Ranjay Krishna
Trans. Assoc. Comput. Linguistics1
2024 The Hard Positive Truth About Vision-Language Compositionality
Amita Kamath, Cheng-Yu Hsieh, Kai-Wei Chang 0001, Ranjay Krishna
ECCV (14)1
2024 Matryoshka Query Transformer for Large Vision-Language Models
abstract
Large Vision-Language Models (LVLMs) typically encode an image into a fixed number of visual tokens (e.g., 576) and process these tokens with a language model. Despite their strong performance, LVLMs face challenges in adapting to varying computational constraints. This raises the question: can we achieve flexibility in the number of visual tokens to suit different tasks and computational resources? We answer this with an emphatic yes. Inspired by Matryoshka Representation Learning, we introduce the Matryoshka Query Transformer (MQT), capable of encoding an image into $m$ visual tokens during inference, where $m$ can be any number up to a predefined maximum. This is achieved by employing a query transformer with $M$ latent query tokens to compress the visual embeddings. During each training step, we randomly select $m \leq M$ latent query tokens and train the model using only these first $m$ tokens, discarding the rest. Combining MQT with LLaVA, we train a single model once, and flexibly and drastically reduce the number of inference-time visual tokens while maintaining similar or better performance compared to training independent models for each number of tokens. Our model, MQT-LLaVA, matches LLaVA-1.5 performance across 11 benchmarks using a maximum of 256 tokens instead of LLaVA’s fixed 576. Reducing to 16 tokens (8x less TFLOPs) only sacrifices the performance by 2.4 points on MMBench. On certain tasks such as ScienceQA and MMMU, we can even go down to only 2 visual tokens with performance drops of just 3\% and 6\% each. Our exploration of the trade-off between the accuracy and computational cost brought about by the number of visual tokens facilitates future research to achieve the best of both worlds.
Wenbo Hu 0006, Zi-Yi Dou, Liunian Harold Li, Amita Kamath, Nanyun Peng 0001, Kai-Wei Chang 0001
NeurIPS4
2023 Text encoders bottleneck compositionality in contrastive vision-language models
abstract
Performant vision-language (VL) models like CLIP represent captions using a single vector.How much information about language is lost in this bottleneck?We first curate CompPrompts, a set of increasingly compositional image captions that VL models should be able to capture (e.g., single object, to ob-ject+property, to multiple interacting objects).Then, we train text-only recovery probes that aim to reconstruct captions from single-vector text representations produced by several VL models.This approach does not require images, allowing us to test on a broader range of scenes compared to prior work.We find that: 1) CLIP's text encoder falls short on more compositional inputs, including object relationships, attribute-object association, counting, and negations; 2) some text encoders work significantly better than others; and 3) text-only recovery performance predicts multimodal matching performance on ControlledImCaps: a new evaluation benchmark we collect and release consisting of fine-grained compositional images and captions.Specifically, our results suggest textonly recoverability is a necessary (but not sufficient) condition for modeling compositional factors in contrastive VL models.We release our datasets and code.
Amita Kamath, Jack Hessel, Kai-Wei Chang 0001
EMNLP1
2023 What's "up" with vision-language models? Investigating their struggle with spatial reasoning
abstract
Recent vision-language (VL) models are powerful, but can they reliably distinguish "right" from "left"?We curate three new corpora to quantify model comprehension of such basic spatial relations.These tests isolate spatial reasoning more precisely than existing datasets like VQAv2, e.g., our What'sUp benchmark contains sets of photographs varying only the spatial relations of objects, keeping their identity fixed (see Figure 1: models must comprehend not only the usual case of a dog under a table, but also, the same dog on top of the same table).We evaluate 18 VL models, finding that all perform poorly, e.g., BLIP finetuned on VQAv2, which nears human parity on VQAv2, achieves 56% accuracy on our benchmarks vs. humans at 99%.We conclude by studying causes of this surprising behavior, finding: 1) that popular vision-language pretraining corpora like LAION-2B contain little reliable data for learning spatial relationships; and 2) that basic modeling interventions like up-weighting preposition-containing instances or fine-tuning on our corpora are not sufficient to address the challenges our benchmarks pose.We are hopeful that these corpora will facilitate further research, and we release our data and code at https://github.com/amitakamath/ whatsup_vlms.
Amita Kamath, Jack Hessel, Kai-Wei Chang 0001
EMNLP1
2022 Towards General Purpose Vision Systems: An End-to-End Task-Agnostic Vision-Language Architecture
abstract
Computer vision systems today are primarily N-purpose systems, designed and trained for a predefined set of tasks. Adapting such systems to new tasks is challenging and often requires nontrivial modifications to the network architecture (e.g. adding new output heads) or training process (e.g. adding new losses). To reduce the time and expertise required to develop new applications, we would like to create general purpose vision systems that can learn and perform a range of tasks without any modification to the architecture or learning process. In this paper, we propose GPV-1, a task-agnostic vision-language architecture that can learn and perform tasks that involve receiving an image and producing text and/or bounding boxes, including classification, localization, visual question answering, captioning, and more. We also propose evaluations of generality of architecture, skill-concept11For this work, we define concepts, skills and tasks as follows: Concepts - nouns (e.g. car, person, dog), Skills - operations that we wish to perform on the given inputs (e.g. classification, object detection, image captioning), Tasks - predefined combinations of a set of skills performed on a set of concepts (e.g. ImageNet classification task involves the skill of image classification across 1000 concepts). transfer, and learning efficiency that may informfuture work on general purpose vision. Our experiments indicate GPV-1 is effective at multiple tasks, reuses some concept knowledge across tasks, can perform the Referring Expressions task zero-shot, and further improves upon the zero-shot performance using a few training samples.
Tanmay Gupta, Amita Kamath, Aniruddha Kembhavi, Derek Hoiem
CVPR2
2022 Webly Supervised Concept Expansion for General Purpose Vision Models
Amita Kamath, Tanmay Gupta, Eric Kolve, Derek Hoiem, Aniruddha Kembhavi
ECCV (36)1
2020 Selective Question Answering under Domain Shift
abstract
To avoid giving wrong answers, question answering (QA) models need to know when to abstain from answering.Moreover, users often ask questions that diverge from the model's training data, making errors more likely and thus abstention more critical.In this work, we propose the setting of selective question answering under domain shift, in which a QA model is tested on a mixture of in-domain and out-of-domain data, and must answer (i.e., not abstain on) as many questions as possible while maintaining high accuracy.Abstention policies based solely on the model's softmax probabilities fare poorly, since models are overconfident on out-of-domain inputs.Instead, we train a calibrator to identify inputs on which the QA model errs, and abstain when it predicts an error is likely.Crucially, the calibrator benefits from observing the model's behavior on out-of-domain data, even if from a different domain than the test data.We combine this method with a SQuADtrained QA model and evaluate on mixtures of SQuAD and five other QA datasets.Our method answers 56% of questions while maintaining 80% accuracy; in contrast, directly using the model's probabilities only answers 48% at 80% accuracy.
Amita Kamath, Robin Jia, Percy Liang
ACL1