EDBT 2026 Demo / reviewers in the wild / expert
Adriana Kovashka
dblp:51/8652
· DBLP profile ↗
60ranked-venue papers
7as first author
33since 2021 · last 2026
0000-0003-1901-9660ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 7 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 42 · 5 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generalizing Sports Feedback Generation by Watching Competitions and Reading Books: A Rock Climbing Case StudyabstractWhile there is rapid progress in video-LLMs with advanced reasoning capabilities, prior work shows that these models struggle on the challenging task of sports feedback generation and require expensive and difficult-to-collect finetuning feedback data for each sport. This limitation is evident from the poor generalization to sports unseen during finetuning. Furthermore, traditional text generation evaluation metrics (e.g., BLEU-4, METEOR, ROUGE-L, BERTScore), originally developed for machine translation and summarization, fail to capture the unique aspects of sports feedback quality. To address the first problem, using rock climbing as our case study, we propose using auxiliary freelyavailable web data from the target domain, such as competition videos and coaching manuals, in addition to existing sports feedback from a disjoint, source domain to improve sports feedback generation performance on the target domain. To improve evaluation, we propose two evaluation metrics: (1) specificity and (2) actionability. Together, our approach enables more meaningful and practical generation of sports feedback under limited annotations. Arushi Rai, Adriana Kovashka |
WACV | 2 |
| 2025 | Multi-party Lexical Alignment in Collaborative Learning with a Teachable Robot
Yuya Asano, Diane J. Litman, Paras Sharma, Daniel Fritsch, Quentin King-Shepard, Timothy Nokes-Malach, Adriana Kovashka, Erin Walker |
AIED (6) | 7 |
| 2025 | Beyond Static Measures: Temporal Analysis of Lexical Alignment in Human-Human Learning With a Teachable Robot
Paras Sharma, Daniel Fritsch, Yuya Asano, Quentin King-Shepard, Tyree Langley, Tristan Maidment, Diane J. Litman, Timothy Nokes-Malach, Adriana Kovashka, Nikki G. Lobczowski, Erin Walker |
AIED (4) | 9 |
| 2025 | Leveraging Large Models to Evaluate Novel Content: A Case Study on Advertisement CreativityabstractEvaluating creativity is challenging, even for humans, not only because of its subjectivity but also because it involves complex cognitive processes. Inspired by work in marketing, we attempt to break down visual advertisement creativity into atypicality and originality. With fine-grained human annotations on these dimensions, we propose a suite of tasks specifically for such a subjective problem. We also evaluate the alignment between state-of-the-art (SoTA) vision language models (VLMs) and humans on our proposed benchmark, demonstrating both the promises and challenges of using VLMs for automatic creativity assessment. Zhaoyi Hou, Adriana Kovashka, Xiang Li 0069 |
EMNLP | 2 |
| 2025 | Probing Logical Reasoning of MLLMs in Scientific DiagramsabstractWe examine how multimodal large language models (MLLMs) perform logical inference grounded in visual information.We first construct a dataset of food web/chain images, along with questions that follow seven structured templates with progressively more complex reasoning involved.We show that complex reasoning about entities in the images remains challenging (even with elaborate prompts) and that visual information is underutilized. Adriana Kovashka |
EMNLP | 2 |
| 2025 | Integrating Audio Narrations to Strengthen Domain Generalization in Multimodal First-Person Action RecognitionabstractFirst-person activity recognition is rapidly growing due to the widespread use of wearable cameras but faces challenges from domain shifts across different environments, such as varying objects or background scenes. We propose a multi-modal framework that improves domain generalization by integrating motion, audio, and appearance features. Key contributions include analyzing the resilience of audio and motion features to domain shifts, using audio narrations for enhanced audio-text alignment, and applying consistency ratings between audio and visual narrations to optimize the impact of audio in recognition during training. Our approach achieves state-of-the-art performance on the ARGO1M dataset, effectively generalizing across unseen scenarios and locations. Cagri Gungor, Adriana Kovashka |
ICASSP | 2 |
| 2025 | Cap: Evaluation of Persuasive and Creative Image GenerationabstractWe address the task of advertisement image generation and introduce three evaluation metrics to assess Creativity, prompt Alignment, and Persuasiveness (CAP) in generated advertisement images. Despite recent advancements in Text-to-Image (T2I) generation and their performance in generating high-quality images for explicit descriptions, evaluating these models remains challenging. Existing evaluation methods focus largely on assessing alignment with explicit, detailed descriptions, but evaluating alignment with visually implicit prompts remains an open problem. Additionally, creativity and persuasiveness are essential qualities that enhance the effectiveness of advertisement images, yet are seldom measured. To address this, we propose three novel metrics for evaluating the creativity, alignment, and persuasiveness of generated images. Our findings reveal that current T2I models struggle with creativity, persuasiveness, and alignment when the input text is implicit messages. We further introduce a simple yet effective approach to enhance T2I models' capabilities in producing images that are better aligned, more creative, and more persuasive. Aysan Aghazadeh, Adriana Kovashka |
ICCV | 2 |
| 2025 | Role Bias in Diffusion Models: Diagnosing and Mitigating through Intermediate DecompositionabstractText-to-image (T2I) diffusion models exhibit impressive photorealistic image generation capabilities, yet they struggle in compositional image generation. In this work, we introduce RoleBench, a benchmark focused on evaluating compositional generalization in action-based relations (e.g., "mouse chasing cat"). We show that state-of-the-art T2I models and compositional generation methods consistently default to frequent reversed relations (i.e., "cat chasing mouse"), a phenomenon we call role collapse. Related works attribute this to the model’s architectural limitation or underrepresentation in the data. Our key insight reveals that while models fail on rare compositions when their inversions are common, they can successfully generate similar intermediate compositions (e.g., "mouse chasing boy"), suggesting that this limitation is also due to the presence of frequent counterparts rather than just the absence of rare compositions. Motivated by this, we hypothesize that directional decomposition can gradually mitigate role collapse. We test this via ReBind, a lightweight framework that teaches role bindings using carefully selected active/passive intermediate compositions. Experiments suggest that intermediate compositions through simple fine-tuning can significantly reduce role collapse, with humans preferring ReBind more than 78% compared to state-of-the-art methods. Our findings highlight the role of distributional asymmetries in compositional failures and offer a simple, effective path for improving generalization. Sina Malakouti, Adriana Kovashka |
NeurIPS | 2 |
| 2025 | Benchmarking VLMs' Reasoning About Persuasive Atypical ImagesabstractVision-language models (VLMs) have shown strong zero-shot generalization across various tasks, especially when integrated with large language models (LLMs). However, their ability to comprehend rhetorical and persuasive visual media, such as advertisements, remains understudied. Ads often employ atypical imagery, using surprising object juxtapositions to convey shared properties. For example, Fig. 1(e) shows a beer with a feather-like texture. This requires advanced reasoning to deduce that this atypical representation signifies the beer's lightness. We introduce three novel tasks, Multi-label Atypicality Classification, Atypicality Statement Retrieval, and Atypical Object Recognition, to benchmark VLMs' understanding of atypicality in persuasive images. We evaluate how well VLMs use atypicality to infer an ad's message and test their reasoning abilities by employing semantically challenging negatives. Finally, we pioneer atypicality-aware verbalization by extracting comprehensive image descriptions sensitive to atypical elements. Findings reveal that: (1) VLMs lack advanced reasoning capabilities compared to LLMs; (2) simple, effective strategies can extract atypicality-aware information, leading to comprehensive image verbalization; (3) atypicality aids persuasive ad understanding. Code and data is available at aysanaghazadeh.github.io/PersuasiveAdVLMBenchmark/ Sina Malakouti, Aysan Aghazadeh, Ashmit Khandelwal, Adriana Kovashka |
WACV | 4 |
| 2025 | Rubric-Constrained Figure Skating ScoringabstractFigure skating automatic scoring is the task of estimating the competition score of a performance video. The tech-nical element score (TES) aggregates the technical quality (grade of execution) and difficulty (base value) scores for each element. Most prior work, adapted from short-term action quality assessment, entangle difficulty and quality, and compute TES for the entire video, reducing inter-pretability for athletes. This is mainly due to a lack of el-ement segmentation and difficulty annotations in existing datasets. Motivated by increasing interpretability, we propose a novel method that implicitly segments a video to produce element-level representations and uses adherence with a natural language rubric to score each element, without needing additional annotations. We compute element-level representations using learnable element queries in a transformer and propose implicit segmentation regularization to encourage element queries to attend to elements rather than background transitions between elements (most of video). Additionally, we use the element list (sequence of elements) to isolate difficulty, just like judges who receive the rou-tine list in advance, so we can focus on the more critical problem of how well elements are done. These components significantly improve interpretability, scoring precision, and ranking capability. Code is released at htt ps: //arushirail.github.io/rcs-project. Arushi Rai, Adriana Kovashka |
WACV | 2 |
| 2024 | Incorporating Geo-Diverse Knowledge into Prompting for Increased Geographical Robustness in Object RecognitionabstractExisting object recognition models have been shown to lack robustness in diverse geographical scenarios due to domain shifts in design and context. Class representations need to be adapted to more accurately reflect an object concept under these shifts. In the absence of training data from target geographies, we hypothesize that geographically diverse descriptive knowledge of categories can enhance robustness. For this purpose, we explore the feasibility of probing a large language model for geography-based object knowledge, and we examine the effects of integrating knowledge into zero-shot and learnable soft prompting with CLIP. Within this exploration, we propose geog-raphy knowledge regularization to ensure that soft prompts trained on a source set of geographies generalize to an un-seen target set. Accuracy gains over prompting baselines on DollarStreet while training only on Europe data are up to +2.8/1.2/1.6 on target data from Africa/Asia/Americas, and +4.6 overall on the hardest classes. Competitive performance is shown vs. few-shot target training, and analysis is provided to direct future study of geographical robustness. Kyle Buettner, Sina Malakouti, Xiang Li 0069, Adriana Kovashka |
CVPR | 4 |
| 2024 | VEIL: Vetting Extracted Image Labels from In-the-Wild Captions for Weakly-Supervised Object DetectionabstractThe use of large-scale vision-language datasets is limited for object detection due to the negative impact of label noise on localization.Prior methods have shown how such large-scale datasets can be used for pretraining, which can provide initial signal for localization, but is insufficient without clean bounding-box data for at least some categories.We propose a technique to "vet" labels extracted from noisy captions, and use them for weakly-supervised object detection (WSOD), without any bounding boxes.We analyze and annotate the types of label noise in captions in our Caption Label Noise dataset, and train a classifier that predicts if an extracted label is actually present in the image or not.Our classifier generalizes across dataset boundaries and across categories.We compare the classifier to nine baselines on five datasets, and demonstrate that it can improve WSOD without label vetting by 30% (31.2 to 40.5 mAP when evaluated on PASCAL VOC). Arushi Rai, Adriana Kovashka |
EACL (1) | 2 |
| 2024 | Quantifying the Gaps Between Translation and Native Perception in Training for Multimodal, Multilingual RetrievalabstractThere is a scarcity of multilingual visionlanguage models that properly account for the perceptual differences that are reflected in image captions across languages and cultures.In this work, through a multimodal, multilingual retrieval case study, we quantify the existing lack of model flexibility.We empirically show performance gaps between training on captions that come from native German perception and captions that have been either machinetranslated or human-translated from English into German.To address these gaps, we further propose and evaluate caption augmentation strategies.While we achieve mean recall improvements (+1.3), gaps still remain, indicating an open area of future work for the community. Kyle Buettner, Adriana Kovashka |
EMNLP | 2 |
| 2024 | Investigating the Role of Attribute Context in Vision-Language Models for Object Recognition and DetectionabstractVision-language alignment learned from image-caption pairs has been shown to benefit tasks like object recognition and detection. Methods are mostly evaluated in terms of how well object class names are learned, but captions also contain rich attribute context that should be considered when learning object alignment. It is unclear how methods use this context in learning, as well as whether models succeed when tasks require attribute and object understanding. To address this gap, we conduct extensive analysis of the role of attributes in vision-language models. We specifically measure model sensitivity to the presence and meaning of attribute context, gauging influence on object embeddings through unsupervised phrase grounding and classification via description methods. We further evaluate the utility of attribute context in training for open-vocabulary object detection, fine-grained text-region retrieval, and attribution tasks. Our results show that attribute context can be wasted when learning alignment for detection, attribute meaning is not adequately considered in embeddings, and describing classes by only their attributes is ineffective. A viable strategy that we find to increase benefits from attributes is contrastive training with adjective-based negative captions. Kyle Buettner, Adriana Kovashka |
WACV | 2 |
| 2024 | Boosting Weakly Supervised Object Detection using Fusion and Priors from Hallucinated DepthabstractDespite recent attention to depth for various tasks, it is still an unexplored modality for weakly-supervised object detection (WSOD). We propose an amplifier method for enhancing the performance of WSOD by integrating depth information. Our approach can be applied to different WSOD methods based on multiple-instance learning, without necessitating additional annotations or inducing large computational cost. Our proposed method employs monocular depth estimation to obtain hallucinated depth information, which is then incorporated into a Siamese WSOD network using contrastive loss and fusion. By analyzing the relationship between language context and depth, we calculate depth priors to identify the bounding box proposals that may contain an object of interest. These depth priors are then utilized to update the list of pseudo ground-truth boxes, or adjust the confidence of per-box predictions. We evaluate our proposed method on three datasets (COCO, PASCAL VOC, and Conceptual Captions) by implementing it on top of two state-of-the-art WSOD methods, and we demonstrate a substantial enhancement in performance. Cagri Gungor, Adriana Kovashka |
WACV | 2 |
| 2023 | Decoding Symbolism in Language ModelsabstractThis work explores the feasibility of eliciting knowledge from language models (LMs) to decode symbolism, recognizing something (e.g., roses) as a stand-in for another (e.g., love).We present our evaluative framework, Symbolism Analysis (SymbA), which compares LMs (e.g., RoBERTa, GPT-J) on different types of symbolism and analyzes the outcomes along multiple metrics.Our findings suggest that conventional symbols are more reliably elicited from LMs while situated symbols are more challenging.Results also reveal the negative impact of the bias in pre-trained corpora.We further demonstrate that a simple re-ranking strategy can mitigate the bias and significantly improve model performances to be on par with human performances in some cases. Meiqi Guo, Rebecca Hwa, Adriana Kovashka |
ACL (1) | 3 |
| 2023 | Semi-Supervised Domain Generalization for Object Detection via Language-Guided Feature Alignment
Sina Malakouti, Adriana Kovashka |
BMVC | 2 |
| 2023 | Hypernymization of named entity-rich captions for grounding-based multi-modal pretrainingabstractNamed entities are ubiquitous in text that naturally accompanies images, especially in domains such as news or Wikipedia articles. In previous work, named entities have been identified as a likely reason for low performance of image-text retrieval models pretrained on Wikipedia and evaluated on named entities-free benchmark datasets. Because they are rarely mentioned, named entities could be challenging to model. They also represent missed learning opportunities for self-supervised models: the link between named entity and object in the image may be missed by the model, but it would not be if the object were mentioned using a more common term. In this work, we investigate hypernymization as a way to deal with named entities for pretraining grounding-based multi-modal models and for fine-tuning on open-vocabulary detection. We propose two ways to perform hypernymization: (1) a “manual” pipeline relying on a comprehensive ontology of concepts, and (2) a “learned” approach where we train a language model to learn to perform hypernymization. We run experiments on data from Wikipedia and from The New York Times. We report improved pretraining performance on objects of interest following hypernymization, and we show the promise of hypernymization on open-vocabulary detection, specifically on classes not seen during training. Giacomo Nebbia, Adriana Kovashka |
ICMR | 2 |
| 2023 | Towards Shape-regularized Learning for Mitigating Texture Bias in CNNsabstractCNNs have emerged as powerful techniques for object recognition. However, the test performance of CNNs is contingent on the similarity to training distribution. Existing methods focus on data augmentation to address out-of-domain generalization. In contrast, we enforce a shape bias by encouraging our model to learn features that correlate with those learned from the shape of the object. We show that explicit shape cues enable CNNs to learn features that are robust to unseen image manipulations i.e. novel textures with the same semantic content. Our models are validated on Toys4K dataset which consists of 4179 3D objects and image pairs. To quantify texture bias, we synthesize dataset variants called Style (style-transfer with GANs), CueConflict (conflicting texture & semantics), and Scrambled datasets (obfuscating semantics by scrambling pixel blocks). Our experiments show that the benefits of using shape is not subject to specific shape representations like point clouds, rather the same benefits can be obtained from a simpler representation such as the distance transform. Harsh Sinha, Adriana Kovashka |
ICMR | 2 |
| 2023 | Complementary Cues from Audio Help Combat Noise in Weakly-Supervised Object DetectionabstractWe tackle the problem of learning object detectors in a noisy environment, which is one of the significant challenges for weakly-supervised learning. We use multimodal learning to help localize objects of interest, but unlike other methods, we treat audio as an auxiliary modality that assists to tackle noise in detection from visual regions. First, we use the audio-visual model to generate new "ground-truth" labels for the training set to remove noise between the visual features and noisy supervision. Second, we propose an "indirect path" between audio and class predictions, which combines the link between visual and audio regions, and the link between visual features and predictions. Third, we propose a sound-based "attention path" which uses the benefit of complementary audio cues to identify important visual regions. We use contrastive learning to perform region-based audio-visual instance discrimination, which serves as an intermediate task and benefits from the complementary cues from audio to boost object classification and detection performance. We show that our methods, which update noisy ground truth and provide indirect and attention paths, greatly boosting performance on the AudioSet and VGGSound datasets compared to single-modality predictions, even ones that use contrastive learning. Our method outperforms previous weakly-supervised detectors for the task of object detection by reaching the state-of-art on AudioSet, and our sound localization module performs better than several state-of-art methods on AudioSet and MUSIC. Cagri Gungor, Adriana Kovashka |
WACV | 2 |
| 2023 | How to Practice VQA on a Resource-limited Target DomainabstractVisual question answering (VQA) is an active research area at the intersection of computer vision and natural language understanding. One major obstacle that keeps VQA models that perform well on benchmarks from being as successful on real-world applications, is the lack of annotated Image–Question–Answer triplets in the task of interest. In this work, we focus on a previously overlooked perspective, which is the disparate effectiveness of transfer learning and domain adaptation methods depending on the amount of labeled/unlabeled data available. We systematically investigated the visual domain gaps and question-defined textual gaps, and compared different knowledge transfer strategies under unsupervised, self-supervised, semi-supervised and fully-supervised adaptation scenarios. We show that different methods have varied sensitivity and requirements for data amount in the target domain. We conclude by sharing the best practice from our exploration regarding transferring VQA models to resource-limited target domains. Rebecca Hwa, Adriana Kovashka |
WACV | 3 |
| 2023 | Learning to Overcome Noise in Weak Caption Supervision for Object DetectionabstractWe propose the first mechanism to train object detection models from weak supervision in the form of captions at the image level. Language-based supervision for detection is appealing and inexpensive: many blogs with images and descriptive text written by human users exist. However, there is significant noise in this supervision: captions do not mention all objects that are shown, and may mention extraneous concepts. We first propose a technique to determine which image-caption pairs provide suitable signal for supervision. We further propose several complementary mechanisms to extract image-level pseudo labels for training from the caption. Finally, we train an iterative weakly-supervised object detection model from these image-level pseudo labels. We use captions from four datasets (COCO, Flickr30K, MIRFlickr1M, and Conceptual Captions) whose level of noise varies. We evaluate our approach on two object detection datasets. Weighting the labels extracted from different captions provides a boost over treating all captions equally. Further, our primary proposed technique for inferring pseudo labels for training at the image level, outperforms alternative techniques under a wide variety of settings. Both techniques generalize to datasets beyond the one they were trained on. Mesut Erhan Unal, Keren Ye, Christopher Thomas 0004, Adriana Kovashka, Wei Li 0044, Danfeng Qin, Jesse Berent |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Building a Reinforcement Learning Environment from Limited Data to Optimize Teachable Robot Interventions
Tristan Maidment, Mingzhi Yu, Nikki G. Lobczowski, Adriana Kovashka, Erin Walker, Diane J. Litman, Timothy Nokes-Malach |
EDM | 4 |
| 2022 | Characterizing User Susceptibility to COVID-19 Misinformation on Twitter
Xian Teng, Yu-Ru Lin, Wen-Ting Chung, Adriana Kovashka |
ICWSM | 5 |
| 2022 | Comparison of Lexical Alignment with a Teachable Robot in Human-Robot and Human-Human-Robot InteractionsabstractYuya Asano, Diane Litman, Mingzhi Yu, Nikki Lobczowski, Timothy Nokes-Malach, Adriana Kovashka, Erin Walker. Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2022. Yuya Asano, Diane J. Litman, Mingzhi Yu, Nikki G. Lobczowski, Timothy Nokes-Malach, Adriana Kovashka, Erin Walker |
SIGDIAL | 6 |
| 2021 | A Case Study of the Shortcut Effects in Visual Commonsense ReasoningabstractVisual reasoning and question-answering have gathered attention in recent years. Many datasets and evaluation protocols have been proposed; some have been shown to contain bias that allows models to ``cheat'' without performing true, generalizable reasoning. A well-known bias is dependence on language priors (frequency of answers) resulting in the model not looking at the image. We discover a new type of bias in the Visual Commonsense Reasoning (VCR) dataset. In particular we show that most state-of-the-art models exploit co-occurring text between input (question) and output (answer options), and rely on only a few pieces of information in the candidate options, to make a decision. Unfortunately, relying on such superficial evidence causes models to be very fragile. To measure fragility, we propose two ways to modify the validation data, in which a few words in the answer choices are modified without significant changes in meaning. We find such insignificant changes cause models' performance to degrade significantly. To resolve the issue, we propose a curriculum-based masking approach, as a mechanism to perform more robust training. Our method improves the baseline by requiring it to pay attention to the answers as a whole, and is more effective than prior masking strategies. Keren Ye, Adriana Kovashka |
AAAI | 2 |
| 2021 | Linguistic Structures As Weak Supervision for Visual Scene Graph GenerationabstractPrior work in scene graph generation requires categorical supervision at the level of triplets—subjects and objects, and predicates that relate them, either with or without bounding box information. However, scene graph generation is a holistic task: thus holistic, contextual supervision should intuitively improve performance. In this work, we explore how linguistic structures in captions can benefit scene graph generation. Our method captures the information provided in captions about relations between individual triplets, and context for subjects and objects (e.g. visual properties are mentioned). Captions are a weaker type of supervision than triplets since the alignment between the exhaustive list of human-annotated subjects and objects in triplets, and the nouns in captions, is weak. However, given the large and diverse sources of multimodal data on the web (e.g. blog posts with images and captions), linguistic supervision is more scalable than crowdsourced triplets. We show extensive experimental comparisons against prior methods which leverage instance- and image-level supervision, and ablate our method to show the impact of leveraging phrasal and sequential context, and techniques to improve localization of subjects and objects. Keren Ye, Adriana Kovashka |
CVPR | 2 |
| 2021 | Domain-Robust VQA With Diverse Datasets and Methods but No Target LabelsabstractThe observation that computer vision methods overfit to dataset specifics has inspired diverse attempts to make object recognition models robust to domain shifts. However, similar work on domain-robust visual question answering methods is very limited. Domain adaptation for VQA differs from adaptation for object recognition due to additional complexity: VQA models handle multimodal inputs, methods contain multiple steps with diverse modules resulting in complex optimization, and answer spaces in different datasets are vastly different. To tackle these challenges, we first quantify domain shifts between popular VQA datasets, in both visual and textual space. To disentangle shifts between datasets arising from different modalities, we also construct synthetic shifts in the image and question domains separately. Second, we test the robustness of different families of VQA methods (classic two-stream, transformer, and neuro-symbolic methods) to these shifts. Third, we test the applicability of existing domain adaptation methods and devise a new one to bridge VQA domain gaps, adjusted to specific VQA models. To emulate the setting of real-world generalization, we focus on unsupervised domain adaptation and the open-ended classification task formulation. Tristan Maidment, Ahmad Diab, Adriana Kovashka, Rebecca Hwa |
CVPR | 4 |
| 2021 | Detecting Persuasive Atypicality by Modeling Contextual CompatibilityabstractWe propose a new approach to detect atypicality in persuasive imagery. Unlike atypicality which has been studied in prior work, persuasive atypicality has a particular purpose to convey meaning, and relies on understanding the common-sense spatial relations of objects. We propose a self-supervised attention-based technique which captures contextual compatibility, and models spatial relations in a precise manner. We further experiment with capturing common sense through the semantics of co-occurring object classes. We verify our approach on a dataset of atypicality in visual advertisements, as well as a second dataset capturing atypicality that has no persuasive intent. Meiqi Guo, Rebecca Hwa, Adriana Kovashka |
ICCV | 3 |
| 2021 | Breaking Shortcuts by Masking for Robust Visual ReasoningabstractVisual reasoning is a challenging but important task that is gaining momentum. Examples include reasoning about what will happen next in film, or interpreting what actions an image advertisement prompts. Both tasks are "puzzles" which invite the viewer to combine knowledge from prior experience, to find the answer. Intuitively, providing external knowledge to a model should be helpful, but it does not necessarily result in improved reasoning ability. An algorithm can learn to find answers to the prediction task yet not perform generalizable reasoning. In other words, models can leverage "shortcuts" between inputs and desired outputs, to bypass the need for reasoning. We develop a technique to effectively incorporate external knowledge, in a way that is both interpretable, and boosts the contribution of external knowledge for multiple complementary metrics. In particular, we mask evidence in the image and in retrieved external knowledge. We show this masking successfully focuses the method's attention on patterns that generalize. To properly understand how our method utilizes external knowledge, we propose a novel side evaluation task. We find that with our masking technique, the model can learn to select useful knowledge pieces to rely on. Keren Ye, Adriana Kovashka |
WACV | 3 |
| 2021 | Image retrieval with mixed initiative and multimodal feedback
Nils Murrugarra-Llerena, Adriana Kovashka |
Comput. Vis. Image Underst. | 2 |
| 2021 | Predicting Visual Political Bias Using Webly Supervised Data and an Auxiliary Task
Christopher Thomas 0004, Adriana Kovashka |
Int. J. Comput. Vis. | 2 |
| 2021 | Interpreting the Rhetoric of Visual AdvertisementsabstractVisual media have important persuasive power, but prior computer vision approaches have predominantly ignored the persuasive aspects of images. In this work, we propose a suite of data and techniques that enable progress on understanding the messages that visual advertisements convey. We make available a dataset of 64,832 image ads and 3,477 video ads, annotated with ten types of information: the topic and sentiment of the ad; whether it is funny, exciting, or effective; what action it prompts the viewer to do, and what arguments it provides for why this action should be taken; symbolic associations that the ad relies on; the metaphorical object transformations on which especially creative ads rely; and the climax in video ads. We develop methods that use multimodal cues, i.e., both visuals and slogans, for both the image and video domains. Our methods rely on finding poignant content spatially and temporally. We also examine the creative story construction in ads: for videos, we learn to predict when the climax occurs (if any), and how effective the story is; for images, we analyze how object transformations in ads metaphorically depict product properties. Keren Ye, Narges Honarvar Nazari, Zaeem Hussain, Adriana Kovashka |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2020 | SpotPatch: Parameter-Efficient Transfer Learning for Mobile Object Detection
Keren Ye, Adriana Kovashka, Mark Sandler 0002, Menglong Zhu, Andrew G. Howard, Marco Fornoni |
ACCV (6) | 2 |
| 2020 | Preserving Semantic Neighborhoods for Robust Cross-Modal Retrieval
Christopher Thomas 0004, Adriana Kovashka |
ECCV (18) | 2 |
| 2020 | Context for Object Detection via Lightweight Global and Mid-level RepresentationsabstractWe propose an approach for explicitly capturing context in object detection. We model visual and geometric relationships between object regions, but also model the global scene as a first-class participant. In contrast to prior approaches, both the context we rely on, as well as our proposed mechanism for belief propagation over regions, is lightweight. We also experiment with capturing similarities between regions at a semantic level, by modeling class co-occurrence and linguistic similarity between class names. We show that our approach significantly outperforms Faster R-CNN, and performs competitively with a much more costly approach that also models context. Mesut Erhan Unal, Adriana Kovashka |
ICPR | 2 |
| 2019 | Cross-Modality Personalization for RetrievalabstractExisting captioning and gaze prediction approaches do not consider the multiple facets of personality that affect how a viewer extracts meaning from an image. While there are methods that consider personalized captioning, they do not consider personalized perception across modalities, i.e. how a person's way of looking at an image (gaze) affects the way they describe it (captioning). In this work, we propose a model for modeling cross-modality personalized retrieval. In addition to modeling gaze and captions, we also explicitly model the personality of the users providing these samples. We incorporate constraints that encourage gaze and caption samples on the same image to be close in a learned space; we refer to this as content modeling. We also model style: we encourage samples provided by the same user to be close in a separate embedding space, regardless of the image on which they were provided. To leverage the complementary information that content and style constraints provide, we combine the embeddings from both networks. We show that our combined embeddings achieve better performance than existing approaches for cross-modal retrieval. Nils Murrugarra-Llerena, Adriana Kovashka |
CVPR | 2 |
| 2019 | Cap2Det: Learning to Amplify Weak Caption Supervision for Object DetectionabstractLearning to localize and name object instances is a fundamental problem in vision, but state-of-the-art approaches rely on expensive bounding box supervision. While weakly supervised detection (WSOD) methods relax the need for boxes to that of image-level annotations, even cheaper supervision is naturally available in the form of unstructured textual descriptions that users may freely provide when uploading image content. However, straightforward approaches to using such data for WSOD wastefully discard captions that do not exactly match object names. Instead, we show how to squeeze the most information out of these captions by training a text-only classifier that generalizes beyond dataset boundaries. Our discovery provides an opportunity for learning detection models from noisy but more abundant and freely-available caption data. We also validate our model on three classic object detection benchmarks and achieve state-of-the-art WSOD performance. Our code is available at https://github.com/yekeren/Cap2Det. Keren Ye, Adriana Kovashka, Wei Li 0044, Danfeng Qin, Jesse Berent |
ICCV | 3 |
| 2019 | Predicting the Politics of an Image Using Webly Supervised DataabstractThe news media shape public opinion, and often, the visual bias they contain is evident for human observers. This bias can be inferred from how different media sources portray different subjects or topics. In this paper, we model visual political bias in contemporary media sources at scale, using webly supervised data. We collect a dataset of over one million unique images and associated news articles from left- and right-leaning news sources, and develop a method to predict the image's political leaning. This problem is particularly challenging because of the enormous intra-class visual and semantic diversity of our data. We propose a two-stage method to tackle this problem. In the first stage, the model is forced to learn relevant visual concepts that, when joined with document embeddings computed from articles paired with the images, enable the model to predict bias. In the second stage, we remove the requirement of the text domain and train a visual classifier from the features of the former model. We show this two-stage approach facilitates learning and outperforms several strong baselines. We also present extensive qualitative results demonstrating the nuances of the data. Christopher Thomas 0004, Adriana Kovashka |
NeurIPS | 2 |
| 2018 | Asking Friendly Strangers: Non-Semantic Attribute Transfer
Nils Murrugarra-Llerena, Adriana Kovashka |
AAAI | 2 |
| 2018 | Artistic Object Recognition by Unsupervised Style Adaptation
Christopher Thomas 0004, Adriana Kovashka |
ACCV (3) | 2 |
| 2018 | Persuasive Faces: Generating Faces in Advertisements
Christopher Thomas 0004, Adriana Kovashka |
BMVC | 2 |
| 2018 | Image Retrieval with Mixed Initiative and Multimodal Feedback
Nils Murrugarra-Llerena, Adriana Kovashka |
BMVC | 2 |
| 2018 | Story Understanding in Video Advertisements
Keren Ye, Kyle Buettner, Adriana Kovashka |
BMVC | 3 |
| 2018 | Equal But Not The Same: Understanding the Implicit Relationship Between Persuasive Images and Text
Rebecca Hwa, Adriana Kovashka |
BMVC | 3 |
| 2018 | ADVISE: Symbolism and External Knowledge for Decoding Advertisements
Keren Ye, Adriana Kovashka |
ECCV (15) | 2 |
| 2017 | Confidence and Diversity for Active Selection of Feedback in Image Retrieval
Bhavin Modi, Adriana Kovashka |
BMVC | 2 |
| 2017 | Automatic Understanding of Image and Video AdvertisementsabstractThere is more to images than their objective physical content: for example, advertisements are created to persuade a viewer to take a certain action. We propose the novel problem of automatic advertisement understanding. To enable research on this problem, we create two datasets: an image dataset of 64,832 image ads, and a video dataset of 3,477 ads. Our data contains rich annotations encompassing the topic and sentiment of the ads, questions and answers describing what actions the viewer is prompted to take and the reasoning that the ad presents to persuade the viewer (What should I do according to this ad, and why should I do it?), and symbolic references ads make (e.g. a dove symbolizes peace). We also analyze the most common persuasive strategies ads use, and the capabilities that computer vision systems should have to understand these strategies. We present baseline classification results for several prediction tasks, including automatically answering questions about the messages of the ads. Zaeem Hussain, Xiaozhong Zhang, Keren Ye, Christopher Thomas 0004, Zuha Agha, Nathan Ong, Adriana Kovashka |
CVPR | 8 |
| 2017 | Detecting Sexually Provocative ImagesabstractWhile the abundance of visual content available on the Internet, and the easy access to such content by all users allows us to find relevant content quickly, it also poses challenges. For example, if a parent wants to restrict the visual content which their child can see, this content needs to either be automatically tagged as offensive or not, or a computer vision algorithm needs to be trained to detect offensive content. One type of potentially offensive content is sexually explicit or provocative imagery. An image may be sexually provocative if it portrays nudity, but the sexual innuendo could also be contained in the body posture or facial expression of the human subject shown in the photo. Existing methods simply analyze skin exposure, but fail to capture the hidden intent behind images. Thus, they are unable to capture several important ways in which an image might be sexually provocative, hence offensive to children. We propose to address this problem by extracting a unified feature descriptor constituting the percentage of skin exposure, the body posture of the human in the image, and his/her gestures and facial expressions. We learn to predict these cues, then train a hierarchical model which combines them. We show in experiments that this model more accurately detects sexual innuendos behind images. Debashis Ganguly, Mohammad H. Mofrad, Adriana Kovashka |
WACV | 3 |
| 2017 | Learning Attributes from Human GazeabstractWhile semantic visual attributes have been shown useful for a variety of tasks, many attributes are difficult to model computationally. One of the reasons for this difficulty is that it is not clear where in an image the attribute lives. We propose to tackle this problem by involving humans more directly in the process of learning an attribute model. We ask humans to examine a set of images to determine if a given attribute is present in them, and we record where they looked. We create gaze maps for each attribute, and use these gaze maps to improve attribute prediction models. For test images we do not have gaze maps available, so we predict them based on models learned from collected gaze maps for each attribute of interest. Compared to six baselines, we improve prediction accuracies on attributes of faces and shoes, and we show how our method might be adapted for scene images. We demonstrate additional uses of our gaze maps for visualization of attribute models and learning "schools of thought" between users in terms of their understanding of the attribute. Nils Murrugarra-Llerena, Adriana Kovashka |
WACV | 2 |
| 2016 | Seeing Behind the Camera: Identifying the Authorship of a PhotographabstractWe introduce the novel problem of identifying the photographer behind a photograph. To explore the feasibility of current computer vision techniques to address this problem, we created a new dataset of over 180,000 images taken by 41 well-known photographers. Using this dataset, we examined the effectiveness of a variety of features (low and high-level, including CNN features) at identifying the photographer. We also trained a new deep convolutional neural network for this task. Our results show that high-level features greatly outperform low-level features. We provide qualitative results using these learned models that give insight into our method's ability to distinguish between photographers, and allow us to draw interesting conclusions about what specific photographers shoot. We also demonstrate two applications of our method. Christopher Thomas 0004, Adriana Kovashka |
CVPR | 2 |
| 2016 | Adapting attributes by selecting features similar across domainsabstractAttributes are semantic visual properties shared by objects. They have been shown to improve object recognition and to enhance content-based image search. While attributes are expected to cover multiple categories, e.g. a dalmatian and a whale can both have "smooth skin", we find that the appearance of a single attribute varies quite a bit across categories. Thus, an attribute model learned on one category may not be usable on another category. We show how to adapt attribute models towards new categories. We ensure that positive transfer can occur between a source domain of categories and a novel target domain, by learning in a feature subspace found by feature selection where the data distributions of the domains are similar. We demonstrate that when data from the novel domain is limited, regularizing attribute models for that novel domain with models trained on an auxiliary domain (via Adaptive SVM) improves the accuracy of attribute prediction. Adriana Kovashka |
WACV | 2 |
| 2015 | Discovering Attribute Shades of Meaning with the Crowd
Adriana Kovashka, Kristen Grauman |
Int. J. Comput. Vis. | 1 |
| 2015 | WhittleSearch: Interactive Image Search with Relative Attribute Feedback
Adriana Kovashka, Devi Parikh, Kristen Grauman |
Int. J. Comput. Vis. | 1 |
| 2013 | Attribute Pivots for Guiding Relevance Feedback in Image SearchabstractIn interactive image search, a user iteratively refines his results by giving feedback on exemplar images. Active selection methods aim to elicit useful feedback, but traditional approaches suffer from expensive selection criteria and cannot predict in formativeness reliably due to the imprecision of relevance feedback. To address these drawbacks, we propose to actively select "pivot" exemplars for which feedback in the form of a visual comparison will most reduce the system's uncertainty. For example, the system might ask, "Is your target image more or less crowded than this image?" Our approach relies on a series of binary search trees in relative attribute space, together with a selection function that predicts the information gain were the user to compare his envisioned target to the next node deeper in a given attribute's tree. It makes interactive search more efficient than existing strategies-both in terms of the system's selection time as well as the user's feedback effort. Adriana Kovashka, Kristen Grauman |
ICCV | 1 |
| 2013 | Attribute Adaptation for Personalized Image SearchabstractCurrent methods learn monolithic attribute predictors, with the assumption that a single model is sufficient to reflect human understanding of a visual attribute. However, in reality, humans vary in how they perceive the association between a named property and image content. For example, two people may have slightly different internal models for what makes a shoe look "formal", or they may disagree on which of two scenes looks "more cluttered". Rather than discount these differences as noise, we propose to learn user-specific attribute models. We adapt a generic model trained with annotations from multiple users, tailoring it to satisfy user-specific labels. Furthermore, we propose novel techniques to infer user-specific labels based on transitivity and contradictions in the user's search history. We demonstrate that adapted attributes improve accuracy over both existing monolithic models as well as models that learn from scratch with user-specific data alone. In addition, we show how adapted attributes are useful to personalize image search, whether with binary or relative attributes. Adriana Kovashka, Kristen Grauman |
ICCV | 1 |
| 2012 | Relative Attributes for Enhanced Human-Machine CommunicationabstractWe propose to model relative attributes that capture the relationships between images and objects in terms of human-nameable visual properties. For example, the models can capture that animal A is 'furrier' than animal B, or image X is 'brighter' than image B. Given training data stating how object/scene categories relate according to different attributes, we learn a ranking function per attribute. The learned ranking functions predict the relative strength of each property in novel images. We show how these relative attribute predictions enable a variety of novel applications, including zero-shot learning from relative comparisons, automatic image description, image search with interactive feedback, and active learning of discriminative classifiers. We overview results demonstrating these applications with images of faces and natural scenes. Overall, we find that relative attributes enhance the precision of communication between humans and computer vision algorithms, providing the richer language needed to fluidly "teach" a system about visual concepts. Devi Parikh, Adriana Kovashka, Amar Parkash, Kristen Grauman |
AAAI | 2 |
| 2012 | WhittleSearch: Image search with relative attribute feedbackabstractWe propose a novel mode of feedback for image search, where a user describes which properties of exemplar images should be adjusted in order to more closely match his/her mental model of the image(s) sought. For example, perusing image results for a query “black shoes”, the user might state, “Show me shoe images like these, but sportier.” Offline, our approach first learns a set of ranking functions, each of which predicts the relative strength of a nameable attribute in an image (`sportiness', `furriness', etc.). At query time, the system presents an initial set of reference images, and the user selects among them to provide relative attribute feedback. Using the resulting constraints in the multi-dimensional attribute space, our method updates its relevance function and re-ranks the pool of images. This procedure iterates using the accumulated constraints until the top ranked images are acceptably close to the user's envisioned target. In this way, our approach allows a user to efficiently “whittle away” irrelevant portions of the visual feature space, using semantic language to precisely communicate her preferences to the system. We demonstrate the technique for refining image search for people, products, and scenes, and show it outperforms traditional binary relevance feedback in terms of search speed and accuracy. Adriana Kovashka, Devi Parikh, Kristen Grauman |
CVPR | 1 |
| 2011 | Actively selecting annotations among objects and attributesabstractWe present an active learning approach to choose image annotation requests among both object category labels and the objects' attribute labels. The goal is to solicit those labels that will best use human effort when training a multi-class object recognition model. In contrast to previous work in active visual category learning, our approach directly exploits the dependencies between human-nameable visual attributes and the objects they describe, shifting its requests in either label space accordingly. We adopt a discriminative latent model that captures object-attribute and attribute-attribute relationships, and then define a suitable entropy reduction selection criterion to predict the influence a new label might have throughout those connections. On three challenging datasets, we demonstrate that the method can more successfully accelerate object learning relative to both passive learning and traditional active learning approaches. Adriana Kovashka, Sudheendra Vijayanarasimhan, Kristen Grauman |
ICCV | 1 |
| 2010 | Learning a hierarchy of discriminative space-time neighborhood features for human action recognitionabstractRecent work shows how to use local spatio-temporal features to learn models of realistic human actions from video. However, existing methods typically rely on a predefined spatial binning of the local descriptors to impose spatial information beyond a pure “bag-of-words” model, and thus may fail to capture the most informative space-time relationships. We propose to learn the shapes of space-time feature neighborhoods that are most discriminative for a given action category. Given a set of training videos, our method first extracts local motion and appearance features, quantizes them to a visual vocabulary, and then forms candidate neighborhoods consisting of the words associated with nearby points and their orientation with respect to the central interest point. Rather than dictate a particular scaling of the spatial and temporal dimensions to determine which points are near, we show how to learn the class-specific distance functions that form the most informative configurations. Descriptors for these variable-sized neighborhoods are then recursively mapped to higher-level vocabularies, producing a hierarchy of space-time configurations at successively broader scales. Our approach yields state-of-the-art performance on the UCF Sports and KTH datasets. Adriana Kovashka, Kristen Grauman |
CVPR | 1 |