Christopher Thomas 0004

dblp:21/4235-4 · also Chris Thomas 0004, Christopher Lee Thomas · DBLP profile ↗
← Back
27ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0002-3226-396XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 7 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 6 since 2021Software engineering, systems software and programming languages · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LAMP: Learning Universal Adversarial Perturbations for Multi-Image Tasks via Pre-trained Models
abstract
Multimodal Large Language Models (MLLMs) have achieved remarkable performance across vision-language tasks. Recent advancements allow these models to process multiple images as inputs. However, the vulnerabilities of multi-image MLLMs remain unexplored. Existing adversarial attacks focus on single-image settings and often assume a white-box threat model which is impractical in many real-world scenarios. This paper introduces LAMP, a black-box method for learning UAPs targeting multi-image MLLMs. LAMP applies an attention-based constraint that which prevents the model from effectively aggregating information across images. LAMP also introduces a novel cross-image contagious constraint that forces perturbed tokens to influence clean tokens to spread adversarial effects without requiring all inputs to be modified. Additionally, an index-attention suppression loss creates a robust position invariant attack. Experimental results show that LAMP outperforms SOTA baselines and achieves the highest attack success rates across multiple vision-language tasks.
Alvi Md. Ishmam, Najibul Haque Sarker, Zaber Ibn Abdul Hakim, Christopher Thomas 0004
AAAI4
2026 SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal Models
abstract
Aafiya Shamshad Hussain, Gaurav Srivastava, Alvi Md Ishmam, Zaber Ibn Abdul Hakim, Chris Thomas. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Aafiya Hussain, Gaurav Srivastava 0012, Alvi Md. Ishmam, Zaber Ibn Abdul Hakim, Christopher Thomas 0004
ACL (1)5
2026 Designing Multi-Robot Ground Video Sensemaking with Public Safety Professionals
abstract
Videos from fleets of ground robots can advance public safety by providing scalable situational awareness and reducing professionals’ burden. Yet little is known about how to design and integrate multi-robot videos into public safety workflows. Collaborating with six police agencies, we examined how such videos could be made practical. In Study 1, we present the first testbed for multi-robot ground video sensemaking. The testbed includes 38 events of interest relevant to public safety, a dataset of 20 robot patrol videos (10 day/night pairs) covering EoI types, and 6 design requirements aimed at improving current video sensemaking practices. In Study 2, we built MRVS, a tool that augments multi-robot patrol video streams with a prompt-engineered video understanding model. Participants reported reduced manual workload and greater confidence with LLM-based explanations, while noting concerns about false alarms and privacy. We conclude with implications for designing future multi-robot video sensemaking tools.
Puqi Zhou, Ali Asgarov, Aafiya Hussain, Wonjoon Park, Amit Paudyal, Sameep Shrestha, Chia-Wei Tang, Michael F. Lighthiser, Michael R. Hieb, Xuesu Xiao, Christopher Thomas 0004, Sungsoo Ray Hong
CHI11
2025 Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval
abstract
Cross-modal image-text retrieval is challenging because of the diverse possible associations between content from different modalities.Traditional methods learn a single-vector embedding to represent semantics of each sample, but struggle to capture nuanced and diverse relationships that can exist across modalities.Setbased approaches, which represent each sample with multiple embeddings, offer a promising alternative, as they can capture richer and more diverse relationships.In this paper, we show that, despite their promise, these set-based representations continue to face issues including sparse supervision and set collapse, which limits their effectiveness.To address these challenges, we propose Maximal Pair Assignment Similarity to optimize one-to-one matching between embedding sets which preserve semantic diversity within the set.We also introduce two loss functions to further enhance the representations: Global Discriminative Loss to enhance distinction among embeddings, and Intra-Set Divergence Loss to prevent collapse within each set.Our method achieves state-of-theart performance on MS-COCO and Flickr30k without relying on external data.
Hani Alomari, Anushka Sivakumar, Christopher Thomas 0004
ACL (1)4
2025 Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language Models
abstract
Large Vision-Language Models (LVLMs) have achieved strong performance on visionlanguage tasks, particularly Visual Question Answering (VQA).While prior work has explored unimodal biases in VQA, the problem of selection bias in Multiple-Choice Question Answering (MCQA), where models may favor specific option tokens (e.g., "A") or positions, remains underexplored.In this paper, we investigate both the presence and nature of selection bias in LVLMs through fine-grained MCQA benchmarks spanning easy, medium, and hard difficulty levels, defined by the semantic similarity of the options.We further propose an inference-time logit-level debiasing method that estimates an ensemble bias vector from general and contextual prompts and applies confidence-adaptive corrections to the model's output.Our method mitigates bias without retraining and is compatible with frozen LVLMs.Extensive experiments across several state-ofthe-art models reveal consistent selection biases that intensify with task difficulty, and show that our mitigation approach significantly reduces bias while improving accuracy in challenging settings.This work offers new insights into the limitations of LVLMs in MCQA and presents a practical approach to improve their robustness in fine-grained visual reasoning.Datasets and code are available at: https://github.com/ Atabuzzaman/Selection-Bias-of-LVLMs Yellow-headed Blackbird B. This image shows a Yellow-headed Blackbird bird which has a black body, bright yellow head… C.This is a Rolls-Royce Phantom Drophead Coupe Convertible 2012 car which has a distinctive… D. This image shows a hot and sour soup food item which has a rich broth, egg strands, tofu, and … A. This image shows a Yellow-headed Blackbird bird which has a black body, bright yellow head… B. This image shows a Brown Creeper bird with small, dark eyes, a slender, upturned beak, and… C.This image shows a Florida Jay bird which has a small, blue-gray body, a white chest and belly… D. This image shows a Tree Sparrow bird which has a brown and white body, a chestnut cap … A. This image shows a Rusty Blackbird bird which has rusty-brown plumage in non-breeding… B This image shows a Red-winged Blackbird bird which has a black body, distinctive red and… C.This image shows a Brewer Blackbird bird which has a sleek black body, iridescent feathers… D. This image shows a Yellow-headed Blackbird bird which has a black body, bright yellow head… A. This is a Spyker C8 Coupe 2009 car which has a distinctive propeller logo, sleek aerodynamic … Hard Easy Medium
Md. Atabuzzaman, Ali Asgarov, Christopher Thomas 0004
EMNLP3
2025 Flexible-length Text Infilling for Discrete Diffusion Models
abstract
Discrete diffusion models are a new class of text generation models that offer advantages such as bidirectional context, parallelizable generation, and flexible prompting compared to autoregressive models.However, a critical limitation has been the inability to perform flexible-length or flexible-position text infilling without access to ground-truth positional data.We introduce DDOT 1 (Discrete Diffusion with Optimal Transport Position Coupling), a discrete diffusion model that overcomes this limitation by jointly denoising token values and token positions using a novel sample-level optimal transport coupling.This coupling preserves relative token ordering while dynamically adjusting the positions and lengths of infilled segments.DDOT is orthogonal to existing discrete text diffusion methods and is compatible with various pretrained text denoisers.On text-infilling benchmarks such as One-Billion-Word and Yelp, DDOT outperforms naive diffusion baselines and achieves performance on par with state-of-the-art nonautoregressive models, while improving training efficiency and prompting flexibility.
Anushka Sivakumar, Chia-Wei Tang, Christopher Thomas 0004
EMNLP4
2025 Advancing Chart Question Answering with Robust Chart Component Recognition
abstract
Chart comprehension presents significant challenges for machine learning models due to the diverse and intricate shapes of charts. Existing multimodal methods often over-look these visual features or fail to integrate them effectively for Chart Question Answering. To address this, we introduce CHARTFORMER, a unified framework that enhances chart component recognition by accurately identifying and classifying components such as bars, lines, pies, titles, legends, and axes. Additionally, we propose a novel Question-guided Deformable Co-Attention (QDCAt) mechanism, which fuses chart features encoded by Chart-former with the given question, leveraging the question's guidance to ground the correct answer. Extensive experiments demonstrate a 3.2% improvement in mAP over the baselines for chart component recognition. For ChartQA and OpenCQA tasks, our approach achieves improvements of 15.4% in accuracy and 0.8 in BLEU score, respectively, underscoring the robustness of our solution for detailed visual data interpretation across various applications.11The source code and dataset are publicly available at https://github.com/VT-NLP/chartQA
Hanwen Zheng, Christopher Thomas 0004, Lifu Huang
WACV3
2024 Beyond Grounding: Extracting Fine-Grained Event Hierarchies across Modalities
abstract
Events describe happenings in our world that are of importance. Naturally, understanding events mentioned in multimedia content and how they are related forms an important way of comprehending our world. Existing literature can infer if events across textual and visual (video) domains are identical (via grounding) and thus, on the same semantic level. However, grounding fails to capture the intricate cross-event relations that exist due to the same events being referred to on many semantic levels. For example, the abstract event of "war'' manifests at a lower semantic level through subevents "tanks firing'' (in video) and airplane "shot'' (in text), leading to a hierarchical, multimodal relationship between the events. In this paper, we propose the task of extracting event hierarchies from multimodal (video and text) data to capture how the same event manifests itself in different modalities at different semantic levels. This reveals the structure of events and is critical to understanding them. To support research on this task, we introduce the Multimodal Hierarchical Events (MultiHiEve) dataset. Unlike prior video-language datasets, MultiHiEve is composed of news video-article pairs, which makes it rich in event hierarchies. We densely annotate a part of the dataset to construct the test benchmark. We show the limitations of state-of-the-art unimodal and multimodal baselines on this task. Further, we address these limitations via a new weakly supervised model, leveraging only unannotated video-article pairs from MultiHiEve. We perform a thorough evaluation of our proposed method which demonstrates improved performance on this task and highlight opportunities for future research. Data: https://github.com/hayyubi/multihieve
Hammad A. Ayyubi, Christopher Thomas 0004, Lovish Chum, Rahul Lokesh, Long Chen 0016, Yulei Niu, Xudong Lin 0003, Xuande Feng, Jaywon Koo, Sounak Ray, Shih-Fu Chang
AAAI2
2024 MetaSumPerceiver: Multimodal Multi-Document Evidence Summarization for Fact-Checking
abstract
Fact-checking real-world claims often requires reviewing multiple multimodal documents to assess a claim's truthfulness, which is a highly laborious and time-consuming task.In this paper, we present a summarization model designed to generate claim-specific summaries useful for fact-checking from multimodal, multi-document datasets.The model takes inputs in the form of documents, images, and a claim, with the objective of assisting in fact-checking tasks.We introduce a dynamic perceiver-based model that can handle inputs from multiple modalities of arbitrary lengths.To train our model, we leverage a novel reinforcement learning-based entailment objective to generate summaries that provide evidence distinguishing between different truthfulness labels.To assess the efficacy of our approach, we conduct experiments on both an existing benchmark and a new dataset of multi-document claims that we contribute.Our approach outperforms the SOTA approach by 4.6% in the claim verification task on the MOCHEG dataset and demonstrates strong performance on our new Multi-News-Fact-Checking dataset.
Ting-Chih Chen, Chia-Wei Tang, Christopher Thomas 0004
ACL (1)3
2024 Semantic Shield: Defending Vision-Language Models Against Backdooring and Poisoning via Fine-Grained Knowledge Alignment
abstract
In recent years there has been enormous interest in vision-language models trained using self-supervised objectives. However, the use of large-scale datasets scraped from the web for training also makes these models vulnerable to potential security threats, such as backdooring and poisoning attacks. In this paper, we propose a method for mitigating such attacks on contrastively trained vision-language models. Our approach leverages external knowledge extracted from a language model to prevent models from learning correlations between image regions which lack strong alignment with external knowledge. We do this by imposing constraints to enforce that attention paid by the model to visual regions is proportional to the alignment of those regions with external knowledge. We conduct extensive experiments using a variety of recent backdooring and poisoning attacks on multiple datasets and architectures. Our results clearly demonstrate that our proposed approach is highly effective at defending against such attacks across multiple settings, while maintaining model utility and without requiring any changes at inference time.
Alvi Md. Ishmam, Christopher Thomas 0004
CVPR2
2024 M3D: MultiModal MultiDocument Fine-Grained Inconsistency Detection
abstract
Validating claims from misinformation is a highly challenging task that involves understanding how each factual assertion within the claim relates to a set of trusted source materials. Existing approaches often make coarse-grained predictions but fail to identify the specific aspects of the claim that are troublesome and the specific evidence relied upon. In this paper, we introduce a method and new benchmark for this challenging task. Our method predicts the fine-grained logical relationship of each aspect of the claim from a set of multimodal documents, which include text, image(s), video(s), and audio(s). We also introduce a new benchmark (M^3DC) of claims requiring multimodal multidocument reasoning, which we construct using a novel claim synthesis technique. Experiments show that our approach significantly outperforms state-of-the-art baselines on this challenging task on two benchmarks while providing finer-grained predictions, explanations, and evidence.
Chia-Wei Tang, Ting-Chih Chen, Kiet Nguyen, Kazi Sajeed Mehrab, Alvi Md. Ishmam, Christopher Thomas 0004
EMNLP6
2024 Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
Junzhang Liu, Zhecan Wang, Hammad A. Ayyubi, Haoxuan You, Christopher Thomas 0004, Rui Sun 0011, Shih-Fu Chang, Kai-Wei Chang 0001
ACM Multimedia5
2024 JourneyBench: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated Images
abstract
Existing vision-language understanding benchmarks largely consist of images of objects in their usual contexts.As a consequence, recent multimodal large language models can perform well with only a shallow visual understanding by relying on background language biases. Thus, strong performance on these benchmarks does not necessarily correlate with strong visual understanding. In this paper, we release JourneyBench, a comprehensive human-annotated benchmark of generated images designed to assess the model's fine-grained multimodal reasoning abilities across five tasks: complementary multimodal chain of thought, multi-image VQA, imaginary image captioning, VQA with hallucination triggers, and fine-grained retrieval with sample-specific distractors.Unlike existing benchmarks, JourneyBench explicitly requires fine-grained multimodal reasoning in unusual imaginary scenarios where language bias and holistic image gist are insufficient. We benchmark state-of-the-art models on JourneyBench and analyze performance along a number of fine-grained dimensions. Results across all five tasks show that JourneyBench is exceptionally challenging for even the best models, indicating that models' visual reasoning abilities are not as strong as they first appear. We discuss the implications of our findings and propose avenues for further research.
Zhecan Wang, Junzhang Liu, Chia-Wei Tang, Hani Alomari, Anushka Sivakumar, Rui Sun 0011, Md. Atabuzzaman, Hammad A. Ayyubi, Haoxuan You, Alvi Md. Ishmam, Kai-Wei Chang 0001, Shih-Fu Chang, Christopher Thomas 0004
NeurIPS14
2023 Learning to Overcome Noise in Weak Caption Supervision for Object Detection
abstract
We propose the first mechanism to train object detection models from weak supervision in the form of captions at the image level. Language-based supervision for detection is appealing and inexpensive: many blogs with images and descriptive text written by human users exist. However, there is significant noise in this supervision: captions do not mention all objects that are shown, and may mention extraneous concepts. We first propose a technique to determine which image-caption pairs provide suitable signal for supervision. We further propose several complementary mechanisms to extract image-level pseudo labels for training from the caption. Finally, we train an iterative weakly-supervised object detection model from these image-level pseudo labels. We use captions from four datasets (COCO, Flickr30K, MIRFlickr1M, and Conceptual Captions) whose level of noise varies. We evaluate our approach on two object detection datasets. Weighting the labels extracted from different captions provides a boost over treating all captions equally. Further, our primary proposed technique for inferring pseudo labels for training at the image level, outperforms alternative techniques under a wide variety of settings. Both techniques generalize to datasets beyond the one they were trained on.
Mesut Erhan Unal, Keren Ye, Christopher Thomas 0004, Adriana Kovashka, Wei Li 0044, Danfeng Qin, Jesse Berent
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Fine-Grained Visual Entailment
Christopher Thomas 0004, Shih-Fu Chang
ECCV (36)1
2022 Weakly-Supervised Temporal Article Grounding
abstract
Long Chen, Yulei Niu, Brian Chen, Xudong Lin, Guangxing Han, Christopher Thomas, Hammad Ayyubi, Heng Ji, Shih-Fu Chang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Long Chen 0016, Yulei Niu, Brian Chen 0001, Xudong Lin 0003, Guangxing Han, Christopher Thomas 0004, Hammad A. Ayyubi, Heng Ji 0001, Shih-Fu Chang
EMNLP6
2021 InfoSurgeon: Cross-Media Fine-grained Information Consistency Checking for Fake News Detection
abstract
Yi Fung, Christopher Thomas, Revanth Gangi Reddy, Sandeep Polisetty, Heng Ji, Shih-Fu Chang, Kathleen McKeown, Mohit Bansal, Avi Sil. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yi R. Fung 0001, Christopher Thomas 0004, Revanth Gangi Reddy, Sandeep Polisetty, Heng Ji 0001, Shih-Fu Chang, Kathy McKeown, Mohit Bansal, Avirup Sil
ACL/IJCNLP (1)2
2021 Predicting Visual Political Bias Using Webly Supervised Data and an Auxiliary Task
Christopher Thomas 0004, Adriana Kovashka
Int. J. Comput. Vis.1
2020 Preserving Semantic Neighborhoods for Robust Cross-Modal Retrieval
Christopher Thomas 0004, Adriana Kovashka
ECCV (18)1
2019 Predicting the Politics of an Image Using Webly Supervised Data
abstract
The news media shape public opinion, and often, the visual bias they contain is evident for human observers. This bias can be inferred from how different media sources portray different subjects or topics. In this paper, we model visual political bias in contemporary media sources at scale, using webly supervised data. We collect a dataset of over one million unique images and associated news articles from left- and right-leaning news sources, and develop a method to predict the image's political leaning. This problem is particularly challenging because of the enormous intra-class visual and semantic diversity of our data. We propose a two-stage method to tackle this problem. In the first stage, the model is forced to learn relevant visual concepts that, when joined with document embeddings computed from articles paired with the images, enable the model to predict bias. In the second stage, we remove the requirement of the text domain and train a visual classifier from the features of the former model. We show this two-stage approach facilitates learning and outperforms several strong baselines. We also present extensive qualitative results demonstrating the nuances of the data.
Christopher Thomas 0004, Adriana Kovashka
NeurIPS1
2018 Artistic Object Recognition by Unsupervised Style Adaptation
Christopher Thomas 0004, Adriana Kovashka
ACCV (3)1
2018 Persuasive Faces: Generating Faces in Advertisements
Christopher Thomas 0004, Adriana Kovashka
BMVC1
2017 Automatic Understanding of Image and Video Advertisements
abstract
There is more to images than their objective physical content: for example, advertisements are created to persuade a viewer to take a certain action. We propose the novel problem of automatic advertisement understanding. To enable research on this problem, we create two datasets: an image dataset of 64,832 image ads, and a video dataset of 3,477 ads. Our data contains rich annotations encompassing the topic and sentiment of the ads, questions and answers describing what actions the viewer is prompted to take and the reasoning that the ad presents to persuade the viewer (What should I do according to this ad, and why should I do it?), and symbolic references ads make (e.g. a dove symbolizes peace). We also analyze the most common persuasive strategies ads use, and the capabilities that computer vision systems should have to understand these strategies. We present baseline classification results for several prediction tasks, including automatically answering questions about the messages of the ads.
Zaeem Hussain, Xiaozhong Zhang, Keren Ye, Christopher Thomas 0004, Zuha Agha, Nathan Ong, Adriana Kovashka
CVPR5
2016 Seeing Behind the Camera: Identifying the Authorship of a Photograph
abstract
We introduce the novel problem of identifying the photographer behind a photograph. To explore the feasibility of current computer vision techniques to address this problem, we created a new dataset of over 180,000 images taken by 41 well-known photographers. Using this dataset, we examined the effectiveness of a variety of features (low and high-level, including CNN features) at identifying the photographer. We also trained a new deep convolutional neural network for this task. Our results show that high-level features greatly outperform low-level features. We provide qualitative results using these learned models that give insight into our method's ability to distinguish between photographers, and allow us to draw interesting conclusions about what specific photographers shoot. We also demonstrate two applications of our method.
Christopher Thomas 0004, Adriana Kovashka
CVPR1
2015 Application of Slow Intelligence Framework for Smart Pet Care System Design
abstract
This article presents the design of a smart pet care system based on the slow intelligence framework for providing pets with suitable living conditions that closely mirror their natural habitat.By integrating heterogeneous information from various sensing data, the smart environment-aware pet care system can adaptively adjust the setting of temperature and humidity that best fits for the pet through iterative slow intelligence computation.Simulations of two case studies were provided to illustrate the application of the proposed system for pets such as snakes and dogs.The simulation results demonstrate the feasibility of the proposed approach to the design of smart pet care systems.
Shi-Kuo Chang, Wen-Hui Chen, Wen-Chyi Lin, Christopher Thomas 0004
SEKE4
2015 Application of Slow Intelligence Framework for Smart Pet Care System Design
abstract
This article presents the design of a smart pet care system based on the slow intelligence framework for providing pets with suitable living conditions that closely mirror their natural habitat. By integrating heterogeneous information from various sensing data, the smart environment-aware pet care system can adaptively adjust the setting of temperature and humidity that best fits the pet through iterative slow intelligence computation. Simulations of two case studies were provided to illustrate the application of the proposed system for pets such as snakes and dogs. The simulation results demonstrate the feasibility of the proposed approach to the design of smart pet care systems.
Shi-Kuo Chang, Wen-Hui Chen, Wen-Chyi Lin, Christopher Thomas 0004
Int. J. Softw. Eng. Knowl. Eng.4
2014 TBI-Doc: Generating Patient & Clinician Reports from Brain Imaging Data
abstract
The TBI-Doc prototype demonstrates the feasibility of automatically producing draft case reports for a new brain imaging technology, High Definition Fiber Tracking (HDFT). Here we describe the ontology for the HDFT domain, the system architecture and our goals for future research and development.
Pamela W. Jordan, Nancy L. Green, Christopher Thomas 0004, Susan Holm
INLG3