EDBT 2026 Demo / reviewers in the wild / expert
Christopher Thomas 0004
dblp:21/4235-4 · also Chris Thomas 0004, Christopher Lee Thomas
· DBLP profile ↗
27ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0002-3226-396XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 7 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 6 since 2021Software engineering, systems software and programming languages · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LAMP: Learning Universal Adversarial Perturbations for Multi-Image Tasks via Pre-trained ModelsabstractMultimodal Large Language Models (MLLMs) have achieved remarkable performance across vision-language tasks. Recent advancements allow these models to process multiple images as inputs. However, the vulnerabilities of multi-image MLLMs remain unexplored. Existing adversarial attacks focus on single-image settings and often assume a white-box threat model which is impractical in many real-world scenarios. This paper introduces LAMP, a black-box method for learning UAPs targeting multi-image MLLMs. LAMP applies an attention-based constraint that which prevents the model from effectively aggregating information across images. LAMP also introduces a novel cross-image contagious constraint that forces perturbed tokens to influence clean tokens to spread adversarial effects without requiring all inputs to be modified. Additionally, an index-attention suppression loss creates a robust position invariant attack. Experimental results show that LAMP outperforms SOTA baselines and achieves the highest attack success rates across multiple vision-language tasks. Alvi Md. Ishmam, Najibul Haque Sarker, Zaber Ibn Abdul Hakim, Christopher Thomas 0004 |
AAAI | 4 |
| 2026 | SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal ModelsabstractAafiya Shamshad Hussain, Gaurav Srivastava, Alvi Md Ishmam, Zaber Ibn Abdul Hakim, Chris Thomas. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Aafiya Hussain, Gaurav Srivastava 0012, Alvi Md. Ishmam, Zaber Ibn Abdul Hakim, Christopher Thomas 0004 |
ACL (1) | 5 |
| 2026 | Designing Multi-Robot Ground Video Sensemaking with Public Safety ProfessionalsabstractVideos from fleets of ground robots can advance public safety by providing scalable situational awareness and reducing professionals’ burden. Yet little is known about how to design and integrate multi-robot videos into public safety workflows. Collaborating with six police agencies, we examined how such videos could be made practical. In Study 1, we present the first testbed for multi-robot ground video sensemaking. The testbed includes 38 events of interest relevant to public safety, a dataset of 20 robot patrol videos (10 day/night pairs) covering EoI types, and 6 design requirements aimed at improving current video sensemaking practices. In Study 2, we built MRVS, a tool that augments multi-robot patrol video streams with a prompt-engineered video understanding model. Participants reported reduced manual workload and greater confidence with LLM-based explanations, while noting concerns about false alarms and privacy. We conclude with implications for designing future multi-robot video sensemaking tools. Puqi Zhou, Ali Asgarov, Aafiya Hussain, Wonjoon Park, Amit Paudyal, Sameep Shrestha, Chia-Wei Tang, Michael F. Lighthiser, Michael R. Hieb, Xuesu Xiao, Christopher Thomas 0004, Sungsoo Ray Hong |
CHI | 11 |
| 2025 | Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal RetrievalabstractCross-modal image-text retrieval is challenging because of the diverse possible associations between content from different modalities.Traditional methods learn a single-vector embedding to represent semantics of each sample, but struggle to capture nuanced and diverse relationships that can exist across modalities.Setbased approaches, which represent each sample with multiple embeddings, offer a promising alternative, as they can capture richer and more diverse relationships.In this paper, we show that, despite their promise, these set-based representations continue to face issues including sparse supervision and set collapse, which limits their effectiveness.To address these challenges, we propose Maximal Pair Assignment Similarity to optimize one-to-one matching between embedding sets which preserve semantic diversity within the set.We also introduce two loss functions to further enhance the representations: Global Discriminative Loss to enhance distinction among embeddings, and Intra-Set Divergence Loss to prevent collapse within each set.Our method achieves state-of-theart performance on MS-COCO and Flickr30k without relying on external data. Hani Alomari, Anushka Sivakumar, Christopher Thomas 0004 |
ACL (1) | 4 |
| 2025 | Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language ModelsabstractLarge Vision-Language Models (LVLMs) have achieved strong performance on visionlanguage tasks, particularly Visual Question Answering (VQA).While prior work has explored unimodal biases in VQA, the problem of selection bias in Multiple-Choice Question Answering (MCQA), where models may favor specific option tokens (e.g., "A") or positions, remains underexplored.In this paper, we investigate both the presence and nature of selection bias in LVLMs through fine-grained MCQA benchmarks spanning easy, medium, and hard difficulty levels, defined by the semantic similarity of the options.We further propose an inference-time logit-level debiasing method that estimates an ensemble bias vector from general and contextual prompts and applies confidence-adaptive corrections to the model's output.Our method mitigates bias without retraining and is compatible with frozen LVLMs.Extensive experiments across several state-ofthe-art models reveal consistent selection biases that intensify with task difficulty, and show that our mitigation approach significantly reduces bias while improving accuracy in challenging settings.This work offers new insights into the limitations of LVLMs in MCQA and presents a practical approach to improve their robustness in fine-grained visual reasoning.Datasets and code are available at: https://github.com/ Atabuzzaman/Selection-Bias-of-LVLMs Yellow-headed Blackbird B. This image shows a Yellow-headed Blackbird bird which has a black body, bright yellow head… C.This is a Rolls-Royce Phantom Drophead Coupe Convertible 2012 car which has a distinctive… D. This image shows a hot and sour soup food item which has a rich broth, egg strands, tofu, and … A. This image shows a Yellow-headed Blackbird bird which has a black body, bright yellow head… B. This image shows a Brown Creeper bird with small, dark eyes, a slender, upturned beak, and… C.This image shows a Florida Jay bird which has a small, blue-gray body, a white chest and belly… D. This image shows a Tree Sparrow bird which has a brown and white body, a chestnut cap … A. This image shows a Rusty Blackbird bird which has rusty-brown plumage in non-breeding… B This image shows a Red-winged Blackbird bird which has a black body, distinctive red and… C.This image shows a Brewer Blackbird bird which has a sleek black body, iridescent feathers… D. This image shows a Yellow-headed Blackbird bird which has a black body, bright yellow head… A. This is a Spyker C8 Coupe 2009 car which has a distinctive propeller logo, sleek aerodynamic … Hard Easy Medium Md. Atabuzzaman, Ali Asgarov, Christopher Thomas 0004 |
EMNLP | 3 |
| 2025 | Flexible-length Text Infilling for Discrete Diffusion ModelsabstractDiscrete diffusion models are a new class of text generation models that offer advantages such as bidirectional context, parallelizable generation, and flexible prompting compared to autoregressive models.However, a critical limitation has been the inability to perform flexible-length or flexible-position text infilling without access to ground-truth positional data.We introduce DDOT 1 (Discrete Diffusion with Optimal Transport Position Coupling), a discrete diffusion model that overcomes this limitation by jointly denoising token values and token positions using a novel sample-level optimal transport coupling.This coupling preserves relative token ordering while dynamically adjusting the positions and lengths of infilled segments.DDOT is orthogonal to existing discrete text diffusion methods and is compatible with various pretrained text denoisers.On text-infilling benchmarks such as One-Billion-Word and Yelp, DDOT outperforms naive diffusion baselines and achieves performance on par with state-of-the-art nonautoregressive models, while improving training efficiency and prompting flexibility. Anushka Sivakumar, Chia-Wei Tang, Christopher Thomas 0004 |
EMNLP | 4 |
| 2025 | Advancing Chart Question Answering with Robust Chart Component RecognitionabstractChart comprehension presents significant challenges for machine learning models due to the diverse and intricate shapes of charts. Existing multimodal methods often over-look these visual features or fail to integrate them effectively for Chart Question Answering. To address this, we introduce CHARTFORMER, a unified framework that enhances chart component recognition by accurately identifying and classifying components such as bars, lines, pies, titles, legends, and axes. Additionally, we propose a novel Question-guided Deformable Co-Attention (QDCAt) mechanism, which fuses chart features encoded by Chart-former with the given question, leveraging the question's guidance to ground the correct answer. Extensive experiments demonstrate a 3.2% improvement in mAP over the baselines for chart component recognition. For ChartQA and OpenCQA tasks, our approach achieves improvements of 15.4% in accuracy and 0.8 in BLEU score, respectively, underscoring the robustness of our solution for detailed visual data interpretation across various applications.11The source code and dataset are publicly available at https://github.com/VT-NLP/chartQA Hanwen Zheng, Christopher Thomas 0004, Lifu Huang |
WACV | 3 |
| 2024 | Beyond Grounding: Extracting Fine-Grained Event Hierarchies across ModalitiesabstractEvents describe happenings in our world that are of importance. Naturally, understanding events mentioned in multimedia content and how they are related forms an important way of comprehending our world. Existing literature can infer if events across textual and visual (video) domains are identical (via grounding) and thus, on the same semantic level. However, grounding fails to capture the intricate cross-event relations that exist due to the same events being referred to on many semantic levels. For example, the abstract event of "war'' manifests at a lower semantic level through subevents "tanks firing'' (in video) and airplane "shot'' (in text), leading to a hierarchical, multimodal relationship between the events. In this paper, we propose the task of extracting event hierarchies from multimodal (video and text) data to capture how the same event manifests itself in different modalities at different semantic levels. This reveals the structure of events and is critical to understanding them. To support research on this task, we introduce the Multimodal Hierarchical Events (MultiHiEve) dataset. Unlike prior video-language datasets, MultiHiEve is composed of news video-article pairs, which makes it rich in event hierarchies. We densely annotate a part of the dataset to construct the test benchmark. We show the limitations of state-of-the-art unimodal and multimodal baselines on this task. Further, we address these limitations via a new weakly supervised model, leveraging only unannotated video-article pairs from MultiHiEve. We perform a thorough evaluation of our proposed method which demonstrates improved performance on this task and highlight opportunities for future research. Data: https://github.com/hayyubi/multihieve Hammad A. Ayyubi, Christopher Thomas 0004, Lovish Chum, Rahul Lokesh, Long Chen 0016, Yulei Niu, Xudong Lin 0003, Xuande Feng, Jaywon Koo, Sounak Ray, Shih-Fu Chang |
AAAI | 2 |
| 2024 | MetaSumPerceiver: Multimodal Multi-Document Evidence Summarization for Fact-CheckingabstractFact-checking real-world claims often requires reviewing multiple multimodal documents to assess a claim's truthfulness, which is a highly laborious and time-consuming task.In this paper, we present a summarization model designed to generate claim-specific summaries useful for fact-checking from multimodal, multi-document datasets.The model takes inputs in the form of documents, images, and a claim, with the objective of assisting in fact-checking tasks.We introduce a dynamic perceiver-based model that can handle inputs from multiple modalities of arbitrary lengths.To train our model, we leverage a novel reinforcement learning-based entailment objective to generate summaries that provide evidence distinguishing between different truthfulness labels.To assess the efficacy of our approach, we conduct experiments on both an existing benchmark and a new dataset of multi-document claims that we contribute.Our approach outperforms the SOTA approach by 4.6% in the claim verification task on the MOCHEG dataset and demonstrates strong performance on our new Multi-News-Fact-Checking dataset. Ting-Chih Chen, Chia-Wei Tang, Christopher Thomas 0004 |
ACL (1) | 3 |
| 2024 | Semantic Shield: Defending Vision-Language Models Against Backdooring and Poisoning via Fine-Grained Knowledge AlignmentabstractIn recent years there has been enormous interest in vision-language models trained using self-supervised objectives. However, the use of large-scale datasets scraped from the web for training also makes these models vulnerable to potential security threats, such as backdooring and poisoning attacks. In this paper, we propose a method for mitigating such attacks on contrastively trained vision-language models. Our approach leverages external knowledge extracted from a language model to prevent models from learning correlations between image regions which lack strong alignment with external knowledge. We do this by imposing constraints to enforce that attention paid by the model to visual regions is proportional to the alignment of those regions with external knowledge. We conduct extensive experiments using a variety of recent backdooring and poisoning attacks on multiple datasets and architectures. Our results clearly demonstrate that our proposed approach is highly effective at defending against such attacks across multiple settings, while maintaining model utility and without requiring any changes at inference time. Alvi Md. Ishmam, Christopher Thomas 0004 |
CVPR | 2 |
| 2024 | M3D: MultiModal MultiDocument Fine-Grained Inconsistency DetectionabstractValidating claims from misinformation is a highly challenging task that involves understanding how each factual assertion within the claim relates to a set of trusted source materials. Existing approaches often make coarse-grained predictions but fail to identify the specific aspects of the claim that are troublesome and the specific evidence relied upon. In this paper, we introduce a method and new benchmark for this challenging task. Our method predicts the fine-grained logical relationship of each aspect of the claim from a set of multimodal documents, which include text, image(s), video(s), and audio(s). We also introduce a new benchmark (M^3DC) of claims requiring multimodal multidocument reasoning, which we construct using a novel claim synthesis technique. Experiments show that our approach significantly outperforms state-of-the-art baselines on this challenging task on two benchmarks while providing finer-grained predictions, explanations, and evidence. Chia-Wei Tang, Ting-Chih Chen, Kiet Nguyen, Kazi Sajeed Mehrab, Alvi Md. Ishmam, Christopher Thomas 0004 |
EMNLP | 6 |
| 2024 | Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
Junzhang Liu, Zhecan Wang, Hammad A. Ayyubi, Haoxuan You, Christopher Thomas 0004, Rui Sun 0011, Shih-Fu Chang, Kai-Wei Chang 0001 |
ACM Multimedia | 5 |
| 2024 | JourneyBench: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated ImagesabstractExisting vision-language understanding benchmarks largely consist of images of objects in their usual contexts.As a consequence, recent multimodal large language models can perform well with only a shallow visual understanding by relying on background language biases. Thus, strong performance on these benchmarks does not necessarily correlate with strong visual understanding. In this paper, we release JourneyBench, a comprehensive human-annotated benchmark of generated images designed to assess the model's fine-grained multimodal reasoning abilities across five tasks: complementary multimodal chain of thought, multi-image VQA, imaginary image captioning, VQA with hallucination triggers, and fine-grained retrieval with sample-specific distractors.Unlike existing benchmarks, JourneyBench explicitly requires fine-grained multimodal reasoning in unusual imaginary scenarios where language bias and holistic image gist are insufficient. We benchmark state-of-the-art models on JourneyBench and analyze performance along a number of fine-grained dimensions. Results across all five tasks show that JourneyBench is exceptionally challenging for even the best models, indicating that models' visual reasoning abilities are not as strong as they first appear. We discuss the implications of our findings and propose avenues for further research. Zhecan Wang, Junzhang Liu, Chia-Wei Tang, Hani Alomari, Anushka Sivakumar, Rui Sun 0011, Md. Atabuzzaman, Hammad A. Ayyubi, Haoxuan You, Alvi Md. Ishmam, Kai-Wei Chang 0001, Shih-Fu Chang, Christopher Thomas 0004 |
NeurIPS | 14 |
| 2023 | Learning to Overcome Noise in Weak Caption Supervision for Object DetectionabstractWe propose the first mechanism to train object detection models from weak supervision in the form of captions at the image level. Language-based supervision for detection is appealing and inexpensive: many blogs with images and descriptive text written by human users exist. However, there is significant noise in this supervision: captions do not mention all objects that are shown, and may mention extraneous concepts. We first propose a technique to determine which image-caption pairs provide suitable signal for supervision. We further propose several complementary mechanisms to extract image-level pseudo labels for training from the caption. Finally, we train an iterative weakly-supervised object detection model from these image-level pseudo labels. We use captions from four datasets (COCO, Flickr30K, MIRFlickr1M, and Conceptual Captions) whose level of noise varies. We evaluate our approach on two object detection datasets. Weighting the labels extracted from different captions provides a boost over treating all captions equally. Further, our primary proposed technique for inferring pseudo labels for training at the image level, outperforms alternative techniques under a wide variety of settings. Both techniques generalize to datasets beyond the one they were trained on. Mesut Erhan Unal, Keren Ye, Christopher Thomas 0004, Adriana Kovashka, Wei Li 0044, Danfeng Qin, Jesse Berent |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Fine-Grained Visual Entailment
Christopher Thomas 0004, Shih-Fu Chang |
ECCV (36) | 1 |
| 2022 | Weakly-Supervised Temporal Article GroundingabstractLong Chen, Yulei Niu, Brian Chen, Xudong Lin, Guangxing Han, Christopher Thomas, Hammad Ayyubi, Heng Ji, Shih-Fu Chang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Long Chen 0016, Yulei Niu, Brian Chen 0001, Xudong Lin 0003, Guangxing Han, Christopher Thomas 0004, Hammad A. Ayyubi, Heng Ji 0001, Shih-Fu Chang |
EMNLP | 6 |
| 2021 | InfoSurgeon: Cross-Media Fine-grained Information Consistency Checking for Fake News DetectionabstractYi Fung, Christopher Thomas, Revanth Gangi Reddy, Sandeep Polisetty, Heng Ji, Shih-Fu Chang, Kathleen McKeown, Mohit Bansal, Avi Sil. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yi R. Fung 0001, Christopher Thomas 0004, Revanth Gangi Reddy, Sandeep Polisetty, Heng Ji 0001, Shih-Fu Chang, Kathy McKeown, Mohit Bansal, Avirup Sil |
ACL/IJCNLP (1) | 2 |
| 2021 | Predicting Visual Political Bias Using Webly Supervised Data and an Auxiliary Task
Christopher Thomas 0004, Adriana Kovashka |
Int. J. Comput. Vis. | 1 |
| 2020 | Preserving Semantic Neighborhoods for Robust Cross-Modal Retrieval
Christopher Thomas 0004, Adriana Kovashka |
ECCV (18) | 1 |
| 2019 | Predicting the Politics of an Image Using Webly Supervised DataabstractThe news media shape public opinion, and often, the visual bias they contain is evident for human observers. This bias can be inferred from how different media sources portray different subjects or topics. In this paper, we model visual political bias in contemporary media sources at scale, using webly supervised data. We collect a dataset of over one million unique images and associated news articles from left- and right-leaning news sources, and develop a method to predict the image's political leaning. This problem is particularly challenging because of the enormous intra-class visual and semantic diversity of our data. We propose a two-stage method to tackle this problem. In the first stage, the model is forced to learn relevant visual concepts that, when joined with document embeddings computed from articles paired with the images, enable the model to predict bias. In the second stage, we remove the requirement of the text domain and train a visual classifier from the features of the former model. We show this two-stage approach facilitates learning and outperforms several strong baselines. We also present extensive qualitative results demonstrating the nuances of the data. Christopher Thomas 0004, Adriana Kovashka |
NeurIPS | 1 |
| 2018 | Artistic Object Recognition by Unsupervised Style Adaptation
Christopher Thomas 0004, Adriana Kovashka |
ACCV (3) | 1 |
| 2018 | Persuasive Faces: Generating Faces in Advertisements
Christopher Thomas 0004, Adriana Kovashka |
BMVC | 1 |
| 2017 | Automatic Understanding of Image and Video AdvertisementsabstractThere is more to images than their objective physical content: for example, advertisements are created to persuade a viewer to take a certain action. We propose the novel problem of automatic advertisement understanding. To enable research on this problem, we create two datasets: an image dataset of 64,832 image ads, and a video dataset of 3,477 ads. Our data contains rich annotations encompassing the topic and sentiment of the ads, questions and answers describing what actions the viewer is prompted to take and the reasoning that the ad presents to persuade the viewer (What should I do according to this ad, and why should I do it?), and symbolic references ads make (e.g. a dove symbolizes peace). We also analyze the most common persuasive strategies ads use, and the capabilities that computer vision systems should have to understand these strategies. We present baseline classification results for several prediction tasks, including automatically answering questions about the messages of the ads. Zaeem Hussain, Xiaozhong Zhang, Keren Ye, Christopher Thomas 0004, Zuha Agha, Nathan Ong, Adriana Kovashka |
CVPR | 5 |
| 2016 | Seeing Behind the Camera: Identifying the Authorship of a PhotographabstractWe introduce the novel problem of identifying the photographer behind a photograph. To explore the feasibility of current computer vision techniques to address this problem, we created a new dataset of over 180,000 images taken by 41 well-known photographers. Using this dataset, we examined the effectiveness of a variety of features (low and high-level, including CNN features) at identifying the photographer. We also trained a new deep convolutional neural network for this task. Our results show that high-level features greatly outperform low-level features. We provide qualitative results using these learned models that give insight into our method's ability to distinguish between photographers, and allow us to draw interesting conclusions about what specific photographers shoot. We also demonstrate two applications of our method. Christopher Thomas 0004, Adriana Kovashka |
CVPR | 1 |
| 2015 | Application of Slow Intelligence Framework for Smart Pet Care System DesignabstractThis article presents the design of a smart pet care system based on the slow intelligence framework for providing pets with suitable living conditions that closely mirror their natural habitat.By integrating heterogeneous information from various sensing data, the smart environment-aware pet care system can adaptively adjust the setting of temperature and humidity that best fits for the pet through iterative slow intelligence computation.Simulations of two case studies were provided to illustrate the application of the proposed system for pets such as snakes and dogs.The simulation results demonstrate the feasibility of the proposed approach to the design of smart pet care systems. Shi-Kuo Chang, Wen-Hui Chen, Wen-Chyi Lin, Christopher Thomas 0004 |
SEKE | 4 |
| 2015 | Application of Slow Intelligence Framework for Smart Pet Care System DesignabstractThis article presents the design of a smart pet care system based on the slow intelligence framework for providing pets with suitable living conditions that closely mirror their natural habitat. By integrating heterogeneous information from various sensing data, the smart environment-aware pet care system can adaptively adjust the setting of temperature and humidity that best fits the pet through iterative slow intelligence computation. Simulations of two case studies were provided to illustrate the application of the proposed system for pets such as snakes and dogs. The simulation results demonstrate the feasibility of the proposed approach to the design of smart pet care systems. Shi-Kuo Chang, Wen-Hui Chen, Wen-Chyi Lin, Christopher Thomas 0004 |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2014 | TBI-Doc: Generating Patient & Clinician Reports from Brain Imaging DataabstractThe TBI-Doc prototype demonstrates the feasibility of automatically producing draft case reports for a new brain imaging technology, High Definition Fiber Tracking (HDFT). Here we describe the ontology for the HDFT domain, the system architecture and our goals for future research and development. Pamela W. Jordan, Nancy L. Green, Christopher Thomas 0004, Susan Holm |
INLG | 3 |