VLDB 2026 Research / reviewers in the wild / expert
Artemis Panagopoulou
dblp:290/2230
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Vision and language · 63% 3D vision · 21% Trustworthy machine learning · 12% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
multimodal reasoning |
1.4 | 2 | 2025 | Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D · EMNLP 2025 Visual Goal-Step Inference using wikiHow · EMNLP (1) 2021 |
Computer vision › Vision and language
visual programming |
0.9 | 1 | 2025 | ViUniT: Visual Unit Tests for More Robust Visual Programming · CVPR 2025 |
Computer vision › Vision and language
visual reasoning |
0.9 | 1 | 2025 | ViUniT: Visual Unit Tests for More Robust Visual Programming · CVPR 2025 |
Multimedia analysis and retrieval › multimedia analysis
cross-media analysis |
0.9 | 1 | 2025 | Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D · EMNLP 2025 |
Computer vision › Vision and language › 3d vision and language
3d captioning |
0.8 | 1 | 2024 | ULIP-2: Towards Scalable Multimodal Pre-Training for 3D Understanding · CVPR 2024 |
Computer vision › 3D vision › 3d object recognition
3d object classification |
0.8 | 1 | 2024 | ULIP-2: Towards Scalable Multimodal Pre-Training for 3D Understanding · CVPR 2024 |
Computer vision › 3D vision › geometric deep learning
3d representation learning |
0.8 | 1 | 2024 | ULIP-2: Towards Scalable Multimodal Pre-Training for 3D Understanding · CVPR 2024 |
Computer vision › Vision and language
3d vision and language |
0.8 | 1 | 2024 | ULIP-2: Towards Scalable Multimodal Pre-Training for 3D Understanding · CVPR 2024 |
Computer vision › Vision and language
cross-modal alignment |
0.8 | 1 | 2024 | X-InstructBLIP: A Framework for Aligning Image, 3D, Audio, Video to LLMs and its Emergent Cross-Modal Reasoning · ECCV (45) 2024 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.8 | 1 | 2024 | X-InstructBLIP: A Framework for Aligning Image, 3D, Audio, Video to LLMs and its Emergent Cross-Modal Reasoning · ECCV (45) 2024 |
Computer vision › 3D vision › 3d object recognition › 3d object classification
zero-shot 3d classification |
0.8 | 1 | 2024 | ULIP-2: Towards Scalable Multimodal Pre-Training for 3D Understanding · CVPR 2024 |
Machine learning › Trustworthy machine learning › interpretability
concept bottleneck model |
0.7 | 1 | 2023 | Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image Classification · CVPR 2023 |
Machine learning › Trustworthy machine learning
interpretability |
0.7 | 1 | 2023 | Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image Classification · CVPR 2023 |
Computer vision › Vision and language › vision-language model
vision-language model application |
0.7 | 1 | 2023 | Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image Classification · CVPR 2023 |
Machine learning › Representation and self-supervised learning
multimodal representation learning |
0.2 | 1 | 2024 | X-InstructBLIP: A Framework for Aligning Image, 3D, Audio, Video to LLMs and its Emergent Cross-Modal Reasoning · ECCV (45) 2024 |
Computer vision › Video understanding and tracking › human action analysis
action understanding |
0.1 | 1 | 2021 | Visual Goal-Step Inference using wikiHow · EMNLP (1) 2021 |
Methods — techniques the papers use, named apart from their topics
cross-modal evaluation · 1.7contrastive learning · 1.7reinforcement learning · 0.9language model · 0.9image synthesis · 0.9large multimodal model · 0.8large language model · 0.8instruction tuning · 0.8contrastive pre-training · 0.8CLIP · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ViUniT: Visual Unit Tests for More Robust Visual ProgrammingabstractProgramming based approaches to reasoning tasks have substantially expanded the types of questions models can answer about visual scenes. Yet on benchmark visual reasoning data, when models answer correctly, they produce incorrect programs 33% of the time. These models are often right for the wrong reasons and risk unexpected failures on new data. Unit tests play a foundational role in ensuring code correctness and could be used to repair such failures. We propose Visual Unit Testing (ViUniT), a framework to improve the reliability of visual programs by automatically generating unit tests. In our framework, a unit test is represented as a novel image and answer pair meant to verify the logical correctness of a program produced for a given query. Our method leverages a language model to create unit tests in the form of image descriptions and expected answers, followed by image synthesis to produce corresponding images. We conduct a comprehensive analysis of what constitutes an effective visual unit test suite, exploring unit test generation, sampling strategies, image generation methods, and varying the number of programs and unit tests. Additionally, we introduce four applications of visual unit tests: best program selection, answer refusal, re-prompting, and unsupervised reward formulations for reinforcement learning. Experiments with two models across three datasets in visual question answering and image-text matching demonstrate that ViUniT improves model performance by 11.4 points in accuracy. Notably, it enables 7B open-source language models to outperform gpt-4o-mini in visual program generation by an average of 7.7 points and reduces the occurrence of programs that are correct for the wrong reasons by 40%. Artemis Panagopoulou, Honglu Zhou, Silvio Savarese, Caiming Xiong, Chris Callison-Burch, Mark Yatskar, Juan Carlos Niebles |
CVPR | 1 |
| 2025 | Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3DabstractArtemis Panagopoulou, Le Xue, Honglu Zhou, Silvio Savarese, Ran Xu, Caiming Xiong, Chris Callison-Burch, Mark Yatskar, Juan Carlos Niebles. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Artemis Panagopoulou, Le Xue, Honglu Zhou, Silvio Savarese, Ran Xu 0001, Caiming Xiong, Chris Callison-Burch, Mark Yatskar, Juan Carlos Niebles |
EMNLP | 1 |
| 2024 | ULIP-2: Towards Scalable Multimodal Pre-Training for 3D UnderstandingabstractRecent advancements in multimodal pretraining have shown promising efficacy in 3D representation learning by aligning multimodal features across 3D shapes, their 2D counterparts, and language descriptions. However, the methods used by existing frameworks to curate such multimodal data, in particular language descriptions for 3D shapes, are not scalable, and the collected language descriptions are not diverse. To address this, we introduce ULIP-2, a simple yet effective tri-modal pretraining framework that leverages large multimodal models to automatically generate holistic language descriptions for 3D shapes. It only needs 3D data as input, eliminating the need for any manual 3D annotations, and is therefore scalable to large datasets. ULIP-2 is also equipped with scaled-up backbones for better multimodal representation learning. We conduct experiments on two large-scale 3D datasets, Objaverse and ShapeNet, and augment them with tri-modal datasets of 3D point clouds, images, and language for training ULIP-2. Experiments show that ULIP-2 demonstrates substantial benefits in three downstream tasks: zero-shot 3D classification, standard 3D classification with fine-tuning, and 3D captioning (3D-to-language generation). It achieves a new SOTA of 50.6% (top-1) on Objaverse-LVIS and 84.7% (top-1) on ModelNet40 in zero-shot classification. In the ScanObjectNN benchmark for standard fine-tuning, ULIP-2 reaches an overall accuracy of 91.5% with a compact model of only 1.4 million parameters. ULIP-2 sheds light on a new paradigm for scalable multimodal 3D representation learning without human annotations and shows significant improvements over existing baselines. The code and datasets are released at https://github.com/salesforce/ULIP. Le Xue, Ning Yu 0006, Shu Zhang 0007, Artemis Panagopoulou, Junnan Li 0001, Roberto Martin Martin, Jiajun Wu 0001, Caiming Xiong, Ran Xu 0001, Juan Carlos Niebles, Silvio Savarese |
CVPR | 4 |
| 2024 | X-InstructBLIP: A Framework for Aligning Image, 3D, Audio, Video to LLMs and its Emergent Cross-Modal Reasoning
Artemis Panagopoulou, Le Xue, Ning Yu 0006, Junnan Li 0001, Dongxu Li 0003, Shafiq R. Joty, Ran Xu 0001, Silvio Savarese, Caiming Xiong, Juan Carlos Niebles |
ECCV (45) | 1 |
| 2023 | Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image ClassificationabstractConcept Bottleneck Models (CBM) are inherently interpretable models that factor model decisions into humanreadable concepts. They allow people to easily understand why a model is failing, a critical feature for high-stakes applications. CBMs require manually specified concepts and often under-perform their black box counterparts, preventing their broad adoption. We address these shortcomings and are first to show how to construct high-performance CBMs without manual specification of similar accuracy to black box models. Our approach, Language Guided Bottlenecks (LaBo), leverages a language model, GPT-3, to define a large space of possible bottlenecks. Given a problem domain, LaBo uses GPT-3 to produce factual sentences about categories to form candidate concepts. LaBo efficiently searches possible bottlenecks through a novel submodular utility that promotes the selection of discriminative and diverse information. Ultimately, GPT-3's sentential concepts can be aligned to images using CLIP, to form a bottleneck layer. Experiments demonstrate that LaBo is a highly effective prior for concepts important to visual recognition. In the evaluation with 11 diverse datasets, LaBo bottlenecks excel at few-shot classification: they are 11.7% more accurate than black box linear probes at 1 shot and comparable with more data. Overall, LaBo demonstrates that inherently interpretable models can be widely applied at similar, or better, performance than black box approaches.11Code and data are available at https://github.com/YueYANG1996/LaBo Yue Yang 0006, Artemis Panagopoulou, Shenghao Zhou, Daniel Jin, Chris Callison-Burch, Mark Yatskar |
CVPR | 2 |
| 2021 | Visual Goal-Step Inference using wikiHowabstractUnderstanding what sequence of steps are needed to complete a goal can help artificial intelligence systems reason about human activities.Past work in NLP has examined the task of goal-step inference for text.We introduce the visual analogue.We propose the Visual Goal-Step Inference (VGSI) task, where a model is given a textual goal and must choose which of four images represents a plausible step towards that goal.With a new dataset harvested from wikiHow consisting of 772,277 images representing human actions, we show that our task is challenging for state-of-theart multimodal models.Moreover, the multimodal representation learned from our data can be effectively transferred to other datasets like HowTo100m, increasing the VGSI accuracy by 15 -20%.Our task will facilitate multimodal reasoning about procedural events. Yue Yang 0006, Artemis Panagopoulou, Qing Lyu 0001, Li Zhang 0039, Mark Yatskar, Chris Callison-Burch |
EMNLP (1) | 2 |
| 2021 | Self-Supervised Optical Flow with Spiking Neural Networks and Event Based CamerasabstractOptical flow can be leveraged in robotic systems for obstacle detection where low latency solutions are critical in highly dynamic settings. While event-based cameras have changed the dominant paradigm of sending by encoding stimuli into spike trails, offering low bandwidth and latency, events are still processed with traditional convolutional networks in GPUs defeating, thus, the promise of efficient low capacity low power processing that inspired the design of event sensors. In this work, we introduce a shallow spiking neural network for the computation of optical flow consisting of Leaky Integrate and Fire neurons.Optical flow is predicted as the synthesis of motion orientation selective channels. Learning is accomplished by Back-propapagation Through Time. We present promising results on events recorded in real "in the wild" scenes that has the capability to use only a small fraction of the energy consumed in CNNs deployed on GPUs. Kenneth Chaney, Artemis Panagopoulou, Chankyu Lee, Kaushik Roy 0001, Kostas Daniilidis |
IROS | 2 |