VLDB 2026 Research / reviewers in the wild / expert
Nikolaos Spanos
dblp:381/3860
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0001-2691-0956ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Trustworthy machine learning · 46% Generative modeling · 30% Vision and language · 20% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › interpretability
counterfactual explanation |
0.9 | 1 | 2025 | V-CECE: Visual Counterfactual Explanations via Conceptual Edits · NeurIPS 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | V-CECE: Visual Counterfactual Explanations via Conceptual Edits · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
image editing |
0.9 | 1 | 2025 | V-CECE: Visual Counterfactual Explanations via Conceptual Edits · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | V-CECE: Visual Counterfactual Explanations via Conceptual Edits · NeurIPS 2025 |
Computer vision › Vision and language
multimodal reasoning |
0.9 | 1 | 2025 | Analyze-Prompt-Reason: A Collaborative Agent-Based Framework for Multi-Image Vision-Language Reasoning · ACM Multimedia 2025 |
Machine learning › Trustworthy machine learning › interpretability › counterfactual explanation
visual counterfactual explanation |
0.9 | 1 | 2025 | V-CECE: Visual Counterfactual Explanations via Conceptual Edits · NeurIPS 2025 |
Computer vision › Image recognition and object detection
image classification |
0.3 | 1 | 2025 | V-CECE: Visual Counterfactual Explanations via Conceptual Edits · NeurIPS 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.3 | 1 | 2025 | Analyze-Prompt-Reason: A Collaborative Agent-Based Framework for Multi-Image Vision-Language Reasoning · ACM Multimedia 2025 |
Methods — techniques the papers use, named apart from their topics
vision transformer · 0.9prompt engineering · 0.9large vision-language model · 0.9diffusion model · 0.9agent-based framework · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Analyze-Prompt-Reason: A Collaborative Agent-Based Framework for Multi-Image Vision-Language ReasoningabstractWe present a Collaborative Agent-Based Framework for Multi-Image Reasoning. Our approach tackles the challenge of interleaved multimodal reasoning across diverse datasets and task formats by employing a dual-agent system: a language-based PromptEngineer, which generates context-aware, task-specific prompts, and a VisionReasoner, a large vision-language model (LVLM) responsible for final inference. The framework is fully automated, modular, and training-free, enabling generalization across classification, question answering, and free-form generation tasks involving one or multiple input images. We evaluate our method on 18 diverse datasets from the 2025 MIRAGE Challenge (Track A), covering a broad spectrum of visual reasoning tasks including document QA, visual comparison, dialogue-based understanding, and scene-level inference. Our results demonstrate that LVLMs can effectively reason over multiple images when guided by informative prompts. Notably, Claude 3.7 achieves near-ceiling performance on challenging tasks such as TQA (99.13% accuracy), DocVQA (96.87%), and MMCoQA (75.28 ROUGE-L). We also explore how design choices-such as model selection, shot count, and input length-influence the reasoning performance of different LVLMs. Angelos Vlachos, Giorgos Filandrianos, Maria Lymperaiou, Nikolaos Spanos, Ilias Mitsouras, Vasileios Karampinis, Athanasios Voulodimos |
ACM Multimedia | 4 |
| 2025 | V-CECE: Visual Counterfactual Explanations via Conceptual EditsabstractRecent black-box counterfactual generation frameworks fail to take into account the semantic content of the proposed edits, while relying heavily on training to guide the generation process. We propose a novel, plug-and-play black-box counterfactual generation framework, which suggests step-by-step edits based on theoretical guarantees of optimal edits to produce human-level counterfactual explanations with zero training. Our framework utilizes a pre-trained image editing diffusion model, and operates without access to the internals of the classifier, leading to an explainable counterfactual generation process. Throughout our experimentation, we showcase the explanatory gap between human reasoning and neural model behavior by utilizing both Convolutional Neural Network (CNN), Vision Transformer (ViT) and Large Vision Language Model (LVLM) classifiers, substantiated through a comprehensive human evaluation. Nikolaos Spanos, Maria Lymperaiou, Giorgos Filandrianos, Konstantinos Thomas, Athanasios Voulodimos, Giorgos B. Stamou |
NeurIPS | 1 |