Shuai Wang 0054

dblp:42/1503-54 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-1595-3619ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Understanding Art & Culture
abstract
Art and cultural heritage objects carry visual, textual, relational, and symbolic meaning that cannot be reduced to standard image understanding tasks. This tutorial presents computational methods for studying fine art and cultural artifacts, organized around three themes: relationality, meaning, and recognizability. We cover multimodal and graph-based representation learning for fine art analysis, knowledge-retrieval and agentic reasoning frameworks for artwork interpretation, and instance-level recognition in cultural heritage settings, as well as large-scale museum benchmarks and synthetic data generation strategies. Beyond these technical contributions, the tutorial foregrounds the cultural dimension of multimedia research, inviting discussion on current approaches for studying art and culture and directions for future research.
Piera Riccio, Selina Khan, Ludovica Schaerf, Shuai Wang 0054, Athanasios Efthymiou, Noa Garcia, Nanne van Noord
ICMR4
2026 A-MAR: Agent-based Multimodal Art Retrieval for Fine-Grained Artwork Understanding
abstract
Understanding artworks requires multi-step reasoning over visual content and cultural, historical, and stylistic context. While recent multimodal large language models show promise in artwork explanation, they rely on implicit reasoning and internalized knowledge, limiting interpretability and explicit evidence grounding. We propose A-MAR, an Agent-based Multimodal Art Retrieval framework that explicitly conditions retrieval on structured reasoning plans. Given an artwork and a user query, A-MAR first decomposes the task into a structured reasoning plan that specifies the goals and evidence requirements for each step. Retrieval is then conditioned on this plan, enabling targeted evidence selection and supporting step-wise, grounded explanations. To evaluate agent-based multimodal reasoning within the art domain, we introduce ArtCoT-QA. This diagnostic benchmark features multi-step reasoning chains for diverse art-related queries, enabling a granular analysis that extends beyond simple final answer accuracy. Experiments on SemArt and Artpedia show that A-MAR consistently outperforms static, non-planned retrieval and strong MLLM baselines in final explanation quality, while evaluations on ArtCoT-QA further demonstrate its advantages in evidence grounding and multi-step reasoning ability. These results highlight the importance of reasoning-conditioned retrieval for knowledge-intensive multimodal understanding and position A-MAR as a step toward interpretable, goal-driven AI systems, with particular relevance to cultural industries. The code and data are available at: https://github.com/ShuaiWang97/A-MAR.
Shuai Wang 0054, Hongyi Zhu 0004, Yixian Shen, Chengxi Zeng, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg, Marcel Worring
ICMR1
2026 SAM3-LiteText: An Anatomical Study of the SAM3 Text Encoder for Efficient Vision-Language Segmentation
abstract
Vision-language segmentation models such as SAM3 enable flexible, prompt-driven visual grounding, but inherit large, general-purpose text encoders originally designed for open-ended language understanding. In practice, segmentation prompts are short, structured, and semantically constrained, leading to substantial over-provisioning in text encoder capacity and persistent computational and memory overhead. In this paper, we perform a large-scale anatomical analysis of text prompting in vision–language segmentation, covering 404,796 real prompts across multiple benchmarks. Our analysis reveals severe redundancy: most context windows are underutilized, vocabulary usage is highly sparse, and text embeddings lie on a low-dimensional manifold despite high-dimensional representations. Motivated by these findings, we propose SAM3-LiteText, a lightweight text encoding framework that replaces the original SAM3 text encoder with a compact MobileCLIP student that is optimized by knowledge distillation. Extensive experiments on image and video segmentation benchmarks show that SAM3-LiteText reduces text encoder parameters by up to 88%, substantially reducing static memory footprint, while maintaining segmentation performance comparable to the original model. Code: https://github.com/SimonZeng7108/efficientsam3/tree/sam3_litetext.
Chengxi Zeng, Yuxuan Jiang 0015, Ge Gao 0005, Shuai Wang 0054, Duolikun Danier, Bin Zhu 0006, Stevan Rudinac, David Bull 0001, Fan Zhang 0017
ICMR4
2026 Agent-Based Query Reformulation: Simulating Feedback and Mitigating Negation Blindness in Interactive Image Retrieval
abstract
Interactive image retrieval overcomes the limitations of single-turn search by allowing users to refine their intent through dialogue. However, developing robust retrieval systems is currently hindered by reliance on pre-generated question-answer pairs that do not capture actual retrieval results during the conversation, preventing the system from dynamically adjusting its strategy to real-time errors. In this paper, we first address this limitation by proposing a Multimodal Conversational Search Simulation framework. This closed-loop environment enables a user simulator to have direct interaction with the retrieval results and generate feedback about the most relevant images. We further propose an agent-based image retrieval system that tracks user preferences in multi-turn interactions and summarizes the historical information into a reformulated query. Leveraging this dynamic environment, we conduct an extensive exploration of query reformulation strategies based on Large Language Models (LLMs) and identify a persistent yet underexplored failure mode in conversational image search: the inability of dense retrievers to process negation and exclusion constraints (e.g., “not red”). We analyze this phenomenon and propose effective mitigation strategies based on zero-shot learning and supervised fine-tuning. Finally, to synthesize these insights into a robust system, we propose a transition from passive query rewriting to Agent-Based Query Reformulation. Unlike traditional methods that merely mimic human conversation, our approach treats the reformulator as a strategic agent optimized for ranking performance. We introduce a novel pipeline that fine-tunes an LLM using Direct Preference Optimization (DPO) on retrieval rewards, effectively enabling the model to learn the specific “dialect” of the search engine. Extensive experiments demonstrate that our agent-based approach significantly outperforms standard baselines, particularly in complex scenarios involving negation and exclusion.
Hongyi Zhu 0004, Shuai Wang 0054, Jia-Hong Huang, Yixian Shen, Stevan Rudinac, Evangelos Kanoulas
ICMR2
2025 ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding
abstract
Visual art understanding requires joint modeling of multiple perspectives and contextual inference rooted in cultural, historical, and stylistic knowledge. Recent multimodal large language models (MLLMs) demonstrate strong performance in generic captioning, primarily based on object recognition and training on large-scale generic data. They struggle in providing captions incorporating the multiple perspectives that fine art demands. In this work, we introduce ArtRAG, a novel training-free framework that integrates structured knowledge into a retrieval-augmented generation (RAG) pipeline for multi-perspective artwork explanation. ArtRAG automatically constructs an Art Context Knowledge Graph (ACKG) from domain-specific textual sources, organizing entities such as artists, themes, movements, and historical events into a rich, interpretable knowledge graph. At inference time, a multi-granular structured context retriever selects semantically and topologically relevant subgraphs to guide explanation generation. This approach enables MLLMs to produce contextually grounded, multi-perspective descriptions. Experiments on the SemArt and Artpedia datasets demonstrate that ArtRAG outperforms existing heavily trained baselines. Human evaluations further confirm ArtRAG's ability to generate coherent, informative, and culturally enriched interpretations of artworks.
Shuai Wang 0054, Ivona Najdenkoska, Hongyi Zhu 0004, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg, Marcel Worring
ACM Multimedia1
2024 Prototype-Enhanced Hypergraph Learning for Heterogeneous Information Networks
Shuai Wang 0054, Athanasios Efthymiou, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg, Marcel Worring
MMM (3)1
2023 Towards Open-Vocabulary Video Instance Segmentation
abstract
Video Instance Segmentation (VIS) aims at segmenting and categorizing objects in videos from a closed set of training categories, lacking the generalization ability to handle novel categories in real-world videos. To address this limitation, we make the following three contributions. First, we introduce the novel task of Open-Vocabulary Video Instance Segmentation, which aims to simultaneously segment, track, and classify objects in videos from open-set categories, including novel categories unseen during training. Second, to benchmark Open-Vocabulary VIS, we collect a Large-Vocabulary Video Instance Segmentation dataset (LV-VIS), that contains well-annotated objects from 1,196 diverse categories, significantly surpassing the category size of existing datasets by more than one order of magnitude. Third, we propose an efficient Memory-Induced Transformer architecture, OV2Seg, to first achieve Open-Vocabulary VIS in an end-to-end manner with near real-time inference speed. Extensive experiments on LV-VIS and four existing VIS datasets demonstrate the strong zero-shot generalization ability of OV2Seg on novel categories. The dataset and code are released here https://github.com/haochenheheda/LVVIS.
Xu Tang 0007, Yao Hu 0002, Cilin Yan, Weidi Xie, Shuai Wang 0054, Efstratios Gavves
ICCV7