EDBT 2026 Demo / reviewers in the wild / expert
Hongyi Zhu 0004
dblp:147/8584-4
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0006-0298-0905ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A-MAR: Agent-based Multimodal Art Retrieval for Fine-Grained Artwork UnderstandingabstractUnderstanding artworks requires multi-step reasoning over visual content and cultural, historical, and stylistic context. While recent multimodal large language models show promise in artwork explanation, they rely on implicit reasoning and internalized knowledge, limiting interpretability and explicit evidence grounding. We propose A-MAR, an Agent-based Multimodal Art Retrieval framework that explicitly conditions retrieval on structured reasoning plans. Given an artwork and a user query, A-MAR first decomposes the task into a structured reasoning plan that specifies the goals and evidence requirements for each step. Retrieval is then conditioned on this plan, enabling targeted evidence selection and supporting step-wise, grounded explanations. To evaluate agent-based multimodal reasoning within the art domain, we introduce ArtCoT-QA. This diagnostic benchmark features multi-step reasoning chains for diverse art-related queries, enabling a granular analysis that extends beyond simple final answer accuracy. Experiments on SemArt and Artpedia show that A-MAR consistently outperforms static, non-planned retrieval and strong MLLM baselines in final explanation quality, while evaluations on ArtCoT-QA further demonstrate its advantages in evidence grounding and multi-step reasoning ability. These results highlight the importance of reasoning-conditioned retrieval for knowledge-intensive multimodal understanding and position A-MAR as a step toward interpretable, goal-driven AI systems, with particular relevance to cultural industries. The code and data are available at: https://github.com/ShuaiWang97/A-MAR. Shuai Wang 0054, Hongyi Zhu 0004, Yixian Shen, Chengxi Zeng, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg, Marcel Worring |
ICMR | 2 |
| 2026 | Agent-Based Query Reformulation: Simulating Feedback and Mitigating Negation Blindness in Interactive Image RetrievalabstractInteractive image retrieval overcomes the limitations of single-turn search by allowing users to refine their intent through dialogue. However, developing robust retrieval systems is currently hindered by reliance on pre-generated question-answer pairs that do not capture actual retrieval results during the conversation, preventing the system from dynamically adjusting its strategy to real-time errors. In this paper, we first address this limitation by proposing a Multimodal Conversational Search Simulation framework. This closed-loop environment enables a user simulator to have direct interaction with the retrieval results and generate feedback about the most relevant images. We further propose an agent-based image retrieval system that tracks user preferences in multi-turn interactions and summarizes the historical information into a reformulated query. Leveraging this dynamic environment, we conduct an extensive exploration of query reformulation strategies based on Large Language Models (LLMs) and identify a persistent yet underexplored failure mode in conversational image search: the inability of dense retrievers to process negation and exclusion constraints (e.g., “not red”). We analyze this phenomenon and propose effective mitigation strategies based on zero-shot learning and supervised fine-tuning. Finally, to synthesize these insights into a robust system, we propose a transition from passive query rewriting to Agent-Based Query Reformulation. Unlike traditional methods that merely mimic human conversation, our approach treats the reformulator as a strategic agent optimized for ranking performance. We introduce a novel pipeline that fine-tunes an LLM using Direct Preference Optimization (DPO) on retrieval rewards, effectively enabling the model to learn the specific “dialect” of the search engine. Extensive experiments demonstrate that our agent-based approach significantly outperforms standard baselines, particularly in complex scenarios involving negation and exclusion. Hongyi Zhu 0004, Shuai Wang 0054, Jia-Hong Huang, Yixian Shen, Stevan Rudinac, Evangelos Kanoulas |
ICMR | 1 |
| 2025 | Gradient Weight-normalized Low-rank Projection for Efficient LLM TrainingabstractLarge Language Models (LLMs) have shown remarkable performance across various tasks, but the escalating demands on computational resources pose significant challenges, particularly in the extensive utilization of full fine-tuning for downstream tasks. To address this, parameter-efficient fine-tuning (PEFT) methods have been developed, but they often underperform compared to full fine-tuning and struggle with memory efficiency. In this work, we introduce Gradient Weight-Normalized Low-Rank Projection (GradNormLoRP), a novel approach that enhances both parameter and memory efficiency while maintaining comparable performance to full fine-tuning. GradNormLoRP normalizes the weight matrix to improve gradient conditioning, facilitating better convergence during optimization. Additionally, it applies low-rank approximations to the weight and gradient matrices, significantly reducing memory usage during training. Extensive experiments demonstrate that our 8-bit GradNormLoRP reduces optimizer memory usage by up to 89.5\% and enables the pre-training of large LLMs, such as LLaMA 7B, on consumer-level GPUs like the NVIDIA RTX 4090, without additional inference costs. Moreover, GradNormLoRP outperforms existing low-rank methods in fine-tuning tasks. For instance, when fine-tuning the RoBERTa model on all GLUE tasks with a rank of 8, GradNormLoRP achieves an average score of 80.65, surpassing LoRA's score of 79.23. These results underscore GradNormLoRP as a promising alternative for efficient LLM pre-training and fine-tuning. Jia-Hong Huang, Yixian Shen, Hongyi Zhu 0004, Stevan Rudinac, Evangelos Kanoulas |
AAAI | 3 |
| 2025 | MaCP: Minimal yet Mighty Adaptation via Hierarchical Cosine ProjectionabstractWe present a new adaptation method MaCP, Minimal yet Mighty adaptive Cosine Projection, that achieves exceptional performance while requiring minimal parameters and memory for fine-tuning large foundation models.Its general idea is to exploit the superior energy compaction and decorrelation properties of cosine projection to improve both model efficiency and accuracy.Specifically, it projects the weight change from the low-rank adaptation into the discrete cosine space.Then, the weight change is partitioned over different levels of the discrete cosine spectrum, and each partition's most critical frequency components are selected.Extensive experiments demonstrate the effectiveness of MaCP across a wide range of single-modality tasks, including natural language understanding, natural language generation, text summarization, as well as multimodality tasks such as image classification and video understanding.MaCP consistently delivers superior accuracy, significantly reduced computational complexity, and lower memory requirements compared to existing alternatives. Yixian Shen, Qi Bi, Jia-Hong Huang, Hongyi Zhu 0004, Andy D. Pimentel, Anuj Pathania |
ACL (1) | 4 |
| 2025 | ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art UnderstandingabstractVisual art understanding requires joint modeling of multiple perspectives and contextual inference rooted in cultural, historical, and stylistic knowledge. Recent multimodal large language models (MLLMs) demonstrate strong performance in generic captioning, primarily based on object recognition and training on large-scale generic data. They struggle in providing captions incorporating the multiple perspectives that fine art demands. In this work, we introduce ArtRAG, a novel training-free framework that integrates structured knowledge into a retrieval-augmented generation (RAG) pipeline for multi-perspective artwork explanation. ArtRAG automatically constructs an Art Context Knowledge Graph (ACKG) from domain-specific textual sources, organizing entities such as artists, themes, movements, and historical events into a rich, interpretable knowledge graph. At inference time, a multi-granular structured context retriever selects semantically and topologically relevant subgraphs to guide explanation generation. This approach enables MLLMs to produce contextually grounded, multi-perspective descriptions. Experiments on the SemArt and Artpedia datasets demonstrate that ArtRAG outperforms existing heavily trained baselines. Human evaluations further confirm ArtRAG's ability to generate coherent, informative, and culturally enriched interpretations of artworks. Shuai Wang 0054, Ivona Najdenkoska, Hongyi Zhu 0004, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg, Marcel Worring |
ACM Multimedia | 3 |
| 2025 | Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
Jia-Hong Huang, Hongyi Zhu 0004, Yixian Shen, Stevan Rudinac, Evangelos Kanoulas |
MMM (4) | 2 |
| 2025 | SSH: Sparse Spectrum Adaptation via Discrete Hartley TransformationabstractYixian Shen, Qi Bi, Jia-hong Huang, Hongyi Zhu, Andy D. Pimentel, Anuj Pathania. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yixian Shen, Qi Bi, Jia-Hong Huang, Hongyi Zhu 0004, Andy D. Pimentel, Anuj Pathania |
NAACL (Long Papers) | 4 |
| 2025 | Interactive Image Retrieval Meets Query Rewriting with Large Language and Vision Language ModelsabstractImage search is a pivotal task in multi-media and computer vision, finding applications across diverse domains, ranging from internet search to medical diagnostics. Conventional image search systems operate by accepting textual or visual queries and retrieving the top-relevant candidate results from the database. However, prevalent methods often rely on single-turn procedures, introducing potential inaccuracies and limited recall. These methods also face challenges, such as vocabulary mismatch and the semantic gap, constraining their overall effectiveness. To address these issues, we propose an interactive image retrieval system capable of refining queries based on user relevance feedback in a multi-turn setting. This system incorporates an image captioner based on a vision-language model (VLM) to enhance the quality of text-based queries, resulting in more informative queries with each iteration. Moreover, we introduce a denoiser based on a large language model (LLM) to refine text-based query expansions, mitigating inaccuracies in image descriptions generated by captioning models. To evaluate our system, we curate a new dataset by adapting the MSR-VTT and MSVD video retrieval datasets to the image retrieval task, offering multiple relevant ground-truth images for each query. Through comprehensive experiments, we validate the effectiveness of our proposed system against baseline methods, achieving state-of-the-art performance with a notable 10% improvement in terms of recall. Our contributions encompass the development of an innovative interactive image retrieval system, the integration of an LLM-based denoiser, the curation of a meticulously designed evaluation dataset, and thorough experimental validation. Hongyi Zhu 0004, Jia-Hong Huang, Yixian Shen, Stevan Rudinac, Evangelos Kanoulas |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Enhancing Interactive Image Retrieval With Query Rewriting Using Large Language Models and Vision Language ModelsabstractImage search stands as a pivotal task in multimedia and computer vision, finding applications across diverse domains, ranging from internet search to medical diagnostics. Conventional image search systems operate by accepting textual or visual queries, retrieving the top-relevant candidate results from the database. However, prevalent methods often rely on single-turn procedures, introducing potential inaccuracies and limited recall. These methods also face the challenges, such as vocabulary mismatch and the semantic gap, constraining their overall effectiveness. To address these issues, we propose an interactive image retrieval system capable of refining queries based on user relevance feedback in a multi-turn setting. This system incorporates a vision language model (VLM) based image captioner to enhance the quality of text-based queries, resulting in more informative queries with each iteration. Moreover, we introduce a large language model (LLM) based denoiser to refine text-based query expansions, mitigating inaccuracies in image descriptions generated by captioning models. To evaluate our system, we curate a new dataset by adapting the MSR-VTT video retrieval dataset to the image retrieval task, offering multiple relevant ground truth images for each query. Through comprehensive experiments, we validate the effectiveness of our proposed system against baseline methods, achieving state-of-the-art performance with a notable 10% improvement in terms of recall. Our contributions encompass the development of an innovative interactive image retrieval system, the integration of an LLM-based denoiser, the curation of a meticulously designed evaluation dataset, and thorough experimental validation. Hongyi Zhu 0004, Jia-Hong Huang, Stevan Rudinac, Evangelos Kanoulas |
ICMR | 1 |
| 2024 | Exquisitor at the Video Browser Showdown 2024: Relevance Feedback Meets Conversational Search
Omar Shahbaz Khan, Hongyi Zhu 0004, Ujjwal Sharma 0001, Evangelos Kanoulas, Stevan Rudinac, Björn Þór Jónsson 0001 |
MMM (4) | 2 |