Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ivona Najdenkoska

dblp:297/4696 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0001-6852-0609ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 27% Deep learning architectures and training · 27% Transfer learning and domain adaptation · 20%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%
Databases, data mining, and information retrieval
1 paper
Knowledge graphs · 100%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
art image understanding
0.912025
ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding · ACM Multimedia 2025
Machine learning › Deep learning architectures and training
positional encoding
0.912025
TULIP: Token-length Upgraded CLIP · ICLR 2025
Machine learning › Deep learning architectures and training › positional encoding
relative positional encoding
0.912025
TULIP: Token-length Upgraded CLIP · ICLR 2025
Computer vision › Vision and language
vision-language model
0.912025
TULIP: Token-length Upgraded CLIP · ICLR 2025
Machine learning › Generative modeling
diffusion model
0.812024
Context Diffusion: In-Context Aware Image Generation · ECCV (77) 2024
Natural language and speech › Language models and text generation
in-context generation
0.812024
Context Diffusion: In-Context Aware Image Generation · ECCV (77) 2024
Machine learning › Transfer learning and domain adaptation › few-shot learning
cross-modal few-shot learning
0.712023
Meta Learning to Bridge Vision and Language Models for Multimodal Few-Shot Learning · ICLR 2023
Machine learning › Transfer learning and domain adaptation
meta-learning
0.712023
Meta Learning to Bridge Vision and Language Models for Multimodal Few-Shot Learning · ICLR 2023
Visualization and visual analytics
model visualization
0.512021
GCNIllustrator: Illustrating the Effect of Hyperparameters on Graph Convolutional Networks · ACM Multimedia 2021
Knowledge graphs
knowledge graph construction
0.312025
ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding · ACM Multimedia 2025
Machine learning › Graph learning › graph neural network
graph convolutional network
0.112021
GCNIllustrator: Illustrating the Effect of Hyperparameters on Graph Convolutional Networks · ACM Multimedia 2021

Methods — techniques the papers use, named apart from their topics

retrieval-augmented generation · 1.7multimodal large language model · 1.7visual analytics · 1.0knowledge distillation · 0.9in-context conditioning · 0.8meta-learning · 0.7
YearPublicationVenuePosition
2025 TULIP: Token-length Upgraded CLIP
abstract
We address the challenge of representing long captions in vision-language models, such as CLIP. By design these models are limited by fixed, absolute positional encodings, restricting inputs to a maximum of 77 tokens and hindering performance on tasks requiring longer descriptions. Although recent work has attempted to overcome this limit, their proposed approaches struggle to model token relationships over longer distances and simply extend to a fixed new token length. Instead, we propose a generalizable method, named TULIP, able to upgrade the token length to any length for CLIP-like models. We do so by improving the architecture with relative position encodings, followed by a training procedure that (i) distills the original CLIP text encoder into an encoder with relative position encodings and (ii) enhances the model for aligning longer captions with images. By effectively encoding captions longer than the default 77 tokens, our model outperforms baselines on cross-modal tasks such as retrieval and text-to-image generation. The code repository is available at https://github.com/ivonajdenkoska/tulip.
Ivona Najdenkoska, Mohammad Mahdi Derakhshani, Yuki Markus Asano, Nanne van Noord, Marcel Worring, Cees Snoek
ICLR1
2025 ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding
abstract
Visual art understanding requires joint modeling of multiple perspectives and contextual inference rooted in cultural, historical, and stylistic knowledge. Recent multimodal large language models (MLLMs) demonstrate strong performance in generic captioning, primarily based on object recognition and training on large-scale generic data. They struggle in providing captions incorporating the multiple perspectives that fine art demands. In this work, we introduce ArtRAG, a novel training-free framework that integrates structured knowledge into a retrieval-augmented generation (RAG) pipeline for multi-perspective artwork explanation. ArtRAG automatically constructs an Art Context Knowledge Graph (ACKG) from domain-specific textual sources, organizing entities such as artists, themes, movements, and historical events into a rich, interpretable knowledge graph. At inference time, a multi-granular structured context retriever selects semantically and topologically relevant subgraphs to guide explanation generation. This approach enables MLLMs to produce contextually grounded, multi-perspective descriptions. Experiments on the SemArt and Artpedia datasets demonstrate that ArtRAG outperforms existing heavily trained baselines. Human evaluations further confirm ArtRAG's ability to generate coherent, informative, and culturally enriched interpretations of artworks.
Shuai Wang 0054, Ivona Najdenkoska, Hongyi Zhu 0004, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg, Marcel Worring
ACM Multimedia2
2024 Context Diffusion: In-Context Aware Image Generation
Ivona Najdenkoska, Animesh Sinha, Abhimanyu Dubey, Dhruv Mahajan 0001, Vignesh Ramanathan, Filip Radenovic
ECCV (77)1
2023 Meta Learning to Bridge Vision and Language Models for Multimodal Few-Shot Learning
Ivona Najdenkoska, Xiantong Zhen, Marcel Worring
ICLR1
2023 Open-Ended Medical Visual Question Answering Through Prefix Tuning of Language Models
Tom van Sonsbeek, Mohammad Mahdi Derakhshani, Ivona Najdenkoska, Cees Snoek, Marcel Worring
MICCAI (5)3
2022 LifeLonger: A Benchmark for Continual Disease Classification
Mohammad Mahdi Derakhshani, Ivona Najdenkoska, Tom van Sonsbeek, Xiantong Zhen, Dwarikanath Mahapatra, Marcel Worring, Cees Snoek
MICCAI (2)2
2022 Uncertainty-aware report generation for chest X-rays by variational topic inference
abstract
Automating report generation for medical imaging promises to minimize labor and aid diagnosis in clinical practice. Deep learning algorithms have recently been shown to be capable of captioning natural photos. However, doing a similar thing for medical data, is difficult due to the variety in reports written by different radiologists with fluctuating levels of knowledge and experience. Current methods for automatic report generation tend to merely copy one of the training samples in the created report. To tackle this issue, we propose variational topic inference, a probabilistic approach for automatic chest X-ray report generation. Specifically, we introduce a probabilistic latent variable model where a latent variable defines a single topic. The topics are inferred in a conditional variational inference framework by aligning vision and language modalities in a latent space, with each topic governing the generation of one sentence in the report. We further adopt a visual attention module that enables the model to attend to different locations in the image while generating the descriptions. We conduct extensive experiments on two benchmarks, namely Indiana U. Chest X-rays and MIMIC-CXR. The results demonstrate that our proposed variational topic inference method can generate reports with novel sentence structure, rather than mere copies of reports used in training, while still achieving comparable performance to state-of-the-art methods in terms of standard language generation criteria.
Ivona Najdenkoska, Xiantong Zhen, Marcel Worring, Ling Shao 0001
Medical Image Anal.1
2021 Variational Topic Inference for Chest X-Ray Report Generation
Ivona Najdenkoska, Xiantong Zhen, Marcel Worring, Ling Shao 0001
MICCAI (3)1
2021 GCNIllustrator: Illustrating the Effect of Hyperparameters on Graph Convolutional Networks
abstract
An increasing number of real-world applications are using graph-structured datasets, imposing challenges to existing machine learning algorithms. Graph Convolutional Networks (GCNs) are deep learning models, specifically designed to operate on graphs. One of the most tedious steps in training GCNs is the choice of the hyperparameters, especially since they exhibit unique properties compared to other neural models. Not only machine learning beginners, but also experienced practitioners often have difficulties to properly tune their models. We hypothesize that having a tool that visualizes the effect of hyperparameters choice on the performance can accelerate the model development and improve the understanding of these black-box models. Additionally, observing clusters of certain nodes helps to empirically understand how a given prediction was made due to the feature propagation step of GCNs. Therefore, this demo introduces GCNIllustrator - a web-based visual analytics tool for illustrating the effect of hyperparameters on the predictions in a citations graph.
Ivona Najdenkoska, Jeroen den Boef, Justo van der Werf, Reinier de Ridder, Fajar Fathurrahman, Marcel Worring
ACM Multimedia1