Nanne van Noord

dblp:123/5104 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0002-5145-3603ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021
YearPublicationVenuePosition
2026 Understanding Art & Culture
abstract
Art and cultural heritage objects carry visual, textual, relational, and symbolic meaning that cannot be reduced to standard image understanding tasks. This tutorial presents computational methods for studying fine art and cultural artifacts, organized around three themes: relationality, meaning, and recognizability. We cover multimodal and graph-based representation learning for fine art analysis, knowledge-retrieval and agentic reasoning frameworks for artwork interpretation, and instance-level recognition in cultural heritage settings, as well as large-scale museum benchmarks and synthetic data generation strategies. Beyond these technical contributions, the tutorial foregrounds the cultural dimension of multimedia research, inviting discussion on current approaches for studying art and culture and directions for future research.
Piera Riccio, Selina Khan, Ludovica Schaerf, Shuai Wang 0054, Athanasios Efthymiou, Noa Garcia, Nanne van Noord
ICMR8
2025 EMPLACE: Self-Supervised Urban Scene Change Detection
abstract
Urban change is a constant process that influences the perception of neighbourhoods and the lives of the people within them. The field of Urban Scene Change Detection (USCD) aims to capture changes in street scenes using computer vision and can help raise awareness of changes that make it possible to better understand the city and its residents. Traditionally, the field of USCD has used supervised methods with small scale datasets. This constrains methods when applied to new cities, as it requires labour-intensive labeling processes and forces a priori definitions of relevant change. In this paper we introduce AC-1M the largest USCD dataset by far of over 1.1M images, together with EMPLACE, a self-supervising method to train a Vision Transformer using our adaptive triplet loss. We show EMPLACE outperforms SOTA methods both as a pre-training method for linear fine-tuning as well as a zero-shot setting. Lastly, in a case study of Amsterdam, we show that we are able to detect both small and large changes throughout the city and that changes uncovered by EMPLACE, depending on size, correlate with housing prices - which in turn is indicative of inequity.
Tim Alpherts, Sennay Ghebreab, Nanne van Noord
AAAI3
2025 TULIP: Token-length Upgraded CLIP
abstract
We address the challenge of representing long captions in vision-language models, such as CLIP. By design these models are limited by fixed, absolute positional encodings, restricting inputs to a maximum of 77 tokens and hindering performance on tasks requiring longer descriptions. Although recent work has attempted to overcome this limit, their proposed approaches struggle to model token relationships over longer distances and simply extend to a fixed new token length. Instead, we propose a generalizable method, named TULIP, able to upgrade the token length to any length for CLIP-like models. We do so by improving the architecture with relative position encodings, followed by a training procedure that (i) distills the original CLIP text encoder into an encoder with relative position encodings and (ii) enhances the model for aligning longer captions with images. By effectively encoding captions longer than the default 77 tokens, our model outperforms baselines on cross-modal tasks such as retrieval and text-to-image generation. The code repository is available at https://github.com/ivonajdenkoska/tulip.
Ivona Najdenkoska, Mohammad Mahdi Derakhshani, Yuki Markus Asano, Nanne van Noord, Marcel Worring, Cees Snoek
ICLR4
2024 Drawing Insights: Sequential Representation Learning in Comics
Sam Titarsolej, Neil Cohn, Nanne van Noord
BMVC3
2024 Find the Cliffhanger: Multi-modal Trailerness in Soap Operas
Carlo Bretti, Pascal Mettes, Hendrik Vincent Koops, Daan Odijk, Nanne van Noord
MMM (2)5
2024 GO4Align: Group Optimization for Multi-Task Alignment
abstract
This paper proposes **GO4Align**, a multi-task optimization approach that tackles task imbalance by explicitly aligning the optimization across tasks. To achieve this, we design an adaptive group risk minimization strategy, comprising two techniques in implementation: (i) dynamical group assignment, which clusters similar tasks based on task interactions; (ii) risk-guided group indicators, which exploit consistent task correlations with risk information from previous iterations. Comprehensive experimental results on diverse benchmarks demonstrate our method's performance superiority with even lower computational costs.
Qi Wang 0009, Zehao Xiao, Nanne van Noord, Marcel Worring
NeurIPS4
2023 Cross-modal Scalable Hyperbolic Hierarchical Clustering
abstract
Hierarchical clustering is a natural approach to discover ontologies from data. Yet, existing approaches are hampered by their inability to scale to large datasets and the discrete encoding of the hierarchy. We introduce scalable Hyperbolic Hierarchical Clustering (sHHC) which overcomes these limitations by learning continuous hierarchies in hyperbolic space. Our hierarchical clustering is of high quality and can be obtained in a fraction of the runtime.Additionally, we demonstrate the strength of sHHC on a downstream cross-modal self-supervision task. By using the discovered hierarchies from sound and vision to construct continuous hierarchical pseudo-labels we can efficiently optimize a network for activity recognition and obtain competitive performance compared to recent self-supervised learning models. Our findings demonstrate the strength of Hyperbolic Hierarchical Clustering and its potential for Self-Supervised Learning.
Nanne van Noord
ICCV2
2023 Prototype-based Dataset Comparison
abstract
Dataset summarisation is a fruitful approach to dataset inspection. However, when applied to a single dataset the discovery of visual concepts is restricted to those most prominent. We argue that a comparative approach can expand upon this paradigm to enable richer forms of dataset inspection that go beyond the most prominent concepts.To enable dataset comparison we present a module that learns concept-level prototypes across datasets. We leverage self-supervised learning to discover these prototypes without supervision, and we demonstrate the benefits of our approach in two case-studies. Our findings show that dataset comparison extends dataset inspection and we hope to encourage more works in this direction. Code and usage instructions available at https://github.com/Nanne/ProtoSim
Nanne van Noord
ICCV1
2023 MATTE: Multi-task multi-scale attention
abstract
In this work, we propose a general method for learning task and scale based attention representations in Multi-Task Learning (MTL) for vision. It relies on learning and maintaining cross-task and cross-scale representations of visual information, whose interaction contributes to a symmetrical improvement across the entire task pool. Apart from learning data representations, we additionally optimize for the most beneficial interaction between tasks and their representations at different scales. Our method adds an attention modulated feature as residual information to the processing of each scale stage within the model, including the final layer of task outputs. We empirically show the effectiveness of our method through experiments with current multi-modal and multi-scale architectures on diverse MTL datasets. We evaluate MATTE on high and low level vision MTL problems, against MTL and single task learning (STL) counterparts. For all experiments we report solid performance improvements in both qualitative and quantitative performance.
Gjorgji Strezoski, Nanne van Noord, Marcel Worring
Comput. Vis. Image Underst.2
2022 Hyperbolic Image Segmentation
abstract
For image segmentation, the current standard is to perform pixel-level optimization and inference in Euclidean output embedding spaces through linear hyperplanes. In this work, we show that hyperbolic manifolds provide a valuable alternative for image segmentation and propose a tractable formulation of hierarchical pixel-level classification in hyperbolic space. Hyperbolic Image Segmentation opens up new possibilities and practical benefits for segmentation, such as uncertainty estimation and boundary information for free, zero-label generalization, and increased performance in low-dimensional output embeddings.
Mina Ghadimi Atigh, Julian Schoep, Erman Acar, Nanne van Noord, Pascal Mettes
CVPR4
2022 Extending CLIP for Category-to-Image Retrieval in E-Commerce
Mariya Hendriksen, Maurits J. R. Bleeker, Svitlana Vakulenko, Nanne van Noord, Ernst Kuiper, Maarten de Rijke
ECIR (1)4
2021 Inside Out Visual Place Recognition
Sarah Ibrahimi, Nanne van Noord, Tim Alpherts, Marcel Worring
BMVC2
2021 Stylistic Multi-Task Analysis of Ukiyo-e Woodblock Prints
Selina J. Khan, Nanne van Noord
BMVC2
2021 Automatic Annotations and Enrichments for Audiovisual Archives
abstract
The practical availability of Audiovisual Processing tools to media scholars and heritage institutions remains limited, despite all the technical advancements of recent years. In this article we present the approach chosen in the CLARIAH project to increase this availability, we discuss the challenges encountered, and introduce the technical solutions we are implementing. Through three use cases focused on the enrichment of AV archives, Pose Analysis, and Automatic Speech Recognition, we demonstrate the potential and breadth of using Audiovisual Processing for archives and Digital Humanities research.
Nanne van Noord, Christian Gosvig Olesen, Roeland Ordelman, Julia Noordegraaf
ICAART (1)1
2020 OmniEyes: Analysis and Synthesis of Artistically Painted Eyes
Gjorgji Strezoski, Rogier Knoester, Nanne van Noord, Marcel Worring
MMM (1)3
2019 Many Task Learning With Task Routing
abstract
Typical multi-task learning (MTL) methods rely on architectural adjustments and a large trainable parameter set to jointly optimize over several tasks. However, when the number of tasks increases so do the complexity of the architectural adjustments and resource requirements. In this paper, we introduce a method which applies a conditional feature-wise transformation over the convolutional activations that enables a model to successfully perform a large number of tasks. To distinguish from regular MTL, we introduce Many Task Learning (MaTL) as a special case of MTL where more than 20 tasks are performed by a single model. Our method dubbed Task Routing (TR) is encapsulated in a layer we call the Task Routing Layer (TRL), which applied in an MaTL scenario successfully fits hundreds of classification tasks in one model. We evaluate on 5 datasets and the Visual Decathlon (VD) challenge against strong baselines and state-of-the-art approaches.
Gjorgji Strezoski, Nanne van Noord, Marcel Worring
ICCV2
2019 Learning Task Relatedness in Multi-Task Learning for Images in Context
abstract
Multimedia applications often require concurrent solutions to multiple tasks. These tasks hold clues to each-others solutions, however as these relations can be complex this remains a rarely utilized property. When task relations are explicitly defined based on domain knowledge multi-task learning (MTL) offers such concurrent solutions, while exploiting relatedness between multiple tasks performed over the same dataset. In most cases however, this relatedness is not explicitly defined and the domain expert knowledge that defines it is not available. To address this issue, we introduce Selective Sharing, a method that learns the inter-task relatedness from secondary latent features while the model trains. Using this insight, we can automatically group tasks and allow them to share knowledge in a mutually beneficial way. We support our method with experiments on 5 datasets in classification, regression, and ranking tasks and compare to strong baselines and state-of-the-art approaches showing a consistent improvement in terms of accuracy and parameter counts. In addition, we perform an activation region analysis showing how Selective Sharing affects the learned representation.
Gjorgji Strezoski, Nanne van Noord, Marcel Worring
ICMR2
2017 Learning scale-variant and scale-invariant features for deep image classification
abstract
Convolutional Neural Networks (CNNs) require large image corpora to be trained on classification tasks. The variation in image resolutions, sizes of objects and patterns depicted, and image scales, hampers CNN training and performance, because the task-relevant information varies over spatial scales. Previous work attempting to deal with such scale variations focused on encouraging scale-invariant CNN representations. However, scale-invariant representations are incomplete representations of images, because images contain scale-variant information as well. This paper addresses the combined development of scale-invariant and scale-variant representations. We propose a multi-scale CNN method to encourage the recognition of both types of features and evaluate it on a challenging image classification task involving task-relevant characteristics at multiple scales. The results show that our multi-scale CNN outperforms single-scale CNN. This leads to the conclusion that encouraging the combined development of a scale-invariant and scale-variant representation in CNNs is beneficial to image recognition performance.
Nanne van Noord, Eric O. Postma
Pattern Recognit.1