Gertjan J. Burghouts

dblp:84/2061 · DBLP profile ↗
← Back
34ranked-venue papers
13as first author
13since 2021 · last 2025
0000-0001-6265-7276ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 8 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Near, far: Patch-ordering enhances vision foundation models' scene understanding
abstract
We introduce NeCo: Patch Neighbor Consistency, a novel self-supervised training loss that enforces patch-level nearest neighbor consistency across a student and teacher model. Compared to contrastive approaches that only yield binary learning signals, i.e. "attract" and "repel", this approach benefits from the more fine-grained learning signal of sorting spatially dense features relative to reference patches. Our method leverages differentiable sorting applied on top of pretrained representations, such as DINOv2-registers to bootstrap the learning signal and further improve upon them. This dense post-pretraining leads to superior performance across various models and datasets, despite requiring only 19 hours on a single GPU. This method generates high-quality dense feature encoders and establishes several new state-of-the-art results such as +2.3 % and +4.2% for non-parametric in-context semantic segmentation on ADE20k and Pascal VOC, +1.6% and +4.8% for linear segmentation evaluations on COCO-Things and -Stuff and improvements in the 3D understanding of multi-view consistency on SPair-71k, by more than 1.5%.
Valentinos Pariza, Mohammadreza Salehi, Gertjan J. Burghouts, Francesco Locatello, Yuki Markus Asano
ICLR3
2025 Information gathering in POMDPs using active inference
Erwin Walraven, Joris Sijs, Gertjan J. Burghouts
Auton. Agents Multi Agent Syst.3
2024 Graph Neural Networks for Learning Equivariant Representations of Neural Networks
abstract
Neural networks that process the parameters of other neural networks find applications in domains as diverse as classifying implicit neural representations, generating neural network weights, and predicting generalization errors. However, existing approaches either overlook the inherent permutation symmetry in the neural network or rely on intricate weight-sharing patterns to achieve equivariance, while ignoring the impact of the network architecture itself. In this work, we propose to represent neural networks as computational graphs of parameters, which allows us to harness powerful graph neural networks and transformers that preserve permutation symmetry. Consequently, our approach enables a single model to encode neural computational graphs with diverse architectures. We showcase the effectiveness of our method on a wide range of tasks, including classification and editing of implicit neural representations, predicting generalization performance, and learning to optimize, while consistently outperforming state-of-the-art methods. The source code is open-sourced at https://github.com/mkofinas/neural-graphs.
Miltiadis Kofinas, Boris Knyazev 0001, Yunlu Chen, Gertjan J. Burghouts, Efstratios Gavves, Cees Snoek, David W. Zhang
ICLR5
2024 Guided SAM: Label-Efficient Part Segmentation
Sabina van Rooij, Gertjan J. Burghouts
ICPR (29)2
2024 Anticipating Future Object Compositions Without Forgetting
Youssef Zahran, Gertjan J. Burghouts, Yke Bauke Eisma
ICPR (30)2
2024 Towards Probabilistic Inductive Logic Programming with Neurosymbolic Inference and Relaxation
abstract
Abstract Many inductive logic programming (ILP) methods are incapable of learning programs from probabilistic background knowledge, for example, coming from sensory data or neural networks with probabilities. We propose Propper, which handles flawed and probabilistic background knowledge by extending ILP with a combination of neurosymbolic inference, a continuous criterion for hypothesis selection (binary cross-entropy) and a relaxation of the hypothesis constrainer (NoisyCombo). For relational patterns in noisy images, Propper can learn programs from as few as 8 examples. It outperforms binary ILP and statistical models such as a graph neural network.
Fieke Hillerström, Gertjan J. Burghouts
Theory Pract. Log. Program.2
2023 Self-Guided Diffusion Models
abstract
Diffusion models have demonstrated remarkable progress in image generation quality, especially when guidance is used to control the generative process. However, guidance requires a large amount of image-annotation pairs for training and is thus dependent on their availability and correctness. In this paper, we eliminate the need for such annotation by instead exploiting the flexibility of self-supervision signals to design a framework for self-guided diffusion models. By leveraging a feature extraction function and a self-annotation function, our method provides guidance signals at various image granularities: from the level of holistic images to object boxes and even segmentation masks. Our experiments on single-label and multi-label image datasets demonstrate that self-labeled guidance always outperforms diffusion models without guidance and may even surpass guidance based on ground-truth labels. When equipped with self-supervised box or mask proposals, our method further generates visually diverse yet semantically consistent images, without the need for any class, box, or segment label annotation. Self-guided diffusion is simple, flexible and expected to profit from deployment at scale.
Vincent Tao Hu, David W. Zhang, Yuki Markus Asano, Gertjan J. Burghouts, Cees Snoek
CVPR4
2023 Efficient Transfer by Robust Label Selection and Learning with Pseudo-Labels
abstract
Semi-supervised techniques have been successful in reducing the amount of labels needed to train a neural network. Often these techniques focus on making the most out of the given labels and exploiting the unlabeled data. Instead of considering the labels as a given, we start with focusing on how to select good labels for efficient semi-supervised learning with pseudo-labels. We propose CLaP: Clustering, Label selection and Pseudo-labels. It clusters an unlabeled dataset on an embedding, that was pretrained on large scale dataset (ImageNet1k). We use the cluster centers for querying initial labels that are both representative and diverse. We use samples consistent over multiple clustering runs as pseudo-labels. We propose a mixed loss on both the initial labels and the pseudo-labeled data to train a neural network. With CLaP the samples to be annotated are automatically selected, reducing the load by a human annotator to search for these samples. We demonstrate that CLaP outperforms of state-of-the-art methods in few-shot transfer tasks on full datasets by 20% on 1-shot to 3% on 5-shot. It also outperforms the accuracy of the state-of-the-art in the BSCD-FSL benchmark up to 23%, depending on the dataset and amount of labels.
Wyke Huizinga, Maarten Kruithof, Gertjan J. Burghouts, Klamer Schutte
ICIP3
2023 Unlocking Slot Attention by Changing Optimal Transport Costs
abstract
Slot attention is a powerful method for object-centric modeling in images and videos. However, its set-equivariance limits its ability to handle videos with a dynamic number of objects because it cannot break ties. To overcome this limitation, we first establish a connection between slot attention and optimal transport. Based on this new perspective we propose MESH (Minimize Entropy of Sinkhorn): a cross-attention module that combines the tiebreaking properties of unregularized optimal transport with the speed of regularized optimal transport. We evaluate slot attention using MESH on multiple object-centric learning benchmarks and find significant improvements over slot attention in every setting.
David W. Zhang, Simon Lacoste-Julien, Gertjan J. Burghouts, Cees Snoek
ICML4
2022 Multiset-Equivariant Set Prediction with Approximate Implicit Differentiation
David W. Zhang, Simon Lacoste-Julien, Gertjan J. Burghouts, Cees Snoek
ICLR4
2022 Maximum Class Separation as Inductive Bias in One Matrix
abstract
Maximizing the separation between classes constitutes a well-known inductive bias in machine learning and a pillar of many traditional algorithms. By default, deep networks are not equipped with this inductive bias and therefore many alternative solutions have been proposed through differential optimization. Current approaches tend to optimize classification and separation jointly: aligning inputs with class vectors and separating class vectors angularly. This paper proposes a simple alternative: encoding maximum separation as an inductive bias in the network by adding one fixed matrix multiplication before computing the softmax activations. The main observation behind our approach is that separation does not require optimization but can be solved in closed-form prior to training and plugged into a network. We outline a recursive approach to obtain the matrix consisting of maximally separable vectors for any number of classes, which can be added with negligible engineering effort and computational overhead. Despite its simple nature, this one matrix multiplication provides real impact. We show that our proposal directly boosts classification, long-tailed recognition, out-of-distribution detection, and open-set recognition, from CIFAR to ImageNet. We find empirically that maximum separation works best as a fixed bias; making the matrix learnable adds nothing to the performance. The closed-form implementation and code to reproduce the experiments are available on github.
Tejaswi Kasarla, Gertjan J. Burghouts, Max van Spengler, Elise van der Pol, Rita Cucchiara, Pascal Mettes
NeurIPS2
2021 Set Prediction without Imposing Structure as Conditional Density Estimation
David W. Zhang, Gertjan J. Burghouts, Cees Snoek
ICLR2
2021 Independent Prototype Propagation for Zero-Shot Compositionality
abstract
Humans are good at compositional zero-shot reasoning; someone who has never seen a zebra before could nevertheless recognize one when we tell them it looks like a horse with black and white stripes. Machine learning systems, on the other hand, usually leverage spurious correlations in the training data, and while such correlations can help recognize objects in context, they hurt generalization. To be able to deal with underspecified datasets while still leveraging contextual clues during classification, we propose ProtoProp, a novel prototype propagation graph method. First we learn prototypical representations of objects (e.g., zebra) that are independent w.r.t. their attribute labels (e.g., stripes) and vice versa. Next we propagate the independent prototypes through a compositional graph, to learn compositional prototypes of novel attribute-object combinations that reflect the dependencies of the target distribution. The method does not rely on any external data, such as class hierarchy graphs or pretrained word embeddings. We evaluate our approach on AO-Clevr, a synthetic and strongly visual dataset with clean labels, UT-Zappos, a noisy real-world dataset of fine-grained shoe types, and C-GQA, a large-scale object detection dataset modified for compositional zero-shot learning. We show that in the generalized compositional zero-shot setting we outperform state-of-the-art results, and through ablations we show the importance of each part of the method and their contribution to the final results. The code is available on github.
Frank Ruis, Gertjan J. Burghouts, Doina Bucur
NeurIPS2
2016 Recognizing Stress Using Semantics and Modulation of Speech and Gestures
abstract
This paper investigates how speech and gestures convey stress, and how they can be used for automatic stress recognition. As a first step, we look into how humans use speech and gestures to convey stress. In particular, for both speech and gestures, we distinguish between stress conveyed by the intended semantic message (e.g. spoken words for speech, symbolic meaning for gestures), and stress conveyed by the modulation of either speech and gestures (e.g. intonation for speech, speed and rhythm for gestures). As a second step, we use this decomposition of stress as an approach for automatic stress prediction. The considered components provide an intermediate representation with intrinsic meaning, which helps bridging the semantic gap between the low level sensor representation and the high level context sensitive interpretation of behavior. Our experiments are run on an audiovisual dataset with service-desk interactions. The final goal is having a surveillance system that would notify when the stress level is high and extra assistance is needed. We find that speech modulation is the best performing intermediate level variable for automatic stress prediction. Using gestures increases the performance and is mostly beneficial when speech is lacking. The two-stage approach with intermediate variables performs better than baseline feature level or decision level fusion.
Iulia Lefter, Gertjan J. Burghouts, Léon J. M. Rothkrantz
IEEE Trans. Affect. Comput.2
2014 Complex threat detection: Learning vs. rules, using a hierarchy of features
abstract
Theft of cargo from a truck or attacks against the driver are threats hindering the day to day operations of trucking companies. In this work we consider a system, which is using surveillance cameras mounted on the truck to provide an early warning for such evolving threats. Low-level processing involves tracking people and calculating motion features. Intermediate-level processing provides kinematics and localisation, activity descriptions and threat stage estimates. At the high level, we compare threat detection performed with a statistical trained SVM based classifier against a rule based system. Results are promising, and show that the best system depends on the scenario.
Gertjan J. Burghouts, Paulien van Slingerland, R. J. M. ten Hove, Richard den Hollander, Klamer Schutte
AVSS1
2014 A unified approach to the recognition of complex actions from sequences of zone-crossings
Gerard Sanroma, Jose Luis Patino, Gertjan J. Burghouts, Klamer Schutte, James M. Ferryman
Image Vis. Comput.3
2014 Requirements for multimedia metadata schemes in surveillance applications for security
Jeroen van Rest, F. A. Grootjen, Marc Grootjen, Remco Wijn, Olav Aarts, M. L. Roelofs, Gertjan J. Burghouts, Henri Bouma, Lejla Alic, Wessel Kraaij
Multim. Tools Appl.7
2014 Selection of negative samples and two-stage combination of multiple features for action detection in thousands of videos
Gertjan J. Burghouts, Klamer Schutte, Henri Bouma, Richard den Hollander
Mach. Vis. Appl.1
2013 Activity recognition and localization on a truck parking lot
abstract
In this paper we present a set of activity recognition and localization algorithms that together assemble a large amount of information about activities on a parking lot. The aim is to detect and recognize events that may pose a threat to truck drivers and trucks. The algorithms perform zone-based activity learning, individual action recognition and group detection. Visual sensor data, from one camera, have been recorded for 23 realistic scenarios of different complexities. The scene is complicated and causes uncertain and false position estimates. We also present a situational assessment ontology which serves the algorithms with relevant knowledge about the observed scene (e.g. information about objects, vulnerabilities and historical data). The algorithms are tested with real tracking data and the evaluations show promising results. The accuracies are 90 % for zone-based activity learning, 71 % for individual action recognition and 66 % for group detection (i.e. merging of people).
Maria Andersson, Jose Luis Patino, Gertjan J. Burghouts, Adam Flizikowski, Murray Evans, David Gustafsson, Henrik Petersson, Klamer Schutte, James M. Ferryman
AVSS3
2013 Improved action recognition by combining multiple 2D views in the bag-of-words model
abstract
Action recognition is a hard problem due to the many degrees of freedom of the human body and the movement of its limbs. This is especially hard when only one camera viewpoint is available and when actions involve subtle movements. For instance, when looked from the side, checking one's watch may look very similar to crossing one's arms. In this paper, we investigate how much the recognition can be improved when multiple views are available. The novelty is that we explore various combination schemes within the robust and simple bag-of-words (BoW) framework, from early fusion of features to late fusion of multiple classifiers. In new experiments on the publicly available IXMAS dataset, we learn that action recognition can be improved significantly already by only adding one viewpoint. We demonstrate that the state-of-the-art on this dataset can be improved by 5% - achieving 96.4% accuracy - when multiple views are combined. Cross-view invariance of the BoW pipeline can be improved by 32% with intermediate-level fusion.
Gertjan J. Burghouts, Pieter Eendebak, Henri Bouma, Johan-Martijn ten Hove
AVSS1
2013 A Search Engine for Retrieval and Inspection of Events with 48 Human Actions in Realistic Videos
Gertjan J. Burghouts, Leo de Penning, Maarten Kruithof, Patrick Hanckmann, Johan-Martijn ten Hove, S. Landsmeer, S. P. van den Broek, Richard den Hollander, C. van Leeuwen, S. Korzec, Henri Bouma, Klamer Schutte
ICPRAM1
2013 Soft-Assignment Random-forest with an Application to Discriminative Representation of Human Actions in Videos
abstract
The bag-of-features model is a distinctive and robust approach to detect human actions in videos. The discriminative power of this model relies heavily on the quantization of the video features into visual words. The quantization determines how well the visual words describe the human action. Random forests have proven to efficiently transform the features into distinctive visual words. A major disadvantage of the random forest is that it makes binary decisions on the feature values, and thus not taking into account uncertainties of the values. We propose a soft-assignment random forest, which is a generalization of the random forest, by substitution of the binary decisions inside the tree nodes by a sigmoid function. The slope of the sigmoid models the degree of uncertainty about a feature's value. The results demonstrate that the soft-assignment random forest improves significantly the action detection accuracy compared to the original random forest. The human actions that are hard to detect — because they involve interactions with or manipulations of some (typically small) item — are structurally improved. Most prominent improvements are reported for a person handing, throwing, dropping, hauling, taking, closing or opening some item. Improvements are achieved for the state-of-the-art on the IXMAS and UT-Interaction datasets by using the soft-assignment random forest.
Gertjan J. Burghouts
Int. J. Pattern Recognit. Artif. Intell.1
2013 Spatio-temporal layout of human actions for improved bag-of-words action detection
Gertjan J. Burghouts, Klamer Schutte
Pattern Recognit. Lett.1
2013 A comparative study on automatic audio-visual fusion for aggression detection using meta-information
Iulia Lefter, Léon J. M. Rothkrantz, Gertjan J. Burghouts
Pattern Recognit. Lett.3
2012 Automatic Audio-Visual Fusion for Aggression Detection Using Meta-information
abstract
We propose a new method for audio-visual sensor fusion and apply it to automatic aggression detection. While a variety of definitions of aggression exist, in this paper we see it as any kind of behavior that has a disturbing effect on others. We have collected multi- and unimodal assessments by humans, who have given aggression scores on a 3 point scale. There are no trivial fusion algorithms to predict the multimodal labels from the unimodal labels. We propose an intermediate step to discover the structure in the fusion process. We call these meta-features and we find a set of five which have an impact on the fusion process. We use simple state of the art low level audio and video features to predict the level of aggression in audio and video, and we also predict the three most feasible meta-features. We show the significant positive impact of adding the meta-features on predicting the multimodal label as compared to standard fusion techniques like feature and decision level fusion.
Iulia Lefter, Gertjan J. Burghouts, Léon J. M. Rothkrantz
AVSS2
2012 Learning the fusion of audio and video aggression assessment by meta-information from human annotations
Iulia Lefter, Gertjan J. Burghouts, Léon J. M. Rothkrantz
FUSION2
2012 Correlations between 48 human actions improve their detection
Gertjan J. Burghouts, Klamer Schutte
ICPR1
2011 Reasoning About Threats: From Observables to Situation Assessment
abstract
We propose a mechanism to assess threats that are based on observables. Observables are properties of persons, i.e., their behavior and interaction with other persons and objects. We consider observables that can be extracted from sensor signals and intelligence. In this paper, we discuss situation assessment that is based on observables for threat assessment. In the experiments, the assessment is evaluated for scenarios that are relevant to antiterrorism and crowd control. The experiments are performed within an evaluation framework, where the setup is such that conclusions can be drawn concerning: 1) the accuracy and robustness of an architecture to assess situations with respect to threats; and 2) the architecture's dependence of the underlying observables in terms of their false positive and negative rates. One of the interesting conclusions is that discriminative assessment of threatening situations can be achieved by combining generic observables. Situations can be assessed with a precision of 90% at a false positive and negative rate of 15% using only eight learning examples. In a real-world experiment at a large train station, we have classified various types of crowd dynamics. Using simple video features of shape and motion, we have proposed a scheme to translate such features into observables that can be classified by a conditional random field (CRF). The implemented CRF shows to classify successfully the crowd dynamics up to 80 % accuracy.
Gertjan J. Burghouts, Jan-Willem Marck
IEEE Trans. Syst. Man Cybern. Part C1
2009 Performance evaluation of local colour invariants
Gertjan J. Burghouts, Jan-Mark Geusebroek
Comput. Vis. Image Underst.1
2009 Material-specific adaptation of color invariant features
Gertjan J. Burghouts, Jan-Mark Geusebroek
Pattern Recognit. Lett.1
2007 The Distribution Family of Similarity Distances
abstract
Assessing similarity between features is a key step in object recognition and scene categorization tasks. We argue that knowledge on the distribution of distances generated by similarity functions is crucial in deciding whether features are similar or not. Intuitively one would expect that similarities between features could arise from any distribution. In this paper, we will derive the contrary, and report the theoretical result that $L_p$-norms --a class of commonly applied distance metrics-- from one feature vector to other vectors are Weibull-distributed if the feature values are correlated and non-identically distributed. Besides these assumptions being realistic for images, we experimentally show them to hold for various popular feature extraction algorithms, for a diverse range of images. This fundamental insight opens new directions in the assessment of feature similarity, with projected improvements in object and scene recognition algorithms. Erratum: The authors of paper have declared that they have become convinced that the reasoning in the reference is too simple as a proof of their claims. As a consequence, they withdraw their theorems.
Gertjan J. Burghouts, Arnold W. M. Smeulders, Jan-Mark Geusebroek
NIPS1
2006 Color Textons for Texture Recognition
abstract
Texton models have proven to be very discriminative for the recognition of grayvalue images taken from rough textures. To further improve the discriminative power of the distinctive texton models of Varma and Zisserman (VZ model) (IJCV, vol. 62(1), pp. 61-81, 2005), we propose two schemes to exploit color information. First, we incorporate color information directly at the texton level, and apply color invariants to deal with straightforward illumination effects as local intensity, shading and shadow. But, the learning of representatives of the spatial structure and colors of textures may be hampered by the wide variety of apparent structure-color combinations. Therefore, our second contribution is an alternative approach, where we weight grayvalue-based textons with color information in a post-processing step, leaving the originalVZ algorithm intact. We demonstrate that the color-weighted textons outperform the VZ textons as well as the color invariant textons. The color-weighted textons are specifically more discriminative than grayvalue-based textons when the size of the example image set is reduced. When using 2 example images only, recognition performance is 85.6%, which is an improvement over grayvaluebased textons of 10%. Hence, incorporating color in textons facilitates the learning of textons. 1
Gertjan J. Burghouts, Jan-Mark Geusebroek
BMVC1
2006 Quasi-periodic spatiotemporal filtering
abstract
This paper presents the online estimation of temporal frequency to simultaneously detect and identify the quasiperiodic motion of an object. We introduce color to increase discriminative power of a reoccurring object and to provide robustness to appearance changes due to illumination changes. Spatial contextual information is incorporated by considering the object motion at different scales. We combined spatiospectral Gaussian filters and a temporal reparameterized Gabor filter to construct the online temporal frequency filter. We demonstrate the online filter to respond faster and decay faster than offline Gabor filters. Further, we show the online filter to be more selective to the tuned frequency than Gabor filters. We contribute to temporal frequency analysis in that we both identify ("what") and detect ("when") the frequency. In color video, we demonstrate the filter to detect and identify the periodicity of natural motion. The velocity of moving gratings is determined in a real world example. We consider periodic and quasiperiodic motion of both stationary and nonstationary objects.
Gertjan J. Burghouts, Jan-Mark Geusebroek
IEEE Trans. Image Process.1
2005 The Amsterdam Library of Object Images
Jan-Mark Geusebroek, Gertjan J. Burghouts, Arnold W. M. Smeulders
Int. J. Comput. Vis.2