VLDB 2026 Research / reviewers in the wild / expert
Erwin M. Bakker
dblp:46/4861
· DBLP profile ↗
30ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-2624-5271ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 since 2021Computer networks · 2 · 2 first-authorSystems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ProtoSimNet: Exploiting prototype similarity for few-shot learningabstractMany few-shot learning (FSL) approaches exploit Euclidean Distance as the metric, as it provides a stable and well-understood geometric structure for prototype-based classification. Despite these advantages, we observe that Euclidean Distance metric may suffer from a metric-level limitation under few-shot supervision. Meanwhile, with limited supervision, cosine-based scoring often leads to a metric over-smoothing phenomenon, where similarity scores among competing prototypes become overly compressed, producing weak effective margins and less stable decision boundaries, resulting in more ambiguous predictions when distinguishing fine-grained novel categories. In this work, we identify this metric-level limitation and propose a problem-driven prototype similarity framework, termed ProtoSimNet, for FSL which explicitly performs score-level distribution shaping of prototype similarities under data-scarce episodic supervision. ProtoSimNet introduces a cosine sharpening transformation, combined with an elastic regularization, which together enhance discriminability by amplifying the margins between high- and low-confidence prototype similarity scores while regularizing the encoder toward compact and less redundant feature embeddings. Our method strengthens the reliability of cosine-based classification and improves the robustness of prototype matching when only a few labeled samples are available. Experimental results on FSL benchmark datasets show that ProtoSimNet achieves consistent improvements over standard Prototypical Networks and competitive or superior performance compared with representative state-of-the-art methods. Nan Pu, Erwin M. Bakker, Michael S. Lew |
Neurocomputing | 3 |
| 2023 | COCA: COllaborative CAusal Regularization for Audio-Visual Question AnsweringabstractAudio-Visual Question Answering (AVQA) is a sophisticated QA task, which aims at answering textual questions over given video-audio pairs with comprehensive multimodal reasoning. Through detailed causal-graph analyses and careful inspections of their learning processes, we reveal that AVQA models are not only prone to over-exploit prevalent language bias, but also suffer from additional joint-modal biases caused by the shortcut relations between textual-auditory/visual co-occurrences and dominated answers. In this paper, we propose a COllabrative CAusal (COCA) Regularization to remedy this more challenging issue of data biases. Specifically, a novel Bias-centered Causal Regularization (BCR) is proposed to alleviate specific shortcut biases by intervening bias-irrelevant causal effects, and further introspect the predictions of AVQA models in counterfactual and factual scenarios. Based on the fact that the dominated bias impairing model robustness for different samples tends to be different, we introduce a Multi-shortcut Collaborative Debiasing (MCD) to measure how each sample suffers from different biases, and dynamically adjust their debiasing concentration to different shortcut correlations. Extensive experiments demonstrate the effectiveness as well as backbone-agnostic ability of our COCA strategy, and it achieves state-of-the-art performance on the large-scale MUSIC-AVQA dataset. Mingrui Lao, Nan Pu, Yu Liu 0012, Erwin M. Bakker, Michael S. Lew |
AAAI | 5 |
| 2023 | Multi-Domain Lifelong Visual Question Answering via Self-Critical DistillationabstractVisual Question Answering (VQA) has achieved significant success over the last few years, while most studies focus on training a VQA model on a stationary domain (e.g., a given dataset). In real-world application scenarios, however, these methods are often inefficient because VQA systems are always supposed to extend their knowledge and meet the ever-changing demands of users. In this paper, we introduce a new and challenging multi-domain lifelong VQA task, dubbed MDL-VQA, which encourages the VQA model to continuously learn across multiple domains while mitigating the forgetting on previously-learned domains. Furthermore, we propose a novel replay-free Self-Critical Distillation (SCD) framework tailor-made for MDL-VQA, which alleviates forgetting issue via transferring previous-domain knowledge from teacher to student models. First, we propose to introspect the teacher's understanding over original and counterfactual samples, thereby creating informative instance-relevant and domain-relevant knowledge for logits-based distillation. Second, on the side of feature-based distillation, we propose to introspect the reasoning behavior of student model to establish the harmful domain-specific knowledge acquired in current domain, and further leverage the metric learning strategy to encourage student to learn useful knowledge in new domain. Extensive experiments demonstrate that SCD framework outperforms state-of-the-art competitors with different training orders. Mingrui Lao, Nan Pu, Yu Liu 0012, Zhun Zhong, Erwin M. Bakker, Nicu Sebe, Michael S. Lew |
ACM Multimedia | 5 |
| 2023 | Dual selective knowledge transfer for few-shot classificationabstractAbstract Few-shot learning aims at recognizing novel visual categories from very few labelled examples. Different from the existing few-shot classification methods that are mainly based on metric learning or meta-learning, in this work we focus on improving the representation capacity of feature extractors. For this purpose, we propose a new two-stage dual selective knowledge transfer (DSKT) framework, to guide models towards better optimization. Specifically, we first exploit an improved multi-task learning approach to train a feature extractor with robust representation capability as a teacher model. Then, we design an effective dual selective knowledge distillation method, which enables the student model to selectively learn knowledge from the teacher model and current samples, thereby improving the student model’s ability to generalize on unseen classes. Extensive experimental results show that our DSKT achieves competitive performances on four well-known few-shot classification benchmarks. Nan Pu, Mingrui Lao, Erwin M. Bakker, Michael S. Lew |
Appl. Intell. | 4 |
| 2023 | Deep Learning for Instance Retrieval: A SurveyabstractIn recent years a vast amount of visual content has been generated and shared from many fields, such as social media platforms, medical imaging, and robotics. This abundance of content creation and sharing has introduced new challenges, particularly that of searching databases for similar content - Content Based Image Retrieval (CBIR) - a long-established research area in which improved efficiency and accuracy are needed for real-time retrieval. Artificial intelligence has made progress in CBIR and has significantly facilitated the process of instance search. In this survey we review recent instance retrieval works that are developed based on deep learning algorithms and techniques, with the survey organized by deep feature extraction, feature embedding and aggregation methods, and network fine-tuning strategies. Our survey considers a wide variety of recent methods, whereby we identify milestone work, reveal connections among various methods and present the commonly used benchmarks, evaluation results, common challenges, and propose promising future directions. Wei Chen 0072, Yu Liu 0012, Weiping Wang 0002, Erwin M. Bakker, Theodoros Georgiou 0001, Paul W. Fieguth, Li Liu 0002, Michael S. Lew |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Meta Reconciliation Normalization for Lifelong Person Re-IdentificationabstractLifelong person re-identification (LReID) is a challenging and emerging task, which concerns the ReID capability on both seen and unseen domains after learning across different domains continually. Existing works on LReID are devoted to introducing commonly-used lifelong learning approaches, while neglecting a serious side effect caused by using normalization layers in the context of domain-incremental learning. In this work, we aim to raise awareness of the importance of training proper batch normalization layers by proposing a new meta reconciliation normalization (MRN) method specifically designed for tackling LReID. Our MRN consists of grouped mixture standardization and additive rectified rescaling components, which are able to automatically maintain an optimal balance between domain-dependent and domain-independent statistics, and even adapt MRN for different testing instances. Furthermore, inspired by synaptic plasticity in human brain, we present a MRN-based meta-learning framework for mining the meta-knowledge shared across different domains, even without replaying any previous data, and further improve the model's LReID ability with theoretical analyses. Our method achieves new state-of-the-art performances on both balanced and imbalanced LReID benchmarks. Nan Pu, Yu Liu 0012, Wei Chen 0072, Erwin M. Bakker, Michael S. Lew |
ACM Multimedia | 4 |
| 2022 | Deep hashing with self-supervised asymmetric semantic excavation and margin-scalable constraintabstractDue to its effectivity and efficiency, deep hashing approaches are widely used for large-scale visual search. However, it is still challenging to produce compact and discriminative hash codes for images associated with multiple semantics for two main reasons, 1) similarity constraints designed in most of the existing methods are based upon an oversimplified similarity assignment(i.e., 0 for instance pairs sharing no label, 1 for instance pairs sharing at least 1 label), 2) the exploration in multi-semantic relevance are insufficient or even neglected in many of the existing methods. These problems significantly limit the discrimination of generated hash codes. In this paper, we propose a novel Deep Hashing with Self-Supervised Asymmetric Semantic Excavation and Margin-Scalable Constraint(SADH) approach to cope with these problems. SADH implements a self-supervised network to sufficiently preserve semantic information in a semantic feature dictionary and a semantic code dictionary for the semantics of the given dataset, which efficiently and precisely guides a feature learning network to preserve multi-label semantic information using an asymmetric learning strategy. By further exploiting semantic dictionaries, a new margin-scalable constraint is employed for both precise similarity searching and robust hash code generation. Extensive empirical research on four popular benchmarks validates the proposed method and shows it outperforms several state-of-the-art approaches. The source codes URL of our SADH is: http://github.com/SWU-CS-MediaLab/SADH . Song Wu 0003, Zhihao Dou, Erwin M. Bakker |
Neurocomputing | 4 |
| 2022 | Multi-label enhancement based self-supervised deep cross-modal hashing
Xitao Zou, Song Wu 0003, Erwin M. Bakker |
Neurocomputing | 3 |
| 2022 | Multi-label modality enhanced attention based self-supervised deep cross-modal hashing
Xitao Zou, Song Wu 0003, Erwin M. Bakker |
Knowl. Based Syst. | 4 |
| 2021 | Lifelong Person Re-Identification via Adaptive Knowledge AccumulationabstractPerson re-identification (ReID) methods always learn through a stationary domain that is fixed by the choice of a given dataset. In many contexts (e.g., lifelong learning), those methods are ineffective because the domain is continually changing in which case incremental learning over multiple domains is required potentially. In this work we explore a new and challenging ReID task, namely lifelong person re-identification (LReID), which enables to learn continuously across multiple domains and even generalise on new and unseen domains. Following the cognitive processes in the human brain, we design an Adaptive Knowledge Accumulation (AKA) framework that is endowed with two crucial abilities: knowledge representation and knowledge operation. Our method alleviates catastrophic forgetting on seen domains and demonstrates the ability to generalize to unseen domains. Correspondingly, we also provide a new and large-scale benchmark for LReID. Extensive experiments demonstrate our method outperforms other competitors by a margin of 5.8% mAP in generalising evaluation. The codes will be available at https://github.com/TPCD/LifelongReID. Nan Pu, Wei Chen 0072, Yu Liu 0012, Erwin M. Bakker, Michael S. Lew |
CVPR | 4 |
| 2021 | Integrating information theory and adversarial learning for cross-modal retrievalabstractAccurately matching visual and textual data in cross-modal retrieval has been widely studied in the multimedia community. To address these challenges posited by the heterogeneity gap and the semantic gap, we propose integrating Shannon information theory and adversarial learning. In terms of the heterogeneity gap, we integrate modality classification and information entropy maximization adversarially. For this purpose, a modality classifier (as a discriminator) is built to distinguish the text and image modalities according to their different statistical properties. This discriminator uses its output probabilities to compute Shannon information entropy, which measures the uncertainty of the modality classification it performs. Moreover, feature encoders (as a generator) project uni-modal features into a commonly shared space and attempt to fool the discriminator by maximizing its output information entropy. Thus, maximizing information entropy gradually reduces the distribution discrepancy of cross-modal features, thereby achieving a domain confusion state where the discriminator cannot classify two modalities confidently. To reduce the semantic gap, Kullback-Leibler (KL) divergence and bi-directional triplet loss are used to associate the intra- and inter-modality similarity between features in the shared space. Furthermore, a regularization term based on KL-divergence with temperature scaling is used to calibrate the biased label classifier caused by the data imbalance issue. Extensive experiments with four deep models on four benchmarks are conducted to demonstrate the effectiveness of the proposed approach. Wei Chen 0072, Yu Liu 0012, Erwin M. Bakker, Michael S. Lew |
Pattern Recognit. | 3 |
| 2021 | Multi-label semantics preserving based deep cross-modal hashing
Xitao Zou, Erwin M. Bakker, Song Wu 0003 |
Signal Process. Image Commun. | 3 |
| 2020 | On the Exploration of Incremental Learning for Fine-grained Image Retrieval
Wei Chen 0072, Yu Liu 0012, Weiping Wang 0002, Tinne Tuytelaars, Erwin M. Bakker, Michael S. Lew |
BMVC | 5 |
| 2020 | Dual Gaussian-based Variational Subspace Disentanglement for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is a challenging and essential task in night-time intelligent surveillance systems. Except for the intra-modality variance that RGB-RGB person re-identification mainly overcomes, VI-ReID suffers from additional inter-modality variance caused by the inherent heterogeneous gap. To solve the problem, we present a carefully designed dual Gaussian-based variational auto-encoder (DG-VAE), which disentangles an identity-discriminable and an identity-ambiguous cross-modality feature subspace, following a mixture-of-Gaussians (MoG) prior and a standard Gaussian distribution prior, respectively. Disentangling cross-modality identity-discriminable features leads to more robust retrieval for VI-ReID. To achieve efficient optimization like conventional VAE, we theoretically derive two variational inference terms for the MoG prior under the supervised setting, which not only restricts the identity-discriminable subspace so that the model explicitly handles the cross-modality intra-identity variance, but also enables the MoG distribution to avoid posterior collapse. Furthermore, we propose a triplet swap reconstruction (TSR) strategy to promote the above disentangling process. Extensive experiments demonstrate that our method outperforms state-of-the-art methods on two VI-ReID datasets. Codes will be available at https://github.com/TPCD/DG-VAE. Nan Pu, Wei Chen 0072, Yu Liu 0012, Erwin M. Bakker, Michael S. Lew |
ACM Multimedia | 4 |
| 2020 | Self-constraining and attention-based hashing network for bit-scalable cross-modal retrieval
Xitao Zou, Erwin M. Bakker, Song Wu 0003 |
Neurocomputing | 3 |
| 2019 | Domain Uncertainty Based On Information Theory for Cross-Modal Hash RetrievalabstractCross-modal hash retrieval has received considerable interest in the area of deep learning. Here hash codes of data of different modalities are learned where pair-wise loss functions control feature similarity in a shared embedding space. In this paper we improve on feature similarity by using Shannon's information entropy with respect to the modality information that is left in learning superior hash codes. We introduce a novel network for predicting the domain from the learned features while the protagonist network uses a loss function based on Shannon's information entropy to learn to maximize the domain uncertainty and therefore the information content. Additionally, according to the number of common labels between each similar image-text pair, we define a multi-level similarity matrix as supervisory information, which constrains all similar pairs with different weights. We show with extensive experiments that our novel approach to domain uncertainty leads to a cross-modal hash retrieval that outperforms the state-of-the-art. Wei Chen 0072, Nan Pu, Yu Liu 0012, Erwin M. Bakker, Michael S. Lew |
ICME | 4 |
| 2019 | Learning a Domain-Invariant Embedding for Unsupervised Person Re-identificationabstractPerson re-identification (Re-ID) aims at matching images of the same person where images are captured by non-overlapping camera views distributed at different locations. To solve this problem, most recent works require a large pre-labeled dataset for training a deep model. These methods are not always suitable for real-world applications, because the latter often lack labeled data. In order to tackle this drawback, we proposed a novel Domain-Invariant Embedding Network (DIEN) to learn a domain-invariant embedding (DIE) feature by introducing a multi-loss joint learning with Recurrent Top- Down Attention (RTDA) mechanism. Due to the improvement in traditional triplet loss, our proposed model can benefit from both source-domain (labeled) data and target-domain (unlabeled) data. Furthermore, the resulting DIE feature not only has improved class discrimination but also robustness to domain shift. We compared our method with recent competitive algorithms and also evaluated the effectiveness of the proposed modules. Nan Pu, Theodoros Georgiou 0001, Erwin M. Bakker, Michael S. Lew |
IJCNN | 3 |
| 2019 | CycleMatch: A cycle-consistent embedding network for image-text matching
Yu Liu 0012, Yanming Guo, Li Liu 0002, Erwin M. Bakker, Michael S. Lew |
Pattern Recognit. | 4 |
| 2018 | CNN-RNN: a large-scale hierarchical image classification frameworkabstractObjects are often organized in a semantic hierarchy of categories, where fine-level categories are grouped into coarse-level categories according to their semantic relations. While previous works usually only classify objects into the leaf categories, we argue that generating hierarchical labels can actually describe how the leaf categories evolved from higher level coarse-grained categories, thus can provide a better understanding of the objects. In this paper, we propose to utilize the CNN-RNN framework to address the hierarchical image classification task. CNN allows us to obtain discriminative features for the input images, and RNN enables us to jointly optimize the classification of coarse and fine labels. This framework can not only generate hierarchical labels for images, but also improve the traditional leaf-level classification performance due to incorporating the hierarchical information. Moreover, this framework can be built on top of any CNN architecture which is primarily designed for leaf-level classification. Accordingly, we build a high performance network based on the CNN-RNN paradigm which outperforms the original CNN (wider-ResNet) and also the current state-of-the-art. In addition, we investigate how to utilize the CNN-RNN framework to improve the fine category classification when a fraction of the training data is only annotated with coarse labels. Experimental results demonstrate that CNN-RNN can use the coarse-labeled training data to improve the classification of fine categories, and in some cases it even surpasses the performance achieved by fully annotated training data. This reveals that, CNN-RNN can alleviate the challenge of specialized and expensive annotation of fine labels. Yanming Guo, Yu Liu 0012, Erwin M. Bakker, Yuanhao Guo, Michael S. Lew |
Multim. Tools Appl. | 3 |
| 2018 | Bag of Surrogate Parts Feature for Visual RecognitionabstractConvolutional neural networks (CNNs) have attracted significant attention in visual recognition. Several recent studies have shown that, in addition to the fully connected layers, the features derived from the convolutional layers of CNNs can also achieve promising performance in image classification tasks. In this paper, we propose a new feature from the convolutional layers, called Bag of Surrogate Parts (BoSP), and its spatial variant, Spatial-BoSP (S-BoSP). The main idea is, we assume the feature maps in the convolutional layers as surrogate parts, and densely sample and assign image regions to these surrogate parts by observing the activation values. Together with BoSP/S-BoSP, we further propose another two schemes to enhance the performance: scale pooling and global-part prediction. Scale pooling aims to handle the objects with different scales and deformations, and global-part prediction combines the predictions of global and part features. By conducting extensive experiments on generic object, fine-grained object and scene datasets, we find the proposed scheme can not only achieve superior performance to the fully connected feature, but also produces competitive or, in some cases, remarkably better performance than the state of the art. Yanming Guo, Yu Liu 0012, Songyang Lao, Erwin M. Bakker, Michael S. Lew |
IEEE Trans. Multim. | 4 |
| 2017 | Learning a Recurrent Residual Fusion Network for Multimodal Matching
Yu Liu 0012, Yanming Guo, Erwin M. Bakker, Michael S. Lew |
ICCV | 3 |
| 2017 | Deep binary codes for large scale image retrieval
Song Wu 0003, Ard Oerlemans, Erwin M. Bakker, Michael S. Lew |
Neurocomputing | 3 |
| 2017 | A comprehensive evaluation of local detectors and descriptors
Song Wu 0003, Ard Oerlemans, Erwin M. Bakker, Michael S. Lew |
Signal Process. Image Commun. | 3 |
| 2013 | An evaluation of content-based duplicate image detection methods for web searchabstractThe world wide web is filled with billions of images and duplicates of images can frequently be found on many websites. These duplicates can be exact copies or differ slightly in their visual content. In this paper we provide a comparative study on how well content-based duplicate image detection methods are able to detect the duplicates of a query image. We conduct a survey to better understand in which ways such images on the internet differ from each other and use these observations to form a realistic and challenging duplicate image detection scenario. The methods we evaluate in our study are representative techniques from the research literature. In our evaluation, we target the performance of each method in relation to their descriptor size, description time and matching time, to assess their feasibility of application to large image collections (> 1 million). Bart Thomee, Mark J. Huiskes, Erwin M. Bakker, Michael S. Lew |
ICME | 3 |
| 2010 | TOP-SURF: a visual words toolkitabstractTOP-SURF is an image descriptor that combines interest points with visual words, resulting in a high performance yet compact descriptor that is designed with a wide range of content-based image retrieval applications in mind. TOP-SURF offers the flexibility to vary descriptor size and supports very fast image matching. In addition to the source code for the visual word extraction and comparisons, we also provide a high level API and very large pre-computed codebooks targeting web image content for both research and teaching purposes. Bart Thomee, Erwin M. Bakker, Michael S. Lew |
ACM Multimedia | 2 |
| 2009 | Deep exploration for experiential image retrievalabstractExperiential image retrieval systems aim to provide the user with a natural and intuitive search experience. The goal is to empower the user to navigate large collections based on his own needs and preferences, while simultaneously providing him with an accurate sense of what the database has to offer. In this paper we integrate a new browsing mechanism called deep exploration with the proven technique of retrieval by relevance feedback. In our approach, relevance feedback focuses the search on relevant regions, while deep exploration facilitates transparent navigation to promising regions of feature space that would normally remain unreachable. Optimal feature weights are determined automatically based on the evidential support for the relevance of each single feature. To achieve efficient refinement of the search space, images are ranked and presented to the user based on their likelihood of being useful for further exploration. Bart Thomee, Mark J. Huiskes, Erwin M. Bakker, Michael S. Lew |
ACM Multimedia | 3 |
| 2008 | Using an artificial imagination for texture retrievalabstractOur goal is to determine if artificially imagined or synthesized images can be beneficial to interactive visual search. We present a novel approach for using artificially imagined images in relevance feedback. Since the search engine constructs the synthetic images itself, any feedback given by the user on these images allows it to obtain a better understanding of what the user is looking for than it would from feedback on database images alone. We evaluated and compared our image synthesis approach with a normal Rocchio-based system on a well-known texture database with real users. Bart Thomee, Mark J. Huiskes, Erwin M. Bakker, Michael S. Lew |
ICPR | 3 |
| 2007 | The Automatic Transformation of Linked List Data Structures
Sven Groot, Harmen L. A. van der Spek, Erwin M. Bakker, Harry A. G. Wijshoff |
PACT | 3 |
| 1993 | Prefix Routing Schemes in Dynamic Networks
Erwin M. Bakker, Jan van Leeuwen, Richard B. Tan |
Comput. Networks ISDN Syst. | 1 |
| 1993 | Uniform d-emulations of rings, with an application to distributed virtual ring constructionabstractAbstract Emulations are a special kind of structure‐preserving mappings between processor interconnection networks (graphs). In this paper, the notion of emulation is generalized to the notion of d‐emulation for any d ≧ 1. Several problems concerning the complexity of (uniform) d‐emulations are studied. It is shown that the problem of deciding for arbitrary graphs H and G whether H can be uniformly d‐emulated on G is NP‐complete for every fixed d ϵ N. Also, it is shown that the problem of deciding for arbitrary rings R and planar graphs G whether R can be uniformly 2‐emulated on G is NP‐complete. Further, a constructive proof is given of the fact that there always exists a uniform 3‐emulation of a ring R on a graph G = (V, E) if |VRy| = k. |VG| and k ϵ N. This uniform 3‐emulation of R on G is used as the basis for the Distributed Virtual Ring Construction Algorithm that only uses 2|E| messages of O(1) bits and has time complexity ≦ 2(|V| − 1). A traversal of the constructed virtual ring costs 2(|V| − 1) messages. © 1993 John Wiley & Sons, Inc. Erwin M. Bakker, Jan van Leeuwen |
Networks | 1 |