VLDB 2026 Research / reviewers in the wild / expert
Joost van de Weijer 0001
dblp:67/3379
· DBLP profile ↗
157ranked-venue papers
18as first author
61since 2021 · last 2026
0000-0002-9656-9706ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 112 · 10 first-author · 48 since 2021Graphics, computer vision, multimedia, augmented reality and games · 98 · 14 first-author · 29 since 2021Databases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Leveraging Semantic Attribute Binding for Free-Lunch Color Control in Diffusion ModelsabstractRecent advances in text-to-image (T2I) diffusion models have enabled remarkable control over various attributes, yet precise color specification remains a fundamental challenge. Existing approaches, such as ColorPeel, rely on model personalization, requiring additional optimization and limiting flexibility in specifying arbitrary colors. In this work, we introduce ColorWave, a novel training-free approach that achieves exact RGB-level color control in diffusion models without fine-tuning. By systematically analyzing the cross-attention mechanisms within IP-Adapter, we uncover an implicit binding between textual color descriptors and reference image features. Leveraging this insight, our method rewires these bindings to enforce precise color attribution while preserving the generative capabilities of pretrained models. Our approach maintains generation quality and diversity, outperforming prior methods in accuracy and applicability across diverse object categories. Through extensive evaluations, we demonstrate that ColorWave establishes a new paradigm for structured, color-consistent diffusion-based image synthesis. Héctor Laria Mantecon, Alexandra Gomez-Villa, Jiang Qin, Muhammad Atif Butt, Bogdan Raducanu, Javier Vazquez-Corral, Joost van de Weijer 0001, Kai Wang 0060 |
WACV | 7 |
| 2026 | StyleDiffusion: Prompt-Embedding Inversion for Text-Based EditingabstractA significant research effort is focused on exploiting the outstanding capacities of pretrained diffusion models for image editing. Approaches either fine tune the model, or invert the image in the latent space of the pretrained model. However, they suffer from two problems: (i) unsatisfactory results in selected regions and unexpected changes in non selected regions, and (ii) the need for careful text prompt editing: the prompt should include all visual objects in the input image. To address this, we propose two improvements: (i) only optimizing the input of the value linear network in the cross-attention layers is sufficiently powerful to reconstruct a real image, and (ii) attention regularization to preserve the object-like attention maps after reconstruction and editing, enabling accurate style editing without causing significant structural change. We further improve the editing technique used for the unconditional branch of classifier-free guidance as used by P2P. Extensive experimental prompt-editing results on a variety of images demonstrate qualitatively and quantitatively that our method has editing capabilities superior to those of existing and concurrent works. Our StyleDiffusion code is available at https://github.com/sen-mao/StyleDiffusion. Senmao Li, Joost van de Weijer 0001, Taihang Hu, Fahad Shahbaz Khan, Qibin Hou, Yaxing Wang, Jian Yang 0003, Ming-Ming Cheng |
Comput. Vis. Media | 2 |
| 2026 | Training-free image inversion for one-step diffusion modelsabstractIn this work, we introduce a novel training-free inversion (TFinv) framework for one-step diffusion models, addressing key challenges in real image inversion and editing. We first identify two critical factors hampering real-image inversion and editing: (1) Initial Latent Editability, which is related to the distance between the initial noise and the ideal Gaussian distribution, and (2) Caption Gap, which means the alignment between text captions and image representations. Both factors influence inversion efficiency and the editability of one-step diffusion models. Then, we propose two novel techniques: iterative noise alignment (iterNA), which minimizes the distribution gap to align with the normal Gaussian distribution, and suffix learning (suffL), which enhances text-to-image caption alignment by introducing learned suffix prompt tokens. These techniques enable precise inversion of input images into their initial noise representations and facilitate image editing. Furthermore, we propose a mask-based editing technique for localized edits while preserving background integrity. Comprehensive experiments on the PIE-Bench dataset validate that our method TFinv not only achieves state-of-the-art performance in one-step diffusion editing, but also significantly outperforms existing multistep approaches in efficiency. Senmao Li, Yaxing Wang, Shiqi Yang 0002, Kai Wang 0060, Joost van de Weijer 0001 |
Pattern Recognit. | 6 |
| 2025 | The Art of Deception: Color Visual Illusions and Diffusion ModelsabstractVisual illusions in humans arise when interpreting out- of-distribution stimuli: if the observer is adapted to certain statistics, perception of outliers deviates from reality. Recent studies have shown that artificial neural networks (ANNs) can also be deceived by visual illusions. This revelation raises profound questions about the nature of visual information. Why are two independent systems, both human brains and ANNs, susceptible to the same illusions? Should any ANN be capable of perceiving visual illusions? Are these perceptions a feature or a flaw? In this work, we study how visual illusions are encoded in diffusion models. Remarkably, we show that they present human-like brightness/color shifts in their latent space. We use this fact to demonstrate that diffusion models can predict visual illusions. Furthermore, we also show how to generate new unseen visual illusions in realistic images using text-to-image diffusion models. We validate this ability through psychophysical experiments that show how our model-generated illusions also fool humans. Alexandra Gomez-Villa, Kai Wang 0060, C. Alejandro Párraga, Bartlomiej Twardowski, Jesús Malo, Javier Vazquez-Corral, Joost van de Weijer 0001 |
CVPR | 7 |
| 2025 | One-Way Ticket: Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion ModelsabstractText-to-Image (T2I) diffusion models have made remarkable advancements in generative modeling; however, they face a trade-off between inference speed and image quality, posing challenges for efficient deployment. Existing distilled T2I models can generate high-fidelity images with fewer sampling steps, but often struggle with diversity and quality, especially in one-step models. From our analysis, we observe redundant computations in the UNet encoders. Our findings suggest that, for T2I diffusion models, decoders are more adept at capturing richer and more explicit semantic information, while encoders can be effectively shared across decoders from diverse time steps. Based on these observations, we introduce the first Time-independent Unified Encoder (TiUE) for the student model UNet architecture, which is a loop-free image generation approach for distilling T2I diffusion models. Using a one-pass scheme, TiUE shares encoder features across multiple decoder time steps, enabling parallel sampling and significantly reducing inference time complexity. In addition, we incorporate a KL divergence term to regularize noise prediction, which enhances the perceptual realism and diversity of the generated images. Experimental results demonstrate that TiUE outperforms state-of-the-art methods, including LCM, SD-Turbo, and SwiftBrushv2, producing more diverse and realistic results while maintaining the computational efficiency. https://github.com/sen-mao/Loopfree Senmao Li, Lei Wang 0118, Kai Wang 0060, Jiehang Xie, Joost van de Weijer 0001, Fahad Shahbaz Khan, Shiqi Yang 0002, Yaxing Wang, Jian Yang 0003 |
CVPR | 6 |
| 2025 | Replay-free Online Continual Learning with Self-Supervised MultiPatchesabstractOnline Continual Learning (OCL) methods train a model on a non-stationary data stream where only a few examples are available at a time, often leveraging replay strategies.However, usage of replay is sometimes forbidden, especially in applications with strict privacy regulations.Therefore, we propose Continual MultiPatches (CMP), an effective plugin for existing OCL self-supervised learning strategies that avoids the use of replay samples.CMP generates multiple patches from a single example and projects them into a shared feature space, where patches coming from the same example are pushed together without collapsing into a single point.CMP surpasses replay and other SSL-based strategies on OCL streams, challenging the role of replay as a go-to solution for self-supervised OCL.Code available at https://github.com/giacomo-cgn/cmp . Giovanni A. Cignoni, Andrea Cossu, Alexandra Gomez-Villa, Joost van de Weijer 0001, Antonio Carta |
ESANN | 4 |
| 2025 | Ask and Remember: A Questions-Only Replay Strategy for Continual Visual Question Answering
Imad Eddine Marouf, Enzo Tartaglione, Stéphane Lathuilière, Joost van de Weijer 0001 |
ICCV | 4 |
| 2025 | InterLCM: Low-Quality Images as Intermediate States of Latent Consistency Models for Effective Blind Face RestorationabstractDiffusion priors have been used for blind face restoration (BFR) by fine-tuning diffusion models (DMs) on restoration datasets to recover low-quality images. However, the naive application of DMs presents several key limitations.
(i) The diffusion prior has inferior semantic consistency (e.g., ID, structure and color.), increasing the difficulty of optimizing the BFR model;
(ii) reliance on hundreds of denoising iterations, preventing the effective cooperation with perceptual losses, which is crucial for faithful restoration.
Observing that the latent consistency model (LCM) learns consistency noise-to-data mappings on the ODE-trajectory and therefore shows more semantic consistency in the subject identity, structural information and color preservation,
we propose $\textit{InterLCM}$ to leverage the LCM for its superior semantic consistency and efficiency to counter the above issues.
Treating low-quality images as the intermediate state of LCM, $\textit{InterLCM}$ achieves a balance between fidelity and quality by starting from earlier LCM steps.
LCM also allows the integration of perceptual loss during training, leading to improved restoration quality, particularly in real-world scenarios.
To mitigate structural and semantic uncertainties, $\textit{InterLCM}$ incorporates a Visual Module to extract visual features and a Spatial Encoder to capture spatial details, enhancing the fidelity of restored images.
Extensive experiments demonstrate that $\textit{InterLCM}$ outperforms existing approaches in both synthetic and real-world datasets while also achieving faster inference speed. Code and models will be publicly available. Senmao Li, Kai Wang 0060, Joost van de Weijer 0001, Fahad Shahbaz Khan, Chunle Guo, Shiqi Yang 0002, Yaxing Wang, Jian Yang 0003, Ming-Ming Cheng |
ICLR | 3 |
| 2025 | One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single PromptabstractText-to-image generation models can create high-quality images from input prompts. However, they struggle to support the consistent generation of identity-preserving requirements for storytelling. Existing approaches to this problem typically require extensive training in large datasets or additional modifications to the original model architectures. This limits their applicability across different domains and diverse diffusion model configurations. In this paper, we first observe the inherent capability of language models, coined $\textit{context consistency}$, to comprehend identity through context with a single prompt. Drawing inspiration from the inherent $\textit{context consistency}$, we propose a novel $\textit{training-free}$ method for consistent text-to-image (T2I) generation, termed "One-Prompt-One-Story" ($\textit{1Prompt1Story}$). Our approach $\textit{1Prompt1Story}$ concatenates all prompts into a single input for T2I diffusion models, initially preserving character identities. We then refine the generation process using two novel techniques: $\textit{Singular-Value
Reweighting}$ and $\textit{Identity-Preserving Cross-Attention}$, ensuring better alignment with the input description for each frame. In our experiments, we compare our method against various existing consistent T2I generation approaches to demonstrate its effectiveness, through quantitative metrics and qualitative assessments. Code is available at https://github.com/byliutao/1Prompt1Story. Kai Wang 0060, Senmao Li, Joost van de Weijer 0001, Fahad Shahbaz Khan, Shiqi Yang 0002, Yaxing Wang, Jian Yang 0003, Ming-Ming Cheng |
ICLR | 4 |
| 2025 | No Task Left Behind: Isotropic Model Merging with Common and Task-Specific SubspacesabstractModel merging integrates the weights of multiple task-specific models into a single multi-task model. Despite recent interest in the problem, a significant performance gap between the combined and single-task models remains. In this paper, we investigate the key characteristics of task matrices -- weight update matrices applied to a pre-trained model -- that enable effective merging. We show that alignment between singular components of task-specific and merged matrices strongly correlates with performance improvement over the pre-trained model. Based on this, we propose an isotropic merging framework that flattens the singular value spectrum of task matrices, enhances alignment, and reduces the performance gap. Additionally, we incorporate both common and task-specific subspaces to further improve alignment and performance. Our proposed approach achieves state-of-the-art performance on vision and language tasks across various sets of tasks and model scales. This work advances the understanding of model merging dynamics, offering an effective methodology to merge models without requiring additional training. Daniel Marczak, Simone Magistri, Sebastian Cygert, Bartlomiej Twardowski, Andrew D. Bagdanov, Joost van de Weijer 0001 |
ICML | 6 |
| 2025 | Improving Continual Learning Performance and Efficiency with Auxiliary ClassifiersabstractContinual learning is crucial for applying machine learning in challenging, dynamic, and often resource-constrained environments. However, catastrophic forgetting — overwriting previously learned knowledge when new information is acquired — remains a major challenge. In this work, we examine the intermediate representations in neural network layers during continual learning and find that such representations are less prone to forgetting, highlighting their potential to accelerate computation. Motivated by these findings, we propose to use auxiliary classifiers (ACs) to enhance performance and demonstrate that integrating ACs into various continual learning methods consistently improves accuracy across diverse evaluation settings, yielding an average 10% relative gain. We also leverage the ACs to reduce the average cost of the inference by 10-60% without compromising accuracy, enabling the model to return the predictions before computing all the layers. Our approach provides a scalable and efficient solution for continual learning. Filip Szatkowski, Yaoyue Zheng, Fei Yang 0004, Tomasz Trzcinski, Bartlomiej Twardowski, Joost van de Weijer 0001 |
ICML | 6 |
| 2025 | Covariances for Free: Exploiting Mean Distributions for Training-free Federated LearningabstractUsing pre-trained models has been found to reduce the effect of data heterogeneity and speed up federated learning algorithms. Recent works have explored training-free methods using first- and second-order statistics to aggregate local client data distributions at the server and achieve high performance without any training. In this work, we propose a training-free method based on an unbiased estimator of class covariance matrices which only uses first-order statistics in the form of class means communicated by clients to the server. We show how these estimated class covariances can be used to initialize the global classifier, thus exploiting the covariances without actually sharing them. We also show that using only within-class covariances results in a better classifier initialization. Our approach improves performance in the range of 4-26% with exactly the same communication cost when compared to methods sharing only class means and achieves performance competitive or superior to methods sharing second-order statistics with dramatically less communication overhead. The proposed method is much more communication-efficient than federated prompt-tuning methods and still outperforms them. Finally, using our method to initialize classifiers and then performing federated fine-tuning or linear probing again yields better performance. Code is available at https://github.com/dipamgoswami/FedCOF. Dipam Goswami, Simone Magistri, Kai Wang 0060, Bartlomiej Twardowski, Andrew D. Bagdanov, Joost van de Weijer 0001 |
NeurIPS | 6 |
| 2025 | Accurate and Efficient Low-Rank Model Merging in Core SpaceabstractIn this paper, we address the challenges associated with merging low-rank adaptations of large neural networks. With the rise of parameter-efficient adaptation techniques, such as Low-Rank Adaptation (LoRA), model fine-tuning has become more accessible. While fine-tuning models with LoRA is highly efficient, existing merging methods often sacrifice this efficiency by merging fully-sized weight matrices. We propose the Core Space merging framework, which enables the merging of LoRA-adapted models within a common alignment basis, thereby preserving the efficiency of low-rank adaptation while substantially improving accuracy across tasks. We further provide a formal proof that projection into Core Space ensures no loss of information and provide a complexity analysis showing the efficiency gains. Extensive empirical results demonstrate that Core Space significantly improves existing merging techniques and achieves state-of-the-art results on both vision and language tasks while utilizing a fraction of the computational resources. Codebase is available at https://github.com/apanariello4/core-space-merging. Aniello Panariello, Daniel Marczak, Simone Magistri, Angelo Porrello, Bartlomiej Twardowski, Andrew D. Bagdanov, Simone Calderara, Joost van de Weijer 0001 |
NeurIPS | 8 |
| 2025 | Free-Lunch Color-Texture Disentanglement for Stylized Image GenerationabstractRecent advances in Text-to-Image (T2I) diffusion models have transformed image generation, enabling significant progress in stylized generation using only a few style reference images. However, current diffusion-based methods struggle with \textit{fine-grained} style customization due to challenges in controlling multiple style attributes, such as color and texture. This paper introduces the first tuning-free approach to achieve free-lunch color-texture disentanglement in stylized T2I generation, addressing the need for independently controlled style elements for the Disentangled Stylized Image Generation (DisIG) problem. Our approach leverages the \textit{Image-Prompt Additivity} property in the CLIP image embedding space to develop techniques for separating and extracting Color-Texture Embeddings (CTE) from individual color and texture reference images. To ensure that the color palette of the generated image aligns closely with the color reference, we apply a whitening and coloring transformation to enhance color consistency. Additionally, to prevent texture loss due to the signal-leak bias inherent in diffusion training, we introduce a noise term that preserves textural fidelity during the Regularized Whitening and Coloring Transformation (RegWCT). Through these methods, our Style Attributes Disentanglement approach (SADis) delivers a more precise and customizable solution for stylized image generation. Experiments on images from the WikiArt and StyleDrop datasets demonstrate that, both qualitatively and quantitatively, SADis surpasses state-of-the-art stylization methods in the DisIG task. Jiang Qin, Alexandra Gomez-Villa, Senmao Li, Shiqi Yang 0002, Yaxing Wang, Kai Wang 0060, Joost van de Weijer 0001 |
NeurIPS | 7 |
| 2025 | Multi-Class Textual-Inversion Secretly Yields a Semantic-Agnostic ClassifierabstractWith the advent of large pre-trained vision-language models such as CLIP, prompt learning methods aim to enhance the transferability of the CLIP model. They learn the prompt given few samples from the downstream task given the specific class names as prior knowledge, which we term as semantic-aware classification. However, in many realistic scenarios, we only have access to few samples and no knowledge of the class names (e.g., when considering instances of classes). This challenging scenario represents the semantic-agnostic discriminative case. Text-to-Image (T2I) personalization methods aim to adapt T2I models to unseen concepts by learning new tokens and endowing these tokens with the capability of generating the learned concepts. These methods do not require knowledge of class names as a semantic-aware prior. Therefore, in this paper, we first explore Textual Inversion and reveal that the new concept tokens possess both generation and classification capabilities by regarding each category as a single concept. However, learning classifiers from single-concept textual inversion is limited since the learned tokens are sub-optimal for the discriminative tasks. To mitigate this issue, we propose Multi-Class textual inversion, which includes a discriminative regularization term for the token updating process. Using this technique, our method MC-TI achieves stronger Semantic-Agnostic Classification while preserving the generation capability of these modifier tokens given only few samples per category. In the experiments, we extensively evaluate MC-TI on 12 datasets covering various scenarios, which demonstrates that MC-TI achieves superior results in terms of both classification and generation outcomes. Kai Wang 0060, Fei Yang 0004, Bogdan Raducanu, Joost van de Weijer 0001 |
WACV | 4 |
| 2025 | CEDL+: Exploiting evidential deep learning for continual out-of-distribution detection
Eduardo Aguilar 0001, Bogdan Raducanu, Petia Radeva, Joost van de Weijer 0001 |
Expert Syst. Appl. | 4 |
| 2025 | Exemplar-Free Continual Learning of Vision Transformers via Gated Class-Attention and Cascaded Feature Drift CompensationabstractAbstract Vision transformers (ViTs) have achieved remarkable successes across a broad range of computer vision applications. As a consequence, there has been increasing interest in extending continual learning theory and techniques to ViT architectures. We propose a new method for exemplar-free class incremental training of ViTs. The main challenge of exemplar-free continual learning is maintaining plasticity of the learner without causing catastrophic forgetting of previously learned tasks. This is often achieved via exemplar replay which can help recalibrate previous task classifiers to the feature drift which occurs when learning new tasks. Exemplar replay, however, comes at the cost of retaining samples from previous tasks which for many applications may not be possible. To address the problem of continual ViT training, we first propose gated class-attention to minimize the drift in the final ViT transformer block. This mask-based gating is applied to class-attention mechanism of the last transformer block and strongly regulates the weights crucial for previous tasks. Importantly, gated class-attention does not require the task-ID during inference, which distinguishes it from other parameter isolation methods. Secondly, we propose a new method of feature drift compensation that accommodates feature drift in the backbone when learning new tasks. The combination of gated class-attention and cascaded feature drift compensation allows for plasticity towards new tasks while limiting forgetting of previous ones. Extensive experiments performed on CIFAR-100, Tiny-ImageNet and ImageNet100 demonstrate that our exemplar-free method obtains competitive results when compared to rehearsal based ViT methods.(Code: https://github.com/OcraM17/GCAB-CFDC ) Marco Cotogni, Fei Yang 0004, Claudio Cusano, Andrew D. Bagdanov, Joost van de Weijer 0001 |
Int. J. Comput. Vis. | 5 |
| 2024 | Resurrecting Old Classes with New Data for Exemplar-Free Continual LearningabstractContinual learning methods are known to suffer from catastrophic forgetting, a phenomenon that is particularly hard to counter for methods that do not store exemplars of previous tasks. Therefore, to reduce potential drift in the feature extractor, existing exemplar-free methods are typically evaluated in settings where the first task is significantly larger than subsequent tasks. Their performance drops drastically in more challenging settings starting with a smaller first task. To address this problem of feature drift estimation for exemplar-free methods, we propose to adversarially perturb the current samples such that their embeddings are close to the old class prototypes in the old model embedding space. We then estimate the drift in the embedding space from the old to the new model using the perturbed images and compensate the prototypes accordingly. We exploit the fact that adversarial samples are transferable from the old to the new feature space in a continual learning setting. The generation of these images is simple and computationally cheap. We demonstrate in our experiments that the proposed approach better tracks the movement of prototypes in embedding space and outperforms existing methods on several standard continual learning benchmarks as well as on fine-grained datasets. Code is available at https://github.com/dipamgoswami/ADC. Dipam Goswami, Albin Soutif-Cormerais, Sandesh Kamath, Bartlomiej Twardowski, Joost van de Weijer 0001 |
CVPR | 6 |
| 2024 | ColorPeel: Color Prompt Learning with Diffusion Models via Color and Shape Disentanglement
Muhammad Atif Butt, Kai Wang 0060, Javier Vazquez-Corral, Joost van de Weijer 0001 |
ECCV (7) | 4 |
| 2024 | Exemplar-Free Continual Representation Learning via Learnable Drift Compensation
Alexandra Gomez-Villa, Dipam Goswami, Kai Wang 0060, Andrew D. Bagdanov, Bartlomiej Twardowski, Joost van de Weijer 0001 |
ECCV (7) | 6 |
| 2024 | Enhancing Perceptual Quality in Video Super-Resolution Through Temporally-Consistent Detail Synthesis Using Diffusion Models
Claudio Rota, Marco Buzzelli, Joost van de Weijer 0001 |
ECCV (12) | 3 |
| 2024 | Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion ModelsabstractThe success of recent text-to-image diffusion models is largely due to their capacity to be guided by a complex text prompt, which enables users to precisely describe the desired content. However, these models struggle to effectively suppress the generation of undesired content, which is explicitly requested to be omitted from the generated image in the prompt. In this paper, we analyze how to manipulate the text embeddings and remove unwanted content from them. We introduce two contributions, which we refer to as soft-weighted regularization and inference-time text embedding optimization. The first regularizes the text embedding matrix and effectively suppresses the undesired content. The second method aims to further suppress the unwanted content generation of the prompt, and encourages the generation of desired content. We evaluate our method quantitatively and qualitatively on extensive experiments, validating its effectiveness. Furthermore, our method is generalizability to both the pixel-space diffusion models (i.e. DeepFloyd-IF) and the latent-space diffusion models (i.e. Stable Diffusion). Senmao Li, Joost van de Weijer 0001, Taihang Hu, Fahad Shahbaz Khan, Qibin Hou, Yaxing Wang, Jian Yang 0003 |
ICLR | 2 |
| 2024 | Elastic Feature Consolidation For Cold Start Exemplar-Free Incremental LearningabstractExemplar-Free Class Incremental Learning (EFCIL) aims to learn from a sequence of tasks without having access to previous task data. In this paper, we consider the challenging Cold Start scenario in which insufficient data is available in the first task to learn a high-quality backbone. This is especially challenging for EFCIL since it requires high plasticity, which results in feature drift which is difficult to compensate for in the exemplar-free setting. To address this problem, we propose a simple and effective approach that consolidates feature representations by regularizing drift in directions highly relevant to previous tasks and employs prototypes to reduce task-recency bias. Our method, called Elastic Feature Consolidation (EFC), exploits a tractable second-order approximation of feature drift based on an Empirical Feature Matrix (EFM). The EFM induces a pseudo-metric in feature space which we use to regularize feature drift in important directions and to update Gaussian prototypes used in a novel asymmetric cross entropy loss which effectively balances prototype rehearsal with data from new tasks. Experimental results on CIFAR-100, Tiny-ImageNet, ImageNet-Subset and ImageNet-1K demonstrate that Elastic Feature Consolidation is better able to learn new tasks by maintaining model plasticity and significantly outperform the state-of-the-art. Simone Magistri, Tomaso Trinci, Albin Soutif-Cormerais, Joost van de Weijer 0001, Andrew D. Bagdanov |
ICLR | 4 |
| 2024 | IterInv: Iterative Inversion for Pixel-Level T2I ModelsabstractLarge-scale text-to-image diffusion models have been a ground-breaking development in generating convincing images following an input text prompt. The goal of image editing research is to give users control over the generated images by modifying the text prompt. Current image editing techniques predominantly hinge on DDIM inversion as a prevalent practice rooted in Latent Diffusion Models (LDM). However, the large pretrained T2I models working on the latent space suffer from losing details due to the first compression stage with an autoencoder mechanism. Instead, other mainstream T2I pipeline working on the pixel level, such as Imagen and DeepFloyd-IF, circumvents the above problem. They are commonly composed of multiple stages, typically starting with a text-to-image stage and followed by several super-resolution stages. In this pipeline, the DDIM inversion fails to find the initial noise and generate the original image given that the super-resolution diffusion models are not compatible with the DDIM technique. According to our experimental findings, iteratively concatenating the noisy image as the condition is the root of this problem. Based on this observation, we develop an iterative inversion (IterInv) technique for this category of T2I models and verify IterInv with the open-source DeepFloyd-IF model. Specifically, IterInv employ NTI as the inversion and reconstruction of low-resolution image generation. In stages 2 and 3, we update the latent variance at each timestep to find the deterministic inversion trace and promote the reconstruction process. By combining our method with a popular image editing method, we prove the application prospects of IterInv. The code will be released upon acceptance. The code is available at https://github.com/Tchuanm/IterInv.git Chuanming Tang, Kai Wang 0060, Joost van de Weijer 0001 |
ICME | 3 |
| 2024 | Token Merging for Training-Free Semantic Binding in Text-to-Image SynthesisabstractAlthough text-to-image (T2I) models exhibit remarkable generation capabilities,
they frequently fail to accurately bind semantically related objects or attributes
in the input prompts; a challenge termed semantic binding. Previous approaches
either involve intensive fine-tuning of the entire T2I model or require users or
large language models to specify generation layouts, adding complexity. In this
paper, we define semantic binding as the task of associating a given object with its
attribute, termed attribute binding, or linking it to other related sub-objects, referred
to as object binding. We introduce a novel method called Token Merging (ToMe),
which enhances semantic binding by aggregating relevant tokens into a single
composite token. This ensures that the object, its attributes and sub-objects all share
the same cross-attention map. Additionally, to address potential confusion among
main objects with complex textual prompts, we propose end token substitution as
a complementary strategy. To further refine our approach in the initial stages of
T2I generation, where layouts are determined, we incorporate two auxiliary losses,
an entropy loss and a semantic binding loss, to iteratively update the composite
token to improve the generation integrity. We conducted extensive experiments to
validate the effectiveness of ToMe, comparing it against various existing methods
on the T2I-CompBench and our proposed GPT-4o object binding benchmark. Our
method is particularly effective in complex scenarios that involve multiple objects
and attributes, which previous methods often fail to address. The code will be
publicly available at https://github.com/hutaihang/ToMe Taihang Hu, Joost van de Weijer 0001, Hongcheng Gao, Fahad Shahbaz Khan, Jian Yang 0003, Ming-Ming Cheng, Kai Wang 0060, Yaxing Wang |
NeurIPS | 3 |
| 2024 | Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model InferenceabstractOne of the main drawback of diffusion models is the slow inference time for image generation. Among the most successful approaches to addressing this problem are distillation methods. However, these methods require considerable computational resources. In this paper, we take another approach to diffusion model acceleration. We conduct a comprehensive study of the UNet encoder and empirically analyze the encoder features. This provides insights regarding their changes during the inference process. In particular, we find that encoder features change minimally, whereas the decoder features exhibit substantial variations across different time-steps. This insight motivates us to omit encoder computation at certain adjacent time-steps and reuse encoder features of previous time-steps as input to the decoder in multiple time-steps. Importantly, this allows us to perform decoder computation in parallel, further accelerating the denoising process. Additionally, we introduce a prior noise injection method to improve the texture details in the generated image. Besides the standard text-to-image task, we also validate our approach on other tasks: text-to-video, personalized generation and reference-guided generation. Without utilizing any knowledge distillation technique, our approach accelerates both the Stable Diffusion (SD) and DeepFloyd-IF model sampling by 41$\%$ and 24$\%$ respectively, and DiT model sampling by 34$\%$, while maintaining high-quality generation performance. Our code will be publicly released. Senmao Li, Taihang Hu, Joost van de Weijer 0001, Fahad Shahbaz Khan, Shiqi Yang 0002, Yaxing Wang, Ming-Ming Cheng, Jian Yang 0003 |
NeurIPS | 3 |
| 2024 | Plasticity-Optimized Complementary Networks for Unsupervised Continual LearningabstractContinuous unsupervised representation learning (CURL) research has greatly benefited from improvements in self-supervised learning (SSL) techniques. As a result, existing CURL methods using SSL can learn high-quality representations without any labels, but with a notable performance drop when learning on a many-tasks data stream. We hypothesize that this is caused by the regularization losses that are imposed to prevent forgetting, leading to a suboptimal plasticity-stability trade-off: they either do not adapt fully to the incoming data (low plasticity), or incur significant forgetting when allowed to fully adapt to a new SSL pretext-task (low stability). In this work, we propose to train an expert network that is relieved of the duty of keeping the previous knowledge and can focus on performing optimally on the new tasks (optimizing plasticity). In the second phase, we combine this new knowledge with the previous network in an adaptation-retrospection phase to avoid forgetting and initialize a new expert with the knowledge of the old network. We perform several experiments showing that our proposed approach outperforms other CURL exemplar-free methods in few- and many-task split settings. Furthermore, we show how to adapt our approach to semi-supervised continual learning (Semi-SCL) and show that we surpass the accuracy of other exemplar-free Semi-SCL methods and reach the results of some others that use exemplars. Alexandra Gomez-Villa, Bartlomiej Twardowski, Kai Wang 0060, Joost van de Weijer 0001 |
WACV | 4 |
| 2024 | MineGAN++: Mining Generative Models for Efficient Knowledge Transfer to Limited Data Domains
Yaxing Wang, Abel Gonzalez-Garcia, Chenshen Wu, Luis Herranz, Fahad Shahbaz Khan, Shangling Jui, Jian Yang 0003, Joost van de Weijer 0001 |
Int. J. Comput. Vis. | 8 |
| 2024 | Projected Latent Distillation for Data-Agnostic Consolidation in distributed continual learningabstractIn continual learning applications on-the-edge multiple self-centered devices (SCD) learn different local tasks independently, with each SCD only optimizing its own task. Can we achieve (almost) zero-cost collaboration between different devices? We formalize this problem as a Distributed Continual Learning (DCL) scenario, where SCDs greedily adapt to their own local tasks and a separate continual learning (CL) model perform a sparse and asynchronous consolidation step that combines the SCD models sequentially into a single multi-task model without using the original data. Unfortunately, current CL methods are not directly applicable to this scenario. We propose Data-Agnostic Consolidation (DAC), a novel double knowledge distillation method which performs distillation in the latent space via a novel Projected Latent Distillation loss. Experimental results show that DAC enables forward transfer between SCDs and reaches state-of-the-art accuracy on Split CIFAR100, CORe50 and Split TinyImageNet, both in single device and distributed CL scenarios. Somewhat surprisingly, a single out-of-distribution image is sufficient as the only source of data for DAC. Antonio Carta, Andrea Cossu, Vincenzo Lomonaco, Davide Bacciu, Joost van de Weijer 0001 |
Neurocomputing | 5 |
| 2024 | Main product detection with graph networks for fashionabstractAbstract Computer vision has established a foothold in the online fashion retail industry. Main product detection is a crucial step of vision-based fashion product feed parsing pipelines, focused on identifying the bounding boxes that contain the product being sold in the gallery of images of the product page. The current state-of-the-art approach does not leverage the relations between regions in the image, and treats images of the same product independently, therefore not fully exploiting visual and product contextual information. In this paper, we propose a model that incorporates Graph Convolutional Networks (GCN) that jointly represent all detected bounding boxes in the gallery as nodes. We show that the proposed method is better than the state-of-the-art, especially, when we consider the scenario where title-input is missing at inference time and for cross-dataset evaluation, our method outperforms previous approaches by a large margin. Vacit Oguz Yazici, Arnau Ramisa, Luis Herranz, Joost van de Weijer 0001 |
Multim. Tools Appl. | 5 |
| 2024 | Conditional Diffusion Model With Spatial-Frequency Refinement for SAR-to-Optical Image TranslationabstractThe presence of speckles and geometric distortions poses a serious challenge to the visual interpretation of synthetic aperture radar (SAR) images. SAR-to-optical (S2O) image translation technology provides a feasible solution and has attracted increasing attention. Restricted by substantial gaps between optical and SAR images, current S2O translation methods unavoidably result in geometric distortions, target missing, and generating low-fidelity images, thereby limiting subsequent cross-modal applications. In this article, we propose an augmented conditional denoising diffusion probabilistic model with spatial-frequency refinement (SFDiff) for high-fidelity S2O image translation. SFDiff progressively narrows the gap between synthesized and real images in both spatial and frequency perspectives, showcasing notable performance in terms of quality and consistency. Specifically, to incorporate rich spatial content priors provided by SAR images, we design an SAR context prior extractor (SCPE) with denoising enhancement to extract multiscale conditional representations, thereby aiding SFDiff in capturing more descriptive cues for S2O translation. In addition, a spatial-frequency complementary learning (SFCL) module is designed to learn spatial semantics and simultaneously enhances informative frequency components and global dependencies. Furthermore, SFDiff is optimized using the joint spatial-frequency refinement loss, facilitating iterative refinement in both spatial and frequency domains to enhance content consistency and fidelity in the synthesized images. Based on the experimental findings from the UNICORN dataset and the SEN12 dataset, SFDiff maintains a high level of content and structural consistency, resulting in visually appealing translation results that surpass the state-of-the-art (SOTA) methods. In particular, SFDiff exhibits excellent performance in preserving small targets and details, which is crucial in cross-modal detection applications. Jiang Qin, Kai Wang 0060, Bin Zou 0001, Lamei Zhang, Joost van de Weijer 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | 3D-Aware Multi-Class Image-to-Image Translation with NeRFsabstractRecent advances in 3D-aware generative models (3D-aware GANs) combined with Neural Radiance Fields (NeRF) have achieved impressive results. However no prior works investigate 3D-aware GANs for 3D consistent multiclass image-to-image (3D-aware 121) translation. Naively using 2D-121 translation methods suffers from unrealistic shape/identity change. To perform 3D-aware multiclass 121 translation, we decouple this learning process into a multiclass 3D-aware GAN step and a 3D-aware 121 translation step. In the first step, we propose two novel techniques: a new conditional architecture and an effective training strategy. In the second step, based on the well-trained multiclass 3D-aware GAN architecture, that preserves view-consistency, we construct a 3D-aware 121 translation system. To further reduce the view-consistency problems, we propose several new techniques, including a U-net-like adaptor network design, a hierarchical representation constrain and a relative regularization loss. In exten-sive experiments on two datasets, quantitative and qualitative results demonstrate that we successfully perform 3D-aware 121 translation with multi-view consistency. Code is available in 3DI2I. Senmao Li, Joost van de Weijer 0001, Yaxing Wang, Fahad Shahbaz Khan, Meiqin Liu 0002, Jian Yang 0003 |
CVPR | 2 |
| 2023 | Endpoints Weight Fusion for Class Incremental Semantic SegmentationabstractClass incremental semantic segmentation (CISS) focuses on alleviating catastrophic forgetting to improve discrimination. Previous work mainly exploits regularization (e.g., knowledge distillation) to maintain previous knowledge in the current model. However, distillation alone often yields limited gain to the model since only the representations of old and new models are restricted to be consistent. In this paper, we propose a simple yet effective method to obtain a model with a strong memory of old knowledge, named Endpoints Weight Fusion (EWF). In our method, the model containing old knowledge is fused with the model retaining new knowledge in a dynamic fusion manner, strengthening the memory of old classes in ever changing distributions. In addition, we analyze the relationship between our fusion strategy and a popular moving average technique EMA, which reveals why our method is more suitable for class-incremental learning. To facilitate parameter fusion with closer distance in the parameter space, we use distillation to enhance the optimization process. Furthermore, we conduct experiments on two widely used datasets, achieving state-of-the-art performance. Jia-Wen Xiao, Chang-Bin Zhang, Jiekang Feng, Xialei Liu, Joost van de Weijer 0001, Ming-Ming Cheng |
CVPR | 5 |
| 2023 | Augmented Box Replay: Overcoming Foreground Shift for Incremental Object DetectionabstractIn incremental learning, replaying stored samples from previous tasks together with current task samples is one of the most efficient approaches to address catastrophic forgetting. However, unlike incremental classification, image replay has not been successfully applied to incremental object detection (IOD). In this paper, we identify the overlooked problem of foreground shift as the main reason for this. Foreground shift only occurs when replaying images of previous tasks and refers to the fact that their background might contain foreground objects of the current task. To overcome this problem, a novel and efficient Augmented Box Replay (ABR) method is developed that only stores and replays foreground objects and thereby circumvents the foreground shift problem. In addition, we propose an innovative Attentive RoI Distillation loss that uses spatial attention from region-of-interest (RoI) features to constrain current model to focus on the most important information from old model. ABR significantly reduces forgetting of previous classes while maintaining high plasticity in current classes. Moreover, it considerably reduces the storage requirements when compared to standard image replay. Comprehensive experiments on Pascal-VOC and COCO datasets support the state-of-the-art performance of our model1. Yang Cong, Dipam Goswami, Xialei Liu, Joost van de Weijer 0001 |
ICCV | 5 |
| 2023 | ICICLE: Interpretable Class Incremental Continual LearningabstractContinual learning enables incremental learning of new tasks without forgetting those previously learned, resulting in positive knowledge transfer that can enhance performance on both new and old tasks. However, continual learning poses new challenges for interpretability, as the rationale behind model predictions may change over time, leading to interpretability concept drift. We address this problem by proposing Interpretable Class-InCremental LEarning (ICICLE), an exemplar-free approach that adopts a prototypical part-based approach. It consists of three crucial novelties: interpretability regularization that distills previously learned concepts while preserving user-friendly positive reasoning; proximity-based prototype initialization strategy dedicated to the fine-grained setting; and task-recency bias compensation devoted to prototypical parts. Our experimental results demonstrate that ICICLE reduces the interpretability concept drift and outperforms the existing exemplar-free methods of common class-incremental learning when applied to concept-based models. Dawid Rymarczyk, Joost van de Weijer 0001, Bartosz Zielinski 0001, Bartlomiej Twardowski |
ICCV | 2 |
| 2023 | Planckian Jitter: countering the color-crippling effects of color jitter on self-supervised training
Simone Zini, Alexandra Gomez-Villa, Marco Buzzelli, Bartlomiej Twardowski, Andrew D. Bagdanov, Joost van de Weijer 0001 |
ICLR | 6 |
| 2023 | Dynamic Prompt Learning: Addressing Cross-Attention Leakage for Text-Based Image EditingabstractLarge-scale text-to-image generative models have been a ground-breaking development in generative AI, with diffusion models showing their astounding ability to synthesize convincing images following an input text prompt. The goal of image editing research is to give users control over the generated images by modifying the text prompt. Current image editing techniques are susceptible to unintended modifications of regions outside the targeted area, such as on the background or on distractor objects which have some semantic or visual relationship with the targeted object. According to our experimental findings, inaccurate cross-attention maps are at the root of this problem. Based on this observation, we propose $\textit{Dynamic Prompt Learning}$ ($DPL$) to force cross-attention maps to focus on correct $\textit{noun}$ words in the text prompt. By updating the dynamic tokens for nouns in the textual input with the proposed leakage repairment losses, we achieve fine-grained image editing over particular objects while preventing undesired changes to other image regions. Our method $DPL$, based on the publicly available $\textit{Stable Diffusion}$, is extensively evaluated on a wide range of images, and consistently obtains superior results both quantitatively (CLIP score, Structure-Dist) and qualitatively (on user-evaluation). We show improved prompt editing results for Word-Swap, Prompt Refinement, and Attention Re-weighting, especially for complex multi-object scenes. Kai Wang 0060, Fei Yang 0004, Shiqi Yang 0002, Muhammad Atif Butt, Joost van de Weijer 0001 |
NeurIPS | 5 |
| 2023 | FeCAM: Exploiting the Heterogeneity of Class Distributions in Exemplar-Free Continual LearningabstractExemplar-free class-incremental learning (CIL) poses several challenges since it prohibits the rehearsal of data from previous tasks and thus suffers from catastrophic forgetting. Recent approaches to incrementally learning the classifier by freezing the feature extractor after the first task have gained much attention. In this paper, we explore prototypical networks for CIL, which generate new class prototypes using the frozen feature extractor and classify the features based on the Euclidean distance to the prototypes. In an analysis of the feature distributions of classes, we show that classification based on Euclidean metrics is successful for jointly trained features. However, when learning from non-stationary data, we observe that the Euclidean metric is suboptimal and that feature distributions are heterogeneous. To address this challenge, we revisit the anisotropic Mahalanobis distance for CIL. In addition, we empirically show that modeling the feature covariance relations is better than previous attempts at sampling features from normal distributions and training a linear classifier. Unlike existing methods, our approach generalizes to both many- and few-shot CIL settings, as well as to domain-incremental settings. Interestingly, without updating the backbone network, our method obtains state-of-the-art results on several standard continual learning benchmarks. Code is available at https://github.com/dipamgoswami/FeCAM. Dipam Goswami, Bartlomiej Twardowski, Joost van de Weijer 0001 |
NeurIPS | 4 |
| 2023 | Attribution-aware Weight Transfer: A Warm-Start Initialization for Class-Incremental Semantic SegmentationabstractIn class-incremental semantic segmentation (CISS), deep learning architectures suffer from the critical problems of catastrophic forgetting and semantic background shift. Although recent works focused on these issues, existing classifier initialization methods do not address the background shift problem and assign the same initialization weights to both background and new foreground class classifiers. We propose to address the background shift with a novel classifier initialization method which employs gradient-based attribution to identify the most relevant weights for new classes from the classifier’s weights for the previous background and transfers these weights to the new classifier. This warm-start weight initialization provides a general solution applicable to several CISS methods. Furthermore, it accelerates learning of new classes while mitigating forgetting. Our experiments demonstrate significant improvement in mIoU compared to the state-of-the-art CISS methods on the Pascal-VOC 2012, ADE20K and Cityscapes datasets. Dipam Goswami, René Schuster, Joost van de Weijer 0001, Didier Stricker |
WACV | 3 |
| 2023 | Casting a BAIT for offline and online source-free domain adaptation
Shiqi Yang 0002, Yaxing Wang, Luis Herranz, Shangling Jui, Joost van de Weijer 0001 |
Comput. Vis. Image Underst. | 5 |
| 2023 | Generative Multi-Label Zero-Shot LearningabstractMulti-label zero-shot learning strives to classify images into multiple unseen categories for which no data is available during training. The test samples can additionally contain seen categories in the generalized variant. Existing approaches rely on learning either shared or label-specific attention from the seen classes. Nevertheless, computing reliable attention maps for unseen classes during inference in a multi-label setting is still a challenge. In contrast, state-of-the-art single-label generative adversarial network (GAN) based approaches learn to directly synthesize the class-specific visual features from the corresponding class attribute embeddings. However, synthesizing multi-label features from GANs is still unexplored in the context of zero-shot setting. When multiple objects occur jointly in a single image, a critical question is how to effectively fuse multi-class information. In this work, we introduce different fusion approaches at the attribute-level, feature-level and cross-level (across attribute and feature-levels) for synthesizing multi-label features from their corresponding multi-label class embeddings. To the best of our knowledge, our work is the first to tackle the problem of multi-label feature synthesis in the (generalized) zero-shot setting. Our cross-level fusion-based generative approach outperforms the state-of-the-art on three zero-shot benchmarks: NUS-WIDE, Open Images and MS COCO. Furthermore, we show the generalization capabilities of our fusion approach in the zero-shot detection task on MS COCO, achieving favorable performance against existing methods. Sanath Narayan, Salman Khan 0001, Fahad Shahbaz Khan, Ling Shao 0001, Joost van de Weijer 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Class-Incremental Learning: Survey and Performance Evaluation on Image ClassificationabstractFor future learning systems, incremental learning is desirable because it allows for: efficient resource usage by eliminating the need to retrain from scratch at the arrival of new data; reduced memory usage by preventing or limiting the amount of data required to be stored - also important when privacy limitations are imposed; and learning that more closely resembles human learning. The main challenge for incremental learning is catastrophic forgetting, which refers to the precipitous drop in performance on previously learned tasks after learning a new one. Incremental learning of deep neural networks has seen explosive growth in recent years. Initial work focused on task-incremental learning, where a task-ID is provided at inference time. Recently, we have seen a shift towards class-incremental learning where the learner must discriminate at inference time between all classes seen in previous tasks without recourse to a task-ID. In this paper, we provide a complete survey of existing class-incremental learning methods for image classification, and in particular, we perform an extensive experimental evaluation on thirteen class-incremental methods. We consider several new experimental scenarios, including a comparison of class-incremental methods on multiple large-scale image classification datasets, an investigation into small and large domain shifts, and a comparison of various network architectures. Marc Masana, Xialei Liu, Bartlomiej Twardowski, Mikel Menta, Andrew D. Bagdanov, Joost van de Weijer 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Trust Your Good Friends: Source-Free Domain Adaptation by Reciprocal Neighborhood ClusteringabstractDomain adaptation (DA) aims to alleviate the domain shift between source domain and target domain. Most DA methods require access to the source data, but often that is not possible (e.g., due to data privacy or intellectual property). In this paper, we address the challenging source-free domain adaptation (SFDA) problem, where the source pretrained model is adapted to the target domain in the absence of source data. Our method is based on the observation that target data, which might not align with the source domain classifier, still forms clear clusters. We capture this intrinsic structure by defining local affinity of the target data, and encourage label consistency among data with high local affinity. We observe that higher affinity should be assigned to reciprocal neighbors. To aggregate information with more context, we consider expanded neighborhoods with small affinity values. Furthermore, we consider the density around each target sample, which can alleviate the negative impact of potential outliers. In the experimental results we verify that the inherent structure of the target features is an important source of information for domain adaptation. We demonstrate that this local structure can be efficiently captured by considering the local neighbors, the reciprocal neighbors, and the expanded neighborhood. Finally, we achieve state-of-the-art performance on several 2D image and 3D point cloud recognition datasets. Shiqi Yang 0002, Yaxing Wang, Joost van de Weijer 0001, Luis Herranz, Shangling Jui, Jian Yang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Self-Training for Class-Incremental Semantic SegmentationabstractIn class-incremental semantic segmentation, we have no access to the labeled data of previous tasks. Therefore, when incrementally learning new classes, deep neural networks suffer from catastrophic forgetting of previously learned knowledge. To address this problem, we propose to apply a self-training approach that leverages unlabeled data, which is used for rehearsal of previous knowledge. Specifically, we first learn a temporary model for the current task, and then, pseudo labels for the unlabeled data are computed by fusing information from the old model of the previous task and the current temporary model. In addition, conflict reduction is proposed to resolve the conflicts of pseudo labels generated from both the old and temporary models. We show that maximizing self-entropy can further improve results by smoothing the overconfident predictions. Interestingly, in the experiments, we show that the auxiliary data can be different from the training data and that even general-purpose, but diverse auxiliary data can lead to large performance gains. The experiments demonstrate the state-of-the-art results: obtaining a relative gain of up to 114% on Pascal-VOC 2012 and 8.5% on the more challenging ADE20K compared to previous state-of-the-art methods. Lu Yu 0004, Xialei Liu, Joost van de Weijer 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Attention Distillation: self-supervised vision transformer students need more guidance
Kai Wang 0060, Fei Yang 0004, Joost van de Weijer 0001 |
BMVC | 3 |
| 2022 | Positive Pair Distillation Considered Harmful: Continual Meta Metric Learning for Lifelong Object Re-Identification
Kai Wang 0060, Chenshen Wu, Andrew D. Bagdanov, Xialei Liu, Shiqi Yang 0002, Shangling Jui, Joost van de Weijer 0001 |
BMVC | 7 |
| 2022 | MVMO: A Multi-Object Dataset for Wide Baseline Multi-View Semantic SegmentationabstractWe present MVMO (Multi-View, Multi-Object dataset): a synthetic dataset of 116,000 scenes containing randomly placed objects of 10 distinct classes and captured from 25 camera locations in the upper hemisphere. MVMO comprises photorealistic, path-traced image renders, together with semantic segmentation ground truth for every view. Unlike existing multi-view datasets, MVMO features wide baselines between cameras and high density of objects, which lead to large disparities, heavy occlusions and view-dependent object appearance. Single view semantic segmentation is hindered by self and inter-object occlusions that could benefit from additional viewpoints. Therefore, we expect that MVMO will propel research in multi-view semantic segmentation and cross-view semantic transfer. We also provide baselines that show that new research is needed in such fields to exploit the complementary information of multi-view setups1. Aitor Alvarez-Gila, Joost van de Weijer 0001, Yaxing Wang, Estíbaliz Garrote |
ICIP | 2 |
| 2022 | Distilling GANs with Style-Mixed Triplets for X2I Translation with Limited Data
Yaxing Wang, Joost van de Weijer 0001, Lu Yu 0004, Shangling Jui |
ICLR | 2 |
| 2022 | Visual Transformers with Primal Object Queries for Multi-Label Image ClassificationabstractMulti-label image classification is about predicting a set of class labels that can be considered as orderless sequential data. Transformers process the sequential data as a whole, therefore they are inherently good at set prediction. The first vision-based transformer model, which was proposed for the object detection task introduced the concept of object queries. Object queries are learnable positional encodings that are used by attention modules in decoder layers to decode the object classes or bounding boxes using the region of interests in an image. However, inputting the same set of object queries to different decoder layers hinders the training: it results in lower performance and delays convergence. In this paper, we propose the usage of primal object queries that are only provided at the start of the transformer decoder stack. In addition, we improve the mixup technique proposed for multi-label classification. The proposed transformer model with primal object queries improves the state-of-the-art class wise F1 metric by 2.1% and 1.8%; and speeds up the convergence by 79.0% and 38.6% on MS-COCO and NUS-WIDE datasets respectively. Vacit Oguz Yazici, Joost van de Weijer 0001 |
ICPR | 2 |
| 2022 | Attracting and Dispersing: A Simple Approach for Source-free Domain AdaptationabstractWe propose a simple but effective source-free domain adaptation (SFDA) method. Treating SFDA as an unsupervised clustering problem and following the intuition that local neighbors in feature space should have more similar predictions than other features, we propose to optimize an objective of prediction consistency. This objective encourages local neighborhood features in feature space to have similar predictions while features farther away in feature space have dissimilar predictions, leading to efficient feature clustering and cluster assignment simultaneously. For efficient training, we seek to optimize an upper-bound of the objective resulting in two simple terms. Furthermore, we relate popular existing methods in domain adaptation, source-free domain adaptation and contrastive learning via the perspective of discriminability and diversity. The experimental results prove the superiority of our method, and our method can be adopted as a simple but strong baseline for future research in SFDA. Our method can be also adapted to source-free open-set and partial-set DA which further shows the generalization ability of our method. Code is available in https://github.com/Albert0147/AaD_SFDA. Shiqi Yang 0002, Yaxing Wang, Kai Wang 0060, Shangling Jui, Joost van de Weijer 0001 |
NeurIPS | 5 |
| 2022 | Class-Balanced Active Learning for Image ClassificationabstractActive learning aims to reduce the labeling effort that is required to train algorithms by learning an acquisition function selecting the most relevant data for which a label should be requested from a large unlabeled data pool. Active learning is generally studied on balanced datasets where an equal amount of images per class is available. However, real-world datasets suffer from severe imbalanced classes, the so called long-tail distribution. We argue that this further complicates the active learning process, since the imbalanced data pool can result in suboptimal classifiers. To address this problem in the context of active learning, we proposed a general optimization framework that explicitly takes class-balancing into account. Results on three datasets showed that the method is general (it can be combined with most existing active learning algorithms) and can be effectively applied to boost the performance of both informative and representative-based active learning methods. In addition, we showed that also on balanced datasets our method1generally results in a performance gain. Javad Zolfaghari Bengar, Joost van de Weijer 0001, Laura Lopez-Fuentes, Bogdan Raducanu |
WACV | 2 |
| 2021 | HCV: Hierarchy-Consistency Verification for Incremental Implicitly-Refined Classification
Kai Wang 0060, Xialei Liu, Luis Herranz, Joost van de Weijer 0001 |
BMVC | 4 |
| 2021 | When Deep Learners Change Their Mind: Learning Dynamics for Active Learning
Javad Zolfaghari Bengar, Bogdan Raducanu, Joost van de Weijer 0001 |
CAIP (1) | 3 |
| 2021 | TransferI2I: Transfer Learning for Image-to-Image Translation from Small DatasetsabstractImage-to-image (I2I) translation has matured in recent years and is able to generate high-quality realistic images. However, despite current success, it still faces important challenges when applied to small domains. Existing methods use transfer learning for I2I translation, but they still require the learning of millions of parameters from scratch. This drawback severely limits its application on small domains. In this paper, we propose a new transfer learning for I2I translation (TransferI2I). We decouple our learning process into the image generation step and the I2I translation step. In the first step we propose two novel techniques: source-target initialization and self-initialization of the adaptor layer. The former finetunes the pretrained generative model (e.g., StyleGAN) on source and target data. The latter allows to initialize all non-pretrained network parameters without the need of any data. These techniques provide a better initialization for the I2I translation step. In addition, we introduce an auxiliary GAN that further facilitates the training of deep I2I systems even from small datasets. In extensive experiments on three datasets, (Animal faces, Birds, and Foods), we show that we outperform existing methods and that mFID improves on several datasets with over 25 points. Our code is available at: https://github.com/yaxingwang/TransferI2I. Yaxing Wang, Héctor Laria Mantecon, Joost van de Weijer 0001, Laura Lopez-Fuentes, Bogdan Raducanu |
ICCV | 3 |
| 2021 | Generalized Source-free Domain AdaptationabstractDomain adaptation (DA) aims to transfer the knowledge learned from a source domain to an unlabeled target domain. Some recent works tackle source-free domain adaptation (SFDA) where only a source pre-trained model is available for adaptation to the target domain. However, those methods do not consider keeping source performance which is of high practical value in real world applications. In this paper, we propose a new domain adaptation paradigm called Generalized Source-free Domain Adaptation (G-SFDA), where the learned model needs to perform well on both the target and source domains, with only access to current unlabeled target data during adaptation. First, we propose local structure clustering (LSC), aiming to cluster the target features with its semantically similar neighbors, which successfully adapts the model to the target domain in the absence of source data. Second, we propose sparse domain attention (SDA), it produces a binary domain specific attention to activate different feature channels for different domains, meanwhile the domain attention will be utilized to regularize the gradient during adaptation to keep source information. In the experiments, for target performance our method is on par with or better than existing DA and SFDA methods, specifically it achieves state-of-the-art performance (85.4%) on VisDA, and our method works well for all domains after adapting to single or multiple target domains. Code is available in https://github.com/Albert0147/G-SFDA. Shiqi Yang 0002, Yaxing Wang, Joost van de Weijer 0001, Luis Herranz, Shangling Jui |
ICCV | 3 |
| 2021 | Exploiting the Intrinsic Neighborhood Structure for Source-free Domain AdaptationabstractDomain adaptation (DA) aims to alleviate the domain shift between source domain and target domain. Most DA methods require access to the source data, but often that is not possible (e.g. due to data privacy or intellectual property). In this paper, we address the challenging source-free domain adaptation (SFDA) problem, where the source pretrained model is adapted to the target domain in the absence of source data. Our method is based on the observation that target data, which might no longer align with the source domain classifier, still forms clear clusters. We capture this intrinsic structure by defining local affinity of the target data, and encourage label consistency among data with high local affinity. We observe that higher affinity should be assigned to reciprocal neighbors, and propose a self regularization loss to decrease the negative impact of noisy neighbors. Furthermore, to aggregate information with more context, we consider expanded neighborhoods with small affinity values. In the experimental results we verify that the inherent structure of the target features is an important source of information for domain adaptation. We demonstrate that this local structure can be efficiently captured by considering the local neighbors, the reciprocal neighbors, and the expanded neighborhood. Finally, we achieve state-of-the-art performance on several 2D image and 3D point cloud recognition datasets. Code is available in https://github.com/Albert0147/SFDA_neighbors. Shiqi Yang 0002, Yaxing Wang, Joost van de Weijer 0001, Luis Herranz, Shangling Jui |
NeurIPS | 3 |
| 2021 | Controlling biases and diversity in diverse image-to-image translation
Yaxing Wang, Abel Gonzalez-Garcia, Luis Herranz, Joost van de Weijer 0001 |
Comput. Vis. Image Underst. | 4 |
| 2021 | Saliency for free: Saliency prediction as a side-effect of object recognition
Carola Figueroa Flores, David Berga, Joost van de Weijer 0001, Bogdan Raducanu |
Pattern Recognit. Lett. | 3 |
| 2021 | ACAE-REMIND for online continual learning with compressed feature replay
Kai Wang 0060, Joost van de Weijer 0001, Luis Herranz |
Pattern Recognit. Lett. | 2 |
| 2021 | On Implicit Attribute Localization for Generalized Zero-Shot LearningabstractZero-shot learning (ZSL) aims to discriminate images from unseen classes by exploiting relations to seen classes via their attribute-based descriptions. Since attributes are often related to specific parts of objects, many recent works focus on discovering discriminative regions. However, these methods usually require additional complex part detection modules or attention mechanisms. In this paper, 1) we show that common ZSL backbones (without explicit attention nor part detection) can implicitly localize attributes, yet this property is not exploited. 2) Exploiting it, we then propose SELAR, a simple method that further encourages attribute localization, surprisingly achieving very competitive generalized ZSL (GZSL) performance when compared with more complex state-of-the-art methods. Our findings provide useful insight for designing future GZSL methods, and SELAR provides an easy to implement yet strong baseline. Shiqi Yang 0002, Kai Wang 0060, Luis Herranz, Joost van de Weijer 0001 |
IEEE Signal Process. Lett. | 4 |
| 2021 | Distributed Learning and Inference With Compressed ImagesabstractModern computer vision requires processing large amounts of data, both while training the model and/or during inference, once the model is deployed. Scenarios where images are captured and processed in physically separated locations are increasingly common (e.g. autonomous vehicles, cloud computing, smartphones). In addition, many devices suffer from limited resources to store or transmit data (e.g. storage space, channel capacity). In these scenarios, lossy image compression plays a crucial role to effectively increase the number of images collected under such constraints. However, lossy compression entails some undesired degradation of the data that may harm the performance of the downstream analysis task at hand, since important semantic information may be lost in the process. Moreover, we may only have compressed images at training time but are able to use original images at inference time (i.e. test), or vice versa, and in such a case, the downstream model suffers from covariate shift. In this paper, we analyze this phenomenon, with a special focus on vision-based perception for autonomous driving as a paradigmatic scenario. We see that loss of semantic information and covariate shift do indeed exist, resulting in a drop in performance that depends on the compression rate. In order to address the problem, we propose dataset restoration, based on image restoration with generative adversarial networks (GANs). Our method is agnostic to both the particular image compression method and the downstream task; and has the advantage of not adding additional cost to the deployed models, which is particularly important in resource-limited devices. The presented experiments focus on semantic segmentation as a challenging use case, cover a broad range of compression rates and diverse datasets, and show how our method is able to significantly alleviate the negative effects of compression on the downstream visual task. Sudeep Katakol, Basem Elbarashy, Luis Herranz, Joost van de Weijer 0001, Antonio M. López 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Semantic Drift Compensation for Class-Incremental LearningabstractClass-incremental learning of deep networks sequentially increases the number of classes to be classified. During training, the network has only access to data of one task at a time, where each task contains several classes. In this setting, networks suffer from catastrophic forgetting which refers to the drastic drop in performance on previous tasks. The vast majority of methods have studied this scenario for classification networks, where for each new task the classification layer of the network must be augmented with additional weights to make room for the newly added classes. Embedding networks have the advantage that new classes can be naturally included into the network without adding new weights. Therefore, we study incremental learning for embedding networks. In addition, we propose a new method to estimate the drift, called semantic drift, of features and compensate for it without the need of any exemplars. We approximate the drift of previous tasks based on the drift that is experienced by current task data. We perform experiments on fine-grained datasets, CIFAR100 and ImageNet-Subset. We demonstrate that embedding networks suffer significantly less from catastrophic forgetting. We outperform existing methods which do not require exemplars and obtain competitive results compared to methods which store exemplars. Furthermore, we show that our proposed SDC when combined with existing methods to prevent forgetting consistently improves results. Lu Yu 0004, Bartlomiej Twardowski, Xialei Liu, Luis Herranz, Kai Wang 0060, Yongmei Cheng, Shangling Jui, Joost van de Weijer 0001 |
CVPR | 8 |
| 2020 | MineGAN: Effective Knowledge Transfer From GANs to Target Domains With Few ImagesabstractOne of the attractive characteristics of deep neural networks is their ability to transfer knowledge obtained in one domain to other related domains. As a result, high-quality networks can be trained in domains with relatively little training data. This property has been extensively studied for discriminative networks but has received significantly less attention for generative models. Given the often enormous effort required to train GANs, both computationally as well as in the dataset collection, the re-use of pretrained GANs is a desirable objective. We propose a novel knowledge transfer method for generative models based on mining the knowledge that is most beneficial to a specific target domain, either from a single or multiple pretrained GANs. This is done using a miner network that identifies which part of the generative distribution of each pretrained GAN outputs samples closest to the target domain. Mining effectively steers GAN sampling towards suitable regions of the latent space, which facilitates the posterior finetuning and avoids pathologies of other methods such as mode collapse and lack of flexibility. We perform experiments on several complex datasets using various GAN architectures (BigGAN, Progressive GAN) and show that the proposed method, called MineGAN, effectively transfers knowledge to domains with few target images, outperforming existing methods. In addition, MineGAN can successfully transfer knowledge from multiple pretrained GANs. Our code is available at: \url{https://github.com/yaxingwang/MineGAN}. Yaxing Wang, Abel Gonzalez-Garcia, David Berga, Luis Herranz, Fahad Shahbaz Khan, Joost van de Weijer 0001 |
CVPR | 6 |
| 2020 | Semi-Supervised Learning for Few-Shot Image-to-Image TranslationabstractIn the last few years, unpaired image-to-image translation has witnessed Remarkable progress. Although the latest methods are able to generate realistic images, they crucially rely on a large number of labeled images. Recently, some methods have tackled the challenging setting of few-shot image-to-image ranslation, reducing the labeled data requirements for the target domain during inference. In this work, we go one step further and reduce the amount of required labeled data also from the source domain during training. To do so, we propose applying semi-supervised learning via a noise-tolerant pseudo-labeling procedure. We also apply a cycle consistency constraint to further exploit the information from unlabeled images, either from the same dataset or external. Additionally, we propose several structural modifications to facilitate the image translation task under these circumstances. Our semi-supervised method for few-shot image translation, called \emph{SEMIT}, achieves excellent results on four different datasets using as little as 10\% of the source labels, and matches the performance of the main fully-supervised competitor using only 20\% labeled data. Our code and models are made public at: \url{https://github.com/yaxingwang/SEMIT}. Yaxing Wang, Salman Khan 0001, Abel Gonzalez-Garcia, Joost van de Weijer 0001, Fahad Shahbaz Khan |
CVPR | 4 |
| 2020 | Orderless Recurrent Models for Multi-Label ClassificationabstractRecurrent neural networks (RNN) are popular for many computer vision tasks, including multi-label classification. Since RNNs produce sequential outputs, labels need to be ordered for the multi-label classification task. Current approaches sort labels according to their frequency, typically ordering them in either rare-first or frequent-first. These imposed orderings do not take into account that the natural order to generate the labels can change for each image, e.g. first the dominant object before summing up the smaller objects in the image. Therefore, in this paper, we propose ways to dynamically order the ground truth labels with the predicted label sequence. This allows for the faster training of more optimal LSTM models for multi-label classification. Analysis evidences that our method does not suffer from duplicate generation, something which is common for other models. Furthermore, it outperforms other CNN-RNN models, and we show that a standard architecture of an image encoder and language decoder trained with our proposed loss obtains the state-of-the-art results on the challenging MS-COCO, WIDER Attribute and PA-100K and competitive results on NUS-WIDE. Vacit Oguz Yazici, Abel Gonzalez-Garcia, Arnau Ramisa, Bartlomiej Twardowski, Joost van de Weijer 0001 |
CVPR | 5 |
| 2020 | Learning to Rank for Active Learning: A Listwise ApproachabstractActive learning emerged as an alternative to alleviate the effort to label huge amount of data for data-hungry applications (such as image/video indexing and retrieval, autonomous driving, etc.). The goal of active learning is to automatically select a number of unlabeled samples for annotation (according to a budget), based on an acquisition function, which indicates how valuable a sample is for training the model. The learning loss method is a task-agnostic approach which attaches a module to learn to predict the target loss of unlabeled data, and select data with the highest loss for labeling. In this work, we follow this strategy but we define the acquisition function as a learning to rank problem and rethink the structure of the loss prediction module, using a simple but effective listwise approach. Experimental results on four datasets demonstrate that our method outperforms recent state-of-the-art active learning approaches for both image classification and regression tasks. Minghan Li 0003, Xialei Liu, Joost van de Weijer 0001, Bogdan Raducanu |
ICPR | 3 |
| 2020 | RATT: Recurrent Attention to Transient Tasks for Continual Image CaptioningabstractResearch on continual learning has led to a variety of approaches to mitigating catastrophic forgetting in feed-forward classification networks. Until now surprisingly little attention has been focused on continual learning of recurrent models applied to problems like image captioning. In this paper we take a systematic look at continual learning of LSTM-based models for image captioning. We propose an attention-based approach that explicitly accommodates the transient nature of vocabularies in continual image captioning tasks -- i.e. that task vocabularies are not disjoint. We call our method Recurrent Attention to Transient Tasks (RATT), and also show how to adapt continual learning approaches based on weight regularization and knowledge distillation to recurrent continual learning problems. We apply our approaches to incremental image captioning problem on two new continual learning benchmarks we define using the MS-COCO and Flickr30 datasets. Our results demonstrate that RATT is able to sequentially learn five captioning tasks while incurring no forgetting of previously learned ones. Riccardo Del Chiaro, Bartlomiej Twardowski, Andrew D. Bagdanov, Joost van de Weijer 0001 |
NeurIPS | 4 |
| 2020 | DeepI2I: Enabling Deep Hierarchical Image-to-Image Translation by Transferring from GANsabstractImage-to-image translation has recently achieved remarkable results. But despite current success, it suffers from inferior performance when translations between classes require large shape changes. We attribute this to the high-resolution bottlenecks which are used by current state-of-the-art image-to-image methods. Therefore, in this work, we propose a novel deep hierarchical Image-to-Image Translation method, called DeepI2I. We learn a model by leveraging hierarchical features: (a) structural information contained in the bottom layers and (b) semantic information extracted from the top layers. To enable the training of deep I2I models on small datasets, we propose a novel transfer learning method, that transfers knowledge from pre-trained GANs. Specifically, we leverage the discriminator of a pre-trained GANs (i.e. BigGAN or StyleGAN) to initialize both the encoder and the discriminator and the pre-trained generator to initialize the generator of our model. Applying knowledge transfer leads to an alignment problem between the encoder and generator. We introduce an adaptor network to address this. On many-class image-to-image translation on three datasets (Animal faces, Birds, and Foods) we decrease mFID by at least 35% when compared to the state-of-the-art. Furthermore, we qualitatively and quantitatively demonstrate that transfer learning significantly improves the performance of I2I systems, especially for small datasets. Finally, we are the first to perform I2I translations for domains with over 100 classes. Yaxing Wang, Lu Yu 0004, Joost van de Weijer 0001 |
NeurIPS | 3 |
| 2020 | Mix and Match Networks: Cross-Modal Alignment for Zero-Pair Image-to-Image Translation
Yaxing Wang, Luis Herranz, Joost van de Weijer 0001 |
Int. J. Comput. Vis. | 3 |
| 2020 | Object proposals for salient object segmentation in videos
Rahma Kalboussi, Aymen Azaza, Joost van de Weijer 0001, Mehrez Abdellaoui, Ali Douik |
Multim. Tools Appl. | 3 |
| 2020 | Variable Rate Deep Image Compression With Modulated AutoencoderabstractVariable rate is a requirement for flexible and adaptable image and video compression. However, deep image compression methods (DIC) are optimized for a single fixed rate-distortion (R-D) tradeoff. While this can be addressed by training multiple models for different tradeoffs, the memory requirements increase proportionally to the number of models. Scaling the bottleneck representation of a shared autoencoder can provide variable rate compression with a single shared autoencoder. However, the R-D performance using this simple mechanism degrades in low bitrates, and also shrinks the effective range of bitrates. To address these limitations, we formulate the problem of variable R-D optimization for DIC, and propose modulated autoencoders (MAEs), where the representations of a shared autoencoder are adapted to the specific R-D tradeoff via a modulation network. Jointly training this modulated autoencoder and the modulation network provides an effective way to navigate the R-D operational curve. Our experiments show that the proposed method can achieve almost the same R-D performance of independent models with significantly fewer parameters. Fei Yang 0004, Luis Herranz, Joost van de Weijer 0001, José Antonio Iglesias Guitián, Antonio M. López 0001, Mikhail G. Mozerov |
IEEE Signal Process. Lett. | 3 |
| 2019 | Learning Metrics From Teachers: Compact Networks for Image EmbeddingabstractMetric learning networks are used to compute image embeddings, which are widely used in many applications such as image retrieval and face recognition. In this paper, we propose to use network distillation to efficiently compute image embeddings with small networks. Network distillation has been successfully applied to improve image classification, but has hardly been explored for metric learning. To do so, we propose two new loss functions that model the communication of a deep teacher network to a small student network. We evaluate our system in several datasets, including CUB-200-2011, Cars-196, Stanford Online Products and show that embeddings computed using small student networks perform significantly better than those computed using standard networks of similar size. Results on a very compact network (MobileNet-0.25), which can be used on mobile devices, show that the proposed method can greatly improve Recall@1 results from 27.5\% to 44.6\%. Furthermore, we investigate various aspects of distillation for embeddings, including hint and attention layers, semi-supervised learning and cross quality distillation. (Code is available at https://github.com/yulu0724/EmbeddingDistillation). Lu Yu 0004, Vacit Oguz Yazici, Xialei Liu, Joost van de Weijer 0001, Yongmei Cheng, Arnau Ramisa |
CVPR | 4 |
| 2019 | Active Learning for Deep Detection Neural NetworksabstractThe cost of drawing object bounding boxes (i.e. labeling) for millions of images is prohibitively high. For instance, labeling pedestrians in a regular urban image could take 35 seconds on average. Active learning aims to reduce the cost of labeling by selecting only those images that are informative to improve the detection network accuracy. In this paper, we propose a method to perform active learning of object detectors based on convolutional neural networks. We propose a new image-level scoring process to rank unlabeled images for their automatic selection, which clearly outperforms classical scores. The proposed method can be applied to videos and sets of still images. In the former case, temporal selection rules can complement our scoring process. As a relevant use case, we extensively study the performance of our method on the task of pedestrian detection. Overall, the experiments show that the proposed method performs better than random selection. Hamed H. Aghdam, Abel Gonzalez-Garcia, Antonio M. López 0001, Joost van de Weijer 0001 |
ICCV | 4 |
| 2019 | Learning the Model Update for Siamese TrackersabstractSiamese approaches address the visual tracking problem by extracting an appearance template from the current frame, which is used to localize the target in the next frame. In general, this template is linearly combined with the accumulated template from the previous frame, resulting in an exponential decay of information over time. While such an approach to updating has led to improved results, its simplicity limits the potential gain likely to be obtained by learning to update. Therefore, we propose to replace the handcrafted update function with a method which learns to update. We use a convolutional neural network, called UpdateNet, which given the initial template, the accumulated template and the template of the current frame aims to estimate the optimal template for the next frame. The UpdateNet is compact and can easily be integrated into existing Siamese trackers. We demonstrate the generality of the proposed approach by applying it to two Siamese trackers, SiamFC and DaSiamRPN. Extensive experiments on VOT2016, VOT2018, LaSOT, and TrackingNet datasets demonstrate that our UpdateNet effectively predicts the new target template, outperforming the standard linear update. On the large-scale TrackingNet dataset, our UpdateNet improves the results of DaSiamRPN with an absolute gain of 3.9% in terms of success score. Lichao Zhang 0001, Abel Gonzalez-Garcia, Joost van de Weijer 0001, Martin Danelljan, Fahad Shahbaz Khan |
ICCV | 3 |
| 2019 | SDIT: Scalable and Diverse Cross-domain Image TranslationabstractRecently, image-to-image translation research has witnessed remarkable progress. Although current approaches successfully generate diverse outputs or perform scalable image transfer, these properties have not been combined into a single method. To address this limitation, we propose SDIT: Scalable and Diverse image-to-image translation. These properties are combined into a single generator. The diversity is determined by a latent variable which is randomly sampled from a normal distribution. The scalability is obtained by conditioning the network on the domain attributes. Additionally, we also exploit an attention mechanism that permits the generator to focus on the domain-specific attribute. We empirically demonstrate the performance of the proposed method on face mapping and other datasets beyond faces. Yaxing Wang, Abel Gonzalez-Garcia, Joost van de Weijer 0001, Luis Herranz |
ACM Multimedia | 3 |
| 2019 | Self-supervised blur detection from synthetically blurred scenes
Aitor Alvarez-Gila, Adrian Galdran, Estíbaliz Garrote, Joost van de Weijer 0001 |
Image Vis. Comput. | 4 |
| 2019 | Exploiting Unlabeled Data in CNNs by Self-Supervised Learning to RankabstractFor many applications the collection of labeled data is expensive laborious. Exploitation of unlabeled data during training is thus a long pursued objective of machine learning. Self-supervised learning addresses this by positing an auxiliary task (different, but related to the supervised task) for which data is abundantly available. In this paper, we show how ranking can be used as a proxy task for some regression problems. As another contribution, we propose an efficient backpropagation technique for Siamese networks which prevents the redundant computation introduced by the multi-branch network architecture. We apply our framework to two regression problems: Image Quality Assessment (IQA) and Crowd Counting. For both we show how to automatically generate ranked image sets from unlabeled data. Our results show that networks trained to regress to the ground truth targets for labeled data and to simultaneously learn to rank unlabeled data obtain significantly better, state-of-the-art results for both IQA and crowd counting. In addition, we show that measuring network uncertainty on the self-supervised proxy task is a good measure of informativeness of unlabeled data. This can be used to drive an algorithm for active learning and we show that this reduces labeling effort by up to 50 percent. Xialei Liu, Joost van de Weijer 0001, Andrew D. Bagdanov |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Saliency for fine-grained object recognition in domains with scarce training data
Carola Figueroa Flores, Abel Gonzalez-Garcia, Joost van de Weijer 0001, Bogdan Raducanu |
Pattern Recognit. | 3 |
| 2019 | Sparse Data Interpolation Using the Geodesic Distance Affinity SpaceabstractIn this letter, we adapt the geodesic distance-based recursive filter to the sparse data interpolation problem. The proposed technique is general and can be easily applied to any kind of sparse data. We demonstrate its superiority over other interpolation techniques in three experiments for qualitative and quantitative evaluation. In addition, we compare our method with the popular interpolation algorithm presented in the paper on EpicFlow optical flow, which is intuitively motivated by a similar geodesic distance principle. The comparison shows that our algorithm is more accurate and considerably faster than the EpicFlow interpolation technique. Mikhail G. Mozerov, Fei Yang 0004, Joost van de Weijer 0001 |
IEEE Signal Process. Lett. | 3 |
| 2019 | One-View Occlusion Detection for Stereo Matching With a Fully Connected CRF ModelabstractIn this paper, we extend the standard belief propagation (BP) sequential technique proposed in the tree-reweighted sequential method [15] to the fully connected CRF models with the geodesic distance affinity. The proposed method has been applied to the stereo matching problem. Also a new approach to the BP marginal solution is proposed that we call one-view occlusion detection (OVOD). In contrast to the standard winner takes all (WTA) estimation, the proposed OVOD solution allows to find occluded regions in the disparity map and simultaneously improve the matching result. As a result we can perform only one energy minimization process and avoid the cost calculation for the second view and the left-right check procedure. We show that the OVOD approach considerably improves results for cost augmentation and energy minimization techniques in comparison with the standard one-view affinity space implementation. We apply our method to the Middlebury data set and reach state-of-the-art especially for median, average and mean squared error metrics. Mikhail G. Mozerov, Joost van de Weijer 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Synthetic Data Generation for End-to-End Thermal Infrared TrackingabstractThe usage of both off-the-shelf and end-to-end trained deep networks have significantly improved the performance of visual tracking on RGB videos. However, the lack of large labeled datasets hampers the usage of convolutional neural networks for tracking in thermal infrared (TIR) images. Therefore, most state-of-the-art methods on tracking for TIR data are still based on handcrafted features. To address this problem, we propose to use image-to-image translation models. These models allow us to translate the abundantly available labeled RGB data to synthetic TIR data. We explore both the usage of paired and unpaired image translation models for this purpose. These methods provide us with a large labeled dataset of synthetic TIR sequences, on which we can train end-to-end optimal features for tracking. To the best of our knowledge, we are the first to train end-to-end features for TIR tracking. We perform extensive experiments on the VOT-TIR2017 dataset. We show that a network trained on a large dataset of synthetic TIR data obtains better performance than one trained on the available real TIR data. Combining both data sources leads to further improvement. In addition, when we combine the network with motion features, we outperform the state of the art with a relative gain of over 10%, clearly showing the efficiency of using synthetic data to train end-to-end TIR trackers. Lichao Zhang 0001, Abel Gonzalez-Garcia, Joost van de Weijer 0001, Martin Danelljan, Fahad Shahbaz Khan |
IEEE Trans. Image Process. | 3 |
| 2018 | Metric Learning for Novelty and Anomaly Detection
Marc Masana, Idoia Ruiz, Joan Serrat 0002, Joost van de Weijer 0001, Antonio M. López 0001 |
BMVC | 4 |
| 2018 | Leveraging Unlabeled Data for Crowd Counting by Learning to RankabstractWe propose a novel crowd counting approach that leverages abundantly available unlabeled crowd imagery in a learning-to-rank framework. To induce a ranking of cropped images, we use the observation that any sub-image of a crowded scene image is guaranteed to contain the same number or fewer persons than the super-image. This allows us to address the problem of limited size of existing datasets for crowd counting. We collect two crowd scene datasets from Google using keyword searches and query-by-example image retrieval, respectively. We demonstrate how to efficiently learn from these unlabeled datasets by incorporating learning-to-rank in a multi-task network which simultaneously ranks images and estimates crowd density maps. Experiments on two of the most challenging crowd counting datasets show that our approach obtains state-of-the-art results. Xialei Liu, Joost van de Weijer 0001, Andrew D. Bagdanov |
CVPR | 2 |
| 2018 | Mix and Match Networks: Encoder-Decoder Alignment for Zero-Pair Image TranslationabstractWe address the problem of image translation between domains or modalities for which no direct paired data is available (i.e. zero-pair translation). We propose mix and match networks, based on multiple encoders and decoders aligned in such a way that other encoder-decoder pairs can be composed at test time to perform unseen image translation tasks between domains or modalities for which explicit paired samples were not seen during training. We study the impact of autoencoders, side information and losses in improving the alignment and transferability of trained pairwise translation models to unseen translations. We show our approach is scalable and can perform colorization and style transfer between unseen combinations of domains. We evaluate our system in a challenging cross-modal setting where semantic segmentation is estimated from depth images, without explicit access to any depth-semantic segmentation training pairs. Our model outperforms baselines based on pix2pix and CycleGAN models. Yaxing Wang, Joost van de Weijer 0001, Luis Herranz |
CVPR | 2 |
| 2018 | Transferring GANs: Generating Images from Limited Data
Yaxing Wang, Chenshen Wu, Luis Herranz, Joost van de Weijer 0001, Abel Gonzalez-Garcia, Bogdan Raducanu |
ECCV (6) | 4 |
| 2018 | Learning Illuminant Estimation from Object RecognitionabstractIn this paper we present a deep learning method to estimate the illuminant of an image. Our model is not trained with illuminant annotations, but with the objective of improving performance on an auxiliary task such as object recognition. To the best of our knowledge, this is the first example of a deep learning architecture for illuminant estimation that is trained without ground truth illuminants. We evaluate our solution on standard datasets for color constancy, and compare it with state of the art methods. Our proposal is shown to outperform most deep learning methods in a cross-dataset evaluation setup, and to present competitive results in a comparison with parametric solutions. Marco Buzzelli, Joost van de Weijer 0001, Raimondo Schettini |
ICIP | 2 |
| 2018 | Weakly Supervised Domain-Specific Color Naming Based on AttentionabstractThe majority of existing color naming methods focuses on the eleven basic color terms of the English language. However, in many applications, different sets of color names are used for the accurate description of objects. Labeling data to learn these domain-specific color names is an expensive and laborious task. Therefore, in this article we aim to learn color names from weakly labeled data. For this purpose, we add an attention branch to the color naming network. The attention branch is used to modulate the pixel-wise color naming predictions of the network. In experiments, we illustrate that the attention branch correctly identifies the relevant regions. Furthermore, we show that our method obtains state-of-the-art results for pixel-wise and image-wise classification on the EBAY dataset and is able to learn color names for various domains. Lu Yu 0004, Yongmei Cheng, Joost van de Weijer 0001 |
ICPR | 3 |
| 2018 | Rotate your Networks: Better Weight Consolidation and Less Catastrophic ForgettingabstractIn this paper we propose an approach to avoiding catastrophic forgetting in sequential task learning scenarios. Our technique is based on a network reparameterization that approximately diagonalizes the Fisher Information Matrix of the network parameters. This reparameterization takes the form of a factorized rotation of parameter space which, when used in conjunction with Elastic Weight Consolidation (which assumes a diagonal Fisher Information Matrix), leads to significantly better performance on lifelong learning of sequential tasks. Experimental results on the MNIST, CIFAR-100, CUB-200 and Stanford-40 datasets demonstrate that we significantly improve the results of standard elastic weight consolidation, and that we obtain competitive results when compared to the state-of-the-art in lifelong learning without forgetting. Xialei Liu, Marc Masana, Luis Herranz, Joost van de Weijer 0001, Antonio M. López 0001, Andrew D. Bagdanov |
ICPR | 4 |
| 2018 | Image-to-image translation for cross-domain disentanglementabstractDeep image translation methods have recently shown excellent results, outputting high-quality images covering multiple modes of the data distribution. There has also been increased interest in disentangling the internal representations learned by deep methods to further improve their performance and achieve a finer control. In this paper, we bridge these two objectives and introduce the concept of cross-domain disentanglement. We aim to separate the internal representation into three parts. The shared part contains information for both domains. The exclusive parts, on the other hand, contain only factors of variation that are particular to each domain. We achieve this through bidirectional image translation based on Generative Adversarial Networks and cross-domain autoencoders, a novel network component. Our model offers multiple advantages. We can output diverse samples covering multiple modes of the distributions of both domains, perform domain- specific image transfer and interpolation, and cross-domain retrieval without the need of labeled data, only paired images. We compare our model to the state-of-the-art in multi-modal image translation and achieve better results for translation on challenging datasets as well as for cross-domain retrieval on realistic datasets. Abel Gonzalez-Garcia, Joost van de Weijer 0001, Yoshua Bengio |
NeurIPS | 2 |
| 2018 | Memory Replay GANs: Learning to Generate New Categories without ForgettingabstractPrevious works on sequential learning address the problem of forgetting in discriminative models. In this paper we consider the case of generative models. In particular, we investigate generative adversarial networks (GANs) in the task of learning new categories in a sequential fashion. We first show that sequential fine tuning renders the network unable to properly generate images from previous categories (i.e. forgetting). Addressing this problem, we propose Memory Replay GANs (MeRGANs), a conditional GAN framework that integrates a memory replay generator. We study two methods to prevent forgetting by leveraging these replays, namely joint training with replay and replay alignment. Qualitative and quantitative experimental results in MNIST, SVHN and LSUN datasets show that our memory replay approach can generate competitive images while significantly mitigating the forgetting of previous categories. Chenshen Wu, Luis Herranz, Xialei Liu, Yaxing Wang, Joost van de Weijer 0001, Bogdan Raducanu |
NeurIPS | 5 |
| 2018 | Color Naming for Multi-color Fashion Items
Vacit Oguz Yazici, Joost van de Weijer 0001, Arnau Ramisa |
WorldCIST (3) | 2 |
| 2018 | Context proposals for saliency detection
Aymen Azaza, Joost van de Weijer 0001, Ali Douik, Marc Masana |
Comput. Vis. Image Underst. | 2 |
| 2018 | Review on computer vision techniques in emergency situations
Laura Lopez-Fuentes, Joost van de Weijer 0001, Manuel González Hidalgo, Harald Skinnemoen, Andrew D. Bagdanov |
Multim. Tools Appl. | 2 |
| 2018 | Scale coding bag of deep features for human attribute and action recognitionabstractMost approaches to human attribute and action recognition in still images are based on image representation in which multi-scale local features are pooled across scale into a single, scale-invariant encoding. Both in bag-of-words and the recently popular representations based on convolutional neural networks, local features are computed at multiple scales. However, these multi-scale convolutional features are pooled into a single scale-invariant representation. We argue that entirely scale-invariant image representations are sub-optimal and investigate approaches to scale coding within a bag of deep features framework. Our approach encodes multi-scale information explicitly during the image encoding stage. We propose two strategies to encode multi-scale information explicitly in the final image representation. We validate our two scale coding techniques on five datasets: Willow, PASCAL VOC 2010, PASCAL VOC 2012, Stanford-40 and Human Attributes (HAT-27). On all datasets, the proposed scale coding approaches outperform both the scale-invariant method and the standard deep features of the same network. Further, combining our scale coding approaches with standard deep features leads to consistent improvement over the state of the art. Fahad Shahbaz Khan, Joost van de Weijer 0001, Rao Muhammad Anwer, Andrew D. Bagdanov, Michael Felsberg, Jorma Laaksonen |
Mach. Vis. Appl. | 2 |
| 2018 | Beyond Eleven Color Names for Image Understanding
Lu Yu 0004, Lichao Zhang 0001, Joost van de Weijer 0001, Fahad Shahbaz Khan, Yongmei Cheng, C. Alejandro Párraga |
Mach. Vis. Appl. | 3 |
| 2017 | 3D color charts for camera spectral sensitivity estimation
Rada Deeb, Damien Muselet, Mathieu Hébert, Alain Trémeau, Joost van de Weijer 0001 |
BMVC | 5 |
| 2017 | RankIQA: Learning from Rankings for No-Reference Image Quality AssessmentabstractWe propose a no-reference image quality assessment (NR-IQA) approach that learns from rankings (RankIQA). To address the problem of limited IQA dataset size, we train a Siamese Network to rank images in terms of image quality by using synthetically generated distortions for which relative image quality is known. These ranked image sets can be automatically generated without laborious human labeling. We then use fine-tuning to transfer the knowledge represented in the trained Siamese Network to a traditional CNN that estimates absolute image quality from single images. We demonstrate how our approach can be made significantly more efficient than traditional Siamese Networks by forward propagating a batch of images through a single network and backpropagating gradients derived from all pairs of images in the batch. Experiments on the TID2013 benchmark show that we improve the state-of-theart by over 5%. Furthermore, on the LIVE benchmark we show that our approach is superior to existing NR-IQA techniques and that we even outperform the state-of-the-art in full-reference IQA (FR-IQA) methods without having to resort to high-quality reference images to infer IQA. Xialei Liu, Joost van de Weijer 0001, Andrew D. Bagdanov |
ICCV | 2 |
| 2017 | Domain-Adaptive Deep Network Compression
Marc Masana, Joost van de Weijer 0001, Luis Herranz, Andrew D. Bagdanov, José M. Álvarez 0004 |
ICCV | 2 |
| 2017 | TEX-Nets: Binary Patterns Encoded Convolutional Neural Networks for Texture RecognitionabstractRecognizing materials and textures in realistic imaging conditions is a challenging computer vision problem. For many years, local features based orderless representations were a dominant approach for texture recognition. Recently deep local features, extracted from the intermediate layers of a Convolutional Neural Network (CNN), are used as filter banks. These dense local descriptors from a deep model, when encoded with Fisher Vectors, have shown to provide excellent results for texture recognition. The CNN models, employed in such approaches, take RGB patches as input and train on a large amount of labeled images. We show that CNN models, which we call TEX-Nets, trained using mapped coded images with explicit texture information provide complementary information to the standard deep models trained on RGB patches. We further investigate two deep architectures, namely early and late fusion, to combine the texture and color information. Experiments on benchmark texture datasets clearly demonstrate that TEX-Nets provide complementary information to standard RGB deep network. Our approach provides a large gain of 4.8%, 3.5%, 2.6% and 4.1% respectively in accuracy on the DTD, KTH-TIPS-2a, KTH-TIPS-2b and Texture-10 datasets, compared to the standard RGB network of the same architecture. Further, our final combination leads to consistent improvements over the state-of-the-art on all four datasets. Rao Muhammad Anwer, Fahad Shahbaz Khan, Joost van de Weijer 0001, Jorma Laaksonen |
ICMR | 3 |
| 2017 | Bandwidth Limited Object Recognition in High Resolution ImageryabstractThis paper proposes a novel method to optimize bandwidth usage for object detection in critical communication scenarios. We develop two operating models of active information seeking. The first model identifies promising regions in low resolution imagery and progressively requests higher resolution regions on which to perform recognition of higher semantic quality. The second model identifies promising regions in low resolution imagery while simultaneously predicting the approximate location of the object of higher semantic quality. From this general framework, we develop a car recognition system via identification of its license plate and evaluate the performance of both models on a car dataset that we introduce. Results are compared with traditional JPEG compression and demonstrate that our system saves up to one order of magnitude of bandwidth while sacrificing little in terms of recognition performance. Laura Lopez-Fuentes, Andrew D. Bagdanov, Joost van de Weijer 0001, Harald Skinnemoen |
WACV | 3 |
| 2017 | Improved Recursive Geodesic Distance Computation for Edge Preserving FilterabstractAll known recursive filters based on the geodesic distance affinity are realized by two 1D recursions applied in two orthogonal directions of the image plane. The 2D extension of the filter is not valid and has theoretically drawbacks, which lead to known artifacts. In this paper, a maximum influence propagation method is proposed to approximate the 2D extension for the geodesic distance-based recursive filter. The method allows to partially overcome the drawbacks of the 1D recursion approach. We show that our improved recursion better approximates the true geodesic distance filter, and the application of this improved filter for image denoising outperforms the existing recursive implementation of the geodesic distance. As an application, we consider a geodesic distance-based filter for image denoising. Experimental evaluation of our denoising method demonstrates comparable and for several test images better results, than state-of-the-art approaches, while our algorithm is considerably faster with computational complexity O(8P). Mikhail G. Mozerov, Joost van de Weijer 0001 |
IEEE Trans. Image Process. | 2 |
| 2016 | Hierarchical part detection with deep neural networksabstractPart detection is an important aspect of object recognition. Most approaches apply object proposals to generate hundreds of possible part bounding box candidates which are then evaluated by part classifiers. Recently several methods have investigated directly regressing to a limited set of bounding boxes from deep neural network representation. However, for object parts such methods may be unfeasible due to their relatively small size with respect to the image. We propose a hierarchical method for object and part detection. In a single network we first detect the object and then regress to part location proposals based only on the feature representation inside the object. Experiments show that our hierarchical approach outperforms a network which directly regresses the part locations. We also show that our approach obtains part detection accuracy comparable or better than state-of-the-art on the CUB-200 bird and Fashionista clothing item datasets with only a fraction of the number of part proposals. Esteve Cervantes, Andrew D. Bagdanov, Marc Masana, Joost van de Weijer 0001 |
ICIP | 5 |
| 2016 | Combining Holistic and Part-based Deep Representations for Computational Painting CategorizationabstractAutomatic analysis of visual art, such as paintings, is a challenging inter-disciplinary research problem. Conventional approaches only rely on global scene characteristics by encoding holistic information for computational painting categorization. We argue that such approaches are sub-optimal and that discriminative common visual structures provide complementary information for painting classification. Rao Muhammad Anwer, Fahad Shahbaz Khan, Joost van de Weijer 0001, Jorma Laaksonen |
ICMR | 3 |
| 2015 | From Emotions to Action Units with Hidden and Semi-Hidden-Task LearningabstractLimited annotated training data is a challenging problem in Action Unit recognition. In this paper, we investigate how the use of large databases labelled according to the 6 universal facial expressions can increase the generalization ability of Action Unit classifiers. For this purpose, we propose a novel learning framework: Hidden-Task Learning. HTL aims to learn a set of Hidden-Tasks (Action Units) for which samples are not available but, in contrast, training data is easier to obtain from a set of related Visible-Tasks (Facial Expressions). To that end, HTL is able to exploit prior knowledge about the relation between Hidden and Visible-Tasks. In our case, we base this prior knowledge on empirical psychological studies providing statistical correlations between Action Units and universal facial expressions. Additionally, we extend HTL to Semi-Hidden Task Learning (SHTL) assuming that Action Unit training samples are also provided. Performing exhaustive experiments over four different datasets, we show that HTL and SHTL improve the generalization ability of AU classifiers by training them with additional facial expression data. Additionally, we show that SHTL achieves competitive performance compared with state-of-the-art Transductive Learning approaches which face the problem of limited training data by using unlabelled test samples during training. Adria Ruiz, Joost van de Weijer 0001, Xavier Binefa |
ICCV | 2 |
| 2015 | Compact color-texture description for texture classification
Fahad Shahbaz Khan, Rao Muhammad Anwer, Joost van de Weijer 0001, Michael Felsberg, Jorma Laaksonen |
Pattern Recognit. Lett. | 3 |
| 2015 | Recognizing Actions Through Action-Specific Person DetectionabstractAction recognition in still images is a challenging problem in computer vision. To facilitate comparative evaluation independently of person detection, the standard evaluation protocol for action recognition uses an oracle person detector to obtain perfect bounding box information at both training and test time. The assumption is that, in practice, a general person detector will provide candidate bounding boxes for action recognition. In this paper, we argue that this paradigm is suboptimal and that action class labels should already be considered during the detection stage. Motivated by the observation that body pose is strongly conditioned on action class, we show that: 1) the existing state-of-the-art generic person detectors are not adequate for proposing candidate bounding boxes for action classification; 2) due to limited training examples, the direct training of action-specific person detectors is also inadequate; and 3) using only a small number of labeled action examples, the transfer learning is able to adapt an existing detector to propose higher quality bounding boxes for subsequent action classification. To the best of our knowledge, we are the first to investigate transfer learning for the task of action-specific person detection in still images. We perform extensive experiments on two benchmark data sets: 1) Stanford-40 and 2) PASCAL VOC 2012. For the action detection task (i.e., both person localization and classification of the action performed), our approach outperforms methods based on general person detection by 5.7% mean average precision (MAP) on Stanford-40 and 2.1% MAP on PASCAL VOC 2012. Our approach also significantly outperforms the state of the art with a MAP of 45.4% on Stanford-40 and 31.4% on PASCAL VOC 2012. We also evaluate our action detection approach for the task of action classification (i.e., recognizing actions without localizing them). For this task, our approach, without using any ground-truth person localization at test time, outperforms on both data sets state-of-the-art methods, which do use person locations. Fahad Shahbaz Khan, Jiaolong Xu, Joost van de Weijer 0001, Andrew D. Bagdanov, Rao Muhammad Anwer, Antonio M. López 0001 |
IEEE Trans. Image Process. | 3 |
| 2015 | Accurate Stereo Matching by Two-Step Energy MinimizationabstractIn stereo matching, cost-filtering methods and energy-minimization algorithms are considered as two different techniques. Due to their global extent, energy-minimization methods obtain good stereo matching results. However, they tend to fail in occluded regions, in which cost-filtering approaches obtain better results. In this paper, we intend to combine both the approaches with the aim to improve overall stereo matching results.We show that a global optimization with a fully connected model can be solved by cost-filtering methods. Based on this observation, we propose to perform stereo matching as a two-step energy-minimization algorithm. We consider two Markov random field (MRF) models: 1) a fully connected model defined on the complete set of pixels in an image and 2) a conventional locally connected model. We solve the energy-minimization problem for the fully connected model, after which the marginal function of the solution is used as the unary potential in the locally connected MRF model. Experiments on the Middlebury stereo data sets show that the proposed method achieves the state-of-the-arts results. Mikhail G. Mozerov, Joost van de Weijer 0001 |
IEEE Trans. Image Process. | 2 |
| 2015 | Global Color Sparseness and a Local Statistics Prior for Fast Bilateral FilteringabstractThe property of smoothing while preserving edges makes the bilateral filter a very popular image processing tool. However, its non-linear nature results in a computationally costly operation. Various works propose fast approximations to the bilateral filter. However, the majority does not generalize to vector input as is the case with color images. We propose a fast approximation to the bilateral filter for color images. The filter is based on two ideas. First, the number of colors, which occur in a single natural image, is limited. We exploit this color sparseness to rewrite the initial non-linear bilateral filter as a number of linear filter operations. Second, we impose a statistical prior to the image values that are locally present within the filter window. We show that this statistical prior leads to a closed-form solution of the bilateral filter. Finally, we combine both ideas into a single fast and accurate bilateral filter for color images. Experimental results show that our bilateral filter based on the local prior yields an extremely fast bilateral filter approximation, but with limited accuracy, which has potential application in real-time video filtering. Our bilateral filter, which combines color sparseness and local statistics, yields a fast and accurate bilateral filter approximation and obtains the state-of-the-art results. Mikhail G. Mozerov, Joost van de Weijer 0001 |
IEEE Trans. Image Process. | 2 |
| 2014 | Regularized Multi-Concept MIL for weakly-supervised facial behavior categorization
Adria Ruiz, Joost van de Weijer 0001, Xavier Binefa |
BMVC | 2 |
| 2014 | Adaptive Color Attributes for Real-Time Visual TrackingabstractVisual tracking is a challenging problem in computer vision. Most state-of-the-art visual trackers either rely on luminance information or use simple color representations for image description. Contrary to visual tracking, for object recognition and detection, sophisticated color features when combined with luminance have shown to provide excellent performance. Due to the complexity of the tracking problem, the desired color feature should be computationally efficient, and possess a certain amount of photometric invariance while maintaining high discriminative power. This paper investigates the contribution of color in a tracking-by-detection framework. Our results suggest that color attributes provides superior performance for visual tracking. We further propose an adaptive low-dimensional variant of color attributes. Both quantitative and attribute-based evaluations are performed on 41 challenging benchmark color sequences. The proposed approach improves the baseline intensity-based tracker by 24 % in median distance precision. Furthermore, we show that our approach outperforms state-of-the-art tracking methods while running at more than 100 frames per second. Martin Danelljan, Fahad Shahbaz Khan, Michael Felsberg, Joost van de Weijer 0001 |
CVPR | 4 |
| 2014 | Scale Coding Bag-of-Words for Action RecognitionabstractRecognizing human actions in still images is a challenging problem in computer vision due to significant amount of scale, illumination and pose variation. Given the bounding box of a person both at training and test time, the task is to classify the action associated with each bounding box in an image. Most state-of-the-art methods use the bag-of-words paradigm for action recognition. The bag-of-words framework employing a dense multi-scale grid sampling strategy is the de facto standard for feature detection. This results in a scale invariant image representation where all the features at multiple-scales are binned in a single histogram. We argue that such a scale invariant strategy is sub-optimal since it ignores the multi-scale information available with each bounding box of a person. This paper investigates alternative approaches to scale coding for action recognition in still images. We encode multi-scale information explicitly in three different histograms for small, medium and large scale visual-words. Our first approach exploits multi-scale information with respect to the image size. In our second approach, we encode multi-scale information relative to the size of the bounding box of a person instance. In each approach, the multi-scale histograms are then concatenated into a single representation for action classification. We validate our approaches on the Willow dataset which contains seven action categories: interacting with computer, photography, playing music, riding bike, riding horse, running and walking. Our results clearly suggest that the proposed scale coding approaches outperform the conventional scale invariant technique. Moreover, we show that our approach obtains promising results compared to more complex state-of-the-art methods. Fahad Shahbaz Khan, Joost van de Weijer 0001, Andrew D. Bagdanov, Michael Felsberg |
ICPR | 2 |
| 2014 | Painting-91: a large scale database for computational painting categorization
Fahad Shahbaz Khan, Shida Kunz, Joost van de Weijer 0001, Michael Felsberg |
Mach. Vis. Appl. | 3 |
| 2014 | Multi-Illuminant Estimation With Conditional Random FieldsabstractMost existing color constancy algorithms assume uniform illumination. However, in real-world scenes, this is not often the case. Thus, we propose a novel framework for estimating the colors of multiple illuminants and their spatial distribution in the scene. We formulate this problem as an energy minimization task within a conditional random field over a set of local illuminant estimates. In order to quantitatively evaluate the proposed method, we created a novel data set of two-dominant-illuminant images comprised of laboratory, indoor, and outdoor scenes. Unlike prior work, our database includes accurate pixel-wise ground truth illuminant information. The performance of our method is evaluated on multiple data sets. Experimental results show that our framework clearly outperforms single illuminant estimators as well as a recently proposed multi-illuminant estimation approach. Shida Kunz, Christian Riess, Joost van de Weijer 0001, Elli Angelopoulou |
IEEE Trans. Image Process. | 3 |
| 2014 | Semantic Pyramids for Gender and Action RecognitionabstractPerson description is a challenging problem in computer vision. We investigated two major aspects of person description: 1) gender and 2) action recognition in still images. Most state-of-the-art approaches for gender and action recognition rely on the description of a single body part, such as face or full-body. However, relying on a single body part is suboptimal due to significant variations in scale, viewpoint, and pose in real-world images. This paper proposes a semantic pyramid approach for pose normalization. Our approach is fully automatic and based on combining information from full-body, upper-body, and face regions for gender and action recognition in still images. The proposed approach does not require any annotations for upper-body and face of a person. Instead, we rely on pretrained state-of-the-art upper-body and face detectors to automatically extract semantic information of a person. Given multiple bounding boxes from each body part detector, we then propose a simple method to select the best candidate bounding box, which is used for feature extraction. Finally, the extracted features from the full-body, upper-body, and face regions are combined into a single representation for classification. To validate the proposed approach for gender recognition, experiments are performed on three large data sets namely: 1) human attribute; 2) head-shoulder; and 3) proxemics. For action recognition, we perform experiments on four data sets most used for benchmarking action recognition in still images: 1) Sports; 2) Willow; 3) PASCAL VOC 2010; and 4) Stanford-40. Our experiments clearly demonstrate that the proposed approach, despite its simplicity, outperforms state-of-the-art methods for gender and action recognition. Fahad Shahbaz Khan, Joost van de Weijer 0001, Rao Muhammad Anwer, Michael Felsberg, Carlo Gatta |
IEEE Trans. Image Process. | 2 |
| 2013 | Evaluating the Impact of Color on Texture Recognition
Fahad Shahbaz Khan, Joost van de Weijer 0001, Sadiq Ali, Michael Felsberg |
CAIP (1) | 2 |
| 2013 | Discriminative Color DescriptorsabstractColor description is a challenging task because of large variations in RGB values which occur due to scene accidental events, such as shadows, shading, specularities, illuminant color changes, and changes in viewing geometry. Traditionally, this challenge has been addressed by capturing the variations in physics-based models, and deriving invariants for the undesired variations. The drawback of this approach is that sets of distinguishable colors in the original color space are mapped to the same value in the photometric invariant space. This results in a drop of discriminative power of the color description. In this paper we take an information theoretic approach to color description. We cluster color values together based on their discriminative power in a classification problem. The clustering has the explicit objective to minimize the drop of mutual information of the final representation. We show that such a color description automatically learns a certain degree of photometric invariance. We also show that a universal color representation, which is based on other data sets than the one at hand, can obtain competing performance. Experiments show that the proposed descriptor outperforms existing photometric invariants. Furthermore, we show that combined with shape description these color descriptors obtain excellent results on four challenging datasets, namely, PASCAL VOC 2007, Flowers-102, Stanford dogs-120 and Birds-200. Rahat Khan, Joost van de Weijer 0001, Fahad Shahbaz Khan, Damien Muselet, Christophe Ducottet, Cécile Barat |
CVPR | 2 |
| 2013 | An Active Contour Model for Speech Balloon Detection in ComicsabstractComic books constitute an important cultural heritage asset in many countries. Digitization combined with subsequent comic book understanding would enable a variety of new applications, including content-based retrieval and content retargeting. Document understanding in this domain is challenging as comics are semi-structured documents, combining semantically important graphical and textual parts. Few studies have been done in this direction. In this work we detail a novel approach for closed and non-closed speech balloon localization in scanned comic book pages, an essential step towards a fully automatic comic book understanding. The approach is compared with existing methods for closed balloon localization found in the literature and results are presented. Christophe Rigaud, Jean-Christophe Burie, Jean-Marc Ogier, Dimosthenis Karatzas, Joost van de Weijer 0001 |
ICDAR | 5 |
| 2013 | Intrinsic image evaluation on synthetic complex scenesabstractScene decomposition into its illuminant, shading, and reflectance intrinsic images is an essential step for scene understanding. Collecting intrinsic image groundtruth data is a laborious task. The assumptions on which the ground-truth procedures are based limit their application to simple scenes with a single object taken in the absence of indirect lighting and interreflections. We investigate synthetic data for intrinsic image research since the extraction of ground truth is straightforward, and it allows for scenes in more realistic situations (e.g, multiple illuminants and interreflections). With this dataset we aim to motivate researchers to further explore intrinsic image decomposition in complex scenes. Shida Kunz, Marc Serra, Joost van de Weijer 0001, Robert Benavente, María Vanrell 0001, Olivier Penacchio, Dimitris Samaras |
ICIP | 3 |
| 2013 | Towards multispectral data acquisition with hand-held devicesabstractWe propose a method to acquire multispectral data with handheld devices with front-mounted RGB cameras. We propose to use the display of the device as an illuminant while the camera captures images illuminated by the red, green and blue primaries of the display. Three illuminants and three response functions of the camera lead to nine response values which are used for reflectance estimation. Results are promising and show that the accuracy of the spectral reconstruction improves in the range of 30-40% over the spectral reconstruction based on a single illuminant. Furthermore, we propose to compute sensor-illuminant aware linear basis by discarding the part of the reflectances that falls in the sensor-illuminant null-space. We show experimentally that optimizing reflectance estimation on these new basis functions decreases the RMSE significantly over basis functions that are independent to sensor-illuminant. We conclude that, multispectral data acquisition is potentially possible with consumer hand-held devices such as tablets, mobiles, and laptops, opening up applications which are currently considered to be unrealistic. Rahat Khan, Joost van de Weijer 0001, Dimosthenis Karatzas, Damien Muselet |
ICIP | 2 |
| 2013 | Coloring Action Recognition in Still Images
Fahad Shahbaz Khan, Rao Muhammad Anwer, Joost van de Weijer 0001, Andrew D. Bagdanov, Antonio M. López 0001, Michael Felsberg |
Int. J. Comput. Vis. | 3 |
| 2012 | Color attributes for object detectionabstractState-of-the-art object detectors typically use shape information as a low level feature representation to capture the local structure of an object. This paper shows that early fusion of shape and color, as is popular in image classification, leads to a significant drop in performance for object detection. Moreover, such approaches also yields suboptimal results for object categories with varying importance of color and shape. In this paper we propose the use of color attributes as an explicit color representation for object detection. Color attributes are compact, computationally efficient, and when combined with traditional shape features provide state-of-the-art results for object detection. Our method is tested on the PASCAL VOC 2007 and 2009 datasets and results clearly show that our method improves over state-of-the-art techniques despite its simplicity. We also introduce a new dataset consisting of cartoon character images in which color plays a pivotal role. On this dataset, our approach yields a significant gain of 14% in mean AP over conventional state-of-the-art methods. Fahad Shahbaz Khan, Rao Muhammad Anwer, Joost van de Weijer 0001, Andrew D. Bagdanov, María Vanrell 0001, Antonio M. López 0001 |
CVPR | 3 |
| 2012 | Harmony Potentials - Fusing Global and Local Scale for Semantic Image Segmentation
Xavier Boix, Josep M. Gonfaus, Joost van de Weijer 0001, Andrew D. Bagdanov, Joan Serrat 0002, Jordi Gonzàlez 0001 |
Int. J. Comput. Vis. | 3 |
| 2012 | Modulating Shape Features by Color Attention for Object Recognition
Fahad Shahbaz Khan, Joost van de Weijer 0001, María Vanrell 0001 |
Int. J. Comput. Vis. | 2 |
| 2012 | Improving Color Constancy by Photometric Edge WeightingabstractEdge-based color constancy methods make use of image derivatives to estimate the illuminant. However, different edge types exist in real-world images, such as material, shadow, and highlight edges. These different edge types may have a distinctive influence on the performance of the illuminant estimation. Therefore, in this paper, an extensive analysis is provided of different edge types on the performance of edge-based color constancy methods. First, an edge-based taxonomy is presented classifying edge types based on their photometric properties (e.g., material, shadow-geometry, and highlights). Then, a performance evaluation of edge-based color constancy is provided using these different edge types. From this performance evaluation, it is derived that specular and shadow edge types are more valuable than material edges for the estimation of the illuminant. To this end, the (iterative) weighted Gray-Edge algorithm is proposed in which these edge types are more emphasized for the estimation of the illuminant. Images that are recorded under controlled circumstances demonstrate that the proposed iterative weighted Gray-Edge algorithm based on highlights reduces the median angular error with approximately 25 percent. In an uncontrolled environment, improvements in angular error up to 11 percent are obtained with respect to regular edge-based color constancy. Arjan Gijsenij, Theo Gevers, Joost van de Weijer 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Discriminative compact pyramids for object and scene recognition
Noha M. Elfiky, Fahad Shahbaz Khan, Joost van de Weijer 0001, Jordi Gonzàlez 0001 |
Pattern Recognit. | 3 |
| 2011 | Object recoloring based on intrinsic image estimationabstractObject recoloring is one of the most popular photo-editing tasks. The problem of object recoloring is highly under-constrained, and existing recoloring methods limit their application to objects lit by a white illuminant. Application of these methods to real-world scenes lit by colored illuminants, multiple illuminants, or interreflections, results in unrealistic recoloring of objects. In this paper, we focus on the recoloring of single-colored objects presegmented from their background. The single-color constraint allows us to fit a more comprehensive physical model to the object. We demonstrate that this permits us to perform realistic recoloring of objects lit by non-white illuminants, and multiple illuminants. Moreover, the model allows for more realistic handling of illuminant alteration of the scene. Recoloring results captured by uncalibrated cameras demonstrate that the proposed framework obtains realistic recoloring for complex natural images. Furthermore we use the model to transfer color between objects and show that the results are more realistic than existing color transfer methods. Shida Kunz, Joost van de Weijer 0001 |
ICCV | 2 |
| 2011 | Portmanteau Vocabularies for Multi-Cue Image RepresentationabstractWe describe a novel technique for feature combination in the bag-of-words model of image classification. Our approach builds discriminative compound words from primitive cues learned independently from training images. Our main observation is that modeling joint-cue distributions independently is more statistically robust for typical classification problems than attempting to empirically estimate the dependent, joint-cue distribution directly. We use Information theoretic vocabulary compression to find discriminative combinations of cues and the resulting vocabulary of portmanteau words is compact, has the cue binding property, and supports individual weighting of cues in the final image representation. State-of-the-art results on both the Oxford Flower-102 and Caltech-UCSD Bird-200 datasets demonstrate the effectiveness of our technique compared to other, significantly more complex approaches to multi-cue image representation Fahad Shahbaz Khan, Joost van de Weijer 0001, Andrew D. Bagdanov, María Vanrell 0001 |
NIPS | 2 |
| 2011 | Describing Reflectances for Color Segmentation Robust to Shadows, Highlights, and TexturesabstractThe segmentation of a single material reflectance is a challenging problem due to the considerable variation in image measurements caused by the geometry of the object, shadows, and specularities. The combination of these effects has been modeled by the dichromatic reflection model. However, the application of the model to real-world images is limited due to unknown acquisition parameters and compression artifacts. In this paper, we present a robust model for the shape of a single material reflectance in histogram space. The method is based on a multilocal creaseness analysis of the histogram which results in a set of ridges representing the material reflectances. The segmentation method derived from these ridges is robust to both shadow, shading and specularities, and texture in real-world images. We further complete the method by incorporating prior knowledge from image statistics, and incorporate spatial coherence by using multiscale color contrast information. Results obtained show that our method clearly outperforms state-of-the-art segmentation methods on a widely used segmentation benchmark, having as a main characteristic its excellent performance in the presence of shadows and highlights at low computational cost. Eduard Vazquez, Ramón Baldrich, Joost van de Weijer 0001, María Vanrell 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Computational Color Constancy: Survey and ExperimentsabstractComputational color constancy is a fundamental prerequisite for many computer vision applications. This paper presents a survey of many recent developments and state-of-the-art methods. Several criteria are proposed that are used to assess the approaches. A taxonomy of existing algorithms is proposed and methods are separated in three groups: static methods, gamut-based methods, and learning-based methods. Further, the experimental setup is discussed including an overview of publicly available datasets. Finally, various freely available methods, of which some are considered to be state of the art, are evaluated on two datasets. Arjan Gijsenij, Theo Gevers, Joost van de Weijer 0001 |
IEEE Trans. Image Process. | 3 |
| 2010 | Harmony potentials for joint classification and segmentationabstractHierarchical conditional random fields have been successfully applied to object segmentation. One reason is their ability to incorporate contextual information at different scales. However, these models do not allow multiple labels to be assigned to a single node. At higher scales in the image, this yields an oversimplified model, since multiple classes can be reasonable expected to appear within one region. This simplified model especially limits the impact that observations at larger scales may have on the CRF model. Neglecting the information at larger scales is undesirable since class-label estimates based on these scales are more reliable than at smaller, noisier scales. To address this problem, we propose a new potential, called harmony potential, which can encode any possible combination of class labels. We propose an effective sampling strategy that renders tractable the underlying optimization problem. Results show that our approach obtains state-of-the-art results on two challenging datasets: Pascal VOC 2009 and MSRC-21. Josep M. Gonfaus, Xavier Boix, Joost van de Weijer 0001, Andrew D. Bagdanov, Joan Serrat 0002, Jordi Gonzàlez 0001 |
CVPR | 3 |
| 2010 | The Impact of Color on Bag-of-Words Based Object RecognitionabstractIn recent years several works have aimed at exploiting color information in order to improve the bag-of-words based image representation. There are two stages in which color information can be applied in the bag-of-words framework. Firstly, feature detection can be improved by choosing highly informative color-based regions. Secondly, feature description, typically focusing on shape, can be improved with a color description of the local patches. Although both approaches have been shown to improve results the combined merits have not yet been analyzed. Therefore, in this paper we investigate the combined contribution of color to both the feature detection and extraction stages. Experiments performed on two challenging data sets, namely Flower and Pascal VOC 2009; clearly demonstrate that incorporating color in both feature detection and extraction significantly improves the overall performance. David Augusto Rojas Vigo, Fahad Shahbaz Khan, Joost van de Weijer 0001, Theo Gevers |
ICPR | 3 |
| 2010 | Generalized Gamut Mapping using Image Derivative Structures for Color ConstancyabstractThe gamut mapping algorithm is one of the most promising methods to achieve computational color constancy. However, so far, gamut mapping algorithms are restricted to the use of pixel values to estimate the illuminant. Therefore, in this paper, gamut mapping is extended to incorporate the statistical nature of images. It is analytically shown that the proposed gamut mapping framework is able to include any linear filter output. The main focus is on the local n -jet describing the derivative structure of an image. It is shown that derivatives have the advantage over pixel values to be invariant to disturbing effects (i.e. deviations of the diagonal model) such as saturated colors and diffuse light. Further, as the n -jet based gamut mapping has the ability to use more information than pixel values alone, the combination of these algorithms are more stable than the regular gamut mapping algorithm. Different methods of combining are proposed. Based on theoretical and experimental results conducted on large scale data sets of hyperspectral, laboratory and real-world scenes, it can be derived that (1) in case of deviations of the diagonal model, the derivative-based approach outperforms the pixel-based gamut mapping, (2) state-of-the-art algorithms are outperformed by the n -jet based gamut mapping, (3) the combination of the different n -jet based gamut mappings provide more stable solutions, and (4) the fusion strategy based on the intersection of feasible sets provides better color constancy results than the union of the feasible sets. Arjan Gijsenij, Theo Gevers, Joost van de Weijer 0001 |
Int. J. Comput. Vis. | 3 |
| 2009 | Physics-based edge evaluation for improved color constancyabstractEdge-based color constancy makes use of image derivatives to estimate the illuminant. However, different edge types exist in real-world images such as shadow, geometry, material and highlight edges. These different edge types may have a distinctive influence on the performance of the illuminant estimation. Arjan Gijsenij, Theo Gevers, Joost van de Weijer 0001 |
CVPR | 3 |
| 2009 | Top-down color attention for object recognitionabstractGenerally the bag-of-words based image representation follows a bottom-up paradigm. The subsequent stages of the process: feature detection, feature description, vocabulary construction and image representation are performed independent of the intentioned object classes to be detected. In such a framework, combining multiple cues such as shape and color often provides below-expected results. Fahad Shahbaz Khan, Joost van de Weijer 0001, María Vanrell 0001 |
ICCV | 2 |
| 2009 | Learning Color Names for Real-World ApplicationsabstractColor names are required in real-world applications such as image retrieval and image annotation. Traditionally, they are learned from a collection of labeled color chips. These color chips are labeled with color names within a well-defined experimental setup by human test subjects. However, naming colors in real-world images differs significantly from this experimental setting. In this paper, we investigate how color names learned from color chips compare to color names learned from real-world images. To avoid hand labeling real-world images with color names, we use Google Image to collect a data set. Due to the limitations of Google Image, this data set contains a substantial quantity of wrongly labeled data. We propose several variants of the PLSA model to learn color names from this noisy data. Experimental results show that color names learned from real-world images significantly outperform color names learned from labeled color chips for both image retrieval and image annotation. Joost van de Weijer 0001, Cordelia Schmid, Jakob Verbeek, Diane Larlus |
IEEE Trans. Image Process. | 1 |
| 2008 | Image Segmentation in the Presence of Shadows and Highlights
Eduard Vazquez, Joost van de Weijer 0001, Ramón Baldrich |
ECCV (4) | 2 |
| 2007 | Learning Color Names from Real-World ImagesabstractWithin a computer vision context color naming is the action of assigning linguistic color labels to image pixels. In general, research on color naming applies the following paradigm: a collection of color chips is labelled with color names within a well-defined experimental setup by multiple test subjects. The collected data set is subsequently used to label RGB values in real-world images with a color name. Apart from the fact that this collection process is time consuming, it is unclear to what extent color naming within a controlled setup is representative for color naming in real-world images. Therefore we propose to learn color names from real-world images. Furthermore, we avoid test subjects by using Google Image to collect a data set. Due to limitations of Google Image this data set contains a substantial quantity of wrongly labelled data. The color names are learned using a PLSA model adapted to this task. Experimental results show that color names learned from real-world images significantly outperform color names learned from labelled color chips on retrieval and classification. Joost van de Weijer 0001, Cordelia Schmid, Jakob Verbeek |
CVPR | 1 |
| 2007 | Using High-Level Visual Information for Color ConstancyabstractWe propose to use high-level visual information to improve illuminant estimation. Several illuminant estimation approaches are applied to compute a set of possible illuminants. For each of them an illuminant color corrected image is evaluated on the likelihood of its semantic content: is the grass green, the road grey, and the sky blue, in correspondence with our prior knowledge of the world. The illuminant resulting in the most likely semantic composition of the image is selected as the illuminant color. To evaluate the likelihood of the semantic content, we apply probabilistic latent semantic analysis. The image is modelled as a mixture of semantic classes, such as sky, grass, road, and building. The class description is based on texture, position and color information. Experiments show that the use of high-level information improves illuminant estimation over a purely bottom-up approach. Furthermore, the proposed method is shown to significantly improve semantic class recognition performance. Joost van de Weijer 0001, Cordelia Schmid, Jakob Verbeek |
ICCV | 1 |
| 2007 | Applying Color Names to Image DescriptionabstractPhotometric invariance is a desired property for color image descriptors. It ensures that the description has a certain robustness with respect to scene incidental variations such as changes in viewpoint, object orientation, and illuminant color. A drawback of photometric invariance is that the discriminative power of the description reduces while increasing the photometric invariance. In this paper, we look into the use of color names for the purpose of image description. Color names are linguistic labels that humans attach to colors. They display a certain amount of photometric invariance, and as an additional advantage allow the description of the achromatic colors, which are undistinguishable in a photometric invariant representation. Experiments on an image classification task show that color description based on color names outperforms description based on photometric invariants. Joost van de Weijer 0001, Cordelia Schmid |
ICIP (3) | 1 |
| 2007 | Edge-Based Color ConstancyabstractColor constancy is the ability to measure colors of objects independent of the color of the light source. A well-known color constancy method is based on the gray-world assumption which assumes that the average reflectance of surfaces in the world is achromatic. In this paper, we propose a new hypothesis for color constancy namely the gray-edge hypothesis, which assumes that the average edge difference in a scene is achromatic. Based on this hypothesis, we propose an algorithm for color constancy. Contrary to existing color constancy algorithms, which are computed from the zero-order structure of images, our method is based on the derivative structure of images. Furthermore, we propose a framework which unifies a variety of known (gray-world, max-RGB, Minkowski norm) and the newly proposed gray-edge and higher order gray-edge algorithms. The quality of the various instantiations of the framework is tested and compared to the state-of-the-art color constancy methods on two large data sets of images recording objects under a large number of different light sources. The experiments show that the proposed color constancy algorithms obtain comparable results as the state-of-the-art color constancy methods with the merit of being computationally more efficient. Joost van de Weijer 0001, Theo Gevers, Arjan Gijsenij |
IEEE Trans. Image Process. | 1 |
| 2006 | Coloring Local Feature Extraction
Joost van de Weijer 0001, Cordelia Schmid |
ECCV (2) | 1 |
| 2006 | Blur Robust and Color Constant Image DescriptionabstractAn important class of color constant image descriptors is based on image derivatives. These derivative-based image descriptors have a major drawback: they are sensitive to changes of image blur. Image blur has various causes such as being out-of-focus, motion of the camera or the object, and inaccurate acquisition settings. Since image blur is a frequently occurring image degradation, it is desirable for object description to be robust to its variations. We propose a set of descriptors which are both robust with respect to blurring effects, and invariant to illuminant color changes. Experiments on retrieval tasks show that the newly proposed object descriptors outperform existing descriptors in the presence of blurring effects. Joost van de Weijer 0001, Cordelia Schmid |
ICIP | 1 |
| 2006 | Boosting Color Saliency in Image Feature DetectionabstractThe aim of salient feature detection is to find distinctive local events in images. Salient features are generally determined from the local differential structure of images. They focus on the shape-saliency of the local neighborhood. The majority of these detectors are luminance-based, which has the disadvantage that the distinctiveness of the local color information is completely ignored in determining salient image features. To fully exploit the possibilities of salient point detection in color images, color distinctiveness should be taken into account in addition to shape distinctiveness. In this paper, color distinctiveness is explicitly incorporated into the design of saliency detection. The algorithm, called color saliency boosting, is based on an analysis of the statistics of color image derivatives. Color saliency boosting is designed as a generic method easily adaptable to existing feature detectors. Results show that substantial improvements in information content are acquired by targeting color salient features. Joost van de Weijer 0001, Theo Gevers, Andrew D. Bagdanov |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2006 | Robust photometric invariant features from the color tensorabstractLuminance-based features are widely used as low-level input for computer vision applications, even when color data is available. The extension of feature detection to the color domain prevents information loss due to isoluminance and allows us to exploit the photometric information. To fully exploit the extra information in the color data, the vector nature of color data has to be taken into account and a sound framework is needed to combine feature and photometric invariance theory. In this paper, we focus on the structure tensor, or color tensor, which adequately handles the vector nature of color images. Further, we combine the features based on the color tensor with photometric invariant derivatives to arrive at photometric invariant features. We circumvent the drawback of unstable photometric invariants by deriving an uncertainty measure to accompany the photometric invariant derivatives. The uncertainty is incorporated in the color tensor, hereby allowing the computation of robust photometric invariant features. The combination of the photometric invariance theory and tensor-based features allows for detection of a variety of features such as photometric invariant edges, corners, optical flow, and curvature. The proposed features are tested for noise characteristics and robustness to photometric changes. Experiments show that the proposed features are robust to scene incidental events and that the proposed uncertainty measure improves the applicability of full invariants. Joost van de Weijer 0001, Theo Gevers, Arnold W. M. Smeulders |
IEEE Trans. Image Process. | 1 |
| 2005 | Boosting Saliency in Color Image FeaturesabstractThe aim of salient point detection is to find distinctive events in images. Salient features are generally determined from the local differential structure of images. They focus on the shape saliency of the local neighborhood. The majority of these detectors is luminance based which has the disadvantage that the distinctiveness of the local color information is completely ignored. To fully exploit the possibilities of color image salient point detection, color distinctiveness should be taken into account next to shape distinctiveness. In this paper color distinctiveness is explicitly incorporated into the design of saliency detection. The algorithm, called color saliency boosting, is based on an analysis of the statistics of color image derivatives. Isosalient color derivatives can be closely approximated by ellipsoidal surfaces in color derivative space. Based on this remarkable statistical finding, isosalient derivatives are transformed by color boosting to have equal impact on the saliency. Color saliency boosting is designed as a generic method easily adaptable to existing feature detectors. Results show that substantial improvements in information content are acquired by targeting color salient features. Further, the generality of the method is illustrated by applying color boosting to multiple existing saliency methods. Joost van de Weijer 0001, Theo Gevers |
CVPR (1) | 1 |
| 2005 | Color constancy based on the Grey-edge hypothesisabstractA well-known color constancy method is based on the Grey-World assumption i.e. the average reflectance of surfaces in the world is achromatic. In this article we propose a new hypothesis for color constancy, namely the Grey-Edge hypothesis assuming that the average edge difference in a scene is achromatic. Based on this hypothesis, we propose an algorithm for color constancy. Recently, the Grey-World hypothesis and the max-RGB method were shown to be two instantiations of a Minkowski norm based color constancy method. Similarly we also propose a more general version of the Grey-Edge hypothesis which assumes that the Minkowsky norm of derivatives of the reflectance of surfaces is achromatic. The algorithms are tested on a large data set of images under different illuminants, and the results show that the new method outperforms the Grey-World assumption and the max-RGB method. Results are comparable to more elaborate algorithms, however at lower computational costs. Joost van de Weijer 0001, Theo Gevers |
ICIP (2) | 1 |
| 2005 | Least Squares and Robust Estimation of Local Image Structure
Joost van de Weijer 0001, Rein van den Boomgaard |
Int. J. Comput. Vis. | 1 |
| 2005 | Edge and Corner Detection by Photometric Quasi-InvariantsabstractFeature detection is used in many computer vision applications such as image segmentation, object recognition, and image retrieval. For these applications, robustness with respect to shadows, shading, and specularities is desired. Features based on derivatives of photometric invariants, which we will call full invariants, provide the desired robustness. However, because computation of photometric invariants involves nonlinear transformations, these features are Instable and, therefore, impractical for many applications. We propose a new class of derivatives which we refer to as quasi-invariants. These quasi-invariants are derivatives which share with full photometric invariants the property that they are insensitive for certain photometric edges, such as shadows or specular edges, but without the inherent instabilities of full photometric invariants. Experiments show that the quasi-invariant derivatives are less sensitive to noise and introduce less edge displacement than full invariant derivatives. Moreover, quasi-invariants significantly outperform the full invariant derivatives in terms of discriminative power. Joost van de Weijer 0001, Theo Gevers, Jan-Mark Geusebroek |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | Robust optical flow from photometric invariants
Joost van de Weijer 0001, Theo Gevers |
ICIP | 1 |
| 2003 | Color Edge Detection by Photometric Quasi-InvariantsabstractPhotometric invariance is used in many computer vision applications. The advantage of photometric invariance is the robustness against shadows, shading and illumination conditions. However, the drawbacks of photometric invariance are the loss of discriminative power and the inherent instabilities caused by the nonlinear transformations to compute the invariants. In this paper, we propose a new class of derivatives which we refer to as photometric quasi-invariants. These quasi-invariants share with full invariants the nice property that they are robust against photometric edges, such as shadows or specular edges. Further, these quasi-invariants do not have the inherent instabilities of full photometric invariants. We will apply these quasi-invariant derivatives in the context of photometric invariant edge detection and classification. Experiments show that the quasi-invariant derivatives are stable and they significantly outperform the full invariant derivatives in discriminative power. Joost van de Weijer 0001, Theo Gevers, Jan-Mark Geusebroek |
ICCV | 1 |
| 2003 | Fast anisotropic Gauss filteringabstractWe derive the decomposition of the anisotropic Gaussian in a one-dimensional (1-D) Gauss filter in the x-direction followed by a 1-D filter in a nonorthogonal direction phi. So also the anisotropic Gaussian can be decomposed by dimension. This appears to be extremely efficient from a computing perspective. An implementation scheme for normal convolution and for recursive filtering is proposed. Also directed derivative filters are demonstrated. For the recursive implementation, filtering an 512 x 512 image is performed within 40 msec on a current state of the art PC, gaining over 3 times in performance for a typical filter, independent of the standard deviations and orientation of the filter. Accuracy of the filters is still reasonable when compared to truncation error or recursive approximation error. The anisotropic Gaussian filtering method allows fast calculation of edge and ridge maps, with high spatial and angular accuracy. For tracking applications, the normal anisotropic convolution scheme is more advantageous, with applications in the detection of dashed lines in engineering drawings. The recursive implementation is more attractive in feature detection applications, for instance in affine invariant edge and ridge detection in computer vision. The proposed computational filtering method enables the practical applicability of orientation scale-space analysis. Jan-Mark Geusebroek, Arnold W. M. Smeulders, Joost van de Weijer 0001 |
IEEE Trans. Image Process. | 3 |
| 2002 | Fast Anisotropic Gauss Filtering
Jan-Mark Geusebroek, Arnold W. M. Smeulders, Joost van de Weijer 0001 |
ECCV (1) | 3 |
| 2001 | Local Mode FilteringabstractLinear filters have two major drawbacks. First, edges in the image are smoothed with increasing filter size. Second, by extending the filters to multi-channel data, correlation between the channels is lost. Only a few researchers have explored the possibilities of mode filtering to overcome these problems. Mode filtering is motivated from both a local histogram with tonal scale and a robust statistics point of view. The tonal scale is proved to be equal to the scale of the error norm function within the robust statistics framework. Instead of the more commonly studied global mode, our focus is on the local mode. It preserves edges and details and is easily extensible to multi-channel data. A generalization of the spatial Gaussian filtering to a spatial and tonal Gaussian filter is used to iterate to the local mode. Results on color images include successful noise attenuation while preserving edges and detail by local mode filtering. Joost van de Weijer 0001, Rein van den Boomgaard |
CVPR (2) | 1 |
| 2001 | Color mode filteringabstractIn this paper mode filtering of color images is explored. An existing framework based on local histograms is extended to multi-channel images. Within this framework three color mode operations are proposed; 1. Global mode operation for edge sharpening, noise reduction and small object removal, 2. Constrained mode operation for white noise filtering while preserving detail, and 3. Uncertain data mode filtering to incorporate prior knowledge about the certainty of the measurements into the mode computation. Results obtained for a variety of images indicate the feasibility of color mode filtering. Joost van de Weijer 0001, Theo Gevers |
ICIP (1) | 1 |
| 2001 | Curvature Estimation in Oriented Patterns Using Curvilinear Models Applied to Gradient Vector FieldsabstractCurved oriented patterns are dominated by high frequencies and exhibit zero gradients on ridges and valleys. Existing curvature estimators fail here. The characterization of curved oriented patterns based on translation invariance lacks an estimation of local curvature and yields a biased curvature-dependent confidence measure. Using parameterized curvilinear models we measure the amount of local gradient energy along the model gradient as a function of model curvature. Minimizing the residual energy yields a closed-form solution for the local curvature estimate and the corresponding confidence measure. We show that simple curvilinear models are applicable in the analysis of a wide variety of curved oriented patterns. Joost van de Weijer 0001, Lucas J. van Vliet, Piet W. Verbeek, Michael van Ginkel |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2000 | Colour Constancy from Hyper-Spectral DataabstractThis paper aims for color constant identification of object colors through the analysis of spectral color data. New computational color models are proposed which are not only invariant to illumination variations (color constancy) but also robust to a change in viewpoint and object geometry (color invariance). Color constancy and invariance is achieved by spectral imaging using a white reference, and based on color ratio’s (without a white reference). From the theoretical and experimental results it is concluded that the proposed computational methods for color constancy and invariance are highly robust to a change in SPD of the light source as well as a change in the pose of the object. Theo Gevers, Harro M. G. Stokman, Joost van de Weijer 0001 |
BMVC | 3 |
| 1998 | Improved curvature and anisotropy estimation for curved line bundlesabstractThe gradient-square tensor describes the orientation dependence of the squared directional derivative in images. The ratio of eigenvalues is a measure of local anisotropy. For an area showing shift invariance along some orientation (think of a piece of straight rail track) one of the tensor eigenvalues is zero. In practical situations (think of a piece of curved rail track) rotation invariance (perhaps around a remote center) occurs more often than shift invariance. Then curvature contributes to the smallest eigenvalue. In order to avoid this we deform a local area in such a way that the rotational symmetry becomes a translational one. Next the gradient square tensor defined on the transformed area, is expressed in derivatives of the original area. A curvature corrected anisotropy measure is defined. The correction turns out to be simple and straightforward. An average-curvature estimate for the area results as a valuable by product. Piet W. Verbeek, Lucas J. van Vliet, Joost van de Weijer 0001 |
ICPR | 3 |