EDBT 2026 Demo / reviewers in the wild / expert
Kai Wang 0060
dblp:78/2022-60
· DBLP profile ↗
27ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0002-9605-8279ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 5 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Leveraging Semantic Attribute Binding for Free-Lunch Color Control in Diffusion ModelsabstractRecent advances in text-to-image (T2I) diffusion models have enabled remarkable control over various attributes, yet precise color specification remains a fundamental challenge. Existing approaches, such as ColorPeel, rely on model personalization, requiring additional optimization and limiting flexibility in specifying arbitrary colors. In this work, we introduce ColorWave, a novel training-free approach that achieves exact RGB-level color control in diffusion models without fine-tuning. By systematically analyzing the cross-attention mechanisms within IP-Adapter, we uncover an implicit binding between textual color descriptors and reference image features. Leveraging this insight, our method rewires these bindings to enforce precise color attribution while preserving the generative capabilities of pretrained models. Our approach maintains generation quality and diversity, outperforming prior methods in accuracy and applicability across diverse object categories. Through extensive evaluations, we demonstrate that ColorWave establishes a new paradigm for structured, color-consistent diffusion-based image synthesis. Héctor Laria Mantecon, Alexandra Gomez-Villa, Jiang Qin, Muhammad Atif Butt, Bogdan Raducanu, Javier Vazquez-Corral, Joost van de Weijer 0001, Kai Wang 0060 |
WACV | 8 |
| 2026 | Training-free image inversion for one-step diffusion modelsabstractIn this work, we introduce a novel training-free inversion (TFinv) framework for one-step diffusion models, addressing key challenges in real image inversion and editing. We first identify two critical factors hampering real-image inversion and editing: (1) Initial Latent Editability, which is related to the distance between the initial noise and the ideal Gaussian distribution, and (2) Caption Gap, which means the alignment between text captions and image representations. Both factors influence inversion efficiency and the editability of one-step diffusion models. Then, we propose two novel techniques: iterative noise alignment (iterNA), which minimizes the distribution gap to align with the normal Gaussian distribution, and suffix learning (suffL), which enhances text-to-image caption alignment by introducing learned suffix prompt tokens. These techniques enable precise inversion of input images into their initial noise representations and facilitate image editing. Furthermore, we propose a mask-based editing technique for localized edits while preserving background integrity. Comprehensive experiments on the PIE-Bench dataset validate that our method TFinv not only achieves state-of-the-art performance in one-step diffusion editing, but also significantly outperforms existing multistep approaches in efficiency. Senmao Li, Yaxing Wang, Shiqi Yang 0002, Kai Wang 0060, Joost van de Weijer 0001 |
Pattern Recognit. | 5 |
| 2025 | The Art of Deception: Color Visual Illusions and Diffusion ModelsabstractVisual illusions in humans arise when interpreting out- of-distribution stimuli: if the observer is adapted to certain statistics, perception of outliers deviates from reality. Recent studies have shown that artificial neural networks (ANNs) can also be deceived by visual illusions. This revelation raises profound questions about the nature of visual information. Why are two independent systems, both human brains and ANNs, susceptible to the same illusions? Should any ANN be capable of perceiving visual illusions? Are these perceptions a feature or a flaw? In this work, we study how visual illusions are encoded in diffusion models. Remarkably, we show that they present human-like brightness/color shifts in their latent space. We use this fact to demonstrate that diffusion models can predict visual illusions. Furthermore, we also show how to generate new unseen visual illusions in realistic images using text-to-image diffusion models. We validate this ability through psychophysical experiments that show how our model-generated illusions also fool humans. Alexandra Gomez-Villa, Kai Wang 0060, C. Alejandro Párraga, Bartlomiej Twardowski, Jesús Malo, Javier Vazquez-Corral, Joost van de Weijer 0001 |
CVPR | 2 |
| 2025 | One-Way Ticket: Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion ModelsabstractText-to-Image (T2I) diffusion models have made remarkable advancements in generative modeling; however, they face a trade-off between inference speed and image quality, posing challenges for efficient deployment. Existing distilled T2I models can generate high-fidelity images with fewer sampling steps, but often struggle with diversity and quality, especially in one-step models. From our analysis, we observe redundant computations in the UNet encoders. Our findings suggest that, for T2I diffusion models, decoders are more adept at capturing richer and more explicit semantic information, while encoders can be effectively shared across decoders from diverse time steps. Based on these observations, we introduce the first Time-independent Unified Encoder (TiUE) for the student model UNet architecture, which is a loop-free image generation approach for distilling T2I diffusion models. Using a one-pass scheme, TiUE shares encoder features across multiple decoder time steps, enabling parallel sampling and significantly reducing inference time complexity. In addition, we incorporate a KL divergence term to regularize noise prediction, which enhances the perceptual realism and diversity of the generated images. Experimental results demonstrate that TiUE outperforms state-of-the-art methods, including LCM, SD-Turbo, and SwiftBrushv2, producing more diverse and realistic results while maintaining the computational efficiency. https://github.com/sen-mao/Loopfree Senmao Li, Lei Wang 0118, Kai Wang 0060, Jiehang Xie, Joost van de Weijer 0001, Fahad Shahbaz Khan, Shiqi Yang 0002, Yaxing Wang, Jian Yang 0003 |
CVPR | 3 |
| 2025 | Anchor Token Matching: Implicit Structure Locking for Training-Free AR Image EditingabstractText-to-image generation has seen groundbreaking advancements with diffusion models, enabling high-fidelity synthesis and precise image editing through cross-attention manipulation. Recently, autoregressive (AR) models have re-emerged as powerful alternatives, leveraging next-token generation to match diffusion models. However, existing editing techniques designed for diffusion models fail to translate directly to AR models due to fundamental differences in structural control. Specifically, AR models suffer from spatial poverty of attention maps and sequential accumulation of structural errors during image editing, which disrupt object layouts and global consistency. In this work, we introduce Implicit Structure Locking (ISLock), the first training-free editing strategy for AR visual models. Rather than relying on explicit attention manipulation or fine-tuning, ISLock preserves structural blueprints by dynamically aligning self-attention patterns with reference images through the Anchor Token Matching (ATM) protocol. By implicitly enforcing structural consistency in latent space, our method ISLock enables structure-aware editing while maintaining generative autonomy. Extensive experiments demonstrate that ISLock achieves high-quality, structure-consistent edits without additional training and is superior or comparable to conventional editing techniques. Our findings pioneer the way for efficient and flexible AR-based image editing, further bridging the performance gap between diffusion and autoregressive generative models. The code will be publicly available at https://github.com/hutaiHang/ATM Taihang Hu, Kai Wang 0060, Yaxing Wang, Jian Yang 0003, Ming-Ming Cheng |
ICCV | 3 |
| 2025 | InterLCM: Low-Quality Images as Intermediate States of Latent Consistency Models for Effective Blind Face RestorationabstractDiffusion priors have been used for blind face restoration (BFR) by fine-tuning diffusion models (DMs) on restoration datasets to recover low-quality images. However, the naive application of DMs presents several key limitations.
(i) The diffusion prior has inferior semantic consistency (e.g., ID, structure and color.), increasing the difficulty of optimizing the BFR model;
(ii) reliance on hundreds of denoising iterations, preventing the effective cooperation with perceptual losses, which is crucial for faithful restoration.
Observing that the latent consistency model (LCM) learns consistency noise-to-data mappings on the ODE-trajectory and therefore shows more semantic consistency in the subject identity, structural information and color preservation,
we propose $\textit{InterLCM}$ to leverage the LCM for its superior semantic consistency and efficiency to counter the above issues.
Treating low-quality images as the intermediate state of LCM, $\textit{InterLCM}$ achieves a balance between fidelity and quality by starting from earlier LCM steps.
LCM also allows the integration of perceptual loss during training, leading to improved restoration quality, particularly in real-world scenarios.
To mitigate structural and semantic uncertainties, $\textit{InterLCM}$ incorporates a Visual Module to extract visual features and a Spatial Encoder to capture spatial details, enhancing the fidelity of restored images.
Extensive experiments demonstrate that $\textit{InterLCM}$ outperforms existing approaches in both synthetic and real-world datasets while also achieving faster inference speed. Code and models will be publicly available. Senmao Li, Kai Wang 0060, Joost van de Weijer 0001, Fahad Shahbaz Khan, Chunle Guo, Shiqi Yang 0002, Yaxing Wang, Jian Yang 0003, Ming-Ming Cheng |
ICLR | 2 |
| 2025 | One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single PromptabstractText-to-image generation models can create high-quality images from input prompts. However, they struggle to support the consistent generation of identity-preserving requirements for storytelling. Existing approaches to this problem typically require extensive training in large datasets or additional modifications to the original model architectures. This limits their applicability across different domains and diverse diffusion model configurations. In this paper, we first observe the inherent capability of language models, coined $\textit{context consistency}$, to comprehend identity through context with a single prompt. Drawing inspiration from the inherent $\textit{context consistency}$, we propose a novel $\textit{training-free}$ method for consistent text-to-image (T2I) generation, termed "One-Prompt-One-Story" ($\textit{1Prompt1Story}$). Our approach $\textit{1Prompt1Story}$ concatenates all prompts into a single input for T2I diffusion models, initially preserving character identities. We then refine the generation process using two novel techniques: $\textit{Singular-Value
Reweighting}$ and $\textit{Identity-Preserving Cross-Attention}$, ensuring better alignment with the input description for each frame. In our experiments, we compare our method against various existing consistent T2I generation approaches to demonstrate its effectiveness, through quantitative metrics and qualitative assessments. Code is available at https://github.com/byliutao/1Prompt1Story. Kai Wang 0060, Senmao Li, Joost van de Weijer 0001, Fahad Shahbaz Khan, Shiqi Yang 0002, Yaxing Wang, Jian Yang 0003, Ming-Ming Cheng |
ICLR | 2 |
| 2025 | Covariances for Free: Exploiting Mean Distributions for Training-free Federated LearningabstractUsing pre-trained models has been found to reduce the effect of data heterogeneity and speed up federated learning algorithms. Recent works have explored training-free methods using first- and second-order statistics to aggregate local client data distributions at the server and achieve high performance without any training. In this work, we propose a training-free method based on an unbiased estimator of class covariance matrices which only uses first-order statistics in the form of class means communicated by clients to the server. We show how these estimated class covariances can be used to initialize the global classifier, thus exploiting the covariances without actually sharing them. We also show that using only within-class covariances results in a better classifier initialization. Our approach improves performance in the range of 4-26% with exactly the same communication cost when compared to methods sharing only class means and achieves performance competitive or superior to methods sharing second-order statistics with dramatically less communication overhead. The proposed method is much more communication-efficient than federated prompt-tuning methods and still outperforms them. Finally, using our method to initialize classifiers and then performing federated fine-tuning or linear probing again yields better performance. Code is available at https://github.com/dipamgoswami/FedCOF. Dipam Goswami, Simone Magistri, Kai Wang 0060, Bartlomiej Twardowski, Andrew D. Bagdanov, Joost van de Weijer 0001 |
NeurIPS | 3 |
| 2025 | From Cradle to Cane: A Two-Pass Framework for High-Fidelity Lifespan Face AgingabstractFace aging has become a crucial task in computer vision, with applications ranging from entertainment to healthcare. However, existing methods struggle with achieving a realistic and seamless transformation across the entire lifespan, especially when handling large age gaps or extreme head poses. The core challenge lies in balancing $age\ accuracy$ and $identity\ preservation$—what we refer to as the $Age\text{-}ID\ trade\text{-}off$. Most prior methods either prioritize age transformation at the expense of identity consistency or vice versa. In this work, we address this issue by proposing a $two\text{-}pass$ face aging framework, named $Cradle2Cane$, based on few-step text-to-image (T2I) diffusion models. The first pass focuses on solving $age\ accuracy$ by introducing an adaptive noise injection ($AdaNI$) mechanism. This mechanism is guided by including prompt descriptions of age and gender for the given person as the textual condition.
Also, by adjusting the noise level, we can control the strength of aging while allowing more flexibility in transforming the face.
However, identity preservation is weakly ensured here to facilitate stronger age transformations.
In the second pass, we enhance $identity\ preservation$ while maintaining age-specific features by conditioning the model on two identity-aware embeddings ($IDEmb$): $SVR\text{-}ArcFace$ and $Rotate\text{-}CLIP$. This pass allows for denoising the transformed image from the first pass, ensuring stronger identity preservation without compromising the aging accuracy.
Both passes are $jointly\ trained\ in\ an\ end\text{-}to\text{-}end\ way\$. Extensive experiments on the CelebA-HQ test dataset, evaluated through Face++ and Qwen-VL protocols, show that our $Cradle2Cane$ outperforms existing face aging methods in age accuracy and identity consistency.
Additionally, $Cradle2Cane$ demonstrates superior robustness when applied to in-the-wild human face images, where prior methods often fail. This significantly broadens its applicability to more diverse and unconstrained real-world scenarios. Code is available at https://github.com/byliutao/Cradle2Cane. Dafeng Zhang, Gengchen Li, Shizhuo Liu, Yongqi Song, Senmao Li, Shiqi Yang 0002, Boqian Li, Kai Wang 0060, Yaxing Wang |
NeurIPS | 9 |
| 2025 | Free-Lunch Color-Texture Disentanglement for Stylized Image GenerationabstractRecent advances in Text-to-Image (T2I) diffusion models have transformed image generation, enabling significant progress in stylized generation using only a few style reference images. However, current diffusion-based methods struggle with \textit{fine-grained} style customization due to challenges in controlling multiple style attributes, such as color and texture. This paper introduces the first tuning-free approach to achieve free-lunch color-texture disentanglement in stylized T2I generation, addressing the need for independently controlled style elements for the Disentangled Stylized Image Generation (DisIG) problem. Our approach leverages the \textit{Image-Prompt Additivity} property in the CLIP image embedding space to develop techniques for separating and extracting Color-Texture Embeddings (CTE) from individual color and texture reference images. To ensure that the color palette of the generated image aligns closely with the color reference, we apply a whitening and coloring transformation to enhance color consistency. Additionally, to prevent texture loss due to the signal-leak bias inherent in diffusion training, we introduce a noise term that preserves textural fidelity during the Regularized Whitening and Coloring Transformation (RegWCT). Through these methods, our Style Attributes Disentanglement approach (SADis) delivers a more precise and customizable solution for stylized image generation. Experiments on images from the WikiArt and StyleDrop datasets demonstrate that, both qualitatively and quantitatively, SADis surpasses state-of-the-art stylization methods in the DisIG task. Jiang Qin, Alexandra Gomez-Villa, Senmao Li, Shiqi Yang 0002, Yaxing Wang, Kai Wang 0060, Joost van de Weijer 0001 |
NeurIPS | 6 |
| 2025 | Multi-Class Textual-Inversion Secretly Yields a Semantic-Agnostic ClassifierabstractWith the advent of large pre-trained vision-language models such as CLIP, prompt learning methods aim to enhance the transferability of the CLIP model. They learn the prompt given few samples from the downstream task given the specific class names as prior knowledge, which we term as semantic-aware classification. However, in many realistic scenarios, we only have access to few samples and no knowledge of the class names (e.g., when considering instances of classes). This challenging scenario represents the semantic-agnostic discriminative case. Text-to-Image (T2I) personalization methods aim to adapt T2I models to unseen concepts by learning new tokens and endowing these tokens with the capability of generating the learned concepts. These methods do not require knowledge of class names as a semantic-aware prior. Therefore, in this paper, we first explore Textual Inversion and reveal that the new concept tokens possess both generation and classification capabilities by regarding each category as a single concept. However, learning classifiers from single-concept textual inversion is limited since the learned tokens are sub-optimal for the discriminative tasks. To mitigate this issue, we propose Multi-Class textual inversion, which includes a discriminative regularization term for the token updating process. Using this technique, our method MC-TI achieves stronger Semantic-Agnostic Classification while preserving the generation capability of these modifier tokens given only few samples per category. In the experiments, we extensively evaluate MC-TI on 12 datasets covering various scenarios, which demonstrates that MC-TI achieves superior results in terms of both classification and generation outcomes. Kai Wang 0060, Fei Yang 0004, Bogdan Raducanu, Joost van de Weijer 0001 |
WACV | 1 |
| 2024 | GRIF-DM: Generation of Rich Impression Fonts Using Diffusion ModelsabstractFonts are integral to creative endeavors, design processes, and artistic productions. The appropriate selection of a font can significantly enhance artwork and endow advertisements with a higher level of expressivity. Despite the availability of numerous diverse font designs online, traditional retrieval-based methods for font selection are increasingly being supplanted by generation-based approaches. These newer methods offer enhanced flexibility, catering to specific user preferences and capturing unique stylistic impressions. However, current impression font techniques based on Generative Adversarial Networks (GANs) necessitate the utilization of multiple auxiliary losses to provide guidance during generation. Furthermore, these methods commonly employ weighted summation for the fusion of impression-related keywords. This leads to generic vectors with the addition of more impression keywords, ultimately lacking in detail generation capacity. In this paper, we introduce a diffusion-based method, termed GRIF-DM, to generate fonts that vividly embody specific impressions, utilizing an input consisting of a single letter and a set of descriptive impression keywords. The core innovation of GRIF-DM lies in the development of dual cross-attention modules, which process the characteristics of the letters and impression keywords independently but synergistically, ensuring effective integration of both types of information. Our experimental results, conducted on the MyFonts dataset, affirm that this method is capable of producing realistic, vibrant, and high-fidelity fonts that are closely aligned with user specifications. This confirms the potential of our approach to revolutionize font generation by accommodating a broad spectrum of user-driven design requirements. Our code is publicly available at https://github.com/leitro/GRIF-DM. Lei Kang 0002, Fei Yang 0004, Kai Wang 0060, Mohamed Ali Souibgui, Lluís Gómez i Bigorda, Alicia Fornés, Ernest Valveny, Dimosthenis Karatzas |
ECAI | 3 |
| 2024 | ColorPeel: Color Prompt Learning with Diffusion Models via Color and Shape Disentanglement
Muhammad Atif Butt, Kai Wang 0060, Javier Vazquez-Corral, Joost van de Weijer 0001 |
ECCV (7) | 2 |
| 2024 | Exemplar-Free Continual Representation Learning via Learnable Drift Compensation
Alexandra Gomez-Villa, Dipam Goswami, Kai Wang 0060, Andrew D. Bagdanov, Bartlomiej Twardowski, Joost van de Weijer 0001 |
ECCV (7) | 3 |
| 2024 | IterInv: Iterative Inversion for Pixel-Level T2I ModelsabstractLarge-scale text-to-image diffusion models have been a ground-breaking development in generating convincing images following an input text prompt. The goal of image editing research is to give users control over the generated images by modifying the text prompt. Current image editing techniques predominantly hinge on DDIM inversion as a prevalent practice rooted in Latent Diffusion Models (LDM). However, the large pretrained T2I models working on the latent space suffer from losing details due to the first compression stage with an autoencoder mechanism. Instead, other mainstream T2I pipeline working on the pixel level, such as Imagen and DeepFloyd-IF, circumvents the above problem. They are commonly composed of multiple stages, typically starting with a text-to-image stage and followed by several super-resolution stages. In this pipeline, the DDIM inversion fails to find the initial noise and generate the original image given that the super-resolution diffusion models are not compatible with the DDIM technique. According to our experimental findings, iteratively concatenating the noisy image as the condition is the root of this problem. Based on this observation, we develop an iterative inversion (IterInv) technique for this category of T2I models and verify IterInv with the open-source DeepFloyd-IF model. Specifically, IterInv employ NTI as the inversion and reconstruction of low-resolution image generation. In stages 2 and 3, we update the latent variance at each timestep to find the deterministic inversion trace and promote the reconstruction process. By combining our method with a popular image editing method, we prove the application prospects of IterInv. The code will be released upon acceptance. The code is available at https://github.com/Tchuanm/IterInv.git Chuanming Tang, Kai Wang 0060, Joost van de Weijer 0001 |
ICME | 2 |
| 2024 | Token Merging for Training-Free Semantic Binding in Text-to-Image SynthesisabstractAlthough text-to-image (T2I) models exhibit remarkable generation capabilities,
they frequently fail to accurately bind semantically related objects or attributes
in the input prompts; a challenge termed semantic binding. Previous approaches
either involve intensive fine-tuning of the entire T2I model or require users or
large language models to specify generation layouts, adding complexity. In this
paper, we define semantic binding as the task of associating a given object with its
attribute, termed attribute binding, or linking it to other related sub-objects, referred
to as object binding. We introduce a novel method called Token Merging (ToMe),
which enhances semantic binding by aggregating relevant tokens into a single
composite token. This ensures that the object, its attributes and sub-objects all share
the same cross-attention map. Additionally, to address potential confusion among
main objects with complex textual prompts, we propose end token substitution as
a complementary strategy. To further refine our approach in the initial stages of
T2I generation, where layouts are determined, we incorporate two auxiliary losses,
an entropy loss and a semantic binding loss, to iteratively update the composite
token to improve the generation integrity. We conducted extensive experiments to
validate the effectiveness of ToMe, comparing it against various existing methods
on the T2I-CompBench and our proposed GPT-4o object binding benchmark. Our
method is particularly effective in complex scenarios that involve multiple objects
and attributes, which previous methods often fail to address. The code will be
publicly available at https://github.com/hutaihang/ToMe Taihang Hu, Joost van de Weijer 0001, Hongcheng Gao, Fahad Shahbaz Khan, Jian Yang 0003, Ming-Ming Cheng, Kai Wang 0060, Yaxing Wang |
NeurIPS | 8 |
| 2024 | Plasticity-Optimized Complementary Networks for Unsupervised Continual LearningabstractContinuous unsupervised representation learning (CURL) research has greatly benefited from improvements in self-supervised learning (SSL) techniques. As a result, existing CURL methods using SSL can learn high-quality representations without any labels, but with a notable performance drop when learning on a many-tasks data stream. We hypothesize that this is caused by the regularization losses that are imposed to prevent forgetting, leading to a suboptimal plasticity-stability trade-off: they either do not adapt fully to the incoming data (low plasticity), or incur significant forgetting when allowed to fully adapt to a new SSL pretext-task (low stability). In this work, we propose to train an expert network that is relieved of the duty of keeping the previous knowledge and can focus on performing optimally on the new tasks (optimizing plasticity). In the second phase, we combine this new knowledge with the previous network in an adaptation-retrospection phase to avoid forgetting and initialize a new expert with the knowledge of the old network. We perform several experiments showing that our proposed approach outperforms other CURL exemplar-free methods in few- and many-task split settings. Furthermore, we show how to adapt our approach to semi-supervised continual learning (Semi-SCL) and show that we surpass the accuracy of other exemplar-free Semi-SCL methods and reach the results of some others that use exemplars. Alexandra Gomez-Villa, Bartlomiej Twardowski, Kai Wang 0060, Joost van de Weijer 0001 |
WACV | 3 |
| 2024 | Diffusion-based network for unsupervised landmark detection
Kai Wang 0060, Chuanming Tang, Jianlin Zhang 0001 |
Knowl. Based Syst. | 2 |
| 2024 | Conditional Diffusion Model With Spatial-Frequency Refinement for SAR-to-Optical Image TranslationabstractThe presence of speckles and geometric distortions poses a serious challenge to the visual interpretation of synthetic aperture radar (SAR) images. SAR-to-optical (S2O) image translation technology provides a feasible solution and has attracted increasing attention. Restricted by substantial gaps between optical and SAR images, current S2O translation methods unavoidably result in geometric distortions, target missing, and generating low-fidelity images, thereby limiting subsequent cross-modal applications. In this article, we propose an augmented conditional denoising diffusion probabilistic model with spatial-frequency refinement (SFDiff) for high-fidelity S2O image translation. SFDiff progressively narrows the gap between synthesized and real images in both spatial and frequency perspectives, showcasing notable performance in terms of quality and consistency. Specifically, to incorporate rich spatial content priors provided by SAR images, we design an SAR context prior extractor (SCPE) with denoising enhancement to extract multiscale conditional representations, thereby aiding SFDiff in capturing more descriptive cues for S2O translation. In addition, a spatial-frequency complementary learning (SFCL) module is designed to learn spatial semantics and simultaneously enhances informative frequency components and global dependencies. Furthermore, SFDiff is optimized using the joint spatial-frequency refinement loss, facilitating iterative refinement in both spatial and frequency domains to enhance content consistency and fidelity in the synthesized images. Based on the experimental findings from the UNICORN dataset and the SEN12 dataset, SFDiff maintains a high level of content and structural consistency, resulting in visually appealing translation results that surpass the state-of-the-art (SOTA) methods. In particular, SFDiff exhibits excellent performance in preserving small targets and details, which is crucial in cross-modal detection applications. Jiang Qin, Kai Wang 0060, Bin Zou 0001, Lamei Zhang, Joost van de Weijer 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Dynamic Prompt Learning: Addressing Cross-Attention Leakage for Text-Based Image EditingabstractLarge-scale text-to-image generative models have been a ground-breaking development in generative AI, with diffusion models showing their astounding ability to synthesize convincing images following an input text prompt. The goal of image editing research is to give users control over the generated images by modifying the text prompt. Current image editing techniques are susceptible to unintended modifications of regions outside the targeted area, such as on the background or on distractor objects which have some semantic or visual relationship with the targeted object. According to our experimental findings, inaccurate cross-attention maps are at the root of this problem. Based on this observation, we propose $\textit{Dynamic Prompt Learning}$ ($DPL$) to force cross-attention maps to focus on correct $\textit{noun}$ words in the text prompt. By updating the dynamic tokens for nouns in the textual input with the proposed leakage repairment losses, we achieve fine-grained image editing over particular objects while preventing undesired changes to other image regions. Our method $DPL$, based on the publicly available $\textit{Stable Diffusion}$, is extensively evaluated on a wide range of images, and consistently obtains superior results both quantitatively (CLIP score, Structure-Dist) and qualitatively (on user-evaluation). We show improved prompt editing results for Word-Swap, Prompt Refinement, and Attention Re-weighting, especially for complex multi-object scenes. Kai Wang 0060, Fei Yang 0004, Shiqi Yang 0002, Muhammad Atif Butt, Joost van de Weijer 0001 |
NeurIPS | 1 |
| 2022 | Attention Distillation: self-supervised vision transformer students need more guidance
Kai Wang 0060, Fei Yang 0004, Joost van de Weijer 0001 |
BMVC | 1 |
| 2022 | Positive Pair Distillation Considered Harmful: Continual Meta Metric Learning for Lifelong Object Re-Identification
Kai Wang 0060, Chenshen Wu, Andrew D. Bagdanov, Xialei Liu, Shiqi Yang 0002, Shangling Jui, Joost van de Weijer 0001 |
BMVC | 1 |
| 2022 | Attracting and Dispersing: A Simple Approach for Source-free Domain AdaptationabstractWe propose a simple but effective source-free domain adaptation (SFDA) method. Treating SFDA as an unsupervised clustering problem and following the intuition that local neighbors in feature space should have more similar predictions than other features, we propose to optimize an objective of prediction consistency. This objective encourages local neighborhood features in feature space to have similar predictions while features farther away in feature space have dissimilar predictions, leading to efficient feature clustering and cluster assignment simultaneously. For efficient training, we seek to optimize an upper-bound of the objective resulting in two simple terms. Furthermore, we relate popular existing methods in domain adaptation, source-free domain adaptation and contrastive learning via the perspective of discriminability and diversity. The experimental results prove the superiority of our method, and our method can be adopted as a simple but strong baseline for future research in SFDA. Our method can be also adapted to source-free open-set and partial-set DA which further shows the generalization ability of our method. Code is available in https://github.com/Albert0147/AaD_SFDA. Shiqi Yang 0002, Yaxing Wang, Kai Wang 0060, Shangling Jui, Joost van de Weijer 0001 |
NeurIPS | 3 |
| 2021 | HCV: Hierarchy-Consistency Verification for Incremental Implicitly-Refined Classification
Kai Wang 0060, Xialei Liu, Luis Herranz, Joost van de Weijer 0001 |
BMVC | 1 |
| 2021 | ACAE-REMIND for online continual learning with compressed feature replay
Kai Wang 0060, Joost van de Weijer 0001, Luis Herranz |
Pattern Recognit. Lett. | 1 |
| 2021 | On Implicit Attribute Localization for Generalized Zero-Shot LearningabstractZero-shot learning (ZSL) aims to discriminate images from unseen classes by exploiting relations to seen classes via their attribute-based descriptions. Since attributes are often related to specific parts of objects, many recent works focus on discovering discriminative regions. However, these methods usually require additional complex part detection modules or attention mechanisms. In this paper, 1) we show that common ZSL backbones (without explicit attention nor part detection) can implicitly localize attributes, yet this property is not exploited. 2) Exploiting it, we then propose SELAR, a simple method that further encourages attribute localization, surprisingly achieving very competitive generalized ZSL (GZSL) performance when compared with more complex state-of-the-art methods. Our findings provide useful insight for designing future GZSL methods, and SELAR provides an easy to implement yet strong baseline. Shiqi Yang 0002, Kai Wang 0060, Luis Herranz, Joost van de Weijer 0001 |
IEEE Signal Process. Lett. | 2 |
| 2020 | Semantic Drift Compensation for Class-Incremental LearningabstractClass-incremental learning of deep networks sequentially increases the number of classes to be classified. During training, the network has only access to data of one task at a time, where each task contains several classes. In this setting, networks suffer from catastrophic forgetting which refers to the drastic drop in performance on previous tasks. The vast majority of methods have studied this scenario for classification networks, where for each new task the classification layer of the network must be augmented with additional weights to make room for the newly added classes. Embedding networks have the advantage that new classes can be naturally included into the network without adding new weights. Therefore, we study incremental learning for embedding networks. In addition, we propose a new method to estimate the drift, called semantic drift, of features and compensate for it without the need of any exemplars. We approximate the drift of previous tasks based on the drift that is experienced by current task data. We perform experiments on fine-grained datasets, CIFAR100 and ImageNet-Subset. We demonstrate that embedding networks suffer significantly less from catastrophic forgetting. We outperform existing methods which do not require exemplars and obtain competitive results compared to methods which store exemplars. Furthermore, we show that our proposed SDC when combined with existing methods to prevent forgetting consistently improves results. Lu Yu 0004, Bartlomiej Twardowski, Xialei Liu, Luis Herranz, Kai Wang 0060, Yongmei Cheng, Shangling Jui, Joost van de Weijer 0001 |
CVPR | 5 |