Yaxing Wang

dblp:166/2465 · DBLP profile ↗
← Back
43ranked-venue papers
12as first author
32since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 10 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 6 first-author · 11 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 StyleDiffusion: Prompt-Embedding Inversion for Text-Based Editing
abstract
A significant research effort is focused on exploiting the outstanding capacities of pretrained diffusion models for image editing. Approaches either fine tune the model, or invert the image in the latent space of the pretrained model. However, they suffer from two problems: (i) unsatisfactory results in selected regions and unexpected changes in non selected regions, and (ii) the need for careful text prompt editing: the prompt should include all visual objects in the input image. To address this, we propose two improvements: (i) only optimizing the input of the value linear network in the cross-attention layers is sufficiently powerful to reconstruct a real image, and (ii) attention regularization to preserve the object-like attention maps after reconstruction and editing, enabling accurate style editing without causing significant structural change. We further improve the editing technique used for the unconditional branch of classifier-free guidance as used by P2P. Extensive experimental prompt-editing results on a variety of images demonstrate qualitatively and quantitatively that our method has editing capabilities superior to those of existing and concurrent works. Our StyleDiffusion code is available at https://github.com/sen-mao/StyleDiffusion.
Senmao Li, Joost van de Weijer 0001, Taihang Hu, Fahad Shahbaz Khan, Qibin Hou, Yaxing Wang, Jian Yang 0003, Ming-Ming Cheng
Comput. Vis. Media6
2026 MaTe3D: Mask-Guided Text-based 3D-Aware Portrait Editing
Kangneng Zhou, Daiheng Gao, Xuan Wang 0009, Jie Zhang 0090, Peng Zhang 0080, Xusen Sun, Longhao Zhang, Shiqi Yang 0002, Bang Zhang, Liefeng Bo, Yaxing Wang, Ming-Ming Cheng
Int. J. Comput. Vis.11
2026 Training-free image inversion for one-step diffusion models
abstract
In this work, we introduce a novel training-free inversion (TFinv) framework for one-step diffusion models, addressing key challenges in real image inversion and editing. We first identify two critical factors hampering real-image inversion and editing: (1) Initial Latent Editability, which is related to the distance between the initial noise and the ideal Gaussian distribution, and (2) Caption Gap, which means the alignment between text captions and image representations. Both factors influence inversion efficiency and the editability of one-step diffusion models. Then, we propose two novel techniques: iterative noise alignment (iterNA), which minimizes the distribution gap to align with the normal Gaussian distribution, and suffix learning (suffL), which enhances text-to-image caption alignment by introducing learned suffix prompt tokens. These techniques enable precise inversion of input images into their initial noise representations and facilitate image editing. Furthermore, we propose a mask-based editing technique for localized edits while preserving background integrity. Comprehensive experiments on the PIE-Bench dataset validate that our method TFinv not only achieves state-of-the-art performance in one-step diffusion editing, but also significantly outperforms existing multistep approaches in efficiency.
Senmao Li, Yaxing Wang, Shiqi Yang 0002, Kai Wang 0060, Joost van de Weijer 0001
Pattern Recognit.3
2025 3DFaceController: Region-Controllable Face Synthesis via Decomposed and Recomposed Neural Radiance Fields
Kangneng Zhou, Yaxing Wang, Shuang Song 0005, Jie Zhang 0090, Ping Li 0016
CVM (2)2
2025 One-Way Ticket: Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models
abstract
Text-to-Image (T2I) diffusion models have made remarkable advancements in generative modeling; however, they face a trade-off between inference speed and image quality, posing challenges for efficient deployment. Existing distilled T2I models can generate high-fidelity images with fewer sampling steps, but often struggle with diversity and quality, especially in one-step models. From our analysis, we observe redundant computations in the UNet encoders. Our findings suggest that, for T2I diffusion models, decoders are more adept at capturing richer and more explicit semantic information, while encoders can be effectively shared across decoders from diverse time steps. Based on these observations, we introduce the first Time-independent Unified Encoder (TiUE) for the student model UNet architecture, which is a loop-free image generation approach for distilling T2I diffusion models. Using a one-pass scheme, TiUE shares encoder features across multiple decoder time steps, enabling parallel sampling and significantly reducing inference time complexity. In addition, we incorporate a KL divergence term to regularize noise prediction, which enhances the perceptual realism and diversity of the generated images. Experimental results demonstrate that TiUE outperforms state-of-the-art methods, including LCM, SD-Turbo, and SwiftBrushv2, producing more diverse and realistic results while maintaining the computational efficiency. https://github.com/sen-mao/Loopfree
Senmao Li, Lei Wang 0118, Kai Wang 0060, Jiehang Xie, Joost van de Weijer 0001, Fahad Shahbaz Khan, Shiqi Yang 0002, Yaxing Wang, Jian Yang 0003
CVPR9
2025 Not All Parameters Matter: Masking Diffusion Models for Enhancing Generation Ability
abstract
The diffusion models, in early stages focus on constructing basic image structures, while the refined details, including local features and textures, are generated in later stages. Thus the same network layers are forced to learn both structural and textural information simultaneously, significantly differing from the traditional deep learning architectures (e.g., ResNet or GANs) which captures or generates the image semantic information at different layers. This difference inspires us to explore the time-wise diffusion models. We initially investigate the key contributions of the U-Net parameters to the denoising process and identify that properly zeroing out certain parameters (including large parameters) contributes to denoising, substantially improving the generation quality on the fly. Capitalizing on this discovery, we propose a simple yet effective method—termed "MaskUNet"— that enhances generation quality with negligible parameter numbers. Our method fully leverages timestep- and sample-dependent effective U-Net parameters. To optimize MaskUNet, we offer two fine-tuning strategies: a training-based approach and a training-free approach, including tailored networks and optimization functions. In zero-shot inference on the COCO dataset, MaskUNet achieves the best FID score and further demonstrates its effectiveness in downstream task evaluations. Project page: https://gudaochangsheng.github.io/MaskUnet-Page/
Lei Wang 0118, Senmao Li, Jianye Wang, Yaxing Wang, Jian Yang 0003
CVPR7
2025 Anchor Token Matching: Implicit Structure Locking for Training-Free AR Image Editing
abstract
Text-to-image generation has seen groundbreaking advancements with diffusion models, enabling high-fidelity synthesis and precise image editing through cross-attention manipulation. Recently, autoregressive (AR) models have re-emerged as powerful alternatives, leveraging next-token generation to match diffusion models. However, existing editing techniques designed for diffusion models fail to translate directly to AR models due to fundamental differences in structural control. Specifically, AR models suffer from spatial poverty of attention maps and sequential accumulation of structural errors during image editing, which disrupt object layouts and global consistency. In this work, we introduce Implicit Structure Locking (ISLock), the first training-free editing strategy for AR visual models. Rather than relying on explicit attention manipulation or fine-tuning, ISLock preserves structural blueprints by dynamically aligning self-attention patterns with reference images through the Anchor Token Matching (ATM) protocol. By implicitly enforcing structural consistency in latent space, our method ISLock enables structure-aware editing while maintaining generative autonomy. Extensive experiments demonstrate that ISLock achieves high-quality, structure-consistent edits without additional training and is superior or comparable to conventional editing techniques. Our findings pioneer the way for efficient and flexible AR-based image editing, further bridging the performance gap between diffusion and autoregressive generative models. The code will be publicly available at https://github.com/hutaiHang/ATM
Taihang Hu, Kai Wang 0060, Yaxing Wang, Jian Yang 0003, Ming-Ming Cheng
ICCV4
2025 w+: Extending Classifier-Free Guidance in Diffusion Models for Real Image Inversion
Kaihua Li, Senmao Li, Yaxing Wang, Boqian Li, Gen Xu, Wanming Hao
ICIG (2)5
2025 InterLCM: Low-Quality Images as Intermediate States of Latent Consistency Models for Effective Blind Face Restoration
abstract
Diffusion priors have been used for blind face restoration (BFR) by fine-tuning diffusion models (DMs) on restoration datasets to recover low-quality images. However, the naive application of DMs presents several key limitations. (i) The diffusion prior has inferior semantic consistency (e.g., ID, structure and color.), increasing the difficulty of optimizing the BFR model; (ii) reliance on hundreds of denoising iterations, preventing the effective cooperation with perceptual losses, which is crucial for faithful restoration. Observing that the latent consistency model (LCM) learns consistency noise-to-data mappings on the ODE-trajectory and therefore shows more semantic consistency in the subject identity, structural information and color preservation, we propose $\textit{InterLCM}$ to leverage the LCM for its superior semantic consistency and efficiency to counter the above issues. Treating low-quality images as the intermediate state of LCM, $\textit{InterLCM}$ achieves a balance between fidelity and quality by starting from earlier LCM steps. LCM also allows the integration of perceptual loss during training, leading to improved restoration quality, particularly in real-world scenarios. To mitigate structural and semantic uncertainties, $\textit{InterLCM}$ incorporates a Visual Module to extract visual features and a Spatial Encoder to capture spatial details, enhancing the fidelity of restored images. Extensive experiments demonstrate that $\textit{InterLCM}$ outperforms existing approaches in both synthetic and real-world datasets while also achieving faster inference speed. Code and models will be publicly available.
Senmao Li, Kai Wang 0060, Joost van de Weijer 0001, Fahad Shahbaz Khan, Chunle Guo, Shiqi Yang 0002, Yaxing Wang, Jian Yang 0003, Ming-Ming Cheng
ICLR7
2025 One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt
abstract
Text-to-image generation models can create high-quality images from input prompts. However, they struggle to support the consistent generation of identity-preserving requirements for storytelling. Existing approaches to this problem typically require extensive training in large datasets or additional modifications to the original model architectures. This limits their applicability across different domains and diverse diffusion model configurations. In this paper, we first observe the inherent capability of language models, coined $\textit{context consistency}$, to comprehend identity through context with a single prompt. Drawing inspiration from the inherent $\textit{context consistency}$, we propose a novel $\textit{training-free}$ method for consistent text-to-image (T2I) generation, termed "One-Prompt-One-Story" ($\textit{1Prompt1Story}$). Our approach $\textit{1Prompt1Story}$ concatenates all prompts into a single input for T2I diffusion models, initially preserving character identities. We then refine the generation process using two novel techniques: $\textit{Singular-Value Reweighting}$ and $\textit{Identity-Preserving Cross-Attention}$, ensuring better alignment with the input description for each frame. In our experiments, we compare our method against various existing consistent T2I generation approaches to demonstrate its effectiveness, through quantitative metrics and qualitative assessments. Code is available at https://github.com/byliutao/1Prompt1Story.
Kai Wang 0060, Senmao Li, Joost van de Weijer 0001, Fahad Shahbaz Khan, Shiqi Yang 0002, Yaxing Wang, Jian Yang 0003, Ming-Ming Cheng
ICLR7
2025 From Cradle to Cane: A Two-Pass Framework for High-Fidelity Lifespan Face Aging
abstract
Face aging has become a crucial task in computer vision, with applications ranging from entertainment to healthcare. However, existing methods struggle with achieving a realistic and seamless transformation across the entire lifespan, especially when handling large age gaps or extreme head poses. The core challenge lies in balancing $age\ accuracy$ and $identity\ preservation$—what we refer to as the $Age\text{-}ID\ trade\text{-}off$. Most prior methods either prioritize age transformation at the expense of identity consistency or vice versa. In this work, we address this issue by proposing a $two\text{-}pass$ face aging framework, named $Cradle2Cane$, based on few-step text-to-image (T2I) diffusion models. The first pass focuses on solving $age\ accuracy$ by introducing an adaptive noise injection ($AdaNI$) mechanism. This mechanism is guided by including prompt descriptions of age and gender for the given person as the textual condition. Also, by adjusting the noise level, we can control the strength of aging while allowing more flexibility in transforming the face. However, identity preservation is weakly ensured here to facilitate stronger age transformations. In the second pass, we enhance $identity\ preservation$ while maintaining age-specific features by conditioning the model on two identity-aware embeddings ($IDEmb$): $SVR\text{-}ArcFace$ and $Rotate\text{-}CLIP$. This pass allows for denoising the transformed image from the first pass, ensuring stronger identity preservation without compromising the aging accuracy. Both passes are $jointly\ trained\ in\ an\ end\text{-}to\text{-}end\ way\$. Extensive experiments on the CelebA-HQ test dataset, evaluated through Face++ and Qwen-VL protocols, show that our $Cradle2Cane$ outperforms existing face aging methods in age accuracy and identity consistency. Additionally, $Cradle2Cane$ demonstrates superior robustness when applied to in-the-wild human face images, where prior methods often fail. This significantly broadens its applicability to more diverse and unconstrained real-world scenarios. Code is available at https://github.com/byliutao/Cradle2Cane.
Dafeng Zhang, Gengchen Li, Shizhuo Liu, Yongqi Song, Senmao Li, Shiqi Yang 0002, Boqian Li, Kai Wang 0060, Yaxing Wang
NeurIPS10
2025 Free-Lunch Color-Texture Disentanglement for Stylized Image Generation
abstract
Recent advances in Text-to-Image (T2I) diffusion models have transformed image generation, enabling significant progress in stylized generation using only a few style reference images. However, current diffusion-based methods struggle with \textit{fine-grained} style customization due to challenges in controlling multiple style attributes, such as color and texture. This paper introduces the first tuning-free approach to achieve free-lunch color-texture disentanglement in stylized T2I generation, addressing the need for independently controlled style elements for the Disentangled Stylized Image Generation (DisIG) problem. Our approach leverages the \textit{Image-Prompt Additivity} property in the CLIP image embedding space to develop techniques for separating and extracting Color-Texture Embeddings (CTE) from individual color and texture reference images. To ensure that the color palette of the generated image aligns closely with the color reference, we apply a whitening and coloring transformation to enhance color consistency. Additionally, to prevent texture loss due to the signal-leak bias inherent in diffusion training, we introduce a noise term that preserves textural fidelity during the Regularized Whitening and Coloring Transformation (RegWCT). Through these methods, our Style Attributes Disentanglement approach (SADis) delivers a more precise and customizable solution for stylized image generation. Experiments on images from the WikiArt and StyleDrop datasets demonstrate that, both qualitatively and quantitatively, SADis surpasses state-of-the-art stylization methods in the DisIG task.
Jiang Qin, Alexandra Gomez-Villa, Senmao Li, Shiqi Yang 0002, Yaxing Wang, Kai Wang 0060, Joost van de Weijer 0001
NeurIPS5
2025 MaskDiffusion: Boosting Text-to-Image Consistency with Conditional Mask
Yupeng Zhou, Daquan Zhou, Yaxing Wang, Jiashi Feng, Qibin Hou
Int. J. Comput. Vis.3
2024 Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models
abstract
The success of recent text-to-image diffusion models is largely due to their capacity to be guided by a complex text prompt, which enables users to precisely describe the desired content. However, these models struggle to effectively suppress the generation of undesired content, which is explicitly requested to be omitted from the generated image in the prompt. In this paper, we analyze how to manipulate the text embeddings and remove unwanted content from them. We introduce two contributions, which we refer to as soft-weighted regularization and inference-time text embedding optimization. The first regularizes the text embedding matrix and effectively suppresses the undesired content. The second method aims to further suppress the unwanted content generation of the prompt, and encourages the generation of desired content. We evaluate our method quantitatively and qualitatively on extensive experiments, validating its effectiveness. Furthermore, our method is generalizability to both the pixel-space diffusion models (i.e. DeepFloyd-IF) and the latent-space diffusion models (i.e. Stable Diffusion).
Senmao Li, Joost van de Weijer 0001, Taihang Hu, Fahad Shahbaz Khan, Qibin Hou, Yaxing Wang, Jian Yang 0003
ICLR6
2024 Token Merging for Training-Free Semantic Binding in Text-to-Image Synthesis
abstract
Although text-to-image (T2I) models exhibit remarkable generation capabilities, they frequently fail to accurately bind semantically related objects or attributes in the input prompts; a challenge termed semantic binding. Previous approaches either involve intensive fine-tuning of the entire T2I model or require users or large language models to specify generation layouts, adding complexity. In this paper, we define semantic binding as the task of associating a given object with its attribute, termed attribute binding, or linking it to other related sub-objects, referred to as object binding. We introduce a novel method called Token Merging (ToMe), which enhances semantic binding by aggregating relevant tokens into a single composite token. This ensures that the object, its attributes and sub-objects all share the same cross-attention map. Additionally, to address potential confusion among main objects with complex textual prompts, we propose end token substitution as a complementary strategy. To further refine our approach in the initial stages of T2I generation, where layouts are determined, we incorporate two auxiliary losses, an entropy loss and a semantic binding loss, to iteratively update the composite token to improve the generation integrity. We conducted extensive experiments to validate the effectiveness of ToMe, comparing it against various existing methods on the T2I-CompBench and our proposed GPT-4o object binding benchmark. Our method is particularly effective in complex scenarios that involve multiple objects and attributes, which previous methods often fail to address. The code will be publicly available at https://github.com/hutaihang/ToMe
Taihang Hu, Joost van de Weijer 0001, Hongcheng Gao, Fahad Shahbaz Khan, Jian Yang 0003, Ming-Ming Cheng, Kai Wang 0060, Yaxing Wang
NeurIPS9
2024 Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference
abstract
One of the main drawback of diffusion models is the slow inference time for image generation. Among the most successful approaches to addressing this problem are distillation methods. However, these methods require considerable computational resources. In this paper, we take another approach to diffusion model acceleration. We conduct a comprehensive study of the UNet encoder and empirically analyze the encoder features. This provides insights regarding their changes during the inference process. In particular, we find that encoder features change minimally, whereas the decoder features exhibit substantial variations across different time-steps. This insight motivates us to omit encoder computation at certain adjacent time-steps and reuse encoder features of previous time-steps as input to the decoder in multiple time-steps. Importantly, this allows us to perform decoder computation in parallel, further accelerating the denoising process. Additionally, we introduce a prior noise injection method to improve the texture details in the generated image. Besides the standard text-to-image task, we also validate our approach on other tasks: text-to-video, personalized generation and reference-guided generation. Without utilizing any knowledge distillation technique, our approach accelerates both the Stable Diffusion (SD) and DeepFloyd-IF model sampling by 41$\%$ and 24$\%$ respectively, and DiT model sampling by 34$\%$, while maintaining high-quality generation performance. Our code will be publicly released.
Senmao Li, Taihang Hu, Joost van de Weijer 0001, Fahad Shahbaz Khan, Shiqi Yang 0002, Yaxing Wang, Ming-Ming Cheng, Jian Yang 0003
NeurIPS8
2024 MineGAN++: Mining Generative Models for Efficient Knowledge Transfer to Limited Data Domains
Yaxing Wang, Abel Gonzalez-Garcia, Chenshen Wu, Luis Herranz, Fahad Shahbaz Khan, Shangling Jui, Jian Yang 0003, Joost van de Weijer 0001
Int. J. Comput. Vis.1
2024 Harnessing the potential of large language models in medical education: promise and pitfalls
abstract
OBJECTIVES: To provide balanced consideration of the opportunities and challenges associated with integrating Large Language Models (LLMs) throughout the medical school continuum. PROCESS: Narrative review of published literature contextualized by current reports of LLM application in medical education. CONCLUSIONS: LLMs like OpenAI's ChatGPT can potentially revolutionize traditional teaching methodologies. LLMs offer several potential advantages to students, including direct access to vast information, facilitation of personalized learning experiences, and enhancement of clinical skills development. For faculty and instructors, LLMs can facilitate innovative approaches to teaching complex medical concepts and fostering student engagement. Notable challenges of LLMs integration include the risk of fostering academic misconduct, inadvertent overreliance on AI, potential dilution of critical thinking skills, concerns regarding the accuracy and reliability of LLM-generated content, and the possible implications on teaching staff.
Trista M. Benítez, Yueyuan Xu, J. Donald Boudreau, Alfred Wei Chieh Kow, Fernando Bello, Le Van Phuoc, Gilberto Ka-Kit Leung, Yanyan Lan, Yaxing Wang, Davy Cheng, Tien Yin Wong, Kevin C. Chung
J. Am. Medical Informatics Assoc.11
2023 Adaptive Texture Filtering for Single-Domain Generalized Segmentation
abstract
Domain generalization in semantic segmentation aims to alleviate the performance degradation on unseen domains through learning domain-invariant features. Existing methods diversify images in the source domain by adding complex or even abnormal textures to reduce the sensitivity to domain-specific features. However, these approaches depends heavily on the richness of the texture bank and training them can be time-consuming. In contrast to importing textures arbitrarily or augmenting styles randomly, we focus on the single source domain itself to achieve the generalization. In this paper, we present a novel adaptive texture filtering mechanism to suppress the influence of texture without using augmentation, thus eliminating the interference of domain-specific features. Further, we design a hierarchical guidance generalization network equipped with structure-guided enhancement modules, which purpose to learn the domain-invariant generalized knowledge. Extensive experiments together with ablation studies on widely-used datasets are conducted to verify the effectiveness of the proposed model, and reveal its superiority over other state-of-the-art alternatives.
Mingjia Li 0001, Yaxing Wang, Chuan-Xian Ren, Xiaojie Guo 0001
AAAI3
2023 3D-Aware Multi-Class Image-to-Image Translation with NeRFs
abstract
Recent advances in 3D-aware generative models (3D-aware GANs) combined with Neural Radiance Fields (NeRF) have achieved impressive results. However no prior works investigate 3D-aware GANs for 3D consistent multiclass image-to-image (3D-aware 121) translation. Naively using 2D-121 translation methods suffers from unrealistic shape/identity change. To perform 3D-aware multiclass 121 translation, we decouple this learning process into a multiclass 3D-aware GAN step and a 3D-aware 121 translation step. In the first step, we propose two novel techniques: a new conditional architecture and an effective training strategy. In the second step, based on the well-trained multiclass 3D-aware GAN architecture, that preserves view-consistency, we construct a 3D-aware 121 translation system. To further reduce the view-consistency problems, we propose several new techniques, including a U-net-like adaptor network design, a hierarchical representation constrain and a relative regularization loss. In exten-sive experiments on two datasets, quantitative and qualitative results demonstrate that we successfully perform 3D-aware 121 translation with multi-view consistency. Code is available in 3DI2I.
Senmao Li, Joost van de Weijer 0001, Yaxing Wang, Fahad Shahbaz Khan, Meiqin Liu 0002, Jian Yang 0003
CVPR3
2023 Task Offloading and Resource Allocation for Edge-Cloud Collaborative Computing
Yaxing Wang, Jia Hao 0006, Baoqi Huang
ICA3PP (5)1
2023 Provable Multi-instance Deep AUC Maximization with Stochastic Pooling
abstract
This paper considers a novel application of deep AUC maximization (DAM) for multi-instance learning (MIL), in which a single class label is assigned to a bag of instances (e.g., multiple 2D slices of a CT scan for a patient). We address a neglected yet non-negligible computational challenge of MIL in the context of DAM, i.e., bag size is too large to be loaded into GPU memory for backpropagation, which is required by the standard pooling methods of MIL. To tackle this challenge, we propose variance-reduced stochastic pooling methods in the spirit of stochastic optimization by formulating the loss function over the pooled prediction as a multi-level compositional function. By synthesizing techniques from stochastic compositional optimization and non-convex min-max optimization, we propose a unified and provable muli-instance DAM (MIDAM) algorithm with stochastic smoothed-max pooling or stochastic attention-based pooling, which only samples a few instances for each bag to compute a stochastic gradient estimator and to update the model parameter. We establish a similar convergence rate of the proposed MIDAM algorithm as the state-of-the-art DAM algorithms. Our extensive experiments on conventional MIL datasets and medical datasets demonstrate the superiority of our MIDAM algorithm. The method is open-sourced at https://libauc.org/.
Dixian Zhu, Bokun Wang, Zhi Chen 0025, Yaxing Wang, Milan Sonka, Xiaodong Wu 0001, Tianbao Yang
ICML4
2023 Casting a BAIT for offline and online source-free domain adaptation
Shiqi Yang 0002, Yaxing Wang, Luis Herranz, Shangling Jui, Joost van de Weijer 0001
Comput. Vis. Image Underst.2
2023 Trust Your Good Friends: Source-Free Domain Adaptation by Reciprocal Neighborhood Clustering
abstract
Domain adaptation (DA) aims to alleviate the domain shift between source domain and target domain. Most DA methods require access to the source data, but often that is not possible (e.g., due to data privacy or intellectual property). In this paper, we address the challenging source-free domain adaptation (SFDA) problem, where the source pretrained model is adapted to the target domain in the absence of source data. Our method is based on the observation that target data, which might not align with the source domain classifier, still forms clear clusters. We capture this intrinsic structure by defining local affinity of the target data, and encourage label consistency among data with high local affinity. We observe that higher affinity should be assigned to reciprocal neighbors. To aggregate information with more context, we consider expanded neighborhoods with small affinity values. Furthermore, we consider the density around each target sample, which can alleviate the negative impact of potential outliers. In the experimental results we verify that the inherent structure of the target features is an important source of information for domain adaptation. We demonstrate that this local structure can be efficiently captured by considering the local neighbors, the reciprocal neighbors, and the expanded neighborhood. Finally, we achieve state-of-the-art performance on several 2D image and 3D point cloud recognition datasets.
Shiqi Yang 0002, Yaxing Wang, Joost van de Weijer 0001, Luis Herranz, Shangling Jui, Jian Yang 0003
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 MVMO: A Multi-Object Dataset for Wide Baseline Multi-View Semantic Segmentation
abstract
We present MVMO (Multi-View, Multi-Object dataset): a synthetic dataset of 116,000 scenes containing randomly placed objects of 10 distinct classes and captured from 25 camera locations in the upper hemisphere. MVMO comprises photorealistic, path-traced image renders, together with semantic segmentation ground truth for every view. Unlike existing multi-view datasets, MVMO features wide baselines between cameras and high density of objects, which lead to large disparities, heavy occlusions and view-dependent object appearance. Single view semantic segmentation is hindered by self and inter-object occlusions that could benefit from additional viewpoints. Therefore, we expect that MVMO will propel research in multi-view semantic segmentation and cross-view semantic transfer. We also provide baselines that show that new research is needed in such fields to exploit the complementary information of multi-view setups1.
Aitor Alvarez-Gila, Joost van de Weijer 0001, Yaxing Wang, Estíbaliz Garrote
ICIP3
2022 Distilling GANs with Style-Mixed Triplets for X2I Translation with Limited Data
Yaxing Wang, Joost van de Weijer 0001, Lu Yu 0004, Shangling Jui
ICLR1
2022 Attracting and Dispersing: A Simple Approach for Source-free Domain Adaptation
abstract
We propose a simple but effective source-free domain adaptation (SFDA) method. Treating SFDA as an unsupervised clustering problem and following the intuition that local neighbors in feature space should have more similar predictions than other features, we propose to optimize an objective of prediction consistency. This objective encourages local neighborhood features in feature space to have similar predictions while features farther away in feature space have dissimilar predictions, leading to efficient feature clustering and cluster assignment simultaneously. For efficient training, we seek to optimize an upper-bound of the objective resulting in two simple terms. Furthermore, we relate popular existing methods in domain adaptation, source-free domain adaptation and contrastive learning via the perspective of discriminability and diversity. The experimental results prove the superiority of our method, and our method can be adopted as a simple but strong baseline for future research in SFDA. Our method can be also adapted to source-free open-set and partial-set DA which further shows the generalization ability of our method. Code is available in https://github.com/Albert0147/AaD_SFDA.
Shiqi Yang 0002, Yaxing Wang, Kai Wang 0060, Shangling Jui, Joost van de Weijer 0001
NeurIPS2
2022 A novel framework for image-to-image translation and image compression
Fei Yang 0004, Yaxing Wang, Luis Herranz, Yongmei Cheng, Mikhail G. Mozerov
Neurocomputing2
2021 TransferI2I: Transfer Learning for Image-to-Image Translation from Small Datasets
abstract
Image-to-image (I2I) translation has matured in recent years and is able to generate high-quality realistic images. However, despite current success, it still faces important challenges when applied to small domains. Existing methods use transfer learning for I2I translation, but they still require the learning of millions of parameters from scratch. This drawback severely limits its application on small domains. In this paper, we propose a new transfer learning for I2I translation (TransferI2I). We decouple our learning process into the image generation step and the I2I translation step. In the first step we propose two novel techniques: source-target initialization and self-initialization of the adaptor layer. The former finetunes the pretrained generative model (e.g., StyleGAN) on source and target data. The latter allows to initialize all non-pretrained network parameters without the need of any data. These techniques provide a better initialization for the I2I translation step. In addition, we introduce an auxiliary GAN that further facilitates the training of deep I2I systems even from small datasets. In extensive experiments on three datasets, (Animal faces, Birds, and Foods), we show that we outperform existing methods and that mFID improves on several datasets with over 25 points. Our code is available at: https://github.com/yaxingwang/TransferI2I.
Yaxing Wang, Héctor Laria Mantecon, Joost van de Weijer 0001, Laura Lopez-Fuentes, Bogdan Raducanu
ICCV1
2021 Generalized Source-free Domain Adaptation
abstract
Domain adaptation (DA) aims to transfer the knowledge learned from a source domain to an unlabeled target domain. Some recent works tackle source-free domain adaptation (SFDA) where only a source pre-trained model is available for adaptation to the target domain. However, those methods do not consider keeping source performance which is of high practical value in real world applications. In this paper, we propose a new domain adaptation paradigm called Generalized Source-free Domain Adaptation (G-SFDA), where the learned model needs to perform well on both the target and source domains, with only access to current unlabeled target data during adaptation. First, we propose local structure clustering (LSC), aiming to cluster the target features with its semantically similar neighbors, which successfully adapts the model to the target domain in the absence of source data. Second, we propose sparse domain attention (SDA), it produces a binary domain specific attention to activate different feature channels for different domains, meanwhile the domain attention will be utilized to regularize the gradient during adaptation to keep source information. In the experiments, for target performance our method is on par with or better than existing DA and SFDA methods, specifically it achieves state-of-the-art performance (85.4%) on VisDA, and our method works well for all domains after adapting to single or multiple target domains. Code is available in https://github.com/Albert0147/G-SFDA.
Shiqi Yang 0002, Yaxing Wang, Joost van de Weijer 0001, Luis Herranz, Shangling Jui
ICCV2
2021 Exploiting the Intrinsic Neighborhood Structure for Source-free Domain Adaptation
abstract
Domain adaptation (DA) aims to alleviate the domain shift between source domain and target domain. Most DA methods require access to the source data, but often that is not possible (e.g. due to data privacy or intellectual property). In this paper, we address the challenging source-free domain adaptation (SFDA) problem, where the source pretrained model is adapted to the target domain in the absence of source data. Our method is based on the observation that target data, which might no longer align with the source domain classifier, still forms clear clusters. We capture this intrinsic structure by defining local affinity of the target data, and encourage label consistency among data with high local affinity. We observe that higher affinity should be assigned to reciprocal neighbors, and propose a self regularization loss to decrease the negative impact of noisy neighbors. Furthermore, to aggregate information with more context, we consider expanded neighborhoods with small affinity values. In the experimental results we verify that the inherent structure of the target features is an important source of information for domain adaptation. We demonstrate that this local structure can be efficiently captured by considering the local neighbors, the reciprocal neighbors, and the expanded neighborhood. Finally, we achieve state-of-the-art performance on several 2D image and 3D point cloud recognition datasets. Code is available in https://github.com/Albert0147/SFDA_neighbors.
Shiqi Yang 0002, Yaxing Wang, Joost van de Weijer 0001, Luis Herranz, Shangling Jui
NeurIPS2
2021 Controlling biases and diversity in diverse image-to-image translation
Yaxing Wang, Abel Gonzalez-Garcia, Luis Herranz, Joost van de Weijer 0001
Comput. Vis. Image Underst.1
2020 MineGAN: Effective Knowledge Transfer From GANs to Target Domains With Few Images
abstract
One of the attractive characteristics of deep neural networks is their ability to transfer knowledge obtained in one domain to other related domains. As a result, high-quality networks can be trained in domains with relatively little training data. This property has been extensively studied for discriminative networks but has received significantly less attention for generative models. Given the often enormous effort required to train GANs, both computationally as well as in the dataset collection, the re-use of pretrained GANs is a desirable objective. We propose a novel knowledge transfer method for generative models based on mining the knowledge that is most beneficial to a specific target domain, either from a single or multiple pretrained GANs. This is done using a miner network that identifies which part of the generative distribution of each pretrained GAN outputs samples closest to the target domain. Mining effectively steers GAN sampling towards suitable regions of the latent space, which facilitates the posterior finetuning and avoids pathologies of other methods such as mode collapse and lack of flexibility. We perform experiments on several complex datasets using various GAN architectures (BigGAN, Progressive GAN) and show that the proposed method, called MineGAN, effectively transfers knowledge to domains with few target images, outperforming existing methods. In addition, MineGAN can successfully transfer knowledge from multiple pretrained GANs. Our code is available at: \url{https://github.com/yaxingwang/MineGAN}.
Yaxing Wang, Abel Gonzalez-Garcia, David Berga, Luis Herranz, Fahad Shahbaz Khan, Joost van de Weijer 0001
CVPR1
2020 Semi-Supervised Learning for Few-Shot Image-to-Image Translation
abstract
In the last few years, unpaired image-to-image translation has witnessed Remarkable progress. Although the latest methods are able to generate realistic images, they crucially rely on a large number of labeled images. Recently, some methods have tackled the challenging setting of few-shot image-to-image ranslation, reducing the labeled data requirements for the target domain during inference. In this work, we go one step further and reduce the amount of required labeled data also from the source domain during training. To do so, we propose applying semi-supervised learning via a noise-tolerant pseudo-labeling procedure. We also apply a cycle consistency constraint to further exploit the information from unlabeled images, either from the same dataset or external. Additionally, we propose several structural modifications to facilitate the image translation task under these circumstances. Our semi-supervised method for few-shot image translation, called \emph{SEMIT}, achieves excellent results on four different datasets using as little as 10\% of the source labels, and matches the performance of the main fully-supervised competitor using only 20\% labeled data. Our code and models are made public at: \url{https://github.com/yaxingwang/SEMIT}.
Yaxing Wang, Salman Khan 0001, Abel Gonzalez-Garcia, Joost van de Weijer 0001, Fahad Shahbaz Khan
CVPR1
2020 GANwriting: Content-Conditioned Generation of Styled Handwritten Word Images
Lei Kang 0002, Pau Riba, Yaxing Wang, Marçal Rusiñol, Alicia Fornés, Mauricio Villegas
ECCV (23)3
2020 DeepI2I: Enabling Deep Hierarchical Image-to-Image Translation by Transferring from GANs
abstract
Image-to-image translation has recently achieved remarkable results. But despite current success, it suffers from inferior performance when translations between classes require large shape changes. We attribute this to the high-resolution bottlenecks which are used by current state-of-the-art image-to-image methods. Therefore, in this work, we propose a novel deep hierarchical Image-to-Image Translation method, called DeepI2I. We learn a model by leveraging hierarchical features: (a) structural information contained in the bottom layers and (b) semantic information extracted from the top layers. To enable the training of deep I2I models on small datasets, we propose a novel transfer learning method, that transfers knowledge from pre-trained GANs. Specifically, we leverage the discriminator of a pre-trained GANs (i.e. BigGAN or StyleGAN) to initialize both the encoder and the discriminator and the pre-trained generator to initialize the generator of our model. Applying knowledge transfer leads to an alignment problem between the encoder and generator. We introduce an adaptor network to address this. On many-class image-to-image translation on three datasets (Animal faces, Birds, and Foods) we decrease mFID by at least 35% when compared to the state-of-the-art. Furthermore, we qualitatively and quantitatively demonstrate that transfer learning significantly improves the performance of I2I systems, especially for small datasets. Finally, we are the first to perform I2I translations for domains with over 100 classes.
Yaxing Wang, Lu Yu 0004, Joost van de Weijer 0001
NeurIPS1
2020 Mix and Match Networks: Cross-Modal Alignment for Zero-Pair Image-to-Image Translation
Yaxing Wang, Luis Herranz, Joost van de Weijer 0001
Int. J. Comput. Vis.1
2019 SDIT: Scalable and Diverse Cross-domain Image Translation
abstract
Recently, image-to-image translation research has witnessed remarkable progress. Although current approaches successfully generate diverse outputs or perform scalable image transfer, these properties have not been combined into a single method. To address this limitation, we propose SDIT: Scalable and Diverse image-to-image translation. These properties are combined into a single generator. The diversity is determined by a latent variable which is randomly sampled from a normal distribution. The scalability is obtained by conditioning the network on the domain attributes. Additionally, we also exploit an attention mechanism that permits the generator to focus on the domain-specific attribute. We empirically demonstrate the performance of the proposed method on face mapping and other datasets beyond faces.
Yaxing Wang, Abel Gonzalez-Garcia, Joost van de Weijer 0001, Luis Herranz
ACM Multimedia1
2018 Mix and Match Networks: Encoder-Decoder Alignment for Zero-Pair Image Translation
abstract
We address the problem of image translation between domains or modalities for which no direct paired data is available (i.e. zero-pair translation). We propose mix and match networks, based on multiple encoders and decoders aligned in such a way that other encoder-decoder pairs can be composed at test time to perform unseen image translation tasks between domains or modalities for which explicit paired samples were not seen during training. We study the impact of autoencoders, side information and losses in improving the alignment and transferability of trained pairwise translation models to unseen translations. We show our approach is scalable and can perform colorization and style transfer between unseen combinations of domains. We evaluate our system in a challenging cross-modal setting where semantic segmentation is estimated from depth images, without explicit access to any depth-semantic segmentation training pairs. Our model outperforms baselines based on pix2pix and CycleGAN models.
Yaxing Wang, Joost van de Weijer 0001, Luis Herranz
CVPR1
2018 Transferring GANs: Generating Images from Limited Data
Yaxing Wang, Chenshen Wu, Luis Herranz, Joost van de Weijer 0001, Abel Gonzalez-Garcia, Bogdan Raducanu
ECCV (6)1
2018 Memory Replay GANs: Learning to Generate New Categories without Forgetting
abstract
Previous works on sequential learning address the problem of forgetting in discriminative models. In this paper we consider the case of generative models. In particular, we investigate generative adversarial networks (GANs) in the task of learning new categories in a sequential fashion. We first show that sequential fine tuning renders the network unable to properly generate images from previous categories (i.e. forgetting). Addressing this problem, we propose Memory Replay GANs (MeRGANs), a conditional GAN framework that integrates a memory replay generator. We study two methods to prevent forgetting by leveraging these replays, namely joint training with replay and replay alignment. Qualitative and quantitative experimental results in MNIST, SVHN and LSUN datasets show that our memory replay approach can generate competitive images while significantly mitigating the forgetting of previous categories.
Chenshen Wu, Luis Herranz, Xialei Liu, Yaxing Wang, Joost van de Weijer 0001, Bogdan Raducanu
NeurIPS4
2015 Automated Layer Segmentation of 3D Macular Images Using Hybrid Methods
Chuang Wang 0005, Yaxing Wang, Djibril Kaba, Zidong Wang 0001, Xiaohui Liu 0001, Yongmin Li 0001
ICIG (1)2
2015 Segmentation of Intra-retinal Layers in 3D Optic Nerve Head Images
Chuang Wang 0005, Yaxing Wang, Djibril Kaba, Haogang Zhu, Zidong Wang 0001, Xiaohui Liu 0001, Yongmin Li 0001
ICIG (3)2