EDBT 2026 Demo / reviewers in the wild / expert
Shiqi Yang 0002
dblp:78/8501-2
· DBLP profile ↗
18ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0002-4141-7190ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 6 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MaTe3D: Mask-Guided Text-based 3D-Aware Portrait Editing
Kangneng Zhou, Daiheng Gao, Xuan Wang 0009, Jie Zhang 0090, Peng Zhang 0080, Xusen Sun, Longhao Zhang, Shiqi Yang 0002, Bang Zhang, Liefeng Bo, Yaxing Wang, Ming-Ming Cheng |
Int. J. Comput. Vis. | 8 |
| 2026 | Training-free image inversion for one-step diffusion modelsabstractIn this work, we introduce a novel training-free inversion (TFinv) framework for one-step diffusion models, addressing key challenges in real image inversion and editing. We first identify two critical factors hampering real-image inversion and editing: (1) Initial Latent Editability, which is related to the distance between the initial noise and the ideal Gaussian distribution, and (2) Caption Gap, which means the alignment between text captions and image representations. Both factors influence inversion efficiency and the editability of one-step diffusion models. Then, we propose two novel techniques: iterative noise alignment (iterNA), which minimizes the distribution gap to align with the normal Gaussian distribution, and suffix learning (suffL), which enhances text-to-image caption alignment by introducing learned suffix prompt tokens. These techniques enable precise inversion of input images into their initial noise representations and facilitate image editing. Furthermore, we propose a mask-based editing technique for localized edits while preserving background integrity. Comprehensive experiments on the PIE-Bench dataset validate that our method TFinv not only achieves state-of-the-art performance in one-step diffusion editing, but also significantly outperforms existing multistep approaches in efficiency. Senmao Li, Yaxing Wang, Shiqi Yang 0002, Kai Wang 0060, Joost van de Weijer 0001 |
Pattern Recognit. | 4 |
| 2025 | One-Way Ticket: Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion ModelsabstractText-to-Image (T2I) diffusion models have made remarkable advancements in generative modeling; however, they face a trade-off between inference speed and image quality, posing challenges for efficient deployment. Existing distilled T2I models can generate high-fidelity images with fewer sampling steps, but often struggle with diversity and quality, especially in one-step models. From our analysis, we observe redundant computations in the UNet encoders. Our findings suggest that, for T2I diffusion models, decoders are more adept at capturing richer and more explicit semantic information, while encoders can be effectively shared across decoders from diverse time steps. Based on these observations, we introduce the first Time-independent Unified Encoder (TiUE) for the student model UNet architecture, which is a loop-free image generation approach for distilling T2I diffusion models. Using a one-pass scheme, TiUE shares encoder features across multiple decoder time steps, enabling parallel sampling and significantly reducing inference time complexity. In addition, we incorporate a KL divergence term to regularize noise prediction, which enhances the perceptual realism and diversity of the generated images. Experimental results demonstrate that TiUE outperforms state-of-the-art methods, including LCM, SD-Turbo, and SwiftBrushv2, producing more diverse and realistic results while maintaining the computational efficiency. https://github.com/sen-mao/Loopfree Senmao Li, Lei Wang 0118, Kai Wang 0060, Jiehang Xie, Joost van de Weijer 0001, Fahad Shahbaz Khan, Shiqi Yang 0002, Yaxing Wang, Jian Yang 0003 |
CVPR | 8 |
| 2025 | InterLCM: Low-Quality Images as Intermediate States of Latent Consistency Models for Effective Blind Face RestorationabstractDiffusion priors have been used for blind face restoration (BFR) by fine-tuning diffusion models (DMs) on restoration datasets to recover low-quality images. However, the naive application of DMs presents several key limitations.
(i) The diffusion prior has inferior semantic consistency (e.g., ID, structure and color.), increasing the difficulty of optimizing the BFR model;
(ii) reliance on hundreds of denoising iterations, preventing the effective cooperation with perceptual losses, which is crucial for faithful restoration.
Observing that the latent consistency model (LCM) learns consistency noise-to-data mappings on the ODE-trajectory and therefore shows more semantic consistency in the subject identity, structural information and color preservation,
we propose $\textit{InterLCM}$ to leverage the LCM for its superior semantic consistency and efficiency to counter the above issues.
Treating low-quality images as the intermediate state of LCM, $\textit{InterLCM}$ achieves a balance between fidelity and quality by starting from earlier LCM steps.
LCM also allows the integration of perceptual loss during training, leading to improved restoration quality, particularly in real-world scenarios.
To mitigate structural and semantic uncertainties, $\textit{InterLCM}$ incorporates a Visual Module to extract visual features and a Spatial Encoder to capture spatial details, enhancing the fidelity of restored images.
Extensive experiments demonstrate that $\textit{InterLCM}$ outperforms existing approaches in both synthetic and real-world datasets while also achieving faster inference speed. Code and models will be publicly available. Senmao Li, Kai Wang 0060, Joost van de Weijer 0001, Fahad Shahbaz Khan, Chunle Guo, Shiqi Yang 0002, Yaxing Wang, Jian Yang 0003, Ming-Ming Cheng |
ICLR | 6 |
| 2025 | One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single PromptabstractText-to-image generation models can create high-quality images from input prompts. However, they struggle to support the consistent generation of identity-preserving requirements for storytelling. Existing approaches to this problem typically require extensive training in large datasets or additional modifications to the original model architectures. This limits their applicability across different domains and diverse diffusion model configurations. In this paper, we first observe the inherent capability of language models, coined $\textit{context consistency}$, to comprehend identity through context with a single prompt. Drawing inspiration from the inherent $\textit{context consistency}$, we propose a novel $\textit{training-free}$ method for consistent text-to-image (T2I) generation, termed "One-Prompt-One-Story" ($\textit{1Prompt1Story}$). Our approach $\textit{1Prompt1Story}$ concatenates all prompts into a single input for T2I diffusion models, initially preserving character identities. We then refine the generation process using two novel techniques: $\textit{Singular-Value
Reweighting}$ and $\textit{Identity-Preserving Cross-Attention}$, ensuring better alignment with the input description for each frame. In our experiments, we compare our method against various existing consistent T2I generation approaches to demonstrate its effectiveness, through quantitative metrics and qualitative assessments. Code is available at https://github.com/byliutao/1Prompt1Story. Kai Wang 0060, Senmao Li, Joost van de Weijer 0001, Fahad Shahbaz Khan, Shiqi Yang 0002, Yaxing Wang, Jian Yang 0003, Ming-Ming Cheng |
ICLR | 6 |
| 2025 | From Cradle to Cane: A Two-Pass Framework for High-Fidelity Lifespan Face AgingabstractFace aging has become a crucial task in computer vision, with applications ranging from entertainment to healthcare. However, existing methods struggle with achieving a realistic and seamless transformation across the entire lifespan, especially when handling large age gaps or extreme head poses. The core challenge lies in balancing $age\ accuracy$ and $identity\ preservation$—what we refer to as the $Age\text{-}ID\ trade\text{-}off$. Most prior methods either prioritize age transformation at the expense of identity consistency or vice versa. In this work, we address this issue by proposing a $two\text{-}pass$ face aging framework, named $Cradle2Cane$, based on few-step text-to-image (T2I) diffusion models. The first pass focuses on solving $age\ accuracy$ by introducing an adaptive noise injection ($AdaNI$) mechanism. This mechanism is guided by including prompt descriptions of age and gender for the given person as the textual condition.
Also, by adjusting the noise level, we can control the strength of aging while allowing more flexibility in transforming the face.
However, identity preservation is weakly ensured here to facilitate stronger age transformations.
In the second pass, we enhance $identity\ preservation$ while maintaining age-specific features by conditioning the model on two identity-aware embeddings ($IDEmb$): $SVR\text{-}ArcFace$ and $Rotate\text{-}CLIP$. This pass allows for denoising the transformed image from the first pass, ensuring stronger identity preservation without compromising the aging accuracy.
Both passes are $jointly\ trained\ in\ an\ end\text{-}to\text{-}end\ way\$. Extensive experiments on the CelebA-HQ test dataset, evaluated through Face++ and Qwen-VL protocols, show that our $Cradle2Cane$ outperforms existing face aging methods in age accuracy and identity consistency.
Additionally, $Cradle2Cane$ demonstrates superior robustness when applied to in-the-wild human face images, where prior methods often fail. This significantly broadens its applicability to more diverse and unconstrained real-world scenarios. Code is available at https://github.com/byliutao/Cradle2Cane. Dafeng Zhang, Gengchen Li, Shizhuo Liu, Yongqi Song, Senmao Li, Shiqi Yang 0002, Boqian Li, Kai Wang 0060, Yaxing Wang |
NeurIPS | 7 |
| 2025 | Free-Lunch Color-Texture Disentanglement for Stylized Image GenerationabstractRecent advances in Text-to-Image (T2I) diffusion models have transformed image generation, enabling significant progress in stylized generation using only a few style reference images. However, current diffusion-based methods struggle with \textit{fine-grained} style customization due to challenges in controlling multiple style attributes, such as color and texture. This paper introduces the first tuning-free approach to achieve free-lunch color-texture disentanglement in stylized T2I generation, addressing the need for independently controlled style elements for the Disentangled Stylized Image Generation (DisIG) problem. Our approach leverages the \textit{Image-Prompt Additivity} property in the CLIP image embedding space to develop techniques for separating and extracting Color-Texture Embeddings (CTE) from individual color and texture reference images. To ensure that the color palette of the generated image aligns closely with the color reference, we apply a whitening and coloring transformation to enhance color consistency. Additionally, to prevent texture loss due to the signal-leak bias inherent in diffusion training, we introduce a noise term that preserves textural fidelity during the Regularized Whitening and Coloring Transformation (RegWCT). Through these methods, our Style Attributes Disentanglement approach (SADis) delivers a more precise and customizable solution for stylized image generation. Experiments on images from the WikiArt and StyleDrop datasets demonstrate that, both qualitatively and quantitatively, SADis surpasses state-of-the-art stylization methods in the DisIG task. Jiang Qin, Alexandra Gomez-Villa, Senmao Li, Shiqi Yang 0002, Yaxing Wang, Kai Wang 0060, Joost van de Weijer 0001 |
NeurIPS | 4 |
| 2024 | Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model InferenceabstractOne of the main drawback of diffusion models is the slow inference time for image generation. Among the most successful approaches to addressing this problem are distillation methods. However, these methods require considerable computational resources. In this paper, we take another approach to diffusion model acceleration. We conduct a comprehensive study of the UNet encoder and empirically analyze the encoder features. This provides insights regarding their changes during the inference process. In particular, we find that encoder features change minimally, whereas the decoder features exhibit substantial variations across different time-steps. This insight motivates us to omit encoder computation at certain adjacent time-steps and reuse encoder features of previous time-steps as input to the decoder in multiple time-steps. Importantly, this allows us to perform decoder computation in parallel, further accelerating the denoising process. Additionally, we introduce a prior noise injection method to improve the texture details in the generated image. Besides the standard text-to-image task, we also validate our approach on other tasks: text-to-video, personalized generation and reference-guided generation. Without utilizing any knowledge distillation technique, our approach accelerates both the Stable Diffusion (SD) and DeepFloyd-IF model sampling by 41$\%$ and 24$\%$ respectively, and DiT model sampling by 34$\%$, while maintaining high-quality generation performance. Our code will be publicly released. Senmao Li, Taihang Hu, Joost van de Weijer 0001, Fahad Shahbaz Khan, Shiqi Yang 0002, Yaxing Wang, Ming-Ming Cheng, Jian Yang 0003 |
NeurIPS | 7 |
| 2023 | Dynamic Prompt Learning: Addressing Cross-Attention Leakage for Text-Based Image EditingabstractLarge-scale text-to-image generative models have been a ground-breaking development in generative AI, with diffusion models showing their astounding ability to synthesize convincing images following an input text prompt. The goal of image editing research is to give users control over the generated images by modifying the text prompt. Current image editing techniques are susceptible to unintended modifications of regions outside the targeted area, such as on the background or on distractor objects which have some semantic or visual relationship with the targeted object. According to our experimental findings, inaccurate cross-attention maps are at the root of this problem. Based on this observation, we propose $\textit{Dynamic Prompt Learning}$ ($DPL$) to force cross-attention maps to focus on correct $\textit{noun}$ words in the text prompt. By updating the dynamic tokens for nouns in the textual input with the proposed leakage repairment losses, we achieve fine-grained image editing over particular objects while preventing undesired changes to other image regions. Our method $DPL$, based on the publicly available $\textit{Stable Diffusion}$, is extensively evaluated on a wide range of images, and consistently obtains superior results both quantitatively (CLIP score, Structure-Dist) and qualitatively (on user-evaluation). We show improved prompt editing results for Word-Swap, Prompt Refinement, and Attention Re-weighting, especially for complex multi-object scenes. Kai Wang 0060, Fei Yang 0004, Shiqi Yang 0002, Muhammad Atif Butt, Joost van de Weijer 0001 |
NeurIPS | 3 |
| 2023 | Casting a BAIT for offline and online source-free domain adaptation
Shiqi Yang 0002, Yaxing Wang, Luis Herranz, Shangling Jui, Joost van de Weijer 0001 |
Comput. Vis. Image Underst. | 1 |
| 2023 | Trust Your Good Friends: Source-Free Domain Adaptation by Reciprocal Neighborhood ClusteringabstractDomain adaptation (DA) aims to alleviate the domain shift between source domain and target domain. Most DA methods require access to the source data, but often that is not possible (e.g., due to data privacy or intellectual property). In this paper, we address the challenging source-free domain adaptation (SFDA) problem, where the source pretrained model is adapted to the target domain in the absence of source data. Our method is based on the observation that target data, which might not align with the source domain classifier, still forms clear clusters. We capture this intrinsic structure by defining local affinity of the target data, and encourage label consistency among data with high local affinity. We observe that higher affinity should be assigned to reciprocal neighbors. To aggregate information with more context, we consider expanded neighborhoods with small affinity values. Furthermore, we consider the density around each target sample, which can alleviate the negative impact of potential outliers. In the experimental results we verify that the inherent structure of the target features is an important source of information for domain adaptation. We demonstrate that this local structure can be efficiently captured by considering the local neighbors, the reciprocal neighbors, and the expanded neighborhood. Finally, we achieve state-of-the-art performance on several 2D image and 3D point cloud recognition datasets. Shiqi Yang 0002, Yaxing Wang, Joost van de Weijer 0001, Luis Herranz, Shangling Jui, Jian Yang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Positive Pair Distillation Considered Harmful: Continual Meta Metric Learning for Lifelong Object Re-Identification
Kai Wang 0060, Chenshen Wu, Andrew D. Bagdanov, Xialei Liu, Shiqi Yang 0002, Shangling Jui, Joost van de Weijer 0001 |
BMVC | 5 |
| 2022 | Attracting and Dispersing: A Simple Approach for Source-free Domain AdaptationabstractWe propose a simple but effective source-free domain adaptation (SFDA) method. Treating SFDA as an unsupervised clustering problem and following the intuition that local neighbors in feature space should have more similar predictions than other features, we propose to optimize an objective of prediction consistency. This objective encourages local neighborhood features in feature space to have similar predictions while features farther away in feature space have dissimilar predictions, leading to efficient feature clustering and cluster assignment simultaneously. For efficient training, we seek to optimize an upper-bound of the objective resulting in two simple terms. Furthermore, we relate popular existing methods in domain adaptation, source-free domain adaptation and contrastive learning via the perspective of discriminability and diversity. The experimental results prove the superiority of our method, and our method can be adopted as a simple but strong baseline for future research in SFDA. Our method can be also adapted to source-free open-set and partial-set DA which further shows the generalization ability of our method. Code is available in https://github.com/Albert0147/AaD_SFDA. Shiqi Yang 0002, Yaxing Wang, Kai Wang 0060, Shangling Jui, Joost van de Weijer 0001 |
NeurIPS | 1 |
| 2021 | Generalized Source-free Domain AdaptationabstractDomain adaptation (DA) aims to transfer the knowledge learned from a source domain to an unlabeled target domain. Some recent works tackle source-free domain adaptation (SFDA) where only a source pre-trained model is available for adaptation to the target domain. However, those methods do not consider keeping source performance which is of high practical value in real world applications. In this paper, we propose a new domain adaptation paradigm called Generalized Source-free Domain Adaptation (G-SFDA), where the learned model needs to perform well on both the target and source domains, with only access to current unlabeled target data during adaptation. First, we propose local structure clustering (LSC), aiming to cluster the target features with its semantically similar neighbors, which successfully adapts the model to the target domain in the absence of source data. Second, we propose sparse domain attention (SDA), it produces a binary domain specific attention to activate different feature channels for different domains, meanwhile the domain attention will be utilized to regularize the gradient during adaptation to keep source information. In the experiments, for target performance our method is on par with or better than existing DA and SFDA methods, specifically it achieves state-of-the-art performance (85.4%) on VisDA, and our method works well for all domains after adapting to single or multiple target domains. Code is available in https://github.com/Albert0147/G-SFDA. Shiqi Yang 0002, Yaxing Wang, Joost van de Weijer 0001, Luis Herranz, Shangling Jui |
ICCV | 1 |
| 2021 | Exploiting the Intrinsic Neighborhood Structure for Source-free Domain AdaptationabstractDomain adaptation (DA) aims to alleviate the domain shift between source domain and target domain. Most DA methods require access to the source data, but often that is not possible (e.g. due to data privacy or intellectual property). In this paper, we address the challenging source-free domain adaptation (SFDA) problem, where the source pretrained model is adapted to the target domain in the absence of source data. Our method is based on the observation that target data, which might no longer align with the source domain classifier, still forms clear clusters. We capture this intrinsic structure by defining local affinity of the target data, and encourage label consistency among data with high local affinity. We observe that higher affinity should be assigned to reciprocal neighbors, and propose a self regularization loss to decrease the negative impact of noisy neighbors. Furthermore, to aggregate information with more context, we consider expanded neighborhoods with small affinity values. In the experimental results we verify that the inherent structure of the target features is an important source of information for domain adaptation. We demonstrate that this local structure can be efficiently captured by considering the local neighbors, the reciprocal neighbors, and the expanded neighborhood. Finally, we achieve state-of-the-art performance on several 2D image and 3D point cloud recognition datasets. Code is available in https://github.com/Albert0147/SFDA_neighbors. Shiqi Yang 0002, Yaxing Wang, Joost van de Weijer 0001, Luis Herranz, Shangling Jui |
NeurIPS | 1 |
| 2021 | Refine for Semantic Segmentation Based on Parallel Convolutional Network with Attention Model
Gang Peng 0002, Shiqi Yang 0002, Hao Wang 0123 |
Neural Process. Lett. | 2 |
| 2021 | On Implicit Attribute Localization for Generalized Zero-Shot LearningabstractZero-shot learning (ZSL) aims to discriminate images from unseen classes by exploiting relations to seen classes via their attribute-based descriptions. Since attributes are often related to specific parts of objects, many recent works focus on discovering discriminative regions. However, these methods usually require additional complex part detection modules or attention mechanisms. In this paper, 1) we show that common ZSL backbones (without explicit attention nor part detection) can implicitly localize attributes, yet this property is not exploited. 2) Exploiting it, we then propose SELAR, a simple method that further encourages attribute localization, surprisingly achieving very competitive generalized ZSL (GZSL) performance when compared with more complex state-of-the-art methods. Our findings provide useful insight for designing future GZSL methods, and SELAR provides an easy to implement yet strong baseline. Shiqi Yang 0002, Kai Wang 0060, Luis Herranz, Joost van de Weijer 0001 |
IEEE Signal Process. Lett. | 1 |
| 2018 | Parallel Convolutional Networks for Image Recognition via a Discriminator
Shiqi Yang 0002, Gang Peng 0002 |
ACCV (1) | 1 |