Yingying Deng

dblp:50/1407 · DBLP profile ↗
← Back
16ranked-venue papers
9as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Z-Magic: Zero-shot Multiple Attributes Guided Image Creator
abstract
The customization of multiple attributes has gained popularity with the rising demand for personalized content creation. Despite promising empirical results, the contextual coherence between different attributes has been largely overlooked. In this paper, we argue that subsequent attributes should follow the multivariable conditional distribution introduced by former attribute creation. In light of this, we reformulate multi-attribute creation from a conditional probability theory perspective and tackle the challenging zero-shot setting. By explicitly modeling the dependencies between attributes, we further enhance the coherence of generated images across diverse attribute combinations. Furthermore, we identify connections between multi-attribute customization and multi-task learning, effectively addressing the high computing cost encountered in multi-attribute synthesis. Extensive experiments demonstrate that Z-Magic outperforms existing models in zero-shot image generation, with broad implications for AI-driven design and creative applications.
Yingying Deng, Fan Tang, Weiming Dong
CVPR1
2025 Multi-Turn Consistent Image Editing
abstract
Many real-world applications, such as interactive photo retouching, artistic content creation, and product design, require flexible and iterative image editing. However, existing image editing methods primarily focus on achieving the desired modifications in a single step, which often struggles with ambiguous user intent, complex transformations, or the need for progressive refinements. As a result, these methods frequently produce inconsistent outcomes or fail to meet user expectations. To address these challenges, we propose a multi-turn image editing framework that enables users to iteratively refine their edits, progressively achieving more satisfactory results. Our approach leverages flow matching for accurate image inversion and a dual-objective Linear Quadratic Regulators (LQR) for stable sampling, effectively mitigating error accumulation. Additionally, by analyzing the layer-wise roles of transformers, we introduce a adaptive attention highlighting method that enhances editability while preserving multi-turn coherence. Extensive experiments demonstrate that our framework significantly improves edit success rates and visual fidelity compared to existing methods.
Zijun Zhou, Yingying Deng, Weiming Dong, Fan Tang
ICCV2
2025 FireFlow: Fast Inversion of Rectified Flow for Image Semantic Editing
abstract
Though Rectified Flows (ReFlows) with distillation offer a promising way for fast sampling, its fast inversion transforms images back to structured noise for recovery and following editing remains unsolved. This paper introduces FireFlow, an embarrassingly simple yet effective zero-shot approach that inherits the startling capacity of ReFlow-based models (such as FLUX) in generation while extending its capabilities to accurate inversion and editing in 8 steps. We first demonstrate that a carefully designed numerical solver is pivotal for ReFlow inversion, enabling accurate inversion and reconstruction with the precision of a second-order solver while maintaining the practical efficiency of a first-order Euler method. This solver achieves a $3\times$ runtime speedup compared to state-of-the-art ReFlow inversion and editing techniques while delivering smaller reconstruction errors and superior editing results in a training-free mode. The code is available at this-URL.
Yingying Deng, Changwang Mei, Fan Tang
ICML1
2025 A Novel Calibration of Precipitable Water Vapor From HY-2B Scanning Microwave Radiometer by Integrating GNSS and ERA5 Data for Offshore Areas
abstract
The Chinese Haiyang-2B (HY-2B) satellite is equipped with a scanning microwave radiometer (SMR) for global marine atmospheric water vapor detection. To calibrate precipitable water vapor (PWV) derived from the SMR accurately, a novel calibration method, referred to as Spatial Correction for Water Vapor (SC-WV), is proposed. This method integrates PWV data obtained from the Global Navigation Satellite System (GNSS) with the fifth generation of the European Centre for Medium-Range Weather Forecasts atmospheric reanalysis products (ERA5). By leveraging the ERA5 data, the spatial variations in water vapor surrounding the GNSS station are calculated, and the GNSS-derived PWV is subsequently corrected to obtain precise measurements at the SMR grid nodes. The calibration of SMR PWV within a 200 km radius around GNSS stations is conducted using GNSS PWV data observed at 47 island stations and 44 coastal stations from the International GNSS Service (IGS) in 2021. Both the proposed SC-WV method and the traditional inverse distance weighting (IDW) method are employed for comparison. The results demonstrate the applicability of the SC-WV method for both GNSS island and coastal stations, enabling the accurate calibration of the SMR PWV at each grid node surrounding the GNSS stations. On the basis of the GNSS PWV measurements obtained from island and coastal stations, the average bias values for the SMR PWV were 0.5 and 0.8 mm, respectively, with corresponding root mean square errors (RMSEs) of 2.6 and 2.8 mm respectively. The consistent accuracy indices observed provide evidence of the robustness and validity of the SC-WV method.
Shijie Fan, Jianfei Zang, Yingying Deng, Yanxiong Liu
IEEE Trans. Geosci. Remote. Sens.4
2024 Z*: Zero-shot Style Transfer via Attention Reweighting
abstract
Despite the remarkable progress in image style transfer, formulating style in the context of art is inherently subjective and challenging. In contrast to existing methods, this study shows that vanilla diffusion models can directly extract style information and seamlessly integrate the generative prior into the content image without retraining. Specifically, we adopt dual denoising paths to represent content/style references in latent space and then guide the content image denoising process with style latent codes. We further reveal that the cross-attention mechanism in latent diffusion models tends to blend the content and style images, resulting in stylized outputs that deviate from the original content image. To overcome this limitation, we introduce a cross-attention reweighting strategy. Through theoretical analysis and experiments, we demonstrate the effectiveness and superiority of the diffusion-based zero-shot §_tyle transfer via attention reweighting, Z -STAR.
Yingying Deng, Fan Tang, Weiming Dong
CVPR1
2024 CapHuman: Capture Your Moments in Parallel Universes
abstract
We concentrate on a novel human-centric image synthesis task, that is, given only one reference facial photograph, it is expected to generate specific individual images with diverse head positions, poses, facial expressions, and illuminations in different contexts. To accomplish this goal, we argue that our generative model should be capable of the following favorable characteristics: (1) a strong visual and semantic understanding of our world and human society for basic object and human image generation. (2) generalizable identity preservation ability. (3) flexible and fine-grained head control. Recently, large pre-trained text-to-image diffusion models have shown remarkable results, serving as a powerful generative foundation. As a basis, we aim to unleash the above two capabilities of the pre-trained model. In this work, we present a new framework named CapHuman. We embrace the “encode then learn to align” paradigm, which enables generalizable identity preservation for new individuals without cumbersome tuning at inference. CapHuman encodes identity features and then learns to align them into the latent space. Moreover, we introduce the 3D facial prior to equip our model with control over the human head in a flexible and 3D-consistent manner. Extensive qualitative and quantitative analyses demonstrate our CapHuman can produce well-identity-preserved, photo-realistic, and high-fidelity portraits with content-rich representations and various head renditions, superior to established baselines. Code and checkpoint will be released at https://github.com/VamosC/CapHuman.
Chao Liang 0002, Fan Ma, Linchao Zhu, Yingying Deng, Yi Yang 0001
CVPR4
2024 Exploring the Temporal Consistency of Arbitrary Style Transfer: A Channelwise Perspective
abstract
Arbitrary image stylization by neural networks has become a popular topic, and video stylization is attracting more attention as an extension of image stylization. However, when image stylization methods are applied to videos, unsatisfactory results that suffer from severe flickering effects appear. In this article, we conducted a detailed and comprehensive analysis of the cause of such flickering effects. Systematic comparisons among typical neural style transfer approaches show that the feature migration modules for state-of-the-art (SOTA) learning systems are ill-conditioned and could lead to a channelwise misalignment between the input content representations and the generated frames. Unlike traditional methods that relieve the misalignment via additional optical flow constraints or regularization modules, we focus on keeping the temporal consistency by aligning each output frame with the input frame. To this end, we propose a simple yet efficient multichannel correlation network (MCCNet), to ensure that output frames are directly aligned with inputs in the hidden feature space while maintaining the desired style patterns. An inner channel similarity loss is adopted to eliminate side effects caused by the absence of nonlinear operations such as softmax for strict alignment. Furthermore, to improve the performance of MCCNet under complex light conditions, we introduce an illumination loss during training. Qualitative and quantitative evaluations demonstrate that MCCNet performs well in arbitrary video and image style transfer tasks. Code is available at https://github.com/kongxiuxiu/MCCNetV2.
Xiaoyu Kong, Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Yongyong Chen, Zhenyu He 0001, Changsheng Xu
IEEE Trans. Neural Networks Learn. Syst.2
2022 StyTr2: Image Style Transfer with Transformers
abstract
The goal of image style transfer is to render an image with artistic features guided by a style reference while maintaining the original content. Owing to the locality in convolutional neural networks (CNNs), extracting and maintaining the global information of input images is difficult. Therefore, traditional neural style transfer methods face biased content representation. To address this critical issue, we take long-range dependencies of input images into account for image style transfer by proposing a transformer-based approach called StyTr2. In contrast with visual transformers for other vision tasks, StyTr2 contains two different transformer encoders to generate domain-specific sequences for content and style, respectively. Following the encoders, a multi-layer transformer decoder is adopted to stylize the content sequence according to the style sequence. We also analyze the deficiency of existing positional encoding methods and propose the content-aware positional encoding (CAPE), which is scale-invariant and more suitable for image style transfer tasks. Qualitative and quantitative experiments demonstrate the effectiveness of the proposed StyTr2 compared with state-of-the-art CNN-based and flow-based approaches. Code and models are available at https://github.com/diyiiyiii/StyTR-2.
Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Xingjia Pan, Changsheng Xu
CVPR1
2022 Transformers in computational visual media: A survey
abstract
Transformers, the dominant architecture for natural language processing, have also recently attracted much attention from computational visual media researchers due to their capacity for long-range representation and high performance. Transformers are sequence-to-sequence models, which use a self-attention mechanism rather than the RNN sequential structure. Thus, such models can be trained in parallel and can represent global information. This study comprehensively surveys recent visual transformer works. We categorize them according to task scenario: backbone design, high-level vision, low-level vision and generation, and multimodal learning. Their key ideas are also analyzed. Differing from previous surveys, we mainly focus on visual transformer methods in low-level vision and generation. The latest works on backbone design are also reviewed in detail. For ease of understanding, we precisely describe the main contributions of the latest works in the form of tables. As well as giving quantitative comparisons, we also present image results for low-level vision and generation tasks. Computational costs and source code links for various important works are also given in this survey to assist further development.
Yifan Xu 0008, HuaPeng Wei, Minxuan Lin, Yingying Deng, Kekai Sheng, Mengdan Zhang, Fan Tang, Weiming Dong, Feiyue Huang, Changsheng Xu
Comput. Vis. Media4
2022 A Comparative Study of CNN- and Transformer-Based Visual Style Transfer
HuaPeng Wei, Yingying Deng, Fan Tang, Xingjia Pan, Weiming Dong
J. Comput. Sci. Technol.2
2021 Arbitrary Video Style Transfer via Multi-Channel Correlation
abstract
Video style transfer is attracting increasing attention from the artificial intelligence community because of its numerous applications, such as augmented reality and animation production. Relative to traditional image style transfer, video style transfer presents new challenges, including how to effectively generate satisfactory stylized results for any specified style while maintaining temporal coherence across frames. Towards this end, we propose a Multi-Channel Correlation network (MCCNet), which can be trained to fuse exemplar style features and input content features for efficient style transfer while naturally maintaining the coherence of input videos to output videos. Specifically, MCCNet works directly on the feature space of style and content domain where it learns to rearrange and fuse style features on the basis of their similarity to content features. The outputs generated by MCC are features containing the desired style patterns that can further be decoded into images with vivid style textures. Moreover, MCCNet is also designed to explicitly align the features to input and thereby ensure that the outputs maintain the content structures and the temporal continuity. To further improve the performance of MCCNet under complex light conditions, we also introduce illumination loss during training. Qualitative and quantitative evaluations demonstrate that MCCNet performs well in arbitrary video and image style transfer tasks. Code is available at https://github.com/diyiiyiii/MCCNet.
Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Changsheng Xu
AAAI1
2021 Exploring the Representativity of Art Paintings
abstract
Art painting evaluation is sophisticated for a novice with no or limited knowledge on art criticism, and history. In this study, we propose the concept ofrepresentativityto evaluate paintings instead of using professional concepts, such as genre, media, and style, which may be confusing to non-professionals. We define the concept of representativity to evaluate quantitatively the extent to which a painting can represent the characteristics of an artists creations. We begin by proposing a novel deep representation of art paintings, which is enhanced by style information through a weighted pooling feature fusion module. In contrast to existing feature extraction approaches, the proposed framework embeds painting styles, and authorship information, and learns specific artwork characteristics in a single framework. Subsequently, we propose a graph-based learning method for representativity learning, which considers intra-category, and extra-category information. In view of the significance of historical factors in the art domain, we introduce the creation time of a painting into the learning process. User studies demonstrate our approach helps the public effectively access the creation characteristics of artists through sorting paintings by representativity from highest to lowest.
Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Feiyue Huang, Oliver Deussen, Changsheng Xu
IEEE Trans. Multim.1
2020 Arbitrary Style Transfer via Multi-Adaptation Network
abstract
Arbitrary style transfer is a significant topic with research value and application prospect. A desired style transfer, given a content image and referenced style painting, would render the content image with the color tone and vivid stroke patterns of the style painting while synchronously maintaining the detailed content structure information. Style transfer approaches would initially learn content and style representations of the content and style references and then generate the stylized images guided by these representations. In this paper, we propose the multi-adaptation network which involves two self-adaptation (SA) modules and one co-adaptation (CA) module:the SA modules adaptively disentangle the content and style representations, i.e., content SA module uses position-wise self-attention to enhance content representation and style SA module uses channel-wise self-attention to enhance style representation; the CA module rearranges the distribution of style representation based on content representation distribution by calculating the local similarity between the disentangled content and style features in a non-local fashion. Moreover, a new disentanglement loss function enables our network to extract main style patterns and exact content structures to adapt to various input images, respectively. Various qualitative and quantitative experiments demonstrate that the proposed multi-adaptation network leads to better results than the state-of-the-art style transfer methods.
Yingying Deng, Fan Tang, Weiming Dong, Feiyue Huang, Changsheng Xu
ACM Multimedia1
2019 Selective clustering for representative paintings selection
Yingying Deng, Fan Tang, Weiming Dong, Fuzhang Wu, Oliver Deussen, Changsheng Xu
Multim. Tools Appl.1
2013 AOPUT: A recommendation framework based on social activities and content interests
abstract
Content consuming and sharing are two most important user activities in social networking sites (SNSs). Lots of studies have been conducted on content recommendation using users' common interests. However, little has been done to help users to select friends and share content within their social networks. In this paper, we contribute a recommendation framework AOPUT to recommend both content and friend list for sharing to users leveraging content and social information in SNSs. It consists of two recommendation components: Recder and ShareAider. Recder generates content recommendations by connecting users with common interests. An improved Jaccard similarity is proposed to improve the Collaborative Filtering (CF) recommendation quality. ShareAider recommends a friend list to users when they want to share content with their friends. CF method and a social-based method are compared and the combination of them are explored to achieve better results. AOPUT is evaluated on a real world social network. The experimental results show that (1) Recder can provide better recommendation quality than the traditional CF method thanks to the improved Jaccard similarity; (2) social-based method performs better than CF since the sharing behavior in SNSs are highly dominated by users' social preferences, and the combination of these two methods performs better than each of them individually.
Yingying Deng, Tun Lu, Huanhuan Xia, Dongsheng Li 0002, Tiejiang Liu, Xianghua Ding, Ning Gu 0001
CSCWD1
2004 Possibilistic-clustering-based MR brain image segmentation with accurate initialization
abstract
Magnetic resonance image analysis by computer is useful to aid diagnosis of malady. We present in this paper a automatic segmentation method for principal brain tissues. It is based on the possibilistic clustering approach, which is an improved fuzzy c-means clustering method. In order to improve the efficiency of clustering process, the initial value problem is discussed and solved by combining with a histogram analysis method. Our method can automatically determine number of classes to cluster and the initial values for each class. It has been tested on a set of forty MR brain images with or without the presence of tumor. The experimental results showed that it is simple, rapid and robust to segment the principal brain tissues.
Qingmin Liao, Yingying Deng, Weibei Dou, Su Ruan, Daniel Bloyet
VCIP2