EDBT 2026 Demo / reviewers in the wild / expert
Yingying Deng
dblp:50/1407
· DBLP profile ↗
16ranked-venue papers
9as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Z-Magic: Zero-shot Multiple Attributes Guided Image CreatorabstractThe customization of multiple attributes has gained popularity with the rising demand for personalized content creation. Despite promising empirical results, the contextual coherence between different attributes has been largely overlooked. In this paper, we argue that subsequent attributes should follow the multivariable conditional distribution introduced by former attribute creation. In light of this, we reformulate multi-attribute creation from a conditional probability theory perspective and tackle the challenging zero-shot setting. By explicitly modeling the dependencies between attributes, we further enhance the coherence of generated images across diverse attribute combinations. Furthermore, we identify connections between multi-attribute customization and multi-task learning, effectively addressing the high computing cost encountered in multi-attribute synthesis. Extensive experiments demonstrate that Z-Magic outperforms existing models in zero-shot image generation, with broad implications for AI-driven design and creative applications. Yingying Deng, Fan Tang, Weiming Dong |
CVPR | 1 |
| 2025 | Multi-Turn Consistent Image EditingabstractMany real-world applications, such as interactive photo retouching, artistic content creation, and product design, require flexible and iterative image editing. However, existing image editing methods primarily focus on achieving the desired modifications in a single step, which often struggles with ambiguous user intent, complex transformations, or the need for progressive refinements. As a result, these methods frequently produce inconsistent outcomes or fail to meet user expectations. To address these challenges, we propose a multi-turn image editing framework that enables users to iteratively refine their edits, progressively achieving more satisfactory results. Our approach leverages flow matching for accurate image inversion and a dual-objective Linear Quadratic Regulators (LQR) for stable sampling, effectively mitigating error accumulation. Additionally, by analyzing the layer-wise roles of transformers, we introduce a adaptive attention highlighting method that enhances editability while preserving multi-turn coherence. Extensive experiments demonstrate that our framework significantly improves edit success rates and visual fidelity compared to existing methods. Zijun Zhou, Yingying Deng, Weiming Dong, Fan Tang |
ICCV | 2 |
| 2025 | FireFlow: Fast Inversion of Rectified Flow for Image Semantic EditingabstractThough Rectified Flows (ReFlows) with distillation offer a promising way for fast sampling, its fast inversion transforms images back to structured noise for recovery and following editing remains unsolved. This paper introduces FireFlow, an embarrassingly simple yet effective zero-shot approach that inherits the startling capacity of ReFlow-based models (such as FLUX) in generation while extending its capabilities to accurate inversion and editing in 8 steps. We first demonstrate that a carefully designed numerical solver is pivotal for ReFlow inversion, enabling accurate inversion and reconstruction with the precision of a second-order solver while maintaining the practical efficiency of a first-order Euler method. This solver achieves a $3\times$ runtime speedup compared to state-of-the-art ReFlow inversion and editing techniques while delivering smaller reconstruction errors and superior editing results in a training-free mode. The code is available at this-URL. Yingying Deng, Changwang Mei, Fan Tang |
ICML | 1 |
| 2025 | A Novel Calibration of Precipitable Water Vapor From HY-2B Scanning Microwave Radiometer by Integrating GNSS and ERA5 Data for Offshore AreasabstractThe Chinese Haiyang-2B (HY-2B) satellite is equipped with a scanning microwave radiometer (SMR) for global marine atmospheric water vapor detection. To calibrate precipitable water vapor (PWV) derived from the SMR accurately, a novel calibration method, referred to as Spatial Correction for Water Vapor (SC-WV), is proposed. This method integrates PWV data obtained from the Global Navigation Satellite System (GNSS) with the fifth generation of the European Centre for Medium-Range Weather Forecasts atmospheric reanalysis products (ERA5). By leveraging the ERA5 data, the spatial variations in water vapor surrounding the GNSS station are calculated, and the GNSS-derived PWV is subsequently corrected to obtain precise measurements at the SMR grid nodes. The calibration of SMR PWV within a 200 km radius around GNSS stations is conducted using GNSS PWV data observed at 47 island stations and 44 coastal stations from the International GNSS Service (IGS) in 2021. Both the proposed SC-WV method and the traditional inverse distance weighting (IDW) method are employed for comparison. The results demonstrate the applicability of the SC-WV method for both GNSS island and coastal stations, enabling the accurate calibration of the SMR PWV at each grid node surrounding the GNSS stations. On the basis of the GNSS PWV measurements obtained from island and coastal stations, the average bias values for the SMR PWV were 0.5 and 0.8 mm, respectively, with corresponding root mean square errors (RMSEs) of 2.6 and 2.8 mm respectively. The consistent accuracy indices observed provide evidence of the robustness and validity of the SC-WV method. Shijie Fan, Jianfei Zang, Yingying Deng, Yanxiong Liu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Z*: Zero-shot Style Transfer via Attention ReweightingabstractDespite the remarkable progress in image style transfer, formulating style in the context of art is inherently subjective and challenging. In contrast to existing methods, this study shows that vanilla diffusion models can directly extract style information and seamlessly integrate the generative prior into the content image without retraining. Specifically, we adopt dual denoising paths to represent content/style references in latent space and then guide the content image denoising process with style latent codes. We further reveal that the cross-attention mechanism in latent diffusion models tends to blend the content and style images, resulting in stylized outputs that deviate from the original content image. To overcome this limitation, we introduce a cross-attention reweighting strategy. Through theoretical analysis and experiments, we demonstrate the effectiveness and superiority of the diffusion-based zero-shot §_tyle transfer via attention reweighting, Z -STAR. Yingying Deng, Fan Tang, Weiming Dong |
CVPR | 1 |
| 2024 | CapHuman: Capture Your Moments in Parallel UniversesabstractWe concentrate on a novel human-centric image synthesis task, that is, given only one reference facial photograph, it is expected to generate specific individual images with diverse head positions, poses, facial expressions, and illuminations in different contexts. To accomplish this goal, we argue that our generative model should be capable of the following favorable characteristics: (1) a strong visual and semantic understanding of our world and human society for basic object and human image generation. (2) generalizable identity preservation ability. (3) flexible and fine-grained head control. Recently, large pre-trained text-to-image diffusion models have shown remarkable results, serving as a powerful generative foundation. As a basis, we aim to unleash the above two capabilities of the pre-trained model. In this work, we present a new framework named CapHuman. We embrace the “encode then learn to align” paradigm, which enables generalizable identity preservation for new individuals without cumbersome tuning at inference. CapHuman encodes identity features and then learns to align them into the latent space. Moreover, we introduce the 3D facial prior to equip our model with control over the human head in a flexible and 3D-consistent manner. Extensive qualitative and quantitative analyses demonstrate our CapHuman can produce well-identity-preserved, photo-realistic, and high-fidelity portraits with content-rich representations and various head renditions, superior to established baselines. Code and checkpoint will be released at https://github.com/VamosC/CapHuman. Chao Liang 0002, Fan Ma, Linchao Zhu, Yingying Deng, Yi Yang 0001 |
CVPR | 4 |
| 2024 | Exploring the Temporal Consistency of Arbitrary Style Transfer: A Channelwise PerspectiveabstractArbitrary image stylization by neural networks has become a popular topic, and video stylization is attracting more attention as an extension of image stylization. However, when image stylization methods are applied to videos, unsatisfactory results that suffer from severe flickering effects appear. In this article, we conducted a detailed and comprehensive analysis of the cause of such flickering effects. Systematic comparisons among typical neural style transfer approaches show that the feature migration modules for state-of-the-art (SOTA) learning systems are ill-conditioned and could lead to a channelwise misalignment between the input content representations and the generated frames. Unlike traditional methods that relieve the misalignment via additional optical flow constraints or regularization modules, we focus on keeping the temporal consistency by aligning each output frame with the input frame. To this end, we propose a simple yet efficient multichannel correlation network (MCCNet), to ensure that output frames are directly aligned with inputs in the hidden feature space while maintaining the desired style patterns. An inner channel similarity loss is adopted to eliminate side effects caused by the absence of nonlinear operations such as softmax for strict alignment. Furthermore, to improve the performance of MCCNet under complex light conditions, we introduce an illumination loss during training. Qualitative and quantitative evaluations demonstrate that MCCNet performs well in arbitrary video and image style transfer tasks. Code is available at https://github.com/kongxiuxiu/MCCNetV2. Xiaoyu Kong, Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Yongyong Chen, Zhenyu He 0001, Changsheng Xu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | StyTr2: Image Style Transfer with TransformersabstractThe goal of image style transfer is to render an image with artistic features guided by a style reference while maintaining the original content. Owing to the locality in convolutional neural networks (CNNs), extracting and maintaining the global information of input images is difficult. Therefore, traditional neural style transfer methods face biased content representation. To address this critical issue, we take long-range dependencies of input images into account for image style transfer by proposing a transformer-based approach called StyTr2. In contrast with visual transformers for other vision tasks, StyTr2 contains two different transformer encoders to generate domain-specific sequences for content and style, respectively. Following the encoders, a multi-layer transformer decoder is adopted to stylize the content sequence according to the style sequence. We also analyze the deficiency of existing positional encoding methods and propose the content-aware positional encoding (CAPE), which is scale-invariant and more suitable for image style transfer tasks. Qualitative and quantitative experiments demonstrate the effectiveness of the proposed StyTr2 compared with state-of-the-art CNN-based and flow-based approaches. Code and models are available at https://github.com/diyiiyiii/StyTR-2. Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Xingjia Pan, Changsheng Xu |
CVPR | 1 |
| 2022 | Transformers in computational visual media: A surveyabstractTransformers, the dominant architecture for natural language processing, have also recently attracted much attention from computational visual media researchers due to their capacity for long-range representation and high performance. Transformers are sequence-to-sequence models, which use a self-attention mechanism rather than the RNN sequential structure. Thus, such models can be trained in parallel and can represent global information. This study comprehensively surveys recent visual transformer works. We categorize them according to task scenario: backbone design, high-level vision, low-level vision and generation, and multimodal learning. Their key ideas are also analyzed. Differing from previous surveys, we mainly focus on visual transformer methods in low-level vision and generation. The latest works on backbone design are also reviewed in detail. For ease of understanding, we precisely describe the main contributions of the latest works in the form of tables. As well as giving quantitative comparisons, we also present image results for low-level vision and generation tasks. Computational costs and source code links for various important works are also given in this survey to assist further development. Yifan Xu 0008, HuaPeng Wei, Minxuan Lin, Yingying Deng, Kekai Sheng, Mengdan Zhang, Fan Tang, Weiming Dong, Feiyue Huang, Changsheng Xu |
Comput. Vis. Media | 4 |
| 2022 | A Comparative Study of CNN- and Transformer-Based Visual Style Transfer
HuaPeng Wei, Yingying Deng, Fan Tang, Xingjia Pan, Weiming Dong |
J. Comput. Sci. Technol. | 2 |
| 2021 | Arbitrary Video Style Transfer via Multi-Channel CorrelationabstractVideo style transfer is attracting increasing attention from the artificial intelligence community because of its numerous applications, such as augmented reality and animation production. Relative to traditional image style transfer, video style transfer presents new challenges, including how to effectively generate satisfactory stylized results for any specified style while maintaining temporal coherence across frames. Towards this end, we propose a Multi-Channel Correlation network (MCCNet), which can be trained to fuse exemplar style features and input content features for efficient style transfer while naturally maintaining the coherence of input videos to output videos. Specifically, MCCNet works directly on the feature space of style and content domain where it learns to rearrange and fuse style features on the basis of their similarity to content features. The outputs generated by MCC are features containing the desired style patterns that can further be decoded into images with vivid style textures. Moreover, MCCNet is also designed to explicitly align the features to input and thereby ensure that the outputs maintain the content structures and the temporal continuity. To further improve the performance of MCCNet under complex light conditions, we also introduce illumination loss during training. Qualitative and quantitative evaluations demonstrate that MCCNet performs well in arbitrary video and image style transfer tasks. Code is available at https://github.com/diyiiyiii/MCCNet. Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Changsheng Xu |
AAAI | 1 |
| 2021 | Exploring the Representativity of Art PaintingsabstractArt painting evaluation is sophisticated for a novice with no or limited knowledge on art criticism, and history. In this study, we propose the concept ofrepresentativityto evaluate paintings instead of using professional concepts, such as genre, media, and style, which may be confusing to non-professionals. We define the concept of representativity to evaluate quantitatively the extent to which a painting can represent the characteristics of an artists creations. We begin by proposing a novel deep representation of art paintings, which is enhanced by style information through a weighted pooling feature fusion module. In contrast to existing feature extraction approaches, the proposed framework embeds painting styles, and authorship information, and learns specific artwork characteristics in a single framework. Subsequently, we propose a graph-based learning method for representativity learning, which considers intra-category, and extra-category information. In view of the significance of historical factors in the art domain, we introduce the creation time of a painting into the learning process. User studies demonstrate our approach helps the public effectively access the creation characteristics of artists through sorting paintings by representativity from highest to lowest. Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Feiyue Huang, Oliver Deussen, Changsheng Xu |
IEEE Trans. Multim. | 1 |
| 2020 | Arbitrary Style Transfer via Multi-Adaptation NetworkabstractArbitrary style transfer is a significant topic with research value and application prospect. A desired style transfer, given a content image and referenced style painting, would render the content image with the color tone and vivid stroke patterns of the style painting while synchronously maintaining the detailed content structure information. Style transfer approaches would initially learn content and style representations of the content and style references and then generate the stylized images guided by these representations. In this paper, we propose the multi-adaptation network which involves two self-adaptation (SA) modules and one co-adaptation (CA) module:the SA modules adaptively disentangle the content and style representations, i.e., content SA module uses position-wise self-attention to enhance content representation and style SA module uses channel-wise self-attention to enhance style representation; the CA module rearranges the distribution of style representation based on content representation distribution by calculating the local similarity between the disentangled content and style features in a non-local fashion. Moreover, a new disentanglement loss function enables our network to extract main style patterns and exact content structures to adapt to various input images, respectively. Various qualitative and quantitative experiments demonstrate that the proposed multi-adaptation network leads to better results than the state-of-the-art style transfer methods. Yingying Deng, Fan Tang, Weiming Dong, Feiyue Huang, Changsheng Xu |
ACM Multimedia | 1 |
| 2019 | Selective clustering for representative paintings selection
Yingying Deng, Fan Tang, Weiming Dong, Fuzhang Wu, Oliver Deussen, Changsheng Xu |
Multim. Tools Appl. | 1 |
| 2013 | AOPUT: A recommendation framework based on social activities and content interestsabstractContent consuming and sharing are two most important user activities in social networking sites (SNSs). Lots of studies have been conducted on content recommendation using users' common interests. However, little has been done to help users to select friends and share content within their social networks. In this paper, we contribute a recommendation framework AOPUT to recommend both content and friend list for sharing to users leveraging content and social information in SNSs. It consists of two recommendation components: Recder and ShareAider. Recder generates content recommendations by connecting users with common interests. An improved Jaccard similarity is proposed to improve the Collaborative Filtering (CF) recommendation quality. ShareAider recommends a friend list to users when they want to share content with their friends. CF method and a social-based method are compared and the combination of them are explored to achieve better results. AOPUT is evaluated on a real world social network. The experimental results show that (1) Recder can provide better recommendation quality than the traditional CF method thanks to the improved Jaccard similarity; (2) social-based method performs better than CF since the sharing behavior in SNSs are highly dominated by users' social preferences, and the combination of these two methods performs better than each of them individually. Yingying Deng, Tun Lu, Huanhuan Xia, Dongsheng Li 0002, Tiejiang Liu, Xianghua Ding, Ning Gu 0001 |
CSCWD | 1 |
| 2004 | Possibilistic-clustering-based MR brain image segmentation with accurate initializationabstractMagnetic resonance image analysis by computer is useful to aid diagnosis of malady. We present in this paper a automatic segmentation method for principal brain tissues. It is based on the possibilistic clustering approach, which is an improved fuzzy c-means clustering method. In order to improve the efficiency of clustering process, the initial value problem is discussed and solved by combining with a histogram analysis method. Our method can automatically determine number of classes to cluster and the initial values for each class. It has been tested on a set of forty MR brain images with or without the presence of tumor. The experimental results showed that it is simple, rapid and robust to segment the principal brain tissues. Qingmin Liao, Yingying Deng, Weibei Dou, Su Ruan, Daniel Bloyet |
VCIP | 2 |