Yefei Sheng

dblp:367/1944 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0002-2547-4646ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DreamStory: character identity-driven and spatial layout-guided diffusion for visual story generation
Zhuo Sheng, Yefei Sheng
Multim. Syst.2
2026 AdaEdit: Adaptive Diffusion Model for Invisible Target Oriented Text-Conditioned Image Editing
abstract
Text-conditioned image editing aims to modify a source image into a target image according to a specified text description, tackling two core challenges: locating target editing regions and ensuring consistency in non-target editing areas. Existing approaches utilize manual selection or cross-modal attention to define editing regions and deploy diffusion models to generate edited images. Despite these recent advancements, two problems remain. First, current methods fail to locate editing areas described in the text but invisible in the image. Second, they struggle to ensure spatial consistency in non-targeted regions due to the global noise addition along with excessive denoising during the diffusion process. To overcome these limitations, we propose AdaEdit, which comprises an adaptive mask localization module and an adaptive denoising strategy for text-conditioned image editing. AdaEdit can accurately identify the editing area via the measurement of cross-modal semantic mismatch, even when the visual details are not explicitly described in the text inputs. The adaptive denoising strategy applies varying noise levels to differentiate between targeted and non-targeted regions, enhancing the stability and consistency of the non-edited areas. Extensive experiments demonstrate that our proposed method achieves excellent performance on MS-COCO, MagicBrush, and Laion. We also expand our application to iterative editing tasks, thereby extending its utility for generalized editing scenarios.
Yefei Sheng, Jie Wang 0061, Ming Tao 0002, Bing-Kun Bao
ACM Trans. Multim. Comput. Commun. Appl.1
2025 InstantPainting: Expanding GANs for Efficient Text-Conditioned Image Generation Platform
abstract
Text-conditioned image generation enables cross-modal comprehension. Recent emergence of many platforms have found applications in diverse domains like assisted designing and video gaming. However, there still exist challenges in existing platforms due to their expensive training and time-consuming generation processes. In this paper, we introduce an efficient text-conditioned image generation platform, termed InstantPainting. Unlike existing platforms based on large-scale pre-trained diffusion models, InstantPainting expands generative adversarial networks (GANs) to achieve efficient generation by using only about three percent pre-training data of other platforms. Compared to existing platforms, InstantPainting achieves the following functions at a very low deployment cost and approximately 4 to 5 times faster generation speeds: (1) Multi-category and multi-size image generation (2) Image stylization and controlled generation (3) Creative generation, including the generation of poetry pictures and counterfactual images. The proposed platform provides web application implementations for PC and mobile, users can create high-quality images directly through the user interface.
Bing-Kun Bao, Yefei Sheng, Jie Wang 0061, Sisi You
AAAI2
2025 D2Gaussian: Dynamic Control with Discretized 3D View Modeling for Text-Driven 3D Gaussian Splatting Editing
abstract
Current advances in text-driven 3D scene editing tasks typically render the 3D representations into multi-view images and modify the images with the text instructions. Context consistency across multiple views and cross-modal consistency in the single-view are the keys to effective 3D editing. Accordingly, existing methods introduce additional image constraints and apply pre-trained 2D editing models. However, they fix the same text instruction across all views and freeze the pre-trained 2D model for single-view editing, leading to deficient modeling of 3D scene views and results in inconsistent generations with visual artifacts. To address these limitations, we introduce a discretized 3D view modeling method and a diffusion-based multi-view consistent editing pipeline for text-driven 3D gaussian splatting editing, abbreviated as D2Gaussian. Specifically, our approach constructs a codebook that encodes continuous 3D view information into discrete token embeddings to model the spatial feature expressions. Then, the token embeddings are proposed to guide and finetune the diffusion-based image editing model with the dynamic addition of control conditions, yielding a multi-view consistent editing pipeline. Finally, we introduce a 3D editing dataset generation approach along with a 3D-CLIP-SIM metric to form a benchmark, 3D-MagicBrush, to provide more diverse evaluation scenarios for future 3D editing works. Experiments demonstrate that our method achieves better visual results and multi-view consistency than previous state-of-the-art methods.
Yefei Sheng, Jie Wang 0061, Ming Tao 0002, Bing-Kun Bao
ACM Multimedia1
2024 Semantic Distance Adversarial Learning for Text-to-Image Synthesis
abstract
Text-to-Image (T2I) synthesis is a cross-modality task that requires a text description as input to generate a realistic and semantically consistent image. To guarantee semantic consistency, previous studies regenerate text descriptions from synthetic images and align them with the given descriptions. However, the existing redescription modules lack explicit modeling of their training objectives, which is crucial for reliable measurement of semantic distance between redescriptions and given text inputs. Consequently, the aligned text redescriptions suffer from training bias caused by the emergence of adversarial image samples, unseen semantics, and mistaken contents from low-quality synthesized images. To this end, we propose a SEMantic distance Adversarial learning (SEMA) framework for Text-to-Image synthesis which strengthens semantic consistency from two aspects: 1) We introduce adversarial learning between the image generator and the text redescription module to mutually promote or demote the quality of generated image or text instances. This learning model ensures accurate redescription of image contents, thus diminishing the generation of adversarial image samples. 2) We introduce two-fold semantic distance discrimination (SEM distance) to characterize semantic relevance between matching text or image pairs. The unseen semantics and mistaken contents will be penalized with a large SEM distance. The proposed discrimination method also simplifies the model training process with no need to optimize multiple discriminators. Experimental results on CUB Birds 200 and MS-COCO datasets show that the proposed model outperforms the state-of-the-art methods.
Yefei Sheng, Bing-Kun Bao, Yi-Ping Phoebe Chen, Changsheng Xu
IEEE Trans. Multim.2
2024 ISF-GAN: Imagine, Select, and Fuse with GPT-Based Text Enrichment for Text-to-Image Synthesis
abstract
Text-to-Image synthesis aims to generate an accurate and semantically consistent image from a given text description. However, it is difficult for existing generative methods to generate semantically complete images from a single piece of text. Some works try to expand the input text to multiple captions via retrieving similar descriptions of the input text from the training set but still fail to fill in missing image semantics. In this article, we propose a GAN-based approach to Imagine, Select, and Fuse for Text-to-image synthesis, named ISF-GAN. The proposed ISF-GAN contains Imagine Stage and Select and Fuse Stage to solve the above problems. First, the Imagine Stage proposes a text completion and enrichment module. This module guides a GPT-based model to enrich the text expression beyond the original dataset. Second, the Select and Fuse Stage selects qualified text descriptions and then introduces a cross-modal attentional mechanism to interact these different sentence embeddings with the image features at different scales. In short, our proposed model enriches the input text information for completing missing semantics and introduces a cross-modal attentional mechanism to maximize the utilization of enriched text information to generate semantically consistent images. Experimental results on CUB, Oxford-102, and CelebA-HQ datasets prove the effectiveness and superiority of the proposed network. Code is available at https://github.com/Feilingg/ISF-GAN
Yefei Sheng, Ming Tao 0002, Jie Wang 0061, Bing-Kun Bao
ACM Trans. Multim. Comput. Commun. Appl.1