Changshuo Wang 0003

dblp:297/9002-3 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0002-9970-191XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 FashionDPO: Fine-tune Fashion Outfit Generation Model using Direct Preference Optimization
abstract
Personalized outfit generation aims to construct a set of compatible and personalized fashion items as an outfit. Recently, generative AI models have received widespread attention, as they can generate fashion items for users to complete an incomplete outfit or create a complete outfit. However, they have limitations in terms of lacking diversity and relying on the supervised learning paradigm. Recognizing this gap, we propose a novel framework FashionDPO, which fine-tunes the fashion outfit generation model using direct preference optimization. This framework aims to provide a general fine-tuning approach to fashion generative models, refining a pre-trained fashion outfit generation model using automatically generated feedback, without the need to design a task-specific reward function. To make sure that the feedback is comprehensive and objective, we design a multi-expert feedback generation module which covers three evaluation perspectives, i.e., quality, compatibility and personalization. Experiments on two established datasets, i.e., iFashion and Polyvore-U, demonstrate the effectiveness of our framework in enhancing the model's ability to align with users' personalized preferences while adhering to fashion compatibility principles. Our code and model checkpoints are available at https://github.com/Yzcreator/FashionDPO.
Mingzhe Yu, Yunshan Ma 0002, Lei Wu 0002, Changshuo Wang 0003, Xue Li 0010, Lei Meng 0001
SIGIR4
2024 Scene Sketch-to-Image Synthesis Based on Multi-Object Control
abstract
Scene sketch-to-image synthesis is a challenging task, especially when the sketches contain multiple objects of different classes. Existing methods interfere between different classes of objects when generating images from scene sketches, making it difficult to synthesis images with accurate object classes. In this paper, we propose a scene sketch-to-image generation method based on multi-object control, which can generate high-quality and class-accurate images from scene sketches and text prompts. We propose a sampling strategy based on segmentation mask and independent denoising, which can accurately control the classes of foreground objects and make foreground objects and background more harmonized. Our method is based on a pre-trained diffusion model without additional training overhead. Experiments on SketchyCOCO and SketchyScene datasets demonstrate that our method’s capacity to generate realistic complex images from scene sketches and text prompts.
Zhenwei Cheng, Lei Wu 0002, Changshuo Wang 0003, Xiangxu Meng
ICASSP3
2024 Text-Guided Multi-region Scene Image Editing Based on Diffusion Model
Lei Wu 0002, Changshuo Wang 0003, Pei Dong
ICIC (11)3
2024 LayoutDM: Precision Multi-Scale Diffusion for Layout-to-Image
abstract
In recent research literature, the layout to image domain has gained significant traction. Research based on GANs can generate complete images, but there are still issues with insufficient details and overall quality not being high. Diffusion models, on the other hand, challenges still exist in ensuring the quality of image details in complex scene regions, as well as ensuring the smooth expression of the overall semantic information. To address these challenges, we propose LayoutDM based on multi-scale diffusion, which employs the Parallel Sampling Module to enhance local precision and ensures global semantic coherence through the Semantic Coherence Module. Significantly, our approach generates within the visible space, progressively revealing more details and semantic information. Simultaneously, our method enhances layout part handling via parallel region-wise clip guidance, achieving strong zero-shot generation without direct training samples. Tests on COCO-stuff and VG datasets confirm that our approach achieves both fine-grained object generation and overall visual effectiveness.
Mingzhe Yu, Lei Wu 0002, Changshuo Wang 0003, Lei Meng 0001, Xiangxu Meng
ICME3
2024 InstantAS: Minimum Coverage Sampling for Arbitrary-Size Image Generation
abstract
In recent years, diffusion models have dominated the field of image generation with their outstanding generation quality. However, pre-trained large-scale diffusion models are generally trained using fixed-size images, and fail to maintain their performance at different aspect ratios. Existing methods for generating arbitrary-size images based on diffusion models face several issues, including the requirement for extensive finetuning or training, sluggish sampling speed, and noticeable edge artifacts. This paper presents the InstantAS method for arbitrary-size image generation. This method performs non-overlapping minimum coverage segmentation on the target image, minimizing the generation of redundant information and significantly improving sampling speed. To maintain the consistency of the generated image, we also proposed the Inter-Domain Distribution Bridging method to integrate the distribution of the entire image and suppress the separation of diffusion paths in different regions of the image. Furthermore, we propose the dynamic semantic guided cross-attention method, allowing for the control of different regions using different semantics. Experimental results show that InstantAS has better fusion capabilities compared to previous arbitrary-size image generation methods and is far ahead in sampling speed compared to them.
Changshuo Wang 0003, Mingzhe Yu, Lei Wu 0002, Lei Meng 0001, Xiang Li 0177, Xiangxu Meng
ACM Multimedia1
2023 Letter Embedding Guidance Diffusion Model for Scene Text Editing
abstract
Scene text editing(STE) aims to modify the text in the scene image to the target text while retaining the original style. Existing models are based on GAN, where the source image and the target text are input only once during the generation process, and this approach could not fully obtain the style of the source image and content of the target text. In this paper, we propose an STE method based on the classifier-free guidance diffusion model. To our best knowledge, our model is the first work that developed diffusion models to handle the STE task. Specifically, we divide the STE task into multiple steps and extract style information and text content information in each step. In addition, we introduce the letter embedding method as guidance. We experimentally prove that our method outperforms other STE models in terms of overall realism and maintaining glyphs.
Changshuo Wang 0003, Lei Wu 0002, Xu Chen 0031, Xiang Li 0177, Lei Meng 0001, Xiangxu Meng
ICME1
2023 Compositional Zero-Shot Artistic Font Synthesis
abstract
Recently, many researchers have made remarkable achievements in the field of artistic font synthesis, with impressive glyph style and effect style in the results. However, due to less exploration in style disentanglement, it is difficult for existing methods to envision a kind of unseen style (glyph-effect) compositions of artistic font, and thus can only learn the seen style compositions. To solve this problem, we propose a novel compositional zero-shot artistic font synthesis gan (CAFS-GAN), which allows the synthesis of unseen style compositions by exploring the visual independence and joint compatibility of encoding semantics between glyph and effect. Specifically, we propose two contrast-based style encoders to achieve style disentanglement due to glyph and effect intertwining in the image. Meanwhile, to preserve more glyph and effect detail, we propose a generator based on hierarchical dual styles AdaIN to reorganize content-styles representations from structure to texture gradually. Extensive experiments demonstrate the superiority of our model in generating high-quality artistic font images with unseen style compositions against other state-of-the-art methods. The source code and data is available at moonlight03.github.io/CAFS-GAN/.
Xiang Li 0177, Lei Wu 0002, Changshuo Wang 0003, Lei Meng 0001, Xiangxu Meng
IJCAI3
2023 Anything to Glyph: Artistic Font Synthesis via Text-to-Image Diffusion Model
abstract
The automatic generation of artistic fonts is a challenging task that attracts many research interests. Previous methods specifically focus on glyph or texture style transfer. However, we often come across creative fonts composed of objects in posters or logos. These fonts have proven to be a challenge for existing methods as they struggle to generate similar designs. This paper proposes a novel method for generating creative artistic fonts using a pre-trained text-to-image diffusion model. Our model takes a shape image and a prompt describing an object as input and generates an artistic glyph image consisting of such objects. Specifically, we introduce a novel heatmap-based weak position constraint method to guide the positioning of objects in the generated image, and we also propose the Latent Space Semantic Augmentation Module that improves other information while constraining object position. Our approach is unique in that it can preserve the object’s original shape while constraining its position. And our training method requires only a small quantity of generated data, making it an efficient unsupervised learning approach. Experimental results demonstrate that our method can generate various glyphs, including Chinese, English, Japanese, and symbols, using different objects. We also conducted qualitative and quantitative comparisons with various position control methods for the diffusion model. The results indicate that our approach outperforms other methods in terms of visual quality, innovation, and user evaluation.
Changshuo Wang 0003, Lei Wu 0002, Xiaole Liu, Xiang Li 0177, Lei Meng 0001, Xiangxu Meng
SIGGRAPH Asia1