Shuyang Gu

dblp:180/3316 · DBLP profile ↗
← Back
28ranked-venue papers
7as first author
20since 2021 · last 2025
0000-0003-4535-2280ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 15 · 4 first-author · 12 since 2021Theory of computation · 7 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization
abstract
Generating visually appealing images is fundamental to modern text-to-image generation models. A potential solution to better aesthetics is direct preference optimization (DPO), which has been applied to diffusion models to improve general image quality including prompt alignment and aesthetics. Popular DPO methods propagate preference labels from clean image pairs to all the intermediate steps along the two generation trajectories. However, preference labels provided in existing datasets are blended with layout and aesthetic opinions, which would disagree with aesthetic preference. Even if aesthetic labels were provided (at substantial cost), it would be hard for the two-trajectory methods to capture nuanced visual differences at different steps. To improve aesthetics economically, this paper uses existing generic preference data and introduces step-by-step preference optimization (SPO) that discards the propagation strategy and allows fine-grained image details to be assessed. Specifically, at each denoising step, we 1) sample a pool of candidates by denoising from a shared noise latent, 2) use a step-aware preference model to find a suitable win-lose pair to supervise the diffusion model, and 3) randomly select one from the pool to initialize the next denoising step. This strategy ensures that diffusion models focus on the subtle, fine-grained visual differences instead of layout aspect. We find that aesthetics can be significantly enhanced by accumulating these improved minor differences. When fine-tuning Stable Diffusion v1.5 and SDXL, SPO yields significant improvements in aesthetics compared with existing DPO methods while not sacrificing image-text alignment compared with vanilla models. Moreover, SPO converges much faster than DPO methods due to the use of more correct preference labels provided by the step-aware preference model. Code and models are available at https://github.com/RockeyCoss/SPO.
Yuhui Yuan, Shuyang Gu, Tiankai Hang, Mingxi Cheng
CVPR3
2025 DesignDiffusion: High-Quality Text-to-Design Image Generation with Diffusion Models
abstract
In this paper, we present DesignDiffusion, a simple yet effective framework for the novel task of synthesizing design images from textual descriptions. A primary challenge lies in generating accurate and style-consistent textual and visual content. Existing works in a related task of visual text generation often focus on generating text within given specific regions, which limits the creativity of generation models, resulting in style or color inconsistencies between textual and visual elements if applied to design image generation. To address this issue, we propose an end-to-end, one-stage diffusion-based framework that avoids intricate components like position and layout modeling. Specifically, the proposed framework directly synthesizes textual and visual design elements from user prompts. It utilizes a distinctive character embedding derived from the visual text to enhance the input prompt, along with a character localization loss for enhanced supervision during text generation. Furthermore, we employ a self-play Direct Preference Optimization fine- tuning strategy to improve the quality and accuracy of the synthesized visual text. Extensive experiments demonstrate that DesignDiffusion achieves state-of-the-art performance in design image generation.
Jianmin Bao, Shuyang Gu, Dong Chen 0003, Wengang Zhou 0001, Houqiang Li
CVPR3
2025 Incorporating Pre-Trained Diffusion Models in Solving the Schrödinger Bridge Problem
abstract
This paper aims to unify Score-based Generative Models (SGMs), also known as Diffusion models, and the Schrödinger Bridge (SB) problem through three reparameterization techniques: Iterative Proportional Mean-Matching (IPMM), Iterative Proportional Terminus-Matching (IPTM), and Iterative Proportional Flow-Matching (IPFM). These techniques significantly accelerate and stabilize the training of SB-based models. Furthermore, the paper introduces novel initialization strategies that use pre-trained SGMs to effectively train SB-based models. By using SGMs as initialization, we leverage the advantages of both SB-based models and SGMs, ensuring efficient training of SB-based models and further improving the performance of SGMs. Extensive experiments demonstrate the significant effectiveness and improvements of the proposed methods. We believe this work contributes to and paves the way for future research on generative models.
Zhicong Tang, Tiankai Hang, Shuyang Gu, Dong Chen 0003, Baining Guo
ECAI3
2025 Improved Noise Schedule for Diffusion Training
abstract
Diffusion models have emerged as the de facto choice for generating high-quality visual signals across various domains. However, training a single model to predict noise across various levels poses significant challenges, necessitating numerous iterations and incurring significant computational costs. Various approaches, such as loss weighting strategy design and architectural refinements, have been introduced to expedite convergence and improve model performance. In this study, we propose a novel approach to design the noise schedule for enhancing the training of diffusion models. Our key insight is that the importance sampling of the logarithm of the Signal-to-Noise ratio ($\log \text{SNR}$), theoretically equivalent to a modified noise schedule, is particularly beneficial for training efficiency when increasing the sample frequency around $\log \text{SNR}=0$. This strategic sampling allows the model to focus on the critical transition point between signal dominance and noise dominance, potentially leading to more robust and accurate predictions.We empirically demonstrate the superiority of our noise schedule over the standard cosine schedule.Furthermore, we highlight the advantages of our noise schedule design on the ImageNet benchmark, showing that the designed schedule consistently benefits different prediction targets. Our findings contribute to the ongoing efforts to optimize diffusion models, potentially paving the way for more efficient and effective training paradigms in the field of generative AI.
Tiankai Hang, Shuyang Gu, Jianmin Bao, Fangyun Wei, Dong Chen 0003, Xin Geng 0001, Baining Guo
ICCV2
2025 A Dual-Level Game-Theoretic Approach for Collaborative Learning in UAV-Assisted Heterogeneous Vehicle Networks
abstract
Knowledge diversity and knowledge forgetting are two major issues in sustaining collaborative learning within heterogeneous vehicle networks. These issues become especially severe when vehicles possess varying sensing capabilities, computational resources, and domain expertise, leading to fragmented learning and unstable knowledge retention over time. To address these challenges, we propose a dual-level game-theoretic approach. We first formulate a new metric, Utility-of-Information (UoI), to characterize the features of knowledge learning, retention, and consolidation. Based on this metric, we design a game-theoretic dual-level approach, which comprises a lower-level coalition formation game where vehicles self-organize into “teacher-student” coalitions based on their UoI profiles, and an upper-level UAV resource allocation game where vehicle coalitions compete for limited communication resources. To optimize both levels of the game, we design a unified reinforcement learning-based framework that enables adaptive searching for optimization under dynamic network conditions. Experimental results demonstrate that our approach effectively addresses knowledge diversity and significantly mitigates the effects of knowledge forgetting in UAV-assisted heterogeneous vehicle networks.
Jun Huang 0002, Qiang Duan 0002, Yanxiao Zhao, Shuyang Gu
IPCCC6
2025 VolumeDiffusion: Feed-forward text-to-3D generation with efficient volumetric encoder
abstract
This work presents VolumeDiffusion, a novel feed-forward text-to-3D generation framework that directly synthesizes 3D objects from textual descriptions. It bypasses the conventional score distillation loss based or text-to-image-to-3D approaches. To scale up the training data for the diffusion model, a novel 3D volumetric encoder is developed to efficiently acquire feature volumes from multi-view images. The 3D volumes are then trained on a diffusion model for text-to-3D generation using a 3D U-Net. This research further addresses the challenges of inaccurate object captions and high-dimensional feature volumes. The proposed model, trained on the public Objaverse dataset, demonstrates promising outcomes in producing diverse and recognizable samples from text prompts. Notably, it empowers finer control over object part characteristics through textual cues, fostering model creativity by seamlessly combining multiple concepts within a single object. This research significantly contributes to the progress of 3D generation by introducing an efficient, flexible, and scalable representation methodology.
Zhicong Tang, Shuyang Gu, Chunyu Wang 0001, Ting Zhang 0002, Jianmin Bao, Dong Chen 0003, Baining Guo
Graph. Model.2
2025 CCA: collaborative competitive agents for image editing
Tiankai Hang, Shuyang Gu, Dong Chen 0003, Xin Geng 0001, Baining Guo
Frontiers Comput. Sci.2
2024 InstructDiffusion: A Generalist Modeling Interface for Vision Tasks
abstract
We present InstructDiffusion, a unified and generic framework for aligning computer vision tasks with hu-man instructions. Unlike existing approaches that integrate prior knowledge and pre-define the output space (e.g., categories and coordinates) for each vision task, we cast diverse vision tasks into a human-intuitive image-manipulating pro-cess whose output space is a flexible and interactive pixel space. Concretely, the model is built upon the diffusion process and is trained to predict pixels according to user instructions, such as encircling the man's left shoulder in red or applying a blue mask to the left car. InstructDiffusion could handle a variety of vision tasks, including understanding tasks (such as segmentation and keypoint de-tection) and generative tasks (such as editing and enhance-ment) and outperforms prior methods on novel datasets. This represents a solid step towards a generalist modeling interface for vision tasks, advancing artificial general intelligence in the field of computer vision.
Zigang Geng, Binxin Yang, Tiankai Hang, Shuyang Gu, Ting Zhang 0002, Jianmin Bao, Zheng Zhang 0022, Houqiang Li, Han Hu 0001, Dong Chen 0003, Baining Guo
CVPR5
2024 FontStudio: Shape-Adaptive Diffusion Model for Coherent and Consistent Font Effect Generation
Xinzhi Mu, Li Chen 0033, Shuyang Gu, Jianmin Bao, Dong Chen 0003, Ji Li 0006, Yuhui Yuan
ECCV (58)4
2023 RODIN: A Generative Model for Sculpting 3D Digital Avatars Using Diffusion
abstract
This paper presents a 3D diffusion model that automatically generates 3D digital avatars represented as neural radiance fields (NeRFs). A significant challenge for 3D diffusion is that the memory and processing costs are prohibitive for producing high-quality results with rich details. To tackle this problem, we propose the roll-out diffusion network (RODIN), which takes a 3D NeRF model represented as multiple 2D feature maps and rolls out them onto a single 2D feature plane within which we perform 3D-aware diffusion. The RODIN model brings much-needed computational efficiency while preserving the integrity of 3D diffusion by using 3D-aware convolution that attends to projected features in the 2D plane according to their original relationships in 3D. We also use latent conditioning to orchestrate the feature generation with global coherence, leading to high-fidelity avatars and enabling semantic editing based on text prompts. Finally, we use hierarchical synthesis to further enhance details. The 3D avatars generated by our model compare favorably with those produced by existing techniques. We can generate highly detailed avatars with realistic hairstyles and facial hair. We also demonstrate 3D avatar generation from image or text, as well as text-guided editability.
Tengfei Wang 0002, Bo Zhang 0025, Ting Zhang 0002, Shuyang Gu, Jianmin Bao, Tadas Baltrusaitis, Jingjing Shen, Dong Chen 0003, Fang Wen 0001, Qifeng Chen 0001, Baining Guo
CVPR4
2023 Paint by Example: Exemplar-based Image Editing with Diffusion Models
abstract
Language-guided image editing has achieved great success recently. In this paper, we investigate exemplar-guided image editing for more precise control. We achieve this goal by leveraging self-supervised training to disentangle and re-organize the source image and the exemplar. However, the naive approach will cause obvious fusing artifacts. We carefully analyze it and propose a content bottleneck and strong augmentations to avoid the trivial solution of directly copying and pasting the exemplar image. Meanwhile, to ensure the controllability of the editing process, we design an arbitrary shape mask for the exemplar image and leverage the classifier-free guidance to increase the similarity to the exemplar image. The whole framework involves a single forward of the diffusion model without any iterative optimization. We demonstrate that our method achieves an impressive performance and enables controllable editing on in-the-wild images with high fidelity. The code and pretrained models are available at https://github.com/Fantasy-Studio/Paint-by-Example.
Binxin Yang, Shuyang Gu, Bo Zhang 0025, Ting Zhang 0002, Xuejin Chen, Xiaoyan Sun 0001, Dong Chen 0003, Fang Wen 0001
CVPR2
2023 Efficient Diffusion Training via Min-SNR Weighting Strategy
abstract
Denoising diffusion models have been a mainstream approach for image generation, however, training these models often suffers from slow convergence. In this paper, we discovered that the slow convergence is partly due to conflicting optimization directions between timesteps. To address this issue, we treat the diffusion training as a multi-task learning problem, and introduce a simple yet effective approach referred to as Min-SNR-γ. This method adapts loss weights of timesteps based on clamped signal-to-noise ratios, which effectively balances the conflicts among timesteps. Our results demonstrate a significant improvement in converging speed, 3.4× faster than previous weighting strategies. It is also more effective, achieving a new record FID score of 2.06 on the ImageNet 256 × 256 benchmark using smaller architectures than that employed in previous state-of-the-art. The code is available at https://github.com/TiankaiHang/Min-SNR-Diffusion-Training.
Tiankai Hang, Shuyang Gu, Jianmin Bao, Dong Chen 0003, Han Hu 0001, Xin Geng 0001, Baining Guo
ICCV2
2023 Profit maximization in social networks and non-monotone DR-submodular maximization
Shuyang Gu, Chuangen Gao, Weili Wu 0001
Theor. Comput. Sci.1
2022 A Binary Search Double Greedy Algorithm for Non-monotone DR-submodular Maximization
Shuyang Gu, Chuangen Gao, Weili Wu 0001
AAIM1
2022 Vector Quantized Diffusion Model for Text-to-Image Synthesis
abstract
We present the vector quantized diffusion (VQ-Diffusion) model for text-to-image generation. This method is based on a vector quantized variational autoencoder (VQ-VAE) whose latent space is modeled by a conditional variant of the recently developed Denoising Diffusion Probabilistic Model (DDPM). We find that this latent-space method is well-suited for text-to-image generation tasks because it not only eliminates the unidirectional bias with existing methods but also allows us to incorporate a mask-and-replace diffusion strategy to avoid the accumulation of errors, which is a serious problem with existing methods. Our experiments show that the VQ-Diffusion produces significantly better text-to-image generation results when compared with conventional autoregressive (AR) models with similar numbers of parameters. Compared with previous GAN-based text-to-image methods, our VQ-Diffusion can handle more complex scenes and improve the synthesized image quality by a large margin. Finally, we show that the image generation computation in our method can be made highly efficient by reparameterization. With traditional AR methods, the text-to-image generation time increases linearly with the output image resolution and hence is quite time consuming even for normal size images. The VQ-Diffusion allows us to achieve a better trade-off between quality and speed. Our experiments indicate that the VQ-Diffusion model with the reparameterization is fifteen times faster than traditional AR methods while achieving a better image quality. The code and models are available at https://github.com/cientgu/VQ-Diffusion.
Shuyang Gu, Dong Chen 0003, Jianmin Bao, Fang Wen 0001, Bo Zhang 0025, Dongdong Chen 0001, Lu Yuan 0001, Baining Guo
CVPR1
2022 StyleSwin: Transformer-based GAN for High-resolution Image Generation
abstract
Despite the tantalizing success in a broad of vision tasks, transformers have not yet demonstrated on-par ability as ConvNets in high-resolution image generative modeling. In this paper, we seek to explore using pure transformers to build a generative adversarial network for high-resolution image synthesis. To this end, we believe that local attention is crucial to strike the balance between computational efficiency and modeling capacity. Hence, the proposed generator adopts Swin transformer in a style-based architecture. To achieve a larger receptive field, we propose double attention which simultaneously leverages the context of the local and the shifted windows, leading to improved generation quality. Moreover, we show that offering the knowledge of the absolute position that has been lost in window-based transformers greatly benefits the generation quality. The proposed StyleSwin is scalable to high resolutions, with both the coarse geometry and fine structures benefit from the strong expressivity of transformers. However, blocking artifacts occur during high-resolution synthesis because performing the local attention in a block-wise manner may break the spatial coherency. To solve this, we empirically investigate various solutions, among which we find that employing a wavelet discriminator to examine the spectral discrepancy effectively suppresses the artifacts. Extensive experiments show the superiority over prior transformer-based GANs, especially on high resolutions, e.g.,$1024 \times$1024. The StyleSwin, without complex training strategies, excels over StyleGAN on CelebA-HQ 1024, and achieves on-par performance on FFHQ-1024, proving the promise of using transformers for high-resolution image generation. The code and pretrained models are available at https://github.com/microsoft/StyleSwin.
Bowen Zhang 0010, Shuyang Gu, Bo Zhang 0025, Jianmin Bao, Dong Chen 0003, Fang Wen 0001, Baining Guo
CVPR2
2022 Adaptive seeding for profit maximization in social networks
Chuangen Gao, Shuyang Gu, Jiguo Yu, Hai Du, Weili Wu 0001
J. Glob. Optim.2
2021 High-Fidelity and Arbitrary Face Editing
abstract
Cycle consistency is widely used for face editing. However, we observe that the generator tends to find a tricky way to hide information from the original image to satisfy the constraint of cycle consistency, making it impossible to maintain the rich details (e.g., wrinkles and moles) of non-editing areas. In this work, we propose a simple yet effective method named HifaFace to address the above-mentioned problem from two perspectives. First, we relieve the pressure of the generator to synthesize rich details by directly feeding the high-frequency information of the input image into the end of the generator. Second, we adopt an additional discriminator to encourage the generator to synthesize rich details. Specifically, we apply wavelet transformation to transform the image into multi-frequency domains, among which the high-frequency parts can be used to recover the rich details. We also notice that a fine-grained and wider-range control for the attribute is of great importance for face editing. To achieve this goal, we propose a novel attribute regression loss. Powered by the proposed framework, we achieve high-fidelity and arbitrary face editing, outperforming other state-of-the-art approaches.
Yue Gao 0006, Fangyun Wei, Jianmin Bao, Shuyang Gu, Dong Chen 0003, Fang Wen 0001, Zhouhui Lian
CVPR4
2021 Learnable Sampling 3D Convolution for Video Enhancement and Action Recognition
abstract
A key challenge in video enhancement and action recognition is to fuse useful information from neighboring frames. Recent works suggest establishing accurate correspondences between neighboring frames before fusing temporal information. However, the generated results heavily depend on the quality of correspondence estimation. This paper proposes a more robust solution: sampling and fusing multi-level features across neighborhood frames to generate the results. Based on this idea, we introduce a new module to improve the capability of 3D convolution, namely, learnable sampling 3D convolution (LS3D-Conv). We add learnable 2D offsets to 3D convolution, aiming to sample locations on spatial feature maps across frames. The offsets can be learned for specific tasks. The LS3D-Conv can flexibly replace 3D convolution layers in existing 3D networks and get new architectures, which learns the sampling at multiple feature levels. The experiments on video interpolation, video super-resolution, video denoising, and action recognition demonstrate the effectiveness of our approach.
Shuyang Gu, Jianmin Bao, Dong Chen 0003
ICME1
2021 A constrained two-stage submodular maximization
Shuyang Gu, Chuangen Gao, Weili Wu 0001, Dachuan Xu 0001
Theor. Comput. Sci.2
2020 GIQA: Generated Image Quality Assessment
Shuyang Gu, Jianmin Bao, Dong Chen 0003, Fang Wen 0001
ECCV (11)1
2020 Data placement in distributed data centers for improved SLA and network cost
Yuqi Fan 0001, Chen Wang 0059, Shuyang Gu, Weili Wu 0001, Ding-Zhu Du
J. Parallel Distributed Comput.4
2020 Interaction-aware influence maximization and iterated sandwich method
Chuangen Gao, Shuyang Gu, Jiguo Yu, Weili Wu 0001, Dachuan Xu 0001
Theor. Comput. Sci.2
2019 Interaction-Aware Influence Maximization and Iterated Sandwich Method
Chuangen Gao, Shuyang Gu, Jiguo Yu, Weili Wu 0001, Dachuan Xu 0001
AAIM2
2019 A Two-Stage Constrained Submodular Maximization
Shuyang Gu, Chuangen Gao, Weili Wu 0001, Dachuan Xu 0001
AAIM2
2019 Mask-Guided Portrait Editing With Conditional GANs
abstract
Portrait editing is a popular subject in photo manipulation.The Generative Adversarial Network (GAN) advances the generating of realistic faces and allows more face editing. In this paper, we argue about three issues in existing techniques: diversity, quality, and controllability for portrait synthesis and editing. To address these issues, we propose a novel end-to-end learning framework that leverages conditional GANs guided by provided face masks for generating faces. The framework learns feature embeddings for every face component (e.g., mouth, hair, eye), separately, contributing to better correspondences for image translation, and local face editing. With the mask, our network is available to many applications, like face synthesis driven by mask, face Swap+ (including hair in swapping), and local manipulation. It can also boost the performance of face parsing a bit as an option of data augmentation.
Shuyang Gu, Jianmin Bao, Hao Yang 0036, Dong Chen 0003, Fang Wen 0001, Lu Yuan 0001
CVPR1
2019 Robust Profit Maximization with Double Sandwich Algorithms in Social Networks
abstract
Social networks are becoming important dissemination platforms, and a large body of works have been performed on viral marketing, but most are to maximize the benefits associated with the number of active nodes. In this paper, we study the benefits related to interactions among activated nodes. Furthermore, due to the uncertainty in edge probability estimates in social networks, we propose the robust profit maximization problem to have the best solution in the worst case of probability settings. We design a double sandwich algorithm to this problem and further improve the algorithm with sampling method such that it increases robustness of the output. Through real data sets, we verify the effectiveness of our proposed algorithm.
Chuangen Gao, Shuyang Gu, Hongwei Du 0001, Smita Ghosh
ICDCS2
2018 Arbitrary Style Transfer With Deep Feature Reshuffle
abstract
This paper introduces a novel method by reshuffling deep features (i.e., permuting the spacial locations of a feature map) of the style image for arbitrary style transfer. We theoretically prove that our new style loss based on reshuffle connects both global and local style losses respectively used by most parametric and non-parametric neural style transfer methods. This simple idea can effectively address the challenging issues in existing style transfer methods. On one hand, it can avoid distortions in local style patterns, and allow semantic-level transfer, compared with neural parametric methods. On the other hand, it can preserve globally similar appearance to the style image, and avoid wash-out artifacts, compared with neural non-parametric methods. Based on the proposed loss, we also present a progressive feature-domain optimization approach. The experiments show that our method is widely applicable to various styles, and produces better quality than existing methods.
Shuyang Gu, Congliang Chen, Jing Liao 0001, Lu Yuan 0001
CVPR1