Yuming Jiang 0003

dblp:52/146-3 · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0001-7653-4015ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 7 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
abstract
Video generation has witnessed significant advancements, yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align with human perceptions; 2) An ideal evaluation system should provide insights to inform future developments of video generation. To this end, we present VBench++, a comprehensive benchmark suite that dissects "video generation quality" into specific, hierarchical, and disentangled dimensions, each with tailored prompts and evaluation methods. VBench++ has several appealing properties: 1) Comprehensive Dimensions: VBench++ comprises 16 dimensions in text-to-video generation (e.g., subject identity inconsistency, motion smoothness, temporal flickering, and spatial relationship, etc). The evaluation metrics with fine-grained levels reveal individual models' strengths and weaknesses. 2) Human Alignment: We also provide a dataset of human preference annotations to validate our benchmarks' alignment with human perception, for each evaluation dimension respectively. 3) Valuable Insights: We look into current models' ability across various evaluation dimensions, and various content types. We also investigate the gaps between video and image generation models. 4) Versatile Benchmarking: VBench++ is designed to evaluate a wide range of video generation tasks, including text-to-video and image-to-video. We introduce a high-quality Image Suite with an adaptive aspect ratio to enable fair evaluations across different image-to-video generation settings. Beyond assessing technical quality, VBench++ evaluates the trustworthiness of video generative models, providing a more holistic view of model performance. 5) Full Open-Sourcing: We fully open-source VBench++, including all prompts, the Image Suite, evaluation methods, generated videos, and human preference annotations.
Fan Zhang 0045, Yinan He, Jiashuo Yu, Ziyue Dong, Qianli Ma 0008, Nattapol Chanpaisit, Chenyang Si, Yuming Jiang 0003, Yaohui Wang 0001, Ying-Cong Chen, Limin Wang 0002, Dahua Lin, Yu Qiao 0001, Ziwei Liu 0002
IEEE Trans. Pattern Anal. Mach. Intell.10
2026 AniFeats: Animate 3D Feature Meshes for Character Video Generation
abstract
Generating high-quality character animation videos is a fascinating yet challenging task. Existing methods use geometry guidance signals like skeletons, normal maps, or depth maps in a diffusion model to generate character videos from a single reference image. Although these approaches have shown encouraging results, they solely rely on cross attention layers to extract geometry guidance which inevitably leads to temporal inconsistencies and reduced quality. In this paper, we present a novel framework AniFeats to generate high-quality character animation videos. In contrast to existing methods, our key insight is to incorporate explicit features on 3D character meshes during the video generation to achieve significantly improved temporal consistency. Specifically, AniFeats extracts detailed features from the reference image, projects them onto 3D feature meshes based on SMPL-X, and utilizes rendered feature maps from the animated 3D feature meshes as guidance throughout the generation process. This approach directly links local patterns in the input image to those in the output video, effectively strengthening temporal coherence. Extensive experiments demonstrate that AniFeats generates high-quality, temporally consistent character animations with remarkably enhanced realism.
Beijia Lu, Zekai Gu, Zhiyang Dou, Haotian Yuan 0008, Chenyang Si, Yuming Jiang 0003, Yuan Liu 0025, Wenping Wang 0001, Ziwei Liu 0002
IEEE Trans. Vis. Comput. Graph.8
2025 LaVie: High-Quality Video Generation with Cascaded Latent Diffusion Models
Yaohui Wang 0001, Xin Ma 0031, Shangchen Zhou, Yi Wang 0074, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang 0001, Yuwei Guo 0002, Tianxing Wu 0002, Chenyang Si, Yuming Jiang 0003, Cunjian Chen, Chen Change Loy, Bo Dai 0002, Dahua Lin, Yu Qiao 0001, Ziwei Liu 0002
Int. J. Comput. Vis.14
2024 VideoBooth: Diffusion-based Video Generation with Image Prompts
abstract
Text-driven video generation witnesses rapid progress. However, merely using text prompts is not enough to depict the desired subject appearance that accurately aligns with users' intents, especially for customized content creation. In this paper, we study the task of video generation with image prompts, which provide more accurate and direct content control beyond the text prompts. Specifically, we propose a feed-forward framework VideoBooth, with two dedicated designs: 1) We propose to embed image prompts in a coarse-to-fine manner. Coarse visual embeddings from image encoder provide high-level encodings of image prompts, while fine visual embeddings from the proposed attention injection module provide multi-scale and detailed encoding of image prompts. These two complementary embeddings can faithfully capture the desired appearance. 2) In the attention injection module at fine level, multi-scale image prompts are fed into different cross-frame attention layers as additional keys and values. This extra spatial in-formation refines the details in the first frame and then it is propagated to the remaining frames, which maintains temporal consistency. Extensive experiments demonstrate that Video Booth achieves state-of-the-art performance in gener-ating customized high-quality videos with subjects specified in image prompts. Notably, VideoBooth is a generalizable framework where a single model works for a wide range of image prompts with only feed-forward passes.
Yuming Jiang 0003, Tianxing Wu 0002, Shuai Yang 0001, Chenyang Si, Dahua Lin, Yu Qiao 0001, Chen Change Loy, Ziwei Liu 0002
CVPR1
2024 VBench: Comprehensive Benchmark Suite for Video Generative Models
abstract
Video generation has witnessed significant advance-ments, yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align with human perceptions; 2) An ideal eval-uation system should provide insights to inform future de-velopments of video generation. To this end, we present VBench, a comprehensive benchmark suite that dissects “video generation quality” into specific, hierarchical, and disentangled dimensions, each with tailored prompts and evaluation methods. VBench has three appealing proper-ties: 1) Comprehensive Dimensions: VBench comprises 16 dimensions in video generation (e.g., subject identity in-consistency, motion smoothness, temporal flickering, and spatial relationship, etc.). The evaluation metrics with fine-grained levels reveal individual models' strengths and weaknesses. 2) Human Alignment: We also provide a dataset of human preference annotations to validate our benchmarks' alignment with human perception, for each evaluation dimension respectively. 3) Valuable Insights: We look into current models' ability across various evaluation dimensions, and various content types. We also investi-gate the gaps between video and image generation models. We will open-source VBench, including all prompts, evaluation methods, generated videos, and human preference an-notations, and also include more video generation models in VBench to drive forward the field of video generation.
Yinan He, Jiashuo Yu, Fan Zhang 0045, Chenyang Si, Yuming Jiang 0003, Yuanhan Zhang, Tianxing Wu 0002, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang 0001, Limin Wang 0002, Dahua Lin, Yu Qiao 0001, Ziwei Liu 0002
CVPR6
2024 FreeU: Free Lunch in Diffusion U-Net
abstract
In this paper, we uncover the untapped potential of dif-fusion U-Net, which serves as a “free lunch” that substan-tially improves the generation quality on the fly. We initially investigate the key contributions of the U-Net architecture to the denoising process and identify that its main backbone primarily contributes to denoising, whereas its skip connections mainly introduce high-frequency features into the de-coder module, causing the potential neglect of crucial functions intrinsic to the backbone network. Capitalizing on this discovery, we propose a simple yet effective method, termed “FreeU”, which enhances generation quality without additional training or finetuning. Our key insight is to strategi-cally re-weight the contributions sourced from the U-Net's skip connections and backbone feature maps, to leverage the strengths of both components of the U-Net architec-ture. Promising results on image and video generation tasks demonstrate that our FreeU can be readily integrated to ex-isting diffusion models, e.g., Stable Diffusion, DreamBooth and ControlNet, to improve the generation quality with only a few lines of code. All you need is to adjust two scaling factors during inference.
Chenyang Si, Yuming Jiang 0003, Ziwei Liu 0002
CVPR3
2024 GroupDiff: Diffusion-Based Group Portrait Editing
Yuming Jiang 0003, Nanxuan Zhao, Qing Liu 0017, Krishna Kumar Singh, Shuai Yang 0001, Chen Change Loy, Ziwei Liu 0002
ECCV (34)1
2024 FreeInit: Bridging Initialization Gap in Video Diffusion Models
Tianxing Wu 0002, Chenyang Si, Yuming Jiang 0003, Ziwei Liu 0002
ECCV (3)3
2024 ReVersion: Diffusion-Based Relation Inversion from Images
abstract
Diffusion models gain increasing popularity for their generative capabilities.Recently, there have been surging needs to generate customized images by inverting diffusion models from exemplar images, and existing inversion methods mainly focus on capturing object appearances (i.e., the "look").However, how to invert object relations, another important pillar in the visual world, remains unexplored.In this work, we propose the Relation Inversion task, which aims to learn a specific relation (represented as "relation prompt") from exemplar images.Specifically, we learn a relation prompt with a frozen pre-trained text-to-image diffusion model.The learned relation prompt can then be applied to generate relation-specific images with new objects, backgrounds, and styles.To tackle the Relation Inversion task, we propose the ReVersion Framework.Specifically, we propose a novel "relation-steering contrastive learning" scheme to steer the relation prompt towards relation-dense regions, and disentangle it away from object appearances.We further devise "relation-focal importance sampling" to emphasize high-level interactions over low-level appearances (e.g., texture, color).To comprehensively evaluate this new task, we contribute the ReVersion Benchmark, which provides various exemplar images with diverse
Tianxing Wu 0002, Yuming Jiang 0003, Kelvin C. K. Chan, Ziwei Liu 0002
SIGGRAPH Asia3
2024 ReliTalk: Relightable Talking Portrait Generation from a Single Video
Haonan Qiu, Zhaoxi Chen 0009, Yuming Jiang 0003, Hang Zhou 0009, Xiangyu Fan 0002, Lei Yang 0045, Wayne Wu, Ziwei Liu 0002
Int. J. Comput. Vis.3
2024 Talk-to-Edit: Fine-Grained 2D and 3D Facial Editing via Dialog
abstract
Facial editing is to manipulate the facial attributes of a given face image. Nowadays, with the development of generative models, users can easily generate 2D and 3D facial images with high fidelity and 3D-aware consistency. However, existing works are incapable of delivering a continuous and fine-grained editing mode (e.g., editing a slightly smiling face to a big laughing one) with natural interactions with users. In this work, we propose Talk-to-Edit, an interactive facial editing framework that performs fine-grained attribute manipulation through dialog between the user and the system. Our key insight is to model a continual "semantic field" in the GAN latent space. 1) Unlike previous works that regard the editing as traversing straight lines in the latent space, here the fine-grained editing is formulated as finding a curving trajectory that respects fine-grained attribute landscape on the semantic field. 2) The curvature at each step is location-specific and determined by the input image as well as the users' language requests. 3) To engage the users in a meaningful dialog, our system generates language feedback by considering both the user request and the current state of the semantic field. We demonstrate the effectiveness of our proposed framework on both 2D and 3D-aware generative models. We term the semantic field for the 3D-aware models as "tri-plane" flow, as it corresponds to the changes not only in the color space but also in the density space. We also contribute CelebA-Dialog, a visual-language facial editing dataset to facilitate large-scale study. Specifically, each image has manually annotated fine-grained attribute annotations as well as template-based textual descriptions in natural language. Extensive quantitative and qualitative experiments demonstrate the superiority of our framework in terms of 1) the smoothness of fine-grained editing, 2) the identity/attribute preservation, and 3) the visual photorealism and dialog fluency. Notably, the user study validates that our overall system is consistently favored by around 80% of the participants.
Yuming Jiang 0003, Tianxing Wu 0002, Xingang Pan, Chen Change Loy, Ziwei Liu 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Collaborative Diffusion for Multi-Modal Face Generation and Editing
abstract
Diffusion models arise as a powerful generative tool recently. Despite the great progress, existing diffusion models mainly focus on uni-modal control, i.e., the diffusion process is driven by only one modality of condition. To further unleash the users' creativity, it is desirable for the model to be controllable by multiple modalities simultaneously, e.g. generating and editing faces by describing the age (text-driven) while drawing the face shape (mask-driven). In this work, we present Collaborative Diffusion, where pre-trained uni-modal diffusion models collaborate to achieve multi-modal face generation and editing without re-training. Our key insight is that diffusion models driven by different modalities are inherently complementary regarding the latent denoising steps, where bilateral connections can be established upon. Specifically, we propose dynamic diffuser, a meta-network that adaptively hallucinates multimodal denoising steps by predicting the spatial-temporal influence functions for each pre-trained uni-modal model. Collaborative Diffusion not only collaborates generation capabilities from uni-modal diffusion models, but also integrates multiple uni-modal manipulations to perform multi-modal editing. Extensive qualitative and quantitative experiments demonstrate the superiority of our framework in both image quality and condition consistency.
Kelvin C. K. Chan, Yuming Jiang 0003, Ziwei Liu 0002
CVPR3
2023 Text2Performer: Text-Driven Human Video Generation
abstract
Text-driven content creation has evolved to be a transformative technique that revolutionizes creativity. Here we study the task of text-driven human video generation, where a video sequence is synthesized from texts describing the appearance and motions of a target performer. Compared to general text-driven video generation, human-centric video generation requires maintaining the appearance of synthesized human while performing complex motions. In this work, we present Text2Performer to generate vivid human videos with articulated motions from texts. Text2Performer has two novel designs: 1) decomposed human representation and 2) diffusion-based motion sampler. First, we decompose the VQVAE latent space into human appearance and pose representation in an unsupervised manner by utilizing the nature of human videos. In this way, the appearance is well maintained along the generated frames. Then, we propose continuous VQ-diffuser to sample a sequence of pose embeddings. Unlike existing VQ-based methods that operate in the discrete space, continuous VQ-diffuser directly outputs the continuous pose embeddings for better motion modeling. Finally, motion-aware masking strategy is designed to mask the pose embeddings spatial-temporally to enhance the temporal coherence. Moreover, to facilitate the task of text-driven human video generation, we contribute a Fashion-Text2Video dataset with manually annotated action labels and text descriptions. Extensive experiments demonstrate that Text2Performer generates high-quality human videos (up to 512 × 256 resolution) with diverse appearances and flexible motions. Our project page is https://yumingj.github.io/projects/Text2Performer.html
Yuming Jiang 0003, Shuai Yang 0001, Tong Liang Koh, Wayne Wu, Chen Change Loy, Ziwei Liu 0002
ICCV1
2023 UnitedHuman: Harnessing Multi-Source Data for High-Resolution Human Generation
abstract
Human generation has achieved significant progress. Nonetheless, existing methods still struggle to synthesize specific regions such as faces and hands. We argue that the main reason is rooted in the training data. A holistic human dataset inevitably has insufficient and low-resolution information on local parts. Therefore, we propose to use multisource datasets with various resolution images to jointly learn a high-resolution human generative model. However, multi-source data inherently a) contains different parts that do not spatially align into a coherent human, and b) comes with different scales. To tackle these challenges, we propose an end-to-end framework, UnitedHuman, that empowers continuous GAN with the ability to effectively utilize multi-source data for high-resolution human generation. Specifically, 1) we design a Multi-Source Spatial Transformer that spatially aligns multi-source images to full-body space with a human parametric model. 2) Next, a continuous GAN is proposed with global-structural guidance and CutMix consistency. Patches from different datasets are then sampled and transformed to supervise the training of this scale-invariant generative model. Extensive experiments demonstrate that our model jointly learned from multi-source data achieves superior quality than those learned from a holistic dataset. Project page: https://unitedhuman.github.io/.
Jianglin Fu, Shikai Li, Yuming Jiang 0003, Kwan-Yee Lin, Wayne Wu, Ziwei Liu 0002
ICCV3
2023 Reference-Based Image and Video Super-Resolution via $C^{2}$-Matching
abstract
Reference-based Super-Resolution (Ref-SR) has recently emerged as a promising paradigm to enhance a low-resolution (LR) input image or video by introducing an additional high-resolution (HR) reference image. Existing Ref-SR methods mostly rely on implicit correspondence matching to borrow HR textures from reference images to compensate for the information loss in input images. However, performing local transfer is difficult because of two gaps between input and reference images: the transformation gap (e.g.scale and rotation) and the resolution gap (e.g.HR and LR). To tackle these challenges, we propose$C^{2}$-Matching in this work, which performs explicit robust matching crossing transformation and resolution. 1) To bridge the transformation gap, we propose a contrastive correspondence network, which learns transformation-robust correspondences using augmented views of the input image. 2) To address the resolution gap, we adopt teacher-student correlation distillation, which distills knowledge from the easier HR-HR matching to guide the more ambiguous LR-HR matching. 3) Finally, we design a dynamic aggregation module to address the potential misalignment issue between input images and reference images. In addition, to faithfully evaluate the performance of Reference-based Image Super-Resolution (Ref Image SR) under a realistic setting, we contribute the Webly-Referenced SR (WR-SR) dataset, mimicking the practical usage scenario. We also extend$C^{2}$-Matching to Reference-based Video Super-Resolution (Ref VSR) task, where an image taken in a similar scene serves as the HR reference image. Extensive experiments demonstrate that our proposed$C^{2}$-Matching significantly outperforms state of the arts by up to 0.7dB on the standard CUFED5 benchmark and also boosts the performance of video super-resolution by incorporating the$C^{2}$-Matching component into Video SR pipelines. Notably,$C^{2}$-Matching also shows great generalizability on WR-SR dataset as well as robustness across large scale and rotation transformations. Codes and datasets are available athttps://github.com/yumingj/C2-Matching.
Yuming Jiang 0003, Kelvin C. K. Chan, Xintao Wang 0002, Chen Change Loy, Ziwei Liu 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 StyleGAN-Human: A Data-Centric Odyssey of Human Generation
Jianglin Fu, Shikai Li, Yuming Jiang 0003, Kwan-Yee Lin, Chen Qian 0006, Chen Change Loy, Wayne Wu, Ziwei Liu 0002
ECCV (16)3
2022 Text2Human: text-driven controllable human image generation
abstract
Generating high-quality and diverse human images is an important yet challenging task in vision and graphics. However, existing generative models often fall short under the high diversity of clothing shapes and textures. Furthermore, the generation process is even desired to be intuitively controllable for layman users. In this work, we present a text-driven controllable framework, Text2Human, for a high-quality and diverse human generation. We synthesize full-body human images starting from a given human pose with two dedicated steps. 1) With some texts describing the shapes of clothes, the given human pose is first translated to a human parsing map. 2) The final human image is then generated by providing the system with more attributes about the textures of clothes. Specifically, to model the diversity of clothing textures, we build a hierarchical texture-aware codebook that stores multi-scale neural representations for each type of texture. The codebook at the coarse level includes the structural representations of textures, while the codebook at the fine level focuses on the details of textures. To make use of the learned hierarchical codebook to synthesize desired images, a diffusion-based transformer sampler with mixture of experts is firstly employed to sample indices from the coarsest level of the codebook, which then is used to predict the indices of the codebook at finer levels. The predicted indices at different levels are translated to human images by the decoder learned accompanied with hierarchical codebooks. The use of mixture-of-experts allows for the generated image conditioned on the fine-grained text input. The prediction for finer level indices refines the quality of clothing textures. Extensive quantitative and qualitative evaluations demonstrate that our proposed Text2Human framework can generate more diverse and realistic human images compared to state-of-the-art methods. Our project page is https://yumingj.github.io/projects/Text2Human.html. Code and pretrained models are available at https://github.com/yumingj/Text2Human.
Yuming Jiang 0003, Shuai Yang 0001, Haonan Qiu, Wayne Wu, Chen Change Loy, Ziwei Liu 0002
ACM Trans. Graph.1
2021 Robust Reference-Based Super-Resolution via C2-Matching
abstract
Reference-based Super-Resolution (Ref-SR) has recently emerged as a promising paradigm to enhance a low-resolution (LR) input image by introducing an additional high-resolution (HR) reference image. Existing Ref-SR methods mostly rely on implicit correspondence matching to borrow HR textures from reference images to compensate for the information loss in input images. However, performing local transfer is difficult because of two gaps between input and reference images: the transformation gap (e.g. scale and rotation) and the resolution gap (e.g. HR and LR). To tackle these challenges, we propose C2-Matching in this work, which produces explicit robust matching crossing transformation and resolution. 1) For the transformation gap, we propose a contrastive correspondence network, which learns transformation-robust correspondences using augmented views of the input image. 2) For the resolution gap, we adopt a teacher-student correlation distillation, which distills knowledge from the easier HR-HR matching to guide the more ambiguous LR-HR matching. 3) Finally, we design a dynamic aggregation module to address the potential misalignment issue. In addition, to faithfully evaluate the performance of Ref-SR under a realistic setting, we contribute the Webly-Referenced SR (WR-SR) dataset, mimicking the practical usage scenario. Extensive experiments demonstrate that our proposed C2-Matching significantly outperforms state of the arts by over 1dB on the standard CUFED5 benchmark. Notably, it also shows great generalizability on WR-SR dataset as well as robustness across large scale and rotation transformations1.
Yuming Jiang 0003, Kelvin C. K. Chan, Xintao Wang 0002, Chen Change Loy, Ziwei Liu 0002
CVPR1
2021 Talk-to-Edit: Fine-Grained Facial Editing via Dialog
abstract
Facial editing is an important task in vision and graphics with numerous applications. However, existing works are incapable to deliver a continuous and fine-grained editing mode (e.g., editing a slightly smiling face to a big laughing one) with natural interactions with users. In this work, we propose Talk-to-Edit, an interactive facial editing framework that performs fine-grained attribute manipulation through dialog between the user and the system. Our key insight is to model a continual "semantic field" in the GAN latent space. 1) Unlike previous works that regard the editing as traversing straight lines in the latent space, here the fine-grained editing is formulated as finding a curving trajectory that respects fine-grained attribute landscape on the semantic field. 2) The curvature at each step is location-specific and determined by the input image as well as the users’ language requests. 3) To engage the users in a meaningful dialog, our system generates language feedback by considering both the user request and the current state of the semantic field.We also contribute CelebA-Dialog, a visual-language facial editing dataset to facilitate large-scale study. Specifically, each image has manually annotated fine-grained attribute annotations as well as template-based textual descriptions in natural language. Extensive quantitative and qualitative experiments demonstrate the superiority of our framework in terms of 1) the smoothness of fine-grained editing, 2) the identity/attribute preservation, and 3) the visual photorealism and dialog fluency. Notably, user study validates that our overall system is consistently favored by around 80% of the participants. Our project page is https://www.mmlab-ntu.com/project/talkedit/.
Yuming Jiang 0003, Xingang Pan, Chen Change Loy, Ziwei Liu 0002
ICCV1