Wendi Zheng

dblp:242/3819 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
12since 2021 · last 2026
0009-0008-7313-839XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation
abstract
Visual generative models have achieved remarkable progress in synthesizing photorealistic images and videos, yet aligning their outputs with human preferences across critical dimensions remains a persistent challenge. Though reinforcement learning from human feedback offers promise for preference alignment, existing reward models for visual generation face limitations, including black-box scoring without interpretability and potentially resultant unexpected biases. We present VisionReward, a general framework for learning human visual preferences in both image and video generation. Specifically, we employ a hierarchical visual assessment framework to capture fine-grained human preferences, and leverages linear weighting to enable interpretable preference learning. Furthermore, we propose a multi-dimensional consistent strategy when using VisionReward as a reward model during preference optimization for visual generation. Experiments show that VisionReward can significantly outperform existing image and video reward models on both machine metrics and human evaluation. Notably, VisionReward surpasses VideoScore by 17.2% in preference prediction accuracy, and text-to-video models with VisionReward achieve a 31.6% higher pairwise win rate compared to the same models using VideoScore.
Jiazheng Xu, Yuanming Yang, Wenbo Duan, Shen Yang 0001, Qunlin Jin, Shurun Li, Jiayan Teng, Zhuoyi Yang, Wendi Zheng, Xiao Liu 0036, Ming Ding 0004, Shiyu Huang 0001, Xiaotao Gu, Minlie Huang, Jie Tang 0001, Yuxiao Dong
AAAI13
2025 CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
abstract
We present CogVideoX, a large-scale text-to-video generation model based on diffusion transformer, which can generate 10-second continuous videos that align seamlessly with text prompts, with a frame rate of 16 fps and resolution of 768 x 1360 pixels. Previous video generation models often struggled with limited motion and short durations. It is especially difficult to generate videos with coherent narratives based on text. We propose several designs to address these issues. First, we introduce a 3D Variational Autoencoder (VAE) to compress videos across spatial and temporal dimensions, enhancing both the compression rate and video fidelity. Second, to improve text-video alignment, we propose an expert transformer with expert adaptive LayerNorm to facilitate the deep fusion between the two modalities. Third, by employing progressive training and multi-resolution frame packing, CogVideoX excels at generating coherent, long-duration videos with diverse shapes and dynamic movements. In addition, we develop an effective pipeline that includes various pre-processing strategies for text and video data. Our innovative video captioning model significantly improves generation quality and semantic alignment. Results show that CogVideoX achieves state-of-the-art performance in both automated benchmarks and human evaluation. We publish the code and model checkpoints of CogVideoX along with our VAE model and video captioning model at https://github.com/THUDM/CogVideo.
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding 0004, Shiyu Huang 0001, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Guanyu Feng, Da Yin, Yean Cheng, Bin Xu 0001, Xiaotao Gu, Yuxiao Dong, Jie Tang 0001
ICLR3
2025 ZeroFlow: Overcoming Catastrophic Forgetting is Easier than You Think
abstract
Backpropagation provides a generalized configuration for overcoming catastrophic forgetting. Optimizers such as SGD and Adam are commonly used for weight updates in continual learning and continual pre-training. However, access to gradient information is not always feasible in practice due to black-box APIs, hardware constraints, or non-differentiable systems, a challenge we refer to as the gradient bans. To bridge this gap, we introduce ZeroFlow, the first benchmark designed to evaluate gradient-free optimization algorithms for overcoming forgetting. ZeroFlow examines a suite of forward pass-based methods across various algorithms, forgetting scenarios, and datasets. Our results show that forward passes alone can be sufficient to mitigate forgetting. We uncover novel optimization principles that highlight the potential of forward pass-based methods in mitigating forgetting, managing task conflicts, and reducing memory demands. Additionally, we propose new enhancements that further improve forgetting resistance using only forward passes. This work provides essential tools and insights to advance the development of forward-pass-based methods for continual learning.
Tao Feng 0014, Didi Zhu, Hangjie Yuan, Wendi Zheng, Jie Tang 0001
ICML5
2024 Inf-DiT: Upsampling Any-Resolution Image with Memory-Efficient Diffusion Transformer
Zhuoyi Yang, Heyang Jiang, Wenyi Hong, Jiayan Teng, Wendi Zheng, Yuxiao Dong, Ming Ding 0004, Jie Tang 0001
ECCV (83)5
2024 CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion
Wendi Zheng, Jiayan Teng, Zhuoyi Yang, Jidong Chen, Xiaotao Gu, Yuxiao Dong, Ming Ding 0004, Jie Tang 0001
ECCV (77)1
2024 Relay Diffusion: Unifying diffusion process across resolutions for image synthesis
abstract
Diffusion models achieved great success in image synthesis, but still face challenges in high-resolution generation. Through the lens of discrete cosine transformation, we find the main reason is that *the same noise level on a higher resolution results in a higher Signal-to-Noise Ratio in the frequency domain*. In this work, we present Relay Diffusion Model (RDM), which transfers a low-resolution image or noise into an equivalent high-resolution one for diffusion model via blurring diffusion and block noise. Therefore, the diffusion process can continue seamlessly in any new resolution or model without restarting from pure noise or low-resolution conditioning. RDM achieves state-of-the-art FID on CelebA-HQ and sFID on ImageNet 256$\times$256, surpassing previous works such as ADM, LDM and DiT by a large margin. All the codes and checkpoints are open-sourced at \url{https://github.com/THUDM/RelayDiffusion}.
Jiayan Teng, Wendi Zheng, Ming Ding 0004, Wenyi Hong, Jianqiao Wangni, Zhuoyi Yang, Jie Tang 0001
ICLR2
2024 MACRec: A Multi-Agent Collaboration Framework for Recommendation
abstract
LLM-based agents have gained considerable attention for their decision-making skills and ability to handle complex tasks.Recognizing the current gap in leveraging agent capabilities for multiagent collaboration in recommendation systems, we introduce MACRec, a novel framework designed to enhance recommendation systems through multi-agent collaboration.Unlike existing work on using agents for user/item simulation, we aim to deploy multiagents to tackle recommendation tasks directly.In our framework, recommendation tasks are addressed through the collaborative efforts of various specialized agents, including Manager, User/Item Analyst, Reflector, Searcher, and Task Interpreter, with different working flows.Furthermore, we provide application examples of how developers can easily use MACRec on various recommendation tasks, including rating prediction, sequential recommendation, conversational recommendation, and explanation generation of recommendation results.The framework and demonstration video are publicly available at https://github.com/wzf2000/MACRec.
Zhefan Wang 0001, Yuanqing Yu, Wendi Zheng, Weizhi Ma, Min Zhang 0006
SIGIR3
2023 CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
Wenyi Hong, Ming Ding 0004, Wendi Zheng, Xinghan Liu, Jie Tang 0001
ICLR3
2023 GLM-130B: An Open Bilingual Pre-trained Model
Aohan Zeng, Xiao Liu 0036, Zhengxiao Du, Hanyu Lai, Ming Ding 0004, Zhuoyi Yang, Yifan Xu 0014, Wendi Zheng, Weng Lam Tam, Zixuan Ma, Jidong Zhai, Zhiyuan Liu 0001, Peng Zhang 0077, Yuxiao Dong, Jie Tang 0001
ICLR9
2023 Graph over-parameterization: Why the graph helps the training of deep graph convolutional network
Yucong Lin, Silu Li, Jiaxing Xu, Wendi Zheng
Neurocomputing6
2022 CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers
abstract
Development of transformer-based text-to-image models is impeded by its slow generation and complexity, for high-resolution images. In this work, we put forward a solution based on hierarchical transformers and local parallel autoregressive generation. We pretrain a 6B-parameter transformer with a simple and flexible self-supervised task, a cross-modal general language model (CogLM), and fine-tune it for fast super-resolution. The new text-to-image system, CogView2, shows very competitive generation compared to concurrent state-of-the-art DALL-E-2, and naturally supports interactive text-guided editing on images.
Ming Ding 0004, Wendi Zheng, Wenyi Hong, Jie Tang 0001
NeurIPS2
2021 CogView: Mastering Text-to-Image Generation via Transformers
abstract
Text-to-Image generation in the general domain has long been an open problem, which requires both a powerful generative model and cross-modal understanding. We propose CogView, a 4-billion-parameter Transformer with VQ-VAE tokenizer to advance this problem. We also demonstrate the finetuning strategies for various downstream tasks, e.g. style learning, super-resolution, text-image ranking and fashion design, and methods to stabilize pretraining, e.g. eliminating NaN losses. CogView achieves the state-of-the-art FID on the blurred MS COCO dataset, outperforming previous GAN-based models and a recent similar work DALL-E.
Ming Ding 0004, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou 0005, Da Yin, Junyang Lin, Xu Zou 0001, Zhou Shao, Hongxia Yang, Jie Tang 0001
NeurIPS4