Zhekai Chen

dblp:361/2345 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Generative modeling · 55% Vision and language · 27% Language models and text generation · 19%
Computer graphics and multimedia
4 papers
Visual content generation and editing · 81% Image and video processing · 19%

Topics — the 17 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
autoregressive model
1.722025
Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation · NeurIPS 2025
TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation · NeurIPS 2025
Machine learning › Generative modeling › autoregressive model
autoregressive image generation
0.912025
TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation · NeurIPS 2025
Natural language and speech › Language models and text generation › decoding › decoding strategy
parallel decoding
0.912025
Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation · NeurIPS 2025
Machine learning › Generative modeling › cross-modal generation
story visualization
0.912025
AutoStory: Generating Diverse Storytelling Images with Minimal Human Efforts · Int. J. Comput. Vis. 2025
Natural language and speech › Language models and text generation
test-time scaling
0.912025
TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
AutoStory: Generating Diverse Storytelling Images with Minimal Human Efforts · Int. J. Comput. Vis. 2025
Visual content generation and editing › image generation
autoregressive image generation
0.912025
TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation · NeurIPS 2025
Visual content generation and editing › image editing
image morphing
0.912025
Framer: Interactive Frame Interpolation · ICLR 2025
Visual content generation and editing › image generation
text-to-image generation
0.912025
Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation · NeurIPS 2025
Image and video processing
video frame interpolation
0.912025
Framer: Interactive Frame Interpolation · ICLR 2025
Computer vision › Vision and language › image captioning › description generation
detailed image description
0.812024
Image Textualization: An Automatic Framework for Generating Rich and Detailed Image Descriptions · NeurIPS 2024
Machine learning › Generative modeling
diffusion model
0.812024
FreeCompose: Generic Zero-Shot Image Composition with Diffusion Prior · ECCV (17) 2024
Computer vision › Vision and language
image captioning
0.812024
Image Textualization: An Automatic Framework for Generating Rich and Detailed Image Descriptions · NeurIPS 2024
Computer vision › Vision and language › vision-language dataset
vision-language dataset construction
0.812024
Image Textualization: An Automatic Framework for Generating Rich and Detailed Image Descriptions · NeurIPS 2024
Visual content generation and editing › image editing
image compositing
0.812024
FreeCompose: Generic Zero-Shot Image Composition with Diffusion Prior · ECCV (17) 2024
Visual content generation and editing › video generation › controllable video generation
trajectory control
0.312025
Framer: Interactive Frame Interpolation · ICLR 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.212024
Image Textualization: An Automatic Framework for Generating Rich and Detailed Image Descriptions · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

speculative decoding · 1.7resampling-based potential selection · 1.7jacobi iteration · 1.7diffusion model · 1.7denoising · 1.7clustering-based diversity search · 1.7adaptive batch size schedule · 1.7zero-shot image composition · 1.5keypoint-based trajectory control · 0.9vision expert models · 0.8multimodal large language model · 0.8
YearPublicationVenuePosition
2025 Framer: Interactive Frame Interpolation
abstract
We propose Framer for interactive frame interpolation, which targets producing smoothly transitioning frames between two images as per user creativity. Concretely, besides taking the start and end frames as inputs, our approach supports customizing the transition process by tailoring the trajectory of some selected keypoints. Such a design enjoys two clear benefits. First, incorporating human interaction mitigates the issue arising from numerous possibilities of transforming one image to another, and in turn enables finer control of local motions. Second, as the most basic form of interaction, keypoints help establish the correspondence across frames, enhancing the model to handle challenging cases (e.g., objects on the start and end frames are of different shapes and styles). It is noteworthy that our system also offers an "autopilot" mode, where we introduce a module to estimate the keypoints and refine the trajectory automatically, to simplify the usage in practice. Extensive experimental results demonstrate the appealing performance of Framer on various applications, such as image morphing, time-lapse video generation, cartoon interpolation, etc. The code, model, and interface are publicly accessible at https://github.com/aim-uofa/Framer.
Wen Wang 0015, Qiuyu Wang, Kecheng Zheng, Hao Ouyang, Zhekai Chen, Biao Gong, Hao Chen 0041, Yujun Shen, Chunhua Shen
ICLR5
2025 TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation
abstract
Scaling visual generation models is essential for real-world content creation, yet requires substantial training and computational expenses. Alternatively, test-time scaling has garnered growing attention due to resource efficiency and promising performance. In this work, we present the first general test-time scaling framework for visual auto-regressive (VAR) models, TTS-VAR, modeling the generation process as a path searching problem. Inspired by VAR's hierarchical coarse-to-fine multi-scale generation, our framework integrates two key components: (i) At coarse scales, we observe that generated tokens are hard for evaluation, possibly leading to erroneous acceptance of inferior samples or rejection of superior samples. Noticing that the coarse scales contain sufficient structural information, we propose clustering-based diversity search. It preserves structural variety through semantic feature clustering, enabling later selection on samples with higher potential. (ii) In fine scales, resampling-based potential selection prioritizes promising candidates using potential scores, which are defined as reward functions incorporating multi-scale generation history. To dynamically balance computational efficiency with exploration capacity, we further introduce an adaptive descending batch size schedule throughout the causal generation process. Experiments on the powerful VAR model Infinity2B show a notable 8.7% GenEval score improvement (0.69→0.75). Key insights reveal that early-stage structural features effectively influence final quality, and resampling efficacy varies across generation scales.
Zhekai Chen, Ruihang Chu, Yukang Chen, Shiwei Zhang 0001, Yujie Wei 0001, Yingya Zhang, Xihui Liu
NeurIPS1
2025 Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
abstract
We present Wan-Move, a simple and scalable framework that brings motion control to video generative models. Existing motion-controllable methods typically suffer from coarse control granularity and limited scalability, leaving their outputs insufficient for practical use. We narrow this gap by achieving precise and high-quality motion control. Our core idea is to directly make the original condition features motion-aware for guiding video synthesis. To this end, we first represent object motions with dense point trajectories, allowing fine-grained control over the scene. We then project these trajectories into latent space and propagate the first frame's features along each trajectory, producing an aligned spatiotemporal feature map that tells how each scene element should move. This feature map serves as the updated latent condition, which is naturally integrated into the off-the-shelf image-to-video model, e.g., Wan-I2V-14B, as motion guidance without any architecture change. It removes the need for auxiliary motion encoders and makes fine-tuning base models easily scalable. Through scaled training, Wan-Move generates 5-second, 480p videos whose motion controllability rivals Kling 1.5 Pro's commercial Motion Brush, as indicated by user studies. To support comprehensive evaluation, we further design MoveBench, a rigorously curated benchmark featuring diverse content categories and hybrid-verified annotations. It is distinguished by larger data volume, longer video durations, and high-quality motion annotations. Extensive experiments on MoveBench and the public dataset consistently show Wan-Move's superior motion quality. Code, models, and benchmark data are made available.
Ruihang Chu, Yefei He, Zhekai Chen, Shiwei Zhang 0001, Xiaogang Xu 0002, Dingdong Wang, Hongwei Yi, Xihui Liu, Hengshuang Zhao, Yu Liu 0063, Yingya Zhang, Yujiu Yang 0001
NeurIPS3
2025 Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation
abstract
As a new paradigm of visual content generation, autoregressive text-to-image models suffer from slow inference due to their sequential token-by-token decoding process, often requiring thousands of model forward passes to generate a single image. To address this inefficiency, we propose Speculative Jacobi-Denoising Decoding (SJD2), a framework that incorporates the denoising process into Jacobi iterations to enable parallel token generation in autoregressive models. Our method introduces a next-clean-token prediction paradigm that enables the pre-trained autoregressive models to accept noise-perturbed token embeddings and predict the next clean tokens through low-cost fine-tuning. This denoising paradigm guides the model towards more stable Jacobi trajectories. During inference, our method initializes token sequences with Gaussian noise and performs iterative next-clean-token-prediction in the embedding space. We employ a probabilistic criterion to verify and accept multiple tokens in parallel, and refine the unaccepted tokens for the next iteration with the denoising trajectory. Experiments show that our method can accelerate generation by reducing model forward passes while maintaining the visual quality of generated images.
Yao Teng, Fuyun Wang, Zhekai Chen, Yu Wang 0002, Zhenguo Li, Weiyang Liu, Difan Zou, Xihui Liu
NeurIPS4
2025 AutoStory: Generating Diverse Storytelling Images with Minimal Human Efforts
Wen Wang 0015, Canyu Zhao, Hao Chen 0041, Zhekai Chen, Kecheng Zheng, Chunhua Shen
Int. J. Comput. Vis.4
2024 FreeCompose: Generic Zero-Shot Image Composition with Diffusion Prior
Zhekai Chen, Wen Wang 0015, Zhen Yang 0009, Zeqing Yuan, Hao Chen 0041, Chunhua Shen
ECCV (17)1
2024 Image Textualization: An Automatic Framework for Generating Rich and Detailed Image Descriptions
abstract
Image description datasets play a crucial role in the advancement of various applications such as image understanding, text-to-image generation, and text-image retrieval. Currently, image description datasets primarily originate from two sources. One source is the scraping of image-text pairs from the web. Despite their abundance, these descriptions are often of low quality and noisy. Another way is through human labeling. Datasets such as COCO are generally very short and lack details. Although detailed image descriptions can be annotated by humans, the high cost limits their quantity and feasibility. These limitations underscore the need for more efficient and scalable methods to generate accurate and detailed image descriptions. In this paper, we propose an innovative framework termed Image Textualization, which automatically produces high-quality image descriptions by leveraging existing mult-modal large language models (MLLMs) and multiple vision expert models in a collaborative manner. We conduct various experiments to validate the high quality of the descriptions constructed by our framework. Furthermore, we show that MLLMs fine-tuned on our dataset acquire an unprecedented capability of generating richer image descriptions, substantially increasing the length and detail of their output with even less hallucinations.
Renjie Pi, Jianshu Zhang 0003, Rui Pan 0002, Zhekai Chen, Tong Zhang 0001
NeurIPS5