VLDB 2026 Research / reviewers in the wild / expert
Liuhan Chen
dblp:359/4146
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0002-6083-8529ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Next Patch Prediction for AutoRegressive Visual GenerationabstractAutoregressive models, built based on the Next Token Prediction (NTP) paradigm, show great potential in developing a unified framework that integrates both language and vision tasks. Pioneering works introduce NTP to autoregressive visual generation tasks. In this work, we rethink the NTP for autoregressive image generation and extend it to a novel Next Patch Prediction (NPP) paradigm. Our key idea is to group and aggregate image tokens into patch tokens with higher information density. By using patch tokens as a more compact input sequence, the autoregressive model is trained to predict the next patch, significantly reducing computational costs. To further exploit the natural hierarchical structure of image data, we propose a multi-scale coarse-to-fine patch grouping strategy. With this strategy, the training process begins with a large patch size and ends with vanilla NTP where the patch size is 1x1, thus maintaining the original inference process without modifications. Extensive experiments across a diverse range of model sizes demonstrate that NPP could reduce the training cost to around 0.6 times while improving image generation quality by up to 1.0 FID score on the ImageNet 256x256 generation benchmark. Notably, our method retains the original autoregressive model architecture without introducing additional trainable parameters or specifically designing a custom image tokenizer, offering a flexible and plug-and-play solution for enhancing autoregressive visual generation. Yatian Pang, Peng Jin 0001, Bin Lin 0014, Chaoran Feng 0001, Zhenyu Tang 0004, Liuhan Chen, Francis E. H. Tay, Ser-Nam Lim, Harry Yang, Li Yuan 0007 |
AAAI | 8 |
| 2025 | WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion ModelabstractVideo Variational Autoencoder (VAE) encodes videos into a low-dimensional latent space, becoming a key component of most Latent Video Diffusion Models (LVDMs) to reduce model training costs. However, as the resolution and duration of generated videos increase, the encoding cost of Video VAEs becomes a limiting bottleneck in training LVDMs. Moreover, the block-wise inference method adopted by most LVDMs can lead to discontinuities of latent space when processing long-duration videos. The key to addressing the computational bottleneck lies in decomposing videos into distinct components and efficiently encoding the critical information. Wavelet transform can decompose videos into multiple frequency-domain components and improve the efficiency significantly, we thus propose Wavelet Flow VAE (WF-VAE), an autoencoder that leverages multi-level wavelet transform to facilitate low-frequency energy flow into latent representation. Furthermore, we introduce a method called Causal Cache, which maintains the integrity of latent space during block-wise inference. Compared to state-of-the-art video VAEs, WF-VAE demonstrates superior performance in both PSNR and LPIPS metrics, achieving 2× higher throughput and 4× lower memory consumption while maintaining competitive reconstruction quality. Our code and models are available at https://github.com/PKU-YuanGroup/WF-VAE. Zongjian Li, Bin Lin 0014, Liuhan Chen, Xinhua Cheng, Shenghai Yuan 0002, Li Yuan 0007 |
CVPR | 4 |
| 2025 | Identity-Preserving Text-to-Video Generation by Frequency DecompositionabstractIdentity-Preserving text-to-video (IPT2V) generation aims to create high-fidelity videos with consistent human identity. It is an important task in video generation but remains an open problem for generative models. This paper pushes the technical frontier of IPT2V in two directions that have not been resolved in the literature: (1) A tuning-free pipeline without tedious case-by-case finetuning, and (2) A frequency-aware heuristic identity-preserving Diffusion Transformer (DiT)-based control scheme. To achieve these goals, we propose ConsisID, a tuning-free DiT-based controllable IPT2V model to keep human-identity consistent in the generated video. Inspired by prior findings in frequency analysis of vision/diffusion transformers, it employs identity-control signals base on frequency domain, since facial features can be decomposed into low-frequency global features (e.g., profile, proportions) and high-frequency intrinsic features (e.g., identity markers that remain unaffected by pose changes). Extensive experiments demonstrate that our frequency-aware heuristic scheme provides an optimal control solution for DiT-based models, making strides towards more effective IPT2V. Shenghai Yuan 0002, Jinfa Huang, Xianyi He, Yunyang Ge, Yujun Shi, Liuhan Chen, Jiebo Luo 0001, Li Yuan 0007 |
CVPR | 6 |
| 2025 | OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion ModelabstractVariational Autoencoder (VAE), compressing videos into latent representations, is a crucial preceding component of Latent Video Diffusion Models (LVDMs). However, most LVDMs utilize 2D image VAE, which only compresses video spatially. This will lead to temporally redundant representations, reducing the efficiency of LVDMs. To eliminate this issue, we propose an omni-dimension compression VAE, named OD-VAE, which can temporally and spatially compress videos based on 3D-Causal-CNN architecture. To obtain a better trade-off between video reconstruction quality and compression speed, we further introduce and analyze four model variants of OD-VAE. In addition, a novel initialization method is designed to train our OD-VAE more efficiently, and a novel inference strategy is proposed to enable OD-VAE to handle videos of arbitrary length with limited GPU memory. Comprehensive experiments on video reconstruction and LVDM-based video generation demonstrate the effectiveness and efficiency of our proposed methods. The source code and models are available at here. Liuhan Chen, Zongjian Li, Bin Lin 0014, Qian Wang 0062, Shenghai Yuan 0002, Xinhua Cheng, Li Yuan 0007 |
ICME | 1 |
| 2024 | Prompt2Poster: Automatically Artistic Chinese Poster Creation from Prompt OnlyabstractAs a critical component in graphic design, artistic posters are widely applied in the advertising and entertainment industry, thus the automatic poster creation from user-provided prompts has become increasingly desired recently. Although existing Text2Image methods create impressive images aligned with given prompts, they fail to generate ideal artistic posters, especially with Chinese texts. To create desired artistic Chinese posters including an aligned background, reasonable layouts, and stylized graphical texts from given prompts only, we propose an automatic poster creation framework, named Prompt2Poster. Our framework utilizes the capacity of the powerful Large Language Model (LLM) to extract user intention from provided prompts and generate the aligned background. Although only taking a user prompt as the input, linguistic, visual, and geometrical information is fully utilized in the framework, bringing the ability to fit different distributions. To achieve the use of multi-modal information in the framework, two carefully designed modules, Controllable Layout Generator (CLG) and Graphical Text Generator (GTG) are proposed, leading to accurate and pleasurable visual results. Comprehensive experiments demonstrate that our Prompt2Poster achieves superior performance, especially in text quality and visual harmony. Yunyang Ge, Liuhan Chen, Haiyang Zhou, Qian Wang 0062, Xinhua Cheng, Li Yuan 0007 |
ACM Multimedia | 3 |
| 2023 | End-to-end XY Separation for Single Image Blind DeblurringabstractSingle image blind deblurring, only exploiting a blurry observation to reconstruct the sharp image, is a popular yet challenging low-level vision task. Current state-of-the-art deblurring networks mainly follow the coarse-to-fine strategy for architecture design and utilize U-net or its variant, XYDeblur, as the basic units. However, the one-encoder-one-decoder and the recently proposed one-encoder-two-decoder structures of basic units both fail to comprehensively take advantage of the directional separability of 2D deblurring, which increases the learning content of networks, thus leading to performance degradation. To thoroughly decouple the deblurring into two spatially orthogonal parts, we propose a novel substitution for U-net and its variant, called XYU-net. Specifically, it consists of two structurally identical U-nets, named XU-net and YU-net. They share orthogonal parameters by rotating kernels and focus on restoring a 2D blurry image in two spatially orthogonal directions respectively, which not only brings efficiency enhancement but also maintains parameter number. To further reduce the graphics memory demand of XYU-net, we transfer some non-linear transform modules (NLTM) from the outside of the network to its inside and propose the modified version, called MXYU-net. Experimental results on three large blurry image datasets demonstrate the efficiency of XYU-net and MXYU-net compared with U-net and XYDeblur, both as standalone models and as basic units of advanced U-net-based deblurring networks. Liuhan Chen, Yirou Wang, Yongyong Chen |
ACM Multimedia | 1 |