VLDB 2026 Research / reviewers in the wild / expert
Tianrui Zhu
dblp:359/3888
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2026
0009-0002-5146-019XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 100% | |
| Artificial intelligence
1 paper |
Generative modeling · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Visual content generation and editing › video generation
diffusion-based video generation |
1.0 | 1 | 2026 | UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors · ACM Trans. Graph. 2026 |
Visual content generation and editing
video generation |
1.0 | 1 | 2026 | UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors · ACM Trans. Graph. 2026 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | KV-Edit: Training-Free Image Editing for Precise Background Preservation · ICCV 2025 |
Machine learning › Generative modeling › diffusion model
image editing |
0.9 | 1 | 2025 | KV-Edit: Training-Free Image Editing for Precise Background Preservation · ICCV 2025 |
Machine learning › Generative modeling › diffusion model › image editing
training-free image editing |
0.9 | 1 | 2025 | KV-Edit: Training-Free Image Editing for Precise Background Preservation · ICCV 2025 |
Visual content generation and editing
image editing |
0.3 | 1 | 2025 | KV-Edit: Training-Free Image Editing for Precise Background Preservation · ICCV 2025 |
Methods — techniques the papers use, named apart from their topics
inversion-free method · 1.7diffusion transformer · 1.7KV cache · 1.7stochastic condition masking · 1.0diffusion prior · 1.0decoupled gated LoRA · 1.0cross-modal self-attention · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion PriorsabstractRecent progress has shown that video diffusion models (VDMs) can be repurposed to solve various multimodal graphics tasks. However, existing approaches predominantly train separate models for each specific problem setting. This practice locks models into fixed input-output mappings, and typically ignores the joint correlations across modalities. In this paper, we present UniVidX , a unified multimodal framework designed to leverage VDM priors to enable versatile video generation. Our goal is to (i) master diverse pixel-aligned tasks by formulating them as conditional generation problems within multimodal space, (ii) adapt to modality-specific distributions without compromising the backbone's native priors, and (iii) ensure cross-modal consistency during synthesis. Concretely, we propose three key designs: 1) Stochastic Condition Masking (SCM): by randomly partitioning modalities into clean conditions and noisy targets during training, we enable the model to learn omni-directional conditional generation rather than fixed mappings. 2) Decoupled Gated LoRA (DGL): we attach per-modality LoRAs and activate them when a modality serves as a generation target, thereby preserving the VDM's strong priors. 3) Cross-Modal Self-Attention (CMSA): we explicitly share keys/values across modalities while maintaining modality-specific queries, facilitating information exchange and inter-modal alignment. We validate our framework by instantiating it in two domains: 1) UniVid-Intrinsic for RGB videos and their intrinsic maps (albedo, irradiance, normal), and 2) UniVid-Alpha for blended RGB videos and their constituent RGBA layers. Experimental results demonstrate that both models achieve performance competitive with state-of-the-art methods across distinct tasks. Notably, they exhibit robust generalization capabilities in in-the-wild scenarios, even when trained on limited datasets of fewer than 1k videos. Houyuan Chen, Hong Li 0016, Xianghao Kong, Tianrui Zhu, Shaocong Xu, Weiqing Xiao, Yuwei Guo 0002, Chongjie Ye, Lvmin Zhang, Hao Zhao 0002, Anyi Rao |
ACM Trans. Graph. | 4 |
| 2025 | KV-Edit: Training-Free Image Editing for Precise Background PreservationabstractBackground consistency remains a significant challenge in image editing tasks. Despite extensive developments, existing works still face a trade-off between maintaining similarity to the original image and generating content that aligns with the target. Here, we propose KV-Edit, a training-free approach that uses KV cache in DiTs to maintain background consistency, where background tokens are preserved rather than regenerated, eliminating the need for complex mechanisms or expensive training, ultimately generating new content that seamlessly integrates with the background within user-provided regions. We further explore the memory consumption of the KV cache during editing and optimize the space complexity to $O(1)$ using an inversion-free method. Our approach is compatible with any DiT-based generative model without additional training. Experiments demonstrate that KV-Edit significantly outperforms existing approaches in terms of both background and image quality, even surpassing training-based methods. Project webpage is available at https://xilluill.github.io/projectpages/KV-Edit Tianrui Zhu, Jiawei Shao, Yansong Tang |
ICCV | 1 |