Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Tianrui Zhu

dblp:359/3888 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0009-0002-5146-019XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
2 papers
Visual content generation and editing · 100%
Artificial intelligence
1 paper
Generative modeling · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing › video generation
diffusion-based video generation
1.012026
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors · ACM Trans. Graph. 2026
Visual content generation and editing
video generation
1.012026
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors · ACM Trans. Graph. 2026
Machine learning › Generative modeling
diffusion model
0.912025
KV-Edit: Training-Free Image Editing for Precise Background Preservation · ICCV 2025
Machine learning › Generative modeling › diffusion model
image editing
0.912025
KV-Edit: Training-Free Image Editing for Precise Background Preservation · ICCV 2025
Machine learning › Generative modeling › diffusion model › image editing
training-free image editing
0.912025
KV-Edit: Training-Free Image Editing for Precise Background Preservation · ICCV 2025
Visual content generation and editing
image editing
0.312025
KV-Edit: Training-Free Image Editing for Precise Background Preservation · ICCV 2025

Methods — techniques the papers use, named apart from their topics

inversion-free method · 1.7diffusion transformer · 1.7KV cache · 1.7stochastic condition masking · 1.0diffusion prior · 1.0decoupled gated LoRA · 1.0cross-modal self-attention · 1.0
YearPublicationVenuePosition
2026 UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
abstract
Recent progress has shown that video diffusion models (VDMs) can be repurposed to solve various multimodal graphics tasks. However, existing approaches predominantly train separate models for each specific problem setting. This practice locks models into fixed input-output mappings, and typically ignores the joint correlations across modalities. In this paper, we present UniVidX , a unified multimodal framework designed to leverage VDM priors to enable versatile video generation. Our goal is to (i) master diverse pixel-aligned tasks by formulating them as conditional generation problems within multimodal space, (ii) adapt to modality-specific distributions without compromising the backbone's native priors, and (iii) ensure cross-modal consistency during synthesis. Concretely, we propose three key designs: 1) Stochastic Condition Masking (SCM): by randomly partitioning modalities into clean conditions and noisy targets during training, we enable the model to learn omni-directional conditional generation rather than fixed mappings. 2) Decoupled Gated LoRA (DGL): we attach per-modality LoRAs and activate them when a modality serves as a generation target, thereby preserving the VDM's strong priors. 3) Cross-Modal Self-Attention (CMSA): we explicitly share keys/values across modalities while maintaining modality-specific queries, facilitating information exchange and inter-modal alignment. We validate our framework by instantiating it in two domains: 1) UniVid-Intrinsic for RGB videos and their intrinsic maps (albedo, irradiance, normal), and 2) UniVid-Alpha for blended RGB videos and their constituent RGBA layers. Experimental results demonstrate that both models achieve performance competitive with state-of-the-art methods across distinct tasks. Notably, they exhibit robust generalization capabilities in in-the-wild scenarios, even when trained on limited datasets of fewer than 1k videos.
Houyuan Chen, Hong Li 0016, Xianghao Kong, Tianrui Zhu, Shaocong Xu, Weiqing Xiao, Yuwei Guo 0002, Chongjie Ye, Lvmin Zhang, Hao Zhao 0002, Anyi Rao
ACM Trans. Graph.4
2025 KV-Edit: Training-Free Image Editing for Precise Background Preservation
abstract
Background consistency remains a significant challenge in image editing tasks. Despite extensive developments, existing works still face a trade-off between maintaining similarity to the original image and generating content that aligns with the target. Here, we propose KV-Edit, a training-free approach that uses KV cache in DiTs to maintain background consistency, where background tokens are preserved rather than regenerated, eliminating the need for complex mechanisms or expensive training, ultimately generating new content that seamlessly integrates with the background within user-provided regions. We further explore the memory consumption of the KV cache during editing and optimize the space complexity to $O(1)$ using an inversion-free method. Our approach is compatible with any DiT-based generative model without additional training. Experiments demonstrate that KV-Edit significantly outperforms existing approaches in terms of both background and image quality, even surpassing training-based methods. Project webpage is available at https://xilluill.github.io/projectpages/KV-Edit
Tianrui Zhu, Jiawei Shao, Yansong Tang
ICCV1