Kaidong Zhang

dblp:172/6472 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021
YearPublicationVenuePosition
2026 StableV2V: Stabilizing Shape Consistency in Video-to-Video Editing
abstract
Recent advancements in generative artificial intelligence have significantly promoted content creation and editing, where prevailing studies further extend this exciting progress to video editing. These studies mainly transfer the inherent motion patterns from the source videos to the edited ones, where they often produce inferior results with inconsistency to user intentions, especially when shape changes between the edited and original objects might occur, due to the lack of particular alignments between the delivered motions and edited content. To address this limitation, we present a shape-consistent video editing method, namely StableV2V. Our method decomposes the entire editing pipeline into several sequential procedures, where we first edit the initial video frame, then simulate the shape-aware alignment between the delivered motions and edited sequence, and propagate the edited content to all other frames based on such alignment. Furthermore, we curate a testing benchmark, namely DAVIS-Edit, to offer a comprehensive evaluation of video editing, considering various types of prompts and difficulties. Experimental results and analyses illustrate the superior performance, visual consistency, and inference efficiency of our proposed method compared to existing state-of-the-art video editing studies.
Chang Liu 0165, Kaidong Zhang, Yunwei Lan, Dong Liu 0002
IEEE Trans. Circuits Syst. Video Technol.3
2026 Drim-NeRF: Diffusion-Based Restoration for Improving Neural Radiance Fields
abstract
The rendering degradations produced by Neural Radiance Field (NeRF) is a long-standing but complex issue in the field of 3D implicit representation, which arises from a multitude of intricate causes and was not entirely solved by designing complicated scene parameterization methods before. In this paper, we present a diffusion-based restoration method for improving Neural Radiance Field (Drim-NeRF). We consider the NeRF enhancement issue from a low-level restoration perspective by viewing all types of rendering artifacts as a specific degradation model added to clean ground truths. By leveraging the powerful prior knowledge encapsulated in diffusion model, we could restore the high-realism improved renderings conditioned on the raw low-quality rendering counterparts. To further ensure the multi-view consistent rendering enhancement, we innovatively propose to adopt optical flow warping to reduce temporal inconsistency and employ feature-wrapping in VAE decoder to improve fidelity. Our proposed method is easy to implement and agnostic to various NeRF backbones. We conduct extensive experiments on challenging large-scale urban scenes and unbounded 360-degree scenes, as well as other baselines and datasets and achieve substantial qualitative and quantitative improvements, both in the restoration quality and the multi-view consistency perspective.
Ganlin Yang, Kaidong Zhang, Jingjing Fu, Dong Liu 0002
IEEE Trans. Circuits Syst. Video Technol.2
2026 LaCon: Late-Constraint Controllable Visual Generation
abstract
Diffusion models have demonstrated impressive abilities in generating photo-realistic and creative images. To offer more controllability for the generation process of diffusion models, previous studies normally adopt extra modules to integrate condition signals by manipulating the intermediate features of the noise predictors, where they often fail in conditions not seen in the training. Although subsequent studies are motivated to handle multi-condition control, they are mostly resource-consuming to implement, where more generalizable and efficient solutions are expected for controllable visual generation. In this paper, we present a late-constraint controllable visual generation method, namely LaCon, which enables generalization across various modalities and granularities for each single-condition control. LaCon establishes an alignment between the external condition and specific diffusion timesteps, and guides diffusion models to produce conditional results based on this built alignment. Experimental results on prevailing benchmark datasets illustrate the promising performance and generalization capability of LaCon under various conditions and settings. Ablation studies analyze different components in LaCon, illustrating its great potential to offer flexible condition controls for different backbones.
Chang Liu 0165, Kaidong Zhang, Yunwei Lan, Dong Liu 0002
IEEE Trans. Image Process.3
2025 RoboPearls: Editable Video Simulation for Robot Manipulation
Tang Tao, Likui Zhang, Youpeng Wen, Kaidong Zhang, Jiawang Bian, Tianyi Yan, Kun Zhan, Peng Jia 0007, Hefeng Wu, Xiaodan Liang
ICCV4
2025 $A_{0}$: An Affordance-Aware Hierarchical Model for General Robotic Manipulation
Rongtao Xu, Youpeng Wen, Haoting Yang, Jianzheng Huang, Zhe Li 0008, Kaidong Zhang, Liqiong Wang, Yuxuan Kuang, Meng Cao 0002, Feng Zheng 0001, Xiaodan Liang
ICCV9
2025 RoBridge: A Hierarchical Architecture Bridging Cognition and Execution for General Robotic Manipulation
abstract
Operating robots in open-ended scenarios with diverse tasks is a crucial research and application direction in robotics. While recent progress in natural language processing and large multimodal models has enhanced robots' ability to understand complex instructions, robot manipulation still faces the procedural skill dilemma and the declarative skill dilemma in open environments. Existing methods often compromise cognitive and executive capabilities. To address these challenges, in this paper, we propose RoBridge, a hierarchical intelligent architecture for general robotic manipulation. It consists of a high-level cognitive planner (HCP) based on a large-scale pre-trained vision-language model (VLM), an invariant operable representation (IOR) serving as a symbolic bridge, and a generalist embodied agent (GEA). RoBridge maintains the declarative skill of VLM and unleashes the procedural skill of reinforcement learning, effectively bridging the gap between cognition and execution. RoBridge demonstrates significant performance improvements over existing baselines, achieving a 75% success rate on new tasks and an 83% average success rate in sim-to-real generalization using only five real-world data samples per task. This work represents a significant step towards integrating cognitive reasoning with physical execution in robotic systems, offering a new paradigm for general robotic manipulation.
Kaidong Zhang, Rongtao Xu, Pengzhen Ren, Junfan Lin, Hefeng Wu, Xiaodan Liang
ICCV1
2025 Surfer: A World Model-Based Framework for Vision-Language Robot Manipulation
abstract
Considering how to make the model accurately understand and follow natural language instructions and perform actions consistent with world knowledge is a key challenge in robot manipulation. This mainly includes human fuzzy instruction reasoning and the following of physical knowledge. Therefore, the embodied intelligence agent must have the ability to model world knowledge from training data. However, most existing vision and language robot manipulation methods mainly operate in less realistic simulators and language settings and lack explicit modeling of world knowledge. To bridge this gap, we introduce a novel and simple robot manipulation framework, called Surfer. It is based on the world model, treats robot manipulation as a state transfer of the visual scene, and decouples it into two parts: action and scene. Then, the generalization ability of the model on new instructions and new scenes is enhanced by explicit modeling of the action and scene prediction in multimodal information. In addition, we built a robot manipulation simulation platform that supports physics execution based on the MuJoCo physics engine. It can automatically generate demonstration training data and test data, effectively reducing labor costs. To conduct a comprehensive and systematic evaluation of the visual-language understanding and physical execution of the manipulation model, we also created a robotic manipulation benchmark with different difficulty levels, called SeaWave. It contains four visual-language manipulation tasks of different difficulty levels and can provide a standardized testing platform for embedded AI agents in multimodal environments. Overall, we hope Surfer can freely surf in the robot's SeaWave benchmark. Extensive experiments show that Surfer consistently outperforms all baselines significantly in all manipulation tasks. On average, Surfer achieved a success rate of 54.74% on the defined four levels of manipulation tasks, exceeding the best baseline performance of 51.07%. The simulator, code, and benchmarks are released at https://pzhren.github.io/Surfer.
Pengzhen Ren, Kaidong Zhang, Hetao Zheng, Yuhang Wen 0001, Fengda Zhu, Shikui Ma, Xiaodan Liang
IEEE Trans. Neural Networks Learn. Syst.2
2024 PIVOT-R: Primitive-Driven Waypoint-Aware World Model for Robotic Manipulation
abstract
Language-guided robotic manipulation is a challenging task that requires an embodied agent to follow abstract user instructions to accomplish various complex manipulation tasks. Previous work generally maps instructions and visual perceptions directly to low-level executable actions, neglecting the modeling of critical waypoints (e.g., key states of “close to/grab/move up” in action trajectories) in manipulation tasks. To address this issue, we propose a PImitive-driVen waypOinT-aware world model for Robotic manipulation (PIVOT-R) that focuses solely on the prediction of task-relevant waypoints. Specifically, PIVOT-R consists of a Waypoint-aware World Model (WAWM) and a lightweight action prediction module. The former performs primitive action parsing and primitive-driven waypoint prediction, while the latter focuses on decoding low-level actions. Additionally, we also design an asynchronous hierarchical executor (AHE) for PIVOT-R, which can use different execution frequencies for different modules of the model, thereby helping the model reduce computational redundancy and improve model execution efficiency. Our PIVOT-R outperforms state-of-the-art (SoTA) open-source models on the SeaWave benchmark, achieving an average relative improvement of 19.45% across four levels of instruction tasks. Moreover, compared to the synchronously executed PIVOT-R, the execution efficiency of PIVOT-R with AHE is increased by 28-fold, with only a 2.9% drop in performance. These results provide compelling evidence that our PIVOT-R can significantly improve both the performance and efficiency of robotic manipulation.
Kaidong Zhang, Pengzhen Ren, Bingqian Lin, Junfan Lin, Shikui Ma, Hang Xu 0004, Xiaodan Liang
NeurIPS1
2024 Exploiting Optical Flow Guidance for Transformer-Based Video Inpainting
abstract
Transformers have been widely used for video processing owing to the multi-head self attention (MHSA) mechanism. However, the MHSA mechanism encounters an intrinsic difficulty for video inpainting, since the features associated with the corrupted regions are degraded and incur inaccurate self attention. This problem, termed query degradation, may be mitigated by first completing optical flows and then using the flows to guide the self attention, which was verified in our previous work - flow-guided transformer (FGT). We further exploit the flow guidance and propose FGT++ to pursue more effective and efficient video inpainting. First, we design a lightweight flow completion network by using local aggregation and edge loss. Second, to address the query degradation, we propose a flow guidance feature integration module, which uses the motion discrepancy to enhance the features, together with a flow-guided feature propagation module that warps the features according to the flows. Third, we decouple the transformer along the temporal and spatial dimensions, where flows are used to select the tokens through a temporally deformable MHSA mechanism, and global tokens are combined with the inner-window local tokens through a dual-perspective MHSA mechanism. FGT++ is experimentally evaluated to be outperforming the existing video inpainting networks qualitatively and quantitatively.
Kaidong Zhang, Jialun Peng, Jingjing Fu, Dong Liu 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Toward Interactive Image Inpainting via Robust Sketch Refinement
abstract
One tough problem of image inpainting is to restore complex structures in the corrupted regions. It motivates interactive image inpainting which leverages additional hints, e.g., sketches, to assist the inpainting process. A sketch is simple and intuitive for end users to provide, but meanwhile has free forms with much randomness. Such randomness may confuse the inpainting models, and incur severe artifacts in completed images. To better facilitate image inpainting with sketch guidance, we propose a two-stage image inpainting system, termed SketchRefiner. The first stage of our approach serves as a data provider that simulates real sketches and derives the capability of sketch calibration from the simulated data. In the second stage, our approach aligns the sketch guidance with the inpainting process so as to elevate image inpainting with sketches. We also propose a real-world test protocol to address the evaluation of inpainting methods upon practical applications with user sketches. Experimental results on three prevailing benchmark datasets, i.e., CelebA-HQ, Places2, and ImageNet, and the proposed test protocol demonstrate the state-of-the-art performance of our approach, and its great potentials upon real-world applications. Further analyses illustrate that our approach effectively utilizes sketch information as guidance and eliminates the artifacts due to the free-form sketches.
Chang Liu 0165, Shunxin Xu, Jialun Peng, Kaidong Zhang, Dong Liu 0002
IEEE Trans. Multim.4
2022 Inertia-Guided Flow Completion and Style Fusion for Video Inpainting
abstract
Physical objects have inertia, which resists changes in the velocity and motion direction. Inspired by this, we introduce inertia prior that optical flow, which reflects object motion in a local temporal window, keeps unchanged in the adjacent preceding or subsequent frame. We propose a flow completion network to align and aggregate flow features from the consecutive flow sequences based on the inertia prior. The corrupted flows are completed under the supervision of customized losses on reconstruction, flow smoothness, and consistent ternary census transform. The completed flows with high fidelity give rise to significant improvement on the video inpainting quality. Nevertheless, the existing flow-guided cross-frame warping methods fail to consider the lightening and sharpness variation across video frames, which leads to spatial incoherence after warping from other frames. To alleviate such problem, we propose the Adaptive Style Fusion Network (ASFN), which utilizes the style information extracted from the valid regions to guide the gradient refinement in the warped regions. Moreover, we design a data simulation pipeline to reduce the training difficulty of ASFN. Extensive experiments show the superiority of our method against the state-of-the-art methods quantitatively and qualitatively. The project page is at https://github.com/hitachinsk/ISVI.
Kaidong Zhang, Jingjing Fu, Dong Liu 0002
CVPR1
2022 Flow-Guided Transformer for Video Inpainting
Kaidong Zhang, Jingjing Fu, Dong Liu 0002
ECCV (18)1