Xiaodong Chen 0009

dblp:70/4319-9 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0003-1624-2680ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MultiPaint: A Unified Framework for Multi-Task, Multi-Object, and Multi-Condition Video Inpainting
abstract
Video inpainting modifies local regions in video while ensuring spatial and temporal coherence. However, existing methods-both traditional and recent diffusion-based ones-face key limitations: they lack unified support for both insertion and completion, and are restricted to single-object inpainting, making it difficult to handle multi-object scenarios involving grounding and interaction. In this article, we propose MultiPaint, a unified framework for multi-task, multi-object, and multi-condition video inpainting. First, we introduce dual-branch adapters to unify the insertion and completion tasks within a single model. Moreover, we propose a test-time scheduled feature composition strategy that enables multi-object inpainting with user-specified locations while better preserving interactions among objects, a setting that has been insufficiently addressed in prior work. Additionally, we introduce a multi-condition inpainting scheme that integrates text-guided, image-guided, and keyframe-guided modes via dynamic frame masking, providing more controllability in appearance customization. Extensive experiments show that MultiPaint achieves state-of-the-art performance on object insertion and scene completion among the recent works. We further demonstrate its versatility in downstream tasks including grounded video generation, object editing, object removal, image-guided inpainting, and long video inpainting.
Zheng Gu 0001, Xin Tao 0001, Pengfei Wan 0001, Xiaodong Chen 0009, Jing Liao 0001
IEEE Trans. Vis. Comput. Graph.6
2025 Few-shot exemplar-driven inpainting with parameter-efficient diffusion fine-tuning
abstract
Text-to-image diffusion models have demonstrated impressive capabilities in image generation and have been effectively applied to image inpainting. While text prompt provides an intuitive guidance for conditional inpainting, users often seek the ability to inpaint a specific object with customized appearance by providing an exemplar image. Unfortunately, existing methods struggle to achieve high fidelity in exemplar-driven inpainting. To address this, we use a plug-and-play low-rank adaptation (LoRA) module based on a pretrained text-driven inpainting model. The LoRA module is dedicated to learn the exemplar-specific concepts through few-shot fine-tuning, bringing improved fitting capability to customized exemplar images, without intensive training on large-scale datasets. Additionally, we introduce GPT-4V prompting and prior noise initialization techniques to further facilitate the fidelity in inpainting results. In brief, the denoising diffusion process first starts with the noise derived from a composite exemplar–background image, and is subsequently guided by an expressive prompt generated from the exemplar using the GPT-4V model. Extensive experiments demonstrate that our method achieves state-of-the-art performance, qualitatively and quantitatively, offering users an exemplar-driven inpainting tool with enhanced customization capability.
Zheng Gu 0001, Wenyue Hao, Yi Wang 0066, Huaiyu Cai, Xiaodong Chen 0009
Frontiers Inf. Technol. Electron. Eng.6
2025 DDM: A Metric for Comparing 3D Shapes Using Directional Distance Fields
abstract
Qualifying the discrepancy between 3D geometric models, which could be represented with either point clouds or triangle meshes, is a pivotal issue with board applications. Existing methods mainly focus on directly establishing the correspondence between two models and then aggregating point-wise distance between corresponding points, resulting in them being either inefficient or ineffective. In this paper, we propose DDM, an efficient, effective, robust, and differentiable distance metric for 3D geometry data. Specifically, we construct DDM based on the proposed implicit representation of 3D models, namely directional distance field (DDF), which defines the directional distances of 3D points to a model to capture its local surface geometry. We then transfer the discrepancy between two 3D geometric models as the discrepancy between their DDFs defined on an identical domain, naturally establishing model correspondence. To demonstrate the advantage of our DDM, we explore various distance metric-driven 3D geometric modeling tasks, including template surface fitting, rigid registration, non-rigid registration, scene flow estimation and human pose optimization. Extensive experiments show that our DDM achieves significantly higher accuracy under all tasks. As a generic distance metric, DDM has the potential to advance the field of 3D geometric modeling.
Junhui Hou, Xiaodong Chen 0009, Hongkai Xiong, Wenping Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 GeoUDF: Surface Reconstruction from 3D Point Clouds via Geometry-guided Distance Representation
abstract
We present a learning-based method, namely GeoUDF, to tackle the long-standing and challenging problem of reconstructing a discrete surface from a sparse point cloud. To be specific, we propose a geometry-guided learning method for UDF and its gradient estimation that explicitly formulates the unsigned distance of a query point as the learnable affine averaging of its distances to the tangent planes of neighboring points on the surface. Besides, we model the local geometric structure of the input point clouds by explicitly learning a quadratic polynomial for each point. This not only facilitates upsampling the input sparse point cloud but also naturally induces unoriented normal, which further augments UDF estimation. Finally, to extract triangle meshes from the predicted UDF we propose a customized edge-based marching cube module. We conduct extensive experiments and ablation studies to demonstrate the significant advantages of our method over state-of-the-art methods in terms of reconstruction accuracy, efficiency, and generality. The source code is publicly available at https://github.com/rsy6318/GeoUDF.
Junhui Hou, Xiaodong Chen 0009, Ying He 0001, Wenping Wang 0001
ICCV3
2023 Uni-paint: A Unified Framework for Multimodal Image Inpainting with Pretrained Diffusion Model
abstract
Recently, text-to-image denoising diffusion probabilistic models (DDPMs) have demonstrated impressive image generation capabilities and have also been successfully applied to image inpainting. However, in practice, users often require more control over the inpainting process beyond textual guidance, especially when they want to composite objects with customized appearance, color, shape, and layout. Unfortunately, existing diffusion-based inpainting methods are limited to single-modal guidance and require task-specific training, hindering their cross-modal scalability. To address these limitations, we propose Uni-paint, a unified framework for multimodal inpainting that offers various modes of guidance, including unconditional, text-driven, stroke-driven, exemplar-driven inpainting, as well as a combination of these modes. Furthermore, our Uni-paint is based on pretrained Stable Diffusion and does not require task-specific training on specific datasets, enabling few-shot generalizability to customized images. We have conducted extensive qualitative and quantitative evaluations that show our approach achieves comparable results to existing single-modal methods while offering multimodal inpainting capabilities not available in other methods. Code is available at https://github.com/ysy31415/unipaint.
Xiaodong Chen 0009, Jing Liao 0001
ACM Multimedia2
2023 CorrI2P: Deep Image-to-Point Cloud Registration via Dense Correspondence
abstract
Motivated by the intuition that the critical step of localizing a 2D image in the corresponding 3D point cloud is establishing 2D-3D correspondence between them, we propose the first feature-based dense correspondence framework for addressing the challenging problem of 2D image-to-3D point cloud registration, dubbed CorrI2P. CorrI2P is mainly composed of three modules, i.e., feature embedding, symmetric overlapping region detection, and pose estimation through the established correspondence. Specifically, given a pair of a 2D image and a 3D point cloud, we first transform them into high-dimensional feature spaces and feed the resulting features into a symmetric overlapping region detector to determine the region where the image and point cloud overlap. Then we use the features of the overlapping regions to establish dense 2D-3D correspondence, on which EPnP within RANSAC is performed to estimate the camera pose, i.e., translation and rotation matrices. Experimental results on KITTI and NuScenes datasets show that our CorrI2P outperforms state-of-the-art image-to-point cloud registration methods significantly. The code will be publicly available athttps://github.com/rsy6318/CorrI2P.
Yiming Zeng 0002, Junhui Hou, Xiaodong Chen 0009
IEEE Trans. Circuits Syst. Video Technol.4
2021 Learning to estimate smooth and accurate semantic correspondence
Huaiyuan Xu, Xiaodong Chen 0009, Jiaqi Xi, Jing Liao 0001
Neurocomputing2
2021 Disocclusion-type aware hole filling method for view synthesis
Xiaodong Chen 0009, Haitao Liang, Huaiyuan Xu, Huaiyu Cai, Yi Wang 0066
Multim. Tools Appl.1
2017 Fast CU partition strategy for HEVC based on Haar wavelet
abstract
As the latest video coding standard, high‐efficiency video coding (HEVC) achieves better performance and supports higher resolution compared with the predecessor standard, H.264/advanced video coding (AVC). Intra‐coding is an important feature in HEVC standard, which reduces the spatial redundancy significantly, due to the flexible coding structure, and high density of angular prediction modes. However, the improvement on coding efficiency is obtained at the expense of the extraordinary computation complexity. This study presents a novel coding unit (CU) partitioning technique for HEVC. By using a fast texture complexity detection method, which is based on two‐dimensional Haar wavelet transform, texture complexity for each CU can be extracted. According to the Haar wavelet coefficients obtained, an early CU splitting termination is proposed to decide whether a CU should be decomposed into four lower dimensions CUs or not. Experimental results demonstrate that the fast CU partition strategy achieves better trade‐off between rate‐distortion performance and complexity reduction than the previous algorithms. Compared with the reference software HM16.7, the proposed algorithm can lessen the encoding time up to 46.22% on average, with a negligible bit rate increase of 0.45%, and quality losses lower than 0.04 dB, respectively.
Xuebin Sun, Xiaodong Chen 0009, Yi Wang 0066, Daoyin Yu
IET Image Process.2