VLDB 2026 Research / reviewers in the wild / expert
Dianbing Xi
dblp:332/3729
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2026
0009-0003-1668-1082ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PFAvatar: Pose-Fusion 3D Personalized Avatar Reconstruction from Real-World Outfit-of-the-Day PhotosabstractWe propose PFAvatar (Pose-Fusion Avatar), a new method that reconstructs high-quality 3D avatars from Outfit of the Day (OOTD) photos, which exhibit diverse poses, occlusions, and complex backgrounds. Our method consists of two stages: (1) fine-tuning a pose-aware diffusion model from few-shot OOTD examples and (2) distilling a 3D avatar represented by a neural radiance field (NeRF). In the first stage, unlike previous methods that segment images into assets (e.g. garments, accessories) for 3D assembly, which is prone to inconsistency, we avoid decomposition and directly model the full-body appearance. By integrating a pre-trained ControlNet for pose estimation and a novel Condition Prior Preservation Loss (CPPL), our method enables end-to-end learning of fine details while mitigating language drift in few-shot training. Our method completes personalization in just 5 minutes, achieving a 48x speed-up compared to previous approaches. In the second stage, we introduce a NeRF-based avatar representation optimized by canonical SMPL-X space sampling and Multi-Resolution 3D-SDS. Compared to mesh-based representations that suffer from resolution-dependent discretization and erroneous occluded geometry, our continuous radiance field can preserve high-frequency textures (e.g., hair) and handle occlusions correctly through transmittance. Experiments demonstrate that PFAvatar outperforms state-of-the-art methods in terms of reconstruction fidelity, detail preservation, and robustness to occlusions/truncations, advancing practical 3D avatar generation from real-world OOTD albums. In addition, the reconstructed 3D avatars support downstream applications such as virtual try-on, animation, and human video reenactment, further demonstrating the versatility and practical value of our approach. Dianbing Xi, Guoyuan An, Jingsen Zhu, Ruiyuan Zhang, Jiayuan Lu, Yuchi Huo, Rui Wang 0004 |
AAAI | 1 |
| 2026 | OmniVDiff: Omni Controllable Video Diffusion for Generation and UnderstandingabstractIn this paper, we propose a novel framework for controllable video diffusion, OmniVDiff , aiming to synthesize and comprehend multiple video visual content in a single diffusion model. To achieve this, OmniVDiff treats all video visual modalities in the color space to learn a joint distribution, while employing an adaptive control strategy that dynamically adjusts the role of each visual modality during the diffusion process, either as a generation modality or a conditioning modality. Our framework supports three key capabilities: (1) Text-conditioned video generation, where all modalities are jointly synthesized from a textual prompt; (2) Video understanding, where structural modalities are predicted from rgb inputs in a coherent manner; and (3) X-conditioned video generation, where video synthesis is guided by finegrained inputs such as depth, canny and segmentation. Extensive experiments demonstrate that OmniVDiff achieves state-of-the-art performance in video generation tasks and competitive results in video understanding. Its flexibility and scalability make it well-suited for downstream applications such as video-to-video translation, modality adaptation for visual tasks, and scene reconstruction. Dianbing Xi, Jiepeng Wang 0005, Yuanzhi Liang, Xi Qiu, Yuchi Huo, Rui Wang 0004, Chi Zhang 0012, Xuelong Li 0001 |
AAAI | 1 |
| 2025 | IntrinsicControlNet: Cross-Distribution Image Generation with Real and Unreal
Jiayuan Lu, Rengan Xie, Zhizhen Wu, Dianbing Xi, Qi Ye 0001, Rui Wang 0004, Hujun Bao, Yuchi Huo |
ICCV | 5 |
| 2025 | Inverse Rendering using Multi-Bounce Path Tracing and Reservoir SamplingabstractWe introduce MIRReS, a novel two-stage inverse rendering framework that
jointly reconstructs and optimizes explicit geometry, materials, and lighting
from multi-view images. Unlike previous methods that rely on implicit irradiance fields or oversimplified ray tracing, our method begins with an initial
stage that extracts an explicit triangular mesh. In the second stage, we refine this representation using a physically-based inverse rendering model
with multi-bounce path tracing and Monte Carlo integration. This enables our method to accurately estimate indirect illumination effects, including self-shadowing and internal reflections, leading to a more precise
intrinsic decomposition of shape, material, and lighting. To address the
noise issue in Monte Carlo integration, we incorporate reservoir sampling,
improving convergence and enabling efficient gradient-based optimization
with low sample counts. Through both qualitative and quantitative assessments across various scenarios, especially those with complex shadows,
we demonstrate that our method achieves state-of-the-art decomposition
performance. Furthermore, our optimized explicit geometry seamlessly
integrates with modern graphics engines supporting downstream applications such as scene editing, relighting, and material editing. Yuxin Dai, Qi Wang 0111, Jingsen Zhu, Dianbing Xi, Yuchi Huo, Chen Qian 0006, Ying He 0001 |
ICLR | 4 |
| 2025 | AniTex: Light-Geometry Consistent PBR Material Generation for Animatable ObjectsabstractHigh-quality Physically-Based Rendering (PBR) materials are crucial for visual realism in 3D asset creation, yet existing methods primarily target static objects, leading to challenges in maintaining multi-frame consistency for animatable entities. To tackle this issue, we introduce AniTex, the first generative pipeline that utilizes diffusion models to synthesize high-quality PBR materials for animatable objects based on text prompts. The pipeline consists of three key stages: First, sequences of RGB images are generated using a video diffusion model conditioned on depth, normals, irradiance, and motion vectors to ensure temporal coherence and geometric alignment across multiple frames and viewpoints. Second, these RGB image sequences are decomposed into per-view, per-frame PBR material maps (albedo, roughness, metallic) by a specialized Intrinsic Diffusion Model (IDM), which is conditioned on the RGB images along with consistent geometry and lighting cues to disentangle material from illumination. Finally, these per-view, per-frame PBR maps are hierarchically blended. This process first ensures temporal coherence within each view’s frame sequence, then amalgamates these into globally consistent PBR materials for the animatable object, maintaining overall temporal coherence and visual consistency throughout its animation. Extensive experiments show that AniTex produces more realistic PBR materials for both static and animated objects, outperforming baseline methods in visual appeal. Jieting Xu, Guoyuan An, Rengan Xie, Dianbing Xi, Wenjun Song, Rui Wang 0004, Yuchi Huo |
SIGGRAPH Asia | 7 |
| 2023 | I2-SDF: Intrinsic Indoor Scene Reconstruction and Editing via Raytracing in Neural SDFsabstractIn this work, we present I2-SDF, a new method for intrinsic indoor scene reconstruction and editing using differentiable Monte Carlo raytracing on neural signed distance fields (SDFs). Our holistic neural SDF-based frame-work jointly recovers the underlying shapes, incident radiance and materials from multi-view images. We introduce a novel bubble loss for fine-grained small objects and error-guided adaptive sampling scheme to largely improve the reconstruction quality on large-scale indoor scenes. Further, we propose to decompose the neural radiance field into spatially-varying material of the scene as a neural field through surface-based, differentiable Monte Carlo raytracing and emitter semantic segmentations, which enables physically based and photorealistic scene relighting and editing applications. Through a number of qualitative and quantitative experiments, we demonstrate the superior quality of our method on indoor scene reconstruction, novel view synthesis, and scene editing compared to state-of-the-art baselines. Our project page is at https://jingsenzhu.github.io/i2-sdf. Jingsen Zhu, Yuchi Huo, Qi Ye 0001, Fujun Luan, Jifan Li, Dianbing Xi, Lisha Wang, Rui Tang 0015, Wei Hua 0002, Hujun Bao, Rui Wang 0004 |
CVPR | 6 |
| 2022 | SGW-Based Multi-task Learning in Vision Tasks
Ruiyuan Zhang, Yuyao Chen, Dianbing Xi, Yuchi Huo, Chao Wu 0001 |
ACCV (4) | 4 |
| 2022 | Learning-based Inverse Rendering of Complex Indoor Scenes with Differentiable Monte Carlo RaytracingabstractIndoor scenes typically exhibit complex, spatially-varying appearance from global illumination, making inverse rendering a challenging ill-posed problem. This work presents an end-to-end, learning-based inverse rendering framework incorporating differentiable Monte Carlo raytracing with importance sampling. The framework takes a single image as input to jointly recover the underlying geometry, spatially-varying lighting, and photorealistic materials. Specifically, we introduce a physically-based differentiable rendering layer with screen-space ray tracing, resulting in more realistic specular reflections that match the input photo. In addition, we create a large-scale, photorealistic indoor scene dataset with significantly richer details like complex furniture and dedicated decorations. Further, we design a novel out-of-view lighting network with uncertainty-aware refinement leveraging hypernetwork-based neural radiance fields to predict lighting outside the view of the input photo. Through extensive evaluations on common benchmark datasets, we demonstrate superior inverse rendering quality of our method compared to state-of-the-art baselines, enabling various applications such as complex object insertion and material editing with high fidelity. Code and data will be made available at https://jingsenzhu.github.io/invrend Jingsen Zhu, Fujun Luan, Yuchi Huo, Zihao Lin 0007, Dianbing Xi, Rui Wang 0004, Hujun Bao, Jiaxiang Zheng, Rui Tang 0015 |
SIGGRAPH Asia | 6 |