VLDB 2026 Research / reviewers in the wild / expert
Xudong Xu
dblp:210/2741
· DBLP profile ↗
22ranked-venue papers
4as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 7 since 2021Systems, architecture and hardware · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Neural Network Approach to Sonar Target Localization and Trajectory Tracking
Bingru Li, Ouming Ye, Xudong Xu |
J. Supercomput. | 4 |
| 2025 | MeshCoder: LLM-Powered Structured Mesh Code Generation from Point CloudsabstractReconstructing 3D objects into editable programs is pivotal for applications like reverse engineering and shape editing. However, existing methods often rely on limited domain-specific languages (DSLs) and small-scale datasets, restricting their ability to model complex geometries and structures. To address these challenges, we introduce MeshLLM, a novel framework that reconstructs complex 3D objects from point clouds into editable Blender Python scripts. We develop a comprehensive set of expressive Blender Python APIs capable of synthesizing intricate geometries. Leveraging these APIs, we construct a large-scale paired object-code dataset, where the code for each object is decomposed into distinct semantic parts. Subsequently, we train a multimodal large language model (LLM) that translates 3D point cloud into executable Blender Python scripts. Our approach not only achieves superior performance in shape-to-code reconstruction tasks but also facilitates intuitive geometric and topological editing through convenient code modifications. Furthermore, our code-based representation enhances the reasoning capabilities of LLMs in 3D shape understanding tasks. Together, these contributions establish MeshLLM as a powerful and flexible solution for programmatic 3D shape reconstruction and understanding. Bingquan Dai, Li Ray Luo, Qihong Tang, Xinyu Lian, Minghan Qin, Xudong Xu, Bo Dai 0002, Haoqian Wang, Zhaoyang Lyu, Jiangmiao Pang |
NeurIPS | 8 |
| 2025 | MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial ReasoningabstractThe ability of robots to interpret human instructions and execute manipulation tasks necessitates the availability of task-relevant tabletop scenes for training. However, traditional methods for creating these scenes rely on time-consuming manual layout design or purely randomized layouts, which are limited in terms of plausibility or alignment with the tasks. In this paper, we formulate a novel task, namely task-oriented tabletop scene generation, which poses significant challenges due to the substantial gap between high-level task instructions and the tabletop scenes. To support research on such a challenging task, we introduce \textbf{MesaTask-10K}, a large-scale dataset comprising approximately 10,700 synthetic tabletop scenes with \emph{manually crafted layouts} that ensure realistic layouts and intricate inter-object relations. To bridge the gap between tasks and scenes, we propose a \textbf{Spatial Reasoning Chain} that decomposes the generation process into object inference, spatial interrelation reasoning, and scene graph construction for the final 3D layout. We present \textbf{MesaTask}, an LLM-based framework that utilizes this reasoning chain and is further enhanced with DPO algorithms to generate physically plausible tabletop scenes that align well with given task descriptions. Exhaustive experiments demonstrate the superior performance of MesaTask compared to baselines in generating task-conforming tabletop scenes with realistic layouts. Jinkun Hao, Naifu Liang, Xudong Xu, Weipeng Zhong, Ran Yi 0002, Yichen Jin, Zhaoyang Lyu, Feng Zheng 0001, Lizhuang Ma, Jiangmiao Pang |
NeurIPS | 4 |
| 2025 | InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic LayoutsabstractThe advancement of Embodied AI heavily relies on large-scale, simulatable 3D scene datasets characterized by scene diversity and realistic layouts.However, existing datasets typically suffer from limitations in data scale or diversity, sanitized layouts lacking small items, and severe object collisions.To address these shortcomings, we introduce \textbf{InternScenes}, a novel large-scale simulatable indoor scene dataset comprising approximately 40,000 diverse scenes by integrating three disparate scene sources, \ie, real-world scans, procedurally generated scenes, and designer-created scenes, including 1.96M 3D objects and covering 15 common scene types and 288 object classes.We particularly preserve massive small items in the scenes, resulting in realistic and complex layouts with an average of 41.5 objects per region.Our comprehensive data processing pipeline ensures simulatability by creating real-to-sim replicas for real-world scans, enhances interactivity by incorporating interactive objects into these scenes, and resolves object collisions by physical simulations.We demonstrate the value of InternScenes with two benchmark applications: scene layout generation and point-goal navigation. Both show the new challenges posed by the complex and realistic layouts. More importantly, InternScenes paves the way for scaling up the model training for both tasks, making the generation and navigation in such complex scenes possible. We commit to open-sourcing the data and benchmarks to benefit the whole community. Weipeng Zhong, Peizhou Cao, Yichen Jin, Li Ray Luo, Wenzhe Cai, Jingli Lin, Zhaoyang Lyu, Xudong Xu, Bo Dai 0002, Jiangmiao Pang |
NeurIPS | 10 |
| 2025 | Splp-yolo: an all-weather real-time detector for airport runway foreign object debris
Binru Li, Ouming Ye, Xudong Xu |
J. Supercomput. | 5 |
| 2025 | Correction: Splp-yolo: an all-weather real-time detector for airport runway foreign object debris
Bingru Li, Ouming Ye, Xudong Xu |
J. Supercomput. | 5 |
| 2025 | AnySplat: Feed-forward 3D Gaussian Splatting from Unconstrained ViewsabstractWe introduce AnySplat, a feed-forward network for novel-view synthesis from uncalibrated image collections. In contrast to traditional neural-rendering pipelines that demand known camera poses and per-scene optimization, or recent feed-forward methods that buckle under the computational weight of dense views—our model predicts everything in one shot. A single forward pass yields a set of 3D Gaussian primitives encoding both scene geometry and appearance, and the corresponding camera intrinsics and extrinsics for each input image. This unified design scales effortlessly to casually captured, multi-view datasets without any pose annotations. In extensive zero-shot evaluations, AnySplat matches the quality of pose-aware baselines in both sparse- and dense-view scenarios while surpassing existing pose-free approaches. Moreover, it greatly reduces rendering latency compared to optimization-based neural fields, bringing real-time novel-view synthesis within reach for unconstrained capture settings. Project page: https://city-super.github.io/anysplat/. Lihan Jiang, Yucheng Mao, Linning Xu, Tao Lu 0005, Kerui Ren, Yichen Jin, Xudong Xu, Mulin Yu, Jiangmiao Pang, Feng Zhao 0004, Dahua Lin, Bo Dai 0002 |
ACM Trans. Graph. | 7 |
| 2024 | DiffMorpher: Unleashing the Capability of Diffusion Models for Image MorphingabstractDiffusion models have achieved remarkable image generation quality surpassing previous generative models. However, a notable limitation of diffusion models, in comparison to GANs, is their difficulty in smoothly interpolating between two image samples, due to their highly unstructured latent space. Such a smooth interpolation is intriguing as it naturally serves as a solution for the image mor-phing task with many applications. In this work, we address this limitation via DiffMorpher, an approach that enables smooth and natural image interpolation by harnessing the prior knowledge of a pretrained diffusion model. Our key idea is to capture the semantics of the two images by fitting two LoRAs to them respectively, and interpolate between both the LoRA parameters and the latent noises to ensure a smooth semantic transition, where correspon-dence automatically emerges without the need for annotation. In addition, we propose an attention interpolation and injection technique, an adaptive normalization adjustment method, and a new sampling schedule to further enhance the smoothness between consecutive images. Extensive experiments demonstrate that DiffMorpher achieves starkly better image morphing effects than previous methods across a variety of object categories, bridging a critical functional gap that distinguished diffusion models from GANs. Kaiwen Zhang 0015, Yifan Zhou 0001, Xudong Xu, Bo Dai 0002, Xingang Pan |
CVPR | 3 |
| 2024 | TELA: Text to Layer-Wise 3D Clothed Human Generation
Junting Dong, Zehuan Huang, Xudong Xu, Jingbo Wang 0003, Sida Peng, Bo Dai 0002 |
ECCV (25) | 4 |
| 2024 | RoomTex: Texturing Compositional Indoor Scenes via Iterative Inpainting
Qi Wang 0105, Ruijie Lu, Xudong Xu, Jingbo Wang 0003, Michael Yu Wang, Bo Dai 0002, Dan Xu 0002 |
ECCV (68) | 3 |
| 2024 | Boosting 3D object generation through PBR materialsabstractBoosting Normal BoostingFig. 1.Overview.Given a single image, the existing image-to-3D generative models always synthesize 3D meshes with flawed geometry and RGB textures only.Our method not only boosts existing approaches with PBR materials, empowering relighting under various lighting conditions, but also boosts the object's normal maps, capturing more intricate details and better aligning with the given image.Notably, we fine-tune Stable Diffusion to estimate the albedo map from the single-view RGB image and lift it to multi-view albedo maps for a complete albedo UV. Xudong Xu, Bo Dai 0002 |
SIGGRAPH Asia | 2 |
| 2024 | HyperStyle3D: Text-Guided 3D Portrait Stylization via HypernetworksabstractPortrait stylization is a long-standing task enabling extensive applications. Although 2D-based methods have made great progress in recent years, real-world applications such as metaverse and games often demand 3D content. On the other hand, the requirement of 3D data, which is costly to acquire, significantly impedes the development of 3D portrait stylization methods. In this paper, inspired by the success of 3D-aware GANs that bridge 2D and 3D domains with 3D fields as the intermediate representation for rendering 2D images, we propose a novel method, dubbed HyperStyle3D, based on 3D-aware GANs for 3D portrait stylization. At the core of our method is a hyper-network learned to manipulate the parameters of the generator in a single forward pass. It not only offers a strong capacity to handle multiple styles with a single model, but also enables flexible fine-grained stylization that affects only texture, shape, or local part of the portrait. While the use of 3D-aware GANs bypasses the requirement of 3D data, we further alleviate the necessity of style images with the CLIP model being the style guidance. We conduct an extensive set of experiments across the style, attribute, and shape, and meanwhile, measure the 3D consistency. These experiments demonstrate the superior capability of our HyperStyle3D model in rendering 3D-consistent images in diverse styles, deforming the face shape, and editing various attributes. Zhuo Chen 0060, Xudong Xu, Yichao Yan, Wenhan Zhu, Wayne Wu, Bo Dai 0002, Xiaokang Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Voxurf: Voxel-based Efficient and Accurate Neural Surface Reconstruction
Jiaqi Wang 0003, Xingang Pan, Xudong Xu, Christian Theobalt, Ziwei Liu 0002, Dahua Lin |
ICLR | 4 |
| 2023 | Sounding Bodies: Modeling 3D Spatial Sound of Humans Using Body Pose and AudioabstractWhile 3D human body modeling has received much attention in computer vision, modeling the acoustic equivalent, i.e. modeling 3D spatial audio produced by body motion and speech, has fallen short in the community. To close this gap, we present a model that can generate accurate 3D spatial audio for full human bodies. The system consumes, as input, audio signals from headset microphones and body pose, and produces, as output, a 3D sound field surrounding the transmitter's body, from which spatial audio can be rendered at any arbitrary position in the 3D space. We collect a first-of-its-kind multimodal dataset of human bodies, recorded with multiple cameras and a spherical array of 345 microphones. In an empirical evaluation, we demonstrate that our model can produce accurate body-induced sound fields when trained with a suitable loss. Dataset and code are available online. Xudong Xu, Dejan Markovic, Jake Sandakly, Todd Keebler, Steven Krenn, Alexander Richard |
NeurIPS | 1 |
| 2022 | A Conditional Point Diffusion-Refinement Paradigm for 3D Point Cloud Completion
Zhaoyang Lyu, Zhifeng Kong, Xudong Xu, Liang Pan, Dahua Lin |
ICLR | 3 |
| 2021 | Visually Informed Binaural Audio Generation without Binaural AudiosabstractStereophonic audio, especially binaural audio, plays an essential role in immersive viewing environments. Recent research has explored generating visually guided stereophonic audios supervised by multi-channel audio collections. However, due to the requirement of professional recording devices, existing datasets are limited in scale and variety, which impedes the generalization of supervised methods in real-world scenarios. In this work, we propose PseudoBinaural, an effective pipeline that is free of binaural recordings. The key insight is to carefully build pseudo visual-stereo pairs with mono data for training. Specifically, we leverage spherical harmonic decomposition and head-related impulse response (HRIR) to identify the relationship between spatial locations and received binaural audios. Then in the visual modality, corresponding visual cues of the mono data are manually placed at sound source positions to form the pairs. Compared to fully-supervised paradigms, our binaural-recording-free pipeline shows great stability in cross-dataset evaluation and achieves comparable performance under subjective preference. Moreover, combined with binaural recordings, our method is able to further boost the performance of binaural audio generation under supervised settings1. Xudong Xu, Hang Zhou 0009, Ziwei Liu 0002, Bo Dai 0002, Xiaogang Wang 0001, Dahua Lin |
CVPR | 1 |
| 2021 | A Shading-Guided Generative Implicit Model for Shape-Accurate 3D-Aware Image SynthesisabstractThe advancement of generative radiance fields has pushed the boundary of 3D-aware image synthesis. Motivated by the observation that a 3D object should look realistic from multiple viewpoints, these methods introduce a multi-view constraint as regularization to learn valid 3D radiance fields from 2D images. Despite the progress, they often fall short of capturing accurate 3D shapes due to the shape-color ambiguity, limiting their applicability in downstream tasks. In this work, we address this ambiguity by proposing a novel shading-guided generative implicit model that is able to learn a starkly improved shape representation. Our key insight is that an accurate 3D shape should also yield a realistic rendering under different lighting conditions. This multi-lighting constraint is realized by modeling illumination explicitly and performing shading with various lighting conditions. Gradients are derived by feeding the synthesized images to a discriminator. To compensate for the additional computational burden of calculating surface normals, we further devise an efficient volume rendering strategy via surface tracking, reducing the training and inference time by 24% and 48%, respectively. Our experiments on multiple datasets show that the proposed approach achieves photorealistic 3D-aware image synthesis while capturing accurate underlying 3D shapes. We demonstrate improved performance of our approach on 3D shape reconstruction against existing methods, and show its applicability on image relighting. Our code is available at https://github.com/XingangPan/ShadeGAN. Xingang Pan, Xudong Xu, Chen Change Loy, Christian Theobalt, Bo Dai 0002 |
NeurIPS | 2 |
| 2021 | Generative Occupancy Fields for 3D Surface-Aware Image SynthesisabstractThe advent of generative radiance fields has significantly promoted the development of 3D-aware image synthesis. The cumulative rendering process in radiance fields makes training these generative models much easier since gradients are distributed over the entire volume, but leads to diffused object surfaces. In the meantime, compared to radiance fields occupancy representations could inherently ensure deterministic surfaces. However, if we directly apply occupancy representations to generative models, during training they will only receive sparse gradients located on object surfaces and eventually suffer from the convergence problem. In this paper, we propose Generative Occupancy Fields (GOF), a novel model based on generative radiance fields that can learn compact object surfaces without impeding its training convergence. The key insight of GOF is a dedicated transition from the cumulative rendering in radiance fields to rendering with only the surface points as the learned surface gets more and more accurate. In this way, GOF combines the merits of two representations in a unified framework. In practice, the training-time transition of start from radiance fields and march to occupancy representations is achieved in GOF by gradually shrinking the sampling region in its rendering process from the entire volume to a minimal neighboring region around the surface. Through comprehensive experiments on multiple datasets, we demonstrate that GOF can synthesize high-quality images with 3D consistency and simultaneously learn compact and smooth object surfaces. Our code is available at https://github.com/SheldonTsui/GOF_NeurIPS2021. Xudong Xu, Xingang Pan, Dahua Lin, Bo Dai 0002 |
NeurIPS | 1 |
| 2020 | Sep-Stereo: Visually Guided Stereophonic Audio Generation by Associating Source Separation
Hang Zhou 0009, Xudong Xu, Dahua Lin, Xiaogang Wang 0001, Ziwei Liu 0002 |
ECCV (12) | 2 |
| 2019 | Recursive Visual Sound Separation Using Minus-Plus NetabstractSounds provide rich semantics, complementary to visual data, for many tasks. However, in practice, sounds from multiple sources are often mixed together. In this paper we propose a novel framework, referred to as MinusPlus Network (MP-Net), for the task of visual sound separation. MP-Net separates sounds recursively in the order of average energy, removing the separated sound from the mixture at the end of each prediction, until the mixture becomes empty or contains only noise. In this way, MP-Net could be applied to sound mixtures with arbitrary numbers and types of sounds. Moreover, while MP-Net keeps removing sounds with large energy from the mixture, sounds with small energy could emerge and become clearer, so that the separation is more accurate. Compared to previous methods, MP-Net obtains state-of-the-art results on two large scale datasets, across mixtures with different types and numbers of sounds. Xudong Xu, Bo Dai 0002, Dahua Lin |
ICCV | 1 |
| 2019 | Vision-Infused Deep Audio InpaintingabstractMulti-modality perception is essential to develop interactive intelligence. In this work, we consider a new task of visual information-infused audio inpainting, i.e. synthesizing missing audio segments that correspond to their accompanying videos. We identify two key aspects for a successful inpainter: (1) It is desirable to operate on spectrograms instead of raw audios. Recent advances in deep semantic image inpainting could be leveraged to go beyond the limitations of traditional audio inpainting. (2) To synthesize visually indicated audio, a visual-audio joint feature space needs to be learned with synchronization of audio and video. To facilitate a large-scale study, we collect a new multi-modality instrument-playing dataset called MUSIC-Extra-Solo (MUSICES) by enriching MUSIC dataset [51]. Extensive experiments demonstrate that our framework is capable of inpainting realistic and varying audio segments with or without visual contexts. More importantly, our synthesized audio segments are coherent with their video counterparts, showing the effectiveness of our proposed Vision-Infused Audio Inpainter (VIAI). Code, models, dataset and video results are available at https://github.com/Hangz-nju-cuhk/ Vision-Infused-Audio-Inpainter-VIAI. Hang Zhou 0009, Ziwei Liu 0002, Xudong Xu, Ping Luo 0002, Xiaogang Wang 0001 |
ICCV | 3 |
| 2017 | A novel DDPG method with prioritized experience replayabstractRecently, a state-of-the-art algorithm, called deep deterministic policy gradient (DDPG), has achieved good performance in many continuous control tasks in the MuJoCo simulator. To further improve the efficiency of the experience replay mechanism in DDPG and thus speeding up the training process, in this paper, a prioritized experience replay method is proposed for the DDPG algorithm, where prioritized sampling is adopted instead of uniform sampling. The proposed DDPG with prioritized experience replay is tested with an inverted pendulum task via OpenAI Gym. The experimental results show that DDPG with prioritized experience replay can reduce the training time and improve the stability of the training process, and is less sensitive to the changes of some hyperparameters such as the size of replay buffer, minibatch and the updating rate of the target network. Yuenan Hou, Lifeng Liu, Xudong Xu |
SMC | 4 |