VLDB 2026 Research / reviewers in the wild / expert
Qingshan Xu 0001
dblp:32/9530-1
· DBLP profile ↗
33ranked-venue papers
10as first author
29since 2021 · last 2026
0000-0003-0405-3962ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 10 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 6 first-author · 19 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Personalize Your Gaussian: Consistent 3D Scene Personalization from a Single ImageabstractPersonalizing 3D scenes from a single reference image enables intuitive user-guided editing, which requires achieving both multi-view consistency across perspectives and referential consistency with the input image. However, these goals are particularly challenging due to the viewpoint bias caused by the limited perspective provided in a single image. Lacking the mechanisms to effectively expand reference information beyond the original view, existing methods of image-conditioned 3DGS personalization often suffer from this viewpoint bias and struggle to produce consistent results. Therefore, in this paper, we present Consistent Personalization for 3D Gaussian Splatting (CP-GS), a framework that progressively propagates the single-view reference appearance to novel perspectives. In particular, CP-GS integrates pre-trained image-to-3D generation and iterative LoRA fine-tuning to extract and extend the reference appearance, and finally produces faithful multi-view guidance images and the personalized 3DGS outputs through a view-consistent generation process guided by geometric cues. Extensive experiments on real-world scenes show that our CP-GS effectively mitigates the viewpoint bias, achieving high-quality image-conditioned 3DGS personalization that significantly outperforms existing methods. Xuanyu Yi, Qingshan Xu 0001, Yuan Zhou 0016, Long Chen 0016, Hanwang Zhang |
AAAI | 3 |
| 2026 | Pushing Rendering Boundaries: Hard Gaussian Splattingabstract3D Gaussian Splatting (3DGS) has demonstrated impressive Novel View Synthesis (NVS) results in a real-time rendering manner. During training, it relies heavily on the average magnitude of view-space positional gradients to grow Gaussians to reduce rendering loss. However, this average operation smooths the positional gradients from different viewpoints and rendering errors from different pixels, hindering the growth and optimization of many defective Gaussians. This leads to strong spurious artifacts in some areas. To address this problem, we propose Hard Gaussian Splatting, dubbed HGS, which considers multi-view significant positional gradients and rendering errors to grow hard Gaussians that fill the gaps of classical Gaussian Splatting on 3D scenes, thus achieving superior NVS results. In detail, we present positional gradient driven HGS, which leverages multi-view significant positional gradients to uncover hard Gaussians. Moreover, we propose rendering error guided HGS, which identifies noticeable pixel rendering errors and potentially over-large Gaussians to jointly mine hard Gaussians. By growing and optimizing these hard Gaussians, our method helps to resolve blurring and needle-like artifacts. Experiments on various datasets demonstrate that our method achieves state-of-the-art rendering quality while maintaining real-time efficiency, yielding LPIPS improvements of 5.1%, 19.7% and 6.3% on Mip-NeRF360, Tanks&Temples and Deep Blending, respectively. Qingshan Xu 0001, Jiequan Cui, Xuanyu Yi, Yuan Zhou 0016, Yew-Soon Ong, Hanwang Zhang |
AAAI | 1 |
| 2026 | NeuSpring: Neural Spring Fields for Reconstruction and Simulation of Deformable Objects from VideosabstractIn this paper, we aim to create physical digital twins of deformable objects under interaction. Existing methods focus more on the physical learning of current state modeling, but generalize worse to future prediction. This is because existing methods ignore the intrinsic physical properties of deformable objects, resulting in the limited physical learning in the current state modeling. To address this, we present NeuSpring, a neural spring field for the reconstruction and simulation of deformable objects from videos. Built upon spring-mass models for realistic physical simulation, our method consists of two major innovations: 1) a piecewise topology solution that efficiently models multi-region spring connection topologies using zero-order optimization, which considers the material heterogeneity of real-world objects. 2) a neural spring field that represents spring physical properties across different frames using a canonical coordinate-based neural network, which effectively leverages the spatial associativity of springs for physical learning. Experiments on real-world datasets demonstrate that our NeuSping achieves superior reconstruction and simulation performance for current state modeling and future prediction, with Chamfer distance improved by 20% and 25%, respectively. Qingshan Xu 0001, Jiao Liu 0006, Shangshu Yu, Yuan Zhou 0016, Junbao Zhou, Jiequan Cui, Yew-Soon Ong, Hanwang Zhang |
AAAI | 1 |
| 2026 | DragNeXt: Rethinking Drag-Based Image EditingabstractDrag-Based Image Editing (DBIE), which allows users to manipulate images by directly dragging objects within them, has recently attracted much attention from the community. However, it faces two key challenges: (i) point-based drag is often highly ambiguous and difficult to align with user intentions; (ii) current DBIE methods primarily rely on alternating between motion supervision and point tracking, which is not only cumbersome but also fails to produce high-quality results. These limitations motivate us to explore DBIE from a new perspective---unifying it as a Latent Region Optimization (LRO) problem that aims to use region-level geometric transformations to optimize latent code to realize drag manipulation. Thus, by specifying the areas and types of geometric transformations, we can effectively address the ambiguity issue. We also propose a simple yet effective editing framework, dubbed DragNeXt. It solves LRO through Progressive Backward Self-Intervention (PBSI), simplifying the overall procedure of the alternating workflow while further enhancing quality by fully leveraging region-level structure information and progressive guidance from intermediate drag states. We validate DragNeXt on our NextBench, and extensive experiments demonstrate that our proposed method can significantly outperform existing approaches. Yuan Zhou 0016, Junbao Zhou, Qingshan Xu 0001, Kesen Zhao, Hao Fei 0001, Richang Hong, Hanwang Zhang |
AAAI | 3 |
| 2026 | A Hierarchical Prior Mining Approach for Non-Local Multi-View StereoabstractAs a fundamental problem in computer vision, multi-view stereo (MVS) aims at recovering the 3D geometry of the target from a set of 2D images. However, the reconstructed quality is significantly impacted by the presence of low-textured areas. In this paper, we propose a Hierarchical Prior Mining (HPM) framework for non-local multi-view stereo. Different from most existing works dedicated to focusing on local information and only using a single prior, HPM captures non-local structural cues and leverages multi-source priors for geometry recovery. Based on the framework, we first propose HPM-MVS, which obtains precise initial hypotheses through non-local operations, simultaneously constructing a better planar prior model in an HPM framework to further facilitate hypothesis generation. In addition, we futher propose HPM-MVS++, which excavates the structured region information of images and spatial geometric relationships of hypotheses as prior knowledge. Then, it incorporates them into probabilistic graphical models, ultimately deducing two novel multi-view matching costs. This significantly enhances the robustness to challenging situations and improves the completeness of the reconstruction. Experimental results on the ETH3D and Tanks & Temples have verified the superior performance and strong generalization capability of our approach. Jiaqi Yang 0002, Yanan He, Chunlin Ren, Qingshan Xu 0001, Siwen Quan, Xiyu Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Reconstructing Sparse-View Indoor Scenes in View Space With Global Monocular Prior AlignmentabstractAlthough 3D Gaussian Splatting (3DGS) has greatly advanced the development of novel view synthesis (NVS) and surface reconstruction tasks, it still faces serious challenges in indoor scenes with only sparse views due to the lack of initialization point cloud. Existing GS-like methods are geometrically initialized using point clouds from Structure from Motion (SfM), and their performance is heavily limited by the point cloud quality. However, in indoor scenes containing large areas of textureless or weakly textured regions, the SfM algorithm struggles to generate point clouds in these areas, especially when only a limited number of input views are available. This situation can cause the GS optimization process to overfit the photometric error, resulting in severe geometric degradation in regions lacking proper initialization. To address the above problem, we propose a novel sparse-view 3DGS method using global scale-aligned monocular depth to initialize Gaussian primitives and optimize them in view space, named VGA-GS. First, we design a global scale-aligned monocular depth initialization strategy, which provides geometric priors in textureless or weakly textured regions while avoiding the local geometric distortions in vanilla alignments. Second, we propose a novel Gaussian primitive representation based on view space and a refinement strategy based on pixel errors, which can constrain each primitive to always be within the field of view where it is observed to be fully optimized. Finally, we extract the surface prior from monocular depth to regularize the rendering depth and normal, and further propagate the constraints of the training view to the neighboring pseudo-views via image warping. Our method achieves state-of-the-art reconstruction performance in both textureless regions and local details for indoor scenes with sparse views. Moreover, it is highly scalable, supporting integration with various 3DGS rasterizers and depth estimation methods. We release the code at https://github.com/XT5un/VGA-GS. Xiaotian Sun 0005, Qingshan Xu 0001, Mengyin Liu, Cheng Wang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | STGC-NeRF: Spatial-Temporal Geometric Consistency for LiDAR Neural Radiance Fields in Dynamic ScenesabstractWhile Neural Radiance Fields (NeRFs) have advanced the frontiers of novel view synthesis (NVS) using LiDAR data, they still struggle in dynamic scenes. Due to the low frequency and sparsity characteristics of LiDAR point clouds, it is challenging to spontaneously learn a dynamic and consistent scene representation from posed scans. In this paper, we propose STGC-NeRF, a novel LiDAR NeRF method that combines spatial-temporal geometry consistency to enhance the reconstruction of dynamic scenes. First, we propose a temporal geometry consistency regularization to enhance the regression of time-varying scene geometries from low-frequency LiDAR sequences. By estimating the pointwise correspondences between synthetic (or real) and real frames at different times, we convert them into various forms of temporal supervision. This alleviates the inconsistency caused by moving objects in dynamic scenes. Second, to improve the reconstruction of sparse LiDAR data, we propose spatial geometric consistency constraints. By computing multiple neighborhood feature descriptors incorporating geometric and contextual information, we capture structural geometry information from sparse LiDAR data. This helps encourage consistent direction, smoothness, and detail of the local surface. Extensive experiments on the KITTI-360 and nuScenes datasets demonstrate that STGC-NeRF outperforms state-of-the-art methods in both geometry and intensity accuracy for dynamic LiDAR scene reconstruction. Shangshu Yu, Xiaotian Sun 0005, Wen Li 0005, Qingshan Xu 0001, Zhimin Yuan, Rui She 0001, Cheng Wang 0003 |
AAAI | 4 |
| 2025 | Boosting Adversarial Transferability through Augmentation in Hypothesis SpaceabstractAdversarial examples can mislead deep neural networks with subtle perturbations, causing them to make incorrect predictions. Notably, adversarial examples crafted for one model can also deceive other models, a phenomenon known as the transferability of adversarial examples. To improve transferability, existing studies have designed increasingly complex mechanisms, but the improvements achieved remain relatively limited and are often difficult to adapt to other modalities, further restricting the scalability of these methods. In this work, we observe a mirroring relationship between model generalization and adversarial example transferability. Motivated by this observation, we propose an augmentation-based attack, called OPS (OperatorPerturbation-based Stochastic optimization), which constructs a stochastic optimization problem by input transformation operators and random perturbations, and solves this problem to generate adversarial examples with better transferability. Extensive experiments on both images and 3D point clouds demonstrate that OPS significantly outperforms existing state-of-the-art methods in terms of both performance and cost, showcasing the universality and superiority of our approach. The code is available at https://github.com/the-full/OPS. Weiquan Liu, Qingshan Xu 0001, Shijun Zheng, Shujun Huang, Chenglu Wen, Cheng Wang 0003 |
CVPR | 3 |
| 2025 | CARE Transformer: Mobile-Friendly Linear Visual Transformer via Decoupled Dual InteractionabstractRecently, large efforts have been made to design efficient linear-complexity visual Transformers. However, current linear attention models are generally unsuitable to be deployed in resource-constrained mobile devices, due to suffering from either few efficiency gains or significant accuracy drops. In this paper, we propose a new deCoupled duAl-interactive lineaR attEntion (CARE) mechanism, revealing that features’ decoupling and interaction can fully unleash the power of linear attention. We first propose an asymmetrical feature decoupling strategy that asymmetrically decouples the learning process for local inductive bias and long-range dependencies, thereby preserving sufficient local and global information while effectively enhancing the efficiency of models. Then, a dynamic memory unit is employed to maintain critical information along the network pipeline. Moreover, we design a dual interaction module to effectively facilitate interaction between local inductive bias and long-range information as well as among features at different layers. By adopting a decoupled learning way and fully exploiting complementarity across features, our method can achieve both high efficiency and accuracy. Extensive experiments on ImageNet-1K, COCO, and ADE20K datasets demonstrate the effectiveness of our approach, e.g., achieving 78.4/82.1% top-1 accuracy on ImagegNet-1K at the cost of only 0.7/1.9 GMACs. Codes will be released on github. Yuan Zhou 0016, Qingshan Xu 0001, Jiequan Cui, Junbao Zhou, Richang Hong, Hanwang Zhang |
CVPR | 2 |
| 2025 | Nautilus: Locality-Aware Autoencoder for Scalable Mesh GenerationabstractTriangle meshes are fundamental to 3D applications, enabling efficient modification and rasterization while maintaining compatibility with standard rendering pipelines. However, current automatic mesh generation methods typically rely on intermediate representations that lack the continuous surface quality inherent to meshes. Converting these representations into meshes produces dense, suboptimal outputs. Although recent autoregressive approaches demonstrate promise in directly modeling mesh vertices and faces, they are constrained by the limitation in face count, scalability, and structural fidelity. To address these challenges, we propose Nautilus, a locality-aware autoencoder for artist-like mesh generation that leverages the local properties of manifold meshes to achieve structural fidelity and efficient representation. Our approach introduces a novel tokenization algorithm that preserves face proximity relationships and compresses sequence length through locally shared vertices and edges, enabling the generation of meshes with an unprecedented scale of up to 5,000 faces. Furthermore, we develop a Dual-stream Point Conditioner that provides multi-scale geometric guidance, ensuring global consistency and local structural fidelity by capturing fine-grained geometric features. Extensive experiments demonstrate that Nautilus significantly outperforms state-of-the-art methods in both fidelity and scalability. The project page is at https://nautilusmeshgen.github.io. Xuanyu Yi, Haohan Weng, Qingshan Xu 0001, Xiaokang Wei, Xianghui Yang, Chunchao Guo, Long Chen 0016, Hanwang Zhang |
ICCV | 4 |
| 2025 | On Path to Multimodal Generalist: General-Level and General-BenchabstractThe Multimodal Large Language Model (MLLM) is currently experiencing rapid growth, driven by the advanced capabilities of language-based LLMs. Unlike their specialist predecessors, existing MLLMs are evolving towards a Multimodal Generalist paradigm. Initially limited to understanding multiple modalities, these models have advanced to not only comprehend but also generate across modalities. Their capabilities have expanded from coarse-grained to fine-grained multimodal understanding and from supporting singular modalities to accommodating a wide array of or even arbitrary modalities. To assess the capabilities of various MLLMs, a diverse array of benchmark test sets has been proposed. This leads to a critical question: Can we simply assume that higher performance across tasks indicates a stronger MLLM capability, bringing us closer to human-level AI? We argue that the answer is not as straightforward as it seems. In this project, we introduce an evaluation framework to delineate the capabilities and behaviors of current multimodal generalists. This framework, named General-Level, establishes 5-scale levels of MLLM performance and generality, offering a methodology to compare MLLMs and gauge the progress of existing systems towards more robust multimodal generalists and, ultimately, towards AGI (Artificial General Intelligence). Central to our framework is the use of Synergy as the evaluative criterion, categorizing capabilities based on whether MLLMs preserve synergy across comprehension and generation, as well as across multimodal interactions. To evaluate the comprehensive abilities of various generalists, we present a massive multimodal benchmark, General-Bench, which encompasses a broader spectrum of skills, modalities, formats, and capabilities, including over 700 tasks and 325,800 instances. The evaluation results that involve over 100 existing state-of-the-art MLLMs uncover the capability rankings of generalists, highlighting the challenges in reaching genuine AI. We expect this project to pave the way for future research on next-generation multimodal foundation models, providing a robust infrastructure to accelerate the realization of AGI. Project Page: https://generalist.top/, Leaderboard: https://generalist.top/leaderboard/, Benchmark: https://huggingface.co/General-Level/. Hao Fei 0001, Yuan Zhou 0016, Juncheng Li 0006, Xiangtai Li, Qingshan Xu 0001, Bobo Li 0001, Shengqiong Wu, Yaoting Wang, Junbao Zhou, Jiahao Meng, Liangtao Shi, Minghe Gao, Daoan Zhang, Zhiqi Ge, Siliang Tang, Kaihang Pan, Yaobo Ye, Haobo Yuan, Tao Zhang 0042, Weiming Wu, Tianjie Ju, Zixiang Meng, Shilin Xu 0001, Liyu Jia, Meng Luo 0010, Jiebo Luo 0001, Tat-Seng Chua, Shuicheng Yan, Hanwang Zhang |
ICML | 5 |
| 2025 | Looks Great, Functions Better: Physics Compliance Text-to-3D Shape GenerationabstractText-to-3D shape generation has shown great promise in generating novel 3D content based on given text prompts. However, existing generative methods mainly consider geometric or visual plausibility while ignoring functionality for the generated 3D shapes. This greatly hinders the practicality of generated 3D shapes in real-world applications. Towards physical AI, we propose Fun3D, a physics-compliant functional text-to-3D shape generation method. By analyzing the solid mechanics of generated 3D shapes, we reveal that the 3D shapes generated by existing text-to-3D generation methods are impractical for real-world applications, as the generated 3D shapes do not comply with the physical laws. To this end, we leverage 3D diffusion models to provide 3D shape priors and design a data-driven differentiable physics layer to optimize 3D shape priors with solid mechanics. This allows us to optimize geometry efficiently and learn physical information about 3D shapes at the same time. Experimental results demonstrate that our method can consider both geometric plausibility and functional requirement, further bridging 3D virtual modeling and physical worlds to advance physical AI. Qingshan Xu 0001, Jiao Liu 0006, Melvin Wong, Caishun Chen, Yew-Soon Ong |
IJCNN | 1 |
| 2025 | Lightweight and Accurate Multi-View Stereo With Confidence-Aware Diffusion ModelabstractTo reconstruct the 3D geometry from calibrated images, learning-based multi-view stereo (MVS) methods typically perform multi-view depth estimation and then fuse depth maps into a mesh or point cloud. To improve the computational efficiency, many methods initialize a coarse depth map and then gradually refine it in higher resolutions. Recently, diffusion models achieve great success in generation tasks. Starting from a random noise, diffusion models gradually recover the sample with an iterative denoising process. In this paper, we propose a novel MVS framework, which introduces diffusion models in MVS. Specifically, we formulate depth refinement as a conditional diffusion process. Considering the discriminative characteristic of depth estimation, we design a condition encoder to guide the diffusion process. To improve efficiency, we propose a novel diffusion network combining lightweight 2D U-Net and convolutional GRU. Moreover, we propose a novel confidence-based sampling strategy to adaptively sample depth hypotheses based on the confidence estimated by diffusion model. Based on our novel MVS framework, we propose two novel MVS methods, DiffMVS and CasDiffMVS. DiffMVS achieves competitive performance with state-of-the-art efficiency in run-time and GPU memory. CasDiffMVS achieves state-of-the-art performance on DTU, Tanks & Temples and ETH3D. Fangjinhua Wang, Qingshan Xu 0001, Yew-Soon Ong, Marc Pollefeys |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | PSDF: Prior-Driven Neural Implicit Surface Learning for Multi-View ReconstructionabstractSurface reconstruction has traditionally relied on the Multi-View Stereo (MVS)-based pipeline, which often suffers from noisy and incomplete geometry. This is due to that although MVS has been proven to be an effective way to recover the geometry of the scenes, especially for locally detailed areas with rich textures, it struggles to deal with areas with low texture and large variations of illumination where the photometric consistency is unreliable. Recently, Neural Implicit Surface Reconstruction (NISR) combines surface rendering and volume rendering techniques and bypasses the MVS as an intermediate step, which has emerged as a promising alternative to overcome the limitations of traditional pipelines. While NISR has shown impressive results on simple scenes, it remains challenging to recover delicate geometry from uncontrolled real-world scenes which is caused by its underconstrained optimization. To this end, the framework PSDF is proposed which resorts to external geometric priors from a pretrained MVS network and internal geometric priors inherent in the NISR model to facilitate high-quality neural implicit surface learning. Specifically, the visibility-aware feature consistency loss and depth prior-assisted sampling based on external geometric priors are introduced. These proposals provide powerfully geometric consistency constraints and aid in locating surface intersection points, thereby significantly improving the accuracy and delicate reconstruction of NISR. Meanwhile, the internal prior-guided importance rendering is presented to enhance the fidelity of the reconstructed surface mesh by mitigating the biased rendering issue in NISR. Extensive experiments on Tanks and Temples datasets show that PSDF achieves state-of-the-art performance on complex uncontrolled scenes. Wanjuan Su, Chen Zhang 0043, Qingshan Xu 0001, Wenbing Tao |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | PG-NeuS: Robust and Efficient Point Guidance for Multi-View Neural Surface ReconstructionabstractRecently, learning multi-view neural surface reconstruction with the supervision of point clouds or depth maps has been a promising way. However, due to weak perception and underutilization of prior information, current methods still struggle with the challenges of limited accuracy and excessive time complexity. In addition, prior data perturbation is also an important yet rarely considered issue, often resulting in distorted geometry. To address these challenges, we propose a novel point-guided method named PG-NeuS, which achieves accurate and efficient reconstruction while robustly coping with point noise. Specifically, the aleatoric uncertainty of the point cloud is modeled to capture the noise distribution, estimating the reliability of each point and enhancing robustness against noise. Moreover, a Neural Projection module is proposed to connect points and images, adding geometric constraints to the implicit surface and achieving more precise point guidance. To better compensate for geometric bias between volume rendering and point modeling, we additionally design a Bias network that leverages the geometric information in high-fidelity points to enhance detail representation. Benefiting from the effective point guidance, the proposed PG-NeuS achieves an 11x speed increase and a 33.3% accuracy improvement compared to NeuS on DTU, even with a lightweight network. Extensive experiments show that our method yields high-quality surfaces with high efficiency, especially for fine-grained details and smooth regions, outperforming the state-of-the-art methods. Moreover, it exhibits strong robustness to noisy data and sparse data. Chen Zhang 0043, Wanjuan Su, Qingshan Xu 0001, Xinyao Liao, Wenbing Tao |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | SPEAL: Skeletal Prior Embedded Attention Learning for Cross-Source Point Cloud RegistrationabstractPoint cloud registration, a fundamental task in 3D computer vision, has remained largely unexplored in cross-source point clouds and unstructured scenes. The primary challenges arise from noise, outliers, and variations in scale and density. However, neglected geometric natures of point clouds restricts the performance of current methods. In this paper, we propose a novel method termed SPEAL to leverage skeletal representations for effective learning of intrinsic topologies of point clouds, facilitating robust capture of geometric intricacy. Specifically, we design the Skeleton Extraction Module to extract skeleton points and skeletal features in an unsupervised manner, which is inherently robust to noise and density variances. Then, we propose the Skeleton-Aware GeoTransformer to encode high-level skeleton-aware features. It explicitly captures the topological natures and inter-point-cloud skeletal correlations with the noise-robust and density-invariant skeletal representations. Next, we introduce the Correspondence Dual-Sampler to facilitate correspondences by augmenting the correspondence set with skeletal correspondences. Furthermore, we construct a challenging novel cross-source point cloud dataset named KITTI CrossSource for benchmarking cross-source point cloud registration methods. Extensive quantitative and qualitative experiments are conducted to demonstrate our approach’s superiority and robustness on both cross-source and same-source datasets. To the best of our knowledge, our approach is the first to facilitate point cloud registration with skeletal geometric priors. Kezheng Xiong, Maoji Zheng, Qingshan Xu 0001, Chenglu Wen, Cheng Wang 0003 |
AAAI | 3 |
| 2024 | Global and Hierarchical Geometry Consistency Priors for Few-Shot NeRFs in Indoor ScenesabstractIt is challenging for Neural Radiance Fields (NeRFs) in the few-shot setting to reconstruct high-quality novel views and depth maps in 360° outward-facing indoor scenes. The captured sparse views for these scenes usually contain large viewpoint variations. This greatly reduces the potential consistency between views, leading NeRFs to degrade a lot in these scenarios. Existing methods usually leverage pre-trained depth prediction models to improve NeRFs. However, these methods cannot guarantee geometry consistency due to the inherent geometry ambiguity in the pretrained models, thus limiting NeRFs' performance. In this work, we present p2 NeRF to capture global and hierarchical geometry consistency priors from pretrained models, thus facilitating few-shot NeRFs in 360° outward-facing indoor scenes. On the one hand, we propose a matching-based geometry warm-up strategy to provide global geometry consistency priors for NeRFs. This effectively avoids the overfitting of early training with sparse inputs. On the other hand, we propose a group depth ranking loss and ray weight mask regularization based on the monocular depth estimation model. This provides hierarchical geometry consistency priors for NeRFs. As a result, our approach can fully leverage the geometry consistency priors from pretrained models and help few-shot NeRFs achieve state-of-the-art performance on two challenging indoor datasets. Our code is released at https://github.com/XT5un/P2NeRF. Xiaotian Sun 0005, Qingshan Xu 0001, Cheng Wang 0003 |
CVPR | 2 |
| 2024 | Diffusion Time-step Curriculum for One Image to 3D GenerationabstractScore distillation sampling (SDS) has been widely adopted to overcome the absence of unseen views in reconstructing 3D objects from a single image. It leverages pretrained 2D diffusion models as teacher to guide the reconstruction of student 3D models. Despite their remarkable success, SDS-based methods often encounter geometric artifacts and texture saturation. We find out the crux is the overlooked indiscriminate treatment of diffusion time-steps during optimization: it unreasonably treats the student-teacher knowledge distillation to be equal at all time-steps and thus entangles coarse-grained and fine-grained modeling. Therefore, we propose the Diffusion Time-step Curriculum one-image-to-3D pipeline (DTC123), which involves both the teacher and student models collaborating with the time-step curriculum in a coarse-to-fine manner. Extensive experiments on NeRF4, RealFusion15, GSO and Level50 benchmark demonstrate that DTC123 can produce multiview consistent, high-quality, and diverse 3D assets. Codes and more generation demos will be released in https://github.com/yxymessi/DTC123. Xuanyu Yi, Zike Wu, Qingshan Xu 0001, Pan Zhou 0002, Joo-Hwee Lim, Hanwang Zhang |
CVPR | 3 |
| 2024 | Few-Shot NeRF by Adaptive Rendering Loss Regularization
Qingshan Xu 0001, Xuanyu Yi, Jianyao Xu, Wenbing Tao, Yew-Soon Ong, Hanwang Zhang |
ECCV (66) | 1 |
| 2024 | Mining and Transferring Feature-Geometry Coherence for Unsupervised Point Cloud RegistrationabstractPoint cloud registration, a fundamental task in 3D vision, has achieved remarkable success with learning-based methods in outdoor environments. Unsupervised outdoor point cloud registration methods have recently emerged to circumvent the need for costly pose annotations. However, they fail to establish reliable optimization objectives for unsupervised training, either relying on overly strong geometric assumptions, or suffering from poor-quality pseudo-labels due to inadequate integration of low-level geometric and high-level contextual information. We have observed that in the feature space, latent new inlier correspondences tend to cluster
around respective positive anchors that summarize features of existing inliers. Motivated by this observation, we propose a novel unsupervised registration method termed INTEGER to incorporate high-level contextual information for reliable pseudo-label mining. Specifically, we propose the Feature-Geometry Coherence Mining module to dynamically adapt the teacher for each mini-batch of data during training and discover reliable pseudo-labels by considering both high-level feature representations and low-level geometric cues. Furthermore, we propose Anchor-Based Contrastive Learning to facilitate contrastive learning with anchors for a robust feature space. Lastly, we introduce a Mixed-Density Student to learn density-invariant features, addressing challenges related to density variation and low overlap in the outdoor scenario. Extensive experiments on KITTI and nuScenes datasets demonstrate that our INTEGER achieves competitive performance in terms of accuracy and generalizability. Kezheng Xiong, Haoen Xiang, Qingshan Xu 0001, Chenglu Wen, Jonathan Jun Li, Cheng Wang 0003 |
NeurIPS | 3 |
| 2024 | MVGamba: Unify 3D Content Generation as State Space Sequence ModelingabstractRecent 3D large reconstruction models (LRMs) can generate high-quality 3D content in sub-seconds by integrating multi-view diffusion models with scalable multi-view reconstructors. Current works further leverage 3D Gaussian Splatting as 3D representation for improved visual quality and rendering efficiency. However, we observe that existing Gaussian reconstruction models often suffer from multi-view inconsistency and blurred textures. We attribute this to the compromise of multi-view information propagation in favor of adopting powerful yet computationally intensive architectures (\eg, Transformers).
To address this issue, we introduce MVGamba, a general and lightweight Gaussian reconstruction model featuring a multi-view Gaussian reconstructor based on the RNN-like State Space Model (SSM). Our Gaussian reconstructor propagates causal context containing multi-view information for cross-view self-refinement while generating a long sequence of Gaussians for fine-detail modeling with linear complexity.
With off-the-shelf multi-view diffusion models integrated, MVGamba unifies 3D generation tasks from a single image, sparse images, or text prompts. Extensive experiments demonstrate that MVGamba outperforms state-of-the-art baselines in all 3D content generation scenarios with approximately only $0.1\times$ of the model size. The codes are available at \url{https://github.com/SkyworkAI/MVGamba}. Xuanyu Yi, Zike Wu, Qiuhong Shen, Qingshan Xu 0001, Pan Zhou 0002, Joo-Hwee Lim, Shuicheng Yan, Xinchao Wang, Hanwang Zhang |
NeurIPS | 4 |
| 2023 | Hierarchical Prior Mining for Non-local Multi-View StereoabstractAs a fundamental problem in computer vision, multi-view stereo (MVS) aims at recovering the 3D geometry of a target from a set of 2D images. Recent advances in MVS have shown that it is important to perceive non-local structured information for recovering geometry in low-textured areas. In this work, we propose a Hierarchical Prior Mining for Non-local Multi-View Stereo (HPM-MVS). The key characteristics are the following techniques that exploit non-local information to assist MVS: 1) A Non-local Extensible Sampling Pattern (NESP), which is able to adaptively change the size of sampled areas without becoming snared in locally optimal solutions. 2) A new approach to leverage non-local reliable points and construct a planar prior model based on K-Nearest Neighbor (KNN), to obtain potential hypotheses for the regions where prior construction is challenging. 3) A Hierarchical Prior Mining (HPM) framework, which is used to mine extensive non-local prior information at different scales to assist 3D model recovery, this strategy can achieve a considerable balance between the reconstruction of details and low-textured areas. Experimental results on the ETH3D and Tanks & Temples have verified the superior performance and strong generalization capability of our method. Our code will be available at https://github.com/CLinvx/HPM-MVS. Chunlin Ren, Qingshan Xu 0001, Shikun Zhang, Jiaqi Yang 0002 |
ICCV | 2 |
| 2023 | Edge-Aware Spatial Propagation Network for Multi-view Depth Estimation
Qingshan Xu 0001, Wanjuan Su, Wenbing Tao |
Neural Process. Lett. | 2 |
| 2023 | Multi-Scale Geometric Consistency Guided and Planar Prior Assisted Multi-View StereoabstractIn this paper, we propose some efficient multi-view stereo methods for accurate and complete depth map estimation. We first present our basic methods with Adaptive Checkerboard sampling and Multi-Hypothesis joint view selection (ACMH & ACMH+). Based on our basic models, we develop two frameworks to deal with the depth estimation of ambiguous regions (especially low-textured areas) from two different perspectives: multi-scale information fusion and planar geometric clue assistance. For the former one, we propose a multi-scale geometric consistency guidance framework (ACMM) to obtain the reliable depth estimates for low-textured areas at coarser scales and guarantee that they can be propagated to finer scales. For the latter one, we propose a planar prior assisted framework (ACMP). We utilize a probabilistic graphical model to contribute a novel multi-view aggregated matching cost. At last, by taking advantage of the above frameworks, we further design a multi-scale geometric consistency guided and planar prior assisted multi-view stereo (ACMMP). This greatly enhances the discrimination of ambiguous regions and helps their depth sensing. Experiments on extensive datasets show our methods achieve state-of-the-art performance, recovering the depth estimation not only in low-textured areas but also in details. Related codes are available at https://github.com/GhiXu. Qingshan Xu 0001, Weihang Kong, Wenbing Tao, Marc Pollefeys |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | LGP-MVS: combined local and global planar priors guidance for indoor multi-view stereo
Weihang Kong, Qingshan Xu 0001, Wanjuan Su, Wenbing Tao |
Vis. Comput. | 2 |
| 2022 | Geo-Neus: Geometry-Consistent Neural Implicit Surfaces Learning for Multi-view ReconstructionabstractRecently, neural implicit surfaces learning by volume rendering has become popular for multi-view reconstruction. However, one key challenge remains: existing approaches lack explicit multi-view geometry constraints, hence usually fail to generate geometry-consistent surface reconstruction. To address this challenge, we propose geometry-consistent neural implicit surfaces learning for multi-view reconstruction. We theoretically analyze that there exists a gap between the volume rendering integral and point-based signed distance function (SDF) modeling. To bridge this gap, we directly locate the zero-level set of SDF networks and explicitly perform multi-view geometry optimization by leveraging the sparse geometry from structure from motion (SFM) and photometric consistency in multi-view stereo. This makes our SDF optimization unbiased and allows the multi-view geometry constraints to focus on the true surface optimization. Extensive experiments show that our proposed method achieves high-quality surface reconstruction in both complex thin structures and large smooth regions, thus outperforming the state-of-the-arts by a large margin. Qiancheng Fu, Qingshan Xu 0001, Yew-Soon Ong, Wenbing Tao |
NeurIPS | 2 |
| 2022 | Sparse prior guided deep multi-view stereo
Yuhang Qi, Wanjuan Su, Qingshan Xu 0001, Wenbing Tao |
Comput. Graph. | 3 |
| 2022 | Learning Inverse Depth Regression for Pixelwise Visibility-Aware Multi-View Stereo Networks
Qingshan Xu 0001, Wanjuan Su, Yuhang Qi, Wenbing Tao, Marc Pollefeys |
Int. J. Comput. Vis. | 1 |
| 2022 | Uncertainty Guided Multi-View Stereo Network for Depth EstimationabstractDeep learning has greatly promoted the development of multi-view stereo in recent years. However, how to measure the reliability of the estimated depth map for practical applications and make reasonable depth hypothesis sampling for the cost volume building in the coarse-to-fine architecture are still unresolved crucial problems. To this end, an Uncertainty Guided multi-view Network (UGNet) is proposed in this paper. In order to enable the network to perceive the uncertainty, an uncertainty-aware loss function is introduced, which not only can infer uncertainty implicitly in an unsupervised manner but also can reduce the bad impact of high uncertainty regions and the erroneous labels in the training set during training. Moreover, an uncertainty-based depth hypothesis sampling strategy is further proposed to adaptively determine the depth search range of each pixel for finer stages, which helps to generate more rational depth intervals compared with other methods and build more compact cost volumes without redundancy. Experimental results on DTU dataset, BlendedMVS dataset, Tanks and Temples dataset and ETH3D high-res benchmark show that our method achieves promising reconstruction results compared with other state-of-the-art methods. Wanjuan Su, Qingshan Xu 0001, Wenbing Tao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Learning Inverse Depth Regression for Multi-View Stereo with Correlation Cost VolumeabstractDeep learning has shown to be effective for depth inference in multi-view stereo (MVS). However, the scalability and accuracy still remain an open problem in this domain. This can be attributed to the memory-consuming cost volume representation and inappropriate depth inference. Inspired by the group-wise correlation in stereo matching, we propose an average group-wise correlation similarity measure to construct a lightweight cost volume. This can not only reduce the memory consumption but also reduce the computational burden in the cost volume filtering. Based on our effective cost volume representation, we propose a cascade 3D U-Net module to regularize the cost volume to further boost the performance. Unlike the previous methods that treat multi-view depth inference as a depth regression problem or an inverse depth classification problem, we recast multi-view depth inference as an inverse depth regression task. This allows our network to achieve sub-pixel estimation and be applicable to large-scale scenes. Through extensive experiments on DTU dataset and Tanks and Temples dataset, we show that our proposed network with Correlation cost volume and Inverse DEpth Regression (CIDER1), achieves state-of-the-art results, demonstrating its superior performance on scalability and accuracy. Qingshan Xu 0001, Wenbing Tao |
AAAI | 1 |
| 2020 | Planar Prior Assisted PatchMatch Multi-View StereoabstractThe completeness of 3D models is still a challenging problem in multi-view stereo (MVS) due to the unreliable photometric consistency in low-textured areas. Since low-textured areas usually exhibit strong planarity, planar models are advantageous to the depth estimation of low-textured areas. On the other hand, PatchMatch multi-view stereo is very efficient for its sampling and propagation scheme. By taking advantage of planar models and PatchMatch multi-view stereo, we propose a planar prior assisted PatchMatch multi-view stereo framework in this paper. In detail, we utilize a probabilistic graphical model to embed planar models into PatchMatch multi-view stereo and contribute a novel multi-view aggregated matching cost. This novel cost takes both photometric consistency and planar compatibility into consideration, making it suited for the depth estimation of both non-planar and planar regions. Experimental results demonstrate that our method can efficiently recover the depth information of extremely low-textured areas, thus obtaining high complete 3D models and achieving state-of-the-art performance. Qingshan Xu 0001, Wenbing Tao |
AAAI | 1 |
| 2019 | Multi-Scale Geometric Consistency Guided Multi-View StereoabstractIn this paper, we propose an efficient multi-scale geometric consistency guided multi-view stereo method for accurate and complete depth map estimation. We first present our basic multi-view stereo method with Adaptive Checkerboard sampling and Multi-Hypothesis joint view selection (ACMH). It leverages structured region information to sample better candidate hypotheses for propagation and infer the aggregation view subset at each pixel. For the depth estimation of low-textured areas, we further propose to combine ACMH with multi-scale geometric consistency guidance (ACMM) to obtain the reliable depth estimates for low-textured areas at coarser scales and guarantee that they can be propagated to finer scales. To correct the erroneous estimates propagated from the coarser scales, we present a novel detail restorer. Experiments on extensive datasets show our method achieves state-of-the-art performance, recovering the depth estimation not only in low-textured areas but also in details. Qingshan Xu 0001, Wenbing Tao |
CVPR | 1 |
| 2019 | Efficient large-scale geometric verification for structure from motion
Qingshan Xu 0001, Wenbing Tao, Delie Ming |
Pattern Recognit. Lett. | 1 |