Yining Xu 0001

dblp:183/5633-1 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0001-2472-3460ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PUNO: A Neural Operator Framework for Point Cloud Upsampling
abstract
We propose PUNO, a novel deep operator-based framework for point cloud upsampling, addressing the challenge of reconstructing high-resolution geometries from sparse point clouds. PUNO generalizes the neural operators proven effective in image super-resolution to 3D point cloud upsampling. Moreover, it first designs a network for point cloud tasks to achieve vertex displacement and manifold parameterization, thereby forming a coarse geometric representation that is compatible with super-resolution neural operators. This is followed by iterative kernel integral approximations in the function space and backprojection to generate the target coordinates, fully utilizing the high-frequency information in the function space. Unlike prior work, PUNO performs transformations in both the data domain and the function domain, with the solution space containing richer basis functions, yielding finer results that mitigate the ill-posed nature of sparse data. It also benefits global continuity. Extensive experiments demonstrate its superior accuracy, robustness, and generalization ability.
Zijian Xiao, Yining Xu 0001, Yingjie Huang 0001, Li Yao 0003
AAAI2
2026 SaF-AD: Saliency-Adaptive and Feature-Consistent Diffusion for Industrial Anomaly Detection
abstract
Diffusion models have recently shown strong potential for reconstr-uction-based unsupervised anomaly detection (UAD). However, industrial UAD across diverse object categories remains challenging: subtle defects often overlap with intrinsic structural details, and commonly used uniform or semantically agnostic perturbations can induce two failure modes—identity shortcut (copying uncorrupted content, resulting in misleading residuals) and semantic drift (over-smoothed yet structurally inconsistent restorations). We propose SaF-AD, a saliency-adaptive and feature-consistent diffusion framework to mitigate these issues. First, Saliency-Adaptive Perturbation Masking (SAPM) applies soft, saliency-guided masking to adaptively corrupt informative regions while avoiding hard boundaries, encouraging structure-aware reconstruction instead of background redundancy. Second, Progressive Anchor Decoupling (PAD) progressively adjusts the masking preference during training to reduce persistent anchors and prevent shortcut learning, forcing reconstruction of salient structures from diverse contextual cues. Third, Hierarchical Semantic Feature Consistency (HSFC) regularizes multi-level features on corrupted regions using a frozen backbone, improving semantic coherence while preserving fine-grained details. Experiments on MVTec-AD and VisA show that SaF-AD achieves competitive image-level detection performance and more consistent gains on pixel-level anomaly localization.
Yuanyang Zhang, Zirui Luo, Kaixi Xu, Yining Xu 0001, Li Yao 0003
ICMR5
2026 Ref2Inpaint: 3D Gaussian Inpainting via Visibility-Aware Mask Refinement and VLM-Guided Reference Retrieval
abstract
3D scene inpainting aims to restore geometrically and texturally consistent content after object removal, enabling immersive scene editing and virtual content creation. Despite rapid progress in neural 3D reconstruction and rendering (e.g., Neural Radiance Fields and 3D Gaussian Splatting), achieving accurate and artifact-free 3D completion remains challenging. In particular, (i) imprecise 2D masks yield unreliable inpainting scopes, (ii) selecting high-quality 2D reference views for lifting to 3D is difficult due to view-dependent perceptual fidelity, and (iii) integrating 2D priors into 3D often introduces blurred textures and structural artifacts. These issues can accumulate and amplify as inconsistencies are fused into the 3D representation. We propose Ref2Inpaint, a geometry-aware and reference-guided framework for high-quality 3D scene inpainting. First, our Visibility-Aware Mask Refinement aggregates cross-view visibility cues to suppress erroneous masked regions and establish a spatially consistent inpainting scope. Second, our VLM-Guided Reference Retrieval combines geometric filtering with VLM-based quality ranking to select high-fidelity, cross-view consistent references for 3D initialization and inpainting guidance. Finally, a Two-Stage Structural Densification progressively reconstructs missing geometry from coarse layouts to fine-grained details, reducing floaters and boundary artifacts while improving structural plausibility. Our work demonstrates that advanced retrieval mechanisms can significantly alleviate the texture inconsistency issue in generative 3D tasks. Extensive experiments on both real and synthetic scenes demonstrate that Ref2Inpaint achieves superior visual fidelity, geometric coherence, and multi-view consistency compared to state-of-the-art methods.
Yining Xu 0001, Yuanyang Zhang, Jingjiao You, Yingjie Huang 0001, Jianbo Mei, Li Yao 0003
ICMR1
2026 GlassSplat: Geometric Consistency and Pruning for Reflection-Free 3D Scene Reconstruction
abstract
Rendering high-fidelity 3D scenes is crucial for immersive applications like virtual reality and digital twins. However, standard 3D Gaussian Splatting (3DGS) relies heavily on multi-view consistency, making it fragile in real-world scenarios plagued by glass reflections. These reflections often manifest as geometric "floaters" or severe texture artifacts, obscuring the true background. Existing solutions, which typically employ single-image priors or NeRF-based in-painting, often lack explicit 3D constraints or rely on synthetic data, failing to generalize to complex environments. To address these challenges, we first present a novel benchmark dataset of 8 real-world scenes, capturing physically paired reflective and reflection-free images. Building on this, we propose GlassSplat, a robust framework designed to eliminate view-dependent artifacts and recover clean transmission geometry. Our method initializes with a reflection prior and introduces an Affine-Based Exposure Correction module to align global photometric inconsistencies. To distinguishing valid geometry from virtual outliers, we incorporate an Epipolar Consistency Loss and an uncertainty-weighted Depth Regularization. Finally, to physically purge residual noise, we devise a Visibility-Aware Pruning strategy that dynamically filters artifacts based on multi-view statistics. Extensive experiments demonstrate that GlassSplat significantly outperforms state-of-the-art approaches, effectively recovering a clean, artifact-free 3D scene representation.
Jingjiao You, Yuanyang Zhang, Yining Xu 0001, Li Yao 0003, Cunjian Chen, Tien-Tsin Wong
ICMR3