EDBT 2026 Demo / reviewers in the wild / expert
Li Yao 0003
dblp:83/6976-3
· DBLP profile ↗
14ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0003-2930-8407ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PUNO: A Neural Operator Framework for Point Cloud UpsamplingabstractWe propose PUNO, a novel deep operator-based framework for point cloud upsampling, addressing the challenge of reconstructing high-resolution geometries from sparse point clouds. PUNO generalizes the neural operators proven effective in image super-resolution to 3D point cloud upsampling. Moreover, it first designs a network for point cloud tasks to achieve vertex displacement and manifold parameterization, thereby forming a coarse geometric representation that is compatible with super-resolution neural operators. This is followed by iterative kernel integral approximations in the function space and backprojection to generate the target coordinates, fully utilizing the high-frequency information in the function space. Unlike prior work, PUNO performs transformations in both the data domain and the function domain, with the solution space containing richer basis functions, yielding finer results that mitigate the ill-posed nature of sparse data. It also benefits global continuity. Extensive experiments demonstrate its superior accuracy, robustness, and generalization ability. Zijian Xiao, Yining Xu 0001, Yingjie Huang 0001, Li Yao 0003 |
AAAI | 4 |
| 2026 | Mamba-Driven Multi-View Discriminative Clustering via Global-Local Cross-View Sequence ModelingabstractMulti-view clustering (MVC) has recently garnered increasing attention for its ability to partition unlabeled samples into distinct clusters by leveraging complementary and consistent information from different views. Existing MVC methods primarily combine deep neural networks with contrastive learning for cross-view representation learning, yet often overlook the inherent global-local structural relationships among samples. While GNN-based methods capture local structures, they struggle to model global dependencies, leading to inferior inter-cluster separability. In contrast, Transformer-based methods excel at global aggregation but suffer from quadratic complexity, and their attention smoothing effect weakens fine-grained local structures, resulting in suboptimal intra-cluster compactness. To address these limitations, we propose a novel end-to-end MVC framework called Mamba-Driven Multi-View Discriminative Clustering via Global-Local Cross-View Sequence Modeling (MGLC). By flexibly constructing multi-view sequences, MGLC fully exploits the efficient sequence modeling capabilities of Mamba to jointly model cross-view dependencies and global-local structural relationships among samples. Furthermore, MGLC introduces a Cross-Mamba Fusion module to dynamically integrate cross-view and global-local structural representations. Additionally, MGLC incorporates a Dual Calibration Contrastive Learning module, guided by high-confidence pseudo-labels, that adaptively refines both feature and semantic representations while mitigating false negatives among semantically similar samples. Extensive comparative experiments and ablation studies demonstrate the effectiveness of MGLC. Yuanyang Zhang, Xinhang Wan, Jie Xu 0044, Cunjian Chen, Tien-Tsin Wong, Li Yao 0003, Yijie Lin 0001 |
AAAI | 7 |
| 2026 | DiVE: Decoupling Intra-layer Visual Evidence for Mitigating Hallucinations in Large Vision-Language ModelsabstractRecent Large Vision-Language Models (LVLMs) have achieved significant progress yet frequently suffer from visual hallucinations, often stemming from an over-reliance on language priors rather than visual evidence.Existing decoding-based approaches often rely on input perturbations to weaken language priors, but they do not explicitly decouple visual evidence from mixed vision-language representations.To address these limitations, we propose DiVE (Decoupling intra-layer Visual Evidence).DiVE dynamically identifies layers enriched with visual information and performs intra-layer decoupling to extract aggregated visual evidence.By suppressing this evidence to construct a language-priordominated reference distribution, DiVE employs contrastive decoding to calibrate the output logits, thereby mitigating hallucinations.Extensive experiments across diverse LVLM architectures demonstrate that DiVE achieves state-of-the-art performance among decodingbased methods on multiple benchmarks.Crucially, it eliminates the latency of an extra forward pass, offering a lightweight and efficient solution. Li Lin 0011, Hui Jiao, Li Yao 0003, Tien-Tsin Wong, Hanqian Wu |
ACL (1) | 4 |
| 2026 | SaF-AD: Saliency-Adaptive and Feature-Consistent Diffusion for Industrial Anomaly DetectionabstractDiffusion models have recently shown strong potential for reconstr-uction-based unsupervised anomaly detection (UAD). However, industrial UAD across diverse object categories remains challenging: subtle defects often overlap with intrinsic structural details, and commonly used uniform or semantically agnostic perturbations can induce two failure modes—identity shortcut (copying uncorrupted content, resulting in misleading residuals) and semantic drift (over-smoothed yet structurally inconsistent restorations). We propose SaF-AD, a saliency-adaptive and feature-consistent diffusion framework to mitigate these issues. First, Saliency-Adaptive Perturbation Masking (SAPM) applies soft, saliency-guided masking to adaptively corrupt informative regions while avoiding hard boundaries, encouraging structure-aware reconstruction instead of background redundancy. Second, Progressive Anchor Decoupling (PAD) progressively adjusts the masking preference during training to reduce persistent anchors and prevent shortcut learning, forcing reconstruction of salient structures from diverse contextual cues. Third, Hierarchical Semantic Feature Consistency (HSFC) regularizes multi-level features on corrupted regions using a frozen backbone, improving semantic coherence while preserving fine-grained details. Experiments on MVTec-AD and VisA show that SaF-AD achieves competitive image-level detection performance and more consistent gains on pixel-level anomaly localization. Yuanyang Zhang, Zirui Luo, Kaixi Xu, Yining Xu 0001, Li Yao 0003 |
ICMR | 6 |
| 2026 | SAND: Semantic-Aware Anomaly Detection with Region-Consistent Memory for Noisy TrainingabstractIndustrial anomaly detection is typically trained under the assumption of clean normal data, yet real-world manufacturing datasets are often contaminated by unlabeled defects and background outliers. Global memory-based detectors are especially fragile in such noisy regimes: semantic mismatch across regions can induce erroneous retrieval, while boundary-crossing patches yield ambiguous pseudo supervision. We propose SAND, a semantic-aware framework for unsupervised anomaly detection under noisy training data that leverages training-only semantic priors derived from pre-computed semantic region partitions. SAND introduces a region-consistent memory design that constrains retrieval to the same semantic region, with a global fallback when region assignment is uncertain, reducing cross-region aliasing. To further stabilize learning, we develop purity-aware soft weighting with adaptive thresholds to downweight boundary-ambiguous pseudo labels. We additionally propose a region-wise representation regularizer that enforces intra-region compactness and inter-region separation, thereby mitigating anomaly contamination of prototypes. Extensive experiments on MVTec AD and VisA under multiple noise protocols, including realistic background perturbations, demonstrate consistent improvements in both image-level detection and pixel-level localization, particularly in multi-object scenarios. Notably, SAND achieves these gains without semantic priors or test-time retrieval, supporting efficient inference for practical deployment. Our results highlight the importance of semantic provenance in noisy anomaly detection and suggest a practical pathway toward robust industrial deployment. Kaixi Xu, Yuanyang Zhang, Zirui Luo, Li Yao 0003 |
ICMR | 6 |
| 2026 | Ref2Inpaint: 3D Gaussian Inpainting via Visibility-Aware Mask Refinement and VLM-Guided Reference Retrievalabstract3D scene inpainting aims to restore geometrically and texturally consistent content after object removal, enabling immersive scene editing and virtual content creation. Despite rapid progress in neural 3D reconstruction and rendering (e.g., Neural Radiance Fields and 3D Gaussian Splatting), achieving accurate and artifact-free 3D completion remains challenging. In particular, (i) imprecise 2D masks yield unreliable inpainting scopes, (ii) selecting high-quality 2D reference views for lifting to 3D is difficult due to view-dependent perceptual fidelity, and (iii) integrating 2D priors into 3D often introduces blurred textures and structural artifacts. These issues can accumulate and amplify as inconsistencies are fused into the 3D representation. We propose Ref2Inpaint, a geometry-aware and reference-guided framework for high-quality 3D scene inpainting. First, our Visibility-Aware Mask Refinement aggregates cross-view visibility cues to suppress erroneous masked regions and establish a spatially consistent inpainting scope. Second, our VLM-Guided Reference Retrieval combines geometric filtering with VLM-based quality ranking to select high-fidelity, cross-view consistent references for 3D initialization and inpainting guidance. Finally, a Two-Stage Structural Densification progressively reconstructs missing geometry from coarse layouts to fine-grained details, reducing floaters and boundary artifacts while improving structural plausibility. Our work demonstrates that advanced retrieval mechanisms can significantly alleviate the texture inconsistency issue in generative 3D tasks. Extensive experiments on both real and synthetic scenes demonstrate that Ref2Inpaint achieves superior visual fidelity, geometric coherence, and multi-view consistency compared to state-of-the-art methods. Yining Xu 0001, Yuanyang Zhang, Jingjiao You, Yingjie Huang 0001, Jianbo Mei, Li Yao 0003 |
ICMR | 7 |
| 2026 | GlassSplat: Geometric Consistency and Pruning for Reflection-Free 3D Scene ReconstructionabstractRendering high-fidelity 3D scenes is crucial for immersive applications like virtual reality and digital twins. However, standard 3D Gaussian Splatting (3DGS) relies heavily on multi-view consistency, making it fragile in real-world scenarios plagued by glass reflections. These reflections often manifest as geometric "floaters" or severe texture artifacts, obscuring the true background. Existing solutions, which typically employ single-image priors or NeRF-based in-painting, often lack explicit 3D constraints or rely on synthetic data, failing to generalize to complex environments. To address these challenges, we first present a novel benchmark dataset of 8 real-world scenes, capturing physically paired reflective and reflection-free images. Building on this, we propose GlassSplat, a robust framework designed to eliminate view-dependent artifacts and recover clean transmission geometry. Our method initializes with a reflection prior and introduces an Affine-Based Exposure Correction module to align global photometric inconsistencies. To distinguishing valid geometry from virtual outliers, we incorporate an Epipolar Consistency Loss and an uncertainty-weighted Depth Regularization. Finally, to physically purge residual noise, we devise a Visibility-Aware Pruning strategy that dynamically filters artifacts based on multi-view statistics. Extensive experiments demonstrate that GlassSplat significantly outperforms state-of-the-art approaches, effectively recovering a clean, artifact-free 3D scene representation. Jingjiao You, Yuanyang Zhang, Yining Xu 0001, Li Yao 0003, Cunjian Chen, Tien-Tsin Wong |
ICMR | 5 |
| 2026 | Structure-Aware Conditional Diffusion Generation for Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering (IMVC) has attracted increasing attention in recent years, owing to the prevalence of missing data in real-world multi-view scenarios. Existing imputation-based IMVC methods partially mitigate the impact of missing information but still face three key limitations: (i) overlooking latent structural relationships among samples, which leads to imputed representations deviating from the true distribution; (ii) decoupling imputation from clustering, which reduces the discriminability of the recovered representations; and (iii) exhibiting low efficiency, which makes it difficult to balance recovery quality and inference speed under complex missing scenarios. To address these issues, we propose a Structure-Aware Conditional Diffusion Generation (SACDG) framework. During training, SACDG first models local structural relationships via adaptive neighborhood graphs and injects them as conditional priors into the diffusion model, where a cross-attention mechanism integrates these priors into the noise prediction process to learn structure-aware generative capability. Meanwhile, a semantic distribution alignment module is introduced to leverage pseudo-labels for enforcing cross-view consistency, thereby enhancing semantic discriminability. During inference, SACDG integrates cross-view structural information through cross-view adjacency fusion to guide the reverse denoising trajectory, and employs deterministic DDIM sampling to efficiently and stably recover the representations of missing views. Extensive comparative experiments and ablation studies on multiple benchmark datasets demonstrate that SACDG achieves superior clustering performance and improved efficiency over state-of-the-art methods. Our code is available athttps://github.com/zhangyuanyang21/SACDG. Yuanyang Zhang, Yijie Lin 0001, Xinhang Wan, Jie Xu 0044, Li Yao 0003, Weiqing Yan, Chang Tang |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2025 | Incomplete Multi-view Clustering via Diffusion Contrastive GenerationabstractIncomplete multi-view clustering (IMVC) has garnered increasing attention in recent years due to the common issue of missing data in multi-view datasets. The primary approach to address this challenge involves recovering the missing views before applying conventional multi-view clustering methods. Although imputation-based IMVC methods have achieved significant improvements, they still encounter notable limitations: 1) heavy reliance on paired data for training the data recovery module, which is impractical in real scenarios with high missing data rates; 2) the generated data often lacks diversity and discriminability, resulting in suboptimal clustering results. To address these shortcomings, we propose a novel IMVC method called Diffusion Contrastive Generation (DCG). Motivated by the consistency between the diffusion and clustering processes, DCG learns the distribution characteristics to enhance clustering by applying forward diffusion and reverse denoising processes to intra-view data. By performing contrastive learning on a limited set of paired multi-view samples, DCG can align the generated views with the real views, facilitating accurate recovery of views across arbitrary missing view scenarios. Additionally, DCG integrates instance-level and category-level interactive learning to exploit the consistent and complementary information available in multi-view data, achieving robust and end-to-end clustering. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches. Yuanyang Zhang, Yijie Lin 0001, Weiqing Yan, Li Yao 0003, Xinhang Wan, Guanzhou Ke, Jie Xu 0044 |
AAAI | 4 |
| 2024 | MicroMamba: State Space Model with Partitioned Window Scan for Micro-Expression Recognition
Jiateng Liu, Li Yao 0003 |
MMAsia | 4 |
| 2024 | LGSur-Net: A Local Gaussian Surface Representation Network for Upsampling Highly Sparse Point CloudabstractAbstract We introduce LGSur‐Net, an end‐to‐end deep learning architecture, engineered for the upsampling of sparse point clouds. LGSur‐Net harnesses a trainable Gaussian local representation by positioning a series of Gaussian functions on an oriented plane, complemented by the optimization of individual covariance matrices. The integration of parametric factors allows for the encoding of the plane's rotational dynamics and Gaussian weightings into a linear transformation matrix. Then we extract the feature maps from the point cloud and its adjoining edges and learn the local Gaussian depictions to accurately model the shape's local geometry through an attention‐based network. The Gaussian representation's inherent high‐order continuity endows LGSur‐Net with the natural ability to predict surface normals and support upsampling to any specified resolution. Comprehensive experiments validate that LGSur‐Net efficiently learns from sparse data inputs, surpassing the performance of existing state‐of‐the‐art upsampling methods. Our code is publicly available at https://github.com/Rangiant5b72/LGSur-Net . Zijian Xiao, Li Yao 0003 |
Comput. Graph. Forum | 3 |
| 2023 | Multi-Level Implicit Function for Detailed Human Reconstruction by Relaxing SMPL ConstraintsabstractAbstract Aiming at enhancing the rationality and robustness of the results of single‐view image‐based human reconstruction and acquiring richer surface details, we propose a multi‐level reconstruction framework based on implicit functions. This framework first utilizes the predicted SMPL model (Skinned Multi‐Person Linear Model) as a prior to further predict consistent 2.5D sketches (depth map and normal map), and then obtains a coarse reconstruction result through an Implicit Function fitting network (IF‐Net). Subsequently, with a pixel‐aligned feature extraction module and a fine IF‐Net, the strong constraints imposed by SMPL are relaxed to add more surface details to the reconstruction result and remove noise. Finally, to address the trade‐off between surface details and rationality under complex poses, we propose a novel fusion repair algorithm that reuses existing information. This algorithm compensates for the missing parts of the fine reconstruction results with the coarse reconstruction results, leading to a robust, rational, and richly detailed reconstruction. The final experiments prove the effectiveness of our method and demonstrate that it achieves the richest surface details while ensuring rationality. The project website can be found at https://github.com/MXKKK/2.5D‐MLIF . Xikai Ma, Yiqing Teng, Li Yao 0003 |
Comput. Graph. Forum | 4 |
| 2019 | Fast and high-quality virtual view synthesis from multi-view plus depth videos
Li Yao 0003, Yingdong Han |
Multim. Tools Appl. | 1 |
| 2019 | 2D-to-3D conversion using optical flow based depth generation and cross-scale hole filling algorithm
Li Yao 0003, Zhukui Liu, Bingfeng Wang |
Multim. Tools Appl. | 1 |