VLDB 2026 Research / reviewers in the wild / expert
Yingna Wu
dblp:332/3576
· DBLP profile ↗
10ranked-venue papers
0as first author
10since 2021 · last 2026
0009-0005-4858-3852ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unsupervised Multi-View Visual Anomaly Detection via Progressive Homography-Guided AlignmentabstractUnsupervised visual anomaly detection from multi-view images presents a significant challenge: distinguishing genuine defects from benign appearance variations caused by viewpoint changes. Existing methods, often designed for single-view inputs, treat multiple views as a disconnected set of images, leading to inconsistent feature representations and a high false-positive rate. To address this, we introduce ViewSense-AD (VSAD), a novel framework that learns viewpoint-invariant representations by explicitly modeling geometric consistency across views. At its core is our Multi-View Alignment Module (MVAM), which leverages homography to project and align corresponding feature regions between neighboring views. We integrate MVAM into a View-Align Latent Diffusion Model (VALDM), enabling progressive and multi-stage alignment during the denoising process. This allows the model to build a coherent and holistic understanding of the object's surface from coarse to fine scales. Furthermore, a lightweight Fusion Refiner Module (FRM) enhances the global consistency of the aligned features, suppressing noise and improving discriminative power. Anomaly detection is performed by comparing multi-level features from the diffusion model against a learned memory bank of normal prototypes. Extensive experiments on the challenging RealIAD and MANTA datasets demonstrate that VSAD sets a new state-of-the-art, significantly outperforming existing methods in pixel, view, and sample-level visual anomaly detection, proving its robustness to large viewpoint shifts and complex textures. Xintao Chen, Xiaohao Xu, Bozhong Zheng, Yingna Wu |
AAAI | 5 |
| 2025 | Towards Visual Discrimination and Reasoning of Real-World Physical Dynamics: Physics-Grounded Anomaly DetectionabstractHumans detect real-world object anomalies by perceiving, interacting, and reasoning based on object-conditioned physical knowledge. The long-term goal of Industrial Anomaly Detection (IAD) is to enable machines to autonomously replicate this skill. However, current IAD algorithms are largely developed and tested on static, semantically simple datasets, which diverge from real-world scenarios where physical understanding and reasoning are essential. To bridge this gap, we introduce the Physics Anomaly Detection (Phys-AD) dataset, the first large-scale, real-world, physics-grounded video dataset for industrial anomaly detection. Collected using a real robot arm and motor, Phys-AD provides a diverse set of dynamic, semantically rich scenarios. The dataset includes more than 6400 videos across 22 real-world object categories, interacting with robot arms and motors, and exhibits 47 types of anomalies. Anomaly detection in Phys-AD requires visual reasoning, combining both physical knowledge and video content to determine object abnormality. We benchmark state-of-the-art anomaly detection methods under three settings: unsupervised AD, weakly-supervised AD, and video-understanding AD, highlighting their limitations in handling physics-grounded anomalies. Additionally, we introduce the Physics Anomaly Explanation (PAEval) metric, designed to assess the ability of visual-language foundation models to not only detect anomalies but also provide accurate explanations for their underlying physical causes. Our project is available at https://guyao2023.github.io/Phys-AD/. Wenqiao Li, Yao Gu, Xintao Chen, Xiaohao Xu, Xiaonan Huang, Yingna Wu |
CVPR | 7 |
| 2025 | Multi-Sensor Object Anomaly Detection: Unifying Appearance, Geometry, and Internal PropertiesabstractObject anomaly detection is essential for industrial quality inspection, yet traditional single-sensor methods face critical limitations. They fail to capture the wide range of anomaly types, as single sensors are often constrained to either external appearance, geometric structure, or internal properties. To overcome these challenges, we introduce MulSen-AD, the first high-resolution, multi-sensor anomaly detection dataset tailored for industrial applications. MulSen-AD unifies data from RGB cameras, laser scanners, and lock-in infrared thermography, effectively capturing external appearance, geometric deformations, and internal defects. The dataset spans 15 industrial products with diverse, real-world anomalies. We also present MulSen-AD Bench, a benchmark designed to evaluate multi-sensor methods, and propose MulSen-TripleAD, a decision-level fusion algorithm that integrates these three modalities for robust, unsupervised object anomaly detection. Our experiments demonstrate that multi-sensor fusion substantially outperforms single-sensor approaches, achieving 96.1% AUROC in object-level detection accuracy. These results highlight the importance of integrating multi-sensor data for comprehensive industrial anomaly detection. The dataset and code are available at https://github.com/ZZZBBBZZZ/MulSen-ADtosupportfurtherresearch. Wenqiao Li, Bozhong Zheng, Xiaohao Xu, Jinye Gan, Fading Lu, Na Ni, Xiaonan Huang, Shenghua Gao, Yingna Wu |
CVPR | 11 |
| 2025 | 3DSignDiff: Towards 3D Sign Language Gesture GenerationabstractSign language is vital for communication within the deaf and hard-of-hearing community. Deep learning techniques, such as Generative Adversarial Networks and diffusion models, have shown promise in generating sign language video content. However, existing methods typically focus on 2D sign language video generation, with limited exploration into 3D sign language generation. A prominent issue is the lack of an available 3D sign language dataset for training and evaluating 3D sign language video generation models. To address this, we introduce a new dataset called How2Gesture, which contains long sequences of 3D sign language actions along with detailed text translations. We also propose a novel 3D sign language generation paradigm, named 3DSignDiff. This model is based on a Transformer encoder, directly conditioning on long text sequences, and uses a diffusion model as a denoiser. It also employs text context enhancement techniques to strengthen the guidance of text on sign language actions. Additionally, an advanced joint generation approach is used to generate human sign language actions in stages. Extensive experiments conducted on the open-source How2Gesture dataset validate the superiority of our model. How2Geature dataset is available at https://github.com/lifeisawar41/3Dsigndiff. Ronghao Yu, Yun Liu 0029, Xiyue Bai, Yingna Wu |
ICASSP | 5 |
| 2025 | Bridging 3D Anomaly Localization and Repair Via High-Quality Continuous Geometric Representationabstract3D point cloud anomaly detection is essential for robust vision systems but is challenged by pose variations and complex geometric anomalies. Existing patch-based methods often suffer from geometric fidelity issues due to discrete voxelization or projection-based representations, limiting fine-grained anomaly localization. We introduce Pose-Aware Signed Distance Field (PASDF), a novel framework that integrates 3D anomaly detection and repair by learning a continuous, pose-invariant shape representation. PASDF leverages a Pose Alignment Module for canonicalization and a SDF Network to dynamically incorporate pose, enabling implicit learning of high-fidelity anomaly repair templates from the continuous SDF. This facilitates precise pixel-level anomaly localization through an Anomaly-Aware Scoring Module. Crucially, the continuous 3D representation in PASDF extends beyond detection, facilitating in-situ anomaly repair. Experiments on Real3D-AD and Anomaly-ShapeNet demonstrate state-of-the-art performance, achieving high object-level AUROC scores of 80.2% and 90.0%, respectively. These results highlight the effectiveness of continuous geometric representations in advancing 3D anomaly detection and facilitating practical anomaly region repair. The code is available at https://github.com/ZZZBBBZZZ/PASDF to support further research. Bozhong Zheng, Jinye Gan, Xiaohao Xu, Xintao Chen, Wenqiao Li, Xiaonan Huang, Na Ni, Yingna Wu |
ICCV | 8 |
| 2025 | Progressive self-supervised learning: A pre-training method for crowd counting
Yao Gu, Yingna Wu, Guangping Xie, Na Ni |
Pattern Recognit. Lett. | 3 |
| 2024 | Towards Scalable 3D Anomaly Detection and Localization: A Benchmark via 3D Anomaly Synthesis and A Self-Supervised Learning NetworkabstractRecently, 3D anomaly detection, a crucial problem in-volving fine-grained geometry discrimination, is getting more attention. However, the lack of abundant real 3D anomaly data limits the scalability of current models. To enable scalable anomaly data collection, we propose a 3D anomaly synthesis pipeline to adapt existing large-scale 3D models for 3D anomaly detection. Specifically, we construct a synthetic dataset, i.e., Anomaly-ShapeNet, based on ShapeNet. Anomaly-ShapeNet consists of 1600 point cloud samples under 40 categories, which provides a rich and varied collection of data, enabling efficient training and enhancing adaptability to industrial scenarios. Meanwhile, to enable scalable representation learning for 3D anomaly localization, we propose a self-supervised method, i.e., It-erative Mask Reconstruction Network (IMRNet). During training, we propose a geometry-aware sample module to preserve potentially anomalous local regions during point cloud down-sampling. Then, we randomly mask out point patches and sent the visible patches to a trans-former for reconstruction-based self-supervision. During testing, the point cloud repeatedly goes through the Mask Reconstruction Network, with each iteration's output be-coming the next input. By merging and contrasting the final reconstructed point cloud with the initial input, our method successfully locates anomalies. Experiments show that IMRNet outperforms previous state-of-the-art methods, achieving 66.1% in I-AUC on our Anomaly-ShapeNet dataset and 72.5% in I-AUC on ReaI3D-AD dataset. Our benchmark will be released at https://github.com/Chopper-233/Anomaly-ShapeNet. Wenqiao Li, Xiaohao Xu, Yao Gu, Bozhong Zheng, Shenghua Gao, Yingna Wu |
CVPR | 6 |
| 2023 | Prototype calibration for long tailed recognitionabstractIn the real world, data distribution usually presents imbalanced characteristics, such as long-tailed distribution, which is generally divided into the head and tail classes. For tail classes, image features cannot be represented well due to insufficient training samples. It is a vital task to learn discriminative image representation on imbalanced data distribution. In our work, through exploring prototype information, we propose a prototype-based contrastive learning(PCL) loss and prototype-based feature augmentation(PFA) module to improve the accuracy of the classifier on the imbalanced dataset. Specifically, we utilize the classifier parameters to generate learnable embeddings, which can be regarded as the class centers after using metric learning. The PFA module generates the image features of each tail class with the help of head class information. We validate our approach on common long-tailed benchmarks. Our results indicate that the PCL and PFA make the classification model achieve significant performance boosts on these benchmarks. Zhongan Wang, Yingna Wu |
ICME | 3 |
| 2022 | HARNet: Hierarchical adaptive regression with location recovery for crowd counting
Na Ni, Guangping Xie, Aichun Zhu, Yingna Wu |
Neurocomputing | 5 |
| 2022 | ArCo: Attention-reinforced transformer with contrastive learning for image captioningabstractImage captioning is a significant step toward achieving automatic interactions between humans and computers, in which a textual sequence of the content of an image is generated. Recently, the transformer-based encoder–decoder paradigm has made great achievements in image captioning. This method is usually trained with a cross-entropy loss function. However, for various captions of images with the same meaning, the computed losses may be different. The result is that the descriptions of images tend to be consistent, which limits the diversity of image captioning. In this paper, we present an attention-reinforced transformer, a transformer-based architecture for image captioning. The architecture improves the image encoding stage, which exploits the relationships between image regions by integrating a feature attention block (FAB). During the training phase, we trained the model with a combination of cross-entropy loss and contrastive loss. We experimentally explored the performance of ArCo and other fully attentive models. We also validated the baseline of the transformer for image captioning with different pre-trained models. Our proposed approach was demonstrated to achieve a new state-of-the-art performance on the offline ‘Karpathy’ test split and online test server. Zhongan Wang, Zirong Zhai, Yingna Wu |
Image Vis. Comput. | 4 |