Yaping Yan

dblp:154/7651 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-1933-732XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 8 since 2021
YearPublicationVenuePosition
2026 Toward Free-Form Local Feature Matching
abstract
Existing feature matching methods are strongly coupled to their pre-defined position priors. For instance, sparse matchers are coupled to keypoints, and semi-dense matchers are coupled to grids. The coupled position prior dictates the distribution of matching points and imposes inherent limitations on the matcher. Consequently, sparse matchers suffer from a reliance on keypoint repeatability, while semi-dense matchers lack texture-based precision. Our preliminary work RCM leverages the keypoint prior in the source image and the grid prior in the target image, ensuring texture-based precision with keypoints while eliminating reliance on repeatability. However, RCM still relies heavily on keypoints in the source image, inheriting limitations such as sparsity and poor distribution in challenging scenes. To address these challenges, we introduce RCM+, which presents a novel free-form matching paradigm. By combining a position-agnostic encoder with a parameter-free decoder, we decouple the matcher from any position prior. As a result, the free-form matcher can match arbitrary input positions in a zero-shot manner, including detected keypoints, lines, edges, grids of any resolution, user-specified points, and more. This paradigm offers exceptional flexibility, allowing users to select position priors based on scene properties without retraining. Thus, RCM+ can leverage the advantages of various position priors without over-relying on any single prior, avoiding limitations in specific scenarios. To better match multiple position priors, we propose the Balancer, which reconciles all input position priors to achieve a more favorable point distribution for downstream tasks. Additionally, we enhance the view switcher and conflict-free matching layer introduced in RCM, further improving matching quality. Comprehensive experiments demonstrate the excellent performance, efficiency, and flexibility of RCM+, underscoring its promising potential for applications.
Xiaoyong Lu, Songlin Du, Yaping Yan, Xiaobo Lu, Takeshi Ikenaga
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 SceneGlue: Scene-Aware Transformer for Feature Matching Without Scene-Level Annotation
abstract
Local feature matching plays a critical role in understanding the correspondence between cross-view images. However, traditional methods are constrained by the inherent local nature of feature descriptors, limiting their ability to capture non-local scene information that is essential for accurate cross-view correspondence. In this paper, we introduce SceneGlue, a scene-aware feature matching framework designed to overcome these limitations. SceneGlue leverages a hybridizable matching paradigm that integrates implicit parallel attention and explicit cross-view visibility estimation. The parallel attention mechanism simultaneously exchanges information among local descriptors within and across images, enhancing the scene’s global context. To further enrich the scene awareness, we propose the Visibility Transformer, which explicitly categorizes features into visible and invisible regions, providing an understanding of cross-view scene visibility. By combining explicit and implicit scene-level awareness, SceneGlue effectively compensates for the local descriptor constraints. Notably, SceneGlue is trained using only local feature matches, without requiring scene-level groundtruth annotations. This scene-aware approach not only improves accuracy and robustness but also enhances interpretability compared to traditional methods. Extensive experiments on applications such as homography estimation, pose estimation, image matching, and visual localization validate SceneGlues superior performance. The source code is available at https://github.com/songlindu/ SceneGlue.
Songlin Du, Xiaoyong Lu, Yaping Yan, Guobao Xiao, Xiaobo Lu, Takeshi Ikenaga
IEEE Trans. Circuits Syst. Video Technol.3
2026 Progressive disentanglement for robust image anomaly detection
Yaping Yan, Yuanrui Zeng, Songlin Du
Vis. Comput.1
2025 Exploring the Relationship Between Samples and Masks for Robust Defect Localization
abstract
Defect detection aims to detect and localize regions out of the normal distribution. The previous approaches often explicitly incorporate the defect detection concept, such as by utilizing self-supervised ground truth or manually defined feature comparison. The aforementioned processes involve modeling the distribution of normal samples, and they rely on the modeled normality for accurate inference. This reliance may hinder their ability to generalize to unseen test scenarios or the test set that deviates from the training distribution. In this paper, we propose a one-stage framework that detects defective patterns directly without the modeling process. This ability is adopted through the joint efforts of three parties: a generative adversarial network (GAN), a newly proposed scaled pattern loss, and a dynamic correction mechanism that allows the network to self-correct. In training, explicit information that could indicate the position of defects is intentionally excluded to avoid learning any direct mapping. Experimental results show that the proposed method performs superior in comparison with the previous SOTA methods in various test scenarios.
Jiang Lin, Fanxiu Sun, Yaping Yan
AAAI4
2025 Label distribution learning with structured manifold subspace
Yaping Yan, Yunlong Tang 0004, Yongxin Jiang, Songlin Du
Neurocomputing1
2025 Tex2Sem: Learning From Textures to Semantics for Robust Semantic Correspondence
abstract
Recent advances in semantic correspondence have witnessed growing interest in vision foundation models, particularly stable diffusion (SD) and self-distillation with no labels (DINO). However, existing methods underutilize the matching potential of SD and DINOv2 features and show similar background interference patterns. They lack texture-to-semantic learning and intra- and inter-image feature interaction. This study proposes Tex2Sem, a framework learning from textures to semantics, to address the two problems. For the first problem, we propose a texture-to-semantic learning paradigm that achieves texture-semantic trade-offs on features and correlation maps, including progressive fusion and correlation map computation. The SD and DINOv2 features are aggregated from textures to semantics to produce multi-stage progressive fusion features. The resulting multi-stage progressive fusion correlation maps improve semantic correspondence significantly. For the second problem, MamFormer, a hybrid architecture of Mamba-2 and Transformer, is proposed to improve intra- and inter-image feature aggregation and interaction. It enhances foreground focus and background suppression. Given the high computational cost of processing all-stage progressive fusion features, the terminal-stage aggregation and interaction mechanism (TAIM) is proposed to enhance feature learning efficiency. Experiments demonstrate that Tex2Sem achieves state-of-the-art performance on SPair-71k, AP-10K, and PF-PASCAL. Furthermore, Tex2Sem shows remarkable generalization capabilities in cross-species, cross-family, and cross-dataset matching and demonstrates the potential for applications in video swap and human pose estimation. Code is available at https://github.com/wzhlearning/Tex2Sem.
Zenghui Wang 0009, Songlin Du, Yaping Yan, Guobao Xiao, Xiaobo Lu
IEEE Trans. Circuits Syst. Video Technol.3
2024 A Comprehensive Augmentation Framework for Anomaly Detection
abstract
Data augmentation methods are commonly integrated into the training of anomaly detection models. Previous approaches have primarily focused on replicating real-world anomalies or enhancing diversity, without considering that the standard of anomaly varies across different classes, potentially leading to a biased training distribution. This paper analyzes crucial traits of simulated anomalies that contribute to the training of reconstructive networks and condenses them into several methods, thus creating a comprehensive framework by selectively utilizing appropriate combinations. Furthermore, we integrate this framework with a reconstruction-based approach and concurrently propose a split training strategy that alleviates the overfitting issue while avoiding introducing interference to the reconstruction process. The evaluations conducted on the MVTec anomaly detection dataset demonstrate that our method outperforms the previous state-of-the-art approach, particularly in terms of object classes. We also generate a simulated dataset comprising anomalies with diverse characteristics, and experimental results demonstrate that our approach exhibits promising potential for generalizing effectively to various unseen anomalies encountered in real-world scenarios.
Jiang Lin, Yaping Yan
AAAI2
2023 ParaFormer: Parallel Attention Transformer for Efficient Feature Matching
abstract
Heavy computation is a bottleneck limiting deep-learning-based feature matching algorithms to be applied in many real-time applications. However, existing lightweight networks optimized for Euclidean data cannot address classical feature matching tasks, since sparse keypoint based descriptors are expected to be matched. This paper tackles this problem and proposes two concepts: 1) a novel parallel attention model entitled ParaFormer and 2) a graph based U-Net architecture with attentional pooling. First, ParaFormer fuses features and keypoint positions through the concept of amplitude and phase, and integrates self- and cross-attention in a parallel manner which achieves a win-win performance in terms of accuracy and efficiency. Second, with U-Net architecture and proposed attentional pooling, the ParaFormer-U variant significantly reduces computational complexity, and minimize performance loss caused by downsampling. Sufficient experiments on various applications, including homography estimation, pose estimation, and image matching, demonstrate that ParaFormer achieves state-of-the-art performance while maintaining high efficiency. The efficient ParaFormer-U variant achieves comparable performance with less than 50% FLOPs of the existing attention-based models.
Xiaoyong Lu, Yaping Yan, Bin Kang, Songlin Du
AAAI2
2023 MSFORMER: Multi-Scale Transformer with Neighborhood Consensus for Feature Matching
abstract
Existing feature matching methods tend to extract feature descriptors by feeding down-sampled feature maps into a Transformer that is unable to extend feature scales, leading to false correspondences between small-size objects. This paper proposes MSFormer, which uses Transformers situated in different branches to obtain feature descriptors. In one branch, convolutions are integrated into self-attention layers elegantly to compensate for the lack of the local structure information. In another branch, a multi-scale Transformer is proposed through injecting heterogeneous receptive field sizes into tokens. Additionally, a neighborhood consensus mechanism is proposed by re-ranking initial matches to make a constraint of geometric consensus on neighborhood feature descriptors. Extensive experiments on indoor and outdoor pose estimations show that MSFormer outperforms existing state-of-the- art methods by a large margin.
Yaping Yan, Dong Liang 0008, Songlin Du
ICASSP2
2023 Scene-Aware Feature Matching
abstract
Current feature matching methods focus on point-level matching, pursuing better representation learning of individual features, but lacking further understanding of the scene. This results in significant performance degradation when handling challenging scenes such as scenes with large viewpoint and illumination changes. To tackle this problem, we propose a novel model named SAM, which applies attentional grouping to guide Scene-Aware feature Matching. SAM handles multi-level features, i.e., image tokens and group tokens, with attention layers, and groups the image tokens with the proposed token grouping module. Our model can be trained by ground-truth matches only and produce reasonable grouping results. With the sense-aware grouping guidance, SAM is not only more accurate and robust but also more interpretable than conventional feature matching models. Sufficient experiments on various applications, including homography estimation, pose estimation, and image matching, demonstrate that our model achieves state-of-the-art performance.
Xiaoyong Lu, Yaping Yan, Songlin Du
ICCV2
2023 Bi-SCM: bidirectional spiking cortical model with adaptive unsharp masking for mammography image enhancement
Yaping Yan, Hongjuan Zhang, Songlin Du, Yide Ma
Multim. Tools Appl.1
2022 JointFusionNet: Parallel Learning Human Structural Local and Global Joint Features for 3D Human Pose Estimation
Zhiwei Yuan, Yaping Yan, Songlin Du, Takeshi Ikenaga
ICANN (4)2
2022 Learning from Noisy Labels via Meta Credible Label Elicitation
abstract
Deep neural networks (DNNS) can easily overfit to noisy data, which leads to a significant degradation of performance. Previous efforts are primarily made by label correction or sample selection to alleviate supervision problem. To distinguish between noisy labels and clean labels, we propose a meta-learning framework which could gradually elicit credible labels via the meta-gradient descent step under the guidance of potentially non-noisy samples. Specifically, by exploiting the topological information of feature space, we can automatically estimate label confidence with a meta-learner. An iterative procedure is designed to select the most trustworthy noisy-labeled instances to generate pseudo labels. Then we train DNNs with pseudo supervision and original noisy super vision, which learns sufficiency and robustness properties in a joint learning objective. Experimental results on benchmark classification datasets show the superiority of our approach against the state-of-the-art methods.
Ziyang Gao, Yaping Yan, Xin Geng 0001
ICIP2
2022 More Than Accuracy: An Empirical Study of Consistency Between Performance and Interpretability
Dong Liang 0008, Rong Quan, Songlin Du, Yaping Yan
PRICAI (3)5
2020 Accumulated and aggregated shifting of intensity for defect detection on micro 3D textured surfaces
Yaping Yan, Shun'ichi Kaneko, Hirokazu Asano
Pattern Recognit.1
2018 Accumulated Aggregation Shifting Based on Feature Enhancement for Defect Detection on 3D Textured Low-Contrast Surfaces
abstract
Detecting defects on 3D textured low-contrast surfaces plays an important role in product quality control. However, because of the affects from uneven distributions of materials, irregular textures, and unclear boundaries between defects and background, this is still a challenging problem. In this paper, a saliency-guided defect detection method, named accumulated aggregation shifting (AAS) model, is proposed to iteratively shift brightness of pixels based on their defective probability. And then, the output sequences of AAS at different iterations can be formalized as linear distribution or exponential distribution through statistical analysis. Finally, by utilizing the risk minimization method, we theoretically determine a reasonable threshold to classify all pixels as defective ones or defect-free ones. This method models defect detection problem under a probabilistic framework. And only a handful of samples are needed for parameter optimization. Experiments on a real-world image dataset for an industrial surface defect detection task demonstrate the effectiveness of our approach.
Yaping Yan, Hirokazu Asano, Shun'ichi Kaneko
ICPR1
2017 LAP: a bio-inspired local image structure descriptor and its applications
Songlin Du, Yaping Yan, Yide Ma
Multim. Tools Appl.2
2016 When spatial distribution unites with spatial contrast: an effective blind image quality assessment model
abstract
Blind image quality assessment (BIQA), which aims to estimate the perceptual quality of images without any reference information, is a very important yet challenging task. Although human visual system is sensitive to degradations on both spatial contrast and spatial distribution, most of the existing structural degradation based BIQA models consider only one of them. This study introduces a novel BIQA model by taking into account degradations on both contrast and spatial distribution. First, the authors construct a multi‐threshold local tetra pattern (MTLTrP) instead of local binary pattern to measure the changes on spatial distribution. Second, Weber–Laplacian of Gaussian (WLOG) operator, which responds to intensity contrast in a small spatial neighbourhood, is proposed to extract local contrast features. Finally, the joint statistics of MTLTrP and WLOG are utilised for BIQA model learning. Experimental results on three large benchmark databases demonstrate that the proposed model outperforms state‐of‐the‐art BIQA models, as well as with several well‐known full reference quality assessment methods.
Yaping Yan, Songlin Du, Hongjuan Zhang, Yide Ma
IET Image Process.1
2015 Quantum-Accelerated Fractal Image Compression: An Interdisciplinary Approach
abstract
Fractal image compression (FIC) is one of the most widely approved image compression approaches for its high compression ratio and quality of retrieved images. However, FIC suffers from high computational cost in searching local self-similarities in natural image. Although many papers aiming at speeding up FIC have been published, they use pre-processing tools or approximation methods. Reducing the intrinsic computational complexity of FIC is still an open problem. Since quantum mechanics based Grover's quantum search algorithm (QSA) is able to achieve square-root speedup over classical algorithms in unsorted database searching, we propose an interdisciplinary approach by using Grover's QSA to reduce the intrinsic computational complexity of FIC in this letter. In particular, both domain blocks and range blocks are represented as quantum states, then Grover's QSA is employed to search the most similar domain block for each range block under the criterion of maximizing quantum fidelity between these two kinds of quantum states. Without sacrificing compression ratio, experimental results show that the execution time of the proposed method is 100 times shorter than that of the baseline FIC. Moreover, retrieved images from our proposal are also less distorted than those from other state-of-the-art FIC approaches.
Songlin Du, Yaping Yan, Yide Ma
IEEE Signal Process. Lett.2