Tangfei Liao

dblp:363/8236 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
0009-0001-7948-6829ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Scalable and Generalizable Correspondence Pruning via Geometry-Consistent Pre-Training
abstract
Two-view correspondence pruning aims to identify reliable correspondences for camera pose estimation, serving as a fundamental step in many 3D vision tasks. Existing methods rely on geometric consistency to seek true correspondences (inliers) from numerous false correspondences (outliers). In this learning paradigm, outliers severely affect the representation learning of inliers, resulting in models that are neither robust nor generalizable. To address this issue, we propose a geometry-consistent pre-training paradigm that sculpts scalable and generalizable representations free from outlier interference. The paradigm features two appealing properties. 1) Implementation of geometry-consistent pre-training. We introduce masked inlier reconstruction as a pretext task and develop a simple yet effective pre-training framework based on a masked autoencoder. Specifically, due to the irregular and unordered nature of correspondences, which lack explicit positional information, we adopt a dual-branch structure that separately reconstructs the keypoints of two images. This enables indirect reconstruction of 4D correspondences, where keypoints from the paired image provide positional prompts. 2) Unified correspondence encoder. We propose a simple dual-stream encoder with built-in consensus interaction, providing a unified, extensible architecture that enhances representation learning. Extensive experiments demonstrate that our method, GeneralPruner, consistently outperforms state-of-the-art approaches in terms of robustness and generalization across various downstream tasks. Specifically, our method achieves 10.76%, 11.84%, and 8.65% performance gains in camera pose estimation, visual localization, and 3D registration, respectively. To the best of our knowledge, we are the first work to introduce a pre-training framework tailored for correspondence pruning, offering a more universal and scalable solution.
Tangfei Liao, Xiaoqin Zhang 0002, Tao Wang 0052, Min Li 0052, Guobao Xiao, Mang Ye
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 CorrMoE: Mixture of Experts with De-Stylization Learning for Cross-Scene and Cross-Domain Correspondence Pruning
abstract
Establishing reliable correspondences between image pairs is a fundamental task in computer vision, underpinning applications such as 3D reconstruction and visual localization. Although recent methods have made progress in pruning outliers from dense correspondence sets, they often hypothesize consistent visual domains and overlook the challenges posed by diverse scene structures. In this paper, we propose CorrMoE, a novel correspondence pruning framework that enhances robustness under cross-domain and cross-scene variations. To address domain shift, we introduce a De-stylization Dual Branch, performing style mixing on both implicit and explicit graph features to mitigate the adverse influence of domain-specific representations. For scene diversity, we design a Bi-Fusion Mixture of Experts module that adaptively integrates multi-perspective features through linear-complexity attention and dynamic expert routing. Extensive experiments on benchmark datasets demonstrate that CorrMoE achieves superior accuracy and generalization compared to state-of-the-art methods. The code and pre-trained models are available at https://github.com/peiwenxia/CorrMoE.
Peiwen Xia, Tangfei Liao, Danhuai Zhao, Jianjun Ke, Kaihao Zhang, Tong Lu 0002, Tao Wang 0052
ECAI2
2025 PTH-Net: Dynamic Facial Expression Recognition Without Face Detection and Alignment
abstract
Pyramid Temporal Hierarchy Network (PTH-Net) is a new paradigm for dynamic facial expression recognition, applied directly to raw videos, without face detection and alignment. Unlike the traditional paradigm, which focus only on facial areas and often overlooks valuable information like body movements, PTH-Net preserves more critical information. It does this by distinguishing between backgrounds and human bodies at the feature level, offering greater flexibility as an end-to-end network. Specifically, PTH-Net utilizes a pre-trained backbone to extract multiple general features of video understanding at various temporal frequencies, forming a temporal feature pyramid. It then further expands this temporal hierarchy through differentiated parameter sharing and downsampling, ultimately refining emotional information under the supervision of expression temporal-frequency invariance. Additionally, PTH-Net features an efficient Scalable Semantic Distinction layer that enhances feature discrimination, helping to better identify target expressions versus non-target ones in the video. Finally, extensive experiments demonstrate that PTH-Net performs excellently in eight challenging benchmarks, with lower computational costs compared to previous methods. The source code is available at https://github.com/lm495455/PTH-Net.
Min Li 0052, Xiaoqin Zhang 0002, Tangfei Liao, Guobao Xiao
IEEE Trans. Image Process.3
2024 VSFormer: Visual-Spatial Fusion Transformer for Correspondence Pruning
abstract
Correspondence pruning aims to find correct matches (inliers) from an initial set of putative correspondences, which is a fundamental task for many applications. The process of finding is challenging, given the varying inlier ratios between scenes/image pairs due to significant visual differences. However, the performance of the existing methods is usually limited by the problem of lacking visual cues (e.g., texture, illumination, structure) of scenes. In this paper, we propose a Visual-Spatial Fusion Transformer (VSFormer) to identify inliers and recover camera poses accurately. Firstly, we obtain highly abstract visual cues of a scene with the cross attention between local features of two-view images. Then, we model these visual cues and correspondences by a joint visual-spatial fusion module, simultaneously embedding visual cues into correspondences for pruning. Additionally, to mine the consistency of correspondences, we also design a novel module that combines the KNN-based graph and the transformer, effectively capturing both local and global contexts. Extensive experiments have demonstrated that the proposed VSFormer outperforms state-of-the-art methods on outdoor and indoor benchmarks. Our code is provided at the following repository: https://github.com/sugar-fly/VSFormer.
Tangfei Liao, Xiaoqin Zhang 0002, Li Zhao 0005, Tao Wang 0047, Guobao Xiao
AAAI1
2024 CorrAdaptor: Adaptive Local Context Learning for Correspondence Pruning
abstract
In the fields of computer vision and robotics, accurate pixel-level correspondences are essential for enabling advanced tasks such as structure-from-motion and simultaneous localization and mapping. Recent correspondence pruning methods usually focus on learning local consistency through k-nearest neighbors, which makes it difficult to capture robust context for each correspondence. We propose CorrAdaptor, a novel architecture that introduces a dual-branch structure capable of adaptively adjusting local contexts through both explicit and implicit local graph learning. Specifically, the explicit branch uses KNN-based graphs tailored for initial neighborhood identification, while the implicit branch leverages a learnable matrix to softly assign neighbors and adaptively expand the local context scope, significantly enhancing the model’s robustness and adaptability to complex image variations. Moreover, we design a motion injection module to integrate motion consistency into the network to suppress the impact of outliers and refine local context learning, resulting in substantial performance improvements. The experimental results on extensive correspondence-based tasks indicate that our CorrAdaptor achieves state-of-the-art performance both qualitatively and quantitatively.
Yuping He, Tangfei Liao, Xiaoqiu Xu, Tao Wang 0052, Tong Lu 0002
ECAI4
2024 Dual-STI: Dual-path spatial-temporal interaction learning for dynamic facial expression recognition
Min Li 0052, Xiaoqin Zhang 0002, Chenxiang Fan, Tangfei Liao, Guobao Xiao
Inf. Sci.4
2024 A novel non-pretrained deep supervision network for polyp segmentation
Zhenni Yu, Li Zhao 0005, Tangfei Liao, Xiaoqin Zhang 0002, Geng Chen 0001, Guobao Xiao
Pattern Recognit.3
2024 Multi-Prior Driven Network for RGB-D Salient Object Detection
abstract
Most existing RGB-D salient object detection (SOD) methods rely on high-quality depth images. However, their performance is limited when processing low-quality depth maps. This paper exploits more complementary image priors to guide the model to learn on variable depth maps, and a novel multi-prior driven network called MPDNet is proposed for RGB-D SOD. MPDNet utilizes four processing pipelines to process RGB images and other priors, which include an RGB image processing pipeline, a depth map processing pipeline, a fine-grained and gradient prior processing pipeline, and an edge learning pipeline. Specifically, fine-grained and gradient priors are input to the same processing pipeline. For the depth maps, fine-grained and gradient priors, a prior channel attention module utilizes the channel attention mechanism to filter noises and highlights the salient cues. The RGB image processing pipeline uses a multi-feature progressive enhancement module to fuse and enhance features from depth maps. And a multi-feature prediction decoder decodes initial salient masks. In the edge learning pipeline, edge prior serves as an edge label and is captured by an edge capture module. Finally, the clear salient masks are obtained by fusing the salient information from the four pipelines. The experimental results on six benchmarks indicate that the proposed method outperforms thirteen state-of-the-art methods in six evaluation metrics.
Xiaoqin Zhang 0002, Yuewang Xu, Tao Wang 0052, Tangfei Liao
IEEE Trans. Circuits Syst. Video Technol.4
2023 SGA-Net: A Sparse Graph Attention Network for Two-View Correspondence Learning
abstract
Establishing reliable correspondences between two images is a fundamental and important task in computer vision. This paper proposes a novel network called Sparse Graph Attention Network (SGA-Net), to capture rich contextual information of sparse graphs for feature matching task. Specifically, a graph attention block is proposed to enhance the representational ability of graph-structured features. The proposed block introduces a novel normalization technique for graph-structured features to embed global information into each edge feature, and it adopts the squeeze-and-excitation mechanism to capture graph-wise contextual information. Meanwhile, to further obtain interesting structural information of sparse graphs, a novel sparse graph transformer is developed based on multi-headed self-attention mechanism, while maintaining permutation-equivariance. Additionally, considering that the graph contexts in shallow layers are not fully exploited, a simple graph-context fusion block is introduced to adaptively capture topological information from different layers by implicitly modeling the interdependence between these graph contexts. The proposed SGA-Net can search dependable candidates among the putative correspondences and simultaneously estimate accurate camera poses for two-view geometry estimation. Extensive experiments on outlier removal and camera pose estimation tasks have demonstrated that the proposed SGA-Net outperforms state-of-the-art methods on both outdoor and indoor benchmarks (i.e., YFCC100M and SUN3D). SGA-Net achieves a mAP5° of 58.88% without RANSAC on the outdoor dataset, and it achieves a precision increase of 13.45% and 7.34% compared with the state-of-the-art result on outdoor and indoor datasets, respectively.
Tangfei Liao, Xiaoqin Zhang 0002, Yuewang Xu, Ziwei Shi, Guobao Xiao
IEEE Trans. Circuits Syst. Video Technol.1