Guobao Xiao

dblp:160/7340 · DBLP profile ↗
← Back
89ranked-venue papers
15as first author
66since 2021 · last 2026
0000-0003-2928-8100ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 45 · 4 first-author · 33 since 2021Artificial intelligence and machine learning · 44 · 12 first-author · 29 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Scalable and Generalizable Correspondence Pruning via Geometry-Consistent Pre-Training
abstract
Two-view correspondence pruning aims to identify reliable correspondences for camera pose estimation, serving as a fundamental step in many 3D vision tasks. Existing methods rely on geometric consistency to seek true correspondences (inliers) from numerous false correspondences (outliers). In this learning paradigm, outliers severely affect the representation learning of inliers, resulting in models that are neither robust nor generalizable. To address this issue, we propose a geometry-consistent pre-training paradigm that sculpts scalable and generalizable representations free from outlier interference. The paradigm features two appealing properties. 1) Implementation of geometry-consistent pre-training. We introduce masked inlier reconstruction as a pretext task and develop a simple yet effective pre-training framework based on a masked autoencoder. Specifically, due to the irregular and unordered nature of correspondences, which lack explicit positional information, we adopt a dual-branch structure that separately reconstructs the keypoints of two images. This enables indirect reconstruction of 4D correspondences, where keypoints from the paired image provide positional prompts. 2) Unified correspondence encoder. We propose a simple dual-stream encoder with built-in consensus interaction, providing a unified, extensible architecture that enhances representation learning. Extensive experiments demonstrate that our method, GeneralPruner, consistently outperforms state-of-the-art approaches in terms of robustness and generalization across various downstream tasks. Specifically, our method achieves 10.76%, 11.84%, and 8.65% performance gains in camera pose estimation, visual localization, and 3D registration, respectively. To the best of our knowledge, we are the first work to introduce a pre-training framework tailored for correspondence pruning, offering a more universal and scalable solution.
Tangfei Liao, Xiaoqin Zhang 0002, Tao Wang 0052, Min Li 0052, Guobao Xiao, Mang Ye
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 LLHA-Net: A hierarchical attention network for two-view correspondence learning
Shuyuan Lin, Xiao Chen 0021, Guobao Xiao, Feiran Huang
Pattern Recognit.5
2026 HSENet: Hierarchical semantic-enriched network for multi-modal image fusion
Rui Ming, Songlin Du, Lianghua He, Guobao Xiao
Pattern Recognit.6
2026 PTCNet: Pure transformer network for two-view correspondence pruning
Xiangyang Miao, Shunxing Chen, Shiping Wang, Junwen Guo, Fengying Wu, Guobao Xiao, Zongjuan Li
Pattern Recognit.6
2026 DS2D: Decoupling feature guidance with state-space diffusion for infrared-visible image fusion
Zongjuan Li, Shiping Wang, Fengying Wu, Guobao Xiao
Pattern Recognit.7
2026 SceneGlue: Scene-Aware Transformer for Feature Matching Without Scene-Level Annotation
abstract
Local feature matching plays a critical role in understanding the correspondence between cross-view images. However, traditional methods are constrained by the inherent local nature of feature descriptors, limiting their ability to capture non-local scene information that is essential for accurate cross-view correspondence. In this paper, we introduce SceneGlue, a scene-aware feature matching framework designed to overcome these limitations. SceneGlue leverages a hybridizable matching paradigm that integrates implicit parallel attention and explicit cross-view visibility estimation. The parallel attention mechanism simultaneously exchanges information among local descriptors within and across images, enhancing the scene’s global context. To further enrich the scene awareness, we propose the Visibility Transformer, which explicitly categorizes features into visible and invisible regions, providing an understanding of cross-view scene visibility. By combining explicit and implicit scene-level awareness, SceneGlue effectively compensates for the local descriptor constraints. Notably, SceneGlue is trained using only local feature matches, without requiring scene-level groundtruth annotations. This scene-aware approach not only improves accuracy and robustness but also enhances interpretability compared to traditional methods. Extensive experiments on applications such as homography estimation, pose estimation, image matching, and visual localization validate SceneGlues superior performance. The source code is available at https://github.com/songlindu/ SceneGlue.
Songlin Du, Xiaoyong Lu, Yaping Yan, Guobao Xiao, Xiaobo Lu, Takeshi Ikenaga
IEEE Trans. Circuits Syst. Video Technol.4
2026 Spatially Aware Adaptive Diffusion: Unifying Low-Resolution Image Fusion and Super-Resolution
abstract
Low-resolution visible-infrared image fusion and super-resolution (LRVIF) are critical for enhancing image quality in low-resolution scenarios, yet limited information in the input images often constrains performance. To address these challenges, we propose SaDiff, a spatially-aware adaptive diffusion model that introduces diffusion processes into LRVIF for the first time, representing a major breakthrough in the field. Leveraging the generative capabilities of diffusion models, our approach unifies and enhances image fusion and super-resolution within a cohesive framework. A key component of SaDiff is the Spatial Residual Adaptation Block, which extends the diffusion process by dynamically adapting feature representations to spatial variations in the local regions of the input images. This module maximally preserves crucial information from the input images, such as texture details and contrast, while effectively suppressing noise, ensuring robust and context-aware feature refinement. Then we further propose Direct Diffusion Synthesis, a novel mechanism that utilizes noise predictions during diffusion to generate fused images, enabling joint training of the fusion and super-resolution networks. Additionally, a Cross-Feature Fusion Module integrates texture and contrast details, producing super-resolution fused images with improved clarity and structural integrity. Extensive experiments show that SaDiff achieves state-of-the-art performance, offering a robust and unified solution to infrared-visible image fusion and super-resolution. The code for the proposed method will be made available at https://github.com/guobaoxiao/SaDiff.
Jiajia Fu, Zhenni Yu, Haosheng Chen 0001, Songlin Du, Changcai Yang, Lianghua He, Guobao Xiao
IEEE Trans. Circuits Syst. Video Technol.7
2026 Revisiting Semantic Correspondence: When Feature Aggregation Hurts Structural Integrity
abstract
Semantic correspondence seeks to establish matches between different instances of the same category. A common paradigm for this task leverages high-quality features from stable diffusion (SD) and DINOv2. However, we identify a widely overlooked yet critical issue: common feature aggregation disrupts the structural integrity of SD features, degrading semantic matching performance. We revisit and analyze this phenomenon and propose structure-aware aggregation (SAA) for SD features as a direct replacement for common feature aggregation methods. SAA uses filtering to decompose SD features into fine texture details and coarse contour structures. It aggregates only the texture components while preserving the contours. This divide-and-conquer mechanism enables SAA to significantly enhance the performance of state-of-the-art semantic correspondence models without increasing trainable parameters or computational overhead. Extensive qualitative and quantitative experiments confirm our analysis and validate the effectiveness of SAA. Moreover, SAA generalizes well to geometric, cross-species, and cross-family semantic correspondence tasks. Code is available at https://github.com/wzhlearning/SAA.
Zenghui Wang 0009, Songlin Du, Xiaobo Lu, Guobao Xiao
IEEE Trans. Image Process.5
2026 Simpler is Better: Feature Guard and Interaction for Semantic Correspondence
abstract
Semantic correspondence establishes keypoint correspondences between different instances of the same category. Fusing texture and semantic features from vision foundation models like stable diffusion (SD) and DINO significantly improves matching performance. However, we found an unnoticed yet essential problem: current feature fusion enhances the edge and semantic information in SD features with fine textures and DINOv2 features with fine semantics, but it destroys the semantic and structural information in SD features with weak and coarse semantics. We propose guard features (GuFT), a simple yet efficient method, to prevent feature degradation. Moreover, matching methods designed for traditional deep neural networks can be simplified based on two key insights: 1) vision foundation models provide rich visual knowledge; and 2) GuFT yields high-quality feature descriptors. We propose a bottleneck-style non-shared aggregation and backward interaction (NABI) module to efficiently capture intra- and inter-feature relationships, instead of common self- and cross-attention. The resulting framework, SimBetter, embodies a "simpler is better" design philosophy. It achieves state-of-the-art results with lower computation on SPair-71k, AP-10K, and PF-PASCAL, excelling in geometry-aware, cross-species, cross-family, and cross-dataset tasks. SimBetter also shows excellent potential in the applications of image-video semantic correspondence and sticker editing. Code is available at https://github.com/wzhlearning/SimBetter.
Zenghui Wang 0009, Songlin Du, Guobao Xiao
IEEE Trans. Image Process.3
2026 Trend and Order Features for Semi-Supervised Time-Series Classification via Multitask Learning
abstract
Multitask learning with a pretext task has excelled in time-series classification task lacking labeled data. The key to multitask learning is to build a pretext task and learn the most representative feature from the raw time series. In this article, we propose trend and order features for semi-supervised time-series classification via multitask learning (TOFL). Specifically, we propose a simple but effective pretext task-self-sequence order prediction (SOP)-to discover the order relation. In addition, we design a gradual trend fusion (GTF) block concatenating different trend features as the shared backbone network basis element to obtain high-quality trend features for the SOP task. Finally, we not only theoretically analyze the uniform stability and generalization error of TOFL but also evaluate the results compared with state-of-the-art (SOTA) supervised and semi-supervised methods on the 128 UCR datasets and three real-world datasets. TOFL demonstrates a high level of competitiveness and, in most cases, closely matches or even surpasses SOTA methods in terms of accuracy. The source code and data of TOFL are freely available at: https://github.com/Sample-design-alt/TOFL.
Xuanhui Yan, Guobao Xiao, Shilin Zhou 0001
IEEE Trans. Neural Networks Learn. Syst.3
2025 SMR-Net: Semantic-Guided Mutually Reinforcing Network for Cross-Modal Image Fusion and Salient Object Detection
abstract
This paper introduces a lightweight Semantic-guided Mutually Reinforcing network (SMR-Net) for the tasks of cross-modal image fusion and salient object detection (SOD). The core concept of SMR-Net is to leverage semantics for directing the mutual reinforcing between image fusion and SOD. Specifically, a Progressive Cross-modal Interaction (PCI) image fusion subnetwork is designed to exploit local interactions via convolution operations and extend to global interactions utilizing spatial and channel attention mechanisms. Subsequently, a cross-modal Bit-Plane Slicing-based SOD subnetwork (BPS) is developed by incorporating the fused image as a third modality. This component employs bit-plane slicing and the deformable convolution technique to effectively extract irregular semantic information embedded in fusion features. The refined semantic information then guides the feature extraction process of the source modalities in a reweighted fashion. By cascading these two subnetworks, BPS leverages final semantic results to direct PCI towards focusing more on semantic information. Ultimately, through this semantic-guided mutual enhancement process, SMR-Net excels in both producing high-quality fused images and achieving effective salient object detection. Our extensive experiments on image fusion and SOD tasks convincingly demonstrate the superiority of our network over existing state-of-the-art alternatives without introducing noticeable computational costs. Compared to nearest competitors, our method demonstrates a stronger generalization ability with 26% fewer parameters.
Guobao Xiao, Zebin Lin, Rui Ming
AAAI1
2025 SAM-TTT: Segment Anything Model via Reverse Parameter Configuration and Test-Time Training for Camouflaged Object Detection
abstract
This paper introduces a new Segment Anything Model (SAM) that leverages reverse parameter configuration and test-time training to enhance its performance on Camouflaged Object Detection (COD), named SAM-TTT. While most existing SAM-based COD models primarily focus on enhancing SAM by extracting favorable features and amplifying its advantageous parameters, a crucial gap is identified: insufficient attention to adverse parameters that impair SAM's semantic understanding in downstream tasks. To tackle this issue, the Reverse SAM Parameter Configuration Module is proposed to effectively mitigate the influence of adverse parameters in a train-free manner by configuring SAM's parameters. Building on this foundation, the T-Visioner Module is unveiled to strengthen advantageous parameters by integrating Test-Time Training layers, originally developed for language tasks, into vision tasks. Test-Time Training layers represent a new class of sequence modeling layers characterized by linear complexity and an expressive hidden state. By integrating two modules, SAM-TTT simultaneously suppresses adverse parameters while reinforcing advantageous ones, significantly improving SAM's semantic understanding in COD task. Our experimental results on various COD benchmarks demonstrate that the proposed approach achieves state-of-the-art performance, setting a new benchmark in the field. The code will be available at https://github.com/guobaoxiao/SAM-TTT.
Zhenni Yu, Li Zhao 0005, Guobao Xiao, Xiaoqin Zhang 0002
ACM Multimedia3
2025 COMPrompter: reconceptualized segment anything model with multiprompt network for camouflaged object detection
Xiaoqin Zhang 0002, Zhenni Yu, Li Zhao 0005, Deng-Ping Fan, Guobao Xiao
Sci. China Inf. Sci.5
2025 SSDFusion: A scene-semantic decomposition approach for visible and infrared image fusion
Rui Ming, Yixian Xiao, Guolong Zheng, Guobao Xiao
Pattern Recognit.5
2025 CLG-Net: Rethinking Local and Global Perception in Lightweight Two-View Correspondence Learning
abstract
Correspondence learning aims to identify correct correspondences from the initial correspondence set and estimate camera pose between a pair of images. At present, Transformer-based methods have make notable progress in the correspondence learning task due to their powerful non-local information modeling capabilities. However, these methods seem to neglect local structures during feature aggregation from all query-key pairs, resulting in computational inefficiency and inaccurate correspondence identification. To address this issue, we propose a novel Context-aware Local and Global interaction Transformer (CLGFormer), a lightweight Transformer-based module with dual-branches that address local and global context perception in attention mechanisms. CLGFormer explores the relationship between neighborhood consistency observed in correspondences and context-aware weights appearing in vanilla attention and introduces an attention-style convolution operator. On top of that, CLGFormer also incorporates a cascaded operation that splits full features into multiple subsets and then feeds to the attention heads, which not only reduces computational costs but also enhances attention diversity. At last, we also introduce a feature recombination operate with high jointness and a lightweight channel attention module. The culmination of our efforts is the Context-aware Local and Global interaction Network (CLG-Net), which accurately estimates camera pose and identifies inliers. Through rigorous experiments, we demonstrate that our CLG-Net network outperforms existing state-of-the-art methods while exhibiting robust generalization capabilities across various scenarios. Code will be available athttps://github.com/guobaoxiao/CLG.
Minjun Shen, Guobao Xiao, Changcai Yang, Junwen Guo, Lei Zhu 0002
IEEE Trans. Circuits Syst. Video Technol.2
2025 Tex2Sem: Learning From Textures to Semantics for Robust Semantic Correspondence
abstract
Recent advances in semantic correspondence have witnessed growing interest in vision foundation models, particularly stable diffusion (SD) and self-distillation with no labels (DINO). However, existing methods underutilize the matching potential of SD and DINOv2 features and show similar background interference patterns. They lack texture-to-semantic learning and intra- and inter-image feature interaction. This study proposes Tex2Sem, a framework learning from textures to semantics, to address the two problems. For the first problem, we propose a texture-to-semantic learning paradigm that achieves texture-semantic trade-offs on features and correlation maps, including progressive fusion and correlation map computation. The SD and DINOv2 features are aggregated from textures to semantics to produce multi-stage progressive fusion features. The resulting multi-stage progressive fusion correlation maps improve semantic correspondence significantly. For the second problem, MamFormer, a hybrid architecture of Mamba-2 and Transformer, is proposed to improve intra- and inter-image feature aggregation and interaction. It enhances foreground focus and background suppression. Given the high computational cost of processing all-stage progressive fusion features, the terminal-stage aggregation and interaction mechanism (TAIM) is proposed to enhance feature learning efficiency. Experiments demonstrate that Tex2Sem achieves state-of-the-art performance on SPair-71k, AP-10K, and PF-PASCAL. Furthermore, Tex2Sem shows remarkable generalization capabilities in cross-species, cross-family, and cross-dataset matching and demonstrates the potential for applications in video swap and human pose estimation. Code is available at https://github.com/wzhlearning/Tex2Sem.
Zenghui Wang 0009, Songlin Du, Yaping Yan, Guobao Xiao, Xiaobo Lu
IEEE Trans. Circuits Syst. Video Technol.4
2025 PTH-Net: Dynamic Facial Expression Recognition Without Face Detection and Alignment
abstract
Pyramid Temporal Hierarchy Network (PTH-Net) is a new paradigm for dynamic facial expression recognition, applied directly to raw videos, without face detection and alignment. Unlike the traditional paradigm, which focus only on facial areas and often overlooks valuable information like body movements, PTH-Net preserves more critical information. It does this by distinguishing between backgrounds and human bodies at the feature level, offering greater flexibility as an end-to-end network. Specifically, PTH-Net utilizes a pre-trained backbone to extract multiple general features of video understanding at various temporal frequencies, forming a temporal feature pyramid. It then further expands this temporal hierarchy through differentiated parameter sharing and downsampling, ultimately refining emotional information under the supervision of expression temporal-frequency invariance. Additionally, PTH-Net features an efficient Scalable Semantic Distinction layer that enhances feature discrimination, helping to better identify target expressions versus non-target ones in the video. Finally, extensive experiments demonstrate that PTH-Net performs excellently in eight challenging benchmarks, with lower computational costs compared to previous methods. The source code is available at https://github.com/lm495455/PTH-Net.
Min Li 0052, Xiaoqin Zhang 0002, Tangfei Liao, Guobao Xiao
IEEE Trans. Image Process.5
2025 MambaMatch: Establishing Reliable Correspondences via Multi-Scale State Space Model
abstract
Correspondence pruning aims to identify inliers from correspondences severely disturbed by outliers. Although Transformers and graph neural networks have shown impressive results in this field, they are either limited by a narrow receptive field or encounter quadratic computational complexity. To tackle this challenge, this work pioneers the integration of state space model into correspondence pruning task, proposing a Mamba-based framework named MambaMatch. Specifically, to address the limitations of the Mamba architecture in local consensus modeling, we proposes a multi-scale scanning strategy. It first employs an adaptive clustering algorithm to map origin correspondences into spatially coherent feature clusters, constructing a dual-representation space encompassing both full-scale and clustered-scale features. Bidirectional scan operations are then performed at both scales: 1) full-scale scan preserves global structural context, and 2) clustered-scale scan enhances local consistency. Subsequently, a Multi-Scale Interaction layer is designed to dynamically fuse dual-scale features via a cross-attention mechanism, further integrated with a Gated Feed-Forward Network to significantly improve the network's feature discrimination capability. Extensive experiments validate that MambaMatch surpasses state-of-the-art approaches across multiple benchmarks for two-view geometry estimation. Furthermore, MambaMatch exhibits robust generalization across diverse scenarios, tasks, and feature extractors. The source code is available at: https://github.com/mxyttkx/MambaMatch.
Xiangyang Miao, Shunxing Chen, Shiping Wang, Songlin Du, Lianghua He, Guobao Xiao
IEEE Trans. Image Process.7
2025 Spatial Clustering Guided Two-View Multi-Structural Deterministic Geometric Model Fitting
abstract
This paper addresses the two-view geometric model fitting problem on the multi-structural data with severe outliers for providing reliable and consistent fitting results. The key idea is to adopt spatial clustering to guide deterministically sample minimum subsets. Specifically, we firstly improve the effectiveness of spatial clustering with good neighbors that preserve the consensus of neighborhood elements and neighborhood topology, for enhancing the quality of sampled minimum subsets. Then we further design a multi-scale fusion strategy, which not only boosts more high-quality minimum subsets, but also enables our method to cover all model instances in data. Moreover, we propose a simple and effective model selection algorithm to estimate the parameters of model instances in data. The final proposed method is able to guarantee fast, accurate and stable model fitting results for the multi-structural data. In addition, we construct two large labeled datasets, for homography and fundamental matrix estimation, respectively. Experimental results on real images from six datasets show the significant superiority of the proposed method on both accuracy and speed over several state-of-the-art alternatives. Especially for the MS-COCO-F and YFCC100M-F datasets, the proposed method yields a performance boost of over three times on segmentation error, parameter error and the CPU time.
Guobao Xiao
IEEE Trans. Image Process.1
2025 Topology Learning for Two-View Correspondence Filtering
abstract
In this paper, we propose a novel neural network called Topology Learning Network (TL-Net), that exploits local and global geometric relation by topology graphs to handle the problem of correspondence filtering in complex scenes. Specifically, we first design a Multi-level Topology Encoder (MLTE), which fuses local and global topology graphs by a channel attention, to sufficiently extract the geometric relation among correspondences. MLTE not only includes local topology graphs by gathering the information of relative motion and multi-resolution group convolution, but also includes a global topology graph by aggregating the information of the similarity and the Graph Laplacian. In addition, inspired by Transformer, we design the backbone of TL-Net to generate enriched fdeature maps for correspondence filtering. Meanwhile, by simplifying the global context aggregation, we maintain the lightweight of the backbone, introducing the superiority of Transformer while avoiding extra parameters and calculations. Empirical experiments on several computer vision tasks show that the performance and generalization ability of TL-Net are significantly superior to the state of the art methods. Notably, on relative pose estimation, we achieve 5.63% and 5.03% mAP improvements under an error threshold of$5^{\circ }$outdoors and indoors, respectively.
Ziwei Shi, Xiangyang Miao, Guobao Xiao, Songlin Du, Zheng Wang 0044, Heng Tao Shen
IEEE Trans. Multim.3
2024 Graph Context Transformation Learning for Progressive Correspondence Pruning
abstract
Most of existing correspondence pruning methods only concentrate on gathering the context information as much as possible while neglecting effective ways to utilize such information. In order to tackle this dilemma, in this paper we propose Graph Context Transformation Network (GCT-Net) enhancing context information to conduct consensus guidance for progressive correspondence pruning. Specifically, we design the Graph Context Enhance Transformer which first generates the graph network and then transforms it into multi-branch graph contexts. Moreover, it employs self-attention and cross-attention to magnify characteristics of each graph context for emphasizing the unique as well as shared essential information. To further apply the recalibrated graph contexts to the global domain, we propose the Graph Context Guidance Transformer. This module adopts a confident-based sampling strategy to temporarily screen high-confidence vertices for guiding accurate classification by searching global consensus between screened vertices and remaining ones. The extensive experimental results on outlier removal and relative pose estimation clearly demonstrate the superior performance of GCT-Net compared to state-of-the-art methods across outdoor and indoor datasets.
Junwen Guo, Guobao Xiao, Shiping Wang, Jun Yu 0002
AAAI2
2024 VSFormer: Visual-Spatial Fusion Transformer for Correspondence Pruning
abstract
Correspondence pruning aims to find correct matches (inliers) from an initial set of putative correspondences, which is a fundamental task for many applications. The process of finding is challenging, given the varying inlier ratios between scenes/image pairs due to significant visual differences. However, the performance of the existing methods is usually limited by the problem of lacking visual cues (e.g., texture, illumination, structure) of scenes. In this paper, we propose a Visual-Spatial Fusion Transformer (VSFormer) to identify inliers and recover camera poses accurately. Firstly, we obtain highly abstract visual cues of a scene with the cross attention between local features of two-view images. Then, we model these visual cues and correspondences by a joint visual-spatial fusion module, simultaneously embedding visual cues into correspondences for pruning. Additionally, to mine the consistency of correspondences, we also design a novel module that combines the KNN-based graph and the transformer, effectively capturing both local and global contexts. Extensive experiments have demonstrated that the proposed VSFormer outperforms state-of-the-art methods on outdoor and indoor benchmarks. Our code is provided at the following repository: https://github.com/sugar-fly/VSFormer.
Tangfei Liao, Xiaoqin Zhang 0002, Li Zhao 0005, Tao Wang 0047, Guobao Xiao
AAAI5
2024 BCLNet: Bilateral Consensus Learning for Two-View Correspondence Pruning
abstract
Correspondence pruning aims to establish reliable correspondences between two related images and recover relative camera motion. Existing approaches often employ a progressive strategy to handle the local and global contexts, with a prominent emphasis on transitioning from local to global, resulting in the neglect of interactions between different contexts. To tackle this issue, we propose a parallel context learning strategy that involves acquiring bilateral consensus for the two-view correspondence pruning task. In our approach, we design a distinctive self-attention block to capture global context and parallel process it with the established local context learning module, which enables us to simultaneously capture both local and global consensuses. By combining these local and global consensuses, we derive the required bilateral consensus. We also design a recalibration block, reducing the influence of erroneous consensus information and enhancing the robustness of the model. The culmination of our efforts is the Bilateral Consensus Learning Network (BCLNet), which efficiently estimates camera pose and identifies inliers (true correspondences). Extensive experiments results demonstrate that our network not only surpasses state-of-the-art methods on benchmark datasets but also showcases robust generalization abilities across various feature extraction techniques. Noteworthily, BCLNet obtains significant improvement gains over the second best method on unknown outdoor dataset, and obviously accelerates model training speed.
Xiangyang Miao, Guobao Xiao, Shiping Wang, Jun Yu 0002
AAAI2
2024 Exploring Deeper! Segment Anything Model with Depth Perception for Camouflaged Object Detection
abstract
This paper introduces a new Segment Anything Model with Depth Perception (DSAM) for Camouflaged Object Detection (COD). DSAM exploits the zero-shot capability of SAM to realize precise segmentation in the RGB-D domain. It consists of the Prompt-Deeper Module and the Finer Module. The Prompt-Deeper Module utilizes knowledge distillation and the Bias Correction Module to achieve the interaction between RGB features and depth features, especially using depth features to correct erroneous parts in RGB features. Then, the interacted features are combined with the box prompt in SAM to create a prompt with depth perception. The Finer Module explores the possibility of accurately segmenting highly camouflaged targets from a depth perspective. It uncovers depth cues in areas missed by SAM through mask reversion, self-filtering, and self-attention operations, compensating for its defects in the COD domain. DSAM represents the first step towards the SAM-based RGB-D COD model. It maximizes the utilization of depth features while synergizing with RGB features to achieve multimodal complementarity, thereby overcoming the segmentation limitations of SAM and improving its accuracy in COD. Experimental results on COD benchmarks demonstrate that DSAM achieves excellent segmentation performance and reaches the state-of-the-art (SOTA) on COD benchmarks with less consumption of training resources. The code will be available at https://github.com/guobaoxiao/DSAM.
Zhenni Yu, Xiaoqin Zhang 0002, Li Zhao 0005, Yi Bin, Guobao Xiao
ACM Multimedia5
2024 Progressive correspondence learning by effective multi-channel aggregation
Xin Liu 0091, Shunxing Chen, Guobao Xiao, Changcai Yang, Riqing Chen
Neurocomputing3
2024 Local neighbor propagation on graphs for mismatch removal
Hanlin Guo, Guobao Xiao, Lumei Su, Jiaxing Zhou, Dahan Wang
Inf. Sci.2
2024 Dual-STI: Dual-path spatial-temporal interaction learning for dynamic facial expression recognition
Min Li 0052, Xiaoqin Zhang 0002, Chenxiang Fan, Tangfei Liao, Guobao Xiao
Inf. Sci.5
2024 T-Net++: Effective Permutation-Equivariance Network for Two-View Correspondence Pruning
abstract
We propose a conceptually novel, flexible, and effective framework (named T-Net++) for the task of two-view correspondence pruning. T-Net++ comprises two unique structures: the "-'' structure and the "|'' structure. The "-'' structure utilizes an iterative learning strategy to process correspondences, while the "|'' structure integrates all feature information of the "-'' structure and produces inlier weights. Moreover, within the "|'' structure, we design a new Local-Global Attention Fusion module to fully exploit valuable information obtained from concatenating features through channel-wise and spatial-wise relationships. Furthermore, we develop a Channel-Spatial Squeeze-and-Excitation module, a modified network backbone that enhances the representation ability of important channels and correspondences through the squeeze-and-excitation operation. T-Net++ not only preserves the permutation-equivariance manner for correspondence pruning, but also gathers rich contextual information, thereby enhancing the effectiveness of the network. Experimental results demonstrate that T-Net++ outperforms other state-of-the-art correspondence pruning methods on various benchmarks and excels in two extended tasks.
Guobao Xiao, Xin Liu 0091, Xiaoqin Zhang 0002, Jiayi Ma 0001, Haibin Ling
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Latent Semantic Consensus for Deterministic Geometric Model Fitting
abstract
Estimating reliable geometric model parameters from the data with severe outliers is a fundamental and important task in computer vision. This paper attempts to sample high-quality subsets and select model instances to estimate parameters in the multi-structural data. To address this, we propose an effective method called Latent Semantic Consensus (LSC). The principle of LSC is to preserve the latent semantic consensus in both data points and model hypotheses. Specifically, LSC formulates the model fitting problem into two latent semantic spaces based on data points and model hypotheses, respectively. Then, LSC explores the distributions of points in the two latent semantic spaces, to remove outliers, generate high-quality model hypotheses, and effectively estimate model instances. Finally, LSC is able to provide consistent and reliable solutions within only a few milliseconds for general multi-structural model fitting, due to its deterministic fitting nature and efficiency. Compared with several state-of-the-art model fitting methods, our LSC achieves significant superiority for the performance of both accuracy and speed on synthetic data and real images.
Guobao Xiao, Jun Yu 0002, Jiayi Ma 0001, Deng-Ping Fan, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 SSL-Net: Sparse semantic learning for identifying reliable correspondences
Shunxing Chen, Guobao Xiao, Ziwei Shi, Junwen Guo, Jiayi Ma 0001
Pattern Recognit.2
2024 Semantic attention-based heterogeneous feature aggregation network for image fusion
Zhiqiang Ruan, Guobao Xiao, Jiayi Ma 0001
Pattern Recognit.3
2024 A novel non-pretrained deep supervision network for polyp segmentation
Zhenni Yu, Li Zhao 0005, Tangfei Liao, Xiaoqin Zhang 0002, Geng Chen 0001, Guobao Xiao
Pattern Recognit.6
2024 MSGA-Net: Progressive Feature Matching via Multi-Layer Sparse Graph Attention
abstract
Feature matching is an essential computer vision task that requires the establishment of high-quality correspondences between two images. Constructing sparse dynamic graphs and extracting contextual information by searching for neighbors in feature space is a prevalent strategy in numerous previous works. Nonetheless, these works often neglect the potential connections between dynamic graphs from different layers, leading to underutilization of available information. To tackle this issue, we introduce a Sparse Dynamic Graph Interaction block for feature matching. This innovation facilitates the implicit establishment of dependencies by enabling interaction and aggregation among dynamic graphs across various layers. In addition, we design a novel Multiple Sparse Transformer to enhance the capture of the global context from the sparse graph. This block selectively mines significant global contextual information along spatial and channel dimensions, respectively. Ultimately, we present the Multi-layer Sparse Graph Attention Network (MSGA-Net), a framework designed to predict probabilities of correspondences as inliers and to recover camera poses. Experimental results demonstrate that our proposed MSGA-Net surpasses state-of-the-art methods on challenging indoor and outdoor datasets. Code will be available at https://github.com/gongzhepeng/MSGA-Net.
Zhepeng Gong, Guobao Xiao, Ziwei Shi, Riqing Chen, Jun Yu 0002
IEEE Trans. Circuits Syst. Video Technol.2
2024 Second-Order Proximity Guided Sampling Consensus for Robust Model Fitting
abstract
Robust model fitting plays a critical role in artificial intelligence and computer vision, with its performance primarily depends on the utilization of sampling algorithms. However, existing sampling algorithms become less effective when initial correspondences between two images are corrupted by a large number of outliers, especially in the presence of multi-structure data. In this paper, we propose a novel sampling algorithm (called SPGSC) for robust model fitting, where minimal subsets are sampled with the guidance of the second-order proximity measure, which involves global geometric relationships instead of local consistency relationships. Specifically, we first propose a second-order proximity measure to facilitate graph construction, which helps detect a potential inlier from input data as the first datum (i.e., the seed datum) of a minimal subset. After that, we propose a second-order proximity based initial minimal subset generation strategy, which is able to choose a certain number of minimal subsets by the seed data for efficiently producing significant model hypotheses. Furthermore, to achieve better fitting performance, we propose a maximum spanning tree based refinement (MSTR) strategy, which is used to refine the previous sampled minimal subsets and improve the effectiveness and efficiency of the sampling process. Experimental results on three vision tasks (i.e., two-view based motion segmentation, affine matrix based segmentation, and 3D motion segmentation) show the superiority of the proposed SPGSC in comparison with other state-of-the-art algorithms.
Hanlin Guo, Guobao Xiao, Lumei Su, Tianyou Li, Dahan Wang, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.2
2024 Transformer-Based Multimodal Emotional Perception for Dynamic Facial Expression Recognition in the Wild
abstract
Dynamic expression recognition in the wild is a challenging task due to various obstacles, including low light condition, non-positive face, and face occlusion. Purely vision-based approaches may not suffice to accurately capture the complexity of human emotions. To address this issue, we propose a Transformer-based Multimodal Emotional Perception (T-MEP) framework capable of effectively extracting multimodal information and achieving significant augmentation. Specifically, we design three transformer-based encoders to extract modality-specific features from audio, image, and text sequences, respectively. Each encoder is carefully designed to maximize its adaptation to the corresponding modality. In addition, we design a transformer-based multimodal information fusion module to model cross-modal representation among these modalities. The unique combination of self-attention and cross-attention in this module enhances the robustness of output-integrated features in encoding emotion. By mapping the information from audio and textual features to the latent space of visual features, this module aligns the semantics of the three modalities for cross-modal information augmentation. Finally, we evaluate our method on three popular datasets (MAFW, DFEW, and AFEW) through extensive experiments, which demonstrate its state-of-the-art performance. This research offers a promising direction for future studies to improve emotion recognition accuracy by exploiting the power of multimodal features.
Xiaoqin Zhang 0002, Min Li 0052, Guobao Xiao
IEEE Trans. Circuits Syst. Video Technol.5
2024 DHM-Net: Deep Hypergraph Modeling for Robust Feature Matching
abstract
We present a novel deep hypergraph modeling architecture (called DHM-Net) for feature matching in this paper. Our network focuses on learning reliable correspondences between two sets of initial feature points by establishing a dynamic hypergraph structure that models group-wise relationships and assigns weights to each node. Compared to existing feature matching methods that only consider pair-wise relationships via a simple graph, our dynamic hypergraph is capable of modeling nonlinear higher-order group-wise relationships among correspondences in an interaction capturing and attention representation learning fashion. Specifically, we propose a novel Deep Hypergraph Modeling block, which initializes an overall hypergraph by utilizing neighbor information, and then adopts node-to-hyperedge and hyperedge-to-node strategies to propagate interaction information among correspondences while assigning weights based on hypergraph attention. In addition, we propose a Differentiation Correspondence-Aware Attention mechanism to optimize the hypergraph for promoting representation learning. The proposed mechanism is able to effectively locate the exact position of the object of importance via the correspondence aware encoding and simple feature gating mechanism to distinguish candidates of inliers. In short, we learn such a dynamic hypergraph format that embeds deep group-wise interactions to explicitly infer categories of correspondences. To demonstrate the effectiveness of DHM-Net, we perform extensive experiments on both real-world outdoor and indoor datasets. Particularly, experimental results show that DHM-Net surpasses the state-of-the-art method by a sizable margin. Our approach obtains an 11.65% improvement under error threshold of 5° for relative pose estimation task on YFCC100M dataset. Code will be released at https://github.com/CSX777/DHM-Net.
Shunxing Chen, Guobao Xiao, Junwen Guo, Qiangqiang Wu, Jiayi Ma 0001
IEEE Trans. Image Process.2
2024 Multi-Stage Network With Geometric Semantic Attention for Two-View Correspondence Learning
abstract
The removal of outliers is crucial for establishing correspondence between two images. However, when the proportion of outliers reaches nearly 90%, the task becomes highly challenging. Existing methods face limitations in effectively utilizing geometric transformation consistency (GTC) information and incorporating geometric semantic neighboring information. To address these challenges, we propose a Multi-Stage Geometric Semantic Attention (MSGSA) network. The MSGSA network consists of three key modules: the multi-branch (MB) module, the GTC module, and the geometric semantic attention (GSA) module. The MB module, structured with a multi-branch design, facilitates diverse and robust spatial transformations. The GTC module captures transformation consistency information from the preceding stage. The GSA module categorizes input based on the prior stage's output, enabling efficient extraction of geometric semantic information through a graph-based representation and inter-category information interaction using Transformer. Extensive experiments on the YFCC100M and SUN3D datasets demonstrate that MSGSA outperforms current state-of-the-art methods in outlier removal and camera pose estimation, particularly in scenarios with a high prevalence of outliers. Source code is available at https://github.com/shuyuanlin.
Shuyuan Lin, Xiao Chen 0021, Guobao Xiao, Hanzi Wang, Feiran Huang, Jian Weng 0001
IEEE Trans. Image Process.3
2024 Fusion-Embedding Siamese Network for Light Field Salient Object Detection
abstract
Light field salient object detection (SOD) has shown remarkable success and gained considerable attention from the computer vision community. Existing methods usually employ a single-/two-stream network to detect saliency. However, these methods can only handle up to two different modalities at a time, preventing them from being able to fully explore the rich information in multi-modal light field derived data. To address this, we propose the first joint multi-modal learning framework, called FES-Net, for light field SOD, which can take rich inputs not limited to two modalities. Specifically, we propose an attention-aware adaptation module to first transform the multi-modal inputs for use in our joint learning framework. The transformed inputs are then fed to a Siamese network along with multiple embedded feature fusion modules to extract informative multi-modal features. Finally, we predict saliency maps from the high-level extracted features using a saliency decoder module. Our joint multi-modal learning framework effectively resolves the limitations of existing methods, providing efficient and effective multi-modal learning that can fully explore the valuable information in light field data for accurate saliency detection. Furthermore, we improve the performance by introducing the Transformer as our backbone network. To the best of our knowledge, the improved version of our model, called FES-Trans, is the first attempt to address the challenging light field SOD with the powerful Transformer technique. Extensive experiments on benchmark datasets demonstrate that our models are superior light field SOD approaches and outperform cutting-edge models remarkably.
Geng Chen 0001, Huazhu Fu, Tao Zhou 0002, Guobao Xiao, Keren Fu, Yong Xia 0001, Yanning Zhang 0001
IEEE Trans. Multim.4
2023 Learning matrix factorization with scalable distance metric and regularizer
Shiping Wang, Yunhe Zhang 0001, Xincan Lin, Lichao Su, Guobao Xiao, William Zhu 0001, Yiqing Shi
Neural Networks5
2023 JRA-Net: Joint representation attention network for correspondence learning
Ziwei Shi, Guobao Xiao, Linxin Zheng, Jiayi Ma 0001, Riqing Chen
Pattern Recognit.2
2023 SGA-Net: A Sparse Graph Attention Network for Two-View Correspondence Learning
abstract
Establishing reliable correspondences between two images is a fundamental and important task in computer vision. This paper proposes a novel network called Sparse Graph Attention Network (SGA-Net), to capture rich contextual information of sparse graphs for feature matching task. Specifically, a graph attention block is proposed to enhance the representational ability of graph-structured features. The proposed block introduces a novel normalization technique for graph-structured features to embed global information into each edge feature, and it adopts the squeeze-and-excitation mechanism to capture graph-wise contextual information. Meanwhile, to further obtain interesting structural information of sparse graphs, a novel sparse graph transformer is developed based on multi-headed self-attention mechanism, while maintaining permutation-equivariance. Additionally, considering that the graph contexts in shallow layers are not fully exploited, a simple graph-context fusion block is introduced to adaptively capture topological information from different layers by implicitly modeling the interdependence between these graph contexts. The proposed SGA-Net can search dependable candidates among the putative correspondences and simultaneously estimate accurate camera poses for two-view geometry estimation. Extensive experiments on outlier removal and camera pose estimation tasks have demonstrated that the proposed SGA-Net outperforms state-of-the-art methods on both outdoor and indoor benchmarks (i.e., YFCC100M and SUN3D). SGA-Net achieves a mAP5° of 58.88% without RANSAC on the outdoor dataset, and it achieves a precision increase of 13.45% and 7.34% compared with the state-of-the-art result on outdoor and indoor datasets, respectively.
Tangfei Liao, Xiaoqin Zhang 0002, Yuewang Xu, Ziwei Shi, Guobao Xiao
IEEE Trans. Circuits Syst. Video Technol.5
2023 Learning for Feature Matching via Graph Context Attention
abstract
Establishing reliable correspondences via a deep learning network is an important task in remote sensing, photogrammetry and other computer vision fields. It usually requires mining the relationship among correspondences to aggregate both local and global context. However, current methods are insufficient to effectively acquire context information with high reliability. In this paper, we propose a graph context attention based network (called GCA-Net) to capture and leverage abundant contextual information for feature matching. Specifically, we design a graph context attention block which generates multi-path graph contexts and softly fuses them to combine respective advantages. In addition, for building the graph context containing stronger representation ability and outlier resistance ability, we further design a local-global channel mining block to gather context information by focusing on the significant part as well as to mine dependencies among channels of correspondences in both local and global aspects. The proposed GCA-Net is able to effectively infer the probability of correspondences being inliers or outliers and estimate the essential matrix meanwhile. Extensive experimental results for outlier removal and relative pose estimation demonstrate that GCA-Net outperforms the state-of-art methods on both outdoor and indoor datasets (i.e.,YFCC100M and SUN3D). In addition, experiments extended to remote sensing and point cloud scenes also demonstrate the powerful generalization capability of our network.
Junwen Guo, Guobao Xiao, Shunxing Chen, Shiping Wang, Jiayi Ma 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Density-Guided Incremental Dominant Instance Exploration for Two-View Geometric Model Fitting
abstract
Existing two-view multi-model fitting methods typically follow a two-step manner, i.e., model generation and selection, without considering their interaction. Therefore, in the first step, these methods have to generate a considerable number of instances in order to cover all desired ones, which not only offers no guarantees, but also introduces unnecessary expensive calculations. To address this challenge, this study presents a new algorithm, termed as D2Fitting, that incrementally explores dominant instances. Particularly, rather than viewing model generation and selection as two disjoint parts, D2Fitting fully considers their interaction, and thus performs these two subroutines alternatively under a simple yet effective optimization framework. This design can avoid generating too many redundant instances, thus reducing computational overhead and allowing the proposed D2Fitting being real-time. Meanwhile, we further design a novel density-guided sampler to sample high-quality minimal subsets during the model generation process, so as to fully exploit the spatial distribution of the input data. Also, to mitigate the influence of noise on the subsets sampled by the proposed sampler, a global-residual optimization strategy is investigated for the minimal subset refinement. With all the ingredients mentioned above, the proposed D2Fitting can accurately estimate the number and parameters of geometric models and efficiently segment the input data simultaneously. Extensive experiments on several public datasets demonstrate the significant superiority of D2Fitting over several state-of-the-arts.
Zizhuo Li, Jiayi Ma 0001, Guobao Xiao
IEEE Trans. Image Process.3
2023 PGFNet: Preference-Guided Filtering Network for Two-View Correspondence Learning
abstract
Accurate correspondence selection between two images is of great importance for numerous feature matching based vision tasks. The initial correspondences established by off-the-shelf feature extraction methods usually contain a large number of outliers, and this often leads to the difficulty in accurately and sufficiently capturing contextual information for the correspondence learning task. In this paper, we propose a Preference-Guided Filtering Network (PGFNet) to address this problem. The proposed PGFNet is able to effectively select correct correspondences and simultaneously recover the accurate camera pose of matching images. Specifically, we first design a novel iterative filtering structure to learn the preference scores of correspondences for guiding the correspondence filtering strategy. This structure explicitly alleviates the negative effects of outliers so that our network is able to capture more reliable contextual information encoded by the inliers for network learning. Then, to enhance the reliability of preference scores, we present a simple yet effective Grouped Residual Attention block as our network backbone, by designing a feature grouping strategy, a feature grouping manner, a hierarchical residual-like manner and two grouped attention operations. We evaluate PGFNet by extensive ablation studies and comparative experiments on the tasks of outlier removal and camera pose estimation. The results demonstrate outstanding performance gains over the existing state-of-the-art methods on different challenging scenes. The code is available at https://github.com/guobaoxiao/PGFNet.
Xin Liu 0091, Guobao Xiao, Riqing Chen, Jiayi Ma 0001
IEEE Trans. Image Process.2
2023 Efficient and Differentiable Low-Rank Matrix Completion With Back Propagation
abstract
The low-rank matrix completion has gained rapidly increasing attention from researchers in recent years for its efficient recovery of the matrix in various fields. Numerous studies have exploited the popular neural networks to yield low-rank outputs under the framework of low-rank matrix factorization. However, due to the discontinuity and nonconvexity of rank function, it is difficult to directly optimize the rank function via back propagation. Although a large number of studies have attempted to find relaxations of the rank function, e.g., Schatten-$p$norm, they still face the following issues when updating parameters via back propagation: 1) These methods or surrogate functions are still non-differentiable, bringing obstacles to deriving the gradients of trainable variables. 2) Most of these surrogate functions perform singular value decomposition upon the original matrix at each iteration, which is time-consuming and blocks the propagation of gradients. To address these problems, in this paper, we develop an efficient block-wise model dubbed differentiable low-rank learning (DLRL) framework that adopts back propagation to optimize the Multi-Schatten-$p$norm Surrogate (MSS) function. Distinct from the original optimization of this surrogate function, the proposed framework avoids singular value decomposition to admit the gradient propagation and builds a block-wise learning scheme to minimize values of Schatten-$p$norms. Accordingly, it speeds up the computation and makes all parameters in the proposed framework learnable according to a predefined loss function. Finally, we conduct substantial experiments in terms of image recovery and collaborative filtering. The experimental results verify the superiority of the proposed framework in terms of both runtimes and learning performance compared with other state-of-the-art low-rank optimization methods. Our codes are available athttps://github.com/chenzl23/DLRL.
Zhaoliang Chen, Guobao Xiao, Shiping Wang
IEEE Trans. Multim.3
2023 Correspondence Attention Transformer: A Context-Sensitive Network for Two-View Correspondence Learning
abstract
Seeking reliable correspondences then recovering camera poses from a set of putative correspondences extracted from two images of the same scene is a fundamental problem in computer vision. Recent advances have demonstrated that this problem can be effectively solved by using a deep architecture based on the multi-layer perceptron, where the context normalization is designed to make the network permutation-equivariant and embed global information in the sparse point data. However, the context normalization simply normalizes the feature maps according to their distribution and treats each correspondence equally, leading to difficulties in adequately capturing scene geometry encoded by the inliers, especially in case of severe outliers. To address this issue, this paper designs a context-sensitive network based on the self-attention mechanism, termed as correspondence attention transformer (CAT), to enhance the consistent geometry information of inliers and simultaneously suppress outliers during embedding global information. In particular, we design an attention-style structure to aggregate features from all correspondences, i.e., a spatial attention namely CAT-S, which provides each correspondence with information exchange from others in the putative set. To capture the contextual information in a more comprehensive and robust way, we also introduce a multi-head mechanism in our structure to exploit the geometrical context from different aspects. Moreover, considering the high memory request in spatial attention, we propose a covariance normalized channel attention CAT-C in our framework, which can largely reduce the memory consumption and parameter scale, but it asks for eigenvalue decomposition in each attention block thus resulting in more runtime. Anyway, these two attention mechanisms can realize information exchange from the spatial or channel aspect, which both contribute to constructing the geometrical context between inliers and encourage the network to pay more attention to the feature subset about potential inliers. Extensive experiments have been conducted over both indoor and outdoor datasets on the tasks of camera pose estimation, outlier removal, and image registration, which demonstrate the superiority of our method that realizes a large performance improvement compared with the current state-of-the-art approaches.
Jiayi Ma 0001, Aoxiang Fan, Guobao Xiao, Riqing Chen
IEEE Trans. Multim.4
2022 SRDFM: Siamese Response Deep Factorization Machine to improve anti-cancer drug recommendation
abstract
Predicting the response of cancer patients to a particular treatment is a major goal of modern oncology and an important step toward personalized treatment. In the practical clinics, the clinicians prefer to obtain the most-suited drugs for a particular patient instead of knowing the exact values of drug sensitivity. Instead of predicting the exact value of drug response, we proposed a deep learning-based method, named Siamese Response Deep Factorization Machines (SRDFM) Network, for personalized anti-cancer drug recommendation, which directly ranks the drugs and provides the most effective drugs. A Siamese network (SN), a type of deep learning network that is composed of identical subnetworks that share the same architecture, parameters and weights, was used to measure the relative position (RP) between drugs for each cell line. Through minimizing the difference between the real RP and the predicted RP, an optimal SN model was established to provide the rank for all the candidate drugs. Specifically, the subnetwork in each side of the SN consists of a feature generation level and a predictor construction level. On the feature generation level, both drug property and gene expression, were adopted to build a concatenated feature vector, which even enables the recommendation for newly designed drugs with only chemical property known. Particularly, we developed a response unit here to generate weighted genetic feature vector to simulate the biological interaction mechanism between a specific drug and the genes. For the predictor construction level, we built this level integrating a factorization machine (FM) component with a deep neural network component. The FM can well handle the discrete chemical information and both low-order and high-order feature interactions could be sufficiently learned. Impressively, the SRDFM works well on both single-drug recommendation and synergic drug combination. Experiment result on both single-drug and synergetic drug data sets have shown the efficiency of the SRDFM. The Python implementation for the proposed SRDFM is available at at https://github.com/RanSuLab/SRDFM Contact: [email protected], [email protected] and [email protected].
Ran Su, Guobao Xiao, Leyi Wei
Briefings Bioinform.4
2022 Feature Matching via Motion-Consistency Driven Probabilistic Graphical Model
Jiayi Ma 0001, Aoxiang Fan, Xingyu Jiang 0005, Guobao Xiao
Int. J. Comput. Vis.4
2022 RANet: A relation-aware network for two-view correspondence learning
Guorong Lin, Xin Liu 0091, Fangfang Lin, Guobao Xiao, Jiayi Ma 0001
Neurocomputing4
2022 DA-Net: Dual-attention network for multivariate time series classification
Xuanhui Yan, Shiping Wang, Guobao Xiao
Inf. Sci.4
2022 Learning Two-View Correspondences and Geometry Using Neighbor-Aware Network
abstract
Finding reliable correspondences between two images is a fundamental problem in remote-sensing image registration. In the face of sparse and unordered correspondences, the previous learning-based works often focus on global information and ignore valuable local information. To gather rich local information, we propose a simple and effective approach, the principle of which is to establish a local neighborhood structure for putative correspondences and then extract and aggregate neighborhood context. Specifically, we design a residual block with an innovative normalization operation and then we construct a sub-net, by introducing a competition mechanism in the local neighborhood to enhance the expression ability of features. Finally, we propose a learning-based network, to improve the performance of outlier rejection by extracting neighborhood context and global context. Extensive experiments on relative pose estimation demonstrate that the proposed network surpasses current state-of-the-art approaches on both challenging indoor and outdoor datasets.
Ziwei Shi, Guobao Xiao, Jiayi Ma 0001
IEEE Geosci. Remote. Sens. Lett.3
2022 CSDA-Net: Seeking reliable correspondences by channel-Spatial difference augment network
Shunxing Chen, Linxin Zheng, Guobao Xiao, Jiayi Ma 0001
Pattern Recognit.3
2022 BSCA-Net: Bit Slicing Context Attention network for polyp segmentation
Jichun Wu, Guobao Xiao, Junwen Guo, Geng Chen 0001, Jiayi Ma 0001
Pattern Recognit.3
2022 Triplet Relationship Guided Sampling Consensus for Robust Model Estimation
abstract
RANSAC (RANdom SAmple Consensus) is a widely used robust estimator for estimating a geometric model from feature matches in an image pair. Unfortunately, it becomes less effective when initial input feature matches (i.e., input data) are corrupted by a large number of outliers. In this paper, we propose a new robust estimator (called TRESAC) for model estimation, where data subsets are sampled with the guidance of the triplet relationships, which involve high relevance and local geometric consistency. Each triplet consists of three data, whose relationships satisfy spatial consistency constraints. Therefore, the triplet relationships can be used to effectively initialize and refine the sampling process. With the advantage of the triplet relationships, TRESAC significantly alleviates the influence of outliers and also improves the computational efficiency of model estimation. Experimental results on four challenging datasets show that TRESAC can achieve superior performance on both estimation accuracy and computational efficiency against several other state-of-the-art methods.
Hanlin Guo, Yang Lu 0009, Guobao Xiao, Shuyuan Lin, Hanzi Wang
IEEE Signal Process. Lett.3
2022 Robust Feature Matching for Remote Sensing Image Registration via Guided Hyperplane Fitting
abstract
Feature matching is a fundamental problem in feature-based remote sensing image registration. Due to the ground relief variations and imaging viewpoint changes, remote sensing images often involve local distortions, leading to difficulties in high-accuracy image registration. To address this issue, in this article, we propose a robust feature matching method called First Neighbor Relation Guided (FNRG) for remote sensing image registration via guided hyperplane fitting. The key idea of FNRG is to exploit the first neighbor relation of feature points between two images for seeking consistent seeds in a parameter-free manner. To boost more consistent matches based on the consistent seeds, we formulate the feature matching problem into an affine hyperplane fitting problem by imposing the motion consistency, and then we design a hyperplane updating strategy to refine the fitting model. We also introduce a locality preserving structure-based cost function to promote the matching performance of the hyperplane updating strategy. Our method can mine consistent matches from thousands of putative ones within only a few milliseconds, and it also can handle the data with a large-scale change, rotation, or severe nonrigid deformation. Extensive experiments on the remote sensing image data sets with different types of image transformations show that the proposed method achieves significant superiority over several state-of-the-art methods.
Guobao Xiao, Huan Luo 0001, Leyi Wei, Jiayi Ma 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Co-Clustering on Bipartite Graphs for Robust Model Fitting
abstract
Recently, graph-based methods have been widely applied to model fitting. However, in these methods, association information is invariably lost when data points and model hypotheses are mapped to the graph domain. In this paper, we propose a novel model fitting method based on co-clustering on bipartite graphs (CBG) to estimate multiple model instances in data contaminated with outliers and noise. Model fitting is reformulated as a bipartite graph partition behavior. Specifically, we use a bipartite graph reduction technique to eliminate some insignificant vertices (outliers and invalid model hypotheses), thereby improving the reliability of the constructed bipartite graph and reducing the computational complexity. We then use a co-clustering algorithm to learn a structured optimal bipartite graph with exact connected components for partitioning that can directly estimate the model instances (i.e., post-processing steps are not required). The proposed method fully utilizes the duality of data points and model hypotheses on bipartite graphs, leading to superior fitting performance. Exhaustive experiments show that the proposed CBG method performs favorably when compared with several state-of-the-art fitting methods.
Shuyuan Lin, Hailing Luo, Yan Yan 0001, Guobao Xiao, Hanzi Wang
IEEE Trans. Image Process.4
2022 MSA-Net: Establishing Reliable Correspondences by Multiscale Attention Network
abstract
In this paper, we propose a novel multi-scale attention based network (called MSA-Net) for feature matching problems. Current deep networks based feature matching methods suffer from limited effectiveness and robustness when applied to different scenarios, due to random distributions of outliers and insufficient information learning. To address this issue, we propose a multi-scale attention block to enhance the robustness to outliers, for improving the representational ability of the feature map. In addition, we also design a novel context channel refine block and a context spatial refine block to mine the information context with less parameters along channel and spatial dimensions, respectively. The proposed MSA-Net is able to effectively infer the probability of correspondences being inliers with less parameters. Extensive experiments on outlier removal and relative pose estimation have shown the performance improvements of our network over current state-of-the-art methods with less parameters on both outdoor and indoor datasets. Notably, our proposed network achieves an 11.7% improvement at error threshold 5° without RANSAC than the state-of-the-art method on relative pose estimation task when trained on YFCC100M dataset.
Linxin Zheng, Guobao Xiao, Ziwei Shi, Shiping Wang, Jiayi Ma 0001
IEEE Trans. Image Process.2
2021 T-Net: Effective Permutation-Equivariant Network for Two-View Correspondence Learning
abstract
We develop a conceptually simple, flexible, and effective framework (named T-Net) for two-view correspondence learning. Given a set of putative correspondences, we reject outliers and regress the relative pose encoded by the essential matrix, by an end-to-end framework, which is consisted of two novel structures: "−" structure and "|" structure. " − " structure adopts an iterative strategy to learn correspondence features. "|" structure integrates all the features of the iterations and outputs the correspondence weight. In addition, we introduce Permutation-Equivariant Context Squeeze-and-Excitation module, an adapted version of SE module, to process sparse correspondences in a permutation-equivariant way and capture both global and channel-wise contextual information. Extensive experiments on outdoor and indoor scenes show that the proposed T-Net achieves state-of-the-art performance. On outdoor scenes (YFCC100M dataset), T-Net achieves an mAP of 52.28%, a 34.22% precision increase from the best-published result (38.95%). On indoor scenes (SUN3D dataset), T-Net (19.71%) obtains a 21.82% precision increase from the best-published result (16.18%). Source code: https://github.com/x-gb/T-Net.
Guobao Xiao, Linxin Zheng, Jiayi Ma 0001
ICCV2
2021 iDNA-ABT: advanced deep learning model for detecting DNA methylation with adaptive features and transductive information maximization
abstract
MOTIVATION: DNA methylation plays an important role in epigenetic modification, the occurrence, and the development of diseases. Therefore, identification of DNA methylation sites is critical for better understanding and revealing their functional mechanisms. To date, several machine learning and deep learning methods have been developed for the prediction of different DNA methylation types. However, they still highly rely on manual features, which can largely limit the high-latent information extraction. Moreover, most of them are designed for one specific DNA methylation type, and therefore cannot predict multiple methylation sites in multiple species simultaneously. In this study, we propose iDNA-ABT, an advanced deep learning model that utilizes adaptive embedding based on Bidirectional Encoder Representations from Transformers (BERT) together with transductive information maximization (TIM). RESULTS: Benchmark results show that our proposed iDNA-ABT can automatically and adaptively learn the distinguishing features of biological sequences from multiple species, and thus perform significantly better than the state-of-the-art methods in predicting three different DNA methylation types. In addition, TIM loss is proven to be effective in dichotomous tasks via the comparison experiment. Furthermore, we verify that our features have strong adaptability and robustness to different species through comparison of adaptive embedding and six handcrafted feature encodings. Importantly, our model shows great generalization ability in different species, demonstrating that our model can adaptively capture the cross-species differences and improve the predictive performance. For the convenient use of our method, we further established an online webserver as the implementation of the proposed iDNA-ABT. AVAILABILITY AND IMPLEMENTATION: Our proposed iDNA-ABT and data are freely accessible via http://server.wei-group.net/iDNA_ABT and our source codes are available for downloading in the GitHub repository (https://github.com/YUYING07/iDNA_ABT). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Junru Jin, Guobao Xiao, Li-Zhen Cui 0001, Rao Zeng, Leyi Wei
Bioinform.4
2021 Segmentation by Continuous Latent Semantic Analysis for Multi-structure Model Fitting
Guobao Xiao, Hanzi Wang, Jiayi Ma 0001, David Suter
Int. J. Comput. Vis.1
2021 SCSA-Net: Presentation of two-view reliable correspondence learning via spatial-channel self-attention
Xin Liu 0091, Guobao Xiao, Luanyuan Dai, Changcai Yang, Riqing Chen
Neurocomputing2
2021 TSSN-Net: Two-step Sparse Switchable Normalization for learning correspondences with heavy outliers
Guobao Xiao, Shiping Wang
Neurocomputing2
2021 Seeded random walk for multi-view semi-supervised classification
Shiping Wang, Zhewen Wang, Kart-Leong Lim, Guobao Xiao, Wenzhong Guo
Knowl. Based Syst.4
2021 Mining consistent correspondences using co-occurrence statistics
Guobao Xiao, Shiping Wang, Han Wang 0001, Jiayi Ma 0001
Pattern Recognit.1
2021 Point2CN: Progressive two-view correspondence learning via information fusion
Xin Liu 0091, Guobao Xiao, Riqing Chen
Signal Process.2
2021 MANet: Multi-Scale Attention Network for Correspondence Learning
abstract
Establishing reliable correspondences from a putative correspondence set is a challenging task. Most of state-of-the-art methods utilize the local context and global context to address the task. However, the local and global context often contains large number of outliers, which have a negative impact on capturing scene geometry. In this paper, we propose a Multi-scale Attention Network (called MANet), which introduces the attention mechanism for feature matching, to improve the ability of capturing scene geometry. Specifically, we first fuse the features of low and high levels by an multi-scale strategy network to enhance the representative ability of features. Then, we propose an attentive PointCN block and an attentive pooling layer, to discriminatively capture global context and local context information, respectively. We demonstrate through extensive experiments on both indoor and outdoor datasets that MANet provides a significant improvement in the performance of the two-view geometry and correspondences accuracy compared to the state-of-the-art methods.
Yukai Chen, Linxin Zheng, Xin Liu 0091, Guobao Xiao
IEEE Signal Process. Lett.4
2020 Meta-GDBP: a high-level stacked regression model to improve anticancer drug response prediction
abstract
Anticancer drug response prediction plays an important role in personalized medicine. In particular, precisely predicting drug response in specific cancer types and patients is still a challenge problem. Here we propose Meta-GDBP, a novel anticancer drug-response model, which involves two levels. At the first level of Meta-GDBP, we build four optimized base models (BMs) using genetic information, chemical properties and biological context with an ensemble optimization strategy, while at the second level, we construct a weighted model to integrate the four BMs. Notably, the weights of the models are learned upstream, thus the parameter cost is significantly reduced compared to previous methods. We evaluate the Meta-GDBP on Genomics of Drug Sensitivity in Cancer (GDSC) and the Cancer Cell Line Encyclopedia (CCLE) data sets. Benchmarking results demonstrate that compared to other methods, the Meta-GDBP achieves a much higher correlation between the predicted and the observed responses for almost all the drugs. Moreover, we apply the Meta-GDBP to predict the GDSC-missing drug response and use the CCLE-known data to validate the performance. The results show quite a similar tendency between these two response sets. Particularly, we here for the first time introduce a biological context-based frequency matrix (BCFM) to associate the biological context with the drug response. It is encouraging that the proposed BCFM is biologically meaningful and consistent with the reported biological mechanism, further demonstrating its efficacy for predicting drug response. The R implementation for the proposed Meta-GDBP is available at https://github.com/RanSuLab/Meta-GDBP.
Ran Su, Guobao Xiao, Leyi Wei
Briefings Bioinform.3
2020 A two-step hypergraph reduction based fitting method for unbalanced data
Guobao Xiao, Yan Yan 0001, Hanzi Wang
Pattern Recognit. Lett.1
2020 Deterministic Model Fitting by Local-Neighbor Preservation and Global-Residual Optimization
abstract
Geometric model fitting has been widely used in many computer vision tasks. However, it remains as a challenging task when handing multiple-structural data contaminated by noises and outliers. Most previous work on model fitting cannot guarantee the consistency of their solutions due to their randomness, precluding them from many real-world applications. In this research, we propose a fast two-view approximately deterministic model fitting scheme (called LGF), to provide consistent solutions for multiple-structural data. The proposed LGF scheme starts from defining preference function by preserving local neighborhood relationship, and then adopts the min-hash technique to roughly sample subsets. By this way, it is able to cover all model instances in data in the parameter space with a high probability. After that, LGF refines the previous sampled subsets by globalresidual optimization. Furthermore, we propose a simple yet effective model selection framework to estimate the number and the parameters of model instances in data. Extensive experiments on real images show that the proposed LGF scheme is able to observe superior or very competitive performance on both accuracy and speed over several state-of-the-art model fitting methods.
Guobao Xiao, Jiayi Ma 0001, Shiping Wang, Chang Wen Chen
IEEE Trans. Image Process.1
2020 Learning a Layout Transfer Network for Context Aware Object Detection
abstract
We present a context aware object detection method based on a retrieve-and-transform scene layout model. Given an input image, our approach first retrieves a coarse scene layout from a codebook of typical layout templates. In order to handle large layout variations, we use a variant of the spatial transformer network to transform and refine the retrieved layout, resulting in a set of interpretable and semantically meaningful feature maps of object locations and scales. The above steps are implemented as a Layout Transfer Network which we integrate into Faster RCNN to allow for joint reasoning of object detection and scene layout estimation. Extensive experiments on three public datasets verified that our approach provides consistent performance improvements to the state-of-the-art object detection baselines on a variety of challenging tasks in the traffic surveillance and the autonomous driving domains.
Tao Wang 0047, Xuming He 0001, Yuanzheng Cai, Guobao Xiao
IEEE Trans. Intell. Transp. Syst.4
2019 Hypergraph Optimization for Multi-Structural Geometric Model Fitting
abstract
Recently, some hypergraph-based methods have been proposed to deal with the problem of model fitting in computer vision, mainly due to the superior capability of hypergraph to represent the complex relationship between data points. However, a hypergraph becomes extremely complicated when the input data include a large number of data points (usually contaminated with noises and outliers), which will significantly increase the computational burden. In order to overcome the above problem, we propose a novel hypergraph optimization based model fitting (HOMF) method to construct a simple but effective hypergraph. Specifically, HOMF includes two main parts: an adaptive inlier estimation algorithm for vertex optimization and an iterative hyperedge optimization algorithm for hyperedge optimization. The proposed method is highly efficient, and it can obtain accurate model fitting results within a few iterations. Moreover, HOMF can then directly apply spectral clustering, to achieve good fitting performance. Extensive experimental results show that HOMF outperforms several state-of-the-art model fitting methods on both synthetic data and real images, especially in sampling efficiency and in handling data with severe outliers.
Shuyuan Lin, Guobao Xiao, Yan Yan 0001, David Suter, Hanzi Wang
AAAI2
2019 Superpixel-Guided Two-View Deterministic Geometric Model Fitting
Guobao Xiao, Hanzi Wang, Yan Yan 0001, David Suter
Int. J. Comput. Vis.1
2019 Robust geometric model fitting based on iterative Hypergraph Construction and Partition
Guobao Xiao, Hanzi Wang, Yan Yan 0001, Liming Zhang 0002
Neurocomputing1
2019 Searching for Representative Modes on Hypergraphs for Robust Geometric Model Fitting
abstract
In this paper, we propose a simple and effective geometric model fitting method to fit and segment multi-structure data even in the presence of severe outliers. We cast the task of geometric model fitting as a representative mode-seeking problem on hypergraphs. Specifically, a hypergraph is first constructed, where the vertices represent model hypotheses and the hyperedges denote data points. The hypergraph involves higher-order similarities (instead of pairwise similarities used on a simple graph), and it can characterize complex relationships between model hypotheses and data points. In addition, we develop a hypergraph reduction technique to remove "insignificant" vertices while retaining as many "significant" vertices as possible in the hypergraph. Based on the simplified hypergraph, we then propose a novel mode-seeking algorithm to search for representative modes within reasonable time. Finally, the proposed mode-seeking algorithm detects modes according to two key elements, i.e., the weighting scores of vertices and the similarity analysis between vertices. Overall, the proposed fitting method is able to efficiently and effectively estimate the number and the parameters of model instances in the data simultaneously. Experimental results demonstrate that the proposed method achieves significant superiority over several state-of-the-art model fitting methods on both synthetic data and real images.
Hanzi Wang, Guobao Xiao, Yan Yan 0001, David Suter
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Robust procedural model fitting with a new geometric similarity estimator
Zongliang Zhang, Jonathan Li 0001, Yulan Guo, Xin Li 0003, Yangbin Lin, Guobao Xiao, Cheng Wang 0003
Pattern Recognit.6
2018 Simple Iterative Clustering on Graphs for Robust Model Fitting
abstract
In this paper, we propose a novel method, simple iterative clustering on graphs (SICG), to deal with robust model fitting problems. Specifically, we first construct a graph, where each vertex denotes a model hypothesis and each edge represents the similarity between two model hypotheses, for model fitting. We then propose a simple iterative clustering algorithm, which adapts the k-medoids clustering algorithm, to intuitively estimate model instances in data. The proposed SICG method is able to effectively fit and segment multiple-structure data contaminated with a large number of outliers and noises. Experimental results show that SICG achieves superior fitting results over several state-of-the-art model fitting methods on real images.
Hailing Luo, Guobao Xiao, Hanzi Wang
VCIP2
2018 Conceptual space based model fitting for multi-structure data
Guobao Xiao, Hailing Luo, Bo Li 0006, Yan Yan 0001, Hanzi Wang
Neurocomputing1
2017 A Hierarchical Voting Scheme for Robust Geometric Model Fitting
Guobao Xiao, Yan Yan 0001, Hanzi Wang
ICIG (1)2
2017 Weighted median-shift on graphs for geometric model fitting
abstract
In this paper, we deal with geometric model fitting problems on graphs, where each vertex represents a model hypothesis, and each edge represents the similarity between two model hypotheses. Conventional median-shift methods are very efficient and they can automatically estimate the number of clusters. However, they assign the same weighting scores to all vertices of a graph, which can not show the discriminability on different vertices. Therefore, we propose a novel weighted median-shift on graphs method (WMSG) to fit and segment multiple-structure data. Specifically, we assign a weighting score to each vertex according to the distribution of the corresponding inliers. After that, we shift vertices towards the weighted median vertices iteratively to detect modes. The proposed method can adaptively estimate the number of model instances and deal with data contaminated with a large number of outliers. Experimental results on both synthetic data and real images show the advantages of the proposed method over several state-of-the-art model fitting methods.
Hanzi Wang, Guobao Xiao, Yan Yan 0001, Liming Zhang 0002
ICIP3
2017 Efficient guided hypothesis generation for multi-structure epipolar geometry estimation
Taotao Lai, Hanzi Wang, Yan Yan 0001, Guobao Xiao, David Suter
Comput. Vis. Image Underst.4
2016 Message Passing on the Two-Layer Network for Geometric Model Fitting
Guobao Xiao, Yan Yan 0001, Hanzi Wang
ACCV (1)2
2016 Superpixel-Based Two-View Deterministic Fitting for Multiple-Structure Data
Guobao Xiao, Hanzi Wang, Yan Yan 0001, David Suter
ECCV (6)1
2016 Conceptual space based gross outlier removal for geometric model fitting
abstract
In this paper, we propose an efficient and robust gross outlier removal method, called the Conceptual Space based Gross Outlier Removal (CSGOR) method, to remove gross outliers for geometric model fitting. In the proposed method, each data point is mapped to a conceptual space by computing the preference of "good" model hypotheses. In the conceptual space, the distributions of inliers and gross outliers are significantly different. Specifically, inliers of each model instance are distributed in a subspace and they are far away from the origin of the conceptual space, while gross outliers are distributed near the origin. In this manner, the problem of densely gross outlier removal is formulated as a binary classification problem. The main advantage of the proposed method is that it can handle data with a large proportion of outliers and effectively remove gross outliers in data. Experimental results on both synthetic and real data have demonstrated the efficiency and effectiveness of the proposed method.
Guobao Xiao, Bo Li 0006, Yan Yan 0001, Hanzi Wang
ICARCV3
2016 Rapid hypothesis generation by combining residual sorting with local constraints
Taotao Lai, Hanzi Wang, Yan Yan 0001, Dahan Wang, Guobao Xiao
Multim. Tools Appl.5
2016 Hypergraph modelling for geometric model fitting
Guobao Xiao, Hanzi Wang, Taotao Lai, David Suter
Pattern Recognit.1
2016 Mode seeking on graphs for geometric model fitting via preference analysis
Guobao Xiao, Hanzi Wang, Yan Yan 0001, Liming Zhang 0002
Pattern Recognit. Lett.1
2015 Mode-Seeking on Hypergraphs for Robust Geometric Model Fitting
abstract
In this paper, we propose a novel geometric model fitting method, called Mode-Seeking on Hypergraphs (MSH), to deal with multi-structure data even in the presence of severe outliers. The proposed method formulates geometric model fitting as a mode seeking problem on a hypergraph in which vertices represent model hypotheses and hyperedges denote data points. MSH intuitively detects model instances by a simple and effective mode seeking algorithm. In addition to the mode seeking algorithm, MSH includes a similarity measure between vertices on the hypergraph and a "weight-aware sampling" technique. The proposed method not only alleviates sensitivity to the data distribution, but also is scalable to large scale problems. Experimental results further demonstrate that the proposed method has significant superiority over the state-of-the-art fitting methods on both synthetic data and real images.
Hanzi Wang, Guobao Xiao, Yan Yan 0001, David Suter
ICCV2
2015 An Outlier Removal Method by Statistically Analyzing Hypotheses for Geometric Model Fitting
Yuewei Na, Guobao Xiao, Hanzi Wang
ICIG (1)2
2014 Combining preference analysis with local constraints for rapid hypothesis generation
abstract
Hypothesis generation is crucial to many robust model fitting methods. In this paper, we propose an effective hypothesis generation method by adopting conditional sampling with local constraints. We choose data to generate hypotheses according to sampling weights, which are computed according to ordered residual indices. To sample a minimal subset, we randomly choose a seed datum, compute sampling weights of all data with regard to the seed datum, search the neighborhood set of the seed datum by using the sampling weights, and then sample the remaining data of the minimal subset from the neighborhood set. It has two advantages to consider the neighboring information in guided sampling: It raises the probability of generating all-inlier minimal subsets and it reduces the computational loads in hypotheses generation. The proposed method shows good performance in fundamental matrix estimation using real image pairs.
Taotao Lai, Dahan Wang, Guobao Xiao, Hanzi Wang
ICARCV3