Guoxia Xu

dblp:206/7165 · DBLP profile ↗
← Back
36ranked-venue papers
8as first author
27since 2021 · last 2025
0000-0002-0036-8820ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 since 2021Computer networks · 3 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 PVSSNet: Progressive Feature Interaction Visual State-Space Network for Multispectral Pansharpening
abstract
Pansharpening involves extracting spectral information from multispectral images and structural details from panchromatic images, then fusing them to produce high-resolution multispectral remote sensing images. However, high-resolution multispectral images often suffer from spectral or structural information loss. In this paper, we introduce a pansharpening algorithm based on a Progressive Feature Interaction Visual State Space Network. It enables interaction between local and global features of multispectral and panchromatic images and facilitates the injection of spectral and spatial details through distinct attention modules. This approach effectively preserves both spectral characteristics and spatial structure through inter-branch information interaction and complementation. Additionally, by integrating a visual state space network, the proposed model achieves deep reconstruction of multi-scale global information, enhancing robustness and generalization. Extensive experimental results demonstrate that the proposed network achieves highly competitive performance in both visual assessments and objective metric evaluations.
Guoxia Xu, Zhenwei Xu, Lizhen Deng, Hu Zhu
IEEE Geosci. Remote. Sens. Lett.1
2025 TDSF-Net: Tensor Decomposition-Based Subspace Fusion Network for Multimodal Medical Image Classification
abstract
Data from multimodalities bring complementary information for deep learning-based medical image classification models. However, data fusion methods simply concatenating features or images barely consider the correlations or complementarities among different modalities and easily suffer from exponential growth in dimensions and computational complexity when the modality increases. Consequently, this article proposes a subspace fusion network with tensor decomposition (TD) to heighten multimodal medical image classification. We first introduce a Tucker low-rank TD module to map the high-level dimensional tensor to the low-rank subspace, reducing the redundancy caused by multimodal data and high-dimensional features. Then, a cross-tensor attention mechanism is utilized to fuse features from the subspace into a high-dimension tensor, enhancing the representation ability of extracted features and constructing the interaction information among components in the subspace. Extensive comparison experiments with state-of-the-art (SOTA) methods are conducted on one self-established and three public multimodal medical image datasets, verifying the effectiveness and generalization ability of the proposed method. The code is available at https://github.com/1zhang-yi/TDSFNet.
Yi Zhang 0111, Guoxia Xu, Meng Zhao 0001, Hao Wang 0003, Fan Shi 0001, Shengyong Chen
IEEE Trans. Neural Networks Learn. Syst.2
2024 Multi-Perspective Text-Guided Multimodal Fusion Network for Brain Tumor Segmentation
Huanping Zhang, Yi Zhang 0111, Guoxia Xu, Jiangpeng Zheng, Meng Zhao 0001
PRCV (14)3
2024 SSP-Net: A Siamese-Based Structure-Preserving Generative Adversarial Network for Unpaired Medical Image Enhancement
abstract
Recently, unpaired medical image enhancement is one of the important topics in medical research. Although deep learning-based methods have achieved remarkable success in medical image enhancement, such methods face the challenge of low-quality training sets and the lack of a large amount of data for paired training data. In this article, a dual input mechanism image enhancement method based on Siamese structure (SSP-Net) is proposed, which takes into account the structure of target highlight (texture enhancement) and background balance (consistent background contrast) from unpaired low-quality and high-quality medical images. Furthermore, the proposed method introduces the mechanism of the generative adversarial network to achieve structure-preserving enhancement by jointly iterating adversarial learning. Experiments comprehensively illustrate the performance in unpaired image enhancement of the proposed SSP-Net compared with other state-of-the-art techniques.
Guoxia Xu, Hao Wang 0003, Marius Pedersen, Meng Zhao 0001, Hu Zhu
IEEE Trans. Comput. Biol. Bioinform.1
2024 Deep Tensor Evidence Fusion Network for Sentiment Classification
abstract
Recently, a multimodal sentiment analysis of social media has attracted increasing attention, and its core idea is to discovery heuristic fusion strategy to analyze the sentiment orientations over heterogeneous multimodal source from a learned compact multimodal representation. The existing multimodal fusion techniques not only struggle to achieve full heterogeneous data interaction, but also they are unable to dynamically assess the quality of various modal data to determine predictability. In this article, we present a novel deep tensor evidence fusion (DTEF) network for multimodal sentiment classification. First, we propose a common view evaluation network that uses a long short-term memory (LSTM) network and a tensor-based neural network to extract rich intermodal and intramodal information. Then, we propose a unique time cue evaluation network that takes advantage of the temporal granularity associated with numerous pattern sequences. To make reliable decisions, we finally incorporate uncertainty through the trusted fusion layer, which improves the accuracy and robustness of sentimental classification. Our model is validated using the CMU Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) and CMU Multimodal Corpus of Sentiment Intensity (CMU-MOSI) datasets, and the experimental findings demonstrate the superior performance of the proposed network in terms of accuracy compared with the state-of-the-art methods.
Guoxia Xu, Xiaokang Zhou, Jung Yoon Kim, Hu Zhu, Lizhen Deng
IEEE Trans. Comput. Soc. Syst.2
2024 Trusted Multimodal Socio-Cyber Sentiment Analysis Based on Disentangled Hierarchical Representation Learning
abstract
The rapid development of the digital age has led to a qualitative leap in social media. To meet the cognitive needs of users, social media platforms have been mining users’ private information and disseminating information through various means. However, these platforms lack effective management of information release and various forms of emotional expressions make public propaganda increasingly diverse and complex. Therefore, accurately identifying the relationships between multimodal data poses a challenge. An effective modal representation must consider both the consistency of multimodal data and the complementarity of single-modal data. However, existing methods focus on fusing different modal features into a unified feature representation, while neglecting to evaluate the reliability of prediction results. In this article, we disentangle the consistency and complementarity in the fused representation problem of multimodal data. We construct the modal private task (unique) by using the Dirichlet distribution and evidence theory to solve the uncertainty of each modal prediction. The model can output the uncertainty of prediction and learn complementary information through the fusion of decision layers. At the same time, we construct the modal common task using a low-rank tensor fusion model to learn consistent features. Finally, we compare the model with the current mainstream methods on three public datasets, and the experimental results show that the performance of our method reaches the level of current advanced algorithms.
Guoxia Xu, Lizhen Deng, Yansheng Li 0001, Yantao Wei, Xiaokang Zhou, Hu Zhu
IEEE Trans. Comput. Soc. Syst.1
2024 HQG-Net: Unpaired Medical Image Enhancement With High-Quality Guidance
abstract
Unpaired medical image enhancement (UMIE) aims to transform a low-quality (LQ) medical image into a high-quality (HQ) one without relying on paired images for training. While most existing approaches are based on Pix2Pix/CycleGAN and are effective to some extent, they fail to explicitly use HQ information to guide the enhancement process, which can lead to undesired artifacts and structural distortions. In this article, we propose a novel UMIE approach that avoids the above limitation of existing methods by directly encoding HQ cues into the LQ enhancement process in a variational fashion and thus model the UMIE task under the joint distribution between the LQ and HQ domains. Specifically, we extract features from an HQ image and explicitly insert the features, which are expected to encode HQ cues, into the enhancement network to guide the LQ enhancement with the variational normalization module. We train the enhancement network adversarially with a discriminator to ensure the generated HQ image falls into the HQ domain. We further propose a content-aware loss to guide the enhancement process with wavelet-based pixel-level and multiencoder-based feature-level constraints. Additionally, as a key motivation for performing image enhancement is to make the enhanced images serve better for downstream tasks, we propose a bi-level learning scheme to optimize the UMIE task and downstream tasks cooperatively, helping generate HQ images both visually appealing and favorable for downstream tasks. Experiments on three medical datasets verify that our method outperforms existing techniques in terms of both enhancement quality and downstream task performance. The code and the newly collected datasets are publicly available at https://github.com/ChunmingHe/HQG-Net.
Chunming He, Kai Li 0012, Guoxia Xu, Jiangpeng Yan, Longxiang Tang, Yulun Zhang 0001, Yaowei Wang 0001, Xiu Li 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 DM-Fusion: Deep Model-Driven Network for Heterogeneous Image Fusion
abstract
Heterogeneous image fusion (HIF) is an enhancement technique for highlighting the discriminative information and textural detail from heterogeneous source images. Although various deep neural network-based HIF methods have been proposed, the most widely used single data-driven manner of the convolutional neural network always fails to give a guaranteed theoretical architecture and optimal convergence for the HIF problem. In this article, a deep model-driven neural network is designed for this HIF problem, which adaptively integrates the merits of model-based techniques for interpretability and deep learning-based methods for generalizability. Unlike the general network architecture as a black box, the proposed objective function is tailored to several domain knowledge network modules to model the compact and explainable deep model-driven HIF network termed DM-fusion. The proposed deep model-driven neural network shows the feasibility and effectiveness of three parts, the specific HIF model, an iterative parameter learning scheme, and data-driven network architecture. Furthermore, the task-driven loss function strategy is proposed to achieve feature enhancement and preservation. Numerous experiments on four fusion tasks and downstream applications illustrate the advancement of DM-fusion compared with the state-of-the-art (SOTA) methods both in fusion quality and efficiency. The source code will be available soon.
Guoxia Xu, Chunming He, Hao Wang 0003, Hu Zhu, Weiping Ding 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Degradation-Resistant Unfolding Network for Heterogeneous Image Fusion
abstract
Heterogeneous image fusion (HIF) techniques aim to enhance image quality by merging complementary information from images captured by different sensors. Among these algorithms, deep unfolding network (DUN)-based methods achieve promising performance but still suffer from two issues: they lack a degradation-resistant-oriented fusion model and struggle to adequately consider the structural properties of DUNs, making them vulnerable to degradation scenarios. In this paper, we propose a Degradation-Resistant Unfolding Network (DeRUN) for the HIF task to generate high-quality fused images even in degradation scenarios. Specifically, we introduce a novel HIF model for degradation resistance and derive its optimization procedures. Then, we incorporate the optimization unfolding process into the proposed DeRUN for end-to-end training. To ensure the robustness and efficiency of DeRUN, we employ a joint constraint strategy and a lightweight partial weight sharing module. To train DeRUN, we further propose a gradient direction-based entropy loss with powerful texture representation capacity. Extensive experiments show that DeRUN significantly outperforms existing methods on four HIF tasks, as well as downstream applications, with cheaper computational and memory costs.
Chunming He, Kai Li 0012, Guoxia Xu, Yulun Zhang 0001, Runze Hu, Zhenhua Guo 0001, Xiu Li 0001
ICCV3
2023 Weakly-Supervised Concealed Object Segmentation with SAM-based Pseudo Labeling and Multi-scale Feature Grouping
abstract
Weakly-Supervised Concealed Object Segmentation (WSCOS) aims to segment objects well blended with surrounding environments using sparsely-annotated data for model training. It remains a challenging task since (1) it is hard to distinguish concealed objects from the background due to the intrinsic similarity and (2) the sparsely-annotated training data only provide weak supervision for model learning. In this paper, we propose a new WSCOS method to address these two challenges. To tackle the intrinsic similarity challenge, we design a multi-scale feature grouping module that first groups features at different granularities and then aggregates these grouping results. By grouping similar features together, it encourages segmentation coherence, helping obtain complete segmentation results for both single and multiple-object images. For the weak supervision challenge, we utilize the recently-proposed vision foundation model, ``Segment Anything Model (SAM)'', and use the provided sparse annotations as prompts to generate segmentation masks, which are used to train the model. To alleviate the impact of low-quality segmentation masks, we further propose a series of strategies, including multi-augmentation result ensemble, entropy-based pixel-level weighting, and entropy-based image-level selection. These strategies help provide more reliable supervision to train the segmentation model. We verify the effectiveness of our method on various WSCOS tasks, and experiments demonstrate that our method achieves state-of-the-art performance on these tasks.
Chunming He, Kai Li 0012, Yachao Zhang 0001, Guoxia Xu, Longxiang Tang, Yulun Zhang 0001, Zhenhua Guo 0001, Xiu Li 0001
NeurIPS4
2023 Encoder Activation Diffusion and Decoder Transformer Fusion Network for Medical Image Segmentation
Xueru Li, Guoxia Xu, Meng Zhao 0001, Fan Shi 0001, Hao Wang 0003
PRCV (13)2
2023 Cooperative linear regression model for image set classification
Yu-Feng Yu 0001, Xian-Liang Wang, Long Chen 0001, Yingxu Wang 0002, Guoxia Xu
Expert Syst. Appl.5
2023 Learning the Distribution-Based Temporal Knowledge With Low Rank Response Reasoning for UAV Visual Tracking
abstract
In recent years, the constraint based correlation filter has shown good performance in unmanned aerial vehicle (UAV) tracking, which gains a lot popularity in many intelligence transportation applications. In this work, a distribution-based temporal knowledge driven method is proposed to leverage the temporal translation property in UAV tracking. Instead of focusing on the traditional issues in the correlation filter, we provide a new method of learning parametric distribution on temporal knowledge by Wasserstein distance which is successfully embedded to solve the problem of temporal degeneration in learning process of tracking. Furthermore, we approximate optimal response reasoning with low-rank constraint over response consistency. Furthermore, the proposed method is solved by a simple iterative scheme with alternating direction multiplication ADMM algorithm. We demonstrate the superior tracking performance in several public standard UAV tracking benchmarks compared with state-of-the-art algorithms.
Guoxia Xu, Hao Wang 0003, Meng Zhao 0001, Marius Pedersen, Hu Zhu
IEEE Trans. Intell. Transp. Syst.1
2023 Unpaired Self-supervised Learning for Industrial Cyber-Manufacturing Spectrum Blind Deconvolution
abstract
Cyber-Manufacturing combines industrial big data with intelligent analysis to find and understand the intangible problems in decision-making, which requires a systematic method to deal with rich signal data. With the development of spectral detection and photoelectric imaging technology, spectral blind deconvolution has achieved remarkable results. However, spectral processing is limited by one-dimensional signal, and there is no available structural information with few training samples. Moreover, in the majority of practical applications, it is entirely feasible to gather unpaired spectrum dataset for training. This training method of unpaired learning is practical and valuable. Therefore, a two-stage deconvolution scheme combining self supervised learning and feature extraction is proposed in this paper, which generates two complementary paired sets through self supervised learning to extract the final deconvolution network. In addition, a new deconvolution network is designed for feature extraction. The spectrum is pre-trained through spectral feature extraction and noise estimation network to improve the training efficiency and meet the assumed noise characteristics. Experimental results show that this method is effective in dealing with different types of synthetic noise.
Lizhen Deng, Guoxia Xu, Jiaqi Pi, Hu Zhu, Xiaokang Zhou
ACM Trans. Internet Techn.2
2022 A Self-paced Learning based Transfer Model for Hypergraph Matching
Hu Zhu, Guoxia Xu, Lizhen Deng
Inf. Sci.3
2022 Kernel embedding transformation learning for graph matching
Yu-Feng Yu 0001, Long Chen 0001, Ke-Kun Huang, Hu Zhu, Guoxia Xu
Pattern Recognit. Lett.5
2022 A Dual Stream Spectrum Deconvolution Neural Network
abstract
With the development of spectral detection and photoelectric imaging, multiband spectrum is always degraded by the random noise and band overlap during the acquisition of spectrum devices. Owing to the fixed spectrum degradation model, the existing spectrum deconvolution technologies are sensitive to the handcrafted model designed and manually selected parameters. The fundamental cause of these limitations during spectral analysis is that spectral processing is limited by 1-D signal without structural information available and insufficient training samples. In this article, a dual stream neural network is proposed to reconstruct the original infrared spectroscopy, which effectively strengthens the capability to represent the feature of infrared spectrum. A novel activation function is proposed to realize the function of the dual stream network. Furthermore, a heuristic learning strategy from the perspective of balanced self-paced learning is exploited to help network train from simple to difficult, resolving the problem of high sample repeatability. Compared with other traditional methods, the experimental results show that our network can achieve state-of-the-art reconstruction result and fairly excellent performance in terms of the corresponding index within synectics and real spectrum experiments.
Lizhen Deng, Guoxia Xu, Yanyu Dai, Hu Zhu
IEEE Trans. Ind. Informatics2
2022 FCFusion: Fractal Componentwise Modeling With Group Sparsity for Medical Image Fusion
abstract
Multimodal image fusion is the process of combing relevant biological information that can be used for automated industrial application. In this article, we present a novel framework combining fractal constraint with group sparsity to achieve the optimal fusion quality. First, we adopt the idea of patch division and componentwise separation to perceive the fractal characteristics across multimodality sources. Then, to preserve the spatial information against the redundancy of component-entanglement, the group sparsity is proposed. A dual variable weighting rule is inherently embedded to mitigate the overfitting across the component penalty. Furthermore, the alternating direction method of multipliers is conducted to the proposed model optimization. The experiments show that our model has a better performance in quantitative visual quality and qualitative evaluation analysis. Finally, a real segmentation application of positron emission tomography/computed tomography image fusion proves the effectiveness of our algorithm.
Guoxia Xu, Xiaoxue Deng, Xiaokang Zhou, Marius Pedersen, Lucia Cimmino, Hao Wang 0003
IEEE Trans. Ind. Informatics1
2022 PcGAN: A Noise Robust Conditional Generative Adversarial Network for One Shot Learning
abstract
Traffic sign classification plays a vital role in autonomous vehicles for its powerful capability in information representation. However, the low-quality data of traffic signs captured by in-vehicle cameras often inevitably bring inherent challenges to the one-shot classification task. Apart from the problem of data degradation, learning-based classification techniques of real traffic signs also come across the challenges of intra-class and inter-class data imbalance from the training data. To overcome the aforementioned problems, we propose an end-to-end degradation robust deep model, termed PcGAN, to classify traffic signs in a manner of few-shot learning. The proposed PcGAN models the joint distribution between the degraded traffic signal data and the corresponding prototypes from both degradation removal and generation perspectives by two alternating optimized modules, which ensures the generalization of the learned embedding of latent space for novel tasks. A multi-task loss function is designed to improve the robustness of PcGAN. Numerous experiments comprehensively demonstrate that the accuracy of our proposed PcGAN is improved by 5% compared with other state-of-the-art (SOTA) approaches in few-shot classification.
Lizhen Deng, Chunming He, Guoxia Xu, Hu Zhu, Hao Wang 0003
IEEE Trans. Intell. Transp. Syst.3
2022 Bilateral Weighted Regression Ranking Model With Spatial-Temporal Correlation Filter for Visual Tracking
abstract
Many discriminative correlation filter (DCF)-based methods have successfully leveraged the guidance for solving two problems (i.e., the boundary effect and temporal filtering degradation) as a model prior to visual tracking. The intuitive motivation of these methods is to control the degeneration of the updating loss of the objective function with a structural framework. While these methods rely mostly on various regularization items, they always ignore the loss from data fidelity term. Therefore, we propose a bilateral weighted regression ranking model termed as BWRR. Here, we resort to two procedures for solving the above problems. First, BWRR introduces a bilateral constraint into the data fidelity term to control the loss of rows and columns of the filter learning data term. The weighted matrices could impose an adaptive penalty for large data loss during the learning process to avoid the model degradation problem. Second, the data of the updated weighted matrices is not directly applied to the calculation of the filter during each iteration. Instead, a new weighted product matrix is obtained by ranking and numerical transformation for updating the filter. We show that the proposed model converts the original correlation filter regression problem into a regression-with-ranking problem, thus avoiding the problem of positive and negative sample imbalance. Overall, the BWRR model is iteratively solved by the alternating direction method of multipliers(ADMM). Qualitative and quantitative evaluations demonstrate the effectiveness and superiority of our proposed method by extensive and quantitative experiments on the OTB, VOT, and UAV datasets.
Hu Zhu, Guoxia Xu, Lizhen Deng, Yueying Cheng, Aiguo Song
IEEE Trans. Multim.3
2021 Video smoke removal based on low-rank tensor completion via spatial-temporal continuity constraint
abstract
Abstract Smoke has a very bad effect on the outdoor vision system. Not only are the videos with poor visual effects obtained, but also the quality and structure of the videos are reduced. In this paper, we propose a video smoke removal method based on low‐rank tensor completion via spatial‐temporal continuity constraint. The proposed method is based on the smoke mixing model and consider the sparseness of smoke and the global and local consistency of clean video. Then, the optimal solution of the smoke removal algorithm model is quickly realized by the Alternating Direction Method of Multiplier. Finally, we evaluate the experiment results of real‐world data and simulated data from the visual effects and objective indicators. And the experiment results show that our proposed algorithm can achieve better smoke removal results.
Hu Zhu, Guoxia Xu, Lizhen Deng
Concurr. Comput. Pract. Exp.2
2021 Vector co-occurrence morphological edge detection for colour image
abstract
Abstract Morphological edge detection is a principal component in pattern recognition and machine vision. Traditional edge detection operators only take pixel mutual into consideration. However, the edges are influenced not only by pixel mutual but also by the boundary characteristics. Here, the vector co‐occurrence morphological edge detection operator is proposed, which takes the pixel and boundary information both into consideration. The vector co‐occurrence algorithm is exploited to resist the influence of the noise points and detect the edges from the colour image rather than the grey image. And, we lead to define a precise definition of the manner of sorting high‐dimensional data for the colour image. The experiment results always illustrate the advancement and practicability of our methods against the baseline method. In terms of experiments, the BSDS500 dataset is introduced to compare and analyse with other algorithms. Based on the standard benchmark index evaluation in the BSDS500 dataset, the ODS and AP of various algorithms are compared and analysed.
Chunming He, Yu-Feng Yu 0001, Guoxia Xu, Hu Zhu, Lizhen Deng
IET Image Process.4
2021 RoDeRain: Rotational Video Derain via Nonconvex and Nonsmooth Optimization
Lizhen Deng, Guoxia Xu, Hu Zhu, Bing-Kun Bao
Mob. Networks Appl.2
2021 Infrared small target detection via adaptive M-estimator ring top-hat transformation
Lizhen Deng, Jieke Zhang, Guoxia Xu, Hu Zhu
Pattern Recognit.3
2021 Dual Calibration Mechanism Based L2, p-Norm for Graph Matching
abstract
Unbalanced geometric structure caused by variations with deformations, rotations and outliers is a critical issue that hinders correspondence establishment between image pairs in existing graph matching methods. To deal with this problem, in this work, we propose a dual calibration mechanism (DCM) for establishing feature points correspondence in graph matching. In specific, we embed two types of calibration modules in the graph matching, which model the correspondence relationship in point and edge respectively. The point calibration module performs unary alignment over points and the edge calibration module performs local structure alignment over edges. By performing the dual calibration, the feature points correspondence between two images with deformations and rotations variations can be obtained. To enhance the robustness of correspondence establishment, the L2,p-norm is employed as the similarity metric in the proposed model, which is a flexible metric due to setting the different p values. Finally, we incorporate the dual calibration and L2,p-norm based similarity metric into the graph matching model which can be optimized by an effective algorithm, and theoretically prove the convergence of the presented algorithm. Experimental results in the variety of graph matching tasks such as deformations, rotations and outliers evidence the competitive performance of the presented DCM model over the state-of-the-art approaches.
Yu-Feng Yu 0001, Guoxia Xu, Ke-Kun Huang, Hu Zhu, Long Chen 0001, Hao Wang 0003
IEEE Trans. Circuits Syst. Video Technol.2
2021 Tensor Field Graph-Cut for Image Segmentation: A Non-Convex Perspective
abstract
Image segmentation is a key component of image analysis, which refers to the process of partitioning the image into multiple segments. Graph cut is widely used in image segmentation by constructing a graph that the minimal cut of this graph would lead to partition the corresponding pixels of the different objects. In this paper, we reconstruct the graph cut problem as a special non-convex optimization problem instead of the traditional maximum flow problem. We extend this non-convex problem to the hypergraph method and combine it with a tensor field based on a directional bilateral filter bank to achieve segmentation in grayscale images. Accordingly, an efficient minimization algorithm is proposed to solve this non-convex problem with global convergence. Furthermore, we have selected the data of BSDS300 and BSDS500 as tests. Experimental results and evaluation index tests further demonstrate the superiority of the proposed method.
Hu Zhu, Jieke Zhang, Guoxia Xu, Lizhen Deng
IEEE Trans. Circuits Syst. Video Technol.3
2021 Joint Transformation Learning via the L2, 1-Norm Metric for Robust Graph Matching
abstract
Establishing correspondence between two given geometrical graph structures is an important problem in computer vision and pattern recognition. In this paper, we propose a robust graph matching (RGM) model to improve the effectiveness and robustness on the matching graphs with deformations, rotations, outliers, and noise. First, we embed the joint geometric transformation into the graph matching model, which performs unary matching over graph nodes and local structure matching over graph edges simultaneously. Then, the L2,1-norm is used as the similarity metric in the presented RGM to enhance the robustness. Finally, we derive an objective function which can be solved by an effective optimization algorithm, and theoretically prove the convergence of the proposed algorithm. Extensive experiments on various graph matching tasks, such as outliers, rotations, and deformations show that the proposed RGM model achieves competitive performance compared to the existing methods.
Yu-Feng Yu 0001, Guoxia Xu, Min Jiang 0003, Hu Zhu, Dao-Qing Dai, Hong Yan 0001
IEEE Trans. Cybern.2
2020 Kernelized dual regression incorporating local information for image set classification
Xian-Liang Wang, Jiao Du, Guoxia Xu, Ignazio Passero, Hao Wang 0003, Yu-Feng Yu 0001
Pattern Recognit. Lett.3
2020 Image Correspondence With CUR Decomposition-Based Graph Completion and Matching
abstract
Establishing correspondence between pictorial descriptions of two images is an important task and can be treated as graph matching problem. However, the process of extracting a favourable graph structure from raw images for matching is influenced by cluttered backgrounds and deformations, which may result in the abundance of noisy graph structures. This paper addresses the problem of point set correspondence and presents a robust graph matching method which recovers the correspondence matches among the graph nodes in a CUR based factorization framework. The graph representation in terms of CUR, inherently preserves the actual nodes connection in sparse manner, this particularly renders the complex space-time realization of affinities among graph nodes. The reformulation of graph matching in terms of small CUR factorization matrices, allows to compute and relax the partially observed graphs, without observing the whole large-scale graph matrix. In particular, we propose two variants of this approach, first, approximating the matching matrix from small CUR observed graph structure, and second, completing the graph structure with higher order CUR form to find correspondence. The CUR based matching algorithms are realized by computing set of compatibility coefficients from pairwise matching graphs and further conducting the probability relaxation procedure to find the matching confidences among nodes. Experiments and analysis on synthetic and natural images dataset prove the effectiveness of proposed methods against state-of-the-art methods. We also explore CUR matching for non-rigid moving object in a video sequence to demonstrate the potential application of graph matching to video analysis.
Sheheryar Khan, Mehmood Nawaz, Guoxia Xu, Hong Yan 0001
IEEE Trans. Circuits Syst. Video Technol.3
2020 DSPNet: A Lightweight Dilated Convolution Neural Networks for Spectral Deconvolution With Self-Paced Learning
abstract
In the fields of industry research, infrared spectrometers are widely used in diverse applications. However, the spectrum often suffers from band overlap and random noise due to the distortion caused by the point spread function, especially for aging instruments. The problem of reconstructing the clear spectrum from the degraded spectrum is called spectrum deconvolution. Traditional partial differential equation (PDE) methods rely on distribution assumptions in the reconstructed process. This restriction makes PDE methods sensitive to tackle complex instrumental broadening effect in the dispersive IR spectrometers. Also, we need to spend much time setting the parameters of PDE models manually. These problems intuitively degrade the performances of PDE methods. In this article, we propose an end-to-end neural network framework for spectral deconvolution problem. The novelty of this article lies in its strong robustness from dilated deconvolution and self-paced learning procedure to challenge the complicated degraded spectra. Actually, the deconvolution problem is tailored to a dense prediction problem in this article. Inspired by the extensive use and excellent effects of dilated convolutions in dense prediction, a lightweight dilated convolution module is given to detect the overlaps of degraded spectra. Experimental results demonstrate that the proposed solution has an outstanding performance against many other approaches. Such improvements have the potential to facilitate industrial applications and further exploration of an unknown chemical mixture. Our framework has a good performance on feature extracting and spectrum reconstruction, even in the case of low signal-to-noise ratio.
Hu Zhu, Yiming Qiao, Guoxia Xu, Lizhen Deng, Yu-Feng Yu 0001
IEEE Trans. Ind. Informatics3
2020 TNLRS: Target-Aware Non-Local Low-Rank Modeling With Saliency Filtering Regularization for Infrared Small Target Detection
abstract
Recently, infrared small target detection problem has attracted substantial attention. Many works based on local low-rank model have been proven to be very successful for enhancing the discriminability during detection. However, these methods construct patches by traversing local images and ignore the correlations among different patches. Although the calculation is simplified, some texture information of the target is ignored, and targets of arbitrary forms cannot be accurately identified. In this paper, a novel target-aware method based on a non-local low-rank model and saliency filter regularization is proposed, with which the newly proposed detection framework can be tailored as a non-convex optimization problem, therein enabling joint target saliency learning in a lower dimensional discriminative manifold. More specifically, non-local patch construction is applied for the proposed target-aware low-rank model. By combining similar patches, we reconstruct them together to achieve a better generalization of non-local spatial sparsity constraints. Furthermore, to encourage target saliency learning, our proposed saliency filtering regularization term based on entropy is restricted to lie between the background and foreground. The regularization of the saliency filtering locally preserves the contexts from the target and surrounding areas and avoids the deviated approximation of the low-rank matrix. Finally, a unified optimization framework is proposed and solved with the alternative direction multiplier method (ADMM). Experimental evaluations of real infrared images demonstrate that the proposed method is more robust under different complex scenes compared with some state-of-the-art methods.
Hu Zhu, Haopeng Ni, Shiming Liu, Guoxia Xu, Lizhen Deng
IEEE Trans. Image Process.4
2020 Multimodal Fusion Method Based on Self-Attention Mechanism
abstract
Multimodal fusion is one of the popular research directions of multimodal research, and it is also an emerging research field of artificial intelligence. Multimodal fusion is aimed at taking advantage of the complementarity of heterogeneous data and providing reliable classification for the model. Multimodal data fusion is to transform data from multiple single-mode representations to a compact multimodal representation. In previous multimodal data fusion studies, most of the research in this field used multimodal representations of tensors. As the input is converted into a tensor, the dimensions and computational complexity increase exponentially. In this paper, we propose a low-rank tensor multimodal fusion method with an attention mechanism, which improves efficiency and reduces computational complexity. We evaluate our model through three multimodal fusion tasks, which are based on a public data set: CMU-MOSI, IEMOCAP, and POM. Our model achieves a good performance while flexibly capturing the global and local connections. Compared with other multimodal fusions represented by tensors, experiments show that our model can achieve better results steadily under a series of attention mechanisms.
Hu Zhu, Yingying Hua, Guoxia Xu, Lizhen Deng
Wirel. Commun. Mob. Comput.5
2019 Singular value decomposition based recommendation using imputed data
Xiaofeng Yuan, Lixin Han, Subin Qian, Guoxia Xu, Hong Yan 0001
Knowl. Based Syst.4
2019 Dilated-aware discriminative correlation filter for visual tracking
Guoxia Xu, Hu Zhu, Lizhen Deng, Lixin Han, Yujie Li 0001, Huimin Lu 0001
World Wide Web1
2018 Discriminative tracking via supervised tensor learning
Guoxia Xu, Sheheryar Khan, Hu Zhu, Lixin Han, Michael Kwok-Po Ng, Hong Yan 0001
Neurocomputing1
2017 An online spatio-temporal tensor learning model for visual tracking and its applications to facial expression recognition
Sheheryar Khan, Guoxia Xu, Hong Yan 0001
Expert Syst. Appl.2