Xiongxin Tang

dblp:353/7153 · DBLP profile ↗
← Back
25ranked-venue papers
0as first author
25since 2021 · last 2026
0009-0000-7037-3097ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 12 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 LIRNet: Boosting the performance for unified low-light image restoration
Chao Yin 0001, Fan Ji, Xiongxin Tang, Fanjiang Xu
Expert Syst. Appl.4
2026 RIR-Agent: An interactive framework for effective and adaptive restoration of remote sensing imagery
Junyu Liu, Tianyu Li 0003, Lanyue Liang, Gang Fu 0003, Guoqing Wang 0001, Quan Rui, Xiongxin Tang, Shuyuan Zhu, Yang Yang 0002
Expert Syst. Appl.7
2026 Towards continual low-light image enhancement through causal inference
Fan Ji, Jiangmeng Li, Xiongxin Tang, Fanjiang Xu
Neural Networks4
2026 Think Twice Before Determining: Toward Scene-Aware Visual Reasoning for Mirror Detection
abstract
Mirror detection (MD) aims to overcome interference caused by reflections and locate mirror regions. Existing methods focus on designing components to explicitly establish the associations between physical entities and corresponding imagings, or utilizing rotation to construct symmetric consistency. We observe that: a) incomplete and incorrect correspondence between entities and imagings; b) other physical materials (e.g., glass) exhibit characteristics partially similar to mirrors, causing confusion when they co-occur; c) complex interfering factors (e.g., occlusion) and reflection mechanisms may expand vector space several times over. To address these issues in a unified manner, we formulate the scene-aware visual reasoning network (SVRNet) based on visual prompts. Specifically, we construct the prototype-guided prompt chain reasoning (PPCR) that generates a mixed chain of thought reasoning based on maximal difference heterogeneous prototypes to construct comprehensive spatial location and semantic perception. Noise may accumulate gradually through the chain, and crucial clues may also disappear. Therefore, we design the prompt evolution (PE) to filter out noise and enhance the coupling between prompts. We further develop the mixture of prompt injection expert (MPIE) to dynamically select the optimal injection strategy in the low-rank space based on specific scene. Due to reflection interference and random parameter space introducing potential ambiguity, we formulate the three-way evidence-aware (TEA) loss to quantify the uncertainty, thereby providing reliable predictions. To leverage historical knowledge and further disentangle representations, we propose the frequency prototype contrastive (FPC) loss for learning more generalizable features across images. Finally, we relabel 25,828 images and formulate the first point-supervised MD framework. Extensive experiments conducted on four mirror benchmarks under three settings demonstrate that our method surpasses state-of-the-art approaches. Promising results are also achieved on six related benchmarks, showing its generality.
Mingfeng Zha, Guoqing Wang 0001, Yunqiang Pei, Tianyu Li 0003, Xiongxin Tang, Jiayi Ma 0001, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Circuits Syst. Video Technol.5
2026 Expose Camouflage in the Water: Underwater Camouflaged Instance Segmentation and Dataset
abstract
With the development of underwater exploration and marine protection, underwater vision tasks are widespread. Due to the degraded underwater environment, characterized by color distortion, low contrast, and blurring, camouflaged instance segmentation (CIS) faces greater challenges in accurately segmenting objects that blend closely with their surroundings. Traditional camouflaged instance segmentation methods, trained on terrestrial-dominated datasets with limited underwater samples, may exhibit inadequate performance in underwater scenes. To address these issues, we introduce the first underwater camouflaged instance segmentation (UCIS) dataset, abbreviated as UCIS4K, which comprises 3,953 images of camouflaged marine organisms with instance-level annotations. In addition, we propose an Underwater Camouflaged Instance Segmentation network based on Segment Anything Model (UCIS-SAM). Our UCIS-SAM includes three key modules. First, the Channel Balance Optimization Module (CBOM) enhances channel characteristics to improve underwater feature learning, effectively addressing the model's limited understanding of underwater environments. Second, the Frequency Domain True Integration Module (FDTIM) is proposed to emphasize intrinsic object features and reduce interference from camouflage patterns, enhancing the segmentation performance of camouflaged objects blending with their surroundings. Finally, the Multi-scale Feature Frequency Aggregation Module (MFFAM) is designed to strengthen the boundaries of low-contrast camouflaged instances across multiple frequency bands, improving the model's ability to achieve more precise segmentation of camouflaged objects. Extensive experiments on the proposed UCIS4K and public benchmarks show that our UCIS-SAM outperforms state-of-the-art approaches. The code and dataset are released at https://github.com/wchchw/UCIS4K.
Chuhong Wang, Hua Li 0012, Chongyi Li, Huazhong Liu, Xiongxin Tang, Sam Kwong
IEEE Trans. Image Process.5
2026 3D-UIR: 3D Gaussian for Underwater 3D Scene Reconstruction via Physics-Based Appearance-Medium Decoupling
abstract
Novel view synthesis for underwater scene reconstruction presents unique challenges due to complex light-media interactions. Optical scattering and absorption in water body bring inhomogeneous medium attenuation interference that disrupts conventional volume rendering assumptions of uniform propagation medium. While 3D Gaussian Splatting (3DGS) offers real-time rendering capabilities, it struggles with underwater inhomogeneous environments where scattering media introduces artifacts and inconsistent appearance. In this study, we propose a physics-based framework that disentangles object appearance from water medium effects through tailored Gaussian modeling. Our approach introduces appearance embeddings, which are explicit medium representations for backscatter and attenuation, enhancing scene consistency. In addition, we propose a depth-guided optimization strategy that leverages pseudo-depth maps as supervision with depth regularization and scale penalty terms to improve geometric fidelity. By integrating the proposed appearance and medium modeling components via an underwater imaging model, our approach achieves both high-quality novel view synthesis and physically accurate scene restoration. Experiments demonstrate our significant improvements in rendering quality and restoration accuracy over existing methods. The project page is available at https://bilityniu.github.io/3D-UIR.
Jieyu Yuan, Yuanlin Zhang 0009, Chunle Guo, Xiongxin Tang, Ruixing Wang, Chongyi Li
IEEE Trans. Image Process.5
2025 BIAWDiff: Enhancing Low-Light Images with Bio-Inspired Attention and Wavelet Diffusion
abstract
Low-light image enhancement aims to improve visual quality under challenging lighting conditions while preserving details and color fidelity. Existing traditional algorithms and deep learning approaches, often struggle with balancing brightness enhancement and detail preservation, leading to issues such as overexposure, artifacts, and loss of high-frequency details. To address these challenges, we propose a novel method, Bio-Inspired Attention and Wavelet Diffusion (BIAWDiff), that integrates Retinex theory with bio-inspired attention and wavelet-based diffusion models to enhance low-light Images. BIAWDiff consists of three key modules: the Initial Light Restoration (ILR) module for brightness enhancement and noise reduction, the Rod Cell-Inspired Attention Refinement (RCAR) module for luminance refinement, and the Detail Refinement (DR) module for restoring high-frequency details. Experimental results demonstrate that BIAWDiff outperforms existing techniques, achieving superior results in brightness enhancement, noise reduction, and detail preservation, with an average PSNR increase of 5.1% and SSIM improvement of 3.2% on paired datasets, and a reduction in NIQE and BRISQUE by 7.4% and 8.6% on unpaired datasets.
Hanxiang Yang, Xiongxin Tang, Fengge Wu, Fanjiang Xu
ICASSP4
2025 Amplitude-Guidance Low-Light Image Enhancement with Frequency-based Channel Attention
abstract
Low-light image enhancement aims to improve lightness and eliminate degradation caused by low light. However, most current methods struggle to effectively handle the mixed degradations of both brightness and structure, leading to structural distortions and insufficient brightness enhancement. Additionally, existing Fourier-based methods learn amplitude and phase independently, yet overlook the intrinsic connection between brightness and structure. In this paper, we propose an amplitude-guidance low-light image enhancement network, which utilizes the Fourier transform to extract the amplitude and phase of images and reconstruct them using the network. Considering that uneven brightness distribution in images can lead to varying levels of structural degradation, we design an Amplitude-Guidance Self-Attention (AGSA) that uses amplitude to guide phase recovery, enabling it to handle different levels of structure degradation. Additionally, to further improve the enhancement capability of our network, we design a Frequency-based Channel Attention (FCA) that preserves more frequency information when compressing channels. Extensive experiments demonstrate the superiority of our proposed network over existing SOTA methods.
Xiongxin Tang, Fanjiang Xu, Hanxiang Yang
ICASSP2
2025 DMKPN: Image Deblurring Under Multi-Factor Aliasing Diffusion Degradation
abstract
Image degradation results from a combination of factors. Recently, CNN-based image deblurring methods have made significant progress, but they rely heavily on the accuracy of paired data, which is impractical to collect for every camera. To address this, we propose a physical model for natural images that applies to various cameras. This model considers the diffusion effects of multiple factors during degradation and effectively simulates the degraded state of natural images. We then design the Defocus Map-based Kernel Prediction Network (DMKPN) for adaptive image quality enhancement. Specifically, we develop a DM-Attention Block to assist kernel prediction under the guidance of the defocus map and design Multi-Scale Modulation to filter information at each scale, making the most of image context. Additionally, Multi-Scale Loss is introduced to enhance network robustness. Experiments demonstrate that our method exhibits strong spatial adaptability and generates high-quality images with sharp edges.
Xiongxin Tang, Hanxiang Yang, Fanjiang Xu
ICASSP2
2025 Frequency-Domain Guided Multiple Parallel Kernels Network for Low-Light Remote Sensing Image Enhancement
abstract
Due to dark environments, optical aberrations, etc, the remote sensing images are often submerged under low contrast degradation, which greatly hinders their practical applications for agricultural management and other related tasks. The surface features of remote sensing images are often continuously distributed in space, thus, the sizes of the network’s receptive fields and its ability to learn long-range dependencies are crucial for restoring low-light remote sensing images. Existing methods based on CNN provide limited receptive fields, while Transformer-based methods are constrained by their quadratic computational complexity. To cope with these issues, we propose a novel low-light remote sensing image enhancement network that combines multi-scale receptive fields with frequency-domain attention. Specifically, this network employs multiple parallel kernels of varying sizes to learn multi-scale local features in the spatial domain and complements frequency-domain information to learn global long-range correlations, which achieves local-global feature extraction and further facilitates subsequent degraded images enhancement. We have conducted extensive experiments to demonstrate that our network outperforms existing methods quantitatively and achieves exceptional visual performance, which fully highlights the effectiveness and superiority of our method in enhancing low-light remote sensing images.
Jingxuan Zhou, Xiongxin Tang, Fanjiang Xu
ICASSP4
2025 ASFST:Adaptive Spectral Filters Sparse Transformer for Hyperspectral Image Denoising
Ruijie Chen, Xiongxin Tang, Fanjiang Xu
ICIC (6)2
2025 Learning Adaptive High-Frequency Semantic Guidance for Low-light Image Enhancement
abstract
The low-light image enhancement has always been an important yet challenging task, which attracts significant attention in many fields. However, prior methods either ignore integrating semantic priors or depend on the masks generated by the pre-trained segmentation model. This way is complex and inevitably leads to inaccurate masks when facing unseen scenarios, which may be incompatible with the original feature and result in suboptimal performance. To address this issue, we first consider the high-frequency physical prior is more related to structural and textural properties, which embrace the rich semantic clues and can adaptively assist the learning process under various scenarios. Inspired by this, we propose the high-frequency semantic-aware guidance framework (HighFreNet) to leverage the guidance of semantic information tailored for enhancing low-light images. Specifically, the core parts are the novel Frequency-based Semantic Embedding Module (FSEM) and the Spatial-based Semantic Embedding Module (SSEM), which are separately designed to fully exploit the structure knowledge to modulate the original representation from frequency and spatial perspectives. Extensive experiments showcase that our method significantly outperforms the state-of-the-art methods on five benchmark datasets both in natural and remote sensing environments.
Jingxuan Zhou, Jiangmeng Li, Xiongxin Tang, Fanjiang Xu
ICME5
2025 SNRFour: Rethinking the SNR Guidance for Low-Light Image Enhancement from the Frequency Perspective
Fan Ji, Xiongxin Tang, Fanjiang Xu
PRCV (8)3
2025 Continual Test-Time Adaptation for Single Image Defocus Deblurring via Causal Siamese Networks
Jiangmeng Li, Xiongxin Tang, Bing Su 0001, Fanjiang Xu, Hui Xiong 0001
Int. J. Comput. Vis.4
2025 Unlocking spatial textures: Gradient-guided pansharpening for enhancing multispectral imagery
Lanyue Liang, Tianyu Li 0003, Guoqing Wang 0001, Lin Mei 0001, Xiongxin Tang, Chaofan Qiao, Dongyu Xie
Neurocomputing5
2025 Heterogeneous Experts and Hierarchical Perception for Underwater Salient Object Detection
abstract
Existing underwater salient object detection (USOD) methods design fusion strategies to integrate multimodal information, but lack exploration of modal characteristics. To address this, we separately leverage the RGB and depth branches to learn disentangled representations, formulating the heterogeneous experts and hierarchical perception network (HEHP). Specifically, to reduce modal discrepancies, we propose the hierarchical prototype guided interaction (HPI), which achieves fine-grained alignment guided by the semantic prototypes, and then refines with complementary modalities. We further design the mixture of frequency experts (MoFE), where experts focus on modeling high- and low-frequency respectively, collaborating to explicitly obtain hierarchical representations. To efficiently integrate diverse spatial and frequency information, we formulate the four-way fusion experts (FFE), which dynamically selects optimal experts for fusion while being sensitive to scale and orientation. Since depth maps with poor quality inevitably introduce noises, we design the uncertainty injection (UI) to explore high uncertainty regions by establishing pixel-level probability distributions. We further formulate the holistic prototype contrastive (HPC) loss based on semantics and patches to learn compact and general representations across modalities and images. Finally, we employ varying supervision based on branch distinctions to implicitly construct difference modeling. Extensive experiments on two USOD datasets and four relevant underwater scene benchmarks validate the effect of the proposed method, surpassing state-of-the-art binary detection models. Impressive results on seven natural scene benchmarks further demonstrate the scalability.
Mingfeng Zha, Guoqing Wang 0001, Yunqiang Pei, Tianyu Li 0003, Xiongxin Tang, Chongyi Li, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Image Process.5
2024 Learning Semantic-aware Retinex Network with Spatial-Frequency Interaction for Low-light Image Enhancement
abstract
Retinex-based methods have achieved significant progress in enhancing low-light images benefits for its disentanglement property. However, existing methods either ignore the semantic priors or randomly leverage them only in the spatial domain, which leads to insufficient coupling and limits the performance gains. Considering the lightness mainly exists in the amplitude component and the rest is related to the phase component, making it optimal to combine the Retinex decomposition with the Fourier transform to achieve customized restoration. In this paper, we propose a novel method RetinexFour tailored for low-light image enhancement. Specifically, it consists of Phase-Guided Multi-head Self-Attention (PG-MSA) and Local Spatial Attention (LSA) to allow for the reconstruction of structure from spatial-frequency perspectives. To achieve exposure correction, we introduce Selective Amplitude feature Fusion (SAFF) by combining the original and complementary amplitude to achieve global lightness adjustment. Extensive experiments demonstrate the superiority of our method over other SOTA methods on four benchmark datasets.
Hanxiang Yang, Xiongxin Tang, Fanjiang Xu
ICME4
2024 Learning Frequency-Aware Representation For Low-light Image Enhancement
abstract
The low-light image enhancement (LLIE) aims to improve image brightness and alleviate the degradation caused by low-light conditions. Recently, many researchers have explored the frequency information for LLIE. Within the frequency domain, amplitude indicates brightness, while phase represents structural details. Based on this observation, numerous methods have been proposed to learn features in the Fourier space and achieved impressive performance. However, these methods ignore the interactions between different frequencies and suffer from a cumbersome two-stage training process. In this paper, we introduce a simple yet effective one-stage frequency-aware network, FANet, comprising two core modules: frequency self-attention block (FSAB) and frequency filter block (FFB). FSAB utilizes the self-attention mechanism to separately learn the illumination and structure degradations based on the Fourier prior. Motivated by the fact that different frequencies contribute to LLIE to varying extents, we propose to pay more attention to important frequencies. To this end, the frequency filter mechanism is applied to capture global frequency information in FFB, dynamically focusing on the crucial frequencies and further improving both amplitude and phase features for LLIE. We validate our proposed approach on the LOL-v1, LOL-v2-real, and LOL-v2-synthetic datasets using PSNR and SSIM metrics. The quantitative and qualitative experiments demonstrate the superiority and effectiveness of our method against the other state-of-the-art methods.
Xiongxin Tang
IJCNN2
2024 HQPAFT: Enhancing Low-Light Images with High-Quality Priors and Advanced Feature Transformations Using Only Normal Light Images
Hanxiang Yang, Xiongxin Tang, Fanjiang Xu
PRICAI (3)4
2024 Optical Imaging Degradation Simulation and Transformer-Based Image Restoration for Remote Sensing
abstract
Due to atmospheric turbulence, optical system limitations, satellite platform jitter, and other reasons, remote sensing images inevitably undergo different degrees of degradation. Employing the deep learning method to improve the on-orbit image quality faces many challenges such as lack of data, limited computing resources, network architecture design, and so on. Among these factors, establishing a physics-guided dataset during the image restoration stage and avoiding unforeseen effects such as ringing pose a significant challenge for remote sensing image restoration. This letter proposes an optical imaging degradation simulation model and Transformer-based algorithm to improve remote sensing image quality. First, we model the degradation result from phase to image of optical remote sensing imaging using Zernike Polynomials, thus, a large-scale paired dataset is constructed. Then, a multi-level feature fusion transformer is introduced to mitigate the defect during restoration. The proposed algorithm incorporates a multi-level feature fusion module to fuse feature information from multi-scales effectively. Additionally, a multi-level space and frequency loss function is introduced to enhance the learning of high-frequency information to ensure that the edge suppresses noise amplification and ringing effects during recovery. Finally, experimental results on synthetic data show that our method improved by 25.4% and 22.3% with the blurred images on the PSNR index and SSIM index. Visual results on the GaoFen-1/2A PMS images have enhanced clarity and suppressed artifacts such as ringing which demonstrate the effectiveness and capability of our proposed method.
Hua Wei 0007, Kun Gao 0001, Qiuyan Tang, Xiongxin Tang, Fanjiang Xu
IEEE Geosci. Remote. Sens. Lett.5
2024 Towards a Flexible Semantic Guided Model for Single Image Enhancement and Restoration
abstract
Low-light image enhancement (LLIE) investigates how to improve the brightness of an image captured in illumination-insufficient environments. The majority of existing methods enhance low-light images in a global and uniform manner, without taking into account the semantic information of different regions. Consequently, a network may easily deviate from the original color of local regions. To address this issue, we propose a semantic-aware knowledge-guided framework (SKF) that can assist a low-light enhancement model in learning rich and diverse priors encapsulated in a semantic segmentation model. We concentrate on incorporating semantic knowledge from three key aspects: a semantic-aware embedding module that adaptively integrates semantic priors in feature representation space, a semantic-guided color histogram loss that preserves color consistency of various instances, and a semantic-guided adversarial loss that produces more natural textures by semantic priors. Our SKF is appealing in acting as a general framework in the LLIE task. We further present a refined framework SKF++ with two new techniques: (a) Extra convolutional branch for intra-class illumination and color recovery through extracting local information and (b) Equalization-based histogram transformation for contrast enhancement and high dynamic range adjustment. Extensive experiments on various benchmarks of LLIE task and other image processing tasks show that models equipped with the SKF/SKF++ significantly outperform the baselines and our SKF/SKF++ generalizes to different models and scenes well. Besides, the potential benefits of our method in face detection and semantic segmentation in low-light conditions are discussed.
Yuhui Wu 0001, Guoqing Wang 0001, Shaochong Liu, Yang Yang 0002, Wei Liu 0005, Xiongxin Tang, Shuhang Gu, Chongyi Li, Heng Tao Shen
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Dual Domain Perception and Progressive Refinement for Mirror Detection
abstract
Mirror detection aims to discover mirror regions in images to avoid misidentifying reflected objects. Existing methods mainly mine clues from spatial domain. We observe that the frequencies inside and outside the mirror region are distinctive. Besides, the low-frequency representing the feature semantics can help to locate the mirror region, and the high-frequency representing the details can refine it. Motivated by this, we introduce frequency guidance and propose the dual domain perception progressive refinement network (DPRNet) to mine dual-domain information. Specifically, we first decouple the images into high-frequency and low-frequency components by Laplace pyramid and vision Transformer, respectively, and design the frequency interaction alignment (FIA) module to integrate frequency features to initially localize the mirror region. To handle scale variations, we propose the multi-order feature perception (MOFP) module to adaptively aggregate adjacent features with progressive and gating mechanisms. We further propose the separation-based difference fusion (SDF) module to establish associations between entities and imagings and discover the correct boundary to mine the complete mirror region. Extensive experiments show that DPRNet outperforms the state-of-the-art method by an average of 3% with only about one-fifth of the parameters and FLOPs on four datasets. Our DPRNet also achieves promising performance on remote sensing and camouflage scenarios, validating its generalization. The code is available athttps://github.com/winter-flow/DPRNet.
Mingfeng Zha, Feiyang Fu, Yunqiang Pei, Guoqing Wang 0001, Tianyu Li 0003, Xiongxin Tang, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Circuits Syst. Video Technol.6
2024 Physics-Guided Optical Simulation and PSF Analysis for Remote Sensing Images Deblurring
abstract
The presence of blur is prevalent in satellite remote sensing images (RSIs), and its detrimental impact on downstream applications cannot be overlooked. Current deep learning approaches for image deblurring have gained substantial attention due to their effectiveness and fast inference speed. However, these methods often heavily rely on extensive paired training datasets and lack interpretability. Existing deblurring datasets primarily include regular scenes while remote sensing images exhibit distinct blurring mechanisms. Consequently, deep learning methods lacking prior physical knowledge can only tackle the image deblurring problem in specific scenarios, but hard to achieve satisfactory results on remote sensing images. To address these problems, it is essential to construct a remote sensing image dataset that incorporates the realistic causes of blurriness and integrate prior knowledge into the methods. In this work, we first analyze the satellite imaging system and use Zernike polynomials to approximate the optical aberrations to simulate the RSI blurring process which ensures the proposed dataset adhering solid physical principles. Moreover, we propose a novel physics-guided RSI deblurring (PGRSID) network that integrates an explicit Wiener deconvolution process in both spatial and deep feature space. This integration better leverages the physical interpretation to facilitate effective learning for the RSI deblurring network. We further incorporate denoise loss and cycle consistency loss in the objective function to facilitate the model’s learning process for RSI deblurring. Extensive experiments are conducted on both our synthetic dataset and real GF-1A/PMS data. Qualitative and quantitative experiment results highlight the effectiveness and superiority of our physics-guided deblurring network for satellite RSI.
Fan Ji, Jiangmeng Li, Xiongxin Tang, Fanjiang Xu
IEEE Trans. Geosci. Remote. Sens.5
2023 Model Driven Deep Unfolding Network for Extreme Low-Light Image Enhancement and Denoising
abstract
Low visibility and severe noise are two main degradations in extreme low-light images. Nevertheless, existing low-light image enhancement methods often fail to handle real low-light images with strong noise. To address this issue, We propose a deep unfolding network based on the robust Retinex model with an additional noise term. In particular, we design an optimization model with implicit priors and employ the proximal gradient descent (PGD) technique to alternately solve three iterative sub-problems of the optimization model in a data-driven manner. The proposed method combines the interpretability of model-based methods with the speed and strong fitting ability of learning-based methods. In addition, we collect an extreme low-light sRGB image dataset (E-LOL) containing noisy low/normal-light image pairs. Extensive experimental results demonstrate that our method outperforms state-of-the-art methods in enhancing noisy low-light images and obtains better-exposed illumination, richer colors and textures.
Fanjiang Xu, Xiongxin Tang, Quan Zheng 0004
IJCNN3
2023 A lightweight vision transformer with symmetric modules for vision tasks
abstract
Transformer-based networks have demonstrated their powerful performance in various vision tasks. However, these transformer-based networks are heavyweight and cannot be applied to edge computing (mobile) devices. Despite that the lightweight transformer network has emerged, several problems remain, i.e., weak feature extraction ability, feature redundancy, and lack of convolutional inductive bias. To address these three problems, we propose a lightweight visual transformer (Symmetric Former, SFormer), which contains two novel modules (Symmetric Block and Symmetric FFN). Specifically, we design Symmetric Block to expand feature capacity inside the module and enhance the long-range modeling capability of attention mechanism. To increase the compactness of the model and introduce inductive bias, we introduce convolutional cheap operations to design Symmetric FFN. We compared the SFormer with existing lightweight transformers on several vision tasks. Remarkably, on the image recognition task of ImageNet [13], SFormer gains 1.2% and 1.6% accuracy improvements compared to PVTv2-b0 and Swin Transformer, respectively. On the semantic segmentation task of ADE20K [64], SFormer delivers performance improvements of 0.2% and 0.7% compared to PVTv2-b0 and Swin Transformer, respectively. On the cityscapes dataset [11], SFormer delivers performance improvements of 2.5% and 4.2% compared to PVTv2-b0 and Swin Transformer, respectively. The code is open-source and available at: https://github.com/ISCLab-Bistu/Symmetric_Former.git.
Shengjun Liang, Mingxin Yu, Wenshuai Lu, Xinglong Ji, Xiongxin Tang, Rui You
Intell. Data Anal.5