EDBT 2026 Demo / reviewers in the wild / expert
Guangwei Gao
dblp:118/3484
· DBLP profile ↗
91ranked-venue papers
20as first author
64since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 45 · 10 first-author · 37 since 2021Artificial intelligence and machine learning · 39 · 8 first-author · 24 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Security and privacy · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scene Prior Filtering for Depth Super-Resolution
Zhengxue Wang, Zhiqiang Yan 0001, Ming-Hsuan Yang 0001, Jinshan Pan, Guangwei Gao, Ying Tai, Jian Yang 0003 |
Int. J. Comput. Vis. | 5 |
| 2026 | Correction: Scene Prior Filtering for Depth Super-Resolution
Zhengxue Wang, Zhiqiang Yan 0001, Ming-Hsuan Yang 0001, Jinshan Pan, Guangwei Gao, Ying Tai, Jian Yang 0003 |
Int. J. Comput. Vis. | 5 |
| 2026 | UCPNet: An Ultra-Lightweight Cross-Perception Network for Real-Time Semantic Segmentation
Guoan Xu, Juncheng Li 0003, Heyou Chang, Guangwei Gao |
Image Vis. Comput. | 5 |
| 2026 | Cross-modal hashing under the federated learning framework
Qinghua Huang, Jiahuan Lu, Fei Wu 0004, Guangchuan Peng, Guangwei Gao, Xiaoyuan Jing |
J. Vis. Commun. Image Represent. | 5 |
| 2026 | HazeRes-DFDet: Haze-Resilient Depth-Frequency Detector for foggy drone images
Guangwei Gao, Jiucheng Xie |
Pattern Recognit. | 3 |
| 2026 | INSERTION: From traditional incremental learning to open-world stream learning
Yanchao Li 0001, Hongwei Dou, Guanxiao Li, Guangwei Gao, Huiyu Zhou 0001 |
Pattern Recognit. | 4 |
| 2026 | Noise-tolerant scheme and explicit regularizer for deep active learning with noisy oracles
Yanchao Li 0001, Ziteng Xie, Hongwu Zhong, Guangwei Gao |
Pattern Recognit. | 4 |
| 2026 | Diffusion-based Laplacian frequency-aware network for low-light image enhancement
Juncheng Li 0003, Guangwei Gao, Chia-Wen Lin |
Pattern Recognit. | 4 |
| 2026 | AEA-FIRM: Adaptive Elastic Alignment With Fine-Grained Representation Mining for Text-Based Aerial Pedestrian RetrievalabstractUnmanned aerial vehicles (UAVs) have garnered significant attention due to their operational flexibility, enabling expanded application scenarios across diverse fields. The Text-Based Pedestrian Retrieval (TBPR) task aims to identify corresponding images from textual descriptions, yet existing research has primarily focused on ground-level views. To broaden the applicability of TBPR systems, we introduce aerial-view analysis and propose a novel Text-Based Aerial Pedestrian Retrieval (TBAPR) task. This task introduces unique challenges, particularly the dual gaps in cross-view (aerial vs. ground) and cross-modal (text vs. image) matching, which are more complex than traditional TBPR or aerial-ground pedestrian understanding tasks. To address these challenges, we propose an Adaptive Elastic Alignment Network with FIne-Grained Representation Mining (AEA-FIRM). Our framework tackles the cross-view gap through an AEA loss that adaptively prioritizes critical semantic features while dynamically aligning textual and aerial semantics under challenging conditions. Concurrently, the FIRM module refines visual-linguistic representations by mining fine-grained pedestrian attributes and explicitly textualizing them for cross-modal matching verification. Extensive experiments demonstrate that AEA-FIRM achieves state-of-the-art performance, outperforming existing TBPR methods by 4.87% in Rank-1 accuracy. Our code and dataset are available at https://github.com/xbdxwyh/AEA-FIRM-main.git. Yihao Wang 0010, Meng Yang 0001, Rui Cao 0003, Guangwei Gao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Transformer-Progressive Mamba Network for Lightweight Image Super-ResolutionabstractRecently, Mamba-based super-resolution (SR) methods have demonstrated the ability to capture global receptive fields with linear complexity, addressing the quadratic computational cost of Transformer-based SR approaches. However, existing Mamba-based methods lack fine-grained transitions across different modeling scales, which limits the efficiency of feature representation. In this paper, we propose T-PMambaSR, a lightweight SR framework that integrates window-based self-attention with Progressive Mamba. By enabling interactions among receptive fields of different scales, our method establishes a fine-grained modeling paradigm that progressively enhances feature representation without introducing additional computational cost. Furthermore, we introduce an Adaptive High-Frequency Refinement Module (AHFRM) to recover high-frequency details lost during Transformer and Mamba processing. Extensive experiments demonstrate that T-PMambaSR progressively enhances the model's receptive field and expressiveness, achieving competitive performance with recent Transformer- or Mamba-based methods while incurring lower computational cost. The code is available at https://github.com/IVIPLab/T-PMambaSR. Sichen Guo, Yuanyang Liu, Guangwei Gao, Jian Yang 0003, Chia-Wen Lin |
IEEE Trans. Image Process. | 4 |
| 2026 | SCASeg: Strip Cross-Attention for Efficient Semantic SegmentationabstractThe Vision Transformer (ViT) has achieved notable success in computer vision, with its variants widely validated across various downstream tasks, including semantic segmentation. However, as general-purpose visual encoders, ViT backbones often do not fully address the specific requirements of task decoders, highlighting opportunities for designing decoders optimized for efficient semantic segmentation. This paper proposes Strip Cross-Attention (SCASeg), an innovative decoder head specifically designed for semantic segmentation. Instead of relying on the conventional skip connections, we utilize lateral connections between encoder and decoder stages, leveraging encoder features as Queries in cross-attention modules. Additionally, we introduce a Cross-Layer Block (CLB) that integrates hierarchical feature maps from various encoder and decoder stages to form a unified representation for Keys and Values. The CLB also incorporates the local perceptual strengths of convolution, enabling SCASeg to capture both global and local context dependencies across multiple layers, thus enhancing feature interaction at different scales and improving overall efficiency. To further optimize computational efficiency, SCASeg compresses the channels of queries and keys into one dimension, creating strip-like patterns that reduce memory usage and increase inference speed compared to traditional vanilla cross-attention. Experiments show that SCASeg's adaptable decoder delivers competitive performance across various setups, outperforming leading segmentation architectures on benchmark datasets, including ADE20K, Cityscapes, COCO-Stuff 164k, and Pascal VOC2012, even under diverse computational constraints. Guoan Xu, Jiaming Chen 0001, Wenfeng Huang, Wenjing Jia, Guangwei Gao, Guo-Jun Qi |
IEEE Trans. Image Process. | 5 |
| 2026 | Dual-Domain Modulation Network for Lightweight Image Super-ResolutionabstractLightweight image super-resolution (SR) aims to reconstruct high-resolution images from low-resolution images under limited computational costs. We find existing frequencybased SR methods cannot balance the reconstruction of overall structures and high-frequency parts. Meanwhile, these methods are inefficient for handling frequency features and unsuitable for lightweight SR. In this paper, we show introducing both wavelet and Fourier information allows our model to consider both highfrequency features and overall SR structure reconstruction while reducing costs. Specifically, we propose a Dual-domain Modulation Network that integrates both wavelet and Fourier information for enhanced frequency modeling. Unlike existing methods that rely on a single frequency representation, our design combines wavelet-domain modulation via a Wavelet-domain Modulation Transformer (WMT) with global Fourier supervision, enabling complementary spectral learning well-suited for lightweight SR. Experimental results show that our method achieves a comparable PSNR of SRFormer [1] and MambaIR [2] while with less than 50% and 60% of their FLOPs and achieving inference speeds 15.4× and 5.4× faster, respectively, demonstrating the effectiveness of our method on SR quality and lightweight. Heng Guo 0003, Yuefeng Hou, Guangwei Gao, Zhanyu Ma |
IEEE Trans. Multim. | 4 |
| 2025 | DORNet: A Degradation Oriented and Regularized Network for Blind Depth Super-ResolutionabstractRecent RGB-guided depth super-resolution methods have achieved impressive performance under the assumption of fixed and known degradation (e.g., bicubic downsampling). However, in real-world scenarios, captured depth data often suffer from unconventional and unknown degradation due to sensor limitations and complex imaging environments (e.g., low reflective surfaces, varying illumination). Consequently, the performance of these methods significantly declines when real-world degradation deviate from their assumptions. In this paper, we propose the Degradation Oriented and Regularized Network (DORNet), a novel framework designed to adaptively address unknown degradation in real-world scenes through implicit degradation representations. Our approach begins with the development of a self-supervised degradation learning strategy, which models the degradation representations of low-resolution depth data using routing selection-based degradation regularization. To facilitate effective RGB-D fusion, we further introduce a degradation-oriented feature transformation module that selectively propagates RGB content into the depth data based on the learned degradation priors. Extensive experimental results on both real and synthetic datasets demonstrate the superiority of our DORNet in handling unknown degradation, outperforming existing methods. Zhengxue Wang, Zhiqiang Yan 0001, Jinshan Pan, Guangwei Gao, Kai Zhang 0008, Jian Yang 0003 |
CVPR | 4 |
| 2025 | OCTAMamba: A State-Space Model Approach for Precision OCTA Vasculature SegmentationabstractOptical Coherence Tomography Angiography (OCTA) is a crucial imaging technique for visualizing retinal vasculature and diagnosing eye diseases such as diabetic retinopathy and glaucoma. However, precise segmentation of OCTA vasculature remains challenging due to the multi-scale vessel structures and noise from poor image quality and eye lesions. In this study, we proposed OCTAMamba, a novel U-shaped network based on the Mamba architecture, designed to segment vasculature in OCTA accurately. OCTAMamba integrates a Quad Stream Efficient Mining Embedding Module for local feature extraction, a Multi-Scale Dilated Asymmetric Convolution Module to capture multi-scale vasculature, and a Focused Feature Recalibration Module to filter noise and highlight target areas. Our method achieves efficient global modeling and local feature extraction while maintaining linear complexity, making it suitable for low-computation medical applications. Extensive experiments on the OCTA 3M, OCTA 6M, and ROSSA datasets demonstrated that OCTAMamba outperforms state-of-the-art methods, providing a new reference for efficient OCTA segmentation. Code is available at https://github.com/zs1314/OCTAMamba Shun Zou, Zhuo Zhang 0007, Guangwei Gao |
ICASSP | 3 |
| 2025 | High-Frequency Semantic Enhancement in Compressed Scenarios for Robust Visual and Machine Vision ApplicationsabstractWith the growing demand for video processing in both human and machine vision, optimizing post-processing techniques has become a crucial challenge. To address the limitations of current post-processing techniques in these domains, this paper introduces a novel post-processing method that enhances high-frequency information through semantic enhancement, significantly improving performance in both domains. We propose a High Semantic Extraction (HSE) model to capture more recognizable details, and design a High-Frequency Semantic Fusion (HFSF) strategy that preserves critical details while suppressing noise. Experimental results demonstrate that our method effectively enhances performance in object detection, semantic segmentation, and video quality, achieving a significant advancement in optimizing video processing for both human and machine vision. Keren He, Guangwei Gao, Jinjia Zhou |
ICIP | 3 |
| 2025 | MambaMIC: An Efficient Baseline for Microscopic Image Classification with State Space ModelsabstractIn recent years, CNN and Transformer-based methods have made significant progress in Microscopic Image Classification (MIC). However, existing approaches still face the dilemma between global modeling and efficient computation. While the Selective State Space Model (SSM) can simulate long-range dependencies with linear complexity, it still encounters challenges in MIC, such as local pixel forgetting, channel redundancy, and lack of local perception. To address these issues, we propose a simple yet efficient vision backbone for MIC tasks, named MambaMIC. Specifically, we introduce a Local-Global dual-branch aggregation module: the MambaMIC Block, designed to effectively capture and fuse local connectivity and global dependencies. In the local branch, we use local convolutions to capture pixel similarity, mitigating local pixel forgetting and enhancing perception. In the global branch, SSM extracts global dependencies, while Locally Aware Enhanced Filter reduces channel redundancy and local pixel forgetting. Additionally, we design a Feature Modulation Interaction Aggregation Module for deep feature interaction and key feature re-localization. Extensive benchmarking shows that MambaMIC achieves state-of-the-art performance across five datasets. code is available at https://zs1314.github.io/MambaMIC. Shun Zou, Zhuo Zhang 0007, Guangwei Gao |
ICME | 4 |
| 2025 | Fraesormer: Learning Adaptive Sparse Transformer for Efficient Food RecognitionabstractIn recent years, Transformer has witnessed significant progress in food recognition. However, most existing approaches still face two critical challenges in lightweight food recognition: (1) the quadratic complexity and redundant feature representation from interactions with irrelevant tokens; (2) static feature recognition and single-scale representation, which overlook the unstructured, non-fixed nature of food images and the need for multi-scale features. To address these, we propose an adaptive and efficient sparse Transformer architecture (Fraesormer) with two core designs: Adaptive Top-k Sparse Partial Attention (ATK-SPA) and Hierarchical Scale-Sensitive Feature Gating Network (HSSFGN). ATK-SPA uses a learnable Gated Dynamic Top-K Operator (GDTKO) to retain critical attention scores, filtering low query-key matches that hinder feature aggregation. It also introduces a partial channel mechanism to reduce redundancy and promote expert information flow, enabling local-global collaborative modeling. HSSFGN employs gating mechanism to achieve multi-scale feature representation, enhancing contextual semantic information. Extensive experiments show that Fraesormer outperforms state-of-the-art methods. code is available at https://zs1314.github.io/Fraesormer. Shun Zou, Mingya Zhang, Shipeng Luo, Zhihao Chen 0014, Guangwei Gao |
ICME | 6 |
| 2025 | Learning Dual-Domain Multi-Scale Representations for Single Image DerainingabstractExisting image deraining methods typically rely on single-input, single-output, and single-scale architectures, which overlook the joint multi-scale information between external and internal features. Furthermore, single-domain representations are often too restrictive, limiting their ability to handle the complexities of real-world rain scenarios. To address these challenges, we propose a novel Dual-Domain Multi-Scale Representation Network (DMSR). The key idea is to exploit joint multi-scale representations from both external and internal domains in parallel while leveraging the strengths of both spatial and frequency domains to capture more comprehensive properties. Specifically, our method consists of two main components: the Multi-Scale Progressive Spatial Refinement Module (MPSRM) and the Frequency Domain Scale Mixer (FDSM). The MPSRM enables the interaction and coupling of multi-scale expert information within the internal domain using a hierarchical modulation and fusion strategy. The FDSM extracts multi-scale local information in the spatial domain, while also modeling global dependencies in the frequency domain. Extensive experiments show that our model achieves state-of-the-art performance across six benchmark datasets. Shun Zou, Mingya Zhang, Shipeng Luo, Guangwei Gao, Guo-Jun Qi |
ICME | 5 |
| 2025 | Cross Paradigm Representation and Alignment Transformer for Image DerainingabstractTransformer-based networks have achieved strong performance in low-level vision tasks like image deraining by utilizing spatial or channel-wise self-attention. However, irregular rain patterns and complex geometric overlaps challenge single-paradigm architectures, necessitating a unified framework to integrate complementary global-local and spatial-channel representations. To address this, we propose a novel Cross Paradigm Representation and Alignment Transformer (CPRAformer). Its core idea is the hierarchical representation and alignment, leveraging the strengths of both paradigms (spatial-channel and global-local) to aid image reconstruction. It bridges the gap within and between paradigms, aligning and coordinating them to enable deep interaction and fusion of features. Specifically, we use two types of self-attention in the Transformer blocks: sparse prompt channel self-attention (SPC-SA) and spatial pixel refinement self-attention (SPR-SA). SPC-SA enhances global channel dependencies through dynamic sparsity, while SPR-SA focuses on spatial rain distribution and fine-grained texture recovery. To address the feature misalignment and knowledge differences between them, we introduce the Adaptive Alignment Frequency Module (AAFM), which aligns and interacts with features in a two-stage progressive manner, enabling adaptive guidance and complementarity. This reduces the information gap within and between paradigms. Through this unified cross-paradigm dynamic interaction framework, we achieve the extraction of the most valuable interactive fusion information from the two paradigms. Extensive experiments demonstrate that our model achieves state-of-the-art performance on eight benchmark datasets and further validates CPRAformer's robustness in other image restoration tasks and downstream applications. Shun Zou, Juncheng Li 0003, Guangwei Gao, Guo-Jun Qi |
ACM Multimedia | 4 |
| 2025 | Self-Supervised Selective-Guided Diffusion Model for Old-Photo Face RestorationabstractOld-photo face restoration poses significant challenges due to compounded degradations such as breakage, fading, and severe blur. Existing pre-trained diffusion-guided methods either rely on explicit degradation priors or global statistical guidance, which struggle with localized artifacts or face color. We propose Self-Supervised Selective-Guided Diffusion (SSDiff), which leverages pseudo-reference faces generated by a pre-trained diffusion model under weak guidance. These pseudo-labels exhibit structurally aligned contours and natural colors, enabling region-specific restoration via staged supervision: structural guidance applied throughout the denoising process and color refinement in later steps, aligned with the coarse-to-fine nature of diffusion. By incorporating face parsing maps and scratch masks, our method selectively restores breakage regions while avoiding identity mismatch. We further construct VintageFace, a 300-image benchmark of real old face photos with varying degradation levels. SSDiff outperforms existing GAN-based and diffusion-based methods in perceptual quality, fidelity, and regional controllability. Code link: https://github.com/PRIS-CV/SSDiff. Heng Guo 0003, Guangwei Gao, Zhanyu Ma |
NeurIPS | 4 |
| 2025 | Tri-Perspective View Decomposition for Geometry Aware Depth Completion and Super-ResolutionabstractDepth completion and super-resolution are crucial tasks for comprehensive RGB-D scene understanding, as they involve reconstructing the precise 3D geometry of a scene from sparse or low-resolution depth measurements. However, most existing methods either rely solely on 2D depth representations or directly incorporate raw 3D point clouds for compensation, which are still insufficient to capture the fine-grained 3D geometry of the scene. In this paper, we introduce Tri-Perspective View Decomposition (TPVD) frameworks that can explicitly model 3D geometry. To this end, (1) TPVD ingeniously decomposes the original 3D point cloud into three 2D views, one of which corresponds to the sparse or low-resolution depth input. (2) For sufficient geometric interaction, TPV Fusion is designed to update the 2D TPV features through recurrent 2D-3D-2D aggregation. (3) By adaptively searching for TPV affinitive neighbors, two additional refinement heads are developed for these two tasks to further improve the geometric consistency. Meanwhile, we build novel datasets named TOFDC for depth completion and TOFDSR for depth super-resolution. Both datasets are acquired using time-of-flight (TOF) sensors and color cameras on smartphones. Extensive experiments on TOFDC, KITTI, NYUv2, SUN RGBD, VKITTI, TOFDSR, RGB-D-D, Lu, and Middlebury datasets indicate that our TPVD outperforms previous depth completion and super-resolution methods, reaching the state of the art. Zhiqiang Yan 0001, Kun Wang 0042, Xiang Li 0041, Guangwei Gao, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Balanced Multi-modal Learning with Hierarchical Fusion for Fake News Detection
Fei Wu 0004, Guangwei Gao, Yimu Ji 0001, Xiaoyuan Jing |
Pattern Recognit. | 3 |
| 2025 | IEMFormer: Internal and External Multi-Fusion Transformer for Indoor RGB-D Semantic SegmentationabstractEffectively fusing and complementing RGB and depth modalities while mitigating image noise is a critical challenge in the RGB-D semantic segmentation task. In this paper, we propose a novel Internal and External Multi-fusion Transformer (IEMFormer) to address this issue. IEMFormer incorporates stage-specific fusion strategies to enhance modal complementarity. For internal fusion, we integrate a fusion unit within the traditional Transformer block, combining matching tokens from both modalities on a pixel-by-pixel basis. For external fusion, the proposed External Adaptive Cross-modal Fusion (EACF) module filters dual-modal features across both spatial and channel dimensions, serving the purpose of adaptively weighting complementary channel information and robustly aggregating spatial patterns from both modalities, thereby facilitating the integration of multimodal information. Additionally, the Global Self-attention Guided Fusion (GSGF) module in the decoder refines the fused features from earlier stages, effectively suppressing noise. This is achieved by leveraging high-level semantic features to guide the refinement and incorporating an active noise suppression mechanism to prevent overfitting to dominant, noisy features. Extensive experiments on the NYUv2 and SUN RGB-D datasets demonstrate that IEMFormer achieves highly competitive performance in accurately understanding indoor scenes. Kaidi Hu, Wei Li 0111, Guangwei Gao, Ruigang Yang |
IEEE Signal Process. Lett. | 3 |
| 2025 | AgentPolyp: Accurate Polyp Segmentation via Image Enhancement AgentabstractCaptured polyp images often suffer from degradation, such as dim lighting, blur, and overexposure. Direct segmentation is prone to artifact diffusion, which significantly degrades the performance of downstream segmentation algorithms and leads to inaccurate boundary delineation. Addressing these varied degradations requires a dynamic, intelligent process that diagnoses and applies targeted corrections. We present AgentPolyp, a novel framework driven by an intelligent agent that integrates CLIP-based semantic guidance and dynamic image enhancement with a lightweight segmentation network. The agent adaptively selects reinforcement learning strategies to perform context-aware denoising, contrast adjustment, and artifact reduction. This selection process is continuously optimized through a feedback loop that includes quality assessment, ensuring the optimization and enhancement of downstream segmentation. This approach addresses degradation complexity and feature compatibility issues, offering a deployable solution for endoscopic polyp analysis. Pu Wang 0008, Guangwei Gao, Youshan Zhang, Zhuoran Zheng |
IEEE Signal Process. Lett. | 3 |
| 2025 | HFS-SAM2: Segment Anything Model 2 With High-Frequency Feature Supplementation for Camouflaged Object DetectionabstractCamouflaged Object Detection (COD) aims to identify objects seamlessly blended with their backgrounds. While effective solutions exist for camouflaged animals, detecting camouflaged plants presents unique challenges and remains an open problem. This work introduces a novel Plant Camouflage Detection (PCD) method leveraging the Segment Anything Model 2 (SAM2). Our approach enhances the vanilla SAM2 decoder with specialized frequency-aware modules to improve the performance on PCD. Specifically, we employ a laplacian pyramid to extract high-frequency image components and introduce a High-Frequency Supplementation (HFS) module to enhance crucial spatial details for identifying camouflaged plants. The Multi-Scale Extraction (MSE) module is leveraged to capture rich multi-scale information, after which the features from the last three encoder layers are fused through a Cross-Layer Aggregation (CLA) module to obtain the aggregated high-level semantic features. A Semantic Gap Reduction (SGR) module is further proposed to bridge the semantic gap between high-level and shallow features during fusion. Finally, a Reverse Feature Mining (RFM) module is designed to highlight complementary regions and fine details. Extensive experiments on five datasets, encompassing both plant and animal camouflage detection, demonstrate the superior performance of our method compared to state-of-the-art approaches. Zihuang Wu, Xinyu Xiong, Guangwei Gao, Hongwei Li 0017 |
IEEE Signal Process. Lett. | 3 |
| 2025 | ReviveDiff: A Universal Diffusion Model for Restoring Images in Adverse Weather ConditionsabstractImages captured in challenging environments-such as nighttime, smoke, rainy weather, and underwater-often suffer from significant degradation, resulting in a substantial loss of visual quality. The effective restoration of these degraded images is critical for the subsequent vision tasks. While many existing approaches have successfully incorporated specific priors for individual tasks, these tailored solutions limit their applicability to other degradations. In this work, we propose a universal network architecture, dubbed "ReviveDiff", which can address various degradations and restore images to their original quality by enhancing and restoring their details. Our approach is inspired by the observation that, unlike degradation caused by movement or electronic issues, quality degradation under adverse conditions primarily stems from natural media (such as fog, water, and low luminance), which generally preserves the original structures of objects. To restore the quality of such images, we leveraged the latest advancements in diffusion models and developed ReviveDiff to restore image quality from both macro and micro levels across some key factors determining image quality, such as sharpness, distortion, noise level, dynamic range, and color accuracy. We rigorously evaluated ReviveDiff on seven benchmark datasets covering five types of degrading conditions: Rainy, Underwater, Low-light, Smoke, and Nighttime Hazy. Our experimental results demonstrate that ReviveDiff outperforms the state-of-the-art methods both quantitatively and visually. Wenfeng Huang, Guoan Xu, Wenjing Jia, Stuart W. Perry, Guangwei Gao |
IEEE Trans. Image Process. | 5 |
| 2025 | S2AFormer: Strip Self-Attention for Efficient Vision TransformerabstractThe Vision Transformer (ViT) has achieved remarkable success in computer vision due to its powerful token mixer, which effectively captures global dependencies among all tokens. However, the quadratic complexity of standard self-attention with respect to the number of tokens severely hampers its computational efficiency in practical deployment. Although recent hybrid approaches have sought to combine the strengths of convolutions and self-attention to improve the performance-efficiency trade-off, the costly pairwise token interactions and heavy matrix operations in conventional self-attention remain a critical bottleneck. To overcome this limitation, we introduce S2AFormer, an efficient Vision Transformer architecture built around a novel Strip Self-Attention (SSA) mechanism. Our design incorporates lightweight yet effective Hybrid Perception Blocks (HPBs) that seamlessly fuse the local inductive biases of CNNs with the global modeling capability of Transformer-style attention. The core innovation of SSA lies in simultaneously reducing the spatial resolution of the key ( $K$ ) and value ( $V$ ) tensors while compressing the channel dimension of the query ( $Q$ ) and key ( $K$ ) tensors. This joint spatial-and-channel compression dramatically lowers computational cost without sacrificing representational power, achieving an excellent balance between accuracy and efficiency. We extensively evaluate S2AFormer on a wide range of vision tasks, including image classification (ImageNet-1K), semantic segmentation (ADE20K), and object detection/instance segmentation (COCO). Experimental results consistently show that S2AFormer delivers substantial accuracy improvements together with superior inference speed and throughput across both GPU and non-GPU platforms, establishing it as a highly competitive solution in the landscape of efficient Vision Transformers. Guoan Xu, Wenfeng Huang, Wenjing Jia, Jiamao Li, Guangwei Gao, Guo-Jun Qi |
IEEE Trans. Image Process. | 5 |
| 2025 | Efficient Image Super-Resolution With Feature Interaction Weighted Hybrid NetworkabstractLightweight image super-resolution aims to reconstruct high-resolution images from low-resolution images using low computational costs. However, existing methods result in the loss of middle-layer features due to activation functions. To minimize the impact of intermediate feature loss on reconstruction quality, we propose a Feature Interaction Weighted Hybrid Network (FIWHN), which comprises a series of Wide-residual Distillation Interaction Block (WDIB) as the backbone. Every third WDIB forms a Feature Shuffle Weighted Group (FSWG) by applying mutual information shuffle and fusion. Moreover, to mitigate the negative effects of intermediate feature loss, we introduce Wide Residual Weighting units within WDIB. These units effectively fuse features of varying levels of detail through a Wide-residual Distillation Connection (WRDC) and a Self-Calibrating Fusion (SCF). To compensate for global feature deficiencies, we incorporate a Transformer and explore a novel architecture to combine CNN and Transformer. We show that our FIWHN achieves a favorable balance between performance and efficiency through extensive experiments on low-level and high-level tasks. Juncheng Li 0003, Guangwei Gao, Weihong Deng, Jian Yang 0003, Guo-Jun Qi, Chia-Wen Lin |
IEEE Trans. Multim. | 3 |
| 2025 | Attention-Guided Multiscale Interaction Network for Face Super-ResolutionabstractRecently, CNN and Transformer hybrid networks demonstrated excellent performance in face super-resolution (FSR) tasks. Because of numerous features at different scales in hybrid networks, how to fuse these multiscale features and promote their complementarity is crucial for enhancing FSR. However, existing hybrid network-based FSR methods ignore this, only simply combining the Transformer and CNN. To address this issue, we propose an attention-guided multiscale interaction network (AMINet), which incorporates local and global feature interactions, as well as encoder–decoder phase feature interactions. Specifically, we propose a local and global feature interaction (LGFI) module to promote the fusion of global features and the local features extracted from different receptive fields by our residual depth feature extraction (RDFE) module. Additionally, we propose a selective kernel attention fusion (SKAF) module to adaptively select fusions of different features within the LGFI and encoder–decoder phases. Our above design allows the free flow of multiscale features from within modules and between the encoder and decoder, which can promote the complementarity of different scale features to enhance FSR. Comprehensive experiments confirm that our method consistently performs well with less computational consumption and faster inference. Xujie Wan, Guangwei Gao, Huimin Lu 0001, Jian Yang 0003, Chia-Wen Lin |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | MFPNet: A Multi-scale Feature Propagation Network for Lightweight Semantic Segmentation
Guoan Xu, Wenjing Jia, Ligeng Chen, Guangwei Gao |
ICANN (3) | 5 |
| 2024 | An efficient feature reuse distillation network for lightweight image super-resolution
Chunying Liu, Guangwei Gao |
Comput. Vis. Image Underst. | 2 |
| 2024 | A Decoder Structure Guided CNN-Transformer Network for face super-resolutionabstractAbstract Recent advances in deep convolutional neural networks have shown improved performance in face super‐resolution through joint training with other tasks such as face analysis and landmark prediction. However, these methods have certain limitations. One major limitation is the requirement for manual marking information on the dataset for multi‐task joint learning. This additional marking process increases the computational cost of the network model. Additionally, since prior information is often estimated from low‐quality faces, the obtained guidance information tends to be inaccurate. To address these challenges, a novel Decoder Structure Guided CNN‐Transformer Network (DCTNet) is introduced, which utilises the newly proposed Global‐Local Feature Extraction Unit (GLFEU) for effective embedding. Specifically, the proposed GLFEU is composed of an attention branch and a Transformer branch, to simultaneously restore global facial structure and local texture details. Additionally, a Multi‐Stage Feature Fusion Module is incorporated to fuse features from different network stages, further improving the quality of the restored face images. Compared with previous methods, DCTNet improves Peak Signal‐to‐Noise Ratio by 0.23 and 0.19 dB on the CelebA and Helen datasets, respectively. Experimental results demonstrate that the designed DCTNet offers a simple yet powerful solution to recover detailed facial structures from low‐quality images. Rui Dou, Xujie Wan, Heyou Chang, Guangwei Gao |
IET Comput. Vis. | 6 |
| 2024 | A fully automatic adjacent key-points localization framework for minimal repeated pattern detection in printed fabric images
Qiyan Zang, Jian Zhang 0082, Liling Bo, Guangwei Gao, Heng Zhang 0001, Hongran Li, Zhaoman Zhong |
Knowl. Based Syst. | 5 |
| 2024 | Cycle mapping with adversarial event classification network for fake news detection
Fei Wu 0004, Yujian Feng, Guangwei Gao, Yimu Ji 0001, Xiaoyuan Jing |
Multim. Tools Appl. | 4 |
| 2024 | EWT: Efficient Wavelet-Transformer for single image denoising
Juncheng Li 0003, Bodong Cheng, Guangwei Gao, Jun Shi 0004, Tieyong Zeng |
Neural Networks | 4 |
| 2024 | Efficient Image Classification via Structured Low-Rank Matrix Factorization RegressionabstractIn real-world applications involving sparse coding and low-rank matrix recovery problems, linear regression methods usually struggle to effectively capture the structured correlations present in data matrices. This limitation arises from representation approaches that treat images as vectors and handle testing samples individually, overlooking these correlations. To address these challenges, we propose a novel approach that leverages the low-rank property to capture the global and intrinsic structure of residual and coefficient matrices, departing from the assumption of independent and identically distributed (I.I.D) data. Our method introduces nonconvex and nonsmooth low-rank matrix regression models guided by the extended matrix variate power exponential distribution (M.P.E.D). By incorporating factorization strategies into the regression coefficient matrix and utilizing the Schatten-$p$norm with three distinct values of$p$, we enhance computational efficiency. Our formulation enables efficient subproblem solving through the introduction of auxiliary variables and the use of singular value threshold operators. We achieve closed-form solutions using the proposed multi-variable alternating direction method of multipliers (ADMM). Theoretical analysis establishes the local convergence properties and computational complexity of our optimization algorithm. Furthermore, we conduct numerical experiments on various image datasets, including face, object, and digital, to demonstrate the superior performance and computational efficiency of our methods compared to several related regression approaches. The source codes for our method are available athttps://github.com/ZhangHengMin/TIFS_SLRMFR. Hengmin Zhang, Jian Yang 0003, Jianjun Qian, Guangwei Gao, Xiangyuan Lan, Zhiyuan Zha, Bihan Wen |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | HAFormer: Unleashing the Power of Hierarchy-Aware Features for Lightweight Semantic SegmentationabstractBoth Convolutional Neural Networks (CNNs) and Transformers have shown great success in semantic segmentation tasks. Efforts have been made to integrate CNNs with Transformer models to capture both local and global context interactions. However, there is still room for enhancement, particularly when considering constraints on computational resources. In this paper, we introduce HAFormer, a model that combines the hierarchical features extraction ability of CNNs with the global dependency modeling capability of Transformers to tackle lightweight semantic segmentation challenges. Specifically, we design a Hierarchy-Aware Pixel-Excitation (HAPE) module for adaptive multi-scale local feature extraction. During the global perception modeling, we devise an Efficient Transformer (ET) module streamlining the quadratic calculations associated with traditional Transformers. Moreover, a correlation-weighted Fusion (cwF) module selectively merges diverse feature representations, significantly enhancing predictive accuracy. HAFormer achieves high performance with minimal computational overhead and compact model size, achieving 74.2% mIoU on Cityscapes and 71.1% mIoU on CamVid test datasets, with frame rates of 105FPS and 118FPS on a single 2080Ti GPU. The source codes are available at https://github.com/XU-GITHUB-curry/HAFormer. Guoan Xu, Wenjing Jia, Ligeng Chen, Guangwei Gao |
IEEE Trans. Image Process. | 5 |
| 2024 | Cross-Receptive Focused Inference Network for Lightweight Image Super-ResolutionabstractRecently, Transformer-based methods have shown impressive performance in single image super-resolution (SISR) tasks due to the ability of global feature extraction. However, the capabilities of Transformers that need to incorporate contextual information to extract features dynamically are neglected. To address this issue, we propose a lightweight Cross-receptive Focused Inference Network (CFIN) that consists of a cascade of CT Blocks mixed with CNN and Transformer. Specifically, in the CT block, we first propose a CNN-based Cross-Scale Information Aggregation Module (CIAM) to enable the model to better focus on potentially helpful information to improve the efficiency of the Transformer phase. Then, we design a novel Cross-receptive Field Guided Transformer (CFGT) to enable the selection of contextual information required for reconstruction by using a modulated convolutional kernel that understands the current semantic information and exploits the information interaction within different self-attention. Extensive experiments have shown that our proposed CFIN can effectively reconstruct images using contextual information, and it can strike a good balance between computational cost and model performance as an efficient model. Juncheng Li 0003, Guangwei Gao, Weihong Deng, Jiantao Zhou 0001, Jian Yang 0003, Guo-Jun Qi |
IEEE Trans. Multim. | 3 |
| 2024 | Boundary-Guided Lightweight Semantic Segmentation With Multi-Scale Semantic ContextabstractLightweight semantic segmentation plays an essential role in image signal processing that is beneficial to many multimedia applications, such as self-driving, robotic vision, and virtual reality. Due to the powerful capability to encode image details and semantics, many lightweight dual-resolution networks have been proposed in recent years for semantic segmentation. In spite of achieving remarkable progresses, they often ignore semantic context ranged from different scales. Furthermore, most of them always neglect the object boundaries, serving as a significant assistance for lightweight semantic segmentation. To alleviate these problems, this paper develops a Boundary-guide dual-resolution lightweight network with multi-scale Semantic Context, called BSCNet, for semantic segmentation. Specifically, to enhance the capability of feature representation, an Extremely Lightweight Pyramid Pooling Module (ELPPM) is designed to capture multi-scale semantic context at the top of low-resolution branch of BSCNet. In addition, to increase feature similarity of the same object while keeping feature discrimination of different objects, pixel information is propagated throughout the entire object area using a simple Boundary Auxiliary Fusion Module (BAFM), where the predicted object boundaries are served as high-level guidance to refine low-level convolutional features. The comprehensive experimental results have demonstrated that our BSCNet is simple and effective, achieving state-of-the-art trade-off in terms of segmentation accuracy and running efficiency on CityScapes, CamVid, and KITTI datasets. Quan Zhou 0004, Guangwei Gao, Bin Kang, Weihua Ou, Huimin Lu 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | PFT-SSR: Parallax Fusion Transformer for Stereo Image Super-ResolutionabstractStereo image super-resolution aims to boost the performance of image super-resolution by exploiting the supplementary information provided by binocular systems. Although previous methods have achieved promising results, they did not fully utilize the information of cross-view and intra-view. To further unleash the potential of binocular images, in this letter, we propose a novel Transformer-based parallax fusion module called Parallax Fusion Transformer (PFT). PFT employs a Cross-view Fusion Transformer (CVFT) to utilize cross-view information and an Intra-view Refinement Transformer (IVRT) for intra-view feature refinement. Meanwhile, we adopted the Swin Transformer as the backbone for feature extraction and SR reconstruction to form a pure Transformer architecture called PFT-SSR. Extensive experiments and ablation studies show that PFT-SSR achieves competitive results and outperforms most SOTA methods. Source code is available at https://github.com/MIVRC/PFT-PyTorch. Hansheng Guo, Juncheng Li 0013, Guangwei Gao, Zhi Li 0080, Tieyong Zeng |
ICASSP | 3 |
| 2023 | Semi-supervised cross-modal hashing via modality-specific and cross-modal graph convolutional networks
Fei Wu 0004, Guangwei Gao, Yimu Ji 0001, Xiaoyuan Jing, Zhiguo Wan |
Pattern Recognit. | 3 |
| 2023 | CTCNet: A CNN-Transformer Cooperation Network for Face Image Super-ResolutionabstractRecently, deep convolution neural networks (CNNs) steered face super-resolution methods have achieved great progress in restoring degraded facial details by joint training with facial priors. However, these methods have some obvious limitations. On the one hand, multi-task joint learning requires additional marking on the dataset, and the introduced prior network will significantly increase the computational cost of the model. On the other hand, the limited receptive field of CNN will reduce the fidelity and naturalness of the reconstructed facial images, resulting in suboptimal reconstructed images. In this work, we propose an efficient CNN-Transformer Cooperation Network (CTCNet) for face super-resolution tasks, which uses the multi-scale connected encoder-decoder architecture as the backbone. Specifically, we first devise a novel Local-Global Feature Cooperation Module (LGCM), which is composed of a Facial Structure Attention Unit (FSAU) and a Transformer block, to promote the consistency of local facial detail and global facial structure restoration simultaneously. Then, we design an efficient Feature Refinement Module (FRM) to enhance the encoded features. Finally, to further improve the restoration of fine facial details, we present a Multi-scale Feature Fusion Unit (MFFU) to adaptively fuse the features from different stages in the encoder procedure. Extensive evaluations on various datasets have assessed that the proposed CTCNet can outperform other state-of-the-art methods significantly. Source code will be available at https://github.com/IVIPLab/CTCNet. Guangwei Gao, Zixiang Xu, Juncheng Li 0003, Jian Yang 0003, Tieyong Zeng, Guo-Jun Qi |
IEEE Trans. Image Process. | 1 |
| 2023 | Lightweight Real-Time Semantic Segmentation Network With Efficient Transformer and CNNabstractIn the past decade, convolutional neural networks (CNNs) have shown prominence for semantic segmentation. Although CNN models have very impressive performance, the ability to capture global representation is still insufficient, which results in suboptimal results. Recently, Transformer achieved huge success in NLP tasks, demonstrating its advantages in modeling long-range dependency. Recently, Transformer has also attracted tremendous attention from computer vision researchers who reformulate the image processing tasks as a sequence-to-sequence prediction but resulted in deteriorating local feature details. In this work, we propose a lightweight real-time semantic segmentation network called LETNet. LETNet combines a U-shaped CNN with Transformer effectively in a capsule embedding style to compensate for respective deficiencies. Meanwhile, the elaborately designed Lightweight Dilated Bottleneck (LDB) module and Feature Enhancement (FE) module cultivate a positive impact on training from scratch simultaneously. Extensive experiments performed on challenging datasets demonstrate that LETNet achieves superior performances in accuracy and efficiency balance. Specifically, It only contains 0.95M parameters and 13.6G FLOPs but yields 72.8% mIoU at 120 FPS on the Cityscapes test set and 70.5% mIoU at 250 FPS on the CamVid test dataset using a single RTX 3090 GPU. Source code will be available athttps://github.com/IVIPLab/LETNet. Guoan Xu, Juncheng Li 0003, Guangwei Gao, Huimin Lu 0001, Jian Yang 0003, Dong Yue 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | Occluded Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) aims to match person images between the visible and near-infrared modalities. Previous VI-ReID methods are based on holistic pedestrian images and achieve excellent performance. However, in real-world scenarios, images captured by visible and near-infrared cameras usually contain occlusions. The performance of these methods degrades significantly due to the loss of information of discriminative features from the occlusion of the images. We define visible-infrared person re-identification in this occlusion scene as Occluded VI-ReID, where only partial content information of pedestrian images can be used to match images of different modalities from different cameras. In this paper, we propose a matching framework for occlusion scenes, which contains a local feature enhance module (LFEM) and a modality information fusion module (MIFM). LFEM adopts Transformer to learn features of each modality, and adjusts the importance of patches to enhance the representation ability of local features of the non-occluded areas. MIFM utilizes a co-attention mechanism to infer the correlation between each image for reducing the difference between modalities. We construct two occluded VI-ReID datasets, namely Occluded-SYSU-MM01 and Occluded-RegDB datasets. Our approach outperforms existing state-of-the-art methods on two occlusion datasets, while remains top performance on two holistic datasets. Yujian Feng, Yimu Ji 0001, Fei Wu 0004, Guangwei Gao, Yang Gao 0001, Tianliang Liu, Shangdong Liu, Xiaoyuan Jing, Jiebo Luo 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Context-Patch Representation Learning With Adaptive Neighbor Embedding for Robust Face Image Super-ResolutionabstractRepresentation learning steered robust face image super-resolution (FSR) methods have attracted extensive attention in the past few decades. Most previous methods were devoted to exploiting the local position patches in the training set for FSR. However, they usually overlooked the sufficient usage of the contextual information around the testing patches, which are useful for stable representation learning. In this article, we attempt to utilize the context-patch around the testing patch and propose a method named context-patch representation learning with adaptive neighbor embedding (CRL-ANE) for FSR. On one hand, we simultaneously use the testing position patch and its adjacent ones for stable representation weight learning. This contextual information can compensate for recovering missing details in the target patch. On the other hand, for each input patch set, due to its inherent facial structural properties, we design an adaptive neighbor embedding strategy to elaborately and adaptively choose primary candidates for more accurate reconstruction. These two improvements enable the proposed method to achieve better SR performance than some of the other methods. Qualitative and quantitative experiments on some benchmarks have validated the superiority of the proposed method over some state-of-the-art methods. Guangwei Gao, Yi Yu 0001, Huimin Lu 0001, Jian Yang 0003, Dong Yue 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | JDSR-GAN: Constructing an Efficient Joint Learning Network for Masked Face Super-ResolutionabstractWith the growing importance of preventing the COVID-19 virus in cyber-manufacturing security, face images obtained in most video surveillance scenarios are usually low resolution together with mask occlusion. However, most of the previous face super-resolution solutions can not efficiently handle both tasks in one model. In this work, we consider both tasks simultaneously and construct an efficient joint learning network, called JDSR-GAN, for masked face super-resolution tasks. Given a low-quality face image with mask as input, the role of the generator composed of a denoising module and super-resolution module is to acquire a high-quality high-resolution face image. The discriminator utilizes some carefully designed loss functions to ensure the quality of the recovered face images. Moreover, we incorporate the identity information and attention mechanism into our network for feasible correlated feature expression and informative feature learning. By jointly performing denoising and face super-resolution, the two tasks can complement each other and attain promising performance. Extensive qualitative and quantitative results show the superiority of our proposed JDSR-GAN over some competitive methods. Guangwei Gao, Fei Wu 0004, Huimin Lu 0001, Jian Yang 0003 |
IEEE Trans. Multim. | 1 |
| 2023 | FBSNet: A Fast Bilateral Symmetrical Network for Real-Time Semantic SegmentationabstractReal-time semantic segmentation, which can be visually understood as the pixel-level classification task on the input image, currently has broad application prospects, especially in the fast-developing fields of autonomous driving and drone navigation. However, the huge burden of calculation together with redundant parameters are still the obstacles to its technological development. In this article, we propose a Fast Bilateral Symmetrical Network (FBSNet) to alleviate the above challenges. Specifically, FBSNet employs a symmetrical encoder-decoder structure with two branches, semantic information branch and spatial detail branch. The Semantic Information Branch (SIB) is the main branch with semantic architecture to acquire the contextual information of the input image and meanwhile acquire sufficient receptive field. While the Spatial Detail Branch (SDB) is a shallow and simple network used to establish local dependencies of each pixel for preserving details, which is essential for restoring the original resolution during the decoding phase. Meanwhile, a Feature Aggregation Module (FAM) is designed to effectively combine the output of these two branches. Experimental results of Cityscapes and CamVid show that the proposed FBSNet can strike a good balance between accuracy and efficiency. Specifically, it obtains 70.9% and 68.9% mIoU along with the inference speed of 90 fps and 120 fps on these two test datasets, respectively, with only 0.62 million parameters on a single RTX 2080Ti GPU. The code is available athttps://github.com/IVIPLab/FBSNet. Guangwei Gao, Guoan Xu, Juncheng Li 0003, Yi Yu 0001, Huimin Lu 0001, Jian Yang 0003 |
IEEE Trans. Multim. | 1 |
| 2023 | Lightweight Feature De-redundancy and Self-calibration Network for Efficient Image Super-resolutionabstractIn recent years, thanks to the inherent powerful feature representation and learning abilities of the convolutional neural network (CNN), deep CNN-steered single image super-resolution approaches have achieved remarkable performance improvements. However, these methods are often accompanied by large consumption of computing and memory resources, which is difficult to be adopted in real-world application scenes. To handle this issue, we design an efficient Feature De-redundancy and Self-calibration Super-resolution network (FDSCSR). In particular, a Feature De-redundancy and Self-calibration Block (FDSCB) is proposed to reduce the repetitive feature information extracted by the model and further enhance the efficiency of the model. Then, based on FDSCB, a Local Feature Fusion Module is presented to elaborately utilize and fuse the feature information extracted by each FDSCB. Abundant experiments on benchmarks have demonstrated that our FDSCSR achieves superior performance with relatively less computational consumption and storage resource than other state-of-the-art approaches. The code is available at https://github.com/IVIPLab/FDSCSR . Zhengxue Wang, Guangwei Gao, Juncheng Li 0003, Huimin Lu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Feature Distillation Interaction Weighting Network for Lightweight Image Super-resolutionabstractConvolutional neural networks based single-image superresolution (SISR) has made great progress in recent years. However, it is difficult to apply these methods to real-world scenarios due to the computational and memory cost. Meanwhile, how to take full advantage of the intermediate features under the constraints of limited parameters and calculations is also a huge challenge. To alleviate these issues, we propose a lightweight yet efficient Feature Distillation Interaction Weighted Network (FDIWN). Specifically, FDIWN utilizes a series of specially designed Feature Shuffle Weighted Groups (FSWG) as the backbone, and several novel mutual Wide-residual Distillation Interaction Blocks (WDIB) form an FSWG. In addition, Wide Identical Residual Weighting (WIRW) units and Wide Convolutional Residual Weighting (WCRW) units are introduced into WDIB for better feature distillation. Moreover, a Wide-Residual Distillation Connection (WRDC) framework and a Self-Calibration Fusion (SCF) unit are proposed to interact features with different scales more flexibly and efficiently. Extensive experiments show that our FDIWN is superior to other models to strike a good balance between model performance and efficiency. The code is available at https://github.com/IVIPLab/FDIWN. Guangwei Gao, Juncheng Li 0003, Fei Wu 0004, Huimin Lu 0001, Yi Yu 0001 |
AAAI | 1 |
| 2022 | Cross-Resolution Person Re-Identification via Deep Group-Aware Representation LearningabstractPerson re-identification (Re-ID) aims to identify the same person from samples taken by different cameras. However, in practical application scenarios, due to the quality of the camera equipment and the distance between the camera and the pedestrian, the captured pedestrian images usually have different resolutions, which will cause the mismatch problem of person Re-ID. To mitigate the resolution discrepancy issue, in this paper, we propose a method called deep group-aware representation learning (DGRL) for effective Re-ID. Firstly, We use the residual Transformer block in the feature extraction stage to thoroughly extract richer local and global information from variable resolution shallow images. Then our proposed multi-layered group-aware representation (MGAR) scheme can generate diverse representations different from the main branch, thereby improving the representation capability of the deeply embedded features. In addition, we calculate the kullback leibler divergence loss (KLDivLoss) values on the probability prediction outputs of any two branches, forcing the entire network to be well optimized. Plenty of evaluations on four benchmark datasets have demonstrated the effectiveness of our method. Xiang Ye, Guangwei Gao |
ICPR | 2 |
| 2022 | Lightweight Bimodal Network for Single-Image Super-Resolution via Symmetric CNN and Recursive TransformerabstractSingle-image super-resolution (SISR) has achieved significant breakthroughs with the development of deep learning. However, these methods are difficult to be applied in real-world scenarios since they are inevitably accompanied by the problems of computational and memory costs caused by the complex operations. To solve this issue, we propose a Lightweight Bimodal Network (LBNet) for SISR. Specifically, an effective Symmetric CNN is designed for local feature extraction and coarse image reconstruction. Meanwhile, we propose a Recursive Transformer to fully learn the long-term dependence of images thus the global information can be fully used to further refine texture details. Studies show that the hybrid of CNN and Transformer can build a more efficient model. Extensive experiments have proved that our LBNet achieves more prominent performance than other state-of-the-art methods with a relatively low computational cost and memory consumption. The code is available at https://github.com/IVIPLab/LBNet. Guangwei Gao, Zhengxue Wang, Juncheng Li 0003, Yi Yu 0001, Tieyong Zeng |
IJCAI | 1 |
| 2022 | Towards end-to-end container code recognition
Guangwei Gao |
Multim. Tools Appl. | 3 |
| 2022 | Dual-aligned unsupervised domain adaptation with graph convolutional networks
Fei Wu 0004, Pengfei Wei 0001, Guangwei Gao, Changhui Hu 0001, Qi Ge, Xiaoyuan Jing |
Multim. Tools Appl. | 3 |
| 2022 | JSPNet: Learning joint semantic & instance segmentation of point clouds via feature self-similarity and cross-task probability
Feng Chen 0047, Fei Wu 0004, Guangwei Gao, Yimu Ji 0001, Guoping Jiang, Xiaoyuan Jing |
Pattern Recognit. | 3 |
| 2022 | Multi-feature sparse similar representation for person identification
Meng Yang 0001, Kangyin Ke, Guangwei Gao |
Pattern Recognit. | 4 |
| 2022 | Hierarchical Deep CNN Feature Set-Based Representation Learning for Robust Cross-Resolution Face RecognitionabstractCross-resolution face recognition (CRFR), which is important in intelligent surveillance and biometric forensics, refers to the problem of matching a low-resolution (LR) probe face image against high-resolution (HR) gallery face images. Existing shallow learning-based and deep learning-based methods focus on mapping the HR-LR face pairs into a joint feature space where the resolution discrepancy is mitigated. However, little works consider how to extract and utilize the intermediate discriminative features from the noisy LR query faces to further mitigate the resolution discrepancy due to the resolution limitations. In this study, we desire to fully exploit the multi-level deep convolutional neural network (CNN) feature set for robust CRFR. In particular, our contributions are threefold. (i) To learn more robust and discriminative features, we desire to adaptively fuse the contextual features from different layers. (ii) To fully exploit these contextual features, we design a feature set-based representation learning (FSRL) scheme to collaboratively represent the hierarchical features for more accurate recognition. Moreover, FSRL utilizes the primitive form of feature maps to keep the latent structural information, especially in noisy cases. (iii) To further promote the recognition performance, we desire to fuse the hierarchical recognition outputs from different stages. Meanwhile, the discriminability from different scales can also be fully integrated. By exploiting these advantages, the efficiency of the proposed method can be delivered. Experimental results on several face datasets have verified the superiority of the presented algorithm to the other competitive CRFR approaches. Guangwei Gao, Yi Yu 0001, Jian Yang 0003, Guo-Jun Qi, Meng Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | MSCFNet: A Lightweight Network With Multi-Scale Context Fusion for Real-Time Semantic SegmentationabstractIn recent years, how to strike a good trade-off between accuracy, inference speed, and model size has become the core issue for real-time semantic segmentation applications, which plays a vital role in real-world scenarios such as autonomous driving systems and drones. In this study, we devise a novel lightweight network using a multi-scale context fusion (MSCFNet) scheme, which explores an asymmetric encoder-decoder architecture to alleviate these problems. More specifically, the encoder adopts some developed efficient asymmetric residual (EAR) modules, which are composed of factorization depth-wise convolution and dilation convolution. Meanwhile, instead of complicated computation, simple deconvolution is applied in the decoder to further reduce the amount of parameters while still maintaining the high segmentation accuracy. Also, MSCFNet has branches with efficient attention modules from different stages of the network to well capture multi-scale contextual information. Then we combine them before the final classification to enhance the expression of the features and improve the segmentation efficiency. Comprehensive experiments on challenging datasets have demonstrated that the proposed MSCFNet, which contains only 1.15M parameters, achieves 71.9% Mean IoU on the Cityscapes testing dataset and can run at over 50 FPS on a single Titan XP GPU configuration. Guangwei Gao, Guoan Xu, Yi Yu 0001, Jin Xie 0001, Jian Yang 0003, Dong Yue 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Leaning compact and representative features for cross-modality person re-identification
Guangwei Gao, Hao Shao, Fei Wu 0004, Meng Yang 0001, Yi Yu 0001 |
World Wide Web | 1 |
| 2021 | Lightweight Image Super-Resolution with Multi-Scale Feature Interaction NetworkabstractRecently, the single image super-resolution (SISR) approaches with deep and complex convolutional neural network structures have achieved promising performance. However, those methods improve the performance at the cost of higher memory consumption, which is difficult to be applied for some mobile devices with limited storage and computing resources. To solve this problem, we present a lightweight multi-scale feature interaction network (MSFIN). For lightweight SISR, MSFIN expands the receptive field and adequately exploits the informative features of the low-resolution observed images from various scales and interactive connections. In addition, we design a lightweight recurrent residual channel attention block (RRCAB) so that the network can benefit from the channel attention mechanism while being sufficiently lightweight. Extensive experiments on some benchmarks have confirmed that our proposed MSFIN can achieve comparable performance against the state-of-the-arts with a more lightweight model. Zhengxue Wang, Guangwei Gao, Juncheng Li 0003, Yi Yu 0001, Huimin Lu 0001 |
ICME | 2 |
| 2021 | Adaptive deformable convolutional network
Feng Chen 0047, Fei Wu 0004, Guangwei Gao, Qi Ge, Xiaoyuan Jing |
Neurocomputing | 4 |
| 2021 | Constructing multilayer locality-constrained matrix regression framework for noise robust face super-resolution
Guangwei Gao, Yi Yu 0001, Jin Xie 0001, Jian Yang 0003, Meng Yang 0001, Jian Zhang 0002 |
Pattern Recognit. | 1 |
| 2021 | Short-term Load Forecasting by Using Improved GEP and Abnormal Load RecognitionabstractLoad forecasting in short term is very important to economic dispatch and safety assessment of power system. Although existing load forecasting in short-term algorithms have reached required forecast accuracy, most of the forecasting models are black boxes and cannot be constructed to display mathematical models. At the same time, because of the abnormal load caused by the failure of the load data collection device, time synchronization, and malicious tampering, the accuracy of the existing load forecasting models is greatly reduced. To address these problems, this article proposes a Short-Term Load Forecasting algorithm by using Improved Gene Expression Programming and Abnormal Load Recognition (STLF-IGEP_ALR). First, the Recognition algorithm of Abnormal Load based on Probability Distribution and Cross Validation is proposed. By analyzing the probability distribution of rows and columns in load data, and using the probability distribution of rows and columns for cross-validation, misjudgment of normal load in abnormal load data can be better solved. Second, by designing strategies for adaptive generation of population parameters, individual evolution of populations and dynamic adjustment of genetic operation probability, an Improved Gene Expression Programming based on Evolutionary Parameter Optimization is proposed. Finally, the experimental results on two real load datasets and one open load dataset show that compared with the existing abnormal data detection algorithms, the algorithm proposed in this article have higher advantages in missing detection rate, false detection rate and precision rate, and STLF-IGEP_ALR is superior to other short-term load forecasting algorithms in terms of the convergence speed, MAE, MAPE, RSME, and R 2 . Song Deng, Fulin Chen, Xia Dong, Guangwei Gao, Xindong Wu 0001 |
ACM Trans. Internet Techn. | 4 |
| 2021 | Robust Facial Image Super-Resolution by Kernel Locality-Constrained Coupled-Layer RegressionabstractSuper-resolution methods for facial image via representation learning scheme have become very effective methods due to their efficiency. The key problem for the super-resolution of facial image is to reveal the latent relationship between the low-resolution ( LR ) and the corresponding high-resolution ( HR ) training patch pairs. To simultaneously utilize the contextual information of the target position and the manifold structure of the primitive HR space, in this work, we design a robust context-patch facial image super-resolution scheme via a kernel locality-constrained coupled-layer regression (KLC2LR) scheme to obtain the desired HR version from the acquired LR image. Here, KLC2LR proposes to acquire contextual surrounding patches to represent the target patch and adds an HR layer constraint to compensate the detail information. Additionally, KLC2LR desires to acquire more high-frequency information by searching for nearest neighbors in the HR sample space. We also utilize kernel function to map features in original low-dimensional space into a high-dimensional one to obtain potential nonlinear characteristics. Our compared experiments in the noisy and noiseless cases have verified that our suggested methodology performs better than many existing predominant facial image super-resolution methods. Guangwei Gao, Huimin Lu 0001, Yi Yu 0001, Heyou Chang, Dong Yue 0001 |
ACM Trans. Internet Techn. | 1 |
| 2021 | Chinese Image Captioning via Fuzzy Attention-based DenseNet-BiLSTMabstractChinese image description generation tasks usually have some challenges, such as single-feature extraction, lack of global information, and lack of detailed description of the image content. To address these limitations, we propose a fuzzy attention-based DenseNet-BiLSTM Chinese image captioning method in this article. In the proposed method, we first improve the densely connected network to extract features of the image at different scales and to enhance the model’s ability to capture the weak features. At the same time, a bidirectional LSTM is used as the decoder to enhance the use of context information. The introduction of an improved fuzzy attention mechanism effectively improves the problem of correspondence between image features and contextual information. We conduct experiments on the AI Challenger dataset to evaluate the performance of the model. The results show that compared with other models, our proposed model achieves higher scores in objective quantitative evaluation indicators, including BLEU , BLEU , METEOR, ROUGEl, and CIDEr. The generated description sentence can accurately express the image content. Huimin Lu 0001, Rui Yang 0018, Zhenrong Deng, Yonglin Zhang, Guangwei Gao, Rushi Lan |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2020 | Cross-View Image Synthesis with Deformable Convolution and Attention Mechanism
Songsong Wu, Hao Tang 0005, Fei Wu 0004, Guangwei Gao, Xiaoyuan Jing |
PRCV (1) | 5 |
| 2020 | Cross-resolution face recognition with pose variations via multilayer locality-constrained structural orthogonal procrustes regression
Guangwei Gao, Yi Yu 0001, Meng Yang 0001, Heyou Chang, Dong Yue 0001 |
Inf. Sci. | 1 |
| 2020 | Semi-supervised Dual-Branch Network for image classification
Jiaming Chen 0008, Meng Yang 0001, Guangwei Gao |
Knowl. Based Syst. | 3 |
| 2020 | Local and global aligned spatiotemporal attention network for video-based person re-identification
Li Cheng 0006, Xiaoyuan Jing, Xiaoke Zhu, Changhui Hu 0001, Guangwei Gao, Songsong Wu |
Multim. Tools Appl. | 5 |
| 2020 | A novel biclustering of gene expression data based on hybrid BAFS-BSA algorithm
Yan Cui 0007, Huacheng Gao, Yinqiu Liu, Guangwei Gao |
Multim. Tools Appl. | 6 |
| 2020 | Locality-constrained feature space learning for cross-resolution sketch-photo face recognition
Guangwei Gao, Yannan Wang, Heyou Chang, Huimin Lu 0001, Dong Yue 0001 |
Multim. Tools Appl. | 1 |
| 2020 | Face image super-resolution with pose via nuclear norm regularized structural orthogonal Procrustes regression
Guangwei Gao, Meng Yang 0001, Huimin Lu 0001, Wankou Yang, Hao Gao 0005 |
Neural Comput. Appl. | 1 |
| 2020 | Unsupervised visual domain adaptation via discriminative dictionary evolution
Songsong Wu, Guangwei Gao, Fei Wu 0004, Xiaoyuan Jing |
Pattern Anal. Appl. | 2 |
| 2020 | Adaptive Convolution Local and Global Learning for Class-Level Joint Representation of Facial Recognition With a Single Sample Per Data SubjectabstractDue to the absence of training samples and intraclass variation, the extraction of discriminative facial features and construction of powerful classifiers have bottlenecks in improving the performance of facial recognition (FR) with a single sample per data subject (SSPDS). In this paper, we propose to learn regional adaptive convolution features that are locally and globally discriminative to facial identity and robust to facial variation. Then, a novel class-level joint representation framework is presented to exploit the distinctiveness and class-level commonality of different facial features. In the proposed class-level joint representation with regional adaptive convolution features (CJR-RACF), both discriminative facial features that are robust to facial variations and powerful representations for classification with generic facial variations have been fully exploited. Furthermore, the gallery discrimination is extracted by our proposed weight-embedded supervision in the training phase (denoted by CJR-RACFw), which is conducive to more specific features for FR with SSPDS. CJR-RACF and CJR-RACFw have been evaluated on several popular databases, including the large-scale CMU Multi-PIE, LFW, Megaface, and VGGFace datasets. Experimental results demonstrate the much higher robustness and effectiveness of the proposed methods compared to the state-of-the-art methods. Meng Yang 0001, Xing Wang 0012, LinLin Shen, Guangwei Gao |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2019 | Lednet: A Lightweight Encoder-Decoder Network for Real-Time Semantic SegmentationabstractThe extensive computational burden limits the usage of CNNs in mobile devices for dense estimation tasks. In this paper, we present a lightweight network to address this problem, namely LEDNet, which employs an asymmetric encoder-decoder architecture for the task of real-time semantic segmentation. More specifically, the encoder adopts a ResNet as backbone network, where two new operations, channel split and shuffle, are utilized in each residual block to greatly reduce computation cost while maintaining higher segmentation accuracy. On the other hand, an attention pyramid network (APN) is employed in the decoder to further lighten the entire network complexity. Our model has less than 1M parameters, and is able to run at over 71 FPS in a single GTX 1080Ti GPU. The comprehensive experiments demonstrate that our approach achieves state-of-the-art results in terms of speed and accuracy trade-off on CityScapes dataset. Yu Wang 0109, Quan Zhou 0004, Jian Xiong 0005, Guangwei Gao, Xiaofu Wu, Longin Jan Latecki |
ICIP | 5 |
| 2019 | Fuzzy Bilinear Latent Canonical Correlation Projection for Feature Learning
Yun-Hao Yuan 0001, Yun Li 0010, Jipeng Qiang, Jianping Gou, Guangwei Gao, Bin Li 0006 |
ICONIP (1) | 6 |
| 2019 | Feature extraction based on graph discriminant embedding and its applications to face recognition
Guangwei Gao |
Soft Comput. | 3 |
| 2019 | Multi-scale deep context convolutional neural networks for semantic segmentation
Quan Zhou 0004, Guangwei Gao, Weihua Ou, Huimin Lu 0001, Longin Jan Latecki |
World Wide Web | 3 |
| 2018 | Locality-regularized linear regression discriminant analysis for feature extraction
Zhenqiu Shu, Guangwei Gao, Chengshan Qian |
Inf. Sci. | 4 |
| 2018 | "Like charges repulsion and opposite charges attraction" law based multilinear subspace analysis for face recognition
Fei Wu 0004, Xiaoyuan Jing, Songsong Wu, Guangwei Gao, Qi Ge, Ruchuan Wang 0001 |
Knowl. Based Syst. | 4 |
| 2018 | Semi-Supervised Cross-View Projection-Based Dictionary Learning for Video-Based Person Re-IdentificationabstractVideo-based person re-identification (re-id) has attracted a lot of research interest. When facing dramatic growth in new pedestrian videos, existing video-based person re-id methods usually need large quantities of labeled pedestrian videos to train a discriminative model. In practice, labeling large quantities of pedestrian videos is a costly and time-consuming task, which will limit the application of these methods in the real environment. Therefore, it is valuable and necessary to investigate how to learn a discriminative re-id model by using limited labeled training pedestrian videos. In this paper, we propose a semi-supervised cross-view projection-based dictionary learning (SCPDL) approach for video-based person re-id. Specifically, SCPDL jointly learns a pair of feature projection matrices and a pair of dictionaries by integrating the information contained in labeled and unlabeled pedestrian videos. With the learned feature projection matrices, the influence of variations within each video to the re-id can be reduced. With the learned dictionary pair, pedestrian videos from two different cameras can be converted into coding coefficients in a common representation space, such that the differences between different cameras can be bridged. In the learning process, the labeled pedestrian videos are used to ensure that the learned dictionaries have favorable discriminability; the large quantities of unlabeled pedestrian videos are used to ensure that SCPDL can better capture the variations between pedestrian videos, such that the learned dictionaries can own stronger representative capability. Experiments on two public pedestrian sequence data sets (iLIDS-VID and PRID 2011) demonstrate the effectiveness of the proposed approach. Xiaoke Zhu, Xiaoyuan Jing, Liang Yang 0002, Xinge You, Dan Chen 0001, Guangwei Gao, Yunhong Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2017 | Locality-Constrained Iterative Matrix Regression for Robust Face Hallucination
Guangwei Gao, Huijuan Pang, Cailing Wang, Dong Yue 0001 |
ICONIP (3) | 1 |
| 2017 | A Method of Pedestrian Re-identification Based on Multiple Saliency Features
Cailing Wang, Yechao Xu, Guangwei Gao, Song Tang 0001, Xiaoyuan Jing |
ICONIP (6) | 3 |
| 2017 | Learning robust and discriminative low-rank representations for face recognition with occlusion
Guangwei Gao, Jian Yang 0003, Xiaoyuan Jing, Fumin Shen, Wankou Yang, Dong Yue 0001 |
Pattern Recognit. | 1 |
| 2016 | Locality-constrained matrix regression for position-patch based face hallucinationabstractPosition-patch based face hallucination approaches have been proposed to replace the probabilistic graph-based or manifold learning-based models recently. In this paper, we propose a novel position-based face hallucination method based on locality-constrained matrix regression (LcMR). LcMR uses nuclear norm to characterize the reconstruction error straightforward, thus preserving the essential structural information of the input. On the other hand, LcMR imposes a locality constraint onto the combination coefficients to reach sparsity and locality simultaneously. The locality constraint can derive an analytical solution to the optimization problem. Moreover, LcMR can be solved using alternating direction method of multipliers. Experimental results demonstrate the superiority of the proposed method over some state-of-the-art approaches. Guangwei Gao, Xiaoyuan Jing, Quan Zhou 0004, Songsong Wu, Dong Yue 0001 |
ICIP | 1 |
| 2016 | Parameterless reconstructive discriminant analysis for feature extraction
Guangwei Gao |
Neurocomputing | 2 |
| 2016 | Regional deep learning model for visual tracking
Guoxing Wu, Guangwei Gao, Chunxia Zhao |
Neurocomputing | 3 |
| 2016 | Binary code learning via optimal class representations
Xiang Zhou 0008, Fumin Shen, Yang Yang 0002, Guangwei Gao, Yuan Wang 0003 |
Neurocomputing | 4 |
| 2014 | A novel sparse representation based framework for face image super-resolution
Guangwei Gao, Jian Yang 0003 |
Neurocomputing | 1 |
| 2014 | Integration of multiple orientation and texture information for finger-knuckle-print verification
Guangwei Gao, Jian Yang 0003, Jianjun Qian, Lin Zhang 0014 |
Neurocomputing | 1 |
| 2013 | Discriminative histograms of local dominant orientation (D-HLDO) for biometric image feature extraction
Jianjun Qian, Jian Yang 0003, Guangwei Gao |
Pattern Recognit. | 3 |
| 2013 | Reconstruction Based Finger-Knuckle-Print Verification With Score Level Adaptive Binary FusionabstractRecently, a new biometrics identifier, namely finger knuckle print (FKP), has been proposed for personal authentication with very interesting results. One of the advantages of FKP verification lies in its user friendliness in data collection. However, the user flexibility in positioning fingers also leads to a certain degree of pose variations in the collected query FKP images. The widely used Gabor filtering based competitive coding scheme is sensitive to such variations, resulting in many false rejections. We propose to alleviate this problem by reconstructing the query sample with a dictionary learned from the template samples in the gallery set. The reconstructed FKP image can reduce much the enlarged matching distance caused by finger pose variations; however, both the intra-class and inter-class distances will be reduced. We then propose a score level adaptive binary fusion rule to adaptively fuse the matching distances before and after reconstruction, aiming to reduce the false rejections without increasing much the false acceptances. Experimental results on the benchmark PolyU FKP database show that the proposed method significantly improves the FKP verification accuracy. Guangwei Gao, Lei Zhang 0006, Jian Yang 0003, Lin Zhang 0014, David Zhang 0001 |
IEEE Trans. Image Process. | 1 |