Guangqian Guo

dblp:260/3364 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-8940-1382ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Boosting Segment Anything Model to Generalize Visually Non-Salient Scenarios
abstract
Segment Anything Model (SAM), known for its remarkable zero-shot segmentation capabilities, has garnered significant attention in the community. Nevertheless, its performance is challenged when dealing with what we refer to as visually non-salient scenarios, where there is low contrast between the foreground and background. In these cases, existing methods often cannot capture accurate contours and fail to produce promising segmentation results. In this paper, we propose Visually Non-Salient SAM (VNS-SAM), aiming to enhance SAM's perception of visually non-salient scenarios while preserving its original zero-shot generalizability. We achieve this by effectively exploiting SAM's low-level features through two designs: Mask-Edge Token Interactive decoder and Non-Salient Feature Mining module. These designs help the SAM decoder gain a deeper understanding of non-salient characteristics with only marginal parameter increments and computational requirements. The additional parameters of VNS-SAM can be optimized within 4 hours, demonstrating its feasibility and practicality. In terms of data, we established VNS-SEG, a unified dataset for various VNS scenarios, with more than 35K images, in contrast to previous single-task adaptations. It is designed to make the model learn more robust VNS features and comprehensively benchmark the model's segmentation performance and generalizability on VNS scenarios. Extensive experiments across various VNS segmentation tasks demonstrate the superior performance of VNS-SAM, particularly under zero-shot settings, highlighting its potential for broad real-world applications. Codes and datasets are publicly available at https://guangqian-guo.github.io/VNS-SAM/.
Guangqian Guo, Pengfei Chen 0004, Boqiang Zhang, Shan Gao 0003
IEEE Trans. Image Process.1
2026 ThermalGate-GS: Frequency-Gated Graph Splatting for Thermal Novel View Synthesis
abstract
Thermal infrared imaging is pivotal for all-weather 3D perception, yet analyzing thermal information remains a formidable challenge due to the complexity of heat conduction. Unlike visible light, heat conduction acts as a natural low-pass filter that suppresses high-frequency textural details, causing severe geometric ambiguities and "ghosting" artifacts in standard 3D reconstruction pipelines. To accurately model the inherently diffusive thermal field for high-fidelity reconstruction, we propose Frequency-Gated Graph Splatting (ThermalGate-GS), a framework that explicitly decouples the scene into diffusive thermal distributions (low-frequency) and sharp structural boundaries (high-frequency). Within this framework, we introduce a novel Frequency-Gated Anisotropic Diffusion mechanism. Specifically, the frequency-gating module utilizes extracted high-frequency structural cues to determine spatially-adaptive gating weights. Subsequently, these weights drive an anisotropic diffusion process that dynamically regulates thermal feature propagation, promoting smoothness on object surfaces while suppressing cross-boundary bleeding. Finally, these spectrally refined features are employed to regress 3D Gaussian attributes, substantially alleviating the ambiguity in thermal reconstruction. Extensive experiments demonstrate that ThermalGate-GS achieves state-of-the-art performance, with a notable 7.94 dB PSNR improvement on the ThermoScenes benchmark over prior physics-inspired baselines.
Yaoxing Wang, Wenkang Chen, Guangqian Guo, Chaowei Wang, Yan Di, Shan Gao 0003
IEEE Trans. Image Process.4
2025 Segment Any-Quality Images with Generative Latent Space Enhancement
abstract
Despite their success, Segment Anything Models (SAMs) experience significant performance drops on severely degraded, low-quality images, limiting their effectiveness in real-world scenarios. To address this, we propose Gle-SAM, which utilizes Generative Latent space Enhancement to boost robustness on low-quality images, thus enabling generalization across various image qualities. Specifically, we adapt the concept of latent diffusion to SAM-based segmentation frameworks and perform the generative diffusion process in the latent space of SAM to reconstruct high-quality representation, thereby improving segmentation. Additionally, we introduce two techniques to improve compatibility between the pre-trained diffusion model and the segmentation framework. Our method can be applied to pre-trained SAM and SAM2 with only minimal additional learnable parameters, allowing for efficient optimization. We also construct the LQSeg dataset with a greater diversity of degradation types and levels for training and evaluating the model. Extensive experiments demonstrate that GleSAM significantly improves segmentation robustness on complex degradations while maintaining generalization to clear images. Furthermore, GleSAM also performs well on unseen degradations, underscoring the versatility of our approach and dataset.
Guangqian Guo, Xuehui Yu, Yaoxing Wang, Shan Gao 0003
CVPR1
2025 SAM-COD+: SAM-Guided Unified Framework for Weakly-Supervised Camouflaged Object Detection
abstract
Most Camouflaged Object Detection (COD) methods heavily rely on mask annotations, which are time-consuming and labor-intensive to acquire. Existing weakly-supervised COD approaches exhibit significantly inferior performance compared to fully-supervised methods and struggle to simultaneously support all the existing types of camouflaged object labels, including scribbles, bounding boxes, and points. Even for Segment Anything Model (SAM), it is still problematic to handle the weakly-supervised COD and it typically encounters challenges of prompt compatibility of the scribble labels, extreme response, semantically erroneous response, and unstable feature representations, producing unsatisfactory results in camouflaged scenes. To mitigate these issues, we propose a unified COD framework in this paper, termed SAM-COD, which is capable of supporting arbitrary weakly-supervised labels. Our SAM-COD employs a prompt adapter to handle scribbles as prompts based on SAM. Meanwhile, we introduce response filter and semantic matcher modules to improve the quality of the masks obtained by SAM under COD prompts. To alleviate the negative impacts of inaccurate mask predictions, a new strategy of prompt-adaptive knowledge distillation is utilized to ensure a reliable feature representation. To validate the effectiveness of our approach, we have conducted extensive empirical experiments on three mainstream COD benchmarks. The results demonstrate the superiority of our method against state-of-the-art weakly-supervised and even fully-supervised methods. Our source codes and trained models will be publicly released.
Pengxu Wei, Guangqian Guo, Shan Gao 0003
IEEE Trans. Circuits Syst. Video Technol.3
2025 Go Deep or Broad? Exploit Hybrid Network Architecture for Weakly Supervised Object Classification and Localization
abstract
Weakly supervised object classification and localization are learned object classes and locations using only image-level labels, as opposed to bounding box annotations. Conventional deep convolutional neural network (CNN)-based methods activate the most discriminate part of an object in feature maps and then attempt to expand feature activation to the whole object, which leads to deteriorating the classification performance. In addition, those methods only use the most semantic information in the last feature map, while ignoring the role of shallow features. So, it remains a challenge to enhance classification and localization performance with a single frame. In this article, we propose a novel hybrid network, namely deep and broad hybrid network (DB-HybridNet), which combines deep CNNs with a broad learning network to learn discriminative and complementary features from different layers, and then integrates multilevel features (i.e., high-level semantic features and low-level edge features) in a global feature augmentation module. Importantly, we exploit different combinations of deep features and broad learning layers in DB-HybridNet and design an iterative training algorithm based on gradient descent to ensure the hybrid network work in an end-to-end framework. Through extensive experiments on caltech-UCSD birds (CUB)-200 and imagenet large scale visual recognition challenge (ILSVRC) 2016 datasets, we achieve state-of-the-art classification and localization performance.
Shan Gao 0003, Guangqian Guo, Hanqiao Huang, C. L. Philip Chen
IEEE Trans. Neural Networks Learn. Syst.2
2024 Just a Hint: Point-Supervised Camouflaged Object Detection
Dian Shao, Guangqian Guo, Shan Gao 0003
ECCV (35)3
2024 SAM-COD: SAM-Guided Unified Framework for Weakly-Supervised Camouflaged Object Detection
Pengxu Wei, Guangqian Guo, Shan Gao 0003
ECCV (35)3
2024 P2P: Transforming from Point Supervision to Explicit Visual Prompt for Object Detection and Segmentation
Guangqian Guo, Dian Shao, Sha Meng, Shan Gao 0003
IJCAI1
2024 Save the Tiny, Save the All: Hierarchical Activation Network for Tiny Object Detection
abstract
Tiny object detection (TOD) remains a challenging problem due to the extremely small size and weak feature presentations of tiny objects. Many effective methods have improved the detection of small objects below$32\times 32$pixels to some extent, but the performance is still poor for the tiny objects below$16\times 16$pixels. In this paper, we find that the aliasing between the features and object scales, namely feature-scale-aliasing, leads to the misalignment between feature subspaces and detection subspaces, and thus results in the interference of features, especially for tiny objects. To alleviate this, we propose a Hierarchical Activation (HA) method to obtain scale-specific feature subspaces by activating object features at different scales hierarchically. To this end, we design a Scale-Guided Feature Activation (SGFA) to decompose the original object-aliasing feature spaces into a group of scale-specific feature subspaces by scale-guided activation maps. Then, Scale-Specific Feature re-Coupling (SSFC) is used to enhance the feature subspaces by adaptively aggregating the feature subspaces from different groups. In addition, we propose to complement the scale-specific detailed information by a designed Detailed Information Compensation (DIC) method. Implementing HA, a multi-scale keypoint-based detector is constructed to improve the tiny object detection, referred to as Hierarchical Activation Network (HANet). Extensive experiments are carried out on three tiny object detection datasets, e.g., TinyPerson, AI-TOD, and TinyCOCO. Our HANet achieves 58.45%$AP_{50}^{all}$, 22.1%$AP$, and 15.76%$AP$on TinyPerson, AI-TOD, and TinyCOCO, respectively, showing a significant performance gain over the competitors.
Guangqian Guo, Pengfei Chen 0004, Xuehui Yu, Zhenjun Han, Qixiang Ye, Shan Gao 0003
IEEE Trans. Circuits Syst. Video Technol.1
2024 Effective Rotate: Learning Rotation-Robust Prototype for Aerial Object Detection
abstract
Aerial images often depict objects with arbitrary orientations, which pose challenges for conventional object detectors to detect and classify. To address this issue, rotation-equivariant Convolutional Neural Networks (CNNs) have been proposed to extract rotation-equivariant features. However, the orientation encoding in these networks is often unstable and noisy, deteriorating detection performance. In this paper, we first analyze the rotation-equivariant network. Then, we propose a Rotation-robust Prototype Generation (RPG) method, which consists of two parts, stabilization module and enhancement module. In stabilization module, we generate rotation-robust prototypes to increase the stability of cyclic shifts. In enhancement module, we use the obtained prototype to improve the response of the features to object semantics. The RPG method can be used as a plug-and-play module in both one-stage and two-stage detectors. With only 30 lines of code, we achieve an average 1% improvement on four challenging datasets, including DOTA-V1.5, DOTA-v1.0, DIOR-R, and HRSC2016.
Chaowei Wang, Guangqian Guo, Chang Liu 0047, Dian Shao, Shan Gao 0003
IEEE Trans. Geosci. Remote. Sens.2