EDBT 2026 Demo / reviewers in the wild / expert
Shan Gao 0003
dblp:67/4510-3
· DBLP profile ↗
19ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0002-3427-5825ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Boosting Segment Anything Model to Generalize Visually Non-Salient ScenariosabstractSegment Anything Model (SAM), known for its remarkable zero-shot segmentation capabilities, has garnered significant attention in the community. Nevertheless, its performance is challenged when dealing with what we refer to as visually non-salient scenarios, where there is low contrast between the foreground and background. In these cases, existing methods often cannot capture accurate contours and fail to produce promising segmentation results. In this paper, we propose Visually Non-Salient SAM (VNS-SAM), aiming to enhance SAM's perception of visually non-salient scenarios while preserving its original zero-shot generalizability. We achieve this by effectively exploiting SAM's low-level features through two designs: Mask-Edge Token Interactive decoder and Non-Salient Feature Mining module. These designs help the SAM decoder gain a deeper understanding of non-salient characteristics with only marginal parameter increments and computational requirements. The additional parameters of VNS-SAM can be optimized within 4 hours, demonstrating its feasibility and practicality. In terms of data, we established VNS-SEG, a unified dataset for various VNS scenarios, with more than 35K images, in contrast to previous single-task adaptations. It is designed to make the model learn more robust VNS features and comprehensively benchmark the model's segmentation performance and generalizability on VNS scenarios. Extensive experiments across various VNS segmentation tasks demonstrate the superior performance of VNS-SAM, particularly under zero-shot settings, highlighting its potential for broad real-world applications. Codes and datasets are publicly available at https://guangqian-guo.github.io/VNS-SAM/. Guangqian Guo, Pengfei Chen 0004, Boqiang Zhang, Shan Gao 0003 |
IEEE Trans. Image Process. | 6 |
| 2026 | ThermalGate-GS: Frequency-Gated Graph Splatting for Thermal Novel View SynthesisabstractThermal infrared imaging is pivotal for all-weather 3D perception, yet analyzing thermal information remains a formidable challenge due to the complexity of heat conduction. Unlike visible light, heat conduction acts as a natural low-pass filter that suppresses high-frequency textural details, causing severe geometric ambiguities and "ghosting" artifacts in standard 3D reconstruction pipelines. To accurately model the inherently diffusive thermal field for high-fidelity reconstruction, we propose Frequency-Gated Graph Splatting (ThermalGate-GS), a framework that explicitly decouples the scene into diffusive thermal distributions (low-frequency) and sharp structural boundaries (high-frequency). Within this framework, we introduce a novel Frequency-Gated Anisotropic Diffusion mechanism. Specifically, the frequency-gating module utilizes extracted high-frequency structural cues to determine spatially-adaptive gating weights. Subsequently, these weights drive an anisotropic diffusion process that dynamically regulates thermal feature propagation, promoting smoothness on object surfaces while suppressing cross-boundary bleeding. Finally, these spectrally refined features are employed to regress 3D Gaussian attributes, substantially alleviating the ambiguity in thermal reconstruction. Extensive experiments demonstrate that ThermalGate-GS achieves state-of-the-art performance, with a notable 7.94 dB PSNR improvement on the ThermoScenes benchmark over prior physics-inspired baselines. Yaoxing Wang, Wenkang Chen, Guangqian Guo, Chaowei Wang, Yan Di, Shan Gao 0003 |
IEEE Trans. Image Process. | 9 |
| 2025 | Segment Any-Quality Images with Generative Latent Space EnhancementabstractDespite their success, Segment Anything Models (SAMs) experience significant performance drops on severely degraded, low-quality images, limiting their effectiveness in real-world scenarios. To address this, we propose Gle-SAM, which utilizes Generative Latent space Enhancement to boost robustness on low-quality images, thus enabling generalization across various image qualities. Specifically, we adapt the concept of latent diffusion to SAM-based segmentation frameworks and perform the generative diffusion process in the latent space of SAM to reconstruct high-quality representation, thereby improving segmentation. Additionally, we introduce two techniques to improve compatibility between the pre-trained diffusion model and the segmentation framework. Our method can be applied to pre-trained SAM and SAM2 with only minimal additional learnable parameters, allowing for efficient optimization. We also construct the LQSeg dataset with a greater diversity of degradation types and levels for training and evaluating the model. Extensive experiments demonstrate that GleSAM significantly improves segmentation robustness on complex degradations while maintaining generalization to clear images. Furthermore, GleSAM also performs well on unseen degradations, underscoring the versatility of our approach and dataset. Guangqian Guo, Xuehui Yu, Yaoxing Wang, Shan Gao 0003 |
CVPR | 6 |
| 2025 | SAM-COD+: SAM-Guided Unified Framework for Weakly-Supervised Camouflaged Object DetectionabstractMost Camouflaged Object Detection (COD) methods heavily rely on mask annotations, which are time-consuming and labor-intensive to acquire. Existing weakly-supervised COD approaches exhibit significantly inferior performance compared to fully-supervised methods and struggle to simultaneously support all the existing types of camouflaged object labels, including scribbles, bounding boxes, and points. Even for Segment Anything Model (SAM), it is still problematic to handle the weakly-supervised COD and it typically encounters challenges of prompt compatibility of the scribble labels, extreme response, semantically erroneous response, and unstable feature representations, producing unsatisfactory results in camouflaged scenes. To mitigate these issues, we propose a unified COD framework in this paper, termed SAM-COD, which is capable of supporting arbitrary weakly-supervised labels. Our SAM-COD employs a prompt adapter to handle scribbles as prompts based on SAM. Meanwhile, we introduce response filter and semantic matcher modules to improve the quality of the masks obtained by SAM under COD prompts. To alleviate the negative impacts of inaccurate mask predictions, a new strategy of prompt-adaptive knowledge distillation is utilized to ensure a reliable feature representation. To validate the effectiveness of our approach, we have conducted extensive empirical experiments on three mainstream COD benchmarks. The results demonstrate the superiority of our method against state-of-the-art weakly-supervised and even fully-supervised methods. Our source codes and trained models will be publicly released. Pengxu Wei, Guangqian Guo, Shan Gao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Go Deep or Broad? Exploit Hybrid Network Architecture for Weakly Supervised Object Classification and LocalizationabstractWeakly supervised object classification and localization are learned object classes and locations using only image-level labels, as opposed to bounding box annotations. Conventional deep convolutional neural network (CNN)-based methods activate the most discriminate part of an object in feature maps and then attempt to expand feature activation to the whole object, which leads to deteriorating the classification performance. In addition, those methods only use the most semantic information in the last feature map, while ignoring the role of shallow features. So, it remains a challenge to enhance classification and localization performance with a single frame. In this article, we propose a novel hybrid network, namely deep and broad hybrid network (DB-HybridNet), which combines deep CNNs with a broad learning network to learn discriminative and complementary features from different layers, and then integrates multilevel features (i.e., high-level semantic features and low-level edge features) in a global feature augmentation module. Importantly, we exploit different combinations of deep features and broad learning layers in DB-HybridNet and design an iterative training algorithm based on gradient descent to ensure the hybrid network work in an end-to-end framework. Through extensive experiments on caltech-UCSD birds (CUB)-200 and imagenet large scale visual recognition challenge (ILSVRC) 2016 datasets, we achieve state-of-the-art classification and localization performance. Shan Gao 0003, Guangqian Guo, Hanqiao Huang, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | ShapeMatcher: Self-Supervised Joint Shape Canonicalization, Segmentation, Retrieval and DeformationabstractIn this paper, we present ShapeMatcher, a unified self-supervised learning framework for joint shape canonicalization, segmentation, retrieval and deformation. Given a partially-observed object in an arbitrary pose, we first canonicalize the object by extracting point-wise affine-invariant features, disentangling inherent structure of the object with its pose and size. These learned features are then leveraged to predict semantically consistent part segmentation and corresponding part centers. Next, our lightweight retrieval module aggregates the features within each part as its retrieval token and compare all the tokens with source shapes from a pre-established database to identify the most geometrically similar shape. Finally, we deform the retrieved shape in the deformation module to tightly fit the input object by harnessing part center guided neural cage deformation. The key insight of ShapeMaker is the simultaneous training of the four highly-associated processes: canonicalization, segmentation, retrieval, and deformation, leveraging cross-task consistency losses for mutual supervision. Extensive experiments on synthetic datasets PartNet, ComplementMe, and real-world dataset Scan2CAD demonstrate that ShapeMatcher surpasses competitors by a large margin. Code is released at https://github.com/Det1999/ShapeMaker. Yan Di, Chenyangguang Zhang, Chaowei Wang, Ruida Zhang, Guangyao Zhai, Xiangyang Ji, Shan Gao 0003 |
CVPR | 9 |
| 2024 | Just a Hint: Point-Supervised Camouflaged Object Detection
Dian Shao, Guangqian Guo, Shan Gao 0003 |
ECCV (35) | 4 |
| 2024 | SAM-COD: SAM-Guided Unified Framework for Weakly-Supervised Camouflaged Object Detection
Pengxu Wei, Guangqian Guo, Shan Gao 0003 |
ECCV (35) | 4 |
| 2024 | P2P: Transforming from Point Supervision to Explicit Visual Prompt for Object Detection and Segmentation
Guangqian Guo, Dian Shao, Sha Meng, Shan Gao 0003 |
IJCAI | 6 |
| 2024 | Save the Tiny, Save the All: Hierarchical Activation Network for Tiny Object DetectionabstractTiny object detection (TOD) remains a challenging problem due to the extremely small size and weak feature presentations of tiny objects. Many effective methods have improved the detection of small objects below$32\times 32$pixels to some extent, but the performance is still poor for the tiny objects below$16\times 16$pixels. In this paper, we find that the aliasing between the features and object scales, namely feature-scale-aliasing, leads to the misalignment between feature subspaces and detection subspaces, and thus results in the interference of features, especially for tiny objects. To alleviate this, we propose a Hierarchical Activation (HA) method to obtain scale-specific feature subspaces by activating object features at different scales hierarchically. To this end, we design a Scale-Guided Feature Activation (SGFA) to decompose the original object-aliasing feature spaces into a group of scale-specific feature subspaces by scale-guided activation maps. Then, Scale-Specific Feature re-Coupling (SSFC) is used to enhance the feature subspaces by adaptively aggregating the feature subspaces from different groups. In addition, we propose to complement the scale-specific detailed information by a designed Detailed Information Compensation (DIC) method. Implementing HA, a multi-scale keypoint-based detector is constructed to improve the tiny object detection, referred to as Hierarchical Activation Network (HANet). Extensive experiments are carried out on three tiny object detection datasets, e.g., TinyPerson, AI-TOD, and TinyCOCO. Our HANet achieves 58.45%$AP_{50}^{all}$, 22.1%$AP$, and 15.76%$AP$on TinyPerson, AI-TOD, and TinyCOCO, respectively, showing a significant performance gain over the competitors. Guangqian Guo, Pengfei Chen 0004, Xuehui Yu, Zhenjun Han, Qixiang Ye, Shan Gao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Effective Rotate: Learning Rotation-Robust Prototype for Aerial Object DetectionabstractAerial images often depict objects with arbitrary orientations, which pose challenges for conventional object detectors to detect and classify. To address this issue, rotation-equivariant Convolutional Neural Networks (CNNs) have been proposed to extract rotation-equivariant features. However, the orientation encoding in these networks is often unstable and noisy, deteriorating detection performance. In this paper, we first analyze the rotation-equivariant network. Then, we propose a Rotation-robust Prototype Generation (RPG) method, which consists of two parts, stabilization module and enhancement module. In stabilization module, we generate rotation-robust prototypes to increase the stability of cyclic shifts. In enhancement module, we use the obtained prototype to improve the response of the features to object semantics. The RPG method can be used as a plug-and-play module in both one-stage and two-stage detectors. With only 30 lines of code, we achieve an average 1% improvement on four challenging datasets, including DOTA-V1.5, DOTA-v1.0, DIOR-R, and HRSC2016. Chaowei Wang, Guangqian Guo, Chang Liu 0047, Dian Shao, Shan Gao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Tracking without Label: Unsupervised Multiple Object Tracking via Contrastive Similarity LearningabstractUnsupervised learning is a challenging task due to the lack of labels. Multiple Object Tracking (MOT), which inevitably suffers from mutual object interference, occlusion, etc., is even more difficult without label supervision. In this paper, we explore the latent consistency of sample features across video frames and propose an Unsupervised Contrastive Similarity Learning method, named UCSL, including three contrast modules: self-contrast, cross-contrast, and ambiguity contrast. Specifically, i) self-contrast uses intra-frame direct and inter-frame indirect contrast to obtain discriminative representations by maximizing self-similarity. ii) Cross-contrast aligns cross- and continuous-frame matching results, mitigating the persistent negative effect caused by object occlusion. And iii) ambiguity contrast matches ambiguous objects with each other to further increase the certainty of subsequent object association through an implicit manner. On existing benchmarks, our method outperforms the existing unsupervised methods using only limited help from ReID head, and even provides higher accuracy than lots of fully supervised methods. Sha Meng, Dian Shao, Jiacheng Guo, Shan Gao 0003 |
ICCV | 4 |
| 2021 | A Graphical Social Topology Model for RGB-D Multi-Person TrackingabstractTracking multiple persons is a challenging task especially when persons move in groups and occlude one another. Existing research have investigated the problems of group division and segmentation; however, lacking overall person-group topology modeling limits the ability to handle complex person and group dynamics. We propose a Graphical Social Topology (GST) model in the RGB-D data domain, and estimate object group dynamics by jointly modeling the group structure and states of persons using RGB-D topological representation. With our topology representation, moving persons are not only assigned to groups, but also dynamically connected with each other, which enables in-group individuals to be correctively associated and the cohesion of each group to be precisely modeled. Using the learned typical topology pattern and group online update modules, we infer the birth/death and merging/splitting of dynamic groups. With the GST model, the proposed multi-person tracker can naturally facilitate the occlusion problem by treating the occluded object and other in-group members as a whole, while leveraging overall state transition. Experiments on different RGB-D and RGB datasets confirm that the proposed multi-person tracker improves the state-of-the-arts. Shan Gao 0003, Qixiang Ye, Li Liu 0002, Arjan Kuijper, Xiangyang Ji |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Monocular Piecewise Depth Estimation in Dynamic Scenes by Exploiting Superpixel RelationsabstractIn this paper, we propose a novel and specially designed method for piecewise dense monocular depth estimation in dynamic scenes. We utilize spatial relations between neighboring superpixels to solve the inherent relative scale ambiguity (RSA) problem and smooth the depth map. However, directly estimating spatial relations is an ill-posed problem. Our core idea is to predict spatial relations based on the corresponding motion relations. Given two or more consecutive frames, we first compute semi-dense (CPM) or dense (optical flow) point matches between temporally neighboring images. Then we develop our method in four main stages: superpixel relations analysis, motion selection, reconstruction, and refinement. The final refinement process helps to improve the quality of the reconstruction at pixel level. Our method does not require per-object segmentation, template priors or training sets, which ensures flexibility in various applications. Extensive experiments on both synthetic and real datasets demonstrate that our method robustly handles different dynamic situations and presents competitive results to the state-of-the-art methods while running much faster than them. Henrique Morimitsu, Shan Gao 0003, Xiangyang Ji |
ICCV | 3 |
| 2018 | 3-Stream Convolutional Networks for Video Action Recognition with Hybrid Motion FieldabstractTwo-stream based architectures for video action recognition exhibit great success recently. They encode the appearance with RGB frame, and the motion with optical flow. It is observed that optical flow depicts pixel-level motion field, focusing much on detail information, is hard to tackle the large displacement. In fact, human always focus the global motion rather than pixel-level motion. Inspired by this, we propose a novel 3-stream network structure with a spatial ConvNet, a pixel-level temporal ConvNet and a block-level temporal ConvNet. Integrating multi-granularity motion representation significantly outperforms single pixel-level motion field based architectures. Further, we can obtain the block-level motion vector field from compressed videos without extra calculation. We address missing and noisy motion patterns of motion vector field with intra-encoded block rectifying and flow guided filtering, building a hybrid motion field for our block-level temporal ConvNet. Our approach obtains state-of-the-art accuracy on UCF101 (95.27%) and HMDB 51 (69.21 %). Wukui Yang, Shan Gao 0003, Wenran Liu, Xiangyang Ji |
MMSP | 2 |
| 2017 | Beyond Group: Multiple Person Tracking via Minimal Topology-Energy-VariationabstractTracking multiple persons is a challenging task when persons move in groups and occlude each other. Existing group-based methods have extensively investigated how to make group division more accurately in a tracking-by-detection framework; however, few of them quantify the group dynamics from the perspective of targets' spatial topology or consider the group in a dynamic view. Inspired by the sociological properties of pedestrians, we propose a novel socio-topology model with a topology-energy function to factor the group dynamics of moving persons and groups. In this model, minimizing the topology-energy-variance in a two-level energy form is expected to produce smooth topology transitions, stable group tracking, and accurate target association. To search for the strong minimum in energy variation, we design the discrete group-tracklet jump moves embedded in the gradient descent method, which ensures that the moves reduce the energy variation of group and trajectory alternately in the varying topology dimension. Experimental results on both RGB and RGB-D data sets show the superiority of our proposed model for multiple person tracking in crowd scenes. Shan Gao 0003, Qixiang Ye, Junliang Xing, Arjan Kuijper, Zhenjun Han, Jianbin Jiao, Xiangyang Ji |
IEEE Trans. Image Process. | 1 |
| 2015 | Real-Time Multipedestrian Tracking in Traffic Scenes via an RGB-D-Based Layered Graph ModelabstractMultipedestrian tracking in traffic scenes is challenging due to cluttered backgrounds and serious occlusions. In this paper, we propose a layered graph model in image (RGB) and depth (D) domains for real-time robust multipedestrian tracking. The motivation is to investigate high-level constraints in RGB-D data association and to improve the optimization from the trajectory level to the layer level. To construct a layered graph, we define constraints in the depth domain so that pedestrian objects in the image domain are assigned to proper layers. We use pedestrian detection responses in the RGB domain as graph nodes, and we integrate 3-D motion, appearance, and depth features as graph edges. An online updating depth factor is defined to describe the depth relationships among the observations in and out of the layers, and the occlusion issue is processed with an analytical layer-level strategy. With a heuristic label switching algorithm, multiple pedestrian objects are optimally associated and tracked. Experiments and comparison on five public data sets show that our proposed approach significantly reduces pedestrian's ID switch and improves tracking accuracy in the cases of serious occlusions. Shan Gao 0003, Zhenjun Han, Ce Li 0005, Qixiang Ye, Jianbin Jiao |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2014 | Depth Structure Association for RGB-D Multi-target TrackingabstractMulti-target tracking in outdoor scenes plays an important role in many computer vision applications. Most previous work on visual information based multi-target tracking does not incorporate depth information and the absence of depth information often leads to mismatching or tracking failures. In this paper, we propose a Depth Structure Association (DSA) approach for RGB-D data based multi-target tracking. DSA encodes depth information in a chain structure, the structure is used by DSA together with appearance and motion information to address object occlusion issues in outdoor scenes. Additionally, the use of DSA has the advantages of regulating a much smaller solution space, greatly reducing the computational complexity. Experimental results on three datasets demonstrate that our DSA approach can significantly reduce object mismatch and tracking failure for long term occlusions. Shan Gao 0003, Zhenjun Han, David S. Doermann, Jianbin Jiao |
ICPR | 1 |
| 2014 | Locality-Constrained Sparse Reconstruction for Trajectory ClassificationabstractTrajectory classification has been extensively investigated in recent years, however, problems remain when processing incomplete trajectories of noises and local variations. In this paper, we propose a Locality-constrained Sparse Reconstruction (LSR) approach that explores both sparsity and local adaptability for robust trajectory classification. A trajectory dictionary with locality constrains is constructed with track lets partitioned from collected trajectories by control points of cubic B-spline curves. On the dictionary, the proposed LSR is used to calculate a discriminate code matrix. Then, a loss weighted decoding strategy is employed to perform multi-class trajectory classification. In addition, the approach can be used for anomalous trajectory detection with a thresholding strategy. Experiments on two datasets show that the results of the LSR approach improve the state of the art. Ce Li 0005, Zhenjun Han, Qixiang Ye, Shan Gao 0003, Lijin Pang, Jianbin Jiao |
ICPR | 4 |