Shenglong Hu

dblp:242/9630 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
10since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Spectral Clustering for Community Detection of Multi-Layer Networks
abstract
Community extraction for multi-layer networks is a fundamental problem. Community extraction methods for multi-layer networks typically rely on a single-layer network, weights for each layer, and an additional clustering step. Current methods usually utilize average or empirical weights for each layer, which may not be reasonable. To address the challenge of assigning better weights to the layers, this paper proposes a Spectral Clustering Community Detection (SCCD) optimization model based on a unified similarity matrix. Specifically, the unified similarity matrix is defined via an optimization model that uses a weighted combination of similarity matrices from each layer, where the layer weights are learned under simplex constraints. A rank constraint is further imposed on the Laplacian matrix of the unified similarity matrix, which intrinsically leads to the division of the nodes into the desired number of clusters. Then, the SCCD optimization model is proposed, and its ability to generate community partitions is shown. An alternating minimization algorithm with simplex projections and closed-form updates is developed to solve the model, and its convergence is proven. Finally, numerical experiments are conducted on both synthetic and real-world multi-layer networks. Compared with several methods reported in the references, our method improves average NMI by 6.78% and ARI by 6.50% over the strongest baselines on the AUCs, UCI mfeat, Wikipedia, and Primary school networks—see Table III.
Gu-Yan Ni, Xiaojun Duan, Shenglong Hu
IEEE Trans. Knowl. Data Eng.7
2025 Learning Deep Frequency Degradation Prior for Remote Sensing Spatio-temporal Fusion
abstract
Existing deep learning-based remote sensing spatiotemporal fusion (STF) relies on a data-driven paradigm without considering the degradation prior modeling from the coarseto fine-resolution images. This makes the learned model easy to overfit to the training dataset, resulting in poor domain generalization across different datasets. To this end, this paper presents a deep frequency degradation prior to STF, dubbed as DeepFDP. The DeepFDP is based on a statistical observation that the frequency feature distributions of the tokens from the coarse- and fine-resolution images in different datasets have a low intra-resolution variance and a high inter-resolution variance with compact well-separated clusters. Therefore, the DeepFDP first designs a frequency proximity module to learn the mapping function between the token frequency representations of the coarse- and fine-resolution images, which is to narrow their feature distribution difference. Since the statistical properties of the token frequency representations are independent of the land-cover classes, the DeepFDP can be learned on a training set of limited image pairs without extra-supervision signals, which has a favorable zero-shot generalization capability across different datasets. Then, to faithfully recover the land-surface details, a high-frequency feature modulation module is designed that uses the fine-resolution image as guidance to progressively learn the multi-scale residual features in a coarse-to-fine fashion, yielding the fused features with rich high-frequency details. Finally, the progressively-fused features at each stage are fed into a hybrid fusion module, yielding the fine-resolution image prediction. Extensive evaluations on LGC and CIA datasets demonstrate favorable performance of the DeepFDP over state-of-the-art methods. Especially, the DeepFDP also shows a good zero-shot generalization performance when training on LGC and testing on CIA.
Yiting Bian, Shenglong Hu, Huihui Song 0003, Kaihua Zhang 0001
ICASSP2
2025 Continuously Learning Video-level Object Tokens for Robust UAV tracking
abstract
Due to the dynamic changes in flight motion and viewpoint, the objects in unmanned aerial vehicle (UAV) tracking scenarios often suffer from drastic appearance variations. Existing UAV trackers often leverage a frame-level matching mechanism, which measures the appearance similarity between the object template and the search frame. The drastic object appearance variations degrade the learned model, leading to drift issue. To this end, this paper presents a video-level UAV tracking framework that focuses on Continuously Learning (CL) effective and efficient spatio-temporal object tokens for robust tracking, dubbed as CLTrack. Specifically, the CLTrack first learns a series of spatio-temporal object tokens via a dynamic filtering module (DFM), which encodes more consensus object appearance information from each frame. Afterwards, a spatio-temporal enhancement module (STEM) is designed via cascading a temporal and a spatial attention to fully interact with the selected tokens with stable long-range spatio-temporal context information of the tracked object. Finally, to ensure the learned model encodes the rich context information without catastrophic forgetting, a video-level tracking loss is designed to supervise feature learning from the whole video frames. Extensive experiments on three UAV benchmarks including UAV123, DTB70 and VisDrone2018 demonstrate that the proposed CLTrack achieves state-of-the-art performance.
Shenglong Hu, Gang Dong, Lingyan Liang, Dongchao Wen, Kaihua Zhang 0001
ICASSP2
2025 Learning Joint Appearance and Shape Co-Representations for Co-Saliency Detection
abstract
Existing leading Co-saliency Detection (CoD) framework aims to segment the co-salient objects by learning the consensus visual representation of the foreground objects. However, despite different categories, some distractors may have similar appearance to the co-salient objects, such as Apples vs. Bananas have similar color and textures. This makes it challenging to distinguish the distractors only through learning the co-salient object appearance representations. To address this issue, we propose a joint appearance and shape co-representation learner for CoD, dubbed as ASCoD. The ASCoD is composed of a Co-Appearance learning Module (CoAM) and a Co-Shape learning Module (CoSM). The CoAM first learns a co-salient object appearance embedding that encodes the global cross-image and spatial context information. Then, this embedding is set as a co-appearance prototype, which guides the model to enhance the features to highlight the co-salient object regions. Afterwards, we design the CoSM that is a cross-attention module, among which the key and the value encode the shape information from a set of salient tokens dynamically selected by a Co-Shape Prototype generation Module (CSPM). Finally, through jointly optimizing the cascaded CoAM and CoSM, the optimal appearance and shape co-representations are achieved, marrying the merits of both appearance and shape co-representations that are not only robust to co-salient objects appearance variations, but also can well discriminate the co-salient objects from the distractors with similar appearance. Extensive evaluations on three challenging benchmarks including CoCA, CoSOD3k and CoSal2015, demonstrate superiority of the ASCoD to a variety of state-of-the-art CoD methods.
Guanting Guo, Shenglong Hu, Huihui Song 0003, Kaihua Zhang 0001
ICASSP2
2025 Open-Vocabulary Saliency-Guided Progressive Refinement Network for Unsupervised Video Object Segmentation
abstract
Existing leading unsupervised video object segmentation (UVOS) paradigm often leverages a dual-stream architecture with motion and appearance branches, where only the motion cues from optical flow are used as a guide to locating the primary foreground objects. When suffering from challenging factors such as static scenes, fast camera shaking, severe motion blur, etc., the estimated optical flow is noisy with low quality, leading to erroneous primary foreground objects estimation. To address this issue, we propose an open-vocabulary saliency-guided progressive refinement network for UVOS, dubbed as OVSNet. It is observed that most of the primary foreground objects also demonstrate saliency characteristics in the appearance branch. Based on this, our OVSNet complements motion cues with saliency cues predicted by a series of foundation models equipped with favorable zero-shot generalization capabilities. Specifically, we first leverage the off-the-shelf contrastive vision-language pre-training (CLIP) and CLIPSeg to generate an OVS attention map as saliency cues. Then, the saliency cues together with motion cues prompt the segment anything model (SAM) to generate a location map. In the location process, we design two lightweight adapters to fine-tune SAM, which makes SAM well adapt to the downstream UVOS task. Finally, the location map generated by SAM is used to progressively guide object representation refinement in the appearance branch, ultimately achieving an accurate segmentation mask prediction. Extensive evaluations on DAVIS-16, FBMS, and YouTube-Objects demonstrate the favorable performance of our OVSNet over the state-of-the-art methods.
Zhidong Han, Shenglong Hu, Huihui Song 0003, Kaihua Zhang 0001
ICASSP2
2025 Spatio-Semantic Prompt guided Adaptive Segment Anything for Remote Sensing Change Detection
abstract
Existing leading remote sensing change detection (RSCD) often takes a semantic-agnostic learning paradigm, which uses a binary ground-truth mask as supervision for model training. Despite the demonstrated success, due to the intrinsic characteristic of extremely complicated scene changes in RS images, this paradigm is prone to be misled by irrelevant semantic category changes, leading to a noisy CD mask prediction. To address this issue, this paper presents a Spatio-Semantic Prompt (SSP) guided adaptive Segment Anything Model (SAM) for RSCD, dubbed as SSP-SAM. The SSP-SAM introduces sparse textual and dense mask prompts into SAM to encode the task-specific semantic knowledge for RSCD. Specifically, we first encode the powerful textual semantic knowledge using Contrastive Language-Image Pre-training (CLIP) to determine the desired change semantic category. Then, we design a spatial dense prompt module that yields an attention map as prompt features to further refine the desired changed regions. Subsequently, we fine-tune the SAM through an adaptor to integrate the spatial-semantic prompt cues, yielding a coarse CD mask prediction. Finally, guided by the coarse CD mask, a multi-scale mask attention mechanism is adopted to learn the refined semantic representations of the changed targets, predicting the accurate CD mask. Extensive experiments on a variety of benchmark datasets demonstrate that the proposed SSP-SAM achieves state-of-the-art performance.
Shenglong Hu, Zhidong Han, Gang Dong, Lingyan Liang, Dongchao Wen, Kaihua Zhang 0001
ICASSP1
2025 Group-wise Semantic-enhanced Interaction Network for Remote Sensing Spatio-Temporal Fusion
abstract
Remote sensing spatio-temporal fusion (STF) aims at fusing temporally-dense coarse-resolution images and temporally-sparse fine-resolution images to reconstruct high spatio-temporal resolution images. Multi-band remote sensing images are often accepted as inputs for STF that have complementary characteristics for high-fidelity land surface reconstruction, yet the existing STF framework often treats different bands uniformly without considering the statistical correlation between different bands, resulting in unsatisfying results. To address this problem, this paper presents a group-wise semantic-enhanced interactive network for STF, dubbed as GSINet. Based on statistical observations, the feature correlation between the visible-light group and the invisible-light group is weak, while the intra-group correlation is strong. Therefore, the GSINet first separates the inputs into visible-light and invisible-light groups, which are fed into different branches with independent encoders for feature extraction and fusion. Afterwards, to address the issue of land cover changes between the prediction coarse- and reference fine-resolution images, a Semantic-Enhancement Fusion Module (SEFM) is designed to interact with the features from the same group with enhanced semantic information captured in an unsupervised learning manner. Then, the semantic-enhanced fused features from different bands are fed into an Interleaved Cross-attention Module (ICM) for further fusion. Finally, the output fusion features fully encode the intra- and inter-group information, which are fed into the decoder, reconstructing the spatio-temporal high-resolution images. Extensive experiments on CIA and LGC benchmark datasets demonstrate that the GSINet outperforms a variety of state-of-the-art methods in terms of multiple metrics.
Baoluo Zhu, Shenglong Hu, Huihui Song 0003, Kaihua Zhang 0001
ICASSP2
2025 Joint Feature Learning and Mixing via State Space Model for Remote Sensing Change Detection
abstract
Remote sensing change detection (RSCD) aims to identify change areas between bi-temporal images of the same location captured at different points in time. The existing siamese framework adopts a shared-weight strategy to process each bitemporal image independently. Due to the lack of inter-image information interaction, this strategy exhibits limited capability to target change perception and discrimination when faced with small targets or ambiguous changes such as low-covering change areas and irregular morphology in real-world complex scenes. To this end, we propose a one-stream framework using the State Space (SS) Model Mamba to jointly perform feature learning and mixing for RSCD, dubbed as SSCD. Specifically, by leveraging the long-range modeling capability and linear computational complexity of the SS model Mamba, the SSCD leverages a unified approach to feature extraction and information integration through simultaneously processing the bi-temporal images. This enables the model to intensively mutually guide to extract discriminative change features. In addition, to recover more spatial details, we design a texture enhancement module that makes full use of the selective scan modeling capability of the Mamba to enhance the texture features in different directions. Without bells and whistles, our SSCD achieves the state-of-the-art performance on three benchmark datasets including SYSU-CD, LEVIR-CD, and LEVIR+-CD.
Shenglong Hu, Huihui Song 0003, Kaihua Zhang 0001
ICME2
2024 Dual temporal memory network with high-order spatio-temporal graph learning for video object segmentation
Jiaqing Fan, Shenglong Hu, Kaihua Zhang 0001, Bo Liu 0005
Image Vis. Comput.2
2024 A Low-Rank Tensor Completion Method via Strassen-Ottaviani Flattening
abstract
Abstract. In this paper, a tensor completion method is proposed based on the Strassen–Ottaviani flattening, which can reveal the underlying tensor rank intrinsically. The resulting tensor completion optimization problem is formulated by spectral functions (convex or nonconvex) as surrogates for the rank function. An exact recovery result for the nuclear norm surrogate is given. An efficient method is proposed for this problem, and adaptive parameters for the spectral functions are allowed during the iterations. The global convergence is established under mild assumptions on the spectral functions and the adaptive parameter updates. In particular, we show that the class of weighted nuclear norms (either with nonincreasing or nondecreasing weights), the [Formula: see text]-sparsity index, and a class of weighted nuclear norms with adaptive weights, which are all widely employed in the literature, all fulfill the assumptions, and thus the global convergence is valid without any assumption. Numerical experiments on color images show that the proposed methods are promising, and always return images with better quality than some state-of-the-art methods.
Shenglong Hu, Zheng-Hai Huang
SIAM J. Imaging Sci.2