Songsong Duan

dblp:326/3792 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0003-2983-4044ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Toward Universal Semantic Communication via Matchable Semantic Subspace Transmission
abstract
Semantic communication targets reliable task execution at the receiver under stringent bandwidth and channel constraints. However, existing communication paradigms either focus on bit-level signal reconstruction, impeding the balance between task efficacy and bandwidth efficiency, or are limited by fixed vocabularies and lack generalization when facing unknown categories and open scenarios. To this end, we propose Universal Semantic Communication (UniSC), an open-vocabulary semantic communication framework that formulates transmission as a Matchable Semantic Subspace Transmission (MSST) problem. In this work, "universal" refers to the ability to handle arbitrary text-defined semantic categories beyond fixed vocabularies, rather than universality across all vision tasks. The transmitted representation is explicitly constrained to preserve cross-modal matchability after noisy transmission, rather than merely supporting latent recovery or closed-set inference. Concretely, UniSC comprises a Visual Semantic Engine (VSE), a Semantic Squeeze Network (SSN), a Noise-Adaptive Semantic Re-expansion (NASR) module, and a VLM-based Decoder. VSE and SSN project images into a compact semantic subspace for transmission. This subspace is optimized to preserve both robustness and cross-modal matchability under channel corruption. NASR denoises and lifts the received features back into a semantically complete visual space, from which the VLM-based Decoder performs open-category inference by matching arbitrary text queries rather than relying on a fixed classifier head. The VLM-based Decoder employs a Text Semantic Engine (TSE) to map natural language to text embeddings and, via a learnable Text-Visual Bridge (TVB), aligns them with the reconstructed visual structure for cross-modal matching. To improve cross-modal alignment and transmission robustness, a two-stage training strategy first establishes cross-modal anchors and then optimizes end-to-end robustness and compactness. Extensive experiments on semantic segmentation benchmarks demonstrate that UniSC achieves strong generalization and state-of-the-art performance under harsh channel conditions, outperforming existing methods in both low-SNR and extreme-compression regimes.
Xi Yang 0011, Songsong Duan, Nannan Wang 0001
IEEE Trans. Image Process.3
2025 Dual Information Purification for Lightweight SAR Object Detection
abstract
Synthetic aperture radar (SAR) object detection requires accurate identification and localization of targets at various scales within SAR images. However, background clutter and speckle noise can obscure key features and mislead the knowledge distillation process. To address these challenges, we introduce the Dual Information Purification Knowledge Distillation (DIPKD) method, which improves the performance of the student model through three key strategies: denoising, enrichment, and decoupling. First, our Selective Noise Suppression (SNS) technique reduces speckle noise in global features by minimizing misleading information from the teacher model. Second, the Knowledge Level Decoupling (KLD) module separates features into target and non-target knowledge, balancing feature mapping and reducing background noise to enhance the extraction of critical information for the student model. Finally, the Reverse Information Transfer (RIT) module refines intermediate features in the student model, compensating for the loss of detailed local information. Experimental results demonstrate that DIPKD significantly outperforms existing distillation techniques in SAR object detection, achieving 60.2% and 51.4% mAP scores on the SSDD and HRSID datasets, respectively. Additionally, the student model shows performance improvements of 1.3% and 2.9% over the teacher model, highlighting the effectiveness of the information purification approach.
Xi Yang 0011, Songsong Duan, De Cheng
AAAI3
2025 Multi-Label Prototype Visual Spatial Search for Weakly Supervised Semantic Segmentation
abstract
Existing Weakly Supervised Semantic Segmentation (WSSS) relies on the CNN-based Class Activation Map (CAM) and Transformer-based self-attention map to generate class-specific masks for semantic segmentation. However, CAM and self-attention maps usually cause incomplete segmentation due to classification bias issue. To address this issue, we propose a Multi-Label Prototype Visual Spatial Search (MuP-VSS) method with a spatial query mechanism. Specifically, MuP-VSS consists of two key components: multi-label prototype representation and multi-label prototype optimization. The former designs a global embedding to learn the global tokens from the images, and then proposes a Prototype Embedding Module (PEM) to interact with patch tokens to understand the local semantic information. The latter utilizes the exclusivity and consistency principles of the multi-label prototypes to design three prototype losses to optimize them, which contain cross-class prototype (CCP) contrastive loss, cross-image prototype (CIP) contrastive loss, and patch-to-prototype (P2P) consistency loss. CCP loss models exclusivity of multi-label prototypes learned from a single image to enhance the discriminative properties of each class better. CCP loss learns the consistency of the same class-specific prototypes extracted from multiple images to enhance the semantic consistency. P2P loss is proposed to control the semantic response of the prototype to the image patches. Experimental results on Pascal VOC 2012 and MS COCO show that MuP-VSS significantly outperforms recent methods and achieves state-of-the-art performance.
Songsong Duan, Xi Yang 0011, Nannan Wang 0001
CVPR1
2025 풟ℐℋ-CLIP: Unleashing the Diversity of Multi-Head Self-Attention for Training-Free Open-Vocabulary Semantic Segmentation
Songsong Duan, Xi Yang 0011, Nannan Wang 0001
ICCV1
2025 Scale-Consistent Learnable PnP Network for Space Target Pose Estimation
abstract
Due to constraints from imaging devices, the most effective method for estimating the pose of space target RGB images is to establish a 2-D–3-D correspondence and then use the perspective-n-point (PnP) algorithm for pose recovery. However, traditional PnP algorithms are not differentiable, which hinders their integration with neural network training. Although recent work attempts to make PnP partially differentiable during the 2-D–3-D matching stage by mathematical methods, this leads to increased inevitably computational costs. To this end, we propose a scale-consistent learnable PnP (SCLP) network that facilitates end-to-end pose estimation for space targets. Our method incorporates a sparse keypoint learnable PnP (SKL-PnP) layer within a multiscale network, enabling PnP to function as a differentiable layer that integrates seamlessly with preceding neural components. Additionally, we also sample the 2-D–3-D correspondences to obtain sparse keypoint pairs, achieving a lightweight single-stage 6-D pose estimation algorithm. To manage the significant scale variations in space target images, we introduce Gaussian perception sampling (GPS) by assigning instances to different pyramid levels based on size. Furthermore, we propose a scale consistency regularization (SCR) module that aligns downsized feature maps with original ones to better address scale differences. Experimental results demonstrate that our approach achieves superior accuracy and efficiency on the SPEED and SwissCube datasets, showing significant improvements over state-of-the-art methods.
Xi Yang 0011, Jingyuan Wang 0002, Songsong Duan
IEEE Trans. Geosci. Remote. Sens.3
2025 SCIR: A Weakly Supervised Contextual Instance Refinement Method for Remote Sensing Object Detection
abstract
Most weakly supervised object detection (WSOD) methods currently prioritize the top-scoring object instance from proposals to train the corresponding object detector. The detector tends to focus on the entire object by analyzing the contextual information around the most discriminative activation regions. However, the traditional selective search algorithm yields top-scoring instance proposals that cover only a portion of the object, thereby diminishing the detector’s sensitivity to objects with a wide range of scale variations in remote sensing images (RSIs), particularly tiny object clusters. To address this issue, this paper proposes a novel WSOD method called SAM-guided proposal generation with Contextual Instance Refinement (SCIR) method for detecting objects with significant scale variations in RSIs. Specifically, a SAM-guided proposal generator (SPG) module is designed to generate high-quality proposals based on SAM masks rather than the traditional top-scoring one. Our SPG combines WSOD’s advantage of mining classification clues through inexact supervision with SAM’s capability of pre-learned world knowledge to provide automatic prompts for SAM. Meanwhile, we propose a contextual enhanced feature extractor (CEFE) module to capture the global context of the visual scene, which further activates the feature representation of the entire object. Finally, the feature map from CEFE and proposals generated by SPG are fed into a context-perceived instance refinement (CPIR) module. Our CPIR aims to shift the attention of the detection network from the local feature portion to the entire object by integrating local and global contextual information. Extensive experiments on the challenging DOTA and DIOR datasets demonstrate that our proposed SCIR achieves state-of-the-art performance and is quite effective on multi-scale object issues.
Xi Yang 0011, Zhongyuan Zhou, Songsong Duan, Dong Yang 0012
IEEE Trans. Geosci. Remote. Sens.3
2025 Lightweight RGB-D Salient Object Detection From a Speed-Accuracy Tradeoff Perspective
abstract
Current RGB-D methods usually leverage large-scale backbones to improve accuracy but sacrifice efficiency. Meanwhile, several existing lightweight methods are difficult to achieve high-precision performance. To balance the efficiency and performance, we propose a Speed-Accuracy Tradeoff Network (SATNet) for Lightweight RGB-D SOD from three fundamental perspectives: depth quality, modality fusion, and feature representation. Concerning depth quality, we introduce the Depth Anything Model to generate high-quality depth maps,which effectively alleviates the multi-modal gaps in the current datasets. For modality fusion, we propose a Decoupled Attention Module (DAM) to explore the consistency within and between modalities. Here, the multi-modal features are decoupled into dual-view feature vectors to project discriminable information of feature maps. For feature representation, we develop a Dual Information Representation Module (DIRM) with a bi-directional inverted framework to enlarge the limited feature space generated by the lightweight backbones. DIRM models texture features and saliency features to enrich feature space, and employ two-way prediction heads to optimal its parameters through a bi-directional backpropagation. Finally, we design a Dual Feature Aggregation Module (DFAM) in the decoder to aggregate texture and saliency features. Extensive experiments on five public RGB-D SOD datasets indicate that the proposed SATNet excels state-of-the-art (SOTA) CNN-based heavyweight models and achieves a lightweight framework with 5.2 M parameters and 415 FPS. The code is available at https://github.com/duan-song/SATNet.
Songsong Duan, Xi Yang 0011, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.1
2024 Pro2SAM: Mask Prompt to SAM with Grid Points for Weakly Supervised Object Localization
Xi Yang 0011, Songsong Duan, Nannan Wang 0001, Xinbo Gao 0001
ECCV (69)2
2024 An Effective and Lightweight Hybrid Network for Object Detection in Remote Sensing Images
abstract
In recent years, general object detection on nature images leverages convolutional neural networks (CNNs) and vision transformers (ViT) to achieve great progress and high accuracy. However, unlike nature scenarios, remote sensing systems usually deploy a large number of edge devices. Therefore, this condition encourages detectors to have lower parameters than high-complexity neural networks. To this end, we propose an effective and lightweight detection framework of hybrid network, which enhances representation learning to balance efficiency and accuracy of model. Specifically, to compensate low precision caused by lightweight neural networks, we design a boundary-aware context (BAC) module and a frequency self-attention refinement (FSAR) module to improve detector performance in a hybrid structure. The BAC module enhances the local features of the object by fusing the spatial context information with the original image, which not only improves the model accuracy but also effectively solves the problem of multiscale objects. To alleviate the interference of complex background, the FSAR module adopts an adaptive technology to filter out redundant information at different frequencies to improve overall detection performance. The comprehensive experiments on remote sensing datasets, i.e., NWPU VHR-10, LEVIR, and RSOD, indicate that the proposed method achieves state-of-the-art performance and balances between model size and accuracy.
Xi Yang 0011, Songsong Duan
IEEE Trans. Geosci. Remote. Sens.3
2022 Multi-Modality Diversity Fusion Network with Swintransformer for RGB-D Salient Object Detection
abstract
Multi-modality complementary information brings new impetus and innovation to saliency object detection (SOD). However, most existing RGB-D SOD methods either indiscriminately handle RGB features and depth features or only take depth features as additional information of RGB subnet-work, ignoring the different roles of two modalities for SOD tasks. To tackle this issue, we propose a novel multi-modality diversity fusion network with SwinTransformer (M2DFNet) for RGB-D SOD from the perspective of the different status of multi-modality, which adequately explores the roles of RGB and depth modalities. To this end, a triple-diversity supervision mechanism (TDSM) and a diversity fusion module (DFM) are designed to parse the function of different modalities. Besides, we designed a dense decoder (DSD) to integrate multi-scale features and transfer gain information from top to bottom, which can improve the performance of SOD. Extensive experiments on five benchmark datasets demonstrate that the proposed M2DFNet outperforms 17 other state-of-the-art (SOTA) RGB-D SOD methods.
Songsong Duan, Chenxing Xia, Xiuju Gao, Bin Ge 0001, Hanling Zhang, Kuanching Li
ICIP1
2022 DAST: Depth-Aware Assessment and Synthesis Transformer for RGB-D Salient Object Detection
Chenxing Xia, Songsong Duan, Xianjin Fang, Bin Ge 0001, Xiuju Gao, Jianhua Cui
PRICAI (2)2
2022 GCENet: Global contextual exploration network for RGB-D salient object detection
Chenxing Xia, Songsong Duan, Xiuju Gao, Rongmei Huang, Bin Ge 0001
J. Vis. Commun. Image Represent.2
2022 HDNet: Multi-Modality Hierarchy-Aware Decision Network for RGB-D Salient Object Detection
abstract
RGB-D Salient object detection (SOD) is a pixel-level dense prediction task, which can highlight the prominent object in the scene. Recently, Convolution Neural Network (CNN) is widely applied in SOD to generate multi-level features, which are complementary to each other. However, most methods ignore the unique characteristics of multi-level features (high-level and low-level features). Given the effective employment of multi-level features, we propose a novel multi-modality hierarchy-aware decision network (HDNet) by embedding a Swin Transformer as an encoder. The proposed HDNet contains three primary designs: (1) a Swin Transformer encoder is employed instead of a CNN to learn long-range dependencies; (2) a hierarchy-aware feature decision mechanism (HFDM) is proposed to exploit effective local detail cues of low-level features and global semantic information of high-level features, which consists of two sub-modules, namely low-hierarchy edge module (LEM) and high-hierarchy region module (HRM); (3) a decision-based fusion module (DFM) is designed to fuse RGB and depth features under the attribute of multi-level features generated from HFDM. Experiments on five public benchmarks verify that our framework has better performance than the other 18 state-of-the-art algorithms.
Chengxing Xia, Songsong Duan, Bin Ge 0001, Hanling Zhang, Kuanching Li
IEEE Signal Process. Lett.2
2022 DMINet: dense multi-scale inference network for salient object detection
Chenxing Xia, Xiuju Gao, Bin Ge 0001, Songsong Duan
Vis. Comput.5