Songhe Deng

dblp:359/4180 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2026
0009-0005-9843-2083ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
YearPublicationVenuePosition
2026 HitBack: Transformer With Hierarchical-Semantic Cross Attention and Background Contrast for Weakly Supervised Wildlife Semantic Segmentation
abstract
Monitoring wildlife behavior and population changes is critical for conservation efforts. However, specialized analysis of large volumes of wildlife images is extremely chal lenging, necessitating the use of artificial intelligence techniques to automatically detect, segment, and classify species captured by trap cameras. Despite the increasing use of AI in wildlife monitoring, challenges with data quality and availability persist. The Snapshot Serengeti (SS) dataset only has image-level labels and very few bounding box labels, and there's no dataset with pixel-level labels due to the significant annotation costs. To this end, we create and release the large-scale Semantic Segmentation for Snapshots of the Serengeti (S4) dataset, consisting of 24K high-quality images across 47 species with precise masks, for both common and rare species. This dataset serves as a resource for developing semantic segmentation algorithms in wildlife studies. Additionally, we introduce HitBack, a novel method leveraging Hierarchical-Semantic Cross Attention (HCA) and Background Contrast (BC) for weakly supervised semantic segmentation (WSSS). The HCA module is used to capture both the shared and distinct features across species, and the BC module is designed to enhance foreground activation by ensuring consistency in the backgrounds. Extensive experiments on the newly proposed S4benchmark show that, our HitBack presents competitive performance when compared with the state-of-the-art models. The mIoU of HitBack is +10.4%, +14.7%, and +18.4% higher than that of ToCo, SIPE, and MCTformer, respectively. In addition, our HitBack even obtains performance that surpasses the fully supervised and semi-supervised methods when annotation data is limited. Code and datasets will be available at Github.
Puxuan Xie, Xinshao Wang, Songhe Deng, Weizhao He, LinLin Shen
IEEE Trans. Multim.4
2025 CLIMS++: Cross Language Image Matching with Automatic Context Discovery for Weakly Supervised Semantic Segmentation
Jinheng Xie, Songhe Deng, Xianxu Hou, Zhaochuan Luo, LinLin Shen, Yawen Huang, Yefeng Zheng 0001, Zheng Shou 0001
Int. J. Comput. Vis.2
2024 Multi-Scale Dynamic and Hierarchical Relationship Modeling for Facial Action Units Recognition
abstract
Human facial action units (AUs) are mutually related in a hierarchical manner, as not only they are associated with each other in both spatial and temporal domains but also AUs located in the same/close facial regions show stronger relationships than those of different facial regions. While none of existing approach thoroughly model such hi-erarchical inter-dependencies among AUs, this paper proposes to comprehensively model multi-scale AU-related dynamic and hierarchical spatiotemporal relationship among AUs for their occurrences recognition. Specifically, we first propose a novel multi-scale temporal differencing network with an adaptive weighting block to explicitly capture facial dynamics across frames at different spatial scales, which specifically considers the heterogeneity of range and mag-nitude in different AUs' activation. Then, a two-stage strategy is introduced to hierarchically model the relationship among AUs based on their spatial distribution (i.e., local and cross-region AU relationship modelling). Experimental results achieved on BP4D and DISFA show that our approach is the new state-of-the-art in the field of AU occurrence recognition. Our code is publicly available at https://github.com/CVI-SZU/MDHR.
Zihan Wang 0005, Siyang Song, Songhe Deng, Weicheng Xie 0001, LinLin Shen
CVPR4
2024 APSeg: Auto-Prompt Network for Cross-Domain Few-Shot Semantic Segmentation
abstract
Few-shot semantic segmentation (FSS) endeavors to segment unseen classes with only a few labeled samples. Current FSS methods are commonly built on the assumption that their training and application scenarios share similar domains, and their performances degrade significantly while applied to a distinct domain. To this end, we propose to leverage the cutting-edge foundation model, the segment Anything Model (SAM), for generalization enhancement. The SAM however performs unsatisfactorily on domains that are distinct from its training data, which primarily comprise natural scene images, and it does not support automatic segmentation of specific semantics due to its interactive prompting mechanism. In our work, we introduce APSeg, a novel auto-prompt network for cross-domain few-shot semantic segmentation (CD-FSS), which is designed to be auto-prompted for guiding cross-domain segmentation. Specifically, we propose a Dual Prototype Anchor Transformation (DPAT) module that fuses pseudo query prototypes extracted based on cycle-consistency with support prototypes, allowing features to be transformed into a more stable domain-agnostic space. Additionally, a Meta Prompt (MPG) module is introduced to automatically generate prompt embeddings, eliminating the need for manual visual prompts. We build an efficient model which can be applied directly to target domains without fine-tuning. Extensive experiments on four cross-domain datasets show that our model outperforms the state-of-the-art CD-FSS method by 5.24% and 3.10% in average accuracy on 1-shot and 5-shot settings, respectively.
Weizhao He, Yang Zhang 0012, LinLin Shen, Songhe Deng
CVPR6
2024 Tune-an-Ellipse: CLIP Has Potential to Find what you Want
abstract
Visual prompting of large vision language models such as CLIP exhibits intriguing zero-shot capabilities. A manually drawn red circle, commonly used for highlighting, can guide CLIP's attention to the surrounding region, to identify specific objects within an image. Without precise object proposals, however, it is insufficient for localization. Our novel, simple yet effective approach, i.e., Differentiable Visual Prompting, enables CLIP to zero-shot localize: given an image and a text prompt describing an object, we first pick a rendered ellipse from uniformly distributed anchor ellipses on the image grid via visual prompting, then use three loss functions to tune the ellipse coefficients to encap-sulate the target region gradually. This yields promising ex-perimental results for referring expression comprehension without precisely specified object proposals. In addition, we systematically present the limitations of visual prompting inherent in CLIP and discuss potential solutions.
Jinheng Xie, Songhe Deng, Bing Li 0024, Yawen Huang, Yefeng Zheng 0001, Jürgen Schmidhuber, Bernard Ghanem, LinLin Shen, Zheng Shou 0001
CVPR2
2024 Enhancing Unsupervised Semantic Segmentation Through Context-Aware Clustering
abstract
Despite the great progress of semantic segmentation with supervised learning, annotating large amounts of pixel-wise labels is, however, very expensive and time-consuming. To this end, Unsupervised Semantic Segmentation(USS) has been proposed to learn semantic segmentation, without any form of annotations. This approach involves dense prediction of semantics which is however challenging due to the unreliable nature of local representations. To solve this problem, we propose a newly context-aware unsupervised semantic segmentation framework, which aims to enhance the unsupervised semantic segmentation by leveraging contextual knowledge within and across images. In particular, we introduce a training strategy based on our Pyramid Semantic Guidance (PSG), which utilizes holistic semantics on pyramid views to guide pixel clustering with a siamese network-based framework. Additionally, we introduce a Context-Aware Embedding (CAE) module to fuse global features with low-level geometrical and appearance representations. We evaluate our method on the COCO-Stuff dataset and achieved competitive results compared to both the convolutional and ViT-based USS methods. Specifically, we attain significant improvements of +4.5% and +5% mIoU for Stuff and all class segmentation respectively, compared to previous approaches that employ unsupervised convolutional backbones.
Yuan Wang 0083, Junliang Chen 0002, Songhe Deng, Zhi Wang 0001, LinLin Shen, Wenwu Zhu 0001
IEEE Trans. Multim.4
2023 QA-CLIMS: Question-Answer Cross Language Image Matching for Weakly Supervised Semantic Segmentation
abstract
Class Activation Map (CAM) has emerged as a popular tool for weakly supervised semantic segmentation (WSSS), allowing the localization of object regions in an image using only image-level labels. However, existing CAM methods suffer from under-activation of target object regions and false-activation of background regions due to the fact that a lack of detailed supervision can hinder the model's ability to understand the image as a whole. In this paper, we propose a novel Question-Answer Cross-Language-Image Matching framework for WSSS (QA-CLIMS), leveraging the vision-language foundation model to maximize the text-based understanding of images and guide the generation of activation maps. First, a series of carefully designed questions are posed to the VQA (Visual Question Answering) model with Question-Answer Prompt Engineering (QAPE) to generate a corpus of both foreground target objects and backgrounds that are adaptive to query images. We then employ contrastive learning in a Region Image Text Contrastive (RITC) network to compare the obtained foreground and background regions with the generated corpus. Our approach exploits the rich textual information from the open vocabulary as additional supervision, enabling the model to generate high-quality CAMs with a more complete object region and reduce false-activation of background regions. We conduct extensive analysis to validate the proposed method and show that our approach performs state-of-the-art on both PASCAL VOC 2012 and MS COCO datasets.
Songhe Deng, Jinheng Xie, LinLin Shen
ACM Multimedia1