Gangshan Wu

dblp:78/1123 · DBLP profile ↗
← Back
10ranked-venue papers in the field
0as first author
5since 2021 · last 2024
0000-0003-1391-1762ORCID · corroborated

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 6Other / Interdisciplinary · 4
YearPublicationVenuePosition
2024 RGB-D Video Object Segmentation via Enhanced Multi-store Feature Memory
abstract
The RGB-Depth (RGB-D) Video Object Segmentation (VOS) aims to integrate the fine-grained texture information of RGB with the spatial geometric clues of depth modality, boosting the performance of segmentation. However, off-the-shelf RGB-D segmentation methods fail to fully explore cross-modal information and suffer from object drift during long-term prediction. In this paper, we propose a novel RGB-D VOS method via multi-store feature memory for robust segmentation. Specifically, we design the hierarchical modality selection and fusion, which adaptively combines features from both modalities. Additionally, we develop a segmentation refinement module that effectively utilizes the Segmentation Anything Model (SAM) to refine the segmentation mask, ensuring more reliable results as memory to guide subsequent segmentation tasks. By leveraging spatio-temporal embedding and modality embedding, mixed prompts and fused images are fed into SAM to unleash its potential in RGB-D VOS. Experimental results show that the proposed method achieves state-of-the-art performance on the latest RGB-D VOS benchmark.
Boyue Xu, Ruichao Hou, Tongwei Ren, Gangshan Wu
ICMR4
2023 ADNet: An Asymmetric Dual-Stream Network for RGB-T Salient Object Detection
abstract
RGB-Thermal salient object detection (RGB-T SOD) aims to locate salient objects in images that include both RGB and thermal information. Previous approaches often suggest designing a symmetric network structure to tackle the challenge of dealing with low-quality RGB or thermal images. However, we contend that RGB and thermal modalities possess different numbers of channels and disparities in information density. In this paper, we propose a novel asymmetric dual-stream network (ADNet). Specifically, we leverage an asymmetric backbone to extract four stages of RGB features and four stages of thermal features. To enable effective interaction among low-level features in the first two stages, we introduce the Channel-Spatial Interaction (CSI) module. In the last two stages, deep features are enhanced using the Self-Attention Enhancement (SAE) module. Experimental results on the VT5000, VT1000, and VT821 datasets attest to the superior performance of our proposed ADNet compared to state-of-the-art methods.
Yaqun Fang, Ruichao Hou, Jia Bei, Tongwei Ren, Gangshan Wu
MMAsia5
2023 RGB-D Tracking via Hierarchical Modality Aggregation and Distribution Network
abstract
The integration of dual-modal features has been pivotal in advancing RGB-Depth (RGB-D) tracking. However, current trackers are less efficient and focus solely on single-level features, resulting in weaker robustness in fusion and slower speeds that fail to meet the demands of real-world applications. In this paper, we introduce a novel network, denoted as HMAD (Hierarchical Modality Aggregation and Distribution), which addresses these challenges. HMAD leverages the distinct feature representation strengths of RGB and depth modalities, giving prominence to a hierarchical approach for feature distribution and fusion, thereby enhancing the robustness of RGB-D tracking. Experimental results on various RGB-D datasets demonstrate that HMAD achieves state-of-the-art performance. Moreover, real-world experiments further validate HMAD’s capacity to effectively handle a spectrum of tracking challenges in real-time scenarios.
Boyue Xu, Ruichao Hou, Jia Bei, Tongwei Ren, Gangshan Wu
MMAsia6
2022 Reproducibility Companion Paper: Human Object Interaction Detection via Multi-level Conditioned Network
abstract
To support the replication of ?Human Object Interaction Detection via Multi-level Conditioned Network", which was presented at ICMR'20, this companion paper provides the details of the artifacts. Human Object Interaction Detection (HOID) aims to recognize fine-grained object-specific human actions, which demands the capabilities of both visual perception and reasoning. In this paper, we explain the file structure of the source code and publish the details of our experiments settings. We also provide a program for component analysis to assist other researchers with experiments on alternative models that are not included in our experiments. Moreover, we provide a demo program for facilitating the use of our model.
Yunqing He, Xu Sun 0009, Tongwei Ren, Gangshan Wu, Maria Sinziana Astefanoaei, Andreas Leibetseder
ICMR5
2021 Hybrid Improvements in Multimodal Analysis for Deep Video Understanding
abstract
The Deep Video Understanding Challenge (DVU) is a task that focuses on comprehending long duration videos which involve many entities. Its main goal is to build relationship and interaction knowledge graph between entities to answer relevant questions. In this paper, we improved the joint learning method which we previously proposed in many aspects, including few shot learning, optical flow feature, entity recognition, and video description matching. We verified the effectiveness of these measures through experiments.
Beibei Zhang 0005, Fan Yu 0003, Yaqun Fang, Tongwei Ren, Gangshan Wu
MMAsia5
2020 Human Object Interaction Detection via Multi-level Conditioned Network
abstract
As one of the essential problems in scene understanding, human object interaction detection (HOID) aims to recognize fine-grained object-specific human actions, which demands the capabilities of both visual perception and reasoning. Existing methods based on convolutional neural network (CNN) utilize diverse visual features for HOID, which are insufficient for complex human object interaction understanding. To enhance the reasoning capablity of CNN, we propose a novel multi-level conditioned network that fuses extra spatial-semantic knowledge with visual features. Specifically, we construct a multi-branch CNN as backbone for multi-level visual representation. We then encode extra knowledge including human body structure and object context as condition to dynamically influence the feature extraction of CNN by affine transformation and attention mechanism. Finally, we fuse the modulated multimodal features to distinguish the interactions. The proposed method is evaluated on two most frequently-used benchmarks, HICO-DET and V-COCO. The experiment results show that our method is superior to the state-of-the-arts.
Xu Sun 0009, Xinwen Hu, Tongwei Ren, Gangshan Wu
ICMR4
2018 Object Trajectory Proposal via Hierarchical Volume Grouping
abstract
Object trajectory proposal aims to locate category-independent object candidates in videos with a limited number of trajectories,i.e.,bounding box sequences. Most existing methods, which derive from combining object proposal with tracking, cannot handle object trajectory proposal effectively due to the lack of comprehensive objectness measurement through analyzing spatio-temporal characteristics over a whole video. In this paper, we propose a novel object trajectory proposal method using hierarchical volume grouping. Specifically, we first represent a given video with hierarchical volumes by mapping hierarchical regions with optical flow. Then, we filter the short volumes and background volumes, and combinatorially group the retained volumes into object candidates. Finally, we rank the object candidates using a multi-modal fusion scoring mechanism, which incorporates both appearance objectness and motion objectness, and generate the bounding boxes of the object candidates with the highest scores as the trajectory proposals. We validated the proposed method on a dataset consisting of 200 videos from ILSVRC2016-VID. The experimental results show that our method is superior to the state-of-the-art object trajectory proposal methods.
Xu Sun 0009, Yuantian Wang, Tongwei Ren, Zhi Liu 0003, Zhengjun Zha, Gangshan Wu
ICMR6
2015 Topic Modeling in Semantic Space with Keywords
abstract
A common and convenient approach for user to describe his information need is to provide a set of keywords. Therefore, the technique to understand the need becomes crucial. In this paper, for the information need about a topic or category, we propose a novel method called TDCS(Topic Distilling with Compressive Sensing) for explicit and accurate modeling the topic implied by several keywords. The task is transformed as a topic reconstruction problem in the semantic space with a reasonable intuition that the topic is sparse in the semantic space. The latent semantic space could be mined from documents via unsupervised methods, e.g. LSI. Compressive sensing is leveraged to obtain a sparse representation from only a few keywords. In order to make the distilled topic more robust, an iterative learning approach is adopted. The experiment results show the effectiveness of our method. Moreover, with only a few semantic concepts remained for the topic, our method is efficient for subsequent text mining tasks.
Xiaojia Pu, Rong Jin 0001, Gangshan Wu, Dingyi Han, Gui-Rong Xue
CIKM3
2012 Semiconducting bilinear deep learning for incomplete image recognition
abstract
Image recognition with incomplete data is a well-known hard problem in multimedia content analysis. This paper proposes a novel deep learning technique called semiconducting bilinear deep belief networks (SBDBN) by referencing human's visual cortex and intelligent perception. Inheriting from deep models, SBDBN simulates the laminar structure of human's cerebral cortex and the neural loop in human's visual areas. To address the special difficulties of image recognition with incomplete data, we design a novel second-order deep architecture with semiconducting restricted boltzmann machines. Moreover, two peaks activation of human's perception is implemented by three learning stages of semiconducting bilinear discriminant initialization, greedy layer-wise reconstruction, and global fine-tuning. Owing to exploiting the embedding information according to the reliable features rather than any completion of missing features, the proposed SBDBN has demonstrated outstanding recognition ability on two standard datasets and one constructed dataset, comparing with both incomplete image recognition techniques and existing deep learning models.
Shenghua Zhong, Yan Liu 0004, Korris Fu-Lai Chung, Gangshan Wu
ICMR4
2012 S-SIFT: A Shorter SIFT without Least Discriminability Visual Orientation
abstract
Detection and description of local features are a classical problem in image processing and multimedia content analysis. Based on the in homogeneity of visual orientation in human visual system, we propose a novel algorithm S-SIFT to detect and describe local image features. In three stages of S-SIFT, the information from the least discriminability orientation is omitting. Compared with the standard SIFT algorithm, S-SIFT has lower dimension and provides a faster key point matching. Experiments on the standard dataset demonstrate that our algorithm yields comparable or even better results for feature detection and matching tasks.
Shenghua Zhong, Yan Liu 0004, Gangshan Wu
Web Intelligence3