VLDB 2026 Research / reviewers in the wild / expert
Gangshan Wu
dblp:78/1123
· DBLP profile ↗
157ranked-venue papers
1as first author
84since 2021 · last 2026
0000-0003-1391-1762ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 114 · 61 since 2021Artificial intelligence and machine learning · 68 · 50 since 2021Databases, data management, data science and information retrieval · 10 · 5 since 2021Systems, architecture and hardware · 8 · 1 since 2021Theory of computation · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Computer networks · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RASR: Retrieval-Augmented Super Resolution for Practical Reference-based Image Restoration
Shuning Xu, Xiangyu Chen 0006, Dell Zhang, Jiantao Zhou 0001, Jie Tang 0006, Gangshan Wu, Jie Liu 0040 |
ISCAS | 7 |
| 2026 | GLAD: Generative Language-Assisted Visual Tracking for Low-Semantic Templates
Xingyu Luo, Yidong Cai, Jie Liu 0040, Jie Tang 0006, Gangshan Wu, Limin Wang 0002 |
Int. J. Comput. Vis. | 5 |
| 2026 | Spatiotemporal Predictive Pre-training for Robotic Motor Control
Jiange Yang, Bei Liu 0001, Jianlong Fu, Bocheng Pan, Gangshan Wu, Limin Wang 0002 |
Int. J. Comput. Vis. | 5 |
| 2026 | Learning frequency and memory-aware prompts for multi-modal object tracking
Boyue Xu, Ruichao Hou, Tongwei Ren, Dongming Zhou 0001, Gangshan Wu, Jinde Cao |
Pattern Recognit. | 5 |
| 2026 | Cross-View and Cross-Modal Contrastive Learning for Radar Object DetectionabstractFrequency-modulated continuous-wave radar is a cornerstone of advanced driver assistance systems thanks to its low cost and resilience to adverse weather. Yet the absence of explicit semantics makes radar annotation difficult, and the scarcity of large-scale labeled data limits the performance of radar perception models. To address this issue, we propose a self-supervised framework for object detection directly from Range–Azimuth– Doppler (RAD) cubes that learns transferable representations from unlabeled radar data. Specifically, we introduce cross-view contrastive learning to model correspondences among complementary views of the RAD cube, encouraging the network to capture spatial structure from multiple perspectives. In addition, an auxiliary cross-modal contrastive objective distills semantic knowledge from vision into radar. The joint objective integrates cross-view and cross-modal signals to strengthen radar feature representations. We further extend the framework to cross-domain pretraining using datasets from different sources. Experimental results demonstrate that the proposed method significantly improves radar object detection performance, especially with limited labeled data. Qiaolong Qian, Ruichao Hou, Haoyu Qin, Gangshan Wu |
IEEE Signal Process. Lett. | 5 |
| 2026 | HyPSAM: Hybrid Prompt-Driven Segment Anything Model for RGB-Thermal Salient Object DetectionabstractRGB-thermal salient object detection (RGB-T SOD) aims to identify prominent objects by integrating complementary information from RGB and thermal modalities. However, learning the precise boundaries and complete objects remains challenging due to the intrinsic insufficient feature fusion and the extrinsic limitations of data scarcity. In this paper, we propose a novel hybrid prompt-driven segment anything model (HyPSAM), which leverages the zero-shot generalization capabilities of the segment anything model (SAM) for RGB-T SOD. Specifically, we first propose a dynamic fusion network (DFNet) that generates high-quality initial saliency maps as visual prompts. DFNet employs dynamic convolution and multi-branch decoding to facilitate adaptive cross-modality interaction, overcoming the limitations of fixed-parameter kernels and enhancing multi-modal feature representation. Moreover, we propose a plug-and-play refinement network (P2RNet) which serves as a general optimization strategy to guide SAM in refining saliency maps by using hybrid prompts. The text prompt ensures reliable modality input, while the mask and box prompts enable precise salient object localization. Extensive experiments on three public datasets demonstrate that our method achieves state-of-the-art performance. Notably, HyPSAM has remarkable versatility, seamlessly integrating with different RGB-T SOD methods to achieve significant performance gains, thereby highlighting the potential of prompt engineering in this field. The code and results of our method are available at: https://github.com/milotic233/HyPSAM. Ruichao Hou, Tongwei Ren, Dongming Zhou 0001, Gangshan Wu, Jinde Cao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Text-Guided Nonverbal Enhancement Based on Modality-Invariant and -Specific Representations for Video Speaking Style RecognitionabstractVideo speaking style recognition (VSSR) aims to classify different types of conversations in videos, contributing significantly to understanding human interactions. A significant challenge in VSSR is the inherent similarity among conversation videos, which makes it difficult to distinguish between different speaking styles. Existing VSSR methods commit to providing available multimodal information to enhance the differentiation of conversation videos. Nevertheless, treating each modality equally leads to a suboptimal result for these methods due to text is inherently more aligned with conversation understanding compared to nonverbal modalities. To address this issue, we propose a text-guided nonverbal enhancement method, TNvE, which is composed of two core modules: 1) a text-guided nonverbal representation selection module employs cross-modal attention based on modality-invariant representations, picking out critical nonverbal information via textual guide; and 2) a modality-invariant and -specific representation decoupling module incorporates modality-specific representations and decouples them from modality-invariant representations, enabling a more comprehensive understanding of multimodal data. The former module encourages multimodal representations close to each other, while the latter module provides unique characteristics of each modality as a supplement. Extensive experiments are conducted on long-form video understanding datasets to demonstrate that TNvE is highly effective for VSSR, achieving a new state-of-the-art. Beibei Zhang 0005, Tongwei Ren, Gangshan Wu |
AAAI | 3 |
| 2025 | In-the-wild Audio Spatialization with Flexible Text-guided LocalizationabstractBinaural audio enriches immersive experiences by enabling the perception of the spatial locations of sounding objects in AR, VR, and embodied AI applications. While existing audio spatialization methods can generally map any available monaural audio to binaural audio signals, they often lack the flexible and interactive control needed in complex multi-object user-interactive environments. To address this, we propose a Text-guided Audio Spatialization (TAS) framework that utilizes diverse text prompts and evaluates our model from unified generation and comprehension perspectives. Due to the limited availability of high-quality, large-scale stereo data, we construct the SpatialTAS dataset, which encompasses 376,000 simulated binaural audio samples to facilitate the training of our model. Our model learns binaural differences guided by 3D spatial location and relative position prompts, enhanced with flipped-channel audio. Experimental results show that our model can generate high quality binaural audios for various audio types on both simulated and real-recorded datasets. Besides, we establish an assessment model based on Llama-3.1-8B, which evaluates the semantic accuracy of spatial locations through a spatial reasoning task. Results demonstrate that by utilizing text prompts for flexible and interactive control, we can generate binaural audio with both high quality and semantic consistency in spatial locations. Tianrui Pan, Jie Liu 0040, Zewen Huang, Jie Tang 0006, Gangshan Wu |
ACL (1) | 5 |
| 2025 | CATANet: Efficient Content-Aware Token Aggregation for Lightweight Image Super-ResolutionabstractTransformer-based methods have demonstrated impressive performance in low-level visual tasks such as Image Super-Resolution (SR). However, its computational complexity grows quadratically with the spatial resolution. A series of works attempt to alleviate this problem by dividing Low-Resolution images into local windows, axial stripes, or dilated windows. SR typically leverages the redundancy of images for reconstruction, and this redundancy appears not only in local regions but also in long-range regions. However, these methods limit attention computation to content-agnostic local regions, limiting directly the ability of attention to capture long-range dependency. To address these issues, we propose a lightweight Content-Aware Token Aggregation Network (CATANet). Specifically, we propose an efficient Content-Aware Token Aggregation module for aggregating long-range content-similar tokens, which shares token centers across all image tokens and updates them only during the training phase. Then we utilize intra-group self-attention to enable long-range information interaction. Moreover, we design an inter-group cross-attention to further enhance global information interaction. The experimental results show that, compared with the state-of-the-art cluster-based method SPIN, our method achieves superior performance, with a maximum PSNR improvement of 0.33dB and nearly double the inference speed. Xin Liu 0012, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
CVPR | 4 |
| 2025 | AutoLUT: LUT-Based Image Super-Resolution with Automatic Sampling and Adaptive Residual LearningabstractIn recent years, the increasing popularity of Hi-DPI screens has driven a rising demand for high-resolution images. However, the limited computational power of edge devices poses a challenge in deploying complex super-resolution neural networks, highlighting the need for efficient methods. While prior works have made significant progress, they have not fully exploited pixel-level information. Moreover, their reliance on fixed sampling patterns limits both accuracy and the ability to capture fine details in low-resolution images. To address these challenges, we introduce two plug-and-play modules designed to capture and leverage pixel information effectively in Look-Up Table (LUT) based super-resolution networks. Our method introduces Automatic Sampling (AutoSample), a flexible LUT sampling approach where sampling weights are automatically learned during training to adapt to pixel variations and expand the receptive field without added inference cost. We also incorporate Adaptive Residual Learning (AdaRL) to enhance inter-layer connections, enabling detailed information flow and improving the network’s ability to reconstruct fine details. Our method achieves significant performance improvements on both MuLUT and SPF-LUT while maintaining similar storage sizes. Specifically, for MuLUT, we achieve a PSNR improvement of approximately +0.20 dB improvement on average across five datasets. For SPF-LUT, with more than a 50% reduction in storage space and about a 2/3 reduction in inference time, our method still maintains performance comparable to the original. The code is available at https://github.com/SuperKenVery/AutoLUT. Yuheng Xu, Xin Liu 0012, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
CVPR | 6 |
| 2025 | Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy ConditioningabstractLearning from multiple domains is a primary factor that influences the generalization of a single unified robot system. In this paper, we aim to learn the trajectory prediction model by using broad out-of-domain data to improve its performance and generalization ability. Trajectory model is designed to predict any-point trajectories in the current frame given an instruction and can provide detailed control guidance for robotic policy learning. To handle the diverse out-of-domain data distribution, we propose a sparsely-gated MoE (Top-1 gating strategy) architecture for trajectory model, coined as Tra-MoE. The sparse activation design enables good balance between parameter cooperation and specialization, effectively benefiting from large-scale out-of-domain data while maintaining constant FLOPs per token. In addition, we further introduce an adaptive policy conditioning technique by learning 2D mask representations for predicted trajectories, which is explicitly aligned with image observations to guide action prediction more flexibly. We perform extensive experiments on both simulation and real-world scenarios to verify the effectiveness of Tra-MoE and adaptive policy conditioning technique. We also conduct a comprehensive empirical study to train Tra-MoE, demonstrating that our Tra-MoE consistently exhibits superior performance compared to the dense baseline model, even when the latter is scaled to match Tra-MoE’s parameter count. Jiange Yang, Haoyi Zhu, Gangshan Wu, Tong He 0001, Limin Wang 0002 |
CVPR | 4 |
| 2025 | LDMWSeg: Latent Diffusion Models for Weakly Supervised Medical Image Segmentation
Zuxian Huang, Gangshan Wu |
ICIC (28) | 2 |
| 2025 | KAN-SAM: Kolmogorov-Arnold Network Guided Segment Anything Model for RGB-T Salient Object DetectionabstractExisting RGB-thermal salient object detection (RGB-T SOD) methods aim to identify visually significant objects by leveraging both RGB and thermal modalities to enable robust performance in complex scenarios, but they often suffer from limited generalization due to the constrained diversity of available datasets and the inefficiencies in constructing multi-modal representations. In this paper, we propose a novel prompt learning-based RGB-T SOD method, named KAN-SAM, which reveals the potential of visual foundational models for RGB-T SOD tasks. Specifically, we extend Segment Anything Model 2 (SAM2) for RGB-T SOD by introducing thermal features as guiding prompts through efficient and accurate Kolmogorov-Arnold Network (KAN) adapters, which effectively enhance RGB representations and improve robustness. Furthermore, we introduce a mutually exclusive random masking strategy to reduce reliance on RGB data and improve generalization. Experimental results on benchmarks demonstrate superior performance over the state-of-the-art methods. Ruichao Hou, Tongwei Ren, Gangshan Wu |
ICME | 4 |
| 2025 | Towards Practical Real-Time Low-Latency Music Source SeparationabstractIn recent years, significant progress has been made in the field of deep learning for music demixing. However, there has been limited attention on real-time, low-latency music demixing, which holds potential for various applications, such as hearing aids, audio stream remixing, and live performances. Additionally, a notable tendency has emerged towards the development of larger models, limiting their applicability in certain scenarios. In this paper, we introduce a lightweight real-time low-latency model called Real-Time Single-Path TFC-TDF UNET (RT-STT), which is based on the Dual-Path TFC-TDF UNET (DTTNet). In RT-STT, we propose a feature fusion technique based on channel expansion. We also demonstrate the superiority of single-path modeling over dual-path modeling in real-time models. Moreover, we investigate the method of quantization to further reduce inference time. RT-STT exhibits superior performance with significantly fewer parameters and shorter inference times compared to state-of-the-art models. Junyu Wu, Jie Liu 0040, Tianrui Pan, Jie Tang 0006, Gangshan Wu |
ICME | 5 |
| 2025 | MotionRAG: Motion Retrieval-Augmented Image-to-Video GenerationabstractImage-to-video generation has made remarkable progress with the advancements in diffusion models, yet generating videos with realistic motion remains highly challenging. This difficulty arises from the complexity of accurately modeling motion, which involves capturing physical constraints, object interactions, and domain-specific dynamics that are not easily generalized across diverse scenarios. To address this, we propose MotionRAG, a retrieval-augmented framework that enhances motion realism by adapting motion priors from relevant reference videos through Context-Aware Motion Adaptation (CAMA). The key technical innovations include: (i) a retrieval-based pipeline extracting high-level motion features using video encoder and specialized resamplers to distill semantic motion representations; (ii) an in-context learning approach for motion adaptation implemented through a causal transformer architecture; (iii) an attention-based motion injection adapter that seamlessly integrates transferred motion features into pretrained video diffusion models. Extensive experiments demonstrate that our method achieves significant improvements across multiple domains and various base models, all with negligible computational overhead during inference. Furthermore, our modular design enables zero-shot generalization to new domains by simply updating the retrieval database without retraining any components. This research enhances the core capability of video generation systems by enabling the effective retrieval and transfer of motion priors, facilitating the synthesis of realistic motion dynamics. Chenhui Zhu, Yilu Wu, Gangshan Wu, Limin Wang 0002 |
NeurIPS | 4 |
| 2025 | Transferring Foundation Models for Generalizable Robotic ManipulationabstractImproving the generalization capabilities of general-purpose robotic manipulation in real world has long been a significant challenge. Existing approaches often rely on collecting large-scale robotic data which is costly and time-consuming. However, due to insufficient diversity of data, they typically suffer from limiting their capability in open-domain scenarios with new objects and diverse environments. In this paper, we propose a novel paradigm that effectively leverages language-reasoning segmentation mask generated by internet-scale foundation models, to condition robot manipulation tasks. By integrating the mask modality, which incorporates semantic, geometric, and temporal correlation priors derived from vision foundation models, into the end-to-end policy model, our approach can effectively and robustly perceive object pose and enable sample-efficient generalization learning, including new object instances, semantic categories, and unseen backgrounds. We first introduce a series of foundation models to ground natural language demands across multiple tasks. Secondly, we develop a two-stream 2D policy model based on imitation learning, which processes raw images and object masks to predict robot actions with a local-global perception manner. Extensive real-world experiments conducted on a Franka Emika robot and a low-cost dual-arm robot demonstrate the effectiveness of our proposed paradigm and policy. Demos can be found in link 1 or link 2 and our code will be released at https://github.com/MCG-NJU/TPM. Jiange Yang, Wenhui Tan, Chuhao Jin, Keling Yao, Bei Liu 0001, Jianlong Fu, Ruihua Song, Gangshan Wu, Limin Wang 0001 |
WACV | 8 |
| 2025 | FSDM: An efficient video super-resolution method based on Frames-Shift Diffusion Model
Chao Chen 0026, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
Neural Networks | 5 |
| 2025 | CycleACR: Cycle Modeling of Actor-Context Relations for Video Action DetectionabstractThe relation modeling between actors and scene context advances video action detection where the correlation of multiple actors makes their action recognition challenging. Existing studies model each actor and scene relation to improve action recognition. However, the scene variations and background interference limit their effectiveness. In this paper, we propose to select actor-related scene context, rather than directly laveraging raw video scenario, to improve relation modeling. We develop a Cycle Actor-Context Relation network (CycleACR) where there is a symmetric graph that models the actor and context relations in a bidirectional form. Specifically, our CycleACR is constituted of two modules: 1) Actor-to-Context Reorganization (A2C-R), which adaptively collects actor features for context feature reorganizations, and 2) Context-to-Actor Enhancement (C2A-E), which dynamically utilizes the reorganized context features for actor feature enhancement. Stacking multiple CycleACR modules is able to effectively capture the high-order relation and efficiently exchange useful information between actors and context. To fully exploit time-dependent and holistic context information, we further design a parallel local and global temporal context modeling branch. The outputs of the two branches are integrated as the final context-enhanced actor feature representations. Finally, we propose a context-aware memory bank for long-term relation modeling. The proposed bank can effectively store actor-related scene context from other clips without additional memory overhead. Compared to existing designs that focus on C2A-E, our CycleACR introduces the core design of A2C-R for more effective relation modeling. This cycle modeling enablesour CycleACR to achieve state-of-the-art performance on two popular action detection datasets: AVA (40.6 mAP) and UCF101-24 (84.7 mAP). We also provide ablation studies and visualizations to show how our cycle actor-context relation modeling improves video action detection. Zhan Tong, Yibing Song, Gangshan Wu, Limin Wang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | JointFormer: A Unified Framework With Joint Modeling for Video Object SegmentationabstractCurrent prevailing Video Object Segmentation (VOS) methods follow the pipeline of extraction-then-matching, which first extracts features on current and reference frames independently, and then performs dense matching between them. This decoupled pipeline limits information propagation between frames to high-level features, and fails to capture fine-grained details for matching. Furthermore, the pixel-wise matching lacks holistic target understanding, making it prone to disturbance by similar distractors. To address these issues, we propose a unified VOS framework, coined JointFormer, for jointly modeling feature extraction, correspondence matching, and a compressed memory. The Joint Modeling Block leverages attention operations to simultaneously extract and propagate the target information from the reference frame to the current frame and a compressed memory token.This joint modeling scheme enables extensive multi-layer propagation beyond high-level feature space and facilitates robust instance-distinctive feature learning. In addition, to incorporate the long-term and holistic target information, we introduce a compressed memory token with a customized online updating mechanism, which aggregates target features and performs temporal information propagation in a frame-wise manner, enhancing the global modeling consistency. Our JointFormer achieves a new state-of-the-art performance on the DAVIS 2017 val/test-dev (89.7% and 87.6%) benchmarks and the YouTube-VOS 2018/2019 val (87.0% and 87.0%) benchmarks. To demonstrate the generalizability of JointFormer, it is further evaluated on four new benchmarks with various challenges, including MOSE for complex scenes, VISOR for egocentric videos, VOST for complex transformations, and LVOS for long-term videos. Without specific design to address these unusual difficulties, our model achieves the best performance across all benchmarks when compared with several current best models, illustrating its excellent generalization and robustness. Further extensive ablations and visualizations indicate our JointFormer enables more comprehensive and effective feature learning and matching. Jiaming Zhang 0010, Yutao Cui, Gangshan Wu, Limin Wang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Group Visual Relation DetectionabstractIn this paper, we propose a novel visual relation detection task, named Group Visual Relation Detection (GVRD), for detecting visual relations whose subjects and/or objects are groups (GVRs), inspired by the observation that groups are common in image semantic representation. GVRD can be deemed as an evolution over the existing visual relation detection task that limits both subjects and objects of visual relations as individuals. We propose a Simultaneous Group Relation Prediction (SGRP) method that can simultaneously predict groups and predicates to address GVRD. SGRP contains an Entity Construction (EC) module, a Feature Extraction (FE) module, and a Group Relation Prediction (GRP) module. Specifically, the EC module constructs instances, group candidates, and phrase candidates; the FE module extracts visual, location and semantic features for these entities; and the GRP module simultaneously predicts groups and predicates, and generates the GVRs. Moreover, we construct a new dataset, named COCO-GVR, to facilitate solutions to GVRD task, which consists of 9,570 images from COCO dataset and 31,855 manually labeled GVRs. We test and validate the performance of SGRP by extensive experiments on COCO-GVR dataset. It shows that SGRP outperforms the baselines generated from the state-of-the-art visual relation detection and scene graph generation methods. Fan Yu 0003, Beibei Zhang 0005, Tongwei Ren, Gangshan Wu, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | Sketch and Refine: Towards Fast and Accurate Lane DetectionabstractLane detection is to determine the precise location and shape of lanes on the road. Despite efforts made by current methods, it remains a challenging task due to the complexity of real-world scenarios. Existing approaches, whether proposal-based or keypoint-based, suffer from depicting lanes effectively and efficiently. Proposal-based methods detect lanes by distinguishing and regressing a collection of proposals in a streamlined top-down way, yet lack sufficient flexibility in lane representation. Keypoint-based methods, on the other hand, construct lanes flexibly from local descriptors, which typically entail complicated post-processing. In this paper, we present a “Sketch-and-Refine” paradigm that utilizes the merits of both keypoint-based and proposal-based methods. The motivation is that local directions of lanes are semantically simple and clear. At the “Sketch” stage, local directions of keypoints can be easily estimated by fast convolutional layers. Then we can build a set of lane proposals accordingly with moderate accuracy. At the “Refine” stage, we further optimize these proposals via a novel Lane Segment Association Module (LSAM), which allows adaptive lane segment adjustment. Last but not least, we propose multi-level feature integration to enrich lane feature representations more efficiently. Based on the proposed “Sketch-and-Refine” paradigm, we propose a fast yet effective lane detector dubbed “SRLane”. Experiments show that our SRLane can run at a fast speed (i.e., 278 FPS) while yielding an F1 score of 78.9%. The source code is available at: https://github.com/passerer/SRLane. Chao Chen 0026, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
AAAI | 5 |
| 2024 | SportsHHI: A Dataset for Human-Human Interaction Detection in Sports VideosabstractVideo-based visual relation detection tasks, such as video scene graph generation, play important roles in fine-grained video understanding. However, current video visual relation detection datasets have two main limitations that hinder the progress of research in this area. First, they do not explore complex human-human interactions in multi-person scenarios. Second, the relation types of existing datasets have relatively low-level semantics and can be often recognized by appearance or simple prior information, without the need for detailed spatio-temporal context reasoning. Nevertheless, comprehending high-level interactions between humans is crucial for understanding complex multi-person videos, such as sports and surveillance videos. To address this issue, we propose a new video visual relation detection task: video human-human interaction detection, and build a dataset named SportsHHI for it. SportsHHI contains 34 high-level interaction classes from basketball and volleyball sports. 118,075 human bounding boxes and 50,649 interaction instances are annotated on 11,398 keyframes. To benchmark this, we propose a two-stage baseline method and conduct extensive experiments to reveal the key factors for a successful human-human interaction detector. We hope that SportsHHI can stimulate research on human interaction understanding in videos and promote the development of spatio-temporal context modeling techniques in video visual relation detection. Tao Wu 0020, Runyu He, Gangshan Wu, Limin Wang 0002 |
CVPR | 3 |
| 2024 | Asymmetric Masked Distillation for Pre-Training Small Foundation ModelsabstractSelf-supervised foundation models have shown great potential in computer vision thanks to the pre-training paradigm of masked autoencoding. Scale is a primary factor influencing the performance of these foundation models. However, these large foundation models often result in high computational cost. This paper focuses on pre-training relatively small vision transformer models that could be efficiently adapted to downstream tasks. Specifically, taking inspiration from knowledge distillation in model compression, we propose a new asymmetric masked distillation (AMD) framework for pre-training relatively small models with au-toencoding. The core of AMD is to devise an asymmetric masking strategy, where the teacher model is enabled to see more context information with a lower masking ratio, while the student model is still equipped with a high masking ratio. We design customized multi-layer feature alignment between the teacher encoder and student encoder to regularize the pre-training of student MAE. To demonstrate the effectiveness and versatility of AMD, we apply it to both ImageMAE and VideoMAE for pre-training relatively small ViT models. AMD achieved 84.6% classification accuracy on IN1K using the ViT-B model. And AMD achieves 73.3% classification accuracy using the ViT-B model on the Something-in-Something V2 dataset, a 3.7% improvement over the original ViT-B model from VideoMAE. We also transfer AMD pre-trained models to downstream tasks and obtain consistent performance improvement over the original masked autoencoding. The code and models are available at https://github.com/MCG-NJU/AMD. Bingkun Huang, Sen Xing, Gangshan Wu, Yu Qiao 0001, Limin Wang 0002 |
CVPR | 4 |
| 2024 | Dual DETRs for Multi-Label Temporal Action DetectionabstractTemporal Action Detection (TAD) aims to identify the action boundaries and the corresponding category within untrimmed videos. Inspired by the success of DETR in object detection, several methods have adapted the query-based framework to the TAD task. However, these approaches primarily followed DETR to predict actions at the instance level (i.e., identify each action by its center point), leading to sub-optimal boundary localization. To address this issue, we propose a new Dual-level query-based TAD framework, namely DualDETR, to detect actions from both instance-level and boundary-level. Decoding at different levels requires semantics of different granularity, therefore we introduce a two-branch decoding structure. This structure builds distinctive decoding processes for different lev-els, facilitating explicit capture of temporal cues and se-mantics at each level. On top of the two-branch design, we present a joint query initialization strategy to align queries from both levels. Specifically, we leverage encoder propos-als to match queries from each level in a one-to-one man-ner. Then, the matched queries are initialized using position and content prior from the matched action proposal. The aligned dual-level queries can refine the matched proposal with complementary cues during subsequent decoding. We evaluate DualDETR on three challenging multi-label TAD benchmarks. The experimental results demonstrate the su-perior performance of DualDETR to the existing state-of-the-art methods, achieving a substantial improvement under det-mAP and delivering impressive results under seg-mAP. Jing Tan 0002, Gangshan Wu, Limin Wang 0002 |
CVPR | 4 |
| 2024 | GTPT: Group-Based Token Pruning Transformer for Efficient Human Pose Estimation
Jie Liu 0040, Jie Tang 0006, Gangshan Wu, Yanbing Chou |
ECCV (69) | 4 |
| 2024 | RGB-D Video Object Segmentation via Enhanced Multi-store Feature MemoryabstractThe RGB-Depth (RGB-D) Video Object Segmentation (VOS) aims to integrate the fine-grained texture information of RGB with the spatial geometric clues of depth modality, boosting the performance of segmentation. However, off-the-shelf RGB-D segmentation methods fail to fully explore cross-modal information and suffer from object drift during long-term prediction. In this paper, we propose a novel RGB-D VOS method via multi-store feature memory for robust segmentation. Specifically, we design the hierarchical modality selection and fusion, which adaptively combines features from both modalities. Additionally, we develop a segmentation refinement module that effectively utilizes the Segmentation Anything Model (SAM) to refine the segmentation mask, ensuring more reliable results as memory to guide subsequent segmentation tasks. By leveraging spatio-temporal embedding and modality embedding, mixed prompts and fused images are fed into SAM to unleash its potential in RGB-D VOS. Experimental results show that the proposed method achieves state-of-the-art performance on the latest RGB-D VOS benchmark. Boyue Xu, Ruichao Hou, Tongwei Ren, Gangshan Wu |
ICMR | 4 |
| 2024 | RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual CuesabstractWhile existing Audio-Visual Speech Separation (AVSS) methods primarily concentrate on the audio-visual fusion strategy for two-speaker separation, they demonstrate a severe performance drop in the multi-speaker separation scenarios. Typically, AVSS methods employ guiding videos to sequentially isolate individual speakers from the given audio mixture, resulting in notable missing and noisy parts across various segments of the separated speech. In this study, we propose a simultaneous multi-speaker separation framework that can facilitate the concurrent separation of multiple speakers within a singular process. We introduce speaker-wise interactions to establish distinctions and correlations among speakers. Experimental results on the VoxCeleb2 and LRS3 datasets demonstrate that our method achieves state-of-the-art performance in separating mixtures with 2, 3, 4, and 5 speakers, respectively. Additionally, our model can utilize speakers with complete audio-visual information to mitigate other visual-deficient speakers, thereby enhancing its resilience to missing visual cues. We also conduct experiments where visual information for specific speakers is entirely absent or visual frames are partially missing. The results demonstrate that our model consistently outperforms others, exhibiting the smallest performance drop across all settings involving 2, 3, 4, and 5 speakers. Tianrui Pan, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
ACM Multimedia | 5 |
| 2024 | AWT: Transferring Vision-Language Models via Augmentation, Weighting, and TransportationabstractPre-trained vision-language models (VLMs) have shown impressive results in various visual classification tasks.
However, we often fail to fully unleash their potential when adapting them for new concept understanding due to limited information on new classes.
To address this limitation, we introduce a novel adaptation framework, AWT (Augment, Weight, then Transport). AWT comprises three key components: augmenting inputs with diverse visual perspectives and enriched class descriptions through image transformations and language models; dynamically weighting inputs based on the prediction entropy; and employing optimal transport to mine semantic correlations in the vision-language space.
AWT can be seamlessly integrated into various VLMs, enhancing their zero-shot capabilities without additional training and facilitating few-shot learning through an integrated multimodal adapter module.
We verify AWT in multiple challenging scenarios, including zero-shot and few-shot image classification, zero-shot video action recognition, and out-of-distribution generalization. AWT consistently outperforms the state-of-the-art methods in each setting. In addition, our extensive studies further demonstrate AWT's effectiveness and adaptability across different VLMs, architectures, and scales. Yuyang Ji, Gangshan Wu, Limin Wang 0002 |
NeurIPS | 4 |
| 2024 | Weakly supervised instance segmentation via peak mining and filteringabstractAbstract Learning the full extent of pixel‐level instance response in a weakly supervised manner remains unsatisfactory. Peak response maps (PRMs) localizes the discriminative object regions but cannot provide complete instance information, suffering from incomplete segmentation and unreliable mask prediction by noisy proposal retrieval. This work tackles this challenging problem by mining diverse class peak responses that include more discriminative and complete object regions and retrieving more reliable proposals from noisy segment proposal galleries. First, the existing method is enhanced with two more classification branches, thus contributing to more diverse and abundant instance regions from peak response maps. The mined class peak responses from two of the branches are then merged to generate more complete peak response maps by a clustering approach in their deep feature space. Then, instance segmentation masks are retrieved from a noisy object segment proposal gallery with class confidence, which is calculated by a normal classifier to obtain cleaner mask prediction. Finally, the pseudo‐supervision can be used to train an instance segmentation network in a fully supervised manner. Experiments on the PASCAL VOC 2012 dataset and COCO dataset show that the approach works effectively and outperforms other counterparts by a margin of more than 6 %, 4%, and 3% with the mean average precision (mAP) at IoU threshold of 0.25, 0.5 and 0.75, respectively. Zuxian Huang, Dongsheng Pan, Gangshan Wu |
IET Image Process. | 3 |
| 2024 | Dual Graph Networks for Pose Estimation in Crowded Scenes
Gangshan Wu, Limin Wang 0002 |
Int. J. Comput. Vis. | 2 |
| 2024 | MixFormer: End-to-End Tracking With Iterative Mixed AttentionabstractVisual object tracking often employs a multi-stage pipeline of feature extraction, target information integration, and bounding box estimation. To simplify this pipeline and unify the process of feature extraction and target information integration, in this paper, we present a compact tracking framework, termed as MixFormer, built upon transformers. Our core design is to utilize the flexibility of attention operations, and we propose a Mixed Attention Module (MAM) for simultaneous feature extraction and target information integration. This synchronous modeling scheme allows us to extract target-specific discriminative features and perform extensive communication between target and search area. Based on MAM, we build our MixFormer trackers simply by stacking multiple MAMs and placing a localization head on top. Specifically, we instantiate two types of MixFormer trackers, a hierarchical tracker MixCvT, and a non-hierarchical simple tracker MixViT. For these two trackers, we investigate a series of pre-training methods and uncover the different behaviors between supervised pre-training and self-supervised pre-training in our MixFormer trackers. We also extend the masked autoencoder pre-training to our MixFormer trackers and design the new competitive TrackMAE pre-training technique. Finally, to handle multiple target templates during online tracking, we devise an asymmetric attention scheme in MAM to reduce computational cost, and propose an effective score prediction module to select high-quality templates. Our MixFormer trackers set a new state-of-the-art performance on seven tracking benchmarks, including LaSOT, TrackingNet, VOT2020, GOT-10 k, OTB100, TOTB and UAV123. In particular, our MixViT-L achieves AUC scores of 73.3% on LaSOT, 86.1% on TrackingNet and 82.8% on TOTB. Yutao Cui, Cheng Jiang 0005, Gangshan Wu, Limin Wang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | STMixer: A One-Stage Sparse Action DetectorabstractTraditional video action detectors typically adopt the two-stage pipeline, where a person detector is first employed to generate actor boxes and then 3D RoIAlign is used to extract actor-specific features for action recognition. This detection paradigm requires multi-stage training and inference, and the feature sampling is only constrained inside the box, failing to effectively leverage richer context information outside. Recently, several query-based action detectors have been proposed to predict action instances in an end-to-end manner. However, they still lack adaptability in feature sampling and decoding, thus suffering from the issues of inferior performance or slower convergence. In this paper, we propose two core designs for a more flexible one-stage sparse action detector. First, we present a query-based adaptive feature sampling module, which endows the detector with the flexibility of mining a group of discriminative features from the entire spatio-temporal domain. Second, we devise a decoupled feature mixing module, which dynamically attends to and mixes video features along the spatial and temporal dimensions respectively for better feature decoding. Based on these designs, we instantiate two detection pipelines, that is, STMixer-K for keyframe action detection and STMixer-T for action tubelet detection. Without bells and whistles, our STMixer detectors obtain the state-of-the-art results on five challenging spatio-temporal action detection benchmarks for keyframe action detection or action tube detection. Tao Wu 0020, Mengqi Cao, Ziteng Gao, Gangshan Wu, Limin Wang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Jointly modeling association and motion cues for robust infrared UAV tracking
Boyue Xu, Ruichao Hou, Jia Bei, Tongwei Ren, Gangshan Wu |
Vis. Comput. | 5 |
| 2023 | From Coarse to Fine: Hierarchical Pixel Integration for Lightweight Image Super-resolutionabstractImage super-resolution (SR) serves as a fundamental tool for the processing and transmission of multimedia data. Recently, Transformer-based models have achieved competitive performances in image SR. They divide images into fixed-size patches and apply self-attention on these patches to model long-range dependencies among pixels. However, this architecture design is originated for high-level vision tasks, which lacks design guideline from SR knowledge. In this paper, we aim to design a new attention block whose insights are from the interpretation of Local Attribution Map (LAM) for SR networks. Specifically, LAM presents a hierarchical importance map where the most important pixels are located in a fine area of a patch and some less important pixels are spread in a coarse area of the whole image. To access pixels in the coarse area, instead of using a very large patch size, we propose a lightweight Global Pixel Access (GPA) module that applies cross-attention with the most similar patch in an image. In the fine area, we use an Intra-Patch Self-Attention (IPSA) module to model long-range pixel dependencies in a local patch, and then a spatial convolution is applied to process the finest details. In addition, a Cascaded Patch Division (CPD) strategy is proposed to enhance perceptual quality of recovered images. Extensive experiments suggest that our method outperforms state-of-the-art lightweight SR methods by a large margin. Code is available at https://github.com/passerer/HPINet. Jie Liu 0040, Chao Chen 0026, Jie Tang 0006, Gangshan Wu |
AAAI | 4 |
| 2023 | CoMAE: Single Model Hybrid Pre-training on Small-Scale RGB-D DatasetsabstractCurrent RGB-D scene recognition approaches often train two standalone backbones for RGB and depth modalities with the same Places or ImageNet pre-training. However, the pre-trained depth network is still biased by RGB-based models which may result in a suboptimal solution. In this paper, we present a single-model self-supervised hybrid pre-training framework for RGB and depth modalities, termed as CoMAE. Our CoMAE presents a curriculum learning strategy to unify the two popular self-supervised representation learning algorithms: contrastive learning and masked image modeling. Specifically, we first build a patch-level alignment task to pre-train a single encoder shared by two modalities via cross-modal contrastive learning. Then, the pre-trained contrastive encoder is passed to a multi-modal masked autoencoder to capture the finer context features from a generative perspective. In addition, our single-model design without requirement of fusion module is very flexible and robust to generalize to unimodal scenario in both training and testing phases. Extensive experiments on SUN RGB-D and NYUDv2 datasets demonstrate the effectiveness of our CoMAE for RGB and depth representation learning. In addition, our experiment results reveal that CoMAE is a data-efficient representation learner. Although we only use the small-scale and unlabeled training set for pre-training, our CoMAE pre-trained models are still competitive to the state-of-the-art methods with extra large-scale and supervised RGB dataset pre-training. Code will be released at https://github.com/MCG-NJU/CoMAE. Jiange Yang, Sheng Guo 0005, Gangshan Wu, Limin Wang 0002 |
AAAI | 3 |
| 2023 | LinK: Linear Kernel for LiDAR-based 3D PerceptionabstractExtending the success of 2D Large Kernel to 3D perception is challenging due to: 1. the cubically-increasing overhead in processing 3D data; 2. the optimization difficulties from data scarcity and sparsity. Previous work has taken the first step to scale up the kernel size from 3 × 3 × 3 to 7 × 7 × 7 by introducing block-shared weights. However, to reduce the feature variations within a block, it only employs modest block size and fails to achieve larger kernels like the 21 × 21 × 21. To address this issue, we propose a new method, called LinK, to achieve a wider-range perception receptive field in a convolution-like manner with two core designs. The first is to replace the static kernel matrix with a linear kernel generator, which adaptively provides weights only for non-empty voxels. The second is to reuse the pre-computed aggregation results in the overlapped blocks to reduce computation complexity. The proposed method successfully enables each voxel to perceive context within a range of 21 × 21 × 21. Extensive experiments on two basic perception tasks, 3D object detection and 3D semantic segmentation, demonstrate the effectiveness of our method. Notably, we rank 1st on the public leaderboard of the 3D detection benchmark of nuScenes (LiDAR track), by simply incorporating a LinK-based backbone into the basic detector, CenterPoint. We also boost the strong segmentation baseline's mIoU with 2.7% in the SemanticKITTI test set. Code is available at https://github.com/MCG-NJU/LinK. Tao Lu 0005, Haisong Liu, Gangshan Wu, Limin Wang 0002 |
CVPR | 4 |
| 2023 | STMixer: A One-Stage Sparse Action DetectorabstractTraditional video action detectors typically adopt the two-stage pipeline, where a person detector is first employed to generate actor boxes and then 3D RoIAlign is used to extract actor-specific features for classification. This detection paradigm requires multi-stage training and inference, and cannot capture context information outside the bounding box. Recently, a few query-based action detectors are proposed to predict action instances in an end-to-end manner. However, they still lack adaptability in feature sampling and decoding, thus suffering from the issues of inferior performance or slower convergence. In this paper, we propose a new one-stage sparse action detector, termed STMixer. STMixer is based on two core designs. First, we present a query-based adaptive feature sampling module, which endows our STMixer with the flexibility of mining a set of discriminative features from the entire spatiotemporal domain. Second, we devise a dual-branch feature mixing module, which allows our STMixer to dynamically attend to and mix video features along the spatial and the temporal dimension respectively for better feature decoding. Coupling these two designs with a video backbone yields an efficient end-to-end action detector. Without bells and whistles, our STMixer obtains the state-of-the-art results on the datasets of AVA, UCF101-24, and JHMDB. Tao Wu 0020, Mengqi Cao, Ziteng Gao, Gangshan Wu, Limin Wang 0002 |
CVPR | 4 |
| 2023 | Extracting Motion and Appearance via Inter-Frame Attention for Efficient Video Frame InterpolationabstractEffectively extracting inter-frame motion and appearance information is important for video frame interpolation (VFI). Previous works either extract both types of information in a mixed way or devise separate modules for each type of information, which lead to representation ambiguity and low efficiency. In this paper, we propose a new module to explicitly extract motion and appearance information via a unified operation. Specifically, we rethink the information process in inter-frame attention and reuse its attention map for both appearance feature enhancement and motion information extraction. Furthermore, for efficient VFI, our proposed module could be seamlessly integrated into a hybrid CNN and Transformer architecture. This hybrid pipeline can alleviate the computational complexity of inter-frame attention as well as preserve detailed low-level structure information. Experimental results demonstrate that, for both fixed- and arbitrary-timestep interpolation, our method achieves state-of-the-art performance on various datasets. Meanwhile, our approach enjoys a lighter computation overhead over models with close performance. The source code and models are available at https://github.com/MCG-NJU/EMA-VFI. Youxin Chen, Gangshan Wu, Limin Wang 0002 |
CVPR | 5 |
| 2023 | Single Image Super-Resolution with Sequential Multi-axis Blocked Attention
Bincheng Yang, Gangshan Wu |
ICANN (3) | 2 |
| 2023 | Robust Object Modeling for Visual TrackingabstractObject modeling has become a core part of recent tracking frameworks. Current popular tackers use Transformer attention to extract the template feature separately or interactively with the search region. However, separate template learning lacks communication between the template and search regions, which brings difficulty in extracting discriminative target-oriented features. On the other hand, interactive template learning produces hybrid template features, which may introduce potential distractors to the template via the cluttered search regions. To enjoy the merits of both methods, we propose a robust object modeling framework for visual tracking (ROMTrack), which simultaneously models the inherent template and the hybrid template features. As a result, harmful distractors can be suppressed by combining the inherent features of target objects with search regions’ guidance. Target-related features can also be extracted using the hybrid template, thus resulting in a more robust object modeling framework. To further enhance robustness, we present novel variation tokens to depict the ever-changing appearance of target objects. Variation tokens are adaptable to object deformation and appearance variations, which can boost overall performance with negligible computation. Experiments show that our ROMTrack sets a new state-of-the-art on multiple benchmarks. Yidong Cai, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
ICCV | 4 |
| 2023 | Efficient Video Action Detection with Token Dropout and Context RefinementabstractStreaming video clips with large-scale video tokens impede vision transformers (ViTs) for efficient recognition, especially in video action detection where sufficient spatiotemporal representations are required for precise actor identification. In this work, we propose an end-to-end framework for efficient video action detection (EVAD) based on vanilla ViTs. Our EVAD consists of two specialized designs for video action detection. First, we propose a spatiotemporal token dropout from a keyframe-centric perspective. In a video clip, we maintain all tokens from its keyframe, preserve tokens relevant to actor motions from other frames, and drop out the remaining tokens in this clip. Second, we refine scene context by leveraging remaining tokens for better recognizing actor identities. The region of interest (RoI) in our action detector is expanded into temporal domain. The captured spatiotemporal actor identity representations are refined via scene context in a decoder with the attention mechanism. These two designs make our EVAD efficient while maintaining accuracy, which is validated on three benchmark datasets (i.e., AVA, UCF101-24, JHMDB). Compared to the vanilla ViT backbone, our EVAD reduces the overall GFLOPs by 43% and improves real-time inference speed by 40% with no performance degradation. Moreover, even at similar computational costs, our EVAD can improve the performance by 1.1 mAP with higher resolution inputs. Code is available at https://github.com/MCG-NJU/EVAD. Zhan Tong, Yibing Song, Gangshan Wu, Limin Wang 0002 |
ICCV | 4 |
| 2023 | SportsMOT: A Large Multi-Object Tracking Dataset in Multiple Sports ScenesabstractMulti-object tracking (MOT) in sports scenes plays a critical role in gathering players statistics, supporting further applications, such as automatic tactical analysis. Yet existing MOT benchmarks cast little attention on this domain. In this work, we present a new large-scale multi-object tracking dataset in multiple sports scenes, coined as SportsMOT, where all players on the court are supposed to be tracked. It consists of 240 video sequences, over 150K frames (almost 15x MOT17) and over 1.6M bounding boxes (3x MOT17) collected from 3 sports categories, including basketball, volleyball and football. Our dataset is characterized with two key properties: 1) fast and variable-speed motion and 2) similar yet distinguishable appearance. We expect SportsMOT to encourage the MOT trackers to promote in both motion-based association and appearance-based association. We benchmark several state-of-the-art trackers and reveal the key challenge of SportsMOT lies in object association. To alleviate the issue, we further propose a new multi-object tracking framework, termed as MixSort, introducing a MixFormer-like structure as an auxiliary association model to prevailing tracking-by-detection trackers. By integrating the customized appearance-based association with the original motion-based association, MixSort achieves state-of-the-art performance on SportsMOT and MOT17. Based on MixSort, we give an in-depth analysis and provide some profound insights into SportsMOT. Yutao Cui, Chenkai Zeng, Yichun Yang, Gangshan Wu, Limin Wang 0002 |
ICCV | 5 |
| 2023 | MTNet: Learning Modality-aware Representation with Transformer for RGBT TrackingabstractThe ability to learn robust multi-modality representation has played a critical role in the development of RGBT tracking. However, the regular fusion paradigm and the invariable tracking template remain restrictive to the feature interaction. In this paper, we propose a modality-aware tracker based on transformer, termed MTNet. Specifically, a modality- aware network is presented to explore modality-specific cues, which contains both channel aggregation and distribution module (CADM) and spatial similarity perception module (SSPM). A transformer fusion network is then applied to capturing global dependencies to reinforce instance representations. To estimate the precise location and tackle the challenges, such as scale variation and deformation, we design a trident prediction head and a dynamic update strategy which jointly maintain a reliable template for facilitating inter-frame communication. Extensive experiments validate that the proposed method achieves satisfactory results compared with the state-of-the-art competitors on three RGBT benchmarks while reaching real-time speed. Ruichao Hou, Boyue Xu, Tongwei Ren, Gangshan Wu |
ICME | 4 |
| 2023 | Video Frame Interpolation with Densely Queried Bilateral CorrelationabstractVideo Frame Interpolation (VFI) aims to synthesize non-existent intermediate frames between existent frames. Flow-based VFI algorithms estimate intermediate motion fields to warp the existent frames. Real-world motions' complexity and the reference frame's absence make motion estimation challenging. Many state-of-the-art approaches explicitly model the correlations between two neighboring frames for more accurate motion estimation. In common approaches, the receptive field of correlation modeling at higher resolution depends on the motion fields estimated beforehand. Such receptive field dependency makes common motion estimation approaches poor at coping with small and fast-moving objects. To better model correlations and to produce more accurate motion fields, we propose the Densely Queried Bilateral Correlation (DQBC) that gets rid of the receptive field dependency problem and thus is more friendly to small and fast-moving objects. The motion fields generated with the help of DQBC are further refined and up-sampled with context features. After the motion fields are fixed, a CNN-based SynthNet synthesizes the final interpolated frame. Experiments show that our approach enjoys higher accuracy and less inference time than the state-of-the-art. Source code is available at https://github.com/kinoud/DQBC. Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
IJCAI | 4 |
| 2023 | Deep Video Understanding with Video-Language ModelabstractPre-trained video-language models (VLMs) have shown superior performance in high-level video understanding tasks, analyzing multi-modal information, aligning with Deep Video Understanding Challenge (DVUC) requirements.In this paper, we explore pre-trained VLMs' potential in multimodal question answering for long-form videos. We propose a solution called Dual Branches Video Modeling (DBVM), which combines knowledge graph (KG) and VLMs, leveraging their strengths and addressing shortcomings.The KG branch recognizes and localizes entities, fuses multimodal features at different levels, and constructs KGs with entities as nodes and relationships as edges.The VLM branch applies a selection strategy to adapt input movies into acceptable length and a cross-matching strategy to post-process results providing accurate scene descriptions.Experiments conducted on the DVUC dataset validate the effectiveness of our DBVM. Yaqun Fang, Fan Yu 0003, Ruiqi Tian, Tongwei Ren, Gangshan Wu |
ACM Multimedia | 6 |
| 2023 | Lightweight Super-Resolution Head for Human Pose EstimationabstractHeatmap-based methods have become the mainstream method for pose estimation due to their superior performance. However, heatmap-based approaches suffer from significant quantization errors with downscale heatmaps, which result in limited performance and the detrimental effects of intermediate supervision. Previous heatmap-based methods relied heavily on additional post-processing to mitigate quantization errors. Some heatmap-based approaches improve the resolution of feature maps by using multiple costly upsampling layers to improve localization precision. To solve the above issues, we creatively view the backbone network as a degradation process and thus reformulate the heatmap prediction as a Super-Resolution (SR) task. We first propose the SR head, which predicts heatmaps with a spatial resolution higher than the input feature maps (or even consistent with the input image) by super-resolution, to effectively reduce the quantization error and the dependence on further post-processing. Besides, we propose SRPose to gradually recover the HR heatmaps from LR heatmaps and degraded features in a coarse-to-fine manner. To reduce the training difficulty of HR heatmaps, SRPose applies SR heads to supervise the intermediate features in each stage. In addition, the SR head is a lightweight and generic head that applies to top-down and bottom-up methods. Extensive experiments on the COCO, MPII, and CrowdPose datasets show that SRPose outperforms the corresponding heatmap-based approaches. Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
ACM Multimedia | 4 |
| 2023 | ADNet: An Asymmetric Dual-Stream Network for RGB-T Salient Object DetectionabstractRGB-Thermal salient object detection (RGB-T SOD) aims to locate salient objects in images that include both RGB and thermal information. Previous approaches often suggest designing a symmetric network structure to tackle the challenge of dealing with low-quality RGB or thermal images. However, we contend that RGB and thermal modalities possess different numbers of channels and disparities in information density. In this paper, we propose a novel asymmetric dual-stream network (ADNet). Specifically, we leverage an asymmetric backbone to extract four stages of RGB features and four stages of thermal features. To enable effective interaction among low-level features in the first two stages, we introduce the Channel-Spatial Interaction (CSI) module. In the last two stages, deep features are enhanced using the Self-Attention Enhancement (SAE) module. Experimental results on the VT5000, VT1000, and VT821 datasets attest to the superior performance of our proposed ADNet compared to state-of-the-art methods. Yaqun Fang, Ruichao Hou, Jia Bei, Tongwei Ren, Gangshan Wu |
MMAsia | 5 |
| 2023 | RGB-D Tracking via Hierarchical Modality Aggregation and Distribution NetworkabstractThe integration of dual-modal features has been pivotal in advancing RGB-Depth (RGB-D) tracking. However, current trackers are less efficient and focus solely on single-level features, resulting in weaker robustness in fusion and slower speeds that fail to meet the demands of real-world applications. In this paper, we introduce a novel network, denoted as HMAD (Hierarchical Modality Aggregation and Distribution), which addresses these challenges. HMAD leverages the distinct feature representation strengths of RGB and depth modalities, giving prominence to a hierarchical approach for feature distribution and fusion, thereby enhancing the robustness of RGB-D tracking. Experimental results on various RGB-D datasets demonstrate that HMAD achieves state-of-the-art performance. Moreover, real-world experiments further validate HMAD’s capacity to effectively handle a spectrum of tracking challenges in real-time scenarios. Boyue Xu, Ruichao Hou, Jia Bei, Tongwei Ren, Gangshan Wu |
MMAsia | 6 |
| 2023 | MixFormerV2: Efficient Fully Transformer TrackingabstractTransformer-based trackers have achieved strong accuracy on the standard benchmarks. However, their efficiency remains an obstacle to practical deployment on both GPU and CPU platforms. In this paper, to overcome this issue, we propose a fully transformer tracking framework, coined as \emph{MixFormerV2}, without any dense convolutional operation and complex score prediction module. Our key design is to introduce four special prediction tokens and concatenate them with the tokens from target template and search areas. Then, we apply the unified transformer backbone on these mixed token sequence. These prediction tokens are able to capture the complex correlation between target template and search area via mixed attentions. Based on them, we can easily predict the tracking box and estimate its confidence score through simple MLP heads. To further improve the efficiency of MixFormerV2, we present a new distillation-based model reduction paradigm, including dense-to-sparse distillation and deep-to-shallow distillation. The former one aims to transfer knowledge from the dense-head based MixViT to our fully transformer tracker, while the latter one is used to prune some layers of the backbone. We instantiate two types of MixForemrV2, where the MixFormerV2-B achieves an AUC of 70.6\% on LaSOT and AUC of 56.7\% on TNL2k with a high GPU speed of 165 FPS, and the MixFormerV2-S surpasses FEAR-L by 2.7\% AUC on LaSOT with a real-time CPU speed. Yutao Cui, Tianhui Song, Gangshan Wu, Limin Wang 0002 |
NeurIPS | 3 |
| 2023 | Webly-supervised semantic segmentation via curriculum learning
Zuxian Huang, Gangshan Wu, Limin Wang 0002 |
Comput. Vis. Image Underst. | 2 |
| 2023 | LIP: Local Importance-Based Pooling
Ziteng Gao, Limin Wang 0002, Gangshan Wu |
Int. J. Comput. Vis. | 3 |
| 2023 | Temporal Perceiver: A General Architecture for Arbitrary Boundary DetectionabstractGeneric Boundary Detection (GBD) aims at locating the general boundaries that divide videos into semantically coherent and taxonomy-free units, and could serve as an important pre-processing step for long-form video understanding. Previous works often separately handle these different types of generic boundaries with specific designs of deep networks from simple CNN to LSTM. Instead, in this paper, we presentTemporal Perceiver, a general architecture with Transformer, offering a unified solution to the detection of arbitrary generic boundaries, ranging from shot-level, event-level, to scene-level GBDs. The core design is to introduce a small set of latent feature queries as anchors to compress the redundant video input into a fixed dimension via cross-attention blocks. Thanks to this fixed number of latent units, it greatly reduces the quadratic complexity of attention operation to a linear form of input frames. Specifically, to explicitly leverage the temporal structure of videos, we construct two types of latent feature queries: boundary queries and context queries, which handle the semantic incoherence and coherence accordingly. Moreover, to guide the learning of latent feature queries, we propose an alignment loss on the cross-attention maps to explicitly encourage the boundary queries to attend on the top boundary candidates. Finally, we present a sparse detection head on the compressed representation, and directly output the final boundary detection results without any post-processing module. We test our Temporal Perceiver on a variety of GBD benchmarks. Our method obtains the state-of-the-art results on all benchmarks with RGB single-stream features: SoccerNet-v2 (81.9 percent average-mAP), Kinetics-GEBD (86.0 percent average-f1), TAPOS (73.2 percent average-f1), MovieScenes (51.9 percent AP and 53.1 percent$M_{iou}$) and MovieNet (53.3 percent AP and 53.2 percent$M_{iou}$), demonstrating the generalization ability of our Temporal Perceiver. To further pursue a general GBD model, we combined various tasks to train a class-agnostic Temporal perceiver and evaluate its performance across all benchmarks. Results show that the class-agnostic Perceiver achieves comparable detection accuracy and even better generalization ability compared to dataset-specific Temporal Perceiver. Jing Tan 0002, Gangshan Wu, Limin Wang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | APP-Net: Auxiliary-Point-Based Push and Pull Operations for Efficient Point Cloud RecognitionabstractAggregating neighbor features is essential for point cloud neural network. In the existing work, each point in the cloud may inevitably be selected as the neighbors of multiple aggregation centers, as all centers will gather neighbor features from the whole point cloud independently. Thus, each point has to participate in the calculation repeatedly, generating redundant duplicates in the memory, leading to intensive computation costs and memory consumption. Meanwhile, to pursue higher accuracy, previous methods often rely on a complex local aggregator to extract fine geometric representation, further slowing down the processing pipeline. To address these issues, we propose a new local aggregator of linear complexity for point cloud analysis, coined as APP. Specifically, we introduce an auxiliary container as an anchor to exchange features between the source point and the aggregating center. Each source point pushes its feature to only one auxiliary container, and each center point pulls features from only one auxiliary container. This avoids the re-computation issue of each source point. To facilitate the learning of the local structure of point cloud, we use an online normal estimation module to provide explainable geometric information to enhance our APP modeling capability. Our built network is more efficient than all the previous baselines with a clear margin while still consuming a lower memory. Experiments on classification and semantic segmentation demonstrate that APP-Net reaches comparable accuracies to other networks. In the classification task, it can process more than 10,000 samples per second with less than 10GB of memory on a single GPU. We will release the code at https://github.com/MCG-NJU/ APP-Net. Tao Lu 0005, Chunxu Liu, Youxin Chen, Gangshan Wu, Limin Wang 0002 |
IEEE Trans. Image Process. | 4 |
| 2023 | Efficient Single-image Super-resolution Using Dual path Connections with Multiple scale LearningabstractDeep convolutional neural networks have been demonstrated to be effective for single-image super-resolution in recent years. On the one hand, residual connections and dense connections have been used widely to ease forward information and backward gradient flows to boost performance. However, current methods use residual connections and dense connections separately in most network layers in a sub-optimal way. On the other hand, although various networks and methods have been designed to improve computation efficiency, save parameters, or utilize training data of multiple scale factors for each other to boost performance, they either do super-resolution in high-resolution space to have a high computation cost or cannot share parameters between models of different scale factors to save parameters and inference time. To tackle these challenges, we propose an efficient single-image super-resolution network using dual path connections with multiple scale learning (EMSRDPN). By introducing dual path connections inspired by Dual path Networks into EMSRDPN, it uses residual connections and dense connections in an integrated way in most network layers. Dual path connections have the benefits of both reusing common features of residual connections and exploring new features of dense connections to learn a good representation for single-image super-resolution. To utilize the feature correlation of multiple scale factors, EMSRDPN shares all network units in low-resolution space between different scale factors to learn shared features and only uses a separate reconstruction unit for each scale factor, which can utilize training data of multiple scale factors to help each other to boost performance, meanwhile, which can save parameters and support shared inference for multiple scale factors to improve efficiency. Experiments show EMSRDPN achieves better performance and comparable or even better parameter and inference efficiency over state-of-the-art methods. Code will be available at https://github.com/yangbincheng/EMSRDPN . Bincheng Yang, Gangshan Wu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Negative Sample Matters: A Renaissance of Metric Learning for Temporal GroundingabstractTemporal grounding aims to localize a video moment which is semantically aligned with a given natural language query. Existing methods typically apply a detection or regression pipeline on the fused representation with the research focus on designing complicated prediction heads or fusion strategies. Instead, from a perspective on temporal grounding as a metric-learning problem, we present a Mutual Matching Network (MMN), to directly model the similarity between language queries and video moments in a joint embedding space. This new metric-learning framework enables fully exploiting negative samples from two new aspects: constructing negative cross-modal pairs in a mutual matching scheme and mining negative pairs across different videos. These new negative samples could enhance the joint representation learning of two modalities via cross-modal mutual matching to maximize their mutual information. Experiments show that our MMN achieves highly competitive performance compared with the state-of-the-art methods on four video grounding benchmarks. Based on MMN, we present a winner solution for the HC-STVG challenge of the 3rd PIC workshop. This suggests that metric learning is still a promising method for temporal grounding via capturing the essential cross-modal correlation in a joint embedding space. Code is available at https://github.com/MCG-NJU/MMN. Zhenzhi Wang 0001, Limin Wang 0002, Tao Wu 0020, Gangshan Wu |
AAAI | 5 |
| 2022 | MixFormer: End-to-End Tracking with Iterative Mixed AttentionabstractTracking often uses a multistage pipeline of feature extraction, target information integration, and bounding box estimation. To simplify this pipeline and unify the process of feature extraction and target information integration, we present a compact tracking framework, termed as MixFormer, built upon transformers. Our core design is to utilize the flexibility of attention operations, and propose a Mixed Attention Module (MAM) for simultaneous feature extraction and target information integration. This synchronous modeling scheme allows to extract target-specific discriminative features and perform extensive communication between target and search area. Based on MAM, we build our MixFormer tracking framework simply by stacking multiple MAMs with progressive patch embedding and placing a localization head on top. In addition, to handle multiple target templates during online tracking, we devise an asymmetric attention scheme in MAM to reduce computational cost, and propose an effective score prediction module to select high-quality templates. Our MixFormer sets a new state-of-the-art performance on five tracking benchmarks, including LaSOT, TrackingNet, VOT2020, GOT-10k, and UAV123. In particular, our MixFormer-L achieves NP score of 79.9% on LaSOT, 88.9% on TrackingNet and EAO of 0.555 on VOT2020. We also perform in-depth ablation studies to demonstrate the effectiveness of simultaneous feature extraction and information integration. Code and trained models are publicly available at https://github.com/MCG-NJU/MixFormer. Yutao Cui, Cheng Jiang 0005, Limin Wang 0002, Gangshan Wu |
CVPR | 4 |
| 2022 | Two Strategies Toward Lightweight Image Super-ResolutionabstractRecent convolution neural networks (CNNs) have achieved remarkable success in lightweight image super-resolution (LISR). The goal of LISR is to restore more accurate details with less model capacity. However, we observe two phenomena in current micro-architectures, one is the lack of consistent learning ability of high-frequency components, the other is large residual problem which does harm to the stability of residual learning. To tackle the two issues, we propose two strategies, namely global-guided attention strategy (GGAS) and channel-wise scaling strategy (CWSS), which can significantly improve the performance of the state-of-the-arts with negligible overheads. Zongcai Du, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
ICASSP | 4 |
| 2022 | Pyramid Fusion Attention Network For Single Image Super-ResolutionabstractRecently, convolutional neural network (CNN) has made a mighty advance in image super-resolution (SR). Most recent models exploit attention mechanism (AM) to focus on high-frequency information. However, these methods exclusively consider interdependencies among channels or spatials, leading to equal treatment of channel-wise or spatial-wise features thus hindering the power of AM. In this paper, we propose a pyramid fusion attention network (PFAN) to tackle this problem. Specifically, a novel pyramid fusion attention (PFA) is developed where stacked residual blocks are employed to model the relationship between pixels among all channels, and pyramid fusion structure is adopted to expand receptive field. Besides, a progressive backward fusion strat-egy is introduced to make full use of hierarchical features, which are beneficial to obtaining more contextual representations. Comprehensive experiments demonstrate the superiority of our proposed PFAN against state-of-the-art methods. Zongcai Du, Jie Tang 0006, Gangshan Wu |
ICASSP | 5 |
| 2022 | Hierarchical Feature Aggregation Network for Deep Image CompressionabstractExisting CNN-based methods for image compression extract features through serially connected high-to-low (encoder) or low-to-high (decoder) resolution stages, leading to insufficient utilization of hierarchical features. To solve this problem, we present a hierarchical feature aggregation network (HFAN) for generating more informative latent representations. In detail, we propose two strategies, namely inter-stage feature aggregation and intra-stage feature aggregation. The inter-stage feature aggregation integrates multi-scale information thereby producing more contextual features. The intra-stage aggregation fuses features within the same stage to enrich representations of one specific resolution. Besides, we incorporate a lightweight pixel-wise attention mechanism to further enhance the discriminative ability of our network. Extensive experiments demonstrate that our HFAN achieves superior performance over state-of-the-art methods without a hyperprior variational autoencoder. Zongcai Du, Jie Tang 0006, Gangshan Wu |
ICASSP | 5 |
| 2022 | MIRNet: A Robust RGBT Tracking Jointly with Multi-Modal Interaction and RefinementabstractRGBT tracking attempts to design a robust all-weather tracker by integrating the complementary features of visible and thermal spectrums. To explore the latent interdependencies across modalities, we propose a novel real-time tracker named MIR-Net, which contains a multi-modal interaction module (MIM) and a refinement mechanism (RM), thereby adaptively merging multi-modal features and achieving precise scale estimation. Specifically, to enhance instance representation in low-quality modality, the MIM reinforces discriminative features from one modality to another in a bidirectional way. Considering the negative effects of unreliable modality, we further introduce a gate function in MIM to filter redundancy. To address the problem of random drifting and estimate the precise scale in the online tracking, we present a well-designed RM that combines optical flow and refinement network. Comprehensive experiments on two public RGBT benchmarks validate that our tracker outperforms the state-of-the-art methods. Ruichao Hou, Tongwei Ren, Gangshan Wu |
ICME | 3 |
| 2022 | Reproducibility Companion Paper: Human Object Interaction Detection via Multi-level Conditioned NetworkabstractTo support the replication of ?Human Object Interaction Detection via Multi-level Conditioned Network", which was presented at ICMR'20, this companion paper provides the details of the artifacts. Human Object Interaction Detection (HOID) aims to recognize fine-grained object-specific human actions, which demands the capabilities of both visual perception and reasoning. In this paper, we explain the file structure of the source code and publish the details of our experiments settings. We also provide a program for component analysis to assist other researchers with experiments on alternative models that are not included in our experiments. Moreover, we provide a demo program for facilitating the use of our model. Yunqing He, Xu Sun 0009, Tongwei Ren, Gangshan Wu, Maria Sinziana Astefanoaei, Andreas Leibetseder |
ICMR | 5 |
| 2022 | Heterogeneous Learning for Scene Graph GenerationabstractScene Graph Generation (SGG) task aims to construct a graph structure to express objects and their relationships in a scene at a holistic level. Due to the neglect of heterogeneity of feature spaces between objects and relations, coupling of feature representations becomes obvious in current SGG methods, which results in large intra-class variation and inter-class ambiguity. In order to explicitly emphasize the heterogeneity in SGG, we propose a plug-and-play Heterogeneous Learning Branch (HLB), which enhances the independent representation capability of relation features. The HLB actively obscures the interconnection between objects and relation feature spaces via gradient reversal, with the assistance of a link prediction module as information barrier and an Auto Encoder for information preservation. To validate the effectiveness of HLB, we apply HLB to typical SGG methods in which the feature spaces are either homogeneous or semi-heterogeneous, and conduct evaluation on VG-150 dataset. The experimental results demonstrate that HLB significantly improves the performance of all these methods in the common evaluation criteria for SGG task. Yunqing He, Tongwei Ren, Jinhui Tang 0001, Gangshan Wu |
ACM Multimedia | 4 |
| 2022 | Multimodal Analysis for Deep Video Understanding with Video Language TransformerabstractThe Deep Video Understanding Challenge (DVUC) is aimed to use multiple modality information to build high-level understanding of video, involving tasks such as relationship recognition and interaction detection. In this paper, we use a joint learning framework to simultaneously predict multiple tasks with visual, text, audio and pose features. In addition, to answer the queries of DVUC, we design multiple answering strategies and use video language transformer which learns cross-modal information for matching videos with text choices. The final DVUC result shows that our method ranks first for group one of movie-level queries, and ranks third for both of group one and group two of scene-level queries. Beibei Zhang 0005, Yaqun Fang, Tongwei Ren, Gangshan Wu |
ACM Multimedia | 4 |
| 2022 | SpotFormer: A Transformer-based Framework for Precise Soccer Action SpottingabstractAction spotting and classification consist in detecting the exact moments at which events occur in long videos. The current mainstream spotting practices generally use a two-stage pipeline that performs feature collection and integration, then salient action detection and postprocessing. Following that, we present SpotFormer, a simple yet effective framework, capable of precise action spotting. Specifically, we employ several most advanced backbone networks as auxiliary feature extractors, and reduce feature dimensionality in a straightforward and efficient way. The frame-wise features are fed into a transformer-based spotting network devised to leverage spatiotemporal information. We obtain 0.609 tight mAP score via model ensemble and achieve the state-of-the-art performance on the SoccerNet-v2 dataset. Mengqi Cao, Min Yang 0011, Yilu Wu, Gangshan Wu, Limin Wang 0002 |
MMSP | 6 |
| 2022 | IAA-VSR: An iterative alignment algorithm for video super-resolution
Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
Appl. Intell. | 3 |
| 2022 | Fully convolutional online tracking
Yutao Cui, Cheng Jiang 0005, Limin Wang 0002, Gangshan Wu |
Comput. Vis. Image Underst. | 4 |
| 2022 | Cross-Domain Gated Learning for Domain Generalization
Dapeng Du, Jiawei Chen 0009, Yuexiang Li, Kai Ma 0002, Gangshan Wu, Yefeng Zheng 0001, Limin Wang 0002 |
Int. J. Comput. Vis. | 5 |
| 2021 | A Closer Look at Few-Shot Video Classification: A New Baseline and Benchmark
Zhenxi Zhu, Limin Wang 0002, Sheng Guo 0005, Gangshan Wu |
BMVC | 4 |
| 2021 | TDN: Temporal Difference Networks for Efficient Action RecognitionabstractTemporal modeling still remains challenging for action recognition in videos. To mitigate this issue, this paper presents a new video architecture, termed as Temporal Difference Network (TDN), with a focus on capturing multi-scale temporal information for efficient action recognition. The core of our TDN is to devise an efficient temporal module (TDM) by explicitly leveraging a temporal difference operator, and systematically assess its effect on short-term and long-term motion modeling. To fully capture temporal information over the entire video, our TDN is established with a two-level difference modeling paradigm. Specifically, for local motion modeling, temporal difference over consecutive frames is used to supply 2D CNNs with finer motion pattern, while for global motion modeling, temporal difference across segments is incorporated to capture long-range structure for motion feature excitation. TDN provides a simple and principled temporal modeling framework and could be instantiated with the existing CNNs at a small extra computational cost. Our TDN presents a new state of the art on the Something-Something V1 & V2 datasets and is on par with the best performance on the Kinetics-400 dataset. In addition, we conduct in-depth ablation studies and plot the visualization results of our TDN, hopefully providing insightful analysis on temporal difference modeling. We release the code at https://github.com/MCG-NJU/TDN. Limin Wang 0002, Zhan Tong, Gangshan Wu |
CVPR | 4 |
| 2021 | CGA-Net: Category Guided Aggregation for Point Cloud Semantic SegmentationabstractPrevious point cloud semantic segmentation networks use the same process to aggregate features from neighbors of the same category and different categories. However, the joint area between two objects usually only occupies a small percentage in the whole scene. Thus the networks are well- trained for aggregating features from the same category point while not fully trained on aggregating points of different categories. To address this issue, this paper proposes to utilize different aggregation strategies between the same category and different categories. Specifically, it presents a customized module, termed as Category Guided Aggregation (CGA), where it first identifies whether the neighbors belong to the same category with the center point or not, and then handles the two types of neighbors with two carefully-designed modules. Our CGA presents a general network module and could be leveraged in any existing semantic segmentation network. Experiments on three different backbones demonstrate the effectiveness of our method. Tao Lu 0005, Limin Wang 0002, Gangshan Wu |
CVPR | 3 |
| 2021 | Learning Discriminative Features for Semi-Supervised Anomaly DetectionabstractAnomaly detection is the task of identifying unusual samples in data. Typically anomaly detection is defined on an unlabeled dataset that is assumed most of the samples are normal and others are anomalies. However, in industrial practice, one may have access to a part of annotated data. This gives us the potential for semi-supervised learning. In addition, existing methods assume all training data is normal and neglect the impact of a small number of anomalous samples. In this paper, we consolidate the model’s discriminative power by introducing a transfer learning scheme to anomaly detection, thereby the model suffers less perturbation caused by pollution. We also propose a novel loss function to further adapt to semi-supervised data scenario. We ensure that the contribution of pollution can be well suppressed and reach a harmonious balance in magnitude of loss/gradient between unlabeled and labeled samples. Experiments on three publicly available datasets show that our method achieves state-of-the-art results. Jie Tang 0006, Yishun Dou, Gangshan Wu |
ICASSP | 4 |
| 2021 | Lightweight Human Pose Estimation under Resource-Limited ScenesabstractRecent research on human pose estimation has achieved significant improvement. However, most existing methods tend to pursue higher scores on benchmark datasets using complex architecture, ignoring the deployment costs in practice. In this paper, we investigate the problem of lightweight human pose estimation under resource-limited scenes.We first redesign a lightweight bottleneck block with two concepts: depthwise convolution and attention mechanism. And then, based on the lightweight block, we present a single-stage Lightweight Pose Network (LPN). Our small network LPN-50 only has 2.7M parameters and 1.0G FLOPs, which is much more lightweight than other popular networks. In order to overcome the training barrier, we propose an iterative training strategy that can give full play to our LPNs’ potential to get more accurate predicted results. We empirically demonstrate the effectiveness and efficiency of our methods on the benchmark dataset: the COCO keypoint detection dataset. Besides, we show the speed superiority of our lightweight network at inference time on a non-GPU platform. Specifically, our LPN-50 can achieve 68.7 in AP score on the COCO test-dev set, with 17 FPS inference speed on an Intel i7-8700K (6 cores) CPU machine. Jie Tang 0006, Gangshan Wu |
ICASSP | 3 |
| 2021 | Mutual Supervision for Dense Object DetectionabstractThe classification and regression head are both indispensable components to build up a dense object detector, which are usually supervised by the same training samples and thus expected to have consistency with each other for detecting objects accurately in the detection pipeline. In this paper, we break the convention of the same training samples for these two heads in dense detectors and explore a novel supervisory paradigm, termed as Mutual Supervision (MuSu), to respectively and mutually assign training samples for the classification and regression head to ensure this consistency. MuSu defines training samples for the regression head mainly based on classification predicting scores and in turn, defines samples for the classification head based on localization scores from the regression head. Experimental results show that the convergence of detectors trained by this mutual supervision is guaranteed and the effectiveness of the proposed method is verified on the challenging MS COCO benchmark. We also find that tiling more anchors at the same location benefits detectors and leads to further improvements under this training scheme. We hope this work can inspire further researches on the interaction of the classification and regression task in detection and the supervision paradigm for detectors, especially separately for these two heads. Ziteng Gao, Limin Wang 0002, Gangshan Wu |
ICCV | 3 |
| 2021 | Self Supervision to Distillation for Long-Tailed Visual RecognitionabstractDeep learning has achieved remarkable progress for visual recognition on large-scale balanced datasets but still performs poorly on real-world long-tailed data. Previous methods often adopt class re-balanced training strategies to effectively alleviate the imbalance issue, but might be a risk of over-fitting tail classes. The recent decoupling method overcomes over-fitting issues by using a multi-stage training scheme, yet, it is still incapable of capturing tail class information in the feature learning stage. In this paper, we show that soft label can serve as a powerful solution to incorporate label correlation into a multi-stage training scheme for long-tailed recognition. The intrinsic relation between classes embodied by soft labels turns out to be helpful for long-tailed recognition by transferring knowledge from head to tail classes.Specifically, we propose a conceptually simple yet particularly effective multi-stage training scheme, termed as Self Supervised to Distillation (SSD). This scheme is composed of two parts. First, we introduce a self-distillation framework for long-tailed recognition, which can mine the label relation automatically. Second, we present a new distillation label generation module guided by self-supervision. The distilled labels integrate information from both label and data domains that can model long-tailed distribution effectively. We conduct extensive experiments and our method achieves the state-of-the-art results on three long-tailed recognition benchmarks: ImageNet-LT, CIFAR100-LT and iNaturalist 2018. Our SSD outperforms the strong LWS baseline by from 2.7% to 4.5% on various datasets. Limin Wang 0002, Gangshan Wu |
ICCV | 3 |
| 2021 | MultiSports: A Multi-Person Video Dataset of Spatio-Temporally Localized Sports ActionsabstractSpatio-temporal action detection is an important and challenging problem in video understanding. The existing action detection benchmarks are limited in aspects of small numbers of instances in a trimmed video or low-level atomic actions. This paper aims to present a new multi-person dataset of spatio-temporal localized sports actions, coined as MultiSports. We first analyze the important ingredients of constructing a realistic and challenging dataset for spatio-temporal action detection by proposing three criteria: (1) multi-person scenes and motion dependent identification, (2) with well-defined boundaries, (3) relatively fine-grained classes of high complexity. Based on these guide-lines, we build the dataset of MultiSports v1.0 by selecting 4 sports classes, collecting 3200 video clips, and annotating 37701 action instances with 902k bounding boxes. Our datasets are characterized with important properties of high diversity, dense annotation, and high quality. Our Multi-Sports, with its realistic setting and detailed annotations, exposes the intrinsic challenges of spatio-temporal action detection. To benchmark this, we adapt several baseline methods to our dataset and give an indepth analysis on the action detection results in our dataset. We hope our MultiSports can serve as a standard benchmark for spatio-temporal action detection in the future. Our dataset website is at https://deeperaction.github.io/multisports/. Yixuan Li 0002, Runyu He, Zhenzhi Wang 0001, Gangshan Wu, Limin Wang 0002 |
ICCV | 5 |
| 2021 | Relaxed Transformer Decoders for Direct Action Proposal GenerationabstractTemporal action proposal generation is an important and challenging task in video understanding, which aims at detecting all temporal segments containing action in-stances of interest. The existing proposal generation approaches are generally based on pre-defined anchor windows or heuristic bottom-up boundary matching strategies. This paper presents a simple and efficient framework (RTD-Net) for direct action proposal generation, by re-purposing a Transformer-alike architecture. To tackle the essential visual difference between time and space, we make three important improvements over the original transformer detection framework (DETR). First, to deal with slowness prior in videos, we replace the original Transformer en-coder with a boundary attentive module to better capture long-range temporal information. Second, due to the ambiguous temporal boundary and relatively sparse annotations, we present a relaxed matching scheme to relieve the strict criteria of single assignment to each groundtruth. Finally, we devise a three-branch head to further improve the proposal confidence estimation by explicitly predicting its completeness. Extensive experiments on THUMOS14 and ActivityNet-1.3 benchmarks demonstrate the effectiveness of RTD-Net, on both tasks of temporal action proposal generation and temporal action detection. Moreover, due to its simplicity in design, our framework is more efficient than previous proposal generation methods, without non-maximum suppression post-processing. The code and models are made available at https://github.com/MCG-NJU/RTD-Action. Jing Tan 0002, Jiaqi Tang 0001, Limin Wang 0002, Gangshan Wu |
ICCV | 4 |
| 2021 | Target Adaptive Context Aggregation for Video Scene Graph GenerationabstractThis paper deals with a challenging task of video scene graph generation (VidSGG), which could serve as a structured video representation for high-level understanding tasks. We present a new detect-to-track paradigm for this task by decoupling the context modeling for relation prediction from the complicated low-level entity tracking. Specifically, we design an efficient method for frame-level VidSGG, termed as Target Adaptive Context Aggregation Network (TRACE), with a focus on capturing spatio-temporal context information for relation recognition. Our TRACE framework streamlines the VidSGG pipeline with a modular design, and presents two unique blocks of Hierarchical Relation Tree (HRTree) construction and Target-adaptive Context Aggregation. More specific, our HRTree first provides an adpative structure for organizing possible relation candidates efficiently, and guides context aggregation module to effectively capture spatio-temporal structure information. Then, we obtain a contextualized feature representation for each relation candidate and build a classification head to recognize its relation category. Finally, we provide a simple temporal association strategy to track TRACE detected results to yield the video-level VidSGG. We perform experiments on two VidSGG benchmarks: ImageNet-VidVRD and Action Genome, and the results demonstrate that our TRACE achieves the state-of-the-art performance. The code and models are made available at https://github.com/MCG-NJU/TRACE. Yao Teng, Limin Wang 0002, Zhifeng Li 0001, Gangshan Wu |
ICCV | 4 |
| 2021 | MGSampler: An Explainable Sampling Strategy for Video Action RecognitionabstractFrame sampling is a fundamental problem in video action recognition due to the essential redundancy in time and limited computation resources. The existing sampling strategy often employs a fixed frame selection and lacks the flexibility to deal with complex variations in videos. In this paper, we present a simple, sparse, and explainable frame sampler, termed as Motion-Guided Sampler (MGSampler). Our basic motivation is that motion is an important and universal signal that can drive us to adaptively select frames from videos. Accordingly, we propose two important properties in our MGSampler design: motion sensitive and motion uniform. First, we present two different motion representations to enable us to efficiently distinguish the motion-salient frames from the background. Then, we devise a motion-uniform sampling strategy based on the cumulative motion distribution to ensure the sampled frames evenly cover all the important segments with high motion salience. Our MGSampler yields a new principled and holistic sampling scheme, that could be incorporated into any existing video architecture. Experiments on five benchmarks demonstrate the effectiveness of our MGSampler over the previous fixed sampling strategies, and its generalization power across different backbones, video models, and datasets. The code is available at https://github.com/MCG-NJU/MGSampler. Yuan Zhi, Zhan Tong, Limin Wang 0002, Gangshan Wu |
ICCV | 4 |
| 2021 | Spatial-Temporal Human-Object Interaction DetectionabstractIn this paper, we propose a new instance-level human-object interaction detection task on videos called ST-HOID, which aims to distinguish fine-grained human-object interactions (HOIs) and the trajectories of subjects and objects. It is motivated by the fact that HOI is crucial for human-centric video content understanding. To solve ST-HOID, we propose a novel method consisting of an object trajectory detection module and an interaction reasoning module. Furthermore, we construct the first dataset named VidOR-HOID for ST-HOID evaluation, which contains 10,831 spatial-temporal HOI instances. We conduct extensive experiments to evaluate the effectiveness of our method. The experimental results demonstrate that our method outperforms the baselines generated by the state-of-the-art methods of image human-object interaction detection, video visual relation detection and video human-object interaction recognition. Xu Sun 0009, Yunqing He, Tongwei Ren, Gangshan Wu |
ICME | 4 |
| 2021 | Reproducibility Companion Paper: Visual Relation of Interest DetectionabstractIn this companion paper, we provide the details of the reproducibility artifacts of the paper "Visual Relation of Interest Detection" presented at MM'20. Visual Relation of Interest Detection (VROID) aims to detect visual relations that are important for conveying the main content of an image. In this paper, we explain the file structure of the source code and publish the details of our ViROI dataset, which can be used to retrain the model with custom parameters. We also detail the scripts for component analysis and comparison with other methods and list the parameters that can be modified for custom training and inference. Fan Yu 0003, Tongwei Ren, Jinhui Tang 0001, Gangshan Wu, Jingjing Chen 0001, Zhenzhong Kuang |
ACM Multimedia | 5 |
| 2021 | Joint Learning for Relationship and Interaction Analysis in Video with Multimodal Feature FusionabstractTo comprehend long duration videos, the deep video understanding (DVU) task is proposed to recognize interactions on scene level and relationships on movie level and answer questions on these two levels. In this paper, we propose a solution to the DVU task which applies joint learning of interaction and relationship prediction and multimodal feature fusion. Our solution handles the DVU task with three joint learning sub-tasks: scene sentiment classification, scene interaction recognition and super-scene video relationship recognition, all of which utilize text features, visual features and audio features, and predict representations in semantic space. Since sentiment, interaction and relationship are related to each other, we train a unified framework with joint learning. Then, we answer questions for video analysis in DVU according to the results of the three sub-tasks. We conduct experiments on the HLVU dataset to evaluate the effectiveness of our method. Beibei Zhang 0005, Fan Yu 0003, Yanxin Gao, Tongwei Ren, Gangshan Wu |
ACM Multimedia | 5 |
| 2021 | Hybrid Improvements in Multimodal Analysis for Deep Video UnderstandingabstractThe Deep Video Understanding Challenge (DVU) is a task that focuses on comprehending long duration videos which involve many entities. Its main goal is to build relationship and interaction knowledge graph between entities to answer relevant questions. In this paper, we improved the joint learning method which we previously proposed in many aspects, including few shot learning, optical flow feature, entity recognition, and video description matching. We verified the effectiveness of these measures through experiments. Beibei Zhang 0005, Fan Yu 0003, Yaqun Fang, Tongwei Ren, Gangshan Wu |
MMAsia | 5 |
| 2021 | Cross-Modal Pyramid Translation for RGB-D Scene Recognition
Dapeng Du, Limin Wang 0002, Gangshan Wu |
Int. J. Comput. Vis. | 4 |
| 2021 | SADRNet: Self-Aligned Dual Face Regression Networks for Robust 3D Dense Face Alignment and ReconstructionabstractThree-dimensional face dense alignment and reconstruction in the wild is a challenging problem as partial facial information is commonly missing in occluded and large pose face images. Large head pose variations also increase the solution space and make the modeling more difficult. Our key idea is to model occlusion and pose to decompose this challenging task into several relatively more manageable subtasks. To this end, we propose an end-to-end framework, termed as Self-aligned Dual face Regression Network (SADRNet), which predicts a pose-dependent face, a pose-independent face. They are combined by an occlusion-aware self-alignment to generate the final 3D face. Extensive experiments on two popular benchmarks, AFLW2000-3D and Florence, demonstrate that the proposed method achieves significant superior performance over existing state-of-the-art methods. Zeyu Ruan, Changqing Zou, Longhai Wu, Gangshan Wu, Limin Wang 0002 |
IEEE Trans. Image Process. | 4 |
| 2020 | Residual Feature Aggregation Network for Image Super-ResolutionabstractRecently, very deep convolutional neural networks (CNNs) have shown great power in single image super-resolution (SISR) and achieved significant improvements against traditional methods. Among these CNN-based methods, the residual connections play a critical role in boosting the network performance. As the network depth grows, the residual features gradually focused on different aspects of the input image, which is very useful for reconstructing the spatial details. However, existing methods neglect to fully utilize the hierarchical features on the residual branches. To address this issue, we propose a novel residual feature aggregation (RFA) framework for more efficient feature extraction. The RFA framework groups several residual modules together and directly forwards the features on each local residual branch by adding skip connections. Therefore, the RFA framework is capable of aggregating these informative residual features to produce more representative features. To maximize the power of the RFA framework, we further propose an enhanced spatial attention (ESA) block to make the residual features to be more focused on critical spatial contents. The ESA block is designed to be lightweight and efficient. Our final RFANet is constructed by applying the proposed RFA framework with the ESA blocks. Comprehensive experiments demonstrate the necessity of our RFA framework and the superiority of our RFANet over state-of-the-art SISR methods. Jie Liu 0040, Wenjie Zhang 0006, Yuting Tang, Jie Tang 0006, Gangshan Wu |
CVPR | 5 |
| 2020 | Belief Map Enhancement Network for Accurate Human Pose Estimation
Jie Liu 0040, Yishun Dou, Wenjie Zhang 0006, Jie Tang 0006, Gangshan Wu |
ECAI | 5 |
| 2020 | Actions as Moving Points
Yixuan Li 0002, Limin Wang 0002, Gangshan Wu |
ECCV (16) | 4 |
| 2020 | Boundary-Aware Cascade Networks for Temporal Action Segmentation
Zhenzhi Wang 0001, Ziteng Gao, Limin Wang 0002, Zhifeng Li 0001, Gangshan Wu |
ECCV (25) | 5 |
| 2020 | Context-Aware RCNN: A Baseline for Action Detection in Videos
Jianchao Wu, Zhanghui Kuang, Limin Wang 0002, Wayne Zhang 0001, Gangshan Wu |
ECCV (25) | 5 |
| 2020 | Low Complexity Single Image Super-Resolution with Channel Splitting and Fusion NetworkabstractRecently, deep convolutional neural networks (CNNs) have made remarkable progress on single image super-resolution (SISR). However, many of these methods use very deep or wide convolutional layers to achieve good performance, which treat all feature channels indiscriminately and neglect the difference among the contribution of each channel to the output results. In this paper, we propose a low complexity solution based on channel splitting and fusion network (CSFN) to address this problem. Our method uses channel splitting and channel fusion to enhance feature maps and make full use of valuable information, and then multiple residual channel splitting and fusion blocks (CSFB) are cascaded to continuously extract more important information for reconstruction. To further minimize redundant parameters and improve efficiency, we adopt group and recursive con-volutional layer strategy in CSFB. Experiments demonstrate that our proposed CSFN could achieve higher performance with low computational complexity than most state-of-the-art methods. Minqiang Zou, Jie Tang 0006, Gangshan Wu |
ICASSP | 3 |
| 2020 | Human Object Interaction Detection via Multi-level Conditioned NetworkabstractAs one of the essential problems in scene understanding, human object interaction detection (HOID) aims to recognize fine-grained object-specific human actions, which demands the capabilities of both visual perception and reasoning. Existing methods based on convolutional neural network (CNN) utilize diverse visual features for HOID, which are insufficient for complex human object interaction understanding. To enhance the reasoning capablity of CNN, we propose a novel multi-level conditioned network that fuses extra spatial-semantic knowledge with visual features. Specifically, we construct a multi-branch CNN as backbone for multi-level visual representation. We then encode extra knowledge including human body structure and object context as condition to dynamically influence the feature extraction of CNN by affine transformation and attention mechanism. Finally, we fuse the modulated multimodal features to distinguish the interactions. The proposed method is evaluated on two most frequently-used benchmarks, HICO-DET and V-COCO. The experiment results show that our method is superior to the state-of-the-arts. Xu Sun 0009, Xinwen Hu, Tongwei Ren, Gangshan Wu |
ICMR | 4 |
| 2020 | Memory Recursive Network for Single Image Super-ResolutionabstractRecently, extensive works based on convolutional neural network (CNN) have shown great success in single image super-resolution (SISR). In order to improve the SISR performance while reducing the number of model parameters, some methods adopt multiple recursive layers to enhance the intermediate features. However, in the recursive process, these methods only use the output features of current stage as the input of the next stage and neglect the output features of historical stages, which degrades the performance of the recursive blocks. The long-term dependencies can only be learned implicitly during the recursive processes. To address these issues, we propose the memory recursive network (MRNet) to make full use of the output features at each stage. The proposed MRNet utilizes a memory recursive module (MRM) to generate features for each recursive stage, and then these features are fused by our proposed ShuffleConv block. Specifically, MRM adopts a memory updater block to explicitly model the long-term dependencies between the output features of historical recursive stages. The output features from the memory updater will be used as the input of the next recursive stage and will be continuously updated during the recursions. To reduce the number of parameters and ease the training difficulty, we introduce a ShuffleConv module to fuse the features from different recursive stages, which is much more effective than using plain convolutional combinations. Comprehensive experiments demonstrate that the proposed MRNet achieves state-of-the-art SISR performance while using much fewer parameters. Jie Liu 0040, Minqiang Zou, Jie Tang 0006, Gangshan Wu |
ACM Multimedia | 4 |
| 2020 | Visual Relation of Interest DetectionabstractIn this paper, we propose a novel Visual Relation of Interest Detection (VROID) task, which aims to detect visual relations that are important for conveying the main content of an image, motivated from the intuition that not all correctly detected relations are really "interesting" in semantics and only a fraction of them really make sense for representing the image main content. Such relations are named Visual Relations of Interest (VROIs). VROID can be deemed as an evolution over the traditional Visual Relation Detection (VRD) task that tries to discover all visual relations in an image. We construct a new dataset to facilitate research on this new task, named ViROI, which contains 30,120 images each with VROIs annotated. Furthermore, we develop an Interest Propagation Network (IPNet) to solve VROID. IPNet contains a Panoptic Object Detection (POD) module, a Pair Interest Prediction (PaIP) module and a Predicate Interest Prediction (PrIP) module. The POD module extracts instances from the input image and also generates corresponding instance features and union features. The PaIP module then predicts the interest score of each instance pair while the PrIP module predicts that of each predicate for each instance pair. Then the interest scores of instance pairs are combined with those of the corresponding predicates as the final interest scores. All VROI candidates are sorted by final interest scores and the highest ones are taken as final results. We conduct extensive experiments to test effectiveness of our method, and the results show that IPNet achieves the best performance compared with the baselines on visual relation detection, scene graph generation and image captioning. Fan Yu 0003, Tongwei Ren, Jinhui Tang 0001, Gangshan Wu |
ACM Multimedia | 5 |
| 2020 | Reproducibility Companion Paper: Instance of Interest DetectionabstractTo support the replication of "Instance of Interest Detection", which was presented at MM'19, this companion paper provides the details of the artifacts. Instance of Interest Detection (IOID) aims to provide instance-level user interest modeling for image semantic description. In this paper, we explain the file structure of the source code and publish the details of our IOID dataset, which can be used to retrain the model with custom parameters. We also provide a program for component analysis to help other researchers to do experiments with alternative models that are not included in our experiments. Moreover, we provide a demo program for using our model easily. Fan Yu 0003, Tongwei Ren, Jinhui Tang 0001, Gangshan Wu, Jingjing Chen 0001, Michael Riegler 0001 |
ACM Multimedia | 6 |
| 2020 | Steganographer Detection via Multi-Scale Embedding Probability EstimationabstractSteganographer detection aims to identify the guilty user who utilizes steganographic methods to hide secret information in the spread of multimedia data, especially image data, from a large amount of innocent users on social networks. A true embedding probability map illustrates the probability distribution of embedding secret information in the corresponding images by specific steganographic methods and settings, which has been successfully used as the guidance for content-adaptive steganographic and steganalytic methods. Unfortunately, in real-world situation, the detailed steganographic settings adopted by the guilty user cannot be known in advance. It thus becomes necessary to propose an automatic embedding probability estimation method. In this article, we propose a novel content-adaptive steganographer detection method via embedding probability estimation. The embedding probability estimation is first formulated as a learning-based saliency detection problem and the multi-scale estimated map is then integrated into the CNN to extract steganalytic features. Finally, the guilty user is detected via an efficient Gaussian vote method with the extracted steganalytic features. The experimental results prove that the proposed method is superior to the state-of-the-art methods in both spatial and frequency domains. Shenghua Zhong, Yuantian Wang, Tongwei Ren, Mingjie Zheng 0002, Yan Liu 0004, Gangshan Wu |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2019 | Translate-to-Recognize Networks for RGB-D Scene RecognitionabstractCross-modal transfer is helpful to enhance modality-specific discriminative power for scene recognition. To this end, this paper presents a unified framework to integrate the tasks of cross-modal translation and modality-specific recognition, termed as Translate-to-Recognize Network TRecgNet. Specifically, both translation and recognition tasks share the same encoder network, which allows to explicitly regularize the training of recognition task with the help of translation, and thus improve its final generalization ability. For translation task, we place a decoder module on top of the encoder network and it is optimized with a new layer-wise semantic loss, while for recognition task, we use a linear classifier based on the feature embedding from encoder and its training is guided by the standard cross-entropy loss. In addition, our TRecgNet allows to exploit large numbers of unlabeled RGB-D data to train the translation task and thus improve the representation power of encoder network. Empirically, we verify that this new semi-supervised setting is able to further enhance the performance of recognition network. We perform experiments on two RGB-D scene recognition benchmarks: NYU Depth v2 and SUN RGB-D, demonstrating that TRecgNet achieves superior performance to the existing state-of-the-art methods, especially for recognition solely based on a single modality. Dapeng Du, Limin Wang 0002, Gangshan Wu |
CVPR | 5 |
| 2019 | Learning Actor Relation Graphs for Group Activity RecognitionabstractModeling relation between actors is important for recognizing group activity in a multi-person scene. This paper aims at learning discriminative relation between actors efficiently using deep models. To this end, we propose to build a flexible and efficient Actor Relation Graph (ARG) to simultaneously capture the appearance and position relation between actors. Thanks to the Graph Convolutional Network, the connections in ARG could be automatically learned from group activity videos in an end-to-end manner, and the inference on ARG could be efficiently performed with standard matrix operations. Furthermore, in practice, we come up with two variants to sparsify ARG for more effective modeling in videos: spatially localized ARG and temporal randomized ARG. We perform extensive experiments on two standard group activity recognition datasets: the Volleyball dataset and the Collective Activity dataset, where state-of-the-art performance is achieved on both datasets. We also visualize the learned actor graphs and relation features, which demonstrate that the proposed ARG is able to capture the discriminative relation information for group activity recognition. Jianchao Wu, Limin Wang 0002, Jie Guo 0001, Gangshan Wu |
CVPR | 5 |
| 2019 | LIP: Local Importance-Based PoolingabstractSpatial downsampling layers are favored in convolutional neural networks (CNNs) to downscale feature maps for larger receptive fields and less memory consumption. However, for discriminative tasks, there is a possibility that these layers lose the discriminative details due to improper pooling strategies, which could hinder the learning process and eventually result in suboptimal models. In this paper, we present a unified framework over the existing downsampling layers (e.g., average pooling, max pooling, and strided convolution) from a local importance view. In this framework, we analyze the issues of these widely-used pooling layers and figure out the criteria for designing an effective downsampling layer. According to this analysis, we propose a conceptually simple, general, and effective pooling layer based on local importance modeling, termed as Local Importance-based Pooling (LIP). LIP can automatically enhance discriminative features during the downsampling procedure by learning adaptive importance weights based on inputs. Experiment results show that LIP consistently yields notable gains with different depths and different architectures on ImageNet classification. In the challenging MS COCO dataset, detectors with our LIP-ResNets as backbones obtain a consistent improvement (≥1.4%) over the vanilla ResNets, and especially achieve the current state-of-the-art performance in detecting small objects under the single-scale testing scheme1. Ziteng Gao, Limin Wang 0002, Gangshan Wu |
ICCV | 3 |
| 2019 | Dynamically Visual Disambiguation of Keyword-based Image SearchabstractDue to the high cost of manual annotation, learning directly from the web has attracted broad attention. One issue that limits their performance is the problem of visual polysemy. To address this issue, we present an adaptive multi-model framework that resolves polysemy by visual disambiguation. Compared to existing methods, the primary advantage of our approach lies in that our approach can adapt to the dynamic changes in the search results. Our proposed framework consists of two major steps: we first discover and dynamically select the text queries according to the image search results, then we employ the proposed saliency-guided deep multi-instance learning network to remove outliers and learn classification models for visual disambiguation. Extensive experiments demonstrate the superiority of our proposed approach. Yazhou Yao, Zeren Sun, Fumin Shen, Li Liu 0004, Limin Wang 0002, Fan Zhu 0001, Lizhong Ding 0001, Gangshan Wu, Ling Shao 0001 |
IJCAI | 8 |
| 2019 | Video Visual Relation Detection via Multi-modal Feature FusionabstractVideo visual relation detection is a meaningful research problem, which aims to build a bridge between dynamic vision and language. In this paper, we propose a novel video visual relation detection method with multi-model feature fusion. First, we detect objects on each frame densely with the state-of-the-art video object detection model, flow-guided feature aggregation (FGFA), and generate object trajectories by linking the temporally independent objects with Seq-NMS and KCF tracker. Next, we break the relation candidates, i.e., co-occurrent object trajectory pairs, into short-term segments and predict relations with spatial-temporal feature and language context feature. Finally, we greedily associate the short-term relation segments into complete relation instances. The experiment results show that our proposed method outperforms other methods by a large margin, which also earned us the first place in visual relation detection task of Video Relation Understanding Challenge (VRU), ACMMM 2019. Xu Sun 0009, Tongwei Ren, Yuan Zi, Gangshan Wu |
ACM Multimedia | 4 |
| 2019 | Hierarchical Visual Relationship DetectionabstractActing as a bridge between vision and language, visual relationship detection (VRD) aims to represent objects and their interactions in an image with several relationship triplets. Nevertheless, the conventional VRD task shows little consideration for the penalization of incorrect relationship predictions, which in turn undermines its support for image understanding applications. In this paper, we propose a novel VRD task named hierarchical visual relationship detection (HVRD), which encourages predictions with abstract yet compatible relationship triplets when the confidence level of the specific image content is relatively low. Meanwhile, HVRD can handle the inevitable ambiguity of groundtruth annotation in VRD. Based on this, we propose a HVRD method, consisting of hierarchical object detection and hierarchical predicate detection. It can effectively detect the hierarchical visual relationships by exploiting both object concept hierarchy and predicate concept hierarchy with order embedding. We also propose the first datasets for HVRD evaluation, H-VRD and H-VG, by expanding the relationship category spaces of VRD and VG datasets to hierarchical ones respectively. The experimental results show that our method is superior to the state-of-the-art baselines. Xu Sun 0009, Yuan Zi, Tongwei Ren, Jinhui Tang 0001, Gangshan Wu |
ACM Multimedia | 5 |
| 2019 | Crowd Counting via Multi-layer RegressionabstractCrowd counting aims to estimate the number of persons in a crowd image--a challenge until this day--as congestion degree varies, people's appearances may seem different. To address this problem, we propose a novel crowd counting method named Multi-layer Regression Network (MRNet), which consists of a multi-layer recognition branch and several density regressors. In practice, the recognition branch recognizes the congestion degree of the regions in a crowd image, then disintegrates the image into background and several crowd regions layer by layer, each regions are assigned different congestion degrees. In each layer, the recognized crowd regions with the specific congestion degree are delivered to a regressor with the corresponding density prior for crowd density estimation. The generated density maps at all layers are integrated to obtain the final density map for crowd density estimation. To date, MRNet is the first method to estimate crowd densities on crowd regions with different regressors. We conduct a comprehensive evaluation of MRNet on four typical datasets in comparison with nine state-of-the-art methods. By using multi-layer regression, MRNet achieves significant improvement in crowd counting accuracy, and outperforms the state-of-the-art methods. Chun Tao, Tongwei Ren, Jinhui Tang 0001, Gangshan Wu |
ACM Multimedia | 5 |
| 2019 | Instance of Interest DetectionabstractIn this paper, we propose a novel task named Instance of Interest Detection (IOID) to provide instance-level user interest modeling for image semantic description. IOID focuses on extracting the instances which are beneficial to represent image content, while other related tasks such as saliency analysis, attention model and instance segmentation extract the regions attracting visual attention or with a predefined category. To this end, we propose a Cross-influential Network for IOID, which integrates both visual saliency and semantic context. Moreover, we contribute the first dataset IOID evaluation, which consists of 45,000 images from MSCOCO with manually annotated instances of interest. Our method outperforms the state-of-the-art baselines on this dataset. Fan Yu 0003, Tongwei Ren, Jinhui Tang 0001, Gangshan Wu |
ACM Multimedia | 5 |
| 2019 | Personalized Recommendation of Photography Based on Deep Learning
Zhixiang Ji, Jie Tang 0006, Gangshan Wu |
MMM (1) | 3 |
| 2019 | Exploring overall opinions for document level sentiment classification with structural SVM
Xiaojia Pu, Gangshan Wu, Chunfeng Yuan |
Multim. Syst. | 2 |
| 2019 | User-aware topic modeling of online reviews
Xiaojia Pu, Gangshan Wu, Chunfeng Yuan |
Multim. Syst. | 2 |
| 2018 | A Parallel Method for All-Pair SimRank Similarity Computation
Xingkun Gao, Jie Tang 0006, Gangshan Wu |
ICA3PP (1) | 4 |
| 2018 | Depth Images Could Tell us More: Enhancing Depth Discriminability for RGB-D Scene RecognitionabstractRecently depth-modal information has been witnessed effectively in computer vision community, especially for scene analysis related tasks. However, it still suffers severely from depth data scarcity as well as improperly transferring pre-trained RGB models to fit depth-modal data. In this study, we propose a novel two-step training strategy to address these problems and focus on enhancing the recognition power for depth-modal images in RGB-D scene recognition task. Specifically, we build an effective “Res-U” architecture on a GAN (generative adversarial networks) based RGB-to-depth modality translation model, which is endowed with both short and long skips for residual learning. On one hand, this could first well pre-train a depth-modal-specific discriminator network from scratch in an unsupervised manner, which is effectively transformed for the subsequent recognition task instead of directly fitting pre-trained RGB model to depth-specific one. On the other hand, new depth images with helpful perturbations, generated from the modality translation model, help argument the original training set and regularize the learning process in some sense. This two-step training strategy makes it more effective for training a modal-specific network to discriminate depth scenes. Besides, we extensively explore the modality translation network to investigate the effects in recognizing depth-modal scenes, which encourages a reasonable way to take full advantage of multi-modalities. The proposed method achieves state-of-the-art accuracy on NYU Depth v2 and SUN RGB-D benchmark datasets, especially on depth data only evaluation. Dapeng Du, Tongwei Ren, Gangshan Wu |
ICME | 4 |
| 2018 | Object Trajectory Proposal via Hierarchical Volume GroupingabstractObject trajectory proposal aims to locate category-independent object candidates in videos with a limited number of trajectories,i.e.,bounding box sequences. Most existing methods, which derive from combining object proposal with tracking, cannot handle object trajectory proposal effectively due to the lack of comprehensive objectness measurement through analyzing spatio-temporal characteristics over a whole video. In this paper, we propose a novel object trajectory proposal method using hierarchical volume grouping. Specifically, we first represent a given video with hierarchical volumes by mapping hierarchical regions with optical flow. Then, we filter the short volumes and background volumes, and combinatorially group the retained volumes into object candidates. Finally, we rank the object candidates using a multi-modal fusion scoring mechanism, which incorporates both appearance objectness and motion objectness, and generate the bounding boxes of the object candidates with the highest scores as the trajectory proposals. We validated the proposed method on a dataset consisting of 200 videos from ILSVRC2016-VID. The experimental results show that our method is superior to the state-of-the-art object trajectory proposal methods. Xu Sun 0009, Yuantian Wang, Tongwei Ren, Zhi Liu 0003, Zhengjun Zha, Gangshan Wu |
ICMR | 6 |
| 2018 | HeterStyle: A Heterogeneous Video Style Transfer ApplicationabstractVideo style transfer aims to synthesize a stylized video that preserves the content of a given video and is rendered in the style of a reference image.A key issue in video style transfer is how to balance video content preservation and reference style rendering, in order to avoid over-stylization with serious video content loss or under-stylization with unrecognized reference style. In this demonstration, we illustrate a novel video style transfer application, named HeterStyle, which can stylize different regions in the video with adaptive intensities.The core algorithm of HeterStyle application is our proposed heterogeneous video style transfer method, which minimizes a heterogeneous style transfer loss function considering content, style and temporal consistency in a Convolutional Neural Networks based optimization framework.With the HeterStyle application, a user can easily generate the stylized videos with good video content preservation and reference style rendering. Jingfan Guo, Tongwei Ren, Yahong Han, Lei Huang 0004, Gangshan Wu |
ACM Multimedia | 6 |
| 2018 | Robust and Real-Time Visual Tracking Based on Complementary Learners
Xingzhou Luo, Dapeng Du, Gangshan Wu |
MMM (2) | 3 |
| 2018 | Adaptive video object proposals by a context-aware model
Wenjing Geng, Chunlong Zhang, Gangshan Wu |
Multim. Tools Appl. | 3 |
| 2018 | Adaptive saliency cuts
Yuantian Wang, Tongwei Ren, Shenghua Zhong, Yan Liu 0004, Gangshan Wu |
Multim. Tools Appl. | 5 |
| 2017 | Video salient object detection via cross-frame cellular automataabstractSalient object detection aims to detect the attractive objects on images and videos. In this paper, we propose a novel salient object detection method for videos based on cross-frame cellular automata. Given a video, we first represent the video frames with super-pixels, and construct a saliency propagation network among super-pixels within a frame and between adjacent frames based on their appearance similarities and temporal coherency. Second, we initialize the saliency map of each frame with the fusion of two saliency maps generated by appearance and motion features independently. Finally, we utilize cellular automata updating to propagate saliency among super-pixels iteratively and generate the coherent saliency maps with complete objects. The experimental results show that our method outperforms the state-of-the-art methods on different types of videos. Jingfan Guo, Tongwei Ren, Lei Huang 0004, Ming-Ming Cheng, Gangshan Wu |
ICME | 6 |
| 2017 | Object trajectory proposalabstractWe propose a novel method for video object proposal to generate sequences of bounding boxes for each object candidate in videos, namely object trajectory proposals. Unlike the image-based methods that produce object proposals independently in each video frame, our method generates temporally consistent proposals in the form of object trajectories, which is crucial for subsequent analysis of object appearance and motion characteristics. Given a video, we extract motion seeds through estimating the outliers of global motion, and generate moving object trajectory proposals from the seeds. By ignoring the motion outliers, we consistently sample bounding boxes with pruning to form static object trajectory proposals. Finally, we rank both the moving and static object trajectory proposals under a unified scoring mechanism. The experimental results show that our method can effectively generate object trajectory proposals and outperform the state-of-the-art methods. Xindi Shang, Tongwei Ren, Hanwang Zhang, Gangshan Wu, Tat-Seng Chua |
ICME | 4 |
| 2017 | Deep convolutional neural networks for pedestrian detection with skip poolingabstractWith the big success of deep convolutional neural networks (CNN) in image classification task, many proposal based networks are proposed to detect given objects in an image. Faster R-CNN is such a network that uses a region proposal network (RPN) to generate nearly cost-free region proposals, which has shown excellent performance in ILSVRC and MS COCO datasets. However, Faster R-CNN does not behave so well for the task of pedestrian detection since the images in popular pedestrian detection datasets have more complicated background and contain a lot of small foreground objects. In this work, we leverage the RPN architecture of Faster R-CNN and extend it to a multi-layer version combined with skip pooling to tackle the pedestrian detection problem. Skip pooling is a kind of network connection that combines multiple ROI pooling results from lower layers to form a single input to a higher layer while bypassing intermediate layers. We comprehensively evaluate our network, referred to as SP-CNN, on the Caltech pedestrian detection benchmark and KITTI object detection benchmark. Our method achieves state-of-the-art accuracy on Caltech dataset and presents a comparable result on KITTI dataset while maintaining a good speed. Jie Liu 0040, Xingkun Gao, Nianyuan Bao, Jie Tang 0006, Gangshan Wu |
IJCNN | 5 |
| 2017 | Sentiment analysis with the exploration of overall opinion sentencesabstractWith the rapid growth of opinionated contents, e.g. product reviews, sentiment analysis has drawn much attention from the researchers. The most fundamental task of sentiment analysis is document sentiment classification which aims to predict the overall sentiment (e.g. positive or negative) towards the opinion target in a review. There are usually various opinion sentences towards different aspects with different sentiments. Among them, the overall opinion towards the whole target should be more deterministic in document sentiment prediction. However, most existing methods treat all the sentences equally, thus, they may encounter difficulty especially when the sentiments of most aspect opinion sentences differ from the overall sentiment. To address this, we propose a novel method for document sentiment classification which adequately explores the effect of overall opinion sentences. The method is extended from structural SVM, and the overall opinion sentences are taken as the hidden variables for document sentiment. Experiments on several standard product review datasets show the effectiveness of our method. Xiaojia Pu, Gangshan Wu, Chunfeng Yuan |
IJCNN | 2 |
| 2017 | Multi-modal deep feature learning for RGB-D object detection
Yuncheng Li, Gangshan Wu, Jiebo Luo 0001 |
Pattern Recognit. | 3 |
| 2016 | User-oriented stereo video refocusing by computational cinematographic modelabstractRefocusing, as a most popular photographic technique, is widely welcomed in both photography and cinematography. Currently, most refocusing effects in movies are implemented manually using professional devices or computer graphics methods. Few solutions are designed for the videos captured by common users. In this paper, we propose a user-oriented method to facilitate the video refocusing for daily life, with an extra requirement for only stereo cameras which are quite popular today. Given a stereo video, one can select the desired focusing part in a single frame, and the user selected intention will be tracked along with the timeline. Then the depth-of-field (DoF) based blurring effect can be generated by scene depth estimation based on the stereo techniques and the proposed cinematographic models in computational photography. The performance of the refocused videos compares favourably with the ones generated by digital single lens reflex (DSLR). To evaluate the proposed method, we conduct the experiments on several challenging stereo video datasets. The quantitative and subjective evaluations both show that the proposed method can achieve attractive and highly aesthetic performance with few user interactions. Wenjing Geng, Dapeng Du, Tongwei Ren, Gangshan Wu |
ICME | 4 |
| 2016 | Scalable Single-Source SimRank Computation for Large GraphsabstractSimRank is an effective similarity measure between vertices in a graph, which has become a fundamental technique in graph analytics. Despite its popularity, computation of SimRank is often costly in both space and time, especially with the ever growing scale of graph data nowadays. In this paper, we focus on the computation of Single-Source SimRank: given a query vertex, return the similarities between this vertex and any other vertices in the graph. The traditional centralized SimRank algorithms are not efficient for this problem. To fully utilize the computing power of modern distributed systems, we propose sssSimRank, an efficient distributed algorithm based on the random walk model. Our algorithm achieves scalability via minimizing the total number, the space cost, and the matching time of random walks. We implement our approach on the popular distributed processing platform Spark. Experimental results demonstrate the effectiveness, efficiency and scalability of our method. Xingkun Gao, Nianyuan Bao, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
ICPADS | 5 |
| 2016 | Context-Aware Video Object ProposalsabstractRecent advances in object proposals have been achieved obvious performance to speed up sliding window based object detection or recognition. However, the spatial-temporal object proposal of multi-objects in video is still a challenging problem. Applying the existing image methods frame by frame will result in three defects. First, no guarantee to keep the consistent proposal results, i.e., it is hard to avoid omitting objects even in consecutive or similar sequences. Second, the latent information contained in time dimension would not be made best use of to improve the detection rate. Third, due to the motion blur caused by motion flow, the efficiency of object proposals relying on contour or edge features would be definitely degraded. In this paper, we propose an efficient method for video object proposals. By introducing image method into context-aware framework, we get the improved detection rate compared to the frame by frame usage, while keeping a controllable computing efficiency. Firstly, the bounding boxes produced by image proposals are used as the input. Then the candidate windows are scored with contextual information by generating motion-based mapping boxes. To evaluate the multi-object proposal results, we build a specific dataset. Experiments show that the proposed method can improve the detection rate of the original image method, and especially achieve better performance when proposing a small set of bounding boxes. Wenjing Geng, Gangshan Wu |
ICPADS | 2 |
| 2016 | Object proposals using SVM-based integrated modelabstractUtilizing object proposals as a preprocessing procedure has been shown its significance in many multimedia computing tasks. Most state-of-the-art methods devoted to finding a generic objectness measure for rating the possibilities of the initial sliding windows with or without objects. In fact, the object criteria vary from one objectness measure to another, which leads to the definite bottleneck for the single method. By observing the performance of the state-of-the-art in the large dataset, an integrated objectness model is proposed in this paper by accumulating the advantages from selected state-of-the-art techniques. First, the initial bounding boxes are generated by the strategy as same as the method with the highest object detection rate and slowest intersection over union drop. Second, these candidate boxes are re-scored based on each method's objectness system. Then, a score feature is obtained for each bounding box. A support vector machines (SVM) is utilized to train a general model on the training set constructed from a series of score vectors and the probabilistic scores for the testing boxes are predicted according to the learned model. The final proposals are ranked on account of the predicted scores. The evaluation on the challenging PASCAL VOC 2007 dataset shows that the proposed method has dominant concentration with better performance compared to the single state-of-the-art method. Wenjing Geng, Shuzhen Li, Tongwei Ren, Gangshan Wu |
IJCNN | 4 |
| 2016 | How important is location information in saliency detection of natural images
Tongwei Ren, Yan Liu 0004, Ran Ju, Gangshan Wu |
Multim. Tools Appl. | 4 |
| 2015 | Topic Modeling in Semantic Space with KeywordsabstractA common and convenient approach for user to describe his information need is to provide a set of keywords. Therefore, the technique to understand the need becomes crucial. In this paper, for the information need about a topic or category, we propose a novel method called TDCS(Topic Distilling with Compressive Sensing) for explicit and accurate modeling the topic implied by several keywords. The task is transformed as a topic reconstruction problem in the semantic space with a reasonable intuition that the topic is sparse in the semantic space. The latent semantic space could be mined from documents via unsupervised methods, e.g. LSI. Compressive sensing is leveraged to obtain a sparse representation from only a few keywords. In order to make the distilled topic more robust, an iterative learning approach is adopted. The experiment results show the effectiveness of our method. Moreover, with only a few semantic concepts remained for the topic, our method is efficient for subsequent text mining tasks. Xiaojia Pu, Rong Jin 0001, Gangshan Wu, Dingyi Han, Gui-Rong Xue |
CIKM | 3 |
| 2015 | A New Data Replication Scheme for PVFS2
Nianyuan Bao, Jie Tang 0006, Gangshan Wu |
ICA3PP (3) | 4 |
| 2015 | Pre-stack Kirchhoff Time Migration on Hadoop and Spark
Jie Tang 0006, Gangshan Wu |
ICA3PP (3) | 4 |
| 2015 | A Dynamic Extension and Data Migration Method Based on PVFS
Jie Tang 0006, Gangshan Wu |
ICA3PP (2) | 4 |
| 2015 | StereoSnakes: Contour Based Consistent Object Extraction for Stereo ImagesabstractConsistent object extraction plays an essential role for stereo image editing with the population of stereoscopic 3D media. Most previous methods perform segmentation on entire images for both views using dense stereo correspondence constraints. We find that for such kind of methods the computation is highly redundant since the two views are near-duplicate. Besides, the consistency may be violated due to the imperfectness of current stereo matching algorithms. In this paper, we propose a contour based method which searches for consistent object contours instead of regions. It integrates both stereo correspondence and object boundary constraints into an energy minimization framework. The proposed method has several advantages compared to previous works. First, the searching space is restricted in object boundaries thus the efficiency significantly improved. Second, the discriminative power of object contours results in a more consistent segmentation. Furthermore, the proposed method can effortlessly extend existing single-image segmentation methods to work in stereo scenarios. The experiment on the Adobe benchmark shows superior extraction accuracy and significant improvement of efficiency of our method to state-of-the-art. We also demonstrate in a few applications how our method can be used as a basic tool for stereo image editing. Ran Ju, Tongwei Ren, Gangshan Wu |
ICCV | 3 |
| 2015 | Saliency cuts based on adaptive triple thresholdingabstractSalient object detection attracts much attention for its effectiveness in numerous applications. However, how to effectively produce a high quality binary mask from a saliency map, named saliency cuts, is still an open problem. In this paper, we propose a novel saliency cuts approach using unsupervised seeds generation and GrabCut algorithm. With the input of a saliency map, we produce seeds for segmentation using adaptive triple thresholding, and feed the seeds to GrabCut algorithm. Finally, a high quality object mask is generated by iteratively optimization. The experimental results show that the proposed approach is competent to the task of saliency cuts and outperforms the state-of-the-art methods. Shuzhen Li, Ran Ju, Tongwei Ren, Gangshan Wu |
ICIP | 4 |
| 2015 | Adaptive integration of depth and color for objectness estimationabstractThe goal of objectness estimation is to predict a moderate number of proposals of all possible objects in a given image with high efficiency. Most existing works solve this problem solely in conventional 2D color images. In this paper, we demonstrate that the depth information could benefit the estimation as a complementary cue to color information. After detailed analysis of depth characteristics, we present an adaptively integrated description for generic objects, which could take full advantages of both depth and color. With the proposed objectness description, the ambiguous area, especially the highly textured regions in original color maps, can be effectively discriminated. Meanwhile, the object boundary areas could be further emphasized, which leads to a more powerful objectness description. To evaluate the performance of the proposed approach, we conduct the experiments on two challenging datasets. The experimental results show that our proposed objectness description is more powerful and effective than state-of-the-art alternatives. Ling Ge, Tongwei Ren, Gangshan Wu |
ICME | 4 |
| 2015 | EventBuilder: Real-time Multimedia Event Summarization by Visualizing Social MediaabstractDue to the ubiquitous availability of smartphones and digital cameras, the number of photos/videos online has increased rapidly. Therefore, it is challenging to efficiently browse multimedia content and obtain a summary of an event from a large collection of photos/videos aggregated in social media sharing platforms such as Flickr and Instagram. To this end, this paper presents the EventBuilder system that enables people to automatically generate a summary for a given event in real-time by visualizing different social media such as Wikipedia and Flickr. EventBuilder has two novel characteristics: (i) leveraging Wikipedia as event background knowledge to obtain more contextual information about an input event, and (ii) visualizing an interesting event in real-time with a diverse set of social media activities. According to our initial experiments on the YFCC100M dataset from Flickr, the proposed algorithm efficiently summarizes knowledge structures based on the metadata of photos/videos and Wikipedia articles. Rajiv Ratn Shah, Anwar Dilawar Shaikh, Yi Yu 0001, Wenjing Geng, Roger Zimmermann, Gangshan Wu |
ACM Multimedia | 6 |
| 2015 | Flat3D: Browsing Stereo Images on a Conventional Screen
Wenjing Geng, Ran Ju, Tongwei Ren, Gangshan Wu |
MMM (1) | 5 |
| 2015 | Optimal Point Movement for Covering Circular Regions
Danny Ziyi Chen, Xuehou Tan, Haitao Wang 0001, Gangshan Wu |
Algorithmica | 4 |
| 2015 | Depth-aware salient object detection using anisotropic center-surround difference
Ran Ju, Yang Liu 0007, Tongwei Ren, Ling Ge, Gangshan Wu |
Signal Process. Image Commun. | 5 |
| 2014 | Depth saliency based on anisotropic center-surround differenceabstractMost previous works on saliency detection are dedicated to 2D images. Recently it has been shown that 3D visual information supplies a powerful cue for saliency analysis. In this paper, we propose a novel saliency method that works on depth images based on anisotropic center-surround difference. Instead of depending on absolute depth, we measure the saliency of a point by how much it outstands from surroundings, which takes the global depth structure into consideration. Besides, two common priors based on depth and location are used for refinement. The proposed method works within a complexity of O(N) and the evaluation on a dataset of over 1000 stereo images shows that our method outperforms state-of-the-art. Ran Ju, Ling Ge, Wenjing Geng, Tongwei Ren, Gangshan Wu |
ICIP | 5 |
| 2014 | OBSIR: Object-based stereo image retrievalabstractRecent years, the stereo image has become an emerging media in the field of 3D technology, which leads to an urgent demand of stereo image retrieval. In this paper, we attempt to introduce a framework for object-based stereo image retrieval (OBSIR), which retrieves images containing the similar objects to the one captured in the query image by the user. The proposed approach consists of both online and offline procedures. In the offline procedure, we propose a salient object segmentation method making use of both color and depth to extract objects from each image. The extracted objects are then represented by multiple visual feature descriptors. In order to improve the image search efficiently, we construct an approximate nearest neighbor (ANN) index using cluster-based locality sensitive hashing (LSH). In the online stage, the user may supply the query object by selecting a region of interest (ROI) in the query image, or clicking one of the objects recommended by the salient object detector. For the image retrieval evaluation we build a new dataset containing over 10K stereo images. The experiments on this dataset show that the proposed method can effectively recommend the correct object and the final retrieval result is also better than other baseline methods. Wenjing Geng, Ran Ju, Yang Yang 0222, Tongwei Ren, Gangshan Wu |
ICME | 6 |
| 2014 | Image annotation by modeling Supporting Region Graph
Qiaojin Guo, Ning Li 0013, Gangshan Wu |
Appl. Intell. | 4 |
| 2014 | Image Relevance Prediction Using Query-Context Bag-of-Object Retrieval ModelabstractImage search reranking and image research result summarization are two effective approaches which enhance text-based image search results using visual information. Since the existing approaches optimize search relevance in terms of average performance, they usually cannot achieve satisfactory results for some particular classes of queries, like “object queries,” which is defined as the queries with the intent of searching for some kinds of objects. One possible reason is that the generic approaches such as , , are mostly built based on the global statistics of images as features while ignoring the fact that the relevance between the image and the query sometimes depends on an image patch instead of the whole image. In this paper, we therefore design a novel bag-of-object retrieval model to predict image relevance, which is particularly effective for object queries. First, we construct an object vocabulary containing query-relative objects by mining frequent object patches from the result image collection of the expanded query set. After representing each image as a bag of objects, our retrieval model can be derived from a risk-minimization framework for language modeling. To demonstrate the effectiveness of the proposed model, this paper also present two related applications: for image search reranking, we adopt a supervised framework to combine multiple ranking features from different assumptions; for image search result summarization, we propose a two-step ranking process which optimizes not only representativeness but also image attractiveness. The experimental results show that the proposed methods can significantly outperform the existing approaches. Yang Yang 0222, Linjun Yang, Gangshan Wu, Shipeng Li 0001 |
IEEE Trans. Multim. | 3 |
| 2013 | Integrating image segmentation and annotation using supervised PLSAabstractIn this work, we propose an integrated framework for image segmentation and annotation based on topic models. We first employ probabilistic latent semantic analysis (PLSA) to discover the topics of image regions and pixels, both utilized in our framework to segment images into different regions and to improve the annotation accuracy. Furthermore, we propose a supervised version of PLSA, SPLSA, as a new graphical model in order to accommodate the annotation results into the segmentation process to improve its performance. We compare the proposed SPLSA model with PLSA on their image segmentation performance, and also evaluate the image annotation performance of our methods on different datasets. The experimental results prove that the supervised information in SPLSA improves the segmentation results, and the image annotation accuracy is higher than the state-of-art methods including Conditional Random Fields. Qiaojin Guo, Ning Li 0013, Gangshan Wu |
ICIP | 4 |
| 2013 | Effective local stereo matching by extended triangular interpolationabstractIn this paper, we propose an effective local stereo matching method based on extended triangular interpolation. The whole image is covered with a triangular mesh by performing triangulation on a set of initial support points. Since the disparity interpolation in some areas is ineffective, especially in the triangle which appears across the boundaries of objects or is formed by initial matching outliers, we formulate a new matching model based on the Bayesian rule to address this challenge. In the model, we introduce the concept of the triangle's reliability, and utilize the disparity planes determined by the neighboring triangles as the Bayesian prior. With the model, the disparity interpolation can be effectively performed. Experiments on the standard stereo data sets show that the method is effective, especially in dealing with the boundaries of objects. Chunrong Xia, Yang Yang 0222, Ran Ju, Gangshan Wu |
ICME | 4 |
| 2013 | Approximation algorithms for cutting a convex polyhedron out of a sphere
Xuehou Tan, Gangshan Wu |
Theor. Comput. Sci. | 2 |
| 2012 | Optimal Point Movement for Covering Circular Regions
Danny Ziyi Chen, Xuehou Tan, Haitao Wang 0001, Gangshan Wu |
ISAAC | 4 |
| 2012 | Semiconducting bilinear deep learning for incomplete image recognitionabstractImage recognition with incomplete data is a well-known hard problem in multimedia content analysis. This paper proposes a novel deep learning technique called semiconducting bilinear deep belief networks (SBDBN) by referencing human's visual cortex and intelligent perception. Inheriting from deep models, SBDBN simulates the laminar structure of human's cerebral cortex and the neural loop in human's visual areas. To address the special difficulties of image recognition with incomplete data, we design a novel second-order deep architecture with semiconducting restricted boltzmann machines. Moreover, two peaks activation of human's perception is implemented by three learning stages of semiconducting bilinear discriminant initialization, greedy layer-wise reconstruction, and global fine-tuning. Owing to exploiting the embedding information according to the reliable features rather than any completion of missing features, the proposed SBDBN has demonstrated outstanding recognition ability on two standard datasets and one constructed dataset, comparing with both incomplete image recognition techniques and existing deep learning models. Shenghua Zhong, Yan Liu 0004, Korris Fu-Lai Chung, Gangshan Wu |
ICMR | 4 |
| 2012 | A bag-of-objects retrieval model for web image searchabstractImage search reranking has been an active research topic in recent years to boost the performance of the existing web image search engine which is mostly based on textual metadata of images. Various approaches have been proposed to rerank images for general queries and argue that, they may not necessarily be optimal for queries in specific domain, e.g., object queries, since the reranking algorithms are operated on whole images, instead of the relevant parts of images. In this paper, we propose a novel bag-of-objects retrieval model for image search reranking of object queries. Firstly, we employ a common object discovery algorithm to discover query-relevant objects from the search results returned by text-based image search engine. Then, the query and its result images are represented as a language model on the query relevant object vocabulary, based on which the ranking function can be derived. As the common object discovery is unreliable and may introduce noises, we propose to incorporate the attributes of the discovered objects, e.g., size, position, etc., into the ranking function through a linear model, and the weights on the object attributes can be learned. The experiments on two subsets of Web Queries dataset comprising object queries demonstrate that our approach can significantly outperform the existing reranking methods on object queries. Yang Yang 0222, Linjun Yang, Gangshan Wu, Shipeng Li 0001 |
ACM Multimedia | 3 |
| 2012 | S-SIFT: A Shorter SIFT without Least Discriminability Visual OrientationabstractDetection and description of local features are a classical problem in image processing and multimedia content analysis. Based on the in homogeneity of visual orientation in human visual system, we propose a novel algorithm S-SIFT to detect and describe local image features. In three stages of S-SIFT, the information from the least discriminability orientation is omitting. Compared with the standard SIFT algorithm, S-SIFT has lower dimension and provides a faster key point matching. Experiments on the standard dataset demonstrate that our algorithm yields comparable or even better results for feature detection and matching tasks. Shenghua Zhong, Yan Liu 0004, Gangshan Wu |
Web Intelligence | 3 |
| 2011 | Image Annotation with Multiple QuantizationabstractImage annotation plays an important role in image retrieval and understanding. Various techniques have been proposed for assigning keywords to images. One of the most frequently used methods is to search annotated images with similar visual features, and keywords are transfered to new coming images. This leads to the problem of nearest neighbor search, which is a hot topic of pattern recognition, information retrieval, and data compression. In this paper we proposed a fast and effective method for retrieving similar images from large collections of annotated images. The proposed technique employs discrete cosine transform and regular lattice quantization to encode images and search similar images directly with the corresponding codes. This technique is evaluated on image annotation. Similar images are retrieved by utilizing our encoding strategy, and keywords are assigned by utilizing traditional label transfer mechanism. Experimental results show that our method provides competitive performance with traditional methods, and mean while provides one scalable framework for annotating large collections of image dataset. Qiaojin Guo, Ning Li 0013, Gangshan Wu |
ICIG | 4 |
| 2011 | A Robust and Compact Descriptor Based on Center-Symmetric LBPabstractCenter-symmetric local binary pattern (CS-LBP) is a novel texture feature which utilizes texture to describe the local regions. It combines the good property of local binary pattern (LBP) and SIFT. It has been extended to a region descriptor and achieved promising performance in many applications. However, it is sensitive to noise and less efficient due to its high dimensional descriptor vector. Due to these, we propose a novel descriptor based on CS-LBP operator denoted as PCA-CS-LBP. Our proposed descriptor achieves better noise robustness using the difference of pixels instead of the rough comparing of pixels. Besides, PCA is employed and applied to generate a more compact representation. Comparisons between our descriptor and standard CS-LBP descriptor are given on a standard image matching dataset. Experimental results show that our descriptor is outperforms the standard CS-LBP descriptor in most cases. Jinwei Xiao, Gangshan Wu |
ICIG | 2 |
| 2011 | Supervised LDA for Image AnnotationabstractRegion-based Image Annotation has received increasing attention in recent years. Topic models such as probabilistic Latent Semantic Analysis (PLSA) and Latent Dirichlet Allocation (LDA) have shown great success in object recognition and localization. In this paper, we introduce a supervised topic model for region-based image annotation. Images are segmented into superpixels, and visual features are extracted from each superpixel region. Boosted classifiers are then trained for each class, and the output of boosted classifiers are quantized as boosted visual words. The proposed model builds a generative model on both visual words and corresponding class labels. We tested the model on the 21-class MSRC dataset. Experimental results show that our model improves the annotation performance comparing with boosted classifiers. Qiaojin Guo, Ning Li 0013, Gangshan Wu |
SMC | 4 |
| 2010 | Rapid image retargeting based on curve-edge grid representationabstractImage retargeting technique attracts more and more attention for convenient image display on mobile devices. However, current methods can't well balance the retargeting efficiency and effectiveness, which limits their applications on the mobile devices with low computing ability. In this paper, we propose a novel image retargeting approach by combining uniform sampling and structure-aware curve-edge grid representation. We first decompose the original image into curve-edge grids by dynamic programming, and then generate the target image by uniformly sampling the pixels within the grids. The simplicity of retargeting procedure and sampling strategy enables our approach to easily achieve good computational efficiency. Furthermore, the constraint of curve-edge grid representation ensures important content emphasis and image structure preservation in the target image. Experiments on different images demonstrate the effectiveness and efficiency of our approach. Tongwei Ren, Yan Liu 0004, Gangshan Wu |
ICIP | 3 |
| 2010 | Automatic image retargeting evaluation based on user perceptionabstractAs image retargeting techniques have attracted more and more attention for effective image display on mobile devices, quality evaluation of image retargeting is required. To address the lack of automatic evaluation techniques in retargeting, this paper proposes a user perception based framework to automatically evaluate the quality of the target image against the original image. In the framework, the pixels in the original image and the target image are first approximately order-preserved matched by dynamic programming. Based on the pixel matching result, several features are extracted to describe the user requirements and further adapted to fit user perception in retargeting. Finally, the overall score of the target image quality is calculated by integrating the scores in different evaluation aspects. Experiments demonstrate the effectiveness of the proposed framework. Tongwei Ren, Gangshan Wu |
ICIP | 2 |
| 2010 | Interective Point Clouds Fairing on Many-Core SystemabstractThis Paper proposes an interactive point clouds fairing algorithm running on many-core system. The algorithm is composed of four steps. Firstly, a k nearest neighbor searching method was designed which could fully utilize the computing ability of GPU. Secondly, a parallel Gaussian weighted normal estimation was put forward. Thirdly, a weighted fairing method was proposed to get better result especially for the unevenly distributed point clouds. The whole algorithm was implemented on NVIDIA GPU using CUDA. Experimental results show that the algorithm could achieve interactive fairing of large size point clouds with good quality. Jie Tang 0006, Gangshan Wu, Zhongliang Gong |
ISPA | 2 |
| 2009 | A Motion-Insensitive Dissolve Detection Method with SURFabstractAs dissolve is the most common gradual shot transition, dissolve detection plays an important role in video segmentation which is the fundamental step for efficient video indexing and retrieval. However, the existing detection methods easily confuse dissolve with camera motion or object motion when using global features. Besides, when using local features' change tendency, they can' t get accurate trajectories to reflect this. In this paper, we propose a SURF feature based dissolve detection algorithm which can well differentiate dissolve from motion appearances. We get the trajectories of SURF key points by matching between two successive frames. Then a candidate set of dissolves is obtained according to the distribution of the starting points and ending points on the trajectory, filtering part of motions. Dissolves are further located by analyzing the curve of the proportion of subtrajectories with monotonous variation through each frame. Experiments demonstrate the effectiveness and efficiency of the proposed method. Yaqiong Wang, Yang Yang 0222, Tongwei Ren, Gangshan Wu |
ICIG | 4 |
| 2009 | Image retargeting based on global energy optimizationabstractThis paper proposes a novel image retargeting technique based on global energy optimization. Most existing methods enhance the high energy parts of the original image by pre-defined strategies or local optimization based iterations. They can not achieve the global optimal effect in energy retainment. To solve this problem, our approach formulates image retargeting as a global optimization problem on energy. We first calculate the energy map of the original image. Then, we utilize a constrained linear programming to maximize the retained energy in retargeting. Finally, we propose a pixel fusion based method to generate the retargeted image. To make it more feasible in implementation, we further provide two strategies to reduce the time cost of our approach. We demonstrate the proposed approach by comparing with typical image retargeting methods. Tongwei Ren, Yan Liu 0004, Gangshan Wu |
ICME | 3 |
| 2009 | Image retargeting using multi-map constrained region warpingabstractImage retargeting aims to adapt images to various screens with small sizes and arbitrary aspect ratios. In this paper, we propose a novel image retargeting approach based on region warping, which emphasizes the image parts with important content while reducing the visual distortion over the whole image. First, the original image is decomposed into homogeneous regions and further represented by curve-edge trapezoid meshes. Then, two kinds of energy maps, importance map and sensitivity map, are calculated by visual attention model and weighted gradient map respectively. With mesh representation and energy map constraints, image retargeting is formulated to a constrained optimization problem of mesh vertexes relocation. Finally, the target image is generated by separately warping the regions based on the deduced optimal solution. The experiments on different images demonstrate the effective and efficiency of our algorithm. Tongwei Ren, Yan Liu 0004, Gangshan Wu |
ACM Multimedia | 3 |
| 2008 | Constrained sampling for image retargetingabstractIn this paper, we present a new approach for retargeting large images to mobile devices with small screens. As the core of image retargeting, information fidelity is adequately considered in terms of reservations of salient regions, edge integrity, and image layout. By taking these aspects as constraints, image retargeting is formulated as a constrained sampling task. Each pixel in image is first represented with a vector encoding the constraints. Then, pixels with the same vector values combine to form blocks, and the original image is thus converted into a graph representation. Thereafter, the sampling ratio of each block is determined with a balanced minimum cost flow algorithm. Final result is generated by an interpolated sampling scheme and direct scaling. Experiments demonstrate the effectiveness of the proposed approach. Tongwei Ren, Yanwen Guo 0001, Gangshan Wu, Fuyan Zhang |
ICME | 3 |
| 2003 | On Intra-page and Inter-page Semantic Analysis of Web Pages
Jun Wang 0018, Gangshan Wu, Hiroshi Tsuda |
PACLIC | 3 |
| 2000 | Synchronization validation mechanism in multimedia document presentationabstractSynchronization is the major processing task during multimedia presentation, which defines the time relation between multimedia objects. There are many synchronization specification methods, and each has different advantages and disadvantages. Those methods work well for the simple and specific synchronization case, but when considering large complex multimedia presentation, there lacks an effective method to deal with the complex time relation. In this paper, a synchronization point based specification method (SPM) is proposed, and based on it, a static and dynamic validation mechanism for the synchronization relation is given. In SPM, the synchronization relation is defined in synchronization points (SP). Multimedia presentation control is based on SP's status during presentation. According to the order of SPs, SPM gives the static and dynamic validation algorithm, which can avoid the conflict relation between SPs. SPM gives a way to define the nondetermined synchronization relation between multimedia objects, and proposes static and dynamic validation mechanisms for complex synchronization control. This synchronization point based mechanism can extend the capability of synchronization definition and control in a multimedia system. Gangshan Wu, Fuyan Zhang |
SMC | 1 |