Xiangbo Shu

dblp:169/3410 · DBLP profile ↗
← Back
115ranked-venue papers
14as first author
79since 2021 · last 2026
0000-0003-4902-4663ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 76 · 6 first-author · 51 since 2021Artificial intelligence and machine learning · 50 · 9 first-author · 35 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 3 · 2 since 2021Security and privacy · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Beyond Quadratic: Linear-Time Change Detection with RWKV
abstract
Existing paradigms for remote sensing change detection are caught in a trade-off: CNNs excel at efficiency but lack global context, while Transformers capture long-range dependencies at a prohibitive computational cost. This paper introduces ChangeRWKV, a new architecture that reconciles this conflict. By building upon the Receptance Weighted Key Value (RWKV) framework, our ChangeRWKV uniquely combines the parallelizable training of Transformers with the linear-time inference of RNNs. Our approach core features two key innovations: a hierarchical RWKV encoder that builds multi-resolution feature representation, and a novel Spatial-Temporal Fusion Module (STFM) engineered to resolve spatial misalignments across scales while distilling fine-grained temporal discrepancies. ChangeRWKV not only achieves state-of-the-art performance on the LEVIR-CD benchmark, with an 85.46% IoU and 92.16% F1 score, but does so while drastically reducing parameters and FLOPs compared to previous leading methods. This work demonstrates a new, efficient, and powerful paradigm for operational-scale change detection.
Gensheng Pei, Tao Chen 0012, Xia Yuan, Haofeng Zhang 0001, Xiangbo Shu, Yazhou Yao
AAAI6
2026 Spatiotemporal-Untrammelled Mixture of Experts for Multi-Person Motion Prediction
abstract
Comprehensively and flexibly capturing the complex spatio-temporal dependencies of human motion is critical for multi-person motion prediction. Existing methods grapple with two primary limitations: i) Inflexible spatiotemporal representation due to reliance on positional encodings for capturing spatiotemporal information. ii) High computational costs stemming from the quadratic time complexity of conventional attention mechanisms. To overcome these limitations, we propose the Spatiotemporal-Untrammelled Mixture of Experts (ST-MoE), which flexibly explores complex spatio-temporal dependencies in human motion and significantly reduces computational cost. To adaptively mine complex spatio-temporal patterns from human motion, our model incorporates four distinct types of spatiotemporal experts, each specializing in capturing different spatial or temporal dependencies. To reduce the potential computational overhead while integrating multiple experts, we introduce bidirectional spatiotemporal Mamba as experts, each sharing bidirectional temporal and spatial Mamba in distinct combinations to achieve model efficiency and parameter economy. Extensive experiments on four multi-person benchmark datasets demonstrate that our approach not only outperforms state-of-art in accuracy but also reduces model parameter by 41.38% and achieves a 3.6× speedup in training.
Zheng Yin, Chengjian Li, Xiangbo Shu, Meiqi Cao, Rui Yan 0010, Jinhui Tang 0001
AAAI3
2026 AS-CAR: adaptive topology evolution with semantic alignment for continual action recognition
Xingyu Zhu 0008, Xiangbo Shu, Binqian Xu, Jinhui Tang 0001
Sci. China Inf. Sci.2
2026 Spatio-Temporal Decoupled Knowledge Compensator for Few-Shot Action Recognition
abstract
Few-Shot Action Recognition (FSAR) is a challenging task that requires recognizing novel action categories with a few labeled videos. Recent works typically apply semantically coarse category names as auxiliary contexts to guide the learning of discriminative visual features. However, such context provided by the action names is too limited to provide sufficient background knowledge for capturing novel spatial and temporal concepts in actions. In this paper, we propose DiST, an innovative Decomposition-incorporation framework for FSAR that makes use of decoupled Spatial and Temporal knowledge provided by large language models to learn expressive multi-granularity prototypes. In the decomposition stage, we decouple vanilla action names into diverse spatio-temporal attribute descriptions (action-related knowledge). Such commonsense knowledge complements semantic contexts from spatial and temporal perspectives. In the incorporation stage, we propose Spatial/Temporal Knowledge Compensators (SKC/TKC) to discover discriminative object-level and frame-level prototypes, respectively. In SKC, object-level prototypes adaptively aggregate important patch tokens under the guidance of spatial knowledge. Moreover, in TKC, frame-level prototypes utilize temporal attributes to assist in inter-frame temporal relation modeling. These learned prototypes thus provide transparency in capturing fine-grained spatial details and diverse temporal patterns. Experimental results show DiST achieves state-of-the-art results on five standard FSAR datasets.
Hongyu Qu, Xiangbo Shu, Rui Yan 0010, Hailiang Gao, Wenguan Wang, Jinhui Tang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 MN-AQA: Multi-stage neuro-symbolic action quality assessment for explainable diving scoring
Huilin Ge, Bingying Hu, Xiangbo Shu, Meiqi Cao, Zheng Wang 0007
Pattern Recognit.4
2026 Locality-Aware Cross-Modal Correspondence Learning for Dense Audio-Visual Events Detection
abstract
Dense-localization Audio-Visual Events (DAVE) aims to identify time boundaries and corresponding categories for events that are both audible and visible in a long video, where events may co-occur and exhibit varying durations. However, complex audio-visual scenes often involve asynchronization between modalities, making accurate localization challenging. Existing DAVE solutions extract audio and visual features through unimodal encoders, and fuse them via dense cross-modal interaction. However, independent unimodal encodingstruggles to emphasize shared semantics between modalitieswithout cross-modal guidance, while dense cross-modal attention mayover-attend to semantically unrelated audio-visual features. To address these problems, we present LOCO, a Locality-aware cross-modal Correspondence learning framework for DAVE. LOCO leverages the local temporal continuity of audio-visual events as important guidance to filter irrelevant cross-modal signals and enhance cross-modal alignment throughout both unimodal and cross-modal encoding stages. i) Specifically, LOCO applies Local Correspondence Feature (LCF) Modulation to enforce unimodal encoders to focus on modality-shared semantics by modulating agreement between audio and visual features based on local cross-modal coherence. ii) To better aggregate cross-modal relevant features, we further customize Local Adaptive Cross-modal (LAC) Interaction, which dynamically adjusts attention regions in a data-driven manner. This adaptive mechanism focuses attention on local event boundaries and accommodates varying event durations. By incorporating LCF and LAC, LOCO provides solid performance gains and outperforms existing DAVE methods. The source code will be released.
Ling Xing 0003, Hongyu Qu, Rui Yan 0010, Xiangbo Shu, Jinhui Tang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Federated Unsupervised Skeletal Action Recognition From Condensation to Expansion
Jinjin Gong, Binqian Xu, Jiachao Zhang, Xiangbo Shu
IEEE Trans. Inf. Forensics Secur.4
2026 ThinkMatter: Panoramic-Aware Instructional Semantics for Monocular Vision-and-Language Navigation
abstract
Vision-and-Language Navigation in continuous environments (VLN-CE) requires an embodied robot to navigate the target destination following the natural language instruction. Most existing methods use panoramic RGB-D cameras for 360° observation of environments. However, these methods struggle in real-world applications because of the higher cost of panoramic RGB-D cameras. This paper studies a low-cost and practical VLN-CE setting, e.g., using monocular cameras of limited field of view, which means "Look Less" for visual observations and environment semantics. In this paper, we propose a ThinkMatter framework for monocular VLN-CE, where we motivate monocular robots to "Think More" by 1) generating novel views and 2) integrating instruction semantics. Specifically, we achieve the former by the proposed 3DGS-based panoramic generation to render novel views at each step, based on past observation collections. We achieve the latter by the proposed enhancement of the occupancy-instruction semantics, which integrates the spatial semantics of occupancy maps with the textual semantics of language instructions. These operations promote monocular robots with wider environment perceptions as well as transparent semantic connections with the instruction. Both extensive experiments in the simulators and real-world environments demonstrate the effectiveness of ThinkMatter, providing a promising practice for real-world navigation.
Guangzhao Dai, Shuo Wang 0015, Hao Zhao 0002, Bin Zhu 0006, Qianru Sun, Xiangbo Shu
IEEE Trans. Image Process.6
2026 Attack-Augmented Mixing-Contrastive Skeletal Representation Learning
abstract
Contrastive learning facilitates the acquisition of informative skeleton representations for unsupervised action recognition by leveraging effective positive and negative sample pairs. However, most existing methods construct these pairs through weak or strong data augmentations, which typically rely on random appearance alterations of skeletons. While such augmentations are somewhat effective, they introduce semantic variations only indirectly and face two inherent limitations. First, simply modifying the appearance of skeletons often fails to reflect meaningful semantic variations. Second, random perturbations can unintentionally blur the boundary between positive and negative pairs, weakening the contrastive objective. To address these challenges, we propose an attack-driven augmentation framework that explicitly introduces semantic-level perturbations. This approach facilitates the generation of hard positives while guiding the model to mine more informative hard negatives. Building on this idea, we present Attack-Augmented Mixing-Contrastive Skeletal Representation Learning (A2MC), a novel framework that focuses on contrasting hard positive and hard negative samples for more robust representation learning. Within A2MC, we design an Attack-Augmentation (Att-Aug) module that integrates both targeted (attack-based) and untargeted (augmentation-based) perturbations to generate informative hard positive samples. In parallel, we propose the Positive-Negative Mixer (PNM), which blends hard positive and negative features to synthesize challenging hard negatives. These are then used to update a mixed memory bank for more effective contrastive learning. Comprehensive evaluations across three public benchmarks demonstrate that our approach, termed A2MC, achieves performance on par with or exceeding existing state-of-the-art methods.
Binqian Xu, Xiangbo Shu, Jiachao Zhang, Rui Yan 0010, Guosen Xie
IEEE Trans. Image Process.2
2026 Location Matters: Frequency-Spatial Dual-Space Adaptation for Cross-Domain Few-Shot Segmentation
abstract
Current cross-domain few-shot semantic segmentation (CD-FSS) methods tend to overlook a fundamental yet domain-agnostic prior: the spatial correspondence between support and query images driven by the task itself. Unlike semantic similarity, this spatial correlation arises from the consistent structural layout of foreground objects across domains. To exploit this structural prior, we propose a novel frequency-spatial dual space adaptation (FDSA) framework, to learn domain-invariant structures and task-specific priors by jointly suppressing domain-specific redundancy in frequency domain and reinforcing geometric priors in spatial domain. Specifically, FDSA consists of two sequential modules, i.e., the frequency structural adapter (FSA) and the spatial geometry adapter (SGA). FSA performs image modulation in the frequency domain by emphasizing low-frequency foreground semantics and attenuating high-frequency noise, thus maintaining structural integrity of these input images. By contrast, SGA leverages handcrafted local descriptors to extract keypoints from both support and query images, generating Gaussian-based geometric priors that highlight desirable aligned regions. Additionally, we introduce spatial-guided SAM refinement (SSR) to extend our spatial geometric prior into the Segment Anything Model (SAM). SSR generates a soft Gaussian point prompt centered on the coarse mask, enabling SAM to refine segmentation masks without manual intervention. This integration effectively bridges task-specific localization with high-quality segmentation. Extensive experiments on four standard CD-FSS benchmarks demonstrate that our method achieves new state-of-the-art performance. Code is available at https://github.com/CVL-hub/FDSA.git.
Guolei Sun, Yong Li 0032, Hongsong Wang 0001, Xiangbo Shu, Guosen Xie
IEEE Trans. Image Process.5
2026 MUST: Multi-Scale Structural-Temporal Link Prediction Model for UAV Ad Hoc Networks
abstract
Predicting future connections among Unmanned Aerial Vehicles (UAVs) is one of the fundamental tasks in UAV Ad Hoc Networks (UANETs). In adversarial environments where UAV operational information is unavailable, future link prediction must rely solely on the observed historical topo logical data. However, the highly dynamic and sparse nature of UANET topologies poses substantial challenges in capturing structural and temporal link formation patterns. Most existing link prediction methods focus only on single-scale structural features while neglecting the effects of network sparsity, thus limiting their performance when applied to UANETs. In this paper, we propose MUST, a Multi-scale Structural-Temporal link prediction model for UANETs. In our model, multi-scale structural representations are learned using a weighted graph attention network combined with multi-scale pooling, capturing features at the levels of individual UAVs, UAV communities, and the entire network, which are then fused via concatenation. Then, a stacked long short-term memory network is employed to learn the temporal dynamics of these multi-scale structural features. To address the impact of network sparsity, we develop a tailored loss function that emphasizes the contribution of existing links during training. We validate the performance of MUST using several UANET datasets generated through simulations. Extensive experimental results demonstrate that MUST achieves state-of-the-art link prediction performance in highly dynamic and sparse UANETs.
Cunlai Pu, Fangrui Wu, Rajput Ramiz Sharafat, Guangzhao Dai, Xiangbo Shu
IEEE Trans. Knowl. Data Eng.5
2026 Dual-Attention Video Representation Learning for Parameter Efficient Text-Video Retrieval
abstract
Recently, significant progress has been made in text-video retrieval by adapting CLIP to the text-video domain. To mitigate the computational overhead of full fine-tuning, current studies have shifted focus toward parameter-efficient fine-tuning strategies, such as Adapters, Prompts, and LoRA. However, most methods generally obtain the global video representation by frame-level aggregation,i.e.,pooling operation over [CLS] token in each frame, which fails to capture the local details. Meanwhile, these methods process all visual tokens indiscriminately when learning video representations, overlooking the interference introduced by redundant tokens. To tackle these issues, this paper proposes aDual-Attention Video Representation Learning(DAVRL) framework which synergistically combines dual-attention based global-local interactions with token selection. Specifically, we incorporate a dual-attention global-local interaction paradigm and formulate a progressive cross-frame interaction module, which implement global interactions between well-designed learnable global video tokens and all visual tokens, and progressive local interactions from single-frame to multi-frame scales. Moreover, we propose a semantic-aware token selection module to prune the spatial-temporal redundant tokens, under the guidance of both global and local semantics. As such, our DAVRL framework can learn more informative yet discriminative global video representations while mitigating the interference of redundant tokens. Extensive experiments demonstrate the superiority of our DAVRL on diverse datasets with only0.32% tunable parameters.
Ziyi Bian, Yadan Luo, Fangzhi Zhu, Xiangbo Shu, Zheng Zhang 0006
IEEE Trans. Multim.4
2025 3D-aware Select, Expand, and Squeeze Token for Aerial Action Recognition
abstract
Aerial Action Recognition (AAR) in videos captured by Unmanned Aerial Vehicles (UAVs) plays a vital role in numerous applications. However, current methods related to traditional action recognition primarily cater to fixed or near cameras, and rarely consider the movement disturbance of UAVs, including their varying attitudes and positions. Those characteristics of aerial videos bring moving objects in small regions compared to broad backgrounds and relative movement to the motion of objects, which reflect more sparse and disturbed semantic information for AAR. To address these issues, we present a novel framework, dubbed 3D-Tok, to Select, Expand, and Squeeze original visual tokens for obtaining compact yet diverse semantic-enhanced tokens. In particular, we present a 3D-token selector (3TS) to select complex yet diverse tokens in three channels, capturing the semantic awareness of moving objects in comparatively small regions. Additionally, to get rid of disturbed semantic information caused by the UAV flight, we present an Expand-Squeeze Converter (ESC) to adaptively expand and squeeze the 3D-selected tokens constrained by contrastive loss, thereby suppressing the semantic-irrelevant information and reinforce semantic-relevant information via the interpolation converting. By involving the token selecting, expanding, and squeezing into an all-in-one framework, 3D-Tok shows significant improvements on the UAV-Human dataset(↑9.5%), RoCoG-v2 dataset (↑23.5%), and Drone-Action dataset (↑5.7%).
Luying Peng, Xiangbo Shu, Yazhou Yao, Guosen Xie
AAAI2
2025 Kernel-Aware Graph Prompt Learning for Few-Shot Anomaly Detection
abstract
Few-shot anomaly detection (FSAD) aims to detect unseen anomaly regions with the guidance of very few normal support images from the same class. Existing FSAD methods usually find anomalies by directly designing complex text prompts to align them with visual features under the prevailing large vision-language model paradigm. However, these methods, almost always, neglect intrinsic contextual information in visual features, e.g., the interaction relationships between different vision layers, which is an important clue for detecting anomalies comprehensively. To this end, we propose a kernel-aware graph prompt learning framework, termed as KAG-prompt, by reasoning the cross-layer relations among visual features for FSAD. Specifically, a kernel-aware hierarchical graph is built by taking the different layer features focusing on anomalous regions of different sizes as nodes, meanwhile, the relationships between arbitrary pairs of nodes stand for the edges of the graph. By message passing over this graph, KAG-prompt can capture cross-layer contextual information, thus leading to more accurate anomaly prediction. Moreover, to integrate the information of multiple important anomaly signals in the prediction map, we propose a novel image-level scoring method based on multi-level information fusion. Extensive experiments on MVTecAD and VisA datasets show that KAG-prompt achieves state-of-the-art FSAD results for image-level/pixel-level anomaly detection.
Fenfang Tao, Guosen Xie, Fang Zhao 0006, Xiangbo Shu
AAAI4
2025 Seeing What Matters: Empowering CLIP with Patch Generation-to-Selection
abstract
The CLIP model has demonstrated significant advancements in aligning visual and language modalities through large-scale pre-training on image-text pairs, enabling strong zero-shot classification and retrieval capabilities on various domains. However, CLIP’s training remains computationally intensive, with high demands on both data processing and memory. To address these challenges, recent masking strategies have emerged, focusing on the selective removal of image patches to improve training efficiency. Although effective, these methods often compromise key semantic information, resulting in suboptimal alignment between visual features and text descriptions. In this work, we present a concise yet effective approach called Patch Generation-to-Selection (CLIP-PGS) to enhance CLIP’s training efficiency while preserving critical semantic con tent. Our method introduces a gradual masking process in which a small set of candidate patches is first pre-selected as potential mask regions. Then, we apply Sobel edge detection across the entire image to generate an edge mask that prioritizes the retention of the primary object areas. Finally, similarity scores between the candidate mask patches and their neighboring patches are computed, with optimal transport normalization refining the selection process to ensure a balanced similarity matrix. Our approach, CLIP-PGS, sets new state-of-the-art results in zero-shot classification and retrieval tasks, achieving superior performance in robustness evaluation and language compositionality benchmarks.
Gensheng Pei, Tao Chen 0012, Xinhao Cai, Xiangbo Shu, Tianfei Zhou, Yazhou Yao
CVPR5
2025 Cycle-Consistent Learning for Joint Layout-to-Image Generation and Object Detection
Xinhao Cai, Qiuxia Lai, Gensheng Pei, Xiangbo Shu, Yazhou Yao, Wenguan Wang
ICCV4
2025 Exploiting Frequency Dynamics for Enhanced Multimodal Event-Based Action Recognition
Meiqi Cao, Xiangbo Shu, Xin Jiang 0010, Rui Yan 0010, Yazhou Yao, Jinhui Tang 0001
ICCV2
2025 Tensor-Aggregated LoRA in Federated Fine-Tuning
Binqian Xu, Xiangbo Shu, Jiachao Zhang, Yazhou Yao, Guosen Xie, Jinhui Tang 0001
ICCV3
2025 CA2C: A Prior-Knowledge-Free Approach for Robust Label Noise Learning via Asymmetric Co-Learning and Co-Training
Mengmeng Sheng, Zeren Sun, Tianfei Zhou, Xiangbo Shu, Jinshan Pan, Yazhou Yao
ICCV4
2025 Learning Clustering-based Prototypes for Compositional Zero-Shot Learning
abstract
Learning primitive (i.e., attribute and object) concepts from seen compositions is the primary challenge of Compositional Zero-Shot Learning (CZSL). Existing CZSL solutions typically rely on oversimplified data assumptions, e.g., modeling each primitive with a single centroid primitive presentation, ignoring the natural diversities of the attribute (resp. object) when coupled with different objects (resp. attribute). In this work, we develop ClusPro, a robust clustering-based prototype mining framework for CZSL that defines the conceptual boundaries of primitives through a set of diversified prototypes. Specifically, ClusPro conducts within-primitive clustering on the embedding space for automatically discovering and dynamically updating prototypes. To learn high-quality embeddings for discriminative prototype construction, ClusPro repaints a well-structured and independent primitive embedding space, ensuring intra-primitive separation and inter-primitive decorrelation through prototype-based contrastive learning and decorrelation learning. Moreover, ClusPro effectively performs prototype clustering in a non-parametric fashion without the introduction of additional learnable parameters or computational budget during testing. Experiments on three benchmarks demonstrate ClusPro outperforms various top-leading CZSL solutions under both closed-world and open-world settings. Our code is available at CLUSPRO.
Hongyu Qu, Jianan Wei, Xiangbo Shu, Wenguan Wang
ICLR3
2025 Reliable and Diverse Hierarchical Adapter for Zero-shot Video Classification
abstract
Adapting pre-trained vision-language models to downstream tasks has emerged as a novel paradigm for zero-shot learning. Existing test-time adaptation (TTA) methods such as TPT attempt to fine-tune visual or textual representations to accommodate downstream tasks but still require expensive optimization costs. To this end, Training-free Dynamic Adapter (TDA) maintains a cache containing visual features for each category in a parameter-free manner and measures sample confidence based on prediction entropy of test samples. Inspired by TDA, this work aims to develop the first training-free adapter for zero-shot video classification. Capturing the intrinsic temporal relationships within video data to construct and maintain the video cache is key to extending TDA to the video domain. In this work, we propose a reliable and diverse Hierarchical Adapter for zero-shot video classification, which consists of Frame-level Cache Refiner and Video-level Cache Updater. Before each video sample enters the corresponding cache, it needs to be refined at frame level based on prediction entropy and temporal probability difference. Due to the limited capacity of the cache, we update the cache during inference based on the principle of diversity. Experiments on four popular video classification benchmarks demonstrate the effectiveness of Hierarchical Adapter. The code is available at https://github.com/Gwxer/Hierarchical-Adapter.
Wenxuan Ge, Rui Yan 0010, Hongyu Qu, Guosen Xie, Xiangbo Shu
IJCAI6
2025 OmniGaze: Reward-inspired Generalizable Gaze Estimation in the Wild
abstract
Current 3D gaze estimation methods struggle to generalize across diverse data domains, primarily due to $\textbf{i)}$ $\textit{the scarcity of annotated datasets}$, and $\textbf{ii)}$ $\textit{the insufficient diversity of labeled data}$. In this work, we present OmniGaze, a semi-supervised framework for 3D gaze estimation, which utilizes large-scale unlabeled data collected from diverse and unconstrained real-world environments to mitigate domain bias and generalize gaze estimation in the wild. First, we build a diverse collection of unlabeled facial images, varying in facial appearances, background environments, illumination conditions, head poses, and eye occlusions. In order to leverage unlabeled data spanning a broader distribution, OmniGaze adopts a standard pseudo-labeling strategy and devises a reward model to assess the reliability of pseudo labels. Beyond pseudo labels as 3D direction vectors, the reward model also incorporates visual embeddings extracted by an off-the-shelf visual encoder and semantic cues from gaze perspective generated by prompting a Multimodal Large Language Model to compute confidence scores. Then, these scores are utilized to select high-quality pseudo labels and weight them for loss computation. Extensive experiments demonstrate that OmniGaze achieves state-of-the-art performance on five datasets under both in-domain and cross-domain settings. Furthermore, we also evaluate the efficacy of OmniGaze as a scalable data engine for gaze estimation, which exhibits robust zero-shot generalization on four unseen datasets.
Hongyu Qu, Jianan Wei, Xiangbo Shu, Yazhou Yao, Wenguan Wang, Jinhui Tang 0001
NeurIPS3
2025 Vision-centric Token Compression in Large Language Model
abstract
Real-world applications are stretching context windows to hundreds of thousand of tokens while Large Language Models (LLMs) swell from billions to trillions of parameters. This dual expansion send compute and memory costs skyrocketing, making $\textit{token compression}$ indispensable. We introduce Vision Centric Token Compression ($\textbf{Vist}$), a $\textit{slow–fast}$ compression framework that mirrors human reading: the $\textit{fast}$ path renders distant tokens into images, letting a $\textbf{frozen, lightweight vision encoder}$ skim the low-salience context; the $\textit{slow}$ path feeds the proximal window into the LLM for fine-grained reasoning. A Probability-Informed Visual Enhancement (PVE) objective masks high-frequency tokens during training, steering the Resampler to concentrate on semantically rich regions—just as skilled reader gloss over function words. On eleven in-context learning benchmarks, $\textbf{Vist}$ achieves the same accuracy with 2.3$\times$ fewer tokens, cutting FLOPs by 16\% and memory by 50\%. This method delivers remarkable results, outperforming the strongest text encoder-based compression method CEPE by $\textbf{7.6}$\% on average over benchmarks like TriviaQA, NQ, PopQA, NLUI, and CLIN, setting a new standard for token efficiency in LLMs. The project is at https://github.com/CSU-JPG/VIST.
Ling Xing 0003, Alex Jinpeng Wang, Rui Yan 0010, Xiangbo Shu, Jinhui Tang 0001
NeurIPS4
2025 You Only Communicate Once: One-shot Federated Low-Rank Adaptation of MLLM
abstract
Multimodal Large Language Models (MLLMs) with Federated Learning (FL) can quickly adapt to privacy-sensitive tasks, but face significant challenges such as high communication costs and increased attack risks, due to their reliance on multi-round communication. To address this, One-shot FL (OFL) has emerged, aiming to complete adaptation in a single client-server communication. However, existing adaptive ensemble OFL methods still need more than one round of communication, because correcting heterogeneity-induced local bias relies on aggregated global supervision, meaning they still do not achieve true one-shot communication. In this work, we make the first attempt to achieve true one-shot communication for MLLMs under OFL, by investigating whether implicit (i.e., initial rather than aggregated) global supervision alone can effectively correct local training bias. Our key finding from the empirical study is that imposing directional supervision on local training substantially mitigates client conflicts and local bias. Building on this insight, we propose YOCO, in which directional supervision with sign-regularized LoRA B enforces global consistency, while sparsely regularized LoRA A preserves client-specific adaptability. Experiments demonstrate that YOCO cuts communication to $\sim$0.03\% of multi-round FL while surpassing those methods in several multimodal scenarios and consistently outperforming all one-shot competitors.
Binqian Xu, Haiyang Mei, Zechen Bai, Jinjin Gong, Rui Yan 0010, Guosen Xie, Yazhou Yao, Basura Fernando, Xiangbo Shu
NeurIPS9
2025 Structure-aware contrastive learning for glomerulus segmentation in renal pathology
Xiangbo Shu, Yuhui Zheng, Jin Ding, Xianghui Fu
Image Vis. Comput.3
2025 Maximum dissimilarity channel complementary reconstruction for convolutional efficiency
Shougang Ren, Xingjian Gu, Xiangbo Shu
Multim. Syst.5
2025 SODA: Structural Pre-Decoupling and Co-Aligning for Video Compositional Representation
abstract
The core challenge in video compositional representation lies in jointly identifying atomic actions (verbs) and their associated objects (nouns), while generalizing to unseen verb-noun combinations. Existing approaches often adopt a shared backbone with multi-head classifiers, leading to semantic entanglement and recognition imbalance, especially under ViT-based paradigms. To tackle these issues, we propose a novel framework, Structural Pre-Decoupling and Co-Aligning (SODA), which structurally decouples the early-stage learning paths of verbs and nouns, mitigating semantic interference and recognition imbalance. A core component is the Divide-and-Conquer Disentanglement (DCD) module, comprising two parallel paths: an Object-Removed Verb (OR-V) path that explicitly suppresses object appearance by integrating frame-difference features for short-term motion and trajectory cues for long-term dynamics, and an Object-Centric Noun (OC-N) path that constructs adaptive key-frame representations in the guidance of dynamic features from OR-V. On top of these, a Compositional Co-Alignment (CCA) strategy is further introduced that aligns atomic and compositional representations in a shared semantic space through contrastive learning, capturing implicit commonsense associations of verb-noun pairs while preserving the discriminative power of each stream. Extensive experiments on both standard and zero-shot compositional action recognition benchmarks validate the effectiveness and generalization of our approach.
Wenxuan Ge, Henghao Zhao, Xiangbo Shu
IEEE Signal Process. Lett.5
2025 Revisiting Few-Shot Compositional Action Recognition With Knowledge Calibration
abstract
The primary challenge of Few-shot Compositional Action Recognition (FSCAR) lies in effectively generalizing to and identifying unseen compositions (i.e., motions and objects) from only a few labeled videos. However, current approaches typically evaluate FSCAR as a subsidiary task of standard CAR, ignoring the insufficient generalization of the former in the regime of more significant distribution bias and limited data. To this end, we thoroughly revisit FSCAR, explicitly acknowledging the crucial role of fine-tuning, and propose a novel trio-tuning-testing framework to alleviate the problem of compositional generalization in few-shot scenarios. Specifically, we devise an effective inner-to-outer baseline, namely Motion-Object Composer (MOC), to hierarchically learn comprehensive representations from concepts to compositions. Furthermore, we propose the Trio-Knowledge Calibration (TKC) strategy to calibrate the inference, by transferring the prior visual and language knowledge learned from fine-tuning. Extensive experiments demonstrate the state-of-the-art performance of the approach compared to current competitive methods.
Hongyu Qu, Xiangbo Shu
IEEE Signal Process. Lett.3
2025 Toward Physically Stable Motion Generation: A New Paradigm of Human Pose Representation
abstract
In machine learning, generating realistic human motion is paramount for a range of applications that require lifelike movements. Traditional methods have often overlooked the adherence to physical principles, leading to motion sequences that exhibit unrealistic behaviors such as foot sliding, penetration, and floating. These issues are particularly pronounced in complex tasks like dance choreography, which demand a higher degree of fidelity and realism. To address these challenges, we introduce RF-Rotation, a novel approach to human pose representation that strategically repositions the root joint of the SMPL model to align with both feet, while representing other joints through recursive bone rotations. It not only aligns more closely with the natural dynamics of human movement but also integrates an advanced contact predictor to ascertain the ground contact status of both feet, thereby preventing physically implausible movements on feet. We note that RF-Rotation is compatible with any motion generation tasks, including dance choreography, text-to-motion synthesis, and motion prediction, and can be seamlessly integrated into existing frameworks without modifications. Extensive experiments across three distinct tasks demonstrate the superior performance of RF-Rotation in enhancing the realism and stability of generated motion sequences. This method can significantly reduce foot sliding, floating, and penetration issues, without affecting computational efficiency, underscores its potential to set new standards in human motion generation.
Qiongjie Cui, Zhenyu Lou, Zhenbo Song, Xiangbo Shu
IEEE Trans. Circuits Syst. Video Technol.4
2025 Appearance-Agnostic Representation Learning for Compositional Action Recognition
abstract
The discussion of compositional generalization in action recognition,i.e., Compositional Action Recognition (CAR), has recently received increasing attention. CAR challenges models to recognize unseen combinations of actions and objects, with the primary challenge being the distribution shift from training to testing. Most previous approaches for CAR incorporate supplementary object annotations (e.g. bounding boxes and objects categories) to learn an instance-centric dynamic representation. However, these methods inevitably introduce stronger visual inductive bias, including object appearance and background bias, that impact generalization performance, particularly in out-of-distribution scenarios. To this end, this work attempts to construct an appearance-agnostic de-biased representation by leveraging the powerful segmentation capability of Segment Anything Model (SAM), which is the first exploration of SAM in the field of compositional action recognition. Specifically, we propose a novel SAM-driven Appearance-Agnostic Representation Learning (A2RL) framework for CAR, which contains two effective sub-modules: Fore-Back Mask (FBM) and Dynamic Relation Modeling (DRM). In FBM, we design a fine-grained instance-invisible and background-removed masking strategy to effectively weaken the strong connection between visual cues and action labels, as well as minimize the impact of irrelevant factors. In DRM, we explore the potential association between subjects and objects involved in one action and then build appearance-agnostic relational descriptors for dynamic modeling. Extensive experiments demonstrate the generalization ability of this work. Notably, FBM achieves significant improvements in all three compositional settings without adding any additional model parameters. The proposed also gains state-of-the-art performance in comparison with the most recent methods in CAR.
Xiangbo Shu, Rui Yan 0010, Zhewei Tu, Jinhui Tang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Eliminating Semantic Ambiguity in Human Pose Estimation via Stable Feature Upsampling
abstract
Human pose estimation is a challenging research task in the computer vision community due to the semantic ambiguity problem caused by inevitable occlusions, varying body shapes, and complex articulations. Although deep learning-based methods have significantly improved the performance of this task, existing feature upsampling operations,e.g., bilinear interpolation and transposed convolution, within current convolutional neural networks and Transformer frameworks suffer from a multitude of limitations, including the inability to adapt to specific tasks and the loss of fine-grained semantic details. In this work, we propose a simple yet effective two-step stable feature upsampling (SIU) strategy that addresses these limitations by leveraging a learnable and efficient upsampling operation. Specifically, wefirstapply periodic shuffling to increase the resolution of the feature maps.Secondly, we utilize convolution layers to adjust the size of feature channels to match those of the input feature maps. The proposed SIU enables the entire network to adapt to the specific feature requirements of the human pose estimation task, making it more effective in preserving spatial information. Quantitatively, extensive experimental results on the challenging COCO-WholeBody dataset validate that our approach outperforms state-of-the-art methods accurately and efficiently, and possesses strong transferability, making it applicable to a wide range of baselines. Moreover, the qualitative results validate that SIU can effectively eliminate the semantic ambiguity problem in challenging pose scenarios, such as occlusions and overlapping1.
Rui Yan 0010, Xiangbo Shu, Pingcheng Dong, Long Chen 0016, Xiaoyu Du 0002
IEEE Trans. Circuits Syst. Video Technol.4
2025 STPM: Spatial-Temporal Token Pruning and Merging for Complex Activity Recognition
abstract
Lightweight video representation techniques have advanced significantly for simple activity recognition, but they still encounter several issues when applied to complex activity recognition: 1) The presence of numerous individuals and varying spatial positions makes it difficult for traditional token pruning methods to maintain accuracy. 2) Simply discarding entire frames may result in the loss of crucial clues. 3) To maintain parallel computing, applying the same pruning rate to every frame leads to significant redundancy in frames with low information content. To this end, we propose a lightweight and novel Spatial-Temporal Token Pruning and Merging (STPM) framework, specifically designed for complex action videos where human actors occupy a small spatial resolution within video frames. Our framework considers two critical factors: semantic importance and spatial-temporal redundancy, to further reduce overhead. For semantic importance, STPM captures class-specific attention scores by learning multiple class tokens within the transformer to guide token pruning. For spatial-temporal redundancy, STPM employs an anchor graph and temporal attention to perform spatial and temporal token merging, preserving appearance and temporal cues while eliminating semantic duplication and redundancy. We conduct extensive experiments on JRDB-PAR primarily using recently introduced video transformer backbones, e.g., MViT and ViT. Our framework achieves similar results while requiring 40% less computation.
Yumeng Su, Jiachao Zhang, Rui Yan 0010, Pengpeng Li 0001, Guosen Xie, Xiangbo Shu
IEEE Trans. Circuits Syst. Video Technol.6
2025 MoCoLSK: Modality-Conditioned High-Resolution Downscaling for Land Surface Temperature
abstract
Land surface temperature (LST) is a critical parameter for environmental studies, but directly obtaining high spatial resolution LST data remains challenging due to the spatiotemporal tradeoff in satellite remote sensing. Guided LST downscaling has emerged as an alternative solution to overcome these limitations, but current methods often neglect spatial nonstationarity, and there is a lack of an open-source ecosystem for deep learning methods. In this article, we propose the modality-conditioned large selective kernel (MoCoLSK) network, a novel architecture that dynamically fuses multimodal data through modality-conditioned projections. MoCoLSK achieves a confluence of dynamic receptive field adjustment and multimodal feature fusion, leading to enhanced LST prediction accuracy. Furthermore, we establish the GrokLST Project, a comprehensive open-source ecosystem featuring the GrokLST dataset, a high-resolution (HR) benchmark, and the GrokLST toolkit, an open-source PyTorch-based toolkit encapsulating MoCoLSK alongside 40+ state-of-the-art approaches. Extensive experimental results validate MoCoLSK’s effectiveness in capturing complex dependencies and subtle variations within multispectral data, outperforming existing methods in LST downscaling. Our code, dataset, and toolkit are available athttps://github.com/GrokCV/GrokLST.
Qun Dai, Chunyang Yuan, Yimian Dai, Yuxuan Li 0004, Xiang Li 0041, Kang Ni, Jianhui Xu, Xiangbo Shu, Jian Yang 0003
IEEE Trans. Geosci. Remote. Sens.8
2025 Multi-Granularity Aggregation Network for Remote Sensing Few-Shot Segmentation
abstract
Few-shot semantic segmentation (FSS) aims to segment a query image using a limited number of densely annotated support images from the same category. Most existing conventional FSS methods are tailored for coping with images from natural scenarios. Unlike natural images, remote sensing images usually have a similar background context among the support and query image pairs, and more severe intraclass inconsistency exists due to overhead shooting views. However, facing such realistic and challenging remote sensing FSS tasks, the existing methods seldom consider these intrinsic characteristics from a unified viewpoint, thus leading to inferior results. To solve the above dilemma, we propose a multi-granularity aggregation network (MGANet) to progressively capture multi-granularity discriminative information, for tackling the remote sensing FSS task. Specifically, MGANet consists of a multi-granularity similarity (MGS) module and an adaptive multiprototype aggregation (AMPA) module. To fully utilize background context, MGS extracts multi-granularity support and query feature maps from the backbone network to calculate a holistic correlation by incorporating the background information. Next, to alleviate the intraclass inconsistency of remote sensing images, AMPA decomposes the support foreground region into mainstay and auxiliary subregions by the guidance of reverse prediction on support features, thus generating three types of prototypes by masked average pooling (MAP) on these paired features and masks. Furthermore, these multiprototypes are collaboratively interacted with the query features to pursue reinforced discriminative features, relying on prototype-aware slot attention (PASA). Extensive experiments on iSAID-$5^{i}$and LoveDA-$2^{i}$demonstrate well the superiority of the proposed MGANet. The source code is available athttps://github.com/CVL-hub/MGANet/.
Shi-Feng Peng, Guosen Xie, Fang Zhao 0006, Xiangbo Shu, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Uncertainty-Aware Transformer for Referring Camouflaged Object Detection
abstract
Referring camouflaged object detection (Ref-COD) is a recently proposed task, aiming to segment specified camouflaged objects by leveraging visual reference, i.e., a small set of referring images with salient target objects. Ref-COD poses a considerable challenge due to the difficulty of discerning camouflaged objects from their highly similar backgrounds, as well as the significant feature differences between the camouflaged objects and the provided visual reference. To tackle the above dilemma, we propose a novel uncertainty-aware transformer for the Ref-COD task, termed UAT. UAT first utilizes a cross-attention mechanism to align and integrate visual reference to guide camouflaged feature learning, and then models dependencies between patches in a probabilistic manner to learn predictive uncertainty and excavate discriminative camouflaged features. Specifically, we first design a referring feature aggregation (RFA) module to align and incorporate referring features with camouflaged features, guiding targeted specific feature learning within the feature space of camouflaged images. Then, to enhance multi-level feature extraction, we develop a cross-attention encoder (CAE) to integrate global information and multi-scale semantics between adjacent layers to excavate critical camouflage cues. More importantly, we propose a transformer probabilistic decoder (TPD) to model the dependencies between patches as Gaussian random variables to capture uncertainty-aware camouflaged features. Extensive experiments on the golden Ref-COD benchmark demonstrate the superiority of UAT over existing state-of-the-art competitors. The proposed UAT also achieves competitive performance on several conventional COD datasets, further demonstrating its scalability. The source code is available at https://github.com/CVL-hub/UAT.
Ranwan Wu, Tian-Zhu Xiang, Guosen Xie, Rongrong Gao, Xiangbo Shu, Fang Zhao 0006, Ling Shao 0001
IEEE Trans. Image Process.5
2025 Client-Unbiased Skeletal Action Recognizer in Federated Learning
abstract
Edge sensor devices generate vast amounts of user data, but centralized processing poses privacy risks. Federated Learning addresses this by decentralizing training. However, applying Federated Learning directly to skeleton videos fails to preserve motion dynamics and suffers from client heterogeneity bias. To address these limitations, we propose CSAR-a Client-Unbiased Skeletal Action Recognizer for Federated Learning-which tackles two core challenges: motion dynamics preservation and classifier bias mitigation. Specifically, CSAR employs a Model Calibration Loss during client training to align client-server representations and reduce drift. On the server, it generates class-balanced spatiotemporal federated features through Prototypical Gaussian Sampling, subsequently refined via a Motion-aware Differential Loss to capture kinematic properties. These features enable retraining of a globally debiased recognizer that achieves accuracy comparable to real-data-trained models. Further stabilization is achieved through Knowledge Matching, which enhances global understanding. Experiments under natural and label heterogeneity confirm that CSAR outperforms state-of-the-art methods.
Xingyu Zhu 0008, Xiangbo Shu, Jinhui Tang 0001
IEEE Trans. Image Process.2
2025 GPT4Ego: Unleashing the Potential of Pre-Trained Models for Zero-Shot Egocentric Action Recognition
abstract
Vision-Language Models (VLMs), pre-trained on large-scale datasets, have shown impressive performance in various visual recognition tasks. This advancement paves the way for notable performance in some egocentric tasks, Zero-Shot Egocentric Action Recognition (ZS-EAR), entailing VLMs zero-shot to recognize actions from first-person videos enriched in more realistic human-environment interactions. Typically, VLMs handle ZS-EAR as a global video-text matching task, which often leads to suboptimal alignment of vision and linguistic knowledge. We propose a refined approach for ZS-EAR using VLMs, emphasizing fine-grained concept-description alignment that capitalizes on the rich semantic and contextual details in egocentric videos. In this work, we introduce a straightforward yet remarkably potent VLM framework,akaGPT4Ego, designed to enhance the fine-grained alignment of concept and description between vision and language. Specifically, we first propose a new Ego-oriented Text Prompting (EgoTP$\spadesuit$) scheme, which effectively prompts action-related text-contextual semantics by evolving word-level class names to sentence-level contextual descriptions by ChatGPT with well-designed chain-of-thought textual prompts. Moreover, we design a new Ego-oriented Visual Parsing (EgoVP$\clubsuit$) strategy that learns action-related vision-contextual semantics by refining global-level images to part-level contextual concepts with the help of SAM. Extensive experiments demonstrate GPT4Ego significantly outperforms existing VLMs on three large-scale egocentric video benchmarks, i.e., EPIC-KITCHENS-100 (33.2%$\uparrow$$_{\bm {+9.4}}$), EGTEA (39.6%$\uparrow$$_{\bm {+5.5}}$), and CharadesEgo (31.5%$\uparrow$$_{\bm {+2.6}}$). In addition, benefiting from the novel mechanism of fine-grained concept and description alignment, GPT4Ego can sustainably evolve with the advancement of ever-growing pre-trained foundational models. We hope this work can encourage the egocentric community to build more investigation into pre-trained vision-language models.
Guangzhao Dai, Xiangbo Shu, Rui Yan 0010, Jiachao Zhang
IEEE Trans. Multim.2
2025 Hierarchical Motion-Enhanced Matching Framework for Few-Shot Action Recognition
abstract
Few-Shot Action Recognition (FSAR) aims to recognize novel class action with limited annotated training data from the same class. Most FSAR methods subconsciously follow the few-shot image classification solutions by solely focusing on appearance-level matching between support and query videos, such as part-level matching, frame-level matching, and segment-level matching. However, these methods, almost always, have two main limitations: 1) generally ignore the relationship among these part-, frame- and segment-level features and 2) may mismatch the same class actions under fast-term and slow-term dynamics. To this end, we present a novel Hierarchical Motion-enhanced Matching (HM${^{2}}$) framework to hierarchically learn the relation-aware multi-modal features, and jointly promote the multi-modal matching, including appearance-level matching on segments, frames, and parts, as well as the motion-level matching on dynamics. Specifically, we first propose a new Hierarchical Tokenizer (HT) to learn multi-modal features, namely utilizing a hierarchical Transformer to learn appearance-level features, along with a Slow-Fast Aware Motion (SFAM) strategy to learn motion-level features covering fast- and slow-term dynamics. Next, we propose a new Relation-aware Matcher (RM) to match the multi-modal features, by leveraging a Hierarchical Relational Graph Convolutional Network (H-RGCN) to capture the relationship among these appearance-level features. Further, a Dual Sample-to-Class Matching (DSCM) strategy is proposed to measure the bidirectional similarities among appearance- and motion-modal features by sample-to-class matching and class-to-sample matching. Extensive experiments on four golden FSAR datasets demonstrate significant performance improvements of HM${^{2}}$compared with the state-of-the-art methods.
Hailiang Gao, Guosen Xie, Rui Yan 0010, Qiongjie Cui, Hongyu Qu, Xiangbo Shu
IEEE Trans. Multim.6
2025 MVP-Shot: Multi-Velocity Progressive-Alignment Framework for Few-Shot Action Recognition
abstract
Recent few-shot action recognition (FSAR) methods typically perform semantic matching on learned discriminative features to achieve promising performance. However, most FSAR methods focus on single-scale (e.g., frame-level, segment-level,etc.) feature alignment, which ignores that human actions with the same semantic may appear at different velocities. To this end, we develop a novel Multi-Velocity Progressive-alignment (MVP-Shot) framework to progressively learn and align semantic-related action features at multi-velocity levels. Concretely, a Multi-Velocity Feature Alignment (MVFA) module is designed to measure the similarity between features from support and query videos with different velocity scales and then merge all similarity scores in a residual fashion. To avoid the multiple velocity features deviating from the underlying motion semantic, our proposed Progressive Semantic-Tailored Interaction (PSTI) module injects velocity-tailored text information into the video feature via feature interaction on channel and temporal domains at different velocities. The above two modules compensate for each other to make more accurate query sample predictions under the few-shot settings. Experimental results show our method outperforms current state-of-the-art methods on multiple standard few-shot benchmarks (i.e., HMDB51, UCF101, Kinetics, SSv2-full, and SSv2-small).
Hongyu Qu, Rui Yan 0010, Xiangbo Shu, Hailiang Gao, Guosen Xie
IEEE Trans. Multim.3
2025 Prompt-Guided Prototype-Aware Commonality and Discrimination Learning for Zero-Shot Skeleton-Based Action Recognition
abstract
Zero-Shot Skeleton-Based Action Recognition (ZSSAR) is an emerging research field focused on developing alignment models that connect skeleton movements with action definitions, thus enabling generalization to unobserved actions. Current methods often employ generative models to reconstruct cross-modal features or enhance mutual information across modalities for alignment. However, when applied to unseen action categories, these models often neglect the inherent consistency among basic actions, thereby diminishing their generalization capabilities. Furthermore, imprecise annotations fail to capture the rich semantic details of actions, resulting in misalignment. Inspired by human cognitive processes and chain of thought, we argue that integrating prior information about human actions with intrinsic commonality knowledge of basic actions is essential for ZSSAR. To actualize this, we propose a novel method termed Prompt-guided Prototype-aware Commonality and Discrimination Learning (PP-CDL). This method utilize the comprehensive world knowledge contained in LLMs, employing tailored prompts to partition seen action categories into distinct, non-overlapping prototype spaces that embody the commonality knowledge of basic actions. Subsequently, we introduce the Inter- and Intra-Prototype Discriminating (I2PD) module and the Intra-Prototype Commonality Mining (IPCM) module. The I2PD amplifies the distinctiveness of knowledge within prototypes, furnishing a personalized search space for the recognition of unseen actions. In contrast, the IPCM models the shared commonality concept within prototypes, bolstering the consistency between skeleton action representations and corresponding text knowledge representations. Experiments on different skeleton action benchmarks demonstrate the significant improvement of our method over existing alternatives.
Xingyu Zhu 0008, Xiangbo Shu, Jinhui Tang 0001
IEEE Trans. Multim.2
2025 Coarse-Fine Nested Network for Weakly Supervised Group Activity Recognition
abstract
Weakly supervised group activity recognition (WSGAR) aims at identifying the overall behavior of multiple persons without any fine-grained supervision information (including individual position and action label). Traditional methods usually adopt a person-to-whole way: detect persons via off-the-shelf detectors, obtain person-level features, and integrate into the group-level features for training the classifier. However, these methods are unflexible due to serious reliance on the quality of detectors. To get rid of the detector, recent works learn several prototype tokens from noisy grid features with learnable weights directly, which treat all the local visual information equally and bring in redundant and ambiguous information to some extent. To this end, we propose a novel coarse-fine nested network (CFNN) to coarsely localize the key visual patches of activity and further finely learn the local features, as well as the global features. Specifically, we design a nested interactor (NI) to progressively model the spatiotemporal interactions of the learnable global token. According to the cue of spatial interaction in NI, we localize several key visual patches via a new coarse-grained spatial localizer (CSL). Then, we finally encode these localized visual patches with the help of global spatiotemporal dependency via a new fine-grained spatiotemporal selector (FSS). Extensive experiments on Volleyball and NBA datasets demonstrate the effectiveness of the proposed CFNN compared with the existing competitive methods. Code is available at: https://github.com/gexiaojingshelby/CFNN.
Xiaojing Ge, Rui Yan 0010, Xiangbo Shu, Keke Chen, Guosen Xie
IEEE Trans. Neural Networks Learn. Syst.3
2025 Attribute Prompt Alignment Network for Zero-Shot Learning
abstract
In the vanilla zero-shot learning (ZSL) paradigm, category attributes is the key for knowledge generalizable transfer from seen to unseen classes. By contrast, the current contrastive language-image pretraining (CLIP) model relies on the category names to achieve a more general ZSL-like prediction. When vanilla ZSL meets general CLIP, however, most existing methods on both sides struggle to benefit from each other. In this brief, we resort to attribute prompt tuning (APT) for improving the knowledge transferability from the pretrained CLIP model to the downstream ZSL framework for pursuing desirable feature representations. Our approach, termed as attribute prompt alignment network (APAN), leverages APT for cross-network feature alignment (CFA). In this way, we can investigate the effects of CLIP to vanilla ZSL task in the era of large model by the two branch APAN architecture. Specifically, APT takes as an input the templates of class attribute descriptions to produce attribute prompts, which are further used to both guide the localizations of visual regions across two frozen feature extraction networks, through a visual-semantic interaction attention. This enables APAN to progressively refine and align these cross-network features, thus resulting in generalizable feature representations that can capture fine-grained attribute information. For CFA, we simply introduce prediction alignment loss that constrains the predictions from these two cross-network visual features. Experimental results on three benchmark datasets well demonstrate that APAN outperforms the state-of-the-art methods by absorbing generalizable knowledge from CLIP models.
Guosen Xie, Ting Guo 0004, Xiangbo Shu, Fang Zhao 0006, Zheng Zhang 0006, Ling Shao 0001
IEEE Trans. Neural Networks Learn. Syst.4
2025 Global-Local Multiple Granularity Learning for Cross-Modality Visible-Infrared Person Reidentification
abstract
Cross-modality visible-infrared person reidentification (VI-ReID), which aims to retrieve pedestrian images captured by both visible and infrared cameras, is a challenging but essential task for smart surveillance systems. The huge barrier between visible and infrared images has led to the large cross-modality discrepancy and intraclass variations. Most existing VI-ReID methods tend to learn discriminative modality-sharable features based on either global or part-based representations, lacking effective optimization objectives. In this article, we propose a novel global-local multichannel (GLMC) network for VI-ReID, which can learn multigranularity representations based on both global and local features. The coarse- and fine-grained information can complement each other to form a more discriminative feature descriptor. Besides, we also propose a novel center loss function that aims to simultaneously improve the intraclass cross-modality similarity and enlarge the interclass discrepancy to explicitly handle the cross-modality discrepancy issue and avoid the model fluctuating problem. Experimental results on two public datasets have demonstrated the superiority of the proposed method compared with state-of-the-art approaches in terms of effectiveness.
Liyan Zhang 0001, Guodong Du 0005, Fan Liu 0003, Huawei Tu, Xiangbo Shu
IEEE Trans. Neural Networks Learn. Syst.5
2025 Triplet Contrastive Representation Learning for Unsupervised Vehicle Re-Identification
abstract
Part feature learning plays a crucial role in achieving fine-grained semantic understanding in unsupervised vehicle re-identification. However, existing approaches directly model part and global features, which can easily lead to severe gradient vanishing issues due to their unequal feature information and unreliable pseudo-labels. To address this problem, in this article, we propose a triplet contrastive representation learning (TCRL) framework, which leverages cluster features to bridge the part features and global features for unsupervised vehicle re-identification. Specifically, TCRL devises three memory banks to store the instance/cluster features and proposes a proxy contrastive loss (PCL) to make contrastive learning between adjacent memory banks, thus presenting the associations between the part and global features as a transition of the part-cluster and cluster-global associations. Since the cluster memory bank copes with all the vehicle features, it can summarize them into a discriminative feature representation. To deeply exploit the instance/cluster information, TCRL proposes two additional loss functions. For the instance-level feature, a hybrid contrastive loss (HCL) re-defines the sample correlations by approaching the positive instance features and pushing all negative instance features away. For the cluster-level feature, a weighted regularization cluster contrastive loss (WRCCL) refines the pseudo labels by penalizing the mislabeled images according to the instance similarity. Extensive experiments show that TCRL outperforms many state-of-the-art unsupervised vehicle re-identification approaches.
Fei Shen 0004, Xiaoyu Du 0002, Liyan Zhang 0002, Xiangbo Shu, Jinhui Tang 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Leveraging Frame- and Feature-level Progressive Augmentation for Semi-supervised Action Recognition
abstract
Semi-supervised action recognition is a challenging yet prospective task due to its low reliance on costly labeled videos. One high-profile solution is to explore frame-level weak/strong augmentations for learning abundant representations, inspired by the FixMatch framework dominating the semi-supervised image classification task. However, such a solution mainly brings perturbations in terms of texture and scale, leading to the limitation in learning action representations in videos with spatiotemporal redundancy and complexity. Therefore, we revisit the creative trick of weak/strong augmentations in FixMatch and then propose, to the best of our knowledge, a novel Frame- and Feature-level augmentation FixMatch (dubbed as F 2 -FixMatch) framework to learn more abundant action representations for being robust to complex and dynamic video scenarios. Specifically, we design a new Progressive Augmentation mechanism that implements the weak/strong augmentations first at the frame level, and further implements the perturbation at the feature level, to obtain abundant four types of augmented features in broader perturbation spaces. Moreover, we present an evolved Multihead Pseudo-Labeling scheme to promote the consistency of features across different augmented versions based on the pseudo labels. We conduct extensive experiments on several public datasets to demonstrate that our F 2 -FixMatch achieves the performance gain compared with current state-of-the-art methods. The source codes of F 2 -FixMatch are publicly available at https://github.com/zwtu/F2FixMatch .
Zhewei Tu, Xiangbo Shu, Rui Yan 0010, Zhenxing Liu 0001, Jiachao Zhang
ACM Trans. Multim. Comput. Commun. Appl.2
2024 DTS-TPT: Dual Temporal-Sync Test-time Prompt Tuning for Zero-shot Activity Recognition
Rui Yan 0010, Hongyu Qu, Xiangbo Shu, Jinhui Tang 0001, Tieniu Tan
IJCAI3
2024 AdaFPP: Adapt-Focused Bi-Propagating Prototype Learning for Panoramic Activity Recognition
abstract
Panoramic Activity Recognition (PAR) aims to identify multi-granul-arity behaviors performed by multiple persons in panoramic scenes, including individual activities, group activities, and global activities. Previous methods 1) heavily rely on manually annotated detection boxes in training and inference, hindering further practical deployment; or 2) directly employ normal detectors to detect multiple persons with varying size and spatial occlusion in panoramic scenes, blocking the performance gain of PAR. To this end, we consider learning a detector adapting varying-size occluded persons, which is optimized along with the recognition module in the all-in-one framework. Therefore, we propose a novel Adapt-Focused bi-Propagating Prototype learning (AdaFPP) framework to jointly recognize individual, group, and global activities in panoramic activity scenes by learning an adapt-focused detector and multi-granularity prototypes as the pretext tasks in an end-to-end way. Specifically, to accommodate the varying sizes and spatial occlusion of multiple persons in crowed panoramic scenes, we introduce a panoramic adapt-focuser, achieving the size-adapting detection of individuals by comprehensively selecting and performing fine-grained detections on object-dense sub-regions identified through original detections. In addition, to mitigate information loss due to inaccurate individual localizations, we introduce a bi-propagation prototyper that promotes closed-loop interaction and informative consistency across different granularities by facilitating bidirectional information propagation among the individual, group, and global levels. Extensive experiments demonstrate the significant performance of AdaFPP and emphasize its powerful applicability for PAR.
Meiqi Cao, Rui Yan 0010, Xiangbo Shu, Guangzhao Dai, Yazhou Yao, Guosen Xie
ACM Multimedia3
2024 DoFIT: Domain-aware Federated Instruction Tuning with Alleviated Catastrophic Forgetting
abstract
Federated Instruction Tuning (FIT) advances collaborative training on decentralized data, crucially enhancing model's capability and safeguarding data privacy. However, existing FIT methods are dedicated to handling data heterogeneity across different clients (i.e., client-aware data heterogeneity), while ignoring the variation between data from different domains (i.e., domain-aware data heterogeneity). When scarce data needs supplementation from related fields, these methods lack the ability to handle domain heterogeneity in cross-domain training. This leads to domain-information catastrophic forgetting in collaborative training and therefore makes model perform sub-optimally on the individual domain. To address this issue, we introduce DoFIT, a new Domain-aware FIT framework that alleviates catastrophic forgetting through two new designs. First, to reduce interference information from the other domain, DoFIT finely aggregates overlapping weights across domains on the inter-domain server side. Second, to retain more domain information, DoFIT initializes intra-domain weights by incorporating inter-domain information into a less-conflicted parameter space. Experimental results on diverse datasets consistently demonstrate that DoFIT excels in cross-domain collaborative training and exhibits significant advantages over conventional FIT methods in alleviating catastrophic forgetting. Code is available at [this link](https://github.com/1xbq1/DoFIT).
Binqian Xu, Xiangbo Shu, Haiyang Mei, Zechen Bai, Basura Fernando, Zheng Shou 0001, Jinhui Tang 0001
NeurIPS2
2024 Rethinking attribute localization for zero-shot learning
Shuhuang Chen, Shiming Chen 0002, Guosen Xie, Xiangbo Shu, Xinge You, Xuelong Li 0001
Sci. China Inf. Sci.4
2024 Dilation-erosion for single-frame supervised temporal action localization
Yan Song 0005, Fanming Wang, Yang Zhao 0035, Xiangbo Shu, Yan Rui
Multim. Tools Appl.5
2024 Motion-Aware Mask Feature Reconstruction for Skeleton-Based Action Recognition
abstract
Despite recent advancements in masked skeleton modeling and visual-language pre-training, no method has yet been proposed to explore capturing and utilizing the rich semantic information embedded in both modalities for enhanced action recognition. To address this challenge, we propose a novel Motion-Aware Mask Feature Reconstruction (MMFR) method for the challenging task of skeleton-based action recognition. MMFR ingeniously integrates masked skeleton feature reconstruction with visual-language pre-trained model within a consolidated framework, aiming to leverage the synergistic potential of both domains. Specifically, It employs visual-language model to infuse semantic understanding into the skeleton feature reconstruction process via probability distribution distillation. Moreover, we introduce a multi-granularity semantic contrast module that refines vision-text alignment precision and augments contextual information for accurate mask reconstruction. Extensive experiments demonstrate MMFR’s superiority in skeleton-based action recognition, as well as its efficacy in zero-shot scenarios.
Xingyu Zhu 0008, Xiangbo Shu, Jinhui Tang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Semantic-Disentangled Transformer With Noun-Verb Embedding for Compositional Action Recognition
abstract
Recognizing actions performed on unseen objects, known as Compositional Action Recognition (CAR), has attracted increasing attention in recent years. The main challenge is to overcome the distribution shift of "action-objects" pairs between the training and testing sets. Previous works for CAR usually introduce extra information (e.g. bounding box) to enhance the dynamic cues of video features. However, these approaches do not essentially eliminate the inherent inductive bias in the video, which can be regarded as the stumbling block for model generalization. Because the video features are usually extracted from the visually cluttered areas in which many objects cannot be removed or masked explicitly. To this end, this work attempts to implicitly accomplish semantic-level decoupling of "object-action" in the high-level feature space. Specifically, we propose a novel Semantic-Decoupling Transformer framework, dubbed as DeFormer, which contains two insightful sub-modules: Objects-Motion Decoupler (OMD) and Semantic-Decoupling Constrainer (SDC). In OMD, we initialize several learnable tokens incorporating annotation priors to learn an instance-level representation and then decouple it into the appearance feature and motion feature in high-level visual space. In SDC, we use textual information in the high-level language space to construct a dual-contrastive association to constrain the decoupled appearance feature and motion feature obtained in OMD. Extensive experiments verify the generalization ability of DeFormer. Specifically, compared to the baseline method, DeFormer achieves absolute improvements of 3%, 3.3%, and 5.4% under three different settings on STH-ELSE, while corresponding improvements on EPIC-KITCHENS-55 are 4.7%, 9.2%, and 4.4%. Besides, DeFormer gains state-of-the-art results either on ground-truth or detected annotations.
Rui Yan 0010, Xiangbo Shu, Zhewei Tu, Guangzhao Dai, Jinhui Tang 0001
IEEE Trans. Image Process.3
2024 Spatiotemporal Decouple-and-Squeeze Contrastive Learning for Semisupervised Skeleton-Based Action Recognition
abstract
Contrastive learning has been successfully leveraged to learn action representations for addressing the problem of semisupervised skeleton-based action recognition. However, most contrastive learning-based methods only contrast global features mixing spatiotemporal information, which confuses the spatial- and temporal-specific information reflecting different semantic at the frame level and joint level. Thus, we propose a novel spatiotemporal decouple-and-squeeze contrastive learning (SDS-CL) framework to comprehensively learn more abundant representations of skeleton-based actions by jointly contrasting spatial-squeezing features, temporal-squeezing features, and global features. In SDS-CL, we design a new spatiotemporal-decoupling intra-inter attention (SIIA) mechanism to obtain the spatiotemporal-decoupling attentive features for capturing spatiotemporal specific information by calculating spatial- and temporal-decoupling intra-attention maps among joint/motion features, as well as spatial- and temporal-decoupling inter-attention maps between joint and motion features. Moreover, we present a new spatial-squeezing temporal-contrasting loss (STL), a new temporal-squeezing spatial-contrasting loss (TSL), and the global-contrasting loss (GL) to contrast the spatial-squeezing joint and motion features at the frame level, temporal-squeezing joint and motion features at the joint level, as well as global joint and motion features at the skeleton level. Extensive experimental results on four public datasets show that the proposed SDS-CL achieves performance gains compared with other competitive methods.
Binqian Xu, Xiangbo Shu, Jiachao Zhang, Guangzhao Dai, Yan Song 0005
IEEE Trans. Neural Networks Learn. Syst.2
2023 UATVR: Uncertainty-Adaptive Text-Video Retrieval
abstract
With the explosive growth of web videos and emerging large-scale vision-language pre-training models, e.g., CLIP, retrieving videos of interest with text instructions has attracted increasing attention. A common practice is to transfer text-video pairs to the same embedding space and craft cross-modal interactions with certain entities in specific granularities for semantic correspondence. Unfortunately, the intrinsic uncertainties of optimal entity combinations in appropriate granularities for cross-modal queries are understudied, which is especially critical for modalities with hierarchical semantics, e.g., video, text, etc. In this paper, we propose an Uncertainty-Adaptive Text-Video Retrieval approach, termed UATVR, which models each lookup as a distribution matching procedure. Concretely, we add additional learnable tokens in the encoders to adaptively aggregate multi-grained semantics for flexible high-level reasoning. In the refined embedding space, we represent text-video pairs as probabilistic distributions where prototypes are sampled for matching evaluation. Comprehensive experiments on four benchmarks justify the superiority of our UATVR, which achieves new state-of-the-art results on MSR-VTT (50.8%), VATEX (64.5%), MSVD (49.7%), and DiDeMo (45.8%). The code is available at https://github.com/bofang98/UATVR.
Bo Fang 0003, Yu Zhou 0015, YuXin Song 0001, Weiping Wang 0005, Xiangbo Shu, Xiangyang Ji, Jingdong Wang 0001
ICCV7
2023 MUP: Multi-granularity Unified Perception for Panoramic Activity Recognition
abstract
Panoramic activity recognition is required to jointly identify multi-granularity human behaviors including individual actions, group activities, and global activities in multi-person videos. Previous methods encode these behaviors hierarchically through multiple stages, which disturb the inherent co-occurrence across multi-granularity behaviors in the same scene. To this end, we propose a novel Multi-granularity Unified Perception (MUP) framework that perceives different granularity behaviors universally to explore the co-occurrence motion pattern via the same parameters in an end-to-end fashion. To be specific, the proposed framework stacks three Unified Motion Encoding (UME) blocks for modeling multiple granularity behaviors with shared parameters. UME block mines intra-relevant and cross-relevant semantics synchronously from input feature sequences via Intra-granularity Motion Embedding (IME) and Cross-granularity Motion Prototyping (CMP). In particular, IME aims to model the interactions among visual features within each granularity based on the attention mechanism. CMP aims to aggregate features across different granularities (i.e., person to group) via several learnable prototypes. Extensive experiments demonstrate that MUP outperforms the state-of-the-art methods on JRDB-PAR and has satisfactory interpretability.
Meiqi Cao, Rui Yan 0010, Xiangbo Shu, Jiachao Zhang, Jinpeng Wang 0001, Guosen Xie
ACM Multimedia3
2023 Foreground/Background-Masked Interaction Learning for Spatio-temporal Action Detection
abstract
Spatio-temporal Action Detection (SAD) aims to recognize the multi-class actions, and meanwhile locate their spatio-temporal occurrence in untrimmed videos. Besides relying on the inherent inter-actor interactions, most previous SAD approaches model actor interactions between multi-actors and the whole frames or special parts (e.g., objects/hands). However, such approaches are relatively graceless by 1) roughly treating all various actors to equivalently interact with frames/parts or by 2) sumptuously borrowing multiple costly detectors to acquire the special parts. To solve the above dilemma, we propose a novel Foreground/Background-masked Interaction Learning (dubbed as FBI Learning) framework to learn the multi-actor features by attentively interacting with the hands-down foreground and background frames. Specifically, we first design a new Mask-guided Cross Attention (MCA) mechanism that calculates the masked cross-attentions to capture the compact relations between the actors and foreground/background regions. Next, we present a new Actor-guided Feature Aggregation (AFA) scheme that integrates foreground- and background-interacted actor features with the learnable actor-based weights. Finally, we construct a long-term feature bank that associates temporal context information to facilitate action classification. Extensive experiments are conducted on commonly available UCF101-24, MultiSports, and AVA v2.1/v2.2 datasets, which illustrate the competitive performance of FBI Learning against the state-of-the-art methods.
Keke Chen, Xiangbo Shu, Guosen Xie, Rui Yan 0010, Jinhui Tang 0001
ACM Multimedia2
2023 Slowfast Diversity-aware Prototype Learning for Egocentric Action Recognition
abstract
Egocentric Action Recognition (EAR) is required to recognize both the interacting objects (noun) and the motion (verb) against cluttered backgrounds with distracting objects. For capturing interacting objects, traditional approaches heavily rely on luxury object annotations or detectors, though a few works heuristically enumerate the fixed sets of verb-constrained prototypes to roughly exclude the background. For capturing motion, the inherent variations of motion duration among egocentric videos with different lengths are almost ignored. To this end, we propose a novel Slowfast Diversity-aware Prototype learning (SDP) to effectively capture interacting objects by learning compact yet diverse prototypes, and adaptively capture motion in either long-time video or short-time video. Specifically, we present a new Part-to-Prototype (P2P) scheme to learn prototypes from raw videos covering the interacting objects by refining the semantic information from part level to prototype level. Moreover, for adaptively capturing motion, we design a new Slow-Fast Context (SFC) mechanism that explores the Up/Down augmentations for the prototype representation at the semantic level to strengthen the transient dynamic information in short-time videos and eliminate the redundant dynamic information in long-time videos, which are further fine-complemented via the slow- and fast-aware attentions. Extensive experiments demonstrate SDP outperforms state-of-the-art methods on two large-scale egocentric video benchmarks, i.e., EPIC-KITCHENS-100 and EGTEA.
Guangzhao Dai, Xiangbo Shu, Rui Yan 0010, Jinhui Tang 0001
ACM Multimedia2
2023 Pedestrian-specific Bipartite-aware Similarity Learning for Text-based Person Retrieval
abstract
Text-based person retrieval is a challenging task that aims to search pedestrian images with the same identity according to language descriptions. Current methods usually indiscriminately measure the similarity between text and image by matching global visual-textual features and matched local region-word features. However, these methods underestimate the key cue role of mismatched region-word pairs and ignore the problem of low similarity between matched region-word pairs. To alleviate these issues, we propose a novel Pedestrian-specific Bipartite-aware Similarity Learning (PBSL) framework that efficiently reveals the plausible and credible levels of contribution of pedestrian-specific mismatched and matched region-word pairs towards overall similarity. Specifically, to focus on mismatched region-word pairs, we first develop a new co-interactive attention that utilizes cross-modal information to guide the extraction of pedestrian-specific information in a single modality. We then design a negative similarity regularization mechanism to use the negative similarity score as a bias to correct the overall similarity. Additionally, to enhance the contribution of matched region-word pairs, we introduce graph networks to aggregate and propagate local information of pedestrian-specific, using overall visual-textual similarity to evaluate locally matched region-word pairs for weight refinement. Finally, extensive experiments are conducted on the CUHK-PEDES, ICFG-PEDES, and RSTPReid datasets to demonstrate the competitive performance of the proposed PBSL in the text-based person retrieval task.
Fei Shen 0004, Xiangbo Shu, Xiaoyu Du 0002, Jinhui Tang 0001
ACM Multimedia2
2023 BCMask: a finer leaf instance segmentation with bilayer convolution mask
Xingjian Gu, Yongjie Zhu, Shougang Ren, Xiangbo Shu
Multim. Syst.4
2023 Supervised Learning Strategy for Spiking Neurons Based on Their Segmental Running Characteristics
Xingjian Gu, Xiangbo Shu
Neural Process. Lett.6
2023 Multi-Granularity Anchor-Contrastive Representation Learning for Semi-Supervised Skeleton-Based Action Recognition
abstract
In the semi-supervised skeleton-based action recognition task, obtaining more discriminative information from both labeled and unlabeled data is a challenging problem. As the current mainstream approach, contrastive learning can learn more representations of augmented data, which can be considered as the pretext task of action recognition. However, such a method still confronts three main limitations: 1) It usually learns global-granularity features that cannot well reflect the local motion information. 2) The positive/negative pairs are usually pre-defined, some of which are ambiguous. 3) It generally measures the distance between positive/negative pairs only within the same granularity, which neglects the contrasting between the cross-granularity positive and negative pairs. Toward these limitations, we propose a novel Multi-granularity Anchor-Contrastive representation Learning (dubbed as MAC-Learning) to learn multi-granularity representations by conducting inter- and intra-granularity contrastive pretext tasks on the learnable and structural-link skeletons among three types of granularities covering local, context, and global views. To avoid the disturbance of ambiguous pairs from noisy and outlier samples, we design a more reliable Multi-granularity Anchor-Contrastive Loss (dubbed as MAC-Loss) that measures the agreement/disagreement between high-confidence soft-positive/negative pairs based on the anchor graph instead of the hard-positive/negative pairs in the conventional contrastive loss. Extensive experiments on both NTU RGB+D and Northwestern-UCLA datasets show that the proposed MAC-Learning outperforms existing competitive methods in semi-supervised skeleton-based action recognition tasks.
Xiangbo Shu, Binqian Xu, Liyan Zhang 0001, Jinhui Tang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Progressive Instance-Aware Feature Learning for Compositional Action Recognition
abstract
In order to enable the model to generalize to unseen "action-objects" (compositional action), previous methods encode multiple pieces of information (i.e., the appearance, position, and identity of visual instances) independently and concatenate them for classification. However, these methods ignore the potential supervisory role of instance information (i.e., position and identity) in the process of visual perception. To this end, we present a novel framework, namely Progressive Instance-aware Feature Learning (PIFL), to progressively extract, reason, and predict dynamic cues of moving instances from videos for compositional action recognition. Specifically, this framework extracts features from foreground instances that are likely to be relevant to human actions (Position-aware Appearance Feature Extraction in Section III-B1), performs identity-aware reasoning among instance-centric features with semantic-specific interactions (Identity-aware Feature Interaction in Section III-B2), and finally predicts instances' position from observed states to force the model into perceiving their movement (Semantic-aware Position Prediction in Section III-B3). We evaluate our approach on two compositional action recognition benchmarks, namely, Something-Else and IKEA-Assembly. Our approach achieves consistent accuracy gain beyond off-the-shelf action recognition algorithms in terms of both ground truth and detected position of instances.
Rui Yan 0010, Lingxi Xie, Xiangbo Shu, Liyan Zhang 0001, Jinhui Tang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 HiGCIN: Hierarchical Graph-Based Cross Inference Network for Group Activity Recognition
abstract
Group activity recognition (GAR) is a challenging task aimed at recognizing the behavior of a group of people. It is a complex inference process in which visual cues collected from individuals are integrated into the final prediction, being aware of the interaction between them. This paper goes one step further beyond the existing approaches by designing a Hierarchical Graph-based Cross Inference Network (HiGCIN), in which three levels of information, i.e., the body-region level, person level, and group-activity level, are constructed, learned, and inferred in an end-to-end manner. Primarily, we present a generic Cross Inference Block (CIB), which is able to concurrently capture the latent spatiotemporal dependencies among body regions and persons. Based on the CIB, two modules are designed to extract and refine features for group activities at each level. Experiments on two popular benchmarks verify the effectiveness of our approach, particularly in the ability to infer with multilevel visual cues. In addition, training our approach does not require individual action labels to be provided, which greatly reduces the amount of labor required in data annotation.
Rui Yan 0010, Lingxi Xie, Jinhui Tang 0001, Xiangbo Shu, Qi Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Com-STAL: Compositional Spatio-Temporal Action Localization
abstract
Spatio-temporal action localization aims to locate the spatial and temporal positions of actors and classify their actions. However, prior research overlooks the fact that human actions often interact with novel objects in real-world scenarios, which neglects the various combinations of action-object, and considerably limits the generalization of the developed models. In this paper, we study the action-object combinations by researching multi-modal vision information of them. To this end, we propose a novel compositional spatio-temporal action localization (Com-STAL) task, which features non-overlapping action-object combinations in their training and test sets. Based on this, we construct a compositional action localization dataset (Com-AD). Beyond that, we propose a simple yet effective framework, Instance-Centric Interaction Network (ICIN), to reduce invalid induction biases within the visual modality and alleviate the combined distribution bias issue by leveraging additional modal information. The extensive experiment results on Com-AD demonstrate superior action localization performance of ICIN.
Shaomeng Wang, Rui Yan 0010, Guangzhao Dai, Yan Song 0005, Xiangbo Shu
IEEE Trans. Circuits Syst. Video Technol.6
2023 A simple yet effective image stitching with computational suture zone
Jiachao Zhang, Yunbin Huang, Yanming Yu, Xiangbo Shu
Vis. Comput.6
2022 PNP: Robust Learning from Noisy Labels by Probabilistic Noise Prediction
abstract
Label noise has been a practical challenge in deep learning due to the strong capability of deep neural networks in fitting all training data. Prior literature primarily resorts to sample selection methods for combating noisy labels. However, these approaches focus on dividing samples by order sorting or threshold selection, inevitably introducing hyperparameters (e.g., selection ratio / threshold) that are hard-to-tune and dataset-dependent. To this end, we propose a simple yet effective approach named PNP (Probabilistic Noise Prediction) to explicitly model label noise. Specifically, we simultaneously train two networks, in which one predicts the category label and the other predicts the noise type. By predicting label noise probabilistically, we identify noisy samples and adopt dedicated optimization objectives accordingly. Finally, we establish a joint loss for network update by unifying the classification loss, the auxiliary constraint loss, and the in-distribution consistency loss. Comprehensive experimental results on synthetic and realworld datasets demonstrate the superiority of our proposed method. The source code and models have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/PNP.
Zeren Sun, Fumin Shen, Qiong Wang 0003, Xiangbo Shu, Yazhou Yao, Jinhui Tang 0001
CVPR5
2022 Look Less Think More: Rethinking Compositional Action Recognition
abstract
Compositional action recognition which aims to identify the unseen combinations of actions and objects has recently attracted wide attention. Conventional methods bring in additional cues (e.g., dynamic motions of objects) to alleviate the inductive bias between the visual appearance of objects and the human action-level labels. Besides, compared with non-compositional settings, previous methods only pursue higher performance in compositional settings, which can not prove their generalization ability. To this end, we firstly rethink the problem and design a more generalized metric (namely Drop Ratio) and a more practical setting to evaluate the compositional generalization of existing action recognition algorithms. Beyond that, we propose a simple yet effective framework, Look Less Think More (LLTM), to reduce the strong association between visual objects and action-level labels (Look Less), and then discover the commonsense relationships between object categories and human actions (Think More). We test the rationality of the proposed Drop Ratio and Practical setting by comparing several popular action recognition methods on SSV2. Besides, the proposed LLTM achieves state-of-the-art performance on SSV2 with different settings.
Rui Yan 0010, Xiangbo Shu, Junhao Zhang 0001, Yonghua Pan, Jinhui Tang 0001
ACM Multimedia3
2022 Spatiotemporal Perturbation Based Dynamic Consistency for Semi-supervised Temporal Action Detection
Yan Song 0005, Rui Yan 0010, Xiangbo Shu
MMM (1)4
2022 Progressive enhancement network with pseudo labels for weakly supervised temporal action localization
Yan Song 0005, Xiangbo Shu
J. Vis. Commun. Image Represent.4
2022 Skip-attention encoder-decoder framework for human motion prediction
Xiangbo Shu, Rui Yan 0010, Jiachao Zhang, Yan Song 0005
Multim. Syst.2
2022 Wavelet-Attention CNN for image classification
Xiangbo Shu
Multim. Syst.3
2022 Spatiotemporal Co-Attention Recurrent Neural Networks for Human-Skeleton Motion Prediction
abstract
Human motion prediction aims to generate future motions based on the observed human motions. Witnessing the success of Recurrent Neural Networks (RNN) in modeling sequential data, recent works utilize RNNs to model human-skeleton motions on the observed motion sequence and predict future human motions. However, these methods disregard the existence of the spatial coherence among joints and the temporal evolution among skeletons, which reflects the crucial characteristics of human motions in spatiotemporal space. To this end, we propose a novel Skeleton-Joint Co-Attention Recurrent Neural Networks (SC-RNN) to capture the spatial coherence among joints, and the temporal evolution among skeletons simultaneously on a skeleton-joint co-attention feature map in spatiotemporal space. First, a skeleton-joint feature map is constructed as the representation of the observed motion sequence. Second, we design a new Skeleton-Joint Co-Attention (SCA) mechanism to dynamically learn a skeleton-joint co-attention feature map of this skeleton-joint feature map, which can refine the useful observed motion information to predict one future motion. Third, a variant of GRU embedded with SCA collaboratively models the human-skeleton motion and human-joint motion in spatiotemporal space by regarding the skeleton-joint co-attention feature map as the motion context. Experimental results of human motion prediction demonstrate that the proposed method outperforms the competing methods.
Xiangbo Shu, Liyan Zhang 0001, Guo-Jun Qi, Wei Liu 0005, Jinhui Tang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Coherence Constrained Graph LSTM for Group Activity Recognition
abstract
This work aims to address the group activity recognition problem by exploring human motion characteristics. Traditional methods hold that the motions of all persons contribute equally to the group activity, which suppresses the contributions of some relevant motions to the whole activity while overstating some irrelevant motions. To address this problem, we present a Spatio-Temporal Context Coherence (STCC) constraint and a Global Context Coherence (GCC) constraint to capture the relevant motions and quantify their contributions to the group activity, respectively. Based on this, we propose a novel Coherence Constrained Graph LSTM (CCG-LSTM) with STCC and GCC to effectively recognize group activity, by modeling the relevant motions of individuals while suppressing the irrelevant motions. Specifically, to capture the relevant motions, we build the CCG-LSTM with a temporal confidence gate and a spatial confidence gate to control the memory state updating in terms of the temporally previous state and the spatially neighboring states, respectively. In addition, an attention mechanism is employed to quantify the contribution of a certain motion by measuring the consistency between itself and the whole activity at each time step. Finally, we conduct experiments on two widely-used datasets to illustrate the effectiveness of the proposed CCG-LSTM compared with the state-of-the-art methods.
Jinhui Tang 0001, Xiangbo Shu, Rui Yan 0010, Liyan Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Expansion-Squeeze-Excitation Fusion Network for Elderly Activity Recognition
abstract
This work focuses on the task of elderly activity recognition, which is a challenging task due to the existence of individual actions and human-object interactions in elderly activities. Thus, we attempt to effectively aggregate the discriminative information of actions and interactions from both RGB videos and skeleton sequences by attentively fusing multi-modal features. Recently, some nonlinear multi-modal fusion approaches are proposed by utilizing nonlinear attention mechanism that is extended from Squeeze-and-Excitation Networks (SENet). Inspired by this, we propose a novel Expansion-Squeeze-Excitation Fusion Network (ESE-FN) to effectively address the problem of elderly activity recognition, which learns modal and channel-wise Expansion-Squeeze-Excitation (ESE) attentions for attentively fusing the multi-modal features in the modal and channel-wise ways. Specifically, ESE-FN firstly implements the modal-wise fusion with the Modal-wise ESE Attention (M-ESEA) to aggregate discriminative information in modal-wise way, and then implements the channel-wise fusion with the Channel-wise ESE Attention (C-ESEA) to aggregate the multi-channel discriminative information in channel-wise way (referring toFigure 1). Furthermore, we design a new Multi-modal Loss (ML) to keep the consistency between the single-modal features and the fused multi-modal features by adding the penalty of difference between the minimum prediction losses on single modalities and the prediction loss on the fused modality. Finally, we conduct experiments on a largest-scale elderly activity dataset, i.e., ETRI-Activity3D (including 110,000+ videos, and 50+ categories), to demonstrate that the proposed ESE-FN achieves the best accuracy compared with the state-of-the-art methods. In addition, more extensive experimental results show that the proposed ESE-FN is also comparable to the other methods in terms of normal action recognition task.
Xiangbo Shu, Jiawen Yang, Rui Yan 0010, Yan Song 0005
IEEE Trans. Circuits Syst. Video Technol.1
2022 X-Invariant Contrastive Augmentation and Representation Learning for Semi-Supervised Skeleton-Based Action Recognition
abstract
Semi-supervised skeleton-based action recognition is a challenging problem due to insufficient labeled data. For addressing this problem, some representative methods leverage contrastive learning to obtain more features from the pre-augmented skeleton actions. Such methods usually adopt a two-stage way: first randomly augment samples, and then learn their representations via contrastive learning. Since skeleton samples have already been randomly augmented, the representation ability of the subsequent contrastive learning is limited due to the inconsistency between the augmentations and representations. Thus, we propose a novel X-invariant Contrastive Augmentation and Representation learning (X-CAR) framework to thoroughly obtain rotate-shear-scale (X for short) invariant features by learning augmentations and representations of skeleton sequences in a one-stage way. In X-CAR, a new Adaptive-combination Augmentation (AA) mechanism is designed to rotate, shear, and scale the skeletons by learnable controlling factors in an adaptive way rather than a random way. Here, such controlling factors are also learned in the whole contrastive learning process, which can facilitate the consistency between the learned augmentations and representations of skeleton sequences. In addition, we relax the pre-definition of positive and negative samples to avoid the confusing allocation of ambiguous samples, and present a new Pull-Push Contrastive Loss (PPCL) to pull the augmenting skeleton close to the original skeleton, while push far away from the other skeletons. Experimental results on both NTU RGB+D and North-Western UCLA datasets show that the proposed X-CAR achieves better accuracy compared with other competitive methods in the semi-supervised scenario.
Binqian Xu, Xiangbo Shu, Yan Song 0005
IEEE Trans. Image Process.2
2022 Position-Aware Participation-Contributed Temporal Dynamic Model for Group Activity Recognition
abstract
Group activity recognition (GAR) aiming at understanding the behavior of a group of people in a video clip has received increasing attention recently. Nevertheless, most of the existing solutions ignore that not all the persons contribute to the group activity of the scene equally. That is to say, the contribution from different individual behaviors to group activity is different; meanwhile, the contribution from people with different spatial positions is also different. To this end, we propose a novel Position-aware Participation-Contributed Temporal Dynamic Model (P2CTDM), in which two types of the key actor are constructed and learned. Specifically, we focus on the behaviors of key actors, who maintain steady motions (long moving time, called long motions) or display remarkable motions (but closely related to other people and the group activity, called flash motions) at a certain moment. For capturing long motions, we rank individual motions according to their intensity measured by stacking optical flows. For capturing flash motions that are closely related to other people, we design a position-aware interaction module (PIM) that simultaneously considers the feature similarity and position information. Beyond that, for capturing flash motions that are highly related to the group activity, we also present an aggregation long short-term memory (Agg-LSTM) to fuse the outputs from PIM by time-varying trainable attention factors. Four widely used benchmarks are adopted to evaluate the performance of the proposed P2CTDM compared to the state of the art.
Rui Yan 0010, Xiangbo Shu, Chengcheng Yuan, Qi Tian 0001, Jinhui Tang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2021 Attention-aware conditional generative adversarial networks for facial age synthesis
Xiahui Chen, Yunlian Sun, Xiangbo Shu
Neurocomputing3
2021 Hierarchical Long Short-Term Concurrent Memory for Human Interaction Recognition
abstract
In this work, we aim to address the problem of human interaction recognition in videos by exploring the long-term inter-related dynamics among multiple persons. Recently, Long Short-Term Memory (LSTM) has become a popular choice to model individual dynamic for single-person action recognition due to its ability to capture the temporal motion information in a range. However, most existing LSTM-based methods focus only on capturing the dynamics of human interaction by simply combining all dynamics of individuals or modeling them as a whole. Such methods neglect the inter-related dynamics of how human interactions change over time. To this end, we propose a novel Hierarchical Long Short-Term Concurrent Memory (H-LSTCM) to model the long-term inter-related dynamics among a group of persons for recognizing human interactions. Specifically, we first feed each person's static features into a Single-Person LSTM to model the single-person dynamic. Subsequently, at one time step, the outputs of all Single-Person LSTM units are fed into a novel Concurrent LSTM (Co-LSTM) unit, which mainly consists of multiple sub-memory units, a new cell gate, and a new co-memory cell. In the Co-LSTM unit, each sub-memory unit stores individual motion information, while this Co-LSTM unit selectively integrates and stores inter-related motion information between multiple interacting persons from multiple sub-memory units via the cell gate and co-memory cell, respectively. Extensive experiments on several public datasets validate the effectiveness of the proposed H-LSTCM by comparing against baseline and state-of-the-art methods.
Xiangbo Shu, Jinhui Tang 0001, Guo-Jun Qi, Wei Liu 0005, Jian Yang 0003
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Host-Parasite: Graph LSTM-in-LSTM for Group Activity Recognition
abstract
This article aims to tackle the problem of group activity recognition in the multiple-person scene. To model the group activity with multiple persons, most long short-term memory (LSTM)-based methods first learn the person-level action representations by several LSTMs and then integrate all the person-level action representations into the following LSTM to learn the group-level activity representation. This type of solution is a two-stage strategy, which neglects the "host-parasite" relationship between the group-level activity ("host") and person-level actions ("parasite") in spatiotemporal space. To this end, we propose a novel graph LSTM-in-LSTM (GLIL) for group activity recognition by modeling the person-level actions and the group-level activity simultaneously. GLIL is a "host-parasite" architecture, which can be seen as several person LSTMs (P-LSTMs) in the local view or a graph LSTM (G-LSTM) in the global view. Specifically, P-LSTMs model the person-level actions based on the interactions among persons. Meanwhile, G-LSTM models the group-level activity, where the person-level motion information in multiple P-LSTMs is selectively integrated and stored into G-LSTM based on their contributions to the inference of the group activity class. Furthermore, to use the person-level temporal features instead of the person-level static features as the input of GLIL, we introduce a residual LSTM with the residual connection to learn the person-level residual features, consisting of temporal features and static features. Experimental results on two public data sets illustrate the effectiveness of the proposed GLIL compared with state-of-the-art methods.
Xiangbo Shu, Liyan Zhang 0001, Yunlian Sun, Jinhui Tang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2020 Web-Supervised Network with Softly Update-Drop Training for Fine-Grained Visual Classification
abstract
Labeling objects at the subordinate level typically requires expert knowledge, which is not always available from a random annotator. Accordingly, learning directly from web images for fine-grained visual classification (FGVC) has attracted broad attention. However, the existence of noise in web images is a huge obstacle for training robust deep neural networks. In this paper, we propose a novel approach to remove irrelevant samples from the real-world web images during training, and only utilize useful images for updating the networks. Thus, our network can alleviate the harmful effects caused by irrelevant noisy web images to achieve better performance. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is much superior to state-of-the-art webly supervised methods. The data and source code of this work have been made anonymously available at: https://github.com/z337-408/WSNFGVC.
Chuanyi Zhang, Yazhou Yao, Huafeng Liu 0004, Guosen Xie, Xiangbo Shu, Tianfei Zhou, Zheng Zhang 0006, Fumin Shen, Zhenmin Tang
AAAI5
2020 Social Adaptive Module for Weakly-Supervised Group Activity Recognition
Rui Yan 0010, Lingxi Xie, Jinhui Tang 0001, Xiangbo Shu, Qi Tian 0001
ECCV (8)4
2020 Data-driven Meta-set Based Fine-Grained Visual Recognition
abstract
Constructing fine-grained image datasets typically requires domain-specific expert knowledge, which is not always available for crowd-sourcing platform annotators. Accordingly, learning directly from web images becomes an alternative method for fine-grained visual recognition. However, label noise in the web training set can severely degrade the model performance. To this end, we propose a data-driven meta-set based approach to deal with noisy web images for fine-grained recognition. Specifically, guided by a small amount of clean meta-set, we train a selection net in a meta-learning manner to distinguish in- and out-of-distribution noisy images. To further boost the robustness of the model, we also learn a labeling net to correct the labels of in-distribution noisy data. In this way, our proposed method can alleviate the harmful effects caused by out-of-distribution noise and properly exploit the in-distribution noisy samples for training. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is much superior to state-of-the-art noise-robust methods.
Chuanyi Zhang, Yazhou Yao, Xiangbo Shu, Zechao Li, Zhenmin Tang, Qi Wu 0001
ACM Multimedia3
2020 Storyboard relational model for group activity recognition
abstract
This work concerns how to effectively recognize the group activity performed by multiple persons collectively. As known, Storyboards (i.e., medium shot, close shot) jointly describe the whole storyline of a movie in a compact way. Likewise, the actors in small subgroups (similar to Storyboards) of a group activity scene contribute a lot to such group activity and develop more compact relationships among them within subgroups. Inspired by this, we propose a Storyboard Relational Model (SRM) to address the problem of Group Activity Recognition by splitting and reintegrating the group activity based on the small yet compact Storyboards. SRM mainly consists of a Pose-Guided Pruning (PGP) module and a Dual Graph Convolutional Networks (Dual-GCN) module. Specifically, PGP is designed to refine a series of Storyboards from the group activity scene by leveraging the attention ranges of individuals. Dual-GCN models the compact relationships among actors in a Storyboard. Experimental results on two widely-used datasets illustrate the effectiveness of the proposed SRM compared with the state-of-the-art methods.
Boning Li, Xiangbo Shu, Rui Yan 0010
MMAsia2
2020 Cross Fusion for Egocentric Interactive Action Recognition
Haiyu Jiang, Yan Song 0005, Xiangbo Shu
MMM (1)4
2020 A Novel CNN Architecture for Real-Time Point Cloud Recognition in Road Environment
Duyao Fan, Yazhou Yao, Yunfei Cai, Xiangbo Shu, Wankou Yang
PRCV (1)4
2020 Deep multi-person kinship matching and recognition for family photos
Mengyin Wang, Xiangbo Shu, Jiashi Feng, Xun Wang 0007, Jinhui Tang 0001
Pattern Recognit.2
2020 CAN-GAN: Conditioned-attention normalized GAN for face age synthesis
Chenglong Shi, Jiachao Zhang, Yazhou Yao, Yunlian Sun, Huaming Rao, Xiangbo Shu
Pattern Recognit. Lett.6
2020 Deep supervised feature selection for social relationship recognition
Mengyin Wang, Xiaoyu Du 0002, Xiangbo Shu, Xun Wang 0007, Jinhui Tang 0001
Pattern Recognit. Lett.3
2020 Facial Age Synthesis With Label Distribution-Guided Generative Adversarial Network
abstract
The existing research work on facial age synthesis has been mostly focused on long-term aging (e.g., over an age span of 10 years or more). In this paper, we employ generative adversarial networks (GANs) as a tool to investigate age synthesis over different age spans. Compared with long-term aging, short-term age synthesis suffers from the reduced amount of available training data, which can severely hinder the model training. We conduct a series of experiments to validate this. To facilitate short-term age synthesis, we further propose label distribution-guided generative adversarial network (ldGAN), where each sample is associated with an age label distribution (ALD) rather than a single age group. Accordingly, each sample can contribute not only to the learning of its own age group but also to neighbouring groups' learning. This is useful when addressing short-term aging to cope with the reduced amount of training data. In addition, unlike one-hot encoding which treats age groups as independent from one another, ldGAN can well capture the correlation among different age groups, so that smooth aging sequences can be achieved. The ALD model is integrated into GAN with a two-step process. Firstly, instead of the traditional one-hot encoding, ALD is applied as the condition of the generator. Secondly, we add a sequence of label distribution learners on top of several multi-scale discriminators, with the aim of minimizing the label distribution learning loss when optimizing both the generator and discriminators. Both qualitative and quantitative evaluations are conducted to assess ldGAN's ability in dealing with two core issues of face aging, i.e., aging effect generation and identity preservation. The obtained experimental results demonstrate the effectiveness of ldGAN in both learning short-term aging patterns and coping with the lack of training data.
Yunlian Sun, Jinhui Tang 0001, Xiangbo Shu, Zhenan Sun, Massimo Tistarelli
IEEE Trans. Inf. Forensics Secur.3
2020 Bi-Modal Progressive Mask Attention for Fine-Grained Recognition
abstract
Traditional fine-grained image recognition is required to distinguish different subordinate categories (e.g., birds species) based on the visual cues beneath raw images. Due to both small inter-class variations and large intra-class variations, it is desirable to capture the subtle differences between these sub-categories, which is crucial but challenging for fine-grained recognition. Recently, language modality aggregation has been proved as a successful technique to improve visual recognition in the experience. In this paper, we introduce an end-to-end trainable Progressive Mask Attention (PMA) model for fine-grained recognition by leveraging both visual and language modalities. Our Bi-Modal PMA model can not only stage-by-stage capture the most discriminative part in the visual modality by our mask-based fashion, but also explore the out-of-visual-domain knowledge from the language modality in an interactional alignment paradigm. Specifically, at each stage, a self-attention module is proposed to attend to the key patch from images or text descriptions. Besides, a query-relational module is designed to seize the key words/phrases of texts and further bridge the connection between two modalities. Later, the learned representations of bi-modality from multiple stages are aggregated as the final features for recognition. Our Bi-Modal PMA model only needs raw images and raw text descriptions, without requiring bounding boxes/part annotations in images or key word annotations in texts. By conducting comprehensive experiments on fine-grained benchmark datasets, we demonstrate that the proposed method achieves superior performance over the competing baselines, on either vision and language bi-modality or single visual modality.
Kaitao Song, Xiu-Shen Wei, Xiangbo Shu, Renjie Song, Jianfeng Lu 0003
IEEE Trans. Image Process.3
2019 Temporal Action Localization Based on Temporal Evolution Model and Multiple Instance Learning
Minglei Yang 0007, Yan Song 0005, Xiangbo Shu, Jinhui Tang 0001
MMM (2)3
2019 Image annotation refinement via 2P-KNN based group sparse reconstruction
Qian Ji, Liyan Zhang 0001, Xiangbo Shu, Jinhui Tang 0001
Multim. Tools Appl.3
2019 Social Anchor-Unit Graph Regularized Tensor Completion for Large-Scale Image Retagging
abstract
Image retagging aims to improve the tag quality of social images by completing the missing tags, rectifying the noise-corrupted tags, and assigning new high-quality tags. Recent approaches simultaneously explore visual, user and tag information to improve the performance of image retagging by mining the tag-image-user associations. However, such methods will become computationally infeasible with the rapidly increasing number of images, tags and users. It has been proven that the anchor graph can significantly accelerate large-scale graph-based learning by exploring only a small number of anchor points. Inspired by this, we propose a novel Social anchor-Unit GrAph Regularized Tensor Completion (SUGAR-TC) method to efficiently refine the tags of social images, which is insensitive to the scale of data. First, we construct an anchor-unit graph across multiple domains (e.g., image and user domains) rather than traditional anchor graph in a single domain. Second, a tensor completion based on Social anchor-Unit GrAph Regularization (SUGAR) is implemented to refine the tags of the anchor images. Finally, we efficiently assign tags to non-anchor images by leveraging the relationship between the non-anchor units and the anchor units. Experimental results on a real-world social image database well demonstrate the effectiveness and efficiency of SUGAR-TC, outperforming the state-of-the-art methods.
Jinhui Tang 0001, Xiangbo Shu, Zechao Li, Yu-Gang Jiang 0001, Qi Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Deep Ordinal Hashing With Spatial Attention
abstract
Hashing has attracted increasing research attention in recent years due to its high efficiency of computation and storage in image retrieval. Recent works have demonstrated the superiority of simultaneous feature representations and hash functions learning with deep neural networks. However, most existing deep hashing methods directly learn the hash functions by encoding the global semantic information, while ignoring the local spatial information of images. The loss of local spatial structure makes the performance bottleneck of hash functions, therefore limiting its application for accurate similarity retrieval. In this paper, we propose a novel deep ordinal hashing (DOH) method, which learns ordinal representations to generate ranking-based hash codes by leveraging the ranking structure of feature space from both local and global views. In particular, to effectively build the ranking structure, we propose to learn the rank correlation space by exploiting the local spatial information from fully convolutional network and the global semantic information from the convolutional neural network simultaneously. More specifically, an effective spatial attention model is designed to capture the local spatial information by selectively learning well-specified locations closely related to target objects. In such hashing framework, the local spatial and global semantic nature of images is captured in an end-to-end ranking-to-hashing manner. Experimental results conducted on three widely used datasets demonstrate that the proposed DOH method significantly outperforms the state-of-the-art hashing methods.
Lu Jin 0001, Xiangbo Shu, Kai Li 0005, Zechao Li, Guo-Jun Qi, Jinhui Tang 0001
IEEE Trans. Image Process.2
2018 Participation-Contributed Temporal Dynamic Model for Group Activity Recognition
abstract
Group activity recognition, a challenging task that a number of individuals occur in the scene of activity while only a small subset of them participate in, has received increasing attentions. However, most of the previous methods model all the individuals' actions equivalently while ignoring a fact that not all of them are contributed to the discrimination of group activity. That is to say, only a small number of key actors (participants) play important roles in the whole group activity. Inspired by this, we explore a new "One to Key" idea to progressively aggregate temporal dynamics of key actors with different participation degrees over time from each person. Here, we focus on two types of key actors in the whole activity, who steadily move in the whole process (long moving time) or intensely move (but closely related to the group activity) at a significant moment. Based on this, we propose a novel Participation-Contributed Temporal Dynamic Model (PC-TDM) to recognize group activity, which mainly consists of a "One" network and a "One to Key" network. Specifically, "One" network aims at modeling the individual dynamic of each person. "One to Key" network feeds the outputs from the "One" network into a Bidirectional LSTM (Bi-LSTM) according to the order of individual's moving time. Subsequently, each output state of Bi-LSTM weighted by a trainable time-varying attention factor is aggregated by going through LSTM one-by-one. Experimental results on two benchmarks demonstrate that the proposed method improves group activity recognition performance compared to the state-of-the-arts.
Rui Yan 0010, Jinhui Tang 0001, Xiangbo Shu, Zechao Li, Qi Tian 0001
ACM Multimedia3
2018 Global and Local C3D Ensemble System for First Person Interactive Action Recognition
Lingling Fa, Yan Song 0005, Xiangbo Shu
MMM (2)3
2018 A Feature Selection Method for Projection Twin Support Vector Machine
Rui Yan 0010, Qiaolin Ye, Liyan Zhang 0001, Xiangbo Shu
Neural Process. Lett.5
2018 Personalized Age Progression with Bi-Level Aging Dictionary Learning
abstract
Age progression is defined as aesthetically re-rendering the aging face at any future age for an individual face. In this work, we aim to automatically render aging faces in a personalized way. Basically, for each age group, we learn an aging dictionary to reveal its aging characteristics (e.g., wrinkles), where the dictionary bases corresponding to the same index yet from two neighboring aging dictionaries form a particular aging pattern cross these two age groups, and a linear combination of all these patterns expresses a particular personalized aging process. Moreover, two factors are taken into consideration in the dictionary learning process. First, beyond the aging dictionaries, each person may have extra personalized facial characteristics, e.g., mole, which are invariant in the aging process. Second, it is challenging or even impossible to collect faces of all age groups for a particular person, yet much easier and more practical to get face pairs from neighboring age groups. To this end, we propose a novel Bi-level Dictionary Learning based Personalized Age Progression (BDL-PAP) method. Here, bi-level dictionary learning is formulated to learn the aging dictionaries based on face pairs from neighboring age groups. Extensive experiments well demonstrate the advantages of the proposed BDL-PAP over other state-of-the-arts in term of personalized age progression, as well as the performance gain for cross-age face verification by synthesizing aging faces.
Xiangbo Shu, Jinhui Tang 0001, Zechao Li, Hanjiang Lai, Liyan Zhang 0001, Shuicheng Yan
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Image Classification With Tailored Fine-Grained Dictionaries
abstract
In this paper, we propose a novel fine-grained dictionary learning method for image classification. To learn a high-quality discriminative dictionary, three types of multispecific subdictionaries, i.e., class-specific dictionaries (CSDs), universal dictionary (UD), and family-specific dictionaries (FSDs), are simultaneously uncovered. Here, CSDs and UD, respectively, model the patterns for each class and the patterns irrespective of any class. FSDs can help reveal the shared patterns between multiple image classes, by filling the gap between the patterns in CSDs and UD. The dependence among image classes is revealed by the shared FSDs, and a common FSD can be assigned to several classes to represent their residual. Finally, the most discriminative FSD for each class is identified by minimizing the sparse reconstruction error. Extensive experiments are conducted on different widely used data sets for image classification. The results demonstrate the superior performance of the proposed method over some state-of-the-art methods.
Xiangbo Shu, Jinhui Tang 0001, Guo-Jun Qi, Zechao Li, Yu-Gang Jiang 0001, Shuicheng Yan
IEEE Trans. Circuits Syst. Video Technol.1
2017 Face Aging with Contextual Generative Adversarial Nets
abstract
Face aging, which renders aging faces for an input face, has attracted extensive attention in the multimedia research. Recently, several conditional Generative Adversarial Nets (GANs) based methods have achieved great success. They can generate images fitting the real face distributions conditioned on each individual age group. However, these methods fail to capture the transition patterns, e.g., the gradual shape and texture changes between adjacent age groups. In this paper, we propose a novel Contextual Generative Adversarial Nets (C-GANs) to specifically take it into consideration. The C-GANs consists of a conditional transformation network and two discriminative networks. The conditional transformation network imitates the aging procedure with several specially designed residual blocks. The age discriminative network guides the synthesized face to fit the real conditional distribution. The transition pattern discriminative network is novel, aiming to distinguish the real transition patterns with the fake ones. It serves as an extra regularization term for the conditional transformation network, ensuring the generated image pairs to fit the corresponding real transition pattern distribution. Experimental results demonstrate the proposed framework produces appealing results by comparing with the state-of-the-art and ground truth. We also observe performance gain for cross-age face verification.
Si Liu 0001, Yao Sun 0004, Defa Zhu, Renda Bao, Wei Wang 0108, Xiangbo Shu, Shuicheng Yan
ACM Multimedia6
2017 Computational face reader based on facial attribute estimation
Xiangbo Shu, Yunfei Cai, Liyan Zhang 0001, Jinhui Tang 0001
Neurocomputing1
2017 Tri-Clustered Tensor Completion for Social-Aware Image Tag Refinement
abstract
Social image tag refinement, which aims to improve tag quality by automatically completing the missing tags and rectifying the noise-corrupted ones, is an essential component for social image search. Conventional approaches mainly focus on exploring the visual and tag information, without considering the user information, which often reveals important hints on the (in)correct tags of social images. Towards this end, we propose a novel tri-clustered tensor completion framework to collaboratively explore these three kinds of information to improve the performance of social image tag refinement. Specifically, the inter-relations among users, images and tags are modeled by a tensor, and the intra-relations between users, images and tags are explored by three regularizations respectively. To address the challenges of the super-sparse and large-scale tensor factorization that demands expensive computing and memory cost, we propose a novel tri-clustering method to divide the tensor into a certain number of sub-tensors by simultaneously clustering users, images and tags into a bunch of tri-clusters. And then we investigate two strategies to complete these sub-tensors by considering (in)dependence between the sub-tensors. Experimental results on a real-world social image database demonstrate the superiority of the proposed method compared with the state-of-the-art methods.
Jinhui Tang 0001, Xiangbo Shu, Guo-Jun Qi, Zechao Li, Meng Wang 0001, Shuicheng Yan, Ramesh Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2016 Recurrent Face Aging
abstract
Modeling the aging process of human face is important for cross-age face verification and recognition. In this paper, we introduce a recurrent face aging (RFA) framework based on a recurrent neural network which can identify the ages of people from 0 to 80. Due to the lack of labeled face data of the same person captured in a long range of ages, traditional face aging models usually split the ages into discrete groups and learn a one-step face feature transformation for each pair of adjacent age groups. However, those methods neglect the in-between evolving states between the adjacent age groups and the synthesized faces often suffer from severe ghosting artifacts. Since human face aging is a smooth progression, it is more appropriate to age the face by going through smooth transition states. In this way, the ghosting artifacts can be effectively eliminated and the intermediate aged faces between two discrete age groups can also be obtained. Towards this target, we employ a twolayer gated recurrent unit as the basic recurrent module whose bottom layer encodes a young face to a latent representation and the top layer decodes the representation to a corresponding older face. The experimental results demonstrate our proposed RFA provides better aging faces over other state-of-the-art age progression methods.
Wei Wang 0108, Zhen Cui 0001, Yan Yan 0002, Jiashi Feng, Shuicheng Yan, Xiangbo Shu, Nicu Sebe
CVPR6
2016 Computational Face Reader
Xiangbo Shu, Liyan Zhang 0001, Jinhui Tang 0001, Guosen Xie, Shuicheng Yan
MMM (1)1
2016 Age progression: Current technologies and applications
Xiangbo Shu, Guosen Xie, Zechao Li, Jinhui Tang 0001
Neurocomputing1
2016 Kinship-Guided Age Progression
Xiangbo Shu, Jinhui Tang 0001, Hanjiang Lai, Zhiheng Niu, Shuicheng Yan
Pattern Recognit.1
2016 Instance-Aware Hashing for Multi-Label Image Retrieval
abstract
Similarity-preserving hashing is a commonly used method for nearest neighbor search in large-scale image retrieval. For image retrieval, deep-network-based hashing methods are appealing, since they can simultaneously learn effective image representations and compact hash codes. This paper focuses on deep-network-based hashing for multi-label images, each of which may contain objects of multiple categories. In most existing hashing methods, each image is represented by one piece of hash code, which is referred to as semantic hashing. This setting may be suboptimal for multi-label image retrieval. To solve this problem, we propose a deep architecture that learns instance-aware image representations for multi-label image data, which are organized in multiple groups, with each group containing the features for one category. The instance-aware representations not only bring advantages to semantic hashing but also can be used in category-aware hashing, in which an image is represented by multiple pieces of hash codes and each piece of code corresponds to a category. Extensive evaluations conducted on several benchmark data sets demonstrate that for both the semantic hashing and the category-aware hashing, the proposed method shows substantial improvement over the state-of-the-art supervised and unsupervised hashing methods.
Hanjiang Lai, Pan Yan, Xiangbo Shu, Yunchao Wei, Shuicheng Yan
IEEE Trans. Image Process.3
2016 Generalized Deep Transfer Networks for Knowledge Propagation in Heterogeneous Domains
abstract
In recent years, deep neural networks have been successfully applied to model visual concepts and have achieved competitive performance on many tasks. Despite their impressive performance, traditional deep networks are subjected to the decayed performance under the condition of lacking sufficient training data. This problem becomes extremely severe for deep networks trained on a very small dataset, making them overfitting by capturing nonessential or noisy information in the training set. Toward this end, we propose a novel generalized deep transfer networks (DTNs), capable of transferring label information across heterogeneous domains, textual domain to visual domain. The proposed framework has the ability to adequately mitigate the problem of insufficient training images by bringing in rich labels from the textual domain. Specifically, to share the labels between two domains, we build parameter- and representation-shared layers. They are able to generate domain-specific and shared interdomain features, making this architecture flexible and powerful in capturing complex information from different domains jointly. To evaluate the proposed method, we release a new dataset extended from NUS-WIDE at http://imag.njust.edu.cn/NUS-WIDE-128.html. Experimental results on this dataset show the superior performance of the proposed DTNs compared to existing state-of-the-art methods.
Jinhui Tang 0001, Xiangbo Shu, Zechao Li, Guo-Jun Qi, Jingdong Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2015 Personalized Age Progression with Aging Dictionary
abstract
In this paper, we aim to automatically render aging faces in a personalized way. Basically, a set of age-group specific dictionaries are learned, where the dictionary bases corresponding to the same index yet from different dictionaries form a particular aging process pattern cross different age groups, and a linear combination of these patterns expresses a particular personalized aging process. Moreover, two factors are taken into consideration in the dictionary learning process. First, beyond the aging dictionaries, each subject may have extra personalized facial characteristics, e.g. mole, which are invariant in the aging process. Second, it is challenging or even impossible to collect faces of all age groups for a particular subject, yet much easier and more practical to get face pairs from neighboring age groups. Thus a personality-aware coupled reconstruction loss is utilized to learn the dictionaries based on face pairs from neighboring age groups. Extensive experiments well demonstrate the advantages of our proposed solution over other state-of-the-arts in term of personalized aging progression, as well as the performance gain for cross-age face verification by synthesizing aging faces.
Xiangbo Shu, Jinhui Tang 0001, Hanjiang Lai, Luoqi Liu, Shuicheng Yan
ICCV1
2015 Task-Driven Feature Pooling for Image Classification
abstract
Feature pooling is an important strategy to achieve high performance in image classification. However, most pooling methods are unsupervised and heuristic. In this paper, we propose a novel task-driven pooling (TDP) model to directly learn the pooled representation from data in a discriminative manner. Different from the traditional methods (e.g., average and max pooling), TDP is an implicit pooling method which elegantly integrates the learning of representations into the given classification task. The optimization of TDP can equalize the similarities between the descriptors and the learned representation, and maximize the classification accuracy. TDP can be combined with the traditional BoW models (coding vectors) or the recent state-of-the-art CNN models (feature maps) to achieve a much better pooled representation. Furthermore, a self-training mechanism is used to generate the TDP representation for a new test image. A multi-task extension of TDP is also proposed to further improve the performance. Experiments on three databases (Flower-17, Indoor-67 and Caltech-101) well validate the effectiveness of our models.
Guosen Xie, Xu-Yao Zhang, Xiangbo Shu, Shuicheng Yan, Cheng-Lin Liu 0001
ICCV3
2015 Partially Common-Semantic Pursuit for RGB-D Object Recognition
abstract
For the RGB-D object recognition task, the robust and rich representations can boost the performance. Most works employ feature learning approaches to learn specific representation for the RGB and depth modalities independently, while some directly learn common property. Different from them, this paper proposes a novel supervised feature learning method for RGB-D object recognition, named Partially Common-Semantic Learning (PCSL), which jointly captures the complementary and consistency semantic information from RGB and depth modalities. The complementary information is revealed by the individual modality, while the consistency is exploited by both modalities simultaneously. In PCSL, Reconstruction Independent Component Analysis (RICA) is extended to integrate the supervised information and learn both of the complementary and partially shared common semantic information. The proposed approach is evaluated on two public RGB-D datasets and achieves better performance than several state-of-the-art methods.
Lu Jin 0001, Zechao Li, Xiangbo Shu, Shenghua Gao, Jinhui Tang 0001
ACM Multimedia3
2015 Deep Face Beautification
abstract
The beautification of human photos usually requires professional editing softwares, which are difficult for most users. In this technical demonstration, we propose a deep face beautification framework, which is able to automatically modify the geometrical structure of a face so as to boost the attractiveness. A learning based approach is adopted to capture the underlying relations between the facial shape and the attractiveness via training the Deep Beauty Predictor (DBP). Relying on the pre-trained DBP, we construct the BeAuty SHaper (BASH) to infer the "flows" of landmarks towards the maximal aesthetic level. BASH modifies the facial landmarks with the direct guidance of the beauty score estimated by DBP.
Jianshu Li, Luoqi Liu, Xiangbo Shu, Shuicheng Yan
ACM Multimedia4
2015 Weakly-Shared Deep Transfer Networks for Heterogeneous-Domain Knowledge Propagation
abstract
In recent years, deep networks have been successfully applied to model image concepts and achieved competitive performance on many data sets. In spite of impressive performance, the conventional deep networks can be subjected to the decayed performance if we have insufficient training examples. This problem becomes extremely severe for deep networks with powerful representation structure, making them prone to over fitting by capturing nonessential or noisy information in a small data set. In this paper, to address this challenge, we will develop a novel deep network structure, capable of transferring labeling information across heterogeneous domains, especially from text domain to image domain. This weakly-shared Deep Transfer Networks (DTNs) can adequately mitigate the problem of insufficient image training data by bringing in rich labels from the text domain.
Xiangbo Shu, Guo-Jun Qi, Jinhui Tang 0001, Jingdong Wang 0001
ACM Multimedia1
2015 What Shall I Look Like after N Years?
abstract
"What shall I look like after N years?" In this paper, we present an Auto Age Progression system, which automatically renders a series of aging faces in the future age ranges and generates an aging sequence (aging video) covering the entire life for an individual input. In the offline stage, a set of age-range specific dictionaries are learned from the constructed database, where the dictionary bases corresponding to the same index yet from different dictionaries form a particular aging process pattern across different age groups, and a linear combination of these patterns expresses a particular personalized aging process. In the online stage, for an input face of an individual, our system renders the aging faces corresponding to different age ranges through the aging dictionaries, and then generates an age progression by the presented face morphing technology.
Xiangbo Shu, Jinhui Tang 0001, Luoqi Liu, Zhiheng Niu, Shuicheng Yan
ACM Multimedia1
2015 Deep kinship verification
abstract
To improve the performance of kinship verification, we propose a novel deep kinship verification (DKV) model by integrating excellent deep learning architecture into metric learning. Unlike most existing shallow models based on metric learning for kinship verification, we employ a deep learning model followed by a metric learning formulation to select nonlinear features, which can find the appropriate project space to ensure the margin of negative sample pairs (i.e. parent and child without kinship relation) as large as possible and the margin of positive sample pairs (i.e. parent and child with kinship relation) as small as possible. Experimental results show that our method achieves satisfactory performance on two widely-used benchmarks, i.e. KFW-I and KFW-II.
Mengyin Wang, Zechao Li, Xiangbo Shu, Jingdong Wang 0001, Jinhui Tang 0001
MMSP3